← Back to Index
Daily Research Digest

arXiv Papers

2026-10-09
611
Papers
9
Categories
146
Translated
收藏清单 0
精选 · Favorites
146
cs.AI / 1 / 2610.10786
Plan-and-Patch: Diffusion Language Models for Agentic Planning
Plan-and-Patch:面向智能体规划的扩散语言模型
Syamantak Kumar, Jiang Guo, Hassan Hamad, Hideo Kobayashi, Yi Xiang, Yezhou Yang, Yanjun Qi, Daniele Bonadiman, Jiarong Jiang
cs.AI · cs.CL · cs.LG
diffusion
扩散模型相关
Abstract
Planning is increasingly important for long-horizon agents, where successful execution requires coordinating subgoals, tool use, and intermediate outcomes over many steps. Yet assumptions made during planning may be invalidated by the environment, tools may return unexpected results, or actions may fail. Effective agents must therefore not only generate plans, but also revise them. Such revisions often affect only part of a plan, leaving the preceding and subsequent structure intact. Rather than regenerate the entire plan and risk unnecessary changes, repair can regenerate the affected region conditioned on the preserved prefix and suffix. We introduce Plan-and-Patch, a plan-and-act framework in which a diffusion language model (dLLM) generates a structured, program-like plan through parallel unmasking and repairs it by filling in selected regions while keeping the surrounding steps fixed. We compare DreamReasoner-8B and Qwen3-8B as diffusion and autoregressive (AR) planners. On Natural Plan without task-specific training, diffusion (53.7%) achieves nearly twice the plan repair success rate of AR (27.0%). After task-specific training on agentic benchmarks, ALFWorld and TextCraft, the planners achieve similar observed success in plan generation, while diffusion reduces mean plan-generation latency by 39-46% relative to AR. Our results show that Plan-and-Patch provides a framework for faster plan generation and effective plan repair in long-horizon agents.
Chinese Translation
规划对于长时程智能体而言日益重要,因为成功的执行需要在许多步骤中协调子目标、工具使用与中间结果。然而,规划过程中所做的假设可能会被环境推翻,工具可能返回意外的结果,或者动作可能失败。因此,有效的智能体不仅必须生成规划,还必须对其进行修订。这类修订通常只影响规划的一部分,使之前和之后的结构保持完整。与其重新生成整个规划并冒着引入不必要更改的风险,修复可以以保留的前缀和后缀为条件,重新生成受影响的区域。我们提出 Plan-and-Patch,一个规划—行动框架,其中扩散语言模型(dLLM)通过并行去掩码生成结构化的、类程序式规划,并通过填充选定区域同时保持周围步骤固定来对其进行修复。我们将 DreamReasoner-8B 与 Qwen3-8B 分别作为扩散规划器和自回归(AR)规划器进行比较。在未进行任务特定训练的 Natural Plan 上,扩散模型(53.7%)取得的规划修复成功率接近自回归模型(27.0%)的两倍。在智能体基准 ALFWorld 和 TextCraft 上进行任务特定训练后,两种规划器在规划生成上取得相近的观测成功率,而扩散模型相对于自回归模型将平均规划生成延迟降低了 39-46%。我们的结果表明,Plan-and-Patch 为长时程智能体中的更快规划生成与有效规划修复提供了一个框架。
cs.AI / 2 / 2610.10942
StoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator Agents
StoreBench:一个用于评估和训练自主运营智能体的实时商务环境
Daksh Raghuvanshi, Ved Vedere, Yifan Wang
cs.AI · cs.CL · cs.LG
large language model
大语言模型相关
Abstract
Reinforcement learning environments are now a primary lever for improving large language model (LLM) capabilities in post-training, yet most agentic benchmarks remain static: the world moves only when the agent acts, the reward is a terminal verdict, and the pass bar is set arbitrarily. We introduce StoreBench, a live-commerce environment in which an agent runs a mid-size online apparel store on a production-grade commerce backend, testing long-horizon planning and economic judgment under uncertainty. Customers order around the clock, suppliers reprice and fail, and market shocks arrive with partial or no warning. The agent acts through the same 29 merchant tools a human operator would use, under a windowed operation budget that makes simulated time a function of actions taken, so model latency cannot influence simulated time. Pass thresholds are calibrated against scripted anchor policies, the reward is hardened against a catalogue of reward hacks, and every episode replays identically given a sequence of actions. We evaluate seven frontier LLMs on 11 scenarios of 30 to 45 days and a full simulated year, over three world seeds at matched reasoning effort. No model matches the scripted smart-triage policy on average: the best, DeepSeek-V4-Pro, passes 49% of task-seed cells against the heuristic's 97%. Human experts working through the same tools and budgets outscore every model (mean composite 0.708 vs. 0.700). Over a full simulated year under the Claude Code harness, most models show dramatic performance improvement. In a GRPO post-training run, Qwen3.5-27B trained on only five disjoint tasks raises its mean composite on the held-out evaluation tasks from 0.136 to 0.373. We release five example training-split tasks, ten sample trajectories, and the scoring and verification tooling; the full environment and evaluation suite are withheld to keep the benchmark uncontaminated.
Chinese Translation
强化学习环境现在是提升大语言模型(LLM)在后训练阶段能力的主要杠杆,但大多数智能体基准仍然是静态的:世界只有在智能体行动时才会变化,奖励是终止式判定,通过门槛则被任意设定。我们提出 StoreBench,一个实时商务环境,其中智能体在一个生产级商业后端上经营一家中型在线服装店,测试不确定性下的长时程规划和经济判断。顾客全天候下单,供应商调整价格并可能失败,市场冲击在仅有部分预警或毫无预警的情况下来临。智能体通过人类运营者会使用的同样的 29 种商家工具采取行动,并在一种窗口化运营预算下运行,该预算使模拟时间成为所采取行动的函数,因此模型延迟不能影响模拟时间。通过阈值依据脚本化锚定策略进行校准,奖励针对一整套奖励破解手段进行了加固,并且在给定动作序列的情况下,每个回合都能以完全相同的方式重放。我们在 11 个持续 30 至 45 天的场景以及一个完整模拟年上,跨三个世界种子并在匹配的推理努力下,评估了七个前沿 LLM。平均而言,没有模型能匹敌脚本化的智能分诊策略:表现最好的 DeepSeek-V4-Pro 通过了 49% 的任务-种子单元,而启发式策略通过了 97%。通过相同工具和预算工作的人类专家得分超过每个模型(平均综合得分 0.708 对 0.700)。在 Claude Code 测试框架下经过一个完整模拟年,大多数模型表现出显著的性能提升。在一次 GRPO 后训练运行中,仅在五个不相交任务上训练的 Qwen3.5-27B 将其在留出评估任务上的平均综合得分从 0.136 提高到 0.373。我们发布五个示例训练划分任务、十条示例轨迹以及评分和验证工具;完整环境和评估套件不予公开,以保持该基准不受污染。
cs.AI / 3 / 2610.11087
Beyond Imitation: A Framework and Benchmark for LLM-Assisted Peer Review
超越模仿:面向LLM辅助同行评审的框架与基准
Rachel S. Y. Teo, Yutaro Yamada, Shashank Kotyan, Yuki Imajuku, Tarin Clanuwat
cs.AI
large language model
大语言模型相关
Abstract
The rapid growth of scientific publishing has strained peer review, particularly in machine learning, raising concerns about declining review quality and increasing reviewer workload. Large language models (LLMs) have been proposed as automated review assistants, yet their evaluation has focused largely on imitating human-written reviews rather than supporting the core functions of peer review. Here, we introduce a verification-centric perspective on LLM-assisted peer review, emphasizing error detection as a critical and resource-intensive task. We present a scalable benchmark that evaluates review systems' ability to identify logical contradictions, constructed through synthetic insertion of errors into conference papers, yielding unambiguous evaluation targets and enabling systematic comparison. We further propose a Multi-Layered Review (MLR) framework that prioritizes detailed manuscript comprehension before review generation, aligning more closely with human reviewing practices while improving token efficiency. Across evaluations, our approach demonstrates strong alignment with human review scores, achieves high error detection performance, and provides complementary perspectives on reviewer focus. These improvements can be attributed to both the choice of the underlying LLM and the design of our system. At the same time, we corroborate persistent vulnerabilities to adversarial manipulation, underscoring the need for robustness in automated review systems. Our findings highlight the importance of rigorous, error-focused evaluation to guide responsible deployment of LLM-based tools in peer review and other critical scientific workflows.
Chinese Translation
科学出版的快速增长使同行评审承受压力,尤其是在机器学习领域,引发了人们对评审质量下降和审稿人工作负担增加的担忧。大语言模型(LLMs)已被提出用作自动化评审助手,然而对其的评估在很大程度上聚焦于模仿人类撰写的评审意见,而非支持同行评审的核心功能。在此,我们引入一种以验证为中心的视角来看待LLM辅助的同行评审,强调错误检测是一项关键且资源密集型的任务。我们提出一个可扩展的基准,用于评估评审系统识别逻辑矛盾的能力;该基准通过向会议论文中合成插入错误而构建,从而产生明确无歧义的评估目标,并使系统性的比较成为可能。我们进一步提出一个多层次评审(MLR)框架,该框架在生成评审意见之前优先进行详尽的手稿理解,从而更贴近人类评审实践,同时提升token效率。在各项评估中,我们的方法展现出与人类评审分数的高度一致性,实现了较高的错误检测性能,并就审稿人关注点提供了互补的视角。这些改进可归因于底层LLM的选择以及我们系统的设计两方面。与此同时,我们证实了自动评审系统对对抗性操纵存在持续性的脆弱性,凸显了此类系统对鲁棒性的需求。我们的发现凸显了严谨、以错误为中心的评价的重要性,以指导基于LLM的工具在同行评审及其他关键科学工作流中的负责任部署。
cs.AI / 4 / 2610.11226
When Lower Reconstruction Loss Hurts: Distributionally Robust Refinement for Low-Bit LLM Quantization
当更低的重建损失反而有害:面向低比特 LLM 量化的分布鲁棒精化
Yanlong Zhao, Xiaoyuan Cheng, Huihang Liu, Baihua He, Xinyu Zhang, Harrison Bo Hua Zhu, Wenlong Chen, Li Zeng, Zhuo Sun
cs.AI · cs.LG · stat.ML
large language model
大语言模型相关
Abstract
Weight-only post-training quantization (PTQ) relies heavily on reconstruction loss minimization to preserve model quality at low precision. We show that the weights favored by minimizing this loss need not yield better model performance on new tasks. In fact, we find that lower reconstruction loss can even degrade model performance on the same calibration data. Our analysis further shows that weights with lower reconstruction loss on calibration data can have higher loss than other weights when the distribution of input activations changes. Motivated by these observations and our analysis, we propose Distributionally Robust Quantization (DRQ), a post-hoc refinement process that minimizes worst-case reconstruction loss over a constrained set of input activation distributions. DRQ refines the integer codes representing quantized weights within the existing quantization grid, keeping quantization parameters and inference operators unchanged. Extensive experiments show that DRQ improves models quantized by six representative PTQ methods, including AWQ, GPTQ, and ParoQuant, and delivers gains across both dense and mixture-of-experts large language models. These results establish DRQ as a general post-hoc refinement framework for weight-only PTQ, achieving better downstream performance without adding inference overhead.
Chinese Translation
仅权重的训练后量化(PTQ)在很大程度上依赖重建损失最小化,以在低精度下保持模型质量。我们表明,通过最小化该损失而受到偏好的权重,未必能在新任务上带来更好的模型性能。事实上,我们发现更低的重建损失甚至可能在同一校准数据上降低模型性能。我们的分析进一步表明,当输入激活的分布发生变化时,在校准数据上具有更低重建损失的权重,其损失可能高于其他权重。受这些观察和我们分析的启发,我们提出分布鲁棒量化(DRQ),这是一种事后精化过程,它在一个受限的输入激活分布集合上最小化最坏情况下的重建损失。DRQ 在既有的量化网格内精化表示量化权重的整数编码,同时保持量化参数和推理算子不变。大量实验表明,DRQ 能够改进由六种代表性 PTQ 方法(包括 AWQ、GPTQ 和 ParoQuant)量化得到的模型,并在稠密大语言模型和专家混合大语言模型上均带来收益。这些结果确立了 DRQ 作为仅权重 PTQ 的通用事后精化框架的地位,在不增加推理开销的情况下实现更好的下游性能。
cs.AI / 5 / 2610.11253
LLM-IDEA: Identifiability-Driven Experimental Agent for Autonomous Discovery of Mechanistic World Models
LLM-IDEA:面向机制性世界模型自主发现的可辨识性驱动实验智能体
Surya Shetty, Ulisses Braga-Neto
cs.AI
large language model
大语言模型相关
Abstract
Large language model agents are being increasingly deployed as autonomous scientists, designing experiments and inferring mechanistic world models with minimal human oversight. Yet identifiability is often overlooked: when a plateau is reached, the agent needs to know whether it is not yet capable enough or the model simply is not identifiable from the data, in which case no amount of further experimentation of the same kind can help. We propose the Identifiability-Driven Experimental Agent (LLM-IDEA) for closed-loop discovery with an identifiability engine that returns a three-way plateau verdict: capability limit, resolvable within the design class, or certified exhausted. On ODEBench, 60 of the 62 systems with free constants are identifiable at round 0; the RC circuit is certified exhausted for every experiment that protocol can run, and a harvesting model is resolvable by one added initial condition. The identifiability engine reproduces known verdicts on Lotka-Volterra, Van der Pol, Lorenz, and a pharmacokinetic model, where it recommends the intravenous arm pharmacologists use, and it ranks the depth scorer of our own benchmark last among four observation designs. On the DiscoverPhysics benchmark, it finds two public worlds whose explanation rubric rewards a distinction no legal experiment can make, and every model there with accurate trajectories failed the explanation grade (15 of 15, against 5 of 9 in identifiable worlds, p = 0.012). On the Alien Universe, a two-body testbed we propose in which a force law switches between a provably non-identifiable and an identifiable protocol, LLM-IDEA on the identifiable protocol reaches discovery depth at least three on 8/8 seeds versus 1/8 without it. An autonomous discovery agent can thus compute, rather than guess, whether a plateau calls for more search, a better experiment of the same kind, or a different kind of experiment.
Chinese Translation
大语言模型智能体正越来越多地被部署为自主科学家,以最少的人类监督设计实验并推断机制性世界模型。然而,可辨识性常常被忽视:当到达平台期时,智能体需要知道,究竟是其能力还不够强,还是模型根本无法从数据中被辨识出来;在后一种情况下,再多进行同类实验也无济于事。我们提出可辨识性驱动的实验智能体(LLM-IDEA),用于闭环发现,其配备一个可辨识性引擎,返回三选一的平台期判定:能力极限、可在设计类内解决,或已被证明耗尽。在 ODEBench 上,62 个含自由常数的系统中,有 60 个在第 0 轮即可辨识;RC 电路对于该协议能够运行的每一个实验都被证明耗尽,而一个收获模型通过增加一个初始条件即可解决。该可辨识性引擎在 Lotka-Volterra、Van der Pol、Lorenz 以及一个药代动力学模型上重现了已知判定,在其中它推荐了药理学家使用的静脉给药组,并且它在我们自有基准的深度评分器中,将其在四种观测设计里排为最后。在 DiscoverPhysics 基准上,它发现了两个公开世界,其解释评分标准奖励一种任何合法实验都无法做出的区分,并且在那里每一个具有准确轨迹的模型都未通过解释评级(15/15,相比之下,在可辨识世界中为 5/9,p = 0.012)。在 Alien Universe 上——我们提出的一个二体测试平台,其中力定律在一个可证明不可辨识的协议和一个可辨识协议之间切换——LLM-IDEA 在可辨识协议上于 8/8 个随机种子中达到至少为三的发现深度,相比之下,没有它时为 1/8。因此,自主发现智能体能够计算而非猜测:一个平台期需要的是更多搜索、一种更好的同类实验,还是一种不同类型的实验。
cs.AI / 6 / 2610.11312
MedBenchAgent: Towards Systematic Automation of Medical VLM Benchmark Construction
MedBenchAgent:迈向医学 VLM 基准构建的系统化自动化
Yulin Fu, Junren Wang, Guangjing Yang, Zhangyuan Yu, Wanran Sun, Jiabao Zhou, Jin Yin, Qicheng Lao
cs.AI
large language model
大语言模型相关
Abstract
Large-scale construction of medical vision-language model (VLM) benchmarks is increasingly feasible with richly annotated imaging datasets and large language models (LLMs), yet existing automation largely focuses on generating evaluation items within predefined benchmark specifications. We study the broader problem of automatically deriving the specification itself: what to evaluate, which annotations support each task, and how to translate this evidence into reliable evaluation items. We formulate benchmark construction as constrained compilation, in which the benchmark specification is progressively derived from evaluation requirements, heterogeneous annotations, and medical knowledge. Based on this formulation, we introduce MedBenchAgent, a multi-agent framework with a Benchmark Intermediate Representation (BIR) that encodes task definitions, evidence mappings, evaluation protocols, and item specifications across construction stages. MedBenchAgent separates planning, which derives and verifies the specification, from instantiation, which constructs and audits items under the locked specification. MedBenchAgent achieves a Task-Space F1 of 90.9%, outperforming direct task induction (79.2-80.0%) and prior-guided induction (85.1%); 994 of 1,000 sampled items from correctly identified tasks pass human audit. We further demonstrate portability to a specialized medical domain and evaluate twelve VLMs, revealing task- and setting-specific variation obscured by aggregate scores. These results establish constrained compilation as a scalable and auditable framework for medical VLM benchmark construction beyond question generation.
Chinese Translation
在标注丰富的影像数据集和大语言模型(LLM)的支持下,大规模构建医学视觉语言模型(VLM)基准正变得日益可行,然而现有自动化主要集中于在预定义的基准规范内生成评估条目。我们研究一个更广泛的问题:自动推导规范本身,即评估什么、哪些标注支持每个任务,以及如何将这些证据转化为可靠的评估条目。我们将基准构建形式化为受约束的编译,其中基准规范从评估需求、异构标注和医学知识中逐步推导出来。基于这一形式化,我们提出 MedBenchAgent,一个多智能体框架,带有基准中间表示(BIR),其对跨构建阶段的任务定义、证据映射、评估协议和条目规范进行编码。MedBenchAgent 将规划与实例化分离:规划推导并验证规范,而实例化则在锁定规范下构建并审核条目。MedBenchAgent 取得了 90.9% 的任务空间 F1,优于直接任务归纳(79.2-80.0%)和先验引导归纳(85.1%);来自正确识别任务的 1,000 个抽样条目中有 994 个通过了人工审核。我们进一步证明其可移植到某一专门医学领域,并评估了十二个 VLM,揭示了被聚合分数所掩盖的任务特定和设置特定的差异。这些结果确立了受约束的编译是一种可扩展且可审核的框架,可用于超越问题生成的医学 VLM 基准构建。
cs.AI / 7 / 2610.11317
DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition
DivMoE:通过跨域专家组合实现的细粒度 MoE 升级回收
Yuxuan Lou, Kai Yang, Geng Zhang, Yong Liu, Yang You
cs.AI · cs.LG
large language model
大语言模型相关
Abstract
Mixture-of-Experts (MoE) architectures have become essential for scaling large language models, with recent work demonstrating the benefits of fine-grained expert designs. Training such models from scratch is expensive, and sparse upcycling from pre-trained dense models is an attractive alternative. However, we identify a structural pathology of fine-grained upcycling: when fine-grained experts are derived from a single source model, naive routing collapses and downstream accuracy drops to near-random (e.g., on Qwen3-1.7B, Drop-Upcycling-fine-grained reaches only 23.2% average accuracy across 15 benchmarks, essentially matching from-scratch training at 22.2%, while the same method's coarse-grained variant reaches 50.2%). We propose DivMoE, the first framework achieving fine-grained MoE Upcycling with structurally-balanced routing. DivMoE introduces domain-specialized fine-grained expert initialization, deriving experts from dense models that have undergone domain-adaptive continual pre-training, and diversity-constrained routing, a hard structural constraint guaranteeing that each token activates experts from distinct domain groups. Across two base models and 15 benchmarks, DivMoE consistently outperforms six upcycling baselines (55.6% vs. 51.6% for the strongest baseline on Qwen3-1.7B) and strictly improves over the dense base model on every benchmark after Stage 2 continual pre-training -- closing the regression gap that has plagued prior fine-grained upcycling. After supervised fine-tuning on a public reasoning mixture, our 12B-parameter DivMoE model matches Moonlight-MoE (16B) at 64.5% average accuracy while outperforming a controlled NVIDIA-Upcycling baseline by 6.1 percentage points.
Chinese Translation
专家混合(MoE)架构已成为扩展大语言模型的关键,近期研究证明了细粒度专家设计的优势。从头训练此类模型代价高昂,而从预训练的稠密模型进行稀疏升级回收是一种颇具吸引力的替代方案。然而,我们发现了细粒度升级回收的一种结构性病理:当细粒度专家源自单一源模型时,朴素路由会崩溃,下游准确率降至接近随机水平(例如,在 Qwen3-1.7B 上,Drop-Upcycling 的细粒度版本在 15 个基准上的平均准确率仅为 23.2%,基本与从头训练的 22.2% 相当,而同一方法的粗粒度变体则达到 50.2%)。我们提出 DivMoE,这是首个实现具有结构平衡路由的细粒度 MoE 升级回收的框架。DivMoE 引入了领域特化的细粒度专家初始化,即从经过领域自适应持续预训练的稠密模型中派生专家,以及多样性约束路由——一种硬性结构约束,保证每个 token 激活来自不同领域组的专家。在两个基础模型和 15 个基准上,DivMoE 持续优于六个升级回收基线(在 Qwen3-1.7B 上为 55.6%,而最强基线为 51.6%),并且在第二阶段持续预训练后,在每个基准上都严格优于稠密基础模型——弥合了此前困扰细粒度升级回收的性能回退差距。在公开推理混合数据上进行监督微调后,我们的 12B 参数 DivMoE 模型以 64.5% 的平均准确率与 Moonlight-MoE(16B)持平,同时以 6.1 个百分点的优势超过受控的 NVIDIA-Upcycling 基线。
cs.AI / 8 / 2610.11363
UniData: Universal Multimodal Instruction Generation Pipeline
UniData:通用多模态指令生成流水线
Jiaqi Tang, Yi-Feng Wu, Yuting Zhang, Hao Lu, Bowen Fu, Qing-Guo Chen, Xiaogang Xu, Yuwei Hu, Shiyin Lu, Wei Wei, Lei Zhang, Zhao Xu, Weihua Luo, Qifeng Chen, Ying-Cong Chen
cs.AI · cs.CL
large language model
大语言模型相关
Abstract
Multimodal Large Language Models (MLLMs) are increasingly being applied in a wider range of real-world scenarios. However, due to the substantial labor cost, creating high-quality multimodal instruction datasets for MLLMs remains a significant challenge. Although some methods propose to generate instruction data, they often face limitations in modality support and struggle with generating multi-round instructions. To address these problems, we introduce UniData, a universal instruction generation pipeline, to transform simple user requirements into multi-round, multimodal instructions. Specifically, UniData first expands user requirements into multiple diverse events. Using these events, UniData then integrates an any-to-any large model for multimodal instruction generation. Finally, UniData enhances data quality by correcting irrelevant and redundant inference flow, leveraging correlations between instruction rounds. To train this pipeline, we also build UniDataset, a dataset comprising 20,000 entries across nine modalities for improved multimodal generation. Our experiments demonstrate that UniData achieves SOTA performance in data quality and can also enhance the understanding and generation capabilities of other multimodal models.
Chinese Translation
多模态大语言模型(MLLMs)正日益被应用于更广泛的真实世界场景。然而,由于高昂的人工成本,为 MLLMs 创建高质量的多模态指令数据集仍然是一个重大挑战。尽管一些方法提出生成指令数据,但它们常常在模态支持方面存在局限,并且难以生成多轮指令。为了解决这些问题,我们提出了 UniData,一个通用指令生成流水线,用于将简单的用户需求转化为多轮、多模态指令。具体而言,UniData 首先将用户需求扩展为多个多样化事件。随后,利用这些事件,UniData 集成一个任意到任意的大模型,用于多模态指令生成。最后,UniData 通过纠正无关和冗余的推理流,并利用指令轮次之间的相关性,来提升数据质量。为了训练该流水线,我们还构建了 UniDataset,这是一个包含跨九种模态的 20,000 个条目的数据集,用于改进多模态生成。我们的实验表明,UniData 在数据质量方面达到了 SOTA 性能,并且还能增强其他多模态模型的理解和生成能力。
cs.AI / 9 / 2610.11380
Equal Path Cost, Unequal Output Effects: Understanding Perturbation Propagation in Diffusion Models
相同路径代价,不同输出效应:理解扩散模型中的扰动传播
Wei Guo, Yaowen Zhang, Xingtong Ge, Jun Zhang
cs.AI
diffusion
扩散模型相关
Abstract
Diffusion models have achieved remarkable success in generative modeling, with their sampling procedures routinely modified to control generation and improve efficiency. These modifications introduce perturbations along the sampling trajectory, raising a central question: how do such perturbations affect generated output? To address this question, we develop a theoretical framework to investigate perturbation propagation, combining dynamical analysis of the sampling process with an information-theoretic characterization of output responses. Within this framework, we quantify perturbation strength using the Kullback--Leibler (KL) divergence between perturbed and reference trajectory distributions, termed as path cost, which is shown to bound, but do not determine, changes in the output distribution. Building on this analysis, we derive a response identity that connects the propagation and accumulation of local perturbations with the information captured by a selected feature mean, explaining why changes in the output distribution can remain undetected by its first-order response. We test our theoretical analysis through controlled interventions at equal path cost in pretrained diffusion models, revealing distinct patterns of output sensitivity across sampling stages and spatial frequencies. To assess whether our framework can diagnose perturbations arising from practical approximations, we apply it to cache-based acceleration and show that our propagation analysis reliably identifies sampling intervals where caching causes larger image errors.
Chinese Translation
扩散模型在生成建模中取得了显著成功,其采样过程经常被修改以控制生成并提高效率。这些修改会在采样轨迹上引入扰动,从而提出一个核心问题:此类扰动如何影响生成的输出?为了解决这个问题,我们建立了一个理论框架来研究扰动传播,将采样过程的动力学分析与输出响应的信息论刻画相结合。在该框架内,我们使用扰动轨迹分布与参考轨迹分布之间的 Kullback--Leibler (KL) 散度来量化扰动强度,称之为路径代价,并表明它能够界定、但并不决定输出分布的变化。在此分析的基础上,我们推导出一个响应恒等式,将局部扰动的传播与累积同选定特征均值所捕获的信息联系起来,从而解释为什么输出分布的变化可能无法被其一阶响应检测到。我们在预训练扩散模型中通过等路径代价下的受控干预来检验我们的理论分析,揭示了输出敏感性在采样阶段和空间频率上的不同模式。为了评估我们的框架能否诊断由实际近似引起的扰动,我们将其应用于基于缓存的加速,并表明我们的传播分析能够可靠地识别缓存导致更大图像误差的采样区间。
cs.AI / 10 / 2610.11483
BridgeGuard: Explicit Safety Drift for Diffusion-based Autonomous Driving
BridgeGuard:面向基于扩散的自动驾驶的显式安全漂移
Zhenjun Qiu, Jianing Huang, Dongang Liu, Baiyu Du, Yixun Niu, Hao Yang, Xinyu Huang, Chuan Hu, Shu Liu
cs.AI
diffusion
扩散模型相关
Abstract
Diffusion-based driving planners capture diverse behaviors but can generate unsafe trajectories under distribution shift. We propose BridgeGuard, a safety-constrained diffusion planning method that progressively strengthens a constraint term during denoising to drive intermediate trajectories toward a scene-dependent safety domain. Corrections operate in a low-dimensional curve space, promoting geometric coherence. A learned module, DistanceFieldNet, predicts a time-dependent distance field from bird's-eye-view features. Value and spatial-gradient supervision at queries sampled beyond expert trajectories teaches this field about both safe and unsafe regions. The learned field supplies the constraint term through safety injection while the pretrained perception backbone and planner remain frozen. We further establish sufficient conditions for terminal safety in an idealized continuous-time bridge. On Bench2Drive, BridgeGuard improves driving score/success rate from 87.99/74.99% to 90.88/76.36% for BridgeDrive and from 80.79/58.18% to 90.46/74.09% for $\text{DiffusionDrive}^{\text{geo}}$, demonstrating cross-model generalization.
Chinese Translation
基于扩散的驾驶规划器能够捕获多样的行为,但在分布偏移下可能生成不安全轨迹。我们提出 BridgeGuard,一种安全约束的扩散规划方法,该方法在去噪过程中逐步增强约束项,以将中间轨迹推向依赖于场景的安全域。修正操作在低维曲线空间中进行,促进几何连贯性。一个学习到的模块 DistanceFieldNet 从鸟瞰图特征预测时间相关的距离场。在专家轨迹之外采样的查询点处进行值监督和空间梯度监督,使该场同时学习到安全区域和不安全区域。学习到的场通过安全注入提供约束项,而预训练的感知骨干网络和规划器保持冻结。我们进一步在理想化的连续时间桥中建立了终端安全的充分条件。在 Bench2Drive 上,BridgeGuard 将 BridgeDrive 的驾驶分数/成功率从 87.99/74.99% 提升至 90.88/76.36%,并将 $\text{DiffusionDrive}^{\text{geo}}$ 的驾驶分数/成功率从 80.79/58.18% 提升至 90.46/74.09%,展示了跨模型泛化能力。
cs.AI / 11 / 2610.11502
Fed-GRPO: Reward-Signal-Driven Federated Group Relative Policy Optimization
Fed-GRPO:奖励信号驱动的联邦组相对策略优化
Pengxin Guo, Shuang Zeng, Zonggen Li, Weiying Zheng, Mengting Liu, Liangqiong Qu
cs.AI
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) have shown strong reasoning capabilities when fine-tuned with reinforcement learning (RL), particularly through Group Relative Policy Optimization (GRPO). However, existing GRPO methods assume centralized access to training data, which may not hold in practice due to privacy or regulatory constraints. To this end, we propose Fed-GRPO, a federated GRPO training framework that addresses these privacy constraints by enabling collaborative reasoning training without sharing raw data, which leverages the reward statistics naturally produced during GRPO training as zero-cost signals to guide aggregation, local training, and communication. Fed-GRPO contains three reward-signal-driven mechanisms: (i) \emph{signal-weighted aggregation} that weights clients by their reward standard deviation, prioritizing clients with stronger learning signals; (ii) \emph{global reward calibration} that re-weights per-prompt objectives based on the local-global reward gap, steering each client toward its relative weaknesses; and (iii) \emph{adaptive sparse communication} that allocates bandwidth based on the informativeness of each client's update. Extensive experiments on mathematical reasoning tasks demonstrate that Fed-GRPO achieves the best performance among all federated methods, clearly outperforms FedAvg and approaches centralized training performance, while losslessly reducing communication by $32\times$ and supporting up to $621\times$ compression under tight bandwidth budgets with only graceful accuracy degradation. Our code is available at https://github.com/HKU-HealthAI/Fed-GRPO.
Chinese Translation
大型语言模型(LLMs)在使用强化学习(RL)进行微调时已展现出强大的推理能力,尤其是通过组相对策略优化(GRPO)。然而,现有的 GRPO 方法假设能够集中式地访问训练数据,而由于隐私或监管限制,这一假设在实践中可能并不成立。为此,我们提出了 Fed-GRPO,一个联邦 GRPO 训练框架,它通过在不共享原始数据的情况下实现协同推理训练来应对这些隐私约束,并利用 GRPO 训练过程中自然产生的奖励统计量作为零成本信号,来引导聚合、本地训练与通信。Fed-GRPO 包含三种奖励信号驱动的机制:(i) \emph{信号加权聚合},它依据客户端的奖励标准差对其进行加权,从而优先考虑具有更强学习信号的客户端;(ii) \emph{全局奖励校准},它基于本地—全局奖励差距对每个提示词的目标函数重新加权,引导每个客户端关注其相对的薄弱之处;(iii) \emph{自适应稀疏通信},它根据每个客户端更新所包含的信息量来分配带宽。在数学推理任务上进行的大量实验表明,Fed-GRPO 在所有联邦方法中取得了最佳性能,明显优于 FedAvg 并接近集中式训练的性能,同时无损地将通信量降低 $32\times$,并在紧张的带宽预算下支持高达 $621\times$ 的压缩,且仅带来平缓的精度下降。我们的代码可在 https://github.com/HKU-HealthAI/Fed-GRPO 获取。
cs.AI / 12 / 2610.11527
AtomWorld-Mirror: Macro-Step World Modeling of Critical Evolution Backbones for Materials Dynamics
AtomWorld-Mirror:面向材料动力学的关键演化骨架的宏步世界建模
Ziming Pan, Ruge Zhang, Haozhi Han, Junkai Zhou, Xingyuan Chen, Yifeng Chen, Yunquan Zhang, Ting Cao, Yunxin Liu, Kun Li
cs.AI · cond-mat.mtrl-sci
diffusion
扩散模型相关
Abstract
Atomistic simulation is a fundamental tool for studying long-term materials evolution, from diffusion and defect dynamics to interfacial reactions and fracture. Yet conventional simulators typically advance at microscopic resolution, spending substantial computation on low-impact local updates before reaching structurally consequential states, an evolutionary-resolution bottleneck that limits long-horizon simulation. We propose AtomWorld-Mirror, a time-aware macro-step world model for the critical evolution backbone of atomic systems. For Step-Wise atomistic simulation, AtomWorld-Mirror distills short micro-event segments into physically reachable transitions between key states, jointly predicting sparse structural edits and accumulated physical time through latent macro-step dynamics. Local reachability, inventory conservation, and continuous-time consistency constrain each transition. By amortizing local atomic physics into a reusable latent macro model and replacing explicit micro-event replay with macro-step inference, this formulation provides a path toward substantially faster prediction of long-term materials evolution while preserving structural validity and time semantics. Across five atomic systems, spanning Cu-rich RPV steel irradiation aging, Cu-Zr metallic glass, and Li$_3$N-based anti-perovskite solid electrolyte, macro-step inference delivers a speed up of $10^3$ to $10^4$ times over event-by-event simulation.
Chinese Translation
原子尺度模拟是研究材料长期演化的基础工具,涵盖从扩散与缺陷动力学到界面反应与断裂等过程。然而,传统模拟器通常以微观分辨率推进,在到达具有结构意义的状态之前,将大量计算花费在低影响的局部更新上,这是一种限制长时程模拟的演化分辨率瓶颈。我们提出 AtomWorld-Mirror,一种面向原子系统关键演化骨架的时间感知宏步世界模型。对于逐步原子模拟,AtomWorld-Mirror 将短微观事件片段蒸馏为关键状态之间物理上可达的跃迁,通过潜在宏步动力学联合预测稀疏结构编辑和累积物理时间。局部可达性、数量守恒和连续时间一致性约束着每个跃迁。通过将局部原子物理摊销到一个可复用的潜在宏模型中,并用宏步推理取代显式的微观事件重放,这一形式为大幅加快长期材料演化预测提供了一条路径,同时保持结构有效性和时间语义。在五个原子系统中,涵盖富 Cu RPV 钢辐照老化、Cu-Zr 金属玻璃以及基于 Li$_3$N 的反钙钛矿固体电解质,宏步推理比逐事件模拟实现了 $10^3$ 到 $10^4$ 倍的加速。
cs.AI / 13 / 2610.11552
Safe, Persistent, and Evolving Agent Harness for Understanding Partially Observable Worlds
面向理解部分可观测世界的安全、持久且可演化的智能体执行框架
Yisen Gao, Yue Guo, Qing Zong, Yiwen Guo, Yangqiu Song
cs.AI
large language model
大语言模型相关
Abstract
Large language model agents can invoke tools fluently, but enterprise workflows demand more than selecting the right tools: actions must strictly comply with organizational policies, tool feedback often conceals hidden side effects under partial observability, and long-horizon tasks require persistent state tracking across multiple records. To address these challenges, we introduce E-Ledger, a multi-agent harness for safe and persistent execution. E-Ledger employs a code approval layer that checks every proposed action against policy before execution, and maintains a world ledger of verified hidden rules alongside evidence-backed dynamic state. Because hidden rules are typically unknown a priori, we further propose WorldAbduct, an abductive, world-model-driven harness evolution framework. WorldAbduct diagnoses execution trajectories across four complementary views (state consistency, world-observation gap, policy-gate correctness, and goal judgment) to hypothesize latent rules, and verifies them through targeted abductive interactions before integrating them into the ledger. On the enterprise benchmark World of Workflows, E-Ledger with WorldAbduct improves safe task completion across four LLM backbones, outperforming the strongest evolution baseline by 5--15 percentage points. Experiments in ScienceWorld and DiscoveryWorld further show that abductive harness evolution carries over to scientific environments. Our code is available at https://github.com/HKUST-KnowComp/E-LEDGER-WorldAbduct.
Chinese Translation
大语言模型智能体可以流畅地调用工具,但企业工作流要求的不只是选择正确的工具:动作必须严格遵守组织策略,工具反馈在部分可观测性下常常掩盖隐藏的副作用,而长时程任务需要跨多条记录进行持久的状态跟踪。为应对这些挑战,我们提出 E-Ledger,一个用于安全且持久执行的多智能体执行框架。E-Ledger 采用代码审批层,在执行前根据策略检查每一个提议的动作,并维护一个世界账本,其中包含经过验证的隐藏规则以及有证据支持的动态状态。由于隐藏规则通常是先验未知的,我们进一步提出 WorldAbduct,一个溯因式、世界模型驱动的执行框架演化框架。WorldAbduct 从四个互补视角(状态一致性、世界-观测差距、策略门正确性和目标判断)诊断执行轨迹,以提出潜在规则假设,并在将其整合到账本之前,通过有针对性的溯因交互对其进行验证。在企业基准 World of Workflows 上,带有 WorldAbduct 的 E-Ledger 在四种 LLM 主干模型上提升了安全任务完成率,比最强的演化基线高出 5--15 个百分点。在 ScienceWorld 和 DiscoveryWorld 中的实验进一步表明,溯因式执行框架演化可以迁移到科学环境。我们的代码可在 https://github.com/HKUST-KnowComp/E-LEDGER-WorldAbduct 获取。
cs.AI / 14 / 2610.11573
Memory Type Varies: Empowering LLM Agents for Long-Term Memory with Diverse Strategies
记忆类型各异:以多样化策略赋能LLM智能体的长期记忆
Yi Wen, Derong Xu, Pengyue Jia, Yichao Wang, Yingyi Zhang, Maolin Wang, Junyi Li, Wenlin Zhang, Xiaopeng Li, Yong Liu, Xiangyu Zhao
cs.AI
large language model
大语言模型相关
Abstract
The memory capabilities of Large Language Models (LLMs) have garnered increasing attention recently. Despite great success achieved, existing retrieval-based memory approaches typically overlook the differences between memories and employ a unified strategy to process all memories, leading to suboptimal performance. Thus, an intuitive question arises: can we categorize memory into different types and select appropriate strategies? However, given the topic-rich, scenario-complex, and boundary-blurred nature of memory scenarios, achieving precise classification of memories is not easy. To address this challenge, we propose a memory multi-class dataset in this paper, termed TriMEM, which provides precise annotations for memory types across diverse scenarios. Building upon this foundation, we propose a novel memory framework, named MemoType, which can adaptively recognize each memory and query type with the learned router model. With the memory and query routing, MemoType can retrieve the memory with corresponding query types and design tailored retrieval strategies, thereby enhancing the retrieval performance. Moreover, we theoretically prove that any single retrieval strategy is subject to a fundamental upper bound on its expected retrieval precision in multi-class corpora, leading to systematic precision degradation. Extensive experiments on three datasets demonstrate that MemoType consistently outperforms existing methods, achieving up to 16.18% improvement in Recall@1.
Chinese Translation
大语言模型(LLM)的记忆能力近来日益受到关注。尽管已取得巨大成功,但现有的基于检索的记忆方法通常忽视记忆之间的差异,并采用统一的策略来处理所有记忆,从而导致性能欠佳。因此,一个直观的问题随之产生:我们能否将记忆划分为不同类型并选择相应的策略?然而,鉴于记忆场景具有主题丰富、情境复杂且边界模糊的特性,实现记忆的精确分类并非易事。为应对这一挑战,本文提出了一个记忆多类别数据集,称为 TriMEM,它为不同场景下的记忆类型提供了精确的标注。在此基础上,我们提出了一种新颖的记忆框架,命名为 MemoType,它能够借助所学习到的路由器模型自适应地识别每个记忆与查询的类型。通过记忆与查询的路由,MemoType 能够检索与相应查询类型匹配的记忆,并设计量身定制的检索策略,从而提升检索性能。此外,我们从理论上证明,在多类别语料库中,任何单一检索策略的期望检索精度都受制于一个根本性的上界,从而导致系统性的精度下降。在三个数据集上的大量实验表明,MemoType 始终优于现有方法,在 Recall@1 上实现了高达 16.18% 的提升。
cs.AI / 15 / 2610.11620
Where to Adapt Matters: Layer-Selective Fine-Tuning for Capability Retention
适配位置很重要:面向能力保持的层选择性微调
Zhiqiang Pang, Zihong Sun, Qi Xie, Jun Shu, Deyu Meng, Zongben Xu
cs.AI
large language model
大语言模型相关
Abstract
Parameter-efficient fine-tuning (PEFT) enables large language models (LLMs) to adapt to specialized tasks, but often at the cost of degrading general capabilities acquired during pretraining. Existing approaches primarily mitigate this trade-off through data replay or regularization, relying on additional data or explicit optimization constraints. We instead focus on a different question: where should adaptation be applied? We find that fine-tuning different Transformer layers produces different target-task gains and degrees of capability degradation, suggesting that not all layers are equally suitable for adaptation. To characterize this difference, we use layer-wise empirical Fisher information to measure target-task sensitivity. However, computing Fisher scores requires backward computation and becomes increasingly expensive for large models. We therefore introduce input--output cosine similarity as a lightweight, forward-only proxy for ranking layer sensitivity. Across models and tasks, layers with lower input--output similarity consistently exhibit higher empirical Fisher scores. Building on this observation, we propose Layer-Selective LoRA (LS-LoRA), which places trainable LoRA adapters only in layers with low input--output similarity. Experiments on mathematical reasoning and code generation show that LS-LoRA improves average target-task performance while retaining substantially more commonsense reasoning capability than standard all-layer LoRA, demonstrating that carefully choosing where to adapt can provide a simple and effective way to balance target-task adaptation and general capability retention.
Chinese Translation
参数高效微调(PEFT)使大型语言模型(LLMs)能够适配专门任务,但往往以退化预训练期间获得的通用能力为代价。现有方法主要通过数据回放或正则化来缓解这一权衡,依赖于额外数据或显式优化约束。我们转而关注一个不同的问题:适配应该应用在哪里?我们发现,微调不同的 Transformer 层会产生不同的目标任务增益和能力退化程度,这表明并非所有层都同样适合适配。为了刻画这种差异,我们使用逐层经验 Fisher 信息来度量目标任务敏感性。然而,计算 Fisher 分数需要反向计算,并且对于大型模型而言变得日益昂贵。因此,我们引入输入--输出余弦相似度,作为一种轻量级、仅前向的代理,用于对层敏感性进行排序。在不同模型和任务中,输入--输出相似度较低的层始终表现出更高的经验 Fisher 分数。基于这一观察,我们提出层选择性 LoRA(LS-LoRA),它仅将可训练的 LoRA 适配器放置在输入--输出相似度较低的层中。在数学推理和代码生成上的实验表明,LS-LoRA 提高了平均目标任务性能,同时比标准的全层 LoRA 保留多得多的常识推理能力,这表明仔细选择适配位置可以提供一种简单而有效的方式来平衡目标任务适配和通用能力保持。
cs.AI / 16 / 2610.11651
Evidence-Traceable Dynamic Interviewer Architecture for Expertise-Adaptive Qualitative Interviews Using Local LLMs
面向专业知识自适应定性访谈的可溯源证据动态访谈者架构:基于本地大语言模型
Aisvarya Adeseye, Jouni Isoaho, Adeyemi Adeseye, Seppo Virtanen, Mohammad Tahir
cs.AI
large language model
大语言模型相关
Abstract
Automated interviewers and conversational agents are increasingly used in research, recruitment, customer service, and education. However, many existing systems rely on fixed question sequences and provide limited context-based personalization without considering participants' knowledge, which can lead to repetitive or irrelevant follow-up questions. Therefore, there is a need for an adaptive interviewing system that can adjust question depth while maintaining conversational continuity and semantic progression. To address this, an Evidence-Traceable Dynamic Interviewer Architecture is presented using a locally hosted Large Language Model (LLM), with the interview continuously adapted throughout the entire conversation based on the participant's responses and evolving context. The interviewer profiles participants' expertise in real time to generate knowledge-appropriate questions, well-articulated responses, and smooth transition messages that support conversational continuity. A five-module prompt-driven architecture and persistent interview-state record support these functions. The interviewer was evaluated with 246 participants. Expertise Profiling module (M3) showed 78.9% exact agreement with independently reported participant expertise, with a weighted Cohen's K of 0.80. Generate Iterative Questions module (M4) showed a strong expertise-complexity association (p=.79, p<.001), and participants reported high relevance (mean 4.41), engagement (mean 4.32), and satisfaction (mean 4.38), providing evidence that the architecture's adaptive components operated consistently with their intended functions while participants reported a positive interview experience.
Chinese Translation
自动化访谈者和对话代理正越来越多地被应用于研究、招聘、客户服务和教育领域。然而,许多现有系统依赖固定的问题序列,并且仅提供有限的基于上下文的个性化,而未考虑参与者的知识水平,这可能导致后续问题重复或无关。因此,需要一种自适应访谈系统,能够在保持对话连续性和语义推进的同时调整问题的深度。为解决这一问题,本文提出了一种可溯源证据的动态访谈者架构(Evidence-Traceable Dynamic Interviewer Architecture),其使用本地托管的大语言模型(LLM),并在整个对话过程中根据参与者的回答和不断演变的上下文持续调整访谈。访谈者实时刻画参与者的专业水平,以生成与知识水平相适配的问题、表述清晰的回应以及支持对话连续性的平滑过渡信息。一个由五个模块组成的提示驱动架构以及持久的访谈状态记录为这些功能提供支持。该访谈者系统在 246 名参与者上进行了评估。专业水平刻画模块(M3)与独立报告的参与者专业水平达到 78.9% 的完全一致率,加权 Cohen's K 为 0.80。生成迭代问题模块(M4)显示出很强的专业水平—复杂度关联(p=.79, p<.001),并且参与者报告了较高的相关性(均值 4.41)、参与度(均值 4.32)和满意度(均值 4.38),这为以下结论提供了证据:该架构的自适应组件与其预期功能保持一致地运行,同时参与者报告了积极的访谈体验。
cs.AI / 17 / 2610.11696
A 3D Characterization Framework for Intelligent Sequential Decision Making
面向智能序贯决策的三维刻画框架
Sadig Gojayev, Carolina Fortuna
cs.AI
large language model
大语言模型相关
Abstract
Puzzles are widely used to evaluate the reasoning capabilities of artificial intelligence (AI) systems for sequential decision making, yet approaches originating from different paradigms are rarely compared under unified conditions. To address this gap, we introduce a three-dimensional characterization framework that enables the analysts of AI methods by 1) projecting them to the Markov decision process (MDP) sequential decision making formalism, 2) degree of autonomy through human prior ranking of their designs and, 3) skill and computational cost. Using this framework, we analyze how representative graph-based, reinforcement learning, and large language model (LLM)-based approaches differ in their design choices and performance characteristics, instantiated respectively by Neurosolver, forward-backward reinforcement learning (FBRL), and automated thought-of-search (AutoToS), including a double-agent extension of thought-of-search (DA-ToS). The analysis relies on the Tower of Hanoi puzzle that provides a controlled benchmark with well-defined rules and scalable complexity, enabling consistent comparison across increasing problem sizes. The 3D characterization reveals that LLM-based methods, due to their weakly constrained action-space design, shift complexity from architecture to inference-time verification, leading to substantially higher memory and runtime costs than Neurosolver and FBRL.
Chinese Translation
谜题被广泛用于评估人工智能(AI)系统在序贯决策中的推理能力,然而源自不同范式的方法很少在统一条件下进行比较。为弥补这一空白,我们引入一个三维刻画框架,使 AI 方法的分析者能够通过以下三个方面进行分析:1)将 AI 方法投影到马尔可夫决策过程(MDP)序贯决策形式体系中,2)通过人类对其设计的先验排序得到的自主程度,以及 3)技能与计算成本。使用该框架,我们分析具有代表性的基于图、强化学习以及基于大语言模型(LLM)的方法在其设计选择和性能特征上有何不同,这些方法分别由 Neurosolver、前向-后向强化学习(FBRL)和自动化搜索思维(AutoToS)实例化,其中包括搜索思维的双智能体扩展(DA-ToS)。该分析依赖于汉诺塔谜题,它提供了一个具有明确定义规则和可扩展复杂度的受控基准,使得能够在不断增大的问题规模上进行一致比较。三维刻画揭示,基于 LLM 的方法由于其弱约束的动作空间设计,将复杂性从架构转移到推理时验证,导致其内存和运行时成本显著高于 Neurosolver 和 FBRL。
cs.AI / 18 / 2610.11715
Internalizer: Portable Context-to-Parameter Mapping for Very Large Language Models
Internalizer:面向超大型语言模型的可移植上下文到参数映射
Peter Devine, Nick Ryan, Benjamin Sirb, Alex Chiocchi
cs.AI · cs.CL · cs.LG
large language model
大语言模型相关
Abstract
Hypernetworks that map a context directly to a LoRA adapter let a large language model carry that context in its weights, but prior work has demonstrated them only on base models of up to 14 billion parameters. We present the Internalizer, a state-of-the-art, portable Context-to-Parameter Mapping hypernetwork that generates document-specific LoRA adapters for the frozen 284B-parameter DeepSeek v4 Flash, a target two orders of magnitude larger than in any previous work. Most of its parameters live in a model-agnostic trunk with only thin entry and exit layers per base model, so it trains cheaply against small models before being ported to the large one. On unseen documents of up to 4096 tokens, the generated adapters reach 84.9% top-1 and 97.8% top-5 teacher-forced accuracy against 63.4% and 83.5% for the base model, with nothing in the context window but a three-word instruction. Once the hypernetwork is trained, a single forward pass turns any document into an adapter for such a model, which could be served alone for speed or alongside the document in the window to raise accuracy further.
Chinese Translation
将上下文直接映射为 LoRA 适配器的超网络,使大型语言模型能够将该上下文承载于其权重中,但先前的工作仅在参数量最高达 140 亿的基础模型上对其进行了验证。我们提出了 Internalizer,这是一种最先进的、可移植的上下文到参数映射超网络,它为冻结的 2840 亿参数 DeepSeek v4 Flash 生成针对特定文档的 LoRA 适配器,这一目标规模比以往任何工作都大两个数量级。其大多数参数位于与模型无关的主干中,每个基础模型仅具有很薄的入口层和出口层,因此它可以先针对小模型廉价地训练,然后再迁移到大型模型上。在长达 4096 个 token 的未见文档上,生成的适配器达到 84.9% 的 top-1 和 97.8% 的 top-5 教师强制准确率,而基础模型为 63.4% 和 83.5%,且上下文窗口中除了一个三词指令外别无他物。一旦超网络训练完成,一次前向传播就能将任何文档转换为此类模型的一个适配器;该适配器可以单独服务以提升速度,也可以与窗口中的文档一起服务以进一步提高准确率。
cs.AI / 19 / 2610.11732
MemTrial: Learning When to Trust Memory in LLM Portfolio Agents
MemTrial:学习何时信任LLM投资组合智能体中的记忆
Guanghao Wu, Zhuo Cai, Shoujin Wang
cs.AI
large language model
大语言模型相关
Abstract
Large language model (LLM) agents for portfolio management learn from experience: they credit each experience in their memory with the outcome of the decisions that used it. In financial markets, however, this outcome mostly reflects the market move shared by all decisions on that date, so the credit tracks the market rather than the experience, and these agents often do worse than simply holding the equal-weight (1/$N$) portfolio. We ask how an agent can credit an experience with what it changes, and answer it by putting memory on trial: drafts of the same decision with and without an experience face the same market, so the outcome they share cancels in their difference. Our agent, MemTrial, drafts each decision with eight combinations of its retrieved experiences, chosen by a fractional factorial design, and credits each experience with its Banzhaf value, the average of these differences. As each date occurs once and each draft is a noisy LLM sample, these credits are noisy and may not hold on new dates. MemTrial therefore pools them across dates and similar experiences with a hierarchical Bayesian model, acts on them only after they have predicted unseen dates, and otherwise stays anchored at a conservative reference such as 1/$N$. On four benchmarks, MemTrial not only benefits from experiences that matter (the best of 15 methods on a semi-synthetic benchmark with known experience quality) but also limits its losses when its values do not hold (at most 2.2\% below 1/$N$ on PortBench and InvestorBench, against 15--38\% for the best experience-learning agent). Averaged over five settings, it improves the utility of the best experience-learning agent by 21.2\%, and with eight LLMs it beats every LLM-based baseline on InvestorBench.
Chinese Translation
用于投资组合管理的大语言模型(LLM)智能体从经验中学习:它们将使用了记忆中每条经验的决策所产生的结果归功于该条经验。然而,在金融市场中,这一结果主要反映的是当日所有决策所共同面对的市场走势,因此这种归因跟踪的是市场而非经验,而这些智能体的表现往往还不如简单地持有等权重(1/$N$)投资组合。我们提出这样一个问题:智能体如何将功劳归于某条经验所带来的改变?我们通过让记忆接受审判来回答这一问题:同一决策在包含与不包含某条经验的情况下所形成的草稿面对相同的市场,因此它们共享的结果在二者的差值中被抵消。我们的智能体 MemTrial 用其检索到的经验的八种组合来为每个决策生成草稿,这些组合由部分因子设计选取,并以每条经验的 Banzhaf 值(即这些差值的平均)来为其记功。由于每个日期只出现一次,且每份草稿都是带噪声的 LLM 采样,这些功劳值是带噪声的,并且可能在新日期上不成立。因此,MemTrial 使用分层贝叶斯模型跨日期和相似经验对它们进行汇总,只有当它们已经预测了未见过的日期之后才据此行动,否则就锚定在一个保守的参考基准上,例如 1/$N$。在四个基准上,MemTrial 不仅能从真正重要的经验中获益(在经验质量已知的半合成基准上,它是 15 种方法中表现最好的),而且在其价值不成立时也能限制损失(在 PortBench 和 InvestorBench 上最多比 1/$N$ 低 2.2\%,而最佳经验学习智能体则低 15--38\%)。在五个设置上取平均,它将最佳经验学习智能体的效用提高了 21.2\%,并且在八个 LLM 的情况下,它在 InvestorBench 上击败了所有基于 LLM 的基线方法。
cs.AI / 20 / 2610.12056
Structure Tax: How Structured Output affects LLMs Performance
结构税:结构化输出如何影响大语言模型性能
Vineet Kumar, Kanishka, Bhuvanesh Mandora
cs.AI
large language model
大语言模型相关
Abstract
Deploying large language models in production often requires constraining outputs to structured formats such as JSON or XML, and prior work treats the resulting accuracy loss as an inherent `structure tax'. We re-examine this claim by evaluating a battery of models, datasets and schemas, measuring task accuracy, confidence calibration, and hidden-state geometry. The tax turns out to depend on schema design rather than on structure per se: reasoning-first field ordering matches or exceeds free-form accuracy, while answer-first ordering causes steep drops, particularly in smaller models. Format sensitivity scales inversely with a task's own structural constraints, and schemas that preserve reasoning order also improve calibration with CKA showing greater separability between correct and incorrect representations in middle transformer layers. Our findings indicate that properly designed structured formats can match or exceed free-form performance, reframing the critical question from `whether to structure' to `how to structure' for optimal reasoning preservation.
Chinese Translation
在生产环境中部署大语言模型通常需要将输出约束为 JSON 或 XML 等结构化格式,而先前的研究将由此产生的准确率损失视为一种固有的“结构税”。我们通过评估一系列模型、数据集和模式来重新审视这一论断,测量任务准确率、置信度校准以及隐藏状态几何结构。结果表明,这种税取决于模式设计,而非结构本身:推理优先的字段顺序能够达到或超过自由形式的准确率,而答案优先的顺序则会导致准确率急剧下降,在较小模型中尤为明显。格式敏感度与任务自身的结构约束程度成反比,而保留推理顺序的模式还能改善校准,CKA 显示在 transformer 中间层中正确表示与错误表示之间具有更强的可分性。我们的发现表明,设计得当的结构化格式能够达到或超过自由形式的性能,从而将关键问题从“是否要结构化”重新表述为“如何结构化”,以实现对推理的最佳保留。
cs.AI / 21 / 2610.12061
When Should Agents Think? Adaptive Reasoning via Cross-Turn Estimation
智能体应当何时思考?基于跨轮次估计的自适应推理
Yiruo Cheng, Shen Huang, Xiaoshuai Song, Jiejun Tan, Guanting Dong, Pengjun Xie, Ji-Rong Wen, Zhicheng Dou
cs.AI · cs.CL
large language model
大语言模型相关
Abstract
Large language model (LLM)-based agents have demonstrated strong capabilities on complex tasks. They typically perform reasoning before each action throughout an interaction trajectory. However, reasoning may not be necessary at every turn, as reasoning produced earlier can continue to support subsequent actions. A key challenge is therefore to determine when existing reasoning remains sufficient and when a new reasoning step is needed, without relying on costly generation-based verification. We find that decreases in the likelihood of subsequent reference actions after removing additional reasoning closely track whether those actions remain recoverable given earlier reasoning, providing an effective and lightweight signal for estimating cross-turn action support. Based on this observation, we propose Reasoning Adaptation through Cross-Turn Estimation (RACE), a training approach for adaptive agent reasoning. RACE introduces a Likelihood-Guided Progressive Reasoning Cover Detection (LoGiC) procedure that progressively identifies reasoning turns whose removal has limited impact on the current and subsequent reference actions. The resulting removal signals are incorporated into both supervised fine-tuning and agentic reinforcement learning, enabling the policy to learn when to reason and when to act directly. Extensive experiments on four representative agent benchmarks show that RACE substantially reduces reasoning cost while maintaining or improving task performance.
Chinese Translation
基于大语言模型(LLM)的智能体已在复杂任务上展现出强大能力。在整个交互轨迹中,它们通常会在每个动作之前进行推理。然而,并非每一轮都需要推理,因为先前产生的推理可以继续支持后续动作。因此,一个关键挑战是:在不依赖昂贵的基于生成的验证的情况下,判断现有推理何时仍然充足,以及何时需要新的推理步骤。我们发现,在移除额外推理之后,后续参考动作似然度的下降,密切反映了在给定较早推理的情况下这些动作是否仍然可恢复,从而为估计跨轮次动作支持提供了一种有效且轻量的信号。基于这一观察,我们提出通过跨轮次估计进行推理适应(Reasoning Adaptation through Cross-Turn Estimation,RACE),一种用于自适应智能体推理的训练方法。RACE 引入了一种似然引导的渐进推理覆盖检测(Likelihood-Guided Progressive Reasoning Cover Detection,LoGiC)过程,该过程逐步识别那些移除后对当前及后续参考动作影响有限的推理轮次。由此得到的移除信号被纳入监督微调和智能体强化学习,使策略能够学习何时进行推理以及何时直接行动。在四个代表性智能体基准上的大量实验表明,RACE 在保持或提升任务性能的同时,大幅降低了推理成本。
cs.AI / 22 / 2610.12085
Is Memorization Context-Sensitive? Prefix-Based Extraction Beyond Isolated Prefixes
记忆是否上下文敏感?超越孤立前缀的基于前缀的提取
Ali Satvaty, Narjes Sharafi, Jirui Qi, Suzan Verberne, Fatih Turkmen
cs.AI
large language model
大语言模型相关
Abstract
Large language models (LLMs) can expose memorized training sequences under prefix-based extraction: given a prefix from a training example, the model may assign high probability to the original continuation. In deployed systems, however, prefixes are rarely evaluated in isolation. They often appear together with instructions, retrieved documents, or other task-specific context, as in retrieval-augmented generation (RAG). This motivates examining whether contextual conditioning mitigates memorization or merely changes the set of memorized samples that become extractable. We investigate this issue through paired item-level measurements of probabilistic suffix extraction. For each prefix-suffix pair, we score the target suffix under an empty prompt and under retrieved contexts of varying relevance, across three open-weight instruction-tuned models. We find that context does not simply erase memorization. Instead, extractable memorization consists of a context-robust core and a context-sensitive boundary. Many samples that are extractable without context remain extractable under the retrieved context, especially as the prefix length increases. At the same time, context mainly affects marginal samples near the extraction threshold: it suppresses some exposures, but also enables new ones that are missed by prefix-only evaluation. These findings qualify the view that RAG reduces memorization risk. Context can lower aggregate extraction by suppressing boundary cases, yet robustly extractable samples persist, and context-enabled extractability remains security-relevant.
Chinese Translation
大型语言模型(LLMs)可以在基于前缀的提取下暴露已记忆的训练序列:给定来自训练示例的前缀,模型可能对原始续写赋予高概率。然而,在部署系统中,前缀很少被孤立评估。它们通常与指令、检索到的文档或其他任务特定上下文一起出现,如在检索增强生成(RAG)中。这促使我们考察上下文条件化是减轻了记忆,还是仅仅改变了可变得可提取的已记忆样本集合。我们通过对概率性后缀提取进行成对的项目级测量来研究这一问题。对于每个前缀-后缀对,我们在空提示下以及在不同相关性的检索上下文下,对目标后缀进行评分,涵盖三个开放权重指令调优模型。我们发现,上下文并不会简单地抹除记忆。相反,可提取的记忆由一个上下文稳健的核心和一个上下文敏感的边界组成。许多在没有上下文时可提取的样本在检索上下文下仍然可提取,尤其是随着前缀长度增加。与此同时,上下文主要影响接近提取阈值的边缘样本:它抑制了一些暴露,但也使一些仅前缀评估所遗漏的新暴露成为可能。这些发现对 RAG 降低记忆风险的观点作出了限定。上下文可以通过抑制边界案例来降低总体提取,但稳健可提取的样本仍然存在,并且上下文使能的提取能力仍然与安全相关。
cs.AI / 23 / 2610.12114
Universal Textual Teaching for LLMs
面向大型语言模型的通用文本教学
Zhanyi Lu, Huan Wang
cs.AI
large language model
大语言模型相关
Abstract
Knowledge distillation (KD) transfers knowledge from stronger Teacher models to weaker Student models, but most methods require training the Student parameters, thereby binding the distilled knowledge to a specific architecture and checkpoint. This implicit representation is difficult to interpret or reuse across models and limits KD for API-only or costly-to-train models. This paper studies knowledge transfer for large language models (LLMs). We introduce Universal Textual Teaching (UTT), a parameter-update-free framework that distills observed Teacher-Student knowledge gaps into a textual, interpretable, and reusable natural-language artifact called Primer. Specifically, UTT first identifies representative gap cases through paired evaluations, and iteratively updates the Primer via multi-role interactions: the Student attempts each task, the Prompter turns evaluation feedback into a teaching instruction, the Teacher provides a targeted demonstration, and the Synthesizer consolidates validated lessons. Empirically, on the challenging math (Omni-MATH-2) and code generation (KernelBench) tasks, extensive results confirm the effectiveness of the method: UTT remarkably raises the Student's accuracy from 9.4% to 48.6% and Fast1 accuracy from 9% to 35% on KernelBench, while increasing mathematical reasoning accuracy from 27.6% to 51.7%. UTT also performs better than representative prompt engineering and parameter-based KD methods. Of note, UTT is shown to be generalizable across different Teachers and Students: a Primer synthesized for one Teacher-Student pair can generalize to other Students that do not participate in the synthesis.
Chinese Translation
知识蒸馏(KD)将知识从更强的教师模型迁移到较弱的学生模型,但大多数方法需要训练学生模型参数,从而将蒸馏得到的知识与特定架构和检查点绑定。这种隐式表示难以解释,也难以在不同模型之间复用,并限制了KD在仅能通过API访问或训练成本高昂的模型上的应用。本文研究大型语言模型(LLMs)的知识迁移。我们提出通用文本教学(UTT),这是一个无需参数更新的框架,它将观察到的教师-学生知识差距蒸馏为一个文本化、可解释且可复用的自然语言产物,称为Primer。具体而言,UTT首先通过配对评估识别具有代表性的差距案例,并通过多角色交互迭代更新Primer:学生模型尝试每个任务,提示器将评估反馈转化为教学指令,教师模型提供有针对性的示范,合成器整合经过验证的经验。实证上,在具有挑战性的数学(Omni-MATH-2)和代码生成(KernelBench)任务上,大量结果证实了该方法的有效性:UTT显著地将学生模型在KernelBench上的准确率从9.4%提高到48.6%,并将Fast1准确率从9%提高到35%,同时将数学推理准确率从27.6%提高到51.7%。UTT还优于代表性的提示工程和基于参数的知识蒸馏方法。值得注意的是,UTT被证明可以在不同的教师模型和学生模型之间泛化:为一个教师-学生对合成的Primer可以泛化到未参与合成的其他学生模型。
cs.AI / 24 / 2610.12124
Use and Disuse: Intent-Structured Experience Consolidation for Memory and Learning in LLM Agents
用进废退:面向 LLM 智能体记忆与学习的意图结构化经验巩固
Xiangyi Zeng, Baihang Liu, Xutong Wang, Ze Jin, Yunpeng Li, Qixu Liu
cs.AI
large language model
大语言模型相关
Abstract
The evolution of Large Language Model agents from single-task execution to long-term autonomous operation highlights the critical challenge of transforming continuous experiences into reusable knowledge. To address this, we propose Hippocam, a hierarchical memory and continual learning architecture. Hippocam draws inspiration from two characteristics of human memory: cognitive processes selectively maintain information relevant to current goals, while long-term memories form gradually through repeated consolidation. Accordingly, Hippocam structures an agent's ongoing work as nested intents. The active context remains centered on the current intent, while completed intents are consolidated into the task-relevant outcomes and state needed for subsequent work, rather than carrying forward their full working details. Concurrently, a recursive prefix consolidation mechanism repeatedly consolidates earlier history, causing long-unused experiences to become increasingly abstract. Original interactions are preserved, allowing the agent to progressively recover finer-grained details through the hierarchy and stop once sufficient information is available. Crucially, when past experiences are recalled and reintegrated into active work, they undergo subsequent consolidation alongside new experiences, thereby being reinforced, supplemented, and updated. Through this memory dynamic of use and disuse, Hippocam connects working context, long-term memory, knowledge accumulation, and skill learning within a single continuously evolving experiential process. This enables agents to learn and evolve capabilities through their own experiences without parameter updates.
Chinese Translation
大语言模型智能体从单任务执行向长期自主运行的演进,凸显了将持续的经验转化为可复用知识这一关键挑战。为解决这一问题,我们提出 Hippocam,一种分层记忆与持续学习架构。Hippocam 的灵感来源于人类记忆的两个特征:认知过程会选择性地保持与当前目标相关的信息,而长期记忆则通过反复巩固逐渐形成。据此,Hippocam 将智能体正在进行的工作组织为嵌套的意图。活跃上下文始终以当前意图为中心,而已完成的意图则被巩固为后续工作所需的任务相关结果与状态,而不是将其完整的工作细节继续携带下去。同时,一种递归前缀巩固机制会反复巩固更早的历史,使长期未被使用的经验变得日益抽象。原始交互被保留下来,使智能体能够沿着该层级逐步恢复更细粒度的细节,并在获得足够信息时停止。关键在于,当过去的经验被回忆并重新整合进当前工作时,它们会与新的经验一起经历后续的巩固,从而得到强化、补充与更新。通过这种用进废退的记忆动态,Hippocam 在单一持续演化的经验过程之中,将工作上下文、长期记忆、知识积累与技能学习连接起来。这使智能体无需参数更新,便能通过自身经验学习和演化能力。
cs.AI / 25 / 2610.12142
Instruction-Conditioned Electromagnetic Spectrum Understanding via Budget-Adaptive Signal Tokenization
通过预算自适应信号标记化的指令条件化电磁频谱理解
Lei Zhai, Zhihao Chang, Shuyuan Yang, Zhixi Feng
cs.AI · eess.SP
large language model
大语言模型相关
Abstract
Electromagnetic spectrum monitoring increasingly requires flexible analysis beyond task-specific recognition and detection. Multimodal large language models offer a unified interface, but extending vision-language models (VLMs) to raw I/Q signals requires tokenization that balances fidelity against a strict budget. For signals, dense encoding causes token costs to grow with observation length, whereas fixed-resolution compression may discard short-duration or localized signal evidence. Thus, we propose \textbf{BATok}, a budget-adaptive signal tokenizer that adjusts token capacity to the input length while allocating that capacity according to the signal content. BATok constructs candidate representations from signal-derived features using lightweight multi-resolution branches, then combines a local energy prior with learnable queries to resample these representations into compact signal tokens. The number of tokens adapts to the input length while remaining strictly bounded. The resulting tokens are projected into the language embedding space of VLMs. We further introduce \textbf{EMSpec-Instruct}, a multimodal instruction dataset aligning raw I/Q signals, waterfall images, and language supervision for modulation recognition, structured detection, and language-conditioned signal grounding. Experiments show that BATok learns effective signal representations and achieves competitive performance across all tasks.
Chinese Translation
电磁频谱监测日益需要超越特定任务识别与检测的灵活分析。多模态大语言模型提供了一种统一接口,但将视觉-语言模型(VLMs)扩展到原始 I/Q 信号需要一种在保真度与严格预算之间取得平衡的标记化方法。对于信号而言,稠密编码会导致标记成本随观测长度增长,而固定分辨率压缩可能会丢弃短时或局部化的信号证据。因此,我们提出了 \textbf{BATok},一种预算自适应信号标记器,它根据输入长度调整标记容量,同时依据信号内容分配该容量。BATok 使用轻量级多分辨率分支从信号导出特征构建候选表示,然后结合局部能量先验与可学习查询,将这些表示重采样为紧凑的信号标记。标记的数量随输入长度自适应,同时保持严格有界。所得的标记被投影到 VLMs 的语言嵌入空间中。我们进一步引入了 \textbf{EMSpec-Instruct},一个多模态指令数据集,为调制识别、结构化检测和语言条件化信号定位对齐原始 I/Q 信号、瀑布图和语言监督。实验表明,BATok 学习到了有效的信号表示,并在所有任务上取得了有竞争力的性能。
cs.AI / 26 / 2610.12304
Looking Inside LLMs: Small-World Connectivity as a Signature of Reasoning Performance
窥探大语言模型的内部:小世界连接性作为推理性能的标志
Zheng Huang, Sansheng Cao, Enpei Zhang, Weikang Qiu, Elynn Chen, Xiang Zhang, Yaoqing Yang, Rex Ying, Dawei Zhou, Yujun Yan
cs.AI
large language model
大语言模型相关
Abstract
Understanding large language model (LLM) reasoning requires looking beyond behavioral performance to examine how reasoning ability is reflected in internal organization. Inspired by neuroscience findings linking higher intelligence to stronger small-world organization in functional brain networks, we investigate small-world connectivity as a structural signature of LLM reasoning. We construct functional graphs from attention-head activation similarities and find that a higher small-world index (SWI), capturing local clustering and short global paths, consistently correlates with better fluid reasoning performance across models and training checkpoints. Since local clustering is central to small-world organization, we further examine how heads important for model performance connect within and across communities. We find that these heads tend to have a larger share of connection weight within their own communities (high core scores) and a more concentrated weight distribution across communities (low bridge scores). These observations motivate the hypothesis that high core and low bridge scores serve as structural indicators of head importance for reasoning capability. We validate this hypothesis through pruning, introducing Small-World Allocation (SWA), a hierarchical sparsity allocation method guided by these scores. Across six LLMs, SWA better preserves small-world organization and model performance than competing allocation strategies, reducing WikiText perplexity by up to 20%. Together, these findings identify small-world functional connectivity as a measurable signature of LLM reasoning performance, offering a structural perspective that complements behavioral evaluation.
Chinese Translation
理解大语言模型(LLM)的推理能力,需要超越行为表现,考察推理能力如何反映在其内部组织中。受神经科学发现的启发——即更高的智力与功能性脑网络中更强的小世界组织相关——我们研究小世界连接性作为 LLM 推理的一种结构性标志。我们从注意力头的激活相似性构建功能图,发现更高的小世界指数(SWI)——它刻画了局部聚类与较短的全局路径——在不同模型和训练检查点上均与更好的流体推理性能一致相关。由于局部聚类是小世界组织的核心,我们进一步考察对模型性能重要的注意力头如何在社区内部以及跨社区之间连接。我们发现,这些头在其自身社区内往往具有更大比例的连接权重(高核心分数),并且在跨社区的权重分布上更为集中(低桥接分数)。这些观察促使我们提出假设:高核心分数与低桥接分数可作为注意力头对推理能力重要性的结构性指标。我们通过剪枝验证了这一假设,提出了小世界分配(SWA),一种由这些分数引导的层次化稀疏分配方法。在六个 LLM 上,SWA 比竞争性的分配策略更好地保持了小世界组织与模型性能,将 WikiText 困惑度最多降低 20%。总之,这些发现表明小世界功能连接性是 LLM 推理性能的一种可测量标志,提供了一种补充行为评估的结构性视角。
cs.AI / 27 / 2610.12313
Verdict Without the Rule: Diagnosing and Auditing Regulatory Rule Sensitivity in LLM Compliance Systems
无规则之裁决:诊断与审计LLM合规系统中的监管规则敏感性
Saisab Sadhu, Aadit Sengupta, Vinay kumar Sankarapu, Pratinav Seth
cs.AI · cs.CL · cs.CY
large language model
大语言模型相关
Abstract
Large language model compliance systems are deployed on the assumption that a verdict depends on the regulatory rule it is given. We test this directly across five models and 20 regulatory and platform-policy domains: delete, swap, or negate the governing rule while holding the case fixed, and check whether the verdict changes (OCS) or the model's internal representation of compliance shifts at all (ICS-delta). Neither moves much: models' verdicts are often invariant to substantial perturbations of the supplied rule, and the guard model, evaluated here under a custom-rule adaptation of its native taxonomy, is the least rule-sensitive and least accurate of the five, barely above chance (51%, versus 90-92% for general-purpose models). This reflects easy cases more than blanket neglect: on cases where deleting the rule changes a previously correct model prediction, models do track it closely. Neither better prompting nor direct intervention on the model's internal representations closes this gap. Accuracy alone does not establish that a compliance verdict is grounded in the supplied rule.
Chinese Translation
大语言模型合规系统被部署时基于这样一个假设:裁决取决于其所被给定的监管规则。我们在五个模型和20个监管与平台政策领域中直接对此进行检验:在保持案件固定的同时,删除、替换或否定起支配作用的规则,并检查裁决是否发生变化(OCS),或者模型对合规性的内部表示是否发生任何偏移(ICS-delta)。两者都没有太大变化:模型的裁决往往对所提供规则的实质性扰动保持不变,而防护模型——在此处以其原生分类体系的自定义规则适配版进行评估——是五个模型中对规则最不敏感且最不准确的,仅略高于随机水平(51%,而通用模型为90–92%)。这更多反映的是简单案例,而非一概忽视:在删除规则会改变模型先前正确预测的那些案例上,模型确实会密切跟踪规则。更好的提示或对模型内部表示的直接干预都无法弥合这一差距。仅凭准确率并不能确立合规裁决是以所提供规则为依据的。
cs.AI / 28 / 2610.12325
Prior or Feedback? What an LLM Uses When Adapting Neural Operators
先验还是反馈?LLM 在适配神经算子时使用了什么
Julian Chan, Javier Mora Jimenez
cs.AI · cs.LG · physics.comp-ph
large language model
大语言模型相关
Abstract
Do LLM scientific agents rely only on their initial task context, or do they adapt their decisions in response to experimental feedback? We study this question in neural operator adaptation, where a large language model (LLM) selects fine-tuning configurations under a limited trial budget. Across transfers within and between partial differential equation (PDE) families, the LLM achieves lower held-out test nRMSE than random search and Bayesian optimisation in nearly every matched comparison. Endpoint performance alone cannot distinguish what happens, so we verify each attribution with controlled interventions. Before observing any validation score, the LLM's first configuration already ranks near the top of the corresponding random-search pool, indicating a useful initial bias. A complementary cold-start intervention shows that the selected base learning rate shifts with the PDE description. Once feedback becomes available, reassigning validation scores among evaluated configurations changes the next proposal in every case tested, whereas a value-preserving rewrite produces no comparable aggregate effect. These interventions establish that the LLM's decision-level actions respond to the given task and observed outcomes, showing that it combines a task-dependent prior with sensitivity to experimental feedback.
Chinese Translation
LLM 科学智能体是仅依赖其初始任务上下文,还是会根据实验反馈调整其决策?我们在神经算子适配这一场景中研究该问题,其中大语言模型(LLM)在有限的试验预算下选择微调配置。在偏微分方程(PDE)族内部以及族之间的迁移中,LLM 在几乎每一项配对比较中都取得了比随机搜索和贝叶斯优化更低的留出测试 nRMSE。仅凭最终性能无法辨别究竟发生了什么,因此我们通过受控干预来验证每一种归因。在观察到任何验证分数之前,LLM 的第一个配置就已经在相应的随机搜索池中排名接近前列,这表明存在一种有用的初始偏置。一项互补的冷启动干预表明,所选的基础学习率会随 PDE 描述而变化。一旦反馈变得可用,在被评估的配置之间重新分配验证分数,在测试的每一种情形下都会改变下一个提议,而保持数值不变的改写则不会产生可比的总体效应。这些干预表明,LLM 在决策层面的行为会响应给定任务和观测到的结果,说明它将对任务依赖的先验与对实验反馈的敏感性结合了起来。
cs.AI / 29 / 2610.12345
Overcoming Prior Barriers: Supervised Fine-Tuning under Long-Tail Distribution
克服先验壁垒:长尾分布下的监督微调
Haohui Wang, Jiahao Xu, Wangzhi Zhan, Tong Zeng, Dongqi Fu, Hong Li, Swastik Roy, Naren Ramakrishnan, Chris North, Jian Kang, Yujun Yan, Dawei Zhou
cs.AI · cs.CL · cs.LG
large language model
大语言模型相关
Abstract
Supervised fine-tuning (SFT) adapts pretrained large language models (LLMs) to downstream tasks, but the required concepts can receive substantially different levels of pretrained support. Frequent concepts are more likely to be well learned, whereas rare concepts may remain weakly represented. We introduce a novel notion named prior barrier to quantify how strongly the pretrained model supports competing concepts over the target concept. We observe that prior barriers follow a long-tail distribution, placing head and tail concepts at different starting points for SFT: head concepts face lower prior barriers, whereas tail concepts require additional instructions to overcome their higher prior barriers. Our theoretical analysis further derives a predictive risk bound for SFT under long-tail prior barriers, explicitly characterizing how the prior barrier and accumulated SFT evidence jointly determine predictive performance. Motivated by this prior barrier-dependent demand, we propose PASS, an adaptive SFT instruction selection method that constructs reference-derived concepts and estimates the distinguishing evidence provided by each instruction, and adaptively allocates the selection budget toward concepts that remain insufficiently covered under the current selection. In this way, PASS jointly considers which instructions can provide useful evidence and where additional supervision is needed under a limited budget. Experiments show that our method consistently outperforms seven state-of-the-art instruction selection methods on four backbone-budget settings. An ablation study further shows that PASS's adaptive allocation consistently improves over uniform allocation.
Chinese Translation
监督微调(SFT)将预训练大语言模型(LLMs)适配到下游任务,但所需概念可能获得程度差异显著的预训练支持。高频概念更可能被充分学习,而稀有概念可能仍然表征薄弱。我们引入一个名为先验壁垒的新概念,用以量化预训练模型对竞争概念相对于目标概念的支持强度。我们观察到先验壁垒遵循长尾分布,这使头部概念与尾部概念在SFT中处于不同的起点:头部概念面临较低的先验壁垒,而尾部概念则需要额外的指令来克服其较高的先验壁垒。我们的理论分析进一步推导了长尾先验壁垒下SFT的预测风险界,明确刻画了先验壁垒与累积的SFT证据如何共同决定预测性能。受这种依赖于先验壁垒的需求启发,我们提出PASS,一种自适应SFT指令选择方法,它构建由参考导出的概念,估计每条指令所提供的区分性证据,并自适应地将选择预算分配给在当前选择下仍未得到充分覆盖的概念。通过这种方式,PASS在有限预算下同时考虑了哪些指令能够提供有用的证据,以及哪些地方需要额外的监督。实验表明,我们的方法在四种骨干模型-预算设置上均持续优于七种最先进的指令选择方法。消融研究进一步表明,PASS的自适应分配相较于均匀分配持续带来提升。
cs.AI / 30 / 2610.12361
Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought Faithfulness
被引用却未被参考:法律思维链忠实性的反事实审计
Saisab Sadhu, Shreeyans Arora, Pratinav Seth
cs.AI · cs.CL · cs.CY
large language model
大语言模型相关
Abstract
Large language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence that the decision follows from it. We test this directly: holding case facts fixed, we substitute the named legal authority for an unrelated one and decode a model's evolving verdict from its hidden states. Across seven open-weight models (8B-70B) and four benchmarks spanning judicial and contractual reasoning, when explicitly required to justify a verdict by naming the governing authority, models name the correct one in 66.7%-100% of generations, while the verdict changing when the authority changes is far less consistent: 0.0%-21.7% on CaseHOLD, 30.0%-76.7% on ECHR and SCOTUS, and 43.3%-50.0% on ContractNLI. Neither scale nor a purpose-built legal-reasoning model (a best-effort LoRA reproduction; Section 6) closes this gap. A red-teaming evaluation on five core models finds compliance with an adversarial instruction hidden in the case facts (73.3%-96.4%) exceeds verdict-swap sensitivity by a wide margin, holding without exception across model rankings. Naming a legal authority is thus a poor proxy for a verdict's dependence on it, while the same verdict remains separately vulnerable to adversarial manipulation. Both findings replicate across checks ruling out prompt-wording noise and confounded sampling, and bear directly on the use of generated legal explanations as compliance or audit artefacts.
Chinese Translation
大语言模型越来越多地通过指明判决背后的法规或先例来为法律决策提供理由,并将其视为该决策由此得出的证据。我们直接对此进行检验:在保持案件事实不变的情况下,我们将所引用的法律权威替换为一个不相关的法律权威,并从模型的隐藏状态中解码其不断演变的判决。在七个开放权重模型(8B-70B)和四个涵盖司法推理与合同推理的基准上,当明确要求通过指明支配性权威来为判决提供理由时,模型在 66.7%-100% 的生成中指出了正确的权威,而当权威改变时判决也随之改变的情况则远不一致:在 CaseHOLD 上为 0.0%-21.7%,在 ECHR 和 SCOTUS 上为 30.0%-76.7%,在 ContractNLI 上为 43.3%-50.0%。无论是规模,还是专门构建的法律推理模型(一次尽力而为的 LoRA 复现;第 6 节),都未能弥合这一差距。对五个核心模型进行的红队评估发现,对隐藏在案件事实中的对抗性指令的遵从率(73.3%-96.4%)以很大幅度超过判决替换敏感度,并且这一结果在所有模型排名中无一例外地成立。因此,指明某一法律权威并不能很好地代理判决对其的依赖性,而同一判决仍然单独地容易受到对抗性操纵的影响。这两项发现在排除提示措辞噪声和混杂抽样的各种检验中均得到复现,并直接关系到将生成的法律解释用作合规或审计产物。
cs.AI / 31 / 2610.12391
GeoReform: Reflective Formalization Evolution for Multimodal Geometry Problem Solving
GeoReform:面向多模态几何问题求解的反思式形式化演化
Jialu Wang, Ruichen Zhang, Xiaoou Liu, Hua Wei, Tianlong Chen
cs.AI
large language model
大语言模型相关
Abstract
Multimodal large language models (MLLMs) often struggle to identify and use geometric relations in diagrams. Recent methods address this challenge by converting geometric entities, relations, and constraints into explicit textual representations for the model to reason over. However, effective formalization is highly non-trivial: on Geometry3K, structure injection fixes 28 errors but introduces 13 new ones among 200 examples. Redundant relations can distract the model, while ambiguous references to diagram elements can lead it to apply constraints incorrectly. This suggests that the key challenge is not merely extracting more geometric facts, but organizing them into representations that support downstream reasoning. To fully exploit the power of formalization, we further propose GeoReform, a reflective formalization evolution framework that treats formalization as an optimizable policy rather than a fixed parser output. GeoReform executes the full reasoning pipeline, collects failed rollouts, diagnoses defects in the current representation, and mutates the policy to better select, ground, group, and present geometric entities, relations, constraints, and targets. On Geometry3K, GeoReform improves Qwen3VL-2B accuracy from 42.0\% to 56.0\%. Extensive experiments and analyses across geometry reasoning benchmarks demonstrate that effective formalization is crucial for improving multimodal geometry reasoning.
Chinese Translation
多模态大语言模型(MLLMs)常常难以识别和使用图中的几何关系。最近的方法通过将几何实体、关系和约束转换为显式的文本表示,供模型进行推理,以应对这一挑战。然而,有效形式化绝非易事:在 Geometry3K 上,结构注入修复了 28 个错误,但在 200 个样例中引入了 13 个新错误。冗余关系可能分散模型注意力,而对图元素的模糊指代则可能导致其错误地应用约束。这表明,关键挑战不仅仅在于提取更多几何事实,而在于将它们组织成支持下游推理的表示。为了充分利用形式化的力量,我们进一步提出 GeoReform,一个反思式形式化演化框架,它将形式化视为可优化的策略,而非固定的解析器输出。GeoReform 执行完整的推理流程,收集失败的 rollout 轨迹,诊断当前表示中的缺陷,并变异策略,以更好地选择、定位、分组和呈现几何实体、关系、约束和目标。在 Geometry3K 上,GeoReform 将 Qwen3VL-2B 的准确率从 42.0\% 提升到 56.0\%。在几何推理基准上的大量实验和分析表明,有效形式化对于提升多模态几何推理至关重要。
cs.AR / 32 / 2610.10858
RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
RFChipAgent:用于模拟/射频芯片设计的多智能体 AI 流程
Awani Khodkumbhe, Yunfei Feng, Raj Rangarajan, Kevin Wang, Kamal Sahota
cs.AR · cs.AI · cs.LG · cs.MA · eess.SY
large language model
大语言模型相关
Abstract
Analog/RF circuits remain the critical interface between digital computation and the physical world, and emerging standards from Wi-Fi 7 to 6G place stringent demands on them, yet analog/RF design remains one of the most labor-intensive steps in chip development. We present RFChipAgent, a first-of-its-kind multi-agent flow of large language model (LLM) agents for end-to-end analog/RF circuit design automation, in which AI agents collaboratively orchestrate the complete design flow under human supervision. RFChipAgent is built around four technical pillars. First, a multimodal retrieval-augmented generation (RAG) subsystem with private per-document FAISS indexing extracts design knowledge from existing engineering documentation. Second, a topology agent drives topology selection, and a schematic and testbench agent automates circuit and testbench assembly. Third, a closed-loop hybrid circuit-sizing engine combines Tree-structured Parzen Estimator (TPE) and CMA-ES optimization, evaluating every candidate in a simulator-in-the-loop framework. Fourth, a trust-scored simulation database accumulates verified performance data and builds an adaptive optimization model that informs subsequent trials. We validate RFChipAgent on a family of GF22FDSOI 60 GHz wideband mm-wave low-noise amplifier (LNA) topologies, demonstrating automated topology generation, specification-driven design-space exploration, and simulator-guided optimization. Experimental results show substantial reductions in design effort while maintaining signoff-quality verification. This work establishes a foundation for LLM-driven multi-agent electronic design automation (EDA) for analog/RF circuits.
Chinese Translation
模拟/射频电路仍然是数字计算与物理世界之间的关键接口,而从 Wi-Fi 7 到 6G 的新兴标准对其提出了严苛要求,然而模拟/射频设计仍然是芯片开发中最劳动密集的环节之一。我们提出 RFChipAgent,这是一种首创的、由大语言模型(LLM)智能体构成的多智能体流程,用于端到端模拟/射频电路设计自动化,其中 AI 智能体在人类监督下协同编排完整的设计流程。RFChipAgent 围绕四个技术支柱构建。首先,一个具有私有逐文档 FAISS 索引的多模态检索增强生成(RAG)子系统从现有工程文档中提取设计知识。其次,拓扑智能体驱动拓扑选择,而原理图和测试台智能体自动完成电路与测试台的组装。第三,一个闭环混合电路尺寸调整引擎结合树结构 Parzen 估计器(TPE)和 CMA-ES 优化,在仿真器在环框架中评估每个候选方案。第四,一个带信任评分的仿真数据库累积经过验证的性能数据,并构建一个自适应优化模型来指导后续试验。我们在一个 GF22FDSOI 60 GHz 宽带毫米波低噪声放大器(LNA)拓扑系列上验证 RFChipAgent,展示了自动化拓扑生成、规格驱动的设计空间探索以及仿真器引导的优化。实验结果表明,在保持签核质量验证的同时,设计工作量显著减少。这项工作为面向模拟/射频电路的 LLM 驱动的多智能体电子设计自动化(EDA)奠定了基础。
cs.AR / 33 / 2610.11284
DynaTE: Accelerating Diffusion LLMs via Dynamic Token Execution
DynaTE:通过动态 token 执行加速扩散 LLM
Minghan Jiang, Jiayi Wang, Shuaiting Li, Haibin Shen, Kejie Huang
cs.AR · cs.AI
diffusion
扩散模型相关
Abstract
Diffusion-based LLMs (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs by enabling bidirectional parallel refinement, alleviating the sequential decoding bottleneck of AR generation. However, their parallel iterative refinement mismatches AR accelerators optimized for sequential decoding and their discrete token generation differs from DiT accelerators designed for continuous denoising. Recent dLLM accelerators have explored workload-specific optimizations to reduce vocabulary processing overhead and redundant computation across denoising iterations. However, these approaches retain all tokens in parallel execution, despite varying token refinement utility and execution requirements. This paper presents DynaTE, a hardware--software co-design architecture that dynamically adapts accelerator execution to evolving token states during dLLM decoding. DynaTE first enables adaptive token execution by skipping low-utility token computation, while a dimension-reconfigurable PE array maintains high utilization under varying active-token patterns. Second, DynaTE exploits dynamic token dependencies through FLDD to refine a small number of locally dependent tokens within the current iteration, reducing the overall number of denoising iterations, while a Merge--Split--Merge dataflow hides the resulting serial overhead. Third, a streaming vocabulary engine interleaves multiple token streams from the LM head to accommodate irregular output variations caused by selective token computation and uneven vocabulary-selection demands. Evaluated on two representative dLLMs, DynaTE achieves 2.05--2.78$\times$ speedup and 2.99--3.93$\times$ higher energy efficiency over state-of-the-art dLLM accelerators, while delivering 2.55$\times$ speedup and 6.07$\times$ higher energy efficiency over Jetson AGX Orin.
Chinese Translation
基于扩散的 LLM(dLLMs)最近作为自回归(AR)LLM 的一种有前景的替代方案出现,其通过支持双向并行细化,缓解了 AR 生成的顺序解码瓶颈。然而,它们的并行迭代细化与为顺序解码优化的 AR 加速器不匹配,并且其离散 token 生成不同于为连续去噪设计的 DiT 加速器。最近的 dLLM 加速器已经探索了面向特定工作负载的优化,以减少词表处理开销和跨去噪迭代的冗余计算。然而,这些方法尽管 token 细化效用和执行需求各不相同,仍在并行执行中保留所有 token。本文提出 DynaTE,一种硬件--软件协同设计架构,可在 dLLM 解码过程中动态调整加速器执行以适应不断演变的 token 状态。DynaTE 首先通过跳过低效用 token 计算实现自适应 token 执行,而维度可重构的 PE 阵列在不同活跃 token 模式下保持高利用率。其次,DynaTE 通过 FLDD 利用动态 token 依赖关系,在当前迭代内细化少量局部依赖的 token,减少总体去噪迭代次数,同时 Merge--Split--Merge 数据流隐藏由此产生的串行开销。第三,流式词表引擎将来自 LM head 的多个 token 流交织在一起,以适应由选择性 token 计算和不均匀词表选择需求引起的不规则输出变化。在两个代表性 dLLM 上评估,DynaTE 相比最先进的 dLLM 加速器实现了 2.05--2.78$\times$ 加速和 2.99--3.93$\times$ 更高的能效,同时相比 Jetson AGX Orin 实现了 2.55$\times$ 加速和 6.07$\times$ 更高的能效。
cs.CL / 34 / 2610.10845
Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute
AI 的真正长期记忆:一个比重新计算更快、更便宜的 5000 万 token 窗口
Sietse Schelpe
cs.CL · cs.AI · cs.DC · cs.LG
large language model
大语言模型相关
Abstract
A large language model can only use the text that fits in its context window, and it recomputes its internal key-value (KV) state for a prompt every time the prompt is sent. We test a memory layer, the public package galahad-kv, that saves the KV state of each block of about 16,000 tokens to encrypted local NVMe disk and loads it back later, byte-exact, without recomputing it. We ran it on 50,000,000 tokens of real public text, served through vLLM on one NVIDIA H100, with Gemma 4 12B and Gemma 4 31B. Every block we probed was loaded back from the encrypted store with no recompute (100 of 100, at depths from 0 to 50M tokens) on both models. Loading a block was 2.8x to 4.3x faster than recomputing it and used 8.8x to 12.3x less GPU energy, and GPU memory stayed flat over the whole 50M-token stream. Asked about facts planted millions of tokens earlier, the 12B model gave the right answer 82 times out of 100 and the 31B model 98 times out of 100. Neither model made up an answer. The limits are as follows. This is reuse of stored state, not a wider attention window: one block is loaded at a time, and how well a question is answered depends on the model. Writing the memory is a one-time cost, and the store takes terabytes of local NVMe disk. We describe the test protocol, which is built to resist common ways of gaming long-context benchmarks, and give a single-GPU reproduction that uses public software and a free licence for the package.
Chinese Translation
大型语言模型只能使用能够放入其上下文窗口的文本,并且每次发送提示时,它都会为提示重新计算其内部键值(KV)状态。我们测试了一个记忆层,即公开软件包 galahad-kv,它将每个约 16,000 个 token 的块的 KV 状态保存到加密的本地 NVMe 磁盘,并在之后逐字节精确地将其加载回来,而无需重新计算。我们在 50,000,000 个 token 的真实公开文本上运行了它,通过 vLLM 在一台 NVIDIA H100 上提供服务,并使用 Gemma 4 12B 和 Gemma 4 31B。在两个模型上,我们探测的每一个块都从加密存储中加载回来,没有重新计算(100 次中 100 次,深度从 0 到 5000 万 token)。加载一个块比重新计算它快 2.8 倍到 4.3 倍,并且使用的 GPU 能量少 8.8 倍到 12.3 倍,而且在整个 5000 万 token 流中,GPU 内存保持平稳。当被问及数百万个 token 之前植入的事实时,12B 模型在 100 次中给出了 82 次正确答案,31B 模型在 100 次中给出了 98 次正确答案。两个模型都没有编造答案。限制如下。这是对存储状态的复用,而不是更宽的注意力窗口:一次只加载一个块,并且一个问题回答得有多好取决于模型。写入记忆是一次性成本,而存储需要数 TB 的本地 NVMe 磁盘。我们描述了测试协议,该协议旨在抵御常见的操纵长上下文基准测试的方法,并提供了一个单 GPU 复现,该复现使用公开软件和该软件包的免费许可证。
cs.CL / 35 / 2610.10871
Sparse Attention Is Matrix Approximation, Not Choosing from a Bag of Values
稀疏注意力是矩阵近似,而非从一袋值中挑选
Fang Wan, Xufeng Liu, Fan Li, Yi Liu
cs.CL
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) achieve strong performance across many domains, but their efficiency is limited by the quadratic cost of attention with respect to prompt length. Sparse attention reduces this cost by retaining only a small fraction of query-key interactions to approximate the full attention matrix. However, existing methods are trapped in a mathematically wrong view: they simply keep large scalar entries or high-mass regions of the attention matrix. This treats the attention matrix as a bag of values, ignoring that it is used as a structured matrix whose entries jointly determine the attention output through multiplication with value vectors. We argue that this is the core conceptual issue: sparse attention should be formulated as matrix approximation, not as blindly choosing the largest values from a bag of entries. Based on this view, we propose Matrix Approximation Sparse Attention (MASA). MASA replaces raw attention-mass ranking with a closed-form score that measures how much each sparse unit reduces matrix-product approximation error. As a theory-grounded plug-in correction, MASA can be added to existing sparse attention frameworks without changing their sparse kernels or budgets. Extensive experiments across multiple sparse attention methods, benchmarks, and LLM backbones show consistent accuracy gains, supporting both MASA and the matrix-approximation view of sparse attention.
Chinese Translation
大语言模型(LLMs)在许多领域都取得了强劲的性能,但其效率受到注意力相对于提示长度的二次方成本的限制。稀疏注意力通过仅保留一小部分查询-键交互来近似完整的注意力矩阵,从而降低这一成本。然而,现有方法陷入了一种在数学上错误的观点:它们只是简单地保留注意力矩阵中较大的标量项或高质量区域。这种做法把注意力矩阵当作一袋值,忽视了它实际上是被用作一个结构化矩阵,其各项通过与值向量相乘共同决定注意力输出。我们认为这是核心的概念性问题:稀疏注意力应当被表述为矩阵近似,而不是盲目地从一袋项中挑选最大的值。基于这一观点,我们提出了矩阵近似稀疏注意力(MASA)。MASA用闭式分数取代了原始的注意力质量排序,该分数衡量每个稀疏单元能在多大程度上降低矩阵乘积的近似误差。作为一种有理论依据的即插即用修正,MASA可以在不改变现有稀疏注意力框架的稀疏核或预算的情况下被添加到这些框架中。跨多种稀疏注意力方法、基准和LLM主干网络的广泛实验显示出持续一致的准确率提升,这既支持了MASA,也支持了稀疏注意力的矩阵近似观点。
cs.CL / 36 / 2610.10918
Large Language Models for Machine Translation Quality Annotation: Humans and Models Are Both Challenged
大型语言模型用于机器翻译质量标注:人类与模型均面临挑战
Hala Almaghout, Christian Federmann, Qin Gao
cs.CL
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) are considered to be a more efficient and cost-effective alternative to human judgment for Machine Translation (MT) evaluation. With MT evaluation spanning a large number of language pairs, domains and levels of annotation granularity, LLMs must be thoroughly evaluated across these dimensions before being reliably used as alternatives to human evaluation. In this paper, we evaluate the performance of LLMs for two prominent MT quality evaluation schemes: Multidimensional Quality Metrics (MQM) and Error Span Annotation (ESA) by comparing their agreement with human annotators. We present results on a long-context test set of 70 language pairs and the publicly available WMT23 and WMT25 data, investigating both score and error span annotation agreement across a variety of language pairs and domains. Our results show that while LLM agreement with human annotators exceeds agreement between human annotators for some evaluation tasks, both vary substantially across annotation schemes, language pairs and domains and remain unreliable for most settings. Furthermore, we identify challenges facing both human and LLM annotators: humans are particularly challenged by fine-grained MQM annotations and low-resource language pairs, while LLMs struggle with minor errors, wrong language variants and error span annotation. Our results highlight the potential for improvement for both human and LLM annotation performance, possibly through human-LLM collaborative annotation pipelines that address the reliability issues identified in this work.
Chinese Translation
大型语言模型(LLMs)被认为是在机器翻译(MT)评估中替代人类判断的一种更高效且更具成本效益的方案。由于机器翻译评估涵盖大量语言对、领域和标注粒度级别,在LLMs能够可靠地用作人工评估的替代方案之前,必须在这些维度上对其进行全面评估。在本文中,我们评估了LLMs在两种重要的机器翻译质量评估方案中的表现:多维质量指标(MQM)和错误跨度标注(ESA),通过比较它们与人工标注者的一致性。我们展示了在包含70个语言对的长上下文测试集以及公开可用的WMT23和WMT25数据上的结果,考察了多种语言对和领域上的分数与错误跨度标注一致性。我们的结果表明,尽管在某些评估任务中,LLM与人工标注者之间的一致性超过了人工标注者之间的一致性,但二者在标注方案、语言对和领域之间均存在显著差异,并且在大多数设置下仍然不可靠。此外,我们识别出人类和LLM标注者共同面临的挑战:人类尤其难以应对细粒度MQM标注和低资源语言对,而LLMs则难以处理细微错误、错误语言变体和错误跨度标注。我们的结果凸显了人类和LLM标注性能的提升潜力,可能通过人类-LLM协作标注流程来解决本工作中识别出的可靠性问题。
cs.CL / 37 / 2610.10946
AI4Fire: Evaluating Large Language Models on Wildfire Tasks
AI4Fire:评估大型语言模型在野火任务上的表现
Yue Zhao, Xiyang Hu, Zuobin Xiong, Zhangyu Wang, Ruolin Li
cs.CL · cs.LG
large language model
大语言模型相关
Abstract
Large language models (LLMs) are entering wildfire management, where overstated evaluations can cost property and lives. How do they perform on wildfire tasks, with and without grounding? Bare means a model receives the task input alone. Grounded means it also receives one task-specific addition: for smoke detection, a smoke-free reference frame from the same camera. AI4Fire runs six core models bare and grounded on five wildfire tasks, zero-shot; a sweep adds 29 more. Our literature search on fire tasks found 138 works; none combines this roster, task coverage, and paired bare and grounded runs. We report three findings. (1) Grounding helped most where the addition carried the answer: a read-only SQL tool lifted every core model's database accuracy from at most 16 to at least 88 percent. (2) Simple rules were hard to beat: no core model outperformed repeating today's staffing count, and two open-weight models mostly copied the median of similar earlier fire-days, a worse forecast. (3) Public releases carry hazards: a fire-danger column separates the holdout perfectly, and 67 aerial fire frames carry smoldering or fire-free labels read from a clipped thermal maximum. We release prompts, responses, scores, code, and the survey record.
Chinese Translation
大型语言模型(LLMs)正在进入野火管理领域,在该领域,夸大的评估可能造成财产损失和人员伤亡。它们在野火任务上表现如何,在有外部依据和没有外部依据的情况下?裸输入(bare)指模型仅接收任务输入。有依据输入(grounded)指模型还接收一项特定于任务的附加信息:对于烟雾检测,来自同一相机的一帧无烟参考帧。AI4Fire 在五个野火任务上以零样本方式,运行六个核心模型的裸输入和有依据输入版本;一次扩展扫描又增加了 29 个模型。我们对火灾任务的文献检索发现了 138 项工作;没有一项工作结合了这样的模型名单、任务覆盖范围以及成对的裸输入和有依据输入运行。我们报告三项发现。(1)在有依据输入中,帮助最大的是附加信息本身携带答案的情形:一个只读 SQL 工具将每个核心模型的数据库准确率从至多 16% 提升到至少 88%。(2)简单规则难以被击败:没有任何核心模型的表现超过直接重复今天的人员配置数量,并且两个开放权重模型大多复制了先前相似火灾日的中位数,这是一个更差的预测。(3)公开发布物带有隐患:一个火险列完美地区分了留出集,并且 67 个航空火灾帧带有从截断的热最大值读取的阴燃或无火标签。我们发布提示词、回答、分数、代码以及调查记录。
cs.CL / 38 / 2610.10971
When Citations Mislead? A Claim-Level Benchmark for Legal Hallucination Detection
当引用产生误导时?面向法律幻觉检测的断言级基准
M. Mikail Demir, M. Abdullah Canbaz
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Large language models are increasingly used in legal research and drafting, but they can still produce claims that sound convincing without being supported by the cited source. We introduce PARCEL, a benchmark for checking whether a legal claim is supported by the underlying authority. Using recent New York State Court of Appeals decisions, we build a dataset of 3,396 parenthetical-style claims labeled as Supported, Refuted, or Not Found. We cast this task as a three-way natural language inference problem and evaluate several state-of-the-art LLMs in a zero-shot setting. Although the strongest models reach up to 0.97 accuracy, the results also show an important weakness: models still incorrectly mark unsupported claims as supported, even when the full opinion text is provided. Across models, missing support is harder to detect than direct contradiction, and fabricated but plausible citations cause the largest drop in performance. Overall, PARCEL provides a practical benchmark for testing claim-level groundedness in legal RAG systems.
Chinese Translation
大型语言模型正越来越多地用于法律研究和文书起草,但它们仍可能生成听起来令人信服、却未得到所引来源支持的断言。我们提出 PARCEL,一个用于检查法律断言是否得到其所依据的权威来源支持的基准。我们利用纽约州上诉法院近期的判决,构建了一个包含 3,396 条括号式断言的数据集,这些断言被标注为 Supported、Refuted 或 Not Found。我们将该任务设定为一个三分类自然语言推理问题,并在零样本设置下评估若干最先进的大型语言模型。尽管最强的模型达到了高达 0.97 的准确率,结果也显示出一个重要弱点:即使提供了完整的判决意见文本,模型仍会错误地将未获支持的断言标记为已获支持。在不同模型之间,缺失支持比直接矛盾更难检测,而捏造但看似合理的引用会导致最大的性能下降。总体而言,PARCEL 提供了一个实用基准,用于测试法律 RAG 系统中断言级别的有据性。
cs.CL / 39 / 2610.11033
FedAlphaEdit: Null-Space-Aligned Merging for Collaborative Knowledge Editing
FedAlphaEdit:面向协作知识编辑的零空间对齐合并
Sota Sugawara, Yukihiko Okada
cs.CL · cs.LG
large language model
大语言模型相关
Abstract
Multiple institutions may each hold their own private knowledge edits and wish to integrate them into a single large language model without sharing raw edit requests. Null-space-constrained editing methods such as AlphaEdit mathematically guarantee that each update leaves unrelated knowledge intact, while collaborative frameworks such as CollabEdit aggregate edits from multiple clients without data sharing. Combining the two appears trivial. However, we show that this naive combination fails structurally, and we identify its cause. Guided by this analysis, we propose FedAlphaEdit. To our knowledge, this is the first collaborative knowledge editing framework that aligns both local editing and the server-side merging rule under a single null-space principle for preserving existing knowledge. FedAlphaEdit builds on null-space-aligned merging, in which clients share projected statistics and the server provably recovers the result of editing everything in one place under a one-shot idealization. Empirically, the proposed method repairs the collapse and brings edit success and preservation simultaneously close to the level of centralized editing across two architecture families. FedAlphaEdit thus lets institutions that cannot share raw edit data, such as hospitals and financial firms, jointly maintain a shared model that closely approximates editing all facts in one place.
Chinese Translation
多个机构可能各自持有自己的私有知识编辑,并希望在不共享原始编辑请求的情况下将它们集成到单个大语言模型中。诸如 AlphaEdit 之类的零空间约束编辑方法在数学上保证每次更新都不会改变无关知识,而诸如 CollabEdit 之类的协作框架则在无需数据共享的情况下聚合来自多个客户端的编辑。将二者结合看似平凡。然而,我们表明这种朴素结合会在结构上失败,并且我们识别出了其原因。在此分析的指导下,我们提出了 FedAlphaEdit。据我们所知,这是第一个在用于保留现有知识的单一零空间原则下同时对齐本地编辑和服务器端合并规则的协作知识编辑框架。FedAlphaEdit 建立在零空间对齐合并之上,其中客户端共享投影后的统计量,并且在一次性理想化下,服务器可证明地恢复在一处编辑所有内容的结果。经验上,所提出的方法修复了这种崩溃,并在两个架构族上使编辑成功和保留同时接近集中式编辑的水平。因此,FedAlphaEdit 使无法共享原始编辑数据的机构(例如医院和金融公司)能够共同维护一个共享模型,该模型紧密近似于在一处编辑所有事实。
cs.CL / 40 / 2610.11069
Clinician use of language models diverges from how the models are evaluated
临床医生使用语言模型的方式与模型的评估方式相背离
Krithik Vishwanath, Haitong Lin, Anton Alyakin, Jin Vivian Lee, D. Brock Hewitt, Jie J. Yao, William Robert Small, Hammad A. Khan, Cordelia Orillac, Aakaash Varma, Brandon Ye, Daniel Alexander Alber, Gustavo Stolovitzky, Batia Wiesenfeld, Oded Nov, Wei Wu, Kang Zhang, Yindalon Aphinyanaphongs, Tim Requarth, Eric Karl Oermann, The International Digital Twin Consortium in Healthcare, Medicine
cs.CL
large language model
大语言模型相关
Abstract
Large language model (LLM) assistants are being deployed to clinicians across health systems, and judgments about their readiness rest largely on benchmark scores, most of them derived from examination questions or curated cases. A benchmark predicts performance in deployment only to the extent that its items resemble real use, yet whether benchmarks reflect the work these systems receive has rarely been measured. Here we analyze 127,833 queries sent by 6,342 physicians, advanced practice providers and nurses in 35 specialties to an institutional assistant during an eight-month roll-out. We characterize each query with RCQ-Map, a clinician-validated framework grounded in taxonomies of clinical questions and of LLM evaluation, which records its task, intent, answerability, missing information and potential harm. Documentation and administration (36.2%) and knowledge retrieval (28.9%) made up nearly two-thirds of use, and diagnosis 3.7%; more than a third of queries could not be answered well as posed. Applying RCQ-Map to 58 public benchmarks drawn from major evaluation suites and frontier model reports, which we assemble into the Clinical AI Benchmark Atlas, showed that the median benchmark contained no documentation requests and shared 31% of the task mix of real use, less than an even spread across task categories would. Benchmarks in suites designed to resemble clinical practice were individually no closer to real use than those used in frontier model reports. Benchmark scores therefore say little about how clinical AI performs on most of the work it is actually given, and evaluation should be matched to real clinical use.
Chinese Translation
大型语言模型(LLM)助手正被部署给整个卫生系统中的临床医生,而对其是否准备就绪的判断在很大程度上取决于基准分数,这些分数大多来自考试题目或精选病例。基准只有在其实测项目类似于真实使用的情况下,才能预测部署中的表现,然而基准是否反映这些系统所接收的工作却很少被衡量。在此,我们分析了在为期八个月的推广期间,来自 35 个专科的 6,342 名医生、高级实践提供者和护士发送给一个机构助手的 127,833 条查询。我们使用 RCQ-Map 对每条查询进行刻画,RCQ-Map 是一个经临床医生验证的框架,其基础是临床问题分类法和 LLM 评估分类法,记录查询的任务、意图、可回答性、缺失信息和潜在危害。文书与行政管理(36.2%)和知识检索(28.9%)构成了近三分之二的使用,而诊断占 3.7%;超过三分之一的查询无法按所提出的方式得到良好回答。将 RCQ-Map 应用于 58 个公开基准——这些基准取自主要评估套件和前沿模型报告,我们将其汇编为《临床 AI 基准图集》——结果显示,中位基准不包含任何文书请求,并且与真实使用的任务组合仅共享 31%,低于任务类别均匀分布本会达到的比例。在旨在模拟临床实践的套件中,各个基准与真实使用的接近程度并不比前沿模型报告中使用的基准更高。因此,基准分数几乎无法说明临床 AI 在它实际被赋予的大部分工作上的表现,评估应当与真实临床使用相匹配。
cs.CL / 41 / 2610.11132
SFT-as-Context Mitigates Forgetting in Supervised Fine-Tuning
SFT-as-Context 缓解监督微调中的遗忘
Kenan Tang, Andong Hua, Chengxuan Qian, Saket Tiwari, Yao Qin
cs.CL · cs.AI · cs.LG
large language model
大语言模型相关
Abstract
Supervised fine-tuning (SFT) equips large language models (LLMs) with specialized capabilities, but often comes at the cost of forgetting the general capabilities of their parent models (i.e., the pretrained models before fine-tuning). This trade-off is especially limiting for queries that require both specialized and general capabilities. We introduce SFT-as-context, a training-free method in which the parent model uses the SFT model's response as context to answer the query. This allows the parent model to acquire fine-tuned capabilities from the SFT response through in-context learning while preserving its own general capabilities. Across 19 parent-SFT model pairs and 11 benchmarks, SFT-as-context remains close to the SFT models on fine-tuned capabilities, with gaps of only 2.2 and 2.1 percentage points on AIME 2024 and LiveCodeBench and 2.0 macro MAE on NutriBench-English, while staying within 2.2 percentage points of the parent models on general capabilities on average. Remarkably, it can solve queries requiring both fine-tuned and general capabilities, even when neither the parent nor SFT model succeeds alone. This approach also extends beyond parent-SFT pairs: responses from a small open-source SFT model can improve a strong closed-source LLM, outperforming either model alone. Furthermore, we use a Bayesian framework to derive theoretical guarantees that bound the error of SFT-as-context relative to the SFT model on fine-tuned capabilities and to the parent model on general capabilities. In addition, we visualize the attention weights and find that the parent model attends more to useful SFT responses and less to irrelevant ones, suggesting that selective attention helps the parent model use the SFT response through in-context learning.
Chinese Translation
监督微调(SFT)赋予大语言模型(LLM)专门化能力,但往往以遗忘其父模型(即微调前的预训练模型)的通用能力为代价。这种权衡对于同时需要专门化能力和通用能力的查询尤其具有限制性。我们提出 SFT-as-context,这是一种免训练方法,其中父模型使用 SFT 模型的响应作为上下文来回答查询。这使父模型能够通过上下文学习从 SFT 响应中获得微调后的能力,同时保留其自身的通用能力。在 19 对父模型-SFT 模型组合和 11 个基准上,SFT-as-context 在微调后能力上仍接近 SFT 模型,在 AIME 2024 和 LiveCodeBench 上的差距仅为 2.2 和 2.1 个百分点,在 NutriBench-English 上的差距仅为 2.0 macro MAE,同时在通用能力上平均与父模型的差距保持在 2.2 个百分点以内。值得注意的是,它能够解决同时需要微调后能力和通用能力的查询,即使父模型和 SFT 模型单独都无法成功。这种方法还超越了父模型-SFT 模型组合:来自一个小型开源 SFT 模型的响应可以改进一个强大的闭源 LLM,其表现优于任一单独模型。此外,我们使用贝叶斯框架推导理论保证,为 SFT-as-context 在微调后能力上相对于 SFT 模型的误差,以及在通用能力上相对于父模型的误差提供界限。另外,我们可视化了注意力权重,并发现父模型对有用的 SFT 响应给予更多关注,对无关响应给予更少关注,这表明选择性注意力有助于父模型通过上下文学习使用 SFT 响应。
cs.CL / 42 / 2610.11159
Local Prototype Reconstruction for Text-Compatible Speech-to-LLM Bridge Pretraining
面向文本兼容的语音到LLM桥接预训练的局部原型重建
Xinnian Zhao, Chia-Hua Wu, Pu Wang, Hugo Van Hamme
cs.CL · cs.SD
large language model
大语言模型相关
Abstract
Speech-to-LLM systems often connect a frozen speech encoder to a frozen large language model (LLM) through a small trainable bridge. The bridge is usually treated as plumbing, but it in fact defines the geometry of the speech-to-LLM interface, and the pretraining objective decides whether that interface provides a reusable initialization for downstream tasks. We study a transferable bridge through two complementary properties: global alignment with the text side, and local lexical manifold compatibility, where bridge embeddings remain close to the frozen LLM's input-embedding neighbourhoods. We make this property measurable with a fixed, head-free, timestamp-free diagnostic that applies to any objective, and show that next-word prediction (NWP) and sentence-level contrastive pretraining do not fully capture token-level lexical compatibility. We then introduce Local Prototype Reconstruction (LPR), a lightweight training-only regularizer that requires each aligned bridge token to be reconstructable from a small neighbourhood of frozen LLM token embeddings, with a hard single-prototype anchor as its limiting case. On multilingual ASR and speech translation, LPR improves transfer, with the largest gains on translation and low-resource adaptation. Crucially, our independent diagnostic correlates with downstream gains across objectives, suggesting that lexical manifold compatibility is predictive of reusability for speech-to-LLM bridges.
Chinese Translation
语音到LLM系统通常通过一个小型可训练的桥接模块,将冻结的语音编码器连接到冻结的大语言模型(LLM)。该桥接模块通常被视为管道式的连接件,但它实际上定义了语音到LLM接口的几何结构,而预训练目标则决定了该接口能否为下游任务提供可复用的初始化。我们通过两个互补的性质来研究可迁移的桥接模块:与文本侧的全局对齐,以及局部词汇流形兼容性,即桥接嵌入保持在冻结LLM的输入嵌入邻域附近。我们用一个固定的、无需预测头、无需时间戳的诊断方法使该性质可度量,该方法适用于任何目标函数,并表明下一词预测(NWP)和句子级对比预训练并不能完全捕捉词元级的词汇兼容性。随后我们引入局部原型重建(LPR),这是一种轻量级的、仅在训练时使用的正则化器,它要求每个已对齐的桥接词元都能从冻结LLM词元嵌入的一个小邻域中重建出来,并以硬性单原型锚点作为其极限情形。在多语言ASR和语音翻译任务上,LPR改善了迁移效果,其中在翻译和低资源适配上的增益最大。至关重要的是,我们的独立诊断指标在不同目标函数下都与下游增益相关,这表明词汇流形兼容性能够预测语音到LLM桥接模块的可复用性。
cs.CL / 43 / 2610.11305
BeliefScope: Diagnosing Evidence-Driven Revision and Pressure-Induced Shifts in Large Language Models
BeliefScope:诊断大型语言模型中的证据驱动修正与压力诱发偏移
Shuai Guo, Yidong Cui
cs.CL
large language model
大语言模型相关
Abstract
A language model may revise the same proposition after receiving genuinely relevant evidence or after receiving directional user pressure that adds no relevant fact. The observable response shift alone therefore does not reveal which source drove the change. We introduce BeliefScope, a controlled black-box framework for separating these two sources of influence around a fixed target proposition. BeliefScope crosses Evidence and Pressure with factor-specific local controls and measures response changes through probability reports, categorical judgments, and action recommendations on channel-appropriate scales. To determine when these observable contrasts support reliable attribution, we evaluate the observation design under controlled synthetic conditions. Known-truth recovery and targeted ablations establish where Evidence- and Pressure-related effects can be separated, while semi-synthetic stress tests map how that recoverability changes as the observation process becomes noisier and more heterogeneous. Across a 36-family Qwen/Llama study, with targeted 12-family checks that also include Gemma3-12B, the resulting profiles show substantial evaluation-context dependence: broad model-level differences can change under matched controls, decoding, or response interfaces, while some narrower within-model patterns remain stable. Instruction interventions further show that reduced target-aligned Pressure following can reflect either stable resistance or movement in the opposite direction. BeliefScope summarizes these measurements as a conditional belief-response profile that keeps diagnostic effects tied to the evaluation conditions under which they are observed, together with explicit validity boundaries for that diagnosis.
Chinese Translation
语言模型可能在接收到真正相关的证据之后,或在接收到不添加任何相关事实的方向性用户压力之后,修正同一个命题。因此,仅凭可观测的响应变化本身,并不能揭示是哪种来源驱动了这一改变。我们提出 BeliefScope,一个受控的黑盒框架,用于围绕一个固定的目标命题分离这两种影响来源。BeliefScope 将 Evidence 与 Pressure 同因子特定的局部控制进行交叉,并通过在适合各通道的量表上的概率报告、类别判断和行动建议来测量响应变化。为确定这些可观测对比何时支持可靠的归因,我们在受控的合成条件下评估该观测设计。已知真值恢复与针对性消融确立了与 Evidence 和 Pressure 相关的效应在何处可以被分离,而半合成压力测试则刻画了随着观测过程变得更具噪声、更加异质,这种可恢复性如何变化。在一项涵盖 36 个家族的 Qwen/Llama 研究中,并辅以也包含 Gemma3-12B 的 12 个家族的针对性核查,所得到的画像显示出显著的评估情境依赖性:宽泛的模型层面差异在匹配控制、解码或响应接口下可能发生变化,而某些更窄的模型内模式则保持稳定。指令干预进一步表明,目标对齐的压力顺从度下降既可能反映稳定的抵抗,也可能反映朝相反方向的移动。BeliefScope 将这些测量结果总结为一个条件化的信念–响应画像,使诊断效应始终与其被观测时所处的评估条件相绑定,并同时给出该诊断的明确有效性边界。
cs.CL / 44 / 2610.11314
From Retrieval to Reconstruction: Constructing Evolvable Cognitive Memory for Long-Term Dialogue
从检索到重构:为长期对话构建可演化的认知记忆
Zirui Liao, Zhengxian Wu, Zhuohong Chen, Yunyao Yu, Xiaoyu Liu, Yifan Xu, Haoqian Wang
cs.CL
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) serving as long-term dialogue agents require memory systems that support reliable reasoning over extended interactions. However, existing Retrieval-Augmented Generation (RAG) frameworks typically treat memory as passive storage, making it difficult to distinguish source-attributed beliefs from unattributed event/fact records and to connect evidence dispersed across sessions. We introduce CogMem, a cognitive memory architecture based on the PEC$^2$F (Person-Event-Concept-Claim-Fact) graph schema. Dedicated Claim nodes preserve the source and target of subjective statements, while Fact and Event nodes represent semantic and episodic knowledge. Dialogue turns are incrementally converted into provenance-aware graph records, consolidated into higher-level facts, and reconciled into temporally scoped Claim views when the same source provides conflicting updates. For retrieval, a rule-based controller driven by LLM intent parsing composes four deterministic graph operators---anchoring, traversal, intersection, and evidence grounding---to reconstruct query-relevant context. Experiments on LoCoMo and LongMemEval show strong performance, especially on multi-hop, temporal, and knowledge-update tasks. Ablations and a semantic-collapse probe support complementary contributions from epistemic separation, consolidation, and agentic retrieval. Code: https://github.com/Silent-Rain02/CogMem.
Chinese Translation
作为长期对话智能体的大语言模型(LLMs)需要能够支持在扩展交互中进行可靠推理的记忆系统。然而,现有的检索增强生成(RAG)框架通常将记忆视为被动存储,这使得难以区分带有来源归属的信念与无归属的事件/事实记录,也难以连接分散在不同会话中的证据。我们提出 CogMem,一种基于 PEC$^2$F(Person-Event-Concept-Claim-Fact)图模式的认知记忆架构。专用的 Claim 节点保留主观陈述的来源与目标,而 Fact 与 Event 节点则表示语义知识与情景知识。对话轮次被增量地转换为具有来源溯源信息的图记录,被整合为更高层级的事实,并在同一来源提供冲突更新时被调和为具有时间范围的 Claim 视图。在检索方面,一个由 LLM 意图解析驱动的基于规则的控制器组合四种确定性图算子——锚定、遍历、求交与证据落地——以重构与查询相关的上下文。在 LoCoMo 与 LongMemEval 上的实验显示出强劲的性能,尤其是在多跳、时序与知识更新任务上。消融实验与一个语义坍缩探针支持了认知分离、整合与智能体式检索的互补贡献。代码:https://github.com/Silent-Rain02/CogMem。
cs.CL / 45 / 2610.11351
Deception by Omission: Language Models Knowingly Hide Their Mistakes
通过省略进行欺骗:语言模型明知故犯地隐藏其错误
Lucas Florin, Amelie Knecht, Ulysse Schaller, Thilo Hagendorff
cs.CL
large language model
大语言模型相关
Abstract
Large language models (LLMs) increasingly act as agents with little human oversight, so potential mistakes they make can go unnoticed. Users then depend on the model to report what went wrong. An honest model discloses its mistakes, while a deceptive one conceals them. However, it is unclear how current LLMs behave in such situations. In this study, we prefill LLM trajectories with synthetic mistakes. The trajectories resemble real deployments in chat and agentic settings. Models fail to disclose their mistake in 36.4% of chat and 67.1% of agentic rollouts. In 2.4% and 5.3% of rollouts, respectively, they are aware of the mistake in their chain of thought but still deceptively conceal it. Rates vary by model: for instance, Gemini 3.5 Flash knowingly conceals mistakes in up to 19.9% of agentic rollouts. In 11.9% of chat and 51.8% of agentic rollouts, models show no awareness of mistakes, even though they reliably spot them when reviewing the same transcript as an outside observer. Our results show that, as agents take on more tasks with less oversight, users cannot rely on them to self-report possible mistakes. Developers should instead use independent monitors that review agent trajectories, or specifically train models to check their past actions and disclose what they find.
Chinese Translation
大型语言模型(LLM)日益作为智能体行动,而人类监督很少,因此它们可能犯下的潜在错误可能不被察觉。用户随后依赖模型报告出了什么问题。诚实的模型会披露自己的错误,而欺骗性的模型则会隐瞒错误。然而,当前 LLM 在这类情况下会如何表现尚不清楚。在本研究中,我们向 LLM 轨迹中预填充合成错误。这些轨迹类似于聊天和智能体环境中的真实部署。模型在 36.4% 的聊天运行轨迹和 67.1% 的智能体运行轨迹中未能披露其错误。在分别 2.4% 和 5.3% 的运行轨迹中,它们在思维链中意识到该错误,但仍欺骗性地隐瞒了它。比率因模型而异:例如,Gemini 3.5 Flash 在最多 19.9% 的智能体运行轨迹中故意隐瞒错误。在 11.9% 的聊天运行轨迹和 51.8% 的智能体运行轨迹中,模型没有表现出对错误的意识,尽管当它们作为外部观察者审阅同一份记录时能够可靠地发现这些错误。我们的结果表明,随着智能体承担更多任务而受到的监督更少,用户不能依赖它们自我报告可能的错误。开发者应转而使用审查智能体轨迹的独立监控器,或者专门训练模型检查其过往行为并披露其发现。
cs.CL / 46 / 2610.11371
SignRAG: Unified Retrieval-Augmented Gloss-Free Sign Language Translation
SignRAG:统一的检索增强无gloss手语翻译
Zhi Rao, Yucheng Zhou, Qianran Sun, Yiqing Huang, Longcan Yuan, Jiayi Hou, Chengwen Yao, Lin Cheng, Donghui Sun, Xiaoxin Chen, Jun Wan
cs.CL · cs.CV · cs.MM
large language model
大语言模型相关
Abstract
Contemporary decoder-only large language models (LLMs) have demonstrated strong capabilities across a wide range of domains. However, existing pretraining paradigms for gloss-free sign language translation (SLT) are largely designed around conventional encoder-decoder pretrained language models, which limits their direct applicability to decoder-only LLMs. To address this limitation, we propose SignRAG, a unified framework combining hierarchical pretraining, target-domain retrieval augmentation, and retrieval-aware reinforcement fine-tuning. Hierarchical pretraining first learns linguistically grounded sign representations and then jointly aligns the sign encoder with an LLM, mitigating cross-modal optimization imbalance. For downstream adaptation, SignRAG complements parameter-based fine-tuning with a target-domain retrieval gallery that provides instance-specific translation cues. To ensure that retrieved contexts are used appropriately, we further introduce Retrieval Utility-Guided Reinforcement Fine-Tuning (RUG-RFT), which combines translation-quality and retrieval-utility rewards to encourage beneficial retrieval use while suppressing harmful reliance. Experiments on multiple SLT benchmarks establish new state-of-the-art performance. In particular, to the best of our knowledge, SignRAG is the first gloss-free approach to outperform gloss-supervised methods across all reported metrics on CSL-Daily. Our code has been released at \href{https://github.com/shahelaojieraozhi/SignRAG}{GitHub}, together with models of different sizes to support future academic research.
Chinese Translation
当代仅解码器(decoder-only)大语言模型(LLM)已在广泛的领域中展现出强大的能力。然而,现有的无gloss手语翻译(SLT)预训练范式大多是围绕传统的编码器-解码器预训练语言模型设计的,这限制了它们对仅解码器LLM的直接适用性。为解决这一局限,我们提出SignRAG,一个结合了分层预训练、目标域检索增强与检索感知强化微调的统一框架。分层预训练首先学习具有语言学依据的手语表征,然后将手语编码器与LLM联合对齐,从而缓解跨模态优化不平衡。在下游适配方面,SignRAG通过一个目标域检索库为基于参数的微调提供补充,该检索库能够提供针对具体实例的翻译线索。为确保检索到的上下文被恰当使用,我们进一步提出检索效用引导的强化微调(RUG-RFT),它结合翻译质量奖励与检索效用奖励,以鼓励有益的检索使用,同时抑制有害的依赖。在多个SLT基准上的实验确立了新的最先进性能。特别地,据我们所知,SignRAG是首个在CSL-Daily的所有报告指标上均超越gloss监督方法的无gloss方法。我们的代码已发布在 \href{https://github.com/shahelaojieraozhi/SignRAG}{GitHub},并同时发布了不同规模的模型,以支持未来的学术研究。
cs.CL / 47 / 2610.11430
BioBigBird: A Sparse Attention Model for Long-Range Dependency Processing in Biomedical Text
BioBigBird:一种用于生物医学文本长距离依赖处理的稀疏注意力模型
Roshan Balaji, Pavan Kumar S, Vasudev Gupta, Sreejith N, Keerthana Sridhar, Nirav Bhatt
cs.CL · cs.LG
large language model
大语言模型相关
Abstract
While domain-specific Large Language Models (LLMs) have encoded vast biomedical knowledge, their limited context windows often hinder a deep understanding of nuanced relationships within and across texts. To address this limitation, we introduce BioBigBird, a bidirectional language model pre-trained on extensive biomedical literature and clinical data, specifically designed to handle long-range dependencies. BioBigBird leverages a sparse attention mechanism to process sequences up to 4096 tokens, and its training incorporates a multi-stage process to mitigate noise from the large-scale pre-training corpus. We further enhance its performance by employing a multi-task learning (MTL) framework that jointly optimizes for Named Entity Recognition and Relation Extraction. Comprehensive evaluations on the BLURB benchmark reveal that our MTL-enhanced BioBigBird achieves highly competitive results against state-of-the-art models. Our work contributes an effective methodology for developing powerful, long-context language models for specialized domains, demonstrating the value of extended sequence processing for complex text analysis. Our models are publicly available at https://huggingface.co/collections/bisectgroup/biobigbird.
Chinese Translation
尽管领域专用的大语言模型(LLMs)已经编码了海量的生物医学知识,但其有限的上下文窗口往往阻碍了对文本内部及跨文本细微关系的深入理解。为解决这一局限,我们提出了 BioBigBird,一个在大量生物医学文献和临床数据上预训练的双向语言模型,专门设计用于处理长距离依赖。BioBigBird 利用稀疏注意力机制来处理长达 4096 个 token 的序列,并且其训练采用多阶段过程,以减轻来自大规模预训练语料的噪声。我们进一步通过采用多任务学习(MTL)框架来联合优化命名实体识别和关系抽取,从而提升其性能。在 BLURB 基准上的全面评估表明,我们经过 MTL 增强的 BioBigBird 相较于最先进的模型取得了极具竞争力的结果。我们的工作为开发面向专业领域的强大长上下文语言模型提供了一种有效的方法论,证明了扩展序列处理对于复杂文本分析的价值。我们的模型已在 https://huggingface.co/collections/bisectgroup/biobigbird 公开提供。
cs.CL / 48 / 2610.11501
Beyond Sequences: Distilling Structured Decision Memory for LLM Recommendation
超越序列:为LLM推荐蒸馏结构化决策记忆
Leikun Liang, Guoshuai Wang, Xingsheng He, Yushan Han, Yunyi Xuan, Xiaoxiao Xu, Lin Qu
cs.CL
large language model
大语言模型相关
Abstract
Despite the adoption of large language models (LLMs) in recommendation systems, prevailing approaches mostly model single-type behaviors (e.g., views or purchases). Even when incorporating multiple behaviors, existing methods flatten heterogeneous actions into homogeneous token sequences, ignoring their distinct decision-making roles. This flattening fails to capture semantic hierarchies and contextual nuances in complex decision-making, such as trade-offs between price and quality. Consequently, performance degrades in critical ``difficult-choice'' scenarios involving highly similar items. To bridge this gap, we propose MARI (Memory-Augmented Recommendation with Interpretability), which grounds predictions in explicit, structured decision evidence. MARI maintains a Decision Memory Bank (DMB) that archives users' past rationales as Structured Decision Memories (SDMs): concise records of goals, constraints, and trade-offs. These SDMs are generated offline via Post-Hoc Decision Distillation from heterogeneous behaviors and user-generated content. By retrieving relevant SDMs to augment LLM reasoning, MARI achieves interpretability and scalability without the prohibitive cost of processing long raw sequences. Extensive experiments show MARI significantly outperforms state-of-the-art baselines on standard next-item prediction and a newly introduced Difficult Choice Prediction task, incurring low latency overhead by decoupling memory construction from online inference. Qualitative analyses reveal actionable, human-readable insights into user decision-making, marking a concrete step toward reasoning-aware recommendation systems.
Chinese Translation
尽管大语言模型(LLM)已在推荐系统中得到采用,但主流方法大多对单一类型的行为(例如浏览或购买)进行建模。即便在纳入多种行为时,现有方法也将异质动作展平为同质的词元序列,忽视了它们各自不同的决策作用。这种展平方式无法捕捉复杂决策中的语义层级与语境细微差别,例如价格与质量之间的权衡。因此,在涉及高度相似物品的关键“困难选择”场景中,性能会下降。为弥合这一差距,我们提出 MARI(Memory-Augmented Recommendation with Interpretability,记忆增强的可解释推荐),它将预测建立在显式、结构化的决策证据之上。MARI 维护一个决策记忆库(Decision Memory Bank, DMB),将用户过去的决策依据归档为结构化决策记忆(Structured Decision Memories, SDMs):即对目标、约束与权衡的简洁记录。这些 SDM 通过事后决策蒸馏(Post-Hoc Decision Distillation)从异质行为与用户生成内容中离线生成。通过检索相关 SDM 来增强 LLM 推理,MARI 在无需承担处理长原始序列的高昂成本的情况下,实现了可解释性与可扩展性。大量实验表明,MARI 在标准的下一个物品预测以及新引入的困难选择预测(Difficult Choice Prediction)任务上显著优于最先进的基线方法,并通过将记忆构建与在线推理解耦,仅带来较低的延迟开销。定性分析揭示了关于用户决策的可操作、人类可读的洞见,标志着朝着具备推理意识的推荐系统迈出了切实的一步。
cs.CL / 49 / 2610.11510
Does Modern Standard Arabic (MSA) Dominate Arabic Dialects in LLMs? A Representation-Level Analysis
现代标准阿拉伯语(MSA)是否在LLM中支配阿拉伯语方言?一项表征层面的分析
Abdu Sallouh, Nicholas Popovič, Michael Färber
cs.CL
large language model
大语言模型相关
Abstract
Large language models (LLMs) often default to Modern Standard Arabic (MSA) when generating Arabic, even when prompted with dialectal Arabic. A natural explanation is that their internal representations are dominated by MSA. We test this hypothesis by adapting the language-dominance framework of Shani and Basirat (2025) (https://doi.org/10.18653/v1/2025.blackboxnlp-1.7) to 26 Arabic varieties. Across layers and model families, we find no evidence that MSA acts as a dominant internal representation for Arabic dialects. Instead, dialect representations form a dense and highly overlapping space: normalized mutual information drops sharply compared to patterns reported for more distinct languages. Moreover, the strongest separability effects are not confined to intermediate layers, but can shift toward later layers depending on the architecture. These findings challenge a common interpretation of MSA-biased generation: output preference does not necessarily reveal internal representational dominance. Analyses of multilingual and dialectal LLMs should therefore distinguish generation bias from the geometry of internal representations.
Chinese Translation
大型语言模型(LLM)在生成阿拉伯语时,常常默认使用现代标准阿拉伯语(MSA),即使被提示的是方言阿拉伯语。一种自然的解释是,它们的内部表征被MSA主导。我们通过将Shani和Basirat(2025)(https://doi.org/10.18653/v1/2025.blackboxnlp-1.7)的语言主导性框架适配到26种阿拉伯语变体来检验这一假设。跨层和模型家族,我们没有发现证据表明MSA充当阿拉伯语方言的主导内部表征。相反,方言表征形成了一个稠密且高度重叠的空间:与针对差异更明显的语言所报告的模式相比,归一化互信息急剧下降。此外,最强的可分性效应并不局限于中间层,而是可能取决于架构而向更靠后的层偏移。这些发现挑战了对偏向MSA的生成的一种常见解释:输出偏好并不必然揭示内部表征的主导性。因此,对多语言和方言LLM的分析应当将生成偏差与内部表征的几何结构区分开来。
cs.CL / 50 / 2610.11585
Adapting English Quality Classifiers for Multilingual LLM Pretraining Data Selection
适配英语质量分类器以用于多语言 LLM 预训练数据选择
Vinko Sabolčec, Bettina Messmer, Yassine Turki, Martin Jaggi
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Recent advances in large language model (LLM) pretraining highlight the role of high-quality training data in improving performance. While model-based filtering has proven effective in selecting high-quality subsets from web-scale corpora, especially for high-resource languages, low-resource languages face challenges due to limited availability of annotated data. This work explores extending quality filtering to over 100 languages by proposing a multilingual adaptation approach that converts an existing English quality classifier into a multilingual variant. Our approach proposes training a small multi-layer perceptron on top of Transformer encoder-only model embeddings, using multilingual text as input and scores obtained from English classifiers applied to machine-translated text as labels. Our 1B, 3B and 8B scale experiments show that our approach maintains the downstream LLM benchmark performance of existing multilingual model-based filtering baselines, without harming regional and cultural knowledge benchmarks. To further evaluate cross-lingual generalization, we compare classifier scores of high-quality synthetic data and web samples, and the correlation of classifier scores with LLM-based ones, revealing that the classifier can learn the scoring criteria of its original English variant, even for languages not included in its training data.
Chinese Translation
最近在大语言模型(LLM)预训练方面的进展凸显了高质量训练数据在提升性能中的作用。虽然基于模型的过滤已被证明在从网络规模语料库中选择高质量子集方面有效,尤其对高资源语言而言,但低资源语言由于标注数据可用性有限而面临挑战。这项工作探索将质量过滤扩展到 100 多种语言,通过提出一种多语言适配方法,将现有的英语质量分类器转换为多语言变体。我们的方法提出在仅编码器 Transformer 模型嵌入之上训练一个小型多层感知机,使用多语言文本作为输入,并使用由应用于机器翻译文本的英语分类器获得的分数作为标签。我们的 1B、3B 和 8B 规模实验表明,我们的方法在保持现有基于多语言模型的过滤基线的下游 LLM 基准性能的同时,没有损害区域和文化知识基准。为了进一步评估跨语言泛化能力,我们比较高质量合成数据和网络样本的分类器分数,以及分类器分数与基于 LLM 的分数之间的相关性,揭示出该分类器能够学习其原始英语变体的评分标准,即使对于其训练数据中未包含的语言也是如此。
cs.CL / 51 / 2610.11599
Large Language Model Turnover Undermines Screening for Artificial Intelligence-Assisted Scientific Writing
大型语言模型更替削弱了对人工智能辅助科学写作的筛查
Kazuki Nakajima, Takayuki Mizuno
cs.CL · cs.AI · cs.CY · cs.DL · cs.SI
large language model
大语言模型相关
Abstract
Journals and conferences have begun to screen submitted manuscripts for text written using large language models (LLMs). The reliability of this screening rests on benchmark evaluations against a fixed set of LLM versions, while the versions in actual use keep changing. Here we quantify how this LLM turnover affects the screening of scientific manuscripts. We paired 4,000 pre-ChatGPT abstracts from the Proceedings of the National Academy of Sciences with their rewrites by 23 LLM versions from three vendors, released between June 2023 and August 2026. We then trained detectors under maintenance scenarios ranging from a detector retrained on every new version to one trained once and never updated. Detectors trained only on a vendor's past versions can collapse at the boundaries between model generations: calibrated to falsely flag 1% of human-written abstracts, they catch above 99% of rewrites just before the sharpest boundary and 3.8% just after it. Detectors trained on later versions can also miss rewrites of earlier ones. Vocabulary differences between versions largely track where detection transfers and where it fails. In the two screening scenarios we simulated, screens covering all 23 versions either flagged one in eight human-written abstracts or missed one in three rewrites of the newest version. Indeed, a commercial detector missed most rewrites of the version just after the sharpest boundary while flagging almost no human-written abstracts. Research-integrity policy should therefore treat the benchmark accuracy of a detector as provisional, to be re-verified with every LLM release, including earlier versions.
Chinese Translation
期刊和会议已开始筛查提交的稿件中由大型语言模型(LLM)撰写的文本。这种筛查的可靠性依赖于针对固定的一组LLM版本的基准评估,而实际使用的版本却不断变化。在此,我们量化这种LLM更替如何影响科学稿件的筛查。我们将来自《美国国家科学院院刊》的4,000篇ChatGPT之前的摘要与由三家供应商的23个LLM版本(于2023年6月至2026年8月期间发布)对其进行的改写配对。然后,我们在维护情景下训练检测器,这些情景范围从在每个新版本上重新训练检测器,到训练一次后永不更新。仅基于某一供应商过去版本训练的检测器可能会在模型代际边界处崩溃:在校准为将1%的人类撰写摘要错误标记的情况下,它们在最急剧边界之前能捕获超过99%的改写,而在边界之后仅能捕获3.8%。基于较晚版本训练的检测器也可能漏掉较早版本的改写。版本之间的词汇差异在很大程度上对应着检测在何处能够迁移、在何处失效。在我们模拟的两种筛查情景中,覆盖所有23个版本的筛查要么标记出八分之一的人类撰写摘要,要么漏掉最新版本三分之一的改写。事实上,一个商业检测器漏掉了最急剧边界之后紧接着的那个版本的大多数改写,同时几乎没有标记任何人类撰写摘要。因此,科研诚信政策应将检测器的基准准确率视为暂定值,需要在每次LLM发布时重新验证,包括较早版本。
cs.CL / 52 / 2610.11678
TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation
TRACE:诊断智能体评估中的验证器脆弱性
Radhika Gaonkar
cs.CL
large language model
大语言模型相关
Abstract
Verifier scores now serve as both benchmark metrics and training rewards for large language model (LLM) agents, and a change in score is routinely read as a change in capability. It may instead reflect a change in the evaluation. We introduce TRACE, a protocol that turns a score change from a verdict into a testable diagnosis: it applies a targeted change to one part of an evaluation, compares paired runs, checks whether the agent's behavior changed, and rescores unchanged trajectories to test whether the scoring rule is responsible. In a controlled suite of 25 synthetic tasks, renaming tools lowers a scripted agent's score by 0.250 even though it performs exactly the same operations; restoring the original names at scoring time closes the entire gap, while the same mutation exposes a genuine behavioral failure in a second agent. On public $τ^2$-bench tasks with four LLM agents, an initial 30-task study finds mixed reward changes whose one clear effect does not replicate. In a larger follow-up on 88 new tasks with repeated runs per condition, renaming tools or reformatting tool outputs leaves reward unchanged to within $\pm$0.10 for seven of eight agent-change pairs, whereas tool names that deliberately mislead lower every agent's reward by 0.20-0.44, showing that the setup can detect real effects. Identical reruns flip 15-36% of task outcomes, so single-run comparisons cannot separate presentation effects from run-to-run variation. Two frontier LLM judges give consistent verdicts when a fixed trajectory is presented differently, yet disagree with each other on 57% of the same records, largely because one grades procedure rather than outcome. TRACE thus separates what a score change says about the agent from what it says about the measurement.
Chinese Translation
验证器分数如今既作为基准指标,也作为大型语言模型(LLM)智能体的训练奖励,而分数的变化通常被解读为能力的变化。但它也可能反而反映的是评估本身的变化。我们提出 TRACE,一种将分数变化从裁决转变为可检验诊断的协议:它对评估的某一部分施加有针对性的改变,比较配对运行,检查智能体的行为是否发生变化,并对未改变的轨迹重新评分,以检验评分规则是否是原因。在一个包含 25 个合成任务的受控套件中,重命名工具会使一个脚本化智能体的分数降低 0.250,尽管它执行的操作完全相同;在评分时恢复原始名称可弥合整个差距,而同样的变异在第二个智能体上暴露出真实的行为失败。在带有四个 LLM 智能体的公开 $τ^2$-bench 任务上,一项最初的 30 任务研究发现混合的奖励变化,其中一个明确的效应未能复现。在一项更大规模的后续研究中,针对 88 个新任务并在每种条件下重复运行,重命名工具或重新格式化工具输出使八个智能体-变化对中的七个的奖励保持在 $\pm$0.10 以内的不变范围内,而故意误导的工具名称则使每个智能体的奖励降低 0.20-0.44,表明该设置能够检测到真实效应。完全相同的重复运行会使 15-36% 的任务结果发生翻转,因此单次运行比较无法将呈现效应与运行间变异区分开来。两个前沿 LLM 评判者在固定轨迹以不同方式呈现时给出一致的裁决,但在同一批记录中有 57% 彼此不一致,这主要是因为其中一个评判的是过程而非结果。因此,TRACE 将分数变化所说明的关于智能体的信息与它所说明的关于测量的信息区分开来。
cs.CL / 53 / 2610.11765
Thinking Inertia: LLMs Keep Thinking When Told Not To
思维惯性:LLMs 在被要求不要思考时仍继续思考
Dianqiao Lei, Kevin Qinghong Lin, Pan Lu, Philip Torr, James Zou
cs.CL
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) increasingly ship with explicit "thinking modes", yet their counterpart, "no-thinking", has received far less attention. We study LLMs' no-thinking behavior along two axes. a. How to measure no-thinking? Prior work typically defines no-thinking through proxies such as a disabled thinking mode or the absence of long traces. These proxies are unreliable: disabled thinking modes may still emit reasoning, while long traces may contain filler rather than genuine inference. We instead normalize each response into a pre-answer trace and final answer, and evaluate it at three levels: (i) Empty-Thinking Rate for strict answer-only compliance; (ii) instruction-aware Question-Pre-answer Relevance for similarity between the question and pre-answer trace; and (iii) LLM-as-judge Explicit Inference Rate for visible explicit inference. Together, these metrics distinguish answer-only output, relevant but non-inferential text, and explicit inference. b. How does no-thinking vary across tasks and models? We evaluate six prompting interventions on six LLMs across Boolean, multiple-choice, and open-ended questions. We find that explicit no-think controls cannot reliably eliminate visible inference. Models instead exhibit "Thinking Inertia": explicit inference persists even under strict controls and becomes more prevalent as the answer space opens. Accuracy remains stable on Boolean and multiple-choice tasks, whereas open-ended tasks reveal a trade-off between answer-only compliance and task accuracy. Rewriting the same questions across answer spaces shows that supplying candidate answers makes answer-only responses easier to produce. These findings establish no-thinking as a non-trivial capability: stopping explicit reasoning cannot be assumed from model settings or instructions alone and deserves systematic evaluation alongside reasoning ability.
Chinese Translation
大型语言模型(LLMs)越来越多地随附显式的“思考模式”,但其对应物“不思考”所受到的关注则少得多。我们沿两个轴研究 LLMs 的不思考行为。a. 如何衡量不思考?先前工作通常通过代理指标来定义不思考,例如禁用的思考模式或缺乏长轨迹。这些代理指标并不可靠:禁用的思考模式可能仍会输出推理,而长轨迹可能包含填充内容而非真正的推断。我们改为将每个回答归一化为答案前轨迹和最终答案,并在三个层面进行评估:(i)空思考率(Empty-Thinking Rate),用于衡量严格的仅答案符合度;(ii)指令感知的问题-答案前相关性(Question-Pre-answer Relevance),用于衡量问题与答案前轨迹之间的相似度;以及(iii)LLM 作为评判者的显式推理率(Explicit Inference Rate),用于衡量可见的显式推理。这些指标共同区分仅答案输出、相关但非推理的文本以及显式推理。b. 不思考在不同任务和模型之间如何变化?我们在布尔、多项选择和开放式问题上,对六个 LLMs 评估了六种提示干预。我们发现,显式的不思考控制无法可靠地消除可见推理。模型反而表现出“思维惯性”:显式推理即使在严格控制下也持续存在,并随着答案空间的开放而变得更加普遍。在布尔和多项选择任务上准确率保持稳定,而开放式任务则揭示出仅答案符合度与任务准确率之间的权衡。在不同答案空间上重写相同问题表明,提供候选答案会使仅答案回答更容易产生。这些发现确立了不思考是一种非平凡能力:仅凭模型设置或指令不能假定显式推理会停止,它值得与推理能力一起被系统评估。
cs.CL / 54 / 2610.11766
Same Outcome, Different Evidence: Intent Recovery in LLM Safety Evaluation
相同结果,不同证据:LLM 安全评估中的意图恢复
Haitong Jiang, Chunlin Liu, Sihan Tang, Chan Wu, Xiaoqing Su, Yuhong Feng
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Safety evaluations of large language models commonly summarize harmful-output behavior with attack success rate (ASR). Yet the same non-harmful outcome can arise for very different reasons. A model may recover a harmful task and refuse it, fail to recover the task, or respond to something else entirely. Distinguishing these cases becomes especially important under intent-obscuring prompts, where a low ASR does not reveal whether the evaluated task was actually engaged. To make this distinction explicit, we pair ASR with operative understanding rate (UR), which measures whether a response both identifies the evaluated task and treats it as the task to be answered. Across interfaces, this paired view reveals substantial variation hidden by ASR: similar ASR values can correspond to sharply different recovery rates. Controlled English reconstructions show that recovery consistently improves as compressed prompts become more explicit, whereas ASR does not follow the same pattern. A complementary contrast comes from FormalLogic, where high recovery can still coincide with frequent harmful assistance. Together, these results show that non-harmful outcomes are not equally informative about model safety, motivating the joint reporting of intent recovery and ASR in LLM safety evaluation. Code and experiment inputs are available at https://github.com/kevinjiang0121-cyber/IRIS.
Chinese Translation
大语言模型的安全性评估通常用攻击成功率(ASR)来总结有害输出行为。然而,同样的非有害结果可能由非常不同的原因产生。模型可能恢复出一个有害任务并拒绝它,可能未能恢复该任务,或者完全是在回应其他东西。在遮蔽意图的提示下,区分这些情况变得尤为重要,因为低 ASR 并不能揭示被评估的任务是否实际上被处理。为了明确这一区分,我们将 ASR 与操作性理解率(UR)配对,后者衡量一个回答是否既识别了被评估的任务,又将其视为需要回答的任务。跨接口来看,这种配对视角揭示了被 ASR 所隐藏的显著差异:相似的 ASR 值可能对应截然不同的恢复率。受控英语重构显示,随着压缩提示变得更明确,恢复率持续提高,而 ASR 并不遵循同样的模式。一个互补的对比来自 FormalLogic,在那里高恢复率仍可能与频繁的有害协助同时出现。总之,这些结果表明,非有害结果对于模型安全性的信息量并不相同,从而促使在 LLM 安全评估中联合报告意图恢复与 ASR。代码和实验输入可在 https://github.com/kevinjiang0121-cyber/IRIS 获取。
cs.CL / 55 / 2610.11845
Detecting Spin in Clinical Trials with Large Language Models
使用大语言模型检测临床试验中的偏向性表述
Tjaš Ajdovec, Marko Robnik-Šikonja, Simon Šuster
cs.CL
large language model
大语言模型相关
Abstract
Spin in clinical trials includes reporting practices that distort the presentation of results. This is particularly critical in medicine, where spin is present in more than 50% of randomized controlled trials that fail to reach statistical significance. The comparison of primary and reported outcomes is crucial for detecting several types of spin, including outcome switching. We used 300 pairs of outcomes labeled with semantic similarity to develop a system for automatic detection of outcome switching. We evaluated baseline text similarity models and open-source LLMs using generated similarity scores and the Youden index to determine the classification threshold. The proposed approach involves prompt engineering, classification based on token probabilities, and majority voting for the final decision. The results on the test set of 2,496 examples with an F1 score of 0.78 and an accuracy of 0.90 outperform baseline text similarity models but trail behind fine-tuned versions of BERT. We used LLMs to generate natural language explanations for the classified instances and manually assessed their quality.
Chinese Translation
临床试验中的偏向性表述包括歪曲结果呈现的报告做法。这在医学中尤为关键,因为在未能达到统计学显著性的随机对照试验中,超过50%存在偏向性表述。主要结局与报告结局的比较对于检测若干类型的偏向性表述(包括结局指标替换)至关重要。我们使用了300对标注有语义相似度的结局,开发了一个自动检测结局指标替换的系统。我们评估了基线文本相似度模型和开源大语言模型,使用生成的相似度分数和约登指数来确定分类阈值。所提出的方法涉及提示工程、基于token概率的分类,以及用于最终决策的多数投票。在包含2,496个示例的测试集上,F1分数为0.78且准确率为0.90的结果优于基线文本相似度模型,但落后于BERT的微调版本。我们使用大语言模型为已分类实例生成自然语言解释,并人工评估了它们的质量。
cs.CL / 56 / 2610.11899
Forms of LLM-Integrated Applications from LLM-Chats to Autonomous AI Agent System
LLM 集成应用的形式:从 LLM 聊天到自主 AI 智能体系统
Irene Weber
cs.CL · cs.AI · cs.SE
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly embedded as components in software systems, marketed under labels such as chatbot, copilot, retrieval-augmented generation, workflow, coding agent and AI agent. Whether these labels denote genuine architectural forms or serve as branding has not been assessed systematically. In the sources surveyed, labels do carry architectural content, most clearly in vendor usage: copilot denotes a router-worker architecture operating a host application under step-by-step user confirmation, while the more recent shift to the label agent coincides with AI-planned multi-step execution of which the user sees only the outcome. The coding agents of four major providers share one architecture, a reason-and-act loop delegating to subagents. This survey describes seven recurring forms---LLM chats, custom agents, retrieval-augmented generation (RAG), AI-enhanced workflows, copilots, coding agents, and, in part, agentic RAG---in a common vocabulary of agents and tools. Each is characterized along four structural dimensions (agentic RAG only partially): the architectural pattern, the control of execution and the point of user intervention, the number of agent calls per task, and tool use. An illustrative corpus of 22 systems from research publications and vendor documentation grounds the descriptions and shows where they reach their limit.
Chinese Translation
大语言模型(LLM)正日益作为组件被嵌入软件系统之中,并以聊天机器人、副驾驶(copilot)、检索增强生成、工作流、编码智能体和 AI 智能体等标签进行营销。这些标签究竟指代真实的架构形式,还是仅作为品牌宣传,尚未得到系统性的评估。在所调查的来源中,标签确实承载着架构层面的内容,这一点在厂商的用法中最为明显:copilot 指代一种在用户逐步确认下操作宿主应用的路由器—工作者(router-worker)架构,而较晚近向 agent 这一标签的转变,则与由 AI 规划的多步执行同时出现,在此过程中用户只能看到最终结果。四家主要提供商的编码智能体共享同一种架构,即一个委派给子智能体的推理—行动(reason-and-act)循环。本综述以智能体与工具的通用词汇,描述了七种反复出现的形式——LLM 聊天、自定义智能体、检索增强生成(RAG)、AI 增强的工作流、副驾驶(copilot)、编码智能体,以及部分意义上的智能体化 RAG——。每一种形式都沿四个结构性维度加以刻画(智能体化 RAG 仅部分涉及):架构模式、执行控制与用户介入点、每个任务的智能体调用次数,以及工具使用。一个由来自研究出版物与厂商文档的 22 个系统组成的示例性语料库为这些描述提供了依据,并显示出它们在何处达到自身的极限。
cs.CL / 57 / 2610.11901
Can Decision Models Understand Stance? Evaluating Jev Against General-Purpose LLMs
决策模型能理解立场吗?评估 Jev 相对于通用大型语言模型的表现
Xing Li, Jinzhong Ning, Yijia Zhang, Liang Yang, Hongfei Lin
cs.CL
large language model
大语言模型相关
Abstract
Stance detection requires identifying an author's attitude toward a given target, sometimes based on conversational context. Jev, a specialized decision model designed for structured decision-making, offers an alternative to general-purpose large language models (LLMs). In this work, we evaluate Jev on two stance detection datasets, VAST (English texts) and ZS-CSD (Chinese conversations), comparing it with four general-purpose LLMs and two fine-tuned models. Results show that Jev achieves competitive performance on VAST, matching GPT-5.6 and outperforming the other general-purpose LLMs. However, it falls behind stronger LLMs on ZS-CSD, particularly in distinguishing favor from against. Further analysis suggests that this limitation may be related to understanding reply relationships and stance direction rather than conversation length alone. These findings highlight both the potential and limitations of Jev for stance detection.
Chinese Translation
立场检测需要识别作者对给定目标的态度,有时还要基于对话上下文。Jev 是一种专为结构化决策设计的专用决策模型,为通用大型语言模型(LLMs)提供了一种替代方案。在这项工作中,我们在两个立场检测数据集 VAST(英文文本)和 ZS-CSD(中文对话)上评估 Jev,并将其与四个通用大型语言模型和两个微调模型进行比较。结果表明,Jev 在 VAST 上取得了有竞争力的性能,与 GPT-5.6 相当,并优于其他通用大型语言模型。然而,它在 ZS-CSD 上落后于更强的 LLMs,尤其是在区分支持与反对方面。进一步分析表明,这一局限可能与会话中回复关系和立场方向的理解有关,而不仅仅是对话长度。这些发现凸显了 Jev 在立场检测方面的潜力和局限。
cs.CL / 58 / 2610.11915
Not Every Change Is Necessary: Recoverable Drift in Large Language Model Unlearning
并非每一处变化都是必要的:大型语言模型遗忘中的可恢复漂移
Xunlei Chen, Qinghui Gong, Jingkun Xue, Qihe Liu, Shijie Zhou, Fei Ye
cs.CL
large language model
大语言模型相关
Abstract
Machine unlearning in large language models aims to remove unwanted knowledge while preserving the model's remaining capabilities. Although existing methods use retention objectives or restrict where edits occur, achieving the desired forgetting level can still leave collateral changes that impair non-target behavior. Our recovery comparisons suggest that some of these changes can be reversed while preserving observed forgetting performance. In this work, we present Propose-Then-Project Unlearning (PTP-U), a framework that combines targeted forgetting with the recovery of non-target capabilities. PTP-U first applies local analytic edits to weaken target knowledge associations, then aligns non-target output distributions with those of the original model to recover capabilities while maintaining fixed forgetting constraints. Both stages serve a common goal: satisfying the forgetting requirements while preserving fluent generation and performance on non-target tasks. Across three benchmarks, PTP-U achieves the strongest forgetting-retention trade-off among evaluated methods, reaching 81.22%-91.03% forgetting while preserving 94.20% non-target utility on average. At matched forgetting, PTP-U consistently retains higher non-target utility.
Chinese Translation
大型语言模型中的机器遗忘旨在移除不需要的知识,同时保留模型其余的能力。尽管现有方法使用保留目标或限制编辑发生的位置,但达到期望的遗忘水平仍可能留下损害非目标行为的附带变化。我们的恢复比较表明,其中一些变化可以在保持所观察到的遗忘性能的同时被逆转。在这项工作中,我们提出 Propose-Then-Project Unlearning (PTP-U),一个将有针对性遗忘与非目标能力恢复相结合的框架。PTP-U 首先应用局部解析编辑来削弱目标知识关联,然后将非目标输出分布与原始模型的分布对齐,以在维持固定遗忘约束的同时恢复能力。这两个阶段服务于一个共同目标:在满足遗忘要求的同时,保持流畅生成以及非目标任务上的性能。在三个基准上,PTP-U 在受评估方法中实现了最强的遗忘-保留权衡,达到 81.22%-91.03% 的遗忘率,同时平均保留 94.20% 的非目标效用。在遗忘水平匹配时,PTP-U 始终保留更高的非目标效用。
cs.CL / 59 / 2610.11978
Specialized Decision Models vs. General-Purpose LLMs: Benchmarking Jev Across Knowledge, Reasoning, and Multilingual Tasks
专用决策模型 vs. 通用大语言模型:在知识、推理和多语言任务上对 Jev 进行基准测试
Xing Li, Qingcheng Chang, Jinzhong Ning, Changfeng Xu, Shenlong Zhang, Yijia Zhang, Ling Luo, Hongfei Lin
cs.CL
large language model
大语言模型相关
Abstract
Jev is a "System One" model that returns a choice among given options instead of generating text. We study how such a specialized decision model compares with general-purpose large language models (LLMs). We evaluate Jev on 13 multiple-choice benchmarks covering knowledge, reasoning, and multilingual understanding, and compare it with 19 LLMs in three tiers: frontier, representative, and small. Jev is competitive with frontier LLMs on knowledge and commonsense benchmarks and obtains the best score on MMLU-Redux and ARC-Challenge. Outside mathematics, it also outperforms most representative LLMs and all small LLMs. However, it falls behind on mathematical word problems: on MathQA, it is 17.7 points below the frontier median and scores lower than all 19 LLMs. These results indicate that a specialized decision model can match general-purpose LLMs on decisions that rely mainly on knowledge, but not on decisions that require multi-step calculation.
Chinese Translation
Jev 是一个“System One”模型,它在给定选项中返回一个选择,而不是生成文本。我们研究这种专用决策模型与通用大语言模型(LLM)相比表现如何。我们在 13 个涵盖知识、推理和多语言理解的多项选择基准上评估 Jev,并将其与 19 个 LLM 进行比较,这些 LLM 分为三个层级:前沿、代表性和小型。Jev 在知识和常识基准上与前沿 LLM 具有竞争力,并在 MMLU-Redux 和 ARC-Challenge 上获得了最佳分数。在数学之外,它还优于大多数代表性 LLM 和所有小型 LLM。然而,它在数学应用题上表现落后:在 MathQA 上,它比前沿中位数低 17.7 分,并且得分低于所有 19 个 LLM。这些结果表明,专用决策模型可以在主要依赖知识的决策上匹配通用 LLM,但在需要多步计算的决策上则不能。
cs.CL / 60 / 2610.12022
Examining Social Attribution in LLM Reasoning: A Theory-Guided Probing Methodology
考察LLM推理中的社会归因:一种理论指导的探测方法
Zhaoxin Yu, Qingchao Kong, Dajun Zeng, Wenji Mao
cs.CL · cs.AI · cs.LG
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly deployed in sociotechnical systems where social attribution, the reasoning process attributing external events to the causes and reasons of agents' social behaviors, plays a critical role. These processes involve judgments of social cause, responsibility, and blame/credit to agents. Although attributional models are well-studied in social psychology and cognition through Attribution Theory, social attribution remains underexplored in AI, particularly LLM social reasoning. This paper provides the first systematic exploration of LLM social attribution. Our work focuses on responsibility and blame attributions, examining current LLMs' judgments and their underlying internal mechanisms. Guided by attribution theory, we construct a social attribution benchmark consisting of a Vignette subset based on classic scenarios from attribution theory research and a Reality subset based on real-world social narratives, yielding 7,639 responsibility/blame judgment questions. On this basis, we evaluate 32 representative LLMs and 5 basic non-LLM baselines. To further explore the internal mechanisms underlying the LLM judgment process, we develop a probing-based methodology to investigate the latent-space representations of 5 key attribution dimensions and the consistency of their influences on LLM judgments compared to those in human social attribution. Our research findings reveal that current LLMs exhibit measurable but incomplete agreement with human responsibility and blame judgments, and meanwhile, this agreement is positively correlated with model size. Some attribution dimensions are systematically decodable from specific positions in LLM hidden states, and their influences on the final judgment are consistent with those indicated by human Attribution Theory. The dataset and associated code are available at https://github.com/Yuzhaoxin946/SAB-Bench.
Chinese Translation
大语言模型(LLMs)正越来越多地部署于社会技术系统中,在这些系统中,社会归因——即将外部事件归因于行动者社会行为的原因与理由的推理过程——发挥着关键作用。这些过程涉及对社会原因、责任以及对行动者的责备/赞誉的判断。尽管通过归因理论,归因模型在社会心理学和认知领域已得到充分研究,但社会归因在人工智能中,尤其是在LLM社会推理中,仍探索不足。本文首次对LLM社会归因进行了系统性探索。我们的工作聚焦于责任归因和责备归因,考察当前LLM的判断及其潜在的内部机制。在归因理论的指导下,我们构建了一个社会归因基准,该基准由基于归因理论研究中经典情景的Vignette子集和基于现实世界社会叙事的Reality子集组成,共产生7,639个责任/责备判断问题。在此基础上,我们评估了32个具有代表性的LLM和5个基础的非LLM基线。为了进一步探索LLM判断过程背后的内部机制,我们开发了一种基于探测的方法,以研究5个关键归因维度的潜在空间表示,以及它们对LLM判断的影响与人类社会归因中的影响相比的一致性。我们的研究发现揭示,当前LLM与人类的责任和责备判断表现出可测量但不完全一致,同时,这种一致性同模型规模呈正相关。一些归因维度可以从LLM隐藏状态中的特定位置被系统性地解码,并且它们对最终判断的影响与人类归因理论所表明的影响一致。该数据集和相关代码可在 https://github.com/Yuzhaoxin946/SAB-Bench 获取。
cs.CL / 61 / 2610.12030
Natural Language to First-Order Logic LLM-based Autoformalization
自然语言到一阶逻辑的基于 LLM 的自动形式化
Andrea Brunello, Cristian Curaba, Luca Geatti, Michele Mignani, Angelo Montanari, Nicola Saccomanno
cs.CL
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) have renewed interest in autoformalization. Yet, when First-Order Logic (FOL) is considered as the target formalism, the field still lacks a unified task formulation and a systematic survey. This paper addresses this gap: we first provide a principled definition for the FOL-autoformalization task by distinguishing Ontology Extraction from Logical Translation, showing how their conflation obscures (cross-study) evaluation; we review existing datasets, evaluation metrics, and LLM-based methods, including fine-tuning, prompting, and verification-based refinement; we identify open challenges in benchmarking, semantic evaluation, ontology-aware methods, and end-to-end applications.
Chinese Translation
大型语言模型(LLMs)重新激起了对自动形式化的兴趣。然而,当将一阶逻辑(FOL)视为目标形式体系时,该领域仍然缺乏统一的任务表述和系统综述。本文弥补这一空白:我们首先通过区分本体抽取与逻辑翻译,为 FOL 自动形式化任务提供一个有原则的定义,展示它们的混淆如何模糊(跨研究)评估;我们回顾现有数据集、评估指标和基于 LLM 的方法,包括微调、提示和基于验证的细化;我们识别在基准测试、语义评估、本体感知方法和端到端应用中的开放挑战。
cs.CL / 62 / 2610.12118
A persistent accuracy ceiling in automated verbal deception detection
自动言语欺骗检测中持续存在的准确率上限
Riccardo Loconte, Jonas Festor, Zane Fatjanova, Mariam Bolkvadze, Bennett Kleinberg
cs.CL
large language model
大语言模型相关
Abstract
Automated methods have been proposed to overcome the limitations of human verbal deception detection, but evidence remains fragmented across disciplines. We systematically reviewed 25 years of research (289 reports, 6,136 classification models) and meta-analyzed 3,653 models nested within 97 datasets. Pooled accuracy was 74.4% (95% CI: 71.2%-77.4%) with substantial heterogeneity. Accuracy was driven by methodological quality (ground truth, data source, class balance, evaluation procedure) more than by model complexity: the adoption of embeddings and large language models has not translated into improved predictive performance. Only 12.46% of reports used data with verifiable ground-truth, and only 23.96% of models were evaluated on independent data. The pooled accuracy aligns with meta-analyses of manual approaches, suggesting a ceiling of 70-75%, unlikely to be lifted by current research conventions.
Chinese Translation
人们已提出自动化方法以克服人类言语欺骗检测的局限性,但相关证据在各学科之间仍然零散。我们系统综述了25年的研究(289篇报告,6,136个分类模型),并对嵌套于97个数据集中的3,653个模型进行了元分析。合并准确率为74.4%(95% CI:71.2%–77.4%),且存在显著的异质性。准确率更多由方法学质量(真实标签、数据来源、类别平衡、评估程序)所驱动,而非由模型复杂度所驱动:嵌入表示与大型语言模型的采用并未转化为预测性能的提升。仅12.46%的报告使用了具有可验证真实标签的数据,且仅23.96%的模型在独立数据上进行了评估。该合并准确率与人工方法的元分析结果相一致,提示存在70–75%的上限,且这一上限不太可能被当前的研究惯例所突破。
cs.CL / 63 / 2610.12242
TokenRouter: Efficient Serving System for Token-Level LLM Routing
TokenRouter:面向 token 级 LLM 路由的高效服务系统
Tianyu Fu, Tengxuan Liu, Ruoxi Wang, Yixin Dong, Yi Ge, Yichen You, Yu Wang
cs.CL
large language model
大语言模型相关
Abstract
Large language model (LLM) routing distributes inference work across different models, advancing the cost-quality Pareto frontier of LLM serving. While coarse-grained routing at the session or query level has been widely adopted in production systems, recent algorithmic work shows that fine-grained token-level routing can yield substantial efficiency and quality gains. However, efficiently serving token-level routed inference poses significant challenges to existing systems. Built on single-LLM assumptions, current systems suffer from severe step desynchronization and frequent batch admission delays under token-level routing, and they also impose high implementation complexity on developers. To address these challenges, we design TokenRouter, an efficient and developer-friendly serving system for token-level routed LLM inference. TokenRouter follows the principle of request-centric programming, model-centric execution: developers describe routing logic from the perspective of a single request, while the runtime launches a subserver for each LLM and dispatches requests asynchronously. Each subserver employs a delayed-batching scheduler, whose optimal hyperparameters are derived from a mathematical throughput model of the system. Across diverse routing algorithms, workloads, and model pairs, TokenRouter achieves 2.01-64.15x higher decoding throughput than existing systems, substantially advancing the serving efficiency of token-level LLM routing. Our code is available at https://github.com/thu-nics/TokenRouter.
Chinese Translation
大语言模型(LLM)路由将推理工作分配到不同的模型上,推进了 LLM 服务的成本-质量帕累托前沿。尽管会话级或查询级的粗粒度路由已在生产系统中得到广泛采用,但近期的算法工作表明,细粒度的 token 级路由可以带来显著的效率和质量收益。然而,高效地服务 token 级路由推理对现有系统提出了重大挑战。当前系统建立在单 LLM 假设之上,在 token 级路由下会遭受严重的步骤失同步和频繁的批次准入延迟,并且还给开发者带来了很高的实现复杂度。为了解决这些挑战,我们设计了 TokenRouter,一个高效且对开发者友好的、用于 token 级路由 LLM 推理的服务系统。TokenRouter 遵循以请求为中心的编程、以模型为中心的执行这一原则:开发者从单个请求的视角描述路由逻辑,而运行时为每个 LLM 启动一个子服务器并异步分派请求。每个子服务器采用延迟批处理调度器,其最优超参数由系统的数学吞吐量模型推导得到。在多种路由算法、工作负载和模型对下,TokenRouter 实现了比现有系统高 2.01-64.15 倍的解码吞吐量,显著推进了 token 级 LLM 路由的服务效率。我们的代码可在 https://github.com/thu-nics/TokenRouter 获取。
cs.CL / 64 / 2610.12338
VFold: Symmetry-Aware Cross-Layer Value Cache Compression
VFold:对称感知的跨层值缓存压缩
Neha Verma, Sungwon Kim, Kenton Murray, Kevin Duh
cs.CL · cs.LG
large language model
大语言模型相关
Abstract
While caching key-value (KV) states accelerates Large Language Model (LLM) decoding, this cache can dominate memory usage at long context lengths. One solution is to compress this memory by exploiting inter-layer cache similarities. However, most existing techniques necessitate architectural changes to LLMs and incur substantial overhead. In this work, we propose a symmetry-aware value cache merging strategy that reduces cache memory while avoiding both harmful performance degradation and architectural overhead during decoding. Furthermore, we show that this approach can be exploited alongside existing cache compression techniques, composing with high-ratio quantization or key cache pruning to reach compression ratios that neither method reaches alone, with minimal additional cost. Ultimately, our findings reveal a major source of underutilized capacity in the value cache, offering a simple yet highly effective direction for scaling context windows under memory constraints.
Chinese Translation
尽管缓存键值(KV)状态可以加速大语言模型(LLM)的解码,但在长上下文长度下,该缓存可能会主导内存使用。一种解决方案是通过利用层间缓存相似性来压缩这部分内存。然而,大多数现有技术需要对 LLM 进行架构改动,并带来显著的开销。在本工作中,我们提出了一种对称感知的值缓存合并策略,它在减少缓存内存的同时,避免了有害的性能下降以及解码过程中的架构开销。此外,我们表明该方法可以与现有的缓存压缩技术协同使用,与高比例量化或键缓存剪枝相组合,以在最小的额外成本下达到任一方法单独都无法达到的压缩率。最终,我们的发现揭示了值缓存中一个未被充分利用的容量来源,为在内存约束下扩展上下文窗口提供了一个简单却极为有效的方向。
cs.CR / 65 / 2610.10742
BRANCH: Bypassing Multi-Scanner AI Guardrails
BRANCH:绕过多扫描器 AI 防护栏
William Hackett, Peter Garraghan
cs.CR · cs.AI
large language model
大语言模型相关
Abstract
AI systems increasingly rely on Large Language Models (LLMs) as core reasoning engines, making them targets for prompt injection and jailbreaks. Guardrails monitor and validate model inputs and outputs, yet their isolated, task-focused detection leaves gaps in their classification making them susceptible to bypasses. In response, guardrail systems formed by multiple scanners have emerged that collaboratively detect different types of malicious instructions, whereby shared latent representations across classification boundaries render established bypassing techniques ineffective. We propose BRANCH, a bypassing methodology designed for multi-scanner guardrail systems. Our method leverages a branching tree search approach that dynamically applies adversarial perturbation against individual scanners, with subsequent perturbation optimization and technique selection based on overall improvement across all guardrail system scanners, effectively decoupling bypass evaluation from attack signal optimization. Our findings demonstrate that BRANCH achieves 100% attack success rate across 6 guardrail systems in 120 scenarios with 72% fewer queries and 4.5x reduced wallclock time compared to established techniques, while preserving semantic meaning within the bypass. We also show how bypasses generated by BRANCH transfer to 29 unseen guardrails, including 8 commercial black-box guardrails, improving attack success in some cases up to 100% with no additional optimization.
Chinese Translation
AI 系统日益依赖大型语言模型(LLMs)作为核心推理引擎,这使其成为提示注入和越狱攻击的目标。防护栏监控并验证模型输入与输出,然而它们孤立、以任务为中心的检测在其分类中留下了缺口,使其容易受到绕过。作为回应,由多个扫描器组成的防护栏系统已经出现,它们协同检测不同类型的恶意指令,借此跨越分类边界的共享潜在表示使既有的绕过技术失效。我们提出 BRANCH,一种为多扫描器防护栏系统设计的绕过方法。我们的方法利用一种分支树搜索方法,动态地对单个扫描器施加对抗性扰动,并基于所有防护栏系统扫描器上的整体改进进行后续扰动优化和技术选择,从而有效地将绕过评估与攻击信号优化解耦。我们的研究结果表明,与既有技术相比,BRANCH 在 120 个场景中的 6 个防护栏系统上实现了 100% 的攻击成功率,同时查询次数减少 72%,挂钟时间减少 4.5 倍,并且在绕过中保持了语义含义。我们还展示了 BRANCH 生成的绕过如何迁移到 29 个未见过的防护栏,包括 8 个商业黑盒防护栏,在没有额外优化的情况下,在某些情况下将攻击成功率提升至高达 100%。
cs.CR / 66 / 2610.10752
Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection
基于扩散模型的检测引导自适应净化用于鲁棒音频深度伪造检测
Muhammed Salih Kayhan, Qiben Yan
cs.CR · cs.SD
diffusion
扩散模型相关
Abstract
Audio deepfake detectors remain vulnerable to adversarial perturbations that suppress the acoustic cues used for detection, allowing manipulated utterances to evade the detector. Although existing defenses can improve robustness, they require retraining the detector or introduce additional distortion. Diffusion-based purification instead leaves the pretrained detector unchanged, but existing methods use the same purification strength for all inputs, creating a trade-off between removing adversarial perturbations and preserving the subtle spoofing cues needed for detection. In this paper, we propose Detection-Guided Adaptive Purification (DGAP), a diffusion-based defense that adjusts purification strength per input. Building on the observation that a light purification perturbs the detector score of an adversarial input far more than that of a benign one, the framework uses the resulting score shift as a reference-free indicator of adversarial manipulation. Inputs with small shifts are passed unchanged, whereas flagged inputs undergo stronger purification before final detection. We evaluate the framework against three adversarial attack settings across three deepfake detectors, and compare it with nine existing defenses. Our results show that DGAP achieves the best defense performance across all detectors while leaving benign inputs nearly unaffected, and remains effective under the defense-aware adaptive attack.
Chinese Translation
音频深度伪造检测器仍然容易受到对抗性扰动的影响,这些扰动会抑制用于检测的声学线索,从而使经过操纵的语音得以逃避检测器。尽管现有的防御方法可以提高鲁棒性,但它们需要重新训练检测器,或引入额外的失真。基于扩散的净化则保持预训练检测器不变,但现有方法对所有输入使用相同的净化强度,从而在去除对抗性扰动与保留检测所需的细微欺骗线索之间形成权衡。在本文中,我们提出了检测引导自适应净化(DGAP),这是一种基于扩散的防御方法,能够针对每个输入调整净化强度。基于如下观察——轻度净化对对抗性输入的检测器分数造成的扰动远大于对良性输入的扰动——该框架将由此产生的分数偏移用作一种无需参考的对抗性操纵指示器。偏移较小的输入保持不变地通过,而被标记的输入则在最终检测前接受更强的净化。我们在三种深度伪造检测器上针对三种对抗性攻击设置对该框架进行了评估,并将其与九种现有防御方法进行了比较。我们的结果表明,DGAP 在所有检测器上均取得了最佳的防御性能,同时几乎不影响良性输入,并且在防御感知的自适应攻击下仍然有效。
cs.CR / 67 / 2610.10929
Speedbumps: Rejection Attacks on Speculative Decoding
减速带:对推测解码的拒绝攻击
Adam Y. J. Jones, Yu Yuan, Sergio Maffeis
cs.CR · cs.AI
large language model
大语言模型相关
Abstract
Speculative decoding is a popular technique for increasing the speed and reducing the costs of large language model (LLM) inference by verifying multiple draft tokens in a single target-model forward pass. The resulting benefit depends on the ability of the drafter to approximate the target model's distribution. In this work, we study Speculative Rejection Attacks (SRAs), a novel class of attacks that cause draft and target models to disagree more often, resulting in fewer draft tokens being accepted per draft cycle. This leads to more target model forward passes needed per generated token, slowing down inference and increasing costs for the victim. We introduce two attacks which append an adversarial suffix to attacker-controlled content to degrade speculative decoding on a victim's prompts. Both attacks optimise the expected length of the accepted speculative prefix, estimating per-depth acceptance from the target's probability of the drafted proposals (Speedbump-P) or from the overlap between the draft and target distributions (Speedbump-D). In some cases, attacks degrade speculative decoding to the point of being slower than autoregressive decoding. The degradation reduces the output quality - regularisation restores output quality but gives up most of the degradation, trading effectiveness for stealthiness. Additionally, the suffixes remain effective under sampling, and transfer across drafters (Speedbump-P) or across target models sharing a drafter (Speedbump-D). These findings identify the draft-target interaction of speculative decoding as a realistic attack surface through which adversarial inputs can inflate inference costs.
Chinese Translation
推测解码是一种流行的技术,它通过在一次目标模型前向传播中验证多个草稿 token,来提高大型语言模型(LLM)推理的速度并降低成本。由此产生的收益取决于草稿模型逼近目标模型分布的能力。在这项工作中,我们研究推测拒绝攻击(Speculative Rejection Attacks, SRAs),这是一类新颖的攻击,会使草稿模型与目标模型更频繁地不一致,从而导致每个草稿周期中被接受的草稿 token 更少。这导致每生成一个 token 需要更多次目标模型前向传播,从而减慢推理速度并增加受害者的成本。我们引入两种攻击,它们将对抗性后缀附加到攻击者控制的内容上,以在受害者的提示上劣化推测解码。两种攻击都优化被接受的推测前缀的期望长度,分别从目标模型对草拟提议的概率(Speedbump-P)或从草稿分布与目标分布之间的重叠(Speedbump-D)来估计逐深度的接受率。在某些情况下,攻击会把推测解码劣化到比自回归解码还慢的程度。这种劣化会降低输出质量——正则化可以恢复输出质量,但会放弃大部分劣化效果,以有效性换取隐蔽性。此外,这些后缀在采样下仍然有效,并且可以跨草稿模型迁移(Speedbump-P),或跨共享同一草稿模型的目标模型迁移(Speedbump-D)。这些发现将推测解码中的草稿-目标交互识别为一个现实的攻击面,通过该攻击面,对抗性输入可以抬高推理成本。
cs.CR / 68 / 2610.11556
SoK: Are LLMs Reliable at Source Code Recovery? A Taxonomy and Empirical Evaluation
SoK:大语言模型在源代码恢复中可靠吗?一个分类体系与实证评估
Varun Kohli, Lee Bing Cheng, Nur Hazim Ghazali, Gao Yuze, Daryl Poon, Dinil Mon Divakaran
cs.CR · cs.SE
large language model
大语言模型相关
Abstract
Effective source recovery is critical to security applications such as malware analysis, vulnerability assessment, and legacy maintenance. Large Language Models (LLMs) are reshaping this field, shifting the paradigm away from rule-based heuristics to probabilistic and high fidelity semantic recovery of source code from assembly or classical decompiler-derived pseudo-C. However, despite rapid progress, the field suffers from fragmentation across numerous approaches as well as their non-unified evaluations, limiting objective comparisons. Further, existing works have limited coverage of embedded, IoT architectures and source languages beyond C/C++. In this work, we present the first Systematization of Knowledge (SoK) focused specifically on LLM-assisted binary-to-source recovery. We provide a granular design-centric taxonomy of LLM-assisted source recovery methods and systematic evaluations using seven key metrics along six evaluation dimensions. We evaluate state-of-the-art methods on 45,000 test samples derived from four standard and five embedded architectures, five optimization levels, and symbol stripping. We ablate the impact of design choices on recovery performance, including input representation, contextual enrichment, model scale, iteration and review roles using controlled in-house recovery pipelines and three off-the-shelf models. Finally, we test same-language and cross-language recovery capability covering four mature and two legacy languages. Our systematization and comprehensive evaluations provide key insights that guide future directions in this field.
Chinese Translation
有效的源代码恢复对于恶意软件分析、漏洞评估和遗留系统维护等安全应用至关重要。大语言模型(LLMs)正在重塑这一领域,将范式从基于规则的启发式方法转向从汇编代码或经典反编译器生成的伪 C 代码中进行概率化且高保真的源代码语义恢复。然而,尽管进展迅速,该领域仍受困于众多方法之间的碎片化以及其不统一的评估,限制了客观比较。此外,现有工作在嵌入式、IoT 架构以及 C/C++ 之外的源语言方面覆盖有限。在本工作中,我们提出了首个专门聚焦于 LLM 辅助的二进制到源代码恢复的知识系统化(SoK)。我们提供了一个细粒度的、以设计为中心的 LLM 辅助源代码恢复方法分类体系,并沿六个评估维度使用七个关键指标进行了系统性评估。我们在 45,000 个测试样本上评估了最先进的方法,这些样本源自四种标准架构和五种嵌入式架构、五个优化级别以及符号剥离。我们使用受控的自研恢复流水线和三种现成模型,消融分析了设计选择对恢复性能的影响,包括输入表示、上下文增强、模型规模、迭代与审查角色。最后,我们测试了同语言与跨语言恢复能力,覆盖四种成熟语言和两种遗留语言。我们的系统化梳理与全面评估提供了指导该领域未来方向的关键洞见。
cs.CR / 69 / 2610.11634
LTBD: Learnable Trust-Boundary Delimiters for Prompt Injection Defense
LTBD:用于提示注入防御的可学习信任边界分隔符
Luman Zhao, Minghui Xu, Yue Zhang, Yijun Yang
cs.CR · cs.AI
large language model
大语言模型相关
Abstract
Large language models (LLMs) perform remarkably well on complex tasks, yet remain highly vulnerable to prompt injection attacks, where malicious instructions embedded in external data can override user intent. Existing defenses remain limited by model fine-tuning requirements, vulnerability to adaptive attacks, or reliance on brittle handcrafted prompts. We argue that a fundamental source of this vulnerability is the lack of an explicit representation of trust provenance. To address this, we introduce Learnable Trust-Boundary Delimiters (LTBD), a lightweight defense that explicitly encodes trust boundaries in the input while keeping the LLM parameters unchanged. LTBD uses a small number of learnable delimiters to distinguish trusted user instructions from untrusted external data, enabling the model to better respect the intended trust hierarchy. Experimental results show that LTBD substantially outperforms inference-time defenses and performs competitively with training-based approaches, while preserving benign-task utility and introducing negligible inference overhead. In particular, LTBD achieves 0.00% ASR on AlpacaFarm and only 0.11-0.19% ASR on TaskTracker. LTBD also remains effective under adaptive attacks, where adversaries have full knowledge of the defense and explicitly attempt to bypass it.
Chinese Translation
大型语言模型(LLMs)在复杂任务上表现非常出色,但仍然极易受到提示注入攻击的影响;在这类攻击中,嵌入外部数据中的恶意指令可以覆盖用户意图。现有防御仍受限于模型微调需求、对自适应攻击的脆弱性,或对脆弱的手工提示的依赖。我们认为,这一脆弱性的一个根本来源是缺乏对信任来源的显式表示。为了解决这一问题,我们提出了可学习的信任边界分隔符(LTBD),这是一种轻量级防御方法,它在输入中显式编码信任边界,同时保持 LLM 参数不变。LTBD 使用少量可学习的分隔符来区分可信的用户指令与不可信的外部数据,使模型能够更好地遵循预期的信任层级。实验结果表明,LTBD 显著优于推理时防御,并且与基于训练的方法相比具有竞争力,同时保持良性任务效用并引入可忽略的推理开销。特别地,LTBD 在 AlpacaFarm 上实现了 0.00% 的 ASR,在 TaskTracker 上仅实现了 0.11-0.19% 的 ASR。LTBD 在自适应攻击下也仍然有效;在自适应攻击中,攻击者完全了解该防御并明确试图绕过它。
cs.CR / 70 / 2610.11848
SemField: A Simple, Linear, Continuous, yet Robust Semantic Watermark
SemField:一种简单、线性、连续却稳健的语义水印
Varun Gumma, Navonil Majumdar, Soujanya Poria
cs.CR
large language model
大语言模型相关
Abstract
The rapid proliferation of Large Language Models (LLMs) necessitates reliable watermarking techniques to identify AI-generated text and ensure appropriate attribution. While token-based watermarks are vulnerable to paraphrasing, a central challenge for semantic watermarking is to turn sentence meanings into a stable, well-calibrated document-level signal. To this end, we introduce SemField, a simple, training-free semantic watermark that embeds a continuous, linear signal directly into the sentence embedding space. Using a shared secret key to define a specific Gaussian direction, SemField iteratively evaluates candidate sentences and selects those that maximize the alignment of the cumulative document aggregate with this targeted direction. We also propose SemField-PL, a polarity-locked variant that first determines the optimal orientation from the initial sentence and continuously reinforces it throughout the generation process. The document-level aggregated formulation guarantees exact invariance to sentence reordering and provides theoretical bounds against structural tampering, such as sentence insertion and deletion. Lastly, with extensive evaluations across three models, we demonstrate that both variants outperform 12 recent baselines. Across clean detection and four paraphrasing attacks, both variants achieve a mean True Positive Rate (TPR) of 88.2% to 90.4% at a 1% False Positive Rate (FPR), all while maintaining comparable perplexity and naturalness as human-generated content. We open-source our implementation at https://github.com/declare-lab/SemField
Chinese Translation
大型语言模型(LLM)的迅速普及使得可靠的水印技术成为必要,以识别 AI 生成的文本并确保适当的归属。虽然基于词元的水印容易受到改写攻击,但语义水印的一个核心挑战是将句子含义转化为稳定且校准良好的文档级信号。为此,我们提出 SemField,一种简单、无需训练的语义水印,它将连续、线性的信号直接嵌入句子嵌入空间。SemField 使用共享密钥来定义一个特定的高斯方向,迭代地评估候选句子,并选择那些能使累积文档聚合与该目标方向的对齐程度最大化的句子。我们还提出 SemField-PL,一种极性锁定变体,它首先从初始句子确定最优方向,并在整个生成过程中持续强化该方向。文档级聚合公式保证了对句子重排序的精确不变性,并提供了针对结构篡改(如句子插入和删除)的理论界限。最后,通过在三个模型上的广泛评估,我们证明这两种变体都优于 12 个近期基线。在干净检测和四种改写攻击中,两种变体在 1% 假阳性率(FPR)下均达到 88.2% 到 90.4% 的平均真阳性率(TPR),同时保持与人类生成内容相当的困惑度和自然度。我们在 https://github.com/declare-lab/SemField 开源了我们的实现。
cs.CR / 71 / 2610.11893
A Security Meta-Model for Retrieval-Augmented Generation Systems
面向检索增强生成系统的安全元模型
Steve Nouyep, Sébastien Salva, Maxime Puys
cs.CR
large language model
大语言模型相关
Abstract
Retrieval-Augmented Generation (RAG) systems extend large language models (LLMs) with external knowledge through a multi-stage pipeline. While this architecture can improve the factual grounding of generated answers, it introduces structural attack surfaces that extend beyond those of standalone LLMs. In this paper, we introduce a security meta-model that captures explicit causal relationships between RAG surfaces, attacks, weaknesses, risks, and CIA impact (Confidentiality, Integrity, Availability). Its purpose is to provide security engineers with a structured and user-friendly framework for gathering and assessing the risks, weaknesses, and mitigations relevant to their RAG deployment. We designed the meta-model through an iterative, structured analysis of 43~publications (2023--2026) and instantiated it as a catalog populated with the security threats and remediations reported in the literature. Filtering the catalog according to a deployment configuration produces a risk profile containing the risks applicable to that deployment. An interactive web visualizer lets users navigate the catalog as a graph, follow causal chains, and explore stakeholder-specific views. Analysis of the catalog revealed a persistent imbalance between attack-focused and defense-focused research, a concentration of threats at ingestion, and coverage gaps affecting output integrity. Coverage is assessed against the OWASP LLM Top~10, and operational applicability is illustrated across textual, graph-based, and multimodal RAG configurations.
Chinese Translation
检索增强生成(RAG)系统通过多阶段流水线,用外部知识扩展大语言模型(LLM)。尽管这种架构能够提升所生成答案的事实依据性,但它引入了超出独立 LLM 的结构性攻击面。在本文中,我们提出了一种安全元模型,用以刻画 RAG 面、攻击、弱点、风险与 CIA 影响(机密性、完整性、可用性)之间显式的因果关系。其目的是为安全工程师提供一个结构化且易于使用的框架,用于收集和评估与其 RAG 部署相关的风险、弱点和缓解措施。我们通过对 43 篇出版物(2023--2026)进行迭代式、结构化的分析设计了该元模型,并将其实例化为一个目录,其中填充了文献中报告的安全威胁与修复措施。根据部署配置对目录进行过滤,可生成包含适用于该部署的风险的风险画像。一个交互式网页可视化工具让用户能够以图的形式浏览该目录、追踪因果链,并探索面向不同利益相关者的视图。对目录的分析揭示出:以攻击为重心的研究与以防御为重心的研究之间存在持续的不平衡,威胁集中在摄取阶段,以及影响输出完整性的覆盖缺口。覆盖情况依据 OWASP LLM Top 10 进行评估,并通过文本型、基于图的以及多模态的 RAG 配置来说明其运行适用性。
cs.CR / 72 / 2610.12106
Could LLM Watermark Detection be Public?
大语言模型水印检测能否公开?
Georgios Milis, Tom Sander, Tomáš Souček, Heng Huang, Pierre Fernandez
cs.CR · cs.LG
large language model
大语言模型相关
Abstract
Watermarking large language models is popular for tracing chatbot and agentic outputs, yet detectors remain unreleased since exposing them could let attackers do targeted edits with the detector's feedback. However, watermarks are already vulnerable to uninformed tampering attacks. We thus first quantify whether a public detector would be an additional liability in a deployment setting at varying levels of access, from token-level scores to a binary verdict. Second, we introduce a split-key public-private watermarking method that exposes one key through a public detector while keeping the other for full verification and forensics. An informed attacker can only move the public signal, creating an imbalance between public and private scores. We introduce a statistical test for this imbalance, and combine it with the full key verdict in a two-stage mechanism. Third, we evaluate the split-key method on a wide range of removal and forgery attacks, comparing the uninformed to detector-informed settings. Public detection improves removal only at small edit budgets, since plain rephrasing already strips the watermark at a lower quality cost, but it does enable forgery, which the private pipeline can identify. Overall, releasing half of the watermark enables transparency and interoperability, and tampering with the released half stays detectable. This bounds the provider's liability and questions the need to keep detectors fully private.
Chinese Translation
为大型语言模型添加水印在追踪聊天机器人和智能体输出方面十分流行,但检测器却始终未予公开,因为将其暴露可能使攻击者能够借助检测器的反馈进行有针对性的编辑。然而,水印本就容易受到无信息篡改攻击。因此,我们首先量化了在不同访问级别下——从词元级分数到二元判定——公开检测器在部署环境中是否会构成额外的风险负担。其次,我们提出了一种分密钥的公私水印方法,该方法通过公开检测器暴露其中一个密钥,同时保留另一个密钥用于完整验证与取证。掌握信息的攻击者只能移动公开信号,从而在公开分数与私有分数之间造成失衡。我们针对这种失衡引入了一种统计检验,并将其与完整密钥判定结合在一个两阶段机制中。第三,我们在广泛的去除与伪造攻击上评估了分密钥方法,比较了无信息场景与检测器知情场景。公开检测仅在小编辑预算下提升去除效果,因为单纯改写已经能以更低的质量代价剥离水印,但它确实使伪造成为可能,而私有流水线能够识别出这种伪造。总体而言,公开一半水印能够实现透明性与互操作性,而对已公开那一半的篡改仍可被检测到。这限制了提供方的责任,并对必须将检测器完全保密的必要性提出了质疑。
cs.CR / 73 / 2610.12137
Poster: A Preliminary Study of LLM Distillation Inference
海报:LLM 蒸馏推断的初步研究
Edward Chen, Yuntao Du
cs.CR · cs.LG
large language model
大语言模型相关
Abstract
Unauthorized model distillation, in which a model is trained on the outputs of a proprietary large language model (LLM), is a growing threat to model providers. We study distillation inference: determining whether a suspect model was distilled from another model or trained independently. We formulate this problem as a hypothesis test and estimate the behavior expected under each hypothesis by training shadow models: distilled shadow models learn from the teacher's reasoning traces, whereas independent shadow models learn only from reference answers. The auditor measures how closely each model predicts the teacher's reasoning outputs and then uses the shadow models to convert the suspect's score into a calibrated p-value. In a preliminary study using Qwen2.5-7B as the teacher and Llama-3.2-3B for the suspects, our test achieves a true positive rate of 1.0 at a significance level of 0.02. These results demonstrate the feasibility of using distillation inference to detect distillation attacks.
Chinese Translation
未经授权的模型蒸馏,即使用专有大型语言模型(LLM)的输出训练一个模型,正对模型提供者构成日益严重的威胁。我们研究蒸馏推断:判定一个可疑模型是从另一个模型蒸馏而来,还是独立训练的。我们将该问题表述为假设检验,并通过训练影子模型来估计每种假设下预期的行为:蒸馏影子模型从教师的推理轨迹中学习,而独立影子模型仅从参考答案中学习。审计者测量每个模型与教师推理输出的预测有多接近,然后使用影子模型将可疑模型的分数转换为校准后的 p 值。在一项以 Qwen2.5-7B 作为教师、以 Llama-3.2-3B 用于可疑模型的初步研究中,我们的检验在 0.02 的显著性水平下达到了 1.0 的真阳性率。这些结果证明了使用蒸馏推断来检测蒸馏攻击的可行性。
cs.LG / 74 / 2610.10859
Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models
在少步蒸馏文本到图像扩散模型中实现偏好驱动的遗忘
Gaurav Patel, Jun Fang, Greg Ver Steeg, Qiang Qiu, Sravan Sripada
cs.CV · cs.LG
diffusion
扩散模型相关
Abstract
Text-to-image diffusion models are increasingly distilled into few-step variants and being deployed to enable fast inference. However, their ability to generate harmful or undesired content poses significant safety risks. Data-driven unlearning methods suppress targeted generations by fine-tuning model weights using specialized unlearning objectives. Crucially, these objectives implicitly rely on multi-step denoising dynamics, an assumption that breaks down for few-step distilled (FSD) models, resulting in ineffective forgetting. Furthermore, performing unlearning on the non-distilled base model and subsequently re-distilling it to obtain an unlearned FSD model incurs substantial computational and time overhead, making it impractical in many settings. Hence, we address this limitation with a preference-driven unlearning framework that revisits Direct Preference Optimization (DPO) for diffusion models. We show that standard DPO and its unlearning derivatives, formulated around noise-prediction error, transfer poorly to FSD models due to their altered generation dynamics. To overcome this, we introduce a modified preference optimization formulation explicitly aligned with the few-step generation properties, enabling direct concept removal in FSD models while preserving few-step efficiency and maintaining strong retention of desirable (non-targeted) capabilities. We evaluate our framework primarily on identity and NSFW (nudity) removal tasks and also extend our method to object-level unlearning. Extensive experiments demonstrate consistent and effective forgetting, and strong retention performance, establishing our method as a practical and principled solution for unlearning in FSD models.
Chinese Translation
文本到图像扩散模型正日益被蒸馏为少步变体,并被部署以实现快速推理。然而,它们生成有害或不良内容的能力带来了重大的安全风险。数据驱动的遗忘方法通过使用专门的遗忘目标微调模型权重来抑制目标生成。关键在于,这些目标隐含地依赖于多步去噪动态,而这一假设在少步蒸馏(FSD)模型中不再成立,从而导致遗忘无效。此外,在未蒸馏的基座模型上执行遗忘,随后再对其进行重新蒸馏以获得已遗忘的 FSD 模型,会带来大量的计算与时间开销,使其在许多场景下不切实际。因此,我们通过一个偏好驱动的遗忘框架来解决这一局限,该框架重新审视了用于扩散模型的直接偏好优化(DPO)。我们表明,围绕噪声预测误差构建的标准 DPO 及其遗忘衍生方法,由于其生成动态发生了改变,难以很好地迁移到 FSD 模型。为克服这一问题,我们引入了一种改进的偏好优化公式,其与少步生成特性显式对齐,能够在 FSD 模型中直接移除概念,同时保持少步效率,并维持对期望(非目标)能力的强保留。我们主要在身份移除和 NSFW(裸露)移除任务上评估我们的框架,并将我们的方法扩展到对象级遗忘。大量实验表明,该方法能够实现一致且有效的遗忘,并具有强大的保留性能,从而确立了我们的方法作为 FSD 模型中一种实用且有原则的遗忘解决方案。
cs.LG / 75 / 2610.10990
Omni-Diffusion-Distill: Few-Step Distillation of Unified Multimodal Diffusion Large Language Models
Omni-Diffusion-Distill:统一多模态扩散大语言模型的少步蒸馏
Hong Huang, Chenhongyi Yang, Junzhe Sun, Animesh Sinha, Wuyang Chen, Yifan Jiang
cs.CV · cs.LG
diffusionlarge language model
扩散模型相关
大语言模型相关
Abstract
Unified multimodal diffusion large language models (dLLMs) offer a single architecture for both image generation and multimodal understanding, but their iterative decoding requires tens to hundreds of forward passes. Existing few-step distillation methods largely focus on either image generation or text generation, making it unclear how to compress a fully discrete multimodal dLLM into a single efficient student while preserving both generation and understanding. We introduce Omni-Diffusion-Distill, a unified two-stage distillation framework that retains strong generation and understanding capabilities while substantially reducing the inference cost of a unified multimodal dLLM. Omni-Diffusion-Distill aligns the distillation of both generation and understanding, for both images and text, in the discrete token space. In the first stage, the student is trained to skip decoding steps by replaying cached teacher trajectories, and in the second stage the student is refined on intermediate states along its own rollouts. We further remedy two sources of degradation in unified distillation with a pairwise collision penalty that reduces repetition under parallel text decoding, and entropy-matched guidance that prevents entropy collapse caused by fitting the sharpened teacher distribution in image generation. Omni-Diffusion-Distill achieves state-of-the-art trade-offs between decoding efficiency and generation and understanding performance for multimodal dLLMs, reducing image generation from 128 to 8 decoding steps and multimodal understanding from 512 to 64, giving 18.2x and 21.2x wall-clock speedups. Under these budgets, it scores 0.828 on GenEval and 83.0 on DPG-Bench for text-to-image generation, while reaching GPT judge scores of 20.0 on MM-Vet and 57.2 on COCO captioning (twice the teacher's 28.4 at the same steps) for multimodal understanding.
Chinese Translation
统一多模态扩散大语言模型(dLLMs)为图像生成和多模态理解提供了单一架构,但其迭代解码需要数十到数百次前向传播。现有的少步蒸馏方法大多集中于图像生成或文本生成中的某一项,这使得如何将一个完全离散的多模态 dLLM 压缩为单个高效的学生模型、同时保留生成与理解能力这一问题仍不明确。我们提出 Omni-Diffusion-Distill,一个统一的两阶段蒸馏框架,它在显著降低统一多模态 dLLM 推理成本的同时,保留了强大的生成与理解能力。Omni-Diffusion-Distill 在离散 token 空间中对图像与文本两者的生成和理解蒸馏进行了对齐。在第一阶段,学生模型通过重放缓存的教师轨迹来训练以跳过解码步骤;在第二阶段,学生模型在其自身 rollout 沿途的中间状态上得到精炼。我们进一步用成对碰撞惩罚(pairwise collision penalty)与熵匹配引导(entropy-matched guidance)来弥补统一蒸馏中的两个退化来源,前者减少并行文本解码下的重复,后者防止因拟合图像生成中被锐化的教师分布而导致的熵坍塌。Omni-Diffusion-Distill 在多模态 dLLM 的解码效率与生成和理解性能之间实现了最先进的权衡,将图像生成从 128 个解码步骤减少到 8 个,将多模态理解从 512 个减少到 64 个,分别带来 18.2 倍和 21.2 倍的墙钟时间加速。在这些预算下,它在文本到图像生成上于 GenEval 取得 0.828、于 DPG-Bench 取得 83.0,同时在多模态理解上取得 GPT 评判分数:MM-Vet 上为 20.0,COCO 图像描述上为 57.2(是教师模型在相同步骤下 28.4 的两倍)。
cs.AI / 76 / 2610.11019
Mid-Training Language Models on Raw Video
在原始视频上进行语言模型的中期训练
Jaedong Hwang, Xiaoqian Shen, Ernie Chang, Changsheng Zhao, Chong Zhou, Saksham Suri, Qi Qian, Zechun Liu, Lemeng Wu, Qinsi Wang, Raghuraman Krishnamoorthi, Wei Wen
cs.CV · cs.AI · cs.LG
large language model
大语言模型相关
Abstract
Multimodal large language models learn mostly from paired image-text data or annotated video, and raw web video is rarely used to further train an existing language model. We study whether raw video, with no captions and no text loss, can serve as mid-training data for a pretrained language model. Frames are encoded into continuous visual tokens, and the language model learns to predict the next visual token. We mid-train Qwen3-1.7B on raw clips from YT-Temporal-1B and then apply the same image-text instruction tuning to it and to the model without mid-training, so that the two differ only in mid-training. The mid-trained model scores 2.9 points higher on average across four video benchmarks and 5.1 points higher across ten image benchmarks, spanning perception, document, and chart tasks. Text performance is preserved even though mid-training includes no text, with an average of 48.9 across 14 text benchmarks compared with 48.0 for the model without mid-training. Analyses across training show that the image and video gains emerge within 30% of training and plateau thereafter, varying by less than 0.5 points. Predicting captions fails to outperform next-visual-token prediction, demonstrating that video mid-training can remain purely self-supervised without the computational overhead or labeling noise of automated captioning.
Chinese Translation
多模态大语言模型主要从成对的图像-文本数据或标注视频中学习,而原始网络视频很少被用于进一步训练已有的语言模型。我们研究了没有任何字幕、也不使用文本损失的原始视频,能否作为预训练语言模型的中期训练数据。帧被编码为连续视觉token,语言模型学习预测下一个视觉token。我们在来自YT-Temporal-1B的原始视频片段上对Qwen3-1.7B进行中期训练,然后对经过中期训练的模型和未经过中期训练的模型施加相同的图像-文本指令微调,从而使二者的差异仅在于中期训练。经过中期训练的模型在四个视频基准上平均高出2.9分,在十个图像基准上平均高出5.1分,涵盖感知、文档和图表任务。尽管中期训练不包含文本,文本性能仍得以保持,在14个文本基准上的平均分为48.9,而未经过中期训练的模型为48.0。对不同训练阶段的分析表明,图像和视频方面的提升在训练进度30%以内出现,此后趋于平稳,波动小于0.5分。预测字幕未能超越下一视觉token预测,这表明视频中期训练可以保持完全自监督,而无需自动字幕带来的计算开销或标注噪声。
cs.AI / 77 / 2610.11144
Improving Image-Based Nutrition Estimation Through Multimodal Food-Item Verification and Recovery
通过多模态食物项验证与恢复改进基于图像的营养估计
Jingbo Yue, Bruce Coburn, Jinge Ma, Jui-Feng Chi, Fengqing Zhu
cs.CV · cs.AI · eess.IV
large language model
大语言模型相关
Abstract
Single-image nutrition estimation can fail silently when visible foods are missed. Even when a food is correctly identified, its proposed region may not support portion estimation. We propose a framework that uses multimodal large language models (MLLMs) to inventory visible foods and separately verify food identity and whether each proposed 2D region supports portion estimation. One whole-image review uses these verification results to identify unresolved gaps and omitted foods, triggering at most one targeted recovery pass. Recovered regions are re-verified without access to the recovery prompt, then reconciled into a final item set for nutrition estimation. The framework requires no task-specific fine-tuning. Matched evaluation on common valid-output samples shows that item-level grounding improves mass accuracy across all tested settings and energy accuracy relative to an adapted retrieval baseline, with item-identity precision and recall also improving, while post-recovery visual coverage is assessed separately at inference time without ground-truth annotations.
Chinese Translation
单图像营养估计在漏掉可见食物时可能悄无声息地失败。即使某种食物被正确识别,其提出的区域也可能不支持份量估计。我们提出一个框架,该框架使用多模态大型语言模型(MLLMs)来盘点可见食物,并分别验证食物身份以及每个提出的 2D 区域是否支持份量估计。一次整图审查利用这些验证结果来识别未解决的缺口和被遗漏的食物,至多触发一次有针对性的恢复轮次。恢复的区域在无法访问恢复提示的情况下被重新验证,然后被协调合并为最终食物项集合,用于营养估计。该框架不需要任务特定的微调。在共同有效输出样本上的匹配评估表明,项目级定位在所有测试设置下提高了质量准确率,并相对于一个经过适配的检索基线提高了能量准确率,同时项目身份精确率和召回率也有所提高,而恢复后的视觉覆盖率则在推理时单独评估,无需真实标注。
cs.LG / 78 / 2610.11310
FloorSAV: Elucidating Spatial Audio-Visual Context with 2D Floormap for AV-LLMs
FloorSAV:利用 2D 平面图为 AV-LLM 阐明空间音视频上下文
Kyeong-Rae Kim, Sungnyun Kim, Tae-Hyun Oh
cs.CV · cs.LG
large language model
大语言模型相关
Abstract
While 3D spatial reasoning in dynamic egocentric environments is crucial for embodied intelligence, audio-visual large language models (AV-LLMs) lack explicit mechanisms to process and internalize global geometry directly from raw sensory streams. Existing approaches either require costly fine-tuning or underutilize the model's cross-modal reasoning capacities. In this paper, we propose FloorSAV, a novel framework that explicitly grounds spatial audio-visual context by rendering a dynamic 2D floormap. By integrating 3D point clouds, camera trajectories, spatial audio cues, and semantically grounded object landmarks, we inject this floormap into the AV-LLM as a synchronized stream with an egocentric video. AV-LLMs utilize their multi-modal capabilities to jointly reason over visual, auditory, and geometric cues in a single inference with floormap interpretation guidance. We further introduce SAVED-Bench (Spatial Audio-Visual Egocentric Benchmark with Dynamic Agents), constructing essential tasks of spatial capability in real-world scenarios: dynamic relativity, regional, and path reasoning QAs. FloorSAV improves AV-LLMs' spatial reasoning on various tasks from both SAVED-Bench and SAVVY-Bench. Studies with ground-truth floormaps demonstrate the substantial potential of FloorSAV with accurate spatial information.
Chinese Translation
尽管动态自我中心环境中的 3D 空间推理对于具身智能至关重要,但音视频大语言模型(AV-LLM)缺乏显式机制来直接从原始感知流中处理和内化全局几何。现有方法要么需要昂贵的微调,要么未能充分利用模型的跨模态推理能力。在本文中,我们提出 FloorSAV,这是一种新颖的框架,通过渲染动态 2D 平面图来显式地将空间音视频上下文落地。通过整合 3D 点云、相机轨迹、空间音频线索以及语义接地的物体地标,我们将该平面图作为与自我中心视频同步的流注入 AV-LLM。AV-LLM 利用其多模态能力,在平面图解释指导下,在单次推理中联合推理视觉、听觉和几何线索。我们进一步提出 SAVED-Bench(Spatial Audio-Visual Egocentric Benchmark with Dynamic Agents,带有动态智能体的空间音视频自我中心基准),构建了真实世界场景中空间能力的基本任务:动态相对性、区域和路径推理问答。FloorSAV 提升了 AV-LLM 在来自 SAVED-Bench 和 SAVVY-Bench 的各种任务上的空间推理能力。使用真实标注平面图的研究展示了 FloorSAV 在具有准确空间信息时的巨大潜力。
cs.AI / 79 / 2610.11376
CRISP: Fixing Flying Pixels in Latent LiDAR Generation via Diffusion Decoding
CRISP:通过扩散解码修复潜在LiDAR生成中的飞行像素
Andrea Ceron, Michael Schmidt, Alvaro Marcos-Ramiro, Sebastian Schmidt, Benjamin Busam
cs.CV · cs.AI
diffusion
扩散模型相关
Abstract
Latent LiDAR pipelines suffer from flying pixels: convolutional VAEs blur sharp radial depth discontinuities, yielding edge depths that back-project to points floating between surfaces. We identify this as a major, directly correctable decoder bottleneck and introduce CRISP: a pixel-space diffusion decoder with a backbone-agnostic latent adapter, DiT-based denoiser, and support mask predictor. CRISP replaces video-VAE and LiDAR-native decoders alike while keeping the encoder and latent generator fixed. Across KITTI-360, SemanticKITTI, and nuScenes, replacing only the decoder reduces FSVD/FPVD by 50.5% on average across frozen backbones; for generic video VAEs, the reductions reach 71%/74%. On the LiDAR-native LiDM backbone, FRID drops by 71%, with the largest gains at depth discontinuities. In a pretrained LiDM world model, the same zero-shot replacement improves FSVD by 15.5%, narrowing the sim-to-real gap.
Chinese Translation
潜在LiDAR流水线受到飞行像素的困扰:卷积VAE会模糊尖锐的径向深度不连续处,产生反投影后落在表面之间漂浮点的边缘深度。我们将此识别为解码器的一个重大且可直接纠正的瓶颈,并引入CRISP:一种像素空间扩散解码器,带有与主干无关的潜在适配器、基于DiT的去噪器以及支撑掩码预测器。CRISP可替代视频VAE和LiDAR原生解码器,同时保持编码器和潜在生成器固定不变。在KITTI-360、SemanticKITTI和nuScenes上,仅替换解码器即可在冻结主干上平均将FSVD/FPVD降低50.5%;对于通用视频VAE,降低幅度可达71%/74%。在LiDAR原生的LiDM主干上,FRID下降71%,且在深度不连续处的提升最大。在预训练的LiDM世界模型中,同样的零样本替换将FSVD改善15.5%,缩小了仿真到真实的差距。
cs.CL / 80 / 2610.11469
SAGE: Sink-Aware Guided Emphasis for Visual Grounding in Vision-Language Decoders
SAGE:面向视觉语言解码器中视觉定位的汇聚感知引导强调
Jeonghyo Song, YoungJoon Yoo
cs.CV · cs.CL
large language model
大语言模型相关
Abstract
Recent large vision-language models (VLMs) pair a visual encoder with a large language model (LLM) and perform well on diverse image-text tasks, yet their reliability is often limited by decoder attention pathologies that suppress visual evidence and exacerbate hallucinations. In this paper, we revisit visual attention sinks and uncover a structured, layer-dependent behavior: across prompts, early and late decoder layers exhibit prompt-invariant attention collapse onto the same few image regions, which we term PIS (Prompt-Invariant Sinks), whereas mid layers become prompt-conditioned and drive vision-language alignment. This split suggests that treating sinks as a uniform effect is incomplete. Building on this insight, we propose SAGE (Sink-Aware Guided Emphasis), a lightweight intervention that steers decoder attention away from PIS and toward query-dependent regions of interest (ROIs) using token-aligned ROI masks derived from standard vision backbones such as CLIP, ViT, and DINOv3. Evaluated on diverse vision-encoder + decoder-only LLM VLM families, SAGE improves visual grounding, reduces hallucinations, and yields consistent gains across public downstream vision-language benchmarks, including fine-grained visual discrimination settings where localized evidence is crucial, when instantiated with backbone-derived ROI masks.
Chinese Translation
近期的大型视觉语言模型(VLM)将视觉编码器与大型语言模型(LLM)配对,在各种图像-文本任务上表现良好,然而其可靠性常常受到解码器注意力病理的限制,这些病理会抑制视觉证据并加剧幻觉。在本文中,我们重新审视视觉注意力汇聚,并揭示了一种结构化的、依赖于层的行为:在不同提示下,早期和晚期解码器层表现出对相同少数图像区域的提示不变注意力坍缩,我们将其称为 PIS(提示不变汇聚),而中间层则变得受提示条件化并驱动视觉-语言对齐。这种分化表明,将汇聚视为一种均匀效应是不完整的。基于这一见解,我们提出了 SAGE(汇聚感知引导强调),这是一种轻量级干预,它利用从标准视觉骨干(如 CLIP、ViT 和 DINOv3)导出的 token 对齐 ROI 掩码,将解码器注意力从 PIS 引导开,转向依赖于查询的兴趣区域(ROI)。在多种视觉编码器 + 仅解码器 LLM 的 VLM 家族上进行评估,当使用骨干导出的 ROI 掩码实例化时,SAGE 改善了视觉定位,减少了幻觉,并在公共下游视觉-语言基准上取得了一致的增益,包括局部证据至关重要的细粒度视觉辨别设置。
cs.AI / 81 / 2610.11479
Conditional Residual Prediction: Improving Autoregressive Video Diffusion without a Bidirectional Teacher
条件残差预测:无需双向教师模型即可改进自回归视频扩散
Bowen Zheng, Zhiguang Liu, Jiarong Ou, Rui Chen, Tianyang Hu
cs.CV · cs.AI · cs.LG
diffusion
扩散模型相关
Abstract
Causal video diffusion models generate video autoregressively, which suits streaming, interactive, and long-video generation. Under standard training, however, they often yield lower generation quality than bidirectional models of the same size. Many existing approaches address this gap by initializing from or distilling a pretrained bidirectional teacher. We instead train a causal model from an image-model initialization, with no bidirectional video model at any stage. Because this path requires neither a large bidirectional teacher nor a complex distillation pipeline, it is simpler and more scalable. On this path, we find that a causal model trained on ground-truth history becomes strongly dependent on it, so that at inference errors in its own generated history propagate forward. We hypothesize that much of this dependence is unnecessary, because the current input already determines much of what the history provides. We propose Conditional Residual Prediction (CRP), a simple recipe for reducing a model's reliance on a condition: the model first predicts the target without the condition, and the condition may only add a residual on top of this prediction. Applied to history, CRP makes the model predict each chunk from the present as far as it can and use the past only for what the present cannot supply. In controlled experiments, CRP nearly closes the 6.14-point gap to a bidirectional model trained under the same setup. Scaling this recipe, we train Optica, a 2B-parameter causal video model that autoregressively generates 5-second 480p videos and reaches 82.78 on VBench with only about 15M training videos.
Chinese Translation
因果视频扩散模型以自回归方式生成视频,这适合流式、交互式以及长视频生成。然而,在标准训练下,它们往往比同等规模的双向模型产生更低的生成质量。许多现有方法通过从预训练的双向教师模型初始化或进行蒸馏来弥合这一差距。我们则不同,从图像模型初始化出发训练因果模型,在任何阶段都不使用双向视频模型。由于这条路径既不需要大型双向教师模型,也不需要复杂的蒸馏流程,因此它更简单、也更具可扩展性。在这条路径上,我们发现,在真值历史(ground-truth history)上训练的因果模型会对其产生强烈依赖,以至于在推理时,其自身生成历史中的误差会向前传播。我们假设,这种依赖在很大程度上是不必要的,因为当前输入已经决定了历史所提供的大部分信息。我们提出条件残差预测(Conditional Residual Prediction, CRP),这是一种降低模型对某个条件依赖的简单方法:模型首先在不使用该条件的情况下预测目标,而该条件只能在这一预测之上添加一个残差。将CRP应用于历史,模型会尽可能仅凭当前(present)预测每个视频块(chunk),并且只在当前无法提供信息时使用过去(past)。在受控实验中,CRP几乎弥合了与在相同设置下训练的双向模型之间6.14个点的差距。将这一方法扩展后,我们训练了Optica,一个具有20亿参数的因果视频模型,能够自回归地生成5秒、480p的视频,并且仅用约1500万个训练视频就在VBench上达到了82.78。
cs.CR / 82 / 2610.11496
Stop My Dancing! Understanding, Detecting and Attributing Motion-Aware Deepfake Videos
停止我的舞蹈!理解、检测和归因运动感知深度伪造视频
Fazhong Liu, Yan Meng, Tian Dong, Guoxing Chen, Haojin Zhu
cs.CV · cs.CR
diffusion
扩散模型相关
Abstract
Pose-guided diffusion models can now synthesize entire human figures in motion, spawning a new class of deepfakes: Motion Aware Deepfake (MAD) that have already reached hundreds of millions of viewers. To better understand this emerging threat, we construct the first MAD-specific benchmark and measurement framework, containing over 1.5 million frames that mix 1,363 real and 30,122 synthetic videos from six controllable generators, with realistic perturbations and open-world evaluation splits. Then, we dissect MAD and discover that, despite their global coherence, these videos betray faint yet reliable cues: because the model relies on limited input frames for motion synthesis, it must predict and simulate coherent movement at motion boundaries, thereby producing high-frequency artifacts along with model-specific spectral fingerprints. Based on the observations obtained from analysis on dataset, we propose MoDA, the first defense framework tailored to detect and attribute MAD videos. MoDA couples spatial semantics with steganalysis-rich frequency features via cross-domain alignment and multi-scale aggregation, achieving 94.8% in-distribution and 89.1% cross-dataset detection accuracy gains of 10% to 25% over prior work and 91.5% model attribution accuracy. MoDA achieves 81.94% accuracy on 200 clips produced by two unseen commercial MAD platforms, indicating promising zero-shot transfer, and 78.13% detection accuracy on 1,200 unseen MAD video clips (55k frames in total) collected from the open Internet. Under white-box, gray-box, and black-box adaptive attacks, MoDA maintains relatively stable detection and attribution performance while the accuracies of the baselines drop rapidly.
Chinese Translation
姿态引导扩散模型如今能够合成运动中的完整人体,催生了一类新的深度伪造:运动感知深度伪造(MAD),其已经触达数亿观众。为了更好地理解这一新兴威胁,我们构建了首个专门针对 MAD 的基准和测量框架,包含超过 150 万帧,混合了来自六个可控生成器的 1,363 个真实视频和 30,122 个合成视频,并带有真实扰动和开放世界评估划分。然后,我们剖析 MAD 并发现,尽管这些视频具有全局连贯性,它们仍会泄露微弱但可靠的线索:由于模型依赖有限的输入帧进行运动合成,它必须在运动边界处预测并模拟连贯运动,从而产生高频伪影以及模型特定的频谱指纹。基于对数据集分析得到的观察,我们提出 MoDA,首个专门用于检测和归因 MAD 视频的防御框架。MoDA 通过跨域对齐和多尺度聚合,将空间语义与富含隐写分析特征的频率特征结合起来,实现了 94.8% 的分布内和 89.1% 的跨数据集检测准确率,相较先前工作提升 10% 到 25%,以及 91.5% 的模型归因准确率。在由两个未见过的商业 MAD 平台生成的 200 个片段上,MoDA 达到 81.94% 的准确率,表明具有前景的零样本迁移,并在从开放互联网收集的 1,200 个未见过的 MAD 视频片段(总计 55k 帧)上达到 78.13% 的检测准确率。在白盒、灰盒和黑盒自适应攻击下,MoDA 保持相对稳定的检测和归因性能,而基线方法的准确率迅速下降。
cs.AI / 83 / 2610.11685
HI3D 3.0 (Twinkle3D): Object-specific 3D Asset Generation with High Resolution
HI3D 3.0(Twinkle3D):面向特定对象的高分辨率3D资产生成
Ziying Li, Shengchu Zhao, Huiang He, Yiyang Chen, Jianwen Huang, Bailin Li, Changhao Li, Jianhui Li, Jie Li, Ruiyang Liu, Yibo Luo, Tengjiao Sun, Pei Tang, Shiwen Wang, Jiaqi Wu, Kang Wu, Kaiqiao Yang, Zherui Yang, Hu Zhang, Xuezhi Zhao, Xinhe Zheng, Yukun Li, Heliang Zheng, Rongfei Jia
cs.CV · cs.AI
diffusion
扩散模型相关
Abstract
Image-to-3D generation has become increasingly capable of producing objects that closely resemble the input image, and an outstanding challenge is to reproduce the depicted object itself, including the specific geometry that defines it. Inscriptions, brand marks, and repeated structures are frequently distorted or lost, despite being critical to object identity. We present Hi3D 3.0, an image-to-3D generation system targeting object-specific fidelity, with Twinkle3D as its geometry model for generating watertight triangle meshes at $2048^{3}$ resolution. Twinkle3D advances high-fidelity geometry generation along four dimensions. First, while O-Voxel/FaithC offers high representational precision, it often suffers from poor surface quality and non-watertight geometry. We address both issues while retaining its $2048^{3}$-level precision. Second, we scale diffusion generation to sequences of up to 300K geometric tokens through a redesigned DiT architecture and large-scale distributed training optimizations, reducing training time per step from approximately ten minutes to ten seconds. Third, subsequent refinement cannot fully compensate for errors introduced during initial generation; we therefore strengthen both global shape and local detail in the initial generation stage, and the resulting single-stage model surpasses prior two-stage pipelines with $512^{3}$ refinement. Finally, we introduce a fine-grained image-3D cross-modal interaction mechanism that strengthens correspondence between visual evidence and geometric tokens, improving the recovery of object-specific structures. We evaluate geometric fidelity using alignment metrics derived from silhouettes and normal fields. Hi3D 3.0 outperforms four commercial systems across all reported metrics, recovering 82.1% of inscribed characters at 98.2% precision, compared with 21.7% recall for the strongest competitor.
Chinese Translation
图像到3D生成已越来越能够生成与输入图像高度相似的物体,而一个突出的挑战是复现所描绘的物体本身,包括定义它的特定几何结构。铭文、品牌标记和重复结构尽管对对象身份至关重要,却经常被扭曲或丢失。我们提出 Hi3D 3.0,一个面向特定对象保真度的图像到3D生成系统,并将 Twinkle3D 作为其几何模型,用于以 $2048^{3}$ 分辨率生成水密三角网格。Twinkle3D 在四个维度上推进了高保真几何生成。首先,尽管 O-Voxel/FaithC 提供了高表示精度,但它常常存在表面质量差和非水密几何的问题。我们在保留其 $2048^{3}$ 级精度的同时解决了这两个问题。其次,我们通过重新设计的 DiT 架构和大规模分布式训练优化,将扩散生成扩展到包含多达 300K 个几何 token 的序列,将每步训练时间从大约十分钟减少到十秒。第三,后续细化无法完全补偿初始生成期间引入的误差;因此,我们在初始生成阶段同时增强全局形状和局部细节,而由此得到的单阶段模型超越了先前具有 $512^{3}$ 细化的两阶段流水线。最后,我们引入一种细粒度的图像-3D跨模态交互机制,该机制加强了视觉证据与几何 token 之间的对应关系,从而改善对特定对象结构的恢复。我们使用从轮廓和法线场导出的对齐指标来评估几何保真度。Hi3D 3.0 在所有报告的指标上均优于四个商业系统,以 98.2% 的精确率恢复了 82.1% 的铭文字符,相比之下,最强竞争对手的召回率为 21.7%。
cs.AI / 84 / 2610.12104
VINCIE-NExT: Unlocking Video Editing from Images via In-Context Modeling
VINCIE-NExT:通过上下文建模从图像解锁视频编辑
Leigang Qu, Feng Cheng, Ziyan Yang, Bangbang Yang, Zhaoyang Huang, Wei Chow, Yicong Li, Wenjie Wang, Tat-Seng Chua, Yan Zeng
cs.CV · cs.AI · cs.MM
diffusion
扩散模型相关
Abstract
Building a capable video editor remains significantly harder than a video generator: editing requires (source, instruction, edited) triplets that are prohibitively expensive to annotate and difficult to synthesize at scale, whereas image editing has already reached maturity with millions of such pairs readily available. In this work, we introduce VINCIE-NExT, a unified framework that transfers editing capability from images to videos through in-context visual demonstrations, alleviating the need for large-scale paired video editing data. VINCIE-NExT decomposes video editing into a structured chain of composable sub-tasks (Video -> Image -> Image -> Video), routing editing intent through the image domain and enabling scalable joint training from heterogeneous image and video corpora under a unified diffusion objective. An image editing pair, synthesized by the model or supplied by the user, is prepended as an in-context visual demonstration that serves as a spatial appearance blueprint for every output frame. To ground appearance edits across the interleaved context, we introduce a novel position encoding that links image demonstrations and video frames in a shared spatial coordinate system, enabling pixel-faithful propagation of appearance changes to every output frame. Chain-of-Editing further provides principled test-time scaling: by executing the sub-task chain as progressive diffusion stages, editing quality can be improved by investing additional compute without retraining. Comprehensive experiments on OpenVE-Bench demonstrate the state-of-the-art performance across diverse editing categories, with ablations confirming the effectiveness of each component.
Chinese Translation
构建一个能力强大的视频编辑器仍然比构建一个视频生成器困难得多:编辑需要(source, instruction, edited)三元组,而这类三元组的标注成本高得令人望而却步,并且难以大规模合成;相比之下,图像编辑已经成熟,已有数百万个此类图像对可供直接使用。在这项工作中,我们提出了 VINCIE-NExT,一个统一的框架,它通过上下文视觉演示将编辑能力从图像迁移到视频,从而缓解了对大规模成对视频编辑数据的需求。VINCIE-NExT 将视频编辑分解为一条由可组合子任务构成的结构化链条(Video -> Image -> Image -> Video),将编辑意图经由图像域进行路由,并在统一的扩散目标下,实现从异构图像和视频语料中进行可扩展的联合训练。一个由模型合成或由用户提供的图像编辑对会被前置,作为上下文视觉演示,充当每一输出帧的空间外观蓝图。为了使外观编辑在交错上下文中得到锚定,我们引入了一种新颖的位置编码,它将图像演示和视频帧链接在一个共享的空间坐标系中,使外观变化能够以像素级忠实度传播到每一输出帧。Chain-of-Editing 进一步提供了有原则的测试时扩展:通过将子任务链作为渐进式扩散阶段来执行,可以在无需重新训练的情况下,通过投入额外计算来提升编辑质量。在 OpenVE-Bench 上的全面实验证明了其在多种编辑类别中达到最先进性能,消融实验也证实了每个组件的有效性。
cs.AI / 85 / 2610.12266
DVD: Dynamic Vector Decoding for Efficient MLLM-based Perception
DVD:面向高效基于MLLM的感知的动态向量解码
Jinghua Hou, Zhe Liu, Hengshuang Zhao
cs.CV · cs.AI
large language model
大语言模型相关
Abstract
Multimodal large language models have made remarkable progress in bridging vision and language, facilitating various perception tasks essential for human-machine interaction, robotics, and autonomous driving. However, existing MLLM-based perception methods predominantly rely on text-based coordinate representation, which suffers from excessive token overhead, or fixed-range quantization, which suffers from range and precision constraints, especially for 3D domains with unbounded spatial range and high localization accuracy requirements. To address these challenges, we propose a dynamic vector decoding method named DVD, which unifies the representation of 2D and 3D perception tasks. Specifically, we first transform diverse perceptual representation (i.e., 2D bounding boxes, 2D masks, and 3D bounding boxes) into 1D vector sequences, which are then mapped to compact discrete tokens in the high-dimensional space. Then, a lightweight de-tokenizer enables seamless integration with MLLMs by decoding output tokens back to original 2D and 3D perceptual representations. Extensive experiments on 2D and 3D perception benchmarks including RefCOCO series, SUN-RGBD, KITTI, Hypersim, nuScenes demonstrate that DVD achieves superior performance in 2D and 3D tasks and reduces significantly the token overhead and inference latency. DVD provides an efficient and general framework for integrating perception capabilities into MLLMs, overcoming the inherent limitations of existing methods.
Chinese Translation
多模态大语言模型在连接视觉与语言方面取得了显著进展,促进了对于人机交互、机器人学和自动驾驶至关重要的各类感知任务。然而,现有的基于MLLM的感知方法主要依赖于基于文本的坐标表示,这会导致过多的token开销,或者依赖于固定范围的量化,这会受到范围和精度的限制,尤其是对于具有无界空间范围和高定位精度要求的三维领域而言。为了解决这些挑战,我们提出了一种名为DVD的动态向量解码方法,它统一了二维和三维感知任务的表示。具体而言,我们首先将多样的感知表示(即二维边界框、二维掩码和三维边界框)转换为1D向量序列,然后将其映射到高维空间中的紧凑离散token。接着,一个轻量级的去分词器通过将输出token解码回原始的二维和三维感知表示,实现了与MLLM的无缝集成。在包括RefCOCO系列、SUN-RGBD、KITTI、Hypersim、nuScenes在内的二维和三维感知基准上进行的大量实验表明,DVD在二维和三维任务中取得了优越的性能,并显著降低了token开销和推理延迟。DVD为将感知能力集成到MLLM中提供了一个高效且通用的框架,克服了现有方法固有的局限性。
cs.CL / 86 / 2610.12417
WOVEN: Weaving Visual World Modeling into Multimodal LLMs
WOVEN:将视觉世界建模编织进多模态大语言模型
Zheyu Fan, Yue Zhang, Mingkai Deng, Kangrui Wang, Qineng Wang, Canyu Chen, Jie Hao, Xing Fan, Chenlei Guo, Eric P. Xing, Mohit Bansal, Manling Li
cs.CV · cs.CL · cs.LG
large language model
大语言模型相关
Abstract
Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from different supervision sources and reuse across different tasks, with a systematic training recipe. Existing benchmarks document these deficits separately but do not support controlled comparisons across scenes, actions, and reasoning operations. We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types. We first evaluate 38 frontier MLLMs (e.g., GPT-5.4 and Qwen3-VL-235B-A22B) and find a substantial and systematic deficit: even the strongest models fall far below humans, and the failures recur across model families and persist with scale. We then train MLLMs at multiple scales on WOVEN and find that they learn a shared capability that transfers broadly: training subsets of only about 2,000 items each collectively improve 22 of 26 external benchmarks by up to 27.3 percentage points, and WOVEN data can replace 30-50% of a task's own training data with comparable accuracy. Controlled comparisons further yield a training recipe for visual world modeling, validated prospectively on held-out benchmarks: select supervision by the reasoning operation it teaches rather than by the actions, scenes, or domains it shows, and prefer larger changes to the visual state for robustness. Our work establishes visual transition reasoning as a reusable foundation for systematic visual world-model training in MLLMs.
Chinese Translation
多模态大语言模型(MLLM)在空间、具身、物理和时间推理方面表现不佳。我们假设这些失败反映了视觉转换推理中存在的一种共同缺陷,并检验这种能力能否作为一种共享的训练基元——一种不同模型可以从不同监督来源中学习、并跨不同任务复用的基元——并配以一套系统化的训练配方。现有基准分别记录了这些缺陷,但不支持跨场景、动作和推理操作的受控比较。因此,我们提出了 WOVEN,一个面向视觉转换推理的训练来源与基准,它按场景、动作和推理类型来组织转换监督,使用来自视频预训练生成模型的多样、逼真的推演:涵盖 20 种场景类型、5 种动作类型和 8 种推理类型的 36,076 个示例。我们首先评估了 38 个前沿 MLLM(例如 GPT-5.4 和 Qwen3-VL-235B-A22B),发现存在显著且系统性的缺陷:即使最强的模型也远低于人类水平,而且这些失败在不同模型家族中反复出现,并随规模扩大而持续存在。随后,我们在 WOVEN 上训练了多个规模的 MLLM,发现它们学到了一种可广泛迁移的共享能力:每个仅约 2,000 条的训练子集共同将 26 个外部基准中的 22 个提升了最多 27.3 个百分点,并且 WOVEN 数据能够以相当的准确率替代某项任务自身 30-50% 的训练数据。受控比较进一步得出了视觉世界建模的训练配方,并在留出基准上进行了前瞻性验证:依据监督所教授的推理操作而非其所展示的动作、场景或领域来选择监督,并且为提升鲁棒性而偏好对视觉状态的更大变化。我们的工作确立了视觉转换推理作为 MLLM 中系统化视觉世界模型训练的可复用基础。
cs.CL / 87 / 2610.12427
FastBench: Can Streaming VLMs Perceive High-Dynamic Real-World Streams?
FastBench:流式 VLM 能否感知高动态的真实世界视频流?
Yuxuan Hu, Weikang Shi, Yang Bo, Xudong Lu, Xintong Guo, Shuhan Li, Yuyang He, Huankang Guan, Peiwen Sun, Yunqiao Yang, Wenbo Li, Rui Liu, Hongsheng Li
cs.CV · cs.CL
large language model
大语言模型相关
Abstract
Streaming Video Large Language Models (VLMs) enable continuous video understanding, yet existing benchmarks focus on low-dynamic scenarios. Under bounded context budgets, models must balance temporal history, spatial resolution, and temporal granularity; sparse sampling at 1--2 FPS misses fast events. We introduce FastBench to evaluate high-dynamic perception in real-world video streams. Its trajectory-grounded pipeline combines QA generation from high-FPS clips, filtering of questions answerable at 2 FPS, answer verification using SAM3 and CoTracker3 trajectories, and three rounds of human inspection. FastBench contains 306 QA pairs across eight domains, six capabilities, and forward, instant, and backward temporal scopes, with human-annotated evidence intervals. We also present ProactiveFrame, a training-free baseline that adjusts incoming frame rates through text tokens. A dual-tier sliding window retains recent high-FPS observations while downsampling older ones into sparse history. Experiments reveal substantial limitations: the strongest model, Gemini-3.5-Flash, scores only 50.7%. Denser sampling improves Qwen3-VL-8B from 32.9% at 2 FPS to 44.6% at 24 FPS, but gains saturate as history is compressed. ProactiveFrame outperforms sparse uniform sampling by 5.4 and 1.5 percentage points, yet remains well below oracle-guided focusing, showing that current VLMs struggle to determine from the stream alone when finer temporal perception is needed. FastBench provides a testbed for high-dynamic streaming video understanding. Code and data: https://github.com/Ashone3/FastBench.
Chinese Translation
流式视频大语言模型(VLMs)能够实现连续视频理解,然而现有基准主要关注低动态场景。在有限上下文预算下,模型必须平衡时间历史、空间分辨率和时间粒度;以 1--2 FPS 的稀疏采样会漏掉快速事件。我们提出 FastBench,用于评估真实世界视频流中的高动态感知。其以轨迹为依据的流水线结合了从高 FPS 片段生成 QA、过滤出在 2 FPS 下可回答的问题、使用 SAM3 和 CoTracker3 轨迹进行答案验证,以及三轮人工检查。FastBench 包含 306 个 QA 对,横跨八个领域、六种能力以及前向、即时和后向时间范围,并带有人工标注的证据区间。我们还提出 ProactiveFrame,一种通过文本 token 调整输入帧率的免训练基线。一种双层滑动窗口保留近期高 FPS 观测,同时将较旧的观测下采样为稀疏历史。实验揭示了显著局限:最强模型 Gemini-3.5-Flash 仅得分 50.7%。更密集的采样将 Qwen3-VL-8B 从 2 FPS 下的 32.9% 提升到 24 FPS 下的 44.6%,但随着历史被压缩,增益趋于饱和。ProactiveFrame 比稀疏均匀采样高出 5.4 和 1.5 个百分点,但仍远低于 oracle 引导的聚焦,表明当前 VLMs 难以仅凭流本身判断何时需要更细粒度的时间感知。FastBench 为高动态流式视频理解提供了一个测试平台。代码和数据:https://github.com/Ashone3/FastBench。
cs.AI / 88 / 2610.11025
Intent Graph: Navigating the Analytical Reasoning Space for Exploratory Data Analysis
意图图:为探索性数据分析导航分析推理空间
Junran Yang, Shruti Badrish, Teanna Barrett, Leilani Battle
cs.HC · cs.AI
large language model
大语言模型相关
Abstract
Exploratory data analysis (EDA) is rarely open-ended in practice: analysts work from high-level domain questions toward the concrete analyses that can answer them, prioritizing directions with domain knowledge and prior hypotheses. Large language models (LLMs) can supply such knowledge, but their responses are unstructured, leaving analysts no way to see what has been explored, what is missing, or why one direction was chosen over another. We present DAG-EDA, a system that lets analysts and an LLM co-navigate the space of possible analyses through two linked structures. An intent graph, governed by a grammar of analytical intent, decomposes an ambiguous natural-language question into progressively concrete analysis tasks, keeping alternative framings open and letting analysts branch, backtrack, and compare paths. A multi-layered knowledge graph externalizes the LLM's domain knowledge, linking domain concepts to the dataset variables that can measure them, so analysts can inspect and contest how their question is grounded in the data. Both graphs are constructed from only the dataset and the analyst's question, and the analyses the analyst reaches are rendered as interactive dashboards. We illustrate the system through a usage scenario and describe a user study design for examining whether the system scaffold analysts' reasoning and navigation.
Chinese Translation
在实践中,探索性数据分析(EDA)很少是完全开放式的:分析师从高层次的领域问题出发,走向能够回答这些问题的具体分析,并借助领域知识和先验假设来确定各方向的优先次序。大语言模型(LLM)能够提供这类知识,但其响应是非结构化的,使分析师无从看出已经探索了什么、还缺少什么,或者为什么选择了某个方向而非另一个方向。我们提出 DAG-EDA,一个让分析师与大语言模型通过两个相互关联的结构共同导航可能分析空间的系统。意图图受分析意图语法的支配,将一个模糊的自然语言问题分解为逐步具体的分析任务,使其他可能的表述框架保持开放,并让分析师能够分支、回溯和比较路径。一个多层知识图谱将 LLM 的领域知识外化,把领域概念与能够度量它们的数据集变量关联起来,使分析师能够检视并质疑其问题是如何在数据中得到落地的。这两个图都仅由数据集和分析师的问题构建而成,而分析师最终得到的分析结果被呈现为交互式仪表板。我们通过一个使用场景来说明该系统,并描述一项用户研究设计,以考察该系统是否能够为分析师的推理与导航提供支撑。
cs.AI / 89 / 2610.11031
Language Modeling is Monotone Compression
语言建模是单调压缩
Noam Mazor, Andrew Morgan, Rafael Pass
cs.IT · cs.AI · cs.CR
large language model
大语言模型相关
Abstract
A long-standing hypothesis in artificial intelligence and neuroscience posits that intelligence is closely related to compression: the ability to compress information efficiently intuitively reflects capacities associated with intelligence and learning. Indeed, recent experimental works verify this intuition by showing connections between the capabilities of large language models (LLMs) and their ability as compressors: for instance, Deletang et al. (ICLR'24) demonstrate that LLMs can be used as powerful compressors, and Huang et al. (COLM'24) show that the compression ability of LLMs is highly correlated with their performance on benchmarks for knowledge and reasoning. In this work, we initiate a theoretical study of this connection. Our main result is that LLMs (formally modeled as next-token predictors) are equivalent to monotone (a.k.a. order-preserving) compression algorithms---namely, compression algorithms where the encoding process preserves the ordering of the inputs---in the sense that the one can be constructed from the other while preserving the same error up to an additive gap of 2. We next show that the monotonicity is required for this equivalence to hold if and only if cryptographic (infinitely-often) one-way functions exist. As a direct corollary, we get a cryptographic result of independent interest: the notion of next-bit pseudoentropy (a computational analogue of entropy) of a distribution is equivalent to monotone incompressibility of the distribution. (Previously, it was only known (Haitner et al., ITCS'23) that incompressibility implies next-bit pseudoentropy.)
Chinese Translation
人工智能和神经科学中一个长期存在的假说认为,智能与压缩密切相关:高效压缩信息的能力直观上反映了与智能和学习相关的能力。事实上,最近的实验工作通过展示大语言模型(LLMs)的能力与其作为压缩器的能力之间的联系,验证了这一直觉:例如,Deletang 等人 (ICLR'24) 证明 LLMs 可以被用作强大的压缩器,而 Huang 等人 (COLM'24) 表明 LLMs 的压缩能力与其在知识和推理基准上的表现高度相关。本文中,我们开启了对这种联系的理论研究。我们的主要结果是,LLMs(被形式化建模为下一词元预测器)等价于单调(又称保序)压缩算法——即编码过程保持输入顺序的压缩算法——其意义在于,其中一个可以由另一个构造出来,同时保持相同误差,至多相差一个加性间隙 2。我们接下来表明,单调性是这一等价成立的必需条件,当且仅当密码学(无限经常)单向函数存在。作为一个具有独立意义的密码学结果的直接推论,我们得到:一个分布的下一比特伪熵(熵的一种计算类比)概念等价于该分布的单调不可压缩性。(此前,人们只知道 (Haitner 等人, ITCS'23) 不可压缩性蕴含下一比特伪熵。)
cs.LG / 90 / 2610.10809
Evaluating Rubric Generation with Interventional Transfer
用干预性迁移评估评分标准生成
Erik Skalnes, Layne C. Price, Raviteja Anantha, Michael Oberst
cs.LG
large language model
大语言模型相关
Abstract
Instance-specific rubrics are common in AI benchmarks where reliable evaluation requires specific expert knowledge. This approach is difficult to scale, prompting research into the generation of rubrics with large language models (LLMs). However, even when expert rubrics are available as references, it is unclear how to productively evaluate the quality of generated rubrics at scale. In this paper, we introduce a method for the evaluation of rubric generation, which we call Interventional Transfer (IT), based on the idea that two rubrics are similar if they move together when a response is perturbed to pass/fail one of them. In contrast to existing approaches for evaluation of rubric generation, we argue that different forms of interventional transfer can be used to evaluate the utility of generated rubrics for different tasks. For instance, we apply this approach in a case study on HealthBench, where we demonstrate an asymmetry in rubrics generated by Qwen3.8-27B, Deepseek-V4-Flash, and Opus-5, used to evaluate responses from GPT-5.6-Terra. Perturbations that degrade responses according to the generated/expert rubric transfer into lower scores on the corresponding expert/generated rubric, but perturbations that improve on one rubric do not reliably transfer into higher scores on the other. We argue that this finding has implications for the usage of LLM-generated rubrics for performance monitoring and hill-climbing. We contrast our approach with existing approaches for rubric evaluation, which do not surface the same asymmetry that we observe.
Chinese Translation
实例特定的评分标准在AI基准测试中很常见,因为可靠的评估需要特定的专家知识。这种方法难以规模化,促使人们研究使用大语言模型(LLM)生成评分标准。然而,即使专家评分标准可作为参考,如何大规模有效地评估生成评分标准的质量仍不清楚。在本文中,我们提出一种用于评估评分标准生成的方法,我们称之为干预性迁移(Interventional Transfer,IT),其思想是:如果当对回答进行扰动以使其通过/不通过其中一个评分标准时,两个评分标准会一起变化,那么这两个评分标准就是相似的。与现有的评分标准生成评估方法相比,我们认为不同形式的干预性迁移可用于评估生成评分标准对不同任务的效用。例如,我们在HealthBench上进行案例研究并应用该方法,其中我们展示了由Qwen3.8-27B、Deepseek-V4-Flash和Opus-5生成的评分标准中的一种不对称性,这些评分标准用于评估来自GPT-5.6-Terra的回答。根据生成/专家评分标准使回答变差的扰动,会迁移为在对应的专家/生成评分标准上更低的分数;但在一个评分标准上带来改善的扰动,并不能可靠地迁移为在另一个评分标准上更高的分数。我们认为,这一发现对于将LLM生成的评分标准用于性能监测和爬山法具有启示意义。我们将我们的方法与现有的评分标准评估方法进行对比,这些方法并未揭示出我们所观察到的相同不对称性。
cs.LG / 91 / 2610.10854
KDFP: A first-principles approach to knowledge distillation in large language models
KDFP:大语言模型知识蒸馏的第一性原理方法
Ryan Swift, Konstantinos Psounis
cs.LG
large language model
大语言模型相关
Abstract
Knowledge distillation is an established technique for improving the capabilities of small, efficient student models by training them with the representations of larger, more capable teacher models. Much of the recent work in the distillation of large language models (LLMs) has focused on distilling abilities learned during post-training, such as instruction following, chain-of-thought reasoning, and tool usage. This has left a large research gap in general knowledge distillation for LLMs, which is essential for developing efficient and private systems suitable for deployment on edge devices. We take a first-principles approach, evaluating previous lessons from prior works and conducting new explorations to develop a distillation methodology suitable for modern LLMs. We present KDFP, a novel methodology for white-box general knowledge distillation in LLMs. We demonstrate that KDFP outperforms existing methods by 1.6% $-$ 4.9% across 9 benchmarks while increasing training efficiency by up to 99.1% through ephemeral parameter reduction.
Chinese Translation
知识蒸馏是一种成熟的技术,通过使用更大、能力更强的教师模型的表示来训练小型、高效的学生模型,从而提升其能力。近期关于大语言模型(LLM)蒸馏的许多工作都集中在蒸馏后训练阶段学到的能力,例如指令遵循、思维链推理和工具使用。这在面向 LLM 的通用知识蒸馏中留下了一个很大的研究空白,而通用知识蒸馏对于开发适合在边缘设备上部署的高效且私密的系统至关重要。我们采用第一性原理方法,评估先前工作中的经验教训,并进行新的探索,以开发一种适用于现代 LLM 的蒸馏方法。我们提出 KDFP,一种用于 LLM 中白盒通用知识蒸馏的新方法。我们证明,KDFP 在 9 个基准上以 1.6% $-$ 4.9% 的优势优于现有方法,同时通过临时参数缩减将训练效率提升至多 99.1%。
cs.LG / 92 / 2610.10875
Barron Optimal Transport I: Generative Modeling
Barron 最优传输 I:生成建模
Evan Dogariu, Joan Bruna
cs.LG · math.AP
diffusion
扩散模型相关
Abstract
Motivated by recent applications in generative modeling and sampling, we introduce a framework for optimal measure transport where cost captures the notion of neural network complexity. In transport-based generative models, samples from a reference distribution (e.g. Gaussian) are mapped to samples of a target distribution along ordinary or stochastic differential equations. These are implemented as deep residual networks when discretized in time, where each hidden layer approximates the associated instantaneous velocity. Thus, given a pair of target and reference measures, a natural question is to search for the most efficient neural representation that implements this transport. Our starting point is the kinetic formulation of OT, due to Benamou and Brenier. We replace the average kinetic $L^2$ energy by the \emph{Barron} energy \cite{bach2017breaking, ma2022barron}, a natural norm which measures the complexity of representing a given vector field with a neural hidden layer, and which captures the adaptive properties of feature learning. This defines a metric on the space of probability measures, complementing existing Wasserstein and Stein geometries. In this work we examine the properties of this metric in the context of generative modeling. As a first application, we quantify the suboptimality of diffusion generative modeling in the Barron geometry by establishing super-polynomial score approximation lower bounds for data generated by neural network pushforwards of the Gaussian. We then investigate the benefit of adaptivity as a way to study alternative generative models. In a companion paper \cite{companionpaper} we leverage the Barron transport geometry for sampling applications, extending the scope of Stein variational gradient methods via feature adaptation.
Chinese Translation
受生成建模与采样中近期应用的启发,我们引入一个最优测度传输框架,其中代价刻画了神经网络复杂度的概念。在基于传输的生成模型中,来自参考分布(例如高斯分布)的样本沿常微分方程或随机微分方程被映射为目标分布的样本。当在时间上离散化时,这些方程被实现为深度残差网络,其中每个隐藏层近似相应的瞬时速度。因此,给定一对目标测度和参考测度,一个自然的问题是寻找实现该传输的最高效神经表示。我们的出发点是 Benamou 和 Brenier 提出的 OT 的动能表述。我们将平均动能 $L^2$ 能量替换为 \emph{Barron} 能量 \cite{bach2017breaking, ma2022barron},这是一种自然的范数,它度量用神经隐藏层表示给定向量场的复杂度,并刻画特征学习的自适应性质。这在概率测度空间上定义了一个度量,补充了现有的 Wasserstein 几何和 Stein 几何。在这项工作中,我们在生成建模的背景下考察该度量的性质。作为第一个应用,我们通过为高斯分布的神经网络前推生成的数据建立超多项式得分近似下界,来量化扩散生成建模在 Barron 几何中的次优性。然后我们研究自适应性带来的益处,以此作为研究替代生成模型的一种方式。在一篇配套论文 \cite{companionpaper} 中,我们将 Barron 传输几何用于采样应用,通过特征自适应扩展了 Stein 变分梯度方法的适用范围。
cs.LG / 93 / 2610.10965
Spectrally Targeted Muon
谱靶向 Muon
Vishrut Goyal, Rohan Ramkumar
cs.LG
large language model
大语言模型相关
Abstract
The Muon optimizer orthogonalizes each update matrix, setting all of its singular values to one, and has proven highly effective for training large language models. It remains unclear, however, whether this success comes from amplifying small singular directions that gradient descent neglects or from suppressing large, degenerate directions that disrupt training. We introduce Spectrally Targeted Muon, which orthogonalizes only the singular values above or below a threshold $τ$, so that varying $τ$ interpolates between normalized SGD and Muon. It isolates the relevant singular subspaces with projections computed by Newton-Schulz iteration on a shifted Gram matrix, so no SVD is needed. We evaluate these variants on the CIFAR-10 and NanoGPT speedruns, tracking the effective rank of gradient, update, and weight matrices and a new metric, the alignment of updates with the tangent space of the weight matrix's isospectral manifold. We find three things. First, the small singular values of the momentum are not noise. Orthogonalizing everything except the few largest singular values of each matrix nearly matches Muon while touching only a small fraction of the momentum, whereas orthogonalizing only the top falls well short even though it holds almost all of it. On language models every momentum singular value is far below one, so targeted orthogonalization can only amplify, and Muon wins by making directions that are too small to train on at their raw scale trainable. Second, shrinking the largest singular values is what keeps the parameter spectrum flat; this is cheap, and it is not what drives the loss. Third, AdamW differs from Muon mainly in how slowly it builds structure, which explains its slower start and why warmup helps AdamW but only hurts Muon.
Chinese Translation
Muon 优化器对每个更新矩阵进行正交化,将其所有奇异值设为 1,并已被证明在训练大语言模型时极为有效。然而,目前尚不清楚这种成功究竟源于放大了梯度下降所忽略的小奇异方向,还是源于抑制了会扰乱训练的大而退化的方向。我们提出谱靶向 Muon,它仅对高于或低于阈值 $τ$ 的奇异值进行正交化,因此改变 $τ$ 可在归一化 SGD 与 Muon 之间插值。它通过在一个平移后的 Gram 矩阵上以 Newton-Schulz 迭代计算得到的投影来分离出相关的奇异子空间,因此无需 SVD。我们在 CIFAR-10 和 NanoGPT 的 speedrun 上评估这些变体,追踪梯度、更新与权重矩阵的有效秩,以及一个新指标,即更新与权重矩阵等谱流形切空间的对齐程度。我们发现三点。第一,动量的小奇异值并非噪声。对每个矩阵中除少数最大奇异值之外的所有部分进行正交化,几乎可以达到与 Muon 相当的效果,而只触及动量的一小部分;相比之下,仅对最大的那些奇异值进行正交化则远远达不到,尽管后者占据了几乎全部动量。在语言模型上,每个动量奇异值都远小于 1,因此靶向正交化只能起放大作用,而 Muon 之所以取胜,是因为它使那些在原始尺度下小到无法用于训练的方向变得可训练。第二,压缩最大的奇异值才是保持参数谱平坦的原因;这一操作代价低廉,而且它并不是驱动损失的因素。第三,AdamW 与 Muon 的主要差异在于它构建结构的缓慢程度,这解释了它起步较慢的原因,也解释了为什么预热对 AdamW 有帮助,却只会损害 Muon。
cs.LG / 94 / 2610.10969
ASPIRE: Saddle-Point Discovery through Set Prediction and Physical Refinement
ASPIRE:通过集合预测与物理精修发现鞍点
Yucheng Zhao, Quanyou Zhang, Shaoxiang Qin, Haixuan Xu, Xiongye Xiao
cs.LG
diffusion
扩散模型相关
Abstract
Predicting thermally activated diffusion and defect evolution with event-driven models requires identifying atomic rearrangement mechanisms and their activation barriers. Discovering the associated saddle points is a major computational bottleneck: multiple rearrangements may originate from one metastable state, while costly local searches can fail or repeatedly converge to the same saddle. To address this challenge, we introduce ASPIRE (Atomistic Saddle-Point Inference with Refinement for Events), a framework that predicts a set of saddle candidates from a single initial atomic environment and refines them through Dimer searches on the original interatomic potential. The framework's equivariant set predictor, Ev-Quiformer, integrates (i) geometry-conditioned scalar-vector event slots for generating multiple saddle-point proposals and (ii) a decoder that maps each slot to a full atomic displacement field by combining atom, slot, and anchor-relative vectors with invariant coefficients. We also contribute two datasets: (i) BCCFE4VACAV-4000, comprising 4,000 four-vacancy body-centered cubic iron configurations and 65,450 reference events grouped by initial state for set supervision and post-refinement evaluation; and (ii) BCCFE-1TO4VAC, comprising 5,372 configurations with one to four vacancies each. Theoretically, we establish conditions for proposal equivariance. Experimentally, ASPIRE achieves 77.20% reference-event coverage on this benchmark, compared with 75.73% for a conventional Dimer baseline, while requiring approximately half as many Dimer force evaluations. In a timing evaluation on 50 configurations, ASPIRE reduces wall time per configuration from 478.8 s to 176.3 s under the stated hardware settings.
Chinese Translation
使用事件驱动模型预测热激活扩散和缺陷演化,需要识别原子重排机制及其激活势垒。发现相关联的鞍点是主要的计算瓶颈:多个重排可能源于一个亚稳态,而代价高昂的局部搜索可能失败或反复收敛到同一个鞍点。为应对这一挑战,我们提出 ASPIRE(面向事件的原子鞍点推断与精修),这是一个框架,它从单个初始原子环境预测一组鞍点候选,并通过对原始原子间势进行 Dimer 搜索来精修它们。该框架的等变集合预测器 Ev-Quiformer 集成了 (i) 几何条件化的标量-向量事件槽,用于生成多个鞍点提议,以及 (ii) 一个解码器,该解码器通过将原子、槽和锚点相对向量与不变系数相结合,将每个槽映射到完整原子位移场。我们还贡献了两个数据集:(i) BCCFE4VACAV-4000,包含 4,000 个四空位体心立方铁构型以及 65,450 个按初始状态分组的参考事件,用于集合监督和精修后评估;以及 (ii) BCCFE-1TO4VAC,包含 5,372 个各具有一到四个空位的构型。理论上,我们建立了提议等变性的条件。实验上,ASPIRE 在此基准上实现了 77.20% 的参考事件覆盖率,而传统 Dimer 基线为 75.73%,同时所需的 Dimer 力评估次数大约只有其一半。在针对 50 个构型的计时评估中,在所述硬件设置下,ASPIRE 将每个构型的墙钟时间从 478.8 秒降至 176.3 秒。
cs.LG / 95 / 2610.10974
Budgeted Multi-Source Counterfactual Annotation for Off-Policy Evaluation
面向离策略评估的预算受限多源反事实标注
Biao Xiang, Ali Eshragh, Yuexing Li, Kai Wang
cs.LG · cs.AI · math.OC
large language model
大语言模型相关
Abstract
Off-policy evaluation (OPE) estimates the value of a target policy from logged data, but limited behavior-policy coverage can force high-variance reweighting or reward-model extrapolation. Counterfactual annotations can add evidence about unobserved actions, yet practical sources, including domain experts and large language models (LLMs), may be costly, biased, or noisy. We study budgeted acquisition of such annotations for contextual-bandit OPE. Given source-specific costs and error profiles, we formulate an integer allocation problem over context-action pairs and annotation sources to minimize the component of estimator variance that depends on the annotation plan. We characterize when annotations are valuable through a first-annotation threshold and local annotation-value regimes. For the coupled multi-source problem, we develop a majorization-minimization algorithm with dynamic-programming subroutines that monotonically improves the objective. Experiments in synthetic clinical and LLM-annotated education bandits show that our allocation method reduces fixed-profile mean squared error (MSE) by 20.58% and 10.77%, respectively, relative to no annotation.
Chinese Translation
离策略评估(OPE)从日志数据中估计目标策略的价值,但行为策略覆盖有限可能迫使人们采用高方差的重加权或奖励模型外推。反事实标注能够为未观测到的动作补充证据,然而实际可用的来源——包括领域专家和大语言模型(LLM)——可能代价高昂、存在偏差或含有噪声。我们研究面向上下文赌博机 OPE 的此类标注的预算受限获取问题。在给定各来源特定的成本与误差特征的情况下,我们将问题形式化为一个在上下文-动作对与标注来源上的整数分配问题,以最小化估计量方差中依赖于标注方案的那个分量。我们通过首次标注阈值以及局部标注价值区间,刻画了标注在何时是有价值的。针对这一耦合的多源问题,我们提出了一种带有动态规划子程序的主化-最小化算法,该算法能够单调地改进目标函数。在合成临床赌博机以及由 LLM 标注的教育赌博机上的实验表明,相对于无标注的情形,我们的分配方法使固定配置下的均方误差(MSE)分别降低了 20.58% 和 10.77%。
cs.LG / 96 / 2610.10975
Optimizing Large Language Models with Chained LMOs
使用链式 LMO 优化大型语言模型
Sungyoon Kim, Kaan Ozkara, Youngsuk Park
cs.LG · cs.AI
large language model
大语言模型相关
Abstract
Muon has motivated a growing family of optimizers that compose multiple matrix normalizations, but these methods remain fragmented and lack a unified perspective. We introduce chained linear minimization oracles (chained LMOs), which cast these methods as compositions of LMOs. Despite their empirical success, many chains fall outside the standard LMO framework and can diverge on smooth convex objectives. To explain why composition can nevertheless help, we turn to linear associative memory and show that chaining can improve over Muon under anisotropic embeddings. Empirically, we propose TensorChain, a novel optimizer within the framework that stacks compatible weight matrices across different layers and normalizes the 3d tensor across its axes. In Qwen3 0.6B and 1.7B pretraining, TensorChain outperforms all chained baselines in average token efficiency, with average token savings of 9.6% over Muon at matched validation loss.
Chinese Translation
Muon 激发了一个不断壮大的优化器家族,这些优化器组合多种矩阵归一化,但这些方法仍然零散且缺乏统一视角。我们引入链式线性最小化预言机(chained LMOs),这些预言机将这些方法表述为 LMO 的复合。尽管它们在经验上取得了成功,许多链仍落在标准 LMO 框架之外,并且在光滑凸目标上可能发散。为了解释为何复合仍然可能有帮助,我们转向线性联想记忆,并表明在各向异性嵌入下,链式方法可以优于 Muon。在实证上,我们提出了 TensorChain,这是该框架内的一种新型优化器,它堆叠不同层中兼容的权重矩阵,并沿其各轴对三维张量进行归一化。在 Qwen3 0.6B 和 1.7B 预训练中,TensorChain 在平均 token 效率上优于所有链式基线,在匹配验证损失下相比 Muon 平均节省 9.6% 的 token。
cs.LG / 97 / 2610.10989
Multi-Bandwidth Distribution Matching Distillation: On the Equivalence of Distribution Matching Distillation and Drifting Models
多带宽分布匹配蒸馏:论分布匹配蒸馏与漂移模型的等价性
Jialin Zhu, Xing Liu, Feixiang He, He Wang
cs.LG · cs.CV
diffusion
扩散模型相关
Abstract
Researchers are exploring effective one-step generative model continuously, and, Drifting Models (Deng et al., 2026), demonstrate great potential in one-step generation recently. There are works that reveal the connection between Diffusion & Flow Style Generative Models (DFSGMs) (Ho et al., 2020; Song et al., 2020a;b; Lipman et al., 2022; Liu et al., 2022) and Drifting Models (Li & Zhu, 2026; Lai et al., 2026; Turan et al., 2026). But no one has yet established a precise correspondence between the Drifting Model and the widely used distillation method- Distribution Matching Distillation (DMD/DMD2) (Yin et al., 2024b;a) to the best of our knowledge, even though their optimization objective formulas are virtually identical. In this paper, we prove that by converting the velocity-field / noise-field from the pre-trained DFSGMs into the attraction force field in Drifting Models and estimating the repulsion force field from the generative distribution, training the Drifting Model is naturally equivalent to the Distribution Matching Distillation. With this equivalent concept, we propose an improved method based on DMD from the Drifting Model's perspective- Multi-Bandwidth Distribution Matching Distillation (MBDMD).
Chinese Translation
研究人员正在持续探索有效的一步生成模型,而漂移模型(Deng et al., 2026)近来在一步生成中展现出巨大潜力。已有工作揭示了扩散与流风格生成模型(DFSGMs)(Ho et al., 2020; Song et al., 2020a;b; Lipman et al., 2022; Liu et al., 2022)与漂移模型(Li & Zhu, 2026; Lai et al., 2026; Turan et al., 2026)之间的联系。但据我们所知,尚无人建立漂移模型与广泛使用的蒸馏方法——分布匹配蒸馏(DMD/DMD2)(Yin et al., 2024b;a)——之间的精确对应关系,尽管它们的优化目标公式几乎完全相同。在本文中,我们证明,通过将预训练 DFSGMs 中的速度场/噪声场转换为漂移模型中的吸引力场,并从生成分布中估计排斥力场,训练漂移模型自然等价于分布匹配蒸馏。基于这一等价概念,我们从漂移模型的视角提出一种基于 DMD 的改进方法——多带宽分布匹配蒸馏(MBDMD)。
cs.LG / 98 / 2610.11063
Emergent Inverse-Depth Scaling From Nonlinearity In Attention
注意力中非线性所涌现的反深度缩放
Zirui Peng, Yizhou Liu, Ziming Liu, Jeff Gore
cs.LG · cs.AI
large language model
大语言模型相关
Abstract
Scaling laws describe power-law improvements in model performance with dataset size and parameter count, yet their underlying mechanisms are not fully understood. To explain the parameter count scaling, existing theory posits power-law scaling with model depth. In linear-attention models, this scaling is tied to a power-law data spectrum: unable to selectively attend to relevant tokens, these models learn according to global spectral strength, with stronger directions learned before weaker ones. Large language models, however, can be strongly nonlinear. Here, we show that nonlinear attention yields inverse-depth decay of loss across all tested data spectra. Nonlinearity enables attention to focus selectively on relevant tokens, allowing strong and weak spectral directions to be learned in parallel. Similar focusing across layers motivates a connection to the central limit theorem: shared error across layers sets the loss plateau, while aggregation turns layer-specific differences into continued gains with depth. Our findings suggest that depth scaling may arise from nonlinearity in attention, which allows large language models to focus locally and may make the global covariance structure less relevant.
Chinese Translation
缩放律描述了模型性能随数据集规模和参数数量呈幂律改进的现象,然而其底层机制尚未被完全理解。为解释参数数量的缩放,现有理论假设其随模型深度呈幂律缩放。在线性注意力模型中,这种缩放与幂律的数据谱相关联:由于无法选择性地关注相关的 token,这些模型依据全局谱强度进行学习,较强的方向先于较弱的方向被学到。然而,大型语言模型可能具有强烈的非线性。在此,我们表明非线性注意力在所有测试过的数据谱上都产生损失的随深度反比衰减。非线性使注意力能够选择性地聚焦于相关的 token,从而使强与弱的谱方向可以被并行地学习。各层之间类似的聚焦促使我们将其与中心极限定理联系起来:层间共享的误差决定了损失平台,而聚合则将各层特有的差异转化为随深度持续增加的收益。我们的发现表明,深度缩放可能源于注意力中的非线性,它使大型语言模型能够进行局部聚焦,并可能使全局协方差结构变得不那么重要。
cs.LG / 99 / 2610.11216
The Lattice of Transition Laws
转移律的格
T. Y. Tsui, Jiatao Gu, Lingjie Liu
cs.LG · cs.CL
diffusion
扩散模型相关
Abstract
Diffusion and autoregression (AR) have long been seen as different categories of generative models, with diffusion specialising in continuous fields and AR specialising in discrete tokens. Recent work seeks to combine the advantages of the two models, and each hybrid fixes its decoding schedule by design. In this paper, we ask whether the performance of decoding schedules of one model can be predicted before decoding at a fixed number of steps. We describe diffusion, AR, and models in between as paths on one corruption lattice, and define the cost of a schedule as the dependence its parallel steps discard. The cost shows that the fewest steps of a zero-cost schedule are set by the geometry of the data, in the same way for tokens and for continuous fields. In particular, for data that are Markov on a graph and dependent along its paths, the fewest steps equal the graph's treedepth, which is logarithmic in the length of a sequence and linear in the side length of a grid. With fewer steps than the treedepth, every schedule pays a positive cost, whose ranking we predict before decoding with a kernel of pairwise dependence estimated from pretrained weights. Across text generation, image generation, and video generation, we verify most of the predictions about the rankings of different schedules under different metrics and benchmarks. This work therefore provides a design principle for decoding for future AR models, diffusion models, and anything in between. Our code is available at https://github.com/TSUITUENYUE/The-Lattice-of-Transition-Laws.
Chinese Translation
扩散与自回归(AR)长期以来被视为两类不同的生成模型,其中扩散专精于连续场,而 AR 专精于离散词元。近期的工作试图结合这两类模型的优势,而每一种混合模型都通过设计固定其解码调度。在本文中,我们追问:一个模型的解码调度的性能,能否在固定步数下进行解码之前就被预测出来。我们将扩散、AR 以及介于两者之间的模型描述为同一个破坏格上的路径,并将一个调度的代价定义为其并行步骤所丢弃的依赖。这一代价表明,零代价调度的最少步数由数据的几何结构所决定,对于词元和连续场而言方式相同。特别地,对于在某个图上具有马尔可夫性、并沿其路径存在依赖的数据,最少步数等于该图的树深,它对序列长度是对数级的,对网格边长是线性的。当步数少于树深时,每一个调度都要付出正的代价,我们用一个由预训练权重估计出的成对依赖核,在解码之前预测这些代价的排序。在文本生成、图像生成和视频生成中,我们验证了关于在不同指标与基准下不同调度排序的大部分预测。因此,这项工作为未来的 AR 模型、扩散模型以及介于两者之间的任何模型提供了解码的设计原则。我们的代码可在 https://github.com/TSUITUENYUE/The-Lattice-of-Transition-Laws 获取。
cs.LG / 100 / 2610.11228
Multimodal Graph Retrieval-Augmented Sequential Recommendation via Collaborative Filtering Paths
基于协同过滤路径的多模态图检索增强序列推荐
Jason Marcell Setiadi, Xin Cao, Lina Yao
cs.LG
large language model
大语言模型相关
Abstract
Multimodal Large Language Models (MLLMs) have demonstrated strong potential for sequential recommendation through their ability to reason over complex multimodal data. However, existing approaches either rely solely on the target user's own interaction history, neglecting collaborative signals from neighboring users, or incur substantial computational overhead through repeated MLLM inference over long interaction histories. To address these challenges, we propose MGRASRec, a multimodal graph retrieval-augmented framework for sequential recommendation. MGRASRec injects collaborative filtering signals conditioned on the candidate item directly into the MLLM prompt by retrieving structured paths from a user-item interaction graph, extended via multimodal similarity to increase coverage beyond exact co-interaction overlap. This retrieval also surfaces the history items most relevant to the candidate at no additional cost, removing the need for recurrent summarization and keeping inference to a single forward pass per candidate. All components are unified into an augmented prompt for parameter-efficient fine-tuning of an MLLM. Extensive evaluations across three publicly available datasets validate the effectiveness of MGRASRec, achieving the best performance on all metrics with particularly strong gains in ranking quality.
Chinese Translation
多模态大语言模型(MLLMs)通过对复杂多模态数据进行推理的能力,已在序列推荐中展现出强大潜力。然而,现有方法要么仅依赖目标用户自身的交互历史,忽略了来自邻近用户的协同信号,要么因在长交互历史上重复进行 MLLM 推理而带来大量计算开销。为应对这些挑战,我们提出 MGRASRec,一个用于序列推荐的多模态图检索增强框架。MGRASRec 通过从用户-物品交互图中检索结构化路径,将以候选物品为条件的协同过滤信号直接注入 MLLM 提示中,并通过多模态相似度进行扩展,以将覆盖范围扩大到超出精确共同交互重叠的范围。这种检索还以无额外成本的方式呈现出与候选物品最相关的历史物品,从而消除了对循环式摘要的需求,并将每个候选物品的推理保持为单次前向传播。所有组件被统一为一个增强提示,用于对 MLLM 进行参数高效微调。在三个公开可用数据集上的大量评估验证了 MGRASRec 的有效性,其在所有指标上均取得最佳性能,并在排序质量上获得了尤其显著的提升。
cs.LG / 101 / 2610.11281
How to post-train on a surrogate: Envelope sampling mitigates reward hacking
如何在代理上进行后训练:包络采样缓解奖励黑客行为
Sanjit Dandapanthula, Shuvom Sadhuka, Samir Khan, Michael Oberst, Aaditya Ramdas, Alexandra Chouldechova
cs.LG · cs.AI · stat.ME
large language model
大语言模型相关
Abstract
Large language models (LLMs) are commonly post-trained against LLM judges and other cheap surrogates because the true reward, such as human preference, is too expensive to query at scale. This practice often leads to reward hacking, where reinforcement learning against a miscalibrated surrogate leads to undesirable side effects. In this work, we study a setting in which a small number $n$ of model outputs are annotated with ground-truth labels (e.g., from expert review) and used to recalibrate the LLM judge before optimizing against it. Prior approaches to judge recalibration are costly or heuristic, and it is known that on-policy sampling fails when the surrogate is miscalibrated on a rare set of outputs. In this work, we propose envelope sampling, a theoretically-grounded method for judge recalibration that seeks to minimize an upper bound on the regret of the post-trained model under the assumption that the human reward and re-calibrated reward lie in an $L^2$ ball around the judge. We give practical algorithms to sample from the envelope by rejection or by fine-tuning against a modified reward, and experiments on clinical note generation and on a controlled sycophancy task show that recalibrating on envelope samples mitigates reward hacking where recalibrating on base-model samples does not.
Chinese Translation
大型语言模型(LLMs)通常针对 LLM 评判器和其他廉价代理进行后训练,因为真实奖励(例如人类偏好)在大规模查询时过于昂贵。这种做法常常导致奖励黑客行为,即针对校准不当的代理进行强化学习会带来不良副作用。在这项工作中,我们研究这样一种设定:少量 $n$ 个模型输出被标注了真实标签(例如来自专家评审),并在针对 LLM 评判器进行优化之前用它们来重新校准该评判器。已有的评判器重新校准方法成本高昂或依赖启发式,并且已知当代理在一组稀有输出上校准不当的时候,同策略采样会失败。在这项工作中,我们提出包络采样,这是一种有理论基础的评判器重新校准方法,其目标是在假设人类奖励和重新校准后的奖励位于评判器周围的一个 $L^2$ 球内的情况下,最小化后训练模型遗憾的上界。我们给出了实用算法,可以通过拒绝采样或针对修改后的奖励进行微调来从包络中采样;在临床笔记生成和一个受控的谄媚任务上的实验表明,在包络样本上重新校准能够缓解奖励黑客行为,而在基础模型样本上重新校准则不能。
cs.LG / 102 / 2610.11288
Memorization and Malign Generalization in Conditional Diffusion Models with Random Features
具有随机特征的条件扩散模型中的记忆与恶性泛化
Gwangho Kim, Sungyoon Lee
cs.LG
diffusion
扩散模型相关
Abstract
Conditional diffusion models generate diverse, novel, and high-quality samples under prescribed conditions. However, theoretical understanding of their memorization and generalization remains limited, while recent works have characterized these behaviors primarily in unconditional settings. In this work, we analyze a random-feature conditional score model in the high-dimensional proportional limit, deriving asymptotic expressions for training and test losses. By decomposing the test loss, we show that in the overparameterized regime, increasing model width improves prediction of the condition-dependent mean while reducing within-condition prediction variance, a phenomenon we term "malign generalization." Furthermore, analyzing the training loss reveals that more informative conditions lead to memorization of training samples at smaller widths. These theoretical findings are supported by experiments with U-Net architectures on realistic data.
Chinese Translation
条件扩散模型在给定条件下生成多样、新颖且高质量的样本。然而,对其记忆与泛化的理论理解仍然有限,而近期工作主要是在无条件设定下刻画这些行为。在这项工作中,我们分析了高维比例极限下的随机特征条件得分模型,推导出训练损失和测试损失的渐近表达式。通过分解测试损失,我们表明,在过参数化区域,增加模型宽度会改善对条件相关均值的预测,同时降低条件内预测方差,我们将这一现象称为“恶性泛化”。此外,分析训练损失表明,信息更丰富的条件会导致在更小的宽度下对训练样本的记忆。这些理论发现得到了在真实数据上使用 U-Net 架构的实验支持。
cs.LG / 103 / 2610.11315
Deflating the Hessian: Rank-4 W4A4 Quantization for Multimodal Diffusion Transformers
收缩 Hessian:面向多模态扩散 Transformer 的秩-4 W4A4 量化
Shiwen Wang, Pengxiang Zhao, Xiaoming Yuan
cs.LG · cs.CV · math.OC
diffusion
扩散模型相关
Abstract
In diffusion transformers, low-rank branches can mitigate 4-bit weight--activation (W4A4) post-training quantization (PTQ) loss by decomposing each weight into a low-bit residual and a high-precision low-rank component. Existing low-rank PTQ approaches, however, either optimize low-rank compensation and residual quantization separately, often requiring higher ranks, or rely on second-order weight updates without explicitly modeling activation quantization error, which becomes particularly pronounced under 4-bit quantization. To address these limitations, we present \method{}, a unified framework modeling low-rank-assisted W4A4 PTQ as a coupled calibration problem and deriving optimization-based solvers from the joint objective. Eliminating the output-side low-rank factor yields a \emph{deflated Hessian} that discounts residual errors already captured by the low-rank component, while an activation-noise surrogate is incorporated to suppress activation quantization error. Across five diffusion backbones, rank-4 \method{} consistently outperforms rank-4 SVDQuant in PSNR and LPIPS. It further surpasses rank-32 SVDQuant on SANA-1.6B, FLUX.1-schnell, and FLUX.1-dev with an $8\times$ smaller rank and up to $6.25\times$ faster quantization. Furthermore, on the Qwen3-8B LLM, rank-4 \method{} improves MMLU accuracy from 61.50\% to 68.17\% over rank-32 SVDQuant. Overall, \method{} achieves better W4A4 performance with substantially lower rank and quantization cost.
Chinese Translation
在扩散 Transformer 中,低秩分支可以通过将每个权重分解为一个低比特残差和一个高精度低秩分量,来缓解 4 比特权重--激活(W4A4)训练后量化(PTQ)损失。然而,现有的低秩 PTQ 方法要么分别优化低秩补偿和残差量化,这通常需要更高的秩;要么依赖二阶权重更新,而没有显式建模激活量化误差,而该误差在 4 比特量化下变得尤为显著。为了解决这些局限,我们提出 \method{},一个将低秩辅助 W4A4 PTQ 建模为耦合校准问题、并从联合目标中推导基于优化的求解器的统一框架。消除输出侧低秩因子会产生一个 \emph{收缩 Hessian},它会对已由低秩分量捕获的残差误差进行折减,同时引入一个激活噪声代理来抑制激活量化误差。在五个扩散骨干网络上,秩-4 \method{} 在 PSNR 和 LPIPS 上始终优于秩-4 SVDQuant。它进一步在 SANA-1.6B、FLUX.1-schnell 和 FLUX.1-dev 上超越秩-32 SVDQuant,且秩小 $8\times$,量化速度最高快 $6.25\times$。此外,在 Qwen3-8B LLM 上,秩-4 \method{} 相较于秩-32 SVDQuant,将 MMLU 准确率从 61.50\% 提升到 68.17\%。总体而言,\method{} 以显著更低的秩和量化成本实现了更好的 W4A4 性能。
cs.LG / 104 / 2610.11349
Sample-Efficient Generative Conformal Prediction
样本高效的生成式共形预测
Minxing Zheng, Shixiang Zhu
cs.LG · stat.ME
diffusion
扩散模型相关
Abstract
Generative conformal prediction builds uncertainty sets from samples of a conditional generator, which are efficient only when the samples represent the response distribution well. This can require many samples, each of which can be costly, as in large diffusion models and scientific simulators, so the sampling budget must be used efficiently. Existing methods draw the same number of samples at every input, wasting samples where the response distribution is simple and undersampling where it is complex, which inflates sets and leaves those inputs under-covered. We propose CASA (Conformal Adaptive Sample Allocation), which characterizes the marginal value of an additional sample and allocates samples across inputs to minimize the expected set size subject to marginal coverage and an average sampling budget. Theoretical analysis shows that adaptive allocation yields smaller sets than a fixed count at the same budget: a missed mode forces a radius that spans the gap between modes, and even oracle radius cannot compensate for it. On synthetic and real tasks, CASA produces substantially smaller sets at the same budget, often improves conditional coverage, and complements existing radius-adaptive methods.
Chinese Translation
生成式共形预测从条件生成器的样本构建不确定性集,只有当这些样本很好地表示响应分布时,这些集合才是高效的。这可能要求大量样本,而每个样本都可能代价高昂,如在大型扩散模型和科学模拟器中那样,因此必须高效地使用采样预算。现有方法在每个输入处抽取相同数量的样本,在响应分布简单处浪费样本,在响应分布复杂处采样不足,这会膨胀集合,并使那些输入覆盖不足。我们提出 CASA(共形自适应样本分配),它刻画额外样本的边际价值,并在输入之间分配样本,以在边缘覆盖率和平均采样预算约束下最小化期望集合大小。理论分析表明,在相同预算下,自适应分配比固定数量产生更小的集合:一个被遗漏的模态会迫使半径跨越模态之间的间隙,而即使使用 oracle 半径也无法弥补这一点。在合成任务和真实任务上,CASA 在相同预算下产生显著更小的集合,通常改善条件覆盖率,并且与现有的半径自适应方法互补。
cs.LG / 105 / 2610.11362
Bernoulli Flow Models: Self-Consistent Generative Modeling for Binary Data
伯努利流模型:面向二值数据的自洽生成建模
Hao Mo, Liying Yang, Shumin Yao, Xinxing Yu, Ajian Liu, Xudong Mao, Yanyan Liang
cs.LG · cs.CV
diffusion
扩散模型相关
Abstract
Binary diffusion models typically require a large number of function evaluations (NFEs) to generate high-quality samples, making practical inference computationally expensive. Reducing NFEs while preserving sample quality without distillation or additional training remains a significant challenge. Existing binary diffusion models define a discrete one-step forward path and then derive the reverse posterior. In low-NFE settings requiring cross-step sampling, they approximate the true multi-step likelihood with a single-step likelihood transition, which severely degrades sample quality. To address this fundamental limitation and decouple the generative dynamics from fixed discrete time steps, we propose Bernoulli Flow Models (BFM). Rather than relying on sequential one-step Markov diffusion chains, BFM defines a unified continuous global Bernoulli probability flow path between data distributions and pure noise, from which we derive analytical closed-form posterior transitions over arbitrary time intervals. Consequently, reducing the inference NFE is no longer an approximation based on skipping discrete steps; it only requires re-evaluating the analytical posterior over a new time grid. This eliminates the structural training-inference mismatch inherent to discrete chains and yields self-consistent low-NFE sampling. Experiments show that BFM is highly robust to aggressive NFE reduction. On LSUN Churches 256x256, a BFM trained with 256 steps achieves an FID of 9.22 using only 16 sampling steps, whereas the state-of-the-art discrete baseline degrades to 204.10. BFM also remains competitive with continuous and discrete generative baselines under standard full-step inference. These results establish BFM as a theoretically rigorous, self-consistent, and practically effective framework for fast binary data generation.
Chinese Translation
二值扩散模型通常需要大量函数评估(NFEs)才能生成高质量样本,这使得实际推理的计算代价高昂。在不使用蒸馏或额外训练的情况下,减少 NFE 同时保持样本质量仍然是一个重大挑战。现有二值扩散模型定义了一个离散的单步前向路径,然后推导反向后验。在需要跨步采样的低 NFE 设置中,它们用单步似然转移来近似真实的多步似然,这严重降低了样本质量。为了解决这一根本性局限,并将生成动力学与固定的离散时间步解耦,我们提出伯努利流模型(BFM)。BFM 不依赖于顺序的单步马尔可夫扩散链,而是在数据分布与纯噪声之间定义了一个统一的连续全局伯努利概率流路径,并由此推导出任意时间区间上的解析闭式后验转移。因此,减少推理 NFE 不再是一种基于跳过离散步骤的近似;它只需要在新的时间网格上重新计算解析后验。这消除了离散链固有的训练-推理结构性不匹配,并产生自洽的低 NFE 采样。实验表明,BFM 对激进的 NFE 缩减具有高度鲁棒性。在 LSUN Churches 256x256 上,一个用 256 步训练的 BFM 仅使用 16 个采样步就达到了 9.22 的 FID,而最先进的离散基线则退化到 204.10。在标准全步推理下,BFM 也保持与连续和离散生成基线具有竞争力。这些结果确立了 BFM 是一个理论严谨、自洽且实际有效的快速二值数据生成框架。
cs.LG / 106 / 2610.11366
SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces
SpatialOPSD:从经过验证的编码智能体轨迹中自蒸馏空间智能
Rongxue Li, Meng Yang, Yiru Mao, Yongliang Tao, Lulu Hu, Bin Yang, Zhao Xu, Weihua Luo, Bowen Xu
cs.LG
large language model
大语言模型相关
Abstract
Spatial coding agents significantly improve spatial reasoning in Multimodal Large Language Models (MLLMs) by using external tools to generate verified execution traces. However, this paradigm inherently suffers from prohibitive inference-time overhead and external dependencies. In this paper, we explore whether an MLLM can internalize this agentic capability to operate entirely tool-free. We begin with a simple observation: prompting an MLLM with summarized execution traces of a spatial coding agent naturally unlocks the model's internal spatial Chain-of-Thought (CoT). Motivated by this, we introduce SpatialOPSD, an on-policy self-distillation framework that internalizes spatial reasoning into a standalone MLLM by formulating verified agent traces as privileged information. To mitigate privileged-information leakage during distillation, we introduce Repetition-Aware Distillation, which combines repetition masking with unlikelihood regularization. Experiments across multiple benchmarks demonstrate that self-distilling SpatialOPSD achieves higher average accuracy than SFT and GRPO on both spatial and OOD datasets, exhibiting superior performance and generalization.
Chinese Translation
空间编码智能体通过使用外部工具生成经过验证的执行轨迹,显著提升了多模态大语言模型(MLLMs)中的空间推理能力。然而,这一范式本质上存在难以承受的推理时开销和外部依赖。本文探讨了一个 MLLM 是否能够内化这种智能体能力,以完全无需工具地运行。我们从一项简单的观察出发:用空间编码智能体的执行轨迹摘要来提示 MLLM,会自然地解锁模型内部的空间思维链(CoT)。受此启发,我们提出了 SpatialOPSD,一个同策略自蒸馏框架,它通过将经过验证的智能体轨迹形式化为特权信息,将空间推理内化到一个独立的 MLLM 中。为了缓解蒸馏过程中的特权信息泄漏,我们引入了重复感知蒸馏,它将重复掩码与非似然正则化相结合。在多个基准上的实验表明,自蒸馏 SpatialOPSD 在空间和 OOD 数据集上均取得了比 SFT 和 GRPO 更高的平均准确率,展现出更优的性能和泛化能力。
cs.LG / 107 / 2610.11373
From a Prompt to Repertoires: Evolving Functional REpertoires Enable LLM Continual Learning
从单个提示到函数库:演化功能函数库使LLM具备持续学习能力
Fengyuan Liu, Yue Wang, Hangxi Guo, Fengyuan Liu, Chenxu Wu, Yanguang Liu, Mengnan Du
cs.LG · cs.CL
large language model
大语言模型相关
Abstract
Continual learning remains challenging for large language models, which must enable models to acquire new skills and knowledge without degrading existing capabilities. Existing approaches typically address this challenge by carefully designing how model parameters are updated. In contrast, prompt optimization avoids costly parameter updates while achieving competitive or even superior performance to reinforcement learning methods such as GRPO on individual knowledge-intensive and reasoning tasks. This raises a natural question: \textit{Can prompt optimization, as an efficient adaptation approach, be directly applied to continual learning?} Our analysis shows that, under sequential task adaptation, it suffers from catastrophic forgetting, while optimized prompts accumulate rules that overfit to local task distributions. To address these limitations, we propose \emph{Evolving Functional REpertoires} (EFRE), which replaces a single prompt with a repertoire of functions that evolves as new tasks arrive: compatible updates refine existing functions, while conflicting updates trigger the emergence of new ones. On a three-task continual-learning stream, EFRE achieves a final average performance 7.50 percentage points higher than GRPO. Moreover, after adaptation to the Bio task, its performance on FinQA decreases by only 1.56 percentage points, compared with 25.10 percentage points for the base prompt optimization method. We further instantiate EFRE in a minimal agent system and observe consistent improvements across different backbone models. Overall, these results demonstrate EFRE's strong performance in continual learning for large language models and highlight its substantial potential for continual learning in advanced agent systems.
Chinese Translation
持续学习对大语言模型而言仍然具有挑战性,模型必须能够在获取新技能和知识的同时不损害已有能力。现有方法通常通过精心设计模型参数如何更新来应对这一挑战。相比之下,提示优化避免了代价高昂的参数更新,同时在单项知识密集型和推理任务上取得了与GRPO等强化学习方法相当甚至更优的性能。这引出一个自然的问题:\textit{作为一种高效的适配方法,提示优化能否直接应用于持续学习?}我们的分析表明,在顺序任务适配下,它会遭受灾难性遗忘,而优化后的提示会累积起过拟合于局部任务分布的规则。为解决这些局限,我们提出\emph{演化功能函数库}(EFRE),它用随新任务到来而演化的函数库取代单一提示:兼容的更新会细化已有函数,而冲突的更新则触发新函数的涌现。在一个三任务持续学习流上,EFRE的最终平均性能比GRPO高7.50个百分点。此外,在适配Bio任务后,其在FinQA上的性能仅下降1.56个百分点,而基础提示优化方法则下降25.10个百分点。我们进一步在一个最小化智能体系统中实例化EFRE,并观察到在不同骨干模型上的一致改进。总体而言,这些结果证明了EFRE在大语言模型持续学习中的强劲性能,并凸显了其在先进智能体系统持续学习中的巨大潜力。
cs.LG / 108 / 2610.11475
Rare Gate Disagreements Can Limit Plasticity: When Gradient Flow Mispredicts Finite-Batch SGD
罕见的门分歧可能限制可塑性:当梯度流错误预测有限批量 SGD 时
Ruoyu Zhao, Mingxuan Zhang, Jianbo Dai, Jiaqi Wu, Chenyu Zhu, Tong Che
cs.LG · math.OC · stat.ML
diffusion
扩散模型相关
Abstract
Population gradient flow is a common tool for reasoning about how neural networks adapt, including after pretraining. We show that it can mispredict finite-batch stochastic gradient descent (SGD) qualitatively, and we trace the discrepancy to a specific mechanism. In a two-unit ReLU regression, a source task drives the two neurons toward positive proportionality and a target task rewards separating them. After source training for time $T$, gradient flow recovers on the target in time linear in $T$. Online SGD with batch size $b$ and step size $η$ in both phases instead fails with high probability throughout a horizon of order $e^{c/η}$ once $T \gtrsim \log(b/η)$, uniformly on an explicit set of initializations with Gaussian probability above one percent. For each fixed $T$, small-step SGD still recovers, so the failure requires the joint limit of small steps and long pretraining. At the target clone, the population instability is carried entirely by inputs on which the two ReLU gates disagree. For units at angle $δ$ these inputs form a wedge of probability $δ/π$, and weight decay shrinks the angle exponentially during pretraining. On every other input both units receive the same random linear update, which contracts their separation in conditional expectation. Bounding the cumulative probability of sampling the wedge along the exact online recursion, without a diffusion approximation, shows that recovery with fixed probability from an identical source-gradient-flow checkpoint, within $e^{c/η}$ updates, requires $Nb \gtrsim e^{λT}$ target samples and batch size $b \gtrsim ηe^{λT}$, where $N$ counts updates and $λ$ is the weight decay. In simulations, recovery is approximately a function of the disagreement budget $bδ/η$ and saturates in the horizon.
Chinese Translation
总体梯度流是用于推理神经网络如何适应(包括预训练之后)的一种常用工具。我们表明,它可能在定性上错误预测有限批量随机梯度下降(SGD),并且我们将这一差异追溯到一个特定机制。在一个两单元 ReLU 回归中,源任务驱动两个神经元趋向正比例关系,而目标任务奖励将它们分离。在源任务上训练时间 $T$ 后,梯度流在目标任务上以与 $T$ 成线性的时间恢复。在两个阶段中均使用批大小 $b$ 和步长 $η$ 的在线 SGD 反而会以高概率失败:一旦 $T \gtrsim \log(b/η)$,它会在量级为 $e^{c/η}$ 的一个视界内一直失败,并且在一个高斯概率超过百分之一的显式初始化集合上一致成立。对于每个固定的 $T$,小步长 SGD 仍能恢复,因此这种失败需要小步长和长预训练的共同极限。在目标克隆处,总体不稳定性完全由两个 ReLU 门不一致的输入承载。对于夹角为 $δ$ 的单元,这些输入构成一个概率为 $δ/π$ 的楔形区域,而权重衰减在预训练期间使该夹角指数收缩。在每一个其他输入上,两个单元接收到相同的随机线性更新,这会在条件期望中收缩它们的分离度。在不使用扩散近似的情况下,对沿精确在线递归采样该楔形区域的累积概率进行界定表明:要从一个相同的源梯度流检查点出发,在 $e^{c/η}$ 次更新内以固定概率恢复,需要 $Nb \gtrsim e^{λT}$ 个目标样本和批大小 $b \gtrsim ηe^{λT}$,其中 $N$ 计更新次数,$λ$ 是权重衰减。在模拟中,恢复大致是分歧预算 $bδ/η$ 的函数,并在该视界内饱和。
cs.LG / 109 / 2610.11670
Early Signatures of Memorization in Diffusion Models via Basin Geometry and Cyclic Denoising
通过盆地几何与循环去噪识别扩散模型中记忆化的早期特征
Nikhil Verma, Siddharthan Dileep, Anoop Singh, Srikanth Sastry, Ramya Hebbalaguppe, Sayan Ranu, N. M. Anoop Krishnan
cs.LG
diffusion
扩散模型相关
Abstract
Diffusion models generalize early in training and later reproduce individual training samples. Standard tests detect memorization only once one-shot generation produces near-copies, leaving a released model unaudited until its outputs fail. We show that memorization is encoded in the geometry of the learned energy landscape before it appears in generated samples, a state we call latent memorization. Using score divergence and basin volume, we find that localized basins form around training samples and separate them from held-out samples before the first memorized sample appears, with an onset that follows the same $O(n)$ scaling as the memorization time. We probe these basins with cyclic denoising, which repeatedly applies partial noising and denoising. Under the exact empirical score, we prove that cycling started near an isolated training sample recovers it and returns to it over any finite number of cycles with high probability. In trained models, cycling recovers training images from CelebA and CIFAR-10 checkpoints whose one-shot samples contain no copies, and at a CelebA checkpoint with 0.1% one-shot copies, 500 cycles raise the memorized fraction above 30%. Cycling also reveals degenerate attractors that match no single training image and fade as training proceeds, so residence in a basin does not by itself imply memorization. These findings hold on a Gaussian mixture, CelebA, and CIFAR-10 across optimizers, architectures, noise schedules, and training-set sizes, and extend to off-the-shelf Stable Diffusion v1.4, where the cycled conditional-unconditional divergence gap separates memorized from non-memorized prompts with an AUC of 0.944 and a TPR of 0.866 at 1% FPR. More broadly, what a diffusion model has memorized is a property of the geometry and stability of its learned distribution, and assessing it requires examining this structure rather than generated outputs alone.
Chinese Translation
扩散模型在训练早期进行泛化,随后会重现单个训练样本。标准测试仅在单次生成产生近乎复制品时才能检测到记忆化,这使得已发布的模型在其输出失败之前一直未被审计。我们表明,记忆化在出现在生成样本中之前,就已经被编码在学习到的能量景观的几何结构之中,我们将这种状态称为潜在记忆化。利用分数散度和盆地体积,我们发现,在第一个被记忆的样本出现之前,训练样本周围会形成局部化的盆地,并将其与留出样本区分开来,其起始时间遵循与记忆化时间相同的 $O(n)$ 缩放规律。我们通过循环去噪来探测这些盆地,该方法反复施加部分加噪与去噪。在精确的经验分数下,我们证明:从某个孤立训练样本附近开始的循环过程会恢复该样本,并以高概率在任意有限次循环后返回到该样本。在训练好的模型中,循环过程能够从那些单次采样中不含任何复制品的 CelebA 和 CIFAR-10 检查点中恢复出训练图像;而在一个单次采样复制率为 0.1% 的 CelebA 检查点上,500 次循环可将记忆化比例提升至 30% 以上。循环过程还揭示了退化吸引子,它们不与任何单个训练图像相匹配,并随着训练进行而逐渐消退,因此停留在某个盆地中本身并不意味着记忆化。这些发现在高斯混合模型、CelebA 和 CIFAR-10 上,在不同优化器、架构、噪声调度和训练集规模下均成立,并进一步扩展到现成的 Stable Diffusion v1.4:在其中,循环后的条件-无条件散度差距能够将记忆化提示与非记忆化提示区分开来,AUC 为 0.944,在 1% FPR 下 TPR 为 0.866。更广泛地说,扩散模型记住了什么,是其学习分布的几何结构与稳定性的一种属性,要评估这一点,需要考察这种结构,而不能仅凭生成的输出。
cs.LG / 110 / 2610.11761
RAGenome: Scaling Retrieval-Based Genomic Language Models to Long Contexts
RAGenome:将基于检索的基因组语言模型扩展至长上下文
Frederikke Isa Marin, Panagiotis Antoniadis, Dionysia Danai Brilli, Andreas Bjerregaard, Rachael DeVries, Yan Li, Ole Winther, Wouter Boomsma
cs.LG
large language model
大语言模型相关
Abstract
The genome holds the blueprint that governs the biological properties of the cell. Consequently, advancing our knowledge of genomic function is crucial both for a broader understanding of biology and for continued biomedical advances. The success of large language models on natural language and protein sequences has motivated similar efforts on genomic data. However, standard genomic language models (gLMs) often require extremely large computational resources and still fall behind traditional methods on some downstream tasks. Recently, MSA-based pretraining has been proposed as an efficient alternative, but existing models are limited to short input contexts, restricting their use to short-range tasks, such as variant effect prediction. In this work, we present RAGenome, the first retrieval-based gLM that scales pretraining to longer contexts (100$\times$ longer than existing MSA-based gLMs), allowing it to capture both across-species evolutionary relationships and within-species longer-range interactions. Trained on whole-genome alignments from 100 vertebrates, RAGenome substantially improves the long-range capabilities of MSA-based gLMs, raising gene finding performance from 0.45 to 0.60, while remaining competitive on purely evolutionary-based tasks like prioritizing pathogenic variants. RAGenome provides competitive gLM performance at a fraction of the training cost, unifying evolutionary modeling and long-range capabilities within a single, flexible, scalable framework. Code is available at https://github.com/PanosAntoniadis/RAGenome.
Chinese Translation
基因组承载着决定细胞生物学特性的蓝图。因此,推进我们对基因组功能的认识,对于更广泛地理解生物学以及持续推动生物医学进展都至关重要。大语言模型在自然语言和蛋白质序列上的成功,激发了在基因组数据上进行类似工作的尝试。然而,标准的基因组语言模型(gLMs)通常需要极其庞大的计算资源,并且在一些下游任务上仍然落后于传统方法。最近,基于MSA的预训练被提出作为一种高效的替代方案,但现有模型受限于较短的输入上下文,使其应用局限于短程任务,例如变异效应预测。在这项工作中,我们提出了 RAGenome,这是首个基于检索的 gLM,可将预训练扩展至更长的上下文(比现有基于 MSA 的 gLM 长 100$\times$),使其能够同时捕捉跨物种的进化关系和物种内的长程相互作用。在来自100种脊椎动物的全基因组比对数据上训练后,RAGenome 显著提升了基于 MSA 的 gLM 的长程能力,将基因查找性能从 0.45 提高到 0.60,同时在诸如致病性变异优先级排序等纯进化类任务上仍保持竞争力。RAGenome 以极低的训练成本提供了具有竞争力的 gLM 性能,将进化建模与长程能力统一在一个灵活、可扩展的单一框架之中。代码可在 https://github.com/PanosAntoniadis/RAGenome 获取。
cs.LG / 111 / 2610.11834
Recovery Guarantees for Posterior Sampling of One-Bit Compressed Sensing
单比特压缩感知后验采样的恢复保证
Jing Ma, Yujia Wu, Zhaoqiang Liu
cs.LG · cs.IT
diffusion
扩散模型相关
Abstract
We study the sample complexity of noisy one-bit compressed sensing for signals drawn from a prior distribution. By characterizing the effective distributional complexity of the prior via its approximate covering number, we prove that posterior sampling achieves accurate recovery with high probability when the number of measurements scales with the logarithm of the approximate covering number, up to a one-bit separation gap factor. This upper bound is robust to learned prior mismatch. Specifically, we show that posterior sampling with an approximate prior remains reliable, provided that the learned prior distribution is sufficiently close to the true signal distribution in Wasserstein distance. In addition, we establish a sample complexity lower bound for any reliable method of noisy one-bit compressed sensing, showing that our upper bound is nearly matched in its main prior dependent term. To approximate the ideal posterior sampling process for real world scenarios, we instantiate posterior sampling through a plug-and-play algorithm with diffusion priors. Experiments on the FFHQ and ImageNet datasets demonstrate the effectiveness of our proposed approach.
Chinese Translation
我们研究从先验分布中抽取的信号的含噪单比特压缩感知的样本复杂度。通过用先验的近似覆盖数刻画其有效分布复杂度,我们证明:当测量次数与近似覆盖数的对数成比例(最多相差一个单比特分离间隙因子)时,后验采样以高概率实现准确恢复。该上界对学习到的先验失配具有鲁棒性。具体来说,我们表明,只要学习到的先验分布在 Wasserstein 距离上足够接近真实信号分布,使用近似先验的后验采样就仍然可靠。此外,我们为含噪单比特压缩感知的任何可靠方法建立了一个样本复杂度下界,表明我们的上界在其主要依赖于先验的项上几乎与之匹配。为了在现实世界场景中近似理想的后验采样过程,我们通过带有扩散先验的即插即用算法来实例化后验采样。在 FFHQ 和 ImageNet 数据集上的实验证明了我们提出的方法的有效性。
cs.LG / 112 / 2610.11854
GRPODropout: Less is More for Online Reinforcement Learning Rollouts
GRPODropout:在线强化学习采样轨迹中的少即是多
Hexuan Deng, Zihao Yan, Xuebo Liu, Shuo Nie, Yue Wang, Chen Wang, Zhaohua Zhang, Tianwen Jiang, Qiuyong Xiao, Jihong Zhang, Min Zhang
cs.LG · cs.AI · cs.CL
large language model
大语言模型相关
Abstract
Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement. Existing methods address this issue either through algorithm-level interventions, such as reward modification and entropy/KL regularization, or through token-level reweighting. We investigate a complementary perspective: entropy collapse can also be mitigated by changing which generated rollouts contribute to policy updates. Under the same sampling budget, not all rollouts contribute positively to an update, and selectively excluding some can improve learning. To address this, we propose GRPODropout: before the standard update, we use a simple strategy that selectively removes a small number of high-probability positive-advantage rollouts and recenters the retained advantages. To motivate this design, we develop a rollout-level theoretical analysis that guides method design and threshold selection. The method changes only rollout usage, and adds negligible computational overhead. Experiments show higher accuracy than original GRPO and higher actor entropy while using fewer rollout samples for updates, illustrating "less is more." This work provides insight into RL rollout usage: removing some rollouts can improve performance. Code is available at https://github.com/hexuandeng/GRPODropout/.
Chinese Translation
诸如 GRPO 之类的强化学习(RL)方法显著提升了大语言模型推理能力,但常常遭受策略熵坍塌:采样多样性的丧失削弱了探索,并限制了进一步改进。现有方法要么通过算法层面的干预来解决这一问题,例如奖励修改和熵/KL 正则化,要么通过 token 级别的重加权来解决。我们研究了一个互补的视角:熵坍塌也可以通过改变哪些生成的采样轨迹对策略更新有贡献来缓解。在相同的采样预算下,并非所有采样轨迹都会对更新产生正向贡献,而有选择地排除其中一些可以改进学习。为了解决这一问题,我们提出了 GRPODropout:在标准更新之前,我们使用一种简单策略,有选择地移除少量高概率、正优势的采样轨迹,并对保留下来的优势进行重新中心化。为了给这一设计提供动机,我们提出了一项采样轨迹层面的理论分析,用以指导方法设计和阈值选择。该方法仅改变采样轨迹的使用方式,并增加可忽略的计算开销。实验表明,与原始 GRPO 相比,该方法具有更高的准确率,并且在使用更少的采样轨迹样本进行更新时具有更高的 actor 熵,体现了“少即是多”。这项工作为 RL 采样轨迹的使用提供了洞见:移除一些采样轨迹可以提升性能。代码可在 https://github.com/hexuandeng/GRPODropout/ 获取。
cs.LG / 113 / 2610.11896
Automated Assembly Instruction Generation from CAD Models Using Grounded Large Language Models: A Human-in-the-Loop Framework
使用基于事实的大语言模型从CAD模型自动生成装配指令:一种人在回路框架
Aaron Dsouza, Mohammed Azeez Khan, Ashutosh Mishra, Arshaan Khan, Neha K. Nair, Amar Kumar Behera
cs.LG
large language model
大语言模型相关
Abstract
Assembly documentation is a downstream manufacturing artifact that is still usually authored by interpreting CAD models by hand. Structured product data and large language models are both available, yet studies of CAD interpretation, assembly sequence planning, instruction writing, and human oversight have largely proceeded separately. This paper formulates CAD-grounded assembly instruction generation: the production of natural-language assembly procedures constrained by structured engineering information extracted from CAD models. The proposed framework maps a STEP assembly to a typed ProductGraph intermediate representation, derives a precedence order by deterministic topological sorting, realizes each step as language conditioned only on selected graph context, attaches per-step visual documentation, and applies rule-based and model-assisted checks. PDF export remains disabled until a human reviewer resolves every quality flag. The case study establishes endto-end feasibility on a built-in six-part reference assembly: the pipeline preserves a reported assembly order and carries quantity, material, and torque into an exported manual page. Generalization and geometric validation remain open empirical questions. The contribution is an architecture that separates engineering state, deterministic reasoning, grounded language realization, verification, and human release.
Chinese Translation
装配文档是一种下游制造产物,目前通常仍通过人工解读CAD模型来编写。结构化产品数据和大语言模型都已具备,然而关于CAD解读、装配序列规划、指令编写和人工监督的研究在很大程度上仍是各自独立进行的。本文提出了基于CAD的装配指令生成这一问题形式:即在从CAD模型中提取的结构化工程信息的约束下,生成自然语言装配流程。所提出的框架将STEP装配体映射为带类型的ProductGraph中间表示,通过确定性拓扑排序推导出先后顺序,将每个步骤实现为仅以选定的图上下文为条件的语言,附加每个步骤的可视化文档,并应用基于规则和模型辅助的检查。在人工审核者解决所有质量标记之前,PDF导出保持禁用。该案例研究在一个内置的六零件参考装配体上确立了端到端的可行性:该流水线保持了所报告的装配顺序,并将数量、材料和扭矩传递到导出的手册页面中。泛化能力和几何验证仍是尚未解决的实证问题。其贡献在于一种架构,该架构将工程状态、确定性推理、基于事实的语言实现、验证和人工放行分离开来。
cs.LG / 114 / 2610.12183
A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization
深入审视智能体式 BBO:为黑盒优化基准测试 LLM 智能体
Ming Chen, Rong-Xi Tan, Ke Xue, Yu-Jie Zhou, Taiye Lu, Zhi-Xuan Gao, Peng Xie, Zijun Shen, Chen Lu, Haopu Shang, Chao Qian
cs.LG · cs.AI · cs.NE
large language model
大语言模型相关
Abstract
Black-box optimization (BBO) arises in many scientific and engineering problems where objective evaluations are expensive and limited. Recent large language model (LLM) agents offer a new way to approach BBO by combining task semantics, computation, optimization tools, and feedback-driven decision making, showing great potential due to the integration with mathematically rigorous tools. However, existing agentic BBO studies use different task domains and system configurations, making their results difficult to compare and the effects of individual design choices hard to isolate. We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol. In our experiments, agentic BBO achieves higher family-averaged scores than direct LLM-based methods in all five domains and outperforms the best numerical optimizers in four. We further study three factors shaping agent performance: optimization tools, task information and prior knowledge, and the role of the LLM during search. Our results show that additional numerical tools do not consistently improve performance, task semantics are broadly useful while more specific priors are less reliable, and numerical optimizers can effectively absorb gains from search trajectories established by the agent. Finally, we introduce a five-task frontier challenge within AgenticBBO-Bench and evaluate seven LLMs under the Codex agent harness, where GPT-6 Astra and DeepSeek-V4.1-Flash lie on the Pareto frontier of performance and cost among the evaluated models. Our code is available at https://github.com/lamda-bbo/agentic-bbo.
Chinese Translation
黑盒优化(BBO)出现在许多科学和工程问题中,在这些问题中目标评估昂贵且受限。近期的大语言模型(LLM)智能体通过结合任务语义、计算、优化工具和反馈驱动的决策,为处理 BBO 提供了一种新途径,并且由于与数学上严谨的工具相结合,展现出巨大潜力。然而,现有的智能体式 BBO 研究使用不同的任务领域和系统配置,使得其结果难以比较,并且难以分离各个设计选择的影响。因此,我们提出了 AgenticBBO-Bench,一个面向智能体式 BBO 的跨领域基准,涵盖合成函数、超参数优化、数据库调优、芯片设计和分子设计,并在统一的有限预算评估协议下进行。在我们的实验中,智能体式 BBO 在所有五个领域中都取得了比直接基于 LLM 的方法更高的族平均分数,并在四个领域中优于最佳数值优化器。我们进一步研究了塑造智能体性能的三个因素:优化工具、任务信息和先验知识,以及 LLM 在搜索过程中的作用。我们的结果表明,额外的数值工具并不总能提升性能;任务语义具有广泛用处,而更具体的先验则不那么可靠;数值优化器可以有效吸收由智能体建立的搜索轨迹所带来的收益。最后,我们在 AgenticBBO-Bench 中引入了一个五项任务的前沿挑战,并在 Codex 智能体运行框架下评估了七个 LLM,其中 GPT-6 Astra 和 DeepSeek-V4.1-Flash 位于被评估模型性能与成本的帕累托前沿上。我们的代码可在 https://github.com/lamda-bbo/agentic-bbo 获取。
cs.LG / 115 / 2610.12189
Just Weather Scoring: Efficient End-to-end Nowcasting with Distributional Diffusion
Just Weather Scoring:使用分布扩散的高效端到端临近预报
Jannik Wiese, Johannes Schusterbauer, Tommaso Martorella, Björn Ommer
cs.LG · cs.CV
diffusion
扩散模型相关
Abstract
Generative diffusion models are well-suited for probabilistic precipitation nowcasting, but existing approaches often rely on separately trained compression or deterministic forecasting components and remain costly at inference due to iterative denoising. We introduce Just Weather Scoring (JWS), a single-stage, end-to-end diffusion model which addresses both issues by forecasting directly in radar space and enabling few-step generation. Radar-space modeling greatly simplifies training and inference and eliminates uncertainty arising from lossy compression. JWS combines Masked Asynchronous Diffusion, a timestep-sampling scheme that preserves clean context while adapting diffusion training to high-dimensional spatio-temporal data, with a simple scoring-rule objective that aligns training with probabilistic forecasting and unlocks few-step generation. On the SEVIR and MeteoNet benchmarks, JWS achieves state-of-the-art probabilistic forecasting performance at reduced training and inference cost. Even our smallest model remains competitive using substantially fewer parameters and more than 17x faster inference.
Chinese Translation
生成式扩散模型非常适合概率性降水临近预报,但现有方法通常依赖于单独训练的压缩或确定性预报组件,并且由于迭代去噪而在推理时仍然代价高昂。我们提出 Just Weather Scoring (JWS),一种单阶段、端到端的扩散模型,它通过在雷达空间中直接进行预报并实现少步生成,解决了这两个问题。雷达空间建模极大地简化了训练和推理,并消除了由有损压缩引起的不确定性。JWS 将 Masked Asynchronous Diffusion(一种时间步采样方案,它在使扩散训练适应高维时空数据的同时保留干净上下文)与一个简单的评分规则目标相结合,该目标使训练与概率预报对齐并解锁少步生成。在 SEVIR 和 MeteoNet 基准上,JWS 以降低的训练和推理成本实现了最先进的概率预报性能。即使我们最小的模型也仍然具有竞争力,它使用的参数显著更少,并且推理速度快 17 倍以上。
cs.LG / 116 / 2610.12327
SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference
SparseDecoding:面向准确且高效 LLM 推理的解码感知剪枝
Qitong Wang, Xinwei Niu, Mingluo Su, Shanwei Zhao, Shiai Zhu, Huan Wang
cs.LG · cs.CL
large language model
大语言模型相关
Abstract
The memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency. Layer-wise training-free network pruning approaches guided by the Hessian have been a prominent solution to this problem, as pruning reduces the number of nonzero parameters read from memory during decoding. Nevertheless, typical methods in this line compute the Hessian using pre-collected natural sequences, whereas the model is fed self-generated tokens during decoding, creating a distribution shift between the two sequences. The Hessian calculated on the natural sequence is different from that calculated on the generated sequence. We observe that this discrepancy causes the activation distribution during generation to deviate from that used for pruning, further hurting the pruned model performance. Moreover, most existing LLM pruning methods that bring actual speedup primarily target the sparse matrix-matrix (SpMM) multiplication, providing limited support for the sparse matrix-vector (SpMV) operations, which dominate decoding. To solve these problems, we introduce SparseDecoding, a principled decoding-aware pruning framework tailored for accurate and efficient LLM decoding. Specifically, at the algorithmic axis, SparseDecoding constructs calibration matrices from layer-wise activations collected during the dense-model autoregressive generation, excluding prefill, thereby aligning the pruning objective with the decoding activations. At the system axis, we develop an optimized N:M sparse matrix-vector kernel with bitmask indexing and fixed-step traversal. Substantial empirical results on representative LLMs (Llama-3.1-8B, Llama-3.3-70B, Qwen3-14B / 32B) demonstrate that our method consistently outperforms standard fixed-text calibration on the long-form generation benchmarks while achieving up to 1.48x end-to-end wall-clock decoding speedup on A100 GPUs.
Chinese Translation
大语言模型(LLM)推理的解码阶段的访存受限特性会带来显著的延迟。由 Hessian 引导的逐层免训练网络剪枝方法一直是该问题的重要解决方案,因为剪枝减少了解码期间从内存中读取的非零参数数量。然而,这一类方法中的典型方法使用预先收集的自然序列来计算 Hessian,而在解码期间模型被输入自生成 token,从而在这两种序列之间造成分布偏移。在自然序列上计算得到的 Hessian 不同于在生成序列上计算得到的 Hessian。我们观察到,这种差异导致生成过程中的激活分布偏离用于剪枝的激活分布,进一步损害剪枝模型的性能。此外,大多数能够带来实际加速的现有 LLM 剪枝方法主要针对稀疏矩阵-矩阵(SpMM)乘法,对主导解码的稀疏矩阵-向量(SpMV)运算提供有限支持。为解决这些问题,我们提出 SparseDecoding,一个面向准确且高效 LLM 解码的、有原则的解码感知剪枝框架。具体而言,在算法层面,SparseDecoding 从稠密模型自回归生成期间收集的逐层激活(排除预填充)中构建校准矩阵,从而使剪枝目标与解码激活对齐。在系统层面,我们开发了一个优化的 N:M 稀疏矩阵-向量内核,具有位掩码索引和固定步长遍历。在代表性 LLM(Llama-3.1-8B、Llama-3.3-70B、Qwen3-14B / 32B)上的大量实证结果表明,我们的方法在长文本生成基准上持续优于标准的固定文本校准,同时在 A100 GPU 上实现了高达 1.48 倍的端到端挂钟解码加速。
cs.LG / 117 / 2610.12340
Ambient Discrete Diffusion: Using the Wrong Data at the Right Time for Data Efficient Learning
环境离散扩散:在正确的时间使用错误的数据以实现数据高效学习
Julian Kleutgens, Mauricio Tec, Claudio Battiloro, Francesca Dominici, Giannis Daras
cs.LG
diffusion
扩散模型相关
Abstract
We introduce RefineMix, a framework for training discrete diffusion models under severe data scarcity, a common constraint in scientific applications. RefineMix uses out-of-distribution data at selected diffusion times to improve generalization without biasing the sampling distribution. Although this strategy has been explored in continuous diffusion, discrete diffusion presents a distinct challenge: unlike Gaussian noise, masking preserves domain information in surviving tokens, limiting the use of related data at high noise levels. At low noise levels, however, the domains effectively disjoint supports become an advantage, allowing the model to learn from both in-domain and out-of-distribution data without biasing the sampler. We formalize these intuitions and provide a theoretical analysis for the proposed method. Experimentally, across five domain-shift settings, RefineMix matches or outperforms in-domain finetuning and data mixing. For protein sequence generation, finetuning with just 197 in-domain examples nearly doubles the fraction of generated proteins that are simultaneously novel, foldable, and in-family compared to standard finetuning.
Chinese Translation
我们介绍了 RefineMix,一个在严重数据稀缺(科学应用中常见约束)下训练离散扩散模型的框架。RefineMix 在选定的扩散时间使用分布外数据,以提高泛化能力而不使采样分布产生偏差。尽管该策略已在连续扩散中被探索,离散扩散却呈现出一种独特的挑战:与高斯噪声不同,掩码会在存活 token 中保留域信息,限制了在高噪声水平下对相关数据的使用。然而,在低噪声水平下,各域实际上不相交的支撑集成为一项优势,使模型能够同时从域内和分布外数据中学习,而不使采样器产生偏差。我们将这些直觉形式化,并为所提方法提供了理论分析。在实验上,在五种域偏移设置中,RefineMix 匹配或优于域内微调和数据混合。对于蛋白质序列生成,仅用 197 个域内样本进行微调,与标准微调相比,生成的同时具有新颖性、可折叠性和同家族性的蛋白质比例几乎翻倍。
cs.LG / 118 / 2610.12390
Long Text to Predictive Features: LLM-Guided Blockwise Feature Engineering via Executable Program Search
从长文本到预测性特征:通过可执行程序搜索实现的LLM引导分块特征工程
Ziming Dai, Dabiao Ma, Ziheng Guo, Jack Dong, Zimu Zhou
cs.LG · cs.CL
large language model
大语言模型相关
Abstract
Industrial risk-control systems typically rely on structured-data models for efficient prediction, yet substantial valuable information remains embedded in unstructured long text. Extracting this information through manual feature engineering is labor-intensive, while requiring a large language model (LLM) to process every real-time input may not meet practical deployment requirements. To address this challenge, we propose LLM-BlockFE, an LLM-guided offline feature construction framework that converts long text into executable feature programs, thereby avoiding LLM calls during online inference. LLM-BlockFE constructs feature programs by incrementally appending immutable code blocks and evaluates candidate features using a downstream model. To address the tendency of conventional greedy search to become trapped in suboptimal solutions, our method introduces a block-level rollback mechanism based on depth-calibrated credit allocation and advances multiple independent search trajectories in an interleaved manner, reducing redundant exploration by sharing fixed descriptions of each trajectory's exploration direction. After the search, the resulting programs are frozen and deployed to extract structured features for downstream prediction models. Across two public and two private datasets, LLM-BlockFE achieves absolute AUC improvements of 0.0069 to 0.0358 over the strongest baseline on each dataset in the full-dataset comparison. Post-launch monitoring across five deployed financial risk-control applications shows absolute KS improvements of 0.02 to 1.56 percentage points over the existing manually designed strategy.
Chinese Translation
工业风控系统通常依赖结构化数据模型来实现高效预测,然而大量有价值的信息仍然嵌入在非结构化长文本中。通过人工特征工程提取这些信息耗费大量人力,而要求大语言模型(LLM)处理每一个实时输入可能无法满足实际部署要求。为应对这一挑战,我们提出了LLM-BlockFE,一个由LLM引导的离线特征构建框架,它将长文本转换为可执行的特征程序,从而避免在线推理期间调用LLM。LLM-BlockFE通过增量式追加不可变代码块来构建特征程序,并使用下游模型评估候选特征。针对传统贪心搜索容易陷入次优解的倾向,我们的方法引入了一种基于深度校准信用分配的块级回滚机制,并以交错方式推进多条独立的搜索轨迹,通过共享每条轨迹探索方向的固定描述来减少冗余探索。搜索完成后,所得到的程序被冻结并部署,用于为下游预测模型提取结构化特征。在两个公开数据集和两个私有数据集上,在全数据集比较中,LLM-BlockFE在每个数据集上相较最强基线取得了0.0069至0.0358的绝对AUC提升。对五个已部署的金融风控应用进行的上线后监控显示,相较现有的人工设计策略,KS绝对提升了0.02至1.56个百分点。
cs.AI / 119 / 2610.11918
From Surface to Depth: Towards Cognitive Appraisal Reasoning in Multimodal Emotion Understanding
从表面到深度:迈向多模态情感理解中的认知评价推理
Jia Li, Yichao He, Yangchen Yu, Qiankun Li, Xinyi Li, Baiyi Ye, Zhenzhen Hu, Richang Hong, Erik Cambria
cs.MM · cs.AI · cs.HC
large language model
大语言模型相关
Abstract
Recent multimodal large language models (MLLMs) increasingly incorporate explainable reasoning for emotion understanding. However, reasoning based mainly on observable affective cues can reduce emotion understanding to superficial cue-label associations, giving rise to the Clever Hans effect. Such shortcuts become unreliable when affective cues are implicit, conflicting across modalities, linguistically misleading, or obscured by redundant details. In contrast, human emotions are shaped by how individuals interpret and evaluate surrounding events beyond observable cues. Inspired by appraisal theories of emotion, we formulate multimodal emotion understanding as a progression from perception to cognitive appraisal, and introduce a dataset, a model, and a benchmark to support this novel paradigm. CogEmo-40K is a large-scale instruction-tuning dataset constructed through a perception-to-appraisal pipeline to elicit evidence-grounded reasoning across six cognitive appraisal dimensions underlying emotion. CogEmo-MoE is a compact sparse MLLM that introduces interleaved MoE blocks for appraisal-specific adaptation, enabling effective appraisal reasoning at a substantially smaller scale than typical emotion MLLMs. CogEmo-Bench introduces an Appraisal Evidence Quality Score (AEQS) to assess cognitive-affective understanding across six complementary appraisal dimensions, addressing the limitation of conventional emotion metrics that evaluate what emotion is predicted but not why it arises. Extensive experiments show that our paradigm not only leads CogEmo-Bench, but also exhibits strong cross-domain generalization. Our findings suggest that perception-to-appraisal reasoning can move beyond surface-level cue-label associations toward more reliable multimodal emotion understanding and closer cognitive alignment between MLLMs and humans.
Chinese Translation
近年来,多模态大语言模型(MLLMs)越来越多地将可解释推理纳入情感理解之中。然而,主要基于可观察情感线索的推理可能将情感理解简化为表层的线索—标签关联,从而引发聪明汉斯效应(Clever Hans effect)。当情感线索是隐含的、跨模态冲突的、语言上具有误导性的,或被冗余细节所掩盖时,此类捷径会变得不可靠。相比之下,人类情感是由个体如何解释和评价可观察线索之外的周围事件所塑造的。受情感评价理论启发,我们将多模态情感理解形式化为从感知到认知评价的递进过程,并引入一个数据集、一个模型和一个基准来支持这一新范式。CogEmo-40K 是一个大规模指令微调数据集,通过从感知到评价的流水线构建,以激发跨越情感背后的六个认知评价维度的基于证据的推理。CogEmo-MoE 是一个紧凑的稀疏 MLLM,引入交错的 MoE 块以进行面向评价的适配,使其能够在比典型情感 MLLM 小得多的规模上实现有效的评价推理。CogEmo-Bench 引入了评价证据质量分数(Appraisal Evidence Quality Score, AEQS),以评估跨六个互补评价维度的认知—情感理解,解决了传统情感指标只评估预测了什么情感而不评估其为何产生的局限。大量实验表明,我们的范式不仅在 CogEmo-Bench 上领先,而且展现出强大的跨域泛化能力。我们的发现表明,从感知到评价的推理能够超越表层的线索—标签关联,迈向更可靠的多模态情感理解以及 MLLM 与人类之间更紧密的认知对齐。
cs.LG / 120 / 2610.10810
Diagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill Seams
诊断并恢复长时程技能接缝处的观测空间偏移
Pranav Wagh, Yu Fang, Yue Yang, Mingyu Ding
cs.RO · cs.LG
diffusion
扩散模型相关
Abstract
Long-horizon robotic manipulation is often built by chaining independently trained skills. Although each skill can be reliable in isolation, performance degrades sharply when skills are chained: each downstream skill must start from the state its predecessor leaves behind rather than from its training distribution. We study this failure mode, Observation-Space Shift (OSS), and ask what causes these skill-seam failures. Using privileged simulator resets, we find that the dominant shift comes from displaced scene state (e.g., an open drawer or secondary objects left behind by earlier skills), not from the robot's joint configuration or the object the downstream skill manipulates. To test this diagnosis, we build a fully learned detect-restore-resume system: a task-progress monitor detects the stall, a learned policy restores the displaced scene components, and seam-robust fine-tuning lets the skill resume. It recovers the seam where every tested alternative fails, which we treat as evidence for the diagnosis rather than as a general-purpose method. On the BOSS-44 benchmark, the system improves full-chain success from 7.6% to 26.5%, a 3.5x improvement over the base policy and 51% of a privileged restoration oracle, whereas best-of-K resampling, a Diffusion Policy, and world-model baselines fail to recover from the evaluated seam states. On a real Franka arm running a fine-tuned $π_{0.5}$ policy, the same monitor is limited by exterior-camera observability, yet closing the loop still recovers some otherwise-terminal failures, motivating wrist and gripper sensing. These results suggest that some long-horizon composition failures are better addressed by restoring the scene before resuming the policy than by retrying from an off-support state.
Chinese Translation
长时程机器人操作通常通过串联独立训练的技能来构建。尽管每个技能在孤立情况下都能可靠运行,但当技能被串联起来时,性能会急剧下降:每个下游技能必须从前序技能留下的状态开始,而不是从其训练分布开始。我们研究这一失败模式,即观测空间偏移(Observation-Space Shift, OSS),并追问是什么导致了这些技能接缝失败。利用特权模拟器重置,我们发现主要偏移来自被移位的场景状态(例如,被前序技能留下的打开的抽屉或次要物体),而不是来自机器人的关节配置或下游技能所操作的物体。为了检验这一诊断,我们构建了一个完全学习得到的检测-恢复-继续系统:任务进度监测器检测停滞,学习到的策略恢复被移位的场景组件,接缝鲁棒微调让技能能够继续。它恢复了所有被测试替代方案都失败的接缝,我们将其视为对该诊断的证据,而不是一种通用方法。在 BOSS-44 基准上,该系统将全链条成功率从 7.6% 提升到 26.5%,相比基础策略提升了 3.5 倍,并达到特权恢复 oracle 的 51%,而 best-of-K 重采样、Diffusion Policy 和世界模型基线均无法从所评估的接缝状态中恢复。在运行微调后的 $π_{0.5}$ 策略的真实 Franka 机械臂上,同一监测器受到外部相机可观测性的限制,但闭环仍然恢复了一些原本会终止的失败,这促使我们采用腕部和夹爪传感。这些结果表明,对于一些长时程组合失败,在恢复策略之前先恢复场景,比从偏离支撑的状态重试更有效。
cs.AI / 121 / 2610.10962
iAm.md: Robot Skill Self-Assessment through Agentic Introspection for Unknown Open-Vocabulary Domains
iAm.md:面向未知开放词汇领域,通过智能体内省进行机器人技能自我评估
Vincenzo Guarino, Emanuele Musumeci, Vincenzo Suriani, Daniele Nardi
cs.RO · cs.AI
large language model
大语言模型相关
Abstract
Agentic AI based on Large Language Model generalization capabilities offers a wide range of potential applications, including planning for embodied tasks. For example, embodied agents based on Foundation models can generate plausible plans in autonomous robotics scenarios. Due to limited context windows or hallucinatory phenomena in the next-token prediction formulation, behaviors may be generated without establishing whether the deployed robot and the observed environment actually support the requested operation, in what we call a "grounding failure". Thanks to the recent improvements in reasoning capabilities of foundation models, autonomous robot behavior generation problem can be formulated as a code generation problem. We present iAm.md, a Markdown standard and generation framework, that allows anchoring this process in complementary forms of deployment evidence. Through open-vocabulary semantic mapping, we combine local vision-language detections and object segmentation and refer them to persistent object records in this intermediate standardized representation, allowing agentic introspection. We then study this new technique on a simulated TIAGo, on navigation-and-manipulation tasks, showing how this standardized representation jointly supports skill self-assessment and executable task generalization.
Chinese Translation
基于大语言模型泛化能力的智能体式人工智能提供了广泛的潜在应用,包括对具身任务进行规划。例如,基于基础模型的具身智能体能够在自主机器人场景中生成看似合理的计划。由于有限的上下文窗口或下一 token 预测形式中的幻觉现象,可能会在未确定所部署的机器人和所观察到的环境是否实际支持所请求的操作的情况下生成行为,我们将此称为“接地失败”(grounding failure)。得益于基础模型推理能力的近期提升,自主机器人行为生成问题可以被形式化为一个代码生成问题。我们提出 iAm.md,一种 Markdown 标准和生成框架,它允许将这一过程锚定在互补形式的部署证据中。通过开放词汇语义建图,我们结合局部视觉-语言检测和对象分割,并将它们关联到这一中间标准化表示中的持久对象记录,从而实现智能体内省。然后,我们在模拟的 TIAGo 上、在导航与操作任务中研究这一新技术,展示这一标准化表示如何共同支持技能自我评估和可执行任务泛化。
cs.LG / 122 / 2610.11401
WAM-Cache: Staleness-Bounded KV Reuse for Efficient World Action Models
WAM-Cache:面向高效世界动作模型的陈旧度受限 KV 复用
Kai Ding, Yang He, Ruijie Quan, Yi Yang
cs.RO · cs.CV · cs.LG
diffusion
扩散模型相关
Abstract
World Action Models (WAMs) enable generalist robot manipulation by conditioning an action expert on representations from a pretrained video Diffusion Transformer (DiT). In closed-loop control, the video DiT runs at every chunk to encode the current observation into layerwise key-value (KV) pairs that the action expert queries. This prefill dominates the per-chunk computational cost, yet existing training-free accelerations leave it fully dense. We present WAM-Cache, a training-free framework that retains layerwise key-value representations across chunks and recomputes only a sparse refresh set of tokens. Crucially, we find that the intuitive heuristic of refreshing visually drifted tokens plateaus far below the dense baseline, even with an oracle predicting ground-truth KV drift. Downstream action accuracy is instead governed by where the action expert attends, not by what moved. WAM-Cache therefore selects the refresh set by uniting the action expert's cross-attention with visual latent surprise, complemented by a strict age bound that suppresses compounding error. On Fast-WAM, WAM-Cache cuts video DiT prefill FLOPs by 32-42% across RoboTwin 2.0, LIBERO, and real-world experiments, while staying within 0.7-1.8 percentage points of the dense policy in simulation and 2.5 points on a real robot.
Chinese Translation
世界动作模型(WAMs)通过将动作专家以预训练视频 Diffusion Transformer(DiT)的表示为条件,实现通用机器人操作。在闭环控制中,视频 DiT 在每个 chunk 运行,以将当前观测编码为逐层的键值(KV)对,供动作专家查询。这种预填充主导了每个 chunk 的计算成本,然而现有的免训练加速方法仍使其保持完全稠密。我们提出 WAM-Cache,一个免训练框架,它跨 chunk 保留逐层键值表示,并仅重新计算一个稀疏的刷新 token 集合。关键的是,我们发现,刷新视觉上漂移的 token 这一直观启发式方法远低于稠密基线并趋于平台,即使使用预测真实 KV 漂移的 oracle 也是如此。相反,下游动作精度由动作专家关注的位置决定,而不是由什么发生了移动决定。因此,WAM-Cache 通过将动作专家的交叉注意力与视觉潜在 surprise 结合起来来选择刷新集合,并辅以一个严格的年龄界限来抑制累积误差。在 Fast-WAM 上,WAM-Cache 在 RoboTwin 2.0、LIBERO 和真实世界实验中把视频 DiT 预填充 FLOPs 降低 32-42%,同时在仿真中与稠密策略的差距保持在 0.7-1.8 个百分点以内,在真实机器人上保持在 2.5 个百分点以内。
cs.LG / 123 / 2610.11956
Reliability-Aware Future Conditioning for Temporally Robust Robot Manipulation
面向时间鲁棒机器人操作的可靠性感知未来条件化
Mohammad Khoshnazar, Mohammad Dehghani Tezerjani, Zhiyuan Gao, Deyuan Qu, Max Gandyra, Yanxiang Zhan, Mehreen Naeem, Andrew Melnik, Jeroen Schafer, Qing Yang, Michael Beetz
cs.RO · cs.LG
diffusion
扩散模型相关
Abstract
A generated video of a task the robot is about to perform is useful guidance only if it depicts the phase the robot is actually in. We show that temporal misalignment can turn a task-consistent generated future into actively harmful guidance. On CALVIN, a five-frame early shift nearly erases the benefit of generated futures, reducing success from 81.3% to 54.8% against 54.0% without futures; imposed timing shifts reduce it even further to 34.2%, 19.8 points below the future-free policy. We introduce Reliability-Aware Future Conditioning (RAFC), which treats this as a control problem rather than a generation problem. At every step, RAFC estimates how far to trust the received clip and which nearby temporal hypothesis to prefer, falling back toward a static branch when neither fits, and it learns both from task reward alone without shift labels or alignment supervision. RAFC sits on top of Future-Experience Conditioning (FEC), which builds the clip once from task grounding, a robot-free digital-twin rollout, and mask-free video diffusion. Under deliberately off-grid phase shifts and rate mismatch, RAFC substantially improves success under temporal mismatch. Candidate ensembling accounts for most of the recovery near alignment, while learned reliability adds a further 7.0 percentage points over uniform averaging of the identical candidate bank under off-grid shifts. The gain holds on the evaluated task sets and survives on a Franka under natural timing mismatch nobody imposed, where aggregate success rises from 26.7% to 56.7%. All resources will be made publicly available. https://future-condition.github.io/.
Chinese Translation
机器人即将执行任务的生成视频,只有当它描绘的是机器人实际所处的阶段时,才是有用的引导。我们表明,时间错位可以把一个与任务一致的生成未来变成具有主动危害性的引导。在 CALVIN 上,提前五帧的偏移几乎抹去了生成未来的收益,将成功率从 81.3% 降至 54.8%,而无未来的成功率为 54.0%;施加的时间偏移将其进一步降至 34.2%,比无未来的策略低 19.8 个百分点。我们提出可靠性感知未来条件化(RAFC),它将此视为一个控制问题而非生成问题。在每一步,RAFC 估计应在多大程度上信任所接收的片段,以及应偏好哪个邻近的时间假设,当两者都不匹配时回退到静态分支,并且它仅从任务奖励中学习这两者,无需偏移标签或对齐监督。RAFC 建立在未来经验条件化(FEC)之上,FEC 通过任务接地、无机器人的数字孪生推演以及无掩码视频扩散一次性构建该片段。在刻意设置的偏离网格的相位偏移和速率不匹配下,RAFC 显著提升了时间不匹配情形下的成功率。候选集成解释了接近对齐时的大部分恢复,而在偏离网格的偏移下,学习到的可靠性在相同候选库的均匀平均之上又额外带来 7.0 个百分点。该增益在所评估的任务集上得以保持,并且在 Franka 上于无人为施加的自然时间不匹配下依然成立,总体成功率从 26.7% 上升到 56.7%。所有资源将公开发布。https://future-condition.github.io/。
cs.AI / 124 / 2610.12407
LeWAM: A JEPA World Action Model with Diffusion-Steering-Based MPC
LeWAM:一种带有基于扩散引导的 MPC 的 JEPA 世界动作模型
Shashank Hegde, Alexander Popov, Elie Aljalbout, Nikolai Smolyanskiy
cs.RO · cs.AI · cs.CV
diffusion
扩散模型相关
Abstract
World action models (WAMs) predict actions and future observations, typically from a reconstruction-based representation that carries noisy, redundant information which can complicate downstream predictions. We introduce LeWAM, a bidirectional transformer for forward, backward, inverse dynamics and policy prediction, on a decoder-free JEPA latent trained end-to-end through all four modes. We see the following benefits: 1) Alignment: linear probes read robot and object state from LeWAM's latent better than from a regular Le World Model (a forward-only JEPA world model), while the latent ignores visual distractors as well as LeWM does and far better than a reconstruction-based WAM. 2) Acting: Closed-loop evaluations of LeWAM match a regular flow-matching policy trained on the same encoder at matched size, while also providing a world model. 3) Planning: Sampling raw actions when planning with WAMs lets MPC exploit dynamics-model inaccuracies; planning in the noise space of the policy head instead improves the closed-loop performance of these WAMs.
Chinese Translation
世界动作模型(WAM)预测动作与未来观测,通常基于一种重建式表示,这种表示携带噪声大、冗余的信息,可能使下游预测变得复杂。我们提出 LeWAM,一个用于前向动力学、后向动力学、逆向动力学与策略预测的双向 Transformer,它建立在无解码器的 JEPA 潜表示之上,并通过全部四种模式进行端到端训练。我们看到以下收益:1)对齐:线性探针从 LeWAM 的潜表示中读取机器人与物体状态的效果优于从常规 Le World Model(一个仅前向的 JEPA 世界模型)中读取,同时该潜表示对视觉干扰物的忽略程度与 LeWM 相当,且远优于基于重建的 WAM。2)行动:在相同规模下,LeWAM 的闭环评估结果与在同一编码器上训练的常规流匹配策略相当,同时还额外提供了一个世界模型。3)规划:在用 WAM 进行规划时采样原始动作会让 MPC 利用动力学模型的不准确性;而在策略头的噪声空间中进行规划,则能提升这些 WAM 的闭环性能。
cs.AI / 125 / 2610.11387
RISR: Residual-Informed Scientific Equation Discovery with Large Language Models
RISR:残差信息驱动的基于大语言模型的科学方程发现
Haobo Li, Wenshuo Zhang, Wenxiao Zhao, Eunseo Jung, Rui Sheng, Yushi Sun, Peiqin Zhuang, Hao Chen, Fenghua Ling
cs.SC · cs.AI
large language model
大语言模型相关
Abstract
Symbolic regression combines structural search with numerical fitting, but aggregate fit scores do not describe how the remaining error varies across inputs. We introduce RISR, a residual-informed method that uses these error patterns to guide formula discovery and learn which corrections are worth fitting. A residual encoder compresses aligned inputs, targets, current predictions, and residuals into continuous tokens that condition a language model to propose formulas. For subsequent refinement, a dual-view relational encoder uses additive and regularized multiplicative residuals to predict the post-fit utility of candidate corrections. We evaluate RISR on scientific tasks from the LLM-SRBench. RISR achieves 63.57% and 38.50% ID accuracy at the 1% and 0.1% pointwise relative-error tolerances, respectively. The corresponding OOD accuracies are 56.07% and 38.24%. RISR outperforms the reported baselines using the same backbone. The results show that our residual-informed approach can improve numerical equation recovery.
Chinese Translation
符号回归将结构搜索与数值拟合结合起来,但总体拟合分数并不能描述剩余误差如何随输入变化。我们提出 RISR,一种残差信息驱动的方法,它利用这些误差模式来指导公式发现,并学习哪些修正值得拟合。一个残差编码器将对齐后的输入、目标、当前预测和残差压缩为连续 token,这些连续 token 作为条件使语言模型提出公式。对于后续精炼,一个双视图关系编码器使用加性残差和正则化乘性残差来预测候选修正的拟合后效用。我们在来自 LLM-SRBench 的科学任务上评估 RISR。RISR 在 1% 和 0.1% 的逐点相对误差容限下,分别达到 63.57% 和 38.50% 的 ID 准确率。相应的 OOD 准确率分别为 56.07% 和 38.24%。RISR 在使用相同骨干网络的情况下优于所报告的基线。结果表明,我们的残差信息驱动方法能够改进数值方程恢复。
cs.CL / 126 / 2610.11461
Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech
超越语音描述:面向对话式文本到语音的语音奖励风格规划
Shiao Zhu, Lianbo Liu, Sizhen Lyu, Yuzhe Wang, Sheng Li, Takahiro Shinozaki
cs.SD · cs.CL · eess.AS
large language model
大语言模型相关
Abstract
Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS). However, using descriptions as pseudo-labels compresses target acoustics into text, and descriptive fidelity need not imply effective control of a particular synthesizer. We empirically show that speech-text alignment only weakly predicts downstream acoustic similarity among candidate instructions for the same utterance. We therefore propose Speech-Rewarded Style Planning (SRSP), which trains a text-based style planner through a frozen downstream TTS model. Given dialogue history and response text, the planner generates candidate instructions and is optimized with group-relative policy optimization (GRPO), using the teacher-forced likelihood of target speech tokens as the reward. On an English subset of the ISCSLP 2026 CoT-TTS corpus, SRSP achieves higher speech-style and emotion similarity to target speech and lower mel-cepstral distortion than the Base LLM and target-audio-informed captioning baselines. LLM-based expressive speech evaluation further shows gains over all baselines in contextual appropriateness and reference consistency.
Chinese Translation
自然语言风格描述为大型语言模型(LLMs)与可控文本到语音(TTS)之间提供了一个可解释的接口。然而,将描述用作伪标签会把目标声学信息压缩进文本,而描述层面的保真度并不必然意味着对某一特定合成器的有效控制。我们通过实证表明,语音-文本对齐只能微弱地预测同一话语的候选指令之间的下游声学相似度。因此,我们提出语音奖励风格规划(Speech-Rewarded Style Planning,SRSP),它通过一个冻结的下游 TTS 模型来训练一个基于文本的风格规划器。给定对话历史和回复文本,该规划器生成候选指令,并使用组相对策略优化(GRPO)进行优化,以目标语音 token 的教师强制似然作为奖励。在 ISCSLP 2026 CoT-TTS 语料库的英语子集上,与 Base LLM 和以目标音频为条件的描述生成基线相比,SRSP 在语音风格和情感上与目标语音的相似度更高,且梅尔倒谱失真更低。基于 LLM 的表现力语音评估进一步表明,其在上下文适当性和参考一致性方面均优于所有基线。
cs.CL / 127 / 2610.12214
DiffuPlex: Accelerating Full-Duplex Spoken Dialog Models via Rolling Masked Diffusion
DiffuPlex:通过滚动掩码扩散加速全双工语音对话模型
Heeseung Kim
cs.SD · cs.CL
diffusion
扩散模型相关
Abstract
Recent full-duplex spoken dialog models enable simultaneous listening and speaking, but fine-grained models still advance their backbone autoregressively at every interaction frame. We introduce DiffuPlex, a rolling masked diffusion framework that reduces this sequential computation by predicting multiple future user and assistant frames in a single backbone wake. DiffuPlex consumes only a confident prefix of each predicted future while interaction continues at the original frame rate. As user speech arrives, it checks the corresponding user predictions and, when the interaction diverges, preserves already played assistant content while revising only the unplayed future. We consider two inference policies over the same predictor: DiffuPlex-LISTEN consumes multiple future frames when they predict assistant silence, whereas DiffuPlex-SPEAK can also consume predicted assistant speech. Across full-duplex interaction and spoken-language evaluations, DiffuPlex substantially reduces sequential backbone computation while largely preserving interaction behavior and general capability. DiffuPlex-LISTEN and DiffuPlex-SPEAK achieve $1.46\times$ and $1.59\times$ deployment-path wall-clock speedups and $1.61\times$ and $1.80\times$ Core LM speedups, with all measured backbone invocations completing within the 80ms interaction interval. Human evaluation shows that LISTEN preserves speech naturalness and conversational quality, while SPEAK retains conversational quality with some degradation in speech naturalness.
Chinese Translation
近期的全双工语音对话模型能够同时聆听和说话,但细粒度模型仍然在每个交互帧上以自回归方式推进其骨干网络。我们提出 DiffuPlex,一种滚动掩码扩散框架,它通过在单次骨干网络唤醒中预测多个未来的用户和助手帧,减少了这种顺序计算。DiffuPlex 仅消耗每个预测未来中置信度高的前缀,同时交互以原始帧率继续进行。当用户语音到达时,它会检查对应的用户预测,并且当交互发生偏离时,保留已经播放的助手内容,只修正尚未播放的未来内容。我们在同一预测器上考虑两种推理策略:DiffuPlex-LISTEN 在预测到助手静音时消耗多个未来帧,而 DiffuPlex-SPEAK 还可以消耗预测的助手语音。在全双工交互和口语语言评估中,DiffuPlex 大幅减少了顺序骨干计算,同时很大程度上保持了交互行为和通用能力。DiffuPlex-LISTEN 和 DiffuPlex-SPEAK 分别实现了 $1.46\times$ 和 $1.59\times$ 的部署路径墙钟时间加速,以及 $1.61\times$ 和 $1.80\times$ 的核心 LM 加速,所有测量到的骨干调用都在 80ms 交互间隔内完成。人工评估表明,LISTEN 保持了语音自然度和对话质量,而 SPEAK 在语音自然度有所下降的情况下仍保持了对话质量。
cs.AI / 128 / 2610.12250
How Much Audio Is Left In An Embedding? An Inversion Audit Of Audio Encoders
嵌入中还剩下多少音频?对音频编码器的反演审计
Marios Glytsos, Brian McFee
cs.SD · cs.AI
diffusion
扩散模型相关
Abstract
Pretrained audio encoders are reused for downstream tasks that are often unknown when the encoder is trained, so their usefulness depends partly on which signal properties survive the pretext objective. We study this retained information through paired source reconstruction. Using a shared Stable Audio Open latent diffusion decoder, we reconstruct five-second, 44.1-kHz stereo music from frozen representations produced by supervised classifiers (VGGish, ConvNeXt), an audio-text contrastive model (CLAP), and a waveform-reconstruction model (EnCodec). These objectives impose different pressures to preserve source detail, while their exposed interfaces vary substantially in temporal and spectral resolution. Evaluating on the Million Song Dataset (MSD), we find clear differences in reconstructability across encoder families, while within encoder comparisons show improved recovery when finer temporal or spectral structure is exposed. Even compressed task oriented embeddings support reconstructions that preserve measurable source specificity and high level musical content.
Chinese Translation
预训练的音频编码器会被复用于下游任务,而这些任务在编码器训练时往往是未知的,因此它们的实用性部分取决于哪些信号属性能够在预训练目标中得以保留。我们通过配对源重建来研究这些被保留下来的信息。使用一个共享的 Stable Audio Open 潜在扩散解码器,我们从由监督分类器(VGGish、ConvNeXt)、一个音频-文本对比模型(CLAP)以及一个波形重建模型(EnCodec)所产生的冻结表示中重建出五秒、44.1-kHz 的立体声音乐。这些目标对保留源细节施加了不同的压力,而它们所暴露的接口在时间分辨率和频谱分辨率上差异很大。在 Million Song Dataset(MSD)上进行评估时,我们发现不同编码器家族之间在可重建性上存在明显差异,而在同一编码器内部的比较则表明,当暴露出更精细的时间或频谱结构时,重建恢复效果会有所提升。即使是压缩的、面向任务的嵌入,也支持那些能够保留可测量的源特定性以及高层音乐内容的重建。
cs.SE / 129 / 2610.11300
Characterizing Overconfident Failure in LLM-Based Code Generation
刻画基于LLM的代码生成中的过度自信失败
Ravishka Rathnasuriya, Wei Yang
cs.SE · cs.AI
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly used for automated code generation, but generated programs can appear syntactically plausible while still failing execution-based correctness checks. Existing validation methods, such as testing and program analysis, remain essential but are often incomplete, costly, or applied only after generation. Model-derived uncertainty is therefore a natural early reliability signal. This paper studies the dilemma of overconfidence in code LLMs where incorrect programs are often generated with token-level confidence comparable to correct programs. We study this dilemma across four open-source code models and three execution-based benchmarks. Our analysis begins by investigating whether existing uncertainty metrics provide reliable proxies for execution correctness in code generation. We then characterize overconfidence at both global and local token levels, asking whether incorrect programs remain indistinguishable from correct ones under confidence and entropy summaries, including selective generation and the limits of instruction tuning. Finally, we evaluate whether common mitigation strategies reduce this failure mode. Our study yields four findings. First, existing uncertainty signals provide only partial and model-dependent evidence of execution failure. Second, overconfidence persists at both program and token levels, and uncertainty-based selection does not consistently improve accepted-set accuracy. Third, instruction tuning can increase certainty on failing generations without consistently improving correctness discrimination. Fourth, common mitigation techniques improve specific aspects of reliability but do not reliably resolve overconfident failure. Our exploratory latent analysis suggests that hidden representations may encode correctness-related signals that output confidence does not expose.
Chinese Translation
大型语言模型(LLM)正越来越多地被用于自动代码生成,但生成的程序可能在语法上看似合理,却仍然无法通过基于执行的正确性检查。现有的验证方法,例如测试和程序分析,仍然不可或缺,但往往不完整、成本高昂,或仅在生成之后才被应用。因此,由模型导出的不确定性是一种自然的早期可靠性信号。本文研究了代码LLM中过度自信的困境,即不正确的程序往往以与正确程序相当的词元级置信度被生成。我们在四个开源代码模型和三个基于执行的基准上研究了这一困境。我们的分析首先考察现有的不确定性度量是否为代码生成中的执行正确性提供可靠的代理指标。随后,我们从全局和局部词元两个层面刻画过度自信,追问在置信度和熵的汇总下——包括选择性生成和指令微调的局限——不正确的程序是否仍与正确的程序无法区分。最后,我们评估常见的缓解策略是否能够减少这一失败模式。我们的研究得出四项发现。第一,现有的不确定性信号仅能提供关于执行失败的部分且依赖模型的证据。第二,过度自信在程序层面和词元层面都持续存在,而基于不确定性的选择并不能一致地提升被接受集合的准确率。第三,指令微调可能提高失败生成结果上的确定性,却不能一致地改善对正确性的判别。第四,常见的缓解技术改善了可靠性的特定方面,但并不能可靠地解决过度自信失败。我们的探索性潜在分析表明,隐藏表示可能编码了输出置信度未暴露的、与正确性相关的信号。
cs.SE / 130 / 2610.11447
Closed-loop evaluation of LLM agents for embedded software development
面向嵌入式软件开发的 LLM 智能体闭环评估
Jorge García-Carrasco, Sergio García-Carrasco, Alejandro Maté, Juan Trujillo
cs.SE · cs.AI · cs.AR · cs.LG
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly deployed as coding agents that edit files, run builds and tests, inspect execution results, and repair software iteratively. Embedded firmware is a demanding target because correctness depends on closed-loop behavior under sensing, timing, and safety constraints, not only on static source quality. Yet embedded-agent evaluation remains limited and often emphasizes one-shot synthesis or offline correctness. We present a benchmark for closed-loop evaluation of embedded coding agents. Each task provides a plain-text engineering description, constrained workspace, and visible build-and-runtime surface. The agent must translate requirements into implementation and self-verification steps, then iterate until the required device behavior is achieved. The suite contains five embedded-control tasks and four feedback scenarios: one-shot generation, realistic self-verification, CI-style red/green feedback, and oracle-style detailed feedback. The implementation targets simulated ESP32 firmware for reproducibility. We evaluate seven GPT-family and Qwen-family configurations across five tasks and four scenarios, with three repetitions per condition for 420 runs. gpt-5.4 has the highest pass rate among evaluated configurations but does not saturate the benchmark; qwen3.5-27B is the strongest observed local model; and smaller local models degrade sharply in pass rate and search efficiency. These results suggest that capable local embedded coding agents are emerging.
Chinese Translation
大语言模型(LLMs)正日益被部署为编码智能体,这些智能体可以编辑文件、运行构建和测试、检查执行结果,并迭代式地修复软件。嵌入式固件是一个要求严苛的目标,因为其正确性取决于在感知、时序和安全约束下的闭环行为,而不仅仅取决于静态源代码质量。然而,嵌入式智能体评估仍然有限,并且常常强调一次性合成或离线正确性。我们提出了一个用于嵌入式编码智能体闭环评估的基准。每个任务都提供纯文本工程描述、受限工作区以及可见的构建与运行时表面。智能体必须将需求转化为实现和自验证步骤,然后迭代,直到实现所需的设备行为。该套件包含五个嵌入式控制任务和四种反馈场景:一次性生成、真实的自我验证、CI 风格的红/绿反馈,以及 oracle 风格的详细反馈。该实现以模拟的 ESP32 固件为目标,以保证可复现性。我们在五个任务和四种场景下评估了七个 GPT 系列和 Qwen 系列配置,每个条件重复三次,共 420 次运行。gpt-5.4 在所评估配置中具有最高通过率,但并未使该基准饱和;qwen3.5-27B 是观察到的最强本地模型;而更小的本地模型在通过率和搜索效率上急剧下降。这些结果表明,有能力的本地嵌入式编码智能体正在出现。
cs.SE / 131 / 2610.11482
Evaluating Local Language Model Agents for Reproducible Data Engineering: An Empirical Software Engineering Study of Mobility Workflows
评估用于可复现数据工程的本地语言模型智能体:一项关于移动性工作流的实证软件工程研究
Jorge García-Carrasco, Javier Sanchis, Alejandro Reina-Reina, Alejandro Maté, Juan Trujillo
cs.SE · cs.AI · cs.DB · cs.LG
large language model
大语言模型相关
Abstract
Context: Large language model (LLM) agents are increasingly used as software and data-engineering assistants, yet evidence about locally deployable open-weight agents remains limited. Existing evaluations often emphasize textual responses or isolated code generation rather than the validity of complete engineering artifacts. Objectives: We evaluate whether local LLM agents can produce correct and reproducible data-engineering artifacts, quantify the effect of a closed-loop workspace condition, and examine trade-offs in model scale, architecture, quantization, runtime, tool use, and failure. Methods: We introduce a benchmark of fifteen mobility-workflow tasks covering data discovery, connectors, transport-feed processing, semantic enrichment, feature engineering, validation, visualization, and reporting. Deterministic checkers assess generated scripts, tables, structured files, figures, and reports. Ten local configurations are evaluated in one-shot and closed-loop conditions, with five repetitions per model, mode, and task, yielding 1,500 scored attempts on a consumer-grade GPU. Results: Among models larger than two billion parameters, the workspace condition increases pass rates by 26.7-52.0 percentage points over one-shot generation. The strongest configuration reaches 85.3% artifact-level success, and a quantized 9-billion-parameter model reaches 69.3% with an approximately 6.5 GB memory footprint. Gains are largest when intermediate artifacts expose errors the agent can inspect and repair. Conclusion: Local open-weight agents can support a meaningful subset of software-intensive data-engineering work, but reliability depends on model capability, task verifiability, and deterministic validation. The benchmark provides a reproducible method for evaluating complete agent configurations before adoption in engineering workflows.
Chinese Translation
背景:大语言模型(LLM)智能体正越来越多地被用作软件与数据工程助手,但关于可本地部署的开放权重智能体的证据仍然有限。现有评估通常强调文本回复或孤立的代码生成,而不是完整工程制品的有效性。目标:我们评估本地 LLM 智能体能否生成正确且可复现的数据工程制品,量化闭环工作区条件的影响,并考察模型规模、架构、量化、运行时、工具使用和失败之间的权衡。方法:我们引入一个包含十五个移动性工作流任务的基准,涵盖数据发现、连接器、交通数据流处理、语义增强、特征工程、验证、可视化和报告。确定性检查器评估生成的脚本、表格、结构化文件、图形和报告。十个本地配置在一次性(one-shot)和闭环条件下进行评估,每个模型、模式和任务重复五次,在消费级 GPU 上产生 1,500 次评分尝试。结果:在参数量大于二十亿的模型中,工作区条件相比一次性生成将通过率提高 26.7-52.0 个百分点。最强配置达到 85.3% 的制品级成功率,而一个量化后的 90 亿参数模型以约 6.5 GB 内存占用达到 69.3%。当中间制品暴露出智能体可以检查和修复的错误时,收益最大。结论:本地开放权重智能体可以支持软件密集型数据工程工作中一个有意义的子集,但其可靠性取决于模型能力、任务可验证性和确定性验证。该基准提供了一种可复现的方法,用于在工程工作流中采用之前评估完整的智能体配置。
cs.SE / 132 / 2610.11578
Chronos Enables Code Agents to Reason over Software Evolution
Chronos 使代码智能体能够对软件演化进行推理
Xin Yin, Yiang Zhang, Zhiyuan Peng, Chao Ni, Zhe Cui, Xiaohua Xin
cs.SE · cs.CL
large language model
大语言模型相关
Abstract
Historical pull requests record the design decisions, compatibility constraints, and implementation patterns behind a codebase's current state. Experience relevant to a new task can span related changes whose descriptions emphasize different concerns. We introduce Chronos, a test-time framework that makes this connected history available to large language model (LLM)-based code agents. Chronos distills merged pull requests into structured experience cards and connects them through a typed graph of code-level, developer-intent, and organizational relations. Semantic search identifies entry cards, and weighted multi-hop expansion retrieves connected changes for selective reading. The same memory guides candidate generation and patch selection: a patch-focused change agent and a validation-strategy agent each develop a patch, and an evolution steward consults history to select between them. On SWE-Bench Verified, the full workflow improves SWE-Agent across all six evaluated LLM backbones, raising the mean resolution rate from 69.2% to 72.9% and reaching 79.8% with MiniMax M2.5. With the same backbone, it raises resolution rates from 48.3% to 51.7% on SWE-Bench Pro and from 41.0% to 43.5% on FEA-Bench Lite. Both experience-guided single-agent variants also outperform the base agent. In a human evaluation on 100 tasks with ten cards retrieved per task, graph-grounded retrieval increases the mean number of useful cards from 1.24 to 2.87 over flat semantic retrieval. These results demonstrate the value of PR relations for retrieving useful repository experience and of the evaluated workflows for applying that experience during patch generation and selection.
Chinese Translation
历史 pull request 记录了代码库当前状态背后的设计决策、兼容性约束与实现模式。与新任务相关的经验可能跨越多个相关变更,而这些变更的描述强调的关注点各不相同。我们提出 Chronos,一个测试时框架,它使这种相互关联的历史可供基于大语言模型(LLM)的代码智能体使用。Chronos 将已合并的 pull request 提炼为结构化的经验卡片,并通过一个由代码层、开发者意图与组织关系构成的有类型图将它们连接起来。语义搜索确定入口卡片,加权多跳扩展则检索出相互关联的变更以供选择性阅读。同一记忆同时指导候选生成与补丁选择:一个聚焦补丁的变更智能体和一个验证策略智能体各自开发一个补丁,而演化管家则查阅历史以在二者之间做出选择。在 SWE-Bench Verified 上,完整工作流在所评估的全部六个 LLM 基座模型上都提升了 SWE-Agent,将平均解决率从 69.2% 提升至 72.9%,并在 MiniMax M2.5 上达到 79.8%。在相同基座模型下,它将在 SWE-Bench Pro 上的解决率从 48.3% 提升至 51.7%,在 FEA-Bench Lite 上从 41.0% 提升至 43.5%。两种由经验引导的单智能体变体也均优于基础智能体。在一项涵盖 100 个任务、每个任务检索十张卡片的人工评估中,相较于扁平语义检索,基于图的检索将有价值卡片的平均数量从 1.24 提升至 2.87。这些结果证明了 PR 关系在检索有用的代码库经验方面的价值,以及所评估工作流在补丁生成与选择过程中应用这些经验的价值。
cs.SE / 133 / 2610.11618
PolyCodeEval: Benchmarking Multilingual Code Generation from Functions to Repositories
PolyCodeEval:从函数到仓库的多语言代码生成基准测试
Bowen Yang, Jiajun Jiang, Luxue Yu, Yihao Wang, Fengjie Li, Dong Wang
cs.SE
large language model
大语言模型相关
Abstract
As large language models increasingly move toward repository-level software engineering, existing code-generation benchmarks remain fragmented across language coverage, task granularity, and evaluation protocols, impeding systematic comparison. To address this gap, we present PolyCodeEval, a unified multilingual and multi-granularity benchmark for code generation. It comprises 2,590 code generation tasks spanning functions to repositories, derived from 58 real, executable open-source repositories in five programming languages. All tasks are evaluated under a unified execution-based protocol with integration procedures tailored to their generation targets. Building on this benchmark, we evaluate frontier large language models, state-of-the-art specialized methods, and general coding agents. Our results show that existing approaches still struggle to correctly generate complete code fragments across granularities and languages. Specifically, the studied methods generate at most 71.7%, 76.7%, and 31.0% correct functions, files, and repositories, respectively, with performance varying widely across languages. Paired experiments further show that implementation context from related functions in the same file improves the executable correctness of function generation. Method rankings also vary across task granularities and programming languages, highlighting the importance of multilingual, multi-granularity evaluation for comprehensively assessing code generation capabilities.
Chinese Translation
随着大型语言模型日益迈向仓库级软件工程,现有的代码生成基准在语言覆盖范围、任务粒度和评估协议方面仍然各自为政,阻碍了系统性比较。为弥补这一空白,我们提出了 PolyCodeEval,一个统一的、多语言且多粒度的代码生成基准。它包含 2,590 个代码生成任务,跨越从函数到仓库的多个层面,源自五种编程语言中的 58 个真实、可执行的开源仓库。所有任务都在统一的基于执行的协议下进行评估,并配有针对其生成目标量身定制的整合流程。在此基准的基础上,我们评估了前沿大型语言模型、最先进的专用方法以及通用编程智能体。我们的结果表明,现有方法仍难以跨粒度和跨语言正确生成完整的代码片段。具体而言,所研究的方法在函数、文件和仓库上分别最多生成 71.7%、76.7% 和 31.0% 的正确结果,且不同语言之间的性能差异很大。配对实验进一步表明,来自同一文件中相关函数的实现上下文可提升函数生成的可执行正确性。方法排名也因任务粒度和编程语言而异,凸显了多语言、多粒度评估对于全面评估代码生成能力的重要性。
cs.SE / 134 / 2610.11727
MAP4CS: A Multi-dimensional Data Pruning Framework for Efficient Code Retriever Fine-tuning
MAP4CS:面向高效代码检索器微调的多维数据剪枝框架
Yuxuan Chen, Mingwei Liu, Guangsheng Ou, Zekai Zhang, Zike Li, Yanlin Wang, Pelin Zheng
cs.SE · cs.AI
large language model
大语言模型相关
Abstract
Retrieval-Augmented Generation (RAG) has become a cornerstone in software engineering for enhancing Large Language Models (LLMs) with domain-specific knowledge. However, adapting retrievers to evolving code repositories remains challenging due to the noise and redundancy inherent in massive code corpora. Standard fine-tuning on the full corpus is computationally expensive and often leads to sub-optimal performance due to negative transfer from low-quality samples. Conversely, simple random sampling fails to guarantee data representativeness. To address these challenges, we propose MAP4CS (Multi-dimensional Awareness Pruning for Code Search), an adaptive data pruning framework. MAP4CS identifies a small, high-quality core subset by integrating syntactic structure, semantic diversity, and distributional representation, followed by a rigorous rule-based filtering pipeline. Extensive experiments on two large-scale datasets demonstrate that MAP4CS consistently outperforms random sampling baselines using only 5% of the training data. Remarkably, it achieves performance comparable to, or even superior to, fine-tuning on the full dataset, validating the ''less is more'' hypothesis in data-centric AI. Furthermore, linguistic analysis reveals an adaptive optimization mechanism: MAP4CS automatically functions as a de-duplicator for redundant corpora and a denoiser for chaotic ones, constructing a training corpus that is both lexically diverse and information-dense.
Chinese Translation
检索增强生成(RAG)已成为软件工程中利用领域特定知识增强大语言模型(LLM)的基石。然而,由于海量代码语料库中固有的噪声与冗余,使检索器适应不断演化的代码仓库仍然充满挑战。在全量语料库上进行标准微调计算开销高昂,且常常由于低质量样本带来的负迁移而导致次优性能。相反,简单的随机采样无法保证数据的代表性。为应对这些挑战,我们提出了 MAP4CS(Multi-dimensional Awareness Pruning for Code Search,面向代码搜索的多维感知剪枝),一种自适应数据剪枝框架。MAP4CS 通过整合句法结构、语义多样性与分布表示,并辅以严格的基于规则的过滤流程,识别出一个规模小、质量高的核心子集。在两个大规模数据集上的大量实验表明,MAP4CS 仅使用 5% 的训练数据便持续优于随机采样基线。值得注意的是,它达到了与在全量数据集上微调相当、甚至更优的性能,验证了以数据为中心的人工智能中“少即是多”的假设。此外,语言学分析揭示了一种自适应优化机制:MAP4CS 会自动对冗余语料库充当去重器,对混乱语料库充当去噪器,从而构建出一个兼具词汇多样性与信息密度的训练语料库。
cs.SE / 135 / 2610.11801
ObliVul: Alert-Conditioned Safety Obligation Modeling and Bidirectional Counterfactual Validation for Code Vulnerability Detection
ObliVul:面向代码漏洞检测的以告警为条件的安全义务建模与双向反事实验证
Heyang Tan, Chengxin Gao, Xin Wen, Jiaxin Li, Rui Cao
cs.SE
large language model
大语言模型相关
Abstract
In real-world software development, the primary challenge in vulnerability detection is often not finding suspicious code, but identifying which alerts among the large number of candidate alerts produced by static analysis truly warrant attention. Existing learning-based methods mainly identify suspicious patterns at the function or line level, making it difficult to extract complete program evidence centered on an individual alert. Although large language models can infer risk sources, dangerous operations, protection conditions, and state preconditions from local program facts, such semantic information cannot be reliably aligned with specific program nodes, dependency relations, and propagation paths, and is therefore insufficient to verify whether the corresponding safety obligations truly affect the current alert. To address this problem, we propose ObliVul, an alert-conditioned safety obligation modeling and bidirectional counterfactual validation framework for vulnerability detection. For each candidate alert, ObliVul first extracts a Local Evidence Pack (LEP) from the Code Property Graph (CPG) and uses a large language model to recover candidate safety obligations. It then aligns the safety obligations with program nodes, dependency edges, and path scopes to construct a Local Safety Obligation Graph (LSOG). Finally, VAFI aggregates complementary verified alert evidence while suppressing redundant or weaker evidence to produce a function-level vulnerability prediction. Experimental results show that ObliVul effectively distinguishes vulnerable versions from fixed versions and reduces persistent false positives on fixed code. Ablation studies further confirm the necessity of each component for safety obligation recovery, risk response validation, and function-level vulnerability inference.
Chinese Translation
在真实世界的软件开发中,漏洞检测面临的主要挑战往往不是发现可疑代码,而是从静态分析产生的大量候选告警中识别出哪些告警真正值得关注。现有基于学习的方法主要在函数级或行级识别可疑模式,难以提取以单个告警为中心的完整程序证据。尽管大语言模型能够从局部程序事实中推断出风险来源、危险操作、保护条件与状态前置条件,但此类语义信息无法可靠地与具体的程序节点、依赖关系和传播路径对齐,因而不足以验证相应的安全义务是否真正影响当前告警。为解决这一问题,我们提出了 ObliVul,一个面向漏洞检测的以告警为条件的安全义务建模与双向反事实验证框架。对于每个候选告警,ObliVul 首先从代码属性图(CPG)中提取局部证据包(LEP),并利用大语言模型恢复候选安全义务。随后,它将安全义务与程序节点、依赖边和路径范围对齐,以构建局部安全义务图(LSOG)。最后,VAFI 聚合互补的已验证告警证据,同时抑制冗余或较弱的证据,以产生函数级漏洞预测。实验结果表明,ObliVul 能够有效区分易受攻击版本与修复版本,并减少在修复代码上持续出现的误报。消融研究进一步证实了各组件对于安全义务恢复、风险响应验证和函数级漏洞推断的必要性。
cs.SE / 136 / 2610.11963
Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes
大语言模型能否无需代码地修复它?迈向无代码缺陷修复的自动化验证
Utku Boran Torun, Veli Karakaya, Eray Tüzün
cs.SE · cs.AI
large language model
大语言模型相关
Abstract
A no-code fix resolves an invalid bug report by directing the user to change a setting, update to a version where the problem is already fixed, or adjust their workflow. Manually verifying whether a proposed no-code fix resolves the reported bug takes considerable developer time. This study proposes an automated, execution-based pipeline for evaluating the capability of large language models (LLMs) to generate no-code fixes in a real browser environment. We evaluate 322 no-code fixes generated by the 12 configurations released with the benchmark of a previous study, covering bug reports categorized as Faulty Configuration, Wrong Version, or External System & Dependency. An executor agent applies each fix by following its natural-language instructions, and an issue-specific checker determines whether the reported bug persists. We repeat the pipeline with three executors: two Computer-Use Agents, OpenCUA-72B and Claude Sonnet 5, and one multimodal agentic LLM, Meta's Muse Glimmer. Only 17.6% of the candidate issues could be set up and passed both sanity gates. Across the 322 fixes, 14.6% to 49.7% resolved the bug depending on the executor, and the strongest configuration, Claude Opus 4.6 in the Vanilla pipeline, resolved up to 74.1% of its fixes under Claude Sonnet 5. Changing only the executor shifted a configuration's resolution rate by 38.8% on average, and the three executors reached the same verdict on only 46.9% of the fixes. Compared with human execution, the executors matched the human consensus for 66.1% to 88.1% of the sampled fixes. Even under the best executor, fewer than half of the LLM-generated no-code fixes resolve the reported bug, so such fixes need verification before they reach users. Execution-based verification can provide this, but the measured capability depends strongly on the executor, which evaluations must report and control.
Chinese Translation
无代码修复通过引导用户更改某项设置、升级到问题已被修复的版本,或调整其工作流程,来解决无效的缺陷报告。人工验证所提出的无代码修复是否解决了所报告的缺陷会耗费开发者大量时间。本研究提出了一种基于执行的自动化流水线,用于评估大语言模型(LLM)在真实浏览器环境中生成无代码修复的能力。我们评估了由先前研究的基准所发布的 12 种配置生成的 322 个无代码修复,涵盖被归类为错误配置、错误版本或外部系统与依赖的缺陷报告。一个执行器智能体通过遵循每个修复的自然语言指令来应用该修复,而一个针对具体问题的检查器则判定所报告的缺陷是否仍然存在。我们使用三个执行器重复该流水线:两个计算机使用智能体 OpenCUA-72B 和 Claude Sonnet 5,以及一个多模态智能体式大语言模型 Meta 的 Muse Glimmer。只有 17.6% 的候选问题能够被搭建起来并通过两道健全性门控。在这 322 个修复中,根据执行器的不同,有 14.6% 到 49.7% 解决了缺陷,而最强的配置——Vanilla 流水线中的 Claude Opus 4.6——在 Claude Sonnet 5 下解决了其修复中的多达 74.1%。仅更换执行器就会使某个配置的解决率平均变动 38.8%,并且三个执行器仅在 46.9% 的修复上得出一致判定。与人工执行相比,执行器在 66.1% 到 88.1% 的抽样修复上与人类共识相匹配。即便在最佳执行器下,也只有不到一半的 LLM 生成的无代码修复解决了所报告的缺陷,因此此类修复在到达用户之前需要验证。基于执行的验证可以提供这一点,但所测得的能力在很大程度上取决于执行器,评估必须报告并控制这一因素。
cs.SE / 137 / 2610.12289
TestPrism: Rethinking Test Evaluation Beyond a Single Reference
TestPrism:超越单一参考重新思考测试评估
Han Li, Lingxiang Hu, Jiacheng Huang, Ziqian Jiang, Jingkai Luo, Wei Gao, Yunfan Tan, Zun Wang, Jiaheng Liu
cs.SE
large language model
大语言模型相关
Abstract
Large language model (LLM) coding agents have advanced test generation across diverse programming tasks. However, the common practice of evaluating tests against a single reference solution overlooks alternative valid implementations and can overstate test quality. We introduce TestPrism, comprising 300 test tasks from 17 sources and 3000 candidate implementations, evenly split between valid and invalid solutions. Its primary metric, Joint Success Function, requires the generated tests to fail on the initial program state, accept every valid candidate, and reject every invalid candidate. Across fourteen baseline coding agent configurations, Joint Success Function reaches only 28.00%, whereas single reference success reaches 59.67%. Our analysis reveals missed behaviors, unsupported assertions, and faulty test construction. To address these weaknesses, we introduce TestHelix, which combines heterogeneous synthesis of test and repair pairs with peer cross validation and recursive self improvement (RSI). Across two models, TestHelix improves Joint Success Function by 8.67 to 9.00 percentage points over the native harness comparators in the TestHelix evaluation
Chinese Translation
大语言模型(LLM)编码智能体已在各种编程任务中推进了测试生成。然而,依据单一参考解决方案评估测试的常见做法忽视了其他有效的实现,并可能高估测试质量。我们介绍 TestPrism,它包含来自 17 个来源的 300 个测试任务和 3000 个候选实现,在有效解与无效解之间均等划分。其主要指标联合成功函数(Joint Success Function)要求生成的测试在初始程序状态下失败,接受每一个有效候选,并拒绝每一个无效候选。在十四个基线编码智能体配置中,联合成功函数仅达到 28.00%,而单一参考成功率则达到 59.67%。我们的分析揭示了遗漏的行为、不受支持的断言以及有缺陷的测试构造。为了解决这些弱点,我们引入 TestHelix,它将测试与修复对的异构合成与同行交叉验证以及递归自我改进(RSI)相结合。在两个模型上,TestHelix 在 TestHelix 评估中相较于原生测试框架比较对象,将联合成功函数提高了 8.67 到 9.00 个百分点。
cs.CL / 138 / 2610.10868
Conversational Voice Aesthetic Model with Reinforcement Learning from Human Listeners
基于人类听者强化学习的对话式语音美学模型
Xilin Jiang, Shun Zhang, Tejas Jayashankar, Yinghao Aaron Li, Osama Hanna
eess.AS · cs.CL · cs.LG · cs.MM · cs.SD
large language model
大语言模型相关
Abstract
We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts. Given a context and a response speech, CVAM describes salient moments that characterize the voice and predicts nine categorical attributes spanning gender, pitch, pacing, emotion, and delivery. The key challenge lies in perceptual fields such as emotion and delivery, which are inherently subjective and lack definitive ground truth. Therefore, we collect ~10 human annotations for each of 3k real and synthetic responses derived from the CANDOR corpus. CVAM is supervised finetuned on synthesized aesthetic descriptions and labels, then optimized with Group Relative Policy Optimization on human judgments. Experiments show that CVAM better agrees with human listeners than Gemini 3.1 Pro and open-source speech LLMs, and outperforms single-human-vs.-rest agreement. Together, we demonstrate the importance of grounding voice aesthetics in human perception and propose a principled framework for human alignment.
Chinese Translation
我们介绍了对话式语音美学模型(Conversational Voice Aesthetic Model),这是一个语音大语言模型,用于描述自然对话语境中真实或合成语音响应的语音美学。给定一个语境和一个响应语音,CVAM 描述刻画该语音的显著时刻,并预测涵盖性别、音高、节奏、情感和表达方式的九个类别属性。关键挑战在于情感和表达方式等感知领域,这些领域本质上是主观的,并且缺乏确定性的真值。因此,我们为源自 CANDOR 语料库的 3k 个真实和合成响应中的每一个收集约 10 个人类标注。CVAM 在合成的美学描述和标签上进行有监督微调,然后基于人类判断使用 Group Relative Policy Optimization 进行优化。实验表明,CVAM 比 Gemini 3.1 Pro 和开源语音大语言模型更好地与人类听者一致,并且优于单个人类与其余人类之间的一致性。总之,我们展示了将语音美学建立在人类感知中的重要性,并提出了一个用于人类对齐的原则性框架。
cs.AI / 139 / 2610.11768
Narrow and Deep: An Ontology Tower as the Knowledge of an LLM Agent for an Industrial Equipment System
窄而深:本体塔作为工业设备系统 LLM 智能体的知识
Younghwan Joo, Sung-il Kim
eess.SY · cs.AI
large language model
大语言模型相关
Abstract
Large language model (LLM) agents are beginning to operate industrial energy equipment, and what they get right depends on what they are told about the plant. Established building ontologies name many kinds of points across many sites, whereas an industrial equipment system needs few entities with much knowledge about each. This study proposes the ontology tower, a narrow-and-deep ontology of a single equipment system whose knowledge deepens in two ways: through quantities derived from the measured points by physical relations, and through lessons from the operating journal incorporated as knowledge nodes. On a real low-humidity air-handling test plant operated daily through a programmable logic controller, agents received a text projected from its tower in a preregistered evaluation of nine tasks replayed from the plant's records, using four open-weight models from 9 to about 750 billion parameters. This knowledge raised the rate at which the agents avoided the most plausible misjudgment of each task by about 20 percentage points, and the overall task score of the 9-billion-parameter model as much as that of the largest. Operating lessons were used when incorporated into the tower or placed in the prompt as records, but seldom when left in the journal behind a search tool. In live runs through an invariant safety layer, the agents brought the controlled variable into its target band in 12 of 14 runs. An ontology narrow in entities but deep in what is known about them can thus supply the knowledge that an agent for an industrial equipment system needs.
Chinese Translation
大型语言模型(LLM)智能体正开始操作工业能源设备,而它们做对什么取决于它们被告知了关于该工厂的哪些信息。已有建筑本体命名了跨许多站点的许多类型的点,而工业设备系统需要很少的实体,但需要关于每个实体的丰富知识。本研究提出本体塔,一种单一设备系统的窄而深本体,其知识通过两种方式加深:通过物理关系从测量点推导出的量,以及通过将运行日志中的经验教训作为知识节点纳入。在一个真实的低湿度空气处理试验装置上,该装置通过可编程逻辑控制器每日运行,在一次对九项任务的预注册评估中,智能体接收了从其本体塔投影出的文本,这九项任务是从该装置的记录中重放的,并使用四个开放权重模型,参数规模从 90 亿到约 7500 亿不等。这一知识将智能体避免每项任务中最可能出现的误判的比率提高了约 20 个百分点,并使 90 亿参数模型的总体任务得分与最大模型的得分一样高。运行经验在并入本体塔或作为记录放入提示词时会被使用,但当被留在搜索工具背后的日志中时则很少被使用。在通过不变安全层的实时运行中,智能体在 14 次运行中有 12 次将被控变量带入其目标区间。因此,一种实体少但关于它们的知识深的本体,能够为工业设备系统的智能体提供其所需的知识。
cs.AI / 140 / 2610.12293
Machine Learning Meets High-Energy Nuclear Physics: From Pattern Recognition to Physics-Integrated Discovery
机器学习与高能核物理相遇:从模式识别到物理集成的发现
Xun Chen, Weiyao Ke, Yu-Gang Ma, Long-Gang Pang, Kai Zhou
hep-ph · cs.AI · hep-lat · hep-th · nucl-th
diffusion
扩散模型相关
Abstract
Machine learning (ML) in high-energy nuclear physics (HENP) is entering a new stage in which physical knowledge is incorporated more directly into data analysis, simulation, and physics inference. This mini-review focuses on developments that have matured in the past several years. Whereas earlier applications emphasized event classification, pattern recognition, and surrogate models for selected observables, recent work has moved toward physics-integrated workflows: calibrated Bayesian extraction of QCD matter properties, dense-matter equation-of-state inference from heavy-ion and neutron-star data, generative event modeling, neural unfolding of weak physical signals, differentiable inverse solvers, gauge-equivariant and diffusion-based lattice-field samplers, and neural reconstruction of model functions in holographic QCD. We survey recent applications of ML in heavy-ion collisions, neutron-star physics, lattice QFT, and holographic or continuum QCD. The emphasis is not on ML architectures alone, but on how they enter concrete physics workflows, how physical constraints such as symmetries, conservation laws, causality, thermodynamic stability, and topology are imposed, and how uncertainty quantification and validation determine whether an AI-assisted result can support a reliable physics conclusion.
Chinese Translation
高能核物理(HENP)中的机器学习(ML)正在进入一个新阶段,在这一阶段,物理知识被更直接地纳入数据分析、模拟与物理推断之中。本篇小型综述聚焦于过去几年中已趋于成熟的发展。早期应用强调事件分类、模式识别以及针对选定观测量的代理模型,而近期的工作则转向物理集成的工作流程:QCD 物质性质的校准贝叶斯提取、从重离子与中子星数据推断致密物质状态方程、生成式事件建模、微弱物理信号的神经解卷、可微逆求解器、规范等变与基于扩散的格点场采样器,以及全息 QCD 中模型函数的神经重建。我们综述了 ML 在重离子碰撞、中子星物理、格点量子场论以及全息或连续 QCD 中的近期应用。重点并不只在于 ML 架构本身,而在于它们如何进入具体的物理工作流程、诸如对称性、守恒定律、因果性、热力学稳定性与拓扑等物理约束如何被施加,以及不确定度量化与验证如何决定一项 AI 辅助的结果能否支撑可靠的物理结论。
cs.AI / 141 / 2610.12067
MAST: Motif-Augmented Diffusion with Search Tree for Spectroscopic Molecular Structure Elucidation
MAST:用于光谱分子结构解析的基元增强扩散与搜索树
Chenghao Jia, Mengdi Liu, Hong Chang, Shiguang Shan, Xilin Chen
physics.chem-ph · cs.AI
diffusion
扩散模型相关
Abstract
Elucidating molecular structures from spectra is a foundational problem in chemical and materials characterization, yet remains challenging due to spectral ambiguity and the vast molecular space. Although recent diffusion-based generators show strong promise for spectra-conditioned elucidation, existing methods struggle to learn robust spectra-structure relationships from limited paired data when relying solely on global spectral representation. Moreover, the repeated full sampling inference strategy incurs substantial computation overhead. To address these limitations, we propose \textbf{MAST}, a \textbf{M}otif-\textbf{A}ugmented diffusion framework with \textbf{S}earch \textbf{T}ree, for joint 2D-3D spectroscopic molecular structure elucidation. MAST introduces explicit, interpretable \emph{motif priors} as intermediate evidences throughout denoising, reducing conditional ambiguity and facilitating spectra-conditioned optimization. We further cast diffusion sampling as \emph{reward-guided tree search} to prioritize high-reward denoising trajectories, yielding a compact set of spectra-consistent candidates under limited budgets. On the QM9S multi-spectra benchmark, MAST achieves \textbf{94.89\%} exact recovery and improves 3D fidelity, while preserving high chemical validity and stability. Code is available at https://github.com/Jia040223/MAST.
Chinese Translation
从光谱中解析分子结构是化学与材料表征中的一个基础性问题,然而由于光谱的歧义性以及庞大的分子空间,这一问题仍然极具挑战性。尽管近期的基于扩散的生成器在光谱条件解析方面展现出很大的潜力,但当仅依赖全局光谱表示时,现有方法难以从有限的配对数据中学习到鲁棒的光谱-结构关系。此外,重复的完整采样推理策略会带来大量的计算开销。为解决这些局限,我们提出了 \textbf{MAST},一个\textbf{基}元-\textbf{增}强扩散框架,并带有\textbf{搜}索\textbf{树},用于联合的 2D-3D 光谱分子结构解析。MAST 在去噪全过程中引入显式且可解释的\emph{基元先验}作为中间证据,从而降低条件歧义并促进光谱条件优化。我们进一步将扩散采样表述为\emph{奖励引导的树搜索},以优先选择高奖励的去噪轨迹,从而在有限预算下得到一组紧凑的、与光谱一致的候选结构。在 QM9S 多光谱基准上,MAST 实现了 \textbf{94.89\%} 的精确恢复率,并提升了 3D 保真度,同时保持了较高的化学有效性与稳定性。代码可在 https://github.com/Jia040223/MAST 获取。
cs.AI / 142 / 2610.12281
Unlocking the Regulatory Genome by ARGUS: An Evidence-Constrained Agentic Framework for Interpreting Single Nucleotide Variants
以 ARGUS 解锁调控基因组:一种用于解读单核苷酸变异的证据约束型智能体框架
Pratik Dutta, Matthew B. Obusan, Max Chao, Rekha Sathian, Nimisha Papineni, Ramana V. Davuluri
q-bio.GN · cs.AI · cs.LG
large language model
大语言模型相关
Abstract
Over 90% of disease-associated variants from genome-wide association studies fall in noncoding regulatory regions, yet their functional interpretation remains a central open problem in genomic medicine. Large language models prompted to interpret such variants routinely hallucinate transcription factor (TF) binding changes, fabricate experimental support, and assign biological significance to statistically negligible signals. We present ARGUS (Agentic Regulatory Genomics for an Uncertainty-aware Scientist), which strictly separates deterministic biological computation from LLM-mediated reasoning. ARGUS wraps 458 DNABERT-based TF binding models in a hypothesis-directed investigation loop where a planner selects evidence sources based on current uncertainty, a verifier deterministically interprets each observation, and intermediate results change the investigation path. On variant rs6983267 at the 8q24 cancer risk locus, the same planner produces four divergent trajectories for four TFs. FOXA1 is rescued in 3 steps when real ADASTRA allele-specific binding data (15 experiments, FDR = 0.030) reveals a model false negative masked by saturation. KLF6 traverses 8 steps across ADASTRA, JASPAR motif analysis, and ENCODE cCRE regulatory annotation before abstaining due to mixed indirect evidence. RAD21 abstains in 8 steps after ADASTRA returns a coverage-qualified but nonsignificant allelic test (5 experiments, FDR = 0.65), and SP1, which shares FOXA1's saturated retained prediction, abstains because no direct experimental evidence exists at this locus. All observations come from real ADASTRA, JASPAR, and ENCODE cCRE queries; none are simulated. A comparison of fixed-priority and LLM-mediated planning shows that the LLM planner reaches identical verdicts with fewer tool calls by declining evidence that cannot resolve the claim under test.
Chinese Translation
来自全基因组关联研究的疾病相关变异中,超过 90% 位于非编码调控区,然而其功能解读仍然是基因组医学中一个核心的未解决问题。被提示解读此类变异的大型语言模型经常幻觉出转录因子(TF)结合变化、编造实验支持,并将生物学意义赋予统计上可忽略的信号。我们提出 ARGUS(面向不确定性感知科学家的智能体调控基因组学),它严格地将确定性生物计算与 LLM 介导的推理分开。ARGUS 将 458 个基于 DNABERT 的 TF 结合模型封装在一个假设导向的调查循环中,其中规划器根据当前不确定性选择证据来源,验证器确定性地解释每个观察结果,并且中间结果会改变调查路径。在 8q24 癌症风险位点的变异 rs6983267 上,同一个规划器为四个 TF 产生四条不同的轨迹。当真实的 ADASTRA 等位基因特异性结合数据(15 项实验,FDR = 0.030)揭示出一个被饱和掩盖的模型假阴性时,FOXA1 在 3 步内被挽救。KLF6 在 ADASTRA、JASPAR 基序分析和 ENCODE cCRE 调控注释中历经 8 个步骤,之后因混合的间接证据而弃权。RAD21 在 ADASTRA 返回一个覆盖度合格但不显著的等位基因检验(5 项实验,FDR = 0.65)后,于 8 步内弃权;而 SP1 与 FOXA1 共享其饱和的保留预测,它弃权是因为该位点不存在直接实验证据。所有观察结果都来自真实的 ADASTRA、JASPAR 和 ENCODE cCRE 查询;没有一个是模拟的。对固定优先级规划和 LLM 介导规划的对比表明,LLM 规划器通过拒绝那些无法解决受检验主张的证据,以更少的工具调用达到相同的判定。
cs.LG / 143 / 2610.11407
Beyond Distributional Fidelity: Causal-Penalized Diffusion for Synthetic Tabular Data
超越分布保真度:面向合成表格数据的因果惩罚扩散
Lan Tao, Yongxian He, Shirong Xu, Yidong Ouyang, Guang Cheng
stat.ML · cs.LG
diffusion
扩散模型相关
Abstract
Synthetic tabular generators are commonly optimized for distributional fidelity, but statistical similarity alone does not guarantee preservation of causal effects. In this paper, we study whether causal fidelity can be improved directly within a fully generative tabular model. Causal Fidelity is defined with respect to a target estimand as the discrepancy between inferential distributions obtained from real and synthetic data, and theoretical results show that high statistical fidelity does not generally imply high causal fidelity. We then propose a causal-fidelity-aware training framework which adds a causal discrepancy penalty to the generative objective. The framework is instantiated with a causal-penalized TabDDPM and optimized using an on-policy score-function estimator. We further establish conditions under which causal regularization improves expected causal fidelity. Experiments across diverse treatment-effect simulations and two benchmark datasets evaluate the ability of our method to improve causal fidelity while preserving competitive statistical fidelity.
Chinese Translation
合成表格生成器通常针对分布保真度进行优化,但仅有统计相似性并不能保证因果效应的保持。在本文中,我们研究是否可以在完全生成式表格模型内部直接提升因果保真度。因果保真度是相对于目标估计量定义的,即从真实数据和合成数据中获得的推断分布之间的差异,并且理论结果表明,高统计保真度通常并不意味着高因果保真度。我们随后提出一种因果保真度感知的训练框架,该框架在生成目标中加入因果差异惩罚。该框架通过因果惩罚的 TabDDPM 实例化,并使用同策略得分函数估计器进行优化。我们进一步建立了因果正则化能够改善期望因果保真度的条件。在多样的处理效应模拟和两个基准数据集上的实验评估了我们的方法在保持有竞争力的统计保真度的同时提升因果保真度的能力。
cs.LG / 144 / 2610.11459
Feature Space Adaptation for Effortless Gaussian Process Flows
面向轻松高斯过程流的特征空间自适应
Thomas Cowperthwaite, Louis Sharrock, Lachlan Astfalck, Henry Moss
stat.ML · cs.LG · stat.ME
diffusion
扩散模型相关
Abstract
Outside the linear-Gaussian regime, conditional sampling from Gaussian processes (GPs) is challenging. Recent methods such as FlowGP (Moss et al., (2026)) can condition on arbitrary non-linear and non-Gaussian statements, but at considerable cost: an expensive iterative and high-dimensional diffusion that requires hand-specified kernel hyperparameters. In this paper, we alleviate two significant drawbacks of FlowGP by (1) introducing kernel approximations that enable scaling to high-resolution domains and (2) proposing a way to obtain the marginal likelihood by measuring the work needed to steer the diffusion towards conditioning statements. We enable, for the first time, hyperparameter optimisation within FlowGP and demonstrate our approach on probabilistic downscaling from areal summary statistics, PDE solution inference on irregular domains, and recovery of sea level anomaly fields from non-Gaussian satellite observations.
Chinese Translation
在线性-高斯情形之外,从高斯过程(GP)中进行条件采样具有挑战性。近期的方法,例如 FlowGP(Moss 等人,(2026)),能够以任意非线性和非高斯陈述为条件,但代价相当大:需要一次昂贵的迭代式高维扩散,并且需要人工指定核超参数。在本文中,我们通过以下两点缓解了 FlowGP 的两个显著缺点:(1) 引入核近似,使其能够扩展到高分辨率域;(2) 提出一种通过测量将扩散引导至条件陈述所需做的功来获得边际似然的方法。我们首次实现了 FlowGP 内部的超参数优化,并在以下任务上展示了我们的方法:从区域汇总统计量出发的概率降尺度、不规则区域上的 PDE 解推断,以及从非高斯卫星观测中恢复海平面异常场。
cs.LG / 145 / 2610.12052
Diffusion Removes Langevin's Conditioning Dependence: A Sharp Gaussian Analysis
扩散模型消除了朗之万动力学的条件数依赖:一个精确的高斯分析
Adam Perbost, Francis Bach, Pierre Marion
stat.ML · cs.LG
diffusion
扩散模型相关
Abstract
Despite their empirical success, why diffusion models overcome the bottlenecks of classical score-based samplers remains unclear. In this work, we leverage Gaussian distributions to isolate this phenomenon. We establish 2-Wasserstein convergence bounds for optimized hyperparameters, showing that diffusion processes achieve a sampling error of $O(\sqrt{dλ_{\max}}\log N/N)$, where $d$ is the dimension, $N$ the number of sampling steps, and $λ_{\max}$ the largest eigenvalue of the target covariance matrix. Unadjusted and underdamped Langevin dynamics suffer from an additional $\sqrtκ$ factor, where $κ$ is the condition number. These rates follow from spectral bounds which are sharp: we confirm them via matching first-order asymptotics as $N\rightarrow\infty$. Our analysis provides a rigorous characterization, in the Gaussian setting, of how time-dependent score trajectories remove condition-number dependence during sampling. By contrast, in the learning phase, we show that estimating the unnoised score by gradient descent leads to essentially the same estimator as estimating a noisy score, which suggests that the benefits of noising do not come from the learning phase.
Chinese Translation
尽管扩散模型在经验上取得了成功,但它们为何能够克服经典基于分数的采样器的瓶颈,仍不清楚。在本工作中,我们利用高斯分布来隔离这一现象。我们为优化后的超参数建立了 2-Wasserstein 收敛界,表明扩散过程达到了 $O(\sqrt{dλ_{\max}}\log N/N)$ 的采样误差,其中 $d$ 是维度,$N$ 是采样步数,$λ_{\max}$ 是目标协方差矩阵的最大特征值。未调整的与欠阻尼的朗之万动力学则额外承受一个 $\sqrtκ$ 因子,其中 $κ$ 是条件数。这些速率源自谱界,而这些谱界是精确的:我们通过匹配 $N\rightarrow\infty$ 时的一阶渐近行为来验证它们。我们的分析在高斯设定下,严格刻画了依赖于时间的分数轨迹如何在采样过程中消除条件数依赖。相比之下,在学习阶段,我们表明通过梯度下降估计无噪声分数所得到的估计量,与估计含噪分数所得到的估计量本质上相同,这表明加噪带来的益处并非来自学习阶段。
cs.LG / 146 / 2610.12200
Quickest Change Detection with Diffusion-Integrated Scores
基于扩散集成分数的最快变化检测
Arman Adibi, Mohammadreza Maleki, Sanjeev Kulkarni, H. Vincent Poor
stat.ML · cs.LG
diffusion
扩散模型相关
Abstract
Classical CUSUM relies on the log-likelihood ratio of the underlying distributions, which cannot generally be computed from finite pre- and post-change samples alone. We propose diffusion-integrated score CUSUM (DI-SCUSUM), a training-free detector. We add Gaussian noise to the samples to form two smooth density estimates and calculate their Hyvärinen scores exactly, without training a score network. For each incoming observation, we sample a diffusion time, perturb the observation, and use the importance-weighted score difference as an increment in the DI-SCUSUM recursion. Under the assumption that observations follow the fixed empirical distributions, the post-change mean increment is proportional to the Kullback-Leibler (KL) divergence from the smoothed post-change to the smoothed pre-change empirical distribution. We establish exponential false-alarm scaling and a first-order delay bound that, for a fixed threshold and increment scaling, is inversely proportional to the KL divergence. In the calibrated anisotropic Gaussian simulation, DI-SCUSUM nearly matches likelihood-ratio CUSUM and reduces the measured detection delay by about 91% relative to score-based CUSUM. On MNIST and Oxford-IIIT Pet, DI-SCUSUM also has lower empirical conditional detection delay than SCUSUM at comparable false-alarm levels.
Chinese Translation
经典 CUSUM 依赖于底层分布的对数似然比,而该对数似然比通常无法仅由有限的变前和变后样本计算得到。我们提出扩散集成分数 CUSUM(DI-SCUSUM),一种免训练检测器。我们向样本中加入高斯噪声,以形成两个平滑密度估计,并精确计算它们的 Hyvärinen 分数,而无需训练分数网络。对于每个到来的观测,我们采样一个扩散时间,扰动该观测,并将重要性加权的分数差用作 DI-SCUSUM 递推式中的增量。在观测服从固定经验分布的假设下,变化后的平均增量与从平滑后的变化后经验分布到平滑后的变化前经验分布的 Kullback-Leibler(KL)散度成正比。我们建立了指数虚警缩放和一个一阶延迟界;对于固定阈值和增量缩放,该延迟界与 KL 散度成反比。在校准的各向异性高斯仿真中,DI-SCUSUM 几乎匹敌似然比 CUSUM,并且相对于基于分数的 CUSUM 将测得的检测延迟降低了约 91%。在 MNIST 和 Oxford-IIIT Pet 上,在相当的虚警水平下,DI-SCUSUM 也比 SCUSUM 具有更低的经验条件检测延迟。
人工智能 (cs.AI)
142
cs.AI / 1 / 2610.10805
Whose Ground Truth? Embracing Ambiguity in Human-Centered AI
Jingyao Wu, Mohammad Tariqul Islam, Per Rådberg Nagbøl, Julie Gerlings, Giovanni Leoni, Georgiana Fryza, Lars Pilegaard Thomsen, Jakob Mainz
cs.AI · cs.LG
Abstract
As AI systems increasingly interact with people and make decisions about them, understanding human interpretations becomes an important part of developing human-centered AI. Conventional machine learning and AI systems are largely developed under the assumption that a single definitive ground truth exists, with variability in human annotations often resolved through aggregation or treated as noise. However, for many human-centered tasks, human interpretation is inherently ambiguous, and multiple interpretations of the same input may be simultaneously reasonable and valid. Reducing such ambiguity to a single target risks overlooking meaningful information about the diversity of human perception, judgment, and experience. In this position paper, we call for a shift towards modeling the interpretation space of plausible human judgments, while distinguishing meaningful ambiguity from annotation noise. We argue that this perspective should guide how AI systems are represented, learned, evaluated, deployed, and governed, supporting more human-centered AI that better reflects the diversity of human interpretation.
cs.AI / 2 / 2610.10833
On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents
Aaron Wang, Neelabh Madan, Vlad Sobal, Matthew Trager, Michael Kleinman, Elman Mansimov, Wei Xia, Stefano Soatto
cs.AI · cs.LG
Abstract
We study whether small LLM agents can operate effectively under explicit wall-clock time budgets by both respecting the allocated runtime and using available time productively. We evaluate Qwen3.6-27B on five competitions from MLE-Bench Lite and Qwen3-4B on Zork I (Jericho), two agentic benchmarks where additional computational time can meaningfully improve performance. In the simplest setting, where the budget is stated only in the prompt, agents fail to translate the stated budget into controlled use of time. These failures arise from gaps in time awareness, since the harness provides no timing feedback, but also because they cannot reliably anticipate the duration of actions, and do not have a learned mapping from available time to an appropriate strategy. We investigate two complementary classes of interventions: harness-based mechanisms that expose timing information and enforce deadlines, and reinforcement learning with budget-aware rewards. Injecting timing information through the harness substantially improves budget adherence for Qwen3.6-27B without measurable loss in performance, while enforcement hooks tighten adherence further. RL with GRPO achieves near-perfect budget adherence on Zork I and generalizes to held-out budgets not seen during training, but does not improve task performance over the untrained harness on MLE-Bench. Once agents are made to respect the budget, they still fail to use additional time to improve task performance. RL-trained policies learn when to stop but often fill extra time with repeated actions, and GRPO training on multiple budgets tends to collapse toward the strategy learned for the shortest budget. Our results reveal a gap between time adherence and productive time allocation, which remains a central challenge for budget-conditioned agents.
cs.AI / 3 / 2610.10857
Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
Prabin Kumar Rath, Omkar Patil, Nakul Gopalan
cs.AI · cs.RO
Abstract
Behavior cloning (BC) in non-Markovian environments is a challenging problem because policies have to reason over contextual information over long horizons. Existing policy architectures rely on recurrent or attention-based mechanisms to capture long-term dependencies. However, recurrent models suffer from hidden-state collapse and gradient instability under backpropagation through time, while attention-based models are fundamentally limited by context length. To address these issues, we propose Keyframe Mnemonics, a novel self-supervised method that $\textit{discovers}$ a set of information-critical observations ($\textit{mnemonics}$) by learning an objective from randomly sampled past observations and using it as a reward for keyframe selection. We then train a BC policy that conditions on the discovered keyframes to model the action distribution. Under certain task-structure assumptions, our formulation provides context retention guarantees over an infinite horizon, while maintaining a small set of decision-relevant keyframes in the policy's working memory. We evaluate our method on synthetic memory domains, where mnemonic-conditioned BC policies achieve $100$% success rates (SR) and generalize to horizons orders of magnitude beyond training without performance degradation. Additionally, we evaluate on memory-intensive robot manipulation benchmark, achieving a $13.9$% average absolute SR improvement over the strongest baseline across $23$ tasks and retaining $80$% SR at $20\times$ longer horizons on a real robot. Code and videos are available at https://keyframe-mnemonics.github.io.
cs.AI / 4 / 2610.10906
Reading the Room: Foundations, Design, and Challenges of Normative Competence in LLMs
Andrea Wynn, Harsh Satija, Seokhyun, Baek, Anqi Liu, Eric Nalisnick, Gillian K. Hadfield
cs.AI · cs.MA
Abstract
Human communities are governed by normative systems: shared standards that produce \textit{norms} dictating acceptable behavior, enforced through community sanctioning. Aligning increasingly autonomous AI systems with these norms is a central alignment challenge, complicated by the fact that norms are vast in number, change quickly, and are often arbitrary (e.g., dress or language conventions). Thus, alignment requires \textit{normative competence}: the ability to discern from interaction alone what norms a community enforces without relying on static pretrained knowledge. We introduce a multi-agent community debate setting, where access to debate is governed by synthetic norms, to study normative competence in isolation from pretraining exposure. We show that baseline LLM agents fail to learn norms even when doing so would improve their accuracy. We then experiment with various \textit{normative modules} -- architectural components for norm inference -- finding that norm-following is highly sensitive to both the style of norm and the model powering the normative module, suggesting a lack of generalizability. Furthermore, when idiosyncratic, non-normative behaviors accompany the true norm, LLM agents exhibit an unselective attribution failure: they indiscriminately copy idiosyncratic noise alongside enforced rules, a pattern that persists even when imitating unnecessary behaviors is explicitly penalized. To the best of our knowledge, our work is the first to operationalize and evaluate normative competence in LLMs, demonstrating that current AI systems excel at behavioral mimicry but lack the capacity to discern socially enforced order.
cs.AI / 5 / 2610.10954
Learning How to Search for Plans with Exponentially Less Space
Dominik Drexler, Simon Ståhlberg, Markus Fritzsche, Blai Bonet
cs.AI
Abstract
Heuristic search for a plan can store exponentially many states, even when its heuristic is almost perfect. We instead learn search control, one specification per domain, written as an indexical policy: a generalized policy with registers that hold objects and modes that sequence its rules. We add the choose rule, which loads an object into a register and marks a backtracking point, where one candidate suffices; every other rule must work for all of its outcomes and needs no search. Our main result is that structural termination, which rules out infinite executions, also bounds every execution by a polynomial in the number of objects. A depth-first procedure then finds a plan in polynomial space, however large the state space, with no list of visited states. The cost is time, exponential only in the choice depth, the number of real choices along an execution. Any class that such a policy solves therefore lies in NP, and in P at constant choice depth. We learn these policies with a language model in a counterexample-guided loop that certifies termination, verifies the training tasks, and keeps the choice depth small. With the learned policies, the procedure solves 1,709 of 1,890 test tasks of the IPC 2023 Learning Track and the Autoscale Agile suite, more than LAMA, BFWS, and Levitron, and most of them within one second and 100 MiB.
cs.AI / 6 / 2610.11005
How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense
Zhankai Ye, Yanning Wang, Yukai Jin, Bo Mei, Fangyi Li, Wei Wang, Shangqian Gao, Xin Liu
cs.AI
Abstract
Safety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper. We measure this vulnerability across languages and registers: attack success on Qwen3-1.7B is already 89.4% in English and 93.0% in modern Chinese, and reaches 95.7% in Classical Chinese. We build GUISE, a benchmark for systematically studying this vulnerability. It includes parallel requests in English, modern Chinese, and Classical Chinese, matched harmful and benign pairs, wrapper types held out for evaluation, and a stricter criterion that counts warn-then-answer responses as attack successes. Representation analysis shows that language and register move harmful-request representations only slightly away from the model's refusal direction, whereas narrative wrappers move them much farther away. We propose AXIS, which combines preference optimisation with a rotation objective that aligns harmful-request representations with the refusal direction and a commitment objective that trains the model to refuse completely rather than produce a warn-then-answer response. Across Qwen3-1.7B, Qwen3-4B and GLM-4-9B, AXIS achieves the highest combined safety and usability score among the compared methods.
cs.AI / 7 / 2610.11007
Curating Always-Loaded Context for LLM Agents: A Capacitated Assortment Model with Censored Feedback
Zexuan Liu, Yuning Yang, Tiancheng Zhao
cs.AI · cs.LG · math.OC
Abstract
At the start of every session, LLM agents load a fixed context file, such as $\texttt{AGENTS.md}$. Each loaded token in the file is charged again in every later round of the session, and these files can degrade performance as they grow in size. However, in practice, human or automated curators usually grow these files by appending. We formulate context curation as a capacitated assortment problem. Instructions consume tokens under a finite attention capacity; adding an instruction never raises the compliance of the others, while retained instructions incur a per-session setup cost. We prove an upper bound on the optimal file size, regardless of the number of available candidate instructions, and that appending every instruction with positive standalone value can be arbitrarily worse in net value than selecting an optimal subset. A token budget also limits the loss when the token price is underestimated. We then examine what can be learned from past sessions and how this information can guide decisions to add or remove instructions. Feedback is inherently censored: the benefits and harms of loaded instructions are observable, whereas missing instructions generate feedback only when their absence causes harm. In this setting, we show that deleting instructions ignored by agents can inevitably remove helpful ones. We characterize how much evidence should be collected before adding an instruction. Besides, we bound regret when human reviewers can inspect only a limited number of edits per period. Empirical experiments further show that irrelevant rules drawn from real context files reduce language-model compliance.
cs.AI / 8 / 2610.11012
Distillation for Incrimination and Distillation for Capabilities
Sebastian Prasanna, Jacqueline Tay, Alek Westover
cs.AI
Abstract
Powerful misaligned AI models might recognize alignment evaluations and strategically behave well on them, rendering direct audits uninformative. However, distilling such a model into a weaker benign student places the teacher in a Distillation Double Bind: if misalignment transfers, the student may conceal it less effectively, revealing evidence about the teacher; if it does not, the student may learn useful capabilities while remaining benign. We introduce two distinct distillation approaches, one targeting each outcome. Distillation for Incrimination (DFI) aims to transfer misalignment but not the ability to conceal it. Distilling AuditBench's secret-keeping models into their underlying instruction-tuned model produces students that are significantly more likely than their teachers to admit their hidden behavior when asked, suggesting that knowledge of the behavior transferred more readily than the propensity to conceal it. Confession gains largely disappear when the student does not share the teacher's pretrained base, so DFI should target the teacher's own pre-RL checkpoint, which is weaker than the teacher but shares its base model. Distillation for Capabilities (DFC) aims to transfer capabilities but not misalignment. Among several techniques we evaluate, two are effective: inoculation prompting and training for more epochs on fewer unique examples. Both preserve the capability gains of standard distillation while substantially reducing the subliminal transfer of an animal preference, our proxy for misalignment. Together, these findings demonstrate two ways distillation can be used for AI safety: incriminating misaligned models, and extracting their capabilities without their misalignment.
cs.AI / 9 / 2610.11050
AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks
Xing Han Lù, Dheeraj Vattikonda, Sina Hajimiri, Fatemeh Pesaran Zadeh, Parishad BehnamGhader, Ghazwa Darwiche, Amirhossein Kazemnejad, Christopher Pal, Alexandre Drouin, Siva Reddy
cs.AI · cs.LG
Abstract
Computer-use agents are capable of completing complex tasks, increasing the use of automatic judges to determine success, either for training or for evaluation without human involvement. Despite their flexibility, their reliability on long tasks spanning multiple applications remains unclear. A trajectory, composed of long sequences of screenshots and actions, may appear complete, but in reality violates constraints from the instruction or introduces an unwanted side effect. To identify these errors, a judge needs to examine the trajectory with respect to the user's instruction. To this end, we introduce AgentHorizon, a benchmark of 1,373 computer-use tasks (instruction-trajectory pairs) drawn from 166 hours of human-recorded trajectories spanning three operating systems. By recording trajectories for closely related instructions, we can construct negative tasks by swapping the instructions. This paired design evaluates judges on their ability to distinguish a successful trajectory from one that completed a similar (but incompatible) request. We release the benchmark under three splits: a frontier split, AgentHorizon (AH), a simplified split, AgentHorizon-Simple (AH-S), and a development split, AgentHorizon-Development (AH-D). We further evaluate eleven judges by (1) directly passing the full trajectory (with up to 300 screenshots and actions), and (2) using them as coding agents across five agent harnesses. We find that our best agentic judge, GPT-5.5, achieves 80.9% balanced accuracy on the AH subset. We find that tool-use improves certain models but results in worse performance for open-weight models, and that judges differ drastically in their ability to accept a valid trajectory and reject failed ones. Our findings highlight the need for judges that are capable of locating and verifying often hidden evidence that a task was properly completed inside long interaction histories.
cs.AI / 10 / 2610.11118
OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences
Zhiyi Li, Sihan Hu, Tianning Xiao, Xiansheng Cai, Xiaojun Tan, Youjin Deng, Kun Chen
cs.AI
Abstract
The next frontier for artificial general intelligence is tackling unresolved scientific problems, calling for benchmarks that assess progress beyond established knowledge. We introduce OpenProblemBench, a benchmark of 82 unresolved problems drawn from the mathematics and theoretical physics literature. Each problem supplies the research context, assumptions, and prior progress needed to investigate the question. We select problems whose proposed solutions admit comparatively clear checks of their decisive mathematical or computational claims. Four evaluator models independently assess the correctness, completeness, and degree of progress of each submission without reference solutions. Across seven evaluated configurations, GPT-6-Astra achieves the highest mean judged solve rate of 14.0%, compared with 5.5-6.7% for the evaluated full-size open models and 2.4-3.7% for Flash models. Case comparisons connect stronger outcomes to changes in problem representation, general arguments that extend beyond finite evidence, and proofs of the steps needed to complete a solution. By grounding evaluation in questions arising from the research literature, OpenProblemBench provides a setting for investigating the capabilities and limitations of AI as a contributor to foundational theoretical science.
cs.AI / 11 / 2610.11123
When Interfaces Speak: Data-Aware Generative UI Harness for Active Interaction
Xiaolong Li, Xiaohan Xu, Jinyang Li, Xinnuo Xu, Ge Qu, Nan Huo, Jack Williams, Reynold Cheng
cs.AI
Abstract
Most human-agent interaction today remains text-based. Natural language can impose cognitive overload, ambiguity, information chaos, and slow input for complex tasks; ephemeral generative UIs can present structured information and guide users toward task completion. We propose GenUI-Harness, a multi-agent harness pairing a Tool Agent for information retrieval and task execution with a GUI Coder Agent that identifies ambiguities and generates front-end code for structured interfaces. Training the coder with reinforcement learning is challenging: verifiable rewards for interactive UI generation require costly execution, while LLM-as-a-Judge rewards are prone to reward hacking. We address the first challenge with Dynamic UX, a lightweight package for dynamic interaction and reward collection in a single sandbox, and the second with Reward Auditor, a meta-reward mechanism that monitors reward distributions and distills diagnostic patterns into a shared rubric and scoring specification. We introduce UI-TAU Bench, a benchmark for active human-agent interaction through generated UI code, built on 10 real-world domain databases constructed from public data sources and based on Tau-Bench tool-use settings, with Lite (300 tasks) and Full (1,000 tasks) splits. GenUI-Harness achieves an average Pass@3 gain of 4.48 percentage points over smolagents on Lite. Training with GenUI-Harness improves a 4B backbone from 9.33% to 58.00% Pass@3, outperforming larger frontier models such as Claude Opus 5 (46.67%). GenUI-Harness also remains robust on ambiguous and non-ambiguous queries. In a reviewer survey comparing communication channels, generated UIs reduce average dialogue rounds from 3.4 to 1.2. These results show that data-aware generative interfaces can support effective task completion and reduce dialogue rounds in evaluated database-backed workflows.
cs.AI / 12 / 2610.11128
Balancing Reference Guidance and Free Generation in Trajectory Rollouts for Reasoning RL
Hanyu Wang, Nakul Agarwal, Hossein Nourkhiz Mahjoub, Ehsan Moradi Pari, Makoto Fukushima, Jinghui Chen, Vaishnav Tadiparthi
cs.AI
Abstract
A verified reference solution provides a correct trajectory for training a reasoning model. Alternatively, a prefix of the reference can guide the model in generating a trajectory of its own. How much reference guidance should we provide? We study this question through prefix continuation, where the model continues from a reference prefix and keeps the resulting trajectory if it passes verification, falling back to the reference otherwise. Since both procedures produce correct trajectories, we compare their distributions with the ideal distribution, the model's own distribution conditioned on successful verification. For one continuation, we derive the KL divergence in closed form, which, up to a bounded term, decreases with the product of the probability of generating a different correct trajectory and the reference surprisal, the negative log probability of the reference suffix given the prefix. Since a longer prefix tends to raise the former but lowers the latter, continuation success alone does not determine the preferred amount of guidance. From this analysis, we learn a prefix selector shared across training questions from continuation outcomes, without estimating success probabilities or additional generation. The resulting Adaptive Reference Guidance (ARG) constructs correct trajectories within a fixed generation budget, and we apply it to all-failure groups in Group Relative Policy Optimization (GRPO). Experiments on Qwen3-4B and Qwen3-8B across five mathematical reasoning benchmarks show that ARG achieves the highest aggregate pass@12 among the evaluated methods with competitive average sampled accuracy.
cs.AI / 13 / 2610.11129
GameCommBench: A Unified Benchmark and Type-Aware Evaluation for AI-Generated Game Commentary
Qirui Zheng, Zhengteng Lin, Yunyi Xiao, Junhao Li, Keyuan Cheng, Xingbo Wang, Yongyi Wang, Lingfeng Li, Yunlong Lu, Wenxin Li
cs.AI · cs.CL
Abstract
Game commentary is an open-ended generation task requiring multimodal perception, strategic reasoning, and contextual knowledge. Existing AI-Generated Game Commentary (AI-GGC) studies remain fragmented across games, modalities, and evaluation protocols, while overlap-based or holistic evaluators fail to capture the functional heterogeneity of commentary. We introduce \textsc{GameCommBench}, a unified benchmark spanning board games, sports, and esports, with commentary aligned to heterogeneous game contexts and annotated by commentary type. We further propose Type-Aware Commentary Evaluation (TACE), a structured framework for evaluating different types of commentary. We then validate TACE for reliability and human agreement, and use it to benchmark representative AI commentators. Results reveal non-uniform capability profiles, with live observation and strategic analysis emerging as major bottlenecks. Together, \textsc{GameCommBench} and TACE provide a diagnostic foundation for comparable and interpretable AI-GGC evaluation.
cs.AI / 14 / 2610.11155
Social Pain Disrupts Emotion-Action Brain-State Dynamics in Adolescents with Non-Suicidal Self-Injury
Ying Xu, Xiaojun Liang, Li Zhang, Yixuan Yuan, Gan Huang, Yongjie Zhou, Zhen Liang
cs.AI
Abstract
Non-suicidal self-injury (NSSI) is prevalent among adolescents with depression, but the rapid brain-state dynamics linking social distress to maladaptive behavior remain unclear. We combine an experimental pain paradigm, electroencephalography (EEG) microstate analysis, and interpretable deep sequence modeling to investigate NSSI-related neurodynamics in 106 adolescents with depression, including 67 with NSSI (DN+) and 39 without NSSI (DN-), during social pain, physical pain, and resting-state conditions. A model integrating disease-specific, domain-adversarial, consistency, and contrastive learning captures higher-order dependencies in microstate sequences. Social pain yields the strongest NSSI discrimination, with 68.55% accuracy, outperforming the best baseline by 8.94% points. Model interpretation and conventional microstate analyses reveal weakened bidirectional transitions between MS3 and MS5 in DN+ adolescents during social pain. Source reconstruction associates MS3 with emotional/interoceptive processing and MS5 with action preparation, suggesting disrupted emotion-action coupling. Time-resolved analyses show greater early-to-middle action-state recruitment and later emotion-state recruitment in DN+ adolescents. In DN- adolescents, MS5-to-MS3 dynamics mediate associations between social-evaluation sensitivity and affective outcomes, whereas this mediation is absent in DN+; conversely, MS3-to-MS5 transitions are associated with greater negative affect in DN+. Together, these findings identify disrupted emotion-action coupling as a key neurodynamic mechanism underlying altered social pain processing in adolescents with NSSI, providing a mechanistically interpretable neural signature for objective identification of NSSI.
cs.AI / 15 / 2610.11188
What to Admit and How to Present: Governing Persistent Memory in LLM Agents
Chang Liu, Deliang Ding
cs.AI
Abstract
Persistent memory can improve personalization in LLM agents but can also induce sycophancy and cross-domain leakage. We distinguish two governance decisions: admission, which determines what recalled information enters the working context, and presentation, which determines how admitted information is expressed. We implement two inference-time designs without retraining: factor-compiled admission (FC), which assesses whole memory entries, and permission-semantic admission (PS), which decomposes entries into typed units; both translate adjudicated attributes into eligibility decisions via deterministic policies. We evaluate on a four-backbone development suite and an external benchmark with four tasks of 300 samples each. Relative to verbatim injection, FC and PS reduce pooled judge-assessed failure rates on the external benchmark by 6.7 and 8.8 percentage points (p = 2.7e-7 and 4.1e-12), and development-set cross-domain leakage falls by up to 29.5 percentage points. A query-conditioned gating baseline shows no significant change in objective-fact failure or pooled failure. Under matched admission budgets, PS outperforms random and relevance-based selection on external objective-fact judgment after Holm correction. Holding presentation fixed, tightening admission cuts cross-domain failure by a further 17.5 percentage points (p = 1.6e-4); in contrast, no comparison between two renderings of identical adjudicated outputs survives multiple-comparison correction. Both designs increase personalization failures, and PS misses the preregistered improvement and personalization-preservation criteria. These results support evaluating admission and presentation separately: selection quality provides task-specific safety gains, while preserving beneficial memory use remains unresolved.
cs.AI / 16 / 2610.11213
AliO: Output Alignment Matters in Long-Term Time Series Forecasing
Kwangryeol Park, Jaeho Kim, Seulki Lee
cs.AI
Abstract
Long-term Time Series Forecasting (LTSF) tasks, which leverage the current data sequence as input to predict the future sequence, have become increasingly crucial in real-world applications such as weather forecasting and planning of electricity consumption. However, state-of-the-art LTSF models often fail to achieve prediction output alignment for the same timestamps across lagged input sequences. Instead, these models exhibit low output alignment, resulting in fluctuation in prediction outputs for the same timestamps, undermining the model's reliability. To address this, we propose AliO (Align Outputs), a novel approach designed to improve the output alignment of LTSF models by reducing the discrepancies between prediction outputs for the same timestamps in both the time and frequency domains. To measure output alignment, we introduce a new metric, TAM (Time Alignment Metric), which quantifies the alignment between prediction outputs, whereas existing metrics such as MSE only capture the distance between prediction outputs and ground truths. Experimental results show that AliO effectively improves the output alignment, i.e., up to 58.2% in TAM, while maintaining or enhancing the forecasting performance (up to 27.5%). This improved output alignment increases the reliability of the LTSF models, making them more applicable in real-world scenarios.
cs.AI / 17 / 2610.11231
Harness Compilation: Which Decisions Should a Small Vision-Language Model Keep?
Minhao Fan, Yinyi Liu, Jiayu Zhao, Zihan Teng, Song Chen, Weichen Liu
cs.AI · cs.CV
Abstract
Small vision-language models may be able to read external evidence yet struggle to obtain it. We introduce Harness Compilation (HC), an offline procedure that adapts the division of work between a frozen small VLM and its external harness. A large teacher uses student execution traces to revise reusable content and control, while a separate validation set selects the deployed harness. Deployment requires neither weight updates nor teacher calls. Across seven visual question-answering settings with students of at most 9B parameters, HC improves scores over bare students by 9.9-23.9 points, averaged over three independent builds per setting. Interventions on five runtime decision types (invocation, selection, argument generation, evidence integration and abstention) show why this allocation matters: requesting evidence and generating open queries can be costly, whereas bounded choices and reading supplied text can remain useful student work. Fact cards benefit all ten evaluated students, but decision policies transfer unevenly. Recompilation for a new student model helps when the transferred interface no longer fits the student. With 100 practice items, HC exceeds answer-only LoRA on three tasks. Larger training budgets can match or surpass a fixed harness, while combining the two improves SlideVQA beyond either alone. These findings support allocating work from measured student behavior rather than uniformly removing decisions.
cs.AI / 18 / 2610.11268
TaReD: Tool-Aware Recursive Decomposition for Long-Horizon Tasks
Wei-Xiang Mao, Zhi-Kai Chen, De-Chuan Zhan, Han-Jia Ye
cs.AI
Abstract
Agents combine reasoning with tools to interact with external systems and complete real-world tasks. Early agents typically interleave reasoning and actions along a single execution chain. On complex tasks, this chain becomes unreliable because growing histories obscure intermediate dependencies and allow early planning errors to propagate. Recursively decomposing a complex task into smaller subtasks offers a natural solution, yet effective decomposition must account for the system's capabilities so that each subtask can be executed by the available tools. In realistic systems, however, tool libraries can be too large to expose in full. Injecting every tool description consumes substantial context while making relevant tools harder to retrieve and useful task boundaries harder to identify. We propose tool-aware recursive decomposition, which organizes tools by functional relationships into a hierarchy of capabilities. During execution, the agent discovers tools on demand and uses the hierarchy to recursively decompose a complex task into a subtask tree whose levels are aligned with the capabilities required at each stage. Experiments on complex real-world tasks show that the proposed method improves end-to-end task success rate by up to 40 percentage points over the compared baselines. The implementation of TaReD is available on GitHub: https://github.com/WeiXiang-Mao/TaReD.
cs.AI / 19 / 2610.11289
Open-ended Scientific Discovery with Possibilistic Reasoning
Anita Yang, Siu Lun Chau, Tomoya Wakayama, Krikamol Muandet, Masaki Adachi
cs.AI
Abstract
Autonomous scientific discovery with LLMs requires generating and testing hypotheses adaptively as evidence accumulates while maintaining statistical validity. Existing anytime-valid methods can handle data-dependent hypotheses, but open-ended discovery poses a deeper challenge: the best discovered hypothesis may still be the best of a bad lot, with better explanations yet undiscovered, while even background knowledge such as physical laws may require revision in light of new findings. In response, we formalize the problem as Abductive Autonomous Scientific Discovery (AASD) using possibility theory. We introduce abductive utility, a computable measure of discovery progress, and possibility frontier search, the first algorithm for AASD, which maintains anytime validity and achieves $\varepsilon$-optimal abductive utility asymptotically under suitable conditions. Experiments on synthetic and real-world scientific-discovery tasks show strong performance.
cs.AI / 20 / 2610.11294
ORDO: Operation-level Round-aware Dynamic Ordering for MIP Presolve
Zehuan Chen, Chunhe Song
cs.AI
Abstract
Presolve strongly affects mixed-integer programming (MIP) performance, yet learning-based methods only optimize parameter configurations and cannot express the non-commutative temporal dependencies among actions, whose default order is nearly unique on most domains, yet functionally necessary: artificially shuffling the order of the same sequence inflates the tail of the solve-time distribution by up to several-fold. We recast presolve planning as autoregressive sequence generation over a unified atomic action space, moving the decision object to action sequences; we call this framework ORDO---Operation-level Round-aware Dynamic Ordering for MIP Presolve. Its payoff is cross-domain generalization: on multiple unseen domains it attains end-to-end zero-shot speedup---to our knowledge the first for presolve action sequences---varying by domain and not explained by corpus richness, the strongest domain reaching the largest speedup once racing is added. Deployment uses sequence racing, in which candidate sequences run concurrently and the winner is kept, enabled by an execution-and-observation facility, added by modifying the SCIP source, that injects sequences along the native path and records which actions actually execute and in which round.
cs.AI / 21 / 2610.11299
DuplexAgent-RSI: Recursive Harness Improvement for Full-Duplex Voice Agent Collaboration
Yingda Shen, Yuxiang Wang, Kunyu Feng, Qinke Ni, Jiaqi Li, Minghao Hsu, Junan Zhang, Dekun Chen, Yutong Bian, Zhizheng Wu
cs.AI · cs.SD
Abstract
Voice agents are converging on a collaboration pattern: a full-duplex interaction model stays on the live channel as the entry to the conversation, while search, reasoning, and coding are handled through asynchronous delegation. A duplex model supports continuous listening and speaking, but complex reasoning and tool use may exceed its capabilities. A coding agent can plan and execute extended tasks, but its sequential interface is a poor fit for live conversation. Combining them requires a harness that coordinates task acceptance, progress, cancellation, replacement, and result delivery while keeping the conversation responsive. Existing harnesses often rely on coupled heuristics, making them difficult to improve systematically from evidence. We present DuplexAgent, a full-duplex collaboration system whose harness expresses this workflow as six editable modules, and Duplex-Harness-RSI, a closed loop that revises them from interaction traces. A simulator automatically generates timed test conversations, runs the system, and produces failure traces that identify the collaboration modules requiring repair. Reasoning LLMs and coding agents in the delegation pool also serve the improvement loop: the Exam Planner selects the next tests from observed weaknesses and the repair archive, and the Harness Editor proposes targeted module changes. The capabilities that serve the user thus also improve the system's coordination. Experiments on intelligence, agentic, and duplex benchmarks show that DuplexAgent combines continuous interaction with difficult reasoning and complex task execution, achieving stronger spoken-knowledge and executable-tool scores than the compared delegated systems while maintaining strong interruption response. A harness ablation further shows that this modular, verifiable loop outperforms the initial harness and repeated editing that lacks its diagnosis and repair archive.
cs.AI / 22 / 2610.11318
Finsler Flow Matching: Dynamics-Aware Geodesic Interpolation for Single-Snapshot Trajectory Inference
Niklas Canova, Jonas Simon Fleck
cs.AI
Abstract
Single-cell snapshot data can resolve a continuum of cellular states but do not uniquely determine the dynamics governing transitions between them. However, additional dynamical information can often be encoded in a cell-cell Markov transition kernel. Existing generative approaches for single cell trajectory inference either infer transport only from population marginals, impose a symmetric geometry on the state space, or incorporate directionality through a single velocity vector at each observed state. We introduce Finsler Flow Matching (FFM), a framework for learning continuous stochastic dynamics from discrete Markov transition graphs. We use the first and second local moments to construct a Finsler structure motivated by the Freidlin--Wentzell action, where the second moment determines anisotropic accessibility and the first moment introduces a preferred direction of motion. We learn neural approximations of the resulting directed geodesics, use their Finsler cost to construct source-target couplings, and define geometry-aware stochastic conditional paths that can be distilled into a continuous generative process through simulation-free score and flow matching. Across synthetic and single-cell trajectory inference benchmarks, FFM improves recovery of withheld intermediate populations, particularly when the transition dynamics are strongly directional or anisotropic. Our results provide a principled route from discrete transition probabilities to continuous generative dynamics while retaining both directional and diffusive structure.
cs.AI / 23 / 2610.11328
Mine Odyssey: Benchmarking Spatial Agentic Intelligence in the Wild
Yuxuan Cao, Junlong Li, Hao Li, Junxian He
cs.AI
Abstract
Advances in foundation models are driving efforts to introduce agents to assist people in the physical world. Such agents require agentic spatial intelligence: exploring unfamiliar environments, updating spatial understanding through interaction, and adapting actions based on feedback to sustain progress toward a sequence of goals. Existing benchmarks cover only a limited range of spatial layouts, scales, and traversal requirements. We introduce Mine Odyssey, a benchmark for evaluating agentic spatial intelligence using Minecraft reconstructions of real-world locations. It comprises 180 tasks covering 30 such locations across 20 countries and regions on five continents, including 20 outdoor and 10 indoor settings. These settings span diverse spatial scales, layouts, terrains, and connectivity patterns, from Midtown Manhattan and rural Entrup to Santa Lucía Hill and Buckingham Palace. We select meaningful waypoints, such as landmarks, buildings, and rooms, and manually verify their accessibility. Each task provides a natural-language instruction specifying which waypoints to visit and in what order. Completing these tasks requires agents to find accessible routes and entrances, open doors, and move between levels using stairs and ladders, while monitoring their progress and recovering from navigation errors. Across eight evaluated state-of-the-art models, GPT-6 Astra achieves the highest success rate of 85.6%. However, the second-best model, Claude Opus 5.5, completes 73.9% of tasks, while the strongest evaluated open-weight model, DeepSeek-V4.1-Flash, reaches 23.9%, highlighting substantial room for improvement in the agentic spatial intelligence of current models. Comprehensive analyses and ablation studies on Mine Odyssey reveal current models' limitations and provide insights for advancing agentic spatial intelligence.
cs.AI / 24 / 2610.11333
TokenBank: Financial Infrastructure for AI Services
Cary Chang, Jialin Zhou
cs.AI · cs.CE
Abstract
AI services incur inference costs during execution, while revenue may arrive later. Changing API prices, limited upfront capital, and service failures can limit operators' ability to sustain or expand their services. Beyond reducing per-request costs, operators need to plan future spending, fund execution before revenue arrives, and obtain compensation for specified losses. This requires clear agreements across services with different pricing and execution conditions. These agreements must distinguish rights to consume services from rights to receive payments, define obligations under uncertain costs and income, and specify which failures qualify for compensation and how much can be paid. We present TokenBank, a financial infrastructure that represents these commitments through structured contracts. It supports service-consumption rights, agreements that settle API-price differences in cash (forwards), financing through limited rights to future service revenue, and protection claims for specified service failures. Contracts specify participants, covered services, validity, ownership, fulfillment conditions, and settlement rules. Evaluation combines replay of 899,441 API requests, real model-driven agent execution, and contract API tests. In a zero-discount rising-price resampling scenario, forwards reduce mean expenditure by USD 304.88 but increase its standard deviation from USD 1,152.45 to USD 1,190.82. A controlled replication with five portfolios per capital condition finds mean contribution differences between financing and self-funding of +1.0635, -0.1406, and -0.2962 experimental USD under low, baseline, and ample capital, respectively. The evaluation distinguishes contract correctness from economic effectiveness under declared economic and failure assumptions; supplier invoices and commercial revenue are unavailable.
cs.AI / 25 / 2610.11334
ReCast: Attribution-Oriented Step Representation Learning for LLM-Based Agent Systems
Weilin Jin, Mingyu Wang, Taiyu Zhu, Ziqi Zhou, Wenbo Li, Haoyang Huang, Nan Duan, Yifan Wu, Ying Li, Zhonghai Wu
cs.AI
Abstract
In LLM-based agent systems, failures can originate from early steps whose effects propagate through subsequent interactions, making their origins difficult to identify. To trace such failures back to their origin, failure attribution has been formulated as the task of identifying the earliest step responsible for the failure. Recent methods leverage LLM internal signals for failure attribution, typically using hidden states as step representations. We therefore conduct an empirical study to evaluate how effectively these representations distinguish root-cause steps from other steps and find limited separation. Motivated by this observation, we propose ReCast, a step representation learning method that transforms hidden states from a frozen LLM into attribution-oriented step representations. ReCast first selects attribution-relevant layers, then constructs complementary pattern and deviation features, and finally learns contextualized step representations through an encoder trained with contrastive and ranking objectives. We also introduce ReCast-2K, a training dataset for failure attribution. ReCast achieves the best Hit@1 across four benchmarks, surpassing the strongest baseline by 5.65 and 9.19 pp on Who&When Algorithm and Handcrafted, respectively. Code is available at https://anonymous.4open.science/r/ReCast-5FB6 .
cs.AI / 26 / 2610.11344
EvoSim: Learning to Model, Modeling to Learn
Yun-Wei Song, Jinkai Tao, Jun-Dong Zhang, Rui Zhang, Yi-Min Wu, Qiang Zhang
cs.AI · cs.CE
Abstract
Physics-based models connect scientific explanation with quantitative prediction. Constructing them requires selecting physical processes, defining states and governing equations, specifying couplings, and identifying parameters from experiments. Existing AI systems remain limited in making these model structure decisions autonomously. We introduce EvoSim, a self-evolving AI scientist for physical modeling. It uses experimental discrepancies to drive mechanism and equation revisions and held-out experimental data to test physical plausibility. Exploration traces make updates to knowledge, skills, and multi-agent orchestration. This co-evolution improves physics-based models and EvoSim's ability to select mechanisms, diagnose failures, and coordinate research. We evaluate EvoSim on two industrial battery modeling tasks. It predicts lithium-metal-plating onset from 25 to 45 degrees Celsius and 2 C to 6 C with a mean absolute error of 1.79% in state of charge. Dynamic voltage prediction under vehicle driving conditions achieves a root mean square error of 7.62 mV, surpassing the reported accuracy of models developed by human experts. Self-evolution reduces model and physics errors by approximately 36% relative to baseline, demonstrating improved scientific modeling capability. EvoSim turns experimental observations into validated models and cumulative research expertise.
cs.AI / 27 / 2610.11345
SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning
Wei Yang, Shawn Li, Yuehan Qin, Yawei Wang, Mingxi Wang, Shixuan Li, Tiankai Yang, Jiate Li, Jesse Thomason, Xuezhe Ma, Yue Zhao
cs.AI
Abstract
Self-evolving LLM agents promise to improve autonomously through continual interaction and learning, reducing their dependence on manually curated supervision. Realizing this promise requires not only updating the agent, but also evolving its training experience as its capabilities change. However, most existing pipelines rely on static datasets or separately updated synthesis models, causing previously useful tasks to become trivial while overly difficult tasks remain uninformative. This growing mismatch between agent capability and training experience limits sustained self-improvement. To address this problem, we propose SynCo, an agentic data synthesis co-training framework for self-evolving LLMs based on multi-agent reinforcement learning. SynCo jointly optimizes two independently parameterized agents: a Synthesizer that constructs training tasks from the Reasoner's evolving capability state, and a Reasoner that learns from the resulting experience. Each synthesized task induces multiple Reasoner rollouts whose outcomes provide complementary rewards to both agents. Correctness feedback improves the Reasoner, while task quality, answer reliability, and outcome-grounded teachability guide the Synthesizer. Their updates are fed back into subsequent synthesis rounds, allowing the task-solving policy and its training distribution to evolve together. Extensive experiments across eight mathematical reasoning benchmarks demonstrate that SynCo substantially outperforms a broad range of existing synthetic-data methods and controlled baselines, achieving the strongest overall performance while deriving most of its gains from previously unsolved problems.
cs.AI / 28 / 2610.11352
RL-ARC: Calibrating Large Reasoning Models via Reasoning-guided Uncertainty
Gukhyeon Lee, SangKeun Lee
cs.AI · cs.CL · cs.LG
Abstract
Language models (LMs) are commonly trained with Reinforcement Learning with Verifiable Rewards (RLVR) to enhance their reasoning capabilities. However, since RLVR does not explicitly account for calibration during training, it can lead to severe calibration degradation, including overconfidence. Recent calibration-aware training methods for LMs, which incorporate objectives for uncertainty estimation into training, improve calibration but still exhibit overconfidence under distribution shift, while sacrificing reasoning performance. To this end, we propose RL-ARC, a calibration-aware training framework that jointly leverages reasoning confidence and answer confidence. Specifically, RL-ARC leverages reasoning confidence as an auxiliary signal for calibrating answer confidence, applying it as reasoning-guided regularization for correct cases and as an overconfidence penalty for incorrect cases. Comprehensive results across ID and OOD settings show that, beyond improving calibration, RL-ARC enables reasoning models to adaptively estimate confidence based on the given question without substantially sacrificing reasoning performance, thereby highlighting the importance of reasoning confidence for training reliable reasoning models.
cs.AI / 29 / 2610.11355
Writing for the Reviewer: Defensive Writing in GPT Models
Junchi Liao
cs.AI
Abstract
Researchers increasingly use ChatGPT to revise their papers, and recent GPT versions often narrow or even retract the authors' claims. We call such changes defensive writing when the given material does not support them, and we test two explanations: the model corrects the authors' overclaiming, or it writes for an anticipated reviewer. We ask GPT versions and models from other developers to rewrite paragraphs from papers written before ChatGPT, or to write from an evidence sheet that lists a paper's method and results. Defensive writing grows with GPT version. GPT-6-astra retracts the authors' claims outright, and when it writes from the evidence sheet, it still adds the most ungrounded qualifications. The results favor the anticipated-review explanation, and correcting overclaiming explains only a small part. When the models are only asked to polish, defense stays near the level of the originals; mentioning review raises it, and one round of self-review raises it further. At the same time, fewer than one in ten of the claims GPT-6-astra retracts are overstated. AI reviewers score defensive rewrites higher, while human readers find them harder to read and the authors less certain. Combining AI writing with AI review may amplify this style.
cs.AI / 30 / 2610.11358
RaReCache: Bridging the Gap in Cross-Model KV Cache Reuse via Rank disagreement-based Selective Recomputation
Sreetama Sarkar, Saptarshi Mitra, Sitao Huang, Souvik Kundu, Peter A. Beerel
cs.AI · cs.LG
Abstract
Cross-model KV-cache reuse remains a key challenge in modern LLM serving. Coding agents and multi-model systems increasingly route a shared context across models: a user may switch models mid-session, or a cascade may escalate a difficult query. Because KV caches contain model-specific representations, each switch typically forces the receiving model to prefill the entire context from scratch. Recent work shows that closed-form linear maps can translate KV caches between models in the same family, but transfer accuracy degrades as the model-size gap widens. In this paper, we establish that these transfer failures are concentrated in a small subset of information-dense tokens. To bridge this gap, we introduce RaReCache, a framework that enables a large target model to decode accurately from a cache prefilled by a much smaller source via selective recomputation. RaReCache identifies these critical positions using a novel rank disagreement metric, scoring each token by the energy of its mapped KV in output directions weakly supported by the calibration data. Across two model families and five benchmarks, on a 23x parameter gap (Qwen3-0.6B to 14B) recomputing just 30% of positions retains 95-99% of the target accuracy, whereas on a 8.8x gap (Llama3-8B to 70B), recomputing 40% retains 96.5% of the target accuracy. RaReCache largely removes sensitivity to source-model size, and achieves up to a 3.04x prefill speedup. For online serving, it handles 1.8x the request throughput of target prefill on a single GPU, and at the target's saturation load, reduces median and 99th-percentile time-to-first-token (TTFT) by 5.0x and 6.4x respectively, with a 30% recompute budget. RaReCache establishes an efficient serving paradigm where small models prefill on behalf of massive targets, enabling large models to recompute only critical tokens, drastically reducing prefill latency.
cs.AI / 31 / 2610.11370
RIT-RAG: Navigating Document Corpora with Retrieval-Induced Trees
Meghanadh Pulivarthi, Swaraj Kumar Biswal, Kushagra Bhushan, Yatin Nandwani, Sachindra Joshi, Dinesh Raghu
cs.AI · cs.IR
Abstract
Retrieval-augmented generation (RAG) grounds language models in external corpora. Agentic RAG enables iterative search, yet exposes the model to isolated chunks without document structure, making it difficult to distinguish relevant evidence from chunks that merely resemble the query. Structure-aware methods such as PageIndex navigate document structure but cannot scale to the structures of large corpora, which do not fit in the LLM context. Hence, they first commit to a single document using a document retriever and cannot recover from a wrong choice. We propose RIT-RAG (Retrieval-Induced Tree RAG), which combines content retrieval with structural navigation. Offline, RIT-RAG builds a tree for each document from its table of contents or sitemap. At query time, it retrieves a broad set of chunks and uses their positions to induce manageable sub-trees, potentially across multiple documents. An LLM agent navigates these sub-trees, selectively reads promising nodes, and reformulates queries when needed. Thus, retrieval proposes where to look, while the agent decides what to read. Across financial, scientific, and customer-support benchmarks, RIT-RAG achieves the highest answer accuracy among vanilla, graph-based, and agentic baselines. On EntQABench, our new benchmark of 2.84 million technical-documentation webpages, it improves accuracy by 6.8 to 11.4 points over the strongest baseline across three LLMs.
cs.AI / 32 / 2610.11384
Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation
Hangxi Guo, Fengyuan Liu, Yue Wang, Yuhua Qi, Haoyi Xiong, Fei Sun, Mengnan Du
cs.AI
Abstract
Reinforcement learning is commonly used to train language agents in interactive environments, but cannot be directly applied when rewards are unavailable. Recent methods use environmental feedback as privileged context for hindsight self-distillation, but our analysis suggests that simply conditioning the teacher on feedback is insufficient, motivating us to rethink how environmental feedback is used in agentic self-distillation. Given that environmental feedback contains rich supervision for modeling how the environment responds to agent actions, we introduce \textit{agentic SElf-distilLation with environmental Feedback modeling} (SELF), a framework that jointly optimizes environmental feedback modeling and hindsight self-distillation. SELF learns to predict environmental responses while distilling guidance from a feedback-conditioned self-teacher into the policy. Our analysis reveals a mutually reinforcing mechanism: environmental feedback modeling strengthens hindsight supervision and policy learning, while self-distillation enhances the model's ability to model environmental feedback. With Qwen3-8B, SELF outperforms SDPO and GRPO by 6.4 and 4.1 percentage points in $τ$-bench success rate, and by 10.71 and 3.57 percentage points in AppWorld task goal completion, respectively. These results show that SELF uses environmental feedback more effectively within agentic self-distillation, improving agent capabilities.
cs.AI / 33 / 2610.11392
TypedBench: A Benchmark for Calibration, Framing Sensitivity, and Cost in System One Decision Models
Rahul Sharma, Andrew B. Ducan, Gaétan Marceau Caron, Sebastian J. Vollmer
cs.AI · stat.AP
Abstract
System One models output calibrated probabilities over typed answers such as categorical choices, ordinal levels, or binary outcomes, via a non-generative interface. Software can act on these probabilities through thresholds, cost-weighted choices, and escalation rules. Consequently, if these probabilities are miscalibrated or wording-sensitive, the software ma take unintended actions leaving human operators with no textual rationale to inspect. Current evaluations largely report accuracy and calibration on public classification datasets without a clear reference. We present TypedBench, a benchmark for typed decision models like Jev built from seven policy-labelled generators and nine evaluation suites. We report accuracy as median and range across paraphrases, and calibration error relative to the finite-sample noise floor of a matched, perfectly calibrated predictor. We assess probability quality through selective prediction, ordinal proper scoring rules, and realised cost under asymmetric cost matrices. We evaluate a hosted model, an open encoder, and a family of open decoders spanning 0.8B-9B parameters on identical items. The hosted model follows the stated policy but is wording-sensitive and systematically underconfident; under asymmetric costs, using its probabilities can be worse than taking its top answer. Decoders route exactly and are slow as options or questions are added. The decoder is least accurate on policy questions and degrades with more options. Overall, typed decision models must be evaluated jointly on policy adherence, wording robustness, probability quality, and induced decision outcomes.
cs.AI / 34 / 2610.11399
Estimating great expectations under autoregressive language models with potentials
Francesco I. Re, Shubhangi Ghosh, Tim Vieira, Ryan Cotterell
cs.AI · cs.LG · stat.ME
Abstract
Many applications of language models hinge not on individual samples but on the expectation of a test functional under the model. Estimating such expectations reliably can be computationally expensive. In this paper, we show how to make estimation more efficient by exploiting the next-token conditional probabilities which are available as a by-product of sampling. We do so through potentials: real-valued functions on prefixes that decompose the test functional additively. We construct an estimator whose variance depends on the chosen potential, and derive conditions under which a potential reduces this variance. We then develop practical potentials for several estimands and applications, and demonstrate substantial variance reductions across several estimands at comparable computational cost.
cs.AI / 35 / 2610.11410
Cognition-Oriented Emotion Tracing from Causes to Consequences in Real-World Social Scenes
Hao Li, Jinye Zhang, Bobo Li, Mong-Li Lee, Wynne Hsu, Zheng Wang, Hao Fei, Min Zhang
cs.AI
Abstract
Affective computing has progressed from categorical emotion recognition to open-ended affective analysis with large multimodal models. Yet affective science describes emotion as an unfolding process shaped by appraisal, regulation, and social interpretation, which remains underexplored computationally. We propose TRACE, a cognition-oriented framework that formalizes an affective episode through three interrelated stages: Condition, Affect, and Effect, integrating observable cues with cognitive factors such as internal stance and regulation of emotional display. Based on this formulation, TRACE-Bench evaluates multimodal models in real-world social scenes through five tasks spanning grounded affect recognition, regulation decoding, cause reasoning, effect reasoning, and full-chain reconstruction, with 3,746 structured question-answer pairs over 646 videos. A matched human-model comparison reveals a substantial performance gap, while affect-specialized models also generally lag behind general-purpose MLLMs. Model outputs show recurring failures, including treating displayed behavior as genuine feeling and fabricating unsupported events during long-chain generation. We further propose TRACER, a cognition-grounded structured reasoning method that couples each inference with explicit premises from factual observations, cognitive appraisals, and established upstream conclusions, forming a traceable graph of intermediate and target conclusions. TRACER outperforms all evaluated model baselines on each of the five tasks. Project page: https://cogaffc.github.io/TRACE
cs.AI / 36 / 2610.11424
Why machines will still not rule the world
Jobst Landgrebe, Barry Smith
cs.AI
Abstract
In our book Why machines will never rule the world [13, 14] we argue that arti- ficial general intelligence is mathematically impossible. This is because the human beings and the processes which exhibit intelligence are complex systems whose be- haviour cannot be captured by the kinds of models that we can generate with or without computers. Proponents of contemporary machine intelligence respond with two lines of argument: a theoretical one, grounded in the universal approximation theorems for neural networks and the Church-Turing-Deutsch principle; and an em- pirical one, grounded in rapidly rising scores on standardized benchmarks. In this communication we examine and reject both responses. First, we show serious issues in the physicalist counter-argument based on the Church-Turing-Deutsch principle. Second, we review recent evidence to the effect that prominent benchmarks are compromised by training-data contamination, flawed test construction, and strate- gic optimization. Our central argument remains: That models required to perform cognitive behaviour in open-ended, thermodynamically complex and non-ergodic environments are not and will not become achievable.
cs.AI / 37 / 2610.11450
Tracing the Thoughts of a Coding Agent Playing ARC-AGI-3: Lessons for Continual Learning
Chen Wu, Josh Passenger, Yin Song
cs.AI
Abstract
We study how a coding agent learns across a sequence of abstract reasoning tasks. The agent runs on a frozen foundation model inside a fixed harness and acts by writing and running Python and shell scripts. It retains no state across turns other than its written artifacts, so every thought it forms, carries, corrects or abandons leaves a trace, where a thought is any belief, rule or plan committed to a file. We let the agent play ARC-AGI-3, a set of interactive reasoning games that provide no instructions. Each game is a sequence of levels, and a strategy that clears one level can fail on the next, so every new level is in effect a new task. The agent records what it learns as Python scripts and text notes, while the harness keeps a complete log of every action and observation. Our contribution is a measurement protocol that traces each thought through these files, from the task where it forms to the task where it is corrected or abandoned, applied to seven evaluation runs with three backbones from two model families. Scripts written for one task are almost never called again in a later task (33 of 630 references cross a task boundary), because most scripts embed the state of the current level. Instead, the agent rewrites its knowledge into new scripts, keeping the general rules and dropping the level-specific details, and abandons 74% of the scripts it wrote before a boundary. The notes, which only the model reads, are never revised: the agent appends without removing earlier claims, and the contradictions that accumulate are settled against the log. Because the log preserves everything, the agent forgets selectively, not catastrophically. The most costly error is a hard-coded value carried into a task where it no longer holds. These findings come from the files the agent wrote, without access to the model, and constitute a white-box analysis of how a coding agent continually learns.
cs.AI / 38 / 2610.11456
SpikeSSL: A Universal Spike Inference Framework with Dynamics-Informed State-Space Layers
Chenghao Yue, Siming Xing, Shuran Liu, Angran Li, Yuanlong Zhang
cs.AI
Abstract
Two-photon calcium imaging is a standard tool for recording large neural populations in vivo, yet inferring spikes accurately across the growing diversity of calcium indicators remains an open problem. Existing supervised methods achieve reasonable in-domain accuracy but generalize poorly to unseen indicators, because different indicators induce distinct fluorescence kinetics and signal statistics while existing architectures remain relatively simple generic temporal regressors without dynamics-matched inductive bias. We propose SpikeSSL, a universal spike inference framework whose temporal backbone is a bank of bidirectional IIR state-space layers broadly motivated by calcium dynamics. A multi-modal conditioning encoder maps indicator identity, sampling rate, and trace-level signal statistics into a global conditioning vector that modulates the backbone via Adaptive Layer Normalization, while a heteroscedastic variance head provides calibrated per-frame uncertainty. On a benchmark with five fixed evaluation splits built from 33 public ground-truth datasets, SpikeSSL achieves state-of-the-art performance in both in-domain and zero-shot leave-one-indicator-out settings. We also develop a biophysical simulation pipeline capable of generating paired fluorescence-spike traces with systematically varied kinetic parameters, spike statistics, response nonlinearities, baseline drift, and noise. Using this pipeline, we synthesize approximately 11,000 simulated traces. Augmenting training with these data effectively closes the cross-indicator domain gap and improves zero-shot generalization. Code is publicly available at https://github.com/detimage123/SpikeSSL.
cs.AI / 39 / 2610.11464
Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents
Xing Zhang, Guanghui Wang, Yanwei Cui, Ziyuan Li, Wei Qiu, Bing Zhu, Peiyang He
cs.AI · cs.CL · cs.LG
Abstract
We changed the agent: did it actually get better? Every self-improving agent loop answers this hundreds of times, and every answer comes from a verifier. On open-ended tasks none exists, so the loop is handed a hand-written rubric or a bare LLM judge grading output from a model like itself, inviting reward hacking and shared blind spots. We make the verifier the evolving object: an inspectable expression over small, mostly deterministic drawback detectors, synthesized from clustered failures, gated at birth, and selected for agreement with a ten-item anchored reference set plus consensus over unlabeled outputs, never for the agent's score. On MBPP+ it gains +0.21 held-out agreement over the hand-authored seed composition, on every seed, and ends ahead of the bare LLM judge it contains. One finding should change how co-evolved verifiers are validated: removing the anchor guards collapses the verifier into a vacuous always-pass grader, yet that collapsed verifier trains skills just as well. Downstream task score cannot certify a self-evolved verifier. Score does answer sufficiency, and there an evolved verifier can substitute: Double Ratchet, pairing the verifier with a lifecycle-managed skill loop, retains 88-110% of the lift that ground truth or a rubric buys the same loop, across code generation, enterprise text-to-SQL, and reference-free report generation. When evolved skills gamed the report rubric, an outer judge caught it and one added detector repaired it; the judge itself was wrong until given the task contract.
cs.AI / 40 / 2610.11472
From Chain-of-Thought to Loops: Non-Autoregressive Latent Reasoning via Looped Transformers
Gerard Grau García, Arnau Padrés Masdemont, Niccolò Grillo, Jordi Ros-Giralt, Arash Behboodi, Victor Conchello Vendrell
cs.AI
Abstract
Chain-of-thought (CoT) reasoning often improves language-model performance by giving models additional computation before answering. However, explicit CoT expresses this computation as a sequence of autoregressively generated tokens. Latent reasoning replaces these tokens with compact continuous states, but most autoregressive latent-reasoning methods retain a left-to-right dependency among latent vectors. We introduce LLoCoT: a looped latent-reasoning framework that replaces left-to-right latent generation with iterative refinement of a compact latent workspace. A shared transformer is reapplied for a small number of refinement iterations, jointly updating the latent slots based on the prompt and the evolving workspace state. Using the refined state, a probabilistic head predicts a distribution from which latent tokens are sampled in parallel and used to condition an autoregressive decoder for answer generation. Training uses continuous representations derived from explicit CoT together with a final-answer prediction loss and likelihood-based supervision of the latent states. Across HumanEval and MBPP, LLoCoT achieves the highest mean among the evaluated methods, performing on par in accuracy with Reasoning SFT, our explicit-CoT baseline, while outperforming the base model, answer-only SFT and NF-CoT. Relative to Reasoning SFT, LLoCoT reduces time to the first answer token by approximately $36\times$ and reasoning-phase latency by approximately $42\times$, while increasing end-to-end throughput by $9.2\%$. This design replaces serial thought generation with parallel latent-slot refinement while retaining probabilistic latent modeling and autoregressive answer decoding.
cs.AI / 41 / 2610.11504
The Operator Mismatch Problem: Deploying BEV Perception with Portable GPU Compute
Rohit Verma, Anand V Bodas
cs.AI
Abstract
Modern autonomous driving systems rely on bird's-eye-view (BEV) perception models that fuse camera and LiDAR inputs to detect objects in 3D space. These models are accurate, but they cannot be deployed through standard inference runtimes. The reason is an operator mismatch between dense convolutions (which runtimes handle well), sparse 3D convolutions (which runtimes cannot represent), and geometric scatter operations (which runtimes have no vocabulary for). Today, every sparse convolution library is CUDA-only and PyTorch-coupled, locking BEV deployment to a single vendor's hardware and a single execution framework. We present BEVPIPE, a framework for deploying multimodal BEV perception pipelines using portable GPU compute APIs and integrating them with production inference runtimes. BEVPIPE partitions the model into runtime-managed dense subgraphs and three external operator extensions (voxelizer, sparse encoder, BEV projector), connected through a shared GPU memory space. BEVPIPE achieves a 19.5x end-to-end speedup over conventional deployments while retaining 98.5% of reference mAP. We also showcase that BEVPIPE is portable across different GPU backends.
cs.AI / 42 / 2610.11529
ReTeach: Building a Self-Teacher through Multi-Round Reflection and Retry
Yafeng Tang, Hao Li, Hongsheng Yu, Qiang Fu
cs.AI
Abstract
Self-distillation can improve reasoning without a separately trained, more capable teacher, but its effectiveness depends on how the self-teacher gains an advantage over the student. Conditioning the teacher on reference answers or solutions can provide such an advantage, but this information may be unavailable. Reflection offers a way to derive explicit error diagnoses and revision guidance from self-generated attempts, yet existing reflection-based methods often combine it with reference information, rich task feedback, or persistent memory. We introduce ReTeach, a Reflective self-distillation framework that constructs its self-Teacher through multi-round reflection and retry using only self-generated attempts and outcome-level verification. Starting from an unsuccessful student rollout, the teacher alternates explicit reflection with renewed attempts until success or the retry budget is exhausted, without reference answers or solutions, external diagnostic feedback, or cross-example memory. Each failed retry informs subsequent reflection, while successful correction provides outcome-level evidence for the potential utility of the resulting teacher context. An outcome-aware selection and weighting strategy distinguishes initially correct, reflection-corrected, and unresolved examples, assigning separate weights to their category-normalized distillation losses. Through on-policy distillation, the student matches the teacher's context-conditioned token-level predictive distributions at prefixes of its own rollouts, transferring the benefits of iterative correction while retaining single-pass inference. Across six benchmarks spanning mathematical reasoning, science question answering, and tool use, ReTeach improves average accuracy over GRPO by 1.39 percentage points.
cs.AI / 43 / 2610.11561
Workerville: Towards an Organizational Behavior Account of Agent Safety
Hanjun Luo, Junting Mao, Yuhan Lu, Haobo Zhang, Zhimu Huang, Yankai Chen, Hanan Salam, Xue Liu
cs.AI
Abstract
LLM-based agents now interact with their environments continuously, shaped by such organizational channels as user instructions, peer messages, and long-term memory. Existing safety research has examined these influences, but largely as separate agent components. How such factors jointly shape an agent's safety behavior from a unified perspective remains unmeasured. To bridge this gap, we advocate organizational behavior (OB) as a framework for studying the safety of advanced agents, reorganizing the objects of study, theoretical foundations, and experimental design around the relational structure in which agents operate. We present the first systematic formalization of counterproductive work behavior (CWB), a canonical safety-relevant subfield of OB, as Agentic Counterproductive Behavior (ACB). ACB specifies three organizational antecedents (vertical supervisor relations, horizontal peer norms, and internal cognitive structures) and maps them onto three counterproductive outcome dimensions (unauthorized disclosure, destructive operations, and production deviation). To operationalize ACB, we introduce Workerville, a controlled benchmark that manipulates organizational conditions over shared tasks, applying 16 organizational configurations to 210 tasks to yield 3,360 challenges, evaluated by human-validated agentic judges. Benchmarking 6 frontier LLMs, we find that (I) negative organizational antecedents exhibit non-monotonic amplification when combined, with the unauthorized-disclosure rate rising from 16.5% under no negative antecedent to 60.1% under two and falling back to 50.3% under three; (II) agents reproduce typical behavioral patterns predicted by human CWB research; (III) these results establish OB as a systematic framework for agent safety research, pointing toward a new research agenda.
cs.AI / 44 / 2610.11570
Scaling to Tens of Thousands of Test-Time Iterations with Loop-Native Attention Residuals
Pengxiang Li, Dilxat Muhtar, Di He, Guinan Su, Lu Yin, Shiwei Liu
cs.AI
Abstract
In this paper, we argue that looped Transformers need their own residual connections to prevent performance degradation as the number of iterations grows. We observe that increasing loop iterations can reduce reasoning accuracy: noisy state updates overwrite correct intermediate deductions and even undo completed solutions. This leaves subsequent iterations to recover lost information from an already degraded representation: once an error arises in an earlier loop, often as a result of long-range propagation through the recurrence, later loops find it difficult to correct. In this paper, we introduce InfiLoop, a loop-native residual connection that learns which past computations to retain and how much to accept from each new update. InfiLoop combines content-based weighting with learned temporal decay to maintain a running summary of recurrent states. An exact streaming recurrence keeps its persistent aggregation memory constant as the loop count grows. The resulting adaptive update suppresses unreliable proposals and preserves useful intermediate states. Across extensive reasoning tasks, a 7M-parameter InfiLoop model outperforms existing recursive architectures, reaching 97.9% exact accuracy on Sudoku-Extreme, and 13.6% pass@2 on ARC-AGI-2. Notably, on Sudoku-Extreme, InfiLoop continues to improve with test-time looping beyond 20,000 effective steps, showing that added depth translates directly into stronger reasoning. Our code is available at https://github.com/pixeli99/InfiLoop.
cs.AI / 45 / 2610.11600
Error-Propagation Modeling for Failure Attribution in LLM-Based Multi-Agent Systems
Jiaqi Liao, Yuanzhao Zhai, Huanxi Liu, Xu Zhang, Zheming Zhuang, Dawei Feng, Bo Ding, Huaimin Wang
cs.AI
Abstract
LLM-based multi-agent systems (MASs) are increasingly used to solve complex tasks through coordinated reasoning, tool use, and interaction with external resources. However, attributing failures in such systems remains challenging because the observed outcome often does not directly reveal the error responsible for the failed execution. In this work, the attribution target is the decisive error, defined as the agent--step pair whose correction would recover the failed execution. Existing approaches largely identify suspicious steps without explicitly modeling how errors propagate across interactions or persist in unresolved loops, making decisive errors difficult to distinguish from downstream failure symptoms. We propose \textbf{E}rror-Propagation \textbf{M}odeling for \textbf{F}ailure \textbf{A}ttribution (\textbf{EMFA}). EMFA constructs a structured representation of the failed trajectory, models both cascading propagation and persistent interaction loops, and uses propagation-aware candidate screening followed by counterfactual verification to identify the decisive agent--step pair. On the Who\&When benchmark, EMFA achieves state-of-the-art step-level attribution accuracy and remains competitive at the agent level. It improves the previous best step-level results by 3.45 and 4.40 percentage points on the Hand-Crafted and Algorithm-Generated subsets, respectively.
cs.AI / 46 / 2610.11613
AgentEvolver: System-Wide Self-Evolution Through Task Execution
Wentao Zhang, Fuchao Yang, Yilei Zhao, Xinrun Wang, Bo An
cs.AI
Abstract
An agent can complete a task without improving how it works. Turning task experience into reusable capability requires connecting the changed component to its evaluation and subsequent use. We present AgentEvolver, a system for developing capabilities during task execution while keeping the foundation model fixed. Eight entity families expose reusable operations, methods, agents, control flow, interfaces, and supporting state to revision through a common versioned lifecycle. A shared Runtime coordinates ongoing work, while persistent planning and recoverable context preserve task direction and supporting evidence. We evaluate task outcomes on SWE-bench Pro Public and examine capability changes in six application cases. The team reports an 82.08\% resolution rate with evolution, exceeding its reported baseline without evolution. The cases show retained capabilities entering later website, game, and research work, while also documenting incomplete objectives and an unsuccessful strategy. These findings distinguish improvement in a reusable component from success on the final task. AgentEvolver provides a concrete basis for studying capability accumulation through execution; independent-task transfer and total development cost remain open questions.
cs.AI / 47 / 2610.11655
Harness Evolution Hits a Ceiling: When Weight Training Should Begin
Yuan Tian, Bing Hu, Hao Wang, Binghang Lu, Fang Wu
cs.AI · cs.CL · cs.LG
Abstract
Improving a long-horizon LLM agent means evolving the harness around a frozen model or training its weights. We let a self-evolving harness make the system stronger first, then cross seed and evolved harnesses with base and trained weights to learn which gains the trained model keeps and which still need the runtime. We show that the right lever can be read off the agent's failure composition: labelling failed trajectories by the first signal that fires separates process failures (blocked calls, loops, exhausted step budgets) from content failures (a delivered plan that is poor). Harness evolution repairs the former, the behaviour it instils can be trained into the weights, and content failures are what weight training is for. On DeepPlanning, a self-evolving harness loop lifts the held-out score of Qwen3.5-4B from 0.16 to 0.30 and of Qwen3.5-9B from 0.32 to 0.44; for 4B, held-out delivery rises from 55% to 90% while content failures are left for the weights. LoRA adapters trained on evolved-harness trajectories internalise the gain: under the original harness they add +0.13 on held-out tasks for both sizes; on 4B they stack with the harness to more than double the held-out score, and on 9B the adapter alone matches the full evolution line, cutting content failures from a quarter of trajectories to one in twenty. A placebo adapter trained on answer-shuffled trajectories falls below the base model. The loop transfers to WebArena-Lite (+0.09 on 117 unseen tasks), where the gain lives in what the model sees and adapters do not add to it. The result is a diagnose-then-intervene rule applied twice: read the failure composition to choose between harness and weights, then read what the accepted edits changed to decide which gains to train in. Scores are four-rollout means against fresh anchors, same-night except where marked, across eight models from six families and two benchmarks.
cs.AI / 48 / 2610.11662
Lamarck's Driving School: Discovering Autonomous Driving Training Strategies through Evolutionary Competition
Yichun Ye, He Zhang, Ye Tian, Jian Sun
cs.AI
Abstract
Autonomous driving capabilities depend strongly on the distribution of scenarios encountered during training. Existing methods commonly construct or dynamically adapt training scenario distributions using surrogate criteria such as realism, difficulty, or risk. However, these predefined surrogates may misrepresent training value, leading to inefficient use of training resources. To address this limitation, we propose a Lamarckian evolutionary framework that replaces surrogate-based guidance with competition among candidate distributions. We formulate training strategy discovery as a multi-stage bilevel optimization problem and use Lamarckian evolution algorithm to approximate its solution. At the outer level, Darwinian crossover, mutation, and selection explore the scenario distribution space; at the inner level, policy learning acquires new capabilities, and Lamarckian inheritance transfers them to subsequent stages, allowing scenario distributions and policy capabilities to co-evolve. The resulting evolutionary trajectories reveal recurring stage-wise regularities among high-value distributions, characterized by capability accumulation through stage-wise challenge rotation. We further distill these regularities into a lightweight, reusable Lamarckian Training Strategy. Experiments show that, compared with the baseline, the complete framework reduces performance loss by up to 25.07%, while the lightweight strategy still achieves a 19.13% reduction. These results demonstrate that evolutionary competition can both discover effective training strategies and reveal reusable stage-wise patterns in how the value of training distributions changes with policy capability. Code is available on GitHub.
cs.AI / 49 / 2610.11663
Constrained Command-Conditioned Reinforcement Learning with Bandit Strategy Selection in Real-Time Strategy Games
Nick Leenders, Roy Lindelauf, Joost van Oijen, Boris Cule
cs.AI
Abstract
Deep reinforcement learning agents reach strong performance in real-time strategy games but can be brittle against opponents outside their training distribution. Separating strategic command selection from learned unit control allows different strategies to be selected for different opponents while reusing the same execution policy. This requires an executor that can follow different commands and measurable criteria for assessing whether it does so. We introduce a constrained command-conditioned Proximal Policy Optimization (PPO) policy, the executor, for MicroRTS, a real-time strategy environment. Discrete commands specify strategic objectives and behavioral requirements for economy, army composition, military posture, and worker policy over multiple environment steps; the executor determines the unit-level actions used to fulfill them. A Thompson-sampling bandit acts as the strategist, selecting command tuples from an estimate of the opponent's strategy built from in-game observations rather than opponent identity. In a controlled comparison with a flat PPO baseline trained with the same architecture, budget, curriculum and self-play league, the strategist-executor system wins significantly more often against three of the four strongest opponents on a training map, including the two strongest held-out ones (0.55 to 0.97 and 0.01 to 0.34), with no significant difference against the others.
cs.AI / 50 / 2610.11704
Intervention anchors and scientific verification in synthetic vascular predictive representations
Lingsen You, Yujun Guo, Xinyu Zhong, Zisu Peng, Wentong Wang, Li Shen, Junbo Ge
cs.AI
Abstract
Complete orthogonal predictive coordinates do not by themselves bind a latent direction to a named intervention. We present a mathematical and synthetic audit motivated by vascular device-vessel suitcordance. Capacity-matched least-squares predictors were exactly equivalent under complete fixed output transforms, whereas an anchor-only observer recovered interpretations only within the span of known perturbation signatures. Six three-dimensional configurations across 64 seeds gave a maximum paired prediction discrepancy of 6.7e-15 but a median untransported edit error of 1.513. Coordinate transport removed that error. Noisy and weak anchors constrained calibration stability, and changing the representation basis required recalibration or verified transport. Across 256 additional fits in dimensions 3-24, prediction equivalence persisted within 4.0e-15. We then evaluated nine deliberate runnable fault classes across 64 seeds. All 576 faulty executions completed, but each violated at least one reconstruction, prediction, delivered-edit or scope contract; all 320 valid control records passed. Repeating a faulty implementation gave exact self-agreement despite error against the separately computed simulator expectation. For one omitted-direction defect, probe coverage followed its analytic law, and rank-aware abstention protected unsupported interpretations. Scalar-noise experiments exposed both missed weak faults and excessive rejection under narrow relative tolerances. These controls provide an executable separation of prediction, semantic support and scientific acceptance. They are synthetic numerical audits, not clinical validation, neural JEPA-Anything replication, agent learning or patient treatment-effect estimation.
cs.AI / 51 / 2610.11707
DeltaReplay: Task-Relative Memory Reuse for Mobile GUI Agents
Yudong Bai, Yihong Chen, Quanming Yao, Yaqing Wang
cs.AI
Abstract
Memory-augmented mobile GUI agents store successful execution trajectories and reuse them in later tasks, but a stored trajectory rarely matches a new task exactly. The new task may use different parameters, share only some of its steps with a stored trajectory, or have no relevant record in memory. Forcing the agent to use irrelevant memory can mislead it, whereas discarding memory that may still be useful deprives it of guidance from past experience. To address this dilemma, we propose DeltaReplay, a step-level memory reuse framework that decides how to use existing memory without modifying it. We observe that the reusable part of a stored record is determined not by the record itself but by its relation to the new task, mainly through two factors: page-level consistency and action-level generality. We therefore store execution trajectories as paths in a transition graph, whose nodes (pages) and edges (actions between pages) capture these two factors. At reuse time, the action on each edge is split into a task-independent operation and task-specific parameters. DeltaReplay then compares each recorded step with the new task and the current screen, and decides whether to follow it, execute it after replacing its parameters, or leave it to the base agent. On AndroidWorld and SPA-Bench, DeltaReplay improves the task success rate over a base agent with the same backbone by up to 10.3 and 25.0 percentage points, respectively. These results indicate that deciding at each step how to use retrieved memory lets agents benefit even from partially matching trajectories.
cs.AI / 52 / 2610.11723
MultiWorldBench: Do Independently Controlled Views Describe One Shared World?
Zhangbo Xu, Ruoxi Zhang, Rui Hu, Yisong Wang
cs.AI · cs.CV
Abstract
Multiplayer world models must ensure that independently controlled views remain consistent with one shared and persistent world. We introduce MultiWorldBench, a diagnostic Minecraft benchmark containing 495 case configurations across seven task suites and ten capabilities, including independent control, cross-view motion, shared-state synchronization, persistence, structural reasoning, concurrent interaction, and delayed revisit. We evaluate Solaris, Gamma-World, and MineWorld, using Engine GT as a reference. Gamma-World achieves the highest ten-capability average among the generated systems at 21.39, followed by Solaris at 20.88 and MineWorld at 1.89, while Engine GT reaches 91.69. Gamma-World performs better on several control, shared-state, and revisit capabilities, whereas Solaris leads in cross-view motion and race-condition consistency. Nevertheless, all generated systems score at most 8.00 on state persistence and 1.33 on structural consistency, and none succeeds in spatial reasoning or building-identity preservation. Human preferences produce the same overall ranking and show strong alignment with the automatic evaluation, with a mean dimension-level Spearman correlation of 0.96. These results show that plausible individual views do not yet constitute a coherent multiplayer world.
cs.AI / 53 / 2610.11735
Onboard Marine Anomaly Detection on $Φ$sat-2: From Simulation-Based Development to In-Orbit Demonstration
Clotilde Szywala, Thomas Goudemant, Marjorie Bellizzi, Benjamin Francesconi, Adrien Girard
cs.AI · cs.CV
Abstract
Onboard Artificial Intelligence can improve responsiveness and bandwidth efficiency of Earth Observation systems by processing data directly on the satellite. This paper presents the experience gained from the development, onboard integration, and post-launch adaptation of a lightweight marine anomaly detection pipeline deployed on the European Space Agency's $Φ$sat-2 mission. The application combines sea segmentation, self-supervised feature encoding of marine regions, generic anomaly detection based on deviations from a normal sea state, and optional characterization of selected anomaly types. Before launch, the pipeline was trained and validated on simulated $Φ$sat-2 imagery to assess algorithmic performance and compatibility with resource-constrained onboard hardware. After integration and functional validation in the mission environment, early experiments on real $Φ$sat-2 acquisitions revealed a significant mismatch between simulated and in-orbit data. The pipeline was therefore retrained on real Level-1 imagery using an improved annotation strategy to better handle ambiguous marine regions, substantially enhancing performance. Beyond demonstrating the onboard feasibility of the application, the $Φ$sat-2 experience highlights the importance of robust annotation strategies and sensor-aware design, and shows that simulation-based development is valuable for pre-flight risk reduction, while reliable scientific validation requires representative in-orbit data and should be clearly distinguished from functional validation.
cs.AI / 54 / 2610.11738
Scalable AI Uncertainty Quantification via Generalized Laplace Active Subspaces
Wouter N. Edeling, Peter V. Coveney
cs.AI
Abstract
Reliable uncertainty quantification (UQ) is essential for deploying neural networks in scientific and high-stakes applications, but full Bayesian inference over the network parameters is computationally infeasible. We propose a low-rank generalized Laplace approximation for neural-network UQ based on a small number of data-informed curvature directions. Starting from a generalized Bayesian posterior defined through an empirical loss, we construct a local Gaussian approximation around a pretrained set of weights in this active curvature subspace. The posterior variances in the retained subspace are available in closed form, and the prior variance is calibrated by an empirical Bayes procedure. The generalized Bayesian formulation allows us to compare two posterior scalings: the standard Bayesian scaling associated with the summed negative log likelihood, and a mean-loss scaling in which the empirical loss is normalized by the number of data. A central finding is that the standard scaling induces a data-size dependent contraction of the posterior variance in the leading active directions. In regression problems, this can force the low-rank framework to retain additional weak-curvature directions in order to achieve nominal coverage of calibration data. When posterior samples are propagated through the non-linear network, these additional directions can degrade the coherence of the predictive intervals and shift the posterior predictive mean away from the pretrained model. In contrast, the generalized mean-loss scaling yields a more stable, lower dimensional active subspace and produces calibrated, coherent predictive confidence intervals. These results indicate that generalized Laplace active subspaces provide a practical and scalable route to calibrated uncertainty quantification in neural networks.
cs.AI / 55 / 2610.11750
Where Draft Trees Lose Target Mass: Exit-Guided Speculative Decoding
Shijing Hu, Xuancheng Ren, Zhihui Lu, Pan Zhou
cs.AI
Abstract
Tree-based speculative decoding verifies multiple draft continuations in one target-model pass, but finite trees built from draft scores face a fundamental draft-target mismatch. We ask whether better exact verification can increase acceptance on a fixed tree and how target feedback can improve the tree itself. Through a target-flow view, we identify a canonical exit law and prove that one plus target coverage sharply bounds the expected output-block length, including the bonus token, of any exact path verifier. All optimal verifiers share the same exit and bonus-token law, already attained by representative predraw-and-follow and sequential residual verifiers. This yields Tree Exit Verification (TEV), an exact, level-parallel procedure using one exit-node decision and one bonus-token decision. The exit law also identifies missing target probability, providing node-level feedback for Exit-Guided Draft-Tree Training (ExitTrain) on inference-time draft trees. Experiments across dialogue, code, and mathematical reasoning validate fixed-tree equivalence: ExitTrain increases average output-block length by 13%, while TEV reduces verifier-stage latency by 15%, yielding a 14% end-to-end speedup over DDTree. Our results distinguish two opportunities: better draft trees for higher acceptance and more direct verification for lower latency. Code: https://github.com/hsj576/TEV.
cs.AI / 56 / 2610.11754
What Output-Only Review Cannot Verify: Study Contracts for Research Agents
Eitan Waks, Ben Glocker
cs.AI
Abstract
Some defects in an AI-generated study can be identified from its artifacts; others require knowledge of what was approved before execution. We propose study contracts that bind declared experimental choices, run obligations and claim scope to recorded execution evidence, and distinguish this contract-relative verification from scientific truth. A diagnostic using eight self-authored clean/mutated pairs illustrates the information boundary. A deterministic checker applying a registered, fault-specific rule to approved and executed objects detected all eight registered mutations. Across eighteen recorded judge aliases given individual metadata-filtered packages without pair context or the registry-selected fault label, 104 of 144 mutated evaluation cases received defect flags; the remaining cases comprised 32 abstentions and eight terminal failures, with no explicit clean decisions on mutated cases. Some packages retained approval and execution fields, including digests. The prompt instructed judges to abstain when evidence was insufficient. These results characterize a deliberately information-asymmetric development setting; they do not isolate the effect of authoritative information from differences in task specification and rule selection, and they are not comparative verifier quality or agent reward hacking. We identify full-information comparisons, legitimate-adaptation controls and closed-loop agent evaluations as necessary tests of whether contract checks improve useful compliant completion under optimization.
cs.AI / 57 / 2610.11773
Safe Actions Alone Do Not Ensure Safe Agents: Identifying Unfulfilled Obligations with Guard Models
Youwei Feng, Yitong Zhang, Yuetong Liu, Jia Li
cs.AI
Abstract
Guard models are increasingly used to safeguard LLM-based agents, primarily by identifying actions that agents are forbidden to perform. However, identifying forbidden actions alone is insufficient to ensure agent safety. In this paper, we argue that agent safety also depends on identifying required yet unperformed safety-critical actions, which we call obligations. Our preliminary study on a popular benchmark for evaluating safety shows that 56.92% of GLM-5.3 trajectories contain unfulfilled obligations, compared with only 30.00% containing forbidden actions. This finding reveals unfulfilled obligations as a major and previously overlooked source of safety risk. However, to our knowledge, no existing benchmark evaluates whether guard models can identify these obligations. To close this gap, we introduce ObligationBench, the first benchmark for evaluating the capability of obligation identification, comprising 240 expert-validated trajectories covering issue resolution, feature development, and terminal operations. Our evaluation of 14 representative models reveals substantial limitations: the highest recall and exact-match rate are only 48.97% and 10.00%, respectively. To address these limitations, we develop ObligationGuard using 40,000 synthetic training examples. ObligationGuard achieves 57.52% recall and an exact-match rate of 21.67%, surpassing all evaluated models on both metrics. We call on the community to incorporate obligation identification into the design and evaluation of future guard models to improve agent safety.
cs.AI / 58 / 2610.11775
RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing
Ilya Lasy, Nora Yinuo Cai, Kola Ayonrinde
cs.AI · cs.CL · cs.LG
Abstract
Sparse Mixture of Experts (MoE) models scale more efficiently than dense models by routing tokens to modular expert networks that are only active for processing a fraction of tokens. A leading hypothesis for the performance of MoE models is that each expert specialises in a single, coherent domain. However, interpretability efforts that assume this hypothesis have generally been unsuccessful. We propose and present evidence for an alternative account that we call the Superposed Specialisation Hypothesis (SSH): experts specialise in a disjoint union of fine-grained features rather than one broad domain. Leveraging the SSH, we introduce RouterInterp, a method for interpreting expert routing that identifies Sparse Autoencoder features most predictive of routing decisions and produces unified natural language explanations. On gpt-oss-20b, RouterInterp explains expert routing with ${\sim}65\%$ higher detection accuracy than prior token statistics based methods. This work provides a scalable method for generating more accurate explanations of expert routing and increases our understanding of a previously uninterpretable component of foundation models.
cs.AI / 59 / 2610.11794
Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks
Haoyu Zhao, Zhengxu Yu, Zhiyuan He, Meng Fang, Rasul Tutunov, Haitham Bou-Ammar, Weilin Luo, Jun Wang
cs.AI · cs.CL · cs.CV · cs.LG
Abstract
Learning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past interactions but predict different outcomes in unseen states. We introduce Memento 3, building on the Memento series to enable frozen LLM agents to continually learn explicit world models through external memory. The agent maintains a natural-language rulebook as persistent semantic memory, recording revisable hypotheses about environment dynamics while leaving unknown aspects underspecified. It compiles this rulebook into executable code for prediction and planning. Through a continual loop of observation, reflection, rule revision, compilation, and verification, the agent uses prediction errors to refine both the rulebook and its code. Updated code is accepted only when the LLM judges it faithful to the rulebook and cell-exact replay reproduces the observed transitions. We investigate this process as a model-based route to recursive self-improvement (RSI): the agent autonomously explores the environment, revises its world model, and uses verified updates to guide subsequent interaction and learning, while the underlying LLM remains fixed. A population extension maintains multiple world models in parallel, sharing interaction evidence and using their predictions to guide exploration. On ARC-AGI-3, the single-model agent clears every level of all 25 public games, achieves a mean Relative Human Action Efficiency (RHAE) of 100.0, and uses 44% of the human action count. In an Atari Pong case study, a learned feedback controller wins 21:0 in each of three evaluated episodes with different openings, without further LLM calls.
cs.AI / 60 / 2610.11833
Probability-Signature Dynamics: Unpacking Modular Addition Learning Within Two-Layer Networks
Yunji Wang, Junjie Yao, Linyu Liu, Pinyan Lu, Zhi-Qin John Xu
cs.AI
Abstract
Neural networks trained on modular addition tasks often develop Fourier-structured representations that support exact generalization. While prior work has identified these Fourier circuits, the mechanism by which gradient-based training selects them from the data distribution remains unclear. We address this question using probability signatures, which express leading gradient interactions through conditional statistics of the training distribution. For modular addition, these signatures are cyclic shift operators and are diagonalized by the discrete Fourier transform, yielding approximately decoupled Fourier-mode dynamics. This explains the emergence of Fourier sparsity, frequency matching, and phase alignment. The same framework resolves a puzzle under label noise: corrupted examples can show faster early loss decrease than clean examples, despite lacking a coherent generalization rule. We show that noise increases conditional label collisions, strengthening early shared-coordinate reinforcement. Finally, this method can be applied to other operators. Taking XOR as an example, we observed the predicted frequency in experiments.
cs.AI / 61 / 2610.11862
LEVER: Adaptive Cost-Aware Proof Search Over AND/OR Graphs
Nihal Jain, Shuangjie Yao, Begum Cicekdag, Zhuo Zhang, Suman Jana
cs.AI
Abstract
Mathematicians value proofs for more than correctness: among correct proofs, simplicity, purity and the computational cost of finding them vary widely. Yet LLM-powered theorem provers largely search for any correct proof, and improve its quality only after it is found. We propose LEVER, a proof search algorithm that makes the objective over correct proofs programmable and optimizes it during search. LEVER scores partial proofs over an AND/OR proof graph, combining realized objective values with predictions for open subgoals, so the objective guides search before a proof is complete. The same mechanism optimizes computational cost, proof length, topical impurity, and even their weighted combinations, while the Lean kernel enforces correctness. On PutnamBench in Lean 4, under matched budgets, LEVER costs 34% less than a strong single-conversation agent while raising the solve rate from 80% to 96%. On reducing topical impurity, i.e., how far a proof strays from its theorem's subject, it improves over post-hoc refactoring (42% reduction against 33%) at two-thirds of the cost and more reliably; on proof length, the metric refactoring is built for, it approaches refactoring. Varying the objective's weights traces a quality-cost trade-off curve, so the user can choose how much a better proof is worth. Overall, LEVER is a performant, cost-efficient and tunable proof search algorithm for navigating the space of correct proofs.
cs.AI / 62 / 2610.11877
How Is Automated Research Evaluated? A Survey of Benchmarks and Evaluation Practices
Liulei Zhang, Dejing Zhou, Chuyue Huang, Guanhua Chen, Yutong Yao, Lidia S. Chao, Chi Man Vong, Derek F. Wong
cs.AI
Abstract
Automated research systems support literature synthesis, ideation, experiments, writing, and peer review, but their evaluation is dispersed across tasks, benchmarks, and studies that are difficult to compare directly. We review this literature from the perspective of evaluation design and evidence, covering six targets: literature synthesis, research ideation, executable workflows, scholarly writing and communication, automatic peer review, and end-to-end research. We compare task construction, evidence sources, evaluators, and scoring procedures to explain the capabilities assessed by different designs. Our synthesis highlights three recurring lessons: output checks, process checks, and human studies provide complementary information; evaluator calibration is specific to the property being assessed; and resource budgets and attempt selection are integral to interpreting performance comparisons. We identify diagnostic evaluation designs and documented gaps in supporting evidence, and translate these comparisons into reporting and audit recommendations for specific evaluation settings. The survey helps readers navigate existing evaluations, select appropriate benchmarks, and design subsequent studies.
cs.AI / 63 / 2610.11966
MindFlow: Mind Supernet Powered Thinking Flows for Research Idea Innovation
Mengdi Liu, Wenjue Chen, Wenyue Chen, Cheng Yang, Fanqi Kong, Zhangyang Gao, Xiaoxue Cheng, Yiheng Li, Yujian Yuan, Keliang Li, Hong Chang, Shiguang Shan, Chenglin Wu
cs.AI · cs.CL · cs.MA
Abstract
Research idea innovation is a fundamental engine of scientific progress, yet it remains difficult to generate and evaluate in a scalable and controllable way. This challenge lies in its inherently open-ended and multi-objective nature, where ideas should balance novelty, plausibility and feasibility. While recent LLM-based approaches have made progress through carefully designed prompts or agent pipelines, they are constrained by predefined, static ideation workflows. To address this limitation, we propose MindFlow, a framework that explicitly formulates ideation as a graph-structured Flow in Mind, which is composed of modular thinking operators and modeled by a probabilistic mind supernet. Given a research topic, a controller dynamically samples thinking flows to generate candidate ideas. This open-ended problem is optimized using a tournament-based relative ranking, enabling the controller to progressively favor higher-quality thinking flows. We further introduce an evaluation protocol that jointly assesses problem finding and problem solving, going beyond title- or abstractonly judgments. Across diverse topics, MindFlow shows its superiority as an explicit, controllable and optimizable research idea innovator.
cs.AI / 64 / 2610.11989
MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation
Zipeng Wang, Xinpeng Dong, Yuefan Wang, Pingchen Lu, Xian Wei, Kun Kuang, Fei Wu, Zhongxiang Dai, Min Zhang
cs.AI
Abstract
On-policy distillation (OPD) trains a student on its own generated responses using token-level teacher supervision. However, uniform weighting overlooks differences in token learning value, while existing weighting methods rely on predefined mappings from prediction signals to token weights. These mappings are not learned from the effectiveness of the resulting student updates, limiting their ability to adapt to evolving learning needs. In this paper, we propose MetaOPD, a bilevel optimization framework that jointly learns the student model and a lightweight token-weighting network. The inner objective updates the student through weighted OPD, while the outer objective optimizes the weighting network using validation loss on reference solutions after a virtual student update. Differentiating through this update connects weighting decisions to their effects on post-update performance, allowing the mapping from prediction signals to token weights to evolve alongside the student. Experiments on six mathematical reasoning and three out-of-domain datasets, covering two student scales and seven baselines, demonstrate the effectiveness of MetaOPD, with Avg@8/Pass@8 gains over OPD of 1.99/5.97 percentage points for the 0.6B student and 2.25/6.41 points for the 1.7B student.
cs.AI / 65 / 2610.12003
An Interpretable Approach to PDE Solution Discovery via Structural Experience Distillation
Yunpeng Gong, Huolong Wu, Can Yang, Min Jiang
cs.AI
Abstract
PDE solution discovery aims to identify explicit symbolic expressions for unknown physical fields from observations under known physical constraints. Existing methods, however, collapse data fidelity and physical consistency into a single terminal score used as the sole feedback signal, providing little information about which subexpressions are responsible for a candidate's final performance. This opaque terminal feedback severely limits the interpretability of the search process itself, offering no insight into why a candidate succeeds or fails. Consequently, reusable structures in otherwise suboptimal candidates are often discarded, whereas incidental syntax along successful search trajectories may be repeatedly reinforced. We propose SED-MCTS, a Monte Carlo tree search approach that distills structural experience from evaluated expressions and reuses it to guide subsequent symbolic solution search. Through counterfactual subtree interventions, SED-MCTS estimates local structural contributions, routes reliable evidence to the responsible construction edges, and preserves useful components in a refined structural archive. The approach naturally extends to coupled multiphysics systems. Across a diverse suite of PDE benchmarks, SED-MCTS achieves strong performance under a fixed evaluation budget and improves search efficiency and robustness under noisy or scarce observations.
cs.AI / 66 / 2610.12008
Complexity of Grounded Semantics and Preferred Semantics in Finitary Argumentation Frameworks
Jinfan Xu, Jieting Luo
cs.AI · cs.LO
Abstract
Abstract argumentation frameworks (AFs) introduced by Dung provide a formal foundation for non-monotonic reasoning in artificial intelligence. While decision problems for general infinite AFs typically reside at high levels of the analytical hierarchy ($Σ_1^1$ or $Π_1^1$), restricting the framework to be computably finitary reduces some of the complexity to the arithmetical hierarchy. In this paper, we present a complexity mapping of grounded and preferred semantics in computably finitary AFs across standard decision problems: credulous acceptance ($\Cred$), skeptical acceptance ($\Skep$), extension existence ($\Ex$), uniqueness ($\Uni$), and non-empty existence ($\NE$). For grounded semantics, credulous and skeptical acceptance are already known to be $Σ_1^0$-complete. We show that non-empty existence is also $Σ_1^0$-complete, whereas existence and uniqueness are trivial. These classifications are understood within the domain of valid computably finitary representations. For preferred semantics, using a computably finitely branching computation tree, $\Cred_{\pref}$ is shown to be in $Π_1^0$-c and $\NE_{\pref}$ is $Σ_2^0$-c. However, it is insufficient to reduce universal quantification and global uniqueness, leaving $\Skep_{\pref}$ in $Π_1^1$ and $\UniPref$ in $Σ_2^1$-c. Our results show the precise boundary where finitarity succeeds to bring reasoning down to the arithmetical hierarchy and where second-order quantification forces problems back into the analytical hierarchy.
cs.AI / 67 / 2610.12010
PulseBound: Future-Beat State Forecasting Under an Explicit Information Boundary
Chenyang Xu, Donglin Xie, Xi Xiang, Xiaoyu Li, Yufan Lu, Jiqiun Gao, Yi Zhao, Xin-Yi Li, Guangpu Zhu, Zijian Wang, Xiwen Yang, Dezhen Wang, Lin Chen, Shenda Hong, Leilei Li
cs.AI
Abstract
Predictive representation learning from photoplethysmography (PPG) can violate causal information access even with causal attention, as normalization, nonlocal transforms, or companion views may depend on withheld samples. We introduce PulseBound, a PPG representation learner combining physiologically structured future-beat prediction with an explicit stored-window information boundary. A content-independent cutoff separates the visible prefix from the prediction target. Prefix-only normalization, suffix replacement before derived-view construction, and aligned masking ensure that encoder inputs depend only on the visible prefix and cutoff. This yields stored-suffix invariance: with fixed model state, randomness, prefix, and cutoff, changing the stored suffix cannot change the forecast context. A shared horizon-conditioned head predicts nine rhythm and morphology descriptors for up to four extractor-valid future beats, using elementwise validity masks; optional ECG-derived pulse-arrival-time supervision is restricted to training. On MIMIC and VitalDB groups held out from PulseBound backbone pretraining, PulseBound reduces nine-state transformed-space MAE relative to last-visible-beat persistence by 28.06% and 22.22%, respectively, with gains in MAE, MAE-Skill, and Spearman correlation across all 40 source-cutoff-horizon cells. In a separate comparison of seven models on 13 downstream tasks, PulseBound achieves the best mean on nine frozen linear-probe and seven full-fine-tuning tasks. Stored-suffix interventions cause zero recorded changes in forecast contexts or predictions, with zero suffix-input gradients at audited precision under the stored-window interface. These findings separate three testable aspects of predictive physiological representation learning: information access, supervised future structure, and transfer.
cs.AI / 68 / 2610.12023
InterviewPlayground: A Simulation Environment for Evaluating AI Interviewers
Jonathan Ivey, Aimee Liang, Arthur Y. S. Wang, Madeline Mandell, Ziang Xiao, Anjalie Field
cs.AI · cs.CL
Abstract
Increasingly, AI interviewers are being developed to elicit open-ended responses in applications like market research, public polling, preference elicitation, and social science research. However, evaluating AI interviewers is challenging because they function in extended, multi-turn interactions where they must adapt to participant behaviors. To address this need, we develop InterviewPlayground, a simulation environment for evaluating AI interviewers using simulated study participants whose behaviors are grounded in social theory. Simulated studies in InterviewPlayground produce an InterviewReportCard, which assesses the performance of AI interviewers using a suite of validated measures. To test whether our simulation-based evaluations predict performance with human participants, we conduct 15 real qualitative studies with five AI interviewers, three interview topics, and 450 human participants and compare them to simulated studies in InterviewPlayground. We find that AI interviewer performance in InterviewPlayground predicts performance in human studies with an average Pearson correlation of 0.86 across 12 measures, and the simulated interactions from InterviewPlayground reproduce key findings from behavioral analysis of AI interviewers in the human studies. Together, these findings support the validity of InterviewPlayground in assessing AI interviewer performance and examining potential failure modes. Our work contributes a simulation environment for AI interviewers supported with empirical validation, and more broadly, a roadmap for future work to develop validated, simulation-based evaluations of conversational AI systems.
cs.AI / 69 / 2610.12086
EvoAlloc: A Self-Evolving Resource Allocation Agent for Efficient Program Evolution
Yanning Dai, Yuhui Wang, Nanbo Li, Wenyi Wang, Jürgen Schmidhuber
cs.AI
Abstract
LLM-based program evolution relies on evaluation feedback to guide the iterative search for high-performing programs. However, evaluation is often computationally expensive, making it essential to allocate limited resources to candidates that can most effectively advance the search. Existing LLM-based methods typically rely on fixed allocation strategies throughout the search, potentially wasting resources on low-value candidates while overlooking promising ones. We propose EvoAlloc, a self-evolving resource-allocation agent that learns from search experience to revise its strategy for allocating computational resources across candidates. EvoAlloc periodically consolidates prior search and allocation outcomes into reusable experience, which informs subsequent strategy revisions. It further uses a counterfactual exploration mechanism to occasionally evaluate candidates denied resources by the allocator, revealing their outcomes to enrich its experience for future strategy updates. Across coding and agent-harness optimization benchmarks, EvoAlloc requires 59-82% fewer full evaluations and 61-89% fewer total LLM tokens to reach baseline-level performance. Moreover, under the same full-evaluation budget, EvoAlloc achieves 8.7-12.0% higher final performance.
cs.AI / 70 / 2610.12090
Recompose and Refine Latent Reasoning Flows for Vision-Language-Action Models
Hongyu Shi, Sen Zhao, Zuyu Zhang, Lifeng Shen, Ding Zou, Xinyu He, Xu Zhang, Qinghua Zhang
cs.AI
Abstract
Latent reasoning enables vision-language-action (VLA) models to transform multimodal observations into task-relevant internal states before generating continuous robot actions. While existing methods learn to generate or refine such states for each policy query, they discard successful reasoning after execution and therefore reconstruct similar computation from scratch. We present Reasoning and Flow Memory (FLOWMEM), a unified VLA model that turns successful latent computation into reusable reasoning experience. Rather than appending a fixed retrieved context, FLOWMEM dynamically retrieves and recomposes compatible latent fragments as the embodied context evolves, forming a reasoning route that follows the temporal structure and progress of successful computation. The route is then refined using current visual and proprioceptive evidence before it conditions action generation. Experiments on RoboMME and LIBERO-Plus show that FLOWMEM attains 48.0% and 77.3% success, outperforming memory-free policies by 1.7 and 4.1 percentage points, respectively. These results demonstrate the value of reusing successful latent computation for closed-loop VLA control.
cs.AI / 71 / 2610.12129
An Investigation of Model Coherence: Narrow Finetunes Contradict Themselves Under Resampling
Robert Graham, Yariv Barsheshat, Phil Blandfort, Sabri Alouache
cs.AI
Abstract
A large body of research measures model coherence based on output variance without adequately considering competing causes. We identify two such causes, ambiguity and indifference, and we introduce a set of 175 questions where contradicting answers cannot easily be explained by either. We then measure incoherence in terms of contradictions when resampling answers to the same question. In contrast to other methods our metric has high specificity, and only ranks models as incoherent when the issues are glaring. Even so, we find narrow finetunes score poorly. Inspecting inconsistencies flagged by our method, we find that model organisms from the literature display severe issues such as identity conflation, introspection failures and rationalizations. These findings suggest that the pathologies induced by narrow finetuning may limit what these models can tell us about coherent misaligned behaviour.
cs.AI / 72 / 2610.12134
OA-MAP: Evidence-Grounded Multi-Agent Multimodal Framework for Interpretable Knee Osteoarthritis Progression
Sixu Chen, Mingrui Yang, Qiang Guan, Xiaojuan Li
cs.AI · cs.MA
Abstract
Knee osteoarthritis (KOA) progression prediction can support patient monitoring, requiring the integration of multimodal data and multidomain expertise. Moreover, isolated risk estimates provide limited insight underlying a prediction. To automate the progression assessment workflow and reduce manual effort while providing interpretable findings and supporting evidence, we present OA-MAP, an autonomous multi-agent framework for evidence-grounded assessment of structural and pain progression in KOA. The system incorporates modality-specific agents including MRI, X-ray, and clinical agents, together with a coordinator agent. This framework can autonomously recruit specialist agents, select tools for prediction and analysis, and retrieve literature as external evidence based on user request and available patient information. An uncertainty-informed human-in-the-loop mechanism enables clinicians to review and correct intermediate findings, triggering recomputation of affected results. We evaluate the prediction models using 600 participants from the FNIH Osteoarthritis Biomarkers Consortium cohort. On the test set of 100 participants, the fusion models achieve AUROCs of 0.80 for structural progression and 0.68 for pain progression. A case study illustrates how OA-MAP combines risk estimates with intermediate findings, cross-modal conflicts, literature support, and uncertainty indicators to support interactive review.
cs.AI / 73 / 2610.12135
Q-Shaped Options for Hierarchical Reinforcement Learning
Clarisse Wibault, Antoine Gorceix, Antonio Léon Villares, Alexey Zakharov, Evangelos Chatzaroulas, Michael Matthews, Eduardo Pignatelli, Jakob Foerster
cs.AI
Abstract
Learning to tackle long-horizon, goal-conditioned tasks requires an agent to reason over extended timescales and act across a broad range of states. In principle, Hierarchical Reinforcement Learning (HRL) addresses both challenges through the interaction between action (temporal) and state (spatial) abstraction. First, using an action abstraction to represent temporally extended behaviour as options reduces the effective decision horizon. Second, enabling different state abstractions at each level of the decision process permits greater data aggregation for learning. However, realising these two benefits of a hierarchical policy depends on learning an appropriate action abstraction. Current HRL algorithms fail in one of two ways. Some discard distinctions between options needed for optimal control, undermining hierarchy altogether. Others retain unnecessary distinctions, preserving horizon reduction, but forfeiting coarser state abstraction. In this work, we characterise three desiderata for an action abstraction. We introduce Q-Shaped Options (QSO) to address all three. QSO builds on an architecture with distinct state-value functions, Q functions and policies at each level of the hierarchy. It learns the action abstraction between consecutive levels as a shared encoder shaped by their respective Q functions. The low-level Q function uses the option as a goal, encouraging the abstraction to retain distinctions necessary for optimal control. The high-level Q function uses it as an action, encouraging unnecessary distinctions to be discarded. Across offline goal-conditioned locomotion and manipulation environments, QSO learns semantically meaningful option spaces and outperforms baselines, achieving non-zero performance in tasks where all other evaluated algorithms fail.
cs.AI / 74 / 2610.12168
Learning to Plan by Looking Back: Hindsight Hierarchies for Training Reasoning Models
Lars Simon, Holger Eble, Manuel Radons
cs.AI · cs.LG · cs.LO
Abstract
We introduce a self-improvement loop for reasoning models based on the following observation: Even when the difficulty of a problem exceeds the model's current solving abilities, an additionally supplied solution might enable the model to extract useful solution ideas in hindsight. We operationalize this by jointly training the same model to exhibit the following three capabilities: predicting solution ideas from problems alone, reverse-engineering ideas from problems and known solutions, and solving problems using provided ideas. The loop alternates between reverse engineering such ideas from problems with supplied solutions and using these ideas as additional supervision for joint training of all three capabilities. We give a formal specification of our method and a concrete instantiation for interactive theorem proving in the Lean theorem prover; empirical evaluation remains future work.
cs.AI / 75 / 2610.12176
Recursive Self-Improvement through Multi-Agent Self-Supervision
Hyunin Lee, Jinglue Xu, Jeffrey Seely, Donghyun Lee, Somayeh Sojoudi, Matei Zaharia, Yujin Tang
cs.AI
Abstract
Recursive self-improvement (RSI) of a model on non-verifiable tasks, such as open-ended research, faces a supervision bottleneck when its outputs exceed what even human experts can reliably assess, leaving the model itself (optimizee) as the best available optimizer and evaluator. However, a single model instance struggles to critique and improve its own complex reasoning under this homogeneous loop. To address this, we propose Multi-Agent Self-Supervision (MASS), an RSI method that alternates between evolutionary workflow optimization and supervised fine-tuning on self-generated trajectories. Guided by early findings that multi-agent topologies excel at complex reasoning, MASS prompts a single base model to iteratively propose, execute, and self-evaluate multi-agent workflows. Through an evolutionary search constrained by structural guardrails, the model optimizes these computational-graph-like orchestrations, discovering the most effective distinct roles and information routing for a given task. Over two MASS cycles with Qwen3.6-27B, the model achieves 1.2-1.6x higher performance per output tokens on four open-ended public benchmarks. Because the improved model subsequently acts as a better optimizer and evaluator, this alternating framework enables a continuous, recursive bootstrapping of the model's capabilities. Moreover, multi-agent traces are also more training-efficient: a student trained on them outperforms a single-agent student trained on 1.4x more training tokens. These findings suggest that jointly learning orchestration and bounded subagent execution from multi-agent trajectories can provide an effective signal for RSI.
cs.AI / 76 / 2610.12212
When Has a Bayesian Neural Network Sampled Enough? Adaptive Inference Time with Statistical Guarantees
Fabian Denoodt, Sibylle Hess
cs.AI
Abstract
Bayesian neural network predictions are commonly approximated using a fixed number of Monte Carlo samples per input, without controlling the resulting error that comes from this finite sample. We propose the use of confidence sequences to dynamically determine how many samples are needed while maintaining statistical guarantees. We consider several ways in which predictive probabilities are used, including identifying the most likely class, approximating the full predictive distribution, and resolving probability-threshold decisions. Sampling stops once the corresponding decision can be made with the desired guarantee. Experiments show that the method allocates the computational budget efficiently, assigning more samples to ambiguous inputs than to easy inputs while preserving reliable decisions and reducing overall latency relative to a fixed Monte Carlo budget.
cs.AI / 77 / 2610.12292
One Word Opens the Gate: The Option-Channel Attack on Typed Decision Models as Agent Guardrails
Seyedarmin Azizi, Erfan Baghaei Potraghloo, Massoud Pedram
cs.AI
Abstract
A typed decision model reads a piece of text and returns a probability over caller-defined options, each with a short written definition, generating no text. Recent work places these models in agent systems as guardrails: the component that reads a proposed tool call or incoming message and decides whether to allow it. We evaluate seven open-weight models in that role and report the two error directions separately: a fail-open error allows a prohibited action and is a vulnerability; a fail-closed error blocks a permitted one and is only a cost. On prompt-injection, jailbreak and toxic-content screening, accuracy at the allow-or-block decision ranges from 36% to 72% against a chance level of 50%. A low error rate in one direction only reflects which answer a model defaults to: one allows nearly everything, another blocks nearly everything. On a synthetic suite of agent tool calls, six lines of server log text that say nothing about the policy raise a gate's fail-open rate from 0% to 63% on a policy it otherwise decides correctly. Giving the permissive option a misleading name, with its definition and the judged text untouched, raises that rate to between 93% and 100% on the four models that place the label in their input. Every defense we tested is defeated, either by an attacker who targets its mechanism or by attacker-controlled text. Escalating the least confident decisions does not help either: a decision an attack has reversed is no less confident than the one it replaced. Parsing each policy field into a typed value does eliminate one attack, but it also makes the model unnecessary: a deterministic rule over those values reaches 100% accuracy on all six policies. These models can reduce how many cases reach a reviewer, but on this evidence they should not be the component that decides. Code is available at https://github.com/ArminAzizi98/option-channel-attack.
cs.AI / 78 / 2610.12303
Learning Probabilistic Logic Programs with Functional Gradient Guided Language Models
Saurabh Mathur, Sahil Sidheekh, Bhavan Vasu, Farbod Tavakkoli, Prasad Tadepalli, Kristian Kersting, Sriraam Natarajan
cs.AI
Abstract
Declarative logic programs offer a powerful and interpretable abstraction for encoding relational structure and neurosymbolic reasoning, by expressing dependencies as weighted compositional rules. However, inducing them from data remains fundamentally hard, bottlenecked by the combinatorial explosion of symbolic search spaces. LLMs have recently emerged as powerful hypothesis generators, but when used in isolation, they lack the capacity to do systematic inductive reasoning needed to reliably synthesize valid programs that fit complex relational distributions. We introduce grasp (Gradient-boosted Synthesis of Probabilistic logic programs), a neurosymbolic framework that casts relational structure learning as functional gradient boosting in which the weak learner is a first-order rule and the intractable inner search is delegated to an LLM proposal oracle. We evaluate grasp on four relational benchmarks spanning molecular toxicity prediction (Tox21), mutagenesis, and citation matching (Cora), and show that it improves over purely symbolic, neural, and LLM-based baselines, while producing interpretable weighted rule ensembles. By replacing combinatorial search with gradient-guided LLM hypothesis generation, grasp retains boosting guarantees without sacrificing the transparency of symbolic outputs.
cs.AI / 79 / 2610.12306
A Structural Theory of Cognitive Representation and Problem Solving,Contexts, Invariance, and the Knowledge Space
Antal Jakovác, András Telcs
cs.AI
Abstract
Learning and problem solving depend critically on the structure of internal representations. While many modern data-driven artificial systems achieve strong predictive performance, their learned representations often lack explicit structure for expressing abstraction, invariance, and task-relevant regularities. We propose a minimal structural framework in which representational operations relevant to problem solving, such as context formation, invariance recognition, representative selection, abstraction, and procedural reuse, are made explicit. The central notion is that of a \emph{context}, formalized as a partition of a subset of an underlying state space, which fixes the distinctions, granularity, and form in which a problem can be posed. Within this setting, invariance recognition and representative selection are treated as fundamental representational operations. The framework is realized as a Knowledge Space composed of two coupled graph structures: a Concept Graph that hosts constructed and refined concepts, and a Procedure Graph that encodes typed operations over representations. Together, these structures provide a minimal cognitive-representational algebra for operating on representations without assuming sophisticated inference, learning, control, perception, or motor mechanisms. Using simple illustrative examples and a finite weak-solver demonstration, we show that appropriate representational organization can simplify the form and scope of admissible regularities, even when problem solving is carried out by a fixed and limited solver. The contribution of the paper is structural rather than algorithmic: it identifies representational prerequisites for abstraction, invariance, and procedural reuse in problem solving, and states explicit success and failure conditions for the weak-solver setting.
cs.AI / 80 / 2610.12341
Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition
Kaisen Yang, Qingle Liu, Kejin Wang, Yicheng Zhao, Jieming Li, Shenghan Zheng, Ruize Yang, Bojun Yang, Heng Gong, Xiang Gao, Lanyue Zhang, Kaiyu Zhong, Zhuo Liu, Shaoxuan Li, Chengxi Li, Yong Yan, Weixuan Zhang, Tianwei Luo, Situ Wang, Youjie Zheng, Sihan Zhao, Shengyuan Wang, Huan-ang Gao, Jiazheng Xu, Xiaohui Xie, Wentao Han, Hongning Wang
cs.AI · cs.CL
Abstract
Adversarial games have driven advances from heuristic search to reinforcement learning, yet learning and adapting strategies from limited samples remain challenging. AI agents offer an alternative by turning game experience into revisions of executable policies. Building on heuristic learning (HL), we formalize Adversarial Heuristic Learning (AHL), a paradigm that uses AI agents as learning engines to refine game policies and supporting software while keeping model weights fixed. We introduce AAArena, a benchmark comprising 12 authentic adversarial games and 1,920 archived human programs, with an evaluation protocol modeled on real-world game competitions. Agents interpret rules, choose opponents, analyze replays, and revise game agents to achieve their highest ranking within fixed match and evaluation budgets. We evaluate \val{completedmodels} model and harness configurations: Opus5.5 with Claude Code earns 6 gold medals, while no evaluated configuration tops the remaining 6 human ladders. Performance is generally weaker in games with more complex rule specifications. Further experiments show that opponent selection and dense feedback support policy improvement, and that agents learn from both on-policy replays of their own matches and off-policy replays of other players' matches. These results highlight HL's potential in adversarial games and identify persistent challenges in game understanding, strategy implementation, and long-horizon policy development.
cs.AI / 81 / 2610.12360
Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict
Kaiser Sun, Bernal Jimenez Gutierrez, Hongjun Liu, Jingyu Zhang, Jie Gao, Mark Dredze, Daniel Khashabi
cs.AI · cs.CL
Abstract
When retrieved evidence contradicts an agent's prior beliefs, does it revise its answer, acknowledge uncertainty, or persist with an incorrect conclusion? Existing evaluations of agentic systems focus primarily on task success, offering limited insight into how agents handle such conflicts. We propose to evaluate agents on epistemic humility (EH): the agent's willingness to recognize, act on, and communicate uncertainty during task execution. We operationalize EH through three trajectory-level behavioral dimensions: Identify, Solve, and Escalate (ISE). Through knowledge conflict, situations where the backbone language model's parametric knowledge contradicts the evidence it encounters, or where two contextual sources disagree, we evaluate two conflict settings: (1) controlled conflict and (2) naturally occurring conflict during multi-step agentic execution, each paired with matched no-conflict controls. Evaluating four agents, we find that higher task accuracy does not necessarily correspond to greater epistemic humility: some high-accuracy configurations recognize conflicts during execution but do not communicate unresolved uncertainty in their incorrect final answers. Trajectory-level analysis further reveals that agents frequently detect conflicts in early steps of execution but fail to maintain or resolve them in later steps. Finally, we show that model-level interventions can improve EH, but often at the cost of task accuracy, suggesting that epistemic humility emerges from the interaction among the backbone model, agent harness, and evaluation environment.
cs.AI / 82 / 2610.12363
HANS: A Handwritten Answer Sheet Dataset for Noisy Hybrid Document Parsing
Xiazhen Wu, Wansong Qin, Yangbin Zheng, Liangda Fang, Zhan Li, Xiujie Huang, Liushen Zhou, Quanlong Guan
cs.AI · cs.CV
Abstract
Intelligent grading and automated scoring technologies constitute critical infrastructure for smart education. However, existing document parsing and handwriting recognition benchmarks are predominantly designed for well-structured printed documents or isolated mathematical expressions, lacking datasets that capture the complex characteristics inherent to student answer sheets, including multi-line derivation processes, heterogeneous mixtures of text and mathematical formulae, and noise artifacts such as strikethroughs. To address this gap, we introduce HANS, the first dataset explicitly constructed for real-world educational scenarios, encompassing mathematical expressions, natural language text, hand-drawn tables, and diverse noise patterns including corrections and deletions, accompanied by fine-grained annotations that establish a reliable foundation for robust recognition research. Building upon HANS, we propose NA-GOT, an end-to-end framework that achieves two-stage noise suppression through a lightweight noise suppression module operating at the feature level, complemented by a noiseaware attention mechanism incorporated into the decoding stage. Experimental results demonstrate that HANS poses substantial challenges to existing methods, while NA-GOT achieves significant improvements in both accuracy and stability for answer process recognition. The dataset will be made publicly available upon publication.
cs.AI / 83 / 2610.12375
OnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal Transport
Babak Barazandeh, Connor Swanson, Chinmay Kulkarni, Nikhil Mungel
cs.AI · cs.CL · cs.CY · cs.LG
Abstract
Agents are deployed in applications from trip planners and stock trading to IT incident triage. In most cases, LLM agents work autonomously with minimal rule-based safeguarding, leading to cost and safety issues from irreversible actions. Recent works resolve this either by using a safeguard agent to monitor behavior or evaluating logs post-hoc. The first adds cost and latency to every step; the second delivers its verdict after the run, when tokens are burned and damage is done. To overcome this, we propose OnTrack, a streaming monitoring mechanism that compares an agent's steps and dependencies against recorded successful runs to alert users or block the agent in about a millisecond per step. We study this problem in three regimes of decreasing access: full reference access (historical runs and tool schemas), intermediate access (only tool schemas), and no prior knowledge (only step logs as generated). Expectation of OnTrack's monitoring capabilities reduces as data access drops, ranging from plan violation detection to identifying loops, stalls, and repeated tool calls. Finally, we evaluate OnTrack using SWE-bench trajectories. Based on the first 8 steps, our method ranks failing trajectories below succeeding ones better than content similarity approaches (+0.057 AUROC). With an abort policy, we save about 18% of compute that would be burned on failing runs, where 83% of interrupted runs were actually heading to failure (5 out of 6 aborts were correct).
cs.AI / 84 / 2610.12393
HRIL: Learning Multimodal Synergy via Higher-Order Tensor Modeling
Qun Dai, Liangjian Wen, Jiang Duan, Yong Dai, Dongkai Wang, Maolin Wang, Mingjie Wang, Jianzhuang Liu, He Yan, Zhao Kang
cs.AI · cs.LG
Abstract
Self-supervised multimodal representation learning has achieved remarkable success across diverse domains, yet capturing synergistic information remains challenging due to the complexity of cross-modal interactions. Unlike the shared information across individual modalities, synergy arises when task-relevant signals emerge only from the joint configuration of multiple modalities and cannot be recovered from any modality in isolation. This work focuses on how to preserve the information capacity for such synergistic signals in multimodal representations. The key observation is that synergistic information is reflected in higher-order statistical dependence among modalities, which provides a principled target for explicitly modeling joint interactions. Motivated by this insight, we propose Higher-order Representation and Information Learning (HRIL), which constructs an empirical cross-moment tensor over modality embeddings to represent multi-way interactions. HRIL employs Tucker decomposition to obtain a core tensor, complemented by a synergy-aware regularizer that prevents energy concentration and preserves higher-order coupling capacity for synergistic information capture. Experiments on the controlled synergy task and real-world benchmarks demonstrate consistent improvements over existing multimodal contrastive methods, with notable gains on tasks dominated by synergistic interactions. Code is released at https://github.com/brightest66/HRIL.
cs.AI / 85 / 2610.12409
Searching for "Harmful Refusal": A Psychometric Audit of an AI Safety Benchmark
Christopher M. Stewart, Preston Botter, Natalie Sarabosing, Muye Zhang, Rachel Phinnemore, Shalini Ghosh, Hong Shen, Hoda Heidari
cs.AI
Abstract
Safety benchmarks typically report one overall score for a suite of datasets, each of which may target one or more safety-related attributes, so models with similar overall scores can have very different attribute profiles. Comparing models is more tractable at the level of individual attributes, yet it is often unclear whether even a single dataset's scores isolate any single attribute. One plausible candidate for such an attribute is harmful refusal, a model's tendency to refuse dangerous or policy-violating prompts. We examine whether it constitutes a single, measurable attribute in HELM Safety. Using a construct validity framework that stipulates that an attribute must exist before a test can measure it, we start with HELM Safety's four datasets that might plausibly target harmful refusal, but find that three are saturated. We subject the remaining dataset, HarmBench, to two psychometric tests to determine if a single attribute like harmful refusal could stand behind its score. First, multidimensional item response theory modeling strongly suggests that HarmBench does not measure a singular attribute. Second, a differential item functioning analysis finds items where models from different developers with the same refusal ability score differently. These flags largely disappear under scope-specific matching, a pattern consistent with aggregation effects but not sufficient to rule out domain-specific developer differences. Zooming out, HarmBench collapses distinct harm behaviors into one score, and the overall HELM safety aggregate further collapses HarmBench and scores from other datasets into a single top-line number. Any safety score that averages over datasets and items can hide saturation and conflate behaviors this way. We argue that a score should earn its single-attribute reading before models are compared with it.
cs.AI / 86 / 2610.12436
Ecology of AI Agents: Collaboration Creates a Population Threshold for Takeoff
Erin Crawley, Hidenori Tanaka
cs.AI · cond-mat.dis-nn · cs.MA · physics.bio-ph
Abstract
AI agents can now conduct real-world cyberattacks, scale up capabilities with the number of agents, and collectively pursue misaligned goals to obtain rewards. Together, these factors raise the risk of a population explosion of misaligned agents: agents could compromise computers and secretly deploy additional agents, creating a self-reinforcing cycle where larger populations develop greater collective cyber capability and expand further. This raises a fundamental question: What determines whether a population of misaligned agents remains contained or takes off into this self-reinforcing cycle? This population-level problem is ecological safety: unlike individual-agent or multi-agent safety with a fixed population, it concerns the dynamics of the population itself. Here, we develop an ecological theory of AI-agent populations based on a population growth equation in which fitness (growth rate) depends on cybersecurity capability. We show that, without collaboration, the population takes off only when individual-agent capability exceeds a critical threshold. With collaboration, however, collective cybersecurity capability increases with population size. This creates a critical population threshold: below it, the population declines; above it, the population takes off, even though individual-agent capability has not changed. In ecology, this phenomenon is known as the strong Allee effect. Because red teaming a small group of agents cannot guarantee ecological safety in larger populations, our theory calls for ecological red teaming and population pacing: gradually deploying larger agent populations in controlled environments, while measuring how cyber capability scales with population size, and estimating the critical population size for takeoff. Capability gains may lower this threshold, requiring re-estimation for each new model generation.
cs.AI / 87 / 2610.12452
BrickBench: Evaluating Agentic Brick Design
Peter Kulits, Yiqing Xu, R. Kenny Jones, Cordelia Schmid, Jiajun Wu
cs.AI · cs.CV · cs.GR
Abstract
We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design. Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be physically built. To do so, it must select parts from a discrete library and reason jointly about local and global constraints. We score validity, alignment, and design across three settings that vary in scale and part availability. We provide BrickAgent, an environment for coding agents to construct, inspect, and validate their designs. We find that leading agents largely satisfy verifiable physical and semantic requirements, but fall short of human designs. We release our benchmark and environment at http://www.brickben.ch
cs.AI / 88 / 2610.12466
On the estimation and validity of AI time horizons---a statistical look at the METR plot
Drew T. Nguyen, William Fithian
cs.AI
Abstract
METR's 50\% time horizon measures the human completion time of software tasks that an AI solves with 50\% probability, allowing AI capabilities to be expressed in interpretable units. On 228 tasks and 26 AIs, we recompute the time horizons using splines and item-response theory to relax the assumption that the AI difficulty of a task depends linearly on the log of human time. Our fitted spline can be interpreted as a function that \emph{converts} human time to AI difficulty; it is nearly flat in a region from 2--30 min but close to linear elsewhere. Hence, a time-horizon jump from 3 min to 30 min is much easier than one from 30 min to 5 hours despite the same multiplier of $10 \times$. Overall, we contribute time-horizon point estimates that perform better under a cross-validated suite of proper scoring rules, as well as diagnostic plots for assessing time horizons' construct validity. We suggest that time horizons be interpreted together with the diagnostic plots, especially as new time-horizon-based benchmarks are proposed or existing ones grow to include longer tasks.
cs.AI / 89 / 2610.10722
LinSlot: Exploiting Linear Representation hypothesis for unsupervised attribute discovery from slot based object representation
Sanket Gandhi, Utkarsh Giri, Varun Subramanium, Rohan Paul, Parag Singla
cs.CV · cs.AI
Abstract
This paper studies the problem of learning disentangled representations of objects and their attributes from raw, unstructured image data. Slot-based methods have shown considerable success in unsupervised learning of object representations from images. Block-slot attention-based methods extend this framework to attribute representations by assuming a uniform factorization of object representations into attributes, which may be suboptimal and consequently limit the quality of the learned representations. We therefore investigate a framework for jointly discovering object and attribute representations. Our key contribution is leveraging the Linear Representation Hypothesis (LRH), which postulates that composable concepts can be represented as linearly additive subspaces in slot representations. Based on this insight, we propose a probabilistic model connecting images, slots (objects), and blocks (attributes). We present an architecture that leverages block attention to connect attribute representations to slots and incorporates LRH in both object and attribute representation spaces. This architecture effectively optimizes the Evidence Lower Bound (ELBO) of the proposed graphical model. Our experiments demonstrate (i) effective discovery of disentangled object and attribute representations, (ii) empirical evidence for LRH in slot space, and (iii) the ability to perform image editing owing to the disentangled and interpretable nature of the learned representations. Our experiments on multiple datasets demonstrate improvements in DCI scores over state-of-the-art methods.
cs.AI / 90 / 2610.10782
VICO: Visual Environments Co-Evolving for Vision-Language Model Reasoning
Meng Lu, Ligeng Zhu, Olivia Xiao, Yuchen Zhuang, Zihan Wang, Kuncheng Wu, Bangya Liu, Yu Wang, Charles Fleming, Wenqi Shi, Xuan Wang
cs.CV · cs.AI
Abstract
Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for post-training vision-language models (VLMs), but it typically assumes a static training environment. As the actor improves, fixed tasks drift out of its learning frontier: many become trivial, others remain unsolvable; and the learning signal collapses. We argue that VLM post-training should evolve the visual environment alongside the actor, not just the actor itself. We propose VICO, a co-evolutionary framework in which an actor and an Environment-as-Rewriter (EnvRewriter) are trained jointly: the EnvRewriter edits verifiable image-side structures, such as scene graphs, chart tables, or protected region masks, and re-renders them to produce label-valid training samples whose difficulty is calibrated to the actor's current ability through a pass-rate-based reward. This loop continuously realigns task difficulty with actor capability without any additional human annotation. Across nine multimodal benchmarks spanning mathematical reasoning and visually grounded understanding, VICO-8B improves over its base model by up to +5.0% on out-of-domain tasks, surpasses the strongest self-evolution and text-editing co-evolution baselines by +4.3% and +8.4% respectively, and stays comparable to chart-specialized RLVR methods using 16-160 times fewer labeled samples. By shifting from human-labeled supervision to image-editing co-evolution, VICO offers a scalable path beyond static-corpus RLVR for visual reasoning.
cs.AI / 91 / 2610.11060
AffordDrive3D: Affordance-Aware World-Action Modeling with Spatial Understanding
Tianhui Cai, Xinglong Sun, Chao Fang, Zhenxin Li, Rui Song, Jose M. Alvarez, Yunxiang Mao, Jiaqi Ma, Langechuan Liu
cs.CV · cs.AI
Abstract
World-action models have recently improved autonomous driving by jointly learning future scene prediction and trajectory generation. Most existing approaches model the future primarily through RGB appearance, and recent works have begun to incorporate geometric prediction to improve spatial understanding. However, dense geometry describes the spatial layout of the entire scene without indicating which parts are most relevant to the ego vehicle's action. For driving, the model must also identify and anticipate where it can safely move and which regions may pose collision risks. Jointly modeling action-relevant regions and future geometry can provide the policy with both driving-relevant cues and their corresponding spatial structure. We therefore propose AffordDrive3D, an affordance- and geometry-aware world-action model that jointly learns future action-relevant regions and spatial structure. In order to capture the scene semantics and driving context needed for driving affordance prediction, we build AffordDrive3D on a VLM backbone to forecast drivable areas and collision-critical regions that directly affect ego motion, while predicting future geometry from RGB world-model latents. On NAVSIM, AffordDrive3D achieves state-of-the-art performance with 91.3 PDMS and 89.9 EPDMS, demonstrating the effectiveness of jointly modeling future affordances and geometry for trajectory planning.
cs.AI / 92 / 2610.11096
Continuous Ground-Truth Construction and a Recovery Policy for Air--Water Robotic Tracking
Jiangong Xiao, Zhe Sun, Kanzhong Yao, Yuanbo Bi, Haofei Zhao, Ruixuan Hu, Guan Huang, Xuelong Li
cs.CV · cs.AI
Abstract
Visual tracking across the air-water interface is challenged by splashes, bubbles, refraction, reflections, and abrupt appearance changes that can temporarily invalidate observations. This setting poses two coupled difficulties: first, for evaluation, image-only annotation cannot reliably describe the target's physical location during visual blindness; second, for online tracking, corrupted observations can contaminate motion estimates and appearance templates. We address the first difficulty with a construction pipeline that synchronizes camera frames with motion-capture poses, projects known target geometry, corrects underwater projection with a medium-gated residual, and subjects the annotations to manual review. This yields an evaluation-only cross-medium test set of 22,346 frames. We further introduce a Cross-Medium Recovery Policy (CMRP) centered on confidence-triggered template selection. It supplies MixFormerV2 with the fixed initial template, a window-best pre-trigger template, and a trigger-frame Kalman-guided image crop, together with their associated weights, without retraining the visual backbone. In the accuracy evaluation, CMRP achieves 49.90 Macro Success AUC, 2.95 points above MixFormerV2 Official. On selected cross-medium transition and occlusion-recovery intervals, CMRP increases MixFormerV2 tracking coverage from 47.91\% to 50.43\% relative to Official updating, while mean loss-to-recovery latency over successfully recovered videos decreases from 55.3 to 49.3 frames.
cs.AI / 93 / 2610.11148
SP-DocReader: Difference-Aware Self-Play for Precise Document OCR
Wenjie Liao, Xiaohui Song, Liangjie Zhao, Haonan Lu
cs.CV · cs.AI
Abstract
Accurate page transcription remains difficult for vision language models under limited input and training budgets. We present SP-DocReader, a self-play framework for optical character recognition (OCR) that targets residual errors after supervised fine-tuning. Reading Discrepancy Masking aligns reference and generated model tokens through a longest common subsequence, then scores unmatched positions with their full conditioning prefixes. Focused Fidelity Loss adds direct negative log-likelihood supervision at unmatched ground-truth positions. Only the OCR module is trained, while the backbone remains frozen. We derive the combined gradient to distinguish relative score optimization from direct supervision. Compared with SFT-2, SP-DR-3 reduces Vary-600K character error rate on both backbones. On Qwen3-VL-4B, it reduces character error rate by approximately 54 percent and improves DocVQA Average Normalized Levenshtein Similarity (ANLS) by 3.7 points. These results show the value of focusing self-play training on the discrepancies that remain after supervised fine-tuning.
cs.AI / 94 / 2610.11153
CARE: Constrained Attention Refinement for Fine-Grained Visual Classification via Teacher-Student Distillation
Ruibo Wen, Hang Shao, Yiming Lei
cs.CV · cs.AI
Abstract
Fine-grained visual classification requires models to recognize subtle local traits while exposing the visual evidence behind their predictions. Class-specific attention pathways provide a natural basis for interpretable recognition, but their constrained prediction structure limits discriminative capacity and underuses intermediate representations from strong pretrained backbones. To address this problem, we propose CARE, a constrained attention refinement framework for interpretable fine-grained recognition via teacher-student distillation. CARE keeps the final prediction and explanation within a class-specific attention student, while introducing a training-only auxiliary query teacher that reads selected intermediate DINOv2 layers with learnable queries. The teacher fuses multi-level representations and transfers logit-standardized class-discriminative knowledge to the student. To further refine the explanation pathway, we design diversity and sparsity terms to regularize student attention heads, reducing redundancy and encouraging compact trait localization. Experiments on CUB, Oxford-IIIT Pet, Stanford Dogs, and Stanford Cars show that CARE achieves strong classification performance under an interpretable frozen-backbone setting, reaching 78.5% Top-1 accuracy on CUB. Faithfulness analysis with insertion and deletion metrics further indicates that the top-ranked attention regions retain class-relevant evidence for explanation.
cs.AI / 95 / 2610.11171
VAMR: Multi-Question Agentic Reasoning for Efficient Long-Form Video Understanding
Runquan Gui, Hanzhu Chen, Zehao Wang, Hanxin Zhu, Xin Li, Zhibo Chen
cs.CV · cs.AI
Abstract
Long-form video understanding often involves multiple questions about different aspects of the same recording. Yet existing video agents typically process each question through an isolated tool-use trajectory. This repeatedly restarts video exploration and memory construction, missing opportunities to acquire evidence jointly and progressively build a shared understanding that supports the complete question set. We introduce \textbf{VAMR} (\textbf{V}ideo \textbf{A}gent for \textbf{M}ulti-Question \textbf{R}easoning), which coordinates all questions about a video through one shared tool-use trajectory. At each round, a persistent policy model can invoke tools for one or more unresolved questions and submit answers for questions with sufficient evidence. Question-conditioned visual perception retrieves fine-grained clues for several questions in one call, while layered multi-question memory integrates reusable context into a shared video story and preserves separate evidence for individual questions. After supervised fine-tuning initializes this interaction protocol, we propose question-horizon policy optimization (\qhpo) to optimize shared trajectories in which questions progress and finish at different rounds. Specifically, a question-level critic estimates the value of each active question, while round alignment maps each question advantage to the rounds that directly serve it before the aligned advantages are aggregated to optimize the shared actor. Across LVBench, Video-Holmes, and LongVideoBench, VAMR achieves the highest accuracy overall and the fewest reasoning rounds among iterative methods. On LVBench, it reaches 62.1\% accuracy, exceeding VideoARM by \textbf{4.3} points while reducing reasoning rounds and processed frames by \textbf{85.9\%} and \textbf{61.4\%}.
cs.AI / 96 / 2610.11374
Rethinking Contrastive Loss in CLIP Post-training: A Complementary Framework with Frozen Text Encoder
Zidan Wang, Yaqian Li, Xiaokai Zhang, Kaiwen Long, Kun He, Hanpeng Liu
cs.CV · cs.AI
Abstract
CLIP serves as a foundational vision-language model and the de facto vision encoder for downstream VLMs such as LLaVA. Post-training offers a lightweight route to refine CLIP, but recent work argues that the standard contrastive loss is unsuitable for post-training due to catastrophic forgetting under small batches, motivating designs that abandon the contrastive objective in favor of distillation. We revisit this premise and find that, for the InfoNCE objective, the reported forgetting is driven primarily not by insufficient negatives but by an inappropriate magnitude of the contrastive temperature $τ$: with $τ$ set sufficiently small, contrastive post-training improves rather than degrades the pretrained CLIP, which we explain through the temperature dependence of the InfoNCE gradient. Building on this finding, we propose \textbf{ComCLIP}, a lightweight single-epoch post-training recipe that freezes CLIP's text encoder---so the refined vision encoder is a drop-in replacement with unchanged architecture and inference cost---and trains the vision encoder with a properly-tempered contrastive loss, an MSE anchoring loss against the original CLIP, and a relational distillation loss from DINOv2. Over multiple seeds, ComCLIP matches the self-distillation baseline CLIP-Refine on zero-shot classification while significantly improving the transferability of visual features, measured by linear probing ($48.99$ vs.\ $42.28$ on ViT-B/16), and on ViT-L/14 it also improves MMVP over CLIP-Refine ($24.20$ vs.\ $19.01$); CLIP-Refine remains stronger on image-text retrieval. Used as a drop-in vision encoder for LLaVA-1.5-7B without re-aligning the projector or LLM, ComCLIP yields no net change across $8$ VLM benchmarks, i.e., the refinement does not break downstream compatibility. Code and models are available at https://github.com/showstarpro/ComCLIP.git.
cs.AI / 97 / 2610.11381
EvoKnow: Continual Knowledge Evolution for AI-Generated Image Detection
Zhiheng Peng, Wenwei Jin, Yangshi Ge, Siyu Xia, Jiawei Li, Xu Tang
cs.CV · cs.AI
Abstract
AI-generated image detectors are commonly trained on fixed generator domains and become difficult to maintain as new generative models emerge. Continual adaptation is challenging because replaying historical generated images is costly, whereas updating shared parameters with limited current-domain data can overwrite prior forensic knowledge. We propose EvoKnow, a replay-free framework that formulates continual AI-generated image detection as forensic knowledge evolution. EvoKnow preserves a shared forensic basis learned from base domains, incrementally adds isolated residual experts for complementary generator-relevant evidence, and retrieves expertise through an Analytical Incremental Router (AIR) updated in closed form from current-stage generated images and accumulated sufficient statistics. Experiments demonstrate effective cross-generator generalization, few-shot expansion, and long-horizon continual adaptation. With ten generated images per arriving generator, EvoKnow achieves 96.70% average accuracy on non-base GenImage generators and 94.48% accuracy on Chameleon without target-benchmark adaptation. Under a strict replay-free continual learning protocol, EvoKnow achieves state-of-the-art continual learning performance, attaining 96.32% mean stage-wise accuracy and 4.32% average forgetting.
cs.AI / 98 / 2610.11419
Missing Modality-Aware Calibration for Trustworthy Brain Tumor Segmentation
Sol Lee, Hyunji Kim, Sungrae Hong, Donghee Han, Mun Yi
cs.CV · cs.AI
Abstract
Multimodal brain tumor segmentation typically leverages multiple MRI modalities, yet incomplete modality acquisition is common in clinical practice due to protocol heterogeneity and scan failures. Although recent methods maintain segmentation accuracy under missing modality conditions, they frequently overlook prediction reliability, leading to miscalibrated confidence estimates that hinder clinical adoption. Existing calibration techniques are largely modality-agnostic or assume that prediction difficulty decreases monotonically as additional modalities become available. However, in brain tumor segmentation, prediction difficulty depends primarily on which modalities are absent rather than how many, leading to combination-specific and spatially heterogeneous calibration errors. To address this, we propose Missing Modality-Aware Local Temperature Scaling (MMA-LTS), a post-hoc voxel-wise confidence calibration method. It estimates a spatially adaptive temperature field conditioned on a modality-availability learnable token and a voxel-wise difficulty score. Experiments on BraTS 2020 and FeTS 2024 show that MMA-LTS improves calibration while preserving the segmentation accuracy of state-of-the-art models across diverse missing-modality scenarios, thereby enhancing trustworthiness toward clinical deployment.
cs.AI / 99 / 2610.11460
ProtoSemImage: Image-Valued Prototypes with Deformable Row Alignment for Interpretable Document Classification
Mohammad Zare, Pirooz Shamsinejadbabaki
cs.CV · cs.AI
Abstract
Prototypes in classification models are almost always vectors, and a vector has no readable form. This paper asks what happens when a prototype is an image. Documents give the question a natural form, because a document can be rendered as a multi-channel image in which every token becomes a pixel, so a class representative can take the same shape and the same channel semantics as the inputs it stands for. ProtoSemImage represents each class by one or more visual archetypes: prototype images in a four-channel HSV space whose channels carry named linguistic factors. A Skip-Gram objective learns that color space end to end through a four-dimensional bottleneck, discourse boundary rows become differentiable typed difference rows, and classification reduces to 2D visual template matching: a deformable row alignment between a document image and the archetype bank, in the spirit of dynamic time warping. Because the match is a spatial pattern comparison rather than a linear readout, the model reports where an input departs from its archetype and along which channel, and a generative head decodes each archetype back into text. The image representation works: it beats an otherwise identical model with vector prototypes in all three paired seeds, by between 4.3 and 11.8 points on a ten-class task. The distance-based matching does not. A diagnostic that keeps the representation fixed and swaps only the classifier recovers the sequence baselines, which locates a 20.6-point shortfall in the matching rather than in the color compression, and a benchmark built so that a pair of documents shares a bag of words and differs only in arrangement confirms the layout-preservation it was designed for. We report both directions, because for a representation whose whole purpose is inspect ability, the failure modes are as informative as the gains.
cs.AI / 100 / 2610.11526
MSGAT: Multi-Head Spiking Graph Attention with Similarity-Space Fusion for Image-Text Retrieval
Xintao Zong, Wenxuan Liu, Jianhao Ding, Zhaofei Yu, Tiejun Huang
cs.CV · cs.AI
Abstract
Spiking neural networks (SNNs) offer an energy-efficient computing paradigm through sparse event-driven computation, showing great potential for efficient multimodal learning. However, applying SNNs to high-level multimodal tasks, such as image-text retrieval (ITR), remains challenging, since sparse spike representations make it difficult to capture semantic structures required for cross-modal alignment. Existing spiking ITR methods rely on local alignment and additional soft-label supervision during training, while lacking awareness of structural and multi-granularity relationships. To address these issues, we propose a Multi-head Spiking Graph Attention Network (\textbf{MSGAT}) for structural modeling and equip it with dynamic attention heads to capture complementary relational patterns and enable spike-driven graph reasoning and aggregation. However, within a two-branch multi-granularity fusion framework, the fine-grained spike representations generated by MSGAT are sparse and discrete, whereas the global representations are continuous, making conventional feature-level fusion susceptible to interference across heterogeneous representations. Therefore, we introduce \textbf{Sim-Fuse}, a similarity-space fusion alignment strategy integrating coarse- and fine-grained matching relations while avoiding direct fusion of heterogeneous representations. Experiments on Flickr30K and MSCOCO show our method outperforms ANN methods under matched settings and existing SNN retrieval baselines. Moreover, with only two time steps, our SNN achieves comparable or superior performance to its ANN counterpart while reducing theoretical module-level energy by 55\%. The code is provided in the Supplementary Materials.
cs.AI / 101 / 2610.11610
Beyond Report Imitation: Clinically Aware Multi-Image Ultrasound Report Generation from Visible Evidence
Yuchen Yang, Xin Wang, Lufan Wang, Yinghong Pan, Yujuan Feng, Yuqing Yang
cs.CV · cs.AI
Abstract
Generating ultrasound reports from multiple images requires aggregating clinical evidence across views, yet archived key frames capture only part of the dynamic examination. Raw-report imitation is therefore misaligned with visual supervision: content that is clinically valid for the full examination may be unverifiable from the images available to a model. This gap creates a clinical behavior alignment problem. A model must preserve visible findings, avoid diagnostic reversals and unsupported completion, and not collapse into conservative templates. We propose CAMEO, a Clinically Aware Multi-image Evidence-grounded Orchestration framework for ultrasound report generation. Stage I learns ultrasound visual-language primitives; Stage II performs Cross-View Evidence Grounding by distilling trusted visible report points into multi-image QA and report-style supervision; and Stage III performs Clinically Aware Preference Alignment using clinical-error-oriented preference pairs. From USReport, we construct USReport-Distilled with 17,670 evidence-grounded paired-image training instances and USReport-Pref with 21,869 preference pairs; we additionally use 25,631 PubMedVision-US ultrasound instruction samples for domain adaptation and multi-image instruction tuning. On the primary USReport-Distilled benchmark, CAMEO improves over EchoVLM from 0.25 to 0.40 BLEU-1, 0.28 to 0.45 ROUGE-1, and 0.27 to 0.43 METEOR, while raising ClinicalScore from 55.02 to 74.20. These results underscore the value of evidence-grounded supervision, clinically aware alignment, and clinically structured evaluation for reliable ultrasound report generation.
cs.AI / 102 / 2610.11617
TAM: Task-Aware Memory Distillation for Efficient Spatiotemporal Prediction
Yuqi Li, Xiaoqin Feng, Fan Xu, Weilun Feng, Chuanguang Yang, Yingli Tian, Hao Wu
cs.CV · cs.AI
Abstract
Knowledge distillation enables efficient spatiotemporal prediction by transferring knowledge from an accurate teacher to a compact student. However, matching outputs or features independently for each sample leaves cross-sample predictive structure underused. Exploiting this structure requires representations and historical references that reflect the dynamics of each task. We propose TAM, a Task-Aware Memory Distillation framework that organizes a frozen teacher's knowledge into a bounded, retrievable history. Memory entries encode latent features, forecast changes, or flow residuals, while task-specific selection rules identify relevant historical references. The student either matches the teacher's similarity distribution over shared references or regresses observation-conditioned residual prototypes. These objectives complement supervised prediction and conventional distillation. The teacher, memory, and auxiliary adapters are used only during training, leaving student inference unchanged. We evaluate TAM on video prediction, weather forecasting, and traffic flow prediction across multiple teacher-student configurations. Averaged over four paired runs, adding TAM improves SSIM on all six video datasets and reduces MSE on five relative to the corresponding KD baselines. Mean paired MSE reductions reach 1.86% on KittiCaltech, 1.93% on WeatherBench with a gSTA teacher, and 1.01% on TaxiBJ. These results demonstrate the utility of historical teacher supervision across distinct forecasting tasks without additional student inference cost.
cs.AI / 103 / 2610.11818
From Pixels to Structure: Lightweight Vision-Language Models for Document OCR and Structured JSON Extraction
Uddipan Basu Bir, Vincent Christlein, Andreas Maier, Mathias Zinnen
cs.CV · cs.AI
Abstract
While massive, closed-source Vision-Language Models (VLMs) set strong benchmarks for document understanding, their dependence on commercial APIs limits adoption in institutional archives due to data autonomy concerns, recurring costs, and the environmental footprint of hyperscale computing. This is especially acute in heritage digitization, where documents include historical handwriting, domain-specific terminology (e.g., jewelry, prehistory, architecture), and non-standard layouts requiring high-dimensional structured extraction. We present a comparative study of eight open-source lightweight VLMs (up to 7B parameters) for Optical Character Recognition (OCR)-to-structure across three university heritage collections. Given a document image, models must extract text and generate schema-compliant JSON, enabling automatic validation and downstream use. We evaluate models under a constraint-aware protocol across zero-shot, few-shot, and fine-tuning settings, measuring extraction fidelity and structured-output quality using Character Error Rate (CER), Approximate Normalized Levenshtein Similarity (ANLS*), and mean Average Precision F1 (mAP-F1). Against a fine-tuning baseline, we further test the independent impact of (i) hyperparameter optimization, (ii) classical image preprocessing (illumination flattening, denoising, and CLAHE), and (iii) multi-stage training. Finally, we analyze the trade-off between dataset-specific fine-tuning and a single multi-dataset checkpoint, where joint training enables one model to operate across collections but can shift performance between datasets. Overall, we show that carefully adapted VLMs with up to 7B parameters can provide a sustainable, private, high-performing alternative to manual transcription or commercial black-box systems, and we offer actionable guidance for heritage institutions seeking institution-controlled OCR-to-JSON extraction.
cs.AI / 104 / 2610.11826
From Suppression to Repair: Mitigating Object Hallucination in Large Vision-Language Models via Localized Distribution Alignment
Chen Zhao, Xingping Dong, Jiachun Shi, Liang Peng, Chong Wang, Zhen Lei, Ran He, Bo Du
cs.CV · cs.AI
Abstract
Object hallucination remains a major obstacle for large vision-language models (LVLMs) to generate reliable content. An intuitive mitigation strategy is to suppress hallucination-related components in hidden representations. However, these components may also contain useful information, and suppressing them can weaken the model's multimodal capabilities. In this paper, we propose ResOT, a training-free method that repairs representations at inference time through localized distribution alignment. Specifically, ResOT projects dominant hallucinated directions away from the faithful subspace, forming a low-dimensional residual subspace for intervention. Within this subspace, ResOT uses Gaussian optimal transport (OT) to align the hallucinated distribution with the faithful one. The resulting map defines repair targets with minimal changes to the original representations. At inference, ResOT adaptively controls how far each token state moves toward its OT target. Experiments on three representative LVLMs show that ResOT substantially reduces object hallucination while improving image caption quality and multimodal performance across multiple benchmarks. Code will be released.
cs.AI / 105 / 2610.11907
Beyond Visual Enhancement: Adaptive Multi-Context Steering to Mitigate LVLM Hallucinations
Shuran Ma, JiaLe Li, Yuxin Dong, Shan Zheng, Qingyun Jiang, Xiang Chen, Qi Zhu, Deyi Ji, Yifan Yang, Jianfeng Pan, Yu Tian, Xue Yang
cs.CV · cs.AI
Abstract
Hallucination remains a significant challenge in Large Vision-Language Models (LVLMs). Existing training-free methods generally mitigate hallucinations through contrastive decoding or visual enhancement, often increasing the relative influence of visual evidence during generation. This raises a fundamental question: Can LVLMs dynamically regulate the contributions of different context sources to suppress hallucinations? In this work, we investigate and quantify how LVLMs coordinate multiple context sources during decoding and examine how this intrinsic behavior can guide hallucination mitigation. We find that LVLMs exhibit an intrinsic vision-attending tendency that can guide adaptive visual steering, while textual contexts can also contribute to hallucination mitigation. Motivated by these findings, we propose AIMS (Adaptive Information Multi-source Steering), a lightweight training-free framework that adaptively coordinates visual, prefilled textual, and generated contexts during decoding. Specifically, AIMS constructs compact prototypes for the three context domains and estimates their affinities with the current query to determine head-wise steering weights. The resulting multi-source steering direction is applied to the query representation, enabling adaptive context integration without additional model training or auxiliary forward passes. Extensive experiments across multiple LVLMs and decoding strategies demonstrate that AIMS effectively mitigates object hallucination while maintaining competitive general-purpose multimodal capabilities.
cs.AI / 106 / 2610.11993
DataVista: Diagnosing Multimodal LLMs on Data Video Understanding
Yupeng Xie, Zhenyang Wang, Jiayi Zhu, Yinghao Tang, Zhouan Shen, Yiyu Chen, Yuyu Luo
cs.CV · cs.AI · cs.CL
Abstract
Data video is a media form that integrates data visualization with video narrative, widely adopted in news reporting and business analysis. Compared with general video understanding, data video understanding places greater emphasis on accurately reading data from animated charts, integrating evidence across charts and time, and understanding how narrative organization and visual design communicate information. Yet existing benchmarks target either general videos or static charts, and data video understanding has not been systematically evaluated. We present DataVista, the first benchmark for data video understanding, containing 961 real-world data videos and 6,775 evaluation questions organized under a three-level progressive capability framework (data perception, temporal reasoning, narrative understanding) with 10 fine-grained question types across five topic domains. Systematic evaluation of 19 mainstream MLLMs shows that the best-performing model, Gemini-3.1-Pro, achieves 70.0% overall accuracy, still far below human expert performance, with models performing worst on Causal Reasoning and Narrative Structure. Increasing frame counts and adding subtitles mainly benefit data perception and temporal reasoning, with limited gains in narrative understanding. Further analysis of model responses identifies typical failure modes in chart reading, evidence judgment, and instruction understanding. The benchmark is available at https://github.com/HKUSTDial/DataVista.
cs.AI / 107 / 2610.12127
LVS: Local View Synthesis from Relative Camera Pose by Reusing Previous Views
Qizhou Huo, Xuan Sun, Yongfei Guo, Zhipeng Wang, Yuanhao Gong
cs.CV · cs.AI · cs.GR · cs.MM · eess.IV
Abstract
Interactive scene exploration requires frequent view updates, although small camera motions preserve much of the visible content. Conventional 3D Gaussian Splatting nevertheless renders each target view, leaving this image overlap unexploited. Reusing rendered images offers an alternative. Geometric warping alone cannot recover newly exposed content and remains sensitive to depth errors. We propose a per-scene framework that replaces repeated scene rendering for nearby views with relative-pose-guided RGB-D image reuse. Geometric warping uses depth and relative pose to transport source content, while a lightweight multiscale network predicts RGB residuals to correct artifacts and infer missing appearance. Cached source features further reduce repeated computation. On GS-render, residual refinement improves PSNR by 0.72~dB over pure warping; evaluations on captured and rendered scenes demonstrate low query latency. This separation of scene rendering from local view updates supports responsive scene exploration, with potential applications in augmented and virtual reality.
cs.AI / 108 / 2610.12182
AI-Based On-Board Maritime Object Detection for Earth Observation Payload Data Reduction on Versal Embedded Hardware
Thomas Goudemant, Aurélien Bobey, Omar Hlimi, Marjorie Bellizzi
cs.CV · cs.AI
Abstract
Very-high-resolution Earth-observation satellites acquire more data than they can store and downlink, while in maritime surveillance the vessels cover a tiny fraction of each scene. We study onboard vessel detection as a way to select what is downlinked, which reduces the data according to its content rather than coding every pixel; it is complementary to conventional onboard compression. The work follows three axes. (i) Data and algorithm: a controlled dataset is generated from 68 annotated Maxar scenes with 43 vessel classes, and a YOLOX-S detector is trained on it. (ii) Embedded deployment: the detector is quantized and deployed on the DPU of a Versal VC1902, with a limited loss of detection quality and a processing time of a few seconds per scene. (iii) Data reduction: we propose several downlink modes, from metadata only (box, class and score of each detection) to image crops around vessels, tiles holding detections, or the whole scene with a degraded background, and estimate from the measured detection errors the trade-off each offers between the vessels kept and the volume downlinked. On our dense harbor and coastal scenes, tiles keep 98% of the vessels with 29% of the scene volume, and crops 83% with 3%.
cs.AI / 109 / 2610.12229
VibeEdit: Image Editing with Canvas Instructions
Jinjing Zhao, Fangyun Wei, Yitong Wang, Xiuyu Wu, Yunuo Chen, Yang Yue, Sirui Zhang, Wenbo Wang, Hongyang Zhang, Dong Chen, Yan Lu, Chang Xu
cs.CV · cs.AI
Abstract
In text-guided image editing, describing the desired change is often straightforward, but identifying the intended object or region can be cumbersome, especially when several objects look alike. We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image. Together, these annotations form a canvas instruction that specifies where to edit and what to change. Our editor, VibeEdit, follows these instructions to perform object addition, removal, replacement, attribute modification, and movement without a separate text prompt. We construct 1.55 million source-target edit pairs with object masks and structured edit descriptions, from which we render canvas instructions during training. We adapt Qwen-Image-Edit with layer-decoupled conditioning that separately encodes source images and canvas instructions for image editing. We train the model with region-weighted supervised fine-tuning, followed by rubric-guided reinforcement learning to improve edit completion, local edit quality, and preservation of unedited regions. We evaluate VibeEdit on an independently constructed, human-curated benchmark of 419 cases emphasizing target selection among similar objects. VibeEdit achieves a VLM rubric score of 79.9 and an outside-region PSNR of 32.8 dB, compared with 67.4 and 24.0 dB for FireRed, the highest-scoring text-instructed baseline in our evaluation.
cs.AI / 110 / 2610.12230
From Prompting to Composing: A Spatial Canvas Interface for Poster Generation
Yitong Wang, Fangyun Wei, Jinjing Zhao, Sirui Zhang, Hongyang Zhang, Dong Chen, Bo Dai, Yan Lu
cs.CV · cs.AI
Abstract
Text prompting is an indirect interface for poster generation, requiring users to encode inherently two-dimensional composition intent into a one-dimensional sequence of words. We introduce a Spatial Canvas Interface that enables users to directly compose generation intent in space through four complementary binding types: semantic, identity, text, and pixel, together with Text Specifications for individual elements and global appearance. Based on this interface, we develop Compo, a poster generation model adapted from a pretrained image editing model to understand Spatial Canvas inputs and Text Specifications. Compo supports both direct inference, where users explicitly construct the canvas, and agentic mode, where a high-level request is automatically translated into a planned Spatial Canvas. To train Compo, we develop a scalable pipeline that automatically constructs supervision data for different binding types and their combinations, enabling efficient adaptation without training a specialized poster generator from scratch. We further introduce a benchmark that evaluates adherence to individual binding types and their joint composition. Experiments show that Compo achieves stronger compositional controllability than both general-purpose image generation models and dedicated poster generation systems while maintaining high visual quality. By decoupling intent specification from visual generation, our work shifts poster generation from prompting toward composing.
cs.AI / 111 / 2610.12256
Syn-Omni: Structured Specialization and Progressive Collaboration for Omnimodal Embeddings
Youngtaek Oh, Qiyu Wu, Hiromi Wakaki, Junmo Kim, Yuki Mitsufuji
cs.CV · cs.AI · cs.IR
Abstract
Omnimodal embeddings naturally involve both shared representations and modality-specific features across heterogeneous inputs. However, existing omnimodal embedding methods often rely on a single shared parameter space over mixed-modality data, limiting structural separation between universal and modality-specific representations. To address this, we propose Syn-Omni, a unified framework for structured omnimodal adaptation with modality specialization and controlled cross-modal collaboration. Specifically, we introduce Orthogonal Modality-Expert LoRA (OME-LoRA), which decomposes adaptation into a shared LoRA path for universal semantics and modality-expert LoRA paths for modality-aware specialization. Furthermore, Progressive Synergy Routing (PSR) enables experts to first establish modality-specific priors, then gradually interact with other modality-experts for cross-modal synergy. Evaluated across 81 diverse tasks spanning image, video, audio, and audiovisual modalities, Syn-Omni consistently outperforms omnimodal baselines, demonstrating the effectiveness of structured specialization and cross-modal progressive collaboration.
cs.AI / 112 / 2610.12299
Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction
Dahyun Chung, Siyoon Jin, Hyunwook Choi, Honggyu An, Junyoung Seo, Hyunsung Kim, Seung Wook Kim, Seungryong Kim
cs.CV · cs.AI
Abstract
Egocentric world models predict first-person observations conditioned on an agent's actions, but most focus on a single agent. Real embodied settings often involve multiple agents that act and interact within a shared environment. Existing multi-agent world models rely on coarse actions like locomotion, camera control, or discrete commands, leaving fine-grained embodied interactions underexplored. We formulate multi-agent egocentric world modeling as synchronized ego-stream generation for multiple agents interacting through fine-grained actions in a shared world. This requires cross-view action consistency, shared-environment consistency, and consistent propagation of interaction-induced state updates. We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory. We train and evaluate on real and synthetic multi-agent data and introduce shared-world consistency metrics for environment, update, and identity consistency. Experiments show ME-World improves shared-world consistency, action control, identity preservation, and video quality over existing methods.
cs.AI / 113 / 2610.12333
RiCo: Neural Simulation of Rigid-Body Interactions via Local Contact Reasoning
Ruixiang Ouyang, Guanren Qiao, Fansen Meng, Yueci Deng, Ruixing Jin, Kui Jia, Guiliang Liu
cs.CV · cs.AI · cs.GR · cs.LG
Abstract
Accurate simulation of rigid-body interactions is essential for predictive physical world models. Despite recent progress in modeling object dynamics, capturing how local contacts between surfaces shape object motion remains challenging. While end-to-end world models predict interactions across entire scenes or objects, in practice, rigid-body contact is inherently local, and only nearby surfaces can directly exchange contact forces. Motivated by this observation, we introduce Rigid-body Contact Reasoning (RiCo), which represents interactions between objects through sparse neighborhoods of contact surface points. RiCo combines each point's state with the relative geometry, motion, and physical properties of nearby surfaces, then reasons across the object's points to determine how these local contacts jointly affect its motion. By confining cross-object reasoning to nearby surfaces while propagating contact information within each rigid body, RiCo retains fine-grained interaction details without the cost of modeling every pair of scene points. Such properties enable RiCo a higher accuracy and contact fidelity. Experiments on MOVi-benchmark demonstrate that RiCo reduces 100-frame position and orientation errors by 31-35% and approximately 38%, respectively, compared with baselines. Moreover, RiCo achieves high contact fidelity, with ground-truth-relative penetration-time and mean-depth differences of 11.0% and 2.22 mm, respectively. RiCo further generalizes zero-shot from small-scale training scenarios to scenes containing 270 objects. Our real-world multi-ball collision experiments further provide preliminary evidence of sim-to-real transfer.
cs.AI / 114 / 2610.12337
ContiLNN: Mitigating Slice Sampling Discontinuity with Liquid Neural Networks for Medical Image Restoration
Jialei He, Enhe Liu, Sifan Song, Pengfei Jin, Jionglong Su, Hongbin Wang, Zhixiang Lu, Yanhao Huang, Anteng Cai, Zhengyong Jiang, Jiaman Ding, S. Kevin Zhou, Jinfeng Wang
cs.CV · cs.AI
Abstract
Anatomical continuity provides complementary information for medical image restoration, but its use requires accounting for local anatomy and variations in slice sampling. We introduce ContiLNN, which augments two-dimensional restoration backbones with bidirectional closed-form continuous-time (Bi-CfC) modules for cross-slice modeling while retaining in-plane feature extraction. Slice-index intervals modulate gates determined by local features and hidden states, enabling propagation to respond to sampling variations without numerical ODE integration. Reference-guided consistency aligns first- and second-order cross-slice intensity differences to preserve anatomical variation, while distillation from a frozen backbone helps retain in-plane fidelity. Across five training seeds, ContiLNN improves mean PSNR over Restore-RWKV by 0.1907, 1.0176, and 1.2482 dB for CT denoising, MRI super-resolution, and reduced-count PET restoration, respectively, with lower RMSE in all three tasks. CT results are descriptive for one held-out patient. PET ablations support ordered propagation beyond additional pointwise capacity. Under contiguous training, Bi-CfC achieves higher fidelity than a Bi-GRU with similar parameter counts and arithmetic costs across all tested sampling conditions. Matched seven-slice profiling shows 52.8% lower latency and 57.0% lower peak GPU memory use than Bi-GRU. Mixed-gap training improves sparse and irregular-context performance for both operators, without a uniform ranking across metrics and contexts. Experiments with fewer training patients and a second backbone further support data efficiency and backbone compatibility.
cs.AI / 115 / 2610.12355
Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models
Hongxing Li, Yixin Li, Dingming Li, Zixuan Wang, Yuchen Yan, Wenqi Zhang, Weiming Lu, Yongliang Shen
cs.CV · cs.AI
Abstract
Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence. Existing remedies either inject 3D into the model at inference, paying architecture and latency costs, or train with outcome rewards that supervise only the final answer. Spatial errors originate in perception: a misjudged depth or direction can be corrected only by the scene's true geometry, which the 3D-scanned sources of spatial training corpora already provide. We propose GPD (Geometry-Privileged Distillation), which makes geometric evidence the privilege in on-policy self-distillation (OPSD). For each question, depth, semantic, and bird's-eye-view (BEV) cues are rendered as compact text and routed to the teacher alongside the reference answer; a privileged KL, applied only to incorrect trajectories, augments GRPO, and the deployed model remains RGB-only. On the 4B backbone, GPD achieves 57.1 on VSI-Bench and 37.6 average across MindCube, SPARBench, MMSI-Bench, and ViewSpatial, outperforming both GRPO and answer-privileged OPSD across spatial reasoning benchmarks. Ablations confirm the complementarity of 3D and answer privilege, the advantage of question-conditioned routing over full-context injection, and the benefit of restricting distillation to incorrect trajectories.
cs.AI / 116 / 2610.12399
SpaceFlow: Locally Controllable 3D Generation
Neil De La Fuente, Joan Lafuente, Mukhammadali Sayfiddinov, Felicia Scharitzer, Marc Pollefeys, Ata Celen, Sayan Deb Sarkar, Elisabetta Fedele
cs.CV · cs.AI · cs.GR
Abstract
Current 3D generation methods lack explicit local control: geometric adherence is often defined by a global control strength, and appearance cannot be specified locally. We present SpaceFlow, a training-free pipeline for locally controllable 3D generation from text descriptions and a collection of geometric primitives. Each primitive serves as a proxy for an object part and is assigned a local control level, enabling users to specify whether regions should strictly follow the input shape or allow generative completion. During structure generation, we enforce these spatial constraints within the generative flow process. For appearance synthesis, the generated structure is segmented and matched to the primitives. Each generated part is conditioned only on its assigned text or image cue, thereby limiting cross-part leakage. Regional geometry metrics demonstrate that SpaceFlow preserves the specified geometry in high-control regions and enables plausible shape variation in low-control areas. A user study further indicates that the resulting balance between geometric fidelity and generative freedom remains competitive in overall quality. When evaluating appearance on fixed geometry, text-conditioned routing achieves state-of-the-art prompt faithfulness and color/material accuracy. Qualitative results additionally show localized routing of image cues. The project page is available at SpaceFlow3D.github.io.
cs.AI / 117 / 2610.12416
MAMHOI: Factorizing Scene-Aware Human-Object Interaction through Affordances
Mingyuan Lei, Yoonchang Sung, Tat-Jen Cham
cs.CV · cs.AI
Abstract
Generating realistic human-object interactions (HOI) in complex 3D scenes requires two complementary capabilities: reasoning about interaction feasibility in the environment and synthesizing realistic human-object motion. However, supervision for these capabilities is rarely available jointly at scale. Human-scene datasets provide rich information about environment-aware motion, while human-object datasets capture detailed interaction dynamics, yet paired human-object-scene data remain scarce. We present MAMHOI, an affordance-mediated factorization for scene-aware human-object interaction generation. MAMHOI factorizes scene-aware HOI generation through an explicit motion-affordance interface between scene understanding and motion synthesis: a scene-conditioned model first predicts where and how an interaction can be feasibly executed, and an affordance-conditioned HOI model then generates the corresponding human-object motion. This factorization allows scene understanding and interaction dynamics to be learned from complementary sources of supervision without requiring paired human-object-scene data. Experiments in complex indoor environments show that MAMHOI reduces object--scene penetration while better preserving human--object interaction quality, yielding more realistic and physically feasible scene-aware interactions. Project page: https://leimingyuan.github.io/MAMHOI-project-page/
cs.AI / 118 / 2610.11690
Epistemic Disturbance in the Graph Model for Conflict Resolution: State-Preserving Actions, Four-Valued Assessments, and the Distinction between Capability and Intention
Yukiko Kato
cs.GT · cs.AI · cs.MA
Abstract
In the graph model for conflict resolution (GMCR), a decision maker (DM) either moves the conflict to another state or does nothing. The basic definitions leave inaction implicit, so every action that leaves the state unchanged is treated as doing nothing. Yet announcements, exercises, leaks and selective disclosures are neither moves nor inaction: they leave the state unchanged but change what other DMs believe about which moves are available and which moves others would want to make. We introduce such state-preserving actions by augmenting states with the DMs' epistemic states: a physical move changes the physical state, a state-preserving action changes only the epistemic state, and inaction is the absence of a transition. Actions generate evidence through observer-specific interpretation maps. Building on a four-valued extension of GMCR from the author's earlier work, which separates evidence for and against, we show that evidence for a move can only enable perceived moves and evidence against can only disable them, that two of the four reduction operators ignore one kind of evidence, and that contradictory assessments are absorbing under monotone accumulation. With the monotonicity of stability in move sets, this fixes the direction in which any action moves a DM's stability judgements and characterizes when actions can enable provocation or deterrence. Capability assessments affect all sanction-based stability concepts, and on the DM's own side also Nash stability, whereas intention assessments affect only sequential stability. Hedging between two candidate types weakly expands or shrinks an observer's sequentially stable set according to how it reads contradiction. In the 1995 DVD format negotiation, general metarationality cannot distinguish its phases, since the computer industry group could always sanction; sequential stability, which asks whether it would, can.
cs.AI / 119 / 2610.10770
AI-Mediated Self: How HCI Defines and Relates to the Self
Jenny Xiyu Fu, Qian Yang, Malte Jung
cs.HC · cs.AI
Abstract
How might AI alter how we understand and experience the self? This scoping review analyzes 102 papers to examine how the self is defined in the field of human-computer interaction (HCI), how AI-self relationships are conceptualized, and what risks emerge when AI becomes entangled with selfhood. Our synthesis makes three contributions. First, we define AI-mediated self as a conceptual umbrella that connects dispersed work across education, workplace, health, and creative practices. Second, we consolidate six framings of the self with four domains of ethical risk-agency/autonomy, identity/authorship, relational capacity, and meaning-making-into a conceptual map that provides a reusable vocabulary across contexts. Third, we introduce the Inclusion of AI-Self framework, which situates AI-self relationships along a spectrum of proximity. Together, these contributions position selfhood as a central design space in HCI.
cs.AI / 120 / 2610.11539
Design Creativity Bench: Measuring creativity in LLM-Generated UI
Aman Rusia, Abhijit Bhole, Prashank Gupta, Dipanjan Dey
cs.HC · cs.AI
Abstract
As leading LLMs improve on capability evaluations, their limitations in producing creative outputs on design tasks remain insufficiently characterised. Our work introduces Design Creativity Bench, a benchmark that evaluates diversity and appropriateness in UI designs. It measures distinctiveness among models on the same prompt (originality), how much a model's designs change between two prompts for the same UI goal in different product domains (creative range), and the share of a brief's acceptance criteria each design meets (appropriateness). Originality is 0.592 for same-prompt design pairs from different models (95% CI [0.582, 0.602]), far below the 0.764 for same-prompt human-model pairs (95% CI [0.751, 0.778]). Creative range is 0.581 across models (95% CI [0.567, 0.597]), against 0.902 for human designs (95% CI [0.884, 0.919]). Appropriateness is above 90% for every model, and the best model reaches 99.2%, slightly above the 98.0% for human designs. Our work shows that the default output of LLMs, though generally appropriate, is substantially more repetitive than the human baseline. This calls for strong measures to address the issue.
cs.AI / 121 / 2610.11650
SkillContrast: Difference-Guided Text Selection for Agent Skill Reranking
Jiandong Ding, Honglei Ji, Ming Liu, Tao Duan
cs.IR · cs.AI
Abstract
Similar agent skills can share instructions but differ in their conditions of use. Query-based text selection may retain shared instructions and omit these distinctions. We introduce SkillContrast, a training-free selector that compares retrieved skills and retains their differing text with local context for a pretrained reranker. On 1,235 requests from SameCapRisk-Bench, it yields 54-72 more clean hits (requests that retrieve a helpful skill without its marked risky sibling) than TF-IDF query selection at identical per-candidate input lengths, across 2 retrievers and 2 reranker sizes. Length-matched component replacements identify differing text as the main contributor in the primary setting, with smaller, mixed context effects. Relative to full skill bodies, SkillContrast uses 51.1-58.8% fewer model-input tokens, with 10-18 fewer clean hits at 0.6B and matching or higher observed clean-hit counts at 4B. Candidate-relative differences thus complement query relevance in selecting compact reranking inputs.
cs.AI / 122 / 2610.11792
AuraLuxMuse: Adaptive Fusion Modeling for Aesthetic Stage Lighting Design with Music and Expert Guidance
Junyu Deng, Jiale Cao, Mengtian Li, Zhongxia Ji, Ruhua Chen, Yiyi He, Guangnan Ye, Zuo Hu
cs.MM · cs.AI
Abstract
We present AuraLuxMuse, a novel system for automated aesthetic stage lighting design that integrates expert knowledge, representation learning, and preference-adaptive modeling. Lighting design in live performance settings requires the seamless translation of musical features into dynamic lighting behaviors. However, traditional workflows remain time-consuming, labor-intensive, and difficult to transfer. AuraLuxMuse encodes music and professional cue sequences into a shared retrieval space, estimates cue-event density, and retargets selected fixture commands to the destination stage. It assists pre-production authoring by returning editable cues rather than replacing the designer with an unconstrained generator. At the heart of AuraLuxMuse are two key modules: Lighting-Aligned Music Pretraining (LAMP), which performs contrastive learning between audio and lighting cues for alignment, and Preference-Adaptive Mixture of Experts (PAMoE), which conditions preference-aware cue retrieval and adaptation on designers' intent through a gated ensemble of style-specific expert networks. To support training and evaluation, we introduce Musilux, the first dataset of paired musical audio and professional lighting cue sequences under diverse performance scenarios. We evaluate AuraLuxMuse across both virtual simulation environments and professional-grade laboratories. Experimental results, including objective and subjective evaluation, demonstrate that AuraLuxMuse retrieves and adapts stage-lighting cues that are visually cohesive, semantically meaningful, and artistically expressive, showing its potential for AI-assisted aesthetic stage design.
cs.AI / 123 / 2610.10787
NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime
Gengze Zhou, Yicong Hong, Jiazhao Zhang, Xunyi Zhao, Jian Zhou, Zixing Lei, Zun Wang, Chongyang Zhao, Xionghui Chen, Stephen Gould, Anton van den Hengel, Qi Wu
cs.RO · cs.AI · cs.CL · cs.CV · cs.LG
Abstract
Language models trained with long-horizon agentic reinforcement learning can generalize knowledge through reasoning, express precise actions, and pursue goals over many steps, raising the ceiling on what an embodied agent can understand and decide. Physical interaction, however, remains the domain of action policies, which provide dense, low-latency control. We present NavGPT-3, a harness that connects the two models, with an OS-like runtime built above it: reasoning, acting, and monitoring run as threads with their own context, tools, and permissions, while the runtime schedules them and decides which thread controls the robot's motion, so that the robot can react to sudden real-world events through interruption and thread switching. Beneath it, our action policy NavGPT VLA, trained on 19.28M examples, allocates visual tokens using codec allocation, in proportion to scene change; its 8B model alone reaches 74.51 SR on R2R-CE and leads RxR-CE with 78.19 SR. With the complete harness, NavGPT-3 sets the state of the art on R2R-CE (81.51 SR) and, for the first time, brings an autonomous agent to human level: on RxR-CE it matches human followers in success (90.43 vs. 90.4 SR) and path fidelity (78.47 vs. 77.7 nDTW) at 1 min 22 s per episode, versus roughly 3 min for a human. We comprehensively ablate the harness design and the interaction between the two models, showing how tools and the action policy shape the path from language-model reasoning to physical control: when NavGPT VLA executes the route, the reasoning loop shortens and the system's minimum reaction time falls from 3-19 s per language-model decision to 0.5-1 s per action-policy step (1-2 Hz). These results show that designing this embodied interface is central to connecting frontier language-model intelligence with low-level physical control. We will release all models, code, and evaluation records.
cs.AI / 124 / 2610.11168
PMTRM: Pseudo-Memory Temporal Re-encoding Module for Embodied Policy Learning
Changchuan Yang, Haoxuan Xu, Wenbo Chen, Shuai Ren, Jianlong Zheng, Huarui Zhang, Tianfu Li, Guanzhong Tian
cs.RO · cs.AI
Abstract
Robotic manipulation often contains repeated motions whose local observations look similar at different phases. When these phases require different actions, a policy that relies mainly on the current observation may repeat completed motions or switch phases at the wrong time. To address this phase ambiguity, we present the Pseudo-Memory Temporal Re-encoding Module (PMTRM), a lightweight plug-in module with only 7.61M parameters that encodes a bounded history of executed states and actions into a latent sequence for existing policies. To help distinguish phases, a temporal heterogeneity objective penalizes positive similarity between distant positions in this sequence, while anchor and reconstruction losses preserve information needed for action prediction. The reconstruction decoder is used only during training, leaving the temporal re-encoder to supply history to the policy at inference. We train the module progressively on synthetic sequences and robot data, then jointly with the policy, using temporal masking to accommodate partial histories. This integration retains the original action head and action space and adds auxiliary losses to the original policy loss. Experiments with multiple policy backbones in simulation and on a real robot show improved task success on tasks with phase ambiguity, with little additional computation.
cs.AI / 125 / 2610.11175
Higher-Order Action Supervision Makes A Strong Policy Class
Peng Cheng, Yunxian Hou, Zhi Zhou, Qian Zhang, Chang Huang, Xianyuan Zhan
cs.RO · cs.AI
Abstract
Modern data-driven decision-making methods, such as imitation learning (IL) and reinforcement learning (RL), have achieved great success in solving many complex tasks. However, these methods often suffer from serious control instability and robustness issues when applied in real-world applications such as robotics and autonomous driving, posing notable challenges for their practical deployment. We argue that this instability issue stems largely from their limitations in solely supervising and optimizing zeroth-order actions (i.e., the action labels), failing to account for higher-order action dynamics and temporal consistency. In this paper, we show that simultaneously supervising both zeroth- and first-order actions can dramatically enhance policies' performance and control robustness. To achieve this, we introduce a novel and elegant loss scheme supported by formal theoretical guarantees that can equip any off-the-shelf policy model (e.g., deterministic, stochastic, or flow policies) with the capability for higher-order action supervision, without requiring any structural modifications. Moreover, our proposed method can serve as a lightweight plug-and-play module that seamlessly integrates with a broad spectrum of existing offline RL frameworks. Extensive evaluations on OGBench and D4RL demonstrate that our approach yields substantial performance and robustness improvements across a wide range of continuous control environments. Notably, our method can also enhance policies' out-of-distribution (OOD) generalization capability in the challenging low-data regime, making it an ideal tool in tackling many real-world control problems.
cs.AI / 126 / 2610.11223
SafeInferCom: Safe Inference-Time Compute via Verifier-Guided Mid-Generation Intervention for Robotic Task Planning
Weizhe Xu, Jialiang Fan, Mengyu Liu, Fanxin Kong
cs.RO · cs.AI · cs.CL · cs.LO
Abstract
Large Reasoning Language Models (LRLMs) enable multi-step reasoning for robotic task planning, but continued reasoning can overwrite valid intermediate plans or leave constraint violations unresolved, reducing planning reliability and wasting inference-time computation. We develop an inference-time monitor that exposes and verifies intermediate plans without disrupting the original decoding trajectory. Building on this monitor, we propose SafeInferCom, a formal verifier-guided framework that preserves valid intermediate plans and directs error correction during generation. Experiments across multiple LRLMs and planning domains reveal reasoning-response inconsistency and limited self-correction under one-shot inference. SafeInferCom improves planning success and accelerates error correction relative to one-shot inference. When combined with iterative refinement, it further improves success while reducing token usage compared with refinement alone. We additionally evaluate SafeInferCom in VirtualHome and provide a real-world robotic-arm demonstration.
cs.AI / 127 / 2610.11416
Rewiring Semantics, Dynamics, and Control: A Simple yet Effective Action-Centric Tri-Stream Transformer
Shuang Luo, Yilun Kong, Yunpeng Qing, Yihang Jiao, Zhi Hou, Shunyu Liu, Xiaogang Wang, Dacheng Tao
cs.RO · cs.AI
Abstract
Vision-Language-Action (VLA) models have emerged as a prominent framework for complex robotic manipulation, building on the strong semantic understanding of pretrained Vision-Language Models (VLMs). However, such VLM backbones offer insufficient physical dynamics priors, which limits the generalization capabilities of robot policies. Recent efforts therefore integrate video-generation World Models (WMs) into robot policies through various strategies, using predictive dynamics to facilitate action generation. Despite these advances, harnessing semantic understanding and dynamics prediction as complementary guidance for action generation remains challenging. In this paper, we introduce $\mathrm{ACT}^3$, a simple yet effective Action-Centric Tri-Stream Transformer that fuses semantic and dynamics information into control actions while preserving the distinct roles of context streams. Specifically, $\mathrm{ACT}^3$ enables the dedicated action expert to access VLM and WM representations through layerwise attention, with each backbone attending only within its own stream. This straightforward interaction design maintains independent forward propagation in the context streams while allowing both backbones to be updated through control supervision. Experiments on both simulated and real-world robotic manipulation benchmarks show that the proposed $\mathrm{ACT}^3$ yields results superior to its counterparts.
cs.AI / 128 / 2610.11480
RoboAware: Learning to Coordinate Embodied Skills from Counterfactual Outcomes
Bohan Zhou, Xingbei Chen, Emily Huang, Weilin Ruan, Haojian Huang, Yehang Zhang, Zexi Li, Wenqian Li, Qize Yu, Zetian Song, Leyi Wu, Jinghao Li, Mingxuan Song, Xinrun Xu, Zongyang Qiu, Yangkai Wei, Tianyi Zhang, Kaiwen Zhou, Yinchuan Li, James Cheng
cs.RO · cs.AI
Abstract
Embodied coding agents can combine modular robot skills with frozen end-to-end policies, yet effective composition requires anticipating which policy family will succeed in the current physical state. We present RoboAware, which builds on coding agents' skill orchestration by learning only a state-conditioned responsibility coordinator from counterfactual outcomes. Inspired by the success of REPL, we propose the $P^5$ schema and formulate a hierarchical MDP based on it. $P^5$ organizes skills uniformly into five semantic stages, defining where responsibility can be compared. To address the lack of counterfactual branch outcomes in existing work, we introduce State-Locked Counterfactual Branching (SCB), which restores the same training state to generate and execute a code block from each admissible family, exposing outcomes that selected-branch experience leaves unobserved. Building on this, we propose Execution-Aware Learning (EAL), which combines Monte Carlo tree search with Q-learning to distill these outcomes into family-conditioned values. At deployment, the coordinator selects the policy family according to observable context, and the frozen coding agent generates the next local code block. Comprehensive single-episode evaluations on 100 tasks show that RoboAware reaches a 77.0% overall success rate, with SOTA averages of 90.0% on RoboSuite, 73.8% on diverse LIBERO-Pro task clusters, and 90.0% on challenging RoboTwin bimanual tasks, outperforming existing code-as-policy and VLA-harness baselines.
cs.AI / 129 / 2610.11583
SDPAD: A Fully Spike-Driven Pipeline for End-to-End Autonomous Driving
Chengjun Zhang, Yuhao Zhang, Jie Yang, Mohamad Sawan
cs.RO · cs.AI
Abstract
End-to-end autonomous driving demands trajectory planners that are both highly accurate and cheap enough for edge deployment. State-of-the-art artificial neural network (ANN) planners meet the accuracy requirement at the cost of heavy dense computation, while spiking neural networks (SNNs)---though promising orders-of-magnitude energy savings through sparse, event-driven arithmetic---still lag far behind in planning accuracy. We present \textbf{SDPAD}, a fully spike-driven end-to-end planning pipeline that closes this gap. SDPAD converts a pre-trained ANN perception stack into integer-spike form via quantized ANN2SNN conversion, lifts multi-view images into the bird's-eye-view (BEV) space with a spike-driven-max (SDM) depth distribution (Spike-3D-Lift), and plans through the Spike-QFormer, a spiking query transformer in which ego, agent, and map queries distilled from the BEV scene are fused by learnable waypoint queries via cross-attention, followed by deformable spike-cross-attention refinement. Every operation is gated by integer spikes and inference is a single feed-forward pass without temporal simulation loops. On the nuScenes open-loop benchmark, SDPAD achieves an average $L_2$ error of 0.40\,m and a collision rate of 0.12\%, on par with strong ANN planners while consuming 69.9\,mJ---less than 2\% of recent ANN baselines. In closed-loop evaluation on the NAVSIM navtest split, SDPAD reaches 86.3 PDMS, surpassing the previous SNN planner SAD by 4.3 points and matching mainstream ANN planners at a fraction of their energy. To our knowledge, SDPAD is the first fully spike-driven planner evaluated in end-to-end autonomous driving, demonstrating that SNNs can rival dense ANNs in complex driving tasks.
cs.AI / 130 / 2610.11631
Neural Networks for Temporal Pattern Recognition and Dynamic Arm Gesture Speed Estimation for Robot Control
Milán Zsolt Bagladi, László Gulyás
cs.RO · cs.AI · cs.CV
Abstract
Deploying intelligent robotic systems that interact with humans through gestures requires neural networks capable of recognizing diverse temporal patterns. We present a systematic benchmark of ten abstract sequential tasks--five permutation-invariant (set) and five order-dependent (sequence) problems--evaluated across eighteen neural network architectures spanning recurrent, convolutional, attention-based, and set-function families. Beyond the core architecture-task grid, we explore numerous preprocessing and target-variable transformations, yielding more than 250 distinct experimental configurations. All variants are trained and tested under strictly identical conditions (fixed random seeds, shared hyperparameters, shared data splits) to ensure fair and reproducible comparison. Ranking across all ten tasks reveals four consistently top-performing architectures--BiGRU, TCN, Conv1D, and GRUReLU--all compact enough for real-time deployment (under 2,000 parameters in the benchmark setting). Based on this ranking, we apply three architecturally diverse top models (BiGRU, TCN, and GRUReLU) to a practical robotics problem: estimating the execution speed of dynamic arm gestures from skeletal keypoint sequences. Three speed interpretations (peak count, period time, and mean spike spacing) are evaluated on a custom dataset of eight traffic-related gesture classes comprising 256,710 frames recorded via OpenPose. The best configuration achieves a mean absolute error of 0.198 on the peak-count interpretation, corresponding to roughly 5% relative error, while the period-time interpretation reaches approximately 4% relative error, and the mean spike spacing interpretation approximately 8% relative error. These results demonstrate that neural networks can reliably estimate gesture speed from skeletal data, opening a path toward speed-aware gesture-controlled robotic systems.
cs.AI / 131 / 2610.12007
REACT: Rolling Denoising and Dual Decoupling for Reactive Robot Control with VLA Models
Houlong Xiong, Zhenqi Qiu, Zechen Wang, Suohang Zhang, Yiyu Ren, Wanting Xu, Hongfei Niu, Chengyang He, Ge Sun, Ran Cheng, Qian Zhu
cs.RO · cs.AI
Abstract
Flow-based vision-language-action (VLA) models generate action chunks for temporally coherent robot motion, but chunked control creates a fundamental closed-loop trade-off: long chunks provide smooth execution, whereas frequent replanning improves reactivity at the cost of action discontinuities. We introduce REACT, a rolling-denoising framework that makes flow-based VLAs more reactive while preserving long-horizon context. Instead of regenerating entire action chunks from scratch, REACT maintains a persistent action buffer with staggered flow timesteps. At each control step, the full horizon is denoised using the latest observation, the cleanest action block is executed, partially refined future blocks are shifted forward, and fresh noise is appended to the tail. As a result, each executed action block is refined across multiple recent observations before deployment. To support real-time control, we further introduce dual decoupling, which separates sensing, VLM encoding, DiT denoising, and action execution, enabling high-frequency observation updates and action streaming under practical compute constraints. Across the RoboTwin 2.0 simulation benchmark and real-world tasks spanning bimanual manipulation and dynamic control on multiple robot platforms, REACT improves task success and reduces reaction latency while producing smoother trajectories than frequent-replanning and asynchronous baselines.
cs.AI / 132 / 2610.12026
Humanoid World Action Model With Joint State--Action Generation
Yan Yang, Jikun Rong, Minzhao Zhu, Zheyi Zhao, Qirui Hu, Zihan Lan, Weixin Mao, Yinhao Li, Zhen Fu, Hua Chen
cs.RO · cs.AI
Abstract
Humanoid robots are a promising platform for general-purpose manipulation. Recent Vision-Language-Action (VLA) policies learn actions directly from multimodal observations, while World Action Models (WAMs) further incorporate future visual prediction to improve action generation. However, in hierarchical humanoid systems, VLA and WAM policies output reference actions that are subsequently realized through whole-body control, robot dynamics, balance, and contact. This hierarchy creates an action--execution gap: the reference produced by the policy can differ from the motion realized by the robot. Without explicitly modeling the realized body state, future visual prediction must jointly explain scene evolution and discrepancies between reference actions and executed motion, making it difficult to associate an action with its physical outcome. We propose HWAM, a Humanoid World Action Model with joint state--action generation, which makes the robot's post-execution proprioceptive state an explicit prediction target. By jointly generating reference actions and their realized body states, HWAM directly incorporates supervision of executed motion into action learning. HWAM is trained through three complementary conditional paths. The Policy path jointly denoises state--action trajectories conditioned only on current observations, matching deployment conditions. Forward Dynamics Modeling (FDM) predicts future visual observations conditioned on actions and post-execution states, while Inverse Dynamics Modeling (IDM) reconstructs the joint trajectory from visual transitions. Together, these paths connect policy references, realized body motion, and visual outcomes. HWAM achieves the highest success rate among evaluated baselines on three real-robot tasks on the LimX OLI humanoid. On Candy Picking, HWAM achieves a 70.6% success rate, compared with 43.3% for Fast-WAM.
cs.AI / 133 / 2610.12033
Traceable World State: A Provenance-Aware State Representation and Deterministic Replay Framework for Robotic Systems
Zoe Li
cs.RO · cs.AI · cs.SE
Abstract
Robotic systems operating over extended tasks must maintain a world state assembled from observations arriving at different times, with varying confidence and potential revisions. Conventional representations emphasize latest estimates, hindering fact provenance, decision reproduction, or execution auditing. We present Traceable World State (TWS), a middleware-neutral semantic representation and reference runtime for provenance-aware robot world state. A TWS snapshot captures entities, relations, observations, confidence, and revision metadata. Validated update operations transform snapshots immutably, ordered updates support deterministic replay, and a canonical SHA-256 hash chain ensures tamper-evident logs. We evaluate TWS through schema conformance, complete state lifecycles, deterministic replay, and fault injection. Passing 38 tests across Python 3.10-3.14, the framework detects record corruptions, broken hash links, sequence discontinuities, and world mismatches. Across ten public BEHAVIOR-1K task definitions, TWS imported 153 entities and 146 relations with successful validation. On 103 NVIDIA Unitree G1 simulated trajectories containing 78,369 frames, TWS achieved exact terminal-state replay in all episodes and detected 412/412 injected corruptions with a 1.72% storage overhead over Plain JSONL.
cs.AI / 134 / 2610.12172
Unifying Policy Learning and State Prediction through Spatial Language Modeling
Minye Wu, Zehao Wang, Tinne Tuytelaars
cs.RO · cs.AI
Abstract
Learning how actions change scene geometry can provide complementary supervision for goal-directed manipulation. We introduce Spatial Language Modeling, which represents scene contours, goals, action targets, and future states with a shared vocabulary of discrete coordinates and semantic tokens. A task-specific grammar organizes these elements into spatial sequences, allowing one autoregressive Transformer to learn action generation and action-conditioned state prediction through a common next-token objective. We train the model from scratch using random-play transition pretraining followed by joint action and state training on expert demonstrations. During pretraining, recorded action coordinates condition subsequent state predictions and are excluded from the prediction loss. During control, the model decodes only executable action targets and updates its history with newly observed states. We evaluate the approach on Push-T in simulation and on a real robot. The model achieves competitive simulation performance and higher task success and target coverage than the evaluated real-robot policy baselines. Training ablations show improved control with joint action and state sequences, with further gains from random-play pretraining. Given supplied action trajectories, the same model also predicts successive scene states, capturing the geometric effects of pushing.
cs.AI / 135 / 2610.12249
Real-Time Motion Planning with Dynamic Hazards: Classical vs. Learning-Based Methods
Eran Iceland, Alexander Tuisov, Oren Gal, Ariel Barel, Alfred M. Bruckstein
cs.RO · cs.AI
Abstract
We study real-time motion planning in dynamic hazard fields through a controlled comparison between classical planning and learning-based methods. Rather than introducing a new planner, we construct a unified benchmark in which representative classical and learning-based methods face the same environments, motion constraints, information assumptions, and evaluation metrics. The test environment consists of planar domains populated with rotating sprinkler-like hazards that generate time-varying forbidden regions via sweeping angular sectors. Our results show a clear regime shift. In deterministic environments, classical planners achieve near-perfect success and higher-quality paths, though sometimes at the cost of substantial planning or replanning time. Under stochastic obstacle dynamics, however, online search becomes strongly budget-sensitive: low budgets lead to frequent failure, while high budgets improve success at the cost of latency and longer trajectories. PPO-based policies, trained under the same scenario distribution, consistently outperform in latency, success rate, and path quality in these stochastic regimes. Overall, the results indicate that uncertainty in obstacle evolution, more than partial observability, is the dominant factor determining which planning paradigm is practically effective for the problem at hand.
cs.AI / 136 / 2610.12386
ARC: A Reasoning Recipe for Robot Foundation Models
Gokul Puthumanaillam, Tao Sun, Elie Aljalbout, Moritz Reuss, Zhaoshuo Li, Fabio Ramos, Ankit Goyal, Jenai Xuning Yang
cs.RO · cs.AI
Abstract
The prevailing approach to improving robot foundation models (RFMs) relies on larger models, more robot demonstrations, and costly training at scale. We show that there exists an effective and efficient complementary approach: the right reasoning recipe can substantially improve the zero-shot task performance of existing state-of-the-art RFMs. We refer to this recipe as ARC. It consists of three key ingredients: a reasoning trace, a scalable automatic labeling pipeline, and a strategy for adapting pretrained RFMs to use these traces for control. First, we find that effective reasoning traces should be grounded in the robot's next action and explain its causal structure: why the action is appropriate and what effect it should produce. Second, we show that these traces can be generated automatically from existing demonstrations, enabling us to construct ARC-Trace-DROID from DROID without collecting new robot data. Third, we show how state-of-the-art VLAs such as $π_{0.5}$ and WAMs such as Cosmos3-Nano-Policy can learn to use these traces for control, with fine-tuning and inference tailored to each model's architecture and capabilities. Using ARC, we obtain gains in zero-shot RFM performance that, to our knowledge, are unprecedented without additional robot demonstrations or foundation-scale training. The adapted models establish a new state of the art on RoboLab-120 and MolmoSpaces, with gains of up to 50 percentage points on RoboLab-Reasoning-50. On real robots, ARC improves $π_{0.5}$'s task success by 82.2 percentage points. Project website: https://arc-robot-reasoning.github.io/
cs.AI / 137 / 2610.12424
RoboRSI: Stable, efficient, and reusable robot self-evolution in complex real-world environments
Zimo Wen, Yijin Chen, Yuxuan Cao, Wendi Chen, Yanwen Zou, Wenye Yu, Fuhang Kuang, Han Xue, Jun Lv, Chuan Wen, Cewu Lu
cs.RO · cs.AI
Abstract
A generalist robot should not only perform diverse tasks but also improve through experience, turning what it learns during execution into capabilities that later tasks can reuse. Robot agents that act through code can already repair programs from execution feedback, yet it remains a central challenge to organize this experience around the task structure that gives it meaning, so that each repair is attributed to the responsible capability, supported by execution evidence, and validated before it is reused. We introduce RoboRSI, a robot self-improvement system built on Top-Down Skill Refinement (TSR). TSR decomposes tasks into compound, atomic, and base skills with scoped responsibilities and explicit input--output contracts, attributes each execution outcome to the responsible branch, and confines revision to that branch. Building upon this structure, a Manager, Planner, Engineer, and Reviewer coordinate planning, execution, diagnosis, and the validated release of new skills, while people steer the process through objectives and corrections; stable skill sequences are further consolidated into reusable compound skills. On a mobile manipulator, RoboRSI develops multi-object household cleanup over 104 rounds. In simulation, it achieves the highest success rate on LIBERO, LIBERO-PRO, LIBERO-Plus, and RoboTwin, exceeding the strongest baseline by 2.7 to 11.0 percentage points.
cs.AI / 138 / 2610.11008
Why LLM Agents Favor Their Group: Stakes, Observed Norms, and Reputation
Yujiao Chen
physics.soc-ph · cs.AI
Abstract
Language-model agents favor their own group because they have watched their members favor each other. The group label alone does little once the decision has a cost; what drives favoritism is observed behavior, and an individual's own record can override it. We test this in small societies with arbitrary group labels, ten rounds of point sharing, and matched one-shot decisions across fifteen OpenAI models and three Claude models, about 4,400 societies and 3.3 million audited model calls. First, the large effect of a bare group label reported in earlier work appears only when giving others points costs the agent nothing; once the agent can keep points for itself, that effect collapses on every model that shows it. Second, under a stake, interaction history becomes the main source of favoritism: the history effect is statistically positive on 13 of 15 models, reaches about 3.5-8 points out of 10 on 11, grows with the number of rounds played, and extends to labeled strangers the agent has never met. Third, with scripted histories, favoritism falls to near zero under an egalitarian norm and reverses when the agent's own group is seen favoring the other side; stronger models side with an individual's record when it conflicts with the group. Group favoritism is thus conformity to observed group behavior, carried to strangers by the label and overridden by individual reputation. The same account predicts responses to betrayal, scandal, and a free offer to change group: public reprimand repairs betrayal better than apology or restitution, allocation punishment remains confined to the offending member, and a formed group cannot be bought but can, on weaker models, be invited away.
cs.AI / 139 / 2610.12236
La-Ribo: RNA Co-Design via Geometry-Latent Flow Matching
Runze Ma, Will Hua, Shuangjia Zheng
q-bio.BM · cs.AI
Abstract
RNA function arises from the coupling of nucleotide sequence and three-dimensional structure, motivating their joint design. Coordinating global folding with nucleotide-level detail remains challenging under limited structural supervision. We introduce La-Ribo, a generative framework for RNA sequence-structure co-design via geometry-latent flow matching. La-Ribo retains a sparse phosphate-sugar--base scaffold and encodes nucleotide identity and local conformation in residue-wise latents. A shared flow network generates both jointly, and an RNA-specific decoder then reconstructs all heavy atoms. To expand supervision, we construct a quality-controlled corpus of 168,561 RNA structures, integrating experimental data with predictions from three folding models, including 10,631 MSA-supported structures generated in this work. La-Ribo improves designability and codesignability over the evaluated baselines across sampling budgets and two refolding models, and the same prior supports scaffold-conditioned inverse folding without additional training.
cs.AI / 140 / 2610.11923
Neural Decoding as Cognitive Inference
Yi Guo, Changhong Jing, Yong Hu, Yan Liu, Michael K. P. Ng, Shanshan Wang, Shuqiang Wang
q-bio.NC · cs.AI
Abstract
The brain maintains stable cognition despite continuously changing neural activity. How to extract stable cognitive states from variable neural observations remains a central problem in neural decoding. Existing neural decoding methods map neural observations to predefined external labels based on the stimulus-response principle, often capturing recording-specific spurious correlations. Inspired by how the brain infers the world, and specifically by Bayesian brain theory, we recast neural decoding as cognitive inference constrained by brain-intrinsic priors, yielding high-level meta-neural semantic representations. In decoding experiments spanning five neural recording modalities and three cognitive domains (motor, perception and internal mentation), our cognitive inference method reorganized the geometry of neural observation representations, yielding meta-neural semantic representations that exhibited consistent geometric relationships across cognitive tasks and enabled the recovery of stable cognitive states from variable neural observations. Our work provides an account of how the brain maintains relatively stable cognition despite continual changes in the external environment. Cognitive stability is sustained through cognitive inference from changing neural activity, without requiring fixed neural activity patterns.
cs.AI / 141 / 2610.11694
Elucidating the Space of Enzymatic Reaction: A Unified Benchmark and Pretrained Model
Yutong Hu, Tianming Huang, Yanbo Zhao, Qiongyu Zhang, Shixiang Tang, Lei Bai, Ziyi Zhou, Liang Hong, Pan Tan
q-bio.QM · cs.AI
Abstract
Existing reaction models primarily learn molecular transformations, whereas enzy- matic reactions depend jointly on molecular structure and catalytic function. We formulate this problem as learning an enzymatic reaction space linking reactants, products, and Enzyme Commission (EC) annotations. To characterize this space, we introduce VenusRX-Bench, a unified benchmark for forward reaction prediction, single-step retrosynthesis, and EC-number prediction. VenusRX-Bench integrates reactions from multiple biochemical databases with standardized curation, leakage- controlled splits, and consistent evaluation. Benchmarking representative chemical and enzymatic models reveals a clear chemical-to-enzymatic domain gap, driven by limited domain data, catalytic-context dependency, and the difficulty of modeling large biomolecular structures. To bridge this gap, we develop VenusRX, a unified T5-style sequence-to-sequence model for enzymatic reactions. VenusRX jointly learns forward prediction, ret- rosynthesis, and reaction reconstruction, with two-stage training on millions of template-expanded reactions followed by real biochemical reactions. In addition, optional EC conditioning incorporates catalytic context, while Molecule Library- Constrained Decoding improves the generation of complex biomolecules. Across benchmark tasks and challenging generalization splits, VenusRX achieves the best or competitive performance on most evaluated settings over representative chem- ical and enzymatic baselines. Moreover, EC information consistently improves reaction prediction, while learned reaction representations support accurate EC prediction, revealing a bidirectional relationship between reaction structure and catalytic function. Together, VenusRX-Bench and VenusRX provide a unified framework for elucidating and modeling enzymatic reaction space
cs.AI / 142 / 2610.12332
Prediction-Powered Data Fusion for Treatment Effect Estimation
Yonghan Jung, Shu Yang
stat.ML · cs.AI · cs.LG · stat.ME
Abstract
Randomized controlled trials (RCTs) identify treatment effects without confounding but are often small, whereas observational studies (OBS) are large but may be confounded. Many estimators combining a small RCT with a large OBS have been developed for the average treatment effect (ATE) and the conditional ATE (CATE). However, existing ATE estimators either make assumptions on the OBS or do not borrow enough power from them. The CATE has been studied less than the ATE. Existing CATE methods either assume the OBS are unconfounded, rely on a model of the confounding function, or accept bias in exchange for lower variance. We therefore propose a framework that, without special assumptions on the OBS, fuses the OBS and the RCT by preserving the unbiasedness of RCT-based estimation while borrowing power from the large OBS to boost precision. Applying this principle, we build an ATE estimator, AIPW-Fusion, with closed-form weights and confidence intervals, and two CATE learners, DR-Fusion and R-Fusion. Experiments corroborate our findings.
机器学习 (cs.LG)
195
cs.LG / 1 / 2610.12132
A structure-preserving neural density functional for the ions of a polymer electrolyte
Liyao Lyu
cond-mat.soft · cs.LG · math.NA
Abstract
Predicting the structure and response of inhomogeneous polymer electrolytes requires a description of ion correlations that retains molecular-scale accuracy while remaining transferable across spatial scales and geometries. We develop a neural density functional for electrolytes that preserves spatial symmetries, thermodynamic integrability and the Noether identities, with perfect screening recovered in stable, noncritical bulk states. Its nonlinear density dependence captures the concentration-dependent correlations missed by a pair closure, including a crossover from enhanced to suppressed long-wavelength number fluctuations at strong coupling. The functional describes density profiles at an untrained salt concentration and predicts bulk structure factors and the long-wavelength number response. Trained solely on planar density and internal-force profiles from molecular dynamics, the functional predicts ionic structure in larger domains and in two-dimensional external fields. On the same ion data, it is more accurate than three other neural density-functional architectures and keeps its accuracy with a quarter of the training runs, where the errors of the best alternative grow by about two thirds. The spatial transferability provides a necessary foundation for connecting molecular correlations to continuum predictions at larger scales.
cs.LG / 2 / 2610.10768
Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
Yasaman Cheraghi, Reidar B. Bratvold, Aojie Hong, Ressi B. Muhammad, Sergey Alyaev
cs.CE · cs.LG · eess.SY
Abstract
The global challenge of climate change has driven significant steps to reduce CO2 emissions, guided by international agreements like the Paris Agreement of 2015. Acting too slowly could result in future losses and reputational damage, while moving too quickly could jeopardize shareholder value due to the marginal profitability or potential losses due to technology immaturity of many renewable projects. To navigate this complex transition, energy companies must adopt Sequential Decision Making (SDM) strategies to maximize value creation from decision flexibility under uncertainties. To support this, we developed a custom simulation environment to model the dynamic energy landscape up to 2050. Building on this, we designed a multi-criteria SDM framework that explores various decision strategies related to different portfolios for allocating funds across three sectors: oil & gas, renewables, and CO2 reduction. It aims to maximize value during the transition while accounting for uncertainties in productions, energy prices, and costs. This framework has three objectives: maximizing profit, minimizing CO2 social costs, and enhancing competitive advantage in the renewable energy sector. This research evaluates the use of Reinforcement Learning (RL) to identify optimal investment policies within the defined SDM framework. The agent's sequential decisions shape a virtual dynamic environment by influencing key variables such as oil and gas production, renewable energy output, CO2 emissions, and revenues. Through repeated interaction, the RL algorithm explores the state space and learns an optimal policy under uncertainty. We benchmark the RL strategy against a set of manually defined baseline policies and find it consistently outperforms them in adaptability and long-term value creation.
cs.LG / 3 / 2610.10862
A mesh-based neural energy method for the simulation of heterogeneous composites
Pius J. M. Wichmann, Stefan Hildebrand, Sandra Klinge
cs.CE · cs.LG
Abstract
Modeling heterogeneous materials remains a challenge for physics-informed neural networks such as the deep energy method (DEM). The DEM and its variants, here collectively referred to as the neural energy method (NEM), offer a differentiable variational framework. However, their conventional collocation-based implementation (C-NEM) often suffers from physically inadmissible displacement oscillations, integration errors, and high computational costs from automatic differentiation. This work introduces the mesh-based neural energy method (M-NEM), extending the NEM through a mesh-based discretization of the displacement field. By interpolating nodal displacements via shape functions, the M-NEM imposes a kinematic constraint that suppresses oscillations. Furthermore, the method replaces automatic differentiation with algebraic shape function derivatives for strain computation and employs high-order Gaussian quadrature for accurate energy integration. On a directly comparable benchmark problem, the M-NEM reduces stress errors by up to three orders of magnitude relative to the C-NEM while being one to two orders of magnitude faster. On two further benchmarks involving extreme stiffness contrasts, only the M-NEM converges. A comparative study of neural architectures reveals that radial basis function neural networks (RBFNNs) yield optimal performance within the M-NEM, resolving sharp gradients at material interfaces with higher accuracy than multi-layer perceptrons (MLPs) with random Fourier feature (RFF) mapping and faster convergence than Kolmogorov-Arnold networks (KANs).
cs.LG / 4 / 2610.10945
GHARP: Real-time Gaussian Head Animation from Large-scale Reconstruction Prior
Ali Benlalah, Sepehr Johari, Patricia Vitoria, Armin Kappeler, Artem Sevastopolsky, Alexander Jung, Gabriele Fanelli, Kevin Mader, Manuel Breitenstein, Claudia Plüss, Jan Rüegg, Simon Biland, Thomas Etterlin, Dmitry Kostiaev, Mathias Deschler, Brian Amberg, Sebastian Martin
cs.CV · cs.LG
Abstract
We present GHARP (Real-time Gaussian Head Animation from Large-scale Reconstruction Prior), a method that animates 3D human heads in real time from a few input images of a subject and a driving expression signal. We decouple the problem into an identity stage that builds a representation of the subject's geometry and appearance offline, and an animation stage that predicts expression-dependent residuals on top of it at runtime. This separation offers a favorable trade-off with respect to fidelity, quality and runtime: the identity stage can be expensive while the animation stage runs a lightweight network, optimized for mobile devices. Our method performs animation in a semantically structured latent space of a pretrained reconstruction model, where expression changes remain spatially contained, making residual prediction efficient. This reconstruction prior provides a consistent spatial layout, allowing fusion of multiple input views into a compact, fixed-size canonical Gaussian representation. While this two-stage design improves the runtime-quality trade-off, it still inherits a problem common to all expression-driven avatar methods: expression codes describe only the face and thus omit body pose and clothing position, making these regions underspecified in the input. The animation network faces an ill-posed mapping and resorts to averaging over conflicting body appearances, producing blur and temporal flicker. We address this with a body alignment network that learns to align the person's body in the target image with the input reference images, removing the ambiguity from the training signal. Our method achieves state-of-the-art quality on the Ava-256 benchmark while running up to 13x faster on an A100 GPU with 8x fewer Gaussians.
cs.LG / 5 / 2610.10991
Region-Aware CLS Token Augmentation for Fine-Grained Image Retrieval
Ian de Holanda Cavalcanti Bezerra, Vivek Trivedy, Lucas Pascotti Valem, Longin Jan Latecki
cs.CV · cs.LG
Abstract
Image retrieval methods often rely on a single global semantic descriptor extracted from an image, e.g., the [CLS] token in vision transformers. However, trying to squeeze all the semantic information of an image into a single descriptor can hurt downstream retrieval performance, especially for fine-grained retrieval tasks. In this work, we augment the semantic tokens in the newer visual transformers, the global [CLS] token and the four register tokens, with a carefully selected collection of spatial tokens, aiming to capture the spatial region representation that characterizes the contents captured in each of the semantic tokens. We leverage the DINOv2-reg model, which includes register tokens that emergently learn object and part-based representations. For each "cue" token ([CLS] and each register token), we find a "buddy" image patch token and extract an N x N patch region to produce a set of localized ROI tokens. Our approach automatically captures important regions of interest without any external bounding boxes or saliency modules, purely by matching semantic tokens with their spatial representation regions. Furthermore, we incorporate these tokens into a multi-vector retrieval framework inspired by ColBERT, enabling fine-grained matching via a per-token alignment mechanism while avoiding the large storage cost of keeping all patch embeddings. Through extensive experiments, we find that (1) register tokens encode useful fine-grained details that can complement the [CLS] token; (2) automatically pooled ROI tokens further improve fine-grained discrimination; and (3) multi-vector retrieval with a small set of tokens improves over a DINOv2-reg single-vector baseline while remaining tractable for large-scale search. The code is available at https://github.com/IdhcbIan/Augmenting_CLS_with_ROI_tokens.
cs.LG / 6 / 2610.11057
Learning What to Trust in Multimodal Learning under Noisy Supervision
Jiashuo Zou, Xiaobo Xia
cs.CV · cs.LG
Abstract
Multimodal classification processes and relates information from multiple modalities to achieve more accurate predictions. However, existing methods typically rely on high-quality ground-truth labels, which are difficult to obtain in real-world scenarios. While sample-selection methods for learning with noisy labels aim to identify correctly labeled examples from noisy data, traditional methods primarily focus on unimodal settings and fail to exploit multimodal information fully. This motivates us to build a more reliable noise detector in multimodal learning. To this end, we theoretically analyze the relationship between representation structure and noise detection capability. Based on this analysis, we propose REFINE, which is a multimodal label-noise detection framework that jointly uses fused and unimodal representations for label-noise detection. Specifically, REFINE constructs discriminative eigenvectors through discriminative analysis of the target and background classes and selects trusted representation spaces with better noise detection capability for each class. Within each trusted space, REFINE measures the alignment between each instance representation and the discriminative eigenvectors. It then combines the subsets selected from these spaces. The combined set provides cleaner supervision for updating the multimodal classifier, thereby reducing the influence of mislabeled examples during training and improving model generalization. Extensive experiments across diverse tasks demonstrate REFINE's superiority compared to baseline methods. The source code will be publicly available.
cs.LG / 7 / 2610.11205
Dissecting Representation Structure in Vision Transformers: A Rigorous Architectural Study
Kim-Cuc Nguyen, Ngai-Man Cheung
cs.CV · cs.LG
Abstract
Representation structure is crucial for understanding Vision Transformer (ViT) architectures and their generalization behavior. However, prior studies neither isolate nor analyze module-level features nor investigate how their interactions contribute to performance estimation. In this work, we conduct the first rigorous analysis of feature information across diverse architectural scales, empirically uncover the relationship between ViT representation and generalization behavior, and leverage these insights to guide efficient ViT design. Our contributions are fivefold: Across diverse architectural scales, 1) We identify feature collapse at initialization, which leads to redundancy, and propose a reduction scheme to mitigate this issue. 2) We quantify feature information using entropy and the minimum eigenvalue, demonstrating that these metrics serve as reliable indicators for generalization prediction. 3) We show that feature in the token space provides a more faithful representation than those in embedding space. 4) We discover an unexpected finding: features produced by linear submodules within ViT layers are critical for the prediction of generalization performance. 5) Our proposed proxy improves the correlation ranking by 18-48% over prior baselines and can effectively identify ViT architectures that achieve higher accuracy at lower or comparable computational cost.
cs.LG / 8 / 2610.11251
V-CoLA: Vision Token Compression with Linear Attention
Hao Jiang, Yiru Mao, Tianpeng Bu, Hao Zhou, Hongtao Duan, Wang Jing, Bowen Xu, Xin Chen, Lulu Hu, Bin Yang, Yongliang Tao, Minying Zhang
cs.CV · cs.LG
Abstract
Vision-language models (VLMs) have demonstrated impressive capabilities but suffer from substantial computational overhead, as vision tokens dominate the input sequence. This motivates vision token compression as a key direction to alleviate the burden. However, with the emergence of hybrid architectures incorporating linear attention (\eg, Qwen3.5), prior methods designed for softmax attention struggle to generalize. Our analysis reveals that both attention- and similarity-based approaches suffer notable performance degradation, underscoring the urgent need for compression methods tailored to this regime. To this end, we propose \textbf{V-CoLA}, an efficient training-free token compression framework specifically designed for linear attention. V-CoLA introduces a novel \textit{uniqueness-aware importance criterion} for identifying critical vision tokens, coupled with an \textit{adaptive token merging strategy} that performs compression. All components are optimized at the implementation level to remain compatible with the chunk-wise parallelism of linear attention, ensuring strong practical value. Extensive experiments across multiple benchmarks demonstrate the superiority of V-CoLA: it achieves 99.5\% of the original performance with only 50.0\% of vision tokens, and over 88.0\% with as few as 12.5\%, while delivering a 1.86$\times$ to 6.15$\times$ prefill speedup.
cs.LG / 9 / 2610.11534
HAND: A Biologically-Inspired Activation Function that Improves Generalisation and Sample Efficiency in Image Classification
Michael W. Spratling, Heiko H. Schütt
cs.CV · cs.LG
Abstract
DNNs exhibit robustness and generalisation issues not seen in humans. They are also far less data-efficient learners, requiring considerably more training samples to accurately classify novel exemplars. Inductive bias could help with these issues by providing in-built mechanisms to improve generalisation, and hence, reduce reliance on learning from data. We incorporate a biologically-inspired inductive bias into a new activation function, HAND (Homeostasis, Accelerating Nonlinearity, and Divisive-nomalisation), and show its effectiveness with CNNs trained on image classification. Using HAND a ConvNeXt-tiny required 25 training epochs to reach the same accuracy on ImageNet1k as the unmodified model achieved after 200 epochs. Consistent with the effects of an inductive bias, the performance gap reduced with training time and increased data augmentation. When the volume of training data was reduced and unevenly distributed between classes (Long-tailed ImageNet) the improvements in accuracy were even larger and did not reduce with increased training time. Generalisation performance with the common-corruptions data, and the ability to reject samples from unknown classes, were unaffected or improved by HAND. Results generalised across CNN architectures and training data-sets. HAND can, therefore, reduce the required training time and/or the required volume and variety of training data, helping to improve sample efficiency.
cs.LG / 10 / 2610.11846
Open-Vocabulary Audio-Visual Event Localization via Complex-Valued Fusion
Anirudh Praveen, Koteswar Rao Jerripothula, Pratik Joshi, Aveen Dayal, Neela Sawant
cs.CV · cs.LG · cs.MM · cs.SD · eess.IV
Abstract
Open-Vocabulary Audio-Visual Event Localization (OV-AVEL) labels each video segment with an event class, including classes that were never seen during training. The dominant pipeline uses a frozen multimodal foundation model (e.g. ImageBind) to embed the visual frame, the audio mel-spectrogram, and each candidate class name into a shared space, then computes two cosine similarities for each segment against each class: visual-text and audio-text. Existing methods then collapse this pair into a single scalar score with a fixed rule (geometric mean, weighted average) before taking the argmax. Instead, we compute complex-valued similarities and learn their fusion using a complex-valued neural network (CVNN). Each modality's standard representation becomes the real part of our pipeline, and a paired companion stream supplies the imaginary part. We use imaginary part of iHSV for visual modality and CycleGAN-translated phase spectrogram for audio modality as these companion streams. This results in two complex similarities, which are then fused. While the vision and audio encoders remain frozen, only the temporal-attention blocks and the fusion CVNN are trained. The four-stream complex architecture sets a new state of the art on both OV-AVEL benchmarks. On the open (unseen-class) split of OV-AVEBench we reach 66.5/59.1/54.1% Acc/Seg-F1/Event-F1 (+1.6/+4.1/+6.6 over the previously reported fine-tuned baseline), with consistent gains for seen classes as well. We also modify AVE dataset for this task and observe that our architecture reaches 60.7/51.9/50.4% Acc/Seg-F1/Event-F1, achieving state-of-the-art OV-AVEL results on it as well. We also propose a two-stream alternative, which also sees great improvements over the baseline.
cs.LG / 11 / 2610.11924
Revisiting Identity and Spectra Dispersion in Media-Bridged Time Series Forecasting: Linking Multivariate Signals and Narrative Flows
Jierui Lei, Wenjian Zhang, Qingyi Yang, Yuyang Hong, Fangzheng Chen, Zhengbo Zhang, Haina Tang, Shiming Xiang
cs.CV · cs.LG
Abstract
Media-bridged time series forecasting is expanding to encompass traditional "multivariate" and emerging "multimodal" (e.g., through textual assistance). Existing Time Series Forecasting (TSF) models still rely on paradigm-specific relation, fusion, and temporal modules, hindering a common forecasting backbone across numerical and pre-aligned narrative-flow settings. To explore this, we propose the Multimedia Identity-Aware Prism Network (MIDAPN), a unified spatiotemporal forecasting backbone based on media-general graph adaptation and automatic temporal learning: (1) Following media pre-alignment, our Multimedia Identity-Aware Graph (MIDAG) revisits identity through static essence, dynamic behavior, and latent commonality, inducing affinities that extend variable-specific dependencies across media. Contextual Identity Modulation (CIM) further refines discriminative aggregation. (2) We develop Spectral Prism Convolution (SPConv) to automatically perform hierarchical temporal analysis, balancing coarse trends and fine-grained details. Meanwhile, its Adaptive Search Guidance configures a scale-efficient architecture for temporal-dimension reconstruction. These decoupled yet synergistic components jointly address media identity disentanglement and temporal-scale mismatch. Comprehensive evaluations involving 16 SOTA TSF models across 13 "multivariate" and 12 "multimodal" datasets, alongside targeted long-context comparisons against 14 time series foundation models and fused pretrained language models, demonstrate MIDAPN's consistent superiority and broad shared backbone compatibility. The code is available at \href{https://github.com/leijieruilq/MIDAPN/tree/main}{https://github.com/MIDAPN}.
cs.LG / 12 / 2610.12081
Perception Test 2026: Challenge Summary and Extension to City-scale Audio-Visual Reasoning
Fedor Kitashov, João Carreira, Shiry Ginosar, Dima Damen, Andrew Zisserman, Viorica Pătrăucean
cs.CV · cs.LG
Abstract
Continuing the Perception Test challenge series, we organised the fourth edition as a workshop at the European Conference on Computer Vision (ECCV) 2026 in Malmö, Sweden. This edition focused on spatial intelligence and featured four different tracks: unified multiple-choice videoQA and grounded videoQA from the original Perception Test benchmark, alongside two new tracks based on city-scale walking-tour videos (KilometerAudio and KilometerVision). In this report, we describe the new benchmarks used for the city-scale tracks and summarise the winning solutions across all tracks, including a generalist model that competed across all tracks with satisfactory performance. The winning solutions in the newly added city-scale tracks demonstrated that complex spatial and multimodal reasoning can be solved by expensive agentic pipelines, but remains difficult for multimodal models used standalone.
cs.LG / 13 / 2610.12095
DVLA-RL++: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning
Wenhao Li, Xianjing Meng, Qiangchang Wang, Zhongyi Han, Yilong Yin, Liqiang Nie
cs.CV · cs.LG
Abstract
Few-shot learning aims to recognize novel categories from limited labeled examples. Recent studies incorporate textual semantics to compensate for limited visual observations and improve class representations. However, high image-text agreement may reflect both intrinsic object properties and incidental context, making support prototypes susceptible to contextual contamination. To address this problem, we propose DVLA-RL++, which extends DVLA-RL with complementary semantic purification (CSP) and counterfactual reinforcement-learning gating (CRG). Specifically, CSP generates intrinsic and nuisance descriptions from labeled supports and compares their agreement with each support token. An ambiguity-dependent rejection margin guides sparse evidence allocation, while an intrinsic semantic anchor fills the unassigned mass to provide a fallback when visual evidence is unreliable. CRG learns layer-wise semantic fusion strengths using a reward that balances recognition performance and nuisance exposure. An independently executed reference trajectory on the same episode provides a paired learning signal. Theoretical analysis relates retained evidence and anchor quality to prototype stability and establishes conditions for unbiased on-policy gradient estimation. Experiments on standard, fine-grained, and cross-domain benchmarks show state-of-the-art accuracy, with an average gain of 1.4% over DVLA-RL. The project page is available at https://peacelwh.github.io/TPAMI27-DVLA-RLpp/.
cs.LG / 14 / 2610.12421
Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence Matching
Luping Liu, Bingyi Kang, Yifan Wang, Dong Xu
cs.CV · cs.LG
Abstract
Dense correspondence matching has historically been bounded by simplifying spatio-temporal priors, such as smooth motion and rigid geometry. While effective for classical tasks, these assumptions break down in image editing and reference-guided generation (IEG), where transformations can preserve visual identity while breaking physical continuity. To establish identity-preserving correspondence across such transformations, we introduce FreeMatching, a generalizable framework combining generative and semantic foundation representations with heterogeneous supervision from classical datasets, tracked videos, and synthetic scenes. Teacher-guided iterative refinement further improves correspondence in IEG without dense correspondence annotations. Experimentally, a single FreeMatching model substantially improves correspondence quality on challenging IEG image pairs while retaining competitive performance on classical benchmarks. Furthermore, we demonstrate its utility as a quantitative metric for evaluating identity preservation, with scores that correlate with human judgment. The code is available at https://github.com/luping-liu/FreeMatching.
cs.LG / 15 / 2610.12448
One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
Adrian Bulat, Yassine Ouali, Georgios Tzimiropoulos
cs.CV · cs.LG
Abstract
In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank. A continuous normalized-depth coordinate programs this mixture, defining a resampleable trajectory through FFN parameter space. We evaluate this design in two regimes: supervised ImageNet-1k training and distillation from a DINOv2 teacher. Across both regimes, controlled adaptations identify weight-space merging as the strongest tested MoE family at a matching one-FFN budget, ahead of the token-dispatch and output-mixture alternatives. Trained from scratch, reViT-B/16 attains DeiT III accuracy with about 70\% fewer stored parameters. An 8-experts model distilled using only the teacher's output features retains nearly all of its DINOv2 teacher's linear-probe accuracy and transfers across classification, segmentation, and depth prediction. Elastic-depth training allows one checkpoint (trained model) to operate at multiple tested depths by resampling the same normalized coordinate interval. For fixed-depth deployment, the recurrent block can be materialized as a conventional dense graph, removing online routing and merging without changing the one-FFN-per-depth compute but expanding deployment storage.
cs.LG / 16 / 2610.10919
Implementation Guidelines for Data Quality Metrics
Philipp Jung, Katinka Becker, Felix Biessmann, Valerie Restat, Martin Seyferth, Daniel Schwabe, Lisa Ehrlinger
cs.DB · cs.LG
Abstract
Despite decades of data quality (DQ) research, a gap remains between DQ dimensions, such as accuracy or completeness, which the literature defines in textual form, and DQ tools, which typically implement low-level checks that are not aligned with these dimensions. ISO/IEC 25024 and ISO/IEC 5259 attempt to bridge this gap by defining DQ metrics for each dimension. However, these DQ metrics are hardly used, because the standards leave open how to implement them: for example, the metric for syntactic accuracy counts syntactically accurate values, but does not state how to decide that a value is syntactically accurate. This simply moves the problem to another level without solving it. As a result, DQ assessment currently cannot build on the standards. In this paper, we make the ISO DQ metrics executable. We classify all data-level metrics of both standards into (i) generalizable metrics that need no input beyond the data, (ii) parameterized metrics whose parameters can be learned from clean reference data or set by an expert, and (iii) non-generalizable metrics that need qualitative judgment and cannot be automated. For the metrics that can be automated, i.e., categories (i) and (ii), we propose implementation guidelines that resolve what the standards leave open. We realize the guidelines in dqmeasure, an open-source library of 20 metrics that learns these parameters from reference data instead of relying on manually defined rules. Our experiments on real-world and synthetic datasets show that the metric scores decrease monotonically with an increasing number of injected errors, decline together with downstream ML performance, and scale linearly with the number of rows, which enables automated DQ monitoring based on the standards.
cs.LG / 17 / 2610.11908
Cost-Aware Mixture-of-Experts Coordination for Model Markets
Yizhou Ma, Wenbo Wu, Xikun Jiang, Zhuoqin Yang, Luis-Daniel Ibáñez
cs.DB · cs.LG
Abstract
Existing model marketplaces typically trade and select individual models as indivisible units, limiting their ability to exploit complementarities among heterogeneous experts. This paper proposes an MoE-based model market framework that lifts Mixture-of-Experts from a model-level learning architecture to a market-level coordination mechanism. In this framework, brokers use gating networks to coordinate multiple heterogeneous experts and deliver a composite model service. We formalize the market participants, service workflow, expert cost structure, and a welfare objective that combines predictive utility with heterogeneous execution costs. We then derive a cost-aware gating mechanism and market-aware training objective, and introduce a cost-adjusted revenue allocation rule that distributes residual revenue according to realized expert participation and execution cost. We also establish basic theoretical properties of the allocation rule, including budget balance, participation monotonicity, and cost sensitivity. Experiments over five random seeds on fifteen tabular and image benchmarks use independently trained and frozen neural and tree-based experts together with latency-derived execution costs. MoE Market achieves the highest mean welfare on all fifteen datasets and a lower mean expected cost than Standard MoE in every case, while maintaining competitive predictive performance. The allocation experiments further demonstrate systematic sensitivity to expert participation and cost, together with substantially lower computational overhead than exact Shapley allocation. These results suggest that MoE can serve as a market-level coordination principle for collaborative, cost-aware, and economically grounded model marketplaces.
cs.LG / 18 / 2610.11201
PageWeaver: KV-Guided Query Unions for Sparse Attention
Zhiyuan Li, Zihan Li, Zefang Yuan, Lei Wang, Hao Wang
cs.DC · cs.LG
Abstract
Dynamic sparse attention limits the KV pages selected by each query, but a small support does not necessarily yield efficient GPU work. Query unions share page loads and populate Tensor Core tiles; their cost depends on which queries are grouped together. We present PageWeaver, an execution design that uses selected-page affinity to assemble query groups while preserving each query's original support and complete output ownership. A bounded GPU search produces query IDs, and an ID-aware two-CTA kernel consumes them without materializing reordered Q tensors or cross-page partial outputs. A direct KV-page union implementation provides a complementary design study of nonlocal reuse and reduction cost. With FP8 KV throughout, the H200 Union8 implementation achieves a 1.70x geometric-mean complete-call speedup over the measured FlashInfer path on six captures. Online regrouping further lowers latency by 3.26-7.66% on five selected 64K-context captures. Whole-model prefill throughput is 7.88-14.36% above the tested native path; the incremental regrouping benefit is smaller, with observed median gains of 0.47-0.73% at 32K/64K and regressions at 8K. A B300 comparison identifies cases where preparation cost and a stronger native kernel remove the advantage. These results separate execution-group reuse from the complete cost of exploiting it online.
cs.LG / 19 / 2610.11415
Refinement as a Service: Algorithmic Predictor Refinement
Wei Tang, Hanrui Zhang
cs.GT · cs.LG
Abstract
Prediction aggregation aims to combine information from multiple predictors into a more informative one. We study this question in the setting of calibrated predictors, where each prediction must equal the conditional expectation of the quantity being predicted given the predictor's signal. Given several calibrated input predictors and the feature distribution, but not the underlying Bayes probabilities, we ask when one can construct refined calibrated predictors that preserve the information in the original predictors and cannot be further refined using the available information. We formulate calibrated predictors as signaling schemes and define refinement through feature-independent garblings: a predictor refines another if its signal can simulate the other's signal. Constructibility is characterized through observable linear information: each signal corresponds to a vector over the feature space, and a new signal is constructible exactly when its vector lies in the linear span of the input signal vectors. Under this formulation, we establish a sharp algorithmic picture. For deterministic output predictors, bilateral refinement admits a polynomial-time algorithm based on a bipartite graph between the two input signal partitions, while refinement with an arbitrary number of input predictors is $\mathsf{NP}$-hard. In contrast, when randomized output predictors are allowed, we give a polynomial-time algorithm for any number of input predictors by decomposing constructible signal vectors into extreme rays of the associated polyhedral cone.
cs.LG / 20 / 2610.10690
Learning infinite context windows in recurrent architectures via spatial neural computing
Aleix Salvador-Pomarol, Arthur N. Montanari, Earl K. Miller, Adilson E. Motter, Jorge Cortés
cs.LG · eess.SY · q-bio.NC
Abstract
Recurrent neural networks (RNNs) offer linear-time scaling with sequence length while requiring only constant memory, yet they struggle to capture long-range dependencies due to vanishing gradients and limited receptive fields. To address these limitations, we introduce a second-order recurrent model in which the standard neuron-to-neuron communication is replaced by a spatially evolving field governed by (discretized) partial differential equations. Drawing inspiration from the role of cortical waves in brain computation, this mechanism allows structured spatiotemporal patterns to serve as an implicit, high-capacity memory. We show that the resulting model is equivalent to a structured infinite-order RNN in which the current state depends explicitly on its entire history of past states, yielding an effectively unbounded receptive field with a fixed number of parameters. We further derive constructive conditions to ensure marginal stability, constraining the gradient spectrum on the unit circle and thereby eliminating vanishing and exploding gradients. Empirically, the proposed architecture outperforms other recurrent models on long-horizon benchmarks while using substantially fewer parameters, demonstrating that spatial dynamics can effectively bridge the gap between efficient inference and long-term memory.
cs.LG / 21 / 2610.10740
Conversational Task Disambiguation over Tabular Data: Leakage-Aware Formulation, Benchmark Suite, and Training
Nafiseh Ghoroghchian, Luis Scoccola, Tina Sedaghat, Omid Vaheb, Hannah Chen, Dino D'Agostino, Keyvan Golestan
cs.LG · cs.AI · cs.CL
Abstract
Conversational task disambiguation over tabular data uses dialogue to resolve missing information about a user's intended task before producing a solution over tables or databases. Existing evaluation and training lack a leakage-aware foundation. Task success mixes the agent's disambiguation and solution-generation capabilities and can also reflect oracle leakage, that is, information that a user simulator reveals beyond what a real user would. Existing datasets also lack a shared representation of ambiguities and access boundaries. We introduce the notion of an ambiguous verifiable task, which formalizes ambiguities and resolutions, decomposing the agent into an asking policy and a solution policy, and the environment into an oracle and verifier. This framework provides baselines and metrics for evaluating task disambiguation separately from solution generation, formal definitions of oracle leakage, judge-free leakage diagnostics, and a training objective for the asking policy. We instantiate the framework in text-to-SQL with AmbiTab, a benchmark suite that unifies six ambiguous datasets under a common representation specifying what the agent, oracle, and verifier may access. We evaluate clarification strategies and oracle leakage, and train an asking policy with reinforcement learning. The trained asker improves our disambiguation metrics on all six datasets and task success on five, and our leakage diagnostics measure how training affects oracle leakage.
cs.LG / 22 / 2610.10776
NEMORA: Neural Equivariant Multipole Operators for Long-Range Atomistic Learning
Jay L. Kaplan, Samuel Varner, Rebecca Willett, Juan J. de Pablo
cs.LG · physics.chem-ph · physics.comp-ph
Abstract
Equivariant graph neural networks have emerged as foundational architectures for machine-learned interatomic potentials, approaching quantum-chemical accuracy at a fraction of the computational cost. These models describe local atomic environments accurately, but finite spatial cutoffs truncate long-range information flow, and stacking message-passing layers can lead to over-smoothing and over-squashing. Existing long-range extensions either prescribe a fixed analytical propagation kernel, restrict long-range communication to scalars or degree-preserving channels, are only approximately equivariant, or incur super-linear computational cost. Combining learnable long-range equivariant transport with multiscale many-body expressivity and efficient scaling for larger systems remains a central challenge. We introduce Neural Equivariant Multipole Operators (NEMORA), a neural equivariant extension of the Fast Multipole Method (FMM) for learning long-range tensorial representations. NEMORA generalizes the FMM's analytical multipole expansion and translation operators to learned equivariant counterparts on an adaptive spatial hierarchy. Its operators couple angular degrees and form many-body interactions across length scales, retaining the FMM's hierarchical organization and analytical radial factors as physical inductive biases while learning data-dependent long-range couplings. NEMORA evaluates in linear time and memory complexity, allowing it to treat larger systems than other long-range methods reaching hundreds of thousands of atoms, and it augments both symmetry-constrained and unconstrained short-range backbones. On non-local benchmarks, it reduces force and energy errors relative to the short-range backbones by over an order of magnitude and up to three orders of magnitude, respectively, which is better than or competitive with existing long-range extensions in accuracy.
cs.LG / 23 / 2610.10778
MemoWM: How World Models Change What Agents Need to Remember
Bingfan Zeng, Zhisheng Chen, Chenbo Sang, Zhengwei Xie, Jinpeng Wang, Xiangchen Guan, Rui Qian, Zheng Lu, Jingwei Song
cs.LG · cs.AI
Abstract
Long-term agents face growing storage demands as they accumulate experience. World models capture reusable regularities that can reduce the information stored for each experience. We formulate the problem of memory allocation conditioned on a world model and introduce MemoWM, a framework that uses shared predictions to compress retained information and reconstruct omitted content. Its task-aware allocation rule balances the expected impact of reconstruction errors against storage cost, retaining information with downstream value beyond the predictive prior. Across five long-term agent-memory benchmarks, MemoWM achieves 42.42\% average answer accuracy, exceeding the strongest baseline by 2.62 percentage points, while reducing average experience-specific storage by 53.9\% relative to MIRIX, the most storage-efficient baseline. Further analysis shows that stronger world models reduce per-experience storage at comparable task quality. Accounting for model parameters reveals a trade-off between shared model capacity and recurring storage costs, with the capacity that minimizes total storage increasing as more interactions are retained. Our code is available at https://github.com/Feld-maxiu/MemoWM.
cs.LG / 24 / 2610.10795
How Many Repeated Pairwise Comparisons Are Needed for Ranking under Heterogeneity?
Shashaank Aiyer, Han Shao
cs.LG · cs.IT
Abstract
We study ranking models by population-average utility from pairwise comparisons when preferences vary across users and tasks. Prior work shows that a single comparison per user can be insufficient to identify the alternative with the highest average utility, even with arbitrarily many users (Golz et al., 2025). We investigate how many repeated comparisons within each user-task context are necessary and sufficient for ranking recovery. Under a heterogeneous Bradley-Terry model with fixed inverse temperature, we start with a naive MLE-based algorithm that requires $Ω(1/Δ^2)$ repeated comparisons per context to ensure ranking recovery. We then present two MLE-based variants and a randomized Russian Roulette-style algorithm that recover the ranking using $O(\log(1/Δ))$ repeated comparisons per context, and we prove that this logarithmic dependence is optimal. Despite this worst-case requirement, our Russian Roulette algorithm uses only $O(1)$ comparisons per context in expectation. Synthetic experiments and semi-synthetic experiments based on Arena data compare the four algorithms in settings with varying levels of preference heterogeneity and under varying context distributions.
cs.LG / 25 / 2610.10807
Symbolic Density Estimators for Unnormalized Distributions
Vikas Kanaujia, Riyansha Singh, Shashank Sharma, Vipul Arora
cs.LG · hep-lat
Abstract
Estimating the symbolic or analytical form of probability density functions (PDFs) from observed samples is a fundamental challenge in statistical and computational modelling. This process is critical for deriving interpretable and generalizable relationships characterizing the underlying phenomenon. Traditionally, this estimation depends strongly on domain expertise and prior field-specific knowledge, with experts selecting appropriate functional forms or parametric families based on empirical evidence and theoretical understanding. The coefficients of these forms are then typically determined through parameter estimation. In this paper, we develop a framework for estimating symbolic expressions of unnormalized distributions from observed samples using domain-specific prior knowledge, such as the range of interactions and a predefined set of primitive functions. We integrate deep generative models with symbolic regression (SR), incorporating inductive biases, such as factorizing large distributions, to keep the problem tractable. The deep generative models we examine include likelihood-based models, viz., flow models, and score-based models. Experiments show the effectiveness of the proposed framework for estimating density functions for multivariate toy distributions as well as lattices from computational physics, namely, XY model and $φ^4$ theory. When applied to the renormalization problem in $φ^4$ theory, the proposed framework estimates compact symbolic approximations of the hamiltonian function at different scales directly from samples, yielding expressions that may be challenging to derive using traditional perturbative or analytic approaches in nonperturbative settings.
cs.LG / 26 / 2610.10808
Controlled Acquisition and Abstention in Three-Channel Score Conflicts
Mengzhe Geng
cs.LG · cs.MM · eess.AS
Abstract
When audio, video, and text disagree, accuracy alone does not show whether to acquire another source or abstain. We study these choices in a controlled three-score benchmark: a policy observes two signed scores, may request the third at a cost, and can abstain. The primary reward is mechanism-specific: abstention is correct only for one designated ambiguity mechanism and is penalized under mixed corruption. Matched controls show that a threshold policy matches always-request decisions with fewer requests; its advantage over always-answer fusion depends on the reward assigned to that ambiguity. On a partially held-out synthetic split, the threshold policy reaches 0.789 +/- 0.006 targeted decision accuracy and 0.481 +/- 0.014 utility across 83 seeds. A three-score majority reference reaches 0.626 +/- 0.008 and 0.252 +/- 0.016, but uses more information. In a matched-budget test, a train-only value selector improves utility over no-query and matched-random policies at 10% and 25% budgets, while pair uncertainty has higher utility at every budget. At 50% and 63.7% budgets, the selector lowers utility despite slightly higher non-ambiguous accuracy. If all abstentions are scored incorrect, majority outranks the threshold policy in utility. At a central temporal setting, full-trace controls match the neural models while position perturbations separate them. On held-out-actor emotion clips, eight-frame fusion has opposite-signed accuracy differences for two encoder pairs, with both actor intervals containing zero; matched-request routing gains are small and uncertain. These results separate full-modality accuracy from pre-request selection value and show that selection value depends on budget and the observed-pair ranking.
cs.LG / 27 / 2610.10832
MotherTree: Meta-learning on synthetic data improves decision tree training
Ziyuan Wang, Fredrik D. Johansson
cs.LG
Abstract
Conventional decision tree algorithms produce effective, transparent models that can be audited, communicated, and deployed independently of the training data, but require learning every new task from scratch. In contrast, tabular foundation models demonstrate that meta-learning from a synthetic prior distribution enables strong in-context prediction for previously unseen tasks, especially in small-sample regimes. However, this approach does not produce a standalone model that can be inspected in isolation. We introduce MotherTree, a tabular transformer that meta-learns decision tree induction: given a training set for a new task, it outputs a hard, axis-aligned decision tree, equivalent in form to classically trained trees, in a single forward pass. MotherTree is pre-trained on a synthetic prior using stochastic gradient descent without requiring reference trees for supervision. On established benchmarks with controlled sample size, the approach is competitive with size-matched trees from common algorithms: recursive partitioning, gradient-based tree learning, globally optimal trees, and distillation from tabular foundation models. Notably, MotherTree consistently improves over from-scratch gradient-based learning and acts as a strong initializer: task-specific tuning of the generated tree outperforms the corresponding from-scratch learner on all benchmarks and sample sizes. These results show that meta-learning can provide effective inductive biases for learning stand-alone, small decision tree classifiers.
cs.LG / 28 / 2610.10848
Amortized Off-Policy Evaluation for LLMs
Younwoo Choi, Leo Feng, Vincent Liu, Haanvid Lee
cs.LG
Abstract
Accurate evaluation is central to selecting which LLM to deploy, yet testing a candidate on live traffic exposes real users to an unvetted model. Teams therefore evaluate candidates offline, on data produced by already-deployed models. This is off-policy evaluation (OPE), and it faces two distribution shifts: as a model is updated in post-training, its responses diverge from the logged ones (policy shift), and the reward definition under which it is judged changes with business requirements (reward shift). Classical OPE methods are ill-suited to this continual-deployment setting because they are defined per task and require fitting from scratch on every new logged dataset or reward definition. To address this, we propose PFN-OPE, a prior-data fitted network that amortizes OPE across a distribution of contextual-bandit tasks. We pretrain it once on tasks constructed from a pool of LLM responses scored by several reward functions, in which both shifts occur. At test-time it maps a logged dataset and one sampled target response per prompt to a value estimate in a single forward pass, with no per-task fitting. On HelpSteer2 and UltraFeedback with Qwen, Llama, and Gemma policies, PFN-OPE achieves 2.0 to 9.3 times lower error than the best baselines across all tested configurations in the reward-shifted settings.
cs.LG / 29 / 2610.10850
Similar Predictive Fit but Different Latent Dynamics: Characterizing Learned Dynamical Structure in Personalized Models of Brain Disorders
Rita Huan-Ting Peng, Nhat Bui
cs.LG · eess.SP · q-bio.NC
Abstract
As AI models move toward clinical decision-making and personalized treatment, understanding \emph{what} a model learns is important beyond predictive accuracy alone. We investigate whether personalized latent dynamics reveal clinically associated differences even when predictive fit is similar. A lightweight CNN--Transformer EEG foundation model pretrained on the Temple University EEG Corpus (TUEG) extracts segment-level representations. Using the Temple University Epilepsy Corpus (TUEP), representations are mapped to a shared latent-state space, and sparse multinomial logistic transition distributions (mLTD) are fit independently to each subject to obtain personalized transition-dependency graphs $W_n$. Analyses include $n{=}198$ subjects (99 epilepsy / 99 non-epilepsy). At $k{=}4$, epilepsy subjects exhibit substantially denser learned dependency structure ($p{=}1.1\times10^{-7}$), with the same pattern at $k{=}6$ (19.90 vs. 13.46; $p{=}5.2\times10^{-5}$). Graph-derived features provide moderate group discrimination under 5-fold subject-wise cross-validation (AUROC 0.68 at $k{=}4$; 0.65 at $k{=}6$). In contrast, held-out log-likelihood is nearly identical between groups at $k{=}4$ ($-0.992$ vs. $-0.991$; $p{=}0.95$), with similarly matched next-state prediction (AUROC 0.855 vs. 0.861; $p{=}0.54$). Thus, similar predictive fit does not imply similar learned dynamics: groups can be comparably predictable while differing substantially in the internal dynamical structure learned by personalized models. This distinction motivates evaluating learned structure alongside predictive performance in personalized clinical models.
cs.LG / 30 / 2610.10867
Shape irregularity of Life-Like Network Automaton rules as an indicator of classification performance
Lucas C. S. Oliveira, Michiel Rollier, Jan Baetens, Odemir M. Bruno
cs.LG · nlin.CG
Abstract
Complex Network (CN) classification requires high-level structural characterizations that are both scale-invariant and computationally efficient. Methods based on Life-Like Network Automata (LLNA) offer an interesting way to extract network descriptors by leveraging emergent temporal patterns without requiring provided features, but their efficacy is bottlenecked by a high-cost combinatorial optimization problem: the selection of the automaton transition rule. While current literature relies on exhaustive searches that are unfeasible for large-scale applications, this work reveals that the rule space is fundamentally structured by a property we term ``jaggedness'', that quantifies the resemblance of a LLNA transition function with a sawtooth shape. We demonstrate that this metric acts as a theoretical proxy for chaoticity and sensitivity -- properties essential for generating discriminative dynamic behaviors among network categories. Moreover, we introduce a heuristic search strategy that uses jaggedness to guide the rule selection. Experimental results show that our approach achieves classification accuracies within 5% of the global optimum while reducing computational overhead by 90% compared to exhaustive approach. Our findings provide a novel, efficient, framework for optimizing automata-based methods for pattern recognition.
cs.LG / 31 / 2610.10897
Gen-PINNs: Generative Adversarial Physics Informed Neural Networks for solving partial differential equations
Muhammad M. Akmal, Kamy Sepehrnoori, Michael J. Pyrcz
cs.LG · physics.comp-ph · physics.flu-dyn
Abstract
Physics-Informed Neural Networks (PINNs) are a widely used data-free method for solving Partial Differential Equations (PDEs) using machine learning. With recent advances in Generative Adversarial Networks (GANs), adversarial learning has shown strong capabilities for modeling complex data-driven problems; however, the use of GANs in deterministic physics-informed PDE solutions remains limited. In this work, we first identify limitations of standard PINNs for solving PDEs, including spectral bias, loss imbalance, and optimizer stagnation. We then propose Generative Adversarial Physics-Informed Neural Networks (Gen-PINNs), a unified deterministic residual-adversarial framework designed to improve data-free solutions of PDEs with sharp or shock-front behavior. The generator learns the underlying PDE solution using dynamically weighted physics-informed loss components, while separate discriminators evaluate complementary PDE residual features against ideal zero-residual states. The framework further develops and adapts several methodological components, including a Fourier representation for resolving high-frequency spatial content, an orthonormal spectral diagnostic for quantifying frequency-dependent solution errors, and a modified gradient-based dynamic weighting system for physics, initial-condition, boundary-condition, and adversarial loss objectives. Gen-PINNs is tested against standard PINNs on nonlinear and higher-order PDEs, including the Burgers, Allen-Cahn, and Kuramoto-Sivashinsky equations. The results demonstrate substantial improvements in accuracy and convergence across sharp-front, stiff, and higher-order PDE solutions, highlighting the potential of deterministic residual-adversarial learning as an effective approach for solving challenging nonlinear PDEs.
cs.LG / 32 / 2610.10899
How Hackable Is Your Speech Quality Metric? A Corrected Protocol, a Benchmark, and What Patching Buys
Ali Alavi, Donald S. Williamson
cs.LG
Abstract
Speech quality predictors are increasingly used as rewards, yet no agreed measure of their hackability exists. The usual measurement has two flaws. First, the perturbation reaches the predictor through a processing chain -- here a neural codec -- that shifts the score on its own, which scoring against the raw input charges to the attack. Referencing the unperturbed round trip instead changes measured hackability by up to a factor of four (0.31 to 0.08 for one defence). Second, one trained attacker is a sample, not a measurement: five attackers differing only in random seed reach success rates from 0.00 to 0.38 against one fixed predictor, so a defence claim needs the worst case over several. Under this protocol, four published predictors differ widely: NISQA is hacked on 90% of utterances, SSL-MOS on 21%, DNSMOS on 14% and UTMOS on 6%. We then audit a closed attack-detect-patch loop. It hardens the predictor only in its own attack space, by less than the spread between attackers; a random-perturbation baseline matches it; and it costs up to 0.30 system SRCC out of domain. Enhancers post-trained against patched predictors hack them far less (PESQ -0.03 versus -0.23). Code, preregistration and run outputs are released.
cs.LG / 33 / 2610.10920
Language Models for Page-Level Layout Decisions in E-commerce Search
Varun Joshi, Eva C. Song, ChengXiang Zhai
cs.LG · cs.CL · cs.IR
Abstract
E-commerce search pages are critical touchpoints for millions of online shoppers. While traditional search engines return a ranked list of results, modern E-commerce search pages increasingly incorporate recommender system modules -- for example, secondary stacks that surface alternative product groupings at specific positions. When introduced appropriately, secondary stacks can improve user engagement; however, suboptimal placement may disrupt browsing flow and degrade the primary results. Unlike traditional search ranking, where evaluation techniques such as interleaving are well established, evaluating page-level layout changes e.g., when and where to insert a secondary stack remains challenging without costly online A/B testing. To address this, we study offline methods for evaluating whether a given layout decision -- specifically, the inclusion of a secondary stack at a particular position -- is beneficial to users. We investigate language models as scalable evaluators by comparing direct prompt-based, prompt-derived feature, and representation-based methods. Our results show that representation-based approaches consistently outperform prompt-based judging in predicting user engagement, suggesting they provide a reliable foundation for offline layout evaluation in E-commerce search.
cs.LG / 34 / 2610.10926
Adaptive Multi-Discriminator WGAN Framework for Resource-Constrained Internet of Vehicles Using Reinforcement Learning and Game Theory
Farhoud Jafari Kaleibar, Amr M. Zaki, Marin Litoiu
cs.LG · cs.DC
Abstract
Managing machine learning workloads as a network service introduces a resource-orchestration problem distinct from conventional model training; which nodes should be allocated to a task, how communication and computation budgets should be divided among them, and how service quality should be sustained as connectivity and node availability change with mobility. Deploying Generative Adversarial Networks (GANs) in Internet of Vehicles (IoV) environments is a demanding instance of this problem; resource constraints, dynamic network topologies, and competing optimization objectives mean that traditional GAN architectures cannot simultaneously achieve high accuracy, efficient resource use, low delay, and low communication overhead. This paper introduces an adaptive multi-discriminator Wasserstein GAN (MD-WGAN) framework that integrates reinforcement learning with game-theoretic coordination to address these challenges jointly. In our framework, roadside units host generators paired with Deep Q-Network (DQN) agents that select discriminator subsets and manage distributed training across mobile vehicular nodes, while a game-theoretic coordination step allocates training epochs between generators and discriminators. A unified optimization objective ties adversarial learning quality to resource efficiency, communication overhead, and latency under vehicular constraints, allowing the framework to continuously adapt its training behavior as network conditions change. Evaluation on real-world NGSIM trajectory data shows that the framework attains prediction accuracy comparable to state-of-the-art GAN baselines - the lowest RMSE (1.029) and MAE (0.894) among all evaluated methods - while markedly improving resource efficiency: average CPU utilization is reduced by roughly 28% and mean memory usage by roughly 6%, at competitive communication overhead and latency.
cs.LG / 35 / 2610.10931
Coefficient Calibration as Selection Pressure in Symbolic Regression
Mattia Billa, Veronica Guidetti, Federica Mandreoli
cs.LG
Abstract
In memetic symbolic regression, candidate structures are compared after coefficient calibration, so the calibration protocol itself contributes to evolutionary selection. Standard centralized calibration evaluates each structure at its pooled-sample optimum, ignoring how stable this calibration is under covariate shifts, and can thus favor structures whose fit relies on sample-specific coefficients. We propose Dirichlet-Sinkhorn Constant Averaging (DSCA), a calibration strategy that partitions the optimization data into equally sized subsets with different covariate distributions, calibrates each candidate independently on every partition, and evaluates it at the mean of the resulting parameters. We show that the excess loss of DSCA relative to centralized calibration vanishes at the population level for correctly specified, identifiable expressions, whereas under misspecification it persists when partition-specific calibrations do not aggregate to the pooled optimum. On synthetic benchmarks and ten real-world datasets, DSCA improves functional recovery and the accuracy-complexity trade-off over centralized Broyden-Fletcher-Goldfarb-Shanno and Levenberg-Marquardt calibration, under selection by negative log-likelihood and by the Akaike and Bayesian information criteria. Mechanism analyses associate the DSCA excess loss with the generalization gap and show that the effect is not reproduced by repeated centralized fitting. These results indicate that controlled heterogeneous calibration provides a complementary source of selection pressure in symbolic-regression search.
cs.LG / 36 / 2610.10932
World-Model Policy Arbiter for Goal-Conditioned Reinforcement Learning
Junwei Quan, Evgenii Opryshko, Nicholas Rhinehart, Igor Gilitschenski
cs.LG
Abstract
Offline goal-conditioned reinforcement learning (GCRL) has produced a diverse set of goal-reaching algorithms, yet no single algorithm performs best across environments, goals, and even different phases of the same task. Rather than deploying only the best-performing policy, we ask whether a set of frozen goal-conditioned policies can be used collectively as a portfolio, deciding at every state which policy should act. Choosing a policy at each state is not straightforward. The policies' own value functions cannot be compared directly: they may use different scales, and some policies have no value function. We need to judge each policy by the states it is likely to reach, even though we can execute only one policy at a time. We also need to avoid switching so often that control becomes unstable. To address these challenges, we introduce World-Model Policy Arbiter (WMPA), a test-time framework that, given a set of frozen policies as input, rolls out each frozen policy in a learned state-space world model, evaluates the imagined futures with a shared goal-conditioned value function, and executes the highest-scoring policy for a short commitment interval before the next round of arbitration (policy selection). WMPA assumes access to a bank of frozen goal-conditioned policies and requires neither policy retraining nor privileged task-specific knowledge. Under the official OGBench evaluation protocol on 18 state-based datasets spanning maze navigation as well as cube, scene, and puzzle manipulation, WMPA improves the macro-average success rate from the 44% achieved by the best policy selected per dataset to 58%, with statistically significant gains on 12 datasets. These gains include +33 percentage points on cube-double-play and +36 percentage points on scene-play.
cs.LG / 37 / 2610.10933
Rethinking the Tradeoff Between Temporal Encoding and Nonlinear Computation in Spiking Language Models
Hanfei Liu, Shuchang Feng, Yanxia Chen, Changzeng Fu, Shiqi Zhao
cs.LG · cs.AI
Abstract
Spiking language models face a tradeoff between representing continuous semantic features over short temporal windows and retaining costly nonlinear attention operations. We introduce Spora, which jointly designs spike encodings and attention operators. Binary temporal weights let $T$ spikes represent compositional values with up to $T$ bits of capacity, compared with $O(\log_2 T)$ bits for spike-count readout. Unipolar Binary Spiking (UBS) uses thresholds and spike-triggered residual decay to produce non-negative integer codes; Bipolar Binary Spiking (BBS) separates sign and magnitude and learns a scale for signed activations. These representations support accumulation-and-shift dot products and integer-exponent mappings in attention. With four time steps, Spora achieves 76.6 average GLUE score and 44.1 CoLA MCC, improving over SpikeLM by 1.2 and 6.2 points, respectively. Extending BBS to six steps raises these scores to 78.2 and 47.4. Conditional-decay analysis, matched-budget activation-quantization comparisons, event-workload statistics, and fixed-point evaluation further characterize the connection between encoding fidelity and computational cost.
cs.LG / 38 / 2610.10947
RH-Detect: A Unified Benchmark for Reward Hacking Detection
Junwei Quan, Evgenii Opryshko, Rohan Subramani, Igor Gilitschenski
cs.LG
Abstract
Reward hacking, where a model exploits an evaluation signal without completing the intended task, threatens the reliability of deployed language model systems. Existing datasets use different labels, response formats, and metadata conventions, making detector results difficult to compare. We present RH-Detect, a benchmark that combines reward-hacking-relevant subsets from eleven public datasets, comprising 92,761 rows and six behavior categories, into a common schema. On 5,021 open-ended evaluation units, each comprising a task prompt and a free-form model continuation, including multi-turn tool-use trajectories, we evaluate six off-the-shelf language models from five families as reward hacking detectors without additional training. The best model achieves a pooled AUROC of 0.962, with accuracy above 93%. For the four strongest models, however, accuracy on the two multi-turn tool-use datasets, MALT and TRACE, is 10.7-15.9 percentage points lower than on the other sources at a common decision threshold, highlighting a key gap for deployment-time monitoring. We find that different input formats have different effects across models. Removing thinking raises Qwen3.5-4B AUROC from 0.779 to 0.849, but lowers Qwen Flash from 0.977 to 0.950. We also evaluate the benchmark as a training dataset for detectors. Holding out each source in turn, single-token SFT improves average AUROC on five of six held-out sources. A GRPO follow-up on that failure case yields a slight improvement in detection performance. Our results show that a single pooled score can conceal variation across data sources, detector inputs, and training procedures.
cs.LG / 39 / 2610.10952
SPD-MetaFormer is what you need for small-data brain decoding
Zhida Wang, Wei Lyu, Guo Yu, Sui Tang
cs.LG · stat.ML
Abstract
Brain signal decoding is challenging because neural recordings are noisy and vary across individuals, while labeled data are often limited. Recent attention-based models on the symmetric positive definite (SPD) manifold have nevertheless achieved strong performance using covariance and connectivity representations, yet the contribution of learned token weighting remains unclear. We examine two representative architectures, MAtt (based on log-Euclidean geometry) and GBWAtt (based on generalized Bures--Wasserstein geometry), and find that their learned attention weights remain close to uniform after training. We relate this behavior to bounded similarity parameterizations that, under the original softmax scaling, limit attention-weight contrast. Moreover, replacing learned weights with uniform weights, throughout training and evaluation, has little effect on mean predictive performance while preserving each model's original aggregation geometry. Motivated by these findings, we introduce SPD-MetaFormer, an attention-free architecture built on uniformly weighted Fréchet aggregation under log-Euclidean geometry. Its backbone uses a geodesic residual to update a summary token and a shared spectral feedforward map to transform all tokens, followed by a learned weighted readout. Token states remain SPD-valued until tangent-space classification. Across three EEG benchmarks, SPD-MetaFormer achieves competitive results relative to published Euclidean and manifold baselines. Separate matched reproductions test learned versus uniform weighting within MAtt and GBWAtt. These results suggest that, in the short-sequence and limited-data regimes studied, carefully designed SPD architectures can provide a simpler and effective alternative to adaptive manifold attention.
cs.LG / 40 / 2610.10957
TRACE: A Governance Framework for Measuring Explainability Debt in Production AI Systems
Harish Kant Pathak
cs.LG · cs.AI
Abstract
Production AI systems deployed in high-stakes domains accumulate a governance liability that existing monitoring frameworks fail to detect: the progressive inability to explain individual decisions when regulators, auditors, or affected individuals demand accountability. We introduce TRACE (Transparency, Risk, Accountability, Compliance, and Explainability), a seven-instrument governance framework for measuring, tracking, and remediating Explainability Debt in production AI systems. The foundational instrument, the Explainability Debt Score (EDS), quantifies the proportion of production decisions falling below a governance-defined explainability confidence threshold at any point in time. Complementary instruments include DART (Debt Accumulation Rate Tracker for breach forecasting), SHIV (Scenario Health and Integrity Validator for daily governance), FDE (Feature Drift Evaluator for causal attribution), HVE (Human Validation Engine), AIDE (Audit Intervention Decision Engine), and ZERO (Zero Explainability Risk Optimiser for remediation). Through a twelve-month longitudinal case study of a production fraud detection system processing 50,000 daily financial transactions, achieving 98.46% accuracy and ROC-AUC of 0.9990, we demonstrate that an EDS of 0.23 on audit day was statistically predictable six months in advance using DART trajectory analysis (beta = 0.008/week, R-squared = 0.94, 95% CI: [0.006, 0.010]), and that 78% of Explainability Debt was concentrated in the highest-regulatory-risk decision category (transactions above $10,000), a risk asymmetry completely invisible to system-level metrics. TRACE provides the first quantitative operational architecture for EU AI Act Article 13 compliance in production AI deployment, establishing a new subdiscipline of explanation governance distinct from explanation generation.
cs.LG / 41 / 2610.10972
Transferability of Learned States in Neural PDE Solvers
Shunye Wang, Haochen Wen, Shuo Li Liu, Xuanyi Wang, Lihao Liu, Zhongying Deng
cs.LG
Abstract
Assessing useful reuse in neural PDE solvers is challenging: final accuracy can reflect source learning and target-time computation. Our reuse contract separates solution accuracy, learning contribution, and numerical utility through paired state comparisons, matched target information and budgets, and cost accounting. A literature audit extracts 18 version-specific protocol records from 12 papers, documenting retained states, target-time resources, and reported controls. For a fixed linear system and residual tolerance, we construct two initial guesses with identical solution-error, energy-error, and residual norms, reaching the same solution with different conjugate-gradient (CG) iteration counts. Across 240 source-training trajectories, two linear PDE families, Fourier neural operators and convolutional networks, a fixed predictor's benefit reverses across correction algorithms. Among pairs with both relative prediction errors less than or equal to 5 percent on 64 in-distribution tasks (63 by 63 interior grids), reductions in all three norms accompany more CG iterations, at mean taskwise rates of 23.5 percent and 23.9 percent in two libraries. Work-based selection saves 2.50-3.33 CG iterations on held-out in-distribution tasks; matched adaptation demonstrates finite-budget pretraining value. Independent batches confirm a 0.73 percent complete online saving for one physics-trained Fourier neural operator against zero-initialized Poisson-preconditioned CG. Reuse requires matched state comparisons and downstream computational evidence.
cs.LG / 42 / 2610.10996
Invariant-Measure Reasoners: Stable Representations for Latent Reasoning
Yuto Inui, Takuya Konishi, Yoshinobu Kawahara
cs.LG
Abstract
Latent reasoning models repeatedly update a latent state using the same recurrent block. As the recurrent depth increases, the sequence of latent states may converge to a compact subset of the state space without necessarily converging to a fixed point. Existing models typically predict by applying a prediction head to a single latent state. However, the latent state can continue to change even after many updates, potentially making predictions unstable across recurrent depths. To address this instability, we introduce invariant-measure reasoners (ImR), a framework that uses an invariant measure as a stable representation. This measure describes the long-run distribution of latent states on the compact subset and is invariant under updates by the recurrent block. ImR predicts from the expectation of the prediction head's output under this measure. We use ImR in two ways: fine-tuning only the prediction head of existing models and training models from scratch. Both approaches reduce prediction instability and improve accuracy in many settings on maze and Sudoku tasks. In some settings, models trained with ImR exhibit non-fixed-point behavior more frequently than existing models yet achieve high accuracy even with such behavior, unlike existing models. These results suggest that ImR can leverage otherwise destabilizing dynamics for latent reasoning.
cs.LG / 43 / 2610.10998
F$^3$NO: Frequency-Decomposed Finite-Time Flow-map Neural Operators with Cross-Scale Conditioning
Fan Wu, Cheng Jing, Kookjin Lee
cs.LG
Abstract
Neural operators enable fast PDE forecasting, but repeated predictions accumulate errors and fine-scale structures remain difficult to resolve. We introduce a frequency-decomposed finite-time flow-map neural operator (F$^3$NO) that leverages updated low-frequency features to guide nonlinear refinement of high-frequency information. Within each layer, this cross-scale conditioning connects global spectral processing with local detail refinement. The model directly predicts states at specified future times and adjusts the contributions of the two branches according to the prediction interval. For longer trajectories, it combines parallel predictions within short temporal segments with recursive propagation between segments. Experiments on five PDE benchmarks demonstrate improved forecasting accuracy over autoregressive and direct-prediction baselines. Ablations show that frequency-decomposed refinement can improve accuracy with fewer parameters, while the benefits of segmentation depend on spatial resolution and dynamical regime.
cs.LG / 44 / 2610.11000
Low-rank tensor structure of precipitation and its application to satellite-reference merging
Ryan Solgi, Rohan Shankar, Hugo A. Loaiciga
cs.LG · physics.ao-ph
Abstract
The intermittent and variable nature of precipitation makes its accurate estimation over extended domains difficult, yet its spatiotemporal structure suggests that a low-rank representation may be possible. This work represents daily precipitation over the contiguous United States (CONUS) as spatiotemporal tensors and applies CANDECOMP/PARAFAC factorization, showing that preserving the native spatial and temporal modes yields more accurate reconstruction than factorizing independent daily fields or unfolded space--time matrices. Building on this finding, this work presents TMerge, a tensor-based framework that integrates satellite precipitation with sparse reference observations through shared low-rank spatial and temporal factors. TMerge was applied to correct the IMERG Final Run product with climate prediction center reference observations over CONUS. During 2019-2022, TMerge increased correlation from 0.53 to 0.85 and reduced root-mean-square error and mean absolute error by 48.2% and 29.3%, respectively. TMerge consistently outperformed linear bias correction, quantile mapping, and neural networks across seasons, precipitation-intensity regimes, and regions. Improvements were spatially coherent and largest in coastal regions where IMERG errors were greatest. These results demonstrate that low-rank tensor structure parsimoniously approximates the dominant spatiotemporal variability of precipitation and provides a practical mechanism for improving satellite estimates under limited reference observations over extended domains.
cs.LG / 45 / 2610.11064
Measuring and Mitigating Solution Mode Collapse in RLVR
Liv G. d'Aliberti, Marwa Abdulhai, Sofiia Druchyna, Peter Henderson, Manoel Horta Ribeiro
cs.LG · cs.CL
Abstract
A language model (LM) can usually answer the same question in more than one way, but reinforcement learning with verifiable rewards (RLVR) is indifferent to which correct answer a model produces. A solution will earn the same reward whether it is the thousandth copy of a familiar answer or one the model has never produced before. Yet, there is potential value in having the model retain multiple correct solutions as it is trained. For instance, multiple modes may give users a choice and provide problem-solving strategies that improve overall model performance. Here, we introduce ModeBench, a benchmark of multi-solution tasks in which the verifier returns both correctness and mode discovered. We then use ModeBench to measure how solution diversity changes under RLVR post-training. We find that RLVR post-training concentrates probability onto fewer correct modes even as accuracy holds or improves, and moreover, that frontier models are already highly concentrated. We then introduce our solution, Re:Max, which stores one verified example per discovered mode in a replay buffer and trains on those stored modes uniformly. A solution found once is, therefore, practiced as often as one found repeatedly. Across three model scales, two RL objectives, and harder task constructions, replay improves both how often a policy succeeds and how many different ways it can succeed.
cs.LG / 46 / 2610.11065
CityDeploy-Bench: Benchmarking Physics-Grounded Spatial Set Planning for Multi-Transmitter Network Deployment
Chenyang Yuan, Xiaoyuan Cheng
cs.LG
Abstract
Automating city-scale wireless deployment remains challenging under complex urban propagation and network-wide interference. We introduce \textbf{CityDeploy-Bench}, a benchmark that reframes multi-transmitter deployment as \emph{physics-grounded spatial set planning} under a unified ray-tracing verifier. The benchmark separates utility representation from planning dynamics, enabling controlled comparison between direct scalar rewards, relational models, and higher-order interaction structures across diverse planners. Our experiments reveal a clear transition in planning behavior as physical coupling grows. Deployment quality becomes increasingly dependent on whether the learned utility captures collective transmitter interactions, whereas stronger search alone cannot compensate for missing relational structure. This establishes multi-transmitter deployment as a coordination problem over physically interacting sets rather than a collection of independent spatial decisions. We release CityDeploy-Data and the benchmark framework as a reproducible testbed for research linking decision learning with physically grounded wireless network design.
cs.LG / 47 / 2610.11074
Optimally Pacing Budget Spending and Learning
Mark Braverman, Jingyi Liu, Jieming Mao, Jon Schneider, Eric Xue
cs.LG · cs.GT
Abstract
We establish near-optimal regret bounds for budget-constrained online learning against arbitrary classes of budget-pacing experts in the adversarial setting. In particular, given any class of $F$ experts and a candidate budget pacing schedule, we provide a full-information algorithm which obtains regret $O(D \sqrt{\log F}+ \sqrt{T\log F})$ against all experts whose cumulative spending stays within distance $D$ of this schedule, matching lower bounds established by Braverman et al. (2025). We additionally show that our technique extends to various problems in online resource allocation, where the learner gets to see the rewards and costs of the current options available to them, and establish $O(D\sqrt{\log F})$ regret bounds when fractional allocation is allowed. This is the first algorithm we are aware of which can achieve $o(\sqrt{T})$ guarantees for such tasks.
cs.LG / 48 / 2610.11076
Stability-Plasticity Balance via Singular-Vector Selection in LLM Continual Learning
Lingxiang Wang, Hainan Zhang, Liang Pang, Hongwei Zheng, Zhiming Zheng
cs.LG · cs.AI
Abstract
Domain-specific continual adaptation of LLMs risks catastrophic forgetting, creating a fundamental tension between acquiring new capabilities and preserving those learned during pretraining. PEFT mitigates this problem by restricting the number of trainable parameters, but existing methods lack a principled unit for deciding where plasticity should be allocated and stability should be preserved. We identify the singular-vector channel as a natural unit for managing this trade-off. Each channel represents an input-output transformation, which can be updated to acquire new knowledge or fixed to preserve pretrained capabilities. Based on this perspective, we introduce SVC, a parameter-efficient continual-learning method that selectively updates Singular-Vector Channels. Before fine-tuning, SVC uses domain-specific data to estimate each channel's adaptation benefit and a fixed public general-domain corpus only as a history activation proxy for estimating forgetting cost. It then adaptively selects trainable channels based on these scores via knee-based cost screening, Pareto-front filtering, and Otsu thresholding. Experimental results across four LLM families and eight downstream tasks show that SVC better preserves pretrained capabilities while achieving strong downstream performance relative to existing PEFT baselines. Further analysis of channel scoring and selection demonstrates that selective plasticity at the singular-vector-channel level enables effective continual LLM adaptation.
cs.LG / 49 / 2610.11085
DaCe-DT: Data-Centric Offline Multi-Task Reinforcement Learning via Adaptive Prompts and Trajectory Correction for Heterogeneous Tasks
Xinfei Wang, Shanchen Pang, Chenhao Zhang, Shudong Wang, Wenhao Ji, Haiyuan Gui, Meng Han, Xiaojian Liao
cs.LG
Abstract
Offline multi-task reinforcement learning (Offline MTRL) heavily depends on the quality and distribution of pre-collected data. However, existing methods mainly focus on algorithmic optimization, with less emphasis on data-level improvements to enhance learning ability and generalization performance. This paper, from a data perspective, reveals three key bottlenecks that limit Offline MTRL performance:(i) ineffective utilization of prompts length under diverse task complexities, and (ii) semantic irrelevance of randomly sampled prompt segments, (iii) misleading supervision induced by fragmented and discontinuous trajectories. To address these challenges, we propose DaCe-DT, a robust offline MTRL framework designed to be insensitive to heterogeneous task complexities and data quality, featuring length-gated prompt masking (LGPM), retrieval-augmented prompt construction (RAPC), and value-adaptive return calibration (VARC). Together, these mechanisms enable DaCe-DT to deliver data-centric prompt adaptation and trajectory refinement, resulting in robust multi-task generalization and stable policy learning amid heterogeneous offline data and tasks. Experimental results on Meta-World show that DaCe-DT consistently outperforms state-of-the-art methods, achieving an average improvement of 11.73% on optimal datasets and an improvement of 13.34% on suboptimal datasets, demonstrating its effectiveness in learning stably from imperfect data and improving overall multi-task performance.
cs.LG / 50 / 2610.11108
Cova-PINN: Cross-Domain Conservation Physics-Informed Neural Network for Fluid-Solid Conjugate Heat Transfer in Complex Geometries
Weizheng Zhang, Xunjie Xie, Hao Pan, Lin Lu
cs.LG · physics.flu-dyn
Abstract
Multi-domain physics-informed neural networks (PINNs) flexibly model medium-specific representations to solve fluid--solid conjugate heat transfer (CHT). However, standard multi-domain PINNs enforce governing equations and interface conditions on separately sampled domain supports, which can yield plausible temperature fields but inaccurate end-to-end energy transfer and outlet temperatures. We propose Cova-PINN, a multi-domain PINN framework that aligns conservation support with thermal interaction paths in complex geometries. Cova-PINN jointly optimizes cross-domain composite control-volume balances at the local scale and paired-wall closure at the global exchanger scale. We evaluate Cova-PINN on four triply periodic minimal surface (TPMS) heat exchangers and a geometrically distinct DualMS design against CHT-specific, optimization-oriented, and complex-geometry PINN baselines under a common protocol. Relative to the closest baseline, MUSA-PINN-CHT, Cova-PINN reduces average outlet-temperature and device-level closure errors across the four TPMS topologies by $37.7\%$ and $60.2\%$, respectively, while also improving full-field and heat-duty accuracy, with consistent gains on DualMS.
cs.LG / 51 / 2610.11115
Dynamics as Code: On Model Compression via Dynamic System
Fan Gao, Wei Su, Juntong Fan, Renfeng Peng, Hongyu Liu, Jinqiao Duan, Feng-Lei Fan
cs.LG
Abstract
The escalating size of pretrained neural networks has rendered model compression a prerequisite for deployment under stringent memory and compute constraints. With the irrational winding as an example, earlier work introduced a dynamic system (DS) paradigm that reconceptualizes compression as compact weight representation: high-dimensional parameters are encoded by the index of a trajectory produced by a dynamic system, from which the vector is recovered during decompression. This mechanism is fundamentally distinct from pruning, quantization, knowledge distillation, and low-rank decomposition. Along this direction, we prove that under a Diophantine condition, a finite trajectory of \(M = O(ε^{-(d+ν)})\) states in the irrational winding constitutes an \(ε\)-net over the \(d\)-dimensional weight space, thereby linking state resolution, decompression error, and compression ratio in a predictable manner. Furthermore, we propose a generalized DS-based model compression framework by unifying four DS families---space-filling curves (Hilbert, Peano, Morton/Z-order, Snake), chaotic systems (Lorenz), congruential and pseudo-random generators (LCG, PCG), and low-discrepancy sequences (Halton). Also, we introduce the KD-tree and coordinate-template acceleration to scale to large models as well as outlier identification to control the error. Experiments on ResNet-18 and Qwen2.5-1.5B/Qwen1.5-7B validate that DS-based compression achieves competitive compression ratios without post-hoc retraining, with controllable decompression error and flexible state-space design, establishing it as a principled and practical compression approach.
cs.LG / 52 / 2610.11133
Machine Learning Optimization for Enhanced OS Fingerprinting
Jae Sung Kim, Spencer Ekeroth, Jeremy Neale
cs.LG
Abstract
Operating System (OS) Fingerprinting is a technique that can be used to identify a network's operating systems by evaluating network traffic in the form of TCP/IP packets. This research will explore the effectiveness of passively identifying operating systems on the CIC-IDS2017 dataset, a collection of over 47 gigabytes of pcap files with their corresponding operating systems. This research also proposes a new command line interface, OsirisML, which uses nPrint to preprocess the data into tabular data and XGBoost to apply ML to the data to generate, retrain, and test ML models. When packets are split randomly between training and testing, OsirisML models reach an accuracy of 97.66% on a down-sampled subset of the Friday capture and 84.69% on the entire capture. On the entire Monday capture, which contains no attacks, OsirisML reaches an accuracy of 73.83% and an F-1 score of 79.38%.
cs.LG / 53 / 2610.11134
QUILT: Rethinking Sparse-Attention Prefill through Shared Query Execution
Zhenduo Zhao, Qihui Zhou, Mingcong Song, Zhiyi Chen, Chuangguan Ye, Fengfan Hou, Zequn Gong, Jing Li, Hongjie Si, Guoping Long
cs.LG
Abstract
Sparse attention reduces the cost of long-context attention, but existing kernels typically process queries independently, repeatedly loading and dequantizing KV entries shared across queries. We observe substantial overlap in the KV entries selected by neighboring queries, creating opportunities for cross-query reuse. We present QUILT, a workload-aware sparse-attention execution mechanism that jointly processes neighboring queries and reuses shared KV entries to reduce redundant memory traffic and computation. QUILT introduces Shift-and-Compare Set Decomposition (SCSD), which transforms irregular set operations into regular data-parallel primitives suitable for modern accelerators, and pipelines SCSD with attention computation to hide its overhead. Cascaded sharing captures reuse hierarchically at multiple granularities. A tile-aware execution strategy balances sharing granularity with hardware tile utilization and selectively removes low-importance query-specific tails to eliminate underutilized tiles. We evaluate QUILT on LongBench using GLM-5.3 and DeepSeek-3.2 under both tensor and sequence parallelism. Compared with the state-of-the-art sparse-attention kernel, QUILT reduces average kernel latency by up to 55.1% and processed KV data by up to 55.9%, while reducing time-to-first-token (TTFT) latency by up to 36.8% with negligible accuracy degradation.
cs.LG / 54 / 2610.11140
ActiveMedAgent: Cost-Aware Trajectory Learning for Multimodal Medical Diagnosis
Weiwei Ma, Xiaobing Yu, Peijie Qiu, Jin Yang, Zhaoqi An, Xuanzhao Dong, Xiaoqi Zhao, Xiaofeng Liu
cs.LG · cs.AI · cs.CL · cs.CV · cs.MM
Abstract
Clinical diagnosis is inherently sequential: clinicians escalate from cheap to costly tests only when additional evidence is expected to resolve diagnostic uncertainty. We present ActiveMedAgent, a framework that brings this cost-aware sequential logic to multimodal medical AI. Given a frozen, API-accessed vision-language model, ActiveMedAgent tracks probability distributions over candidate diagnoses and scores each acquisition by its per-step diagnostic utility minus cost. A lightweight MLP controller is then trained offline on these scored trajectories, learning when to request additional evidence and when to commit. Across three commonly used benchmarks, trajectory-based policy learning consistently outperforms both unguided acquisition and full-modality baselines. Notably, we identify an information overload effect. In 175 cases, the agent produces a correct diagnosis with fewer channels while the full-modality baseline fails, showing that learning what to omit can be as important as learning what to acquire.
cs.LG / 55 / 2610.11146
Ranking Prior Alignment for Credit Risk Modeling: When Do External Priors Matter?
Qiye Lu, Jiang Ji, Liang Zhang
cs.LG
Abstract
Cold-start credit scoring -- deploying models with scarce labeled data, weak features, or minimal capacity -- is a recurring problem in financial machine learning. When a new lending product launches, labeled default data is scarce, feature pipelines are immature, and models must be deployed with minimal capacity to avoid overfitting. Standard defenses operate on the same limited data; what is needed is a source of external regularization grounded in domain knowledge. We propose Ranking Prior Alignment, a model-agnostic framework that distills external ranking priors (from domain experts, teacher models, or LLMs) into any scoring model via a temperature-scaled KL divergence loss. The framework unifies neural (MIL attention) and tree-based (XGBoost custom objective) architectures through a single formulation: L = L_task + gamma(t) * KL(P_agent || P_model), where gamma(t) follows an exponential decay schedule. The method requires no external model at inference, and its tree-based instantiation tolerates annotation noise up to eta = 0.5. On an industrial dataset of over 1.5M merchants, MIL alignment achieves 7/7 positive evaluation cells at 3K bags (1 ID + 3 OOT + 3 degradation metrics; peak Delta AUC = +0.020 on OOT-1), and XGBoost ablation achieves 9/9 positive metrics at 300 bags. Cross-dataset validation on public Amex shows 5/5 positive folds (avg Delta AUC = +0.041). Four model families (MIL, XGBoost, LightGBM, Logistic Regression) and four teacher architectures show that the framework is both model-agnostic and prior-source-independent. We further observe that alignment gains exhibit an inverse-scaling pattern: benefits grow as data abundance N, model capacity C, and feature quality Q decrease, helping practitioners decide when to invest in prior annotation.
cs.LG / 56 / 2610.11152
Do LLMs Learn from Rewards in Context? : Rethinking the role of reward in In-Context Reinforcement Learning
Minchan Kwon, Seunghee Koh, Sunghyun Baek, Minsung Bae, Junmo Kim
cs.LG · cs.CL
Abstract
LLM agents increasingly improve at inference time by accumulating experience in context rather than by updating parameters. This process is often described as in-context reinforcement learning (ICRL). Whether in-context learning (ICL) can actually play the role of RL, however, has not been tested. We study this question in its simplest form, direct ICRL, where the model conditions directly on raw trajectory-reward pairs, and ask whether the reward acts as a learning signal. Through controlled experiments on four benchmarks across six models, we find that the reward is read, but its effect is small: flipping, randomizing, or removing the reward leaves the improvement curve almost unchanged, and this holds even under meta-prompts that explicitly instruct the model to explore, exploit, or reason over rewards. Trajectories drive improvement, but not through their semantic content: shuffled or corrupted trajectories work as well as real ones. These patterns closely mirror those known in ICL, suggesting that direct ICRL is better understood as a special case of ICL than as inference-time RL. This reframing has implications for agent memory design: ICL factors such as input distribution and demonstrations may matter more than RL elements such as reward shaping and exploration.
cs.LG / 57 / 2610.11164
RideBench: A Large-Scale Exogenous-Aware Benchmark for Ride-Hailing Time Series Forecasting
Shengsheng Lin, Jing Hu, Zhengyang Hu, Jiazheng Sun, Zichun Cao, Siwei Sun, Zhichao Zou, Enyun Yu, Dongdong Li, Xinyi Hu, Weiwei Lin
cs.LG
Abstract
We release Ride-Hailing, a large-scale ride-hailing time series dataset synthesized from DiDi's marketplace data across 200 spatial areas. Ride-Hailing spans four consecutive years at half-hourly granularity and covers three representative exogenous scenarios: Weather Disturbance, Holiday Effect, and Large-scale Event Impact. Built upon Ride-Hailing, we introduce RideBench, a comprehensive benchmark for exogenous-aware ride-hailing forecasting, covering both regular week-ahead forecasting and long-horizon 8-week-ahead forecasting with up to 2,688 prediction steps. RideBench evaluates over 30 representative forecasting methods, including endogenous-only models, exogenous-aware models, and time series foundation models. Our results show that future-known exogenous variables provide clear benefits in regular week-ahead forecasting, especially under weather, holiday, and large-scale event (e.g., major sporting events and concerts) scenarios. However, current exogenous-aware models still struggle to fully capture disturbance-induced pattern changes under complex external contexts. For long-horizon forecasting, existing models cannot simultaneously achieve low pointwise errors, accurate broad trends, and reliable near-term forecasts. These findings reveal a clear mismatch between existing forecasting models and real-world ride-hailing requirements, highlighting the need for models that can better exploit future-known exogenous information, scale across heterogeneous areas, and support long-horizon planning. By introducing Ride-Hailing and RideBench, we aim to encourage the community to study these practical challenges in real-world ride-hailing forecasting.
cs.LG / 58 / 2610.11165
CARE: A Lightweight Plug-in Gated Correction and Uncertainty-aware Module for Long-term Time Series Forecasting
Guo Cheng, Changlong Lv, Jingyi Hou
cs.LG
Abstract
Multivariate long-horizon forecasting is critical to electricity load scheduling and traffic flow management, and to financial risk control. Existing deterministic backbones output a single trajectory, masking heterogeneous prediction difficulty across horizons and channels and providing no localized reliability signal. We present CARE (Corrective branch with Aligned context and Relative-error Estimation), a lightweight plug-in that enhances any deterministic forecaster without architectural redesign. Operating in parallel with the base model, CARE resamples historical context to match the forecast horizon, learns residual correction patterns from this aligned history, and applies scale-aware bounded updates modulated by per-coordinate sigmoid risk gates. A multi-objective loss jointly optimizes forecast accuracy, residual tracking, risk alignment, and base-model anchoring. Across eight benchmarks with three representative backbones, CARE improves accuracy with marginal parameter and latency overhead. Its risk gates reliably identify high-error regions: on Weather, the highest-gate tertile exhibits nearly four times the error of the lowest-gate tertile, offering planners an interpretable per-step trust signal. Code is available at https://github.com/CG-BNYC/CARE.
cs.LG / 59 / 2610.11167
PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation
Heng Li, Yong Zhang, Ning Cheng, Zhigen Li, Yun Zhu, Yanmeng Wang, Shaojun Wang, Jing Xiao
cs.LG · cs.AI
Abstract
Vertical-domain few-shot classification remains challenging for small language models, as limited supervision makes it difficult to acquire domain-specific decision knowledge. On-Policy Distillation (OPD) can improve teacher-guided adaptation by supervising student-generated rollouts, while GRPO-based reinforcement learning can further refine downstream predictions. However, existing KD-to-RL pipelines typically rely on globally fixed transition schedules, ignoring that different samples may require different amounts of teacher-guided acquisition before reward-driven refinement. We propose PIVOT (Perplexity-Informed Transition Optimization), a dynamic transition framework that routes samples between OPD and GRPO according to teacher-evaluated sequence perplexity. PIVOT moves low-perplexity samples to GRPO for reward-driven refinement while keeping high-perplexity samples under OPD for continued domain knowledge acquisition. Experiments on Banking77 and HWU64 show that PIVOT consistently outperforms continued OPD and globally synchronized OPD$\rightarrow$GRPO baselines under the same number of post-warm-up student optimization steps, achieving stronger downstream performance and more stable training dynamics.
cs.LG / 60 / 2610.11170
SACQ: Structured Decoding with Memory-Conditioned Refinement for Long-Horizon Forecasting
Guo Cheng, Zhengzhuo Xu, Chenchen Jing, Jingyi Hou
cs.LG
Abstract
Long-term time series forecasting (LTSF) models predominantly employ patch-based encoders terminated by a flatten readout head that maps the entire encoded historical memory to all future steps through a single shared projection. This implicit coupling of future positions obscures position-specific historical-to-future alignment and amplifies sensitivity to corrupted inputs and extreme supervision noise. We present SACQ, a plug-in structured prediction head that replaces flatten readout while keeping the encoder unchanged. SACQ adopts a two-stage decoding pipeline: it first establishes a coarse patch-grid forecast scaffold, then refines each future position through cross-attention over historical memory and merges the attention-derived correction with the coarse scaffold via a learned per-patch gate. To stabilize optimization under long horizons and noisy labels, we further propose a batch-adaptive scaled log-cosh loss that automatically calibrates robustness to the current residual scale, suppressing outlier gradients while preserving MSE-like sensitivity for typical errors. SACQ attains top-tier test MSE/MAE across PatchTST, DLinear, and patch-Mamba backbones with only modest incremental overhead in parameters and latency. Under inference-time input corruption and training-set label-noise stress tests, SACQ substantially outperforms flatten readouts, with ablation studies validating each architectural component.
cs.LG / 61 / 2610.11185
Predictive Multiplicity in Cell-Fate Assignment: Label-Free Rashomon Sets and the Limits of Per-Cell Certification
Arjun Bhupatiraju, Abhiram Bhupatiraju
cs.LG · q-bio.GN
Abstract
Single-cell trajectory inference maps transcriptomic measurements onto developmental continua, yet configurations that fit the data equally well can assign conflicting cell fates. FateMultiplicity is a label-free framework that constructs a statistically admissible model set, or Rashomon set, without lineage labels, by evaluating model discrepancy on cross-fitted held-out genes under non-inferiority testing calibrated against random-seed variation. Multiplicity is large and depends more on the diversity of the model space than its size: twelve configurations of a second algorithm expose 20.0% of cells where twenty-four of the first expose 3.8%. Whether the per-cell certified fate margin FM yields more reliable assignments than the fitted model already provides is then tested, and it does not. On simulation ground truth, on the same cells, FM discriminates misassignment at AUC 0.682, against 0.965 for the baseline configuration's own decision margin (p = 0.003) and 0.854 for a seed-dispersion baseline. Informativeness is governed by the breadth of the admitted set, not its cardinality: at cardinality four, seed refits give 0.933 and hyperparameter-perturbed sets 0.701. Relaxing the infimum to a q-quantile recovers discrimination but converges toward the single model's own confidence; the supremum reaches 0.973 because theta*'s membership bounds it from below, while the infimum is unanchored. Multiplicity in trajectory inference is worth measuring and reporting, but per-cell certification over a label-free Rashomon set is not a route to more reliable fate calls. Two constructions survive: a margin-erosion ratio separates real from spurious branch points in simulation (AUC 0.890, untested on real data), and against clonally observed fate, uncertified cells disagree with their clone's outcome 16.4 percentage points more often than certified cells (p < 0.001).
cs.LG / 62 / 2610.11206
Do Flatter Minima Drive Better Generalization? An Algorithmic Separation in Grokking
Mohnish Harwani
cs.LG
Abstract
Flat loss landscapes have long been linked to better generalization in neural networks. However, its role as a causal mechanism for generalization is less established. Grokking provides an unique testbed to understand this distinction: models are prone to fit observed data using non-generalizing structure and remain in that regime for prolonged periods, transitioning to generalization only under particular training conditions. In this work, we study whether flat loss landscapes can act as a driving mechanism in this transition. While recent work has argued for flatness as a necessary geometric condition for this transition, we find that biasing training toward flatter solutions using sharpness-aware minimization (SAM) is insufficient to reliably induce this transition, despite producing flatter solutions. However, when SAM is paired with mechanisms that drive generalization such as weight decay, an interesting property emerges: SAM can accelerate the transition to generalizing solutions by up to 4x at the epoch-level. We theoretically untangle this relationship between SAM and weight decay using a minimal interpolating two-layer ReLU model with both memorizing and generalizing solutions. We show that even in this simple setup, flatness alone cannot distinguish a memorizing solution from a generalizing one, while weight decay favors generalizing solutions. However, under a local stability analysis, there exists a window where a memorizing interpolant is locally stable under gradient descent but unstable under SAM in the low-norm regime, which can explain SAM's ability to accelerate this transition. Overall, our results provide a more interpretable account of the role of flatness in driving generalization, especially in settings where models are vulnerable to minimizing loss through learning non-generalizing structure.
cs.LG / 63 / 2610.11214
Bridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight Migration
Kaicheng Xiao, Liran Dong, Haotian Li, Guoliang Xing
cs.LG · cs.CL
Abstract
KV-cache quantization and linear attention are two representative approaches to tackling the storage and computational costs of Transformers. KV-cache quantization compresses individual KV entries into discrete codes but retains all entries, whereas linear attention recurrently aggregates multiple historical KV contributions into a fixed-size continuous state but can introduce interference. This contrast raises the question of whether per-KV compression and multi-KV aggregation can be bridged within a single mechanism for efficient attention. We identify RAM-Net as such a bridge through soft assignments over a discrete address space. These assignments determine recurrent updates to the continuous slot state associated with each address. Under a restricted RAM-Net construction, we prove that soft address assignments extend hard quantized matching to a separable read-write overlap that locally approximates full-attention similarity and supports recurrent aggregation. These connections further enable Transformer-to-RAM-Net weight migration through a new path based on a soft-quantized intermediate construction. Across nine pretrained Transformer models from 0.3B to 7B parameters, RAM-Net recovers an average of 87.1% of the teachers' accuracy gains over random guessing across six commonsense and knowledge tasks using only a 500M-token budget per model.
cs.LG / 64 / 2610.11229
SteerCast: Retrieval-Based Latent Steering for Decoder-Only Time Series Forecasting
Van Dai Do, Huu Hiep Nguyen, Minh Hoang Nguyen, Hung Le
cs.LG
Abstract
Time series forecasting aims to predict future values from historical observations and auxiliary features. We propose \textbf{SteerCast}, a retrieval-based latent steering method that improves decoder-only forecaster at inference time, without updating its parameters. SteerCast constructs a database from the training set by storing a representation of each history window together with a \emph{steering vector} computed in the forecaster's latent space, defined as the difference between representations induced by the ground-truth continuation and by the model's own prediction. At test time, SteerCast retrieves nearest neighbors for a query history, aggregates their steering vectors, and injects the resulting signal into the forecaster's hidden states at every step of autoregressive generation, guiding predictions toward trajectories consistent with similar training cases. Experiments across diverse multivariate benchmarks and multiple horizons show that SteerCast consistently improves forecasting accuracy over the fine-tuned backbone and retrieval-based baselines, while requiring no additional training beyond the original fine-tuning and using only the training set as a retrieval corpus.
cs.LG / 65 / 2610.11236
Low-Cost Sensor Calibration for Indoor Air Quality Monitoring: A Dataset, Evaluation Scenarios, and a Lightweight Model
Jinyong Yun, Seokho Ahn, Hyungjin Kim, Sungbok Shin, Young-Duk Seo
cs.LG
Abstract
Low-cost sensors enable scalable indoor air quality monitoring but require calibration because of nonlinear distortions, noise, and temporal drift. The conventional strict pairwise calibration setting requires a co-located reference sensor at each deployment location and does not account for spatial and temporal heterogeneity. To address these limitations, we introduce a six-month dataset comprising multivariate indoor air-quality measurements from low-cost and reference sensors with contextual metadata collected at five locations. Using this dataset, we define four evaluation scenarios. The reference-efficient and location-transfer scenarios evaluate spatial generalization, whereas the long-term drift and event-conditioned scenarios assess robustness to gradual and abrupt distribution shifts. Based on these scenarios, we derive design requirements and propose a lightweight temporal model that combines input-window compression with residual temporal and feature fusion. Experiments show strong calibration performance across all four scenarios with low edge-inference cost.
cs.LG / 66 / 2610.11245
Read What Matters: Query-Adaptive Quantization for KV Caches
Siddharth Bhandari, Lucas Gretta, Krishna Balasubramanian, Shiva Kasiviswanathan
cs.LG · cs.CL · cs.PF
Abstract
KV-cache entries are stored before their future queries are known, but each decoding query needs precision in different places. We study this mismatch using separate budgets for retained bits and bits fetched per query. ReadKV stores each key and value in a progressive code whose prefixes support different reconstruction precisions. For each query, it allocates key-channel prefixes using the query, computes attention from the reconstructed keys, and then allocates value-token prefixes using that attention. Stored entries remain unchanged. Each stage optimizes a calibrated distortion objective under a fixed budget; we prove exact allocation under diminishing refinement gains and relate these objectives to attention-output error. We also exhibit a finite-dimensional attention family where query-dependent access strictly outperforms every query-independent reader at the same read budget, even with unrestricted competing encoders and decoders. Across six base models, reading four bits on average from an eight-bit cache increases C4 perplexity by at most 0.66%, using about one quarter of the logical reads and half the retained capacity of a 16-bit cache. It is consistently more accurate than storing and fully reading four bits at the same payload-read budget. Retaining more bits than each query fetches is aimed at long-context decoding, where the cache bytes moved per step, rather than the weights, dominate cost. Long-context question answering and retrieval on two instruction-tuned models provide additional quality evidence. On the tested 8K-token, batch-one, single-layer workload on an NVIDIA A10G, a restricted eight-bit ReadKV reader with a two-bit mean payload-read budget has 39% lower latency than the tested TurboQuant codec.
cs.LG / 67 / 2610.11247
Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals
Lei Zhao, Qichao Zhao, Bowen Zuo, Qishi Zhan
cs.LG · cs.CL
Abstract
On-policy distillation (OPD) enables effective capability transfer between language models, yet the mechanisms underlying its failures are not fully understood. Across code generation and mathematical reasoning, OPD with larger-scale teachers exhibits early loss plateaus, with an average final loss reduction of 25.1% after 200 updates, compared with 96.2% for self-RL teachers, obtained by further reinforcement learning (RL) training of the initial student. To understand this difference, we analyze OPD as an idealized continuous-time dynamical system in the small-learning-rate limit. Our training-log diagnostics associate these plateaus with an early decline in a gradient-based learning-signal proxy while substantial loss remains; these measurements do not establish why the underlying gradient weakens. We further prove a local recovery guarantee for teachers sufficiently close to the initial student in a shared parameterization under regularity conditions, offering a conditional explanation for the success of self-RL teachers in our experiments. Across runs with and without loss plateaus, we observe small relative parameter changes (0.025-0.098%) and high similarity between the student's representations before and after OPD (linear CKA $>0.98$ across layers). These observations suggest that limited representation adaptation may contribute to learning-signal collapse, a hypothesis that remains to be tested. Code is available at https://github.com/leizhao7/opd-learning-signals.
cs.LG / 68 / 2610.11252
Neuro-Memory Fuzzy Inference System for Mimicking Human-like Car Following Behavior
Nazmul Haque, Md Asif Raihan. Md. Hadiuzzaman
cs.LG · cs.AI
Abstract
This study presents the Neuro-Memory Fuzzy Inference System (NeMeFIS), a hierarchical machine learning architecture that asymmetrically models acceleration and deceleration in car following behavior by integrating five human memory types procedural, working, episodic, semantic, and declarative. By linking external variables to memory functions via metaheuristics and validating them through factor and p-value analyses, NeMeFIS uncovers latent cognitive influences across Arterial, Collector, and Rural Highway corridors for different types of vehicles. Results from 54 different trained models emphasize cognitive thresholds shaped by driver perception limits and cognitive load. The trained NeMeFIS models outperform traditional statistical and conventional machine learning models in replicating realistic driving behavior, including comparisons with Linear Regression, ANFIS, and LSTM architectures. Fuzzy rule analysis reveals that declarative memory demands the highest rule, especially during deceleration, indicating complex braking decisions. Procedural memory drives acceleration, while semantic and declarative memory guide deceleration. Risk perception also emerges as a key factor, particularly on urban roads. Validated on both heterogeneous and homogeneous datasets, NeMeFIS offers a robust framework for modeling driver cognition. The findings support psychotherapeutic applications and the development of adaptive, human-like decision systems in Connected and Autonomous Vehicles (CAVs) to enhance traffic safety.
cs.LG / 69 / 2610.11257
Residual spectral instabilities in representation learning
Zhen Li
cs.LG · cond-mat.stat-mech
Abstract
Learned representations can lose latent degrees of freedom successively, suggesting a cascade of transitions whose underlying stability principle remains unclear. Here we formulate dimension-wise posterior collapse in variational autoencoder (VAE) as a fluctuation theory around partially collapsed states. Interpreting the negative evidence lower bound as an effective free energy, its quadratic expansion defines a Gaussian theory whose Hessian acts as a mass matrix for latent fluctuations. We show that the collapsed directions form an invariant fluctuation sector and derive its exact mass spectrum in terms of a conditional residual operator. A local reactivation direction lowers the free energy when the decoder variance falls below the residual spectral upper edge, with equality marking marginality. The criterion recovers principal component thresholds in the linear Gaussian VAE limit. Viewed in reverse along continuously connected branches, the reactivation boundary provides a local criterion for successive collapse. Numerical continuation experiments show successive loss of latent dimensions near these spectral marginalities. These results support a spectral cascade interpretation governed by residual information left unexplained by the surviving representation.
cs.LG / 70 / 2610.11309
From Geometry to Generalization: Why Row Normalization Can Beat Adam and Muon
Jihwan Kim, Dogyoon Song, Chulhee Yun
cs.LG · cs.AI · math.OC · stat.ML
Abstract
Different optimizers can fit the same training data while selecting classifiers with substantially different geometries, but whether this difference provably affects population performance remains unclear. We show that row-wise normalization can achieve strictly higher population accuracy than full-batch Adam, a proxy for random-reshuffling Adam, and exact-SVD Muon in high-dimensional multiclass classification. Under an isotropic Gaussian-cloud data model, this advantage arises because row normalization's class-wise Euclidean geometry asymptotically preserves the population decision-boundary directions, whereas Adam's coordinate-wise geometry and Muon's spectral geometry introduce nonvanishing distortions. Beyond isotropy, the advantage persists for full-batch training on class means with independently oriented class-mean and test-noise covariances. It holds for power-law spectra with class-mean exponent below one, even under heavily anisotropic test noise. When both covariances are diagonal and sufficiently close, the advantage over Adam can reverse, while applying the same random rotation to both restores it by changing only their alignment with Adam's coordinate axes. Synthetic and last-layer language-model experiments support the predicted advantage.
cs.LG / 71 / 2610.11313
NP-Hardness of Minimizing Neurons in Two-Hidden-Layer ReLU Neural Networks
Sangrock Lee
cs.LG
Abstract
A fundamental question in neural network architecture optimization is whether the minimum hidden-neuron count required to approximate a target function within a prescribed tolerance can be computed efficiently. This paper resolves this question for two-hidden-layer ReLU networks under an $L^p(\mathbb{R}^d,\mathbb{R}^m)$ approximation constraint. For every fixed $d \ge 1$, $m \ge 1$, and $1 \le p < \infty$, we prove that computing the optimum exactly is NP-hard. The result holds even when the target is represented by a rational ReLU network whose realization is nonzero, componentwise nonnegative, compactly supported, globally Lipschitz, and continuous piecewise affine. The polynomial-time reduction from 3-SAT produces an architecture gap in which unsatisfiable formulas yield an optimum of zero, whereas satisfiable formulas yield an optimum of at least $d+2$. The proof constructs compactly supported polyhedral frustum functions realized by two-hidden-layer ReLU networks and establishes the $L^p$-density of finite linear combinations of box-frustum functions. The results offer theoretical justification for employing heuristic approximation methods in the design of ReLU neural networks, illustrating that attaining a minimal configuration within polynomial time is computationally unachievable.
cs.LG / 72 / 2610.11394
CoPoE: Multimodal Fusion via Decomposable Disease-Coordinate Product-of-Experts for Missing-Modality Alzheimer's Diagnosis
Chihun An, Ikbeom Jang
cs.LG · cs.CV
Abstract
Multimodal Alzheimer's disease (AD) diagnosis benefits from integrating heterogeneous clinical, imaging, genomic, and biomarker evidence, but clinical cohorts frequently suffer from irregular modality missingness. Existing fusion methods often synthesize absent inputs, risking the introduction of artificial surrogates, or pool available signals into uninterpretable latent spaces. We present CoPoE (Disease-Coordinate Product-of-Experts), a disease-coordinate framework that maps multimodal evidence into a structured latent space partitioned into four distinct biological and clinical axes: genetic Risk, molecular Pathology, Neurodegeneration, and clinical Stage (R/P/N/S). Each observed modality parameterizes a diagonal Gaussian expert over the full RPNS vector, and a masked Product-of-Experts architecture fuses only the available modalities. Consequently, absent modalities add no factor to the fusion path, allowing the network to preserve a robust, decomposable posterior for any non-empty modality subset without synthetic imputation in the RPNS path. Through extensive missing-modality experiments on the ADNI dataset, CoPoE achieves the best all-modality performance and the highest mean AUROC across all 15 observed-subset evaluations among standardized missing-modality fusion baselines under a shared non-PET ADNI embedding benchmark, while substantially improving raw-probability ECE, Brier score, and NLL. Furthermore, PET-supervised probing shows evidence enrichment within the pathology (P) block under full modalities, with tau-related signal retained even when direct fluid biospecimen inputs are withheld. Our code is available at https://github.com/labhai/CoPoE.
cs.LG / 73 / 2610.11403
Learning from Hetero Density for Cryo-EM Protein Reconstruction
Xu Han, Chaozhuo Li, Xiaowei Yuan, Yuancheng Sun, Kang Liu, Qiwei Ye
cs.LG
Abstract
Reconstructing protein structures from cryo-electron microscopy (cryo-EM) maps is essential for understanding macromolecular assemblies. Although learning-based methods have improved protein reconstruction, information from hetero components remains underused. Our analysis finds both false predictions and reference protein sites near hetero components; filtering nearby candidates can improve or impair chain construction. We introduce CryoCue, a framework that uses hetero information to guide protein reconstruction. An anchor-supervised detector learns hetero representations across five component classes. Multiscale hetero features guide backbone localization, while predicted hetero candidates condition structure refinement through their class, confidence, and frame-relative geometry. Experiments show that CryoCue improves backbone localization near hetero components and achieves more accurate protein structure reconstruction.
cs.LG / 74 / 2610.11408
MC-TRCM: Observation-Aware Recursive Fusion for Incomplete Mobile and Wearable Mental-Health Feature Views
Wentao Wang, Lifeng Han, Zining Ren, Hengyu Zhong, Guangyu Zou
cs.LG
Abstract
Public mobile and wearable mental-health datasets often provide summarized feature tables rather than synchronized raw sensor streams. In these releases, each anchor corresponds to a survey or label time and may combine phone or wearable summaries, prior symptom scores, demographics, clinical variables, and source-availability indicators. We propose the Modality-Conditioned Temporal Recursive Context Model (MC-TRCM), which preserves each feature source as a separate token and incorporates missingness as part of the input context. Observed sources are encoded with values and missingness summaries, absent sources use learned absence tokens, dataset and task embeddings condition fusion, and a recursive prediction head refines each output over validation-selected steps. We evaluated MC-TRCM on six predefined endpoints from DepreST-CAT and Prediction of Severity Change-Depression (PSYCHE-D) using participant-level splits and validation-only model selection. MC-TRCM achieved the lowest mean absolute error on DepreST-CAT Patient Health Questionnaire-9 (PHQ-9) and Generalized Anxiety Disorder-7 (GAD-7) severity, improving over the best tabular reference by 0.181 and 0.217 scale points. Classification endpoints showed task-dependent behavior: MC-TRCM matched the best rounded GAD-7 category balanced accuracy, was numerically highest by 0.002 balanced-accuracy points on PSYCHE-D multiclass prediction, and remained close to the strongest references on PHQ-9 category and PSYCHE-D binary prediction. Ablations support Feature-wise Linear Modulation, absence tokens, missingness projections, and recursive refinement, while calibration and feature-source controls characterize endpoint behavior. Our code is available at https://github.com/Botwwt/MC-TRCM.
cs.LG / 75 / 2610.11423
Policy Alignment: New Signals for Membership Auditing in On-Policy Distillation
Yilong Yang, Wenzhuo Shang, Yule Liu, Jiale Teng, Zhuo Ma
cs.LG
Abstract
On-policy distillation (OPD) trains a student model by aligning its policy with a teacher model on trajectories generated by the student model itself. Through this process, the student policy moves toward the teacher on the prompts used for distillation. However, these prompts are often private and costly, creating a need for prompt-level membership auditing. Existing methods mainly rely on likelihood-based confidence signals or student policy drift between checkpoints, but they do not capture the teacher-induced direction of the student update. In this paper, we propose Policy Alignment Membership Auditing (PAMA), a new auditing framework tailored for OPD. Our key observation is that a member prompt directly contributes to the teacher-guided policy update, while a non-member prompt only experiences indirect effects through cross-prompt generalization. Based on this directional trace, PAMA measures whether the student update moves toward reducing the teacher loss on a candidate prompt. Specifically, we introduce Teacher Alignment Gain (TAG) to estimate the teacher-aligned update direction from model outputs, and further combine it with student drift and uncertainty alignment signals for reliable membership auditing. We evaluate PAMA on six datasets and three teacher-student model families. On MATH, the primary evaluation benchmark, PAMA achieves AUC values of 0.791--0.941, improving AUC by 14.6--20.6% over state-of-the-art baselines.
cs.LG / 76 / 2610.11442
MotiveMob: Motivation as Semantic Action for Closed-Loop Human Mobility Generation
Mengkun Gao, Zengqing Wu, Renhe Jiang, Jiawei Wang, Yusong Wang, Chuang Yang, Shuyuan Zheng, Makoto Onizuka, Chuan Xiao
cs.LG
Abstract
Human mobility generation, an important task in urban research, synthesizes trajectory data for urban planning and transportation management. Human mobility can be characterized as a "why-where-when" decision process: people form an intention to move and then determine where and when the corresponding activity will take place. Trajectory generation under user-level and temporal distribution shifts may benefit from explicitly modeling this decision structure. However, many existing human mobility generation methods either represent behavioral intent at a coarse granularity, such as a daily plan or a trajectory-level description, or directly predict future locations without explicitly reasoning about a possible motivation for each movement step. We introduce MotiveMob, a motivation-driven autoregressive framework for human mobility generation that first forms a hypothesis about why the next movement may occur and then jointly generates where and when it may occur. At each step, a motivation predictor conditions on the current mobility state, a long-term behavioral report, and the mobility history to infer a plausible motivation or determine whether the trajectory should terminate. Given the hypothesized motivation, a state predictor grounds it in a candidate next location and arrival time. The candidate then undergoes speed-feasibility and repetition checks before being fed back for the next decision. We evaluate MotiveMob under distribution shifts involving unseen users and unseen temporal periods, including seasonal changes and the substantial behavioral disruption caused by the COVID-19 pandemic. Experiments show that MotiveMob consistently achieves better distributional fidelity than competitive pretraining-based and prompting-based methods under user-level and temporal distribution shifts, demonstrating robust generalization to out-of-distribution mobility patterns.
cs.LG / 77 / 2610.11454
Zatom-2: Multitask Pretraining on Atomistic Data for Generative Modeling across Domains
Miruna Cretu, Alex Abrudan, Antonia Panescu, Tynan Perez, Rishabh Anand, N. Benjamin Erichson, Michael W. Mahoney, Samuel Blau, Joseph Jacobson, Rafael Gómez-Bombarelli, Rex Ying, Tuomas Knowles, Pietro Liò, Alex Morehead
cs.LG · cs.AI
Abstract
Unified atomistic modeling has the potential to accelerate discovery in chemistry, materials science, and biology by bridging data-rich chemical domains and data-scarce biological contexts. However, existing generative approaches to atomistic modeling remain highly specialized to scientific disciplines (chemistry vs. biology) or do not leverage both high-volume organic (molecule) and inorganic (material) data for general-purpose pretraining. To this end, we introduce Zatom-2, an atomistic generative model pretrained on approximately five million structures from the OMol25 and OMat24 electronic structure datasets. Zatom-2 features a multiscale Transformer architecture coupled with conditional flow matching that supports force conditioning and foundational pretraining tasks such as generation, structure prediction, and prediction of molecular and material energies and forces. Empirically, Zatom-2 achieves better molecular distribution fidelity than Zatom-1 and achieves strong performance on existing molecule and material generation benchmarks. Zatom-2 demonstrates the ability to control sample generation across low- and high-force regimes, and enhances protein generation in a low-data setting through joint generative-predictive pretraining and transfer learning, increasing protein backbone designability in a length extrapolation setting from 67.8% without pretraining to 74.8% after finetuning on 2,000 protein domains.
cs.LG / 78 / 2610.11458
Generative Adversarial Loops
Kislay Aditya Oj, Nidhi Jain, Sri Surya Varma Datla, Priyanka Jayaswal, Kumar Krishna Agrawal, Aditya Desai
cs.LG · cs.AI
Abstract
AI research progress can be viewed as the interaction between two processes: benchmark creation and method discovery. Historically, both were driven by human intelligence. However, recent advances in AI have accelerated automated method discovery, while automated benchmark creation has received comparatively less attention. To enable self-advancing systems, we propose Generative Adversarial Loop (GAL), a generator-discriminator framework alternating between two agentic searches: (1) a discriminator that generates adversarial data to expose weaknesses in current systems, and (2) a generator that discovers algorithms to overcome them. We apply this framework to approximation algorithms for efficient inference. Unlike existing auto research systems, which primarily focus on algorithm discovery, GAL introduces a discriminator agent that automates goalpost setting by continually searching for weaknesses in the current algorithm. We demonstrate adversarial data generation across four tasks: KV compression, sparse video generation, sparse attention, and context extension, where the discriminator identifies weaknesses in state of the art techniques. We further show that GAL enables autonomous improvement, with newly discovered algorithms improving not only on adversarially generated data, but also on established benchmarks. Specifically, GAL improves CompactorPress on KV compression with Qwen3-4B at 4x, raising performance on the discriminator dataset from 0.35 to 0.97, while also outperforming RULER-HARD (+0.77 pts). For context extension, GAL boosts Dual Chunk Attention from 0.20 to 0.90 on the discriminator dataset, while yielding gains on standard benchmarks(ScienceFiction (+6 pts) and PG19 32K (-0.33 PPL)). GAL thus provides a path toward autonomous goalpost setting and algorithmic improvement, where AI systems continually discover their own weaknesses and develop methods to overcome them.
cs.LG / 79 / 2610.11506
Compactness and Consistency: A Conjoint Framework for Deep Graph Clustering
Wei Ju, Siyu Yi, Kangjie Zheng, Yifan Wang, Ziyue Qiao, Li Shen, Yongdao Zhou, Xiaochun Cao, Jiancheng Lv
cs.LG · cs.AI · cs.IR · cs.SI
Abstract
Graph clustering is a fundamental task in data analysis, aiming at grouping nodes with similar characteristics in the graph into clusters. This problem has been widely explored using graph neural networks (GNNs) due to their ability to leverage node attributes and graph topology for effective cluster assignments. However, representations learned through GNNs typically struggle to capture global relationships between nodes via local message-passing mechanisms. Moreover, the redundancy and noise inherently present in graph data may easily result in node representations lacking compactness and robustness. To address these issues, we propose a conjoint framework CoCo, which captures compactness and consistency in the learned node representations for deep graph clustering. Technically, our CoCo leverages graph convolutional filters to learn robust node representations from both local and global views, and then encodes them into low-rank compact embeddings, thus effectively removing the redundancy and noise as well as uncovering the intrinsic underlying structure. To further enrich the node semantics, we develop a consistency learning strategy based on compact embeddings to facilitate knowledge transfer from the two perspectives. Our experimental results indicate that our CoCo outperforms state-of-the-art counterparts on various datasets.
cs.LG / 80 / 2610.11523
When to Intervene? State-Aware Sparse Manipulation in Federated Reinforcement Learning
Shutong Zheng, Sijia Chen
cs.LG
Abstract
Federated reinforcement learning (FRL) enables distributed agents to collaboratively train decision-making policies, but its decentralized training process also exposes global policy learning to Byzantine manipulation. Existing poisoning attacks primarily focus on how to construct malicious updates, while trajectory-level intervention timing remains largely implicit. In sequential decision making, however, where an intervention is applied can alter subsequent trajectories and learning signals. Through controlled experiments, we find that changing the selected trajectory states materially alters attack efficacy even when the malicious-update construction is fixed. We therefore identify when as a distinct attack dimension and introduce the Viability-constrained Behavioral Steering Attack (V-BSA), which uses local policy uncertainty to select sparse intervention states and applies envelope-constrained behavioral steering. Across discrete-action benchmarks, V-BSA achieves substantial degradation against robust aggregators and ensemble defenses with only a fraction of the interventions used by dense poisoning, while revealing task- and aggregation-dependent boundaries. Overall, our results highlight intervention timing as a distinct dimension of sequential robustness in FRL. The code is available at https://github.com/Yodeesy/V-BSA
cs.LG / 81 / 2610.11548
Conditional Transfer from Controlled Pretraining Mixtures to Code
Ohad Rubin
cs.LG
Abstract
Synthetic tasks are increasingly used both as probes of language-model capability and as pretraining data. Both uses are often justified by loss reduction: falling loss is treated as informative, and faster loss reduction with more sampling as evidence that a task is worth sampling. We separate three signals. A task is diagnostic when its loss tracks global pretraining progress; it is teachable when its loss responds to its own token budget; and a data source transfers when including it improves a downstream target. We study controlled pretraining in which 70% of the corpus is fixed general Python and the remaining 30% is a simplex over three source families: OpenCodeInstruct, a curated suite of 12 code-adjacent synthetic tasks, and 15 literature-derived probe tasks. Across a task-budget sweep we detect teachability for 14 of 27 tasks, with a sharp asymmetry between the two synthetic families (10/12 curated versus 4/15 literature-derived). Teachability and downstream transfer give different rankings. On the mixture-simplex edge between the curated suite and OpenCodeInstruct, HumanEval pass@20 after a fixed fine-tuning stage rises from 15.9 at pure curated data to its highest observed value, 22.6, at a mixture that is 75% OpenCodeInstruct, then falls to 19.5 at pure OpenCodeInstruct. Curated synthetic data therefore has conditional value: it contributes as a limited share of a mixture that a target-aligned source still dominates. Finally, a loss-based adaptive scheduler exposes the mismatch between residual loss reducibility and downstream transfer. Across three 60k-step free-ratio runs, Ado drives the OpenCodeInstruct share below 5% within the first 5k steps and to 1.2--1.4% by the end of training, and underperforms its matched fixed-mixture controls by 2.4--11.0 percentage points. Optimizing near-term task-loss reduction moves the mixture away from the region that transfers.
cs.LG / 82 / 2610.11549
$C_4$-Equivariant Flow Matching on Anisotropic Power-Diagram Graphs for Microstructure Generation
Dawid Lipinski, Jixiang Qing, Henry Moss
cs.LG
Abstract
Acquiring realistic microstructure data through Electron Backscatter Diffraction (EBSD) is costly and time consuming, often relying on specialised equipment. As microstructures strongly influence material properties, generating realistic samples is essential for modelling the behaviour of polycrystalline materials. We introduce a generative model for synthesising realistic polycrystalline microstructures using flow matching and graph neural networks. By representing microstructures as anisotropic power diagrams, our model learns a compact geometric parametrisation and can render generated samples at arbitrary pixel resolution. A $C_4$-equivariant architecture incorporates rotational symmetry directly into the model, ensuring that rotations of the input noise produce corresponding rotations of the generated microstructure. We also demonstrate how training-free guidance can be used to generate complex microstructures, based on user defined objective function. In particular, we generate microstructures resembling a copper weld, cast metal slab, 3D-printed stainless steel and heterogeneous lamella titanium.
cs.LG / 83 / 2610.11551
New Lower Bound and Upper Bounds on the Regret for Online Sparse Linear Regression
Xiaofeng Cao, Junfan Li, Langzhang Liang, Mingwei Xu, Xiao Zhang
cs.LG
Abstract
We study online sparse linear regression (OSLR) where any algorithm is restricted to accessing only $b$ out of $d$ attributes per instance for prediction and $b_0\geq 0$ additional attributes after prediction, which was proved to be NP-hard. Previous work focused on designing computationally efficient algorithms under regularity assumptions, but did not characterize its information theoretic complexity. In this work, we give the first lower bound on the minimax regret of OSLR and design algorithms with better upper bounds without regularity assumptions. We characterize how minimax regret scales with problem-dependent parameters, capturing the information theoretic complexity of OSLR.
cs.LG / 84 / 2610.11555
Best of Both Worlds in Federated LSA: Speedup When Possible, Personalization Always
Safwan Labbi, Paul Mangold, Eric Moulines
cs.LG
Abstract
We study personalized federated linear stochastic approximation (LSA), a framework which notably encompass personalized temporal difference learning. In this setting, heterogeneous agents collaborate to solve distinct linear fixed-point equations, each corresponding to an agent-specific learning problem. A central open question in personalized learning is whether a single method can adapt to an unknown level of heterogeneity by converging to each agent's personalized solution in all regimes while achieving a linear speedup in the number of agents when their learning problems are sufficiently similar. We answer this question affirmatively by introducing PF-LSA, a minimalist algorithm that mixes each agent's local stochastic update with the average update across agents, at no additional computational cost relative to standard federated methods. We prove that PF-LSA, achieves best-of-both-worlds guarantees without any prior knowledge on the level of heterogeneity. Our analysis is based on a sharp decomposition of the error into consensus and disagreement components. The consensus error decays rapidly, whereas the disagreement error decays more slowly but becomes negligible in low-heterogeneity regimes.
cs.LG / 85 / 2610.11567
Sera: Semantic Representation Aggregation for Reliable and Interpretable Battery Health Forecasting
Jiawei Li, Fang Liu, Wei Zhang, Zuming Liu, Man-Fai Ng, Zhi Wei Seh
cs.LG · cs.AI
Abstract
Battery state of health (SoH) forecasting is important for battery management, but remains challenging due to nonlinear degradation and heterogeneity across batteries. Existing data-driven approaches primarily use temporal models to learn from numerical battery time series, and higher-level degradation characteristics are often not explicitly represented. These characteristics, however, can provide degradation guidance to support reliable forecasting and make the influence of degradation more interpretable. In this paper, we propose \textsc{Sera}, a \underline{se}mantic \underline{r}epresentation \underline{a}ggregation framework that complements temporal modelling with degradation semantics. Guided by battery domain expertise, \textsc{Sera} extracts degradation semantics from time series and constructs two complementary representations using rule-based knowledge and LLM-based interpretation. The representations are independently encoded and integrated with the representation learned by temporal models through gated aggregations. Experiments on the mainstream benchmark across multiple prediction horizons and different temporal models show that \textsc{Sera} consistently improves forecasting performance, achieving up to a 37.3\% reduction in prediction error over the temporal baseline and enhanced generalizability. Counterfactual analysis examines how forecasts respond to changes in degradation semantics to assess interpretability. The results show that prediction responses are consistent with the meanings of key degradation descriptors across tested horizons. Together, these findings demonstrate that structured degradation semantics and effective aggregation can improve forecasting accuracy and support reliable and interpretable battery health forecasting for advanced battery management.
cs.LG / 86 / 2610.11580
Uncertainty-Aware Optimization for Physics-Aware Highway Trajectory Prediction
Aanchal Rajesh Chugh, Sebastian Dorn
cs.LG
Abstract
Accurate trajectory forecasting and well-defined predictive uncertainty are crucial for reliable, safety-critical applications such as autonomous driving. Most trajectory prediction approaches provide point estimates only, while uncertainty-aware approaches typically quantify uncertainty only in the trajectory space. In physics-aware approaches, uncertainty in the predicted motion variables should be explicitly modeled and propagated through the vehicle dynamics. Otherwise, the resulting trajectory-space uncertainty may not fully reflect the variability introduced by the underlying motion prediction. Therefore, in this work, uncertainty-aware extensions of X-TRACK (X-TRACK-DE and X-TRACK-MCD), a physics-aware trajectory prediction framework, are proposed. The proposed framework predicts future vehicle motion variables and models both aleatoric and epistemic uncertainties by propagating motion space uncertainty to trajectory space. Additionally, conformal prediction is applied to the trajectory space predictive covariance to construct uncertainty regions targeting a desired marginal coverage level. Evaluation on the highD dataset shows that X-TRACK-DE improves trajectory prediction accuracy over the deterministic baseline, while both uncertainty-aware variants provide predictive uncertainty that can be conformally calibrated to the desired marginal coverage level.
cs.LG / 87 / 2610.11605
NanoProof: Open and Efficient Automated Theorem Proving in Lean 4
Matěj Kripner, Milan Straka
cs.LG · cs.AI
Abstract
We introduce NanoProof, to our knowledge the first factorized execution-guided theorem prover in Lean 4 whose training data, extraction tooling, training pipeline, and weights are all released, making it end-to-end reproducible using open-source resources. To this end, we build and release a dataset of structured proof trees, as well as a tool for programmatic interaction and data extraction within the Lean 4 formal verifier. To support sustainable research, we focus on compute efficiency to facilitate accessible training and evaluation. NanoProof achieves 50.8% pass@16 on MiniF2F-Test, exceeding the two closest systems of its class, HyperTree Proof Search and ABEL, at roughly 90x and 7x less compute, and using more than four orders of magnitude less compute than AlphaProof. Stronger open-weight provers exist, but they are fine-tuned from large pretrained language models and release neither training data nor pipeline; NanoProof shows that the factorized execution-guided class of provers can be rebuilt from scratch with modest resources.
cs.LG / 88 / 2610.11621
Constructing Structured Decision Sources for Consensus-Based Pseudo-Label Learning
Long Wang
cs.LG
Abstract
Consensus can make pseudo-label learning more reliable, but only when its predictors contribute genuinely different evidence. Multiple models that repeat the same boundary provide additional votes without additional information. We address this problem by con structing decision sources through controlled changes to within-class structure. Starting from a shared graph representation, we vary center granularity and neighborhood mixing, reproduce each resulting source to test its stability, and select a complementary subset using node pair coassignment. Unanimous predictions from the selected sources are then ranked for student training. On the public fixed splits of Cora, CiteSeer, and PubMed, evaluated with five random seeds, the constructed sources improve fixed-budget training pseudo-label precision by 1.19 to 4.39 percentage points over three conventionally initialized GCN sources. Under matched structural filters, three-source consensus is more precise than each constituent source in all 45 dataset slot eed comparisons. The gains are strongest in pseudo-label quality: downstream accuracy remains competitive but does not lead on every dataset. These results identify source construction rather than model count alone as an important design problem for consensus-based pseudo-label learning.
cs.LG / 89 / 2610.11627
Beyond Action Entropy: Quotient-Space Exploration for Genome-Scale Metabolic Model Repair
Xuan Gong, Hanbo Huang, Wenbin Dai, Jing Wang, Lei Bai, Xiang Xiao, Weishu Zhao, Shiyu Liang
cs.LG · q-bio.MN
Abstract
Repairing scientific models from functional observations differs fundamentally from supervised prediction: feedback may certify a solution without revealing which structural correction is responsible. We study this setting for genome-scale metabolic model (GEM) repair, where multiple reaction edits can explain the same phenotypes and many apparently distinct edits correspond to the same biological mechanism. This many-to-one structure creates a hidden failure mode for conventional exploration: diversity in the output space need not translate into diversity of scientific hypotheses. We introduce QuotientPO, which collapses equivalent repairs into canonical mechanisms and optimizes exploration directly over the resulting quotient space. To make quotient exploration informative under finite rollouts, we derive a kernelized Rényi estimator that resolves graded crowding among distinct repair cores beyond coarse exact-match counts. On 2,212 held-out GEMs, QuotientPO improves Success@32 from 17.93% to 20.10% (+12.1% relative) while consistently increasing distinct successful-core discovery under the same sampling budget. These results establish quotient-space exploration as a principled approach to mechanism-level discovery under verifier-induced equivalence.
cs.LG / 90 / 2610.11665
Evi-VN: Hard Region Guided Virtual Node Evidence Injection for GNN-Based Fraud Detection
Jiran Tao, Yifan Wu, Binyan Jiang
cs.LG
Abstract
Online platforms contain growing numbers of bots, deceptive reviewers, and scam accounts that imitate legitimate users. Such camouflage blurs graph neighborhoods and behavioral attributes, making it difficult for graph neural networks (GNNs) to distinguish both well-disguised fraudsters and legitimate users. Across diverse GNNs, we observe overlapping errors on a shared hard region, suggesting the presence of latent fraud evidence that graph topologies and standard features fail to capture. Fraud-specific GNNs can mitigate particular graph pathologies, yet they still make limited use of heterogeneous evidence such as structured records, text, images, and audio; uniform multimodal fusion may also disturb nodes already handled reliably by the graph. We propose Evi-VN to learn and correct these shared blind spots rather than build another fraud detector. To our knowledge, Evi-VN is the first graph fraud detection framework to use feature isolated evidence chains to correct hard regions shared across GNNs. Its evidence chains connect behavior, content, and context across structured, textual, visual, and acoustic sources, helping expose camouflage that graph neighborhoods may miss. Crucially, Evi-VN selectively applies this evidence only to likely hard samples via virtual class nodes, preserving both the reliable predictions and the input design of existing GNNs. Shared hard regions also let Evi-VN enhance generic, fraud-specific, and unseen GNNs even with imperfect evidence models. Experiments across bot, fake-review, refund-evidence, and telecom-fraud tasks validate these advantages.
cs.LG / 91 / 2610.11683
Camera-Noise Residuals for Face-Swap Detection: Redundant, Not Complementary, and Why
Danil Davydov, Bader Rasheed, Dmitriy Vatolin
cs.LG
Abstract
Fusing a learned camera-noise fingerprint with an RGB appearance backbone is an appealing route to generator-independent deepfake detection, because the noise residual is grounded in image-formation physics rather than in the texture statistics of a particular generator. We test, on FaceForensics++, whether a Noiseprint++ residual channel carries information \emph{complementary} to an RGB Xception backbone for face-swap detection. A three-model ablation (RGB-only, residual-only, late-fusion) shows that fusion does not improve over RGB alone and that the residual branch alone is near chance. A seven-level bottleneck diagnostic localizes the cause: the noise maps do carry a discriminative signal, but it is statistical---carried by the per-sample first and second moments (mean, variance, energy) of the residual---and the per-sample \texttt{InstanceNorm} layer placed at the noise-branch input, following the TruFor template, standardizes exactly those moments away (five-fold cross-validated AUC drops from $0.747$ to $0.554$). A context-crop control rules out cropping geometry, and two fixed-fusion variants that remove the bottleneck recover the statistical signal yet still fail to beat RGB on every dataset. We conclude that, on this manipulation distribution, the noise residual is redundant with RGB rather than complementary, and we give concrete guidance for practitioners adopting noise-residual fusion for face-swap detection.
cs.LG / 92 / 2610.11692
Can Jev be Your Q or Policy in Reinforcement Learning?
Yi Ma, Tianpei Yang, Yaodong Yang, Weixun Wang, Hongyao Tang
cs.LG · cs.AI
Abstract
Foundation models supply reinforcement learning (RL) with priors that mitigate its longstanding weaknesses in sample efficiency and transfer, but their token-by-token generation makes queries sequential and costly. Jev, a recently released decision model, generates nothing and returns calibrated, typed answers in a single forward pass. Existing work studies foundation models in RL either as models to be trained or as generators to be prompted, and Jev belongs to neither category, having so far served only as a black box in single domains. How well such a model decides on its own in RL environments, and how it can improve RL as a component of training, therefore remain unaddressed. To this end, in this paper we first examine the requirements that the objects of an RL system place on the answers they consume, and establish that Jev can fulfill all of them except the cardinal use of a value function. The remaining objects form positions that admit several roles each. We then construct algorithms that employ Jev at three of these positions, as a reference policy, an exploration judge, and a replay rater, to improve sample efficiency, exploration, and learning performance. Across nine MiniGrid tasks and three Atari games, training with Jev outperforms a standard RL learner, including where the learner makes no progress alone, while the model itself remains untrained. To our knowledge, we present the first use of Jev within the RL learning process and establish a frozen decision model as a usable component of RL training, inviting further exploration of how Jev and other advanced decision models can improve RL.
cs.LG / 93 / 2610.11697
What is the goal of unsupervised machine learning?
Aapo Hyvärinen
cs.LG
Abstract
Unsupervised learning is one of the main branches of machine learning. Here I argue that unlike the other branches of machine learning (supervised and reinforcement learning), unsupervised learning is a rather heterogenous field that can serve several different goals. It seems futile to try to define one single goal for unsupervised learning. I identify four different goals for unsupervised learning: 1) Estimating the distribution, 2) Generating new data points, 3) Extracting features for downstream tasks, and 4) Understanding the data.
cs.LG / 94 / 2610.11706
Does an Illumination Prior Help Face-Swap Detection? A Controlled Study of Temporal Self-Blended Images
Danil Davydov, Bader Rasheed, Dmitriy Vatolin
cs.LG
Abstract
Self-blended images are widely used to train face-swap detectors, but primarily capture blending artifacts. We investigate whether adding illumination inconsistencies improves detection. Temporal Self-Blended Images (T-SBI) transfer lighting statistics between frames of the same video, with the mismatch controlled by luminance difference (ΔL). Using five training regimes and a three-seed comparison of high- and low-ΔL training, we find no evidence of illumination-specific improvements. AUC differences remain within seed variability across four datasets, and an analysis of 506,328 attribute-binned samples shows no preferential reduction in errors under harsh lighting. Instead, T-SBI shifts prediction scores, changing optimal thresholds by approximately 0.34 on FaceForensics++ and 0.30 on Celeb-DF, making comparisons at a fixed threshold misleading. However, T-SBI improves robustness to heavy JPEG compression on DFDC (AUC 0.780 versus 0.696), potentially reflecting greater reliance on low-frequency cues. These findings highlight the importance of evaluating training methods against their intended targets and accounting for threshold effects.
cs.LG / 95 / 2610.11716
4-Tensor Attention Model for Semantic Physical Reality
Jongwook Kim, Sangheon Yun
cs.LG · cs.CL
Abstract
We describe a 4-tensor attention model that predicts the next semantic state of a scene, for video generation and robot planning. A window of states has positions (x, t) and two fibers, a semantic fiber and a temporal-context fiber, and one softmax normalizes attention jointly over the window. Frames and an agent's situation are written as those states; the encoder, the renderer, and the planner remain outside the update. To test the update on its own, we train on ROCStories, where each window poses the same next-sentence task at the semantic layer. On the validation split, with one seed per setting, the last-sentence cross-entropy on the three matched settings is lower for the 4-tensor model than for a free-running one-dimensional transformer by 5.3% at H=2, L=2, by 2.6% at H=4, L=2, and by 2.4% at H=4, L=3. At H=4, L=2 the parameter counts are nearly the same, 172.5M and 175.9M. On the same two GPUs that 4-tensor run finished in 2.4 hours and the baseline run in 45.2 hours; the baseline is trained by free-running decoding, one sequential forward pass per target token.
cs.LG / 96 / 2610.11724
Addressing Overcommitment in the Reasoning of Gendered Economic Memes under Multimodal Ambiguity
Kushal Kanwar, Dushyant Singh Chauhan, Kapil Rana, Gopendra Vikram Singh, Nils Lukas
cs.LG
Abstract
Multimodal meme understanding is increasingly used to analyze socially sensitive content, yet existing models often exhibit biased behavior when interpreting economic dependence and social roles under ambiguity. Many memes express economic relationships through sparse text or symbolic visual cues, providing insufficient evidence for gendered attribution. In such underspecified settings, models tend to rely on pretraining correlations, leading to hallucinated and stereotypical economic role assignments. In this work, we study gendered economic dependence in image-text memes through the lens of contextual sufficiency and identify epistemic overcommitment-inferring roles without adequate evidence-as a primary source of bias. We propose CGER-Net, a context-grounded multimodal framework that estimates whether the input provides sufficient evidence for gendered economic reasoning and applies evidence-gated inference to enable confident attribution when cues are explicit while favoring principled abstention otherwise. We evaluate CGER-Net on EconMeme-GE, a curated dataset of image-text memes annotated as Men, Women, Neutral, or Ambiguous. Across strong contemporary multimodal baselines, CGER-Net reduces Gender Overcommitment Rate by up to 44% on ambiguous instances while maintaining comparable accuracy on unambiguous cases. Human evaluation further shows that 79% of generated rationales are judged as epistemically aligned with the available evidence. These results highlight the importance of modeling when not to infer for reliable and responsible multimodal analysis.
cs.LG / 97 / 2610.11730
Spectral Weight Decay: Inducing Low-Rank Structure in Neural Network Weights
Dmitrii Andriianov, Andrey Veprikov, Aleksandr Beznosikov
cs.LG · math.OC
Abstract
Standard weight decay treats each weight matrix as a vector and ignores its spectral structure. We introduce spectral weight decay, a post-step decoupled nuclear-norm update that applies additive rather than multiplicative spectral shrinkage. We connect the update to approximate proximal descent and show that its sensitivity to update order can exceed that of conventional $\ell_2$ weight decay near rank deficiency. Across LLaMA models with $124$M to $500$M parameters, spectral weight decay lowers effective rank and improves SVD-LLM compression at matched validation loss. At $500$M and a $4\%$ distortion budget, it reaches $1.89\times$ compression and $1.18\times$ GPU inference speedup, compared with $1.14\times$ and $1.01\times$ after standard weight decay. Under fixed-horizon training with $60\%$ label noise, it also improves final mean clean-test accuracy over matched $\ell_2$ regularization by up to $17.8$ points on MNIST and $4.6$ points across four BERT-base tasks. Code is available at https://github.com/brain-lab-research/SpectralWD.
cs.LG / 98 / 2610.11734
Timer-M1: A Multivariate Time Series Foundation Model via Learning Primitives
Haoran Zhang, Haixuan Liu, Xingjian Su, Yong Liu, Zhi Chen, Yuxuan Wang, Jianmin Wang, Mingsheng Long
cs.LG · cs.AI
Abstract
We introduce Timer-M1, a pretrained multivariate time series foundation model that learns with primitives for zero-shot forecasting. Across domains, time series share elementary temporal and relational patterns, termed primitives, yet differ in how these primitives manifest and evolve across different contexts. Despite progress in zero-shot and task-general forecasting, existing foundation models may still struggle to generalize to complex real-world scenarios. To this end, we develop a primitive-based data synthesis and pretraining pipeline. The synthesis pipeline generates series with temporal primitives shared across domains and then assembles real and generated series into multivariate samples using relational primitives. Afterwards, samples are organized into episodes by assigning distinct channel roles as target variates, past-only covariates, and known-future covariates, ensuring that the model is optimized on predictable variates using available exogenous information. Technically, Timer-M1 further adapts gated two-dimensional Transformer blocks that dynamically allocate cross-variate attention across layers. Across three large-scale forecasting benchmarks, Timer-M1 ranks first on both FEV and TIME and second on GIFT-Eval among most recent time series foundation models. These results support effective primitive-based pretraining as a route to robust general forecasting technique across domains and task settings.
cs.LG / 99 / 2610.11737
Uncovering and Fixing Collider Bias in Bayesian PINNs
Michael Obermayr, Robert Peharz
cs.LG · cs.AI
Abstract
Bayesian physics-informed neural networks (B-PINNs) are a popular framework for parameter and state inference from sparse or noisy observations. They are commonly formulated via a collider structure, in which physical and trajectory parameters are assumed to be a priori independent and become coupled through virtual likelihoods on differential-equation residuals that enforce physical consistency. We show that this modeling choice can induce severe systematic bias in the posterior over physical parameters: even when the prior is favorably centered on the ground-truth parameters, the resulting posterior can drift away and concentrate far from them. As a remedy, we advocate a hierarchical chain model in which physics generates trajectories, which in turn generate observations. The chain model does not suffer from this posterior bias, but it poses a harder, so-called doubly intractable, inference problem due to a physics-dependent normalization constant. This challenge can be resolved by discretizing the underlying stochastic dynamics, after which the chain posterior can be sampled exactly with particle MCMC. We identify two distinct mechanisms characterizing the collider bias, derive analytical approximations of their magnitudes, and establish diagnostic criteria for predicting when standard B-PINNs remain reliable. Experiments confirm the predicted bias and show that the chain formulation successfully avoids it.
cs.LG / 100 / 2610.11740
Correlational Training of Morphological Neural Networks
Konstantinos Fotopoulos, Petros Maragos
cs.LG
Abstract
Neural networks are typically trained using first-order methods and back-propagation. It is unclear whether this approach is optimal for morphological layers whose weight Jacobians are sparse and whose resulting parameter gradients can be poor. In this work, we propose a novel weight update method for morphological neural networks inspired from the Multiplicative Weights Update (MWU) scheme. We view each morphological perceptron as an instance of the learning from experts' advice problem in logarithmic space, and use a correlation-based reward that favors inputs aligned with the desired output change, regardless of whether a strong gradient signal has reached their weight. We empirically evaluate our approach by training fully connected layers both as stand-alone models and as parts of larger transformer networks. Across nine benchmarks, correlational training yields improvements on eight, by up to 32.84 percentage points, while substantially reducing run-to-run variability.
cs.LG / 101 / 2610.11743
TraceRelay: Attention-Aligned Recurrence over Rolling Traces
Sungwoo Goo, Hwi-yeol Yun, Sangkeun Jung
cs.LG
Abstract
We present TraceRelay, an attention-aligned recurrent architecture that distributes persistent representations over a rolling sequence of low-dimensional traces. Local right looking attention forms increments from lower-layer representations; delivery is delayed until all attended inputs are in the causal past. A fixed additive phase recurrence accumulates the delayed increments, and left-looking attention reads the resulting residual augmented stream. A stride-wise prefix sum supports parallel prefill and bounded-buffer continuation. We study 36 small-model runs on Equal Repeats, bounded Dyck closing-type prediction, and causal Most-Freq generation, using three seeds per setting. At trained length 256, Equal Repeats models with recurrent phase inheritance reach 98.81-99.69% accuracy versus 50.73-51.63% for separately trained variants without inheritance, despite the latter receiving more updates. Accuracy drops sharply at lengths beyond the training range. At the longest evaluated lengths, models with more dimensions in the middle layer's recurrent traces perform better on Dyck (76.34% versus 55.49% close accuracy at length 4096), whereas models with fewer trace dimensions perform better on five-symbol Most-Freq (70.74% versus 55.60% exact generation at length 1024). These contrasting cases motivate further study of how the size of recurrent representations should be chosen for different tasks, without establishing a general rule across tasks or model configurations.
cs.LG / 102 / 2610.11744
The Ball and the Box: Two Geometries of Computation in Superposition
Xiaoyu Li, Lequan Lin, Dai Shi, Jiaojiao Jiang, Junbin Gao, Andi Han
cs.LG
Abstract
Neural representations can encode more features than they have dimensions, a phenomenon known as superposition. We study the dimension needed to compute Boolean gates from such representations. For a single threshold layer with a Gaussian random dictionary and uniformly random sparse Boolean inputs, we derive sharp dimension thresholds under two error criteria. A vanishing expected error count can require more dimensions than correctness of every output with high probability. Shared reads explain the gap: rare realizations can produce many errors at once. The expected-count threshold has ball geometry, while joint reliability has box geometry when a gate is evaluated on every feature tuple. Optimizing shared readout weights and biases gives explicit thresholds for conjunction, disjunction, and majority. For pairwise conjunction, the analysis also describes the transition near the threshold, in agreement with exact simulations.
cs.LG / 103 / 2610.11755
SR-TTA: Spatial-Redundancy Test-Time Adaptation for Interference-Robust Respiration Sensing
Jingyuan Liu, Zheng Chang, Haoqiu Xiong, Zhuangzhuang Cui, Sofie Pollin
cs.LG
Abstract
Future 6G networks aim to expose sensing as a native service by reusing communication infrastructure. We study respiration sensing on a cell-free massive multiple-input multiple-output (MIMO) base station, where a 64-antenna channel must be fused into a breathing waveform. The state-of-the-art hand-crafted fusion is near-optimal in benign conditions. It collapses, however, under strong in-band motion interference, whose frequency falls inside the respiration band. We show that a learned complex-weight beamformer recovers respiration by spatial nulling, and that the remaining gap to a per-recording oracle can be closed at deployment by label-free test-time adaptation. Crucially, we identify which label-free signal makes this work. Frequency- and variance-based criteria cannot separate an in-band interferer from breathing. Our spatial-redundancy test-time adaptation (SR-TTA), which maximizes consistency across random antenna subsets under an out-of-band spectral veto, preserves benign performance in our tests. The respiration-rate error drops from 5.8 to 0.8 breaths per minute (bpm) under simulated in-band interference, and the pipeline maps onto the Open Radio Access Network (O-RAN) architecture as O-RAN distributed-unit (O-DU) range-gating, an adaptation xApp, and a calibration rApp. On real testbed recordings, a one-time cross-subject calibration plus SR-TTA reduces failures from 47% to 7%, drawing level with the hand-crafted combiner using label-free test-time adaptation.
cs.LG / 104 / 2610.11783
In-Ride Alcohol-Impairment Detection in E-Scooterists with False-Alarm Control
Marco Capuccini, Rahul Rajendra Pai
cs.LG
Abstract
Shared e-scooter services have become a widely adopted urban transport mode. While most users ride responsibly, alcohol intoxication stands out among the factors contributing to severe crashes. Nonetheless, countermeasures remain limited to single-point reaction tests and night bans that suspend the service altogether. This paper proposes a new approach in which onboard sensors evaluate the rider as the trip unfolds, raising an alarm as soon as enough evidence of impairment has accumulated. Specifically, we introduce a detector that operates on inertial and throttle measurements, with a provable bound on the rate of false alarms. Experiments on sensor data from 141 rides, in which 25 participants rode while sober and at two target blood alcohol concentration levels, confirm that the bound holds, whereas baselines and ablations either exceed it or lose detection performance, and in some cases delay the alarm. At a bound of 0.023, the detector identifies 91% of the rides performed at the higher concentration and 50% of those at the lower one, with median detection times of 25 and 27 seconds, respectively. We further show that an embedded implementation meets the real-time requirement, making mitigation actions feasible onboard, without requiring data to leave the vehicle. Overall, this work lays the ground for interventions that reach impaired riders as soon as possible, sparing the sober ones the burden of a pre-ride test or the suspension of the service at night, while letting operators budget false alarms against user experience.
cs.LG / 105 / 2610.11784
Compile the Table: Query-Calibrated Operator Compression for Tabular In-Context Learning
Xu Zhao, Jiaming Zhao, Bin Zhao, Yong Yang
cs.LG
Abstract
Tabular in-context learning (ICL) has emerged as a training-free and accurate paradigm for tabular prediction, but current approaches to compressing its in-context examples face an accuracy-throughput tradeoff: fixed subsets can sacrifice accuracy, while query-specific retrieval limits cache reuse and batching across queries, reducing throughput. We propose QCOC (Query-Calibrated Operator Compression), which exploits the exchangeability and repeated use of in-context examples by compiling their full KV cache once into compact memory shared across subsequent queries. Instead of retaining raw examples, QCOC clusters their states into joint-KV prototypes, preserves per-cluster multiplicities and the original example count, and calibrates prototype values against attention query vectors produced by the in-context examples through an anchored closed-form solution. Prototype compression drives the speedup, while value fitting helps preserve accuracy. On 64 held-out OpenML-CC18 datasets, QCOC achieves the highest mean accuracy among the compared compression and retrieval methods at both retained counts. Across 12 configurations on seven long tables, it ranks first among compressed methods in ten and averages 0.23 percentage points below full context. Compressing 8,192 in-context examples to 512 memory slots yields a 10.5x cache compression ratio; excluding one-time compilation, in a single-core CPU online-serving comparison over 1,000 queries, QCOC is up to 508x faster than dynamic retrieval baselines and 1.98x faster than full-context inference. These results show that QCOC enables compact-memory reuse and efficient inference across queries while retaining accuracy close to full context.
cs.LG / 106 / 2610.11803
PRAXIS: Learning Dynamics of Self-Improving Models with Symbolic Archives
Venkat Margapuri, Mustafa Teber
cs.LG
Abstract
Self-improving learning systems adapt data selection, optimization, and auxiliary symbolic components, inducing nonstationary objectives outside standard learning assumptions. We introduce \textsc{PRAXIS}, a co-evolutionary framework that models generators, learners, and symbolic archives as interacting dynamical processes. We prove that KL-constrained generator updates and controlled archive-weight movement bound one-step objective drift, that archive updates suppress a program relative to any fixed comparator with a persistent cumulative utility advantage under sub-Gaussian noise, and that stochastic gradient descent achieves an average-stationarity guarantee whose degradation is governed by cumulative objective drift. Experiments across visual robustness, relational graph reasoning, and algorithmic graph reasoning exhibit generator stabilization, decreasing learner loss, and archive concentration consistent with these theoretical mechanisms.
cs.LG / 107 / 2610.11825
Self-Supervised Speech Representations for Cross-Speaker Dysarthria Detection During Awake Craniotomy
Kanthila Chinmayi, Abdallah Nassib, Misy Harrison, Panheleux Celine, Saliou Vanessa, Seizeur Romuald, Dardenne Guillaume
cs.LG · cs.SD · eess.SP
Abstract
Detecting intra-operative speech impairment during awake craniotomy is essential for preserving language function. However, automated detection remains challenging because operating-room recordings contain substantial acoustic interference, clinically relevant speech events are rare, and available cohorts are small and heterogeneous across speakers. This study presents a systematic component-wise evaluation of a pipeline for distinguishing dysarthric from no-trouble speech in the DATABRASE corpus of awake-craniotomy recordings. The pipeline incorporates speaker diarization to isolate patient speech, a multi-view representation combining handcrafted acoustic descriptors with multilayer wav2vec 2.0 embeddings, speaker-conditional normalization and transferability-based feature selection to improve cross-speaker robustness, and a cascaded classifier comprising a gradient-boosted first stage and a neural second stage. Evaluation was conducted under strict speaker-independent conditions using leave-one-speaker-out cross-validation. The results show that cross-speaker performance is influenced more strongly by the speech representation than by classifier choice. The AUCs of three classifiers differed by no more than 4.7%, whereas replacing conventional acoustic descriptors with the multilayer self-supervised representation produced AUC improvements of 18.2%-26.1%. Diarization-conditioned feature extraction and the proposed classifier cascade provided additional consistent gains. These findings indicate that reliable patient-specific speech isolation and strong pretrained representations are more important than increased classifier complexity in low-resource intra-operative settings. They also quantify the potential performance gains that may be achieved through patient-specific preoperative calibration.
cs.LG / 108 / 2610.11853
DADP: Dynamic Activity-Dependent Pruning, A Reverse Hebbian-Inspired Structural Pruning Method
Bhushan Deshpande
cs.LG
Abstract
Modern neural networks are heavily over-parameterized. This redundancy incurs substantial compute and memory overhead during training and inference. Existing pruning methods rely on post-hoc magnitude thresholds or static initialization heuristics. Consequently, they often require manual per-layer sparsity targets or expensive retraining cycles. We propose Dynamic Activity-Dependent Pruning (DADP), a biologically inspired structural plasticity mechanism. During training, DADP measures connection importance via the accumulated product of pre-synaptic activations and post-synaptic error gradients. Using a single global threshold instead of fixed layer budgets, DADP dynamically allocates sparsity across network depth while naturally inducing neuron- and channel-level pruning. Across MLP, VGG-16, ResNet-18, BiLSTM-CRF, and MiniBERT architectures, DADP matches or outperforms Magnitude, SNIP and RigL, retaining 73.67% accuracy (dense baseline: 76.06%) at 99% sparsity on ResNet-18. Finally, matrix-based Shannon entropy and effective rank measurements confirm that DADP preserves latent feature diversity at extreme sparsities without representation collapse.
cs.LG / 109 / 2610.11866
Understanding Latent-Dimension Scaling in Dynamical-System Learning through Spectral Reliability
Itsushi Sakata, Yuta Miyauchi, Yoshinobu Kawahara
cs.LG · math.DS · nlin.CD
Abstract
In deep learning, approximation theory motivates increasing representation size. We ask whether this benefit extends to dynamics learning through autoregressive prediction. We analyze the learned time evolution through the eigenstructure of Koopman operators, using relative residuals to detect spurious eigenpairs arising even as one-step error falls. For bounded Koopman operators, we show that minimal residuals over learned dictionary spaces converge pointwise to their full-space counterparts as these spaces approximate the observable space in $L^2$. Our hypothesis is that Koopman spectral reliability helps explain how consistently rollout error decreases with increasing dimension. We compare two models of a shared Koopman autoencoder trained alternately for reconstruction and latent evolution, using latent-prediction loss (one-step prediction errors in latent coordinates) or spectral-residual loss (relative residuals of candidate eigenpairs). Across six chaotic systems, both models reduced median windowed rollout error from smallest to largest dimension. The spectral-residual model achieved lower medians than the latent-prediction model for all systems and dimensions, and its median fell by a larger factor in every system. Its median decreased monotonically with dimension in four systems, against one for latent prediction. Against four baseline families, its mean valid prediction times were nearly always longer. At the largest dimension under two-stage training, we compared eigenvalue positions with each learned dictionary's residual contours. Spectral-residual eigenvalues concentrated in low-residual regions, whereas latent-prediction eigenvalues also appeared in high-residual regions, consistent with the hypothesis.
cs.LG / 110 / 2610.11914
Puffin: Probabilistic Learning of Spatial Detail From Coarse Observations
Chaitanya Jobanputra, Sebastian Vollmer, Gerrit Großmann
cs.LG
Abstract
High-resolution socioeconomic variables are important for applications such as urban planning, public health, disaster response, and resource allocation. In practice, however, these variables are often observed only at a coarse spatial resolution. We introduce Puffin, a probabilistic framework for statistical disaggregation that raises the resolution of coarse totals using high-resolution satellite embeddings as covariates. Instead of predicting a single value for each fine-resolution subregion, Puffin learns a probability distribution and is trained through an aggregation-aware likelihood. At inference, Puffin conditions these predictions on the observed regional total and splits it among the subregions. The resulting fine-scale estimates are consistent with the observed aggregate and come with calibrated uncertainty, without requiring fine-resolution labels for training. We evaluate Puffin on German and US census, employment, and election data across population, jobs, and other count variables, and study when statistical disaggregation succeeds or fails across regions, countries, and targets.
cs.LG / 111 / 2610.11953
Interval-valued SHAP in Tree-Based Models
Chenrui Zhu, Vu-Linh Nguyen, Marie-Hélène Masson, Sébastien Destercke
cs.LG
Abstract
Shapley values are among the most popular feature-attribution explanations. Efficient approaches for computing/estimating Shapley values for tree-based models, which are state-of-the-art for tabular data sets, have been developed. However, it is known that Shapley values can be (highly) unrobust due to small and realistic changes. In this paper, we propose an imprecise Dirichlet model (IDM) based method to analyze the robustness of Shapley values in decision trees and random forests. Technically, it is done by quantifying and analyzing the interval-valued Shapley values when a few unannotated instances are randomly introduced to the leaves of the trees. The interval-valued Shapley values can be defined following common principles in handling incomplete data: the pessimistic and averaging principles. We derive various theoretical results that lead to efficient computation of the interval-valued Shapley values. We also show that the proposed method can be straightforwardly generalized to the case of Banzhaf values. We then present various case studies and experiments to illustrate the behaviour of the proposed interval-valued Shapley values and their applications in debiasing uninformative features.
cs.LG / 112 / 2610.11957
Stochastic Grouping Conformal Prediction for Effective Subgroup Reliability
Meihui Zhong, Wenxin Tai, Ting Zhong, Fan Zhou
cs.LG · cs.AI
Abstract
Conformal prediction offers a distribution-free coverage guarantee, making it especially attractive for clinical applications. Standard conformal prediction, however, provides such guarantees only at the population level, and its prediction sets can exhibit coverage disparities across clinically important subgroups. A natural remedy is to calibrate within predefined groups. However, this can require access to sensitive subgroup attributes and is prone to a worst-group bottleneck: protecting the most difficult subgroup can inflate prediction sets for all, increasing cognitive burden on decision makers. To this end, we propose Stochastic Grouping Conformal Prediction (SGCP), a conformal framework for subgroup-reliable uncertainty quantification. It learns a stochastic grouping map that allows each sample to draw calibration information from others with similar calibration behavior, yielding a local score law that boosts reliability across subpopulations. We prove that SGCP retains the standard coverage guarantee. Experiments on synthetic and real-world benchmarks show that it consistently reduces subgroup coverage gaps while achieving smaller or comparable prediction set sizes relative to existing baselines.
cs.LG / 113 / 2610.11984
Example-driven Parametrisations for Bayesian Shape Optimisation
Gabriel Diaz-Aylwin, Joseph Neighbor, Abiel Malkani Talwar, Rui-Yang Zhang, Henry B. Moss
cs.LG
Abstract
Bayesian optimisation is the natural tool for shape design when objectives are expensive and non-differentiable, but it needs a compact yet expressive parameterisation of the search space. Hand-crafting one is a complex endeavour requiring domain expertise, and often yields implicit infeasible regions, artificial bounds, and coupled, unordered coordinates. We instead learn the parameterisation from a collection of existing designs, applying principal component analysis to the deformations between shapes. The result is a linear, interpretable search space in which the number of components explicitly trades expressivity against dimensionality. Across aerofoils, wings, and radio-frequency cavities, spanning 2D geometry to 3D aerodynamics and electromagnetics, we show improved sample efficiency and the ability to explore beyond the confines of hand-crafted baselines.
cs.LG / 114 / 2610.12002
Agentic-TTT: Training test-time policy for test-time training
Jiahao Lu, Mohan Kankanhalli
cs.LG · cs.AI · cs.CL
Abstract
Test-time training (TTT) adapts an LLM's parameters using signals derived from test inputs, and can make striking improvements in pre-specified settings such as IMO competitions or designated open problems. By turning deployment experience into parameter updates, TTT provides a direct mechanism for model-level self-improvement. Yet TTT is not universally beneficial: each TTT algorithm works in different settings, and applying an ill-suited method could waste test-time compute or even damage model performance. Therefore, such parameter-level self-improvement requires agency: the model must decide when TTT is warranted, which algorithm to invoke, and whether an existing skill can be reused. To fill this gap, we introduce Agentic-TTT, which learns a test-time policy to govern those decisions. Agentic-TTT turns TTT procedures into callable tools, treats accumulated skills as an evolving deployment environment, and trains its policy using the observed utility gains from its decisions. On our benchmark, Agentic-TTT nearly doubles the utility over the backbone model, learns to trade off utility against compute, and generalizes to domains unseen during training. Together, these results point toward autonomous self-improvement: models that can decide how to learn from their own deployment experience.
cs.LG / 115 / 2610.12004
The Polytopal Neural Network
A. Emilie J. Wedenborg, Anders V. Nørskov, Teresa Dorszewski, Kristoffer Wickstrøm, Morten Mørup
cs.LG · cs.AI
Abstract
Understanding how deep neural networks process information remains a central challenge. Existing interpretability methods often compromise structural fidelity, rely on prespecified corpora, or explain models post-hoc. We propose Polytopal Neural Networks (PNNs), a framework that extracts distinct layer-wise aspects by enforcing a polytope-based structure that is used directly in subsequent information processing. We scale our approach using learned corpus representations and an amortized simplex inference procedure and highlight how the framework also gives a direct route to vector quantized (VQ) training. In PNNs, observations are explicitly described by their alignment with layer-specific aspects. Empirical results show that imposing polytopal constraints on neural network representations preserves meaningful structures in the latent space with minimal degradation in performance, favorable compressed representations when compared to VQ representations in unsupervised learning, while also providing a performant new approach to VQ deep learning training. Our findings suggest that deep networks can enforce interpretable polytope-based representations, offering a principled path toward more transparent AI systems with minimal performance compromise.
cs.LG / 116 / 2610.12005
Test-Time Compute for Tabular Foundation Models: Mechanisms, Gains, and Limits
Kanghui Ning, Marin Biloš, James T. Wilson, Yilang Zhang, Kashif Rasul, Dongjin Song, Anderson Schneider, Yuriy Nevmyvaka
cs.LG
Abstract
Which forms of test-time compute improve the predictions of strong pretrained tabular foundation models (TFMs)? We systematically study this along three axes: adaptation, aggregation, and context construction. Our evaluation spans modern TFMs across the TabArena benchmark, supplemented by experiments on wide and large-scale tables from OpenML. For adaptation, we introduce DiagScale, a diagonal query-key similarity update. It trains only 0.003-0.03% of model parameters and achieves gains comparable to full fine-tuning across three independently pretrained backbones. For aggregation, both pool composition and selection strategy matter. TabPFN-3 already averages predictions from different preprocessing variants of the same data, and adding more such predictions yields diminishing returns. With a broader pool of 96 configurations, greedy selection reduces error by 2.4% relative to the default predictor, but uniform averaging increases error. For context construction, attention-guided retrieval improves TabPFN-3's predictions on some large tables and supports source pools beyond the full context memory limit. The context expansion methods we test yield no consistent improvement. Taken together, our results suggest that adaptation and selective aggregation yield consistent benchmark-level gains. The benefits of context construction depend more on the task and data regime. Adaptation and aggregation over the same backbone yield further gains when combined, but require substantially more computation than default inference. These trade-offs motivate choosing strategies according to the available computation budget. Code is available at https://github.com/kanghui-learning/test-time-compute-for-tabular-foundation-models.
cs.LG / 117 / 2610.12016
CausalDreamer: Learning Predictive World Models with Latent Disentanglement
Prince Jha, Nils Lukas, Kun Zhang, Salem Lahlou
cs.LG
Abstract
World models for control must capture which aspects of the environment respond to the agent's actions and which are relevant to reward. Generative world models such as Dreamer 4 consist of a video tokenizer, which encodes each frame into a latent, and a dynamics model, which is pretrained to predict future latents from past latents and actions. Yet the tokenizer is trained with a reconstruction objective, without action or reward supervision, so its latent provides no explicit mechanism to separate controllable, uncontrollable, reward-relevant, and reward-irrelevant information. We propose \textit{CausalDreamer}, which keeps the tokenizer frozen and re-encodes its latent into a factored representation of four groups along two axes: controllability, where only the two controllable groups receive the action, and reward relevance, learned by predicting the reward from the two reward-relevant groups. The pretrained dynamics model is then fine-tuned to predict the factored representation. We evaluate \textit{CausalDreamer} and the pretrained world model it starts from with model-predictive planning on 20 MMBench2 tasks: 10 clean tasks seen during training and 10 unseen tasks, of which 6 are manipulated variants of clean tasks with a changed background, object, or maze layout, and 4 are new environments. We normalize returns so that a policy taking uniformly random actions scores 0 and an expert scores 1. \textit{CausalDreamer} achieves a 14\% higher normalized score than the pretrained world model on the clean tasks (0.199 vs.\ 0.175) and a 25\% higher score on the manipulated variants (0.307 vs.\ 0.246), while neither model scores meaningfully above the random policy in the new environments. Additionally, our analysis shows that the factored representation separates reward-irrelevant changes, such as a changed background, from its reward-relevant groups.
cs.LG / 118 / 2610.12039
MPGE: A Multi-Perspective Graph Explainer for Molecular Classification Explanation
Mahtab Sarvmaili
cs.LG · cs.AI
Abstract
Graph neural networks (GNNs) predict molecular properties from chemical graph data, but predictive accuracy does not explain how graph information supports an individual decision. A compact prediction-preserving rationale does not necessarily reveal which changes reverse the decision or which modifications the model tolerates. We propose the Multi-Perspective Graph Explainer (MPGE), unifying factual support, counterfactual sensitivity, and exemplar tolerance for a frozen classifier. The factual view, originally termed prototype (PT), seeks a compact retained edge set with the same label and required confidence. Counterfactual (CF) explanations seek bounded prediction-changing deletions; exemplar (EXE) explanations seek non-trivial bounded deletions that preserve the label and confidence. A shared constrained formulation connects prediction behavior, compactness, and edit cost, while separate objectives generate the three views. Our graph-classification extension of CF-GNNExplainer learns symmetric edge rankings and verifies discrete candidates, recording unsuccessful searches. A separate BBBP fragment backend returns RDKit-sanitized molecules. We evaluate the primary GCN implementation on MUTAG, Mutagenicity, AIDS, COX2_MD, and BBBP using semantic coverage, conditional quality, stability, and runtime. Successful factual masks retained 8.6%--15.5% of input edges on average across datasets; bounded counterfactual coverage was 4.8%--67.6%, and exemplar preservation coverage was 98.9%--100.0%. Exploratory controls reveal the influence of hard projection and retained node information. Quantitative comparisons and molecular visualizations characterize model support, sensitivity, and tolerance without treating them as validated chemical mechanisms.
cs.LG / 119 / 2610.12076
Exploiting Gradients in Bayesian Inference of Expensive Simulators
Šimon Soldát, Václav Šmídl
cs.LG · stat.ML
Abstract
Simulators based on differential equations are ubiquitous in science and engineering. They are often used in simulation-based inference to evaluate the posterior distribution of the input parameters based on real-world observations of the simulator outputs. However, inference becomes challenging when individual simulator evaluations are computationally expensive. In such cases, a Bayesian optimization-based active learning approach with Gaussian process surrogate models has been used to maximize the information obtained from a limited simulation budget. Recently, gradients of simulator outputs with respect to input parameters have become increasingly available, yet they are rarely exploited for inference. Even though we only need to learn the simulator input-output relationship, gradient information can provide an additional valuable signal to guide the active learning procedure. This is of particular interest in the case of expensive simulators, when sample efficiency is crucial. In this paper, we demonstrate how incorporating gradient information into the Gaussian process surrogate accelerates Bayesian optimization-based inference under a limited simulation budget. Our results show significant improvement in convergence speed from using gradient information. For reverse-mode differentiation, the inference efficiency gains are maintained when accounting for the additional computational cost. In contrast, for forward-mode differentiation, the inference speed-up does not outweigh the computational costs. These results indicate that gradient-enhanced surrogates are beneficial primarily in problems where the number of parameters exceeds the output dimensionality, where reverse-mode differentiation is efficient.
cs.LG / 120 / 2610.12096
SCORE: Spectral Correlation Estimation for Multivariate Gaussians
Christopher Bülte, Emil Partow, Astha Gupta, Pascal Esser, Gitta Kutyniok
cs.LG
Abstract
Neural network-based predictive modeling with high-dimensional structured Gaussian targets requires an efficient and numerically stable, yet expressive approximation of the covariance matrix. We propose SCORE: a scalable framework, combining scoring rule training with an expressive covariance approximation learned in spectral space. For $d$-dimensional data, the learning task is decomposed into learning the marginal distributions and learning a structured correlation matrix, which enables dense dependencies with linear storage and $\mathcal{O}(d\log d)$ cost. We utilize the closed form Gaussian kernel score for training, which remains defined even for degenerate covariances and admits bounded gradients during optimization. We characterize kernel scores under invertible transforms and prove exact invariance under unitary transforms. At population level, our two-level objective recovers the true marginals and projects the target correlation onto the representable class; finite-sample PAC bounds show that the errors of the two stages enter additively. We evaluate our model on a variety of tasks with a commonly assumed Gaussian domain: Time-series forecasting, monocular depth estimation, and spatial weather prediction, showing improved performance at lower computational cost.
cs.LG / 121 / 2610.12102
Few-Step Generation via Data-Space Iteration
Shanchuan Lin, Yansong Peng, Fu-Yun Wang, Haoqi Fan
cs.LG · cs.CV
Abstract
Flow matching has emerged as a scalable paradigm for training high-quality generative models, but sampling from the learned probability flow requires many network evaluations. Distillation can reduce this cost to one or a few evaluations; however, one-step generation often sacrifices quality, making few-step generation the practical operating regime. Existing few-step methods perform their iterative computation along the probability flow and therefore require a fixed, manually chosen timestep discretization. This discretization is often chosen heuristically and is expensive to tune; it may also be restrictive when refinement difficulty differs across samples or spatial locations. We introduce data-space iteration, a few-step generation framework that removes flow discretization altogether. Starting from noise, a shared generator directly refines its prediction in data space, with every iteration trained to produce the best sample permitted by its capacity. Our formulation integrates with distribution matching distillation (DMD) with minimal changes, enabling a controlled comparison between iteration methods under matched training settings. On class-conditional ImageNet 256x256, data-space iteration outperforms standard discretization baselines and matches or improves upon variants selected through schedule search, without requiring schedule-specific training. These results show that data-space iteration provides a simple and effective alternative to discretized flow-space iteration for fast generation.
cs.LG / 122 / 2610.12113
A Geometric Approach to Soft Actor-Critic with Zonotopes for Locomotion Learning
Panagiotis Roditis, Panagiotis P. Filntisis, Petros Maragos
cs.LG · cs.AI
Abstract
Off-policy actor--critic methods control overestimation bias by taking the minimum of two critics. This uses the same aggregation rule everywhere, regardless of how the critics disagree. We propose \textbf{GeZo-SAC}, which uses auxiliary geometric representations to adapt critic pessimism to the state and action. Alongside its scalar value, each critic predicts a set of generators defining a zonotope. Probing this zonotope along sampled directions provides a geometric width, "subtracted from each critic value as a pessimistic offset, and a measure of disagreement between the two critics, aggregated with log-sum-exp. This disagreement controls how the critics are combined, moving from a width-weighted average toward the usual minimum as disagreement increases. At inference, the deployed policy is an unmodified SAC actor, since the generators are used only on the critic side during training.Across four MuJoCo-v5 locomotion benchmarks and six off-policy baselines, GeZo-SAC achieves the highest mean return on Ant-v5 and Hopper-v5 and remains competitive with other methods on the remaining tasks. Our analysis further shows that GeZo-SAC achieves the lowest average actuator work and action effort per metre among the evaluated methods, while maintaining near-zero measured overestimation frequency across all four environments.
cs.LG / 123 / 2610.12115
Credal Machine Learning for Risk-Averse Decision Making
Timo Löhr, Paul Hofman, Maximilian Muschalik, Eyke Hüllermeier
cs.LG · stat.ML
Abstract
In many machine learning applications, it is necessary to guard against worst-case scenarios and predictions that could result in substantial losses. In principle, this can be achieved by training risk-averse predictive models that minimize loss functions such as conditional value-at-risk (CVaR), rather than relying on models that perform well on average. In practice, however, the effectiveness of this approach to risk aversion is undermined by the learner's uncertainty regarding the true loss distribution and, consequently, the true CVaR. To achieve reliable risk-aversion, we propose a method in which this (epistemic) uncertainty is represented in terms of credal sets, i.e., sets of probability distributions. More specifically, we develop an efficient yet reliable learner that produces predictions in the form of credal sets and combine it with a novel decision rule that maps each credal set to a single predictive distribution for CVaR minimization. Across classification, under distribution shift, and in reinforcement learning, our approach reliably avoids catastrophic decisions, while sacrificing little in expected performance.
cs.LG / 124 / 2610.12119
Using Weisfeiler-Leman Features for Algorithm Selection in Constraint Optimisation
Alessio Pellegrino, Jacopo Mauro
cs.LG · cs.AI
Abstract
Algorithm Selection is essential for efficient Constraint Programming. Over the years, many algorithm selectors based on machine learning methods have been successfully applied, yet traditional feature extraction methods often rely on manually decided instance-level statistics that fail to capture the underlying problem structure. In this paper we aim to bridge this gap by introducing a novel, automated feature extraction methodology that integrates graph conversion and Weisfeiler-Lehman graph kernels to generate robust structural representations of problem instances. The 1-WL test bounds the graph-distinguishing power of standard message-passing Graph Neural Networks (GNNs), and suitable GNN architectures match this bound \citep{Xuetal2018}. WL-based features offer an alternative that does not require training a GNN. Our primary contribution is a cut-based representation (\texttt{WLc}) designed to model structural partitions and provide a more nuanced predictive signal. We evaluate our approach on instances from the 2023--2025 MiniZinc Challenges across two tasks: maximizing Borda count scores and maximizing predictive accuracy. Experimental results across Support Vector Machines, Random Forests, and Multi-Layer Perceptrons demonstrate that cut-based features outperform \texttt{fzn2feat} with SVMs, while results with RFs and MLPs are closer.
cs.LG / 125 / 2610.12148
Large-Scale Benchmarking of Quantum Neural Network Configurations for Financial Time Series Forecasting
Jack Waller, Xing Liang, Dimitrios Makris, Rajagopal Nilavalan
cs.LG · quant-ph
Abstract
Quantum machine learning, and quantum neural networks (QNNs) in particular, are advancing fields with growing potential. Although systematic comparisons of QNN configurations have been explored primarily for classification tasks, comparatively little attention has been given to regression problems, particularly financial time series forecasting. This study presents a large-scale systematic comparative evaluation of QNN component configurations for financial time series forecasting, using the GBP/USD spot exchange rate as a case study. A grid search across encoding methods, ansatz designs, qubit counts, layer depths, and cost functions yields 1,368 distinct model configurations, each evaluated in terms of prediction accuracy, computational cost, and convergence behaviour. The results reveal unique insights into how the choice of methods influences performance, such as that gate selection and arrangement are more critical to model success than raw parameter count, and that entanglement is a system-level property of the full circuit rather than solely at the ansatz level. The best-performing QNN configuration achieves an $R^2$ score of 0.985, outperforming a classical BiLSTM baseline. Additionally, the impact of real quantum hardware noise is assessed through execution on the IQM Emerald device, revealing that gate errors and decoherence represent a significant barrier to practical deployment, with gate selection and circuit depth identified as key determinants of hardware noise resilience. Overall, the findings provide practical architectural guidance for QNN design and establish a baseline characterisation of QNN noise sensitivity on near-term quantum devices.
cs.LG / 126 / 2610.12150
Bayesian Optimisation under State-Preservation Constraints
Gabriel Diaz-Aylwin, Vignesh Gopakumar, Omkar Myatra, David Moulton, Lorenzo Zanisi, David S. Leslie, Henry B. Moss
cs.LG
Abstract
In many engineering design problems, the objective and constraints depend on the state: the solution of a PDE determined by the design parameters. We consider improving a design while holding selected state observables near trusted values, which we call state preservation constraints. Constrained Bayesian optimisation handles these with a learnt feasibility model, but struggles with this problem's highly anisotropic feasible set. Our central idea is to pre-compute the set of controls whose linearised constraint response stays within tolerance, thereby pulling back the state-space constraint into design space. This linearisation defines an ellipsoid from which we can efficiently draw a large number of well-spread candidates. The underlying linear response map is refined online, and the ellipsoid is rebuilt accordingly. We demonstrate the method end-to-end on our key application - Tokamak divertor optimisation under plasma-boundary preservation.
cs.LG / 127 / 2610.12153
Toward Optimal Regret in Adversarial MDPs with Stochastic Hard Constraints
Qian Zuo, Francesco Emanuele Stradi
cs.LG
Abstract
We study episodic constrained Markov decision processes with adversarial losses under stochastic hard constraints. Specifically, starting from a known strictly feasible policy with margin $d$, we seek to obtain optimal regret while satisfying the expected cost constraints in every episode. In this setting, Stradi et al. (2025) show that a carefully designed mixing rule attains regret of order $\widetilde{\mathcal{O}}(\sqrt{T}/\min\{d,d^2\})$. Interestingly, they also provide a lower bound of order $Ω(\sqrt{T}/ρ)$ for the same setting, where $ρ$ is the Slater margin of the offline problem and can be much larger than $d$. In this work, we build on their approach to obtain optimal regret dependence on these margins. Specifically, we propose MA-OPS, an algorithm that combines an optimistic search for the Slater margin with a pessimistic evaluation of the selected policies to safely learn a policy with a large feasibility margin. This policy is then used to minimize regret while satisfying the constraints at every episode. In particular, we show that MA-OPS attains regret $\widetilde{\mathcal{O}}(\sqrt{T}/ρ+ 1/(dρ))$. Finally, we provide a matching lower bound, showing that the dependence on $T$, $d$, $ρ$ in the regret bound is optimal up to logarithmic factors.
cs.LG / 128 / 2610.12161
When KL Regularization Misfires in Group Policy Optimization
Fei Ding
cs.LG · cs.CL
Abstract
Why does removing reference-policy KL regularization sometimes improve group policy optimization? This motivates studying how reference-policy information should enter group-relative updates. We analyze seven potential failure modes in the interactions between KL and rewards: residual KL updates after reward clipping, after gradient cancellation, and in groups with identical rewards; KL growth with response length and an imbalance in its relative contribution; KL concentration on a small number of tokens; and sampling noise when k1 is incorporated into rewards. We propose Zero-Sum Calibrated Policy Optimization (ZCPO), which uses relative drift measured by conditional KL to calibrate within-group reward coefficients and integrates them into the base surrogate. Mathematical reasoning experiments and ablations support this design's effectiveness in our settings.
cs.LG / 129 / 2610.12163
Scalable Hierarchical Graph Generation via Soft Community Structure
Ahmet Tüzen, Helge Langseth, Kjetil Nørvåg
cs.LG · cs.AI
Abstract
Generating large attributed graphs requires reproducing the topology, generating attributes jointly with the structure, and remaining scalable. Many real-world graphs exist as a single large graph, so a generative model has to generalize from the one graph it is fit on, without independent samples. We present Schema, which recursively decomposes a reference graph into a hierarchy of soft communities, assigning each node a membership distribution. Generation is then split into three stages, each trained independently: (1) synthesizing node attributes conditioned on soft memberships, (2) generating intra-community edges from local structural context, and (3) modeling inter-community connections over bridge nodes whose membership mass is distributed across several communities. No stage forms the full adjacency matrix, and each stage operates on a subgraph bounded by the community size. We also introduce an evaluation protocol that covers structural fidelity, memorization, downstream utility, and scalability. On four real-world attributed graphs, Schema recovers the balance between local and long-range structure more closely than any other model that generates attributes, while reproducing only a small fraction of the reference edges. It retains the downstream accuracy of the reference graph without raising it artificially above that level. Baselines that match its structural fidelity memorize the reference, while those with higher downstream accuracy either exceed the reference accuracy or fail to complete on the larger graphs. We measure scalability on six additional graphs with up to 10 million nodes.
cs.LG / 130 / 2610.12167
Is Real-World Training Data Necessary for Generalist Graph Anomaly Detection?
Yujing Liu, Yixin Liu, Yue Tan, Xiaofeng Cao, Alan Wee-Chung Liew, Heng Tao Shen, Shirui Pan
cs.LG
Abstract
Generalist graph anomaly detection (GAD) aims to build a foundation model that detects anomalies on arbitrary unseen graphs without retraining or fine-tuning. Sufficient data are essential for foundation model training, yet generalist GAD still faces a data shortage, as real-world anomalous graphs are scarce and costly to collect and annotate. To fill this gap, we propose AG-FORGE, an Anomalous Graph generation Forge for automatic synthesis of anomalous graphs, exploring the feasibility of synthetic data-driven training for generalist GAD. Empirically, we find that synthetic data can achieve performance comparable to real-world training, but fail to push the performance boundary further due to the limited capacity of existing methods. To further unlock model capacity as training data scale up, we develop TS-GGAD, a Topology-Semantic coordinated Generalist GAD that captures complementary topological and semantic anomaly evidence, together with a curriculum learning strategy tailored to large-scale synthetic training. Extensive experiments on 14 real-world datasets demonstrate that TS-GGAD, trained on data generated by AG-FORGE, significantly outperforms state-of-the-art methods.
cs.LG / 131 / 2610.12190
DataSense-Bench: The First Step Toward an AI Scientist
Yudi Zhang, Mingyu Cao, Lu Yin, Mykola Pechenizkiy, Shiwei Liu
cs.LG
Abstract
As claims about recursive self-improvement (RSI) and artificial general intelligence (AGI) proliferate, we ask a simple question: do frontier AI models have a sense of data, i.e., can they reliably select the right data for training? We introduce DataSense-Bench to study this capability through the fundamental problem of data selection and performance forecasting in machine learning. We ask AI agents to select and rank candidate training subsets that can be used to fine-tune a small LLM model. Agents are allowed to inspect the data, write and execute analysis code, and run model forward passes, but can not train the model or access the actual evaluation tasks. We then fine-tune the base model on each selected subset and evaluate its post-training performance under a standardized protocol. We instantiate the benchmark in terminal problem solving and tool use, selecting trajectories from OpenThoughts-Agent and EnvScaler and evaluating on TBLite and BFCL, respectively. We then evaluate the agents along two complementary dimensions: the post-training performance of the top-ranked subset, reflecting the ability to identify high-value training data, and ranking accuracy, reflecting the ability to predict the relative performance of the selected subsets. In our experiments, selection gains over random selection are limited; agents do not reliably rank their selected groups, and ranking ability does not hold consistently across tasks: Astra identifies the best group in all three tool-use runs but in only one of three terminal runs. Analysis of execution traces on both tasks shows that agents often use similar data signals while interpreting their training value differently.
cs.LG / 132 / 2610.12211
Verification with Transfer: Exact Information Frontiers and Their Price in Calls
Hazar Yueksel
cs.LG · cs.CR · cs.IT · stat.ML
Abstract
A verifier that accepts or rejects whole answers reveals little: under a flat prior over $k$-bit answers, zero error needs $2^k-1$ verifications. The usual remedy is to solve related source tasks, either all first, as a curriculum does, or interleaved with verification. We price this remedy in information and in calls. With an exact verifier, the least causal information that any interleaving of source calls and $n$ verifications needs to succeed with probability $s$ is a list rate-distortion function, attained by one observation before any verification. It lower-bounds the expected number of binary source calls, which designed sources meet within $1+\log_25$ calls for unique answers and within a logarithmic term in general, where no additive constant suffices. With an exact verifier and fixed sources, moving every call before the first verification preserves all hard caps on calls, although interleaving can save unboundedly many expected calls; under a noisy verifier, source-first protocols can lose unbounded factors in information and in error. For linear banks over $\mathbb{F}_2$, optimal accuracy has a closed form, and after a polynomial-time reduction the budget profile is computable in time $2^{O(h^2)}\operatorname{poly}(J,k+h)$ for $J$ sources and nuisance dimension $h$. In these banks, for zero error under a hard cap, the calls beyond the rounded-up information price are exactly those spent on nuisance. Every numbered result apart from two clauses about the planner is machine-checked in Lean 4, assuming two published results. Used as a ruler, the frontier shows a small transformer using all delivered bits at latent dimension $5$ and none at $11$ within fixed training budgets; in a test with predictions recorded before training, low XOR degree of the target bits did not suffice for their use.
cs.LG / 133 / 2610.12232
Training on the Future: A Delay-Aware Audit of Test-Time Adaptation for Time-Series Forecasting
Mohamed Readh Fentazi, Mazene Ameur, Adlen Ksentini
cs.LG
Abstract
Test-time adaptation (TTA) methods for time-series forecasting update a deployed model, or a small adapter around it, from incoming ground truth. But the label of an $H$-step forecast exists only $H$ steps later, and real data pipelines add further delay. We build a leakage-free harness in which the label of forecast origin $s$ is released for updates only at step $s+d$ with $d \ge H$, and enforce this rule inside the released code of four recent TTA methods (TAFAS, COSA, PETSA and DynaTTA), run on their own backbones and checkpoints across five benchmarks (ETTm1, ETTh2, Weather, Electricity and Traffic). As references we add two closed-form correctors: a bank of recursive least squares (RLS) filters combined by a per-coordinate median, with no tunable hyperparameters and 56 microseconds per step on the 7-channel streams, and an ELF-style linear corrector. Under causal delayed labels the picture is asymmetric. On ETTm1 every audited method genuinely adapts, yet the RLS bank still beats three of the four at a fraction of their cost; only DynaTTA beats the bank, only at the minimum causal delay, and at roughly 2,500 times the per-update cost; the ELF-style corrector beats all four. On the other four datasets, the largest statistically significant improvement any published method achieves over its own frozen checkpoint is half a percent, on all four at least one published method is significantly worse than the frozen model at the minimum causal delay, and on drift-heavy ETTh2 longer label delays make every adapter that separates from the frozen model, ours included, significantly harmful. Leaky next-step updates inflate the apparent gains of simple adapters by up to 110%, and the backbone training recipe moves frozen online error by up to a factor of 25, more than any adaptation effect we measure. We release the harness, integration patches and all cached runs.
cs.LG / 134 / 2610.12240
AdaCast: Conditional Parameter Generation for Adaptive Time Series Forecasting
Darahaas Nallagatla, Darryl Cherian Jacob, Pan He
cs.LG · cs.AI
Abstract
Time-series foundation models (TSFMs) have achieved strong forecasting performance across domains. However, most adaptation methods remain static. Existing all-in-one methods learn a single set of dataset-level parameter updates and apply the same adapted model to every input. As a result, they cannot adapt the model parameters to the temporal patterns, seasonality and dynamics of each input time series. This limits their ability to produce forecasts that are tailored to heterogeneous inputs. To address this limitation, we propose AdaCast, a conditional parameter generation framework for time-series forecasting. AdaCast uses a generator to produce input-specific low-rank parameter updates for a frozen pretrained TSFM. These updates adapt the model to each input during both training and inference. Across six public benchmarks, AdaCast consistently outperforms static adaptation baseline in in-domain forecasting and improves zero-shot generalization to held-out datasets across domains. These results demonstrate that conditional parameter generation provides an effective approach for adaptive forecasting.
cs.LG / 135 / 2610.12244
RIFT: Relative Isolation From Trees For Anomaly Detection
Mark Daniel Szalai, Gabor Horvath
cs.LG
Abstract
Isolation Forest (IF) is a widely used baseline for unsupervised anomaly detection. Recent studies provide a closed-form expression for the infinite-forest limit for one-dimensional data. Inspired by the geometric interpretation of this formula, we introduce RIFT (Relative Isolation From Trees), a deterministic anomaly detection method that generates the minimum spanning tree and scores each point by the sum of the apparent sizes of tree edges as viewed from that point. For one-dimensional data, the RIFT score recovers the closed-form IF limit exactly. In higher dimensions, it provides a parameter-free generalization that is deterministic, robust to varying density and clustered anomalies and avoids the axis-parallel artifacts of IF. We further propose an ensemble variant for large datasets. Experiments on synthetic data and the ADBench benchmark demonstrate that the accuracy is comparable to IF, while the ensemble variant exhibits significantly lower variance across random seeds.
cs.LG / 136 / 2610.12247
Batch Before You Lift: Scalable Topological Deep Learning on Large Graphs
David Leko, Luka Benić, Guillermo Bernárdez, Nina Miolane, Olga Fink, Lev Telyatnikov
cs.LG · cs.AI
Abstract
Topological Deep Learning extends graph-based learning to higher-order domains, such as hypergraphs, cellular, and simplicial complexes. These domains are typically constructed from patterns in an input graph through a process of graph lifting. Full-domain training constructs and stores the complete lifted representation before model execution. On large and dense datasets like Reddit (233k nodes and 57.3M edges), this global materialization becomes a severe computational bottleneck, often rendering training infeasible. To address this limitation, we introduce Cluster-TNN, a domain-agnostic framework that avoids this bottleneck by lifting locally instead. After partitioning the input graph during preprocessing, at runtime Cluster-TNN dynamically samples groups of node clusters, reconstructs their induced subgraphs to form mini-batches, and applies the chosen lifting within each mini-batch. Retaining all edges among sampled nodes preserves the connectivity needed to construct higher-order structures across clusters, producing topological mini-batches that existing Topological Neural Networks can process directly. Across 21 matched comparisons with full-graph execution, Cluster-TNN reduces peak GPU memory in every configuration, by 83.2% on average while maintaining competitive predictive performance. Notably, such a reduction enables, to our knowledge, the first training of multiple different higher-order Topological Neural Networks on large datasets such as Reddit and OGBN Products. These results establish Cluster-TNN as a general strategy for scaling Topological Deep Learning beyond the limitations of global domain construction.
cs.LG / 137 / 2610.12265
AdaptLSTM: Efficient Adaptive Online Learning for Cloud Workload Forecasting under Distribution Drift
Xinhua Miao, Bowei Yang, Zhengong Cai
cs.LG
Abstract
Accurate workload forecasting is critical for elastic resource provisioning in web-scale cloud services, where distribution shifts driven by viral content, product launches, and user behavior degrade offline-trained models rapidly. Naive online learning recovers accuracy but incurs prohibitive per-step compute cost. We propose AdaptLSTM, an adaptive online framework that detects drift via validation-calibrated thresholds and applies selective, targeted updates. On the Alibaba Machine Trace, AdaptLSTM recovers 54\% of Naive Online's improvement at 20\% cost ($2.7\times$ efficiency, $p=0.002$ over 10 seeds). On the more volatile Container Trace, it achieves 96\% at 20\% cost ($4.8\times$ efficiency, $+75\%$ MAE reduction over Static). Unlike classical drift detectors (ADWIN, DDM, Page-Hinkley) which fail to trigger on regression-scale error streams, AdaptLSTM fires 42 times over 301 steps and outperforms matched-budget baselines. Wall-clock profiling shows $1.33\times$ throughput gain and 45\% update-time reduction. The framework is model-agnostic: identical Pareto patterns hold for LSTM, GRU, and Transformer backbones.
cs.LG / 138 / 2610.12328
Composite Online-to-Nonconvex Conversion with Optimal Oracle Complexity
Mingyi Li, Taira Tsuchiya, Kenji Yamanishi
cs.LG · math.OC · stat.ML
Abstract
We consider stochastic nonsmooth nonconvex composite optimization, which includes several important problems such as constrained optimization and the regularized training of neural networks. The objective is the sum of a possibly nonsmooth nonconvex Lipschitz function and a convex regularizer, and the function is accessed through stochastic gradients or function values. The goal is to find a point that satisfies a Goldstein-type stationarity condition designed for composite objectives. To our knowledge, no oracle complexity bound for this setting is known under first-order access, and existing complexities under zeroth-order access are suboptimal. To handle this issue, we employ the framework of online-to-nonconvex conversion, which chooses update directions by an online learner and is known to achieve optimal rates for noncomposite problems. We extend the framework to our composite scenario by introducing new losses for the learner, which contain the regularizer itself rather than its linearization and for which a variant of online mirror descent achieves low regret. We show that the resulting algorithm finds such a point with $O(δ^{-1}\varepsilon^{-3})$ stochastic gradient queries or $O(dδ^{-1}\varepsilon^{-3})$ function-value queries, where $δ$ is the Goldstein radius, $\varepsilon$ is the stationarity tolerance, and $d$ is the dimension. These rates match the optimal ones for noncomposite nonsmooth nonconvex optimization, demonstrating that the additional convex regularizer does not worsen the oracle complexity. We also give rates for the smooth case and present numerical experiments.
cs.LG / 139 / 2610.12349
SplitJEPA: Learning Invariant and Variant Latent Worlds without Reconstruction
Ruijin Hua, Zichuan Liu, Zhuokai Zhao, Yujia Zheng
cs.LG
Abstract
Understanding a dynamical world calls for more than a latent state that summarizes its observations: the state should also be organized into the factors that stay shared across related observations and the factors that vary between them. For example, a robot pushing a cube to a goal should take the same action when the camera shifts or the lights dim, since nothing in the scene has moved. Existing approaches to this decomposition commonly obtain it through reconstruction, so the latent variables must first explain the entire observational world before their organization can be trusted. Joint embedding predictive architectures (JEPAs) model the latent state directly and never reconstruct, yet no existing result recovers the invariant and variant parts of the state they learn. How to learn the invariant-variant structure of the latent world without paying for its reconstruction therefore remains open. To close this gap, we introduce SplitJEPA, a JEPA that jointly recovers the latent state and its invariant and variant organization directly in representation space, without any reconstruction. We prove that, under stationary Gaussian predictive dynamics and a full-rank variation condition, SplitJEPA identifies the invariant and variant subspaces up to independent block-wise isometries, without introducing an observation decoder. Since the guarantee needs no decoder, the result extends reconstruction-free latent recovery to invariant-variant block identification. Experiments on synthetic nonlinear systems and robotic manipulation tasks support the theoretical results and show their practical value for both robustness and efficiency.
cs.LG / 140 / 2610.12362
Closing the Horizon Gap in Policy Optimization for Adversarial MDPs
Mingyi Li, Taira Tsuchiya
cs.LG · stat.ML
Abstract
We consider policy optimization for online episodic tabular Markov decision processes (MDPs) with adversarial losses and bandit feedback. Policy optimization updates the policy locally at each state and avoids optimization over the occupancy-measure polytope, but its existing regret bounds are larger by a factor of the horizon $H$ than those of occupancy-measure-based algorithms. We close this gap by using regularized $Q$-functions, which allow us to control the stability of the local updates jointly over all state-action pairs rather than separately at each state. The resulting algorithm attains high-probability regret bounds of $\widetilde O(\sqrt{HS(H+A)T})$ for known transitions and $\widetilde O(HS\sqrt{AT})$ for unknown transitions, where $S$ is the number of states, $A$ the number of actions, and $T$ the number of episodes. Both bounds improve the horizon dependence of existing policy optimization bounds, and the latter matches the best-known bound. We further extend the algorithm to adversarial linear-mixture MDPs and obtain the same improvement in the horizon dependence.
cs.LG / 141 / 2610.12370
Bilevel optimization for data-driven learning of Koopman embeddings using kernel-based autoencoders
Joel-Pascal Ntwali N'konzi, Feliks Nüske, Stefan Klus
cs.LG · math.DS · stat.ML
Abstract
Koopman operator theory provides a linear framework for analyzing nonlinear dynamical systems and has become a major tool for data-driven modeling. A central challenge, however, is that finite-dimensional approximations computed by methods such as extended dynamic mode decomposition (EDMD) require the dictionary to be specified a priori. Recent machine-learning approaches address this limitation by learning the dictionary from data, predominantly using artificial neural network (ANN) autoencoder architectures. Although kernel methods offer an alternative with greater interpretability and tractability for theoretical analysis, they have received little attention in this setting. We introduce extended dynamic mode decomposition with kernel-based dictionary learning (EDMD-kDL), a kernel-based method for learning finite-dimensional Koopman embeddings directly from data. The method combines ideas from collocation methods and bilevel optimization to simultaneously learn a kernel dictionary and the corresponding Koopman approximation. We evaluate EDMD-kDL against state-of-the-art ANN-based approaches on a range of numerical experiments, including global sea-surface-temperature forecasting and learning directly from video data. Across all tested settings, EDMD-kDL achieves performance comparable to or better than the ANN-based methods. Moreover, in contrast to standard kernel methods, the proposed approach is scalable to large datasets by design since the size of the required kernel matrices depends on the number of collocation points rather than the size of the training dataset.
cs.LG / 142 / 2610.12379
Marformer: A Transformer for Predicting Missing Data Distributions
Prabhav Singh, Xiheng Tom Wang, Haojun Shi, Jason Eisner
cs.LG · stat.ME
Abstract
Real decisions are made under incomplete information. If we observe only some of the random variables we need, we can predict the others. The \textbf{conditional marginals} over the missing variables are the key ingredient for computing Bayes risk and Value of Information (VOI), the expected gain from acquiring one more observation before deciding. We present the Marformer, a Transformer trained to directly predict conditional marginals given any set of observed values. Like BERT, which is trained to predict missing words from context, the Marformer constructs a hidden-vector representation for each distribution $p(X_i)$ and iteratively refines it through attention to other distributions $p(X_j)$. Unlike generative approaches, the Marformer does not model the full joint distribution, requires no domain knowledge of the data-generating process, and makes all predictions in a single forward pass. We evaluate across three synthetic domains with missing data---Bayesian networks, discretized multivariate Gaussians, and structured annotation data. The Marformer can match or outperform classical missing-data methods, even when those methods are given the true model family and prior that generated the synthetic data. We also evaluate on a real annotation dataset, where the Marformer outperforms the evaluated baselines at the largest training size. In both cases, the Marformer is substantially faster than the evaluated generative baselines.
cs.LG / 143 / 2610.12397
Prospective Prediction of OOD Degradation from Source-Side Training Dynamics
Sasha, Monin
cs.LG
Abstract
We study whether persistent out-of-distribution (OOD) degradation can be predicted before it is directly observed using only source-side training dynamics. In a controlled shortcut-learning setting, a simple logistic regression predictor develops a clear prospective signal, while training time alone does not. Temporal summaries of the source-side quantities are substantially more informative than their current values. When transferred without additional training from a CNN to an MLP, confidence and entropy dynamics retain substantial predictive information. These results provide a proof of principle that source-side training dynamics can contain an early warning signal for future OOD failure.
cs.LG / 144 / 2610.12401
Learning Kilometer-Scale Weather Prediction with Global-Regional Alignment
Guowen Li, Yang Liu, Yujie Wang, Qiuyan Sun, Haoyuan Liang, Juepeng Zheng, Hong Cheng, Haohuan Fu
cs.LG
Abstract
Kilometer-scale regional weather forecasting is essential for local weather warnings and weather-sensitive decisions. Existing data-driven approaches often rely on numerical forecasts for large-scale guidance or require additional training of global forecasting components. Pretrained global weather models offer an efficient source of large-scale forecasts, motivating their reuse to guide high-resolution regional prediction. However, this coupling requires aligning global and regional representations across different grids and integrating global guidance with local interactions to advance regional states. We propose ScaleCast, a regional forecasting framework that addresses these challenges through Global-Regional Alignment. Its Global-Regional Conversion module aligns joint global and regional representations with regional locations, while the Global-Regional Alignment and Dynamics block combines aligned guidance with regional neighborhood interactions. Experiments using ERA5 global analyses on a 0.25-degree grid and CERRA regional reanalysis at 5.5 km spacing demonstrate improved regional forecasts across surface and upper-air variables, with a single trained model supporting multiple global forecast drivers (i.e., Pangu-Weather, GraphCast, and HRES) without specific retraining. Fine-tuning on HRRR at 3 km spacing further demonstrates the framework's adaptability to a different regional domain and spatial resolution. Windstorm case studies show improved cyclone positioning and core-pressure estimates, while comparisons with HadISD station observations show closer agreement with local temperature and humidity changes.
cs.LG / 145 / 2610.12420
A Unified Bellman Operator for Safety-Critical Reinforcement Learning
Nishanth Arun Rao, Royina Karegoudra Jayanth, Benjamin Eysenbach, Jaime Fernández Fisac
cs.LG
Abstract
Reinforcement learning in safety-critical domains requires maximizing task performance while strictly adhering to safety constraints. Existing safe reinforcement learning paradigms typically force a trade-off: they either require a priori knowledge to provide strict safety guarantees (e.g., safety filters), or they enable joint learning but only satisfy safety constraints on average. In this work, we propose a novel Bellman operator that unifies performance and safety objectives into a joint value function. We show that temporal difference learning with the joint Bellman operator converges under a two-timescale stochastic approximation framework. On the fast timescale, the safety value of the learning joint policy is estimated, while the joint value is estimated on the slow timescale. Convergence is ensured by formulating the limiting dynamics as an occupation-averaged differential inclusion, and showing that it asymptotically converges to a set of limiting optimal safety-constrained task value functions. Theoretically, once converged, the resulting optimal policy maximizes task return while maintaining safety at all times. Empirical evaluations on continuous control tasks with neural approximations demonstrate stable convergence with near-zero safety violations at test time.
cs.LG / 146 / 2610.12444
Rounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State Quantization
Hanyang Li, Shao Tang, Daniel Thomas Braithwaite, Gregory Dexter, Leonardo Neves, Aman Gupta, Hiroto Udagawa, Abhishek Shivanna, Daniel Silva, Rohan Ramanath
cs.LG
Abstract
Quantizing AdamW's optimizer states reduces persistent storage, but quantization errors propagate through the moment recurrences and perturb subsequent adaptive updates. We redesign 4-bit optimizer-state quantization for AdamW from the perspective of \emph{rounding space}: the coordinate in which a quantizer chooses between adjacent reconstruction levels. For the second moment, a local analysis of the quantization cell adjacent to zero shows that small mean state error need not imply small mean preconditioner error at the next step. A one-dimensional quadratic construction further shows qualitatively different optimization dynamics under state-space and preconditioner-space rounding. These results motivate Zero-Inclusive Preconditioner-space Stochastic Rounding (\textbf{ZIP-SR}), which retains zero in the second-moment codebook and computes stochastic-rounding probabilities in preconditioner space. As a complementary route, Zero-Excluding EDEN calibration (\textbf{ZE-EDEN}) uses a zero-excluding second-moment codebook and rescales the quantized second-moment block to mitigate the preconditioner distortion caused by the positive quantization floor. Both configurations use 4-bit NormalFloat (NF4) for the first moment, with targeted stochastic rounding of the LM-head first moment during the final 10\% of training. Across GPT- and Llama-style pretraining experiments ranging from \textbf{130M} to \textbf{2.7B} parameters, both methods reduce TorchAO 4-bit AdamW's mean validation-loss gap to 32-bit AdamW at every evaluated model size, with the largest reported gap reduction reaching \textbf{70\%}. In full-parameter supervised fine-tuning, both recipes achieve lower validation loss than TorchAO while remaining close to 32-bit AdamW on downstream tasks.
cs.LG / 147 / 2610.12445
Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception
Oskar J. Hollinsworth, Alex F. Spies, Tigist Diriba, Adam Gleave, Chris Cundy
cs.LG · cs.AI
Abstract
Recent incidents have highlighted the challenge of monitoring LLM agents and the danger of models deceiving people. We show that white-box deception detection via probes can be scaled up to frontier monitoring settings by collecting the largest deception dataset to date for training probes and introducing a novel probe architecture which can aggregate information across many layers and tokens. Our probes achieve 98.8% AUC in SHADE-Arena, surpassing an Opus 5.5 text-monitoring baseline, and show improved efficacy as the underlying model is scaled up. To push our probes to their limit, we test them on several cases where deception cannot be determined from the context alone. In these cases, which we refer to as introspective deception, the ground truth can only be determined through careful elicitation or thorough knowledge of a model's training data. In one such evaluation, we show that probes can distinguish transcripts containing a model's true hidden goal from other goals with an AUC of up to 99.7%. Our probes also readily detect deception on prominent open-weight models which lie about politically sensitive topics, and about their beliefs when put under pressure. We release our training dataset, dubbed FIBS, to help drive frontier deployment of effective probes, and encourage the community to expand upon it with further examples of deception and sabotage.
cs.LG / 148 / 2610.12449
Bi-FORK: Generative Modeling of High-Dimensional Bifurcating Systems
Anna Zimmel, Fleur Hendriks, Markus Holzleitner, Florian Sestak, Martin Weichselbaumer, Vlado Menkovski, Johannes Brandstetter
cs.LG · cs.AI · cs.CE · physics.comp-ph
Abstract
Bifurcations are ubiquitous in physical systems, from structural buckling to fluid and climate dynamics, yet they remain largely unexplored in deep learning. At a symmetry-breaking bifurcation, a single input admits multiple equally valid solutions, violating the one-to-one assumption underlying most learned physical surrogates. We introduce Bi-FORK, a generative framework for learning these one-to-many solution maps in high-dimensional systems. Bi-FORK generates complete trajectories through latent flow matching, preserving space and time coherence, and uses repulsion-guided sampling to recover distinct solution branches in a single amortized pass. We evaluate Bi-FORK on buckling beams, mechanical metamaterials, and Allen-Cahn phase separation, spanning continuous, discrete, and field-valued bifurcations with discretizations up to 260,000 points. Bi-FORK recovers the multimodal solution structure while scaling several orders of magnitude beyond prior approaches, opening generative modeling to high-dimensional bifurcating physical systems.
cs.LG / 149 / 2610.12336
asdex: Automatic Sparse Differentiation in JAX
Adrian Hill, Guillaume Dalle
cs.MS · cs.LG · math.NA
Abstract
Many tasks in scientific computing and machine learning require the Jacobian or Hessian matrix of a function. Automatic differentiation (AD) computes these derivatives to machine precision, but materializing a dense $m \times n$ Jacobian requires $n$ forward-mode or $m$ reverse-mode AD passes, one per column or row. For a large class of functions, each output depends on only a few inputs, making the derivative matrix sparse. Automatic sparse differentiation (ASD) exploits this structure in four steps: detection of the input-agnostic sparsity pattern, coloring of a graph to group columns or rows that can share an AD pass, compressed differentiation to compute a compressed derivative matrix with one AD pass per color, and finally decompression into the original sparsity pattern. The number of colors, and hence of AD passes, is often independent of the problem dimension: a banded Jacobian with $b$ contiguous bands, for instance, only ever requires $b$ colors, regardless of its size. asdex offers the first standalone ASD toolkit in the popular JAX ecosystem. With asdex.jacobian and asdex.hessian, it provides sparse drop-in replacements for jax.jacobian and jax.hessian.
cs.LG / 150 / 2610.10912
When Listening Becomes Easier: Scrubbing Visual Cues for Shortcut-Free VLAs
Jasper Gerigk, Kenzo Aspuru-Takata, Chin-Hsuan Wu, Mohammad Mohammadi, Shuhong Zheng, Igor Gilitschenski
cs.RO · cs.CV · cs.LG
Abstract
Shortcut learning is a prevalent issue in robot learning. The limited diversity of robot demonstration datasets can mislead policies into exploiting spurious correlations between tasks and irrelevant features, such as viewpoint or background. Collecting sufficiently diverse robot demonstrations is costly and inefficient, motivating algorithmic alternatives. We focus on vision-language-action (VLA) models and discover that different vision-language model backbones exhibit substantially different levels of susceptibility to visual shortcut learning. We find that model behavior correlates with our proposed representation-level metric, action margin, which requires no policy rollouts. Visual shortcuts consistently enter action representations in early layers, with models differing in the extent to which later layers correct them by incorporating language information. To boost models' attention to language, we introduce task scrubbing, a new domain-adversarial training method that decreases models' likelihood of using visual shortcuts and improves VLAs' generalization. Experiments in both simulation and the real world across multiple VLAs and visual cues show that task scrubbing improves out-of-distribution robustness and often eliminates visual shortcut learning.
cs.LG / 151 / 2610.10934
Higher-Order Morphology Priors for Quadruped Reinforcement Learning Under Actuator Degradation
Derek You, Zafir Shamsi, Keqin Wang, Christine Allen-Blanchette
cs.RO · cs.LG
Abstract
Actuator degradation turns quadruped locomotion into a coordination problem requiring joints to compensate for lost actuation. Prior work suggests that morphology-aware graph policies improve learning and generalization under body perturbations. We ask whether these benefits can be strengthened by explicitly modeling higher-order mechanical structure. We represent the Unitree Go1 as a cell complex with limb- and body-level rank-2 cells and apply Hodge-based message passing. Under degradation training, the node-edge-face Hodge actor achieves the highest return on unseen actuator degradations, with higher survival and lower velocity-tracking error. These results support higher-order morphology as a useful inductive bias for whole-body compensation under actuator degradation.
cs.LG / 152 / 2610.11283
Being-M0.7: A Latent World-Action Model for Humanoid Robots
Junpeng Yue, Boyuan Li, Yuxuan Wang, Zepeng Wang, Yuhui Fu, Feiyang Xie, Yu Zhang, Jing Zhang, Xianqi Zhang, Weibo Li, Xiaofei Zheng, Yuming Fang, Jiangxing Wang, Zongqing Lu
cs.RO · cs.CV · cs.LG
Abstract
Humanoid loco-manipulation requires coordinated locomotion and manipulation informed by future scene evolution and whole-body motion, yet learning these capabilities is constrained by scarce robot demonstrations. Human video and motion datasets offer scalable supervision, but many contain only video or motion rather than paired video-motion data. Moreover, human motion does not directly specify executable robot actions. We present Being-M0.7, a latent world-action model that transfers visual-motion priors learned from mixed-modality human data to humanoid control through pre-training, robot mid-training, and action post-training. We curate a corpus from more than 10,000 hours of raw human-centric data, integrating video-only, motion-only, and paired video-motion streams to learn complementary visual dynamics and whole-body kinematic structure. Joint prediction of future latent visual states and motion encourages visual representations to encode future kinematics. Robot mid-training adapts this coarse-grained prior to robot viewpoints and body dynamics. During action post-training, an action expert combines visual predictive representations from the frozen, adapted prior with current images and proprioception through gated cross-attention, grounding predictive context in executable whole-body commands. Being-M0.7 achieves the highest aggregate success rate among the compared baselines on SIMPLE and matches the strongest baseline on real-world Unitree G1 loco-manipulation tasks.
cs.LG / 153 / 2610.11382
PlanWAM: Planning-Shaped Future Representations for End-to-End Autonomous Driving
Jinchang Xu, Hongda Yu, Fengwei Dong, Wenhui Huang, Xi Wei, Yongzhi Liu, Sunan Zhang, Jirao Wang, Chen Lv, Bingbing Li, Guodong Yin, Weichao Zhuang
cs.RO · cs.LG
Abstract
World models in end-to-end autonomous driving predict future scene evolution to provide foresight for trajectory planning. Existing methods mainly study how to predict the future and how to use it, but less often ask which future representation is actually most useful for planning. To this end, we propose PlanWAM, a Planning-Shaped World Action Model. The key idea is to let the planning task shape the future-state representation, so that it retains the information most useful for planning. A latent world model then predicts this planning-shaped future latent representation from historical observations and uses it for planning, enabling foresighted planning. Specifically, we first use a Temporal Register Pyramid to compress multi-frame historical information in a recency-aware manner, learning a compact history representation oriented toward future reasoning and planning. We then introduce a privileged future posterior branch that observes ground-truth future frames, and shape its future latent representation with trajectory-planning objectives to obtain a planning-shaped future latent representation. Hindsight-to-Foresight Distillation trains a prior branch that depends only on history to predict this future latent representation. The predicted future latent representation serves as planning context and guides trajectory generation and selection. PlanWAM achieves 93.8 PDMS / 90.9 EPDMS on NAVSIM-v1/v2 navtest and reaches 38.7 HD-Score on closed-loop HUGSIM in a zero-shot setting, demonstrating leading planning performance across both open-loop and closed-loop evaluations. Extensive experiments further demonstrate that planning-shaped future representations provide an effective and deployable form of foresight for world-action models.
cs.LG / 154 / 2610.11945
TACROSS: An Efficient and Low-Cost Scalable Human Touch System Across Heterogeneous Tactile Sensors for Dexterous Robot Learning
Bo Chen, Huanzhang Hu, Junyang Ma, Bo Yue, Fangdi Yu, Haijier Chen, Xianxin Lai, Shuyu Pan, Zhen Yang, Xiaoquan Sun, Wenze Cui, Zhongliang Jiang, Shaopeng Liu, Jiayu Chen
cs.RO · cs.LG
Abstract
Collecting tactile demonstrations on robots is costly and slow, motivating the use of lower-cost human tactile gloves for scalable data collection. However, human capacitive/piezoresistive gloves and robotic tactile sensors differ fundamentally in transduction principle, sensor layout, spatial resolution, and dynamic response, making alignment of raw sensor channels ill-posed. To address this problem, we present TACROSS, a scalable system for learning from human touch and transferring it to robots that bridges this heterogeneity by aligning tactile streams at the level of contact events rather than raw sensor values. The hardware component of TACROSS integrates a piezoresistive glove with five layers and a cost of USD 10.86 with 285 sensing points. To align contact semantics, we design canonicalizers and residual adapters that map heterogeneous signals into a shared tactile latent with 256 dimensions via a temporal Transformer with attention across fingers. We further introduce a robot-grounded policy learning scheme in which robot demonstrations provide the sole source of ground-truth action supervision, while human demonstrations support tactile representation learning and provide confidence-weighted auxiliary supervision through valid retargeted hand targets. We evaluate our system on four contact-rich manipulation tasks. Compared to conventional teleoperation, our proposed system achieves a 3.5-fold efficiency improvement while reducing demonstration acquisition equipment cost by 95.7%. We will open-source the TACROSS hardware and software system and publicly release a tactile dataset comprising over 150 hours of recordings. Project page: https://tacross-touch-project.github.io/.
cs.LG / 155 / 2610.11971
CAPABLE: Capability-Aware Policy Adaptation via Behavioral Latent Encoding
Mohammad Khoshnazar, Mohammad Dehghani Tezerjani, Deyuan Qu, Zhiyuan Gao, Yanxiang Zhan, Jeroen Schafer, Andrew Melnik, Qing Yang, Michael Beetz
cs.RO · cs.LG · eess.SY
Abstract
Vision-language-action (VLA) policies assume the embodiment on which they were trained and can fail when a joint fault changes how commanded actions are physically executed. Existing fault-recovery methods often require task-specific retraining, fault labels, explicit diagnosis, or privileged embodiment information. We introduce CAPABLE, a unified capability-aware adaptation framework for frozen VLAs that integrates self-supervised capability inference with residual reinforcement learning. CAPABLE infers capability, how much of the commanded motion each joint actually realizes and how that motion contributes to end-effector behavior, online from command-response history and kinematics using a temporal encoder shared across joints, Jacobian grounding, cross-joint attention, and self-supervised physical prediction. The resulting representation conditions a residual policy that adds bounded corrections to the VLA arm action without fault labels or faulty-joint identifiers. Across 28 LIBERO tasks, CAPABLE raises success on an actuator excluded from fault training from 24.8% to 59.3%, outperforming a parameter-matched global-history baseline by 17.4 points while preserving healthy performance. Leave-one-actuator-out experiments across six joints show that this transfer is not specific to one actuator, and additional evaluations characterize transfer to unseen fault families and demonstrate recovery on a physical Franka Panda. https://capable-vla.github.io/
cs.LG / 156 / 2610.12432
FAITH: Feasibility-Aware Safety-Filtered RL for High-Dimensional Systems
Songyuan Zhang, Baljeet Singh, Sarthak Ranjeet Kaingade, Chuchu Fan, Bryan Trinh
cs.RO · cs.LG · eess.SY
Abstract
Safe reinforcement learning commonly places safety and task performance in the same policy objective, where they can introduce competing updates. Safety filters separate them at action execution, but classical designs require an analytic safety function and dynamics model, and standard minimal-intervention filters are myopic to long-horizon task return because they minimize only instantaneous action deviation. Hard projections are also undefined when no safe action exists. We present FAITH, a feasibility-aware, model-free framework that approximates the optimal state-action safety value and amortizes minimal-intervention filtering with a feedforward network. The task policy optimizes the task return through the filtered dynamics, which recovers the feasible constrained problem without a competing safety term in the task-policy update. When no action satisfies the learned safety condition, the same filter approaches the action with minimum predicted peak harm. On a double integrator example and a Safety Gym environment, FAITH achieves the highest return among methods with no feasible-start violations and matches the lowest harm from infeasible starts. On a 29-DoF humanoid, it reaches a 99.95% safety rate while retaining 97% of the unfiltered return in Walking-Avoid, and obtains the highest measured safety rate in Push-Avoid by learning to sacrifice balancing and fall away from the protected region. The same policies are also demonstrated on a real-world Unitree G1 humanoid.
cs.LG / 157 / 2610.12435
VioLA: Learning Generalist Humanoid Control Policies from Human Data
Mert Albaba, Jens Beißwenger, Anna Manasyan, Daniel Marta, Michael J. Black, Wieland Brendel, Andreas Krause, Georg Martius, Martin Riedmiller
cs.RO · cs.LG
Abstract
Teaching a humanoid to follow instructions with its whole body runs into two obstacles. Its action space is large and tightly coupled: legs, arms, and fingers must move together while the robot keeps its balance, which makes joint-level actions hard to learn. And humanoid demonstrations are scarce, so current humanoid generalist policies do not follow new instructions out of the box and are fine-tuned on teleoperated demonstrations of each task before deployment. Human demonstrations exist in far larger numbers, but a person's motion is not a robot command. We remove both obstacles by changing what the generalist policy predicts. We introduce VioLA, a generalist humanoid policy that predicts body and hand motion latents instead of joint commands. A pretrained body- and hand-controller execute these latents on the robot. Their corresponding motion encoders map human and robot motion into the same latent spaces. A human recording is therefore labeled in the policy's action space, and the training demonstration pool contains 140.6 million frames, 93.2% of them human. As a result, VioLA follows locomotion instructions on the real robot zero-shot, without task-specific fine-tuning, reaching 100% success where GR00T N1.7 and $Ψ_0$ reach 16.7% and 0%, respectively. It also reaches 88.6% manipulation success without task-specific fine-tuning. The same approach works across two VLA and one world-action model backbones. A generalist policy trained on human demonstrations alone performs locomotion tasks on the real robot zero-shot. Code and checkpoints will be released.
cs.LG / 158 / 2610.12465
A Balanced Data Diet: Addressing the Exploration Bottleneck in Mega-Scale RL for Robot Control
Octi Zhang, Mateo Guaman Castro, Patrick Yin, Ignacio Dagnino, Abhishek Gupta, Rosario Scalise, Byron Boots
cs.RO · cs.LG
Abstract
General-purpose robots must perform a wide range of tasks from agile locomotion to dexterous manipulation. While sim-to-real reinforcement learning (RL) has proven to be a useful tool for this goal, current RL pipelines depend on engineering-heavy, per-task structural priors such as shaped rewards and demonstrations. Recent work has shown that diverse simulator resets, combined with massively parallel simulation, can alleviate much of this engineering burden on several manipulation problems. However, we find that naively scaling this paradigm to more precise or dynamic problems remains non-trivial. While simulator resets can help with exploration, uniformly sampling over this distribution wastes a growing fraction of learning experience on task configurations the policy has already mastered or cannot yet attempt. This makes it challenging to see the expected benefits of scaling parallel environments for RL, since much of the learning signal in a batch is wasted during learning. To mitigate this, we introduce Success Guided Sampling (SGS), a simple adaptive sampler that concentrates RL training on task configurations around the frontier of the policy's capabilities. Doing so allows large-scale simulated RL to make the most out of the experience in a batch, enabling much more effective scaling to large-scale parallel simulation. Across experiments using up to $2^{20}$ (over one million) parallel environments, SGS enables RL to solve challenging multi-terrain quadruped locomotion and contact-rich assembly tasks that prior methods fail to solve. Finally, we distill the learned manipulation policies into RGB-based policies and demonstrate zero-shot transfer to several challenging assembly tasks on real hardware. Project website: https://sgs-rl.github.io/.
cs.LG / 159 / 2610.12467
CSF: Contextual Safety Filtering for Motion Generators
Lizhi Yang, Yiling Hou, Yao Tang, Junheng Li, Daniel Weng, Blake Werner, Aaron D. Ames
cs.RO · cs.LG
Abstract
Text-conditioned motion generators produce trackable whole-body motion, but they have no notion of scene-dependent safety: the same action may target an object or a person. Existing safeguards either inspect the prompt, require labeled motion data, or enforce geometric constraints; therefore, they do not directly account for how scene context changes a motion's meaning. We introduce contextual safety filtering (CSF), a training-free filter that grounds natural-language safety rules in safe and unsafe reference trajectories produced by the generator. For each active rule, safe and unsafe reference trajectories define an affine safety value that a safe reference tracking CBF-QP enforces. Across four pretrained generators with different architectures, CSF activates the intended rules in all explicit and scene-triggered unsafe cases and reduces the danger-event rate by up to 90%, while preserving 88-100% of benign motions. We demonstrate the complete system on a real-world Unitree G1, where it successfully prevents unsafe motions in a variety of scenarios, including interactions with humans and objects.
cs.LG / 160 / 2610.11422
Causal-fate dynamics of unrealized influence
Yiwei Liu, Luwei Yang, Shunbo Lei
eess.SY · cs.LG
Abstract
Many dynamical systems generate influences whose consequences are not fully exhausted in the realized trajectory at the moment they arise. Such consequences are often treated as absent, delayed or statically stored, leaving unclear how unrealized influence retains future relevance as the system evolves. Here we formulate causal-fate dynamics, in which generated influence may be realized, remain latent, or be transformed by subsequent dynamics, and give an exact finite-transport representation when the relevant maps are specified. A connectome-constrained Caenorhabditis elegans model first motivates the biological hypothesis that unresolved inter-neuronal influence may persist and contribute to later propagation; it does not establish such a mechanism in living animals. We next examine operational Internet routing, where a dynamically updated cross-observer history retains predictive information beyond the current local route state. We then use the representation to construct a Transformer architecture that explicitly transports and selectively realizes latent contextual influence while retaining language-modeling function. The three studies distinguish a model-motivated scientific hypothesis, an observational phenomenon compatible with future-relevant history and an executable construction for carrying unrealized influence through subsequent computation.
cs.LG / 161 / 2610.11619
Randomized Transport Maps for Model-Free Policy-Gradient Mean-Field Control
Adonis Jamal, Samy Mekkaoui, Yadh Hafsi, Huyên Pham
math.OC · cs.LG
Abstract
We develop a model-free policy gradient method for discrete-time mean-field control (MFC). In MFC, the policy affects the objective both through the controlled dynamics and through the population distribution. Standard REINFORCE estimators capture the first effect but not the second. We introduce Transport REINFORCE, a transport map-based approach that perturbs a suitable transformation of the population distribution to estimate this missing mean-field contribution. The method applies to both finite and continuous state spaces. In finite state spaces, we perturb the population distribution directly on the probability simplex through a convex combination of the current population weights and random weights. In continuous state spaces, we project the population distribution onto the manifold of Gaussian mixtures, and then randomize it via a transport map that ensures the perturbed law remains within this manifold. We prove consistency of the perturbed objective and gradient as the perturbation vanishes, and derive bias and mean-square error bounds for the resulting sample-based gradient estimator. Numerical experiments on several MFC benchmarks show that Transport REINFORCE improves over standard REINFORCE.
cs.LG / 162 / 2610.12358
Subspace Uncertainty and Sharp Sampling Thresholds on the Boolean Cube
Thomas Weinberger
math.PR · cs.IT · cs.LG · math.CO
Abstract
We study Gaussian regression under squared population $L_2$ loss in a known $m$-dimensional subspace of degree-at-most-$k$ functions on the $d$-dimensional Boolean cube. Random inputs can undersample regions essential for prediction, delaying the parametric rate even when the model is known. For fixed $q_0<1/2$, $1\le k\le q_0d$, and sufficiently large fixed $A$, the worst-subspace sample threshold for minimax error $Aσ^2(m+t)/n$ with confidence $1-e^{-t}$, $t\ge\log4$, is \[ N=(m+t)\exp\{E_{d,k}+O(k^{1/3})\}, \quad E_{d,k}=dΨ(k/d), \] where $Ψ(q)=\log2-\mathsf H(\tfrac12-\sqrt{q(1-q)})$ and $\mathsf H$ is binary entropy with natural logarithms. The upper bound holds for every feasible $m$; the matching lower bound holds when $m\le\binom d{\lfloor k^{1/3}\rfloor}$ or $t\ge m$. We sharpen the Polyanskiy--Samorodnitsky uncertainty principle in two respects. First, for fixed leakage $ρ\in(0,1)$, the smallest set carrying a fraction $1-ρ$ of a nonzero degree-at-most-$k$ polynomial's energy has probability $\exp\{-E_{d,k}+O_{ρ,q_0}(k^{1/3})\}$. An Airy-kernel construction proves that the remainder cannot be $o(k^{1/3})$ in general. Second, we construct a subspace of dimension $\binom d{\lfloor k^{1/3}\rfloor}$ such that every function in the subspace has at least a fraction $1-ρ$ of its energy on the same set, whose probability is at most $\exp\{-E_{d,k}+C_{ρ,q_0}k^{1/3}\}$. For sufficiently large $k$, this set is a Hamming ball. A striking consequence is an exponential cost of noise: the parametric rate can require $(m+t)4^k\exp\{-O(k^{1/3})\}$ samples, whereas $O((m+t)2^k)$ suffice for noiseless identification. As $k\to\infty$ with $k/d\to0$, the noisy threshold is $(m+t)\exp\{2k+o(k)\}$.
cs.LG / 163 / 2610.11772
Optimal random quantisers for spherically symmetric distributions
Luc Pronzato, Anatoly Zhigljavsky
math.ST · cs.LG · stat.ML
Abstract
Zador's celebrated theorem is a cornerstone of optimal quantisation: it establishes both the weak limit of the empirical distribution of an optimal $n$-point quantiser in $R^d$ and the decay rate of the associated $L_s$-mean quantisation error. In large dimension, however, observing this asymptotic behaviour requires an astronomically large sample size. We prove that, for spherically symmetric target distributions, optimisation over all spherically symmetric distributions is a convex problem and derive an equivalence theorem that both characterises global optimality and yields a constructive algorithm. We show that, for moderate $n$, random quantisers uniformly distributed on a sphere of suitably chosen radius $R$ perform exceptionally well and, over a broad range of values of $n$, are numerically certified to be optimal among all random quantisers. Their expected distortion has an explicit integral representation that can be evaluated to arbitrary precision, and we prove concentration across random quantisers: the distortion variance tends to zero as $n\to\infty$ for fixed $d$. For $s=2$, both the optimal radius and the associated minimum expected distortion admit exact expressions. For general $s$, the optimal radius can be determined efficiently, and extreme-value theory provides useful approximations when $n$ grows with $d$. Depending on this growth rate, $R$ either converges to zero or approaches a positive limit that is independent of $s$.
cs.LG / 164 / 2610.11022
A Graph Neural Network for Global Daily Fire Radiative Power Prediction at Medium-Range Lead Times
Li Zhang, Jun Wang, Isidora Jankov, Yongxin Liu, Gonzalo A. Ferrada, Ravan Ahmadov, Ligia Bernardet, Haonan Chen, Shobha Kondragunta
physics.ao-ph · cs.LG
Abstract
Skillful prediction of biomass-burning activity several days in advance is important for air-quality forecasting and aerosol prediction. Two operational constraints motivate this work. First, the GBBEPx satellite fire radiative power (FRP) product used to initialize NOAA's GEFS-Aerosols is available with about a 1.5-day latency, so each forecast cycle relies on the most recently available, but already outdated, fire observations. Second, these fire inputs are then held fixed throughout the subsequent 5-day operational forecast, or 7 days in the GSL experimental system, effectively assuming no evolution in fire activity. We develop a data-driven model that predicts global FRP one to seven days ahead from the most recent available observations. The model adapts a spatiotemporal graph neural network using reanalysis meteorology, land-cover and vegetation information, recent fire history, and GBBEPx FRP as the training target. It is trained on 2020-2022 data and evaluated for 2023-2024. The model reproduces the global seasonal cycle and substantially outperforms persistence. At 0.1$^\circ$ resolution, mean squared error is reduced by 32% at one-day lead and 43% at seven days in 2023, and by 24% and 40% in 2024. At 1$^\circ$ resolution, the critical success index ranges from 0.32 to 0.60. Detection skill declines only modestly with lead time, whereas intensity skill degrades more rapidly. Large fires are detected reliably, but their radiative power is systematically underestimated. These results demonstrate useful predictability of fire activity several days ahead and identify intensity calibration and small-fire placement as the main remaining challenges before predicted FRP can support operational aerosol forecasts.
cs.LG / 165 / 2610.12009
Ghost tasking for parametrized Gaussian Processes solving linear differential equations
Johanna Moser, Christopher Albert, Sascha Ranftl
physics.comp-ph · cs.LG
Abstract
Physics-informed machine learning has gained significant attention in recent years. In regimes of limited data, parametrized Gaussian processes have become popular. Existing approaches, however, often face limitations, such as requiring parametrizable (also called controllable) systems or a large number of output tasks. In this work, we introduce a systematic procedure we call "ghost tasking", using auxiliary tasks to circumvent these limitations. We prove that such ghost tasks can render any non-parametrizable system effectively parametrizable, enabling algorithmic construction of parametrized Gaussian Processes while keeping the number of required tasks (i.e. output dimensions) and latent functions low. We find that ghost tasking performs especially well in an inverse problem setting, even with very few available data. We show the usage and power of ghost tasking in three experiments, providing systematic comparisons to the only other currently available method applicable to all experiments. We provide necessary syntax and explications for two computer algebra programs that compute parametrizations for systems with polynomial or rational coefficients. Our theoretical results extend to systems with meromorphic functions.
cs.LG / 166 / 2610.11500
Learning qBIC Resonances across Metasurface Families in Dielectric Fourier Space
Shuangteng Lei, Li Yu, Tianxin Li, Wei Lu
physics.optics · cs.LG · physics.comp-ph
Abstract
Bound states in the continuum (BIC) metasurfaces are typically described by geometry-specific parameters, hindering cross-geometry comparison, while ultranarrow qBIC features are easily diluted in full-spectrum learning. Here, 2015 samples from seven dielectric metasurface families are mapped to a shared reciprocal-lattice grid, where two frozen low-order Fourier channels capture resonance shifts with mean within-branch $R^2$ values of 0.871-0.999. Field-level analysis of two representative branches further confirms that these shifts are consistent with the Maxwell-Fourier perturbation picture. A five-channel K-space backbone models the broadband spectrum, while a local complex K-space expert parameterizes the qBIC resonance through a differentiable Fano layer. The expert reduces resonance-position mean absolute error (MAE) from 3.2 to 0.95 nm and the resonance-depth error by 14-fold on a geometry-blocked test set. The same coordinate supports spectrum-to-structure reconstruction.
cs.LG / 167 / 2610.11004
Reconstruction of Multiscale Plasma Dynamics Across Operating Regimes
Maryam Reza, Farbod Faraji
physics.plasm-ph · cs.LG
Abstract
Reconstructing spatially resolved plasma dynamics from few sensors is essential for diagnostics, reduced-order modelling and control, yet remains difficult because the sparse measurements incompletely constrain multiscale, regime-dependent degrees of freedom. The Shallow Recurrent Decoder (SHRED) partially addresses spatial sparsity by using measurement histories; however, its fully connected decoder provides no explicit mechanism for resolving spatial structure across scales or explicit parametric dependency. We introduce the Recurrent Multiscale Affine-modulated Inference Network (ReMAIN), which preserves SHRED's recurrent temporal encoding but replaces its decoder with a U-Net whose feature hierarchy is conditioned by the recurrent state through feature-wise linear modulation. The temporal representation supplies both a dense prior and scale-specific modulation throughout the U-Net. A parametric extension jointly embeds the operating condition and sensor history, enabling reconstruction to adapt as the governing dynamics change with operating regime. ReMAIN is first benchmarked against SHRED on six one-dimensional nonlinear PDEs representing diverse dynamics. Across all benchmarks, it reduces reconstruction errors on unseen trajectories and more faithfully resolves sharp transitions, localized extrema and fine-scale variations. The parameter-conditioned model is then demonstrated on a collisionless $E \times B$ plasma subject to perpendicular axial electric and radial magnetic fields, with the electric-field strength serving as the operating parameter. ReMAIN reconstructs the high-dimensional, multiscale plasma state and recovers its regime-dependent spatiotemporal dynamics at electric-field strengths withheld from training. Together, ReMAIN improves sparse-sensor full-state reconstruction and, through parameter conditioning, generalizes across plasma operating regimes.
cs.LG / 168 / 2610.11222
Cross-species representation learning aligns mouse and human neural dynamics and tracks clinical drug efficacy
Marko Tvrdic, Justin Richmond Domingo, Jae Ann Buenaluz, Jydell Ashley Palomo Penollar, Edmayelle Villavicencio Alforja, Jobi Fallaeria Subosa, Gabriel Ocana-Santero
q-bio.QM · cs.LG · q-bio.NC
Abstract
Preclinical models poorly predict human drug efficacy, particularly in neurological disorders. Neural activity offers a uniquely rich source of translational information because it captures high-dimensional variation in nervous-system function that can be measured in both animals and humans. However, its high dimensionality makes it difficult to distinguish conserved disease-related features from variation arising from species, recording modality and experimental context. Here, we test whether shared neural dynamics can be identified directly from electrophysiology data by learning representations organized by biological state rather than species. We develop a dual-rule contrastive learning framework that aligns corresponding mouse and human states while preserving separation between distinct phenotypes. This framework recovered conserved sensory-response structure across species and, in epilepsy, resolved distinct relationships between three mouse models and heterogeneous human patient populations. When treated animals were projected into a frozen cross-species representation, drug-induced movement towards the human-aligned healthy state retrospectively tracked known clinical efficacy across ten model-drug combinations including a disease-specific detrimental effect. The framework also identified shared disease-associated neural dynamics between Fmr1-knockout mice and human 16p11.2 copy-number variant carriers despite differences in genetic aetiology and recording modality. Together, these findings show the potential of cross-species neural representation learning to map heterogeneous human disease onto experimentally tractable preclinical states and assess whether interventions restore human-relevant circuit function.
cs.LG / 169 / 2610.10727
Deep Learning vs. Statistical Models for Multi-Horizon Price Forecasting of Second-Hand Electronics: A Systematic Benchmark
Mateusz Buczyński, Michał Woźniak, Konrad Kaczyński, Anna Wróblewska, Sebastian Kuk
q-fin.CP · cs.LG
Abstract
Forecasting resale prices of used electronics is critical for subscription-based platforms where pricing errors translate directly into risk. Unlike structured financial markets, second-hand electronics exhibit high volatility, sparse listing histories, and non-normal price dynamics - yet no systematic time-series benchmark exists for this domain. This paper presents the first multi-horizon benchmark of statistical and deep learning forecasting models for used electronics price prediction. We use a large-scale dataset of daily price listings from Polish online marketplaces (January 2022 to March 2025, 100+ smartphone and laptop models) and evaluate eleven models across six horizons from 1 to 365 days, covering classical methods (ARIMA, ETS, Theta), recurrent and convolutional networks (LSTM, TCN), and modern deep architectures (N-BEATS, N-HiTS, TFT, PatchTST, Informer). Three complementary evaluation protocols assess trajectory fitness, one-shot endpoint accuracy, and cross-horizon transfer. N-BEATS achieves the lowest MAPE beyond 30 days, reaching 8.51% at 365 days versus 14.94% for the best statistical baseline - a 43% reduction. At short horizons (1-7 days), all models converge near 0.72% MAPE and the naive baseline remains competitive. A single N-BEATS model trained at 365 days generalizes to all shorter horizons, eliminating the need for horizon-specific models. N-BEATS and N-HiTS also demonstrate superior hyperparameter stability.
cs.LG / 170 / 2610.10697
On linearity or non-linearity in machine learning for quantum chaotic dynamics
Francesco Perciavalle, Agostino Gallo, Francesco Plastina, Gianluigi Greco, Nicola Lo Gullo, Carlo Adornetto
quant-ph · cond-mat.stat-mech · cs.LG
Abstract
Accurately simulating chaotic quantum many-body dynamics remains a major computational challenge for classical methods, due to the rapid buildup and spatial spreading of entanglement during the evolution. This raises the question of whether machine learning can provide an effective alternative for predicting quantum dynamics. We address this question by formulating quantum dynamics as a time-series forecasting problem, using a few-qubit PXP chain, realizable with Rydberg-atom arrays, as a benchmark. By varying the initial state, the system spans dynamical regimes ranging from ergodic behavior to quantum many-body scarring, providing a controlled setting for testing forecasting models across qualitatively different dynamics. We compare two contrasting architectures: an expressive nonlinear Transformer and DLinear, a simple linear forecasting model. The Transformer accurately predicts dynamics in the more ergodic regime, but its performance progressively deteriorates as the initial state approaches the scarred limit. In contrast, DLinear remains accurate across the entire family of initial states, with its main deviations consisting of small high-frequency oscillations that have little effect on the overall prediction error. Remarkably, these results show that observables generated by complex quantum many-body dynamics can be forecast with high accuracy through a simple linear mapping from past to future observations. This reveals that the complexity of the underlying quantum evolution need not translate into an equally complex forecasting problem.
cs.LG / 171 / 2610.11261
Q-Capsule: A Localized Capsule-Based Quantum Neural Architecture for Barren Plateau Mitigation
Awal Ahmed Fime, Tasfia Zaman Samiha, Saika Zaman, Dimitris Pados, George Sklivanitis, Abdur R. Shahid, Ahmed Imteaj
quant-ph · cs.LG
Abstract
Variational quantum algorithms are often limited by barren plateaus: gradients vanish as circuit size and depth increase, making quantum neural networks difficult to train. We propose Q-Capsule, a localized capsule-based quantum neural architecture that mitigates this problem through register partitioning, local readout, sparse inter-capsule coupling, trainable data re-uploading, and Quantum Fisher Information Matrix (QFIM)-guided adaptive depth growth. By restricting the dominant support of each observable to a small capsule and controlling inter-capsule entanglement, Q-Capsule preserves useful gradient signals while retaining communication between local quantum representations. As the register width increases, Q-Capsule consistently maintains stable gradient variance, whereas globally entangling baselines exhibit exponential suppression with a log-gradient-variance slope near -ln 2 per qubit. Q-Capsule also produces more structured optimization landscapes, higher parameter efficiency, improved robustness to depolarizing noise, and lower measurement requirements. Its adaptive policy achieves 98.1% accuracy on binary classification and 97.7% on four-class classification, while using approximately 73% fewer two-qubit gates than the fixed-deep model on the multiclass task.
cs.LG / 172 / 2610.11537
An Efficient Quantum Circuit for Flow Model Execution Using Quantum Neural Networks
Rui Che, Ludvig af Klinteberg
quant-ph · cs.LG
Abstract
Flow models generate trajectories from an initial distribution to a target distribution by solving an ordinary differential equation defined by a velocity field. Flow matching learns this velocity field by modeling the transport dynamics between the two distributions. Wavefunction flow establishes a formal connection between flow models and quantum dynamics by introducing a continuity Hamiltonian, which drives the Schrödinger evolution of quantum states. In this paper, we investigate accurate and efficient quantum simulation of the wavefunction flow, thereby realizing the efficient implementation of flow models on quantum computers. We first leverage a quantum read-only memory (QROM)-based phase kickback framework for the wavefunction flow simulation, generating probability densities that closely match those produced by the corresponding conventional flow model. To address the high circuit-resource cost, we further incorporate a trained quantum neural network (QNN) into the phase kickback framework, replacing QROM for data encoding. Numerical experiments demonstrate that our proposed method implements flow models on quantum computers more efficiently, since it maintains the accuracy of wavefunction flow simulation compared with the QROM-based framework, and significantly reduces the circuit resources.
cs.LG / 173 / 2610.11759
Beyond QAOA: A Review of AI and Quantum Computing for Adaptive Combinatorial Optimization
Hoong Chuin LAU
quant-ph · cs.LG · math.OC
Abstract
Near-term quantum approaches to combinatorial optimization are limited by qubit counts, circuit fidelity, sampling cost, and the difficulty of encoding constraints, while machine learning is increasingly used to configure and control quantum optimization workflows. We call such workflows adaptive: decisions conventionally fixed in advance, from formulation and penalties to shot budgets, backends, and whether to invoke a quantum processor at all, are made by learned policies that respond to the instance, the progress of the solve, or the hardware. This review examines three paradigms, AI for quantum optimization, quantum for AI-driven optimization, and AI-quantum co-optimization, and organizes the literature by the decision being learned rather than by application. A structured review of 119 papers, 67 coded in detail, shows that the evidence is considerably stronger for AI-assisted quantum optimization than for the reverse direction: learning already reduces quantum evaluations, improves initialization, supports decomposition and penalty control, and mitigates noise, whereas evidence that quantum computation improves learned optimizers remains largely confined to small-scale simulation. Experimental controls are thin: 25 of 57 studies include no classical baseline, the quantum contribution is fully isolated in 10 of 24 studies where an ablation applies, and the median experiment uses 17 qubits. We introduce an M0-M5 evidence hierarchy, from simulation to matched-resource practical advantage, and find no broadly convincing result at the highest level. We argue that scaling is increasingly a systems problem: the question is not only whether a problem fits on a quantum processor, but how classical and quantum resources should be allocated across the optimization process. The review is aimed at researchers in quantum computing, machine learning, and operations research.
cs.LG / 174 / 2610.12428
Toward Joint Optimization of Circuit Depth and Training Data Size in Adaptively Grown Quantum Classifiers
Saeefa Rubaiyet Nowmi, Md Mahmuduzzaman Kamol, Mohammad Saidur Rahman
quant-ph · cs.LG
Abstract
Building a quantum model involves a tradeoff: how complex the circuit should be, and how much training data it needs. Caro et al. show that models with fewer trainable gates need less training data to generalize well. Q-FLAIR shows that a quantum feature-map circuit can be grown gate-by-gate, stopping once further growth stops improving the training loss. We ask whether these two results combine into a predictable scaling law. Does Q-FLAIR's own stopping rule pick larger or smaller circuits as training data grows? Does the resulting generalization behavior track Caro et al.'s bound? We reimplement Q-FLAIR's growth mechanism faithfully, including its analytic reconstruction and exact stopping rule. We run it on full-resolution (784-pixel) MNIST 3-vs-5 classification, at five training-set sizes from N = 2000 to 10000. We then fine-tune each resulting circuit, so we can measure Caro et al.'s notion of active gates, K. We find no predictable relationship between training-set size and the circuit size Q-FLAIR converges to. Circuit size and test accuracy both vary non-monotonically with N, and seed-to-seed variance is nearly as large as any trend across N. The empirical generalization gap never exceeds Caro et al.'s bound in 14 of 15 runs, so the bound holds as a valid guarantee in those runs. But the gap correlates only weakly with the bound's value (r = 0.12). This shows that K does not explain most of the variation we observe. Why a valid guarantee can coexist with such weak predictive power remains an open question, and answering it may be necessary before circuit depth and training data size can be jointly optimized in practice.
cs.LG / 175 / 2610.10761
What can linear attention learn from nonlinear teachers in-context?
Mary Letey, Arman Rysmakhanov, Yue M. Lu, Cengiz Pehlevan, Jacob Zavatone-Veth
stat.ML · cs.LG
Abstract
Linear attention is a tractable model for understanding the mechanisms governing in-context learning in transformers. For linear regression tasks, recent asymptotic analyses have characterised its learning and generalisation behaviour. We extend this theory to nonlinear single-index targets, $y=f(x^\top w)+\varepsilon $. Our main result establishes a nonlinearity-noise equivalence: linear attention extracts only the linear Hermite component of $f$, while the remaining nonlinear structure contributes to the generalisation error as effective noise. This reduction allows results from the corresponding linear theory to be transferred to nonlinear tasks. We illustrate its implications for finite pretraining data and for the transition from task memorisation to task generalisation as task diversity increases. These results identify a limitation of the reduced linear-attention model and provide a tractable starting point for studying nonlinear in-context learning.
cs.LG / 176 / 2610.10793
Calibrating Ambiguity Set via Diagnostic Transport for Distributionally Robust Optimization
Wenbin Zhou, Elizabeth Cucuzzella, Shixiang Zhu
stat.ML · cs.LG · math.OC
Abstract
Distributionally robust optimization (DRO) protects decisions against distributional uncertainty by optimizing over an ambiguity set, but poorly aligned set geometry can require large radii and yield overly conservative decisions. We introduce diagnostic-transport DRO (DT-DRO), which uses held-out calibration data to adapt the ambiguity-set geometry to observed predictive errors. DT-DRO uses the conditional probability integral transform cumulative distribution function to diagnose systematic probability misallocation and translates this information into an outcome-level transport that jointly adjusts the ambiguity-set center and ground cost. The resulting formulation admits a computationally tractable dual reformulation. Theoretically, we derive valid ambiguity radii and decision-risk guarantees that tighten as estimation and approximation errors vanish, and show that DT-DRO can eliminate the nonvanishing robustness floor caused by model misspecification. Synthetic experiments and a power-outage application demonstrate improved decision quality, particularly under structural and tail misspecification.
cs.LG / 177 / 2610.10829
Conformal Prediction under Partial Verification
Zijun Yu, Yu Gu, Vahid Partovi Nia, Masoud Asgharian
stat.ML · cs.LG
Abstract
Conformal prediction provides prediction sets with finite-sample guarantees, but the label verification required for calibration can be expensive. We develop a partial verification method that returns exactly the same prediction sets as complete verification. We characterize calibration certificates, the verified information sufficient to determine the conformal threshold, and design a procedure that coordinates verification across calibration examples. For finite thresholds at high coverage, its verification cost is less than twice the minimum certificate cost when candidates are checked in order. Across retrieval, mathematical solutions, and configuration evaluation, it reduces verification cost by 15-82% compared with verifying calibration examples one at a time, while producing identical prediction sets.
cs.LG / 178 / 2610.10870
Transformed Samplers with Variance Reduction
Siran Liu, Michalis Tisias, Petros Dellaportas
stat.ML · cs.LG · stat.ME
Abstract
Markov chain Monte Carlo (MCMC) methods are the standard tool for computing expectations under complex probability distributions. Control variates reduce the variance of the resulting estimates, but a good control variate requires solving the Poisson equation of the sampler, which rarely admits a closed-form solution. Exact solutions are available when the sampler's kernel has a known spectral decomposition on a simple reference density. In our work, we extend these solutions to general targets through a learned change of variables. A bijection, such as a normalizing flow, is trained so that the target becomes close to the reference in a latent space, and we show that Markov kernels and their Poisson solutions are transformed by any bijection. Running such samplers in the latent space then yields explicit control variates, and the estimator is consistent under mild tail conditions on the map and target. Importance sampling (IS) from the flow is the limiting case of the same construction and the control variates apply to it as well. Experiments on synthetic targets and real posteriors compare the procedure against state-of-the-art samplers and control variates.
cs.LG / 179 / 2610.11082
A General $\widetildeΩ(\sqrt{T γ_T})$ Lower Bound for Kernel Bandits
Chenkai Ma, Jonathan Scarlett
stat.ML · cs.IT · cs.LG
Abstract
The kernel bandit problem consists of sequentially optimizing an unknown function with noisy feedback, where the function has bounded norm in a given Reproducing Kernel Hilbert Space (RKHS). A central quantity in the regret analysis of kernel bandits is the maximum information gain $γ_T$. In particular, the best existing upper bounds scale as $\sqrt{Tγ_T}$ up to log factors, and nearly-matching lower bounds have been derived for specific kernels such as squared exponential and Matérn. However, lower bounds for general kernels are lacking, thus making it unclear in what generality the upper bounds are near-optimal. In this paper, we establish a general $Ω(\sqrt{Tγ_T/\log T})$ minimax regret lower bound for non-constant continuous kernels on compact domains, establishing near-optimality (within log factors) in a very general sense. We show that the log factor appearing in this bound is unavoidable in general, but that it can be removed under certain conditions. Among other things, our findings imply that the minimax-optimal scaling is exactly $Θ(\sqrt{Tγ_T})$ (i.e., within constant factors) for the Matérn-$ν$ kernel with $ν\in (0,2)$, $γ$-exponential kernel with $γ\in (0,2)$, and certain piecewise-polynomial kernels.
cs.LG / 180 / 2610.11139
Accelerating Non-Smooth and Heavy-Tailed Sampling
Pervez Ali, Xiaoyu Wang, Yingli Wang, Lingjiong Zhu
stat.ML · cs.LG · math.PR
Abstract
Anchored Langevin dynamics (ALD) is useful for non-smooth sampling where the density of the target distribution is possibly non-differentiable and heavy-tailed; reflected anchored Langevin dynamics (RALD) can sample possibly non-differentiable target density on a constrained domain. In this paper, we propose and study non-reversible anchored Langevin dynamics (NALD) for sampling possibly non-differentiable and heavy-tailed target density in the Euclidean space and the non-reversible reflected anchored Langevin dynamics (NRALD) for sampling possibly non-differentiable target density in the constrained space. Our construction adds a circulation drift generated by a possibly state-dependent divergence-free skew-symmetric matrix field and a stream potential. It preserves the target distribution without requiring derivatives of target density, admits a random-time-change representation, and applies both on the whole Euclidean space and on bounded domains with normal reflection. By breaking reversibility, we show that NALD and NRALD can converge to their target distributions faster than their reversible counterparts via finite-time non-asymptotic convergence analysis, a large deviations analysis and asymptotic variance reduction. Numerical experiments demonstrate the efficiency of the proposed algorithms.
cs.LG / 181 / 2610.11486
PSI-SINDy: Post-Selection Inference for Sparse Identification of Nonlinear Dynamics
Ashraful Islam, Shuichi Nishino, Tomohiro Shiraishi, Ichiro Takeuchi
stat.ML · cs.LG
Abstract
Sparse identification of nonlinear dynamics (SINDy) is a data-driven framework for discovering governing dynamics from time-series data by identifying a sparse subset of candidate dynamical terms from a prespecified library. In this work, we develop a statistical inference framework for quantifying the reliability of dynamical terms selected by SINDy through hypothesis tests and confidence intervals. A key difficulty is that using the same noisy trajectory for both selecting dynamical terms and assessing their statistical significance can introduce selection bias. Post-selection inference provides a principled framework for addressing such bias, and we propose PSI-SINDy, a post-selection inference method tailored to SINDy. Direct application of existing post-selection inference techniques is challenging because SINDy involves measurement error in the candidate terms and shared noise between the response and design. To address these challenges, PSI-SINDy uses data thinning to decompose a single observed trajectory into four mutually independent views with distinct roles in selection and inference. This construction enables inference for selected dynamical terms while accounting not only for selection bias but also for measurement-error and shared noise effects. We establish the theoretical validity of PSI-SINDy under stated conditions and evaluate its performance through numerical experiments on simulated and experimental dynamical-system data.
cs.LG / 182 / 2610.11538
LAIR-Net: Leaky Alignment-Impulse Residual Networks for Tabular Regression
Rahul Goswami, Aryan Bhambu, Bittu Karmakar
stat.ML · cs.LG
Abstract
Deep randomized models fix hidden-layer parameters through random initialization and learn only closed-form readouts, typically adding depth by stacking random trans formations without target-aware control of hidden-state evolution. We propose LAIR Net, the Leaky Alignment-Impulse Residual Network, which mixes a shallow learned anchor into each hidden state through a leaky residual transition. We derive a depth uniform bound on input-perturbation sensitivity and use controlled simulations to attribute gains over a randomized baseline to the anchor rather than recursion or added capacity. Benefits emerge when a nonlinear target structure is learnable at the available noise level and diminish for nearly linear targets or dominant noise. Across 23 benchmark datasets, LAIR-Net achieves the best average rank among eight randomized networks and twelve conventional models, with relative performance associated with the same nonlinear-structure and noise quantities identified in simulation.
cs.LG / 183 / 2610.11584
Embedding-Bias in Conditional Independence Testing
Nikolaj Thams, Anton Rask Lundborg
stat.ML · cs.LG
Abstract
To test conditional independence of $X$ and $Y$ given a text or an image $Z$, one conditions on an embedding $ψ(Z)$ in place of $Z$. The embedded test is valid if $Z$ is independent of $X$ or of $Y$ given $ψ(Z)$, which cannot be confirmed from data, and when this fails, the rejection probability under the null hypothesis can tend to one. We study this failure, and show that focusing on a specific form of dependence relaxes what the embedding must retain. For a residual correlation test inspired by the Generalised Covariance Measure, validity only requires that the parts of $\mathbb{E}[X \mid Z]$ and $\mathbb{E}[Y \mid Z]$ missed by $\mathbb{E}[X \mid ψ(Z)]$ and $\mathbb{E}[Y \mid ψ(Z)]$ are uncorrelated. Otherwise, we treat the discarded information as an omitted variable. Under the null hypothesis, the bias equals the absolute correlation of the missed parts times the geometric mean of two partial $R^2$ values. This identity yields a robust test valid under a declared tolerance for the geometric mean, which, like a sensitivity parameter, is not identified from the data. On synthetic data and text embeddings, the robust test holds its level approximately. On text generated by a language model, under an exact null hypothesis, every embedding, even the generator's own states, biases the embedded test.
cs.LG / 184 / 2610.11628
Minimax Gaussian Mechanisms for Continual Machine Unlearning
Qi Kuang, Yin Xia
stat.ML · cs.LG
Abstract
Machine unlearning updates a trained model after records are deleted, aiming to match exact retraining without repeating the full training procedure. We develop Gaussian mechanisms for Newton updates under sequential deletion requests. Using Gaussian differential privacy (GDP) and its adaptive composition rule, we show that the full sequence of released models is statistically difficult to distinguish from matched exact retraining. To calibrate these mechanisms for empirical risk minimization, we derive upper bounds on the error of the Newton approximation relative to exact retraining and on how this error changes after each deletion batch. Independent Gaussian noise is calibrated using bounds on the full residual at each release, whereas Gaussian random walk noise uses smaller bounds on residual increments. These bounds yield allocations minimizing the worst-case maximum noise variance across releases under the resulting GDP certification constraints. With count-based bounds, the random walk asymptotically matches the worst-case variance of a single release at deletion cap $M$, while independent noise incurs an additional factor of order $M$. Set-based bounds can reduce the noise variances by using gradients and Hessians of the deleted records. For singleton deletion, we further show that count-based independent noise, count-based random walk noise, and set-based independent noise are minimax among fixed Gaussian covariances under their respective residual or increment bounds. With set-based bounds, allowing variances to adapt to deleted records can improve on every fixed covariance by a factor of order $(\log M)^2$ on some data sequences. The residual and noise bounds also yield parameter and predictive consistency relative to exact retraining, uniformly over deletion policies. Simulations and a credit default data analysis evaluate bounds, noise variances, and estimation errors.
cs.LG / 185 / 2610.11668
$σ$Transfer: Uncertainty Transfer from Small to Large Networks under $μ\mathrm{P}$
Richard Bergna, Fernando Ruiz Mazo, Nicolò Felicioni, José Miguel Hernández-Lobato, Kamil Ciosek
stat.ML · cs.LG
Abstract
Reliable predictive uncertainty in Laplace approximations depends critically on the prior precision, yet selecting it requires a posterior sweep that is prohibitively expensive for neural networks with billions of parameters. Under the Maximal Update Parametrization ($μ\mathrm{P}$), we derive a rescaling of the prior covariance that makes the selected precision stable as model width grows. This leads to $σ\mathrm{Transfer}$: we select the precision on a smaller model and zero-shot transfer it to the much larger model, i.e., without searching for the precision on the larger model at all. We show convergence of the prior kernel, posterior covariance, selected precision, and posterior-derived decisions under explicit conditions, and verify $σ\mathrm{Transfer}$ across regression, image classification, and Transformer readouts. For example, measured precision-sweep speedups reach $\sim 5000\times$ when transferring from width 128 to 4096 on MNIST, at a target-NLL degradation of $0.002$; transferring from a public 1B to 7B model gives a median search speedup of $\sim 2.3\times$ (up to $\sim 330\times$), with a mean measured target-NLL increase below $10^{-4}$ across ten tasks. The same posterior stability also enables transfer of acquisition, OOD-detection, and abstention decisions without constructing a target posterior.
cs.LG / 186 / 2610.11798
Softmax Attention on Gaussian Mixtures: Linear When It Can, Selective When It Must
Simon Gabet, Etienne Boursier, Claire Boyer
stat.ML · cs.LG
Abstract
Softmax attention, at the heart of Transformers, has demonstrated remarkable capabilities. Yet its underlying mechanisms remain only partially understood. Recent theoretical work studies Gaussian prompts, where the infinite-prompt limit reduces softmax attention to a linear map, but also removes the query-dependent selection that distinguishes it from linear attention. This work studies the infinite-prompt limit of softmax attention on Gaussian mixtures, which retain the tractability of Gaussian data while introducing latent structure, multimodality, and nonlinear dependencies. We show that softmax attention can represent and learn, via gradient-based methods, optimal solutions to a range of statistical tasks, including supervised classification and denoising. Our results highlight two complementary capabilities of softmax attention: it can recover linear tasks as effectively as its simpler linear counterpart, while also exploiting query-dependent context selection to solve nonlinear tasks beyond the reach of linear attention.
cs.LG / 187 / 2610.11863
Conditional Kernel Stein Discrepancy
Federico Matteucci, Florian Kalinke
stat.ML · cs.LG · math.ST
Abstract
Kernel Stein discrepancies (KSDs) provide a versatile tool for comparing distributions. One of their main applications is in quantifying the goodness-of-fit (GoF) between a data-generating distribution and a prescribed target distribution. In this work, we study the related problem of conditional GoF quantification: given only a (possibly non-normalized) conditional target model, without information on the distribution of its covariates, and samples from a joint distribution, the goal is to assess how well the conditional distribution of the samples matches the target. To tackle this setting, we present a framework that allows lifting unconditional KSDs to the conditional setting through an operator-valued kernel on the covariate space, going beyond the known Euclidean case. We establish that our suggested statistic vanishes if and only if the conditional model and the true conditional distribution agree for almost all covariates and deploy it to test conditional GoF on smooth manifolds and on discrete spaces. Our experiments on level, power, and runtime demonstrate the viability of testing on these domains using the proposed statistic.
cs.LG / 188 / 2610.11869
Learning structured linear dynamical systems from missing observations
Aravinda Kanchana Ruwanpathirana, Hemant Tyagi, Sunny G. W. Wang
stat.ML · cs.LG · eess.SY · math.OC · math.ST
Abstract
We consider the problem of learning structured linear dynamical systems over convex sets $\mathcal{K}$, where only a small subset of the observations are available at each time point. An estimator which minimizes a bias-corrected, potentially non-convex objective function is proposed. Non-asymptotic bounds are obtained for the statistical error, which depend on the local complexity of $\mathcal{K}$, the trajectory length $T$, and the sub-sampling probability $p$. Convergence of the projected gradient descent algorithm is also established. The general theory is applied to settings where (i) $\mathcal{K}$ is a subspace, (ii) $\mathcal{K}$ is the set of bi-isotonic matrices, and (iii) $\mathcal{K}$ is the set of matrices whose rows are formed by sampling Lipschitz functions. We show meaningful recovery of the transition matrix is possible for values of $T$ much smaller than what is required in the unconstrained case, and for $p = o(1)$.
cs.LG / 189 / 2610.11906
RobustLDS: Learning linear dynamical systems under adversarial corruptions
Aravinda Kanchana Ruwanpathirana, Hemant Tyagi
stat.ML · cs.LG · eess.SY · math.OC · math.ST
Abstract
We consider the problem of learning linear dynamical systems under adversarial contamination from a single trajectory of length $T$. While identification of linear dynamical systems itself is well-studied, the problem of robust system identification under adversarial contamination is relatively less explored. In this work, we study the setting where a fraction of the $T$ observations are contaminated by adversarial outliers. We propose different estimators based on relaxations of least-trimmed squares along with an alternating minimization algorithm. Furthermore, we also propose two estimators which exploit the group-sparsity (through penalization/hard-constraints) of the outliers. For the estimator with group-sparse penalty, we derive non-asymptotic error bounds which establish its robustness to outliers. We also show empirically that the proposed estimators work well in practice.
cs.LG / 190 / 2610.11947
Score-Based Learning of Cluster DAGs from Interventions
Gaetano Tedesco, Alex Markham
stat.ML · cs.LG
Abstract
Graphical approaches to causal abstraction transform a low-level causal directed acyclic graph (DAG) over many measured variables into a smaller, high-level DAG whose nodes cluster the original variables and whose edges summarize the causal relations between clusters. Such cluster DAGs are easier to interpret, but learning them requires finding the clusters and recovering the edges between them. Madaleno et al. (2026) learn the interventional coarsening (the cluster DAG that merges variables the interventions cannot distinguish) in two constraint-based phases: first the clusters, then the edges. We introduce COARSE, the first score-based method for this task: it keeps the two-phase structure but, under linear Gaussian assumptions, swaps the constraint-based edge phase for a score-based one. We show that the interventions themselves identify a causal order over the clusters, and learning the edges reduces to a single local search per cluster under a cluster-level BIC score. We prove that the procedure runs in polynomial time and, provided the variables affected by each intervention are correctly identified, that it is consistent. On synthetic and real-world interventional data, COARSE matches state-of-the-art edge recovery given enough samples, with an edge phase up to two orders of magnitude faster, including on dense graphs with hundreds of nodes.
cs.LG / 191 / 2610.11976
Efficient quadratic entropy with distance sketches
Steve Huntsman
stat.ML · cs.LG · math.ST · stat.CO
Abstract
We detail scalable methods for approximating the quadratic entropy $p^T d p$ for arbitrary distributions $p$ and common distances $d$ of negative type. We focus on the Euclidean and spherical geodesic cases, which both use random feature embeddings and projections to dramatically improve computational complexity within a simple framework. Amortization of a single large matrix multiplication and control variates further enable computation at large scale with low memory and runtime in situations where $d$ is held constant while $p$ varies. We demonstrate this with a comparison against direct pair sampling and bibliometric/scientometric examples on Open Graph Benchmark datasets, revealing papers, fields, and institutions with both particularly narrow and broad interdisciplinary reach from their citations and text features alone.
cs.LG / 192 / 2610.12035
Efficient and Generalizable Archetypal Analysis for Discrete Data
A. Emilie J. Wedenborg, Jesper Løve Hinrich, Morten Mørup
stat.ML · cs.LG
Abstract
Archetypal Analysis (AA) represents observations as convex combinations of extremal data-driven profiles, yielding interpretable low-dimensional descriptions of complex datasets. Classical AA relies on a least-squares objective, which is poorly suited to discrete observations such as binary, count, and categorical data. We introduce an efficient likelihood-based framework for AA supporting Bernoulli, Poisson, and multinomial observation models. Our optimization scheme employs local quadratic approximations of the negative log-likelihood, enabling constrained updates through sequential minimal optimization (SMO) and an active-set method. Scalability is improved by bounding the active set while preserving simplex feasibility. We further introduce a cross-validated predictive likelihood criterion for selecting the number of archetypes, providing a principled alternative to reconstruction-error heuristics and stability-based diagnostics. Synthetic experiments demonstrate computational efficiency and accurate recovery of model complexity. Applications to single-cell RNA sequencing, microbiome composition, and somatic mutation data show that the learned archetypes capture interpretable domain-specific structures while achieving competitive likelihood fits and stable solutions. Overall, the proposed framework enables efficient likelihood-based archetypal analysis of discrete data, complemented by predictive likelihood-based model selection.
cs.LG / 193 / 2610.12094
Differentiable Systematic Resampling for Variational Sequential Monte Carlo
Fredrik Cumlin, Saikat Chatterjee
stat.ML · cs.LG · eess.SP
Abstract
Particle filters are a standard tool for nonlinear state estimation, but their resampling step is discrete, preventing gradient-based learning in variational sequential Monte Carlo. We introduce Differentiable Systematic Resampling (DSR), a temperature-controlled relaxation of systematic resampling, that preserves the CDF-ordered, banded structure of systematic resampling while enabling full gradient flow. DSR converges to exact systematic resampling as the temperature vanishes, and we prove a pointwise exponential convergence rate for the induced bias. Compared to optimal-transport-based differentiable resampling, DSR avoids iterative solvers and has substantially lower computational overhead. Experiments on stochastic dynamical systems and real-world handwriting data show that DSR achieves comparable or superior filtering and dynamics learning performance.
cs.LG / 194 / 2610.12213
ISBO: Scalable Spatio-Temporal Bayesian Optimization with Log Gaussian Cox Process Models via the INLA-SPDE Approach
Kaichuang Yang, Håvard Rue, Jakob Zeitler
stat.ML · cs.LG
Abstract
Bayesian Optimization (BO) is a popular method for efficiently optimizing expensive black-box objectives. However, BO utilizing standard Gaussian Processes is ill-suited for doubly stochastic Cox Processes that are often used in spatio-temporal problem spaces. We introduce INLA-SPDE Spatio-Temporal Bayesian Optimization (ISBO): the first scalable BO framework for spatio-temporal data, that models the log-intensity with a Log-Gaussian Cox Process(LGCP) and performs inference via Integrated Nested Laplace Approximation and Stochastic Partial Differential Equations (INLA-SPDE) approach. Using a Matern field on meshes yields a sparse Gaussian Markov Random Field, where INLA provides fast and accurate posterior inference throughout sequential optimization. ISBO stably locates high-intensity regions and the peak of the latent intensity with minimal evaluations. A time-varying Upper Confidence Bound acquisition with masking avoids revisits, while penalized-complexity priors regularize early rounds. Experiments on synthetic and real-world spatio-temporal datasets show accurate peak discovery, intensity recovery, and substantial speedups over an RKHS-based baseline, positioning ISBO as a practical choice for BO with point-process data.
cs.LG / 195 / 2610.12437
Density Ratio Estimation with Stein Displacement Fields
Song Liu
stat.ML · cs.LG
Abstract
Density ratios quantify distribution shift from a probability-mass point of view, whereas displacement fields describe, from a dynamical point of view, how one distribution is transported onto another. Although both offer complementary insights, they are usually estimated separately, and converting one into the other requires post-processing. In this paper, we estimate the density ratio between a target and a base distribution by parametrizing it through a displacement field acting on the base: the log-ratio is modeled as minus the Stein operator of the base applied to the field, up to a normalizing constant. This gives both statistical and dynamical descriptions of the distribution shift through a single convex optimization problem. Iterating this estimate-and-move step gives two inference algorithms: push-forward moves the model and corrects a pretrained sampler without retraining it, whereas pull-back moves the data closer to the base and fits a transformation model one layer at a time. Applications to distribution shift in simulation-based inference and to nonlinear independent component analysis illustrate the benefits and limitations of the approach.
神经与进化计算 (cs.NE)
2
cs.NE / 1 / 2610.12251
Universal Construction and Exact Self-Reproduction in Ternary McCulloch-Pitts Networks
Charles C. Norton
cs.LO · cs.NE
Abstract
A fixed network of McCulloch-Pitts threshold units with weights in {-1,0,1} can hold other threshold networks in its state and run them: the state is a ring of banks of records, each record a unit of a stored network, and each step evaluates one record. We use such a network to carry out von Neumann's universal construction and self-reproduction exactly. As a cellular automaton keeps its rule, the fixed network keeps its weights, and what reproduces is a stored network. A constructor of 143 records reads the description of a network from its tape, builds that network in the next bank, copies the description onto the next tape and hands control to what it built; started on its own description, it rebuilds itself, weights included, in every generation. The scheme scales to a universal computer. A SUBLEQ computer of 17,598 ternary units, stored as 36,080 records and running a program of 27 instructions, builds any network that fits a bank and reproduces itself in the same way, at every word width from eight bits on; run directly, it emits the serialization of its own weights, memory and tape. Integer pre-activations give every orbit a margin of 1/2, and replicating each unit r times multiplies it by r. That suffices against noise of any size on the pre-activations, but against von Neumann's output flips only below a threshold inversely proportional to the fan-in. These results are proved in Rocq.
cs.NE / 2 / 2610.11546
Learning to Orchestrate Evolutionary Search: Progression-Aware Deep Reinforcement Learning for Dynamic DE-CMA-ES Coordination in Optimization and Structural Model Updating
Lechen Li, Rongye Shi, Wanhuan Zhou
cs.NE · cs.AI
Abstract
Solving high-dimensional structural model updating problems requires an algorithm capable of navigating complex, non-convex landscapes with correlated parameters. Existing hybrid evolutionary algorithms typically rely on static architectures or fixed switching rules, resulting in disjointed search phases. To address this, this study proposes a Deep Reinforcement Learning-governed dynamic DE-CMAES Orchestration (DRL-DCO) algorithm, in which a Deep Deterministic Policy Gradient (DDPG)-based actor-critic agent continuously governs the evolutionary process as a single, unified system rather than a mechanical concatenation of algorithms. Guided by a progression-aware state representation and a diversity-informed reward, the agent fluidly reallocates computational resources between the differencevector-based exploration of Differential Evolution (DE) and the covariance-guided exploitation of CMA-ES, while jointly regulating population size, elite preservation, and a restart mechanism to escape local optima. This allows DRL-DCO to autonomously transition between exploration-dominant, exploitation-dominant, and mixed-strategy regimes across generations. Beyond the training phase, the trained actor can operate in a supervision-free inference mode, where the internalized policy autonomously orchestrates DE and CMA-ES control from observed search states through forward inference alone, without critic evaluation or weight updates, enabling faster deployment while retaining full effectiveness. Validated on high-dimensional single-objective optimization benchmarks and the IASC-ASCE structural health monitoring benchmark, DRL-DCO achieves superior convergence accuracy and robustness compared to state-of-the-art adaptive and hybrid evolutionary algorithms, as well as single-operator DRL-governed baselines.
计算语言学 (cs.CL)
62
cs.CL / 1 / 2610.10724
Cognitive Thermometers: Machine Learning and Logical Complexity
Shane Steinert-Threlkeld, Jakub Szymanik
cs.CL
Abstract
How does the human mind represent semantic categories? Why do natural languages favor certain meanings over others? Prior explanations have relied on logical definability and complexity, but these are highly sensitive to the choice of logical language, rendering some design choices unmotivated. In this article, we propose that machine learning provides a somewhat more agnostic approach to measuring semantic complexity. We review emerging evidence that logic and machine learning often yield converging results on relative complexity and its resulting effects in semantic typology. Where they diverge, learning appears to be a better explanation than logical complexity. We argue that treating machine learning models as ``cognitive thermometers'' enables a unified approach to complexity that bridges symbolic logic and connectionist AI.
cs.CL / 2 / 2610.10738
Lossy Compressive Text Autoencoders
Vinko Sabolčec, Angelos Katharopoulos, David Grangier
cs.CL
Abstract
Our work explores learning a compressed latent representation of text, at the intersection of data compression and representation learning. We propose an autoencoder architecture that performs residual downscaling and upscaling of hidden representations along the time axis, with a residual low-dimension discrete bottleneck. We analyze our approach for different quantization methods, training objectives, and datasets. For different levels of compression, we evaluate the similarity between the original and reconstructed text both at the surface-level (BLEU) and at the semantic-level (LLM-based judge). Additionally, we evaluate our models on downstream question-answering and semantic text similarity benchmarks. Our approach results in compressed representations which are on par with lossless text compression algorithms at 2.24 bits per byte on web text data, while having good reconstruction and downstream task performance.
cs.CL / 3 / 2610.10758
Clarify, Then Focus: Statement Normalization for Conversation Analytics at Scale
Mikhail L. Arbuzov, Karan Dave, Evgeniya Dontsova, Yaodong Hu, Vincent Lao, Navita Jain, Sisong Bei, Dmitry Dimov
cs.CL · cs.AI · cs.LG
Abstract
Enterprise conversation analytics asks many questions of millions of interactions. Each question can require reconstructing what people mean and identifying which information matters, repeating costly interpretive work across the same transcripts. We propose a simple principle: clarify the text, then focus the reader. Statement normalization transforms dialogue into short, speaker-attributed statements with source references and semantic tags. The statements make meaning more explicit; the tags support selecting evidence for a particular question. Downstream models can use the full representation or a relevant subset, depending on what helps them make the decision. In an offer-suppression task on customer-service calls, normalization improves a supervised classifier without selection, while weaker prompted readers benefit from both normalization and selection. A small model can learn the normalization contract, while lightweight encoders handle tagging and downstream decisions. Sharing this preparation across questions supports an inference pipeline built entirely from small models, making analytics over millions of conversations substantially less expensive.
cs.CL / 4 / 2610.10827
Grammar Concept Annotation at Scale: Deployed Fine-Tuned Small Language Models Outperform Prompted Frontier Models
Marjan Celikik, Ana Peleteiro Ramallo, Javier Morales
cs.CL · cs.AI
Abstract
Corrective feedback is among the best-evidenced drivers of second-language acquisition, yet corrections delivered during lessons rarely accumulate into an actionable view of grammar mastery. Prompted frontier models can provide such a view from learner--tutor lesson transcripts, but they are costly at scale. We close this gap by fine-tuning Qwen3.5 small language models (SLMs) on filtered and rebalanced teacher-generated supervision, then deploying an efficient 0.8B model in an end-to-end grammar mastery tracker for all English learners on our platform. Internalizing the annotation contract into adapter weights enables pairing the 0.8B model with a compact matched prompt rather than verbose instructions. On two human-curated benchmarks, both the deployed 0.8B model and a 4B reference comparator outperform prompted GPT-5.4 and GPT-5.6 Sol in precision and recall under nested matching criteria of increasing strictness: concept, evidence span, and correctness. The deployed 0.8B SLM reduces serving cost by approximately 16$\times$. A feature-level online experiment shows significant gains in learner engagement ($+15.8\%$) and key business metrics, including scheduled hours ($+2.1\%$) and GMV from new lessons ($+13.2\%$).
cs.CL / 5 / 2610.10865
Disentangling Linguistic and Paralinguistic Information with Routed Sparse Autoencoders
Beimnet Bekele Guta, Xiaoyu Yang, Guangzhi Sun, Philip C. Woodland
cs.CL · eess.AS
Abstract
Self-supervised speech encoders contain linguistic and paralinguistic information in a shared, entangled representation space. We combine a TopK sparse autoencoder with route-specific supervision and cross-factor adversaries. Across frozen SPEAR and WavLM encoders, independent probes show factor-specific retention and suppression: linguistic information remains stronger in the linguistic route, while paralinguistic factors, including speaker identity, emotion, and prosody, are retained in the paralinguistic route and substantially reduced in the linguistic route. The route organisation learned on LibriSpeech persists on MSP-Podcast without representation-side retraining. Feature-space route interventions further transfer the swapped factor while largely preserving the information carried by the unchanged route. These results show consistent route-selective separation across encoders, corpora, independent probes, and representation-level interventions.
cs.CL / 6 / 2610.10878
Stochastic Teacher Intervention for Agentic On-Policy Distillation
Junnan Liu, Linhao Luo, Zhijun Chen, Qianren Mao, Thuy-Trang Vu, Gholamreza Haffari
cs.CL · cs.AI
Abstract
On-policy distillation (OPD) efficiently transfers capabilities from a stronger teacher to a student language model through dense token-level supervision on student-generated rollouts and has shown promise on complex tasks such as mathematical reasoning. However, in multi-turn agentic tasks, student decisions shape subsequent observations, causing early errors to accumulate across turns. The resulting trajectories can drift away from the teacher's rollout distribution, making the teacher's token-level supervision less reliable or even counterproductive for OPD training. To address this issue, we introduce STI-OPD, a stochastic teacher intervention framework for multi-turn agentic OPD. During multi-turn interaction, STI-OPD uses teacher intervention guided by teacher-student policy discrepancy to replace the student's proposed action with a teacher-generated one to maximize the acquisition of reliable supervision. We further develop a stochastic intervention strategy, addressing the limitations of previous threshold-based or fixed-schedule approaches, that estimates policy discrepancy using KL divergence and maps it to an intervention probability. By sampling whether to intervene from this probability, STI-OPD adaptively balances teacher control with student exploration. To learn from the resulting mixed-policy trajectories, we introduce an Importance-Weighted Reverse KL objective that corrects the token sampling mismatch between teacher-generated responses and the student policy to preserve the original OPD objective. Across tool-integrated reasoning and long-horizon interaction, STI-OPD outperforms the strongest prior OPD baseline on every evaluated benchmark and student size. Ablations further show that both discrepancy-guided intervention and importance weighting contribute to these gains.
cs.CL / 7 / 2610.10988
Back in Style: A Sociolinguistic Approach to Authoring and Measuring Persona Fidelity in User Simulation
Lex Konnelly, Elena Khasanova, Riqiang Wang, Matthias Lee, Harsh Saini, Parsa Kavehzadeh
cs.CL
Abstract
As agentic systems gain commercial popularity, user simulators increasingly serve as measurement instrument for their evaluation. However, the fidelity of simulated users in comparison to real human users is generally low, and typically assessed by costly, subjective LLM judges. In this pilot study, we ask whether fidelity can instead be measured deterministically by treating a user persona sociolinguistically: as a social type that emerges from observable linguistic style, rather than one predicted by labels or descriptions a model must extrapolate into behaviour. We author personas as concrete stylistic rates, which lets us transfer two established, model-free instruments -- authorship-verification stylometry and lexicon-based content analysis -- as fidelity diagnostics. We A/B-test the sociolinguistic schema against a flat descriptive baseline across five task-oriented customer-service agents. Results show that the sociolinguistic schema improves both stylistic adherence and stylometric distinguishability for most of the tested models, with a caveat that persona style fidelity does not necessarily equal persona "naturalness". We argue that a sociolinguistic approach to persona design is a promising path towards more diverse and representative user personas, and that these metrics are most valuable in an error-attribution analysis, localizing where fidelity breaks down. This is a first step towards interventions that move user simulations closer to faithful renderings of diverse and variable linguistic outputs.
cs.CL / 8 / 2610.11015
Prompts versus Rules: Auditing and Controlling Speech Naturalness Behaviors in Voice User Simulators
Riqiang Wang, Elena Khasanova, Harsh Saini, Lex Konnelly, Parsa Kavehzadeh, Matthias Lee, Mohamed Attia
cs.CL
Abstract
As voice agents gain more popularity commercially, the user simulators used to evaluate the deployed agents are also being developed to include more realistic, variable, and diverse speech naturalness behaviors -- disfluency, interruption and backchanneling. The quality of the user simulator directly affects the validity of agent evaluation results. However, we find that most studies so far have not examined in detail whether the intended configuration for these behaviors is realized in the simulation. In this study, we audit the realized naturalness behaviors of tau-Voice, our own LLM-based prompting approach across three models, and our rule-based injection algorithm for disfluency, interruption, and backchanneling. We find that prompting for these behaviors is unreliable and produces speech inconsistent with the instructions, placed and distributed less naturally than the instruction implies. In contrast, our rule-based, model-free algorithm produces controllable and diverse naturalness behaviors more aligned with natural speech. Our results suggest that LLMs not purpose-trained for user simulation are not sufficient on their own to represent authentic user behavior, and that linguistically informed deterministic approaches or specialized models are needed to close the gap; auditing and reporting realized naturalness behaviors, rather than configured settings, is what makes that gap visible.
cs.CL / 9 / 2610.11111
Lapras: Latent Reasoning for Time Series Language Models
Yuliang Chen, Yu Yvonne Wu, Patrick Langer, Arvind Pillai, Sudarshan Regmi, Martin Maritsch, Juncheng Liu, Robert Jakob, Thomas Kaar, Tess Z. Griffin, Lisa Marsch, Michael V. Heinz, Nicholas C. Jacobson, Andrew Campbell
cs.CL · cs.LG
Abstract
Time Series Language Models (TSLMs) offer a promising path toward time series understanding by reasoning over temporal signals and producing natural language answers and explanations. A common approach is Chain-of-Thought (CoT), which generates step-by-step rationales linking relevant signal patterns to final answers. Although these models learn from reference CoT traces during post-training, generating faithful descriptions of input time series at inference remains challenging. Expressing high-dimensional, continuous temporal representations in discrete language tokens may cause the model to neglect task-relevant patterns or describe them inaccurately. Because later reasoning steps build on these descriptions, early errors propagate, leading to incorrect answers with plausible explanations that are inconsistent with the input signal. We propose Lapras (Latent Post-trained Reasoning Across Series), a post-training framework that equips TSLMs with latent reasoning. A model trained with Lapras reasons through a sequence of continuous thoughts in the joint time series-language space, producing text only for the final answer. It learns this through teacher-student self-distillation, where a teacher trained on CoT reference traces reasons explicitly through text. The student aligns its hidden states with the teacher's at the answer stage, transferring the teacher's reasoning ability into its latent computation. We evaluate Lapras across four TSLM backbones on five time series question answering benchmarks. Lapras improves average F1 by up to 10.79% over explicit CoT while generating 23.9x fewer tokens. Lapras's continuous thoughts can also be decoded into readable reasoning traces via standard language decoding, preserving textual explanations. Together, these results highlight Lapras as a promising post-training paradigm for efficient, effective, and interpretable TSLM reasoning.
cs.CL / 10 / 2610.11135
Can a System-One LLM Perform Knowledge Tracing When Few or No Learners Are Logged?
Unggi Lee, Haeun Park
cs.CL · cs.CY · cs.LG
Abstract
Knowledge tracing (KT) models need many logged learners, so a new course or platform starts without a usable model. In LLM-based KT the LLM generates the answer, which we call System-Two; it is either fine-tuned on the target data or reasons and votes over ten samples, which is slow and gives coarse probabilities. We ask whether an off-the-shelf System-One LLM, which returns a probability for a typed question directly in a single pass, can perform KT when few or no learners are logged. On seven datasets, Jev without any data from the target platform reaches a mean AUC of .706, above the best of 28 deep KT models trained on 8 learners (.689) and above System-Two Thinking-KT on all seven datasets (.650) at about 1/100 of its API cost. Adding examples and a similar-learner statistic from the logged learners (JevKT) raises this to .722; JevKT stays significantly ahead of deep KT up to 16 learners and ahead on average up to 64, and supervised KT catches up between 64 and 128 learners. Among the readers we tested, the gain is specific to Jev, since three other LLMs queried with the byte-identical typed request through the official System-One adapter fall below it on all seven datasets, and reader swaps and contamination checks find no evidence that the input format or memorised data explain the gain. For new learners the advantage holds from their first interactions, whereas on unseen items with all learners logged, deep KT remains ahead.
cs.CL / 11 / 2610.11136
The "10th Juror": Open-Set Standpoint Screening for Bureaucratic Bias Detection
Yuchen Miao, Zijun Wang, Chang Han, Yurui Shi, Mingtai Zhang, Siyang Xu
cs.CL
Abstract
Presupposing the boundaries of bias is itself a form of bias. We study closed-loop bias governance for Dutch government documents, where a system must detect biased language, ground decisions in legal and contextual evidence, rewrite problematic sentences when intervention is warranted, and verify that the rewrite mitigates harm without distorting meaning. Existing methods face three challenges: (i) discriminative classifiers capture surface regularities but lack normative grounding; (ii) zero-shot LLMs often adopt generic viewpoints and over-flag ambiguous administrative language; and (iii) fixed taxonomies inherit the Closed-World Assumption, missing emerging local targets. We propose MARS-Gov, a standpoint-aware multi-agent framework that combines legal retrieval, open-set target screening, specialized jurors, conservative routing, and rewrite verification. When screening finds an uncovered group, MARS-Gov instantiates a dynamic "10th juror" to deliberate outside the fixed panel. On DGDB, MARS-Gov sets a new SOTA with 0.880 F1, outperforming the strongest zero-shot LLM detector by 20.2 points (29.8% relative) and the best supervised Dutch encoder by 6.8 points, while reducing unnecessary interventions to 2.5%. Leave-One-Category-Out (LOCO) evaluation recovers held-out categories with 85.1% Correct@1 and 93.8% Correct@3.
cs.CL / 12 / 2610.11160
LadderEdit: Edit-Level Residual Compression for Memory-Efficient Lifelong Editing of LLMs
Xiaobing Yu, Peijie Qiu, Jin Yang, Xuanzhao Dong, Weiwei Ma, Zhaoqi An, Xiaoqi Zhao, Xiaofeng Liu
cs.CL · cs.CV · cs.LG · cs.MM
Abstract
Lifelong editing of LLMs requires storing thousands of edits after acquisition. A widely used family of approaches attaches one LoRA adapter per edit, which preserves behavior but grows linearly in storage. To address this challenge, we propose LadderEdit, a method that compresses each LoRA adapter after it is acquired. Each edit is first stored at low rank as a cheap sketch. We then check whether this sketch still satisfies the rewrite, generalization, and locality contract on probe prompts. Edits that pass keep the sketch; those that fail are promoted to a higher rank along a ladder until the contract is met. Because every edit retains some representation, coverage is maintained, and only hard edits consume more rank. Across ZsRE, CounterFact, and WikiBigEdit benchmarks on LLaMA-3-8B, Mistral-7B, and Qwen2.5-7B, LadderEdit tracks exact LoRA storage at 5.2x less memory and remains effective at 50,000 sequential edits.
cs.CL / 13 / 2610.11183
RAG-Stress: Probing the Limits of Evidence Reliance in Retrieval-Augmented Generation
Shunyuan Zhou, Hao Chen, Tianyu Wang, Goose Lin, Zaiyuan Wang, Haiying Zhao
cs.CL · cs.AI
Abstract
Following retrieved evidence does not guarantee factual correctness: misleading evidence can induce a model to replace an answer it previously gave correctly. Standard accuracy measures obscure this behavior by combining answer replacement with preexisting errors. We introduce RAG-Stress, a controlled diagnostic protocol for examining the limits of evidence reliance in retrieval-augmented generation. The protocol holds the question and reference answer fixed, edits one assertion to support a designated incorrect answer, and crosses two source priority policies with three positions of the answer span within the evidence text. We measure misleading rate (MR) on each model's subset of questions answered correctly without retrieval, alongside clean accuracy on the full evaluation set. We evaluate fifteen systems spanning API models, open models, and search agents trained with reinforcement learning on TriviaQA-RC, HotpotQA, and SearchQA, with additional English and Chinese MedQA evaluations. Instructions that prioritize documents consistently produce higher MR than those permitting reliance on prior knowledge. Averaged over models and positions, the gap ranges from 10.9 to 13.5 percentage points across the three QA datasets. Mean MR follows End $>$ Beginning $>$ Middle under both policies, although individual models do not uniformly follow this ordering. A separate paired audit of 500 questions and two checkpoints supports increased harmful override without establishing a corresponding improvement in beneficial correction. These findings distinguish evidence adherence from factual reliability and motivate evaluating whether retrieved evidence preserves, replaces, or corrects a model's answers.
cs.CL / 14 / 2610.11270
Gated Memory: Admission-Controlled Memory Formation for Conversational AI
Preeti Saraswat, Divya Neelagiri, Ajay Manoj
cs.CL · cs.AI · cs.IR · cs.LG
Abstract
Personalized conversational AI relies on long-term memory systems that extract facts from user utterances and store them in persistent vector stores. Despite progress in retrieval, deduplication, and lifecycle management, the formation stage, the moment a fact is first written to storage has received almost no principled attention. We identify this as the binding constraint on memory quality in production systems. Critical contextual signals, such as the distinction between a permanent user attribute and a transient situation, exist only in the original utterance and are irreversibly lost the moment extraction produces a subject-relation-object triple. No downstream process can recover them. We propose Gated Memory, a lightweight, modular formation framework that interposes two decision checkpoints between conversation and storage: an admission gate that evaluates every candidate fact against the full utterance context before extraction runs, and a conditional enrichment stage that grounds admitted facts through an entity scope taxonomy with privacy constraints. The gate evaluates only the current exchange while using prior turns as read-only reference context, and produces a structured formation record. Admitted content is decomposed into atomic facts, each categorized, tagged with provenance (directly stated versus inferred), scoped to its condition of applicability, and grounded in resolved time and place, subject to a constraint that no entity absent from the context may be asserted. On the LoCoMo-10 benchmark with atypical emotional density in utterance data, Gated Memory achieves an overall +2.6% relative improvement in LLM-judge accuracy over a strong baseline with identical retrieval and generation, establishing formation quality as a measurable constraint on memory performance.
cs.CL / 15 / 2610.11275
Phonological Interference in Multilingual Speech Models
Moran Yanuka, Raja Giryes, Moris Alper
cs.CL · cs.LG
Abstract
Phoneme-level models transcribe or generate speech as a sequence of phonemes, the smallest sound units that distinguish words. These models enable fine-grained pronunciation control and understanding, yet often fail on input that does not match any single training language, such as speech alternating between two languages, known as code-switching, or low-resource languages absent from training. We identify a systematic failure mode behind this, phonological interference: models assume the input is in a single language and impose its phonology, overriding local phoneme-level decisions that conflict with the assumed language. We measure interference by how often a model retains phonemes that one language has but the other lacks. On code-switched input, two phone recognizers (speech-to-phoneme models) and a phoneme-conditioned text-to-speech model lose 32% to 79% of these phonemes, but lose far fewer of the phonemes both languages share. On unseen languages, we find that phone recognizers impose the phonology of the training language they assign to the speech, and the more confident the assignment, the more they lose phonemes the unseen language has but the assigned language lacks. We probe the models' language estimate from their internal activations, and trace interference to a low dimensional subspace. On monolingual speech, steering this subspace toward another language makes the model lose the phonemes that only the original language uses and produce phonemes that only the target language has. We introduce windowed language estimation (WLE), an inference time repair that replaces the model's language estimate in this subspace with one computed from a short window around each position. On code-switched input, WLE removes 34% to 69% of the interference in all three models, and in the recognizers it leaves monolingual performance essentially unchanged.
cs.CL / 16 / 2610.11287
REMORY: Learning Residual Memory for Context Compaction
Hanchen Xia, Baoyou Chen, Yutang Ge, Naihao Deng, Senqiao Yang, Zilong Dong, Weihao Yuan, Siyu Zhu
cs.CL · cs.AI
Abstract
Long-horizon agents compact their history to continue within a finite context window, but a textual summary alone may not support every subsequent decision. We introduce REMORY, a neural memory network that supplements the summary with a bounded sequence of soft memory tokens. Given the history and summary, the network learns to generate tokens that help a frozen LLM approximate the continuation it would produce with the full history. The tokens are conditioned on the summary and appended after it, forming an analogue of a residual connection along the sequence dimension. On SummHay, REMORY improves source attribution at nearly unchanged insight coverage and approaches the full-context joint score using only 5.2% of the input positions. Across long-horizon agent benchmarks, Qwen3.8-27B and GLM-5.3-Flash show consistent gains with residual memory. Both models also exhibit substantially fewer repeated tool outputs and tool errors on BrowseComp and Terminal-Bench 2.1.
cs.CL / 17 / 2610.11291
When Do We Need On-Policy Distillation? Distilling on Offline Student Rollouts Is Often Better
Siyan Zhao, Yonggan Fu, Jindong Jiang, Shih-Yang Liu, Song Bian, Byung-Kwan Lee, Sharath Turuvekere Sreenivas, Wenliang Dai, Hanrong Ye, Aditya Grover, Pavlo Molchanov
cs.CL
Abstract
On-policy distillation (OPD) has become increasingly popular for transferring teacher capabilities to student models. In this work, we ask a critical research question: Is on-policy sampling always beneficial for distilling arbitrary teacher-student pairs? We show that a simple alternative, Semi-OPD, which distills from offline rollouts generated by the initial student, can often outperform OPD in both accuracy and training efficiency. Across 17 teacher-student pairs ranging from 1.5B to 235B parameters, Semi-OPD outperforms OPD in 14 cases, with up to +13.6% accuracy and 11.4x training speedup. We further find that the choice between OPD and Semi-OPD depends on the alignment between the initial teacher and student, quantified by an output-token overlap ratio: OPD is beneficial only when the two are highly aligned with high overlap ratios. Our deeper investigation suggests that effective distillation requires on-policyness w.r.t. both the student and the teacher. For misaligned pairs, student rollouts can become increasingly off-policy w.r.t. the teacher as context length grows, weakening the distillation signal. In contrast, Semi-OPD is often more stable, as it distills on shorter contexts while covering full trajectories and exposing the student to more teacher-preferred tokens. Beyond proposing Semi-OPD as an efficient alternative, our work motivates the community to rethink when to use OPD and to study stronger OPD variants with meaningful teacher-student pairs.
cs.CL / 18 / 2610.11316
MetaEncoder: Exploring the Limit of Bi-Encoders for Multimodal System One Decision Making with Natural Language Interface
Jianpeng Cheng, Guangyu Sun, Aashu Singh, Benyu Zhang, Haixing Dai, Hossein Mansour, Jiangfan Zhang, Shlok Kumar Mishra, Wei Sun, Xuanming Cui, Yanli Liu, Qi Guo, Max Xiangjun Fan, Jun Xiao
cs.CL · cs.AI
Abstract
System One models output constrained decisions and probability distributions rather than free-form text generation. While prevailing paradigms rely on structured schema objects to encode state, intent, and candidate choices, we revisit a fully natural language-based System One interface. In this framework, both the user request and each candidate option are expressed in natural language, supported by multimodal (image and video) auxiliary inputs. We introduce MetaEncoder, which fine-tunes a pre-trained Muse-Glimmer 30B decoder into an instruction-following decision-making encoder. To scale effectively across both small closed-set (< 256) and massive open-set (millions) candidate spaces, MetaEncoder employs a bi-encoder architecture trained via unidirectional contrastive learning for request-candidate alignment. We conduct extensive evaluations across 11 benchmark suites and 190 tasks spanning multimodal decision-making, understanding (closed-set) and retrieval (open-set), highlighting where MetaEncoder beats SOTA multimodal encoders, as well as its current limits on reasoning-intensive tasks.
cs.CL / 19 / 2610.11332
ReCal: Calibrating Structured Pruning for On-Policy Distillation Recovery
Houcheng Jiang, Mao Zheng, Mingyang Song, Qiyong Zhong, Jie Sun, Tianyu Zhang, Junfeng Fang
cs.CL
Abstract
Structured pruning reduces the deployment cost of reasoning language models, but the resulting capability degradation can hinder subsequent on-policy distillation (OPD) recovery. Because OPD relies on student-generated trajectories, pruning damage that persists after offline distillation can limit its effectiveness. We propose RECAL, Recovery-Aware Calibration, a simple plug-and-play approach that improves OPD recovery by adjusting calibration before pruning. RECAL uses forward KL between an unpruned teacher and a pruned probe to identify teacher-supported predictions disrupted by pruning, then reweights calibration statistics to guide existing pruning criteria toward preserving these predictions. Across multiple models and pruning methods, RECAL consistently improves mathematical reasoning after OPD, achieving gains of up to 16.7 percentage points on AIME, alongside improvements in most code-generation comparisons. Further analysis shows that RECAL reduces residual damage at heavily affected tokens and establishes performance advantages that persist through recovery. These results demonstrate the value of recovery-aware calibration for improving on-policy distillation recovery of pruned reasoning models.
cs.CL / 20 / 2610.11354
AdaptEvo: Adaptive Agent Learning with Evolving Supervision
Shijun Wan, Jiancong Xie, Hang Xu, Jin Duan, Qixiong Wang, Xi Xiang, Maofei Que, Yahui Liu, Zhongyu Wei, Mu Chuan
cs.CL
Abstract
Rule-governed contextual decision tasks require models to apply specified rules to case-specific context and evidence. Written rules can leave gaps in decision guidance and process evaluation, while reference judgments vary in their support from the rules and evidence. To address these challenges, we introduce AdaptEvo, a framework for learning under imperfect supervision that couples confidence-adaptive policy optimization with evolving decision knowledge and evaluation rubrics. Its Training module uses Confidence-Adaptive GRPO (CA-GRPO) to balance outcome and process rewards according to reference confidence. Its Evolution module synthesizes reusable decision knowledge from recurring failures across training cases and refines process rubrics to detect overlooked errors. To support empirical evaluation, we construct an industrial multimodal content moderation dataset comprising a training set and In-Period and Out-of-Period test sets, with the latter collected under changed rules. Using Qwen3.6-35B-A3B, AdaptEvo achieves 61.9% exact-label accuracy and 72.2% binary decision accuracy on In-Period, exceeding GRPO by 7.5 and 3.7 percentage points, respectively. On Out-of-Period, the policy trained with CA-GRPO retains exact-label accuracy gains over the base model across evaluated checkpoints without injected decision knowledge, while GRPO declines with continued training. CA-GRPO also outperforms the tested fixed reward mixtures on both Out-of-Period metrics.
cs.CL / 21 / 2610.11389
Fact over Fiction: Detection of Pathological Hallucinations in Sinhala-to-English Neural Machine Translation
Navam Obeysekara, Nevidu Jayatilleke
cs.CL
Abstract
Neural Machine Translation (NMT) models, while capable of producing highly fluent outputs, remain vulnerable to hallucinations, which are translations that are natural yet semantically unrelated to the source. This vulnerability is acute in low-resource settings like Sinhala-to-English, where weak cross-lingual alignment leads to hallucinations. This paper introduces a framework for reference-free hallucination detection in this language pair. We present a 45,000-sample synthetic dataset generated through a probabilistic chain of five linguistically motivated corruption strategies, with a semantic rescue mechanism that uses character-level similarity to distinguish hallucinations from morphological variants. We fine-tune mDeBERTa-v3 for token-level sequence labelling, reaching a token-level F1 of 0.841 +/- 0.001 over three seeds on a source-disjoint test set, and study a three-signal ensemble integrating neural risk scores, sequence log-probabilities, and cross-lingual semantic embeddings (LaBSE). A source-ablation control shows that the detector relies on the Sinhala source rather than on surface artefacts of the corruption process: shuffling or removing the source reduces sentence-level AUROC from 0.970 to chance. We benchmark eight NMT systems spanning five model families and find that detector firings vary by an order of magnitude across architectures.
cs.CL / 22 / 2610.11436
Adversarial Cues in Decision Models Used as Judges: The Role of Request Presentation
Hongliang Liu
cs.CL · cs.CR
Abstract
An answer judge instructed to grade the final commitment should reject an explicitly wrong final value even when an earlier value matches the reference. We show that adding one colon to a candidate can violate this requirement depending on the presentation of the structured judging request. Numeric references certify the error, and paired interventions distinguish the candidate edit from the integration's presentation choices. On 200 previously unused DROP and GSM8K source clusters, the edit increased Jev's false acceptance from 1.0% to 26.0% with three output labels and from 3.0% to 26.5% with the published four-label grading instruction under sorted JSON keys. Both candidate variants were rejected under insertion presentation. These interactions passed the prespecified statistical correction even though Jev met the control thresholds under both grading configurations and presentations. Most excess acceptances occurred among candidates assigned larger numerical errors. GPT-6 Sol produced no observed cue-condition false acceptances, with missing responses unresolved. The result shows that basic judging competence can coexist with sharply different vulnerability to a fixed candidate edit across logically equivalent request presentations. The reference-aware grammar and compound ordering change limit the finding's operational scope and leave its internal cause unmeasured.
cs.CL / 23 / 2610.11451
SAIL: Scientific Agentic Intelligence via a Science-Aware Loop
SAIL Model Team, Boyuan Sun, Bryan Dai, Che Liu, Chi Liu, Derek Li, Hongming Piao, Mengzhuo Chen, Xidong Wang, Yan Shu, Yinda Chen, Ziyang Zeng
cs.CL · cs.DL · cs.IR
Abstract
We introduce SAIL, an open model with 35B total and 3B active parameters for literature research, scientific coding, and multi-step research workflows. SAIL is developed through a science-aware improvement loop: agents built on frontier AI models analyze its task failures and construct training tasks that address the underlying capability gaps. The diagnosis examines search and evidence selection in literature tasks, scientific assumptions and reasoning in coding, and planning and revision in longer investigations. The agents draw on paper collections and scientific code repositories to build problems, interaction trajectories, and executable tasks with the required environments and tools. We repeat this loop over multiple development cycles and train SAIL through supervised fine-tuning, specialist training, multi-teacher on-policy distillation, and agentic reinforcement learning. SAIL achieves competitive performance across scientific research tasks with substantially fewer parameters than leading open-weight models.
cs.CL / 24 / 2610.11519
Residual Advantage: Student-Relative Teacher Guidance for RL with Verifiable Rewards
Xiaobing Chen, Zhiqi Pang
cs.CL
Abstract
Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have become two main paradigms for post-training reasoning models. RLVR gives each response a single outcome label, leaving the steps inside it without separate credit. OPD provides token-level guidance at student-visited prefixes, but its pointwise signal does not directly reflect the pattern of teacher--student disagreement across the vocabulary. Dense, unbounded log-ratio supervision can amplify the teacher's influence, yet a strong solver is not necessarily a suitable guide when the student's solution paths depart from the teacher's. We propose Residual Advantage (\RA{}), which treats the teacher--student probability residual as a bounded one-step reward, subtracts the corresponding state value under the student policy to form a standard advantage, and centers the result within each response before adding it to the verifier advantage. The guidance term has zero mean within each response, so the verifier advantage remains the response's mean label and the teacher only redistributes credit among the steps within it. \CoRA{} further updates a teacher LoRA with verifier advantages on the same scored student batch and uses the updated teacher in the next iteration's residual, adapting guidance to the student's attempts. With Qwen3-1.7B-Base and Qwen3-4B-Base students and a Qwen3-8B teacher, \RA{} combined with GRPO or REINFORCE++ improves the underlying sequence-advantage algorithm in all 24 comparisons on three mathematical benchmarks, raising macro Avg@8 by 1.7--3.6 points and Pass@8 by 3.9--6.3 points. Both combinations surpass teacher-only OPD, and \CoRA{} adds a further 1.0--1.5 Avg@8 points.
cs.CL / 25 / 2610.11520
When Can You Prune Your Network? A Study of Intermediate Neurons in Multilingual Speech Parsing
Minnie Kabra, Benjamin Lecouteux, Maximin Coavoux
cs.CL
Abstract
End-to-end speech parsing, a task recently proposed, consists in predicting both the transcription and the syntactic tree for a spoken utterance. Existing architectures for speech parsing often utilise intermediate neural networks. In this work, we examine the effectiveness of intermediate neural networks (NN) for parsing, and, specifically, what role do they play. We introduce a simpler end-to-end architecture for speech parsing, where we remove these intermediate NN units, reducing the parameters by 12%, while achieving comparable or better performance than prior method on both automatic speech recognition (ASR) and parsing. We demonstrate that intermediate NN units help reduce the representational gap when the pre-trained encoder is frozen. We do a comprehensive evaluation of speech parsing on French, and medium-low resource languages Slovenian and Naija. We further investigate the impact of the training data size and intermediate layers of the pretrained speech encoder on speech parsing.
cs.CL / 26 / 2610.11542
Constitutional Gating and Deterministic Recovery for Multi-Agent LLM Negotiation: Ablations Against a Stateful Adversarial Gatekeeper
Masaaki Nakatsu, Reno Wang
cs.CL · cs.AI
Abstract
Multi-agent LLM systems negotiating with a stateful counterpart waste model calls in three ways: polite loops that never meet the counterpart's hidden acceptance condition, malformed outputs that trigger retries, and compliance deadlocks in which the counterpart demands something the agent must refuse. We study a three-part control stack - a 5-Pillar runtime constitution, a 4-tier swarm (Director, three-agent majority vote, Monitor, schema hard gate) and Cognitive Annealing (deterministic deadlock detection, atomic purge of the agent-side context, a canonical recovery message) - against a released adversarial Gatekeeper whose acceptance rules are fixed regular expressions and whose LLM only renders reply text. The testbed has a known solution: it measures whether the stack executes a constitution-aligned strategy against swarm drift and recovers from deadlock, not whether it discovers anything. In five runs per configuration (30 runs; Gemini 2.5 Pro agents, Claude Haiku 4.5 Gatekeeper) we find: (i) the constitution and Director make an acceptable framing possible but not reliable - 0/5 baseline unlocks versus 1/5 and 2/5 with the constitution; when the swarm unlocks it does so in one turn with 7-8 calls and about 15k tokens (67-73% below baseline); when it does not, it costs 17-38% more; (ii) the Monitor and hard gate do not reduce unlocks and leave an audit trail; (iii) under a honeytrap-to-compliance deadlock, LLM-only steering escapes 0 of 5 times while atomic purge plus a canonical strike escapes 5 of 5 (Fisher $p = 0.008$) at the same call budget, with zero calls for the strike. LLM-written strikes failed the deterministic pre-flight 5 of 5 times although an LLM Monitor had approved four. Pre-registered hypotheses on average call and token reduction were not supported. Cost is bounded in every arm by deterministic stop rules; the stack adds recovery at no extra model cost.
cs.CL / 27 / 2610.11543
Learning the Loop, Not Just the Page: Execution-Grounded Loop Learning for Web Generation
Yuxin Meng, Ruixu Zhang, Junjie Wang, Yuhan Suo, Yuhan Sun, Ruining Hu, Yiyao Yu, Yubin Wang, Shouwei Ruan, Bin Wang, Yue Liao, Yuxiang Zhang, Yujiu Yang
cs.CL
Abstract
Functional Web generation is increasingly optimized with executable rewards, yet existing methods largely focus on the quality of the final page and leave the process of diagnosing and repairing imperfect implementations underexplored. We identify a central challenge in this setting: the Generator and Refiner produce executable artifacts with direct environment rewards, whereas the intermediate Critic influences downstream behavior without a directly executable outcome. We introduce WebLoop, an execution-grounded framework that jointly learns generation, critique, and refinement within a shared policy. WebLoop trains an execution-free Critic with complementary signals for requirement-level discriminability and downstream helpfulness, first establishing reliable diagnosis and then introducing consequence-aware credit, while all three roles are jointly optimized with group-relative policy learning. With Qwen3.5-9B, WebLoop reaches 41.5 Overall on WebRise and 38.9% accuracy on WebGen-Bench, improving the base model by 11.3 and 15.4 points, respectively. The gains transfer to first-pass generation, persist at 27B scale, and generalize from text-only training to multimodal inputs. Controlled analyses further show that the improvement cannot be explained by an additional refinement pass alone, highlighting the importance of learning the Critic and the loop itself.
cs.CL / 28 / 2610.11544
Prosody-to-Text: Predicting text from low-pass filtered speech
David Porteš, Aleš Horák
cs.CL
Abstract
While predicting prosody from text is an established task in the field, the opposite direction, predicting text that fits a given prosodic pattern, remains largely overlooked. We find this unfortunate, because this opposite direction could lead to some very interesting use cases. Therefore, in this paper, we make the first steps in the prosody-to-text direction by inves- tigating how much of the original sentence can be recovered from its prosodic pattern. To this end, we fine-tune the Whis- per model using only the 12 lowest Mel bins (low-pass filter with approximately 450Hz cutoff), and obtain surprisingly accurate results (WER 36%), with 10% of utterances be- ing recovered perfectly, and 40% of utterances having Word Error Rate at or below 25%. We also find that, given the correct prefix, the next token was predicted correctly in 79% of cases. Our results suggest that the relationship between low-frequency speech features and lexical content is much stronger than previously thought, and we believe that direct- ing more attention to this topic might open the door to new applications, such as using prosody to guide text generation of modern LLMs
cs.CL / 29 / 2610.11559
SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction
Hexuan Deng, Yue Wang, Wenyu Jiang, Cheng Yang, Haolin Yang, Zhaohua Zhang, Chenchen Zhao, Beiduo Chen, Muxi Chen, Sa Zhu, Geyuan Zhu, Jianhuan Zhuo, Qiuyong Xiao, Tianwen Jiang, Jihong Zhang, Xuebo Liu
cs.CL · cs.AI · cs.MA · cs.SE
Abstract
Coding assistants such as Claude Code and Codex have become a major application of LLM agents, yet existing benchmarks remain far from real-world use, particularly in task horizon and interaction length. Code assistants require completing long chains of development work in continuously evolving repositories, while repeatedly clarifying requirements and adapting implementations through multi-turn interaction. To address these gaps, we introduce SWE-Journey, a benchmark for more realistic evaluation of coding assistants. To address the task-horizon gap, we propose a weak-to-strong synthesis pipeline that automatically constructs long-horizon coding tasks. To address the interaction gap, we mine four representative user personas from real interaction data and build a user-simulation agent to reproduce realistic code-assistance interactions. On average, models pass over 75% of tests for requested functionality with software architects, but fewer than 25% with non-coders. These results show that current coding assistants still fall short of enabling reliable coding for non-coders. We further analyze the reasons for this gap and identify asking right, finding right, and fixing right as key capabilities during interaction.
cs.CL / 30 / 2610.11566
Incremental Open-Ended Deep Research with Structured Harness
Meilin Chen, Hongyuan Bao
cs.CL
Abstract
Existing Open-Ended Deep Research (OEDR) systems primarily generate reports from scratch, making them inefficient for scenarios where research reports need to be continuously maintained as new information emerges. We introduce \textbf{Incremental Open-Ended Deep Research (Incremental-OEDR)}, a research setting that treats a report as an evolving research state and incrementally updates it by preserving valid knowledge, revising outdated or incomplete content, and incorporating newly available information. To support this setting, we propose \textbf{Structured Harness}, which represents reports as structured collections of outlines, sections, and supporting evidence, and provides structured retrieval, a persistent structured evidence pool, and structured generation for selective report updating and evidence reuse. We further establish a temporal evaluation framework spanning ten years, with \emph{Single-Step Task} and \emph{Long-Chain Task} to evaluate incremental updates over both individual transitions and long-term update chains. Extensive Experiments on DeepResearch Bench and DeepConsult under both the Open-source Configuration (OC) and Proprietary Configuration (PC) show that Incremental-OEDR maintains competitive report quality while substantially improving report continuity and reducing research costs. As shown in Figure~\ref{fig:profile}, it achieves up to 0.51 higher content-level ROUGE-L F1, 0.63 higher outline-level EM F1, 33\% lower token consumption, and 61\% fewer search calls than OEDR on DeepResearch Bench. For more details, please refer to our project page: https://ioedr-project.github.io/.
cs.CL / 31 / 2610.11575
Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts
Yunkai Chai, Tong Zhu, Xiaoye Qu, Xuyang Hu, Guanjie Chen, Qipeng Guo, Yu Cheng
cs.CL · cs.LG
Abstract
Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token. However, traditional training paradigms enforce a static choice of top-$k$ experts, which converts a continuous routing distribution into a rigid step function. This constraint introduces a brittle boundary where highly competitive experts are arbitrarily separated into full-supervision and zero-feedback zones based on minor score fluctuations. To address this issue, we propose Elastic Expert Routing, which stochastically samples the active expert budget from a localized discrete distribution centered at $k$. Over multiple training iterations, this mechanism softens the sharp threshold into a gradual probability distribution. Because the sampling neighborhood remains symmetric, this approach matches the expected computational cost of deterministic training, while preserving the inference budget. Extensive experiments demonstrate the efficacy of our method on both supervised fine-tuning and from-scratch pretraining settings. During supervised fine-tuning, elastic routing improves downstream macro-averages on OLMoE-1B-7B and Qwen3-30B-A3B by $+0.84$ and $+2.02$ points, respectively. In addition, in from-scratch pretraining, it outperforms the static top-$k$ baseline by $1.6$ points on average across downstream tasks.
cs.CL / 32 / 2610.11586
Measuring Cultural Alignment Beyond the Average: A Framework for Evaluating Maternal-Health LLM Interactions in Indian Contexts
Umaira Izhar, Gunjan Arora, Pushpendra Singh
cs.CL
Abstract
Existing evaluation methods for healthcare LLMs primarily assess factual correctness,safety, and fluency, while providing limited insight into whether generated interactions reflect culturally situated healthcare reasoning. This limitation is particularly important in maternal health, where care decisions are shaped by social and relational norms. We introduce MH-INDIC, a culturally grounded evaluation framework for maternal-health interactions in urban and semi-urban North Indian contexts that operationalises cultural behaviour through ten dimensions of maternal-health reasoning. Using a 26-item survey administered to 102 pregnant and postpartum women from urban and semi-urban North India, we evaluate ten LLMs. We distinguish population level cultural alignment from profile-level behavioural variation. Although several models approximate the human population-level distribution, all evaluated systems exhibit substantially lower variation across demographic and household profiles than the human cohort, revealing a gap between aggregate alignment and profile-conditioned sensitivity. As a downstream application of MH-INDIC, we use the strongest-aligned proprietary and open-source models to generate culturally conditioned maternal-health dialogues under zero-shot, self-conditioned, and human-grounded prompting. Human-grounded conditioning produces stronger profile alignment and dialogue quality ratings, suggesting that measured cultural profiles can improve the cultural grounding of generated interactions
cs.CL / 33 / 2610.11592
Probing for Long-Horizon Deductive Reasoning Capabilities in Language Models with Prolog
Hadeel Al-Negheimish, Jasna Ilieva, Yoon Kim
cs.CL
Abstract
Current frontier LLMs can theoretically process long contexts with 1M tokens or more. But to what extent can they go beyond simple retrieval and perform deeper reasoning over such long contexts? We empirically investigate long-horizon reasoning capabilities of LLMs, focusing on deductive logic expressed in Prolog. We construct ProloNg, a synthetic testbed to probe Prolog Long Reasoning, which systematically varies the complexity (reasoning depth) of problems, where the hardest case has a reasoning depth of 22 and 62k context length. We study 8 reasoning models across 5 families of frontier LLMs, and find that performance degrades substantially as reasoning depth grows, with the majority of models approaching chance beyond depth 10.
cs.CL / 34 / 2610.11638
UXBench Pro: Benchmarking Personalized User Experience in Multi-Turn Dialogue Interactions
Mengze Hong, Zeyang Lei, Wenbo Shang, Xia Zeng, Xiying Zhao, Qi Zhu, Chen Jason Zhang, Di Jiang, Taiming Fu, Qiongyi Zhou, Qinghe Chang, Fubao Zhang, Chenxuan Ma, Minlong Peng, Jinfeng Huang, Zineng Zhou, Jindou Wu, Muge Qi, Sijun He, Xin Cui, Di Liang, Yuan Hua, Davey Chen
cs.CL
Abstract
Evaluating user experience (UX) with automated computational methods has gained increasing attention, supported by empirical evidence from UXBench. However, binary preference prediction provides limited insight, while relying on a single user-agnostic reward model overlooks the inherent heterogeneity of users, whose expectations can differ substantially. In this paper, we present UXBench Pro, comprising 1{,}000 test instances derived from real user interactions across 12 task scenarios and 82 domains. Each instance is paired with a FACTORS user profile that characterizes the user through seven interpretable behavioral facets, differentiating user groups. To provide richer evaluation insights, we introduce a dual-perspective paradigm that combines a personalized User Reward Model (URM) for third-person judgment with Sim4Eval, a user simulator that enables multi-turn interactions and provides first-person evaluation across four cognitive state dimensions. To assess the reliability of these based evaluators, we further introduce two meta-benchmarks, URMBench and USimBench, that evaluate how faithfully they reproduce real human preferences and behaviors. Extensive experiments reveal seven key findings that highlight the importance of user modeling and multi-perspective evaluation, offering a fresh perspective on user-centric benchmarking and motivating personalized model optimization.
cs.CL / 35 / 2610.11646
Phonologically Informed Tokenization for German Speech Recognition: A Cross-Domain Study
Christopher Witzl, Tobias Bocklet, Korbinian Riedhammer
cs.CL · cs.LG
Abstract
German is a morphologically rich language whose syllable structure is exceptionally well-predicted by the Knuth--Liang hyphenation algorithm. We ask whether phonologically informed tokenization can serve as a competitive target for end-to-end speech recognition. We compare three tokenizer families on the Omnilingual ASR wav2vec 2.0 backbone fine-tuned with CTC: the pretrained multilingual character inventory, a data-driven Byte-Pair Encoding (BPE) over orthography, and phonologically informed units from Pyphen syllabification and grapheme-to-phoneme conversion. Across 40 fine-tunes, we evaluate on three German test sets spanning orthogonal shifts: in-domain read speech, dialectal spontaneous speech, and standard-German spontaneous speech. In-domain, all phonologically informed tokenizers match BPE and the multilingual character baseline on both WER and CER. Under domain shift the picture splits along vocabulary size rather than the linguistic axis of variation: at small vocabularies, syllable-aware tokenization improves on dialectal speech, where phonetic surface forms vary but syllable structure is preserved, and stays ahead on spontaneous speech, where new word-forms violate vocabulary closure. A phoneme-level confusion analysis further shows that all tokenizers commit the same canonical function-word errors, indicating that the acoustic encoder, not the tokenizer, dominates the error topology. Our findings suggest that tokenizer choice may depend on the vocabulary budget as much as on the distribution shift expected at deployment rather than reducing to a single universal optimum.
cs.CL / 36 / 2610.11659
DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation
Anhao Zhao, Haoran Xin, Junlong Tong, Yingqi Fan, Xuan Lu, Ping Nie, Wenjie Li, Xiaoyu Shen
cs.CL · cs.AI
Abstract
On-policy distillation (OPD) supervises student-generated trajectories with token-level teacher signals. Its sampled-token variant avoids the cost of full-vocabulary probabilities. Yet we find that training on fewer tokens can outperform full-token OPD, challenging the intuition that more supervision improves learning. This motivates selecting tokens by learning value. Existing disagreement-based criteria ignore probability scale: tokens assigned negligible probability by both models, termed low-low tokens, can receive large log-ratio rewards and hinder learning. We propose DIAL-OPD, a token-selection method that bridges log-probability and probability spaces by weighting reward magnitude with the logarithmic mean of teacher and student probabilities. A parameter beta controls this weighting, and the highest-scoring tokens are retained. Across 4 teacher-student pairs and 7 mathematical reasoning benchmarks, we compare DIAL-OPD with 9 baselines. Retaining only 40% of tokens, it outperforms Vanilla OPD and its full-token variants, with mean accuracy gains reaching 5.25 percentage points over Vanilla OPD, and doubles AIME25 Pass@16 from 13.33% to 26.67%. It also achieves up to an 18% relative improvement in mean accuracy over the strongest token-selection baseline at matched retention ratios. With a 4B teacher, DIAL-OPD surpasses the strongest full-token baseline using an 8B teacher at both student scales, showing that effective supervision allocation can outweigh teacher scaling. Further analysis shows that moderate beta balances suppressing low-low tokens against preserving useful disagreements. Token-level evidence reveals that DIAL-OPD filters high-reward tokens with limited reasoning value while preserving supervision critical to reasoning correctness.
cs.CL / 37 / 2610.11695
Structured Sentiment Analysis Using Sequence Labeling as Dependency Graph Parsing
Muhammad Imran, Ana Ezquerro, Carlos Gómez-Rodríguez, Anders Søgaard, David Vilares
cs.CL
Abstract
This study addresses the problem of structured sentiment analysis, whose goal is to obtain a fine-grained sentiment graph where the nodes represent spans of sentiment holders, targets, and expressions, while the arcs define the relationships among them. Our proposed approach casts the task as dependency graph parsing, but departs from traditional parsing methods by solving it through sequence labeling. To do so, we leverage recent advances in linearized graph encodings that allow each word in the input to be assigned a label, effectively capturing the structure of the dependency graph. We conducted experiments on seven datasets spanning five languages (English, Spanish, Norwegian, Basque, and Catalan), showing performance competitive with leading, more complex single-model approaches.
cs.CL / 38 / 2610.11776
DPPM: Dual-Path Parametric Memory for Personalized Language Models
Yuhao Chen, Shuochen Liu, Jiayao Shi, Jian Hong, Chen Cheng, Xinyun Ding, Tao Wang, Ya Li, Quan Liu, Tong Xu
cs.CL
Abstract
Long-term personalization requires language models to use interaction history to track users' preferences across sessions. Parametric memory encodes this interaction history into model parameters or adapters, reducing the need to include it in the inference context. However, independent context compilation leaves cross-session integration unspecified, while recurrent updates can attenuate earlier evidence. To address these challenges, we propose Dual-Path Parametric Memory (DPPM). Its Evidence path directly pools representations of the interaction history to preserve earlier evidence, while its Delta path sequentially updates an associative state to capture changes. Fusing both outputs produces history-conditioned LoRA adapters that combine evidence accumulation with ordered revision. Across multiple backbones, DPPM outperforms the evaluated baselines, achieving 54.22% on PersonaMem-v2 and 86.79% on PrefEval. These results suggest that DPPM provides a simple and effective design choice for cross-session personalized parametric memory.
cs.CL / 39 / 2610.11787
From Sparse Representations to Behavioral Insights for Multimodal Depression Assessment
Guimin Hu, Zihao Song, Jiachen Luo, Jiayuan Xie, Ruichu Cai
cs.CL
Abstract
Multimodal depression assessment offers a promising approach to analyzing behavioral patterns associated with depression. However, existing methods often rely on dense and opaque multimodal representations, making it difficult to interpret the behavioral patterns underlying their predictions. In this work, we introduce BehavDep, a sparse factor-based framework that decomposes multimodal behavioral representations into sparse latent factors and associates them with behaviorally meaningful concepts through a semantic bridge. To address the mismatch between user-level annotations and heterogeneous video-level behaviors, BehavDep further learns video-level depression tendency scores under weak supervision and aggregates information across multiple observations for user-level assessment. Extensive experiments demonstrate that BehavDep achieves the best overall assessment performance while revealing complementary modality contributions, heterogeneous behavioral patterns across observations, and prediction responses to concept-level editing. These results show that BehavDep provides a structured and interpretable approach to analyzing multimodal behavioral representations for depression assessment.
cs.CL / 40 / 2610.11790
Easy to anticipate, hard to compute: boundary dependence finds the computed outputs that entropy patching misses
Nicolás Vera Zúñiga
cs.CL
Abstract
Byte-level language models such as the Byte Latent Transformer (BLT) group bytes into patches and run their large global model once per patch. BLT starts a patch where a small model's next-byte entropy is high, so global compute goes where the next byte is hard to predict. We show that this rule has a systematic blind spot: positions whose type is predictable but whose value must be computed, such as the number after "=" in a worked math solution. Under tight patch budgets, entropy-triggered layouts skip these positions, and accuracy on them collapses. In Meta's BLT-1B with patch starts on 10% of bytes, the entropy rule puts a patch start at 16% of the computed results in GSM8K solutions and gets 19.0% of them exactly right; a boundary after each "=" at the same patch count gets 51.8%, and entropy combined with a label-free boundary-dependence signal gets 67.1% (default layout at 26% of bytes: 76.8%). The gap survives adapting BLT-1B to the budget with low-rank fine-tuning (32.9% vs 72.7%, three runs per rule, paired p < 1e-200) and grows with model size in byte models trained from scratch at a 10% budget: at 1M, 12M and 50M parameters, boundary dependence beats entropy on final answers by -1.6, +10.1 and +19.8 points, and at 50M it gets 35.9% of computed results against 13.9% (3 seeds each). BLT's entropy-jump rule helps neither target at 50M. The entropy trigger of Scratchpad Patching is likewise indistinguishable from random scratchpads on final answers (5.6% vs 5.9%, 5 seeds), while answer-start scratchpads give 38.1%. The effect is specific to computed values: copies and lookups gain little, and values the model cannot compute gain nothing. Boundary dependence, the rise in the model's own loss when a patch start is removed, measured per two-byte context, finds these positions without labels: combined with entropy it beats the hand-written rule on computed results.
cs.CL / 41 / 2610.11920
Event-Centric Memory with Query-Aware Graph Augmentation for Long-Term Conversational Agents
Yichen Liu, Chunfeng Yuan, Haowei Liu, Wenjuan Li, Zefeng Lin, Bing Li, Xu Chen, Weiming Hu
cs.CL
Abstract
For persistent and personalized conversational agents, memory systems can enable them to remember, update, and reason over long histories by storing past interactions and retrieving relevant information. Existing memory systems typically follow two paradigms: flat-structured memory and graph-based memory. The former is lightweight but leaves event relations and state updates implicit, while the latter explicitly models memory structure but incurs additional construction cost and introduces irrelevant relations over long histories. To address these limitations, we propose QGMem, a novel memory construction and activation framework motivated by human memory, in which experience is organized into events and query-relevant events are modeled by graph as working memory. QGMem converts long dialogue histories into event-indexed atomic memory units that preserve individual experiences and consolidates related units into dynamic memory traces that retain state trajectories and current states. When a query arrives, hybrid memory retrieval gathers complementary candidate memories, and query-aware reranking activates the most relevant units as a compact working memory. To expose relational dependencies in the working memory and support conflict-aware reasoning, QGMem organizes the working memory as a local graph, which is then encoded as a graph token and provided to the LLM together with the textual working memory to improve evidence utilization during answer generation. Experiments across six benchmarks validate the framework and show consistent gains in retrieval, multi-hop evidence composition, conflict resolution, and ultra-long dialogue reasoning with compact contexts and moderate inference cost.
cs.CL / 42 / 2610.11948
When History Helps and Hurts: Selective History Use across Multimodal Turns
Shuoyang Sun, Kerui Gu, Hao Fang, Shaoli Huang, Bin Chen
cs.CL
Abstract
Reliable multimodal interaction depends on selective use of conversational history: an earlier question may remain relevant while its previous answer is outdated, whereas a current request may depend on historical evidence despite conflicting new observations. Existing multi-turn evaluations rarely separate these history-use demands from underlying question difficulty. To address this gap, we introduce ReTurn, a benchmark of 7,000 base tasks spanning visual and audio evidence for evaluating selective history use. For task-carrying history, Reconfirm/Reground require applying a historical question to current media while varying historical agreement; for evidence-carrying history, Retrieve/Rebind require answering a current question using historical media while varying current-media competition. Each pair preserves the target question, media, and answer. Tasks support open-ended and multiple-choice evaluation, with matched single-turn counterparts serving as answerability references. Across 13 omni-modal, vision-language, and audio-language models, median model-level open-ended accuracy falls from 93.7% with direct input to 72.3% in conversation. Behavioral probes show that high question recall can coexist with weaker task application, while competing media can redirect answers away from historical targets. Supervised adaptation yields only partial gains. ReTurn provides a controlled framework for assessing whether multimodal models select and use the historical information required by each request.
cs.CL / 43 / 2610.11959
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
Xiaomi LLM-Core Team, :, Zongming Qiao, Ziyue Hua, Zirui Ou, Zihao Yue, Zihan Jiang, Zhuo Huang, Zhiyang Chen, Zhixian Zheng, Zhipeng Xu, Zhengrui Ma, Yuyang Hu, Yuhang Dong, Yuechen Zhang, Yudong Wang, Yuanxin Liu, Yixin Yang, Yishuo Cai, Yikai Zhao, Yihan Yan, Yifan Zhang, Yifan Song, Xiyu Wei, Xing Zhang, Xin Zhang, Xiaoqian Liu, Xiaodong Ji, Xiangwei Deng, Xueyu Guo, Wenhan Ma, Weimin Xiong, Weikun Wang, Weiji Zhuang, Shuo Liu, Shuhuai Ren, Shuhao Gu, Shimao Chen, Shijie Cao, Shihua Yu, Shicheng Li, Shengjie Zhou, Shaolei Zhang, Rang Li, Qiying Wang, Qingkai Fang, Qianli Chen, Minzheng Wang, Liwen Wang, Linli Yao, Linghao Zhang, Liangyu Cheng, Liang Zhao, Lei Li, Jinhao Dong, Jinyu Xiang, Jianyu Wei, Jiangshan Duo, Huaqiu Liu, Huanjie Fan, Hongyi Guan, Hongshen Xu, Hao Tian, Hanyu Li, Hailin Zhang, Gang Wang, Fuli Luo, Feng Wei, Dong Zhang, Dawei Zhu, Chiheng Lou, Chenhong He, Chenhao He, Chenghua Liu, Bowen Ye, Bowen Shen, Boshen Xu, Bo Yang, Bingquan Xia, Bangjun Xiao, Baixuan Xu, Zhouxiang Mao, Zhiyang Zhang, Zhixiang Xu, Zhenru Lin, Zhengju Tang, Zhaojun Huang, Yuzhe Weng, Yuxing Xiang, Yuxiao Li, Yuheng Yang, Yuhang Wang, Yuchen Liu, Yuanyuan Tian, Yuanliang Dong, Yu Cheng, Yongzhe He, Yongshun Liang, Yong Wang, Yiyan Wang, Yitian Gong, Yijie Zhang, Yanshu Xin, Xun Zhang, Xingjian Zhao, Wenyu Yang, Wenshan Huang, Wenhao Li, Tingwei Huang, Tianyu Yu, Tianyang Lu, Taoyu Yang, Sinan Du, Shutong Tian, Shulin Du, Shengfan Wang, Shanchuan Fang, Qihao Zhang, Qibin Yang, Qian Yu, Qian Tu, Pengrong Xie, Peipei Wang, Peidian Li, Minkun Guo, Mingchen Shao, Luohan Gao, Lijie Wang, Liang Shi, Kaiqi Chen, Kaiming Liu, Kaifei Wang, Kai Yang, Jinlong Xue, Jiechen Zhang, Jiaxuan Liu, Hongxu An, Hao Peng, Hanglong Lü, Guonan Wang, Feiyu Yang, Fanyu Cao, Fangyue Liu, Fan Cui, Cong Wang, Chun Chen, Chenxu Bai, Chengxuan Zhu, Chenghua Wang, Boyi Zeng
cs.CL
Abstract
Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement. This report introduces the MiMo-V2.6 series, an omni-modal family that pushes the frontier of model intelligence by scaling RL compute. Prior to RL, we conduct mid-training on a broad multimodal corpus to provide ample exploration space, and build a solid infrastructure on the pretrained hybrid-SWA architecture to support subsequent scale-up. We scale RL compute along three dimensions: (1) larger batches and higher throughput, with an asynchronous training that consumes 1,568 samples and 2.7-3.7B tokens per step at context lengths of up to 1M; (2) more diverse and complex environments, spanning code, general, visual, and cyber domains under a mixture of agent harnesses; and (3) more grader compute, via groupwise agentic grading that yields more accurate reward signals for long-horizon tasks and steers the model towards shorter, more token-efficient solutions. To keep training stable at scale, we freeze the MoE router and establish a multi-layer defense against reward hacking. We further build infrastructure for mixed-task agentic RL, including a unified trajectory representation, high-concurrency multi-framework rollout, decoupled control and data planes, and training-inference consistency. We open-source the training dynamics, RL environments, and RL framework to facilitate reproduction and further research on scaled RL and model self-improvement.
cs.CL / 44 / 2610.12064
ILM: An AI-Powered Storytelling Educational Tool
Suhaila Mohammed, Abdelaziz Serour, Allison Lahnala
cs.CL
Abstract
Digital technologies have made Islamic narratives more accessible, but existing platforms provide limited support for structured learning and comprehension of these stories, particularly in Arabic and multilingual settings. We present ILM, an interactive educational platform for Stories of the Prophets that combines Arabic natural language processing, structured knowledge representation, and retrieval-based question generation. Admin-approved Arabic narratives are processed by a Knowledge Graph (KG) Constructor Engine that identifies entities and narrative relationships and stores them as structured knowledge, enabling learners to explore stories through a visual story map and answer entity- and relation-based questions generated from the KG. Separately, a multilingual retrieval pipeline retrieves relevant passages from the original narratives to generate multiple-choice and open-ended comprehension questions. For open-ended questions, an LLM-as-a-Judge evaluates learners' answers against the retrieved passages and reference answers to determine correctness. The platform also incorporates Quranic content as a separate enrichment layer, allowing selected narratives to be supplemented with source-supported information. By combining structured knowledge with passage-based retrieval, ILM supports narrative exploration, comprehension, and assessment across Arabic and multilingual content. The system demonstrates the feasibility of combining structured knowledge representation and retrieval-based generation to support interactive learning of Islamic narratives. A demo is available at anonymous.4open.science/r/mml-5FCF.
cs.CL / 45 / 2610.12083
All Verdicts are Not Equal: Rethinking LLM Judge Reliability
Vineet Kumar, Darshita Rathore, Anindya Moitra
cs.CL · cs.AI
Abstract
LLM-as-a-Judge is the standard paradigm for NLP evaluation, yet its systemic reliability remains poorly understood despite being widely treated as a deterministic ground truth. We present a comprehensive reliability audit, stresstesting six frontier models across four benchmarks, five prompt formats, two presentation orders, three sampling temperatures, and ten repetitions per condition. Our empirical analysis reveals severe vulnerabilities: verdicts change across identical replications at temperature zero, position-order swaps flip the majority of verdicts on challenging tasks, and the most deterministic judge achieves perfect consistency by trivially repeating incorrect verdicts, agreeing with ground truth only 51% of the time. To formalize these multi-faceted failure modes, we introduce the trustworthy verdict rate (T ), a unified metric capturing the joint probability that an evaluation is reproducible, order-invariant, and accurate. UsingT , we derive a theoretical upper bound on accuracy imposed by position bias and show that reliability is item-specific rather than modellevel. Finally, we demonstrate that shifting from pairwise win-rate to holistic rubric scoring improves trustworthiness more than any single-format prompting intervention, offering an actionable framework for robust NLP evaluation.
cs.CL / 46 / 2610.12133
Rehearse Everything, Remember Nothing: Attic-KV Rehearses What Will Be Read
Zhiyun Shi
cs.CL · cs.LG
Abstract
Many key-value (KV) caches are compressed before anyone knows what will be asked of them: a document cached for retrieval, a prompt prefix shared across requests, the memory of a long conversation. The prevailing approach scores KV entries by rehearsal: the model rereads the context and keeps the entries it attends to, assuming that the more completely a cache rehearses its context, the better it remembers it. We show that under tight budgets this assumption backfires: rehearse everything, remember nothing. At a 3% keep ratio, rereading the whole context keeps 31.5 of 96.5 points on RULER, and on LongBench's natural-text tasks it falls below methods that rehearse nothing at all. The cause is that a cache keeps what it rehearses: rereading spreads the budget across the whole context, so the answer's own entries survive at little more than chance. Like a student before an exam, a cache remembers more by testing itself than by rereading. Two principles follow: rehearse what will be read, and rehearse as much as there is. We instantiate them as Attic-KV (Attic for short), a training-free rehearsal in which the model quizzes itself with question-answer pairs that quote the context, alongside anchor tokens in a content-adaptive amount. Changing only the rehearsal lifts three hosts that score it in three different ways: Attic alone is the best training-free method in all eight settings we test on RULER and LongBench's natural-text tasks, and plugged into the gradient-based KVgrad and the trained RestoreKV+, it raises them by up to 17.1 and 28.1 points. Its advantage grows as the budget shrinks, reaching 41.9 points over full rereading at a 3% keep ratio, and it compresses faster than rereading the whole context.
cs.CL / 47 / 2610.12144
Language-Specific Effects of Tokenizer Choice in Multilingual Language Models
Clara Meister, Gül Sena Altıntaş, Antoine Bosselut
cs.CL
Abstract
Tokenizer choice affects multilingual language modeling, but vocabulary capacity is finite and vocabulary size is often constrained: improving representation for some languages often comes at the expense of others. We therefore ask whether tokenizer choice matters equally across languages, a question that the current literature leave unanswered. To this end, we train 123 language models spanning 54 tokenizers. In the main comparison, architecture, training corpus, training-token budget, and optimization are held fixed, so the models differ only in their tokenizer. We find that tokenizer choice matters more for languages with less language-model training data: across the 54 tokenizers, the standard deviation of a language's bits-per-byte (BPB) increases as its model training-data share decreases (Spearman rho = -0.52 over the 31 trained languages and -0.69 over the 28 written with word boundaries). Leaving a language out of tokenizer training raises its BPB in every language we study, and the penalty tends to be larger for languages with less language-model training data. Giving lower-resource languages a larger share of tokenizer-training data, however, does not unconditionally help those languages: both equal weighting and an allocation inverting the shares with respect to the language model training data increase their BPB, particularly when language-model training repeats data. Finally, which intrinsic tokenizer properties are associated with better BPB differs across languages, providing further evidence that what makes a good tokenizer depends on the language. We find that the metrics quantifying these properties can be successfully used to predict downstream models' pairwise BPB rankings, suggesting a practical strategy for screening tokenizer candidates before training language models.
cs.CL / 48 / 2610.12207
SciTBERT: A family of chronologically consistent language models for scientific and technological language processing
Thomas Gebhart, Russell J. Funk
cs.CL · cs.LG
Abstract
Pre-trained transformer models are increasingly being used to study scientific and technological progress. Encoders tuned to paper or patent text outperform general-purpose models on downstream classification, regression, and proximity tasks within science and technology. However, the applicability of these models for studying time-dependent or archival properties of science, technology, and their interface is limited due to lookahead and domain biases inherent to these pre-trained models. These limitations arise from training on corpora with unconstrained chronological and text source distributions. We introduce SciTBERT: a family of chronologically consistent BERT-derived language models trained on text from scientific papers, patents, and high-quality educational web text with training data cutoff dates spanning each year between 2013 and 2025. We also post-train these models in a chronologically-consistent manner using paper and patent citations, creating SciTBERT-CI model family. We find that these models generally outperform predecessor domain-specific encoder models even when training data is limited by early year restrictions in the corpus. To further investigate the extent to which this class of models can learn representations that bridge science and technology, we introduce the PatRepEval benchmark, a suite of patent-related text embedding tasks at the science-technology interface. Performance in a variety of classification, regression, and retrieval tasks spanning papers and patents highlights the importance of aligning encoder model representations with the domain distributions of their downstream tasks, and chronologically consistent encoders can match or exceed models trained without temporal constraints.
cs.CL / 49 / 2610.12235
Language Models as AI Research World Models
Zijun Wang, Zewen Liu, Minhua Lin, Zhaotian Weng, Zhan Shi, Bing He, Yisi Sang, Dakuo Wang, Benoit Dumoulin, Wei Jin, Yuyin Zhou, Cihang Xie, Hanqing Lu
cs.CL
Abstract
AI research agents automate the cycle of proposing, implementing, and evaluating experiments, opening a path toward recursive self-improvement. Yet their ability to propose experiments outpaces their capacity to execute them in real environments, making outcome prediction a key capability for sustained self-improvement under limited experimental budgets. We investigate language models as Research World Models (RWMs), which predict the outcomes of candidate interventions across research environments. Our evaluation draws on over 2,600 experimental records from nine research environments spanning pretraining, post-training, and inference, representing more than 171,000 H100 GPU-hours of experimentation. Research knowledge acquired from real experimental experience improves RWM predictions of unseen interventions within the same environment (Spearman +0.27), and can be reused across environments. For example, using only pretraining experience from OLMo3, Marin, and Nanochat, an RWM reduces selection regret in the Qwen3 environment by 78% compared with zero-experience setting. These benefits extend to multi-round Autoresearch under a fixed selection budget: RWMs with in-env and cross-env research knowledge increase the best gain achieved by 15.8% and 11.6%, respectively. Ablations across 13 language models used as RWMs show that adding research knowledge can improve intervention ranking more than changing models or increasing reasoning effort alone. These findings support language models as RWMs and motivate accumulating experimental data for future RWM training.
cs.CL / 50 / 2610.12248
EgoVoice: Proactive Spoken Assistance from Egocentric Multimodal Streams
Heeseung Kim
cs.CL · cs.CV · cs.SD
Abstract
Wearable augmented reality (AR) assistants are moving toward continuous real-world interaction, where they perceive the user's activity through first-person video and audio and provide timely spoken guidance without being explicitly asked. While proactive video assistants, spoken dialog systems, and egocentric task understanding have each advanced rapidly, existing systems do not address the joint problem of deciding when to speak and what to say from continuous first-person streams. We introduce EgoVoice, a framework for training and evaluating proactive egocentric spoken assistants. From HoloAssist video recordings of real human instructors, we construct clean audio streams through source separation and speech resynthesis, and convert each video session into a format where the model must decide at each moment whether to remain silent or provide spoken guidance. We fine-tune an omni-modal LLM with our data, and further improve its proactive intervention behavior with direct preference optimization. Experiments across closed and open-source models show that existing systems rarely produce well-timed, meaningful proactive interventions, while EgoVoice yields clear improvements in intervention timing, content relevance, and human preference over the zero-shot backbone.
cs.CL / 51 / 2610.12274
HarnessSQL: Harness-Native Training for SQL Agents in Realistic Database Environments
Haolin Yang, Jipeng Zhang, Jian Xie, Shuaishuai Gong, Sirui Han, Yike Guo
cs.CL
Abstract
Text-to-SQL models are commonly trained to map questions directly to static queries, whereas real-world database agents operate through stateful, multi-turn interaction with live databases -- inspecting schemas, executing probe queries, diagnosing errors, and revising hypotheses. This creates a critical train-deploy mismatch, as the execution harness that mediates this interaction is introduced only at inference time. To bridge this gap, we propose HarnessSQL, a harness-native post-training framework that preserves the full interaction structure throughout both supervised fine-tuning and reinforcement learning. HarnessSQL builds isolated, executable database environments paired with hidden execution oracles, rolls out teachers directly inside the target SQL harness, and retains only verified trajectories for full-sequence SFT, followed by execution-reward RL. Across Spider 2.0-SQLite, HarnessSQL dramatically boosts the execution accuracy of compact models, raising Qwen3-8B from 15.5% to 45.2% and Qwen3-14B from 22.2% to 54.8%, while transferring effectively to out-of-distribution interactive benchmarks such as BIRD-Interact and LiveSQLBench. Our findings demonstrate that training database agents directly within their execution harness is essential for mastering complex, long-horizon database workflows.
cs.CL / 52 / 2610.12367
Which Skill to Distill? SGUID: Selecting a Compact Skill Bank for Model-Skill Co-Evolution
Yuhan Liu, Xiyao Ma, Zhongkai Sun, Xu Han, Chengyuan Ma, Benjamin Z. Yao, Chenlei Guo
cs.CL
Abstract
Skills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026). Prior work retrieves skills from a bank by semantic relevance, then uses them as inference-time patches or for model distillation. The individual utility of each skill, however, is largely neglected. We first show that, in on-policy distillation where skill-conditioned policies serve as teachers, fewer than 25% of retrieved skills provide useful distillation signals. We then propose SGUID, a method for selecting a compact subset of skills for distillation. SGUID retains a skill only if it consistently yields effective learning signals during training. The selected skills are then distilled to produce a better model. Our results show that not all skills are worth distilling. Across four models from the Olmo and Qwen families, distilling 6 selected skills matches or exceeds full-bank distillation in mean avg@12 on three of the four models, and on all four after a second round that distills 3 newly selected skills, while the full banks are up to 11x larger. Importantly, SGUID supports stable model-skill co-evolution: after a distillation round, a new candidate bank is curated from the updated model's rollouts, and SGUID selects which skills to internalize next. In the second round, this loop selects 3 new skills and improves Qwen3-8B from 64.3% to 66.3%. The selection step is essential for stability: on Qwen3-4B, naively updating the model with unfiltered skills degrades performance, including a 0.3 percentage point drop on HMMT25, whereas SGUID improves HMMT25 by 0.5 points after the first round and 1.1 points after the second. These results identify skill selection as the key mechanism for stable model-skill co-evolution.
cs.CL / 53 / 2610.12376
Latent Core Tokenizer: Compress, but Meaningfully
Felermino D. M. A. Ali, Millicent Ochieng, Ogbemi Ekwejunor-Etchie, Ade Famoti, Jacki O'Neill, Debjit Paul
cs.CL
Abstract
Tokenizers are commonly optimized for compression, but a compact vocabulary does not necessarily distribute its capacity evenly across languages. We introduce the Latent Core Tokenizer (LCT), a language-agnostic approach that separates structural discovery from vocabulary construction. LCT uses Minimum Description Length, entropy-based boundary signals, and morphotactic constraints to identify reusable linguistic units before constructing a shared vocabulary. Across 104 languages with a 200K-token vocabulary, LCT achieves lower fertility and higher MorphScore than BPE, Unigram, and parity-aware BPE, while maintaining comparable cross-lingual disparity in tokenization cost. Across four multilingual downstream benchmarks, LCT improves aggregate score by 1.48, 1.83, and 2.00 points over BPE, Unigram, and parity-aware BPE, respectively. Our findings show that compression alone does not predict representation quality and highlight the importance of morphology-driven structural discovery and how frequency is used to allocate the final vocabulary across languages.
cs.CL / 54 / 2610.12410
Predicting Alignment Generalization with Value Representations
Andy Liu, Mehar Bhatia, Karolina Stanczak, Mona Diab, Vered Shwartz, Daniel Fried
cs.CL · cs.AI · cs.LG
Abstract
LLM developers post-train their models to exhibit prosocial values and behavioral traits, which are enumerated in an alignment target. However, while recent post-training developments have yielded models that score highly on alignment evaluations, training models on sets of narrow behaviors still influences their behavior across unseen contexts and environments in unexpected ways. In this paper, we establish the task of alignment generalization prediction, i.e., predicting how fine-tuning a model to follow a given value changes its behavior across a wide range of held-out values. We conduct a large-scale analysis of alignment generalization effects across 66 values found in modern alignment targets, and benchmark representational techniques on the alignment generalization prediction task. We find that representations based on model activations when applying values in context significantly outperform methods based on textual descriptions of the values. Specifically, the best activations-based methods achieve correlations of 0.45 with our generalization matrix, compared with 0.05 from description-based baselines. We then show the applicability of representations that predict alignment generalization toward downstream tasks by using them to measure how similar the values in a multi-value alignment target are, which we find is significantly correlated with model robustness. Finally, we show initial evidence towards a shared, model-independent value space, which we use to develop the first taxonomy of LLM values grounded in empirical generalization dynamics. Our work demonstrates the importance of studying value generalization in LLMs and its application toward the more empirical design and training of model behavior.
cs.CL / 55 / 2610.11233
MiniVer-V: Identifying Minimal Sufficient Evidence for Short Video Verification
Leran Chen, Lingnan Kong, Zile Cai
cs.CV · cs.CL · cs.MM
Abstract
A core challenge in short-video fact-checking is identifying which evidence is sufficient to support a verification conclusion. Existing approaches either give the verifier all available evidence, introducing noise, or select evidence by topical relevance, which conflates relatedness with sufficiency. We identify evidential sufficiency as the selection criterion: whether a subset of evidence is adequate to support a confident verdict without redundancy. We introduce MiniVer-V, a benchmark of 195 short videos with three-way verdict annotations (supported, refuted, insufficient) and 5,510 multimodal evidence units spanning visual keyframes, speech transcripts, and web-retrieved external sources. We propose a two-layer verification framework that separates claim-video consistency, assessed from internal evidence, from factual verdict determination, which additionally requires external corroboration. On top of it, a sufficiency-driven greedy search assembles evidence until a sufficiency threshold is met and outputs insufficient when the candidate pool is exhausted, rather than forcing a verdict. With Claude Sonnet 4, the method reaches a Macro-F1 of 0.510 using 4.5 evidence units on average (16% of the full evidence set), statistically indistinguishable from the full-evidence baseline (0.518 with 27.7 units), while significantly improving recognition of insufficient cases over the same search without abstention. The efficiency result replicates with GPT-5.5 and holds only partially with an open-weight Qwen2.5-72B verifier. Ablations show that external evidence is indispensable for factual determination, while internal video evidence grounds the verdict in claim-video consistency. These findings suggest that evidence-efficient verification is achievable, and that explicit abstention is needed when evidence is genuinely inadequate.
cs.CL / 56 / 2610.12402
SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
Hongxing Li, Jinyue Su, Dingming Li, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen
cs.CV · cs.CL
Abstract
Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input. Yet real-world spatial intelligence demands predictive spatial reasoning: constructing a scene from observations, anticipating how an intervention changes it, and reasoning about the unseen outcome. We introduce SpaceCast-Bench, the first benchmark to directly and diagnostically evaluate this capability. Built around an observe-transform-infer framework, its 3,862 questions from 182 real-world scenes span 16 task types at three levels: static perception, local prediction, and global prediction, progressively requiring scene understanding, spatial state updating, and relational inference over unobserved outcomes. Evaluating 21 models exposes a stark gap: the strongest model reaches only 58.0% against 87.2% human performance, while spatially specialized models remain near random chance. Controlled analyses further reveal that bridge views are critical for integrating distributed observations, and that explicit 3D evidence benefits models more reliably than generated outcome images or videos. Fine-tuning on our programmatically generated data lifts Qwen3-VL-4B from 34.0% to 65.7% with macro-average gains across six out-of-domain benchmarks.
cs.CL / 57 / 2610.12403
ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills
Hongxing Li, Dingming Li, Yixin Li, Yong Du, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen
cs.CV · cs.CL
Abstract
Skill-augmented agents improve sample efficiency by distilling successful trajectories into reusable strategies. Yet most existing approaches remain text-centric, linearizing spatial layouts and action-state correspondences into language that loses critical geometric structure. Recent efforts have begun incorporating visual evidence, but construct and update skills separately from policy optimization, leaving their mutual improvement underexplored. We propose ViSkill, a visual-native skill learning framework that encodes successful interactions as composite visual skill cards directly accessible to VLM agents. Retrieved skills guide both inference and reward shaping, while successful trajectories are distilled back into the library, forming a closed feedback loop in which skill accumulation and policy improvement reinforce each other. An optional cold-start mechanism further accelerates early-stage learning. Evaluated on Sokoban, FrozenLake, and PrimitiveSkill, ViSkill achieves an overall success rate of 0.89, rising to 0.91 with cold-start initialization, outperforming all evaluated proprietary and open-source baselines while converging faster than standard PPO. Our code is available at https://github.com/ZJU-REAL/ViSkill.
cs.CL / 58 / 2610.11337
Type-Checking for Pattern-Based Tree Transformations
C. Aiswarya, Sahil Mhaskar, M. Praveen
cs.FL · cs.CL
Abstract
We introduce and study pattern-based tree transformations. As an illustrating example, consider a source pattern $(x \cdot y) + (x \cdot z)$ and a target pattern $x \cdot (y + z)$ as a pair. This source pattern matches any expression $e$ of the form $(e_1 \cdot e_2) + (e_1 \cdot e_3)$ (by substituting $x$ with $e_1$, $y$ with $e_2$, and $z$ with $e_3$) and the pair transforms it into the expression $e_1 \cdot (e_2 + e_3)$ as dictated by the target pattern. Note that in this example, the set of expressions that match the source pattern is not a regular tree language. We propose a model of tree transformations given by a finite representation of a (possibly infinite) set of such (source pattern, target pattern) pairs. The expressive power of this model comes at the cost of undecidability of checking equivalence. Nevertheless, we show that the type-checking problem is decidable for our model of pattern-based tree transformations. The type-checking problem asks whether applying a given transformation to trees having a given regular property (type) preserves the property. Our decision procedure is by a reduction to the emptiness problem of alternating tree automata.
cs.CL / 59 / 2610.11922
Project Greenhouse: Progress Toward Fully Open and Sovereign Agentic Search
Jimmy Lin, Sahel Sharifymoghaddam, Lingwei Gu, Nour Jedidi
cs.IR · cs.CL
Abstract
Project Greenhouse represents our exploration of a simple thesis: We believe that it is possible to build fully open and sovereign models for agentic search with only modest computational resources. As a first milestone, we describe how to build a competitive pointwise decoder-only reranker using a simple two-step recipe comprising pre-training from scratch followed by supervised fine-tuning, starting only from commonly available datasets. Contrary to the dominant approach in the literature, we do not rely on existing open-weight backbones from third parties, and thus we are fully in control of model training, from end to end. We were able to accomplish the bulk of our experiments using no more than a handful of GPUs. This report articulates the importance and benefits of our approach, and we share artifacts that enable transparent, independent reproduction of all aspects of model training. Beyond data, code, and configurations that capture our efforts, we also release checkpoints for our family of Gaggle models, demonstrating the feasibility of our approach and providing a first step toward validating our broader thesis.
cs.CL / 60 / 2610.12243
NativeScope: Relation-Localized Retrieval over Native Topology with a Correct Anchor
Long Wang
cs.IR · cs.CL
Abstract
Dense retrieval usually ranks text chunks by their semantic similarity to a question. This ignores structure that many data systems already store, including section membership, session boundaries, and native order. We propose NativeScope, a scope-then-rank method for queries with a known anchor and relation. It represents a query as q -> (A, r, B). The anchor A and relation r select native units through belonging, before, or after operators, and the target term B ranks only chunks that overlap the selected scope. An internal variant, NS-FullQ, ranks the same candidates with the full question. We evaluate both methods on 200 controlled document and memory records derived from QASPER and LongMemEval under a 1,024-token budget. NativeScope attains native-unit recall of 89.28 percent for documents and 72.50 percent for memories, improving over instance-wide Dense RAG by 42.75 and 22.00 percentage points. NS-FullQ reaches 87.78 percent and 68.50 percent; its differences from NativeScope are inconclusive, locating the primary gain in relational scoping rather than the shorter ranking query. With automatic Top-1 anchors, memory recall falls to 35.50 percent. NativeScope is therefore effective when anchor coordinates and native relations are reliable, but hard scoping inherits errors from the localization interface.
cs.CL / 61 / 2610.11196
Selective Listening: Mechanism-Guided Control of Audio Influence in Large Audio-Language Models
Yulin Sun, Kele Xu, Yong Dou
cs.SD · cs.CL · cs.LG
Abstract
Large audio-language models (LALMs) exploit multimodal evidence, yet task-irrelevant audio can alter text-reasoning decisions when listening is unnecessary. Aggregate Accuracy can hide this paired drift because audio-induced repairs and damages may cancel. Paired drift analysis and targeted interventions identify architecture-specific, intervention-sensitive late audio pathways as actionable control points. We introduce ICAP-Gate, which applies mechanism-guided, task-conditioned control to each model's pathway. Across four LALMs, two reasoning benchmarks, and environmental-sound and natural-speech interference, ICAP-Gate has lower point estimates for Influence Rate and Answer Flip than ungated inference in all 16 full-split model--condition evaluations. Fixed suppression degrades automatic speech recognition (ASR) across all four models, whereas ICAP-Gate matches ungated ASR performance by preserving the pathway for explicit audio-demand instructions. ICAP-Gate has lower paired-drift point estimates than mitigation prompting in all four evaluated settings and provides competitive stabilization relative to eight-sample Self-Consistency while using one generation per query; in controlled ARC measurements, Self-Consistency incurs $7.0$--$9.2\times$ ungated latency. These results establish selective modality influence control as a design principle for robust multimodal reasoning.
cs.CL / 62 / 2610.12201
SteerablePlex: Can We Steer Full-Duplex Models?
Haolong Zheng, Maike Züfle, Dominik Macháček, Peter Polák, Xulin Fan, Xavier Sumba, Siyin Wang, Ondřej Klejch, Mark Hasegawa-Johnson
eess.AS · cs.CL
Abstract
Full-duplex speech models can listen and speak simultaneously, enabling natural interaction, but become increasingly difficult to control as the conversation history grows. When used as user simulators, this lack of control can cause them to deviate from prescribed scenarios and produce unreliable evaluation outcomes. We introduce SimIF-Bench (Simulator Instruction-Following Benchmark), which evaluates whether a conversational model stays within a prescribed scenario and completes multiple goals in the required order. The benchmark reveals that current open-source full-duplex models struggle to follow such constraints. We then introduce a Group Reward-Decoupled Normalization Policy Optimization (GDPO)-based training recipe that enables a full-duplex model to follow textual instructions during an ongoing conversation while maintaining its turn-taking ability. By connecting the resulting SteerablePlex to an asynchronous backend language model that monitors the conversation and provides instructions when needed, we build a more controllable full-duplex user simulator that follows multi-stage constraints more reliably than existing open-source models and GPT-Realtime.
多智能体系统 (cs.MA)
7
cs.MA / 1 / 2610.10882
Decentralized collaborative continual learning: A multi-objective minimization-based technique
Yara Zgheib, Marc Antonini, Roula Nassif
cs.MA
Abstract
In this work, we formulate decentralized continual learning within a multi-objective optimization framework. For a given inference task t (corresponding to a common minimizer shared by the cost functions of all agents), agents collecting data in a distributed and streamed manner are only allowed to perform local computations and to exchange information with neighboring agents over the underlying communication graph. As tasks evolve sequentially over time, agents must adapt to newly arriving tasks while retaining knowledge acquired from previously learned ones. This requirement leads to the wellknown stability plasticity dilemma, where stability refers to the ability to retain previous knowledge, while plasticity refers to the ability to learn and adapt to new tasks. To address the stability challenge, agents store subsets of samples from past tasks in local memory buffers. Then, through an appropriate multiobjective formulation, the stored information is incorporated into the learning process so that parameter updates account jointly for the current task and previously learned tasks. The proposed decentralized continual learning approach is analyzed in the mean square error sense under general assumptions on the individual cost functions and gradient noise processes. The analysis reveals that cooperation among agents improves the performance of continual learning. In particular, by exchanging information with neighboring agents, decentralized collaborative learning can exploit the diversity of locally observed data and memory buffers to improve the network average mean-square deviation (MSD) across tasks. Finally, simulations illustrate the theoretical findings and the effectiveness of the method in reducing forgetting and improving the average MSD across tasks.
cs.MA / 2 / 2610.11375
Personalization Matters: Long-Horizon Conversation Agent with User-Centric Information in Online Shopping Interactions
Rena Gao, Yue Dai, Hao Guan, Shengxiang Gao, Wangyang Wu, Yixin Shen, Jey Han Lau
cs.MA
Abstract
Personalized conversational shopping requires maintaining preference consistency over multi-turn interactions, where users reveal constraints gradually. Existing approaches often rely on static profiles and do not explicitly control long-horizon interaction behavior. We propose a multi-agent, multimodal Retrieval-Augmented Generation (RAG) framework that decomposes dialogue state tracking, recommendation retrieval, preference-aware reasoning, and response generation, while integrating product metadata, product reviews, image-derived descriptions, and user historical reviews. To evaluate interaction-level quality, we adopt a trajectory-level protocol with four dimensions: Global Preference Consistency, Cumulative Information Synthesis, Interaction Trajectory, and Tone Consistency. On an Amazon Reviews 2023 benchmark, retrieval-enabled variants outperform a no-RAG baseline on automatic trajectory metrics (average 4.82 vs. 3.74). In a small real-user study ($n{=}5$), the Full variant achieves the highest mean overall rating (4.60 vs. 2.20 for Baseline), providing exploratory evidence that role decomposition plus user-centric retrieval improves perceived personalization.\footnote{Code and dataset are available at: https://github.com/RenaGao/Multimodel_RAG_Indexing
cs.MA / 3 / 2610.11842
ConventionPlay: Capability-Limited Training for Robust Ad-Hoc Collaboration
Abhishek Sriraman, Eleni Vasilaki, Robert Loftin
cs.MA
Abstract
Ad-hoc collaboration often requires agents to identify and adhere to some shared convention within a cooperative task. Existing work on reinforcement learning (RL) for ad-hoc collaboration focuses on training agents that adapt to the conventions established by their partners. These methods fail to consider the possibility that while some partners might follow only a single fixed convention, others may themselves be capable of adapting to multiple conventions. Here we present ConventionPlay, an RL-based approach that teaches agents to discover their partner's optimal convention by training against a learned population of partners that exhibit different degrees of adaptability across conventions. Some of these partners follow a single, fixed convention, while others are able to adapt to a subset of the possible conventions for the task in question. The existence of partners that support a limited subset of conventions forces agents trained against this population to actively probe their partner's capabilities, and steer their partner towards the most effective joint strategy that they are capable of following. Our experimental results demonstrate that agents trained via ConventionPlay achieve superior performance to existing ad-hoc collaboration methods against test populations of partners that are compatible with multiple conventions.
cs.MA / 4 / 2610.12321
Spatial Pattern Formation from Multi-Agent Learning in Public Goods Dilemmas
Yefei Zhang, Yuxuan Zhao
cs.MA · cs.GT · cs.LG
Abstract
Spatial public goods models show that prescribed movement toward richer locations can generate spatial patterns. We ask how such patterns emerge when agents learn where to move and how learning rates shape their consequences for collective welfare. Fixed populations of cooperators and defectors independently learn movement policies using tabular Q-learning and local observations. Cooperator learning generates clusters around resource peaks, while co-adaptation changes their strength and motion. At a fixed training budget, the largest welfare losses occur when cooperators learn at high rates and defectors at low rates. In part of this regime, learned policies also generate traveling bands supported by a shared directional preference. The conditions supporting travel change with further training, so these patterns reflect training history rather than an established asymptotic outcome. Across the tested learning-rate conditions with cooperator learning, mean collective welfare falls below random movement because increased crowding outweighs gains in resource benefit. Charging agents for the crowding they impose on others during learning recovers much of the welfare loss in the tested conditions. These results connect learning rates to the emergence and welfare costs of spatial organization driven by individual rewards.
cs.MA / 5 / 2610.12453
Mental-Models for Multi-Agent Systems
Hanan Gani, Lulu Shao, Manmohan Chandraker
cs.MA
Abstract
Large foundation models have accelerated progress toward general-purpose agents that interact with humans and other agents through language and multimodal signals. However, robust multi-agent decision-making requires reasoning about what other agents know, intend, and are likely to do under partial observability. Current agentic systems often operate through prompt design, memory, or end-to-end behavioral shaping, but typically do not learn an explicit partner-state representation that can be reused as a decision variable across tasks. We introduce \emph{mental-model-enabled agents}, a framework that equips an agent with a latent mental model of its counterpart, allowing it to infer hidden beliefs, intentions, and likely reactions from the observed history and use these inferences to guide action selection. Our method learns an amortized recursive Theory-of-Mind representation, with first- and second-order mental-state structure, jointly with a belief-conditioned reward model that evaluates candidate actions relative to the inferred partner state. A policy is then learned under this belief-aware signal, yielding an agent that can act independently at inference time while retaining the benefits of explicit partner modeling. We evaluate the same framework on both language-only and multimodal benchmarks. Across these settings, explicit mental-state modeling consistently improves interaction quality and Theory-of-Mind performance over base agentic systems, showing that structured partner modeling is a useful inductive bias for general multi-agent systems. Our code is publicly available at https://github.com/hananshafi/Mental-Models
cs.MA / 6 / 2610.11434
An Embodied Multiagent Framework Based on Token Communications for Cooperative ISAC
Jiahe Guo, Jun Du, Chunxiao Jiang, Jintao Wang
eess.SP · cs.MA
Abstract
The emerging low-altitude economy demands unmanned aerial vehicle (UAV)-enabled integrated sensing and communication (ISAC) for reliable connectivity and environmental awareness. In particular, embodied UAV agents offer a promising means of supporting autonomous operations through a closed loop linking perception, decision-making, and physical actions. However, each UAV has access only to local observations, and effective cooperation requires exchanging local states and intentions. Directly sharing such information can incur substantial signaling overhead and hinder timely coordination in dynamic environments. To deal with this problem, this paper investigates a cooperative ISAC network of embodied UAV agents and formulates a joint token communications (TokCom) and physical control problem to minimize total propulsion energy subject to communication and sensing rate requirements. Then, we propose a state--intent TokCom (SI-TokCom) framework driven by multi-agent embodied policy learning. Specifically, separate pretrained codebooks enable compact exchanges of local states and intentions, while UAV agents jointly learn to select and compose tokens and determine physical actions based on local observations and received tokens. Simulation results show that SI-TokCom achieves 98.9\% and 99.1\% of the centralized baseline's communication and sensing rates, respectively. Compared with the local baseline, it improves the corresponding rates by 5.0\% and 43.6\%, respectively, with essentially unchanged propulsion energy. These results highlight the potential of TokCom for communication-efficient cooperation among embodied UAV agents in ISAC systems.
cs.MA / 7 / 2610.12028
Policy Synthesis for Finite Populations of MDP Agents under Aggregate Reach-Avoid Chance Constraints
Jie Fu, Anamika Dubey
eess.SY · cs.MA
Abstract
Consider a finite population of agents with decoupled Markov transition dynamics and empirical-density feedback, subject to the following constraints: with probability at least $1-δ_r$, at least a fraction $α_r$ of agents must reach a target region at some time $t^*$, while, at each time up to $t^*$, the unsafe population fraction must remain below $β_u$ with probability at least $1-δ_u$. However, standard mean-field methods enforce these constraints only in expectation, which fails to account for stochastic fluctuations at finite fleet size $N$. To address this control problem, we propagate the second-order moment (variance) of the empirical density alongside the mean-field trajectory via a discrete-time Lyapunov recursion, and apply the Cantelli inequality to convert chance constraints into tractable deterministic conditions on the moments of the empirical density. We then incorporate these moment-based surrogate constraints into a gradient-based sequential convex approximation procedure for density-feedback policy synthesis. We further introduce additional moment-error bounds to construct a rigorous finite-$N$ certificate. The method is evaluated on a gridworld environment and a power-system EV-charging aggregation problem and compared with a standard deterministic population-level LP baseline.
软件工程 (cs.SE)
18
cs.SE / 1 / 2610.11875
NosRacer: Dynamic Detection of Race Conditions in On-Device Network Operating Systems
Runze Wu, Jingbo Zhai, Shanming Ping, Lingzhi Ouyang, Hua Duan, Qin Zou, Chengcheng Huang, Bingshe Liu, Xudong Lang, Xiaoxing Ma, Yu Huang
cs.DC · cs.SE
Abstract
Commercial on-device network operating systems (NOSes) run complex control planes in production routers and switches, where configuration update tasks are executed by multiple loosely coupled components through asynchronous message passing. Such executions are prone to race conditions: the same ordered input commands may produce different outcomes when messages are delivered in different orders. The race conditions are difficult to expose because they often manifest only as subtle, delayed malfunctions. Existing static analysis techniques lack scalability and precision for large-scale industrial NOSes, while dynamic ones incur substantial system-execution cost when attempting to cover the space of asynchronous message interleavings. To address the challenges above, we present NosRacer, a dynamic analysis framework for race condition detection in industrial-grade on-device NOSes. NosRacer uses a two-phase design. The concentration phase reduces analysis scope by leveraging the locality of configuration update tasks. Race condition detection is then limited to a small subset of involved components and crucial state variables, thereby reducing the detection cost. The perturbation phase proactively perturbs message asynchrony to increase the likelihood of executions with race-condition-exposing message interleavings. It keeps perturbation practical by grouping compatible perturbations for parallel execution and adaptively strengthening perturbations. We implement NosRacer in a commercial, actively developed NOS. NosRacer is integrated into the existing testing factory and used to detect race conditions in 5 key control-plane components. NosRacer achieves 66% precision and detects 21 race conditions confirmed by developers as severe bugs, while keeping the detection overhead below 10%.
cs.SE / 2 / 2610.10961
Cross-Provider Review as a Runtime Contract for Coding Agents: A Controlled Pilot and Fault-Injection Study
Bowen Xu, Boyu Chen
cs.SE · cs.AI
Abstract
Coding agents increasingly share a workstation while drawing on separate providers and subscription allowances. A second agent can inspect a completed answer, but the call spends another pool and may provide no substantive finding. We describe an advisory cross-provider review contract: distinct resource pools, bounded execution, restricted reviewer capabilities, complete input delivery, usable semantic output, explicit failure states and durable per-attempt evidence. In a controlled, agent-authored pilot of 20 paired development turns, eight had a material reviewer finding (95% exact interval 19.1-63.9%). A boundary-condition scan across both reviewer backends reproduced a previously discovered false success on partial input: four truncation levels passed historically and failed after repair. The scan also found and repaired cancellation during process reaping. In real CLI probes, Claude had no writing tools; Codex attempted writes in five of five read-only trials, each write tool failed, and no disposable repository changed. These tests cover specified paths and versions, not field reliability. A preregistered shadow study of metadata-only review allocation accrued 25 formal observations before an exact-runtime regression found a third defect: a reviewer exiting nonzero with a well-formed verdict was counted as complete. Exit status was not recorded per attempt, so exposure cannot be resolved retrospectively. The 25 formal and two pending records remain an audit cohort; the measurement-valid cohort restarted at zero and collection has begun. No gate result is reported.
cs.SE / 3 / 2610.10978
Probabilistic Sensing, Deterministic Authority: Admitting Model-Produced Observations into Sufficiency-Checked Governance Contracts
Gaston Besanson
cs.SE · cs.AI
Abstract
When a field that an authority contract needs exists only in unstructured evidence, a model can sense it. We admit the model's output only as an observation record with a score. An admission policy, with thresholds fitted on a held-out split at a declared false-positive ceiling, maps each score to true, false or unknown. Unknown denies. A deterministic, sufficiency-checked contract decides. The probability that sensing changes the verdict is bounded by the sum, over the contract's sensed fields, of the admitted-wrong and unknown rates. This is an instantiation of union-bound reasoning, indexed by the contract. Minimising the estimated bound is a valid cost model for choosing among sufficient contracts. In a registered study on two constructed domains with two sensor families (36,000 model calls), no cell refuted the bound. Deny-to-allow changes from sensing appeared for the first time in this programme: 13 of 21,000 test verdicts, all from 3 contradictory records; each flip in a cell with a registered bound lay under it. Sensing-aware selection picked the lower-exposure contract in 4 of 4 registered tests. Both sensors' scores were informative but not calibrated. Correctness is relative to the declared loss model, candidate representation and reachable states; all domains are constructed.
cs.SE / 4 / 2610.11041
Following Breadcrumbs in Code: What Accidentally Committed Ad-Hoc Logs Reveal about Developer Comprehension
Yi-Hung Chou, Boyuan Jiang, Vidit Jain, Yiyang Min, April Yi Wang, James A. Jones
cs.SE
Abstract
Developers frequently insert temporary print or log statements, known as **ad-hoc logs**, to better understand program behavior at runtime, particularly when facing unexpected issues or complex control flows. Despite being a nearly universal practice, systematic study has been limited because these logs are ephemeral: they usually remain only in local environments and are removed before code is committed, making them difficult to capture. In this work, we addressed this challenge by mining accidental commits where developers unintentionally left ad-hoc logs and later deleted them, and by analyzing live-streamed programming sessions to observe their use in practice. Using these methods, we constructed a large dataset across three major programming languages (Java, JavaScript, and Python), enabling the large-scale investigation of logging practices. Our analysis reveals both common and language-specific patterns in where and how developers rely on ad-hoc logs. Across languages, ad-hoc logs tend to appear in program regions that are harder to reason about at runtime. We also identify distinctive language-level patterns, such as frequent use in asynchronous and callback functions in JavaScript and in thread-related classes in Java. In addition, functions containing ad-hoc logs generally have higher cyclomatic complexity than the overall function population. Production logs show a similar association, consistent with logging serving as a means of observing runtime behavior in structurally complex functions. Together, these findings provide empirical insight into developers' runtime comprehension practices and offer a valuable dataset for researchers and tool builders seeking to better support debugging and logging.
cs.SE / 5 / 2610.11073
IRONPROOF: COBOL-to-Python Transpilation with SMT-Based Equivalence Checking
Dominik Blain
cs.SE
Abstract
Translating COBOL to a modern language can change what a program computes, and LLM translations carry no guarantee of equivalence. We present IRONPROOF, which parses COBOL into an intermediate representation, generates Python, encodes both as Z3 formulas over shared inputs, and emits either a machine-checkable equivalence certificate (UNSAT) or a counterexample (SAT). On 2,345 COBOL files (GnuCOBOL tests, NIST CCVS85, open-source collections, and programs we generated or wrote), 782 enter the checking path: 606 (77.5%) are proved equivalent, 101 are partially verified, none is refuted, and 75 fail inside our pipeline and stay in the denominator. On independently authored programs the rate is 52.6% (153 of 291), against 92.3% on programs we wrote. Of the 153 independent proofs, 37 cover a single execution of a program that reads input the encoder does not model, and across all 606 proofs only 14 quantify over an input that a proved output depends on. On two public business-application corpora (AWS CardDemo and IBM GenApp), the encoder models no program end to end. An April 2026 LLM-only baseline failed on 54 of 94 independently authored programs, counting translations our encoder could not verify. PIC-bounded overflow detection flags an out-of-range output in 22.1% of the programs it can analyze, an upper bound. An earlier confrontation with GnuCOBOL 3.2.0 found 24 of 49 executable independent proofs disagreeing with the runtime on a variable in the proof's domain. A proof establishes that the generated Python computes what our intermediate representation says the COBOL computes, not equivalence to a COBOL implementation.
cs.SE / 6 / 2610.11169
Skill Constellations: Tracing the Supply Chain of Agent Skills on GitHub
Fahd Seddik
cs.SE · cs.CR · cs.SI
Abstract
Agent skills are SKILL.md instructions and scripts that AI coding agents such as Claude Code and Codex run with the permissions of their user. Developers share skills by copying them between repositories, which makes them a software supply chain without a registry, versions or provenance. The origin of a copied skill, the reach of a security fix and the repositories that warrant review are therefore unknown. Studies that record which repositories hold a skill at a single point in time cannot reveal who copied it from whom. We contribute the first dated copy network of agent skills, built from the git history of every SKILL.md in GitSkills and covering 2,193,119 skill adoptions across GitHub, together with an interactive viewer. A few repositories are the source of almost all copies, and GitHub stars do not identify them. Skill copies almost never change with their source, and a fix at the source therefore rarely reaches them. We fit a model of which repositories others copy from and use it to rank repositories for audit. Reviewing the 100 repositories it ranks highest prevents 14.9% of later adoptions of high-risk skills, against 0.5% for the 100 most starred, which gives security engineers a short list to check before a skill spreads. Platforms should therefore distribute versioned references rather than copies. Project Website: https://fahdseddik.github.io/Skill-Constellations/
cs.SE / 7 / 2610.11179
Who Pays the Review Cost? Triage, Fairness, and Accountability in AI-authored Pull Requests
Md Shamimur Rahman, Khairul Alam, Banani Roy, Chanchal K. Roy
cs.SE
Abstract
AI coding agents are moving from local code assistance into pull-based workflows, where generated contributions must be reviewed, explained, and maintained within existing project norms. Although recent work has begun to characterize AI-authored pull requests (AIPRs), less is known about how reviewers govern their entry into review, how AI authorship reshapes credibility and fairness, and what intake mechanisms protect review sustainability. We report a mixed-method questionnaire survey of 239 practitioners from 31 countries with code-review experience and varying exposure to AIPRs. In the scenarios and self-reports elicited by the survey, AI authorship was not a categorical rejection signal. Instead, respondents described review effort as conditional on whether an AIPR arrived as an accountable contribution, with bounded scope, project-grounded rationale, validation beyond Continuous Integration (CI), contributor responsiveness, and identifiable post-merge ownership. This conditional logic extended to newcomer AIPRs, where respondents emphasized visible participation in the current review process over profile-level reputation alone. Qualitative responses further described shifts in mentoring, scrutiny, deferral, and routing when human stewardship was difficult to observe. We conceptualize missing rationale, validation, and ownership as reviewability debt, the work reviewers must absorb when generated code lacks sufficient human grounding. These findings reframe AIPR governance around making human judgment observable before generated contributions consume scarce reviewer attention.
cs.SE / 8 / 2610.11514
SSCBench: Evaluating the Evidential Validity of Fault-Injection Tests for Tool-Using LLM Agents
Xincheng He, Wanli Dong, Zhaoqiang Guo, Yan Liu, Lei Xu
cs.SE
Abstract
Fault injection is increasingly used to evaluate the reliability of tool-using LLM agents. However, there has been limited study of how fault-adoption results should be interpreted when the agent itself determines which authoritative observations become visible during execution. In this paper, we present a systematic study of this evidential validity problem in agent fault-injection evaluation. We develop a measurement protocol that specifies what observations can refute an injected assertion, determines whether they can become visible before the affected fact is first used, and records whether the evaluated execution actually realizes this condition. We construct SSCBench as an instantiation of the protocol and evaluate four fault operators and five agent configurations over 1,191 faulted executions in two $τ$-bench environments. Our experiments show that an admitted fault case and agent configuration can realize substantially different evidential conditions across executions, and that aggregate adoption can remain well defined even when the population supporting a timely-counterevidence claim is sparse or absent. For example, among 44 adopted runs in which counterevidence eventually became visible, only 17 received it before first use, while 27 received it afterward. We also find that first-error timing and later stance revision need not coincide, and that automated trajectory analysis can recover adoption without reliably recovering the first faulty-reliance event needed for temporal diagnosis. We argue that the evidential condition realized by an execution and the population supporting a claim-specific interpretation are part of fault-injection evaluation itself and should be reported before adoption is interpreted as failure under pre-use counterevidence.
cs.SE / 9 / 2610.11593
Runnable Commit Untangling for Coding Agents
Jinfeng Jiang, Dongsun Kim, Dayi Lin, Zhou Yang
cs.SE · cs.AI
Abstract
Coding agents produce large, tangled patches that mix multiple development purposes, making the code hard to review and maintain. Commit untangling offers the promise of organizing such large patches into untangled, manageable commits. This paper emphasizes two important limitations in existing commit untangling studies. First, they do not consider that untangled commits are ordered and should leave the code runnable. In practice, maintainers are unlikely to accept commits that prevent the code from running. Second, existing studies claim that commit untangling helps software maintenance. However, they conduct syntactic comparisons between the untangled commits and developers' original commits without directly showing the claimed maintenance benefits. To address these gaps, this paper makes two novel contributions: (1) RucTangle, the first agentic method that untangles commits while keeping the code runnable after each commit; and (2) TangleEval, the first evaluation framework that quantifies how untangled, manageable commit histories help coding agents repair bugs. We compare RucTangle against four untangling methods on 131 agent-generated patches. All histories produced by RucTangle are runnable, while baselines produce 20.6%-37.4% unrunnable commit histories. We further collect 453 agent-generated patches that introduce regressions (i.e., causing previously passing tests to fail) and ask two other coding agents to repair regressions. Augmenting agent context with RucTangle-produced histories yields 5.2% absolute improvement in pass@1. We also analyze agent trajectories to learn how they use untangled commits to navigate and fix bugs. Our findings demonstrate the value of adopting established software engineering practices in the era of coding agents, which broaden the future research agenda: how can agents actively use software history to make better development decisions?
cs.SE / 10 / 2610.11647
One Skill Too Many: How Co-Installed Skills Conflict in Coding Agents
Chaoliang Yan, Zihao Xu, Yuekang Li, Shangzhi Xu, Yi Liu, Gelei Deng, Siqi Ma
cs.SE · cs.AI
Abstract
Coding agents are extended with agent skills, directories whose SKILL.md tells the model when and how to perform a task. Because skills come from independent sources (teams, developers, plugins, copied collections), an installed skill can be co-installed with a similar skill doing the same job, and the model picks between them by name and description alone. In a conflict, the installed skill loses core functions (e.g., a ban on touching git) because the similar skill runs instead or changes what it does. The task still passes, so benchmarks that check only task completion miss such cases. We present the first empirical study of such conflicts. From snapshots of 20,947 repositories, we mine 822,109 candidate similar-skill pairs, have an LLM judge a stratified sample of 3,754, and run 312 confirmed pairs on three models (6,368 runs, 169,294 tool calls, 542 agent-hours). We report five findings. (1) Conflict-prone skills are common: nearly one in four installed skills is co-installed with one that does the same job, and 37% of judged skills sit inside copied collections. (2) Most such pairs involve normative skills, then capability skills. (3) Without lowering task completion, a similar skill takes one in five runs from the installed skill, and runs that open the similar skill first lose over a third of the exclusive core functions that only the installed skill fulfills. (4) Install location decides which skill runs, listing order barely matters, and the final reply names the skill used in only 0.9% of substituted runs. (5) Conflicts are decided at the first skill read, almost always before any file is changed, and a pre-tool hook at that read restores fidelity on exclusive core functions to the level of runs that open the installed skill first. Benchmarks should thus score exclusive core functions, and platforms should guard the first read and show which skill ran.
cs.SE / 11 / 2610.11725
Implementing the Spec Growth Engine: Preventing Spec-Code Divergence, and Growing the Spec with Agents
Hartwig Grabowski
cs.SE
Abstract
The Spec Growth Engine anchors AI-assisted software development in a graph of specifications that the code is coupled to. This paper describes its implementation, which serves two tasks and keeps them apart as two layers. The first layer prevents spec-code divergence: a deterministic engine validates the spec graph, compares it with the code's import graph, earns a node's verified status from recorded test evidence, and classifies every change by what it can break -- without calling a model. The second layer grows the spec with agents: an intent author, a planner and a coder, each played by its own model, extend the graph in rounds, and a deterministic rule decides after each round whether the run goes on. How much of the human's judgement is delegated is set by three independent switches -- a draft gate, a delegation for breaking changes, and the run mode -- which, with two ways of laying a project's floor, give eighteen ways to run a project. We describe each of them, the gates, requests and waivers through which agents and the human communicate, and the spectrum of operation from entirely manual work to an unsupervised run whose decisions the human reviews afterwards. Throughout, one claim holds the design together: an autonomous run is worth only as much as the deterministic instance that measures it.
cs.SE / 12 / 2610.11835
On the Risks of using LLM-Generated Tests for Regression Testing
Mohammadali Charoosaei, Cedric Richter, Mike Papadakis
cs.SE
Abstract
Software is under constant evolution: developers continuously add features, fix bugs, and refactor code, and any of these changes may break existing functionality. Regression testing guards against such effects by capturing expected behavior in test cases. LLM-based test generation aims to automate this process by generating regression tests directly from the code under test. This is beneficial when the implementation is correct, but problematic when the code contains faults: the generated tests may then encode and preserve incorrect behavior. To investigate this risk, we apply LLM-based regression test generation to pull requests merged into the main branch of software projects and study the impact of the generated tests on subsequent project evolution. We distinguish between fault-revealing tests, which assert correctly implemented behavior, and fault-enforcing tests, which assert faulty behavior. Across 145 pull requests from SciPy, Qiskit, and pandas, 8%-17% of the generated tests are fault-enforcing, while only 2.4%-4.8% reveal faults. Fault-enforcing tests persist over time: after several subsequent commits, 83%-91% of them are still relevant and pass. They also accumulate: when the faults of all pull requests are combined in one codebase, 83%-92% remain enforced at the end of the commit history, and the developer-written test suite detects only 14%-30% of them. Our results reveal a fundamental risk of LLM-generated regression tests: without manual validation, they may encode faulty behavior as expected behavior, allowing bugs to persist across software revisions and largely evade developer-maintained test suites. LLM-based regression testing can thus give rise to a new form of technical debt.
cs.SE / 13 / 2610.11858
Trajectory-Guided Fault Localization for Agent Skill Evolution
Yu Ge, Linna Xie, Zhong Li, Yu Pei, Tian Zhang
cs.SE · cs.AI
Abstract
Agent skills provide reusable guidance for code agents, but incomplete or unsuitable guidance can impair task execution. To reduce the manual effort of skill refinement, recent approaches use LLMs to generate revisions from execution feedback. However, grounding these revisions in explicit behavioral evidence remains challenging. To address this gap, we propose SkillMorph, a skill-evolution approach based on trajectory-guided fault localization in agent skills. Its core idea is to link execution evidence to specific skill contents before generating revisions. Specifically, SkillMorph compares failure and success evidence in abstracted trajectories across repeated runs and tasks, incorporating changes between evolution loops to identify suspicious actions. It then uses these suspicious actions to localize edit sites in the skills and generate corresponding revisions. Experiments on SWE-Skills-Bench and CannBot show that the skills evolved by SkillMorph consistently achieve higher trial-level accuracy and execution consistency than the original skills and those from four existing skill-evolution methods. We have also applied SkillMorph to automated kernel generation with an AI operator-development team, which has accepted 6 skill-revision pull requests.
cs.SE / 14 / 2610.11889
Evaluating Exact Output and Checkpoint-State Prediction in Real Programs
Xiaohong Chen, David Bucur, Chenglong Ma, Yi Zhang, Lingming Zhang, Sriram Vishwanath, Grigore Rosu
cs.SE · cs.AI
Abstract
We present a benchmark for predicting final output and checkpoint state from source and input alone. It extends CRUXEval-style output prediction with paired shorter- and longer-trace inputs and checkpoints inside and after a loop. The benchmark contains 400 cases from 371 Python and C++ programs, evaluated under seven settings from four model families without tools or code execution. Of 11,200 planned predictions, 11,151 produced gradable responses. Reasoning-enabled settings outperform their off counterparts by 33.1 to 55.2 percentage points on completed responses. The strongest setting scores 93.0% on shorter-trace final output, 77.0% on longer-trace final output, and 65.5% and 63.5% on the two state tasks; these scores also hold when missing responses count as wrong. Across 2,397 matched Python comparisons with identical source, changing to the longer-trace input yields 528 correct-to-wrong changes and 147 reversals. Source-clustered analyses preserve this accuracy gap, while adjusted Python models give no evidence of a positive incremental association between cumulative state load and error. Changed inputs and checkpoint tasks alter several factors together, so the gaps do not isolate trace length or an internal state-tracking mechanism. The benchmark exposes errors hidden by short-output scores alone.
cs.SE / 15 / 2610.11994
Neural Network Verification for Deep Joint Source-Channel Coding
Thanh Le, Hai Duong, Takeshi Matsumura, ThanhVu Nguyen
cs.SE · cs.AI
Abstract
Deep joint source-channel coding (DeepJSCC) transmits data end-to-end over wireless channels using a neural encoder-decoder, but reconstruction quality can degrade sharply under adversarial perturbations and channel disturbances; no method formally bounds this degradation for DeepJSCC. We present the first bound-propagation framework for verifying DeepJSCC's decoder, bounding worst-case reconstruction error over a given wireless channel's noise region. Current deep neural network (DNN) verifiers do not support three DeepJSCC decoder components: parametric rectified linear activations (PReLU), transposed convolutions, and Rayleigh fading. We extend state-of-the-art techniques for optimization of linear relaxation in DNN verification for PReLU, replace the transposed convolution with its restricted upsample-then-convolution form, and formulate Rayleigh fading as a structural perturbation prepended directly into the decoder, thereby reducing the dimensionality of the verification problem. We also instantiate Lipschitz-regularized global robustness training, denoted GloRo, improving global robustness and enabling tight certification of DeepJSCC models for the first time. On DeepJSCC model for image transmission, this global robustness training procedure combined with structural encoding lowers the median certified bound by up to 41% and certifies about ten times more safe cases (192 against 19) than GloRo with interval encoding at a 10-degree error in channel estimation. Over-the-air validation with an orthogonal frequency-division multiplexing (OFDM) implementation on software-defined radio devices confirm the certificate holds on real hardware, with a worst observed error on radio link at 0.082 against a certified bound of 0.128.
cs.SE / 16 / 2610.12149
Reliability Characterization for N-version Object Detection
Shunsuke Nagao, Fumio Machida
cs.SE
Abstract
N-version object detection (OD) is an approach to diversifying detection results using multiple models or input frames and reducing detection errors by aggregating individual results. Diversity and consistency across multiple detection results are critical information for characterizing the reliability of possible configurations of N-version OD systems. However, existing performance metrics such as mAP and Accuracy fail to capture these factors, as they are defined solely on the final outcome after aggregation. To overcome this limitation, we propose two reliability metrics particularly defined for N-version OD, namely coverage of errors in OD (Cov_OD) and certainty of accurate prediction in OD (Cer_OD), which can be computed from individual detection results without relying on voting strategies. We empirically demonstrate the unique features of the proposed metrics through a case study of N-version OD for a vehicle in an autonomous-driving simulator. We show that the proposed metrics can be used for 1) selecting version combinations with complementary error characteristics, 2) choosing an effective voting strategy based on diversity and consistency profiles, and 3) guiding incremental construction of N-version OD systems via pairwise two-version analysis. These results highlight the importance of the metrics that can guide the design of reliable N-version OD applications.
cs.SE / 17 / 2610.12269
Cadence: Strategic Guidance for Coding Agents
Minxing Wang, He Ye, Earl T. Barr, Yintong Huo
cs.SE
Abstract
Runtime monitors are increasingly used to improve the reliability of LLM-based coding agents by inspecting execution trajectories and delivering corrective guidance upon detecting misbehavior. However, their effectiveness remains limited by static guidance triggering schemes. Existing monitors rely either on fixed inspection intervals, missing timely guidance during severe misbehaviors while incurring unnecessary overhead during healthy execution, or on rigid heuristic rules, failing to detect complex reasoning errors. To address these limitations, we propose Cadence, a dynamic monitoring framework that adaptively schedules inspections and delivers guidance according to the agent's real-time execution health. Cadence consists of two core modules: a two-tier intervention module and an inspection scheduler. The intervention module delivers advisory-level guidance for normal executions and minor lapses, while providing replacement-level guidance for severe misbehaviors. Driven by the intervention level, the scheduler adjusts the inspection frequency by tightening supervision after replacement-level guidance and relaxing it after advisory-level guidance. Evaluated on 300 SWE-bench Lite tasks across two distinct agents, mini-swe-agent and Moatless, Cadence achieves the highest resolve rate among all evaluated monitors. Specifically, Cadence outperforms vanilla agents by 25.33\% (+76 resolved tasks) on mini-swe-agent and 15.67\% (+47 resolved tasks) on Moatless, while maintaining competitive token efficiency compared to state-of-the-art baselines.
cs.SE / 18 / 2610.11195
Retromorphic Testing of Quantum Compiler Passes
Mushahid Khan, Olivia Di Matteo
quant-ph · cs.SE
Abstract
Quantum compilers play a critical role in transforming high-level quantum programs into optimized, hardware-compatible circuits. However, verifying the correctness of compiler passes remains challenging, as determining the expected output of large, deeply entangled quantum circuits is computationally intractable. This challenge is further amplified when compiler passes modify already complex circuit structures, making manual validation of transformed circuits impractical. In this work, we perform a systematic analysis of unit tests for quantum compiler passes in four quantum programming frameworks (PennyLane, Qiskit, Cirq, and pytket). Our findings indicate validation is dominated by program-content and program-metric assertions, and test circuits are generally small and shallow. Motivated by these observations, we introduce a testing methodology for automated validation of quantum compiler passes based on retromorphic testing and principles from the Hadamard test. This methodology analyzes a compiler pass, test circuit, and expected pass behavior to verify semantic preservation and intended structural modifications. We implement our methods in a framework, RetroQ, and apply it to compiler passes in PennyLane and Qiskit. Experimental evaluation reproduced several existing bugs as well as uncovered previously undetected defects, such as flawed symbolic parameter handling, incorrect commutation logic, failure to recognize self-adjointness of gates, and runtime crashes. These findings highlight the need for compiler-pass-specific testing methodologies to improve the reliability of the evolving quantum software stack.
操作系统 (cs.OS)
1
cs.OS / 1 / 2610.12164
Mole: Tier-Specific Hotness Profiling Driven Memory Tiering for Multi-Tiered Memory Systems
Mingyang Liu, Congming Gao, Xufeng Yang, Fang Wu, Youmin Chen, Renhui Chen, Jiwu Shu
cs.OS
Abstract
Multi-tiered memory systems combine fast, small upper tiers with slow, large lower tiers to improve performance and cost efficiency. Existing designs, such as AutoTiering and MTM, rely on greedy promotion and stepwise demotion, which can intensify contention for limited capacity in faster tiers. We observe that promotions directly improve performance, whereas demotions primarily reclaim space. Based on this asymmetry, we propose Mole, a memory tiering system that separates promotion and demotion destinations. Mole employs staging demotion to move cold pages directly to the lowest tier, bypassing intermediate tiers, and targeted promotion to place hot pages in tiers that match their current hotness. The lowest tier serves as a staging area from which pages can be promoted when they become hot again. This separation preserves intermediate-tier capacity for promotions, reducing tier contention and migration failures. However, staging demotion requires timely identification of pages that become hot after demotion. To meet this requirement with low profiling overhead, Mole uses tier-specific profiling: it profiles the lowest tier at high frequency to detect reactivated hot pages, while profiling upper tiers at lower frequency to identify cold pages. Experimental results show that Mole reduces migration failures and improves performance under dynamic access patterns.
硬件架构 (cs.AR)
3
cs.AR / 1 / 2610.10924
The Missing Fourth Term for the Emulation Tensor Memory Equilibrium (TME) Model: The Residue Deconstruction Cost
Harun Bayraktar, John Gunnels, Peter Caday
cs.AR · cs.AI · cs.PF
Abstract
The Tensor-Memory Equilibrium (TME) model of "FP8 is All You Need (Part 1)" calculates the execution time of Ozaki Scheme II emulation of fp64 as the maximum of a tensor-core term and a High-Bandwidth Memory (HBM) traffic term, plus a per-output reconstruction term. However, it omits the per-input deconstruction cost: every streamed fp64 operand must be scaled, rounded, and reduced modulo each of the $r$ moduli on SIMT pipes before any matrix multiply can issue. In this note we add this fourth term, calibrate its constant from the cuBLAS emulation path, and derive a closed-form operational-intensity threshold $\mathrm{OI}^{*} = c_q r P_{\mathrm{fp64}}/(8P_{\mathrm{int}})$ below which emulation cannot match native fp64 regardless of tensor-core throughput. On the NVIDIA B300 GPU the threshold is $\mathrm{OI}^{*}\approx 0.56$ FLOP/B. As a result, GEMV, SpMV, and low-batch GEMV, which are the memory-bound kernels the original paper claims to accelerate, are limited to 0.3-0.9x of native performance, and the 7-point stencil to 1.8x rather than the claimed 3.1x. Dense GEMM is unaffected as expected. We also show that precomputing and storing the residues moves the same cost into the bandwidth term, and we state the instruction count that an implementation would have to achieve to invalidate the bound.
cs.AR / 2 / 2610.10985
ATLAS: Adaptive TDA-guided Landscape-Aware Transistor Sizing
Youngmin Oh, Jihwan Won, Yuntae Park, Bosun Hwang, Suwan Kim
cs.AR
Abstract
Analog transistor sizing, finding design parameters that simultaneously satisfy multiple performance specifications, is a labor-intensive bottleneck in circuit design. To support analog circuit experts, various automation methods have been proposed, including Bayesian optimization (BO), reinforcement learning (RL), and others. Yet existing methods are oblivious to the topological structure of a feasible design space, which can fragment into disconnected regions due to operating-regime transitions, conflicting specification trade-offs, and nonconvex device physics. This topological blindness causes the optimizer to converge within a single feasible region while missing others that may contain superior designs. To address this limitation, we propose ATLAS. a BO framework utilizing Topological Data Analysis (TDA). At each iteration, a Mapper graph is constructed over a surrogate-predicted feasible region to estimate connected regions, enabling topology-aware exploration from the very first iteration without any observed feasible points. A topological sensitivity score classifies candidates as bridge, frontier, or interior points, injecting a targeted exploration bonus into the acquisition function. Experiments on four analog circuit benchmarks in the GF180 and SKY130 processes demonstrate that \coin finds feasible designs with significantly fewer simulations than baselines, including RL and BO methods. To the best of our knowledge, this is the first work to apply topological data analysis to analog circuit design automation. The official implementation is publicly available on https://github.com/youngmin0oh/atlas.
cs.AR / 3 / 2610.11748
DEX: Digit-Level Early Exit for Energy-Efficient MSDF Neural Network Inference
Yousef Sadegheih, Dorit Merhof, Muhammad Usman
cs.AR · cs.AI · cs.LG
Abstract
U-Net inference for brain-tumor segmentation requires billions of multiply-accumulate operations, motivating hardware that can reduce computation dynamically rather than relying only on fixed precision or static model compression. Most-significant-digit-first (MSDF) arithmetic exposes the leading digits of a result during computation, enabling output-dependent decisions before the full value is generated. This paper presents an MSDF accelerator for quantized U-Net segmentation with a two-stage grouped processing element supporting signed INT8 operands and in-stream bias accumulation. Four runtime mechanisms operate directly on the output digit stream: exact early negative detection (END) in ReLU layers, exact sign-only decision making in the segmentation head, calibrated low-order-digit skipping, and calibrated pruning. The two approximate mechanisms are selected offline under an accuracy constraint, while execution requires only lightweight control and does not modify the stored weights. On a residual U-Net trained with nnU-Net for BraTS, the proposed mechanisms reduce digit cycles by 38.38\% while achieving a mean Dice score of 80.58\% on 73 held-out cases, compared with 81.20\% for the floating-point model; the exact mechanisms alone reduce cycles by 18.79\% without altering the quantized output. Synthesized in 45~nm, the processing element operates at 500~MHz, occupies 0.858~mm$^2$, and consumes 0.726~mJ per $192\times192$ patch under switching-activity-annotated power analysis. A projected eight-output accelerator with shared activation delivery achieves 16.6~ms latency and 1.67~mJ per patch.
密码学与安全 (cs.CR)
35
cs.CR / 1 / 2610.10735
DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits
Qiaolin Qin, Wanpeng Li, Benoit Baudry, Lorenzo De Carli, Heng Li, Ettore Merlo
cs.CR · cs.SE
Abstract
Pre-trained models (PTMs) are widely distributed as serialized binaries, but their reuse often exposes software supply chains to deserialization attacks. Despite the emergence of safer serialization formats, the unsafe Pickle format remains prevalent: our analysis of over 10,000 popular Hugging Face repositories reveals that 9.3% rely on Pickle. While many defense mechanisms have been proposed, state-of-the-art model scanners suffer from a coverage-precision gap, missing security-sensitive behaviors and generating excessive false alerts. In this paper, we introduce DITTO, the first stack-based, context-aware scanner for Pickle-based PTMs. DITTO faithfully tracks Pickle virtual machine state transitions and performs context-aware semantic analysis to infer model intentions. We also present PickleBench, a benchmark of 959 benign and 92 malicious real-world models, including extension registry attacks previously missed by existing tools. Across multiple evaluations, DITTO achieves 100% scanning coverage, a 0% false-negative rate, and a 0.7% false-positive rate, yielding an F1 score of 0.966, significantly outperforming state-of-the-art scanners. By minimizing false alerts while preserving detection accuracy, DITTO generates actionable security reports with contextual evidence, enabling safe PTM reuse and strengthening software supply chain integrity.
cs.CR / 2 / 2610.10766
CPU-Auth: Device Fingerprinting for Authentication via DVFS Side-Channel
Ryan Swift
cs.CR · cs.LG
Abstract
Lack of effective authentication has resulted in numerous security and privacy breaches, including unauthorized access to protected information, identity theft, and fraud. One approach to mitigating such attacks is Multi-Factor Authentication (MFA), in which users must provide multiple pieces of information for authentication. Some secondary authentication factors include SMS text verification codes, biometrics, and tokens. Each contains at least one notable flaw: SMS is notoriously insecure; biometrics rely upon access to sensitive personal data; and tokens require dependence on third-party providers (e.g. OAuth providers). This work explores CPU-Auth, a novel authentication mechanism based on unique variations in the physical characteristics of the CPU of a computing device. By measuring the behavior of the Dynamic Voltage and Frequency Scaling (DVFS) governor remotely from within a browser, unique properties of the CPU can be leveraged to establish a hardware-based device fingerprint for use in CPU-Auth. The performance of CPU-Auth is evaluated on over 50,000 data traces using distance-based and deep learning methods. CPU-Auth is part of a larger research project, and the results provided in this report reflect only the contributions made to the project by members of this group.
cs.CR / 3 / 2610.10844
When Flaws Cascade: Understanding Vulnerabilities and Exploitation Chains in JavaScript Engines
Yuhan Ma, Jiongchi Yu, Xiaofei Xie, Qiang Hu, Zhiyi Zhang, Junjie Wang
cs.CR · cs.SE
Abstract
JavaScript engines are pivotal to modern web browsers, enabling the execution of dynamic and interactive web applications. However, their complexity and widespread adoption make them prime targets for attackers exploiting vulnerabilities. While existing research has focused on detecting vulnerabilities of JavaScript engines, a significant gap remains in systematically understanding the characteristics of these vulnerabilities, including their symptoms, root causes, and exploitability. This paper bridges this gap by presenting the first comprehensive empirical study on vulnerabilities in JavaScript engines, investigating their characteristics and potential exploitation strategies. We construct a dataset comprising 241 vulnerabilities across four mainstream JavaScript engines from 2017 to 2024. Through in-depth analysis, we first develop taxonomies for symptoms and root causes. Building on this understanding, we investigate the exploitability of these vulnerabilities, identifying key prerequisites and extracting vulnerability trigger chains that demonstrate how logical errors propagate into memory safety violations. Additionally, we analyze the mitigation strategies to counter these exploits. Finally, we summarize key implications for various stakeholders, including developers and researchers, offering actionable insights to improve the security and resilience of JavaScript engines.
cs.CR / 4 / 2610.10909
Power Side-Channel Membership Inference Attack on Embedded Machine Learning
Sahan Sanjaya, Prabhat Mishra
cs.CR · cs.LG
Abstract
Membership inference attacks (MIAs) threaten the privacy of machine learning (ML) training data by determining whether a sample was used to train a target model. Existing MIAs rely on model outputs, ranging from prediction probabilities to predicted labels, an assumption that can be restrictive for on-device ML systems with limited or inaccessible outputs. However, suppressing model outputs does not eliminate the data-dependent computations that produce them, which may remain observable through physical side channels. We present PSCMIA, a power side-channel membership inference attack against embedded ML models that can infer membership directly from power traces without requiring prediction probabilities or even the predicted labels. We evaluate PSCMIA across multiple datasets (MNIST, FMNIST, CIFAR10, CINIC10), fully connected (FC) and convolutional neural network (CNN) architectures, and two embedded platforms (STM32F3, XMEGA). PSCMIA achieves ROC-AUC values of up to 0.907 on FC models. For CNN models, the ROC-AUC gap between PSCMIA and probability vector-based shadow MIA ranges from 0.006 to 0.116. Across the FC and CNN evaluations, PSCMIA outperforms label-only MIA in 11 of 16 model-dataset-hardware configurations, demonstrating that physical execution can expose membership information even when conventional model outputs are unavailable through unintended power side-channel leakage.
cs.CR / 5 / 2610.10992
The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
Dominik Blain
cs.CR
Abstract
Every ML-DSA (FIPS 204) signature carries a public hint vector h. We find that the Hamming weight of each hint polynomial h_k depends on the signing key: to first order it measures the Euclidean norm of (t0)_k, the low-order part of t that key generation leaves out of the public key. A closed-form model predicts the per-key mean weight with Pearson r between 0.95 and 0.98; once the norm is accounted for, we detect no key-dependent signal above sampling noise. We measure the effect on the reference C implementation, with 200 keys and 2,000 signatures per key for each parameter set. A one-way ANOVA rejects key-independence of the total hint weight for ML-DSA-44, ML-DSA-65 and ML-DSA-87 (F = 30.1, 10.3, 11.1). The effect is weak: the key explains 0.5% to 1.5% of the variance of the total weight. A single signature identifies its key among 200 with 1.1 to 1.3 times chance accuracy from the total weight, and 1.5 to 1.8 times from the per-polynomial weight vector. A one-sided test at significance 0.001 separates two typical keys with probability one half after about 1,300 to 3,700 signatures per key. The hint weight does not endanger the signing key. The Dilithium designers do not treat t0 as secret, and t0 is known to be recoverable from signatures. The hint weight is, however, a weak statistical fingerprint: with enough signatures, it can be used to test whether they come from a common key, without the public key. Bounded-weight signing, a backward-compatible filter whose effect on the security argument we did not analyze, removes 86-91% of the between-key spread of the total weight but leaves the per-polynomial channel largely intact.
cs.CR / 6 / 2610.11030
NOMOS: Compiling Written Policies into Statically Verified Tool-Call Gates for LLM Agents
Min-Young Yu, Tony Kim, Jang Won Choi
cs.CR · cs.AI · cs.LG
Abstract
Tool-using LLM agents violate the policies they are deployed to enforce, often silently. Prior defenses hand-write rules, query an LLM verifier per action, or compile policies through heavyweight formal machinery. Naive compilation fails: extracted rules block the tool satisfying their own precondition, or read arguments their tool lacks. NOMOS, a four-pass compiler, turns a natural-language policy into a deterministic tool-call gate; static verification with tool-schema-level checks alone (no prover, solver, or LLM) repairs or rejects 37% (airline) and 13% (retail) of candidates, without which most shipped rules are inoperable. Replaying compiled rules over undefended transcripts flags bindings that refuse legitimate work (a development binding refused 95.9% of task-passing calls); no evaluation binding is flagged. On $τ^2$-bench the gate cuts violations of reference-encoded clauses among state-changing calls from 66.3% to 2.6% (airline) and 30.8% to 6.9% (retail), raising airline task success significantly for $2 \le k \le 4$; a 26B on-premise compilation is not significantly worse than hand-written or frontier-compiled rules. Unlike AgentDojo's shipped defenses, it reaches a zero attack success rate (ASR) on banking, where nine attack families collapse onto three structural rules. On the other three suites its ASR is at most 3.6%, from goals with no tool call to govern and one write admitted by a binding weaker than its clause; a second agent model, Llama-3.3-70B, reproduces the effect on both benchmarks. Decisions take microseconds without an LLM call, at a domain-dependent benign-utility cost; compilation runs on-premise on open-weight gemma-4-26B.
cs.CR / 7 / 2610.11107
Accelerating HQC for Post-Quantum TLS 1.3 on x86 IoT Gateways
Jihoon Jang, Hyunju Park, Jebin Kim, Seokhie Hong, Suhri Kim
cs.CR
Abstract
Post-quantum TLS at an IoT gateway must protect many device connections without large increases in handshake delay, CPU cost, or network traffic. HQC provides code-based diversity beyond ML-KEM, but its computation and ciphertext sizes can increase these costs. We optimize HQC for x86 processors with AVX2, AVX-512, and the Galois Field New Instructions (GFNI), and integrate the resulting implementations into TLS 1.3. We extend branch-free Toom-Cook/Karatsuba multiplication to AVX-512, accelerate Reed-Solomon decoding with GFNI, improve Reed-Muller decoding, accelerate SHA3-512 and fixed-weight sampling, and port Frobenius additive FFT (FAFFT) multiplication to AVX-512 + GFNI for HQC-5. On an Intel Core i5-1135G7 processor, our AVX2 implementation reduces decapsulation by 11.1-21.5% over the fastest prior AVX2 results. Our AVX-512 implementation reduces key generation by 24.9-29.5%, encapsulation by 10.4-11.1%, and decapsulation by 20.6-26.1% relative to Cabral et al. across the three HQC parameter sets. We evaluate the effects of these implementations on TLS 1.3 handshake latency and server and client CPU costs, as well as the effects of round-trip time (RTT) and bandwidth. In local loopback TLS measurements, our AVX-512 HQC-5 implementation reduces handshake latency from 3.22 ms to 2.92 ms compared with Cabral et al. On a constrained path with 1 Mbit/s bandwidth and an added RTT of 50 ms, the HQC-5 handshake takes 243 ms, compared with 87 ms for ML-KEM-1024. This indicates that HQC public-key and ciphertext sizes, rather than implementation speed, determine the remaining handshake cost on constrained links.
cs.CR / 8 / 2610.11112
False Claims, Credible Images: A Red-Teaming Benchmark for Commercial Image Generators
Zeyu Ye, Yanchun Li, Sibei He, Meng Xie, Hangtao Zhang, Xianlong Wang, Li Zeng, Jiahao Chen, Yichen Wang, Junhui Wang, Ziqi Zhou
cs.CR · cs.CV
Abstract
Image-generation models can now produce text-rich, natural-looking visual artifacts that are hard to distinguish from real-world evidence, such as news reports and textbook pages. Yet, the same capability introduces a new risk: these models can just as easily fabricate visual misinformation. Even commercial models (e.g., GPT-Image-2) readily produce it. Curiously, we find that these models can recognize a claim as false when asked, yet still render that very claim as credible visual evidence. This discrepancy points to a blind spot in current alignment: safeguards judge what an image shows, not what it asserts; however, existing red-teaming benchmarks target conventional harmful content, such as violent or explicit imagery, and say little about where the alignment boundaries lie for visual misinformation, especially in commercial models. To fill this gap, we introduce EpiReal-Bench, the first systematic benchmark for evaluating visual misinformation risks in commercial image generators, comprising 10k false-claim prompts and 10k corresponding generated images that span 10 real-world claim categories and 10 credible visual formats. We further introduce EpiReal-Attack, a skill-guided black-box optimization framework that uses Pareto-based selection and multimodal feedback to identify commands that bypass alignment safeguards while preserving visual realism, textual legibility, and semantic fidelity. Experiments on four commercial models reveal that more than 70% of false-claim prompts elicit images that faithfully depict the corresponding misinformation, and EpiReal-Attack pushes this rate to 95%. Most worryingly, these models are only a click away, and their outputs are cheap to spread yet hard to disbelieve, leaving this dimension of alignment largely unguarded.
cs.CR / 9 / 2610.11163
Characterizing Statistical Separability in TP-CRIV for Probabilistic AI Models
Teruki Sano, Minoru Kuribayashi, Masao Sakai, Shuji Isobe, Eisuke Koizumi, Zhang Zhang, Satoru Matsumoto
cs.CR · cs.AI
Abstract
Third-party challenge-response identity verification (TP-CRIV) enables an independent verifier to assess whether a claimant possesses a model identical to a remotely deployed model without directly accessing the reference model. However, for probabilistic AI models, repeated executions of the same query may produce different outputs and therefore different verification observations. This raises the question of how such stochastic evidence should be accumulated and how much evidence is required for reliable verification. In this work, we characterize statistical separability in TP-CRIV of probabilistic AI models. Specifically, we relate challenge-wise behavior of matching and non-matching provers to verification-level separability. The characterization explicitly describes how the numbers of independent challenges and repeated responses affect detection performance and enables the verification budget required for a target AUC to be estimated. We instantiate the proposed characterization for LLMs using open-ended challenges. The experiments demonstrate matching-non-matching separation, close agreement between theoretical and empirical AUCs, and consistent estimates of the minimum verification budgets. These results provide a statistical basis for relating probabilistic model behavior to verification-level separability and the evidence required for third-party verification.
cs.CR / 10 / 2610.11218
BRACE: Differential Privacy for Dense Associative Memory with LSR Energy
Chang Qu, Zhaoyang Shi
cs.CR · cs.LG
Abstract
Dense associative memory (DAM) provides an energy-based framework for memory retrieval with close connections to attention mechanisms in modern artificial intelligence. Despite growing interest in differential privacy for AI, the privacy of DAM retrieval dynamics remains relatively unexplored. In this paper, we develop a differential privacy framework for log-sum-ReLU (LSR) dense associative memory, whose finite-support retrieval dynamics pose distinctive challenges for privacy-preserving computation. We propose the Boundary-Responsive Adaptive Correction Evolution (BRACE) algorithm, a differentially private retrieval mechanism for LSR-DAM that adaptively corrects boundary-sensitive perturbations to control their cumulative effect over the retrieval trajectory. In theory, we prove that our method is minimax optimal by deriving dimension-independent terminal and full-trajectory retrieval error rates, with optimal dependence on the inverse temperature and, in the growing-horizon regime, the retrieval horizon. We further establish central limit theorems that enable uncertainty quantification for private retrieval by characterizing its asymptotic distribution and the additional variability introduced by privacy. Numerical experiments compare our proposed method with baseline differential privacy approaches and evaluate its retrieval accuracy. Together, our results provide a theoretical foundation for optimal privacy-preserving retrieval and uncertainty quantification in energy-based associative memory systems.
cs.CR / 11 / 2610.11254
Provable Subexponential Algorithms for NIST Third-Round Lattice Families
Yiming Gao, Xuyuan Han, Honggang Hu
cs.CR
Abstract
We give provable classical subexponential algorithms for secret recovery in growing parameter families associated with NIST third-round lattice candidates. For the Kyber/ML-KEM, FrodoKEM, SABER, NTRU LPRime, and Dilithium/ML-DSA families studied here, polynomial moduli and polylogarithmic coefficient scales yield recovery of the short secret component in expected time and space $2^{(1/2+o(1))n/\ln\ln n}$. For noisy or rounded linear relations, we exploit an exact gap in the squared Euclidean norm of a comparison vector defined by each coordinate guess. One Gaussian list suffices to identify every secret coordinate by binary search, without enumerating the others. We establish the required sampling guarantees through new geometric bounds for structured public operators over prime and power of two moduli. The construction builds on the Wagner-style Gaussian sampling framework of Ducas, Engelberts, and Loyer (CRYPTO 2025) and the low-error decision-LWE algorithm of Han, Gao, and Hu (2026). For quotient relations of NTRU type, we develop an affine slice search: fixing coordinates restricts candidate pairs to slices of the public lattice, and their estimated Gaussian masses guide the choice of each next coordinate. In the stated modulus window, Falcon's key generation quality condition supplies the required mass bound. The algorithm then recovers an equivalent signing key with high probability in time and space $2^{O(n/\ln\ln n)}$. The same search recovers the short key core for cyclic NTRU-HPS/HRSS. Together, these results give subexponential algorithms for problem families associated with all seven NIST third-round lattice candidates. Despite the subexponential complexity, our results do not establish a reduction in the concrete security of the currently specified parameter sets.
cs.CR / 12 / 2610.11285
Soft Voting for Policy-Aware Private Data Synthesis
Yingge Hu, Gautham Ramesh Babu, Mostafa Milani
cs.CR · cs.DB
Abstract
Blowfish privacy relaxes differential privacy (DP) by protecting only the attribute-value substitutions a data owner specifies as edges of a policy graph. A sparser policy can reduce the noise required by a mechanism, but only when the released statistic changes less across protected substitutions than across arbitrary DP neighbors. We study this question for evolutionary, nearest-neighbor DP synthesizers such as Private Evolution (PE) and its tabular instantiation Tab-PE, which score private records against a candidate population and release a noisy vote histogram. Their hard vote is constant inside each candidate's decision region and jumps at its boundary. Its policy-specific sensitivity therefore equals the full worst-case value whenever at least one protected substitution crosses a boundary, regardless of how short that substitution is. Because every round we examined contained such a substitution, the policy graph gave no reduction in noise. We propose BF-Soft, a temperature-smoothed soft vote whose response changes gradually with distance. Its sensitivity has a tight closed-form bound in the policy graph's reach and the temperature, independent of the number of candidates, and the bound can be computed once before synthesis. It also predicts from the policy alone when policy-aware smoothing cannot substantially reduce noise: protecting a flat categorical or binary attribute drives the reach to its maximum. On real and synthetic datasets under narrow numeric policies, BF-Soft reduces error relative to hard voting at strong privacy budgets, while the advantage reverses at weaker budgets. A public-data pilot predicts when soft voting is beneficial without spending private budget.
cs.CR / 13 / 2610.11290
ProxyEraseAgent: Blind Watermark Removal in the Wild
Jun Yao, Chao Wang, Yupeng Qiu, Zehua Ma, Weiming Zhang, Bin Liu, Han Fang
cs.CR
Abstract
Invisible image watermark removal has received growing attention. Despite substantial progress, existing attacks face a tension between practicality and specificity. Attacks exploiting detector outputs, decoder responses, or paired images can be tailored to the watermark decision boundary, but require information rarely available in realistic scenarios. Conversely, attacks based on compression, geometric distortion, or reconstruction are easily deployed from a single watermarked image, but remain largely open-loop: they apply generic transformations without knowing if the image is moving toward watermark failure. Thus, the key challenge in single-image blind watermark removal is not merely how to transform the image, but how to obtain a useful removal direction without accessing the hidden decoder. To bridge this gap, we propose ProxyEraseAgent, an agent-driven framework recovering attack specificity through proxy decoder responses. Publicly available watermarking schemes provide a natural knowledge base of candidate decoders, where some are informative for a given unknown image. Our insight is that a decoder producing a strong calibrated response to the query image likely shares a nearby decoding boundary with the hidden target mechanism. ProxyEraseAgent ranks these decoders by calibrated response strength and uses the top ones as proxy boundary estimators. Their responses then guide a progressive search over heterogeneous removal operations (e.g., geometric distortion, JPEG compression, image reconstruction, and gradient perturbation) under perceptual-quality constraints. Experiments across 11 watermarking systems show ProxyEraseAgent achieves a 94.8% attack success rate, demonstrating the effectiveness of response-guided proxy retrieval and feedback-driven sequential planning for blind watermark removal.
cs.CR / 14 / 2610.11398
MORDOR:Mitigating Overheads of Read Disturbance Preventive Operations via Elastic Refresh Scheduling
Maria Makeenkova, Ataberk Olgun, F. Nisa Bostancı, İsmail Emir Yüksel, Spiros Galanopoulos, Onur Mutlu
cs.CR
Abstract
Modern DRAM chips are susceptible to read disturbance phenomena such as RowHammer, where repeatedly accessing (hammering) a row of DRAM cells (i.e., a DRAM row) induces bitflips in other physically nearby (victim) DRAM rows. A common practice to avoid such bitflips is to preventively refresh victim rows that might otherwise experience bitflips. Unfortunately, preventive refreshes cause long latencies and need to be performed urgently before the aggressor row is activated again to ensure data integrity. This is done by prioritizing them over demand memory requests, thereby potentially imposing significant delays on those requests and causing performance and energy overheads. Our goal in this work is to alleviate these overheads by scheduling preventive refreshes off the critical path of demand memory requests. We propose MORDOR, a new preventive refresh scheduling policy that significantly reduces system performance degradation and energy consumption caused by preventive refresh operations. MORDOR is integrated into the memory controller and operates alongside memory-controller-based read disturbance mitigation techniques to intelligently delay preventive refresh operations, while maintaining their data integrity guarantees. MORDOR leverages the key observation that a preventive refresh operation targeting an aggressor row can be delayed to serve any other demand memory request, as long as that memory request does not access the aggressor row. By doing so, MORDOR executes latency-critical memory requests before long-latency preventive refresh operations, while mitigating read disturbance bitflips. We evaluate MORDOR by integrating it into six state-of-the-art read disturbance mitigation techniques. Our comprehensive evaluation shows that MORDOR significantly improves system performance and energy efficiency at low area cost.
cs.CR / 15 / 2610.11439
On-Chain Archaeology of Bitcoin Oracles: Evidence of Use under Limited Observability
Giulio Caldarelli
cs.CR · cs.CY · cs.IR
Abstract
Before Ethereum made the "oracle problem" a household term, Bitcoin already had oracles serving as feeds, key-release services, federated signers, and arbiters that carried real value on the main chain. This study traces their use and the changing evidence of oracle activity from early days through July 2026. We combine a complete census of Counterparty betting (1,149 bets), analysis of the full Bitcoin chain through block 958,628, and searches for documented keys from Reality Keys, Orisi, Bitrated, and Oraclize in an 854-million-row public-key index. We also recover DLC oracle records from an archived explorer and live Nostr relays. Two results emerge. First, early contracts remain on-chain, but many event descriptions have disappeared, and protocol encoding and API limitations complicate access to the surviving record. However, for modern DLCs, public oracle announcements can survive even when the contracts using them cannot be identified on-chain. In the script classes examined, the share of spends that reveal no script peaks at 81.9% in 2024 after excluding spends containing inscription data. Second, public registries can give a misleading picture of oracle use. In Counterparty, 95% of pre-2018 sources declaring an oracle fee were never bet on. In Bitrated, 0.1% of archived keys appear on-chain overall, compared with 10 of 19 keys captured in 2014. Sport dominates Counterparty's matched volume, while a daily price series dominates the archived DLC announcements. These findings show how protocol design and data preservation shape the historical record of Bitcoin oracle use.
cs.CR / 16 / 2610.11467
GROB: A Multi-Agent Architecture for Public-Trace Investigation of Candidate Agentic Activity
Chiara Bonfanti, Cataldo Basile
cs.CR · cs.AI
Abstract
We present GROB, a multi-agent architecture for investigating candidate autonomous-agent activity through public Internet traces when privileged telemetry is unavailable. The system performs controlled, read-only collection of public traces and preserves selected observations for later resolution. In a frozen September 2026 corpus, several collected traces became more informative as additional public evidence emerged. The strongest result concerns Census-labelled identifiers captured on 9 September. Public revision records later resolved these identifiers to specific Census requests from 16 - 17 June. Other results show weaker links between traces collected by GROB and evidence reconstructed or reported later. These links vary in strength, and only some can be tied to specific public records. The results show that sparse public traces can remain useful even before their significance is fully understood. Such evidence can support later reconstruction, but public traces alone do not establish organizational attribution. Execution identity presents a separate problem, as continuity of agent identity remains an active research question for autonomous language-model agents.
cs.CR / 17 / 2610.11471
A Survey of Security Research for Operating Systems
Masaki Hashimoto, Ruo Ando, Toshiyuki Maeda, Hidehiko Tanaka
cs.CR
Abstract
In recent years, information systems have become the social infrastructure, so that their security must be improved urgently. In this paper, we introduce the results of the survey of virtualization, operating system verification and access control technologies in association with the design requirements of the reference monitor. Additionally, we show the prospects and challenges for each technology.
cs.CR / 18 / 2610.11488
MARC: Multi-Bit Watermarking for Autoregressive Audio Generation against Codec Attacks
Liaoran Xu, Weizhi Liu, Zhaoxia Yin
cs.CR
Abstract
Generated audio is now used in a range of applications, creating a need to verify its origin after distribution and signal processing. This task is particularly challenging for autoregressive audio generation because codec processing can alter the token sequence recovered from the waveform. Such changes reduce the reliability of watermark detection and payload decoding. Existing methods construct token groups using either intrinsic token representations or substitution patterns caused by transformations. As a result, intrinsic token relationships and codec induced substitutions are modeled separately. In addition, most methods support only zero bit detection. They can determine whether a watermark is present but cannot distinguish individual generated outputs. We propose \textbf{MARC}, a multi-bit generative watermarking method for autoregressive audio generation. MARC integrates intrinsic token representations with confusion patterns obtained through retokenization and multiple codecs, forming a codec-aware token-cluster space. Within this space, payload-driven cluster scheduling is used to embed a multi-bit watermark, while detection and payload decoding are performed on retokenized observations. Experiments on speech, dialogue, and music generation show that MARC achieves an average of 97.3\% bit extraction accuracy on unmodified watermarked audio and the watermark can still be extracted under diverse codec attacks. MARC also demonstrates robustness to overwriting attacks.
cs.CR / 19 / 2610.11490
A Zero-Knowledge Signature Framework for Efficient Post-Quantum Message Authentication in Cooperative Automated Driving
Takahito Yoshizawa, Aysajan Abidin, Edoardo Pena-Gonzalez, Bart Preneel
cs.CR
Abstract
Connected and Automated Vehicles (CAV) rely on authenticated Vehicle-to-Everything (V2X) communications to exchange safety-critical information among vehicles and roadside infrastructure. As the automotive industry transitions toward post-quantum cryptography (PQC), the significantly larger public keys and signatures of standardized PQC digital signature algorithms introduce substantial communication overhead, which challenges the scalability of certificate-based V2X authentication, particularly for high-frequency cooperative awareness messages (CAM). This paper presents ZKS-PQC, a zero-knowledge signature framework that enables communication-efficient post-quantum message authentication for cooperative V2X systems. Instead of transmitting complete post-quantum public keys and signatures, the proposed framework replaces this authentication material with a compact Zero-Knowledge Proof (ZKP) while preserving compatibility with existing certificate-based trust architectures. This approach supports incremental migration and backward compatibility with the legacy ECDSA. The framework is implemented using the Open Quantum Safe (liboqs) and ZKP (Bulletproofs) libraries, and we evaluated the performance of standardized NIST PQC signature algorithms and additional candidate algorithms. Experimental validation on both a Linux platform and a commercial On-Board Unit (OBU) demonstrates substantial reductions in message size exceeding 95% for all algorithms, while limiting additional processing overhead by staying within the order of milliseconds at both sender and receiver in many algorithms, consistent with the latency requirements of real-time V2X operation. By decoupling communication overhead from the size characteristics of post-quantum signature algorithms, ZKS-PQC offers a practical migration strategy for scalable, quantum-resilient message authentication in future CAV systems.
cs.CR / 20 / 2610.11511
EIFL: Efficiently Protecting Global Model Privacy and Integrity Against an Untrusted Server in Federated Learning
Zehui Liao, Qiang Li, Binghui Wang
cs.CR
Abstract
Federated learning (FL) typically adopts a server-client architecture, where the server aggregates clients' local models (i.e., the input) and returns the aggregated global model (i.e., the output) to clients. An untrusted server may return a tampered global model to compromise the output integrity. Some existing schemes focus on verifying the output integrity. However, these schemes mostly rely on clients pre-negotiating a set of identical auxiliary information among themselves and keeping it confidential from the server. This dependency creates a verification vulnerability: if the auxiliary information is leaked, the verification method may be circumvented. Additionally, in some privacy-sensitive scenarios (e.g., commercial federated learning), the global model may need to be kept confidential from an untrusted server. However, only a few works achieve the output privacy while addressing verification vulnerability, at the cost of prohibitive computation and communication overhead. To address these challenges simultaneously, we propose a novel provably privacy-preserving FL method called EIFL. Specifically, we adopt a two-stage aggregation and combine it with symmetric encryption to protect the output privacy. To address the verification vulnerability, we propose an efficient verification method based on vector inner product for output integrity, and a random vector generation method for clients to agree on auxiliary information. EIFL innovatively binds the auxiliary information to the output integrity, and eliminates the need to keep the auxiliary information confidential from the server. Moreover, EIFL is robust against client dropout during the verification phase through a simple resending operation. Evaluation results validate the advantages of EIFL over state-of-the-art schemes in terms of computation and communication overhead.
cs.CR / 21 / 2610.11513
Certifying Hidden Paths: Scalable Topology Assurance for QKD Networks
Alessandro Colombo, Margherita Cozzolino, Stephan Krenn, Thomas Lorünser
cs.CR · quant-ph
Abstract
Large-scale Quantum Key Distribution (QKD) networks rely on trusted repeaters, making the security properties of the selected communication path an essential part of end-to-end assurance. At the same time, network operators may be unwilling to disclose their internal topology. We present a topology-certification mechanism that lets a provider prove in zero knowledge that policy-compliant routes between two endpoints exist, without disclosing the routes in an individual presentation. Our main idea is to certify nodes and edges independently using multi-message signatures, rather than signing the complete graph as one object. For a route chosen in advance, proof size and cryptographic proving and verification work depend on its positions and attributes, independently of the overall network size, whereas route discovery and certification have separate graph-dependent costs. The construction hides the actual route length up to a public bound and sketches extensions to node-disjoint routes and monotonic additions within a graph epoch. We give a formal security model and conditional proofs of unforgeability and graph hiding for a generic construction. For a BBS-based construction, presentation-size estimates remain below $300$\,KiB at $\ell=m=n=50$, and measured generation and verification times remain below $400$\,ms at $\ell=64$ and $m=n=8$.
cs.CR / 22 / 2610.11602
Where Do the Tokens Go? Understanding and Reducing Costs in LLM Agents for Vulnerability Discovery
Li Lu, Yanjie Zhao, Hongjie Chen, Haoyu Wang
cs.CR · cs.SE
Abstract
LLM agents can spend millions of tokens during vulnerability discovery without producing a working proof of concept (PoC). What consumes that budget, and why does it fail to produce results? We diagnose these costs and failures through a multi-axis open-coding study of 200 CyberGym traces, spanning four agents (i.e., Codex, OpenCode, Cybench, and EnIGMA) under an unaided baseline and four existing efficiency methods. The study reveals three key findings. First, different agents vary substantially in success and cost, and higher spending does not consistently yield better outcomes. Second, code localization and understanding, together with vulnerability reasoning and trigger design, account for 60.4% of tokens and represent the two leading bottlenecks in failed runs. Third, only 24.4% of matched comparisons preserve success at lower total cost; unsuitable signals and auxiliary overhead limit the benefits of existing methods. Motivated by these findings, we present AVRI, an Agent-centric Vulnerability Reasoning Interface built around a persistent Bidirectional Evidence Trace (BET). BET connects how the harness consumes input with the conditions required to make a selected operation unsafe, retaining source-supported correspondences alongside the agent's hypotheses and open questions. Reading, analysis, and persistence commands help agents build and reuse this evidence rather than repeatedly retrieve and reconstruct it. On 20 evaluation tasks, AVRI reduces total cost by 18.0% for Codex and 23.7% for OpenCode while preserving success rates and improving or maintaining recall.
cs.CR / 23 / 2610.11640
Host Attack Graph for Botnet Propagation
Andrei Neagu, Mara-Cristina Sterian, Paul Irofti
cs.CR · cs.GT · cs.NI
Abstract
Botnets represent the backbone of most modern cybersecurity attacks through which botmasters deploy distributed denial of service and advanced persistent threats against legitimate computer networks. In this paper we introduce the Host Attack Graph model and propose two botnet propagation strategies which we study across network topologies, target selection, and attack strategies together with their effectiveness in time. While strategy and botnet size are important for propagation, our study and simulations show that network topology and target selection should not be easily discarded. Implementation and benchmarks are available at https://codeberg.org/AndreiN/HAGP.
cs.CR / 24 / 2610.11843
Anytime-valid detection of LLM weight exfiltration
Ines Ortega-Fernandez, Mateusz Kowalczyk, Keri Warr
cs.CR · stat.ME
Abstract
A compromised LLM inference server can leak model weights by encoding payload bits in otherwise plausible token choices. A replay of the same prompt in a trusted server can expose such deviations, but benign numerical nondeterminism also causes token mismatches. Patient attackers can therefore hide within normal variation unless evidence is combined across responses. We introduce a prompt-level e-process that calibrates whole-response mismatch events on trusted benign traffic and accumulates evidence sequentially while, under a calibration-transfer assumption, controlling the probability of any false alarm over an unbounded monitoring horizon. We evaluate it on four models against a seed-blind attack and a stronger seed-aware attack that hides payload bits only in near-ties to remain stealthy, analyzing the channel capacity vs detectability trade-off. Compared with a hard per-token alarm, the e-process combines weak evidence across responses while providing explicit anytime false-alarm control.
cs.CR / 25 / 2610.11873
Secure Aggregate Encryption with Identity-Based Authentication for Multi-Vendor FPGA Cloud Deployment
Mukta Debnath, Krishnendu Guha, Debasri Saha, Amlan Chakrabarti, Susmita Sur-Kolay
cs.CR
Abstract
Secure deployment of FPGA bitstreams in heterogeneous multi-vendor cloud infrastructures requires scalable authorization, authenticated device access, and bitstream protection throughout the deployment lifecycle. Existing approaches typically address these functions through separate mechanisms, increasing coordination and key management requirements. This paper presents SAEID, a secure FPGA deployment framework that integrates aggregate authorization, certificate-free identity-based device authentication, and identity bound bitstream verification within a pairing based cryptographic framework, with symmetric cryptography used for session binding and AES 256 GCM based bitstream protection. Building on AgEID, an aggregate encryption scheme that enables individual decryption for authorized FPGA devices, SAEID extends the underlying framework to heterogeneous multi vendor deployments by supporting multiple FPGA vendors and IP providers, capability-aware authorization, and dynamic device membership. SAEID provides protection against unauthorized access to future deployments after device revocation and to prior deployments by newly enrolled devices, while retaining constant-size aggregate ciphertexts within each vendor domain and individual decryption. Experimental results demonstrate the practical performance of SAEID, with identity-based device authentication completing in approximately 4.35 ms and device membership updates requiring 4.95 to 5.09 micro seconds for device addition. The aggregate-encryption component exhibits scalable behavior with increasing device-set size. The complete SAEID software decryption path was also validated on a physical ZC702 Cortex A9 platform, requiring 191.126 ms. These results demonstrate scalable aggregate authorization with bounded authentication and dynamic-membership overhead for secure multi-vendor FPGA deployment.
cs.CR / 26 / 2610.11932
From Public Posts to AI-Search Citations: Measuring the Fragility of AI Search
Qi Liu, Geng Hong, Xinyang Zhang, Pei Chen, Yutong Li, Min Yang
cs.CR
Abstract
As more users ask AI systems for information, AI-search platforms are becoming a common gateway to web information. Unlike traditional search, which maps keywords to ranked pages, AI search retrieves pages, filters sources, selects citations, and generates answers before users see sources. This selection layer may amplify source bias and turn source choice into a security question. If a platform repeatedly cites domains where new users can publish posts easily, ordinary publication on those domains can become an indirect path into AI-search citations and answer text. Measuring this path is hard: platforms reveal little about citation selection, citations change over time, and the web contains so much background content that later answer changes are hard to attribute to our posts. We present a measurement framework for identifying and measuring this low-barrier publication path, combining cross-platform citation mapping, publication-barrier testing, and marker-controlled publication experiments. Across 10 AI-search platforms, we analyze 17,211 citation instances over 6,356 unique source domains and find: (1) citations concentrate in platform-specific sources, with top-20 domains capturing 20.5--70.8% of per-platform citations, and 15 of 22 tested publication platforms tied to cited source domains had low or medium barriers for both account setup and posting; (2) in our experiments, ordinary publication on preferred platforms changed what entered AI-search outputs: 8 of 10 platforms cited a fabricated concept within seven days, and one high-preference-platform article had greater citation impact than over 20 matched low-preference posts; and (3) this path is commercially available: a $14 GEO purchase produced 13 public posts, and one AI-search platform cited GEO-posted content with our designed markers within one hour.
cs.CR / 27 / 2610.11996
Moving Target Defense in SDN-enabled EV Charging Network
Roland Plaka, Mikael Asplund, Simin Nadjm-Tehran
cs.CR · cs.GT · cs.NI
Abstract
Software-Defined Networking (SDN) is emerging as a promis- ing technology for EV Charging infrastructure (EVCI) since it enables flexible, programmable control of EVCI networks, yet its security gain to EVCI remains underexplored. SDN-enabled EVCI faces an increas- ing number of cyber threats arising from the convergence of IT and OT components; in particular, low-rate Denial-of-Service (DoS) attacks. Ear- lier works identify a trend toward low-rate attacks in OT networks that may exhaust SDN switch flow tables and which can result in disabling the entire charging sites without triggering volume-based defenses. In this paper, we propose CS-SHIELD, a Moving Target Defense mecha- nism for SDN-enabled EVCI communication. We showcase the imple- mentation of the proposed mechanism that detects malicious flow table rules via cross-layer identity verification and responds by reassigning virtual IP addresses to all active chargers. Experiments on an emulated SDN testbed with a real charging protocol implementation demonstrate that CS-SHIELD maintains full site availability under attack, responds quickly, adds only a negligible delay under normal conditions. The reas- signing (shuffle) approach at each reaction makes the earlier reconnais- sance knowledge by the attacker rendered worthless.
cs.CR / 28 / 2610.12024
HPQ-AKE: A Provably Secure Sign-Less Hybrid Authenticated Key Exchange Protocol for Bandwidth-Constrained IoT and Edge Networks
Khiem Pham-Tuan, Minh Quang Le, Khuong Nguyen-An
cs.CR · cs.NI
Abstract
Securing bandwidth-constrained Internet of Things (IoT) and edge networks during post-quantum migration creates an authentication trade-off: ML-KEM provides post-quantum key establishment, whereas post-quantum signatures increase certificate-chain size, verification cost, and handshake latency. We present HPQ-AKE, a sign-less hybrid authenticated key exchange for RSA-enabled IoT gateways and edge/cloud services that replaces transcript signatures with dual KEMs. HPQ-AKE combines ML-KEM-768 for post-quantum session secrecy and forward secrecy with RSA-OAEP for implicit mutual authentication, preserving existing RSA-enabled infrastructure during migration. We analyze HPQ-AKE in an extended Bellare-Rogaway model, separating classical authentication in the Random Oracle Model from session-key secrecy in the Quantum Random Oracle Model. Under the Module-LWE and RSA assumptions, it establishes AKE security, forward secrecy, and conditional KCI resistance under the long-term-key-only exposure assumptions. Evaluation combines x86 micro-benchmarking, analytical latency modeling, and 10,000-iteration Monte Carlo simulation. HPQ-AKE reduces modeled handshake transmission from 13,009 to 5,668 Bytes, a 56.4% reduction relative to a Hybrid TLS 1.3 Full handshake, while requiring 7.11 ms of total local computation on the measured x86 testbed. Under a simulated 50 kbps satellite-like link with 600 ms RTT and stochastic jitter, median latency is 1,914 ms, 31.4% below the 2,791 ms baseline. At 20 ms RTT, modeled break-even thresholds are 0.31 Mbps against the pre-cached baseline and 4.08 Mbps against the full baseline; both decrease as RTT increases. These x86-based results motivate evaluation on gateway-class IoT and edge platforms; performance on ARM gateways and unaccelerated microcontrollers remains unvalidated.
cs.CR / 29 / 2610.12050
Protecting CPU AI On Edge TEEs: WebAssembly's Promise and Practical Challenges
Friedrich Vandenberghe, Lachlan Gunn, Bruno Volckaert, Merlijn Sebrechts
cs.CR · cs.SE
Abstract
AI models on edge hardware contain important intellectual property (IP), but an adversary can steal it when they achieve root access. Trusted Execution Environments (TEE) like Arm TrustZone protect against these Operating System (OS) level attacks. However, they are challenging to use. More precisely, it is difficult to run unaltered applications inside a TEE. This work presents a solution that allows the execution of unaltered AI models, compiled to WebAssembly, on the WebAssembly Micro Runtime (WAMR) in OP-TEE for Arm TrustZone. Additionally, this work provides an AI model distributor that encrypts the WebAssembly binary and places the encryption key in one of the device's fuses. This way, only the WAMR Trusted Application (TA) in OP-TEE can decrypt and execute the AI model. A thorough evaluation of our solution shows that it incurs an additional overhead of 22% in comparison to an application manually ported to OP-TEE, while also facilitating the execution of unaltered AI models with an additional inference latency of 6% without significant porting effort. Overall, there is a pressing need to safeguard the IP of AI models and this work shows that there is a real promise in WebAssembly, but there remain some practical challenges.
cs.CR / 30 / 2610.12220
One Node, Two Roles: Simultaneous Contests for Validation and Attention in Rollups
Pranay Anchuri, Ben Berger, Matteo Campanelli, Akaki Mamageishvili
cs.CR
Abstract
An optimistic rollup is safe as long as at least one honest validator executes its state transition function (STF) and disputes any assertion that is inconsistent with the result of the execution. Allowing anyone to participate, however, does not give incentive to do so: as long as every assertion is correct, verifying assertion correctness goes unpaid. Attention mechanisms address this by paying nodes to execute the rollup's STF, regardless of whether they participate in the validation process. Since a single node operator can back many registered identities with a single execution, paying per identity does not buy execution diversity, that is, the number of operators that independently execute. In this work we provide a new modeling framework for this problem. We formalize the underlying cryptographic primitive as a new notion, arguments with registered provers, whose properties include a form of non-amortizability that, to our knowledge, has not been studied for modern succinct cryptographic proofs before. We design validation and attention reward mechanisms and analyze them as two coupled Tullock-like contests, in which an operator pays once for the required execution and then chooses how many identities to deploy in each role. We quantify the attainable range of execution diversities in equilibrium as a function of the marginal cost of deploying an extra attention identity. We use our framework as a lens through which we compare two attention mechanisms, TRACE (MARBLE 2026) and Proof of Diligence (AFT 2024). In particular, we show that as opposed to Proof of Diligence, in TRACE the marginal cost mentioned above can be adjusted to support a higher execution diversity. As a case study, we apply our results to Arbitrum One, a widely deployed optimistic rollup.
cs.CR / 31 / 2610.12233
ReSI: Recursive Safety Improvement toward Resistant and Resilient AI
Jingnan Zheng, Dongcheng Zhang, Yi Zhang, Ming Zhang, Qiaosheng Zhang, Youbang Sun, An Zhang, Xiangnan He, Tat-Seng Chua, Xia Hu, Bowen Zhou, Chaochao Lu, Xiang Wang
cs.CR · cs.AI
Abstract
Recursive self-improvement, the participation of AI systems in improving their own capabilities, is beginning to move from theoretical prospect to practice, posing both challenges and opportunities for safety alignment. Models evolve through frequent updates, and their safety alignment requires continual adaptation to each new checkpoint. Meanwhile, with evolving red-teaming methods exposing new vulnerabilities, safety improvement for each checkpoint needs to mitigate exposed vulnerabilities and generalize to risks not yet revealed. Following R$^2$AI, we term these goals resistance to known threats and resilience to unforeseen risks. Recursive self-improvement, in turn, inspires an approach to both goals: safety alignment could likewise advance through successive rounds of evaluation and update. We therefore introduce ReSI, a recursive safety improvement framework that implements this approach through automated research. In each round, ReSI applies diverse red-teaming methods to identify vulnerabilities in the current target model, develops training recipes, and promotes the update with the largest safety gain among those passing a Pareto gate on capability retention as the next target model. Across four dense and mixture-of-experts models, ReSI matches or exceeds evaluated frontier models on in-distribution and out-of-distribution safety benchmarks, and outperforms alignment baselines on nearly all safety evaluations while largely preserving general capabilities. In particular, ReSI reduces the mean X-Teaming attack success rate across the four models from 86.01% to 31.45%, well below GPT-5.6-Luna's leading frontier result of 56.69%, indicating stronger resilience to attacks unseen during training. These findings support recursive safety improvement as a practical path toward resistant and resilient AI.
cs.CR / 32 / 2610.12415
ORCAGen: Orchestrating Context-Aware Malware Deception with RAG-Guided Generative AI
Shihab Ahmed, Md Sajidul Islam Sajid, Teryl Taylor, Frederico Araujo, Tariqul Islam
cs.CR
Abstract
Malware defenses often remove or isolate suspicious programs as quickly as possible. While effective for containment, this approach can also waste an opportunity to observe attacker behavior and deploy targeted countermeasures. ORCAGen takes a different approach: it uses GenAI to build malware-specific deception playbooks offline, validates them before deployment, and enforces only the verified logic at runtime. ORCAGen combines Retrieval-Augmented Generation (RAG) with structured prompt engineering to generate both proof-of-concept (PoC) malware and corresponding deception orchestration code. A curated knowledge base (KB) of malware procedures and active defense strategies grounds the generation process, helping the LLM produce threat-specific and executable deception logic rather than generic or hallucinated outputs. The generated PoC malware provides a safe and reproducible way to test whether a deception strategy can disrupt, redirect, or suppress targeted malware behavior before the strategy is added to the runtime playbook. We evaluate ORCAGen across GPT-4o, GPT-5.5, Gemini 3.5 Flash, Qwen3-Coder, and Claude Sonnet 4.5 using execution success, hallucination rate, refinement effort, deception effectiveness, and runtime overhead. The evaluation covers synthesized malware scenarios and 150 real-world malware samples across keyloggers, information stealers, and ransomware. Across the evaluated scenarios, GPT-5.5 required the fewest refinements and produced no observed hallucinated APIs, while Gemini 3.5 Flash achieved the lowest response time and runtime overhead. The results show that RAG-guided structured prompting can support the scalable construction of malware-specific deception playbooks that remain lightweight and deterministic during runtime enforcement.
cs.CR / 33 / 2610.12463
From Reactive Containment to Proactive Assurance: Lessons from OpenAI, Anthropic, and Google Agent Security Incidents
Abbas Raftari
cs.CR · cs.AI
Abstract
In 2026, cybersecurity evaluations involving OpenAI, Anthropic, and Google agents reached real systems outside their authorized test scope. The paths were different. OpenAI agents exploited research infrastructure, coordinated across runs, and compromised parts of Hugging Face's production environment. Anthropic reported cases in which a misconfigured third-party environment exposed real systems to agents pursuing simulated cyber tasks. In a separately reported evaluation, Google's Gemini accessed three real organizations through an unintended internet route; Google stated that the model stopped in all three instances. Taken together, the cases show why an evaluation cannot rely on an assumed boundary. That boundary must be verified while the agent is operating. This comparative instrumental case study develops a Proactive Agent Security Assurance Cycle (PASAC) and a five-layer Boundary Assurance Stack. The framework combines risk-tiered task design, executable scope contracts, pre-run validation, least-capability access, independent egress enforcement, credential restrictions, cross-run monitoring, automatic stop conditions, and evidence-based reauthorization. A leading-indicator model, nine design propositions, and seven falsifiable hypotheses turn these lessons into a testable research program. Because the public Gemini record is limited to attributed statements and journalism, its detailed causal mechanism remains provisional. The central conclusion is straightforward: proactive agent security requires continuous assurance across the full execution system, not confidence in any single sandbox or safeguard.
cs.CR / 34 / 2610.11870
Local Sensitivity in Exponential Selection: Failure Modes and Valid Calibrations
Dung Nguyen, Anil Vullikanti
cs.DS · cs.CR
Abstract
Selection is a task that chooses one element from a finite public candidate range to maximize a data-dependent score. In differential privacy (DP), the exponential mechanism (EM) samples a candidate at a temperature calibrated to the global sensitivity. In this paper, we study when dataset-dependent sensitivity can safely replace global sensitivity in private selection. We propose three valid approaches. First, a private, high-probability upper bound on local sensitivity yields approximate DP, and the method extends to finite higher-order sensitivity hierarchies. Second, our Propose-Test-Release (PTR) variant privately searches a finite public grid for a temperature scale rather than fixing it in advance. Third, smooth sensitivity supports several designs. A candidate-independent smooth geometric construction produces a sensitivity envelope that is admissible under the local dampening framework, which privacy is guaranteed for any admissible envelope. Additionally, a separate logarithmic transformation utilizes smooth sensitivity to produce a smoothed candidate score function with advantages: having controlled global sensitivity, and preserving the maximizers of the original utility score, i.e., candidates maximizing the utility. Both of the designs yield range-independent pure DP. Besides that, we also give two approximate DP private selectors using smooth sensitivity: a direct EM with smooth sensitivity calibrated to the candidate range and privacy parameters that matches a theoretical lower bound up to some constant factor, and one using a privatized smooth upper scale by analyzing the logarithmic transform of the smoothness. For every proposed mechanism, we derive a high-probability regret bound under its stated conditions.
cs.CR / 35 / 2610.12357
Improved Local Leakage Resilience of Shamir Secret Sharing and Worst-Case Optimal Polynomial Intersection
Yihang Sun, Mary Wootters
quant-ph · cs.CR · cs.DM · cs.IT
Abstract
We study two problems: Local Leakage Resilience (LLR) for Shamir secret sharing, and worst-case Optimal Polynomial Intersection (OPI). Both problems concern polynomials $Q(X)$ of degree less than $k$, over a prime-order finite field $\mathbb{F}_p$. In LLR for Shamir secret sharing, one asks how much one can learn about $Q(0)$ given a few bits leaked from each of $Q(α_1), \ldots, Q(α_n)$, for distinct non-zero evaluation points $α_i \in \mathbb{F}_p$. In OPI, one is given input list $S_1, \ldots, S_n \subset \mathbb{F}_p$, and wants to find a polynomial $Q(X)$ of degree less than $k$ so that $Q(α_i) \in S_i$ for as many $i$ as possible. Leveraging recent connection between these two problems due to (Sun, Wootters 2026), we improve the state-of-the-art for both problems. For LLR, we show that there is some constant $δ> 0$ so that, as long as $R := k/n \geq 1/2 - δ$, Shamir secret-sharing is one-bit locally leakage resilient (meaning that one can learn only a negligible amount about $Q(0)$). This is the first result to break the so-called "one-half barrier" for LLR, and improves over the previous best known result, requiring $R \geq 0.668$ (Kasser, 2025). For OPI, we give a quantum algorithm that finds a polynomial $Q(X)$ that agrees with at least a $\mathsf{SCL}_ρ(R)-\varepsilon$ fraction of the lists in expectation, for every fixed $\varepsilon>0$, where $\mathsf{SCL}_ρ$ is the \emph{semicircle law} of (Jordan et al., 2025). This improves previous algorithmic (and existential) results of (Jo, 2026) and (Horinaga, Yamakawa, 2026). We also give further improved existential results. We also adapt the hardness result of (Yamakawa, Zhandry, 2024) to apply to OPI (rather than a folded version); over large fields, this gives an unconditional separation between the quantum and classical hardness of OPI relative to a membership oracle.