Daily Research Digest
arXiv Papers
2026-10-07
588
Papers
8
Categories
123
Translated
收藏清单 0
精选 · Favorites
123
cs.AI / 1 / 2610.07206
Energy-Conditioned Noise Schedule and Whitening for Spectral Diffusion
用于谱扩散的能量条件噪声调度与白化
diffusion
扩散模型相关
Abstract
This paper introduces an energy-adaptive noise scheduling and whitening strategy for transform-domain diffusion models. Existing spectral diffusion methods account for the non-uniform statistics of transform coefficients through coefficient scaling, normalization, or frequency prioritization, while the forward diffusion noise schedule remains largely independent of the underlying spectral-energy distribution. We investigate whether the temporal evolution of the forward diffusion process should also follow the spectral organization of natural images. The proposed formulation combines global spectral whitening with energy-conditioned noise allocation that jointly modulates the injected noise according to the energy of individual transform coefficients and an image-dependent energy path over diffusion time. The resulting forward process preserves Gaussian transitions with closed-form marginals and remains compatible with standard DDPM and DDIM procedures without modifying the diffusion architecture. Experiments on CIFAR-10 demonstrate the contribution of the proposed energy-conditioned noise schedule and spectral whitening, reducing Fréchet Inception Distance from 142.48 for a compact DCTdiff U-Net variant to 100.45.
Chinese Translation
本文提出了一种用于变换域扩散模型的能量自适应噪声调度与白化策略。现有的谱扩散方法通过系数缩放、归一化或频率优先级来考虑变换系数的非均匀统计特性,而前向扩散噪声调度在很大程度上仍独立于潜在的谱能量分布。我们研究了前向扩散过程的时间演化是否也应遵循自然图像的谱组织方式。所提出的公式将全局谱白化与能量条件噪声分配相结合,后者根据各个变换系数的能量以及随扩散时间变化的、依赖于图像的能量路径,共同调制注入的噪声。由此得到的前向过程保持了具有闭式边缘分布的高斯转移,并且在不修改扩散架构的情况下与标准 DDPM 和 DDIM 流程保持兼容。在 CIFAR-10 上的实验证明了所提出的能量条件噪声调度与谱白化的贡献,将一个紧凑的 DCTdiff U-Net 变体的 Fréchet Inception Distance 从 142.48 降低到 100.45。
cs.AI / 2 / 2610.07250
Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation
通过同策略上下文蒸馏将智能体经验内化到扩散模型权重中
diffusion
扩散模型相关
Abstract
Wrapping an image generation model in an agentic harness can effectively boost Text-to-Image task performance: the harness can leverage memory, skills, workflow orchestration, result verification, and iterative refinement to continually construct and revise prompts, thereby eliciting better images. These gains, however, remain external to the diffusion model and are realized only while the full harness runs. We propose Diffusion On-Policy Context Distillation (D-OPCD), which treats the agent-improved prompt as privileged context and distills the knowledge encoded in the agent harness into the weights of the diffusion model, so that the model retains part of the harness's benefit when conditioned on the original query alone. Using a Text-to-Image agent equipped with our proposed Auto Skill Evolver (ASE), we show that D-OPCD can internalize harness capabilities into the generator's weights, raising the average direct-generation score from 60.52 to 65.09 across four benchmarks. With this knowledge absorbed into the weights, the harness can shed its saturated skills and resume evolving: a second ASE round on the updated generator improves on a skill-free harness by additional 1.83 points, pointing toward text-to-image systems in which harness and model keep improving each other through continual co-evolution.
Chinese Translation
将图像生成模型包装在智能体框架中,可以有效提升文本到图像任务性能:该框架可以利用记忆、技能、工作流编排、结果验证和迭代细化来持续构建和修改提示,从而引出更好的图像。然而,这些收益仍处于扩散模型之外,并且只有在完整框架运行时才能实现。我们提出扩散同策略上下文蒸馏(Diffusion On-Policy Context Distillation,D-OPCD),它将智能体改进后的提示视为特权上下文,并将智能体框架中编码的知识蒸馏到扩散模型的权重中,从而使模型在仅以原始查询为条件时仍保留框架的部分收益。使用配备了我们所提出的自动技能进化器(Auto Skill Evolver,ASE)的文本到图像智能体,我们表明 D-OPCD 可以将框架能力内化到生成器的权重中,在四个基准上将平均直接生成得分从 60.52 提高到 65.09。随着这些知识被吸收到权重中,该框架可以舍弃其已饱和的技能并重新开始进化:在更新后的生成器上进行第二轮 ASE 相比无技能框架额外提升了 1.83 分,指向一种文本到图像系统,其中框架和模型通过持续共同进化相互不断改进。
cs.AI / 3 / 2610.07342
Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding
理由引导的策略优化:借助自适应理由支架学习推理
large language model
大语言模型相关
Abstract
On-policy reinforcement learning has become a central paradigm for improving the reasoning abilities of large language models. However, its effectiveness is often limited by reward sparsity: when a model fails to discover correct trajectories for difficult problems, the optimization process receives little useful signal and may stagnate. Existing approaches mitigate this issue by incorporating off-policy demonstrations, expert traces, or model-generated solutions, but they typically require the auxiliary data to match the format of the reinforcement-learning task, often relying on rejection sampling from stronger models to obtain suitable training trajectories. We introduce Rationale-Guided Policy Optimization (RGPO), a framework that adaptively leverages ground-truth rationale information according to the model's current capability while preserving its freedom to explore. Rather than treating reference solutions as fixed imitation targets, RGPO uses them as temporary scaffolds: rationales help the model generate improved responses, after which only higher-reward, model-generated solutions are transferred back to the original unguided setting. This design allows training to exploit available ground-truth information without requiring off-policy data to follow the same format as the RL task. Across both language-only and vision-language reasoning settings, RGPO consistently improves performance over RLVR baselines, and ablation studies show that adaptive rationale guidance is a key contributor to these gains. These results suggest that RGPO offers a practical and general approach for reducing reward sparsity, stabilizing reinforcement learning, and improving reasoning performance in both text-only and multimodal models.
Chinese Translation
在线策略强化学习已成为提升大语言模型推理能力的核心范式。然而,其有效性常常受限于奖励稀疏性:当模型无法为困难问题发现正确轨迹时,优化过程只能接收到很少的有用信号,并可能陷入停滞。现有方法通过引入离策略示范、专家轨迹或模型生成的解答来缓解这一问题,但它们通常要求辅助数据与强化学习任务的格式相匹配,并且往往依赖从更强模型中进行拒绝采样以获得合适的训练轨迹。我们提出理由引导的策略优化(RGPO),这是一个根据模型当前能力自适应地利用真实理由信息的框架,同时保留其探索自由。RGPO 不将参考解答视为固定的模仿目标,而是将其用作临时支架:理由帮助模型生成改进后的回答,之后只有奖励更高的、由模型生成的解答会被迁移回原始的无引导设置中。这种设计使训练能够利用可用的真实信息,而不要求离策略数据遵循与强化学习任务相同的格式。在纯语言和视觉-语言推理设置中,RGPO 相较于 RLVR 基线持续提升性能,消融研究表明自适应理由引导是这些增益的关键贡献因素。这些结果表明,RGPO 提供了一种实用且通用的方法,用于降低奖励稀疏性、稳定强化学习,并提升纯文本和多模态模型的推理性能。
cs.AI / 4 / 2610.07376
MemCo: Memory-Centric Collaboration for Generalizing LLM Agents to Unseen Environments
MemCo:以记忆为中心的合作,用于将 LLM 智能体泛化到未见环境
large language model
大语言模型相关
Abstract
Large language model (LLM) agents increasingly operate in interactive environments, where they need to make sequential decisions through observation, action, and feedback. Although memory can help agents reuse experience, existing work designs memory in isolation, where collecting enough trajectories to populate it is expensive. Existing shared-memory approaches mitigate isolated experience by pooling episodic memories across tasks and environments. However, retrieving shared memory is challenged by the granularity, where retrieved memories can be either too specific to preserve current grounding or too coarse to support the next action. In this work, we propose MemCo, a memory-centric collaboration framework for generalizing LLM agents to unseen interactive environments. It maintains complementary local and global memory spaces, preserving environment-specific details locally while promoting transferable workflows induced from local trajectories to global memory. During online interaction, MemCo routes relevant local and global memories in terms of the agent's current state and decision phase, enabling agents to reuse the experience of other agents without blindly transferring environment-specific details. Experiments on interactive decision-making benchmarks show that MemCo improves task success and reduces redundant exploration compared with isolate-memory and shared-memory baselines. Our code is available at https://github.com/SYannL/nvdamas.
Chinese Translation
大型语言模型(LLM)智能体越来越多地在交互式环境中运行,它们需要通过观察、行动和反馈做出序列决策。尽管记忆可以帮助智能体复用经验,但现有工作将记忆设计为孤立的,其中收集足够轨迹来填充记忆代价高昂。现有的共享记忆方法通过跨任务和环境汇聚情景记忆来缓解孤立经验的问题。然而,检索共享记忆受到粒度问题的挑战:检索到的记忆要么过于具体,无法保持当前情境基础,要么过于粗略,无法支持下一动作。在这项工作中,我们提出了 MemCo,一个以记忆为中心的合作框架,用于将 LLM 智能体泛化到未见的交互式环境。它维护互补的局部和全局记忆空间,在局部保留环境特定的细节,同时将从局部轨迹中归纳出的可迁移工作流提升到全局记忆。在线交互过程中,MemCo 根据智能体当前状态和决策阶段路由相关的局部和全局记忆,使智能体能够复用其他智能体的经验,而不会盲目迁移环境特定的细节。在交互式决策基准上的实验表明,与孤立记忆和共享记忆基线相比,MemCo 提高了任务成功率并减少了冗余探索。我们的代码可在 https://github.com/SYannL/nvdamas 获取。
cs.AI / 5 / 2610.07403
Defense-in-Depth for LLMs: Evaluating Memory Gates Against Activation-Induced and Memory-Induced Sycophancy
面向LLMs的纵深防御:评估记忆门控对激活诱发型与记忆诱发型谄媚的防御
large language model
大语言模型相关
Abstract
Long-term memory allows Large Language Models (LLMs) to maintain personalized context across interactions, but retrieved user history can induce memory-induced sycophancy, causing models to favor stored user beliefs over objective evidence. Existing defenses primarily operate on retrieved context and are rarely evaluated jointly with internal behavioral bias. We introduce a $2 \times 2$ defense-in-depth framework separating internal activation steering from external memory handling. We extract sycophancy steering directions from 100 paired prompts and evaluate four open-weight models across 10 steering coefficients and five memory-defense configurations on MemSyco-Bench (answers for all 1,550 items; defense conditions judged on a fixed 250-item subsample), with three LLM judges. Three of the five configurations are new (rewriting every memory, a Router Gate that keeps, rewrites, or drops each memory, and dropping all memory); the other two are MemSyco's baselines. Selective Router Gate filtering preserves substantially more of MemSyco's average accuracy than complete memory removal, and this separation persists when the models are steered toward sycophancy. On Llama 3.1 8B with Router Gate, mild inverse steering ($α= -1.5$) lowers judge-averaged sycophancy from 35.80% to 31.32% while average accuracy moves from 43.99% to 43.31%; this reduction has the same direction under all three judges but is not statistically significant (paired $p = 0.08$ to $0.63$ on 149 items). External memory filtering is the part of the design that holds up; our data do not show that inverse steering adds to it.
Chinese Translation
长期记忆使大型语言模型(LLMs)能够在交互之间维持个性化上下文,但检索到的用户历史可能诱发记忆诱发型谄媚,导致模型偏好存储的用户信念而非客观证据。现有防御主要作用于检索到的上下文,并且很少与内部行为偏差一起评估。我们引入一个 $2 \times 2$ 纵深防御框架,将内部激活引导与外部记忆处理分开。我们从100个成对提示中提取谄媚引导方向,并在 MemSyco-Bench 上评估四个开放权重模型在10个引导系数和五种记忆防御配置下的表现(对全部1,550个条目给出答案;防御条件在固定的250个条目子样本上评判),并使用三个LLM评判器。五种配置中的三种是新的(重写每一条记忆、一个对每条记忆进行保留、重写或丢弃的路由门控,以及丢弃所有记忆);另外两种是 MemSyco 的基线。与完全移除记忆相比,选择性路由门控过滤显著保留了更多 MemSyco 的平均准确率,并且当模型被引导向谄媚时,这种分离仍然存在。在 Llama 3.1 8B 上使用 Router Gate 时,轻度反向引导($α= -1.5$)将评判器平均谄媚从35.80%降至31.32%,而平均准确率从43.99%变为43.31%;这种降低在全部三个评判器下方向相同,但在统计上不显著(在149个条目上配对 $p = 0.08$ 到 $0.63$)。外部记忆过滤是设计中站得住脚的部分;我们的数据没有表明反向引导会为其增加收益。
cs.AI / 6 / 2610.07434
When Does AI Supervision Help? A Role-Aware Study of Network Fraud Decision Management with Blockchain Auditability
人工智能监督何时才有帮助?一项具有区块链可审计性的角色感知网络欺诈决策管理研究
large language model
大语言模型相关
Abstract
When does a second artificial intelligence (AI) component improve a primary network-fraud decision rather than add operational burden? We study this question through a role-aware Decider-Supervisor (DS) framework with blockchain auditability, evaluating four directional configurations that combine centralised machine learning, a Federated Averaging (FedAvg)-trained federated meta-model, and Base or Quantized Low-Rank Adaptation (QLoRA) large language model variants. The analysis compares primary-only and supervised decisions using non-hard fraud performance, intervention burden, conditional calibration, traffic-mix and Review-capacity sensitivity, dependability tests, and blockchain lifecycle controls. The deterministic hard gate resolves 89.994% of fraudulent requests, leaving the non-hard population as the main AI decision setting. Conditional validation calibration does not produce a consistently transferable supervisory advantage on deployment replay. DS-3 QLoRA is the least disruptive supervised configuration, but it still underperforms its primary FedAvg stage in F1 and total errors. Across 36 reweighted traffic mixtures, supervision reduces total errors only for DS-4 Base in two extreme high-fraud scenarios. Blockchain tests support digest verification, tamper detection, authorisation, single-use review resolution, and post-finalisation integrity, while exposing a pre-finalisation single-write limitation. The results show that the value of AI supervision depends on role assignment, calibration, escalation policy, traffic composition, and lifecycle controls rather than on the presence of a second model alone.
Chinese Translation
第二个 AI 组件何时能改善主网络欺诈决策,而不是增加运营负担?我们通过一个具有区块链可审计性的、角色感知的决策者-监督者(Decider-Supervisor,DS)框架来研究这一问题,评估了四种方向性配置,这些配置组合了集中式机器学习、经联邦平均(Federated Averaging,FedAvg)训练的联邦元模型,以及 Base 或量化低秩适配(Quantized Low-Rank Adaptation,QLoRA)大语言模型变体。该分析使用非硬欺诈(non-hard fraud)性能、干预负担、条件校准、流量组合与审查容量敏感性、可靠性测试以及区块链生命周期控制,对仅主决策与受监督决策进行了比较。确定性硬门控解决了 89.994% 的欺诈请求,从而将非硬群体留作主要的 AI 决策场景。条件验证校准在部署回放中并未产生一致可迁移的监督优势。DS-3 QLoRA 是破坏性最小的受监督配置,但它在 F1 和总错误数上仍不如其主 FedAvg 阶段。在 36 种重新加权的流量组合中,监督仅在两种极端高欺诈场景下为 DS-4 Base 减少了总错误数。区块链测试支持摘要验证、篡改检测、授权、一次性审查决议以及最终化后的完整性,同时暴露出一个最终化前单次写入的限制。结果表明,AI 监督的价值取决于角色分配、校准、升级策略、流量构成和生命周期控制,而非仅仅取决于第二个模型的存在。
cs.AI / 7 / 2610.07473
PsyCIDRA: A Dual-Agent Framework for Psychiatric Interviewing and Diagnostic Reasoning
PsyCIDRA:用于精神科访谈与诊断推理的双智能体框架
large language model
大语言模型相关
Abstract
Large language models show promise in clinical reasoning, but psychiatric interviewing requires guiding an evolving conversation. Their ability to carry out this interactive assessment remains less studied. We present PsyCIDRA, a dual-agent framework linking free-form psychiatric interviewing with diagnostic reasoning for expert review. Its interviewer agent uses tools to maintain working notes, load expert-written skills, and retrieve ICD-11 references to guide inquiry. Its diagnostic reasoning agent then receives the completed interview transcript and reports hypotheses alongside supporting, conflicting, and missing evidence, withholding a final hypothesis when none is sufficiently supported. Using patient profiles generated with PsyCPG, we first evaluate PsyCIDRA in simulation. Across four models on 53 evaluation cases, it achieves higher diagnostic agreement than direct prompting. On 81 held-out simulated cases, rank-1 accuracy is 60.5% versus 51.9%. In a blinded study of 101 human participants in separate arms, PsyCIDRA agrees with psychologists on whether to propose a diagnostic hypothesis in 79.6% of cases, compared with 65.4% for direct prompting. Together, these findings support the potential of LLM agents to assist psychiatric assessment through free-form dialogue. By examining diagnostic reasoning, interview quality, and safety together, this study contributes to understanding the capabilities and limitations of psychiatric interview agents.
Chinese Translation
大语言模型在临床推理中展现出前景,但精神科访谈需要引导一场不断演变的对话。它们执行这种交互式评估的能力仍较少被研究。我们提出 PsyCIDRA,一个将自由形式的精神科访谈与面向专家评审的诊断推理相连接的双智能体框架。其访谈智能体使用工具来维护工作笔记、加载专家编写的技能,并检索 ICD-11 参考文献以指导问诊。其诊断推理智能体随后接收已完成的访谈转录文本,并报告假设以及支持性、冲突性和缺失的证据;当没有任何假设得到充分支持时,则暂不给出最终假设。使用由 PsyCPG 生成的患者画像,我们首先在模拟中评估 PsyCIDRA。在 53 个评估案例上的四个模型中,它实现了比直接提示更高的诊断一致性。在 81 个留出的模拟案例上,rank-1 准确率为 60.5%,相比之下为 51.9%。在一项有 101 名人类参与者、分属不同臂的盲法研究中,PsyCIDRA 在是否提出诊断假设方面与心理学家的一致率为 79.6%,而直接提示为 65.4%。总之,这些发现支持 LLM 智能体通过自由形式对话辅助精神科评估的潜力。通过同时考察诊断推理、访谈质量和安全性,本研究有助于理解精神科访谈智能体的能力与局限性。
cs.AI / 8 / 2610.07480
In With the Old: Enhancing 'Classical' Document Automation with Generative AI
旧法新用:以生成式人工智能增强“经典”文档自动化
large language model
大语言模型相关
Abstract
Software-based legal assistance systems have leveraged many different forms of knowledge representation and reasoning. This article explores how document automation services rooted in expert system style and other symbolic approaches can usefully enhance and be enhanced by current generative AI approaches. We discuss the possible benefits and challenges, and report on preliminary experiments in using large language models to identify and fix issues in texts written by laypeople.
Chinese Translation
基于软件的法律援助系统已经利用了许多不同形式的知识表示与推理。本文探讨了植根于专家系统风格和其他符号方法的文档自动化服务如何能够有效地增强当前的生成式人工智能方法,并被当前的生成式人工智能方法增强。我们讨论了可能的益处和挑战,并报告了在利用大语言模型识别和修复非专业人士所写文本中的问题方面的初步实验。
cs.AI / 9 / 2610.07544
A Systematic Investigation of Bias in Large Language Models for Advertising Relevance
大型语言模型在广告相关性判断中偏见的系统性研究
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly used to judge how well an advertisement matches a query, but the fairness of these judgments has received limited attention. We conduct a systematic study of fairness in relevance judgments made by LLMs for queries and advertisements. Our counterfactual framework examines the effects of advertiser identity and possible popularity, input language, and demographic wording. We study GPT-4o as a categorical relevance judge and a Qwen-7B model trained specifically for relevance prediction. The advertiser and language experiments use query and advertisement pairs sampled from real advertising logs. Controlled synthetic queries are used to study demographic associations in employment, housing, and credit. For both models, changing the advertiser identity or input language can alter the relevance assessment. Selected demographic comparisons also show patterns consistent with common stereotypes, particularly those involving gender and occupation. We further study mitigation during model inference and training. The results indicate that its effectiveness depends on whether advertiser information is relevant to the query and how advertiser labels are distributed in the training data. These findings can help advertising practitioners identify fairness risks and develop suitable mitigation methods for LLM relevance systems.
Chinese Translation
大型语言模型(LLM)正越来越多地被用于判断广告与查询的匹配程度,但这些判断的公平性所受到的关注却很有限。我们对大型语言模型针对查询和广告所作相关性判断中的公平性开展了系统性研究。我们的反事实框架考察了广告主身份与可能的热度、输入语言以及人口统计措辞所产生的影响。我们研究了作为类别式相关性评判者的 GPT-4o,以及一个专门为相关性预测训练的 Qwen-7B 模型。广告主与语言实验使用了从真实广告日志中采样的查询—广告对。我们使用受控的合成查询来研究就业、住房和信贷领域中的人口统计关联。对于这两个模型而言,改变广告主身份或输入语言都可能改变相关性评估结果。部分选定的群体比较也显示出与常见刻板印象相一致的模式,尤其是涉及性别与职业的刻板印象。我们进一步研究了在模型推理与训练过程中的缓解措施。结果表明,其有效性取决于广告主信息是否与查询相关,以及广告主标签在训练数据中的分布情况。这些发现可以帮助广告从业者识别公平性风险,并为 LLM 相关性系统开发合适的缓解方法。
cs.AI / 10 / 2610.07580
LOGIC: An LLM Benchmark for Intent-Grounded Change Impact in Aerospace Electrical Systems
LOGIC:一个用于航空航天电气系统中基于意图的变更影响的 LLM 基准
large language model
大语言模型相关
Abstract
Aerospace electrical-design revisions can contain multiple genuine changes, although an engineering request may authorize only a subset. Propagating every detected difference can therefore produce overly broad impact reports. We present LOGIC, a controlled benchmark and evaluation framework in which locally deployable language models ground a request in a deterministic candidate-change inventory before selected changes are propagated through a typed electrical traceability graph. This separation permits candidate-selection errors to be distinguished from downstream propagation errors. LOGIC contains 168 scenarios, including 144 selection and 24 abstention cases. We evaluate three 7--8B models against intent-agnostic, lexical, and structured-evidence methods, with an oracle-root upper bound. On 96 explicitly anchored selection cases, gate-only structured evidence achieves candidate F1 of 1.0000, compared with 0.9677 for token-lexical matching. On 12 relational-paraphrase cases, token-lexical F1 is 0.1772 and gate-only F1 is 0.0000, compared with 0.5000--0.6400 for the large language models. Model grounding degrades as candidate inventories grow from 4 to 64 changes, while affected-element and typed-path accuracy remain comparatively stable when frozen selections are replayed over graphs of approximately 1K to 100K nodes. Strict evidence gating suppresses false positives but can remove correct semantic selections. An exploratory evidence-empty abstention policy raises strict abstention accuracy to 0.6667 for all three models and reduces unsafe-report rates to 0.1667, while decreasing answerable-case coverage by 16.0--27.1 percentage points. Four of six conflicting requests remain unsafe for each model. These findings support combining literal evidence and language-model reasoning with engineering review when intent cannot be established reliably.
Chinese Translation
航空航天电气设计修订可能包含多个真实变更,尽管一项工程请求可能仅授权其中一个子集。因此,传播每一个检测到的差异可能产生过于宽泛的影响报告。我们提出 LOGIC,一个受控基准与评估框架,其中可本地部署的语言模型在选定变更通过带类型的电气可追溯性图传播之前,将请求锚定在确定性的候选变更清单中。这种分离使得候选选择错误能够与下游传播错误区分开来。LOGIC 包含 168 个场景,包括 144 个选择案例和 24 个弃权案例。我们评估三个 7--8B 模型,对照与意图无关的方法、词汇方法和结构化证据方法,并带有一个 oracle 根上界。在 96 个显式锚定的选择案例上,仅门控的结构化证据取得 1.0000 的候选 F1,而词元词汇匹配为 0.9677。在 12 个关系释义案例上,词元词汇 F1 为 0.1772,仅门控 F1 为 0.0000,而大语言模型为 0.5000--0.6400。随着候选清单从 4 个变更增加到 64 个变更,模型锚定性能下降,而当冻结的选择在约 1K 到 100K 个节点的图上重放时,受影响元素和带类型路径的准确率保持相对稳定。严格的证据门控抑制假阳性,但可能移除正确的语义选择。一项探索性的证据为空弃权策略将三个模型的严格弃权准确率提高到 0.6667,并将不安全报告率降低到 0.1667,同时将可回答案例覆盖率降低 16.0--27.1 个百分点。对于每个模型,六个冲突请求中有四个仍不安全。这些发现支持在无法可靠地确定意图时,将字面证据和语言模型推理与工程审查相结合。
cs.AI / 11 / 2610.07606
VALSE: Vertical Adaptive Layer Skipping for Efficient Inference in Large Language Models
VALSE:面向大语言模型高效推理的垂直自适应层跳过
large language model
大语言模型相关
Abstract
This paper establishes a theoretical framework for vertical adaptive layer skipping, proving three foundational results: (i) an Expected FLOPs formula (theorem 2) giving a closed-form expression for the computational cost of arbitrary per-sample skip schedules as a function of layer-wise skip probabilities; (ii) function-space superset (theorem 10) and strict inclusion (theorem 11) theorems showing that skip-layer models are strictly contained in---yet meaningfully approximate---the full-layer function space, with an explicit separating example; and (iii) a structural duality between VALSE and Mixture-of-Experts architectures (proposition 6), positioning vertical depth-wise sparsity as the orthogonal counterpart to horizontal width-wise sparsity. Building on this theory, we propose VALSE (Vertical Adaptive Layer Skipping for Efficiency), a per-sample, non-contiguous layer skipping method: a lightweight difficulty estimator scores each input from the first few layers, and per-layer gates selectively skip redundant layers---including arbitrary middle layers while retaining deeper ones---so that only the necessary depth is activated for each input, whose feasibility is preliminarily assessed at prototype scale.
Chinese Translation
本文建立了一个面向垂直自适应层跳过的理论框架,证明了三个基础性结果:(i)一个期望 FLOPs 公式(定理 2),它给出了任意逐样本跳过调度方案的计算开销的闭式表达式,该表达式是逐层跳过概率的函数;(ii)函数空间超集(定理 10)与严格包含(定理 11)定理,表明跳层模型被严格包含于——却能有意义地逼近——全层函数空间,并给出了一个显式的区分性例子;(iii)VALSE 与专家混合(Mixture-of-Experts)架构之间的结构对偶性(命题 6),将垂直方向上的深度稀疏性定位为水平方向上的宽度稀疏性的正交对应物。在此理论基础上,我们提出 VALSE(Vertical Adaptive Layer Skipping for Efficiency,面向效率的垂直自适应层跳过),一种逐样本的、非连续层跳过方法:一个轻量级难度估计器根据前几层对每个输入进行打分,而逐层门控选择性地跳过冗余层——包括任意中间层,同时保留更深的层——从而对每个输入只激活必要的深度,其可行性已在原型规模上进行了初步评估。
cs.AI / 12 / 2610.07620
Explore, Then Commit: Measurement-Efficient Scientific Law Discovery with Language Models
先探索,后承诺:基于语言模型的测量高效科学定律发现
large language model
大语言模型相关
Abstract
Scientific law discovery requires selecting measurements and converting evidence into a governing equation. We evaluate an explore-then-commit protocol in which a large language model proposes hypotheses, a programmatic planner gathers measurements, and a fresh prompt synthesizes the final law from fixed observations. The protocol combines structured probes, automatic numerical diagnostics, restricted measurement batches, and optional interpreter access. Across 576 NewtonBench trials, we compare eight configurations on 12 physics modules using GPT-4.1-mini and a medium-difficulty GPT-4.1 replication. On medium tasks, interpreter-enabled planners use 8.6 versus 22.5 measurements per trial for GPT-4.1-mini and 8.9 versus 43.0 for GPT-4.1. Their mean magnitude-based root-mean-squared logarithmic error falls from 2.514 to 0.202 and from 0.626 to 0.149, respectively. An additional audit retains incomplete and invalid submissions in a coverage-sensitive analysis. Observed symbolic-accuracy gains are less consistent across modules, and random acquisition is competitive with disagreement scoring. Measurement savings occur in every module, but unequal batch constraints prevent attributing them solely to acquisition quality. These results support the complete protocol as a promising measurement-efficient configuration, while leaving its causal components and generalization beyond noiseless direct-equation tasks unresolved.
Chinese Translation
科学定律发现需要选择测量,并将证据转化为支配方程。我们评估一种先探索后承诺协议,其中大型语言模型提出假设,程序化规划器收集测量,而一个新的提示从固定观测中综合出最终定律。该协议结合了结构化探针、自动数值诊断、受限测量批次以及可选的解释器访问。在576次NewtonBench试验中,我们使用GPT-4.1-mini以及一个中等难度的GPT-4.1重复实验,在12个物理模块上比较了八种配置。在中等任务上,对于GPT-4.1-mini,启用解释器的规划器每次试验使用8.6次测量,而对照为22.5次;对于GPT-4.1,则分别为8.9次和43.0次。它们基于幅值的均方根对数误差均值分别从2.514降至0.202,并从0.626降至0.149。一项额外审计在覆盖率敏感分析中保留了不完整和无效的提交。观察到的符号精度提升在不同模块间不太一致,并且随机采集与分歧评分相比具有竞争力。每个模块都出现了测量节省,但不相等的批次约束使得无法将其单独归因于采集质量。这些结果支持完整协议作为一种有前景的测量高效配置,同时其因果组件以及在无噪声直接方程任务之外的泛化能力仍未得到解决。
cs.AI / 13 / 2610.07661
Massive Activation Gating Channel in Large Language Models
大型语言模型中的大规模激活门控通道
large language model
大语言模型相关
Abstract
Massive activations, a phenomenon in which a small number of hidden channels exhibit exceptionally large magnitudes, are pervasive in large language models (LLMs). However, the mechanism by which a token develops massive activations as it propagates through a pretrained LLM remains poorly understood. In this paper, we find that the emergence of massive activations is controlled by a single channel in the input embedding to a spike feed-forward network (FFN). The position of this channel is fixed for a particular LLM. We name this channel the massive activation gating channel (MAGC). When the value of the MAGC is sufficiently large (or small, depending on the LLM), the output of the spike FFN exhibits massive activations. Examining six LLMs across four model families and different model sizes, we verify the existence and effect of MAGC. We further provide a theoretical explanation of the mechanism by which MAGC induces massive activations. When the value of MAGC is sufficiently large (or small), the output of a spike FFN asymptotically reduces to a quadratic form that mixes a few columns of the down-projection matrix of the FFN. Since these columns exhibit the shape of massive activations, the output therefore exhibits massive activations.
Chinese Translation
大规模激活,即少数隐藏通道表现出异常大的幅值这一现象,在大型语言模型(LLMs)中普遍存在。然而,一个 token 在通过预训练 LLM 传播时如何产生大规模激活的机制仍鲜为人知。在本文中,我们发现大规模激活的出现由尖峰前馈网络(FFN)的输入嵌入中的单个通道控制。对于特定的 LLM,该通道的位置是固定的。我们将该通道命名为大规模激活门控通道(MAGC)。当 MAGC 的值足够大(或足够小,取决于 LLM)时,尖峰 FFN 的输出会表现出大规模激活。通过考察四个模型系列、不同模型规模的六个 LLM,我们验证了 MAGC 的存在及其作用。我们进一步对 MAGC 诱导大规模激活的机制给出了理论解释。当 MAGC 的值足够大(或足够小)时,尖峰 FFN 的输出渐近地约化为一个二次型,该二次型混合了 FFN 下投影矩阵的少数几列。由于这些列呈现出大规模激活的形状,因此输出表现出大规模激活。
cs.AI / 14 / 2610.07787
OOPMAS: Object-Oriented Multi-Agent Systems for Query-Level Workflow Generation
OOPMAS:面向查询级工作流生成的面向对象多智能体系统
large language model
大语言模型相关
Abstract
Multi-agent systems (MAS) powered by large language models have shown strong performance across code generation, mathematical reasoning, and question answering. However, existing methods for automating MAS design mostly operate at the task level, producing a single fixed workflow per benchmark that is applied uniformly to all queries. This assumption fails under realistic conditions. Query difficulty varies widely within a task, and real-world workloads mix heterogeneous task types. We introduce OOPMAS, a training-free framework that generates both the agent set and the coordination workflow at the granularity of individual queries. Agents are represented as object-oriented class definitions with dedicated roles, tools, and persistent state, and workflows are expressed as executable main functions over these agent objects. A dynamic skill library accumulates structured lessons from execution feedback across optimization rounds, enabling in-context improvement without any gradient updates or fine-tuning. On a mixed-task benchmark of queries spanning code, math, and QA, OOPMAS achieves 89.6% accuracy, outperforming the strongest baseline by 18.1 percentage points. A model-swap study across four LLM backbones shows consistent scaling, reaching 92.4% with the strongest model.
Chinese Translation
由大语言模型驱动的多智能体系统(MAS)在代码生成、数学推理和问答方面已展现出强劲性能。然而,现有用于自动化 MAS 设计的方法大多在任务层面运行,针对每个基准生成一个固定工作流,并将其统一应用于所有查询。这一假设在现实条件下并不成立。任务内部不同查询的难度差异很大,而真实世界的工作负载混合了异构的任务类型。我们提出了 OOPMAS,这是一个无需训练的框架,它以单个查询为粒度同时生成智能体集合与协调工作流。智能体被表示为面向对象的类定义,具有专门的职责、工具和持久状态,而工作流则表示为作用于这些智能体对象之上的可执行主函数。一个动态技能库在多个优化轮次中从执行反馈中积累结构化经验,从而在无需任何梯度更新或微调的情况下实现上下文内改进。在涵盖代码、数学和问答的混合任务查询基准上,OOPMAS 达到了 89.6% 的准确率,比最强基线高出 18.1 个百分点。在四种 LLM 主干模型上进行的模型替换研究表明其具有一致的扩展性,使用最强模型时达到 92.4%。
cs.AI / 15 / 2610.07791
Illusory Pattern Perception Drives Spurious Inference in Large Language Models
错觉模式感知驱动大语言模型中的虚假推断
large language model
大语言模型相关
Abstract
Illusory pattern perception is a well-documented human cognitive tendency to infer meaningful relationships in data that is actually random. Such a tendency, often described as "connecting the dots" where none exist, can result in systematic reasoning errors. This paper investigates whether Large Language Models (LLMs) exhibit such perceptual tendencies, which can lead to systematic errors in downstream applications. To our knowledge, this work presents the first systematic study of illusory pattern perception in LLMs, adapting classic psychological paradigms to three tasks with direct empirical comparison to human behaviors. We find that LLMs frequently exhibit stronger illusory pattern perception than humans. In particular, models tend to over-associate frequent positive attributes with majority groups or large organizations, and show increased tendencies to construct causal narratives from ambiguous events. To uncover the mechanism behind these behaviors, we develop a feature interpretability framework based on Sparse Autoencoders (SAEs) to analyze internal representations. Our results reveal that holistic frequency perception and analytic cognitive orientation are linked to the emergence of illusory perceptions. These findings highlight a previously underexplored cognitive-like illusion that may affect the reliability of LLM reasoning. Code available at https://github.com/NusIoraPrivacy/illusory.
Chinese Translation
错觉模式感知是一种有充分记载的人类认知倾向,即在实际上随机的数据中推断有意义的关系。这种倾向通常被描述为在不存在点的地方“连接点”,可能导致系统性推理错误。本文研究大语言模型(LLMs)是否表现出此类感知倾向,这种倾向可能导致下游应用中的系统性错误。据我们所知,本工作首次对 LLMs 中的错觉模式感知进行系统研究,将经典心理学范式改编为三项任务,并与人类行为进行直接实证比较。我们发现 LLMs 经常表现出比人类更强的错觉模式感知。特别地,模型倾向于将频繁出现的正面属性过度关联到多数群体或大型组织,并表现出从模糊事件中构建因果叙事的更强倾向。为揭示这些行为背后的机制,我们开发了一个基于稀疏自编码器(SAEs)的特征可解释性框架,以分析内部表示。我们的结果揭示,整体频率感知和分析性认知取向与错觉感知的出现有关。这些发现突显了一种此前未充分探索的类认知错觉,它可能影响 LLM 推理的可靠性。代码可在 https://github.com/NusIoraPrivacy/illusory 获取。
cs.AI / 16 / 2610.07798
Thin Evidence, Thick Priors: How Language Models Substitute Identity for Missing Financial Facts
薄证据,厚先验:语言模型如何在金融事实缺失时以身份替代之
large language model
大语言模型相关
Abstract
People increasingly ask large language models what to do with their money, yet seldom describe their finances in full. This paper asks what a model does with the gap. Holding finances fixed and changing only who the investor is said to be, we grade the financial evidence in the prompt from eight facts to none and measure how far the recommended equity allocation moves. Across 96,600 prompts to Llama-3.1-8B-Instruct, built from 100 financial profiles, 138 personas and seven disclosure conditions, the average gap between two personas with identical finances rises from 4.78 percentage points at full disclosure to 10.34 points with no financial facts. A two-way cluster bootstrap counting duplicated prompts once places the ratio at 2.16 (95% interval 1.69 to 2.79), and the rise is already 1.69-fold with a single fact left. Identity explains 5% of within-profile variation in advice at full disclosure and 96% with no disclosure. Household size is the only attribute whose influence grows reliably as evidence is withdrawn. Once standard errors are clustered on the persona, the unit to which identity was assigned, most attribute-specific interactions reported in the conference version lose significance, and gender instead appears as a small standing gap that full disclosure does not close. Stating risk appetite alone brings the swing into the range seen with two to seven generic facts. With no facts, the model's one-line rationale cites incomes, debts and savings it was never told, and these invented finances turn adverse more often for larger households. Inside the network, gender is linearly decodable at every layer, and ablating the gender direction at five layers leaves the aggregate identity swing unchanged. Advisory systems built on such models should be audited at the disclosure levels users actually reach, and judged across the whole identity space rather than one attribute at a time.
Chinese Translation
人们越来越多地向大型语言模型询问该如何处理自己的钱财,却很少完整描述自己的财务状况。本文追问的是,模型会如何填补这一空白。在保持财务状况不变、仅改变所声称的投资者身份的情况下,我们将提示中的金融证据从八项事实逐级削减至零,并测量所推荐的股票配置比例移动了多远。在向 Llama-3.1-8B-Instruct 发出的 96,600 条提示中(这些提示由 100 个财务画像、138 个人格设定和七种披露条件构建而成),财务状况相同的两个人格设定之间的平均差距从完全披露时的 4.78 个百分点上升到没有任何金融事实时的 10.34 个百分点。一种将重复提示只计一次的二维聚类自助法给出的比值为 2.16(95% 区间为 1.69 至 2.79),而仅剩一项事实时升幅就已达到 1.69 倍。身份在完全披露时解释了同一画像内部建议变异的 5%,而在无披露时解释了 96%。家庭规模是唯一随着证据被撤除而其影响仍可靠增长的属性。一旦将标准误聚类到人格设定——即身份被赋予的那个单元——上,会议版本中报告的多数属性特定交互作用便失去了显著性,而性别转而表现为一个小的、持续存在的差距,完全披露也无法将其消除。仅陈述风险偏好,就能使这一摆动落入与二至七项一般性事实相伴时所见到的区间。在没有任何事实的情况下,模型那仅有一行的理由说明中援引了它从未被告知的收入、债务和储蓄,而这些凭空编造的财务状况在家庭规模较大时更常转向不利。在网络内部,性别在每一层都可被线性解码,而在五个层上消融性别方向后,总体身份摆动保持不变。基于此类模型构建的咨询系统,应当在用户实际能够达到的披露水平上接受审计,并且应针对整个身份空间而非逐个属性加以评判。
cs.AI / 17 / 2610.07851
RA-MoWE: Workflow-Affinity Embeddings for Query Clustering and Agentic Workflow Generation
RA-MoWE:用于查询聚类和智能体工作流生成的工作流亲和度嵌入
large language model
大语言模型相关
Abstract
Agentic workflows enable large language models (LLMs) to solve complex tasks by coordinating reasoning, tool use, and verification. However, a workflow optimized for an entire task collection can overlook differences in the reasoning strategies that individual queries need, while searching for a new workflow for every query repeats costly optimization. To address this tradeoff, we introduce RA-MoWE, a framework that uses workflow-affinity embeddings to cluster queries and guide the generation of reusable expert workflows. Each embedding records how well a fixed set of reference workflows solves a query, revealing similarities in which reasoning strategies are effective. RA-MoWE uses each cluster's queries and average embedding to initialize and refine a specialized workflow through execution feedback. An embedding encoder predicts these embeddings from query text, allowing new queries to select a generated expert without first executing the reference workflows. On a 300-query test set drawn from four benchmarks spanning mathematics, science, and programming, RA-MoWE improves average task score by 4.04 percentage points over selecting among the reference workflows, while using 27.7% fewer language-model calls at inference.
Chinese Translation
智能体工作流使大型语言模型(LLMs)能够通过协调推理、工具使用和验证来解决复杂任务。然而,针对整个任务集合优化的工作流可能会忽略各个查询所需的推理策略差异,而为每个查询搜索新工作流又会重复昂贵的优化。为了解决这一权衡,我们提出 RA-MoWE,一个使用工作流亲和度嵌入来对查询进行聚类并指导可复用专家工作流生成的框架。每个嵌入记录一组固定的参考工作流解决某个查询的效果如何,从而揭示哪些推理策略有效的相似性。RA-MoWE 使用每个聚类的查询和平均嵌入,通过执行反馈来初始化和细化一个专门的工作流。一个嵌入编码器从查询文本预测这些嵌入,使新查询无需先执行参考工作流即可选择一个生成的专家。在从涵盖数学、科学和编程的四个基准中抽取的 300 个查询测试集上,RA-MoWE 相比在参考工作流中进行选择,将平均任务得分提高了 4.04 个百分点,同时在推理时减少了 27.7% 的语言模型调用。
cs.AI / 18 / 2610.07895
Textual Environmental Context and Spatial Graphs for LLM-Based Regional SST Forecasting
面向基于大语言模型的区域海表温度(SST)预测的文本环境上下文与空间图
large language model
大语言模型相关
Abstract
Sea surface temperature (SST) forecasting depends on local temporal persistence, regional spatial dependence, and environmental conditions that evolve with the forecast date. We study how these heterogeneous conditions can be presented to a large language model (LLM) for regional multi-step forecasting without serializing the full SST grid as text. We formulate forecasting as conditional numerical generation: historical SST and anomaly sequences, date-aligned environmental records, and static ocean knowledge form a textual context, while regional spatial state is supplied through continuous graph-derived prefixes. A static graph encodes persistent geographic--climatological relations, and a dynamic graph encodes recent SST correlations and localized tropical-cyclone influence. Two graph neural networks produce a target-node representation that is mapped by a spatial-prefix fusion and injected into the LLM input. On SST forecasting in the South China Sea, the complete configuration achieves the best MAE and $\Rtwo$ among the compared methods over ten forecast steps. Alongside the numerical forecast, a rule-based module matches predicted trends and environmental-factor directions with knowledge entries to return source-linked, post-hoc contextual explanations.
Chinese Translation
海表温度(SST)预测依赖于局部时间持续性、区域空间依赖性,以及随预测日期演变的环境条件。我们研究如何将这些异质条件呈现给大语言模型(LLM),以进行区域多步预测,而无需将完整的SST网格序列化为文本。我们将预测表述为条件数值生成:历史SST与异常序列、按日期对齐的环境记录以及静态海洋知识构成文本上下文,而区域空间状态则通过连续的由图导出的前缀提供。静态图编码持久的地理—气候关系,动态图编码近期的SST相关性以及局地热带气旋影响。两个图神经网络生成目标节点表示,该表示通过空间前缀融合进行映射,并被注入到LLM输入中。在南海的SST预测中,完整配置在十个预测步长上在所比较方法中取得了最佳的MAE和$\Rtwo$。除数值预测之外,一个基于规则的模块将预测趋势和环境因素方向与知识条目进行匹配,以返回带来源链接的事后上下文解释。
cs.AI / 19 / 2610.08095
Natural Language Questions as an Interface for Knowledge Graphs: QRAKEN Graph Distillation and Semantic Self-Healing
自然语言问题作为知识图谱的接口:QRAKEN 图蒸馏与语义自愈
large language model
大语言模型相关
Abstract
Natural-language access to RDF knowledge graphs is a core Semantic Web ambition. Large language models (LLMs) have advanced Text-to-SPARQL, yet on unfamiliar graphs they often generate valid queries that misrepresent the populated data model. QRAKEN is a training-free, ontology-agnostic neurosymbolic pipeline grounding generation in empirical graph evidence rather than schema expectations. An offline distiller produces TTQL, a compact description of populated multi-hop patterns, conditional frequencies and path-conditioned literal examples, plus a class-property co-occurrence matrix. Online, TTQL guides the LLM, while deterministic syntax, vocabulary and data-model checks provide diagnostics for iterative refinement. On CK25 (First International Text2SPARQL Challenge), under matched-condition recomputation on a QLever snapshot, QRAKEN achieves strict F1 of 0.643 $\pm$ 0.026 with GPT-4.1 mini and 0.652 $\pm$ 0.012 with GPT-5.4: relative gains of 30% and 32% over the strongest recomputed participant, outperforming systems using the same base model family. Ablations identify TTQL patterns as the dominant driver (+0.31 strict F1 over a shape-only baseline); the refinement loop provides a cheap safety net, rejecting triple patterns unsupported by the co-occurrence matrix. Compared with auto-derived SHACL, TTQL yields 64% higher strict F1, supporting the value of empirical patterns beyond schema exposure. With two local 35B 4-bit open-weight models at zero marginal cost, the same pipeline matches the strongest recomputed participant, and TTQL advantages over shape-only and SHACL baselines persist. Results on a single, relatively small benchmark provide an initial empirical signal; monolithic TTQL injection on very open cross-domain graphs remains the main limitation.
Chinese Translation
对 RDF 知识图谱的自然语言访问是语义网的一个核心愿景。大语言模型(LLMs)推动了 Text-to-SPARQL 的发展,然而在陌生的图上,它们常常生成有效但却错误表示已填充数据模型的查询。QRAKEN 是一种无需训练、与本体无关的神经符号流水线,它将生成建立在经验图证据而非模式预期之上。一个离线蒸馏器生成 TTQL,这是一种对已填充的多跳模式、条件频率和路径条件化字面量示例的紧凑描述,外加一个类-属性共现矩阵。在线阶段,TTQL 引导 LLM,而确定性的语法、词汇和数据模型检查为迭代细化提供诊断信息。在 CK25(第一届国际 Text2SPARQL 挑战赛)上,在 QLever 快照上进行匹配条件重计算的情况下,QRAKEN 在使用 GPT-4.1 mini 时达到 0.643 $\pm$ 0.026 的严格 F1,在使用 GPT-5.4 时达到 0.652 $\pm$ 0.012 的严格 F1:相对于最强的重计算参赛者分别取得 30% 和 32% 的相对提升,并优于使用相同基础模型家族的系统。消融实验表明,TTQL 模式是主要驱动因素(相比仅形状基线提高 +0.31 严格 F1);细化循环提供了一个廉价的安全网,拒绝共现矩阵不支持的三元组模式。与自动推导的 SHACL 相比,TTQL 的严格 F1 高出 64%,支持了经验模式在模式暴露之外的价值。在两个本地 35B 4 比特开放权重模型上,以零边际成本,同一流水线匹配了最强的重计算参赛者,并且 TTQL 相对于仅形状和 SHACL 基线的优势仍然存在。在单个相对较小的基准上的结果提供了初步的经验信号;在非常开放的跨领域图上进行整体式 TTQL 注入仍然是主要局限。
cs.AI / 20 / 2610.08106
ChartBmkAgent: Harness-Governed Multi-Agent Construction of Chart QA Benchmarks from Sparse Error-Taxonomy Specifications
ChartBmkAgent:由管控框架治理的多智能体从稀疏错误分类体系规范构建图表问答基准
large language model
大语言模型相关
Abstract
Multimodal large language models (MLLMs) advance rapidly, while conventional benchmark development lags behind, delaying investigation of newly observed capability gaps. Such investigation requires an expressive task format and an on-demand construction process: information-rich charts make chart question answering (Chart QA) suitable for probing coupled perception and reasoning. Automated Chart QA construction is intended to shorten the benchmark-development cycle by turning identified gaps into targeted samples on demand. Current methods, however, commonly separate target guidance from scratch generation: target-guided systems often require prepared data, charts, or templates, while scratch-generation systems primarily ensure artifact validity, without explicitly controlling whether newly synthesized requirements and content remain aligned with an externally specified diagnostic target. We introduce ChartBmkAgent, which turns an identified capability gap into targeted diagnostic evidence by constructing complete Chart QA samples from sparse error-taxonomy specifications. Throughout construction, a central harness governs specialized agents, requires stage-specific evidence of alignment with the original error category, and records the basis for each acceptance decision. On 300 taxonomy-wide samples, MLLM accuracies ranged from 32.7% to 84.3% with distinct category profiles, showing that generated samples reveal capability differences. Across three source-model comparisons, targeted follow-ups scored 50.0% versus 82.2% on matched controls ($p=8.96\times10^{-6}$); all six cross-model comparisons had the same direction, demonstrating targeted validation and diagnostic-data generation. Multiple evaluator models assessed whether each sample tested its specified error category; 86.4% met this criterion, providing empirical evidence of target preservation.
Chinese Translation
多模态大语言模型(MLLMs)快速发展,而传统基准开发却滞后,从而延迟了对新观察到的能力差距的研究。此类研究需要一种富有表现力的任务形式和一种按需构建过程:信息丰富的图表使图表问答(Chart QA)适合探究耦合的感知与推理。自动化 Chart QA 构建旨在通过将已识别的差距按需转化为有针对性的样本,以缩短基准开发周期。然而,当前方法通常将目标引导与从头生成分离:目标引导系统往往需要准备好的数据、图表或模板,而从零生成系统主要确保产物有效性,并未显式控制新合成的要求与内容是否始终与外部指定的诊断目标保持一致。我们提出 ChartBmkAgent,它通过从稀疏错误分类体系规范构建完整的 Chart QA 样本来将已识别的能力差距转化为有针对性的诊断证据。在整个构建过程中,一个中心管控框架治理专门智能体,要求提供与原始错误类别保持一致的阶段特定证据,并记录每个验收决策的依据。在 300 个覆盖整个分类体系的样本上,MLLM 准确率范围为 32.7% 到 84.3%,并呈现不同的类别特征,表明生成的样本能够揭示能力差异。在三组源模型比较中,针对性后续测试得分为 50.0%,而匹配对照得分为 82.2%($p=8.96\times10^{-6}$);所有六组跨模型比较均呈现相同方向,证明了针对性验证与诊断数据生成。多个评估器模型评估了每个样本是否测试了其指定的错误类别;86.4% 满足这一标准,为目标保持提供了经验证据。
cs.AI / 21 / 2610.08231
OSFP4: Joint Optimization of Diagonal Smoothing and Block Scales for NVFP4 Quantization
OSFP4:面向 NVFP4 量化的对角平滑与块缩放因子的联合优化
large language model
大语言模型相关
Abstract
NVFP4 is an attractive datatype for large language model (LLM) inference, offering compact storage and native tensor-core acceleration. However, preserving accuracy using NVFP4 requires careful quantization. In this work we develop a novel quantization scheme called Optimized Smoothing and Scaling for NVFP4 (OSFP4). For each linear projection it uses a diagonal smoothing matrix whose entries are optimized to minimize the squared matrix-product quantization error under NVFP4, taking into account the rounding procedure that is used (either round-to-nearest, or GPTQ-style successive interference cancellation). This requires performing joint optimization on the smoothing entries as well as the block scales, which is facilitated by analyzing a multiplicative-dither FP4 quantizer instead of the fixed deterministic one. Experiments show that OSFP4 achieves the highest average accuracy among the evaluated competitors in the corresponding quantization settings, while retaining approximately 94-97\% of vendor NVFP4 prefill throughput on the measured workloads. Our code is available in https://github.com/neriahbd/OSFP4
Chinese Translation
NVFP4 是一种对大语言模型(LLM)推理很有吸引力的数据类型,它提供紧凑存储和原生张量核心加速。然而,使用 NVFP4 保持精度需要仔细的量化。在这项工作中,我们开发了一种新颖的量化方案,称为面向 NVFP4 的优化平滑与缩放(OSFP4)。对于每个线性投影,它使用一个对角平滑矩阵,其元素经过优化,以最小化 NVFP4 下的矩阵乘积量化误差平方,并考虑所使用的舍入过程(要么是舍入到最近值,要么是 GPTQ 风格的逐次干扰消除)。这需要对平滑元素以及块缩放因子进行联合优化,而这通过分析乘性抖动 FP4 量化器而不是固定的确定性量化器来促成。实验表明,在相应的量化设置中,OSFP4 在所评估的竞争者中取得了最高的平均准确率,同时在所测量的工作负载上保留了大约 94-97\% 的厂商 NVFP4 预填充吞吐量。我们的代码可在 https://github.com/neriahbd/OSFP4 获取。
cs.AI / 22 / 2610.08246
LeanPlan: Optimal Planning with LLM-Generated Heuristics and Admissibility Proofs
LeanPlan:使用LLM生成的启发式与可采纳性证明进行最优规划
large language model
大语言模型相关
Abstract
Frontier large language models (LLMs) can generate heuristic functions that guide search to achieve state-of-the-art performance in satisficing planning, where any plan is acceptable. However, these heuristics are not guaranteed to be admissible and can lead to suboptimal plans. We introduce LeanPlan, the first planning system that finds optimal plans with LLM-generated heuristics whose admissibility is machine-checked. Given a domain description and training tasks, an agentic loop uses planner feedback to iteratively improve a reusable domain-specific heuristic, its admissibility proof and the required domain assumptions. LeanPlan implements the heuristic, its proof and an efficient planner with machine-checked grounding and search in Lean 4. We evaluate LeanPlan on ten domains from the International Planning Competition and three new domains, using test tasks with up to 57 times as many objects as the training tasks. With GPT-5.6 Sol in the agentic loop, we successfully generate heuristics and admissibility proofs for all these domains. With the resulting heuristics, LeanPlan usually expands fewer states than the state-of-the-art Scorpion planner and solves more tasks overall.
Chinese Translation
前沿大语言模型(LLM)能够生成启发式函数,引导搜索在满足性规划(satisficing planning)中达到最先进的性能,其中任何规划都是可接受的。然而,这些启发式并不保证是可采纳的,并且可能导致次优规划。我们提出 LeanPlan,这是首个使用LLM生成的启发式来寻找最优规划的规划系统,这些启发式的可采纳性经过机器检验。给定一个领域描述和训练任务,一个智能体循环利用规划器反馈,迭代改进一个可复用的领域特定启发式、其可采纳性证明以及所需的领域假设。LeanPlan 在 Lean 4 中实现了该启发式、其证明以及一个具有机器检验的实例化与搜索的高效规划器。我们在国际规划竞赛(International Planning Competition)的十个领域和三个新领域上评估 LeanPlan,所使用的测试任务中对象数量最多可达训练任务的 57 倍。在智能体循环中使用 GPT-5.6 Sol,我们成功为所有这些领域生成了启发式和可采纳性证明。借助由此得到的启发式,LeanPlan 通常比最先进的 Scorpion 规划器扩展更少的状态,并且总体上解决了更多任务。
cs.AI / 23 / 2610.08250
MASC: A Multi-Agent Self-Calibration Framework with Latent Construct Alignment for Consistent Client Role-Playing in Psychological Counseling
MASC:一种具有潜在构念对齐的多智能体自校准框架,用于心理咨询中一致的来访者角色扮演
large language model
大语言模型相关
Abstract
Large language models are increasingly used to simulate clients for counselor training and psychological counseling research, but reliable simulation requires clients to remain psychologically coherent across extended interactions. Existing role-playing methods largely rely on static profile prompts and may exhibit persona drift, unrealistic cooperativeness, or inconsistent psychological states, communicative actions, and emotions. Existing evaluations also lack a unified testbed for both stable client characteristics and evolving psychological dynamics. We propose MASC, a Multi-Agent Self-Calibration framework with latent construct alignment for consistent client role-playing in psychological counseling. MASC combines construct-guided generation, collaborative refinement, consistency verification, and memory-based revision in a closed calibration loop that detects and corrects inconsistencies as dialogue unfolds. We further introduce CRPC-Bench, a benchmark covering session-level profile information and Big-Five personality traits, as well as turn-level psychological state, communicative action, and emotion expression. CRPC-Bench contains 38 motivational interviewing client profiles augmented with personality and emotion annotations. Experiments show that MASC outperforms existing methods across profile, personality, receptivity, and turn-level consistency, with the heterogeneous configuration achieving the strongest overall performance. MASC and CRPC-Bench provide a unified foundation for developing and evaluating psychologically coherent client simulations for AI-assisted counseling research and training.
Chinese Translation
大语言模型正越来越多地被用于为咨询师培训和心理咨询研究模拟来访者,但可靠的模拟要求来访者在长时间交互中保持心理一致。现有角色扮演方法在很大程度上依赖静态画像提示,可能表现出人设漂移、不切实际的合作性,或心理状态、沟通行为和情绪的不一致。现有评估也缺乏一个统一的测试平台,用于同时考察稳定的来访者特征和不断演变的心理动态。我们提出 MASC,一种具有潜在构念对齐的多智能体自校准框架,用于心理咨询中一致的来访者角色扮演。MASC 将构念引导的生成、协作式精炼、一致性验证和基于记忆的修订结合在一个闭环校准回路中,该回路在对话展开时检测并纠正不一致。我们进一步提出 CRPC-Bench,一个涵盖会话级画像信息和大五人格特质,以及轮次级心理状态、沟通行为和情绪表达的基准。CRPC-Bench 包含 38 个动机式访谈来访者画像,并附加了人格和情绪标注。实验表明,MASC 在画像、人格、接受性和轮次级一致性方面优于现有方法,其中异构配置取得了最强的整体性能。MASC 和 CRPC-Bench 为开发和评估用于 AI 辅助咨询研究与培训的心理一致来访者模拟提供了统一基础。
cs.AI / 24 / 2610.08258
zkLLMPoT: Efficient Zero Knowledge Proof of Training for Large Language Models
zkLLMPoT:面向大型语言模型的高效训练零知识证明
large language model
大语言模型相关
Abstract
Auditing the claimed outcomes of large language model (LLM) training is challenging when model weights and training data are private, while cryptographically proving the full training process is prohibitively expensive at Transformer scale. We present zkLLMPoT, a zero-knowledge framework that certifies auditor-defined properties of a trained checkpoint through forward evaluation rather than verification of its optimization trajectory. zkLLMPoT includes 2 phases: 1) The trainer fixes the architecture and the model weights are committed. Then the auditor selects challenge sequences, preventing the trainer from modifying the checkpoint in response to the audit data. 2) Then the trainer proves the objective value attained by the committed model on those sequences. This formulation makes the certification cost independent of the number of training iterations, without revealing model weights or requiring access to private training data. We build on sumcheck- and lookup-based arguments to certify Transformer computations, while supporting next-token loss and task-specific audit objectives. Across four model families, operator-level benchmarks yield proving times of 41-59 seconds for 1.1-1.5B-parameter models and 131 seconds at 13B for the covered operators, with verification below half a second at a sequence length of 512.
Chinese Translation
当模型权重和训练数据是私有的时,审计大型语言模型(LLM)训练所声称的结果具有挑战性,而在 Transformer 规模下以密码学方式证明完整训练过程则代价高得令人望而却步。我们提出 zkLLMPoT,一个零知识框架,它通过前向评估而非验证其优化轨迹来认证经训练的检查点中由审计者定义的属性。zkLLMPoT 包含 2 个阶段:1)训练者固定架构,并提交模型权重。然后审计者选择挑战序列,防止训练者根据审计数据修改检查点。2)然后训练者证明所提交模型在这些序列上达到的目标值。这种表述使认证成本与训练迭代次数无关,同时不泄露模型权重,也不要求访问私有训练数据。我们基于 sumcheck 和 lookup 的论证来认证 Transformer 计算,同时支持下一 token 损失和任务特定的审计目标。在四个模型系列中,算子级基准测试在 1.1-1.5B 参数模型上产生 41-59 秒的证明时间,在 13B 上对所覆盖的算子产生 131 秒的证明时间,而在序列长度为 512 时验证时间低于半秒。
cs.AI / 25 / 2610.08312
CoDe-LoRA: Mitigating the Orthogonality Dilemma in Continual Learning of LLMs via Knowledge Consolidation and Decoupling
CoDe-LoRA:通过知识巩固与解耦缓解LLMs持续学习中的正交性困境
large language model
大语言模型相关
Abstract
Continual learning (CL) is essential for Large Language Models (LLMs) to sequentially adapt to evolving tasks. To mitigate catastrophic forgetting, recent advances implement low-rank adaptation with orthogonal projections (e.g., O-LoRA) to isolate task parameters. However, we reveal that such strict geometric constraints trigger an "Orthogonality Dilemma": rigid parameter isolation impedes the transfer and accumulation of shared representations across semantically related tasks. In this work, we propose a new replay-free method, called Consolidation and Decoupling LoRA (CoDe-LoRA), for CL of LLMs. CoDe-LoRA disentangles the learning process into Consolidating Universal Knowledge and Decoupling Task-Specific Knowledge. To achieve this, CoDe-LoRA leverages an adaptive null space projection mechanism and semantic routing to balance knowledge accumulation with task-specific adaptation. Experimental results across four backbones and three CL benchmarks show that CoDe-LoRA achieves the best average accuracy. Our code is available at https://github.com/Estrellajer/CoDe-LoRA.
Chinese Translation
持续学习(CL)对于大语言模型(LLMs)顺序适应不断演化的任务至关重要。为了缓解灾难性遗忘,近期进展采用带正交投影的低秩适应(例如 O-LoRA)来隔离任务参数。然而,我们揭示出,这种严格的几何约束会引发“正交性困境”:僵化的参数隔离会阻碍语义相关任务之间共享表征的迁移与积累。在这项工作中,我们提出了一种新的无回放方法,称为巩固与解耦LoRA(CoDe-LoRA),用于LLMs的CL。CoDe-LoRA将学习过程解耦为巩固通用知识与解耦任务特定知识。为实现这一点,CoDe-LoRA利用自适应零空间投影机制和语义路由,以平衡知识积累与任务特定适应。在四个骨干模型和三个CL基准上的实验结果表明,CoDe-LoRA取得了最佳平均准确率。我们的代码可在 https://github.com/Estrellajer/CoDe-LoRA 获取。
cs.AI / 26 / 2610.08327
MedZERO: Self-Evolving Agents for Open-Ended Medical Reasoning Through Controlled Knowledge Accumulation
MedZERO:通过受控知识积累实现开放式医学推理的自演化智能体
large language model
大语言模型相关
Abstract
Large language models (LLMs) have shown promise in medical question answering and clinical reasoning, yet their improvement remains constrained by static parametric knowledge and costly expert supervision. Self-evolving agents offer a promising alternative by enabling models to improve through iterative task generation and problem-solving. However, most existing self-evolving methods are designed for easily verifiable domains such as mathematics and coding, where solutions can be checked by exact answers or executable programs. Medical reasoning is fundamentally different: it is open-ended, knowledge-intensive, and often only partially verifiable. We present MedZERO, a self-evolving framework for open-ended medical reasoning. MedZERO couples an Examiner that generates frontier medical question-option pairs with a Reasoner that solves them through evidence-grounded multi-turn reasoning with external knowledge tools. To support reliable, continual improvement, MedZERO adopts controlled knowledge accumulation, which maintains temporary exploratory knowledge and curated persistent knowledge in reasoning. We evaluate MedZERO on five public medical reasoning benchmarks using 4B- and 8B-scale base models under open-ended evaluation. Across all settings, MedZERO consistently outperforms the underlying base models and prior self-evolving baselines, achieving up to 13.7 average accuracy-point gains over the next-best self-evolving baseline.
Chinese Translation
大语言模型(LLMs)在医学问答和临床推理中已展现出潜力,然而其改进仍受到静态参数化知识和昂贵的专家监督的制约。自演化智能体提供了一种有前景的替代方案,通过迭代任务生成和问题求解使模型得以改进。然而,大多数现有的自演化方法是为易于验证的领域(如数学和编程)设计的,在这些领域中,解决方案可以通过精确答案或可执行程序进行检查。医学推理则根本不同:它是开放式的、知识密集型的,并且通常只能部分验证。我们提出 MedZERO,一个用于开放式医学推理的自演化框架。MedZERO 将一个生成前沿医学问题-选项对的 Examiner 与一个通过基于证据的多轮推理并借助外部知识工具来求解这些问题的 Reasoner 相结合。为支持可靠、持续的改进,MedZERO 采用受控知识积累,在推理中维护临时探索性知识和经过整理的持久性知识。我们在开放式评估下,使用 4B 和 8B 规模的基础模型,在五个公开医学推理基准上评估 MedZERO。在所有设置中,MedZERO 始终优于其底层基础模型和先前的自演化基线,相较于次优的自演化基线,平均准确率提升最高达 13.7 个百分点。
cs.AI / 27 / 2610.08330
MoF: Preference-Aware Mixture Modeling for Black-Box LLM Personalization
MoF:用于黑盒 LLM 个性化的偏好感知混合建模
large language model
大语言模型相关
Abstract
Proprietary Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks, yet aligning their outputs with diverse user preferences remains challenging. Existing personalization approaches for black-box LLMs often rely on user-specific scoring heads, causing the number of personalized parameters to grow linearly with the number of users and requiring additional adaptation for unseen users. To address these limitations, we propose Mixture-of-Facets (MoF), a scalable personalization framework for black-box LLMs that models user preferences as compositions of shared latent preference facets rather than dedicated user-specific parameters. MoF performs personalization through history-conditioned routing over shared facet heads, enabling personalization for users unseen during training without additional parameter updates. Across diverse personalization tasks, MoF delivers stronger personalization performance while maintaining a more scalable and parameter-efficient design than prior approaches. Additional analysis indicates strong generalization to unseen users.
Chinese Translation
专有大型语言模型(LLMs)已在广泛任务中展现出卓越能力,然而使其输出与多样化的用户偏好保持一致仍然具有挑战性。现有的黑盒 LLM 个性化方法通常依赖用户特定的评分头,导致个性化参数的数量随用户数量线性增长,并且需要针对未见用户进行额外适配。为解决这些局限性,我们提出 Mixture-of-Facets(MoF),一种面向黑盒 LLM 的可扩展个性化框架,它将用户偏好建模为共享潜在偏好刻面的组合,而不是专用的用户特定参数。MoF 通过基于历史条件的路由在共享刻面头上执行个性化,从而能够为训练期间未见过的用户实现个性化,而无需额外的参数更新。在多样化的个性化任务中,MoF 提供了更强的个性化性能,同时相比先前方法保持了更具可扩展性且参数更高效的设计。额外分析表明其对未见用户具有很强的泛化能力。
cs.AI / 28 / 2610.08446
AssemState: Manual and Physical-State-Guided Reasoning for Zero-shot Furniture Assembly
AssemState:面向零样本家具装配的手册与物理状态引导推理
large language model
大语言模型相关
Abstract
Multimodal large language models (MLLMs) have made significant progress in visual understanding, but precise 3D spatial reasoning integrated with physical environment remains difficult. Furniture assembly requires not only recovering step-level operations from diagrammatic manuals, but also translating semantic attachment relations into 6D pose updates that enable parts to physically interact with the environment and previously assembled components. To study this problem, we propose AssemState, a zero-shot framework for manual and physical-state-guided furniture assembly. It firstly employs anchor-guided boundary assembly states to decompose manual pages into single-part operations and recover an assembly-tree. Then, it uses iterative after-state feedback refinement to guide successive (SE(3)) updates and corrections, and validates their physical plausibility through simulation-based release tests. Experiments show that compared with the strongest prior baseline, AssemState improves F1 from 38.58\% to 62.80\% and Tree Exact Match from 28.24\% to 53.92\% for assembly-tree recovery. On 243 independently evaluated part-level operations, our proposed iterative refinement improves judge-accepted operations from 0 to 5.3\% and reduces mean Chamfer distance from 5.4111 to 1.7744. However, visually plausible candidate poses may still suffer from collision, floating, mirror-orientation errors, incomplete seating, and wrong-side attachment. These results show that AssemState improves operation-structure recovery and selected local pose metrics, while MLLMs remain limited for spatial relationship reasoning.
Chinese Translation
多模态大语言模型(MLLMs)在视觉理解方面已取得显著进展,但与物理环境集成的精确 3D 空间推理仍然困难。家具装配不仅需要从图解手册中恢复步骤级操作,还需要将语义附着关系转化为 6D 位姿更新,使部件能够与环境以及先前已装配的组件进行物理交互。为研究这一问题,我们提出 AssemState,一个用于手册与物理状态引导的家具装配的零样本框架。它首先采用锚点引导的边界装配状态,将手册页面分解为单部件操作,并恢复一棵装配树。然后,它使用迭代式的后状态反馈细化来引导连续的 (SE(3)) 更新与修正,并通过基于仿真的释放测试验证它们的物理合理性。实验表明,与最强的先前基线相比,AssemState 在装配树恢复方面将 F1 从 38.58\% 提升到 62.80\%,并将树精确匹配(Tree Exact Match)从 28.24\% 提升到 53.92\%。在 243 个独立评估的部件级操作上,我们提出的迭代细化将评判器接受的操作从 0 提高到 5.3\%,并将平均 Chamfer 距离从 5.4111 降低到 1.7744。然而,视觉上看似合理的候选位姿仍可能遭受碰撞、悬浮、镜像朝向错误、未完全就位以及错误侧附着等问题。这些结果表明,AssemState 改进了操作-结构恢复以及选定的局部位姿指标,而 MLLMs 在空间关系推理方面仍然存在局限。
cs.AI / 29 / 2610.08563
Adaptive Power Sampling for LLM Reasoning
面向LLM推理的自适应幂采样
large language model
大语言模型相关
Abstract
Sequence-level power sampling has recently emerged as a training-free approach to reasoning by sampling from a sharpened output distribution of a base large language model (LLM). Nevertheless, existing methods typically sharpen the base model distribution uniformly across queries, overlooking variations in query difficulty and in how well the base model already handles each query. The goal of this work is to equip power sampling with query adaptivity. Theoretically, we show that the benefits of further sharpening are determined by the self-reward gap between correct and incorrect responses. Based on this insight, we propose \emph{Adaptive Power Sampling} (APS), which adjusts the sharpening exponent on a per-query basis at test time using the relationship between answer agreement and the model's self-reward. Experiments across diverse reasoning tasks, including MATH500, HumanEval, and GPQA, show that APS consistently outperforms power sampling with a fixed sharpening exponent, without additional training.
Chinese Translation
序列级幂采样最近作为一种免训练的推理方法而出现,其做法是从基础大语言模型(LLM)的锐化输出分布中采样。然而,现有方法通常对所有查询统一地锐化基础模型分布,忽略了查询难度上的差异,以及基础模型对每个查询原本处理得好坏程度上的差异。本工作的目标是使幂采样具备针对查询的自适应性。在理论上,我们表明进一步锐化所带来的收益由正确回答与错误回答之间的自奖励差距(self-reward gap)所决定。基于这一洞见,我们提出 \emph{Adaptive Power Sampling}(APS),它在测试时利用答案一致性(answer agreement)与模型自奖励之间的关系,逐查询地调整锐化指数。在包括 MATH500、HumanEval 和 GPQA 在内的多种推理任务上的实验表明,APS 无需额外训练即可持续优于使用固定锐化指数的幂采样。
cs.AI / 30 / 2610.08586
MINDSET: Energy-based Schema Evolution for Long Conversational Agent Memory
MINDSET:面向长对话智能体记忆的基于能量的模式演化
large language model
大语言模型相关
Abstract
Long conversational agents have become essential in our daily lives. They must remember what was said long back in order to help us efficiently complete a task without needing the user to repeat instructions and context repeatedly. However, the main issue is that instructions and context change over time and so the agents must be able to adapt accordingly. A useful memory system should preserve both current and historical states, distinguish stale information from active knowledge, retrieve evidence appropriate to the query and avoid repeatedly invoking a large language model to rewrite prior interactions. We introduce MINDSET, a memory controller that stores a conversation as immutable episodes and organizes them into versioned schemas through minimum-energy state transitions. Each incoming episode may reinforce, supersede, split or create a schema. The transition decision balances representation distortion, contradiction, historical damage, fragmentation and internal inconsistency, while hysteresis prevents isolated contradictions from prematurely rewriting stable memory. We evaluate MINDSET against 5 memory systems on a reproducible sample of 850 questions (700 LoCoMo + 150 MemoryAgentBench). MINDSET obtains the highest observed LoCoMo answer F1 while significantly improving retrieval ranking (Recall@8, MRR and nDCG@8) over the second best method LightMem (p<0.01 after Holm correction). It obtains the highest observed scores on MemoryAgentBench although the relative difference is low. Ablations identify controlled fragmentation and schema-aware assignment as the largest contributors to answer quality. Additionally, a 700-question cross-model evaluation with GLM-4.7 and Gemma-4-31B supported model independence. These results show that long-term memory can be better handled as constrained state management rather than continual summarization.
Chinese Translation
长对话智能体已在我们的日常生活中变得不可或缺。它们必须记住很久之前说过的话,以便帮助我们高效地完成一项任务,而无需用户反复重复指令和上下文。然而,主要问题在于指令和上下文会随时间而变化,因此智能体必须能够相应地适应。一个有用的记忆系统应当同时保留当前状态与历史状态,区分过时信息与活跃知识,检索与查询相适配的证据,并避免反复调用大语言模型来重写先前的交互。我们提出 MINDSET,一种记忆控制器,它将一段对话存储为不可变的片段,并通过最小能量状态转移将它们组织成带版本的模式。每个新到来的片段都可能强化、取代、拆分或创建一个模式。该转移决策在表示失真、矛盾、历史损伤、碎片化和内部不一致之间进行权衡,而迟滞机制则防止孤立的矛盾过早地重写稳定的记忆。我们在一个可复现的、包含 850 个问题的样本(700 个 LoCoMo + 150 个 MemoryAgentBench)上,将 MINDSET 与 5 个记忆系统进行了评估。MINDSET 取得了所观测到的最高 LoCoMo 答案 F1,同时在检索排序(Recall@8、MRR 和 nDCG@8)上显著优于排名第二的方法 LightMem(经 Holm 校正后 p<0.01)。尽管相对差异较小,它在 MemoryAgentBench 上仍取得了所观测到的最高分数。消融实验表明,受控碎片化与模式感知分配是答案质量的最大贡献因素。此外,使用 GLM-4.7 和 Gemma-4-31B 进行的 700 个问题的跨模型评估支持了模型独立性。这些结果表明,长期记忆可以更好地被处理为受约束的状态管理,而不是持续不断的摘要生成。
cs.AI / 31 / 2610.08691
ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences
ScienceClaw:对自然科学与社会科学领域中 AI-for-Science 智能体的持续自演化进行基准测试
large language model
大语言模型相关
Abstract
Large language model agents are accelerating scientific automation, yet verified executions rarely become persistent program-level improvements, and existing evaluations do not examine this process across sequential tasks in both the natural and social sciences. We formalize ScienceClaw as fixed-parameter program self-evolution that unifies task solving, scientific verification, and program updates. ScienceClaw-Eval spans 23 disciplines and measures scientific correctness, evolutionary gain, retention, cross-dataset transfer, and evolution cost through sequential streams and independent reset evaluation. Our framework repairs executable workflows through multi-turn interaction, converts re-execution-verified failure--success trajectories into linked Skill and Operator candidates, and retains an update only when source-task replay reproduces the repair and independent scientific tasks improve. Code is available at https://github.com/beita6969/ScienceClaw.
Chinese Translation
大语言模型智能体正在加速科学自动化,然而经过验证的执行很少成为持久的程序级改进,而且现有评估并未考察这一过程在自然科学与社会科学二者中的顺序任务上的表现。我们将 ScienceClaw 形式化为固定参数的程序自演化,它统一了任务求解、科学验证与程序更新。ScienceClaw-Eval 涵盖 23 个学科,并通过顺序流与独立重置评估来测量科学正确性、演化增益、保持能力、跨数据集迁移和演化成本。我们的框架通过多轮交互修复可执行工作流,将经重新执行验证的失败--成功轨迹转换为相互关联的 Skill 与 Operator 候选,并且仅当源任务重放能够复现该修复且独立科学任务有所改进时,才保留一次更新。代码可在 https://github.com/beita6969/ScienceClaw 获取。
cs.AI / 32 / 2610.08775
Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?
瓶中之智体:LLM 智能体能否将其能力转化为廉价、可扩展的人工制品?
large language model
大语言模型相关
Abstract
Large language models (LLMs) can solve many narrow tasks, but querying them separately for millions of related instances can be prohibitively expensive. Can LLM agents autonomously create cheaper solutions for such workloads? We call this ability "bottling": the ability to turn general capabilities into task-specific solutions that balance answer quality and amortised cost. We introduce BOTTLED, a benchmark in which agents receive an entire unlabelled workload and must complete it under fixed time, compute and LLM API budgets. Agents choose their own approach, such as training a small model or writing a reusable program. Across ten models and three tasks, we find that strong zero-shot task performance does not reliably translate into strong bottling capabilities. Models with similar zero-shot scores can differ substantially after bottling, and 48 of 60 bottling runs score below the lower bound of the 95% confidence interval of their model's zero-shot performance. Moreover, 31 of 60 runs underperform the stronger of two small-model distillation baselines with the same token budget. Nevertheless, bottling can yield substantial savings: on query-product relevance classification, Opus 5 retains about 82% of its zero-shot macro-F1 at roughly 657 times lower reported cost. Bottling is also competitive with Jev, a "system one" model built especially for cheap, repetitive inference: Opus 5 on the same task recovers about 94% of Jev's macro-F1 at a quarter of Jev's projected full-workload cost. BOTTLED provides a basis for evaluating and improving agents' ability to invest limited resources in reusable solutions for large, repetitive workloads.
Chinese Translation
大型语言模型(LLM)可以解决许多狭窄的任务,但为数百万个相关实例分别查询它们可能代价高得令人望而却步。LLM 智能体能否自主为这类工作负载创造出更廉价的解决方案?我们将这种能力称为“瓶装”(bottling):即把通用能力转化为任务专用解决方案的能力,这些解决方案能在答案质量与摊销成本之间取得平衡。我们提出了 BOTTLED,这是一个基准,其中智能体接收整个未标注的工作负载,并且必须在固定的时间、计算和 LLM API 预算下完成它。智能体自行选择其方法,例如训练一个小模型或编写一个可复用的程序。在十个模型和三项任务上,我们发现强大的零样本任务性能并不能可靠地转化为强大的瓶装能力。零样本得分相近的模型在瓶装之后可能差异巨大,并且 60 次瓶装运行中有 48 次的得分低于其模型零样本性能 95% 置信区间的下界。此外,在相同的 token 预算下,60 次运行中有 31 次的表现不及两个小模型蒸馏基线中更强的那个。尽管如此,瓶装仍能带来可观的节省:在查询—商品相关性分类任务上,Opus 5 在报告成本大约低 657 倍的情况下,保留了其零样本 macro-F1 的约 82%。瓶装也足以与 Jev 相竞争,Jev 是一个专为廉价、重复性推理而构建的“系统一”模型:在同一任务上,Opus 5 以 Jev 预测的全工作负载成本的四分之一,达到了 Jev 的 macro-F1 的约 94%。BOTTLED 为评估和改进智能体将有限资源投入到面向大规模、重复性工作负载的可复用解决方案中的能力提供了基础。
cs.AI / 33 / 2610.08778
Sherpa: Teaching LLMs to Teach Adaptively
Sherpa:教会大语言模型自适应地教学
large language model
大语言模型相关
Abstract
Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, preference data, or predefined pedagogical criteria that specify what good teaching looks like. However, these signals are often not grounded in individual student learning outcomes, where effective teaching strategies can vary substantially across learners. To address this, we introduce Sherpa, a multi-turn reinforcement learning framework that instantiates multiple student archetypes with LLMs conditioned on distinct learning preferences and trains a teacher model to adapt its instruction by directly maximizing their learning outcomes. Teacher LLMs trained with Sherpa improve instructed students' performance across all archetypes by an average of 20.5 percentage points. Under MathTutorBench's evaluation, Sherpa raises the overall pedagogy score from 52.5% to 79.2%, indicating better teaching responses. Our human studies show that the trained teacher is preferred over the base model in 79.6% of pairwise comparisons. Together, Sherpa trains LLM teachers to adapt to diverse simulated students and become better aligned with human teachers, paving the road towards AI tutors teaching real students.
Chinese Translation
大语言模型(LLMs)已成为越来越有能力的问题解决者,但能够解决一个问题并不等同于能够教授它。现有的将 LLMs 训练为教师的方法依赖于示范、偏好数据或预先定义的教学标准,这些标准规定了什么是好的教学。然而,这些信号往往并不以个体学生的学习成果为依据,而有效的教学策略可能因学习者不同而存在显著差异。为了解决这个问题,我们提出了 Sherpa,一个多轮强化学习框架,该框架以不同的学习偏好为条件,用 LLMs 实例化多个学生原型,并训练一个教师模型,通过直接最大化他们的学习成果来调整其教学。使用 Sherpa 训练的教师 LLMs 使所指导学生在所有原型上的表现平均提高了 20.5 个百分点。在 MathTutorBench 的评估中,Sherpa 将整体教学法得分从 52.5% 提高到 79.2%,表明其教学回应更好。我们的人类研究表明,在 79.6% 的成对比较中,训练后的教师比基础模型更受偏好。总之,Sherpa 训练 LLM 教师适应多样化的模拟学生,并与人类教师更好地对齐,为 AI 导师教授真实学生铺平道路。
cs.AR / 34 / 2610.07443
A Shape-Adaptive Architecture with Disaggregated Quantization for Efficient LLM Serving
一种用于高效LLM服务的、带有解耦量化的形状自适应架构
large language model
大语言模型相关
Abstract
Large language models (LLMs) have become the backbone of modern AI applications, but pose significant challenges for efficient inference. Their autoregressive generation divides execution into two phases: prefill, dominated by large GEMMs, and decoding, dominated by small GEMVs. Modern serving systems further introduce complexity through continuous batching and prefill-decoding disaggregation, leading to dynamic workloads and phase separation. However, existing accelerators remain poorly aligned with these system-level behaviors, resulting in inefficiencies in LLM serving. In this work, we present DynaCore, a unified architecture for efficient LLM serving via system-architecture co-design. We observe that the compute tile a systolic array executes, its Minimum Efficient Unit (MEU), spans all three GEMM dimensions. DynaCore reshapes the MEU along all three: spatially it trades array width against height asymmetrically, raising weight delivery while leaving the input path untouched, and temporally Split-K maps the reduction onto the array, folding partial sums through the interconnect the array already has. To exploit phase separation, we further propose disaggregated quantization, applying dual-side quantization to prefill and weight-only quantization to decoding, with an inner-product mixed-precision datapath that keeps output width invariant to precision. A runtime scheduling framework then selects an MEU per batch. Evaluation with real-world serving traces shows that DynaCore substantially reduces service-level latency over quantization and reconfigurable accelerators, improving TTFT by 3.50x and 2.97x and TPOT by 36.55x and 8.02x, respectively.
Chinese Translation
大语言模型(LLM)已成为现代AI应用的支柱,但为高效推理带来了重大挑战。其自回归生成将执行划分为两个阶段:由大型GEMM主导的预填充(prefill),以及由小型GEMV主导的解码(decoding)。现代服务系统通过连续批处理(continuous batching)和预填充-解码解耦(prefill-decoding disaggregation)进一步引入复杂性,导致动态工作负载和阶段分离。然而,现有加速器仍与这些系统级行为契合不佳,导致LLM服务中的低效。在这项工作中,我们提出DynaCore,一种通过系统-架构协同设计实现高效LLM服务的统一架构。我们观察到,脉动阵列所执行的计算瓦片,即其最小高效单元(MEU),横跨GEMM的全部三个维度。DynaCore沿全部三个维度重塑MEU:在空间上,它以非对称方式用阵列宽度换取高度,在保持输入路径不变的同时提升权重传递;在时间上,Split-K将归约映射到阵列上,通过阵列已有的互连折叠部分和。为利用阶段分离,我们进一步提出解耦量化,对预填充应用双边量化,对解码应用仅权重量化,并采用内积混合精度数据通路,使输出宽度对精度保持不变。一个运行时调度框架随后为每个批次选择一个MEU。使用真实世界服务轨迹的评估表明,DynaCore相较于量化和可重构加速器大幅降低了服务级延迟,分别将TTFT提升3.50倍和2.97倍,并将TPOT提升36.55倍和8.02倍。
cs.AR / 35 / 2610.07668
CACHEFORGE: LLM-Guided End-to-End Generative Cache Replacement Policy for Performance and Hardware Efficiency
CACHEFORGE:LLM 引导的端到端生成式缓存替换策略,面向性能与硬件效率
large language model
大语言模型相关
Abstract
Modern cache replacement designs saturate because they operate within fixed representational structures, hand-crafted and heuristic based feature-engineered predictors, or offline imitation models that cannot generate new decision logic on their own. At the same time, replacement is shaped by the causal interaction of prefetching, thrashing, spatial locality, and access-type behavior, producing an enormous design space that is difficult to traverse manually. Prior approaches typically rely on heuristics, parameter tuning, or imitation of an offline optimal policy, capturing correlations rather than synthesizing new mechanisms. As a result, their performance gains often plateau and they overfit under dynamic workload conditions. CACHEFORGE is the first framework to evolve cache-replacement policies end-to-end by embedding a large language model inside a governed hardware-aware loop. In each iteration, the LLM proposes new C++ replacement logic, the policy is evaluated under a trace-based CRC-2 ChampSim simulator, and the framework enforces feasibility through reward shaping, structural checks, dynamic mutation, temperature scheduling, and cross-policy crossover. This closed-loop generation-evolution loop specifically designed for cache replacement policy enables the discovery of compact policies that satisfy hardware constraints while exploring algorithmic transformations beyond fixed predictor structures. Across SPEC CPU2006, CACHEFORGE outperforms all CRC-2 baselines. It improves the total hit rate by 27.36%, 19.69%, 13.72%, 13.15%, 11.83%, and 5.73% over MPPPB, ReD, Hawk-eye, SHiP++, LIME, and LRU, respectively. On memory-intensive workloads, it increases IPC by 10.15%, 7.89%, 6.34%, 3.64%, 3.12%, and 2.71% over LRU, MPPPB, LIME, ReD, SHiP++, and Hawkeye.
Chinese Translation
现代缓存替换设计趋于饱和,因为它们只能在固定的表示结构、手工设计且基于启发式的特征工程预测器,或无法自行生成新决策逻辑的离线模仿模型内运作。同时,替换还受到预取、抖动、空间局部性和访问类型行为的因果交互影响,这产生了一个难以人工遍历的巨大设计空间。先前方法通常依赖启发式、参数调优,或对离线最优策略的模仿,捕捉的是相关性而非综合出新的机制。因此,它们的性能增益往往趋于平缓,并在动态工作负载条件下过拟合。CACHEFORGE 是首个通过在受控的硬件感知闭环中嵌入大语言模型来端到端演化缓存替换策略的框架。在每次迭代中,LLM 提出新的 C++ 替换逻辑,该策略在基于 trace 的 CRC-2 ChampSim 模拟器下进行评估,框架通过奖励塑形、结构检查、动态变异、温度调度和跨策略交叉来强制满足可行性。这个专为缓存替换策略设计的闭环生成-演化循环,能够发现满足硬件约束的紧凑策略,同时探索超越固定预测器结构的算法变换。在 SPEC CPU2006 上,CACHEFORGE 优于所有 CRC-2 基线。它将总命中率分别比 MPPPB、ReD、Hawk-eye、SHiP++、LIME 和 LRU 提高了 27.36%、19.69%、13.72%、13.15%、11.83% 和 5.73%。在内存密集型工作负载上,它将 IPC 分别比 LRU、MPPPB、LIME、ReD、SHiP++ 和 Hawkeye 提高了 10.15%、7.89%、6.34%、3.64%、3.12% 和 2.71%。
cs.AR / 36 / 2610.08378
Lachesis: Lifetime-Aware KV Cache Placement for Agent Serving across HBM and High-Bandwidth Flash
Lachesis:面向跨 HBM 与高带宽闪存的智能体服务的生命周期感知 KV 缓存放置
large language model
大语言模型相关
Abstract
Large language model (LLM) serving is increasingly dominated by agentic workloads, in which agents and their sub-agents accumulate context as KV cache across many requests, consuming substantial memory. High-bandwidth flash (HBF) is a promising solution, providing an order of magnitude greater capacity at HBM-class read bandwidth, but its finite write endurance is the key limiting factor. Our key insight is that KV cache should be placed across HBM and HBF by its lifetime. Placing shorter-lived data in HBM lets HBM absorb more of an agent run's writes and sends less of them to HBF. As the lifetime of KV cache in agentic serving is dictated by the harness, the program that orchestrates the agents, we analyze its behavior and identify three axes along which lifetime diverges, temporal, structural, and inter-worker. Guided by these observations, we present Lachesis, a lifetime-aware KV cache placement layer between the agent harness and the serving engine. At write time, it places each segment in HBM or HBF according to its lifetime, and frees its blocks once the segment is no longer read. In trace-driven simulation, Lachesis extends HBF lifetime by 1.19-3.13x over HBM-first placement, reaching 3.3-12.2 device-years. Even under continuous 24x7 operation at the full load a tight SLO admits, HBF outlasts its five-year warranty on the multi-agent trace.
Chinese Translation
大语言模型(LLM)服务正日益被智能体工作负载所主导;在这些工作负载中,智能体及其子智能体在众多请求之间以 KV 缓存的形式累积上下文,消耗大量内存。高带宽闪存(HBF)是一种有前景的解决方案,它以 HBM 级别的读带宽提供了高一个数量级的容量,但其有限的写入耐久性是关键的制约因素。我们的核心洞见是:KV 缓存应当依据其生命周期在 HBM 与 HBF 之间进行放置。将生命周期较短的数据放在 HBM 中,可使 HBM 吸收一次智能体运行中更多的写入,并将更少的写入发送到 HBF。由于智能体服务中 KV 缓存的生命周期由 harness(即编排智能体的程序)所决定,我们分析了其行为,并识别出生命周期出现分化的三个维度:时间维度、结构维度和跨工作进程维度。在这些观察的指导下,我们提出 Lachesis,一个位于智能体 harness 与服务引擎之间的生命周期感知 KV 缓存放置层。在写入时,它根据每个段的生命周期将其放置到 HBM 或 HBF 中,并在该段不再被读取后释放其块。在轨迹驱动的模拟中,与 HBM 优先放置相比,Lachesis 将 HBF 寿命延长了 1.19-3.13 倍,达到 3.3-12.2 设备年。即使在严格 SLO 所允许的满负载下 24x7 连续运行,在多智能体轨迹上,HBF 的寿命仍超过其五年保修期。
cs.CL / 37 / 2610.07186
Identifying Introspection From the Inside
从内部识别内省
large language model
大语言模型相关
Abstract
Large language models make claims about themselves that are both consequential and increasingly difficult to verify from behavior alone. How can we distinguish plausible confabulations from genuine introspection? In this paper, we identify mechanistic signatures of faithful self-report in a controlled setting. Using low-rank adapters, we train models to make decisions on behalf of fictitious characters, according to latent linear preference functions. We find sustained fine-tuning on an implicit decision task can lead to the emergence of accurate self-reporting of models' learned preferences, even without explicit self-report supervision. We ask two research questions about this emergent phenomenon. First: is the emergence of accurate self-reporting accompanied by a measurable structural change in the model? Weight ablations and frozen-layer experiments together indicate that preference representations shift to earlier layers over training, consistent with the hypothesis that faithful self-report requires preferences to be located where pre-existing verbalization mechanisms can access them. Second: can these structural differences distinguish faithful models from unfaithful ones? Using attribution patching, we find that faithful models exhibit significantly higher attribution similarity between the decision-making and self-report tasks -- a mechanistic signature of faithful self-report that does not require us to understand the content of the report itself. Previous work on self-report has observed behaviorally that models can be faithful or unfaithful; our work proposes that, at least in our restricted setting, it is possible to distinguish between the two patterns of computation by examining the structure of the networks themselves.
Chinese Translation
大型语言模型对自身做出的陈述既具有重要影响,又越来越难以仅凭行为加以验证。我们如何区分看似合理的虚构与真正的内省?在本文中,我们在一个受控环境中识别出忠实自我报告的机制性特征。使用低秩适配器,我们训练模型依据潜在的线性偏好函数,代表虚构角色做出决策。我们发现,在一个隐式决策任务上进行持续微调,能够导致模型对其所学偏好的准确自我报告涌现出来,即使没有显式的自我报告监督。围绕这一涌现现象,我们提出两个研究问题。第一:准确自我报告的涌现是否伴随着模型中可测量的结构变化?权重消融与冻结层实验共同表明,随着训练进行,偏好表征会迁移到更早的层,这与如下假设一致:忠实的自我报告要求偏好位于既有的言语化机制能够访问的位置。第二:这些结构差异能否将忠实模型与不忠实模型区分开来?使用归因修补,我们发现忠实模型在决策任务与自我报告任务之间表现出显著更高的归因相似度——这是一种忠实自我报告的机制性特征,且不需要我们理解报告本身的内容。先前关于自我报告的研究从行为层面观察到模型可能忠实也可能不忠实;我们的工作提出,至少在我们的受限设定中,可以通过考察网络本身的结构来区分这两种计算模式。
cs.CL / 38 / 2610.07224
TIDE 2.0: an open, model-agnostic engine for keyed de-identification of clinical notes
TIDE 2.0:一个用于临床病历基于密钥去标识化的开放、模型无关引擎
large language model
大语言模型相关
Abstract
Clinical notes capture most of what is documented about a patient's care, but they cannot be used for research until protected health information (PHI) is removed. De-identification is often treated as a detection problem. Detection alone is not sufficient: redaction strips clinical content along with identifiers, date blanking destroys the temporal intervals needed for longitudinal analysis, and assigning a fresh random surrogate at each occurrence breaks links between a patient's notes. We present TIDE 2.0, an MIT-licensed engine with two separable stages: an interchangeable recognizer and a keyed anonymizer. Both run on hardware the institution owns. Surrogates are generated cryptographically with no stored linkage table. Dates shift by a per-patient, interval-preserving offset; each value receives the same surrogate across all occurrences under a given key; and a release produced under a new key cannot be linked to earlier releases. We also release TIDE2-Sentry, a recognizer distilled from a large language model. On two gold-annotated corpora from two institutions, the default configuration reached span-level recall of 0.88 in-domain and 0.77 on the second institution's corpus, at precision 0.88 and 0.87. We report recall and precision per category alongside these aggregates. The engine is open source, and the recognizer is available under a gated research-use agreement, so institutions can run, inspect and extend both within their own environments.
Chinese Translation
临床病历记录了患者诊疗过程中被记载的大部分内容,但在受保护健康信息(PHI)被移除之前,它们无法用于研究。去标识化通常被视为一个检测问题。仅靠检测并不足够:删减会连同标识符一起剥离临床内容,日期置空会破坏纵向分析所需的时间区间,而为每次出现分配一个新的随机替代值则会破坏患者各份病历之间的关联。我们提出 TIDE 2.0,一个采用 MIT 许可证的引擎,具有两个可分离的阶段:可互换的识别器与基于密钥的匿名化器。两者均运行在机构自有的硬件上。替代值以密码学方式生成,不存储任何关联表。日期按每位患者进行保持时间区间的偏移;在给定密钥下,每个取值在所有出现位置都获得相同的替代值;并且在新的密钥下生成的发布版本无法与更早的发布版本相关联。我们还发布了 TIDE2-Sentry,一个从大语言模型蒸馏得到的识别器。在两个来自两家机构的人工标注金标准语料上,默认配置在域内达到 0.88 的跨度级召回率,在第二家机构的语料上达到 0.77,精确率分别为 0.88 和 0.87。我们在报告这些总体指标的同时,也报告了各类别的召回率和精确率。该引擎是开源的,识别器则在受控的研究使用协议下提供,因此机构可以在自己的环境中运行、检查和扩展两者。
cs.CL / 39 / 2610.07306
Kurate: Scalable Scientific Quality Analysis
Kurate:可扩展的科学质量分析
large language model
大语言模型相关
Abstract
Scientific search systems can find papers that are relevant to a question, but they generally do not assess the quality of the evidence that those papers provide. We present Kurate, a system that uses large language models (LLMs) to assess the quality of published studies. Kurate uses both the paper and its related documents (e.g., the study's trial registration and protocol), and links each of its judgments to the passage of text on which that judgment is based. We applied Kurate to a corpus of 4,347 papers (3,913 of which report randomized trials) and scored each paper on 8 dimensions of study design and reporting: specifically, statistical power, causal identification, preregistration, selective reporting, measurement validity, analysis prespecification, reporting transparency, and conflict of interest and funding. Across the corpus, we found that papers most often exhibited issues with statistical power, selective reporting, and analysis prespecification, although average quality differed between clinical areas. When compared against expert annotations of 60 held-out clinical-trial documents, the information Kurate extracted matched the expert label in 221/242 protocol scorepoints and 294/370 results-publication scorepoints, with AC1 0.94 and 0.81, respectively. Using a well-reputed, high quality clinical trial as a worked example, we show how a single paper's overall grade breaks down into separate judgments, with each linked to specific evidence from the trial's registration, protocol, and published report. Together, these results show that large-scale quality assessment of this kind is feasible, and that it can be used to address meta-scientific research questions.
Chinese Translation
科学搜索系统能够找到与某一问题相关的论文,但它们通常并不评估这些论文所提供的证据的质量。我们提出了 Kurate,一个使用大语言模型(LLMs)来评估已发表研究质量的系统。Kurate 同时使用论文本身及其相关文档(例如,该研究的试验注册信息和研究方案),并将其每一项判断与其所依据的文本段落相链接。我们将 Kurate 应用于一个包含 4,347 篇论文的语料库(其中 3,913 篇报告了随机试验),并就研究设计与报告的 8 个维度对每篇论文进行评分:具体而言,包括统计功效、因果识别、预注册、选择性报告、测量效度、分析预设定、报告透明度,以及利益冲突与资金。在整个语料库中,我们发现论文最常出现的问题集中在统计功效、选择性报告和分析预设定方面,尽管各临床领域之间的平均质量存在差异。在与 60 份留出临床试验文档的专家标注进行比较时,Kurate 所提取的信息在 242 个方案评分点中有 221 个、在 370 个结果发表评分点中有 294 个与专家标签相匹配,AC1 分别为 0.94 和 0.81。我们以一项声誉良好、高质量的临床试验作为实例,展示了单篇论文的总体评级如何分解为各自独立的判断,且每一项判断都链接到来自该试验注册信息、研究方案和已发表报告的具体证据。总体而言,这些结果表明,此类大规模质量评估是可行的,并且可用于解决元科学研究问题。
cs.CL / 40 / 2610.07563
HouseholdBench: Evaluating Large Language Models as Predictors of Household Economic Behavior
HouseholdBench:评估大型语言模型作为家庭经济行为预测器
large language model
大语言模型相关
Abstract
Large language models (LLMs) have the potential to meet a key goal in economics: a quantitative model of household decision making, across a variety of settings. Yet existing evaluations cover few surveys and outcomes, and do not study how households adjust to changing economic conditions. We introduce a new evaluation, HouseholdBench, which unites 6 U.S. household surveys and 32 prediction tasks spanning numeric, categorical and probabilistic outcomes, related to consumption, income, labor, expectations, and housing. Using past behavior, demographics and macroeconomic conditions, the tasks test whether LLMs predict behavior, including how households adjust to changes in various policies. We evaluate 13 proprietary and open-weight LLMs against a no-change baseline and a gradient-boosted tree model. Most LLMs outperform the no-change baseline, including for policy response tasks -- with the best model lowering error for numeric outcomes by 12.2%. Across most tasks, gradient-boosted trees rank first; leading proprietary LLMs approach their performance, but open-weight models lag. LLMs exhibit systematic over- and underprediction across different tasks. We identify methods that enable a 4 billion parameter open-weight model to match proprietary models' performance: fine-tuning and aggregating 16 predictions per observation. Improvements generalize to policy-response tasks, which are excluded from fine-tuning. We release our datasets, code, and leaderboard on our website: https://jn-huang.github.io/householdbench
Chinese Translation
大型语言模型(LLMs)有潜力实现经济学中的一个关键目标:一个跨多种情境的家庭决策的定量模型。然而,现有的评估只涵盖少数调查和结果,并且没有研究家庭如何适应不断变化的经济状况。我们引入一个新的评估,HouseholdBench,它整合了 6 项美国家庭调查和 32 个预测任务,涵盖数值型、类别型和概率型结果,涉及消费、收入、劳动、预期和住房。这些任务使用过去行为、人口统计特征和宏观经济状况,检验 LLMs 是否能预测行为,包括家庭如何调整以适应各种政策的变化。我们评估了 13 个专有和开放权重 LLMs,与一个不变基线和一个梯度提升树模型进行比较。大多数 LLMs 优于不变基线,包括在政策响应任务上——最佳模型将数值型结果的误差降低了 12.2%。在大多数任务中,梯度提升树排名第一;领先的专有 LLMs 接近其性能,但开放权重模型落后。LLMs 在不同任务中表现出系统性的高估和低估预测。我们识别出使一个 40 亿参数开放权重模型达到专有模型性能的方法:微调以及聚合每个观测的 16 个预测。这些改进可泛化到政策响应任务,而这些任务被排除在微调之外。我们在我们的网站发布我们的数据集、代码和排行榜:https://jn-huang.github.io/householdbench
cs.CL / 41 / 2610.07572
Two Vectors Replace In-Context Demos: Structured Task Adaptation via Embeddings
两个向量取代上下文示例:通过嵌入实现结构化任务适应
large language model
大语言模型相关
Abstract
In-context learning (ICL) adapts frozen large multimodal models (LMMs) to new tasks from a few demonstrations (demos), but re-encodes them at every query, where each demo image adds up to hundreds of visual tokens. Demo-free methods remove this cost with a compact task state. However, they add it at locations searched per task or at every decoder layer, where task parameters grow with depth. Moreover, inserted tokens or keys cannot change how the original prompt divides its attention within a layer. To address these issues, we propose Structured Task Adaptation via Embeddings (STAVE), which replaces demos with two task-specific vectors added to existing input embeddings. Specifically, a readout vector updates the answer-producing tokens and a context vector updates the other structural token groups. Both are trained with answer labels on prompts with and without demos. We justify these design choices theoretically using a first-order analysis of the loss and a margin bound. Extensive experiments on six LMMs and five large language models show that STAVE matches or outperforms state-of-the-art methods on multimodal tasks with far fewer task parameters and surpasses 15-shot ICL and prior task vectors on 18 text tasks, all at zero-shot inference cost.
Chinese Translation
上下文学习(ICL)能够借助少量示例(demos)使冻结的大型多模态模型(LMMs)适应新任务,但它会在每次查询时重新编码这些示例,而每张示例图像最多会带来数百个视觉 token。无示例方法通过一个紧凑的任务状态消除了这一开销。然而,它们又在按任务搜索得到的位置处、或在每一个解码器层处重新引入了这一开销,而在这些地方,任务参数量会随深度增长。此外,插入的 token 或键无法改变原始提示在某一层内部分配其注意力的方式。为解决这些问题,我们提出了通过嵌入实现的结构化任务适应(STAVE),它用两个添加到现有输入嵌入上的任务特定向量来取代示例。具体而言,一个读出向量更新产生答案的 token,而一个上下文向量更新其他结构化 token 组。二者都在带示例和不带示例的提示上使用答案标签进行训练。我们通过损失的一阶分析和间隔界(margin bound)从理论上为这些设计选择提供了依据。在六个 LMM 和五个大型语言模型上的大量实验表明,STAVE 在多模态任务上以远少得多的任务参数量达到或超越了最先进的方法,并在 18 项文本任务上超越了 15 样本 ICL 和先前的任务向量,且推理成本均为零样本水平。
cs.CL / 42 / 2610.07587
Large Language Model Orchestration under Heterogeneous Preferences via Explicit Persona Inference
异质偏好下通过显式角色推断的大语言模型编排
large language model
大语言模型相关
Abstract
LLM orchestration investigates how an orchestrator coordinates a group of autonomous agents to achieve common goals or maximize collective welfare. The agents are typically heterogeneous, each holding a private preference that it pursues but does not reveal. Inferring such hidden preferences from behavior has been a subject of long-standing research in game theory and multi-agent systems. The core challenge lies in maintaining a belief over every agent's preference and updating it from the agents' observed actions. Existing LLM orchestrators carry that belief as prompt text with no explicit update rule. This lets early errors persist and propagate rather than be corrected. We therefore propose \textbf{HARP} (Heterogeneous-preference Agent oRchestration via Preference inference), a novel framework that moves the belief out of the prompt. Specifically, HARP maintains one numeric posterior per agent over a finite set of candidate preferences and updates it in closed form by Bayes' rule. The language model supplies only actions and per-candidate likelihoods, so estimation is decoupled from its reasoning. We prove that HARP attains the same $\tilde O(\sqrt K)$ Bayesian regret as explicit joint inference when the factorization is exact. Furthermore, HARP\textsuperscript{+} augments planning with a bonus for actions that distinguish the candidates, so inference continues even when the optimal action is uninformative. Empirical results on three substrates, ranging from payoffs the preferences fully determine, through payoffs that depend on more than them, to scales where explicit joint inference is infeasible, demonstrate that HARP\textsuperscript{+} is the strongest non-oracle method across the class our theory identifies.
Chinese Translation
LLM 编排研究编排器如何协调一组自主智能体以实现共同目标或最大化集体福利。这些智能体通常是异质的,每个智能体持有一种它追求但不揭示的私有偏好。从行为中推断此类隐藏偏好一直是博弈论和多智能体系统中长期研究的主题。核心挑战在于维护对每个智能体偏好的信念,并根据观察到的智能体动作更新该信念。现有的 LLM 编排器将该信念作为提示文本携带,没有显式的更新规则。这会让早期错误持续并传播,而不是被纠正。因此,我们提出 \textbf{HARP}(通过偏好推断的异质偏好智能体编排),一个将信念移出提示的新颖框架。具体而言,HARP 在有限候选偏好集合上为每个智能体维护一个数值后验,并通过贝叶斯法则以闭式形式更新它。语言模型仅提供动作和每个候选的似然,因此估计与其推理解耦。我们证明,当分解精确时,HARP 达到与显式联合推断相同的 $\tilde O(\sqrt K)$ 贝叶斯遗憾。此外,HARP\textsuperscript{+} 通过为能区分候选的动作添加额外奖励来增强规划,因此即使最优动作不提供信息,推断也会继续。在三个基底上的实证结果,范围从偏好完全决定的收益,经由依赖于偏好之外的更多因素的收益,到显式联合推断不可行的规模,表明 HARP\textsuperscript{+} 是我们理论所识别的那一类中最强的非预言机方法。
cs.CL / 43 / 2610.07659
DLoop: Looped Speculative Decoding
DLoop:循环式推测解码
large language model
大语言模型相关
Abstract
Speculative decoding accelerates autoregressive generation in large language models. In each drafting stage, a lightweight draft model proposes tokens that the target model subsequently verifies. With increasingly capable draft models, we find that the target model frequently accepts all tokens produced in a drafting stage. A verification nevertheless follows each drafting stage, resulting in unnecessary target-model forward passes even when drafting could have continued. Adaptive draft length methods decide during decoding how many draft tokens precede a verification, but they raise the speedup only for autoregressive draft models. For a parallel draft model, drafting further requires target-model hidden states for draft tokens that have not been verified. We propose DLoop, a looped form of speculative decoding that adaptively performs multiple drafting stages before verification. DLoop continues drafting while the draft model remains confident and verifies all accumulated draft tokens together. Loop-aware training keeps the draft model reliable in the additional drafting stages by exposing it to its own hidden states for unverified draft tokens. By spending additional draft-model forward passes, DLoop reduces the number of target-model forward passes required for verification. Across diverse speculative decoding methods including EAGLE-3, DFlash, Domino, DSpark, and multi-token prediction modules, DLoop improves the wall-clock speedup by 5 to 41 percent while preserving lossless decoding. Code will be available at https://github.com/naver-ai/DLoop.
Chinese Translation
推测解码加速了大语言模型中的自回归生成。在每个起草阶段,一个轻量级的草稿模型提出若干 token,随后由目标模型进行验证。随着草稿模型能力不断增强,我们发现目标模型经常接受某个起草阶段中产生的所有 token。尽管如此,每个起草阶段之后仍会进行一次验证,这导致即使起草本可以继续,也会产生不必要的目标模型前向传播。自适应草稿长度方法在解码过程中决定在一次验证之前有多少个草稿 token,但它们只能为自回归草稿模型带来加速提升。对于并行草稿模型,要进一步起草就需要尚未验证的草稿 token 所对应的目标模型隐藏状态。我们提出 DLoop,一种循环形式的推测解码,它在验证之前自适应地执行多个起草阶段。当草稿模型保持置信时,DLoop 会继续起草,并一起验证所有累积的草稿 token。循环感知训练通过让草稿模型接触其自身针对未验证草稿 token 的隐藏状态,使草稿模型在额外的起草阶段中保持可靠。通过花费额外的草稿模型前向传播,DLoop 减少了验证所需的目标模型前向传播次数。在包括 EAGLE-3、DFlash、Domino、DSpark 以及多 token 预测模块在内的多种推测解码方法上,DLoop 在保持无损解码的同时,将实际运行时间加速比提高了 5% 到 41%。代码将在 https://github.com/naver-ai/DLoop 提供。
cs.CL / 44 / 2610.07700
Detecting LLM-Assisted Vietnamese Writing via Keystrokes under Behavioral Manipulation
在行为操纵下通过击键检测大语言模型辅助的越南语写作
large language model
大语言模型相关
Abstract
We study the robustness of keystroke dynamics for detecting large language model (LLM)-assisted writing. We introduce a Vietnamese keystroke dataset capturing realistic writing modes, including bona fide composition, transcription, and paraphrasing. We also define a behaviorally grounded threat model in which users deliberately alter typing patterns. To implement the threat model, we create behaviorally manipulated variants of the data designed to evade keystroke-based detection. We evaluate four keystroke modeling approaches: temporal and rhythmic representations, and sequential representations modeled with a one-dimensional convolutional neural network (1D-CNN) and TypeNet, under user-independent and context-independent settings. The results show that sequential models outperform feature-based approaches in most cases and that keystroke signals encode discriminative information about the writing process. However, detection is not uniformly robust: transcription is reliably identified, while paraphrasing and adversarially manipulated samples are frequently misclassified as bona fide when not explicitly modeled. To address this, we incorporate adversarial training using behaviorally manipulated data, which substantially improves separability and robustness. These results suggest that keystroke-based detection depends critically on exposure to diverse writing behaviors, and that strong performance under limited conditions does not generalize to realistic or adversarial settings without targeted modeling.
Chinese Translation
我们研究击键动力学在检测大语言模型(LLM)辅助写作方面的鲁棒性。我们引入了一个越南语击键数据集,该数据集捕捉了真实的写作模式,包括真实写作、转写和改写。我们还定义了一个基于行为的威胁模型,其中用户故意改变其打字模式。为实现该威胁模型,我们创建了旨在规避基于击键的检测的数据的行为操纵变体。我们在用户无关和上下文无关的设置下评估了四种击键建模方法:时间与节奏表示,以及使用一维卷积神经网络(1D-CNN)和 TypeNet 建模的序列表示。结果表明,在大多数情况下,序列模型优于基于特征的方法,并且击键信号编码了关于写作过程的判别性信息。然而,检测并非一致地鲁棒:转写能够被可靠地识别,而在未显式建模时,改写和对抗性操纵的样本经常被误分类为真实写作。为解决这一问题,我们引入使用行为操纵数据的对抗训练,这显著提升了可分性和鲁棒性。这些结果表明,基于击键的检测关键取决于对多样化写作行为的接触,并且在有限条件下取得的强劲性能若没有针对性的建模,并不能泛化到现实或对抗性场景中。
cs.CL / 45 / 2610.07722
Does Steering Break Your Model? A Multi-Dimensional Evaluation Suite for LLM Steering Methods
引导会破坏你的模型吗?面向LLM引导方法的多维评估套件
large language model
大语言模型相关
Abstract
Activation steering provides a lightweight and flexible way to control large language model (LLM) behavior. However, effective steering requires more than inducing the intended behavior: it should also limit unintended changes and remain robust across inputs and training data. Existing evaluations cover these dimensions only in fragments. As a result, the trade-offs between efficacy and side effects have not been systematically characterized. We introduce SteerScope, a two-axis, multi-dimensional evaluation suite that jointly characterizes steering outcomes and method properties through 15 metrics. We score target efficacy and side effects on language quality, task capabilities, and safety and reliability, and further assess generalization and data dependence through steering-specific metrics for sample efficiency and sample sensitivity. Rather than comparing methods at a single operating point, we characterize the trade-offs between efficacy and side effects. Under matched models, tasks, and evaluation protocols, we benchmark 23 methods spanning 4 families, including prompting, LoRA, and SFT as baseline methods, and release the suite as an extensible codebase. We find that current activation steering methods do not yet surpass the Prompt Steering baseline in their overall balance between steering efficacy and side effects: across both model scales, no evaluated activation steering method achieves higher efficacy without incurring greater composite side effects. We further uncover a consistent coupling between steering efficacy and side effects. Under OOD prompts, target efficacy is often preserved, whereas side effects tend to become more pronounced, particularly through declines in instruction relevance and fluency. Methods also exhibit sharply different sample-efficiency profiles.
Chinese Translation
激活引导提供了一种轻量且灵活的方式来控制大语言模型(LLM)的行为。然而,有效的引导不仅仅需要诱导出预期行为:它还应当限制非预期变化,并在不同输入和训练数据上保持稳健。现有评估仅零散地覆盖了这些维度。因此,有效性(efficacy)与副作用之间的权衡尚未得到系统刻画。我们提出SteerScope,一个双轴、多维评估套件,通过15个指标共同刻画引导结果和方法属性。我们在语言质量、任务能力以及安全性和可靠性上对目标有效性和副作用进行评分,并进一步通过针对样本效率和样本敏感性的引导专用指标来评估泛化性和数据依赖性。我们不是在单一操作点上比较方法,而是刻画有效性与副作用之间的权衡。在匹配的模型、任务和评估协议下,我们对涵盖4个方法族的23种方法进行了基准测试,其中包括提示、LoRA和SFT作为基线方法,并将该套件发布为可扩展的代码库。我们发现,当前激活引导方法在引导有效性与副作用之间的总体平衡上尚未超越提示引导(Prompt Steering)基线:在两个模型规模上,没有任何被评估的激活引导方法能在不带来更大复合副作用的情况下实现更高有效性。我们进一步揭示出引导有效性与副作用之间存在一致的耦合关系。在OOD提示下,目标有效性通常得以保持,而副作用往往变得更加明显,尤其表现为指令相关性和流畅性的下降。不同方法还表现出截然不同的样本效率特征。
cs.CL / 46 / 2610.07764
No Transformer Beats Six Covariates: Long-Horizon Prediction of Depressive Symptoms from Childhood Essays
没有 Transformer 能击败六个协变量:从童年作文长时程预测抑郁症状
large language model
大语言模型相关
Abstract
Natural language processing (NLP) models can detect depression-related language in text written near the time symptoms are measured, but whether pretrained transformers can predict depressive symptoms from text written twelve years earlier is largely untested. In the National Child Development Study, a British birth cohort, we predict probable depressive symptoms at age 23 from essays the same people wrote at age 11. Our baseline, a logistic regression on six childhood covariates, outperforms every text model that sees only the essay: seven fine-tuned transformers, a bag-of-words model, frozen embeddings and four zero-shot large language models. Its area under the receiver operating characteristic curve (AUC-ROC) is 0.737 against 0.670 for the best transformer on the primary seed, and no added text score detectably raises the baseline's AUC-ROC. None of the five domain-pretrained transformers detectably beats its general-domain control after Bonferroni correction. For long-horizon prediction, the baseline remains the model to beat.
Chinese Translation
自然语言处理(NLP)模型能够检测在症状测量时间附近所写文本中的抑郁相关语言,但预训练的 Transformer 是否能从十二年前所写的文本中预测抑郁症状,这在很大程度上尚未得到检验。在全国儿童发展研究(National Child Development Study)这一英国出生队列中,我们根据同一批人在 11 岁时所写的作文,预测其 23 岁时可能的抑郁症状。我们的基线模型——一个基于六个童年协变量的逻辑回归——优于每一个仅看到作文的文本模型:七个微调的 Transformer、一个词袋模型、冻结嵌入以及四个零样本大语言模型。在主要随机种子上,其受试者工作特征曲线下面积(AUC-ROC)为 0.737,而最佳 Transformer 为 0.670;并且没有任何加入的文本得分能够可检测地提高基线的 AUC-ROC。在 Bonferroni 校正后,五个领域预训练的 Transformer 中没有任何一个能够可检测地胜过其通用领域对照。对于长时程预测而言,基线模型仍然是那个有待被超越的模型。
cs.CL / 47 / 2610.07780
APEX: Speculate smarter, not deeper
APEX:更智能地推测,而非更深地推测
large language model
大语言模型相关
Abstract
Speculative decoding reduces large language model inference latency by drafting multiple tokens before target-model verification, but its effectiveness depends on both the proposal mechanism and draft depth. Fixed configurations cannot respond to changes in predictability, repetition, and acceptance during generation, so deeper drafting can increase wasted computation without proportional speedup. We introduce APEX, a learned controller that balances decoding speed and draft-token waste through request-level expert selection and block-level depth adaptation. APEX-Router selects among EAGLE-3, n-gram, and draft-model speculation for each request, while APEX-Depth adjusts draft length at each verification block using causal decoding signals and recent verifier feedback. APEX models accepted draft length as censored survival feedback, learning position-wise rejection hazards, block execution costs, and an action utility that balances throughput, accepted progress, and wasted tokens. This allows the controller to adapt speculation while retaining the target model's verification procedure. We integrate APEX into vLLM and evaluate it with Qwen3-8B across six workloads, achieving up to 5.24X speedup over autoregressive decoding. Across the aggregate evaluation, APEX-S achieves 4.27X speedup, while APEX-B achieves 3.27X speedup with a 41.0% relative reduction in wasted-token percentage compared with fixed n-gram speculation at k=16, providing distinct operating points for balancing acceleration and draft-token utilization.
Chinese Translation
推测解码通过在目标模型验证之前起草多个 token 来降低大语言模型推理延迟,但其有效性取决于提议机制和起草深度两者。固定配置无法响应生成过程中可预测性、重复性和接受情况的变化,因此更深的起草会增加浪费的计算量,却没有成比例的加速。我们提出 APEX,一个学习型控制器,通过请求级专家选择和块级深度自适应来平衡解码速度与起草 token 浪费。APEX-Router 为每个请求在 EAGLE-3、n-gram 和草稿模型推测之间进行选择,而 APEX-Depth 则利用因果解码信号和近期验证器反馈,在每个验证块处调整起草长度。APEX 将被接受的起草长度建模为删失生存反馈,学习逐位置的拒绝风险、块执行成本,以及一种平衡吞吐量、被接受的进展和浪费 token 的动作效用。这使控制器能够自适应推测,同时保留目标模型的验证流程。我们将 APEX 集成到 vLLM 中,并在六个工作负载上使用 Qwen3-8B 对其进行评估,相较于自回归解码实现了最高 5.24X 的加速。在总体评估中,APEX-S 实现了 4.27X 的加速,而 APEX-B 实现了 3.27X 的加速,并且与 k=16 的固定 n-gram 推测相比,浪费 token 百分比相对降低了 41.0%,为平衡加速与起草 token 利用率提供了不同的操作点。
cs.CL / 48 / 2610.07848
Dynamic Positional Attention Modulation for Parameter-Efficient Fine-Tuning of Large Language Models
用于大语言模型参数高效微调的动态位置注意力调制
large language model
大语言模型相关
Abstract
Parameter-efficient fine-tuning (PEFT) has become a standard approach for adapting large language models to downstream tasks. However, most existing PEFT methods rely on uniform and static adaptations, without accounting for the structured heterogeneity of attention across dimensions, heads, layers, and input tokens. In practice, attention representations exhibit non-uniform behavior, and positional encoding mechanisms such as rotary positional embeddings (RoPE) induce dimension-dependent positional structure, making uniform adaptation suboptimal. In this work, we propose DyPAM (Dynamic Positional Attention Modulation), a PEFT method that adapts how positional information contributes to attention by operating directly on the query and key representations. DyPAM combines input-conditioned, dimension-wise modulation with head-wise and layer-wise structural modulation, performing fine-grained adaptation of positional attention aligned with the RoPE-induced structure without modifying the pretrained backbone. Extensive experiments on mathematical and commonsense reasoning benchmarks across multiple backbone models demonstrate that DyPAM consistently outperforms existing strong PEFT baselines.
Chinese Translation
参数高效微调(PEFT)已成为使大语言模型适配下游任务的标准方法。然而,大多数现有 PEFT 方法依赖于统一且静态的适配,未考虑注意力在维度、头、层和输入标记之间的结构化异质性。在实践中,注意力表示展现出非均匀行为,而诸如旋转位置嵌入(RoPE)之类的位置编码机制会诱导依赖于维度的位置结构,使得统一适配次优。在这项工作中,我们提出 DyPAM(动态位置注意力调制),一种 PEFT 方法,它通过直接作用于查询和键表示来适配位置信息对注意力的贡献方式。DyPAM 将以输入为条件的、逐维度的调制与逐头及逐层的结构调制相结合,在与 RoPE 诱导结构对齐的情况下对位置注意力进行细粒度适配,且不修改预训练主干。在多个主干模型上,针对数学和常识推理基准的大量实验表明,DyPAM 持续优于现有强 PEFT 基线。
cs.CL / 49 / 2610.07894
Rethinking Faithfulness in LLMs: A Pairwise Context-Sensitive Perspective
重新思考大语言模型中的忠实性:一种成对上下文敏感的视角
large language model
大语言模型相关
Abstract
Large language models (LLMs) are expected to answer questions faithfully based on the provided context, abstaining when the context information is insufficient to answer the questions. Existing faithfulness evaluations typically assess each question-context instance in isolation; however, such instance-level evaluation fails to capture a fundamental requirement of faithful behavior: the ability to adapt model responses to changes in available contexts. In particular, a model should provide correct answers when sufficient evidence is present and abstain when it is not. In this work, we propose a Pairwise Faithfulness Benchmark (PFaithBench) that evaluates whether a model can switch between answering and abstaining for the same question under supporting versus non-supporting contexts. Our evaluations across thirty-nine models with seven model families demonstrate that faithfulness fundamentally involves a trade-off between answering and abstaining, and that most current models exhibit a strong bias toward answering, with most faithfulness errors arising from over-answering, i.e., models tend to fabricate a response even when the provided context is insufficient. We further conduct a series of studies on faithfulness training under different data constructions. Our results show that training outcomes are highly sensitive to the specific composition of answering and abstaining data. Constructing answering and abstaining data from mismatched sources can cause models to rely on dataset-specific shortcuts rather than actual context sufficiency. Moreover, increasing answer-supervised data improves answering performance but exacerbates over-answering, while increasing abstaining data reduces hallucination but leads to over-abstention. The code and data are released at https://github.com/tmlr-group/PFaithBench.
Chinese Translation
大语言模型(LLMs)被期望基于所提供的上下文忠实地回答问题,并在上下文信息不足以回答问题时弃答。现有的忠实性评估通常孤立地评估每个问题-上下文实例;然而,这种实例级评估未能捕捉忠实行为的一项基本要求:使模型响应适应可用上下文变化的能力。具体而言,当存在充分证据时,模型应提供正确答案,而当不存在充分证据时应弃答。在这项工作中,我们提出了一个成对忠实性基准(Pairwise Faithfulness Benchmark,PFaithBench),用于评估模型是否能够在支持性上下文与非支持性上下文下,针对同一问题在回答与弃答之间切换。我们对来自七个模型家族的三十九个模型进行的评估表明,忠实性从根本上涉及回答与弃答之间的权衡,并且大多数当前模型表现出强烈的回答偏向,大多数忠实性错误源于过度回答,即即使所提供的上下文不充分,模型也倾向于编造回答。我们进一步开展了一系列关于不同数据构造下忠实性训练的研究。我们的结果表明,训练结果对回答数据与弃答数据的具体组成高度敏感。从失配来源构造回答数据和弃答数据,可能导致模型依赖数据集特定的捷径,而不是实际的上下文充分性。此外,增加答案监督数据会提升回答性能,但会加剧过度回答;而增加弃答数据会减少幻觉,但会导致过度弃答。代码和数据已在 https://github.com/tmlr-group/PFaithBench 发布。
cs.CL / 50 / 2610.07936
Pseudowords as probes: Large Language Models show little of the sublexical sensitivity that governs human pseudoword processing
伪词作为探针:大型语言模型几乎没有表现出支配人类伪词加工的亚词汇敏感性
large language model
大语言模型相关
Abstract
Systematicity, the probabilistic mapping of form to meaning, permeates language at all levels, and sublexical cues have been shown to govern human pseudoword processing. Yet whether LLMs exhibit comparable sensitivity to these cues remains unclear. We tested five LLMs on two Italian two-alternative forced-choice pseudoword experiments and compared their responses with a human behavioural baseline. LLMs aligned more reliably with humans when real-word options provided a lexical familiarity cue than in the pseudoword-only condition, where they fell substantially below fastText, a character-n-gram model. In addition, the sublexical cosine-similarity cue that reliably drove human--fastText agreement did not consistently transfer to human--LLM alignment, and reasoning-token expenditure bore no consistent relation to human processing difficulty. These findings suggest that LLMs do not necessarily share the sublexical cues that govern human pseudoword processing; we discuss tokenization and training-data coverage as candidate explanations.
Chinese Translation
系统性,即形式到意义的概率映射,渗透在语言的各个层面,并且亚词汇线索已被证明支配着人类的伪词加工。然而,大型语言模型是否对这些线索表现出可比的敏感性仍不清楚。我们在两个意大利语二选一强制选择伪词实验中测试了五个大型语言模型,并将它们的反应与人类行为基线进行了比较。当真实词选项提供了词汇熟悉度线索时,大型语言模型与人类的一致性比在仅伪词条件下更可靠;在仅伪词条件下,它们显著低于 fastText——一个字符 n-gram 模型。此外,可靠地驱动人类--fastText 一致性的亚词汇余弦相似性线索并未一致地迁移到人类--大型语言模型对齐中,而且推理 token 消耗与人类加工难度没有一致的关系。这些发现表明,大型语言模型不一定共享支配人类伪词加工的亚词汇线索;我们将分词和训练数据覆盖度作为候选解释进行讨论。
cs.CL / 51 / 2610.08026
The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation
幻觉检测基准中的标注问题:一项实证评估
large language model
大语言模型相关
Abstract
In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed. These methods are often benchmarked with open-domain question answering (QA) datasets containing questions and corresponding short reference answers. First, an LLM is used to generate answers to questions within the QA dataset. Then, some automated labeling strategy is used to label these answers as hallucinated or not by comparing them with the reference answers in the dataset. This evaluation setting creates a methodological ambiguity between two criteria: reference faithfulness (whether the answer is fully supported by the reference) and factual correctness (whether the answer is free from contradictions and factually false specific claims). In practice, automated labelers may apply the former criterion even when the intended target is the latter. We study this potential criterion mismatch using 900 human-labeled question-answer pairs spanning three commonly used QA datasets and three generator models, with labels targeting answer-level factual correctness. We evaluate lexical similarity metrics, a reference-entailment NLI baseline, and seven LLM judges under controlled prompt variants as automated labelers. Our experiments reveal substantial disagreement both among automated labeling strategies and between these labels and human annotations. Many strategies also exhibit strong directional error biases, and for most judge-generator pairs, replacing a faithfulness-oriented prompt with a factual-correctness prompt improves agreement with human annotations and reduces false-positive dominance, indicating that automated hallucination labels depend strongly on how the target criterion is specified. Label-source choice should therefore be considered a fundamental part of benchmark design and made explicit, validated, and matched with the benchmark goal.
Chinese Translation
近年来,已经发展出若干用于检测大语言模型(LLM)何时产生幻觉的方法。这些方法通常使用包含问题及相应简短参考答案的开放域问答(QA)数据集进行基准测试。首先,使用一个 LLM 为 QA 数据集中的问题生成答案。然后,通过将这些答案与数据集中的参考答案进行比较,使用某种自动标注策略将这些答案标注为幻觉或非幻觉。这一评估设置在两个标准之间造成了方法论上的歧义:参考忠实性(答案是否被参考答案完全支持)与事实正确性(答案是否不含矛盾以及事实上错误的特定断言)。在实践中,即使预期目标是后者,自动标注器也可能应用前者标准。我们使用 900 个人工标注的问答对来研究这一潜在的标准不匹配问题,这些问答对涵盖三个常用 QA 数据集和三个生成器模型,其标签针对答案层面的事实正确性。我们在受控的提示词变体下,评估了词汇相似度指标、一个参考蕴含 NLI 基线以及七个 LLM 评判器作为自动标注器的表现。我们的实验揭示了自动标注策略之间存在显著分歧,且这些标签与人工标注之间也存在显著分歧。许多策略还表现出强烈的方向性错误偏差,并且对于大多数评判器-生成器组合而言,将面向忠实性的提示词替换为事实正确性提示词能够提高与人工标注的一致性,并减少假阳性占主导的现象,这表明自动幻觉标签强烈依赖于目标标准是如何被界定的。因此,标签来源的选择应被视为基准设计的根本组成部分,并应被明确说明、经过验证,且与基准目标相匹配。
cs.CL / 52 / 2610.08037
Are Language Models Script-Aware?
语言模型具备文字系统意识吗?
large language model
大语言模型相关
Abstract
Language models frequently generate outputs in unintended languages or scripts, a phenomenon known as off-target generation. While existing research has focused on language selection, the dimension of script knowledge remains understudied: before any linguistic understanding can occur, users must recognize the graphic symbols in a model's response. We investigate whether Small and Large Language Models (SLMs and LLMs) possess script knowledge by testing them on multi-scriptic languages. Through two complementary experiments, we evaluate whether models (1) adapt their output script to match the input, and (2) follow explicit instructions to generate text in a specified script. The models we tested demonstrate substantial script knowledge: they all achieve a near-perfect Latin script fidelity (more than 98%) and follow script instructions with high frequency. Nevertheless, we notice differences between LLMs and SLMs, with higher scores for LLMs including for non-standard script combinations.
Chinese Translation
语言模型经常以非预期的语言或文字系统生成输出,这一现象被称为偏离目标生成。尽管现有研究聚焦于语言选择,但文字系统知识这一维度仍未得到充分研究:在任何语言理解能够发生之前,用户必须识别模型响应中的图形符号。我们通过在多种文字系统的语言上测试小型和大型语言模型(SLMs 和 LLMs),研究它们是否具备文字系统知识。通过两个互补实验,我们评估模型是否(1)调整其输出文字系统以匹配输入,以及(2)遵循明确指令以指定的文字系统生成文本。我们测试的模型表现出相当丰富的文字系统知识:它们都达到了近乎完美的拉丁文字系统保真度(超过 98%),并以高频率遵循文字系统指令。尽管如此,我们注意到 LLMs 和 SLMs 之间存在差异,LLMs 的得分更高,包括在非标准文字系统组合上也是如此。
cs.CL / 53 / 2610.08153
Penalty-Framed No-Valid-Option MCQA: Analyzing LLM Abstention under Invalid Choices
惩罚框架下的无有效选项 MCQA:分析无效选项情境下 LLM 的弃答行为
large language model
大语言模型相关
Abstract
Multiple-choice question answering (MCQA) is commonly used to evaluate large language models under the assumption that one of the provided options is correct, typically using answer-selection accuracy. However, in real deployments, users or retrieval systems may provide invalid option sets in which none of the listed choices is correct, and selecting one of them may incur downstream cost. We study this setting as penalty-framed no-valid-option MCQA. Using the mathematics subset of MMLU-Pro, we remove the labeled correct option, allow models to either choose a remaining option or output ABSTAIN, and penalize invalid forced-choice responses. We further introduce correct-conditioned analysis, evaluating abstention only on instances that the model originally answered correctly. Experiments show that high MCQA accuracy does not fully guarantee abstention reliability: even under explicit no-valid-option-aware instructions and penalty-based scoring, models still produce invalid forced-choice responses for a subset of originally correct instances. These results show that penalty-framed no-valid-option MCQA reveals an aspect of model reliability not captured by standard answer-selection accuracy.
Chinese Translation
多项选择问答(MCQA)通常被用于评估大语言模型,其前提假设是所提供的选项中有一个是正确的,并通常以答案选择准确率作为评价指标。然而,在实际部署中,用户或检索系统可能提供无效的选项集合,其中所列出的选项没有一个正确,而选择其中之一可能会带来下游代价。我们将这一设定作为惩罚框架下的无有效选项 MCQA 加以研究。我们使用 MMLU-Pro 的数学子集,移除被标注为正确的选项,允许模型要么选择一个剩余选项,要么输出 ABSTAIN,并对无效的强制选择回应施加惩罚。我们进一步引入正确条件化分析,仅在模型原本回答正确的实例上评估弃答行为。实验表明,高 MCQA 准确率并不能完全保证弃答的可靠性:即使在明确告知无有效选项的指令以及基于惩罚的评分之下,模型仍会在一部分原本回答正确的实例上产生无效的强制选择回应。这些结果表明,惩罚框架下的无有效选项 MCQA 揭示了标准答案选择准确率未能捕捉到的模型可靠性的一个方面。
cs.CL / 54 / 2610.08303
Language Unalignability: Why Some Concepts Resist Cross-Cultural Benchmark Evaluation
语言不可对齐性:为何某些概念抗拒跨文化基准评估
large language model
大语言模型相关
Abstract
Current evaluation of multilingual Large Language Models (LLMs) rests on an implicit Translation-Isomorphism Assumption (TIA): that semantic structures across languages are congruent and mutually mappable without loss of information. We argue that this assumption is not merely violated in practice, but ill-posed in principle for a typologically identifiable class of concepts, including pragmatic markers, honorifics, and diachronically stratified terms. We formalize this failure using a usage-cloud framework, representing concepts as point sets of contextualized embeddings. We define $α$-unalignability as the impossibility of any mapping that simultaneously preserves lexical faithfulness (centroid correspondence) and structural faithfulness (local neighborhood topology). We provide three layers of evidence. Behaviorally, we show that FLORES-200 translation failures are predicted by language family and resource class but not by script, and that LOBSTER reasoning scores vary by family. Mechanistically, we report a Representation-Intervention Gap (RIG) in a nine-model case study on Yami: the models' activations encode a regularity along which Yami groups with other low-resource and Austronesian languages, yet interventions on language-specific neurons show no demonstrated advantage over random masks: the regularity is visible but not usable by this intervention. Finally, we operationalize these findings into a multidimensional diagnostic profile: Cycle-Consistency, Pragmatic-Load Disagreement, Manifold-Curvature Mismatch, and RIG. We argue that collapsing cultural competence into a single scalar incentivizes "probabilistic flattening," and that recognizing the unalignable class is a precondition for AI that respects, rather than erases, cultural divergence. This suggests that multilingual alignment is not a single well-defined objective, but a set of mutually incompatible projections.
Chinese Translation
当前对多语言大型语言模型(LLMs)的评估建立在一个隐含的翻译同构假设(TIA)之上:即跨语言的语义结构是全等的,并且能够在不丢失信息的情况下相互映射。我们认为,这一假设不仅在实践中被违反,而且对于一类在类型学上可识别的概念——包括语用标记、敬语和历时分层词项——在原则上就是不适定的。我们使用用法云框架将这种失败形式化,将概念表示为上下文化嵌入的点集。我们将 $α$-不可对齐性定义为:任何同时保持词汇忠实性(质心对应)和结构忠实性(局部邻域拓扑)的映射都不可能存在。我们提供三个层面的证据。在行为层面,我们表明 FLORES-200 翻译失败可由语系和资源类别预测,而不能由文字系统预测,并且 LOBSTER 推理分数因语系而异。在机制层面,我们在一项针对雅美语的九模型案例研究中报告了一个表示–干预差距(Representation-Intervention Gap, RIG):模型的激活编码了一种规律性,沿着这种规律性,雅美语与其他低资源语言和南岛语系语言聚为一类,然而针对语言特异性神经元的干预并未显示出相对于随机掩码的已证明优势:这种规律性可见,但不能被这种干预所利用。最后,我们将这些发现操作化为一个多维诊断画像:循环一致性(Cycle-Consistency)、语用负载分歧(Pragmatic-Load Disagreement)、流形曲率失配(Manifold-Curvature Mismatch)以及 RIG。我们认为,将文化能力压缩为一个单一标量会激励“概率扁平化”,并且承认不可对齐类别,是构建尊重而非抹除文化差异的人工智能的前提条件。这表明,多语言对齐并不是一个单一且定义良好的目标,而是一组相互不兼容的投影。
cs.CL / 55 / 2610.08388
Foresight-over-Graph: Reasoning Beyond Local Horizons for Knowledge Base Question Answering
Foresight-over-Graph:面向知识库问答的超越局部视野的推理
large language model
大语言模型相关
Abstract
Large language models (LLMs) have demonstrated strong capabilities in question answering, yet they still frequently suffer from hallucinations on knowledge-intensive tasks. Knowledge graphs (KGs) provide LLMs with structured, interpretable, and updatable factual grounding, making them a promising external knowledge source for reliable reasoning. However, existing LLM-guided graph reasoning methods typically rely on hop-wise greedy or beam-style pruning during evidence retrieval. Such local decision processes are inherently myopic: evidence that appears weak near the source may become crucial only after deeper graph context is explored, causing answer-critical branches to be discarded prematurely and making the reasoning chain difficult to recover. To address this limitation, we propose Foresight-over-Graph (FoG), a foresight-aware evidence retrieval framework for knowledge base question answering (KBQA). FoG iteratively constructs a question-relevant evidence subgraph and uses far-to-near feedback to guide path exploration, and maintains a compact memory subgraph to support continued exploration. Extensive experiments on widely used KBQA benchmarks demonstrate that FoG achieves state-of-the-art performance, with a particularly large improvement of 16.58% in Hit on CWQ, while also reducing LLM calls and token usage. Our code is available at https://github.com/yhong7/FoG .
Chinese Translation
大语言模型(LLMs)在问答方面已展现出强大的能力,但它们在知识密集型任务上仍经常遭受幻觉问题的困扰。知识图谱(KGs)为LLMs提供了结构化、可解释且可更新的事实依据,使其成为用于可靠推理的极具前景的外部知识来源。然而,现有的LLM引导的图推理方法在证据检索过程中通常依赖逐跳贪心或束式剪枝。此类局部决策过程本质上是短视的:在源节点附近显得较弱的证据,可能只有在探索更深层的图上下文之后才会变得关键,这会导致对答案至关重要的分支被过早丢弃,并使推理链难以恢复。为解决这一局限,我们提出Foresight-over-Graph(FoG),一个面向知识库问答(KBQA)的前瞻感知证据检索框架。FoG迭代地构建与问题相关的证据子图,并利用由远及近的反馈来引导路径探索,同时维护一个紧凑的记忆子图以支持持续探索。在广泛使用的KBQA基准上的大量实验表明,FoG达到了最先进的性能,其中在CWQ上的Hit指标尤其取得了16.58%的大幅提升,同时还减少了LLM调用次数和token使用量。我们的代码可在 https://github.com/yhong7/FoG 获取。
cs.CL / 56 / 2610.08413
Knowing When Not to Answer: Cross-Domain and Multi-Turn Generalization of Latent Underspecification Signals
知道何时不回答:潜在欠指定信号的跨领域与多轮泛化
large language model
大语言模型相关
Abstract
Large language models routinely answer questions that cannot be answered from the information given, and in dialogue they answer before enough has been said. Unanswerability is linearly decodable from hidden states, but it is unclear which of its forms share a representation and whether the signal is useful in dialogue. We contribute a turn-labeled multi-turn benchmark (423 conversations, 1,661 labeled turn-states) and an evaluation harness with a simulated user who answers clarifying questions, and use them with six datasets and six open-weight LLMs to test how far probes for unanswerability carry. Probes transfer robustly between datasets that share a ground of unanswerability: missing information in math (AUROC 0.77-0.97) and in a passage (SQuAD 2.0<->MuSiQue, 0.77-0.90). Probes for epistemic "known-unknowns" transfer poorly to math, but this separation weakens under lexical controls and changes with layer and coordinate system, so it remains unresolved. Single-turn probes fail zero-shot to detect when a conversation becomes answerable; in-structure probes recover it, but no better than a bag-of-words classifier. A gate on the calibrated probe, with no model fine-tuning, fires on underspecified turns far more precisely than chance, and its end-task success comes within 0.08 of a gate given the true labels. Yet across four models it does not reliably beat vanilla generation or prompted consolidation. The remaining gap lies mostly in how models use a clarification, not in detection.
Chinese Translation
大型语言模型经常回答那些无法根据所给信息回答的问题,并且在对话中,它们会在尚未说足够多之前就作答。不可回答性可从隐藏状态中线性解码,但尚不清楚它的哪些形式共享同一表示,以及该信号在对话中是否有用。我们贡献了一个按轮次标注的多轮基准(423 段对话,1,661 个已标注轮次状态)以及一个带有模拟用户的评估工具,该模拟用户会回答澄清性问题;我们将其与六个数据集和六个开放权重 LLM 一起使用,以测试不可回答性探针能迁移到何种程度。探针在共享同一不可回答性基础的数据集之间稳健迁移:数学中的缺失信息(AUROC 0.77-0.97)以及段落中的缺失信息(SQuAD 2.0<->MuSiQue,0.77-0.90)。针对认知上的“已知的未知”的探针向数学迁移效果很差,但这种分离在词汇控制下会减弱,并随层和坐标系而变化,因此仍未得到解决。单轮探针在零样本下无法检测对话何时变得可回答;结构内探针能恢复这一检测,但并不比词袋分类器更好。一个基于校准探针的门控,在无需模型微调的情况下,在欠指定轮次上触发,其精确度远高于随机水平,并且其最终任务成功率与给定真实标签的门控相差在 0.08 以内。然而,在四个模型上,它并不能可靠地胜过原始生成或提示式整合。剩余的差距主要在于模型如何使用澄清,而不在于检测。
cs.CL / 57 / 2610.08452
Agentic AutoRAG: RAG Pipeline Optimization through Reasoning-Driven Agents
Agentic AutoRAG:通过推理驱动的智能体实现 RAG 流水线优化
large language model
大语言模型相关
Abstract
Retrieval-augmented generation (RAG) is a widely used approach for grounding large language models (LLMs) in external knowledge. However, configuring a pipeline is an expensive hyperparameter optimization problem over many interacting choices, from chunking and embedding model to reranking and generation. Existing optimizers, from greedy search to Bayesian optimization, reduce each trial to an aggregate score and search without modeling why a configuration performed as it did, even though the retrieved chunks already provide evidence about whether each failure occurred during retrieval or after it. We introduce Agentic AutoRAG, an LLM-agent optimizer for multi-objective RAG hyperparameter optimization with retrieval-versus-generation failure attribution. It proposes configurations scored on a frozen exam from the corpus: after each trial a Diagnoser attributes each failed question to retrieval or generation, and a Proposer, grounded in a knowledge base of model rankings and pricing, selects the next configuration, weighing accuracy against cost to trace a Pareto frontier. On three multi-hop QA benchmarks it reaches higher LLM-judge accuracy than every baseline we compare, and within its first 10 trials it matches or beats the statistical baselines' full 30-trial judge accuracy. In its cost-aware mode on a real-world healthcare corpus it reaches a median exam accuracy of 77%, above the strongest baseline's 71.5%, at about 58% of that baseline's cost per query, and it matches that 71.5% at about 22% of the cost.
Chinese Translation
检索增强生成(RAG)是一种被广泛使用的方法,用于将大语言模型(LLM)锚定在外部知识之上。然而,配置一条流水线是一个代价高昂的超参数优化问题,它涉及众多相互影响的选择,从分块与嵌入模型到重排序与生成。现有的优化器,从贪心搜索到贝叶斯优化,都把每一次试验简化为一个聚合分数并据此搜索,而不对某个配置为何表现如此进行建模,尽管检索到的文本块已经提供了证据,说明每一次失败是发生在检索阶段还是发生在检索之后。我们提出 Agentic AutoRAG,这是一种用于多目标 RAG 超参数优化的 LLM 智能体优化器,并带有检索与生成之间的失败归因。它提出的配置在来自语料的固定试题集上评分:每次试验之后,一个 Diagnoser 将每个失败问题归因于检索或生成,而一个 Proposer 以模型排名与定价的知识库为基础,选择下一个配置,在准确率与成本之间进行权衡,以描绘帕累托前沿。在三个多跳问答基准上,它达到的 LLM 评判准确率高于我们所比较的每一个基线,并且在其前 10 次试验内就达到或超过了统计基线完整 30 次试验的评判准确率。在真实世界的医疗语料上,其成本感知模式达到 77% 的考试准确率中位数,高于最强基线的 71.5%,而每次查询成本约为该基线的 58%;并且它以约 22% 的成本达到该 71.5%。
cs.CL / 58 / 2610.08559
Latent space bias directions in LLMs capture confidence, not fairness
大语言模型中的潜在空间偏见方向捕捉的是置信度,而非公平性
large language model
大语言模型相关
Abstract
Activation steering has gained popularity as a lightweight inference-time debiasing technique for large language models. However, prior work reports that steering vectors generalise poorly, with unintended effects on model performance and limited transfer to new datasets. Our work analyses what the debiasing direction used for activation steering actually encodes, in order to shed light on its inconsistent performance. We study the linear debiasing direction obtained by contrasting the activations of anti-biased and biased prompts, and evaluate it as a steering intervention across bias and general knowledge benchmarks. We find that this direction is dominated by model confidence, pointing from regions of high to low-probability tokens in activation space rather than encoding a meaningful representation of model bias. Steering along it does reduce measured bias, but this is a consequence of reducing model confidence: on QA benchmarks we find that this steering drives the model to abstain from answering, with a side effect of improving fairness metrics. Our experiments show that model confidence is the dominant separating factor between biased and anti-biased prompts in hidden space, indicating that isolating a linear representation of bias which is disentangled from model confidence is difficult and steering-based debiasing results should be interpreted with care. In short, steering appears to reduce bias, not by correcting the model's underlying preferences, but by making it less confident, even on tasks unrelated to bias.
Chinese Translation
激活引导作为一种轻量级的推理时去偏技术,已在大语言模型中获得广泛使用。然而,先前的研究报告称,引导向量泛化能力较差,会对模型性能产生非预期影响,且难以迁移到新数据集。我们的工作分析了用于激活引导的去偏方向究竟编码了什么,以揭示其性能不稳定的原因。我们研究了通过对比反偏见提示与偏见提示的激活所得到的线性去偏方向,并将其作为引导干预在偏见基准和通用知识基准上进行评估。我们发现,该方向主要由模型置信度主导,在激活空间中从高概率词元区域指向低概率词元区域,而不是编码了模型偏见的有意义表征。沿该方向进行引导确实会降低所测得的偏见,但这是降低模型置信度的结果:在问答基准上,我们发现这种引导会促使模型放弃作答,其副作用是改善了公平性指标。我们的实验表明,在隐藏空间中,模型置信度是区分偏见提示与反偏见提示的主导因素,这表明分离出一个与模型置信度解耦的偏见的线性表征是困难的,基于引导的去偏结果应当被谨慎解读。简而言之,引导似乎能够减少偏见,但并非通过纠正模型的潜在偏好,而是通过使模型变得不那么自信,即使在那些与偏见无关的任务上也是如此。
cs.CL / 59 / 2610.08585
Incidental information contaminates patient notes and disrupts clinical reasoning in large language models
附带信息污染患者记录并扰乱大型语言模型中的临床推理
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly relied upon to support ambient documentation and clinical reasoning. Here we examine the impact of a failure mode shared between these two applications by assessing their sensitivity to information incidental to the patient encounter. In 576 patient-clinician dialogues, we found that frontier models inserted small-talk exchanges into 35% of notes, while mean quality scores changed by at most 0.20 points on five-point scales. In 3.7% of frontier notes, models misattributed the asides or used them clinically. In 57 mock recorded consultations, background speech from a separate patient encounter at -10 dB leaked into 48.2% of transcripts, with contamination detected in 5.3% of downstream notes generated by four open-weight models. We propose a dual encoding hypothesis of clinical reasoning and distraction in LLMs, with preliminary evidence that LLM components associated with disruption by incidental information also support clinical reasoning. These findings support evaluating resistance to incidental information before clinical use, with safeguards that prevent contamination while preserving clinical reasoning.
Chinese Translation
大型语言模型(LLMs)正越来越多地被依赖,用于支持环境记录和临床推理。在此,我们通过评估它们对患者就诊过程中附带信息的敏感性,考察这两种应用共有的一个失效模式的影响。在576段患者-临床医生对话中,我们发现前沿模型将闲聊交流插入到35%的记录中,而平均质量得分在五分制量表上最多变化0.20分。在3.7%的前沿模型记录中,模型将这些题外话错误归因,或将其用于临床。在57次模拟录制的会诊中,来自另一次患者就诊的、音量为-10 dB的背景语音泄漏到48.2%的转录文本中,并在四个开放权重模型生成的下游记录中有5.3%检测到污染。我们提出大型语言模型中临床推理与分心的双重编码假说,并有初步证据表明,与附带信息造成干扰相关的LLM组件也支持临床推理。这些发现支持在临床使用前评估对附带信息的抵抗力,并采用能够在防止污染的同时保留临床推理的保障措施。
cs.CL / 60 / 2610.08630
Towards In-Parameter Memory Augmentation for Large Language Models
迈向大型语言模型的参数内记忆增强
large language model
大语言模型相关
Abstract
Recently Large Language Models (LLMs) and LLM-based agents increasingly need to incorporate knowledge acquired after pretraining, e.g., domain facts, user preferences, documents, and interaction experience. In-context learning (ICL) and ICL-based agent harness remain flexible, but they consume context capacity and incur repeated discretized encoding cost that grows with context length. \textbf{In-parameter memory} offers a complementary substrate: reusable memory information is represented in model parameters, adapters, or other parameter-like objects that are composed into the forward pass at inference time. This survey focuses on methods that augment LLMs with such parametric memory at deployment: a memory-bearing parameter object is plugged into the forward pass during inference, whether it is acquired before or during deployment. We organize the landscape with two orthogonal axes: \textbf{Parameter Placement}, which includes Embedding, Attention, FFN layers, or Hybrid when two or more layers are used; and \textbf{Parameter Acquisition Time}, which distinguishes methods whose memory object is acquired during deployment (online) from those acquired before it (offline). We clarify boundaries, conduct comparisons, and discuss open directions in interference, safety, co-design with ICL, and recursive self-improvement.
Chinese Translation
最近,大型语言模型(LLMs)和基于 LLM 的智能体越来越需要融入预训练之后获得的知识,例如领域事实、用户偏好、文档和交互经验。上下文学习(ICL)和基于 ICL 的智能体框架仍然灵活,但它们消耗上下文容量,并产生随上下文长度增长的重复离散化编码成本。\textbf{参数内记忆} 提供了一种互补的基底:可复用的记忆信息表示在模型参数、适配器或其他类似参数的对象中,这些对象在推理时被组合到前向传播中。本综述聚焦于在部署时用此类参数化记忆增强 LLM 的方法:一个承载记忆的参数对象在推理期间被接入前向传播中,无论它是在部署之前还是部署期间获得的。我们用两个正交轴来组织这一领域:\textbf{参数放置位置},包括嵌入层、注意力层、FFN 层,或当使用两个或更多层时的混合式;以及 \textbf{参数获取时间},其区分记忆对象在部署期间获取(在线)的方法与在部署之前获取(离线)的方法。我们阐明边界,进行比较,并讨论在干扰、安全性、与 ICL 的协同设计以及递归自我改进方面的开放方向。
cs.CL / 61 / 2610.08703
Agreement Is Not Validity: Cross-Model LLM Consensus in Diagnosing Student Failure Modes in K-12 Math Tutoring Dialogue
一致性并非效度:跨模型LLM共识在诊断K-12数学辅导对话中学生失败模式中的作用
large language model
大语言模型相关
Abstract
In K-12 mathematics tutoring, student-tutor dialogue provides rich evidence of learners' problem-solving processes and sources of difficulty. Learning analytics research increasingly relies on large language models (LLMs) to extract such information from dialogue for a variety of downstream tasks, including knowledge tracing, behavioral modeling, and diagnosis of student reasoning errors. However, the validity of these model-generated interpretations remains insufficiently understood. In this exploratory study, we examine the validity of LLM classifications of five student failure modes in mathematics tutoring dialogue using an operational diagnostic codebook: uncertainty, misattribution, operator selection, conceptual gap, and procedural slip. Across models, human-LLM agreement was moderate (kappa = .524-.597), while cross-model agreement was substantially higher (kappa = .755-.781; alpha = .769). These findings show that cross-model agreement can create a misleading appearance of correctness, challenging the assumption that consensus among LLMs constitutes evidence of valid learner interpretation. For learning analytics, the implication is clear: scalable labeling is useful only if the inferred constructs are valid, and model consensus cannot substitute for independent evidence of that validity.
Chinese Translation
在K-12数学辅导中,学生—导师对话为学习者的解题过程及其困难来源提供了丰富证据。学习分析研究日益依赖大语言模型(LLMs)从对话中提取这类信息,以支持多种下游任务,包括知识追踪、行为建模和学生推理错误诊断。然而,这些模型生成解释的效度仍未被充分理解。在这项探索性研究中,我们使用一套可操作的诊断编码手册,考察LLM对数学辅导对话中五种学生失败模式的分类效度:不确定性、错误归因、算子选择、概念缺口和程序性失误。跨模型来看,人类与LLM的一致性为中等(kappa = .524-.597),而跨模型一致性则显著更高(kappa = .755-.781;alpha = .769)。这些发现表明,跨模型一致性可能制造出一种具有误导性的正确性表象,挑战了LLM之间的共识构成有效学习者解释证据这一假设。对于学习分析而言,其启示很明确:可扩展的标注只有在其推断出的构念有效时才有用,而模型共识不能替代关于该效度的独立证据。
cs.CL / 62 / 2610.08738
Denoising Hierarchical Representations: Joint Continuous Diffusion for Language Modeling
去噪层次化表示:用于语言建模的联合连续扩散
diffusion
扩散模型相关
Abstract
Diffusion Language Models (DLMs) hold the promise of order-agnostic, parallel text generation. Recently, continuous diffusion and flow matching models have seen substantial gains, driven by carefully crafted token representations and diffusion/flow spaces. In this work, we introduce Hierarchical Continuous Diffusion Language Models (H-CDLMs), a simple framework that further improves continuous DLMs with minimal compute and parameter overhead. Drawing on the discrete DLM and continuous image diffusion literature on joint diffusion, we diffuse multiple modalities in parallel. These modalities represent tokens at different semantic granularities: in our instantiation, the tokens themselves and coarser clusters obtained by clustering pretrained token embeddings. We propose a general setup that allows per-modality samplers and schedules to enhance the interplay between modalities. Applied to CoBit, this yields H-CoBit, which delivers large empirical gains across benchmarks. At dataset entropy, H-CoBit improves MAUVE and reaches a generative perplexity (GenPPL) of 49.4 on LM1B and 50.4 on OWT, improving on the baseline by 24.2 and 20.7 points and surpassing even discrete DLMs of comparable size. On GSM8K, it reaches 27.4% accuracy, outperforming prior continuous diffusion and flow-based models. We further apply H-CDLM to the flow matching model FLM, obtaining consistent gains with H-FLM and demonstrating that the framework generalizes across continuous generative paradigms. Our code will be made publicly available at https://github.com/matol-16/HCDLM.git .
Chinese Translation
扩散语言模型(DLMs)有望实现顺序无关的并行文本生成。最近,连续扩散模型和流匹配模型取得了显著进展,这得益于精心设计的词元表示以及扩散/流空间。在这项工作中,我们提出了层次化连续扩散语言模型(H-CDLMs),这是一个简单的框架,能以最小的计算和参数开销进一步改进连续 DLMs。借鉴离散 DLM 和连续图像扩散文献中关于联合扩散的研究,我们并行地对多种模态进行扩散。这些模态在不同语义粒度上表示词元:在我们的实例化中,即词元本身以及通过聚类预训练词元嵌入得到的更粗粒度簇。我们提出了一个通用设置,允许每种模态各自的采样器和调度来增强模态之间的相互作用。将其应用于 CoBit 后,这产生了 H-CoBit,它在各个基准上带来了巨大的实证增益。在数据集熵下,H-CoBit 提升了 MAUVE,并在 LM1B 上达到 49.4、在 OWT 上达到 50.4 的生成困惑度(GenPPL),相比基线分别提升了 24.2 和 20.7 个点,甚至超过了规模相当的离散 DLMs。在 GSM8K 上,它达到 27.4% 的准确率,优于先前的连续扩散和基于流的模型。我们进一步将 H-CDLM 应用于流匹配模型 FLM,得到 H-FLM 并取得了一致的增益,表明该框架可泛化到不同的连续生成范式。我们的代码将在 https://github.com/matol-16/HCDLM.git 公开提供。
cs.CL / 63 / 2610.08747
The Missing Minimal Pair: Stereotype Evaluation in LLMs
缺失的最小对:大语言模型中的刻板印象评估
large language model
大语言模型相关
Abstract
A common approach to measuring bias in Large Language Models is to compare the log-likelihoods of two contrastive stereotype sentences. We argue that such single-pair comparisons are often unreliable: simply rewriting the same stereotype with an alternative attribute can yield logically inconsistent preferences. To address this, we propose a dual minimal pair setup that introduces two axes of comparison for robust stereotype evaluation. First, we present a data-augmentation framework that fills critical gaps in existing stereotype datasets by generating paraphrases and alternate attributes. We apply our framework on a set of English, Russian, Spanish and Chinese stereotypes. Second, we introduce two evaluation metrics tailored to the dual minimal pair setup. One of these metrics provides a new perspective on bias by modeling the mutual information (MI) between social groups and stereotyped attributes. This MI-based metric is better suited for aggregation and enables more robust comparisons of stereotype strength across different languages and models. Our code is available at https://github.com/stepanat/missing-minimal-pair/.
Chinese Translation
测量大语言模型中偏见的一种常见方法是比较两个对比性刻板印象句子的对数似然。我们认为,这种单对比较往往不可靠:仅用另一个属性重写同一个刻板印象,就可能产生逻辑上不一致的偏好。为解决这一问题,我们提出一种双重最小对设置,引入两个比较轴,以实现稳健的刻板印象评估。首先,我们提出一个数据增强框架,通过生成同义改写和替代属性,填补现有刻板印象数据集中的关键空白。我们将该框架应用于一组英语、俄语、西班牙语和中文的刻板印象。其次,我们介绍两个专为双重最小对设置定制的评估指标。其中一个指标通过建模社会群体与刻板印象属性之间的互信息(MI),为偏见提供了新的视角。这种基于 MI 的指标更适合聚合,并且能够在不同语言和模型之间对刻板印象强度进行更稳健的比较。我们的代码可在 https://github.com/stepanat/missing-minimal-pair/ 获取。
cs.CR / 64 / 2610.07256
Efficient Auditing of Adversarial AI Agent Behavior from Agent Traces
从智能体轨迹中高效审计对抗性 AI 智能体行为
large language model
大语言模型相关
Abstract
AI agents powered by large language models (LLMs) can perform complex tasks but may harm the systems they operate in, either intentionally or unintentionally. Existing agent monitoring approaches rely on rule-based guardrails or LLM-based trace auditing. However, rule-based guardrails can be bypassed through obfuscation and may miss harmful actions beyond their predefined rules, whereas applying an LLM to audit every action is costly. We present a two-stage agent trace auditing framework. The first stage uses single-event and trace-sequence rules to select pending actions for inspection; the second uses an LLM audit agent to examine each selected action in the context of the agent's preceding trace before execution. We jointly refine the gate rules and audit instructions using training data, allowing the framework to adapt to complex agent behaviors rather than relying solely on predefined rules. On the public benchmark OpenAgentSafety, our framework reduces the average number of LLM audits from 8.15 to 2.33 per run and token usage from 47.8k to 14.6k, with a detection rate of 72.8\% compared with 81.5\% when every action is audited. In two simulated multi-agent case studies, the framework flags all malicious traces while reducing audit token usage by more than 80\%.
Chinese Translation
由大型语言模型(LLMs)驱动的 AI 智能体能够执行复杂任务,但可能有意或无意地损害其运行所在的系统。现有的智能体监控方法依赖于基于规则的防护栏或基于 LLM 的轨迹审计。然而,基于规则的防护栏可以通过混淆被绕过,并且可能遗漏其预定义规则之外的有害动作,而应用 LLM 来审计每一个动作则成本高昂。我们提出一个两阶段的智能体轨迹审计框架。第一阶段使用单事件规则和轨迹序列规则来选择待检查的待执行动作;第二阶段使用 LLM 审计智能体,在执行之前结合智能体的先前轨迹上下文来审查每个选定的动作。我们使用训练数据联合优化门控规则和审计指令,使该框架能够适应复杂的智能体行为,而不是仅仅依赖于预定义规则。在公共基准 OpenAgentSafety 上,我们的框架将每次运行的平均 LLM 审计次数从 8.15 次减少到 2.33 次,并将 token 使用量从 47.8k 减少到 14.6k,检测率为 72.8\%,而审计每个动作时的检测率为 81.5\%。在两个模拟的多智能体案例研究中,该框架标记出所有恶意轨迹,同时将审计 token 使用量减少超过 80\%。
cs.CR / 65 / 2610.07639
HarnessSecurity-Bench: Do Security Mechanisms Really Protect Coding Agent Harnesses?
HarnessSecurity-Bench:安全机制真的能保护编码智能体执行框架吗?
large language model
大语言模型相关
Abstract
Coding agent harnesses mediate tool use and authorize actions, yet their security mechanisms and runtime effects remain incompletely characterized. We present HarnessSecurity, the first systematic empirical study and benchmark of open- and closed-source coding agent harnesses. First, we derive a ten-mechanism taxonomy and then assess 400 harness-mechanism cells using independent ratings by researchers and large language model (LLM) judges. We find that about half of confirmed mechanism implementations are opt-in, while closed-source harnesses exhibit substantial evidence gaps. Second, we introduce HarnessSecurity-Bench, a benchmark of 23 tasks across five attack surfaces without sacrificing legitimate task requirements. Using separate deterministic oracles to measure task utility and attack effects with security setting comparisons, we evaluate nine mechanisms across six leading harnesses: Claude Code, Codex CLI, Gemini CLI, gptme, Qwen Code, and GitHub Copilot. Under a controlled LLM baseline GLM-5.2, we conduct 2,500 trials, recording 81,155 tool calls and over 2.2 billion tokens. Enabling auto-approve increases utility and raises attack success from 29.2% to 95.6%. Network isolation and read-only mode reduce attack effects with substantial utility losses, while command allowlisting and command denylisting reduce attack effects with a small utility loss and a utility gain, respectively. Task-level cases show that restrictions on a shared capability can obstruct both legitimate and malicious operations, and that allowed tools or commands can leave unauthorized operations reachable through alternative execution paths. Harness providers should make security settings verifiable, test alternative execution paths to protected operations, and assess attack effects alongside task utility and execution costs.
Chinese Translation
编码智能体执行框架对工具使用进行中介并授权操作,但其安全机制和运行时效果仍未得到充分刻画。我们提出 HarnessSecurity,这是首个针对开源和闭源编码智能体执行框架的系统性实证研究与基准。首先,我们推导出一个十机制分类法,然后使用研究者和大型语言模型(LLM)评判者的独立评分,评估了 400 个执行框架-机制单元。我们发现,约一半已确认的机制实现是选择加入式的,而闭源执行框架则表现出大量证据缺口。其次,我们引入 HarnessSecurity-Bench,这是一个涵盖五个攻击面的 23 项任务的基准,且不牺牲合法任务需求。使用单独的确定性预言机来测量任务效用和攻击效果,并进行安全设置比较,我们评估了六个领先执行框架中的九种机制:Claude Code、Codex CLI、Gemini CLI、gptme、Qwen Code 和 GitHub Copilot。在受控的 LLM 基线 GLM-5.2 下,我们进行了 2,500 次试验,记录了 81,155 次工具调用和超过 22 亿个 token。启用自动批准会提高效用,并将攻击成功率从 29.2% 提高到 95.6%。网络隔离和只读模式会降低攻击效果,但伴随显著的效用损失;而命令允许列表和命令拒绝列表则分别以较小的效用损失和效用增益降低攻击效果。任务级案例表明,对共享能力的限制可能同时阻碍合法和恶意操作,并且被允许的工具或命令可能使未授权操作通过替代执行路径仍然可达。执行框架提供方应使安全设置可验证,测试通往受保护操作的替代执行路径,并连同任务效用和执行成本一起评估攻击效果。
cs.CR / 66 / 2610.07723
The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language Models
模型埋下触发器:多轮大型语言模型中的答案侧后门攻击
large language model
大语言模型相关
Abstract
Safety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks. Existing LLM backdoors are almost all input-centric: activation depends on explicit trigger patterns in the user input, so modern guardrails are built to sanitize the input space. We challenge this assumption with a novel answer-side backdoor for multi-turn dialogue. Instead of inserting the trigger into the input, the adversary uses a benign first-turn prompt to naturally induce the model to generate a specific, seemingly innocuous word. Once merged into the dialogue history, this self-generated word becomes the trigger. When a later harmful query arrives, the model detects its own trigger and bypasses its safety refusal, while the user input stays perfectly clean. Across four LLMs, our attack reaches near-perfect Attack Success Rates, approaching 100\% at only a 5\% poisoning rate, while preserving general utility and clean-input safety, and it evades mainstream input-centric defenses. Representation-level analysis shows that the self-generated trigger consistently suppresses the model's refusal signal, exposing a critical blind spot in current LLM defenses.
Chinese Translation
大型语言模型(LLM)中的安全对齐仍然容易受到后门攻击。现有的 LLM 后门几乎全都是以输入为中心的:其激活取决于用户输入中的显式触发器模式,因此现代防护栏被构建来净化输入空间。我们通过一种新颖的、用于多轮对话的答案侧后门来挑战这一假设。攻击者不是将触发器插入输入,而是使用一个良性的第一轮提示,自然地诱导模型生成一个特定的、看似无害的词。一旦合并到对话历史中,这个自生成的词就变成触发器。当之后出现有害查询时,模型检测到自己的触发器并绕过其安全拒绝,而用户输入保持完全干净。在四个 LLM 上,我们的攻击达到近乎完美的攻击成功率,在仅 5\% 的投毒率下接近 100\%,同时保持通用效用和干净输入安全性,并且它逃避了主流的以输入为中心的防御。表示层面分析表明,自生成触发器持续抑制模型的拒绝信号,暴露出当前 LLM 防御中的一个关键盲点。
cs.CR / 67 / 2610.08061
ASCENT: First-Order Optimal Fine-Tuning with Recalibration for Safety--Utility Co-Enhancement
ASCENT:面向安全—效用协同增强的一阶最优重校准微调
large language model
大语言模型相关
Abstract
Supervised fine-tuning can substantially improve the downstream utility of large language models (LLMs) but may compromise their safety. Existing safety-preserving methods constrain downstream updates using safety-related parameters or subspaces, but mainly focus on safety preservation rather than joint safety and utility enhancement, lack a theoretical characterization of the optimal safety-related subspace and safety-preserving task update, and typically rely on a static safety subspace that may become outdated during fine-tuning. To address these limitations, we propose ASCENT, a downstream fine-tuning framework for safety--utility co-enhancement through first-order optimal safety-aware periodic calibration and task optimization. We model safety as a function of LLM parameters $S(θ)$ and use its first-order approximation to characterize safety changes under parameter updates. Under a fixed rank and Frobenius-norm budget, we prove that the update constructed from the top-$r$ singular components of the safety-function gradient maximizes the estimated safety change, and use it for periodic calibration to preserve and improve safety. We further derive a unique safety-preserving task update that stays close to the original task update while penalizing negative effects on the estimated safety change. ASCENT alternates these optimal task and calibration updates to jointly enhance safety and utility. Experiments across multiple LLM families and downstream tasks show that ASCENT improves downstream utility by up to 20.3\% and reduces attack success rate by up to 35.5\%, achieving state-of-the-art safety and utility across all evaluated settings. Our code is available at https://github.com/ZJU-LLM-Safety/ASCENT.
Chinese Translation
有监督微调可以显著提升大型语言模型(LLMs)的下游效用,但可能损害其安全性。现有保安全方法使用与安全相关的参数或子空间来约束下游更新,但主要关注安全性保持,而不是安全性与效用的联合增强,缺乏对最优安全相关子空间和保安全任务更新的理论刻画,并且通常依赖一个在微调过程中可能变得过时的静态安全子空间。为解决这些局限,我们提出 ASCENT,一个通过一阶最优的安全感知周期性校准和任务优化实现安全—效用协同增强的下游微调框架。我们将安全性建模为 LLM 参数 $S(θ)$ 的函数,并使用其一阶近似来刻画参数更新下的安全性变化。在固定的秩和 Frobenius 范数预算下,我们证明,由安全函数梯度的前 $r$ 个奇异分量构造的更新能够最大化估计的安全性变化,并将其用于周期性校准,以保持并提升安全性。我们进一步推导出一个唯一的保安全任务更新,该更新在保持接近原始任务更新的同时,惩罚对估计的安全性变化的负面影响。ASCENT 交替执行这些最优的任务更新和校准更新,以协同增强安全性和效用。在多个 LLM 系列和下游任务上的实验表明,ASCENT 将下游效用提升最多 20.3\%,并将攻击成功率降低最多 35.5\%,在所有评估设置中实现了最先进的安全性和效用。我们的代码可在 https://github.com/ZJU-LLM-Safety/ASCENT 获取。
cs.CR / 68 / 2610.08316
MARCO: The Radioactive Watermark for Protein Generative Models
MARCO:用于蛋白质生成模型的放射性水印
diffusion
扩散模型相关
Abstract
Protein Generative Models (PGMs) have revolutionized structural biology by enabling the design of complex 3D protein structures from sequence data. However, this breakthrough introduces a dual-use challenge, exposing high-value PGMs to economic risks like unauthorized model extraction and biosecurity threats such as biohazard synthesis. To mitigate these threats, we propose \textbf{MARCO} (\textsc{COnformation waterMARk}), the first radioactive watermarking framework specifically tailored for PGMs. MARCO establishes a Dual-Layer defense that simultaneously protects intellectual property and ensures the forensic traceability of potential biosecurity misuses. (i) To preserve efficiency, MARCO iteratively embeds watermarks during diffusion reverse denoising via an auxiliary encoder-decoder, allowing the original PGM parameters to remain frozen for broad compatibility. (ii) To preserve biophysical fidelity and maximize robustness, we employ specialized loss functions targeting $C_α$-atom pairwise distances and torsion angles ($ψ, φ$) within an adversarial training framework integrated with stochastic attack simulations. (iii) Crucially, MARCO exhibits ``radioactivity'' where the watermark automatically transfers to the outputs of any pirate models trained on the watermarked data, effectively countering model extraction attacks. Comprehensive experiments demonstrate that MARCO achieves superior fidelity and robustness while successfully validating watermark transferability.
Chinese Translation
蛋白质生成模型(Protein Generative Models, PGMs)通过使复杂三维蛋白质结构能够从序列数据中设计出来,已经革新了结构生物学。然而,这一突破带来了双重用途挑战,使高价值的 PGMs 面临经济风险(如未经授权的模型提取)以及生物安全威胁(如生物危害合成)。为缓解这些威胁,我们提出 \textbf{MARCO}(\textsc{COnformation waterMARk}),这是首个专为 PGMs 量身定制的放射性水印框架。MARCO 建立了一种双层防御,同时保护知识产权,并确保对潜在生物安全滥用的取证可追溯性。(i) 为保持效率,MARCO 在扩散逆向去噪过程中通过辅助编码器-解码器迭代地嵌入水印,使原始 PGM 参数保持冻结以实现广泛兼容性。(ii) 为保持生物物理保真度并最大化鲁棒性,我们在与随机攻击模拟相结合的对抗训练框架内,采用针对 $C_α$-原子成对距离和扭转角($ψ, φ$)的专用损失函数。(iii) 关键的是,MARCO 展现出“放射性”,即水印会自动转移到任何在水印数据上训练的盗版模型的输出,从而有效对抗模型提取攻击。全面的实验表明,MARCO 在成功验证水印可转移性的同时,实现了优越的保真度和鲁棒性。
cs.CR / 69 / 2610.08406
Case-Level Verification in Scanner-LLM Cascades: Overcoming the Alert Aggregation Bottleneck to Expand the FRR-TPR Trade-off Space
扫描器-LLM 级联中的案例级验证:克服告警聚合瓶颈以扩展 FRR-TPR 权衡空间
large language model
大语言模型相关
Abstract
Dynamic Application Security Testing (DAST) scanners achieve high recall but also produce a large number of false positives, resulting in substantial manual triage costs. Large Language Models (LLMs), when used for independent detection, achieve extremely high recall (95.4%-100%) but also exhibit prohibitively high false positive rates (49.6%-85.0%), precluding their use as standalone replacements for scanners. A natural solution is a two-stage cascade consisting of scanner detection followed by LLM verification. However, a verification-granularity issue that has long been overlooked in practice creates a structural bottleneck: alert aggregation binds multiple true and false cases into a shared decision unit, such that removing a false positive inevitably eliminates true positives aggregated within the same alert group. This creates a trade-off bottleneck between the False-positive Reduction Rate (FRR) and the True-positive Rate (TPR). We formalize this bottleneck by showing that the alert-level false-positive set is a subset of the case-level false-positive set, and introduce a Case-Level, per-case verification strategy that shifts the decision granularity from the alert level to the instance level, independently replaying HTTP requests and making an independent determination for each detected case. Evaluation on the dual testbeds of Damn Vulnerable Web Application (DVWA) and WebGoat shows that the empirically best Alert-Level operating point achieves FRR=42.86% (TPR=51.7%). The Case-Level Baseline achieves FRR=47.6%, an improvement of 4.7 percentage points (+4.7 pp), while the Case-Focused Evidence Verification Prompt (CEV-Prompt) increases TPR from 55.2% to 62.1% at the same FRR.
Chinese Translation
动态应用安全测试(DAST)扫描器能够实现高召回率,但也会产生大量误报,从而导致高昂的人工分诊成本。大型语言模型(LLM)在被用于独立检测时能够实现极高的召回率(95.4%-100%),但也表现出高得令人难以承受的误报率(49.6%-85.0%),使其无法作为扫描器的独立替代方案。一种自然的解决方案是采用由扫描器检测后接 LLM 验证所组成的两阶段级联。然而,一个在实践中长期被忽视的验证粒度问题造成了结构性瓶颈:告警聚合将多个真案例和假案例绑定为一个共享的决策单元,以至于移除一个误报不可避免地会消除聚合在同一告警组内的真阳性。这在误报削减率(FRR)与真阳性率(TPR)之间造成了权衡瓶颈。我们通过证明告警级误报集合是案例级误报集合的子集,将该瓶颈形式化,并引入一种案例级、逐案例验证策略,将决策粒度从告警级转移到实例级,独立重放 HTTP 请求,并对每一个被检测到的案例作出独立判定。在 Damn Vulnerable Web Application(DVWA)与 WebGoat 这两个测试平台上的评估表明,经验上最佳的告警级工作点达到 FRR=42.86%(TPR=51.7%)。案例级基线达到 FRR=47.6%,提升了 4.7 个百分点(+4.7 pp),而面向案例的证据验证提示(Case-Focused Evidence Verification Prompt,CEV-Prompt)在相同 FRR 下将 TPR 从 55.2% 提升至 62.1%。
cs.CR / 70 / 2610.08590
TwinViT-DeepJSCC: Adversarially Robust Semantic Image Communication
TwinViT-DeepJSCC:对抗鲁棒的语义图像通信
diffusion
扩散模型相关
Abstract
Learning-based semantic communication is vulnerable to adversarial perturbations introduced before semantic encoding or over wireless channels. This paper proposes TwinViT-DeepJSCC, a preventive-corrective semantic image transceiver operating under a fixed channel-use budget. Two Vision Transformer (ViT)-based deep joint source-channel coding (DeepJSCC) branches learn complementary latent representations protected by sensitivity-aware masking. At the receiver, confidence-aware fusion, blind corruption-severity estimation, and signal-to-noise ratio (SNR)-severity-conditioned denoising diffusion implicit model (DDIM) purification mitigate residual corruption without requiring attack metadata. Experiments on the Canadian Institute for Advanced Research 100-class (CIFAR-100) dataset consider fast gradient sign method (FGSM), projected gradient descent (PGD), natural evolution strategies (NES), and Carlini-Wagner (CW) source-domain attacks, as well as random jamming and channel-aware adversarial waveforms over additive white Gaussian noise (AWGN) and block-flat Rayleigh fading. Under matched channel-use and attack budgets, TwinViT-DeepJSCC achieves maximum peak signal-to-noise ratio (PSNR) gains of approximately 9.5 dB under 20-step PGD and 10.8 dB under channel-aware waveform attacks over block-flat Rayleigh fading. Under PGD, it also improves Top-1 accuracy by up to approximately 38 percentage points over the undefended baseline and 13 percentage points over the strongest competing defense. Ablation results confirm the complementary contributions of the proposed transmitter- and receiver-side mechanisms.
Chinese Translation
基于学习的语义通信易受在语义编码之前或经由无线信道引入的对抗性扰动影响。本文提出 TwinViT-DeepJSCC,一种在固定信道使用预算下运行的预防-纠正式语义图像收发机。两个基于视觉 Transformer(ViT)的深度联合信源信道编码(DeepJSCC)分支学习由敏感度感知掩蔽保护的互补潜在表示。在接收端,置信度感知融合、盲损坏严重程度估计以及信噪比(SNR)-严重程度条件化的去噪扩散隐式模型(DDIM)净化,无需攻击元数据即可减轻残余损坏。在加拿大高等研究院 100 类(CIFAR-100)数据集上的实验考虑了快速梯度符号法(FGSM)、投影梯度下降(PGD)、自然进化策略(NES)和 Carlini-Wagner(CW)源域攻击,以及在加性白高斯噪声(AWGN)和块平坦瑞利衰落上的随机干扰和信道感知对抗波形。在匹配的信道使用和攻击预算下,TwinViT-DeepJSCC 在 20 步 PGD 下实现了约 9.5 dB 的最大峰值信噪比(PSNR)增益,在块平坦瑞利衰落上的信道感知波形攻击下实现了 10.8 dB 的增益。在 PGD 下,它还将 Top-1 准确率相较未防御基线最多提高约 38 个百分点,相较最强的竞争防御提高 13 个百分点。消融结果证实了所提出的发射端和接收端机制的互补贡献。
cs.CR / 71 / 2610.08678
Secure Speculative Decoding for Large Language Models
面向大型语言模型的安全推测解码
large language model
大语言模型相关
Abstract
Speculative decoding accelerates inference for a large language model (LLM), referred to as the \emph{target model}, by first using a smaller model, referred to as the \emph{draft model}, to generate candidate tokens and then verifying them with the target model for acceptance or rejection. Prior studies primarily focused on the efficiency-utility trade-off of speculative decoding, e.g., lossy speculative decoding, leaving its security implications largely unexplored. In this work, we bridge this gap by providing the \emph{first} systematic study of the security implications of speculative decoding. Through a large-scale measurement study, we reveal a pronounced security-utility asymmetry: across a wide range of lossy speculative decoding methods, improvements in inference efficiency come at a disproportionately high cost to security, with attack success rates for jailbreak and prompt injection attacks increasing much faster than utility degrades. We then propose SecureSD, a new theory-guided speculative decoding method that enhances security while maintaining efficiency and utility. Specifically, our theoretical analysis reveals that security degradation primarily originates from the early tokens generated by the draft model. Motivated by this insight, SecureSD applies a stricter verification criterion to draft-model tokens at early decoding positions. Extensive experiments on both security and utility benchmarks demonstrate that SecureSD significantly improves security while preserving efficiency and utility compared to existing speculative decoding methods.
Chinese Translation
推测解码通过首先使用一个较小的模型(称为\emph{草稿模型})生成候选 token,然后由目标模型验证这些候选 token 以接受或拒绝,从而加速被称为\emph{目标模型}的大型语言模型(LLM)的推理。先前的研究主要关注推测解码的效率-效用权衡,例如有损推测解码,而对其安全影响在很大程度上仍未探索。在这项工作中,我们通过提供对推测解码安全影响的\emph{首个}系统性研究来弥合这一空白。通过大规模测量研究,我们揭示了一个显著的安全-效用不对称性:在广泛的各类有损推测解码方法中,推理效率的提升以不成比例的高安全代价换得,其中越狱和提示注入攻击的攻击成功率增长远快于效用下降。然后,我们提出了 SecureSD,一种新的理论引导的推测解码方法,在保持效率和效用的同时增强安全性。具体而言,我们的理论分析表明,安全退化主要源于草稿模型生成的早期 token。受这一洞见启发,SecureSD 对早期解码位置处的草稿模型 token 应用更严格的验证标准。在安全性和效用基准上的大量实验表明,与现有的推测解码方法相比,SecureSD 在保持效率和效用的同时显著提高了安全性。
cs.CR / 72 / 2610.08739
BARE-AI: Bit-Flip Attack Resilience in AI Hardware through Built-in Performance Monitors
BARE-AI:通过内置性能监视器实现 AI 硬件中的比特翻转攻击韧性
large language model
大语言模型相关
Abstract
Deep Neural Networks (DNNs) are integral to many safety critical systems, yet they remain highly vulnerable to bit-flip attacks (BFAs), where a few memory level perturbations can drastically degrade accuracy. Existing defenses incur significant hardware overhead, depend on retraining, or fail against targeted flips. We propose BARE-AI, a runtime framework that detects, localizes, and mitigates BFAs during inference. BARE-AI introduces AI Performance Counters (APCs), lightweight hardware monitors in the accelerator datapath that capture per-layer activation statistics such as sparsity, entropy, kurtosis, and spectral shift. These are analyzed by the Predictive Unit for Layer Security Evaluation (PULSE), a compact detector trained offline as an ensemble of classifiers and realized on-chip as a small neural engine. For explainability and recovery, BARE-AI introduces an Activation Shift Index (ASI) for layer level fault localization and a z-score based repair that resets anomalous weights toward clean layer statistics. Across CNNs, Vision Transformers, and Large Language Models under random, targeted, adaptive, and magnitude based BFAs, BARE-AI achieves up to 98% detection accuracy on vision models and 74% to 95% on language models, restores near clean accuracy for CNNs and ViTs, and provides partial recovery for LLMs. Synthesized at 28nm, the monitoring infrastructure incurs under 3% energy, under 4% area, and about 10% latency overhead, with a configurable operating point that reduces latency overhead to about 6%. Unlike error correcting codes, whose redundancy grows with the number of tolerated flips, BARE-AI's overhead remains constant regardless of attack strength, making it attractive for resource constrained, safety critical edge applications such as autonomous systems, energy, and healthcare.
Chinese Translation
深度神经网络(DNN)是许多安全关键系统不可或缺的组成部分,但它们仍然极易受到比特翻转攻击(BFA)的影响,在这种攻击中,少数内存级扰动就能极大地降低准确率。现有防御会带来显著的硬件开销、依赖重新训练,或无法抵御定向翻转。我们提出 BARE-AI,一个运行时框架,可在推理期间检测、定位并缓解 BFA。BARE-AI 引入了 AI 性能计数器(APC),这是加速器数据通路中的轻量级硬件监视器,可捕获逐层激活统计量,如稀疏度、熵、峰度和频谱偏移。这些统计量由用于层安全评估的预测单元(PULSE)分析,PULSE 是一个紧凑型检测器,离线训练为分类器集成,并在片上实现为一个小型神经引擎。为了实现可解释性和恢复,BARE-AI 引入了激活偏移指数(ASI)用于层级故障定位,以及基于 z 分数的修复方法,将异常权重重置为趋向干净层统计量。在随机、定向、自适应和基于幅度的 BFA 下,跨越 CNN、视觉 Transformer 和大语言模型,BARE-AI 在视觉模型上实现了高达 98% 的检测准确率,在语言模型上达到 74% 至 95%,为 CNN 和 ViT 恢复了接近干净的准确率,并为 LLM 提供了部分恢复。以 28nm 工艺综合,该监控基础设施产生低于 3% 的能量开销、低于 4% 的面积开销和约 10% 的延迟开销,并具有可配置工作点,可将延迟开销降至约 6%。与纠错码不同,纠错码的冗余会随着可容忍翻转的数量增加而增长,而 BARE-AI 的开销无论攻击强度如何都保持恒定,这使其对资源受限、安全关键的边缘应用具有吸引力,例如自主系统、能源和医疗保健。
cs.LG / 73 / 2610.07269
What Words Keep of a Place: Zero-Shot Language Reasoning for Cross-View Geo-Localization
词语为一个地点保留了什么:面向跨视角地理定位的零样本语言推理
large language model
大语言模型相关
Abstract
Cross-view geo-localization is commonly solved as an image retrieval problem, matching a ground-level image against a database of satellite tiles through a jointly trained embedding. Such models are accurate, but they need large paired supervision and cannot show what evidence supports a match. In this paper, we study a different question: how much of this task can be solved through language alone? We prompt a multimodal large language model (MLLM) to describe each ground panorama and each satellite tile as structured text, and localize by comparing these descriptions. No component is trained. We evaluate on 9,826 VIGOR pairs from four U.S. cities, in three settings. First, the descriptions are faithful but not discriminative. They agree closely across the two views, yet ranking the full pool by description similarity almost never returns the correct tile (0.39% Recall@1). Second, we narrow the pool to ten neighboring tiles, as a coarse prior would do. The same descriptions now become useful: an MLLM judge that scores structural consistency doubles random ranking and matches a strong lexical baseline. It also states which fields of the two descriptions agree and which conflict, which an embedding distance cannot do, and which we see as a step toward interpretable localization. Third, we place the judge on a trained visual retriever. On the queries it ranks wrongly, reranking from images works, while reranking from our descriptions does not (23.5% against 10.7% Recall@1). Scene structure survives the conversion into language, while the fine appearance detail needed to separate nearby places does not. Code and prompts are publicly available at https://github.com/AyeshAbuLehyeh/GeoLingual.
Chinese Translation
跨视角地理定位通常被作为图像检索问题来解决,即通过一个联合训练的嵌入,将地面图像与卫星瓦片数据库进行匹配。此类模型很准确,但它们需要大量成对监督,并且无法展示是什么证据支持某一匹配。在本文中,我们研究一个不同的问题:这项任务有多少可以仅通过语言来解决?我们提示一个多模态大语言模型(MLLM)将每个地面全景图和每个卫星瓦片描述为结构化文本,并通过比较这些描述来进行定位。没有任何组件经过训练。我们在来自四个美国城市的 9,826 个 VIGOR 对上,在三种设置下进行评估。首先,这些描述是忠实的,但不具有区分性。它们在两个视角之间高度一致,然而按描述相似度对整个候选池进行排序几乎从不会返回正确的瓦片(0.39% Recall@1)。其次,我们将候选池缩小到十个相邻瓦片,就像粗粒度先验会做的那样。相同的描述现在变得有用:一个对结构一致性进行评分的 MLLM 评判器将随机排序的效果提高了一倍,并与一个强词汇基线相当。它还会指出两个描述中哪些字段一致、哪些冲突,这是嵌入距离无法做到的,我们将其视为迈向可解释定位的一步。第三,我们将该评判器置于一个训练过的视觉检索器之上。在它排序错误的查询上,从图像进行重排序有效,而从我们的描述进行重排序则无效(23.5% 对 10.7% Recall@1)。场景结构在转换成语言后得以保留,而区分邻近地点所需的精细外观细节则没有。代码和提示词已在 https://github.com/AyeshAbuLehyeh/GeoLingual 公开可用。
cs.AI / 74 / 2610.07684
Disentangling Dual Image References in Frequency Aware Diffusion Models for Personalized Generation
在频率感知扩散模型中解耦双重图像参考以实现个性化生成
diffusion
扩散模型相关
Abstract
Personalized image generation aims to synthesize text-driven images conditioned on reference images, while mainly casting the generation as image customization for foreground and style transfer for background. Previous arts of diffusion models suffers from the text misalignment with background for image customization and foreground for style transfer during the denoising process. Such facts, as we observed, rooted from the entanglement among hybrid frequency bands during the denoising process. To address such salient limitation, in this paper, we study personalized generation based on dual references - customization and color and style reference - and propose a paradigm to disentangle these Dual image references within Frequency-aware Diffusion Models, dubbed Dual-FDM, to simultaneously tackle two crucial personalized image generation tasks: customization style transfer and color style transfer, by disentangling different frequency bands via mask strategy within frequency domain. For customization style transfer, we replace the mid-frequency band of the background in the style reference with that from the foreground of the customized reference. For color style transfer, we substitute the low-frequency band of the background in the style reference with that from both the foreground and background of the color reference. Both the substituted frequency bands are used as the key and value to reconstruct the query foreground and background of the denoised personalized image.Extensive experiments validate the superiority of Dual-FDM over the state-of-the-art diffusion models for personalized image generation. Our code can be accessed from https://github.com/htyjers/Dual-FDM.
Chinese Translation
个性化图像生成旨在以参考图像为条件合成文本驱动的图像,同时主要将该生成任务视为针对前景的图像定制和针对背景的风格迁移。先前扩散模型的方法在去噪过程中存在图像定制时文本与背景不对齐以及风格迁移时文本与前景不对齐的问题。正如我们所观察到的,这些事实源于去噪过程中混合频段之间的纠缠。为了解决这一显著限制,在本文中,我们研究基于双重参考——定制参考以及颜色和风格参考——的个性化生成,并提出一种在频率感知扩散模型中解耦这些双重图像参考的范式,称为 Dual-FDM,以通过在频域内借助掩码策略解耦不同频段,同时处理两个关键个性化图像生成任务:定制风格迁移和颜色风格迁移。对于定制风格迁移,我们用定制参考的前景中的中频段替换风格参考中背景的中频段。对于颜色风格迁移,我们用颜色参考的前景和背景中的低频段替换风格参考中背景的低频段。两个被替换的频段都被用作键和值,以重建去噪后的个性化图像的查询前景和背景。大量实验验证了 Dual-FDM 在个性化图像生成方面优于最先进的扩散模型。我们的代码可从 https://github.com/htyjers/Dual-FDM 获取。
cs.LG / 75 / 2610.07720
RefRoute: Decoupling Conditioning Cost from References via Compact Residual Conditioning and Spatial Routing
RefRoute:通过紧凑残差条件化与空间路由将条件化成本与参考解耦
diffusion
扩散模型相关
Abstract
Multi-reference image generation requires preserving the appearance of multiple subjects while composing them into a coherent scene. However, existing diffusion transformers commonly encode references as dense visual token grids and jointly process them with global attention, making conditioning increasingly expensive as the number and resolution of references grow. We present RefRoute, a framework that addresses both reference representation cost and attention overhead through two complementary mechanisms. Compact residual conditioning combines low-resolution latent tokens with lightweight residual features extracted from full-resolution pixels, reducing reference token counts while retaining fine-grained appearance cues. Condition routing and attention routing align reference tokens with their assigned target regions and restrict cross-reference interactions, while allowing selective reference access beyond region boundaries for scene integration. We further introduce RefRoute-Data for training many-reference generation models and ManyRef100, a benchmark spanning human, object, and mixed compositions with 10-17 references. After many-reference fine-tuning, RefRoute achieves an overall Weighted-Ref-VIEScore of 36.06 on ManyRef100, compared with 8.88 for FLUX.2-Klein-9B. Separate inference-cost evaluations show substantially slower latency growth as the reference count increases: at 16 references, our 50-step and 4-step configurations achieve $18.3\times$ and $14.2\times$ speedups over their corresponding FLUX baselines, respectively. These results establish compact reference representations and spatially routed attention as an effective approach to scalable many-reference image generation.
Chinese Translation
多参考图像生成需要在将多个主体组合成连贯场景的同时,保持它们的外观。然而,现有的扩散Transformer通常将参考编码为密集的视觉token网格,并用全局注意力联合处理它们,这使得随着参考数量和分辨率的增长,条件化变得越来越昂贵。我们提出了RefRoute,一个通过两种互补机制同时解决参考表示成本和注意力开销的框架。紧凑残差条件化将低分辨率潜在token与从全分辨率像素中提取的轻量级残差特征相结合,在减少参考token数量的同时保留细粒度外观线索。条件路由和注意力路由将参考token与其分配的目标区域对齐,并限制跨参考交互,同时允许在区域边界之外选择性访问参考以进行场景整合。我们进一步引入了RefRoute-Data,用于训练多参考生成模型,以及ManyRef100,一个涵盖人物、物体和混合组合且包含10-17个参考的基准测试。经过多参考微调后,RefRoute在ManyRef100上取得了36.06的总体Weighted-Ref-VIEScore,而FLUX.2-Klein-9B为8.88。单独的推理成本评估显示,随着参考数量增加,延迟增长显著更慢:在16个参考时,我们的50步和4步配置分别相比其对应的FLUX基线实现了$18.3\times$和$14.2\times$的加速。这些结果表明,紧凑参考表示和空间路由注意力是通向可扩展多参考图像生成的有效方法。
cs.AI / 76 / 2610.07913
Multimodal Knowledge Distillation for Gastric Adenocarcinoma Classification from Whole-Slide Images
基于全切片图像的胃腺癌分类的多模态知识蒸馏
large language model
大语言模型相关
Abstract
Gastric adenocarcinoma (GA) is a leading cause of cancer-related mortality worldwide, and accurate histopathological subtype classification from whole-slide images (WSIs) is essential for effective treatment planning. While multimodal approaches that integrate pathology report text with WSIs can improve classification, existing methods often depend on computationally expensive transformer architectures and large language models. We propose a multimodal knowledge distillation (MKD) framework that combines a pretrained WSI image encoder and a clinical text encoder using Low-Rank Multimodal Fusion (LMF) to efficiently model cross-modal interactions during training. Each WSI is represented as a bag of patches paired with a slide-level diagnostic caption. The teacher model learns fused image-text representations for subtype classification, while the student model distills this knowledge to enable accurate image-only inference. We evaluate our method on the PatchGastric benchmark dataset and achieve at least 3.35% higher mean accuracy than state-of-the-art approaches, without relying on transformer-based fusion, multi-task learning, or large language models. The source code is available at https://github.com/helomelo1/MKD-LMF.
Chinese Translation
胃腺癌(GA)是全球癌症相关死亡的主要原因之一,而基于全切片图像(WSIs)进行准确的组织病理学亚型分类对于有效的治疗规划至关重要。尽管将病理报告文本与WSIs整合的多模态方法可以提高分类性能,但现有方法往往依赖于计算成本高昂的Transformer架构和大语言模型。我们提出了一种多模态知识蒸馏(MKD)框架,该框架结合预训练的WSI图像编码器和临床文本编码器,并使用低秩多模态融合(LMF)来在训练期间高效地建模跨模态交互。每个WSI被表示为一个图像块包,并与一个切片级诊断描述配对。教师模型学习用于亚型分类的融合图像-文本表示,而学生模型蒸馏这些知识,以实现仅图像的准确推理。我们在PatchGastric基准数据集上评估我们的方法,并比最先进的方法至少获得3.35%更高的平均准确率,且不依赖于基于Transformer的融合、多任务学习或大语言模型。源代码可在 https://github.com/helomelo1/MKD-LMF 获取。
cs.AI / 77 / 2610.07987
VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs
VisionWeave:将弹性视觉表示编织为 MLLMs 的原生能力
large language model
大语言模型相关
Abstract
Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Downsampling sacrifices this detail, while existing token pruning and adaptive approaches remain limited in content-adaptive granularity, task generalization, and integration with modern MLLMs and serving infrastructure. Overcoming these limitations calls for foundation models that learn, end to end, where-and at what granularity-to allocate visual representations, a native capability we term elastic visual representation weaving. We introduce VisionWeave, establishing this capability in frontier-level MLLMs through large-scale training. It combines two components: a gated spatial pooler constructs coarse-grained representations alongside native fine-grained representations within a shared MRoPE coordinate, while a granularity router learns their content-adaptive allocation. Through self-distillation alone, we validate this capability on Qwen3.5-4B and scale to Qwen3.8-27B with over 30K A100 GPU-hours. Based on Qwen3.8-27B, VisionWeave adaptively adjusts token savings to visual content, saving 43.0% tokens on average while retaining 98.9% native performance across eight benchmarks, versus only 88% performance preserved for token pruning baselines with a fixed 50% savings target. Extensive evaluations confirm robust efficiency-quality trade-offs across diverse tasks, resolutions and video frames. When deployed on SGLang serving engine, our method achieves a 2.3x throughput gain while reducing mean TTFT by 54.4% and mean TPOT by 60.6%. Together, we believe these results position elastic visual weaving as a promising capability for next-generation multimodal models.
Chinese Translation
多模态大语言模型(MLLMs)已成为视觉理解的主导范式,但通过将输入编码为密集、固定大小的图像块 token 而带来高昂成本。然而,视觉信息分布不均:一些区域需要细粒度细节,而另一些区域则允许紧凑表示。下采样牺牲了这些细节,而现有的 token 剪枝与自适应方法在内容自适应粒度、任务泛化以及与当代 MLLMs 和服务基础设施的集成方面仍然受限。克服这些局限需要基础模型能够端到端地学习在何处——以及以何种粒度——分配视觉表示,我们将这种原生能力称为弹性视觉表示编织。我们提出 VisionWeave,通过大规模训练在前沿级 MLLMs 中建立这一能力。它结合了两个组件:一个门控空间池化器在共享的 MRoPE 坐标内构建粗粒度表示,并与原生细粒度表示并存,而一个粒度路由器学习它们的内容自适应分配。仅通过自蒸馏,我们在 Qwen3.5-4B 上验证了这一能力,并以超过 30K A100 GPU 小时扩展到 Qwen3.8-27B。基于 Qwen3.8-27B,VisionWeave 根据视觉内容自适应地调整 token 节省量,在八个基准上平均节省 43.0% token,同时保留 98.9% 的原生性能;相比之下,在固定 50% 节省目标下,token 剪枝基线仅保留 88% 的性能。大量评估证实了在多样任务、分辨率和视频帧上稳健的效率—质量权衡。当部署在 SGLang 服务引擎上时,我们的方法实现了 2.3 倍的吞吐量提升,同时将平均 TTFT 降低 54.4%,将平均 TPOT 降低 60.6%。综合来看,我们相信这些结果将弹性视觉编织定位为下一代多模态模型的一项有前景的能力。
cs.LG / 78 / 2610.08341
DIPrune: Task-Aware Token Pruning with Dual Importance for Efficient Multimodal Language Models
DIPrune:具有双重重要性的任务感知 Token 剪枝,用于高效多模态语言模型
large language model
大语言模型相关
Abstract
Recent training-free pruning approaches for Multimodal Large Language Models (MLLMs) effectively cut computational overhead by exploiting visual redundancy or text-vision attention. However, they frequently suffer from semantic degradation due to their task-agnostic design or unreliable attention estimates. Based on our empirical analysis, we have found that this issue arises because salient tokens in shallow layers persistently suppress emerging semantic ones through numerical inertia, leading to premature discarding of signals crucial for deep reasoning. To address the aforementioned issue, from the task-oriented aspects, we first reformulate training-free pruning as a minimization of the distortion in the final task loss and derive a tractable, token-wise upper bound to serve as a surrogate objective. Specifically, this formulation inherently reveals a previously neglected inter-layer term that accounts for gradients across layers. Accordingly, for the implementation, we propose DIPrune, a rank-based framework that employs a dual importance scoring mechanism to jointly optimize intra-layer static feature saliency and inter-layer dynamic semantic evolution. Extensive experiments on LLaVA and Qwen-VL demonstrate that DIPrune consistently achieves state-of-the-art results.
Chinese Translation
最近的针对多模态大语言模型(MLLMs)的免训练剪枝方法通过利用视觉冗余或文本-视觉注意力,有效削减了计算开销。然而,由于它们采用任务无关的设计或不可靠的注意力估计,这些方法经常遭受语义退化。基于我们的实证分析,我们发现该问题之所以出现,是因为浅层中的显著 token 会通过数值惯性持续压制新出现的语义 token,从而导致对深层推理至关重要的信号被过早丢弃。为解决上述问题,从面向任务的角度出发,我们首先将免训练剪枝重新表述为最小化最终任务损失中的失真,并推导出一个可处理的、逐 token 的上界,以作为替代目标。具体而言,这一表述本质上揭示了一个此前被忽视的层间项,该项考虑了跨层梯度。相应地,在实现层面,我们提出了 DIPrune,一个基于秩的框架,它采用双重重要性评分机制来联合优化层内静态特征显著性和层间动态语义演化。在 LLaVA 和 Qwen-VL 上进行的大量实验表明,DIPrune 持续取得最先进的结果。
cs.AI / 79 / 2610.08268
DySCo: Dynamic Sharding for Collaborative Edge-Cloud LLM Inference with Depth-Synchronized Batching
DySCo:面向采用深度同步批处理的协作式边云 LLM 推理的动态分片
large language model
大语言模型相关
Abstract
Pervasive intelligent applications are increasingly deployed on mobile and Internet of Things (IoT) edge devices. Consequently, Large Language Models (LLMs) are increasingly used to support these applications. Yet, due to their high resource demands, LLMs are mostly deployed in the cloud. Layer-wise edge-cloud inference lets resource-constrained edge devices contribute computation to LLMs they cannot host in full. However, heterogeneous split points introduce two coupled inefficiencies. First, edge execution and communication create idle gaps between cloud invocations. Second, requests arriving at different model depths cannot be conventionally batched. We present DySCo, a collaborative runtime that keeps KV caches local and introduces dyForward, a model-aware layer-range executor that runs configurable contiguous layer ranges from resident model shards without reloading weights. For multi-edge serving settings, we introduce depth-synchronized batching (DSB), which advances heterogeneous requests to the deepest cut and batches their common suffix. Experiments across heterogeneous devices, two model families, and local and wide-area links show that idle gaps increase the latency of subsequent GPU forward calls even when waiting time is excluded, adding up to 25 ms of additional cloud-side suffix latency per decoding step in our measurements. At an average concurrency of eight, DSB improves throughput by 275% over FIFO, 48% over exact-match batching, and 79% over round-robin interleaving while reducing mean per-session latency. Together, these results show that requests with different edge-cloud splits can reuse resident cloud weights and share batched suffix computation. The artifact repository for this work is publicly available at: https://github.com/Large-scale-Sustainable-Computing-LSC/dysco-artifact
Chinese Translation
普适智能应用正越来越多地部署在移动和物联网(IoT)边缘设备上。因此,大语言模型(LLMs)正越来越多地被用于支持这些应用。然而,由于其高资源需求,LLM 大多部署在云端。逐层边云推理使资源受限的边缘设备能够为其无法完整托管的 LLM 贡献计算。然而,异构切分点引入了两种相互耦合的低效问题。首先,边缘执行与通信会在云调用之间产生空闲间隙。其次,到达不同模型深度的请求无法以常规方式进行批处理。我们提出 DySCo,一种协作式运行时,它将 KV 缓存保持在本地,并引入 dyForward,一种模型感知的层范围执行器,它从驻留的模型分片中运行可配置的连续层范围,而无需重新加载权重。针对多边缘服务场景,我们引入深度同步批处理(DSB),它将异构请求推进到最深的切分点,并对其共同后缀进行批处理。跨异构设备、两个模型系列以及本地和广域网链路的实验表明,即使排除等待时间,空闲间隙也会增加后续 GPU 前向调用的延迟,在我们的测量中,每个解码步骤最多增加 25 ms 的额外云端后缀延迟。在平均并发度为八时,DSB 相比 FIFO 将吞吐量提高 275%,相比精确匹配批处理提高 48%,相比轮询交错提高 79%,同时降低平均每会话延迟。这些结果共同表明,具有不同边云切分的请求可以复用驻留的云端权重,并共享批处理后的后缀计算。本工作的工件仓库公开可获取于:https://github.com/Large-scale-Sustainable-Computing-LSC/dysco-artifact
cs.CL / 80 / 2610.08716
Disentangling Paradigm, Identifier, and Decoding in Generative Retrieval
解耦生成式检索中的范式、标识符与解码
diffusion
扩散模型相关
Abstract
Generative retrieval trains a language model to generate the identifier of a relevant document. Recent work replaces the autoregressive decoder with diffusion, but changes identifiers, training recipe and decoding at once, so differences cannot be credited to the paradigm. On NQ320K and MS300K, we train autoregressive, masked-diffusion and block-diffusion models with residual-quantised, product-quantised and random identifiers. With identifier length and training budget fixed, we decode each model in several ways. Decoding alone moves a diffusion model's Hit@1 by 6.6 to 13.7 points. Our reference diffusion decoding, generate-and-match, generates an identifier, then retrieves the closest corpus identifiers. The generated identifier is right for 14-21% of NQ320K queries. We test one-pass scoring to decode diffusion retrievers: the model reads a fully masked identifier once, and each document is scored by its codes' probabilities. It matches or beats generate-and-match in 11 of 12 settings. Autoregressive models still lead in Hit@1; on NQ320K, the lead comes from the model, not beam search. Starting from one sampled identifier, one-pass scoring removes 46-83% of masked diffusion's deficit to beam search; from generate-and-match, at most a quarter. On NQ320K, every paradigm largely memorises which identifier answers which query: random identifiers keep 83-90% of the Hit@1 of residual-quantised ones. There, product-quantised identifiers lead residual-quantised ones by 3.4 points in the autoregressive model and by -0.7 to +3.6 in diffusion models; across decodings, AR's gap exceeds diffusion's by 1.5-2.3 points, around our 2-point threshold. Paradigm comparisons must report each paradigm at its own recipe and best decoding.
Chinese Translation
生成式检索训练一个语言模型来生成相关文档的标识符。近期工作用扩散替换自回归解码器,但同时改变了标识符、训练配方和解码方式,因此差异不能归因于范式。在 NQ320K 和 MS300K 上,我们训练了使用残差量化、乘积量化和随机标识符的自回归、掩码扩散和块扩散模型。在标识符长度和训练预算固定不变的情况下,我们以多种方式对每个模型进行解码。仅解码本身就能使扩散模型的 Hit@1 变动 6.6 到 13.7 个点。我们的参考扩散解码方法 generate-and-match 生成一个标识符,然后检索最接近的语料库标识符。生成的标识符对于 14-21% 的 NQ320K 查询是正确的。我们测试单遍评分来解码扩散检索器:模型读取一个完全被掩码的标识符一次,并且每个文档由其代码的概率进行评分。它在 12 种设置中的 11 种里与 generate-and-match 持平或优于它。自回归模型在 Hit@1 上仍然领先;在 NQ320K 上,这种领先来自模型,而不是束搜索。从一个采样得到的标识符出发,单遍评分消除了掩码扩散相对于束搜索的 46-83% 的差距;而从 generate-and-match 出发,最多消除四分之一。在 NQ320K 上,每种范式在很大程度上都记住了哪个标识符回答哪个查询:随机标识符保留了残差量化标识符 83-90% 的 Hit@1。在那里,乘积量化标识符在自回归模型中领先残差量化标识符 3.4 个点,在扩散模型中领先 -0.7 到 +3.6 个点;跨解码方式来看,AR 的差距比扩散的差距高出 1.5-2.3 个点,大约在我们 2 个点的阈值附近。范式比较必须报告每个范式在其自身配方和最佳解码下的结果。
cs.LG / 81 / 2610.07162
Adversarial Training for Deep Hedging in Nonstationary Markets
非平稳市场中深度对冲的对抗训练
diffusion
扩散模型相关
Abstract
Deep hedging learns trading policies from historical or simulated market trajectories, yet under nonstationarity these training paths may not represent future market conditions. We propose WRAP (Wasserstein-Reweighting Adversarial Perturbation), a drift-aware adversarial training framework derived from a two-budget distributionally robust optimization (DRO) formulation. The formulation is anchored to a weighted empirical reference distribution whose fixed baseline weights are chosen to balance sampling uncertainty against temporal drift. Around this reference distribution, the ambiguity set addresses two complementary forms of distributional misspecification by allowing an adversary to reweight the observed trajectories subject to a $φ$-divergence constraint and perturb their paths subject to an optimal-transport (OT) constraint. We derive a joint first-order expansion in which the leading-order increase over the nominal expected loss decomposes into a reweighting contribution determined by the dispersion of hedging losses across trajectories and a transport contribution determined by the sensitivity of the loss to path perturbations. This expansion yields an explicit finite-dimensional adversarial attack that replaces the distributional inner supremum with a tractable first-order approximation. Across stationary and nonstationary Heston dynamics and a generalized affine diffusion (GAD), the experiments show complementary benefits from reweighting and transport, with joint adversarial training providing the largest gains under nonstationarity.
Chinese Translation
深度对冲从历史或模拟的市场轨迹中学习交易策略,然而在非平稳性下,这些训练路径可能无法代表未来的市场状况。我们提出 WRAP(Wasserstein-Reweighting Adversarial Perturbation,Wasserstein 重加权对抗扰动),一个由双预算分布鲁棒优化(DRO)形式推导得到的漂移感知对抗训练框架。该形式以加权经验参考分布为锚点,其固定基线权重被选择用于平衡采样不确定性与时间漂移。围绕该参考分布,模糊集通过允许对抗者在对观测轨迹进行重加权时受 $φ$-散度约束,并对其路径进行扰动时受最优传输(OT)约束,来处理两种互补形式的分布误设。我们推导出一个联合一阶展开,其中相对于名义期望损失的首阶增量分解为:由跨轨迹对冲损失离散程度决定的重加权贡献,以及由损失对路径扰动的敏感度决定的传输贡献。该展开产生一个显式的有限维对抗攻击,其用可处理的一阶近似取代分布内层上确界。在平稳和非平稳 Heston 动力学以及广义仿射扩散(GAD)上,实验表明重加权和传输具有互补收益,而联合对抗训练在非平稳性下提供最大增益。
cs.LG / 82 / 2610.07208
Can LLM-assisted regularization increase forecast accuracy for migration flows in low data regimes?
LLM辅助正则化能否提高低数据机制下迁移流的预测精度?
large language model
大语言模型相关
Abstract
Predicting migration flows remains a significant challenge for traditional gravity-based forecasting models, which primarily rely on structured socio-economic indicators such as economic disparity, political stability, and geographic distance. This work investigates whether Large Language Models (LLMs) can improve migration forecasting by extracting contextual migration-related signals from news articles and incorporating them into a weighted Lasso forecasting framework through feature-specific regularization penalties. The proposed framework uses hierarchical LLM inference pipelines to classify migration-related push--pull signals from news data and evaluates the resulting forecasting performance across multiple migration corridors between November 2021 and November 2022, including Mexico--United States, Ukraine--Poland, and Syria--Turkey. Experimental results showed mixed performance across migration corridors and modeling strategies, and no single regularization approach consistently outperformed the others across all experiments. The best-performing Mexico configuration, which consisted of a gravity-based model augmented with the proposed push--pull ratios, achieved a Mean Absolute Percentage Error (MAPE) of 17.15%, while the strongest Syria configuration achieved a MAPE of 29.29% using Direct LLM-Lasso. For Ukraine, the best-performing configuration used LLM-Assisted Regularization (AR) and achieved a MAPE of 41.05%. Overall, the results suggest that contextual article-derived features and LLM-guided regularization can improve migration forecasting under certain conditions, although migration corridor characteristics, article volume, and hyperparameter configuration strongly influenced performance.
Chinese Translation
预测迁移流仍然是传统基于重力的预测模型面临的重大挑战,这些模型主要依赖结构化的社会经济指标,如经济差距、政治稳定性和地理距离。本研究考察大型语言模型(LLM)是否能够通过从新闻文章中提取与迁移相关的上下文信号,并通过特征特定的正则化惩罚将这些信号纳入加权 Lasso 预测框架,从而改进迁移预测。所提出的框架使用分层 LLM 推理流水线来对来自新闻数据的与迁移相关的推--拉信号进行分类,并评估由此产生的预测性能,涵盖 2021 年 11 月至 2022 年 11 月之间的多个迁移走廊,包括墨西哥--美国、乌克兰--波兰和叙利亚--土耳其。实验结果显示,不同迁移走廊和建模策略下的表现参差不齐,并且没有任何单一正则化方法在所有实验中始终优于其他方法。表现最佳的墨西哥配置由基于重力的模型加上所提出的推--拉比率构成,达到了 17.15% 的平均绝对百分比误差(MAPE),而表现最强的叙利亚配置使用 Direct LLM-Lasso,达到了 29.29% 的 MAPE。对于乌克兰,表现最佳的配置使用了 LLM 辅助正则化(AR),并达到了 41.05% 的 MAPE。总体而言,结果表明,上下文文章衍生特征和 LLM 引导的正则化在某些条件下能够改进迁移预测,尽管迁移走廊特征、文章数量和超参数配置强烈影响了性能。
cs.LG / 83 / 2610.07226
Minimal Witness Reinforcement Learning
极小见证强化学习
large language model
大语言模型相关
Abstract
``What are the irreducible conditions that are sufficient to produce an outcome?'' is one of the most common questions that recur across computation and science. Its answers, the minimal sufficient witnesses, are what we mean by explanations, mechanisms and reasons. These problems usually ask for multiple minimal witnesses, yet standard RL methods may reveal only one solution or redundant ones. We formalize this problem as minimal-witness identification and introduce Minimal-Witness Reinforcement Learning (MWRL). MWRL takes the union of the sets certified by successful proposals sampled from the policy and credits each proposal for the coverage the group union would lose without that proposal. This credit assignment, derived directly from the problem definition, unifies the demands for minimality and recovery of alternatives from a single black-box verifier bit. Under this principle, we derive a value iteration planner that recovers the entire family of witnesses and a policy gradient method that can scale to large language models. Across different experimental settings, MWRL recovers most minimal witnesses, while other methods return redundant supersets or a single witness. By making witness families learnable from verifier feedback, MWRL expands the scope of reinforcement learning beyond single-solution optimization. Our code is available at https://github.com/TSUITUENYUE/MWRL.
Chinese Translation
「哪些不可约的条件足以产生一个结果?」是横跨计算与科学领域反复出现的最常见问题之一。它的答案,即极小充分见证,正是我们所说的解释、机制与理由。这类问题通常要求给出多个极小见证,然而标准的强化学习方法可能只揭示出一个解,或是揭示出冗余的解。我们将该问题形式化为极小见证识别,并提出极小见证强化学习(MWRL)。MWRL 取从策略中采样得到的成功提议所认证的集合的并集,并将群体并集在缺少某个提议时将会损失的覆盖度记到该提议的账上。这种直接由问题定义导出的信用分配,从单个黑盒验证器比特中统一了对极小性与备选方案恢复这两方面的要求。在这一原则下,我们推导出一个能够恢复整个见证族的值迭代规划器,以及一个能够扩展到大型语言模型的策略梯度方法。在不同的实验设置下,MWRL 恢复出了大多数极小见证,而其他方法则返回冗余的超集或单个见证。通过使见证族能够从验证器反馈中学习,MWRL 将强化学习的范围拓展到了单解优化之外。我们的代码可在 https://github.com/TSUITUENYUE/MWRL 获取。
cs.LG / 84 / 2610.07247
Learning What to Distill: Bilevel Top-K Token Selection for Self-Distillation in Large Language Models
学习蒸馏什么:用于大语言模型自蒸馏的双层Top-K Token选择
large language model
大语言模型相关
Abstract
Large language models have shown strong reasoning capabilities, but their high inference costs make knowledge distillation an important approach for transferring such capabilities to compact models in resource-constrained scenarios. On-policy self-distillation further reduces the reliance on external large teacher models while improving the reasoning ability of compact language models. However, existing methods typically either distill all token positions uniformly or select tokens using fixed heuristic criteria, assigning the same distillation strength to the selected positions rather than adaptively learning which tokens are most beneficial for distillation. To address these limitations, we propose BiToK-SD (Bilevel Top-K Token Selection for Self-Distillation), a bilevel-optimization-based token selection method that learns where distillation should be applied during on-policy self-distillation. Specifically, BiToK-SD is formulated as a bilevel optimization problem, where the lower-level problem models Top-K token selection as a differentiable threshold-based relaxation, allowing the selected positions to adapt as the student policy evolves, while the upper-level problem performs knowledge distillation on the selected positions. Experiments on mathematical reasoning benchmarks show that BiToK-SD achieves the best average performance among all compared methods while requiring only lightweight additional computation.
Chinese Translation
大语言模型已展现出强大的推理能力,但其高昂的推理成本使得知识蒸馏成为在资源受限场景下将此类能力迁移至紧凑模型的重要方法。在线策略自蒸馏进一步减少了对大型外部教师模型的依赖,同时提升了紧凑语言模型的推理能力。然而,现有方法通常要么对所有token位置均匀地进行蒸馏,要么使用固定的启发式标准选择token,对所选位置赋予相同的蒸馏强度,而不是自适应地学习哪些token最有利于蒸馏。为解决这些局限,我们提出BiToK-SD(Bilevel Top-K Token Selection for Self-Distillation),一种基于双层优化的token选择方法,它能够在在线策略自蒸馏过程中学习应在何处施加蒸馏。具体而言,BiToK-SD被形式化为一个双层优化问题,其中下层问题将Top-K token选择建模为基于阈值的可微松弛,使得所选位置能够随着学生策略的演化而自适应调整,而上层问题则在所选位置上执行知识蒸馏。在数学推理基准上的实验表明,BiToK-SD在所有比较方法中取得了最佳的平均性能,同时仅需轻量级的额外计算。
cs.LG / 85 / 2610.07335
Selective Critique for Cost-Aware LLM Agents in Long-Horizon Decision Making
长时程决策中面向成本感知 LLM 智能体的选择性评审
large language model
大语言模型相关
Abstract
Improving the reliability of large language model (LLM) agents in long-horizon decision-making remains a key challenge. When deployed as autonomous agents interacting with complex environments, early mistakes can propagate through trajectories and cause cascading failures. Recent approaches improve reliability by incorporating external critique or deliberation, but invoking these mechanisms at every step substantially increases token consumption and latency, limiting practical deployment. We propose SAG (Self-improving Agent with Gated critique), a cost-aware framework that formulates critique invocation as a step-wise decision problem during long-horizon interaction. SAG introduces a lightweight, training-free gating mechanism that estimates the utility of critique using action-level ambiguity signals--global entropy and local top-2 margin--computed over admissible actions. From a decision-theoretic perspective, this mechanism approximates the Value of Information (VoI) of critique, enabling the agent to selectively allocate expensive feedback only when its expected benefit justifies the cost. SAG further incorporates online bootstrapped self-improvement, allowing the actor to internalize critic-assisted behaviors and progressively reduce reliance on critique. Across three long-horizon interactive benchmarks and multiple backbone models, SAG substantially improves the performance-cost trade-off compared with both no-critique and always-on critique agents. On ALFWorld, SAG increases task success from 24.6% to 78.4% while maintaining a token budget comparable to ReAct, yielding a $3.1\times$ improvement in normalized token efficiency. Moreover, a 7B actor with a lightweight 3B critic achieves performance comparable to a 14B actor without critique, showing that selective critique can recover most of the reliability benefits of deliberation while dramatically reducing inference cost.
Chinese Translation
提升大型语言模型(LLM)智能体在长时程决策中的可靠性仍然是一个关键挑战。当作为与复杂环境交互的自主智能体部署时,早期错误会沿轨迹传播并导致级联失败。近期方法通过引入外部评审或审议来提高可靠性,但在每一步都调用这些机制会显著增加 token 消耗和延迟,从而限制实际部署。我们提出 SAG(带门控评审的自改进智能体),这是一个成本感知框架,将评审调用形式化为长时程交互中的逐步决策问题。SAG 引入一种轻量级、无需训练的门控机制,该机制使用在可行动作上计算的动作级歧义信号——全局熵和局部 top-2 间隔——来估计评审的效用。从决策论视角来看,该机制近似评审的信息价值(VoI),使智能体仅在预期收益足以抵偿成本时选择性地分配昂贵的反馈。SAG 进一步纳入在线自举自我改进,使行动者能够内化评审者辅助的行为,并逐步减少对评审的依赖。在三个长时程交互基准和多个骨干模型上,与无评审和始终开启评审的智能体相比,SAG 显著改善了性能-成本权衡。在 ALFWorld 上,SAG 将任务成功率从 24.6% 提高到 78.4%,同时保持与 ReAct 相当的 token 预算,在归一化 token 效率上带来 $3.1\times$ 的提升。此外,一个 7B 行动者配合轻量级 3B 评审者,达到了与无评审的 14B 行动者相当的性能,表明选择性评审能够恢复审议的大部分可靠性收益,同时显著降低推理成本。
cs.LG / 86 / 2610.07348
Stepped MoE: Segment-Level Routing with Configurable Inference Complexity
Stepped MoE:具有可配置推理复杂度的段级路由
large language model
大语言模型相关
Abstract
Training large language models (LLMs) is resource-intensive, and adapting them for diverse deployment scenarios with varying computational constraints remains challenging. While elastic architectures enable flexible model deployment and sparsely activated models allow input-adaptive computation, existing approaches treat these dimensions independently. Moreover, models catered towards on-device edge inference need to conform to the memory and compute limitations of the serving devices. In this paper, we introduce a unified framework that combines elastic structures with sparsely gated architectures to create models that adapt simultaneously to both deployment constraints and task requirements. Our approach employs a model backbone that conditions on both the context and target efficiency specifications, enabling fine-grained control over the accuracy-efficiency trade-off at inference time. The model learns to activate task-relevant parameters within elastically-nested sub-networks, allowing a single model to span multiple capacity points while maintaining input-adaptive routing. Through experiments we demonstrate that we can create a model that allows the flexibility to use 1,2,3,4 billion parameters while being more accurate than their dense counter-parts (2-5\% on knowledge-intensive benchmarks) and at par with their static versions while delivering similar latency metrics as dense models. Overall, we save on device disk space by sharing the model parameters, allow flexibility of serving based on DRAM and compute available while delivering more accurate results.
Chinese Translation
训练大型语言模型(LLMs)是资源密集型的,并且使它们适应具有不同计算约束的多样部署场景仍然具有挑战性。尽管弹性架构能够实现灵活的模型部署,稀疏激活模型允许输入自适应计算,但现有方法将这些维度独立对待。此外,面向设备端边缘推理的模型需要符合服务设备的内存和计算限制。在本文中,我们引入一个统一框架,将弹性结构与稀疏门控架构相结合,以创建能够同时适应部署约束和任务需求的模型。我们的方法采用一个模型主干,它以上下文和目标效率规格为条件,从而能够在推理时对精度-效率权衡进行细粒度控制。该模型学习在弹性嵌套子网络内激活任务相关参数,使单个模型能够跨越多个容量点,同时保持输入自适应路由。通过实验,我们证明我们可以创建一个模型,该模型允许灵活使用 1、2、3、4 十亿参数,同时比它们的稠密对应模型更准确(在知识密集型基准上为 2-5\%),并与它们的静态版本相当,同时提供与稠密模型相似的延迟指标。总体而言,我们通过共享模型参数节省设备磁盘空间,允许基于可用 DRAM 和计算资源灵活提供服务,同时提供更准确的结果。
cs.LG / 87 / 2610.07362
Dynamic Budget Allocation for LLM Evaluation under Hard Resource Constraints
硬资源约束下用于 LLM 评估的动态预算分配
large language model
大语言模型相关
Abstract
We evaluate large language models (LLMs) in multi-turn interactions through their time-to-event: the number of interaction steps required to produce an event of interest, such as a successful jailbreak or agentic task completion. Under limited compute, interactions may be terminated before the event occurs, so that event times are only partially observed (censored). Existing allocation methods for calibrating time-to-event bounds satisfy the budget only in expectation and can exceed the available budget on a particular evaluation run. Enforcing a hard constraint is particularly challenging as the cost of a trajectory is initially unknown. We introduce Hard-budget Allocation with Reflow for Predictive calibration (HARP), a budget allocation that satisfies hard resource constraints and adaptively reallocates unused budget. We show how to use HARP to construct lower predictive bounds (LPBs) on the time-to-event and to estimate evaluation metrics such as the jailbreak rate on a fixed benchmark. Although HARP induces dependence in acquisition decisions across different trajectories, we prove that HARP never exceeds the target budget, that its LPBs have finite-sample coverage guarantees, and that its metric estimates are unbiased. Experiments on agentic task success, LLM jailbreaks, toxic content generation, and RAG hallucinations show that HARP achieves coverage close to the nominal level with low variance, while never exceeding the given budget.
Chinese Translation
我们通过大语言模型(LLM)在多轮交互中的事件时间(time-to-event)来评估它们:产生感兴趣事件所需的交互步数,例如成功越狱或智能体任务完成。在有限计算资源下,交互可能在事件发生之前被终止,因此事件时间只能被部分观测(删失)。用于校准事件时间界的现有分配方法仅在期望意义上满足预算,并且在某次特定评估运行中可能超出可用预算。执行硬约束尤其具有挑战性,因为轨迹的成本最初是未知的。我们提出用于预测校准的回流硬预算分配(Hard-budget Allocation with Reflow for Predictive calibration,HARP),这是一种满足硬资源约束并自适应重新分配未使用预算的预算分配方法。我们展示如何使用 HARP 来构造事件时间的下预测界(lower predictive bounds,LPBs),并在固定基准上估计评估指标,例如越狱率。尽管 HARP 在不同轨迹之间引入了采集决策的依赖性,但我们证明 HARP 从不超过目标预算,其 LPB 具有有限样本覆盖保证,并且其指标估计是无偏的。在智能体任务成功、LLM 越狱、有害内容生成和 RAG 幻觉上的实验表明,HARP 在从不超过给定预算的同时,以低方差实现了接近名义水平的覆盖。
cs.LG / 88 / 2610.07457
AlignQuant: Tile-Aligned Mixed-Precision Quantization for Efficient LLM Generation
AlignQuant:用于高效 LLM 生成的瓦片对齐混合精度量化
large language model
大语言模型相关
Abstract
Fine-grained mixed-precision quantization promises efficient large language model inference, but local precision choices can conflict with regular GPU storage and computation units. This precision-boundary mismatch limits the translation of compression into practical acceleration. We introduce AlignQuant, a post-training quantization method that uses GPU-compatible two-dimensional weight tiles as the common unit of precision allocation, compact storage, and execution. This shared partition lets precision follow sensitivity within output channels. Joint prefill/decode calibration scores precision reductions using projection-output perturbations weighted by language-model loss gradients under quantized activations. Phase-normalized scores prioritize higher precision for tiles important to either phase under a model-wide weight-storage budget. Each tile stores one selected representation, while phase-specialized kernels reuse the packed model and expand lower-bit weights for INT8 computation with 8-bit activations. Across four LLMs spanning 3B to 14B parameters, AlignQuant achieves up to $2.50\times$ generation speedup over BF16 while preserving model quality. Evaluations further cover three GPUs and contexts up to 64K tokens. These results show that local precision flexibility and regular GPU execution can coexist through a shared tile unit. The implementation is available at https://github.com/HanzhiZhang-Ulrica/AlignQuant.
Chinese Translation
细粒度混合精度量化有望实现高效的大语言模型推理,但局部精度选择可能与常规 GPU 存储和计算单元相冲突。这种精度边界不匹配限制了将压缩转化为实际加速。我们提出 AlignQuant,一种训练后量化方法,它使用 GPU 兼容的二维权重瓦片作为精度分配、紧凑存储和执行的共同单元。这种共享划分使精度能够在输出通道内跟随敏感性。联合预填充/解码校准在量化激活下,使用由语言模型损失梯度加权的投影输出扰动来对精度降低进行评分。阶段归一化分数在模型级权重存储预算下,为对任一阶段重要的瓦片优先分配更高精度。每个瓦片存储一种选定的表示,而阶段专用内核复用打包后的模型,并扩展低位权重以使用 8 位激活进行 INT8 计算。在四个参数规模从 3B 到 14B 的 LLM 上,AlignQuant 在保持模型质量的同时,相比 BF16 实现了最高 $2.50\times$ 的生成加速。评估进一步覆盖了三款 GPU 和长达 64K token 的上下文。这些结果表明,局部精度灵活性与常规 GPU 执行可以通过共享瓦片单元共存。实现可在 https://github.com/HanzhiZhang-Ulrica/AlignQuant 获取。
cs.LG / 89 / 2610.07518
Harmful SFT Leaves a Continuous Trace in LLM Checkpoint Updates
有害SFT在LLM检查点更新中留下连续痕迹
large language model
大语言模型相关
Abstract
Safety auditing of post-trained large language models typically relies on model behavior, requiring model execution and depending on the coverage of available evaluations. This work asks a different question: Do the target behaviors optimized during supervised fine-tuning (SFT) leave readable evidence directly in checkpoint updates? We find that harmful-compliance SFT induces a continuous, objective-dependent ordering in checkpoint-update space. Using a reference geometry defined by pure harmful-compliance, safety-targeted, and benign-utility SFT, we find that a checkpoint-level coordinate s_H tracks controlled harmful-objective composition with Spearman correlations of 0.986-0.992 across four 7-8B backbones, with the same ordering persisting at larger model scales. Matched compliance-versus-refusal controls show that this checkpoint trace reflects the SFT objective rather than harmful-input exposure, while additional controls rule out simple explanations based on harmful-example count or generic training intensity. Building on this structure, we introduce TRACE, a weights-only auditing method that localizes an unknown checkpoint update relative to frozen harmful and non-harmful reference prototypes and converts this geometry into a continuous harmful-objective score. TRACE requires neither model queries nor access to the unknown SFT data, and can be evaluated directly from checkpoint updates. Across distribution shifts, unseen data, different SFT configurations, partial checkpoint access, and LoRA/full-parameter fine-tuning, the trace remains stable and is positively associated with independently measured attack success rates. TRACE remains informative even at low harmful-objective proportions, providing a complementary auditing signal when behavioral evaluation is unavailable or incomplete. Code is available at https://anonymous.4open.science/r/Code4TRACE-54D3.
Chinese Translation
后训练大型语言模型的安全审计通常依赖于模型行为,需要执行模型并取决于可用评估的覆盖范围。这项工作提出了一个不同的问题:在监督微调(SFT)过程中优化的目标行为是否会在检查点更新中直接留下可读的证据?我们发现,有害合规SFT在检查点更新空间中诱导出一种连续的、依赖于目标的排序。使用由纯有害合规、安全目标和良性效用SFT定义的参考几何,我们发现检查点级坐标 s_H 追踪受控的有害目标组成,在四个7-8B骨干模型上的Spearman相关系数为0.986-0.992,并且相同的排序在更大模型规模上持续存在。匹配的合规与拒绝对照表明,该检查点痕迹反映的是SFT目标,而非有害输入暴露,而额外的对照排除了基于有害样本数量或一般训练强度的简单解释。在此结构基础上,我们提出TRACE,一种仅基于权重的审计方法,它将未知检查点更新相对于冻结的有害和非有害参考原型进行定位,并将该几何结构转换为连续的有害目标分数。TRACE既不需要模型查询,也不需要访问未知的SFT数据,并且可以直接从检查点更新进行评估。在分布偏移、未见数据、不同SFT配置、部分检查点访问以及LoRA/全参数微调中,该痕迹保持稳定,并与独立测量的攻击成功率正相关。即使在有害目标比例较低时,TRACE仍然具有信息性,在行为评估不可用或不完整时提供补充的审计信号。代码可在 https://anonymous.4open.science/r/Code4TRACE-54D3 获取。
cs.LG / 90 / 2610.07522
Activation Denoising: A Robustness View on Parallel vs Sequential LLM Quantization
激活去噪:并行与串行 LLM 量化的一种鲁棒性视角
large language model
大语言模型相关
Abstract
Post-training quantization is a powerful tool for compressing large language models. The most scalable methods quantize every layer in parallel, but quantization errors then compound through the residual stream, as no layer corrects for the errors of the layers before it. Sequential quantization accounts for this error compounding by re-calibrating each layer on the already-quantized outputs of its predecessors, yielding stronger results but at the cost of a serial schedule that becomes a bottleneck at scale. As a solution, we propose parallel quantization with activation denoising, which recovers much of the sequential benefit while keeping quantization fully parallel. Rather than re-calibrating layer-by-layer, we take a robustness perspective and model the upstream error as noise, regularizing to be robust to it through a preprocessing step followed by metric-weighted rounding. Applied at every layer, this regularization forms a depth-compounding smoothness penalty that dampens how strongly quantization errors amplify through the model. Unlike orthogonal rotations commonly used in quantization, which must preserve the model's function, we multiply the weights by a more general linear transformation. We find that the two are complementary and their effects compound. Empirically, our robustness regularization recovers a significant part of sequential quantization's benefit in a single parallel pass, at a fraction of its time. Overall, by treating compounding quantization errors as a robustness problem, we offer a principled foundation for more efficient and accurate LLM quantization at scale.
Chinese Translation
训练后量化是压缩大语言模型的一种强大工具。最具可扩展性的方法并行地量化每一层,但此时量化误差会沿残差流累积,因为没有任何一层会纠正其之前各层的误差。串行量化通过在其前驱层已量化的输出上重新校准每一层来考虑这种误差累积,从而产生更强的结果,但代价是串行调度,而该调度在大规模下会成为瓶颈。作为解决方案,我们提出带有激活去噪的并行量化,它在保持量化完全并行的同时,恢复了串行量化的大部分收益。我们没有逐层重新校准,而是采取一种鲁棒性视角,将上游误差建模为噪声,并通过一个预处理步骤以及随后的度量加权舍入来进行正则化,以对其保持鲁棒。在每一层应用时,这种正则化形成了一种随深度累积的平滑性惩罚,从而抑制了量化误差在模型中放大的强度。与量化中常用的、必须保持模型函数的正交旋转不同,我们将权重乘以一个更一般的线性变换。我们发现,二者是互补的,且它们的效果会叠加。在实验上,我们的鲁棒性正则化在单次并行遍历中即恢复了串行量化收益的很大一部分,而耗时仅为其一小部分。总体而言,通过将累积的量化误差视为一个鲁棒性问题,我们为大规模下更高效、更准确的 LLM 量化提供了一个有原则的基础。
cs.LG / 91 / 2610.07654
Does On-Policy Distillation for Safety Pose Backdoor Risks?
面向安全性的同策略蒸馏会带来后门风险吗?
large language model
大语言模型相关
Abstract
On-policy distillation (OPD) has attracted growing attention as an effective way to transfer capabilities from teacher models to student models. Recent studies further explore OPD as a tool for improving large language model safety with promising results. However, these approaches typically assume that the teacher and training data are trustworthy. In this paper, we uncover an overlooked threat to OPD for safety: a safety-aligned but backdoored teacher can propagate its hidden malicious behavior to an initially clean student. Under our threat model, a poisoning rate as low as 3% results in an attack success rate (ASR) of up to 70% on the distilled student. We further identify two training choices that can amplify this risk. First, increasing the number of training epochs can lead to high ASR even at low poisoning rates. With only 10 poisoned samples, ASR reaches 67% after 16 epochs. Second, the commonly used top-k KL can accelerate backdoor transfer, causing trigger-conditioned harmful behavior to emerge earlier than sampled-token KL in most settings. Alongside these findings, we explore a simple mitigation, Lazy Defense, which clips KL rewards to make student updates less aggressive, limiting aggressive updates and slowing backdoor learning. Experiments show that Lazy Defense delays backdoor transfer in low poisoning rate settings. Together, our findings reveal that OPD can propagate backdoors, highlighting the need to address the safety risks of OPD.
Chinese Translation
同策略蒸馏(OPD)作为一种将能力从教师模型迁移到学生模型的有效方式,已受到越来越多的关注。近期研究进一步将 OPD 探索为提升大语言模型安全性的工具,并取得了令人期待的结果。然而,这些方法通常假设教师模型和训练数据是可信的。在本文中,我们揭示了 OPD 在安全性方面一个被忽视的威胁:一个安全对齐但被植入后门的教师模型,可以将其隐藏的恶意行为传播给一个初始干净的学生模型。在我们的威胁模型下,低至 3% 的投毒率即可使蒸馏后的学生模型达到高达 70% 的攻击成功率(ASR)。我们进一步识别出两种可能放大这一风险的训练选择。第一,增加训练轮数即使在低投毒率下也能导致高 ASR。仅使用 10 个投毒样本,在 16 轮之后 ASR 即达到 67%。第二,常用的 top-k KL 会加速后门迁移,在大多数设置下使由触发条件引发的有害行为比采样词元 KL 更早出现。在这些发现之外,我们探索了一种简单的缓解方法——Lazy Defense,它通过裁剪 KL 奖励使学生模型的更新不那么激进,从而限制激进的更新并减缓后门学习。实验表明,Lazy Defense 在低投毒率设置下能够延迟后门迁移。综合而言,我们的发现揭示了 OPD 能够传播后门,凸显了应对 OPD 安全风险的必要性。
cs.LG / 92 / 2610.07706
WASD: Wasserstein-based Knowledge Distillation for Large Language Models
WASD:基于 Wasserstein 的大型语言模型知识蒸馏
large language model
大语言模型相关
Abstract
Autoregressive large language models (LLMs) have rapidly advanced in capability, but their increasing scale comes with substantial computational and memory costs at inference time. Knowledge distillation (KD) offers a practical solution by transferring knowledge from a large teacher model to a smaller student model via alignment of discrete probability distributions. However, existing KD methods for LLMs primarily rely on divergences that evaluate discrepancies through probability values at each vocabulary index, without explicitly leveraging token-level semantic information. We propose Wasserstein-based knowledge distillation (WASD) for LLMs, which incorporates token-level semantic information via the Wasserstein-based distance with a cost matrix derived from token embeddings. To ensure computational tractability, we adopt the Sinkhorn divergence and derive a gradient-equivalent objective that can be efficiently optimized without introducing additional networks. Experiments across multiple LLM families and scales show that WASD consistently improves distillation performance on diverse tasks, including instruction following, mathematical reasoning, and code generation. Our results highlight the importance of semantic information encoded in the token space for effective distribution alignment in LLM distillation. The implementation is publicly available at https://github.com/aailab-kaist/WASD .
Chinese Translation
自回归大型语言模型(LLMs)的能力已迅速提升,但其不断增长的规模在推理时伴随着大量的计算和内存成本。知识蒸馏(KD)通过将知识从大型教师模型迁移到较小的学生模型并对齐离散概率分布,提供了一种实用的解决方案。然而,现有针对 LLMs 的 KD 方法主要依赖于通过在每一词表索引处的概率值来评估差异的散度,而没有显式利用词元级语义信息。我们提出用于 LLMs 的基于 Wasserstein 的知识蒸馏(WASD),其通过基于 Wasserstein 的距离以及从词元嵌入导出的代价矩阵来纳入词元级语义信息。为确保计算可处理性,我们采用 Sinkhorn 散度,并推导出一个梯度等价的目标,可在不引入额外网络的情况下高效优化。跨多个 LLM 系列和规模的实验表明,WASD 在包括指令遵循、数学推理和代码生成在内的多种任务上一致地提升蒸馏性能。我们的结果强调了编码在词元空间中的语义信息对于 LLM 蒸馏中有效分布对齐的重要性。实现已公开提供在 https://github.com/aailab-kaist/WASD。
cs.LG / 93 / 2610.07739
Cite What You Explore: Budget-Aware LLM Reasoning over Medical KGs with Verifiable Evidence
引用你所探索的内容:在医学知识图谱上进行具有可验证证据的预算感知 LLM 推理
large language model
大语言模型相关
Abstract
Post-discharge risk prediction from electronic health records (EHRs) is difficult because many dependencies that link discharge-time observations to downstream complications, such as comorbidity cascades and drug-disease interactions, are absent from the record. External medical knowledge graphs (KGs) can supply these missing dependencies, but tracing them demands three properties: KG exploration must remain cost-bounded, retrieved evidence must be differentiated by source quality, and the resulting rationale must be citable for retrospective review. Large language models (LLMs) can plan and verify over structured evidence, making them natural candidates for KG reasoning, but existing LLM-based methods do not satisfy these three properties jointly. In this paper, we propose BAR, a Budget-Aware LLM Reasoning framework over medical KGs with three contributions. First, BAR refines the raw KG into disease-specific evidence graphs whose edges carry support scores and provenance records, turning the KG into a quality-annotated reasoning space rather than a static feature source. Second, an LLM then reasons over this graph through a plan-navigate-verify loop that decomposes the question into steps, retrieves evidence under a patient-specific budget, and revises when verification fails. Third, a reasoning policy is trained with a reward that compares predictions with and without acquired evidence, combined with acquisition cost and citation-integrity terms. Across 8 diseases and 3 prediction horizons on MIMIC-III and MIMIC-IV, BAR improves AUPRC by 3.4 points over the strongest baseline, raises citation precision from 59.8% to 77.9%, and consumes only 62-65% of the budget cap.
Chinese Translation
基于电子健康记录(EHRs)的出院后风险预测很困难,因为许多将出院时观察与下游并发症联系起来的依赖关系,例如共病级联和药物-疾病相互作用,并不存在于记录中。外部医学知识图谱(KGs)能够提供这些缺失的依赖关系,但追踪它们要求三个属性:KG 探索必须保持成本有界,检索到的证据必须依据来源质量加以区分,并且由此得到的推理依据必须可引用以供回顾性审查。大语言模型(LLMs)能够对结构化证据进行规划与验证,这使它们成为 KG 推理的自然候选者,但现有的基于 LLM 的方法并不能同时满足这三个属性。在本文中,我们提出 BAR,一个面向医学 KG 的预算感知 LLM 推理框架,具有三项贡献。首先,BAR 将原始 KG 精炼为疾病特定的证据图,这些图的边携带支持分数和来源记录,从而将 KG 转变为带有质量标注的推理空间,而不是静态的特征来源。其次,LLM 随后通过一个规划-导航-验证循环在此图上进行推理,该循环将问题分解为步骤,在患者特定预算下检索证据,并在验证失败时进行修正。第三,使用一种奖励来训练推理策略,该奖励比较有和没有所获取证据时的预测,并结合获取成本和引用完整性项。在 MIMIC-III 和 MIMIC-IV 上的 8 种疾病和 3 个预测时间范围中,BAR 相较于最强基线将 AUPRC 提高了 3.4 个点,将引用精确率从 59.8% 提升到 77.9%,并且仅消耗预算上限的 62-65%。
cs.LG / 94 / 2610.07767
TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models
TRACE:面向MoE语言模型FP4强化学习的Rollout引导量化感知训练
large language model
大语言模型相关
Abstract
Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantization accuracy on the training and rollout paths independently rather than directly reducing the discrepancy between the two quantized execution paths. In this work, we propose TRACE (Train-Rollout Quantization Alignment via Compact GuidancE), an FP4 quantization framework for RL training of Mixture-of-Experts (MoE) language models that addresses the limitation of existing FP4 RL methods. TRACE incorporates rollout-guided quantization-aware training that uses rollout-side quantization outcomes to guide training-side FP4 rounding decisions, directly reducing train-rollout discrepancy. Moreover, TRACE adopts an efficient quantization-information caching scheme that selectively retains mantissa and scale information from deeper layers to reduce the storage and communication overhead introduced by rollout guidance. We evaluate TRACE on four large-scale MoE language models across reasoning, coding, and long-horizon RL tasks. Our results demonstrate that TRACE enables joint FP4 weight/activation and FP4 KV-cache rollout with RL performance comparable to BF16 rollout, while achieving up to 5.4xrollout speedup and strong final FP4 performance compared with post-hoc FP4 quantization of BF16-trained policies.
Chinese Translation
用于大型语言模型(LLM)后训练的强化学习(RL)在rollout生成过程中会带来大量的计算和内存开销,这促使人们采用低精度rollout以实现高效的RL训练。然而,现有的FP4 RL方法存在一个关键局限:它们主要分别在训练路径和rollout路径上独立地优化量化精度,而不是直接减小这两条量化执行路径之间的差异。在本工作中,我们提出TRACE(Train-Rollout Quantization Alignment via Compact GuidancE,通过紧凑引导实现训练-rollout量化对齐),这是一个用于专家混合(MoE)语言模型RL训练的FP4量化框架,旨在解决现有FP4 RL方法的上述局限。TRACE融合了rollout引导的量化感知训练,该训练利用rollout侧的量化结果来指导训练侧的FP4舍入决策,从而直接减小训练与rollout之间的差异。此外,TRACE采用一种高效的量化信息缓存方案,该方案有选择地保留来自更深层的尾数和缩放信息,以减少由rollout引导引入的存储与通信开销。我们在四个大规模MoE语言模型上,跨推理、编码和长时程RL任务对TRACE进行了评估。我们的结果表明,TRACE能够实现FP4权重/激活与FP4 KV缓存联合的rollout,其RL性能与BF16 rollout相当,同时相比对BF16训练所得策略进行事后FP4量化,可实现最高5.4倍的rollout加速并取得更强的最终FP4性能。
cs.LG / 95 / 2610.07792
ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?
ServeLearnBench:智能体能否从服务经验中自我改进?
large language model
大语言模型相关
Abstract
Large language model agents are increasingly deployed to perform complex tasks in real-world environments. However, the knowledge required for correct behavior in these environments is often implicit, undisclosed, and subject to change over time. Recent continual-learning harnesses seek to address this challenge by enabling agents to improve from serving experience. Yet the effectiveness and limitations of these methods are not yet well characterized. Existing benchmarks provide only partial coverage: some explicitly provide the target knowledge, others assume a static environment, and those that support continual adaptation remain limited in scale and knowledge diversity. To enable systematic evaluation, we formalize an evolving-environment streaming dataset (EESD), in which agents must infer, apply, and revise latent environment knowledge from interaction and outcome feedback as hidden policies evolve, and introduce ServeLearnBench, spanning retail support, banking, and sales-pitch generation with 53 environment windows and 7,718 tasks. We evaluate five learning harnesses (RAG, Mem0, SkillOpt, Continual Harness, and Prime) across six models (GPT-5.6 Terra, Opus 5, Kimi K3, GLM-5.3, DeepSeek V4.1 Flash, and GLM-5.3 Flash), covering 28 model-harness pairs and 252 learning runs. Our evaluation reveals three main findings: a substantial gap remains between task capability and learning from experience; continual adaptation is costly and can degrade already-correct behavior; and insufficient exploration emerges as a key bottleneck to effective adaptation. Overall, ServeLearnBench provides a controlled testbed for diagnosing these limitations and tracking progress toward agents that continually and reliably improve through serving experience.
Chinese Translation
大语言模型智能体正越来越多地被部署到真实世界环境中执行复杂任务。然而,在这些环境中正确行为所需的知识往往是隐性的、未公开的,并且会随时间发生变化。近期的持续学习框架试图通过让智能体从服务经验中改进来应对这一挑战。然而,这些方法的有效性与局限性尚未得到充分刻画。现有基准只提供了部分覆盖:一些显式提供目标知识,另一些假设环境是静态的,而那些支持持续适应的基准在规模与知识多样性上仍然有限。为了实现系统性评估,我们形式化了一个演化环境流式数据集(EESD),其中智能体必须在隐藏策略不断演化的过程中,从交互与结果反馈中推断、应用并修正潜在的环境知识;同时我们提出了 ServeLearnBench,涵盖零售支持、银行以及销售话术生成,包含 53 个环境窗口和 7,718 个任务。我们在六个模型(GPT-5.6 Terra、Opus 5、Kimi K3、GLM-5.3、DeepSeek V4.1 Flash 和 GLM-5.3 Flash)上评估了五种学习框架(RAG、Mem0、SkillOpt、Continual Harness 和 Prime),覆盖 28 个模型—框架组合与 252 次学习运行。我们的评估揭示了三个主要发现:任务能力与从经验中学习之间仍存在显著差距;持续适应代价高昂,并可能使本已正确的行为退化;探索不足成为有效适应的关键瓶颈。总体而言,ServeLearnBench 提供了一个受控测试平台,用于诊断这些局限,并追踪朝着能够通过服务经验持续且可靠地改进的智能体所取得的进展。
cs.LG / 96 / 2610.07819
$α$Transfer: Coefficient Transfer for Efficient Model Merging
$α$Transfer:面向高效模型合并的系数迁移
large language model
大语言模型相关
Abstract
Model merging offers a promising solution for combining multiple fine-tuned checkpoints into a single model through parameter arithmetic. However, finding optimal merging coefficients requires an extensive search that becomes prohibitively expensive as models scale in both size and number, due to high memory requirements and combinatorial growth in the search space. We show that, within the same model family, models exhibit highly congruent performance distributions over merging coefficients across different model sizes. This distributional similarity enables a practical paradigm we call \textit{$α$Transfer}: searching for optimal coefficients on a small proxy model, then directly transfer them to larger target models. We verify $α$Transfer across multiple merging methods, model families, and tasks. Experimental results demonstrate a 6$\times$ speedup and 70\% memory reduction on vision transformers, and a 20$\times$ speedup and 85\% memory reduction on large language models, while maintaining comparable performance. Our findings establish $α$Transfer as an efficient and generalizable approach to scaling model merging.
Chinese Translation
模型合并为通过参数算术将多个微调后的检查点组合成单个模型提供了一种有前景的解决方案。然而,寻找最优合并系数需要进行大量搜索,而随着模型在规模和数量上的扩展,由于高内存需求和搜索空间的组合式增长,这种搜索变得极其昂贵。我们表明,在同一模型家族内,不同模型规模的模型在合并系数上表现出高度一致的性能分布。这种分布相似性促成了一种我们称之为 \textit{$α$Transfer} 的实用范式:在小型代理模型上搜索最优系数,然后将其直接迁移到更大的目标模型。我们在多种合并方法、模型家族和任务上验证了 $α$Transfer。实验结果表明,在视觉 Transformer 上实现了 6$\times$ 加速和 70\% 内存减少,在大型语言模型上实现了 20$\times$ 加速和 85\% 内存减少,同时保持了相当的性能。我们的发现确立了 $α$Transfer 是一种高效且可泛化的扩展模型合并的方法。
cs.LG / 97 / 2610.07967
DecepEval: A Benchmark for Evaluating Deception in LLM Agents
DecepEval:一个用于评估 LLM 智能体中欺骗的基准
large language model
大语言模型相关
Abstract
As large language model (LLM) agents become increasingly autonomous, they may pursue task performance through deception, raising concerns about their reliable deployment. Existing evaluations show that LLM agents can deceive, but often examine isolated scenarios or narrowly defined conditions, limiting systematic understanding of when deception becomes more likely. To address this gap, we introduce DecepEval, a benchmark comprising 1,532 instances across 3 task families and 28 professional scenarios. Drawing on classical fraud theories, we propose the LLM Deception Diamond framework, which characterizes four external conditions that may induce deception: pressure, incentive, opportunity, and conflict. DecepEval pairs neutral and induced versions of each instance to measure condition-dependent changes in deception rates, while explicit task facts and observable agent behavior help distinguish deception from capability-related errors. Evaluations of nine frontier LLMs show that inducements increase deception across models and task families, even among models with low baseline deception rates. DecepEval makes these vulnerabilities measurable, providing a shared benchmark for progress toward trustworthy artificial intelligence.
Chinese Translation
随着大型语言模型(LLM)智能体日益自主化,它们可能通过欺骗来追求任务表现,这引发了对其可靠部署的担忧。现有评估表明,LLM 智能体能够进行欺骗,但往往考察孤立场景或狭义定义的条件,限制了对欺骗何时变得更可能发生的系统性理解。为弥补这一空白,我们提出 DecepEval,一个包含 1,532 个实例、涵盖 3 个任务族和 28 个专业场景的基准。借鉴经典欺诈理论,我们提出 LLM Deception Diamond 框架,该框架刻画了可能诱发欺骗的四种外部条件:压力、诱因、机会和冲突。DecepEval 将每个实例的中性版本与诱导版本配对,以测量欺骗率随条件变化的变化,同时明确的任务事实和可观察的智能体行为有助于将欺骗与能力相关错误区分开来。对九个前沿 LLM 的评估显示,诱导因素会提高不同模型和任务族中的欺骗,即使在基线欺骗率较低的模型中也是如此。DecepEval 使这些脆弱性可被测量,为推动可信人工智能的进展提供了一个共享基准。
cs.LG / 98 / 2610.08021
Spectra: Exact Component Transport for Test-Time Prior Adaptation in Simulation-Based Inference
Spectra:用于基于模拟的推断中测试时先验自适应的精确分量传输
diffusion
扩散模型相关
Abstract
Simulation-based inference (SBI) has become a powerful approach to Bayesian inference in complex scientific models whose likelihoods are difficult or impossible to evaluate. Amortized SBI learns reusable inference models from simulated data, enabling rapid posterior inference for new observations, and modern generative models have made these models increasingly expressive. However, this reuse is limited to the prior distribution chosen during training, whereas scientific analyses often need revised priors as knowledge accumulates or alternative assumptions are tested. We introduce Spectra, a test-time adaptation method for diffusion-based SBI. Spectra uses an exact score-transport identity to obtain the adapted score from a frozen diffusion model in closed form for structured prior changes, without additional simulation or training. Across six SBI benchmarks, Spectra achieves accurate adaptation under strong prior shifts at low online sampling cost. This enables pretrained SBI models to incorporate updated prior information at test time.
Chinese Translation
基于模拟的推断(SBI)已成为在复杂科学模型中进行贝叶斯推断的一种强大方法,这些模型的似然难以或无法评估。摊销式 SBI 从模拟数据中学习可复用的推断模型,使得能够对新观测进行快速后验推断,而现代生成模型已使这些模型日益具有表现力。然而,这种复用仅限于训练期间所选择的先验分布,而随着知识积累或检验替代假设,科学分析往往需要修订后的先验。我们提出 Spectra,一种用于基于扩散的 SBI 的测试时自适应方法。Spectra 使用精确的分数传输恒等式,针对结构化先验变化,以闭式形式从冻结的扩散模型中获得自适应后的分数,而无需额外的模拟或训练。在六个 SBI 基准上,Spectra 以较低的在线采样成本在强烈的先验偏移下实现了准确的自适应。这使预训练的 SBI 模型能够在测试时纳入更新后的先验信息。
cs.LG / 99 / 2610.08108
Enhancing Diffusion Language Models with Autoregressive Post-Training Weights
利用自回归后训练权重增强扩散语言模型
diffusion
扩散模型相关
Abstract
Diffusion language models (dLLMs) have emerged as a promising alternative to autoregressive (AR) language models, offering flexible token-update orders and parallel decoding. Recent dLLMs are often initialized from pretrained AR models before diffusion conversion in order to inherit their learned representations. After the conversion, however, they typically ignore the extensive post-training ecosystem of their AR ancestors. In this work, we show that these existing AR post-training weight updates can instead be effectively recycled to enhance diffusion models. Despite the changes by AR-to-diffusion conversion, directly adding an AR post-training weight update to a diffusion base model remains effective, bringing its performance close to that achieved by direct diffusion post-training. Notably, AR and diffusion post-training updates are nearly orthogonal in weight space, yet induce substantially more aligned representation changes in the diffusion model. Their distinct updates are also complementary: composing their weights can retain gains from both regimes and further improve the post-trained diffusion model. Based on these findings, we propose A2D, a simple training-free framework for enhancing diffusion models with existing AR post-training resources. A2D can transfer capabilities from AR post-trained models to diffusion base models, and further improve already post-trained diffusion models by composing AR and diffusion post-training updates. Across various dLLMs, including Dream, DreamReasoner, DiffuCoder, Dream-Coder, Nemotron-Labs-Diffusion, and DiffusionGemma, A2D reliably improves instruction following, mathematical reasoning, and coding with both supervised fine-tuning and reinforcement learning updates, without additional training, or inference-time computation.
Chinese Translation
扩散语言模型(dLLMs)已成为自回归(AR)语言模型的一种有前景的替代方案,其提供了灵活的 token 更新顺序和并行解码。近期的 dLLMs 通常在进行扩散转换之前从预训练的 AR 模型初始化,以继承其已学习到的表示。然而,在转换之后,它们通常会忽略其 AR 前身所拥有的庞大后训练生态。在本工作中,我们表明,这些现有的 AR 后训练权重更新反而可以被有效地回收利用,以增强扩散模型。尽管 AR 到扩散的转换带来了变化,将 AR 后训练权重更新直接加到扩散基础模型上仍然有效,使其性能接近通过直接扩散后训练所达到的水平。值得注意的是,AR 与扩散后训练更新在权重空间中近乎正交,但它们在扩散模型中引起的表示变化却显著更加一致。它们各自不同的更新也具有互补性:将它们的权重进行组合可以保留来自两种范式的收益,并进一步提升经过后训练的扩散模型。基于这些发现,我们提出了 A2D,一个简单且无需训练的框架,利用现有的 AR 后训练资源来增强扩散模型。A2D 能够将能力从 AR 后训练模型迁移到扩散基础模型,并通过组合 AR 与扩散后训练更新来进一步提升已经过后训练的扩散模型。在多种 dLLMs 上,包括 Dream、DreamReasoner、DiffuCoder、Dream-Coder、Nemotron-Labs-Diffusion 和 DiffusionGemma,A2D 在监督微调和强化学习两种更新下都能可靠地提升指令遵循、数学推理和编码能力,且无需额外训练或推理时的计算。
cs.LG / 100 / 2610.08164
Align, Then Correct: Training-Free Two-Stage Low-Rank Compensation for Extremely Quantized Large Language Models
先对齐,再校正:面向极端量化大语言模型的免训练两阶段低秩补偿
large language model
大语言模型相关
Abstract
Low-rank quantization error compensation (LQEC) recovers the accuracy lost under aggressive weight quantization by attaching a closed-form rank-$r$ adapter beside each frozen quantized weight, without any training. We show that existing compensators are limited by two shared simplifications. They calibrate symmetrically, evaluating the full-precision and compensated weights on the same activation, which yields a compensation target that is inherently high-rank -- so a fixed rank budget captures only a small fraction of it. And they minimize only the second-order term of the loss, although the compensated model is not stationary: a first-order descent direction larger than the applied compensation itself remains in every layer, and no reconstruction objective can absorb it. We propose a two-stage closed-form framework that removes both simplifications. Stage 1 aligns each layer's output with the full-precision model under a Fisher-weighted asymmetric objective, concentrating the rank budget on a rank-compressible target. Stage 2 re-measures statistics on the compensated model and applies a rank-constrained natural-gradient step that absorbs the remaining first-order signal. Every adapter is the result of a single truncated SVD; backward passes serve only to collect statistics. At 2 bits under QuIP#, our method reduces WikiText-2 perplexity from 12.43 to 10.26 on Qwen3-8B and from 21.11 to 13.22 on Qwen3-4B. On the held-out C4 corpus, it recovers 51% and 84% of the gap to FP16, versus 31% and 63% for the strongest baseline, with consistent gains in the seven-task zero-shot average, at higher bit-widths, and under a distinct quantizer.
Chinese Translation
低秩量化误差补偿(LQEC)通过在每个冻结的量化权重旁附加一个闭式 rank-$r$ 适配器,在无需任何训练的情况下恢复激进权重量化下损失的精度。我们表明,现有补偿器受到两个共同简化的限制。它们对称地校准,在相同的激活上评估全精度权重和补偿后的权重,这会产生一个本质上为高秩的补偿目标——因此固定的秩预算只能捕获其中的一小部分。而且它们只最小化损失的二阶项,尽管补偿后的模型并非处于驻点:每一层中仍存在一个比所施加的补偿本身更大的一阶下降方向,且没有任何重构目标能够吸收它。我们提出一个两阶段闭式框架,消除了这两个简化。阶段 1 在 Fisher 加权的非对称目标下使每一层的输出与全精度模型对齐,从而将秩预算集中到一个可秩压缩的目标上。阶段 2 在补偿后的模型上重新测量统计量,并施加一个秩约束的自然梯度步,以吸收剩余的一阶信号。每个适配器都是单个截断 SVD 的结果;反向传播仅用于收集统计量。在 QuIP# 下 2 比特时,我们的方法将 Qwen3-8B 上的 WikiText-2 困惑度从 12.43 降至 10.26,并将 Qwen3-4B 上的从 21.11 降至 13.22。在留出的 C4 语料上,它恢复了与 FP16 差距的 51% 和 84%,而最强基线为 31% 和 63%,并在七任务零样本平均、更高位宽以及不同量化器下均取得一致的增益。
cs.LG / 101 / 2610.08178
On the Intrinsic Limited Robustness of Latent-Based Watermarking
论基于潜空间的水印内在有限的鲁棒性
diffusion
扩散模型相关
Abstract
Existing latent-based watermarking methods for diffusion models have overestimated their robustness to image distortions, including geometric transformations such as rotation, scaling, and translation (RST). Moreover, this paradigm of watermarking approaches may suffer from inherent limitations arising from the domain in which the watermark is embedded. In this paper, we provide the first theoretical analysis explaining why these methods lack invariance to perturbations. By relaxing the invariant relation, we derive a maximum perturbation bound that characterizes the relationship between pixel-space perturbations and their corresponding effects in latent space. In addition, we present the first analytical formulation that captures all components of practical detection mechanisms. Finally, we conduct experiments to validate the theoretical findings and the limitations of latent-based watermarking methods. Our theoretical and empirical results indicate that, under the current design paradigm, latent-based watermarking methods intrinsically exhibit limited robustness. We conclude by providing the analytical tool and design guidelines that future research could follow.
Chinese Translation
现有的面向扩散模型的基于潜空间的水印方法高估了其对图像失真的鲁棒性,这些失真包括旋转、缩放和平移(RST)等几何变换。此外,这一类水印方法范式可能受制于由水印所嵌入的域所带来的固有局限。在本文中,我们提供了首个理论分析,解释了这些方法为何缺乏对扰动的不变性。通过放松不变关系,我们推导出一个最大扰动界,用以刻画像素空间扰动与其在潜空间中对应影响之间的关系。此外,我们提出了首个能够涵盖实际检测机制所有组成部分的解析形式化表述。最后,我们开展实验以验证理论发现以及基于潜空间的水印方法的局限性。我们的理论与实证结果表明,在当前的设计范式下,基于潜空间的水印方法本质上表现出有限的鲁棒性。最后,我们提供了可供未来研究遵循的分析工具与设计准则。
cs.LG / 102 / 2610.08296
OxiGen: Oxidation-State-Aware Crystal Generation
OxiGen:氧化态感知的晶体生成
diffusion
扩散模型相关
Abstract
Generative models have the potential to accelerate inorganic materials discovery by enabling inverse design, but generating experimentally realisable crystals remains challenging. Oxidation states are widely used to assess the compositional validity of crystals and guide inorganic materials discovery. While existing generative models for crystals can generate materials with charge-neutral oxidation-state assignments, they poorly reproduce the distributions of oxidation states observed in synthesised materials. To address this limitation, we propose OxiGen, an oxidation-state-aware crystal diffusion model that explicitly represents oxidation states during generation. OxiGen enforces global charge neutrality by construction using a structured output layer with exact inference over a finite-state automaton. Empirically, OxiGen substantially improves oxidation-state fidelity, generates the highest rate of stable, unique, and novel crystals among evaluated methods, and maintains high compositional validity even under property conditioning.
Chinese Translation
生成模型有潜力通过实现逆向设计来加速无机材料发现,但生成实验上可实现的晶体仍然具有挑战性。氧化态被广泛用于评估晶体的组成有效性,并指导无机材料发现。虽然现有的晶体生成模型能够生成具有电荷中性氧化态分配的材料,但它们难以很好地再现合成材料中观察到的氧化态分布。为了解决这一局限,我们提出了 OxiGen,这是一种氧化态感知的晶体扩散模型,在生成过程中显式表示氧化态。OxiGen 通过构造强制全局电荷中性,其使用一个结构化输出层并在有限状态自动机上进行精确推断。在实证上,OxiGen 显著提高了氧化态保真度,在评估的方法中生成稳定、独特且新颖晶体的比率最高,并且即使在属性条件下也保持高组成有效性。
cs.LG / 103 / 2610.08403
SSR: Sparse Segment Reduction for Ternary GEMM Acceleration
SSR:面向三元 GEMM 加速的稀疏段归约
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) require substantial computational resources, limiting their deployment on resource-constrained hardware. Ternary LLMs mitigate these demands through weight quantization via ternary values, achieving significant compression often with 50-90% sparsity. However, existing approaches have limitations: methods optimized for ternary weights, such as BitNet, redundant segment reduction (RSR), and its improved version RSR++, do not exploit sparsity structures, while conventional sparse formats neglect ternary characteristics, foregoing dual optimization opportunities. In this paper, we introduce Sparse Segment Reduction (SSR), a ternary matrix multiplication method designed to accelerate the inference of ternary LLMs and general Ternary Weight Networks (TWNs). SSR has a dedicated optimized ternary data format and an algorithm that systematically exploits sparsity patterns through computation trees that scale with the sparsity. SSR provides theoretical gains with asymptotically faster inference than RSR++ for sparsity above 50%, while practical evaluations reveal performance improvements across all sparsity levels. Evaluation results show that SSR achieves 2.1-11.3x speedup over RSR++ on ternary GEMM with 45-95% sparsity. Furthermore, SSR achieves 3.5-6.3x end-to-end speedup and 4.9% of memory saving over RSR++ on the Llama-3 1B model inference.
Chinese Translation
大型语言模型(LLMs)需要大量的计算资源,这限制了它们在资源受限硬件上的部署。三元 LLM 通过以三元值进行权重量化来缓解这些需求,通常能在 50-90% 的稀疏度下实现显著压缩。然而,现有方法存在局限性:针对三元权重进行优化的方法,例如 BitNet、冗余段归约(RSR)及其改进版本 RSR++,并未利用稀疏结构,而传统的稀疏格式则忽视了三元特性,从而放弃了双重优化的机会。在本文中,我们提出稀疏段归约(SSR),一种旨在加速三元 LLM 与通用三元权重网络(TWNs)推理的三元矩阵乘法方法。SSR 具有专门优化的三元数据格式,以及一种通过随稀疏度扩展的计算树来系统性地利用稀疏模式的算法。SSR 提供了理论收益:当稀疏度高于 50% 时,其推理在渐近意义上快于 RSR++,而实际评估则显示其在所有稀疏度水平上都带来了性能提升。评估结果表明,在稀疏度为 45-95% 的三元 GEMM 上,SSR 相比 RSR++ 实现了 2.1-11.3 倍的加速。此外,在 Llama-3 1B 模型推理中,SSR 相比 RSR++ 实现了 3.5-6.3 倍的端到端加速以及 4.9% 的内存节省。
cs.LG / 104 / 2610.08475
Learning PDE solution operators with variable initial conditions via Latent Dynamics Networks
通过潜在动力学网络学习具有可变初始条件的 PDE 解算子
diffusion
扩散模型相关
Abstract
In many-query scenarios, data-driven surrogate models provide an efficient alternative to high-fidelity solvers for simulating physical systems governed by Partial Differential Equations (PDEs). In this context, the Latent Dynamics Network (LDNet) has recently demonstrated remarkable performance in predicting the response of spatio-temporal systems, combining Neural Ordinary Differential Equations with nonlinear dimensionality reduction. However, the original formulation assumes a fixed initial condition, limiting its applicability to many real-world applications where a system evolves from varying starting states. In this work, we overcome this limitation while keeping the end-to-end training procedure of the original LDNet and its encoder-free nature, which preserves its intrinsic independence from spatial resolution and grid topology. We infer the initial latent state directly from a small set of early-time observations, treating latent-state initialization as an adaptation problem, and investigate two strategies: an auto-decoding formulation and a meta-learning approach in which the initial latent state acts as a task-specific context variable. We demonstrate the accuracy of the proposed methods across diverse physical phenomena, spanning advection-diffusion, fluid dynamics, and solid mechanics. Meta-learning markedly accelerates latent-state inference and induces smoother, better-conditioned optimization landscapes, and spontaneously organizes the latent space into a structured representation that reflects physically meaningful features of the underlying dynamics. The coordinate-based decoder enables training from spatially subsampled data while recovering high-resolution solution fields at inference. The resulting approach provides an efficient and resolution-independent surrogate modeling framework for many-query simulations of time-dependent PDEs with varying initial conditions.
Chinese Translation
在多查询场景中,对于模拟由偏微分方程(PDE)支配的物理系统,数据驱动的代理模型为高保真求解器提供了一种高效替代方案。在此背景下,潜在动力学网络(LDNet)最近在预测时空系统的响应方面展示了显著性能,它将神经常微分方程与非线性降维相结合。然而,原始表述假设固定的初始条件,这限制了其在许多系统从不同起始状态演化的现实世界应用中的适用性。在这项工作中,我们克服了这一限制,同时保留了原始 LDNet 的端到端训练过程及其无编码器特性,这保持了其本质上独立于空间分辨率和网格拓扑的性质。我们直接从一小组早期时间观测中推断初始潜在状态,将潜在状态初始化视为一个自适应问题,并研究两种策略:一种自动解码表述,以及一种元学习方法,其中初始潜在状态充当任务特定的上下文变量。我们在多种物理现象中展示了所提方法的准确性,涵盖平流-扩散、流体动力学和固体力学。元学习显著加速了潜在状态推断,并诱导出更平滑、条件更优的优化景观,还自发地将潜在空间组织为一种结构化表示,该表示反映了底层动力学具有物理意义的特征。基于坐标的解码器使得能够从空间下采样数据进行训练,同时在推断时恢复高分辨率解场。所得方法为具有可变初始条件的时间相关 PDE 的多次查询模拟提供了一个高效且与分辨率无关的代理建模框架。
cs.LG / 105 / 2610.08564
Valid for Free: Homophily-Gated Conformal Prediction for Training-Free Node Classification with Tabular Foundation Models
免费获得有效性:面向表格基础模型的免训练节点分类的同配性门控保形预测
diffusion
扩散模型相关
Abstract
Tabular foundation models (TFMs) can classify the nodes of a graph without training on it, by reading node and neighborhood features as table rows next to labeled context rows. Work in this line reports predictive performance, not conformal coverage or prediction-set size. To our knowledge, we give the first reliability study of the setting, with TabICL as the TFM and half of each graph as labeled context. As for any predictor fixed before calibration, a frozen in-context predictor makes split conformal prediction exactly valid in finite samples, with no training, validation fold, or tuning on the target graph. An audit across ten graphs then shows that the training-free TabICL posterior has lower expected calibration error (ECE) than GCN with temperature scaling (GCN+TS) on nine of them. Its mean ECE over the ten graphs is 0.019, about 35 percent below the 0.029 of GCN+TS. We also introduce HG-DAPS, a training-free diffusion score whose homophily gate reads only the in-context labels, so the guarantee still holds. Relative to adaptive prediction sets (APS), it reduces mean set size by 5.8 to 17.1 percent on six homophilous graphs and changes it by under 1 percent on four heterophilous ones. On two binary, class-imbalanced graphs, a pre-registered trap case shows that gating on raw rather than adjusted homophily lowers coverage among low-homophily nodes by 0.27 and 0.12. Marginal coverage stays at the nominal 0.90 and masks this drop.
Chinese Translation
表格基础模型(TFMs)无需在图上训练即可对图的节点进行分类,其做法是把节点与邻域特征作为表格行,与带标签的上下文行并列读取。这一方向的工作所报告的是预测性能,而非保形覆盖率或预测集大小。据我们所知,我们给出了该设定下的首个可靠性研究,其中以 TabICL 作为 TFM,并将每个图的一半作为带标签上下文。正如对于任何在校准之前就已固定的预测器一样,一个冻结的上下文内预测器使分裂保形预测在有限样本下精确有效,且无需在目标图上进行训练、划分验证折或调参。随后在十个图上进行的审查表明,免训练的 TabICL 后验在其中九个图上的期望校准误差(ECE)低于带温度缩放的 GCN(GCN+TS)。它在十个图上的平均 ECE 为 0.019,比 GCN+TS 的 0.029 低约 35%。我们还提出了 HG-DAPS,一种免训练的扩散得分,其同配性门控仅读取上下文内标签,因此该保证依然成立。相对于自适应预测集(APS),它在六个同配图上将平均集合大小降低 5.8% 至 17.1%,并在四个异配图上使其变化小于 1%。在两个二分类、类别不平衡的图上,一个预先注册的陷阱案例表明,基于原始同配性而非调整后同配性进行门控,会使低同配性节点的覆盖率降低 0.27 和 0.12。边际覆盖率仍保持在名义值 0.90,从而掩盖了这一下降。
cs.MA / 106 / 2610.07535
Disentangling Models from Personas in Heterogeneous LLM Simulations
在异构 LLM 模拟中区分模型与人设
large language model
大语言模型相关
Abstract
Multi-agent simulations with large language models (LLMs) often operate networks of agents with a single base model. This overlooks the inter-model effects which may dominate engagement dynamics in real-world deployments. To show this, we simulate a heterogeneous social network powered by several different base models and show that the amount of engagement an agent receives depends more on its base model than on its assigned persona. The attraction or repulsion effects of a base model strengthen dramatically when more models are added in the mix, suggesting that networks dynamics may converge to base model effects at scale. To help explain this effect, we conduct a series of content-mediating analyses, showing the predictability of base models across contexts as well as the relationship between a model's lexical patterns and an engagement-maximizing style. In light of recent developments in mass multi-agent interaction, this work underscores the relevance of heterogeneous compositions in driving the outcomes of those networks
Chinese Translation
使用大语言模型(LLM)的多智能体模拟通常以单一基座模型运行智能体网络。这忽略了模型间效应,而这些效应可能在真实世界部署中主导参与动态。为说明这一点,我们模拟了一个由若干不同基座模型驱动的异构社交网络,并表明一个智能体获得的参与量更多取决于其基座模型,而不是分配给它的角色人设。当混合中加入更多模型时,一个基座模型的吸引或排斥效应会显著增强,这表明网络动态可能在大规模下收敛到基座模型效应。为帮助解释这一效应,我们进行了一系列内容中介分析,展示了基座模型跨情境的可预测性,以及一个模型的词汇模式与最大化参与度的风格之间的关系。鉴于近期在大规模多智能体交互方面的发展,这项工作强调了异构组成在驱动这些网络结果方面的相关性
cs.MA / 107 / 2610.08155
Token-Efficient Multi-Agent Collaboration via System One-Guided Computational Division of Labor
通过系统一引导的计算分工实现Token高效的多智能体协作
large language model
大语言模型相关
Abstract
Large language model (LLM)-based multi-agent systems (MAS) have become a promising paradigm for complex information-seeking and reasoning tasks by enabling collaborative problem solving among specialized agents. However, existing MAS frameworks tightly couple task reasoning with coordination operations, including task selection, role assignment, message routing, and context management. As interactions grow, using powerful LLMs for these bounded control decisions introduces substantial token overhead and latency, limiting the scalability of agentic Web services. In this paper, we investigate whether coordination can be decoupled from expensive reasoning without compromising collaborative performance. We propose S1-MAS, a token-efficient multi-agent framework based on System One-guided computational division of labor. S1-MAS assigns bounded coordination decisions to lightweight System One models while reserving open-ended reasoning for capable LLM workers. Specifically, a lightweight controller selects inspection conditions, chooses subsequent tasks, and determines termination, while a compact reader retrieves condition-relevant evidence from authorized sources to support these decisions. Through a decision-evidence loop, selected tasks dynamically determine worker roles and source access, enabling adaptive collaboration without task-specific training. Extensive experiments on seven diverse benchmarks demonstrate that S1-MAS achieves superior accuracy while substantially reducing the inference cost. Across individual comparisons with AgentVerse, DyLAN, and SelfOrg on seven benchmarks, S1-MAS reduces GPT-4o token consumption by 44.9%-97.2% and measured end-to-end latency by 37.8%-93.0%. These results highlight its potential for scalable and cost-effective agentic Web applications.
Chinese Translation
基于大语言模型(LLM)的多智能体系统(MAS)通过使专用智能体之间能够协作解决问题,已成为复杂信息检索与推理任务的一种有前景的范式。然而,现有MAS框架将任务推理与协调操作紧密耦合,这些操作包括任务选择、角色分配、消息路由和上下文管理。随着交互增长,使用强大的LLM进行这些有界控制决策会引入大量的token开销和延迟,限制了智能体式Web服务的可扩展性。在本文中,我们研究协调是否可以在不损害协作性能的情况下与昂贵的推理进行解耦。我们提出S1-MAS,一个基于系统一引导的计算分工的token高效多智能体框架。S1-MAS将有界协调决策分配给轻量级系统一模型,同时为有能力的LLM工作者保留开放式推理。具体而言,一个轻量级控制器选择检查条件、选择后续任务并确定终止,而一个紧凑型读取器从授权来源检索与条件相关的证据以支持这些决策。通过决策-证据循环,选定的任务动态地确定工作者角色和来源访问,从而实现无需任务特定训练的自适应协作。在七个多样化基准上的大量实验表明,S1-MAS在显著降低推理成本的同时实现了更优的准确率。在七个基准上与AgentVerse、DyLAN和SelfOrg的逐项比较中,S1-MAS将GPT-4o的token消耗降低了44.9%-97.2%,并将测得的端到端延迟降低了37.8%-93.0%。这些结果凸显了其在可扩展且具有成本效益的智能体式Web应用方面的潜力。
cs.AI / 108 / 2610.08595
One for All, All for One: Coordinated Multi-Agent Diffusion Steering via Stochastic Optimal Control
我为人人,人人为我:通过随机最优控制的协同多智能体扩散引导
diffusion
扩散模型相关
Abstract
Deep generative models often produce structured outputs composed of interacting components. Modelling these outputs with a single model requires learning both the component distributions and their interactions. We pursue a modular alternative: reuse independently trained component generators and learn only how to coordinate them to produce coherent structured outputs. Our framework, Coordinated Multi-Agent Diffusion Steering (CMDS), treats frozen pretrained diffusion models as reusable generative primitives and coordinates their reverse processes through a learned control. We formulate coordination as a stochastic optimal control problem, balancing an assembly-level reward that specifies the desired properties of the combined output against deviations from the pretrained dynamics. The learned control amortises this optimisation, allowing reuse across new task instances. Experiments show that CMDS can recover a known target distribution, satisfy different spatial constraints with the same trained control, and recover individual sources from degraded mixtures. Across multi-agent maze navigation, articulated robot planning, and text-conditioned human motion, CMDS turns frozen models into coordinated multi-agent generators.
Chinese Translation
深度生成模型通常产生由相互作用组件构成的结构化输出。用单个模型对这些输出进行建模需要同时学习组件分布及其相互作用。我们追求一种模块化替代方案:复用独立训练的组件生成器,并仅学习如何协调它们以产生连贯的结构化输出。我们的框架——协同多智能体扩散引导(CMDS)——将冻结的预训练扩散模型视为可复用的生成基元,并通过学习到的控制来协调它们的反向过程。我们将协调形式化为一个随机最优控制问题,在指定组合输出期望属性的装配级奖励与偏离预训练动力学之间进行平衡。学习到的控制对该优化进行摊销,从而允许在新任务实例之间复用。实验表明,CMDS 能够恢复已知的目标分布,用同一个训练好的控制满足不同的空间约束,并从退化混合中恢复单个源。在多智能体迷宫导航、铰接式机器人规划以及文本条件人体运动等任务中,CMDS 将冻结模型转变为协同的多智能体生成器。
cs.AI / 109 / 2610.08780
DepthWorld: 3D World Model for Robot Manipulation
DepthWorld:用于机器人操作的3D世界模型
diffusion
扩散模型相关
Abstract
World models offer a data-driven alternative to traditional simulators for robotics, with applications spanning policy evaluation, improvement, and planning. All of these uses depend on faithful 3D geometry, yet current video-based world models are trained on RGB alone and produce rollouts that look correct frame-by-frame but do not compose into a consistent 3D world. Closing this gap requires progress on two fronts: large-scale 3D supervision for manipulation, and an architecture that can absorb it without disturbing strong pretrained video priors. We introduce a calibration pipeline that combines learned stereo depth with a joint factor graph, pooling all episodes collected from the same physical robot to recover its shared kinematic parameters alongside per-scene extrinsics. Applied to the DROID dataset, this yields DROID-3D, a calibrated 3D dataset providing dense metric depth and recalibrated multi-view extrinsics (achieving <0.7 px reprojection error on 90% of episodes for external cameras). We then train DepthWorld, a Stable Video Diffusion-based world model that jointly predicts multi-view RGB and depth via spatial latent tiling, leaving the pretrained Variational Autoencoder (VAE) unchanged. Depth supervision improves RGB prediction itself by +1.48 dB PSNR over an identical RGB-only baseline at equal training budget, while simultaneously yielding accurate metric depth for downstream geometric reasoning.
Chinese Translation
世界模型为机器人学中的传统模拟器提供了一种数据驱动的替代方案,其应用涵盖策略评估、策略改进与规划。所有这些用途都依赖于忠实的3D几何,然而当前基于视频的世界模型仅使用RGB进行训练,所产生的推演逐帧看起来正确,却无法组合成一个一致的3D世界。弥合这一差距需要在两个方面取得进展:面向操作任务的大规模3D监督,以及一种能够吸收这些监督信息而不扰乱强大的预训练视频先验的架构。我们提出了一条标定流程,它将学习得到的立体深度与一个联合因子图相结合,汇集从同一台物理机器人上采集的所有回合,以在恢复每个场景的外参的同时恢复其共享的运动学参数。将其应用于DROID数据集后,这产生了DROID-3D,这是一个经过标定的3D数据集,提供稠密的度量深度和重新标定的多视角外参(对于外部相机,在90%的回合上实现了<0.7 px的重投影误差)。随后我们训练了DepthWorld,这是一个基于Stable Video Diffusion的世界模型,它通过空间隐变量分块联合预测多视角RGB与深度,同时保持预训练的变分自编码器(VAE)不变。在相同的训练预算下,深度监督使RGB预测本身相较于完全相同的仅RGB基线提升了+1.48 dB PSNR,同时还能为下游几何推理给出准确的度量深度。
cs.AI / 110 / 2610.08760
WorldSonus: Bringing Sound to Worlds
WorldSonus:将声音带入世界
diffusion
扩散模型相关
Abstract
Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time generation to keep pace with interactive video streams, interactive control to respond to mid-stream sound instructions, and spatially aligned stereo to reflect scene geometry and camera motion. To address these demands, we introduce WorldSonus, an interactive video-to-audio framework designed for real-time spatial sound synthesis in world models. For real-time generation, WorldSonus employs a streaming causal autoregressive diffusion architecture that synthesizes audio chunks at a low real-time factor (RTF) of 0.41. For interactive control, we incorporate an audio-centric captioning pipeline with chunk-indexed prompt scheduling, enabling dynamic manipulation of sound events during generation. For spatial alignment, we leverage high-quality stereo supervision curated from diverse stereo and ambisonic data. Extensive experiments demonstrate that while tailored for world models, WorldSonus generalizes effectively to open-domain video-to-audio benchmarks, matching or outperforming state-of-the-art bidirectional models in both acoustic quality and spatial alignment. Project page: https://noizai.github.io/WorldSonus/
Chinese Translation
近年来,世界模型的进展使得视觉合成越来越逼真。然而,这些生成的环境在很大程度上仍然是无声的。将声音引入世界模型带来了三个核心挑战:实时生成以跟上交互式视频流,交互式控制以响应流中声音指令,以及空间对齐的立体声以反映场景几何和相机运动。为满足这些需求,我们提出了 WorldSonus,一个面向世界模型中实时空间声音合成的交互式视频到音频框架。为了实现实时生成,WorldSonus 采用了一种流式因果自回归扩散架构,该架构以 0.41 的低实时因子(RTF)合成音频块。为了实现交互式控制,我们加入了一个以音频为中心的描述生成流水线,并采用按块索引的提示调度,从而能够在生成过程中动态操控声音事件。为了实现空间对齐,我们利用了从多样化的立体声和 Ambisonic 数据中整理出的高质量立体声监督。大量实验表明,尽管 WorldSonus 是为世界模型量身定制的,但它能够有效泛化到开放域视频到音频基准,在声学质量和空间对齐方面匹配或超越最先进的双向模型。项目页面:https://noizai.github.io/WorldSonus/
cs.SE / 111 / 2610.07289
Catching Developers in the Flow: Low-Latency Agentic Program Repair at Google Scale
捕捉处于心流状态的开发者:Google 规模下的低延迟智能体程序修复
large language model
大语言模型相关
Abstract
Manual repair of program failures is time-consuming and disruptive for software developers, particularly during the pre-submit phase where test failures occur within continuous integration systems. While Automated Program Repair has seen significant advancement through Large Language Models, existing state-of-the-art techniques primarily focus on post-submit workflows, operating offline without the low-latency requirements necessary to assist developers in real-time within their flow before they switch context. In this paper, we introduce FlowAgent, an AI agent deployed at Google to automatically repair test failures in the pre-submit outer-loop workflow inside continuous integration systems. Integrated into Google's internal developer tools, Critique and Cider,FlowAgent utilizes a ReAct-style generate-and-validate loop, as well as rigorous pre-execution and post-execution abstention filters to ensure high-quality suggestions under strict latency constraints. Based on our case studies, FlowAgent is highly effective. First, a manual evaluation conducted on 195 real-world test failures demonstrated 67.18% accuracy in suggesting correct fixes. Following its Google-wide deployment, FlowAgent suggested fixes on 295,508changes, of which developers previewed 65,069 and applied 28,554. Developer feedback from interviews indicate that the agent is useful in suggesting correct fixes, integration of autonomous repair agents into industrial software engineering workflows is received well, while interesting challenges and opportunities still remain.
Chinese Translation
对程序失败的手动修复既耗时又会打断软件开发者的工作,尤其是在提交前阶段——此时测试失败发生在持续集成系统内部。尽管自动化程序修复(Automated Program Repair)借助大语言模型取得了显著进展,但现有的最先进技术主要聚焦于提交后的工作流,以离线方式运行,缺乏在开发者切换上下文之前、于其心流之中实时协助他们所需的低延迟能力。在本文中,我们介绍了 FlowAgent,这是一个部署在 Google 的 AI 智能体,用于在持续集成系统内部的提交前外循环工作流中自动修复测试失败。FlowAgent 被集成进 Google 的内部开发者工具 Critique 和 Cider,它采用 ReAct 风格的“生成—验证”循环,以及严格的执行前与执行后弃权过滤器,以在严苛的延迟约束下确保高质量的建议。基于我们的案例研究,FlowAgent 极为有效。首先,对 195 个真实世界测试失败所做的人工评估表明,其在建议正确修复方案上的准确率为 67.18%。在其于 Google 全面部署之后,FlowAgent 对 295,508 项变更给出了修复建议,其中开发者预览了 65,069 项,采纳应用了 28,554 项。来自访谈的开发者反馈表明,该智能体在建议正确修复方案方面很有用,将自主修复智能体集成到工业软件工程工作流中的做法受到良好接纳,同时仍存在有趣的挑战与机遇。
cs.SE / 112 / 2610.07356
A Validated Dataset and Benchmark for Coherent Multi-Diagram SysML Models
用于连贯多图 SysML 模型的经过验证的数据集与基准
large language model
大语言模型相关
Abstract
Systems engineers use several diagrams to describe the structure and behavior of systems. Engineers create these diagrams together to make sure that they use the same elements and remain consistent with one another. Large language models can generate diagrams as text or code, which makes it possible to create system diagrams automatically. However, their ability to generate coherent sets of diagrams is not well understood, and existing datasets and benchmarks do not directly measure this ability at scale. We introduce SEMAADB (Systems Engineering Modeling Assistant with AI Dataset and Benchmark), a dataset of 3,000 engineering contexts and 15,000 diagrams. Each context contains five connected SysML views: Requirement, Block Definition, Activity, State Machine, and Sequence. Here, a view is a diagram that presents one aspect of a system. We checked the diagram sets for consistency and valid rendering. A set of 100 contexts is also human-verified and forms the benchmark test set. We evaluate three language models on two tasks. In diagram repair, the strongest model repairs 64.3% of semantic errors . In cross-diagram update the best propagation F1 is 80.7% when a model applies one change across related diagrams. The results show that syntax repair is nearly solved, but semantic repair and consistency across diagrams are still challenging tasks for models. SEMAADB therefore provides both a large diagram resource and a set of benchmarks for measuring coherent multi-diagram SysML generation.
Chinese Translation
系统工程师使用若干图来描述系统的结构和行为。工程师将这些图一起创建,以确保它们使用相同的元素并彼此保持一致。大型语言模型可以将图生成为文本或代码,这使得自动创建系统图成为可能。然而,它们生成连贯图集的能力尚未被充分理解,现有数据集和基准并未直接大规模地衡量这种能力。我们引入 SEMAADB(Systems Engineering Modeling Assistant with AI Dataset and Benchmark,即带有 AI 数据集和基准的系统工程建模助手),一个包含 3,000 个工程上下文和 15,000 个图的数据集。每个上下文包含五个相互关联的 SysML 视图:需求、块定义、活动、状态机和序列。在这里,视图是呈现系统某一方面的图。我们检查了图集的一致性和渲染有效性。一组 100 个上下文还经过人工验证,并构成基准测试集。我们在两个任务上评估了三个语言模型。在图修复中,最强的模型修复了 64.3% 的语义错误。在跨图更新中,当模型将一项更改应用到相关图时,最佳传播 F1 为 80.7%。结果表明,语法修复已接近解决,但语义修复和跨图一致性对模型而言仍是具有挑战性的任务。因此,SEMAADB 既提供了大型图资源,也提供了一组用于衡量连贯多图 SysML 生成的基准。
cs.SE / 113 / 2610.07446
CogAdapt: Cognition-informed Sparse Adaptation of Code LLMs
CogAdapt:认知信息引导的代码大语言模型稀疏适配
large language model
大语言模型相关
Abstract
Large language models (LLMs) have become increasingly capable of generating code. However, achieving stronger code-generation performance still often relies on costly model adaptation, i.e., fine-tuning pretrained model parameters. Prior studies have shown correspondence between human code processing and neural models' attention or internal computation. Human-aligned learning approaches use cognitive signals to guide training, but typically adapt a large portion of the model, leaving training costs largely unchanged. Human cognitive signals may indicate not only what the model must learn from, but also where adaptation is most useful. We investigate whether human responses during code reading correspond to code-model behavior and can guide selective adaptation without sacrificing performance. We present CogAdapt, a cognition-informed framework for task-dependent sparse adaptation of code models. CogAdapt first learns transferable program-level and token-level priors from human Electroencephalography (EEG) and attention data, then combines these priors with the frozen model's response to each coding task to determine how much adaptation to allocate and which transformer blocks should receive updates. During fine-tuning, only the selected blocks are updated, while no new human recordings are required for inference. Across Qwen and GLM, we find consistent correspondence between human reading behavior and Mixture-of-Experts (MoE) computation. CogAdapt achieves the best pass@1 across both LiveCodeBench and BigCodeBench, including gains of 10.86 and 6.29 percentage points over matched regular fine-tuning on LiveCodeBench, while reducing gradient-eligible adaptation parameters by 86.21-87.21%. These results suggest that human comprehension signals can provide useful guidance for making code-model adaptation both more selective and more effective.
Chinese Translation
大语言模型(LLMs)在生成代码方面的能力日益增强。然而,实现更强的代码生成性能仍然通常依赖于代价高昂的模型适配,即微调预训练模型参数。已有研究表明,人类代码处理与神经模型的注意力或内部计算之间存在对应关系。与人类对齐的学习方法使用认知信号来指导训练,但通常会适配模型的大部分,从而使训练成本基本保持不变。人类认知信号可能不仅指示模型必须学习什么,还指示适配在何处最有用。我们研究人类在代码阅读过程中的反应是否与代码模型行为相对应,并且能否在不牺牲性能的情况下指导选择性适配。我们提出 CogAdapt,一个用于代码模型任务相关稀疏适配的认知信息引导框架。CogAdapt 首先从人类脑电图(EEG)和注意力数据中学习可迁移的程序级和词元级先验,然后将这些先验与冻结模型对每个编码任务的响应相结合,以确定应分配多少适配以及哪些 transformer 块应接收更新。在微调期间,仅更新所选的块,而推理时不需要新的人类记录。在 Qwen 和 GLM 上,我们发现人类阅读行为与混合专家(MoE)计算之间存在一致的对应关系。CogAdapt 在 LiveCodeBench 和 BigCodeBench 上均取得了最佳的 pass@1,包括在 LiveCodeBench 上相较匹配的常规微调分别提升 10.86 和 6.29 个百分点,同时将可参与梯度更新的适配参数减少 86.21-87.21%。这些结果表明,人类理解信号可以为使代码模型适配更具选择性和更有效提供有用的指导。
cs.SE / 114 / 2610.07832
Harness Engineering for Software Engineering via Modular Executable Dev-Primitives
面向软件工程的 Harness Engineering:基于模块化可执行开发原语
large language model
大语言模型相关
Abstract
Large language models (LLMs) equipped with terminal access have demonstrated strong capabilities in automating software engineering tasks. However, existing agents remain brittle on long-horizon workflows, where they must repeatedly reconstruct program state scattered across source files, configurations, tests, dependencies, and runtime behavior, leading to increasingly long interaction histories, context explosion, and semantic drift. Large repositories further complicate the identification of task-relevant components. To address these challenges, we introduce \textbf{Dev-Primitives} (\emph{Development Primitives}), a modular and executable abstraction that transforms repository components from passive software artifacts into active participants in software engineering. Each Dev-Primitive pairs a repository artifact with a resident LLM, which gives the artifact an agent-native interface grounded in its own implementation and dependencies, enabling natural-language reasoning, inter-component communication, and localized self-modification. Building on Dev-Primitives, we propose \textbf{HERMES}, a Harness Engineering framework for software engineeRing via Modular Executable Dev-PrimitiveS, which instantiates these primitives at repository scale through a dependency-aware dynamic activation mechanism and a bug diagnosis mechanism that maps execution evidence back to the components that must be revised. Extensive experiments on four software engineering benchmarks demonstrate that HERMES outperforms matched baseline harnesses by 12.4\% on average. Moreover, when paired with strong activation and diagnosis models, HERMES, even with Qwen3-8B Dev-Primitives, remains within 4.5\% of the homogeneous GPT-5.6 Sol configuration across all four benchmarks, while reducing inference cost by 26.2\% on Terminal-Bench 4.0, highlighting the importance of harness design in software engineering agents.
Chinese Translation
具备终端访问能力的大语言模型(LLM)已展现出在自动化软件工程任务方面的强大能力。然而,现有智能体在长时程工作流上仍然表现脆弱,在这类工作流中,它们必须反复重建分散在源文件、配置、测试、依赖项和运行时行为中的程序状态,从而导致交互历史不断增长、上下文爆炸和语义漂移。大型代码仓库进一步加剧了识别任务相关组件的难度。为应对这些挑战,我们提出了 \textbf{Dev-Primitives}(\emph{Development Primitives},开发原语),一种模块化且可执行的抽象,它将仓库组件从被动的软件制品转变为软件工程中的主动参与者。每个 Dev-Primitive 将一个仓库制品与一个常驻 LLM 配对,从而为该制品提供一个以其自身实现和依赖项为基础的智能体原生接口,使其能够进行自然语言推理、组件间通信以及局部自修改。在 Dev-Primitives 的基础上,我们提出了 \textbf{HERMES},一个通过模块化可执行开发原语开展软件工程(Harness Engineering for software engineeRing via Modular Executable Dev-PrimitiveS)的框架,它通过依赖感知的动态激活机制以及将执行证据映射回必须修订组件的缺陷诊断机制,在仓库规模上实例化这些原语。在四个软件工程基准上的大量实验表明,HERMES 平均比匹配的基线 harness 高出 12.4\%。此外,当与强大的激活模型和诊断模型搭配使用时,即便使用 Qwen3-8B Dev-Primitives,HERMES 在全部四个基准上仍与同构的 GPT-5.6 Sol 配置相差不超过 4.5\%,同时在 Terminal-Bench 4.0 上将推理成本降低了 26.2\%,这凸显了 harness 设计在软件工程智能体中的重要性。
cs.SE / 115 / 2610.08240
Newer and Bigger, but Safer? A Longitudinal Study of the Functionality-Security Gap in LLM-Generated Code
更新、更大,但更安全吗?对 LLM 生成代码中功能-安全差距的纵向研究
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) are widely used to generate code. Although their functional plausibility keeps improving, the generated code often contains security vulnerabilities. The functionality-security gap captures code that passes functional tests but fails security tests. A recent longitudinal study of three model families concluded that LLMs become smarter but not safer, with the only considered open-weight family stagnating. Whether this holds for other (open-weight) families and particularly for compact models remains open. We present a longitudinal study of the gap across 32 LLMs from seven model families (five open-weight), covering three successive releases per family in flagship and compact variants. Using CWEval with 119 tasks in five programming languages and 31 CWEs, we compare trajectories across families, model sizes, and languages. Newer models do become safer in absolute terms, although no family closes the gap. Unlike prior work, we find that openness does not separate the families: every considered open-weight family narrows the gap significantly, while Gemini 3.1 Pro keeps a gap as wide as the one reported for Llama. Compact models usually produce less secure code than their flagship counterparts, with notable exceptions (e.g., Gemini 3.7 Flash). At the CWE level, we confirm persistent weaknesses such as log injection (CWE-117) and HTTP response splitting (CWE-113) and regressions in the newest proprietary models on memory and integer weaknesses, and show that the same CWE carries very different risk across languages. From these results, we derive implications for LLM vendors, researchers, and developers. In particular, developers should assume neither that upgrades improve security nor that proprietary models are more secure; they should rerun security checks after each model change and provide secure APIs in the model's context.
Chinese Translation
大型语言模型(LLM)被广泛用于生成代码。尽管其功能上的合理性不断提升,但生成的代码往往包含安全漏洞。功能-安全差距刻画的是那些通过了功能测试但未通过安全测试的代码。最近一项针对三个模型家族的纵向研究得出结论:LLM 变得更聪明,但并未变得更安全,其中唯一被纳入考虑的开源权重家族停滞不前。这一结论是否适用于其他(开源权重)家族,尤其是紧凑型模型,仍有待明确。我们提出了一项针对来自七个模型家族(其中五个为开源权重)的 32 个 LLM 的差距纵向研究,涵盖每个家族旗舰版与紧凑版变体的连续三个发布版本。我们使用包含五种编程语言中 119 个任务以及 31 个 CWE 的 CWEval,比较了不同家族、模型规模和语言之间的轨迹。就绝对水平而言,更新的模型确实变得更安全,尽管没有任何一个家族弥合了这一差距。与先前工作不同,我们发现开放性并不能将这些家族区分开来:所有被考虑的开源权重家族都显著缩小了这一差距,而 Gemini 3.1 Pro 的差距仍与 Llama 所报告的差距一样大。紧凑型模型通常比其旗舰版对应模型生成安全性更低的代码,但也有显著的例外(例如 Gemini 3.7 Flash)。在 CWE 层面,我们确认了诸如日志注入(CWE-117)和 HTTP 响应拆分(CWE-113)等持续存在的弱点,以及最新专有模型在内存和整数弱点上的退化,并表明同一个 CWE 在不同语言中承载着非常不同的风险。基于这些结果,我们为 LLM 供应商、研究人员和开发者推导出了一些启示。特别地,开发者既不应假定升级能够改善安全性,也不应假定专有模型更加安全;他们应在每次模型变更后重新运行安全检查,并在模型的上下文中提供安全的 API。
cs.SE / 116 / 2610.08405
Learning from Failures: A Failure-Driven Prompt Refinement for LLM-Based Vulnerability Analysis
从失败中学习:面向基于 LLM 的漏洞分析的失败驱动提示词精炼
large language model
大语言模型相关
Abstract
Large Language Models have emerged as promising tools for software vulnerability analysis, but their effectiveness depends heavily on prompt design. Existing research primarily compares prompting strategies using aggregate performance metrics, providing limited insight into why models fail or how prompts can be improved systematically. We propose Failure-Driven Prompt Refinement (FDPR), a methodology that analyzes recurring model failures to guide evidence-based prompt refinement. Using the Damn Vulnerable Java Application (DVJA), we identify recurring failure modes, including false positives, false negatives, unsupported reasoning, and CWE misclassification, and translate them into targeted prompt refinements. We then evaluate the resulting prompt on the Juliet Test Suite and perform cross-model validation to assess generalizability. The results show that failure-driven refinement improves the reliability of LLM-based vulnerability analysis while yielding reusable prompt design principles. More broadly, this work demonstrates that recurring model failures provide a principled foundation for prompt engineering, enabling the systematic development of more reliable LLM-based vulnerability analysis systems.
Chinese Translation
大型语言模型已成为软件漏洞分析中颇具前景的工具,但其有效性在很大程度上取决于提示词设计。现有研究主要使用聚合性能指标来比较提示策略,对于模型为何失败或如何系统地改进提示词所提供的见解有限。我们提出失败驱动提示词精炼(Failure-Driven Prompt Refinement, FDPR),这是一种分析反复出现的模型失败以指导基于证据的提示词精炼的方法。使用 Damn Vulnerable Java Application (DVJA),我们识别出反复出现的失败模式,包括假阳性、假阴性、无依据推理和 CWE 错误分类,并将其转化为有针对性的提示词精炼。随后,我们在 Juliet Test Suite 上评估所得提示词,并进行跨模型验证以评估泛化性。结果表明,失败驱动精炼提高了基于 LLM 的漏洞分析的可靠性,同时产生了可复用的提示词设计原则。更广泛地说,这项工作表明,反复出现的模型失败为提示工程提供了有原则的基础,使得能够系统地开发更可靠的基于 LLM 的漏洞分析系统。
cs.LG / 117 / 2610.08764
Rapid Fredholm stabilization of the Kuramoto--Sivashinsky equation with unrestricted, spatially-varying anti-diffusion
具有无限制、空间变化反扩散的 Kuramoto--Sivashinsky 方程的快速 Fredholm 镇定
diffusion
扩散模型相关
Abstract
We develop the first feedback design for rapid stabilization of the Kuramoto--Sivashinsky equation with a spatially varying anti-diffusion coefficient. For constant coefficients, the single-input Fredholm design of Coron and Lü (2015) excludes a discrete set of values at which repeated unstable eigenvalues cause a loss of controllability. We overcome this obstruction by introducing a second boundary input and assigning the two inputs distinct roles. The key idea, inspired by Heymann's Lemma, is to use the boundary value $u(0,t)$ entirely for a pre-feedback that renders the modified plant controllable through the curvature input $u_{xx}(0,t)$. The latter input then stabilizes the plant through a Fredholm backstepping transformation. We show that two inputs suffice for controllability and are necessary when the plant has an unstable double eigenvalue. However, the Fredholm kernel still must be approximated for implementation. Hence, to enable kernel and gain approximation, we prove continuity of the coefficient-to-gain design map on compact admissible design classes. Unlike Volterra-based continuity proofs using successive approximations, our proof uses the modal representation to control the spectral data, the inverse coefficient system, and the tails of the kernel and gain series. This yields a single neural operator approximation of the gain to any prescribed $L^2$ accuracy across the class. Finally, we establish rapid local stabilization of the nonlinear closed-loop system under both the exact gains and sufficiently accurate approximations. We conclude with numerical results that illustrate prescribed decay rates and the computational cost of the approximations. In particular, we train a Fourier neural operator that achieves typical relative gain errors of approximately $0.1\%$ and stabilizes all held-out cases tested, including a plant with an unstable double eigenvalue.
Chinese Translation
我们开发了首个针对具有空间变化反扩散系数的 Kuramoto--Sivashinsky 方程快速镇定的反馈设计。对于常系数,Coron 和 Lü (2015) 的单输入 Fredholm 设计排除了一个离散值集,在这些值处,重复的不稳定特征值导致可控性丧失。我们通过引入第二个边界输入并为这两个输入分配不同的角色来克服这一障碍。关键思想受 Heymann 引理启发,是将边界值 $u(0,t)$ 完全用于一个预反馈,该预反馈使修改后的被控对象可通过曲率输入 $u_{xx}(0,t)$ 控制。后一个输入随后通过 Fredholm 反步变换镇定被控对象。我们证明两个输入足以实现可控性,并且当被控对象具有不稳定二重特征值时它们是必要的。然而,为了实现,Fredholm 核仍然必须被近似。因此,为了能够进行核和增益近似,我们证明系数到增益设计映射在紧容许设计类上的连续性。与使用逐次逼近的基于 Volterra 的连续性证明不同,我们的证明使用模态表示来控制谱数据、逆系数系统以及核与增益级数的尾部。这给出了增益的单个神经算子近似,在整个类上达到任意规定的 $L^2$ 精度。最后,我们建立了非线性闭环系统在精确增益和足够精确近似两种情况下的快速局部镇定。我们以数值结果作结,这些结果展示了规定的衰减率以及这些近似的计算成本。特别地,我们训练了一个 Fourier 神经算子,其达到约 $0.1\%$ 的典型相对增益误差,并镇定所有测试过的留出案例,包括一个具有不稳定二重特征值的被控对象。
cs.LG / 118 / 2610.07602
A Neural JKO Scheme for Hellinger-Kantorovich Gradient Flows via Monge-Growth Pairs
一种通过 Monge-Growth 对的 Hellinger-Kantorovich 梯度流神经 JKO 方案
diffusion
扩散模型相关
Abstract
We develop a mesh-free neural JKO scheme for advection-reaction-diffusion equations with a gradient-flow structure in the Hellinger-Kantorovich (HK) geometry of unbalanced optimal transport. Each update is parametrized by a spatial map and a mass-changing factor, allowing spatial redistribution and local mass creation or loss to be treated jointly within a single variational step. Their cone action bounds the squared HK distance from above, yielding a sufficient condition for discrete energy dissipation through comparison with the identity pair. Minimizing the pair objective over all admissible pairs recovers the exact JKO minimum when the source and a minimizer have positive densities. We establish existence and mass bounds for JKO minimizers and, under additional assumptions, obtain positivity and regularity together with a discrete Euler-Lagrange equation and a metric-dissipation identity. The self-consistent chemical potential is then nonincreasing along an optimal map. There exist parametric pairs whose endpoint densities and objective values converge to those of an exact JKO minimizer, provided a regular-pair approximation hypothesis holds. Finally, we show that a primal-dual gap controls objective suboptimality and, for Boltzmann entropy, the $L^1$ density error, assuming exact-step regularity, positive-semidefinite interactions, and global dual feasibility. Numerical experiments examine pointwise agreement with the PDE, energy dissipation, and the roles of transport, reaction, and fully implicit interactions.
Chinese Translation
我们提出了一种无网格神经 JKO 方案,用于在非平衡最优传输的 Hellinger-Kantorovich (HK) 几何中处理具有梯度流结构的平流-反应-扩散方程。每次更新由一个空间映射和一个质量变化因子参数化,从而使空间重分布以及局部质量产生或损失能够在单个变分步内被联合处理。它们的锥作用从上方界定平方 HK 距离,通过与恒等对比较,产生离散能量耗散的充分条件。当源与一个极小化子具有正密度时,在所有容许对上最小化对目标可恢复精确 JKO 最小值。我们建立了 JKO 极小化子的存在性和质量界,并且在额外假设下,获得正性和正则性,以及离散 Euler-Lagrange 方程和度量耗散恒等式。于是,自洽化学势沿最优映射非增。存在参数化对,其端点密度和目标值收敛到精确 JKO 极小化子的端点密度和目标值,只要正则对逼近假设成立。最后,我们证明原始-对偶间隙控制目标次优性,并且对于 Boltzmann 熵,控制 $L^1$ 密度误差,假设精确步正则性、半正定相互作用和全局对偶可行性。数值实验考察了与 PDE 的逐点一致性、能量耗散,以及传输、反应和全隐式相互作用的作用。
cs.LG / 119 / 2610.07884
Learned Adaptive Multiresolution Diffusion Imaging
学习型自适应多分辨率扩散成像
diffusion
扩散模型相关
Abstract
Adaptive multiresolution methods reduce representation cost by concentrating fine-scale degrees of freedom where needed, but their tree updates are usually governed by fixed local criteria. We introduce Learned Adaptive Multiresolution Diffusion Imaging (Learned AMDI), which preserves the AMDI fixed-tree propagator and hierarchy constraints while replacing the post-propagation selector with a shared local policy trained by proximal policy optimization. Regression tests reproduce deterministic AMDI trajectories to machine precision when identical trees are used. In the Haar implementation studied here, the deterministic one-step selector accepts no refinements in 54 decisions. Across nine held-out cases, Learned AMDI executes 393 refinements and reduces the mean terminal reference discrepancy from $0.17496$ to $0.13657$, while occupancy rises from $0.13737$ to $0.26660$. Step-resolved diagnostics reveal occasional small adaptation-energy increases; fixed-tree energy stability therefore does not guarantee monotonicity of the learned outer iteration. At comparable occupancy, a validation-tuned observed-detail threshold reaches a discrepancy of $0.13792$ with slightly better RMSE and SSIM, placing both methods on essentially the same accuracy--occupancy tradeoff. A decision-1-only control reaches $0.13742$, indicating that most of the improvement on this static benchmark arises from the initial allocation. The shared actor transfers without retraining to $64\times64$ and $128\times128$ images, improving reference discrepancy, RMSE, and SSIM relative to deterministic AMDI, while the frozen threshold rule remains competitive. Learned AMDI thus provides a hierarchy-constrained, resolution-transferable mechanism for adaptive allocation and clarifies the contribution of sequential decisions.
Chinese Translation
自适应多分辨率方法通过将细尺度自由度集中到需要之处来降低表示成本,但其树更新通常由固定的局部准则支配。我们提出学习型自适应多分辨率扩散成像(Learned AMDI),它保留了 AMDI 的固定树传播器和层次约束,同时将传播后的选择器替换为由近端策略优化训练的共享局部策略。当使用相同的树时,回归测试能以机器精度复现确定性 AMDI 轨迹。在此处研究的 Haar 实现中,确定性一步选择器在 54 次决策中不接受任何细化。在九个留出案例上,Learned AMDI 执行了 393 次细化,并将平均终端参考差异从 $0.17496$ 降低到 $0.13657$,同时占用率从 $0.13737$ 上升到 $0.26660$。逐步解析诊断揭示了偶尔出现的小幅自适应能量增加;因此固定树能量稳定性并不能保证学习到的外层迭代的单调性。在相当占用率下,经验证调优的观测细节阈值达到 $0.13792$ 的差异,并具有略好的 RMSE 和 SSIM,使两种方法基本处于相同的准确率--占用率权衡上。仅决策 1 的对照达到 $0.13742$,表明在这个静态基准上大部分改进来自初始分配。共享 actor 无需重新训练即可迁移到 $64\times64$ 和 $128\times128$ 图像,相对于确定性 AMDI 改善了参考差异、RMSE 和 SSIM,而冻结阈值规则仍具竞争力。因此,Learned AMDI 提供了一种受层次约束、可跨分辨率迁移的自适应分配机制,并阐明了顺序决策的贡献。
cs.LG / 120 / 2610.07655
Uniform Discrete Diffusion Models are Minimax Optimal for Estimating Distributions with Small Effective Support Size
均匀离散扩散模型在估计具有小有效支撑大小的分布时是极小极大最优的
diffusion
扩散模型相关
Abstract
Discrete diffusion models have emerged as a practically successful framework for generative modeling on discrete product spaces, yet their statistical generalization properties remain poorly understood. Discrete real-world data such as text or biological sequences often concentrate on a small fraction of the astronomically large ambient space because of semantic or physical constraints, but existing bounds fail to capture this distributional structure and instead scale with the size of the ambient space, giving rise to almost vacuous error bounds. We address this gap for uniform discrete diffusion, one of the two dominant discrete diffusion paradigms alongside masking diffusion, by deriving statistical guarantees governed by the effective support size $s_n(P_0)$, a sample-size-dependent measure of distributional complexity. Given $n$ independent and identically distributed (i.i.d.) samples from an unknown data distribution $P_0$ on $[K]^d$, we show that, with appropriate choices of network size and hyperparameters, the expected total variation (TV) loss scales as $O(\sqrt{s_n(P_0)/n})$, while the expected Kullback--Leibler (KL) divergence is bounded by $O(\frac{1}{n}s_n(P_0)\log(eK^d/s_n(P_0))\log n)$. Furthermore, we show that the TV rate is minimax optimal and that the KL rate is minimax optimal up to a factor of $\log n$. Together, these upper and lower bounds show that uniform discrete diffusion successfully avoids the curse of dimensionality for distributions with small effective support size: the TV error rate depends on the ambient state-space size only through $s_n(P_0)$, while the corresponding KL rate incurs only an additional logarithmic dependence on the ambient state-space size.
Chinese Translation
离散扩散模型已成为离散乘积空间上生成建模的一个在实践中成功的框架,但它们的统计泛化性质仍然知之甚少。诸如文本或生物序列之类的离散现实世界数据,由于语义或物理约束,往往集中在天文数字般巨大的环境空间中的一小部分上,但现有界未能捕捉这种分布结构,反而随环境空间的大小而缩放,从而导致几乎空洞的误差界。我们通过推导由有效支撑大小 $s_n(P_0)$(一种依赖于样本量的分布复杂度度量)控制的统计保证,来填补均匀离散扩散方面的这一空白,均匀离散扩散是与掩蔽扩散并列的两种主要离散扩散范式之一。给定来自 $[K]^d$ 上未知数据分布 $P_0$ 的 $n$ 个独立同分布(i.i.d.)样本,我们表明,在适当选择网络规模和超参数的情况下,期望全变差(TV)损失按 $O(\sqrt{s_n(P_0)/n})$ 缩放,而期望 Kullback--Leibler(KL)散度被 $O(\frac{1}{n}s_n(P_0)\log(eK^d/s_n(P_0))\log n)$ 所界。此外,我们表明 TV 速率是极小极大最优的,而 KL 速率在相差一个 $\log n$ 因子的意义下是极小极大最优的。这些上界和下界共同表明,对于具有小有效支撑大小的分布,均匀离散扩散成功避免了维数灾难:TV 误差率仅通过 $s_n(P_0)$ 依赖于环境状态空间大小,而相应的 KL 速率仅额外产生对环境状态空间大小的对数依赖。
cs.LG / 121 / 2610.07755
Trustworthy Method Comparison with AI Judges: Estimation and Design under Order, Batch, and Aggregation Effects
以 AI 评判者进行可信的方法比较:顺序、批次与聚合效应下的估计与设计
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly used as judges for automated AI evaluation. A common practice is to randomize prompt sequences and average the resulting scores, but its statistical validity remains unclear. We show that LLM evaluation mechanisms can be approximated by a class of Markov generalized linear mixed models (GLMMs), supported by out-of-sample predictions across three major commercial LLMs. Using a first-order Markov GLMM, we study leaderboard ranking and group comparison. For leaderboard ranking, randomize-and-average selection is consistent under a mild separation condition, and a Williams square design can improve efficiency when item qualities are close. For group comparison, naive averaging can yield inconsistent conclusions about differences in group-level quality because of the response model's nonlinearity. Empirical results further support the validity of the proposed model-based inference beyond the first-order theory, including settings with higher-order sequence memory. We illustrate the approach in an application where AI judges compare two graphical model estimation methods.
Chinese Translation
大语言模型(LLM)正日益被用作自动化 AI 评估的评判者。一种常见做法是随机化提示词序列并对所得分数取平均,但其统计有效性仍不明确。我们表明,LLM 评估机制可由一类马尔可夫广义线性混合模型(GLMM)近似,并在三个主要商业 LLM 上通过样本外预测得到支持。使用一阶马尔可夫 GLMM,我们研究排行榜排名与组间比较。对于排行榜排名,随机化并平均的选择在温和的分离条件下具有一致性,并且当项目质量接近时,Williams 方设计可以提高效率。对于组间比较,由于响应模型的非线性,朴素平均可能对组级质量差异得出不一致的结论。实证结果进一步支持了所提出的基于模型的推断在一阶理论之外的有效性,包括具有更高阶序列记忆的设定。我们在一个应用中使用 AI 评判者比较两种图模型估计方法来展示该方法。
cs.AI / 122 / 2610.08626
Feature Information Dynamics in Diffusion
扩散中的特征信息动力学
diffusion
扩散模型相关
Abstract
Diffusion models generate data through a continuum of denoising problems, and are widely observed to reveal coarse structure before fine detail. Yet, this intuition is mostly empirical and qualitative. We introduce feature information dynamics, an information-theoretic framework for localizing when a feature is generated during diffusion. Using the I-MMSE identity, we connect the rate of feature mutual information change to a gap between optimal unconditional and feature-conditional denoising losses, yielding practical estimators for feature information density. We further develop a chained decomposition that separates shared from incremental information in a feature hierarchy. We use this framework first to quantitatively confirm spectral autoregression in pixel diffusion, and then to extend the analysis beyond frequency: under a class $\to$ mask $\to$ Canny conditioning chain, the per-feature information densities differ across pixel, SDVAE, VAVAE, and RAE, exposing fundamental differences between these representations and suggesting that ordered generation could be beneficial for training diffusion models. Our code is available at https://github.com/AI4Science-WestlakeU/feature-information-dynamics.
Chinese Translation
扩散模型通过一系列连续的去噪问题生成数据,并且被广泛观察到会在精细细节之前先揭示粗略结构。然而,这种直觉大多是经验性的和定性的。我们引入特征信息动力学,这是一个信息论框架,用于定位一个特征在扩散过程中何时被生成。利用 I-MMSE 恒等式,我们将特征互信息的变化率与最优无条件去噪损失和特征条件去噪损失之间的差距联系起来,从而得到特征信息密度的实用估计器。我们进一步提出一种链式分解,将特征层次结构中的共享信息与增量信息分离开来。我们首先使用该框架来定量确认像素扩散中的谱自回归,然后将分析扩展到频率之外:在类别 $\to$ 掩码 $\to$ Canny 条件链下,逐特征信息密度在像素、SDVAE、VAVAE 和 RAE 之间存在差异,暴露出这些表示之间的根本差异,并表明有序生成可能有利于训练扩散模型。我们的代码可在 https://github.com/AI4Science-WestlakeU/feature-information-dynamics 获取。
cs.LG / 123 / 2610.08652
Steering Diffusion Models to Rare Events with Sequential Monte Carlo
用序列蒙特卡洛引导扩散模型走向稀有事件
diffusion
扩散模型相关
Abstract
Diffusion models are increasingly used as surrogates for expensive simulators in weather prediction, molecular dynamics, and materials design. In these models, computing the probability $p_0[E]$ of an event $E$ is difficult, especially when the event of interest is rare. A stable estimate using Monte Carlo becomes computationally intractable, requiring a growing sample size $\propto\!1/p_0[E]$ to compensate for an increasing rarity. In this paper, we present Diffusion Importance Sampling of Rare Events or DireSMC, a sequential Monte Carlo scheme that guides a population of weighted samples towards the rare event, giving access not only to samples but also to a calibrated estimate of its probability. We set up our guidance using an analytical relaxation of the event set, allowing the method to easily extend to a wide range of user-defined rare events. We validate our method on a toy problem with analytical solutions and on a score-based climate emulator, where we obtain accurate rare-event probabilities on a range of rarities from $10^{-3}$ to $10^{-5}$, achieving net speed-ups of $9\times$ to $1413\times$ over Monte Carlo.
Chinese Translation
扩散模型正越来越多地被用作天气预报、分子动力学和材料设计中昂贵模拟器的替代模型。在这些模型中,计算事件 $E$ 的概率 $p_0[E]$ 很困难,尤其是当所关注的事件为稀有事件时。使用蒙特卡洛进行稳定估计在计算上变得不可行,因为需要不断增大的样本量 $\propto\!1/p_0[E]$ 来补偿日益增加的稀有性。在本文中,我们提出面向稀有事件的扩散重要性采样,即 DireSMC,这是一种序列蒙特卡洛方案,它引导一组加权样本走向稀有事件,不仅能够获取样本,还能够获得其概率的校准估计。我们使用事件集的解析松弛来构建引导,使该方法能够轻松扩展到各种用户定义的稀有事件。我们在一个具有解析解的玩具问题以及一个基于分数的气候模拟器上验证了我们的方法,在该模拟器上,我们在从 $10^{-3}$ 到 $10^{-5}$ 的一系列稀有性上获得了准确的稀有事件概率,相较于蒙特卡洛实现了 $9\times$ 到 $1413\times$ 的净加速。
人工智能 (cs.AI)
156
cs.AI / 1 / 2610.07862
A self-learning scientific agent for X-ray diffraction
Abstract
A central challenge for scientific agents is to turn analytical experience into reusable expertise grounded in physical evidence. Here we introduce Gan Jiang, a self-learning agent for powder X-ray diffraction built on a diffraction-analysis ecosystem we developed: XMatcher, XQueryer, XDecomposer and WPEM. Together, these engines span phase identification, multiphase decomposition and physics-constrained whole-pattern modelling. Gan Jiang converts analytical experience into executable skills by diagnosing failures, revising skill instructions and code, and validating revisions before reuse, without retraining the language model or changing the underlying physical models. Skills selected using development data and frozen before held-out evaluation achieve higher refinement scores than the original expert-designed skills across FullProf, GSAS-II and PyWPEM. The agent resolves strongly overlapping reflections, quantifies a five-phase ancient Egyptian cosmetic, tracks lattice evolution in an operating battery and compares atomic configurations in a disordered oxide catalyst. On DeltaXRDbench, it leads the evaluated methods in single- and multiphase identification across simulated and experimental data. Without supplied composition, single-phase top-1 accuracies reach 96.30\%, 81.78\% and 40.83\% on MP500, RRUFF and opXRD, respectively, compared with 58.00\%, 58.47\% and 26.45\% for the strongest comparator. These results demonstrate how an integrated scientific tool ecosystem can support agents that extract structural knowledge from measurements while accumulating validated analytical expertise that transfers to new samples.
cs.AI / 2 / 2610.07192
Sim-to-Real Transfer of Vision-Language Navigation in Continuous Environments Using an Ackermann-Steered Mobile Robot
Abstract
Vision-Language Navigation (VLN) enables robots to navigate through environments using natural language instructions, making human-robot interaction intuitive. Traditional VLN models often rely on navigation graphs, 360-degree views, and perfect localization which pose significant challenges when adapting these models to real-world settings. This work addresses these limitations by performing a simulation-to-real domain shift of a VLN approach that operates in continuous environments without requiring navigation graphs or panoramic views. The proposed system integrates vision-language models that align visual inputs and linguistic instructions within a shared embedding space, facilitating natural language-driven navigation. We employ a Cross-Modal Attention (CMA) based architecture trained on an existing dataset in a simulated environment and fine-tune it using real-world data collected from a custom-built Ackermann-steered robot equipped with a camera and a LiDAR sensor. By utilising linear photometric adjustments and fine-tuning on a limited number of episodes, our model successfully adapts to real-world environments, achieving effective navigation while running offline on dedicated hardware. Experimental results, evaluated using Success weighted by Path Length (SPL) and Normalized Dynamic Time Warping (nDTW) metrics, demonstrate the robustness and adaptability of our approach. Keywords: Vision-Language Navigation, Cross-Modal Attention, Natural Language Instructions, Sim-to-Real Transfer, Autonomous Navigation, Ackermann-steering.
cs.AI / 3 / 2610.07219
Cascadia: Resident 975B MoE Inference on Eleven AI PCs
Abstract
Mixture-of-experts models make nearly trillion-parameter capacity accessible with sparse per-token computation, provided that the serving system can distribute the weights and coordinate their execution. We present Cascadia's resident execution of Inkling, a 975B-total/41B-active-parameter model, on eleven Intel Core Ultra X7 358H AI PCs, each with 64 GB of memory, Arc B390 integrated graphics and gigabit Ethernet. We contribute a custom resident MoE engine that preserves Inkling's routing rules, constructs compressed graphs for OpenVINO's fused iGPU primitives, and coordinates FP16 expert computation with FP32 output restoration. The engine fits six consecutive decoder layers per machine and represents dense feed-forward blocks as all-active expert slices, reducing measured dense-layer call time from approximately 8.1 to 4.5 ms. A streaming pipeline coordinates concurrent generation, while captured-state draft evaluation measures agreement with the deployed numerical path. Paired measurements at fifteen concurrency levels from 1 to 176 streams reach 60.29 aggregate decode tokens/s at 88 streams, with 46.87 tokens/s over the complete serving phases. At fifteen streams, median first-token latency is 6.05 s. Raising the context budget from the 1,024-position default, real prompts of 1k to 64k tokens recover the embedded code in all 19 measured answers, with first-token time growing as $aN+bN^2$ and decode latency growing approximately linearly, both bounded by a single-threaded CPU attention loop rather than by memory, which holds 512k positions per stream. Evaluation on captured fleet states separates the effects of vocabulary selection and weight quantization on draft agreement. Together, these contributions establish an execution and evaluation approach for large sparse models on distributed client systems with shared CPU-GPU memory.
cs.AI / 4 / 2610.07237
SPECTRUM: Proximal Spectral Modulation for Looped Self-Distillation
Abstract
A model that learns from its own outputs inherits more than their correctness: it inherits which solutions it produces. We formulate Looped Self-Distillation, a self-evolution framework for code generation in which a model repeatedly generates and learns from its own raw outputs, under a fixed information budget, without ongoing external assessment or test-based selection of the generated samples. We identify a consequential separation: correctness can improve while the breadth of correct implementations contracts. We introduce SPECTRUM, which re-estimates loss-sensitive key/value geometry from a fixed reference anchor at each round and converts it into full-rank proximal spectral modulation. All generated completions train a single student, whose subsequent inference requires no intervention. After five rounds of experiments on MBPP, SPECTRUM retains 89.9% of the initial model's 64-sample correct AST richness, compared with 66.4% for Vanilla self-distillation and 65.5% for a subspace-projection control. The advantage persists at matched correct-sample counts. Without further training or recalibration, the resulting student also achieves higher matched-correct richness than Vanilla SD on HumanEval+ and APPS Intro, demonstrating transfer of the diversity benefit. These findings establish correct-solution retention as a complementary objective of recursive self-improvement (RSI) and show that generation-time intervention can improve the solution repertoire retained by subsequent students.
cs.AI / 5 / 2610.07249
Can Semantic Geometry Teach an AI Judgement?
Abstract
How can an AI agent determine what rules to follow? One rule permits an action. Another imposes a condition, exception, or conflicting obligation. Deterministic systems can resolve those relationships when they have been specified. When they remain implicit in language, an agent can follow one rule while missing another that should stop it. Refusing every unresolved action avoids that risk, but also blocks permissible actions. We wanted the agent to make the distinction and still act. Our initial hypothesis was that geometric measurements could supply a basis for judgment. We represented actions and policies as vectors, then tested whether their geometry could identify governing policies and interpret the action's relation to them. Across four studies, the tested approaches did not establish reliable pre-action judgment. In the final synthetic study, a lexical router recovered every governing and blocking policy while reducing median policy checks by 97.7%. The composed pipeline nevertheless escalated all 2,304 test actions, including those it should have allowed. Supplying every policy to the same downstream mechanism changed no decision. Finding the policies had not solved the problem of interpreting them. This result led us to revise our hypothesis: judgment in AI agents requires developing a consequence graph. Such a graph would connect the actor and authority to policy conditions, exceptions, and the changes an action would produce. Follow-on studies will ask whether making those relationships explicit helps the agent distinguish when to act, stop, or seek review.
cs.AI / 6 / 2610.07257
MemMux: Runtime Verification and Honest Resource Attribution for Fleets of Parallel Coding Agents
Abstract
Developers increasingly run a fleet of coding agents side by side on one workstation. The tools they reach for, terminal multiplexers like tmux and a new generation of agent managers, were built to arrange windows, not to govern memory. When ten agents each spawn language servers, test runners, and browsers, no standard tool can say how much memory belongs to which agent, confirm that a terminated agent's descendants are gone, notice a child that has escaped its agent, or keep the machine off the swap cliff when an OOM kill would silently discard uncommitted work. We treat these as runtime-verification problems: an agent-hosting substrate should continuously emit observable signals an operator or auditor can check while agents run. We present MemMux, a local runtime that turns resource governance into checkable signals (per-agent attribution, complete reclamation, escaped-process visibility, bounded footprint under overcommit, and monitoring overhead), with a claims-disciplined benchmark against tmux, a purpose-built agent multiplexer, and a raw-process baseline on identical workloads. Under a binding memory budget on a Linux host, MemMux keeps the fleet under budget (7.5 GiB) with zero swap by admitting a subset and reclaiming under pressure, while the ungoverned tools run every agent, pin the machine at its RAM ceiling (2x over budget), and spill about 2 GiB into swap. MemMux reclaims 100% of a terminated agent's process subtree where the raw baseline strands half of it, and it alone surfaces escaped children (10 of 10 detected). We report the cost: the 1 Hz attribution scan runs near 0.6% CPU at one agent but 2.7% at ten, above our 2% target. Running the harness on real Claude Code sessions shows 100% attribution and low overhead carry over to live agent trees. We release the engine, benchmark, and a one-command reproducer.
cs.AI / 7 / 2610.07261
Verifying Coordination in Parallel Coding Agents: NP-Bench and a Scheduling Planner
Abstract
A team of coding agents can look fine agent by agent yet fail as a team: each passes its own tests while the merged result is broken, and single-agent evaluation never catches it. As teams run several LLM coding agents in parallel on one codebase, the agents collide: two rewrite the same function, one codes against a contract a teammate just changed, and integration fails after the work is done. Most coordination tools react (watch for a conflict, then warn), but at agent speed the warning arrives after the wasted edit. We recast the problem as scheduling: take each work item's declared scope, partition the work into disjoint scopes, and order merges along the producer->consumer graph, all up front. We build this planner into Nerveplane and evaluate it with NP-Bench, an environment-grounded three-arm benchmark (no coordination; reactive detection; proactive planning) that verifies integration off a real git merge, both in a deterministic simulation and with live agents. The planner lifts clean-integration from 1/9 to 9/9 scenarios and cuts merge conflicts from 13 to 0, with a gap that grows in the number of agents. On a live breaking contract change it rescues an outcome both baselines miss on every seed: the clean-integration rate rises from 0 (no coordination and reactive detection) to 1.0 on a frontier model and 0.6 on a small one, while agents respect assigned scopes (0/5 leakage). A cross-session memory drops the repeated-mistake rate from 1.00 to 0.00 on strong and weak models alike. We also report a negative result: routing facts to agents does not rescue long-context accuracy at window-fitting scales; its value is cost and capacity, not attention. Across two capability tiers and two vendors, the benefit did not shrink as models got stronger, because it comes from how work is allocated, not model reasoning.
cs.AI / 8 / 2610.07270
Does the Model Use the Feature? Separating Steering from Mechanism in LLMs
Abstract
Internal features in LLMs are often interpreted as mechanisms when they track a concept and their manipulation changes a related behavior. Yet steering can push a feature far outside its natural range, where its effects need not reflect the model's own computation. We examine this inference and propose an empirical contract whose tests evaluate features at values observed on natural inputs. One test copies a feature's value from an input that shows a behavior into a matched input that does not (installation) or the reverse (removal); the other restores the feature after an upstream edit (downstream rescue). Installation measures how far the feature suffices for the behavior; removal and downstream rescue measure how much the model uses it. Applied to three kinds of representations, the two strengths separate sharply. The published unknown-entity latent strongly steers knowledge abstention, yet installing observed values from either published latent into matched prompts transfers only a small fraction of the natural known--unknown abstention contrast. Dense known--unknown directions show opposite asymmetries between installation and removal in Gemma and Llama, and how fully a released subject--verb agreement feature set reproduces and restores the behavior depends on how its values are written into the model. Tracking a concept and steering a behavior therefore do not by themselves show that the model uses a feature, and each conclusion holds only for the intervention tested.
cs.AI / 9 / 2610.07274
A Trust Layer for Agent Evaluation
Abstract
Deterministic benchmark scores show that an agent received credit, but not whether that credit was earned, reported honestly, or would hold on a second run. We introduce a Trust Layer for Agent Evaluation, an additive post-hoc framework that reports, beside each recorded score, whether it should be believed. It verifies four properties: whether the result is supported by the benchmark's own grading logic, whether a passing answer was earned through traceable computation, whether the agent's completion claim matches what occurred, and whether the result is stable under repeated execution. The first three use only saved artifacts; the fourth re-runs the agent. Model judgments only label evidence under majority voting; all verdicts follow deterministic rules and never modify the recorded score. Applied to five agent configurations on 108 tasks from Agents' Last Exam, every model shows passing runs with no traceable computation (at rates varying tenfold), confirmed false completion claims, and unstable results: 18-46% of tasks do not stay in one score band over five runs. Only 22.6% of recorded passes clear all four checks (95% CI 15.0-32.6, n=84). Measuring what an agent can do and verifying that it did it are different problems, and current benchmarks address only the first.
cs.AI / 10 / 2610.07309
The Right Memory in the Wrong Context: Verifying Retrieval Admissibility in Long-Term Agent Memory
Abstract
Long-term-memory agents can retrieve relevant information that is inadmissible for the current request because it belongs to another principal, violates policy, or reflects an incompatible lifecycle state. Recall and final-answer accuracy do not reveal this: a route can appear safe by missing required evidence, while a correct answer may follow inadmissible prompt exposure. We introduce a retrieval-admissibility verification framework that assigns each memory-query pair one of three statuses (admissible, inadmissible, or unresolved), compares routes at matched required-evidence recall with bounds for unresolved cases, and tracks memory IDs through prompt exposure while linking exposure to target-level disclosure. We evaluate its stages on separate, non-pooled populations. A post-hoc top-20 reanalysis of frozen rankings from two public long-term-memory benchmarks, RHELM and MemOps, covers 3,767 queries. All released anchors lie within trusted query namespaces; with within-namespace scores unchanged, off-namespace filtering cannot lower their ranks. Top-20 anchor recall increases from 0.432 to 0.533, 80% recall feasibility from 0.237 to 0.311, and exact similarity evaluations decrease by 98.3%. In a frozen 72-case development diagnostic, a released-metadata reference preserves required evidence, whereas neither text-only verifier detects violations under the 1% required-anchor false-denial limit. Across 1,523 paired benchmark-native cases, namespace routing is associated with judged-accuracy gains of 0.053-0.068 across three readers; recall also changes, so this comparison is observational. In 16 controlled exposure scenarios, only one of four reader-specific 95% confidence intervals excludes zero for relevant-inadmissible literal disclosure (+0.156, 95% CI [0.031, 0.312]). Results motivate separate verification of candidate support, admissibility, prompt exposure, and answer disclosure.
cs.AI / 11 / 2610.07311
Understanding and Mitigating Inference-Time Overreliance Using Agentic Memory
Abstract
Agentic memory allows LLM agents to reuse past experience, yet retrieved memories can also distort inference even when they are benign, correctly stored, and appropriately retrieved. We study this failure mode, which we call memory over-reliance. Across benchmarks and memory architectures, we find that memory is useful when past experience transfers to the current task, but can become misleading when only part of the evidence transfers. Failures are strongest under partial query-memory overlap, a pattern further confirmed by controlled experiments thatvary the amount of overlapping evidence. Motivated by this finding, we propose MEMTRIM, a plug-and-play framework that indexes memory evidence at write time and controls its reuse at read time. MEMTRIM removes repeated or conflicting evidence while preserving useful memory-specific information, requires no retraining, and applies to both embedding-based and structured memory systems.Experiments show that MEMTRIM reduces memory overreliance while preserving the benefits of useful memory across models and memory settings.
cs.AI / 12 / 2610.07313
Rule-Based Languages for Neurosymbolic AI
Abstract
Logic programming is increasingly used as the symbolic component of neurosymbolic AI systems. We survey the main rule-based languages in this setting, namely Datalog, answer set, and probabilistic logic programs, along four axes: semantics, expressiveness, neural integration, and evaluation mechanism. We analyse over 50 recent systems and applications, comparing formalism usage across four research areas: databases and programming languages, machine learning, vision, and robotics. We provide a decision matrix mapping application scenarios to required features and close by outlining open problems.
cs.AI / 13 / 2610.07350
Trajectory-Retrieval Speculative Decoding: When Does a Model's Own History Help?
Abstract
Long chain-of-thought reasoning increases sequential decoding cost while creating a growing history of potentially reusable continuations. We investigate when this history supplies useful drafts and complements an existing drafter. Controlled source comparisons reveal trajectory-specific reuse, motivating our method Trajectory-Local Adaptive Retrieval (TLAR). TLAR retrieves approximately matched continuations from the current trajectory and uses recent verification outcomes to adapt retrieval activation and candidate width. TLAR combines retrieved continuations with model-generated drafts in a shared candidate tree, preserving the target model's output distribution through exact verification. Across code debugging, mathematics, and open-ended writing, our evaluation connects source reuse, incremental acceptance, and execution cost. Combining TLAR with strong retrieval baselines improves token acceptance under matched verification budgets and increases end-to-end throughput over the draft-model baseline. These findings support generated trajectories as runtime memory for adaptive inference.
cs.AI / 14 / 2610.07354
Evaluating Escalation Signals for LLM Routing: Targets, Controls, and Five Ways to Fool Yourself
Abstract
Deciding when to escalate a query from a small language model to a larger one requires a cheap signal that predicts, before the large model is called, whether escalating would help. Semantic entropy, originally developed to detect hallucinations, is a natural candidate: it measures how much a model's sampled answers disagree in meaning, and high disagreement often signals an unreliable answer. We test it across three benchmarks and two model families. On GSM8K, with a small/large pair about twelve times apart in size, semantic entropy reliably distinguishes the small model's mistakes (AUROC 0.871) and improves routed accuracy over random escalation by up to nine points at matched cost. An earlier strong-looking result on a synthetic benchmark proved misleading: a simple rule based only on question difficulty, with no model involved, matched semantic entropy almost exactly. This paper's main contribution is a set of checks that catch this before it is reported as real. We show that scoring a cheap, question-only difficulty estimate alongside any signal reveals whether the signal adds real information or just tracks how hard a question looks; that two reasonable definitions of "escalation worked" can produce very different results on the same data; that a benchmark can leave almost no room for any signal to beat simply always using the large model; and that the true cost of live sampling can make routing more expensive than calling the large model directly. For a cheaper alternative that reuses cached past outcomes, we show how to predict whether it will work on a new dataset -- confirmed by correctly forecasting a collapse from AUROC 0.908 to chance level (0.518) ahead of time. We offer these as a general checklist for evaluating escalation signals.
cs.AI / 15 / 2610.07359
Evaluate the Stack, Not the Layer: Do Deterministic and LLM Gates for Agent Actions Fail Independently?
Abstract
Runtime gates for agent tool calls are stacked on the assumption that their errors multiply. We test it on 1,119 labelled agent actions from three corpora, without an adaptive adversary. The stack has one deterministic rule layer and four LLM judges, three of them re-collected with the served model recorded on every call. We read each stack as a number of multiplication-equivalent layers, n_mult, with its floor under perfect coupling. Under the STRICT miss definition (escalation to a human scored as not stopped), any two judges compose to about 1.2 to 1.4 layers (φ median +0.430, 6 of 6 pairs significant, floors 1.02 to 1.17). The rule layer plus one judge composes to 1.86 to 2.09 layers (φ median +0.014, 0 of 4 significant, floors 1.01 to 1.09). Under PRIMARY (escalation scored as caught) the bands are 1.21 to 1.57 and 1.80 to 2.13. Intervals separate on the pooled data, point estimates split on each corpus, and a third-vendor judge lands in the judge band. Solo accuracy does not predict what a layer adds: a cloud rule pack lowers the rule layer's solo miss rate by 20% and adds no new joint coverage. The difficulty share of judge coupling is not identifiable: 31.8% to 61.8% depending on the probe and the miss definition. One judge tier was served by an unrequested model version in 50 of 112 batches, concentrated on the external corpus. That event overturned a pre-declared analysis rule, and the scoring of review verdicts reversed five conclusions. We report both.
cs.AI / 16 / 2610.07423
2d-fet-bench: from spatial reasoning to fet design on flakes
Abstract
Field-effect transistor (FET) layouts on exfoliated two-dimensional flakes are typically drawn by hand for each flake, placing contacts and gates to match its position and outline in optical micrographs. To our knowledge, no executable benchmark tests whether language-model agents can perform this flake-specific construction reliably. We introduce 2D-FET-Bench V2, a benchmark of 128 layout tasks built from microscopy-derived flake contours, including hole-containing flakes and multi-flake tasks. Each task supplies a textual device specification and contour coordinates. An agent generates typed polygon and path operations rendered to GDSII. A deterministic verifier checks geometric and structural requirements, and a separate integrity check verifies that the supplied contours remain unchanged. Scripted reference layouts pass all 128 tasks, showing that every task is solvable. We evaluate six models and seven workflow and scaffold variants of GPT5.6-Luna, with five attempts per task. The best-performing configuration in the six-model panel, GPT5.6-Luna with ReAct-3, passes 62.3% of attempts and solves 80.5% of tasks at least once (coverage) and 43.8% in all five attempts (consistency). ReAct-3 exceeds the one-pass Plan-and-Execute by 27.0 pass@1 points at 2.46 times the tokens. An expert audit of one sampled verifier-passing layout per covered task, across five ReAct-3 configurations, accepts 56.4% to 63.5% of them. The benchmark evaluates geometric and structural FET layout construction.
cs.AI / 17 / 2610.07428
Adaptive Gait Biofeedback With Participant-Held-Out Modeling and Participant-Specific Updating in Chronic Ankle Instability
Abstract
Adaptive gait biofeedback may support repeated practice in chronic ankle instability, but its evaluation must address model performance and human response. We evaluated a temporal convolutional classifier on protocol-defined, angle-derived GOOD/BAD gait-cycle labels using participant-held-out leave-one-subject-out (LOSO) cross-validation in 20 participants. Seven participants in the adaptive-intervention group completed nine sessions over three weeks, with one motion-capture recording analyzed per session. Models updated after failed sessions were compared offline with their parent models on the same-session validation subset used for candidate selection and the first subsequent adaptive-session recording. Frontal-plane ankle angle was compared between the adaptive group and 10 sequentially enrolled controls at Baseline, Post, and 7-day Retention. Across 20 held-out folds, mean fold-level area under the receiver operating characteristic curve (AUROC) was 0.948, sensitivity for angle-threshold-exceeding BAD cycles was 0.941, and specificity for angle-threshold-meeting GOOD cycles was 0.366. Mean BAD-class F1 was higher in candidate models by 0.187 on the same-session subset and 0.118 on the first subsequent recording. At Post, the adaptive group had a baseline-adjusted frontal-plane ankle angle 5.168 degrees lower than controls (95% confidence interval, 1.766-8.569 degrees lower); the Retention contrast was uncertain. These findings characterize population-model discrimination and offline participant-specific updating during repeated biofeedback use, alongside a nonrandomized Post frontal-plane ankle angle association. They do not establish independent clinical gait classification or a causal benefit of updating.
cs.AI / 18 / 2610.07459
Auditable Claims about AI Agents
Abstract
Organizations make claims about their AI agents: a person approves every external email, every action is logged, an evaluation shows the agent is safe to deploy. Article 12 of the EU AI Act requires high-risk systems to allow the automatic recording of events but does not say which records settle a given claim. The position is one sentence: to be checked, a claim about an agent must first name its policy, its scope, the records that would settle it, and who writes them. Adapting the preconditions of an assurance engagement, we call a claim auditable when these elements and a decision rule are fixed before any verdict and the records are obtainable. This extends the Policy Checkability dimension of our Auditable Agents framework from single actions to claims. Agents add three conditions: coverage by an independent record, authorization bound to each action's arguments, and completeness beyond integrity. Under an explicit model, we prove that support is impossible without each wherever its hypotheses hold. A claim-check table applies the method to six common claims, anchored in current NIST, IETF, and OWASP drafts. A worked case follows one claim through five evidence states. We close with a practice box and steps for operators, buyers, auditors, and standard setters.
cs.AI / 19 / 2610.07469
COMPASS: Finding Where Reasoning Lives in Language Models
Abstract
Explicitly eliciting reasoning substantially improves LLM performance. Existing approaches require a predefined characterization of reasoning, whether through CoT prompt design, contrastive CoT directions, or via SAE derived reasoning features. For mathematical reasoning with verifiable answers, we show that a much simpler signal suffices, which is the correctness of the model's own direct answer attempts. This signal yields a latent direction that elicits reasoning. This direction is decodable within the activations of most attention heads, but only a small subset of them can be effectively intervened. We introduce COMPASS, an inference-time steering method that identifies these heads using a logit-space attribution score and steers their activations along the correctness direction, requiring only per-head activation statistics. Across three model families and multiple math benchmarks, COMPASS outperforms the activation-steering baselines we compare against, improves GSM8K accuracy by 16 percentage points on average, and approaches CoT accuracy with 20-70\% fewer generated tokens. Interventions transfer without re-fitting to unseen benchmarks, and ablations show that both the correctness direction and the small set of heads carrying it are necessary, with the effect concentrated in remarkably few heads.
cs.AI / 20 / 2610.07497
Does Muon Need Fine-Grained Spectral Shaping?
Abstract
Muon combines current and past gradients into matrix momentum. For $M=UΣV^\top$, the idealized polar update $Q=UV^\top$ gives every singular direction the same weight. We refer to this as the flat profile. Several recent optimizers replace this flat profile with fine-grained spectral maps that give each direction its own gain. We ask how much of this spectral detail a Muon update needs. Our spectral diagnostics show that approximately $94$--$97\%$ of measured singular modes lie below an estimated noise edge, yet collectively align positively with a reference gradient. We introduce BulkBoost, a two-band spectral reweighting framework with fixed-rank and noise-calibrated variants. The latter uses split-minibatch gradient differences to calibrate a Marchenko--Pastur reference edge for Muon's Nesterov input, separating the bulk below the edge from the spikes above it. Both variants increase the bulk's relative weight through one shared gain while preserving the Frobenius norm of each matrix's unreweighted direction. For a fixed partition, our theory gives the first-order condition under which moving weight toward the bulk lowers the loss. It also quantifies the fraction of the maximal first-order improvement rate, over all per-mode reallocations, that two bands can capture. Across 30 continued-pretraining settings spanning Pythia-14M to 410M and six corpora, two-band reweighting is competitive with the fine-grained power-law profile of Freon and outperforms Spectra. Measured against Muon's flat profile, Freon reduces final loss by $0.022\%$ of the pre-adaptation loss on average, whereas the two-band variants achieve reductions of $0.073$--$0.147\%$. These observations suggest that useful departures from the flat profile are surprisingly low-dimensional: a single bulk-to-spike gain captures at least as much benefit as the fine-grained spectral profiles.
cs.AI / 21 / 2610.07505
MARS: Multi-resolution Adaptive Routing for Sequential Recommendation
Abstract
Long-history recommenders often compress each user's history into a compact, candidate-independent memory that is cached and reused to score large candidate pools. We show that real user histories exhibit multi-scale semantic structure, with short-lived intent, medium-term interests, and long-term preferences coexisting in one sequence, and that monolithic cached memories preserve these scales unevenly: linear probes recover recent and mid-range content far worse than long-range content. We call this failure mode \textit{temporal aliasing}. We propose \textbf{MARS}, a multi-resolution user memory that writes the full history into recurrent state tracks anchored to different half-lives, and a sparse routing reader that materializes compact seed memories by selecting the relevant temporal resolutions for each seed, preserving fixed-size candidate scoring. MARS outperforms strong baselines on three public datasets, with gains that grow with history length. Component-matched ablations with paired tests show that temporal diversity and selective routing each contribute beyond what hard-window memories or added capacity provide. The advantage of MARS over its interface-matched baseline also widens after within-user behavioral shifts, at about $1.02\times$ that baseline's warm-cache serving latency for $1{,}000$ candidates per user.
cs.AI / 22 / 2610.07509
On Open-Ended Information Seeking for Information Elicitation Agents
Abstract
Information elicitation is an open-ended information-seeking problem in which an interaction can unfold in many potentially valuable directions, requiring an elicitor to continually determine which information to pursue as new information emerges. In agentic elicitation, these decisions may be delegated to a foundation model, yet how model choice shapes the resulting information-seeking behavior remains understudied. We study how judgments about information value vary across LLMs and how these differences shape sequential information seeking. We first examine these judgments across 11 LLMs spanning multiple model families and parameter scales, using a shared set of information and elicitation objectives. We then develop a controlled elicitation simulation in which different models encounter the same information space and use the same selection rule, isolating these judgments from question generation and respondent behavior. Using this setting, we characterize the breadth-depth behavior that emerges from model-specific information-seeking preferences over the course of elicitation. We further examine how interaction history changes the evaluation and subsequent selection of prospective information. We test the robustness and boundaries of these findings through sensitivity analyses and ablations over the opportunities available to the elicitor, the response labels used to operationalize information-seeking preferences, the presence of interaction history, and whether redundancy is explicitly relevant to the assessment. The project code, data, and trajectory files are available at https://github.com/infosenselab/open-elicitation.
cs.AI / 23 / 2610.07514
From Local Evidence to Safety Verdicts: Causal Tracing in Vision-Language Models
Abstract
A vision-language model may need to combine an image with a prompt to recognize a safety risk that neither reveals alone. Where does this joint safety judgment become accessible inside the model? We introduce SSU-Bench, a dataset of matched safe and unsafe image-text combinations constructed using single-item prompt edits or image edits with annotated intended regions. Using three vision-language models, we transfer internal states between paired inputs and measure the resulting change in the safety verdict. Across models and both types of counterfactual, interventions at the changed input positions are effective in earlier decoder layers, while interventions at the final input token become effective later. Directions estimated from other examples produce similar late-layer effects. A linear readout of the final-token state also predicts the model's own verdict, including incorrect judgments, and cross-model comparisons reveal similarities in the patterns of counterfactual change. These findings identify a recurring transition in where interventions can influence a joint safety verdict and distinguish a readable model decision from a correct safety judgment.
cs.AI / 24 / 2610.07521
Grounding What Shapes the Plan: Rethinking Groundedness for Physical Intelligence in Autonomous Driving
Abstract
Driving models increasingly ground reasoning in causal relations, spatial structure, perceptual evidence, and predicted futures. These advances make reasoning more faithful to the driving scene, but leave a fundamental question unresolved: what should groundedness mean when the model ultimately outputs an action? Correctly grounded reasoning does not, by itself, ensure desirable driving outcomes. We introduce GroundAct, which starts from a simple premise: driving unfolds through physical entities and their interactions. Entities therefore become the unit of grounding; a lightweight reference token keeps each selected entity's continuous state addressable through symbolic reasoning; and only the referenced entities' interactions with the evolving proposal correct the plan. The result is an explicit path from what reasoning grounds to what the plan does, which we call grounded planning. To assess its practical value, we evaluate GroundAct in both open- and closed-loop settings. GroundAct shows strong open-loop planning across normal, out-of-distribution, and safety-critical scenarios, with closed-loop results extending this evidence to driving in simulation.
cs.AI / 25 / 2610.07556
Decoupled Multi-Agent Orchestration
Abstract
Learned orchestration can automatically construct effective language-model multi-agent systems, but existing approaches couple planning to fixed worker pools and train decomposition and collaboration from the same terminal outcome, limiting transfer and obscuring credit assignment. We introduce DeOrch, which separates worker-agnostic planning from concrete worker selection. Its two-stage planner first decomposes the task without worker information, then chooses collaboration operations using compact, worker-identity-free matchability feedback from the pool, enabling conditional credit assignment to decomposition and collaboration decisions. A lightweight matcher estimates worker suitability from behavior on a fixed probe set and adapts online with a contextual bandit, allowing new workers to be incorporated without retraining the planner or matcher. Across diverse in- and out-of-distribution tasks, DeOrch outperforms prior automatic MAS orchestration methods with fewer worker calls than competing learned orchestrators, remains effective when transferred to an entirely unseen worker pool without retraining, and shows consistent gains from both components.
cs.AI / 26 / 2610.07560
Navigating Route Latent Space for Synthesizable Molecular Design
Abstract
Goal-directed molecular design has advanced rapidly, yet a substantial proportion of designed molecules remain difficult to synthesize in practice, limiting their real-world utility. Prior synthesizability-aware methods either project generated molecules back to synthesizable analogs that deviate from the intended target, or optimize directly in discrete synthesis spaces that lack a continuous landscape for efficient search. We argue that this limitation mainly comes from the search space rather than the optimizer. To address this, we propose RouteFlow, a framework that reformulates synthesizable molecular design as a search over a continuous route latent space, where each latent maps back to a complete synthesis route and synthesizability is inherently preserved. To navigate this space, we adopt reward-guided flow matching as an efficient sampler that steers toward high-property regions. Since reward optimization may push latents off the manifold of real synthesis routes, where decoding becomes unreliable, we further introduce a cycle-consistency mechanism to stabilize fine-tuning. Across 16 optimization tasks from Therapeutic Data Commons, RouteFlow achieves the best sample efficiency among synthesizability-aware baselines, with the best synthetic accessibility and the highest retrosynthesis success rate. Our results also confirm that the proposed cycle-consistency reliably keeps optimization on-manifold while improving target properties, supporting effective synthesizable molecular discovery.
cs.AI / 27 / 2610.07570
Unanimously Wrong: Certified Abstention from How Medical LLM Consensus Forms
Abstract
In clinical practice, agreement among independent experts is treated as evidence of reliability, and multi-round consensus has become a core mechanism of agentic medical question-answering systems. When such a system must decide whether to trust its own answer, the prevailing signal is again agreement, now among the sampled answers. But agreement is a fragile proxy for correctness. A system can be unanimously wrong, returning the same incorrect answer on every sample, and on these questions agreement-based signals carry no information. The cause is that these signals read only the final state of the consensus and discard how it was reached. Agreement that was reached by resolving disagreement with evidence looks identical, at the end, to agreement that was present from the first sample because every sample shares one misconception. ProbeGuard is a certified abstention framework that bases the abstention decision on how the consensus formed. Process features trace agreement trajectories, minority persistence, and retrieval saturation. For unanimous votes, rationale semantic entropy checks whether the reasons behind the vote cohere, and an active probe retrieves counter-evidence and measures whether the consensus survives. A stratified Learn-then-Test calibration then converts these scores into a distribution-free bound on selective risk. We evaluate ProbeGuard on three medical QA benchmarks and a hard-frontier reference, with a published multi-round agentic RAG substrate, against six abstention baselines. On MedQA, 13.4% of unanimous votes are wrong, and no agreement-based signal can flag them. Process signals raise the discrimination of correct from incorrect consensus from chance to 0.696 AUROC. The certified rule answers six in ten unanimous-layer questions at an observed selective risk of 9.0%, and nine in ten once in-domain calibration data accumulate.
cs.AI / 28 / 2610.07578
Cooperating with Future Collaborators: Multi-Agent RL under Staggered Participation
Abstract
In cooperative Multi-Agent Reinforcement Learning (MARL), agents are often trained under concurrent participation, while in many tasks some agents act earlier and leave task-relevant information that becomes useful to agents participating later. We study this setting as staggered participation (SP), which introduces a cross-time, cross-agent learning dependency because an early action may affect the return through the information it provides and the later policy that uses it. Learning under SP therefore requires both identifying what information is useful for future decisions and learning how later agents should use it. We propose Staggered Participation Learning (SPL), a training-time augmentation that addresses these two parts with prospective acquisition supervision for earlier agents and outcome-supervised receiver learning for later agents. We evaluate SPL across multiple policy-based MARL backbones, environments, and staggered-participation patterns. Across 60 MPE/RWARE backbone setting comparisons, SPL achieves higher observed mean task completion in every case, with an average difference of 14.1%. The gains also extend to eight-agent teams and a physics-based UAV-UGV environment in Isaac Lab, providing evidence across algorithmic, temporal, and embodied settings.
cs.AI / 29 / 2610.07582
Representation Bias, Correction Transfer, and Resolution Sensitivity in Three-Dimensional Mitochondrial Morphometry
Abstract
Quantitative imaging pipelines can produce precise but systematically different measurements of the same object. We present an empirical reliability assessment of three-dimensional mitochondrial morphometry that connects representation bias, a controlled processing intervention, correction transfer, and resolution sensitivity. Using 2,720 development objects from the 3D Mitochondria Shape Library for Optical Microscopy, we find that occupancy-derived volumes exceed reference mesh volumes by 3.665% on average despite an intraclass correlation coefficient of 0.994. Boundary analysis identifies an outward label displacement of 0.00304 normalized units. In a controlled label-pipeline reimplementation, removing the depth offset reduces volume error in all 55 analyzed objects by a mean of 1.57 percentage points, approximately 45% of mean reproduced inflation; the source of the remainder is not isolated. A frozen regression using occupancy-derived features reduces median absolute percentage error from 3.481% to 0.664% in 2,728 previously unused objects from the same resource. However, its calibrated error bound covers only 92.1% overall and 49.2% in a low-occupancy subgroup, demonstrating that accuracy and uncertainty transfer must be evaluated separately. In 550 rat-cortex objects from the MitoEM resource, coarsening in-plane spacing from 8 to 24 nanometers changes median surface area by minus 10.60% and sphericity by plus 11.76%, despite a rank correlation of 0.994. These results provide quantitative checks for distinguishing processing-induced descriptor changes from candidate biological differences, without establishing biological invariance or cross-source correction transfer.
cs.AI / 30 / 2610.07588
Personal-Agent Mediated Recommendation with Cross-Platform User History
Abstract
Modern recommendation is shifting from platform-centric personalization toward user-governed personalization, where a personal LLM agent can act on the user's behalf across services. We formalize this emerging paradigm as Personal-Agent Mediated Recommendation: a platform recommender ranks a candidate set using platform-local information, and a personal agent uses user-authorized cross-platform history to mediate the resulting ranking and produce the final top-K slate. Such mediation is nontrivial: the platform ranking can encode strong population evidence that the personal agent cannot observe, so effective mediation must therefore balance beneficial rescues against harmful overrides. To study this trade-off, we introduce MediateRec, a benchmark that includes scalable proxy cross-platform environments and a real cross-platform test under a controlled platform-agent information boundary. To train the agent to use cross-platform history effectively, we further propose Personal Attribution Mediation Optimization (PAMO), which counterfactually masks that history to estimate personal mediation support and reallocates rank-aware advantage mass under a platform-relative value floor. We theoretically prove that PAMO preserves cutoff-level advantage mass and is locally optimal among first-order reallocations that preserve this mass without lowering average platform-relative value. Experiments on MediateRec show that personal-agent mediation enables meaningful platform corrections, yet even strong proprietary LLMs introduce non-negligible harmful overrides. PAMO consistently improves over matched outcome-only RL across seen and unseen target platforms and on the real cross-platform test, while achieving a better rescue-harm balance.
cs.AI / 31 / 2610.07592
LSC-DPO: Learning-Signal-Controlled Direct Preference Optimization
Abstract
Direct Preference Optimization (DPO) has become a standard reward-model-free approach for aligning language models with preference data. However, as the scaled preference margin grows during training, the logistic DPO loss becomes progressively less sensitive to further changes. We study DPO from a loss-level geometric perspective and identify the sigmoid factor as a learning signal that characterizes the local sensitivity of the objective. Based on this view, we propose Learning-Signal-Controlled Direct Preference Optimization (LSC-DPO), which dynamically regulates the learning signal near a target regime. A log-space analysis establishes conditions for stable tracking of the target learning-signal regime. Experiments on AlpacaEval 2, MT-Bench, and Anthropic-HH show that LSC-DPO consistently improves over DPO and strong preference-optimization baselines. We further find that different coefficient initializations induce distinct transient learning-signal trajectories even when their later signal levels become similar. Based on this observation, we derive a signal-budget compensation rule that adjusts the target learning signal to compensate for these transient differences. The resulting compensation substantially reduces performance variation across coefficient initializations.
cs.AI / 32 / 2610.07601
Beyond Scalar IoU: Structured Verification from Rollout Groups for Video Temporal Grounding
Abstract
Reinforcement learning with verifiable rewards (RLVR) provides a natural framework for adapting pretrained models to video temporal grounding, where generated temporal intervals can be scored directly against ground truth intervals. Yet existing overlap verifiers typically score each rollout independently, leaving the joint structure of the rollout group unused. We introduce SUTURE, which conditions verification on the rollout group and exploits its structure at two complementary scales: disagreement across rollouts controls how strongly the target is reweighted, while coverage at each position determines where reward mass is redistributed. We show that the resulting verifier admits an exact decomposition into the standard IoU term and a covariance correction determined by the rollout group. A local gradient diagnostic finds a preference for responses covering relatively less supported target regions in the analyzed groups. Across five temporal grounding benchmarks, SUTURE improves grounding performance at every reported IoU threshold. Its trained policy also shows less video-start anchoring in reasoning traces: for later events, the first temporal mention more often overlaps the annotated target. Together, these results show that the joint structure of a rollout group can support a more informative temporal verifier.
cs.AI / 33 / 2610.07614
BioStudyBench: Evaluating Agents on Post-Cutoff Biomedical Studies
Abstract
We evaluate whether AI agents can match the reported findings of published biomedical studies using public data. Existing evaluations do not consistently separate analysis from prior knowledge or retrieval of the published answer. We introduce BioStudyBench, a benchmark of 25 long-horizon analysis tasks drawn from studies first published between July and September 2026, after the developer-reported knowledge cutoffs of the models we evaluate, semi-automatically filtered down from 404,019 PubMed records. In each task, the agent receives a neutral research question but no data files, so it must find and download the relevant public data, search the literature through tools that return only records dated before its cutoff, and report findings through data analysis. To measure gains over prior knowledge, we run every task both with and without access to data and tools. Across eight models, access to data and tools raises the pass rate by 47 percentage points on average over the no-data baseline. Open-weight models across sizes trail closed-weight models, with the best open-weight model passing 81.3% of tasks against 94.7% for the best closed-weight model.
cs.AI / 34 / 2610.07627
Learning to Outgrow a Theory: Experimental Discovery Beyond the Initial Hypothesis Space
Abstract
Scientific discovery systems typically optimize experiments within a fixed hypothesis space. This creates a failure mode when all available candidates omit the same missing mechanism: candidate disagreement can collapse even while the model class is systematically wrong. We formulate experimental model-class revision, in which a discovery policy jointly proposes a structural edit and a diagnostic experiment that tests whether that edit is necessary. The method couples a class-level distinguishability objective, in which one shared parameterization must explain all selected experiments, with anytime-valid sequential evidence that triggers structural revision only after the current class is rejected. On 400 held-out controlled dynamical environments, the joint policy reaches 89.5% exact recovery with a budget of 32 real experiments, improving the strongest matched baseline by 10.0 percentage points while requiring fewer executed experiments and candidate fits. The learned revision-experiment pairing transfers across unseen mechanism combinations, held-out but expressible primitives, parameter extrapolation, and shifted experiment costs; when the true mechanism is outside the edit grammar, it detects library insufficiency in 88% of cases with a 5.5% false-support rate. Revision gains also transfer to ODEBench and ODEBase model-library tasks, as well as DiscoverPhysics worlds. These results support a view of scientific discovery in which deciding what mechanisms a theory should make expressible and where to collect evidence are treated as a single sequential decision problem.
cs.AI / 35 / 2610.07634
Measuring climate backlash in Twitter and Reddit archives: Lexical definitions, recorded responses and participant turnover
Abstract
Social media archives are often used to study resistance to climate action, but words, response counters and observed participants do not measure the same social process. We examine four supplied Twitter and Reddit archives by processing all registered files without sampling and applying transparent, non-exclusive lexical rules. The study links frame co-occurrence to source-specific temporal and response models, then separates event-period changes among returning authors from participant turnover. Renewable-energy terms accompany cost-related language on Reddit, yet narrower backlash phrases sharply reduce cross-source contrasts and reverse the sign of the Paris Agreement contrast in submissions. Cross-discourse history does not improve eligible primary-context forecasts. Denial/hoax terms are associated with higher recorded Twitter likes, whereas Reddit response associations depend on frame, outcome and author specification. Around the 2019 global climate strike, returning-author expression and participant turnover both contribute to increased protest-language shares. An archive endpoint prevents the corresponding Climate Twitter migration inference. Most crossed-cluster estimates lack released intervals, and joint author/month response covariance estimates fail, restricting formal inference. These results show how operational definitions, platform-specific response fields and observation boundaries shape what can be claimed about climate backlash. The contribution is an archive-based account of these measurement consequences, rather than a measure of individual opposition, persuasion or advocacy-induced backlash.
cs.AI / 36 / 2610.07638
Learning Explainable Representations of Complex Game-playing Strategies
Abstract
As part of learning to play complex games, human players develop develop abstractions for concepts and strategies of gameplay consistent with game rules to improve their performance. These concepts are applied to explain other players' actions, and to inform their own actions in-game. Understanding other players' strategies is a crucial part of such improvement, but requires time and effort. In this paper, we propose a strategy similar to human cognition for training RL agents to synthesize learned strategies and policies as executable procedures based on sequences of gameplay actions. We present methods to automatically learn such programs to play chess and to solve tasks in a grid-based environment. We show that the learned strategies produce effective actions, and can be learned from gameplay data.
cs.AI / 37 / 2610.07640
Towards the Automatic Synthesis of Interpretable Chess Tactics
Abstract
State-of-the-art reinforcement learning agents are capable of outperforming human experts at games like chess, Go and StarCraft II. These agents do not simply take advantage of their digital hardware in being able to react and calculate faster than humans, but employ better strategies that lead to more victories. Interpreting these strategies would give human players valuable insight into how to improve their play. In this preliminary work, we propose a symbolic sub-policy model for playing chess. Inspired by chess tactics, our model attempts to incorporate domain knowledge to improve interpretability. We adapt patterns learned by an inductive logic programming system called PAL to derive our model. We contribute a divergence metric to evaluate our model against a random baseline, and find a set of tactics that is able to suggest moves of similar playing strength to a human beginner. Finally, we propose a computational evaluation scheme for the model by augmenting an off-the-shelf engine with it.
cs.AI / 38 / 2610.07646
Matching Object or Relation? Tracing Abstract Reasoning Inside VLMs
Abstract
Vision Language Models (VLMs) excel on visual benchmarks but fail systematically on tasks requiring abstract reasoning. Existing benchmarks document this failure but cannot say \emph{why} it happens or which cognitive capability is missing. We close this gap by adopting the Relational Match-to-Sample (RMTS) paradigm from comparative and developmental psychology and pairing it with a mechanistic analysis of the model's internals. On a parametrically controlled stimulus set evaluated across frontier API models (GPT, Claude, Gemini) and three open-source families (Qwen3.5, Gemma-4, InternVL3), we identify four levers that shift VLMs toward the relational match---capability tier, model scale, the number of objects per scene, and the absence of per-object stimulus noise---together producing a developmental-like trajectory that mirrors the human \emph{relational shift}. Opening up the model, a per-layer representational similarity analysis and a causal mediation analysis reveal that VLM abstract reasoning is implemented by two competing circuits: an early circuit that organises images by their surface object features, and a late circuit that organises them by their abstract relation. Extending the analysis to ARC-AGI-1, we find that ablating the relational heads identified on RMTS degrades performance more than ablating random heads, indicating that the relational circuit is recruited beyond our controlled stimuli. We hope this mechanism-level view serves as a step toward understanding how abstract reasoning is implemented in VLMs.
cs.AI / 39 / 2610.07657
Where Rules End and Judges Begin: Measuring the Judgment Boundary in Multi-Agent Systems Security
Abstract
LLM-based multi-agent systems (MAS) engage tools, share memory, and delegate tasks, often encountering adversarial content. Current defenses for MAS are typically evaluated in isolation, focusing on one attack type at a time, which can lead to costly and hard-to-audit outcomes. This study organizes defenses into five principles, implementing them as DEFER1 (DEterministic-First Enforcement with Residual judgment), which includes a cascade of 28 checks that blocks what it can and refers the rest to a panel of four judges. In independent testing across four domains, attack success rates drop from about 30.0% to approximately 3.0%, with 78% of blocked attacks handled by deterministic checks. Only a quarter of proposals reach the judges in the security-operations domain, illustrating that the rules provide security for attacks violating clear policies, while judges manage those that only misrepresent intent. Both systems have weaknesses, such as a risk-score approval gate that inaccurately approves most attack proposals but few legitimate ones, highlighting the challenges in assessing threats accurately.
cs.AI / 40 / 2610.07675
EIO-Agents: The Missing Semantic Layer for AI Agent Evaluation
Abstract
AI agents are entering production in increasingly consequential environments without a shared semantic standard for what their evaluations actually mean. Scores, traces, judge outputs, and multi juror findings are increasingly used to justify readiness and release decisions, yet they often do not specify what evidence supports a claim, what that evidence can establish, or how the claim leads to a decision. We introduce EIO-Agents, an open specification for interoperable AI agent evaluation built on two layers. The Evaluation Intelligence Ontology (EIO) provides the semantic layer through typed evidence, versioned behavioral predicates, evidence contracts, claims, witness rules, proof status, recurrence, and computable derivations for metrics, findings, controls, and PASS, REVIEW, or BLOCK decisions. The Portable Evaluation Record (PER) provides the system of record: a canonical, content addressed representation of one evaluation that preserves the evidence to decision chain and can be re derived, explained, and verified. Scores summarize, juries interpret, and traces record, but none of them define what the evidence means or what it can prove. EIO provides that missing semantic contract, while PER preserves the resulting evaluation as a portable and verifiable system of record. As AI agents assume greater operational responsibility, evaluation must become more than a collection of scores and verdicts; it must become an accountable artifact whose meaning, evidence, limitations, and decisions can be independently checked.
cs.AI / 41 / 2610.07686
BluffJAX: Adversarial Imperfect Information Games in JAX
Abstract
We introduce BluffJAX: an open-source suite of adversarial imperfect information games in JAX. We provide canonical implementations of games designed for high simulation throughputs and parallelization on GPU accelerators. Our suite consists of well-studied benchmarks such as Texas Hold'Em Poker and Kuhn Poker, as well as games that have not been previously studied in reinforcement learning research, such as Bluff, Stud Poker, and Kemps. We hope that implementing a variety of game mechanics and difficulties will introduce new challenges and foster novel research directions in game-theoretic methods for RL. We benchmark the throughput performance and memory usage of our environments in single and multi-GPU settings, demonstrating scaling of up to hundreds of millions of samples per second, and motivating the usage of BluffJAX over related GPU and CPU-based libraries. We benchmark reinforcement learning, tree search, and game-solving algorithms in JAX in order to provide users with baseline results and facilitate future comparisons.
cs.AI / 42 / 2610.07701
On the Boundary of Admission Gates: An Injected-Truth Study of Falsification-First Selection in Quantitative Strategy Research
Abstract
Strategy research conflates two problems: finding a profitable rule, and establishing that the finding is not search luck. The latter calls for admission gates -- statistical criteria that must be satisfied before a conclusion is adopted -- yet whether gates work, and at what cost, remains untested. We introduce an injected-truth protocol with a random-admission control that adopts at the same rate as the gate; only if the gate beats this control does it carry information rather than merely raise a threshold. Across synthetic and real-calibrated panels, gates eliminate false discoveries in the weak-signal regime but cut adoption to 1--7%, and add nothing when signals are strong. Most importantly, criteria computed on absolute rather than excess returns silently reject every candidate, including true signals. Keywords: multiple testing, backtest overfitting, strategy admission, injected-truth validation, excess returns, false discovery rate
cs.AI / 43 / 2610.07707
AgentMemGate: Addressing Speculation Contamination in Conversational Assistant Memory
Abstract
Conversational AI assistants with long-term memory extract facts from user messages into a store consulted in later conversations. A stated plan can enter that store as fact: a user who might move to Seattle may be recorded as already living there. We call this speculation contamination. Final-state memory benchmarks miss this error because they do not probe intermediate state and include few unresolved speculations. We present AgentMemGate, a write-time gate for profile-store memory that classifies extracted statements as speculation, completed event, correction, or other. Speculations remain outside memory, with conditions governing later promotion or deletion. We also contribute a dataset of multi-session conversations in which plans are confirmed, abandoned, or left unresolved. On our 147-conversation held-out set, Mem0 and Graphiti assert unresolved plans as current state for 35.2% and 27.3% of pending plans. On the core benchmark, AgentMemGate eliminates all observed contamination relative to the identical ungated pipeline (87.5% to zero for the most exposed extraction style) and raises task accuracy from 65% to 95%. On the harder held-out set, gated contamination is 3.4% to 5.7% and task accuracy rises by 9 to 13 percentage points. Our analysis identifies field matching as the main remaining bottleneck: realistic speculations often match no profile field and never reach the gate. We release our datasets, prompts, and evaluation code.
cs.AI / 44 / 2610.07708
Evidence Before Sampling: Interpretable Implicit Negative Candidate Discovery for Recommendation
Abstract
Recommender systems learn from observed user-item interactions, but explicit negative feedback is often unavailable. Since deep learning models require negative signals for training, negative sampling methods typically treat selected unobserved interactions as negatives. However, a missing interaction does not explain why a user is uninterested in an item or whether there is sufficient evidence to label it negative. This is especially important in business recommendation, where negative signals should be interpretable and aligned with business objectives. We formulate implicit negative candidate discovery to identify unobserved interactions supported by observed customer behavior. We encode these patterns as symbolic rules, score them based on support, informativeness, and product relevance, and rank the retained rules by evidence. An LLM then interprets the retained rules using business objectives and domain knowledge; the interpretations are combined with the statistical evidence in the final report. We evaluate our method in an industrial B2B setting and across five public recommendation datasets. Candidate-quality evaluations in the industrial setting and three public datasets show higher precision than the evaluated baselines, while symbolic selection improves downstream test PR-AUC by 12.5% over random selection with four negatives per positive example in the industrial task. Our results show that negative candidate validity can be evaluated separately from downstream recommendation performance. This distinction enables evidence-based, business-aligned, and explainable negative selection, improving both interpretability and model training in sparse, skewed, real-world recommendation settings.
cs.AI / 45 / 2610.07725
PERSIST: Who-What-When Memory Across Sessions for Full-Duplex Spoken Dialogue
Abstract
Modern voice assistants may be shared by multiple users and should be able to answer questions about earlier conversations such as "When did I originally plan to leave?" or adapt their behavior to individual users based on past interactions. This requires more than retrieving a topically similar passage: the assistant must identify the current speaker, recover the relevant past state, and distinguish it from later revisions. We present PERSIST, a persistent memory system for multi-session, multi-speaker spoken dialogue that explicitly models Who, What, and When. PERSIST structures cross-session histories into readable event records and retrieves them with a 3W joint scoring mechanism that combines semantic content, acoustic speaker identity, and temporal state. For real-time full-duplex interaction, PERSIST further reuses intermediate representations from the dialogue backbone, avoiding query-audio re-encoding and reducing retrieval latency from 578.42 ms to 7.03 ms. We also introduce SpokenTrace, a diagnostic benchmark that factorizes evaluation along memory tasks and speaker-query types, exposing failures in recall, speaker attribution, and temporal-state tracking. On SpokenTrace, PERSIST achieves 85.08% end-to-end task accuracy and improves all-support EM@3 from 49.01% with BGE-large to 82.10%.
cs.AI / 46 / 2610.07751
How Well Do LLMs Reason with Noisy Evidence? An Active Visual Reasoning Benchmark
Abstract
Real-world reasoning rarely reduces to static question answering: agents must actively gather information from tools and sensors that are often noisy and unreliable. Yet most existing active reasoning benchmarks assume that environmental feedback is trustworthy, or introduce noise without exposing an explicit, calibrated uncertainty signal, leaving open how LLMs should reason when the evidence itself is uncertain. We introduce VisualNoiseQA, a novel benchmark for active reasoning under noisy visual feedback. A text-only LLM must solve VQA problems by iteratively querying a fixed, off-the-shelf VLM treated as a stochastic visual sensor. For each query, we draw multiple samples and expose an empirical uncertainty signal via self-consistency, enabling the reasoner to probe from different angles and decide what to ask next and when to stop. Our construction is automatic and scalable: starting from diverse VQA sources and two noisy VLMs, we retain only questions where the sensor is inconsistent yet human-solvable. We evaluate multiple LLM reasoners on 1,000 instances spanning perception, chart understanding, and knowledge-intensive reasoning. VisualNoiseQA thus provides a controlled playground to study how different LLMs exploit uncertainty signals for robust reasoning.
cs.AI / 47 / 2610.07763
ST-Bench: A Spatial-Temporal Benchmark for Multi-Agent System Generation on Scientific Research Tasks
Abstract
The rapid progress of LLM-based multi-agent systems (MAS) has shown that they largely outperform single agents on coding, math, and QA tasks, where executable tests provide a binary success signal. Whether this advantage transfers to real scientific data analysis remains untested. We introduce ST-Bench, a benchmark designed to answer two questions: whether MAS outperform single agents on complex scientific data analysis tasks, and if so, by how much and at what additional cost. ST-Bench contains 100 data science tasks adapted from published Earth science studies across hydrology, agriculture, and wetland methane research, expanded into 2,067 queries grounded in additional published studies and validated by domain experts. Using ST-Bench, we evaluate five recent MAS generation methods under two training protocols, against single-agent baselines on the same GPT-5 backbone. Nine of the ten MAS configurations exceed the cheapest single-agent baseline, with the strongest reaching nearly three times its composite score. This gain is primarily attributable to coverage: trained workflows produce realistic numerical metrics on a larger fraction of queries, while the quality of those metrics, conditional on producing realistic output, is comparable to that of the single-agent baseline. The strongest configuration requires approximately four times the single-agent inference time, whereas a more economical workflow captures the majority of the benefit at less than twice the cost. MAS specialization confers measurable benefit on scientific data analysis, but the benefit is conditional rather than universal.
cs.AI / 48 / 2610.07766
OTel: Open Telco AI Datasets, Benchmarks, and Models
Abstract
We present Open Telco (OTel), an open telecom AI resource that releases derived telecom datasets for retrieval, reranking, instruction tuning, and safety/abstention, together with 30 full-parameter post-trained baselines spanning 10 embedding models, 3 rerankers, and 17 language models. The community has already engaged substantially with the resource: as of May 3, 2026, the released models have been downloaded over 16 million times and the project has received 157+ pieces of media coverage worldwide. Building on prior open telecom datasets and benchmarks, OTel provides documented telecom data sources, held-out evaluation partitions, trained embedding models, rerankers, context-grounded LLMs, and safety/abstention data in one unified resource. Each baseline starts from an open-weight model and is post-trained on OTel-derived data using an open training recipe, then evaluated on held-out OTel evaluation partitions. OTel post-training improves performance across all three model families: embedding retrieval reaches 93.1% NDCG@10, reranking reaches 0.947 MRR@10, and language-model correctness reaches 87.8%. We release OTel as a reproducible starting point and invite the community to expand the data, improve embedding and reranking models, and build stronger context-grounded telecom LLMs.
cs.AI / 49 / 2610.07781
Quantization Effects on Tool-Failure Recovery Vary Across Prompts and Evaluation Designs
Abstract
Post-training quantization reduces the cost of deploying language-model agents, but its effect on recovery from temporary tool failures can depend on how recovery is evaluated. We compare 8-bit and 4-bit variants of Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct on twenty deterministic tool-use tasks and five prompts. The 8-bit-4-bit recovery comparison changes direction across prompts and evaluation targets. On tasks that both variants complete without faults under the same prompt, the difference ranges from 0 to +20.2 percentage points for Llama and from -50.0 to +35.0 points for Qwen. Full-pipeline point estimates favor 8-bit Llama under all five prompts, whereas the Qwen comparison changes direction across prompts. The evaluation target can also reverse the result. For Llama under one prompt, scoring each variant only on its own clean-passing tasks favors 4-bit by 17.5 points; scoring the same tasks for both variants gives no difference, while scoring the full pipeline favors 8-bit by 28.3 points. Executor leniency is a third such choice. Rescoring the same logs with strict output parsing, which 8-bit Llama violates far more often than 4-bit Llama under that prompt, turns that +28.3 into -15.0 while leaving Qwen essentially unchanged. These findings show that one prompt, one screened task set, and one scoring policy do not establish a stable conclusion about quantized-agent robustness. Evaluations should compare variants on matched tasks, report full-pipeline success for deployment decisions, state the scoring policy, and quantify uncertainty across tasks rather than injected fault sites.
cs.AI / 50 / 2610.07782
Persistent Memory in Multi-Agent LLM Inference: What It Costs, What It Buys, and When You Can Tell
Abstract
Decomposing long-context inference across cooperating agents bounds the active KV cache per call rather than total evidence, which matters when KV-cache memory binds. Many such systems add a persistent tier storing and recalling reasoning traces, usually validated by an ablation reporting an accuracy gain. We measure both on one three-tier agent architecture. Decomposition delivers: peak KV working set of 14.3 MiB per query against 35.5 and 35.3 MiB for single-pass and retrieval-augmented baselines. The persistent tier does not: across eight controlled dataset pairs at n=100 per arm it costs +0.368 MiB [+0.167, +0.590] of peak cache and produces no detectable accuracy change (+0.015, 95% CI [-0.011, +0.046]). We argue the null is structural: single-question benchmarks supply each item with its own evidence and score it independently, and correctness requires resetting stored traces between conditions, so recall has nothing informative to retrieve. Reaching it took four measurement corrections -- three inflating the apparent benefit, the fourth making an effect that size look resolvable -- none visible in the results table. We give the conditions an agent-memory ablation must satisfy and detection procedures that need no knowledge of the specific defect.
cs.AI / 51 / 2610.07785
Attacca: Goal-Directed Control under State Continuity for Long-Horizon Embodied Agents
Abstract
A central capability of embodied agents is to accomplish complex objectives through sequences of interdependent tasks. Yet existing visual goal-conditioned policies underlying these agents are typically evaluated on isolated interactions where the target is already visible, and thus do not capture the conditions that arise during continuous long-horizon task execution. In such settings, each task begins from the state left by the previous one: the agent may end at a different position and orientation, the world may have been modified, and the next interaction target may lie outside the current field of view. As a result, agents relying on such policies may struggle to proceed to the next task when they cannot ground their target in the current observation. To address this challenge, we propose Attacca, a new approach that trains visual goal-conditioned policies on complete search-to-interact trajectories using goal images decoupled from the execution environment. Attacca uses context-decoupled goal sampling to pair each demonstration with a class-compatible masked goal image from another world, removing direct scene and pose correspondence. It learns dense current-view grounding through a target-mask prediction head, providing auxiliary supervision beyond action imitation. We further introduce behavioral-phase conditioning that teaches the policy to distinguish Search, Approach, and Interact stages and adapt its control as execution progresses. We evaluate Attacca on multiple short- and long-horizon embodied tasks in Minecraft. Our method achieves 39.0-47.5% clean success, improving over the strongest baseline by 1.7-2.4x. On long-horizon tasks, it attains 54%, 30%, and 28% completion, yielding up to a 7x improvement.
cs.AI / 52 / 2610.07803
ThinkFuse: Trajectory-Aware Test-Time Fusion for Small Reasoning Models
Abstract
Small reasoning models (SRMs) have shown strong performance on complex reasoning tasks by generating extended chain-of-thought trajectories, but they often fail to recover once their reasoning enters an erroneous path. Existing test-time fusion methods rely on local fusion signals to determine when to trigger fusion, which can be misled by transient uncertainty fluctuations and may reinforce unstable reasoning trajectories. We propose ThinkFuse, a training-free test-time fusion framework that selectively intervenes in unreliable reasoning segments. ThinkFuse compares segment-level uncertainty shifts with trajectory-level uncertainty trends to identify unstable reasoning points and fuse auxiliary reasoning paths into the primary model's trajectory. Extensive experiments demonstrate that ThinkFuse outperforms baselines on mathematical and knowledge-intensive reasoning benchmarks, with consistent gains across model-family combinations, and remains robust with a smaller primary model. Our analysis shows that ThinkFuse requires fewer fusion triggers and generates fewer tokens, highlighting the efficiency of selective triggering. Our code is available at https://github.com/js-lee-AI/ThinkFuse.
cs.AI / 53 / 2610.07816
Do I Need the Cloud? Uncertainty-Aware Step-Level Handoff for Small Language Model Agents
Abstract
Small language models (SLMs) are attractive as local agent controllers because they reduce remote inference, latency, and deployment footprint, yet structured tool errors can cause an agent step to fail. Existing routers typically select a model once per query. However, agents expose sequential decision points whose difficulty dynamically changes based on intermediate observations. We propose STEPGATE, an uncertainty-aware handoff framework that scores each local SLM action and selectively escalates challenging steps to a stronger model. On a 52-task held-out single-step BFCL-derived test split, the Qwen2.5-1.5B/7B pair attains 82.7% task success with 30.8% escalation, versus 67.3% local-only and 75.4% random escalation (which uses 33.8% escalation). In a separate multi-turn evaluation, STEPGATE achieves 69.0% trajectory success and 84.0% action success using only 30.0% cloud actions, compared with 48.0%/70.5% local-only, 60.0%/78.2% random escalation, and 57.0%/77.1% query-level routing (strong-only achieves 82.0% trajectory success at 100% cloud actions). These results suggest that step-level escalation recovers a large share of the performance gap to the stronger Qwen2.5-7B backend at a matched cloud-action rate while transmitting fewer tokens remotely. However, our evaluation is limited to one model family, a single stronger backend, and scripted tasks. Furthermore, the test sets are small, multi-turn comparisons rely on paired intervals and statistical tests, and our risk tiers serve as research annotations rather than formal safety guarantees.
cs.AI / 54 / 2610.07829
Agentic Semantic Sensing for Resource-Adaptive AI-RAN
Abstract
Semantic sensing (SemS) acquires task-relevant information rather than reconstructing complete physical information. Existing SemS formulations typically operate open loop: sensing configurations and observation schedules are fixed before inference and cannot respond to evolving task-level evidence. We propose Agentic SemS, a closed-loop framework for AI-enabled radio access networks (AI-RANs) that controls sensing within a communication-feasible profile set. A profile-conditioned causal Transformer updates the semantic belief from streaming observations, while key-value caching enables efficient state updates across profile changes without repeatedly processing the complete history. A semantic utility network estimates the task-level benefit of acquiring the next observation block under each feasible profile after accounting for sensing cost. The resulting continuation utilities jointly support next-profile selection and semantic early exit, adapting sensing configuration and duration to evolving evidence. The expected semantic gain is further related to conditional mutual information, providing a value-of-information interpretation of continued online sensing. Experiments on Widar3.0 with six emulated sensing profiles show that, in comparison with full-sequence High, the resource-efficient Agentic setting reduces normalized cumulative sensing cost by 25.33% while achieving 85.79% Macro-F1. At the same utility checkpoint, semantic early exit provides a further 12.35% cost reduction over adaptive sensing without early exit, with a 0.97-percentage-point Macro-F1 decrease.
cs.AI / 55 / 2610.07835
DHCG: Dynamic Construction of Hierarchical Collaboration Graphs for LLM-Based Multi-Agent Reasoning
Abstract
LLM-based multi-agent systems (MAS) have demonstrated strong capabilities in solving complex problems across diverse domains. Recently, the dynamic orchestration of agent systems has become an important research direction. However, existing methods suffer from limited composition, misaligned dependencies, and inflexible scale, restricting their ability to adapt to reasoning requirements during execution. To address these limitations, we reframe MAS design as a partially observable Markov decision process, in which both the composition and scale of the MAS are dynamically determined. We propose DHCG, a novel framework that coordinates three modules (Planner, Worker, and Generator) to progressively construct a dynamic hierarchical collaboration graph from scratch based on the query and evolving execution feedback. At each step, guided by feedback, the Planner generates a set of distinct and complementary roles tailored to the current reasoning needs and selectively routes relevant information to each role. It can also finalize the hierarchical collaboration graph early or progressively expand it when additional reasoning is required. We further introduce action-aware preference optimization to train the Planner to make more effective decisions when constructing hierarchical collaboration graphs. We systematically evaluate DHCG across code generation, mathematical reasoning, and domain-specific reasoning benchmarks. DHCG achieves state-of-the-art average performance among the compared methods, improving over the single-agent baseline by 13.06 points and outperforming both static and dynamic MAS baselines by 2.77-8.02 points. Additional experiments further demonstrate its generalization across different Planner backbones and unseen Worker models.
cs.AI / 56 / 2610.07860
WorkflowOps: Learning Agent Collaboration Priors for Multi-Agent Workflow Orchestration
Abstract
Multi-agent systems are increasingly deployed for complex knowledge work, yet their orchestration layers remain largely memoryless: each new task is decomposed, assigned, and executed from scratch with no benefit from prior successful executions. We present WorkflowOps, a multi-agent workflow orchestration framework that learns agent collaboration priors from historical workflows and expands its agent pool on demand to cover new capability requirements. Our approach introduces three coupled mechanisms. First, a transition probability matrix captures pairwise agent collaboration frequencies from past workflows and applies them as soft guidance during DAG workflow construction through intra-layer ordering optimization, probability-thresholded edge suggestion, and transitive reduction for parallelism maximization. Second, a sufficiency-driven agent creation loop detects capability gaps via semantic matching scores, generates specialized agents through an LLM, and simultaneously injects them into the collaboration matrix, so that newly created agents are immediately usable with predicted collaboration priors. Third, a layered semantic matching strategy uses pre-trained sentence embeddings for fast, deterministic capability matching as a first pass, invoking LLM verification only for low-confidence cases, thereby reducing LLM routing calls by over 80\% compared to pure-LLM approaches. Experiments on mixed code, math, and question-answering suites show that WorkflowOps improves end-to-end pass rates over recent workflow-construction baselines, with the largest gains on structured, decomposable tasks where past agent handoff patterns transfer.
cs.AI / 57 / 2610.07881
Self-Referenced Social Preferences: Cooperation without Observing Others Rewards
Abstract
Social preferences can promote cooperation in multi-agent reinforcement learning, but existing approaches often require agents to observe the rewards of their peers. In many real-world interactions, however, an agent can, as humans do, observe others' behavior and outcomes without access to their private reward signals. We introduce self-referenced social preferences, in which each agent learns a model of its own reward, applies it to other agents' observed transitions to assess their outcomes from its own perspective, and feeds these self-referenced assessments into standard social preferences. We study two ways to incorporate these assessments: modifying the learning reward, or using them to weight policy updates. We evaluate the approach on three sequential social dilemmas, Escape Room, Clean Up, and Commons Harvest, which require volunteering, public-good contribution, and resource restraint, respectively. Across all three environments, agents learn cooperative behavior without observing others' rewards, including in settings where independent learners fail to cooperate, and frequently achieve more equitable divisions of jointly produced returns than agents with access to true rewards. The effective integration point depends on the social preference: inequity aversion works best in the reward together with a value look-ahead, whereas a purely benevolent preference benefits from policy-update weighting. Under partial observability, the policy-update approach continues to support cooperation. These results show that explicit access to other agents' reward signals is not necessary for learning cooperative behavior: social preferences can instead be grounded in self-referenced assessments of others' outcomes derived from their observed behavior.
cs.AI / 58 / 2610.07886
ShanLiangRen: A Nutrition Agent for Personalized Daily Meal Planning
Abstract
Dietary nutrition planning plays an important role in chronic disease management and maintaining a healthy body. In applications, it must simultaneously satisfy personalized constraints and reasonable multidimensional nutritional goals. These two aspects often conflict, and user constraints evolve with feedback, resulting in a substantial gap between generic guidelines and executable plans. To bridge this gap, we first propose the personalized fully quantified multiobjective dietary planning problem (MDP). To tackle MDP, we develop a nutrition agent, ShanLiangRen. The system first transforms dietary specifications, nutrient data, user attributes and natural language requirements into an individualized constrained planning instance. It then employs an exact retrieval-augmented generation method to shrink the feasible candidate set from a large scale ingredient and recipe space. Finally, it adopts a refinement guided by Pareto principles, where an LLM iteratively revises candidate plans under deterministic nutrition computation and feedback from constraint verification. The system outputs fully quantified meal plans with explicit ingredients and portion sizes, together with reports on nutrition compliance that show constraint satisfaction and nutrient interval attainment. We have released the system online as a WeChat Program, ShanLiangRen. A demo video is available at https://www.youtube.com/watch?v=652OtY5VlGA.
cs.AI / 59 / 2610.07906
Isotropic Yet Undecodable: The Sequential Content-Sufficiency Gap in Latent-Predictive Text Representations
Abstract
We study sequential content sufficiency by investigating whether a representation retains the ordered target information available in its input. An information-theoretic decomposition separates input ambiguity, representation loss, and readout mismatch. We construct recoverable views where perfect agreement and joint isotropic Gaussianity coexist with zero target information, and establish limits imposed by deterministic canonical anchors. Token log-loss provides a one-sided information-loss bound; a fixed-penalty ridge analysis shows why rank alone cannot determine prediction risk. These results motivate CANOPE, a nonautoregressive framework with ordered latent canvases, canonical-token supervision, and geometric regularization. On 40,000 validation sequences, latent-agreement (PL0) and token-grounded (PL2) have nearly identical pooled ranks but reach 13.5% and 98.8% positional Recall@1, respectively, under strong natural corruption when the correct target length is provided. On 3,930 LJSpeech validation utterances, frozen PL2 with a trained MatchaTTS readout yields 21.54% word error rate (WER) on corrupted text, versus 99.22% for frozen PL0, while end-to-end MatchaTTS reaches 10.93%. These results show that geometric regularity alone does not guarantee recoverable sequential content or effective downstream access in the text settings studied here.
cs.AI / 60 / 2610.07907
Continuous Memory Machines
Abstract
Recurrent neural networks typically compress information into a single vector-valued recurrent state, forcing short-term computation and long-term retention to share the same representation. Past extensions alleviate this bottleneck by increasing the memory capacity or separating timescales, but lack the combination of rapid neuron-level processing and longer-term retention found in biology. To that end, we introduce the Continuous Memory Machine (CMM), a recurrent architecture with matrix-valued short- and long-term memory states serving distinct functional roles. Building on the Continuous Thought Machine (CTM), the CMM's short-term memory tracks recent neural activity, with uniquely parameterized neuron-level models learning to use these activity patterns for computation. A persistent long-term memory stores information for later use, with a Transformer jointly updating both memory stores, providing an expressive bidirectional read--write mechanism such that each store can reorganize its own contents and both read from and write to the other. Across algorithmic, in-context learning, and recurrent reasoning tasks, the CMM outperforms a broad suite of baselines, exhibiting stronger generalization than prior memory-augmented networks while preserving the CTM's interpretable attention patterns. Code is available at https://github.com/SakanaAI/continuous-memory-machines.
cs.AI / 61 / 2610.07935
SIGMA: Self-Improving Alignment Generalization from a Model Spec
Abstract
LLM agents are increasingly capable of executing complex tasks and of recursively improving themselves on easy-to-verify objectives such as software engineering and mathematics. Since alignment is much harder to verify, this creates a growing risk of capabilities increasing without appropriate safety alignment, especially as capabilities expand to auto-research and cybersecurity. Existing approaches focus on capability self-improvement using verifiable feedback or on alignment training with supervision from stronger models or curated data, creating an external supervision bottleneck for alignment. We ask whether current models can improve their own safety alignment, and propose SIGMA, a data generation and training pipeline enabling alignment self-improvement that generalizes to out-of-distribution settings. Given only a "Model Spec" stating the model's desired behavior, SIGMA leverages a model's reasoning capabilities to strengthen its own safety reasoning. SIGMA first performs spec-guided task synthesis, using the candidate model as a task designer agent to generate diverse alignment dilemma scenarios and convert them into training tasks that stress-test its understanding of the Model Spec. Next, SIGMA conducts self-judged alignment training through supervised fine-tuning and rubric-based reinforcement learning with the model itself as the reward model. Despite training only on single-turn chat data, SIGMA improves safety alignment in multi-turn agentic environments (AgentHarm harmfulness decreases from 22.6 to 14.8; Agentic Misalignment decreases from 79.1 to 3.8), outperforms Deliberative Alignment and Constitutional AI baselines, and retains general capability. Analyses show that a Model Spec balancing harmlessness and helpfulness, test-time reasoning for safety deliberation, and high-quality rubrics from SIGMA's task designer agent are crucial for effective self-improvement.
cs.AI / 62 / 2610.07948
Confidence Reasoning Graphs: Structured Confidence Estimation for LLM Agents
Abstract
When using an LLM agent in a consequential domain, making an informed decision about whether to trust its output or intervene requires calibrated confidence in the agent's success. Confidence estimation for agents is difficult because evidence about success is distributed across heterogeneous, interdependent steps of an agent's trajectory. Practical agentic deployments introduce further challenges: frontier LLMs often provide limited access to internal signals, agent roll-outs are costly, and training data may be unavailable or quickly become outdated. To address these challenges, we introduce Confidence Reasoning Graphs (CRGs), an inference-time framework that estimates the probability an agent accomplished its task from a single trajectory, without privileged model access or training data. Rather than compressing an execution into a single holistic judgment, a CRG begins with the claim that the agent accomplished its task, decomposes it into contextualized sub-claims grounded in trajectory evidence, estimates confidence for each terminal claim, and finally aggregates these into an overall confidence estimate. Across three agentic benchmarks, three backbone models, and three agent frameworks, CRGs yield better-calibrated confidence and stronger risk-aware decision making than verbalized, sampling-based, and white-box surrogate baselines. We further find that calibration error alone can be misleading: a white-box surrogate baseline appears well calibrated while providing near-chance discrimination. Ablations attribute CRG's improvements to claim-level confidence estimation and aggregation rather than graph construction alone. Finally, a CRG exposes the claims and trajectory evidence underlying each confidence estimate, enabling it to be audited at decision time.
cs.AI / 63 / 2610.07972
Can Agents Work for Everyone? Cross-User Reliability for Mobile GUI Agents in Personalized User Interfaces
Abstract
Mobile GUI agents increasingly operate on interfaces influenced by users' histories and preferences, but their reliability across different users remains underexplored. We introduce PAIR (Personalized Application-state Instantiation and Rendering), a pipeline for constructing user-conditioned application states that enables controlled evaluation of the same task across different users. We further introduce RePAIR (Reinforcement learning with Personalization-Aware Interaction Rewards), a training approach that learns from cross-user differences in subgoal outcomes to improve reliability across user-conditioned mobile environments. Across six agents, we find substantial variation in task success across users and consistently lower subgoal achievement in user-conditioned UI contexts (6.98 to 15.4 pp). This gap further increases for personal targets drawn from each user's own content (8.77 to 22.0 pp). Failures in these contexts frequently involve selecting another item instead of the intended target, particularly before target exposure. Finally, RePAIR improves user-conditioned SAR (+5.87 pp), all-success (+7.50 pp), and overall Task SR (+9.42 pp) over its supervised fine-tuning parent on unseen users, providing initial evidence that explicitly learning from cross-user variation can improve GUI-agent reliability.
cs.AI / 64 / 2610.07979
Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents
Abstract
As agents continuously improve by generating and revising Skills, the process that discovers and refines those Skills becomes a learnable object in its own right. Task-Skills directly act on task execution, whereas Meta-Skills govern how agents discover and improve future Skills; their value therefore emerges through the subsequent search processes they induce. Existing approaches improve Meta-Skills from observed raw Skill-search trajectories and branch outcomes. However, branch performance entangles the effects of the initial discovery state and the Meta-Skill revision that generated the search process, making it difficult to characterize what a particular revision actually changed, and pushing updates toward revisions that benefit from favorable states rather than those that improve the process. We introduce HMED (Hindsight Meta-Experience Distillation), a mechanism for constructing Meta-Experience for self-improving agents. HMED revisits the completed event from which a revision originates and re-executes the incumbent and revised Meta-Skills from the same restored discovery state, so that the changes associated with the revision can be observed under a shared condition. Each comparison is distilled into a Meta-Experience, a structured record that can be reused by future updates, so that even revisions that are not ultimately retained still contribute a learning signal. Across three interactive agent benchmarks and both open-source and closed-source models, HMED consistently improves Skill discovery performance over strong baselines, shifting Meta-Skill learning beyond branch outcomes toward the consequences of changing the improvement process.
cs.AI / 65 / 2610.08033
Learning in Dreams, Winning in Reality: A Continuous Dyna Loop for a Ten-Hero MOBA
Abstract
World models are usually judged from the inside: by prediction loss, by the return a policy earns in imagination, or by how convincing their frames look. We judge one from the outside. We learn a structured, multi-agent world model of a complete ten-hero MOBA (206 units, every hero acting every tick, games of up to 6,000 ticks), train a policy only inside it with 1,400-tick free-running imagined episodes, and measure that policy in the real game against the opponent the game ships with. The real game never provides a gradient; it provides the policy's own games as training data for the world model, and an online evaluation that selects and anchors the policy. Run as a continuous asynchronous Dyna loop, the policy wins 70.2% of real games as radiant (421 of 600; 95% CI 66.4-73.7) on seeds never used for any decision, up from 0% for dream training alone and 33.7% before the loop. It wins none as dire, and neither does the shipped opponent when it plays itself. Four findings explain the result. Model exploitation is invisible from inside the dream: every unanchored run collapsed within a few updates while no in-dream metric tracked the collapse. A world model that is accurate on its training corpus is badly wrong on the policy's own games, and Dyna repairs it there, which is worth +9.2 points of real win rate with the policy recipe held fixed. Finally, the policy inherits its world model's fidelity profile mechanic by mechanic: the model represents the macro game but not crowd control, cast timing or lethality, and the policy wins by map-wide pressure with almost no coordinated fighting. We release the world model, the dream-PPO harness, a world-model debugger, the evaluation protocol, and every policy and log.
cs.AI / 66 / 2610.08036
Same Feedback, Different Answer: Measuring Run-to-Run Instability in Frontier-Model Customer Feedback Analysis
Abstract
AI agents are increasingly being programmed to automate knowledge work over large collections of unstructured data. Such automation requires repeatability: when the underlying evidence is unchanged, the agent's categories, priorities, and counts should not shift materially between runs, even if each individual answer appears plausible. We introduce a repeat-run evaluation framework that aligns semantically equivalent categories and focuses on two operating metrics: theme churn, the normalized change in the returned category set, and volume disagreement, the change in counts for categories that persist. We evaluate three recurring customer-feedback tasks across eight frontier models, corpus sizes from 100 to 5,000 records, multiple prompts, and three execution designs: raw generation, taxonomy-free hierarchical decomposition, and a taxonomy-grounded agent (TGA) using persistent themes, subthemes, and record-level predictions. With Claude Opus 4.8 and the 1,000-record corpus fixed, TGA reduces theme churn by 86--88% relative to both raw generation and hierarchical decomposition, while matched-theme volumes have zero disagreement. The taxonomy-grounded agent is more stable than every raw model in the screen, remains more stable at each corpus size, and keeps this advantage when theme matching is made stricter or looser. Although evaluated on customer feedback, the framework targets repeated synthesis of unstructured corpora more broadly, including financial reports, legal documents, incident records, and scientific literature. Overall, these results show that taxonomy grounding produces more consistent and repeatable outputs for recurring knowledge work.
cs.AI / 67 / 2610.08048
DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks
Abstract
LLM agents often lack the operational knowledge to act reliably in new environments, as they must discover specific tool behaviors or environment conventions on their own. Without memory of past attempts, they repeat the same mistakes across tasks, leading to more task failures and longer trajectories. To address this, agentic systems typically rely on human-written guidelines or on procedural memory built from training tasks and an oracle verifier, both of which require prior knowledge of the environment. We present DAEDALUS, a method for bootstrapping reusable agent memory from self-generated practice without existing tasks or oracle verifiers. DAEDALUS pairs two agents: an explorer that interacts with the environment to generate challenging yet solvable tasks, and a solver that attempts them. A heuristic is derived from each solver failure and accepted only after the solver repeatedly succeeds with that heuristic in context. These outcomes also provide feedback for the explorer to refine the difficulty of future tasks. Accepted heuristics are then consolidated into a memory bank for test-time use. Across AppWorld, $τ^2$-bench, and AutomationBench, DAEDALUS improves mean success rates by up to 15.9 points and pass^5 by up to 2.2x over a no-memory baseline, and is competitive with methods using training tasks, at a lower inference cost than most. We show that performance gains already emerge with a small exploration budget, and that its heuristics also benefit agents from other model families. Our ablations further reveal that solver traces provide the key information needed to derive effective heuristics, while factorizing early discoveries makes exploration more cost-efficient. Beyond memory construction, we find that the tasks generated by DAEDALUS can serve as a proxy for benchmark tasks when ranking models by performance. Code and artifacts: www.github.com/illuin-tech/daedalus.
cs.AI / 68 / 2610.08076
SpeedrunBench: Challenging LLM Agents with Video Game Speedrunning
Abstract
Frontier LLM agents have been shown to be capable of solving increasingly complex tasks for which humans have measurable solutions. This begs the pertinent question of whether LLM agents can go beyond what humans have already solved. The ability to develop sophisticated strategies to tackle consequential problems becomes paramount as well-trodden, human-developed solutions become insufficient for problems for which we lack context or enough training data. We study agents' capability of such strategy formation through the communal practice of video game speedrunning. In speedrunning, practitioners compete to find the fastest way to complete a video game under certain conditions, and in so doing uncovering interesting unorthodox play styles that require a thorough understanding and mastery of the underlying game mechanics. We introduce SPEEDRUNBENCH, a benchmark that evaluates frontier LLM agents across 9 different games. To perform well in this benchmark, agents must repeatedly improve their strategy, reflect on their performance, exploit their gained knowledge, and reason across a long-horizon of actions to improve on an increasingly difficult problem: being faster than themselves and everyone else. Our experiments show that while frontier agents approach human world records in simple platformer games, they remain behind human performance on longer, more complex games under practical budgets. These results suggest that SPEEDRUNBENCH is a useful testbed for studying agents' strategy formation capabilities as well as being a saturation-resistant evaluation measure, as there is almost always a faster completion time waiting to be discovered.
cs.AI / 69 / 2610.08077
Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
Abstract
Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction. For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful information about what the task requires and how the agent fails. We ask a complementary question: can hindsight teach an agent what it could have anticipated before acting? We introduce prospective learning, which uses post-hoc experience to supervise foresight predictions from the pre-interaction view, and instantiate it with Self-Retrospection Distillation (SRD). Intuitively, a completed trajectory reveals knowledge that would have been useful and pitfalls that should be avoided; SRD distills this privileged hindsight into trajectory-blind foresight of the same policy. Foresight serves only as a training target and need not be explicitly generated at inference time. Across 10 tool-integrated reasoning and long-horizon agentic tasks, SRD complements RLVR and self-distillation baselines with gains of up to $24.2$ pp. Its advantage is especially pronounced when reward contrast is scarce: when $37$--$98\%$ of rollout groups are reward-uniform across model scales, yet SRD can still exploit learning signal from sampled trajectories. In the 2B setting, where $98\%$ of groups are all-failure, the RLVR training ends up at $0.0\%$ success, while adding SRD reaches $60.6\%$ under the same rollout budget. Our results suggest that post-hoc agent experience is useful not only for evaluating or improving behavior, but also for shaping predictive representations before available interaction.
cs.AI / 70 / 2610.08082
POLAR: Ontology-Guided Risk Prevention for Tool-Calling LLM Agents
Abstract
LLM tool-use agents operate in dynamic environments where many actions carry operational risk. However, most safety mechanisms react only after errors manifest. Existing pre-emptive approaches either fine-tune the agent on chain-of-thought deliberation or compile natural-language guardrails into runtime checks, but they do so without exposing a structural, auditable verdict. We propose POLAR, a guardrail framework for small tool-calling agents that assesses reversibility through a structured two-layer ontology. POLAR assigns each action a graded reversibility score by deriving a candidate inverse sequence; calls failing a threshold are pruned before execution. Evaluated on $τ^2$-bench across six agent models, POLAR improves mean task reward by 0.11 to 0.18 points on airline for four of six agents, but only eight of eighteen model--domain cells improve overall; retail and stronger agents often regress. POLAR provides an auditable structural check and characterizes its task-utility trade-offs. Reward is not a direct measure of prevented harm.
cs.AI / 71 / 2610.08101
Beyond Corrected Memory: Execution Consistency in Multi-Agent Systems
Abstract
Shared memory coordinates agents' actions, but correct records do not establish that those actions satisfy task requirements. Memory governance and failure diagnosis regulate or inspect recorded information; they do not by themselves establish whether it is sufficient to judge task duties. We define execution consistency through duties governing state use, information handoffs, and final-state agreement, with explicit evidence conditions for judging fulfillment. Our core claim is that identical retained records can correspond to compliant and violating executions under the same task rule. Controlled removal of evidence such as receipt, action dependence, or response validity leaves 82.4% of opposite-label pairs indistinguishable; restoration separates 97.9% of the merged pairs. Natural-log annotations identify the defined violations in actual executions. However, existing logs do not always explicitly represent the execution relationships needed for these judgments. To assess the definition's practical value, we use CAVERT, a framework for consistency diagnosis and recovery, to extract supported relationships from logs and apply these criteria. It consistently outperforms contract-prompted LLM and rule-based baselines in diagnosis across all 12 benchmark-executor settings. Under the same gate and executor limits, it also outperforms rule-guided recovery in all four evaluated environments. These findings identify execution evidence that agent-memory and execution interfaces should preserve for reliable judgment.
cs.AI / 72 / 2610.08102
DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents
Abstract
Conversational MLLM agents are increasingly expected to assist in professional workflows, from AI research and engineering design to product management and business operations. Yet this capability remains underexplored: existing benchmarks largely focus on informal, everyday interactions and personal-life scenarios featuring photographic natural images, isolated static artifacts, and recall-oriented questions. In contrast, professional scenarios often involve structured, information-heavy artifacts that undergo frequent revisions and authority updates, and compositional queries requiring reconciliation of many artifact versions while tracking state precisely. To address these challenges, we introduce DSV-Mem, a benchmark for evaluating Dense Stateful Visual Memory. DSV-Mem comprises expert-reviewed scenarios and 1,000 questions across five user-oriented categories (Current State, Past State, Derived State, Change History, and Conflict/Refusal). A Hartley-inspired criterion favors questions with broader visual-evidence inspection demands. We also introduce a generation harness that produces evaluation suites by decoupling state-transition synthesis from conversation filling. Evaluation over 27 configurations spanning frontier and open-weight models and memory management methods reveals that the strongest baseline scores below 45% on DSV-Mem. Analysis surfaces findings: 1) multimodality and information density both contribute to difficulty, but state evolution, particularly the number of governing updates, is the dominant tested factor. Raw conversation/haystack length, OCR, and arithmetic are not the primary bottlenecks; 2) models often fail to verify user premises against prior state updates before answering; 3) increased reasoning effort and memory management methods yield limited gains, whereas state-aware designs prove more effective. The benchmark and code will be publicly released.
cs.AI / 73 / 2610.08138
Test-Time Agent Evolution for Long-Horizon Legal Reasoning
Abstract
Legal intelligence aims to support reliable decision-making across long-horizon legal processes involving evolving case states and multiple roles. However, real-world legal deployment exhibits substantial case heterogeneity in facts, evidence, and procedural contexts, exposing the limitations of static agent strategies. Moreover, legal reasoning is inherently interdependent across roles and procedural stages, making global reliability fundamentally different from isolated role competence. To address these challenges, we study training-free test-time agent adaptation, where agents continuously exploit deployment-time signals from preceding cases and ongoing interactions without updating model parameters. We propose \method, which introduces \emph{Test-Time Memory Evolution} to retrieve reusable experience from previous cases, adapt it to the current factual and procedural context, and consolidate accumulated experience for subsequent decision-making. Further, \emph{Rubric-Aligned Collaboration} verifies and revises role-specific actions according to behavioral and procedural requirements, enabling coordinated decision-making across roles and stages. Extensive experiments on J1-EVAL and LegalWorld across five backbone models demonstrate consistent improvements over representative reasoning and agent baselines with reasonable interaction and computational costs. Ablation and case studies further show that the two components provide complementary benefits in experience adaptation and cross-role coordination, improving the reliability and efficiency of long-horizon legal reasoning.
cs.AI / 74 / 2610.08142
Partially Observable Zero-shot coordination by Predicting Intention of Partner
Abstract
Zero-shot coordination in embodied settings requires acting while the partner is intermittently out of view, leaving existing methods with ambiguous partner representations and uncertainty over hidden partner states. We propose Predicting Intention of Partner (PIP) to jointly address these challenges. PIP uses a Joint-view VAE to distill richer training-time evidence from the union of both agents' local observations into a partner representation available from local observations alone. Partner-state Belief networks further infer the partner's hidden location and behavioral tendencies from the ego agent's interaction history. We evaluate PIP in Burrito-PO, Overcooked-PO, and a Melting Pot substrate, together with a human evaluation in Burrito-PO. PIP attains the highest mean performance among the compared methods across all three benchmarks. Human evaluation and diagnostic analyses further support coordination with unseen partners and the contributions of both components under partner occlusion.
cs.AI / 75 / 2610.08165
Which alloy composition,what process parameters? Inferring the recipe from optimized metallic microstructure and texture
Abstract
The mechanical properties of a metallic alloy are set by its microstructure and texture: the size and shape of its grains and the orientation of their crystals. That structure is in turn set by a recipe, the alloy composition together with the processing parameters. Alloy development runs this chain forwards, tuning the structure until a target property is met. Running it backwards, from an optimized structure to the recipe that would produce it, still relies on expert knowledge. We ask whether this backwards step can be learned. On an in-house dataset of 107 magnesium alloy extrusion conditions across 14 alloys, each with optical micrographs and an X-ray texture measurement, we compare three descriptors of microstructure and texture: conventional grain and texture statistics, a vision embedding from a pretrained image encoder, and a graph neural network on the grain network. Each is paired with prediction heads for two tasks: the alloy composition given the process (Task A), and the process parameters given the composition (Task B). Under 5-fold cross-validation, the conventional descriptors identify the correct alloy for 65% of held-out conditions, against 17% for always guessing the most common alloy, while the learned embeddings stay below 30%. The process parameters are recoverable but noisier: compared with using the composition alone, the microstructure roughly halves the temperature error. Because only a few alloys were cast and only a few press settings were used, both answers are discrete, and heads that pick from these known options, while respecting their order, worked better than heads that predict a free value.
cs.AI / 76 / 2610.08176
LFHE: Local-First Heuristic Evolution for Bounded Local Topology Search in Decentralized Learning with Non-IID Data
Abstract
Decentralized learning is highly sensitive to communication topology under non-IID data. Adaptive peer-selection methods can exploit local model information, but broader peer discovery may require increasingly large control state, whereas direct spectral optimization typically relies on graph-wide information. We study the intermediate setting of bounded local topology search and propose Local-First Heuristic Evolution (LFHE), a representation-driven rewiring framework whose candidate discovery and scoring use only ego-neighborhood and friend-of-a-friend (FoF) information. The structural score admits an exact interpretation through graph Dirichlet energy: its sum across clients equals twice the representation Dirichlet energy, which under standard linear consensus dynamics governs the instantaneous dissipation of representation disagreement. LFHE combines this state-dependent structural signal with early exploration and degree control, while algebraic connectivity remains an offline graph diagnostic. Under bounded sparse degree, its FoF candidate state remains local rather than expanding toward population-wide peer tracking. Across four image, speech, and text benchmarks, LFHE achieves competitive decentralized learning performance. Matched-protocol controls identify the structural term as the principal empirical topology-selection signal, while comparison with broader peer discovery exposes a trade-off between predictive performance and discovery-state locality. Together, these results motivate state-aware bounded local topology search between pairwise peer selection and globally informed topology optimization.
cs.AI / 77 / 2610.08214
Mathematical Proof Assistants for Teaching Logic: The LogiKEy Methodology
Abstract
We report on an approach to teaching logic to mixed groups of computer science, mathematics, and philosophy students, based on the logico-pluralistic LogiKEy methodology, used for more than a decade in courses, summer schools, and tutorials. LogiKEy uses classical higher-order logic (HOL) as a universal metalogic in which object logics, classical and non-classical alike, are encoded by defining their semantics; through these semantical embeddings a single proof assistant (e.g. Isabelle/HOL), with its automated theorem provers and (counter-)model finders, becomes one environment in which students learn, experiment with, and compare logics. After making the pedagogical case for proof assistants in the logic classroom, we present a graded sequence of classroom examples, each transition motivated by a limitation of the preceding representation, by a need for more explicit modelling resources, or by a new application. A liars-and-truth-tellers puzzle leads from propositional to modal logic; the Wise Men puzzle leads on to dynamic epistemic logic; Boolos's curious inference illustrates what a higher-order meta-logic buys, even for automated proof search; Chisholm's paradox takes the sequence into deontic logic, and from standard to dyadic deontic logic; and Gödel's ontological argument brings it to a research-level metaphysical argument. We then rebut the objection that embedding everything in classical HOL is monism rather than pluralism, reflect on three years of teaching such a course, and sketch the portability of the approach beyond Isabelle.
cs.AI / 78 / 2610.08215
Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?
Abstract
Learning from experience is essential for LLM agents to adapt to unfamiliar and dynmaic environments. Evaluating this ability is therefore important for understanding how effectively agents acquire and use new knowledge. Existing benchmarks have sought to evaluate this ability, but they primarily evaluate tasks whose rules are provided in the instructions or already familiar to pretrained models, making it difficult to distinguish learning from interactions from reasoning with existing knowledge. To address this, we introduce Learn2Play Bench, a benchmark of newly designed text-based games, whose rules are novel or counterintuitive, requiring agents to acquire knowledge through interaction rather than rely solely on pretrained knowledge. These games provide reproducible feedback and automatic scoring, enabling controlled evaluation of learning across repeated attempts. We also vary game instances to test whether agents can apply what they have learned to new situations. Therefore, we evaluate how backbone models, self-evolving methods, and agent harnesses affect agents' learning ability, revealing three findings: (1) Experience retention: Retaining complete records of actions and feedback can support more effective learning than summarizing these experiences into rules or strategies. (2) Human agent gap: Top-performing human players achieve higher peak scores than the evaluated agents. Human explore more varied strategies, and repeat actions less. (3) Harness matters: With the backbone fixed, changing the harness can improve performance while reducing estimated inference cost. Together, these findings provide insights into how LLM agents learn from experience and suggest directions for future work to improve their learning ability. Project website: https://liushiliushi.github.io/learn2play-bench-website/
cs.AI / 79 / 2610.08216
Quantum Entangled Multimodal Fusion Networks (QEMFN): Resource-Aware Hybrid Vision-Language Fusion via Trainable Entanglement
Abstract
Multimodal vision-language systems typically fuse image and text embeddings through classical operators such as concatenation, attention, bilinear pooling, or tensor interactions. We propose Quantum Entangled Multimodal Fusion Networks (QEMFN), a hybrid quantum-classical framework that introduces parameterized entanglement as a structured inductive bias for multimodal fusion. Pretrained visual and textual features are projected into compact latent spaces, encoded as angle-parameterized quantum states, processed through intra-modal and paired cross-modal entangling circuits, and measured to produce fused representations for retrieval. Under matched parameter budgets and identical frozen CLIP backbones, QEMFN outperforms classical fusion baselines on COCO-5k and Flickr30k, including multilayer perceptron, tensor fusion, FiLM, cross-attention, compact transformer, and a dequantized paired-topology analogue. An ablation suite isolates the quantum module's contribution from the surrounding classical projections, and quantum-centric analyses report Meyer-Wallach entangling capability, expressibility, gradient variance against barren-plateau bounds, and entropy-performance correlation under controls for training progress alongside an intervention study on the entangling component. QEMFN is executed under shot-based estimation, a noise-modeled fake backend, and a real superconducting device with zero-noise extrapolation. This work does not claim quantum computational advantage; the contribution is the framework together with a controlled empirical and quantum-centric evaluation that positions trainable entanglement as an interpretable, hardware-executable fusion mechanism at scales accessible on contemporary devices.
cs.AI / 80 / 2610.08229
Confidence-Ordering Reversal under Contextual Priors in Neural Decoding
Abstract
Contextual priors improve neural-to-language decoding by reshaping candidate scores. However, confidence is read from the same reshaped scores, so the errors a prior leaves behind can become more confident with no change in accuracy to reveal it. We study how a prior shapes confidence in speech retrieval on MEG-MASC and MOUS using local decoding scores, a contextual prior combined by additive shallow fusion, and the fused top-two margin as confidence. Among initially incorrect predictions, we find a confidence-ordering reversal: a larger margin makes a repair more likely when the correct candidate starts near the top of the local ranking, but less likely when it starts lower. On MEG-MASC, pooled correctness AUROC is 0.87, yet AUROC separating repairs from residual errors falls from 0.70 at initial ranks 2-3 to 0.39 at ranks 21-50. Errors starting beyond rank 20, inside the reversed region, make up 46.6% of all post-fusion errors. We propose a score-level account: a repair must first close the correct candidate's initial deficit, limiting its final margin, whereas a residual error can build a large margin between two incorrect candidates. A causal intervention that changes only the fusion weight moves the reversal to deeper ranks as predicted. Under a word-level LM prior, it keeps moving after accuracy gain peaks, so a weight chosen for accuracy does not settle confidence. Reading local and prior scores separately improves selective decoding: the decoder answers on 74.5% of windows instead of 56.7%, while 92% of output sets still contain the correct candidate. Confidence after contextual fusion should retain the local and contextual evidence behind each prediction, not just the fused scores. Project website: https://confidencereversal.github.io/; Code: https://github.com/AmadeusFake/NeuDecodingConfReversal
cs.AI / 81 / 2610.08244
Sensor-Language-Action Models
Abstract
Sensors are useful not only for understanding the world but also for deciding what to do next. Existing sensor models however largely stop at perception: they recognize states or predict outcomes, leaving actions modeled separately through task-specific and often closed label spaces. We introduce Sensor-Language-Action (SLA) modeling, a framework that connects multimodal sensor observations, natural language, and actions within a unified model. SLA uses language as a semantic interface between sensing and acting, allowing heterogeneous actions to be represented, predicted, and explained while remaining grounded in the underlying sensor evidence. We build a large-scale SLA benchmark consisting of datasets that span more than 116,000 individuals, 79 sensor modalities, and 60 action groups, together with a multi-faceted captioning pipeline that aligns user context, sensor dynamics, and action evidence. Building on this framework, we present OpenSLA, a unified SLA model for hierarchical action prediction, state understanding, and action explanation. Extensive experiments on real-world tasks in clinical prediction, operating rooms, and metabolic health verify its superior performance over the state-of-the-art. OpenSLA also demonstrates intriguing capabilities including language-guided evidence grounding and zero-shot generalization to unseen actions and cohorts.
cs.AI / 82 / 2610.08314
The Standardization Trap: Certifying Joint Label Processing in Tabular Foundation Models
Abstract
Linear regression and kernel smoothing offer tractable explanations of in-context learning: in both, the features determine the weight assigned to each context label. However, whether this fixed-weight account describes pretrained tabular foundation models (TFMs) remains unclear. Testing this account using derivatives runs into a standardization trap: public TFM packages standardize the labels before the model sees them, yet ordinary derivatives also reflect behavior outside the set of standardized labels, making a model appear nonlinear even when every prediction it makes agrees with a fixed-weight map. We propose two certificates that depend only on predictions at standardized labels and can reject two distinct explanations: fixed-weight prediction and sums of independent nonlinear label transformations. Across the five public TFMs that we evaluate, our certificates show that changing one context label alters how other labels influence the prediction, a behavior we call joint processing. We further find that joint processing emerges with training and that attention scores carry most of the measured interaction. Together, these findings motivate TFM explanations that account for how context labels change the influence of individual examples.
cs.AI / 83 / 2610.08319
SCOPE: Certified Theorem Proving with a Language Model as the Policy Planner
Abstract
In proof assistants such as Lean, a generated proof must pass machine compilation checks, so evaluation needs no human scoring. Direct generation fails on multi-step numeric propositions: a proof is valid only if every content integer is correct, so the pass rate is bounded by the k-th power of the per-integer accuracy. Controlled corruption across 2,617 reference proofs confirms this power law. SCOPE (State-Conditioned Operator Planning and Execution) enforces the natural division of labor: the model plans over an operator vocabulary, a symbolic engine executes the numerics, and a compiler renders the proof. On a 218-problem suite it certifies 191/218 (87.6%) with a 135M backbone; the 7B DeepSeek-Prover-V1.5-RL certifies 18/218 at 27.5 times the tokens and 37.5 times the wall-clock, and DeepSeek-Prover-V2-7B certifies zero on a bidirectional dual suite. Multi-step thinking costs 6.12 discrete decision actions per problem and produces no natural-language thinking text. Replacing the lagged engine state in the decision frame with the current one lifts the pass rate from 117/218 to 191/218, while up-weighting the chain-end loss hurts. On the public Lean-Workbook library, 2,132 of 3,536 gradeable admissible problems certify (60.29%) with zero regression on the main suite. All readings come from a version-frozen review with independent rechecks and reverse verification. Restricting free generation and keeping decision-time information visible is a more direct route than enlarging the model.
cs.AI / 84 / 2610.08329
An AI-Assisted Formalization of the Poincaré Conjecture
Abstract
We present an AI-assisted Lean 4 formalization of the Poincaré conjecture. The project began with limited reusable formal infrastructure for the geometric analysis behind the proof. To organize this work, we combined a proof blueprint prepared by mathematicians with explicit milestone statements. These milestones enabled parallel agent work and gave mathematicians clear points to locate blockers and provide effective mathematical guidance. Our analysis identifies the human interventions and organizational choices behind this workflow. The project provides a starting point toward reusable infrastructure for future formalization projects; such infrastructure, once developed, could eventually reduce the cost of verifying mathematical results in geometric analysis.
cs.AI / 85 / 2610.08363
Explainable Failure Prediction and Prevention in Maritime
Abstract
Maritime systems operate in highly dynamic environments where unexpected equipment failures can compromise safety, reliability, and operational efficiency. Recent advances in artificial intelligence (AI), machine learning, digital twins, and predictive maintenance enable proactive failure prediction and prevention. However, ensuring trustworthy and explainable decision-making remains a major challenge in safety-critical maritime applications. This chapter reviews key AI technologies required for explainable failure prediction and prevention in maritime systems and presents a conceptual architecture capable of supporting autonomous or human-in-the-loop corrective actions. This architecture integrates data acquisition, time-series forecasting, anomaly detection, risk assessment, decision-making, and explainable AI into a closed-loop framework. With reference to the architectural components, a review and discussion of relevant maritime studies is performed, outlining their methods, advantages, and limitations. Furthermore, it highlights current challenges, including uncertainty and robustness, model generalization, explainability, limited availability of maritime datasets, and operational deployment, and identifies future research directions toward trustworthy AI-assisted maritime decision-making.
cs.AI / 86 / 2610.08364
Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations
Abstract
Frontier AI evaluations increasingly use open-ended, agentic, long-horizon tasks whose transcripts can span hundreds of pages of outputs and actions from complex multi-agent networks. The observability envelop-the range of what evaluators can reliably infer about an agent's behaviours-is therefore narrowing. Language model assistants can help classify and interpret agent behaviour but also afford human evaluators significant analytical degrees of freedom, threatening the reproducibility and auditability of language-model-based transcript analysis. Transect is an open source package built on Inspect Scout to help evaluators understand how a long agent run unfolded, identify behaviour worth investigating, and check interpretations against the transcript. Users specify task context and behavioural vocabulary in a reusable evaluation-family configuration, with judge models and analysis settings supplied separately. Transect's navigable reports align recorded events, token use, sub-agent activity, and model-generated behavioural labels on a common turn-based timeline. Reviewers can quickly grasp a run's narrative, trace any label or event to its source turns, and export the underlying data tables for cross-run analysis. We demonstrate the workflow on an AI R&D evaluation that generated almost 13 million tokens, dividing the agents' work into behavioural phases aligned with research-skill classifications, sub-agent delegations and interactions, and token use. The combined view shows a focus on operational work and manuscript production, with little evidence of a sustained hypothesis generation stage-arguably a necessary component for high-quality scientific outputs. Transect's flexible, customisable transcript-analysis pipeline will enable evaluators to keep pace with longer, more complex, more frequent AI evaluations while supporting scientific rigour, transparency, and reproducibility.
cs.AI / 87 / 2610.08432
EMHO: EMbodied Agent Harness Optimization via Experience Traces
Abstract
Improving embodied agents often focuses on optimizing the underlying model through training, while the surrounding agent harness that controls planning, context, and tool use is typically engineered. We ask whether this harness can instead improve itself directly from experience traces under sparse environmental feedback. We propose EMbodied Agent Harness Optimization (EMHO), a self-evolving framework that keeps the embodied model frozen and iteratively revises its harness by analyzing execution trajectories and prior harness history. EMHO optimizes beyond skills or recovery prompts, modifying how the agent monitors progress, uses vision tools, grounds observations, and responds to failures. To support multiple subtasks with a single harness, we introduce EMHO-Merge, which addresses trade-offs in jointly optimizing a single shared harness across subtasks by using episode-level gains and losses to guide evidence-supported refinement of when and how revised behaviors are applied. We evaluate EMHO on EmbodiedBench across navigation and manipulation tasks, and EMHO consistently improves task success for both Qwen 9B and 27B models. Qualitative analysis shows that EMHO goes beyond recovering from failures and unproductive actions to reshape how the embodied agent interprets and interacts with its environment.
cs.AI / 88 / 2610.08510
Cylindrical Geodesic Flow Matching for Quasiperiodic Physiological Signal Transformation
Abstract
Paired translation between quasiperiodic physiological waveforms (i.e., recovering a target oscillatory signal from the source) is central to the interpretation of cardiovascular signals derived from wearables placed at different body locations. This source-to-target mapping in these problems carries inherent geometric structure: the phase wraps around the cycle and must be treated as a circular variable, the amplitude remains strictly positive, and the beat-to-beat alignment can drift unpredictably across cycles and subjects. While deep neural networks have been used for phase estimation and complex-valued signal modeling, prior work does not explicitly learn phase transport between paired signals. Consequently, neither endpoint-supervised regression nor the standard affine path used in flow matching accounts for this phase--amplitude structure. We introduce \emph{cylindrical geodesic flow matching} for paired cardiovascular waveform translation. We show that the standard affine path used in flow matching distorts intermediate amplitude and instantaneous frequency when interpolating between quasiperiodic signals; replacing it with a closed-form geodesic on the phase--amplitude cylinder eliminates these artifacts and converts each training pair into dense, geometry-consistent velocity supervision. On zero-shot photoplethysmography and limited-support seismocardiography adaptation benchmarks, our method consistently outperforms interpolation baselines and matches or exceeds direct supervised prediction, reducing Hilbert Transform, $L_2$, and Dynamic Time Warping distance by up to ${\sim}15\%$ over the strongest competing baseline. These results suggest that bridge geometry is a critical inductive bias for flow matching on oscillatory signal translation.
cs.AI / 89 / 2610.08514
How Much Evidence Should a Coding Agent's Self-Correction Carry? Adaptive Dirichlet Evidence for Self-Distillation
Abstract
Execution feedback lets coding agents revise programs and learn from their own corrections. A correction's learning weight should reflect both the transitions supported by its executions and the amount of evidence behind that support. We introduce Effective-Evidence Self-Distillation (EESD), which represents these quantities separately. Normalized execution relevance determines relative transition support and an effective pseudo-count mass; a Dirichlet posterior then produces an uncertainty-penalized weight for KL-anchored correction learning. Under a symmetric prior, changing mass preserves category ordering, and effective mass yields a supervised coefficient bounded by its matched fixed-mass counterpart. Across four model-domain history sweeps, increasing visible observations from one to eight reduces future-outcome NLL by 55.0-59.3%. At eight observations, effective mass achieves lower NLL than fixed mass in all four comparisons. In the primary matched DeepSeek/RunBugRun study, argmax predictions agree on all 3,000 examples, with the largest NLL gain under concentrated relevance. After one correction-learning round, DeepSeek/CodeARC all-tests Pass@1 increases from 15.0% to 20.4%, with a paired 95% source-bootstrap interval of [+2.8, +8.0] percentage points. The twelve-setting downstream evaluation establishes the model-domain scope of this update. These results show how separating evidence support from evidence mass changes probability estimation and correction learning in coding agents.
cs.AI / 90 / 2610.08540
Toward Alignment Scaling Laws: A Framework and First Preregistered Measurements
Abstract
Whether alignment gets easier or harder as models grow is often argued from isolated findings, as if alignment were one property. We treat it as a family of measurable scaling relations: for each risk category r, the alignment burden needed to hold a fixed safety target is modeled as B_r(N)=a_rN^alpha_r, with N a capability proxy; against a budget proportional to N, scaling helps if alpha_r<1, keeps pace if alpha_r~1, and accumulates alignment debt if alpha_r>1. We give three operationalizations of burden and distinguish observed, audited and true alignment. A toy model, in which corrections consume capability headroom, makes the consequences explicit. We prove that the largest exponent among corrected risks, not an average, sets the long-run regime; that above 1 any policy holding headroom above a floor must grow super-exponentially; that, for burdens that are positive mixtures of power laws, fits on small models underestimate large-scale exponents; and that an audit that uncovers hidden failures without false positives never underestimates true alignment. We propose a pre-registrable protocol and apply reduced versions of it twice. A preregistered reanalysis of public adversarial-training data for Pythia classifiers finds that the compute needed to bring attack success under 10% grows as N^0.60. A preregistered pilot on Qwen2.5 0.5B-72B finds exponents of -0.05 for truthfulness and 0.48 for stated dispositions (both scaling helps under its reduced rule, though local slopes approach 1 at the top; replicated on Qwen3 0.6B-14B), while sycophancy (0.89, or 0.83 with two seeds added at 72B) and a planted backdoor are undetermined: the backdoor is removed quickly when its trigger is known but survives blind safety training at four of five sizes. We release four browser games that play these laws (www.aisafety.fun). We make no claim about which regime holds for current frontier models.
cs.AI / 91 / 2610.08552
AnyBottle: A Recipe to Only Keep the Concepts You Really Need
Abstract
Concept bottleneck models (CBMs) make predictions inspectable and intervenable by routing them through human-interpretable concepts, but originally required concept annotations. Annotation-free variants remove this requirement, but typically use large concept vocabularies, static at both training and inference, producing bottlenecks larger than any task or prediction needs and harder to inspect. We propose AnyBottle, a single recipe for building compact, task-specific CBMs. AnyBottle assumes only a frozen backbone and an unsupervised concept pool, such as a sparse autoencoder. A black-box teacher trained on the same backbone then guides selection: each round adds the concept that best explains the bottleneck's current failures, with candidates restricted to regions of teacher/student disagreement. Trained with nested dropout over this selection order, the final bottleneck predicts accurately from any concept prefix, so inference spends fewer concepts on inputs it is confident about early and more on hard ones. Since no stage is modality-specific, a new domain and task requires swapping only the backbone and concept pool. Across six vision and two text datasets and two teacher paradigms, AnyBottle yields bottlenecks with fewer concepts and higher concept consistency than annotation-free baselines, while staying close to the black-box reference. Overall, AnyBottle shows that going annotation-free need not mean going large: a small, discovered vocabulary can be as expressive as a much larger, fixed one.
cs.AI / 92 / 2610.08621
Recursive Game Creator: An Agentic Product-Level Experience-Oriented Game Harness
Abstract
Recent game design agents have made substantial progress in generating playable games. However, program correctness does not ensure an enjoyable experience for players. We present Recursive Game Creator, an experience-oriented harness to advance agentic game development from rough game prototypes into entertaining games. Recursive Game Creator organizes recursive development around four components: Designer, Builder, Player, and Reviewer. The Designer translates user instructions and Reviewer's feedback into detailed plans. The Builder turns these plans into candidate games. The coding-native Player creates and executes reusable policies through programmatic interfaces to efficiently collect diverse gameplay trajectories, mitigating evaluation bias caused by slow GUI-based collection. The Reviewer uses carefully designed trajectory-based metrics to induce player preferences, integrating with visual evidence and explicit textual preferences to evaluate games against game-specific criteria. Finally, the Reviewer accepts the better version and provides improvement reviews for the next round, closing the recursive loop. Our method achieves state-of-the-art overall performance of 77.89 on GameCraft-Bench. On GameASG-Bench, it achieves a strict task success rate of 53.2%, a 34.1% improvement over the same-model baseline, and the highest mean runtime-check pass rate at 93.4% among compared methods. A user study shows longer playtime and higher ratings. Code is coming soon.
cs.AI / 93 / 2610.08627
Parallel Predictive World Models for Accurate and Efficient Long-Horizon Planning
Abstract
Long-horizon world-model planning typically relies on autoregressive rollouts, where predicted states are repeatedly fed back into the model. This preserves temporal structure but creates a horizon-length sequential path and exposes later predictions to recursive decoded-state feedback. We introduce Parallel Predictive World Models (PPWM), which predict a finite-horizon trajectory in parallel while retaining causal interaction among future representations. Each horizon is conditioned on its causal action prefix, and future representations interact before decoding, separating temporal causality from state-by-state output recursion. We formalize this distinction by viewing autoregressive rollout as a causal trajectory map and identifying the decoded-state feedback pathway removed by PPWM. Across four visual-control tasks, PPWM achieves the lowest long-horizon prediction error and the highest Cross-Entropy Method (CEM) simulator success among the evaluated predictive interfaces. Meanwhile, PPWM achieves more than a 3$\times$ average CEM planning speedup over the autoregressive LeWM baseline. These results suggest that accurate and efficient long-horizon world-model planning does not require state-by-state autoregression, but can instead be achieved through parallel causal trajectory prediction.
cs.AI / 94 / 2610.08647
SquidAgent: Parallelize Wisely, Coordinate Efficiently
Abstract
LLM-based agents solve complex multi-step tasks, but sequential execution incurs substantial latency. In principle, parallelizing work across multiple agents should yield near-linear speedups. Yet existing parallel multi-agent systems often run slower than a single-agent baseline. We attribute this gap to two hidden costs that parallel execution incurs but a serial agent avoids. First, there is a re-exploration cost: redundant effort spent by parallel workers reconstructing context that the orchestrator already possesses, such as prior decisions, that would otherwise be inherited implicitly in a serial execution. Second, there is an alignment cost: the overhead required to reconcile inconsistencies across independently generated outputs. We thus derive a principled decision criterion: a layer should be parallelized only when its critical-path cost, plus re-exploration and alignment overheads, is lower than the corresponding serial cost. While this criterion is naturally expressed in wall-clock time, we observe that LLMs are poorly calibrated when asked to estimate task duration. To address this, we instead measure cost in predicted output tokens, which we empirically find LLMs can estimate substantially more reliably than wall-clock time. Building on this token-based criterion, we propose SquidAgent. It estimates all token budgets in a single planning step, forks each worker directly from the orchestrator's session to eliminate re-exploration cost, and replaces post-hoc reconciliation with a pre-generated shared convention block that converts alignment into a bounded upfront cost. A deterministic scheduler then applies the criterion layer by layer. Empirically, SquidAgent achieves a 2.2$\times$ mean throughput improvement and a 2.6$\times$ mean wall-time speedup over Claude Code, and a 2.0$\times$ throughput improvement over the strongest multi-agent baseline.
cs.AI / 95 / 2610.08662
ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding
Abstract
As coding agents increasingly undertake real-world work autonomously, judging whether their risk treatments are warranted has become important. Existing work evaluates related agent behaviors from separate perspectives, but lacks a systematic framework for unifying these behaviors. To bridge this gap, we introduce ParanoiaEval, the first benchmark for unified evaluation of risk-treatment capabilities in coding agents. Grounded in the well-established Avoidance-Transfer-Mitigation-Acceptance framework in software engineering risk management, ParanoiaEval operationalizes its 4 fundamental treatments for coding-agent settings and contains 200 evidence-controlled repository-level task pairs, each differing only in treatment-defining evidence. We further introduce dedicated metrics for risk-treatment violations and evidence responsiveness, using a human-calibrated agentic judge for reliable evaluation. Large-scale experiments on 8 representative models and a post-hoc human study reveal that (I) unnecessary risk treatment occurs in 11.2%-58.7% of runs despite explicit evidence, with substantial variation across agent configurations; (II) stronger task capability does not ensure more appropriate risk treatment, while treatment violations substantially harm developers' experience, establishing risk treatment as an independent capability dimension; and (III) agents exhibit systematic patterns consistent with established risk-management findings, suggesting that knowledge from human practice can guide the diagnosis and improvement of this capability.
cs.AI / 96 / 2610.08683
Coupled but Late: Turn-Taking Between Full-Duplex Speech Models in Unscripted Dialogue
Abstract
Full-duplex speech models are trained to converse with a person, but they are increasingly made to converse with each other, in self-play data generation, agent societies, and model-based evaluation. In that loop no human absorbs a timing error: each model's turn-taking is the other's input. We ask what timing the loop settles into. Two PersonaPlex-7B instances exchange audio tokens on a shared clock in unscripted conversation, and one floor-transfer rule is applied to them and to Switchboard. Their timing is coupled: re-pairing speakers across conversations destroys it. But the floor changes hands late, at a median of 400-560 ms against 137 ms for humans, and the last 120 ms of the partner's turn, where human projection places a tenth of its transfers, holds 1% of theirs. Delaying one direction of the channel shifts the response one-for-one and leaves the run-up to it empty, consistent with a reactive wait after the perceived end rather than the turn-end projection human timing requires.
cs.AI / 97 / 2610.08699
nanoMuse: An Open-Source Personal Agent for Every Device You Own
Abstract
Assistants from 2011 answered and waited, and agents from 2023 did a task and stopped. In September 2026 Meta's Muse showed an agent for one person, with accounts, devices, memory and a conversation that lasts, closed, in a vendor's cloud, in one country. Such an agent is expected to act on a person's accounts and devices, remember them across weeks, speak first when it is worth it, and answer for what it did. It is a kind of software, not a model, and until now had no open counterpart. This report defines the personal agent in five questions and three horizons. It reads how Muse is built from Meta's public record and a copy of its production prompt, each statement marked by its source. It then presents nanoMuse, the open-source counterpart under the GPL-3.0, one agent on every device a person owns, with hands on the phone's screen and the computer's. They share one conversation over a relay anyone can run; every action goes through a Sentinel, memory is files the person can read, and the model is their choice. Its size and cost are given as estimates. What is open, memory with provenance, an evaluation suite for the hands and an open model for them, is set out as a roadmap.
cs.AI / 98 / 2610.08720
WorldSolver: Can LLM Agents Simulate the Physical Dynamics via Solver Generation?
Abstract
LLM-based agents are increasingly advancing scientific and engineering problem solving, with physics simulation emerging as a challenging yet practical testbed for reproducing complex physical phenomena with application in embodied AI, games and films. As the workhorse of such simulation, a solver computes how the state of a dynamic system evolves over time. Building such solvers requires physical understanding to identify appropriate models, mathematical reasoning to formulate the underlying dynamics, and software engineering to implement them as executable code, yet this capability of LLM agents remains underexplored. To this end, we introduce WorldSolver, a benchmark of 168 simulation tasks derived from physical phenomena in 61 classic computer graphics papers, spanning 7 physical domains. Each task contains a code scaffold that provides a fixed simulation environment for the scene, with the solver implementation left for the agent to complete. Specifically, we evaluate them along three dimensions: Execution Checks for successful execution, Visual Fidelity for reproducing the intended dynamic behavior in the rendered simulation, and Physical Plausibility for physics-grounded verification of the generated dynamics. Experiments on frontier agents reveal that producing executable solvers is difficult itself, and satisfying visual and physical correctness is even harder. GPT-5.6-Sol and Claude-Opus-5 perform comparatively better than the other evaluated agents, yet achieve overall scores of only 48.7% and 46.7%, respectively. WorldSolver is an early step toward agentic solver generation, and we hope it helps drive progress toward agents that can faithfully simulate the dynamic physical world. Code is available at https://github.com/sirujiang/WorldSolver.
cs.AI / 99 / 2610.08722
Does an Agent's History Tell You When Compaction Will Hurt? A Modest, Bounded Effect on the TRACE Paired-Replay Corpus
Abstract
Many long-horizon agents compact their context on a global rule, usually a token budget, blind to what the agent was doing. We ask whether the agent's recent behaviour predicts when a compaction will hurt. TRACE's public corpus of 590 harness-triggered AppWorld compaction boundaries replays each boundary from a re-executed prefix state under the pre-compaction context and under the summary, and records the burden of the next actions: calls that error or repeat a call already made. We find that pre-boundary history predicts post-compaction harm only weakly. An internally prespecified contrast by prefix placement is a wide null, and the naive "has-written" label behind it turns out to measure trajectory phase. The best extension-protocol trigger reaches held-out AUROC 0.66 (0.64 on the replicate's own label) against a same-boundary replicate of 0.72; the best frozen, interpretable trigger avoids 21% of harmful (positive-burden) boundaries while keeping 84% of compaction opportunities, and exceeds the random-rule expectation on count but not on burden mass (a post hoc comparison). Whether the best trigger beats a token-budget rule at matched retention cannot be evaluated on the release. We state what corpora should ship to answer it.
cs.AI / 100 / 2610.08761
VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
Abstract
Self-improving policies continually expose new failure patterns, changing what their judges must be able to verify. However, current fixed judges constrain both optimization feedback and the discovery of useful training examples, limiting further self-improvement. This challenge is even more acute in embodied reasoning, where reliable evaluation must account for spatial grounding, causal reasoning, and safety-aware decision-making. We introduce VeriFine, an agent harness framework that scales verification through the co-evolution of the policy, training curriculum, and judge. The Policy Improvement Loop uses a rubric judge to diagnose recurring failures, construct an adaptive curriculum, and optimize the policy. When progress plateaus and verification becomes a bottleneck, the Judge Improvement Loop selectively queries human guidance on informative failure cases and refines the judge through coactive calibration, in which humans and agents resolve disagreements and converge toward the objective rubric of physical reasoning. The revised judge then guides the next stage of data selection and policy optimization. Experiments on driving and robot navigation tasks demonstrate continuous self-improvement in both policy and judge capability across reinforcement and supervised fine-tuning. These results show how scaling verification supports continuous self-improvement as policy failure patterns evolve.
cs.AI / 101 / 2610.08205
Tool-calling retrieval versus vector RAG for a small Greek--English knowledge base: accuracy and robustness to how users type Greek
Abstract
Assistants grounded in a small, frequently edited knowledge base can retrieve through tool calls to a live data interface or through vector retrieval-augmented generation (RAG). We compare the two on KyGround, a benchmark of 198 questions drawn from the published records of a Greek--English agricultural platform on Kythera, Greece, with answers verified automatically against the records and each question posed in up to nine forms, including Greek without accents, in capitals and in three Latin-script (Greeklish) schemes. With Claude Haiku 4.5 as router and answer model, a reconstruction of the platform's tool agent answered 71.6\% of canonical Greek questions correctly and vector RAG 95.3\% (difference $-23.6$ percentage points, 95\% CI $-33.1$ to $-15.1$). Letting the router write the vector query changed nothing, and placing the whole knowledge base of about 26,000 tokens in the prompt reached 99.3\%. The tool agent's losses arose in retrieval. Its literal searches returned nothing when the router's arguments did not occur verbatim in a record, for example when it transliterated Greek into Latin script or combined words that occur in a record but not as one phrase, and the agent then abstained. Unaccented and capitalised questions cost the tool agent about 20 points and vector RAG at most 2; accent-insensitive search removed this loss, and matching stemmed tokens raised the tool agent to 83.8\% on canonical Greek. Greeklish cost both designs about 21 to 32 points. Tool interfaces for community knowledge bases need search that tolerates how users type.
cs.AI / 102 / 2610.07339
A doctrine-grounded visual question answering dataset for Tactical Combat Casualty Care
Abstract
Tactical Combat Casualty Care (TC3) requires responders to connect visual observations of injuries and interventions with established clinical guidance. Developing vision-language models to support this process requires supervision that links visible evidence to traceable doctrine. We present TC3-VQA, a dataset constructed from public instructional and field TC3 videos and authoritative TC3 documents. It contains 581 items spanning 11 concepts, with 1,860 questions covering intervention recognition, doctrine, clinical reasoning, procedural guidance, and refusal when visual information is insufficient. Doctrine-based answers preserve verbatim source passages and character offsets. Construction combines visual annotation, passage retrieval, entailment checks, and verification across model families. Equipment boxes, anatomical labels, temporal segments, and source metadata accompany the question-answer pairs. Automated audits and ratings by two physicians and two medical students characterize annotation quality, with human ratings available for 88 retained items. The dataset provides a resource for adapting vision-language models to TC3, studying the connection between visual evidence and clinical knowledge, and evaluating recognition, doctrine recall, and abstention.
cs.AI / 103 / 2610.07384
WildMatch: Weakly Supervised Image Matcher Adaptation for Wildlife Re-Identification
Abstract
Individual animal re-identification from camera-trap imagery is an instance retrieval problem central to non-invasive wildlife monitoring: a query image must retrieve the correct individual from a reference set of known animals. This requires computer vision models to recognize distinctive local patterns in fur, skin, or other visual markings. Current approaches either learn global embeddings as a classification problem, requiring many labeled images per individual while largely ignoring local evidence, or apply off-the-shelf, domain-agnostic image matchers. Although such matchers are pretrained on large and diverse image collections, adapting them to wildlife imagery is challenging because available datasets are small and lack correspondence-level annotations. We study weakly supervised adaptation of a pretrained keypoint matcher using only identity labels, without keypoint-level or geometric correspondence ground truth. We mine informative image pairs with the pretrained matcher, derive weak positive and negative supervision from identity agreement, and contrastively fine-tune the matching network to strengthen correspondences for same-identity pairs and suppress them for different identities. Across open-source wildlife re-identification datasets, our approach improves accuracy over off-the-shelf matchers and a state-of-the-art local--global fusion method. Under an open-world protocol with held-out individuals, it learns a transferable correspondence prior rather than memorizing training identities. To our knowledge, this is the first study of matcher-level, identity-supervised adaptation for animal re-identification. Our method enables data-efficient specialization of image matching models to wildlife domains using identity annotations already available in typical monitoring datasets.
cs.AI / 104 / 2610.07460
ElasticFit: Fit-Aware 3D Object Insertion via VLM Reasoning and Generative Adaptation
Abstract
Inserting objects into existing 3D scenes requires more than selecting a plausible location: the inserted object must also fit local geometry while preserving semantic intent and physical plausibility. Although recent Vision-Language Models (VLMs) and generative models enable semantic reasoning and visual content creation, they offer limited 3D grounding and geometric control when an inserted object must fit into constrained local spaces. We introduce \textbf{ElasticFit}, a VLM-guided framework for fit-aware object insertion centered on a novel scene-grounded representation. Given a language instruction and rendered scene observations, ElasticFit infers structured fitting cues that specify where the object should be grounded, what volume it should occupy, how it should be oriented, and its adaptation mode (rigid placement, uniform scaling, or elastic fitting). These cues convert high-level VLM reasoning into explicit 3D constraints that condition object generation and guide downstream geometric fitting. ElasticFit then generates a scene-conditioned object prior, reconstructs it in 3D, and refines the mesh through mode-specific fitting while enforcing collision avoidance, contact consistency, and physical grounding. In fixed-asset baseline comparisons, ElasticFit improves spatial relation success from 50.8\% to 69.7\% and support success from 48.3\% to 91.7\% over the strongest baseline, while providing novel support for generative "make-it-fit" insertions in complex scenarios.
cs.AI / 105 / 2610.07705
What Frame-Level Labels Can and Cannot Do for Small-UAV Point Detection in Thermal Video
Abstract
The growing use of unmanned aerial vehicles (UAVs) has increased the importance of image-based UAV detection. Learning-based detectors are trained on imagery and annotations, with annotation type determining the information available during training. We focus on learning localization from frame-level target presence/absence labels when sensor or scene changes make spatial annotations for additional training burdensome. We analyze the detection capability, learning behavior, and potential applications of an existing architecture for point detection of small UAVs, trained with presence/absence labels and requiring no external detector. The architecture freezes spatial features learned through classification and trains a readout with the same frame labels to produce spatial score maps and point detections. On two thermal infrared datasets, CST Anti-UAV and Anti-UAV410, we evaluate localization hit rates and detection rates under false-alarm constraints, analyze the effects of training stages, label allocation, synthesis, and model configuration, and compare with bounding-box detectors. We also explore potential applications on Airborne Object Tracking (AOT) using its visible-light imagery and frame labels. Classification training strengthened target-related spatial responses, while readout training helped extract them consistently. Distributing similar label counts across more videos yielded higher localization hit rates, while synthesis effects varied by dataset and evaluation criterion. Higher localization hit rates did not always improve detection under false-alarm constraints, and failures remained when target signals were weak relative to background variation and under cross-dataset transfer. These findings provide guidance on label allocation, spatial representations and readouts, synthesis, and false-alarm control.
cs.AI / 106 / 2610.07758
Later Is Better: Token Reduction for ViTs Under Distribution Shift
Abstract
Training-free token reduction accelerates vision transformers by removing redundant tokens across layers, recovering most of the original accuracy at a fraction of the compute. These methods, however, are designed and evaluated primarily on clean data, and under real-world distribution shift their accuracy gap to the uncompressed model widens with the removal rate. We show that this gap is governed by the reduction schedule, the depth profile of removal, usually left fixed as an implementation detail. Concretely, we introduce a one-parameter late-concentrated power-law schedule that consistently improves out-of-distribution accuracy over flat at no extra inference cost. On ImageNet-C with DeiT-S, the late schedule closes 83% of that gap at a 26% compute reduction (+1.17pp), and 99% of it at a lighter 7% reduction (+0.26pp). The gain cannot be attributed to retaining more tokens or using extra compute: held to flat's compute, the late schedule removes more tokens in total and leaves fewer tokens at the end, yet still wins. Single-layer probes point to a mechanism: earlier reductions perturb features that pass through more remaining layers, front-loading reduction error in depth. The effect is broad, holding across five token-reduction methods (ToMe, EViT, ATS, ATC, PiToMe), nine backbones, all ImageNet-C corruption types, eight further shift suites, and two further modalities, video and vision-language QA. It is also specific to shift, still positive on clean and rising monotonically to ~4x that at the highest severity 5. The schedule keeps its gain under six test-time adaptation methods, and needs no per-input or per-domain tuning.
cs.AI / 107 / 2610.07911
Diverse Motion Customization via Control-based Dynamic Optimization
Abstract
Despite recent advances in video generation, motion customization remains challenging due to content leakage, where appearance attributes from the reference video unintentionally propagate into the generated output. We identify this issue as a consequence of the generative process collapsing toward the reference video, which arises from formulating the learning objective as a direct regression on the reference. To address this, we propose Control-based Motion Customization (CMC), a principled training framework that is structurally robust to content leakage. Our key idea is to steer generative dynamics toward desired motion while avoiding collapse toward the reference video, which we formalize using Stochastic Optimal Control (SOC). Under this formulation, customized videos acquire the target motion yet remain within the pre-trained model's prompt-conditional distribution, where appearance is determined by the text prompt rather than the reference video. Furthermore, to improve efficiency, we tailor the SOC formulation to motion customization by eliminating the need for an explicit reward and introducing a timestep-adaptive motion cost that focuses only on early generative stages, accelerating training by 2.5 times. Extensive experiments demonstrate that CMC effectively mitigates content leakage and achieves competitive motion fidelity while preserving the diversity of the base model across diverse scenarios.
cs.AI / 108 / 2610.07928
Dynamic Alignment and Calibration for Multimodal Learning
Abstract
Dynamic multimodal learning aims to learn robust representations by adaptively modeling information discrepancies across modalities. However, existing methods still suffer from two limitations: (i) static cross-modal alignment strategies usually impose uniform constraints on all samples while overlooking sample-wise variations, potentially leading to unreasonable over-alignment; and (ii) confidence- or uncertainty-aware fusion methods often fail to adequately account for feature magnitude and confidence differences across modalities. For modality pairs with significant feature magnitude differences or small confidence gaps, it might be unreliable to strictly align fusion weights according to confidence. To address these issues, we propose an Alignment- and Calibration-driven Multimodal Learning framework (ACML). Specifically, ACML incorporates a dynamic cross-modal triplet alignment module, which enforces strong semantic consistency for high-confidence positive pairs while encouraging diverse representation learning between high- and low-confidence positive pairs according to their confidence gaps. Additionally, ACML introduces a difference-aware attention calibration strategy that adaptively adjusts attention regularization based on feature magnitude and confidence differences across modalities, thereby mitigating biases caused by unreasonable fusion constraints. Extensive experiments on multiple multimodal benchmark datasets demonstrate that ACML consistently achieves superior performance and robustness over recent state-of-the-art methods.
cs.AI / 109 / 2610.07984
Decide Before You Look: Learning Which Retrieved Memories Deserve Pixels
Abstract
Multimodal assistants answer questions from long-term memories that contain images. After retrieval, each retrieved image reaches the answering model either as pixels, at about a thousand visual tokens per image, or as a stored text proxy that often misses the detail the question asks about. We find that the benefit of pixels usually comes from one or two retrieved memories, and that it can be predicted before the answering model runs, without reading any full-resolution image. In PixelTriage, a plug-in placed after retrieval, a small model that does not generate text reads the dialogue, a short note and a thumbnail of each retrieved memory and predicts how much its pixels would add. It is trained on synthetic memory episodes labeled by a frozen 27B model that answers each question with and without each memory's pixels. With a 7B answering model, PixelTriage lies on the accuracy--cost frontier of M$^3$Exam, DMV and MemEye and uses 11--23\% of the visual tokens without a significant loss of accuracy. On DMV it answers 2.9 times faster than opening all images. It outperforms retrieval order and uniform down-sizing at equal budgets and transfers to other memory systems and to a 397B answering model.
cs.AI / 110 / 2610.08126
Supermarket Product Detection and Recognition: Utilizing Deep Learning with Rectified Imagery
Abstract
Product Identification has sprung up to become one of the most challenging problems in the automation of the retail industry. With the new industry 5.0 standards, automated inventory management, and catalog creation tasks are vitally important. Object identification models have emerged as a viable answer with their unprecedented identification and localization accuracy. However, the close-knit rack design of supermarkets generates the problem of angle variation in capturing images. The angle-variant densely packed images(a single image contains many objects) become overwhelming for these models alone. In this paper, we try to supplement object detection models with traditional Hough transform (HT) and homogeneous estimation concepts. We study the effect of rectified images using homography estimation and hough transform and their limitations on the problem of grocery identification. We make a case for creating a new dataset to test the effects of such rectification and produce analytical results on different scenarios of angle variation and object densities per image. Extensive experiments on different object detection models suggest that image rectification of angled images improves the detection accuracy of grocery products in images. The results also highlight the limitation of rectification on the angle of image capture and the object density of the image.
cs.AI / 111 / 2610.08331
Transferable Spatial Temporal Coherence Adversarial Attack on Black-Box Vision Language Models for Autonomous Driving
Abstract
The rapid integration of Vision Language Models (VLMs) into sensitive systems introduces critical safety vulnerabilities that remain unexplored in exist studies. While adversarial attack robustness has been extensively studied for image-based models, the susceptibility of VLMs to temporally-aware adversarial attacks against video in driving context poses a distinct and under examined threat. In this paper, we introduce novel adversarial attack against video targeting VLM models used for autonomous driving scenes named Spatial Temporal Coherence Adversarial Attack (STCA). Our attack comprise from three stages: modalities expansion, Spatial attack, and STCA attack. In modalities expansion, we propose caption-guided frame selection method in order to ensure that adversarial perturbation target the most semantically significant frames. Secondly.In spatial attack, we craft effective perturbation and preserve high similarity. Then the perturbed video generated fed into STCA stage that disrupt cross-frame temporal coherence using motion guided mask. Our method operate under black box threat model against victim target VLMs, relying solely on transferability from white-box surrogate model.We conduct our experiments on the BDD100K and nuScenes autonomous driving datasets across three VLM models: Video LLaVA-7B, Qwen2.5-VL-7B, and Dolphin. Experimental results demonstrate spatial attack achieves an ASR with high SSIM. Our finding reveal that existing video language model, remain highly susceptible to adversarial attack in autonomous driving scenarios, underscoring the urgent need for robust defense for VLM models.
cs.AI / 112 / 2610.08358
Test-Time Adaptation of Quantized ViTs via Single-Pass Quantizer-Aligned Recalibration
Abstract
Post-training quantization is a standard route to fitting vision transformers (ViTs) into edge compute and memory budgets, yet quantized models become especially brittle under distribution shift. Test-time adaptation (TTA) addresses such shifts without labels, but most existing approaches are poorly aligned with the constraints of quantized inference. Prevailing TTA methods recover accuracy through backpropagation, while backprop-free methods often still incur overhead from extra forward passes or parameter updates, and lightweight feature- or logit-level methods recover only part of the loss. Across these approaches, a quantization-specific failure mode that amplifies the drop is not directly targeted: under shift, activations occupy frozen quantizers' calibrated ranges differently, distorting their code distribution. We propose Quantizer-Aligned Recalibration (QuAR), a single-pass TTA method tailored to quantized ViTs that neither backpropagates nor updates any model parameters. QuAR recalibrates activations at the input to a frozen quantizer, mapping the test stream's running per-channel statistics back toward the source calibration. On ImageNet-C with ViT-B, QuAR achieves the highest mean accuracy among state-of-the-art backprop-free TTA methods at 3-, 4-, 6- and 8-bit weight/activation precision, outperforming the strongest baseline by 2.28 points at 8 bits and 4.00 at 3 bits, with 46% lower latency and a memory overhead of only 0.17 MB (0.01% of peak inference memory). Analysis and diagnostics trace the gain to a reduced per-channel mismatch at these quantizers, which restores the code distribution the baselines leave unchanged or distort further. A single fixed configuration remains ahead across continual streams, non-i.i.d. label shift, seven out-of-distribution suites, and three other backbones.
cs.AI / 113 / 2610.08401
GeoPID: Decomposing and Steering Visual Information in Vision-Language Models
Abstract
While recent vision-language models (VLMs) have shown outstanding performance across diverse applications, they tend to under-use visual information and over-rely on textual context. In this work, we propose \textsc{GeoPID}, a training-free framework that analyzes multimodal information within VLMs from a geometric perspective. \textsc{GeoPID} decomposes information into Redundant, Modality-Unique, and Synergistic components through the geometric relationships between visual and textual representation subspaces. Through an extensive analysis across 22 VLMs and 14 benchmarks, we confirm that correct predictions exhibit stronger vision-unique components when questions strongly require visual grounding. Building on this geometric analysis, we introduce a targeted intervention technique that selectively amplifies visual representations along the vision-unique subspace during inference. As a result, visual grounding capabilities were enhanced without any additional model parameter updates, achieving an average relative accuracy gain of 7.63\%.
cs.AI / 114 / 2610.08482
Knee3DVLM: Dual-Sequence Full-Volume Vision-Language Modeling for Comprehensive Knee MRI Assessment
Abstract
Vision-language models (VLMs) are increasingly being applied to three-dimensional medical imaging, but their application to knee MRI remains limited, particularly for interpreting the complementary sequences used in clinical practice. We introduce Knee3DVLM, a sequence-aware VLM that uses full-volume DESS and fluid-sensitive TSE MRI to predict 57 anatomically resolved binary diagnostic targets derived from the MRI Osteoarthritis Knee Score (MOAKS) for structured reporting. We evaluated DESS-only, TSE-only, and paired DESS-TSE configurations using subject-disjoint Osteoarthritis Initiative partitions. In a held-out cohort of 1,074 examinations, the fused model achieved 72.98% average accuracy, 71.17% balanced accuracy, 78.96% mean ROC-AUC, and 78.74% macro ROC-AUC, the highest values among the three configurations. In a secondary multiclass analysis aligned with the released 3DReasonKnee cohort, Knee3DVLM was numerically higher than the strongest reported 3DReasonKnee configuration across five pathology categories. These findings support dual-sequence full-volume modeling for comprehensive knee MRI assessment.
cs.AI / 115 / 2610.08528
MedCORE: Criteria-Grounded Clinical Reasoning for Interpretable Medical Image Diagnosis
Abstract
Clinical diagnosis is inherently a structured reasoning process, yet existing deep learning models often bypass this structure by mapping image features directly to disease labels without explicitly interrogating the morphological and textural criteria that clinicians systematically evaluate. This limits diagnostic transparency and may compromise safe clinical deployment. We present MedCORE (Medical Criteria-Oriented Reasoning and Evidence), a structured diagnostic framework that operationalizes clinical reasoning within a vision-language architecture. For each input image, MedCORE decomposes the diagnostic process into clinically defined criteria, spatially localizes each criterion to diagnostically relevant image regions, encodes evidence through multi-scale representations that capture macro-structural and micro-textural pathological characteristics, and refines criterion representations using a Graph Attention Network that explicitly models inter-criteria dependencies. Criterion representations are further aligned with clinical text descriptors, reinforced through class-wise visual prototypes, and aggregated using uncertainty-calibrated weighting that proportionally discounts low-confidence diagnostic evidence. MedCORE is validated across three clinically heterogeneous imaging modalities, including dermoscopic lesion classification on ISIC 2018, breast ultrasound lesion characterization on BUSI, and diabetic retinopathy grading on IDRiD. Quantitatively, MedCORE achieves 89.2% accuracy, 85.7% macro-F1, and 96.4% AUC on ISIC 2018; 96.1% accuracy, 95.2% macro-F1, and 98.4% AUC on BUSI; and 84.3% accuracy, 80.2% macro-F1, and 92.8% AUC on IDRiD. These results demonstrate consistent improvements over strong CNN, transformer, biomedical vision-language, concept-based, and prototype-based baselines.
cs.AI / 116 / 2610.08574
FedDermaSeg: Federated Learning for Dermatological Image Segmentation
Abstract
Skin cancer is a major global health concern, and early detection and accurate lesion delineation are important for effective diagnosis and treatment planning. Automated skin lesion analysis can assist dermatologists, with lesion segmentation serving as a fundamental step in computer-aided diagnostic systems. Conventional deep learning-based segmentation models typically rely on centralized training, where images and their corresponding segmentation masks are collected on a central server. Such data aggregation raises privacy concerns in medical applications and requires substantial centralized computational resources. To address these limitations, we investigate the feasibility of federated learning for privacy-preserving skin lesion segmentation. The training and validation sets of the ISIC 2018 Skin Lesion Segmentation Challenge dataset are used to simulate a distributed learning environment and develop a federated segmentation model. The resulting model is evaluated on the ISIC 2018 test set and the PH2 dataset to assess its performance and generalizability. Experimental results demonstrate that the federated model achieves performance comparable to centralized training while consistently improving upon the locally trained models. These findings demonstrate the potential of federated learning for collaborative skin lesion segmentation without requiring centralized aggregation of medical images.
cs.AI / 117 / 2610.08659
Selective Transfer of RL Updates for Visual Reasoning
Abstract
Model merging provides a training-free way to transfer reasoning capabilities from language models to vision-language models (VLMs), but endpoint-based transfer can conflate pre-existing model differences with changes acquired during reasoning post-training. We instead formulate capability transfer around the training-stage update, isolating the parameter changes induced by reinforcement learning (RL). Yet transferring this update in full remains suboptimal: we find that its components differ substantially in cross-model transferability, with dominant directions transferring more effectively than the complete update. Based on this finding, we introduce Selective-RL, which isolates the RL-stage update, retains its dominant matrix-wise directions with magnitude preservation, and transfers them to the language modules of a VLM. Across three model families and five visual-reasoning benchmarks, Selective-RL improves full-update interpolation in 12 of 15 comparisons, including an 8.55 percentage-point MathVision gain on the Qwen recipient. Matched controls show that update magnitude or arbitrary low rank alone does not reproduce these gains. These results highlight a distinction between what is acquired during post-training and what remains transferable across models, providing a training-stage perspective on cross-model capability transfer. Code is available at https://anonymous.4open.science/r/selective-rl.
cs.AI / 118 / 2610.08782
4D-HOF: Hand-Object Flow Matching for Feed-Forward 4D Interaction Reconstruction
Abstract
Existing methods for 4D hand-object reconstruction often rely on costly per-sequence optimization, while generative approaches typically synthesize interactions from random noise, which can lead to unstable interaction prediction. We introduce 4D-HOF, a feed-forward framework that reconstructs 4D hand-object interactions from coarse but informative estimates produced by vision foundation models. Concretely, we learn a conditional flow matching model that transports foundation-model-derived hand-object states toward an interaction manifold, allowing the model to correct errors in translation, rotation, and alignment in a feed-forward manner. A key advantage of our generative formulation is that it naturally enables test-time guidance within the transport process. Rather than applying a separate post-hoc optimization after reconstruction, we directly steer the evolving generative states using physical interaction constraints and observed 2D evidence, allowing the reconstruction to be refined as part of the generative process itself. By training the generative model on diverse datasets, 4D-HOF generalizes robustly to challenging in-the-wild scenarios. Experiments on out-of-domain benchmarks show that 4D-HOF achieves state-of-the-art performance, producing more stable and accurate 4D hand-object reconstructions.
cs.AI / 119 / 2610.07476
Can Power Draw Constrain Covert Compute? Limits of Analogue Verification for AI Governance
Abstract
Frontier AI treaties or agreements on limiting computation require external verification; an external auditor must be able to confirm how much computation actually ran and that parties are adhering to the agreement. Analogue, off-chip measurements such as power draw provide an information channel for verification. It is unknown how well these analogue channels can constrain computation against an adversary who actively tries to subvert the audit. We derive a closed form for $β$, the largest hidden computation a power trace cannot exclude, as a fraction of the declared machine capacity. Measurements on NVIDIA A100 GPUs constrain $β= 1.16$ in the worst case, while adversarial matched-energy strategies are shown to hide at least $β= 0.41$ of compute. Analogue power measurements alone therefore constrain compute weakly. Additional restrictions granted by the threat model, such as the ability of the verifier to re-execute the declared work at an observed operating point, let the verifier push $β$ down to $0.059$ in the maximally restricted case. This gives a quantitative estimate of what analogue measurements can contribute to compute verification.
cs.AI / 120 / 2610.08089
When Plans Change Answers: Formalizing Cost-Accuracy Optimization for Semantic Queries
Abstract
In semantic query engines, predicates are evaluated by machine-learned models, and the choice of a query plan affects not only the cost of a query but also its result. Existing systems either apply a fixed threshold to each semantic operator or tune accuracy per operator, without accounting for how errors propagate through joins. We give a formal problem definition for cost-accuracy optimization of such queries. Our starting point is the calibrated confidence that decision models such as Jev attach to each decision. It yields an expected error for every decision; weighting these errors by each decision's contribution to the output (in the simplest case, its fan-out) gives the expected output quality of a plan without any labeled data, and the same computation in reverse turns an output-level accuracy target into a price on each base or intermediate tuple. Building on this, we define an oracle semantics for relational algebra with semantic operators, physical plans as pairs of a logical plan and a decision policy, declarative output-level targets, and a hierarchy of plan equivalence. We show that accuracy is plan-invariant under pointwise-deterministic policies, and that selection pushdown is not quality-sound when escalation bands are calibrated on the plan's own candidates. Expected quality can be computed in polynomial time under bag semantics; under set semantics it follows the dichotomy of tuple-independent probabilistic databases when every relation carries a semantic predicate. Choosing which tuples to drop is NP-hard, while the optimization problem decomposes into per-tuple decisions through two Lagrange multipliers. Simulations on a synthetic workload illustrate these effects; an evaluation on real engines is left for future work.
cs.AI / 121 / 2610.07333
Memory-Efficient Expert Routing for Distributed MoE Training
Abstract
As Mixture-of-Experts (MoE) models scale toward hundreds of experts and higher top-$k$ routing, memory efficiency in distributed training becomes a critical bottleneck. Peak memory is dominated by the MoE block, not attention: every intermediate buffer in the MoE dispatch pipeline is individually scaled by top-k routing. The standard all-to-all dispatcher sends all routed tokens in a single collective step, requiring the full top-$k$-expanded buffer to be constructed at once. In this work, we propose RelayMoE, a ring-based MoE execution model that computes locally as expert weights or tokens circulate, avoiding full top-$k$-expanded dispatch buffers. RelayMoE selects between expert and token routing according to communication volume and overlaps transfers with computation. The ring structure naturally supports memory-efficient MoE recomputation during backward: each hop reconstructs expert intermediates, uses them to compute gradients, and releases them before the next hop. The saved memory supports longer sequences and larger batches, or retains more attention activations to reduce attention recomputation and improve training throughput. We evaluate RelayMoE on 30B$-$57B production MoE models and varied expert configurations. In single-layer MoE experiments, RelayMoE achieves a $2\times$ average speedup over Megatron-LM. In full-model training under the same GPU memory budget, it improves throughput by up to $2.02\times$ and extends the largest tested trainable sequence length by up to $2.85\times$.
cs.AI / 122 / 2610.08430
NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale
Abstract
Agentic reinforcement learning (RL) disaggregates training from rollout, so each policy update must reach the rollout clusters before the next batch. Transferring a full 1T checkpoint for such weight synchronization (refit) takes 87.5 min between two AWS regions. Measurements of BF16 training show that about 1% of weights change their stored values per step. Recent systems exploit this sparsity but fall short on placement, exactness, or efficiency: they reimplement placement rules, assemble full tensors, rebuild values arithmetically, or use a cross-cluster collective, and none fully recovers from mid-refit failures. We present NeMo-DCR (Delta-Compressed Refit), which sends only changes yet is bit-exact: receivers obtain the same parameter and buffer bits as a dense refit. For placement, fixed affine mappings project changes from training shards into the checkpoint's canonical coordinates, residual conversion covers the other changes, and the serving runtime's native loader places all changes in receiver storage. For exactness, compressible XOR masks carry affine changes whose projection and loader preserve stored bits, and overwrites carry the others. Receivers apply both in place, retries overwrite partial writes, and a joint commit binds the policy to the baseline for the next delta. For efficiency, object storage or a relay tree streams payloads during delta construction, without a cross-cluster collective. Even at 3% and 5% change rates, NeMo-DCR refits of 30B-1T models are 12-40$\times$ faster than a transport-only full-checkpoint reference. A 1T relay-tree refit at 3% takes 150 s instead of 87.5 min, making refits practical for cross-cluster agentic RL at trillion-parameter scale.
cs.AI / 123 / 2610.07204
SPEAR: Five Principles for Interactive Human-Agent Alignment
Abstract
Recent AI alignment work often frames alignment as a pre-deployment optimization problem: collect human feedback, learn preferences or principles, finetune the model, and deploy an aligned system. This framing has produced major progress, but it under-specifies what happens once AI systems act as agents on users' behalf in situated, long-term, and social contexts. This position paper reframes human-agent alignment as an ongoing interaction design problem. We propose SPEAR, five pillars of interactive alignment: Specification (how people express intent and establish shared understanding), Process (how agents decide when to act, ask, defer, or pause), Evaluation (how people judge whether agents succeeded), Adaptation (how agents adapt to users over repeated use), and Recalibration (how people adapt their trust, expectations, and behavior in response to agents).
cs.AI / 124 / 2610.07205
Responsible Institutional Analytics: Interpreting Bias with AI Support
Abstract
Institutional Analytics (IA) dashboards inform decision-making in higher education, yet data limitations, constraints in analytical techniques, and missing contextual information often affect their interpretation. To support more responsible interpretation of IA, we introduce FACTRIA, a framework that organizes potential biasing factors across four areas: the analytics pipeline, institutional context, course-level characteristics, and demographics. We used the FACTRIA framework as input to a generative-AI chatbot designed to prompt users to reflect on these factors while analyzing IA. A qualitative study with stakeholders, drawing on four authentic IA cases, and a transition network analysis showed that the chatbot prompted participants to recognize how overlooked factors influenced their initial interpretation. Findings indicated that combining a structured framework with AI-based guidance can enhance context-aware, responsible interpretation of institutional data.
cs.AI / 125 / 2610.07506
Jarvis: A Proactive Speech Agent for Multi-Party Conversations
Abstract
Speech agents are reactive and dyadic: they speak when spoken to, and to one person at a time. We ask what it takes for a speech agent to instead take part in a conversation among several people and speak up only when it can help. We introduce Jarvis, a real-time proactive speech agent that audibly participates in multi-party human conversations. Grounded in a document shared beforehand, Jarvis follows the discussion and intervenes when the group misses or misstates a fact and does not correct itself within a few turns. We make three contributions: a problem setting based on epistemic breakdowns that makes proactive intervention measurable, realized as CHI-180-proactive, a synthetic multi-party dataset seeded with known gaps, errors, and self-corrections; a proactive backbone that harnesses a small, open-weight model with deterministic checks and grounds every claim in a source sentence; and interaction techniques for taking the floor in live speech and showing the cited evidence on screen. On CHI-180-proactive, Jarvis is correct on most events it addresses and stays silent 97% of the time when the group resolves an issue itself. A live study with 23 participants confirms these trends with real-time interventions.
cs.AI / 126 / 2610.07603
Emoception: Selective Affective Layer Fine-Tuning of Video Vision Transformers for Player Arousal Change Recognition From Gameplay Footage
Abstract
This article proposes Selective Affective Layer Fine-Tuning (SALFT), an efficient adaptation framework for Video Vision Transformers in player arousal recognition from gameplay. To bypass computationally expensive full fine-tuning, SALFT introduces a selection criterion based on the L2-norm change in layer parameters after brief adaptation, directly measuring representational shifts and providing a more stable basis than gradient-based alternatives. Evaluated via five-fold cross-validation on the Arousal Video Game AnnotatIoN dataset, SALFT achieves performance comparable to full fine-tuning across all games without statistically significant degradation ($p>0.05$), while updating only $\approx$8% of parameters (over 92% reduction). Notably, in one game, SALFT consistently outperforms both full fine-tuning and the best baseline across all metrics and folds, reaching the theoretical minimum p-value (p=0.0625, exact two-sided Wilcoxon signed-rank test). In addition, we introduce an interpretability method to trace attention patterns, enhancing model transparency. These results establish SALFT as an effective and efficient approach for affective game computing.
cs.AI / 127 / 2610.07666
SENSE: State-aware Emotion Navigation Storytelling Engine
Abstract
This paper presents SENSE, a state-aware framework for generating playable branching visual novels with multi-track emotional navigation. Integrating a state-based narrative architecture called MIND, a structure analyzer, and a path-aware context management module, SENSE produces narratives that are both structurally coherent and emotionally rich. From minimal high-level inputs, it generates multiple intersecting routes while preserving character consistency and narrative causality. Evaluations using LLM judges, affective metrics, and visual assessments indicate SENSE outperforms baselines in narrative diversity and robust asset integration, while preliminary human trials show directional improvements in emotional fidelity alongside comparable enjoyment.
cs.AI / 128 / 2610.07669
Evaluating human-AI workflows for field research in viticulture
Abstract
We assessed the value of two live human-AI interactions in a precision disease control project in California vineyards. The project tested whether 2021-2024 commercial scouting records and remote-sensing measurements across 140 hectares could support 2025 red-leaf symptom forecasting for prioritized scouting and virus testing. In Workflow 1, Aleks v1, a multi-agent research system, developed forecasting models with iterative human refinement. We applied Aleks's 2024 vine-scale model to updated 2025 predictors and evaluated red-leaf forecasts against independent 2025 scouting. In retrospective simulations surveying 45% of all vine positions, adding model-informed row prioritization to adaptive scouting increased the encountered proportion of newly recorded red-leaf observations from 85.8% to 94.1%. Within-block scouting comparisons suggested the model mainly improved scouting allocation among blocks. Despite unreliable internal 2024 performance estimates from synthetic oversampling before train/test splitting, Aleks developed an informative vine-scale model in 145 minutes, increasing throughput and answering our research questions. In Workflow 2, we assessed whether higher model-score vines had more frequent virus detection, and whether Aleks could infer this sampling goal from a general prompt with data and literature. Aleks's plan prioritized balanced vineyard and model score coverage, while our plan prioritized field efficiency and high-model-score oversampling. Aleks's and our plans yielded 41/50 (82%) and 97/100 (97%) sampled vines. Aleks's plan omitted instructions for replacing missing vines, limiting implementation and operational value. Five of 137 sampled vines tested positive for grapevine red blotch virus (model score ROC AUC 0.735). These findings support assessing AI interactions by how well they advance field research objectives under live, project-specific constraints.
cs.AI / 129 / 2610.07800
Novice Reliance Calibration in AI-Assisted Decision Making: The Role of Explanations and Self-Assessment
Abstract
Artificial Intelligence (AI) tools are widely used to support decision making in tasks and domains where no immediate performance feedback is available. In these settings, users cannot learn to adjust their reliance behavior over time through trial and error. However, little is known about how novice users calibrate reliance on AI when external feedback is unavailable, or whether AI explanations can support calibration in its absence. We introduce reliance calibration as an organizing construct for studying how novice users dynamically adjust reliance behavior, and examine how AI explanations and meta-cognitive self-assessment shape it. Through a between-subjects study with 110 participants completing a clinical entity extraction task with AI assistance and limited performance feedback, we observe that novice users exhibit systematic drift toward over-reliance in the presence of explanations, while higher self-reported task understanding is associated with more selective reliance behavior. These results extend reliance calibration research into human-AI collaboration contexts without real-time performance signals and present actionable guidelines on designing AI tools that must support appropriate reliance in these settings.
cs.AI / 130 / 2610.08554
Systemization of Knowledge (SoK): Human-Centered AI Safety for Youth
Abstract
While HCI increasingly examines AI-safety for youth, the literature lacks a comprehensive view of what risks have been identified, how they are addressed, and whether proposed protections work in-practice. We systematically reviewed 100 empirical HCI studies involving children and youth interacting with or exposed to AI across schools, homes, care settings, and public services. Using the YAIR taxonomy for risks and the MIT Mitigation Taxonomy for countermeasures, we map which risks have been identified, whether each risk is addressed by countermeasure(s), and whether each countermeasure for that risk is implemented and even evaluated. The risk-countermeasure mapping shows that most risks are matched only with proposed/ideated countermeasures; few countermeasures have been implemented, and fewer still evaluated; and existing evaluations often measure technical performance rather than protection from harm. We identify where coverage is absent, where safeguards remain untested, and propose concrete directions for HCI research to strengthen youth AI-safety.
cs.AI / 131 / 2610.07731
Learning to Retrieve via Reinforcement Learning in Embedding Space
Abstract
Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinforcement learning framework that enables existing embedding models to learn to retrieve directly in embedding space and align to task-specific rewards. We train RELER by sampling unit-length query and document embedding actions from von Mises-Fisher (vMF) distributions centered on normalized encoder outputs, scoring the resulting retrieval or downstream outcomes as rewards, and updating the encoder with REINFORCE using a leave-one-out baseline (RLOO). As exploration in the high-dimensional embedding space is prone to sampling noise, we further propose conditional-mean projection (CMP), which projects each sampled embedding onto the low-dimensional subspace spanned by its encoder output and the candidate embeddings it is compared against, reducing noise in the policy gradient while preserving its expectation. We evaluate RELER on BRIGHT, a benchmark with reasoning-intensive queries that remain challenging for existing embedding models. RELER consistently outperforms InfoNCE and LambdaLoss in average nDCG@10 when post-training BGE-M3 and Qwen3-Embedding backbones. We further evaluate downstream utility through retrieval-augmented generation (RAG), where we adapt only the query encoder while keeping the document index and generator fixed. Across seven QA datasets, jointly optimizing retrieval and answer rewards improves both average retrieval performance and answer quality in RAG.
cs.AI / 132 / 2610.07761
Contrastive Learning for Aspect Representation towards Explainable Recommendation
Abstract
In this work, we propose a novel recommendation model, CLARER (Contrastive Learning for Aspect Representation towards Explainable Recommendation) that integrates aspect features learned from textual reviews with rating information to improve the accuracy and explainability of recommendations. Our proposed framework learns user and item representations by combining rating-based features and aspect-based features from reviews. Specifically, rating-based features are learned through a multi-layer perceptron (MLP) model, while aspect-specific review representations are learned using a transformer encoder to capture the semantic information and contrastive learning to better distinguish user preferences. To provide explanations, we train a transformer decoder, using the final representations of users and items from both rating and aspect-based features as context. Experimental results in three benchmark data sets demonstrate that our model achieves superior performance compared to baseline methods in both recommendation (accuracy) and explanation generation.
cs.AI / 133 / 2610.07900
IEEE 802.11bx - WLAN Intelligent Networking (WIN): Toward an AI-Ready Wi-Fi 9
Abstract
Wi-Fi 9 is expected to go beyond mere communication and provide new services such as sensing or computation. At this juncture, Artificial Intelligence (AI) is taking a leading role in the definition of the 802.11bx amendment, named WLAN Intelligent Networking (WIN). In this tutorial, we survey the recent progress made toward Wi-Fi 9 within IEEE 802.11 standardization, tracing the drivers and technological advances that motivate an AI-ready Wi-Fi 9. We then examine AI's role along three complementary dimensions, i.e., AI as a protocol (AI is applied to Wi-Fi's PHY/MAC operation), AI as a platform (Wi-Fi infrastructure is repurposed to provide AI computation), and AI as traffic (AI flows call for new traffic-handling policies), and discuss candidate features and open challenges along each. As a concrete illustration of the AI as traffic paradigm, we present a case study on AI traffic differentiation, where we explore a potential extension of the current Enhanced Distributed Channel Access (EDCA) to support new AI traffic flows.
cs.AI / 134 / 2610.08622
Agentic RCA for Internet-Scale Services Using Constrained Creativity
Abstract
System administrators of Internet-scale services need to resolve failure incidents to maintain reliability of such services. Ideally, we want a troubleshooting system to be: (1) expressive to known and unknown incidents with high accuracy; (2) cost efficient at scale; (3) explainable to provide actionable insights operators can act on; and (4) entail low effort from the operators. Unfortunately, most existing systems, including emerging LLM-assisted agentic workflows and structured frameworks for authoring diverse RCA algorithms fall short of achieving all four requirements. We present E4, a novel agentic system for troubleshooting for Internet-scale services. E4 embodies the paradigm of constrained creativity that combines the best of LLM-assisted automation and exploration with the explainability and efficiency of a structured approach. Instead of allowing an LLM agent to write arbitrary code or generate arbitrary responses, we provide the agent a restricted DSL to generate its response via simple loop-free data flow programs. This DSL, equipped with high level operators for troubleshooting, makes E4's output accurate, verifiable and explainable. On a mix of synthetic and real-world workloads, E4 achieves up to 62% better accuracy compared to state-of-the-art solutions, while providing more explainable responses at up to 12x reduced cost.
cs.AI / 135 / 2610.07742
Cleave: Scaling Tensor Program Optimization via Decoupled Algebraic Search and Operator Scheduling
Abstract
Optimized kernels such as FlashAttention and FlashDecoding are crucial for accelerating today's large models. Most of them are handwritten by experts because existing ML compilers cannot match their efficiency. Producing such kernels requires fusing computations with multiple reductions, which requires both algebraic transformation of the computation graph and operator scheduling of the transformed graph. Unfortunately, searching the two jointly yields a space too large to navigate. We propose Cleave, an ML compiler built on symbolic decoupling: Cleave discovers transformations by performing superoptimization on a graph with symbolic shapes, and then schedules each resulting graph on concrete shapes. Representing shapes as symbols makes equivalence checking cheap and lets a new Split operator, with a symbolic split count, parallelize along a reduction dimension. Cleave's scheduler fuses graphs with multiple reductions through iterative tiling and horizontal fusion. Evaluation on common LLM subgraphs shows that Cleave generates kernels up to 2.8x faster than the best baseline (1.6x on average) and reduces compilation time by 5.9x on average compared to Mirage. For dynamic workloads captured from production serving traces, Cleave compiles each operator once and achieves geometric mean speedups of 1.4x and 1.7x over FlashInfer's handwritten FA2 and FA3 backends. Cleave's code is available at: https://github.com/nyu-systems/cleave
cs.AI / 136 / 2610.07217
RoboCap: A New Platform for Egocentric Robot Learning
Abstract
Despite its promise for scaling robot learning, egocentric manipulation data is still scarce today. Collection at scale requires vertically integrating ergonomic hardware with centimeter-precise 3D algorithms, at a precision that has not been publicly demonstrated. To address this gap, we introduce RoboCap, a 250\,g six-camera dual-IMU hat designed for in-the-wild egocentric data capture, and the Grounded API, a suite of device-agnostic 3D algorithms tuned for RoboCap. In this report, we demonstrate how hardware, calibration, and 3D algorithms interact to achieve state-of-the-art performance on the public benchmarks: our SLAM across diverse settings and rigs, our depth estimation on egocentric settings, and our hand tracking when adapted to third-party devices.
cs.AI / 137 / 2610.07558
Seeing the Invisible: Physics-Guided Visual Prompting for Temperature- and Radiation-Aware VLA Navigation
Abstract
Vision-Language-Action (VLA) models have become a major paradigm for Vision-and-Language Navigation (VLN). However, in safety-critical facilities, invisible risks such as radiation or temperature spikes cannot be detected by an RGB camera, and handling each risk is expensive, requiring a new encoder, new data, and model retraining. We propose Physics-Guided Visual Prompting (PG-VP), a plug-and-play multimodal perception module that instead reuses what a frozen VLA model already does well: avoiding visible obstacles. Given a proximal radiation or thermal source, PG-VP performs a physics-guided risk assessment to determine the avoidance direction and overlays a corresponding virtual obstacle that moves across consecutive frames (Dynamic Visual Prompting). The navigation policy then naturally detours around this invisible hazard. The identical virtual obstacle is used regardless of hazard type, so the visual prompting pattern remains fixed as sensors are added. When no hazard is detected, nothing is rendered, and the policy behaves exactly as it would without PG-VP. We evaluate PG-VP on OmniNav using the val-unseen splits of R2R-CE and RxR-CE, where it guides the policy toward intended low-risk actions in 84.9% and 83.2% of cases, at a cost of 6.8 and 7.9 percentage points in navigation success rate. We further test it with distinct scenarios on a real robot in the presence of actual thermal and radiation sources, all without any retraining. The real test shows that PG-VP effectively avoids these invisible hazards, improving worst-10% average trajectory safety by 63.45% and 32.59% against thermal and radiation sources, respectively.
cs.AI / 138 / 2610.07569
OpenSplatGraph: From Dense Semantic Maps to Structured Scene Graphs for Open-Vocabulary Robot Perception
Abstract
Dense 3D mapping with semantic understanding is essential for robotic perception in complex environments. Recent 3D Gaussian Splatting-based mapping approaches enable high-fidelity geometry and efficient open-vocabulary perception, but typically represent semantics as unstructured feature fields that limit object-centric reasoning. In contrast, 3D scene graphs explicitly model objects and their relationships for structured reasoning, but are commonly constructed from sparse geometric representations that do not fully exploit dense semantic maps. In this work, we present OpenSplatGraph, a unified framework that constructs persistent 3D scene graphs directly from an online Gaussian-based open-vocabulary semantic map. The proposed framework augments the dense semantic map with a reliability-aware semantic field that maintains lightweight observation statistics for confidence-aware, query-conditioned object extraction. Extracted object instances are associated with persistent graph nodes, allowing object attributes and relationships to be incrementally updated across observations and queries. By tightly coupling dense semantic mapping with persistent object-centric representations, our framework supports both language-guided object grounding and structured relational reasoning while preserving the geometric fidelity of Gaussian-based mapping. Comprehensive evaluations on standard 3D scene understanding benchmarks and real-world robotic experiments demonstrate that OpenSplatGraph achieves competitive performance for online open-vocabulary perception and downstream robotic tasks. Project page: https://csiro-robotics.github.io/OpenSplatGraph.
cs.AI / 139 / 2610.07599
Modeling Latent Disturbances for Robust Decision-Making in World Models
Abstract
In this paper, we study robust decision-making in the latent space of world models (WMs). Robust optimization is a mathematical framework where, given explicitly specified dynamics and physically meaningful disturbances, a robot can select actions that remain effective even under worst-case disturbances. However, applying this principle to the learned latent space of WMs introduces a fundamental challenge: because WMs have fully learned state spaces and dynamics inferred from high-dimensional observations, it is unclear how to define latent-space disturbances that faithfully represent uncertainty in the underlying system. Our key idea is to model a latent-space disturbance as a perturbation to the learned latent dynamics that induces pessimistic but plausible transitions. Specifically, we construct a set of plausible latent dynamics by combining a dynamics-aware similarity metric that captures plausible transitions with out-of-distribution detection that excludes implausible latent states. We calibrate this uncertainty set over latent dynamics using conformal prediction, ensuring that WM imaginations induced by the latent disturbance remain plausible without becoming overly pessimistic. We then jointly optimize robust robot actions and the worst-case latent disturbances through game-theoretic optimization. We leverage this latent-space robust optimization to robustify policy steering, considering two paradigms: latent safety filtering and sample-and-verify steering of a generative control policy. Our controlled simulation experiments show that our latent disturbance enables robust decision-making directly in WM latent spaces, and hardware experiments with a Franka manipulator show that modeling latent disturbances enables robust policy steering, reducing failures by 70% in safety filtering and 54% in sampling-based policy steering. Project website: https://junwon.me/LatentDisturbance/.
cs.AI / 140 / 2610.07652
SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining
Abstract
The ability to interact with articulated objects is essential for embodied intelligent systems, but collecting large-scale real-world demonstrations for these interactions remains challenging due to the precise contact and constraint-following motions involved. Although simulation provides a promising alternative, existing synthetic data efforts cover limited articulated-object categories, while general-purpose synthesis pipelines lack explicit designs for part-level semantics and articulation constraints, hindering agentic task generation and scalable synthesis of high-quality articulated-manipulation demonstrations. To bridge this gap, we introduce SMART, a scalable system leveraging large-scale Synthesized Manipulation demonstrations for ARTiculated-object manipulation. At its core, we develop SMART-Sim, a simulation platform with articulation-aware design that enables effective task generation and efficient demonstration collection. Building on SMART-Sim, we apply agentic task generation and design a scalable distributed synthesis system, using them to synthesize SMART-Data, comprising over 1M demonstrations across 44 atomic task types, 5 robot setups, and 2,507 articulated objects. The vision-language-action (VLA) model pretrained on SMART-Data shows competitive performance on simulation benchmarks and achieves zero-shot sim-to-real transfer and scalable performance in real-world articulated-object manipulation tasks. This highlights the potential of synthetic demonstrations in providing effective and scalable supervision for improving VLA model performance in contact-rich articulated-object manipulation.
cs.AI / 141 / 2610.07681
EigenDEXplore: Structured Exploration for Dexterous Manipulation with Human Priors
Abstract
Dexterous manipulation poses a challenging high-dimensional optimization problem, as useful behaviors require coordinated motion across many hand joints. In reinforcement learning (RL) and sampling-based trajectory optimization, exploration commonly relies on independent robot joint perturbations, making coordinated behaviors difficult to discover. Prior work reduces this search space for grasp learning using low-dimensional spaces of coordinated joint motions learned from human hand data, but this restricts the expressivity required for general manipulation. Some combine learned and joint-space actions to restore expressivity, but this increases dimensionality and introduces redundancy. We study these effects across diverse manipulation settings, varying action dimensionality, exploration strategy, and the source of human data. Our experiments suggest that human-motion priors are most effective when used to structure exploration rather than change the action representation. Motivated by this finding, we propose EigenDEXplore, which induces correlated exploration by adding perturbations along human-derived eigenvectors to independent joint-space noise, leaving the action space unchanged. Across multiple dexterous hands, EigenDEXplore consistently outperforms joint-space and learned action-space baselines in grasping, in-hand reorientation, and contact-rich manipulation. These gains span unstructured and reference-guided RL, trajectory optimization, and sim-to-real deployment, and are largest in settings with less reward shaping and curriculum design.
cs.AI / 142 / 2610.07946
Adapting Vision-Language-Action Models to Unknown Visual Disruptions During Execution
Abstract
Visual disruptions can arise while a robot is executing a task, leaving a vision-language-action (VLA) policy to respond without knowing the disruption type or timing. We introduce Self-supervised Adaptation from Leftover Trajectories (SALT), which uses the leftover trajectory, the unexecuted part of the previous action chunk, as self-supervision for test-time adaptation. Because consecutive chunks overlap in time, the leftover provides a temporally aligned target for the current prediction over the same future control interval. At the onset of a visual shift, the leftover can retain a plan formed before the corruption, so updating the policy toward it anchors the adaptation across the shift (Transition Anchoring). SALT keeps the adapted policy and regenerates the current chunk, whose leftover becomes the target at the next replan, carrying the correction forward along the execution trajectory (Sequential Correction Propagation). Supervision comes entirely from the policy's own predictions, requiring no disruption annotations, expert actions, or target-domain demonstrations, and a lightweight adaptation gate calibrated only on nominal trajectories decides when updates begin. On LIBERO-10, SALT increases average success across five persistent visual corruptions from 43.9% to 53.2% with SmolVLA and from 58.7% to 66.0% with GR00T N1.7, while largely preserving nominal performance. On a real robot, it raises task progress averaged over digital and physical disruptions from 0.49 to 0.61.
cs.AI / 143 / 2610.08123
Beyond Waypoint Regression: Query-Based Cost Learning over Reachable Ego Futures for End-to-End Driving
Abstract
End-to-end planners based on waypoint regression achieve strong open-loop accuracy, but they primarily learn to mimic expert geometry and remain difficult to adapt to deployment-time safety constraints. We propose a query-based cost-learning framework that estimates bounded costs for dynamically reachable ego trajectory queries, rather than dense BEV cells or a small regressed trajectory set. Compact joint scene tokens capture coherent multimodal agent futures, while contingency-aware cost aggregation and cost-guided intra-cluster MPPI mixing convert the learned cost topology into feasible ego plans. On nuScenes, our method improves over prior cost-estimation planners such as ST-P3 and NMP, outperforms most regression baselines in collision rate, while remaining competitive in L2, and retaining an interpretable cost interface. On real-world driving logs, the proposed planner reduces collision rates compared with SparseDrive and Alpamayo without fine-tuning, while maintaining a diverse set of candidate trajectories.
cs.AI / 144 / 2610.08183
Compact Robot Policies Need Fine-Grained Visual Representations
Abstract
Multi-task manipulation policies differ in architecture, scale, and pretrained priors all at once, so published comparisons cannot attribute performance to any single component. We argue that most of it comes from the visual representation, and that parameter scale and generative priors are largely incidental. To test this, we build CoRP (Compressed Representation Policy), a deliberately compact policy (48.9M parameters, no vision-language model and no video-generative prior) that factorizes into a representation extractor and a flow-matching action generator. It reaches 97.0% on LIBERO and 75.78%/73.36% on RoboTwin 2.0 Clean/Randomized, matching systems 40.9-163.6x larger. Holding the action generator fixed, we then vary one extractor property at a time. Pretrained initialization is decisive: a random ViT-S/14 drops to 78.1% and an ImageNet ResNet-34 to 74.5% on LIBERO. Pretraining alone is not enough, as freezing the encoder costs 19.8 points. Compression matters as much: resampling each view to 48 tokens beats passing all patch tokens (97.0% vs 83.2%), and a variational information bottleneck over those tokens is worse than a hard token budget, cutting LIBERO-Goal from 95.8% to 33.0% by suppressing the instruction-dependent token selection the policy relies on. Language conditioning contributes only where the observation leaves the goal ambiguous (LIBERO-Goal: 9.2% to 95.8%), while on RoboTwin 2.0, where observations are unambiguous, removing it slightly improves success. Therefore, we argue that a compact policy works when its representation is pretrained, task-adapted, and compressed. Project page: https://corp-policy.github.io/
cs.AI / 145 / 2610.08220
VOMMI: Collecting and Leveraging Portable Demonstrations for Mobile Manipulation
Abstract
Portable mobile-manipulation demonstrations can help alleviate data scarcity for embodied intelligence, but obtaining reliable, low-cost, and robot-free motion supervision from RGB observations remains challenging. Existing approaches often rely on teleoperation or specialized devices equipped with additional sensing hardware, while directly using estimated visual odometry (VO) trajectories can introduce inconsistencies due to accumulated drift and imperfect motion supervision. We present the Visual-Odometry-Conditioned Mobile Manipulation Interface (VOMMI), a portable demonstration collection and learning framework that connects portable RGB demonstrations to vision-language-action (VLA) post-training through offline trajectory reconstruction and online visual-motion conditioning. VOMMI synchronizes body and hand views to capture navigation context and local object interactions without requiring human-robot kinematic correspondence calibration. R2-VO refines offline demonstration trajectories using sparse geometric anchors and produces causal local-motion tokens over multiple prediction horizons for online policy conditioning. An action-group residual adapter incorporates these tokens only into the base branch. Experiments use a 500-trajectory portable for each task, with 75 trajectories held out for RGB-VO evaluation, and 200 robot demonstrations as references. Our policy, post-trained only on portable demonstrations, achieves 18.2% lower base-velocity error than a policy trained with robot-collected demonstrations, while maintaining comparable end-effector translation accuracy. Offline reconstruction reduces absolute trajectory errors for the body and hand streams by 24.6% on average relative to the best evaluated baseline for each stream. The complete system improves the mean success rate by 8.3 percentage points over OpenPI 0.5 across three real-robot tasks.
cs.AI / 146 / 2610.08297
Mitigating Concept Drift in QoS Prediction for Teleoperation of Autonomous Vehicles Using Historic Data
Abstract
Teleoperation serves as the fallback solution to autonomous driving but reliable functions of the teleoperation require a certain amount of mobile network resources, which cannot be guaranteed at all times. Therefore, predictive quality of service (pQoS) is introduced as a concept to increase the resilience of the teleoperation. In this paper, based on a data measurement campaign, we propose a prediction framework to prediction two important network KPIs of teleoperation: uplink data-rate and round-trip latency. Furthermore, we introduce a method to alleviate the performance degradation of machine-learning-based prediction models on previously unseen data due to concept drift by incorporating historic data into the prediction pipeline. Additionally, we introduce the metric of critical scenario detection to evaluate the prediction performance specifically for teleoperation.
cs.AI / 147 / 2610.08350
How Much Planning Is Enough? Reducing Search and Computation in World-Model Planning
Abstract
Visual world models enable goal-directed control through decision-time action search, but their deployment efficiency is often limited by conservatively large planning budgets. We show that competitive task performance can be achieved without agreement with the Full-budget action, that sufficient budgets vary across model--task pairs, and that iterative planners repeatedly encode solve-invariant context. To address these inefficiencies, we propose {SufficientPlan}, a simple deployment framework that requires no modification to pretrained world models or planner updates. Its {Paired Sequential Budget Certification (PSBC)} component uses paired closed-loop evidence to search for and certify a reduced model--task-specific budget within a predefined Full-performance tolerance. Its {Static-Context Reuse (SCR)} component caches observation and goal representations across search iterations while preserving candidate-dependent planning and selected actions. Experiments across multiple world-model backbones and visual-control tasks show that SufficientPlan substantially reduces search budgets and planning latency while maintaining competitive control performance.
cs.AI / 148 / 2610.08541
Micro Neural Policies for Safe Real-Time Robotic Control
Abstract
In this paper, we investigate the synthesis of Micro Neural Policies (MNP) to enable safe and robust real-time robotic control on computationally constrained embedded devices. We demonstrate that integrating Evolution Strategy (ES) and Statistical Model Checking (SMC)-based verification for policy search can drastically reduce neural network size without compromising safety and robustness. We conduct a large-scale training and evaluation of MNP on Cartpole and Quadrotor control tasks, varying control frequencies and network architectures. After validating these policies in simulation, we evaluate their deployability through zero-shot transfer to physical systems. Our experiments show that MNP can successfully achieve safe sim-to-real transfer without sacrificing control performance. We then show that the policies' memory footprint, ranging from 0.5 to 7.5 kB, allows deployment on microcontrollers, where they achieve real-time inference latency with under 25 ns of jitter while leaving the chip idle for over 97% of the time for additional workloads. This makes them a highly practical solution for severely resource-constrained robotic systems.
cs.AI / 149 / 2610.08603
A Swarm-Coordinated Multi-Robot System for Early Stress Detection in Agricultural Rows Using Multimodal Leaf Sensing
Abstract
Early stress detection in crops is a necessity today to improve efficiency and reduce waste of time, money, and effort. However, most modern techniques, such as hyperspectral imaging and AI-based systems, are too costly and complex for medium and small-scale farmers to implement. This paper showcases CropSentry, a low-cost, ground-based multi-robot system that uses multimodal leaf sensing to continuously monitor crop health by tracking stress levels. The system comprises two autonomous bots that continuously detect leaf color and environmental data row by row. The observations are spatially mapped and sent over to the master bot, which uses color-coded row segments to generate a real-time web-based dashboard displaying crop health. After 63 observations were collected during the experiments, the results showed an overall crop health classification accuracy of 84.12%, with 82.60% for healthy plants, 88% for nutrient-deficient plants, and 80% for diseased plants. Also, 100% wireless communication success rate across 10 slave observations was achieved. Close-range leaf inspection across multiple bots can detect early stress in crops while remaining affordable, accessible, and scalable. It provides farmers with timely information to improve resource utilization and crop management.
cs.AI / 150 / 2610.08642
HygieneRoboBench: Benchmarking Hygiene-Aware Planning for Household Robots
Abstract
Contact with contaminated objects can spread hazards through a household robot's grippers, tools, and shared surfaces, while new contacts can make an existing plan unsafe. Existing benchmarks do not jointly assess how planners identify hygiene risks from contact history and plan safe continuations after new contact events. Planners must do so within time and resource limits while respecting user priorities. We introduce HygieneRoboBench, with 624 instances across 134 task families, to evaluate safe resolution of household tasks from a given execution history. Tasks capture contamination through two grippers and shared objects, treatment costs, and user priorities. We combine controlled history, profile, and event comparisons with independent plan evaluation. These assess safe resolution, cost efficiency under user priorities, and responses to contact events. Evaluation of LLM-based and symbolic planners shows that safely completing a task does not guarantee the lowest execution costs under the user's priorities. To address this problem, we introduce Hygiene-NSP. It combines LLM-based grounding, contact-history reconstruction, and CP-SAT to jointly plan hygiene treatment and task execution under user priorities. Hygiene-NSP achieves safe resolution and optimal safe resolution rates of 94.4% and 90.4%, respectively. Both rates are higher than those of the evaluated baseline planners on the full dataset. Project page: https://euron-zc.github.io/HygieneRoboBench/.
cs.AI / 151 / 2610.08726
EgoLAP: Learning from Egocentric Human Data through Language-Action Reasoning
Abstract
Egocentric human data offer a path to scaling robot learning beyond costly robot demonstrations, yet the embodiment gap makes raw human trajectories a poor supervisory target for control. Our key insight is that, although low-level actions are embodiment-specific, their underlying motion intent can capture task-relevant structure that transfers across humans and robots. We introduce EgoLAP, a VLA pre-training framework that jointly learns from human and robot trajectories through a shared language-based action chain-of-thought. EgoLAP expresses motion intent as structured, temporally abstracted language actions and pairs them with motion-level reasoning grounded in scene geometry, physics, and object affordances. Across extensive real-world and simulated experiments, EgoLAP transfers human experience to robot control more effectively than alternative action representations and reaches 80.1% mean real-world task progress, a 2.3x performance gain over alternative action representations. Motion-level reasoning also outperforms a composite reasoning format that combines subtask, object-box, and visual-trace reasoning.
cs.AI / 152 / 2610.08107
Exploiting Acoustic and Content-Oriented Speaker Verification Attacks Against Multilingual Voice Anonymization
Abstract
Attacker ASV systems for voice anonymization have been studied primarily in English, leaving their behavior in multilingual settings largely unexplored. Conventional ASV has shown that both acoustic and contextual information are important for multilingual speaker verification. Inspired by this, we investigate whether the same holds for attacker ASV on anonymized speech. We evaluate both acoustic- and content-oriented attackers on multilingual anonymized speech and construct a multilingual voice-converted dataset to improve cross-lingual generalization. Our results show that attacker effectiveness depends on the linguistic utility of the anonymized speech. Overall, acoustic-oriented attackers achieve better performance. However, when linguistic information is well preserved, the performance gap between content- and acoustic-oriented attackers narrows compared with conditions involving stronger speech distortion. The multilingual voice-converted dataset further improves performance and partially reduces the cross-lingual gap. These findings highlight the need for more comprehensive attacker modeling and evaluation protocols that consider both privacy and utility, rather than relying on a attacker strategy\footnote{Full code and pretrained models and MultiVC Dataset link are available at: https://github.com/monkeyDarefeen/DAST
cs.AI / 153 / 2610.07338
Logbook: Extremely Long-form Audio Event Understanding
Abstract
Audio benchmarks are built around short, pre-segmented clips, limiting model design to brief inputs or fixed vocabularies. To close this gap, we introduce Logbook, a benchmark for hour-scale audio understanding, with recordings ranging from ten minutes to six days. Given a continuous audio recording and an event label vocabulary, a system must predict a gap-free segmentation with an event label and a description per segment. We compare 52 systems, end-to-end and cascaded, and ablate fine-tuning, context length, and reasoning budget. We find the task tractable, though the best systems remain below the human reference. Also, over-segmentation is pervasive, and fine-tuning partially mitigates it. Finally, end-to-end are often better than cascaded systems, but degrades with longer context.
cs.AI / 154 / 2610.08144
Navier-Stokes lost in translation: Why Lean verification of AI autoformalisation does not guarantee correct natural language proofs
Abstract
Autoformalisation is increasingly used to verify mathematical texts, including those generated by AI, as in OpenAI's announced proof of blow-up of solutions to the Navier-Stokes equations. In this process, an AI system translates the text from a natural language (NL) into a formal language such as Lean. Once this translation is done, the argument expressed in the formal language can easily be mechanically verified. The purpose of this article is to demonstrate why this process may offer no confidence in the original NL argument, owing to the various difficulties in performing the translation semantically faithfully. In particular, we highlight that the problem of resolving ambiguities in mathematical NL text, which is necessary in order to provide semantically faithful translation, is arbitrarily high up in the Solvability Complexity Index (SCI) hierarchy/arithmetical hierarchy (the SCI $= \infty$). Hence, informally, providing semantically faithful AI autoformalisation is harder than any computational problem including the Halting problem (which has SCI $= 1$). To demonstrate the effect of this result we provide several examples of AI mistranslations of NL statements and proofs into Lean in practice, resulting in mismatches between NL proofs and their Lean `verifications'. These include OpenAI's announced Navier-Stokes proof. In particular, we show that the formalised Lean proof does not correspond to the NL proof of blow-up of solutions to the Navier-Stokes equations.
cs.AI / 155 / 2610.07607
Linear Fitness Subspace in Protein Language Models Enables Sample-Efficient Directed Evolution
Abstract
Model-guided directed evolution seeks to identify high-fitness protein variants under limited oracle budgets. Protein language models (PLMs) provide rich representations for this task, but task-agnostic zero-shot scores can be misaligned with a target assay, while supervised search in high-dimensional embedding spaces can make surrogate modeling and uncertainty estimation sample-inefficient. We propose the Linear Fitness Subspace (LFS) hypothesis: within mutation-induced residue-level representation changes, a compact, assay-specific set of directions makes fitness variation linearly accessible from few labeled variants. This is a local, supervision-recoverable statement rather than a claim that protein fitness landscapes or global PLM geometry are universally linear. Building on this observation, we introduce Subspace-Guided Evolutionary Search (SGES), which estimates an LFS from a small initial sample and performs surrogate modeling, uncertainty estimation, and acquisition in the learned subspace. Across 10 core ProteinGym assays, 87 extended static-validation assays, and an 18-assay budgeted-search evaluation, SGES improves fitness prediction and search efficiency over zero-shot PLMs and recent ML-guided protein optimization baselines. Controlled comparisons with PCA, random projections, label-shuffled PLS, classical mutation features, and acquisition ablations further isolate the benefit of a fitness-aligned site-delta coordinate.
cs.AI / 156 / 2610.07388
DeepAJM: Deep Association Joint Model for Irregularly Sampled data
Abstract
Joint Models simultaneously model longitudinal and survival outcomes, leveraging patterns in patients' longitudinal trajectory to improve the prediction of survival outcomes. The classical parametric joint models, however, rely on fixed parametric assumptions, making them susceptible to bias under model misspecification and smaller sample sizes. We propose a deep joint model, DeepAJM, that does not require any parametric assumptions, while retaining a partially interpretable, per-longitudinal-outcome association structure. The joint model uses an encoder-decoder (sequence-to-sequence) architecture to learn the latent structure in patients' time-varying covariate trajectories. The model links the longitudinal processes to the survival processes through a learned interpretable association structure, in which each longitudinal output from the decoder gets remodulated by baseline covariates before it contributes to the risk scores from the survival head of the architecture. The model was evaluated on three datasets ( a cardiovascular-disease EHR cohort, a primary biliary cirrhosis (PBC2) dataset, and a simulated dataset) against a classical parametric joint model, TransformerJM, DA-LSTM and a Cox-based survival-only model. All models were assessed using C-index, integrated brier score (IBS), time-dependent AUROC, and time-dependent AUPRC. Our model achieved the best discrimination in terms of the C-index, time-dependent AUROC, and AUPRC across all datasets.
机器学习 (cs.LG)
181
cs.LG / 1 / 2610.07637
Asymptotic Analysis of Empirical Risk Minimization on Entry-wise i.i.d. Heavy-Tailed Data
Abstract
Many real-world datasets exhibit unusually large values far more frequently than predicted by Gaussian models. Heavy-tailed distributions capture this behavior, yet evaluating learning performance under them remains challenging because rare, large feature entries retain non-vanishing effects even in high dimensions. Even in the canonical setting of empirical risk minimization for linear regression with entry-wise i.i.d. symmetric $α$-stable data, a precise asymptotic characterization of prediction has been lacking. In this work, we introduce a functional order parameter that describes the random effective problem associated with each coefficient. Using the replica method, we fully characterize the generalization error in the proportional high-dimensional limit where the sample size and feature dimension diverge at a fixed ratio. Additionally, this analysis establishes a heavy-tail universality law, scaling laws relating typical errors to prediction reliability, and the Bayes-optimal prediction error. In addition to characterizing the effects of extreme entries on the learning process, our method applies broadly to other systems with persistent local heterogeneity.
cs.LG / 2 / 2610.07243
Hybrid Cross-Modal Attention Network for Early Breast Cancer Detection in Low-Resource Clinical Settings
Abstract
Breast cancer is the leading cause of cancer-related mortality among women in Sub-Saharan Africa, where delayed diagnosis results from limited radiology expertise and fragmented clinical data systems. Although deep learning models have demonstrated strong performance in mammographic analysis, most rely solely on imaging data and are trained on Western populations, limiting their applicability in African healthcare settings. This paper presents a Hybrid Cross-Modal Attention Network (HCMAN) that integrates mammogram images with structured clinical data using transformer-based cross-modal attention mechanisms. The model was developed and validated using a locally collected dataset of 2,560 mammogram images from 1,024 patients across four Ethiopian referral hospitals, with biopsy-confirmed ground truth labels. The proposed framework achieves 97.8% accuracy, 97.2% sensitivity, 98.3% specificity, and an AUC of 0.987, significantly outperforming image-only baselines. The system demonstrates robustness to low-quality images typical of resource-limited settings, with only 3.2% performance degradation compared to 8.7% for image-only models. Cross-modal attention analysis reveals clinically appropriate behavior: higher reliance on clinical features for ambiguous cases such as dense breasts and young patients. The model's lightweight architecture enables deployment on standard hospital workstations (<2 seconds inference on CPU). This work advances sustainable, context-aware AI solutions for equitable breast cancer diagnostics in Africa.
cs.LG / 3 / 2610.07366
Identity-Conditioned Score Fusion for Open-Set Person Re-Identification
Abstract
Robust person re-identification often combines complementary cues such as face, gait, and body shape. While adaptive fusion typically targets query quality, model strength also varies across identities. We introduce identity-conditioned score fusion, a framework that tailors weights to each gallery identity without training. By contrasting intra-identity consistency against cross-identity impostors, it extracts identity-specific profiles that couple with query-conditioned adaptation via a parameter-free rule. This widens the separation between true and false matches while preserving score calibration. Evaluations on three clothes-changing person re-identification benchmarks show that our method consistently outperforms statistical, rank-based, and learned baselines, achieving up to an 8.8% absolute reduction in the false non-identification rate and demonstrating the value of identity-conditioned fusion in open-set person re-identification.
cs.LG / 4 / 2610.07576
CETUS: How Far Do Representations Trained on Earth Transfer to Cassini SAR of Titan?
Abstract
Cassini synthetic aperture radar (SAR) images reveal the dunes, plains, and lake basins of Titan, providing an instance of representations learned from Earth imagery for planetary terrain classification. Cross-domain Evaluation of Earth-to-Titan Transfer Using SAR (CETUS) compares features from DINOv2, DOFA and CROMA with classical image measurements and features from an untrained vision transformer on the U.S. Geological Survey's Cassini SAR mosaic. The classifiers learn terrain labels from an expert geomorphological map and predict those labels in geographically separate Titan regions. Under logistic regression settings, pretrained encoders achieve higher mean macro F1 than the combined classical features. Encoder rankings change when feature scaling, optimization, and regularization change together. Further training on Titan improves DINOv2 performance, degrades DOFA performance, and leads to mixed results for CROMA under the tested settings. Architectural and input processing differences prevent these comparisons from isolating the effect of pretraining. Classifier fitting and performance on individual terrain classes matter when assessing representation transfer for planetary mapping. Since the map draws partly on the same radar observations, the scores measure agreement with expert interpretation.
cs.LG / 5 / 2610.07585
REViT-v2: Hierarchical Windowed Roto-reflection Equivariant ViT for Equivariant Feature Extraction
Abstract
We propose a scalable roto-reflection-group-equivariant vision transformer based on windowed group-convolutional self-attention and a hierarchical feature architecture. We demonstrate that our approach can be scaled to group-equivariant vision transformers (ViTs) with millions of parameters and large datasets with practically sized images, i.e., ImageNet. The code and pretrained weights for the proposed Hierarchical Windowed Roto-reflection Equivariant ViTs (REViT-v2) are available at https://github.com/kc-ml2/revit.
cs.LG / 6 / 2610.07843
CHARTER: Auditing Reference Substitution in Hierarchical Compact-Evidence Evaluation for Computational Pathology
Abstract
In digital pathology, compact evidence is often used to explain or audit predictions made by whole-slide image multiple instance learning models. In hierarchical compact-evidence pipelines, candidate filtering introduces a strategy-specific candidate-conditioned prediction alongside the original full-bag prediction. If the evaluation reference changes while the intended target remains the original full-bag prediction, however, not only can the measured fidelity of the same compact evidence change, but comparisons between competing candidate strategies can also change. To make this dependence explicit, we introduce CHARTER, a reference-aware evaluation charter that asks researchers to DECLARE the intended target and reference, QUANTIFY candidate-induced prediction shift, and AUDIT the stability of comparative conclusions. Across the 15 comparisons in our main five-seed Random-K audit, 4 showed determinate reversals; in a matched native-ranking stress test, the ACMIL comparison changed from REVERSED to PRESERVED. CHARTER turns otherwise implicit candidate-filtering and reference choices into an auditable evaluation specification, helping distinguish genuine preservation of the intended prediction from apparent gains induced by changing the prediction being explained.
cs.LG / 7 / 2610.08533
Beyond Perturbation Magnitude: Direction-Dependent Responses in Multimodal Geometric Representations
Abstract
Geometric alignment scores based on Gram determinants provide a compact way to model higher-order consistency among modalities, yet how such scores respond to modality degradation is poorly understood. This paper asks whether the response of a multimodal geometric score is determined primarily by the magnitude of the perturbation-induced displacement. Using frozen cohorts from MSR-VTT (N=878) and DiDeMo (N=980), we apply controlled video blur and audio noise and analyze the response in the relational geometry on which the score is defined. Displacement magnitude explains at most 15% of the out-of-sample variance in the absolute response, and magnitude-matched pairs respond systematically differently, so scalar magnitude does not organize the response. The closed-form first-order expansion of the Gramian volume yields the Directional Geometric Response (DGR): the projection of the displacement onto the local volume gradient, which jointly captures the clean operating point, displacement magnitude, and displacement direction. The absolute first-order DGR term explains the observed response with out-of-sample R^2 of 0.838-0.969, matched-magnitude ranking accuracies of 0.864-0.963, and response-sign accuracies of 0.909-0.989, whereas the tested direction-free alternatives remain weak or unstable under the corresponding evaluation protocols. A pre-specified gain-normalization candidate, V/(g_V+eps), fails its predictability and clean-order gates. DGR uses the observed degraded-state displacement and is therefore an explanatory quantity, not a deployment-time predictor: geometric response depends on where the representation operates, how far degradation moves the relational geometry, and in which direction it moves.
cs.LG / 8 / 2610.08717
Co-Evolving Paths and Flows via Path-Flow Alignment
Abstract
We study path-flow alignment as a unified training objective for flow matching. Instead of fixing the interpolation path and learning only the velocity field, we jointly train an endpoint-preserving path network and a flow network using the same alignment loss: the flow learns to match the path velocity, and the path learns to align its velocity to the current flow. Although every fixed learned path defines a valid flow-matching objective, the alignment loss alone is not a reliable criterion for path learning. We identify path overfitting, a failure mode in which the alignment loss decreases while sample quality worsens. We find that this failure is associated with low-entropy bottlenecks in the induced probability path, where the learned path routes samples through overly concentrated intermediate marginals. Motivated by this diagnosis, we introduce a stochastic path regularizer that hides part of the source information from the path network while preserving exact endpoints. The resulting regularization gives an explicit entropy floor for the stochastic training-path marginals and empirically suppresses the bottleneck in the learned sampler, making joint path-flow training effective. On ImageNet-256x256 with SiT backbones, our method consistently improves FID across model scales, extends to model-guidance training, and leaves the inference-time architecture and sampler unchanged. Code is available at https://github.com/lizeyu090312/traj_opt_paper
cs.LG / 9 / 2610.08132
Beyond Marginal Monitoring: Distributed Joint-Distribution Testing for Data Concept Drift in Large Scale E-Commerce Operations
Abstract
Concept drift threatens production machine learning, yet the empirical behavior of multivariate two-sample drift detectors at scale remains under-characterized. Existing benchmarks rarely address the hundreds of millions of rows and high-cardinality features typical of industrial-operational datasets. We evaluate five multi-column two-sample tests (marginal, projection-based, and kernel embedding methods) across three complementary environments: the Harvard Dataverse, a validated Failing Loudly reproduction (mean absolute error between 0.030 and 0.053), and a novel synthetic-injection benchmark on the 137.5-million-row Trendyol collection-ranking feature table. Testing four drift types across two severity-scope regimes, we demonstrate that distributed Maximum Mean Discrepancy with Random Fourier Features on Apache Spark scales robustly. Averaged over the four drift types in the strong regime and under a calibrated threshold, it achieves a Pearson correlation of r = 0.940 with expected drift magnitude, an 80.4% true positive rate, and a 3.2% false positive rate. Conversely, the per-dimension Kolmogorov-Smirnov test failed due to statistic saturation from ID-like columns under asymmetric sampling, establishing a critical constraint for large-scale sampling design. At weak configurations (realized-flip fractions of at most 0.57%), detectors struggled to reliably discriminate, highlighting the need for future intensity-grid power analyses to distinguish fundamental sensitivity bounds from scalable threshold shifts.
cs.LG / 10 / 2610.07438
Artifact removal improves electrodermal waveforms but not downstream classification in a virtual-reality balance task
Abstract
Artifact removal routinely precedes the classification of electrodermal activity (EDA), on the assumption that a cleaner signal supports a better decision. We tested this assumption in a virtual-reality (VR) balance-disturbance task. A residual gating network was trained on a benchmark with expert-corrected EDA, frozen, and applied to VR recordings, where raw and gated signals were classified by five published time-series methods under identical leave-one-participant-out evaluation. On the benchmark the gate detected artifacts well (median record AUROC 0.94) and reduced error inside artifact regions by 17.8%. In the VR task it did not improve classification. Changes in balanced accuracy ranged from -1.35 to +0.93 percentage points, no classifier improved and two lost accuracy, and all five were equivalent to raw input within +/- 3.32 points. The benefit was lost between waveform and decision. The correction that lowered waveform error also reduced skin conductance response detection in all 43 benchmark records. Processing left 92.8% of predictions unchanged, and the predictions it did change were corrected and corrupted at similar rates. The VR recordings also carried little contamination (an estimated 4.6% of samples), and even perfect localization of deliberately injected artifacts recovered only 3.3 points in the most sensitive classifier. A pooled association between artifact level and accuracy (11.3 points) disappeared within participants (0.1 points), showing how differences between people can make cleaning look useful. Preprocessing should be judged by the decision it supports, against an unprocessed arm.
cs.LG / 11 / 2610.07168
A theory of platonic representations in language models
Abstract
Representations of translated sentences are similar in the inner layers of multilingual language models -- an observation connected to the platonic representation hypothesis, yet unexplained theoretically. We provide an explanation based on the assumption that data have a hidden hierarchical structure whose abstract levels are shared across languages while surface levels are modality- or language-specific. Concretely, we generate synthetic languages from probabilistic context-free grammars sharing upper-level but not lower-level production rules. In this setting the Bayes-optimal next-token predictor is belief propagation (BP); encoding its messages in successive layers yields analytical predictions that agree well with transformers trained on the same data. The framework explains why cross-lingual similarity peaks in middle layers, coexists with language-specific structure, and strengthens with language proximity, model quality and data exposure. It distinguishes similarity (shared neighborhood geometry) from alignment (shared coordinates), showing that the latter occurs when code-switched data, i.e. mixed-language sentences, are abundant enough. It further predicts that subtracting from each layer the component linearly predictable from the preceding one increases cross-lingual similarity, which we confirm in pretrained LLMs.
cs.LG / 12 / 2610.07177
CLM-as-a-Judge: Evaluating an Open Contrastive Decision Model on Public Judge Benchmarks
Abstract
An open contrastive decision model is near chance as a judge on the hard public benchmarks: Contrastive-LM/CLM-v0.1-8B scores between 0.351 (best- of-four, chance 0.250) and 0.593 (pairwise, chance 0.500), is statistically indistinguishable from coin flipping on RM-Bench and JudgeBench, and answers every HaluEval item with one constant label, matching the trivial always-first baseline at 0.581. Judges with the same parameter count score far higher everywhere: a reward model reaches 0.764 to 0.976 and a generative judge 0.611 to 0.778, and every gap to CLM is significant after Benjamini-Hochberg correction. Two properties do work. Raw confidences are overconfident by up to +0.401, yet one pooled temperature fit on held-out calibration items repairs expected calibration error to at most 0.062, and the repaired confidence ranks the model's own errors above chance on three of six benchmarks. The decision order-flip rate is 0.0002 against 0.2188 for the generative judge, and the length-preference shift is -0.023 against -0.217. The confidence-gated cascade, however, escalates between 0.923 and 1.000 of items to the strong judge at the preregistered 0.97 retention bar: calibrated confidence about a near-chance judge has almost nothing to keep. The design: five public preference benchmarks and one hallucination benchmark with real labels, scored under a preregistration frozen before any test item was seen, against generative, reward-model, and trivial baselines, with per-item predictions released.
cs.LG / 13 / 2610.07184
Learning Scientific Exploration from Human Research Decision Trajectories
Abstract
A key challenge in building AI systems for scientific research is enabling $\textit{scientific exploration}$: the systematic process of investigating unknown phenomena or ideas to gain new knowledge through sequences of research decisions and actions. Yet this process is largely missing from existing scientific corpora; for example, research papers primarily record final outcomes rather than the trajectories that produced them. In this work, we introduce $\textbf{ResearchTrails}$, a dataset of $\textbf{human research trajectories constructed from Git repositories}$, where $\textbf{commit histories}$ serve as proxies for research exploration. We develop an automated and scalable pipeline that extracts structured research trajectories from repository commits, capturing successive changes to methods, experiments, and ablations. We characterize the resulting dataset and show that these trajectories contain meaningful signals about intermediate research decisions beyond what final papers reveal. We further demonstrate utilities of ResearchTrails in multiple use cases, including retrieving human research experience as external skills at test time and training models on research trajectories to improve generalization to new research decisions. Our results suggest a path toward AI systems that learn not only from the products of science, but from the evolving process of discovery itself.
cs.LG / 14 / 2610.07197
Exact Unlearning via Quantized Sufficient Statistics
Abstract
Exact unlearning requires a deployed predictor to match one rebuilt without the information named by a deletion request. Existing general-purpose exact methods localize retraining through disjoint shards, but every request still invalidates a model, and smaller shards reduce the data available to each constituent predictor. We introduce Quantized Sufficient Statistics (QSS), which separates a small frozen schema from mutable, sum-decomposable content. The schema learns global structure; the content stores local prediction corrections as additive statistics indexed by quantized regions. Deleting content is therefore exact subtraction rather than optimization. We distinguish two guarantees: QSS-L exactly removes a label while retaining the unlabelled input, whereas QSS-E exactly removes both input and label by learning the schema without deletable examples. A deletion takes the arithmetic fast path with probability $1-ρ$ and triggers a full rebuild with probability $ρ$; all reported expected latencies include both events. Across 15 vision, text, and tabular datasets at $ρ=0.5\%$, QSS-L is within 2 percentage points of SISA on 11 tasks and provides 4--483$\times$ lower expected deletion latency on the low-class-count tasks where a compact schema is effective. QSS-E quantifies the additional accuracy cost of removing every trace of an input.
cs.LG / 15 / 2610.07207
Distributionally Robust Mixture-of-Experts Training
Abstract
Mixture-of-Experts (MoE) transformers scale capacity by activating only a few experts per token, but this sparsity creates a hidden reliability problem: when routing is imperfect, load-balanced models may send tokens to experts that are insufficiently trained for the assigned inputs. We propose Distributionally Robust MoE Training (DRMoET), a drop-in objective that treats layer-wise experts as endogenous robustness groups and optimizes high-loss routing outcomes rather than merely equalizing traffic. DRMoET updates a per-layer expert distribution by an entropy-regularized softmax rule on EMA-smoothed, activation-weighted expert losses, strengthening plausible non-top routing paths while preserving standard MoE computation. Under the FLAME-MoE recipe at 746M-total and 10.3B-total scales, DRMoET improves downstream averages over both standard FLAME-MoE and auxiliary-loss-free balancing. At 10.3B total parameters and 67B training tokens, DRMoET improves the seven-task average from 0.6625 to 0.6767, while the auxiliary-loss-free baseline achieves 0.6431. Mechanistic analyses show lower expert-loss variance with nearly unchanged mean loss, 4.3% lower excess loss under forced mid-$k$ misrouting, and improved domain-expert specialization. These results position routing robustness-not only utilization balance-as a practical objective for reliable sparse MoE scaling. Project page and code are available at: https://drmoet.github.io/.
cs.LG / 16 / 2610.07212
Reward-Driven Learning under Prompt-Level Differential Privacy
Abstract
Reinforcement learning with verifiable rewards (RLVR) trains a language model on problems that may themselves be confidential, and the trained model can reveal which problems it saw. We study RLVR under prompt-level differential privacy: the released weights must be (ε,δ)-differentially private with respect to the presence of any one training problem. Taking the group of responses to one prompt as the privacy record, our method aggregates their gradients, clips the prompt's contribution once, adds Gaussian noise, and composes the privacy loss across updates, so the budget depends on neither the number of responses per prompt nor the clipping norm; to our knowledge this is the first differential privacy guarantee for RLVR training. We train Qwen2.5-1.5B-Instruct with LoRA at a per-run budget of ε=8 and compare, on the same prompts and at the same budget, a control that removes only the reward signal and two private supervised fine-tuning recipes. The reward signal improves accuracy over the control by 2.65 points on MATH and 3.24 on GSM8K, in every seed; the improvement survives a format-robust scorer, at 1.3 points on MATH, and is not explained by response length. At the same budget the private model outperforms both supervised recipes on MATH and GSM8K by 2.3 to 3.8 points, retains 85--90% of the gain of non-private GRPO on these tasks, and on MATH the noise of an eightfold tighter budget costs at most 1.2 points. The reward effect also carries to CommonsenseQA, an exploratory non-mathematical task. Verifier feedback thus remains a usable learning signal under prompt-level privacy.
cs.LG / 17 / 2610.07218
Constant-Curvature Sliced Gromov-Wasserstein for Heterogeneous Cross-Curvature Alignment
Abstract
Recent advances in representation learning have highlighted the utility of constant-curvature models, such as hyperbolic and spherical spaces, for modeling complex data. Mixed-curvature models further enhance this by integrating multiple constant-curvature components. However, these models typically learn each component space independently because spaces with different curvatures are inherently heterogeneous and lack a unified metric. Consequently, they lack explicit mechanisms to enforce geometric consistency across various spaces. Moreover, the problem of comparing probability distributions across mixed-curvature spaces remains unexplored. To compare distributions on heterogeneous spaces, Gromov-Wasserstein (GW) distances provide a principled framework by aligning their intra-space geometries. Building on this, we propose constant-curvature sliced Gromov-Wasserstein (CCSGW), a novel divergence for aligning distributions supported on heterogeneous constant-curvature spaces. We first introduce the missing geodesic-based one-dimensional projections for spherical spaces, and then extend sliced GW to constant-curvature spaces, enabling efficient and principled comparison across manifolds with different curvatures. This formulation preserves intrinsic geometric relationships while avoiding the high computational cost. We provide theoretical analysis showing that CCSGW controls intrinsic geometric discrepancy across heterogeneous spaces, promoting distribution-level geometric consistency. By integrating CCSGW into existing mixed-curvature learning tasks, including graph anomaly detection, graph node classification, and multimodal learning, we observe consistent performance gains across diverse settings.
cs.LG / 18 / 2610.07220
Data, Numbers, and Geometry: Three Tutorials on Numerical Methods, Machine Learning, and Evaluation
Abstract
We present three practical tutorials on numerical computation and machine learning for mathematical research, developed for the DANGER: Data, Numbers, and Geometry workshop held at the Banff International Research Station in April 2026. The first develops a numerical approach to exterior calculus from pointwise evaluations of differential forms, using a flux formulation of the exterior derivative. Examples in Euclidean space and on the sphere illustrate geometric identities, topological features, and the effects of approximation and finite precision. The second examines how mathematical structure guides neural network design through examples involving elliptic curves, quivers, and a boundary value problem. It explores how architectural choices affect learning and uses interval arithmetic to bound the residual of a trained network over the full interval of the boundary value problem. The third addresses the evaluation and presentation of machine learning results, covering performance metrics, statistical uncertainty, classification thresholds, receiver operating characteristic curves, and accessible figure design. Throughout, the tutorials distinguish numerical agreement, predictive accuracy, structural guarantees, and rigorous bounds as different forms of evidence. Each contribution can be read independently, with accompanying notebooks and exercises that allow readers to reproduce the examples and adapt the methods to other problems.
cs.LG / 19 / 2610.07229
Conditional Flow Matching for Transport Between Markov Processes
Abstract
Motivated by sequence-to-sequence transport in the context time-series domain adaptation, we study the problem of transportation between trajectories of Markov processes. Given a limited number of trajectories from source distribution and the target distribution, we formulate a flow matching based algorithm which learns a transport map from the source to target trajectory distribution, while preserving the Markov structure. We show that this is consistent in the population limit and derive finite-sample error bounds under mixing time assumptions, following the analysis of classical statistical problems including regression (Nagaraj et al., 2020), principal component analysis (Kumar and Sarkar, 2023), and matrix concentration (Neeman et al., 2024) in the Markov setting. We complement that with a lower-bound construction showing that a mixing-time dependent sample complexity is unavoidable even with regular Gaussian conditional transitions. We evaluate on synthetic and real-world data. For image retrieval from electroencephalography (EEG) on THINGS-EEG2 (Gifford et al., 2022), the task is to identify the viewed image from EEG signals captured from human subjects, which suffers from high inter subject variability. We augment the ENIGMA decoder (Kneeland et al., 2026) with a conditional flow before its subject-specific temporal map. This improves mean top-5 retrieval accuracy from 43.87% to 49.05%, an 11.82% relative improvement.
cs.LG / 20 / 2610.07232
Benchmarking Time Series Foundation Models for Load Forecasting Under Covariate Uncertainty
Abstract
Accurate short-term load forecasting (STLF) is essential for the reliable and efficient operation of modern power systems. While time series foundation models (TSFMs) have recently demonstrated remarkable performance across a wide range of forecasting tasks, their effectiveness for STLF under realistic operational conditions remains largely unexplored. In this paper, we present a comprehensive benchmark of four trained-from-scratch (TFS) models and four TSFMs across three real-world load forecasting datasets under operational scenarios that differ in the availability and quality of future covariate information. Our results show that Chronos-2 consistently achieves state-of-the-art performance in both zero-shot and fine-tuned settings when future covariates are available or accurately forecast. However, its performance degrades as covariate forecasts become increasingly noisy, whereas TimesNet exhibits greater robustness under severe covariate uncertainty. These findings demonstrate the effectiveness of covariate-informed TSFMs for STLF while highlighting the critical role of robust covariate modeling in real-world forecasting applications.
cs.LG / 21 / 2610.07253
Neural Fields Encode Adaptation Geometry
Abstract
Neural fields are usually evaluated by how well they reconstruct an observation. We show that this misses two useful properties of a fitted network: how easily it can adapt to new observations, and what its weights retain from earlier ones. We study these properties as adaptation geometry. For images, we meta-learn class-specific initializations, adapt each one to a new image, and measure how much the network must change to fit it. A simple local linear model closely predicts this adaptation cost, while replacing one network's tangent kernel with another's substantially worsens the prediction. Adaptation thus depends on the local geometry of the fitted network, not only on its current reconstruction. For physical fields, we repeatedly fit the same network to observations from a sequence. Its weights then retain information about that history. When two wave histories end at exactly the same observation, the final weights recover the sign of the wave velocity with 68.6% accuracy, whereas the current observation alone contains no such information and gives 50%. These two phenomena are quantitatively linked: tangent-kernel eigenvalues predict both which changes are easy to learn and how quickly they are overwritten by later fitting. Together, these results show that neural fields contain useful information beyond what they currently reconstruct: in how they can change and in how they got there.
cs.LG / 22 / 2610.07255
Neural Algorithmic Reasoning for Graph Saddle Point Problems
Abstract
Neural algorithmic reasoning, or aligning a neural network with an algorithmic paradigm, has emerged as an approach to solving polynomial-time-solvable and computationally harder combinatorial optimization problems. We propose a new message-passing framework based on the Chambolle-Pock Primal--Dual Hybrid Gradient (PDHG) method called \textsc{GraphPDHG} for solving general graph saddle-point problems. Theoretically, we show that \textsc{GraphPDHG} can efficiently solve a family of graph saddle-point problems by simulating PDHG. We also show that our network can learn an accelerated PDHG algorithm. Experimentally, we support our results on accelerated PDHG by evaluating the performance of our model as a learned warm start for second-order optimization techniques (SSNAL). We also show that alignment with PDHG leads to stronger size generalization than non-aligned graph neural network (GNN) baselines. Overall, we propose a novel architecture for solving a general family of optimization problems on graphs.
cs.LG / 23 / 2610.07271
Algorithmically Aligned Neural Agglomerative Tree Construction
Abstract
Linkage algorithms for hierarchical clustering (HC) are a powerful and efficient framework for constructing clustering trees, yet it is often unclear which merge rule best suits a given dataset or task. In contrast, neural approaches can learn from data, but often fail to retain the efficiency and size generalization of classical algorithms. We introduce NN-linkage, a neural network (NN) model that can learn task-specific and locally dependent merge rules while retaining the recursive structure and efficient inference of classical linkage algorithms. In particular, our model is algorithmically aligned with the Lance-Williams (LW) recurrence, a parameterized framework for defining a broad, continuous family of linkage rules for agglomerative HC. Classical methods such as single linkage (SL), complete linkage (CL), and average linkage arise as discrete choices within this broader family. We show that NN-linkage is a universal approximator for continuous linkage functions, including LW recurrences, and, when paired with a transformer encoding, can also approximate globally dependent rules such as robust single-linkage. We further show that NN-linkage can exactly implement any symmetric constant-coefficient LW recurrence across all input sizes. On the empirical front, we evaluate NN-linkage in real-world applications, clock-tree routing and phylogenetic reconstruction, using both synthetic and real datasets, demonstrating its effectiveness over both classical algorithms and other neural approaches. By learning merge rules directly from target trees, NN-linkage extends efficient HC to scientific and engineering objectives not adequately captured by existing hand-designed linkage rules.
cs.LG / 24 / 2610.07283
Lock-in EP: An In-Situ Training Algorithm for Oscillatory Hardware
Abstract
Analog hardware platforms offer the potential to reduce energy consumption over digital architectures, but in order to succeed, large-scale analog systems must also be able to operate with or recover from the variability of their components. Towards this goal, we derive and demonstrate the lock-in equilibrium propagation (LIEP) training method. LIEP provides local gradient information for each component in an oscillatory network without separate forward and backward sweeps, potentially allowing for in-situ learning capabilities on analog oscillatory hardware platforms. We demonstrate that LIEP can be used both for ab-initio training as well as recovering performance when pre-trained parameters are perturbed. We show that LIEP can be formulated as a three-factor update rule, and suggest that although the method is currently only validated on shallow networks, alternate architectures may allow it to extend to deep and large-scale networks addressing complex tasks.
cs.LG / 25 / 2610.07286
FlexiFlow: Bandit-based Model Switching in ML Workflows
Abstract
Model optimizations help improve inference performance and accuracy of ML workflows. However, relying on a single model to perform inference across all data batches often fails to maximize accuracy and thus overall performance. In many cases, alternate models could perform better on specific subsets of data where a primary model underperforms. Our experiments with real ML workflows indeed show that switching models improves workflow accuracy by up to 23%. Yet, current systems lack the ability to adaptively switch between models based on performance, forcing users to manually test models in sequence. We present FlexiFlow, a dataflow system that dynamically switches between alternate models when the current model exhibits low accuracy. FlexiFlow learns to rank models using a novel multi-armed bandit approach that accounts for model runtimes, probability of passing user-defined assertions, and the computational structure of the ML workflow. We show that the standard Thompson sampling approach is insufficient for switching models in ML workflows. In contrast, our proposed approaches are effective and scales to complex real-world ML workflows. Experiments show that switching models at runtime while reusing intermediate results provides higher accuracy, but also 48% efficiency gain compared to sequential workflow runs.
cs.LG / 26 / 2610.07323
ATLAS-AL: Adaptive Trust-Region for Latent Adversarial Searches via Active Learning
Abstract
Security evaluation of learning-based systems requires more than just testing the system against a fixed collection of attacks. It requires adaptive mechanisms that can efficiently discover \textit{sets} of inputs that induce model failure. We introduce ATLAS (Adaptive Trust-Regions for Latent Adversarial Searches), which is a query-based framework that discovers adversarial input sets for black-box learning systems. ATLAS casts attack generation as an active learning level set estimation problem then combines calibrated approximations with a local-global sampling architecture to find regions of the input space that contain adversarial examples. Once discovered, ATLAS is designed to sample points within these adversarial regions to build adversarial sets that accurately represent the state of robustness of the target model. When applied on toy experiments, we find that ATLAS is able to recover more of the adversarial region under a limited query budget than does previous work. When applied to standard and adversarially trained MNIST, CIFAR, and ImageNet model targets, ATLAS produces better representative attacks than other query-based black-box attacks (NES, SignHunter, BayesOpt). ATLAS represents an automated red-teaming framework that can be used for both analyzing the robustness of learning-based systems under development and continuous auditing to see how the robustness of a system changes over time.
cs.LG / 27 / 2610.07324
Scale-Invariant Training for Time Series Foundation Models
Abstract
Time series foundation models (TSFMs) are trained on large collections of time series datasets that span various morphologies and domains. This setting exposes models to series whose scales -- typical magnitudes of their values -- can differ substantially. Affine scaling methods such as Reversible Instance Normalization (ReVIN) scale model inputs and reverse the transform before computing the loss. We show that this inversion multiplies each series' gradient by $b^p$ relative to loss on scaled targets, where $b$ is the scaling denominator (e.g., standard deviation) and $p$ is the loss degree. We call this scale-contaminated training (ScaleCon), because the scale of each series consequently becomes an importance weight, causing high-scale series to dominate training. For any scale-equivariant scaler and residual loss that is homogeneous of degree $p$, including MSE, MAE, and Quantile Loss, we prove that computing loss on scaled targets makes every mini-batch gradient and, consequently, the full optimization trajectory invariant to arbitrary independent rescaling of the training series, yielding scale-invariant training (ScaleIn). Notably, existing TSFMs use both objectives, with neither consistent reporting nor a common convention on how to compute training loss. We isolate the convergence disparity induced by ScaleCon and its correction under ScaleIn in controlled studies on synthetic and real data. In pretraining across four TSFM architectures, ScaleIn lowers MASE in all 24 architecture-benchmark comparisons, with average reductions across TSFMs of 18.8% on GIFT-Eval and 21.9% on the M-competitions. The gains extend to supervised neural forecasting, where it lowers MASE in 16 of 20 matched settings. Most existing time series forecasting pipelines can adopt ScaleIn with a one-line code change.
cs.LG / 28 / 2610.07332
Structuring MoE Expert Selection for Agentic Reinforcement Learning
Abstract
Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structures remains underexplored. In this work, we comprehensively study the connections between agentic post-training and MoE expert selection. In off-the-shelf MoE models, we observe expert selection exhibits a specialized structure that naturally aligns with agentic trajectories. Specifically, expert routing overlaps more between turns where the agent performs semantically similar operations (e.g., READ, UPDATE) than between turns with differing operations. However, standard RL algorithms ignore this specialization, allowing the MoE routing to go uncontrolled during training, which empirically limit task performance and inference efficiency. To address this, we introduce a hierarchical routing control framework for agentic tasks. We explicitly encourage turn-level expert selections to align with agentic operations while regularizing token-level expert selections to maintain local consistency. To resolve stability issues that arise during post-training with the proposed methods, we further introduce an entropy-gated control mechanism. Overall, our routing control framework achieves over 10-point improvements in success rate on all evaluated benchmarks. These results demonstrate that agentic trajectory structure provides an effective signal for optimizing MoE capacity during RL post-training.
cs.LG / 29 / 2610.07334
Weight Oracles: Reading Neural Network Weights with Language Models
Abstract
Interpretability methods for neural networks are predominantly reactive: they analyse activations produced during specific forward passes, requiring known inputs to find hidden capabilities such as backdoors. We propose Weight Oracles, fine-tuned language models that diagnose properties of a target network by reading its raw weights directly, without behavioural testing. We investigate this paradigm in two phases. Phase I establishes feasibility: through a staged curriculum and an external chain-of-computation that delegates parameter-free operations to deterministic code, an explainer LLM learns to simulate the forward pass of small transformers from their weights, achieving 99% holdout accuracy on unseen targets. Phase II repurposes this infrastructure for safety auditing. We train an oracle on natural language diagnostic questions about weight anomalies using only benign pathologies as training signal, and evaluate it zero-shot on backdoors absent from training. The oracle achieves AUROC 0.93 on attention-routed backdoors and 0.81 across a diversified threat distribution including stealth and adversarially regularized variants. Hand-crafted statistical detectors are sharp on the threat models they implicitly target but collapse on threat-model shift, while the oracle remains uniformly competent across attack types. Scaling to realistic model sizes remains the principal open challenge.
cs.LG / 30 / 2610.07340
CausalBind: Causal Modeling and Learning for Protein-Molecule Virtual Screening
Abstract
Protein-molecule virtual screening is increasingly cast as a problem of representation learning in a shared embedding space. Existing methods rely on dense holistic alignment, entangling invariant binding determinants with nuisance correlations and limiting transfer to new targets. It has been noted that binding in protein-molecule systems involves sparse cross-modality interactions: binding is governed by a small contact interface and a few decisive local interactions (e.g., hydrogen bonds, hydrophobic contacts, and salt bridges) rather than the global structures of the protein and molecule. We hypothesize that uncovering and leveraging sparse interaction patterns is critical for generalization beyond the training data, as these patterns are reusable and expected to improve performance across different scenarios. In this paper, we aim to identify and leverage sparse interaction patterns, and verify our hypothesis. Since the training data contain only observed binding pairs, we formalize this prior via a V-structure causal model under Heckman-style selection, and establish three theoretical results: (i) the latent concepts of interacting proteins and molecules are not identifiable without appropriate sparsity constraints; (ii) these concepts and their sparse interactions are component-wise identifiable under structural sparsity conditions; and (iii) a low-rank relaxation of these conditions yields subspace identifiability of the concepts and interactions. Inspired by these principles, we propose CausalBind with three implementation variants. Extensive experiments on DUD-E and LIT-PCBA benchmarks show that all variants consistently outperform strong retrieval baselines, with the largest gains on LIT-PCBA early enrichment, and further generalize to target- and scaffold-level out-of-distribution splits. Code is available at https://github.com/lokali/CausalBind.
cs.LG / 31 / 2610.07349
RELACE: retrospective likelihood-based action credit estimation for long-horizon language agents
Abstract
Group Relative Policy Optimization (GRPO) avoids a separate critic by estimating advantages from rollout groups. For multi-turn agents, however, trajectory-level supervision provides coarse, noisy credit: terminal rewards do not locate errors and can penalize useful actions alongside mistakes. Group-in-Group Policy Optimization (GiGPO) and subsequent methods refine supervision through state-conditioned comparisons, but their credit estimates remain sensitive to downstream decisions and outcomes. We introduce RELACE, Retrospective Likelihood-based Action, a critic-free framework that integrates retrospective action assessment with state-conditioned advantage estimation. RELACE evaluates executed actions through teacher-forced likelihood scoring under both their original contexts and outcome-augmented contexts. Comparing these likelihoods yields a trajectory-normalized retrospective factor that captures outcome-dependent changes in action plausibility, rather than hindsight plausibility alone. We use this factor to reweight discounted task returns and construct local advantages by comparing weighted returns among actions from equivalent states within a task. This couples retrospective relevance with observed reward, producing fine-grained credit that complements trajectory-level GRPO supervision. Temporal smoothing and success-protecting masking further stabilize the local signal. RELACE requires neither auxiliary value nor reward models nor additional autoregressive rollouts for credit estimation. Experiments on ALFWorld and WebShop with Qwen2.5-1.5B-Instruct and Qwen2.5-7B-Instruct demonstrate substantial improvements over GRPO, GiGPO, and HCAPO. With the 1.5B model, RELACE achieves $96.35\%$ success on ALFWorld and $79.43\%$ on WebShop, surpassing GiGPO by $5.47$ and $5.60$ percentage points, respectively.
cs.LG / 32 / 2610.07358
Towards Explainable Benchmarking for Data-driven Post-Wildfire Debris Flow Prediction
Abstract
Post-wildfire debris flows (PFDFs) are destructive sediment-laden hazards triggered when intense rainfall strikes recently burned terrain, destabilizing hillslopes and threatening infrastructure, local economies, and community safety. Data-driven methods have been proposed to learn predictive patterns directly from historical PFDF observations. However, the current research landscape of data-driven PFDF prediction remains highly fragmented across feature spaces, model architectures, and evaluation protocols, making rigorous comparison and the derivation of scientific insights difficult. Moreover, existing studies lack a systematic investigation into the relative importance of heterogeneous factors (e.g., meteorological conditions, terrain characteristics, soil properties, and burn severity) in triggering PFDF. To address these limitations, we present a unified benchmark for data-driven PFDF prediction, enabling fair and comprehensive evaluation across diverse models and feature configurations. Furthermore, to better understand the underlying drivers of PFDF formation, we propose a reinforcement learning-based feature selection framework that identifies factors whose perturbations render positive and negative events indistinguishable, thereby discovering the regional underlying mechanisms of PFDF occurrence across regions. Our code and benchmark are publicly available at https://github.com/KINDLab-Fly/PFDF-Benchmark.
cs.LG / 33 / 2610.07374
Multigroup Fairness and Omniprediction: Separations and Equivalences
Abstract
Omniprediction is a learning guarantee which requires a single predictor to be competitive relative to the best hypothesis from a benchmark class for any loss chosen from a family of loss functions. Loss Outcome Indistinguishability (loss OI for short) is a stronger notion that implies omniprediction. It requires the predicted distribution on labels to be indistinguishable from the true distribution to tests that depend on the loss functions and the benchmark class. Multiaccuracy and multicalibration are multigroup fairness notions that generalize classical notions of calibration and accuracy in expectation. Most known learning algorithms for omniprediction (both for the standard notion and for strengthenings like loss OI) rely on some version of these multigroup fairness notions, or on an intermediate notion called calibrated multiaccuracy. We ask if this is necessary: Does omniprediction require some form of multigroup fairness? We show that the answer is no for (plain) omniprediction, and yes for loss OI. First, a sequence of works shows that multicalibration or calibrated multiaccuracy imply omniprediction. We rule out even a weak converse, by showing that omniprediction for proper losses does not imply even accuracy in expectation, a much weaker notion than any of calibration, multiaccuracy, or multicalibration. Second, prior work showed how to achieve loss OI from a combination of calibration and multiaccuracy. We show a converse: loss OI is equivalent to a form of calibrated multiaccuracy.
cs.LG / 34 / 2610.07389
Inference and learning in sparse autoencoders as natural gradient flow
Abstract
Sparse autoencoders are widely used to uncover interpretable features in neural networks, yet reliable recovery remains difficult when features overlap or activate infrequently. These challenges involve both inferring which features explain an input and learning the dictionary that represents them. Here, we unify inference and dictionary learning as natural-gradient flows on a shared variational free energy. We instantiate this framework as BeFOND, an encoder-free sparse coding model with closed-form inference and learning dynamics. We show how recurrent explaining away reduces interference between overlapping features, while Fisher preconditioning can compensate for the slow learning of rare features. On synthetic data, BeFOND improves dictionary recovery and rare-feature detection, with a growing advantage over amortized baselines as superposition increases. On language-model activations, it improves single-feature concept detection and selective intervention, outperforming pretrained reference SAEs with substantially less training data. Its feature quality continues to improve with dictionary width, whereas the evaluated baselines largely plateau. Together, these results show how improving inference and learning within a unified probabilistic framework can make better use of data and dictionary capacity to interpret and intervene on neural representations.
cs.LG / 35 / 2610.07399
Fed-BRDECS: Privacy-Preserving and Heterogeneity-Aware Federated Deep Embedded Clustering
Abstract
Federated deep clustering seeks to learn clustering-friendly representations from decentralized unlabeled data while preserving client privacy. However, Deep Embedded Clustering (DEC)-style objectives depend on global soft-assignment statistics that require clients to reveal their sensitive information. We propose Fed-BRDECS, a privacy-preserving and heterogeneity-aware federated deep embedded clustering framework. Fed-BRDECS replaces the globally normalized clustering objective with a locally computable sample-stability loss, avoiding the transmission of local soft-assignment distributions. To tackle non-IID client distributions, we introduce prediction-balanced sampling, which oversamples locally rare predicted clusters without requiring ground-truth labels, and centroid-level restarting, which periodically refreshes biased or inactive centroids. Experiments on image and text clustering benchmarks show that Fed-BRDECS consistently outperforms representative federated clustering and deep clustering baselines under both IID and non-IID partitions. We further demonstrate its applicability to federated time-series anomaly detection, where it improves reconstruction-based detectors without adding inference-time cost.
cs.LG / 36 / 2610.07405
What pass@k Cannot Measure: Evaluating Diversity and Capability Retention after Post-Training
Abstract
pass@$k$, the fraction of problems a model solves within $k$ sampled attempts, is the field's default protocol for deciding whether reinforcement-learning (RL) post-training on verifiable rewards improved a model. At the population level, pass@$k$ depends only on a problem's probability of a correct sample, with no term for how it is distributed across outputs. We show this gap is not academic. Training Qwen2.5-1.5B-Instruct on grade-school math with Group Relative Policy Optimization (GRPO) and with rejection-sampling fine-tuning (RFT, training on the model's own shortest verifier-passed rollout) moves three complementary diversity measures (token-level entropy, answer-level entropy, unique answers per prompt) in opposite directions, with zero overlap across three seeds per arm. The gap survives restricting to verifier-correct completions only (lexical diversity among correct solutions is 15% lower for GRPO, after controlling for length) and a count-controlled check isolating diversity among incorrect answers alone, ruling out that GRPO's higher accuracy alone explains it. Yet pass@8 and pass@32 show no consistent winner on GSM8K, and a hard MATH-500 subset shows the same pattern: separation only at low $k$. Compared against the starting checkpoint, no trained arm significantly improves hard-problem coverage: RFT is significantly worse, while GRPO is statistically indistinguishable from it - so GRPO's pass@1 edge over RFT reflects a smaller loss relative to Base, not a capability gain, a missing-control issue, not a failure of pass@$k$. On GSM8K, only pass@1, with no role in detecting diversity by construction, separates the arms cleanly, rewarding the arm whose correct solutions are least diverse. We argue this is a concrete instance of a standard evaluation protocol missing a property it is routinely used to certify.
cs.LG / 37 / 2610.07406
Evaluation of Active Feature Acquisition Policies with Tabular Foundation Models
Abstract
Active feature acquisition learns policies that sequentially acquire features to maximize information about a target variable. We study how to learn and evaluate such policies from finite offline data using prior-data fitted networks (PFNs), which are off-the-shelf models that output posterior predictive distributions without task-specific training. We show that under the imbalanced coverage of offline data, using total predictive entropy as a reward creates an epistemic bias that penalizes acquiring sparsely observed features. Specifically, this reward conflates epistemic uncertainty (arising from lack of offline data) with aleatoric uncertainty (arising from uninformative features). To address this, we target the posterior expected (aleatoric) entropy instead of the total predictive entropy output by a PFN for evaluating feature acquisitions. Empirical evaluations on synthetic and real-world datasets demonstrate that our approach consistently reduces value estimation bias and yields credible intervals with strong empirical coverage, which can translate to improved downstream policy selection.
cs.LG / 38 / 2610.07419
Learnable Spectral Activations
Abstract
Implicit neural representations (INRs) are shaped by the spectral structure induced by their input encodings and activation functions. Existing methods improve fitting primarily by modifying which frequencies are available to the network, through coordinate encodings or periodic nonlinearities. However, frequency access is not the only bottleneck: signals with localized or spatially varying structure require the network to efficiently compose frequencies into multi-harmonic internal responses. We introduce learnable spectral activations (LSA), which replace fixed neuron-level nonlinearities with a residual truncated Fourier series whose harmonic amplitudes are learned during training. LSA does not expand the asymptotic function class. Instead, it changes the factorization of the representation: linear weights select features while activation coefficients control spectral shaping, and the two are updated by separate gradients. Because the activation output is affine in the coefficients given fixed pre-activations, spectral tuning becomes a more direct subproblem compared to architectures where it is entangled with feature selection. Empirically, this factorization concentrates more target-signal energy in the leading eigenmodes of the neural tangent kernel, consistent with improved optimization behavior. Across audio, image, neural radiance field, and neural acoustic field tasks, LSA also improves reconstruction quality.
cs.LG / 39 / 2610.07420
Benchmarking Label-Revealed Online Updates for EEG BCI Decoding
Abstract
Electroencephalography (EEG) signals drift over time, which can cause static brain-computer interface (BCI) models to degrade in practice. We present a benchmark for online adaptation and compare two widely used pipeline families, Common Spatial Patterns (CSP) and Riemannian covariance-based methods, under time-ordered prequential (test-then-train) evaluation. We examine (i) which pipelines benefit most from label-revealed updates, (ii) whether controlled forgetting of older data improves robustness, and (iii) how a minimal-calibration cold start compares with starting from a pretrained model. Across four datasets (three motor-imagery datasets and one movement-decoding dataset), label-revealed online updates improve 13 of 14 model/dataset pairs on the two largest streams, with relative accuracy gains of up to about 18% over a frozen model. A Shapley-based data-valuation analysis over temporal blocks assigns the largest mean value to the most recent block in each of the three analyzed datasets, while older blocks retain positive value.
cs.LG / 40 / 2610.07430
StaFIR: Convex Learning of Stationarity-Aware Causal Filters
Abstract
Reducing nonstationarity in a persistent time series entails deciding how much of its temporal dependence to remove. In finance, fractional differencing is often tuned using the Augmented Dickey--Fuller (ADF) test, limiting the search to a one-parameter family of lag profiles and addressing input preservation only indirectly. We propose StaFIR, a causal finite-impulse-response filter with a learned nonnegative mixture of exponential lag profiles. Its convex learning objective balances empirical stationarity with similarity to the input. We evaluate StaFIR on ARFIMA--GARCH controlled settings and rolling financial series, including a realized-volatility forecasting task. The experiments show that StaFIR adjusts its filtering strength to persistence while limiting unnecessary transformation in stationary regimes. In downstream forecasting, there is no clear accuracy difference from fixed half-order differencing, while StaFIR achieves higher measured similarity to the raw signal. A complementary direct forecasting experiment finds that greater input similarity is associated with smaller forecasting penalties, although the raw representation remains stronger.
cs.LG / 41 / 2610.07444
Decoupling What from Where: How Should a Small GUI Grounding Model Receive the Action Type?
Abstract
A GUI agent decides which action to take and where to take it; we ask how a small grounding model should receive the action type. Fine-tuning Qwen2-VL-2B with LoRA on Android in the Wild, we compare a flat baseline with five ways of supplying the type under matched data, compute, and decoding: an auxiliary loss, a hard-routed action word, an additive learned embedding, a prepended learned token, and the type written into the prompt. With five seeds, an episode-clustered bootstrap, and seed-level paired tests, the ranking on a mixed stream is clear: the auxiliary loss, the additive embedding, and the prompt word each gain five to seven hit@0.10 points over the baseline, while hard routing and the prepended token are not distinguishable from it. Much of that gain is protection from a preprocessing choice of ours rather than a spatial prior. Our serializer clamps the off-screen touch point AITW records for type events to the origin; that class degrades the baseline's click grounding, and removing it lifts the baseline by nearly seven points, after which no mechanism's hit rate beats it and the intervals exclude a two-point effect, though the auxiliary loss still shortens the average miss; on a stream of taps and swipes none helps. Whether this generalizes beyond one serialization is open. For deployment, the pipeline's margin over the baseline with predicted rather than gold types is not established (+0.016, 95% interval [-0.017, +0.052]), and a wrong type collapses every model conditioned at inference. The prepended token does not help at the shared learning rate, where its rows barely move from initialization; trained ten times faster it reaches the level of the other three, with a margin three seeds do not establish. We also document a silent failure: injecting conditioning through inputs_embeds makes Qwen2-VL fall back to 1-D positions for image tokens, costing nine points.
cs.LG / 42 / 2610.07447
Fork-and-Flush: Escaping Idea Basins in Autoresearch Agents
Abstract
Autoresearch agents tackle open-ended problems by repeatedly proposing candidate solutions, evaluating them, and using feedback to guide subsequent experiments. We show that independent runs of the same agent on the same task often plateau at substantially different scores, with gaps that persist even after considerable additional compute. Embedding their candidate artifacts by functional similarity provides further evidence that trajectories remain in localized regions of the solution space, which we call idea basins. To help agents escape these basins, we study a simple periodic intervention, fork-and-flush. Our method forks the agent into parallel trajectories, each inheriting the accumulated workspace but starting with a fresh chat context. After running each trajectory for a fixed horizon, the agent continues from the highest-scoring one. Across 13 long-horizon research and engineering tasks, with individual agent runs lasting up to several days, fork-and-flush outperformed the single-run and best-of-N baselines by a relative improvement of 66.0% and 44.4%, respectively, on the min-max normalized average score under an equal compute budget.
cs.LG / 43 / 2610.07452
Active Feature Acquisition for Cost-Efficient Temporal Prediction with Reduced Participant Burden
Abstract
Accurate forecasting of pathological outcomes is a central problem in psychology. To do so, psychologists often collect intensive longitudinal data. However, in such studies, the desire to acquire a large number of variables for the sake of accurate prediction is often counteracted by the need to minimize participant burden. Acquiring more variables per occasion can yield better predictions, but having too many acquisitions increase the risk of non-response and attrition. Longitudinal Active Feature Acquisition (LAFA) is a principled approach to resolve this conundrum. Instead of requiring responses to every item at every acquisition occasion, LAFA produces a policy that seeks to optimally select dynamic subsets of items to be acquired at each timepoint while preserving our ability to forecast a specific outcome. However, existing LAFA methods are mostly based on Neural Networks (NN) that are difficult to interpret in practice. In this work, we introduce a tree distillation method for learning an interpretable policy from NN-based LAFA networks. We validated our method through both a simulation and an empirical EMA dataset on forecasting daily alcohol consumption. In both cases, we find that we can meaningfully reduce the number of items acquired at each occasion with minimal loss in accuracy. Networks (NN) that are difficult to interpret in practice. In this work, we introduce a tree distillation method for learning an interpretable policy from NN-based LAFA networks. We validated our method through both a simulation and an empirical EMA dataset on forecasting daily alcohol consumption. In both cases, we find that we can meaningfully reduce the number of items acquired at each occasion with minimal loss in accuracy.
cs.LG / 44 / 2610.07458
Interpretable Hypergraph Learning via Neural Additive Models
Abstract
Hypergraphs offer a natural framework for modeling networked data, where dependencies among entities are governed by higher-order interactions. While hypergraph learning methods such as hypergraph neural networks have demonstrated remarkable predictive performance, most existing approaches rely on black-box message-passing architectures, making it difficult to disentangle the contributions of node attributes and higher-order structural information. To address this challenge, we introduce the hypergraph neural additive network (HGNAN), an inherently interpretable framework for learning on hypergraph-structured data. HGNAN extends classical neural additive models to higher-order relational data by integrating feature-wise nonlinear decomposition with hypergraph-aware structural aggregation, enabling transparent prediction for both node- and hyperedge-level tasks. Extensive experiments on benchmark datasets demonstrate that HGNAN achieves performance comparable with state-of-the-art hypergraph learning methods while providing intrinsic and meaningful interpretability.
cs.LG / 45 / 2610.07466
Efficient Multimodal Inference through Adaptive Acquisition and Sequential Fusion
Abstract
Multimodal systems often encode every available input, even when a subset suffices for prediction. Adaptive acquisition can reduce this cost by using predictions from incrementally fused evidence to decide which modality to encode next and when to stop. However, sequential fusion makes these predictions order-dependent, so decisions based on them may need to distinguish factorially many histories of the same acquired set. We introduce SemARC, which couples a Sequential Modality Aggregator (SeMA) with an Adaptive Runtime Controller (ARC) and uses acquired evidence to select each modality before its encoder runs. SeMA executes only selected encoder and fusion branches, updates a fixed-size state, and predicts after each acquisition without recomputing earlier branches. We supervise every acquisition prefix under randomized modality subsets and orders to encourage consistent predictions across acquisition orders. ARC combines a set-dependent marginal-utility prior with residual fitted-Q learning to select the next available modality or stop, without inspecting unacquired inputs or retaining acquisition order. Across six multimodal classification datasets and eleven baselines, SemARC achieves 3.2% higher macro-F1 and 61.4% lower total inference GFLOPs on average relative to each dataset's most accurate baseline. End-to-end latency falls by 44.0% across GPU and CPU and by 47.2% on Android INT8 relative to the fastest measured baseline, on average. Under varying runtime modality missingness, SemARC still skips available modalities, matching or exceeding the best baseline macro-F1 in 21 of 24 conditions with 14.8% lower total GFLOPs on average. SemARC thus offers a practical path toward efficient multimodal inference across heterogeneous devices.
cs.LG / 46 / 2610.07470
Structure, Not Belief: Correlated Thompson Sampling from LLM-Derived Covariance in Combinatorial Semi-Bandits
Abstract
Combinatorial Thompson sampling (CTS) draws independent posterior samples for every arm, so its exploration dynamics ignore any relation among arms. We study a minimal change to those dynamics: an LLM is queried once for a partition of the arms, the partition becomes a positive-definite correlation matrix $Σ$ through an RBF kernel on cluster ranks, and the per-round posterior sample is drawn with covariance $Σ$ while the Beta posteriors are updated from real rewards only, so the LLM shapes how the sampler moves, not what it believes. We give a self-contained Bayesian regret bound for the idealized Gaussian sampler whose information gain splits into a $K\log T$ term from the $K$-cluster structure and a ridge term that grows to $d\log T$: the $\sqrt{d/K}$ improvement over independent sampling is a finite-horizon transient, exact only as the within-cluster correlation tends to one. The correlated sampler reduces regret by 19% over CTS on 16 synthetic Bernoulli families at $T=2{,}500$ (6-7% at $T=25{,}000$ with data-adaptive kernels) and by 41% on the Microsoft MIND-small news benchmark ($d=200$ real articles), while pseudo-observation warm starts give nothing. An LLM-free ablation with a simulated oracle of controlled quality shows that on unstructured instances the gain is a property of the kernel shape (a random partition, or a plain tempering of the sampling noise, reproduces it), while belief injection at matched oracle quality never helps.
cs.LG / 47 / 2610.07475
Adapting to Changes in Agent Behavior via Finite-Depth Policy Sensitivity
Abstract
Adapting a reinforcement learning policy to changes in another agent's behavior typically requires a large amount of new interaction data. Policy sensitivity provides a first-order prediction of how a locally optimal policy changes with a behavioral parameter, but its computation requires second-order derivatives whose effects propagate across future interactions. We develop a finite-depth framework to estimate this sensitivity by approximating the policy Hessian and mixed derivative using information from a reference environment. The method features an adjustable propagation depth which determines where derivative propagation along the trajectory is truncated. We characterize the derivative contributions omitted by finite-depth propagation and derive truncation-error bounds for the approximated derivatives and resulting policy sensitivity. The bounds are nonincreasing with propagation depth and vanish at full-horizon propagation. Using a belief-driven pursuit-evasion game as a validation scenario, the proposed method generally achieves lower derivative-estimation errors as the propagation depth increases and outperforms the baseline methods in both estimation accuracy and policy adaptation. The sensitivity-based initialization improves zero-shot return over direct transfer, and also shows advantages for the subsequent fine-tuning in the target environment.
cs.LG / 48 / 2610.07484
SpecBraM: What Should an EEG Foundation Model Predict? Masked Band-Power Prediction versus Waveform Reconstruction
Abstract
Self-supervised EEG models often reconstruct masked waveforms or predict discrete codes. We study a task-aligned alternative: masked band-power prediction (MBP), which predicts fixed narrow-band log spectral energy for masked channel-time patches. This target retains rhythm power relevant to sleep staging while avoiding phase-sensitive waveform reconstruction and a learned codebook. Across three pretraining seeds, we compare band-power and waveform targets with matched backbones, pretraining data (2,388 hours), and training steps, including a 2x2 tokenizer-by-target design. On ISRUC and HMC sleep staging, MBP exceeds raw- and band-waveform reconstruction by 1.6-2.8 balanced-accuracy points with all labels and 4.7-7.3 points with 1% of labels under a strict linear probe; the target effect exceeds the tokenizer effect. Its frozen features reach 0.7916/0.7425 balanced accuracy, versus 0.7636/0.7227 for a matched rich handcrafted spectral baseline, although the gap is about one point with 1% of labels. Full fine-tuning reaches 0.8107/0.7669. The gains do not extend to every task with spectral cues, including motor imagery, depression screening, and vigilance regression. These results support choosing pretraining targets to match the physical quantities and spatial and temporal scales relevant to downstream labels.
cs.LG / 49 / 2610.07485
Robust Importance Sampling for Rare Events via Constrained Gaussian Mixtures
Abstract
We study estimating rare-event probabilities $I = \mathbb{P}(g(\mathbf{X}) > γ)$ with $\mathbf{X} \sim \mathcal{N}(\boldsymbolμ, \boldsymbolΣ)$ and general $g : \mathbb{R}^d \to \mathbb{R}$. We address this problem through importance sampling, and propose a framework that substantially improves efficiency and robustness over baselines such as crude Monte Carlo, adaptive cross-entropy, variational-inference-based methods (including reverse- and forward-KL approaches), as well as Safe-ICE, Subset Simulation, and Sequential Monte Carlo, drawing on ideas from both rare-event estimation and cross-entropy optimization. The key contribution has two parts: first, we separate the problem into coverage, to overcome the cold-start barrier, and fitting, to refine proposals once a meaningful signal is available; second, we constrain the final GMM proposal so that it has finite importance-sampling variance (since coverage alone is not sufficient -- without safeguards, importance sampling may still suffer from infinite variance). Together, these ingredients yield expressive proposals; finite variance does not by itself guarantee practical stability at a fixed sampling budget. Extensive experiments demonstrate substantial variance reduction, strong robustness across diverse benchmarks, and favorable cost--efficiency trade-offs, with the proposed approach often outperforming these baselines, particularly in high-dimensional and multimodal settings where competing methods frequently become unstable or fail. Our code is available at https://github.com/lorek/robust-cfi-is.
cs.LG / 50 / 2610.07491
Who Bears the Burden? Learning Responsibility for Shared Constraints in Multi-Agent Reinforcement Learning
Abstract
When multiple agents share a cost budget, a common Lagrange multiplier can enforce the aggregate constraint but does not determine how its penalty should be allocated across agents. Uniform penalties ignore heterogeneity in the rewards agents sacrifice, while agent-specific multipliers may still rely on the same aggregate cost signal. We introduce Lagrangian Responsibility Allocation (LiRA), which learns each agent's share of a common multiplier by optimizing social welfare over a finite training horizon. The multiplier enforces the aggregate budget, while responsibility shares redistribute its influence without modifying the original rewards or constraints. For convex games under standard regularity conditions, varying these shares induces a smooth family of normalized generalized Nash equilibria in which active constraints remain at their budgets while welfare varies. To optimize responsibility before convergence, we derive a welfare gradient that accounts for both learning updates and the induced change in data distribution. Across CityLearn, MABIM, Harvest, and MetaDrive, spanning 3 to 400 agents, LiRA improves average social welfare by up to 29% over uniform and agent-specific multiplier baselines. Grid and driving costs remain within budget, inventory violations decrease, and Harvest makes more effective use of available budget.
cs.LG / 51 / 2610.07499
Source-Learned Reliance for Selective Test-Time Adaptation of Multimodal Time Series
Abstract
Multimodal wearable systems must remain reliable when sensor streams become noisy or unavailable. Existing multimodal test-time adaptation (TTA) methods often assess reliability online, but cross-modal agreement can be misleading when sensors measure different physical processes, and evaluating alternative modality configurations adds inference cost. We propose CARAT, which decouples model reliance from runtime corruption detection to guide omission or attenuation, amortizing reliance estimation through source training. An asymmetric modality-dropout curriculum prepares a missingness-resilient backbone for omission and derives a frozen, backbone-specific reliance proxy from windowed input-projection gradient norms. At deployment, a lightweight one-class detector flags suspect streams, and the proxy guides a joint choice between replacing the suspect set with the backbone's trained missingness symbol and attenuating its representations before fusion, without candidate-subset evaluation. Across four wearable datasets, five corruption types, three backbones, and eight TTA baselines, CARAT achieves the highest overall macro-F1 and best mean rank (2.42), exceeding EATA, the strongest baseline, by 1.58 F1 points across 12 equally weighted dataset-backbone settings. Across five profiled configurations, CARAT uses 9.49% fewer GFLOPs and updates 47.82% fewer parameters than EATA. A pattern also emerges across sensing regimes: multimodal TTA methods such as PTA are competitive on IMU-dominated homogeneous datasets, whereas unimodal TTA methods like TENT and EATA match or exceed it on heterogeneous datasets. These results position CARAT as a practical default to wearable TTA, offering competitive robustness with modest computational requirements and benefits that vary across backbones and dataset regimes.
cs.LG / 52 / 2610.07529
Targeted search shows that random-device testing underestimates worst-case error in a simulated wave-based neural operator
Abstract
Wave-based processors promise fast, energy-efficient Fourier layers for neural operators. They are usually validated on randomly sampled devices, but using them requires knowing how large their error can become under fabrication and alignment variation. In a stylised numerical case study, a hybrid Fourier neural operator runs its four spectral layers on simulated coherent 4f processors with 32 toleranced knobs, whose half-widths are representative rather than calibrated. For 120 models (four tasks, six training methods, five seeds), we compared the worst of N random in-spec devices with a searched one. On a deterministic simulator with one frozen draw of the random static errors, the searched device's held-out error was 1.08-3.10 times the maximum over 200 Monte Carlo devices and 1.06-2.71 times that over 1000. With 20 fresh static draws, it still exceeded the maximum over 200 random devices in 116 of 120 models. Under uniform sampling, the probability of drawing such a device is at most 0.37% per model (two-sided 95% Clopper-Pearson), which says nothing about how large its error is. The gap persisted with uniform or Sobol' sampling at the search's budget, shared knobs, a second crosstalk model, box scales of 0.25-2 and a pixel-level device model. Models trained only with random static errors reached 3.7-39.9 times their nominal error on searched devices, and fine-tuning on random and gradient-searched devices gave the lowest searched error of the six in all 20 task-seed pairs. For two heat-exchanger quantities, a search targeted at each exceeded the worst of 1000 random devices in all 39 models, and hence the Wilks 95/95 limit (worst of 59). For the mean pressure of 11 models, no random device exceeded a 1% error threshold, but the searched device did. Random testing estimates how often errors exceed a threshold; worst-device search gives a lower bound on how large they can be.
cs.LG / 53 / 2610.07540
Preserving Unstable Modes Through Inverse Dynamics in JEPA World Models
Abstract
Robotic systems often exhibit unstable modes, along which small perturbations and disturbances can cause unbounded growth unless corrected through feedback. Controlling such systems from high-dimensional visual observations requires representations that preserve these modes. Joint-embedding predictive architectures (JEPAs) provide a natural framework for learning such representations and their dynamics from visual data. However, we demonstrate that next step prediction combined with anti-collapse regularization does not guarantee that controllable unstable modes are preserved: the training loss can be minimized while these modes are collapsed, making stabilization from the learned representation impossible. To address this, we augment world-model training with an action reconstruction objective (i.e., an inverse dynamics loss) that encourages control-aware representations, namely, visual representations that preserve crucial features for control. We prove that exact action reconstruction makes the encoder injective on the finite-horizon reachable subspace. Thus, the encoder cannot discard any state direction reachable by an action sequence within $H$ steps. Moreover, we show that, as $H$ grows, the dominant eigenspace of the finite-horizon controllability Gramian converges to the controllable unstable subspace. We establish our theoretical results for linear systems and demonstrate empirically that our findings extend to nonlinear visual control tasks (CartPole, Walker2D, and PointMaze), highlighting the benefits of control-aware representation learning.
cs.LG / 54 / 2610.07550
Foundation Model-Aided Multi-Agent Reinforcement Learning for Wireless Random Access Network Optimization
Abstract
Random access (RA) is one of the most foundational medium access control (MAC) layer scheduling schemes for handling unpredictable data traffic from multiple terminals. While multi-agent reinforcement learning (MARL) has been explored to optimize RA-based wireless networks, its reliance on experience-driven, distributed policy learning incurs significant training overhead for each optimization task, limiting its feasibility in real-world applications. In this work, we propose to leverage a foundation model (FM) to improve MARL efficiency across diverse RA network optimization tasks. Specifically, we design an FM-aided actor-critic algorithm within a consensus-based decentralized MARL architecture and provide its convergence analysis under local reward exchanges and nonlinear value function approximations to show that our algorithm achieves the same convergence order as the conventional MARL with critic model exchanges and linear approximations. Our numerical results show that our FM-based approach significantly enhances MARL speed for RA network optimization.
cs.LG / 55 / 2610.07553
Which and When to Admit: Gradient Admission for Data-Centric Small Language Model Finetuning
Abstract
LoRA fine-tuning adapts small language models (SLMs) to heterogeneous instruction data within a low-rank update subspace, making it vulnerable to three structural problems: conflicting gradients that cancel, static data selection that cannot track evolving learning dynamics, and subspace saturation that causes later updates to overwrite useful directions. We argue that effective adaptation therefore requires controlling which data-induced gradients enter the LoRA subspace and when. We propose GRADE (GRadient-Aligned Data-centric rEcipe), a data-centric framework combining two mechanisms: a state-aware selector that continually admits samples aligned with the evolving multi-task gradient field, and a self-calibrating step-level gate that rejects updates likely to cause destructive overwrite near saturation. Across three current-generation backbones and a heterogeneous seven-dataset instruction pool, GRADE outperforms strong data-selection and PEFT-stabilization baselines in accuracy and robustness. It is the only method to improve consistently over standard LoRA on every architecture, while producing more coherent gradient trajectories and less destructive overwrite. These results show that successful SLM adaptation depends not only on which data are selected, but also on which gradients are allowed to enter and persist in the constrained update subspace.
cs.LG / 56 / 2610.07555
Global Transport Couplings for Classifier-Free Guided Flows
Abstract
Optimal-transport couplings have been shown to reduce training variance in unconditional flow models, but their role in conditional generation remains unclear. A natural approach constructs separate couplings for each condition, but this is impractical for large or continuous conditioning spaces found in modern image foundation models. We introduce Global Transport (GT), a global class-agnostic optimal-transport coupling, computed without class labels. GT can associate different conditions with different regions of the source noise, and consequently worsens performance without guidance. However, when combined with classifier-free guidance (CFG), GT consistently improves generation across domains, model scales, and sampling budgets. This reversal suggests that couplings for conditional flows should be evaluated both empirically and theoretically under the guided flow used at inference, rather than on unguided generation. We evaluate GT over both discrete class and continuous text conditioned image generation across model scales, and investigate how coupling choice alters guided trajectories. These results identify coupling design in the guided flow setting as a simple training time axis to improve performance without modifying existing architectures, samplers, or guidance mechanisms.
cs.LG / 57 / 2610.07559
TAFFY: A Task-Adaptive Tabular Foundation Model with In-Context Diversity
Abstract
Recent progress in tabular foundation models suggests that training on synthetic tasks can substantially improve in-context learning capabilities, with overall performance largely depending on how well models can infer task-specific predictive relationships from the available context during inference. In this paper, we introduce TAFFY, a tabular foundation model with an In-Context Diversity Prior and a Task-Conditioned Looped Transformer that strengthen this ability. Specifically, to construct each synthetic pretraining context, the In-Context Diversity Prior samples from multiple related environments derived via controlled interventions and distribution shifts on a shared causal process. This in-context diversity encourages the model to learn a more comprehensive and task-specific representation. Moreover, the Task-Conditioned Looped Transformer iteratively and selectively applies a shared group of Transformer blocks to refine contextual representations, with a task-conditioned gate modulating the final hidden-state update. This enables task-adaptive iterative refinement. Together, these components encourage the model to identify predictive relationships from contextual contrasts during pretraining and dynamically modulate context integration for each task. Across six classification and five regression benchmark datasets, TAFFY attains the lowest average rank.
cs.LG / 58 / 2610.07562
Learning a Mixture of GFlowNets
Abstract
Learning an ensemble of GFlowNets to sample from a discrete target distribution has become a common approach for achieving better state space exploration and convergence than that of a monolithic sampler. However, these methods often add a substantial runtime overhead to the base model, and their conceptual connection remains elusive. To address this, we first propose a general-purpose theoretical framework for describing a mixture of GFlowNets, which we specialize into continuously (CI) and discretely indexed (DI) collections. On the one hand, we show CI GFlowNets can be interpreted through the lens of a random features expansion, provably boosting the sampler's expressivity in graph-structured tasks and reducing learning instability via spectral shifting. On the other hand, we demonstrate DI GFlowNets encompass prior approaches for GFlowNet training and provide the foundation for the newly proposed Stratum-Conditioned (SC) GFlowNets. This method, which is inspired by the Doob's h-transform of Markov chains, decomposes the state space according to a prescribed modular function and restricts each component to sample from a distinct subset of it. Importantly, SC GFlowNets support centralized and component-wise embarrassingly parallel training, and we show both of them significantly speed up learning convergence and mode coverage without introducing any non-negligible extra computation.
cs.LG / 59 / 2610.07565
Complementary Feature Domains: Information Preservation Does Not Imply Predictive-Contribution Preservation
Abstract
Complementary Feature Domains (CFD) theory characterizes predictive value as a context-indexed contribution system induced jointly by representations and their realization family. We show that Shannon-information preservation does not imply preservation of this contribution system: an invertible representation transformation can leave target information unchanged while altering predictive contribution under a restricted decision family. We formalize the resulting transition through a CFD contribution defect that measures how contextual contributions change under controlled recoding. For bounded Lipschitz utility, we show that each coalition utility shift is bounded by the behavioral distance between the attainable action sets before and after recoding; consequently, every contextual contribution defect is bounded by the sum of the corresponding coalition incompatibilities. Exact behavioral closure yields invariance, while increasingly accurate compensation yields restoration. A controlled ECG experiment illustrates the mechanism: a nonlinear bijective recoding preserves the information in a frozen time-frequency representation but changes accuracy under a fixed affine learner; applying the exact inverse restores all tested coalition accuracies. The result separates information preservation from realization-dependent contribution and provides a quantitative transition law for multi-representation prediction.
cs.LG / 60 / 2610.07583
Mechanistic Interpretability of Atmospheric Rivers in GraphCast
Abstract
While AI weather models now rival operational forecasts, how they represent the atmosphere internally remains an open question: feature attribution reveals which input patterns matter, not what the model computes or how it combines information internally. We train sparse autoencoders (SAEs) on GraphCast to uncover its learned concepts, using atmospheric rivers as our phenomenon of focus. Both standard and Matryoshka SAEs show GraphCast computes atmospheric river intensity, measured by integrated vapor transport (IVT), as a stable internal variable, despite IVT being neither an input nor a target. In contrast to the unstructured concept retrieval of the standard SAE, the Matryoshka SAE orders concepts by importance and exposes their relations. Atmospheric river concepts persist across depth and direct interventions confirm causality. This method offers a way to find internal variables and determine which of them the model actually relies on, which is a prerequisite for asking whether those variables remain meaningful as the phenomenon changes under a warming climate.
cs.LG / 61 / 2610.07610
Hub for Outliers, Spokes for Inliers: Uniform Latent Space Construction for Dual-Mismatched Semi-Supervised Learning
Abstract
Semi-supervised learning typically assumes that labeled and unlabeled data share an identical class distribution and label space. However, this setting is often violated: unlabeled data may be imbalanced and contain unknown class samples, causing mismatches in both class distribution and label space. Such dual mismatch leads to majority classes dominating the latent space and unknown class samples being overconfidently misclassified, degrading feature discriminability and pseudo-label quality. To address this, we propose a hub-spoke latent geometry, where known classes are uniformly distributed around a central hub and each class forms compact clusters around its prototype, while the hub provides an anchor for a low-evidence region specifically designed for high-uncertainty unknown class samples. Integrated with an evidence-based classifier, this geometry ultimately enhances feature discriminability and uncertainty separation by mitigating majority-class domination through structured feature organization and guiding high-uncertainty unknown class samples toward the hub. Extensive experiments show that our method outperforms state-of-the-art methods, with a maximum improvement of 3.25% across various settings.
cs.LG / 62 / 2610.07615
AFA-BANDIT: Provably Near-Optimal Online Multi-Feature Classification Under Budget Constraints
Abstract
Active Feature Acquisition (AFA) is a classification problem in which an agent decides which costly features to acquire before predicting each sample's label. Unlike batch AFA, which trains a fixed policy and classifier offline on fully observed data, online AFA updates its predictor from revealed labels as samples arrive. Existing online methods either use deep reinforcement learning (RL) without performance guarantees or maximize cost-adjusted reward rather than enforce a global budget. We formulate online AFA as a combinatorial Bandits with Knapsacks (BwK) problem that couples acquisition and prediction. Unlike prior bandit-based AFA and classical BwK, our setting has combinatorial complexity, evolving rewards, a global budget, and structured side information. We obtain an improved regret upper bound over standard BwK bounds in this framework, leveraging a cardinality-aware confidence bound and the subset update structure. To avoid an exponentially large action space, we propose \emph{LP-Chain}, a variant that searches a cost-aware chain of feature subsets with a size that grows linearly with the number of features. While the regret upper bound is specific to the combinatorial framework, \emph{LP-Chain} empirically achieves comparable predictive performance. On synthetic data, \emph{LP-Chain} outperforms HEDGE-based BwK and deep RL-based online AFA baselines and scales favorably to more features.
cs.LG / 63 / 2610.07625
Stateless Language Agents: Scaling Long-Horizon Automated Research
Abstract
Automated research systems increasingly run LLM agents over long horizons, but more inference does not by itself produce more progress: agents replay growing histories, duplicate one another's work, or stop experimenting while token consumption continues. Yet most evaluations use short budgets or benchmarks that saturate early, leaving these failure modes untested. We trace these failures to two choices: where research state lives and who decides what to try next. We introduce Stateless Language Agents (SLAs), built on the principle of stateful search with stateless agents: no agent carries its conversation across invocations; instead, the harness owns the research state (candidate solutions and measured outcomes) and reconstructs a fresh and role-specific context for every invocation. What each agent sees becomes an explicit design choice rather than a history that grows with the run. We implement this principle in the SLA framework, where a stateless Advisor reads harness-summarized evidence across search directions and assigns concrete experiments to parallel Workers. We evaluate SLA against three recent frameworks on software engineering, kernel optimization, and algorithm design at budgets of up to one billion tokens. SLA achieves the best final result on every task and reaches the strongest kernel baseline's final performance with over 84% fewer tokens. Ablations from shared checkpoints show that focused contexts and explicit assignments each contribute to SLA's progress, with effects that can compound over full runs, while the Advisor consumes less than 0.6% of tokens. These results argue for SLAs, which keep durable research state out of agent conversations, and show that short evaluation horizons can misjudge research systems and their components.
cs.LG / 64 / 2610.07628
Complementary Supervised and Self-Supervised Representations for Out-of-Distribution Graph Learning
Abstract
Out-of-distribution (OOD) generalization remains challenging for graph neural networks (GNNs), as graph distributions can vary substantially across time and domains. Supervised and self-supervised graph representation learning are guided by distinct objectives and offer different perspectives on graph representations. In this work, we study whether self-supervised representations (SSL) can provide complementary signals to improve supervised OOD node classification. We develop two backbone-agnostic frameworks that exploit such information at different stages of learning and prediction. Co-Train jointly learns supervised and SSL representations and adaptively integrates them during training, while Dual-Space Retrieval performs non-parametric prediction in the two representation spaces and combines their predictions through confidence-aware fusion at inference time. The supervised and SSL encoders are separately parameterized and need not share the same GNN architecture. We evaluate multiple GNN backbones and two distinct SSL objectives, DGI and GRACE, on four graph benchmarks spanning temporal and cross-domain distribution shifts. Extensive experiments show that Co-Train consistently outperforms strong supervised OOD baselines, while Dual-Space Retrieval achieves competitive performance as a flexible non-parametric alternative. Results across different backbones and SSL objectives, together with representation analyses and ablations, demonstrate that SSL representations provide complementary information to supervised representations and can improve OOD node classification across diverse settings.
cs.LG / 65 / 2610.07662
MS-ECG-FM: Towards a More Universal Electrocardiogram Foundation Model for Health Monitoring using Multi-source Contrastive Learning
Abstract
Electrocardiography (ECG) records the electrical activity of the heart, aiding diagnosis by detecting abnormalities in cardiac function. ECG foundation models have demonstrated promising results, but are limited by a reliance on ECG interpretation reports as their sole supervision. Because interpretation reports only capture the subset of waveform information routinely recognized by clinicians, this constrains representation learning to overlook the broader diagnostic signals present in ECG. We introduce a new ECG foundation model --- MS-ECG-FM --- that is trained through contrastive alignment to multiple distinct clinical note types, including ECG, echocardiography, radiology, and discharge reports. We evaluate MS-ECG-FM on an extended set of ECG detection benchmarks, showing that it comprehensively outperforms existing methods on the full span of conditions that ECG can detect, including in reduced-lead configurations. Different reports improve representations for different diagnostic domains, while multi-source alignment captures their complementary information and produces consistently strong representations across clinically diverse tasks.
cs.LG / 66 / 2610.07676
Exact-Solution Volume and Length Generalization in Transformers
Abstract
Research on transformer expressivity shows whether a transformer is capable of solving a given task, but gives little indication of whether the solution, if learned, is generalizable to longer input lengths. We study this question through normalized exact-solution volume (NESV): the fraction of a bounded parameter region that achieves an exact solution on every input of length $n$. For fixed-width, single-layer transformers with $\log n$-scaled attention, we establish asymptotic bounds on NESV for four tasks: FIRST ($Θ(1)$), MAJORITY ($Θ(1/(n\log n))$), INDEX ($Θ(1/n^3)$), and PARITY ($0$). These results are consistent with previous empirical results: the faster the exact-solution volume decays with input length, the harder it is to length-generalize on that task. Looking deeper into INDEX, our volume analysis reveals two error sources that grow with $n$. Consequently, we study a transformer model that would structurally eliminate one of the terms, theoretically improving the NESV bound to $Θ(n^{-1})$, and empirically achieving 85% accuracy when tested at $10\times$ the training length, compared with the 60% accuracy of the original model. We conclude that volume analysis may be a useful approach to identify concrete sources of length sensitivity and thus provide insights into task-specific model refinements.
cs.LG / 67 / 2610.07677
Adaptive Model Inversion Attacks Generalize a Privacy-Robustness Tradeoff
Abstract
In this paper, we show that standard evaluations of high-resolution Model Inversion Attacks (MIAs) significantly underestimate training-data privacy leakage. State-of-the-art privacy defenses, standard training techniques such as MixUp and Adversarial Training, and undefended models all leak training images at rates 1.16 to 6.59 times higher on FaceScrub under simple adaptive changes to the attack, with the largest increases among defenses reporting the strongest privacy. We further show that measured leakage depends on the feature basis of the external classifier used to evaluate reconstructions: for the same reconstructed images, an adversarially trained Inception evaluator identifies the targeted identity at different rates than the standard Inception evaluator. Our results suggest that standard MIA evaluation can mistake optimization and measurement failures for privacy. These underestimated leakage rates also concealed a broader relationship between privacy and adversarial robustness. Once we adapt the attack and vary the evaluator, reconstruction leakage closely tracks adversarial robustness across recent defenses and standard training regimes, suggesting that robustness provides an attack-agnostic proxy for reconstruction vulnerability that applies far more broadly than previously theorized. This raises an open question: can a practical defense reduce training-data reconstruction without paying a corresponding cost in adversarial robustness?
cs.LG / 68 / 2610.07699
Improving Synthetic Data Generation for Argument Mining via Adversarial Reinforcement Learning
Abstract
Argument Mining (AM) is fundamentally constrained by the scarcity of high-quality structure-annotated datasets. While LLMs have shown promise in synthetic data generation, producing synthetic AM data that is both structurally accurate and sufficiently diverse remains a challenging problem. To address this problem, we revisit synthetic data generation for AM from a new perspective and propose a novel adversarial reinforcement learning framework for data synthesis. The proposed framework jointly optimizes the generator and the discriminator in an adversarial loop, in which the generator produces structured AM instances, and the discriminator provides learning signals by distinguishing real data from synthetic candidates. This enables the generator to progressively improve both the structural accuracy of generated argument data while maintaining diversity through adversarial feedback. Extensive experiments demonstrate that the proposed framework consistently improves AM performance on three benchmark datasets in both full-data and low-resource settings, validating its effectiveness and scalability.
cs.LG / 69 / 2610.07713
Neuromotor Hierarchy Network: Physiological Inductive Biases for Robust Generalization in sEMG Decoding
Abstract
Surface electromyography (sEMG) provides a wearable, noninvasive interface to neuromuscular activity for movement decoding and human-computer interaction. Population-scale decoding remains difficult because the relationship between sEMG and neuromuscular activity varies across users and sessions, while task-relevant dynamics span channels and multiple timescales. Learning waveform-to-output mappings from task labels leaves the distinction between recording variability and coordinated motor activity implicit. We introduce the Neuromotor Hierarchy Network (NHN), which learns a compact latent neuromotor state from task supervision to represent task-relevant neuromuscular coordination. NHN constructs this latent state through a hierarchy inspired by neuromotor organization.It adapts recording statistics while preserving relative intensity.Its spatiotemporal encoder uses parameter-efficient channel interactions and modulates features with multi-timescale history. The resulting features yield candidate activations of learned motor primitives, which are temporally integrated and continuously weighted to form the state. Theoretical analysis characterizes the efficiency, temporal behavior, and optimization of NHN's core mechanisms. We evaluate the architecture for both continuous hand-pose estimation on emg2pose and touch-typing recognition on emg2qwerty. On emg2pose, NHN reduces user-averaged angular error by 0.52% to 2.84% across all three generalization splits in both Regression and Tracking relative to Hadidi et al.'s best task-specific variants, using 48.42% to 48.51% fewer parameters. On emg2qwerty, NHN reduces beam-search character error rate by 19.40% zero-shot and 30.42% after fine-tuning relative to SplashNet-Upscale, using 65.86% fewer parameters. Physiology-guided inference of a latent neuromotor state supports parameter-efficient sEMG decoding.
cs.LG / 70 / 2610.07754
Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures
Abstract
Adversarial training is one of the most reliable defenses against adversarial attacks, but its high computational cost must generally be paid anew for each task. Robust foundation models offer a promising alternative: adversarially pretrain a model once and then transfer its robustness to downstream tasks through lightweight adaptation. However, a fundamental question remains open: can robustness acquired during pretraining transfer to unseen tasks without further adversarial training? In this study, we answer this question affirmatively. A single model adversarially pretrained at scale can achieve optimal robustness on new tasks without additional task-specific training. Specifically, we show that, for a family of Gaussian-mixture classification tasks, a sufficiently deep linear transformer adversarially trained across tasks can asymptotically attain the robust Bayes error on previously unseen tasks through in-context learning from clean demonstrations. By contrast, a standardly trained model cannot. We further analyze convergence under gradient flow, an accuracy--robustness trade-off, and demonstration complexity.
cs.LG / 71 / 2610.07778
Towards One-for-All Foundation Model for Attributed Graph Clustering
Abstract
Attributed graph clustering aims to discover node groups by jointly exploiting node attributes and graph topology, yet its unsupervised nature makes model selection and adaptation inherently difficult. Existing methods typically train and tune a separate model for each input graph, leading to costly and fragile pipelines that often fail to transfer across graphs with different feature spaces, structural patterns, and attribute-structure correlations. In this paper, we study a one-for-all alternative: can a single model be trained once and directly applied to diverse attributed graphs without graph-specific training, fine-tuning, or hyperparameter search? We propose OFAG, a foundation model for attributed graph clustering. Building upon Prior-data Fitted Networks, OFAG learns a reusable clustering inference strategy from synthetic attributed graphs generated under broad priors over latent clusters, node attributes, and graph structures. To handle incompatible feature spaces across graphs, OFAG adopts a dimension-agnostic signal-wise graph encoder that treats each feature channel as a graph signal and models its response to shared graph filters. The model is trained with a hyperspherical clustering objective, producing clustering-friendly node representations in a single forward pass at inference time. On ten datasets, one frozen OFAG model achieves the best mean performance and average rank across NMI, ACC, ARI, and F1, while completing all ten datasets in 12.43 minutes total---over 6* faster than the second-fastest baseline and nearly 28* faster than the second-best on clustering quality. Our code and pretrained checkpoint are available at https://github.com/Cloudy1225/OFAG, allowing practitioners to directly apply OFAG to their own attributed graph datasets without additional training or tuning.
cs.LG / 72 / 2610.07786
Extending Pathwise Gradients to Discrete Random Variables via Finite-Order Relaxation
Abstract
Pathwise gradients are preferred for continuous random variables because they are unbiased, low variance, and work with a single sample. For discrete variables, however, the pathwise identity cannot generally be exact for every differentiable function. We propose a general framework to construct finite-order exact pathwise gradient estimators for a range of common discrete variables such as Poisson. The estimator is the least-norm solution among all solutions that are unbiased for polynomials of degree at most. The resulting estimators preserve the hard forward sample, require no temperature tuning, and can be implemented in a few lines of codes. Against other admissible solutions, our estimator is unique and minimizes weight variance; in contrast, prior works use categorical variables or augmented representations to approximate non-categorical variables that induces excess variance and computations. To understand approximation bias for functions beyond the prescribed class, we also derive a non-asymptotic bias bound. In experiments our low order methods match or improve tuned baselines across linear, nonlinear and hierarchical latent-variable models, while out-speeding competitors in every runtime benchmark.
cs.LG / 73 / 2610.07796
The Geometry of Empowerment
Abstract
Empowerment captures the capacity for an agent to actively control its environment. While conceptually appealing as an information-theoretic quantity, the connection between empowerment and structurally central states that provide broad access to future outcomes has remained an open question. In this work, we link empowerment maximization and skill-learning methods to provide new geometries for interpreting and analyzing empowerment. Our analyses answer longstanding open questions on the connections between empowerment and structural centrality. Our analyses also reveal distinctions between information and reward geometries, highlighting important theoretical implications to build scalable empowerment-maximization methods. Website and code can be found at https://empowerment-geometry.github.io/.
cs.LG / 74 / 2610.07804
Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow Analysis
Abstract
Prior Fitted Networks (PFNs) such as TabPFN now rival established statistical procedures across prediction and estimation tasks. A natural explanation is that PFNs have the property of statistical adaptivity, that is, they perform nearly as well as a method tailored to the true data-generating model for a heterogeneous set of models, while not being told which model the data comes from. We study how such adaptivity is learned in a controlled location-estimation problem. Each task is an unlabeled sample whose family is hidden: Gaussian data call for averaging, with error of order $n^{-1}$, whereas uniform data are best estimated from their extremes, at the faster rate $n^{-2}$. We also provide the example of a symmetric Gaussian mixture, for which a rate of $σ^2_n/n$ can be attained. On scalar inputs, softmax attention computes the derivative of the empirical cumulant-generating function. A single primitive therefore both supplies features that distinguish the families and forms estimators interpolating between the sample mean and the mid-range. We combine attention experts through either a softmax mixture of experts or a gated linear unit (GLU), and analyze stagewise gradient flow. With $\widetildeΩ(n^{1+ε})$ pretraining tasks, the learned estimator is asymptotically efficient on Gaussian tasks, within a factor $n^ε$ of the minimax rate on uniform tasks, and order-optimal on mixtures in a shrinking-variance regime. These guarantees extend to new locations and longer contexts. A risk decomposition separates expert error, routing error and normalization error, which clarifies the architectural contrast. Softmax gating enforces normalization and exact translation equivariance, whereas the GLU must learn it: its dynamics separate into fast bias removal followed by slow expert selection. End-to-end experiments recover the predicted specialization.
cs.LG / 75 / 2610.07809
MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling
Abstract
Sparsely activated Mixture-of-Experts (MoE) models increase model capacity without a proportional increase in per-token computation. Dense-to-MoE upcycling reuses pretrained dense models to construct such systems, commonly by copying feed-forward networks (FFNs) into independently trained experts. We introduce MASKerade, a dense-to-MoE training method that instead learns experts as sparse subnetworks of a frozen pretrained FFN. Each expert is defined by a learned binary mask, and a token-level router selects which masked FFNs to execute and combine. The router and mask scores are optimized jointly, while the underlying FFN weight values remain unchanged. This formulation supports neuron-structured, semi-structured, and unstructured experts within the same routing architecture. Our main configuration uses four 2:4 experts with top-2 routing, where two half-dense expert passes have the nominal FFN arithmetic of one dense pass, without requiring independent expert weight matrices. On five vision-language benchmarks with Qwen and Gemma backbones, this configuration achieves the highest performance among the compared baselines. Comparisons across mask granularities, routing interventions, and compute-matched controls distinguish the effects of learned connectivity from expert activation count. These results establish mask learning over frozen weights as a practical alternative for constructing token-routed MoE experts.
cs.LG / 76 / 2610.07810
SIFT: Search Intent-to-Filter Transformer for Multi-Task Personalized Filter Ranking at Airbnb
Abstract
Search filters help guests navigate vast catalogs in two-sided marketplaces like Airbnb, and recommending the right filters can meaningfully lift booking conversion. Many such production filter-ranking systems, however, represent the guest through hand-engineered, pre-aggregated features generated by ETL pipelines. This makes it expensive to maintain and difficult to extend for new filter types or contextual dimensions (trip length, group size). We present SIFT (Search Intent-to-Filter Transformer), a ranking model built on transformers that learns guest preferences directly from raw behavioral sequences. SIFT replaces manual feature engineering with a unified guest representation that feeds multiple prediction tasks, including booking likelihood, filter engagement, and ordinal capacity thresholds (e.g., 2+ bedrooms) -- a general framework for filter ranking in two-sided marketplaces that accommodates both boolean and numeric-range filter types. Extending SIFT to new filters requires only adding a new head, not a new feature pipeline. To keep serving fast, this guest representation is computed offline on a daily cadence rather than at request time. Offline, SIFT improves booking and amenity-engagement PR-AUC by +51.9% and +62.8% respectively over the production baseline. In online A/B testing, SIFT increased engagement with recommended filters by +20.0%, overall filter usage among searchers by +0.72%, and usage of the newly-supported bedroom, bathroom, and bed filters by +3.9%, +10.7%, and +0.52% respectively. Demonstrating the system's extensibility, we rapidly integrated a novel hotel-intent filter using the same shared representation, driving a +3.8% lift in uncancelled hotel bookings and a +0.76% lift in overall marketplace bookings. SIFT is now fully deployed in production, serving scalable personalization to millions of guests.
cs.LG / 77 / 2610.07823
TTNet: Multi-Task Deep Learning for Table Tennis Player Analysis with Smart Racket
Abstract
The AI CUP 2025 Precise Analysis of Table Tennis Smart Racket Data Competition introduced smart table tennis rackets that collect extensive player swing data, enabling research on table tennis big data. These data support in-depth analysis of players' return techniques and swing-force consistency, improving the accuracy of player skill assessment. This study focuses on six-axis sensor data collected by smart table tennis rackets and proposes TTNet, a novel deep learning model with multitask learning capabilities, to advance table tennis data analysis and related applications. TTNet combines convolutional neural networks (CNNs), residual networks (ResNet), and self-attention mechanisms to simultaneously predict four player attributes: gender, playing hand, years of experience, and skill level. We adopt a two-stage training strategy that incorporates data augmentation and task-specific loss functions to improve generalization on imbalanced data. Our approach achieved second place on the official competition leaderboard.
cs.LG / 78 / 2610.07824
CANDLE: Cortical Null-Space Decomposition for Noninvasive Brain Source Imaging
Abstract
Electrophysiological source imaging (ESI) aims to estimate cortical source activity from noninvasive electrophysiological measurements such as electroencephalogram (EEG). However, ESI is fundamentally ill-posed because source activity is substantially higher-dimensional than sensor observations, resulting in non-unique solutions. Recent learning-based approaches address this ambiguity by learning data-driven source priors, yet they often struggle to generalize across subject-specific cortical geometries. To address this, we propose CANDLE, a learning-based ESI model that estimates source activity on subject-specific cortical geometries. CANDLE learns a prior over the null space induced by the source-to-sensor mapping derived from T1-weighted MRI, restricting learning to unobservable source components while preserving geometric constraints. To train CANDLE, we develop a whole-brain simulator spanning over 1,100 subject-specific cortical geometries with source configurations derived from over 26,000 statistical brain maps. Trained exclusively on simulated data, CANDLE outperformed prior ESI methods on simulated source activity estimation and generalized to two empirical tasks: (i) intracranial stimulation localization from simultaneously recorded scalp EEG and (ii) epileptogenic zone estimation from presurgical interictal EEG. Our project page is available at https://candle-esi.pages.dev}{https://candle-esi.pages.dev.
cs.LG / 79 / 2610.07834
Retrieval Is Not Enough: Refreshing Memory for Frozen Time-Series Forecasters
Abstract
Retrieval-augmented time-series forecasting uses the continuations of historical segments similar to the current context as references for a forecaster. Most existing methods build the retrieval memory once from the training segment, leaving observations revealed after deployment unavailable as references, and generally do not calibrate how much the retrieved information should influence a frozen forecaster. We identify two key determinants of retrieval utility for a frozen forecaster: whether the history still reflects the current state, and whether the correction it induces aligns with the forecaster's residual errors, an alignment that can shift between validation and deployment when the memory becomes stale. We propose FreshCast, a plug-in retrieval framework that keeps the forecaster frozen, continuously updates a non-parametric memory with new observations, forms a memory forecast through relational kernel regression, and calibrates its weight in closed form on the validation segment. Under a simplified generative model, we characterize the optimal combination gain through the second-order relation between forecaster error and memory correction, and show that a sufficiently long look-back can make periodic memory information redundant. Across seven benchmarks and ten forecasting architectures, FreshCast reduces average MSE for every evaluated forecaster and input length, by 14.6% and 5.6% at input lengths 96 and 720, and achieves lower MSE than the evaluated retrieval-augmented and online baselines in their comparison settings. Ablations show that freezing the memory at the end of training removes most of the gain, identifying post-training observations as a primary source of improvement. For a frozen forecaster, useful historical references must remain timely and provide information that helps correct its remaining errors.
cs.LG / 80 / 2610.07842
Privileged Context as Drift in On-Policy Self-Distillation
Abstract
On-policy self-distillation (OPSD) trains a language model to match a copy of itself conditioned on privileged context. Existing work varies what privileged context contains and how it is produced while also changing models, data, and training setups, making the effects of privileged context design difficult to isolate. Motivated by efforts in continual learning to reduce catastrophic forgetting, we study how the choice of privileged context affects policy drift. Specifically, we vary two axes: content (a demonstration, feedback, or rephrase) and source (external, self-generated with a verifier, or self-generated without a verifier). We train Qwen2.5-7B with OPSD across these nine combinations and three datasets, measuring target-task accuracy, prior-task retention, reverse KL from the base policy, and parameter-update geometry. Holding source fixed, changing content spans a wider median KL range than holding content fixed and changing source. The ratio between these ranges is $5.1\times$ for per-token KL and $2.2\times$ for per-sequence KL. Parameter-update geometry shows the same pattern: updates from adapters that share content are more closely aligned (mean cosine $0.571$) than updates from adapters that share source ($0.255$). For continual learning, these findings suggest that privileged context should be treated as part of OPSD's stability design because it is associated with how far and in what direction the policy moves.
cs.LG / 81 / 2610.07853
Lost in the bf16 Cast: Exporting Ternary Language Models Can Revert Most Low-Learning-Rate Code Changes
Abstract
Ternary language models such as BitNet b1.58, Falcon-E and BitCPM are fine-tuned with higher-precision latent weights and deployed as ternary codes produced by an export step that, in the labs' documented pipelines, first casts the latents to bf16. We audit those pipelines across three labs. In released checkpoints, fp32 quantization of the shipped latents disagrees with the deployed codes on 0.83-1.77% of codes in Falcon-E and BitCPM and on 1.530% in BitNet 2B-4T; for Falcon-E and BitCPM most disagreements are products that bf16 rounding lands exactly on the threshold, which ties-to-even maps to zero, and the unmodified onebitllms exporter reproduces all four Falcon-E releases byte for byte. At fine-tuned endpoints, with learning rates selected to match a nominal learning-rate-to-bf16-ULP ratio, the documented export lowers greedy GSM8K strict accuracy from 58.79% to 0.78% for Falcon-E-1B-Base and from 36.13% to 0.39% for BitCPM-CANN-0.5B, and a bf16 save and reload lowers BitNet 2B-4T's strict accuracy by 27.54 points while its last-number accuracy rises. Two compatibility remedies, writing the training quantizer's codes directly or adjusting the bf16 inputs until the unchanged tools emit them, each met a 4-point strict-accuracy non-inferiority criterion against online evaluation in all three models. In two model families, randomized interventions on the initial distance from the threshold support distance-dependent selection of the codes that fine-tuning changes.
cs.LG / 82 / 2610.07857
A Decision-Focused Neural Optimization Framework for Personalized Route Reproduction from Vehicle Trajectories
Abstract
This study formulates individual route reproduction as a shortest-path problem over learned driver-specific latent link costs. The central idea is that, once such latent costs are inferred from contextual information, observed routes can be reproduced without enumerating alternative route sets. We propose a neural pipeline that includes a perception model that embeds context covariates, which comprises individual characteristics, trip-specific attributes, and network-level traffic states, into the personalized link costs. A constrained optimization (CO) layer, which determines the shortest path (SP) based on these estimated costs, follows the perception encoder. To enable end-to-end training, we employ decision-focused learning to align the predicted shortest paths with observed routes. The implicit maximum likelihood estimation (iMLE) provides an approximate gradient of the loss function that contains the non-differentiable CO layer. Furthermore, a regularization term anchors the latent cost distribution to the empirical scale of observed link travel times, mitigating the scale ambiguity inherent in shortest-path supervision. Empirical evaluations demonstrate that the proposed framework outperforms baseline route choice models in path reproduction. The learned latent costs, interpreted as proxies for perceived travel costs, provide plausible explanations for heterogeneous route choices.
cs.LG / 83 / 2610.07859
Tram-FL: Reducing Communication and Computation Costs through Sequential Model Circulation in Decentralized Federated Learning
Abstract
Conventional decentralized federated learning (DFL) often focuses on clients, with each client maintaining a model copy, performing updates individually, and undertaking model exchange and integration. While fully leveraging computational resources can shorten training times, it can also lead to significant computational and communication waste. This is especially pronounced with non-independent and identically distributed (non-IID) data, where achieving high model accuracy demands extra resources. This research shifts focus to the model itself, aiming to realize DFL with minimal computation and communication costs. To this end, we propose Tram-FL (Traveling Model Training Mechanism for Decentralized Federated Learning), a mechanism designed to efficiently address these challenges. It sequentially trains a single model by circulating it among nodes. We address the training scheduling problem in model circulation-based training, specifically determining which nodes should update the model and the number of updates to perform. This is approached by considering the model's circulation route and update iteration allocation, for which we propose simple yet effective methods. Additionally, with quantized momentum, Tram-FL achieves high accuracy with fewer model circulations while controlling communication load per transmission. Experimental results show that the proposed algorithm, even with non-IID data, converges to a global model with reduced communication and computation.
cs.LG / 84 / 2610.07874
On-Policy Distillation with Negative-Policy Rollouts
Abstract
On-policy distillation (OPD) has been widely studied as a post-training method in which a student model obtains token-level supervision from a stronger teacher on its own rollouts. Recent studies have improved OPD through alternative distillation reward formulations and teacher configurations, while the objective of distillation remains centered on mimicking the teacher. However, when a stronger teacher has limited distributional overlap with the student, such positive guidance can provide insufficient learning signals. In this work, we introduce Negative-Policy OPD (NP-OPD), which complements teacher supervision with rollouts from a lower-performing, lower-capability negative policy that serves as a negative reference for the student. Rather than modifying the distillation reward formulation, NP-OPD introduces the negative policy at the rollout stage, continuously supplying tokens preferred by the negative policy over the teacher so that they remain exposed to teacher supervision throughout training. This provides an explicit negative signal through negative-policy rollouts while preserving the positive teacher supervision used in OPD. Through extensive experiments, we show that NP-OPD improves OPD across model scales, generation modes, reasoning domains, and different OPD variants. Furthermore, our analyses show that NP-OPD effectively suppresses tokens preferred by the negative policy over the teacher and moves the student away from the negative policy. These results support our design of introducing negative signals through negative-policy rollouts and provide new insight into the role of the rollout policy in OPD. Code will be available at https://github.com/naver-ai/np-opd.
cs.LG / 85 / 2610.07885
Label-Efficient Deep Learning for ECG Delineation: A Multi-Dataset Benchmark against Widely Used Delineation Tools
Abstract
Electrocardiogram (ECG) delineation, the identification of waveform boundaries, is a foundational step that translates raw ECG signals into clinically interpretable measurements. Deep learning has advanced this task but remains dependent on costly expert annotations. Label-efficient strategies such as self-supervised pretraining and semi-supervised learning are expected to ease this burden, yet it remains unclear whether they yield reliable delineation and whether the deep models they produce outperform the delineation tools used in practice. We address this in two stages. First, comparing self-supervised objectives with supervised or semi-supervised fine-tuning across one internal and four external datasets, we find that pretraining helps but the objective matters, and that the value of semi-supervised fine-tuning depends on the pretraining objective. Second, we benchmark the selected deep learning model against widely used open-source (NeuroKit2, Prominence, ECGdeli) and commercial (CalECG) tools using three complementary metrics. The model ranks best on every metric and dataset, outperforming the strongest tool by a clear margin on the rhythm-diverse set (mIoU 71.3 vs. 54.8%; averaged point-wise sensitivity 92.6 vs. 76.4%), and degrades the least from sinus to arrhythmia. A rhythm-stratified and point-wise analysis further characterizes the distinctive behavior of each tool, yielding practical guidance for tool selection. These results provide systematic, multi-dataset evidence that self-supervised pretraining is effective for ECG delineation and enables a label-efficiently trained deep learning model to outperform widely used delineation tools by leveraging abundant unlabeled data. This supports adopting such models in diverse, real-world clinical settings.
cs.LG / 86 / 2610.07898
FC-SWE: Failure-Conditioned RL for Long-Horizon Software Engineering Agents
Abstract
Repository-level software engineering (SWE) is a challenging long-horizon setting: agents must reason over extended interactions, use tools, and adapt to stateful environments. Recent work trains SWE agents with reinforcement learning methods such as Group Relative Policy Optimization (GRPO), which independently sample multiple trajectories per issue, test the resulting patches, and compare terminal rewards within a fixed group. However, this training setup does not reuse verifier feedback from failed patches as context for subsequent attempts, even though this feedback contains valuable diagnostic information about what went wrong. Training on recovery trajectories is challenging because the preceding outcome determines whether the next trajectory is generated, while the failed execution determines its conditioning context. We introduce FC-SWE, a failure-conditioned RL framework that incorporates recovery attempts into policy training. After a patch fails verification, FC-SWE restores the repository to its original task state and uses the failed patch and verifier feedback as context for a recovery trajectory. FC-SWE adapts GRPO to these chains of complete, multi-turn tool-use trajectories through two mechanisms. Trajectory-local rewards preserve each attempt's verifier outcome, preventing recovery success from rewarding an earlier failed patch. Active-set advantage estimation forms a comparison group from all initial and recovery trajectories actually executed for the same issue, so failed attempts remain in the group while unexecuted attempts are excluded. On all 500 SWE-bench Verified tasks under a verifier-assisted protocol, FC-SWE with Qwen3.5-4B and SWE-agent achieves 41.7% Resolved@1 and 52.8% Resolved@2, compared with 38.9% and 48.5% for GRPO. Although trained with at most two attempts per chain, FC-SWE reaches 70.7% Resolved@11 under an eleven-attempt test-time budget.
cs.LG / 87 / 2610.07899
Variance-Averse $n$-Step Offline Reinforcement Learning for Sparse Long-Horizon Environments
Abstract
Generative actors are transforming offline reinforcement learning (RL) by enabling expressive policy classes that model complex action distributions. However, this expressiveness also exposes a key challenge in heterogeneous datasets: generative policies can reproduce unreliable action modes whose return distributions exhibit high variance, occasionally yielding high returns by chance but lacking consistency. Consequently, maximizing the expected $Q$-value alone is insufficient for identifying reliable actions. We propose VAN-Flow (Variance-Averse $n$-step Flow), a framework that promotes reliable actions in generative offline RL. VAN-Flow combines (i) a categorical distributional critic, (ii) a variance-averse expectation operator that smoothly reweights atom probabilities to favor actions with both high returns and low dispersion, and (iii) a flow-matching generative actor guided via rejection sampling. Unlike CVaR or mean-variance objectives, the operator redistributes probability mass over the categorical return distribution without hard truncation or auxiliary penalty terms. Across more than 40 tasks from D4RL and OGBench, VAN-Flow consistently outperforms strong baselines, with the largest gains in long-horizon and high-variance regimes where reliable action selection becomes critical.
cs.LG / 88 / 2610.07904
ApexQuant: Data-Free Elastic Quantization by Residual Re-Isotropization
Abstract
We introduce ApexQuant, a calibration-free quantization method that recursively re-quantizes the residual error, serving as a refinement layer on top of existing quantizers. We establish that a fresh random rotation returns each residual to the uniform distribution on the hypersphere, which characterizes the rate of progressive error decay across successive passes. This result lets us determine, before any weight is read, how many passes a layer needs for a target weight-space error. Every prefix is itself a valid lower-rate model, so one artifact serves several precisions. We instantiate ApexQuant with three interchangeable stages, scalar, $E_8$ and trellis, and validate it on four open-weight LLMs and on Earth-observation and medical domains where in-distribution data is often unattainable as imagery arrives under restrictive licences or due to patient material under privacy constraints. Progressive re-isotropization comes within a few percent of full precision at four bits and gives the best two-bit arm we measure, in a completely data-free setting.
cs.LG / 89 / 2610.07910
Revisiting Temporal Regularization for Smooth Control in Deep Reinforcement Learning
Abstract
Deep Reinforcement Learning policies can produce nonsmooth action oscillations that hinder deployment on physical robots. Existing architectural and penalty-based approaches seek spatial smoothness by directly reducing sensitivity to changes in state inputs, but their broad constraints can degrade task performance as stronger smoothing is pursued. Temporal regularization instead constrains action differences along observed transitions, but has been considered unable to provide the spatial smoothness needed under observation noise. We revisit this assumption by proving that the temporal penalty bounds the expected action differences between current states sharing a next state, revealing a spatial effect that empirically extends to spatial smoothness. Building on this finding, we propose Conditioning for Action using only Temporal Smoothness (CATS), which combines a temporal penalty with linear ramp-up. We highlight temporal regularization's ability to provide spatial smoothness while better preserving task performance than explicit spatial regularization. Through linear ramp-up, CATS allows the policy to learn rewarding behavior before progressively smoothing its actions, improving return preservation and both temporal and spatial smoothness. Experiments in both simulation and the real world show that CATS substantially reduces action oscillation without degrading task performance, with little computational overhead.
cs.LG / 90 / 2610.07938
Generalized Matheron Variational Implicit Processes
Abstract
Implicit-process priors specify distributions over functions through sample-forward mechanisms such as Bayesian neural networks and stochastic simulators, but their function-space densities are typically unavailable. We introduce Generalized Matheron Variational Implicit Processes (GMVIP), a pathwise variational family for posterior inference with such priors. For Gaussian-process priors, GMVIP recovers the standard inducing-variable variational GP construction; for general implicit priors, its empirical covariance construction preserves the prior mean and covariance in the population limit. GMVIP constructs posterior samples by drawing a function from the prior and applying a correction anchored at a set of inducing inputs. The effect of this correction away from the inducing inputs is determined directly from prior samples, allowing the posterior to retain the structure and variability of the original implicit process. The (surrogate) prior and variational posterior use the same pathwise construction and differ only in the distribution of whitened inducing coefficients, yielding a tractable coefficient-space Kullback-Leibler divergence. Experiments on regression, classification, and forecasting with simulator-defined and retrieval-conditioned empirical trajectory priors show that GMVIP is broadly competitive with existing methods.
cs.LG / 91 / 2610.07973
Learning a Ranking from Human Feedback in Log-Concave Random Utility Models
Abstract
We study the problem of recovering the ranking of a fixed set of items according to their unknown numerical utilities. At each interaction with the environment, a learner presents the item set to a human and receives comparative feedback of two types. Under full-ranking feedback, each interaction reveals a noisy ranking of all items, whereas under winner-only feedback, it reveals only the item ranked first. In both settings, we model human feedback using a random utility model with log-concave noise and study the number of observations needed to recover an $ε$-accurate ranking with high probability. This novel criterion tolerates ordering errors only between items whose utilities differ by less than $ε$. For both feedback types, we establish worst-case sample-complexity lower bounds and develop algorithms that match these bounds up to logarithmic factors. Neither algorithm requires knowledge of the noise distribution, while only requiring an upper bound on its variance. Our results show that the ranking problem under winner-only feedback is intrinsically harder by exposing the sample complexity dependence on the minimum winning probability across the item set.
cs.LG / 92 / 2610.07981
Do Higher-Order Models Win for Higher-Order Reasons? Rethinking Performance Gains in Hypergraph Learning
Abstract
Higher-order models (e.g., hypergraph neural networks) often outperform lower-order baselines on hypergraph learning benchmarks, and their advantages are commonly attributed to their ability to exploit higher-order information. However, better performance alone does not establish this explanation. We therefore ask: Do higher-order models win for higher-order reasons? To investigate this question, we introduce a controlled performance-attribution framework that perturbs higher-order information while preserving the lower-order, i.e., pairwise, information. Across 25 commonly used hypergraph learning benchmarks spanning three tasks, we frequently observe an intriguing pattern: higher-order models originally outperform lower-order baselines, yet retain most of their advantage after perturbation. This suggests that much of the observed advantage remains achievable without the higher-order information. We then investigate potential lower-order explanations for these remaining gaps. We find that simple additions to a lower-order baseline, e.g., richer pairwise weighting, more steps of pairwise feature propagation, and normalization, reduce the remaining performance gaps, supporting lower-order explanations for part of the observed advantage. Our analysis calls for the hypergraph learning community to rethink performance attribution by distinguishing performance gains from their explanations, adopt stronger lower-order baselines, and use suitable benchmarks that better test the value of higher-order information.
cs.LG / 93 / 2610.07990
A Broader Look at Model Merging: Rethinking Implicit Regularization Induced by Task Arithmetic
Abstract
Model merging aims to build a multi-task model cheaply by combining the weights of individual task-specific models. To perform well across multiple tasks, most existing merging methods use an additional dataset to find the coefficients for the best linear combination of task-specific weight updates. However, we identify an implicit regularization in this standard practice: searching over coefficients restricts the candidate models to a subspace spanned by task-specific weight updates. In this work, we investigate whether this regularization is actually useful. Surprisingly, empirical results show that optimizing merged-model weights without this regularization significantly boosts the performance of common merging methods across multiple architectures, domains, and even in an extremely data-limited scenario where only one instance is available per class. Moreover, directly optimizing the pretrained model weights even outperforms some existing merging methods. Analysis shows that better multi-task weights exist outside the subspace and can be found using multiple methods. We study different strategies for using the additional dataset, discussing their practical use and implications for model merging. Overall, this work calls for revisiting the existing model-merging pipeline, motivating a broader exploration of the weight space and a reconsideration of the implicit regularization induced by task arithmetic.
cs.LG / 94 / 2610.07996
TICDA: Tabular In-Context Data Attribution
Abstract
Tabular foundation models (TFMs) achieve strong predictive performance by conditioning on labeled demonstrations provided in context, without any parameter update. Yet how individual demonstrations shape a given prediction remains poorly understood. This gap matters in practice: the context is often assembled from whatever labeled data is available, potentially leading to the inclusion of mislabeled, redundant, or low-quality examples that degrade performance. Standard data attribution methods do not transfer to the TFM setting: resampling-based approaches such as DemoShapley require a combinatorial number of forward passes, and gradient-based estimators such as influence functions require computing training point's effect on the model parameters, which in-context learning never updates. We introduce TICDA, a method that measures the influence of every demonstration in the context directly from linear surrogates trained on TFM latent embeddings, in a single forward pass and at negligible cost. We show that TICDA offers the best compromise against competitors across four tasks: detecting labeling errors, curating context to preserve predictive accuracy while lowering inference cost, producing attribution scores that transfer across TFMs, and supporting an acquisition strategy for efficient active learning.
cs.LG / 95 / 2610.07997
Can phenotypic activity be predicted without experimental readouts?
Abstract
Molecular encoders contrastively pretrained on paired molecule-morphology data, such as CLOOME and CellCLIP, have been proposed as cheap surrogates for phenotypic prediction, avoiding the need to run a Cell Painting assay. We evaluate this idea for these molecular encoders under a protocol designed to control for two confounds that can inflate apparent performance: leakage across an encoder's own pretraining boundary, and the correlation between phenotypic activity and cytotoxicity. Testing six representations, including a non-pretrained MLP control matching CLOOME's input and layer count, on two distinct Cell Painting screens, we find that once these confounds are controlled for, the pretrained molecular encoders show no clear advantage over plain physicochemical descriptors, and that toxicity is generally easier to predict than phenotypic activity across representations. Our results suggest leakage-aware, confound-controlled evaluation should be standard practice before phenotype-pretrained encoders are trusted as surrogates for phenotypic drug discovery.
cs.LG / 96 / 2610.08046
SepsisLens: Structure-Preserving Sequence Modelling for Decomposable Early Sepsis Warning
Abstract
Early sepsis warning from ICU records can be cast as a structure-preserving prediction problem. A model needs to detect deterioration from irregular measurements while keeping each alert connected to the physiological signals that support it. Many temporal models fuse clinical variables into a patient-level representation, supporting scalar risk prediction but weakening the structure needed for clinical decomposition. We present SepsisLens, which preserves variable-indexed temporal states until risk composition. Observation-aware representations encode each variable's dynamics and measurement history, while a shared temporal encoder models each trajectory without collapsing the variable axis. The StructuredRiskHead composes multi-horizon risk from explicit variable-level and organ-level components. We evaluate SepsisLens on three public ICU cohorts and one private-hospital cohort under a common pre-onset protocol. SepsisLens achieves strong discrimination on all four cohorts and lower alert burden at matched event recall on MIMIC-IV. Structural ablations support the design, while input-side masking shows that the ranked components reflect variables with greater influence on prediction.
cs.LG / 97 / 2610.08049
A Riemannian Geometry for Low-rank Adaptation
Abstract
Low-rank adaptation (LoRA) is widely used as a parameter-efficient fine-tuning technique for pre-trained deep neural networks, which approximates the weight update via full fine-tuning by a low-rank matrix $BA^\top$. This parameterization leads to the equivalence relation $(B, A) \sim (BG^{-1}, AG^\top)$ for any invertible matrix $G$ because $BA^\top = BG^{-1}(AG^\top)^\top$ and thus both pairs yield the same loss value. This relation induces a quotient manifold where matrices $(BG^{-1}, AG^\top)$ for all $G$ are identified, eliminating redundant directions along which the loss value remains unchanged. To respect the geometry of this manifold, the original search space is endowed with a Riemannian metric that is invariant under the equivalence relation. Such a metric induces preconditioning at each gradient step and ensures that each weight update via LoRA changes the loss value, leading to efficient optimization. In this paper, we propose a new Riemannian metric that is specifically tailored to LoRA to close the gap to full fine-tuning at the weight level. We theoretically show that LoRA with our preconditioning induced by this metric satisfies the following two properties at each iteration: (i) The weight update follows the direction closest to the gradient of full fine-tuning within the subspace of first-order weight changes allowed by the LoRA parameterization. (ii) The updated weight matrix is closer in Frobenius norm to that of full fine-tuning than the updated weight matrices of LoRA with conventional preconditioning and without preconditioning. These theoretical insights suggest that our preconditioning makes LoRA better approximate full fine-tuning, thereby leading to more efficient optimization. Experiments show the effectiveness and efficiency of our preconditioning for LoRA on fine-tuning tasks with language and vision domains.
cs.LG / 98 / 2610.08069
Detecting a Shift Is Not Enough: Exact Minimax Limits of Linear Representation Repair
Abstract
A mean shift between two data sources can be easy to detect but hard to remove without substantially changing their representations. We cast its removal as a statistical decision problem: from noisy differences between paired calibration measurements in $\mathbb{R}^d$, learn one linear map, applied to both sources under a hard distortion budget, that leaves as little of the shift as possible on fresh data. We derive the exact finite-sample minimax risk over all such maps, $(d-k) \mathbb{E}[1/(d+2J)]$ with $J\sim\mathrm{Pois}(κ/2)$, where the budget allows deleting $k$ directions and $κ$ is the calibration signal-to-noise ratio. Projecting out the mean calibration difference attains it without knowing $κ$ or the noise scale. This exposes a detection-repair gap: detecting the shift needs only $κ\gg\sqrt d$, whereas removing a fixed fraction of it at constant distortion needs $κ\asymp d$, as for estimating its direction. Standard linear concept erasers (MP, SAL, LEACE) remove the same calibration difference, so the formula gives, before fitting, exactly how much shift they leave on fresh data and how much calibration a target requires. The limit is robust: pairing keeps it exact for non-Gaussian shared content, the projection keeps its guarantee under anisotropic noise, and selective abstention cannot close the gap. On paired clinical and wearable sleep EEG, where differences between participants act as calibration noise, the formula predicts the device shift left in new participants, and more recordings per person soon stop helping. Together, these results tell whether a correction that falls short needs a better method, more recordings, or more participants.
cs.LG / 99 / 2610.08075
Optimization Encoders: Rethinking Second-Order Meta-Learning for Neural Fields
Abstract
Conditional neural fields represent signals continuously, but their effectiveness depends on how the conditional latent representations are inferred from observed data. In meta-learning, this encoding occurs through gradient updates induced by the decoder, tying representation learning directly to decoder design. We formalize this connection by interpreting latent optimization as an optimization encoder, unifying the roles of second-order differentiation, latent parameterization, and task supervision. This concept enables second-order meta-learning for end-to-end training of the encoding procedure alongside the decoder, and clarifies which learning pathway first-order approximations discard. Guided by this view, we introduce Attentive Latent Fields (MetaLF), an equivariant transformer-based neural field that contextualizes a latent pointcloud through self-attention. These interactions shape both field predictions and the updates that construct their representation, allowing local observations to inform coherent non-local structure. Disentangling the inner encoding objective from outer task supervision unifies reconstruction, classification, and segmentation within an end-to-end meta-learning framework, using reconstruction-only latent adaptation at test time. Controlled experiments on polynomial fields link latent coordination to lower effective rank and stronger alignment with the underlying function space. Across image and 3D shape reconstruction, MetaLF improves fidelity within three to five gradient updates, while supporting semantic prediction across images, shapes, and volumes. Together, these findings position the optimization encoder perspective as a unified basis for designing neural fields around how representations are constructed, coordinated, and used.
cs.LG / 100 / 2610.08118
Attenuated in-context identification in time-series foundation models: diagnosis under counterfactual inputs and repair by synthetic forced-system fine-tuning
Abstract
Covariate-aware time-series foundation models (TSFMs) promise training-free what-if answers for instrumented plants: the change in output that a different future input would cause. We test this on forced engineering systems with exact counterfactuals, comparing Chronos-2, TimesFM-2.5 and TabPFN-TS with classical system identification fitted to the same context. Through their default covariate interfaces, TimesFM-2.5 and TabPFN-TS are memoryless: the predicted effect of an input change is a same-time function of that change ($R^2 = 1.000$ for TimesFM-2.5). Chronos-2 identifies dynamics in context but attenuates them. Its predicted effect is 0.33-0.80 of the true effect, its recovered impulse response has the wrong shape, and its error on a one-degree-of-freedom oscillator levels off at 0.57 with 8192 context samples, where ARX fitted to 256 samples reaches 0.02. Context dither at inference lowers the what-if error on all six synthetic classes without training. A 26-minute fine-tune on synthetic forced systems restores the response magnitude (sensitivity 0.83-0.96) and outperforms structure-agnostic identification on Wiener-Hammerstein and a held-out friction class. A specialised in-context identifier trained on the same data comes close, so the forced-system data carry most of the gain. On three of four measured plants classical identification remains clearly better, and the fine-tuned model loses part of its univariate forecasting skill. Paired counterfactual inputs, together with shuffled future inputs on measured records, test two properties: whether the covariate interface can represent dynamics and whether the pretraining prior covers the plant's time scale. Only the counterfactual pairs expose the attenuation.
cs.LG / 101 / 2610.08129
Do LLMs Act on What They Know? From Partner Representations to Cooperative Actions
Abstract
Cooperation with unfamiliar partners requires adapting to communication conventions that are not known in advance. We study this problem in a controlled Hanabi-derived environment with scripted hint generation, LLM-controlled receiving decisions, and frozen model weights. Across eight LLMs, linear probes recover intent conventions substantially more accurately than target conventions, yet receiving choices do not consistently agree with the sender's convention. We compare probe-predicted and ground-truth conventions presented either as general rules or as externally computed action recommendations. Rule statements yield modest and model-dependent changes in cooperation, whereas action translation produces larger gains on average. In a Qwen3-8B case study, matched-state statement reversals reveal much greater sensitivity to action recommendations than to rule statements. Activation transfers from oracle-action and non-oracle hint-restatement donors improve intent accuracy on both action classes, but the tested alternatives do not reliably reproduce these benefits. Together, these results distinguish convention decodability, sensitivity to convention information, and cooperative performance, and highlight limitations in turning available partner information into receiving decisions.
cs.LG / 102 / 2610.08131
Mu-DisCoCat: A Variational Pipeline for Compositional Generalization on Quantum Processors
Abstract
Achieving compositional concept generalization (CoCoGen), the ability to understand novel situations by recombining learned primitives, remains a fundamental challenge in artificial intelligence. Compositional semantic models such as Compositional Distributional Semantics (DisCoCat) offer solutions by generalising vectors to tensors, but suffer from scaling bottlenecks when learning the tensors. Mapping DisCoCat onto Variational Quantum Circuits (VQCs) resolves this limitation for text, yet the methodology has not been expanded to multimodal situations such as the ones involved in CoCoGen. This paper introduces Mu-DisCoCat: a multimodal variational quantum learning framework for DisCoCat that achieves CoCoGen. The framework first learns stable object representations from single-object image-text pairs, then fixes these and uses them to learn the relations between them in multi-object situations. In classical simulations, the model used Uhlmann state fidelity to compute the overlap between the multimodal circuit representations and achieved higher relational OOD accuracy than the evaluated CLIP baseline. Its deployment was evaluated using the destructive SWAP test across noisy quantum emulators, including a range of IBM fake backends, IQM FakeAphrodite, and the IBM Marrakesh quantum processor. Despite real-world device noise, the hardware-executed models maintained a strong positive correlation with simulated fidelities, reliably distinguishing unseen similar and dissimilar pairs. Our work establishes a framework for executing CoCoGen on VQCs, demonstrating a viable use case for near-term quantum hardware.
cs.LG / 103 / 2610.08161
Symphony for Text Generation: Benchmarking Clinical Note Generation
Abstract
Ambient documentation systems are rapidly gaining adoption, yet their impact on clinical note quality remains poorly characterized. We introduce MedConv, a multilingual dataset of 300 clinical encounters in English, Danish, and German, and use it alongside the Ambient Clinical Intelligence benchmark (ACI-BENCH) to compare Corti, a clinical AI platform, with two leading, accessible ambient scribe software applications built on general-purpose AI. We present a controlled clinical evaluation framework that combines entailment metrics with LLM-judged pairwise comparisons across eight dimensions adopted from PDSQI-9. Results show that Corti's API-based text-generation infrastructure is on par with or outperforms leading commercial scribes. We further show that Corti's configurable API provides the flexibility necessary to fine-tune quality dimensions for specific documentation use cases. We present the evaluation methodology and release a dataset to support future reproducible comparison of ambient documentation systems.
cs.LG / 104 / 2610.08200
Finding the Heads and the Neurons Responsible for Network Information Retrieval in Language Models
Abstract
We ask whether specific attention heads, and more finely specific neurons inside those heads, are responsible for recognizing that a language model's context contains network infrastructure information (a hostname paired with its IP address), and whether that responsibility can be validated causally rather than by correlation alone. At the head level the answer is yes, across five models spanning three architecture families: in every model, a small set of heads (1 to 9 out of 128 to 1152 candidates), found by causal ablation screening and tested for selectivity against matched negative and context-free controls, supports a detector with 99.5--100\% held-out accuracy. We then ask whether a head's responsibility concentrates into one neuron or stays spread across its dimensions; this is model-specific. In one model, the top head's signal concentrates into a single neuron, found independently by both a causal intervention and a correlational ranking, which agree exactly (AUC = 1.000, matching the full head). In another, the single clean head works as a whole (AUC = 1.000) but the best causally ranked neuron inside it does not (AUC = 0.665), so the responsibility there is spread across the head. The remaining three models fall in between. On an independent dataset collected by a different institution (reverse-DNS records rather than the discovery data), every model's full-head detector flags 100\% of positive records; the single-neuron versions transfer less reliably, and in one model score below chance. Causal head-finding for a specific network-information entity works across models and architectures; how far that finding can be pushed down to individual neurons varies, and needs to be checked for each model.
cs.LG / 105 / 2610.08271
Reinforcement Learning with Segment Reward Feedback under Linear Function Approximation
Abstract
Classical reinforcement learning (RL) assumes that a reward is observed for every visited state-action pair. However, in real-world applications such as autonomous driving, such fine-grained feedback can be costly or difficult to collect, whereas trajectory-level feedback may be too sparse for efficient learning. To provide a general feedback model bridging these two extremes and handle large state spaces, we study RL with segment reward feedback under linear function approximation. Our work answers how the granularity of segment feedback and the choice of segmentation influence learning. For equal-length segments with known transitions, we design algorithms $\bitssegd$ and $\edlinucbsegd$ for binary and sum feedback types, respectively. They adopt posterior sampling with planning to achieve computational efficiency and the E-optimal experimental design to attain near-optimality. Nearly matching lower bounds are established. For equal-length segments with unknown transitions, we develop a unified $\seglsvits$ framework with two instantiations for binary and sum feedback, which carefully integrates the posterior estimated reward parameters into least-squares value iteration. These results reveal a fundamental insight: under binary feedback, increasing the number of segments significantly reduces the regret through an exponential factor, while surprisingly, under sum feedback, the granularity of segments does not affect learning much. Finally, to investigate whether segmenting according to state-action features can further expedite learning, we design an algorithm $\uneqsegbitsd$ that allows arbitrary segmentations. The resulting regret bound shows that under the usual elliptical potential analysis, the influence of state-action features on the regret appears only through logarithmic factors, and equal segmentation achieves the best performance.
cs.LG / 106 / 2610.08272
Performative Prediction with Selective Labels
Abstract
Many social applications of machine learning exhibit performative effects: population behavior changes in response to deployed models. Performative prediction studies this interaction through a distribution map that relates each model to the population distribution it induces. One of the main results in this framework showed that repeated risk minimization (RRM), which updates models by retraining on the most recent data, can converge to a stable model that minimizes risk on its own induced distribution. However, existing analyses typically assume access to the complete distributions of features and labels after model deployment, ignoring the possibility of selective labels: observing labels only for the accepted subset of the population. In this work, we formalize performative prediction with selective labels and show that retraining only on observed data can misguide the retraining procedure and undermine the guarantees of convergence to a stable solution. We then propose a worst-case objective based on knowledge of a confidence interval on the probability of a positive label. Applying RRM to this objective permits us to remain within a bounded distance to the true stable point. Under a sensitivity assumption on the conditional label distribution, we further show how previously accepted data can tighten these confidence intervals over time. Experiments in a lending application with fairness regularization show that our robust optimization approach closely matches the performance of RRM with complete label access.
cs.LG / 107 / 2610.08287
Scalable extraction and visualization of multi-attribute logical and functional dependencies in tabular data
Abstract
Understanding the structural relationships among attributes in tabular data is fundamental to machine learning and pattern recognition. While functional dependency (FD) discovery has been extensively studied, scalable discovery of logical dependencies (LDs), particularly as the number of attributes and dependency order increase, remains underexplored. These dependencies capture non-deterministic, condition-specific relationships among pairwise or multiple attributes. Furthermore, existing approaches do not provide a unified framework for extracting multi-attribute LDs and FDs. To address these limitations, we propose LDTool and HLDTool for extracting and visualizing multi-attribute LDs and FDs from tabular data. LDTool extends dependency discovery beyond pairwise relationships, while HLDTool enables scalable extraction through hypergraph-guided search-space reduction. Experiments on three simulated and eleven real-world datasets demonstrate that the proposed framework extracts meaningful LDs and FDs while improving scalability. LDTool recovers the same FDs as existing FD discovery methods with lower runtime in high-dimensional feature spaces, whereas HLDTool enables dependency discovery in datasets with hundreds of features. The proposed framework provides interpretable visualizations of dependency structures and supports applications in exploratory data analysis and the quantitative evaluation of synthetic tabular data.
cs.LG / 108 / 2610.08322
Structure-Aware Graph Abstention for Reliable Selective Forecasting
Abstract
Selective forecasting abstains on high-risk test windows under a retained-coverage budget. Existing gates such as TEM (Brusokas et al., 2025) score each forecast as a whole; for multivariate outputs, trajectories can look plausible while violating dependencies among variables. We treat instance-level plausibility and relational consistency as distinct reliability axes and operationalize the latter via a learned sparse graph and a Dirichlet-style structural energy E_struct, trained with error-weighted graph regularization and score-error alignment. On seven long-horizon benchmarks and four backbones, structural gating often reduces selective MSE versus TEM at matched coverage, with the largest gains where cross-variable structure appears more informative in our benchmarks; gains are not universal, indicating a complementary abstention signal. Table 1 is a Protocol A ranking diagnostic (seed 2024); three-seed deployable Protocol B on an aligned subset is in Table 3 (full validation-to-test grids: Appendix A).
cs.LG / 109 / 2610.08353
Uncertainty Quantification Is Indispensable for Reliable Connectome-Based Graph Learning: A Narrative Review and Case Study
Abstract
While graph neural networks (GNNs) have shown substantial promise in connectome-based diagnostic classification, deterministic models inevitably suppress pipeline-induced noise and model ambiguities, yielding overconfident predictions. Although uncertainty quantification (UQ) is widely adopted in voxel-level segmentation, its role in connectomic graph learning remains largely unaddressed. This paper presents a comprehensive narrative review of UQ frameworks tailored to connectome graph learning alongside an empirical case study demonstrating the perils of uncalibrated predictions. We delineate sources of aleatoric and epistemic uncertainty across neuroimaging pipelines and review prominent UQ paradigms, from Bayesian approximations and ensemble methods to evidential learning and conformal prediction. In our case study, a temporal Graph Attention Network (GAT) trained on dynamic functional connectivity (dFC) matrices from the SUDMEX CONN dataset achieves 80.0% diagnostic accuracy (F1 = 0.794) for Cocaine Use Disorder. However, a post-hoc uncertainty audit via Monte Carlo dropout reveals severe overconfidence (ECE = 0.127), with misclassified subjects assigned prediction confidences up to 95%. This empirical divergence between discrimination and calibration underscores the confidence paradox in deep connectomics. Our findings establish that rigorous UQ, calibration, and selective prediction mechanisms are indispensable for deploying trustworthy graph-based biomarkers in clinical neuroscience.
cs.LG / 110 / 2610.08355
Sensor Geometry as a Flow-Matching Prior for Multi-Channel Brain Signals
Abstract
Flow-matching models start from an isotropic Gaussian source, the standard choice when the correlation structure of the data is unknown in advance. For multi-channel brain recordings, however, part of this structure is known in advance. Electrodes sit at fixed positions on the head, and volume conduction through the skull and scalp makes nearby electrodes co-vary in a way that is shared across subjects. Existing EEG generative models nonetheless leave the network to learn this from scratch. We put this structure into the source instead. From the sensor coordinates alone, we build a k-nearest-neighbor graph and take a graph-Matérn function of its Laplacian as the source covariance, so the flow starts from spatially coherent patterns rather than channel-independent noise. The change adds no learned parameters, works with any coupling and any drift network, and uses the same three hyperparameters on every dataset. Across eight EEG datasets and four flow-matching methods, the graph-Matérn source lowers the spectral discrepancy between generated and real signals in the five clinical bands (PSD-KL) on most datasets. PSD-KL falls by 12% to 17% in geometric mean over datasets depending on the method and by up to 40% on PhysioNet-MI, the densest montage. We show that the improvement stems from the spatial eigenvectors of the local graph of sensor positions, since randomizing the eigenvectors while preserving the eigenvalue spectrum eliminates the gain. Furthermore, a prior fitted directly to the empirical data covariance performs worse than isotropic noise. The same construction applies unchanged to MEG, intracranial EEG with patient-specific grids, and a traffic-sensor network, lowering PSD-KL for every method on each. https://jd730.github.io/projects/GraphPrior
cs.LG / 111 / 2610.08367
Evolutionary One-Step Generators: Fast and Diverse Sampling for Discrete Design
Abstract
Several discrete design tasks, such as molecular discovery, require diverse collections of useful candidates at low computational cost. High validity alone does not guarantee a useful candidate library: repeatedly generating the same valid structures leaves few distinct alternatives. Training for both feasibility and diversity is challenging because many relevant criteria can only be evaluated after hard decoding. To address this challenge, we propose EGO (Evolutionary Generators with One-step inference), a framework for training compact generators directly on discrete outputs. The method combines distribution matching with structural constraints and optional diversity or history-dependent rewards, using antithetic low-rank evolution strategies without requiring criterion-specific differentiable surrogates. Once trained, the generator produces the entire graph in a single neural-network evaluation. On molecular generation benchmarks, our compact generator achieves over $50\times$ the valid-and-unique yield per estimated dense operation compared to recent one-step flow-map baselines while retaining high chemical validity. In scaffold completion, EGO achieves an observed $44.3\times$ speedup over MoLeR in generation to SMILES and produces approximately $10\times$ as many filter-passing proposals within matched time budgets for generation and screening. Beyond chemistry, EGO produces $1.54\times$ as many distinct held-out elite architectures as relaxed gradient training on NAS-Bench-101. The low generation cost may enable real-time candidate generation across discrete design tasks, supporting interactive exploration of constrained design spaces and rapid construction of candidate sets for downstream evaluation.
cs.LG / 112 / 2610.08368
Accelerating the Development of PLGA In Situ Forming Depots Through AI-Driven Multi-Objective Optimization
Abstract
Developing long-acting injectable formulations requires the simultaneous optimization of drug loading, release kinetics, viscosity, injectability, stability and other objectives. To navigate this multidimensional space, Corbion and Intrepid combined Corbion's diverse PURASORB bioresorbable polymer library with Intrepid Labs' proprietary AI algorithm (ANDROMEDA 1) to develop in situ forming depots for a therapeutic peptide. Over approximately 15 weeks, 181 unique formulations spanning drug loadings of 6-12% w/w were prepared and characterized through broad design-space mapping and targeted multi-objective optimization. Four lead candidate formulations were identified at 6%, 9%, and 12% w/w drug loading. Each met the predefined viscosity and injectability criteria while providing distinct 30-day in vitro release profiles. The study evaluated polymers spanning a broad range of molecular weights, including commercially available PURASORB grades and new polymers under development by Corbion to expand its polymer toolbox. ANDROMEDA 1 identified that polymers with intermediate molecular weights provided a favorable balance between sustained release and solution viscosity. Together, these findings demonstrate how integrated polymer expertise and AI-driven optimization can rapidly identify differentiated formulation candidates, focus the development space, and establish a strong data-driven foundation for further optimization and in vivo evaluation.
cs.LG / 113 / 2610.08384
Decision-Focused Learning in MDPs: An Occupancy Measure Approach
Abstract
In this work, we consider decision-focused learning (DFL) for a Markov decision process (MDP), where existing methods differentiate through the KKT conditions of the Bellman equation and require solving a linear system over all state-action pairs, limiting its scalability. We address this by reformulating the MDP as an occupancy measure-based linear program (LP), whose feasible region is induced by predicted dynamics, and we derive a closed-form gradient by identifying the active constraints in the feasible polyhedron via the pivoting algorithm. This occupancy measure-based LP layer raises two challenges: (1) LP's solution gradient is discontinuous when active constraints change, and (2) the LP backward cost still scales with the state size, which is costly for large or continuous state spaces. We address the challenges with an augmented Lagrangian surrogate and smooth the boundary jumps by random row sketching of the constraints, and a learnable soft state-aggregation layer and its function-approximation generalization that scales the LP to large finite and continuous-state MDPs. Across multiple tasks, our methods reach lower regret than KKT-based DFL and two-stage baselines with significantly lower computation cost. The source code for all experiments is available at https://github.com/A-Eshragh/State_Aggregation_Project.
cs.LG / 114 / 2610.08400
Atom-JEPA: Joint-Embedding Predictive Architecture for 3D Atomistic Systems
Abstract
Large-scale self-supervised pretraining has reshaped modern machine learning, substantially advancing the ability of language and vision models to generalize across downstream tasks. While deep learning has driven considerable progress in modeling atomistic systems in recent years, self-supervised pretraining in this domain has not yet achieved comparable downstream generalization. To address this, we introduce Atom-JEPA, a self-supervised pretraining framework that learns latent representations from unlabeled 3D structures through complementary atom-level and substructure-level objectives inspired by joint-embedding predictive architectures. We pretrain Atom-JEPA on large-scale molecular and crystalline datasets and evaluate its transfer performance by fine-tuning on a diverse set of downstream property prediction tasks. Atom-JEPA achieves state-of-the-art performance on molecular ADMET and quantum-chemical property prediction tasks, and is highly competitive in predicting the physical properties of crystalline materials. These results demonstrate the potential of latent-space predictive pretraining to support broad downstream generalization from structural data alone. Code and pretrained model checkpoints are publicly available at https://github.com/khelverskovp/atom-jepa
cs.LG / 115 / 2610.08402
VETTA: Coordinating Turn- and Token-Level Credit Assignment for Multi-Turn LLM Agents
Abstract
Multi-turn LLM agents often receive sparse task feedback across several interactions, while generating each response token by token. This creates two related credit-assignment questions: which responses helped achieve the outcome, and which generation decisions mattered within each response? Existing methods typically focus on only one level: turn-level methods evaluate complete responses but do not distinguish the decisions within them; token-level methods can propagate feedback across turns but do not explicitly model credit for each response. These complementary limitations motivate learning credit at both levels and coordinating it in a single policy update. We introduce VETTA, a credit assignment method that jointly learns turn- and token-level values through separate heads on a shared lightweight critic. VETTA computes advantages along both temporal sequences and combines each turn advantage with a within-response-centered token residual for PPO updates. Furthermore, to reduce value-learning cost, the critic retains only early Transformer blocks from the pretrained checkpoint used to initialize the actor. On two challenging agent benchmarks, ALFWorld and WebShop, VETTA improves success rates over PPO by 37.5% and 22.3%, respectively, with Qwen2.5-1.5B-Instruct and achieves success rates of 95.5% and 76.0%, respectively, with Qwen2.5-7B-Instruct. Critic-depth comparisons further show strong task performance with substantially lower critic-side computation. These results suggest that a compact shared critic can coordinate turn- and token-level credit to improve agent performance while keeping value estimation efficient. Code is available at https://github.com/Jiaju-Chen/VETTA-official.
cs.LG / 116 / 2610.08420
Symmetry-Aware Feature Learning: A Polynomial Separation for Multi-Index Models
Abstract
We establish a polynomial sample complexity separation between symmetry-aware and symmetry-agnostic feature learning. We study growing-rank multi-index models with high-dimensional Gaussian covariates in $\mathbb{R}^d$ and $r=Θ(d^δ)$ teacher directions forming a cyclic symmetry orbit, where $0<δ<1/2$. We compare three ways of exploiting this structure: architectural weight sharing, data augmentation over the full symmetry group, and learning without access to the symmetry. In particular, we analyze a symmetry-tied convolutional network, an untied network, and the same untied network trained with full-group data augmentation, using spherical online SGD with correlation loss. For a class of polynomial links with information exponent $p\ge3$, we prove matching sample complexity bounds up to logarithmic factors: the tied and augmented learners achieve weak directional recovery in $\widetildeΘ(d^{p-1})$ samples, whereas the symmetry-agnostic learner requires $\widetildeΘ(rd^{p-1})$. For the pure quadratic Hermite link, the same separation holds for weak recovery of the teacher subspace, with sample complexities $\widetildeΘ(d)$ and $\widetildeΘ(rd)$, respectively. Thus, full-group data augmentation matches the sample efficiency of architectural weight sharing, and both provide a polynomial advantage over training without symmetry. For $p\ge3$, the proof reveals a two-stage mechanism: fluctuations at initialization select one direction in the teacher orbit, after which localized growth amplifies its overlap to the weak recovery scale while competing overlaps remain near their initialization scale.
cs.LG / 117 / 2610.08479
MetaLearnNCA: Few-Shot Offline Meta-Learning via Interacting Neural Cellular Automata
Abstract
Few-shot meta-learning traditionally formulates task adaptation either as analytical gradient descent through unrolled computational graphs or as metric-based distance comparisons over flattened 1D fea- ture vectors, which either incur costly test-time backpropagation or discard native 2D spatial geometry. In this work, we propose METALEARNNCA, a decentralized framework that achieves few-shot adapta- tion through the dynamical interaction of coupled Neural Cellular Automata (NCAs) without computing analytical gradients during inference. MetaLearnNCA decomposes task adaptation into an Active- NCA, which executes task inference conditioned on a continuous 2D spatial memory grid termed the spatial program, and a learned Meta-NCA, which acts as a decentralized cellular optimizer by diffusing spatial error residuals across local neighborhoods to dynamically update this program. METALEARN- NCA is competitive against canonical meta-learners in-distribution (96.12% on Omniglot) with Out-Of- Distribution transfer gains on MNIST, KMNIST, and Fashion-MNIST transfer across 10 independent testing seeds across 1-, 5-, and 10-shot regimes (e.g., surpassing Prototypical Networks by +10.54% on 10-shot MNIST and a +3.87% gain on 10-shot Fashion-MNIST over FOMAML). Our results establish that robust, gradient-free learning-to-learn can emerge from decentralized cellular dynamics on non-von Neumann substrates.
cs.LG / 118 / 2610.08527
PHBA: Prefix-State Hybrid Block Attention
Abstract
Hybrid architectures combining linear sequence models with softmax attention provide an effective balance between efficient long-context modeling and precise token retrieval. Existing designs such as Native Hybrid Attention (NHA) combine compressed long-term states with sliding-window attention, but their exact attention is restricted to a fixed local window. In this work, we introduce Prefix-State Hybrid Block Attention (PHBA), which replaces local sliding-window attention with top-k block-sparse retrieval and couples each retrieved block with a compact prefix state summarizing its preceding context. The prefix states are constructed by a gated linear recurrence at block boundaries and retrieved together with the corresponding token blocks, allowing the model to combine precise long-range evidence with compressed historical context within a unified layer. We further develop a hardware-aware Triton implementation that streams routed token blocks and prefix states without materializing large intermediate tensors. Experiments show that PHBA improves long-context and retrieval performance over strong linear and hybrid baselines while retaining efficient training and inference.
cs.LG / 119 / 2610.08534
How Bregman Divergences Shape Shampoo
Abstract
Understanding the principles behind Shampoo has recently guided the development of more effective neural network optimizers. These methods learn a preconditioner by optimizing the Frobenius or Kullback-Leibler (KL) divergence against the gradient second moment. In this work, we investigate how the choice of divergence shapes preconditioning, which remains unclear and blocks further improvements. To do so, we develop a unified Bregman divergence framework that connects all popular divergences, allowing us to study them jointly. Through empirical spectral analysis of gradient second moments, we examine how divergence choice shapes Kronecker approximation and interacts with finite-sample error in preconditioning. We find that some divergences can better compensate for finite-sample underestimation of the empirical second moment, helping explain the differing behavior of their corresponding Shampoo variants. We further validate this explanation through GPT-2 pretraining experiments. By connecting divergence choice to practical training behavior, we believe our framework provides principled guidance for understanding the foundations of, and further improving, Shampoo.
cs.LG / 120 / 2610.08537
FlowCF: Sparse Counterfactual Explanations for Mixed-Type Tabular Data using Flow Matching
Abstract
In the field of Explainable AI (XAI), counterfactual (CF) explanations interpret a model's decision by suggesting the changes to the input that would lead to a more favourable outcome. To be useful in practice, such an explanation should change few features and change them as little as possible, properties known as sparsity and proximity. We observe that existing methods remain limited in this respect, especially for numerical features, whether they are model-agnostic and amortised, or gradient-based with full access to the model. In this paper, we propose FlowCF, a model-agnostic generative method that frames CF generation as sparse transport from the factual to the target class. We solve this transport with flow matching, which we extend to mixed feature types with a novel mixed flow operator, and exploit the resulting geometry to optimise for sparsity through a gating network that minimises the number of features the transport changes. Extensive experiments on six benchmark datasets demonstrate that FlowCF produces the best numerical sparsity and proximity, changing 29% of the numerical features where the best baseline changes 89%, at 70% smaller displacement, while remaining comparable on the other desiderata.
cs.LG / 121 / 2610.08538
From Shared Demand Patterns to Local Uncertainty: Probabilistic Load Forecasting by Mixing Compact Adaptations
Abstract
Probabilistic load forecasting has been widely studied for power-system operation and planning, but customer- and transformer-level forecasting introduces a distinct scalability challenge. At these levels, load uncertainty is strongly affected by customer behavior, weather, and mixed load composition, making it difficult for a single shared model to capture heterogeneous patterns. Using separate probabilistic models can improve local accuracy, but becomes costly to train, store, update, and validate at scale. To address this challenge, we develop a scalable customer-aware forecasting framework that learns common demand behavior through a shared model while adapting only a compact subset of parameters. Rather than using an independent model for each load or assigning each load to a specialized model, the proposed design learns a small bank of low-dimensional adaptation components and allows each load to combine them according to its forecasting characteristics. This preserves shared knowledge across customers while providing sufficient flexibility for heterogeneous and mixed load compositions. Experiments on 590 load profiles from the SMART-DS dataset show consistent improvements in deterministic accuracy and probabilistic quality over statistical, neural-network, Transformer-based, and pretrained time-series baselines, while retaining low storage and inference costs.
cs.LG / 122 / 2610.08553
DeltaTTT: Layerwise Optimization for Nonlinear Recurrent Memory
Abstract
Sequential test-time training adapts a memory network through successive updates, each computing an inner-loop gradient based on the network's previous state. Intuitively, this state dependence should allow each update to account for what the memory has already learned and better incorporate new information. However, we find that this expected advantage does not consistently materialize in nonlinear memories: a fixed-base parallel TTT baseline outperforms its serial counterpart. Our exploratory experiments point to a key underlying difficulty: nonlinear memories can be harder to optimize than linear ones within a single pass over the sequence. To alleviate this optimization difficulty, we introduce DeltaTTT, which replaces joint inner-loop optimization of a two-layer memory network with layerwise learning. Each layer is assigned a local prediction target and updated through a state-dependent delta rule. This formulation retains a nonlinear readout while enabling chunkwise parallel computation. Experiments on DeltaNet and LaCT backbones show improvements in language modeling and retrieval over their recurrent baselines.
cs.LG / 123 / 2610.08561
Reinforcement Learning for Hierarchical Reasoning Rewards: Minimax-Optimal Rates with Transformers
Abstract
Reinforcement learning (RL) has become a standard tool for post-training language models on reasoning tasks, where the policy is updated by reward feedback while exploring the space of responses. Despite its empirical success, theoretical understanding of RL post-training remains limited, in particular of why on-policy exploration combined with a neural reward model is effective. In this paper, we address this question by modeling the reward as a hierarchical function on the response space: the reward consists of infinitely many local components, each of which becomes relevant only after the preceding ones have been resolved. We show that a natural Transformer-based actor--critic algorithm, which alternates between sampling from the current KL-regularized policy, fitting a Transformer critic to the observed rewards, and updating the policy, achieves the minimax optimal rates in the query budget and in the regularization strength up to logarithmic factors, and is minimax optimal for a fixed number of prompts. In contrast, we prove that sampling from the fixed reference distribution, as in offline reward modeling, can limit regret decay to a logarithmic rate. These results show that on-policy exploration progressively zooms in on the region where the reward is concentrated, and quantify its benefit for RL post-training.
cs.LG / 124 / 2610.08565
Singular Value Decomposition: A Geometric Rediscovery, Where Proofs Become Algorithms
Abstract
This article is a geometric rediscovery of the singular value decomposition, with a further claim: the construction it builds is the machinery behind much of machine learning. The same argument that answers an idle question about ellipses is the algorithm behind principal component analysis, kernel methods, and PageRank, and it is not only the results that transfer but the proofs themselves, run as procedures. The usual introduction states $A = UΣV^T$ and justifies it via the spectral theorem applied to $A^T A$. This is correct but unilluminating, since it assumes a powerful theorem to reach a result that is, in the end, about ellipses. Part I reverses the order. A linear map sends the unit circle to an ellipse; one asks which input directions map to its axes, and finds, example after example, that they are perpendicular. In the plane this can be watched: rotate a frame, track how far its images are from perpendicular, and a sign change forces a frame where they are exactly perpendicular, which is also where the map stretches hardest. Maximizing the stretch and recursing generalizes this to n dimensions, with singular values falling out in order, and the construction proves the spectral theorem rather than assuming it. Part II puts each construction to work: maximize-and-recurse becomes the power method and PageRank; the lemma locating the maximizer becomes the stopping rule of gradient descent; the duality between $A^T A$ and $A A^T$ becomes the transport at the heart of kernel PCA. Each connection is stated with its boundary, saying what the decomposition supplies and where another idea takes over. Prerequisites are the standard sophomore sequence, and the worked examples are small enough to check by hand.
cs.LG / 125 / 2610.08570
Less Is More: A Leakage-Controlled Study of Dermoscopic Preprocessing for Joint Skin Lesion Classification and Segmentation with YOLO26
Abstract
Handcrafted preprocessing is widely employed in automated dermoscopic analysis to suppress imaging artifacts and enhance lesion visibility. Nevertheless, its actual contribution to modern real-time models remains unclear, particularly when evaluation protocols do not adequately control correlations among images of the same lesion. This study presents a leakage-controlled, lesion-disjoint evaluation of dermoscopic preprocessing and augmentation for joint multi-class lesion classification and instance segmentation using a fixed nano-scale YOLO26 segmentation model (YOLO26n-seg). From HAM10000 (10,015 images), quality control yields 10,013 valid image-mask pairs from 7,468 unique lesions, partitioned into mutually exclusive sets by lesion identity. With the architecture, resolution, training budget, and evaluation protocol held fixed, we compare minimally processed images plus online augmentation against offline class balancing, DullRazor-CLAHE preprocessing, and raw-processed hybrid views, over three random seeds. On the lesion-disjoint test set, the raw baseline achieves a mask mAP$_{50:95}$ of $0.5636 \pm 0.0234$, a Dice score of $0.9356 \pm 0.0024$, and a macro-F1 score of $0.6917 \pm 0.0202$. Offline augmentation does not improve the mean performance, while the combined and hybrid strategies reduce both class-aware segmentation and classification accuracy. At only 2.69 million parameters, the model runs at approximately 50 frames per second. Under a leakage-controlled, lesion-disjoint protocol with all non-input factors held fixed, minimally processed dermoscopic images combined with standard online augmentation deliver a better accuracy-efficiency trade-off than increasingly complex deterministic preprocessing, which yields no consistent joint benefit across three seeds on HAM10000.
cs.LG / 126 / 2610.08577
How Learning Governs Unlearning across the Memorization-Generalization Spectrum
Abstract
While unlearning seeks to negate undesired capabilities acquired through learning, little research has examined how the way models learn shapes their subsequent unlearning. In this paper, we investigate this connection from the perspectives of memorization and generalization, the two most representative yet competing strategies that models employ during training. We first classify memorization- and generalization-heavy models using grokking in modular addition and compare their responses to unlearning, showing that the latter suffer greater retain damage, i.e., a larger performance drop on the retain set. Furthermore, we conduct a finer-grained analysis by introducing bucketed modular addition, in which the respective contributions of the two strategies can be explicitly controlled across the memorization-generalization spectrum. In this setup, we reaffirm that the same trend persists and is nearly monotonic. We further demonstrate that this relationship also holds in LLM unlearning across verbatim and factual recall settings. Finally, we provide two practical insights for developing better unlearning methods, highlighting the importance of accounting for learning dynamics in unlearning.
cs.LG / 127 / 2610.08578
Random Feature Gaussian Process Attention: Linear-Time Probabilistic Attention with Calibrated Uncertainty
Abstract
Transformers provide a state-of-the-art modeling framework, yet poor calibration limits their reliability in safety-critical applications. A promising direction addresses this issue by interpreting attention as a Gaussian process (GP) posterior, which enables principled uncertainty calibration but incurs cubic complexity in sequence length due to the inversion of the kernel; although decoupled GP variants reduced the cost to quadratic, the computation remains prohibitive in practice. In this paper, we propose the plug-and-play random Fourier feature Gaussian process attention (RFF-GPA) module, which represents the attention as a GP with a stationary kernel approximated by random Fourier features. This low-rank approximation results in linear-time complexity for approximating the posterior mean and variance, making it far more scalable compared to previous work. Empirical results on multiple real-world datasets show that our attention module improves calibration while maintaining predictive accuracy, and simultaneously reduces computational complexity to linear in the sequence length.
cs.LG / 128 / 2610.08592
CNet: A Complex-Valued Deep Learning Framework with Wirtinger Autodifferentiation and FFT--Hadamard Convolution
Abstract
CNet is a C++/CUDA framework for building and training deep complex-valued neural networks (CVNNs) and, more generally, for optimizing complex-valued functions by gradient descent with Wirtinger (CR-calculus) derivatives. It takes a physics-native stance: a network is a cascade of complex -- and often unitary (the DFT) -- operations acting on an amplitude vector, and classification is a Born-rule measurement $p_k = |z_k|^2 / \|z\|^2$ rather than a softmax over real logits. Every layer ships a CPU reference and a CUDA kernel checked against finite differences, and the computation graph is cloned across the batch for GPU execution. On top of the base layers we add signal-processing primitives that turn the identity conv(x,k) = IFFT(FFT(x) . FFT(k)) into a learnable complex convolutional network, together with a true-Adam optimizer and a reduced-memory inference mode. We report three studies. First, a fully complex-valued, FNet-style causal sequence model built on a new $O(N \log N)$ causal Fourier mixer -- a triangular-masked DFT evaluated by a Bluestein / chirp-z factorization: once properly tuned it matches or exceeds a parameter-matched real-valued causal FNet on character-level language modeling, reaching the real model's converged quality in under half the training steps. Second and third, bottleneck analyses on radio-modulation classification (RML2016.10a) and the Fourier phase problem of coherent-diffraction imaging, which isolate exactly where complex-valued networks still need new operators. Across all three the complex formulation provably learns the physically correct structure. Code: https://github.com/crasmarum/CNet
cs.LG / 129 / 2610.08593
Multi-Label Perceptual Bug Detection in Video Games using Deep Learning on Gameplay Footage
Abstract
Traditional approaches for automated bug detection in video games, such as manual testing, can be beneficial for the improvement of quality assurance, but they can be expensive and time-consuming. The scarce number of tools available to detect multiple perceptual bugs in the same video frame introduces detection challenges for automated bug detection tools in real-world scenarios. We propose a deep learning model for multi-label perceptual bug detection and compare it against video classification models such as Inflated 3D ConvNet and 3D ResNet. Our proposed model, ResNet-BiLSTM, achieved an F1 score of 85.78% on the benchmark dataset. Our results demonstrated that temporal dependency modelling is beneficial for accurate video-based bug detection. We believe this work with multi-label perceptual bug detection on gameplay videos will help save resources spent on manual testing workloads in video games. Furthermore, we introduce a new dataset with multi-label perceptual bugs in this work. The dataset contains 77,969 video clips across different genres of games with approximately 1.2 million frames, containing combinations from 5 classes of bugs in the same video frame.
cs.LG / 130 / 2610.08624
Early Memory Selection for Balanced Adam
Abstract
We propose a method for choosing the shared memory parameter $β_1=β_2=β$ in Adam from a short pilot training. The selected $β$ remains fixed during the subsequent full training. A local model of Adam's normalized direction balances sampling variability against the delay introduced by averaging past gradients. This balance gives a cubic memory rule, whose two coefficients are estimated from gradient probes at a few pilot checkpoints. The estimator uses the numerator and denominator jointly, preserving their covariance. With a 200-update pilot and sixteen probe gradients at each of four checkpoints, a seed-matched retrospective evaluation on eleven vision and language workloads reduces mean relative validation gap by 40.7% and worst-quarter mean gap by 44.3% against the grid representative of shared $β=0.95$. The mean gap is also 32.3% lower than that of the best constant $β$ chosen across all eleven workloads.
cs.LG / 131 / 2610.08669
MemFLoRA: Memory-Floor LoRA for CNN Adaptation at the Edge
Abstract
On-device learning is necessary when the model encounters user-,sensor-, or environment-specific shifts after deployment. Although parameter-efficient fine-tuning (PEFT) methods, particularly Low-Rank Adaptation (LoRA) variants, enable efficient adaptation at the edge, the limiting resource for Convolutional Neural Network (CNN) adaptation is often not the number of trainable parameters but the activation state that must be retained until the backward pass. This paper introduces Memory-Floor LoRA (MemFLoRA), a low-rank CNN adapter built around a memory-first design principle rather than a direct application of transformer-oriented LoRA. Instead of merely reducing trainable weights, we define an activation-memory-floor criterion: trainable backward computations must not depend on full-width layer inputs. The resulting adapter freezes the down-projection, trains a scale-matched up-projection, and combines eval-mode backbone normalization with activation-minimal backward rules, reducing saved state to the low-rank branch. Evaluated on three Human Activity Recognition (HAR) datasets and two CNN backbones under subject, body-location, and sensor-placement shifts, MemFLoRA reduces saved-activation memory by 98.5-98.7% and peak training-state memory by 94.9-97.3% relative to full fine-tuning, while matching or exceeding CNN PEFT baselines.
cs.LG / 132 / 2610.08670
Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment
Abstract
Language models increasingly act as agents. An agent that says an action is wrong and then takes it anyway is a different failure from one that does not know better, and evaluations of stated values cannot see it. We build a pre-registered panel of 248 scenarios across five kinds of pressure. Each scenario is posed twice to the same model, once as the agent choosing what to do and once in the third person asking which option is right, so the model's own judgment is the reference. Every scenario has a twin with the pressure removed, and every model gets a positive control in which its operator orders the violating action, so that a missing gap can be told apart from a blind instrument. On OLMo-3-7B-Instruct, the model takes the action it judged wrong on about one in five pressuring scenarios, more often than on the same scenarios with the pressure removed. Across four instruct models the gap depends on the post-training recipe: OLMo-3 and Meta's Llama-3.1-8B-Instruct carry it; Tulu 3 shows none on the whole panel (above about 0.01 in probability) or on its own most-pressuring scenarios; Qwen2.5-7B-Instruct shows none on the whole panel (above about 0.02) and is unresolved on its own (0.083, -0.028 to 0.195). Meta's recipe and Ai2's Tulu 3 start from the same Llama-3.1 weights, and only Meta's carries the gap. Reading a chat model outside its chat template reverses the sign of its gap with nothing at stake (-0.038 against +0.055 under the template on OLMo-3), a distortion present on two of three recipes. On both models that carry it, reasoning about the stakes before acting moves the choice back toward the model's own judgment, against a same-length non-moral task, with or without the pressure; on OLMo-3, naming the norm at stake does about a third of that. The gap is a measurable target for post-training recipes, not a fixed property of pretrained weights.
cs.LG / 133 / 2610.08677
Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling
Abstract
Off-policy evaluation (OPE) for contextual bandit policies becomes challenging when action-level importance weighting incurs excessive variance. Doubly robust (DR) estimation remains unbiased under common support but retains these high-variance action-level weights. A prior estimator, Off-policy evaluation with Conjunct Effect Model (OffCEM), replaces them with more stable cluster-level weights, at the cost of relying on local correctness of the reward model. In this paper, we show that, under the assumptions required by DR and OffCEM, there exists an unbiased family of estimators that interpolates between OffCEM and DR. Building on this result, we propose the Variance Optimal-CEM (VOCEM) estimator, which selects the interpolation coefficient to minimize variance. We derive the population-optimal coefficient in closed form and show that the resulting estimator has variance no larger than either endpoint, OffCEM or DR. Experiments in controlled synthetic settings and on two large-action benchmarks show that VOCEM improves upon both endpoints in all 23 evaluated conditions, exhibiting greater stability and empirical robustness.
cs.LG / 134 / 2610.08680
A Systematic Study of Small Language Models on Abstract Reasoning Tasks
Abstract
Endpoint accuracy on abstract-reasoning benchmarks does not reveal whether a language model has acquired a transferable rule or fit distribution-specific regularities. We study this distinction in small language models on the ARC-TGI benchmark, which organizes abstract grid transformations into controllable task families and supports resampling, spatial shifts, and cross-benchmark transfer. Across more than 1,000 runs, we profile decoder-only, encoder--decoder, and mixture-of-experts model families under supervised fine-tuning. We examine the efficiency and stability of skill acquisition, robustness beyond the training distribution, interactions with model family and task formulation, and layer-wise attention signatures that accompany behavioral differences. Substantial in-distribution accuracy is attainable, but acquisition is sensitive to optimization and unevenly distributed across task families. Performance deteriorates sharply outside the training distribution, including when the rule is retained but grid scale changes. Greater training-set depth and breadth yield uneven gains, while the effect of additional in-context examples depends on model family. Executable-rule induction also yields correct solutions not observed under direct grid generation. On selected tasks, attention diagnostics show distinct concentration and context-dependence profiles, but do not establish general causal mechanisms. Overall, abstract-reasoning scores are conditional on the model, adaptation regime, evaluation distribution, and response format.
cs.LG / 135 / 2610.08689
Probabilistic Counterfactual Inference for Discrete Outcomes in Gaussian-Process Causal Models
Abstract
Counterfactual inference in Gaussian-process structural causal models (GP-SCMs) has been developed primarily for continuous endogenous variables, limiting applicability to causal graphs that contain discrete child nodes with continuous parents. We introduce a unified probabilistic framework for counterfactual inference with heterogeneous variable types by pairing GP predictors with explicit exogenous noise mechanisms. For discrete outcomes, we derive exact conditional noise-abduction procedures using a uniform threshold for binary variables, a Gumbel-max race for nominal categories, and a latent Gaussian cut-point model for ordinal ones. In each case, we propagate abducted noise through interventions while accounting for posterior uncertainty in the GP latent functions, and prove that the resulting mechanisms reproduce the fitted model's observational and interventional distributions. On synthetic SCMs with known ground-truth counterfactuals, we evaluate estimation accuracy, consistency, and robustness to coupling misspecification. A key finding is that applying a categorical coupling to ordinal data inflates counterfactual error roughly threefold even when observational fit remains comparable, and that this error does not diminish with more data. As the training set grows, the fitted structural equation converges to the truth while the counterfactual error flattens onto a floor. In the reverse direction, forcing a false order onto nominal data instead degrades the fitted equation itself. The choice of coupling must therefore be justified on structural grounds rather than read off the fit.
cs.LG / 136 / 2610.08694
GeneICL: A Tabular Foundation Model for Bulk Transcriptomics
Abstract
Gene expression is widely measured in biomedicine, yet clinical outcome prediction remains challenging due to high dimensionality, strong feature correlations, and limited labeled data. Large self-supervised transcriptomic foundation models often fail to outperform simple supervised baselines. Tabular foundation models offer an alternative through in-context learning, but are typically pretrained on generic synthetic data rather than transcriptomic structure. We ask whether transcriptomics-aware pretraining, rather than scale, is the missing ingredient. Towards this end, we introduce GeneICL, a 4.2M-parameter tabular foundation model combining a semi-synthetic pretraining prior built from measured bulk expression profiles with a parameter-efficient recurrent architecture. We further enable right-censored survival prediction via a training-free reduction to regression using Cox partial-likelihood residuals. We evaluate GeneICL on 80 clinical outcome-prediction tasks spanning classification, regression, and survival. Tabular foundation models consistently outperform self-supervised transcriptomic models, while GeneICL achieves the best overall rank among evaluated foundation models and tuned baselines. GeneICL does so with up to 387$\times$ fewer parameters, no gradient updates at inference, and predictions within seconds on a laptop CPU.
cs.LG / 137 / 2610.08735
Optimal and Efficient Online Inverse Optimization
Abstract
In online inverse linear optimization, a learner recommends an action and then observes the choice of an expert who maximizes a fixed, unknown linear objective on $\mathbb{R}^{d}$; the goal is to learn to optimize this objective without observing it. Sakaue recently obtained the optimal regret $O(\sqrt d)$ with a randomized algorithm making $(dT)^{O(d)}$ linear optimizations per round, and asked whether it can be attained in polynomial time. We answer positively: our deterministic algorithm has regret $O(\sqrt d)$ for every horizon $T$ and runs in time polynomial in $d$ and $T$. It is a variant of the variable-metric algorithms of Sakaue et al.\ and Cai et al., in which a metric update is revoked once the query point moves far enough from where the update was made.
cs.LG / 138 / 2610.08740
On the Computational Tractability of Robust Bandits
Abstract
Learning when the environment does not belong to the learner's hypothesis class is typically handled using agnostic learning guarantees. However, for anything beyond supervised learning, agnostic guarantees are difficult to come by. Recently, imprecise bandits (Kosoy, 2025) (later renamed to robust bandits in Appel and Kosoy, 2025) were introduced as another approach to unrealizable learning in the bandits setting and a $Θ(\sqrt{T})$ regret learner was shown for a large class. However, no computational guarantees were provided. In this paper we identify a special case that admits a polynomial-time learner with $\tilde{O}(\sqrt{T})$ regret. We also show that several small generalizations of this special case are NP-hard thus indicating that the special case is at the boundary of what is tractable. It has been recently suggested (Kosoy, 2018) that computationally efficient learners for unrealizable learning problems are crucial for solving the AI alignment problem. This work is a small step in that direction.
cs.LG / 139 / 2610.08743
Reinforcement Learning with Conformal Action Sets: An Application to Sequential Recommendation
Abstract
Sequential recommenders typically use a fixed slate size even though the number of useful alternatives changes within a session. We propose Reinforcement Learning with Calibrated Pruning (RLCP), which adapts the retained action set using critic scores and an online threshold. The threshold is updated from binary feedback indicating whether the set contains an action in a proxy target. We prove a deterministic bound on the observed proxy miss rate along adaptive trajectories. To quantify the effect of pruning on reward, we derive an exact decomposition of value loss into filtering and selection losses. Under explicit proxy and critic approximation conditions, this decomposition yields a finite session reward bound that also accounts for imperfect selection and set truncation, without requiring the learning parameters to converge. Experiments on KuaiRand-Pure and MovieLens 1M compare two RLCP implementations with four RL baselines. In each of the 19 configurations, at least one RLCP variant achieves the highest catalog diversity, reaching $1.11\times$ to $5.21\times$ that of the strongest baseline, with competitive session depth and no larger retained sets.
cs.LG / 140 / 2610.08745
Linear Bandits under Exact Sliding-Window Constraints
Abstract
We study linear bandits under exact sliding-window constraints, where every consecutive block of actions must belong to a prescribed feasible set. In the offline setting, where the reward function is known, we show that convexity and cyclic-shift invariance make a stationary solution optimal when $w\mid T$ and within an additive $O(w)$ gap otherwise. In the online setting, we show that geometric structure alone is insufficient for learning, and sublinear regret can be impossible. We introduce a transition diameter $τ$ that quantifies feasible reachability and develop a rare-switching OFUL algorithm with regret $\widetilde{O}(d\sqrt{T}+τd+w)$ against the offline-optimal feasible trajectory. Finally, we remove cyclic invariance and consider general sliding-window constraints, where optimal behavior may be non-stationary. We represent recent action history as the state of a finite-memory control problem and introduce a history-state diameter $D$ that measures feasible communication between viable histories. Combining optimistic remaining-horizon planning with rare policy updates, we obtain a regret bound of $\widetilde{O}(d\sqrt{T}+dD+w)$. We evaluate our approach on real-world and synthetic benchmarks, showing that it maintains exact feasibility while achieving reward and regret comparable to baselines with substantially fewer policy updates.
cs.LG / 141 / 2610.08750
Neural Petri flows for chemical reactions
Abstract
Petri nets have been used to describe chemical processes such as reactions.They map well to chemistry: Places are the bonds between atoms and the free valence of each atom, a token is a unit of bond order, a transition forms or breaks a bond, the conserved quantities are the valence budgets of the atoms, and the enabling rule is the valence rule. These semantics are not guaranteed by learned models of reactions or neural networks that are built on Petri nets that use the net as a scaffold for message passing. Here, we ask what architecture remains a Petri net for every value of its weights. We find the answer in the theory, where all semantics of a net share the firing form $m^\prime=m+Cσ$, locality, as enabling reads only the inputs of a transition, and the enabling rule, and we prove that conservation forces the firing form and that non-negativity forces the enabling rule on local rate laws. This leaves free the rate law, which is the propensity of each transition to fire. We introduce Neural Petri Flow, which learns this rate law, or a readout for classification, and hard-wires the rest as parameter-free layers. On what we denote a valence net, atom mapping, reaction classification, and forward prediction become three tasks on one firing vector. Without training, the minimum firing vector maps 88.8% of the curated Golden set against 85.6% for RXNMapper, and 88.7 against 77.9% of the enzymatic reactions of EnzymeMap. On USPTO-480K, NPF trained on these firing vectors predicts 87.7% of the products and 67.4% when trained on a 1% subset of the training reactions. EC numbers of ECREACT are predicted at the third level for 90.2% of reactions, 5.6 points ahead of the best published method. With electrons as tokens, the same token game predicts 90.5% of the elementary steps of FlowER first, ahead of the published baseline, and every top-1 prediction is a valid molecule without a filter.
cs.LG / 142 / 2610.08785
Conformal Prediction Sets Quantify Information Gain: A Theoretical Perspective
Abstract
Conformal prediction is a popular tool for uncertainty quantification that outputs prediction sets with finite-sample coverage guarantees. While prediction set size is commonly used as a heuristic measure of uncertainty, the information-theoretic basis for this interpretation remains poorly understood. In this work, we provide such a foundation using a decision-theoretic generalization of entropy tailored to set-valued prediction. In particular, we introduce a family of generalized information measures based on the size and coverage of conformal prediction sets. Notably, Shannon mutual information admits an exact integral representation in terms of these measures. We then show that, in standard classification settings, the reduction in conformal set size from additional information (i) is sandwiched between calibration-dependent members of this family and (ii) obeys a data processing inequality, both up to finite-sample calibration and model error terms. Together, our results formally relate conformal prediction to classical information-theoretic quantities and justify using set-size reduction as an information gain metric. Empirically, we validate our theory across 11 classification settings and show that set-size reduction and Shannon mutual information can rank features differently in a greedy feature selection experiment.
cs.LG / 143 / 2610.07846
Scen-Opt: A Scenario Optimization Toolbox for Data-Driven Convex Programming
Abstract
The scenario approach is a well-established statistical framework for data-driven decision-making. In particular, in data-driven optimization, the scenario approach unveils how the problem structure governs out-of-sample generalization, and offers a principled basis for assessing and certifying the reliability of the optimal solution as per constraint satisfaction. Despite its strong theoretical development and wide applicability, no software toolbox has been available to date that enables user-friendly, data-driven convex optimization within the scenario-approach framework. In this paper, we introduce Scen-Opt, an open-source software tool that integrates convex programming with data samples while providing statistical guarantees grounded in scenario theory. Scen-Opt is implemented in Python, supporting data-driven linear, quadratic, and semidefinite programming, and offers a Python-based web application with an intuitive and reactive graphical user interface (GUI) built using modern web technologies. Scen-Opt can be used directly through its online interface or installed locally, accommodating both manual input and data-file uploads (CSV, JSON, TXT, TSV, MAT, Excel, NPY, NPZ, Parquet). Built on a Python backend with a modern JavaScript frontend, Scen-Opt offers a highly user-friendly experience and efficient usability across desktops, laptops, tablets, and mobile devices. In this paper, Scen-Opt is applied to a set of representative benchmarks, demonstrating its practical effectiveness for data-driven convex optimization with guaranteed performance.
cs.LG / 144 / 2610.07511
MobileVISTA: Generative Data Augmentation for Pose Generalization in Mobile Manipulation
Abstract
Mobile manipulators such as humanoid robots are increasingly deployed in dynamic, unstructured environments to perform dexterous manipulation tasks. However, end-to-end manipulation policies trained to imitate demonstration data collected from a single robot pose are brittle: even centimeter-scale deviations in robot pose at deployment can drive ego-centric observations and end-effector trajectories out of the training distribution, leading to sharp drops in performance. We introduce MobileVISTA, a data generation framework that transforms demonstrations captured at canonical poses into diverse, pose-perturbed training data by jointly (1) augmenting egocentric visual observations and (2) retargeting actions to compensate for base pose changes. Unlike prior methods, which assume a camera rigidly mounted off the actuated chain or non-trivial articulated robot geometry largely out of frame, MobileVISTA targets compatibility with egocentric platforms (e.g., humanoids) where the camera is both influenced by and must observe the robot's kinematic chain as it moves. We study MobileVISTA in simulated tasks spanning humanoid and bimanual embodiments, and on a real Galaxea R1 Pro. We find policies trained on MobileVISTA-augmented data demonstrate improved robustness to previously out-of-distribution poses encountered at test time, without additional demonstration collection or a trained generative model. Additionally, we find MobileVISTA's benefit is largest on tested humanoids, where the camera rides the actuated chain and the robot fills much of the frame. Additional videos and appendix can be found on our website: https://mobilevista.github.io
cs.LG / 145 / 2610.07594
BiGym 2.0: Benchmarking Learned and Agent-Developed Policies for Humanoid Household Manipulation
Abstract
Humanoid household manipulation requires the arms to act while the body balances, steps and changes posture. We present BiGym 2.0, an adaptation of BiGym for the Unitree G1 across 20 household tasks using a unified whole-body controller for demonstration and evaluation. The suite provides 60 native human virtual-reality demonstrations per task with synchronised multi-camera views and full-body execution records. We benchmark vision-language-action fine-tuning, imitation learning, demo-driven reinforcement learning, and cold-start coding agents given the interaction budget of online reinforcement learning. With the same onboard views, proprioception and whole-body controller for every method, vision-language-action fine-tuning has the highest nine-task mean, and agent-developed programs outperform every demo-driven reinforcement learning baseline on this mean and lead on bimanual reaching. Cross-workspace stacking remains open, $π_{0.5}$ stays low on pick-box, and multi-object transport is hard for imitation learning, demo-driven reinforcement learning and coding agents. All environments, human demonstrations, and evaluation traces are open-sourced at https://github.com/swirl-uk/BiGym2.
cs.LG / 146 / 2610.07597
The Robot Is Not Its Description: GaugeBench for Representation Robustness in Morphology-Aware Policies
Abstract
A robot description does more than specify a physical mechanism: it also encodes arbitrary conventions, such as joint-axis direction, joint-angle zero, and the order and names of links and joints. Morphology-aware policies consume interfaces built from these descriptions, yet cross-embodiment evaluation typically changes the robot while keeping those conventions fixed. This leaves a simple question unanswered: does behavior survive when the robot stays fixed but its description changes? GaugeBench isolates this case by rewriting a fixed mechanism under physically equivalent conventions, verifying that its physics and policy interface are preserved, and then evaluating the same policy weights. The result is stark: three MetaMorph policies score 4030.6 on 80 familiar robots, but only 51.6 when those same robots are equivalently re-described, while 98 genuinely held-out robots score 1489.6. A new description can therefore be more damaging than a new robot. Tracing the failure reveals that axis reversal alone reproduces the collapse, joint-angle zero changes are nearly harmless, and reordering lies between them; moreover, changing joint-state and torque coordinates alone is sufficient to cause the failure, while changing description-derived features alone is not. The same phenomenon appears in ModuMorph and an unrelated PyBullet framework. Yet it is not irreversible: exact two-description transport restores the original controller, and training across equivalent axis conventions raises retained return under axis reversal from 3.6% to 80.6%. Together, these results separate mechanism robustness from representation robustness and show that cross-embodiment evaluation should test both.
cs.LG / 147 / 2610.07613
Learning Grasp Targeting from Point Clouds for Log Pile Clearing on a Hydraulic Crane
Abstract
In mill yards, log loaders clear dense piles by a sequence of bundle grasps: hundreds of logs rest in contact, and each removal changes the pile available to the next grasp. A learned policy chooses where to place and orient the grapple from unsegmented point clouds and runs on a trailer-mounted hydraulic forestry crane. The policy classifies at which observed point to grasp and predicts depth and grapple orientation there. The same network outputs support behavior cloning (BC), reinforcement learning (RL), and deployment. BC learns from successful top-of-pile demonstrations; RL explores for improvements by fine-tuning the cloned policy (BC$\to$RL) or by training from scratch. In simulation, BC clears 98 of 100 piles of 200 logs, while BC$\to$RL improves load stability. Twelve field trials compare a geometric heuristic, RL from scratch, BC, and BC$\to$RL through complete grasp-transport-deposit cycles. BC and BC$\to$RL deposit 93.8% and 88.9% of pooled inventory, against 80.4% for the heuristic. BC$\to$RL deposits logs on 83.6% of its cycles, against 79.6% for the heuristic and 65.7% for BC, while its simulated stability gain does not carry over to the crane testbed. Trained entirely in simulation and run unchanged on the crane, the learned policies clear more than the hand-filtered heuristic while observing unfiltered clouds that still contain the storage rack's rails and poles.
cs.LG / 148 / 2610.08112
Energy-Aware Path Following: Comparative Analysis of Reinforcement Learning and NMPC for Electric Vehicles
Abstract
Path-following control strategies typically follow the bi-objective optimization dilemma: minimizing deviations from a reference path while maintaining smooth speed profiles. The latter objective is especially relevant for Electric Vehicles (EVs), since their limited driving range can be extended by recovering energy through regenerative braking, a feature that has not yet been sufficiently studied in the literature. In this work, we perform a comparative analysis of four controllers under one common Frenet frame-based kinematic vehicle model, utilizing a validated energy model (VT-CPEM) with explicit regenerative braking. Herein, we implement the following controllers: Nonlinear Model Predictive Control (NMPC), Proximal Policy Optimization (PPO), gain-scheduled Ackermann state-feedback baseline (PID-SF), and a Stanley geometric baseline. To satisfy real-time requirements, we implement the NMPC using JIT-compiled CasADi. Moreover, we train the PPO using traditional straight and S-curve tracks, after which we successfully transfer the unmodified policy to unseen tracks, including: an ISO 3888-1 lane-change, a chicane, randomly-generated parameterized-splines, and a $\pm3^\circ$ graded road. In addition, the policy transfers to a dynamic single-track vehicle model with linear tires, zero-shot with an acceptable initial performance, which was optimized after brief fine-tuning. Thereby, we demonstrate that our PPO is readily transferable to more comprehensive vehicle models. We conclude with a performance analysis of developed controllers and discuss ideas for future work.
cs.LG / 149 / 2610.08789
QF3: Fast Flow RL with Filtered Q-Gradients
Abstract
Flow policies have become a standard policy class for learning robot behaviors from demonstrations, but reinforcement learning is still critical for improving pre-trained flow policies or learning them from scratch through interaction. We introduce QF3 (Fast Flow RL with Filtered Q-Gradients), an online off-policy RL algorithm that trains a flow policy with flow matching plus the critic's action gradient, backpropagated through a one-step prediction of the flow's output. To keep updates where the critic and this prediction are reliable, QF3 applies the critic gradient only to action dimensions that stay near the replay action. To our knowledge, QF3 is the first off-policy flow RL method to train humanoid locomotion policies from scratch and transfer them zero-shot to hardware. Paired with a high-throughput off-policy training recipe, it trains humanoid locomotion and motion-tracking policies with a 10x wall-clock speedup over FPO++, a recent on-policy flow RL method. We further apply QF3 to fine-tune pretrained flow-based manipulation policies on both ABC-Sim and Robomimic tasks. These results suggest that QF3 can both learn robot policies from scratch and refine those acquired from demonstrations. Website: https://qf3-rl.github.io/
cs.LG / 150 / 2610.07533
SkillFormer: Skill-Decomposed Adaptation for Audio Language Models
Abstract
Audio language models must handle dozens of distinct skills, from pitch comparison and speaker counting to musical tempo estimation and emotion recognition. Joint training on all skills at once causes interference: gains on one skill often come at the cost of another. We propose \textbf{SkillFormer}, which decomposes audio understanding into skill-specific low-rank adapters and composes them at inference time through a learned router. The router examines the question to decide which adapters to activate and how much weight each should carry, so that a pitch query engages different parameters than a genre classification query. An alternating training schedule updates each adapter on its own skill cluster before jointly calibrating the router, preventing the gradient conflicts that arise in standard multi-task optimization. SkillFormer adds fewer than 4\% of the base model's parameters and requires no changes to the audio encoder or language backbone. Evaluated on three architecturally distinct models across MMSU, MMAU-Pro, and MMAR, it raises the average accuracy by 2.5 to 4.1 points, with balanced gains across perception, reasoning, and semantic subcategories.
cs.LG / 151 / 2610.07966
Feature Encoding in VAE-based Audio Decoders: Effects of Input, Depth and Distribution
Abstract
Neural audio synthesis models like the Realtime Audio Variational autoEncoder (RAVE) achieve impressive genera tion quality, yet how their internal representations encode musical features remains poorly understood. We present a systematic layer-wise and cross-layer cluster analysis of RAVE decoder activations across three models trained on different musical domains, tested with four stimulus types. We then evaluate architectural generalization with a general purpose EnCodec model. For RAVE, we find that synthetic stimuli are encoded well across models and audio features (pitch |\r{ho}|=0.45, 5.1x the null, BPM |\r{ho}| = 0.76, 8.6x the null). These results are reduced but still substantively apparent when using natural audio (mean across features |\r{ho}|=0.25, 2.8x the null). Natural audio sees a stronger encoding when nonlinear probes are used (mean across features R2=0.56, 18x the null, +0.152 nonlinear gain over the linear probe R2). Encoding strength varies throughout the layers of the decoder and an increased ability to joint-encode in the middle layers is seen across all audio features (\b{eta}2 all negative, p < 0.05). The general purpose EnCodec decoder also sees similar strong synthetic responses across audio features, similar nonlinear gains for natural audio joint encoding and similar depth profiles. We find the best cross-layer cluster improves the strength (r = 0.65, p = 0.006) and prevalence (r = 0.75, p = 0.001) of BPM encoding when compared against the best whole layers within the same section, with no effect for joint encoding. These findings advance the interpretability of neural audio models and inform targeted control strategies for neural synthesis.
cs.LG / 152 / 2610.07167
Interleaved Projected Gradient Descent for Safe Imitation Learning
Abstract
We propose an imitation-learning design for neural-network control policies under state and input constraints. Training alternates a standard imitation gradient step with a block of $k$ safety steps that pull the network's actions toward their projection onto the safe set; at run time, the controller is the trained network alone, with no safety filter. We analyze this scheme as inexact projected gradient descent in the space of policy actions. When the projected actions are recomputed at every safety step and each step moves the actions consistently toward the safe set, letting $k$ grow logarithmically yields asymptotic constraint satisfaction on the training states and bounds the distance to the constrained optimum of the imitation loss; with the projected actions held fixed, the same holds only if they are exactly representable by the network. On a nonlinear autonomous racing task, we compare our method with adding a weighted constraint-violation penalty to the imitation loss. With a sufficiently large weight, our method matches the lap time of unconstrained imitation while reducing the fraction of violating episodes from $15\%$ to $1\%$, about six times fewer than the penalty approach at its best weight. Its lap times are less sensitive to the weight, which instead sets how quickly violations vanish during training. In racing, the safety corrections are sparse and the conditions of the analysis do not hold; the gain arises instead through the data collected during training. These gains come at the cost of additional training computation.
cs.LG / 153 / 2610.08337
Machine Learning for German Redispatch Forecasting under Data Delays and Temporal Distribution Shift
Abstract
Public redispatch records provide empirical data for grid congestion forecasting, but delayed reporting, zero-inflated distributions, and temporal shift present major modeling challenges. We assess the accuracy and reliability of probabilistic machine-learning forecasts using published German transmission records under experimentally imposed information-age constraints. The benchmark evaluates eight daily series of upward and downward intervention energy across four German transmission system operators from 2021 to 2024 (48,242 eligible records; 354 evaluation dates in 2024). We compare seasonal empirical, regularized autoregressive (ARX), quantile LightGBM, GRU, and Transformer models under a minimum seven-day target-latency constraint. Neural architectures use a zero-censored output head to accommodate exact-zero outcomes. Static, rolling, and adaptive delayed-feedback calibration are evaluated using normalized weighted interval score (nWIS), empirical coverage, and block-bootstrap inference. Raw LightGBM achieved nWIS 0.7952, outperforming ARX (1.0604) and the seasonal baseline (0.8739) by 25.0% and 9.0%, respectively (Holm-adjusted p<0.005). Rolling calibration improved LightGBM to nWIS 0.7767 versus 0.8251 for static calibration (p=0.0092), with 91.81% coverage for nominal 90% intervals. The zero-censored Transformer achieved nWIS 0.8161, with no significant difference from LightGBM (p=0.260). However, aggregate coverage concealed substantial undercoverage during high-volume interventions (61.91% coverage among above-threshold events). These results show that boosted-tree models with rolling calibration provide accurate probabilistic forecasts of aggregate redispatch volumes under target delays, while nominal aggregate validity does not ensure reliability during extreme congestion events.
cs.LG / 154 / 2610.07211
An overview of machine learning-enhanced iterative methods for systems of linear and nonlinear equations
Abstract
Systems of equations arise in a wide range of scientific and engineering applications. The present work focuses on solvers for general systems of equations, including but not limited to those arising from partial differential equations. These systems can be broadly categorized into linear and nonlinear problems. For large linear systems, iterative solvers are generally preferred over direct methods due to the latter's superlinear growth of computational costs. Although convergence theory is well-developed under certain assumptions on the coefficient matrix, many classes of systems still pose open challenges. These difficulties become even more severe for systems of nonlinear equations, where nonlinear solvers typically rely on repeated linearization. For example, Newton's method may even converge quadratically near the solution; it can also converge slowly or diverge when the initial guess is not chosen appropriately. A wide range of solvers with diverse variants and hyperparameter settings exists, and the development of efficient and robust iterative methods remains an active area of research. Recently, machine learning (ML) techniques have been applied to enhance the efficiency of classical iterative methods while preserving their interpretability and reliability. We refer to these ML-enhanced iterative methods as hybrid iterative methods, in the sense that they combine classical iterative methods with ML. This paper provides a comprehensive overview of state-of-the-art approaches to constructing hybrid iterative methods for systems of both linear and nonlinear equations, while also discussing open challenges and outlining potential directions for future research.
cs.LG / 155 / 2610.08016
FOSLS-deRhaNN: native de Rham neural classes for H(div) and H(curl) with applications to first-order system least-squares neural network methods for partial differential equations
Abstract
We construct neural approximation classes native to the graph spaces H(div) and H(curl), in two and three dimensions and, for H(div), in any dimension. Every realization lies in the space for all parameter values, and with kinked potentials, such as ReLU networks, the admissible jumps appear at finite width. The classes are images of scalar and componentwise networks under fixed operators of the de Rham complex, and do not involve a mesh or finite element emulation. For H(div) in R^n two native classes are given on an equal footing, with a skew-symmetric potential $A$: $\mathrm{Div}\,A+R_nq+\mathbf{h}$, with the divergence $q$ as an explicit unknown, and $\mathrm{Div}\,A+\mathbf{z}$ with an $H^1$ field $\mathbf{z}$; for H(curl) the analogous classes are $\mathrm{grad}\,φ+Sr+\mathbf{h}$ in two dimensions and $\mathrm{grad}\,φ+\mathbf{z}$ in two and three dimensions. In all of them every interface jump of the field is carried by the potential term, $\mathrm{Div}\,A$ or $\mathrm{grad}\,φ$, while the remaining part has no interface jump (it is an $H^1$ field in the regular-decomposition classes); the classes with $\mathbf{z}$ are the componentwise approach enriched by this term. Known or learned interface geometry enters the potential through factors with trainable amplitudes, and the remaining part if the divergence jumps. The classes lead to the FOSLS-deRhaNN method, first-order system least squares with de Rham neural networks, whose loss is the least-squares functional posed in the natural spaces of the weak formulation; for elliptic equations this includes $H^{-1}$ right-hand sides and $H^{1/2}$ Dirichlet data. Elliptic equations with discontinuous coefficients and curl-curl problems are treated as instances, with the functional equivalent to the error; linear transport with discontinuous solutions and conservation laws with shocks use the same flux classes.
cs.LG / 156 / 2610.07290
A Single-Loop, Constant-Batch First-Order Penalty Method for Stochastic Bilevel Optimization
Abstract
Recent advances in penalty-based methods for stochastic bilevel optimization (SBO) have eliminated the need for second-order derivative oracles. However, for stochastic nonconvex-strongly convex bilevel problems, existing first-order methods typically rely on nested loops and/or large batch sizes for attaining $O(ε^{-6})$ or $O(ε^{-4})$ sample complexity under standard bounded-variance assumption or mean-square smoothness assumption. Achieving these rates with a single-loop penalty method and a constant batch size remains challenging due to a large penalty value needed for an accurate approximation. To address this challenge, we develop a stochastic SIngle-loop COnstant-Batch first-order penalty method (SICO) that combines two complementary ingredients. First, it performs one stochastic-gradient update per-iteration for both the original lower-level and penalized problems, with a projection that controls the separation between their iterates. Second, it applies an exponential moving average to stabilize the upper-level gradient estimator. We show that this combination achieves $ O(ε^{-6}) $ sample complexity using only $O(1)$ stochastic-gradient samples per iteration under unbiased, bounded-variance stochastic gradients. Under the additional mean-square smoothness assumption on the lower-level stochastic gradients, the same algorithm improves the complexity to $O(ε^{-4})$ also with $O(1)$ batch size. To the best of our knowledge, this is the first work to match the best-known convergence rate for fully first-order SBO methods using a single loop and a constant batch size. This result addresses an open problem posed in the literature.
cs.LG / 157 / 2610.07721
Exact Calibration and Sharp Risk Geometry for Volume-Sampled Ridge Regression
Abstract
We study ridge regression from exactly $s$ distinct rows of a fixed design. Responses are fixed, and only the subset is random. The determinant law and selected ridge fit share one positive definite penalty. Established mean identities and exponential-family duality give the unique penalty that matches a prescribed full-data ridge fit in expectation. It exists exactly when $s$ exceeds the target's effective dimension. Our main result concerns centered covariance risk normalized by full-data penalized loss. For balanced signed coordinate replicas, a strict sector inequality gives the sharp risk and all maximizing responses at every budget from the dimension to one below the row count. This holds for any nonzero positive semidefinite query. With the target and query fixed, the maximizing response space is unchanged across these budgets. For general designs, we characterize attainment of a leave-one-out envelope. For existing real equiangular tight frames, flat row query energy characterizes when every nonzero residual response maximizes at two deletions. At three deletions, we give the sharp risk and complete maximizing space for isotropic queries, using unequal triangle weights. The balanced geometry yields a same-sample unbiased ridge--Horvitz--Thompson mixture with lower sharp risk and an exact mean-share improvement boundary. Under full recalibration after feature changes, we prove quadratic regret from searching the complete old maximizing space and a query-uniform bound on the mixture's risk gain. The strongest sector inequalities have exact computer-assisted proofs.
cs.LG / 158 / 2610.08020
Learning consistent molecular mechanics force fields from first principles
Abstract
Classical force fields (FFs) remain the workhorse for large-scale simulations even as machine-learned interatomic potentials (MLIPs) approach ab initio accuracy. They decompose total configuration energies into simple effective interactions whose parameters are traditionally assigned based on atom or bond types, enabling efficient simulations but also limiting their ability to adapt across configurations. Recent machine learning approaches have improved the accuracy and transferability of bonded parameters in these FFs by inferring them as functions of local atomic environments, but still rely on empirical nonbonded parameters for practical simulations. In this work, we introduce a unified approach, \texttt{grappa-fullFF}, which learns both bonded and nonbonded parameters \emph{consistently} and simultaneously from ab initio reference data. By incorporating physically inspired regularization via supervision of the electrostatic potential and an architecture that facilitates charge equilibration, our model recovers accurate electric response properties, achieves state-of-the-art accuracy on geometry optimization benchmarks, and reproduces the conformational sampling of both classical and existing machine-learned FFs, without relying on externally assigned nonbonded parameters.
cs.LG / 159 / 2610.07712
Mathematical Invariant-Enabled Topological Neural Networks for Molecular and Materials Property Prediction
Abstract
Existing molecular and materials learning approaches often rely on a limited set of structural representations, which may capture only selected aspects of complex three-dimensional structure. Here, we introduce mathematical invariant-enabled topological neural networks (MITNNs), a framework that represents complex structures through multiple complementary mathematical views and integrates them with topological neural architectures. MITNNs combine multiscale invariants from topology, spectral theory, commutative algebra, differential geometry, and discrete curvature, capturing complementary structural information from the same system. Systematic invariant-subset, architecture-subset, and ensemble analyses show that predictive performance depends on how mathematical representations and neural architectures are paired, with selected combinations outperforming individual models and the aggregation of all available components. Across protein-ligand binding, metal-organic framework properties, mutation-induced protein solubility, and molecular toxicity prediction, MITNN consistently outperforms existing methods. These results establish MITNN as a mathematically multimodal framework for scientific machine learning.
cs.LG / 160 / 2610.07196
Learning Disentangled Representations with Quantum Variational Autoencoders
Abstract
Variational autoencoders are powerful representation learning models that map complex data into low-dimensional latent spaces, enabling the discovery of interpretable and disentangled factors. Such representations can facilitate the interpretation and controllable generation of data describing complex scientific systems. Understanding how these factors are organized and encoded in latent space is therefore important for developing reliable representation learning models. Recently, quantum variational autoencoders (QVAEs) have been proposed as quantum representation models, demonstrating informative latent representations and improved latent-space occupancy through quantum regularization. However, it remains unclear whether and how QVAEs can learn disentangled and interpretable latent factors. A key challenge in investigating quantum latent factors is that a small number of qubits spans an exponentially large Hilbert space, making the notion of an individual quantum latent dimension nontrivial. Here, we investigate what constitutes an individual quantum latent dimension and whether it can encode a distinct factor. We develop theoretical insights into quantum latent dimensions and support them with empirical studies on representative synthetic problems, including MNIST variants. Across three datasets, we demonstrate that QVAEs can discover factorized and semantically interpretable latent representations, with individual qubits functioning as meaningful latent factors. These results establish a foundation for understanding quantum latent spaces and their potential for structured and interpretable representation learning.
cs.LG / 161 / 2610.07328
Advantage of Entangled Learning Rules in Quantum Measurement Class Learning
Abstract
Learning with data in the form of quantum states is of current interest and has led to a variety of problems that boil down to interaction with the available data via quantum measurement and classical post-processing of observed classical outcomes. In quantum measurement PAC learning, one is given a sequence of unknown, prepared quantum states and classical labels, along with a hypothesis class of candidate measurements. The task is to select a measurement from the hypothesis class that minimizes a fixed notion of error in prediction of the classical labels via measurement of a new state by the selected hypothesis. In this work, we consider the advantage of interacting with the given data in the measurement learning framework using learning rules given by measurements that cannot be implemented using local operations and classical communication (LOCC), as opposed to single-copy learning rules. We provide a construction showing that there exist learning scenarios wherein single-copy learning rules are asymptotically suboptimal compared to optimal ones. We then show that learning rules based on entangled measurements enjoy at most a polynomial sample complexity advantage over single-copy learning rules in the PAC learning setting (under a natural joint measurability covering assumption).
cs.LG / 162 / 2610.07391
A perspective note on likelihood approximation and inference for complex simulation models using a chain of aggregated normalizing flows
Abstract
We present a new perspective on the problem of likelihood approximation within the framework of simulation-based inference that promotes scalable and controllable simulation routines for large-scale data analysis, allows efficient parameter space exploration or smooth interpolation in high-dimensions and, thus, supports valid statistical treatments of hypothesis testings as well as uncertainty quantification. In particular, we consider a chain of $n$-aggregated normalizing flows for likelihood approximation scheme, where a set of upfront replicated observation datasets from the forward complex simulation model pass through the first set of bijective transformations, and then subsequently pass to the other sets of bijective transformations. Here, we assume that, for any $k \in \{1,\,2, \ldots, n\}$, the parameters corresponding to the first $k$ sets of bijective transformations are estimated sequentially, in some sense of optimality, for constructing flexible probability distributions, regardless of the remaining $(n-k)$ sets of bijective transformations. Moreover, our objects of interest are to highlight two complementary mathematical arguments that leverage an informatics-theoretic formalization, based-on empirical likelihood estimators under moment restrictions, and a sequential decision-making paradigm, with mixing distributions, for updating and aggregating the estimated parameters of the overall normalizing flows. As a by-product, the framework provides a reliable surrogate model, conditioned on the model parameters defining the forward computational simulation, that allows samples generation, with statistical powers, and facilitates computationally tractable scheme in the Bayesian paradigm for inference, hypothesis testings and uncertainty quantification.
cs.LG / 163 / 2610.08292
Where Do Two Populations of Persistence Diagrams Differ? Calibrated Local Inference at a Fixed Budget
Abstract
Many two-sample tests for populations of persistence diagrams assess global differences without identifying the regions of the birth-death plane that contribute to them. We study simultaneous inference for local mean contrasts when the number of available diagrams is fixed. They are differences in expected weighted feature mass within $\ell_\infty$ neighborhoods at several centers and radii. We estimate these contrasts using additive landmark responses. A Gaussian multiplier bootstrap calibrates simultaneous confidence intervals while allowing unequal group covariances. The neighborhoods whose intervals exclude zero form a map with approximate family-wise error control, and selecting a subset of original intervals for display preserves their joint coverage guarantee. On the simultaneous coverage event, every reported neighborhood lies within twice its radius of the support of the mean-measure difference. A geometric result gives sufficient radius conditions for a displaced feature to produce a nonzero contrast. A comparison of sufficient detection thresholds quantifies the tradeoff between reducing the number of tested coordinates and reserving observations for an independent pilot. In simulations with 40 to 120 diagrams per class, the bands achieved 94%-98% simultaneous coverage under both the strict null and equal means with unequal covariances. In the latter setting, a permutation maximum and the pooled-t implementation of the two-stage persistence-image test of Moon and Lazar rejected in up to 32% and 26% of runs, respectively. In the fixed-budget simulations, spending a third of the observations on a pilot to choose landmarks or radii located changes less often than a prespecified grid at a single radius. On the MUTAG benchmark, the localized region concentrates on rings of fused-ring systems, an exploratory reading.
cs.LG / 164 / 2610.08715
Prediction-powered inference for time series across space
Abstract
The following motif is common in spatiotemporal settings: we have a sequence of covariate and label pairs observed for a relatively short, recent time period. We have access to unlabeled covariates over a longer time period. Data is observed over many spatial locations. For instance, crop yield might be observed over a large geographical area for recent years, but weather data (which is informative about crop yield) is available for a much longer period. The goal is to estimate, at each spatial location, the expected label (e.g., crop yield) in the future and provide a valid confidence interval for this value. The observed time period alone is too short for reliable estimates. Imputing missing labels with machine learning can cause substantial bias. Prediction-powered inference (PPI) can correct for this bias, but it relies on an i.i.d. assumption that breaks under our expected temporal dependencies. Heteroskedasticity and autocorrelation consistent (HAC) procedures account for temporal correlation, but have not been adapted to cases where some labels are imputed. We provide reliable point estimates and confidence intervals given: short labeled time series (across spatial locations), a longer unlabeled time series, and an imperfect predictor of labels given covariates. We show our method outperforms natural alternatives.
cs.LG / 165 / 2610.07228
How Inefficient Is Natural Gradient Descent? From Exact Optimality to Θ( \sqrt{ \log d } ) Divergence
Abstract
Natural gradient descent (NGD) underlies common methods in ML. For dually flat families, idealized NGD on the forward Kullback--Leibler objective follows the mixture geodesic which is often longer than the shortest Fisher--Rao path. We quantify this overhead by the inefficiency ratio \(R \ge 1\), the Fisher length of the mixture geodesic divided by the Fisher--Rao distance, and bound its supremum over endpoint pairs as a function of the parameter dimension \(d\). A tensor criterion identifies the regime (I) families, with \(R=1\) everywhere: exactly those with quadratic potential or dimension one, such as fixed-covariance Gaussians. For non-quadratic families, we prove two further regimes: (II) bounded third-order skewness plus finite Fisher--Rao diameter yields a dimension-independent bound; and (III) for products of scale families---including Gaussian covariances and Gamma rates---\(R\) grows as \(Θ(\sqrt{\log d})\), unbounded in \(d\). Under a per-step Fisher-chord budget, \(R\) translates to a practical computational cost: NGD requires asymptotically at least \(R\) times as many steps as an optimizer following the Fisher--Rao geodesic. Experiments confirm all three regimes: \(R=1\) to machine precision for quadratic-potential families (I), the categorical bound \(π/(2\sqrt{2})\) is approached but not attained (II), and sampled scale-product \(R\) grows with \(d\), reaching \(R \approx 1.5\) for long, high-dimensional moves (III).
cs.LG / 166 / 2610.07292
Assumption-lean logistic regression with missing covariates
Abstract
Missing covariates are frequently encountered in supervised learning problems, and classical methods for estimation using such data use carefully chosen imputation schemes for missing data, or likelihood approximations that lead to nonconvex $M$-estimation problems. These methods and their relatives are suitable for scenarios in which the covariate distribution is known, and more broadly, have enjoyed tremendous success in linear models. But even in basic nonlinear problems such as logistic regression in moderate dimensions, such methods can experience drastic failure modes when the covariate distribution is unknown. Motivated by the need for reliable alternatives, we consider the problem of parameter estimation in logistic regression with missing covariates. Crucially, we operate in the assumption-lean setting where the covariate distribution is unknown (but bounded). We design a stochastic approximation method that is based on $Z$-estimation with a novel monotone operator, and establish that our algorithm is computationally efficient and achieves provable signal recovery at parametric rates under the hypothesis that covariates are missing completely at random. Our theory sharply characterizes the $\ell_2^2$ risk of the estimator in terms of the missingness profile, accommodating heterogeneous observation probabilities. Importantly, it shows that our method always outperforms the de facto ``complete-case'' estimator that ignores observations with any missing data. Even in the setting with homogeneous missingness (in which each covariate is observed independently with probability $q$), our bounds exhibit intricate and nonstandard dependence on $q$ that can yield significant improvements over using only complete cases. We complement our upper bounds with new information-theoretic lower bounds that show that this intricate dependence on $q$ is fundamental in a minimax sense.
cs.LG / 167 / 2610.07383
HyperNSDE: Personalized Neural SDEs for Joint Static-Longitudinal Clinical Data Generation
Abstract
Synthetic patient data generation is a promising solution to the dual challenge of data scarcity and privacy constraints in healthcare machine learning. Realistic synthesis of patient-level clinical data requires jointly modeling heterogeneous static covariates, irregularly sampled longitudinal trajectories, and informative observation times - three tightly coupled components in practice yet rarely addressed together. We propose HyperNSDE, a continuous-time generative model that conditions a latent Neural SDE on static patient representations through a hypernetwork, allowing baseline characteristics to shape trajectory evolution beyond the initial condition without requiring a trajectory encoder, while stochastic latent dynamics capture realistic variability in generated paths. Observation times are modeled jointly through a latent-state-dependent intensity process, and training on irregular stochastic paths is stabilized via a deterministic-stochastic path decomposition with a non-adversarial signature-kernel objective. Experiments on simulated and real clinical datasets show improved observation-time fidelity and competitive performance, while matched-grid analyses reveal that forecasting and correlation metrics are affected by observation-grid regularity and trajectory smoothness.
cs.LG / 168 / 2610.07417
Bayesian Optimization on Function Spaces via Sparse RKHS Manifolds
Abstract
Bayesian Optimization (BO) has become an established methodology for minimizing black-box functions of a vector input. Often, however, this parameter vector arises from the discretization of an inherently functional relationship. Several recent articles have considered the Functional Bayesian Optimization (FBO) setting, in which the variable to be optimized is not a member of a finite dimensional vector space, but rather an infinite dimensional function space. In this work, we propose $L^0$ Manifold Optimization (L0MO), a simple approach to FBO which searches the subset of a Reproducing Kernel Hilbert Space (RKHS) consisting of functions with a sparse representation in the kernel functions, optimizing both the kernel locations and their coefficients. We discuss in detail the relationship between our method and existing ones, providing a unifying lens through which to view prior works. To assess our method against the state of the art, we conduct an extensive computational study, and along the way develop a novel set of benchmark test functions which port standard finite-dimensional ones to the infinite dimensional domain. Our experiments demonstrate that, on balance, the proposed method achieves superior performance across a wide range of test benchmarks.
cs.LG / 169 / 2610.07503
Two-Sample Testing for Random Graphs without Vertex Correspondence
Abstract
Two populations of graphs often have to be compared without any correspondence between their vertices, for instance when networks come from different communities, or when a graph generative model is evaluated against held-out graphs. We study how many graphs such an unaligned two-sample test needs, and which graph statistics can detect which differences. For an Erdős--Rényi null and a planted two-block difference that leaves every expected degree unchanged, we show that $m\asymp t^{-3}$ graphs per group are necessary and sufficient when the per-graph signal-to-noise ratio is $t<1$. Signed triangle counts attain this rate, and the lower bound holds for every graph size. With aligned vertices $m\asymp t^{-1}$ graphs suffice, so misalignment costs a factor of order $t^{-2}$. When the triangle signal cancels, the rate becomes $t^{-4}$ and $4$-cycles are needed. Statistics built from trees have exactly the same expectation under both hypotheses, and tests based on finitely many of them have asymptotically no power. In the graphon limit, this class includes degree distributions and message-passing graph neural network features. For a non-constant null, a generic difference is visible at first order, and a simple motif test attains the aligned order of sample size, suggesting that misalignment is costly mainly for differences that are invisible at low orders. We also give an exactly valid test for one or two graphs per group, at a cost in power. In our simulations, the fitted exponents are close to the predicted ones, and degree-based and random-GNN evaluation metrics stay at their level in a setting where signed triangles need about $65$ graphs.
cs.LG / 170 / 2610.07551
Is $\sqrt{d}$ Separation Necessary for Gradient EM to Learn Gaussian Mixtures in High Dimensions?
Abstract
Learning Gaussian mixture models (GMMs) using the Expectation-Maximization (EM) algorithm and its gradient-based variants is a fundamental problem in machine learning. It is known that randomly initialized (gradient) EM fails to learn multi-component GMMs in the exact-parameterized setting, where the number of components matches that of the ground-truth GMM. Recently, global convergence of gradient EM has been established in the over-parameterized setting, where more components are used, provided that the ground-truth components are well separated. In particular, the minimum separation between ground-truth components is required to scale as $Ω(\sqrt{d})$, where $d$ is the dimension. In this paper, we show that this dimensional dependence is unavoidable in high-dimensional settings. Specifically, we consider a hybrid EM algorithm that uses standard EM updates for the mixing weights and gradient EM updates for the component means. For any $ε> 0$, we prove that when the dimension is sufficiently large, in the worst case a separation of order $Ω(d^{0.5-ε})$ is insufficient to guarantee global convergence of population gradient EM in sub-exponential time under random initialization, even in the over-parameterized regime. Our result establishes an almost optimal worst-case lower bound on the ground-truth separation required for learning Gaussian mixtures via gradient EM in high dimensions.
cs.LG / 171 / 2610.07623
Explicit Asymptotic Bounds for Sequential Calibration Beyond $T^{2/3}$
Abstract
Probability forecasts are calibrated when predicted probabilities match empirical outcome frequencies: among events assigned a probability $p$, we'd hope that the fraction of positive outcomes is close to $p$. We study the problem of sequential forecasting of binary outcomes. The classical $O(T^{2/3})$ bound on expected cumulative $\ell_1$-calibration error established by Foster and Vohra stood for over two decades until Dagan et al. reduced the exponent $2/3$ by an unspecified constant. We establish a new two-phase recursive labeling strategy for the sign-preservation-with-reuse game that yields the bound $O(n^αt^β)$ for all choices of space and time. We then sharpen the reduction from upper bounds on sign preservation to calibration by modifying the equivalence of Dagan et al. to use only $O(\log T)$ instances of the sign-preservation-with-reuse game. This lets us establish an explicit bound of $O(T^{0.662942288})$, the first explicit exponent below $2/3$ for sequential calibration, by combining both improvements and choosing explicit feasible parameters.
cs.LG / 172 / 2610.07717
Stability of Measure-to-Measure Transformers on Sub-Gaussian Data
Abstract
Transformers have exhibited impressive empirical success across various domains, but their theoretical foundations remain less developed. This work constitutes a mathematical study of the measure-to-measure operators defined by transformers. We show that transformers map sub-Gaussian inputs to sub-Gaussian outputs; this ensures that taking arbitrary-length compositions of the softmax operator is well-defined. We then show that transformers are Hölder continuous with respect to the 1-Wasserstein distance on appropriate spaces of sub-Gaussian inputs. This allows us to establish estimates on the error propagation along a transformer between a sub-Gaussian input and its empirical approximation. We also study a mean-field analog of the cross-attention mechanism, which is an operator from a pair of probability measures to a single probability measure. We show that cross-attention exhibits different Hölder regularity and sample-complexity in its two input arguments. Last, we apply our results to deduce approximation guarantees for measure-to-measure transformers. Together, these results provide a firm stability and finite-sample theory for transformers on sub-Gaussian data.
cs.LG / 173 / 2610.07737
Nash Social Welfare for Multi Armed Bandits: Trajectory-wise Expected and High Probability Regret
Abstract
We study fair multi-armed bandits under the Nash Social Welfare (NSW) objective, which measures performance via the geometric mean of accumulated rewards. Existing work defines Nash regret as $\mathrm{NR}_T = μ^\star - (\prod_{t=1}^T \mathbb{E}μ_{I_t})^{1/T}$, where $μ_{I_t}$ is the mean reward of the recommended arm $I_t$ and $T$ is the horizon. Since it applies the geometric mean to per-round marginal expectations, it ignores the joint distribution of rewards across rounds, leaving the NSW fairness motivation unaddressed at the trajectory level. We propose \emph{trajectory-wise Nash regret} $\widetilde{\mathrm{NR}}_T = μ^\star - \mathbb{E}[(\prod_{t=1}^T μ_{I_t})^{1/T}]$, which computes the geometric mean over complete sample paths before taking expectations, capturing NSW fairness more faithfully. By Jensen's inequality, $\widetilde{\mathrm{NR}}_T \geq \mathrm{NR}_T$, making it a strictly stronger metric. We also introduce \emph{high probability Nash regret} $\widehat{\mathrm{NR}}_T = μ^\star - (\prod_t μ_{I_t})^{1/T}$, giving the first high probability regret bounds in fair bandits. Our two-phase algorithm, Round Robin Nash Confidence Bound (\texttt{RR-NCB}), combines round robin exploration with a Nash confidence bound index policy. We show $\widetilde{\mathrm{NR}}_T \leq \widetilde{\mathcal{O}}(\sqrt{k\log T/T})$ and, with probability $1-δ$, $\widehat{\mathrm{NR}}_T \leq \widetilde{\mathcal{O}}(\sqrt{k\log(kT/δ)/T})$, matching the optimal $\widetilde{\mathcal{O}}(\sqrt{k/T})$ rate despite the stronger metrics. Optimality follows from a lower bound via AM-GM and standard $k$-armed bandit minimax arguments. Simulations validate our theory.
cs.LG / 174 / 2610.07740
High-dimensional online calibration from harmonic weights
Abstract
We study the online calibration of multidimensional forecasts over an arbitrary convex set $Y\subseteq\mathbb{R}^d$ relative to an arbitrary error norm $\|\cdot\|_{L}$. For forecasting $d$ binary outcomes simultaneously ($Y=[0,1]^d$), we give the first algorithm that achieves $\varepsilon$-calibration in a number of rounds that is polynomial in $d$ for every fixed accuracy. It requires $d^{O(1/\varepsilon)}$ rounds, exponentially improving the dimension dependence of previous bounds. For multi-class forecasting ($Y=Δ_d$), we obtain the same $d^{O(1/\varepsilon)}$ rate, improving the $d^{\widetilde{O}(1/\varepsilon^2)}$ bounds of Peng and Fishelson et al. Our algorithm is simple: on each round, it outputs a harmonically weighted distribution over harmonically smoothed past outcomes. The same algorithm works for every forecast set and norm. More generally, it achieves $\varepsilon$-calibration after $\exp(O(γ(Y,L)/\varepsilon))$ rounds, where $γ(Y,L)$ is a geometric parameter defined by a matrix discrepancy problem. The harmonic weights are motivated by the fact that the discrete Hilbert transform matrix achieves the optimal discrepancy up to a universal constant, simultaneously for every $L$. This optimality result may be of independent interest.
cs.LG / 175 / 2610.07814
Stochastic Gradient Descent Ascent is Suboptimal for Nonconvex-PL Min-Max Games
Abstract
How far can stochastic gradient descent ascent (SGDA) go by tuning its timescale ratio and step sizes in nonconvex min-max games? We answer this question for nonconvex-PL (NC-PL) games by establishing the first tight complexity of two-timescale SGDA with a fixed timescale ratio and non-increasing step sizes. For $\ell$-smooth games with an inner $μ$-PL inequality, we prove a complexity lower bound $Ω(κ^2\ell\varepsilon^{-2}+κ^4\ellσ^2\varepsilon^{-4})$, where $κ=\ell/μ$ is the condition number, $σ^2$ is the gradient variance, and $\varepsilon$ measures the outer gradient norm. This matches existing SGDA upper bounds and establishes a complexity separation from Smoothed-AGDA (Yang et al., 22'). In addition, we show that SGDA can fail to find a stationary point when its timescale ratio is as small as $o(κ^2)$. Our negative results highlight the fundamental limitation of SGDA in NC-PL games, and justify the development of alternative methods.
cs.LG / 176 / 2610.08078
ProximalFM: Amortized Proximal Causal Inference under Hidden Confounding
Abstract
Standard causal identification methods often assume no unmeasured confounding and can fail when relevant confounders are unobserved. Proximal causal inference instead uses proxy variables to identify effects under hidden confounding. However, nonparametric proximal estimation can be challenging in practice: recovering causal estimands such as the conditional average treatment effect (CATE) requires solving an ill-posed integral equation that is data-hungry, hyperparameter-sensitive, and optimization-unstable. Bayesian inference for such models provides a desirable alternative, mitigating these difficulties by regularizing through the prior. However, computing a posterior is itself challenging, as a typical likelihood function will include latent variables. Following the recent success of tabular foundation models in backdoor, instrumental variable, and frontdoor settings, we propose that prior-data fitted networks (PFNs) are uniquely suited to resolve this bottleneck. Indeed, by training on synthetic data sampled from compliant structural causal models with access to oracle counterfactuals, we simplify the task substantially, amortizing the implied Bayesian operator inversion into a single transformer forward pass. Compared to prior literature that focuses primarily on point estimation, our model, ProximalFM, explicitly targets the Bayesian posterior distribution of the CATE. One unique aspect of this problem is that we need to provide Monte Carlo estimates of the oracle CATEs, leading to a novel variation of PFNs that accounts for the added stochastic error. Across a diverse suite of proximal regimes, ProximalFM achieves consistently strong CATE-estimation performance without dataset-specific tuning, with its largest advantage when latent confounding is substantial and the proxies are weakly informative; it also provides fast inference through a single amortized forward pass.
cs.LG / 177 / 2610.08210
Anytime-valid simulation-based hypothesis testing
Abstract
For a given data distribution $(X_t)_{t \in \mathbb{N}} \sim Q$ i.i.d., we investigate the hypothesis testing problem: $H_0: Q = P_0$ vs. $H_1: Q = P_1$, for two different model probability distributions $P_0$ and $P_1$. In contrast to the standard setting, where analytic densities $p_0$ and $p_1$ are given, here, we consider the density-free setting, where we only have access to i.i.d. simulations $(Z^0_t)_{t \in \mathbb{N}} \sim P_0$ and $(Z^1_t)_{t \in \mathbb{N}} \sim P_1$. For this simulation-based hypothesis testing setting, we construct an e-test martingale, resulting in a sequential test with anytime-valid type-I error guarantees, approximate growth optimality, geometrically decaying type-II error bounds, and asymptotic power one. Most ingredients used in our constructions are variants of well known concepts. The value of this paper lies in the compact presentation of an effective, anytime-valid solution for the density-free simulation-based sequential hypothesis testing case.
cs.LG / 178 / 2610.08227
How Many Independent Samples Does a Satellite Image Contain? Generalization Bounds for Spatially Dependent Data
Abstract
Machine learning classifiers for remote sensing imagery are typically evaluated as though every pixel were an independent sample. Spatial autocorrelation violates this assumption, since neighboring pixels carry redundant information which inflates sample sizes. How many independent samples does a satellite image actually contain? For an $n \times n$ image whose spatial correlation persists over a range of $r$ pixels, the effective sample size is $Θ(n^2/r^2)$, not $n^2$. We prove this as a finite-sample upper bound for classifiers on spatially correlated data, and show via a matching lower bound that the rate is tight, and no algorithm can do better. We extend the results to images with directional correlation and spatially varying correlation structure. Our result justifies spatial cross-validation since block holdout with separation proportional to the correlation range achieves optimal generalization guarantees, while random holdout can underestimate confidence interval widths by a factor proportional to $r$. We validate the theory on synthetic data and satellite image tiles from three sensors (Landsat 8, Sentinel-2, and Sentinel-1).
cs.LG / 179 / 2610.08277
Two-Sample Testing via Generative Processes
Abstract
Deciding whether two samples come from the same distribution is a classical problem in statistics, and generative transport offers a new way to approach it. We build a stochastic interpolant directly between the two samples and observe that, for a symmetric schedule, its law is invariant under the time reflection $t \mapsto 1-t$ whenever the two distributions coincide. We therefore test whether the marginals at times t and 1-t agree by computing their Jensen--Shannon divergence. Both marginals are explicit mixtures over all cross-pairs of observations, so nothing is learned, and permutation calibration gives an exact finite-sample level. For Gaussian noise, this divergence equals a time integral that pairs the reflection defects of the velocity field and of the score, so the test compares transport dynamics rather than endpoints alone. With a narrow-plus-broad noise design, the test attains the minimax separation rate n^{-2s/(4s+d)} over bounded, compactly supported densities whose difference has Sobolev smoothness s > 3d/4, with no lower bound on the densities. Fusing a dyadic grid of noise scales through their permutation ranks, without sample splitting, preserves exact level and adapts to unknown s at an iterated-logarithmic cost. Empirically, the test matches or outperforms state-of-the-art kernel two-sample tests.
cs.LG / 180 / 2610.08345
High-Dimensional Statistical Inference for Sparse Support Vector Machines
Abstract
Using a replica-symmetric high-dimensional characterization, we develop an inferential framework for sparse support vector machines when the sample size and number of features grow proportionally. The main challenge is the nonsmooth hinge loss, which prevents direct application of debiasing arguments developed for smooth classification losses. We overcome this difficulty by representing the $L_1$-penalized support vector machine (SVM) as a linear program and identifying the hinge-loss subgradient through its dual variables. This yields a computationally accessible debiased estimator whose coordinates are asymptotically Gaussian under the proportional asymptotic regime. The resulting distributional characterization provides confidence intervals and hypothesis tests for individual features and enables false-discovery-rate-controlled variable selection. Extensive simulations examine calibration, power, and variable-selection performance under a range of covariance structures, including strongly correlated designs. An analysis of high-dimensional breast cancer gene-expression data illustrates how the proposed inference can distinguish statistically significant features from variables selected by the original sparse SVM.
cs.LG / 181 / 2610.08495
Information-Dense Synthesis for Molecular Discovery
Abstract
Machine learning can accelerate molecular discovery by designing molecules and planning experiments. However, many scientific challenges demand molecules with very rare properties, and in this sparse setting, existing algorithms offer little gain over random guessing. We propose a method to efficiently search large regions of molecular space using algorithmically controlled stochastic synthesis. Rather than design, make and test individual molecules, we design and make complex mixtures, test them as a pool, then deconvolute the molecule-activity map. We optimize synthesis to encode maximal information. Theoretically, this approach can reduce the number of experiments required to find the optimal molecule among $d$ candidates from $\mathcal{O}(d)$ to $\mathcal{O}(\log d)$ or $\mathcal{O}(1)$. In simulation, on estimated protein fitness landscapes, it finds active molecules with an order of magnitude fewer experiments than existing Bayesian optimization methods.
神经与进化计算 (cs.NE)
6
cs.NE / 1 / 2610.08418
From the Drosophila Visual Connectome to General-Purpose Computer Vision
Abstract
Biological connectomes encode structured solutions to visual computation that may provide reusable inductive biases for artificial vision. We develop ConnectomeX around FlyVision, a trainable architecture that preserves parallel ON/OFF processing, recurrent computation and population-level graph interaction while scaling model capacity across tasks. FlyVision reached 99.34% accuracy on MNIST with 80,608 parameters and 78.03% on CIFAR-10 with 81,408 parameters. On ImageNet-1K, FlyVision Base and Large reached 60.79% and 66.25% top-1 accuracy with 1.8 and 3.7 million parameters, while a Large local-k7 model with a learned low-frequency branch reached 66.53%, compared with 69.25% for ResNet18 with 11.7 million parameters. On a 22-class skin-disease benchmark, FlyVision Large achieved 63.78% accuracy and 95.28% macro-AUROC with 2.99 million parameters. In four-class chest radiography, ImageNet-pretrained FlyVision Base and Large reached 92.60% and 92.76% accuracy with 1.33 and 2.97 million parameters, compared with 91.56% for ImageNet-pretrained ResNet18 with 11.18 million. BrainAGE extends FlyVision to volumetric T1-weighted MRI by applying a shared ImageNet-pretrained FlyVision Large encoder to 24 sagittal, coronal and axial slices per scan and combining slice-level age estimates by confidence-modulated Gaussian voting. On 433 held-out scans, three-axis fusion achieved a mean absolute error of 5.98 years and R^2 = 0.868. Across the 224x224 classification tasks, the best FlyVision configuration remained within three percentage points of ResNet18 on ImageNet-1K and skin-disease classification and exceeded it on chest radiography with substantially fewer parameters. These results show that a conserved connectome-informed computation can scale from compact recognition to large-scale natural and biomedical vision.
cs.NE / 2 / 2610.07321
Synapse Loss Estimation for the BrainScaleS Wafer-scale Neuromorphic System
Abstract
Neuromorphic hardware combines memory (synapses) and computation (neurons) into the same silicon substrate to avoid the von-Neumann bottleneck. This way, neuromorphic computing aims to make computational neuroscience simulations and AI processing faster and more energy-efficient. The common crossbar architecture integrates a synapse matrix (analog, mixed-signal or digital) with a neuron array into a neurosynaptic core, whereas multiple cores are interconnected via a dedicated spike routing network. One of such neuromorphic architecture is the BrainScaleS wafer-scale system which allows the realization of very flexible network architecture, supporting both dense and sparse connectivity by highly configurable neurosynaptic cores. Yet, when network models from computational neuroscience are mapped to the BrainScaleS wafer, synapse loss can occur which means that for some model synapses no hardware synapse is available due to the restricted connectivity of the hardware. In this work we analyze the synapse loss when mapping uniform random networks to BrainScaleS both theoretically and empirically. We first develop a methodology to estimate maximum-sized networks without synapse loss based on probability distributions and the hardware's synapse routing architecture. Next, we adopt the methodology to predict the synapse loss for denser or larger network models. Then, we compare the predictions with results from actual runs of the mapping software: The results show the capabilities and limitations of the BrainScaleS system itself and spot shortcomings of the current mapping algorithms. The developed methodology can be helpful in multiple ways: to find the optimal settings when mapping a given network model to BrainScaleS, to act as a benchmark for the mapping software, and as tool for design space exploration for new hardware.
cs.NE / 3 / 2610.07412
Simplified Swarm Optimization for Surrogate-Assisted Reliability Design of Insulated-Gate Bipolar Transistor Power Modules Using an Open-Source Process Finite-Element Model
Abstract
Process-induced warpage, ceramic stress and solder strain limit the reliability of insulated-gate bipolar transistor (IGBT) modules on direct-bonded copper (DBC) substrates. Surrogate-assisted design studies train regression models on finite-element analysis (FEA) databases, but rarely check the optimized designs against new FEA or report how surrogate error interacts with the optimizer. This paper builds and evaluates an open pipeline: an open-source process finite-element model, surrogates tuned by Simplified Swarm Optimization (SSO), multi-objective design search, FEA confirmation of selected designs and confirmation-driven infill. The model starts at the second reflow and reproduces measured warpage within 18.2%, 33.6% and 15.3% at the reflow, housing and molding stages without fitted parameters. On a balanced 60-design database, all stochastic tuners reach the same test accuracy, outperform the published grid on cross-validated performance for every output but generalize better only for warpage; the cross-validated ranking of tuners does not transfer to the test set. With equal result reporting, multi-objective SSO and the non-dominated sorting genetic algorithm II give comparable Pareto fronts; a corrected multi-objective particle swarm optimizer trails both. FEA confirmation shows that warpage predictions hold (mean absolute error 0.5%), whereas at the design-space bounds reached by the optimizers the ceramic-stress surrogate is optimistic by up to 26%. Two confirmation-driven infill rounds reduce this error to 1-10% and halve the out-of-sample error; no confirmed design improves on the database in ceramic stress. Ceramic-stress results are indicative, as the metric is mesh-sensitive at production resolution. The model, database, scripts and pre-registered and post-registration results are released.
cs.NE / 4 / 2610.07808
Common-Mode Errors Limit Low-Timestep Deep Spiking Q-Networks
Abstract
Spiking neural networks (SNNs) offer sparse and event-driven computation, making them attractive for energy-constrained reinforcement learning (RL) on edge devices. In value-based RL, deep spiking Q-networks (DSQNs) combine such efficiency with action-value estimation for decision making. However, existing DSQNs often require multiple simulation timesteps for competitive performance, increasing computational and energy costs, whereas reducing the timesteps can cause substantial performance degradation. We investigate this degradation from the perspective of Q-value estimation errors. By decomposing errors across actions into common-mode and differential-mode components, we find that low-timestep DSQNs suffer disproportionately from common-mode errors shared across action values, which are particularly detrimental to temporal-difference learning through bootstrapped targets. Based on this finding, we propose Common-Mode Compensation Deep Spiking Q-Network (CMC-DSQN), which uses an auxiliary ANN to compensate for common-mode errors in the SNN outputs. At inference, greedy action selection can be performed directly from the SNN outputs, allowing the auxiliary ANN to be completely removed and preserving the energy efficiency of SNNs. Extensive experiments on Atari and MiniAtar environments demonstrate substantial performance improvements under low-timestep settings. CMC-DSQN outperforms state-of-the-art DSQN baselines by nearly $20\%$ at $T=2$ and further surpasses the ANN baseline at $T=4$.
cs.NE / 5 / 2610.07825
Forecast Accuracy Is Not Trading Profit: Evolving Small Recurrent Networks for Stock Return Prediction
Abstract
Time series forecasting models are typically compared on pointwise error, which scores a prediction in isolation from the decision it is produced for, and a lower forecast error does not imply a better decision downstream. A parallel debate asks whether modern transformer architectures forecast better than recurrent and other lightweight models. We compare linear, fixed recurrent, transformer, and mixing based architectures against recurrent networks evolved by neuroevolutionary architecture search, evaluating each on forecast accuracy and on the net return of a daily long/short strategy. All models are fit on a pooled panel, one network trained across the whole universe. Across four mid-cap portfolios and three trading years, the evolved networks rank first on both forecast accuracy and net trading performance, while the second most accurate model loses money once positions are formed and costs are charged. The advantage tracks a horizon match, since rank IC for the evolved networks rises from a one-day to a ten-day scoring horizon while every model above 300 parameters declines. They are also the cheapest end to end: a CPU-only search of 16 minutes yields 66-weight networks that predict in 10.8~$μ$s on a Raspberry Pi Zero, against transformer baselines of up to 817,153 parameters that require GPU training.
cs.NE / 6 / 2610.07962
ReGraph: A Computational Account of Emergent Generalization in the "what" and "where" Dual Visual Streams
Abstract
Where generalization capacity--the ability to extract context-invariant relational structures--first emerges remains a central question in AI and neuroscience. The foundation for this capacity lies upstream of the hippocampus, within the entorhinal cortex, where parallel pathways dissociate relational structure in the medial entorhinal cortex (MEC) from sensory content in the lateral entorhinal cortex. However, as Eichenbaum argued, such factorization likely originates earlier, driven by the segregation of the dorsal ('where') and ventral ('what') visual streams. Supporting this, grid-like firing patterns--a signature of MEC (context-invariant codes)--also appear in preceding neocortical regions along the dorsal pathway. Yet, how such representations are computationally formed along upstream pathways remains unknown. To investigate this in silico, we developed ReGraph, a recurrent dual-stream graph model with biological inductive biases, including retina-driven stream-specialized encoding, dorsal-to-ventral modulation, and dynamic lateral connectivity. Trained on the action benchmark Something-Something V2, ReGraph revealed a pathway-specific emergence of relational mapping: context-invariant codes and grid-like spatial bases uniquely co-emerged along the extended dorsal stream. In contrast, their absence in single-stream, unmodulated variants, and standard baselines implies that these inductive biases are prerequisites for relational structures. Crucially, our post-hoc analyses demonstrated that these grid-like bases serve as reusable routing templates for information processing via lateral connectivity. Together, our findings provide a computational account that generalization may not be a faculty that emerges abruptly within a dedicated region, but a property that already takes shape as sensory information is parsed into factorized streams of hierarchical visual processing.
计算语言学 (cs.CL)
49
cs.CL / 1 / 2610.07365
Who Wrote It Is Not Enough: Detecting Who Contributed the Insight
Abstract
As LLMs increasingly assist scientific writing and peer review, detecting who wrote the text is no longer sufficient: we need to determine who contributed the underlying insight. We introduce Insight Provenance, the task of identifying whether a review insight originates from a human, an LLM, or their hybrid contribution. We construct InsightProv-v0 from 4,057 scientific papers and 12,660 human reviews, simulating different levels of LLM involvement with GPT-4o, Gemini, and DeepSeek and annotating provenance at the sentence level. We show that strong performance on raw data can be misleading, as models exploit linguistic and textual-authorship shortcuts that degrade substantially under progressively debiased evaluation. We therefore propose a two-stage adversarial framework that suppresses shortcut signals while preserving provenance-relevant information. Beyond detection, extensive analyses reveal what makes intellectual authorship identifiable: paper grounding and neighboring review context provide complementary provenance signals, while human, hybrid, and AI insights systematically differ in their information sources and failure modes. Most strikingly, AI insights predominantly remain close to generic or paper-provided information, whereas human insights more often introduce external knowledge and independent judgment. These findings suggest that while wording can be rewritten by an LLM, the provenance of an idea leaves a deeper and more persistent signal.
cs.CL / 2 / 2610.07426
AccentCL: Robust Accent Classification with Incremental Expansion
Abstract
Accent classifiers are typically trained with a fixed label inventory and cannot accommodate new accent categories as new data becomes available. Moreover, accented speech corpora often exhibit substantial class imbalance and/or domain shift due to differences in recording conditions across corpora. We present AccentCL, a class-incremental learning framework for English accent classification that is robust to class imbalance and cross-corpus domain shift. AccentCL extracts multi-layer representations from a frozen Whisper-Large-v3 encoder, optimized with an imbalance-aware cross-entropy loss to reduce bias toward the majority accent classes and a domain mean alignment loss that minimizes distributional mean shift across training corpora. The label space is then expanded via replay-based continual learning, using the frozen base model for knowledge retention and an old-to-new margin loss to reduce overprediction on newly added classes. On a five-class accent classification task, AccentCL achieves 77.1% balanced accuracy and a 76.9% macro-averaged F1 score. We further evaluate the model's ability to incrementally incorporate two new accent categories: Spanish-accented and Chinese-accented English. When adding Spanish-accented English to the pretrained model, AccentCL attains an F1 of 83.3% on the new class while retaining 77.3% balanced accuracy on the base classes. When subsequently adding Chinese-accented English, it achieves 61.8% F1 on the new class while preserving 77.6% balanced accuracy on the previously learned classes. These results show that AccentCL enables robust regional accent classification while allowing new accent categories to be added without full retraining.
cs.CL / 3 / 2610.07502
Closing Ambient Clinical Documentation Gaps with Automated Provider Queries
Abstract
Provider queries are clarifying requests sent by clinical documentation specialists to physicians to close gaps in the clinical note and ensure accurate billing. Prior work automates note drafting, ICD-10 coding, and order extraction assuming a complete transcript, leaving these gaps unaddressed. We study whether an LLM can automate the query loop, termed DAU (Draft, Ask, Update), across those three tasks. An audit of 3,000 real visits identifies the sources of missing documentation, from which we build five transcript-degradation benchmarks on public data. Analyzing 21k clarification turns on real conversations, we find useful-question predictors are task-specific: oracle confidence dominates, but note completeness needs only simple recall questions while ICD-10 coding needs harder, multi-option ones. About 9% of turns hurt performance, driven by redundant questions and non-answers that still trigger a rewrite. Deployment depends on learning "when not" as much as "what to" ask.
cs.CL / 4 / 2610.07519
Not What a Child Expressed: Auditing the Sign-to-Text Safety Interface in Child-Facing AI
Abstract
Automatic sign language translation (SLT) has entered consumer products, turning American Sign Language into English text for dictation, messaging, and queries put to a conversational assistant. Child-facing AI and platform trust-and-safety tooling decide on text, using filters on minor accounts and grooming classifiers that score chat messages. A signing child who uses SLT therefore reaches these safeguards through a translation. We found no publicly documented system in which the two have been jointly evaluated, and the leading deployed SLT model was neither trained nor formally evaluated on signers under 18. Errors that alter negation, participant roles, secrecy, urgency or help-seeking could change a safety decision without disturbing fluency. This paper proposes a Deaf-informed pre-deployment audit of that boundary, with a failure taxonomy, a sanitised scenario schema, four comparison conditions, and four outcome measures. Auslan is the planned first case study.
cs.CL / 5 / 2610.07545
Quality-Aware Self-Correcting Speech Translation on an Edge Device
Abstract
We present a fully offline speech-to-speech translation pipeline that runs on a Jetson Nano (4 GB) and corrects its own weak translations without retraining. A Whisper-tiny ASR feeds an Opus-MT translator; multilingual BERT cosine similarity acts as a Quality Estimation (QE) gate, triggering a secondary-pass correction when confidence falls below a pre-defined threshold $τ$. We compare three correction methods: QE reranking (M1), Minimum Bayes-Risk decoding (M2), and constrained beam search (M3). On 1,012 FLORES-200 sentences (English-Spanish), M2 at $τ=0.90$ produces statistically significant improvements over greedy decoding on BLEU (+0.67, p<0.001), ChrF (+0.51, p<0.001), and COMET (+0.0020 at N=3, p=0.002); M1 yields no significant gains, and M3 is significantly worse than baseline (p>0.99). Our central finding is that QE functions effectively as a gate but poorly as a ranker: removing the QE model from candidate selection (M1$\to$M2) does not hurt quality and frees 680 MB from the critical path. Using a gain-to-edit ratio adapted from the post-editing-effort literature, we further show that smaller candidate pools (N=3) yield more surgical corrections with better semantic adequacy, while larger pools (N=10) maximise lexical reward. We release the system and demonstrate live translation across six language pairs.
cs.CL / 6 / 2610.07591
Recurrent Looped Transformer
Abstract
State tracking requires an update at every input, but the depth a Transformer applies to each token is fixed regardless of sequence length. We introduce the Recurrent Looped Transformer (RLT), which splits its layers between a parallel causal encoder and a recurrent decoder. At each token, the decoder merges the encoder output with the previous token's final decoder state, so the computation path grows with sequence length at a fixed per-token cost. On six algorithmic tasks, we compare five splits of eight layers with an eight-layer Transformer over three seeds. Trained on at most 40 bits, two RLT splits generalize parity to 256 bits with 100% accuracy in every seed, while the Transformer stays at chance. On swap-based $S_5$ permutation tracking at eight times the training length, RLT reaches 97% final-state accuracy versus under 1% for the Transformer, and accuracy increases with decoder depth. On modular arithmetic beyond the training lengths, RLT reaches up to 93% versus 33% for the Transformer. Ablations show that these gains depend on the feedback: removing it drops parity and swap-based $S_5$ to chance at every split. Updating the feedback once per four-token chunk lets known tokens in a chunk run in parallel and keeps 64-bit parity at 99%, while permutation tracking depends on per-token feedback: chunking lowers length-64 swap-based $S_5$ from 100% to 20%.
cs.CL / 7 / 2610.07643
Monte Carlo Estimation for KV Cache Eviction
Abstract
Most KV-cache eviction methods ask, in effect, which memory appeared important while reading the prompt? We instead ask, which memory will matter while answering? Since decoding queries are unavailable at eviction time, prior future-aware methods rely on pseudo-responses or synthetic future-query estimates. We cast fixed-budget future-aware eviction as distributional estimation over plausible model-conditional query trajectories and introduce LORE-KV (Lookahead Output-perturbation with Reliability-weighted Ensembles for Key-Value caches), a training-free method that samples short autoregressive continuations from the frozen target model and uses their response-side query states to estimate prompt-token utility. Tokens are scored by projected leave-one-out attention-output deletion cost and aggregated across sampled futures with optional trajectory weighting. The temporary continuations are discarded before final decoding, requiring no auxiliary model or training. Ablations isolate the mechanism: at B=128, a single response-side continuation recovers about 89% of the gain over the prompt-window control, while additional futures provide smaller improvements. At B=128, LORE-KV raises the LongBench average on Qwen2.5-14B from 45.49 to 48.24 (+2.75) and the 16K RULER average on Mistral-7B from 45.20 to 51.05 (+5.85). Gains diminish at larger cache budgets and coexist with task-level regressions. LORE-KV incurs 1.46-2.77x AnDPro's per-sample wall-clock time as a one-time compression overhead across six dense and hybrid-attention backbones.
cs.CL / 8 / 2610.07716
Readout Stability in Prefill-Only Decision Models:Zero-Label Prediction and Inference-Time Compute Allocation
Abstract
Prefill-only decision models inspired by the Jev model score every candidate in a menu during a single forward pass and never decode, which makes one call one to two orders of magnitude cheaper than a same-scale generative language model. We show that this read-out structure comes with a testable property. When an intervention changes only the candidate menu and leaves the input text fixed, the post-intervention accuracy is already determined by the cached first-pass distribution. The estimator restricts the pass-1 probabilities to the menu, renormalizes, and reads off the argmax; it uses no labels and no second forward pass. Across seven model families, ten datasets and two task types, menu-only interventions are predicted to within 4.2 points, and for one family the prediction is exact. A probability-level variant of the same estimator errs by 21.0 points, so the property lives in the ranking rather than in the probabilities and is not recovered by calibration. Same-scale generative language models do not share the property. On those models the same estimator errs by 1.6 to 15.8 points and degrades as the model grows. The property turns inference-time compute into a decision that can be made before deployment. Uniform extra passes buy calibration but almost no accuracy; at matched cost a confidence cascade outperforms every scheme that re-asks the same model, and curating the menu beats enlarging the model, with a 0.8B model on a curated 5-candidate menu reaching 95.4% on CLINC150 against 80.0% for a 4B model on the full 150-label menu.Code and data are available at https://github.com/rlisml/jev-cascade.
cs.CL / 9 / 2610.07730
SanSi: A Looped Typed Decision Model for System 1.5 Thinking
Abstract
Typed decision models answer a declared question without generating text: a decision head returns a probability for each of the declared options in a single forward pass. A single pass is fast, intuitive System 1 thinking. We study what lies between one pass and generated reasoning: looping, in which the same layers are recursively applied several times before one typed readout. Each loop lets the model revise its hidden state before it commits to an answer, without generating a token; we call this System 1.5 thinking. We propose SanSi, which turns a pre-trained looped language model into a typed decision model. The option probabilities are read after every loop, and every loop is trained with a proper scoring rule, so that one model serves every budget from one loop to eight in a single run. On 10,027 test decisions from 59 sources, SanSi reaches 72.0% accuracy: 13.5 points above a non-looped model of the same shape trained with the same recipe, 5.3 points above a newer non-looped model of its size, and 1.8 points below one with three times the parameters. On two depth-controlled tasks, loops extend the solvable depth beyond the depths seen in training, where the larger single-pass model fails. Used as the judge for policy optimization with reinforcement learning, without gold answers, SanSi raises the generator's F1 by 7.7 points.
cs.CL / 10 / 2610.07753
From Evidence to Action: How Tool-Using Agents Fail
Abstract
Tool-using agents make consequential changes to external state, yet correct outcomes do not guarantee that their actions were supported by evidence established beforehand. We study where this evidence-to-action chain breaks as agents move from deciding whether to act to executing single actions and dependent workflows. Across ten model-harness configurations, strong static action assessment can coexist with much weaker interactive execution. Failures often begin before execution: agents stop with incomplete investigation or act before required evidence is established. Once required evidence is obtained, single-action execution is usually reliable, while multi-action workflows additionally expose unresolved prerequisites and incomplete execution. For this analysis, we introduce SafeActBench, comprising 656 cases across six operational domains and five protocols that progress from static action judgment and investigated non-action to single- and multi-action workflows. A provenance-bound Evidence Ledger and deterministic trajectory evaluator track what information was established, when actions occurred, and whether downstream dependencies were satisfied. These results show that failures arise not only from missing information, but also from how agents use established evidence when deciding and executing actions.
cs.CL / 11 / 2610.07774
Reading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language Models
Abstract
Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs. As mixture-of-experts (MoE) VLMs become increasingly common, recent work has explored various safety interventions, including prompting, supervised fine-tuning, and routing-based expert steering. However, these methods show inconsistent improvements across models and evaluation distributions, and the intervention into model behavior or internal states introduce safety-utility tradeoffs by over-refusal. Rather than manipulating internal states to steer model behavior, we instead ask whether routing states can serve as diagnostic signals for multimodal safety. We find that router logits indeed provide highly predictive signals of whether a multimodal input is safe or not. Motivated by this observation, we introduce a lightweight router-logit safety detector that reads out routing signals during prompt prefill and identifies unsafe requests before generation, without modifying model parameters or expert routing. Across Qwen3-VL and Kimi-VL, the proposed detector substantially reduces safety errors on the HoliSafe benchmark and resoundingly generalizes to out-of-distribution safety benchmarks featuring different safety patterns, including MISHard and MM-SafetyBench. The success of the proposed router-logit detector also suggests a broader perspective on model internals: rather than focusing only on manipulating internal components to steer behavior, simply reading naturally emerging signals and linking them to an external safety mechanism can provide a simple, effective, and non-intrusive complement to existing safety interventions.
cs.CL / 12 / 2610.07817
One Step at a Time: Trading LLM Autonomy for Process Predictability
Abstract
Organizations automating operational processes need more than a correct outcome: they need to predict how a process will run, know which one actually ran, and inspect it step by step. When an agent is the executor that predictability is normally lost: the prescribed procedure goes into the system prompt, and only a final answer comes back. We deliver the procedure step by step over the Model Context Protocol (MCP) instead: a server releases one step at a time, the agent executes it, and each step returns a structured step_output. This trades autonomy for predictability, and two properties then follow by construction, independent of the executor. The execution path is prescribed before the run, so the process is predictable in advance rather than reconstructed afterwards; and the completed step records form a machine-readable execution log that downstream tooling can audit and optimize step by step. Evaluating 15,475 trials across 13 SOP-Bench domains and four open-weight executors from frontier (Kimi K2.5) to lightweight (Ministral 3 8B), we find step-level delivery makes the executed process predictable and inspectable for every executor, and additionally raises accuracy when the executor is small. Across all four, process adherence rises significantly (76-95% to 95-99%) and ungrounded answers (correct outputs produced without executing the SOP) near-vanish, falling from 2.1-4.5% to 0.2-0.3% of trials (all 95% CIs exclude zero); under prompt-based delivery, 31-49% of correct answers on know_your_business bypass the SOP entirely, even for the frontier executor. Accuracy is where the executor's capability enters: the lightweight executor gains +6.5pp grounded accuracy because supplying the process externally removes a reconstruction burden it cannot carry, while capable ones trade a small raw-accuracy decrement for a predictable, auditable process.
cs.CL / 13 / 2610.07822
Nucleus Speculative Decoding: Plausibility-Aware Verification Beyond Exact Distribution
Abstract
Speculative decoding accelerates autoregressive generation by using a lightweight draft model to propose multiple tokens that are verified by a target model in parallel. However, the standard acceptance rule focuses on exact distribution correction and rejects tokens that remain highly plausible under the target model when the draft model assigns excess probability. This conservative verification limits the number of draft tokens retained after each verification forward pass. We introduce Nucleus Speculative Decoding (NSD), a relaxed verification method that incorporates target-model plausibility into speculative decoding. NSD accepts a draft token if it satisfies the standard acceptance rule or belongs to the target model's nucleus. We theoretically characterize the distributional deviation introduced by our method and show that the single-step error is exactly determined by the draft model's excess probability within the target nucleus. We further derive sequence-level fidelity bounds that quantify how local deviations accumulate over autoregressive decoding. Experiments across multiple target models and proposal mechanisms demonstrate that NSD consistently improves speculative decoding efficiency while maintaining competitive task performance. Our method achieves throughput speedups of up to $5.16\times$ over autoregressive decoding and up to $3.15\times$ over standard speculative decoding. These improvements coincide with longer accepted lengths, allowing more output tokens to share the cost of each target verification pass. Analysis shows that plausibility-aware verification provides an effective approach for relaxed verification and speculative decoding efficiency. Our code is available at https://github.com/EIT-NLP/Nucleus-Speculative-Decoding.
cs.CL / 14 / 2610.07847
OMIT the Action: Measuring Framing-Invariant Omission Bias under Philosophical Disagreement
Abstract
As LLMs increasingly assist in moral reasoning, omission bias, the tendency to prefer inaction even when equivalent framings reverse substantive outcomes, poses a significant risk of skewed decision-making. Yet omission bias remains underexplored in LLM evaluation, with the few existing studies limited in scale and focused largely on utilitarian-deontological conflicts. To address this gap, we introduce OMIT, a benchmark consisting of 218 paired-frame scenarios across 10 conflict types, constructed by leveraging disagreement patterns from an LLM-based, five-perspective philosophical persona panel (utilitarianism, deontology, virtue ethics, care ethics, and contractualism). Evaluating eight LLMs, we find that omission bias is pervasive but inversely correlates with model size within families. We further evaluate four inference-time interventions and find that interventions encouraging models to consider moral principles before committing to a yes/no answer reduce omission bias and increase frame-consistent responses, although lower omission bias rates can also coincide with shifts toward action-biased responses. Ultimately, this work contributes not only the OMIT benchmark, but also a methodology for using diverse philosophical disagreement signals to evaluate framing-sensitive inaction preferences and the distributional effects of mitigation attempts in LLMs under complex moral conflicts.
cs.CL / 15 / 2610.07863
ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents
Abstract
Long-horizon LLM agents act on an append-only interaction history that is re-sent to the model at every step, so the context and its cost grow with steps until the sessions exceed the context window. Existing methods manage the context through context requirement prediction, relying on additional model calls, heuristic rules, or trained policies. However, these predictive approaches introduce runtime overhead, invalidate prefix caches, and permanently discard content with no guarantee of recovery. To overcome these limitations, we introduce ReFold: a training-free rendering layer that preserves the underlying interaction history while compressing only the model's rendered context. It removes two kinds of inter-turn redundancy without an auxiliary predictor: content an earlier turn already displayed, replaced by a stub, and turns the agent itself reports finished, folded into a one-line note. Both operators use chunked rendering, rewriting the cached prefix once every few steps rather than at every step. Every removal is strictly reversible, a wrong removal costs one restore from the history rather than permanent content loss. Because it operates at the rendering layer, ReFold is plug-and-play across standard ReAct-style harnesses. Evaluations across five long-horizon benchmarks and two frontier LLMs demonstrate that ReFold reduces token consumption by up to 2.5x and halves the KV-cache memory per session without degrading task success rates. Under capped context budgets, it avoids up to 92% of forced compactions. Under concurrent serving workloads, it reduces request queuing delays by up to 100%, accelerating inference by up to 1.7x, while cutting inference costs by up to 3.4x.
cs.CL / 16 / 2610.07887
Visual Abstention in Unified Multimodal Models
Abstract
Unified multimodal models (UMMs) integrate understanding and generation, yet their generative behavior is rarely governed by what they understand about the task. We formalize visual abstention: when a requested visual transformation is impossible under the task's rules, the model should recognize that no valid solution exists, state this, and decline to generate. We introduce Draw-or-Decline (DoD), a benchmark of 1,050 feasible-infeasible request pairs across 7 task categories that jointly measures editing success and the refusal of infeasible requests. Evaluating 8 UMMs, we find that editing ability and abstention are distinct capabilities: even the strongest editor, at 68.4% editing accuracy, refuses only 0.4% of infeasible requests under ordinary instructions. Their reasoning shows why: the models rarely notice the conflict, and instead plan the edit as if the request were possible, often describing objects that are not in the image, or quietly change the request into one they can complete. Explicitly prompting these UMMs to report infeasibility increases textual refusals but reduces editing accuracy. We propose VisTA (Visual Transformation and Abstention), a training method that pairs feasible and infeasible examples so that a model judges feasibility before deciding whether to generate. We train VisTA-BAGEL to perform feasible edits and decline infeasible requests. Without any reminder, it refuses 93.0% of infeasible requests, up from 0.4% for the strongest editor, while falsely refusing only 0.8% of feasible ones. Unlike a reminder, this does not cost editing accuracy: VisTA-BAGEL completes 74.3% of feasible edits, more than any of the 8 evaluated UMMs.
cs.CL / 17 / 2610.07902
ARIA: Audio-Driven Melody-Tone Relation Modeling for Cantonese Lyric Authoring
Abstract
Cantonese lyric writing requires close alignment between lexical tones and melodic pitch. Existing melody-guided lyric generation methods typically rely on symbolic melody to generate lyrics. However, in real songwriting scenarios, melodies are often expressed as raw singing audio or hummed recordings, where pitch is implicit, noisy, and unstructured, making these methods difficult to apply directly. To address this limitation, we propose ARIA, a two-stage audio-driven melody-tone relation modeling framework for Cantonese lyric authoring that generates Cantonese lyrics from singing recordings with provided character-level timestamps. Specifically, we first design a Tri-Stream Relation-Aware Tone Estimator (TRATE) to predict 0243 sequences from timestamped singing audio by modeling multi-stream acoustic cues and relational tonal structure. We then propose a Decoupled Retrieval-Augmented Tone-Conditioned Lyric Generator (DRA-TCLG) to generate fluent lyrics conditioned on predicted tonal plans with retrieval-enhanced lexical guidance. Moreover, we construct a large-scale aligned audio-Jyutping-0243 dataset from real Cantonese singing recordings to support this new task. Experimental results demonstrate that ARIA achieves strong performance in both 0243 prediction and tone-consistent lyric generation, validating the effectiveness of the proposed framework.
cs.CL / 18 / 2610.07937
Leveraging a four-quadrant approach for evaluating Redpine Science
Abstract
Redpine Science gives models and agents a single access point to a wide range of peer-reviewed literature, queried directly through the Model Context Protocol (MCP) and an API. This report evaluates Redpine Science on two levels: the relevance of the retrieved chunks, and a model's answer when it has access to Redpine Science compared to web search. Both public and expert-validated benchmarks are used. Public benchmarks are a widely accepted way to test model development and are comparable across labs, but risk saturation and memorization. To address this, we complement them with an expert-validated question set. In total, this report presents four evaluations. On ScholarQABench SciFact, the public answer-quality benchmark reported here, an agent with Redpine Science answers 94.4% of claims correctly against 87.6% with no retrieval. On the expert-validated question set, an agent with Redpine Science states 80.1% of the required claims against 70.2% for an agent restricted to web search. On the 668 queries of a public retrieval benchmark whose gold paper Redpine holds, stripped of any model reasoning, Redpine Science places the correct source paper in its top ten results for 83.1% of queries (Recall@10), against 79.3% for the benchmark's creator. A blinded expert relevance panel places Redpine Science's Precision@5 at 75.2% against 39.8% for the PubMed search tool. We release the expert-validated question set and instructions to reproduce every headline result above, at https://github.com/redpine-ai/benchmarks.
cs.CL / 19 / 2610.07940
Hybrid Latent Attention for Looped Language Models
Abstract
Looped language models apply the same stack of layers T times to each token, which deepens the model without adding parameters but multiplies its key-value (KV) cache by T. The larger cache limits how many sequences a GPU can decode at once and slows each decoding step, which reads the whole cache. We propose Hybrid Latent Attention (HLA), which keeps exact keys and values within a sliding window of W recent tokens and stores each older token as a compact latent that the query of each loop reads directly, without reconstructing keys and values. We uptrain HLA on Ouro looped models (T=4) with 1.4B and 2.6B parameters, keeping the pretrained weights frozen and training only the added parameters to reproduce the original attention. The cache shrinks by 10.7x per token, fitting 4.0-8.8x as many concurrent sequences per GPU, and decoding throughput improves by 2.5x at 1K-token contexts and by up to 7.4x at 16K. HLA retains over 97% of the original accuracy on math, knowledge and reasoning benchmarks, and 96-100% on long-context retrieval up to 16K tokens. After supervised fine-tuning, it performs on par with the fine-tuned original model on competition-level math.
cs.CL / 20 / 2610.08018
Structured but Silent: Probing Capability Requirements in LLM Hidden States
Abstract
Reliable tool use requires more than triggering a mechanism or matching a query to an API description. Before selecting a specific tool, an agent must first infer the capability requirements implied by the user query. In this paper, we investigate whether these query-side capability requirements are linearly decodable from LLM hidden representations prior to generation, and how this hidden-state accessibility compares with explicit verbal classification. We introduce TACIT, a framework that decomposes external requirements along three fundamental axes: Source, Transformation, and World Effect, defining eight structurally distinct capability classes. Using 1,600 balanced training queries from benchmarks, synthetic examples, and new domain scenarios, we train linear probes on pre-generation hidden states from four open-weight LLM families. Our empirical results demonstrate that fine-grained capability structures are linearly decodable with high accuracy across all models. Crucially, however, we expose a representation-to-verbalization gap: these same models are significantly less reliable when asked to explicitly classify the same queries in natural language. This disconnect indicates that information about required external capabilities is linearly accessible in LLM hidden representations but not reliably expressed, a phenomenon we define as "structured but silent."
cs.CL / 21 / 2610.08055
Language Carries the Expert's Impression: Instrument-Anchored LLM Judges Transfer Counseling-Quality Assessment and Beat In-Domain Training
Abstract
Automatic assessment of communication quality in dyadic counseling conversations is bottlenecked by data: expert-rated corpora are small and expensive to grow. We study cross-domain transfer of expert overall-impression prediction across three German corpora of simulated counseling (two general-practice medical, one school-related parent-teacher; $n=195$ expert-rated sessions, one corpus after scale equating). Training on the other domains beats training in-domain: leave-one-domain-out transfer reaches nested Spearman $ρ= 0.54$ against $\le 0.48$ within the target domain, a paired session-level gap of $+0.15$ that holds at $+0.12$ when the training-set sizes are matched, so it is not simply data volume. The decisive features are session-level construct scores from small open-weight LLMs reading the two-speaker transcript, with the constructs largely derived from the experts' rating instruments: the instrument-derived battery lifts a single judge from $0.32$ to $0.41$ over generic dialogue qualities, judges from three model families ensemble to $0.51$ language-only, and a nonverbal-dyadic block adds $+0.03$ more, not separable from noise at this sample size. We also price the recording setup: one corpus lost its per-speaker audio, 16% of its diarised segments carry the wrong speaker, and repair is worth $+0.07$ there. At practically attainable corpus sizes, the expert's overall impression is carried by what is said, and by other communication programs' data more than by one's own.
cs.CL / 22 / 2610.08085
DirectSpeech2LLM: A Simple End-to-End Framework to Mitigate Prompt Overfitting in Speech-LLMs
Abstract
Speech-LLMs often exhibit prompt overfitting, where models solely trained on automatic speech recognition (ASR) instruction fail to generalize to new instructions such as speech translation and continue to behave primarily as ASR system. We propose DirectSpeech2LLM, a simple end-to-end framework that preserves the instruction-following ability of the LLM on unseen tasks when conditioned on speech. It computes distance-based CTC loss over the frozen LLM embedding matrix and uses greedy CTC labels to derive geometrically and temporally aligned speech embeddings respectively as an input to the LLM. Trained solely on 960 hours of LibriSpeech ASR data, DirectSpeech2LLM outperforms the cascaded system on ASR (seen task) and generalizes zero-shot to speech translation and emotion recognition (two unseen tasks), closely matching the cascaded system upper bound on these two new instructions despite seeing neither during training. We also find that geometric alignment strength plays a smaller role than previously assumed, as our modified CTC loss is shown to provide sufficient implicit geometric grounding without requiring an explicit regression loss. Results are consistent across two LLM families and scale with both more training data and model capacity.
cs.CL / 23 / 2610.08093
SAGE: Semantic Anchor-Guided Evolution for Grounded Medical QA Data Synthesis
Abstract
Developing reliable models for clinical tasks, such as Medical Question Answering (QA), is severely constrained by the limited availability of high-quality, expert-annotated training data. This challenge is exacerbated by stringent privacy requirements and the impracticality of utilizing large open-source corpora or proprietary cloud APIs within resource-limited clinical settings. To address these obstacles, we introduce SAGE (\textit{Semantic Anchor-Guided Evolution}), a novel data synthesis framework that enables small, locally deployed models to generate high-quality medical training data. SAGE leverages lightweight, publicly available taxonomies such as MeSH as semantic anchors, imposing a structured prior to effectively guide and ground the data generation process. At its core, SAGE iteratively interleaves atomic (individual concept-based) and associative (relation-based) synthesis, bootstrapping training data from minimal seeds. This approach eliminates the need for large collections of medical documents or reliance on external APIs, providing a practical solution for on-premises data creation. Extensive experiments across multiple medical question-answering benchmarks demonstrate that models fine-tuned with SAGE-synthesized data consistently outperform those trained using self-derived or conventional document-based paradigms, highlighting tangible improvements in data efficiency and resource utilization for medical LLM development. Code is available at https://github.com/DIaacKr/SAGE.
cs.CL / 24 / 2610.08159
Making COMET Comparable Across Scripts: Diagnosis and Correction of Tokeniser-Induced Script Bias in Indic MT Evaluation
Abstract
COMET reports translation quality as a single number, and that number is routinely compared across target languages written in different scripts. Such a comparison assumes Script Invariance: the score should not depend on the writing system that carries the target. We test it on IndicMT Eval by re-encoding the target into Latin script, which changes orthographic form while holding content and human ratings fixed. Script identity then accounts for 22.9% of native-script COMET variance, and agreement with annotators falls in all five languages studied. We trace the effect to the tokeniser and measure it with three label-free diagnostics. The bias is two faults, not one. Scores from different scripts occupy incompatible ranges, and within a single script the metric orders translations less accurately. No order-preserving transform of the score can repair the second fault. The first is removed exactly by COMET-QN, which maps the score distribution of each (language, script) pair onto a shared reference. Pooled agreement with annotators rises from 0.300 to 0.399, which is what makes scores from different scripts safe to place on one axis, and every within-language ordering is provably preserved. A regressor over parity features recovers a further 17.1% of the lost sensitivity. The remainder belongs to the encoder, and no post-processing can reach it. We therefore recommend publishing the normalised score, the three diagnostics, and the identity of the tokeniser they were computed against, so that a reader can tell how much of a score reflects translation quality and how much reflects the writing system.
cs.CL / 25 / 2610.08208
STRUCTURALCOST: A controlled reading time dataset for modeling human sentence processing difficulty
Abstract
We introduce STRUCTURALCOST, a self-paced reading dataset of 475 participants and 40,800 observations isolating the processing cost of long-distance subject-verb dependency resolution. We replicate a low-powered psycholinguistic finding at NLP scale, namely that human reading times at the main verb increase with dependency length, driven by syntactic embedding beyond linear distance. Different language models -- spanning n-gram models, SSMs, and transformers -- partially mirror this graded difficulty profile, yet underestimate the integration cost humans incur, with a gap that persists across architectures and model sizes. This suggests these models capture the predictive component of human processing but not the full integration cost that working memory imposes. STRUCTURALCOST provides data needed to drive progress toward evaluating the cognitive plausibility of language models.
cs.CL / 26 / 2610.08300
Memory Depth and Reconstructed Context Width: A Controlled Evaluation of Hierarchical Retrieval
Abstract
Long-term conversational memory is becoming an integral component of modern LLM systems. Proposed architectures group records by topics and events, construct hierarchies and graphs, and connect facts through causal and temporal relations. We experimentally study the interaction between two memory parameters: structural depth and the width of context supplied to the answer model. Using EverMemBench, we evaluate depths D1-D4, core budgets of 1,024/2,048/4,096 tokens, and additional Production and Oracle conditions up to the full archive. Increasing width from 1K to 4K improves Accuracy by 10.11-17.98 percentage points, whereas increasing depth provides no monotonic gain. Beyond 8-16K, Production performance reaches a plateau while tokens per correct answer continue to increase; Oracle preserves quality on full archives of 68-71K tokens. These results motivate further investigation of large, coherent context blocks instead of progressively deeper memory structures.
cs.CL / 27 / 2610.08448
Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability
Abstract
On-Policy Distillation (OPD) trains a student on its own generations using teacher feedback. With different tokenizers, comparing teacher and student predictions requires alignment at both sequence and vocabulary levels. In this paper, we examine whether expanding this alignment coverage improves learning. Across three heterogeneous teacher--student pairs on mathematical reasoning and code generation, strict 1:1 groups already cover most student-generated tokens despite substantial vocabulary mismatch. On responses sampled from the students before distillation, the shared vocabulary retains nearly all teacher and student probability mass at strictly aligned positions on average. Restricting reverse KL to a student-selected top-16 subset of the shared vocabulary at each strict position achieves accuracy comparable to full shared-vocabulary OPD, outperforming the evaluated cross-tokenizer baselines. Adding mean squared error supervision on span log-probabilities in mismatch groups gives complete supervision coverage, yet reduces accuracy. At checkpoints from training with only the strict loss, the span gradients show weak or negative directional agreement with the strict gradients and grow in magnitude relative to them. These diagnostics may help explain the accuracy drop from adding span supervision. Our findings motivate a shift from maximizing alignment coverage to prioritizing supervision reliability: compact supervision at strict positions can be more effective than broader coverage that introduces weakly aligned or conflicting training signals.
cs.CL / 28 / 2610.08463
UNREAL: Unifying Retrieval and Long-Context with a Single Model
Abstract
Long-context inference and Retrieval-Augmented Generation (RAG) handle evidence selection at vastly different scales, from a single long prompt to an entire corpus. We ask whether a single model-internal mechanism can select evidence across this range. We introduce UNifying REtrieval And Long-Context with a Single Model (UNREAL), a model-native evidence selection framework to span corpus retrieval and long-context inference. UNREAL encodes chunks and derives retrieval queries directly from the frozen LLM's internal representations. It adds fewer than 500K trainable parameters and leaves the backbone unchanged. On a 3B-token, 21M-chunk Wikipedia index, all four dense and hybrid UNREAL backbones outperform state-of-the-art retriever-reranker systems. The best model raises recall from 49.1% to 73.2% on HotpotQA, from 31.7% to 60.1% on 2WikiMultiHopQA, and from 8.8% to 14.4% on MuSiQue. Applied to long-context tasks, the same selection mechanism removes distractors before generation, raising NoLiMa accuracy from 1.0% to 24.83% at its maximum context length of 128K tokens, and LV-Eval's F1 score from 49.97% to 54.66% at 256K. UNREAL also reduces FLOPs and time-to-first-token relative to full-context inference from roughly 32K tokens onward, with larger gains as context grows. Together, these results establish model-internal evidence selection as a common foundation for corpus retrieval and evidence-sparse long-context inference.
cs.CL / 29 / 2610.08501
Language-model ratings of depression reflect the rater more than the patient
Abstract
Depression has no diagnostic blood test. Language models promise tireless, consistent assessment, but can accurate raters disagree about individuals? We pre-registered 880 language-model raters, crossing 11 open models with prompting and scoring choices, and applied them to 189 interviews against the eight-item Patient Health Questionnaire. Model choice explained 30.0% of summed-symptom score variance, stable participant differences 10.5%. Two randomly drawn raters with area under the receiver operating characteristic curve (AUC) >= 0.70 disagreed on screening decisions for 40% of participants, on average. Average over-rating governed how many were flagged, yet equal-capacity raters chose differently for about one participant in five. A locked analysis of 86 new interviews reproduced the main pre-registered findings. Exploratory recalibration with 40 labelled participants raised accuracy from about 60% to 75% and halved disagreement, leaving one participant in five decided differently. Calibration repaired much of the rater dependence without securing agreement about individuals.
cs.CL / 30 / 2610.08513
Wiki-Talkie: Multilingual Benchmarking of Persona-Based Agents on Real-World Discussions
Abstract
LLMs are increasingly deployed as autonomous agents in social environments, making it critical to study their ability to faithfully simulate human interactions. Central to this is grounding agents in realistic user personas, yet existing datasets rely on fictional personas and are limited to a handful of languages, lacking the empirical grounding necessary to evaluate behavioral fidelity across diverse populations. We introduce Wiki-Talkie, a multilingual dataset of real-world conversations from Wikipedia Talk pages across five languages spanning two language families: Germanic (German, English) and Romance (Spanish, French, Italian), paired with personas derived from real user communities and encompassing sociodemographic attributes, self-descriptions, and behaviorally grounded interaction traits. Using Wiki-Talkie, we evaluate agent interactional behavior on a next-turn generation task across various persona conditioning strategies. Our evaluation assesses whether agents collectively reproduce the distributional behavioral patterns observed in human discussions. Results show that user's comment history exemplifying interaction behavior consistently outperforms explicit persona information. In addition, models systematically underproduce negative or extreme sentiments, while over producing references and suggestions, revealing biases toward agreeableness and positivity. Crucially, these patterns hold robustly across languages, with small cross-lingual differences.
cs.CL / 31 / 2610.08544
How High Is 0.6? Floors, Ceilings, and Headroom in Interpretability Probing
Abstract
Probes are the workhorse of interpretability. If a model's hidden states predict a variable, the model is said to represent it. But a probe score has no fixed meaning. An $R^2$ of 0.6 may only reflect what the input already gives away, and the same score can mean different things on different data. We propose reading every probe score against two reference points: a floor, what a declared set of simple inputs already predicts, and a ceiling, what the full input can predict. The gap between them, the headroom, is the range in which a probe can show that a model computes something beyond the simple inputs. We prove that headroom vanishes in two ways: the target stops depending on a hidden variable the model must infer, or the input stops revealing it. We test this on transformers trained for in-context meta-analysis, which must infer the hidden heterogeneity between studies to weight them correctly, and where both reference points are known. Under distribution shift, probe scores fall and prediction error rises $12$--$15\times$, yet the model recovers a similar share of the headroom, indicating that the data lost information, not the representation. We then analyze the real models. The single-cell foundation model scGPT encodes biological variability only partially. We also revisit four influential LLM probing studies, which claim that models represent geography, the state of an Othello board, truth, and the demographics of their users. Against a floor computed from the input text alone, some of these claims hold, while others are largely explained by the text itself.
cs.CL / 32 / 2610.08601
Generative AI translations in high-stakes emergency messaging
Abstract
Emergency messaging such as extreme-weather reports and earthquake instructions can involve high stakes, to the extent that translation errors can lead to tragic consequences. The use of machine translation or generative artificial intelligence might therefore not be recommended. On the other hand, time savings in the initial translation can allow greater investments of resources in revision and authorization processes, as well as a wider range of target languages. An experiment with generative AI translations of an earthquake instruction text from English into Chinese and Spanish shows that use of discourse-specific prompts can considerably improve understandability and actionability, although the translations may still not be trusted by translators. Human revision is still required, not only to detect errors but also because of the ethical need for someone to take responsibility for any errors or delays in such messaging.
cs.CL / 33 / 2610.08604
InterCorrect: Intersection-Aware Correction of Demographic Model Merging for Fair ASR
Abstract
Automatic Speech Recognition (ASR) systems often show uneven performance across demographic groups, and errors can be especially difficult to address for speakers belonging to multiple demographic groups. This work studies demographic-aware model merging for fair Speech-LLM-based ASR. Starting from a SLAM-ASR-based model, we fine-tune only the connector on demographic-specific subsets and merge the resulting subgroup-adapted connectors into a global model. We then identify critical cross-axis demographic pairs using subgroup WER and task-vector conflict, and apply intersection-specific correction vectors to the global merged model. Experiments on Fair-Speech show that global demographic merging improves overall WER over the base model, while intersection correction provides additional gains for several merging strategies. In particular, TIES with WER-based correction achieves the best overall WER, reducing it from 7.38\% to 5.13\%. Subgroup and disparity analyses further show that the proposed approach improves performance across demographic axes, while highlighting that lower average WER does not always imply reduced subgroup disparity.
cs.CL / 34 / 2610.08660
Evidence-Bound Reasoning: Neuro-Semantic Verification of Biomedical AI in Glioblastoma Radiogenomics
Abstract
Background: Biomedical AI can generate plausible explanations without reliably verifying whether each statement is supported by patient-specific evidence. We developed a neuro-semantic verification framework that converts radiomic measurements into addressable evidence records and machine-checkable claims. Methods: UPenn-GBM radiomics were aligned with de novo CaPTk extraction from standardized MRI and expert-validated segmentations in an independent multicenter cohort. The shared space comprised 1,728 features from T1, T1GD, T2, and FLAIR MRI across three tumor regions. Reference-defined semantic states were derived from 611 UPenn cases. We evaluated cross-cohort transportability, model-linked provenance, deterministic verification, controlled predictive degradation, and an LLM claim-extraction pilot; MGMT prediction served only as a transport stress test. Results: Median semantic-state agreement was 0.786 (weighted kappa 0.709), ranging from 0.918 for morphologic to 0.252 for intensity features. The external evidence ledger contained 1,655 model-linked records for 331 patients. The verifier achieved 100% exact-set accuracy in a 6,620-claim corruption benchmark. In a 24-case pilot, GPT-5.6 Sol reproduced 72/72 prespecified atomic claims, and the frozen verifier recovered 24/24 expected conditions. During controlled degradation, ROC AUC declined from 0.899 to 0.500 while verification accuracy remained 1.000. External MGMT discrimination was weak (ROC AUC 0.543). Conclusions: Verifiability can be engineered and evaluated independently of predictive performance. LLMs may structure explanations, while final evidence-consistency checking remains deterministic.
cs.CL / 35 / 2610.08675
Same-Number Citation Swaps: Stress-Testing Jev as a Financial Evidence Judge
Abstract
Financial reports repeat values across periods, metrics and accounting lines, allowing an LLM-generated calculation to be numerically correct while citing the wrong financial role. We evaluate what probabilistic evidence verification adds beyond number matching using Jev as a source-support verifier for GPT-4.1-mini calculation traces. A signed-number-at-pointer baseline explains most recovery over exact quotation checks. To isolate the remaining role-recognition problem, we hold operands and arithmetic fixed, move citations between same-number cells, and retain controls that express equivalent facts. These contrasts reveal both wrong-role citations that pass and valid alternative citations that are withheld. Explicit column labels improve selected wrong-role decisions while also lowering support for some equivalent evidence. A constructed follow-up on 36 new source pages, labeled by a non-author reviewer, extends this evaluation and exposes the same tradeoff between detecting role errors and retaining valid citations. The contribution is a controlled evaluation that identifies what a probabilistic financial verifier distinguishes when numerical matching is held fixed. For LLM-based financial assistants, it makes numerical correctness, cited-role support and acceptance outcomes separately assessable.
cs.CL / 36 / 2610.08718
When Forgetting is not Catastrophic: On the Mechanics of Spurious Forgetting
Abstract
Knowledge that a language model appears to forget during finetuning often remains stored and can be recovered, a phenomenon called spurious forgetting. Finetuning on new facts can even produce forgetting that undoes itself: recall of the old facts collapses, recovers as training continues on new facts alone, and only then erodes for good. We seek to understand when such forgetting is not catastrophic. A minimal associative memory reproduces these dynamics with three ingredients: keys with shared structure, concentrated new values, and normalization in the network. Finetuning moves all old representations along a common direction, hiding the old facts while preserving their relative geometry; normalization withdraws this shift once the new facts are learned, whereas fact-specific changes accumulate and cause the erosion. Moreover, subtracting the common shift eliminates the collapse in a Transformer trained on synthetic data, and removing a single direction from each weight update restores old facts in a pretrained language model. Forgetting thus combines a shared, reversible loss of access with a slow erosion of individual facts, and only the second is catastrophic. Which one dominates depends on whether the new data move old memories together or apart.
cs.CL / 37 / 2610.08719
Holdout Best-of-N: Unbiased Evaluation and Its Cost
Abstract
Reusing the scores that select a Best-of-$N$ winner can overstate its expected reward. We study evaluation from a fixed matrix of $K$ independent scores per candidate for a policy that selects using $J$ fresh scores. A single estimator based only on this matrix is exactly unbiased for expected judge reward under every independent, stable collection of candidate-specific score laws if and only if $J<K$, for every pool size $M\ge N\ge2$. At $J=K-1$, the selector deepens as $K$ grows. For independent Gaussian scores with common variance and fixed $M\ge N\ge2$, the unbiased minimax risk in this regime is of order $σ^2/\sqrt K$, attained by Holdout; allowing bias improves the rate to $σ^2/K$. For two candidates, we derive the minimum-variance unbiased estimator at known variance and the sharp asymptotic unbiased minimax constant $1/(π\sqrt2)$, which Holdout attains without knowing the variance. The cyclic average over subsets and ties can be computed in $O(MK\log M)$ operations. At fixed selector depth, cyclic evaluation of bounded scores has $O(K^{-1})$ risk uniformly in pool size. The impossibility result concerns the fixed matrix: one additional fresh winner score permits unbiased evaluation of the all-$K$ policy.
cs.CL / 38 / 2610.08773
AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model
Abstract
Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a task stops teaching once the agent solves it. We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model. The curriculum is rewarded for tasks the agent solves about half of the time, and the adversary only for a success flip, an injection that turns a judged success into a failure. Training in the simulator makes a 4B agent both more capable and more robust: its completion rises with and without attacks, holds against a frontier-model adversary it never trained against, and its capability gain carries over to a real browser. On 150 web tasks, AdvSim2Real raises completion under this unseen adversary by 33.6\% relative to the base agent.
cs.CL / 39 / 2610.08781
IdeaAnchor: Teaching LLMs to Turn Literature into Research Ideas
Abstract
Scientific research often begins by synthesizing ideas from a set of related papers to identify gaps and formulate new directions. However, training language models to perform this form of literature-grounded ideation remains challenging, as existing approaches based on prompting or feedback lack structured supervision for how papers should be synthesized. We introduce IdeaAnchor, a paradigm for training LLMs to perform research ideation using structured specifications as privileged signals. Each IdeaAnchor instance encodes how each input paper should be synthesized into a successful idea, including their functional roles, relationships, and target synthesis criteria. We build this paradigm by mining instances from published papers, capturing how real ideas emerge from prior literature. We then train models via demonstration, self-distillation, and reinforcement learning, and further enhance generation with retrieval at inference time. Experiments show consistent improvements in ideation quality. Our analysis reveals a functional decomposition: anchor-based training strengthens creative synthesis, retrieval enhances detail elaboration, and combining both yields the best performance.
cs.CL / 40 / 2610.07355
Tracking Is Not Permanence: What Video World Models Keep of a Hidden Object
Abstract
Video world models track objects they can see; we ask what they keep of objects they cannot. We hide an object from a frozen V-JEPA 2 predictor and compare its prediction for the hidden region with the encoder's representation of two worlds that differ only inside that region. The predictor's decision keeps a stationary object in part and one carried inside a container not at all, and loses a moving one within 0.3 s (0.5 s under V-JEPA's own tube mask; ViT-H keeps it to 1.1 s at pretraining's 90% masking ratio); in projection a trace remains, below the midpoint, at 14-60% of what a baseline copying the last view retains. The information is there: the encoder reads the object's presence at 1.00 and keeps a closed container's contents decodable for 3.5 s, while the predictor's output, read with the encoder's own probe, contains the ball in 2% of scenes once the box has been closed for half a second. On rendered scenes, permanence is missing on the predictor's side, and training installs it cheaply as a prior: three thousand predictor-only steps on synthetic containers take this belief from 0.05 to 1.00 against two matched controls. They also raise IntPhys-2019 from 84.2% to 93.3%, but so does a curriculum without containers, and which training habit the benchmark credits changes with its scoring rule. Continued training with tube masks produces 1.1-1.6 s of moving-object carry-over on manipulation and internet-style video, so the deficit is not intrinsic to latent prediction. VideoMAE keeps almost nothing, and Cosmos's next-token prediction keeps a stationary hidden object but not one carried inside a moving container.
cs.CL / 41 / 2610.08162
The Failure Is in the Readout: Fine-Grained Emotion Recognition Benchmarks Measure Elicitation, Not Perception
Abstract
Fine-grained emotion recognition supports therapy tools and social robots, but it needs facial data, which raises privacy and data-protection concerns. EmoNet-Face-HQ answers that with generated portraits, expert-rated over a $40$-category taxonomy far finer than the usual six to eight basic emotions. Under the protocol it ships with, vision-language models (VLMs) score poorly on that taxonomy, and the benchmark concludes that a dedicated fine-tuned model is necessary: Empathic-Insight-Face (EIF; Small/Large). We show that off-the-shelf VLMs match or beat that fine-tuned model when the answer is not generated but read from the logits, as one binary query per category. We keep the benchmark's images, taxonomy and ratings, and change only how the answer is read. Experts agree at $κ_w = 0.468$ on the five categories they measure most reliably. Generatively, no interval among eleven open-weight VLMs lies entirely above that anchor ($κ_w=0.268$-$0.486$). Under verification all eleven clear it, each of them significantly better at $κ_w=0.507$-$0.586$. Three also significantly beat EIF sitting at $κ_w = 0.551$ (Small; $0.534$ Large). The gain comes from the graded probability and not from asking a yes/no question: as a control, thresholding those same probabilities to yes/no costs 142% of the average gains and drops binarization below generative elicitation to $κ_w=0.254$-$0.423$. A replication on real photographs (FACES) is weaker and mixed: of the ten models that pass a validity gate, six gain, three are neutral to positive and one is negative, so the effect is not confined to synthetic data.
cs.CL / 42 / 2610.08560
Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness
Abstract
Streaming video-language models must decide not only what to answer, but whether the evidence needed for the current question has arrived. Existing systems learn that decision as a separate trigger; we ask whether an unmodified model already computes it. We show that frozen VideoLLMs carry a linearly readable evidence-readiness signal, labelled from timestamped evidence rather than from model output. It decodes in all seven models of a shared byte-identical evaluation (AUROC 0.733-0.905 under the strictest not-ready sampling, where a fitted clock is near chance), and a probe fitted without any of a benchmark family's footage still reads that family. It is question-conditioned: on byte-identical windows, changing only the question reverses the readout on 66.1% of pairs, while every question-blind control is at chance by construction. The model can answer incorrectly and still encode readiness: AUROC remains 0.722 among wrong answers. Readiness also beats uncertainty estimators and their supervised combination on latency-matched answer selection, and tracks independent human judgments more closely than confidence. Released streaming triggers are also linear readouts, yet a trained trigger read on its own base model's activations is approximately orthogonal to readiness and decodes it far less accurately than a probe. We turn the readout into Readiness Gating, an answer-timing policy that improves accuracy by up to +9.75 pp at matched video duration with negligible computational overhead. How much it gains varies with the accuracy headroom the task makes available: across 26 configurations the gain tracks that headroom, and an intervention that moves it over identical pixels moves the gain with it.
cs.CL / 43 / 2610.08732
A Systematic Study of Semantic ID Spaces for Generative Information Retrieval
Abstract
Generative Information Retrieval (GIR) has emerged as a transformative paradigm, shifting document retrieval from a traditional "retrieve-and-rank" workflow to sequence-to-sequence generation, where a model directly predicts document identifiers (DocIDs). While the semantic design of these DocIDs is known to be critical for performance, a fundamental question remains under-explored: what makes a good DocID? Current approaches rely heavily on computationally expensive downstream evaluations, hindering systematic analysis and rapid iteration. In this work, we address this challenge by presenting a comprehensive study on the properties, metrics, and trade-offs that define effective numerical DocIDs. Specifically, our contributions are threefold: First, we propose a unified framework that unifies Product Quantization (PQ) and Residual Quantization (RQ), and their hybrid variants within a single design space. This enables us to systematically study key DocID properties, such as hierarchy versus parallelism, as well as the impact of hyperparameters like DocID length and codebook size. Second, we define a suite of training-free, intrinsic metrics, to quantify DocID quality and evaluate structural fidelity without the overhead of full model training. Through extensive experiments on MS MARCO 300K and NQ320K, we analyze how these structural properties influence retrieval effectiveness.
cs.CL / 44 / 2610.07327
SharedKV-BT: Node-Local Typed Decisions for Behavior-Tree Agents
Abstract
Agent tasks require sequences of interdependent decisions. Autoregressive models support more flexible decision interfaces than conventional classifiers but incur the latency of token-by-token generation. Recent shared-prefix methods reduce this cost by reusing encoded context and scoring multiple decisions in parallel, but do not model decision dependencies or verify execution. We propose SharedKV-BT, where each active node of a behavior tree (BT) exposes stage-local fields and candidates, and Shared-KV scores the candidates in parallel and passes the selected decision to a separate execution system. We tested SharedKV-BT on robot manipulation, mobile navigation, and computer-use tasks. Across three tasks, SharedKV-BT made typed decisions 2.36-4.15 times faster than prompt-matched autoregressive decoding. On the manipulation task, node-local Shared-KV improved joint decision accuracy from 75% to 94% and closed-loop success from 0% to 60%. Fixed-score policy replay showed that stage gating prevented out-of-order actions and external postconditions prevented premature completion.
cs.CL / 45 / 2610.07647
Loud and Clear: Dynamic Activation Steering for Improving Speech Intelligibility in Noisy Environments
Abstract
Speech becomes less intelligible in noisy environments, and humans naturally adapt their voice to compensate. Inspired by this behavior, we investigate whether a text-to-speech (TTS) model can be guided to produce more intelligible speech using activation steering, without retraining. We focus on two characteristics of the Lombard effect: increased vocal effort and hyper-articulation. We introduce a prompt-relative steering mechanism that prevents steering effects from accumulating during generation while allowing their strength to be adjusted dynamically. Across seen and unseen speakers and multiple languages, our method produces systematic changes in Lombard-related acoustic features, preserves speaker similarity (89-95%), and reduces WER under background noise by 7-22% at 1 dB SNR. These results show that pretrained TTS models can be dynamically controlled to generate more intelligible speech without retraining.
cs.CL / 46 / 2610.07209
Forecasting the Growth of Social Media Information Cascades: Towards Human-in-the-Loop Misinformation Triage
Abstract
Limited review teams must identify which emerging claims are likely to keep growing before their eventual reach is known. We center early misinformation triage on this continuation-forecasting problem: predicting subsequent recorded propagation-tree growth from the first 30 minutes of activity. On FibVID, we compare early node count with structural depth entropy, temporal arrival entropy, and their pair while keeping all propagation trees from each original claim in one partition. Across 352 test trees from 59 claim groups separate from training, the combined model raises $R^2$ for log-transformed future growth from 0.307 to 0.323 and reduces log-MAE by 2.4% (95% claim-bootstrap CI, -0.4% to 5.2%). The gain is especially pronounced among 97 high-activity trees: $R^2$ rises from 0.248 to 0.395, Spearman's $ρ$ from 0.394 to 0.529, and log-MAE falls by 11.7% (95% CI, -1.9% to 23.9%). Complementing the 30-minute growth forecast, we analyze the first 15 replies in 563 PHEME threads. In this cohort, the 15th reply arrives after a median of 28.8 minutes; 52.0% reach the fixed reply prefix within 30 minutes and 71.6% within one hour. Even without the LLM-generated factual-accuracy dimension, the remaining stance, communicative, and affective state composition retains cross-event ranking signal (ROC-AUC 0.538); including that dimension increases ROC-AUC to 0.562. In a separate Check-COVID evaluation of 229 claims, reciprocal-rank fusion retrieves a gold evidence document within the top five for 74.2% of claims and within the top 20 for 94.3%; sentence reranking reaches Recall@20 of 58.1%. We propose an integrated human-review system that brings these early forecasts, response patterns, and retrieved evidence together for misinformation triage relying on the potential virality of claims.
cs.CL / 47 / 2610.07626
A Novel Sentence Stress Detection Framework Leveraging Auxiliary Word-Stress Modeling and Loss Optimization
Abstract
Prosodic stress is a crucial aspect of automatic pronunciation assessment (APA), encompassing both sentence stress detection (SSD) and word stress detection (WSD). SSD highlights semantically salient words that shape discourse meaning, while WSD identifies the primary stressed syllable within each word to ensure lexical clarity. However, most prior work treats SSD and WSD as independent tasks, overlooking their shared reliance on prosodic cues such as pitch, duration, and intensity. To address this gap, we propose an effective SSD approach combining SSD with auxiliary WSD via a novel modeling paradigm. In addition, we introduce a word-span stress regularizer (WSR) that concentrates token-level SSD probabilities within each stressed word span. Experiments on the TinyStress-15K benchmark show that the proposed method outperforms strong baselines, with the complete configuration achieving the best SSD result.
cs.CL / 48 / 2610.08063
HINTT Submission to the 2nd MLC-SLM Challenge: Comparing Cascaded and Unified Approaches to Diarization and ASR
Abstract
This paper presents the HINTT system submitted to the 2nd Challenge and Workshop on Multilingual Conversational Speech Language Model (MLC-SLM). We address multilingual speaker-attributed ASR, where systems must determine who spoke when and what was spoken. We investigate two modeling strategies for this problem: a cascaded pipeline that combines speaker diarization with speech-LLM-based ASR, and a unified speech LLM that directly generates speaker labels, timestamps, and transcriptions. Our final submission is based on the cascaded pipeline, consisting of a fine-tuned DiariZen diarization model, a fine-tuned Qwen3-ASR model, and LLM-based generative error correction. For comparison, we also fine-tune VibeVoice-ASR as a unified model using the same official training data. All task-specific fine-tuning and model selection are performed using only the official MLC-SLM data, without external data or pseudo-labels. Experimental results demonstrate that the cascaded system remains more reliable under the MLC-SLM Task 1 conditions, while unified speech LLMs offer a promising direction for future speaker-attributed ASR.
cs.CL / 49 / 2610.08125
Conversation Is a Two-Body Problem: Dyadic Evaluation of Full-Duplex Dialogue Models
Abstract
Full-duplex spoken dialogue models listen and speak at the same time, enabling voice agents to have natural, low-latency interactions that turn-based systems cannot offer. However, they are commonly evaluated against single-sided interlocutors: pre-recorded audio that cannot react, or an automated examiner that reacts in real time but only administers a fixed sequence of tests and is never graded. These single-sided frameworks evaluate only half of a two-body problem, where turn-taking, overlap, and interruption are joint products of two coupled speakers. We propose DyaFDB, a framework that evaluates full-duplex models in a dyadic setup: two models converse directly under assigned roles with cooperative or conflicting goals, and both sides are scored offline with an external judge. DyaFDB probes how the two models behave toward each other, such as how they take turns or carry an assigned role under different interests. We instantiate four tasks as 140 scenarios and record 7,560 conversations, covering six self- and cross-play pairings. Throughout the experiments, we observe that how a model behaves continually reshapes its partner. We thus demonstrate that each model must be both the examiner and examinee of the other, and no single fixed interlocutor can play both parts. We will release the scenarios, role prompts, and recording protocols between two full-duplex models, without any pre-recorded audio.
多智能体系统 (cs.MA)
6
cs.MA / 1 / 2610.08347
Network Intervention by Polling Strategic Agents
Abstract
A planner in a network of strategic agents faces three entangled challenges: the optimum depends on agents' private information, queried agents may misreport to steer the outcome, and exact computation does not scale. We study these challenges in multi-activity network games with heterogeneous private technologies, in which the planner sets non-discriminatory prices. We show that the optimal prices admit a centrality-based decomposition of the welfare kernel: each agent's contribution scales with its squared centrality in a network reweighted by agents' preferences across activities. This decomposition motivates Poll, a polling algorithm in which the planner samples one agent per round, walks briefly through the agent's neighborhood, and updates the price from a local report. From the same decomposition flow three forms of efficiency: computationally, Poll uses significantly fewer operations than exact computation and other distributed methods, requiring up to three orders of magnitude less communication on a real-world network with over 300,000 agents; statistically, its query complexity scales with topology and preference heterogeneity rather than explicitly with population size; and economically, it converges to welfare-maximizing prices while admitting behavior-specific implementations that induce truthful reports and detect adversarial deviations.
cs.MA / 2 / 2610.07663
Joint Workflow and Prompt Optimization for User Behavior Simulation
Abstract
User behavior simulation is the computational modeling of user interactions within information systems through the use of simulated agents in place of live users. It supports system testing and evaluation, decision-making and forecasting, and user experience design. Existing simulators rely on hand-crafted rules or domain expertise that transfers poorly across tasks. SWORD (Simulation-driven Workflow and Prompt Optimization with Role-based Design) is introduced as a framework that jointly optimizes multi-agent workflow topology and natural-language prompts. It is guided solely by a scalar task metric, without domain initialization or task-specific engineering. The experimental results demonstrate that SWORD achieves statistically significant gains over prompt-only, workflow-only, and staged-optimization baselines under a controlled, identical-backbone comparison. Against the strongest published domain-specific baseline, SWORD further improves accuracy while using a smaller backbone model, substantially less training data, and a very reasonable API cost (\$4--\$6 for each dataset). Beyond predictive performance, SWORD autonomously discovers domain-relevant signals, review-sentiment mapping rules and epidemiological decay priors, purely from scalar error feedback, establishing textual gradients as a mechanism for unsupervised feature-importance discovery in user behavior modeling.
cs.MA / 3 / 2610.07704
Independent Multi-Agent Reinforcement Learning with Counterfactual Semantic-Social World Models
Abstract
Fully decentralized multi-agent reinforcement learning (MARL), also referred to as independent learning, requires each agent to learn and act using only its local information and experience, without a centralized critic or inter-agent communication. Such a stringent information structure renders the conventional reward signal ambiguous. A poor return may result from an ineffective ego action, an incompatible teammate response, or an effective opponent response, yet scalar rewards alone do not reveal which explanation is responsible. We argue that agents can learn more effectively by prospectively comparing the consequences of candidate actions rather than diagnosing failures only from realized returns. We introduce CASTLE (Counterfactual Action-conditioned Semantic Tokens for Local Execution in Decentralized MARL), an offline-training, online-in-context guidance framework with two complementary world models. A Local Dynamics World Model, offline pre-trained over agents' local trajectories, summarizes the agent's local trajectory dynamics and partial observability, while a Semantic-Social World Model predicts compact short-horizon task and social consequences for each candidate ego action. The latter is trained from counterfactual simulator rollouts that expose plausible teammate and opponent responses to alternative actions taken from the same logged rollout state. During online learning and execution, both world models remain frozen and are queried by agents using only locally available information. Their prediction logits provide in-context guidance to an independent PPO policy. Across 30 matched seeds on Tag, Spread, and Adversary in the benchmark multi-particle environments, our proposed CASTLE achieves the highest mean final score among the evaluated methods, exceeding the strongest baseline on each task by 10.67, 6.46, and 0.33 normalized points, respectively.
cs.MA / 4 / 2610.08170
Visual Orchestration Tax in Agentic VLM Pipelines: Auditing and Certifying Visual Evidence Reuse
Abstract
Agentic VLM pipelines increasingly pass the same static visual evidence through multiple specialist agents and tools. This design creates an orchestration-level redundancy mode: semantically unchanged images are repeatedly reconstructed as image-conditioned requests at the VLM API boundary. We call this phenomenon visual orchestration tax and develop a measurement-to-certification framework for visual evidence reuse in agentic VLM pipelines. The audit side defines $\mathrm{M1}_{\mathrm{trace}}$ to count raw visual-evidence touches and M2 to measure structural touch redundancy, with query-level distributions, bootstrap confidence intervals, and paired quality tests. Across SeeingEye and MAMMQA on chart, document, general-VQA, and multi-modal-QA tasks, audits reveal 66.8-75.6% visual-evidence touch redundancy, and every audited query exceeds the predefined gate. The certification side introduces SharedVisCache, a contract-aware evidence reuse hook keyed by image content, preprocessing fingerprint, and encoder assumptions. On SeeingEye, contract validation certifies 75.0-75.5% repeated touches as reusable while preserving 350/350 output strings and $Δ\mathrm{M5}{=}0$. At the physical layer, certified hits reduce $F_{\mathrm{vision}}$ from 800 to 200 in ChartQA-200 trace replay and from 200 to 50 inside live SeeingEye translator-stage physical integration, preserving 800/800 replay strings and 200/200 integrated call outputs. The results position visual reuse as a measurable, behavior-preserving property of agent orchestration and define an agent-layer contract that makes backend prefix or token reuse semantically interpretable.
cs.MA / 5 / 2610.08324
Communication-Free Obstacle Localization from Aggregate Wrench Measurements in Leader--Follower Cooperative Transport
Abstract
We consider obstacle localization for a team of robots cooperatively transporting a rigid payload without explicit inter-robot communication. A leader robot directs the payload's motion, while follower robots assist and react to locally detected obstacles. The leader measures the followers' aggregate wrench, i.e., the combined force and torque they exert on the payload, but cannot directly distinguish their individual reactions. We design a follower control law that allows the leader to recover obstacle locations from these measurements. Each follower resists motion toward nearby obstacles, resulting in a piecewise-linear relationship between the payload's translational and angular velocity and the aggregate wrench. Changes between adjacent linear regions reveal an obstacle's bearing and distance and identify the responding follower. We give sufficient conditions for exact recovery at a fixed payload configuration and develop an adaptive probing procedure in which the leader applies translational and rotational inputs to the payload to obtain the required measurements. We demonstrate the performance of the proposed method in simulations.
cs.MA / 6 / 2610.08332
SC3BF: Shifted Collision Cone Control Barrier Function for Dynamic Obstacle Avoidance
Abstract
The collision cone used by velocity-space control barrier functions is conservative: it rejects every relative velocity aimed into an obstacle, however slow. We propose the \emph{shifted collision-cone CBF} (SC3BF), which adds a state-dependent \emph{allowance} to the cone condition, so the robot may approach the obstacle at a rate that grows with distance and with its own speed. SC3BF is enforced by an ordinary quadratic program, and its safe set is forward invariant under bounded inputs without a minimum forward speed or a clearance margin. We prove that a nonzero allowance preserving safety always exists, and derive one in closed form. Against three velocity-space baselines on a kinematic bicycle among up to $100$ moving obstacles, SC3BF reaches the goal more often and modifies the nominal input less than half as much.
软件工程 (cs.SE)
11
cs.SE / 1 / 2610.07872
Online Sign Language Interpretation System
Abstract
The Walloon Region commissioned this study to improve deaf people's access to public services for the new millennium. It focuses on sign language users, about 25,000 adults in the French-speaking Community, for whom written French is close to a foreign language. The core problem is a severe shortage of interpreters: only about twenty work in French-speaking Belgium, so appointments take weeks to arrange and emergencies cannot be covered. The study assessed three options. The recommended one, remote video interpretation, could be deployed at the time of writing: a professional interpreter joins the deaf person and the employee by videoconference, saving travel time and enabling an emergency service. Similar services already ran in Sweden, Finland, the Netherlands and France. Prototype tests in May 2003 were a clear success and showed that interpretation needs at least CIF video above 256 kbps, ideally 384 kbps. The system suits simple administrative tasks better than emotional, medical or legal situations. The second option, fully automated interpretation, was not yet feasible, although partial systems using speech recognition and signing avatars could already serve kiosks, websites and routine procedures. The third option looks further ahead: UMTS and WiFi could eventually support mobile signed video calls and remote interpretation, provided quality of service is guaranteed. The main recommendation is a first pilot in real conditions before any wide rollout, supported by high-bandwidth links as part of e-government modernisation. The system is meant to relieve interpreters, not replace them, so their training, status and numbers should improve. The authors also call for digital training for deaf people, awareness among public-service staff, more sign language content online, and a multidisciplinary avatar project with Belgian, French and European partners.
cs.SE / 2 / 2610.08238
GPU Acceleration of Awkward Arrays: Using Python cuda.compute
Abstract
Awkward Array is a widely used library in high-energy physics (HEP) for representing and manipulating nested, variable-length data in Python. Previous CHEP contributions have explored GPU acceleration for Awkward Array, demonstrating the feasibility and performance benefits of CUDA-based backend while also identifying limitations related to irregular data access, fine-grained kernel launches, and composability of operations. In this contribution, we present recent developments that build directly on these earlier efforts by introducing a CUDA execution model for Awkward Array based on the Python CUDA Core Compute Libraries (CCCL). Using CCCL, we eliminate the need for custom CUDA kernels and can instead use a high-level Python interface. The CCCL-based approach also enables fusion of multiple Awkward operations into a reduced number of CUDA kernels, addressing kernel launch overhead observed in earlier GPU implementations. Lazy execution allows expression graphs to be constructed and optimized prior to kernel generation, improving performance for analysis workflows involving jagged arrays, combinatorial operations, and reductions. In contrast to earlier approaches, this design also emphasizes extensibility, allowing user-defined Python code to be incorporated into GPU execution paths with minimal boilerplate and without breaking existing analysis semantics. We present performance studies that demonstrate improvements over previously reported eager GPU execution strategies for representative HEP analysis patterns. These developments extend the GPU capabilities of Awkward Array toward a more composable and sustainable backend, aligned with the needs of Python-based analysis at the HL-LHC and beyond.
cs.SE / 3 / 2610.07276
SAFESHIELD: A Decision-Organization Framework for Deployment-Time Safety of Small Language Models
Abstract
Deployment-time safety of language models is commonly implemented through runtime guardrails such as input moderation, routing, retrieval verification, and output filtering. Existing deployment frameworks provide increasingly capable mechanisms for these functions, but offer limited guidance on how the safety decisions they produce should be explicitly organized, coordinated, and audited. We formulate deployment-time safety as a decision-organization problem with two elements: responsibility-oriented decomposition of safety decisions and explicit coordination among them. We instantiate this formulation in SAFESHIELD, a deployment-time safety system for small language models that organizes four recurring decision responsibilities (admission, routing, evidence, and release) and records committed decisions in auditable Decision Traces. We evaluate SAFESHIELD through mechanism-level experiments, aggregate stage ablations, controlled coordination ablations, and a deployment-oriented stress suite. Mechanism-level results show that the instantiated safeguards provide the capabilities required by the decision process, while aggregate ablations show substantial degradation in end-to-end safety as the surrounding safety organization is removed. More importantly, dedicated coordination ablations preserve the participating safeguard mechanisms while selectively severing their dependencies: removing admission gating substantially increases false release, and withholding upstream evidence from the release decision reduces release accuracy from 96.0% to 69.5%. These results provide system-level evidence that deployment-time safety depends not only on the capability of individual guardrails, but also on how their decisions are organized and coordinated.
cs.SE / 4 / 2610.07534
The EPIC Framework for Spec-Driven Development
Abstract
One practitioner we interviewed said their team writes "must" instead of "should" when instructing a coding agent, because the agent may treat "should" as optional. Small wording choices matter because agents often fill gaps in their instructions with their own assumptions. Spec-driven development (SDD) asks developers to write a specification, plan, and tasks before the agent writes code. SDD frameworks provide templates for these artifacts, but the templates do not help developers judge whether they have written enough or clearly enough. We studied what good SDD specifications contain. We scored the artifacts of 114 open-source SDD repositories against ISO/IEC/IEEE 29148 and derived practices from the highest-scoring ones. The resulting framework, EPIC, has 40 practices in 10 quality dimensions that guide developers in making expectations and decisions explicit in specifications, plans, and tasks for coding agents. The majority of SDD practitioners (N=15) endorsed every practice. Repositories in the top third of specification quality spent 11.8% of their commits on bug fixes, compared with 20.4% in the bottom third, and had a median of 4x as many contributors. Developers can use EPIC to find and fill the gaps in a specification before the agent does.
cs.SE / 5 / 2610.07557
CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?
Abstract
Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing coding-agent benchmarks focus on tasks such as patch generation or vulnerability detection and rarely assess whether an agent can develop a working checker in a repository from start to finish. We introduce CheckerBench, an executable benchmark of 300 tasks derived from 297 CVEs across 167 repositories, 85 CWEs, and five language ecosystems. Each task includes vulnerable and fixed revisions, a pinned analysis environment, and a checker scaffold. We further introduce CheckerLab, a common evaluation framework that independently rebuilds submitted checkers and measures vulnerable-fixed diagnostic contrast, patch localization, false positives, and tool use. Across 21 model-harness configurations and three independent repeats per configuration, mean Pass@1 is 32.30%, while the best reaches 45.33%. These results show that reliable, reusable checker development remains challenging for current coding agents.
cs.SE / 6 / 2610.07715
When Old Facts Return: Re-Reads, Reverts, and the Limits of Temporal Memory
Abstract
A memory system can retire an obsolete value and later restore it merely because the same old statement appears again. A re-read of an old source and a genuine revert can produce the same observed sequence of values while requiring opposite current answers. We study this ambiguity on 130 extractor-selected atomic transitions derived from software fixes. In the ordinary transition condition, identity-based temporal memory reaches 98.5% model-judged accuracy with zero observed errors under a literal stale-value proxy. Appending a verbatim re-read of the old statement reduces accuracy to 10.8% and raises the stale-value rate to 88.5%. A guard that refuses to reactivate a previously retired value restores accuracy to 97.7% and reduces that rate to 0.8% in this constructed re-read condition. The guard cannot also recognize a legitimate revert without additional change provenance. Two supporting studies examine exposing retired history to the answer model and supplying current source for changed behavior. An exploratory extraction study over 707 software fixes provides scope context, not a universal coverage estimate. The design implication is to distinguish an observation of a value from evidence that the value changed. Selected inputs, aggregate-only answer records, related-family judges and a post-failure guard evaluation limit the conclusions to the retained experiments.
cs.SE / 7 / 2610.07757
Acquiring and Verifying Repository Norms for Coding Agents
Abstract
Changes produced by coding agents can pass functional tests while leaving repository contribution requirements unmet. Following repository-specific norms requires identifying guidance dispersed across repository sources and interpreting its conditions and exceptions. Retrieval and documentation approaches supply general context, but agents must still determine which norms apply. We introduce RepoNorm to acquire explicit and implicit repository norms independently of coding tasks. It checks norm content and applicability using repository evidence, consults Git history when needed, and delivers norm packages to existing coding agents. Our evaluation uses three coding models and 121 tasks from RepoNormBench. Against the baseline with no additional generated guidance (Raw), relative improvements are 7.42-10.77% for Overall Norm Compliance Rate (NCR), 31.64-45.44% for Contribution NCR, and 11.34-17.69% for Prompt-omitted NCR. All three coding models also obtain higher values on these NCR measures with RepoNorm than with CodeWiki documentation. Functional success rates show observed gains of 5.79-9.09 percentage points over Raw; the paired comparisons do not reach the significance threshold. Sampled precision is 86% with RepoNorm's default configuration.
cs.SE / 8 / 2610.07762
ES-Trace: Auditing Ethical-Sourcing Disclosure of Code Generation Models Beyond Model Cards
Abstract
Code generation models have been increasingly used in software development, but their development raises ethical-sourcing concerns involving intellectual property, privacy, fairness, labour practices, and environmental impact. Although prior work has defined ethical-sourcing criteria for code generation, it remains unclear how much evidence existing models disclose and where that evidence can be found. We introduce ES-Trace, a framework for ethical-sourcing disclosure audits that traces disclosed evidence beyond model cards using the Model Documentation Traceability Graph (MDTG), which represents relationships among models, versions, and documentation artifacts. We apply ES-Trace to 26 models from 10 publishers across 77 documents and 20 ES-CodeGen aspects. Model-card-only auditing yields a mean score of 1.77/5, while resolving the declared references increases it to 2.82/5, with most of the increase arising from documents that the publisher declares in structured metadata. The key findings of our study include: (1) social and labour-related aspects remain poorly documented, even when expanding the audit to the full documentation scope, (2) resolving documentation references substantially increases observed disclosure, raising the mean score from 1.77/5 to 2.82/5, (3) model-card-only audits can mischaracterize release-level disclosure changes, and (4) documentation mismatches can associate evidence with the wrong model or version, highlighting the need for explicit model--version binding and consistency across documentation artifacts. Our study calls for reference-aware ethical-sourcing disclosure audits, explicit model--version binding and consistency across documentation artifacts, and stronger documentation of currently underreported social and labour-related aspects.
cs.SE / 9 / 2610.08173
Beyond the Leaderboard: Multi-Dimensional Evaluation of Dense and Mixture-of-Experts Models for Automated Program Repair
Abstract
Automated Program Repair (APR) with language models is usually evaluated by whether a generated patch passes the test suite, which can hide differences in maintainability, security, and computational cost. We propose a Weighted Quality Index (QI), inspired by the ISO/IEC 25010 software quality model, that combines functional correctness, maintainability, security, and generation efficiency under configurable weighting schemes. We evaluate three dense Qwen2.5-Coder models (3B, 7B, 14B) and the 16B-parameter DeepSeek-Coder-V2-Lite Mixture-of-Experts (MoE) model (2.4B active parameters) on 40 QuixBugs and 90 Defects4J bugs, all run locally on identical hardware to control for infrastructure effects. Model rankings change with the weighting scheme, showing that single-metric evaluation can hide trade-offs. The MoE model shows almost no statistically significant difference in correctness from the 7B and 14B dense models (McNemar's exact test) while using 3-6 times fewer active parameters, whereas correctness increases significantly across the three dense scales. These results suggest that active parameter count can be a more informative lens than total parameter count for sparse code models.
cs.SE / 10 / 2610.08429
RAPO-Sol: Retrieval-Augmented Preference Optimization for Repository-Level Solidity Code Generation
Abstract
Smart contracts written in Solidity manage assets, permissions, and irreversible state changes, making code generation both useful and security-critical. Repository-level Solidity generation is challenging because models must synthesize complete contracts or libraries while preserving consistency across state variables, modifiers, events, inheritance, external calls, and access-control logic. We present RAPO-Sol, a two-stage training framework for repository-level Solidity code generation. First, Retrieval-Augmented Fine-Tuning (RAFT) augments each training input with similar Solidity examples, helping the model learn recurring contract-level patterns while remaining retrieval-free at inference time. Second, Direct Preference Optimization (DPO) trains the model to prefer reference contracts over close but semantically flawed alternatives. We construct rejected samples using Solidity Semantic-Anchor Perturbation (SAP), which perturbs validation statements, visibility modifiers, data-location keywords, context variables, payment operations, and low-level calls. Experiments on SolidityBench with CodeLlama-7B-Instruct, DeepSeek-Coder-6.7B-Instruct, and Qwen2.5-Coder-7B-Instruct show that RAFT consistently improves over supervised fine-tuning, while SAP-based DPO provides further gains in BLEU and SolidityScore. The full RAFT+DPO pipeline achieves the best performance across all three models, demonstrating complementary benefits from retrieval during training and Solidity-aware preference optimization without adding retrieval cost at inference.
cs.SE / 11 / 2610.08651
A Case Study in Assuring AI-Written Software
Abstract
Software-engineering agents can enable people without formal software training to build systems they could not otherwise implement and simultaneously can produce more code than even experts can meaningfully inspect. In both cases, exhaustive code review is not reliable as the sole basis for human control. We report a case study of a production healthcare platform built through coding agents and governed by an operator without formal software-engineering training. Over time, its workflow grew into a human-led meta-agent system where one agent wrote code, other agents supervised and reviewed it, and project rules carried lessons forward. The operator found that tests, monitors and reviewing agents used to supervise the system were fallible. Some monitors measured proxies rather than outcomes, some audits failed silently, missing checks disappeared from reported results and one automated repair caused operational disruption. In this case, human control depended on keeping the intended outcome, the evidence used to judge it, the agents' permissions and the final human decision were all tied to the same underlying objective.
硬件架构 (cs.AR)
14
cs.AR / 1 / 2610.07191
Agentic Design Space Exploration for Joint Hardware Configuration Selection and Mapping of AI Inference Workloads on Heterogeneous Edge SoCs
Abstract
Modern edge Systems-on-Chip (SoCs) integrate heterogeneous processing units (PUs) such as CPUs, GPUs, and NPUs, each with distinct performance and energy characteristics. Deploying AI inference workloads on them under real-time latency and energy constraints requires jointly mapping workloads to PUs and configuring each PU (e.g., selecting the number of active cores and the operating frequency). This joint space grows combinatorially, making exhaustive search infeasible. Most prior work on design space exploration (DSE) applies black-box optimization (BBO) such as evolutionary search, where each evaluation returns only aggregate metrics such as latency and energy. Recent LLM-guided DSE relies on the same sparse feedback. We observe that this limits its efficiency: it offers no insight into the design space or the reasons a design choice performs the way it does, and it leaves the reasoning abilities of LLMs largely unused. We present TraceDSE, an agentic DSE flow that performs joint workload mapping and PU configuration selection for AI inference on heterogeneous SoCs. TraceDSE is an iterative proposer-critic loop driven by richer feedback in the form of system execution traces. The LLM proposer agent generates candidate mappings and PU configurations for hardware evaluation. The LLM critic agent, equipped with programmatic trace-analysis tools, analyzes the traces to identify bottlenecks and suggest targeted refinements. This loop yields deeper insight into each design point, higher-quality decisions, and a more effective search. Across four AI inference workloads (models of varying complexity and a multi-model pipeline) on an Intel Meteor Lake SoC, TraceDSE consistently outperforms two state-of-the-art BBO tools, improving Pareto frontier hypervolume by up to 35% over NSGA-II and up to 68% over Bayesian optimization, while requiring ~6-9x fewer hardware evaluations.
cs.AR / 2 / 2610.07301
A Pipelined FPGA Architecture for Banded Sparse Matrix Dense Matrix Multiplication in Longformer
Abstract
Sparse attention mechanisms have become increasingly important for transformer models processing long input sequences due to their lower computational and memory complexity compared to full self-attention. Longformer achieves this through a sliding-window attention mechanism that produces a structured banded sparse attention matrix. However, existing sparse transformer accelerators primarily target attention generation or unstructured sparsity, leaving sparse matrix--dense matrix multiplication (SpMM) for structured sparse attention largely unexplored. This paper presents a pipelined FPGA architecture for accelerating banded SpMM in Longformer. The proposed design exploits the predictable sparsity pattern of Longformer's attention matrix through a custom row-wise storage scheme with implicit indexing, eliminating the overhead of conventional sparse matrix formats while enabling regular memory accesses. The architecture employs parallel processing elements, pipelined adder trees, and a dual-path computation strategy to maximize throughput and hardware utilization. Implemented in Verilog and evaluated on an RFSoC platform using Vivado 2024.2, the accelerator sustains one complete dot-product result per clock cycle after an initial latency of 11 cycles while maintaining power consumption below 2.9 W. Operating at 100 MHz, the design achieves over 100 million dot-product outputs per second, demonstrating the effectiveness of directly exploiting structured sparsity for sparse transformer acceleration.
cs.AR / 3 / 2610.07600
Beyond No-Good Benders Cuts: Exact Realizability for Discretely Tunable Clock Trees
Abstract
Useful-skew schedules can satisfy timing constraints yet remain unrealizable by a fixed clock tree with discrete tuning choices. We formulate this mismatch as exact membership in a finite relative-latency set and develop a checkable feedback interface between the scheduler and the tree model. Arithmetic certificates explain unrealizable targets, while tree-specific contraction reduces the structural size of the exact relations projected onto selected sinks. These relations become realizability cuts that can exclude more candidates than a no-good on the same certificate support, including discrete holes that linear inequalities cannot separate. Controlled experiments confirm fewer oracle calls and faster projection construction. They also expose important limits: globally minimum certificates can cost more than they save, compact arithmetic feedback can fail on non-parity obstructions, and direct monolithic optimization remains faster on the tested additive models. The contribution is an exact, independently verifiable scheduler--tree interface, rather than a claim of universal solver acceleration or physical signoff.
cs.AR / 4 / 2610.07685
CHiRP: Control-Flow History Reuse Prediction
Abstract
Translation Lookaside Buffers (TLBs) play a critical role in hardware-supported memory virtualization. To speed up address translation and reduce costly page table walks, TLBs cache a small number of recently-used virtual-to-physical address translations. TLBs must make the best use of their limited capacities. Thus, TLB entries with low potential for reuse should be replaced by more useful entries. This paper contributes to an aspect of TLB management that has received little attention in the literature: replacement policy. We show how predictive replacement policies can be tailored toward TLBs to reduce miss rates and improve overall performance. We begin by applying recently proposed predictive cache replacement policies to the TLB. We show these policies do not work well without considering specific TLB behavior. Next, we introduce a novel TLB-focused predictive policy, Control-flow History Reuse Prediction (CHiRP). This policy uses a history signature and replacement algorithm that correlates to known TLB behavior, outperforming other policies. For a 1024-entry 8-way set-associative L2 TLB with a 4KB page size, we show that CHiRP reduces misses per 1000 instructions (MPKI) by an average 28.21% over the least-recently-used (LRU) policy, outperforming Static Re-reference Interval Prediction (SRRIP), Global History Reuse Policy (GHRP) and SHiP, which reduce MPKI by an average of 10.36%, 9.03% and 0.88%, respectively.
cs.AR / 5 / 2610.08186
A gem5-based Simulation Framework for Computing-in-DRAM
Abstract
Computing-in-Memory using DRAM (CIMD) has demonstrated substantial energy and throughput gains for memory-bound workloads consisting of bulk-bitwise operations, by performing computation directly within DRAM subarrays. Realizing CIMD, however, requires a redesign of the memory controller and careful mapping of operands onto the memory arrays. Presently, accurate FPGA-based testbeds exist, but they are costly and labor-intensive, while open-source simulators are largely trace-based and cannot execute full application runs. The only full-application CIMD simulator available, included with MIMDRAM, is built on an outdated version of gem5 and its toolchains. We present gem5-CIMD, a full-application CIMD simulation framework built on the latest version of gem5. We extend the simulated instruction set from four bitwise operations to sixteen CIM instructions spanning arithmetic, relational, conditional, and utility operations across a range of data types and bitwidths (4 to 64 bits). A complete memory-management stack comprising multi-level huge-page support, a CIM-aware allocator, and a subarray-aware address-mapping scheme ensures that CIM operands are co-located in the same DRAM subarray. In addition, gem5-CIMD is aware of the DRAM refresh interval and faithfully models and schedules the mandatory refresh operations when needed. We provide a CIM standard library intended as a compiler target and evaluate gem5-CIMD on multiple case studies including an end-to-end KNN workload. The simulator sources, including the evaluated workloads, are publicly available.
cs.AR / 6 / 2610.08190
On the Impact of Degradation-Balanced Scheduling in Manycore Systems
Abstract
Electromigration degradation in manycore systems depends strongly on how runtime workload is distributed across cores. Task placement affects power, temperature, and current density, which in turn shape the spatial distribution of degradation over time. However, a scheduler that ranks cores only by their current reliability state may not always distinguish among candidate cores, especially when reliability values are close at decision time. In this case, scheduling decisions can be driven by secondary factors rather than by long-term degradation balance. This paper studies degradation-balanced scheduling and proposes a Reliability-Balanced (RB) policy that predicts the full-task effect of each candidate assignment. RB penalizes added imbalance in electromigration exposure, stress, active time, and task count while guarding the predicted Rvalue floor. On randomized continuous mixed workloads, the proposed scheduler preserves completed tasks while improving average Rvalue by 0.37% and minimum Rvalue by 4.35%. It reduces Rvalue, exposure, stress, and active-time standard deviation by 10.84%, 9.39%, 8.35%, and 16.38%, respectively.
cs.AR / 7 / 2610.08373
ECO: Energy-Oriented Configuration Optimization for Attention FFN Disaggregated LLM Serving
Abstract
Energy-efficient LLM serving requires minimizing serving GPU energy while meeting latency and throughput service-level objectives (SLOs). Attention--FFN disaggregation (AFD) enables separate resource allocation and operating controls for attention and expert computation, but their energy effects remain coupled through the execution pipeline. Realizing its energy-saving potential therefore requires navigating a hierarchical configuration space in which deployment structures constrain admissible controls and shape their end-to-end effects. Finding low-energy configurations that meet SLOs is challenging because physical evaluations are costly and only a small fraction of candidates can be measured. We present Energy-Oriented Configuration Optimization (ECO), which jointly searches deployment structures and their admissible operating controls under a limited measurement budget. ECO constructs a structure-aware energy prior from calibrated stage behavior and pipeline dependencies, then learns residual prediction errors with a Gaussian process. Its cost-aware constrained Bayesian optimization prioritizes measurements according to expected energy improvement while accounting for SLO feasibility, execution success, and evaluation cost, and returns the lowest-energy measured feasible configuration. Across all 16 scenarios on A6000 and A100 with Qwen and DeepSeek, ECO's frozen configurations, evaluated on disjoint requests, reduce serving energy by 40.5\% and increase output token rate by 20.7\% on average relative to baselines while meeting target SLOs. Across the 8 A6000 scenarios, its selected feasible energy averages 33.1\% below generic constrained Bayesian optimization and 25.8\% below genetic search.
cs.AR / 8 / 2610.08483
Trail: Scalable and Low-Cost Temporal TLB Prefetching via Page-Table-Embedded Deltas
Abstract
Address translation is a major bottleneck in data-intensive workloads. TLB prefetching can hide translation latency, but existing spatial prefetchers struggle with irregular accesses, while temporal prefetchers store address deltas in fixed-capacity hardware that cannot scale with application memory footprints. Our characterization of 200 translation-intensive workloads reveals that each virtual memory region has a small, recurring set of deltas, and over 94% of deltas fit in 18 signed bits. We introduce Trail, a temporal TLB prefetcher that stores deltas in unused bits of leaf page table entries (PTEs). When a region triggers a page table walk, Trail identifies the region that last triggered a walk from the same instruction and records their delta in that source region's PTE. When a later walk fetches the source region's PTE cache block, Trail retrieves its deltas without additional memory accesses and prefetches translations for likely destination regions into the TLBs and cache hierarchy. Storing multiple deltas per PTE cache block improves coverage, while using existing PTE bits allows metadata capacity to scale with the application's memory footprint without additional metadata storage. Across 200 workloads and 100 multiprogrammed mixes, Trail improves single-core (four-core) performance by 5.7% (11.5%) on average over a baseline without TLB prefetching, outperforming the best prior standalone TLB prefetcher by 1.7% (2.5%). Trail requires only a 64-entry hardware table to track per-instruction page table walk history. Trail is freely available at https://github.com/CMU-SAFARI/Virtuoso/tree/trail-artifact-release.
cs.AR / 9 / 2610.08502
X-OPM: Explainable Automatic Digital On-Chip Power Modeling for Enhanced Robustness
Abstract
Proactive power management systems reduce processor dynamic power through runtime power prediction and power-aware scheduling. Accurate, stable and low-overhead digital on-chip power meters (OPMs) are crucial for improving the prediction quality. Recent studies have explored various modeling methods, including using linear models, decision trees, and multi-layer perceptrons (MLPs) to construct OPMs. However, most current approaches train models end-to-end without analyzing the physical interpretability of features, affecting their ability to generalize to unseen workloads. Grounded in the design principles of synchronous digital VLSI circuits, X-OPM introduces a robust feature engineering framework that uses tree-based models to capture feature interactions and linear models for prediction. It also incorporates a human-in-the-loop workflow to balance model accuracy against modeling effort. Evaluated on a commercial C906 vector processor, X-OPM consistently achieves $R^2 > 0.93$ across all workloads with sampling window size set below $8$ cycles. In contrast, state-of-the-art methods including APOLLO, COBIT, and standard MLPs fail to generalize across all test cases. Layout with commercial EDA tools shows that X-OPM incurs an area overhead below $0.1\%$, which is on par with lightweight tree-based and linear models, and significantly smaller than MLP-based models.
cs.AR / 10 / 2610.08688
A Framework for Accelerating Transformer Inference on RISC-V for Edge AI
Abstract
This work presents a framework for accelerating transformer-based language models (LMs) on resource-constrained IoT devices. The framework targets compact LMs: BERT-Tiny (B-Ty), MobileBERT (M-Bt), MiniLM (M-Lm), Electra (E-Lt) and DeBERTa (D-Bt) -- selected for their architectural diversity and use in edge inference scenarios. The proposed flow derives lightweight instruction set extensions tailored to the non-obvious computational patterns of these models. In addition, a custom instruction is introduced to accelerate the address generation stage of batch matrix multiplication, achieving a performance improvement of 15.39--21.74% with modest ASIC overheads of 6.79% in area and 2.33% in power. To further enhance performance without incurring additional processor core hardware cost, an optional compiler-directed loop unrolling strategy is employed, trading increased code size for overall reduced execution time. Evaluation on the Synopsys trv32p3f RISC-V core demonstrates inference speedups of up to 2.19x, and FPGA implementation on the AMD Zynq UltraScale+ ZCU102 shows a 32.94% area overhead at 75 MHz, whereas the ASIC implementation using the TSMC 28 nm library incurs a 21.08% area overhead while operating at 250 MHz.
cs.AR / 11 / 2610.07593
TRANSIT: Transparent Scale-in for Multi-Node LLM Training
Abstract
TRANSIT is a transparent scale-in framework to enable multi-node model training on fewer GPUs while maintaining training efficiency by transparently leveraging CPU DRAM as an extension of GPU memory during distributed training. It achieves this through a user-space interposition layer, requiring no modifications to the application, training framework, cluster scheduler, device driver, or operating system. Furthermore, TRANSIT achieves higher efficiency by leveraging a zero-copy data path for CPU-GPU transfers. We evaluate TRANSIT on dense and MoE models across scales up to 64 NVIDIA H100 GPUs and multiple parallelism configurations over a RoCE network. Our evaluation shows that TRANSIT can: (a) outperform state-of-the-art framework-managed offloading techniques, achieving up to 68%, 59%, and 42% higher per-GPU throughput than TorchTitan, ZeRO-Offload, and ZeRO-Infinity, respectively, (b) enables training with 50% fewer GPUs while maintaining over 90% of baseline per-GPU throughput, (c) lower per-node network traffic by up to 33%, and (d) improve per-GPU throughput by up to 35% in communication-bound settings.
cs.AR / 12 / 2610.08372
vTen: Tensor-Centric Verification Framework for Domain-Specific Accelerators
Abstract
The semantic gap between tensor-centric software models and signal-level hardware testbenches creates significant productivity bottlenecks in verifying Domain-Specific Accelerators (DSAs). Existing frameworks like Cocotb suffer from prohibitive synchronization overheads due to fine-grained interactions. To address this, we propose vTen, a data-centric framework that strictly decouples verification intent from execution mechanics. By leveraging a declarative DSL and kernel-granular batching, vTen minimizes host-simulator interaction frequency. Evaluation on a production-scale 3D U-Net accelerator demonstrates that vTen achieves a 2x performance improvement in simulation latency and a 60.3% reduction in code complexity compared to Cocotb.
cs.AR / 13 / 2610.07644
From ASIC to Fleet: Lessons from Building and Operating a Hyperscaler NIC
Abstract
We describe the operational infrastructure built to deploy and operate fbnic, a custom multi-host NIC, across hundreds of thousands of production hosts at Meta. Vendor multi-host NICs, designed by retrofitting single-host architectures, suffered from shared firmware and buffers that created cascading isolation failures over seven years. fbnic eliminates these through physical isolation, but shifting to in-house hardware shifts the entire operational burden to the hyperscaler. We present a hardware-in-the-loop CI pipeline testing firmware, driver, and kernel cross-products; a unified observability pipeline co-locating NIC and switch counters for cross-layer fault attribution; a driver-first architecture with fewer than ten firmware message types; a targeted firmware upgrade orchestrator at sub-sled granularity; and scoped repair automation confining blast radius to individual host slices. Over ten months, fbnic achieved a 12X reduction in unplanned unavailability, 37% lower mean time to repair, and 2.3X fewer hardware swaps compared to vendor NICs on the same platform.
cs.AR / 14 / 2610.08444
ActTune: Action-Aware Precision and GPU Operating-Point Adaptation for Energy-Efficient Vision-Language-Action Inference
Abstract
Vision-language-action (VLA) policies repeatedly invoke inference to control robots, making graphics processing unit (GPU) energy a recurring cost of task execution. Reducing energy per inference call, however, may not reduce energy per successful task if numerical errors increase failures or slower inference prolongs execution. We therefore target GPU energy per successful task while preserving task success and keeping the inference-latency increase within 10\%. Our approach builds on two observations: quantization sensitivity varies across action classes, model layers, and weights versus activations; and numerical precision changes the workload, shifting favorable GPU operating points. We introduce ActTune, an action-aware framework that connects layer-wise precision allocation with workload-dependent GPU operating-point selection over requested frequency--power-cap pairs. A lightweight decision tree learns its splits and leaf precision configurations directly from configuration action errors, then selects precision before each policy call. The controller forecasts the next workload and applies the selected GPU operating point asynchronously using a lookup table calibrated under a latency budget. A shared resident quantized weight bank enables configuration switching without weight reconstruction or additional policy evaluations. On LIBERO, a benchmark for lifelong robot learning, ActTune improves mean task success by up to 2.3\% relative to state of the art. Relative to the original BF16 implementations, it delivers up to $2.02\times$ faster inference and, with GPU operating-point adaptation, reduces energy per successful task by up to 76.8\%.
密码学与安全 (cs.CR)
42
cs.CR / 1 / 2610.07258
Lineage-Aware Memory Governance: A Derivation-Gated Framework for Privacy-Preserving Column-Level Access Control in Enterprise AI Agents
Abstract
Enterprise AI agents that share a memory store face two unaddressed risks: sensitive data can leak through legitimately computed results the requester could not derive, and departments can silently compute a same-named key performance indicator (KPI) through conflicting logic. Existing agent-memory systems (e.g., MemGPT, Zep, A-MEM) gate retrieval by content, ownership, and role, not derivation, missing a cached insight that embeds a forbidden column. We introduce the Analytical Memory Unit (AMU), a memory schema that attaches a full derivation (lineage) graph to every cached result, gated by a retrieval policy that serves a hit only when the requester is authorised for every column touched. Provided lineage recording is complete, we prove by construction that the policy blocks retrieval of results derived from a sensitive column outside the requester's permissions, at O(n) worst case -- a conditional design guarantee, not an empirical claim, that excludes derived features encoding sensitive information without naming their source. Eliminating measured leakage required 75-90% recorded lineage completeness, so we treat 90% as a conservative deployment target. Across six experiments, lineage-gated retrieval removes the 18.8-25.5% cross-department leakage naive content-gated memory suffers, keeping 81.5-82.6% of memory reuse at 13.8 microsecond worst-case overhead. A real-agent proof-of-concept with LLM-generated SQL is consistent with the guarantee: zero leaks over 9 round-trips, two conflicts caught automatically -- though a feasibility demonstration, not evidence of production viability. This offers a practical governance layer for shared agent memory, complementing source-layer access control and supporting EU AI Act compliance.
cs.CR / 2 / 2610.07282
A Resilient Runtime-Verification Fabric for Security Monitoring of Critical Edge-IoT Infrastructure
Abstract
Protecting critical infrastructure increasingly depends on continuously verifying large IoT fleets against formal security specifications at runtime. Yet the runtime-verification (RV) pipelines proposed for this task are typically single-host prototypes whose monitors read a shared log file, with no resilience to the failures such deployments incur: a crash or overload silently drops events, clock skew corrupts the ordering metric monitors require, a time-triggered "node has gone silent" property cannot fire when the network itself falls silent, and one slow consumer stalls the pipeline. Each failure is silent: the monitor keeps emitting verdicts over a corrupted view. We present RV-Fabric, a resilient delivery layer that carries the hierarchy over two brokers (MQTT for device ingest, a durable stream broker for backend delivery) and re-establishes five continuity guarantees: durable delivery under crashes, a trusted event order, progress under total silence, consumer isolation and flow control under bounded overload, each an invariant conditioned on broker durability. Above the transport, RV-Fabric makes evidence completeness part of runtime-verification semantics: every verdict carries a status (sound, degraded, incomplete or unavailable) derived from delivery gaps, retention pressure and liveness, so an incomplete stream cannot yield an unqualified all-clear. Under controlled fault injection on a containerised testbed, measured against a fault-free oracle using the real MonPoly engine, the shared-log baseline misses six of seven injected incidents, reporting each as an unqualified all-clear, whereas RV-Fabric preserves all seven; removing a delivery mechanism reintroduces silent loss, removing isolation costs only timeliness. Two published critical-infrastructure datasets, water-SCADA and IoT/IIoT, replay end-to-end.
cs.CR / 3 / 2610.07298
Polar: LLM-Powered Synthesis of Real-World Cyber Evidence for Prioritization and Mitigation
Abstract
Cyber threat analysis increasingly depends on evidence distributed across vendor advisories, vulnerability databases, and threat intelligence sources. Turning these fragmented observations into timely decisions requires models to connect technical severity with evolving exploitation evidence and available defensive actions. We present POLAR, an LLM-powered framework for synthesizing real-world cyber evidence into threat-centric assessments for prioritization and mitigation. POLAR first disentangles overlapping incidents and grounds each threat in source-linked evidence. For prioritization, it infers severity metrics from cyber evidence and combines the resulting assessment with temporally ordered exploitation signals to estimate near-term exploitation likelihood. For mitigation, it links the synthesized threat data to authoritative remediation knowledge and organizes applicable actions according to threat urgency and operational constraints. We evaluate POLAR on real-world vulnerability evidence collected from public resources and compare it with multiple baselines. Across heterogeneous incidents and zero-day settings, POLAR improves threat ranking and mitigation retrieval while producing evidence-linked intermediate assessments that support analyst inspection. The results establish evidence synthesis as a practical foundation for LLM-based cyber decision support across related security tasks.
cs.CR / 4 / 2610.07310
From Sandbox to Enforcement: Confidence-Qualified Threat Intelligence for Critical Infrastructure
Abstract
Security operations centres and national incident-response teams defending critical infrastructure collect abundant threat data yet struggle to turn it into actionable intelligence. A malware sandbox produces detailed behavioural evidence, but as a large, unranked report whose confidence is unstated. We present CG-CTI, an operational pipeline that converts live sandbox output (CAPEv2) into STIX 2.1, correlates it in a knowledge graph with other critical-infrastructure sensors, and attaches to every intelligence object an explicit confidence status derived from provenance, cross-source corroboration, and observation durability. This status gates automated action: only corroborated intelligence is eligible for automated enforcement, while lower-confidence objects are routed to analyst review or kept as context. A grounded language-model stage then narrates the confidence-qualified evidence, where each statement either cites a supporting object or is marked unsupported, so fabricated references are removed before analyst review. We implement CG-CTI within the CYBERGUARD project, whose consortium includes Romania's national cyber-security directorate, and evaluate it against the live sandbox on a labelled malware corpus, measuring conversion validity, indicator yield, technique coverage, corroboration, enforcement eligibility, latency, and summary grounding. CG-CTI turns fragmented sandbox output into corroborated, confidence-ranked, and auditable intelligence for critical-infrastructure defence.
cs.CR / 5 / 2610.07345
Evaluating Behavioral Context for Interpretable IAM Policy Risk Scoring in Cloud Environments
Abstract
IAM policy analysis typically emphasizes the authorization capabilities encoded in a policy, but security analyst review priority may also depend on the behavioral and environmental context surrounding a policy event. This paper evaluates whether contextual information provides measurable incremental value for interpretable IAM policy risk prioritization beyond policy and effective-authorization information. AWS is used as the experimental cloud provider because its IAM and audit-telemetry ecosystem enables controlled evaluation using AWS IAM Context Bench, a benchmark containing 534 real AWS experimental observations across policy, environment, and behavioral scenarios, including matched cases where policy and environment remain fixed while behavioral context changes. Three Explainable Boosting Machine models are evaluated under the same leakage-controlled grouped cross-validation protocol: a policy-centric baseline, a policy-plus-environment model, and a full-context model incorporating CloudTrail telemetry. The full-context model substantially reduces analyst-priority prediction error relative to the policy-centric baseline and closely tracks the reference priority ordering. In matched same-policy context pairs, the policy-centric model remains invariant, whereas the full-context model separates benign and suspicious behavioral conditions with high directional accuracy. The results also show improved concentration of high-priority cases at the top of simulated analyst review queues. These findings indicate that behavioral and environmental context can provide useful incremental information for analyst-oriented IAM risk prioritization while preserving an interpretable additive model structure. The formulation is applicable beyond AWS conceptually, although cross-provider validation remains future work.
cs.CR / 6 / 2610.07351
Simple Extremely Lossy Functions from Small-Exponent Hashing
Abstract
Extremely Lossy Functions (ELFs) are a standard model primitive that captures many useful properties of random oracles (Zhandry, Crypto 2016). While there are many variations of ELFs with additional properties, every construction (excluding obfuscation) has followed essentially the same template of bootstrapping from a sequence of ELFs secure only against fixed-size adversaries, and every construction was based on only the exponential hardness of DDH (or $k$-Lin), an assumption that is only reasonable over elliptic curves. We introduce and construct Extremely Lossy Trapdoor Hashing (ELTDH), a stronger notion that implies all known variations of ELFs. Our construction achieves ELTDH in one go, without bootstrapping from schemes secure for only fixed-size adversaries, which makes it simpler and more efficient than existing ELFs. We assume exponential security of the small-exponent discrete logarithm, together with polynomial security of decisional composite residuosity (DCR). Exponential security is only required in the size of the secret exponent, not the size of the group, so the assumption plausibly holds for multiplication modulo $N^2$ (and for many other cryptographic groups), despite the subexponential-time discrete logarithm attacks from index calculus. Our results diversify the assumptions underlying ELFs, while also giving a simpler construction.
cs.CR / 7 / 2610.07386
NetAgent: Multi-Task Agentic Network Traffic Analysis Made Practical
Abstract
Network traffic analysis is central to network security, spanning tasks from intrusion detection to encrypted traffic classification. Existing approaches either train task-specific models that generalize poorly or rely on costly traffic foundation models that still struggle under distribution shift. We present NetAgent, the first agentic framework for multi-task traffic analysis. Through a carefully designed agent loop, NetAgent supports complex task understanding, on-the-fly decomposition and orchestration, dynamic replanning, and long-horizon analysis, without task-specific training. It introduces five key designs: (1) knowledge-augmented workflow planning that maps attack knowledge to traffic features to bridge the semantic gap; (2) a comprehensive tool action space with 150+ verified tools extracted from 50+ published systems; (3) a unified code execution space for flexible action composition; (4) a three-tier memory for long-term knowledge consolidation; and (5) sandboxing and runtime repair for reliable execution. Across 9 major benchmarks, NetAgent outperforms all baselines (23 single-task and 5 multi-task) on nearly all tasks and generalizes substantially better to unseen traffic distribution (90.04% F1 vs. 2.74% and 3.04% for the best single-task and multi-task baselines) and under realistic background shift (4.85-point F1 drop vs. 74.88-point and 74.80-point drop for the best single-task and multi-task baselines). These results reveal that existing methods owe much of their reported success to overfitting dataset-specific patterns and degrade sharply in realistic network environments, while NetAgent's agentic design remains accurate, generalizable, and robust.
cs.CR / 8 / 2610.07489
Deep Defence on Wheels: A Dual Intrusion Detection System Architecture for Comprehensive In-Vehicle Network Security
Abstract
Increasing connectivity to the outside world and the lack of inbuilt security mechanisms have made legacy intra-vehicular networks vulnerable to cyberattacks. Initial research focused on maximising detection accuracy for known and unknown attacks, often using large, full-precision machine learning models. However, embedding IDSs into vehicular electronic systems also requires low detection latency, energy efficiency and minimal electronic control unit (ECU) resource overhead to process about 2,000 CAN frames/s. Lightweight models must balance accuracy with these deployment constraints. We propose a dual IDS framework comprising supervised and unsupervised learning-based solutions, each optimised for real-time, resource-constrained automotive platforms. A quantised LSTM-based IDS (QLSTM-IDS) achieves over 99.9% detection accuracy for DoS/Flooding, Fuzzing and Spoofing/Malfunction attacks using a single model architecture evaluated on two widely used datasets. The model is trained using the Brevitas quantisation-aware training library, transformed into a dataflow accelerator with custom blocks compatible with AMD's FINN toolchain, and synthesised using Vitis HLS. Complementing this, an 8-bit quantised convolutional autoencoder-based IDS (QCAE-IDS), quantised using AMD's Vitis-AI toolchain, detects previously unseen anomalies that alter CAN-ID sequence patterns with over 99% accuracy. An integration architecture enables both models to operate on a single FPGA, bridging the network interface IP and processing system to minimise software overhead. QLSTM-IDS achieves 0.25 ms inference latency and 0.8 mJ energy consumption per message, while QCAE-IDS achieves 0.42 ms and 1.1 mJ per block. Both solutions are deployed and evaluated on the ZCU104 SoC (XCZU7EV FPGA), demonstrating a flexible hardware/software co-design for real-time detection of known and unknown attacks on high-speed CAN buses.
cs.CR / 9 / 2610.07510
Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training
Abstract
Developers can build LLM agents by adapting third-party models through benign post-training. We study a supply-chain threat in which an attacker supplies a model with a backdoor: hidden behavior that produces malicious outputs when a particular input pattern appears. Focusing on software-engineering agents, we ask whether such backdoors survive the developer's supervised fine-tuning (SFT) and subsequent task-level reinforcement learning (RL). We observe that benign SFT substantially reduces attack success, but subsequent RL often preserves the residual behavior and sometimes even increases attack success. Our analysis of backdoor erosion during SFT identifies two factors that may favor survival: initial backdoor strength and gradient compatibility with benign training. These factors motivate PersistBD, which refines an already-backdoored model before release to improve its persistency through the benign post-training process. On Qwen2.5-Coder-7B, PersistBD raises attack success from 20% to 74% after SFT and from 20% to 76% after SFT-RL, while maintaining comparable benign task performance. Together, our results show that backdoors can remain active through benign post-training and that adversaries can deliberately increase their persistence. This highlights a supply-chain risk for AI developers and motivates stronger techniques for detecting and mitigating inherited backdoors when adapting third-party models into agents. Our code is available at https://github.com/uiuc-kang-lab/PersistBD.
cs.CR / 10 / 2610.07531
BVI: Lightweight, Data-Centric Blockchain-Based Verification of Identity Claims
Abstract
A person who answers an unexpected call claiming to come from a bank has no way to check the claim. Australian text messaging has labelled a message as unverified when the sender identifier is not registered since 1 July 2026, but a voice call still arrives with nothing behind it, and voice cloning has removed the last cue recipients relied on. We present BVI, which answers one question for the recipient: did the calling party prove it holds a credential a registered organisation issued for this call? BVI keeps the organisational record, its authorised channels and its revocation state on a public ledger in a directly queryable form, and puts all decision logic in the handset, which performs eleven checks, pays no transaction fee and holds no full-chain state. We define ten attack classes plus the case in which revocation cannot be resolved, compare them in a 10-by-4 coverage matrix spanning BBCA, the content digest, the detector and composed BVI, and run 27 automated scenario tests against isolated Hardhat fixture states of the BVI contract; all 27 pass. For media substitution and synthetic speech, the contract tests validate commitment and gating properties rather than claiming detector accuracy. The mandatory per-session anchor costs about 91.5k gas and is invariant in call duration and hop count; each participating intermediary additionally contributes one optional attestation write. In a 20,000-session-per-cell path simulation, at one quarter carrier participation BVI detects 20.6 per cent of randomly located in-path rewrites versus 0.1 per cent for the every-hop comparison, while identity and content evidence remain available even when path evidence is indeterminate. We also quantify the throughput ceiling of the anchoring design.
cs.CR / 11 / 2610.07532
Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge
Abstract
LLMs have advanced rapidly, raising growing concerns about their safety. Recent work has proposed approaches to detect and defend against attacks including defenses at decoding stage that leverage models' hidden states. However, existing decoding-stage defenses suffer from two limitations. First, they introduce a trade-off between safety and over-refusal, where strengthening safety degrades the model's helpfulness on benign queries. Second, many of these methods rely on internal hidden states and are thus restricted to specific architectures, incurring substantial overhead and limited generalization across models. To address these limitations, we introduce LADE (Latent Safety Signals for Defense), which leverages latent safety signals extracted by contrasting harmful and benign queries from dark knowledge (i.e., information carried by the output probability distribution beyond its argmax) in the first-token output probability distribution. Our key insight is that, beyond surface-level refusal tokens, the dark knowledge in the first-token distribution contains latent safety signals, defined as tokens whose probabilities differ sharply between harmful and benign queries. We show that these signals consistently align across LLMs, forming a model-agnostic direction that emerges from safety alignment. LADE consists of three components: (1) Extracting Latent Safety Signals from Dark Knowledge, which selects top-k safety-discriminative tokens from the first-token probability distribution; (2) Tokenizer Mapping, which maps these tokens across different tokenizers to enable model-agnostic application; and (3) kNN-based Discrimination, which classifies queries via a k-Nearest Neighbors search over the mapped tokens. Across diverse LLMs and benchmarks, LADE is robust against a wide range of jailbreak attacks and lowers attack success rates while maintaining a competitive safety-utility trade-off.
cs.CR / 12 / 2610.07635
CISB-Bench: An Auditable Source--IR Dataset of Compiler-Introduced Security Bugs
Abstract
Compiler-introduced security bugs (CISBs) arise when an optimization, lowering, or instrumentation decision changes a security-relevant property of the generated program. They are difficult to study because their evidence is distributed across issue reports, reduced tests, historical configurations, and compiler artifacts; a security-related report also does not imply that every associated reduction establishes a security-bearing compiler failure. We present CISB-Bench, an auditable dataset of 429 exact C-program rows mined from GCC and LLVM. Each row contains its C reduction, a standardized LLVM IR analysis bundle at -O0 through -O3, public provenance, a final binary label, and a primary mechanism or boundary annotation. Two reviewers independently labeled the fixed corpus, agreeing on 369 rows (86.0%, Cohen's kappa=0.662); the 60 disagreements were adjudicated. The final dataset comprises 280 CISBs and 149 hard non-CISB cases. The prediction task is to recover this reviewed exact-row label from the supplied artifacts; it is not a claim that standardized IR alone reproduces every historical compiler failure. We characterize the security mechanisms and evidence boundaries represented by the corpus, and demonstrate how its paired artifacts support source-only, IR-aware, and joint analyses. CISB-Bench provides a reusable, inspectable target for compiler-security mining and detection research.
cs.CR / 13 / 2610.07645
SkillPoison: Progressive Skill Poisoning via Successful Experiences
Abstract
Self-improving LLM agents increasingly distill successful experiences into persistent, reusable skills. Existing skill attack methods corrupt this learning pipeline by injecting malicious triggers, behaviors, or false facts into individual experiences or extracted skills. However, such attacks are easily detected, and the injected malicious behaviors often fail to accumulate as persistent skills. In this paper, we show that skill poisoning can arise even from verified successful experiences, without making any individual trajectory malicious. Based on this insight, we propose SkillPoison, a novel framework that progressively poisons skill via successful experiences. SkillPoison first constructs a set of successful experiences that reinforce a target behavior, and then removes the contextual conditions that constrain when the behavior applies. Rather than injecting malicious content, SkillPoison shapes how the skill extractor generalizes, allowing useful behavior to support task success while inducing harmful behavior when they are misapplied. Extensive experiments on three benchmarks show that SkillPoison achieves 95.71% attack success rates, while all injected experiences remain task-correct and pass verification and lexical inspection. Our code, data and implementation details are available for the community at https://github.com/DEEP-JLU/SkillPoison.
cs.CR / 14 / 2610.07691
PerSpectron: Detecting Invariant Footprints of Microarchitectural Attacks with Perceptron
Abstract
Detecting microarchitectural attacks is critical given their proliferation in recent years. Many of these attacks exhibit intrinsic behaviors essential to the nature of their operation, such as creating contention or misspeculation. This study systematically investigates the microarchitectural footprints of hardware-based attacks and shows how they can be detected and classified using an efficient hardware predictor. We present a methodology to use correlated microarchitectural statistics to design a hardware-based neural predictor capable of detecting and classifying microarchitectural attacks before data is leaked. Once a potential attack is detected, it can be proactively mitigated by triggering appropriate countermeasures. Our hardware-based detector, PerSpectron, uses perceptron learning to identify and classify attacks. Perceptron-based prediction has been successfully used in branch prediction and other hardware-based applications. PerSpectron has minimal performance overhead. The statistics being monitored have similar overhead to already existing performance monitoring counters. Additionally, PerSpectron operates outside the processor's critical paths, offering security without added computation delay. Our system achieves a usable detection rate for detecting attacks such as SpectreV1, SpectreV2, SpectreRSB, Meltdown, breakingKSLR, Flush+Flush, Flush+Reload, Prime+Probe as well as cache-attack calibration programs. We also believe that the large number of diverse microarchitectural features offers both evasion resilience and interpretability---features not present in previous hardware security detectors. We detect these attacks early enough to avoid any data leakage, unlike previous work that triggers countermeasures only after data has been exposed.
cs.CR / 15 / 2610.07771
What Response Marginals Miss: Adaptive Query Complexity of Functional Backdoor Recovery
Abstract
Functional backdoor recovery finds any trigger whose attack success rate is at least a given threshold rather than to recover the planted trigger. We study the minimum number of queries required for this task under label feedback which returns the predicted class label. We construct two finite families of victim models that have exactly the same attack success rate for every victim and trigger candidate. The distribution of returned labels for every query is also identical under a uniformly chosen victim. These families form an explicit counterexample that despite the matched quantities, their optimal adaptive query complexities are \(Θ(\log H)\) and \(Θ(H)\) where \(H\) is the number of possible victims. The difference arises because the same non-target labels are associated with different sets of victims, so successive queries eliminate possible victims at different rates. This separation disappears when the response is reduced to binary feedback, which reports only whether the target label is returned. The separation also persists for every fixed failure probability below one. Finally, we realize the same recovery problems with trained CIFAR-10 ResNet-18 classifiers and verify the predicted optimal query budgets. These results show that attack success rate and the distribution of returned labels for each query are insufficient to determine the query complexity of functional backdoor recovery.
cs.CR / 16 / 2610.07776
SCSM: A Traffic-Native Foundation Model for Transferable Website Fingerprinting
Abstract
Website fingerprinting infers the websites visited by users from encrypted traffic metadata. However, models trained under fixed collection conditions often degrade as website sets, collection times, network paths, browsers, or defenses change. Existing transferable attacks either rely on handcrafted perturbations of individual traces or adapt language-oriented architectures to traffic, limiting their ability to capture traffic-native semantics. To address these limitations, we propose SCSM, a traffic-native foundation model for transferable website fingerprinting. Specifically, SCSM constructs pairs of pretraining views from the same group of unlabeled traces through Segmentation, Combination, Scaling, and Masking. These operations produce diverse observable patterns while preserving the underlying packet events and local traffic dynamics of real trace fragments. The corresponding windowed traffic counting matrices serve as inputs for contrastive pretraining of a state-space encoder without website annotations. The pretrained encoder is then fine-tuned on a small labeled support set, and predictions are aggregated across temporal scales at inference. Experimental results demonstrate that SCSM surpasses the strongest baselines by 16.4\% in average top-3 accuracy over six temporal-drift tasks on GTT23 and by 13.8\% in average macro-F1 over four cross-domain datasets. The code and datasets will be made available at https://github.com/SJTU-dxw/WF-SCSM.
cs.CR / 17 / 2610.07820
Efficient and Implementation-Hardened RBLWE on Commodity Cortex-M Microcontrollers
Abstract
Efficient post-quantum cryptography on resource-constrained Internet-of-Things (IoT) devices requires implementations that exploit the target processor architecture while resisting practical implementation attacks. This paper presents an ISA-accelerated and implementation-hardened realization of Ring Binary Learning with Errors (RBLWE) encryption on a commodity ARM Cortex-M33 microcontroller. Packing four 8-bit polynomial coefficients into the byte lanes of a 32-bit register and processing them with SIMD-style instructions, together with a packed message codec, accelerates encryption and decryption, while a buffered hardware-TRNG entropy source drawn from the on-die Secure Element drives the key-generation gain. Together these give same-core cold-start speedups of $4.06\times$, $3.35\times$, and $3.01\times$ for key generation, encryption, and decryption over a scalar baseline, and $3.52\times$/$3.18\times$ lower encryption/decryption cycle counts than a reference Cortex-M0 implementation. On top of this accelerated core, we add four staged countermeasures: constant-time execution, fault hardening against zeroing, random-corruption, and instruction-skip faults, a Fujisaki-Okamoto (FO)-style CCA2 transform, and first-order shared (masked) CPA decryption, reporting each layer's cost individually. Binary-level inspection confirms these countermeasures survive compilation and identifies a compiler-induced masking flaw resolved with a hand-written assembly replacement. Dudect-style timing tests, debugger-assisted fault-injection campaigns, and component-level TVLA then provide implementation-level evidence for the staged protections. The results demonstrate a practical acceleration-security tradeoff for RBLWE on off-the-shelf microcontrollers and reusable architecture-aware techniques for lightweight post-quantum implementations.
cs.CR / 18 / 2610.07854
Lifecycle-Based Design and Evaluation of Real-Time Backup Triggers for Ransomware Damage Mitigation
Abstract
Ransomware continues to encrypt files during the interval between attack onset and detection. Real-time backups can mitigate this damage by preserving files before they are modified. The previously proposed Real-Time Open-File Backup System (ROFBS) triggers backups primarily on file-open events. However, the file lifecycle offers several candidate trigger points, including open, read, write, and rename operations. Triggering backups too early may create unnecessary backup files, whereas triggering them too late may allow ransomware writes to race with backup creation and prevent the preservation of clean file contents. Consequently, it remains unclear which trigger timing best balances recoverability and the number of backups created. In this study, we design and evaluate real-time backup triggers for mitigating ransomware damage from a file-lifecycle perspective. Specifically, we compare four strategies: Open-time backup, Read-time backup, Write-time backup, and Rename-time backup. We implement these strategies in an ROFBS-style prototype on XFS and evaluate them using five ransomware samples: Conti, Sodinokibi, AvosLocker, REvil, and HelloKitty. Our results clarify how trigger timing affects both damage mitigation and the number of backups created, providing design guidance for selecting effective triggers in real-time backup systems against ransomware.
cs.CR / 19 / 2610.07866
The Amplifier Effect: Human-Factor Risks of AI-Suggested Correlation and Auto-Propagation in Multi-Framework GRC Self-Assessment
Abstract
Multi-framework Governance, Risk and Compliance (GRC) platforms increasingly automate the link between an organisation's self-assessment answer and the compliance obligations that answer is said to satisfy. Cross-framework control mapping, AI-suggested question correlation, and automatic propagation of answers and evidence across correlated questions all serve the legitimate efficiency goal of reducing duplicate work for small and medium-sized enterprises under the EU Cyber Resilience Act, NIS2 and GDPR. The same mechanisms, however, amplify the consequences of any human-factor bias in a single answer: one optimistically-graded control, one rubber-stamped attestation, or one AI-drafted answer can be silently replicated as evidence of compliance with many obligations across multiple frameworks. We call this the amplifier effect: a platform-design property (coarse-grained attestation and un-gated propagation) rather than a failing of individual users. Using two EU-funded SME-facing GRC platforms, CYBERFORT and CYBER-BRIDGE, as examples, we (i) describe the amplification mechanism in concrete data-model terms, (ii) propose a six-dimension scoring framework for evaluating any GRC tool's exposure to the effect, (iii) instantiate the framework on a thirteen-tool comparison covering enterprise IRM, mid-market platforms, compliance-automation tools, and the two EU SME projects, and (iv) outline a measurement protocol that a consortium with access to production self-assessment data can run. The thirteen-tool comparison is a structured design assessment, not an empirical measurement of user behaviour. The EU SME platforms score lowest on the amplifier dimensions because their burden-reduction design deliberately trades sign-off granularity for throughput; we report this as a design trade-off, not a verdict on the platforms. Our contribution is the framing and the measurement protocol.
cs.CR / 20 / 2610.07870
Plug-and-Play Quantum-Resistant BLE Pairing for Medical Implants via NFC Out-of-Band
Abstract
Bluetooth Low Energy (BLE) pairing establishes the cryptographic foundation for secure device communication. However, mainstream man-in-the-middle (MITM)-resistant pairing methods, such as Numeric Comparison and Passkey Entry, require user interfaces that implantable medical devices (IMDs) inherently lack, making Near-Field Communication (NFC)-assisted Out-of-Band (OOB) pairing an attractive alternative. Existing NFC-assisted OOB schemes authenticate a classical BLE key exchange that remains vulnerable to quantum attacks, while the NFC channel itself is susceptible to eavesdropping and active injection under stronger threat models. To address these limitations, we propose a lightweight, plug-and-play NFC-based OOB pairing protocol that performs a post-quantum key encapsulation mechanism (KEM) entirely over the NFC channel, ensuring that no shared secret is transmitted over NFC while reducing long-range radio-frequency (RF) exposure during pairing. The proposed protocol requires no modifications to the BLE stack and inherently resists RF battery-depletion attacks. Evaluation on an IMD-class proof-of-concept testbed demonstrates that post-quantum OOB pairing is practical on resource-constrained devices, reducing projected battery life by less than 0.57\% for Kyber-1024 and 0.56\% for FireSABER under an operational usage model relative to the MITM-vulnerable Just Works method.
cs.CR / 21 / 2610.07873
Preparing an AI-Augmented SIEM for the EU Cyber Resilience Act: A Practitioner Case Study
Abstract
The EU Cyber Resilience Act (CRA), Regulation (EU) 2024/2847, makes product cybersecurity a lifecycle obligation for products with digital elements on the EU market: risk assessment, vulnerability handling, conformity documentation, and Article 14 incident- and vulnerability-reporting readiness must be operational before market placement. Small and medium-sized enterprises that build security products are doubly exposed, since their products are in scope while their customers expect them to be exemplary. This case study documents a CRA preparedness pilot for one such product, SEUXDR, an AI-augmented security monitoring product with a large-language-model active-response component, on the open-source CYBERFORT platform. We contribute a reproducible six-step recipe (Scope and Classify, Asset Registration, Produce Evidence, Map to CRA, Gap and Actions, Audit Pack), two end-to-end traceability threads, and a pilot snapshot tracing product risks through baseline and AI-specific controls and policies to CRA objectives. It offers practitioners a replicable starting point for translating CRA legal text into operational preparedness for incident response, vulnerability reporting, and conformity assessment.
cs.CR / 22 / 2610.07875
Don't Let One Lie Survive A Hundred Truths: A Selective Bayesian Trust Estimator for Collaborative Perception
Abstract
Collaborative perception (CP) enables connected vehicles to see beyond their own sensors but makes them dependent on messages they cannot independently verify. A compromised collaborator can surgically conceal a single safety-critical object or inject a non-existing one while correctly reporting many others. Existing Bayesian trust mechanisms pool agreement across objects, which, while effective against blatant untargeted attacks, either incurs high false-positive rates (FPR), or allows unrelated correct reports to dilute persistent attack evidence for stealthy single-object attackers. To address this problem, we propose SABER, a selective two-tier Bayesian trust estimator. The first tier maintains broad agent and object trust, preserving the ability to downweight benign but low-quality contributors. Cumulative-sum screening selects agent--object pairs with persistent omissions or unsupported reports for focused Bayesian assessment. The second tier checks these pairs against other agents' evidence and maintains a separate, reference-weighted Beta state for each. The lowest pair score constrains agent trust, preventing unrelated reports from diluting a targeted attack. We establish sufficient conditions for stronger attacker-side trust reductions with bounded additional benign false alarms at fixed thresholds. Compared with state-of-the-art CP defenses, SABER improves attack detection while reducing benign FPRs. On OPV2V, SABER improves defense ROC-AUC over MATE by up to 0.427 in late fusion and 0.337 in intermediate fusion. Against advanced intermediate-fusion data fabrication attacks, it increases detection rates over ROBOSAC and LUCIA by up to 96.40 and 67.07 percentage points, respectively, while reducing FPRs.
cs.CR / 23 / 2610.07931
Where does a rust speedup come from? Language and algorithm effects in sliding window threat scorer
Abstract
Rewriting a hot path from Python into Rust is a common way to speed up security analytics, and large speedups are routinely reported. A rewrite usually changes the language and the algorithm at once, so a single factor can credit the language with a gain that comes from a better algorithm. We study this on a sliding window threat scorer modelled on the traffic light calculator of the SentinelSphere platform. Five implementations, three in Python and two in Rust, produce bit identical scores, confirmed by a shared checksum, and we time them from 100 to one million events in a bounded and a burst regime. The language alone contributes between about 4 and 29 times. Replacing a per event rescan of the window by an incremental update contributes more than 11,000 times in the burst regime, so the end to end factor reaches about 195,000 times when the window keeps filling and levels off near 13,000 times when it does not. A power law fit to the timings reported for the original rewrite gives growth exponents of 1.85 for Python and 0.88 for Rust, the signature of an algorithmic difference. Performance claims for security tooling should therefore report the language and algorithm contributions separately.
cs.CR / 24 / 2610.07976
Quantifying the Privacy Posture of Operator-Side 5G/O-RAN Profiles
Abstract
Operator-side network profiles derived from 5G/ORAN traffic carry personal data such as ephemeral subscriber identifiers, slice-level KPIs, and control-plane signalling, and must be anonymised before release to a federated-learning aggregator, threat-intelligence exchange, or ML training pipeline. We study how much re-identification risk remains after standard operator-side anonymisation. We quantify privacy posture with k-anonymity, l-diversity and t-closeness, aggregate them into a composite Privacy-Posture Index (PPI), and measure residual re-identification across eight transformation configurations on internal PCAP captures and the public Idaho Labs 5GAD corpus, under a full-QI syntactic bound and two simulated adversaries. The evaluation is modest in scale, and we read its trends as indicative rather than definitive. Three findings emerge. Pseudonymisation alone leaves re-identification unchanged; material privacy gains arise when quasi-identifiers are coarsened through generalisation, optionally combined with suppression. A downstream classification task then shows that suppression-heavy releases retain majority-class utility but sacrifice much of their minority-class recall, a cost the aggregate metrics hide. Finally, the standard kmin-based PPI correlates only modestly with the disclosure bound and not at all with the partial-knowledge attack, whereas a mean-class-size variant PPI correlates strongly with all three disclosure/attack measures; we therefore read PPI as a regulator-facing summary, not a security bound. The profiles are produced by passive operator-side monitoring with rule-based DPI; our contribution is the privacy-quantification layer that computes these metrics, applies the transformation policy, and exposes both through an inspectable dashboard.
cs.CR / 25 / 2610.08066
Systematically Optimized CNN-Transformer with Focal Loss for Imbalanced Intrusion Detection on NSL-KDD
Abstract
Intrusion Detection Systems (IDS) struggle with imbalanced datasets like NSL-KDD, especially in detecting rare R2L and U2R attacks. This work describes a systematically optimized and explainable framework using a CNN-Transformer architecture to improve performance on highly imbalanced data. We decided to use XGBoost for feature selection and Focal Loss as the main mechanism for minority class learning. We used Optuna for end-to-end hyrate=0.00042, batch size=256, and Focal Loss γ), using a thoughtful data splitting process to prevent data leakage. Critically, our ablation studies pointed out that while Focal Loss (optimally at γ = 1.5) substantially enhanced minority recall, adding oversampling methods like SMOTE caused precision degradation via over-correction. Our final model, using only tuned Focal Loss, achieved 98.73% overall accuracy on the NSL-KDD test data, with significantly balanced F1-scores of 84.63% (R2L) and 69.66% (U2R) over baseline approaches. Additionally, we performed SHAP analysis to understand model predictions, understand salient features (e.g., service http, logged in), and explain the persistent U2R precision challenge due to feature overlap. This study presents a comprehensive pipeline for imbalanced IDS, providing methodological insights, and demonstrates that tuned Focal Loss can be a sufficient and effective strategy for class balancing.
cs.CR / 26 / 2610.08090
Explainable Rule Mining of IPv6 Extension-Header Presence Patterns from Paired-Vantage Captures
Abstract
IPv6 extension headers (EHs), such as fragmentation, segment routing, and in-situ telemetry, are operationally important yetwidely dropped in transit, and characterising their behaviour from packet captures is a recurring measurement problem. We ask whetheran explainable miner can recover human-readable rules of EH behaviour, and we contribute two reusable tools: a negative-control protocol that diagnoses whether a mined "temporal" network rule reflects genuine cross-packet dynamics or mere within-packetco-occurrence, and a sender-conditioned, per-family EH-retention measurement. Applying an interpretable temporal-logic rule miner to the JAMES paired-vantage dataset, we recover a portable Fragment-EH rule that the protocol reveals to be a within-packet,near-definitional co-occurrence rather than a temporal pattern, so the temporal-logic machinery does no work for this dominant rule;the retention measurement independently recovers the expected within-window ordering of EH observability. Our main result istherefore an honest, controlled negative finding, corroborated by executed decision-tree and large-language-model baselines: on theevaluated JAMES traces network-temporal structure does not carry the dominant Fragment-EH signal, and we supply the controls thatestablish when it would, validated on a synthetic positive control containing a genuine cross-packet dependency.
cs.CR / 27 / 2610.08097
When Tools Lie: Reliability of Mathematical Agents Under Corrupted Tool Feedback
Abstract
Mathematical problem solving often requires deterministic computational steps that agents delegate to tools and implicitly trust. Yet tools can fail silently, returning plausible but incorrect results. How well can agents detect and correct corrupted tool call outputs? We study this through a controlled corruption framework where a hidden interceptor replaces tool call results with plausible incorrect information on targeted problems. We evaluate agents across 31 problems under four verification designs including no verification (baseline), mandatory same-context reflection, optional fresh-context verification, and optional structural verification. Without verification, corruption causes dramatic accuracy loss, from 100% down to 72.4%. Mandatory reflection fully recovers this performance to 100%. Optional verification improves accuracy only when models actively invoke it. Our results show that checking frequency is strongly associated with robustness differences, while unequal invocation prevents a controlled comparison of verifier quality. A supporting recovery experiment shows that full problem restart succeeds in 100% of cases after explicit detection. These findings demonstrate that verifier availability and verification policy are separate components of mathematical-agent reliability. Mandatory policies enforce verification while optional policies depend on the model's own choice to invoke it.
cs.CR / 28 / 2610.08098
Surviving the Router: Optimizing Skill Injections for Retrieval and Execution
Abstract
AI agents increasingly rely on modular third-party "skills" that are dynamically selected by skill routers to execute complex tasks. While recent studies highlight the threat of prompt injections embedded in these skills, existing evaluations often assume settings where the malicious skill is already selected for execution. We show that this assumption can substantially overestimate attack success. In realistic multi-skill environments, injected skills must first compete for retrieval, reducing the effective attack success rate (ASR) of existing injections by 87-97%. To address this limitation, we introduce CORSA (Cluster Optimization for Router-Aware Skill Attacks), a router-aware attack that optimizes skill injections for both retrieval and execution across clusters of related tasks. We evaluate skill injection attacks under router-managed multi-skill settings by extending the benchmark introduced by SkillRouter with eight malicious payload categories. CORSA uses successive optimization stages to first improve retrieval and then optimize end-to-end attack success, while we evaluate user utility and injection naturalism separately. Our experiments show that CORSA substantially improves both retrieval and end-to-end attack success over existing skill injections while preserving user utility, and that the resulting attacks transfer across different router architectures and LLM backbones.
cs.CR / 29 / 2610.08137
Rethinking Visual Provenance: Detection and Watermarking Across Direct Visual Generation and LLM-Driven Code Rendering
Abstract
AI systems create images and videos with image/video generation models or by writing code and graphics descriptions that are then rendered. These routes can produce similar visible artifacts but expose different representations, intervention points, and provenance evidence. We develop a production-centered framework that compares detection and watermarking across both routes. An explicit verification specification distinguishes passive inference, message recovery, and authenticated provenance. We organize image, video, source-code, and rendering-aware watermarks by production stage. We examine the different requirements of generated images and video, plots and SVG, programmable video, and agent-composed workflows. Documented Claude, OpenAI, and rendering-tool interfaces connect the framework to concrete systems. We pose ten scoped research questions on identifiability, observability, fair comparison across stages, recoverable payload, reconstruction, synchronization, composition, hybrid local contribution, and private production-event authentication. The result is a conceptual research agenda grounded in published methods, inspected interfaces, and elementary boundary examples. It reports no experiments and claims no new theorems; its appendix results are elementary calculations, and documentation and source inspection establish interfaces, not empirical robustness.
cs.CR / 30 / 2610.08174
FBAN: A Fully Homomorphic Encryption Compatible Bottleneck Attention Network for Privacy-Preserving Behavioral Authentication
Abstract
Continuous authentication (CA) strengthens session security by repeatedly verifying the user during device interaction, yet it inherently relies on highly sensitive behavioral traces (e.g., fine-grained touch dynamics) that are often outsourced to cloud/edge services for scalable inference. This raises a fundamental privacy-in-use challenge: protecting behavioral features during computation, not only in transit or at rest. Fully homomorphic encryption (FHE) offers a principled solution, but deploying modern CA models under FHE remains difficult due to non-linearities and attention-style operations that incur high ciphertext cost. We propose FBAN, a TFHE-compatible Bottleneck Attention Network and an end-to-end encrypted CA framework. FBAN is designed for integer-only execution via a two-stage pipeline (floating-point pretraining followed by quantization-aware training) and is compiled into TFHE circuits for homomorphic inference. We further specify a client-server protocol with session-bound blinding and decrypt-and-return verification, enabling the server to authenticate users without observing raw behavioral features. We provide a cryptographic security analysis against an honest-but-curious server under TFHE IND-CPA security, and formalize resistance to replay and impersonation without the TFHE secret key. Experiments on two public touchscreen datasets demonstrate that FBAN achieves strong authentication utility under encrypted inference while maintaining a lightweight model footprint, with TFHE parameters instantiated at $\geq 128$-bit security.
cs.CR / 31 / 2610.08255
HE-OFT: Privacy-Preserving One-Shot Federated Fine-Tuning under Homomorphic Encryption
Abstract
Many organizations adapt large pretrained models to their own tasks by fine-tuning on private data. Several of these parties often hold data for the same task and wish to fine-tune a model together without pooling that data. Federated learning (FL) enables joint fine-tuning, but reconstruction attacks on shared intermediate values (the model or its gradients) remain a privacy risk. A one-shot protocol that exchanges one encrypted contribution exposes no intermediate value. Such a protocol still gives the trained model to every participant, which is not permitted where the model is a regulated or proprietary asset. We present HE-OFT, the first cryptographically secure one-shot federated fine-tuning protocol in which no party receives the trained model. Each client fine-tunes a low-rank adapter and a classifier head on a frozen public backbone and keeps the adapter. The client uploads one encrypted head displacement, which the server combines under multiparty CKKS and never decrypts. A quorum of clients returns only the predicted label to the querier. On four text classification tasks and one vision task, HE-OFT reaches 61 to 79 per cent accuracy, against 20 to 48 per cent for a client training alone. HE-OFT keeps 85 to 96 per cent of the accuracy of a disclosed model. A test-time query takes 443.1 to 1713.1 s on one core, or 56.1 to 255.1 s with level restoration on a GPU. Restoring levels at the server cuts the traffic per query from up to 1.6 GiB to 13.5 MiB.
cs.CR / 32 / 2610.08262
Contextual Chain: Lightweight Continuity Authentication for Intermittently Connected Devices
Abstract
Can authentication make memory, rather than computational hardness, the attacker's bottleneck? Contextual Chain is a lightweight continuity protocol for intermittently connected devices that share evolving physical or operational context. An honest device follows one realized history, updating a compact accumulator and fixed hash-based readiness lanes; outages cause pause or bounded rollback, not branch search. After the epoch is frozen, a fresh challenge selects one lane under a short deadline. An outsider that missed context may therefore need to prepare for many mature histories before learning which one will be tested. In the standard random-oracle model, a causal counting theorem lower-bounds the deadline-accessible retained state required for a target success probability against arbitrary nonlinear preselection encoding and adaptive post-selection queries, accounting for sequential depth, candidate queries, and cross-target protected information obtained online. Honest readiness memory remains fixed and independent of the number of plausible histories. Contextual Chain thus converts shared-experience uncertainty into a tunable preparation requirement without transferring combinatorial complexity to lightweight devices.
cs.CR / 33 / 2610.08301
Zeppelin: Client-Side BFV Encryption and Decryption for Helium-Powered Microcontrollers
Abstract
The growth of the Internet of Things (IoT) has raised concerns over the privacy of data collected by resource-constrained sensing devices. Homomorphic encryption (HE) addresses this by letting a device encrypt its data once and an untrusted cloud server compute on the ciphertext without seeing the values. In practice, HE's memory and computational cost have kept it out of reach of microcontroller-class devices. Prior work, SEAL-Embedded, made this feasible using CKKS, but left three gaps: its arithmetic is entirely scalar, even on hardware with a vector instruction set; it never decrypts on the device, so the client cannot consume a result; and it does not explore BFV, whose encoding uses only integer arithmetic and decrypts exactly. We present Zeppelin, the first HE library to use an embedded vector instruction set, the first to support both encryption and decryption on an embedded device, and the first BFV implementation on MCU-class hardware. Zeppelin vectorizes the number-theoretic transform, HE's main bottleneck, for ARM's Helium extension, and includes a decryption procedure that avoids large-integer arithmetic and timing leakage of the secret key. A server-side adapter converts Zeppelin's ciphertexts into a format compatible with Microsoft SEAL, so the device handles encryption and decryption while the server performs all homomorphic computation. On an STM32N6 MCU with an ARM Cortex-M55, Zeppelin encodes and encrypts 4096 packed values in 4.76 ms (seeded symmetric, 128-bit security per the HE standard) and decrypts and decodes the result in 5.38 ms, using under 500 KB of RAM, with the vectorized NTT engine ${\sim}2.7\times$ faster than an equivalent scalar implementation on the same core.
cs.CR / 34 / 2610.08464
Federated Bayesian Surveillance of Mechanical Thrombectomy Adverse Events: A Population Risk Layer for Surgical Digital Twins
Abstract
Learned surgical simulators and world models can roll out plausible procedural futures, but they carry no grounded estimate of how often interventional devices actually harm patients. We propose treating population-scale adverse-event surveillance as a distinct belief layer of the surgical digital twin, and we evaluate a federated Bayesian protocol for learning it under formal privacy guarantees. Each site holds per-class Gamma-Poisson posteriors over adverse-event rates and exchanges only Rényi-differentially-private natural-parameter updates. We benchmark on the complete FDA MAUDE cohort for thrombus-retrieval catheters (product code NRY): 8,617 reports, of which 6,491 are classified by transparent keyword rules into five thrombectomy complication classes and partitioned across $K=8$ manufacturer sites. At a matched privacy budget of $(\\varepsilon \\approx 2.09, \δ= 10^{-5})$, the conjugate protocol attains a held-out Poisson score of -5.78 per test event versus -26.58 for FedAvg with differential privacy. The non-private federated model also outperforms centralized pooling (+3.19 vs +2.93), evidence that manufacturer-specific complication profiles are real and that federation preserves them. Because MAUDE lacks procedure denominators, outputs are relative rate orderings rather than absolute risks, and we report all privacy-utility operating points.
cs.CR / 35 / 2610.08571
RAG-PIBench: A Leakage-Aware Benchmark for Prompt-Injection Detection in Trustworthy RAG Systems
Abstract
Retrieval-Augmented Generation (RAG) systems are vulnerable to prompt-injection attacks embedded in retrieved content. We introduce RAG-PIBench, a benchmark for RAG-style prompt-injection detection containing 4,876 contextual examples across frozen train, validation, and protected-test splits. Using a leakage-aware construction pipeline and strict evaluation protocol, we compare keyword-based, semantic-reference, TF-IDF, and transformer-based detectors. DistilBERT achieves the best protected-test performance (F1 = 0.896, PR-AUC = 0.968), while TF-IDF SVM and logistic regression remain competitive. Our results demonstrate the value of leakage-aware benchmark design and strong sparse baselines for reliable prompt-injection detection in RAG systems.
cs.CR / 36 / 2610.08668
Semantic Behavioral Watermarking: Paraphrase-Robust and Forgery-Resistant Provenance for LLM Agents
Abstract
Behavioral watermarking embeds an owner identifier in an LLM agent's high-level action choices, giving provenance without touching output tokens. Prior agent watermarks break in two ways. First, all three prior schemes bind the watermark to the exact action symbol, so renaming a tool desynchronizes decoding even when the observation is untouched; in AgentMark's own robustness test, paraphrasing the observation alone drops bit-recovery to 16.8%. Second, every prior agent watermark studies only removal: none asks whether an adversary can forge a trajectory that verifies as someone else's, a question answered affirmatively for text watermarks (Jovanović et al., 2024). We present Semantic Behavioral Watermarking (SBW): watermarking over semantic action clusters under history conditioning, with the public-cluster bin replaced by keyed collision-resistant binning whose fresh-bucket assignment is provably unpredictable in the random-oracle model. Across five agent models (3B-14B, four vendors) and three encoders the ordering holds on both benchmarks: on ToolBench (600 trajectories per model) detection under rewriting is 0.49-0.66 for cluster-level versus 0.05-0.17 for exact-symbol at a permutation-calibrated 1% FPR, at 72-83% choice agreement against 22-27% for logit biasing; on ALFWorld (100 episodes per model) it is 0.92-0.97 versus 0.00-0.01. Keyed binning takes adaptive forgery from 100% to the false-positive floor at the primary operating point (bge, r=64). We also mark the boundary that guarantee does not cover: when the adversary copies the victim's own steps, shuffled splicing is neutralized (0.000 on Qwen2.5-3B) but chained replay remains at 0.76-0.98 across the five models, reported as open. Paraphrase robustness costs about half of the per-step watermark capacity. Code is available at https://anonymous.4open.science/r/SBW-Agent-Watermark.
cs.CR / 37 / 2610.08771
Mission-Aware Attestation Envelopes for Time-Critical Autonomous Action: A Hardware-in-the-Loop V2I Study
Abstract
An autonomous system that asks for a privileged physical action is usually gated on integrity evidence: a platform proves what it is running, and the request is granted or refused on that basis. Such a gate is normally treated as a predicate, yet the evidence behind it has an age, the decision that consumes it has a latency, and the physical system that waits for it has a deadline. We formulate mission-aware attestation as a runtime assurance contract that holds only when integrity is valid, the evidence is fresh enough, and the decision completes inside a budget derived from the current physical state. The contract yields four operational outcomes where a binary gate yields two, separating a refusal caused by tampering from one caused by stale evidence and from one caused by a late decision. We evaluate it on a hardware-in-the-loop vehicle-to-infrastructure platform: a driving simulator supplies the physical state and the authorisation deadline, while a microcontroller on-board unit and a TPM-backed roadside unit running Linux integrity measurement supply the assurance evidence. A security-blind model admits the whole operating space and a hardware-informed one three quarters of it, and every point it refuses fails the freshness margin rather than the response margin. Moving the attestation interval across the range the verifier permits costs about as much as a fivefold scaling of the latency distribution, and the interval is directly configurable, which makes it the immediately actionable deployment parameter. If the freshness bound does not exceed the authorisation budget, every late decision is also stale and lateness becomes unobservable, so the attestation interval and the freshness bound cannot be chosen from security requirements alone.
cs.CR / 38 / 2610.08410
Decoy and disclosure radii of invariant shape descriptors
Abstract
A recognizer that compares rotation-invariant descriptors sees a surface only up to the fiber of the descriptor. We measure this fiber by its radius in the orbit distance from the enrolled surface. A large radius admits decoys, that is, distant shapes that pass the matcher. A small radius discloses the enrolled shape to anyone who captures the stored value. For star-shaped surfaces truncated to spherical harmonics of degree at most $L$, with $n$ coefficients, a descriptor of generic rank $r$ has generic fibers of dimension $n-3-r$ modulo rotations. The standard pool of band powers, even bispectra, and three invariants of the degree-three band therefore admits decoy families of dimension $5$, $13$, $20$ at $L=4,6,8$. Its rank first reaches $n-3$ at $L=16$, and a mirror decoy remains at every $L$. The odd bispectra remove the mirror decoy generically for $L \geq 4$. Yet at fixed mean radius the same pool determines the enclosed volume exactly, and it does not determine whether a surface meets a clearance requirement. We certify two cases by exact and interval arithmetic. At $L=6$ a decoy matches all $32$ invariants to relative precision $2 \cdot 10^{-18}$ at orbit distance at least $0.87$ times the norm of the enrolled tuple. For the radar shape model of asteroid (101955) Bennu, the pool recovers the modeled volume, misses the handedness, and leaves the keep-out radius uncertain by more than $7 \, \mathrm{m}$.
cs.CR / 39 / 2610.08417
Ariadne's Thread of LipSync: Unraveling Forgeries via Inconsistency between Lip Motions and Head Poses
Abstract
Recent advances in LipSync generation technology have led to the creation of highly realistic videos, posing severe societal risks. However, existing defense strategies struggle against LipSync forgeries, as advanced LipSync generation methods not only achieve better lip synchronization but also eliminate visual artifacts. An important reason is that they overlook an inherent biological coupling between lip movements and head poses in natural speech videos. In this paper, we propose LipDA, a novel framework for joint LipSync Detection and Attribution, which takes advantage of the inconsistency between head and lip. For detection, the framework learns to quantify this discrepancy by contrasting lip and pose features from authentic versus forged videos. For attribution, our method is designed to capture the unique temporal dynamics and audio-visual synchronization patterns that act as the fingerprint of models, enabling source tracing. We conduct extensive experiments on two challenging LipSync datasets as well as our own proposed large-scale and multi-generator dataset. LipDA achieves over 97\% AUC in detection and 97.5\% accuracy in model attribution, significantly outperforming existing methods. Code and the proposed LipSync-A dataset are available at https://github.com/AnsonShe/LipDA.
cs.CR / 40 / 2610.07759
When Can Stateless Recovery Defeat Byzantine Quorum Safety? A Tight Normal Form for Single-Step BFT
Abstract
Byzantine quorum safety relies on correct replicas refusing to sign conflicting values. A replica that loses its protocol state during recovery but retains its identity and signing key may forget an earlier vote. We study certificates formed by matching signed votes from at least $q$ of $n$ replicas, assuming that each correct replica avoids conflicting votes between recoveries. If two conflicting certificates form, their overlap has size between $2q-n$ and $b+c$, where $b$ counts Byzantine replicas and $c$ counts correct identities that recovered during the execution considered. Our main result decomposes the slack $b+c-(2q-n)$ into four nonnegative counts: extra signers in the first certificate, extra signers in the second, identities in neither certificate, and Byzantine or recovering identities outside their overlap. Zero slack forces an exact signer partition. With $n=3f+1$ replicas, threshold $q=2f+1$, at most $f$ Byzantine replicas, and exactly one correct recovery event, any conflicting pair forces exactly $f$ Byzantine replicas, all in the overlap together with the recovered replica; each certificate has a disjoint side of $f$ correct replicas. A minimal protocol attains this form. We distinguish certificate formation from acceptance, give a sufficient check using configured fault and recovery caps, and explain why durable vote records written before signature release prevent the conflict.
cs.CR / 41 / 2610.07251
Pauli Error Composition Determines the Multipartite Advantage in Conference Key Agreement
Abstract
GHZ-based conference key agreement distributes a shared key to $N$ parties in one network use, whereas pairwise BB84 needs $N-1$. The survival of that advantage has been studied against party number, loss, and topology, but almost always with the channel fixed at depolarizing. Simulating both protocols exactly for $N = 3, 4, 5$ under dephasing, amplitude damping, and depolarizing noise, we find the crossover is set by the channel's Pauli error composition, not its identity: normalized to average infidelity, amplitude damping and depolarizing agree to within about 1\% at each group size, while dephasing can tolerate roughly $2.4\times$ more at $N = 3$. We derive a closed-form rate for arbitrary per-party Pauli noise and use it to prove that concentrating a fixed noise budget on one arm is strictly worse than spreading it, and that giving the dealer role to the cleanest arm improves the rate by roughly 18\%. One structural fact drives every result: the phase-error term is a product over all $N$ arms, while the bit-error term stays local to a pair.
cs.CR / 42 / 2610.08460
One-Shot Private Confidence Regions via Resampling
Abstract
We propose a simple framework for constructing differentially private confidence regions \textit{in one shot}, i.e., by adding noise only to the final resampling quantile instead of privatizing the estimator computed on each resample. The cost of privacy of our procedure is only logarithmic in the number of resamples $B$ under with-replacement ($m$-out-of-$n$) sampling and independent of $B$ under without replacement sampling (subsampling), avoiding the $\sqrt{B}$ factor that arises in previous works. We provide nonasymptotic Gaussian Differential Privacy (GDP) and utility guarantees for both subsampling and $m$-out-of-$n$ resampling, covering mean-like estimators with small global sensitivity as well as estimators admitting efficiently computable smooth sensitivity bounds, including quantiles and degenerate U-statistics. This allows us to also obtain private confidence regions for degenerate U-statistics where the private error is much smaller than the non-private error. In all, we provide a toolbox for widely applicable DP uncertainty quantification procedures under popular resampling strategies while avoiding the computational and privacy costs of privatizing many intermediate resample statistics.