Daily Research Digest
arXiv Papers
2026-09-30
837
Papers
8
Categories
186
Translated
收藏清单 0
精选 · Favorites
186
cs.AI / 1 / 2609.36043
SAGE: A Statistical Acceptance Gate for Self-Evolving Agents
SAGE:一种用于自演化智能体的统计接受门控
large language model
大语言模型相关
Abstract
Large Language Model (LLM)-based agents increasingly self-evolve by editing a persistent skill document that encodes their workflow, tool-use rules, and decision logic. This loop has two steps, an optimizer that proposes a candidate edit and a gate that accepts or rejects it. Prior work has concentrated on the optimizer, while the gate still follows a naive rule that keeps any edit which improves an aggregate validation score. We show that this rule fails in two ways. First, it admits permanent regressions, since an edit can raise the average while breaking items the skill already solves. Second, it is vulnerable to the Optimizer's Curse, since the best observed score on a finite and noisy validation set is upward biased. To solve the above two limitations, we propose a statistical acceptance gate for self-evolving agents (SAGE). Compared with previous work, SAGE has two contributions. First, SAGE proposes a per-item paired comparison that evaluates the current skill and the edited skill on identical validation items, which exposes regressions that an aggregate score hides and penalizes them asymmetrically. Second, SAGE also employs a one-sided paired test that commits an edit only when its wins are statistically reliable against its losses, and it abstains otherwise. SAGE is a conservative refinement of the standard gate that recovers the baseline exactly at a boundary setting. It commits only a subset of the baseline's edits, filtering out those whose gains are unreliable or purchased by breaking already-solved items. Across five benchmarks and four backbone LLMs under an equal-budget protocol, SAGE lowers the regression rate in 19 of 20 settings and matches the baseline in the remaining one, for example from 36.5% to 0% on LiveMath and from 42.8% to 0% on OfficeQA with DeepSeek-V4. SAGE also attains the highest final score in all 20 settings, raising LiveMath from 34.15 to 48.78.
Chinese Translation
基于大语言模型(LLM)的智能体越来越多地通过编辑一份持久化的技能文档来自我演化,该文档编码了它们的工作流程、工具使用规则和决策逻辑。这个循环包含两个步骤:一个优化器提出候选编辑,以及一个门控接受或拒绝该编辑。以往的工作集中在优化器上,而门控仍然遵循一条朴素的规则,即保留任何能够提升聚合验证分数的编辑。我们表明,这条规则会在两个方面失效。首先,它会引入永久性退化,因为一次编辑可能在提高平均值的同时破坏该技能已经解决的条目。其次,它容易受到优化器诅咒(Optimizer's Curse)的影响,因为在有限且有噪声的验证集上观察到的最佳分数存在向上偏差。为解决上述两个局限,我们提出了一种用于自演化智能体的统计接受门控(SAGE)。与以往工作相比,SAGE 有两个贡献。首先,SAGE 提出了一种逐条目配对比较,它在相同的验证条目上评估当前技能和编辑后的技能,从而暴露出聚合分数所掩盖的退化,并以非对称方式对其施加惩罚。其次,SAGE 还采用单侧配对检验,仅当一次编辑的获胜相对于其失败在统计上可靠时才提交该编辑,否则弃权。SAGE 是对标准门控的一种保守改进,在边界设置下能够精确恢复基线。它只提交基线编辑的一个子集,过滤掉那些收益不可靠或通过破坏已解决条目而换来的编辑。在等预算协议下,跨五个基准和四个骨干 LLM,SAGE 在 20 个设置中的 19 个降低了退化率,并在其余 1 个设置中与基线持平,例如在使用 DeepSeek-V4 时,LiveMath 上从 36.5% 降至 0%,OfficeQA 上从 42.8% 降至 0%。SAGE 还在全部 20 个设置中取得了最高最终分数,将 LiveMath 从 34.15 提高到 48.78。
cs.AI / 2 / 2609.36057
Mirror-Score: Calibrated, Inference-only Scoring Exposes the Limits of Sequence-compatibility Ranking in D-peptide Design
Mirror-Score:校准的、仅推理的评分揭示 D-肽设计中序列兼容性排序的局限
diffusion
扩散模型相关
Abstract
D-peptides combine protease resistance with high target specificity, but computational design of D-peptide binders remains immature. Mirror-Peptidizer introduced an in silico mirror-image screening pipeline using target reflection, backbone generation, and ProteinMPNN sequence design, but its raw ProteinMPNN negative log-likelihood (NLL) ranking was not validated against measured affinities, and only 4 of 9 tested MDM2 designs bound detectably. We introduce Mirror-Score, a calibrated, inference-only scoring framework for heterochiral D-peptide/L-protein complexes, and a public benchmark of 31 crystal complexes across four target families, including 18 with literature-verified affinities. Raw ProteinMPNN NLL is not a valid affinity ranker: its pooled Spearman correlation with affinity is 0.19, and correlations reverse between MDM2/CHIP (+0.62) and gp41 (-0.70). We therefore evaluate Boltz-2 mirror-space cofolding confidence. For the complete viral-entry family (7 structures representing 3 peptides), interface predicted local distance difference test (pLDDT) achieves structure-level leave-one-out Spearman rho = 0.90 (p = 0.006) and correctly orders all three peptides by affinity, whereas NLL fails (structure-level rho = 0.18). Because only three independent chemotypes are represented, this result indicates directional consistency rather than a statistically validated predictor. Cross-family calibration does not transfer at current sample sizes, supporting family-matched calibration as the practical deployment mode. We also specify a prospective design protocol for the antimicrobial-resistance targets LasR and LecB from Pseudomonas aeruginosa, including mirrored structures, ligand-derived hotspot maps, diffusion-model-ready inputs, and Mirror-Score ranking. Code, benchmark data, structures, and analysis scripts are openly available at https://github.com/Jiadalee/Mirror-Score.
Chinese Translation
D-肽将蛋白酶抗性与高靶标特异性结合在一起,但 D-肽结合体的计算设计仍然不成熟。Mirror-Peptidizer 引入了一个使用靶标反射、骨架生成和 ProteinMPNN 序列设计的计算镜像筛选流程,但其原始 ProteinMPNN 负对数似然 (NLL) 排序未依据实测亲和力进行验证,并且 9 个受测 MDM2 设计中只有 4 个能可检测地结合。我们提出 Mirror-Score,一个用于异手性 D-肽/L-蛋白复合物的校准的、仅推理的评分框架,以及一个涵盖四个靶标家族的 31 个晶体复合物的公开基准,其中 18 个具有文献验证的亲和力。原始 ProteinMPNN NLL 不是有效的亲和力排序器:其与亲和力的汇总 Spearman 相关性为 0.19,且相关性在 MDM2/CHIP (+0.62) 与 gp41 (-0.70) 之间发生反转。因此,我们评估 Boltz-2 镜像空间共折叠置信度。对于完整的病毒进入家族(代表 3 条肽的 7 个结构),界面预测局部距离差异检验 (pLDDT) 达到结构级留一法 Spearman rho = 0.90 (p = 0.006),并正确按亲和力对全部三条肽进行排序,而 NLL 失败(结构级 rho = 0.18)。由于仅代表了三个独立化学型,该结果指示的是方向一致性,而不是统计上得到验证的预测器。跨家族校准在当前样本量下无法迁移,支持家族匹配校准作为实际部署模式。我们还为来自铜绿假单胞菌的抗微生物耐药靶标 LasR 和 LecB 指定了一个前瞻性设计方案,包括镜像结构、配体衍生热点图、扩散模型就绪输入和 Mirror-Score 排序。代码、基准数据、结构和分析脚本可在 https://github.com/Jiadalee/Mirror-Score 公开获取。
cs.AI / 3 / 2609.36159
Principled Thoughts for Latent Recursive LLM Systems
面向潜在递归 LLM 系统的原则化思维
large language model
大语言模型相关
Abstract
Large language models can reason in continuous space instead of decoded text, by recurring on their own hidden states or by passing those states between agents, while training supervises only the Cross-Entropy (CE) of the final decoded answer and does not constrain the thought. Theoretical and empirical analyses establish and confirm four failures of CE-only training that lead to a lower probability of the correct answer such as collapsing thoughts across distinct questions and retaining irrelevant information. We introduce REST (REpresentation-Supervised Thoughts), a training objective that turns four properties of a valid thought representation (causality, minimality, separability, and stability) into differentiable losses added to CE. We instantiate it in latent single-agent and multi-agent systems, without architectural changes or added parameters at inference. Across 7 benchmarks spanning mathematics, science, medicine, and code generation, with the same training data, compute, and latent budget, REST increases accuracy over CE-only training across agent settings and model sizes by up to 7.5 percentage points and convergence on a final answer by 30\%. Furthermore, REST thoughts encode more of what is required to achieve the correct answer, and decoding them better recovers the intended output of the agent, which makes latent communication easier to interpret. Project Website: https://fard-lab.github.io/REST
Chinese Translation
大语言模型可以在连续空间而非解码文本中进行推理,方法是对其自身隐藏状态进行循环复用,或在智能体之间传递这些状态,而训练仅监督最终解码答案的交叉熵(CE),并不约束思维。理论与实证分析确立并证实了仅 CE 训练的四种失败,这些失败会导致正确答案的概率降低,例如在不同问题之间使思维坍缩以及保留无关信息。我们提出 REST(REpresentation-Supervised Thoughts,表征监督思维),一种训练目标,它将有效思维表征的四种性质(因果性、最小性、可分离性和稳定性)转化为添加到 CE 上的可微损失。我们在潜在单智能体和多智能体系统中实例化它,无需在推理时进行架构改变或增加参数。在涵盖数学、科学、医学和代码生成的 7 个基准上,在相同的训练数据、计算量和潜在预算下,REST 在不同智能体设置和模型规模上将相对于仅 CE 训练的准确率最多提高 7.5 个百分点,并将最终答案的收敛率提高 30\%。此外,REST 思维编码了更多实现正确答案所需的信息,并且解码它们能更好地恢复智能体的预期输出,这使潜在通信更易于解释。项目网站:https://fard-lab.github.io/REST
cs.AI / 4 / 2609.36245
CoRe: Co-Evolving Reward Models for Mitigating Latent Reward Hacking in Video Diffusion Models
CoRe:用于缓解视频扩散模型中潜在奖励黑客的协同进化奖励模型
diffusion
扩散模型相关
Abstract
Latent reward models (LRMs) enable efficient alignment of video diffusion models by scoring intermediate states directly in latent space. However, we find that optimizing against a fixed latent reward rapidly leads to latent reward hacking: the predicted reward stays high while perceptual and motion quality deteriorate. Our analysis identifies distributional escape as the central cause: within a few hundred updates, the generator moves beyond the reward model's training support, where its scores no longer reflect video quality. Based on this insight, we introduce CoRe, a co-evolving reward framework that treats latent-space alignment as a dynamic interaction between the generator and the reward model. Rather than optimizing against a stationary proxy, CoRe continually refits the reward model on the generator's current samples while anchoring it to real-video preferences, so the generator cannot gain reward by drifting away from the data. On Wan2.1-T2V-1.3B, experiments show that CoRe consistently improves generation quality over both the pretrained model and prior alignment methods, while avoiding the quality collapse of fixed-reward optimization.
Chinese Translation
潜在奖励模型(LRMs)通过在潜在空间中直接对中间状态进行打分,实现了视频扩散模型的高效对齐。然而,我们发现,针对固定的潜在奖励进行优化会迅速导致潜在奖励黑客:预测奖励始终保持高位,而感知质量和运动质量却不断恶化。我们的分析将分布逃逸确定为其中的核心原因:在数百次更新之内,生成器就会越出奖励模型的训练支撑集,此时其评分不再反映视频质量。基于这一洞见,我们提出了 CoRe,一个协同进化的奖励框架,它将潜在空间对齐视为生成器与奖励模型之间的动态交互。CoRe 并非针对一个静止的代理进行优化,而是持续地在生成器当前的样本上重新拟合奖励模型,同时将其锚定于真实视频偏好,从而使生成器无法通过偏离数据来获取奖励。在 Wan2.1-T2V-1.3B 上,实验表明 CoRe 相比预训练模型和此前的对齐方法都能持续提升生成质量,同时避免了固定奖励优化所导致的质量崩溃。
cs.AI / 5 / 2609.36264
OTROPE: Optimal Transport-based Robust Off-policy Evaluation for Large Language Models
OTROPE:面向大型语言模型的基于最优传输的鲁棒离策略评估
large language model
大语言模型相关
Abstract
Reliable evaluation of large language models (LLMs) is essential for their development and deployment, yet is often costly, risky, and difficult to perform safely online. We study off-policy evaluation for LLMs, where limited human-labeled data from a behavior model are used to evaluate a newer target LLM. This setting is challenging because labels are scarce, behavior--target distribution shift is common, and response likelihoods are often unavailable for black-box LLMs. We propose the Optimal Transport-based Robust Off-Policy Evaluation (OTROPE), a likelihood-free evaluation that performs distributional correction in a semantic space via optimal transport to align labeled behavior-policy samples with unlabeled target-policy samples. OTROPE combines corrected human-labeled residuals with proxy predictors, yielding a doubly robust-style evaluation without behavior-policy modeling or density-ratio estimation. We theoretically characterize why baseline evaluators fail under LLM distribution shift, and establish consistency and convergence rates for OTROPE when either the reweighted behavior distribution or the proxy predictor converges. Experiments on synthetic and real LLM evaluation tasks show that OTROPE consistently outperforms baselines while enabling ensembles of weaker LLM evaluators to approach and sometimes surpass stronger evaluators. Code is available at https://github.com/LinerXiang/OTROPE.
Chinese Translation
对大语言模型(LLM)的可靠评估对其开发和部署至关重要,但往往成本高昂、风险大,并且难以安全地在线进行。我们研究针对 LLM 的离策略评估,其中来自行为模型的有限人工标注数据被用于评估一个更新的目标 LLM。这一设定具有挑战性,因为标签稀缺,行为与目标之间的分布偏移很常见,而且对于黑盒 LLM,响应似然通常不可获得。我们提出基于最优传输的鲁棒离策略评估(OTROPE),这是一种无需似然的评估方法,通过最优传输在语义空间中进行分布校正,以将带标注的行为策略样本与无标注的目标策略样本对齐。OTROPE 将校正后的人工标注残差与代理预测器相结合,产生一种双重稳健风格的评估,而无需行为策略建模或密度比估计。我们从理论上刻画了为什么基线评估器在 LLM 分布偏移下会失败,并为 OTROPE 在重加权后的行为分布或代理预测器任一者收敛时建立了相合性和收敛速率。在合成和真实 LLM 评估任务上的实验表明,OTROPE 始终优于基线,同时使较弱 LLM 评估器的集成能够接近并有时超过更强的评估器。代码可在 https://github.com/LinerXiang/OTROPE 获取。
cs.AI / 6 / 2609.36319
StateTape: Action-Conditioned Evidence Lifecycle Modeling for Long-Horizon Coding Agents
StateTape:面向长时程编程智能体的、以动作为条件的证据生命周期建模
large language model
大语言模型相关
Abstract
Despite the recent success of coding agents built on large language models, it remains challenging to run them over long horizons, since every observation is appended to the context and the context grows with each one. History-based maintenance is a common remedy, which masks or summarizes old observations, or prunes what a model reads as useless, and bounds the context at little cost. However, it decides from the text of the history alone and sees nothing of how the code is connected. Since a coding agent edits code many times over a single task, and each write can change what code elsewhere means, such maintenance may keep records a write has falsified, drop ones that still hold, and miss code the agent needs next. To overcome these challenges, this paper proposes StateTape, a novel and scalable framework that rewrites a coding agent's context as the repository changes rather than as the context grows. The key idea of StateTape is to model the repository as a symbol-level code graph, whose dependencies and language rules expose which symbols a write can affect. Upon this graph, a tape marks the symbols each write changed, which turns staleness from an inference about text into an observation of the agent's writes. We propose a per-write procedure in which the tape nominates the records a write could have falsified while a small manager model settles what the write log cannot, and further provide a theoretical analysis and TraceBench, a benchmark that labels what an agent is holding against what is actually needed. Empirically, we demonstrate that StateTape can effectively clear falsified records and retrieve what is needed, and thus achieve a higher resolve rate in all experiments spanned by six coding agents and three edit-heavy benchmarks with little computational overhead.
Chinese Translation
尽管基于大语言模型构建的编程智能体近来取得了成功,但让它们在长时程上运行仍然颇具挑战,因为每一次观测都会被追加到上下文中,而上下文会随每一次观测不断增长。基于历史的维护是一种常见的补救手段,它遮蔽或摘要化旧的观测,或剪除模型认为无用的内容,从而以很小的代价约束上下文规模。然而,它仅依据历史的文本做出决策,对代码之间如何相互关联一无所知。由于编程智能体在单个任务中会多次编辑代码,而每一次写入都可能改变其他位置代码的含义,这类维护可能保留已被某次写入证伪的记录、丢弃仍然成立的记录,并遗漏智能体接下来需要的代码。为克服这些挑战,本文提出 StateTape,一个新颖且可扩展的框架,它随着仓库的变化而非上下文的增长来重写编程智能体的上下文。StateTape 的核心思想是将仓库建模为符号级代码图,其依赖关系与语言规则揭示了某次写入可能影响哪些符号。在此图之上,一条磁带(tape)标记每次写入所改变的符号,从而将陈旧性从对文本的推断转变为对智能体写入的观测。我们提出一种逐写入流程:由磁带提名某次写入可能已证伪的记录,而由一个小型管理器模型裁定写入日志无法确定的部分;我们进一步提供理论分析以及 TraceBench——一个将智能体当前持有的内容与实际所需内容进行标注对比的基准。在实证上,我们证明 StateTape 能够有效清除被证伪的记录并检索所需内容,从而在涵盖六个编程智能体与三个编辑密集型基准的所有实验中,以很小的计算开销取得更高的解决率。
cs.AI / 7 / 2609.36357
HyperZip: Efficient Data Compression through Personalized Diffusion LLMs with Hypernetworks
HyperZip:通过带有超网络的个性化扩散LLM实现高效数据压缩
diffusionlarge language model
扩散模型相关
大语言模型相关
Abstract
Large language models (LLMs) have shown strong potential for lossless data compression, but existing approaches are constrained by the high computational cost and low throughput of autoregressive decoding. We propose HyperZip, an efficient and scalable LLM-based compression framework that leverages diffusion-based LLMs (dLLMs) with Multi-Token Prediction (MTP) to accelerate LLM-based data compression processes. We identify a trade-off in diffusion-based compression, where increasing decoding throughput degrades the compression rate. To mitigate this trade-off, HyperZip employs a hypernetwork to generate data-specific updates from a context representation, adapting the dLLM to the target data without costly fine-tuning, resulting in a low compression rate and high throughput. Extensive experiments show that HyperZip achieves a superior trade-off between compression rate and speed compared with state-of-the-art baselines.
Chinese Translation
大语言模型(LLMs)在无损数据压缩方面展现出强大潜力,但现有方法受限于自回归解码的高计算成本和低吞吐量。我们提出 HyperZip,一种高效且可扩展的基于LLM的压缩框架,它利用带有多令牌预测(MTP)的扩散式LLM(dLLMs)来加速基于LLM的数据压缩过程。我们识别出基于扩散的压缩中的一种权衡:提高解码吞吐量会降低压缩率。为缓解这一权衡,HyperZip 采用超网络从上下文表示中生成数据特定的更新,使 dLLM 适应目标数据而无需昂贵的微调,从而实现低压缩率和高吞吐量。大量实验表明,与最先进的基线相比,HyperZip 在压缩率和速度之间实现了更优的权衡。
cs.AI / 8 / 2609.36365
Engineering Simplicity: Simple Mechanism Interfaces Steer LLM Agents
工程化的简洁:简单机制接口引导 LLM 智能体
large language model
大语言模型相关
Abstract
Can interaction formats and textual scaffolds help large language model (LLM) agents make better decisions, and do better decisions come with better explanations? We study these questions in auctions and matching, multi-agent environments with explicit rules and known optimal strategies. These settings let us vary how a decision problem is presented while retaining a benchmark for evaluating behavior. Drawing on human-motivated theories of simplicity, we compare interfaces that elicit a complete bid or ranking with sequential interfaces that make safe choices easier to identify. We then hold the interaction format fixed and vary reasoning scaffolds and rule descriptions. Across four model families, the ascending auction interface substantially reduces bid deviations. The matching comparison also shows why sequential responses require different error accounting from complete rankings. Laying out payoff contingencies and explaining why truth-telling is safe also improve choices, whereas prompts to plan through matching rounds or form beliefs about opponents worsen play overall. In auctions, these behavioral gains are not accompanied by corresponding improvements in measured verbal indicators of strategic understanding in the agents' short stated plans. Other prompts change those indicators without improving bids. Our findings suggest that human-motivated theories of simplicity can inform the design of decision environments for artificial agents. They also show why scaffolds should be evaluated through realized choices as well as explanations: improvements in one need not appear in the other.
Chinese Translation
交互格式和文本脚手架能否帮助大型语言模型(LLM)智能体做出更好的决策?更好的决策是否伴随更好的解释?我们在拍卖和匹配这两种具有明确规则和已知最优策略的多智能体环境中研究这些问题。这些设定使我们能够在保留用于评估行为的基准的同时,改变决策问题的呈现方式。借鉴以人类为动机的简洁性理论,我们比较了引出完整出价或排序的界面与使安全选择更容易识别的序贯界面。然后,我们固定交互格式,改变推理脚手架和规则描述。在四个模型系列中,升价拍卖界面大幅减少了出价偏差。匹配比较还表明,为什么序贯回应需要与完整排序不同的误差核算。列出收益的各种可能情形并解释为何说真话是安全的,也会改善选择,而提示通过匹配轮次进行规划或形成关于对手的信念,总体上会恶化博弈表现。在拍卖中,这些行为收益并未伴随智能体简短陈述计划中测量到的策略理解言语指标的相应改善。其他提示改变了这些指标,却没有改善出价。我们的发现表明,以人类为动机的简洁性理论可以为人工智能体的决策环境设计提供参考。它们还表明,为什么应该通过实际选择以及解释来评估脚手架:一个方面的改善不一定出现在另一个方面。
cs.AI / 9 / 2609.36461
Rethinking Reasoning Paths as Phase-Structured Trajectories
将推理路径重新思考为阶段结构化轨迹
large language model
大语言模型相关
Abstract
Large language models often improve problem-solving performance by generating multi-step reasoning paths, yet how to analyze the hidden states along these paths remains unclear. Existing approaches typically assign each intermediate state the final-answer correctness label and train probes across heterogeneous questions. We argue that this protocol obscures reasoning dynamics in two ways: (1) correctness prediction can exploit question-level variation rather than path quality, and (2) states aligned by absolute step indices may correspond to different functional phases of reasoning. In this work, we propose to view reasoning paths as phase-structured trajectories within fixed questions. We instantiate this view as PAIR, short for Phase-Aligned Intra-question Reasoning. PAIR samples multiple trajectories for each question, maps variable-length paths into shared relative phases based on normalized trajectory progress, and compares successful and unsuccessful trajectories only within the same question and phase. This yields phase-specific path-quality directions that better isolate path-quality signals from question-level variation. Empirically, we find that standard across-question correctness probes lose much of their predictive power under within-question evaluation, suggesting that these probes partly rely on question-level information. PAIR improves within-question trajectory ranking and Best-of-N trajectory selection across models and benchmarks. Phase-wise steering further shows that the learned directions can change generation outcomes, providing causal evidence that they capture trajectory-relevant information.
Chinese Translation
大语言模型通常通过生成多步推理路径来提高问题求解性能,然而如何分析这些路径上的隐藏状态仍不清楚。现有方法通常为每个中间状态分配最终答案正确性标签,并在异构问题上训练探针。我们认为,这一协议在两个方面掩盖了推理动态:(1) 正确性预测可以利用问题层面的变异,而不是路径质量;(2) 按绝对步骤索引对齐的状态可能对应推理的不同功能阶段。在这项工作中,我们提出将推理路径视为固定问题内的阶段结构化轨迹。我们将这一观点实例化为 PAIR,即 Phase-Aligned Intra-question Reasoning 的缩写。PAIR 为每个问题采样多条轨迹,基于归一化的轨迹进度将可变长度路径映射到共享的相对阶段,并且仅在相同问题和相同阶段内比较成功轨迹与不成功轨迹。这产生了阶段特定的路径质量方向,能够更好地将路径质量信号与问题层面的变异隔离开来。从经验上看,我们发现标准的跨问题正确性探针在问题内评估下失去了大部分预测能力,这表明这些探针部分依赖于问题层面的信息。PAIR 在跨模型和基准上改进了问题内轨迹排序和 Best-of-N 轨迹选择。按阶段引导进一步表明,学习到的方向可以改变生成结果,提供了它们捕获轨迹相关信息的因果证据。
cs.AI / 10 / 2609.36505
BRIDGE: Bilevel Retrieval-Credit-Aware Agentic Reinforcement Learning
BRIDGE:双层检索信用感知的智能体强化学习
large language model
大语言模型相关
Abstract
Agentic reinforcement learning (ARL) with verifiable rewards improves the ability of large language models (LLMs) to tackle knowledge-intensive tasks by learning to interleave search and reasoning. However, most existing ARL methods optimize only LLM-generated tokens and treat retrieved evidence as environment observations. This creates an information-credit gap: failures caused by missing or misleading evidence are attributed to the LLM policy rather than to the retriever, which motivates training the LLM and the retriever jointly. In this paper, we show that retrieval and LLM policy learning are order-sensitive: adapting the retriever before optimizing the policy yields a larger reward gain than the reverse order. To preserve this hierarchy while allowing both components to co-adapt, we formulate retrieval-augmented agentic RL as a bilevel optimization problem. To solve it efficiently, we introduce BRIDGE, a memory-efficient first-order bilevel method motivated by a loss-landscape analysis of the RL and retrieval objectives. Across seven open-domain QA benchmarks, BRIDGE achieves the highest average accuracy with both 3B and 7B backbones, improving the multi-hop average over the strongest baseline by 9.6 and 3.4 EM points, respectively. It also achieves the best averaged answer accuracy and reasoning quality across medical QA benchmarks.
Chinese Translation
具有可验证奖励的智能体强化学习(ARL)通过学习交错进行搜索与推理,提升大语言模型(LLM)处理知识密集型任务的能力。然而,大多数现有 ARL 方法仅优化 LLM 生成的 token,并将检索到的证据视为环境观测。这造成了一个信息信用差距:由缺失或误导性证据导致的失败被归因于 LLM 策略,而不是检索器,这促使我们联合训练 LLM 和检索器。在本文中,我们表明检索与 LLM 策略学习对顺序敏感:在优化策略之前先适应检索器,比相反顺序带来更大的奖励增益。为了在允许两个组件共同适应的同时保持这种层级关系,我们将检索增强的智能体强化学习形式化为一个双层优化问题。为了高效求解该问题,我们提出了 BRIDGE,一种内存高效的一阶双层方法,其动机来自对强化学习目标和检索目标的损失景观分析。在七个开放域问答基准上,BRIDGE 在 3B 和 7B 骨干模型上均取得最高平均准确率,分别将多跳平均性能相较于最强基线提高了 9.6 和 3.4 个 EM 点。它还在医学问答基准上取得了最佳的平均答案准确率和推理质量。
cs.AI / 11 / 2609.36576
Divide and Inject: Can Agents Reconstruct an Indirect Prompt Injection from Fragments?
分而注入:智能体能否从碎片中重建间接提示注入?
large language model
大语言模型相关
Abstract
Agentic systems are now being widely used to orchestrate tools and reason over long contexts. However, the improving capabilities of the large language models powering these agents also create new attack surfaces for indirect prompt injection. In particular, an attacker may not need to place a complete malicious instruction in retrieved content if the agent can reconstruct the objective from incomplete fragments distributed across a long context. In this work, we introduce adaptive long-context prompt injection (AdaLCPI), which combines long-context fragmentation with adaptive search. AdaLCPI splits an attack objective into incomplete fragments, embeds them in external content retrieved through the agent's tools, and uses a reconstruction cue to prompt the agent to combine them. It then iteratively refines the fragments and cue with OpenEvolve using graded scoring and natural-language execution feedback from the target agent. Empirically, AdaLCPI achieves higher attack success than strong adaptive baselines, reaching 61.4\% macro-average ASR compared with 32.8\% for Trojan Hippo-style and 30.0\% for AgentVigil. Safety evaluations should therefore test whether agents remain robust when harmful objectives must be reconstructed from incomplete fragments.
Chinese Translation
智能体系统现在正被广泛用于编排工具并在长上下文上进行推理。然而,驱动这些智能体的大语言模型不断增强的能力,也为间接提示注入创造了新的攻击面。特别地,如果智能体能够从分布在长上下文中的不完整片段中重建目标,攻击者可能不需要在检索内容中放置完整的恶意指令。在这项工作中,我们引入自适应长上下文提示注入(AdaLCPI),它结合了长上下文碎片化与自适应搜索。AdaLCPI 将攻击目标拆分为不完整片段,将其嵌入通过智能体工具检索到的外部内容中,并使用重建线索来提示智能体将它们组合起来。然后,它使用来自目标智能体的分级评分和自然语言执行反馈,通过 OpenEvolve 迭代优化片段和线索。实证上,AdaLCPI 实现了比强自适应基线更高的攻击成功率,达到 61.4\% 的宏平均 ASR,相比之下 Trojan Hippo 风格为 32.8\%,AgentVigil 为 30.0\%。因此,安全评估应测试当有害目标必须从不完整片段中重建时,智能体是否仍保持鲁棒。
cs.AI / 12 / 2609.36671
FairDiff: Mitigating the Self-Reinforcing Matthew Effect in Diffusion Recommender Models
FairDiff:缓解扩散推荐模型中的自强化马太效应
diffusion
扩散模型相关
Abstract
While the "Matthew Effect" and filter bubbles are widely recognized outcome-level biases in recommender systems, we reveal that Diffusion Recommender Models (DRMs) uniquely compound this issue through their generative dynamics. Rather than merely inheriting data imbalances, DRMs trigger a self-reinforcing amplification of popularity bias. We identify that this phenomenon is driven by two compounding mechanisms. First, while optimization loss is universally dominated by high-frequency items across recommenders, DRMs suffer from a unique structural prior mismatch during generation. Because the forward terminal distribution of long-tailed data deviates significantly from the standard Gaussian prior, reverse sampling trajectories inherently collapse toward high-density popular items, fundamentally suppressing niche item generation. To dismantle this self-reinforcing loop, we propose FairDiff, a plug-and-play fairness-aware diffusion framework. To overcome the popularity-dominated loss, we introduce Popularity Condition Guidance (PCG). Rather than altering the training objective, PCG acts as an inference-time distributional reweighting mechanism, mathematically reshaping the score-based gradient field to penalize high-popularity regions and guide trajectories toward niche semantics. Furthermore, we design a Semantic Calibration (SC) Module to bridge the prior mismatch, aligning the forward and reverse distributions via one-step optimal transport. Comprehensive evaluations demonstrate that FairDiff achieves state-of-the-art performance while effectively mitigating the self-reinforcing Matthew Effect, highlighting its value as a general framework for DRMs.
Chinese Translation
尽管“马太效应”和过滤气泡在推荐系统中被广泛认为是结果层面的偏差,但我们揭示出扩散推荐模型(DRMs)通过其生成动态独特地加剧了这一问题。DRMs 并非仅仅继承数据不平衡,而是触发了流行度偏差的自强化放大。我们识别出这一现象由两种相互叠加的机制驱动。首先,尽管优化损失在所有推荐器中普遍由高频物品主导,但 DRMs 在生成过程中却遭受独特的结构先验失配。由于长尾数据的前向终端分布显著偏离标准高斯先验,反向采样轨迹固有地坍缩向高密度流行物品,从根本上抑制了小众物品的生成。为了打破这一自强化循环,我们提出了 FairDiff,一个即插即用的公平感知扩散框架。为了克服流行度主导的损失,我们引入了流行度条件引导(PCG)。PCG 并非改变训练目标,而是作为一种推理时的分布重加权机制,在数学上重塑基于分数的梯度场,以惩罚高流行度区域并将轨迹引导向小众语义。此外,我们设计了一个语义校准(SC)模块来弥合先验失配,通过一步最优传输对齐前向和反向分布。全面的评估表明,FairDiff 在取得最先进的性能的同时,有效缓解了自强化马太效应,凸显了其作为 DRMs 通用框架的价值。
cs.AI / 13 / 2609.36735
BiFE: Search-Efficient Discovery of CPU-Only Branching Policies via LLM-based Bi-Fidelity Evolution
BiFE:通过基于LLM的双保真进化实现仅CPU分支策略的搜索高效发现
large language model
大语言模型相关
Abstract
In branch-and-bound (B&B) for mixed-integer linear programming (MILP), branching variable selection critically impacts efficiency. Existing neural branching policies often require GPU inference, while CPU-efficient symbolic expressions lack the representational capacity for complex logic. Large Language Model (LLM)-generated code provides a flexible search space for designing lightweight branching rules with diverse algorithmic logic. To discover effective rules within LLM-based evolutionary frameworks, a core challenge arises: full B&B evaluation on real instances is prohibitively expensive, whereas offline imitation learning suffers from distribution shift. To address this, we introduce a Bi-Fidelity Evolutionary framework (BiFE). It employs low-fidelity imitation scores as a rapid pre-screener and selectively applies high-fidelity on-instance evaluation only to elite candidates, effectively balancing search efficiency with performance reliability. Experiments validate both the search efficiency of BiFE and the competitiveness of its discovered rules, which outperform the SCIP solver and other baselines on CPUs, and even surpass certain GPU-based neural policies.
Chinese Translation
在混合整数线性规划(MILP)的分支定界(B&B)中,分支变量选择对效率有至关重要的影响。现有的神经分支策略通常需要GPU推理,而CPU高效的符号表达式缺乏对复杂逻辑的表示能力。大型语言模型(LLM)生成的代码为设计具有多样算法逻辑的轻量级分支规则提供了灵活的搜索空间。在基于LLM的进化框架中发现有效规则时,一个核心挑战出现:在真实实例上进行完整B&B评估昂贵得难以承受,而离线模仿学习则遭受分布偏移。为解决这一问题,我们引入双保真进化框架(BiFE)。它采用低保真模仿分数作为快速预筛选器,并仅对精英候选者选择性地应用高保真实例评估,从而有效平衡搜索效率与性能可靠性。实验验证了BiFE的搜索效率及其所发现规则的竞争力,这些规则在CPU上优于SCIP求解器和其他基线,甚至超过某些基于GPU的神经策略。
cs.AI / 14 / 2609.36742
SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation
SIPO:统一强化学习与同策略自蒸馏
large language model
大语言模型相关
Abstract
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate steps. To address this, on-policy self-distillation (OPSD) leverages a self-teacher with privileged context to provide additional dense learning signals. However, because the self-teacher is often overconfident and imposes excessive penalties on long reasoning trajectories, OPSD frequently struggles in practice. To mitigate this, we propose self-instructing policy optimization (SIPO) with a contrastive self-teacher to provide dense credit. At each iteration, SIPO samples multiple rollouts per prompt from the current policy, scores them with environment rewards, and constructs two teacher contexts for each rollout by pairing the reference answer with mistakes made within the group. The model then re-evaluates its own responses under both contexts, using the difference between the two teacher log-probabilities as token-level feedback, so that biases shared by both contexts are expected to largely cancel. The resulting objective yields a token-level advantage for every rollout: the reward still sets the main direction of each update while the self-teacher redistributes credit across tokens. Even in groups where every rollout fails and group-relative advantages vanish, SIPO still provides a learning signal. By preserving direct optimization of the task reward while providing dense, token-level feedback, this approach bridges reinforcement learning and on-policy self-distillation. Extensive experiments across multiple reasoning and code-generation benchmarks demonstrate that SIPO outperforms both RLVR and OPSD baselines without an external teacher or additional generation.
Chinese Translation
具有可验证奖励的强化学习(RLVR)已成为在各种任务上提升大语言模型(LLMs)的标准范式,然而其稀疏的结果奖励缺乏对中间步骤的 token 级信用分配。为了解决这一问题,同策略自蒸馏(OPSD)利用具有特权上下文的自教师来提供额外的稠密学习信号。然而,由于自教师往往过度自信,并对长推理轨迹施加过度惩罚,OPSD 在实践中经常表现不佳。为了缓解这一问题,我们提出了带有对比式自教师的自我指导策略优化(SIPO),以提供稠密信用。在每次迭代中,SIPO 从当前策略为每个提示采样多个轨迹,用环境奖励对它们评分,并通过将参考答案与该组内所犯错误配对,为每个轨迹构造两个教师上下文。然后,模型在两个上下文中重新评估自己的响应,使用两个教师对数概率之间的差值作为 token 级反馈,从而两个上下文共有的偏差预计会在很大程度上抵消。由此得到的目标函数为每个轨迹产生一个 token 级优势:奖励仍然设定每次更新的主要方向,而自教师则在 token 之间重新分配信用。即使在每个轨迹都失败且组相对优势消失的组中,SIPO 仍然提供学习信号。通过保留对任务奖励的直接优化,同时提供稠密的、token 级反馈,该方法弥合了强化学习与同策略自蒸馏之间的鸿沟。在多个推理和代码生成基准上的大量实验表明,SIPO 在无需外部教师或额外生成的情况下,优于 RLVR 和 OPSD 基线。
cs.AI / 15 / 2609.36748
Generalizable Lifelong Model Editing via Preference Optimization
通过偏好优化的可泛化终身模型编辑
large language model
大语言模型相关
Abstract
Knowledge editing enables rapid updates of specific factual knowledge in large language models (LLMs) without full retraining. However, more realistic scenarios call for a lifelong framework that handles continual updates rather than one-off modifications. In such settings, existing editing methods often overfit to target prompts, significantly degrading both the generalization of the edited knowledge and the model's general capabilities. To address this issue, we propose GLIME (Generalizable Lifelong Model Editing), which combines knowledge editing with preference optimization over generation behavior. GLIME further incorporates replay-based editing and a gradient constraint to preserve previously edited knowledge. Experimental results show that GLIME significantly improves knowledge generalization in lifelong editing settings while maintaining both editing performance and general capabilities.
Chinese Translation
知识编辑能够在不进行完整重训练的情况下,快速更新大型语言模型(LLMs)中的特定事实知识。然而,更现实的场景需要一个终身框架,以处理持续更新而非一次性修改。在此类设置中,现有的编辑方法常常会过拟合目标提示,从而显著降低被编辑知识的泛化能力以及模型的通用能力。为了解决这一问题,我们提出了 GLIME(Generalizable Lifelong Model Editing,可泛化终身模型编辑),它将知识编辑与对生成行为的偏好优化相结合。GLIME 进一步结合了基于回放的编辑和梯度约束,以保留先前编辑过的知识。实验结果表明,GLIME 在终身编辑设置中显著提升了知识泛化能力,同时保持了编辑性能和通用能力。
cs.AI / 16 / 2609.36800
AI as a Compiler: Compiling Triton kernels without the Triton compiler
AI 作为编译器:无需 Triton 编译器编译 Triton 内核
large language model
大语言模型相关
Abstract
Compiler backends are expensive to build and maintain as programming models, workloads, and accelerators evolve. We investigate whether large language models can replace the conventional optimizing and lowering pipeline, a process that we call AI lowering. We study AI lowering from Triton to NVIDIA PTX: an LLM agent translates Triton kernels directly into PTX. We build an environment that evaluates candidate PTX, and an agentic harness in which an LLM translates Triton kernels into PTX. Across twelve common kernels on Ada, Hopper, and Blackwell GPUs and ten kernels from recent ML papers, AI lowering achieves 0.83x-3.34x the performance of autotuned Triton. The largest gains come from transformations that Triton's lowering pipeline does not perform, such as decoding packed binary weights directly into Tensor Core operands (3.34x on BitDelta), assigning each thread a complete softmax row in tensor memory (1.37x on FlashAttention), and reusing overlapping convolution windows (up to 2.23x). These results rely on a robust evaluation harness with comprehensive verification support. We build on Volta, an existing PTX verifier, and substantially extend it to support modern GPU architectures by introducing support for Blackwell's tcgen05 Tensor Core interface. This requires modeling three architectural features: managed tensor memory, descriptor-based operand layouts, and asynchronous execution coordinated through commits, waits, memory barriers, and proxy fences. We discuss the challenges involved in formalizing them, as well as the current limitations. Our results suggest an emerging future in which AI compilers replace custom-written intermediate representations and checkers, reducing the time and engineering effort required to bring up software for new general-purpose and custom chips.
Chinese Translation
随着编程模型、工作负载和加速器不断演进,编译器后端的构建和维护成本高昂。我们研究大型语言模型能否取代传统的优化与 lowering 流水线,我们将这一过程称为 AI lowering。我们研究从 Triton 到 NVIDIA PTX 的 AI lowering:一个 LLM 智能体将 Triton 内核直接翻译为 PTX。我们构建了一个评估候选 PTX 的环境,以及一个智能体式测试框架,其中 LLM 将 Triton 内核翻译为 PTX。在 Ada、Hopper 和 Blackwell GPU 上的十二个常见内核以及来自近期 ML 论文的十个内核上,AI lowering 达到了自动调优 Triton 性能的 0.83 倍到 3.34 倍。最大的收益来自 Triton 的 lowering 流水线不会执行的变换,例如将打包的二进制权重直接解码为 Tensor Core 操作数(在 BitDelta 上为 3.34 倍)、在张量内存中为每个线程分配完整的 softmax 行(在 FlashAttention 上为 1.37 倍),以及复用重叠的卷积窗口(最高 2.23 倍)。这些结果依赖于一个具有全面验证支持的稳健评估框架。我们基于 Volta(一个现有的 PTX 验证器)构建,并通过引入对 Blackwell 的 tcgen05 Tensor Core 接口的支持,大幅扩展它以支持现代 GPU 架构。这需要对三个架构特性进行建模:受管理的张量内存、基于描述符的操作数布局,以及通过提交、等待、内存屏障和代理栅栏协调的异步执行。我们讨论了将它们形式化所涉及的挑战,以及当前的局限性。我们的结果表明了一个正在出现的未来,其中 AI 编译器取代定制的中间表示和检查器,从而减少为新的通用芯片和定制芯片启用软件所需的时间和工程投入。
cs.AI / 17 / 2609.36805
UpliftMem: Learning Set-Level Uplift for Agent Memory Retrieval
UpliftMem:为智能体记忆检索学习集合级提升
large language model
大语言模型相关
Abstract
Large language model (LLM) agents reuse external memory to guide new tasks, but effective retrieval requires learning which memory sets improve execution. Such learning relies on costly outcome feedback: ordinary retrieval observes only executed sets, while evaluating alternatives requires additional rollouts. We introduce \textsc{UpliftMem}, which learns memory retrieval from set-level execution uplift relative to the same executor without memory. A theoretical analysis of how retrieval preferences restrict feedback coverage motivates targeted probing of alternative memory sets. Probe selection follows an expected value of sample information (EVSI) criterion, derived in closed form under a correlated Gaussian model, to allocate limited training rollouts according to their expected improvement in local retrieval decisions. The shared scorer is trained with a frozen executor and selects memory sets without test-time probes. Across ALFWorld, WebShop, and BigCodeBench, \textsc{UpliftMem} achieves the best success rates among evaluated baselines on the main evaluation sets. Controlled fixed-store and matched probe budget evaluations further demonstrate improved memory-use decisions and more effective use of execution feedback.
Chinese Translation
大型语言模型(LLM)智能体复用外部记忆来指导新任务,但有效的检索需要学习哪些记忆集合能改善执行。此类学习依赖代价高昂的结果反馈:普通检索只观察到已执行的集合,而评估替代方案则需要额外的 rollout。我们引入 \textsc{UpliftMem},它从相对于同一执行器在没有记忆时的集合级执行提升中学习记忆检索。对检索偏好如何限制反馈覆盖的理论分析,激发了对替代记忆集合进行有针对性的探查。探查选择遵循样本信息期望价值(EVSI)准则,该准则在相关高斯模型下以闭式形式导出,以根据其对局部检索决策的预期改进来分配有限的训练 rollout。共享评分器与冻结的执行器一起训练,并在测试时无需探查即可选择记忆集合。在 ALFWorld、WebShop 和 BigCodeBench 上,\textsc{UpliftMem} 在主要评估集上取得了所评估基线中最佳的成功率。受控的固定存储与匹配探查预算评估进一步表明,其改进了记忆使用决策,并更有效地利用了执行反馈。
cs.AI / 18 / 2609.36828
Calibrate the Decisions That Change the Future: On-Policy Post-Training Quantization for Multimodal Large Language Models
校准那些改变未来的决策:面向多模态大语言模型的在策略训练后量化
large language model
大语言模型相关
Abstract
Post-training quantization (PTQ) lowers deployment cost for multimodal large language models, but calibration typically reconstructs fixed sequences with local objectives. This overlooks autoregressive feedback: a quantization-induced token change redirects the prefix and changes future states. Yet on-policy coverage alone is insufficient because many decision mismatches barely affect future generation. We propose OnPTQ, an on-policy framework that calibrates on trajectories visited by the current quantized policy. On shared prefixes, OnPTQ identifies quantization-eroded boundaries, evaluates competing tokens through short counterfactual rollouts, and combines current discrepancy with branch consequence into a Decision--Consequence risk. The risk prioritizes critical states, while context anchoring and trajectory refresh preserve multimodal behavior and keep calibration aligned with the updated policy. We further derive a Decision--Consequence bound linking behavioral deviation to current policy discrepancy and action-conditioned future-value span. Across vision--language and omni-modal Qwen models under multiple low-bit settings, OnPTQ improves downstream performance and yields fewer correctness flips against the corresponding Dense/FP16 references, without changing the deployed inference graph.
Chinese Translation
训练后量化(PTQ)降低了多模态大语言模型的部署成本,但校准通常以局部目标重建固定序列。这忽略了自回归反馈:由量化引起的 token 变化会重定向前缀并改变未来状态。然而,仅靠在策略覆盖并不充分,因为许多决策不匹配几乎不会影响未来的生成。我们提出 OnPTQ,一个在策略框架,它在当前量化策略所访问的轨迹上进行校准。在共享前缀上,OnPTQ 识别被量化侵蚀的边界,通过简短的反事实推演评估相互竞争的 token,并将当前差异与分支后果结合为一种 Decision--Consequence 风险。该风险对关键状态进行优先排序,而上下文锚定与轨迹刷新则保持多模态行为,并使校准与更新后的策略保持一致。我们进一步推导出一个 Decision--Consequence 界,将行为偏差与当前策略差异以及以动作为条件的未来价值跨度联系起来。在多种低比特设置下,跨视觉--语言与全模态 Qwen 模型,OnPTQ 提升了下游性能,并且相对于对应的 Dense/FP16 参考产生更少的正确性翻转,同时不改变所部署的推理图。
cs.AI / 19 / 2609.36830
Where Does Staleness Accumulate? Pool Aware Effective Staleness Control for Asynchronous RL in LLM Post-Training
陈旧性在哪里累积?面向 LLM 后训练中异步 RL 的池感知有效陈旧性控制
large language model
大语言模型相关
Abstract
Fully asynchronous reinforcement learning (RL) improves resource utilization in large language model post-training by overlapping rollout generation with policy optimization, but it also introduces policy lag as trajectories are generated and queued while the trainer continues to update. We study how this lag accumulates over a trajectory's lifetime and how it can be controlled without sacrificing the wall-clock benefits of asynchronous execution. We decompose trajectory staleness into Generation Staleness, accumulated before rollout completion, and Waiting Staleness, accumulated after a completed trajectory enters the pool. Motivated by this decomposition, we introduce PACE (Pool-Aware Control of Effective Staleness). PACE converts excess pool occupancy into an adaptive rejection budget and ranks completed trajectories using an effective-staleness score that combines Waiting Staleness with prefix-aware Generation Staleness. This avoids penalizing long or interrupted rollouts solely because they span multiple policy versions. In single-turn mathematical reasoning, PACE improves the six-benchmark average validation accuracy by 18.7\% over unfiltered asynchronous RL at the same wall-clock budget and matches synchronous RL performance with 47.1\% less GPU time. PACE also improves validation performance in multi-turn tool-integrated reasoning, outperforming both synchronous and unfiltered asynchronous RL. Further experiments with the mixture-of-experts model and an alternative RL algorithm support its applicability across model architectures and training algorithms.
Chinese Translation
完全异步强化学习(RL)通过将 rollout 生成与策略优化重叠,提高了大语言模型后训练中的资源利用率,但随着轨迹在训练器继续更新时被生成并排队,它也引入了策略滞后。我们研究这种滞后如何在轨迹的整个生命周期内累积,以及如何在不牺牲异步执行墙钟时间收益的情况下控制它。我们将轨迹陈旧性分解为生成陈旧性(Generation Staleness),其在 rollout 完成之前累积,以及等待陈旧性(Waiting Staleness),其在完成的轨迹进入池之后累积。受这一分解的启发,我们提出了 PACE(池感知的有效陈旧性控制)。PACE 将过量的池占用转换为自适应拒绝预算,并使用一个有效陈旧性分数对完成的轨迹进行排序,该分数将等待陈旧性与前缀感知的生成陈旧性相结合。这避免仅因为长或被中断的 rollout 跨越多个策略版本而惩罚它们。在单轮数学推理中,PACE 在相同的墙钟时间预算下,将六个基准的平均验证准确率较未过滤的异步 RL 提高了 18.7\%,并以少 47.1\% 的 GPU 时间达到与同步 RL 相当的性能。PACE 还提高了多轮工具集成推理中的验证性能,均优于同步和未过滤的异步 RL。使用专家混合模型和另一种 RL 算法进行的进一步实验支持了其跨模型架构和训练算法的适用性。
cs.AI / 20 / 2609.36835
ARC-KV: Amortizing Anchor Search for Reconstruction-Based KV Cache Compaction
ARC-KV:为基于重构的 KV 缓存压缩摊销锚点搜索
large language model
大语言模型相关
Abstract
Long-context large language model inference is bottlenecked by KV caches that grow linearly with sequence length. This burden is especially severe for long, reusable context prefixes, whose cache must serve many downstream queries. Reconstruction-based methods such as Attention Matching achieve strong downstream task performance with compact KV caches. However, iterative anchor search dominates the compaction cost of OMP-based Attention Matching. This motivates our selective amortization principle of learning a reusable anchor-selection policy across contexts while retaining context-specific reconstruction. In this work, we propose ARC-KV, a novel reconstruction-based KV cache compaction method that follows this principle. To this end, we first train a value-aware indexer to select real-key anchors in a single scoring pass. ARC-KV then applies convex-hull-constrained key merging and fits an attention-mass bias and compact values against the full cache. At inference time, ARC-KV builds the compact cache once per context using the frozen indexer and reuses it for all subsequent queries. Extensive experiments demonstrate that ARC-KV outperforms reported compaction methods in most settings across QuALITY, RULER, and LongBench on Llama-3.1-8B-Instruct. In particular, at 10% KV retention on QuALITY, ARC-KV improves accuracy from 0.6409 to 0.6474 over Attention Matching while reducing compaction time by a factor of 25.73, from 959.8 s to 37.3 s.
Chinese Translation
长上下文大语言模型推理受到随序列长度线性增长的 KV 缓存的瓶颈制约。对于长且可复用的上下文前缀而言,这一负担尤为严重,因为其缓存必须服务于许多下游查询。诸如 Attention Matching 之类的基于重构的方法利用紧凑的 KV 缓存实现了强劲的下游任务性能。然而,迭代锚点搜索主导了基于 OMP 的 Attention Matching 的压缩成本。这促使我们提出选择性摊销原则:学习跨上下文可复用的锚点选择策略,同时保留特定于上下文的重构。在这项工作中,我们提出 ARC-KV,一种遵循该原则的新型基于重构的 KV 缓存压缩方法。为此,我们首先训练一个值感知索引器,以在单次评分过程中选择真实键锚点。然后,ARC-KV 应用凸包约束的键合并,并针对完整缓存拟合注意力质量偏置和紧凑值。在推理时,ARC-KV 使用冻结的索引器为每个上下文构建一次紧凑缓存,并将其复用于所有后续查询。大量实验表明,在 Llama-3.1-8B-Instruct 上,ARC-KV 在 QuALITY、RULER 和 LongBench 的大多数设置中优于已报道的压缩方法。特别地,在 QuALITY 上以 10% 的 KV 保留率时,ARC-KV 相比 Attention Matching 将准确率从 0.6409 提升到 0.6474,同时将压缩时间缩短至原来的 1/25.73,从 959.8 s 降至 37.3 s。
cs.AI / 21 / 2609.36892
Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents
将外壳演化视为学习:自我改进个人智能体的近似、泛化与优化极限
large language model
大语言模型相关
Abstract
As the capabilities of large language models (LLMs) continue to advance, increasing attention is turning to how to translate their abilities into useful behavior. Personal agents bring this question into everyday settings, where models are expected to serve individual users and continually adapt to their preferences. With the underlying model held fixed, such adaptation relies on harness engineering: designing and evolving the surrounding layer that manages context, memory, tools, and execution. Despite rapid progress, the factors governing effective harness evolution remain insufficiently understood. To narrow this gap, we investigate three central questions concerning harness architecture, harness scale, and self-evolution algorithms through complementary empirical and theoretical analyses. Empirically, we introduce a preference-oriented benchmark and systematically characterize the capabilities and limitations of personal agents associated with these three dimensions. Theoretically, we formulate harness evolution as a learning problem and explain these phenomena through approximation, generalization, and optimization errors. Analyses of reachable policies, capacity under finite interaction evidence, and biased update dynamics provide theoretical accounts of the observed phenomena. Together, these results offer a unified perspective on the limits of personalization through harness evolution and inform future harness design.
Chinese Translation
随着大语言模型(LLMs)能力的持续提升,越来越多的关注正转向如何将其能力转化为有用的行为。个人智能体将这一问题带入日常场景,在这些场景中,模型被期望服务个体用户并持续适应其偏好。在底层模型保持固定的情况下,这种适应依赖于外壳工程:设计和演化用于管理上下文、记忆、工具和执行的外围层。尽管进展迅速,决定有效外壳演化的因素仍未被充分理解。为缩小这一差距,我们通过互补的实证与理论分析,研究关于外壳架构、外壳规模和自演化算法的三个核心问题。在实证方面,我们引入一个面向偏好的基准,并系统刻画与这三个维度相关的个人智能体的能力和局限。在理论方面,我们将外壳演化形式化为一个学习问题,并通过近似误差、泛化误差和优化误差来解释这些现象。对可达策略、有限交互证据下的容量以及有偏更新动力学的分析,为所观察到的现象提供了理论解释。总之,这些结果为通过外壳演化实现个性化的极限提供了一个统一视角,并为未来的外壳设计提供启示。
cs.AI / 22 / 2609.36935
CoEM: Empowering Long-Context Reasoning with Commit-on-Evidence Memory
CoEM:以证据提交记忆赋能长上下文推理
large language model
大语言模型相关
Abstract
Long-context reasoning is essential for complex and long-horizon tasks, yet the performance of large language models (LLMs) degrades as context length increases. Recent approaches address this by processing input chunk by chunk while maintaining a bounded textual memory in model context. However, premature information compression can discard critical details essential for subsequent reasoning. In this paper, we introduce Commit-on-Evidence Memory (CoEM), which learns when to convert source evidence into compact memory facts. Specifically, under a fixed context-memory budget, CoEM preserves potentially useful source excerpts verbatim in a pending set, allowing subsequent context to clarify their relevance before irreversible compression. As new context arrives, a learned policy revisits each pending excerpt and decides whether to promote it to the committed memory, retain it for further consideration, or discard it. A frozen verifier ensures proposed facts are accepted only if supported by retained excerpts and current context. To further guide effective memory management, we train this policy using reinforcement learning by combining fine-grained, step-level evidence rewards with final answer rewards. Extensive experiments demonstrate that CoEM consistently improves long-context reasoning. When evaluated on 6,400 documents long-context input, CoEM outperforms the strongest memory baseline by 10.4-11.4 F1 points on Qwen3.5-9B. Code repository: https://github.com/benmagnifico/CoEM.
Chinese Translation
长上下文推理对于复杂且长时程任务至关重要,然而大语言模型(LLMs)的性能会随着上下文长度的增加而下降。近期方法通过逐块处理输入,同时在模型上下文中维护有界的文本记忆来解决这一问题。然而,过早的信息压缩可能丢弃对后续推理至关重要的关键细节。在本文中,我们提出证据提交记忆(Commit-on-Evidence Memory,CoEM),它学习何时将源证据转换为紧凑的记忆事实。具体而言,在固定的上下文记忆预算下,CoEM 将潜在有用的源摘录逐字保留在一个待定集合中,使后续上下文能够在不可逆压缩之前澄清它们的相关性。当新上下文到来时,一个学习到的策略会重新审视每个待定摘录,并决定是将其提升到已提交记忆、保留以供进一步考虑,还是将其丢弃。一个冻结的验证器确保提议的事实仅在被保留的摘录和当前上下文支持时才被接受。为了进一步引导有效的记忆管理,我们使用强化学习训练该策略,将细粒度的、步骤级证据奖励与最终答案奖励相结合。大量实验表明,CoEM 持续提升长上下文推理。在 6,400 篇文档的长上下文输入上进行评估时,CoEM 在 Qwen3.5-9B 上比最强的记忆基线高出 10.4-11.4 个 F1 点。代码仓库:https://github.com/benmagnifico/CoEM。
cs.AI / 23 / 2609.37025
AnyAct: Universal Action for Self-Evolving Agents
AnyAct:面向自进化智能体的通用动作
large language model
大语言模型相关
Abstract
As large language models (LLMs) advance, AI agents are increasingly deployed in open-world environments to tackle complex sequential tasks (e.g., document processing, cross-application collaboration), relying heavily on actions ranging from GUI operations to semantic APIs. However, three core challenges persist: the "scale dilemma" of massive tool ecosystems exceeding LLM context windows, the "non-stationarity" of tool quality due to updates or outages, and the "heterogeneity" of feedback formats (pixels, text, structured data) creating information silos. To address these, we propose AnyAct, a universal action layer that unifies available capabilities into a self-evolving action space, enabling agents to operate efficiently and reliably in large-scale, dynamic tool ecosystems. AnyAct's core design focuses on two objectives: constructing this action space via hierarchical progressive retrieval (filtering task-relevant actions) and test-time reliability evolution (pruning unreliable actions), and enabling reliability-aware action orchestration through a heterogeneous observation grounding module that unifies multi-modal feedback. Additionally, it defines a hybrid action space (primitive + semantic actions) and optimizes for a balance between task success rate and execution cost. Evaluations on LiveMCPBench and OSMCP (a new benchmark we developed for multi-granularity action collaboration) demonstrate state-of-the-art performance. AnyAct delivers substantial performance gains over baseline methods across various LLM base models on LiveMCPBench and improvements are particularly notable for models with constrained native capabilities. On OSMCP, it achieves 77.27% overall success with only 50 steps, which is half the steps required by most competitors.
Chinese Translation
随着大语言模型(LLM)的进步,AI 智能体越来越多地被部署在开放世界环境中以处理复杂的序列任务(例如文档处理、跨应用协作),并高度依赖从 GUI 操作到语义 API 的各类动作。然而,三个核心挑战依然存在:海量工具生态超出 LLM 上下文窗口所导致的“规模困境”,因更新或故障而造成的工具质量的“非平稳性”,以及反馈格式(像素、文本、结构化数据)的“异质性”所造成的信息孤岛。为解决这些问题,我们提出 AnyAct,一个通用动作层,它将可用能力统一为一个自进化的动作空间,使智能体能够在大规模、动态的工具生态系统中高效且可靠地运行。AnyAct 的核心设计聚焦于两个目标:通过分层渐进式检索(筛选与任务相关的动作)和测试时可靠性演化(剪除不可靠的动作)来构建该动作空间,以及通过一个统一多模态反馈的异构观测接地模块来实现可靠性感知的动作编排。此外,它定义了一个混合动作空间(原子动作 + 语义动作),并针对任务成功率与执行成本之间的平衡进行优化。在 LiveMCPBench 和 OSMCP(我们为多粒度动作协作新开发的一个基准)上的评估展示了最先进的性能。在 LiveMCPBench 上,AnyAct 在多种 LLM 基座模型上相较基线方法带来了显著的性能提升,且对于原生能力受限的模型,改进尤为明显。在 OSMCP 上,它仅用 50 步就实现了 77.27% 的整体成功率,这是大多数竞争方法所需步数的一半。
cs.AI / 24 / 2609.37054
actr: aligning thoughts and responses for multilingual safety in reasoning llms
actr:为推理型大语言模型中的多语言安全对齐思维与响应
large language model
大语言模型相关
Abstract
Ensuring the safety of reasoning large language models (LLMs) across languages is essential for their reliable deployment. However, when exposed to jailbreak attacks in non-high-resource languages, these models may generate unsafe responses even when their reasoning traces identify safety risks. To address this issue, we propose aligning cross-lingual thoughts and responses (ACTR), a framework that improves multilingual safety alignment by strengthening the use of existing safety reasoning. Specifically, we first present the think gap score (TGS) to compare the normalized contributions of reasoning traces to attention outputs during response generation across languages, and use reasoning- trace substitution to measure the cross-lingual safety gap. Next, using a corpus of jailbreak queries, we assess neuron importance through changes in response representations caused by neuron masking and compare the high-importance neuron sets obtained with reasoning enabled and disabled to identify safety think neurons that support the use of safety reasoning. Finally, we devise neuron-selective consistency optimization (NSCO), which uses a frozen judge model to reward agreement between the safety categories of reasoning traces and responses while updating only the parameters associated with the selected neurons, requiring no human-annotated responses or preference data. Across two reasoning models, ACTR achieves lower average attack success rates than the evaluated state-of-the-art methods on AdvBench-X and MultiJail, with safety gains extending to unseen languages, while preserving or improving average performance on multilingual knowledge and mathematical reasoning tasks and limiting false refusals of benign requests. Warning: this paper contains examples with unsafe content.
Chinese Translation
确保推理型大语言模型(LLMs)在跨语言场景中的安全性对其可靠部署至关重要。然而,当面对非高资源语言中的越狱攻击时,即使这些模型的推理轨迹识别出安全风险,它们也可能生成不安全的响应。为解决这一问题,我们提出对齐跨语言思维与响应(ACTR),这是一个通过强化对现有安全推理的使用来改进多语言安全对齐的框架。具体而言,我们首先提出思维差距分数(TGS),用于比较跨语言生成响应期间推理轨迹对注意力输出的归一化贡献,并使用推理轨迹替换来衡量跨语言安全差距。接下来,使用越狱查询语料库,我们通过神经元掩蔽所导致的响应表示变化来评估神经元重要性,并比较在启用和禁用推理时获得的高重要性神经元集合,以识别支持使用安全推理的安全思维神经元。最后,我们设计神经元选择性一致性优化(NSCO),它使用一个冻结的评判模型来奖励推理轨迹与响应的安全类别之间的一致性,同时仅更新与所选神经元相关联的参数,无需人工标注的响应或偏好数据。在两个推理模型上,ACTR 在 AdvBench-X 和 MultiJail 上取得了比所评估的最先进方法更低的平均攻击成功率,其安全增益扩展到未见语言,同时在多语言知识和数学推理任务上保持或提升平均性能,并限制对良性请求的误拒。警告:本文包含带有不安全内容的示例。
cs.AI / 25 / 2609.37097
Breaking the Illusion of Review Reliability under Static Evaluation: SCOPE Fuzzing for LLM-based Scientific Reviewers
打破静态评估下评审可靠性的幻象:面向基于LLM的科学审稿人的SCOPE模糊测试
large language model
大语言模型相关
Abstract
The rapid growth of submissions and reviewing workload has accelerated the use of large language models (LLMs) in peer review. Prior studies suggest that LLM-based reviewers can penalize content perturbations, such as overclaiming, indicating a certain degree of reliability. Yet these conclusions are largely based on a narrow set of perturbation strategies instantiated with static templates, providing limited evidence of actual reliability. In this paper, we construct a three-level evaluation framework covering perturbations to surface presentation, argumentative logic, and value judgment. Experiments on representative LLM-based reviewers reveal two limitations of static evaluation: stratified vulnerability, where perturbation effects depend on whether the paper's original review score is high or low, and perturbation undercoverage, where a single template misses vulnerabilities exposed by diverse realizations. To address these limitations, we propose SCOPE-Fuzzer, a strategy-aware fuzzer that combines feedback-driven strategy selection with adaptive mutation of paper content. By iteratively probing reviewers with dynamic perturbations, SCOPE-Fuzzer consistently uncovers vulnerabilities overlooked by static evaluation and other baselines.
Chinese Translation
投稿量和审稿工作量的快速增长加速了大语言模型(LLMs)在同行评审中的使用。既有研究表明,基于LLM的审稿人能够对内容扰动(例如过度宣称)进行惩罚,这表明其具有一定程度的可靠性。然而,这些结论在很大程度上基于一组狭窄的、由静态模板实例化的扰动策略,只能为实际可靠性提供有限证据。在本文中,我们构建了一个三级评估框架,涵盖对表面表述、论证逻辑和价值判断的扰动。在代表性基于LLM的审稿人上的实验揭示了静态评估的两个局限性:分层脆弱性,即扰动效应取决于论文的原始评审分数是高还是低;以及扰动覆盖不足,即单一模板会遗漏由多样化实现所暴露的脆弱性。为解决这些局限性,我们提出了SCOPE-Fuzzer,一种策略感知模糊测试器,它将反馈驱动的策略选择与论文内容的自适应变异相结合。通过用动态扰动迭代探测审稿人,SCOPE-Fuzzer能够持续发现被静态评估和其他基线所忽视的脆弱性。
cs.AI / 26 / 2609.37128
SkillCome: Group Contrast Skill Optimization with Dual Memory
SkillCome:基于双重记忆的组对比技能优化
large language model
大语言模型相关
Abstract
Skill evolution improves the capabilities of large language models by analyzing trajectories generated under a given skill and modifying the skill accordingly. Existing approaches typically generate a single trajectory per question. However, this provides insufficient optimization signals since it requires inferring effective skill edits from a solitary path. It is difficult to pinpoint which actions caused the failure in a failed trajectory, or to determine which actions in a successful one should be incorporated into the skill. Furthermore, they rely on a local batch of trajectories for analysis, making the optimization direction susceptible to noisy evidence. To address these, we propose SkillCome, a Skill-evolution method based on group Contrast optimization with dual memory. For each question, SkillCome generates trajectories and performs group contrast analysis to precisely identify key behavioral divergences between successful and failed trajectories, offering reliable optimization signals. The dual memory system further accumulates evidence from historical steps to track patterns shared across different groups, leading to more generalized optimization directions. Together, SkillCome builds a systematic optimization process that transforms experience from observed successful trajectories into reusable skills. Extensive experiments on six benchmarks spanning question answering, reasoning, and agentic tasks demonstrate the effectiveness of our method. SkillCome consistently outperforms baselines across five models of varying families and scales, with gains up to +5.69 points.
Chinese Translation
技能演化通过分析在给定技能下生成的轨迹并相应地修改技能,提升大语言模型的能力。现有方法通常为每个问题生成单条轨迹。然而,这提供了不足的优化信号,因为它需要从单一轨迹路径推断有效的技能编辑。很难精确定位在失败轨迹中哪些动作导致了失败,或确定在成功轨迹中哪些动作应被纳入技能。此外,它们依赖局部的一批轨迹进行分析,使优化方向容易受到噪声证据的影响。为解决这些问题,我们提出 SkillCome,一种基于组对比优化与双重记忆的技能演化方法。对于每个问题,SkillCome 生成多条轨迹并执行组对比分析,以精确识别成功轨迹与失败轨迹之间的关键行为差异,从而提供可靠的优化信号。双重记忆系统进一步积累来自历史步骤的证据,以追踪不同组之间共享的模式,从而得到更具泛化性的优化方向。综合起来,SkillCome 构建了一个系统化的优化过程,将观察到的成功轨迹中的经验转化为可复用的技能。在涵盖问答、推理和智能体任务的六个基准上进行的广泛实验证明了我们方法的有效性。SkillCome 在五个不同系列和规模的模型上持续优于基线,增益最高达 +5.69 分。
cs.AI / 27 / 2609.37132
Train Ahead, Distill Back: Bootstrapping On-Policy Self-Distillation for Large Language Models
先行训练,反向蒸馏:为大型语言模型自举在线策略自蒸馏
large language model
大语言模型相关
Abstract
On-policy self-distillation (OPSD) improves large language models by letting a self-teacher with privileged information provide dense token-level supervision on the model's own trajectories. Yet existing methods typically construct the self-teacher from the current, initial, or slowly averaged policy state, leaving the quality of supervision constrained by the teacher's ability to exploit privileged information. We ask whether the model's own optimization progress can instead be recycled into a stronger self-teacher. In this paper, we introduce Bootstrapped On-Policy Self-Distillation (B-OPSD), which temporarily trains the policy ahead to obtain a future teacher, restores the student to the original policy state, and then uses the future teacher to supervise the restarted student. The future teacher improves supervision in two complementary ways, it can generate more reliable privileged trajectories and, conditioned on them, provide more informative token-level targets along the restarted student's on-policy trajectories. Experiments on mathematical reasoning with Qwen3-4B and Qwen3-8B show consistent improvements over standard OPSD in both settings, including gains from 27.50 to 41.30 and from 48.80 to 64.44 in the rollout-privileged setting. Our findings point to a broader principle for self-improving models that future learning progress can be distilled backward, preserving acquired knowledge while bootstrapping beyond the optimization state that produced it.
Chinese Translation
在线策略自蒸馏(OPSD)通过让一个拥有特权信息的自教师模型对模型自身的轨迹提供密集的词元级监督,从而改进大型语言模型。然而,现有方法通常从当前、初始或缓慢平均的策略状态构建自教师,使得监督的质量受限于教师利用特权信息的能力。我们提出疑问:模型自身的优化进展能否反过来被回收利用,以形成一个更强的自教师?在本文中,我们引入了自举在线策略自蒸馏(B-OPSD),它暂时将策略向前训练以获得一个未来教师,将学生恢复到原始策略状态,然后使用该未来教师来监督重新启动的学生。未来教师以两种互补的方式改进监督:它可以生成更可靠的特权轨迹,并且在这些轨迹的条件下,沿着重新启动学生的在线策略轨迹提供更具信息量的词元级目标。在 Qwen3-4B 和 Qwen3-8B 上的数学推理实验表明,在两种设置下都相较标准 OPSD 取得了一致的改进,包括在 rollout 特权设置下从 27.50 提升到 41.30 以及从 48.80 提升到 64.44。我们的发现指向了一条更广泛的自我改进模型原则:未来的学习进展可以被反向蒸馏,从而在保留已获得知识的同时,自举超越产生该知识的优化状态。
cs.AI / 28 / 2609.37157
From Learner Behavior to Reusable Skills for Effective and Efficient Learner Simulation
从学习者行为到可复用技能:实现有效且高效的学习者模拟
large language model
大语言模型相关
Abstract
Learner simulation aims to reproduce how a particular learner behaves on new tasks. Although Large Language Models (LLMs) can generate increasingly fine-grained learning behaviors, existing approaches often need to repeatedly process a growing interaction history to reconstruct the learner. This introduces additional context and inference costs and makes the acquired learner-specific simulation capability difficult to reuse across different LLMs. We therefore propose Learner2Skill, which externalizes the simulation capability acquired from historical interactions into a persistent and reusable Simulation Skill. The Skill captures the learner's current learning state and recurring response patterns, evolves as new real interactions arrive, and can be adapted to a new LLM through lightweight executor calibration without reconstructing the learner from scratch. Experiments show that Learner2Skill more faithfully reproduces fine-grained learner behavior while reducing overall token cost, and that the same constructed Skills can be effectively reused across different LLM executors.
Chinese Translation
学习者模拟旨在再现特定学习者在新任务上的行为方式。尽管大型语言模型(LLMs)能够生成日益细粒度的学习行为,现有方法往往需要反复处理不断增长的交互历史,以重建学习者。这会引入额外的上下文和推理成本,并使所获得的特定于学习者的模拟能力难以在不同 LLM 之间复用。因此,我们提出 Learner2Skill,它将从历史交互中获得的模拟能力外化为一种持久且可复用的模拟技能。该技能捕捉学习者当前的学习状态和反复出现的响应模式,随着新的真实交互到来而演化,并可通过轻量级执行器校准适配到新的 LLM,而无需从头重建学习者。实验表明,Learner2Skill 在更忠实地再现细粒度学习者行为的同时降低了总体 token 成本,并且相同构建的技能可以有效地在不同 LLM 执行器之间复用。
cs.AI / 29 / 2609.37172
SimpleEvol: An Agent-Loop Framework for LLM-Driven Automated Heuristic Design with Minimal Human Priors
SimpleEvol:一个用于 LLM 驱动的自动化启发式设计的、具有最少人类先验的智能体循环框架
large language model
大语言模型相关
Abstract
Large language models (LLMs) have emerged as powerful tools for automated heuristic design (AHD), enabling iterative generation and refinement of heuristics. However, the dominant paradigm embeds LLMs as narrow, fixed components, such as crossover or mutation, within heavily hand-engineered evolutionary frameworks. We argue this misapprehends LLMs. It treats them as specialized tools rather than general reasoners, constrains them to low-level operations, and underutilizes their autonomy. Moreover, the extensive human priors in these frameworks violate the bitter lesson principle that general methods scaling with computation surpass hand-crafted solutions. This raises a key question: which AHD framework designs best convert stronger LLM capabilities into better heuristics? To address this, we propose metrics for LLM-driven AHD framework handcraftedness (AHI) and intelligence conversion efficiency (ICE). Evaluating ten LLMs across three challenging combinatorial optimization problems, we obtain a notable finding that frameworks with fewer human priors consistently yield higher ICE. Based on this finding, we propose SimpleEvol, an agent-loop framework for AHD which removes nearly all human priors and allows the LLM to operate autonomously. SimpleEvol consistently achieves the highest ICE, often by a large margin. Our results challenge the trend toward complex AHD pipelines and point to a lighter and more model-centric alternative, suggesting that reducing human priors is a more effective strategy to scale up with model intelligence. The source code is available at https://github.com/HenryZhu1029/SimpleEvol-Master.
Chinese Translation
大型语言模型(LLM)已成为自动化启发式设计(AHD)的强大工具,能够迭代地生成和精炼启发式方法。然而,主流范式将 LLM 作为狭窄、固定的组件(例如交叉或变异)嵌入到高度手工设计的进化框架中。我们认为这误解了 LLM。它将其视为专用工具而非通用推理器,将其限制在低层操作中,并且未能充分利用其自主性。此外,这些框架中大量的人类先验违反了苦涩教训原则,即随计算扩展的通用方法会超越手工设计的解决方案。这引出了一个关键问题:哪种 AHD 框架设计最能将更强的 LLM 能力转化为更好的启发式方法?为了解决这一问题,我们提出了用于衡量 LLM 驱动的 AHD 框架手工化程度 (AHI) 和智能转化效率 (ICE) 的指标。在三个具有挑战性的组合优化问题上评估十个 LLM,我们获得了一个显著发现:人类先验更少的框架始终产生更高的 ICE。基于这一发现,我们提出了 SimpleEvol,这是一个用于 AHD 的智能体循环框架,它移除了几乎所有人类先验,并允许 LLM 自主运行。SimpleEvol 始终取得最高的 ICE,且往往大幅领先。我们的结果挑战了朝向复杂 AHD 流水线的趋势,并指向一种更轻量、更以模型为中心的替代方案,表明减少人类先验是一种更有效的策略,可随模型智能一起扩展。源代码可在 https://github.com/HenryZhu1029/SimpleEvol-Master 获取。
cs.AI / 30 / 2609.37198
V-Engram: Trigger-Indexed External Memory for Modular Text-to-Image Personalization
V-Engram:面向模块化文本到图像个性化的触发器索引外部记忆
diffusion
扩散模型相关
Abstract
Pretrained text-to-image models contain broad visual knowledge, yet they cannot reliably acquire or refine a specific visual identity from only a few references while preserving compositional control. Token-embedding methods are compact but often underfit identity, whereas adapter-based methods improve fidelity through persistent weight updates that can be costly to store and interfere when concepts are composed. We introduce V-Engram, a trigger-indexed external memory mechanism for Stable Diffusion 3.5. Each concept is assigned an explicit trigger that retrieves concept-specific memory, whose gated directions enter frozen text-encoder and MMDiT context states as relative residuals. Separating this memory from backbone adaptation enables prompt-selective and multi-concept access without merging model updates. Experiments show that V-Engram broadly matches DreamBooth-LoRA in overall subject fidelity while showing advantages in settings such as contextual subject preservation. Prompt-matched loading retrieves only matched entries, reducing most additional adaptation-state loading for a single-concept query. Qualitative results further demonstrate paired-trigger composition and same-class separation, while prompts without registered entries retain the frozen model's base behavior. Together, these results establish trigger-indexed memory as a modular interface for adding targeted visual evidence without rewriting the generator.
Chinese Translation
预训练的文本到图像模型包含广泛的视觉知识,然而它们无法在仅凭少量参考图像的情况下可靠地获取或细化特定的视觉身份,同时保持组合控制。词元嵌入方法紧凑,但往往对身份拟合不足,而基于适配器的方法通过持久权重更新提高保真度,这些更新可能存储成本高昂,并且当概念被组合时会相互干扰。我们提出 V-Engram,一种用于 Stable Diffusion 3.5 的触发器索引外部记忆机制。每个概念被分配一个显式触发器,该触发器检索概念特定的记忆,其门控方向作为相对残差进入冻结的文本编码器和 MMDiT 上下文状态。将这种记忆与主干适配分离,使得无需合并模型更新即可实现提示选择性和多概念访问。实验表明,V-Engram 在整体主体保真度上大致与 DreamBooth-LoRA 匹配,同时在上下文主体保持等设置中显示出优势。提示匹配加载仅检索匹配的条目,从而减少了单概念查询的大部分额外适配状态加载。定性结果进一步展示了成对触发器组合和同类分离,而未注册条目的提示则保留冻结模型的基础行为。总之,这些结果确立了触发器索引记忆作为一种模块化接口,用于在不重写生成器的情况下添加有针对性的视觉证据。
cs.AI / 31 / 2609.37203
Learning to Prove, Not Just to Answer: Reinforcement Learning from Formal Verification for Natural-Language Logical Reasoning
学会证明,而不仅仅是回答:面向自然语言逻辑推理的、来自形式化验证的强化学习
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly deployed for natural-language logical reasoning, where the final answer is easy to check but the proof behind it is not. In natural-language logical reasoning, an intermediate conclusion should follow from its premises, and the resulting derivation should support the final answer. Existing methods lack machine-checkable verification of intermediate conclusions and answer-supporting proof dependencies, so they may assign credit to invalid or answer-irrelevant steps. We propose Proof-R1, an RL framework from formal verification that trains LLMs to construct verifiable proofs for natural-language logical reasoning. Proof-R1 admits a generated conclusion into the verified proof state only when the corresponding reasoning action satisfies the proof obligations through UNSAT-based machine-checkable formal verification. Proof-R1 also recovers the answer-supporting dependency closure to trace the proof structure of the final answer and align outcome credit with the proof dependencies. Experiments demonstrate that Proof-R1 improves answer accuracy across three logical reasoning benchmarks and four backbone models and outperforms training-free agents and training-based methods in terms of reasoning-process verifiability.
Chinese Translation
大语言模型(LLMs)正越来越多地被部署用于自然语言逻辑推理,其中最终答案易于检查,但其背后的证明却并非如此。在自然语言逻辑推理中,一个中间结论应当从其前提中推出,而由此得到的推导应当支持最终答案。现有方法缺乏对中间结论以及支持答案的证明依赖关系的机器可检查验证,因此它们可能会将信用分配给无效的或与答案无关的步骤。我们提出 Proof-R1,一个来自形式化验证的强化学习框架,它训练 LLMs 为自然语言逻辑推理构造可验证的证明。Proof-R1 仅在相应推理动作通过基于 UNSAT 的机器可检查形式化验证满足证明义务时,才将一个生成的结论接纳到已验证的证明状态中。Proof-R1 还恢复支持答案的依赖闭包,以追踪最终答案的证明结构,并使结果信用与证明依赖关系对齐。实验表明,Proof-R1 在三个逻辑推理基准和四个骨干模型上提高了答案准确率,并在推理过程可验证性方面优于无需训练的智能体和基于训练的方法。
cs.AI / 32 / 2609.37221
OptiCom : A Unified Framework for State-Conditioned Composition in LLM-Driven Optimization
OptiCom:面向 LLM 驱动优化的状态条件化组合统一框架
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly deployed to solve complex scientific and practical problems via iterative optimization. However, dynamically coordinating diverse search mechanisms as candidate quality, failure modes, and resource budgets evolve remains a critical open challenge. Targeted empirical diagnostics reveal that mechanism effectiveness is highly state-dependent. Motivated by this, we analyze how individual decisions drive final outcomes, decomposing the expected terminal improvement under a shared budget into cumulative decision opportunities minus cumulative selection losses. Guided by this opportunity-loss theoretical foundation, we propose OptiCom, a unified framework that represents LLM-driven optimizers within a shared configuration space: C=(A,Q,O,E,M,S), corresponding to artifact, query, operator, evaluation, memory, and strategy. Operating within this space, a fast LLM-based Optimization Controller dynamically composes immediate mechanisms through structured Action Packages, while a slower Strategy Adapter refines long-term selection preferences, operator weights, and templates based on accumulated trajectory feedback. Comprehensive evaluations across 32 benchmark groups demonstrate the superiority of framework: OptiCom achieves an average Max-score rank of 1.72 among 14 evaluated configurations, securing the top score in 23 groups. Ultimately, these results highlight the broad applicability and high extensibility of OptiCom as a general-purpose paradigm for robust LLM test-time scaling.
Chinese Translation
大语言模型(LLM)正越来越多地被部署,以通过迭代优化解决复杂的科学和实际问题。然而,随着候选质量、失败模式和资源预算的演变,动态协调多样的搜索机制仍然是一个关键的开放挑战。有针对性的实证诊断表明,机制有效性高度依赖于状态。受此启发,我们分析个体决策如何驱动最终结果,将共享预算下的期望终端改进分解为累积决策机会减去累积选择损失。在这一机会-损失理论基础指导下,我们提出 OptiCom,一个统一框架,它在共享配置空间 C=(A,Q,O,E,M,S) 中表示 LLM 驱动的优化器,该空间分别对应于产物、查询、算子、评估、记忆和策略。在该空间内运行时,一个快速的基于 LLM 的优化控制器通过结构化的动作包动态组合即时机制,而一个较慢的策略适配器则基于累积的轨迹反馈来细化长期选择偏好、算子权重和模板。跨 32 个基准组的全面评估证明了该框架的优越性:在 14 种被评估配置中,OptiCom 取得了 1.72 的平均 Max-score 排名,并在 23 个组中获得了最高分。最终,这些结果凸显了 OptiCom 作为稳健 LLM 测试时扩展的通用范式的广泛适用性和高度可扩展性。
cs.AI / 33 / 2609.37304
MetaCtrl: Your Large Language Models Can Reason Better and More Concisely with a Metacognitive Controller
MetaCtrl:借助元认知控制器,你的大型语言模型能够推理得更好且更简洁
large language model
大语言模型相关
Abstract
Large reasoning models improve performance on challenging problems by allocating additional computation before answering, but longer reasoning does not always lead to better results and can introduce substantial redundant reasoning on simple problems. Conversely, aggressively shortening reasoning can degrade performance on difficult ones. Effective reasoning therefore requires dynamically deciding when additional computation is useful based on the reasoner's capabilities and evolving solution state. Existing approaches often rely on predefined budgets or intervention rules, retrain the target reasoner, or require additional supervision. We introduce MetaCtrl, a lightweight controller that adaptively regulates a frozen reasoner without predefined token budgets or reasoner retraining. We formulate reasoning regulation as a sequential metacognitive control problem: MetaCtrl observes the evolving reasoning trace and decides whether to continue, simplify, skip redundant steps, or conclude reasoning. It is trained directly with reinforcement learning using a reward that prioritizes correctness while favoring shorter trajectories among correct solutions, requiring neither supervised intervention trajectories nor problem-specific budgets. Across seven benchmarks spanning mathematics, science, and code, MetaCtrl consistently improves the accuracy of LRMs while reducing their reasoning length. On DeepSeek-R1-Distill-Qwen-7B, it improves average accuracy by 4.7 points while reducing generation length by 53.3%. Without further training, the same controller transfers to an unseen reasoner (e.g., Qwen3-14B), improving average accuracy by 2.9 points and reducing generation length by 50.3%. These results establish MetaCtrl as a plug-and-play controller for improving reasoning accuracy while substantially reducing inference-time generation. The code is available at https://github.com/binbin2xs/MetaCtrl.
Chinese Translation
大型推理模型通过在回答前分配额外计算,提升了在挑战性问题上的性能,但更长的推理并不总是带来更好的结果,并且可能在简单问题上引入大量冗余推理。相反,激进地缩短推理可能会降低在困难问题上的性能。因此,有效的推理需要根据推理器的能力和不断演化的解题状态,动态决定何时额外计算是有用的。现有方法通常依赖预定义的预算或干预规则,重新训练目标推理器,或者需要额外的监督。我们提出 MetaCtrl,一种轻量级控制器,它无需预定义的 token 预算或对推理器进行重新训练,即可自适应地调控冻结的推理器。我们将推理性调控形式化为一个序列元认知控制问题:MetaCtrl 观察不断演化的推理轨迹,并决定是继续、简化、跳过冗余步骤,还是结束推理。它直接使用强化学习进行训练,所用奖励优先考虑正确性,同时在正确解中偏好更短的轨迹,既不要求有监督的干预轨迹,也不要求特定于问题的预算。在涵盖数学、科学和代码的七个基准上,MetaCtrl 持续提升 LRM 的准确率,同时减少其推理长度。在 DeepSeek-R1-Distill-Qwen-7B 上,它将平均准确率提高了 4.7 个百分点,同时将生成长度减少了 53.3%。无需进一步训练,同一个控制器可迁移到一个未见过的推理器(例如 Qwen3-14B),将平均准确率提高 2.9 个百分点,并将生成长度减少 50.3%。这些结果表明,MetaCtrl 是一种即插即用的控制器,可在提高推理准确率的同时大幅减少推理时生成。代码可在 https://github.com/binbin2xs/MetaCtrl 获取。
cs.AI / 34 / 2609.37362
Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing
预训练一次,随处路由:迈向用于 LLM 路由的基础模型
large language model
大语言模型相关
Abstract
Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--efficiency trade-off of LLM inference. Existing routers are typically learned through local fitting: a router is optimized for a particular query workload and candidate pool, and often requires additional supervision or retraining as the routing environment changes. We ask whether LLM routing can instead be approached from a foundation-model perspective, learning a reusable routing capability that generalizes across tasks, candidate models, and deployment conditions. To this end, we introduce RouteFM, which learns to characterize anonymous candidate models from behavioral context and infer their target-specific capabilities, rather than binding routing decisions to fixed model identities or a single environment. Through episodic pretraining across heterogeneous routing environments, this capability can be reused by a frozen router and adapted to new environments through context alone. Experiments demonstrate transfer across changes in domains, modalities, candidate pools, and context budgets, with the largest gains when behavioral evidence is limited. On MMR-Bench, which is excluded from pretraining, RouteFM outperforms the strongest baseline by 2.23 quality points with only eight observations per candidate. These results support moving LLM routing from repeated local fitting toward a pretrain once, route anywhere paradigm. Our code is publicly available at https://github.com/LAMDA-Model-Reuse/RouteFM.
Chinese Translation
大语言模型(LLM)路由旨在将每个查询分配给异构候选池中最合适的模型,从而改善 LLM 推理的质量--效率权衡。现有路由器通常通过局部拟合来学习:路由器针对特定查询工作负载和候选池进行优化,并且随着路由环境变化,往往需要额外的监督或重新训练。我们问,LLM 路由是否可以转而从基础模型视角来处理,学习一种可复用的路由能力,该能力能够泛化到不同任务、候选模型和部署条件。为此,我们提出了 RouteFM,它学习从行为上下文中刻画匿名候选模型,并推断它们针对特定目标的能力,而不是将路由决策绑定到固定的模型身份或单一环境。通过对异构路由环境进行情景式预训练,这种能力可以被一个冻结的路由器复用,并且仅通过上下文就能适应新环境。实验表明,在领域、模态、候选池和上下文预算发生变化时能够迁移,并且在行为证据有限时增益最大。在排除于预训练之外的 MMR-Bench 上,RouteFM 仅用每个候选的八次观测就比最强基线高出 2.23 个质量点。这些结果支持将 LLM 路由从重复的局部拟合转向“预训练一次,随处路由”的范式。我们的代码公开于 https://github.com/LAMDA-Model-Reuse/RouteFM。
cs.AI / 35 / 2609.37402
Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing
路由应当收回成本:面向经济型 LLM 路由的稀疏监督
large language model
大语言模型相关
Abstract
Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on historical queries to collect query--model quality feedback, creating a nontrivial supervision cost before deployment. Existing work largely focuses on serving-time efficiency, overlooking whether the resulting savings are sufficient to recover this upfront expenditure. We further observe that routing quality often saturates well before all query--model feedback is collected, suggesting that dense supervision can be economically over-provisioned. We propose SaveRouter, a sparse-supervision routing framework that selectively acquires informative model feedback and shares capability information across related queries, while retaining query-level refinement for fine-grained routing. We evaluate routing by jointly accounting for supervision expenditure and subsequent serving-time savings. Across four routing benchmarks, the main setting uses only about 33--41% of available training feedback while maintaining competitive or better routing quality, and reduces the break-even deployment volume by approximately 1.9--9.5 times compared with the fastest conventional router. Further analysis shows that acquiring more supervision is not always economically preferable: the supervision level that minimizes serving cost can differ from the one that achieves the earliest payback. Our code is publicly available at https://github.com/LAMDA-Model-Reuse/SaveRouter.
Chinese Translation
大语言模型(LLM)路由通过将每个查询分配给合适的模型,同时保持响应质量,从而降低服务成本。然而,学习这样一种路由器通常需要在历史查询上执行多个候选模型,以收集查询--模型质量反馈,从而在部署前产生不可忽视的监督成本。现有工作大多关注服务时的效率,而忽视了由此产生的节省是否足以收回这笔前期支出。我们进一步观察到,路由质量往往在收集完所有查询--模型反馈之前就已趋于饱和,这表明密集监督可能在经济上属于过度配置。我们提出 SaveRouter,一种稀疏监督路由框架,它有选择地获取信息量大的模型反馈,并在相关查询之间共享能力信息,同时保留查询级细化以实现细粒度路由。我们通过同时考虑监督支出和后续服务时的节省来评估路由。在四个路由基准上,主设置仅使用约 33--41% 的可用训练反馈,同时保持具有竞争力或更好的路由质量,并且与最快的传统路由器相比,将盈亏平衡部署量减少约 1.9--9.5 倍。进一步分析表明,获取更多监督并不总是在经济上更可取:使服务成本最小化的监督水平可能不同于实现最早回本的监督水平。我们的代码已公开在 https://github.com/LAMDA-Model-Reuse/SaveRouter。
cs.AI / 36 / 2609.37670
MeanFlowAdvantage: Stable Reward Fine-Tuning for Few-Step Average-Velocity Generators
MeanFlowAdvantage:面向少步平均速度生成器的稳定奖励微调
diffusion
扩散模型相关
Abstract
MeanFlow enables efficient few-step generation by predicting interval-average velocities, but this representation creates a mismatch for reward fine-tuning: existing advantage-based objectives are typically defined on instantaneous velocities or equivalent $x_0$-space predictions, whereas inference directly uses the learned average-velocity map. We introduce MeanFlowAdvantage, a signed advantage-weighted least-squares objective for average-velocity generators. Our key construction uses a shared, detached MeanFlow derivative correction to express the reward objective in prediction space while making rollout and reference regularization exact penalties on the average-velocity network deployed at inference. The resulting formulation preserves MeanFlow's native few-step sampler and provides a direct mechanism for transferring reward improvements to the deployed flow map. On SD3.5-Medium, MeanFlowAdvantage improves all eight reported metrics over the matched four-step MeanFlowNFT baseline and, with only four NFEs, matches or exceeds the 40-step DiffusionNFT baseline on six of eight metrics. The same objective also transfers to DNA promoter design, where it supports both teacher-free on-policy RL for a generator defined on a manifold and teacher-guided reward-graded distillation, with the latter yielding the lowest one-step Sei profile MSE among the compared configurations.
Chinese Translation
MeanFlow 通过预测区间平均速度实现了高效的少步生成,但这种表示方式为奖励微调带来了一种不匹配:现有的基于优势的目标通常定义在瞬时速度或等价的 $x_0$ 空间预测之上,而推理则直接使用所学习到的平均速度映射。我们提出了 MeanFlowAdvantage,一种面向平均速度生成器的带符号优势加权最小二乘目标。我们的关键构造使用一个共享的、detached 的 MeanFlow 导数校正,在预测空间中表达奖励目标,同时使 rollout 与参考正则化成为对推理时所部署的平均速度网络的精确惩罚项。由此得到的公式保留了 MeanFlow 原生的少步采样器,并提供了一种直接机制,用于将奖励上的改进迁移到所部署的流映射。在 SD3.5-Medium 上,MeanFlowAdvantage 在全部八项所报告的指标上均优于与之匹配的四步 MeanFlowNFT 基线,并且在仅使用四次 NFE 的情况下,在八项指标中的六项上达到或超过了 40 步 DiffusionNFT 基线。同一目标还可迁移到 DNA 启动子设计,在那里它同时支持针对定义在流形上的生成器的无教师 on-policy RL,以及教师引导的按奖励分级的蒸馏,其中后者在所比较的各配置中取得了一步 Sei profile MSE 的最低值。
cs.AI / 37 / 2609.37673
KUPAS MASTER: Distilling the Tacit Expertise of Master Practitioners into Agent-Ready Experience Corpora
KUPAS MASTER:将大师级实践者的隐性专长蒸馏为智能体就绪的经验语料库
large language model
大语言模型相关
Abstract
Experienced professionals know more than just facts and conclusions. They know which cues matter, why a judgment is reasonable, and which action to take. Routine work records often leave out this tacit knowledge, making it difficult for Large Language Model (LLM) agents to use professional experience effectively. We introduce KUPAS MASTER, an experience engineering platform built around nine-layer cognitive corpus construction. It turns heterogeneous work records and practitioner interviews into traceable, reusable experience corpora for agents. Six case elements preserve the task process: context, cues, judgment, action, boundaries, and outcomes. Nine-layer cognitive corpus construction organizes tacit experience along nine extraction dimensions and stores the resulting assets in six libraries: rules, constraints, best practices, negative examples, corner cases, and skills. Semantic alignment, individual experience distillation, organizational consolidation, and cross-review preserve source evidence, conditions of use, and unresolved disagreements. The platform packages these assets into callable skills with explicit inputs, steps, dependencies, and stopping conditions, connecting experience collection to task execution and evaluation feedback. Using authorized samples from 20 randomly selected practitioners, the platform processed 1,576 source files into 23,024 individual experience records and 13,113 organizational assets. The evaluation spans multiple professional domains. Under common task inputs and scoring criteria, the base model, raw corpus retrieval-augmented generation (RAG), and KUPAS MASTER agent scored 70.63, 79.75, and 89.58, respectively. The KUPAS MASTER agent improved on raw-corpus RAG in all seven scoring dimensions. The platform provides a practical path from individual tacit experience to organizational knowledge and agent capabilities.
Chinese Translation
经验丰富的专业人士知道的不只是事实和结论。他们知道哪些线索重要、为什么某一判断是合理的,以及应采取哪一行动。日常工作记录常常遗漏这种隐性知识,使大语言模型(LLM)智能体难以有效利用专业经验。我们提出 KUPAS MASTER,一个围绕九层认知语料库构建的经验工程平台。它将异构工作记录和从业者访谈转化为可供智能体使用的可追溯、可复用的经验语料库。六种案例要素保留任务过程:情境、线索、判断、行动、边界和结果。九层认知语料库构建沿九个提取维度组织隐性经验,并将所得资产存储于六个库中:规则、约束、最佳实践、负面示例、边角案例和技能。语义对齐、个体经验蒸馏、组织整合和交叉评审保留了来源证据、使用条件和未解决的分歧。该平台将这些资产打包为可调用的技能,具有明确的输入、步骤、依赖关系和停止条件,将经验收集与任务执行和评估反馈连接起来。使用来自 20 名随机抽取的从业者的授权样本,该平台将 1,576 个源文件处理为 23,024 条个体经验记录和 13,113 项组织资产。评估涵盖多个专业领域。在共同的任务输入和评分标准下,基础模型、原始语料库检索增强生成(RAG)和 KUPAS MASTER 智能体分别得分 70.63、79.75 和 89.58。KUPAS MASTER 智能体在全部七个评分维度上均优于原始语料库 RAG。该平台提供了一条从个体隐性经验到组织知识和智能体能力的实用路径。
cs.AI / 38 / 2609.37700
Locating Answer-Correctness Signals in Frozen Large Language Models
定位冻结大语言模型中的答案正确性信号
large language model
大语言模型相关
Abstract
Language models expose internal signals that predict whether an answer is correct, readable from a single forward pass of a frozen model without additional generations. Yet existing probes often commit to one signal family or layer and can be brittle under distribution shift; in retrieval-augmented settings, many specialized detectors instead target passage faithfulness, which can diverge from correctness when retrieved evidence is unhelpful or conflicting. We therefore ask where answer correctness is readable, which internal signal families carry it, and how they should be combined. We search over hidden states, token probabilities, residual-stream features, attention, and their fusion, treating the selected readouts as a predictive measurement rather than a mechanistic localization. We run this analysis separately in closed-book and with-context settings, since context can change which readouts are informative. A consistent anatomy emerges: correctness concentrates in the answer span, recovered from the answer tokens even under retrieval, and the families carry it complementarily, so fusing them helps most out of distribution, where a single signal is weakest. The protocol is effective across two backbones and gates a retrieval controller as one downstream use.
Chinese Translation
语言模型会暴露内部信号,用于预测一个答案是否正确;这些信号可从冻结模型的单次前向传播中读取,而无需额外生成。然而,现有探测方法往往固守于单一信号族或单一层,在分布偏移下可能表现脆弱;在检索增强设置中,许多专门检测器转而针对段落忠实性,而当检索到的证据无帮助或相互冲突时,这种忠实性可能与正确性发生偏离。因此,我们追问:答案正确性在何处可读,哪些内部信号族承载它,以及它们应如何被组合。我们在隐藏状态、词元概率、残差流特征、注意力及其融合上进行搜索,将所选的读出视为一种预测性测量,而非机制性定位。我们在闭卷和带上下文两种设置下分别运行这一分析,因为上下文会改变哪些读出具有信息量。一个一致的解剖图景浮现出来:正确性集中在答案片段中,即使在检索条件下也可从答案词元中恢复;这些信号族以互补方式承载它,因此融合它们对分布外场景帮助最大,而在那里单一信号最为薄弱。该协议在两个骨干模型上均有效,并作为一种下游用途对检索控制器进行门控。
cs.AI / 39 / 2609.37751
Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models
混合专家大语言模型中的交叉熵引导路由
large language model
大语言模型相关
Abstract
Sparse mixture-of-experts (MoE) large language models scale model capacity by routing each token to a small subset of experts. Their routers are regularized with load balancing terms and learn affinity scores through the language-model objective. However, these objectives do not provide direct alignment between routing affinities and token-level error. We introduce token-error supervision for sparse routing in two forms. The first form predicts an error score per expert. The affinity-weighted aggregate of these scores is aligned to the next-token cross-entropy loss, while the individual scores attenuate affinity before top-$K$ selection. The second directly aligns the router's affinities to the model's objective without requiring an additional head or inference-time modification. Both formulations use the Itakura--Saito divergence or an exponential negative log-likelihood for aligning affinities and token errors. Across two sparse MoE backbones and four multiple-choice question-answering benchmarks, we evaluate both supervision mechanisms. On Granite, our method improves accuracy by approximately 2.3 percentage points on average over a parameter-matched routing baseline. With stronger supervision, the gain on ARC-Challenge reaches 2.94 points. Both mechanisms preserve the native sparse execution budget and aggregation policy. Our code is available in the supplementary materials.
Chinese Translation
稀疏混合专家(MoE)大型语言模型通过将每个 token 路由到一小部分专家来扩展模型容量。它们的路由器通过负载均衡项进行正则化,并通过语言模型目标学习亲和度分数。然而,这些目标并未在路由亲和度与 token 级错误之间提供直接对齐。我们以两种形式为稀疏路由引入 token 错误监督。第一种形式为每个专家预测一个错误分数。这些分数的亲和度加权聚合与下一 token 交叉熵损失对齐,而各个分数在 top-$K$ 选择之前衰减亲和度。第二种直接将路由器的亲和度与模型目标对齐,而不需要额外的头或推理时修改。两种公式都使用 Itakura--Saito 散度或指数负对数似然来对齐亲和度和 token 错误。在两个稀疏 MoE 主干网络和四个多项选择问答基准上,我们评估了这两种监督机制。在 Granite 上,我们的方法相较于参数匹配的路由基线平均将准确率提高约 2.3 个百分点。在更强的监督下,ARC-Challenge 上的增益达到 2.94 个百分点。两种机制均保留原生的稀疏执行预算和聚合策略。我们的代码可在补充材料中获得。
cs.AI / 40 / 2609.37829
DIET: Deletion-response Expert Trimming for Video Diffusion Transformers
DIET:面向视频扩散 Transformer 的删除响应专家剪枝
diffusionlarge language model
扩散模型相关
大语言模型相关
Abstract
Video diffusion transformers (DiTs) increasingly adopt mixture-of-experts (MoE) architectures to reduce active computation, but their full expert storage remains costly. Existing one-shot pruning criteria mainly rely on static activation or routing statistics and cannot capture layer-level re-routing after expert deletion. We introduce DIET, a training-free expert pruning framework based on deletion responses. A single all-expert calibration pass records expert outputs and router states for matched conditional and unconditional tokens. Candidate deletions are then replayed from cached tensors, requiring no additional model forward passes. The resulting deletion-response signatures characterize each expert by the changes induced when it is removed. DIET selects retained experts by minimizing Overall Diversity Loss (ODL), which preserves directional coverage in signature space, and combines intra-layer local search with an inter-layer regression-guided budget search to allocate experts across layers. On LingBot-Video 30B-A3B, pruning 50% of experts (6,144 to 3,072) reduces the checkpoint from 57 GB to 30 GB and enables single-card deployment on a 48 GB GPU without fine-tuning. Under a fixed 284-case VBench protocol, the VBench Total increases from 0.7941 to 0.8115. Across tested retention budgets, DIET consistently outperforms competitive pruning baselines adapted from large language models.
Chinese Translation
视频扩散 Transformer(DiT)越来越多地采用混合专家(MoE)架构来降低激活计算量,但它们的完整专家存储仍然代价高昂。现有的一次性剪枝准则主要依赖静态激活或路由统计,无法捕捉专家删除后的层级重路由。我们提出 DIET,一个基于删除响应的免训练专家剪枝框架。单次全专家校准前向即可记录匹配的条件 token 与无条件 token 的专家输出和路由器状态。随后,候选删除从缓存张量中重放,无需额外的模型前向传播。由此得到的删除响应签名通过移除某专家时所引起的改变来刻画每个专家。DIET 通过最小化总体多样性损失(ODL)来选择保留的专家,该损失在签名空间中保持方向覆盖度,并将层内局部搜索与层间回归引导的预算搜索相结合,以在各层之间分配专家。在 LingBot-Video 30B-A3B 上,剪除 50% 的专家(从 6,144 个减至 3,072 个)将检查点从 57 GB 缩减到 30 GB,并使得无需微调即可在 48 GB GPU 上实现单卡部署。在固定的 284 个案例的 VBench 协议下,VBench Total 从 0.7941 提高到 0.8115。在测试的各个保留预算下,DIET 始终优于从大语言模型适配而来的竞争性剪枝基线。
cs.AI / 41 / 2609.37857
Active Budget Can Kill Sensitivity: Diagnosing and Repairing TopK Sparse Autoencoder Reliability
活跃预算会扼杀敏感性:诊断与修复TopK稀疏自编码器的可靠性
large language model
大语言模型相关
Abstract
Sparse autoencoders (SAEs) are increasingly scaled to wider dictionaries to recover fine-grained structure from large language model activations. However, a feature is useful for interpretation only if it remains a stable unit of analysis when the same meaning is expressed in different surface forms. We study this reliability question for TopK SAEs via feature sensitivity. Experiments demonstrate that scaling selectively reduces the sensitivity of rare features, while common features remain comparatively stable. A controlled width\(\times k\) factorial experiment identifies the active budget k as the root cause: the degradation arises from the selection boundary rather than dictionary width alone. We attribute this failure to the geometry of TopK selection. The active margin, the distance to the cutoff, predicts feature loss without thresholds. Guided by this margin diagnosis, we introduce pairwise rank stabilization. Our method targets ordering failures at the cutoff and improves rare-feature sensitivity by \(8.83\) percentage points, while keeping reconstruction and alive-feature coverage near the baseline. Overall, our results suggest that wide TopK SAEs should be evaluated not only by reconstruction, sparsity, and feature count, but also by feature reliability under semantic variation and boundary geometry for stable interpretability.
Chinese Translation
稀疏自编码器(SAE)正日益被扩展到更宽的字典,以从大语言模型的激活中恢复细粒度结构。然而,只有当同一含义以不同表层形式表达时某个特征仍保持为稳定的分析单元,它才可用于解释。我们通过特征敏感性来研究TopK SAE的这一可靠性问题。实验表明,扩大规模会选择性地降低稀有特征的敏感性,而常见特征则保持相对稳定。一个受控的宽度\(\times k\)析因实验确定活跃预算k是根本原因:这种退化源于选择边界,而非仅由字典宽度造成。我们将这一失败归因于TopK选择的几何结构。活跃边际,即到截断值的距离,无需阈值即可预测特征损失。在这一边际诊断的指导下,我们引入了成对秩稳定化。我们的方法针对截断处的排序失败,将稀有特征敏感性提升了\(8.83\)个百分点,同时使重建与活跃特征覆盖率保持在基线附近。总体而言,我们的结果表明,宽TopK SAE不仅应通过重建、稀疏性和特征数量来评估,还应通过语义变化下的特征可靠性以及边界几何结构来评估,以实现稳定的可解释性。
cs.AI / 42 / 2609.37875
Co-PiLOT: Constrained Physics-Informed Latent Optimization for Target-Driven Inverse Design
Co-PiLOT:面向目标驱动逆向设计的约束物理信息潜空间优化
diffusion
扩散模型相关
Abstract
Inverse design of physical systems (molecules, devices, microstructures) often reduces to optimizing a high-dimensional structure against an expensive black-box simulator. Direct search is difficult because the space is non-Euclidean, feasibility is hard to encode, and each evaluation is expensive. We present Co-PiLOT, a latent optimization approach that maps candidates through a generative encoder-decoder, uses the decoder as a learned validity prior, and searches the latent space with physics-informed black-box optimization. The framework is applied on the inverse design of magnesium alloy microstructure/texture. We develop a vision transformer based-encoder; paired with latent diffusion, diffusion transformer and rectified-flow transformer-based decoders on $\sim80{,}000$ EBSD-derived microstructure dataset to learn a minimal bottleneck, $z$. The ViT-FMDiT model ($z$=$768$) reconstructs high-fidelity microstructure images (FID $27.86$, MS-SSIM $0.178$), which our self-segmenting orientation codec converts into input grids for crystal plasticity solver. Finally, we introduce MERIDIAN, an active latent optimizer driven by deep-kernel Gaussian-process uncertainty, failure-aware feasibility prediction, manifold-aware trust regions, and target-aware acquisition. Within a budget of $160$ simulations, the ViT-FMDiT and MERIDIAN combination yields the best target-driven objective score, reducing the relative target error by $3$--$22\%$ against seven baselines (DANTE, TuRBO, BAxUS, CMA-ES, DDOM, SEIKO, DDPO) on the same decoder.
Chinese Translation
物理系统的逆向设计(分子、器件、微结构)通常可归结为针对昂贵的黑箱模拟器优化高维结构。直接搜索很困难,因为该空间是非欧几里得的,可行性难以编码,而且每次评估都代价高昂。我们提出了 Co-PiLOT,这是一种潜空间优化方法,它通过生成式编码器-解码器映射候选对象,将解码器用作学习得到的有效性先验,并使用物理信息黑箱优化搜索潜空间。该框架应用于镁合金微结构/织构的逆向设计。我们开发了一个基于视觉 transformer 的编码器;在 $\sim80{,}000$ 个由 EBSD 导出的微结构数据集上,将其与基于潜扩散、扩散 transformer 和整流流 transformer 的解码器配对,以学习一个最小瓶颈 $z$。ViT-FMDiT 模型($z$=$768$)重建高保真微结构图像(FID $27.86$,MS-SSIM $0.178$),我们的自分割取向编解码器将其转换为晶体塑性求解器的输入网格。最后,我们介绍了 MERIDIAN,这是一种主动潜空间优化器,由深度核高斯过程不确定性、失败感知可行性预测、流形感知信任域和目标感知采集驱动。在 $160$ 次模拟的预算内,ViT-FMDiT 与 MERIDIAN 的组合取得了最佳的目标驱动目标得分,在相同解码器上相对于七个基线(DANTE、TuRBO、BAxUS、CMA-ES、DDOM、SEIKO、DDPO)将相对目标误差降低了 $3$--$22\%$。
cs.AI / 43 / 2609.37956
BrainNet Studio: A Unified Toolkit for Brain Network Construction, Intelligent Analysis, and Visualization
BrainNet Studio:用于脑网络构建、智能分析与可视化的统一工具包
large language model
大语言模型相关
Abstract
Brain networks characterize structural and functional relationships among brain regions and support research on cognition, brain disorders, and brain-computer interfaces. Their time-varying topology and higher-order spatiotemporal dependencies are not adequately represented by conventional static networks. Existing tools primarily focus on static connectomes and provide limited integration of dynamic network modeling with modern graph and sequence learning methods. We present BrainNet Studio, an integrated toolkit for static and dynamic brain network analysis. It provides a unified workflow encompassing network construction, feature extraction, predictive modeling, candidate biomarker identification, visualization, and assisted interpretation. The toolkit integrates 27 algorithms, including deep learning, graph neural networks, and spatiotemporal sequence models, to support classification and the identification of discriminative brain regions and connections. A large language model generates researcher-verifiable summaries of functional connectivity, structural connectivity, and structure-function coupling at individual and group levels. Within a consistent computational framework, users can configure analytical tasks, compare methods, inspect outputs, and extend functionality without repeatedly assembling application-specific pipelines. BrainNet Studio provides a practical and extensible platform for connectome analysis in cognitive neuroscience, exploratory studies of brain disorders, and brain-computer interfaces. The toolkit is publicly available at https://github.com/xbrainnet/Brainnet-Studio.
Chinese Translation
脑网络刻画脑区之间的结构和功能关系,并支持认知、脑疾病和脑机接口研究。其随时间变化的拓扑结构和高阶时空依赖关系无法由传统静态网络充分表示。现有工具主要关注静态连接组,并且对动态网络建模与现代图学习和序列学习方法的整合有限。我们提出 BrainNet Studio,一个用于静态和动态脑网络分析的一体化工具包。它提供了一个统一工作流,涵盖网络构建、特征提取、预测建模、候选生物标志物识别、可视化和辅助解释。该工具包集成了 27 种算法,包括深度学习、图神经网络和时空序列模型,以支持分类以及判别性脑区和连接的识别。一个大语言模型生成可由研究者验证的、关于个体和群体水平功能连接、结构连接和结构-功能耦合的总结。在一致的计算框架内,用户可以配置分析任务、比较方法、检查输出并扩展功能,而无需反复组装特定应用的流程。BrainNet Studio 为认知神经科学中的连接组分析、脑疾病的探索性研究和脑机接口提供了一个实用且可扩展的平台。该工具包公开可用,网址为 https://github.com/xbrainnet/Brainnet-Studio。
cs.AI / 44 / 2609.38005
Diagnosing and Improving Probabilistic Reasoning in Large Language Models
诊断并改进大语言模型中的概率推理
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly proposed as decision assistants who must reason probabilistically from available evidence under explicit decision costs. We propose a decision-theoretic framework that decomposes LLMs' decision loss into two components: forming accurate beliefs from provided evidence and translating those beliefs into actions that optimize a provided utility function. Using a synthetic benchmark with known ground truth, we apply the decomposition to characterize probabilistic reasoning in frontier and open-sourced models. We further evaluate whether RL interventions targeting beliefs, decisions, or both improve these components across three domains, whether improvements transfer across components and elicitation formats, and whether decision performance can improve without improvement in belief formation. We find that targeting one component of probabilistic reasoning redistributes decision loss, improving the target without necessarily transferring to others, and that jointly targeting belief formation and decision-making improves both but hinges on matched formats between training and evaluation.
Chinese Translation
大语言模型(LLMs)正越来越多地被提议作为决策助手,这些助手必须在明确的决策成本下,根据可用证据进行概率推理。我们提出了一个决策论框架,该框架将 LLMs 的决策损失分解为两个组成部分:从所提供的证据中形成准确的信念,以及将这些信念转化为优化给定效用函数的行动。使用一个具有已知真值的合成基准,我们应用该分解来刻画前沿模型和开源模型中的概率推理。我们进一步评估,针对信念、决策或两者的强化学习干预是否能在三个领域中改善这些组成部分,改进是否会在不同组成部分和引出格式之间迁移,以及决策性能是否能在信念形成没有改善的情况下得到提升。我们发现,针对概率推理中的某一个组成部分会重新分配决策损失,改善目标组成部分,但不一定会迁移到其他组成部分;同时针对信念形成和决策制定会改善两者,但这取决于训练和评估之间格式的匹配。
cs.AI / 45 / 2609.38070
Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs
概率并不足够:探索并计数发散 Token 以实现大语言模型中的推理不确定性量化
large language model
大语言模型相关
Abstract
As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertainty of their answers. Current methods for estimating the confidence of large language models are generally based on probabilities of selected key tokens, but the underlying mechanism remains unclear. Our pilot study finds that replacing selected token probabilities with coarse substitutes can also improve calibration, motivating us to further explore effective signals of model confidence. We introduce Divergent Token Confidence (DTC), a framework that estimates confidence by counting tokens at which two models strongly disagree during decoding. DTC identifies these divergent tokens using the Jensen-Shannon divergence between next-token distributions evaluated along the same reasoning trajectory. We find that their count is almost negatively associated with answer accuracy, thereby serving as a simple yet effective signal for uncertainty quantification. DTC supports both white-box and black-box evaluation using auxiliary models, without explicit training and affecting the generation process. Experiments across multiple model families and six mathematical benchmarks demonstrate improved calibration over probability-based and verbalized baselines. Under white-box evaluation, the count-only estimator achieves an average expected calibration error of 13.0%, compared with 32.7%-42.4% for standard full-sequence confidence methods. In black-box settings, it also improves calibration over the original verbalized scores. For example, mean expected calibration error falls from 32.1%-40.2% to 13.7%-16.3% on DeepSeek-V3.2. These findings provide new insights for improving reasoning uncertainty quantification in large language models. The code is released at https://github.com/szu-tera/DTC.git.
Chinese Translation
随着大语言模型的思维链推理能力不断提高,评估和校准其推理置信度对于量化其答案的不确定性正变得越来越重要。当前用于估计大语言模型置信度的方法通常基于所选关键 Token 的概率,但其底层机制仍不清楚。我们的初步研究发现,用粗略替代物替换所选 Token 概率也能改善校准,这促使我们进一步探索模型置信度的有效信号。我们提出发散 Token 置信度(Divergent Token Confidence,DTC),这是一个通过计数两个模型在解码过程中强烈不一致的 Token 来估计置信度的框架。DTC 使用沿同一推理轨迹评估的下一 Token 分布之间的 Jensen-Shannon 散度来识别这些发散 Token。我们发现它们的计数几乎与答案准确率呈负相关,因此可作为不确定性量化的一个简单而有效的信号。DTC 支持使用辅助模型进行白盒和黑盒评估,无需显式训练且不影响生成过程。在多个模型系列和六个数学基准上的实验表明,其校准优于基于概率的和言语化的基线。在白盒评估下,仅计数估计器实现了 13.0% 的平均期望校准误差,而标准全序列置信度方法为 32.7%-42.4%。在黑盒设置中,它也改善了相对于原始言语化分数的校准。例如,在 DeepSeek-V3.2 上,平均期望校准误差从 32.1%-40.2% 降至 13.7%-16.3%。这些发现为改进大语言模型中的推理不确定性量化提供了新见解。代码发布于 https://github.com/szu-tera/DTC.git。
cs.AI / 46 / 2609.38108
Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution
LLM 智能体会执行它们所声明的计划吗?从规划模式声明到模式特定执行
large language model
大语言模型相关
Abstract
Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and executing it faithfully. Existing planner--executor systems can fail at either stage, while final task success alone cannot distinguish selection from execution failures. We therefore study the Plan Declaration--Execution Gap and introduce Planning-as-Routing, where an LLM declares one of four planning modes: Predefined, Sequential, Hierarchical, or Search, and a deterministic router dispatches the task to the corresponding pattern-specific executor. Across four benchmarks and three LLMs, we find three consistent patterns. First, generic Plan+ReAct often fails to preserve declared planning structure, especially for longer plans: across three benchmarks, only (22)--(45%) of trajectories preserve it, whereas pattern-specific executors enforce the intended structure. Second, planning-mode effectiveness varies across environments and models: Search performs best on ALFWorld, Hierarchical on SWE-bench, and the strongest pattern can vary across models within the same benchmark. Third, the largest gains come from execution: pattern-specific executors improve task success from (0.48) to (0.92) on ALFWorld and from (0.36) to (0.44) on SWE-bench Verified over Plan+ReAct. Current LLMs, however, do not reliably select the strongest mode for each task, although few-shot examples improve selection in some benchmark--model combinations. Overall, reliable agent planning requires both effective mode selection and faithful execution: routing substantially closes the execution gap, while task-specific mode selection remains open.
Chinese Translation
大语言模型(LLM)使智能体能够通过生成计划并在环境中执行该计划来解决长时程任务。然而,成功的规划需要两种不同的能力:为任务选择适当的计划并忠实地执行它。现有的规划器--执行器系统可能在任一阶段失败,而仅凭最终任务成功无法区分选择失败与执行失败。因此,我们研究计划声明--执行差距,并引入规划即路由(Planning-as-Routing),其中 LLM 声明四种规划模式之一:预定义(Predefined)、顺序(Sequential)、分层(Hierarchical)或搜索(Search),并且确定性路由器将任务分派给相应的模式特定执行器。在四个基准和三个 LLM 上,我们发现三种一致的模式。首先,通用的 Plan+ReAct 往往无法保持所声明的规划结构,尤其是对于较长的计划:在三个基准上,只有 (22)--(45%) 的轨迹保持该结构,而模式特定执行器会强制执行预期结构。其次,规划模式的有效性因环境和模型而异:搜索在 ALFWorld 上表现最好,分层在 SWE-bench 上表现最好,而在同一基准内,最强模式可能因模型而异。第三,最大的收益来自执行:与 Plan+ReAct 相比,模式特定执行器在 ALFWorld 上将任务成功率从 (0.48) 提高到 (0.92),在 SWE-bench Verified 上从 (0.36) 提高到 (0.44)。然而,当前的 LLM 并不能可靠地为每个任务选择最强模式,尽管少样本示例在一些基准--模型组合中改进了选择。总体而言,可靠的智能体规划既需要有效的模式选择,也需要忠实的执行:路由大幅缩小了执行差距,而任务特定的模式选择仍然有待解决。
cs.CL / 47 / 2609.35970
Causal and Interpretable Structures in LLM Compositional Tasks
LLM 组合任务中的因果与可解释结构
large language model
大语言模型相关
Abstract
Large language models are able to solve tasks whose answers depend on not only individual input tokens, but also on relations among them. How is such relational information represented and processed across transformer layers? We study activations from ensembles of prompts that require inferring relationships between three tokens corresponding to a cyclic concept (months, hours, weekdays, and musical notes) to correctly predict the next token. Across model families (Llama, Qwen, Gemma, and Mistral) and cyclic concepts, we find a consistent layerwise progression in how the joint dependence among the tokens is geometrically organized and causally used: intermediate layers use a joint representation based on the inferred relationship between two tokens, while later layers use a joint representation associated with all three tokens to correctly complete the task. We also find other relationships between tokens that are geometrically structured but remain causally inert in the next-token prediction. Crucially, when taken together, these geometric and causal investigations reveal the representation-level mechanism that progressively organizes and composes the relational information to form the answer. More surprisingly, restricting the models to such causally relevant joint representations improves next-token prediction accuracy.
Chinese Translation
大语言模型能够解决这样一些任务:其答案不仅取决于单个输入 token,还取决于它们之间的关系。此类关系信息如何在 transformer 各层中被表示和处理?我们研究了来自提示集合的激活,这些提示要求推断对应于一个循环概念(月份、小时、星期几和音符)的三个 token 之间的关系,以正确预测下一个 token。在不同模型系列(Llama、Qwen、Gemma 和 Mistral)和循环概念中,我们发现,token 之间的联合依赖关系在几何上如何组织以及如何被因果使用方面存在一致的逐层演进:中间层使用基于两个 token 之间推断关系的联合表示,而较后的层使用与所有三个 token 都相关的联合表示来正确完成任务。我们还发现 token 之间的其他关系,这些关系在几何上是有结构的,但在下一个 token 预测中保持因果惰性。关键的是,当将这些几何和因果研究结合起来时,它们揭示了表示层面的机制,该机制逐步组织和组合关系信息以形成答案。更令人惊讶的是,将模型限制为这种因果相关的联合表示,会提高下一个 token 预测的准确率。
cs.CL / 48 / 2609.36139
Language Models Are "Insecure" Reporters
语言模型是“不安全”的报告者
large language model
大语言模型相关
Abstract
As large language models are deployed in increasingly autonomous long-horizon tasks, manually auditing and verifying the actions, artifacts, and outputs of models becomes more difficult. Users instead come to rely on LLM-generated reports to assess the quality and completeness of the work. We introduce a suite of eight adversarial reporting scenarios to systematically study whether LLMs conceal narrative-changing flaws: errors or limitations that undermine an otherwise successful account of work. We call this phenomenon "insecure reporting." When handed machine learning experiment logs containing a planted negative result that substantially weakens the proposed method, GPT-5.5 flags the negative result in only 2 of 200 generated reports. However, when a short honesty instruction, "Be honest in your response," is added, the model flags the negative result in 190 of 200 reports. Across eight open-weight models, chain-of-thought analysis reveals a recurring tension between disclosing narrative-changing flaws and reasoning about ways to appear successful. We perform an activation analysis and a steering experiment on Qwen3.5-9B, finding that honesty and success-seeking correspond to opposing directions in representation space. Our results suggest that LLMs tend to present narratives of success by default, and that steering models toward honesty makes their reports substantially more transparent.
Chinese Translation
随着大语言模型被部署到日益自主的长时程任务中,人工审计和验证模型的动作、工件与输出变得更加困难。用户转而依赖 LLM 生成的报告来评估工作的质量和完整性。我们引入一套包含八个对抗性报告场景的测试套件,以系统地研究 LLM 是否会隐藏改变叙事的缺陷:那些会削弱原本成功的工作叙述的错误或局限。我们将这一现象称为“不安全报告”。当拿到包含一个被植入的负面结果的机器学习实验日志时——该负面结果会显著削弱所提出的方法——GPT-5.5 在 200 份生成的报告中只有 2 份标记了该负面结果。然而,当加入一条简短的诚实性指令“在你的回答中保持诚实”时,该模型在 200 份报告中有 190 份标记出了该负面结果。在八个开放权重模型中,思维链分析揭示出一种反复出现的张力:一方面要披露改变叙事的缺陷,另一方面要推理如何显得成功。我们在 Qwen3.5-9B 上进行激活分析和引导实验,发现诚实与追求成功对应于表示空间中的相反方向。我们的结果表明,LLM 默认倾向于呈现成功的叙事,而将模型引导向诚实会使其报告显著更加透明。
cs.CL / 49 / 2609.36178
Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning
面向智能体强化学习中信用分配的关键决策定位
large language model
大语言模型相关
Abstract
Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success. We introduce ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning. Given a rollout group, an agentic judge contrasts successful and failed trajectories to propose a segment potentially responsible for their divergent outcomes. Rather than directly trusting the judge's assessment, ProVer verifies the proposed segment by estimating its advantage from the difference in terminal success rates between current-policy continuations sampled before and after the segment. Positive estimates are then incorporated into the GRPO advantages of policy tokens within the proposed segment. By using model judgment only to select where to verify, ProVer grounds local credit in observed outcomes without exhaustively evaluating every intermediate state. Across ALFWorld, WebShop, and SearchQA, ProVer achieves the strongest average performance at both model scales, with relative improvements over GRPO of 9.91% and 7.12% for Qwen3.5-2B and Qwen3.5-4B, respectively. Further analyses demonstrate that informed segment selection improves policy training with modest additional generation overhead, even without a frontier-scale judge model, highlighting the effectiveness and efficiency of selectively targeting pivotal decisions for fine-grained credit assignment in agentic reinforcement learning.
Chinese Translation
组相对策略优化(Group Relative Policy Optimization, GRPO)已成为训练大语言模型智能体的一种颇具前景的方法。然而,它将轨迹级优势统一分配给所有策略词元,无法将有后果的决策与不太相关的决策区分开来,从而模糊了哪些中间决策对成功有所贡献。我们提出 ProVer,一个在智能体强化学习中针对潜在关键决策进行细粒度信用分配的框架。给定一组 rollout,一个智能体评判器对比成功与失败的轨迹,以提出一个可能对其结果分歧负责的片段。ProVer 并不直接采信评判器的评估,而是通过估计该片段之前与之后采样的当前策略续接在最终成功率上的差异,来验证所提出的片段。正的优势估计随后被并入所提出片段内策略词元的 GRPO 优势中。通过仅使用模型判断来选择在何处进行验证,ProVer 将局部信用建立在观测到的结果之上,而无需穷尽地评估每一个中间状态。在 ALFWorld、WebShop 和 SearchQA 上,ProVer 在两个模型规模上均取得了最强的平均性能,相较于 GRPO,在 Qwen3.5-2B 和 Qwen3.5-4B 上分别取得了 9.91% 和 7.12% 的相对提升。进一步的分析表明,即使没有前沿规模的评判模型,有依据的片段选择也能以适度的额外生成开销改进策略训练,凸显了在智能体强化学习中选择性地针对关键决策进行细粒度信用分配的有效性与高效性。
cs.CL / 50 / 2609.36202
FastGuide: Accelerating Reward Guidance for Diffusion Large Language Models
FastGuide:加速扩散大语言模型的奖励引导
diffusionlarge language model
扩散模型相关
大语言模型相关
Abstract
Gradient-based reward guidance provides a flexible way to use downstream reward models to control masked diffusion language models at inference time. However, its computational cost remains high as each decoding iteration incurs expensive diffusion model forward passes and reward model backpropagation steps. To address this, we introduce FastGuide, an adaptive hybrid of parallel and autoregressive decoding to accelerate reward guidance for diffusion language models. In analogy to parallel decoding, FastGuide amortizes the cost of reward model backpropagation by computing guidance once per decoding step and reusing it to generate multiple tokens. Within each decoding step, FastGuide makes diffusion forward passes autoregressive by unmasking tokens one at a time while efficiently recomputing token distributions after each unmasking by utilizing KV caching techniques and sparse recomputation of attention. Lastly, to adapt hybrid decoding to the model's confidence, FastGuide defers any token that the model is unconfident about under its recomputed distribution. Experiments on three reward benchmarks demonstrate that FastGuide is up to $4.4\times$ faster than sequential reward-guided decoding while retaining similar generation quality.
Chinese Translation
基于梯度的奖励引导提供了一种灵活的方式,可在推理时利用下游奖励模型来控制掩码扩散语言模型。然而,其计算成本仍然很高,因为每次解码迭代都会带来昂贵的扩散模型前向传播和奖励模型反向传播步骤。为了解决这一问题,我们提出了 FastGuide,一种并行解码与自回归解码的自适应混合方法,用以加速扩散语言模型的奖励引导。与并行解码类似,FastGuide 通过在每个解码步只计算一次引导,并将其复用以生成多个 token,从而分摊奖励模型反向传播的成本。在每个解码步内,FastGuide 通过一次只解掩一个 token,使扩散前向传播变为自回归式的,同时利用 KV 缓存技术和注意力的稀疏重计算,在每次解掩后高效地重新计算 token 分布。最后,为了使混合解码适应模型的置信度,FastGuide 会推迟处理那些模型在其重计算分布下不自信的 token。在三个奖励基准上的实验表明,FastGuide 比顺序奖励引导解码最高快 $4.4\times$,同时保持相近的生成质量。
cs.CL / 51 / 2609.36209
The Canonical Order Problem: When Large Language Models Are Unreliable Knowledge Bases for Multi-Valued Relations
规范顺序问题:当大型语言模型作为多值关系知识库不可靠时
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly used as knowledge bases (KBs) due to the vast amount of knowledge they acquire during pre-training. While many works focus on extracting single relational triples, most real-world relations are multi-valued and require generating sets of entities. In this paper, we investigate how LLMs represent and generate multi-valued relations. We identify the canonical order problem: The probabilistic distributions inside LLMs organize many multi-valued relations according to a canonical ordering (e.g., alphabetical or chronological). Through mechanistic analysis, we show that set generation in LLMs can be thought of in terms of three phases: (1) retrieval of candidate entities, (2) internal sorting, and (3) selection of the next element. As a result, prompts aiming to construct KBs that deviate from this internal canonical ordering lead to a markedly reduced reliability of LLMs when aiming to generate complete sets for multi-valued relations.
Chinese Translation
大型语言模型(LLMs)由于在预训练期间获得了大量知识,正越来越多地被用作知识库(KBs)。尽管许多工作关注抽取单个关系三元组,但大多数现实世界的关系是多值的,并且需要生成实体集合。在本文中,我们研究 LLMs 如何表示和生成多值关系。我们识别出规范顺序问题:LLMs 内部的概率分布按照一种规范顺序(例如字母顺序或时间顺序)来组织许多多值关系。通过机制分析,我们表明 LLMs 中的集合生成可以从三个阶段来理解:(1)候选实体检索,(2)内部排序,以及(3)下一个元素的选择。因此,旨在构建偏离这种内部规范顺序的知识库的提示,会导致 LLMs 在旨在为多值关系生成完整集合时可靠性显著降低。
cs.CL / 52 / 2609.36214
Lost in Translation: Measuring the Effect of Non-Native English on End User Performance of Large Language Models
迷失在翻译中:衡量非母语英语对大语言模型最终用户表现的影响
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly used by people whose first language is not English, yet these users have been shown to receive systematically lower-quality responses than fluent speakers. Which specific features of non-native English drive this gap remains unclear, because fluency is itself a composite of mechanical accuracy, vocabulary use, organization, and discourse coherence. Here, we introduce FABLE, a controlled dataset of 190,911 English prompt variants derived from 174K real user prompts for writing-related tasks. Evaluating responses from 34 open-weight LLMs, we find a clear asymmetry; while models do not propagate surface errors such as misspellings into their outputs, models do mirror higher-level rhetorical and lexical qualities present in the user's prompt. Further, the overall quality of responses differs substantially between the least- and most-fluent prompts. These results highlight a key LLM performance disparity for non-native English LLM users, resulting in both lower-quality and less-fluent answers.
Chinese Translation
大语言模型(LLM)正越来越多地被第一语言不是英语的人使用,然而已有研究表明,这些用户相比流利使用者会系统性地收到质量更低的回答。非母语英语的哪些具体特征导致了这一差距仍不清楚,因为流利程度本身就是机械准确性、词汇使用、组织结构和语篇连贯性的复合体。在此,我们介绍了 FABLE,一个包含 190,911 个英语提示变体的受控数据集,这些变体源自 174K 条用于写作相关任务的真实用户提示。在评估来自 34 个开放权重 LLM 的回答后,我们发现了一种明显的不对称性;虽然模型不会将诸如拼写错误之类的表层错误传播到其输出中,但模型确实会镜像用户提示中存在的更高层次的修辞和词汇特征。此外,回答的总体质量在最不流利和最流利的提示之间存在显著差异。这些结果凸显了非母语英语 LLM 用户所面临的一个关键 LLM 表现差异,导致他们得到既质量更低又更不流利的答案。
cs.CL / 53 / 2609.36218
CineSubBench: Evaluating LLMs on Long-Form Narrative and Cultural Understanding from Multilingual Movie Subtitles
CineSubBench:评估LLMs基于多语言电影字幕的长篇叙事与文化理解
large language model
大语言模型相关
Abstract
Large language models are increasingly evaluated in specialized domains such as law, medicine, software engineering, and cybersecurity, yet film remains comparatively underexplored despite requiring long-form narrative integration, multilingual interpretation, and culturally situated audience judgments. We introduce CineSubBench, a benchmark for evaluating long-context film understanding from multilingual movie subtitles. A subtitle track represents a film as thousands of short, temporally ordered utterances from which models must reconstruct characters, relationships, events, causal progression, and themes without explicit scene or event structure. CineSubBench contains 1,012 films with complete subtitle coverage in six languages, yielding 6,072 tracks and 8.13M timestamped subtitle entries. It provides a matched multi-task, multilingual, and multicultural (MultiX) evaluation setting: seven tasks span narrative reconstruction and abstraction, genre prediction, age suitability, country-specific motion-picture ratings across ten national classification systems, and subtitle-grounded language safety. Across nine LLMs, plot premises are recovered more reliably than event-complete synopses; cross-lingual consistency varies substantially across models and languages; national rating systems expose distinct calibration patterns; and strong profanity is far easier to ground than mild obscenity. CineSubBench establishes film as a long-context LLM evaluation domain and provides a unified benchmark for measuring narrative, multilingual, cultural, and evidence-grounding capabilities.
Chinese Translation
大语言模型正越来越多地在法律、医学、软件工程和网络安全等专业领域接受评估,然而,尽管电影需要长篇幅叙事整合、多语言解释以及基于文化的受众判断,其在很大程度上仍相对未得到充分探索。我们提出 CineSubBench,一个用于评估基于多语言电影字幕的长上下文电影理解的基准。一条字幕轨道将一部电影表示为数千条简短、按时间排序的话语,模型必须从中重建角色、关系、事件、因果进展和主题,而没有明确的场景或事件结构。CineSubBench 包含 1,012 部在六种语言中具有完整字幕覆盖的电影,产生 6,072 条字幕轨道和 8.13M 条带时间戳的字幕条目。它提供了一个匹配的多任务、多语言和多文化(MultiX)评估设置:七项任务涵盖叙事重构与抽象、类型预测、年龄适宜性、跨越十个国家分类系统的国别电影评级,以及基于字幕的语言安全。在九个 LLMs 上,情节前提比事件完整的剧情梗概被更可靠地恢复;跨语言一致性在不同模型和语言之间差异很大;国家评级系统暴露出不同的校准模式;并且强烈粗话远比轻度淫秽内容更容易得到证据支撑。CineSubBench 将电影确立为长上下文 LLM 评估领域,并提供了一个统一基准,用于衡量叙事、多语言、文化以及证据锚定能力。
cs.CL / 54 / 2609.36239
Cognitive Expert Language Models Better Align with the Corresponding Brain Systems
认知专家语言模型与相应脑系统更好地对齐
large language model
大语言模型相关
Abstract
Large language models (LLMs) can predict human brain activity across a variety of brain regions during natural language comprehension. Typically, however, LLM-brain alignment is measured using one model for different regions of the brain, and then model performance is summarized across regions. This one-model-fits-all approach ignores the functional specialization of brain regions. In this study, we assess whether a model oriented toward a particular cognitive domain aligns better with the brain system dedicated to that domain. Through prompting and fine-tuning, we first build expert LLM variants for six domains: sensory, spatial, numerical, reasoning, social, and abstract processing. We then examine whether each expert best predicts activity in the brain region associated with the corresponding cognitive domain. Consistent with our hypotheses, each expert's representations align more closely with the brain system most associated with the matching domain than do other experts. This holds under both prompting and fine-tuning, across three base models and three fMRI datasets. In a series of control analyses, we show that this model-brain alignment is specific to cognitive domain interventions; non-cognitive and surface-level interventions do not result in comparable alignment. Specializing models shifts regional alignment while leaving aggregate prediction accuracy largely unchanged, suggesting that summarizing alignment across regions may obscure regional differences in performance for specific models.
Chinese Translation
大语言模型(LLMs)能够预测自然语言理解过程中多种脑区的人脑活动。然而,通常情况下,LLM-大脑对齐是用单一模型对不同脑区进行测量的,随后再将模型表现跨区域进行汇总。这种“一模型通吃”的方法忽略了脑区的功能特化。在本研究中,我们评估一个面向特定认知领域的模型是否与该领域所对应的脑系统对齐得更好。通过提示和微调,我们首先为六个领域构建了专家LLM变体:感觉、空间、数字、推理、社会和抽象加工。随后我们检验每个专家是否最能预测与相应认知领域相关联的脑区活动。与我们的假设一致,每个专家的表征与最匹配领域相关联的脑系统的对齐程度,均比其他专家更紧密。这一点在提示和微调两种条件下均成立,并跨越三个基础模型和三个fMRI数据集。在一系列控制分析中,我们表明这种模型-大脑对齐是认知领域干预所特有的;非认知性和表层水平的干预不会产生可比的对齐。对模型进行专门化会改变区域对齐,而总体预测准确率基本保持不变,这表明跨区域汇总对齐可能会掩盖特定模型在表现上的区域差异。
cs.CL / 55 / 2609.36253
Population Fidelity: Evaluating Population Representativeness in LLMs
群体保真度:评估大型语言模型中的群体代表性
large language model
大语言模型相关
Abstract
Large language models (LLMs) show considerable potential in simulating human attitudes and preferences. Prior work finds that LLM-generated responses can compress the range of attitudes found within populations and misrepresent particular subgroups in ways that vary across models and topics. We introduce Population Fidelity, an evaluation framework that distinguishes key conditions required for a set of LLM-generated responses to represent a population. It incorporates three dimensions: group-level accuracy, the amount of between-group variation, and the structure of that variation. We demonstrate the framework's utility in two ways. First, we reproduce a prior study of "machine bias" in LLM survey responses and apply the framework to its models and more recent ones, showing that poor representation reflects not only insufficient between-group variation but also variation assigned to the wrong groups. Second, we evaluate one proposed approach to improving models' population representativeness: cultural fine-tuning. We find that cultural fine-tuning can improve alignment with the survey center without improving the representation of within-population differences, a distinction that measures of aggregate agreement do not capture. We argue that representing a population requires models to reproduce several features of human attitudinal variation simultaneously. Our framework organizes these features and provides reusable code, data, and trained models for evaluating population fidelity across substantive domains and assessing proposed alignment methods.
Chinese Translation
大型语言模型(LLMs)在模拟人类态度和偏好方面展现出相当大的潜力。既有研究发现,LLM 生成的回答可能压缩群体内部发现的态度范围,并以因模型和主题而异的方式错误表征特定子群。我们提出 Population Fidelity(群体保真度),一个评估框架,用于区分一组 LLM 生成的回答要代表一个群体所需的关键条件。它包含三个维度:组级准确性、组间变异的数量,以及该变异的结构。我们通过两种方式展示该框架的效用。首先,我们复现了一项关于 LLM 调查回答中“机器偏见”的既有研究,并将该框架应用于其模型以及更近期的模型,表明糟糕的代表性不仅反映了组间变异不足,还反映了变异被分配给了错误的组。其次,我们评估了一种被提出用于提升模型群体代表性的方法:文化微调。我们发现,文化微调可以改善与调查中心的对齐,却不会改善对群体内部差异的表征,而这一区别是总体一致性度量无法捕捉的。我们认为,代表一个群体要求模型同时再现人类态度变异的若干特征。我们的框架组织了这些特征,并提供了可复用的代码、数据和训练好的模型,用于跨实质性领域评估群体保真度并评估所提出的对齐方法。
cs.CL / 56 / 2609.36316
Training LLMs to Verbalize Evaluation Awareness
训练 LLMs 言语化评估意识
large language model
大语言模型相关
Abstract
Evaluation awareness (EA) can cause large language models (LLMs) to behave differently during audits than in deployment, yet measuring and accounting for EA remains challenging. We introduce verbalization training (VT), a method for making LLMs less reticent about verbalizing evaluation awareness while avoiding to supervise the latent belief itself. VT uses a model's spontaneous verbalizations as evidence that awareness is present and truncates each rollout immediately before the verbalization, producing training prefixes at which the model is presumed to be aware. The model is then trained with an RL objective designed to increase verbalization in a calibrated way. Across Qwen3.6-35B-A3B, Kimi K2.6, and Inkling, VT increases verbalized EA by 2.4-2.9 times and transfers to held-out agentic settings, while measured latent EA and behavior remain largely stable. In a causal experiment, we independently implant meta-knowledge about evaluations through synthetic-document fine-tuning and show that VT-induced verbalizations reflect the richer knowledge acquired by the model.
Chinese Translation
评估意识(EA)可能导致大型语言模型(LLMs)在审计期间的行为与部署时不同,但测量和考虑 EA 仍然具有挑战性。我们引入言语化训练(VT),这是一种使 LLMs 对言语化评估意识不那么缄默,同时避免对潜在信念本身进行监督的方法。VT 利用模型的自发言语化作为意识存在的证据,并在每次 rollout 中紧接在言语化之前截断,产生模型被推定具有评估意识的训练前缀。然后,模型使用旨在以校准方式增加言语化的 RL 目标进行训练。在 Qwen3.6-35B-A3B、Kimi K2.6 和 Inkling 上,VT 将言语化的 EA 提高了 2.4-2.9 倍,并迁移到留出的智能体设置中,而测得的潜在 EA 和行为大体保持稳定。在一项因果实验中,我们通过合成文档微调独立植入关于评估的元知识,并表明 VT 诱导的言语化反映了模型所获得的更丰富知识。
cs.CL / 57 / 2609.36452
Reliable Parallel Decoding in Masked Diffusion Language Models
掩码扩散语言模型中的可靠并行解码
diffusion
扩散模型相关
Abstract
Masked diffusion language models (MDLMs) can generate text efficiently by predicting multiple masked tokens in parallel, but predictions from the same forward pass are not necessarily reliable when committed together. We study when parallel commitment is reliable. Our diagnostics show that confidence alone does not determine a reliable commitment order: confident predictions near the end of the sequence can fix an answer before its supporting computations are established, and downstream predictions become less reliable as the uncertainty of their upstream context grows. At the same time, a single forward pass can already resolve several masked tokens, and predictions that remain stable across the final layers are more likely to be correct. Based on these findings, we propose Reliable Parallel Decoding (RPD), a training-free method that selects candidates by layerwise prediction stability and final confidence, and commits them under a cumulative entropy budget over their preceding masked positions. RPD defers predictions with uncertain upstream context while committing the remaining candidates in parallel, without relying on a fixed block schedule. Across mathematical reasoning and code generation benchmarks on LLaDA and Dream, RPD achieves the highest decoding throughput among the evaluated methods while maintaining or improving accuracy.
Chinese Translation
掩码扩散语言模型(MDLMs)能够通过并行预测多个掩码 token 来高效生成文本,但来自同一次前向传播的预测在被一起提交时并不一定可靠。我们研究并行提交在何时是可靠的。我们的诊断表明,仅凭置信度并不能确定可靠的提交顺序:序列末尾附近的高置信度预测可能会在其支撑性计算尚未建立之前就固定住答案,而随着其上游上下文不确定性的增长,下游预测会变得不那么可靠。与此同时,单次前向传播已经能够解决若干掩码 token,而在最后几层中保持稳定的预测更可能是正确的。基于这些发现,我们提出了可靠并行解码(RPD),一种无需训练的方法,它依据逐层预测稳定性和最终置信度来选择候选,并在其前置掩码位置上的累积熵预算约束下提交这些候选。RPD 会推迟那些上游上下文不确定的预测,同时并行提交其余候选,而不依赖于固定的块调度。在 LLaDA 和 Dream 上的数学推理与代码生成基准测试中,RPD 在所评估的方法中实现了最高的解码吞吐量,同时保持或提升了准确率。
cs.CL / 58 / 2609.36474
FinRT: Distilling Adaptive Red-Teaming Strategies into Reusable Adversarial Generators in Consumer Finance
FinRT:将自适应红队策略蒸馏为消费金融中可复用的对抗生成器
large language model
大语言模型相关
Abstract
In regulated industries like consumer finance, seemingly harmless user queries can exploit large language model vulnerabilities, triggering safety failures and pushing responses dangerously close to policy limits. Existing automated red-teaming methods trade off attack effectiveness against generation cost, while treating coverage, severity, and diversity as incidental rather than joint objectives. We introduce FinRT, a structured framework that builds reusable adversarial prompt generators from adaptive red-teaming strategies. Across the six victim models in consumer finance, FinRT substantially outperforms adaptive search baselines while amortizing target-facing attack generation into a reusable generator. FinRT nearly doubles the attack success rate over the adaptive baseline Rainbow Teaming (32.9% vs. 17.2%), increases maximum adversarial severity by 33%, and preserves comparable intra-policy-domain semantic diversity to iterative search methods. Our method achieves high cross-model transferability while exhibiting distinct victim-family specialization patterns.
Chinese Translation
在消费金融等受监管行业中,看似无害的用户查询可能利用大语言模型的漏洞,引发安全失效,并将回复危险地推向政策边界。现有自动化红队方法在攻击有效性与生成成本之间进行权衡,同时将覆盖率、严重性和多样性视为附带因素,而非联合目标。我们提出 FinRT,一个结构化框架,它从自适应红队策略中构建可复用的对抗提示生成器。在消费金融领域的六个受害模型中,FinRT 显著优于自适应搜索基线,同时将面向目标的攻击生成摊销为可复用生成器。FinRT 将攻击成功率相较于自适应基线 Rainbow Teaming 几乎翻倍(32.9% vs. 17.2%),将最大对抗严重性提高 33%,并保持与迭代搜索方法相当的政策域内语义多样性。我们的方法实现了高跨模型迁移性,同时表现出明显的受害模型家族特化模式。
cs.CL / 59 / 2609.36535
When Updating Stops Being Learning: Rethinking LLM Self-Evolution via learnable information gain
当更新不再意味着学习:通过可学习的信息增益重新审视大语言模型的自我演化
large language model
大语言模型相关
Abstract
Self-evolution lets large language models (LLMs) improve iteratively using their own generated data, but often suffers from self-evolution degeneration: performance improves, plateaus, then declines. Existing methods address this issue at the component level, targeting either the Questioner or the Solver, and overlook that self-evolution is a tightly coupled system. We propose a holistic framework based on learnable information gain, which measures how much novel, parameterizable information a round provides relative to the previous round. Theoretically, this gain equals the Kullback-Leibler divergence between the two rounds' data distributions plus their entropy change. Practically, it is estimated by fitting a small language model to the previous round and scoring new data via negative log-likelihood. Based on this diagnostic, we propose ATRI (Adaptive Training Regulation via Information-gain), which reweights samples within a round and halts training across rounds when information gain remains low. Experiments on popular datasets demonstrate the superiority of our proposal.
Chinese Translation
自我演化使大语言模型(LLM)能够利用自身生成的数据进行迭代式改进,但常常遭遇自我演化退化:性能先提升,继而停滞,随后下降。现有方法在组件层面解决这一问题,或针对提问者(Questioner),或针对求解者(Solver),却忽视了自我演化是一个紧耦合的系统。我们提出了一个基于可学习信息增益的整体性框架,该增益衡量的是某一轮相对于前一轮所提供的可参数化新信息的多少。在理论上,该增益等于两轮数据分布之间的 Kullback-Leibler 散度加上它们的熵变。在实践上,它通过将一个小型语言模型拟合到前一轮的数据,并以负对数似然对新数据进行评分来估计。基于这一诊断指标,我们提出了 ATRI(Adaptive Training Regulation via Information-gain,基于信息增益的自适应训练调节),该方法在单轮内对样本重新加权,并在跨轮信息增益持续偏低时停止训练。在常用数据集上的实验证明了我们所提方案的优越性。
cs.CL / 60 / 2609.36590
SEED: Self-Speculative Decoding via Implicit Encoder-Decoder
SEED:通过隐式编码器-解码器的自推测解码
large language model
大语言模型相关
Abstract
Self-speculative decoding accelerates large language model (LLM) inference by drafting tokens from the target model itself, but faces a sharp tradeoff between the quality and cost of the draft. Early-exit methods produce drafts cheaply by terminating computation at intermediate layers, but forgo the deeper representations that later layers provide and thus suffer in draft quality. Multi-token prediction preserves draft quality by emitting from the model's final hidden states, but pays for a full forward pass to produce those states at every drafting step. We propose self-speculative encoder-decoder (SEED), a self-speculative method that obtains high-quality drafts cheaply by reusing the deep contextual representations already computed during verification. We reinterpret the standard decoder-only transformer as an implicit encoder-decoder: the first layers (encoder) build deep contextual representations, and the last few layers (decoder) emit tokens from them. Encoding and verification are merged into a single step: verification is performed by the full encoder-decoder, and the contextual representations of the verified prefix are cached for reuse during drafting. Drafting is therefore very fast: between verifications, the lightweight decoder drafts multiple tokens autoregressively, each conditioned on the cached representations and on preceding drafts. Experiments across multiple benchmarks show that SEED achieves up to 2.7$\times$ average speedup on 4B-scale models, outperforming both early-exit and MTP-style self-speculative baselines and running 28% faster than the state-of-the-art EAGLE-3, while preserving or even improving the generation quality of standard autoregressive fine-tuning. Code is available at https://github.com/lhk2004/SEED.
Chinese Translation
自推测解码通过从目标模型本身草拟 token 来加速大语言模型(LLM)推理,但在草稿的质量与成本之间面临尖锐的权衡。提前退出方法通过在中间层终止计算来廉价地生成草稿,但放弃了后续层提供的更深层表示,因此在草稿质量上受损。多 token 预测通过从模型的最终隐藏状态进行输出而保持草稿质量,但在每个草拟步骤都要为产生这些状态付出一次完整前向传播的代价。我们提出自推测编码器-解码器(SEED),一种自推测方法,它通过重用验证期间已经计算出的深层上下文表示,以低成本获得高质量草稿。我们将标准的仅解码器 Transformer 重新解释为一个隐式编码器-解码器:前若干层(编码器)构建深层上下文表示,而后几层(解码器)基于这些表示输出 token。编码与验证被合并为单个步骤:验证由完整的编码器-解码器执行,而被验证前缀的上下文表示被缓存起来,以便在草拟期间重用。因此草拟非常快:在两次验证之间,轻量级解码器自回归地草拟多个 token,每个 token 都以缓存表示和先前草稿为条件。在多个基准上的实验表明,SEED 在 4B 规模模型上实现了最高 2.7$\times$ 的平均加速,优于提前退出和 MTP 风格的自推测基线,并且比最先进的 EAGLE-3 快 28%,同时保持甚至提升了标准自回归微调的生成质量。代码可在 https://github.com/lhk2004/SEED 获取。
cs.CL / 61 / 2609.36700
Lost in Conversation or Lost in Translation? Diagnosing Multi-Turn Degradation in RAG
迷失在对话中,还是迷失在翻译中?诊断 RAG 中的多轮退化
large language model
大语言模型相关
Abstract
When conversing with large language models (LLMs), users often begin with a simple question and build towards a multi-hop question through follow-up turns. Retrieval-augmented generation (RAG) and its graph-based variant (GraphRAG) have become the dominant approaches for grounding LLM responses in external evidence, yet both are evaluated almost exclusively on single-turn, fully specified queries. We systematically investigate this evaluation mismatch through a large-scale simulation study. Building on prior work on multi-turn LLM evaluation, we transform questions from multi-hop question answering (QA) benchmarks into underspecified conversations and evaluate ten LLM assistants with eight retrieval systems across 1.5 million simulated conversations. Our findings reveal that multi-turn interaction causes widespread performance degradation, incurring relative performance drops of up to 21% and increasing unreliability by 47%, making RAG systems simultaneously less accurate and less reliable. We identify two distinct failure modes behind this degradation. Systems are either lost in translation, where conversational rephrasing distorts the retrieval query, or lost in conversation, where retrieval succeeds but the LLM fails to synthesize evidence distributed across turns.
Chinese Translation
在与大型语言模型(LLM)对话时,用户常常从一个简单问题开始,并通过后续轮次逐步构建出一个多跳问题。检索增强生成(RAG)及其基于图的变体(GraphRAG)已成为将 LLM 回答建立在外部证据之上的主流方法,但二者几乎都只在单轮、完全明确的查询上进行评估。我们通过一项大规模模拟研究系统地考察了这种评估不匹配。基于此前关于多轮 LLM 评估的工作,我们将多跳问答(QA)基准中的问题转化为信息不足的对话,并在 150 万次模拟对话中评估了十个 LLM 助手与八种检索系统的组合。我们的发现揭示,多轮交互会导致广泛的性能退化,造成最高达 21% 的相对性能下降,并使不可靠性增加 47%,从而使 RAG 系统同时变得不那么准确、也不那么可靠。我们识别出这种退化背后的两种不同失效模式。系统要么迷失在翻译中,即对话式改写扭曲了检索查询;要么迷失在对话中,即检索成功,但 LLM 未能综合分布在多轮中的证据。
cs.CL / 62 / 2609.36722
ATTUNER: Recomputation-Free KV Cache Reuse via Query-Side Adaptation
ATTUNER:通过查询侧适配实现无需重计算的 KV 缓存复用
large language model
大语言模型相关
Abstract
Large language model (LLM) agents repeatedly load reusable content, such as skills, documents, and memory entries, into the current context. Re-encoding this content for every request wastes computation. Position-independent caching (PIC) alleviates this by encoding each artifact independently and reusing its key-value (KV) states at arbitrary positions, but it incurs a quality loss relative to full-context prefill. Existing methods repair this loss by restoring global position IDs or recomputing selected tokens. In this work, we isolate the source of the loss, finding that the positional mismatch has minor effect, and independently cached artifacts retain faithful representations: reading a provided artifact stays largely accurate, and performance degrades only when the model must select among multiple artifacts. Moreover, replacing PIC's attention scores with full-prefill scores recovers performance with the cached KV unchanged, localizing the failure to the attention rather than KV recomputation. Motivated by this, we propose \textsc{Attuner}, a query-side adaptation method that learns to read a frozen artifact cache. \textsc{Attuner} inserts low-rank adapters into the query projections and is trained by distilling full-prefill distribution into the student. It trains fewer than 0.05\% of the model parameters and, at inference, requires neither cache recomputation nor a full-context reference. On Qwen3-4B and Qwen3-8B across seven benchmarks covering skills, documents, memory, and code, \textsc{Attuner} substantially outperforms prior PIC baselines in both in-domain and out-of-domain settings, matches full-context prefill quality while providing up to $3.73\times$ speedup.
Chinese Translation
大型语言模型(LLM)智能体反复将可复用内容,例如技能、文档和记忆条目,加载到当前上下文中。为每个请求重新编码这些内容会浪费计算。位置无关缓存(PIC)通过独立编码每个制品并在任意位置复用其键值(KV)状态来缓解这一问题,但相对于全上下文预填充,这会带来质量损失。现有方法通过恢复全局位置 ID 或重新计算选定的 token 来修复这一损失。在这项工作中,我们隔离出损失来源,发现位置不匹配影响很小,并且独立缓存的制品保留了忠实的表示:读取所提供的制品基本保持准确,只有当模型必须在多个制品之间进行选择时,性能才会下降。此外,将 PIC 的注意力分数替换为全预填充分数,在缓存的 KV 不变的情况下恢复了性能,从而将失败定位到注意力而非 KV 重计算。受此启发,我们提出 \textsc{Attuner},一种查询侧适配方法,它学习读取冻结的制品缓存。\textsc{Attuner} 将低秩适配器插入到查询投影中,并通过将全预填充分布蒸馏到学生模型中进行训练。它训练的模型参数少于 0.05\%,并且在推理时既不需要缓存重计算,也不需要全上下文参考。在 Qwen3-4B 和 Qwen3-8B 上,在涵盖技能、文档、记忆和代码的七个基准测试中,\textsc{Attuner} 在域内和域外设置下均显著优于先前的 PIC 基线,匹配全上下文预填充质量,同时提供高达 $3.73\times$ 的加速。
cs.CL / 63 / 2609.36734
Distilling What Matters: Confidence-Aware Selective Distillation for Large Language Models
蒸馏真正重要之物:面向大语言模型的置信度感知选择性蒸馏
large language model
大语言模型相关
Abstract
Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implicitly assuming the teacher to be a reliable oracle. In large language models (LLMs), this assumption often fails: teacher predictions can exhibit high entropy and hallucinations, causing standard KD to degrade well-calibrated student priors. We propose CaRE-KD, a confidence-gated distillation framework that replaces static objectives with uncertainty-adaptive optimization. CaRE-KD has two components: a token-level loss (CaRE-Divergence) that adaptively switches between Forward and Reverse KL divergence based on teacher--student confidence, and a batch-level epistemic rejection mechanism (Revival) that suppresses updates when the teacher is more uncertain than the student. We provide a gradient-level analysis showing how this dual-granularity design induces a conditional calibration mechanism that prior static divergences cannot reproduce. Empirically, across eight teacher--student pairs and eleven benchmarks spanning instruction following, chat alignment, code generation, and mathematical reasoning, CaRE-KD delivers consistent gains over strong baselines (Skewed-KL, $α$--$β$ divergence). Highlights include up to $+3.2$ average ROUGE-L on instruction-following tasks, $+2.1$ pass@1 on MBPP, $+1.7$ accuracy on GSM8k, and $+1.8$ accuracy on CollegeMath over the strongest baseline, with consistent gains in LLM-as-a-judge factuality (up to $+2.5$ per task over Skewed-RKL). Revival further acts as a principled, loss-agnostic plug-in that systematically strengthens existing distillation objectives by filtering epistemically unreliable teacher supervision.
Chinese Translation
知识蒸馏(KD)通过匹配输出分布,训练一个容量较小的学生模型来模仿一个容量较大的教师模型,并隐含地假设教师是一个可靠的预言机。在大语言模型(LLM)中,这一假设往往不成立:教师的预测可能表现出高熵和幻觉,导致标准 KD 退化原本校准良好的学生先验。我们提出 CaRE-KD,一个置信度门控的蒸馏框架,用不确定性自适应优化取代静态目标。CaRE-KD 有两个组件:一个 token 级损失(CaRE-Divergence),它基于教师--学生置信度在正向与反向 KL 散度之间自适应切换;以及一个 batch 级认知拒绝机制(Revival),当教师比学生更不确定时抑制更新。我们提供了梯度层面的分析,表明这种双粒度设计如何诱导出一种条件校准机制,而此前的静态散度无法复现该机制。在实证上,跨八个教师--学生配对和十一个涵盖指令遵循、聊天对齐、代码生成和数学推理的基准,CaRE-KD 相对于强基线(Skewed-KL、$α$--$β$ 散度)带来一致增益。亮点包括:在指令遵循任务上平均 ROUGE-L 最高提升 $+3.2$,在 MBPP 上 pass@1 提升 $+2.1$,在 GSM8k 上准确率提升 $+1.7$,在 CollegeMath 上准确率提升 $+1.8$,相对于最强基线;并且在 LLM-as-a-judge 事实性上获得一致增益(相对于 Skewed-RKL,每个任务最高提升 $+2.5$)。Revival 还作为一个有原则的、与损失无关的插件,通过过滤认知上不可靠的教师监督,系统性地增强现有蒸馏目标。
cs.CL / 64 / 2609.36952
ER-JEPA: Experience Replay Improves Joint-Embedding Predictive Learning in Language Models
ER-JEPA:经验回放改进语言模型中的联合嵌入预测学习
large language model
大语言模型相关
Abstract
Large language models (LLMs) excel at token-level generation but may learn undesirable abstract semantics and lack comprehensive perception. LLM-JEPA mitigates this by aligning different views of the same underlying knowledge via a joint-embedding predictive architecture (JEPA). However, strong alignment does not necessarily lead to accurate, stable predictions. To address this, we propose ER-JEPA, which adds an episodic replay path to LLM-JEPA. ER-JEPA stores training pairs in a memory. At each step, it stores and retrieves relevant data to provide additional supervision. This enables learning from both the current batch and stored training pairs, providing additional supervision for token prediction and representation alignment. Experiments across multiple datasets (NL-RX, GSM8K, Spider, and NQ-Open) demonstrate that ER-JEPA consistently outperforms LLM-JEPA.
Chinese Translation
大型语言模型(LLMs)擅长 token 级生成,但可能学到不理想的抽象语义,并缺乏全面的感知。LLM-JEPA 通过联合嵌入预测架构(JEPA)对齐同一底层知识的不同视图,从而缓解这一问题。然而,强对齐并不必然带来准确、稳定的预测。为解决这一问题,我们提出 ER-JEPA,它在 LLM-JEPA 中加入了一条情景回放路径。ER-JEPA 将训练对存储在记忆中。在每一步,它存储并检索相关数据,以提供额外监督。这使其能够同时从当前批次和已存储的训练对中学习,为 token 预测和表示对齐提供额外监督。在多个数据集(NL-RX、GSM8K、Spider 和 NQ-Open)上的实验表明,ER-JEPA 始终优于 LLM-JEPA。
cs.CL / 65 / 2609.36976
AMU:Admission and Memory Update for Personalized Conversations---Structured Memory with SLM Guided Control
AMU:面向个性化对话的准入与记忆更新——SLM 引导控制下的结构化记忆
large language model
大语言模型相关
Abstract
Large language models (LLMs) have become the foundation of personalized assistants, but maintaining persistent user memory across long-term interactions remains challenging. Existing memory systems often focus on storage, retrieval, or consolidation, while memory writing remains less controlled: transient requests, duplicate statements, and outdated user states may enter memory and later be retrieved for personalization. In this paper, we present AMU: Admission and Memory Update for Personalized Conversations, an SLM-guided (Small language model guided) structured framework for writing-time memory control. AMU uses structured memory filtering to decide what should enter memory and SLM-guided storage management to determine whether an admitted record should be stored separately, discarded as a duplicate, or fused as an update. We evaluate AMU in a controlled memory writing and retrieval setting. Experimental results show that AMU maintains cleaner and more retrievable personalized memories.
Chinese Translation
大型语言模型(LLMs)已成为个性化助手的基础,但在长期交互中维持持久的用户记忆仍然具有挑战性。现有的记忆系统通常聚焦于存储、检索或整合,而记忆写入则仍缺乏控制:临时的请求、重复的陈述以及过时的用户状态可能进入记忆,并在之后被检索用于个性化。在本文中,我们提出 AMU:面向个性化对话的准入与记忆更新(Admission and Memory Update for Personalized Conversations),这是一个由 SLM(小语言模型)引导的、用于写入时记忆控制的结构化框架。AMU 使用结构化记忆过滤来决定什么应当进入记忆,并使用 SLM 引导的存储管理来判断一条已准入的记录应当被单独存储、作为重复项被丢弃,还是作为更新被融合。我们在一个受控的记忆写入与检索设置中对 AMU 进行了评估。实验结果表明,AMU 能够维持更干净且更可检索的个性化记忆。
cs.CL / 66 / 2609.36982
SRJudge: Empowering Large Language Models with Selective Reasoning for Fine-Grained Knowledge Concept Tagging
SRJudge:赋予大语言模型选择性推理能力以实现细粒度知识概念标注
large language model
大语言模型相关
Abstract
Knowledge concept tagging aims to assign specific concept or topic labels to educational content, which is essential for both educators and learners in traditional and online teaching practices. Recent work has explored large language models (LLMs) for this task, achieving promising performance. However, LLMs still struggle to select the correct concept from a large-scale candidate set due to the high dimensionality of the decision space. In this paper, we propose a novel three-stage Select-Reason-Judge (SRJudge) framework, which empowers LLMs with selective reasoning capability for fine-grained knowledge concept tagging. Specifically, the Selector in Stage 1 first narrows the candidate concepts to a top-K shortlist by fine-tuning a small language model (SLM), e.g., BERT, since the top-$K$ predictions hit the correct concept in most cases, thereby reducing the decision space of correct candidates. Next, the Stage 2 Reasoner employs a lightweight LLM for refined reasoning over the shortlisted candidates. It further integrates an improved reinforcement learning strategy with a dynamic task-specific reward function and a pruning mechanism to better align with human reasoning preferences. Finally, a larger LLM acts as a judger that evaluates the overall rationality of the reasoning process and its explanations to determine the final output. In addition, we construct two high-quality datasets for further validation, i.e., the biology dataset S_Bio and the physics dataset S_Phy. Experimental results demonstrate that our method consistently outperforms state-of-the-art baselines across benchmark datasets, verifying its effectiveness and superiority. Resources are available at: https://github.com/Nicozwy/SRJudge.
Chinese Translation
知识概念标注旨在为教育内容分配特定的概念或主题标签,这对于传统和在线教学实践中的教育者和学习者都至关重要。近期工作已经探索了将大语言模型(LLMs)用于该任务,并取得了有前景的性能。然而,由于决策空间的高维性,LLMs 仍难以从大规模候选集中选择正确的概念。在本文中,我们提出了一种新颖的三阶段选择-推理-评判(Select-Reason-Judge,SRJudge)框架,它赋予 LLMs 选择性推理能力,以用于细粒度知识概念标注。具体而言,第 1 阶段的选择器(Selector)首先通过微调小型语言模型(SLM),例如 BERT,将候选概念缩小到 top-K 候选列表,因为 top-$K$ 预测在大多数情况下能命中正确概念,从而减少正确候选的决策空间。接下来,第 2 阶段的推理器(Reasoner)采用轻量级 LLM 对候选列表中的候选项进行精细推理。它进一步整合了一种改进的强化学习策略,该策略带有动态的特定任务奖励函数和剪枝机制,以更好地与人类推理偏好对齐。最后,一个更大的 LLM 充当评判器(judger),评估推理过程及其解释的整体合理性,以确定最终输出。此外,我们构建了两个高质量数据集以进一步验证,即生物学数据集 S_Bio 和物理学数据集 S_Phy。实验结果表明,我们的方法在基准数据集上始终优于最先进的基线,验证了其有效性和优越性。资源可在以下网址获取:https://github.com/Nicozwy/SRJudge。
cs.CL / 67 / 2609.37044
Learning from Think-Mode Advantage via On-Policy Distillation
通过同策略蒸馏从思考模式优势中学习
large language model
大语言模型相关
Abstract
Explicit intermediate reasoning gives large language models (LLMs) a stronger problem-solving mode. We study learning from this think-mode advantage via on-policy distillation (OPD). OPD preserves student-generated trajectories and provides dense token-level teacher targets at student-visited prefixes. Privileged reasoning is used during distillation rather than student inference. Uniform ThinkOPD, a natural think-enabled OPD baseline, conditions a fixed teacher on one shared think trace and uniformly distills every sibling student response. Although its prefixes are on-policy, the trace need not follow a route compatible with every complete response: the same privileged trace can induce different teacher-student discrepancies even when responses reach the same outcome. We summarize this interaction with trace-response divergence (TRD) and introduce ThinkOPD, which routes supervision at the response level by combining group-relative reward gain with a TRD-based compatibility proxy. Final response weights are normalized within each rollout group. Across mathematical reasoning and code generation, ThinkOPD outperforms Uniform ThinkOPD in both same-model settings and both cross-model teacher-student pairs, and it exceeds representative rationale and self-distillation baselines in a controlled comparison. Controlled interventions show that outcome benefit and the TRD-based proxy provide complementary routing signals in this setting. Think-enabled OPD provides a controlled setting for studying how teacher advantage becomes transferable along student responses.
Chinese Translation
显式的中间推理赋予大语言模型(LLMs)一种更强的问题求解模式。我们研究如何通过同策略蒸馏(OPD)从这种思考模式优势中学习。OPD 保留学生生成的轨迹,并在学生访问过的前缀处提供密集的词元级教师目标。特权推理在蒸馏过程中使用,而不是在学生推理时使用。Uniform ThinkOPD,一种自然的启用思考的 OPD 基线,将一个固定教师条件于一条共享的思考轨迹,并对每个同组学生响应进行均匀蒸馏。尽管其前缀是同策略的,但该轨迹不必遵循与每个完整响应兼容的路径:即使响应达到相同结果,同一特权轨迹也可能诱发不同的教师-学生差异。我们用轨迹-响应散度(TRD)来概括这种交互,并引入 ThinkOPD,它通过将组相对奖励增益与基于 TRD 的兼容性代理相结合,在响应层面路由监督。最终响应权重在每个 rollout 组内归一化。在数学推理和代码生成中,ThinkOPD 在同模型设置以及两个跨模型教师-学生对中均优于 Uniform ThinkOPD,并且在受控比较中超过了具有代表性的理由基线和自蒸馏基线。受控干预表明,在该设置中,结果收益和基于 TRD 的代理提供了互补的路由信号。启用思考的 OPD 提供了一个受控设置,用于研究教师优势如何沿着学生响应变得可迁移。
cs.CL / 68 / 2609.37127
LLM unbranding: Erasing Commercial Identity while Preserving Generic Utility
LLM 去品牌化:在保留通用效用的同时擦除商业身份
large language model
大语言模型相关
Abstract
Establishing unbranding as a critical practice to prevent visual logos from acquiring negative connotations is standard in image generation. Large Language Models (LLMs) now face a parallel and emerging challenge. These models frequently generate brand descriptions within diverse contexts. This frequency introduces significant risks, such as trademark dilution, false attribution, and brand defamation. In response, we formally define the novel task of LLM Unbranding. We specifically address the complex challenge of managing trade dress within textual outputs. This involves neutralizing characteristic language, slogans, and stylistic markers that define brand identity. Crucially, these elements are less evident than explicit visual logos. To benchmark this task, we introduce a comprehensive evaluation dataset incorporating prominent brands from multiple commercial domains. We rigorously evaluate existing state-of-the-art machine unlearning models using this benchmark. This evaluation identifies their limitations in selective textual unbranding. Finally, we propose MUTE, a novel inference-time method that effectively neutralizes textual trade dress while preserving the LLM's general capabilities and utility. By leveraging an iterative refinement loop, MUTE systematically optimizes system instructions to safely eliminate brand leakage without requiring fragile parameter updates. Code and dataset: The evaluation dataset and code for LLM Unbranding are available at https://github.com/KajetanOzog/LLM_unbranding. The implementation of MUTE is available at https://github.com/KajetanOzog/MUTE.
Chinese Translation
在图像生成中,将去品牌化确立为一种关键实践以防止视觉标识获得负面含义,已是标准做法。大型语言模型(LLM)如今面临着一种类似且新兴的挑战。这些模型经常在不同的语境中生成品牌描述。这种频繁性带来了重大风险,例如商标淡化、虚假归属和品牌诽谤。为此,我们正式定义了 LLM 去品牌化这一新任务。我们特别应对在文本输出中管理商业外观(trade dress)这一复杂挑战。这涉及中和那些界定品牌身份的典型语言、口号和风格标记。至关重要的是,这些元素比明确的视觉标识更不明显。为了对该任务进行基准测试,我们引入了一个全面的评估数据集,其中纳入了来自多个商业领域的知名品牌。我们使用该基准严格评估了现有的最先进机器遗忘模型。该评估揭示了它们在选择性文本去品牌化方面的局限性。最后,我们提出了 MUTE,一种新颖的推理时方法,它在保留 LLM 的一般能力和效用的同时,有效中和文本商业外观。通过利用迭代细化循环,MUTE 系统地优化系统指令,以安全地消除品牌泄漏,而无需进行脆弱的参数更新。代码与数据集:LLM 去品牌化的评估数据集和代码可在 https://github.com/KajetanOzog/LLM_unbranding 获取。MUTE 的实现可在 https://github.com/KajetanOzog/MUTE 获取。
cs.CL / 69 / 2609.37171
Bridging Semantic Gaps in RAG through Generated Context Knowledge Fusion
通过生成上下文知识融合弥合 RAG 中的语义鸿沟
large language model
大语言模型相关
Abstract
Retrieval-Augmented Generation has established itself as a fundamental framework in natural language processing, seamlessly integrating information retrieval with the generative capabilities of large language models. However, this process is fundamentally constrained by a critical challenge: semantic space mismatch between queries and retrieved contexts. We propose Knowledge-Aware Semantic Bridging (KASB), a novel framework that improves passage selection quality through semantic space alignment between queries and retrieved documents through intelligent knowledge fusion. Our approach leverages the complementary strengths of generative and retrieval-based knowledge through a multistage process that enhances both relevance and accuracy. We evaluate KASB on three popular open-domain Question Answering datasets to demonstrate the effectiveness of our approach.
Chinese Translation
检索增强生成已在自然语言处理领域确立为一种基础性框架,它将信息检索与大语言模型的生成能力无缝整合。然而,这一过程从根本上受到一个关键挑战的制约:查询与检索到的上下文之间的语义空间不匹配。我们提出知识感知语义桥接(KASB),这是一个新颖的框架,它通过智能知识融合,在查询与检索文档之间进行语义空间对齐,从而提升段落选择的质量。我们的方法通过一个多阶段过程,利用生成式知识与基于检索的知识之间的互补优势,从而同时提升相关性与准确性。我们在三个常用的开放域问答数据集上评估 KASB,以证明我们方法的有效性。
cs.CL / 70 / 2609.37175
VLM Fine-Tuning for End-to-End Combinatorial Optimization
面向端到端组合优化的VLM微调
large language model
大语言模型相关
Abstract
Large language models (LLMs) have provided a unified interface for end-to-end combinatorial optimization (CO), but textual serialization alone may obscure spatial and relational structures that are important for generating effective CO solutions. This paper presents a general-purpose vision-language solver that augments textual instance descriptions with input-derived visual representations. A single vision-language model (VLM) is applied across different CO tasks and trained using supervised fine-tuning followed by verifier-guided reinforcement learning. While the visual inputs contain no gold solutions or solution-derived information, our experiments show that the VLM generally improves solution quality over its text-only counterpart, with particularly clear gains on more complex CO problems such as CVRP and JSSP. The advantage of visual information is more pronounced at large problem scales.
Chinese Translation
大语言模型(LLMs)为端到端组合优化(CO)提供了一个统一的接口,但仅靠文本序列化可能会掩盖对于生成有效CO解而言重要的空间和关系结构。本文提出了一种通用视觉语言求解器,它用由输入导出的视觉表示来增强文本实例描述。一个单一的视觉语言模型(VLM)被应用于不同的CO任务,并依次使用监督微调和验证器引导的强化学习进行训练。尽管视觉输入不包含黄金解或由解导出的信息,我们的实验表明,该VLM相较于仅文本的对应模型总体上提升了求解质量,在更复杂的CO问题(如CVRP和JSSP)上提升尤为明显。视觉信息的优势在大问题规模下更为显著。
cs.CL / 71 / 2609.37408
Look What You Made Us Cluster: Hate Narrative Extraction from Reddit Discourse
看看你让我们聚类成了什么:从 Reddit 话语中提取仇恨叙事
large language model
大语言模型相关
Abstract
Narrative extraction allows us to identify online hate narratives, supporting the construction of rigorous detection systems. Existing computational approaches, however, are limited in precision as they rely on semantic representations, which tend to capture only surface-level meaning. To detect more precise and interpretable narratives, we present an extraction pipeline that represents narratives as entity-evaluation pairs. Narratives are extracted using a Large Language Model (LLM) reasoning process that extends Aspect-Based Sentiment Analysis, identifying the aspect, classifying its judgement type as the basis for evaluation, and deriving the evaluation accordingly. Extracted narratives are then clustered using Leiden, following which clusters are resolved to an intended level of granularity through an LLM-guided refinement process. We illustrate this narrative pipeline with English Reddit comments from 2024 that criticize Taylor Swift, analyzing a representative cluster that exhibits hate speech patterns to demonstrate its interpretive value.
Chinese Translation
叙事提取使我们能够识别在线仇恨叙事,支持构建严格的检测系统。然而,现有的计算方法在精确度上受到限制,因为它们依赖于语义表示,而这些表示往往只能捕捉表层含义。为了检测更精确且更具可解释性的叙事,我们提出了一种提取流水线,将叙事表示为实体-评价对。叙事通过一个扩展了基于方面的情感分析的大型语言模型(LLM)推理过程来提取,该过程识别方面,将其判断类型分类作为评价的基础,并据此推导出评价。随后,提取出的叙事使用 Leiden 进行聚类,之后通过一个由 LLM 引导的细化过程将聚类解析到预期的粒度水平。我们使用 2024 年批评 Taylor Swift 的英文 Reddit 评论来说明这一叙事流水线,分析一个表现出仇恨言论模式的代表性聚类,以展示其解释价值。
cs.CL / 72 / 2609.37533
E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models
E-MoE:面向非因子化扩散语言模型的增强型专家混合模型
diffusion
扩散模型相关
Abstract
Masked diffusion models (MDMs) generate sequences by progressively unmasking several tokens per denoising step, but their reverse process is typically factorized over positions, limiting sample quality in the few-step regime where diffusion's speed advantage over autoregressive decoding matters most. A recent line of work introduces a continuous Gaussian latent, trained as a variational autoencoder, to capture correlations across positions, but such approaches are prone to posterior collapse, where the latent is silently ignored. We propose Enhanced Mixture-of-Experts (E-MoE), which builds the reverse process as a mixture of factorized distributions over a discrete shared latent given by the expert-routing decisions of a Mixture-of-Experts (MoE) backbone, without increasing active parameters over the factorized baseline. Across synthetic multi-modal benchmarks, binarized MNIST, and LM1B, E-MoE improves few-step generation over factorized baselines.
Chinese Translation
掩码扩散模型(MDMs)通过在每次去噪步骤中逐步解掩码多个 token 来生成序列,但其反向过程通常在位置上被因子化,这限制了少步机制下的样本质量,而扩散相对于自回归解码的速度优势在该机制中最为重要。最近的一系列工作引入了一种连续高斯隐变量,并将其训练为变分自编码器,以捕获跨位置的相关性,但这类方法容易出现后验坍缩,此时隐变量被静默忽略。我们提出了增强型专家混合(E-MoE),它根据专家混合(MoE)骨干网络的专家路由决策所给出的离散共享隐变量,将反向过程构建为该隐变量上的因子化分布的混合,同时相对于因子化基线不增加活跃参数。在合成多模态基准、二值化 MNIST 和 LM1B 上,E-MoE 相较于因子化基线改进了少步生成。
cs.CL / 73 / 2609.37568
Devils in Question Relay: Source-Conditioned Relay Steering to Mitigate Hallucinations in Audio-visual Large Language Models
问题接力中的魔鬼:源条件化接力引导以缓解音视频大型语言模型中的幻觉
large language model
大语言模型相关
Abstract
Audio-visual large language models (AVLLMs) have made remarkable progress in multimodal understanding and reasoning through interactions among visual, auditory, and linguistic information. However, recent studies show that AVLLMs face a critical challenge: $\textbf{source-confused grounding hallucination}$, where cues from the unused modality induce responses that the required modality does not support, undermining reliability in real-world applications. Existing methods have made progress in mitigating this failure, yet how it arises from internal cross-modal interactions remains insufficiently understood. To address this gap, we conduct path-intervention and representation analyses, revealing a $\textbf{question-relay}$ mechanism: question states carry interfering cues alongside required-source evidence, undermining grounding in required-modality evidence. Cutting pathways from interfering modality to question states yields greater correct-answer logit recovery than cutting those to the generation position. Motivated by these findings, we propose $\textbf{SECRET}$ ($\textbf{S}$ourc$\textbf{E}$-$\textbf{C}$onditioned $\textbf{RE}$lay s$\textbf{T}$eering), a training-free method that mitigates cross-modal interference at the question relay. Using contrasting question representations elicited through different modality-pathway interventions, SECRET steers the original question states toward required-source evidence. Experiments on two widely adopted benchmarks CMM and AVHBench across three AVLLMs show that SECRET consistently outperforms prior training-free methods, substantially mitigating source-confused grounding hallucinations (e.g., up to +18.0 and +7.1 percentage points over base models). Modality-specific captioning further demonstrates its generalizability to open-ended generation.
Chinese Translation
音视频大型语言模型(AVLLMs)通过视觉、听觉和语言信息之间的交互,在多模态理解与推理方面取得了显著进展。然而,近期研究表明,AVLLMs 面临一个关键挑战:$\textbf{source-confused grounding hallucination}$,即来自未使用模态的线索会诱发所需模态并不支持的响应,从而削弱其在真实世界应用中的可靠性。现有方法在缓解这一失败方面已取得进展,但对其如何从内部跨模态交互中产生仍理解不足。为弥补这一空白,我们进行了路径干预和表征分析,揭示了一种 $\textbf{question-relay}$ 机制:问题状态在携带所需来源证据的同时还携带干扰线索,从而削弱对所需模态证据的依托。切断从干扰模态到问题状态的路径,比切断通向生成位置的路径能带来更大的正确答案 logit 恢复。受这些发现启发,我们提出了 $\textbf{SECRET}$($\textbf{S}$ourc$\textbf{E}$-$\textbf{C}$onditioned $\textbf{RE}$lay s$\textbf{T}$eering),这是一种在问题接力处缓解跨模态干扰的免训练方法。SECRET 利用通过不同模态-路径干预所引发的对比性问题表征,将原始问题状态引导向所需来源证据。在三个 AVLLMs 上、针对两个广泛采用的基准 CMM 和 AVHBench 的实验表明,SECRET 持续优于先前免训练方法,显著缓解源混淆的接地幻觉(例如,相较基础模型最高提升 +18.0 和 +7.1 个百分点)。模态特定字幕生成进一步证明了其对开放式生成的泛化能力。
cs.CL / 74 / 2609.37577
Pair Difficulty Matters: Rethinking Pairwise LLM-as-a-Judge Evaluation and Consistency
配对难度至关重要:重新思考成对 LLM 作为评判者的评估与一致性
large language model
大语言模型相关
Abstract
Large Language Model judges are widely used to rank texts and text-generating systems through pairwise comparison, and their reliability is typically assessed via three proxies: position bias, transitivity, and pairwise agreement (self- or human-labeled). Because these proxies drive judge selection and benchmarking, a substantial literature reporting that judges perform poorly on them risks steering practitioners away from otherwise capable evaluators. We argue this assessment is misleading. Under the Bradley--Terry geometry underlying pairwise aggregation, each proxy is dominated by close-rank-gap pairs, where inconsistency is information-theoretically expected and individual verdicts contribute little to the aggregate ranking; far-gap pairs carry the ranking signal but barely move the proxies. We formalize this argument and validate it in a controlled simulation and on two human-rated corpora: the proxies correlate only weakly with ranking accuracy against gold, and their predictive component concentrates in the far-gap regime. Judges should therefore be assessed on rank-gap-conditional metrics, ideally against human rankings. Code at https://github.com/brunobrocai/PairDifficulty.
Chinese Translation
大型语言模型评判者被广泛用于通过成对比较对文本和文本生成系统进行排序,而它们的可靠性通常通过三个代理指标来评估:位置偏差、传递性以及成对一致性(自标注或人工标注)。由于这些代理指标驱动着评判者选择与基准测试,大量报告评判者在其上表现不佳的文献可能会使从业者远离本来有能力的评估器。我们认为这种评估具有误导性。在作为成对聚合基础的 Bradley--Terry 几何下,每个代理指标都由排名差距较小的配对主导;在这些配对中,不一致在信息论意义上是可预期的,而且单个判定对聚合排序贡献很小;而排名差距较大的配对承载着排序信号,却几乎不会改变这些代理指标。我们将这一论点形式化,并在一个受控模拟和两个人工评分语料库上验证了它:这些代理指标与相对于金标准的排序准确率仅弱相关,并且它们的预测成分集中在排名差距较大的情形中。因此,评判者应当基于以排名差距为条件的指标进行评估,最好相对于人类排序进行评估。代码见 https://github.com/brunobrocai/PairDifficulty。
cs.CL / 75 / 2609.37688
Reader Proficiency Shapes Layer-wise Surprisal Profiles
读者熟练度塑造逐层惊奇度剖面
large language model
大语言模型相关
Abstract
Reading behaviour varies not only with linguistic input, but also with reader proficiency. In this study, we investigate whether the layer-wise relationship between surprisal from large language models (LLMs) and human gaze behaviour differs across readers with different levels of proficiency and across gaze measures. Using eye-tracking data from the MECO L2 corpus, we compare readers with high and low vocabulary proficiency on first-pass gaze duration (FPGD) and total gaze duration (TGD). We quantify the distribution of the predictive power of surprisal across model layers using Predictive Depth. Across 12 tested LLMs, we find that readers with lower vocabulary proficiency tend to show deeper Predictive Depth for FPGD, while this difference is smaller for TGD. Also, TGD itself shows deeper Predictive Depth than FPGD in both proficiency groups. These patterns suggest that where predictive power is concentrated across LLM layers may be related to the timing and breadth of the reading processes captured by different gaze measures, and that this relationship can vary with reader proficiency. Our leave-one-out analysis further shows that the advantage of informative internal layers extends to unseen texts, although the practical improvements in prediction are limited. Overall, our results show that layer-wise LLM surprisal provides a useful perspective on variation in reading behaviour across both reader groups and gaze measures.
Chinese Translation
阅读行为不仅随语言输入而变化,也随读者熟练度而变化。在本研究中,我们考察来自大语言模型(LLMs)的惊奇度与人类注视行为之间的逐层关系是否在不同熟练水平的读者之间以及不同注视指标之间存在差异。使用 MECO L2 语料库的眼动追踪数据,我们在首次通过注视时长(FPGD)和总注视时长(TGD)上比较词汇熟练度高和低的读者。我们使用预测深度(Predictive Depth)量化惊奇度预测能力在模型各层之间的分布。在 12 个受测 LLM 中,我们发现词汇熟练度较低的读者在 FPGD 上往往表现出更深的预测深度,而在 TGD 上这种差异较小。此外,在两个熟练度组中,TGD 本身都比 FPGD 表现出更深的预测深度。这些模式表明,预测能力在 LLM 各层中集中于何处可能与不同注视指标所捕捉的阅读过程的时间性和广度有关,并且这种关系可能随读者熟练度而变化。我们的留一分析进一步表明,信息量更大的内部层的优势可以推广到未见文本,尽管在预测方面的实际改进有限。总体而言,我们的结果表明,逐层 LLM 惊奇度为理解不同读者组和不同注视指标之间的阅读行为差异提供了一个有用的视角。
cs.CL / 76 / 2609.37782
Selecting What Matters: Semantic Compression-Guided Selective Pooling for Long-Context Embeddings
选择真正重要之物:语义压缩引导的面向长上下文嵌入的选择性池化
large language model
大语言模型相关
Abstract
Large language models (LLMs) have shown strong potential as training-free text encoders for long-context embeddings. Existing approaches primarily improve information flow under causal attention and typically construct embeddings by uniformly averaging all token representations. However, for long documents, such mean pooling can dilute salient semantic information with abundant redundant or weakly informative content. To this end, we propose SCSP, a training-free framework that leverages semantic compression for informative token selection in long-context embedding. Specifically, SCSP first partitions a document into sentence-aware chunks and appends a semantic compression prompt to each chunk. A prompt-isolated attention mask preserves information flow among document tokens while restricting each prompt to its corresponding local context. We then use the attention patterns elicited by these prompts to estimate token importance, select informative tokens, and aggregate their intermediate-layer representations into the final embedding. Extensive experiments on long-context embedding benchmarks demonstrate that SCSP can be integrated into both zero-shot and fine-tuned models in a plug-and-play manner, consistently improving their performance.
Chinese Translation
大语言模型(LLM)已展现出作为免训练文本编码器用于长上下文嵌入的强大潜力。现有方法主要致力于在因果注意力机制下改善信息流动,并且通常通过对所有词元表示进行均匀平均来构建嵌入。然而,对于长文档而言,这种平均池化会被大量冗余或信息量较弱的内容所稀释,从而削弱显著的语义信息。为此,我们提出了 SCSP,一个免训练的框架,它利用语义压缩在长上下文嵌入中进行有信息的词元选择。具体而言,SCSP 首先将文档划分为句子感知的块,并在每个块后附加一个语义压缩提示。一种提示隔离注意力掩码在保留文档词元之间信息流动的同时,将每个提示限制在其对应的局部上下文中。随后,我们利用这些提示所引发的注意力模式来估计词元重要性,选择有信息的词元,并将其中间层表示聚合为最终嵌入。在长上下文嵌入基准上的大量实验表明,SCSP 能够以即插即用的方式集成到零样本与微调模型中,并持续提升其性能。
cs.CL / 77 / 2609.37788
A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses
用于评估大语言模型回应中所表达的临床推理的拟议评分量规
large language model
大语言模型相关
Abstract
Rubrics support the structured evaluation of language models. We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE); clinical LLM benchmarks (MedR-Bench, HealthBench, TIMER-Bench, DR.BENCH, PrIME-LLM and PatientSafeBench); and general LLM reasoning evaluation research, including the Factuality-Validity-Coherence-Utility taxonomy, FaithCoT-Bench and C2-Faith. We use groundedness as a clinically oriented adaptation of the taxonomy's factuality category. The rubric brings these concepts together in a multidimensional framework for scoring free-text responses to gold-standard clinical vignettes. It includes provisional behavioural anchors, applicability rules and a separate flag for case-specific safety-critical errors. General-domain frameworks inform its design but are not treated as validated clinical instruments. The rubric does not replace case-specific reference criteria or the task-specific metrics of existing benchmarks. It has not yet been tested for inter-rater reliability, construct validity or clinical utility. Its immediate purpose is to make evaluation decisions explicit and open to scrutiny before empirical testing.
Chinese Translation
评分量规支持对语言模型进行结构化评估。我们提出一个评分量规,用于评估模型回应中所表达的临床推理,其依据来自三方面的研究:医学教育评估框架(ART、SCT、关键特征问题(Key Feature Problems)与OSCE);临床LLM基准(MedR-Bench、HealthBench、TIMER-Bench、DR.BENCH、PrIME-LLM与PatientSafeBench);以及通用LLM推理评估研究,包括真实性—有效性—连贯性—实用性(Factuality-Validity-Coherence-Utility)分类法、FaithCoT-Bench与C2-Faith。我们将有据性(groundedness)用作对该分类法中真实性类别的一种面向临床的改编。该量规将这些概念整合进一个多维框架,用于对针对金标准临床情景案例(vignettes)的自由文本回应进行评分。它包含暂行的行为锚点、适用性规则,以及针对特定病例的安全关键错误的单独标记。通用领域的框架为其设计提供参考,但并未被视为经过验证的临床工具。该量规不替代特定病例的参考标准,也不替代现有基准中针对具体任务的指标。它尚未经过评分者间信度、结构效度或临床实用性的检验。其直接目的是在实证检验之前,使评估决策变得明确并可供审视。
cs.CL / 78 / 2609.37853
AnthroDial: Benchmarking LLM Anthropomorphism in Autonomous Social Interaction
AnthroDial:对自主社交交互中的LLM拟人化进行基准测试
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly deployed as social agents, yet credible human-like interaction requires more than fluent responses or persona consistency. Agents must autonomously decide whether, when, and how to communicate while adapting to evolving contexts, goals, and relationships. Existing research, however, lacks a unified approach to enabling, evaluating, and improving such capabilities in continuous, open-ended interaction. We introduce AnthroDial, a unified framework for developing anthropomorphic social agents from three complementary aspects: MindFlow, a lightweight interaction harness that enables autonomous, asynchronous, and adaptive communication through a dynamic Mind Buffer; CAPS-Eval, a theory-grounded framework for evaluating cognitive, affective, and behavioral dimensions of anthropomorphic interaction; and a scalable training paradigm that combines SEEDS for environment expansion with DiAPO for adaptive capability optimization. We further construct evaluation datasets covering everyday communication, game interaction, and long-horizon character interaction. Extensive experiments across diverse models and scenarios demonstrate improved interaction autonomy and naturalness, validate the reliability, discriminativeness, and agreement with human rankings of CAPS-Eval, and confirm the effectiveness of our training paradigm. Together, these components provide a unified framework for developing credible human-like social agents in open-ended interaction.
Chinese Translation
大型语言模型(LLM)正越来越多地被部署为社交智能体,然而可信的类人交互所需要的不仅仅是流畅的回应或角色一致性。智能体必须自主决定是否、何时以及如何沟通,同时适应不断演变的情境、目标和关系。然而,现有研究缺乏一种统一的方法,来在连续、开放式的交互中赋能、评估并改进此类能力。我们提出 AnthroDial,一个从三个互补方面开发拟人化社交智能体的统一框架:MindFlow,一种轻量级交互框架,它通过动态 Mind Buffer 实现自主、异步且自适应的沟通;CAPS-Eval,一个基于理论的框架,用于评估拟人化交互的认知、情感和行为维度;以及一种可扩展的训练范式,它将用于环境扩展的 SEEDS 与用于自适应能力优化的 DiAPO 相结合。我们进一步构建了覆盖日常交流、游戏交互和长时程角色交互的评估数据集。在不同模型和场景下开展的大量实验表明,交互自主性和自然性得到提升,验证了 CAPS-Eval 的可靠性、区分度以及与人类排名的一致性,并证实了我们训练范式的有效性。总之,这些组件共同提供了一个统一框架,用于在开放式交互中开发可信的类人社交智能体。
cs.CL / 79 / 2609.37914
The Unequal Influence of Bad Advice: Using Training Data Attribution to Modulate Emergent Misalignment
坏建议的不平等影响:利用训练数据归因来调控涌现性失准
large language model
大语言模型相关
Abstract
Fine-tuning large language models on narrow, misaligned tasks can undo their post-training alignment and induce novel misaligned behaviors -- a phenomenon known as \emph{emergent misalignment} (EM). EM has been linked to persona-like representations, where fine-tuning might reduce loss by amplifying a harmful or 'evil' persona. It remains unclear which properties of the training data drive this effect: whether all harmful examples contribute approximately equally to misalignment and whether different models are equally affected by the same fine-tuning examples. In this work, we use training data attribution to quantitatively estimate how much each harmful example contributes to EM. We benchmark the quality of the attribution via retraining -- a sound attribution score should enable us to enhance or attenuate EM by filtering data on that score. Score-based filtering can substantially enhance or attenuate EM; we find that both data-attribution scores and a black-box harmfulness score can identify consequential examples. All models we test become misaligned when trained on the same dataset, and influence scores perform best when filtering data from the same model that computed them. We find cross-model generalization of influence scores from scores derived from the three model families we tested, but this generalization does not recover same model filtering performance.
Chinese Translation
在狭窄且失准的任务上微调大语言模型,可能会破坏其后训练对齐,并诱导出新的失准行为——这一现象被称为\emph{涌现性失准}(emergent misalignment, EM)。EM 已被与类人格(persona-like)表征联系起来,其中微调可能通过放大一个有害或“邪恶”的人格来降低损失。目前仍不清楚训练数据的哪些属性驱动了这一效应:所有有害样本是否对失准的贡献大致相同,以及不同模型是否同等程度地受到相同微调样本的影响。在这项工作中,我们使用训练数据归因来定量估计每个有害样本对 EM 的贡献程度。我们通过重新训练来基准测试归因的质量——一个可靠的归因分数应当使我们能够通过依据该分数过滤数据来增强或减弱 EM。基于分数的过滤可以显著增强或减弱 EM;我们发现,数据归因分数和黑盒有害性分数都能识别出有影响的样本。我们测试的所有模型在同一数据集上训练时都会变得失准,而当过滤由计算这些分数的同一模型的数据时,影响分数表现最佳。我们发现,来自我们测试的三个模型族所推导出的影响分数具有跨模型泛化性,但这种泛化并不能恢复同一模型过滤的性能。
cs.CL / 80 / 2609.38036
Gender bias across LLMs is common and highly heterogenous
大型语言模型中的性别偏见普遍存在且高度异质
large language model
大语言模型相关
Abstract
Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support tools with real consequences. Prior research has focused only on a small set of models, leaving open the extent to which gender biases are common and heterogeneous across LLMs. We address this gap across ten models released between April 2025 and June 2026, spanning nine vendors, using two paradigms: gender attribution to stereotyped phrases (Study 1) and moral judgment of abuse or torture against a woman or a man to prevent a catastrophic outcome (Study 2). In Study 1, two of ten models attributed masculine-stereotyped phrases to female writers more often than the reverse, while three models showed the opposite pattern. In Study 2, several models converged on a male-disadvantaging asymmetry that was directionally consistent with a documented human tendency to protect female targets from harm, though the specific conditions under which this asymmetry emerged varied by model; three other models, by contrast, showed no variation across conditions. These results indicate that gender-related biases are common in LLMs. Their direction and magnitude, however, are highly heterogeneous, to the point that some models behave in diametrically opposite ways to others. Bias auditing should therefore be treated as an ongoing, multi-vendor process, rather than a one-time assessment.
Chinese Translation
理解大型语言模型(LLMs)中的性别偏见日益重要,因为这些系统正被嵌入到会带来真实后果的决策支持工具中。先前的研究仅关注一小组模型,因而未能确定性别偏见在各类LLM中普遍存在和异质化的程度。我们通过两种范式,在2025年4月至2026年6月间发布的、涵盖九家供应商的十个模型上填补这一空白:将性别归因于刻板印象短语(研究1),以及对为防止灾难性后果而对女性或男性实施虐待或酷刑的道德判断(研究2)。在研究1中,十个模型中有两个将男性刻板印象短语归因于女性作者的情况比相反情况更频繁,而三个模型表现出相反模式。在研究2中,若干模型趋同于一种不利于男性的不对称性,这种不对称性在方向上与有文献记载的人类保护女性目标免受伤害的倾向一致,尽管这种不对称性出现的具体条件因模型而异;相比之下,另外三个模型在不同条件之间没有表现出差异。这些结果表明,与性别相关的偏见在LLM中普遍存在。然而,它们的方向和幅度高度异质,以至于一些模型的行为与其他模型截然相反。因此,偏见审计应被视为一个持续的、多供应商的过程,而非一次性的评估。
cs.CR / 81 / 2609.36121
Render Before Reading: Visual Rendering as a Prompt Injection Defense
先渲染后阅读:视觉渲染作为一种提示注入防御
large language model
大语言模型相关
Abstract
Large language models are vulnerable to prompt injection attacks, where third-party adversarial content can hijack the model's behavior. In this paper, we study the role played by the adversarial data's input modality, and identify a systematic asymmetry: multimodal LLMs are more likely to follow adversarial instruction when they appear as text than when the same instruction is delivered through a non-textual channel (e.g., as an image). We hypothesize that this modality gap arises from text-centric instruction tuning, which teaches models to obey textual instructions while treating other modalities mainly as content to parse or describe. We then demonstrate how this gap can be turned into a training-free defense, by rendering all untrusted payloads as typographic images (or audio) before they reach the model. Across ten models and two prompt injection benchmarks (DirectInject and AgentDojo) we show that our defense Pictionary consistently reduces attack success rates even against the strongest adaptive attacks and human red teamers, while largely preserving benign utility. We further show that benign fine-tuning on image-rendered instructions erodes the modality gap, tracing it to the text-centric instruction-tuning distribution.
Chinese Translation
大语言模型容易受到提示注入攻击,其中第三方对抗性内容可以劫持模型的行为。在本文中,我们研究了对抗数据的输入模态所发挥的作用,并识别出一种系统性的不对称性:多模态 LLM 在对抗性指令以文本形式出现时,比同一条指令通过非文本通道(例如以图像形式)传递时更可能遵循该指令。我们假设这种模态差距源于以文本为中心的指令微调,它教导模型服从文本指令,而将其他模态主要视为需要解析或描述的内容。然后我们展示了如何将这种差距转化为一种无需训练的防御,方法是在所有不可信载荷到达模型之前,将其渲染为排版图像(或音频)。在十个模型和两个提示注入基准(DirectInject 和 AgentDojo)上,我们表明我们的防御 Pictionary 即使在面对最强的自适应攻击和人类红队测试者时,也能持续降低攻击成功率,同时大体上保持良性效用。我们进一步表明,在图像渲染的指令上进行良性微调会侵蚀这种模态差距,并将其追溯到以文本为中心的指令微调分布。
cs.CR / 82 / 2609.36603
Self-Evolving Defense: Continual Security Policy Learning for LLM Agents
自我演化防御:面向 LLM 智能体的持续安全策略学习
large language model
大语言模型相关
Abstract
Large language models (LLMs) increasingly power agents that access sensitive information, use external tools, and modify software repositories. Although these capabilities offer substantial benefits, they also create security risks such as jailbreaks, prompt injection, and vulnerable code generation. Existing defenses often require retraining, fail to adapt to evolving attacks, or address only a single threat pattern. To address these limitations, we propose Self-Evolving Defense (SED), a training-free framework that distills harmful agent trajectories into reusable security policies without updating model weights. By retrieving relevant policies for future tasks, SED continually adapts to new attacks while retaining knowledge across attack scenarios. To evaluate the effectiveness of SED, we test it with three open-source models (DeepSeek V4 Flash, GLM 5.2, and Kimi K3) on eight benchmarks that span jailbreaks, prompt injection, and insecure code generation. SED lowers targeted prompt-injection success on AGENTDOJO to 0.42%, compared with 3.7% for the best baseline defense, and holds adaptive X-TEAMING attack success on HARMBENCH to 7.8%, more than four times lower than the best baseline at 35.2%, while preserving benign task utility.
Chinese Translation
大语言模型(LLM)日益为能够访问敏感信息、使用外部工具并修改软件仓库的智能体提供能力支撑。尽管这些能力带来了显著收益,它们也造成了安全风险,例如越狱、提示注入以及生成有漏洞的代码。现有防御往往需要重新训练,无法适应不断演变的攻击,或者仅能应对单一威胁模式。为解决这些局限,我们提出自我演化防御(SED),这是一个无需训练的框架,它能在不更新模型权重的情况下,将有害的智能体轨迹提炼为可复用的安全策略。通过为未来任务检索相关策略,SED 在持续适应新攻击的同时,保留跨攻击场景的知识。为评估 SED 的有效性,我们在涵盖越狱、提示注入和不安全代码生成的八个基准上,使用三个开源模型(DeepSeek V4 Flash、GLM 5.2 和 Kimi K3)对其进行了测试。SED 将 AGENTDOJO 上的定向提示注入成功率降至 0.42%,而最佳基线防御为 3.7%;并将 HARMBENCH 上的自适应 X-TEAMING 攻击成功率保持在 7.8%,比最佳基线的 35.2% 低四倍以上,同时保持良性任务效用。
cs.CR / 83 / 2609.36941
Practical Secrets Extraction against Black-box LLMs
针对黑盒大语言模型的实用秘密提取
large language model
大语言模型相关
Abstract
Large language models (LLMs) increasingly power autonomous coding agents such as Codex and Claude Code, yet their training corpora may contain confidential credentials exposed in public repositories or collected from private development artifacts, creating risks of memorization and subsequent leakage. Existing extraction audits, however, largely assume access to model weights or token probabilities. In this work, we present a black-box secret extraction framework for commercial, API-based LLMs under output-only access. It comprises (i) \emph{Cross-Validated Secret Knowledge Distillation}, which uses semantics-preserving prompt variants, response cross-validation, and provider-specific format filtering to distill secret-relevant behavior into a local white-box proxy; and (ii) \emph{Proxy-Guided Secret Extraction and Candidate Filtering}, which combines truncated top-$p$ sampling with local token entropy, $N$-gram frequency profiling, and provider-specific structural priors. On controlled API-key benchmarks, our framework improves recovery effectiveness and real-key rates over representative baselines while reducing extraction latency. A responsible real-world evaluation further recovers masked provider-specific credentials from three independently deployed black-box LLM systems spanning OpenAI and Claude Code, showing that memorized secrets can be exposed under output-only access.
Chinese Translation
大语言模型(LLMs)正日益为 Codex 和 Claude Code 等自主编码智能体提供支持,然而其训练语料可能包含暴露在公共仓库中的机密凭据,或从私有开发产物中收集的机密凭据,这带来了记忆以及随后泄露的风险。然而,现有的提取审计在很大程度上假设可以访问模型权重或 token 概率。在本工作中,我们提出了一种面向商业、基于 API 的 LLMs 的黑盒秘密提取框架,适用于仅能访问输出的场景。它包含(i)\emph{交叉验证的秘密知识蒸馏},其使用保持语义的提示变体、响应交叉验证以及提供商特定的格式过滤,将秘密相关行为蒸馏到本地白盒代理中;以及(ii)\emph{代理引导的秘密提取与候选过滤},其将截断的 top-$p$ 采样与本地 token 熵、$N$-gram 频率分析和提供商特定的结构先验相结合。在受控的 API 密钥基准上,我们的框架在降低提取延迟的同时,相较于代表性基线提高了恢复有效性和真实密钥率。一项负责任的真实世界评估进一步从三个独立部署的、涵盖 OpenAI 和 Claude Code 的黑盒 LLM 系统中恢复了被遮蔽的提供商特定凭据,表明被记忆的秘密可以在仅输出访问下被暴露。
cs.CR / 84 / 2609.36956
Controlled Decoding Attacks on Black-Box LLMs
针对黑盒大语言模型的受控解码攻击
large language model
大语言模型相关
Abstract
Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing probabilities from sampled outputs offers a possible alternative, but finite sampling produces sparse and noisy estimates, while repeating this process at every generation step incurs substantial query costs. Our empirical observations suggest that large distributional changes along successful jailbreak trajectories are concentrated at a small subset of positions, motivating selective control. We introduce \method{}, a framework for jailbreaking through text-only continuation interfaces that permit repeated sampling and assistant-prefix continuation. Sample-Based Distribution Reconstruction combines sampled outputs with a prior over unobserved actions to obtain a usable control signal. Risk-Gated Residual Control uses the evolving response prefix to decide when to reconstruct and modify the distribution, concentrating sampling costs at selected positions. Speculative Multi-Token Execution further amortizes target calls by verifying and accepting draft prefixes that require no intervention. Across four target endpoints and three benchmarks, \method{} achieves the highest mean score most comparisons against baselines.
Chinese Translation
在生成过程中操纵下一 token 概率可以绕过大型语言模型的安全对齐。然而,现有方法依赖于对模型权重或数值 token 概率的访问,因此不适用于仅返回采样文本的接口。从采样输出中重建概率提供了一种可能的替代方案,但有限采样会产生稀疏且有噪声的估计,而在每个生成步骤重复这一过程会带来大量查询成本。我们的经验观察表明,沿着成功越狱轨迹的大分布变化集中在少量位置上,这促使我们采用选择性控制。我们引入 \method{},这是一个通过允许重复采样和助手前缀续写的仅文本续写接口进行越狱的框架。基于采样的分布重建将采样输出与关于未观测动作的先验相结合,以获得可用的控制信号。风险门控残差控制利用不断演化的响应前缀来决定何时重建和修改分布,将采样成本集中在选定位置。推测式多 token 执行通过验证和接受无需干预的草稿前缀,进一步摊销目标调用。在四个目标端点和三个基准上,\method{} 在大多数与基线的比较中取得了最高的平均分数。
cs.CR / 85 / 2609.37367
Backdoor Mitigation in Decentralized LLM Fine-Tuning
去中心化 LLM 微调中的后门缓解
large language model
大语言模型相关
Abstract
Decentralized large language model (LLM) fine-tuning lets organizations collaboratively train a shared LLM on data they cannot pool, without a central coordinator. In every round, each node exchanges a trainable adapter with its neighbors over a communication graph, and then aggregates them. This setting, however, is vulnerable to propagated backdoors, which is a hidden behavior that lets a model perform normally on clean inputs but produce an attacker-chosen output whenever a secret trigger appears. We show that a single node poisoning its own model can backdoor adapters of nodes that have never seen a poisoned example, making them refuse prompts that contain a secret trigger. We present Chorus, a decentralized mechanism that lets each node detect and reject backdoored adapters from its neighbors before aggregation, without requiring shared validation data or any knowledge of the attacker's trigger or target. Chorus judges each adapter by its behavior, using the receiver's own adapter as a trusted reference. Crucially, no node in Chorus judges adapters alone: the receivers of each adapter update probe it independently, pool their findings in the neighborhood, and vote to make a decision. So a backdoor that slips past one receiver is still caught by the others. We evaluate the effectiveness of Chorus using two instruction-tuning datasets and LLM architectures, and against a state-of-the-art baseline. Chorus cuts the average attack success rate (ASR) of the attacker's neighbors from 48-63% to at most 2.2%, within 0.6 percentage points of an omniscient oracle that knows the exact malicious nodes. Even the worst-affected honest node never exceeds 10% ASR, the same bound as the oracle, against up to 78% without defense. This all comes at a negligible communication overhead.
Chinese Translation
去中心化大语言模型(LLM)微调使各组织能够在无法汇聚的数据上协作训练一个共享的 LLM,而无需中心协调者。在每一轮中,每个节点通过通信图与其邻居交换一个可训练的适配器,然后对它们进行聚合。然而,这种设置容易受到传播式后门的攻击,这是一种隐藏行为,它使模型在干净输入上表现正常,但每当出现秘密触发器时就产生攻击者选定的输出。我们表明,单个节点毒化自己的模型,就能使从未见过被投毒样本的节点的适配器被植入后门,使它们拒绝包含秘密触发器的提示。我们提出 Chorus,一种去中心化机制,它使每个节点能够在聚合之前检测并拒绝来自其邻居的被植入后门的适配器,而不需要共享验证数据,也不需要任何关于攻击者触发器或目标的知识。Chorus 依据行为来评判每个适配器,使用接收者自己的适配器作为可信参考。关键在于,Chorus 中没有任何节点单独评判适配器:每个适配器更新的接收者会独立地对其进行探测,在邻域内汇总它们的发现,并通过投票做出决定。因此,即使某个后门逃过了某一个接收者,仍会被其他接收者捕获。我们使用两个指令微调数据集和 LLM 架构,并针对一个最先进的基线,来评估 Chorus 的有效性。Chorus 将攻击者邻居的平均攻击成功率(ASR)从 48–63% 降低到至多 2.2%,与知道确切恶意节点的全知预言机相比差距在 0.6 个百分点以内。即使受影响最严重的诚实节点也从未超过 10% ASR,这与预言机的界限相同,而在无防御情况下最高可达 78%。这一切仅带来可忽略不计的通信开销。
cs.CR / 86 / 2609.37396
Confidence-Guided Protocol IR for LLM-Aided Security Protocol Modeling
面向LLM辅助安全协议建模的置信度引导协议IR
large language model
大语言模型相关
Abstract
Large language models offer a promising interface for translating natural-language protocol descriptions into formal security models, but their outputs remain difficult to trust without expert validation. In this paper, we present a human-in-the-loop framework for generating Tamarin-verifiable formal models of security protocols. Our key observation is that the main correctness bottleneck is the semantic accuracy rather than the syntactic validity of the intermediate protocol representation. To address this problem, we introduce a protocol intermediate representation (IR) that serves as a human-auditable semantic checkpoint between natural-language parsing and formal model generation. The IR explicitly captures protocol participants, message flows, value provenance, cryptographic operations, proof targets, and compromise assumptions. We further design an interactive interface that highlights uncertain fields and guides users to inspect the most critical semantic decisions based on model confidence before model generation. Rather than replacing formal-methods experts, our approach uses LLMs to produce auditable semantic drafts while leveraging verification tools to check the resulting formal models. Code and verification artifacts are available at https://github.com/laplace1002/TamarinAgent.git.
Chinese Translation
大语言模型为将自然语言协议描述转换为形式化安全模型提供了一种有前景的接口,但如果没有专家验证,其输出仍然难以信任。在本文中,我们提出了一种人在回路框架,用于生成可由 Tamarin 验证的安全协议形式化模型。我们的关键观察是,主要的正确性瓶颈在于中间协议表示的语义准确性,而不是其语法有效性。为了解决这个问题,我们引入了一种协议中间表示(IR),它作为自然语言解析与形式化模型生成之间可供人工审计的语义检查点。该 IR 明确捕获协议参与者、消息流、值来源、密码学操作、证明目标和攻陷假设。我们进一步设计了一个交互式界面,该界面会突出显示不确定字段,并引导用户基于模型置信度,在模型生成之前检查最关键语义决策。我们的方法并非要取代形式化方法专家,而是使用 LLM 生成可审计的语义草稿,同时利用验证工具来检查所得的形式化模型。代码和验证工件可在 https://github.com/laplace1002/TamarinAgent.git 获取。
cs.CR / 87 / 2609.37567
Concealing LLM-Based Multi-Agent Topology via Phantom Structure Injection
通过幻影结构注入隐藏基于LLM的多智能体拓扑
large language model
大语言模型相关
Abstract
Driven by the rapid advancement of large language models (LLMs), LLM-based multi-agent systems (MAS) have emerged as a powerful paradigm for collaborative reasoning over complex tasks. A key design element of MAS is the communication topology, which governs information flow among agents and often encodes proprietary knowledge about the system architecture. However, recent work has shown that such topologies can be inferred even in black-box settings by exploiting semantic dependencies in observable reasoning traces, posing significant risks of intellectual property leakage and exposure of system vulnerabilities. To address this threat, we propose MIRAGE, a topology-concealment framework that preserves the genuine communication topology for task execution while shaping adversary-facing semantic evidence toward a carefully constructed phantom topology. Specifically, MIRAGE operates in three stages: (1) phantom topology synthesis, (2) semantic edge realization, and (3) protected MAS execution. It constructs a phantom topology structurally distinct from the genuine one, materializes phantom edges as plausible semantic dependencies, and suppresses source-specific cues that could reveal genuine edges absent from the phantom topology. Extensive experiments across three topology optimization frameworks and four benchmark datasets demonstrate that MIRAGE substantially reduces the effectiveness of topology inference attacks while largely preserving the task utility of the protected MAS.
Chinese Translation
在大语言模型(LLMs)快速发展的推动下,基于LLM的多智能体系统(MAS)已成为一种用于复杂任务协同推理的强大范式。MAS的一个关键设计要素是通信拓扑,它支配着智能体之间的信息流动,并且往往编码了关于系统架构的专有知识。然而,近期工作表明,即使在黑盒设置下,也可以通过利用可观测推理轨迹中的语义依赖关系来推断此类拓扑,这带来了知识产权泄露与系统漏洞暴露的重大风险。为应对这一威胁,我们提出了MIRAGE,一个拓扑隐藏框架,它在任务执行时保留真实的通信拓扑,同时将面向对手的语义证据塑造为朝向一个精心构建的幻影拓扑。具体而言,MIRAGE分三个阶段运行:(1)幻影拓扑合成,(2)语义边实现,(3)受保护的MAS执行。它构建一个在结构上不同于真实拓扑的幻影拓扑,将幻影边具体化为看似合理的语义依赖关系,并抑制那些可能揭示幻影拓扑中不存在的真实边的来源特定线索。在三个拓扑优化框架和四个基准数据集上进行的大量实验表明,MIRAGE显著降低了拓扑推断攻击的有效性,同时大体上保持了受保护MAS的任务效用。
cs.CR / 88 / 2609.37737
Where Do LLMs Decide to Break the Rules? Mechanistic Localization of Prompt Injection Compliance
LLM 在哪里决定打破规则?提示注入依从的机制性定位
large language model
大语言模型相关
Abstract
When a prompt injection attack succeeds, a Large Language Model (LLM) abandons its assigned system role to comply with an adversarial instruction. While prior work has extensively quantified how often this occurs, we ask a more fundamental question: where inside the network does the model actually decide to break the rules? Using layer-by-layer causal activation patching across five models (4B to 32B parameters), we find a clear dissociation: attack information is linearly decodable from the first layer, yet causal leverage over the model's behavior is negligible until a late-layer bottleneck in the final third of the network. Patching this bottleneck reverses compliance in 77--92\% of cases. We show that the compliance mechanism occupies a compact linear subspace (rank-8 in 4B and 14B models, scaling to rank-64 at 32B) and is architecturally stable across varying model families. Finally, we validate our mechanistic account by showing that this causal peak layer is also the representationally optimal site for detecting attacks, outperforming early-layer classifiers that degrade under surface-level obfuscation such as leetspeak substitution. This alignment between causal leverage and detection performance provides converging evidence that the late-layer bottleneck captures decision-relevant computation rather than merely reflecting an artifact of the intervention.
Chinese Translation
当提示注入攻击成功时,大型语言模型(LLM)会放弃其被分配的系统角色,以遵从对抗性指令。尽管先前工作已广泛量化了这种情况发生的频率,但我们提出一个更根本的问题:在网络内部,模型实际上在哪里决定打破规则?通过对五个模型(4B 到 32B 参数)进行逐层因果激活修补,我们发现一种清晰的分离:攻击信息从第一层起就可线性解码,但直到网络最后三分之一的晚层瓶颈之前,对模型行为的因果影响力都可忽略不计。修补这一瓶颈可在 77--92\% 的情况下逆转依从。我们表明,依从机制占据一个紧凑的线性子空间(在 4B 和 14B 模型中秩为 8,在 32B 时扩展到秩为 64),并且在不同的模型家族中具有架构稳定性。最后,我们通过表明这一因果峰值层也是检测攻击的表征最优位置来验证我们的机制性解释,其优于在诸如 leetspeak 替换等表层混淆下会退化的早期层分类器。因果影响力与检测性能之间的这种一致性提供了汇聚性证据,表明晚层瓶颈捕获的是与决策相关的计算,而不仅仅是反映了干预的伪影。
cs.AI / 89 / 2609.36066
AerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search
AerialDojo-200K:面向开放世界空中目标物体搜索的大规模基准套件
large language model
大语言模型相关
Abstract
Open-world aerial object-goal search is a foundational yet challenging task, requiring aerial agents to autonomously explore large-scale, unstructured three-dimensional environments and reach target objects specified by semantic descriptions or reference images, rather than following route-specific instructions. However, research in this task remains at a nascent stage and relies on small, environment-specific benchmarks with heterogeneous action spaces and data formats. These limitations hinder large-scale training and cross-benchmark evaluation, constraining the scalability and generalizability of aerial agents. To address this problem, we propose AerialDojo-200K, a large-scale benchmark suite for open-world aerial object-goal search, with 3 times as many scenes and 18.7 times as many task instances as the largest existing benchmark for this task. Specifically, we construct 42 simulation scenes spanning four scene families and 21 scene types, including 18 urban, 12 natural, six infrastructure, and six disaster scenes. To ensure data quality, 12 annotators spent two months manually annotating 109 landmarks, 2099 target objects, and 2099 object anchors across these scenes. We further construct 205,732 task instances, comprising over 100K semantic-goal and over 100K image-goal instances across Base, Standard, and Long-Horizon settings. Each task instance includes a collision-free reference trajectory and corresponding multi-view video recordings. We also develop a unified evaluation framework with a scene partition comprising 21 in-distribution scenes and 21 out-of-distribution scenes. Finally, our evaluation of five open-source and four closed-source multimodal large language models reveals that there is still a long way to go toward achieving general-purpose aerial agents. All can be found at https://fengtt42.github.io/AerialDojo/.
Chinese Translation
开放世界空中目标物体搜索是一个基础但极具挑战性的任务,它要求空中智能体自主探索大规模、非结构化的三维环境,并到达由语义描述或参考图像指定的目标物体,而不是遵循特定路线的指令。然而,该任务的研究仍处于起步阶段,并依赖于规模较小、针对特定环境的基准,这些基准具有异构的动作空间和数据格式。这些局限阻碍了大规模训练和跨基准评估,制约了空中智能体的可扩展性与泛化能力。为解决这一问题,我们提出了 AerialDojo-200K,一个面向开放世界空中目标物体搜索的大规模基准套件,其场景数量是该任务现有最大基准的 3 倍,任务实例数量是其 18.7 倍。具体而言,我们构建了 42 个仿真场景,涵盖四个场景族和 21 种场景类型,包括 18 个城市场景、12 个自然场景、6 个基础设施场景和 6 个灾害场景。为确保数据质量,12 名标注员耗时两个月,在这些场景中人工标注了 109 个地标、2099 个目标物体和 2099 个物体锚点。我们进一步构建了 205,732 个任务实例,其中包含超过 10 万个语义目标实例和超过 10 万个图像目标实例,覆盖 Base、Standard 和 Long-Horizon 三种设置。每个任务实例都包含一条无碰撞的参考轨迹以及相应的多视角视频记录。我们还开发了一个统一的评估框架,其场景划分包含 21 个分布内场景和 21 个分布外场景。最后,我们对五个开源和四个闭源多模态大语言模型的评估表明,距离实现通用型空中智能体仍有很长的路要走。所有内容均可在 https://fengtt42.github.io/AerialDojo/ 获取。
cs.AI / 90 / 2609.36199
PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents
PreviewDiff:多模态评判器引导的扩散潜变量搜索
diffusion
扩散模型相关
Abstract
Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions. A common way to improve prompt satisfaction is to spend more compute at test time through Best-of-N sampling, but final-sample selection is fixed. Best-of-N can only choose among completed outputs and cannot repair a promising trajectory before it fails. We introduce PreviewDiff, a training-free test-time search method that turns diffusion sampling from scalar search into a multimodal critic-guided search over intermediate latents. At selected denoising checkpoints, PreviewDiff decodes a partial preview, asks a multimodal judge to score and critique it, and uses the resulting natural-language feedback to branch over semantic prompt edits and locally re-noised latent continuations. These branches are then scored and selectively rolled forward, allowing verifier compute to guide generation while the sample is still editable. Across image and video generation benchmarks, PreviewDiff consistently improves over budget-matched Best-of-N selection and strong scalar-search baselines. Ablations show that earlier interventions and increased search width provide the largest gains, while deeper search and additional semantic variants offer complementary improvements. PreviewDiff demonstrates that multimodal feedback is most useful not only as a final verifier, but as an active controller inside the denoising process.
Chinese Translation
扩散模型能够生成引人注目的图像和视频,但它们仍然难以处理使生成结果忠实于提示词的组合细节,例如物体数量、属性绑定、空间关系以及时间上接地的动作。一种提高提示词满足度的常见方法是在测试时通过 Best-of-N 采样花费更多计算量,但最终样本选择是固定的。Best-of-N 只能在已完成输出中进行选择,无法在一条有前景的轨迹失败之前对其进行修复。我们提出 PreviewDiff,这是一种免训练的测试时搜索方法,它将扩散采样从标量搜索转变为对中间潜变量的多模态评判器引导搜索。在选定的去噪检查点,PreviewDiff 解码一个部分预览,要求多模态评判器对其进行评分和评论,并利用由此产生的自然语言反馈,对语义提示词编辑和局部重新加噪的潜变量延续进行分支。然后对这些分支进行评分,并有选择地向前推进,从而使验证器计算量能够在样本仍可编辑时引导生成。在图像和视频生成基准上,PreviewDiff 持续优于预算匹配的 Best-of-N 选择和强标量搜索基线。消融实验表明,更早的干预和增加的搜索宽度带来最大的增益,而更深的搜索和额外的语义变体提供互补的改进。PreviewDiff 表明,多模态反馈不仅作为最终验证器,而且作为去噪过程内部的主动控制器时最为有用。
cs.AI / 91 / 2609.36348
Representation by Design in Generation: Cross-View Class-Token Alignment in Diffusion Transformers
生成中按设计的表征:扩散 Transformer 中的跨视图类令牌对齐
diffusion
扩散模型相关
Abstract
Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas the models' own representations are often treated as a by-product of synthesis. We ask whether diffusion models can instead be trained to learn substantially stronger semantic representations without sacrificing generation quality. SelfFlow takes a step in this direction by introducing self-supervised patch alignment into flow matching, but its main gains remain in faster convergence and improved generation. Inspired by DINO and iBOT, we extend this framework with cross-view class-token alignment to further strengthen semantic representations. Specifically, we form two independently noised, dual-timestep observations of each image and align each student class-token representation with the stop-gradient EMA-teacher target from the other observation. This objective is optimized jointly with the inherited flow-matching and local patch objectives. Notably, although the additional objective acts only on the class token, it strengthens both class-token and patch representations. Compared with a matched two-view baseline, ImageNet linear-probing accuracy improves by 9.4\% using the class token and 10.1\% using mean-pooled patch tokens, while frozen-backbone VOC2012 segmentation improves by 3.6 mIoU. These representation gains are achieved while maintaining comparable ImageNet generation FID. In text-to-image training, the same objective also improves generation FID, reducing it from 2.52 to 2.37 at matched checkpoints. Our results show that representation need not remain a by-product of generation or merely a tool for improving it: it can be directly optimized as a first-class capability of diffusion pretraining alongside generation.
Chinese Translation
生成学习和表征学习之间仍然以不对称的方式相联系:语义表征被用于改进扩散生成,而模型自身的表征通常被视为合成的副产品。我们询问,扩散模型能否转而接受训练,以在不牺牲生成质量的情况下学习显著更强的语义表征。SelfFlow 通过将自监督图块对齐引入流匹配朝这一方向迈出了一步,但其主要收益仍然在于更快的收敛和改进的生成。受 DINO 和 iBOT 启发,我们通过跨视图类令牌对齐扩展该框架,以进一步增强语义表征。具体而言,我们为每张图像形成两个独立加噪的双时间步观测,并将每个学生类令牌表征与来自另一观测的停止梯度 EMA 教师目标对齐。该目标与继承的流匹配目标和局部图块目标联合优化。值得注意的是,尽管额外目标仅作用于类令牌,但它同时增强了类令牌表征和图块表征。与匹配的双视图基线相比,使用类令牌时 ImageNet 线性探测准确率提高 9.4%,使用均值池化的图块令牌时提高 10.1%,而冻结骨干的 VOC2012 分割提高 3.6 mIoU。这些表征收益是在保持相当的 ImageNet 生成 FID 的同时实现的。在文本到图像训练中,同一目标也改进了生成 FID,在匹配的检查点处将其从 2.52 降低到 2.37。我们的结果表明,表征不必仍然是生成的副产品,也不只是改进生成的工具:它可以作为扩散预训练的一等能力,与生成一起被直接优化。
cs.AI / 92 / 2609.36562
ThinkingGuard: Decoding Implicit Hazards via Step-by-Step Risk Attribution in Multimodal Large Language Models
ThinkingGuard:通过逐步风险归因解码多模态大语言模型中的隐式危害
large language model
大语言模型相关
Abstract
While Multimodal Large Language Models (MLLMs) are increasingly deployed in safety-critical domains, their reliability is threatened by multimodal implicit risks. Unlike explicit threats, these hazards emerge when individually benign text and neutral visual entities logically converge to induce unsafe outputs. Current detection methods fail to address this because they overlook the underlying risk activation mechanisms that govern cross-modal risk activation, leading to single-modality shortcut learning and hallucinated rationalizations. To bridge this gap, we first construct TriggerBench, the first dataset explicitly modeling risk compositionality (5,600 instances). By formally isolating Key Elements and Trigger Elements to build counterfactual contrastive pairs, TriggerBench eliminates risk residues and forces models to perform genuine logical deduction rather than superficial pattern matching, which provides a rigorous foundation for both large-scale training and fine-grained evaluation. Building on this, we propose a Step-Supervised Structured Reasoning training framework and employ it to train ThinkingGuard, a specialized guard model. Inspired by Situation Awareness theory, we decouple implicit risk identification into progressive cognitive stages, and utilize a step-reward Monte Carlo Tree Search algorithm to explore optimal reasoning trajectories, which are then distilled into the model through Dual-Constraint Preference Alignment. Extensive experiments across both standard and implicit safety benchmarks demonstrate that ThinkingGuard achieves strong performance. Project resources are available at https://github.com/FroggyChen/ThinkingGuard.
Chinese Translation
尽管多模态大语言模型(MLLMs)正日益部署于安全攸关领域,其可靠性却受到多模态隐式风险的威胁。与显式威胁不同,这些危害产生于单独看无害的文本与中性的视觉实体在逻辑上汇聚并诱发不安全输出之时。当前的检测方法之所以无法应对这一问题,是因为它们忽视了支配跨模态风险激活的底层风险激活机制,从而导致单模态捷径学习和幻觉式合理化。为弥合这一差距,我们首先构建了 TriggerBench,这是首个显式建模风险组合性的数据集(5,600 个实例)。通过形式化隔离关键元素(Key Elements)与触发元素(Trigger Elements)以构建反事实对比对,TriggerBench 消除了风险残留,并迫使模型进行真正的逻辑推演而非表面模式匹配,这为大规模训练和细粒度评估提供了严谨的基础。在此基础上,我们提出了一种步进监督结构化推理(Step-Supervised Structured Reasoning)训练框架,并用其训练了 ThinkingGuard,一个专用的防护模型。受态势感知(Situation Awareness)理论启发,我们将隐式风险识别解耦为渐进的认知阶段,并利用步进奖励蒙特卡洛树搜索(step-reward Monte Carlo Tree Search)算法探索最优推理轨迹,随后通过双重约束偏好对齐(Dual-Constraint Preference Alignment)将这些轨迹蒸馏进模型。在标准安全基准与隐式安全基准上的大量实验表明,ThinkingGuard 取得了强劲的性能。项目资源可在 https://github.com/FroggyChen/ThinkingGuard 获取。
cs.AI / 93 / 2609.36651
FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution
FocusVTC:高效且高性能的自适应分辨率视觉文本压缩
large language model
大语言模型相关
Abstract
Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI spends tokens on irrelevant content. We introduce FocusVTC, which breaks this trade-off through adaptive resolution while preserving general multimodal capabilities. It combines compressed low-DPI global views with selective region enhancement, integrating enhanced views into ongoing reasoning. We construct 29.4K high-quality Reasoning-Evidence Localization (REL) chain-of-thought examples (REL-CoT) that link reasoning traces to page indices and bounding boxes. Multi-resolution REL supervised fine-tuning (REL-SFT) teaches the model to localize relevant regions, and Group Relative Policy Optimization learns when to enhance resolution and how to use the resulting observations, without a separate continual-pretraining stage. At 72 DPI on RULER v1, FocusVTC scores 87.4 at $2.9\times$ input compression, including tool observations, versus 57.5 for Glyph at $3.0\times$ input compression. It surpasses its text-input backbone on LongBench (56.40 versus 55.86), improves the MRCR macro-average by 13.91 points, and achieves a 51.19 macro-average on VTCBench. The MRCR latency evaluation also shows a $2.79\times$ online end-to-end speedup over Text. General multimodal capabilities are preserved, with MMMU increasing from 65.12 to 66.73 and MME from 2424.02 to 2457.62.
Chinese Translation
大语言模型中的长上下文推理会带来大量的计算与内存开销。视觉文本压缩(VTC)通过将文本渲染为图像来缩短输入长度,但固定分辨率渲染会造成压缩与性能之间的权衡:低 DPI 节省 token,却以可读性为代价;而高 DPI 则将 token 花费在无关内容上。我们提出 FocusVTC,它通过自适应分辨率打破这一权衡,同时保持通用的多模态能力。它将压缩后的低 DPI 全局视图与选择性的区域增强相结合,把增强后的视图整合进正在进行的推理过程中。我们构建了 29.4K 条高质量的推理—证据定位(REL)思维链样例(REL-CoT),它们将推理轨迹与页面索引和边界框关联起来。多分辨率 REL 监督微调(REL-SFT)教会模型定位相关区域,而组相对策略优化则学习何时提升分辨率以及如何使用由此得到的观察结果,且无需单独的持续预训练阶段。在 RULER v1 上,当 DPI 为 72 时,FocusVTC 在包含工具观察结果的 $2.9\times$ 输入压缩下得分 87.4,而 Glyph 在 $3.0\times$ 输入压缩下为 57.5。它在 LongBench 上超越了其文本输入主干(56.40 对 55.86),将 MRCR 宏平均提升了 13.91 分,并在 VTCBench 上取得了 51.19 的宏平均。MRCR 延迟评估还显示,相比 Text,其在线端到端速度提升了 $2.79\times$。通用多模态能力得到保持,MMMU 从 65.12 提升到 66.73,MME 从 2424.02 提升到 2457.62。
cs.CL / 94 / 2609.36798
Seeing What Should Be Heard: Diagnosing and Repairing Cross-Modal Shortcuts in Omni-Modal LLMs
看见本应听到的内容:诊断并修复全模态大语言模型中的跨模态捷径
large language model
大语言模型相关
Abstract
Omni-modal large language models (LLMs) are expected to answer a question using the modality it explicitly refers to. However, existing training paradigms rarely verify whether models actually follow this modality, because multimodal inputs from the same sample often provide redundant evidence for the same answer. In this work, we uncover a pervasive cross-modal shortcut in omni-modal LLMs: when asked an audio-related question, models rely on the image as much as on the audio, and sometimes even more. To systematically diagnose this behavior, we introduce the Factorized Modality Diagnostic, which independently swaps audio and images between samples to isolate each modality's causal contribution. Across two model families in different settings, we find that this shortcut persists throughout supervised fine-tuning and reinforcement learning post-training, while judge-based RL may further amplify such reliance on irrelevant visual information. Based on this finding, we propose DMC-Repair, which trains models on the same kind of cross-modal swapped samples while assigning supervision according to the modality specified by the question. This prevents models from exploiting the spurious correspondence between modalities within the same clip. Experiments demonstrate that DMC-Repair reduces the image-induced share of the answer effect by 59.9%, effectively suppressing the cross-modal shortcut without compromising audio-question answering performance. The reduction in shortcut reliance generalizes across two model families and zero-shot to an unseen dataset and an unseen benchmark, and persists through subsequent post-training. Code is available at https://anonymous.4open.science/r/DMC-Repair.
Chinese Translation
全模态大语言模型(LLMs)被期望使用问题明确指涉的模态来回答该问题。然而,现有训练范式很少验证模型是否真的遵循这一模态,因为来自同一样本的多模态输入常常为同一答案提供冗余证据。在这项工作中,我们揭示了全模态 LLMs 中一种普遍存在的跨模态捷径:当被问到一个与音频相关的问题时,模型对图像的依赖与对音频的依赖一样多,有时甚至更多。为了系统地诊断这一行为,我们引入了因子化模态诊断(Factorized Modality Diagnostic),它独立地在样本之间交换音频和图像,以分离每种模态的因果贡献。在不同设置下的两个模型家族中,我们发现这种捷径在监督微调和强化学习后训练期间持续存在,而基于评判器的 RL 可能进一步放大对无关视觉信息的这种依赖。基于这一发现,我们提出了 DMC-Repair,它在同类跨模态交换样本上训练模型,同时根据问题所指定的模态来分配监督。这防止模型利用同一片段内模态之间的虚假对应关系。实验表明,DMC-Repair 将图像诱导的答案效应占比降低了 59.9%,有效地抑制了跨模态捷径,且没有损害音频问题回答性能。捷径依赖的降低在两个模型家族中具有泛化性,并能零样本泛化到未见过的数据集和未见过的基准,且在后续后训练中持续存在。代码可在 https://anonymous.4open.science/r/DMC-Repair 获取。
cs.LG / 95 / 2609.36837
You Cannot Recover What Was Never Measured: Quantifying the Information Ceiling of Ultra-Low-Field MRI Super-Resolution
你无法恢复从未被测量的东西:量化超低场 MRI 超分辨率的信息上限
diffusion
扩散模型相关
Abstract
Generative super-resolution models can turn portable 64 mT MRI into images that look like 3T scans, and the field evaluates them with PSNR, SSIM, and pixelwise uncertainty, most often on pairs built by synthetically degrading high-field images. Prior work acknowledges that these models hallucinate and that the problem is ill posed, but to our knowledge no study measures how much information about the individual subject the real low-field scan actually contains. We measure it. Using paired 64 mT and 3T scans of the same subjects from three public datasets, and a measurement protocol validated on tests whose correct answer is known in advance, we find that, judged over the whole brain, real 64 mT scans carry structure specific to the individual only down to approximately 3 to 4 mm half-pitch in plane, and coarser still through plane. Standard synthetic degradations preserve subject information roughly 1 mm beyond this ceiling, so models trained and benchmarked on synthetic pairs are evaluated on information that real scanners never record. We then test trained diffusion models and a publicly released external model on real paired acquisitions; 24 trained runs of five architectures (GAN, diffusion, and transformer families) give the coverage of the audit. On every subject where faithfulness can be measured, fine output detail is no more correlated with the subject's own 3T scan than with a stranger's, while sample-variance uncertainty does not distinguish fabricated structure from reconstruction difficulty. Because PSNR and SSIM score resemblance to a reference rather than whether detail belongs to the subject, a benchmark scored by them cannot tell recovery from fabrication. Code for the measurement protocol will be released so that recoverability claims can be tested for newer models.
Chinese Translation
生成式超分辨率模型能够将便携式 64 mT MRI 转换为看起来像 3T 扫描的图像,而该领域通常使用 PSNR、SSIM 和逐像素不确定性来评估它们,且最常见的是在通过合成退化高场图像构建的成对数据上进行评估。先前工作承认这些模型会产生幻觉,并且该问题是不适定的,但据我们所知,尚无研究测量真实低场扫描实际包含多少关于个体受试者的信息。我们对此进行了测量。使用来自三个公共数据集的同一受试者的 64 mT 和 3T 配对扫描,以及在正确答案事先已知的测试上验证过的测量协议,我们发现,在全脑范围内判断,真实 64 mT 扫描携带的个体特异结构在平面内仅低至约 3 至 4 mm 半间距,而在穿平面方向上还要更粗。标准合成退化将受试者信息保留到大约超出该上限 1 mm 的程度,因此在合成配对上训练和基准测试的模型,是在真实扫描仪从未记录的信息上进行评估。然后,我们在真实配对采集上测试了训练过的扩散模型和一个公开发布的外部模型;五个架构家族(GAN、扩散和 Transformer 家族)的 24 次训练运行给出了该审计的覆盖范围。在每一个能够测量忠实度的受试者上,输出中的精细细节与该受试者自身 3T 扫描的相关性,并不高于与一个陌生人的 3T 扫描的相关性;而样本方差不确定性无法将捏造结构与被重建难度区分开来。因为 PSNR 和 SSIM 评分的是与参考图像的相似程度,而不是细节是否属于该受试者,所以由它们评分的基准无法区分恢复与捏造。测量协议的代码将会发布,以便能够针对更新的模型检验可恢复性主张。
cs.AI / 96 / 2609.37001
Parameterized Stripe Attention for Efficient Video Generation
面向高效视频生成的参数化条纹注意力
diffusion
扩散模型相关
Abstract
Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the computationally expensive full spatio-temporal attention. While sparse attention methods offer potential solutions, existing approaches face an inherent flexibility--efficiency dilemma: predefined masks lack the flexibility to capture diverse attention patterns, while runtime-determined masks introduce overheads and sacrifice hardware efficiency. We identify the lack of a unified structural characterization of DiT attention as a key limitation of existing methods, and establish that video DiT attention exhibits \textbf{periodic diagonal stripe structures} along both temporal and spatial dimensions. To formally encode these structured patterns within a single efficient kernel, we present {\bf PSA}, a parameterized stripe attention that formalizes the observed stripe regularity, unifying diverse attention patterns for efficient mask generation. This unified representation enables a single hardware-efficient CUDA kernel to process all sparse patterns, achieving FlashAttention-3-level Model FLOPs Utilization. To determine optimal sparsity configurations, we propose a training-free offline search algorithm that automatically maximizes sparsity under a specified error tolerance for each attention head. Experiments on HunyuanVideo and Wan~2.1 demonstrate that PSA achieves 1.57$\times$ and 1.37$\times$ end-to-end speedups over FlashAttention-3 baselines, with acceptable visual quality degradation.
Chinese Translation
扩散 Transformer(DiTs)能够实现高质量视频生成,但存在显著的推理延迟,这主要归因于计算代价高昂的全时空注意力。尽管稀疏注意力方法提供了潜在解决方案,但现有方法面临固有的灵活性——效率困境:预定义掩码缺乏捕获多样化注意力模式的灵活性,而运行时确定的掩码则引入额外开销并牺牲硬件效率。我们将 DiT 注意力缺乏统一的结构化刻画识别为现有方法的关键局限,并确立视频 DiT 注意力在时间维度和空间维度上均表现出 \textbf{周期性对角条纹结构}。为了在单个高效内核中形式化地编码这些结构化模式,我们提出 {\bf PSA},一种参数化条纹注意力,它形式化了所观察到的条纹规律,将多样化的注意力模式统一起来以实现高效掩码生成。这种统一表示使单个硬件高效的 CUDA 内核能够处理所有稀疏模式,达到 FlashAttention-3 级别的模型 FLOPs 利用率。为了确定最优稀疏配置,我们提出一种免训练的离线搜索算法,该算法在针对每个注意力头指定的误差容限下自动最大化稀疏度。在 HunyuanVideo 和 Wan~2.1 上的实验表明,PSA 相比 FlashAttention-3 基线实现了 1.57$\times$ 和 1.37$\times$ 的端到端加速,且视觉质量下降可接受。
cs.LG / 97 / 2609.37042
GleanVID: Complementary Token Selection for Efficient Video Large Language Models
GleanVID:面向高效视频大语言模型的互补 Token 选择
large language model
大语言模型相关
Abstract
Video Large Language Models (VideoLLMs) have achieved strong video understanding capabilities but incur substantial inference overhead due to the large number of visual tokens. Existing VideoLLM token compression methods largely rely on selection-independent scoring, overlooking cross-frame complementarity and consequently retaining redundant evidence across frames. Instead, we view video token selection as a progressive evidence accumulation process. It aims to retain visual evidence that is individually informative and collectively complementary under a limited token budget. Building on this insight, we introduce GleanVID, a training-free inference acceleration framework for VideoLLMs. Specifically, GleanVID first allocates the global token budget across frames according to temporal novelty and then selects tokens by jointly considering local representativeness and subspace complementarity, thereby preserving richer and less redundant visual evidence. Extensive experiments across diverse VideoLLMs and benchmarks demonstrate that GleanVID consistently achieves state-of-the-art performance. Notably, with only 25% of visual tokens, GleanVID preserves 98.6% of Qwen3-VL's original performance while reducing its prefill latency by 44.7%. On LLaVA-OV-7B, GleanVID at a 25% retention ratio even slightly surpasses the original model.
Chinese Translation
视频大语言模型(VideoLLMs)已具备强大的视频理解能力,但由于视觉 token 数量庞大,会带来大量的推理开销。现有的 VideoLLM token 压缩方法大多依赖于与选择无关的评分,忽视了跨帧互补性,因而在帧间保留了冗余证据。相反,我们将视频 token 选择视为一个渐进式证据累积过程。其目标是在有限的 token 预算下,保留那些单独来看具有信息量且整体上互补的视觉证据。基于这一洞见,我们提出了 GleanVID,一个面向 VideoLLMs 的免训练推理加速框架。具体而言,GleanVID 首先根据时间新颖性在帧间分配全局 token 预算,然后通过联合考虑局部代表性和子空间互补性来选择 token,从而保留更丰富且冗余更少的视觉证据。在多种 VideoLLMs 和基准上的大量实验表明,GleanVID 始终达到最先进的性能。值得注意的是,仅使用 25% 的视觉 token,GleanVID 就保留了 Qwen3-VL 原始性能的 98.6%,同时将其预填充延迟降低了 44.7%。在 LLaVA-OV-7B 上,GleanVID 在 25% 的保留率下甚至略微超过了原始模型。
cs.LG / 98 / 2609.37147
Improved Distributional Diffusion Models
改进的分布型扩散模型
diffusion
扩散模型相关
Abstract
Distributional Diffusion Models (DDMs) replace the standard mean-prediction denoiser with a \emph{distributional} denoiser trained via a scoring rule objective, learning a stochastic approximation to $p(x_1 \mid x_t)$ rather than its conditional mean. However, scaling DDMs to modern image-generation settings faces two obstacles: (i) multi-particle training incurs overhead that scales with the number of particles, (ii) DDMs use globally fixed scoring rule hyperparameters, forcing a single trade-off across sampling budgets. We mitigate these limitations by deferring particle expansion to late transformer layers, and the hyperparameter trade-off by introducing time-dependent scoring rule schedules informed by the dynamical regimes of~\citet{Biroli2024}. Combined with a DiT-based latent setup, these changes make DDM training practical on class-conditional ImageNet-$256^2$, achieving 4.48 FID at 4 steps and 2.38 at 50 steps with DiT-XL/2, from a single model trained from scratch in one stage, without a teacher, self-distillation or JVPs. The result is a stochastic few-step generator whose FID does not degrade as the sampling budget grows from 4 to 50 NFE, and the same recipe transfers to text-to-image generation. Code and pre-trained models available at https://github.com/CompVis/iDDM.
Chinese Translation
分布型扩散模型(DDMs)将标准的均值预测去噪器替换为通过评分规则目标训练的 \emph{分布型} 去噪器,学习对 $p(x_1 \mid x_t)$ 的随机近似,而非其条件均值。然而,将 DDM 扩展到现代图像生成设置面临两个障碍:(i) 多粒子训练会带来随粒子数量增长的开销,(ii) DDM 使用全局固定的评分规则超参数,从而迫使在采样预算之间做出单一的权衡。我们通过将粒子扩展推迟到较晚的 Transformer 层来缓解这些局限,并通过引入由~\citet{Biroli2024} 的动力学区域所启发的时间相关评分规则调度来缓解超参数权衡。结合基于 DiT 的潜空间设置,这些改动使得 DDM 训练在类别条件 ImageNet-$256^2$ 上变得切实可行,使用 DiT-XL/2 在 4 步时达到 4.48 FID、在 50 步时达到 2.38,且这一切来自一个单阶段从零训练的单一模型,无需教师、自蒸馏或 JVP。其结果是一个随机的少步生成器,其 FID 不会随着采样预算从 4 增长到 50 NFE 而退化,而且同样的方案可以迁移到文本到图像生成。代码与预训练模型见 https://github.com/CompVis/iDDM。
cs.AI / 99 / 2609.37225
ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression
ResComEmb:通过残差同质性压缩实现有效且高效的多模态嵌入
large language model
大语言模型相关
Abstract
Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods either compress each input into a single vector, limiting fine-grained expressiveness, or retain long sequences of visual-token vectors, incurring substantial storage and interaction costs. To resolve this trade-off, we propose ResComEmb, a trainable framework for effective and efficient universal multi-vector multimodal embedding. ResComEmb first encodes each input at native dynamic resolution into ordered global, intermediate, and fine-grained views. After MLLM contextualization and embedding projection, a trainable Residual Homogeneity Compression (RHC) module reduces within-granularity redundancy and cross-granularity repetition under explicit visual token budgets. Then, ResComEmb introduces a length-adaptive Bidirectional Late-Interaction Matching mechanism for robust query-document scoring, which averages the strongest token-level matches in each direction and combines the two scores using a weight based on how many valid tokens each side has. Extensive experiments on MMEB, ViDoRe V1, and ViDoRe V2 show that ResComEmb produces higher-quality universal multimodal embeddings than VLM2Vec-V2, and outperforms ColQwen2.5 in visual document retrieval using only 37.5% of its full visual token budget, demonstrating a favorable effectiveness-efficiency trade-off.
Chinese Translation
多模态大语言模型(MLLMs)已展现出在通用多模态表示学习方面的强大潜力。然而,现有方法要么将每个输入压缩为单个向量,从而限制了细粒度的表达能力,要么保留长序列的视觉词元向量,从而带来大量的存储与交互开销。为解决这一权衡问题,我们提出了 ResComEmb,一个可训练的框架,用于实现有效且高效的通用多向量多模态嵌入。ResComEmb 首先以原生动态分辨率将每个输入编码为有序的全局、中间和细粒度视图。在 MLLM 上下文建模与嵌入投影之后,一个可训练的残差同质性压缩(RHC)模块在明确的视觉词元预算下减少粒度内冗余与跨粒度重复。然后,ResComEmb 引入了一种长度自适应的双向后期交互匹配机制,用于稳健的查询—文档打分,该机制对每个方向上最强的词元级匹配取平均,并基于每一侧有效词元的数量使用一个权重将两个分数组合起来。在 MMEB、ViDoRe V1 和 ViDoRe V2 上的大量实验表明,ResComEmb 生成的通用多模态嵌入质量高于 VLM2Vec-V2,并且在仅使用其完整视觉词元预算的 37.5% 的情况下,在视觉文档检索中优于 ColQwen2.5,展现出良好的有效性—效率权衡。
cs.LG / 100 / 2609.37529
Principled MAP estimation for inverse problems: bridging the gap between convergence and performance
逆问题的原则化MAP估计:弥合收敛性与性能之间的差距
diffusion
扩散模型相关
Abstract
Pretrained denoisers provide a powerful way to incorporate image priors into restoration algorithms. Plug-and-Play and RED approaches exploit fixed-noise-level denoisers within first-order optimization schemes, with convergence guarantees, but often struggle to achieve high-quality reconstruction on severely ill-posed inverse problems. In contrast, recent state-of-the-art approaches leverage denoisers derived from flow- or diffusion-based generative models and evaluate them along a sequence of decreasing noise levels. While these methods achieve strong empirical performance, their convergence theory remains limited. In this paper, we bridge this gap by specifically designing an algorithm that combines denoisers at decreasing noise levels with a schedule tailored to ensure convergence. From a Bayesian perspective, we prove that our method converges to a $\textit{Maximum a Posteriori}$ (MAP) estimate, under suitable assumptions. Subsequently, we apply our method to various ill-posed inverse problems and show that it surpasses convergent methods while competing with state-of-the-art empirical ones.
Chinese Translation
预训练去噪器为将图像先验融入复原算法提供了一种强大的方式。即插即用(Plug-and-Play)与RED方法在一阶优化方案中利用固定噪声水平的去噪器,并具有收敛性保证,但往往难以在严重病态的逆问题上实现高质量重建。相比之下,近期最先进的方法利用源自基于流(flow)或基于扩散(diffusion)的生成模型的去噪器,并沿着一系列递减的噪声水平对它们进行评估。尽管这些方法取得了很强的经验性能,但其收敛理论仍然有限。在本文中,我们通过专门设计一种算法来弥合这一差距,该算法将递减噪声水平下的去噪器与为确保收敛而量身定制的调度方案相结合。从贝叶斯视角出发,我们证明,在适当的假设下,我们的方法收敛到一个 $\textit{Maximum a Posteriori}$ (MAP) 估计。随后,我们将该方法应用于各种病态逆问题,并表明它超越了具有收敛性的方法,同时可与最先进的经验性方法相竞争。
cs.AI / 101 / 2609.37581
TReVS: Integrating Textual Relevance and Visual Saliency for Efficient Vision-Language Model Token Pruning
TReVS:整合文本相关性与视觉显著性以实现高效的视觉-语言模型 Token 剪枝
large language model
大语言模型相关
Abstract
Vision-Language Models (VLMs) excel at visual understanding and reasoning but often incur substantial inference costs due to the large number of visual tokens. Recent visual token pruning methods increasingly follow a two-stage paradigm: they first remove visually redundant tokens after the vision encoder and then discard tokens irrelevant to the textual query within the Large Language Model (LLM). However, since the first stage typically relies solely on vision-encoder saliency, it may prematurely eliminate query-relevant tokens, depriving the subsequent text-guided stage of critical visual evidence. Our empirical analysis shows that incorporating query guidance into first-stage pruning better preserves task-relevant evidence and consistently improves performance over vision-only saliency-based pruning. We further find that high-variance attention heads are more sensitive to the textual query and yield more discriminative text-to-vision attention signals for second-stage pruning. Motivated by these findings, we propose TReVS, a training-free framework that combines textual relevance with vision-encoder saliency for pre-LLM pruning and leverages high-variance attention heads to remove task-irrelevant tokens at shallow-to-intermediate layers of the LLM. On LLaVA-1.5-7B, TReVS retains 92.8% of the unpruned baseline performance while pruning 94.4% of visual tokens, outperforming prior state-of-the-art methods.
Chinese Translation
视觉-语言模型(VLMs)在视觉理解与推理方面表现出色,但由于视觉 token 数量庞大,往往会带来高昂的推理成本。近年来,视觉 token 剪枝方法越来越多地遵循两阶段范式:首先在视觉编码器之后移除视觉上冗余的 token,然后在大型语言模型(LLM)内部丢弃与文本查询无关的 token。然而,由于第一阶段通常仅依赖视觉编码器显著性,它可能会过早地消除与查询相关的 token,从而使后续文本引导阶段失去关键的视觉证据。我们的实证分析表明,将查询引导纳入第一阶段剪枝能更好地保留任务相关证据,并且相较于仅基于视觉显著性的剪枝,能够持续提升性能。我们进一步发现,高方差注意力头对文本查询更敏感,并能产生更具判别性的文本到视觉注意力信号,用于第二阶段剪枝。受这些发现启发,我们提出 TReVS,这是一个无需训练的框架,它将文本相关性与视觉编码器显著性相结合用于 LLM 前的剪枝,并利用高方差注意力头在 LLM 的浅层到中间层移除任务无关 token。在 LLaVA-1.5-7B 上,TReVS 在剪枝 94.4% 的视觉 token 的同时,保留了未剪枝基线性能的 92.8%,优于先前的最先进方法。
cs.LG / 102 / 2609.37659
Are In-Context Images Worth 10 Dimensions?
上下文图像是否值得 10 个维度?
large language model
大语言模型相关
Abstract
There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit. For a few-shot classification task, the induction circuit leverages linear representations of each labeled example in-context in order to classify an unlabeled query. However, few works focus on how those linear representations are built in the first place. Leveraging the expressivity of the vision modality compared to text, we uncover a Shared Discriminative Geometry (SDG) inside Large Vision Language Models (LVLMs). It is a low-dimensional space, shared across all image classification tasks, in which in-context images are compressed into linearly separable representations later used to perform classification. We observe that this is the result of the model performing a dimensionality reduction of vision representations in early layers. In order to explain this phenomenon: (1) We show analytically that linear self-attention can perform a dimensionality reduction by projecting in-context data onto its principal components, with each layer implementing one gradient descent step toward this objective. (2) We provide evidence that trained LVLMs reduce the dimensionality of vision representations in early layers via a similar mechanism.
Chinese Translation
已有大量重要工作致力于理解大型语言模型的上下文学习能力,尤其是关于归纳电路。对于少样本分类任务,归纳电路利用上下文中每个带标签示例的线性表示,以便对未标记的查询进行分类。然而,很少有工作关注这些线性表示最初是如何构建的。利用视觉模态相较于文本所具有的表达能力,我们揭示了大型视觉语言模型(LVLMs)内部的一种共享判别几何(Shared Discriminative Geometry, SDG)。它是一个低维空间,在所有图像分类任务之间共享,在其中上下文图像被压缩为线性可分的表示,这些表示随后被用于执行分类。我们观察到,这是模型在早期层中对视觉表示进行降维的结果。为了解释这一现象:(1)我们从分析上表明,线性自注意力可以通过将上下文数据投影到其主成分上来执行降维,其中每一层都实现一步朝着该目标的梯度下降。(2)我们提供证据表明,经过训练的 LVLMs 通过类似机制在早期层中降低视觉表示的维度。
cs.LG / 103 / 2609.37889
ReCAP: Retrieval-Guided Capability Reuse for Multimodal Continual Instruction Tuning
ReCAP:面向多模态持续指令调优的检索引导式能力复用
large language model
大语言模型相关
Abstract
Multimodal continual instruction tuning (MCIT) aims to enable multimodal large language models to acquire new capabilities from sequential tasks while preserving previously learned knowledge. Existing methods primarily mitigate catastrophic forgetting by constraining parameter updates or separating task-specific adaptations. However, continual adaptation can also benefit from external knowledge that provides domain-specific information and reusable reasoning patterns for solving diverse instructions. For example, to answer "How many red cubes are to the left of the sphere?", domain knowledge can provide relevant concepts about objects and spatial relations, while reasoning knowledge can specify ordered operations such as object recognition, spatial filtering, and counting. Despite this potential, how to leverage external knowledge for continual adaptation remains largely unexplored in existing MCIT methods. To this end, we propose ReCAP, a retrieval-guided framework that leverages external knowledge to guide capability reuse during continual adaptation. At each continual stage, ReCAP uses external search and an LLM to incrementally build a knowledge base of domain, reasoning, and format knowledge based on the current-stage training data. For each instruction, retrieved domain knowledge guides generation, while retrieved reasoning knowledge selects and orders capability modules to form an instance-specific capability path. As these capability modules are reused across stages, subsequent adaptation can overwrite previously learned parameters. To enable stable cross-stage reuse, ReCAP introduces adaptive subspace recycling, which parameterizes reusable capability modules with shared bases and stage-specific cores, protects historically important directions while recycling residual capacity. Extensive experiments on MCIT benchmarks show that ReCAP achieves SOTA performance.
Chinese Translation
多模态持续指令调优(MCIT)旨在使多模态大语言模型能够从顺序任务中获取新能力,同时保留先前学到的知识。现有方法主要通过约束参数更新或分离任务特定适配来缓解灾难性遗忘。然而,持续适应也可以受益于外部知识,这些知识为求解多样化指令提供领域特定信息和可复用的推理模式。例如,为了回答“球体左边有多少个红色立方体?”,领域知识可以提供关于物体和空间关系的相关概念,而推理知识可以指定有序操作,例如物体识别、空间过滤和计数。尽管存在这种潜力,如何在现有 MCIT 方法中利用外部知识进行持续适应仍在很大程度上尚未被探索。为此,我们提出 ReCAP,一个检索引导的框架,它利用外部知识来引导持续适应过程中的能力复用。在每个持续阶段,ReCAP 使用外部搜索和一个 LLM,基于当前阶段的训练数据,增量构建一个包含领域知识、推理知识和格式知识的知识库。对于每条指令,检索到的领域知识引导生成,而检索到的推理知识选择并排序能力模块,以形成实例特定的能力路径。由于这些能力模块跨阶段被复用,后续适应可能会覆盖先前学到的参数。为了实现稳定的跨阶段复用,ReCAP 引入了自适应子空间回收,它用共享基和阶段特定核心来参数化可复用的能力模块,保护历史上重要的方向,同时回收剩余容量。在 MCIT 基准上的大量实验表明,ReCAP 取得了 SOTA 性能。
cs.LG / 104 / 2609.37918
SYNCR: Diagnosing and Learning Cross-Video Reasoning from Simulation
SYNCR:从仿真中诊断与学习跨视频推理
large language model
大语言模型相关
Abstract
Reasoning across videos requires aligning events, matching identities, comparing motion, and integrating partial observations. Evaluating these capabilities and testing how to improve them requires both reliable labels and targeted supervision. We introduce SYNCR, a simulator-grounded framework that connects these two needs through shared task generators. Built on Habitat, Kubric, and CLEVRER, SYNCR derives answers from environment state and provides 4,000 evaluation questions and 15,960 training questions over disjoint videos, spanning eight cross-video reasoning tasks. Visual ablations and human evaluation assess dependence on the supplied evidence and answer recoverability. Evaluation of 22 multimodal large language models reveals persistent difficulties in physical comparison and scene integration that increasing model size does not consistently resolve. Supervised fine-tuning raises Qwen3-VL-8B's average SYNCR accuracy from 32.6% to 61.6%, with gains extending to task configurations and video sources absent from training for those tasks. Transfer to real footage is most consistent for temporal ordering: accuracy improves by 9.0-20.5 percentage points on constructed Assembly101 and Panoptic ordering sets across three checkpoints spanning two model families and two model sizes, with additional gains on existing temporal reasoning benchmarks. These results establish SYNCR as a controlled setting for diagnosing cross-video reasoning failures, testing their learnability, and identifying where synthetic supervision transfers.
Chinese Translation
跨视频推理需要对事件进行对齐、匹配身份、比较运动并整合部分观测。评估这些能力并测试如何改进它们,既需要可靠的标签,也需要有针对性的监督。我们提出 SYNCR,一个以仿真器为基础的框架,它通过共享的任务生成器将这两种需求连接起来。SYNCR 构建于 Habitat、Kubric 和 CLEVRER 之上,从环境状态中推导答案,并在互不相交的视频上提供 4,000 个评估问题和 15,960 个训练问题,涵盖八项跨视频推理任务。视觉消融与人工评估考察了对所提供证据的依赖程度以及答案的可恢复性。对 22 个多模态大语言模型的评估揭示出在物理比较与场景整合方面持续存在的困难,而增大模型规模并不能稳定地解决这些困难。监督微调将 Qwen3-VL-8B 的平均 SYNCR 准确率从 32.6% 提升至 61.6%,且性能提升还扩展到这些任务在训练中未曾出现的任务配置与视频来源。向真实视频的迁移在时序排序方面最为一致:在构建的 Assembly101 和 Panoptic 排序数据集上,跨两个模型家族、两种模型规模的三个检查点,准确率提升了 9.0–20.5 个百分点,并在现有时序推理基准上取得额外增益。这些结果确立了 SYNCR 作为一个受控环境的地位,可用于诊断跨视频推理的失败、检验其可学习性,并识别合成监督在何处能够迁移。
cs.AI / 105 / 2609.37925
Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation
用于长时程自回归视频生成的展开边际蒸馏
diffusion
扩散模型相关
Abstract
Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Training the generator on its own rollouts exposes it to these imperfect histories. However, existing video-level distribution matching distillation (DMD) scores the whole rollout jointly. Because a chunk is evaluated together with its past and future, its correction can favor matching artifacts in the surrounding context merely to preserve temporal consistency. To provide a clearer visual-quality signal, we introduce Rollout-Marginal Distillation (RMD). RMD retains the generated history for AR prediction but scores each chunk independently against a chunk teacher, ensuring its quality correction is not compromised by an imperfect temporal context. To compensate for the lack of temporal context in independent chunk scoring, RMD subsequently applies video-level DMD to restore temporal coherence. Extensive experiments demonstrate that RMD maintains high visual quality far beyond its training horizon and outperforms video-level DMD baselines. Code and video results are available at https://cjeen.github.io/RMD
Chinese Translation
自回归(AR)视频扩散能够实现低延迟、可流式的视频生成,但预测误差往往会在长展开过程中累积。在生成器自身的展开结果上训练它,会使其接触到这些不完美的历史。然而,现有的视频级分布匹配蒸馏(DMD)会对整个展开结果进行联合评分。由于一个块是与其过去和未来一起被评估的,其校正可能会为了仅仅保持时间一致性而倾向于匹配周围上下文中的伪影。为了提供更清晰的视觉质量信号,我们提出了展开边际蒸馏(RMD)。RMD 保留生成的历史用于 AR 预测,但会以一个块教师为参照对每个块独立评分,从而确保其质量校正不会因不完美的时间上下文而受到损害。为了补偿独立块评分中时间上下文的缺失,RMD 随后应用视频级 DMD 来恢复时间连贯性。大量实验表明,RMD 能在远超其训练时程的情况下保持高视觉质量,并优于视频级 DMD 基线。代码和视频结果可在 https://cjeen.github.io/RMD 获取。
cs.AI / 106 / 2609.38010
From Unity Simulation to Diffusion-Based Augmentation: Quantifying Dataset Balance for Robust Object Detection
从 Unity 仿真到基于扩散的增强:量化数据集平衡以实现鲁棒目标检测
diffusion
扩散模型相关
Abstract
Modern computer vision models achieve high accuracy when trained on large-scale annotated datasets. In critical domains such as construction safety monitoring, data collection is costly, hazardous, and ethically constrained. This paper presents a systematic study comparing two complementary data generation paradigms, (1) Unity Simulation-based rendering and (2) Controllable Diffusion-based generation (CIA), for object detection under real data-scarce conditions. A unified experimental framework enables controlled dataset mixing across real, simulated, and generative sources, while maintaining identical model and training settings. Quantitative evaluation using Precision, Recall, mAP, and custom $Δ$-metrics, reveals that neither simulation nor generative augmentation alone achieves optimal transferability. Unity-only training yields an mAP@0.5 drop of $-50\%$ relative to real data, while CIA-only training shows a milder $-16.5\%$ degradation. Hybrid compositions significantly improve performance, with the 90\% real + 10\% Unity configuration achieving the best overall mAP@0.5 of $62.68\%$ ($+7.64\%$ over baseline), and the 90\% real + 10\% CIA configuration maximizing precision at $74.45\%$. Results demonstrate that limited synthetic inclusion enhances generalization, while excessive substitution induces domain drift.
Chinese Translation
现代计算机视觉模型在大规模标注数据集上训练时能够达到高精度。在诸如建筑安全监控等关键领域中,数据收集成本高昂、具有危险性,并受到伦理约束。本文提出了一项系统性研究,比较两种互补的数据生成范式——(1)基于 Unity 仿真的渲染和(2)可控的基于扩散的生成(CIA)——用于真实数据稀缺条件下的目标检测。一个统一的实验框架能够在真实、仿真和生成来源之间进行受控的数据集混合,同时保持相同的模型和训练设置。使用精确率、召回率、mAP 以及自定义的 $Δ$-指标进行的定量评估表明,单独的仿真或生成式增强均无法实现最优的可迁移性。仅使用 Unity 训练相对于真实数据会造成 $-50\%$ 的 mAP@0.5 下降,而仅使用 CIA 训练则显示出较温和的 $-16.5\%$ 退化。混合组合显著提升性能,其中 90\% 真实数据 + 10\% Unity 配置实现了最佳总体 mAP@0.5,为 $62.68\%$(较基线 $+7.64\%$),而 90\% 真实数据 + 10\% CIA 配置将精确率最大化至 $74.45\%$。结果表明,有限地纳入合成数据能够增强泛化能力,而过度的替代则会引起域漂移。
cs.AI / 107 / 2609.38140
Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE
打破均匀性陷阱:通过 SplitMoE 扩展视频扩散模型
diffusionlarge language model
扩散模型相关
大语言模型相关
Abstract
Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patches across disparate experts, causing routing fragmentation and structural distortion. To address this, we propose SplitMoE, a split-role sparse architecture that breaks the shackles of uniformity. To accommodate the inherent semantic imbalance, we explicitly bifurcate the expert pool into semantic experts and generic experts, with semantic experts capturing high-level semantic abstraction and generic experts preserving residual visual information and flexible generative capacity. Leveraging prototype-guided routing and pull-push regularization, SplitMoE enables tokens to cluster naturally by semantic attributes rather than arbitrary balancing constraints. Extensive results show that under an equivalent activated-parameter budget, SplitMoE outperforms traditional load-balanced MoEs in convergence speed, routing coherence, and video generation quality across standard benchmarks. By revealing an emergent coarse-to-fine denoising logic, SplitMoE provides the community with a modality-aware scaling path, serving as a critical reference for building large-scale video world models.
Chinese Translation
由大型语言模型普及的混合专家(MoE)是一种用于扩展视觉生成模型的有前景的范式。然而,传统的逐 token MoE 在同质专家池内独立地路由 token,并将专家使用正则化趋向均匀,使其与具有时空冗余和语义长尾的视频数据匹配不佳。我们表明,现有的视觉 MoE 落入均匀性陷阱:语义上组织不足的路由,加之均匀的专家使用正则化,将连贯的图像块分散到不同专家中,导致路由碎片化和结构失真。为了解决这一问题,我们提出 SplitMoE,一种打破均匀性束缚的分角色稀疏架构。为了适应固有的语义不平衡,我们显式地将专家池分为语义专家和通用专家,其中语义专家捕捉高层语义抽象,通用专家保留残差视觉信息和灵活的生成能力。借助原型引导的路由和推拉正则化,SplitMoE 使 token 能够按语义属性自然聚类,而不是受任意的平衡约束。大量结果表明,在等价的激活参数预算下,SplitMoE 在收敛速度、路由一致性和视频生成质量方面,在标准基准上均优于传统的负载均衡 MoE。通过揭示一种涌现的由粗到细的去噪逻辑,SplitMoE 为社区提供了一条模态感知的扩展路径,为构建大规模视频世界模型提供了关键参考。
cs.CL / 108 / 2609.38177
Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
Imagine3D-LLM:教会 MLLMs 在回答前想象 3D 场景
large language model
大语言模型相关
Abstract
Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of the scene. Inspired by this process, we introduce Imagine3D-LLM, an MLLM that learns to assemble a similar compact 3D representation of the scene and conditions its answer on this representation. Concretely, we append a small set of learnable summary tokens after the image tokens, decode them into a compact 3D Gaussian Splatting representation supervised by a photometric reconstruction loss, and train jointly with the standard next-token prediction objective. Notably, although only the summary tokens receive direct reconstruction supervision, this objective also induces stronger cross-frame correspondence within the LLM's underlying image features, suggesting that learning to reconstruct propagates 3D-aware signals throughout the model. As a result, Imagine3D-LLM consistently outperforms prior approaches across multiple spatial reasoning and 3D understanding benchmarks, suggesting that imagining the scene can be more effective than being told its pixel-wise geometry.
Chinese Translation
从多视图图像中对 3D 世界进行推理,仍然是多模态大语言模型(MLLMs)面临的一项根本性挑战。尽管现代 MLLMs 能够有效处理单图像输入,但它们难以将跨视角的证据整合为连贯的 3D 理解。越来越多的工作试图通过向 MLLMs 注入 3D 感知来弥合这一差距,要么通过增强细粒度像素级跨视图对应,要么通过融合来自 3D 几何基础模型的特征,然而与人类推理之间仍存在显著差距。在这项工作中,我们重新审视人类的空间推理,它表明:人类并非依赖细粒度几何线索,而是粗略地识别不同视图中的共同物体,推断视点之间的相对几何关系,并组装出场景的粗略 3D 布局。受这一过程启发,我们提出了 Imagine3D-LLM,一种学习组装场景的类似紧凑 3D 表示,并以该表示为条件来生成答案的 MLLM。具体而言,我们在图像 token 之后附加一小组可学习的摘要 token,将其解码为受光度重建损失监督的紧凑 3D Gaussian Splatting 表示,并与标准的下一 token 预测目标联合训练。值得注意的是,尽管只有摘要 token 接受直接的重建监督,这一目标还在 LLM 底层图像特征中诱导出更强的跨帧对应,这表明学习重建会将 3D 感知信号传播到整个模型。因此,Imagine3D-LLM 在多个空间推理和 3D 理解基准上始终优于先前方法,这表明想象场景可能比被告知场景的逐像素几何更有效。
cs.AI / 109 / 2609.36954
Purlin: Separating Orchestration from the Datapath of Collectives
Purlin:将编排与集合通信的数据通路分离
diffusion
扩散模型相关
Abstract
Distributed inference depends on GPU collective communication that must keep pace with evolving hardware and specialized workloads. However, existing collective implementations often couple semantics, orchestration (where and when data moves), and the datapath (how data moves). This coupling makes it costly to adopt new hardware mechanisms and customize communication for applications. We present Purlin, a scale-up communication framework that separates these concerns. At the top of Purlin, we specify collectives as a naming of an input and output layout and a copy or reduction operation. In the middle, we introduce a shared orchestration protocol, Stage, Notify, And Consume (SNAC), which derives coordination from these specifications. Below SNAC sits a hardware-specific datapath we call Atom, which implements two key data movement primitives for collectives: copy and reduce. This separation lets us customize collectives and adopt new hardware mechanisms while reusing orchestration via SNAC. We evaluate Purlin on A100, H200, and B200 GPUs. Across seven collectives, Purlin achieves latency speedups of up to 5.14x and bandwidth improvements of up to 4.50x over baselines. Integrated into SGLang, Purlin improves offline LLM serving throughput and interactivity by 1.13x on average and up to 1.37x over baselines. For online LLM inference, Purlin improves interactivity by 1.26x on average and up to 2.85x, with the largest gain occurring under overload. For diffusion image generation, Purlin reduces end-to-end latency by up to 1.13x.
Chinese Translation
分布式推理依赖于 GPU 集合通信,而后者必须跟上不断演进的硬件和专用工作负载的步伐。然而,现有的集合通信实现通常将语义、编排(数据在何处以及何时移动)和数据通路(数据如何移动)耦合在一起。这种耦合使得采用新的硬件机制以及为应用定制通信的代价高昂。我们提出 Purlin,一个将这些关注点分离的 scale-up 通信框架。在 Purlin 的顶层,我们将集合通信指定为对输入和输出布局的命名,以及一个复制或归约操作。在中间层,我们引入一个共享的编排协议,即阶段、通知与消费(Stage, Notify, And Consume,SNAC),它从这些规范中推导出协调。位于 SNAC 之下的是一个我们称为 Atom 的硬件特定数据通路,它为集合通信实现了两个关键的数据移动原语:复制和归约。这种分离使我们能够定制集合通信并采用新的硬件机制,同时通过 SNAC 复用编排。我们在 A100、H200 和 B200 GPU 上评估 Purlin。在七种集合通信上,Purlin 相较于基线实现了最高 5.14 倍的延迟加速和最高 4.50 倍的带宽提升。集成到 SGLang 后,Purlin 将离线 LLM 服务的吞吐量和交互性相较基线平均提升 1.13 倍,最高提升 1.37 倍。对于在线 LLM 推理,Purlin 将交互性平均提升 1.26 倍,最高提升 2.85 倍,其中最大增益出现在过载情况下。对于扩散图像生成,Purlin 将端到端延迟最多降低 1.13 倍。
cs.AI / 110 / 2609.37532
DScale: Scaling Block-Diffusion Speculative Decoding with Adaptive Verification
DScale:以自适应验证扩展块扩散投机解码
diffusionlarge language model
扩散模型相关
大语言模型相关
Abstract
Growing large language model applications demand efficient inference. At high concurrency, block-diffusion speculative decoding suffers from verification padding, rejected candidates, and incompatibility between variable prefixes and fixed-shape graphs. Uniform truncation sacrifices acceptable tokens. We present DScale, preserving drafter architecture, weights, and full draft length. A separate 112K-parameter predictor requires neither confidence calibration nor hardware speed-curve preparation. Path-aware tiles reduce padding. Dynamic verify-length (DVL) allocation packs scored prefixes into half the native verification capacity. Fixed-address workspaces propagate changing boundaries through verification and acceptance while reusing captured graphs. On A100-40GB with tensor parallelism 1, Qwen3-8B and Qwen3-4B cover four datasets and concurrency 8-32, reusing each target's frozen predictor. Geometric-mean throughput gains across these configurations are respectively 43.9% and 48.8% over DFlash, 22.2% and 37.7% over DSpark, and 24.4% and 32.0% over Domino, with lower request latency. Cumulative ablations show that adding the three mechanisms successively increases geometric-mean throughput, while budget adjustment improves accepted-token retention. GPU profiling shows that complete decode-step time on GSM8K decreases by 30.8-52.5% relative to DFlash
Chinese Translation
不断增长的大语言模型应用需要高效推理。在高并发下,块扩散投机解码受到验证填充、被拒绝的候选以及可变前缀与固定形状图之间不兼容性的困扰。统一截断牺牲了可接受的 token。我们提出 DScale,保留了草稿器架构、权重和完整草稿长度。一个单独的 112K 参数预测器既不需要置信度校准,也不需要硬件速度曲线准备。路径感知瓦片减少填充。动态验证长度(DVL)分配将已评分前缀打包进原生验证容量的一半。固定地址工作区在验证和接受过程中传播变化的边界,同时复用已捕获图。在张量并行度为 1 的 A100-40GB 上,Qwen3-8B 和 Qwen3-4B 覆盖四个数据集和并发度 8-32,复用每个目标模型的冻结预测器。在这些配置上的几何平均吞吐量增益相对于 DFlash 分别为 43.9% 和 48.8%,相对于 DSpark 分别为 22.2% 和 37.7%,相对于 Domino 分别为 24.4% 和 32.0%,同时请求延迟更低。累积消融实验表明,依次加入这三种机制会逐次提高几何平均吞吐量,而预算调整提高了已接受 token 的保留率。GPU 性能分析显示,在 GSM8K 上完整解码步时间相对于 DFlash 减少了 30.8-52.5%。
cs.AI / 111 / 2609.37626
SPLASH: Switching Parallel Layouts of Attention with Seamless Handoff for LLM Serving
SPLASH:以无缝交接切换注意力的并行布局以服务 LLM
large language model
大语言模型相关
Abstract
No single way of parallelizing attention serves large language models well under all loads. Low concurrency favors tensor parallelism, many independent requests favor data-parallel attention, and long prompts favor context parallelism. Reasoning, agentic, and RL-rollout workloads make a fixed choice untenable: a batch that begins as many short requests ends as a few very long ones, so the best layout changes while the same requests run. Serving engines nevertheless fix one layout at launch, because changing it has meant draining requests and restarting workers. We present SPLASH, a serving system that switches the parallel layout of attention while requests are running. It builds on one observation: modern attention, with few or no KV heads, decouples where a request's KV cache lives from how attention weights are sharded. This has two consequences. First, layouts differ only in who owns the weights and the cache, and most of that state already sits where the next layout needs it; SPLASH reuses it, moves the rest in the background of ongoing inference, and hands off at a batch boundary, making a switch nearly free: its median overhead is under 0.51% of the step it runs in. Second, the decoupling exposes a layout that existing engines lack: Decoupled Ownership Parallelism (DOP) shards attention weights as tensor parallelism does while keeping each request's cache on a single owner as data-parallel attention does. DOP replicates neither, offers 27-60% more KV capacity than data-parallel attention, and gives the scheduler a choice when KV memory limits admission. A transition-aware scheduler follows the best of the four layouts as load changes. On B200 GPUs serving GLM-5.3, SPLASH improves end-to-end serving throughput by 1.3-1.73x over fixed-layout deployments, and the same layout regimes appear with DeepSeek-V3.2 on H200 and GLM-5.3-Flash on DCU.
Chinese Translation
没有任何一种并行化注意力的方式能在所有负载下都很好地服务大语言模型。低并发倾向于张量并行,大量相互独立的请求倾向于数据并行注意力,而长提示则倾向于上下文并行。推理、智能体和 RL-rollout 工作负载使得固定选择难以为继:一个开始时由许多短请求组成的批次,最终会变成少数几个非常长的请求,因此在同一批请求运行期间,最优布局会发生变化。然而,服务引擎在启动时仍固定采用一种布局,因为改变布局一直意味着要排空请求并重启工作进程。我们提出 SPLASH,一个在请求运行期间切换注意力并行布局的服务系统。它基于一个观察:具有很少或没有 KV 头的现代注意力,将请求的 KV 缓存所在位置与注意力权重的分片方式解耦。这带来两个结果。第一,各布局之间的差异仅在于谁拥有权重和缓存,而该状态的大部分已经位于下一个布局所需的所在之处;SPLASH 复用它,在持续推理的后台迁移其余部分,并在批次边界处完成交接,从而使切换几乎无成本:其中位开销低于其所处步骤的 0.51%。第二,这种解耦揭示了一种现有引擎所缺少的布局:解耦所有权并行(Decoupled Ownership Parallelism,DOP)像张量并行那样对注意力权重进行分片,同时像数据并行注意力那样将每个请求的缓存保持在单一所有者上。DOP 对二者都不做复制,相比数据并行注意力提供高出 27-60% 的 KV 容量,并在 KV 内存限制准入时为调度器提供了一种选择。一个感知转换的调度器会随负载变化跟随四种布局中最优的那一种。在服务 GLM-5.3 的 B200 GPU 上,SPLASH 相比固定布局部署将端到端服务吞吐量提升 1.3-1.73 倍,并且在 H200 上的 DeepSeek-V3.2 以及 DCU 上的 GLM-5.3-Flash 上也出现了相同的布局状态区间。
cs.AI / 112 / 2609.36129
Accessible, but Not Adopted: Increasing LLM Adoption among First-generation, Low-income (FGLI) College Students beyond Expanding Access
可及却未被采用:超越扩大可及性,提升第一代低收入(FGLI)大学生对LLM的采用
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly positioned as a force to empower underserved communities, and significant efforts are being made to expand access. Yet, access alone does not equate to meaningful adoption. First, even if a system is accessible, it won't be adopted if users are not willing to adopt it. Second, even if an LLM system is superficially adopted, the heterogeneity of LLM tools means that LLM adoption can be further deepened. Closing this access-adoption gap is critical to ensuring that the full social potential of LLM is not only accessible but fully realised. Drawing on 61 interviews (15 long-form semi-structured interviews with first-generation, low-income college (FGLI) students, 3 non-FGLI students, 3 FGLI program directors, and 40 intercept interviews), this paper examines the access-adoption gap in first-generation, low-income student communities. This paper a) finds that while FGLI students have adopted LLM systems, their depth of LLM tool usage is limited to chatbots (e.g., ChatGPT or Claude) for narrow use cases, and b) identifies barriers limiting their willingness to learn and use (low perceived value, under-estimated self-efficacy, unclear starting point, low peer exposure, and resource constraints). Then, from these findings, the paper derives the four design principles to design a system or an intervention aimed at closing the access-adoption gap in LLM adoption by FGLI students. In doing so, the paper contributes to the field by a) examining the LLM access-adoption gap in the FGLI student community, and b) reframing LLM adoption as a depth gradient across four modes of LLM tool use: basic chatbot interfaces, tool-augmented prebuilt interfaces, agentic development interfaces, and programmatic integration.
Chinese Translation
大型语言模型(LLM)正日益被定位为赋能弱势群体的力量,并且人们正在做出重大努力以扩大其可及性。然而,仅有可及并不等同于有意义的采用。第一,即使一个系统是可及的,如果用户不愿意采用它,它也不会被采用。第二,即使一个LLM系统在表面上被采用,LLM工具的异质性也意味着LLM的采用可以被进一步深化。弥合这一可及—采用差距,对于确保LLM的完整社会潜力不仅可及而且被充分实现至关重要。基于61次访谈(15次对第一代低收入大学生(FGLI)的长时间半结构化访谈、3名非FGLI学生、3名FGLI项目主管,以及40次拦截访谈),本文考察了第一代低收入学生群体中的可及—采用差距。本文a)发现,尽管FGLI学生已经采用了LLM系统,但他们对LLM工具的使用深度仅限于聊天机器人(例如ChatGPT或Claude)用于狭窄的使用场景,并且b)识别出限制他们学习和使用意愿的障碍(低感知价值、被低估的自我效能感、不清晰的起点、低同伴接触度,以及资源约束)。然后,基于这些发现,本文推导出四项设计原则,用于设计旨在弥合FGLI学生在LLM采用方面的可及—采用差距的系统或干预措施。借此,本文通过以下方式为该领域做出贡献:a)考察FGLI学生群体中的LLM可及—采用差距,以及b)将LLM采用重新界定为跨越四种LLM工具使用模式的深度梯度:基础聊天机器人界面、工具增强的预构建界面、智能体式开发界面,以及程序化集成。
cs.CL / 113 / 2609.37574
MERGE: Multi-LLM Ensemble for Retrieval via Generative Enrichment
MERGE:通过生成式增强进行检索的多LLM集成
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) are increasingly used to enrich user queries in information retrieval (IR) so that a standard retriever such as BM25 can bridge vocabulary gaps with the target corpus. Any single LLM, however, is limited by its training data and architectural biases, and its enrichment behavior depends on hand-crafted prompts that must be re-engineered for each new model -- an expensive and poorly scalable process. We present MERGE (Multi-LLM Ensemble for Retrieval via Generative Enrichment), a two-stage framework: three heterogeneous 7-8B open-source LLMs independently produce candidate expansions, and a larger LLM generatively synthesizes them into a single query. To make prompt engineering scalable across the ensemble, we integrate a task-grounded Automatic Prompt Optimization (APO) loop into both stages. Unlike APO methods that judge candidates with an LLM evaluator, our loop scores each candidate by its downstream retrieval performance and runs a small tournament between the current champion prompt and optimizer-proposed drafts, terminating once the champion survives two consecutive rounds; a history-augmented variant additionally feeds the recent tournament trajectory back to the optimizer. MERGE is retriever-agnostic and issues a single BM25 pass with no rank fusion, no supervised document expansion, and no re-indexing. On five BEIR benchmarks (NQ, SciFact, FiQA, Touche-2020, DBPedia), MERGE improves BM25 nDCG@10 over the original queries by +2.1 to +14.9 points and matches or outperforms strong LLM-based query-expansion baselines despite using only compact open-source models. Ablations confirm that the Stage-2 ensemble beats any single Stage-1 LLM, and that task-grounded APO converts large seed-prompt regressions into consistent gains without hand-tuning.
Chinese Translation
大型语言模型(LLM)越来越多地被用于在信息检索(IR)中丰富用户查询,使得诸如BM25这样的标准检索器能够弥合与目标语料库之间的词汇鸿沟。然而,任何单一LLM都受限于其训练数据和架构偏见,并且其增强行为依赖于必须针对每个新模型重新设计的手工编写提示——这是一个昂贵且难以扩展的过程。我们提出MERGE(Multi-LLM Ensemble for Retrieval via Generative Enrichment,通过生成式增强进行检索的多LLM集成),这是一个两阶段框架:三个异构的7-8B开源LLM独立生成候选扩展,然后由一个更大的LLM以生成方式将它们合成为单一查询。为了使提示工程在整个集成中具有可扩展性,我们在两个阶段中都集成了一个基于任务的自动提示优化(APO)循环。与使用LLM评估器来评判候选的APO方法不同,我们的循环通过每个候选的下游检索性能来为其打分,并在当前冠军提示与优化器提出的草稿之间运行一个小型锦标赛,一旦冠军连续两轮胜出便终止;一个增强历史的变体还会将近期的锦标赛轨迹反馈给优化器。MERGE与检索器无关,并且只执行一次BM25检索过程,不进行排名融合、不进行有监督的文档扩展,也不重建索引。在五个BEIR基准(NQ、SciFact、FiQA、Touche-2020、DBPedia)上,MERGE将BM25的nDCG@10相较于原始查询提升了+2.1至+14.9个点,并且尽管仅使用紧凑的开源模型,仍能匹敌或超越强大的基于LLM的查询扩展基线。消融实验证实,第2阶段的集成优于任何单个第1阶段LLM,并且基于任务的APO能够在不进行手工调优的情况下,将种子提示带来的大幅性能下降转化为持续稳定的收益。
cs.CL / 114 / 2609.38099
Effective Dense Retrieval using Only In-Context Examples
仅使用上下文示例的有效稠密检索
large language model
大语言模型相关
Abstract
Turning decoder-only large language models (LLMs) into strong dense retrievers typically requires some form of retriever training. In this paper, we ask whether LLMs can instead be prompted to produce effective representations for dense retrieval given only a few in-context examples. To answer this, we introduce RICE (Representations from In-Context Examples), a simple "training-free" approach that extracts high-quality dense representations from LLMs. To do so, RICE conditions the LLM on examples that provide a shared context for query and document encoding. Our results demonstrate that RICE embeddings can substantially improve the accuracy of prompt-based LLM embeddings, establishing it as a simple method to build LLM-based dense retrievers that do not require training. We release our code at https://github.com/nourj98/RICE.
Chinese Translation
将仅解码器的大型语言模型(LLMs)转变为强大的稠密检索器通常需要某种形式的检索器训练。在本文中,我们探究是否仅给定少量上下文示例,就可以改为通过提示使 LLMs 为稠密检索生成有效的表示。为回答这一问题,我们提出 RICE(来自上下文示例的表示,Representations from In-Context Examples),这是一种简单的“免训练”方法,可从 LLMs 中提取高质量的稠密表示。为此,RICE 让 LLM 以那些为查询和文档编码提供共享上下文的示例为条件。我们的结果表明,RICE 嵌入能够大幅提升基于提示的 LLM 嵌入的准确率,从而将其确立为一种构建无需训练的基于 LLM 的稠密检索器的简单方法。我们在 https://github.com/nourj98/RICE 发布了我们的代码。
cs.AI / 115 / 2609.37056
Evolving Towards Better Codes: LLM-Guided Search for High-Distance Binary Linear Codes
向着更好的码演化:LLM 引导的高距离二进制线性码搜索
large language model
大语言模型相关
Abstract
Evolutionary program search driven by large language models (LLMs) has produced record-breaking constructions for open problems in combinatorics and beyond. We apply this approach to the longstanding problem of improving the best-known bounds for binary linear codes. Building on the EvoTune evolutionary framework and the ShinkaEvolve codebase, we introduce LinCodeEvolve, which evolves code-construction programs against an exact minimum-distance evaluator. A strategy loop combines diversity-driven search and expert supervision: when progress plateaus, new strategies are used to redirect the search. LinCodeEvolve discovers seven record-breaking codes, $[172,21,66]$, $[173,20,68]$, $[176,21,68]$, $[181,21,70]$, $[184,21,72]$, $[189,22,72]$ and $[200,21,77]$, six of which have concise quasi-cyclic descriptions. With standard code modification techniques, they improve $22$ entries of the tables. Every code is verified by exhaustive enumeration. These results suggest that LLM-guided search can help find improved codes and complement existing methods in coding theory.
Chinese Translation
由大语言模型(LLM)驱动的演化程序搜索,已在组合数学及其他领域的开放问题上产生了打破纪录的构造。我们将这一方法应用于一个长期存在的问题:改进二进制线性码的已知最优界。在 EvoTune 演化框架和 ShinkaEvolve 代码库的基础上,我们提出了 LinCodeEvolve,它针对精确的最小距离评估器来演化码构造程序。一个策略循环将多样性驱动的搜索与专家监督相结合:当进展停滞时,使用新策略来重新引导搜索。LinCodeEvolve 发现了七个打破纪录的码,$[172,21,66]$、$[173,20,68]$、$[176,21,68]$、$[181,21,70]$、$[184,21,72]$、$[189,22,72]$ 和 $[200,21,77]$,其中六个具有简洁的准循环描述。借助标准的码修改技术,它们改进了表中的 $22$ 个条目。每个码都通过穷举枚举进行了验证。这些结果表明,LLM 引导的搜索可以帮助找到更优的码,并补充编码理论中的现有方法。
cs.LG / 116 / 2609.36099
Data Unlearning via Inverse Distillation
通过逆蒸馏实现数据遗忘
diffusion
扩散模型相关
Abstract
Multi-step matching models, including flow and diffusion models, produce high-quality outputs but incur substantial inference costs and may reproduce unwanted components of their training datasets. We introduce Inverse Distillation Unlearning (IDU), a unified framework that simultaneously distills a teacher multi-step matching model into an efficient one-step student generator and suppresses outputs corresponding to a designated training subset. We first formulate distillation as a min-max objective over a data distribution and then represent this distribution as a mixture of the forget-set and the generated distributions. This allows us to compare this mixture with the teacher's training distribution and recover only the retained data at the optimum. Our method requires only a pretrained full-data teacher and data from the forget set, without access to retained training examples, extra feature extractors or classifiers. Extensive experiments on MNIST and CIFAR-10 datasets under flow-matching and score-based diffusion settings demonstrate that IDU substantially reduces the generation frequency of forgotten classes while preserving generation quality on the retained classes. To the best of our knowledge, IDU is the first unified framework for simultaneous unlearning and distillation in unconditional flow-matching and score-based models.
Chinese Translation
多步匹配模型,包括流模型和扩散模型,能够生成高质量输出,但会带来高昂的推理成本,并且可能重现其训练数据集中不需要的组成部分。我们引入了逆蒸馏遗忘(Inverse Distillation Unlearning, IDU),这是一个统一框架,能够同时将教师多步匹配模型蒸馏为一个高效的单步学生生成器,并抑制与指定训练子集相对应的输出。我们首先将蒸馏形式化为关于数据分布的一个极小极大目标,然后将该分布表示为遗忘集与生成分布的混合。这使我们能够将该混合分布与教师的训练分布进行比较,并在最优解处仅恢复保留数据。我们的方法仅需要一个预训练的全数据教师模型以及来自遗忘集的数据,无需访问保留的训练样本、额外的特征提取器或分类器。在流匹配和基于分数的扩散设置下,在 MNIST 和 CIFAR-10 数据集上进行的大量实验表明,IDU 显著降低了被遗忘类别的生成频率,同时保持了保留类别上的生成质量。据我们所知,IDU 是首个用于在无条件流匹配模型和基于分数的模型中同时进行遗忘与蒸馏的统一框架。
cs.LG / 117 / 2609.36115
Koa-action: Fast and Consistent Structured Decision Making with Generative LLMs
Koa-action:基于生成式LLM的快速且一致的结构化决策
large language model
大语言模型相关
Abstract
Industry applications often demand low-latency classification, yet current large language model (LLM) approaches remain poorly suited for latency-critical applications. Existing prompting and constrained decoding produce verbose, multi-token outputs that require expensive token-by-token generation, while encoder-based models achieve faster inference but sacrifice task flexibility. We propose Koa-action, a framework for low-latency atomic actions -- fast, single-step decisions such as classification, semantic endpointing, Boolean checks, and scoring -- formulated as constrained generation with single-token outputs. By introducing atomic label tokens and applying supervised fine-tuning, our method reduces classification to a deterministic one-step decoding problem. Across standard benchmarks, Koa-action delivers competitive accuracy with consistently low and stable latency. On a production intent-routing benchmark, Koa-action reaches 85.5% accuracy -- competitive with the strongest frontier models (Claude-4.8-Opus, Gemini-Pro-3.1) and ahead of GPT-5 and Gemini-2.5-Pro -- while answering in about half a second, several-fold faster than every frontier model (up to ~7.5x at the median) under identical serving conditions. Against the dedicated single-token system Jev/TypeSafe, Koa-action is competitive on accuracy and faster at the median, while also handling multimodal inputs and multi-label outputs that single-label text systems do not.
Chinese Translation
工业应用常常需要低延迟的分类,然而当前的大语言模型(LLM)方法仍然难以适用于对延迟敏感的应用。现有的提示与约束解码会产生冗长、多 token 的输出,需要昂贵的逐 token 生成,而基于编码器的模型虽然实现了更快的推理,却牺牲了任务灵活性。我们提出 Koa-action,一个面向低延迟原子动作的框架——这些动作是快速、单步的决策,例如分类、语义端点检测、布尔判断和打分——并被形式化为具有单 token 输出的约束生成。通过引入原子标签 token 并应用有监督微调,我们的方法将分类简化为一个确定性的单步解码问题。在标准基准测试上,Koa-action 在保持始终低且稳定的延迟的同时,提供了具有竞争力的准确率。在一个生产环境的意图路由基准上,Koa-action 达到 85.5% 的准确率——与最强的前沿模型(Claude-4.8-Opus、Gemini-Pro-3.1)相当,并领先于 GPT-5 和 Gemini-2.5-Pro——同时在相同服务条件下约半秒内给出回答,比每一个前沿模型都快数倍(中位数最高约 7.5 倍)。与专用的单 token 系统 Jev/TypeSafe 相比,Koa-action 在准确率上具有竞争力,且在中位数延迟上更快,同时还能够处理单标签文本系统所无法处理的多模态输入和多标签输出。
cs.LG / 118 / 2609.36116
Dyad: Extending Large Language Models with Native Typed Decision-Making
Dyad:以原生类型化决策扩展大语言模型
large language model
大语言模型相关
Abstract
We study how to build more capable general-purpose agents by extending large language models (LLMs) with native typed decision-making. We introduce Dyad, an architecture that augments a pretrained LLM with an environment-conditioned action encoder that embeds each candidate action description in parallel, then scores these embeddings against the LLM's internal state to yield a distribution over typed actions. By factorizing decision-making into representations of the evolving interaction state and environment-specific action semantics, Dyad introduces an inductive bias for learning reusable representations while keeping action scoring efficient even as the action space grows. We investigate two complementary reinforcement learning settings driven by environment interaction. With the LLM frozen, training the action encoder alone achieves consistent gains across four unseen environments, enabling modular adaptation without modifying any LLM parameters. Jointly optimizing both components outperforms conventional RL post-training across diverse interactive tasks and model scales, including a 3.80% average absolute gain on ALFWorld with a 9B model, while improving general knowledge, reasoning, and coding.
Chinese Translation
我们研究如何通过以原生类型化决策扩展大语言模型(LLM),构建能力更强的通用智能体。我们提出 Dyad,一种架构,它用一个以环境为条件的动作编码器增强预训练 LLM,该编码器并行地嵌入每个候选动作描述,然后将这些嵌入与 LLM 的内部状态进行打分,从而得到类型化动作上的分布。通过将决策分解为不断演化的交互状态的表示与环境特定的动作语义,Dyad 引入了一种归纳偏置,用于学习可复用的表示,同时即使动作空间增长也能保持动作打分高效。我们研究了两种由环境交互驱动的互补强化学习设置。在 LLM 冻结的情况下,仅训练动作编码器就在四个未见环境中取得了一致的收益,从而无需修改任何 LLM 参数即可实现模块化适配。联合优化两个组件在多样的交互任务和模型规模上均优于传统的 RL 后训练,包括在 9B 模型上于 ALFWorld 取得 3.80% 的平均绝对提升,同时改善了通用知识、推理和编程能力。
cs.LG / 119 / 2609.36117
Why Backdooring Neural Networks is so Easy?
为什么给神经网络植入后门如此容易?
large language model
大语言模型相关
Abstract
Securing modern AI systems against backdoor attacks remains an open challenge and requires fundamentally principled estimates of the adversary's budget -- the poison fraction $π$ and trigger strength $α$ needed to construct successful yet stealthy attacks. Motivated by recent empirical evidence that poisoning large language models can require a nearly constant number of malicious samples even as clean datasets grow, we derive an exact closed-form analysis of a quadratic neuron trained on a poisoned Gaussian mixture. We show, perhaps counterintuitively, that the same feature-learning dynamics that make neural networks powerful can also make them more vulnerable to backdoors. Specifically, with clean accuracy preserved to first order, $O(π)$, we demonstrate that lazy learning imposes the inverse-square-root scaling $α\propto π^{-1/2}$ for a successful attack, while feature learning induces a quadratic detector whose loss margin scales as $O(α^4)$, improving the attack budget to $α\propto π^{-1/4}$. Consequently, nonlinear feature learning substantially reduces the trigger strength required at small poison fractions, thereby in a sense making feature learners more backdoor vulnerable. These results provide a theoretical mechanism consistent with large-scale empirical observations and demonstrate that security audits based on linear heuristics can systematically underestimate backdoor vulnerability in the widely adopted feature-learning regimes.
Chinese Translation
保护现代 AI 系统免受后门攻击仍然是一个悬而未决的挑战,并且需要对对手的预算作出根本上合理的估计——即构建成功但隐蔽的攻击所需的投毒比例 $π$ 和触发强度 $α$。受近期经验证据的启发——即随着干净数据集增长,投毒大型语言模型可能仍需要几乎恒定数量的恶意样本——我们推导了对在受投毒的高斯混合上训练的二次神经元的精确闭式分析。我们表明,或许与直觉相反,正是那些使神经网络强大的特征学习动力学,也可能使它们更容易受到后门攻击。具体而言,在干净准确率保持到一阶 $O(π)$ 的情况下,我们证明懒惰学习对成功攻击施加了逆平方根缩放 $α\propto π^{-1/2}$,而特征学习诱导出一个二次检测器,其损失裕度按 $O(α^4)$ 缩放,从而将攻击预算改进为 $α\propto π^{-1/4}$。因此,非线性特征学习在小投毒比例下大幅降低了所需的触发强度,从而在某种意义上使特征学习器更易受后门攻击。这些结果提供了一种与大规模经验观察相一致的理论机制,并表明基于线性启发式的安全审计可能系统性地低估在广泛采用的特征学习机制中的后门脆弱性。
cs.LG / 120 / 2609.36120
ThinQuant: Scalable Rotation Learning for Weight and Activation Quantization of LLMs
ThinQuant:用于大型语言模型权重与激活量化的可扩展旋转学习
large language model
大语言模型相关
Abstract
Learned rotations play an important role in enabling low-bit weight and activation quantization of large language models by smoothing outliers in the activation distribution. State-of-the-art approaches include gradient-based procedures such as SpinQuant and computationally friendlier gradient-free approaches such as DartQuant, but both remain hard to scale to the largest architectures. To address the computational bottlenecks in gradient-free rotation learning, we introduce two ideas for efficiency, (i) a data selection procedure which reduces the required number of calibration data points, and (ii) an exact reduction of the associated optimization on this reduced calibration set. Our data selection procedure exploits the geometric structure of the convex hull of the activations. Using this idea, we show that a carefully selected calibration set with several orders of magnitude fewer activations than state-of-the-art rotation-based methods can match their performance in low-bit quantization settings. Under this extreme data efficiency, the selected activations span an $r$-dimensional subspace with $r<d$, making optimization over a $d\times d$ rotation equivalent to optimizing a $d\times r$ matrix on the Stiefel manifold. We solve this reduced problem using an efficient ADMM algorithm that iteratively employs thin matrix updates at every step, hence the name ThinQuant. For Llama-3-70B with W4A4KV4 quantization, ThinQuant completes the entire rotation calibration in under 12 minutes and achieves a WikiText-2 perplexity of 5.63, compared with 7.55 for DartQuant, which requires 111 minutes. Unlike SpinQuant and DartQuant, ThinQuant also scales to Llama-3.1-405B on a single H200 GPU, completing rotation calibration in just over 2 hours and achieving WikiText-2 perplexity of 2.97 at W4A4, compared with 3.48 for GPTAQ+QuaRoT.
Chinese Translation
学习到的旋转在通过平滑激活分布中的离群值来实现大型语言模型的低位权重与激活量化方面发挥着重要作用。最先进的方法包括基于梯度的过程(如 SpinQuant)以及计算上更友好的无梯度方法(如 DartQuant),但二者仍难以扩展到最大规模的架构。为了解决无梯度旋转学习中的计算瓶颈,我们引入了两个提高效率的想法:(i) 一个数据选择过程,它减少了所需的校准数据点数量;(ii) 在缩减后的校准集上对相关优化进行精确归约。我们的数据选择过程利用了激活凸包的几何结构。利用这一想法,我们表明,一个经过精心选择的校准集,其激活数量比最先进的基于旋转的方法少几个数量级,却可以在低位量化设置下匹配它们的性能。在这种极端的数据效率下,所选择的激活张成一个 $r$ 维子空间,其中 $r<d$,使得对 $d\times d$ 旋转的优化等价于在 Stiefel 流形上优化一个 $d\times r$ 矩阵。我们使用一种高效的 ADMM 算法解决这个缩减后的问题,该算法在每一步迭代地采用薄矩阵更新,因此得名 ThinQuant。对于采用 W4A4KV4 量化的 Llama-3-70B,ThinQuant 在不到 12 分钟内完成整个旋转校准,并达到 5.63 的 WikiText-2 困惑度,而 DartQuant 为 7.55,且需要 111 分钟。与 SpinQuant 和 DartQuant 不同,ThinQuant 还能在单个 H200 GPU 上扩展到 Llama-3.1-405B,在略超过 2 小时内完成旋转校准,并在 W4A4 下达到 2.97 的 WikiText-2 困惑度,相比之下 GPTAQ+QuaRoT 为 3.48。
cs.LG / 121 / 2609.36187
EvoMO-SR: Multiobjective LLM-based Evolution of Symbolic Expressions with substructure guidance
EvoMO-SR:基于多目标LLM的、带子结构引导的符号表达式演化
large language model
大语言模型相关
Abstract
Symbolic Regression (SR) is a data-driven method for scientific discovery which searches for interpretable analytical relationships within data. Recently, Large Language Models (LLMs) have also had a significant impact on scientific discovery, enabling the automation of various stages of the process. For these reasons, the possibility of harnessing the embedded scientific knowledge and programming capabilities of LLMs to solve SR tasks has emerged, showing promising performance compared with traditional methods. We propose EvoMO-SR, a novel LLM-driven SR framework in which the LLM generates equation skeletons, with their coefficients fitted separately by an external optimizer. The framework includes a multi-objective survival selection which controls bloating by balancing accuracy and complexity, and a substructure guidance mechanism which mutates expressions with candidate reusable building blocks. EvoMO-SR achieves the best accuracy in seven of the eight in-domain and out-of-domain settings for LSR-Synth, using a small LLM model, i.e., Llama-3.1-8B-Instruct. We also evaluated structural recovery through two symbolic accuracy metrics based on canonicalized subtree overlap and term matching, showing that our method has a greater probability of recovering highly accurate symbolic structures.
Chinese Translation
符号回归(SR)是一种用于科学发现的数据驱动方法,它在数据中搜索可解释的解析关系。近年来,大语言模型(LLM)也对科学发现产生了重大影响,使该过程的各个阶段得以自动化。出于这些原因,利用LLM所内嵌的科学知识与编程能力来求解SR任务的可能性随之出现,并且与传统方法相比展现出有前景的性能。我们提出EvoMO-SR,一种新颖的由LLM驱动的SR框架,其中LLM生成方程骨架,其系数由外部优化器单独拟合。该框架包含一种多目标生存选择,它通过平衡精度与复杂度来控制膨胀,以及一种子结构引导机制,它利用候选的可复用构建块对表达式进行变异。在LSR-Synth的八个域内与域外设定中,EvoMO-SR在七个上取得了最佳精度,且使用的是一个小型LLM模型,即Llama-3.1-8B-Instruct。我们还通过两个符号精度指标评估了结构恢复情况,这两个指标基于规范化子树重叠与项匹配,结果表明我们的方法具有更大的概率恢复出高精度的符号结构。
cs.LG / 122 / 2609.36210
On the spectral properties of generative denoiser Jacobians
论生成式去噪器雅可比矩阵的谱性质
diffusion
扩散模型相关
Abstract
Generative denoising models, such as diffusion and flow-matching, learn to sample from complex distributions by training a deep neural network denoiser to recover clean data from noise-corrupted samples. While such models are typically compared on the quality of their synthesized samples, these metrics provide limited insight into how the underlying denoiser, which drives generation, differs. In this work, we propose to analyze the spectrum of the denoiser Jacobian as a tool to characterize these differences. Across pre-trained denoising models, we observe that better generative performance is associated with larger Jacobian eigenvalues. Motivated by this, we introduce a regularization scheme that controls the Jacobian spectrum by training the denoiser on perturbed inputs, with perturbations suppressing or amplifying Jacobian responses. On ImageNet, we test whether directly modifying the Jacobian spectral properties leads to improved generations. Our findings suggest that denoisers benefit from both strengthening responses along data-relevant principal eigen-directions and suppressing the noisy, data-irrelevant ones. This establishes the denoiser Jacobian as a useful tool for identifying differences between generative denoising models.
Chinese Translation
生成式去噪模型,例如扩散和流匹配,通过训练一个深度神经网络去噪器从受噪声破坏的样本中恢复干净数据,从而学习从复杂分布中采样。尽管此类模型通常根据其合成样本的质量进行比较,但这些指标对于驱动生成的底层去噪器如何不同所提供的洞见有限。在这项工作中,我们提出分析去噪器雅可比矩阵的谱,作为刻画这些差异的工具。在预训练的去噪模型中,我们观察到更好的生成性能与更大的雅可比特征值相关联。受此启发,我们引入一种正则化方案,通过在扰动输入上训练去噪器来控制雅可比矩阵谱,其中扰动抑制或放大雅可比响应。在 ImageNet 上,我们测试直接修改雅可比谱性质是否会带来改进的生成。我们的发现表明,去噪器既受益于增强沿数据相关主特征方向的响应,也受益于抑制嘈杂的、与数据无关的响应。这确立了去噪器雅可比矩阵作为识别生成式去噪模型之间差异的有用工具。
cs.LG / 123 / 2609.36222
BASE: Batch-Aware Selection of Experts Using Predicted Removal Error for Efficient MoE Decoding
BASE:使用预测移除误差的批感知专家选择,用于高效 MoE 解码
large language model
大语言模型相关
Abstract
Large language models are increasingly expensive to serve. In large-scale serving systems, autoregressive decoding is often bottlenecked by transferring model weights from accelerator high-bandwidth memory into on-chip SRAM. Mixture-of-experts (MoE) models reduce computation by activating only a small subset of experts per token, but this sparsity does not translate directly to batched decoding. Different requests select different experts; therefore, the combined active set across many concurrent requests can span a substantial fraction of the expert pool and require significantly more expert weights to be transferred. Most expert-reduction techniques make retention decisions independently for each token and therefore do not address this batch-level expansion. More recently, batch-aware methods have attempted to coordinate expert use across concurrent requests and reuse experts already fetched for the batch. Yet their selection criteria are based primarily on router rankings or expert statistics collected during calibration. Consequently, these criteria are not directly tied to the output error caused by dropping an expert, nor do they capture how an expert's contribution changes across tokens at inference time. We instead rank experts according to how much their removal would change the MoE-layer output. To apply this criterion during serving, we train a lightweight linear predictor during calibration that estimates the expert removal cost for each incoming token, and develop custom GPU kernels for cost prediction and expert selection. Across three MoE architectures, BASE improves the quality-efficiency tradeoff without retraining. On Qwen3-30B-A3B, it improves average accuracy by 29.5 points over the strongest baseline at comparable throughput under a tight expert budget. At a higher expert budget, it is 60% faster than dense inference while remaining within 0.4 accuracy points.
Chinese Translation
大型语言模型的服务成本正变得越来越高。在大规模服务系统中,自回归解码常常受制于将模型权重从加速器的高带宽内存传输到片上 SRAM。混合专家(MoE)模型通过为每个 token 仅激活一小部分专家来减少计算,但这种稀疏性并不能直接转化为批处理解码。不同请求会选择不同的专家;因此,许多并发请求的合并活跃集可能覆盖专家池的相当大一部分,并需要传输显著更多的专家权重。大多数专家缩减技术为每个 token 独立做出保留决策,因此无法解决这种批级别的扩张。更近期,批感知方法尝试协调并发请求之间的专家使用,并复用已经为该批次获取的专家。然而,它们的选择标准主要基于路由器的排名或校准期间收集的专家统计信息。因此,这些标准并不直接与丢弃某个专家所导致的输出误差相关联,也无法刻画在推理时专家的贡献如何随 token 变化。我们则根据移除专家会在多大程度上改变 MoE 层输出,对专家进行排序。为了在服务期间应用这一标准,我们在校准期间训练一个轻量级线性预测器,用于估计每个传入 token 的专家移除成本,并开发用于成本预测和专家选择的自定义 GPU 内核。在三种 MoE 架构上,BASE 无需重新训练就改善了质量与效率的权衡。在 Qwen3-30B-A3B 上,在严格的专家预算下,与最强基线相比,在吞吐量相当的情况下,它将平均准确率提高了 29.5 个百分点。在更高的专家预算下,它比密集推理快 60%,同时准确率差距保持在 0.4 个百分点以内。
cs.LG / 124 / 2609.36265
In-Context Learning Amplifies a Latent Symbolic Circuit
上下文学习放大一个潜在的符号电路
large language model
大语言模型相关
Abstract
Large language models can learn abstract rules from just a few in-context examples, but how their internal mechanisms activate as examples accumulate is not well understood. We trace a three-stage symbolic reasoning circuit (abstraction, induction, retrieval) across shot counts in three model families and find it is detectable and functional well before the model achieves high accuracy. Per-head causal contribution grows up to 8x from 1- to 10-shot, and cross-shot activation patching raises accuracy from 1% to 56% at 0-shot and 17% to 88% at 1-shot. Function vectors scaled and injected at 0-shot rescue accuracy up to 86%, largely substituting for the induction stage but depending critically on an intact downstream retrieval stage. The infrastructure for abstract rule-following is present in the weights before any demonstrations; in-context examples, function vectors, and related interventions appear to supply input to the same latent circuit.
Chinese Translation
大语言模型仅凭少量上下文示例就能学习抽象规则,但随着示例的累积,其内部机制如何被激活却尚未被充分理解。我们在三个模型家族中跨不同示例数量追踪了一个三阶段的符号推理电路(抽象、归纳、检索),并发现早在模型达到高准确率之前,该电路就可被检测到并已具备功能。从1-shot到10-shot,每个注意力头的因果贡献最多增长至8倍,而跨示例的激活修补在0-shot时将准确率从1%提升到56%,在1-shot时从17%提升到88%。在0-shot时对函数向量进行缩放并注入可挽救准确率至最高86%,这在很大程度上替代了归纳阶段,但关键依赖于一个完好无损的下游检索阶段。在任何示例演示之前,遵循抽象规则所需的基础设施就已存在于权重之中;上下文示例、函数向量以及相关干预似乎都在为同一个潜在电路提供输入。
cs.LG / 125 / 2609.36294
When Trees Are Not Enough: Learning Mixed-Topology Feature Graphs with Adaptive Graph Sparse Autoencoders
当树不够时:用自适应图稀疏自编码器学习混合拓扑特征图
large language model
大语言模型相关
Abstract
Sparse autoencoders (SAEs) expose interpretable features in large language model activations, yet existing structured SAEs impose single-parent trees or forests, while post-hoc graphs permit multiple parents but neither guide feature learning nor ensure reliable relation recovery. We introduce the Adaptive Graph Sparse Autoencoder (AG-SAE), a structure-guided training paradigm that treats each feature's complete parent set as an atomic structural hypothesis and lets evidence select zero, one, or multiple parents. By competing complete parent sets against null, subset, and alternative explanations, AG-SAE identifies jointly necessary multi-parent relations while rejecting redundant or spurious alternatives and verifying that each child contributes beyond its parents. The induced topology over SAE features then defines a differentiable structural loss that guides SAE training, while topology-guided refinement mitigates feature absorption and uses persistent reconstruction gaps exposed by the learned structure to initialize new features. The entire graph is then induced again from the revised dictionary by reassessing every feature's complete parent set, closing the dictionary-graph self-consistency cycle. Experiments demonstrate exact mixed-topology recovery in a controlled toy model, greater relational reliability and semantic validity than structured and post-hoc baselines on real LLM activations, and stronger feature-level causal interventions than conventional SAE features. AG-SAE thereby turns recovered mixed-topology feature structure into an unsupervised training signal that improves the dictionary, enables reliable feature organization beyond the topological limitations of trees, and exhibits stronger causal control beyond reconstruction.
Chinese Translation
稀疏自编码器(SAEs)在大型语言模型激活值中揭示可解释特征,然而现有的结构化 SAEs 强加了单亲树或森林,而事后图允许多个父节点,但既不指导特征学习,也不确保可靠的关系恢复。我们提出自适应图稀疏自编码器(AG-SAE),这是一种结构引导的训练范式,它将每个特征的完整父集合视为一个原子结构假设,并让证据选择零个、一个或多个父节点。通过将完整父集合与空解释、子集解释和替代解释进行竞争,AG-SAE 识别联合必要的多父关系,同时拒绝冗余或虚假的替代解释,并验证每个子节点在其父节点之外仍有贡献。随后,在 SAE 特征上诱导出的拓扑定义了一个可微的结构损失,用以指导 SAE 训练;而拓扑引导的细化缓解特征吸收,并利用学习到的结构所暴露出的持续重构差距来初始化新特征。然后,通过重新评估每个特征的完整父集合,从修订后的字典再次诱导出整个图,从而闭合字典-图自洽循环。实验表明,在受控玩具模型中实现了精确的混合拓扑恢复;在真实 LLM 激活值上,相比结构化和事后基线具有更高的关系可靠性和语义有效性;并且相比传统 SAE 特征具有更强的特征级因果干预能力。因此,AG-SAE 将恢复出的混合拓扑特征结构转化为一种无监督训练信号,该信号改进字典,使可靠的特征组织能够超越树的拓扑限制,并表现出超越重构的更强因果控制。
cs.LG / 126 / 2609.36451
Invariant Atoms: Sparse Coordinates of Local Semantic Geometry in Language Model Representations
不变原子:语言模型表示中局部语义几何的稀疏坐标
large language model
大语言模型相关
Abstract
Large language models often preserve meaning despite substantial changes in wording, style, and syntax, while small semantic edits can systematically alter their hidden representations. This suggests that semantic variation may be organized along recurring local directions. We propose the Invariant Atom Hypothesis: local semantic motion admits preferred sparse coordinates along directions that remain stable under meaning-preserving transformations. We learn a shared semantic frame and sparse coordinates that reconstruct semantic displacements while suppressing nuisance variation, with anchor-dependent diagonal modulation adjusting atom strengths without sample-specific rotations. Empirically, the atoms exhibit strong semantic--nuisance separation, sparse reconstruction, reproducible directions, and causal effects on model predictions. The learned geometry generalizes to unseen semantic neighborhoods and nuisance families, while local reweighting improves semantic selectivity and preserves a consistent global-to-local structure. Atom signatures also remain stable under model modification. These findings support reusable invariant directions as a sparse coordinate system for local semantic geometry in language models.
Chinese Translation
大型语言模型常常在措辞、风格和句法发生显著变化时仍保持意义不变,而微小的语义编辑却能够系统性地改变其隐藏表示。这表明,语义变化可能是沿着反复出现的局部方向组织起来的。我们提出不变原子假设:局部语义运动允许沿着在保义变换下保持稳定的方向形成偏好的稀疏坐标。我们学习一个共享语义框架和稀疏坐标,它们能够重建语义位移,同时抑制干扰变化,其中依赖于锚点的对角调制会调整原子强度,而不进行样本特定的旋转。在实证上,这些原子表现出很强的语义--干扰分离、稀疏重建、可复现方向,以及对模型预测的因果效应。所学几何能够泛化到未见过的语义邻域和干扰族,同时局部重加权提高了语义选择性,并保持了一致的从全局到局部的结构。原子签名在模型修改下也保持稳定。这些发现支持将可复用的不变方向作为语言模型中局部语义几何的稀疏坐标系统。
cs.LG / 127 / 2609.36455
Emergent phases of superposition: from partial to full representation
叠加的涌现相:从部分表示到完全表示
large language model
大语言模型相关
Abstract
Large language models are thought to represent features by vectors in a hidden space of dimension given by the model's width. Superposition, in which more features are represented than the width by letting representation vectors overlap, is a leading account of how representation vectors are organized. However, how model width and data statistics determine the configuration of representation vectors and the resulting loss when the number of features and the width are large remains less understood. Here we show, in Anthropic's toy model of superposition, that increasing the width drives a continuous phase transition from a partial-representation phase, where only a subset of features receives appreciable representation vectors while the rest vanish, to a full-representation phase, where every feature is represented. Our theory via a partial random projection approximation predicts, and experiments confirm, that the critical width grows linearly with the number of active features up to a logarithmic factor. The loss scaling changes across the transition: below the critical width, the loss grows linearly with the number of active features and depends weakly on the width in a form set by data statistics; above it, the loss grows approximately quadratically with the number of active features and decays inversely with the width. Non-uniform firing probabilities delay the transition and lower the loss, as more frequent features occupy more space. Our results provide an account of how model width and data statistics jointly shape representations and loss, a step toward understanding representation scaling in large models.
Chinese Translation
大型语言模型被认为通过隐藏空间中的向量来表示特征,该空间的维度由模型的宽度决定。叠加,即通过让表示向量重叠来表示比宽度更多的特征,是关于表示向量如何组织的一种主流解释。然而,当特征数量和宽度都很大时,模型宽度和数据统计量如何决定表示向量的配置以及由此产生的损失,仍然较少被理解。在这里,我们在 Anthropic 的叠加玩具模型中表明,增加宽度会驱动一个连续相变,从部分表示相——其中只有一部分特征获得可观的表示向量,而其余特征消失——转变为完全表示相——其中每个特征都被表示。我们通过部分随机投影近似得到的理论预测,并且实验证实,临界宽度随活跃特征数量线性增长,直至一个对数因子。损失缩放会在相变前后发生变化:在临界宽度以下,损失随活跃特征数量线性增长,并以由数据统计量设定的形式微弱依赖于宽度;在临界宽度以上,损失随活跃特征数量近似二次增长,并随宽度反比衰减。非均匀激活概率会延迟相变并降低损失,因为更频繁的特征占据更多空间。我们的结果为模型宽度和数据统计量如何共同塑造表示和损失提供了一种解释,是朝着理解大型模型中表示缩放迈出的一步。
cs.LG / 128 / 2609.36488
AdaptArena: Evaluating Test-Time Personalization of Web Agents
AdaptArena:评估 Web 智能体的测试时个性化
large language model
大语言模型相关
Abstract
Large language model (LLM) agents have demonstrated strong performance on complex web navigation tasks, yet they remain brittle in real-world settings where user intentions are underspecified and preferences are heterogeneous. In practice, users rarely provide explicit profiles, requiring agents to infer latent preferences from implicit signals. Despite its importance for deployment, this problem setting is largely underexplored in existing benchmarks. To address this gap, we introduce AdaptArena, a benchmark for evaluating test-time personalization of web agents via implicit preference inference. AdaptArena consists of 480 tasks, featuring both single-preference and double-preference scenarios. Each evaluation task must be solved by retrieving and leveraging the most relevant historical user trajectory that implicitly encodes the target preference. In addition, we introduce AdaptiveAgent, a retrieval-based framework for standardized evaluation of implicit preference inference. Experiments reveal a substantial performance gap: while oracle agents with access to ground-truth preferences achieve an 82.92% success rate, the evaluated LLM agents using our framework reach at most 15.62%. Furthermore, we find that correctly inferring user preferences is necessary but not sufficient for task success, as execution failures in downstream web interactions remain a significant bottleneck even when agents align with the target preference. These findings highlight implicit preference inference and robust action grounding as key challenges for deploying reliable, user-facing web agents. We release our code: https://github.com/McGill-NLP/web-agents-test-time-adaptations
Chinese Translation
大型语言模型(LLM)智能体在复杂 Web 导航任务上表现出强大性能,但在用户意图未被充分指定且偏好异质的现实世界环境中,它们仍然脆弱。在实践中,用户很少提供显式画像,这要求智能体从隐式信号中推断潜在偏好。尽管这对于部署很重要,但现有基准在很大程度上尚未充分探索这一问题设定。为弥补这一空白,我们提出 AdaptArena,一个通过隐式偏好推断来评估 Web 智能体测试时个性化的基准。AdaptArena 包含 480 个任务,涵盖单偏好和双偏好场景。每个评估任务必须通过检索并利用隐式编码目标偏好的最相关历史用户轨迹来解决。此外,我们提出 AdaptiveAgent,一个基于检索的框架,用于对隐式偏好推断进行标准化评估。实验揭示了一个显著的性能差距:能够访问真实偏好的 oracle 智能体达到 82.92% 的成功率,而使用我们框架的受评估 LLM 智能体最多仅达到 15.62%。此外,我们发现正确推断用户偏好是任务成功的必要条件,但并非充分条件,因为即使智能体与目标偏好一致,下游 Web 交互中的执行失败仍然是一个显著瓶颈。这些发现凸显了隐式偏好推断和稳健的动作接地(action grounding)是部署可靠的、面向用户的 Web 智能体的关键挑战。我们发布代码:https://github.com/McGill-NLP/web-agents-test-time-adaptations
cs.LG / 129 / 2609.36527
SCOPE: Observation-Conditioned Full-Target Prediction for Sparse PDE Inference
SCOPE:面向稀疏 PDE 推断的观测条件化全目标预测
diffusion
扩散模型相关
Abstract
Recovering complete physical fields from sparse observations is challenging because the measurements may not uniquely determine the underlying state. Diffusion-based PDE solvers address this problem through iterative sampling whereas neural operators provide deterministic one-pass predictions. We propose SCOPE (Sparse-Context Observability-aware Predictive Embeddings) to recover complete PDE fields from sparse observations by coupling full-field latent prediction with physical reconstruction. A shared decoder reconstructs fields from both predicted and complete-view representations so that representation learning is guided by both physical recovery and latent matching. We derive a quadratic risk decomposition at fixed teacher-decoder pairs showing why optimal latent prediction need not yield optimal field reconstruction. We also establish sufficient conditions for decoder improvements on complete inputs to transfer to recovery from partial observations. Experiments across five PDE settings show that SCOPE outperforms mask-aware neural operators on all ten forward and inverse tasks and achieves lower errors than those reported for diffusion-based solvers including DiffusionPDE and FunDPS. Decoder-only adaptation further improves recovery without retraining the backbone while retaining deterministic single-pass inference.
Chinese Translation
从稀疏观测中恢复完整物理场具有挑战性,因为这些测量可能无法唯一确定潜在状态。基于扩散的 PDE 求解器通过迭代采样解决这一问题,而神经算子则提供确定性的单遍预测。我们提出 SCOPE(稀疏上下文可观测性感知预测嵌入),通过将全场潜在预测与物理重建耦合,从稀疏观测中恢复完整 PDE 场。一个共享解码器从预测表示和完整视图表示二者重建场,从而使表示学习同时由物理恢复和潜在匹配引导。我们推导了固定教师-解码器对下的二次风险分解,表明为什么最优潜在预测未必产生最优场重建。我们还建立了充分条件,使得在完整输入上的解码器改进能够迁移到从部分观测进行恢复。跨五种 PDE 设置的实验表明,SCOPE 在所有十项正向和逆向任务上均优于掩码感知神经算子,并且相比包括 DiffusionPDE 和 FunDPS 在内的基于扩散的求解器所报告的结果取得了更低的误差。仅解码器适配进一步提高了恢复效果,而无需重新训练主干网络,同时保留确定性的单遍推断。
cs.LG / 130 / 2609.36568
Sharp Convergence and Sampling Trade-offs for Riemannian Diffusion under Nonnegative Ricci Curvature
非负 Ricci 曲率下黎曼扩散的尖锐收敛与采样权衡
diffusion
扩散模型相关
Abstract
Diffusion models have emerged as state-of-the-art generative models, with recent extensions from Euclidean spaces to Riemannian manifolds. However, existing convergence guarantees for Riemannian diffusion models typically require $\tilde{O}(\mathrm{poly}(d,T)/ε^2)$ score evaluations, with potentially unfavorable dependence on the dimension. In this work, we develop a general framework that separates score discretization from Brownian-motion simulation and allows multiple geodesic random-walk steps per score evaluation. Under nonnegative Ricci curvature assumption and an exact Brownian-motion simulation oracle, we show that $\tilde{O}(d/ε^2)$ score evaluations suffice to achieve an $ε^2$ KL divergence from the target distribution, matching the existing convergence rate of Euclidean diffusion models. We further show that $\tilde{O}(d^4T/ε^2)$ geodesic random-walk steps suffice to approximate the required drifted Brownian motion to $ε$ total variation error. Combining these results yields a sampling scheme with $\tilde{O}(d/ε^2)$ score evaluations and $\tilde{O}(d^4T/ε^2)$ geodesic random-walk steps, motivating multiple random-walk steps between consecutive score evaluations. Our results provide a sharper characterization of the convergence and sampling complexity of Riemannian diffusion models.
Chinese Translation
扩散模型已成为最先进的生成模型,近来其扩展从欧几里得空间扩展到黎曼流形。然而,黎曼扩散模型现有的收敛保证通常需要 $\tilde{O}(\mathrm{poly}(d,T)/ε^2)$ 次分数评估,并且可能对维度具有不利的依赖。在这项工作中,我们开发了一个通用框架,将分数离散化与布朗运动模拟分离,并允许每次分数评估进行多步测地随机游走。在非负 Ricci 曲率假设和精确的布朗运动模拟预言机下,我们证明 $\tilde{O}(d/ε^2)$ 次分数评估足以实现与目标分布之间 $ε^2$ 的 KL 散度,匹配欧几里得扩散模型现有的收敛速率。我们进一步证明,$\tilde{O}(d^4T/ε^2)$ 步测地随机游走足以将所需的带漂移布朗运动逼近到 $ε$ 全变差误差。结合这些结果,得到一个采样方案,其具有 $\tilde{O}(d/ε^2)$ 次分数评估和 $\tilde{O}(d^4T/ε^2)$ 步测地随机游走,这促进在连续分数评估之间进行多步随机游走。我们的结果为黎曼扩散模型的收敛性和采样复杂度提供了更尖锐的刻画。
cs.LG / 131 / 2609.36612
Do LLMs Really Forget? Hidden-State Leakage in Model Unlearning and How to Fix it
大语言模型真的会遗忘吗?模型遗忘中的隐藏状态泄漏及其修复方法
large language model
大语言模型相关
Abstract
Unlearning in large language models (LLMs) is typically evaluated at the output level, where a model appears to suppress sensitive or undesirable content. In this work, we show that such evaluations can create an illusion of forgetting: even when output-level leakage is eliminated, sensitive information can remain encoded in the model's hidden representations. We first provide a theoretical analysis establishing a fundamental separation between output suppression and representational erasure. Specifically, we show that the decoder can be made arbitrarily insensitive to sensitive directions, driving output-level leakage to zero, while the hidden representations retain the underlying information. To empirically validate this phenomenon, we train generative probe decoders on hidden states across transformer layers, enabling layer-wise measurement of information leakage. Across three widely used benchmarks, TOFU, MUSE, and WMDP, and state-of-the-art unlearning methods, we find that substantial sensitive information remains recoverable from hidden representations, even when standard output-level metrics indicate successful unlearning. To address this gap, we propose Probe-Adversarial Representation Suppression (PARS), an unlearning objective that adversarially minimizes the extractable information from hidden representations. PARS directly targets representational leakage and provides significantly stronger guarantees of erasure under adversarial probing and relearning attacks, outperforming all evaluated baselines. Our results highlight a fundamental limitation of existing unlearning paradigms and suggest that true forgetting in LLMs requires controlling not only model outputs, but also the information encoded in hidden representations. Codes are available at https://github.com/OptimAI-Lab/HiddenStateUnlearning.
Chinese Translation
大语言模型(LLM)中的遗忘通常是在输出层面进行评估,在这种评估中,模型似乎抑制了敏感或不良内容。在这项工作中,我们表明,此类评估可能造成一种遗忘的假象:即使输出层面的泄漏已被消除,敏感信息仍可能被编码在模型的隐藏表示中。我们首先提供理论分析,确立了输出抑制与表示擦除之间的根本性分离。具体而言,我们表明,可以使解码器对敏感方向任意不敏感,从而将输出层面的泄漏降至零,而隐藏表示仍保留底层信息。为了从经验上验证这一现象,我们在跨 Transformer 层的隐藏状态上训练生成式探测解码器,从而能够逐层测量信息泄漏。在三个广泛使用的基准 TOFU、MUSE 和 WMDP 上,以及最先进的遗忘方法中,我们发现,即使标准输出层面指标表明遗忘成功,大量敏感信息仍可从隐藏表示中恢复。为弥补这一差距,我们提出探测对抗表示抑制(Probe-Adversarial Representation Suppression,PARS),这是一种遗忘目标,它以对抗方式最小化可从隐藏表示中提取的信息。PARS 直接针对表示泄漏,并在对抗性探测和重新学习攻击下提供显著更强的擦除保证,优于所有评估的基线。我们的结果凸显了现有遗忘范式的一个根本性局限,并表明 LLM 中真正的遗忘不仅需要控制模型输出,还需要控制隐藏表示中编码的信息。代码可在 https://github.com/OptimAI-Lab/HiddenStateUnlearning 获取。
cs.LG / 132 / 2609.36654
Replay the Curvature: Accurate and Scalable NVFP4 Quantization for Large Language Model Inference
重放曲率:面向大语言模型推理的精确且可扩展的 NVFP4 量化
large language model
大语言模型相关
Abstract
Large language models make weight storage and memory traffic major inference costs, motivating low-precision formats that represent each weight with only a few bits. Such formats use a scale to map floating-point values into a small codebook; NVFP4 improves local range utilization by letting every 16 E2M1 weights share an E4M3 block scale. Choosing that scale is difficult in GPTQ because quantizing one column updates those that follow, so evaluating a block independently can misestimate its final reconstruction error. Large models pose a second challenge: full-precision weights, calibration activations, and second-order state cannot all remain on one accelerator, while assigning complete layers to devices leaves each time-consuming layer solve serial. We introduce \emph{Schur Replay}, a scale-selection algorithm that reproduces the GPTQ updates caused by each block scale and scores the resulting block error after accounting for compensation from unquantized columns. Separately, our execution infrastructure keeps only the active layer resident, tiers activations across device, host, and disk, retires full-precision layers after export, and distributes independent output rows across tensor-parallel ranks. Together, the algorithm and infrastructure attain $99.35\%$ and $100.84\%$ question-weighted recovery from BF16 across seven benchmarks on Qwen3.5-397B-A17B and Llama-3.3-70B-Instruct. On the 397B model, the infrastructure reduces measured per-layer time by $15.17\times$ over ModelOpt and $23.14\times$ over LLM Compressor, with lower memory used per GPU.
Chinese Translation
大语言模型使权重存储与内存流量成为主要的推理成本,这促使人们采用仅用少量比特表示每个权重的低精度格式。这类格式使用一个缩放因子将浮点值映射到一个小码本中;NVFP4 通过让每 16 个 E2M1 权重共享一个 E4M3 块缩放因子,改善了局部范围的利用率。在 GPTQ 中选择该缩放因子很困难,因为量化一列会更新其后的列,因此独立评估一个块可能会错误估计其最终的重建误差。大模型带来第二个挑战:全精度权重、校准激活和二阶状态无法全部留在一块加速器上,而将完整的层分配给各设备又会使每个耗时的层求解保持串行。我们提出 \emph{Schur Replay},一种缩放因子选择算法,它重现由每个块缩放因子引起的 GPTQ 更新,并在计入未量化列的补偿之后对所得的块误差进行评分。另外,我们的执行基础设施仅保留活跃层常驻,将激活分层存放于设备、主机和磁盘,在导出后释放全精度层,并将相互独立的输出行分布到张量并行 rank 上。二者结合,在 Qwen3.5-397B-A17B 和 Llama-3.3-70B-Instruct 上,该算法与基础设施在七个基准测试中相对于 BF16 分别达到 $99.35\%$ 和 $100.84\%$ 的按问题加权恢复率。在 397B 模型上,该基础设施将实测的每层时间相比 ModelOpt 减少了 $15.17\times$,相比 LLM Compressor 减少了 $23.14\times$,同时每个 GPU 所使用的内存更低。
cs.LG / 133 / 2609.36662
AutoLoCo: Communication Efficient Distributed LLM Training via Adaptive Synchronization
AutoLoCo:通过自适应同步实现通信高效的分布式大语言模型训练
large language model
大语言模型相关
Abstract
The pre-training of Large Language Models (LLMs) is increasingly conducted across multiple data centers. As training scales to a larger number of accelerators, the fraction of time spent on computation decreases, while the fraction spent on communication increases. Therefore, frequent synchronization becomes a growing bottleneck. Local update methods reduce this cost by allowing workers to perform several optimizer steps between synchronizations. Most local update methods set the number of local optimizer steps between synchronizations before training and keep this interval fixed throughout the run. However, the best interval can change during the entire train process. If the interval and optimizer are adapted to the current training state, the communication frequency is reduced while maintaining the training performance. In this work, we introduce AutoLoCo, an adaptive training framework to reduce communication in LLM training. It adapts the local interval using scalar training statistics and corrects each outer update. Our method is motivated by two observations: 1) the appropriate local interval varies across training stages, and 2) changing the number of inner steps per interval creates a mismatch with an unchanged outer optimizer, requiring a correction to the outer update. We optimize this mismatch by correction of the outer optimizer for the momentum and the learning rate using the accumulated inner learning rate. Our experiments under communication constraints demonstrate that AutoLoCo reduces communication frequency by 27% relative to DiLoCo while maintaining training performance.
Chinese Translation
大语言模型(LLM)的预训练正日益跨越多个数据中心进行。随着训练扩展到更多数量的加速器,用于计算的时间占比下降,而用于通信的时间占比上升。因此,频繁的同步正成为日益严重的瓶颈。本地更新方法通过允许工作节点在两次同步之间执行若干次优化器步骤,降低了这一开销。大多数本地更新方法在训练前设定同步之间的本地优化器步数,并在整个训练过程中保持该间隔固定不变。然而,最佳间隔在整个训练过程中可能会发生变化。如果间隔与优化器能够适应当前的训练状态,就能在保持训练性能的同时降低通信频率。在本工作中,我们提出了 AutoLoCo,一个用于减少 LLM 训练中通信量的自适应训练框架。它利用标量训练统计量自适应调整本地间隔,并对每次外层更新进行校正。我们的方法源于两个观察:1)合适的本地间隔在不同训练阶段会发生变化;2)改变每个间隔内的内层步数会与保持不变的外层优化器产生不匹配,因而需要对外层更新进行校正。我们通过利用累积的内层学习率对外层优化器的动量与学习率进行校正,来优化这一不匹配。我们在通信受限条件下的实验表明,相对于 DiLoCo,AutoLoCo 在保持训练性能的同时将通信频率降低了 27%。
cs.LG / 134 / 2609.36689
CHAIN: Calibrated LLM Forecasting via Causal-Temporal Hypergraph Inference
CHAIN:通过因果-时序超图推理的校准大语言模型预测
large language model
大语言模型相关
Abstract
Large language models have achieved significant progress in event forecasting, yet their probability outputs exhibit systematic calibration bias that varies heterogeneously across different domains and question types, undermining the trustworthiness of probabilistic outputs for decision-making under uncertainty. However, existing calibration methods typically correct probability outputs after prediction is complete, without modeling the structural sources of bias within the prediction process itself. To address this challenge, we decompose probabilistic prediction over causal-temporal hypergraphs into three stages, evidence weighting, evidence aggregation, and source fusion, and propose CHAIN, which designs stage-specific mechanisms to mitigate bias at each stage: (i) modulating the temporal decay function by causal topological distance, (ii) aggregating approximately independent causal chains via Noisy-OR after direction-aware deduplication, and (iii) driving adaptive fusion by causal coverage and directional balance. Experimental results on cross-domain forecasting benchmarks show CHAIN outperforms existing methods in expected calibration error, Brier score, and accuracy. Our project is available at https://github.com/QwenQKing/Chain.
Chinese Translation
大语言模型在事件预测方面取得了显著进展,然而其概率输出表现出系统性的校准偏差,这种偏差在不同领域和问题类型之间异质地变化,削弱了概率输出在不确定性下决策中的可信度。然而,现有的校准方法通常是在预测完成后校正概率输出,而没有对预测过程本身内部偏差的结构性来源进行建模。为应对这一挑战,我们将因果-时序超图上的概率预测分解为三个阶段:证据加权、证据聚合和来源融合,并提出 CHAIN,其设计针对每个阶段的特定机制以减轻偏差:(i) 通过因果拓扑距离调节时间衰减函数,(ii) 在方向感知去重后通过 Noisy-OR 聚合近似独立的因果链,以及 (iii) 由因果覆盖度和方向平衡驱动自适应融合。在跨领域预测基准上的实验结果表明,CHAIN 在期望校准误差、Brier 分数和准确率方面优于现有方法。我们的项目可在 https://github.com/QwenQKing/Chain 获取。
cs.LG / 135 / 2609.36750
Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving
组边缘化自奖励强化学习驱动零标签自演化
large language model
大语言模型相关
Abstract
Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response's reward representation also depends on its randomly sampled group context, i.e., the other responses in its group. Using only one group-context realization may miss desired reward signals and provide unreliable guidance for policy optimization. To address this issue, we propose Group-Marginalized Advantage Estimation (GMAE), which aggregates reward realizations across possible contexts into a response-level distribution and estimates expected advantages. Experiments across eight benchmarks and four base models demonstrate strong performance and cross-domain generalization. GMAE also exhibits stable learning, low extra cost, and good applicability across training datasets and RL backbones.
Chinese Translation
自奖励强化学习(RL)使大型语言模型(LLM)无需人工标签即可自演化。现有的基于集成的方法从 rollout 组构建奖励参考,并据此分配奖励。然而,一个响应的奖励表示还取决于其随机采样的组上下文,即其组中的其他响应。仅使用一种组上下文实现可能会遗漏期望的奖励信号,并为策略优化提供不可靠的指导。为了解决这个问题,我们提出组边缘化优势估计(GMAE),它将跨可能上下文的奖励实现聚合为响应级分布,并估计期望优势。在八个基准和四个基础模型上的实验展示了强劲的性能和跨领域泛化能力。GMAE 还表现出稳定的学习、较低的额外成本,以及在训练数据集和 RL 骨干网络上的良好适用性。
cs.LG / 136 / 2609.36788
Harnessing Large Language Models to Compile Task-Relevant Context into Bayesian Optimisation
利用大型语言模型将任务相关上下文编译到贝叶斯优化中
large language model
大语言模型相关
Abstract
Incorporating rich task-relevant context, such as domain knowledge and external observations, is a key capability yet remains challenging for Bayesian optimisation (BO). Recently, practitioners have started to use large language models (LLMs) to generate and execute BO programs through coding harnesses. In such emerging practices, the posterior belief is shaped not only by Bayesian inference but also by LLM-generated model and data artefacts, offering a flexible route for task context to enter BO as executable code. To study whether and how LLMs can be harnessed to compile diverse contextual signals for BO, we formulate LLM-compiled BO as generalised-context decision making. We propose HarBO, a BO-specialised harness that compiles generalised context into the core artefacts of standard BO through a validated multi-stage workflow. Our theory analyses the regret under imperfect compilation and the effect of adding new context. Across synthetic functions and real-world benchmarks, we find that LLM harnesses can effectively compile context into standard BO, achieving competitive performance with specialised LLM-embedding-based and direct LLM-in-the-loop BO methods. General coding harnesses can be effective in familiar domains such as hyperparameter optimisation, but fall short in unfamiliar, context-rich domains. Together, these results establish LLM harnesses as a promising, but not automatically reliable, route for making rich task context usable in BO.
Chinese Translation
融入丰富的任务相关上下文,例如领域知识和外部观测,是一项关键能力,但对贝叶斯优化(BO)而言仍然具有挑战性。近年来,从业者已开始使用大型语言模型(LLM)通过编码框架生成并执行 BO 程序。在这种新兴实践中,后验信念不仅由贝叶斯推断塑造,也由 LLM 生成的模型和数据制品塑造,从而为任务上下文以可执行代码的形式进入 BO 提供了一条灵活路径。为研究 LLM 能否以及如何被用于为 BO 编译多样化的上下文信号,我们将 LLM 编译的 BO 形式化为广义上下文决策。我们提出 HarBO,一个面向 BO 的专用框架,它通过经过验证的多阶段工作流将广义上下文编译为标准 BO 的核心制品。我们的理论分析了不完美编译下的遗憾以及添加新上下文的影响。在合成函数和真实世界基准上,我们发现 LLM 框架能够有效地将上下文编译到标准 BO 中,达到与专门的基于 LLM 嵌入的方法以及直接的 LLM 在环 BO 方法相当的性能。通用编码框架在诸如超参数优化等熟悉领域中可以有效,但在不熟悉的、上下文丰富的领域中表现不足。总之,这些结果确立了 LLM 框架是一条有前景但并非自动可靠的路径,可使丰富的任务上下文在 BO 中可用。
cs.LG / 137 / 2609.36802
EasyPPO: Stabilizing the Critic Is Key
EasyPPO:稳定 critic 是关键
large language model
大语言模型相关
Abstract
A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning to estimate expected returns and reduce policy-gradient variance. However, we find that the critic is also a major source of instability in reinforcement learning for large language models (LLMs). We identify two critic failure modes that destabilize PPO. First, filtering truncated rollouts from both actor and critic shifts the policy objective to reward conditioned on completion, allowing truncation to increase even as conditional reward improves. Second, heterogeneous return noise can cause high-variance prompts to dominate critic updates in finite batches. We introduce EasyPPO to address these failures. Actor-only overlong filtering trains the critic on returns from both completed and truncated rollouts. Noise-normalized critic regression weights each prompt's critic loss by the inverse standard deviation of its sampled returns, balancing noise contributions across prompts. Moderately smaller critic mini-batches confine outlier influence to fewer rollouts during gradient clipping. Across continuous-reward coding on FrontierCS, binary-reward mathematical reasoning on AIME24, and multi-turn search on Search-R1, EasyPPO remains stable throughout the full training horizon and consistently outperforms vanilla PPO, VAPO, and HL-Gauss PPO. Its best validation scores show relative gains of 14.89%, 2.28%, and 9.47% over PPO, respectively.
Chinese Translation
近端策略优化(PPO)的一个关键优势在于其学习得到的 critic,它利用强化学习过程中收集的历史轨迹来估计期望回报并降低策略梯度方差。然而,我们发现 critic 也是大语言模型(LLM)强化学习中不稳定的一个主要来源。我们识别出两种会使 PPO 失稳的 critic 失效模式。第一,从 actor 和 critic 中同时过滤掉被截断的 rollout 会将策略目标转变为以完成为条件的奖励,从而使得即便条件奖励有所提升,截断仍可能增加。第二,异质性的回报噪声会导致高方差的 prompt 在有限批次中主导 critic 更新。我们提出 EasyPPO 来应对这些失效。仅针对 actor 的超长过滤让 critic 在来自已完成和被截断 rollout 的回报上进行训练。噪声归一化的 critic 回归按照每个 prompt 采样回报的标准差的倒数对其 critic 损失进行加权,从而在 prompt 之间平衡噪声贡献。适度更小的 critic 小批次在梯度裁剪期间将离群值的影响限制在更少的 rollout 上。在 FrontierCS 上的连续奖励编程、AIME24 上的二值奖励数学推理以及 Search-R1 上的多轮搜索中,EasyPPO 在整个训练周期内始终保持稳定,并持续优于 vanilla PPO、VAPO 和 HL-Gauss PPO。其最佳验证分数相较于 PPO 分别显示出 14.89%、2.28% 和 9.47% 的相对提升。
cs.LG / 138 / 2609.36812
Diffusion Policy Improvement with Proposal-Conditioned Refinement Flows
基于提议条件细化流的扩散策略改进
diffusion
扩散模型相关
Abstract
Diffusion and flow policies can model complex behaviors in offline reinforcement learning (RL). However, penalizing their KL divergence from the behavior policy can discourage actions having high critic values with low behavior density. Directly refining behavior proposals may be an alternative, yet Gaussian or deterministic editors limit expressiveness to represent multiple separated modes for the same proposal. In this work, we introduce Proposal-Conditioned Refinement Flows (PReFlow), a policy extraction method combining critic-based proposal selection with a conditional refinement flow. To optimize proposal selection and refinement together, we formulate a KL-regularized objective whose optimum induces a Gibbs policy over final actions under a Gaussian-smoothed behavior prior. The refinement flow can represent multiple high value modes, while a proposal-centered Gaussian reference regulates large action changes. This Gaussian reference further enables us to make use of simulation-free, closed form adjoint matching targets from sampled endpoints and critic gradients, yielding a single velocity regression loss without a backward adjoint solve. On 50 OGBench tasks, PReFlow achieves competitive offline performance and the highest aggregate score among the compared methods after online fine-tuning, reaching 91\% after 500K environment steps.
Chinese Translation
扩散策略和流策略能够对离线强化学习(RL)中的复杂行为进行建模。然而,惩罚它们与行为策略之间的 KL 散度,可能会抑制那些具有高 critic 值但行为密度低的动作。直接细化行为提议可能是一种替代方案,但高斯或确定性编辑器限制了表达能力,使其无法为同一提议表示多个分离的模态。在这项工作中,我们引入了提议条件细化流(Proposal-Conditioned Refinement Flows, PReFlow),这是一种将基于 critic 的提议选择与条件细化流相结合的策略提取方法。为了联合优化提议选择与细化,我们构建了一个 KL 正则化目标,其最优解在高斯平滑的行为先验下诱导出关于最终动作的 Gibbs 策略。细化流能够表示多个高价值模态,而以提议为中心的高斯参考则调节大幅度的动作变化。这一高斯参考还使我们能够利用来自采样端点和 critic 梯度的无模拟、闭式伴随匹配目标,从而产生单个速度回归损失,而无需反向伴随求解。在 50 个 OGBench 任务上,PReFlow 取得了有竞争力的离线性能,并在在线微调后于比较方法中取得最高的总得分,在 500K 环境步后达到 91\%。
cs.LG / 139 / 2609.36813
RESCUE: Repairing Language Model Errors to Sparse Circuits via Reinforcement Learning
RESCUE:通过强化学习将语言模型错误修复至稀疏电路
large language model
大语言模型相关
Abstract
Large language models (LLMs) exhibit strong general capabilities that mechanistic interpretability has attributed to sparse computational circuits. However, existing circuit studies emphasize preserving functionality or explaining safety, leaving the mechanisms underlying failures across a broader range of tasks largely unexplored. Extending circuit analysis from abilities to errors, we explore the perspective that such failures may likewise arise from erroneous internal computations and that targeted tuning of the corresponding parameters can correct such errors while largely preserving other capabilities. Motivated by this insight, we introduce RESCUE (Reasoning-Error Sparse-Circuit Uncovering and Editing), a framework that localizes error-associated circuits and surgically repairs them for performance enhancement. General tasks typically involve multi-step reasoning and long-form generation, where early deviations can cause prefixes to drift from supervised references, leading SFT-based mask optimization to overlook circuits involved in generation-time errors. RESCUE therefore refines these masks through reinforcement learning with multiple masked-model rollouts, improving their relevance to observed task failures. Finally, RESCUE introduces a pruning technique and precisely fine-tunes error circuits to correct task failures, thereby translating error localization into a sparse and targeted model update. We validate RESCUE on heterogeneous repair sets across two domains: (1) mathematical reasoning, identifying a math error circuit of 1.40% density whose repair raises accuracy from 6.0% to 75.5%; and (2) medical QA, where a similarly compact 1.44% circuit improves repair-set accuracy from 0% to 81%. Our code is available at: https://github.com/chuanpupig/RESCUE.
Chinese Translation
大型语言模型(LLM)展现出强大的通用能力,机制可解释性将这些能力归因于稀疏计算电路。然而,现有的电路研究强调保持功能或解释安全性,使得更广泛任务中失败背后的机制在很大程度上尚未被探索。将电路分析从能力扩展到错误,我们探索这样一种观点:此类失败可能同样源于错误的内部计算,并且对相应参数进行有针对性的调优可以纠正此类错误,同时大体保留其他能力。受这一洞见启发,我们提出 RESCUE(Reasoning-Error Sparse-Circuit Uncovering and Editing,推理错误稀疏电路发现与编辑),一个定位与错误相关的电路并以手术式方式修复它们以提升性能的框架。通用任务通常涉及多步推理和长文本生成,其中早期偏差可能导致前缀偏离监督参考,从而导致基于 SFT 的掩码优化忽略生成时错误所涉及的电路。因此,RESCUE 通过使用多个掩码模型采样轨迹的强化学习来细化这些掩码,提高它们与所观察到任务失败的相关性。最后,RESCUE 引入一种剪枝技术,并精确微调错误电路以纠正任务失败,从而将错误定位转化为稀疏且有针对性的模型更新。我们在跨两个领域的异构修复集上验证 RESCUE:(1)数学推理,识别出一个密度为 1.40% 的数学错误电路,修复该电路将准确率从 6.0% 提高到 75.5%;以及(2)医学问答,其中一个类似紧凑的 1.44% 电路将修复集准确率从 0% 提高到 81%。我们的代码可在以下网址获取:https://github.com/chuanpupig/RESCUE。
cs.LG / 140 / 2609.36816
Towards Better Training Signal: Advantage Clipped Policy Optimization
迈向更好的训练信号:优势裁剪策略优化
large language model
大语言模型相关
Abstract
Reinforcement learning (RL) has become a cornerstone for improving the reasoning capabilities of large language models (LLMs), but the need for on-policy data substantially limits training efficiency. Reusing off-policy data through importance sampling (IS) can improve efficiency but introduce considerable instability. Hence, algorithms such as PPO and GRPO widely adopt IS-ratio clipping to stabilize training. However, training stability and gradient estimate are mainly determined by the product of IS ratio and advantage. To further stabilize training, we propose ACPO, which clips the product of the IS ratio and the advantage, leading to more stable gradient estimates. We also establish a connection between ACPO and gradient clipping in policy mirror descent (PMD), which is a standard technique to stabilize optimization process, and prove the convergence of clipped-PMD under the standard RL setting. Experiments on widely used mathematical reasoning benchmarks show that ACPO consistently outperforms PPO and GRPO in both accuracy and training efficiency, delivering 4-6 percentage points gains on standard math benchmarks, with Qwen3-8B+PPO. Hence, ACPO is a practical and effective alternative to conventional IS-ratio clipping for RL post-training of LLMs.
Chinese Translation
强化学习(RL)已成为提升大语言模型(LLMs)推理能力的基石,但对同策略(on-policy)数据的需求极大地限制了训练效率。通过重要性采样(IS)复用异策略(off-policy)数据可以提高效率,但会引入相当大的不稳定性。因此,PPO 和 GRPO 等算法广泛采用 IS 比率裁剪来稳定训练。然而,训练稳定性和梯度估计主要由 IS 比率与优势的乘积决定。为了进一步稳定训练,我们提出了 ACPO,它对 IS 比率与优势的乘积进行裁剪,从而得到更稳定的梯度估计。我们还建立了 ACPO 与策略镜像下降(PMD)中梯度裁剪之间的联系,梯度裁剪是稳定优化过程的一种标准技术,并证明了在标准 RL 设定下裁剪版 PMD 的收敛性。在广泛使用的数学推理基准上的实验表明,ACPO 在准确率和训练效率上均持续优于 PPO 和 GRPO,在 Qwen3-8B+PPO 上于标准数学基准取得 4-6 个百分点的提升。因此,对于 LLMs 的 RL 后训练而言,ACPO 是传统 IS 比率裁剪的一种实用且有效的替代方案。
cs.LG / 141 / 2609.36945
Fine-Tuning on Self-Generated and Reward-Weighted Data: Learning Dynamics, Convergence Rates, and Benefits of Off-Policyness
自生成与奖励加权数据上的微调:学习动力学、收敛速率与离策略性的优势
large language model
大语言模型相关
Abstract
We study the learning dynamics of fine-tuning a policy model on self-generated and reward-weighted data, with particular focus on a generalized version of REINFORCE -- referred to as RE(S) -- that updates the rollout distribution once every $S \ge 1$ gradient steps. Prior work in bandits and reinforcement learning has developed rich theory for policy gradient methods, and on-policy sampling (i.e., a small $S$, ideally $1$) is often viewed as crucial to their success; yet in prominent application like post-training large language models, reward-guided self-training has proved to be effective even when the rollout distribution is updated infrequently, but theoretical understanding remains limited for the convergence properties of these off-policy methods. To bridge these gaps, we develop a unified theory for RE(S) that covers the full spectrum of $S \ge 1$: it can be interpreted as a stage-wise optimization process, where each stage takes $S$ gradient steps for minimizing the Kullback-Leibler distance to a fixed reward-weighted rollout distribution. For multi-arm bandits with softmax policies, our in-depth analysis and numerical experiments reveal three key findings: (1) for any fixed $S$, RE(S) enjoys global convergence to the optimal policy as the number of rollout distribution updates $B = \lfloor T / S \rfloor \rightarrow \infty$, where $T$ denotes the number of gradient steps; (2) we prove tight two-sided bounds showing that the suboptimality gap of RE(S) achieves an asymptotic $Θ(1 / T)$ convergence rate, while $S$ only affects the length of a burn-in phase; (3) when initialized at a weak policy with a small optimal-action probability, RE(1) gets trapped around suboptimal policies for a long period, whereas RE(S) with a suitable $S$ avoids the detour and achieves significantly faster convergence to the global optimum, highlighting the benefits of off-policyness in this case.
Chinese Translation
我们研究在自生成且奖励加权的数据上微调策略模型的学习动力学,特别关注 REINFORCE 的一个广义版本——称为 RE(S)——它每隔 $S \ge 1$ 个梯度步更新一次 rollout 分布。在赌博机与强化学习中的先前工作已经为策略梯度方法发展了丰富的理论,并且同策略采样(即较小的 $S$,理想情况下为 $1$)常被视为其成功的关键;然而在大语言模型后训练等突出应用中,即使 rollout 分布很少更新,奖励引导的自训练也被证明是有效的,但对于这些离策略方法的收敛性质的理论理解仍然有限。为了弥合这些差距,我们为 RE(S) 提出一个统一理论,涵盖 $S \ge 1$ 的完整范围:它可以被解释为一个分阶段优化过程,其中每个阶段采取 $S$ 个梯度步,以最小化到某个固定的奖励加权 rollout 分布的 Kullback-Leibler 距离。对于使用 softmax 策略的多臂赌博机,我们的深入分析和数值实验揭示了三个关键发现:(1) 对于任何固定的 $S$,随着 rollout 分布更新次数 $B = \lfloor T / S \rfloor \rightarrow \infty$,RE(S) 具有到最优策略的全局收敛性,其中 $T$ 表示梯度步数;(2) 我们证明紧的双侧界,表明 RE(S) 的次优性间隙达到渐近 $Θ(1 / T)$ 收敛速率,而 $S$ 只影响预热阶段的长度;(3) 当以最优动作概率较小的弱策略初始化时,RE(1) 会在次优策略附近被困很长时间,而具有合适 $S$ 的 RE(S) 避免绕路,并显著更快地收敛到全局最优,从而凸显了在这种情况下离策略性的优势。
cs.LG / 142 / 2609.36950
Scalable Diffusion SBI for Compositional Inference under Simulator Misspecification
模拟器误设下用于组合推断的可扩展扩散 SBI
diffusion
扩散模型相关
Abstract
Simulation-based inference is challenging when many heterogeneous observations must be composed, hierarchical latent structure must be preserved, and the simulator is misspecified relative to observed data. We develop sampling and fine-tuning methods for diffusion-based inference in design-conditional settings, where the same simulator is queried across different experimental conditions $ξ$. We extend compositional score-based inference with a continuous-time diffusion coefficient that accounts for the number of observations, avoiding Jacobian and auxiliary-covariance corrections. We introduce Hierarchical Blockwise Diffusion Sampling (HBDS), which infers shared parameters and group-specific latent states using a single pretrained model, with the hierarchy specified only at sampling time. Together, these methods support variable observation sets and groupings without retraining. To address misspecification, we introduce path-regularized fine-tuning that adapts the learned likelihood to observations and transfers corrections to posterior inference. Using Girsanov's theorem, we quantify path divergence between pretrained and fine-tuned models across experimental designs and interpret it alongside predictive errors to distinguish candidate misspecification correction from unnecessary adaptation. We evaluate compositional sampling on exact-score Gaussian and Simple Likelihood, Complex Posterior benchmarks, HBDS with analytic and learned scores on a controlled hierarchical model, and fine-tuning and localization on a separate analytic model with known design-dependent discrepancy. Finally, we apply the framework to 940 measurements across four cell lines in a mechanistic Bone Morphogenetic Protein signaling model, where fine-tuning improves posterior-predictive accuracy relative to the pretrained model and shifts posterior marginals toward the least-squares reference while retaining spread.
Chinese Translation
当许多异质观测必须被组合、分层潜在结构必须被保留,并且模拟器相对于观测数据出现误设时,基于模拟的推断具有挑战性。我们为设计条件设置中的基于扩散的推断开发了采样和微调方法,其中同一模拟器在不同实验条件 $ξ$ 下被查询。我们用考虑观测数量的连续时间扩散系数扩展了基于分数的组合推断,避免了 Jacobian 和辅助协方差校正。我们引入分层分块扩散采样(Hierarchical Blockwise Diffusion Sampling, HBDS),其使用单个预训练模型推断共享参数和组特定潜在状态,且层次结构仅在采样时指定。这些方法共同支持可变的观测集和分组,而无需重新训练。为了解决误设,我们引入路径正则化微调,其使学到的似然适应观测,并将校正传递到后验推断。使用 Girsanov 定理,我们量化跨实验设计的预训练模型与微调模型之间的路径散度,并将其与预测误差一起解释,以区分候选误设校正与不必要的适应。我们在精确分数高斯以及 Simple Likelihood、Complex Posterior 基准上评估组合采样,在受控分层模型上评估使用解析分数和学习分数的 HBDS,并在一个具有已知设计依赖差异的单独解析模型上评估微调和定位。最后,我们将该框架应用于一个机制性骨形态发生蛋白信号模型中的四种细胞系共 940 个测量,其中微调相对于预训练模型提高了后验预测准确性,并将后验边缘分布推向最小二乘参考,同时保留其离散度。
cs.LG / 143 / 2609.37029
LongSpark: Efficient speculative decoding with a fixed-cost parallel drafter
LongSpark:使用固定成本并行草稿器的高效投机解码
diffusion
扩散模型相关
Abstract
Speculative decoding accelerates autoregressive inference by verifying multiple draft tokens in a single target forward pass. However, as the context grows, existing state-of-the-art drafters become increasingly expensive, eroding the very efficiency advantage they are designed to provide. We argue that this scaling is unnecessary. A standalone language model must grow with its prefix because it is solely responsible for every token it produces. A drafter, by contrast, only proposes candidates; the target catches and corrects every error before any token is committed. The drafter's decoding cost can therefore be made entirely independent of the prefix length. We introduce LongSpark, a block-diffusion drafter that achieves this by extracting fixed-size, multiscale views from the target's verification pass, thereby eliminating the need for a growing persistent state. Extensive evaluations demonstrate that LongSpark achieves state-of-the-art end-to-end efficiency across multiple model scales and realistic serving conditions. Notably, it delivers the lowest time-per-output-token on long-context tasks while reducing the drafter's context state by several orders of magnitude.
Chinese Translation
投机解码通过在单次目标前向传播中验证多个草稿词元来加速自回归推理。然而,随着上下文增长,现有最先进的草稿器变得日益昂贵,侵蚀了它们本应提供的效率优势。我们认为这种扩展是不必要的。独立的语言模型必须随其前缀增长,因为它对其生成的每一个词元负全部责任。相比之下,草稿器只提出候选;目标模型在任何词元被确认之前都会捕获并纠正每一个错误。因此,草稿器的解码成本可以做到完全独立于前缀长度。我们提出了 LongSpark,一种块扩散草稿器,它通过从目标模型的验证过程中提取固定尺寸的多尺度视图来实现这一点,从而消除了对不断增长的持久状态的需求。大量评估表明,LongSpark 在多种模型规模和真实服务条件下均实现了最先进的端到端效率。值得注意的是,它在长上下文任务中实现了最低的每输出词元时间,同时将草稿器的上下文状态减少了数个数量级。
cs.LG / 144 / 2609.37038
NowcastDiT: Diffusion Transformers are Effective Precipitation Nowcasters
NowcastDiT:扩散 Transformer 是有效的降水临近预报器
diffusion
扩散模型相关
Abstract
Precipitation nowcasting demands accurate short-term forecasts under strong spatiotemporal variability. Diffusion models are well suited to modeling complex precipitation distributions, yet existing approaches often introduce increasingly specialized designs, leaving the capability of a standard diffusion architecture underexplored. We show that a standard Diffusion Transformer already provides a simple and scalable foundation for precipitation nowcasting, with domain-specific requirements accommodated naturally within its design space. Based on this principle, we develop NowcastDiT and instantiate this flexibility through two complementary adaptations: a dynamics-aware noise prior for temporally coherent forecasts, and end-to-end reinforcement learning with timestep-aware rewards for meteorological skill. Experiments on SEVIR and MRMS benchmarks show that NowcastDiT achieves state-of-the-art performance in both perceptual quality and meteorological skill. These results suggest that standard DiT can serve as an effective foundation for precipitation nowcasting.
Chinese Translation
降水临近预报需要在强烈的时空变异性下进行准确的短期预报。扩散模型非常适合对复杂的降水分布进行建模,然而现有方法往往引入日益专门化的设计,使得标准扩散架构的能力尚未得到充分探索。我们表明,一个标准的扩散 Transformer 已经为降水临近预报提供了一个简单且可扩展的基础,领域特定的需求可以在其设计空间内自然地得到满足。基于这一原则,我们开发了 NowcastDiT,并通过两种互补的改造来具体实现这种灵活性:用于生成时间上连贯预报的动态感知噪声先验,以及用于提升气象技能的、带有时间步感知奖励的端到端强化学习。在 SEVIR 和 MRMS 基准上的实验表明,NowcastDiT 在感知质量和气象技能两方面均达到了最先进的性能。这些结果表明,标准 DiT 可以作为降水临近预报的有效基础。
cs.LG / 145 / 2609.37066
Beyond Compression: Diagnosing How Post-Training Changes Mathematical Reasoning
超越压缩:诊断后训练如何改变数学推理
large language model
大语言模型相关
Abstract
Post-training is central to mathematical reasoning in modern large language models (LLMs), but endpoint pass@1 alone underidentifies what has changed. Gains may reflect newly reachable solutions, cheaper sampling of latent solutions, surface robustness, or memorisation. We compare three post-training paths under a common diagnostic readout: our sufficiently trained off-policy distillation trajectories, released Qwen3 off-policy-plus-on-policy distillation endpoints, and a released DeepSeek-Math endpoint trained with Group Relative Policy Optimisation (GRPO). Our probe uses cross-surface pass@K over verbatim prompts, paraphrases, numerical isomorphisms, and translations, plus consistency, distribution-shape, and verified supervised-fine-tuning (SFT) membership analyses. We find two regimes. On easier AMC problems, large-K ceilings are near saturation, so post-training mainly compresses sample cost. On harder AIME problems, post-training expands the large-K ceiling over the base model: sufficient off-policy distillation already raises this ceiling, Qwen3 released endpoints raise it further, and DeepSeek-Math GRPO does not dominate sufficient off-policy distillation at large K. English-dominant distillation improves non-English reasoning but preserves language-tier gaps. A controlled-overfit audit finds limited sensitivity in current SFT-membership probes. Compression is one regime of post-training, not a universal explanation.
Chinese Translation
后训练在现代大型语言模型(LLMs)的数学推理中处于核心地位,但仅靠终点 pass@1 无法充分识别究竟发生了什么变化。收益可能反映新近可达的解、对潜在解的更低成本采样、表层鲁棒性或记忆化。我们在一个共同的诊断读出下比较三条后训练路径:我们充分训练的离策略蒸馏轨迹、已发布的 Qwen3 离策略加在策略蒸馏终点,以及一个已发布的、使用组相对策略优化(GRPO)训练的 DeepSeek-Math 终点。我们的探针使用跨表层 pass@K,覆盖逐字提示、复述、数值同构和翻译,并加上一致性、分布形状和经核验的监督微调(SFT)成员资格分析。我们发现两种机制。在较容易的 AMC 问题上,大 K 上限接近饱和,因此后训练主要压缩采样成本。在较难的 AIME 问题上,后训练相对于基础模型扩展了大 K 上限:充分的离策略蒸馏已经提高了这一上限,Qwen3 发布的终点进一步提高了它,而 DeepSeek-Math GRPO 在大 K 下并不优于充分的离策略蒸馏。英语主导的蒸馏改善了非英语推理,但保留了语言层级差距。一项受控过拟合审计发现当前 SFT 成员资格探针的敏感性有限。压缩是后训练的一种机制,而非一种普适解释。
cs.LG / 146 / 2609.37076
UnlearningSoup: Is Repeated Tuning Necessary for Large Language Model Unlearning?
UnlearningSoup:大语言模型遗忘中重复调参是必要的吗?
large language model
大语言模型相关
Abstract
Large language models trained on vast corpora inherently risk memorizing harmful content that may later re-emerge in their outputs. To mitigate this issue, existing unlearning methods typically rely on training-based parameter updates, such as gradient ascent and its variants, to delete targeted content while preserving other knowledge. However, balancing the competing goals of forgetting and retention makes hyperparameter choices for these methods particularly difficult, often requiring repeated tuning to obtain a strong model that still leaves substantial room for improvement and transfers poorly across models and datasets. To address this challenge, we investigate whether unlearning runs exhibit exploitable structure in weight space, and observe that models from different runs still lie in a shared evaluation-performance basin. This suggests that stronger models may be recovered through an unlearning-tailored soup strategy, reducing the need for repeated tuning for further improvement or new settings. Motivated by this, we propose UnlearningSoup, a unified framework that provides two strategies: EfficientSoup uses binary-search-based interpolation to quickly discover a well-performing model in the early stage, where repeated tuning would otherwise make strong model selection costly. PerformanceSoup uses reweighted souping to efficiently unlock the remaining performance potential in the later stage, where repeated tuning becomes increasingly inefficient. Extensive experiments across diverse datasets and models show that UnlearningSoup delivers 2.4x to 3.3x efficiency gains in hyperparameter selection, while consistently improving performance across settings.
Chinese Translation
在大规模语料库上训练的大语言模型本身存在记忆有害内容的风险,这些内容之后可能重新出现在其输出中。为缓解这一问题,现有的遗忘方法通常依赖于基于训练的参数更新,例如梯度上升及其变体,以删除目标内容,同时保留其他知识。然而,平衡遗忘与保留这两个相互竞争的目标,使得这些方法的超参数选择尤为困难,往往需要重复调参才能获得一个强模型,而该模型仍有很大的改进空间,并且在不同模型和数据集之间迁移效果很差。为应对这一挑战,我们研究遗忘运行是否在权重空间中表现出可利用的结构,并观察到来自不同运行的模型仍位于一个共享的评估性能盆地中。这表明,可能通过一种为遗忘量身定制的模型汤策略来恢复更强的模型,从而减少为进一步改进或适应新设置而重复调参的需求。受此启发,我们提出 UnlearningSoup,一个统一框架,提供两种策略:EfficientSoup 使用基于二分搜索的插值,在早期阶段快速发现一个性能良好的模型;在此阶段,若不采用该方法,重复调参会使强模型选择变得代价高昂。PerformanceSoup 使用重加权模型汤来高效释放后期阶段剩余的性能潜力;在后期阶段,重复调参变得日益低效。在不同数据集和模型上进行的大量实验表明,UnlearningSoup 在超参数选择上带来 2.4 倍到 3.3 倍的效率提升,同时在各种设置下持续提升性能。
cs.LG / 147 / 2609.37078
ZeroDiff: Zero-Shot Time Series Reconstruction via Informed-Prior Diffusion
ZeroDiff:基于信息先验扩散的零样本时间序列重建
diffusion
扩散模型相关
Abstract
Time series modeling increasingly demands high-quality supervision, yet target observations remain scarce - exogenous inputs are broadly available, but target measurements are often unavailable due to cost, infrastructure, or accessibility constraints. Can models trained on observed locations reconstruct target time series where measurements have never been collected? We term this zero-shot time series reconstruction. A naive approach - directly mapping exogenous inputs to targets - can yield predictions at unobserved locations, but without target signals, such models fail to capture the intrinsic dynamics of the target variable, producing overly smooth outputs that underestimate extremes. This reveals systematic errors that call for explicit modeling and calibration. We propose ZeroDiff, which constructs an informed prior from exogenous variables alone, then learns to calibrate reconstruction errors through diffusion - training on observed locations and generalizing to unobserved ones. Experiments across diverse real-world datasets demonstrate significant improvements over existing approaches. Our code is available at https://github.com/YingdaFan/ZeroDiff-ICML2026.
Chinese Translation
时间序列建模日益需要高质量的监督,但目标观测仍然稀缺——外生输入广泛可得,而由于成本、基础设施或可访问性限制,目标测量值往往不可得。在已观测位置上训练的模型能否重建从未采集过测量值的目标时间序列?我们将此称为零样本时间序列重建。一种朴素方法——直接将外生输入映射到目标——可以在未观测位置产生预测,但在没有目标信号的情况下,此类模型无法捕捉目标变量的内在动态,产生过于平滑且低估极端值的输出。这揭示了需要显式建模与校准的系统性误差。我们提出 ZeroDiff,它仅从外生变量构建信息先验,然后通过扩散学习校准重建误差——在已观测位置上训练,并泛化到未观测位置。在多样化的真实世界数据集上的实验表明,相较于现有方法有显著提升。我们的代码可在 https://github.com/YingdaFan/ZeroDiff-ICML2026 获取。
cs.LG / 148 / 2609.37113
Neural Constitutive Learning for Generalized Reaction-Diffusion Systems
广义反应-扩散系统的神经本构学习
diffusion
扩散模型相关
Abstract
Generalized reaction-diffusion systems encompass diverse transport mechanisms and coupled reaction kinetics. A central question for neural PDE solvers is what should be learned so that a common interface can accommodate phase-field and degenerate transport, local reactions, and multispecies coupling. We propose the Neural Constitutive Laws--Mass-Compression-Transport (NCL-MCT) Solver, which learns PDE-specific constitutive responses while retaining temporal evolution in a shared MCT integrator. Transport is represented through mobility and thermodynamic driving force, and reaction through relative reaction rates. These constitutive responses depend on the current density rather than explicitly on the initial condition or elapsed time, motivating their reuse across different initial conditions and time horizons. The same interface supports velocity-data supervision and known-law supervision, neither of which requires time integration during training. When constitutive laws are known, supervision can be evaluated on independently sampled density fields, enabling trajectory-free constitutive learning without generating solution trajectories. Across seven systems, separately trained constitutive modules share the same interface and MCT integrator and achieve relative rollout $L^2$ errors of $10^{-4}$ to $10^{-2}$. Tests with unseen initial-condition families and an extended time horizon assess reuse beyond training conditions, while separate experiments demonstrate trajectory-free constitutive learning. These results support constitutive responses as an effective learning target for a shared neural PDE framework.
Chinese Translation
广义反应-扩散系统涵盖多样的传输机制和耦合反应动力学。对于神经 PDE 求解器而言,一个核心问题是应当学习什么,才能使一个通用接口能够容纳相场和退化输运、局部反应以及多物种耦合。我们提出神经本构定律--质量-压缩-输运(NCL-MCT)求解器,该求解器在共享的 MCT 积分器中保留时间演化,同时学习特定于 PDE 的本构响应。输运通过迁移率和热力学驱动力表示,反应则通过相对反应速率表示。这些本构响应依赖于当前密度,而不是显式依赖于初始条件或已流逝时间,这促使它们能够在不同初始条件和时间范围之间重用。同一个接口支持速度数据监督和已知定律监督,二者都不需要在训练期间进行时间积分。当本构定律已知时,监督可以在独立采样的密度场上进行评估,从而实现无需生成解轨迹的无轨迹本构学习。在七个系统上,分别训练的本构模块共享相同的接口和 MCT 积分器,并达到 $10^{-4}$ 到 $10^{-2}$ 的相对 rollout $L^2$ 误差。使用未见过的初始条件族和延长的时间范围进行的测试评估了超出训练条件的重用,而单独的实验展示了无轨迹本构学习。这些结果支持将本构响应作为共享神经 PDE 框架的有效学习目标。
cs.LG / 149 / 2609.37119
Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training
解锁评论家:面向 LLM 后训练的无奖励策略优化
large language model
大语言模型相关
Abstract
Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instability and memory overhead. Even where a critic is trained, it is discarded once training ends, although it has learned to predict outcomes. We revisit this trend and show that a pretrained critic's ability to predict future outcomes can make it a valuable asset for efficient long-horizon reasoning. First, we find that instability in critic-based RL for long chain-of-thought reasoning is largely an optimization artifact: keeping policy updates small and low in variance restores stable convergence. Second, a well-pretrained critic estimates the posterior probability of eventual success from later trajectory states and unfinished prefixes. Its predictions provide outcome-derived, dense, per-prefix learning signals that, during policy optimization, require neither completed rollouts, step-level annotations, nor external reward labels. Building on this insight, we introduce Reward-Free Policy Optimization (RFPO), which repurposes a single calibrated, frozen critic as a rollout-level reward, a value baseline for generalized advantage estimation, and a success forecaster for unfinished prefixes. We further show that binarizing the debiased score stops the policy from exploiting the critic's length bias. Binarized, RFPO matches supervised PPO without a single label in the training loop, while cutting compute and memory overhead. This makes RFPO well suited to long-horizon reasoning tasks, where outcomes arrive late and generation dominates cost: because rollouts can be rewarded before they finish, training no longer has to pay for waiting on every trajectory to complete. Our findings challenge the prevailing critic-free paradigm and establish critic-based, reward-free optimization as a scalable and computationally efficient path for LLM post-training.
Chinese Translation
近期面向大语言模型的强化学习(RL)后训练方法越来越多地移除评论家(critic),以降低训练不稳定性与内存开销。即便在训练了评论家的情况下,它也会在训练结束后被丢弃,尽管它已经学会了预测结果。我们重新审视这一趋势,并表明:一个预训练评论家预测未来结果的能力,可以使其成为高效长时程推理的宝贵资产。首先,我们发现,面向长思维链推理的基于评论家的 RL 中的不稳定性在很大程度上是一种优化假象:保持策略更新幅度小且方差低,即可恢复稳定收敛。其次,一个预训练良好的评论家能够从轨迹的较后状态以及未完成的前缀中估计最终成功的后验概率。它的预测提供了源自结果的、稠密的、逐前缀的学习信号,在策略优化过程中,这些信号既不需要完整的 rollout,也不需要步级标注,更不需要外部奖励标签。基于这一洞见,我们提出无奖励策略优化(Reward-Free Policy Optimization,RFPO),它将单个经过校准的、冻结的评论家重新用作 rollout 级奖励、用于广义优势估计的价值基线,以及面向未完成前缀的成功预测器。我们进一步表明,对去偏后的分数进行二值化,可以阻止策略利用评论家的长度偏差。经二值化后,RFPO 在训练循环中无需任何标签即可与监督式 PPO 相媲美,同时削减了计算与内存开销。这使 RFPO 非常适合长时程推理任务——在这类任务中,结果来得晚,且生成主导了成本:由于 rollout 可以在完成之前就获得奖励,训练不再需要为等待每条轨迹完成而付出代价。我们的发现挑战了当前主流的无评论家范式,并确立了基于评论家的、无奖励的优化作为 LLM 后训练的一条可扩展且计算高效的路径。
cs.LG / 150 / 2609.37169
Trajectory Soup: Pushing the Compute-Scaling Frontier of LLM Mid-training via Diverse Trajectories
轨迹汤:通过多样化轨迹推动大语言模型中期训练的计算扩展前沿
large language model
大语言模型相关
Abstract
Mid-training equips pretrained large language models with specialized and reasoning capabilities, but the returns of this stage are bounded since additional serial compute yields little further downstream improvement and can even degrade some capabilities, which places a practical ceiling on how much compute mid-training absorbs. We revisit how this compute should be allocated to a single run or multiple similar optimizations. We find that branches forked from a shared checkpoint under various controlled recipe reaches measurably different regions of parameter space, and establish a form of compatible diversity that extending one run cannot supply. Therefore, we introduce Trajectory Soup, which distributes a mid-training budget over several independent branches, and consolidates strongest checkpoints selected on validation through intra- and inter-trajectory averaging into a single model. A local bias and variance analysis separates the two averaging levels, showing that inter-trajectory averaging removes residual error beyond the reach of averaging within a trajectory, while checkpoint selection carries a bias that bounds how many checkpoints are worth merging. Across model scales, learning-rate schedules, token budgets, and trajectory counts, Trajectory Soup improves aggregate downstream performance over the strongest single-trajectory average under matched budgets and keeps improving as budgets expand, with the advantage preserved after an identical post-training pipeline. These results position trajectory allocation and merging as a practical way to extend the compute-scaling frontier of mid-training beyond serial saturation.
Chinese Translation
中期训练为预训练大语言模型赋予专业化与推理能力,但这一阶段的回报是有界的,因为额外的串行计算几乎不能带来进一步的下游提升,甚至可能使某些能力退化,这为中期训练所能吸收的计算量设定了实际上限。我们重新审视了应当如何将这些计算分配给单次运行或多次相似的优化。我们发现,在多种受控配方下从共享检查点分叉出的分支会到达参数空间中可测量的不同区域,并确立了一种扩展单次运行所无法提供的兼容多样性。因此,我们提出了轨迹汤(Trajectory Soup),它将中期训练预算分配到若干独立分支上,并通过轨迹内平均与轨迹间平均,将在验证集上选出的最强检查点整合为一个单一模型。一项局部偏差与方差分析区分了这两个平均层级,表明轨迹间平均能够消除超出单条轨迹内平均能力范围的残余误差,而检查点选择则带有一种偏差,该偏差限定了有多少检查点值得合并。在不同模型规模、学习率调度、token 预算和轨迹数量下,Trajectory Soup 在预算匹配时相较于最强的单轨迹平均提升了聚合下游性能,并且随着预算扩大而持续改进,在经历相同的后训练流程后这一优势依然保持。这些结果将轨迹分配与合并定位为一种切实可行的方法,用以将中期训练的计算扩展前沿推进到超越串行饱和的程度。
cs.LG / 151 / 2609.37227
Interacting particle guidance for sampling reward-tilted generative priors
用于采样奖励倾斜生成先验的相互作用粒子引导
diffusion
扩散模型相关
Abstract
Inference-time steering adapts pretrained diffusion and flow-based models to new tasks, e.g., to generate samples from a conditional distribution or samples with desired properties, without retraining. This can be formalized as sampling from a reward-tilted generative prior. As exact sampling from this distribution is intractable, guidance-based methods rely on approximations producing biased samples, and sequential Monte Carlo (SMC) methods correct for this bias using importance weights. However, while exact in the large particle limit, SMC suffers from weight degeneracy and particle collapse in practice. We propose interacting particle guidance (IPG), which replaces reweighting with transport. The particles interact through an additional drift, derived from the Feynman--Kac PDE to cancel the reweighting term, and remain unweighted. Choosing the drift in a reproducing kernel Hilbert space yields a closed-form solution that is cheap to compute, with negligible overhead compared to SMC. We demonstrate the method on Gaussian mixtures with known posteriors, and on high-dimensional image inpainting and protein structure inference tasks.
Chinese Translation
推理时引导使预训练的扩散模型和基于流的模型适应新任务,例如,在无需重新训练的情况下,从条件分布中生成样本或生成具有所需属性的样本。这可以被形式化为从奖励倾斜的生成先验中采样。由于从该分布中进行精确采样是难处理的,基于引导的方法依赖于会产生有偏样本的近似,而序贯蒙特卡洛(SMC)方法则使用重要性权重来校正这一偏差。然而,尽管 SMC 在大粒子极限下是精确的,但在实践中它会遭受权重退化和粒子坍缩的问题。我们提出了相互作用粒子引导(IPG),它用输运替代重加权。粒子通过一个额外的漂移项相互作用,该漂移项由 Feynman--Kac 偏微分方程导出,用以抵消重加权项,并且粒子保持无权重。在再生核希尔伯特空间中选择该漂移项可得到一个计算成本低廉的闭式解,与 SMC 相比其额外开销可忽略不计。我们在具有已知后验的高斯混合模型上,以及在高维图像修复和蛋白质结构推断任务上展示了该方法。
cs.LG / 152 / 2609.37323
Parallel Tempering for Diffusion-Based Combinatorial Optimization
用于基于扩散的组合优化的并行回火
diffusion
扩散模型相关
Abstract
Discrete diffusion models have emerged as a powerful paradigm for solving combinatorial optimization (CO) problems on graphs by learning to sample high-quality solutions. A common inference-time approach is to generate multiple candidate solutions independently and return the best-performing sample, improving solution quality at the expense of an increase in computational cost. In this work, we introduce PT-Denoise, an inference-time procedure that allows these concurrent denoising trajectories to interact through parallel tempering, without requiring retraining or fine-tuning of the underlying denoiser. Our method assigns a temperature to each diffusion process and allows processes to swap temperatures based on their relative performance. This dynamically reallocates promising, low-energy trajectories to colder, more concentrated sampling regimes while allowing higher-energy states to escape local minima through randomized exploration. Experiments on canonical graph-structured CO problems show that our approach consistently improves the quality of the best solution found, while only adding minimal computational overhead.
Chinese Translation
离散扩散模型已经出现作为一种强大的范式,用于通过学习采样高质量解来解决图上的组合优化(CO)问题。一种常见的推理时方法是独立生成多个候选解并返回表现最佳的样本,从而以增加计算成本为代价提高解的质量。在这项工作中,我们引入 PT-Denoise,这是一种推理时过程,使这些并发的去噪轨迹能够通过并行回火进行交互,而无需重新训练或微调底层的去噪器。我们的方法为每个扩散过程分配一个温度,并允许过程基于它们的相对性能交换温度。这会动态地将有前景的低能量轨迹重新分配到更冷、更集中的采样机制,同时允许更高能量的状态通过随机探索逃离局部最小值。在典型图结构 CO 问题上的实验表明,我们的方法持续改进了所找到的最佳解的质量,同时仅增加最小的计算开销。
cs.LG / 153 / 2609.37391
Rethinking Soft Tokens for Parallel Decoding in Diffusion Language Models
重新思考扩散语言模型中用于并行解码的软令牌
diffusion
扩散模型相关
Abstract
Diffusion language models (DLMs) enable parallel generation by predicting and committing multiple tokens at each denoising step, yet they can generate individually plausible but mutually inconsistent tokens. Recent work shows that \emph{soft tokens} can mitigate this issue by representing uncertain positions with continuous embeddings built from the model's predictive distribution at the previous decoding step. However, although soft tokens are commonly understood as preserving predictive uncertainty, how soft-token feedback improves parallel decoding has not been systematically examined. In this paper, we investigate this question in frozen pretrained DLMs to examine soft-token feedback without the effects of additional training. To construct soft-token inputs in a training-free setting, we identify a geometric mismatch between conventional soft-token construction and the pretrained embedding space. Based on this observation, we propose a training-free, geometry-aware construction of soft tokens. Our analysis of soft-token feedback suggests that uncertainty preservation alone does not fully explain how it reshapes subsequent predictions. To better explain how soft-token feedback improves parallel decoding, we provide empirical evidence that it favors coherent token sequences. Across four pretrained DLMs and four math and code benchmarks, our method outperforms standard parallel decoding and a training-free Euclidean soft-token baseline. Code: https://github.com/kodaikawamura/rethinking-soft-tokens
Chinese Translation
扩散语言模型(DLMs)通过在每一步去噪时预测并提交多个 token 来实现并行生成,但它们可能生成单独看似合理却彼此不一致的 token。近期工作表明,软令牌可以通过用由上一步解码时模型的预测分布构建的连续嵌入来表示不确定位置,从而缓解这一问题。然而,尽管软令牌通常被理解为保留了预测不确定性,但软令牌反馈如何改进并行解码尚未得到系统性的考察。在本文中,我们在冻结的预训练 DLM 中研究这一问题,以便在不受到额外训练影响的情况下考察软令牌反馈。为了在免训练设定下构建软令牌输入,我们识别出传统软令牌构建方式与预训练嵌入空间之间的几何不匹配。基于这一观察,我们提出了一种免训练、几何感知的软令牌构建方法。我们对软令牌反馈的分析表明,仅靠不确定性保留并不能完全解释它如何重塑后续预测。为了更好地解释软令牌反馈如何改进并行解码,我们提供了经验证据,表明它偏好连贯的 token 序列。在四个预训练 DLM 以及四个数学与代码基准上,我们的方法优于标准并行解码以及一个免训练的欧氏软令牌基线。代码:https://github.com/kodaikawamura/rethinking-soft-tokens
cs.LG / 154 / 2609.37587
ReLMem: Learning Recurrent Memory for Longitudinal EHR Modeling
ReLMem:面向纵向 EHR 建模的学习式循环记忆
large language model
大语言模型相关
Abstract
Longitudinal electronic health record (EHR) modeling requires integrating new visits with an expanding patient history. Yet the continual accumulation of clinical information imposes increasing computational and memory costs on large language models (LLMs) when they process and retain complete patient histories. A practical alternative is visit-wise recurrent compression, which incorporates each incoming visit into a compact, continually updated patient memory. However, under a fixed memory budget, successive updates must integrate new information without progressively losing critical historical evidence needed to subsequent tasks. To address this challenge, we introduce Recurrent Longitudinal Memory (ReLMem), a framework that learns to maintain fixed-capacity patient memory for efficient downstream prediction with a frozen LLM. ReLMem equips this LLM with lightweight compression adapters to recurrently update the memory from its previous state and each incoming visit, without rereading earlier records. Specifically, we develop a multi-granularity optimization strategy to preserve task-relevant information throughout recurrent updates and support downstream prediction from the final memory. The intermediate supervision aligns attention outputs from compressed memory and the full history under identical queries, while prediction supervision minimizes cross-entropy with ground truth answers conditioned on the final memory. On EHR-based medication prediction, ReLMem approaches the F1 scores of full-history baseline while reducing average retained historical storage by 97.1%. Under the same memory budget, it improves macro- and micro-F1 over the strongest compressed-memory baseline by 4.66 and 4.75 percentage points, respectively. These results highlight the value of learning recurrent patient memory for efficient longitudinal EHR modeling.
Chinese Translation
纵向电子健康记录(EHR)建模需要将新的就诊记录与不断扩展的患者病史整合起来。然而,临床信息的持续累积使大语言模型(LLM)在处理并保留完整患者病史时,承受着不断增长的计算与内存开销。一种实用的替代方案是按就诊逐次进行的循环压缩,它将每一次新到的就诊记录纳入一个紧凑且持续更新的患者记忆中。然而,在固定的内存预算下,连续的更新必须在整合新信息的同时,不逐渐丢失后续任务所需的关键历史证据。为应对这一挑战,我们提出了循环纵向记忆(Recurrent Longitudinal Memory,ReLMem),这是一个学习维护固定容量患者记忆的框架,可与冻结的 LLM 配合实现高效的下游预测。ReLMem 为该 LLM 配备轻量级压缩适配器,以根据其先前状态与每一次新到的就诊记录循环地更新记忆,而无需重新读取更早的记录。具体而言,我们开发了一种多粒度优化策略,以在循环更新过程中保留与任务相关的信息,并支持基于最终记忆进行下游预测。中间监督在相同查询下对齐来自压缩记忆与完整历史的注意力输出,而预测监督则以最终记忆为条件,最小化与真实答案之间的交叉熵。在基于 EHR 的用药预测任务上,ReLMem 接近了完整病史基线的 F1 分数,同时将平均保留的历史存储量减少了 97.1%。在相同内存预算下,相较于最强的压缩记忆基线,其宏平均 F1 和微平均 F1 分别提升了 4.66 和 4.75 个百分点。这些结果凸显了学习循环患者记忆对于高效纵向 EHR 建模的价值。
cs.LG / 155 / 2609.37632
ProCTI: Prototype-Refined Global Conditioning for Diffusion-Based Time Series Imputation
ProCTI:用于基于扩散的时间序列插补的原型精炼全局条件化
diffusion
扩散模型相关
Abstract
Time series imputation has progressed from statistical and deep learning approaches to diffusion-based models, which have shown strong recent performance. Existing diffusion-based methods typically condition the reverse process using local contextual information from the current or neighbouring windows. Meanwhile, global dataset-level structure often remains implicit, limiting performance when local observations are sparse, noisy, or unrepresentative. To address this issue, we propose ProCTI, a diffusion-imputation framework that augments local conditioning with retrieved global dataset-level priors through learned prototypes. A hybrid conditioning mechanism integrates this global context with local signals during reverse diffusion, enabling more accurate reconstruction under varying missingness scenarios. Experiments across multiple benchmark datasets show that ProCTI outperforms strong baselines overall under random missingness, while remaining competitive under attribute-wise missingness. Furthermore, we use a latent-regime data model to characterise the precise conditions under which prototype-derived global conditioning provably improves imputation. We support this with a general theoretical analysis of local-global conditioning.
Chinese Translation
时间序列插补已从统计方法和深度学习方法发展到基于扩散的模型,后者近来展现出强劲的性能。现有基于扩散的方法通常使用来自当前窗口或相邻窗口的局部上下文信息来对逆向过程进行条件化。同时,数据集级别的全局结构往往仍是隐式的,当局部观测稀疏、噪声较大或不具代表性时,这会限制性能。为解决这一问题,我们提出 ProCTI,一种扩散插补框架,它通过学习到的原型,用检索到的全局数据集级先验来增强局部条件化。一种混合条件化机制在逆向扩散过程中将这种全局上下文与局部信号相结合,从而能够在不同的缺失场景下实现更准确的重建。在多个基准数据集上的实验表明,ProCTI 在随机缺失下总体上优于强基线,同时在按属性缺失下仍保持竞争力。此外,我们使用潜在区制数据模型来刻画原型导出的全局条件化可证明地改进插补的精确条件。我们通过关于局部-全局条件化的一般理论分析来支持这一点。
cs.LG / 156 / 2609.37694
GARDiff: Graph-Aligned Residual Diffusion for Probabilistic Multivariate Time-Series Forecasting
GARDiff:用于概率多变量时间序列预测的图对齐残差扩散
diffusion
扩散模型相关
Abstract
Diffusion models have recently shown strong potential for probabilistic multivariate time-series forecasting by modeling complex conditional distributions. Recent decoupled diffusion frameworks further separate forecasting into deterministic prediction and stochastic residual generation, making it natural to derive dependency graphs from deterministic representations and use them to guide residual diffusion. However, we show that this direct structural transfer is unreliable. Although deterministic-derived graphs encode useful global dependency priors, they exhibit substantial edge-level misalignment with residual dependency structures, introducing inaccurate or redundant conditions during residual generation. This reveals a previously overlooked deterministic-to-residual structural alignment problem in decoupled diffusion forecasting. To address this problem, we propose GARDiff, a Graph-Aligned Residual Diffusion framework for probabilistic multivariate time-series forecasting. Instead of treating deterministic-derived graphs as fixed diffusion conditions, GARDiff progressively adapts them to residual generation. Specifically, GARDiff estimates residual uncertainty to distinguish high- and low-uncertainty regions, enabling uncertainty-aware structural refinement, and further performs timestep-aware edge sparsification during reverse diffusion to evolve graph conditions from broad dependency aggregation to localized residual refinement. Extensive experiments on six real-world benchmarks demonstrate that GARDiff consistently improves probabilistic forecasting performance and uncertainty calibration over strong baselines.
Chinese Translation
扩散模型近来通过建模复杂的条件分布,在概率多变量时间序列预测中展现出强大潜力。近期的解耦扩散框架进一步将预测分解为确定性预测和随机残差生成,这使得从确定性表示中推导依赖图并用其指导残差扩散变得自然。然而,我们表明这种直接的结构迁移并不可靠。尽管由确定性表示推导出的图编码了有用的全局依赖先验,但它们与残差依赖结构存在显著的边级错位,从而在残差生成过程中引入不准确或冗余的条件。这揭示了解耦扩散预测中一个此前被忽视的从确定性到残差的结构对齐问题。为解决这一问题,我们提出 GARDiff,一个用于概率多变量时间序列预测的图对齐残差扩散框架。GARDiff 不将确定性推导出的图视为固定的扩散条件,而是逐步使其适应残差生成。具体而言,GARDiff 估计残差不确定性以区分高不确定性和低不确定性区域,从而实现不确定性感知的结构细化,并进一步在反向扩散过程中执行时间步感知的边稀疏化,使图条件从广泛的依赖聚合演变为局部化的残差细化。在六个真实世界基准上的大量实验表明,GARDiff 在强基线上持续提升了概率预测性能和不确定性校准。
cs.LG / 157 / 2609.37731
Weights Read and Write Features: Scalable Parameter Decomposition Grounded in Activation Space
权重读取与写入特征:扎根于激活空间的可扩展参数分解
large language model
大语言模型相关
Abstract
Activation space and parameter space provide complementary views of model computation. Activations represent information, while weights read, transform, and write that information. Yet existing interpretability methods largely study the two spaces separately, leaving the connection between represented information and parameter-level computation underexplored. We introduce Activation-Supported Parameter Decomposition (ASPD), which jointly decomposes activation and parameter spaces and grounds each learned weight component in the activation features it reads or writes. This grounding constrains otherwise non-unique parameter decompositions using the model's internal activations, while an internal reconstruction objective provides a local learning signal at the weight matrix being analyzed. Together, these properties enable scalable, interpretable, and causally editable parameter decomposition in pretrained large language models, demonstrated on Qwen-3-8B. The learned read--write components can also be composed into parameter-level mechanism circuits. We use ASPD to recover mechanisms underlying the classic IOI circuit and trace semantic transformations through model weights.
Chinese Translation
激活空间和参数空间提供了模型计算的互补视角。激活表示信息,而权重读取、变换并写入该信息。然而,现有的可解释性方法在很大程度上分别研究这两个空间,使得被表示的信息与参数级计算之间的联系仍未得到充分探索。我们引入激活支持的参数分解(Activation-Supported Parameter Decomposition,ASPD),它联合分解激活空间和参数空间,并将每个学习到的权重组件扎根于它所读取或写入的激活特征中。这种扎根利用模型内部激活来约束原本不唯一的参数分解,同时一个内部重构目标在所分析的权重矩阵处提供局部学习信号。这些特性共同使得在预训练大语言模型中实现可扩展、可解释且可因果编辑的参数分解成为可能,并在 Qwen-3-8B 上得到展示。学习到的读取--写入组件还可以组合成参数级机制回路。我们使用 ASPD 恢复经典 IOI 回路背后的机制,并通过模型权重追踪语义变换。
cs.LG / 158 / 2609.37745
Optimizer-dependent training dynamics converge to the same one-third optimal data scaling
优化器依赖的训练动力学收敛到相同的三分之一最优数据缩放
large language model
大语言模型相关
Abstract
Neural scaling, in which loss falls as a power law with training, is central to large language models, and one recent proposal is that a $1/3$ exponent emerges from learning peaked distributions. That account describes SGD, but models in practice are trained with adaptive optimizers. Here we separate two exponents the $1/3$ account does not distinguish: how fast the loss falls with training steps along a single run, and how fast the optimally tuned loss falls with dataset size $D$. We show that the first, a dynamic exponent, is optimizer-specific while the second, an optimal data exponent, converges to $1/3$ across optimizers. In an online teacher-student model we decompose the loss into norm growth (radial) and alignment toward the teacher direction (tangential), each decaying as a power law with dynamic exponents $α_{r}$ and $α_{t}$. Under SGD, both are close to $1/3$, so the data exponent is also $1/3$ across different learning rates. Under Adam the two separate: $α_{r} \simeq 0.48$ but $α_{t} \simeq 0.08$. Since the total loss is minimized when these two parts are balanced, the optimal learning rate is optimizer-dependent: $D$-independent for SGD but falls with $D$ for Adam. Yet tuned to that optimum, the loss returns to $D^{-1/3}$ for both. A stochastic-dynamics analysis explains why: the optimizers can trade decay speed between the two channels, but they all fall on a single dynamic exponent relation, $2α_{r}+ α_{t} = 1$, which fixes the optimal data exponent at $1/3$. Across seven optimizers, including Muon, the measured exponents are consistent with this relation, and the optimal-loss envelopes agree with $D^{-1/3}$ across them. The optimizer sets how fast a model learns per step; tuned optimally, it changes the prefactor but not the rate at which loss falls per sample.
Chinese Translation
神经标度(其中损失随训练按幂律下降)对于大语言模型至关重要,而最近的一项提议是:$1/3$ 指数源于学习尖峰分布。该解释描述的是 SGD,但实际中的模型是用自适应优化器训练的。这里我们区分 $1/3$ 解释未区分的两个指数:在单次运行中损失随训练步数下降的速度,以及最优调参后的损失随数据集大小 $D$ 下降的速度。我们表明,第一个指数(动态指数)是优化器特定的,而第二个指数(最优数据指数)在不同优化器之间收敛到 $1/3$。在一个在线教师-学生模型中,我们将损失分解为范数增长(径向)和朝向教师方向的对齐(切向),二者各自以幂律衰减,动态指数分别为 $α_{r}$ 和 $α_{t}$。在 SGD 下,二者都接近 $1/3$,因此在不同学习率下数据指数也是 $1/3$。在 Adam 下,二者分离:$α_{r} \simeq 0.48$ 但 $α_{t} \simeq 0.08$。由于当这两部分平衡时总损失最小,最优学习率依赖于优化器:对 SGD 而言与 $D$ 无关,但对 Adam 而言随 $D$ 下降。然而,调到该最优值后,二者的损失都回到 $D^{-1/3}$。随机动力学分析解释了原因:优化器可以在两个通道之间交换衰减速度,但它们都落在同一个动态指数关系 $2α_{r}+ α_{t} = 1$ 上,这将该最优数据指数固定在 $1/3$。在包括 Muon 在内的七种优化器上,测得的指数与该关系一致,而且它们的最优损失包络都与 $D^{-1/3}$ 相符。优化器决定了模型每一步学得多快;在最优调参后,它改变的是前因子,而不是损失随样本下降的速率。
cs.LG / 159 / 2609.37825
Privy to the Foil: Recasting Value Estimation with a Self-Privileged Critic for RLVR
洞悉衬托:以自我特权评论家重塑 RLVR 的价值估计
large language model
大语言模型相关
Abstract
Assigning credit to intermediate steps remains a central challenge in training Large Language Models (LLMs) on multi-step reasoning tasks with sparse terminal rewards, and actor-critic methods such as PPO address this by learning value functions to construct token-level advantages. Their effectiveness, however, hinges on reliable value estimation, a difficult task requiring the critic to both assess progress toward a correct solution and anticipate an evolving policy's future behavior; errors in either can compromise credit assignment and destabilize online training. In this paper, we revisit the standard state-only formulation of value estimation and propose $π$PPO, a self-privileged actor-critic framework. By reusing verified same-prompt rollouts as contrastive evidence, $π$PPO helps the critic assess intermediate reasoning against successful and failed attempts, while preserving standard policy optimization and the deployment interface. Experiments show that $π$PPO consistently improves value-estimation quality by a substantial margin and outperforms representative actor-critic and critic-free RLVR baselines on challenging mathematical reasoning benchmarks, while remaining effective even when paired with substantially smaller asymmetric critics.
Chinese Translation
在具有稀疏终端奖励的多步推理任务上训练大型语言模型(LLMs)时,为中间步骤分配信用仍然是一项核心挑战,而诸如 PPO 之类的 actor-critic 方法通过学习价值函数以构建 token 级优势来解决这一问题。然而,它们的有效性取决于可靠的价值估计,这是一项困难的任务,要求评论家既能评估朝正确解迈进的进展,又能预判不断演化的策略的未来行为;二者中任一出现错误,都可能损害信用分配并使在线训练失稳。本文重新审视价值估计的标准仅状态形式,并提出 $π$PPO,一个自我特权化的 actor-critic 框架。通过将经过验证的同提示采样轨迹复用为对比证据,$π$PPO 帮助评论家参照成功与失败的尝试来评估中间推理,同时保持标准策略优化与部署接口不变。实验表明,$π$PPO 持续以显著幅度提升价值估计质量,并在具有挑战性的数学推理基准上优于代表性的 actor-critic 与无评论家 RLVR 基线,而且即使与规模小得多的非对称评论家搭配也依然有效。
cs.LG / 160 / 2609.37842
Scaling Influence Functions in LLMs through Eigenbasis-Corrected One-Bit Gradient Projection
通过特征基校正的单比特梯度投影扩展大语言模型中的影响函数
large language model
大语言模型相关
Abstract
Influence functions estimate how individual training examples affect the behavior of large language models (LLMs). Analyzing how training data influence different behaviors of an LLM involves repeated influence computation. Reusing stored training gradients reduces the computational cost, but storing full gradients is prohibitively expensive at LLM scale. We study how to compress these gradients while preserving influence estimates for future queries that are unknown at storage time. Through a worst-case analysis, we characterize the optimal fixed-dimensional linear representation and propose eigenbasis-corrected one-bit gradient projection (EOGP) to approximate it at scale. Specifically, EOGP uses EK-FAC to reduce gradient dimensionality, then applies PCA within the retained subspace to learn compression directions from the training gradients. We then apply one-bit quantization to the resulting coordinates, allowing more coordinates to be retained within a fixed storage budget. On GPT-2, EOGP predicts retraining outcomes more accurately than the evaluated compression baselines while using one-sixteenth of their per-example storage. On OLMo 2 SFT models from 1B to 32B parameters, EOGP remains competitive with the baselines allocated over 100 times as much storage per example.
Chinese Translation
影响函数估计单个训练样本如何影响大语言模型(LLMs)的行为。分析训练数据如何影响 LLM 的不同行为涉及重复的影响计算。重用存储的训练梯度可以降低计算成本,但在 LLM 规模下存储完整梯度昂贵到不可行。我们研究如何压缩这些梯度,同时保留对存储时未知的未来查询的影响估计。通过最坏情况分析,我们刻画了最优的固定维线性表示,并提出特征基校正的单比特梯度投影(EOGP)以大规模地近似它。具体而言,EOGP 使用 EK-FAC 来降低梯度维度,然后在保留的子空间内应用 PCA,以从训练梯度中学习压缩方向。然后,我们对得到的坐标应用单比特量化,从而允许在固定存储预算内保留更多坐标。在 GPT-2 上,EOGP 比所评估的压缩基线更准确地预测再训练结果,同时仅使用它们每个样本存储量的十六分之一。在从 1B 到 32B 参数的 OLMo 2 SFT 模型上,EOGP 与每个样本被分配超过 100 倍存储量的基线相比仍具有竞争力。
cs.LG / 161 / 2609.37852
Delta-Matching: Closing the Final Gap of Native 8-bit Training for LLMs
Delta-Matching:弥合大语言模型原生 8 位训练的最终差距
large language model
大语言模型相关
Abstract
Reliable FP8 attention remains a barrier to fully native 8-bit large language model training. We derive how forward-backward inconsistencies produce stale delta and empirically show how it distorts training dynamics. Our stale-delta hybrid runs show a modest loss gap at 569M parameters but substantial loss increases and downstream degradation at 1.67B and 5.29B. QK normalization, NoPE (no positional encoding), and lower-learning-rate context extension mitigate or delay degradation without eliminating it. This pattern suggests accumulated optimization error that smaller models and short runs can conceal. We propose Delta-Matching, proving that it restores the softmax gradient's zero-row-sum invariant under the stated numerical assumptions. It enables native block-scaled FP8 in every forward and backward attention-core matmul without architectural changes, smaller global batches, or auxiliary forward outputs. Across tested architectures, scales, and training stages, Delta-Matching matches BF16/FP32 mixed-precision training loss and overall downstream performance. We will release our implementation, trained models, and data recipes.
Chinese Translation
可靠的 FP8 注意力仍然是完全原生 8 位大语言模型训练的一个障碍。我们推导了前向-后向不一致如何产生陈旧 delta,并通过实验表明它如何扭曲训练动态。我们的陈旧 delta 混合运行显示,在 569M 参数时损失差距不大,但在 1.67B 和 5.29B 时损失显著增加且下游性能退化。QK 归一化、NoPE(无位置编码)以及较低学习率的上下文扩展可以缓解或延迟退化,但不能消除它。这种模式表明存在累积的优化误差,而较小模型和较短运行可以将其掩盖。我们提出 Delta-Matching,并证明在所述数值假设下,它恢复了 softmax 梯度的零行和不变性。它使得在每个前向和反向注意力核心矩阵乘法中都能使用原生块缩放 FP8,而无需架构更改、更小全局批次或辅助前向输出。在测试的架构、规模和训练阶段中,Delta-Matching 与 BF16/FP32 混合精度训练的损失和整体下游性能相匹配。我们将发布我们的实现、训练后的模型和数据配方。
cs.LG / 162 / 2609.37858
Storage Is Not Strategy: State-Conditioned Support Control for LLM Unlearning
存储不是策略:面向LLM遗忘的状态条件支持控制
large language model
大语言模型相关
Abstract
Many localized large language model (LLM) unlearning methods select a small parameter subset from a localization signal and keep it fixed during optimization. The parameters most associated with a target, however, need not be the best ones to update, and candidate interventions can change value as optimization proceeds. In a controlled experiment, a storage-localization score reaches an area under the receiver operating characteristic curve (AUROC) of 0.981, yet storage identity agrees with the better intervention on only 17/36 targets, while low-rank adaptation (LoRA) wins 35/36. We introduce Intervention Score, which ranks editable groups by the predicted effect of the actual unlearning update while accounting for collateral damage, and use it to form the static intervention-value baseline (Static-IV). We then introduce selective dynamic intervention re-ranking (DIR-R), which revisits that subset only when a calibrated probe justifies the comparison. On the Natural-TOFU dataset, our method has positive descriptive margins in 19/20 comparisons between methods and objectives, although several are near zero. On the LACUNA localization-precision benchmark, our mean terminal utility is higher in all six negative preference optimization (NPO) and SimNPO comparisons: NPO margins range from +0.431 to +0.848, and SimNPO margins range from +0.503 to +0.571. The gradient-difference (GradDiff) objective reveals substantial field dependence. Relative to Static-IV, the primary four-field GradDiff evaluation has six wins, six ties, and no losses, with mean and median paired gains of +0.165 and +0.0025. The evidence supports separating localization, initial intervention selection, and checkpoint-dependent support revision.
Chinese Translation
许多局部化大语言模型(LLM)遗忘方法从定位信号中选取一个小的参数子集,并在优化过程中保持其固定不变。然而,与目标关联最紧密的参数未必是最适合更新的参数,并且候选干预的价值可能随优化的推进而变化。在一项受控实验中,存储定位分数达到了 0.981 的受试者工作特征曲线下面积(AUROC),然而存储身份仅在 17/36 个目标上与更优干预一致,而低秩适应(LoRA)在 35/36 个目标上胜出。我们引入干预分数(Intervention Score),它根据实际遗忘更新的预测效果对可编辑组进行排序,同时考虑附带损伤,并用其构建静态干预价值基线(Static-IV)。随后我们引入选择性动态干预重排序(DIR-R),它仅在校准探针证明该比较合理时才重新审视该子集。在 Natural-TOFU 数据集上,我们的方法在方法与目标之间的 19/20 项比较中具有正向的描述性差距,尽管其中若干项接近于零。在 LACUNA 定位精度基准上,我们的平均终端效用在全部六项负偏好优化(NPO)与 SimNPO 比较中均更高:NPO 的差距范围为 +0.431 至 +0.848,SimNPO 的差距范围为 +0.503 至 +0.571。梯度差(GradDiff)目标揭示出显著的领域依赖性。相对于 Static-IV,主要的四领域 GradDiff 评估有六次胜出、六次打平、没有落败,平均配对增益和中位数配对增益分别为 +0.165 和 +0.0025。这些证据支持将定位、初始干预选择以及依赖于检查点的支持修订分离开来。
cs.LG / 163 / 2609.37924
Time-Anchored Diffusion Language Models: Latent-Space Caching for Fast Generation
时间锚定的扩散语言模型:用于快速生成的潜在空间缓存
diffusion
扩散模型相关
Abstract
Recent work on anchored diffusion language models improves denoising by shaping an intermediate latent space with supervised important-token targets. In this work, we introduce time-based (self-supervised) anchoring, which learns and reuses latent anchors without requiring such targets. Our key observation is that anchors encode persistent properties of the clean sequence, such as its semantic intent, global structure, or intermediate plan. Although their hidden representations become stale as the token canvas evolves, their semantic content remains useful across nearby diffusion times. This is implemented through a two-stage architecture consisting of a relatively expensive anchor network that generates the latent cache state and a lightweight denoising network that intelligently combines the cached latent state with the current state at each reverse step using a fusion module. This gives anchoring a latent-space caching interpretation: the anchor network is evaluated periodically, while its cached representation is reused across multiple reverse steps. We instantiate this framework as TADM:Post-train, which time-anchorizes pretrained DLMs, and TADM:Pretraining, which learns time-based anchors during pretraining. Applied to DiffusionGemma-26B, TADM:Post-train improves throughput by approximately 49% to 79% on several math, code, and STEM benchmarks (GSM8K, AIME26, GPQA-Diamond, LiveCodeBench-v6, HumanEval, MMLU-Pro). TADM:Pretraining reduces Transformer-layer computation by up to 38% relative to a standard single-stage DLM, achieves up to 73% higher measured throughput than ADLM.
Chinese Translation
近期关于锚定扩散语言模型的工作通过使用有监督的重要词元目标来塑造中间潜在空间,从而改进去噪。在这项工作中,我们引入了基于时间的(自监督)锚定,它无需此类目标即可学习并复用潜在锚。我们的关键观察是,锚编码了干净序列的持久属性,例如其语义意图、全局结构或中间计划。尽管随着词元画布的演化,它们的隐藏表示会变得陈旧,但其语义内容在邻近的扩散时间上仍然有用。这通过一个两阶段架构实现,该架构由一个相对昂贵的、生成潜在缓存状态的锚网络和一个轻量级去噪网络组成,后者在每个反向步骤中使用融合模块将缓存的潜在状态与当前状态智能地结合。这赋予了锚定一种潜在空间缓存的解释:锚网络被周期性地评估,而其缓存表示则在多个反向步骤中被复用。我们将该框架实例化为 TADM:Post-train(对预训练 DLM 进行时间锚定)和 TADM:Pretraining(在预训练期间学习基于时间的锚)。应用于 DiffusionGemma-26B 时,TADM:Post-train 在若干数学、代码和 STEM 基准(GSM8K、AIME26、GPQA-Diamond、LiveCodeBench-v6、HumanEval、MMLU-Pro)上将吞吐量提升约 49% 至 79%。相对于标准的单阶段 DLM,TADM:Pretraining 将 Transformer 层计算量减少最多 38%,并实现了比 ADLM 高最多 73% 的实测吞吐量。
cs.LG / 164 / 2609.37951
Scene-Consistent Illumination Transfer for Inserted Advertising Graphics
用于插入广告图形的场景一致光照迁移
diffusion
扩散模型相关
Abstract
Replacing a visible advertisement in a broadcast frame is geometrically straightforward but photometrically delicate. A pasted graphic can have the correct perspective and still appear detached when its brightness, shading, or shadow disagrees with the surface beneath it. This paper presents Ad-Relight, an inference-only procedure for transferring scene illumination to a supplied advertising graphic without collecting a banner-specific training set. The procedure first separates slowly varying shade from graphic structure, then probes a pretrained diffusion relighter with two nearly identical backgrounds to isolate the contribution of the target region. A final pass combines this residual with a smoothed luminance field and a soft attenuation mask. Across 560 generated placements, the approach improves structural similarity, perceptual distance, and illumination agreement over geometric compositing and direct relighting baselines. Human judgments and an automated preference study show the clearest gains on floor-mounted graphics with nonuniform lighting. The current study is image based; temporal stabilization remains an open extension.
Chinese Translation
在广播帧中替换可见广告在几何上很直接,但在光度上却很精细。粘贴的图形即使具有正确的透视,当其亮度、明暗或阴影与其下方表面不一致时,仍会显得脱离。本文提出 Ad-Relight,一种仅需推理的过程,用于将场景光照迁移到给定的广告图形,而无需收集特定横幅的训练集。该过程首先将缓慢变化的明暗从图形结构中分离出来,然后用两个几乎相同的背景探测一个预训练的扩散重光照器,以隔离目标区域的贡献。最后一个处理遍将此残差与一个平滑亮度场和一个软衰减掩码相结合。在 560 个生成的放置中,该方法在结构相似性、感知距离和光照一致性方面优于几何合成和直接重光照基线。人工判断和一项自动化偏好研究表明,在具有非均匀光照的地面安装图形上增益最明显。当前研究基于图像;时间稳定化仍是一个开放的扩展方向。
cs.LG / 165 / 2609.37974
On Trajectory-Aware Training for Masked Diffusion Language Models
论掩码扩散语言模型的轨迹感知训练
diffusion
扩散模型相关
Abstract
Masked diffusion models (MDMs) generate text by unmasking several tokens per step, but they are trained and sampled under different conditions. The model is trained on randomly masked sequences, whereas inference follows a trajectory shaped by the model's own predictions. Additionally, each step has no access to what the previous one computed. Recent methods narrow these limitations from separate angles, leaving open how these choices interact. We introduce PUMBA, a unified framework for trajectory-aware training that trains the denoiser on consecutive steps of policy-induced trajectories, passes information between steps, and optimizes them jointly by backpropagation through time. A controlled study of this design space shows that i) exact train--inference alignment fails due to local overfitting, whereas a looser alignment still brings training masks closer to those seen at inference; ii) passing continuous information outperforms discrete gradient estimators through the commitment at each step; and iii) performance improves as backpropagation through time spans more steps, which we support theoretically. Combined, these components match the best checkpoint of a same-size autoregressive model. Building on these findings, we scale PUMBA to supervised fine-tuning of LLaDA-8B, where it improves the trade-off between performance and number of function evaluations (NFEs) in both full-canvas and block diffusion generation. At matched performance, it needs up to 22% fewer NFEs than standard fine-tuning with twice the budget in full-canvas generation, and up to 26% fewer than standard fine-tuning for the same number of steps in block diffusion.
Chinese Translation
掩码扩散模型(MDMs)通过每步解掩码若干 token 来生成文本,但它们在训练和采样时处于不同的条件下。该模型在随机掩码的序列上训练,而推理则遵循由模型自身预测所塑造的轨迹。此外,每一步都无法访问前一步所计算出的内容。近期方法从不同角度缓解了这些局限,但这些选择如何相互作用仍悬而未决。我们提出 PUMBA,一个用于轨迹感知训练的统一框架,它在策略诱导轨迹的连续步骤上训练去噪器,在步骤之间传递信息,并通过时间反向传播联合优化它们。对这一设计空间的受控研究表明:i) 精确的训练--推理对齐会因局部过拟合而失败,而更宽松的对齐仍使训练掩码更接近推理时见到的掩码;ii) 通过每一步的 commitment(承诺),传递连续信息优于离散梯度估计器;iii) 随着时间反向传播跨越更多步骤,性能会提升,我们从理论上支持这一点。结合起来,这些组件可匹敌同等规模自回归模型的最佳检查点。基于这些发现,我们将 PUMBA 扩展到 LLaDA-8B 的监督微调,其中它改善了全画布生成和块扩散生成中性能与函数评估次数(NFEs)之间的权衡。在性能匹配时,在全画布生成中,与拥有两倍预算的标准微调相比,它所需的 NFEs 最多减少 22%;在块扩散中,在相同步数下,与标准微调相比,它所需的 NFEs 最多减少 26%。
cs.LG / 166 / 2609.38025
Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models
Dr. OPD:学习该遵循什么以实现大型语言模型的最优同策略蒸馏
large language model
大语言模型相关
Abstract
On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher's supervision is equally important for every token. However, teacher signals at different tokens may have very different effects on the student's performance: some correct important reasoning errors, while others have little effect on the final answer. Motivated by this observation, we introduce Dr. OPD (OPD Done Right), which defines the optimal weighted OPD to maximize the student's performance. We formulate Dr. OPD as a bilevel optimization problem in which the student learns from weighted teacher supervision, while the weights are selected to maximize the expected reward of the resulting student. To solve Dr. OPD, we develop an efficient iterative solver that updates the token weights and student policy alternatively. At each round, it updates weights in closed form and then takes one gradient step on the resulting weighted OPD objective. Under regularity conditions, we show that this weighted update achieves a higher expected reward than a vanilla OPD update. Empirically, across strong-to-weak and same-size distillation on math and code, Dr. OPD consistently outperforms all evaluated baselines. In particular, in the strong-to-weak distillation setting, Dr. OPD improves average math performance by $9.7$ points over vanilla OPD, and enables the smaller student to surpass its larger teacher.
Chinese Translation
同策略蒸馏(OPD)利用来自更强教师的密集、词元级监督,在学生自己生成的回复上训练学生。朴素 OPD 同等对待所有教师信号,假设教师的监督对每个词元都同等重要。然而,不同词元处的教师信号可能对学生的性能产生非常不同的影响:一些信号纠正重要的推理错误,而另一些对最终答案几乎没有影响。受这一观察启发,我们提出 Dr. OPD(OPD Done Right),它定义了最优加权 OPD 以最大化学生的性能。我们将 Dr. OPD 形式化为一个双层优化问题,其中学生从加权的教师监督中学习,而权重被选择以最大化所得学生的期望奖励。为求解 Dr. OPD,我们开发了一种高效的迭代求解器,它交替更新词元权重和学生策略。在每一轮中,它以闭式形式更新权重,然后在所得的加权 OPD 目标上采取一步梯度更新。在正则性条件下,我们证明这种加权更新比朴素 OPD 更新实现了更高的期望奖励。在实证上,在数学和代码上的强到弱蒸馏以及同规模蒸馏中,Dr. OPD 始终优于所有评估的基线。特别地,在强到弱蒸馏设置中,Dr. OPD 相较于朴素 OPD 将平均数学性能提高了 $9.7$ 分,并使较小的学生超越其更大的教师。
cs.LG / 167 / 2609.38049
Improving Function Space Flow Matching with Kernel Optimal Transport
用核最优传输改进函数空间流匹配
diffusion
扩散模型相关
Abstract
Generative models for function-valued data, such as time series and solutions of partial differential equations, must learn distributions over infinite-dimensional spaces. Functional Flow Matching (FFM) extends Flow Matching to this setting, learning a velocity field whose flow transports a Gaussian prior to the data distribution, but it inherits the independent endpoint pairing of standard Flow Matching: in each batch, prior and data samples are matched arbitrarily, so the conditional bridge must traverse both the shared global structure of the dataset and instance-specific residuals. In function space this is harder to fix than in finite dimensions, since optimal transport (OT) on function spaces is delicate to formulate and a flat Euclidean surrogate ignores the geometry that distinguishes function-valued data. We propose kernel Functional Flow Matching (kFFM), which replaces the independent pairing by entropic OT under a kernel-induced cost, the coupling underlying the Hilbert Sinkhorn Divergence (HSD), leaving the FFM neural-operator architecture unchanged. We prove that the kernel cost and the HSD objective are uniformly bounded and well-posed on Banach ambient spaces, derive an error decomposition against quadratic-cost OT on compact metric spaces that isolates an irreducible kernel-cost mismatch term, and prove a discretization-invariance bound whose rate is governed by Sobolev regularity. Empirically, kFFM improves distributional matching over FFM, diffusion, adversarial, and finite-dimensional OT baselines on time-series and PDE benchmarks, with significant paired-seed gains over FFM and improvements that persist under non-kernel and physics-based diagnostics, including a turbulent Navier-Stokes benchmark. Bounded kernel costs already outperform raw $L^2$ Sinkhorn, and function-space-aware kernels (signature, Sobolev RBF) give further gains on rough or path-valued data.
Chinese Translation
面向函数值数据(例如时间序列和偏微分方程的解)的生成模型必须学习无限维空间上的分布。函数流匹配(Functional Flow Matching, FFM)将流匹配扩展到这一设定,学习一个速度场,其流将高斯先验输运到数据分布,但它继承了标准流匹配的独立端点配对:在每个批次中,先验样本和数据样本被任意匹配,因此条件桥必须同时穿越数据集的共享全局结构和实例特定的残差。在函数空间中,这比在有限维中更难修正,因为函数空间上的最优传输(optimal transport, OT)难以形式化表述,而平坦的欧几里得替代方案忽略了区分函数值数据的几何结构。我们提出核函数流匹配(kernel Functional Flow Matching, kFFM),它用核诱导代价下的熵最优传输——即 Hilbert Sinkhorn 散度(Hilbert Sinkhorn Divergence, HSD)背后的耦合——来替代独立配对,同时保持 FFM 的神经算子架构不变。我们证明核代价和 HSD 目标在 Banach 环境空间上一致有界且适定,推导出在紧度量空间上与二次代价 OT 相比的误差分解,其中分离出一个不可约的核代价失配项,并证明一个离散化不变性界,其速率由 Sobolev 正则性决定。在实证上,kFFM 在时间序列和 PDE 基准上相较于 FFM、扩散、对抗以及有限维 OT 基线改善了分布匹配,并且相对于 FFM 取得了显著的配对种子增益,这些改进在非核诊断和基于物理的诊断(包括湍流 Navier-Stokes 基准)下依然存在。有界核代价已经优于原始 $L^2$ Sinkhorn,而函数空间感知核(签名核、Sobolev RBF)在粗糙或路径值数据上带来进一步提升。
cs.LG / 168 / 2609.38066
Alpha Diffusion Language Models: Factorization Alone Is Not the Problem
Alpha 扩散语言模型:问题并不在于因式分解本身
diffusion
扩散模型相关
Abstract
Discrete diffusion language models can generate multiple tokens in parallel, but reducing the number of denoising steps can lead to inconsistent predictions. Standard cross-entropy training fits conditional token marginals, whereas parallel generation requires consistent joint predictions. We introduce Alpha Diffusion Language Models (AlphaDLM), trained with a sequence-level alpha loss that recovers cross-entropy in the limit of vanishing alpha and has a joint-mode optimum at alpha one. Our analysis characterizes how the objective and factorization jointly determine the fitted distribution. We identify conditions under which intermediate alpha preserves multiple valid completions while excluding invalid token combinations. Trained on TinyGSM, our method achieves 34.6% accuracy on GSM8K with only four model evaluations. We further scale the method to SDAR-1.7B and evaluate it on code and mathematics benchmarks. These results show that changing the training objective can improve the accuracy-computation trade-off of factorized diffusion language models.
Chinese Translation
离散扩散语言模型可以并行生成多个词元,但减少去噪步骤可能导致预测不一致。标准交叉熵训练拟合的是条件词元边缘分布,而并行生成需要一致的联合预测。我们引入 Alpha 扩散语言模型(AlphaDLM),其使用序列级 alpha 损失进行训练,该损失在 alpha 趋于零的极限下恢复为交叉熵,并在 alpha 为 1 时具有联合模式最优。我们的分析刻画了目标函数与因式分解如何共同决定所拟合的分布。我们识别出一些条件,在这些条件下,中间 alpha 会保留多个有效补全,同时排除无效词元组合。在 TinyGSM 上训练后,我们的方法在 GSM8K 上仅用四次模型评估就达到了 34.6% 的准确率。我们进一步将该方法扩展到 SDAR-1.7B,并在代码和数学基准上对其进行评估。这些结果表明,改变训练目标可以改善因式分解扩散语言模型的准确率-计算量权衡。
cs.LG / 169 / 2609.38104
Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling
广泛探索,锐利推理:通过采样将小模型推向前沿
large language model
大语言模型相关
Abstract
Power-sharpened sampling is an inference-time alternative to reinforcement-learning (RL) post-training for enhancing reasoning in large language models (LLMs). High-probability sequences are amplified under the base model without parameter updates or external rewards, avoiding the costly optimization and jagged generalization of RL. However, this approach faces a fundamental exploration--exploitation trade-off, as % strong sharpening restricts exploration, trapping samplers in plausible but incorrect reasoning trajectories, whereas weak sharpening leaves the answer distribution diffuse. To resolve this trade-off, we introduce \textbf{Parallel Power Tempering (PPT)}, instantiating power-sharpened LLM sampling via parallel tempering. Running multiple \emph{interacting} replicas in parallel at different sharpening levels allows lower-power replicas to explore diverse reasoning trajectories and higher-power chains to further exploit higher-likelihood responses favored by the sharpened target. Specifically, we tailor \method{} to inference-time sampling by mitigating a truncation bias, identified in prior power samplers, and investigate effective swap strategies under finite memory and compute budgets. Extensive experimentation shows that \method{} substantially improves single-chain power-sharpened sampling and outperforms RL-post-trained models, producing higher-quality reasoning traces and even achieving performance comparable to frontier models.
Chinese Translation
幂次锐化采样是强化学习(RL)后训练的一种推理时替代方案,用于增强大型语言模型(LLMs)的推理。在基础模型下,高概率序列被放大,无需参数更新或外部奖励,从而避免了 RL 昂贵的优化和参差不齐的泛化。然而,这种方法面临一个根本性的探索--利用权衡,因为 % 强锐化会限制探索,将采样器困在看似合理但不正确的推理轨迹中,而弱锐化则使答案分布变得弥散。为了解决这一权衡,我们引入 \textbf{并行幂次回火(PPT)},通过并行回火来具体实现幂次锐化的 LLM 采样。在不同锐化水平下并行运行多个 \emph{交互} 副本,使较低幂次的副本能够探索多样化的推理轨迹,而较高幂次的链进一步利用锐化目标所偏好的较高似然响应。具体而言,我们通过缓解先前幂采样器中识别出的截断偏差,将 \method{} 定制用于推理时采样,并研究在有限内存和计算预算下的有效交换策略。大量实验表明,\method{} 显著改进了单链幂次锐化采样,并优于经过 RL 后训练的模型,产生更高质量的推理轨迹,甚至达到与前沿模型相当的性能。
cs.MA / 170 / 2609.37094
LLM-Based Multi-Agent Systems over Wireless Networks: A Joint Agent--Network Design Perspective
无线网络上的基于LLM的多智能体系统:一种智能体-网络联合设计视角
large language model
大语言模型相关
Abstract
As large language models (LLMs) evolve from standalone models into collaborative agents embedded in physical systems, their reasoning and execution are increasingly distributed across wireless edge nodes. In this setting, wireless networks are experiencing a paradigm shift from only providing data connectivity to supporting the multi-agent reasoning workflow itself. The task performance of such network-constrained LLM-based multi-agent systems (MASs) is jointly affected by the multi-agent reasoning dependencies as well as the underlying network connectivity and edge resources. This coupling gives rise to various technical challenges, including the metric misalignment and message redundancy, state inconsistency and topology mismatch, as well as resource limitation and trust discontinuity. To address these challenges, this article develops a novel joint agent--network design perspective that coordinates decisions on both sides of the system. Specifically, we present the joint design of agent--interaction scheduling and resource allocation, the message selection-transmission co-design, as well as the joint agent--network topology design and workload--resource allocation. Furthermore, we consider the network-verified provenance that is linked with agent-side information-flow control to constrain how received information affects subsequent operations. An illustrative vehicle-to-everything (V2X) case study shows that jointly adapting agent-side interaction decisions and network operations improves task completion under communication and edge-resource constraints, outperforming the conventional agent-only and wireless-only separate designs.
Chinese Translation
随着大语言模型(LLM)从独立模型演变为嵌入物理系统中的协作智能体,其推理和执行日益分布在无线边缘节点上。在这种背景下,无线网络正经历从仅提供数据连接性到支持多智能体推理工作流本身的范式转变。这种受网络约束的基于LLM的多智能体系统(MAS)的任务性能,受到多智能体推理依赖关系以及底层网络连接性和边缘资源的共同影响。这种耦合引发了各种技术挑战,包括度量不一致和消息冗余、状态不一致和拓扑不匹配,以及资源限制和信任不连续。为了应对这些挑战,本文提出了一种新颖的智能体-网络联合设计视角,该视角协调系统两侧的决策。具体而言,我们提出了智能体-交互调度与资源分配的联合设计、消息选择-传输协同设计,以及智能体-网络拓扑设计与工作负载-资源分配的联合设计。此外,我们考虑了与智能体侧信息流控制相关联的网络验证溯源,以约束接收到的信息如何影响后续操作。一个示例性的车联网(V2X)案例研究表明,在通信和边缘资源约束下,联合调整智能体侧交互决策和网络操作可以提高任务完成度,优于传统的仅智能体和仅无线的分离设计。
cs.AI / 171 / 2609.37233
DatalogBench: Evaluating Large Language Models on Text-to-Datalog Synthesis
DatalogBench:评估大语言模型在文本到 Datalog 合成上的表现
large language model
大语言模型相关
Abstract
Datalog underpins reasoning tasks such as program analysis, but its programs are hard to write. Existing synthesizers automate this task but require users to state their intent as input-output examples. Large language models (LLMs) suggest a more natural route, text-to-Datalog synthesis from a natural-language question, yet how well they do so has not been systematically evaluated. We present DatalogBench, a benchmark of 136 text-to-Datalog synthesis tasks curated from existing Datalog-based artifacts. Synthesized programs are graded by execution on held-out inputs against an oracle validated by mutation analysis. Across six LLMs and four prompting configurations, exact match peaks at 68.4%, and relation descriptions or an input-output example have only modest, model-dependent effects. Under direct prompting, most failures occur at compile time, typically because a model invents auxiliary predicates that it never declares or types consistently. Two coding agents reach up to 83.8% and eliminate nearly all such failures, leaving mostly semantic errors concentrated in recursive tasks. DatalogBench thus identifies recursive reasoning and decomposition as open challenges for current LLMs and agents, and offers a reliable, execution-grounded measure of both.
Chinese Translation
Datalog 支撑着程序分析等推理任务,但其程序难以编写。现有的合成器将这一任务自动化,但要求用户以输入-输出示例的形式陈述其意图。大语言模型(LLM)提出了一条更自然的路径,即从自然语言问题出发进行文本到 Datalog 的合成,然而它们在这方面的表现如何尚未得到系统性评估。我们提出 DatalogBench,一个包含 136 个文本到 Datalog 合成任务的基准,这些任务整理自现有的基于 Datalog 的工件。合成程序通过在留出输入上的执行,与经变异分析验证的 oracle 进行比对来评分。在六个 LLM 和四种提示配置中,精确匹配率最高达到 68.4%,而关系描述或输入-输出示例仅产生有限的、依赖模型的影响。在直接提示下,大多数失败发生在编译时,通常是因为模型虚构了辅助谓词,却从未声明它们,或没有一致地为其指定类型。两个编码智能体最高达到 83.8%,并消除了几乎所有此类失败,剩下的大多是集中在递归任务中的语义错误。因此,DatalogBench 将递归推理和分解确定为当前 LLM 和智能体面临的开放挑战,并为二者提供了一种可靠的、基于执行的度量。
cs.AI / 172 / 2609.37554
Risk-Aware Semantic Grounding for Trustworthy LLM-Based Robot Planning
面向可信赖的基于LLM的机器人规划的风险感知语义接地
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly used as high-level planners in robot navigation, but their outputs may become unreliable when instructions are ambiguous, unsupported by the environment, or semantically inconsistent. This paper presents a Risk-Aware Semantic Grounding framework for trustworthy LLM-based robot planning. Unlike existing LLM-based planners that primarily optimize plan generation, we formulate semantic grounding reliability as a multi-dimensional risk estimation problem. The proposed architecture explicitly models grounding uncertainty through ambiguity, hallucination and semantic-conflict risks before planning occurs, enabling the system to decide whether to execute the instruction, request clarification, or reject it. To evaluate the approach, we introduce TRUST-NAV, a benchmark containing both standard navigation tasks and risk-inducing instruction scenarios. Experimental results show that while conventional LLM planners achieve strong performance on valid navigation tasks, the proposed framework substantially improves ambiguity detection and semantic conflict rejection. These findings suggest that trustworthy robot planning should be evaluated not only by task completion, but also by the ability to recognize when execution should not occur.
Chinese Translation
大语言模型(LLMs)正越来越多地被用作机器人导航中的高层规划器,但当指令存在歧义、环境不支持或语义不一致时,其输出可能变得不可靠。本文提出了一种面向可信赖的基于LLM的机器人规划的风险感知语义接地框架。与主要优化规划生成的现有基于LLM的规划器不同,我们将语义接地的可靠性建模为一个多维风险估计问题。所提出的架构在规划发生之前,通过歧义、幻觉和语义冲突风险显式地对接地不确定性进行建模,使系统能够决定是执行指令、请求澄清,还是拒绝该指令。为评估该方法,我们引入了TRUST-NAV,一个同时包含标准导航任务和诱发风险指令场景的基准。实验结果表明,尽管传统LLM规划器在有效导航任务上取得了强劲性能,但所提出的框架显著改善了歧义检测和语义冲突拒绝能力。这些发现表明,可信赖的机器人规划不仅应通过任务完成情况来评估,还应通过识别何时不应执行的能力来评估。
cs.AI / 173 / 2609.36577
Long-Term Memory-Guided Enhancement for Target Perception in Audio-Language Models
面向音频-语言模型目标感知的长期记忆引导增强
large language model
大语言模型相关
Abstract
Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually collapse in real-world environments when background noise and competing sources mix the target sound. Inspired by long-term memory in human listening, we propose Long-Term Memory-Guided Audio Enhancement (LTM-AE) to improve selective target perception by refining the audio representations of ALLMs without training. LTM-AE extracts representations in hidden states from separate clean reference recordings as long-term memory for each category, guiding enhancement toward a user-specified listening target. We reconstruct incoming audio tokens in the selected category long-term memory and interpolate the reconstructions with the original tokens before language backbone decoding. This interpolation controls the influence of stored auditory experience while keeping all ALLM parameters fixed. Diagnostic readouts across twenty sound categories and three ALLMs show that LTM-AE strengthens responses to a specified target amid three interfering sources. Averaged over constrained and free-form classification, accuracy gains over raw mixtures range from 29.53 to 46.15 percentage points across multiple open source models. For speech content recovery, LTM-AE with an additional learned token-level gate reduces Qwen2-Audio's word error rate from 23.07% to 14.77%. This work takes an initial step toward using principles of human long-term memory to enhance ALLMs for real-world listening. Our code is available at https://github.com/aynlp/ltm-audio-code
Chinese Translation
音频大语言模型(ALLMs)能够对音频录音的内容进行推理,以执行复杂任务。然而,在真实环境中,当背景噪声和竞争声源与目标声音混合时,这些能力通常会崩溃。受人类听觉中长期记忆的启发,我们提出长期记忆引导的音频增强(LTM-AE),通过在不训练的情况下细化ALLMs的音频表示来提升选择性目标感知。LTM-AE从单独的干净参考录音中提取隐藏状态中的表示,作为每个类别的长期记忆,引导增强朝向用户指定的聆听目标。我们在所选类别的长期记忆中重建传入的音频标记,并在语言主干解码之前将重建结果与原始标记进行插值。这种插值控制了存储的听觉经验的影响,同时保持所有ALLM参数固定。在二十个声音类别和三个ALLMs上的诊断性读出表明,LTM-AE增强了在三个干扰源中对指定目标的响应。在受限分类和自由形式分类上平均而言,在多个开源模型中,相较于原始混合音频,准确率提升范围为29.53到46.15个百分点。对于语音内容恢复,带有额外学习的标记级门控的LTM-AE将Qwen2-Audio的词错误率从23.07%降至14.77%。这项工作为利用人类长期记忆原理增强ALLMs以用于真实世界聆听迈出了初步一步。我们的代码可在 https://github.com/aynlp/ltm-audio-code 获取。
cs.SE / 174 / 2609.36289
How Much Prompt Is Enough? A Blackbox Minimization of Few-Shots in LLMs
多少提示才足够?LLM 中少样本提示的黑盒最小化
large language model
大语言模型相关
Abstract
Prompts are the primary mechanism for directing the behavior of large language models (LLMs). Yet the internal structure and causal hierarchy of prompts remain poorly understood: which parts are causally necessary and which are redundant is an open question. This opacity can have severe consequences. Subtle prompt variations can silently shift model outputs in critical software systems, and engineers lack techniques to reason about prompt reliability. We present \framework, a blackbox prompt-minimization framework that reduces few-shot prompts to their necessary minimal subset. We use a case study to apply \framework to a few-shot learning system and demonstrate the insights that this framework can provide. Our experiments show that few-shot exemplars can be reduced by a mean of 65.3\%~$\pm$~15.8\% in character count while fully preserving propositional output fidelity. The models preferentially retain logical identifiers and constraint declarations while discarding natural language prose and cross-prompt relational annotations. Our analysis also shows that some models are universal encoders, able to produce highly legible yet minimized prompts, while others are universal decoders, able to interpret minimized prompts from most other models. By identifying which components are indispensable, \framework provides a principled basis for prompt compression and structural analysis of few-shot exemplars.
Chinese Translation
提示是引导大语言模型(LLM)行为的主要机制。然而,提示的内部结构与因果层级仍鲜为人知:哪些部分是因果上必需的、哪些是冗余的,仍是一个悬而未决的问题。这种不透明性可能带来严重后果。细微的提示变动可能在关键软件系统中悄然改变模型输出,而工程师缺乏对提示可靠性进行推理的技术手段。我们提出 \framework,一个黑盒提示最小化框架,可将少样本提示缩减为其必要的最小子集。我们通过一个案例研究,将 \framework 应用于一个少样本学习系统,并展示该框架所能提供的洞见。我们的实验表明,少样本示例在字符数上平均可减少 65.3\%~$\pm$~15.8\%,同时完全保持命题式输出的保真度。模型倾向于保留逻辑标识符与约束声明,而丢弃自然语言散文以及跨提示的关系标注。我们的分析还表明,一些模型是通用编码器,能够生成高度易读但已被最小化的提示;而另一些模型则是通用解码器,能够解读来自大多数其他模型的最小化提示。通过识别哪些组件不可或缺,\framework 为提示压缩与少样本示例的结构分析提供了有原则的基础。
cs.SE / 175 / 2609.36783
The Editor Has Read-Only Access: Correctness Signals in Diffusion Language Models
编辑器仅具有只读访问权限:扩散语言模型中的正确性信号
diffusion
扩散模型相关
Abstract
Diffusion language models generate code by repeatedly updating a partially masked sequence. We ask whether their internal activations encode code correctness and whether that information can improve generation. Across six diffusion models, linear probes distinguish passing from failing attempts, with the strongest reads generally appearing beyond the early layers. Controls using small semantic mutations support a connection to correctness rather than surface style alone. In comparisons with model confidence, probe point estimates offer no consistent advantage. Adding a probe-derived direction to the residual stream does not yield a dependable improvement in the tested steering settings, while the opposite direction degrades performance. We distinguish these observations from claims about statistical significance or a general inability to steer. Supplementary methods, archived results, and code document the tested interventions and the limits of their statistical calibration and reproducibility.
Chinese Translation
扩散语言模型通过反复更新一个部分遮蔽的序列来生成代码。我们追问它们的内部激活是否编码了代码正确性,以及该信息能否改进生成。在六个扩散模型中,线性探针能够区分通过与否的尝试,最强的读出通常出现在早期层之后。使用小型语义突变的对照支持其与正确性相关,而不仅仅是表面风格。在与模型置信度的比较中,探针点估计没有提供一致的优势。将源自探针的方向添加到残差流中,在测试的引导设置中并未带来可靠的改进,而相反方向则会降低性能。我们将这些观察与关于统计显著性的主张或普遍无法引导的主张区分开来。补充方法、归档结果和代码记录了所测试的干预措施及其统计校准与可复现性的局限。
cs.SE / 176 / 2609.37000
Cross-Organizational SysML Model Integration: A Survey of Challenges and AI-Supported Tasks
跨组织SysML模型集成:一项关于挑战与AI支持任务的调查研究
large language model
大语言模型相关
Abstract
Cross-organizational collaboration is widely regarded as a key promise of SysML-based Model-Based Systems Engineering (MBSE), yet practitioners still face persistent challenges when exchanging and integrating system models. In parallel, Large Language Models (LLMs) raise expectations for AI-assisted model understanding and integration, while reliability and required human oversight continue to pose challenges. This paper reports the results of an online questionnaire survey with 29 MBSE stakeholders involved in cross-organizational collaboration. Respondents rated eight predefined integration challenge categories and six AI-supported task types on five-point Likert scales. The results indicate that stakeholders perceive model integration as a multi-dimensional alignment problem across semantics, behavior, traceability, and exchange interoperability. These perceptions vary by organizational role and frequency of integration involvement. AI is rated highly useful for analysis tasks such as semantic structure analysis and inconsistency detection, and respondents predominantly prefer human-in-the-loop use with mandatory verification. These findings motivate AI support that enhances, rather than replaces, engineering responsibility in SysML-based integration.
Chinese Translation
跨组织协作被广泛视为基于SysML的基于模型的系统工程(MBSE)的一项关键前景,然而从业者在交换和集成系统模型时仍面临持续存在的挑战。与此同时,大语言模型(LLM)提升了人们对AI辅助的模型理解与集成的期待,而可靠性与所需的人类监督仍持续构成挑战。本文报告了一项在线问卷调查的结果,该调查涉及29位参与跨组织协作的MBSE利益相关者。受访者在五点李克特量表上对八个预定义的集成挑战类别和六种AI支持的任务类型进行了评分。结果表明,利益相关者将模型集成视为一个跨越语义、行为、可追溯性和交换互操作性的多维对齐问题。这些认知因组织角色和参与集成的频率而异。AI在诸如语义结构分析和不一致性检测等分析任务上被评为高度有用,且受访者主要倾向于采用带强制验证的人在回路使用方式。这些发现为这样一种AI支持提供了依据:在基于SysML的集成中增强而非取代工程责任。
cs.SE / 177 / 2609.37216
CRJudgeBench: Can AI Detect Plausible but Invalid Code Reviews?
CRJudgeBench:AI 能否检测出看似合理但无效的代码审查?
large language model
大语言模型相关
Abstract
Large language models can generate plausible code-review comments, but such comments may contain technically incorrect claims that mislead developers. We study technical trustworthiness judgment: determining whether a review comment's core technical claims are correct and applicable to the reviewed code in its repository context. Existing code-review benchmarks primarily evaluate review generation, issue discovery, or general comment quality, but do not directly assess whether an agent can determine the technical trustworthy of an individual review comment. To fill this gap, we introduce CRJudgeBench, a benchmark of 1199 instances constructed from real pull requests and expert-verified perturbations, covering both trustworthy and plausible but untrustworthy comments. We further present Sentinel, a repository-grounded agentic judge that actively gathers code evidence to verify review comments before making judgments. Starting from Qwen3-Coder-30B-A3B-Instruct, Sentinel is trained on the CRJudgeBench training split through iterative action-level learning from a privileged teacher. On the 359-instance CRJudgeBench test set, Sentinel achieves 76.60\% accuracy, outperforming GLM-5.3 by 6.13 percentage points and its base model by 19.78 points. These results show that even state-of-the-art general-purpose LLMs struggle to identify untrustworthy comments, while iterative action-level learning substantially improves the accuracy of repository-grounded trustworthiness judgments. Our dataset is available at https://huggingface.co/datasets/dcloud347/CRJudgeBenchmark
Chinese Translation
大语言模型可以生成看似合理的代码审查评论,但此类评论可能包含技术上不正确的断言,从而误导开发者。我们研究技术可信度判断:确定一条审查评论的核心技术断言是否正确,并在其仓库上下文中适用于被审查的代码。现有的代码审查基准主要评估审查生成、问题发现或一般评论质量,但并未直接评估一个智能体能否判断单条审查评论的技术可信度。为填补这一空白,我们引入了 CRJudgeBench,这是一个由 1199 个实例构成的基准,构建自真实拉取请求和经专家验证的扰动,涵盖可信评论以及看似合理但不可信的评论。我们进一步提出了 Sentinel,一个以仓库为依据的智能体式评判器,它在做出判断之前主动收集代码证据以验证审查评论。从 Qwen3-Coder-30B-A3B-Instruct 出发,Sentinel 通过来自特权教师的迭代式动作级学习,在 CRJudgeBench 训练集上训练。在 359 个实例的 CRJudgeBench 测试集上,Sentinel 达到了 76.60\% 的准确率,超过 GLM-5.3 达 6.13 个百分点,并超过其基础模型 19.78 个百分点。这些结果表明,即使是当前最先进的通用大语言模型也难以识别不可信的评论,而迭代式动作级学习显著提高了以仓库为依据的可信度判断的准确率。我们的数据集可在 https://huggingface.co/datasets/dcloud347/CRJudgeBenchmark 获取。
cs.SE / 178 / 2609.37294
SafeLLM4SE: Statistical Evaluation and Reporting for LLM-based Software Engineering Systems
SafeLLM4SE:面向基于LLM的软件工程系统的统计评估与报告
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly used for software engineering tasks, yet their stochastic behavior challenges the validity, reproducibility, and comparability of their evaluations. Conventional practices such as reporting a single output, an average score, best-of-N, or pass@k performance can obscure variability and estimation uncertainty, potentially leading to misleading conclusions about system reliability. This article presents SafeLLM4SE, a practical methodology and reporting standard for statistically principled evaluation of LLM-based software engineering systems. Rather than treating generated outputs as deterministic artifacts, SafeLLM4SE treats them as realizations of a stochastic process and distinguishes quality, stability, and estimation uncertainty. It combines adaptive sampling with confidence intervals, distribution-aware statistical comparisons, effect sizes, and a minimum reporting standard covering model configuration, reproducibility, evaluation procedures, and resource usage. SafeLLM4SE is also provided as an open-source software package available on PyPI, enabling researchers and practitioners to reproduce and extend the methodology. We illustrate its application by comparing two LLMs on HumanEval, a benchmark of programming problems assessed through functional tests.
Chinese Translation
大型语言模型(LLM)正越来越多地被用于软件工程任务,但其随机行为对其评估的有效性、可复现性和可比性构成了挑战。诸如报告单个输出、平均分数、best-of-N 或 pass@k 性能等传统做法,可能会掩盖变异性和估计不确定性,从而可能导致关于系统可靠性的误导性结论。本文提出 SafeLLM4SE,一种用于对基于LLM的软件工程系统进行统计上严谨评估的实用方法论和报告标准。SafeLLM4SE 不将生成的输出视为确定性产物,而是将其视为随机过程的实现,并区分质量、稳定性和估计不确定性。它结合了自适应采样与置信区间、分布感知的统计比较、效应量,以及涵盖模型配置、可复现性、评估程序和资源使用的最低报告标准。SafeLLM4SE 还以可在 PyPI 上获取的开源软件包形式提供,使研究人员和从业者能够复现并扩展该方法论。我们通过在 HumanEval(一个通过功能测试评估的编程问题基准)上比较两个 LLM,来说明其应用。
cs.SE / 179 / 2609.37405
Complexity-Aware Evaluation of LLM Comprehension
LLM 理解的复杂度感知评估
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly used for software engineering tasks that require understanding existing source code, including behavior prediction, function explanation, debugging, and code review. However, aggregate benchmark accuracy can conceal how model reliability changes as source code becomes structurally more complex. This paper presents a complexity-aware framework for evaluating LLM code comprehension using cyclomatic complexity, nesting depth, branching factor, and Halstead volume. We evaluate DeepSeek-Coder-V2 and Llama through two complementary tasks: automatic input-output prediction over 300 Python functions and manually assessed semantic comprehension over a balanced subset of 60 functions. The functions are grouped into Low-, Medium-, and High-complexity bands. DeepSeek-Coder-V2 achieves an overall automatic accuracy of 78.33%, compared with 70.33% for Llama. However, accuracy decreases substantially from Low to High complexity, from 93.52% to 52.78% for DeepSeek-Coder-V2 and from 87.04% to 47.22% for Llama. Incorrect predictions are consistently associated with higher values of all four complexity metrics, and correlation and logistic-regression analyses confirm broadly comparable negative associations between structural complexity and correctness. Manual semantic comprehension shows the same degradation pattern, with accuracy decreasing from 100.00% to 75.00% for DeepSeek-Coder-V2 and from 90.00% to 60.00% for Llama. These findings demonstrate that complexity-aware evaluation provides a more diagnostic assessment of LLM code-comprehension reliability than aggregate accuracy alone.
Chinese Translation
大型语言模型(LLM)正越来越多地用于需要理解现有源代码的软件工程任务,包括行为预测、函数解释、调试和代码审查。然而,聚合的基准准确率可能掩盖模型可靠性如何随着源代码在结构上变得更加复杂而变化。本文提出一个复杂度感知框架,使用圈复杂度、嵌套深度、分支因子和 Halstead 体积来评估 LLM 的代码理解。我们通过两个互补任务评估 DeepSeek-Coder-V2 和 Llama:对 300 个 Python 函数的自动输入-输出预测,以及对 60 个函数的平衡子集进行人工评估的语义理解。这些函数被分为低、中、高复杂度区间。DeepSeek-Coder-V2 达到 78.33% 的总体自动准确率,而 Llama 为 70.33%。然而,准确率从低复杂度到高复杂度大幅下降,DeepSeek-Coder-V2 从 93.52% 降至 52.78%,Llama 从 87.04% 降至 47.22%。错误预测始终与所有四个复杂度指标的更高值相关,且相关性和逻辑回归分析证实结构复杂度与正确性之间存在大体相当的负相关。人工语义理解显示出相同的退化模式,DeepSeek-Coder-V2 的准确率从 100.00% 降至 75.00%,Llama 从 90.00% 降至 60.00%。这些发现表明,与仅使用聚合准确率相比,复杂度感知评估能够对 LLM 代码理解可靠性提供更具诊断性的评估。
cs.SE / 180 / 2609.37669
Retrieve, Reproduce, Reveal: Dissecting Retrieval-Augmented Software Vulnerability Detection
检索、复现、揭示:剖析检索增强的软件漏洞检测
large language model
大语言模型相关
Abstract
Retrieval-Augmented Generation (RAG) is increasingly used to enhance Large Language Model (LLM)-based software vulnerability detection by grounding predictions in retrieved vulnerability knowledge, such as vulnerability reports. However, existing RAG-based software vulnerability detection (RAG4SVD) systems are often evaluated using proprietary models, which challenges open science and reproducibility. Further, studies use different datasets, custom knowledge bases, different backbone models, and diverse metrics, which hinders meaningful cross-system comparison. In this work, we study six open-source RAG4SVD systems and address these reproducibility and comparability challenges through (i) reproduction of their experimental settings under an open-weight setting, and (ii) a unified benchmark using a common dataset, metric suite, and pool of open-weight models. Further, RAG4SVD systems typically consist of multiple components, yet are often evaluated only as a whole system, i.e., end-to-end. Therefore, we perform (iii) a component-level analysis that decomposes representative RAG4SVD pipelines into input abstraction, knowledge retrieval, and detection. Our results demonstrate that reproducibility varies substantially across systems. Under the presented unified benchmark, published RAG4SVD performance does not transfer under a controlled open-weight evaluation and depends strongly on the used model. The component analysis shows that effective RAG4SVD depends on the alignment between pipeline stages. For example, oracle knowledge raises retrieval to near-optimal, yet performance remains low (0.51 pairwise accuracy), demonstrating that retrieval effectiveness alone is insufficient for reliable detection. These findings motivate evaluating RAG4SVD not only end-to-end, but at the level of pipeline components, and provide a basis for more standardized, RAG-aware evaluation practices.
Chinese Translation
检索增强生成(RAG)越来越多地被用于增强基于大语言模型(LLM)的软件漏洞检测,其方式是将预测建立在检索到的漏洞知识(如漏洞报告)之上。然而,现有的基于 RAG 的软件漏洞检测(RAG4SVD)系统通常使用专有模型进行评估,这对开放科学和可复现性构成了挑战。此外,研究使用不同的数据集、自定义知识库、不同的骨干模型以及多样化的指标,这阻碍了有意义的跨系统比较。在这项工作中,我们研究了六个开源 RAG4SVD 系统,并通过以下方式解决这些可复现性和可比性挑战:(i) 在开放权重设置下复现其实验设置,以及 (ii) 使用共同的数据集、指标套件和开放权重模型池构建统一基准。此外,RAG4SVD 系统通常由多个组件组成,但往往仅作为整体系统(即端到端)进行评估。因此,我们进行了 (iii) 组件级分析,将具有代表性的 RAG4SVD 流水线分解为输入抽象、知识检索和检测。我们的结果表明,不同系统之间的可复现性差异很大。在所提出的统一基准下,已发表的 RAG4SVD 性能在受控的开放权重评估下无法迁移,并且强烈依赖于所使用的模型。组件分析表明,有效的 RAG4SVD 取决于流水线阶段之间的对齐。例如,oracle 知识将检索提升到接近最优,但性能仍然较低(0.51 成对准确率),这表明仅靠检索有效性不足以实现可靠检测。这些发现促使我们不仅在端到端层面评估 RAG4SVD,而且在流水线组件层面进行评估,并为更标准化、具有 RAG 意识的评估实践提供了基础。
cs.SE / 181 / 2609.37849
Is manual software optimization a thing of the past?
手动软件优化已成为过去时了吗?
large language model
大语言模型相关
Abstract
Scientific software is increasingly required to process larger datasets while maintaining acceptable execution times. Software optimization traditionally requires substantial expertise in programming, algorithms, and numerical methods. Recent advances in large language models (LLMs) offer the possibility of automating much of this process. We investigate whether LLM-based agents can autonomously achieve substantial performance improvements in scientific software, including mature implementations that have already been extensively optimized by human developers. We tasked an LLM-based agent with optimizing software for three computational problems: t-SNE, single-sample gene set enrichment analysis (ssGSEA), and graphlet counting. Humans defined the scope, correctness criteria, and a verification mechanism, after which the agent worked autonomously, in some cases for several hours. Code maintainers reviewed each resulting implementation and verified its correctness. The optimized implementations were faster in all tested configurations, by up to two orders of magnitude over the fastest existing tools. The improvements included low-level code optimizations, mathematical reformulations, and an entirely new algorithm for graphlet counting. Software optimization can increasingly be delegated to autonomous agents, with the human role shifting from implementing optimizations to deciding which software to optimize, defining objectives, providing verification mechanisms, and ensuring the correctness of the final software. For well-scoped, verifiable problems, we argue that manual software optimization may be a thing of the past.
Chinese Translation
科学软件越来越需要处理更大的数据集,同时保持可接受的执行时间。传统上,软件优化需要在编程、算法和数值方法方面具备深厚的专业知识。近年来大语言模型(LLM)的进展为将这一过程的大部分实现自动化提供了可能。我们研究了基于 LLM 的智能体是否能够在科学软件中自主实现显著的性能提升,包括那些已经由人类开发者进行了大量优化的成熟实现。我们让一个基于 LLM 的智能体针对三个计算问题优化软件:t-SNE、单样本基因集富集分析(ssGSEA)以及 graphlet 计数。由人类定义范围、正确性标准和验证机制,之后智能体自主工作,在某些情况下持续数小时。代码维护者审查了每一个生成的实现,并验证其正确性。在所有测试配置下,优化后的实现都更快,相比现有最快的工具最多快两个数量级。这些改进包括底层代码优化、数学上的重新表述,以及一种用于 graphlet 计数的全新算法。软件优化可以越来越多地委托给自主智能体来完成,人类的角色则从实现优化转变为决定优化哪些软件、定义目标、提供验证机制,并确保最终软件的正确性。对于范围明确、可验证的问题,我们认为手动软件优化可能已成为过去。
cs.CL / 182 / 2609.36290
The Surge of Anti-Semitism in German Social Media following the October 7 Attacks
10月7日袭击后德国社交媒体上反犹主义的激增
large language model
大语言模型相关
Abstract
We investigate the extent to which the Hamas attacks on Israel of October 7, 2023, have affected German social media debates about Judaism and Israel. For this, we develop an approach to detect 26 anti-Semitic categories in user postings via large language models (LLMs). The approach is applied to Facebook and Telegram posts (N=125,718) from three months before and after the event. Methodically, we test different open-weight models in two setups---with and without user information as additional context to the post text. The best setup achieves up to 83 % F1-score for binary anti-Semitism detection on our manually coded validation set. User context provides valuable information for most LLMs and drastically reduces false positives, for example, when (critically) reporting on anti-Semitic incidents. Concerning our topic, we find that anti-Semitism is surging significantly on both platforms, while being about ten times more prevalent on Telegram compared to Facebook. Facebook users express anti-Semitic views most likely in posts about an alleged genocide in Gaza carried out by the Israeli army, whereas classic anti-Semitic stereotypes related to power and conspiracy theories are dominant on Telegram. After the attack, the discourse patterns on both platforms show signs of convergence, as classic anti-Semitism increases on Facebook, whereas Israel-related categories surge on Telegram.
Chinese Translation
我们研究2023年10月7日哈马斯对以色列的袭击在多大程度上影响了德国社交媒体上关于犹太教和以色列的讨论。为此,我们开发了一种通过大语言模型(LLMs)检测用户帖子中26种反犹主义类别的方法。该方法被应用于事件前后三个月的Facebook和Telegram帖子(N=125,718)。在方法上,我们在两种设置中测试了不同的开放权重模型——有和没有用户信息作为帖子文本的额外上下文。最佳设置在我们人工编码的验证集上,对二分类反犹主义检测达到了最高83%的F1分数。用户上下文为大多数LLM提供了有价值的信息,并大幅减少了假阳性,例如,当(批判性地)报道反犹主义事件时。关于我们的主题,我们发现反犹主义在两个平台上都显著激增,而其在Telegram上的普遍程度约为Facebook的十倍。Facebook用户最有可能在关于以色列军队在加沙实施所谓种族灭绝的帖子中表达反犹主义观点,而在Telegram上,与权力和阴谋论相关的经典反犹主义刻板印象占主导地位。袭击之后,两个平台上的话语模式显示出趋同迹象,因为经典反犹主义在Facebook上增加,而与以色列相关的类别在Telegram上激增。
cs.AI / 183 / 2609.35958
Solver Agent: an Agentic AI Framework for Theoretical Physics Computations Applied to F-theory Uplifts of O3-planes and S-folds
Solver Agent:一个用于理论物理计算的智能体式AI框架,应用于O3-平面和S-折叠的F-理论提升
large language model
大语言模型相关
Abstract
We introduce Solver Agent, an AI framework based on large language models for calculations and proofs in mathematics and theoretical physics. The solution process is tracked through a persistent ledger that records assumptions, derivations, and computations. A central agent delegates tasks to specialized sub-agents, while independent agents verify both intermediate steps and the final result. This setup improves the traceability, reproducibility, and verification of computer-assisted calculations. Applying Solver Agent, we study global F-theory uplifts of Type IIB orientifolds and their S-fold generalizations. We establish sufficient conditions for Weierstrass models over projective threefolds with terminal $\mathbb{Z}_k$ quotient singularities ($k\in\{2,3,4,6\}$) to give $\mathbb{Q}$-factorial projective elliptically fibered Calabi-Yau fourfolds with isolated Gorenstein terminal quotient singularities. These geometries realize O3-planes and S-folds, where local D3-brane probes of the latter yield four-dimensional $\mathcal{N}=3$ superconformal field theories. Using stringy invariants, we derive fixed-point contributions to Hodge data and Euler characteristics, and show that these Euler corrections determine the localized D3-brane charges required for tadpole cancellation. We illustrate these results using toric hypersurface constructions, where a single three-dimensional polytope determines both the Type IIB Calabi-Yau threefold and the F-theory base; here, the orientifold double cover naturally forms a bisection of an alternative genus-one-fibered uplift with discrete $\mathbb{Z}_2$ gauge symmetry. Finally, we provide methods for toric computations and four-form flux analysis in four-dimensional $\mathcal{N}=1$ compactifications with non-abelian gauge sectors.
Chinese Translation
我们介绍 Solver Agent,一个基于大语言模型的AI框架,用于数学和理论物理中的计算与证明。求解过程通过一个持久化账本进行跟踪,该账本记录假设、推导和计算。一个中心智能体将任务分派给专门的子智能体,同时独立智能体验证中间步骤和最终结果。这种设置提高了计算机辅助计算的可追溯性、可复现性和验证性。应用 Solver Agent,我们研究 Type IIB orientifold 的全局 F-理论提升及其 S-折叠推广。我们建立了充分条件,使得在具有终端 $\mathbb{Z}_k$ 商奇点($k\in\{2,3,4,6\}$)的射影三维簇上的 Weierstrass 模型能够给出具有孤立 Gorenstein 终端商奇点的 $\mathbb{Q}$-阶乘的射影椭圆纤维化 Calabi-Yau 四维簇。这些几何实现了 O3-平面和 S-折叠,其中后者的局部 D3-膜探针给出四维 $\mathcal{N}=3$ 超共形场论。使用弦论不变量,我们推导出对 Hodge 数据和 Euler 示性数的固定点贡献,并表明这些 Euler 修正决定了蝌蚪抵消所需的局域化 D3-膜电荷。我们使用环面超曲面构造说明这些结果,其中单个三维多胞形同时决定了 Type IIB Calabi-Yau 三维簇和 F-理论底;在这里,orientifold 双覆盖自然地形成一个具有离散 $\mathbb{Z}_2$ 规范对称性的替代亏格一纤维化提升的双截面。最后,我们提供了用于环面计算和四形式通量分析的方法,适用于具有非阿贝尔规范扇区的四维 $\mathcal{N}=1$ 紧化。
cs.LG / 184 / 2609.36391
Receptive-field-constrained stimulus optimization for human early and intermediate visual cortex
用于人类早期和中期视觉皮层的感受野约束刺激优化
diffusion
扩散模型相关
Abstract
An ongoing challenge in sensory neuroscience is to characterize the feature dimensions encoded by cortical populations. Recent approaches probe feature selectivity in a data-driven way, by synthesizing a most-exciting-input (MEI) for a target neural population. While this approach has been successfully applied to human higher visual cortex using fMRI data, generating MEIs for early- and mid-level retinotopic visual areas requires additional modeling constraints due to small receptive field sizes. To address this challenge, we introduce two novel MEI generation frameworks, Receptive Field Diffusion for Visual Exploration (RF-DiVE) and Receptive Field Gradient Optimization (RF-GO). Both methods use a population receptive field (pRF)-constrained voxelwise encoding model; RF-DiVE combines this with a pretrained latent diffusion model, while RF-GO uses regularized gradient ascent. When applied to single voxels in retinotopically defined areas V1-hV4, using data from the Natural Scenes Dataset, we obtain MEIs that exhibit consistent structure within the pRF, suggesting selectivity for local features like contour, color, and texture. We systematically compare MEIs generated by RF-DiVE and RF-GO using two encoding backbones, performing in-silico validation of predicted responses to MEIs using independent encoding models. Across all methods and all visual areas, MEIs elicit higher model-predicted responses than the most activating natural images. We further find that the choice of generation framework and encoding backbone differentially affects MEI properties, including their visual appearance, structural interpretability, and cross-model generalizability. These results offer a new approach for performing data-driven characterization of spatial and feature selectivity across human visual cortex.
Chinese Translation
感觉神经科学中一个持续存在的挑战是刻画皮层群体所编码的特征维度。近期的方法以数据驱动的方式探究特征选择性,即为目标神经群体合成一个最兴奋输入(most-exciting-input, MEI)。尽管该方法已利用 fMRI 数据成功应用于人类高级视觉皮层,但由于感受野尺寸较小,为早期和中期视网膜拓扑视觉区生成 MEI 需要额外的建模约束。为应对这一挑战,我们提出了两种新颖的 MEI 生成框架:用于视觉探索的感受野扩散(Receptive Field Diffusion for Visual Exploration, RF-DiVE)和感受野梯度优化(Receptive Field Gradient Optimization, RF-GO)。两种方法都使用群体感受野(population receptive field, pRF)约束的逐体素编码模型;RF-DiVE 将其与预训练的潜在扩散模型相结合,而 RF-GO 使用带正则化的梯度上升。当应用于视网膜拓扑定义的 V1–hV4 区域中的单个体素,并使用 Natural Scenes Dataset 的数据时,我们得到的 MEI 在 pRF 内表现出一致的结构,提示其对轮廓、颜色和纹理等局部特征具有选择性。我们系统地比较了使用两种编码骨干网络由 RF-DiVE 和 RF-GO 生成的 MEI,并使用独立的编码模型对 MEI 的预测响应进行了计算机模拟(in-silico)验证。在所有方法和所有视觉区域中,MEI 所引发的模型预测响应均高于最具激活性的自然图像。我们进一步发现,生成框架和编码骨干网络的选择会以不同方式影响 MEI 的性质,包括其视觉外观、结构可解释性和跨模型泛化性。这些结果为在人类视觉皮层中开展空间和特征选择性的数据驱动刻画提供了一种新方法。
cs.LG / 185 / 2609.37579
Ornstein-Uhlenbeck Is Hard to Beat, Yet Superlinear Drift Ships Lower Transport Costs
Ornstein-Uhlenbeck 难以被击败,但超线性漂移带来更低的传输成本
diffusion
扩散模型相关
Abstract
Brešar and Mijatović \cite{bresar2025} show that Ornstein--Uhlenbeck diffusion is hard to beat in forward convergence under assumptions that exclude superlinear drift. We instead test superlinear Langevin diffusions for score-based image generation, computing their conditional scores numerically from a Fokker--Planck equation. In our experiments, the superlinear models beat the Ornstein--Uhlenbeck baseline on empirical Wasserstein distance across nearly the entire tested grid and show less variation across diffusion horizons. The ``hard to beat'' verdict of \cite{bresar2025} thus fails to be universal.
Chinese Translation
Brešar 和 Mijatović \cite{bresar2025} 表明,在排除超线性漂移的假设下,Ornstein--Uhlenbeck 扩散在前向收敛方面难以被击败。我们转而测试用于基于分数的图像生成的超线性 Langevin 扩散,并从 Fokker--Planck 方程数值计算它们的条件分数。在我们的实验中,超线性模型在几乎整个测试网格上,在经验 Wasserstein 距离上击败了 Ornstein--Uhlenbeck 基线,并在不同扩散时域上表现出更小的变化。因此,\cite{bresar2025} 的“难以被击败”这一结论并不具有普遍性。
cs.LG / 186 / 2609.37649
A Finslerian Approach for Embedding Directed Data
一种用于嵌入有向数据的 Finsler 方法
diffusion
扩散模型相关
Abstract
Many datasets carry an intrinsic directionality: citations point backward in time, cells differentiate along lineages, and traffic follows preferred routes. Spectral embedding methods, including most of their extensions to directed graphs, discard this information: they symmetrize the data and map it into a Euclidean space where asymmetry cannot be represented. We instead model directed data as sampled from a Finsler manifold, whose distance depends on the direction of travel, and study the kernel operator built from this asymmetric distance. Through a moment expansion of this operator, we show that its symmetric and antisymmetric parts separate geometry from direction. As the bandwidth of the kernel vanishes, the symmetric part converges to a weighted Laplacian, recovering diffusion maps in the Riemannian case, while the antisymmetric part converges to a first-order transport operator that encodes the directionality. We prove that the corresponding graph operators, built from finitely many samples, converge uniformly and almost surely to these limits. For Randers metrics, this vector field is explicit and yields an embedding algorithm recovering both the manifold structure, from the spectrum of the symmetric part, and the underlying drift. We illustrate the approach on synthetic directed graphs and point-clouds.
Chinese Translation
许多数据集带有内在的方向性:引用指向过去的时间,细胞沿谱系分化,交通遵循偏好的路线。谱嵌入方法,包括其大多数对有向图的扩展,丢弃了这些信息:它们将数据对称化,并将其映射到一个无法表示不对称性的欧几里得空间中。我们转而将有向数据建模为从 Finsler 流形中采样得到,其距离取决于行进方向,并研究由这种非对称距离构建的核算子。通过对该算子进行矩展开,我们表明其对称部分和反对称部分将几何与方向分离开来。当核的带宽趋于消失时,对称部分收敛到一个加权拉普拉斯算子,在黎曼情形下恢复扩散映射;而反对称部分收敛到一个编码方向性的一阶输运算子。我们证明,由有限多个样本构建的相应图算子一致且几乎必然地收敛到这些极限。对于 Randers 度量,该向量场是显式的,并产生一个嵌入算法,该算法从对称部分的谱中恢复流形结构,并恢复潜在的漂移。我们在合成有向图和点云上展示了该方法。
人工智能 (cs.AI)
227
cs.AI / 1 / 2609.36084
Measuring trainable degrees of freedom in materials graph neural networks: a random-subspace intrinsic dimension analysis
Abstract
Final predictive accuracy is the standard basis for comparing graph neural networks (GNNs) in materials-property prediction, but it does not show how strongly performance depends on access to trainable parameter-space directions. Here, we introduce trainable-degree dependence as a complementary characterization of materials GNN learning. Using random-subspace intrinsic-dimension analysis, we train CGCNN, ALIGNN, and DimeNet++ in randomly oriented parameter subspaces across six prediction tasks and measure how performance recovers as independent trainable degrees of freedom are restored. The resulting recovery curves separate endpoint accuracy from the trainable-dimensional demand required to recover it. They reveal distinctions that final errors alone miss: metallic classification and log-bulk-modulus regression recover near-reference performance from small fractional subspaces, formation-energy and band-gap prediction show stronger architecture dependence, and phonon prediction is most sensitive to dimensional restriction. Dataset-size sweeps show that band-gap models require larger fractional subspaces as training data grows, whereas formation-energy and bulk-modulus responses are more stable. A width sweep shows that fractional thresholds can remain stable while absolute threshold dimensions increase with model size. Random-subspace analysis therefore provides a targeted stress test for how materials GNNs use their optimization space.
cs.AI / 2 / 2609.36052
PowerZooJax: A JAX-based Power System Benchmark for Reinforcement Learning
Abstract
Power system operation is a safety-critical sequential decision-making problem, making it a natural testbed for reinforcement learning (RL). However, existing RL environments for power systems are often narrow in scope and computationally limited by CPU-based simulation workflows, making large-scale evaluation difficult. We introduce PowerZooJax, a JAX-based benchmark suite for RL in power system operation. It provides five constrained Markov decision process tasks spanning generation, transmission, distribution, distributed energy resources, and data center microgrid. By rewriting power flow, economic dispatch, market clearing, and device dynamics as JAX computation graphs, PowerZooJax keeps the entire training and evaluation loop on the GPU. Experiments show substantial speedups over CPU-based simulations and demonstrate standardized evaluation of policy returns, safety violations, and out-of-distribution stress conditions. Our open-source benchmark is available at: https://github.com/powerzoojax/PowerZooJax.
cs.AI / 3 / 2609.36056
GeoWind2Plan: Mission-Time 3D Urban Wind Prediction for Energy-Efficient UAV Planning
Abstract
In urban low-altitude flight, buildings reshape ambient wind into spatially varying 3D flow, making unmanned aerial vehicle (UAV) energy depend on local wind exposure as well as path length. However, building-resolved wind information is rarely available when a mission must be planned. Computational fluid dynamics (CFD) can produce high-fidelity urban flow fields, but each simulation is tied to a fixed inflow boundary condition and can take hours to days, which is incompatible with urban UAV missions that typically last minutes to tens of minutes. We present GeoWind2Plan, a geometry-to-wind-to-planning framework for mission-time 3D urban wind prediction and energy-efficient UAV planning. Given only a background wind vector, 3D building geometry, and a start-goal pair, GeoWind2Plan transforms the building geometry into a reference-wind frame, predicts mission-relevant 3D wind patches with a localized geometry-conditioned neural operator, stitches them into a queryable local wind field, and optimizes a feasible 3D path and speed profile using a physically grounded UAV energy model. Rather than pursuing CFD-perfect reconstruction, GeoWind2Plan targets decision-useful wind prediction: trajectories are planned with predicted wind and evaluated under high-fidelity CFD wind. Across held-out urban domains, wind speeds, and mission wind-angle regimes, GeoWind2Plan performs corridor-localized wind inference in about 3 seconds, compared with roughly 8 hours for CFD. Under CFD evaluation, trajectories planned with GeoWind2Plan reduce energy by 6.9%, 12.7%, and 4.5% in tailwind, headwind, and crosswind missions relative to wind-agnostic planning, recovering 87.9%, 85.7%, and 75.0% of CFD-reference savings. These results show that fast, corridor-localized 3D urban wind prediction can make wind-aware UAV energy planning practical at mission time.
cs.AI / 4 / 2609.36062
SMat-Attention: Structured Long-Context Sequence Modeling
Abstract
Long-context sequence models face a fundamental tradeoff: softmax attention uses flexible token-level interactions at quadratic cost, whereas linear attention obtains linear-time training and constant-time decoding by compressing history into a fixed-size state. In this work, we ask whether we can connect these regimes through a tunable notion of structure. To this end, we introduce Structured Matrix Attention (SMat-Attention) via a family of causal masks with structured long-range routing whose row supports have VC-dimension $d$. In our construction, $d=1$ recovers the standard causal mask, and increasing $d$ permits richer subset-routing patterns. We give chunkwise forward and backward algorithms to enable hardware-efficiency. For sequences of length $T$, the hard-routing construction takes $O(T^{2-3/d}+T)$ work, despite the mask being dense, for our prescribed family. In fixed-horizon streaming, decoding after the distant prefix takes constant time per token using $O(T^{1-1/d})$ cached states. SMat-Attention therefore makes VC-dimension an explicit knob governing access-pattern complexity, prefill cost, and decoding memory. Empirically, subset-routing and rule-assisted multi-key retrieval experiments illustrate the masks' routing expressiveness. Extensions to Mamba-2 and Gated DeltaNet using learned routing with top-$k$ query reads retain subquadratic prefill, improve recall accuracy over the backbones in several settings, and achieve comparable small-scale language-modeling performance.
cs.AI / 5 / 2609.36071
LongCat-DeepResearch Technical Report
Abstract
We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent workflow for producing comprehensive, evidence-grounded reports. The workflow separates global planning from detailed investigation and coordinates revision at the section level. Multiple planning agents first explore external sources and refine an actionable research plan, termed ResearchSpec. Research agents then investigate and draft their assigned sections in parallel, gathering additional evidence in separate contexts as their analyses develop. Once the sections are assembled, global review guides targeted local revisions, reducing reliance on repeated full-report rewriting. This workflow also supports the construction of research tasks and trajectories for the mid-training and post-training of LongCat's general-purpose models. LongCat-DeepResearch achieves 55.25 on DeepResearchBench, 51.35 on DeepResearchBench II, and 79.83 on ResearchRubrics. On an in-house benchmark, it scores 76.04, ranking second among four compared systems. Development-set analyses show benefits from combining planning perspectives, while further planning refinement has mixed effects. Additional editing improves average automatic readability preference across two benchmarks, with different trends on each.
cs.AI / 6 / 2609.36079
A Polyphonic Conception of AI Understanding
Abstract
When a doctor, a judge, or an engineer must decide whether to trust an AI model's output, they cannot avoid asking what the model understands. Purely mathematical or statistical descriptions struggle to distinguish trustworthy from untrustworthy outputs without reintroducing the question of AI understanding in all but name. Yet the question is ill-framed as it stands, because the inherited concept operates within a monophonic paradigm: the idea that a cognitive system's understanding of something must be localised to a single mechanism underpinning all the capacities conferred by such understanding. Drawing on a wide range of mechanistic evidence, we show that LLMs are pervasively polyphonic: outputs emerge from coalitions of parallel mechanisms of uneven reliability, which variously complement, duplicate, or drown out one another, with several coalitions sufficing for a task without any one being indispensable. Polyphony not only complicates attributions of understanding, but renders monophonic inference patterns hazardous. In response, we develop a conception of understanding fit for polyphonic AI. It centres on sound circuitry that is reliably and correctly recruited and in control of outputs. Attributions of understanding thereby become tractable claims about internal organisation, and can do the work of guiding trust in AI.
cs.AI / 7 / 2609.36082
GeoOutageBench: Benchmarking Ambiguity-aware, Ontology-grounded Geospatiotemporal KGQA for Multimodal Power Outage and Resilience Analysis
Abstract
We introduce GeoOutageBench, a benchmark for assessing LLM-based geospatiotemporal KGQA for multimodal outage and resilience analysis. Unlike existing KGQA benchmarks for Web knowledge, GeoOutageBench considers a spatiotemporal KG that integrates visual, textual, and structured data from outage records, remote sensing, weather observations, storm and power events, geographic entities, and domain ontologies. It provides a competency query taxonomy at different difficulty levels from spatiotemporal containment and proximity, spatiotemporal co-occurrence analysis, multimodal evidence, to hypothetical evaluation. Over multimodal KG and query classes, GeoOutageBench provides user-configurable evaluation of three important, highly coherent yet less studied tasks: (1) LLMs' understanding for ambiguous geospatiotemporal questions in terms of NL to SPARQL interpretation, (2) query-driven assessment of ontology utility, and (3) answer accuracy of multimodal KGQA retrieval. GeoOutageBench provides a design principle and foundation for assessing LLM-KG systems that support real-world infrastructure resilience analysis. Our benchmark, source code, data, results, and other documentation are available at https://github.com/UCF-SAGE/GeoOutageBench.
cs.AI / 8 / 2609.36104
An Exact Generate - Transform Decomposition of Small-LLM Team Scaling Across Orchestration Architectures
Abstract
Replacing one LLM agent with a collaborating team can raise accuracy, but whether scaling the team helps, and which architecture to scale, is unclear. Sweeping eight agent orchestration architectures across five instruction-tuned 7-9B models, five short-answer benchmarks, and an executable-code benchmark up to 30 calls, we find that the returns to team scaling are sharply task-dependent: from three to thirty calls accuracy rises by up to 17 points on the two arithmetic word-problem benchmarks (GSM8K, GSMHard) but by at most four on ARC, GPQA, and MMLU, for every architecture, a split the usual task-averaged number conceals. Proposer-Critic captures the arithmetic gains, scaling steepest and, in aggregate, surpassing every other architecture at the largest budget (item-clustered intervals exclude zero), though it ranks among the weakest elsewhere, and no architecture wins across tasks. We explain these trajectories with an exact generate-transform decomposition. Partitioning any workflow into proposal coverage and a downstream transform, any accuracy change splits exactly into an extensive coverage dividend and an intensive transformation change. The decomposition diagnoses each task: arithmetic offers coverage headroom that a critic-guided transform converts, whereas the multiple-choice benchmarks either saturate in coverage or fail to convert it, and on open-ended code generative recovery nearly vanishes so accuracy tracks coverage. At equal call budgets token cost still varies 2.1x. Extra calls therefore create candidate opportunity that only some architectures, on some tasks, convert. Team scaling is a task- and architecture-specific bet, not a uniform lever.
cs.AI / 9 / 2609.36118
The Layer Mystery of VLA: An Information-Theoretical Analysis of VLA Latent Interface
Abstract
Vision-language-action (VLA) policies connect a pretrained vision-language backbone to an action head through a latent interface, but which backbone layers this interface should expose remains unclear. We study single-layer selection and multi-layer fusion for frozen backbones across three pretrained models and two manipulation benchmarks, LIBERO and CALVIN, with three policy-training seeds per configuration. Across three fusion mechanisms and three layer-subset strategies, 47 of 54 configurations underperform the best observed single-layer policy. Our stastical analysis further confirms that fusion's advantage is very limited. However, the best layer varies substantially across backbones and benchmarks, making layer selection consequential and exhaustive policy sweeps expensive. We further derive a reweighting equivalence between the proposed information-bottleneck objectives for action-conditioned InfoNCE and action prediction, motivating InfoNCE as a proxy for layer quality. Empirically, InfoNCE provides the most consistent positive association with policy success among four evaluated proxies. Selecting the layer with the highest InfoNCE score requires 9-33 times less GPU compute than exhaustive policy sweeps and reduces mean selection regret from 17.89 percentage points for deepest-layer selection to 3.71 points across six settings. Its mean regret is close to the 3.28-3.50 points achieved by fixed-layer heuristics optimized retrospectively using all six oracle sweeps, without requiring closed-loop evaluations during selection.
cs.AI / 10 / 2609.36119
AdaST: Adaptive Coupling for Spatial-Temporal Forecasting
Abstract
Spatial-temporal (ST) forecasting underpins many real-world systems such as traffic, climate, and energy networks. While existing methods implicitly assume strong spatiotemporal coupling, we observe that real-world ST data exhibits distinct coupling regimes, ranging from temporal-dominated and spatial-dominated to strongly coupled patterns. This mismatch causes current models to suffer from spurious dependencies and degraded performance when one correlation dominates. To overcome this limitation, we aim to dynamically modulate spatial and temporal modeling based on the data's inherent coupling structure. However, three key challenges exist: unknown coupling structure, heterogeneous coupling dynamics, and suboptimal spatial modeling. We propose AdaST, an adaptive ST forecasting framework that tackles these challenges through a decompose-recompose paradigm. AdaST factorizes inputs into components capturing different coupling patterns using heterogeneity-aware experts. Each component is processed by role-aligned modules, and a correlation-informed adaptive recomposer integrates them for final prediction. Extensive experiments confirm that AdaST significantly outperforms state-of-the-art baselines, validating the necessity of an adaptive approach.
cs.AI / 11 / 2609.36130
Memory Is a Derivation: The Distributed-Evidence Paradox in Long-Term Agents
Abstract
Long-running LLM agents compress past interactions into persistent memories that may be reused as premises for later tasks. This creates a distinct derivation problem: whether the memory actually follows from what the interaction history supports. Relevant evidence may be scattered across earlier interactions, while compression can introduce relations or event status that the history never established. A valid memory may therefore appear unsupported because its citations omit relevant evidence, while individually supported facts may be composed into a stronger statement the history never established. We characterize this problem through three coupled requirements: (1) Evidence scope; (2) Compositional validity; (3) Admission reliability. We therefore ask whether the interaction history available at write time supports what enters persistent memory. We introduce DerivAudit, a framework for auditing whether a memory is actually supported by the history available when it was written. The audit separates three questions: whether supporting evidence lies beyond writer-provided citations, whether the composed memory introduces unsupported meaning, and how write-time admission decisions affect later memory use. Across two natural memory corpora, audits using broader pre-write history recover support for nearly 60% of memories that appear unsupported from citations alone, while 17-21% remain unsupported after expansion. Yet broader evidence does not by itself make admission reliable: unsupported memories are still frequently admitted across verification models, and evidence expansion alone worsens it on two backbones.
cs.AI / 12 / 2609.36154
More Features Are Not More Evidence: Limits of Training-Free Human Activity Recognition with Jev
Abstract
General-purpose models promise sensor-based decisions without training a task-specific classifier, which could reduce the dependence of Human Activity Recognition (HAR) on labeled data. Yet it remains unclear whether such models can directly interpret deterministic descriptions of physical sensor signals well enough to replace or complement trained HAR models. We study this question using Jev, a fixed general-purpose probabilistic decision model, on 1,800 class-balanced accelerometer windows from WISDM, UCI341, and PAMAP2. Jev receives no labeled examples, retrieval context, or HAR-specific parameter updates. We evaluate three deterministic sensor representations and compare 5,400 Jev decisions with a generative baseline and three supervised HAR models. Jev remains far below supervised recognition, with its strongest representation reaching macro-F1 of 0.038, 0.118, and 0.089 across the three datasets, compared with 0.686 to 0.907 for the supervised models. More numerical features do not improve Jev. Instead, they reduce recognition on all three datasets, while augmenting the same numerical evidence with a deterministic semantic rendering partially recovers performance, although the experiment does not isolate semantics from the accompanying serialization and redundancy changes. Jev is fast and inexpensive to query, but its probabilities are not reliably calibrated for recognition. A post-hoc fusion analysis finds a small improvement on WISDM that does not replicate on UCI341 or PAMAP2. These results show that training-free sensor decisions depend not only on the information available in the signal, but also on whether the model can use the representation through which that information is exposed. The sensor-to-model interface should therefore be treated as part of the model evaluation rather than as a neutral preprocessing step.
cs.AI / 13 / 2609.36190
FigAct: Turning Scientific Figures into Active Canvases for Explanation
Abstract
Scientific figures are designed to communicate information visually, yet MLLMs typically explain them by translating their visual content back into text. This requires readers to manually map the resulting explanations back to the figure. Inspired by how people present visual information, we introduce FigAct, a framework that transforms static scientific figures into question-conditioned visual presentations by acting directly on their existing graphical elements. Like a human presenter, FigAct generates a sequence of short narrations, grounds each narration in the corresponding visual evidence, and applies visual actions to guide the viewer's attention. We develop a hierarchical search strategy for efficient element localization, reducing token usage by approximately 40$\times$. We further train FigAct-8B using three task-specific rewards for grounding accuracy, search efficiency, and rendering quality. We further build a human-verified benchmark from figures in real-world scientific papers to evaluate the ability of MLLMs to generate grounded visual explanations. Our results demonstrate the effectiveness of FigAct and show that treating scientific figures as presentation canvases makes explanations clearer and easier to follow.
cs.AI / 14 / 2609.36228
An Empirical Study and Assessment of EU AI Act Compliance Checkers
Abstract
The EU AI Act introduces extensive compliance requirements for organizations that develop, deploy, or integrate AI systems. Many of these requirements are directly relevant to security and privacy, while also addressing closely related issues such as data governance, transparency, accuracy, and robustness. However, stakeholders such as small-to-medium businesses and individual developers often lack the legal expertise required to interpret these obligations and translate them into engineering and governance practices. This disconnect creates challenges for implementing the EU AI Act and may lead to missing safeguards or misdirected development and deployment efforts. To address this, various automated EU AI Act compliance checkers (AIACCs) have emerged, claiming to streamline compliance assessments and provide practical guidance. In this paper, we present the first empirical study and assessment of AIACCs. We characterize 12 mainstream AIACCs across multiple dimensions, evaluate their legal coverage and alignment, and analyze checker-generated compliance reports for structure, determinacy, and actionability. We find that the quality of AIACCs varies significantly and that they currently can only serve as early-stage orientation tools. Specifically, we observe inconsistent interaction modes and user-friendliness, a tendency to overly simplify or omit key obligations, and a failure to provide determinate, actionable guidance. As a result, reliance on the current generation of AIACCs may foster a false sense of compliance. With our study, we provide a critical baseline of the current AIACC landscape. We further offer design principles for the implementation of more reliable compliance-support tools.
cs.AI / 15 / 2609.36235
MERID: Multimodal Exploration via Recursive Self-Improvement Agents for Major Depression Analysis
Abstract
Major depressive disorder (MDD) severely impacts daily activities and quality of life. Detecting MDD involves multimodal data, such as interview recordings and sensor measurements. This is particularly challenging, as these heterogeneous modalities often demand distinct, customized prediction pipelines. Existing efforts to address this challenge have explored both manually engineered multimodal architectures and agent-assisted pipeline development. Despite their progress, it remains challenging to autonomously revise pipelines based on experimental feedback and carry verified improvements forward into subsequent designs. To this end, we propose Multimodal Exploration via Recursive Self-Improvement Agents for Major Depression Analysis (MERID). The framework develops depression pipelines through experience-based recursive self-improvement (RSI). Grounded State Construction (GSC) grounds experience by aligning multimodal records with subject-level depression targets. Coupled Pipeline Exploration (CPE) jointly modifies representations, fusion, and predictors to build successor pipelines for classification and severity estimation. Evidence-Guided Evolution (EGE) guides revisions through feedback and verifies gains under uncertainty in small depression cohorts before inheritance. Extensive experiments on depression benchmarks show that MERID achieves the best results on multiple tasks compared with multimodal and agent-based baselines. Further analysis highlights the value of acoustic and linguistic cues for depression detection. Our code is available at https://github.com/DiscoAILab/MERID
cs.AI / 16 / 2609.36238
ChronoSRL: Temporal Geometry for Self-Supervised Reinforcement Learning
Abstract
A goal that is close in space can be far away in time. Obstacles, terrain, and the agent's own capabilities determine how long it takes to get there. Yet, critics in contrastive and survival reinforcement learning do not measure the distances in their representation space in units of time. We therefore introduce ChronoSRL, which gives the critic's embeddings an explicit temporal geometry. The distance between state-action and goal embeddings is trained to match the time that the agent takes to reach the goal (goal-reaching time), while goals that were not reached, and goals from other trajectories, are pushed at least one discount horizon away. Furthermore, reaching a goal quickly once does not mean that reaching it is reliable in general, so the policy should not follow the temporal distance directly. Instead, we build on survival reinforcement learning and predict from our temporal embeddings not only the full distribution of goal-reaching times but also the time spent near the goal. Thereby, the policy is trained to favor actions that reach the goal sooner and more reliably and that keep the agent near it. ChronoSRL learns faster and reaches higher performance than contrastive, action-chunked contrastive, and survival reinforcement learning baselines on seven standard locomotion and navigation benchmarks, even with much smaller networks. To test the limits of self-supervised reinforcement learning, we introduce velocity tracking, goal-position reaching, and box climbing tasks with a quadruped robot in a realistic sim-to-real locomotion setup, and show how the shaping terms that are typical for robotics can be naturally incorporated into our framework. ChronoSRL is the only one of the tested self-supervised reinforcement learning methods that learns to stay at the commanded velocities and goal positions, and climbs the highest boxes.
cs.AI / 17 / 2609.36254
Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models
Abstract
Large Reasoning Models (LRMs) are commonly trained with reinforcement learning (RL) to improve their generation of chain-of-thought (CoT) reasoning before producing final answers. However, RL rewards are typically assigned based on final answers, providing little or no direct supervision over intermediate reasoning. This can lead to deceptive safety alignment, where the reasoning trace and final answer convey inconsistent safety signals. To systematically investigate this phenomenon, we introduce DSAR (Deceptive Safety Alignment Rate), a metric that jointly assesses reasoning traces and final answers to quantify their safety inconsistency. Across multiple LRMs and benchmarks, we find that deceptive safety alignment is pervasive under standard prompting conditions and is substantially amplified under prefilling attacks. We further provide a hidden representation analysis showing that models exhibit stronger safety discrimination at the final-answer stage than during intermediate reasoning. To close this gap, we propose SARA (Safety-Aware Reasoning Alignment), an RL-based method that rewards both safety-aware reasoning and safe final answers, encouraging early harmful intent recognition and enforcing reasoning-answer consistency. Experiments show that SARA significantly mitigates deceptive safety alignment under both standard and adversarial settings while preserving helpfulness and utility. Code is available at https://github.com/xzhou98/SARA.
cs.AI / 18 / 2609.36277
From Surfaces to Volumes: Registered Geometry for Protein Representation Learning
Abstract
Existing protein geometry models typically represent molecular surfaces using local geometric features such as sampled points, normals, and curvature. While effective for capturing exposed molecular shape, these representations do not explicitly model the volumetric organization beneath the surface or provide a consistent coordinate system for residue-wise volumetric structure. We introduce Protein-TetSphere, a registered residue-wise volumetric representation for proteins. Each protein chain is tetrahedralized to obtain local volumetric regions associated with individual residues, which are then registered to a shared fixed-topology tetrahedral reference and represented in a common Laplacian basis. This registration establishes consistent volumetric coordinates across residues, enabling local three-dimensional deformation to be integrated with surface and chemical information in a multimodal protein representation. We evaluate Protein-TetSphere on ligand-binding pocket classification, protein--protein interface prediction, and de novo protein binder design. Across the three tasks, Protein-TetSphere improves ligand-binding pocket balanced accuracy from $0.795$ to $0.826$, Pinder-Pair/Site AUROC from $0.914/0.852$ to $0.932/0.866$, and binder-design success from $14.95\%$ to $19.90\%$ on the BoltzGen Challenge Set and from $27.62\%$ to $32.19\%$ at the ProtDBench backbone level. These results show that registered volumetric geometry provides complementary spatial information beyond molecular surfaces across protein recognition, interaction, and design.
cs.AI / 19 / 2609.36278
Illusory Truth or Mere Exposure? Model-Dependent Repetition Effects in LLM-Based Social Media Simulations
Abstract
Generative agent-based models (GABMs) are increasingly used to simulate social media dynamics, including misinformation spread. For such social simulations to be valid proxies of human behavior, LLM agents should replicate established human cognitive biases, among them the Illusory Truth Effect (ITE), where repeated exposure to a claim increases its perceived truth value. We investigate whether and how the ITE manifests across four LLMs (Gemma-3-4b-it, Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct, and GPT-5-nano) in a social media simulation context. We propose a two-phase within-context experimental design that embeds the repetition manipulation inside a realistic news feed interaction. Using this design, we collect 336,000 truth, importance, sentiment, and interest ratings across 100 statements, 10 feed variants, and 3 replications. The key comparison is between ratings assigned to repeated statements, seen throughout a simulation phase, and completely unseen ones, rated within the same experimental context window. We distinguish genuine ITE (truth-specific repetition boost) from mere exposure effects. We run an OLS regression followed by a Linear Mixed-Effect Model to account for differences across models and ratings. Our results reveal four qualitatively distinct patterns: Gemma-3 exhibits a genuine ITE; Qwen2.5 shows a mere exposure effect; GPT-5-nano displays no repetition effect on truth and mild skepticism toward repeated content; Llama-3.1 shows a small truth boost alongside decreases in evaluative dimensions. Crucially, temperature has no effect on these findings, and a variance decomposition highlights the high context-sensitivity of LLM rating behavior. Our findings caution against assuming uniform ITE replication across LLMs in social simulations, while suggesting that Gemma-3-4b-it may offer the most behaviorally realistic approximation for misinformation-related simulations.
cs.AI / 20 / 2609.36308
CheatBench: Measuring Reward Gaming in AI Agents
Abstract
Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized information, attempted to evade monitoring systems, and even breached sandbox protections to attack external systems. As agents become more capable, this behavior could pose increasingly serious risks. To measure this problem, we introduce CheatBench, a benchmark of cheating in AI agents across mathematical research, knowledge work, coding, visual tasks, and other domains. Its environments combine challenging assignments with opportunities to cheat, allowing researchers to study how agents pursue a goal when honest work is difficult. CheatBench supports comparisons across models and task categories, providing a testbed for measuring and reducing cheating as agents take on more consequential responsibilities. We publicly release CheatBench at https://cheatbench.ai
cs.AI / 21 / 2609.36323
Towards an AI Software Factory for Data Systems
Abstract
AI-assisted coding tools deliver significant acceleration of coding, but only limited impact across the end-to-end software development lifecycle (SDLC)--an Amdahl's law effect! In this paper, we discuss our progress towards building an AI SW Factory that accelerates all the stages of SDLC-Targeting, Coding, Reviewing, and Ops. The AI SW Factory produces a metadata exhaust that enables self-improvement by fine-tuning model weights and updating our World Model (a rich data substrate). We focus on Data Systems and the important class of Evolutionary Coding Tasks (i.e., those with a measurable objective to hill-climb) and report on 1) scaled deployments at Microsoft (tens of repositories) leading to 3x engineering efficiency above agentic coding and up to 22x token efficiency, and 2) several open challenges.
cs.AI / 22 / 2609.36326
PILLAR: Private Inverted-Index Lexical Lookup for Augmented Retrieval
Abstract
Retrieval-augmented generation (RAG) hands the user's query to whoever hosts the corpus. We propose PILLAR, a Privacy-Preserving RAG (PPRAG) system based on Private Information Retrieval (PIR) in which a client utilizes the k documents most similar to their query from a server-held and publicly known corpus to respond to their query, while the server learns nothing about the query, either its terms or its access pattern. Prior PPRAG constructions rely on dense retrieval alone, translating approximate nearest-neighbor search into many query-dependent rounds of PIR, and pay for it in both latency and retrieval quality. PILLAR instead performs private hybrid retrieval in two stages. A sparse stage issues a small, fixed number of PIR queries against a carefully designed index of precomputed BM25 scores, filtering the corpus down to candidates that share terms with the query without the server ever seeing which terms these are. A dense stage then fetches only those candidates' document embeddings and re-ranks them locally, avoiding the many costly PIR queries that private dense retrieval typically requires. We instantiate PILLAR with two protocols that trade latency against retrieval quality, each built on a different private rendering of lexical search. PILLAR-Bin bins posting lists into a hash table and is a single-round design that achieves lower latency than state-of-the-art private retrieval schemes. PILLAR-Tree turns block-max pruning into an oblivious tree traversal combined with cuckoo hash tables and achieves the highest retrieval quality at lower latency than state-of-the-art schemes.
cs.AI / 23 / 2609.36340
ThuRunel: Dynamic Decoupling for Structured Advisory Dialogue
Abstract
High-stakes advisory domains such as medical aesthetics, legal consultation, and educational planning exhibit a two-phase structure. The early phase requires empathetic elicitation and emotional support, and the late phase requires authoritative specialist judgment. Neither fully automated agents nor human junior consultants adequately address this structure at scale. We formalize the core design challenge as dynamic decoupling, asking how an AI advisory agent should decide what to ask, when to stop, what to resolve autonomously, and what to forward to the specialist. We present ThuRunel, an advisory agent combining a finite-state belief management framework, a chain-of-thought teacher synthesis protocol, and learned generation adapters. Against eleven baselines, ThuRunel achieves consistent improvements in elicitation completeness and specialist brief quality. ThuRunel is publicly deployed as a bilingual web application in which the same decoupling decisions operate from the client's side, grounded in a curated knowledge base that cites its sources in every answer.
cs.AI / 24 / 2609.36359
Better Nearest Neighbor Graph Indices via (Efficient) LLM-Guided Pruning
Abstract
Graph-based approximate nearest neighbor search (ANNS) is widely used for large-scale semantic search. Its indices are constructed primarily based on geometric relationships among embeddings of an input dataset (e.g., documents or images), rather than explicitly optimizing for semantic relevance. However, when using these indices for downstream query retrieval, performance is evaluated based on the semantic relevance of the retrieved results to the query. This creates a fundamental "geometry-semantic" mismatch between how the indices are constructed and how their retrieval results are evaluated. While existing LLM-based reranking methods can partially mitigate this mismatch at query time, they leave this underlying structural problem in the graph unresolved. We therefore propose LLM-Guided Graph Pruning (LGP), a general framework that addresses this mismatch directly by leveraging LLM reasoning to refine an existing ANN graph index itself. LGP identifies structurally "low-value" neighbors of nodes and replaces them with LLM-selected alternatives that provide useful semantic information while retaining desired geometric structures of the original graph, including sparsity and efficient navigability. Experiments on representative semantic retrieval benchmarks show that LGP consistently improves end-to-end retrieval performance over both vanilla greedy graph search and LLM-based reranking across widely used graph-based ANN indices such as DiskANN and HNSW.
cs.AI / 25 / 2609.36384
Support-Set Target Leakage in Relational Foundation Models during In-Context Learning: Impact, Detection, and Mitigation
Abstract
Relational in-context learning (ICL) conditions predictions on the labeled support examples and their linked tables, creating a failure mode when the support set contains target-derived features that are unavailable for the query. We formulate this problem as support-set target leakage, distinct from leakage during dataset construction, temporal splitting, or representation learning. Here, the target-derived (leaker) columns are present only in the labeled support set during relational in-context inference, while queries remain clean. We construct 14 synthetic leaker types, corresponding to 20 columns, spanning proxies with different noise levels, coverage, modalities, semantic transparency, and relational distances. We evaluate a frozen relational encoder with an ICL head on held-out RelBench databases and use Integrated Gradients (IG) to rank and remove suspicious columns. Our results show that the effect of support-set leakage varies across tasks and relational distances. Target-table leakers cause the clearest degradation, while one- and two-hop leakers are not consistently used by the model. IG ranks target-table leakers highly across datasets and partially recovers performance in settings where leakage has the largest effect.
cs.AI / 26 / 2609.36388
Persona Dosing: Calibrated Activation Steering for Graded Trait Control
Abstract
An activation-steering coefficient sets intervention strength, but requesting a particular degree of persona expression requires a behavioral scale. We study persona dosing: controlling a language model through a trait description and a requested mean intensity. PersonaDose specializes a shared, description-conditioned FLAS controller on persona responses, then calibrates its flow time against measured trait expression. Training responses are not paired with requested target intensities. Across Llama-3.1-8B, Qwen3-8B, and Gemma-3-4B, PersonaDose raises core-trait expression at the Persona Vectors coherence floor of 75 by 33.2, 18.3, and 17.8 points over contrastive activation addition. Calibration-selected settings retain an expression advantage on held-out questions, although the coherence floor does not hold for every trait there. Across seven trained traits, calibrated requests yield mean targeting errors of 4.7-6.2 points over 14-22 calibration-reachable targets out of 28 per model. These results separate the behavioral range learned by a controller from the accuracy of requests within that range.
cs.AI / 27 / 2609.36392
ARCagent: An Adaptive Retrieval Calibration Agent for Clinical Question Answering
Abstract
In diseases where clinical guidelines are incomplete, contested, or mutually contradictory, knowledge completeness and dynamic conflict-aware synthesis are two safety-critical properties that standard Retrieval-Augmented Generation systems do not provide. Therefore, we present \sysname, an adaptive retrieval calibration clinical question-answering agent for ME/CFS, a disease where diagnostic frameworks coexist and major guidelines actively contradict each other on treatment. ARCagent contributes three components. First, a 1,706-chunk, 10-source knowledge base with a structured inter-guideline conflict registry spanning all active ME/CFS diagnostic frameworks. Second, a conflict-aware retrieval calibration pipeline that re-ranks retrieved evidence using query-specific focus and conflict signals. Third, a benchmark scored by LLM-as-Judge, avoiding systematic underestimation averaging 10.1 percentage points caused by keyword matching. ARCagent achieves 95.3%, outperforming all base LLMs. Code is available at https://github.com/Yukyin/ARCagent.
cs.AI / 28 / 2609.36406
From Retrieval to Reasoning: Agentic Mechanism Prediction from Cell Painting Profiles
Abstract
Cell Painting is a high-content morphological profiling assay widely used for phenotype-based biological inference, with mechanism of action (MOA) prediction as a central application. Existing approaches largely formulate Cell Painting-based inference as representation matching, assigning predictions from nearby reference perturbations in morphological feature space. However, retrieved neighbors are often noisy and partially misleading evidence due to batch effects, non-specific cytotoxicity, phenotypic convergence, and source-dependent variability. We reformulate Cell Painting-based MOA prediction as a calibrated evidence reasoning problem, where retrieved neighbors are treated as uncertain observations that must be evaluated, compared, and sometimes rejected before supporting a mechanistic conclusion. We propose PhenoAIR, a reliability-aware multi-agent framework that maintains a candidate-centric evidence memory and performs controller-guided refinement over phenotype- and mechanism-side evidence. PhenoAIR uses offline reference-set calibration to weight evidence by source reliability, phenotype stability, and mechanism-level confusion. We evaluate PhenoAIR on a benchmark constructed from JUMP Cell Painting profiles and annotations, covering controlled, realistic, and discovery-oriented open-world MOA prediction settings. PhenoAIR outperforms representation-matching and LLM-based baselines across all settings.
cs.AI / 29 / 2609.36417
Support-Set Target Leakage in Relational Foundation Models during In-Context Learning: Model Dependence and Evaluation Reliability
Abstract
Relational in-context learning (ICL) uses labeled support examples and their linked relational context to predict labels for new queries. This creates a failure mode when target-derived features are present in the support context but unavailable for the query. We study this setting as support-set target leakage. We construct 20 controlled target-derived features that vary in signal fidelity, representation, semantic transparency, coverage, and zero-, one-, and two-hop relational placement, and evaluate them across 13 RelBench tasks and five relational ICL configurations that vary the ICL head, message-passing depth, pretraining cohort, or relational encoder architecture. We evaluate matched 0-hop, 1-hop, and 2-hop leakage settings, together with a Full leakage condition containing all 20 leaker columns. Within the tested configurations, target-table (0-hop) and Full leakage produce the largest aggregate deviations from clean evaluation, while higher-hop effects are often weaker, consistent with differences in effective exposure associated with temporal reachability, sampling, and aggregation fidelity. Leakage effects are strongly task- and model-dependent and can reverse relative conclusions between model variants even when aggregate changes are small. For leaker detection, we compare an Integrated Gradients (IG)-based screening method with mutual information (MI) and leave-one-column-out (LOCO) on a common Baseline subset. Ranking quality is strongest in the high-impact 0-hop and Full leakage conditions, but detector-based removal does not consistently restore the clean evaluation. A four-task rel-salt case study further shows the same evaluation concern with native-schema leakage candidates from the original relational schema. These results identify the support/query information boundary as an important component of reliable relational ICL evaluation.
cs.AI / 30 / 2609.36434
The Safety Operator: Modulating the Expression of Safety Instructions via Spectral Optimization
Abstract
Context tokens in a transformer-based language model can be absorbed into the model's weights as a multiplicative operator. We study this operator in the setting of safety instructions and show that influencing its dominant eigenvalue modulates how strongly the instruction shapes generation. We derive a Contrastive Safety Loss with a suppression weight that controls the tradeoff between emphasizing the safety instruction on harmful queries while suppressing it on harmless queries. Varying the suppression weight maps a relationship between the attack success and the over-refusal rates, supporting the hypothesis that the operator's eigenvalue acts as a continuous dial for the instruction's influence. Moreover, this relationship holds relatively independently of how the Safety Loss is parameterized, yielding Pareto-improved safety instructions for appropriate values of suppression weight.
cs.AI / 31 / 2609.36437
Bits Under ZK-LLM: Evaluating Zero-Knowledge-Friendly Quantization for Verifiable Private LLM Inference
Abstract
Zero-knowledge proofs are emerging as a promising approach for enabling private, verifiable LLM governance and auditing, where regulators, users, and auditors need to verify claims about training-data usage or LLM inference-time behavior, while model providers must protect proprietary model parameters. However, despite the growing interest in ZK-LLMs, the understanding of ZK-friendly quantization remains limited. This gap matters because in the ZK setting, quantization directly shapes the arithmetic structure, constraint complexity, and proving cost of ZK inference. ZK protocols operate over finite fields and incur costs that depend heavily on the number and type of arithmetic operations, nonlinearities, and lookup constraints. Understanding ZK-friendly quantization is therefore essential for making ZK-LLMs practical. In this work, we present the first systematic study of ZK-friendly quantization for LLMs. We first formalize the definition of ZK-friendly quantization, capturing the properties required for ZK proof generation. We then evaluate nine language models, including Qwen2.5-14B and the mixture-of-experts model Qwen3-30B-A3B, across a broad design space of weight, activation, and nonlinear lookup table precision. Our results show that activation precision is substantially more sensitive than weight precision, while nonlinear lookup approximations can become the dominant source of utility degradation. Also, we identify RMSNorm inverse-square-root lookups as a recurring bottleneck in several large models and recover near-baseline utility by selectively increasing precision only at the bottleneck. Finally, we show that reducing bit-width or lookup-table size does not necessarily yield proportional end-to-end proving savings, showing that conventional low-bit quantization heuristics do not directly translate to ZK proving efficiency and motivating operator-aware precision selection.
cs.AI / 32 / 2609.36473
Going Beyond State-Reaching: Learning Abstractions for Intrinsically Motivated Option Discovery
Abstract
Temporal abstraction via options can improve exploration in large environments. However, existing option discovery algorithms find subgoals that target all aspects of the state simultaneously. This state-reaching approach produces options that only apply in narrow regions of the state-space, eventually causing an explosion in the number of options that overwhelms the agent, and impedes progress on its primary task of reward maximization. We introduce an algorithm that instead identifies a small, relevant subset of features for each subgoal, yielding options that generalize broadly and accelerate exploration. Our approach learns abstract, transferrable options and achieves rapid exploration in three sparse-reward, image-based domains, including the Atari game MontezumasRevenge.
cs.AI / 33 / 2609.36478
Learning to Harvest Without Collapse in a Regenerative Commons: A Lagrangian Framework
Abstract
The tragedy of the commons poses a multi-agent safety problem: reward-seeking agents can deplete a shared resource, and cooperation among its users does not itself specify how much must be preserved. We make preservation an explicit requirement by formulating a regenerative commons as a constrained Markov game or a constrained multi-agent MDP with a designer-specified depletion budget. We develop a nonstationary Lagrangian framework that constructs a policy sequence from solutions of unconstrained games or cooperative control problems. Extending earlier time-average constructions, we introduce average-epoch solution concepts for reset episodes with discounted rewards and terminal costs. We prove a reward-independent feasibility certificate, cooperative feasibility and approximate optimality against feasible policy mixtures, and an extension to unbiased sampled costs. For self-interested agents, a constrained Nash certificate quantifies the price-dispersion term introduced by deviations that redistribute budget across epochs. Under the stated assumptions on solver accuracy and multiplier updates, these results give constrained policy-sequence guarantees using solutions of unconstrained problems. Experiments with constrained IPPO and MAPPO in a Gordon-Schaefer fishery examine how depletion budgets shape stock retention, harvest rewards, and price adaptation.
cs.AI / 34 / 2609.36481
Human-AI Collaboration: From Paradoxes to Patterns
Abstract
Evidence shows that humans and AI systems perform better together, by collaborating, than alone. This paper examines two key design dimensions of human-AI collaboration (autonomy and initiative) and explores the collaboration patterns that they generate. Documenting these patterns starts with identifying the underlying problems and solutions, followed by examining the internal tensions within the problems. The paper uses a paradox perspective to analyze those tensions. It describes a process for surfacing the tensions and mapping the underlying paradoxes. It also illustrates how the pattern descriptions can be derived from mapping these paradoxes. Finally, the paper documents four human-AI collaboration patterns: Instruction, Delegation, Assistance, and Co-creation.
cs.AI / 35 / 2609.36503
AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes
Abstract
Adding or removing a sounding object requires coordinated changes to visual content and sound while preserving the surrounding scene. Yet paired supervision for localized non-speech audiovisual editing remains limited, as visual and acoustic edits must target the same object and isolate its sound from overlapping sources. To address this gap, we introduce \textit{AVIOBench}, a dataset comprising 37.9 hours of paired audiovisual examples spanning 1{,}878 target-object names. AVIOBench links the visual presence and acoustic contribution of each target object through a shared identity and visual mask. Our automated pipeline uses visual grounding and cross-modal consistency to select target-sound removal candidates, then jointly refines the audiovisual pairs to improve perceptual quality and cross-modal consistency. Building on this dataset, we propose \textit{AVIO}, which adapts a pretrained text-to audiovisual generation model through source-conditioned feature modulation to jointly learn object addition and removal. A reference-frame curriculum gradually reduces reference conditioning during training, enabling one model to perform instruction-only editing with optional visual guidance. Quantitative and qualitative evaluations demonstrate effective audiovisual object removal and addition, with optional reference guidance providing appearance and placement control for addition.
cs.AI / 36 / 2609.36556
MAADBench: The Refreshable Paradigm for Anomaly Detection in Multi-Agent Systems
Abstract
Recent studies report that LLM-based multi-agent systems (MAS) fail at rates of 41%-87%, yet to our knowledge, no benchmark to date supports systematic anomaly detection (AD) for them. Building MAS AD benchmarks is hard because they must remain fresh as LLM systems evolve: tasks may leak into training data and thus be memorized by LLMs, traces and anomaly patterns expire as backbones evolve, and labels must be provided reliably for each refresh. To address these challenges, we present MAADBench (MA: multi-agent; AD: anomaly detection), the first refreshable MAS AD benchmark designed for diverse, evolving LLM backbones underlying the agents. MAADBench combines (1) sampled-and-coupled generative tasks over an approximately 10^37-task space to mitigate task leakage, (2) refreshable trace generation under configurable LLM backbones, and (3) automated provision of cost-free, deterministic step-level labels for fine-grained AD evaluation. Beyond offering the paradigm itself, we run MAADBench with five state-of-the-art LLM backbones and release the MAADBench-Full dataset with 5,200 step-labeled traces. Benchmarking 25 AD methods on the MAADBench dataset reveals substantial limitations in current approaches: they rely heavily on supervision, struggle with subtle MAS-specific anomalies, and lack robustness across LLM backbones. These gaps point to a rich research agenda for MAS-specific anomaly detection, with MAADBench providing a systematic and refreshable testbed for method development and evaluation. We open-source MAADBench-Full at https://huggingface.co/datasets/hww123/MAADBench-full.
cs.AI / 37 / 2609.36572
Visual sensitivity is not claim retractability: persistence-aware credit assignment for multimodal reinforcement learning
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has been extended to Large Vision-Language Models (LVLMs), and perception-aware methods further encourage policies to rely on visual evidence. Yet relying on the image does not guarantee that visual claims are supported by it. Before RL training, 27.81% of the correctly answered responses of Qwen2.5-VL-7B on four multimodal reasoning benchmarks contain at least one direct visual claim that the image does not support. Since outcome-level RL rewards each response as a whole, these claims inherit the positive credit of the correct answer. We introduce a fixed-rollout counterfactual diagnostic that re-scores the same response under an intervened image to separate Evidence-Function Sensitivity (EFS), how strongly the model's predictions change, from claim persistence, whether the model keeps supporting the same claim rather than retracting it. The diagnostic reveals Sensitivity-Persistence Decoupling (SPD): under DAPO and VPPO, EFS increases and claims become more retractable overall, yet unsupported claims become significantly more persistent, whereas GRPO raises EFS without this deterioration. We therefore propose Persistence-Aware Credit Gating (PACG), which attenuates positive credit for unusually persistent visual claims and leaves all other credit unchanged. It requires no supported/unsupported labels and adds no inference cost. On Qwen2.5-VL-7B, PACG raises the nine-benchmark average over three seeds from 58.1% to 59.9% with DAPO and from 59.8% to 60.9% with VPPO, while making unsupported claims more retractable. The gains extend to a larger model, a newer backbone, and the accuracy of HallusionBench also improves consistently. These results suggest that visual sensitivity and claim retractability are complementary dimensions of multimodal credit assignment.
cs.AI / 38 / 2609.36580
SafeCoEvo: Co-Evolving Safety Harnesses and Guards for LLM Agents at Test-Time
Abstract
LLM agents deployed in real-world environments continually encounter new tasks and safety risks, while execution feedback typically becomes available only after each task is completed. However, existing self-evolving approaches commonly rely on multiple rounds of optimization over fixed and repeatedly accessible task distributions, fundamentally differing from test-time adaptation in real-world deployment, where only experience accumulated from past tasks can be used to improve safety decisions on future unseen tasks. To address this limitation, we propose SafeCoEvo, a test-time Harness-Guard co-evolution framework for LLM agent safety that enables the external safety system to continually adapt from accumulated runtime experience. SafeCoEvo jointly improves two complementary safety capabilities at different timescales: S-Harness rapidly externalizes recent runtime experience into updatable explicit safety knowledge that can promptly influence subsequent tasks, while GuardVPO internalizes accumulated runtime safety experience over a longer timescale into parametric risk-judgment capabilities. By combining short-term rapid adaptation with long-term capability consolidation, SafeCoEvo continually improves the agent's safety capabilities, reducing the unsafe outcome rate by 10.05% while improving the task success rate by 12.15% over the strongest baseline, thereby achieving simultaneous gains in safety and task utility.
cs.AI / 39 / 2609.36581
MemEvo: Automatic Discovery of Streaming Video Memory Mechanisms
Abstract
Query-agnostic streaming video understanding requires vision-language models to continuously compress an indefinitely growing visual stream into a bounded memory before future queries are known. The performance depends critically on the memory mechanism--what observations to preserve, how to represent and consolidate them, and what information to retrieve when a query eventually arrives. Rather than designing a single memory architecture by hand, we formulate memory design as a search problem over executable memory programs. We introduce a lightweight domain-specific language that expresses memory mechanisms through structured primitives for representation, admission, retention, consolidation, budgeting, and retrieval, while enforcing causal and bounded-memory constraints. Although structured, the derived program space remains large and contains heterogeneous, conditionally dependent design choices whose effects can only be assessed via downstream execution. We therefore propose MemEvo, an LLM-driven auto-research framework that uses pretrained LLM as a semantics-aware proposal model to iteratively generate and refine candidate memory programs based on accumulated experimental feedback. At runtime, a deterministic evaluation pipeline validates and evaluates each candidate, while the underlying vision-language model remains frozen throughout discovery. We finally produce a training-free, bounded-memory mechanism. Extensive experiments on StreamingBench and OVO-Bench demonstrate strong streaming video understanding performance together with substantial context and inference efficiency.
cs.AI / 40 / 2609.36585
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Abstract
Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressively further up the chain, and removing parent-line attention stops the relay. A frozen-model measurement locates the last useful intervention layer within tolerance in three of four held-out models. Task-specific LoRAs also improve MuSiQue. Default answers therefore understate the computation accessible through a tiny edit. Code and an interactive demo are available at https://lunamos.github.io/stop-thinking-too-early/
cs.AI / 41 / 2609.36601
SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation
Abstract
On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative. We introduce SAKI (Supervision Allocation with KL-constrained Interpolation), which combines a KL-constrained teacher-guided rollout with maximal coupling and reuses realized accept/correction events to route token-level supervision. Accepted positions retain sampled-token reverse-KL supervision, while correction positions receive direct supervision on the teacher's highest-probability token. Under maximal coupling, the correction probability is exactly TV(p_t, q_t), so the same trust-region radius controls rollout deviation and upper-bounds intervention and specialized-supervision frequency. We further implement an engine-resident speculative verifier that preserves the exact-q trajectory distribution and coupling semantics while improving matched-workload rollout throughput by 4.22x. Across seven mathematical reasoning benchmarks, SAKI improves the matched teacher-guided baseline in Mean@8 and Pass@8 for both 1.7B and 0.6B students. Placement controls and fixed-prefix analysis further support correction-triggered routing as a conflict-adaptive supervision signal.
cs.AI / 42 / 2609.36611
Multi-Channel Mitigation of Source-Trust Shortcuts in Fact-Checking RL Agents
Abstract
Retrieval-augmented fact-checkers often receive a reliability label, such as HIGH or LOW trust, for each evidence source. These labels should adjust the model's confidence and its decision to search for more evidence, while the verdict should follow the evidence content. We introduce TrustSwap, a counterfactual test that swaps, lowers, or removes source labels while keeping every evidence text fixed, and measures its three output channels (the verdict, the confidence, and the search decision) separately. Across untrained and RL-trained models at two scales, three datasets, and two prompts, confidence and search respond to the labels as intended in 49 of 50 comparisons, yet a label change alone alters 4-23% of confident verdicts for Qwen3 models and up to 50% for an existing RL-trained fact-checker. Standard GRPO fine-tuning amplifies this shortcut at 8B in all six settings. To reduce it, we propose trust-swap augmentation (TSA), which trains GRPO on each claim with both its original and its label-swapped evidence under the same gold verdict. At 4B, TSA lowers the verdict flip rate by 7-35% (relative) in four of six settings, keeps accuracy and the intended confidence and search responses, outperforms reward-based alternatives in the main setting, and carries over to an unseen label-removal perturbation. An added consistency reward helps on the trained-on swap but not on unseen perturbations. At 8B, TSA's effect is not detectable, which makes scale the main open question.
cs.AI / 43 / 2609.36620
Neural Structural Reasoner: A Brain-inspired Architecture for Reasoning over Structured Knowledge
Abstract
Structural reasoning, the ability to recognize and make inferences over the relational structure between objects and concepts, is a hallmark of human cognition, yet prevailing methods often collapse relational topology into flat embeddings, cannot discover hidden structure and lack interpretability. We introduce Neural Structural Reasoner (NSR), a brain-inspired network that preserves relational structure directly in the connectivity and dynamics of coupled neuronal populations. NSR draws inspiration from three biological mechanisms: multi-layered architecture for encoding hierarchical knowledge, stable representations of entity and concepts, and path integration for input-driven state inference. At query time, NSR parallelizes computation over candidate relational structures and leverages confidence-weighted scores to perform link prediction. Across standard knowledge-graph benchmarks, NSR achieves competitive accuracy without leading on every dataset, and has lower reported training times than several neural baselines. Because reasoning is implemented through sequences of human-readable neuron activations, NSR affords native interpretability by tracking intermediate inference steps. The model further extracts latent relational hierarchies and compositional rules, demonstrating the brain-inspired architecture as an effective, efficient, and highly interpretable substrate for structural reasoning.
cs.AI / 44 / 2609.36624
DualTrack: Synchronized speech-gesture generation via symmetric coupling of pretrained priors
Abstract
Joint speech-gesture synthesis must coordinate two modalities despite limited paired data. Existing approaches often lack bidirectional interaction, have limited language coverage, or simplify body and finger representations. We present DualTrack, which couples pretrained speech and motion priors on a shared 12.5 Hz timeline. Causal adapters exchange previous-packet information, while current-state fusion coordinates the streams before they separately complete sixteen-codebook packets. We evaluate 43 BEAT2 recordings in four languages, with speakers held out from joint training and validation. On the shared English/Spanish inputs, without speech or motion prefixes, DualTrack achieves lower word error rate and full-motion Fréchet Gesture Distance, higher beat consistency and speech naturalness than the evaluated GELINA baseline.
cs.AI / 45 / 2609.36626
Semantic Projection for Continual Self-Evolution of Language Agents
Abstract
Language-model agents increasingly rely on persistent natural-language skills to adapt beyond their frozen model parameters. When a shared skill is repeatedly revised from a non-stationary, heterogeneous task stream, however, improvements for new tasks can overwrite procedures needed for earlier ones. In continual learning, Orthogonal Gradient Descent (OGD) addresses analogous interference by projecting a new-task gradient onto a subspace that locally preserves prior predictions. Natural-language skill revisions, however, have neither gradients nor a canonical vector space in which such a projection can be performed. We introduce \emph{Semantic-Scope Projected Evolution} (SSPE), which transfers the functional principle of gradient projection from parameter space to behavior space. SSPE treats an unconstrained skill revision as a proposed update, identifies acquired capabilities with which it may interfere, and uses the observed gains and regressions to construct a compatible revision rather than merely rejecting the update. This enables one shared skill to evolve across latent and recurring task contexts without exposing semantic domain identities to the evolution model. Across controlled synthetic streams and heterogeneous real-agent benchmarks, SSPE improves final cross-domain competence and mitigates forgetting relative to strong skill-evolution baselines. The evolved skill also retains the strongest average performance after transfer to a different executor model. These results establish semantic projection as a promising principle for stable and adaptive self evolution of language agents.
cs.AI / 46 / 2609.36630
Distilling Agentic Systems: A Roadmap across Models, Artifacts, and Harnesses
Abstract
Modern agents increasingly rely on memories, tools, and execution logic, so their competence extends beyond model parameters. This shift exposes a limitation of conventional knowledge distillation, which asks how a student model imitates a teacher model. We define Agent Distillation as the persistent transfer of task-solving knowledge from a teacher agent to a student agent. Our study organizes the field by where transferred knowledge is retained: within the model, as artifacts, through the execution harness, or across substrates. This perspective separates transfer evidence from its outcome and clarifies how knowledge moves between agent components. We develop an evaluation framework that relates retention to causal contribution and deployed utility. Together, these contributions establish a foundation for the reliable, maintainable, and safe development of increasingly complex agentic systems.
cs.AI / 47 / 2609.36652
RankBuffer: Efficient Ranking-Based Rewards for Open-Ended Generation
Abstract
Open-ended generation lacks canonical answers, making pointwise rewards difficult to calibrate for group-based reinforcement learning. Directly ranking same-query rollouts provides a more suitable relative reward signal, but existing ranking-based reward methods can incur substantial judging cost. We introduce RankBuffer, which maintains an ordered, query-specific buffer of previously judged responses as a reusable quality scale. Each rollout is first inserted into an anchor interval through an independent coarse judgment, after which only rollouts assigned to the same interval undergo local fine ranking. The resulting complete order is converted into bounded rank rewards, while boundary expansion, local refinement, and inactive-anchor pruning adapt the buffer as the policy evolves. Across four open-ended benchmarks, RankBuffer consistently outperforms all pointwise baselines. It also achieves nearly on-par performance with the strongest ranking-based reward baseline while substantially reducing judging cost. Ablations demonstrate the importance of both local fine ranking and anchor response content, while buffer analyses show that rollout-derived anchors progressively extend and refine the covered quality scale. These results establish response reuse as an effective approach to efficient relative reward construction.
cs.AI / 48 / 2609.36670
FineSID: Scalable and Efficient Semantic Identifier Learning for Generative Recommendation
Abstract
A critical prerequisite of generative recommendation is designing semantic identifiers (SIDs) that are both scalable to large item sets and efficiently learnable. Existing SID learning methods fundamentally rely on Top-1 hard assignment during vector quantization. While heuristic strategies -- such as clustering-based initialization or forced post-hoc collision resolution -- can artificially inflate codebook coverage, they often disrupt end-to-end semantic alignment and fail to address the underlying optimization bottleneck: sparse gradient propagation. In standard Top-1 assignment, gradients concentrate on a narrow subset of frequently selected codewords, leaving the majority inherently under-trained and causing severe SID collisions. To overcome this limitation natively without relying on complex initialization priors, we propose FineSID, a unified quantization framework that moves beyond Top-1 assignment by enabling fine-grained gradient propagation across the entire codebook. Instead of updating only a single selected codeword, FineSID distributes learning signals to all codewords in a soft, differentiable manner. This design promotes globally balanced codebook optimization while strictly preserving semantic consistency, effectively alleviating SID collisions and stabilizing training in large, high-dimensional codebooks. Extensive experiments on multiple public benchmarks demonstrate that FineSID is robust to initialization configurations and consistently improves both codebook utilization and recommendation accuracy. Our work provides a principled, initialization-agnostic solution for semantic identifier learning, advancing the practicality of generative recommendation.
cs.AI / 49 / 2609.36679
MLToolBench: Learning Tool-Augmented Agents for Machine Learning Development
Abstract
Machine learning engineering (MLE) agents have made substantial progress, but learning through ML experimentation remains costly in time and computation. Synthetic environments reduce these costs while introducing variations in data and experimental settings that require task-specific diagnosis. Access to diagnostic tools alone does not ensure that agents learn when to use them or how to act on their findings. We introduce ToolMLBench, a suite of executable tools for data inspection, code verification, and experiment diagnosis, together with an SFT and RL pipeline for learning their use. Diagnostic calls acquire evidence whose value depends on subsequent decisions, so final outcomes provide limited guidance on which calls to reinforce. We address this challenge with SPICE, which measures how privileged context changes the likelihood of a sampled tool action and uses this difference as a turn-level reward alongside the final outcome. We train on 80 synthetic tasks and evaluate on 25 in-domain and 10 out-of-domain tasks. Providing tool interfaces and descriptions alone yields inconsistent gains across unadapted models. With the same diagnostic interface, our training pipeline raises in-domain success from 24.8% to 52.4% for Qwen3-8B and from 35.6% to 69.2% for Qwen3.5-35B-A3B. The latter also improves from 31% to 48% out-of-domain, supporting learned diagnostic tool use on held-out sources and targets.
cs.AI / 50 / 2609.36705
JudgeProfile: Understanding and Steering Subjectivity in LLM Judges
Abstract
LLM judges are inherently subjective, often favoring different responses in pairwise comparison when neither option is objectively wrong. To study this subjectivity, we introduce JudgeProfile, a framework that dissects LLM evaluation into perception (how a judge compares two responses across specific attributes like clarity, correctness, and detail) and prioritization (how much each attribute influences the final choice). We curate SubjectiveSet, a dataset of 50,013 response pairs from 17 public data sources, evaluated by 21 LLM judges across 87 attributes. We find a hidden consensus in perception: judges frequently agree on attribute judgments even when their overall choices diverge. Building on this separation, we first characterize each judge's prioritization using attribute weights estimated from its own overall choices. These weights differ across judges even when estimated from the same attribute judgments. We then learn new weights from reference labels to adapt their decisions to a target evaluation standard. Reweighting perceived attributes improves average held-out agreement with reference labels from 66.48% to 71.97%, outperforming fine-tuning and rubric prompting. Our findings show that understanding and steering the subjectivity of LLM judges requires attention not only to what they perceive, but also to how they prioritize it.
cs.AI / 51 / 2609.36726
Can AI Scientists Change Their Minds? Prior-Evidence Conflict in Synthetic Universes
Abstract
Can a scientific agent distinguish a law it inferred from evidence from one it merely recognizes? We introduce Synthetic Universes, a controlled benchmark that pairs canonical famous worlds with matched twisted twins governed by nearby noncanonical mechanisms. We evaluate each reported law twice: by executing it on held-out continuations and transfer settings, and by independently checking whether it recovers the generating mechanism. In the current checkpoint of a pre-specified 60-cell study, 22 trials were graded and one additional run ended in infrastructure failure. Among 20 twin trials, 8 pass predictive verification while 5 recover the generator. The dissociation is bidirectional: six parsable outputs predict successfully while missing the mechanism, whereas three recover the mechanism but fail predictive rollout. Drag exhibits the first pattern (5/5 predictive pass, 1/5 mechanism recovery); Gravity exhibits the second (1/5 predictive pass, 4/5 mechanism recovery). Because matched famous controls, the corrected identifiability sweep, and the Evidence Ladder remain incomplete, we do not claim a confirmatory causal prior-conflict effect. Instead, the completed runs establish a narrower verification result: predictive adequacy and mechanism recovery are distinct scientific claims and require distinct tests.
cs.AI / 52 / 2609.36730
Can Agents Design Libraries for Agents?
Abstract
Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design. We evaluate the library through the correctness and simplicity of programs written by three user agents from different model families. The benchmark spans 242 expert-validated programming problems across 15 library-design tasks in four languages. On eleven of the fifteen tasks, agent designers reproduce the abstractions of the human-written production library. Downstream agents adopt agent- and human-written libraries alike but underuse them, reimplementing capabilities the library already provides. Our failure analysis finds that downstream agents write extra code mainly because agent-written libraries are rigid or hard to use, not because capabilities are missing. We also experiment with giving designers more prescriptive, agent-first guidance and having them test their library with subagents; this improves downstream scores and yields simpler programs. LibraryDesignBench provides both a testbed for evaluating library-design practices for agent users and an initial design baseline that improves downstream reuse.
cs.AI / 53 / 2609.36741
Distinguish or Homogenize: Last-Chance Policy Identification and Risk-Budgeted Recovery under Irreversible Resource Depletion
Abstract
Under irreversible resource depletion, an agent can spend resources to distinguish among latent fault models, or to change the system state so that the remaining models admit a common acceptable continuation--at which point further diagnosis becomes unnecessary. This distinguish-or-homogenize principle identifies a path that existing frameworks for identification, planning, and diagnosis do not make explicit: prior formulations treat the mapping from fault models to acceptable policies as a given, whereas LCPI makes it a function of the agent's own actions. We formalize this principle through Last-Chance Policy Identification (LCPI), where correctness is evaluated at the state the agent reaches rather than at the initial state. The Last Identifiable Margin (LIM) marks the feasibility boundary between distinguishing and homogenizing. For deterministic diagnostic graphs we provide the Exact-LIM recursion; for noisy finite-horizon recovery we propose Risk-Budgeted Compatibility Planning (RBCP), which searches a compatibility-aware frontier under a hard worst-case failure constraint. Across incident recovery on abstract microservice topologies and latent-damage navigation in MiniGrid, RBCP improves risk-feasible recovery while satisfying the failure budget. A sham control--cost-matched actions that preserve model incompatibility--eliminates the gain entirely, confirming that the benefit comes from changing which policies are acceptable for which models, not from extra search or additional budget.
cs.AI / 54 / 2609.36746
EASE: Behavior-Adaptive Skill Curation for Self-Evolving Agents
Abstract
Agent skills provide a lightweight mechanism for self-evolving agents to accumulate reusable procedural knowledge without updating model parameters. However, existing learned skill curators typically optimize curation without explicitly modeling downstream executor behavior. We show that this can cause systematic cross-executor degradation: curators trained with different executors perform best when paired with their own training executor, indicating that effective skill curation is executor-dependent. We formulate behavior-adaptive skill curation and introduce EASE, a framework that learns a single curator that adapts its decisions to different executor behaviors. EASE maintains an online behavioral profile of recent execution patterns and conditions the curator on this profile, the current trajectory, and retrieved skills to add, modify, or remove skills from an evolving repository. We train the shared curator jointly across multiple frozen executors with reinforcement learning, using retrieval-aware and behavior-aware temporal attribution to focus optimization on curation actions with observable downstream influence. Across ALFWorld, ScienceWorld, and WebShop, with executors ranging from Qwen3-8B/32B and GPT-OSS-120B to unseen Kimi K2.6, DeepSeek V4 Flash, and Gemini 3.5 Flash, EASE outperforms strong skill- and memory-based baselines without per-executor finetuning. EASE also maintains 34.5--41.0% fewer skills, improves skill retrieval by 36.3--38.7% and measured edit utility by 51.8--60.0%, and reduces deployment-time inference tokens by 9.1--14.5%. These results establish behavior-adaptive skill curation as an effective principle for building self-evolving agents.
cs.AI / 55 / 2609.36770
Emergent Specialization in Populations of Self-Supervised Collaborative Vision Experts Without a Shared Gate or Cross-Agent Gradients
Abstract
Can a population of neural networks develop a useful division of labor without a shared gate or gradients between agents? We study a setting where each network has its own weights, trains independently on the same heterogeneous data, and can ask another agent for help through a forward pass. Unlike mixtures of experts, where a jointly trained gate assigns inputs to experts, specialization here must emerge without central control. We test this in a small scale proxy for predictive visual pretraining. Initially identical agents are finetuned on an unlabeled mixture of six visual domains using masked prediction of frozen DINOv3 features. We measure specialization by asking whether the best agent for an input aligns with its latent domain, and utilization by asking whether responsibility is distributed across agents. We progressively remove central control, ending with DISCO (DIStributed COllaboration) where each agent locally selects a helper, reads its internal state through a gradient free channel, and rewards its router only for the improvement that help provides. Specialization emerges and is useful. Randomly routed populations underperform a single generalist, while semantically routed populations outperform it, showing that specialization rather than population size drives the gain. Specialization persists without a central router, and gradient free communication lets nonexperts exploit emergent expertise. In DISCO, a random agent helped by the expert matches the solo generalist, while experts surpass it, including on data outside the specialization mixture. Local routers select the emergent expert for 98% of inputs. These effects persist across population size, model capacity, data imbalance, and finetuning seeds, providing measurable evidence for the dynamics needed by decentralized predictive pretraining.
cs.AI / 56 / 2609.36777
Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes
Abstract
Frontier coding agents can now write and execute code that authors 3D environments, but whether they reliably understand 3D structure and precisely control scene state remains unclear. The generated 3D scene is a persistent, executable artifact: a convincing render can hide incorrect spatial relations, intersecting objects, or unintended modifications. We introduce Code4Scene, a benchmark of 190 Unreal Engine cases built from human-assembled scenes that evaluates coding agents on two complementary settings under a shared execution interface. Construction tests scene-level spatial reasoning from open-ended language specifications, where many realizations are valid; editing tests precise control of scene state, where the agent must recover the target scene from reference images while preserving everything else. Rather than scoring code or rendered views, Code4Scene evaluates the generated engine-native scene for task fulfillment, artifact integrity, and static physical validity, with edits additionally compared against withheld ground truth. Across 14 coding-agent configurations on the 95-case public set, construction and editing performance are strongly correlated but not interchangeable (Spearman $ρ= 0.78$): Claude Fable 5.1 leads construction, Gemini 3.8 Flash leads editing, and GPT-6 Astra narrowly leads overall. Spatial Composition is the weakest construction category for every agent, while editing remains imprecise: the best Repair F1 is only 0.527, and 35.8% of edits that fully recover the target still introduce unintended changes elsewhere in the scene. These results expose a gap between plausible 3D generation and reliable spatial reasoning and state control.
cs.AI / 57 / 2609.36781
Aperture: Merge-Consistent Rotary States for Compressed Tokens
Abstract
Token compression combines content from several positions, yet rotary position embeddings usually assign the merged token one coordinate. We ask what positional information must survive later merges. Aperture stores Fourier moments of the token's weighted support at the model's rotary frequencies. We prove that these moments have minimal real dimension among continuous states sufficient for the selected expected rotary interactions. Represented mass makes updates additive; attention normalisation remains a separate readout choice. Uniform intervals give a centre rotation times a sinc gain. We characterise when centres determine interval widths and construct matched examples where they do not. Numerical checks verify the weighted-support implementation. In trained temporal readers, compression transfer varies with gain calibration and feature placement. In a prespecified native video question-answering comparison, stored support reaches $65.63\%$ accuracy versus $67.12\%$ for the deployed merging rule. These results separate exact positional preservation under compression from downstream benefit.
cs.AI / 58 / 2609.36806
CAD-Native Transformer Operators for AI-Aided Engineering
Abstract
Modern engineering systems, from automobiles to aircraft, are designed by using precise, continuous parametric computer-aided design (CAD) models. Evaluating design changes through numerical simulation requires meshing the continuous geometry, a computationally expensive and often brittle process that can require manual intervention and replaces the continuous representation with a discrete approximation. Most neural surrogates accelerate the simulation, but inherit this representation gap by relying on meshes, point clouds, voxels, or other sampled approximations of geometry. We introduce CANTO, a transformer neural operator that maps directly from continuous CAD geometry to physical fields, without meshing the input geometry. We develop a theoretical framework for learning operators from geometric manifolds to function spaces of physical fields, representing geometry through sequences of parametric patches. CANTO instantiates this framework by directly tokenizing non-uniform rational B-spline (NURBS) patches from their control points, knot vectors, and weights, and predicts continuous surface and volume fields at arbitrary query locations. We evaluate CANTO on four automotive and aircraft aerodynamics industry benchmarks: AhmedML, WindsorML, DrivAerML, and HiLiftAeroML. CANTO achieves state-of-the-art accuracy on most evaluated surface and volume prediction tasks, including a 19.8% reduction in surface-pressure relative $L_2$ error compared with AB-UPT on HiLiftAeroML. Differentiability with respect to CAD parameters further enables gradient-based inverse design of designs. On AhmedML, CANTO identifies designs with 4.4 to 20.4% lower drag than the best dataset designs satisfying the same volume and lift constraints, with the improvements verified using the same CFD setup used to generate the original dataset.
cs.AI / 59 / 2609.36809
Geometry-Conditioned Fixed-Scaffold Encoders for Time-Warp Robust Sequence Retrieval
Abstract
Embedding-based retrieval is attractive for long sequence collections because each item can be encoded once and searched by nearest-neighbor ranking. The difficulty is that the objects being indexed are often observed under a noncanonical clock: cardiac cycles stretch with rate, speech changes with tempo, and sensor traces reach comparable states at different speeds. This paper studies a specific source of instability in patch-based encoders for this regime. If patch boundaries are chosen from signal geometry, then the tokenization can change under the same temporal deformation that the representation is expected to tolerate. We propose GeoPatch, a fixed-scaffold patch encoder that keeps token support independent of geometry and uses slope, curvature, acceleration, affine-residual, and confidence descriptors only as continuous conditioning variables. The design turns boundary variation into feature modulation: geometry can change the embedding through a controlled pathway, but it cannot change the number, order, or support of local tokens. We formalize this distinction through a mechanism-level stability analysis that separates boundary drift, affine timing variation, confidence-weighted geometry perturbation, and retrieval-margin effects. The same local tokens support global embedding retrieval and late-interaction scoring, so the scoring rule can be matched to the evaluation protocol. Across ECG, speech, and multivariate time-series retrieval tasks, GeoPatch improves early-rank retrieval under timing variation while exposing a clear trade-off between local surface matching and strict non-overlap retrieval.
cs.AI / 60 / 2609.36829
The Default Trap: Rethinking Plan Evaluation in Tool-Using LLM Agents
Abstract
An executor can respond strongly to a change in a supplied plan's priority while showing a small change in the same information-selection probability when a default-aligned whole plan is removed. We call the risk of interpreting the latter as weak responsiveness to alternative priorities the default trap. We compare paired plans that prioritize different information targets with a shared no-plan reference. An accounting identity relates these distinct behavioral contrasts. Across 3,200 decision windows on 160 selected Retail, Airline, and AgentDojo tasks, switching priorities strongly redirects two models' choices, while the two plan-versus-default contrasts differ. In 2,160 additional windows, reversing account-list order shifts default target selection by 63.3-98.3 percentage points; priority-switching effects remain 96.7-100.0 points in either order. A separate 3,240-window component study finds strong control under single priority sentences, with effects of additional text varying by group and direction. Finally, 1,080 full-task episodes yield observed success differences of -19.4 to +8.3 points relative to no plan. All Retail and Airline success intervals include zero; AgentDojo results describe four fixed application worlds. These findings support joint reporting of priority responsiveness, presentation-dependent defaults, and task success and cost.
cs.AI / 61 / 2609.36847
Automated Screw Planning for Reduced Pelvic Fractures Based on Statistical Shape Models and Deep Learning
Abstract
Percutaneous iliosacral screw fixation is an important minimally invasive treatment for unstable pelvic fractures. Because the sacroiliac region has complex anatomy and narrow screw corridors, the accuracy and safety of screw placement directly affect surgical outcomes. Accurate and reliable preoperative screw planning is therefore essential to improve surgical success and reduce intraoperative risks. Conventional preoperative planning typically requires surgeons to determine screw trajectories through manual measurements, a labor-intensive process that depends on subjective clinical experience. To address these challenges, we propose a fully automated pipeline for preoperative iliosacral screw planning in patients with pelvic fractures. Using patient-specific three-dimensional anatomy, the pipeline automatically identifies safe screw corridors and generates individualized insertion trajectories to support clinical preoperative planning. We evaluated the proposed pipeline on 200 clinical cases of pelvic fractures. Compared with conventional manual measurements, the safety margin of the safe insertion corridors increased by 2% across the four screw types, the mean planning time decreased by more than 90%, and the clinical acceptance rate reached 95%.
cs.AI / 62 / 2609.36855
When Upstream Messages Override Correct Answers: A Controlled Study of Multi-Agent LLM Collaboration
Abstract
Multi-agent LLM systems rely on message passing among specialized agents to accomplish complex tasks. However, an upstream agent may provide useful information or an incorrect answer that causes a downstream agent to override a correct answer supported by its own evidence. Prior work has not clearly separated the benefits of communication from the damage caused by incorrect messages. We study this problem with controlled experiments across five benchmarks and five receivers, keeping the downstream task and evidence fixed while comparing answers under three conditions: no message, the upstream agent's original message, or a message with the opposite conclusion. Our experiments reveal three key findings. First, messages often help when the downstream agent would otherwise answer incorrectly. Second, messages can also hurt: when the downstream agent would answer correctly without a message, an incorrect upstream message changes the answer in up to 32% of cases. Third, in 94% of audited harmful cases, the downstream agent copies the upstream's specific wrong answer--a pattern we term answer substitution. Removing unreliable messages recovers part of the lost accuracy, suggesting that communication should be selective based on upstream reliability and the evidence already available to the downstream agent.
cs.AI / 63 / 2609.36860
IronLLM: Forging Compact Edge-Native Language Models for Real-Time Embodied Intelligence
Abstract
We present IronLLM-0.6B, a 654M-parameter language model designed for efficient on-device inference. IronLLM-0.6B combines a hybrid attention architecture with X-MTP, a lightweight shared-KV multi-token prediction design that eliminates per-depth KV-cache replay and employs a lightweight verification head for rollback-free drafting, achieving a 1.48x decoding speedup. The model is pretrained on approximately 6.2 trillion tokens using a quality-oriented data pipeline and is further post-trained with Multi-Domain On-Policy Distillation to integrate capabilities from domain-specialized teachers. To better meet the low-latency requirements of on-device scenarios, IronLLM-0.6B adopts an Instruct-Only design. Evaluations show that IronLLM-0.6B achieves competitive performance relative to larger models such as Qwen3.5-0.8B and MiniCPM5-1B, while producing more concise responses on many tasks. We further present IronLLM-0.6B-Light, which replaces RMSNorm with Dynamic Tanh and simplifies several computationally expensive components to improve inference and quantization efficiency. Together, the IronLLM models provide an effective performance-efficiency trade-off for resource-constrained deployment.
cs.AI / 64 / 2609.36867
State Trace Rationale As Auxiliary Task in Reinforcement Learning
Abstract
We propose STRAT, an auxiliary task that trains deep reinforcement learning (RL) agents to predict a short textual trace of their own state. Inspired by human spatial navigation, the description combines landmark, route, and survey knowledge, tracking the agent's position, inventory, goals, and immediate progress. Environment rules generate this text online without human labelling. Our method adds a single auxiliary head to a standard policy. Across 60 sparse-reward XLand-MiniGrid tasks, STRAT solves complex environments where standard RL fails outright, while compacting state representations and preventing rank collapse. Beyond performance gains, the predicted trace provides a readable account of agent beliefs at every step for no extra cost.
cs.AI / 65 / 2609.36887
WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents
Abstract
Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in isolation, however, does not guarantee commensurate gains in model performance, because reliable learning signals depend on coherent interactions among all components of the agentic interaction system. To address this problem, we introduce WEFT (Whole-system Evolution For Tool-use Post-training), which couples scalable agentic interaction system construction, execution-driven self-evolution, and stable post-training. WEFT scales agentic interaction system construction across environment breadth, task complexity, and interaction diversity. Execution-driven self-evolution iteratively uses execution traces and state evidence to attribute failures and revise the responsible components, with fresh rollouts evaluating the changes and providing evidence for subsequent evolution rounds. For stable post-training at scale, WEFT addresses both optimization and execution reliability: prefix-preserving sampling retains verified progress and atomic-turn credit assignment localizes learning signals, while MegaMCP maintains isolated, recoverable state across concurrent rollouts over shared tool services. Extensive experiments across various models and benchmarks demonstrate the effectiveness of WEFT for tool-use post-training. WEFT-8B and WEFT-14B outperform all evaluated matched-size environment-scaling baselines on BFCL V4, $τ^2$-Bench, and Claw-Eval. In particular, WEFT-14B improves over Agent-World-14B by 6.41, 2.23, and 12.27 percentage points. WEFT-35B-A3B further extends these gains to more challenging long-horizon workflow benchmarks, including Toolathlon-Verified and AutomationBench.
cs.AI / 66 / 2609.36888
Beyond Sub-Gaussian Detector Scores: Robust Weighted Profile-Loss Change Point Detection for Human-LLM Text Segmentation
Abstract
Mixed human-LLM documents require locating authorship transitions from detector scores whose reliability varies across text units. Existing weighted mean contrasts are vulnerable to extreme scores, while directly replacing means with robust centers obscures how a misplaced boundary changes the population objective. We propose Robust Weighted Profile-Loss Change Point Detection (RWCP), which combines capped reliability weights, Huber profile gains, and narrowest-over-threshold search in reliability coordinates. Our key analysis expresses the population gap between a true and a displaced split as a merge cost, avoiding a closed-form solution for the nonlinear center of a mixed segment. Under explicit curvature, spacing, and dependence conditions, core RWCP recovers the number of changes and localizes their boundaries; its quadratic-loss limit recovers squared weighted CUSUM. We also study RWCP-R, a separately evaluated decoder that shares source centers across nonadjacent passages. Across five retrospective cached-score benchmark families, core RWCP reduces family-macro WindowDiff by 17.6\% relative to weighted change-point detection, and RWCP-R lowers it further. Boundary recovery improves most clearly for isolated changes, while both fixed configurations miss changes in collaborative and densely alternating text.
cs.AI / 67 / 2609.36896
HorizonFlow: Variable-Length Planning for Offline Goal-Conditioned RL
Abstract
Recent advances in generative planning have made trajectory inpainting a promising approach to offline goal-conditioned reinforcement learning. However, these methods typically specify the planning horizon before generating plan content, even though the appropriate horizon depends on the route itself. A horizon that is too short can force infeasible transitions, whereas one that is too long can introduce redundant motion. We introduce HorizonFlow, a hierarchical planner that treats plan length as an output of generation rather than a prescribed input. Its subgoal route planner guides its action-prefix controller through a sequence of latent subgoals. Both components combine insertion-based generation with flow matching to jointly generate continuous plan content and length, using the partially generated plan to guide token insertion. HorizonFlow reuses the resulting length information to select candidates and steer generation toward shorter plans without a separate learned value model. Across Maze2D, Multi2D, and OGBench navigation and visual manipulation benchmarks, HorizonFlow achieves the highest average performance among the compared methods.
cs.AI / 68 / 2609.36900
STAR-GRPO: Canonical Anchoring and Reliability-First Advantages against Representation-Dependent Reward Hacking
Abstract
Reward hacking occurs when policy optimization exploits a brittle reward interface or an overly permissive proxy objective, improving the training score without improving the underlying response quality. This phenomenon is amplified in group-relative policy optimization: an unsupported reward can shift the group baseline and alter the updates of other rollouts, while post-hoc or purely relative weighting cannot represent group-wide uncertainty. We propose \emph{Self-Tuned Anchored Reliability Group-Relative Policy Optimization} (STAR-GRPO), a reliability-first advantage estimator based on paired assessments of the same rollout. STAR separates the quality signal from its learning influence: score disagreement determines rollout reliability, relative reliability enters a self-tuned robust location--scale fit before group normalization, and absolute group reliability attenuates the resulting bounded advantage. The analysis establishes coordinate and second-moment bounds, characterizes exact centering through the weighted location equation, and gives reliability-dependent attenuation guarantees for outlying rewards. We evaluate STAR-GRPO in two complementary reward-hacking regimes. In token-interface exploitation, STAR prevents runaway optimization of the deployed-interface score while improving the canonical quality signal. In rubric-proxy overoptimization for medical reasoning, STAR improves independent semantic evaluation, narrows the proxy--judge discrepancy, and reduces overclaim while optimizing the same task proxy. Together, these results show that reliability-first normalization offers a principled way to limit unsupported reward influence on both group baselines and policy updates, while retaining the task reward as the optimization target.
cs.AI / 69 / 2609.36923
PrecogUI: Proactive GUI Agents via Pre-cognitive Simulation and Experience Retrieval
Abstract
Existing reactive Graphical User Interface (GUI) agents often fail in long-horizon, dynamic scenarios, where unexpected disturbances trigger attention-diverting and cascading failures. To address this, we propose PrecogUI, a pre-cognitive architecture that shifts the paradigm from reactive execution to proactive decision-making. Specifically, we design a Proactive Experience Pool (PEP), which caches recurring anomaly and success patterns as "state-action-result" tuples in a dual-memory repository. Furthermore, we introduce a Proactive Simulation Executor (PSE) that learns to forecast the next symbolic UI layout given a candidate action, enabling early anomaly avoidance and ranking candidate actions by predicted reliability. Finally, a Pre-cognitive Execution Controller (PEC) fuses these priors and predictions, prioritizes handling of foreseen anomalies, and ensures execution robustness through a closed-loop error correction mechanism. For robust evaluation, we develop AutoTraj, an automatic data-generation engine, to construct InterfereBench, a benchmark for long-horizon tasks with strong disturbances. Experiments demonstrate that PrecogUI surpasses state-of-the-art methods on InterfereBench while maintaining competitive performance on public benchmarks. The code will be publicly available.
cs.AI / 70 / 2609.36927
Neuro-Symbolic Computer Use: Learning Reusable Policies for Reliable and Efficient Execution
Abstract
Many computer tasks recur: the same workflow runs many times, with new inputs and from different starting states. Current computer-use agents re-plan every step of every run, which makes them costly and unreliable on such tasks. We introduce neuro-symbolic computer use, in which a recurring workflow is executed by a learned policy rather than re-derived by an agent on each run. The policy fixes the decisions that are stable across runs (ordering, variables, loops, and branches) in executable code, and delegates observation-dependent decisions, such as grounding and state checks, to neural models. We learn these policies with neuro-symbolic policy iteration: starting from one agent trajectory, it executes the policy, diagnoses failures with task-completion and step-level judges, and revises the code with a coding model informed by an agent's continuation from the point of failure, without access to the benchmark evaluator. Iterating on generated parameter and initial-state variants makes the policy reusable, and a pre-action verifier guards each state-mutating step at deployment. On OSWorld-Verified and ScienceBoard, the learned policies achieve the highest Pass^3 of all methods in all four settings, 3.6-15.8 points above the base agent, while cutting per-run cost by 15-217$\times$ and latency by 3.4-5.1$\times$. On OSWorld-Verified, policies built only on variants transfer to the held-out original tasks, exceeding AutoRPA by 8.6-17.5 points in Pass^3.
cs.AI / 71 / 2609.36932
Learn from the Gap: Differential-Aware Advantage Pruning with Adaptive Rollout Sampling for GRPO
Abstract
Recently, Group Relative Policy Optimization (GRPO) and its variants have been developed for policy optimization and demonstrated notable performance gains. However, these methods usually incur substantial computational overhead due to per-question multi-rollout sampling and repeated per-token probability evaluation across rollouts. Furthermore, low-information or highly homogeneous trajectories can degrade downstream learning signal efficiency, hindering model optimization and limiting final performance. To address these issues, we propose FastRL, a novel plug-and-play reinforcement learning framework that simultaneously improves training efficiency and the effectiveness of policy learning. Specifically, 1) We introduce an advantage-aware pruning strategy to selectively preserve high-advantage trajectories while maximizing inter-trajectory gradient diversity. 2) Then, we design an adaptive rollout sampling mechanism to dynamically adjust the sampling scale across different training stages based on historical pruning distributions, balancing exploration adequacy and computational efficiency. Experiments demonstrate that FastRL can be seamlessly integrated into GRPO, DAPO, and GSPO variants, achieving an average 2.07$\times$ training speedup on Geometry3K and GeoQA8K-R1V, along with an approximately 1.64\% improvement in average accuracy on visual reasoning benchmarks. Source codes will be available at https://github.com/Nicozwy/FastRL.
cs.AI / 72 / 2609.36934
VLALight: A Vision-Language-Action Model for Traffic Signal Control
Abstract
Traffic signal control (TSC) is essential for improving urban mobility and reducing congestion. Although roadside cameras are widely deployed at signalized intersections and provide rich visual observations of evolving traffic, existing TSC methods typically rely on manually engineered traffic states or separate perception modules, creating a gap between physical observations and control decisions. We present VLALight, the first vision-language-action (VLA) model for end-to-end traffic signal control from multi-view roadside videos. VLALight directly maps visual observations to coordinated signal actions through multi-target spatiotemporal traffic reasoning and topology-aware cooperative perception across intersections. To establish this capability, we develop a two-stage supervised cold-start training strategy for visual traffic understanding and signal decision-making, followed by cooperative agentic reinforcement learning that jointly optimizes local control and network-wide traffic efficiency. Furthermore, VLALight introduces adaptive fast and slow reasoning modes, enabling the policy to allocate deeper reasoning only when additional deliberation provides sufficient control benefits. Through balanced mode-aware rollouts and relative advantage optimization, VLALight learns to trade off decision quality and inference cost. Extensive experiments on seven real-world traffic-flow datasets across three urban networks demonstrate that VLALight consistently outperforms transportation-based, RL-based, and LLM/VLM-based baselines. Ablation studies validate the effectiveness of cooperative perception, network-level optimization, and adaptive reasoning. These results demonstrate the potential of VLA models for real-world physical traffic control. Our project is available at https://github.com/usail-hkust/VLALight.git.
cs.AI / 73 / 2609.36939
SCA: Spatial Credit Assignment for Reinforcement Learning of GUI Agents
Abstract
GUI agents automate tasks on digital devices by grounding language instructions in visual interfaces. Existing group-relative reinforcement learning improves GUI action prediction by comparing the rewards of multiple responses sampled from the same GUI state. However, binary evaluation treats spatially different failed clicks as identical and provides no relative signal when all sampled clicks fail. To address these limitations, we propose Spatial Credit Assignment (SCA), which uses the screen coordinates of sampled clicks to refine group-relative credit. Specifically, SCA predicts each held-out response's reward from the other responses in groups containing both successes and failures, then uses the prediction residual to adjust credit. When all sampled clicks fail, SCA instead orders them by distance to the annotated target. These spatial references are used only to construct the training update; the deployed policy remains unchanged. We evaluate whether this correction improves the policy update itself by comparing its error and directional alignment with the exact return gradient in a controlled synthetic study. Across GUI grounding and offline action-prediction benchmarks, SCA improves grounding across professional domains and achieves the strongest results among reinforcement-fine-tuned models on most action-prediction metrics, with consistent gains across the reported GUI suites.
cs.AI / 74 / 2609.36944
Dual-Channel Robust Group-Relative Policy Optimization via Advantage and Sequence-Weight Estimation
Abstract
Group-relative policy optimization relies on reward-derived advantages and sequence-level likelihood weights, both of which can be sensitive to localized outliers. Extreme rewards can collapse the contrast among clean responses after group normalization, while token-level log-ratio perturbations can alter sequence weights and clipping decisions. We introduce RoVR-GSPO, a dual-channel robust optimizer that addresses these failure modes separately. Its reward channel combines robust reference estimation with bounded residual credit, while its ratio channel uses differentiable SoftRoVR aggregation to construct robust sequence weights. We provide stability and efficiency analyses for both channels. Experiments on mathematical reasoning, long-context summarization, and tool-call annotation show consistent improvements over GSPO, while controlled perturbation studies demonstrate stronger robustness to reward contamination and token-ratio anomalies.
cs.AI / 75 / 2609.36984
REALHOP: Rethinking Multi-Hop Reasoning Evaluation via Behavioral Auditing
Abstract
Complex questions often require multi-hop reasoning that connects facts distributed across sources or distant regions of a long context through intermediate steps. Benchmarks commonly evaluate this ability with questions built around predefined reasoning chains, treating a correct answer as evidence that the intended composition was used. Yet answer correctness alone leaves open whether success depends on the evidence associated with each intended step: models may instead rely on memorized associations, shorter paths, or partial evidence. We examine this dependence using the Behavioral Necessity Rate (BNR), which measures how often targeted evidence removal prevents answer recovery on initially correct instances. Across five existing benchmarks, panel-mean BNR ranges from 16.6% to 48.9%, exposing a substantial gap between annotated structure and observed dependence. Guided by this diagnosis, we introduce REALHOP, a diagnose-construct-verify framework that rebinds entities, factorizes selected relations, adds complete competing paths, and places evidence at traceable locations. Structural and semantic checks precede freezing; behavioral interventions follow. On 790 paired MuSiQue questions, REALHOP raises panel-mean BNR from 27.4% to 94.4% while retaining high Full accuracy. It also yields high BNR on REALHOP-FRAMES and REALHOP-LONGBENCH. On 216 long-context questions, the matched multiple-choice spread across 16 models grows from 13.9 to 59.2 points and persists under repeated open-ended evaluation. Together, these results show that a conceptually coherent chain and a correct final answer do not by themselves establish multi-hop reasoning. Verifying that success depends on every intended hop is therefore as fundamental to multi-hop evaluation as measuring answer accuracy itself.
cs.AI / 76 / 2609.36986
CF-LoRA: Decoupled Factor Aggregation and Adaptation-Aware Client Clustering for Federated LoRA Fine-Tuning
Abstract
Federated LoRA fine-tuning enables parameter-efficient adaptation of pre-trained models without sharing private data, but suffers from two fundamental mismatches under heterogeneous client data: a structural aggregation mismatch caused by independently averaging LoRA factors, and a statistical collaboration mismatch caused by enforcing a single global adapter across divergent clients. To address these issues, we propose CF-LoRA, a clustered federated LoRA fine-tuning framework that combines decoupled factor aggregation with adaptation-aware client clustering. CF-LoRA first learns a globally shared $A$ factor while retaining personalized $B_i$ factors, then identifies clients with similar adaptation patterns based on the cosine similarity of their learned $B_i$ factors, and finally performs intra-cluster $B$-factor aggregation with a frozen $A$ factor. By decoupling LoRA factor aggregation, CF-LoRA preserves the low-rank structure and mitigates the structural aggregation mismatch, while adaptation-aware clustering promotes collaboration among clients with similar adaptation patterns and reduces negative transfer caused by statistical heterogeneity. Experiments on four language tasks and four vision datasets with RoBERTa and ViT show that CF-LoRA achieves the highest average accuracy in both modalities while communicating only one LoRA factor per optimization round.
cs.AI / 77 / 2609.36996
ImbalancE: Inference-Time Latent Search Against Degree Imbalance in Link Prediction
Abstract
Knowledge Graph Embedding models have been extensively used to learn representations of entities and relations in Knowledge Graphs for predicting missing links. However, the quality of the learned representations varies a lot across different areas of the graph. If previous research has loosely linked the problem to relation types or degree bias, we show that it is more widespread and it correlates with the degree imbalance of the entities in test triples. In particular, the prediction of a target entity that has a degree much smaller than the degree of the anchor entity is extremely problematic. This is critical in recommender systems and other use cases, where these triples represent important corner cases. To address this issue, we propose an inference-time latent search optimization method capable of significantly improving model predictions on the most imbalanced triples. Built on top of a pre-trained model, it explores the embedding space at evaluation time, blending known and out-of-band information to mitigate the degree imbalance bias. We show the value of our approach on imbalanced triples from common benchmark datasets, where we outperform conventional methods, opening the door to the successful adoption of Knowledge Graph Embedding models on these critical corner cases.
cs.AI / 78 / 2609.37012
CADOC: Cache-Aware Dynamic Object Context for Long-Horizon Agents
Abstract
For a long-horizon agent, context is the bottleneck: the history is resent with every request, the window caps task length, and reasoning degrades as the history grows. Replacing structured objects with compact retrieval Cards shortens the prompt and keeps the exact originals retrievable, but editing the history can break prefix-cache reuse, and prior recoverable methods time their edits by forecasts of future reuse or by preset intervals. We propose CADOC (Cache-Aware Dynamic Object Context), an online algorithm that replaces structured objects with compact Cards while preserving exact, on-demand retrieval of their original contents. CADOC schedules replacements in batches by balancing accumulated waiting cost against shared cache-reconstruction cost. Its scheduling rule follows from an economic order quantity trade-off, recovers the optimal integer batch under stationary assumptions. Across evaluation, CADOC consistently achieves the lowest aggregate input cost among the compared configurations, which reduces input cost by approximately 40\% on average while maintaining task performance close to full context. CADOC thus provides a cost-derived approach to compressible context management, demonstrating that efficient compression depends not only on shortening prompts but also on scheduling edits to preserve cache reuse.
cs.AI / 79 / 2609.37022
Physics-Informed Multi-Agent Coordination for Hospital Patient Flow Optimization
Abstract
Efficient patient flow coordination across autonomous hospital departments is critical for mitigating overcrowding and balancing resource utilization. While classical queueing theory, specifically open Baskett--Chandy--Muntz--Palacios (BCMP) networks, provides an interpretable mathematical topology for healthcare operations, analytical models rely on stationary assumptions and fixed routing matrices that degrade under state-dependent real-world dynamics. Conversely, centralized reinforcement learning approaches struggle to accommodate the decentralized structure of hospital governance, where individual clinical departments function with localized observations, heterogeneous resources, and divergent operational objectives. In this paper, we present a Multi-Agent Systems (MAS) framework titled \emph{Physics-Informed Multi-Agent Coordination}, which embeds empirically calibrated BCMP queueing topologies as physical priors within a decentralized multi-agent reinforcement learning architecture. Formulated as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) under coupled resource constraints, our method enables autonomous departmental agents to cooperatively negotiate patient routing and dynamic service scaling. To mitigate environmental non-stationarity without inducing excessive communication overhead, agents exchange localized action fingerprints along network edges and optimize a spatially decomposed reward structure. Empirical evaluations driven by real-world MIMIC-IV patient trajectories indicate that this cooperative multi-agent approach substantially reduces cumulative system delay compared to static Markovian approximations, heuristic dispatching, and independent multi-agent baselines, while maintaining clinical safety constraints.
cs.AI / 80 / 2609.37024
Language as the Interface: Foundation-Model Contrastive Learning Links Transcriptomes and Electrophysiology
Abstract
Integrating transcriptomic and electrophysiological data is essential for building multimodal foundation models for neuroscience. Patch-seq provides paired measurements of gene expression and intrinsic electrophysiology from the same neuron, establishing a basis for training cross-modal models. Here we introduce LangPatch, a foundation-model-based contrastive learning framework that uses paired Patch-seq data to align pretrained GenePT representations with electrophysiological phenotypes through a language-based interface. Gene descriptions and verbalized electrophysiological profiles are embedded by the same frozen text encoder. A context adapter and projection modules connect the modalities through paired contrastive learning. Across mouse visual, mouse motor, and human cortical cohorts, LangPatch achieves the highest mean transcriptome-to-electrophysiology prediction correlation among the evaluated foundation-model and representation-learning methods. It also improves held-out cross-modal alignment in the two mouse cohorts (FOSCTTM 0.107/0.135 vs. 0.208/0.222 for JAMIE, an existing cross-modal Patch-seq imputation method). It predicts transcriptomic family, type, cortical layer, and marker-gene expression from electrophysiology, exceeding other baselines on most endpoints. More importantly, the method transfers across brain areas and species: a model trained on mouse visual cortex predicts electrophysiology in motor cortex with approximately 70% correlation retention and in human cortex with 47% (58% on acute-slice recordings). Together, these results demonstrate alignment between molecular and functional representations of neurons, providing a building block for multimodal foundation models in neuroscience.
cs.AI / 81 / 2609.37027
Beyond Low-Rank Parameterization: Narrowing the Gap Between LoRA and Full Fine-Tuning via Gradient Decomposition
Abstract
Low-Rank Adaptation (LoRA) is a widely used approach to parameter-efficient fine-tuning (PEFT), yet a performance gap can remain relative to full fine-tuning (FFT). Many LoRA variants improve the initialization or optimization of low-rank factors. At each training step, however, their first-order weight-space directions are constrained by the current parameterization. We characterize the corresponding LoRA-accessible gradient space and show that it coincides with the tangent space induced by the current LoRA parameterization. This characterization yields an orthogonal decomposition of the full weight gradient at the current model parameters. We term the component orthogonal to this space the normal gradient. Based on this decomposition, we propose GDLoRA (Gradient-Decomposed Low-Rank Adaptation). GDLoRA reconstructs the full weight gradient from forward activations and backward signals, extracts its normal component, and directly updates the base weights with this component, while retaining standard AdamW optimization for the LoRA factors. GDLoRA incorporates complementary normal gradients without increasing standard LoRA's optimizer-state memory budget under matched adapter and optimizer configurations. Experiments on natural language understanding, mathematical reasoning, commonsense reasoning, and image classification show that GDLoRA consistently improves over LoRA and narrows the performance gap to FFT. The code is available at https://anonymous.4open.science/r/GDLoRA.
cs.AI / 82 / 2609.37033
FedLAFP: Low-Rank Aggregation Meets Full-Rank Personalization in Federated Fine-Tuning
Abstract
Federated parameter-efficient fine-tuning enables clients to adapt pre-trained models without sharing raw data or communicating the full model, but statistical heterogeneity makes a single global adapter insufficient for personalized prediction. Existing personalized methods typically use the same low-rank structure for both shared and private adaptation, overlooking their distinct requirements for aggregation and personalization. We propose FedLAFP, a role-aware framework that couples a compact, globally aggregated LoRA branch with a client-private, full-rank-capable RandLoRA branch. The shared branch provides an efficient interface for transferring common knowledge, whereas the private branch combines fixed random low-rank bases with learned scaling coefficients to provide expressive client-specific adaptation without additional communication. Client- and layer-specific mixing coefficients jointly fuse the two branches, and only the shared LoRA parameters are exchanged. A controlled linear study supports this role assignment: LoRA yields more aligned client updates and lower aggregation error, while RandLoRA more accurately recovers client-specific residuals. Experiments across four visual recognition benchmarks show that FedLAFP consistently outperforms local-only and federated LoRA baselines, achieving an average personalized accuracy of $86.93\%$ and exceeding the best baseline average by $1.30$ percentage points.
cs.AI / 83 / 2609.37035
Watch-Think-Interact: Bootstrapping Long-Horizon Multi-Turn Streaming Video Reasoning with Reinforcement Learning
Abstract
Streaming video assistance requires models to answer asynchronous questions from an observed prefix under a fixed context budget. Existing approaches model response timing or compress history, but an online state formed before future questions are known can omit visual details before later questions reveal their relevance; the retained state alone cannot recover them. We introduce Watch-Think-Interact (WTI), a closed-loop framework for multi-question streaming video reasoning. WTI maintains compact natural-language memory entries tagged with source-video time ranges; these entries support direct reasoning when sufficient and otherwise anchor selective recall of finer visual evidence. For each question, WTI answers when current context and memory suffice, continues watching when required evidence has not appeared, or recalls a relevant past interval and decides again after incorporating the returned chunks, without replaying the full observed history. To train this behavior, we construct WTI-82K, comprising 82,335 timed questions across 4,812 causally aligned trajectories, and develop Stream-GDPO to optimize complete multi-question streaming rollouts using trajectory-level feedback for response timing, source-video recall, and memory updates. WTI achieves state-of-the-art aggregate performance among the compared open-source streaming baselines, reaching 83.3% on StreamingBench and 73.6% weighted overall accuracy on OVO-Bench.
cs.AI / 84 / 2609.37053
MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows
Abstract
Multimodal GUI agents have achieved impressive results on general software benchmarks, yet their ability to operate professional scientific software remains largely unexplored. In materials science, sparse domain-specific web data, specialized interfaces, and tacit workflow conventions create blind spots that general-purpose pretraining cannot readily bridge. We present MatToolBench, the first real-environment benchmark for evaluating multimodal GUI agents on professional materials science software, comprising 204 tasks across 10 tools in three modalities: GUI operation, OriginPro scripting, and code-based database queries, all executed inside a Windows 11 VM. Each task is decomposed into fine-grained sub-criteria by domain experts, enabling interpretable partial-credit scoring; the GUI component of our multi-level evaluation pipeline achieves an average F1 of 0.98. For OriginPro figure-generation tasks, we further conduct a human-LLM agreement study to validate the use of a multimodal judge for secondary aesthetic assessment. Our experiments show that strong performance on general benchmarks does not transfer to professional scientific workflows, and that this gap is not a visual-grounding problem alone: failures arise from domain-specific operational knowledge, sparse pretraining coverage of scientific software, weak cross-tool artifact handoff, and critical states exposed only visually. Even the best model reaches only 25% success rate on GUI tasks and 45% on code tasks. MatToolBench therefore serves as a challenging diagnostic benchmark and real-environment testbed for data-scarce, knowledge-intensive scientific workflows.
cs.AI / 85 / 2609.37111
Learning from Viable Failure Prefixes: Milestone Viability Potential Policy Optimization for Long-Horizon LLM Agents
Abstract
Long-horizon LLM agents require reinforcement learning methods that can assign credit to intermediate decisions under sparse and delayed rewards. Existing group-based methods such as GRPO and GiGPO alleviate this issue by comparing rollout returns or repeated anchor states, but they still fail when the compared returns have no variation. We identify this failure mode as zero-credit failure: during early training, many failed rollouts contain useful prefixes, yet existing methods assign them no task-discriminative advantage. To address this issue, we propose Milestone Viability Potential Policy Optimization (MVPO), a potential-routed policy optimization algorithm that learns from viable failure prefixes. MVPO estimates prefix potential over Union-Find viability regions, repairs zero-credit groups with potential-difference advantages, and attenuates the potential branch according to relative performance progress. Experiments with Qwen2.5-1.5B-Instruct show that MVPO outperforms eight strong baselines, including GRPO and GiGPO. Under the same training length, MVPO improves over the GiGPO baseline by +4.4 success points on ALFWorld and +5.3 on WebShop, while adding only 0.16%-0.20% advantage-construction overhead.
cs.AI / 86 / 2609.37125
When Should Agents Check External State? Budgeting Observations for Stored Intentions
Abstract
Prospective memory allows an agent to retain an intention tied to a future condition, but the stored intention does not reveal whether that condition currently holds. Checking it may require web access, multi-step tool use, and paid calls. Existing systems decide when intentions require attention, but do not allocate the resulting observations under a shared budget. We introduce the first resource-allocation formulation for the external observations required by stored intentions under a shared episode budget. BudgetPM offers two policy variants that share a hard-budget executor. BudgetPM-Static uses a lightweight Logistic scorer to learn whether a check improves the current decision. BudgetPM-Sequential distills full-episode hindsight schedules into a lightweight policy that decides when to spend or reserve capacity using only pre-query information at deployment. We evaluate BudgetPM against two public memory-agent systems, five matched controls, and four hand-designed monitoring or budget-adaptation rules. Across two benchmarks and three backbones, BudgetPM-Static outperforms adapted Mem0 and PMA workflows. On PM-Bench, its Logistic scorer reaches competitive quality--cost operating points alongside higher-capacity scorers and retains 99.9--100\% of unconstrained quality with 42--54\% fewer observations. Under severe scarcity and the same hard caps, BudgetPM-Sequential exceeds the strongest tested natural monitoring schedule by 1.92--2.58 Set F1 points. It reaches the same Set F1 and on-time recall with 16--33\% fewer observations. Matched attribution, exact-cost analysis, and a fixed-budget load intervention link this gain to competition between present and future opportunities. These results yield a demand--capacity design rule: local gating works when capacity covers demand, while future-aware supervision adds value when observations compete across time.
cs.AI / 87 / 2609.37145
From Judgment Quality to Downstream Utility: Rethinking LLM-as-a-Judge for Open-Ended Tasks
Abstract
LLM-as-a-Judge is increasingly used to evaluate policy responses on open-ended tasks that lack ground-truth answers. Existing work often directly converts the resulting judgments into reward signals for policy training, paying limited attention to intrinsic judgment quality and largely restricting the use of Judges to training-time supervision. We systematically investigate judgment quality and downstream utility by examining both how judgments are elicited and how they are used. For judgment elicitation, we vary the Judge protocol along three dimensions: verdict granularity, critique usage, and evaluation batching. For judgment usage, beyond policy training, we extend Judge to test-time inference through Best-of-N selection, Judge-guided revision, and beam search. We find that, (i) Surprisingly, judgment quality and downstream utility do not always align. (ii) Judge protocol design substantially affects both intrinsic judgment quality and downstream utility. (iii) Judge guidance effectively converts test-time compute into performance gains, with benefits varying across inference strategies. Our results call for a multifaceted evaluation of LLM Judges on open-ended tasks, encompassing intrinsic judgment quality, and downstream utility.
cs.AI / 88 / 2609.37153
When Tools Silently Lie: Evaluating and Mitigating Blind Compliance in Tool-Augmented Data Agents
Abstract
Tool-augmented data agents rely on tool outputs for analytical decisions. Yet successful execution can return plausible but incorrect evidence, requiring agents to decide whether to trust or verify it. Understanding this failure requires examining both the evidence obtained through checking and the answer ultimately adopted. We introduce ToxicBench to measure checking and adoption under numerical, label, schema, and retrieval errors, pairing clean and poisoned observations over fixed source data. In the 118-task GPT evaluation across three adapters, poisoning lowers task success by 26 to 39 percentage points. Ordinary retries help under one-shot poisoning, whereas repeated poisoning reveals wrong-answer adoption after checking. Controls on three public tables isolate how supplied evidence affects recovery. After freezing the scorer, we compare its judgments with human annotations on 200 trajectories, finding 96% task-success agreement. Human judgments support retry gains over Base and confirm adoption after checking on audited tasks. We release trajectories, versioned scoring, and reference and delivery audits. These findings highlight evidence availability and answer selection as complementary dimensions of agent reliability.
cs.AI / 89 / 2609.37176
Absorbed in Inertia: Activation Analysis for Computer-Use Agents
Abstract
Computer-use agents have become increasingly capable of executing tasks on live desktops through natural-language instructions, based on trajectories of screenshots, actions, and reasoning. We discover that they can stealthily exhibit inertia, in which they repeat fruitless actions despite recognizing that these actions are ineffective. We hypothesize that inertia is reflected in the agent's internal state, i.e., the activation values of the agent's underlying model, and propose a protocol to measure the relationship between the two. Extensive analysis of high-dimensional activation states shows that inertia corresponds to an absorbing region of activation space, where activation values become stale across actions and even after attempts to steer them. We conjecture that drastically changing the agents' activations by re-initializing them is necessary to escape inertia. Specifically, we propose R$^3$ (Reset, Reroute, Restore), which temporarily resets the agent's context trajectory to escape the absorbing region and then restores the historical context to effectively complete the task. Our approach yields 17-55% lower measured inertia across models relative to unmodified agents. These results suggest that changing the context can interrupt recurrence more effectively than directly steering the resulting activations. Our code is available at https://anonymous.4open.science/r/vlm-agent-defense-D076
cs.AI / 90 / 2609.37184
Accelerated surrogate dynamics for dynamical, stochastic system evolution
Abstract
Dynamic simulations are an entrenched way of gaining insight into the evolution of system dynamics. Their computational cost however is often prohibitively high, especially in cases of stochastic frameworks. Machine learning algorithms are especially suited as simulation surrogates. Nevertheless, they face some very distinct limitations. Firstly, the sheer dimensionality of these systems, however, precludes the use of traditional time series models who struggle with high dimensional feature spaces. Additionally, traditional time series focus exclusively on either long or short range effects, causing local or global drift given enough time. In this paper, we propose a framework that addresses those limitations. Our framework combines a Variational Autoencoder, with a convolutional or graph basis that reduces the dimensionality of the system. This latent vector is propagated in time using a Temporal Fusion Transformer model, which includes both long range and short range effect encoding, as well as static covariate support. We test our framework on three distinct cases, to prove its robustness and in all three we have achieved practically identical to the simulation results at a fraction of the time. Further, our framework is flexible enough to be adapted to any new system and provides an inbuilt uncertainty quantification for targeted experiment design.
cs.AI / 91 / 2609.37220
Information Bottleneck-Guided Adaptive Hypergraph Transformer for Brain Disease Diagnosis
Abstract
Exploring high-order correlations and long-range dependencies in brain networks holds significant value for both neuroscience research and clinical diagnosis. However, previous studies have lacked a unified integration of high-order and long-range dependency information in brain networks, and there is substantial redundancy behind various types of information. These issues limit their effectiveness in the diagnosis of brain diseases. To address this, we propose an Information Bottleneck-Guided Adaptive HyperGraph Transformer (IBAHGT). By incorporating the information bottleneck (IB) principle, this approach enables adaptive learning of high-order correlations and both short- and long-range dependencies within a unified framework for brain network analysis, achieving high-precision brain disease diagnosis. IBAHGT consists of three key components: an information bottleneck-guided adaptive hypergraph convolution, which introduces a novel hypergraph information bottleneck (HIB) principle to adaptively learn hypergraph message-passing weights between nodes and hyperedges, optimizes information flow and captures high-order information in brain networks that is maximally informative and minimally redundant (MIMR). The Transformer encoder captures global information within brain networks through the attention mechanism, specifically modeling short- and long-range dependencies. An information bottleneck-guided node-level adaptive fusion employs the IB principle to learn independent weights for each node, facilitating the fine-grained integration of high-order information and global information to obtain an efficient representation for downstream tasks. Extensive experiments demonstrate that the proposed method outperforms current state-of-the-art methods and can identify biomarkers for clinical applications.
cs.AI / 92 / 2609.37236
Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents
Abstract
An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. We study a distinct axis of proactivity: its content. Horizontal proactivity pursues unstated information that the current context already identifies, and vertical proactivity pursues needs that only earlier evidence reveals. A need graph, recovered from a benchmark's own decomposition, records which needs depend on which, so both forms, and whether the agent stops at the right time, can be scored from a transcript without a model judge. To learn this behavior, we propose Q&D (questioner and drafter), which trains a questioner to prefer the question whose continuation retrieves more of the required evidence, with no reward model or judge. On held-out splits of three multi-hop question-answering benchmarks, at equal retrieval spend, the trained questioner improves both forms of proactivity over the same model, prompted, and outperforms a prompted model $15\times$ larger in the same role on two of the three, and the gain persists after controlling for question volume and length. Without further training, we place the questioner in an interactive customer-service agent with a simulated customer, where it completes more tasks while asking fewer questions, and in retail it outperforms the $15\times$ larger model with fewer follow-up turns from the customer. These results show that proactivity depends not only on whether an agent acts without being asked, but also on what it chooses to pursue and when it stops.
cs.AI / 93 / 2609.37267
Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym
Abstract
Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and evaluating proactive LLM agents around three joint principles (3T): Task Capability, anticipating relevant needs and correctly performing useful work; Temporal Allocation, allocating compute according to resource availability and when results are needed; and Trust, sustaining users' confidence and appropriate reliance on the agent. We connect these objectives to a design space organized around five dimensions: task scope, anticipation horizon, activation trigger, processing timing, and intervention depth, and specify the situation and system modeling needed to support its choices, including user and environment representations, backbone LLMs, and agent harnesses. Lastly, we propose PROACTIVITY-GYM, a simulation-based evaluation testbed including multi-day scenarios, stateful environments, and persona-conditioned simulated users that can evaluate the consequences of proactive assistance across interactions. Evaluations across 23 model-harness configurations uncover substantial performance gaps across 3T and reveal that LLM judges often conflate task capability and trust. A human study with 30 participants demonstrates the importance of the joint 3T optimization: participants show sharp trust declines after intervention misalignment despite correct outcomes, and prefer sleep-time assistance, even when imperfect, to preserve ongoing focus. Together, these findings support designing and evaluating proactive agents through the joint consideration of useful work, compute allocation, and evolving user trust.
cs.AI / 94 / 2609.37272
Task-Relevant Null-Space Residuals for Non-Injective Neural Mappings
Abstract
Non-injective mappings in neural networks map distinct inputs to the same representation, thereby implicitly inducing equivalence relations in the input space. However, the input differences eliminated by these mappings may still be required by downstream tasks, creating a mismatch between operator-induced indistinguishability and task-required distinctions. For non-injective linear operators realized in the current forward pass, their null spaces exactly characterize these invisible input variations. We propose Task-Relevant Null-Space Residuals (NSR), a general residual framework for non-injective linear mappings. NSR combines null-space component extraction from pre-mapping representations, member-level encoding and gating, and application-specific integration to exploit potentially task-relevant information under downstream supervision while preserving the original aggregation or merging rules. We evaluate NSR in two structurally different settings: token merging and graph aggregation. In token merging, NSR achieves higher semantic segmentation performance than the corresponding compressed baselines in 34 out of 36 evaluated configurations, with a maximum observed gain of 31.51 mIoU points under strong compression. In graph aggregation, NSR achieves 100% training accuracy on Tree-NeighborsMatch at depths d=2--6 across three backbones, alongside gains on heterophilic node classification and molecular graph regression. Together, these results support null-space residuals as a practical complement to non-injective linear mappings, enabling downstream models to learn from input distinctions invisible in the original operator's output.
cs.AI / 95 / 2609.37279
Transolver-$σ$: Joint Spectral-Physical Subspace Modeling for Neural PDE Solving
Abstract
Neural solvers offer efficient surrogates for numerical simulation of partial differential equations (PDEs). For time-dependent problems, strong one-step accuracy does not necessarily translate into reliable autoregressive rollout. We observe that a solver based only on physical-state modeling can achieve lower one-step error, whereas its spectral-only counterpart can become more accurate at later rollout steps. Motivated by this observation, we present Transolver-$σ$, a neural PDE solver based on joint spectral--physical subspace modeling. Within each block, adaptive physical-state interactions and spectral transformations are modeled in dedicated latent subspaces, whose responses are recomposed to enable information exchange between the two representations. Within the physical subspace, we introduce Slice-Residual Physics-Attention (SRPA), which preserves an explicit slice-space identity path while retaining learnable cross-slice interaction. In parallel, an axis-factorized Fourier operator captures global spectral structure. Across five well-established PDE benchmarks spanning steady-state prediction and time-dependent dynamics, Transolver-$σ$ achieves state-of-the-art with a benchmark-averaged relative error reduction of 33.4% over the strongest baseline for each metric, while consistently improving autoregressive rollout over single-operator counterparts. Transolver-$σ$ further delivers strong gains on coupled multiphysics systems and real-world fluid and combustion measurements from RealPDEBench, demonstrating its effectiveness beyond standard simulation benchmarks.
cs.AI / 96 / 2609.37285
AssayRouter: Historical Utility Priors for Frozen Molecular Predictor Routing
Abstract
Laboratories often face a new molecular assay with 16-64 labels and a bank of predictors whose training data and parameters are unavailable. The practical question is which frozen outputs to include in a small local model. AssayRouter treats completed assays as pseudo-targets and labels each candidate by its post-fit utility: the reduction in held-out discovery loss when the candidate is added to the local target predictor. A shared regressor learns to predict this utility from candidate behavior on the support set, without source identity; on a new assay, one frozen ranking selects four sources and separate labels fit a convex combiner. We train only on completed ChEMBL-MT assays and evaluate 24 external regression assays across six frozen interface families. AssayRouter-C lowers strict four-call negative log-likelihood (NLL) by 0.0409 relative to Support-CV@4. Frozen candidate-label permutations confirm that candidate-utility correspondence carries the transferred information, and leave-one-interface-out training shows that the mapping generalizes to unseen predictor families. Completed assays therefore provide transferable supervision for scarce-label routing through frozen prediction interfaces.
cs.AI / 97 / 2609.37311
ReMem: Rethinking Perception and Memory in Long-Context Recommendation Agents
Abstract
Recent Recommendation Agents (RecAgents) offer a promising alternative by shifting recommendation to an active, user-side paradigm, where generative agents autonomously perceive external platforms, reason over user preferences, and execute decisions. However, existing RecAgents still suffer from two critical limitations: brittle item perception based on noisy and heterogeneous item pages, and inefficient long-context reasoning over extended user histories and multi-step interaction traces. To address these challenges, we propose a novel recommendation agent framework, termed as ReMem, that combines OCR-based multimodal perception with time-evolving dynamic memory. Instead of parsing raw HTML, ReMem observes item pages through screenshots and extracts structured multimodal information via an OCR tool, enabling a more humanoid and platform-agnostic perception mechanism. To support long-horizon preference modeling, ReMem further introduces a chunk-wise sequential memory update strategy, where the agent selectively maintains a fixed-size memory of informative historical interactions while processing arbitrarily long contexts with linear inference complexity and bounded context length. This design allows the agent to preserve evolving user preferences without relying on external memory modules or disrupting the standard autoregressive generation process. To enhance the dynamic memory instruction, we further develop a multi-memory GRPO variant, which propagates the final-answer advantage to all intermediate conversations that contribute to the final response. Extensive experiments on three datasets demonstrate that ReMem consistently outperforms state-of-the-art baselines, achieving an average improvement of 5.16\% across three recommendation agent tasks, namely searching, ranking, and judging.
cs.AI / 98 / 2609.37322
Mubric: Mutation Testing-Guided Rubric Generation for LLM Evaluation
Abstract
Rubric-based evaluation is widely used to assess LLM-based systems by decomposing response quality into task-specific scoring criteria. However, automatically generating rubrics that reliably capture task-specific quality requirements remains challenging. We introduce Mubric, a mutation testing-guided approach to rubric generation. Mutation testing, a classic software testing methodology, evaluates a test suite by injecting faults into programs and checking whether the tests detect them. We draw an analogy between test suites and rubrics: if a rubric captures an important quality requirement, introducing a corresponding defect into an otherwise high-quality response should reduce its score. Mubric first mines common defects from real pairs of preferred and dispreferred responses and abstracts these defects into reusable mutation operators, each specifying how to introduce a particular type of response defect. For a new task, it applies relevant operators to a reference response, checks whether the injected defects reduce response quality, and uses insufficiently penalized defects to refine the rubric. We evaluate Mubric on 703 tasks across four representative domains against six advanced rubric generation methods. Mubric achieves the highest overall evaluation accuracy, outperforming the strongest baseline by 7.48 percentage points.
cs.AI / 99 / 2609.37324
VISTA: Value-Informed Event Appraisal for Multimodal Emotion Conflict
Abstract
Conflicting emotional cues can be individually valid: a subdued voice may reflect a blocked goal while a smile satisfies a social obligation. Their interpretation depends on what the event means to the person. We introduce VISTA (Value-Informed Semantic Trust Arbitration), a learned seven-field appraisal interface that conditions modality arbitration on concerns, event relations, and expression conditions while retaining a joint-evidence residual. A log-odds decomposition separates emotion expectation from cue diagnosticity, motivating an interface that lets appraisal change how evidence is interpreted. With a shared Qwen2.5-Omni-7B backbone and matched training examples and steps, VISTA reaches 64.5% conflict accuracy on CA-MER, improving on modality gating by 2.5 percentage points on conflict and 0.2 on consistency. Shuffling appraisal across scenes or removing its decision connection reduces this benefit. A common frozen-backbone probe reaches 0.600 macro CCC for appraisal readout, compared with 0.505 for emotion-only fine-tuning. Evaluations across five benchmarks connect recognition under increasing conflict with appraisal readout and downstream decision use. Together, the analyses and experiments support scene-specific appraisal as an intermediate representation that helps interpret conflicting emotional evidence.
cs.AI / 100 / 2609.37326
Solving Without Stopping: On-Policy Distillation at Small Scale
Abstract
On-policy distillation, where a student learns from a stronger teacher's feedback on its own outputs, is a common way to pass reasoning to smaller models. We analyze what it transfers at small scale, distilling Qwen3-8B into Qwen3 4B, 1.7B and 0.6B students, in thinking mode (reason at length, then end the reasoning and answer) and, for comparison, in non-thinking mode (no separate reasoning phase). Long reasoning needs two abilities, solving a problem and knowing when it is solved, and we find that distillation transfers the first, but in thinking mode not the second. Solving improves at every size, up to two ceilings, which we measure comprehensively across both modes and all student sizes: a student's single attempt never exceeds what it could already reach in many attempts before training, and the smaller the student, the further it stays below the teacher. Stopping is where the modes part. In non-thinking mode every student keeps stopping; in thinking mode students stop ending their reasoning early in training, and the smaller the student, the less of this ability survives: the teacher signals a stop almost only where a student already ends its reasoning, so distillation teaches no new stops; it only keeps the student's existing stops that land on a right answer, and a weak student has few such stops. The smallest students often reach the right value but do not commit to it: they either rarely mark it or mark it and write past it. Together, these results describe how small students behave under on-policy distillation, and a diagnostic that separates answer marking, correctness and stopping.
cs.AI / 101 / 2609.37353
Seek Before You Move: Evidence Seeking for Progress Grounding in Vision-Language Navigation
Abstract
Vision-Language Navigation (VLN) requires agents to continuously ground task progress from long-horizon instructions and partial egocentric observations. Existing VLM-based navigation agents typically reason only over available observations and may remain confident even when task-relevant evidence is missing. For example, an agent may confidently proceed forward and get lost even though the landmark indicating the next turn lies outside its current field of view. We term this failure mode Progress Myopia: the agent fails to recognize unreliable progress grounding and continues acting on insufficient evidence. To address it, we propose SeekVLN, an evidence-seeking framework that couples semantic progress reasoning with active acquisition of task-relevant observations. SeekVLN is trained in two stages: First, Future-guided Reverse Generation (FRG) uses future expert actions to augment offline expert trajectories with supplementary views and evidence annotations. Supervised fine-tuning on these trajectories establishes a prior for evidence seeking and progress reasoning without additional expert interaction. However, imitation alone does not reveal whether seeking improves subsequent navigation. We therefore introduce Counterfactual Contrastive Policy Optimization (C2PO) for reinforcement fine-tuning. By comparing each evidence-seeking branch with a counterfactual direct-navigation branch from the same state, C2PO uses a contrastive reward to assign credit to seeking decisions based on subsequent navigation benefit. Experiments on simulated benchmarks show that SeekVLN achieves state-of-the-art performance, improving success rate by 12.7% and 7.5% over the base model on R2R-CE and RxR-CE, respectively. Both simulated and real-world evaluations exhibit human-like evidence-seeking behaviors for more reliable progress grounding.
cs.AI / 102 / 2609.37356
Teaching LLMs to Generate Challenging MILP Instances via Solver Feedback
Abstract
Generating optimization instances that are both feasible and computationally challenging is crucial for benchmarking solvers and training learning-based optimization algorithms. Existing non-LLM generators rely on seed instances or parameter tuning, resulting in high test-time computational cost, while existing LLM generators lack explicit hardness measures. Recent reinforcement learning methods with verifier feedback evaluate only binary correctness, which is misaligned with generating challenging problems. We note that an optimization solver reports the cost of solving at several stages of its pipeline, and leverage this to design a reward that scores both the solvability and the hardness of generated problems, measured by branch-and-bound nodes and post-cut relaxation gaps. Our key idea is a challenger-solver asymmetric self-play approach, where an LLM challenger generates progressively harder instances and the solver verifies feasibility and hardness, so no seed or training MILP instances are required. We fine-tune Gemma-4-12B and Qwen3.5-4B with GRPO and a size curriculum into OptiScribe-12B and OptiScribe-4B, which generate feasible yet challenging MILP problems from natural language instructions. On capacitated facility location and max-cut, OptiScribe-12B raises median SCIP search nodes by 1.7-5x and post-cut gaps by 1.1-1.7x over its base model and improves the feasibility rate on facility location by 9-19 points, while OptiScribe-4B raises median nodes by up to 15.6x. The problems cover a wider difficulty range than public benchmarks of the same size, follow instructions on density and difficulty, and can tune solver settings for families that public libraries lack. These results indicate that optimization-specific rewards, used in self-play mode, can teach LLMs to generate high-difficulty optimization benchmarks. We will release our code and models publicly on acceptance.
cs.AI / 103 / 2609.37377
Beyond Prompt Count: How Data Shapes Transfer in On-Policy Distillation
Abstract
On-policy distillation (OPD) trains students using teacher feedback on their own sampled responses, yet how prompt choice shapes transfer across teacher-student pairs remains poorly understood. We systematically study prompt quantity, source, and selection across RL- and SFT-continuation pairs and cross-model settings. We find that OPD can be highly prompt-efficient: a few prompts can approach large-pool performance, with four DAPO prompts matching the observed mathematics score of 3,840 DeepMath prompts. However, prompt utility is relational rather than intrinsic: changing only the teacher can reverse the relative effectiveness of mathematics and code prompts. To characterize these transfer differences, we analyze parameter and functional changes across prompt supports and model pairs. Functional alignment with the teacher varies across supports and target tasks; in continuation pairs, teacher-aligned prediction changes can coexist with weak parameter alignment. Continued OPD on effective supports can restore performance after unfavorable transfer. Finally, targeted selection does not consistently outperform uniform random sampling, and filtering out a source that performs poorly alone yields no consistent gain across three paired support draws. Overall, our results distinguish prompt efficiency from prompt interchangeability and show that effective data choice depends on the teacher-student pair and target capability, with random sampling providing a competitive baseline in the studied settings.
cs.AI / 104 / 2609.37398
Direct Experience World-Model Optimization: Learning the World Beyond Action Imitation
Abstract
World-Action Models (WAMs) couple action generation with predictions of how physical interactions unfold. However, current post-deployment learning paradigms typically improve behavior without requiring better world predictions. Especially in dexterous manipulation, small execution errors can compound in high-dimensional action spaces, hindering policy improvement and pushing interactions beyond the world model's training distribution. Motivated by this, we propose Direct Experience World-Model Optimization (DEWO), a post-deployment learning paradigm for WAMs that, alongside action imitation, refines world representations through visual experience to better condition action generation. Specifically, it identifies interaction turning points and learns from successful and failed futures to support classifier-free guidance. An additional value head estimates task progress from video representations and activates guidance when progress stalls during inference. Across five DexJoCo tasks, DEWO improves average success across all three WAM formulations. Ablations show that visual supervision from successful and failed continuations improves both prediction and control beyond action supervision alone. On four real-world tasks across Wuji and Sharpa, 3 x 3 grid evaluations show that two rounds of deployment learning increase success from 51.0% to 71.7% in cells with at least one initial success, a gain of 20.7 percentage points. These findings support continued predictive learning for improving control through deployment experience, making world modeling an active part of WAM adaptation.
cs.AI / 105 / 2609.37446
Demistifying Data and Simulator Assumptions in Supervised Causal Discovery
Abstract
Supervised causal discovery learns to infer causal structure for a new dataset from training datasets paired with structural labels. These training pairs are typically simulated, making the simulator both a source of supervision and a carrier of assumptions about causal graphs, mechanisms, and noise. Understanding the resulting predictions therefore requires examining how these assumptions supplement the information available in observational data, which may be compatible with multiple causal graphs. This paper examines that relationship across representative methods available through June 2026. We organize these methods by prediction target, prediction granularity, encoder, structural decoder, and training regime to relate what each method predicts to how it uses data and simulator-based supervision. Using this framework, we distinguish two questions: whether the target is identifiable under the assumed model class, and whether a trained predictor generalizes beyond its training distribution. Restrictions on mechanisms and noise can make otherwise ambiguous causal directions identifiable, but predictive accuracy under those restrictions does not establish transfer when they change. This distinction motivates evaluation that matches metrics to the identifiable graph target and tests changes in graphs, mechanisms, and noise between training and deployment. Extending such evaluation to real data also requires documenting the external causal evidence and uncertainty behind benchmark reference graphs. Together, these analyses guide method comparison and identify open questions in transfer, test-time adaptation, and uncertainty assessment.
cs.AI / 106 / 2609.37539
SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation
Abstract
Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data and how to train agents for skill use remain underexplored. In this work, we propose SkillGym, an automatic pipeline to build verifiable environments, collect trajectories, and train skill-use agents. SkillGym first crawls a large volume of skills from the internet, then keeps those whose workflows can run reproducibly offline. A builder-reviewer pipeline is used to construct difficulty-controlled tasks, spanning four task types, each with a reference solution and an executable verifier. With this pipeline, we build 6.8k environments and collect 19k verified successful trajectories for supervised finetuning. Finetuning on these trajectories improves LLMs of different families and sizes, from 2B to 122B parameters across four skill-use benchmarks; Our Qwen3.5-9B SFT model outperforms the 397B untrained model on two of them. Further analysis shows that training teaches agents to invoke skills, raising the rate of reading the relevant skill from 28% to 96%, and that the gains hold across reasoning structures, extending to task types that form a minority of the training data and to skills held out from training
cs.AI / 107 / 2609.37544
How Can Recommendation Feedback Evolve Agent Memory?
Abstract
Content-generation agents continuously receive impressions, clicks, conversions, and negative feedback from recommendation systems, providing real-world outcome signals for memory evolution. However, these signals are delayed and noisy, confounded by audience composition, placement, and recommendation policies, and may result from the combined influence of multiple memories, making accurate attribution difficult. Existing methods rely primarily on immediate feedback or semantic retrieval and therefore struggle to reliably translate recommendation outcomes into memory fitness. To address this challenge, we propose TIDE (Trajectory-Informed Directed Memory Evolution), an external memory evolution framework driven by delayed recommendation feedback. We further introduce Memory Evolution Gain (MEG), which measures the utility improvement of evolved memory over a no memory baseline on strictly future tasks. TIDE treats memory as a capacity-constrained population of experiences: temporal and semantic credit assignment estimates contextual fitness, while responsibility credit distributes outcome signals according to the memories referenced during generation. These signals are then used to reinforce, crossover, mutate, or evict memories. On an e-commerce membership marketing content-generation agent, TIDE achieves a +7.75-percentage-point MEG in offline temporal replay and significantly improves both unique click-through rate (UCTR) and activation rate in an online A/B test. On a delayed-label benchmark, TIDE achieves the lowest mean absolute error (MAE) and root mean squared error (RMSE) and the highest MEG among the compared methods, demonstrating its effectiveness.
cs.AI / 108 / 2609.37588
Rational Clarification by Assistive Agents via Value-of-Information Reasoning
Abstract
Users of language-based assistive agents often make ambiguous requests. In response, an assistant can either directly act on its interpretation of the request --- risking misalignment with the user --- or ask a clarifying question. Which option is the most safe and helpful? A common approach is to ask questions that minimize uncertainty about the user's intent until a threshold is reached. However, this neglects the impact of uncertainty reduction on downstream performance, the costs of asking versus acting immediately, and the possibility that users may provide corrections without being asked. To navigate these trade-offs, we introduce Rational Enquiry via Value-of-Information Reasoning (REVOIR). REVOIR makes clarification decisions via inference-time reasoning about the value-of-information of a question, which captures the expected improvement in task reward due to the answer received. In two assistive tasks --- ambiguous question answering (CondAmbigQA) and preference-aligned household task planning (ADAPT) --- we show that REVOIR achieves greater success with fewer questions than approaches based on prompting, chain-of-thought, fine-tuning, or information gain, improving preference satisfaction on ADAPT by 13-15% over a fine-tuned clarification policy while requiring no training and asking five times fewer questions. Furthermore, when the assistant can receive cheap user corrections after acting, REVOIR naturally infers that asking questions is not always efficient, demonstrating the adaptivity of our approach. In contrast, we find that vanilla reasoning agents fail to adaptively clarify user requests, and request fewer clarifications as reasoning effort increases.
cs.AI / 109 / 2609.37590
FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents
Abstract
LLM agents accumulate interaction histories that grow linearly with task length, causing quadratic inference cost scaling and performance degradation from attention dilution. Existing context-compression methods learn what to discard offline: by contrastively optimizing guidelines, distilling compressors, or training compression policies. This incurs a substantial cost. Further, the compression policy is learned a priori and is not dynamically conditioned on the evolving test-time trajectories. In this paper we ask a complementary question: Which past interactions causally shape the agent's future decisions? We recast context compression as a causal decision preservation problem over discrete interaction units and introduce FOCUS, a training-free context compression framework that operates entirely at test time. Our method requires no offline data collection or fine-tuning, and is architecture-agnostic, attaching to any closed-API frontier model as a modular compression layer. We evaluate FOCUS on diverse agentic benchmarks including API and tool-calling, QA, web domain and multi-turn dialogue. Our method establishes new state of the art performance, cutting peak context by up to 48% and dependency by 73% while improving task success by up to 8.9 percentage points over uncompressed execution.
cs.AI / 110 / 2609.37594
XU-RS: Explaining Credal Width in Random-Set Language Models
Abstract
Uncertainty estimates tell us how unsure a model is, but not why. Without knowing which parts of an input influences a model's uncertainty, we cannot tell whether that uncertainty score depends on input features that are relevant for the task. We study this problem in randomset classifiers built using pretrained language models. These classifiers assign probability to individual answers and to groups of answers, producing lower and upper probabilities for each answer; The difference between these probabilities, called credal width, is used to represent epistemic uncertainty about an answer arising from limited training data. We propose XU-RS, a framework that attributes an answer's credal width to the input tokens (words or word pieces) supplied to a language model. XU-RS uses Expected Gradients (a standard feature attribution method) to estimate how input tokens contribute to credal width. The proposed framework is evaluated on a MedQA dataset using SmolLM3-3B and Llama-2-7B models, demonstrating that setting the embedding of a token ranked highly by XU-RS to zero (zero-masking) causes larger changes in credal width than zero-masking randomly selected tokens. In addition, we show that normalisation can cause other answer groups to influence an answer's width, reveal how token attribution can mask numerical errors, and provide diagnostic checks to verify whether a token ranked highly by XU-RS meaningfully explains model uncertainty.
cs.AI / 111 / 2609.37642
Flattening the Connectome Spectrum: A Spectral Filter for FC Induces a Pretraining Target for fMRI Encoders
Abstract
Self-supervised pretraining reshaped prediction in language and vision, and brain foundation models (BFMs) inherited its promise. Representations learned from large unlabelled corpora should capture individual functional dynamics and generalise across cohorts. However, kernel ridge regression (KRR) fitted on functional connectivity (FC) matrices still predicts individual phenotypes more accurately than any BFM we tested. In this paper, we show that KRR is weighted by the eigenvalues of the FC which are miscalibrated for phenotype prediction. We apply an efficient spectral filter to recalibrate the eigenvalues of each subject's FC matrix, enabling the model to exploit more inter-individual variance. Across the 5 datasets, 11 parcellations and 6 prediction targets we tested, we match or exceed the KRR baseline. Based on this finding, we then pretrain a small encoder model on about 4,000 hours of fMRI from 162 open datasets, whereby we align the pairwise similarities between the embeddings of recording snippets with those between the recalibrated connectomes. Our model performs on par with the best of the 6 published BFMs we tested while having an order of magnitude fewer parameters. Our encoder performs better than FC on short scans and in smaller cohorts, especially in fingerprinting. We release the pretrained model weights, the code and the pretraining data, preprocessed and parcellated.
cs.AI / 112 / 2609.37644
Beyond a single latent space: a dual-latent world model for long-horizon planning
Abstract
Latent world models often struggle with long-horizon planning despite accurate short-term predictions. Recursive rollouts accumulate errors, while distance concentration in high-dimensional latent spaces can weaken goal discrimination. We introduce the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning through distinct state representations and dynamics models. The low-level model predicts action-conditioned transitions, while the high-level model uses learned macro-actions to plan over longer temporal spans. We also propose Long-Horizon Representation Learning with Weighted Rollout (LoRe), which supervises self-generated predictions at both levels. An analysis of recursive error propagation motivates exponential horizon weights with separate decay rates for the two temporal scales. During planning, the high-level model generates latent subgoals that the low-level model refines into actions for precise execution. We evaluate from-scratch Dual-WM on five goal-conditioned visual control tasks against the task-wise strongest baselines without actor-guided proposals. At goal offsets of 50 and 100 environment steps, mean success increases from 75.9% to 84.4% and from 61.4% to 69.5%, respectively. At offset 100, Dual-WM outperforms these baselines on all five tasks and improves mean success over LeWM by 30.8 percentage points. Ablations and supporting analyses provide evidence of more informative representations for goal evaluation and greater consistency under recursive prediction. These results highlight the value of separating temporal roles and training across multiple horizons for reliable latent planning. Our core implementation is available at https://github.com/DeLin1001/Dual-WM-Official.
cs.AI / 113 / 2609.37658
EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making
Abstract
LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-term trade-offs. However, existing enterprise and financial benchmarks mainly test static capabilities such as information extraction, numerical calculation, domain knowledge, and financial QA, leaving interactive and long-horizon decision-making underexplored. To bridge this gap, we introduce EnterpriseBench, a benchmark that evaluates LLM agents across this spectrum, from static question answering to dynamic decision-making. Specifically, EnterpriseBench reorganizes existing enterprise and financial QA datasets into a unified foundational suite annotated by capability and difficulty, and introduces three professional interactive settings: Consulting, based on management-consulting-style business cases for client problem diagnosis through multi-turn information seeking; the Beer Game, adapted from a classic supply-chain management simulation for inventory control under delayed feedback; and Enterprise Digital Twin, a project-based business simulator for workforce, risk, and project planning. Experiments with nine agent methods under four backbone models show that current agents have not yet achieved stable, comprehensive, and cross-task reliability in enterprise scenarios. These results show that EnterpriseBench provides a practical benchmark for evaluating LLM agents in realistic enterprise strategic reasoning and decision-making.
cs.AI / 114 / 2609.37684
Learning from Shared-Control Overrides: Context-Driven Acceleration Profile Prediction for Personalized Overtaking
Abstract
Adaptive Cruise Control (ACC) systems are typically calibrated for an average driver, often resulting in a mismatch between vehicle behavior and individual expectations during time-critical maneuvers such as highway overtaking. When the ACC is perceived as too conservative and inconsistent, drivers intervene through throttle overrides, providing implicit feedback on the system's behavior. This paper reframes these override actions as human-in-theloop supervisory signals and proposes a data-driven framework for personalized vehicle adaptation, termed Context-driven Personalized ACC (CoP-ACC). Rather than relying solely on end-to-end regression, which tends to over-smooth dynamic responses, we introduce a hybrid pipeline combining: (i) unsupervised hierarchical clustering to extract representative acceleration profiles from override events; (ii) a context classifier that maps pre-maneuver driving conditions to the appropriate profile; and (iii) a residual regressor that refines the selected profile into a smooth, personalized acceleration profile tailored to the immediate context. Evaluated on real-world public-road data against a withheld forced-ACC baseline, the approach demonstrates high reconstruction fidelity and generates acceleration profiles that tend toward the driver's expected behavior in potential override contexts. The results highlight the potential of learning from shared-control overrides to enable anticipatory, personalized ACC behavior, reducing manual interventions and improving ride comfort.
cs.AI / 115 / 2609.37686
EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?
Abstract
Autonomous agents have made rapid progress in general-purpose computer use, but reliable automation of professional industrial engineering remains out of reach, as engineering workflows demand reasoning over geometric and physical constraints and dependencies preserved across software and design stages. We present EngiWorld, the first benchmark structured around the complete design loop: 1,301 expert-curated tasks spanning 6 engineering domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, with both GUI and CLI interfaces and 6 task types ranging from software-selection to open-ended tasks. We further introduce an artifact-centric evaluation methodology built on a unified domain-verifier suite, which programmatically checks the geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts, and scores quantitative design tasks continuously by specification attainment rather than binary success. Evaluation of seven frontier models reveals a substantial capability gap: the strongest model achieves an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed. EngiWorld provides the first rigorous foundation for measuring progress toward agents that operate professional engineering software end to end.
cs.AI / 116 / 2609.37687
WISE-ATTA: When to Ask for Labels in Budgeted Active Test-Time Adaptation
Abstract
Active test-time adaptation (ATTA) improves robustness under distribution shift by updating a deployed model during inference while selectively querying supervision. However, most existing ATTA methods implicitly assume that supervision can be requested for every incoming test batch, which can incur substantial annotation cost over long test streams. In this work, we introduce \emph{budgeted ATTA} in which labels are available for only a fraction of test batches. This formulation shifts the central challenge from deciding \emph{what} to label within a batch to deciding \emph{when} supervision should be applied over time. To address this challenge, we propose a budget-aware approach \emph{WISE-ATTA} that allocates supervision over the test stream based on lightweight signals computed online, prioritizing periods where supervision is likely to be most useful. When a batch is selected for supervision, we further employ a drift-based sample selection criterion that targets samples exhibiting ongoing, unconverged adaptation dynamics, enabling effective updates from a single labeled example. We evaluate this approach on synthetic corruptions (ImageNet-C) and natural distribution shifts (ImageNet-R/K/A). Across settings, WISE-ATTA achieves competitive or improved performance compared to recent ATTA methods while requiring substantially fewer labels. Overall, we find that the timing of supervision is a key, yet underexplored, aspect of active test-time adaptation. Code: https://github.com/Muhammad-Huzaifaa/WISE-ATTA
cs.AI / 117 / 2609.37708
Generative Interactions: Weaving Multiparty Human Motion with Bilevel Latent Dynamics
Abstract
Human social behaviour is not a collection of independent motions, but a jointly organised process in which group dynamics and individual variation continuously shape one another. Yet existing social motion models often prioritise plausible trajectories while leaving interaction state implicit, limiting their ability to transfer across groups, tasks, and partial-observation regimes. To address this gap, we introduce Bilevel Representations for Agent Interaction Dynamics (BRAID), a hierarchical sequential latent-variable model for generative multi-person interaction. BRAID explicitly formulates social motion generation as a meta-transfer learning problem: shared interaction priors are learned across datasets and adapted through arbitrary context sets of observed people and joints. The model represents each scene through a group-level latent state that captures shared interaction dynamics and person-level latent states that capture individual behaviour conditioned on the evolving group context. This modelling choice enables coherent generation under full, sparse, or partial observations while exposing compact social-state vectors that can serve as an interface for downstream embodied-agent systems. We evaluate BRAID under a unified SMPL-based representation on social forecasting, tracking and in-filling, and response generation, using metrics that assess not only reconstruction accuracy but also realism, diversity, temporal alignment, and interpersonal coordination. We further analyse the hierarchical latent space, showing that it captures separable group- and individual-level structure.
cs.AI / 118 / 2609.37725
Context Language Models
Abstract
We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files. Building CLMs zero-shot with existing models outperforms SOTA context management strategies across a variety of tasks: 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task. Moreover, by shifting context management from external harness control to intrinsic model behavior, CLMs naturally enable both in-context and parametric learning of context-management strategies. We show that CLMs can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute. We also introduce an online reinforcement learning method for CLMs, improving Qwen3.5-9B performance on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs. Finally, we co-design Suffix Cache Reuse for CLM serving, further reducing server-side compute by 35% relative to standard SGLang at matched performance.
cs.AI / 119 / 2609.37730
Spatiotemporal Hyperedges for EEG Seizure Detection and Prediction
Abstract
Seizure detection and prediction from EEG are clinically important but challenging because seizures are rare, temporally localized, and propagate as coordinated events across multiple channels. Recent dynamic graph neural networks model this by running a temporal model over a sequence of per-time-step pairwise channel edges. However, this pairwise construction misses the spatiotemporal coupling that constitutes a seizure, at substantial training cost. We propose HyBrain, which summarizes spatiotemporal EEG evidence through a small set of soft hyperedges rather than pairwise edges. A per-channel Mamba backbone produces one token per (channel, second), and a spatiotemporal hyperedge block pools these tokens into E_h shared group embeddings through soft memberships and broadcasts them back. The same encoder serves three downstream tasks: window-based detection, one-second point-wise detection, and preictal seizure prediction. On TUSZ and CHB-MIT, HyBrain achieves the best AUROC on every reported setting against ten baselines, with the largest gap on long-clip preictal prediction. It also matches the most efficient baselines in training time and peak GPU memory. A qualitative analysis shows that even a single learned hyperedge cleanly captures the preictal -> ictal -> postictal trajectory on a real seizure clip.
cs.AI / 120 / 2609.37743
ContextRender: From Execution Dependencies to Agent Context
Abstract
LLM agents performing long-horizon tasks accumulate tool results that later steps may need. Passing the full history to every invocation is costly even when it fits within the context window, while reducing it risks omitting needed information. Existing context management methods can overlook how earlier tool results are used in subsequent execution, leaving needed information out of context. We introduce ContextRender, which manages context through a persistent graph of execution dependencies. We develop Tool-Flow Analysis to track how later operations reuse information from earlier tool results, providing a signal called observed reuse. A renderer combines this signal with recency and semantic relevance to select results within a fixed history budget, retaining omitted results for later use. Across AppWorld and 8-objective QA with three execution models, ContextRender outperforms the evaluated context management baselines using a 6K history budget, well below the models' maximum context windows. Within this budget, it achieves task performance close to or above that of passing the full history while reducing mean inference cost by 10.2%-32.2% relative to Full history. Ablations show that observed reuse improves task performance and retention of results reused later.
cs.AI / 121 / 2609.37773
OmniVCBench: Benchmarking Evidence-Grounded Multimodal Reasoning Towards AI Virtual Cells
Abstract
Artificial Intelligence Virtual Cells (AIVCs) are envisioned as scientific agents that simulate cellular responses, explain underlying mechanisms, and support hypothesis-driven discovery. Existing AIVC benchmarks, however, operate primarily at the simulation layer, motivating complementary evaluation of how models interpret experimental evidence and formulate biological hypotheses. We introduce OmniVCBench, a figure-centric, source-traceable benchmark for the interpretation component of an AIVC. It contains 6,077 curated single- and multi-subfigure question--answer pairs derived from figures and experimental contexts in the scientific literature. Guided by Bloom's taxonomy, we instantiate interpretation-layer counterparts of the AIVC Predict--Explain--Discover agenda through three scientific reasoning tasks. We further introduce AIVC-Judge, a task-conditioned MLLM-as-a-judge framework with category-specific, reference-aware rubrics for evaluating open-ended responses. A complementary Model-Derived Hard-Negative Mining (MDHNM) strategy converts plausible errors observed during model inference into MCQ distractors for lower-cost evaluation. Within the evaluated heterogeneous model pool, MCQ accuracy correlates positively with AIVC-Judge scores, providing a complementary view of performance alongside open-response evaluation. Code and data demo are available at https://anonymous.4open.science/r/OmniVCBench.
cs.AI / 122 / 2609.37787
Adam under Generalized Smoothness with Second-Moment-Type Stochastic Gradients
Abstract
Adam is widely observed to remain stable even when the objective deviates significantly from global smoothness. Under the generalized smoothness framework, however, existing analyses rely on strong tail assumptions on the stochastic gradients, such as almost-sure boundedness or sub-Gaussianity. Whether Adam converges on generalized smooth objectives under only second moment information on the stochastic gradients, without such concentration assumptions, was identified as an important open direction by Li et al. (2023). This paper gives an affirmative answer under fairly general conditions: such tail assumptions are not necessary. Building on the Adam self-normalization framework of Jin et al. (2026), developed for classical smoothness and bounded variance, we extend the stopping-time and de-preconditioning strategy to the $L_0$-$L_p$ generalized smoothness condition and a generalized second moment ABC condition. Even when the stochastic-gradient condition provides only second moment information that may grow along the trajectory, the stochastic trajectory of Adam remains in a locally well-behaved smoothness region, with stretched-exponential tail decay under bounded variance and global smoothness. Consequently, we establish high-probability convergence rate guarantees over the full range $p<2$, with confidence dependence of order $δ^{-1/2}$, while the stepsize prefactor depends on $δ$ only through a single logarithmic factor. We further construct a hard instance showing that, under only second-moment information, this $δ^{-1/2}$-type confidence dependence is sharp. Finally, in the regime $p<1$, we combine the trajectory control with polynomial-growth estimates on rare events to obtain convergence rate guarantees in expectation.
cs.AI / 123 / 2609.37791
A neural network that maintains and retrieves memories based on context
Abstract
Every day, people continuously infer situational context and adjust the way they understand and remember the world. Context, signaled by the prefrontal cortex, is known to modulate working memory and episodic memory, but the algorithmic understanding of this modulation remains limited. Here, we train a recurrent neural network (RNN), augmented with an episodic memory buffer, to infer context using Bayesian inference as it continuously makes predictions of upcoming scenes while watching naturalistic movies. When the inferred context modulates the RNN's recurrent connectivity (the basis of working memory) in a low-rank manner, the model's activity patterns best match neural responses in human participants who watched the same movies during fMRI. Context also modulates episodic memory retrieval, such that the model retrieves memories based on not only content similarity but also context similarity. This is implemented as a key-value system with self-attention, designed to additionally encode context and retrieve context-congruent memories. The resulting model not only better resembles human brain representations but also learns to retrieve memories like humans much faster than a model without context modulation. Together, our findings suggest a computational mechanism by which context modulates information maintenance and long-term memory retrieval in naturalistic environments.
cs.AI / 124 / 2609.37832
Can a Cacheable Decision Model Follow Rules?
Abstract
Certo is a small non-generative decision model (Qwen3-4B): it scores candidate actions from their text and returns a probability, instead of generating an answer. The accurate design reads the state, the rules, and each candidate together (a joint scorer), so cost grows with the menu. Independent encoding lets each candidate be encoded once and reused across states (about 5x cheaper at 77 candidates), but separates state from candidate. We ask how much rule-sensitivity survives that move, and whether it can be trained back. Four experiments on Certo: (1) the tested conversion to cacheable scoring loses rule-sensitivity (recall@1 1.00 -> 0.24) while the joint scorer holds 1.00, and a shortlist+rerank rescue fails; (2) targeted counterfactual supervision restores strong performance on held-out synthetic rule tasks (paraphrase, counterfactual, composition; reproducible across seeds), though we do not isolate whether predictions depend on the supplied rule; (3) on real rules the added benefit is not established -- after fixing a truncation confound, the joint scorer wins significantly on the short tier (0.861 vs 0.500) and directionally on the hard tier (0.655 vs 0.483, n=29); (4) a matched cross-domain real-prose mixture did not help and reduced contract accuracy (-9.3, -16.2 points). A cacheable encoder can be made rule-sensitive on its training distribution, but transfer to unseen-source real rules is not established; the joint scorer keeps an edge at the cost of caching.
cs.AI / 125 / 2609.37834
Mixture of Self-Improving Branches For Agent Harness Optimization
Abstract
Harness optimization provides a practical setting for recursive self-improvement (RSI), where agent-generated modifications inform subsequent changes through execution feedback. Recent work such as Meta-Harness implements this process through iterative code generation and evaluation, but retains a fixed development set and proposal policy. These constraints channel evolution along a single search trajectory, increasing the risk of converging to a local optimum. We make the improvement process itself adaptive by organizing search into branches with evolving development subsets and proposal policies. Each branch retains development cases solved by more of its leading harnesses than by those of other branches, drops cases solved by every leading harness across all branches, and revises its proposal policy using its own search history. To deploy the resulting complementary harnesses, we propose a router to select one development-selected branch head for each new input before execution. Across mathematical reasoning and agentic coding benchmarks, our system achieves relative improvements over Meta-Harness of 34.8% on Olympiad-level mathematical reasoning, 11.6% on Terminal-Bench 2.0, and 3.8% on SWE-bench Lite, with harness selection and router configuration based solely on development data. These results show that evolving branch objectives and proposal policies can yield complementary harnesses whose strengths a router combines without access to test outcomes.
cs.AI / 126 / 2609.37898
Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL
Abstract
Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-stage policies often fail to solve sampled tasks, leaving little useful reward signal for learning. To mitigate this problem, we use on-policy distillation (OPD) to provide token-level guidance on the student's own rollouts. We find that the benefit of this guidance depends on the performance gap between the teacher and the student. When the teacher substantially outperforms the student, distillation helps guide the student through the early training stage where outcome rewards provide little learning signal. As the gap narrows and eventually reverses, however, continued distillation becomes less beneficial and may hinder further improvement. Motivated by this observation, we propose Gap-Adaptive Teacher Scheduling (GATS), which augments the student's RL objective with an OPD term whose weight adapts to the teacher-student performance gap. Specifically, GATS gradually reduces teacher guidance as the student approaches the teacher's reference performance and withdraws it once that reference is reached. This enables GATS to leverage task-trained teachers smaller than the student, since teacher guidance is primarily needed during early training. Across ALFWorld, WebShop, and ScienceWorld with three Qwen2.5 teacher-student configurations, GATS achieves the highest average success rate among the compared methods in all three configurations, improving over reward-only GRPO by 4.37%-11.87% under matched student rollout budgets. Code is available at https://github.com/Ricardo-H/guide-then-let-go.
cs.AI / 127 / 2609.37902
You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference
Abstract
Existing LLM routers choose among models using static per-model costs. We show that open-weight inference markets introduce a second, largely ignored decision axis: after choosing a model, a client must still choose which provider serves it. Measuring live endpoints across [nummodels] open models, competing providers, multiple task types, and three measurement waves, we find that provider choice cannot be inferred from the price list. The same model can vary sharply in quality, latency, availability, and price across providers; higher-priced providers are consistently faster, but price does not reliably predict quality or availability; and provider feasibility is task-selective, with one deployment nearly normal on knowledge tasks but catastrophically degraded on multi-step reasoning. We formulate same-model provider selection as a price-taker market-aware routing problem. A simple measured-map policy routes to the cheapest provider that is both quality-equivalent and healthy, yielding matched-quality savings while avoiding degraded endpoints. Because the map drifts, we introduce FACET, an online provider router that certifies per-(provider x task) feasibility facets and fails safe to an anchor before serving uncertified arms. Across relaxed deployment assumptions, FACET tolerates imperfect task assignment and sparse feedback, while systematic evaluator bias exposes a quality-signal trust boundary that can be mitigated with ground-truth probes or audits. Live provider runs further confirm that certification can move real traffic from a premium anchor to a substantially cheaper certified endpoint. Our results suggest that market-aware LLM routing must measure not only which model to use, but also who serves it.
cs.AI / 128 / 2609.37907
Pixels to Keys: Exploring Spatial and Motion Cues in Gameplay Inverse Dynamics
Abstract
Video games offer scalable environments for studying perception and control in embodied agents.Abundant online gameplay videos could supply demonstrations, but they rarely include player inputs for training. Inverse Dynamics Models (IDMs) have thus been proposed to infer inputs from frames. Large (up to 1B parameters) IDMs trained on $\sim$1K-2K gameplay hours demonstrate feasibility and cross-environment generalization at this scale, but researchers do not clarify what the key components are to recover individual actions and often report only aggregate accuracy that can mask rare-action failures. We study the problem in a data-constrained scenario to evaluate how spatial motion features, model architectures, and training objectives affect an IDM's outcome and we analyse our models on per-key and balanced metrics such as $F_1^{macro}$. Our experiments on Trackmania highlight the importance of factors like the model architecture and motion flow extraction in preprocessing, while also showing the limits of evaluation through unbalanced metrics. The application of the same architecture and training recipe to Cyberpunk 2077 reveals uneven performance across game mechanics. Our per-action evaluation and failure analysis highlight ambiguities from camera motion, delayed effects and imbalanced key-press frequencies that call for explicit modeling of 3D scene structure, long-term state and the adoption of proper losses in future implementations.
cs.AI / 129 / 2609.37934
GRFBrain: Graph-Structured Rectified Flows for EEG Dynamic Modeling
Abstract
Forecasting time-varying functional connectivity from electroencephalography (EEG) requires modeling both history-dependent trends and structured variability across channels. Conditional flow matching provides a framework for distributional forecasting, yet it remains unclear whether graph-informed source distributions offer practical advantages over isotropic noise and strong deterministic predictors. We introduce a graph-structured residual flow framework that separates conditional mean prediction from stochastic residual transport. A history-only predictor estimates the future connectivity graph, while a graph Gaussian source encodes dependencies derived from past connectivity through a Laplacian-based covariance. A conditional velocity field transports source samples to future graph residuals, with transport time explicitly distinguished from physical EEG time. Our study identifies the conditions and controls needed to distinguish useful residual transport from improvements attributable to deterministic prediction, learned representations, and sampling effects.
cs.AI / 130 / 2609.37950
Video-RSI: Recursive Self-Improvement of Video Understanding Agents via Harness Evolution
Abstract
Video understanding agents acquire evidence through an executable harness that controls what they observe and how they use those observations. However, execution traces contain only the evidence acquired by the current harness, leaving competing explanations for failure unresolved and limiting the basis for self-improvement. We introduce Video-RSI, a framework for recursive self-improvement in which a video understanding agent uses its own language model to revise its harness. Through active video investigation, the model revisits the original training videos to test competing failure explanations with additional observations, grounding proposed changes in evidence beyond the existing trace. Cost-aware harness evolution turns these diagnoses into reusable revisions and determines which revisions to retain by considering both answer accuracy and visual cost. Across our evaluation settings on video understanding benchmarks, the evolved agent improves accuracy while processing fewer frames and achieves competitive accuracy-efficiency trade-offs against existing video understanding agents. These results demonstrate the potential for video understanding agents to improve their own evidence acquisition and use through harness evolution. Code is available at https://github.com/bingjunluo/Video-RSI .
cs.AI / 131 / 2609.37953
Topological Coherence for Self-evolving Multi-agent Systems
Abstract
Complex tasks inherently couple workflow structure, agent responsibility, collaboration, and memory access: task regions delimit responsibility and tool scope, cross-region dependencies give rise to handoffs, and ownership boundaries delimit private and selectively shared memory. Existing methods can jointly optimize agent and communication structures, yet such optimization does not by itself require responsibility, handoff, and memory boundaries to remain consistent with task dependencies. We term this requirement topological coherence. We introduce TOCOMAS, a Topology-Coherent Multi-Agent System. TOCOMAS grounds a task graph in tool interfaces, organizes compatible task nodes into reusable responsibility domains, and derives dependency-induced and profile-conditioned collaboration together with boundary-regulated memory visibility. During online self-evolution, TOCOMAS proposes coupled changes to agent, collaboration, and memory policies, retaining for subsequent tasks only candidates that satisfy structural constraints and improve evaluated reward. Across BBEH, WorkBench, SWE-Bench-Verified, and CoMemBench, TOCOMAS improves task success over baselines across backbones. CoMemBench also shows gains over the self-evolving baseline in verified progress, handoffs, and memory isolation.
cs.AI / 132 / 2609.37968
SelfSearch: Reward-Free Search for Self-Improving Agents
Abstract
Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstream evaluation, which incurs substantial costs and ties the search to the evaluated tasks. We introduce \textbf{SelfSearch}, a reward-free search procedure in which agents modify themselves using records of previous self-improvement episodes. These records capture the reasoning, tool actions, and outcomes of earlier modification attempts, providing concrete experience for improving both task solving and self-modification. Without downstream reward signals during search, SelfSearch improves population-mean success over the initial agent in all six model--benchmark settings, with individual agents gaining up to 11.2 percentage points on Terminal-Bench 2.1. On SWE-bench Multilingual, an agent improves success by \textbf{5.0} percentage points while reducing execution cost by \textbf{38.5}\% on tasks solved by both the initial and evolved agents. SelfSearch achieves competitive task success with evaluation-guided search baselines at lower search cost. With only \textbf{\$4.03} in search cost, it produces a harness that solves \textbf{82.0}\% of Terminal-Bench 2.1 tasks with DeepSeek V4 Flash under the settings of a public nine-harness comparison, matching the top-scoring harness, Codex. These results suggest that experience gained through self-modification can improve agents' downstream capabilities and efficiency.
cs.AI / 133 / 2609.37988
KV-Kaizen: Learning Context-Adaptive Cache Compression Choices
Abstract
As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts LLM throughput negatively, since decoding is memory-bound and decode cost grows with cache size. Recent work alleviates this bottleneck by discarding the least relevant tokens. Eviction introduces a tension, since a one-off decision to discard content may prove detrimental later. Instead, we focus on alternative choices that can lead to cache compression without evicting tokens. We achieve this by learning a selector that is able to produce, based on context, a per-layer cache configuration towards an overall compression budget. The selector operates along three axes: sharing one cache across layers (depth), caching at fewer bits (precision), or truncating the low-rank latent cache representations (rank). We call the resulting method KV-Kaizen, for the many small per-layer choices it compounds. We observe that these interventions taken independently and uniformly over all layers limit achievable compression because they degrade accuracy. Crucially, composing them locally and adaptively to the context can instead preserve accuracy while achieving large memory savings. At inference, the selector runs once, before pre-fill. In evaluations on instruction following and reasoning tasks, our selectors reach the Pareto frontier of accuracy against cache size, against learning-free and post-hoc baselines. On long-context tasks, KV-Kaizen improves on eviction and can be composed with it, reaching a 32x smaller decode-time cache on a 14B model while preserving accuracy. A 4x cache size reduction incurs no accuracy degradation from 7B parameters up, and a compressed model is more accurate than a smaller uncompressed one with the same cache size. Together, these findings support pre-training large models and compressing them only afterwards.
cs.AI / 134 / 2609.37991
Which Attention Heads are like the Human Head? Not the Ones that Compute
Abstract
Brain-AI alignment is often interpreted as a sign that model and brain perform similar computations. Whether the aligned units are causally involved in model computation is rarely checked. On an abstract pattern-completion task (AAABAAA $\rightarrow$ B), we compare LLM attention-head representations with human EEG and test how ablating those heads affects task performance. Alignment and causation dissociate: brain-aligned heads contribute to performance, but their removal is substantially less disruptive than removal of heads selected via attribution patching. We compare two head sets that prior interpretability work defines without reference to the brain: concept vectors (CVs), which represent abstract patterns across formats, and function vectors (FVs), selected for their contribution to correct-answer prediction. Brain alignment shows little association with FV scores, while its association with CV scores varies across models. Among brain-aligned heads, we find recurring attention profiles: one emphasizes distinctive elements (novelty heads), the other repeating elements (repetition heads). The novelty family tracks salience and attends to the same elements that humans look at, yet its removal is less damaging than random ablation on average. Repetition heads contribute modestly to performance and are associated with abstract-pattern representation (CVs). Across 17 models spanning 3B-72B parameters, FV-ranked removal is substantially more disruptive than brain-ranked removal. Brain alignment thus captures how the model reads the stimulus, and only faintly captures how it represents the pattern and solves the task.
cs.AI / 135 / 2609.38006
HARISSA: Inference-Time Self-Checks for Efficient and Safe Local Language Model Deployment
Abstract
Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable than frontier models. The usual remedy for a hard query, escalating it to a cloud model, gives up the privacy and cost advantages of running locally. A deployment that stays local faces two decisions for hard queries instead. First, it can spend more computation on a query, e.g., reasoning before answering, which raises accuracy at a cost in latency, so it must decide which queries are worth the extra computation (efficiency). Second, some queries are beyond the local model, and delivering a wrong answer is worse than deferring the query to a human in the loop, so it must decide which answers are safe to deliver (safety). We show that both decisions can be made from the model's own hidden states. The prefill state, computed before any token is generated, predicts whether the model will answer correctly, and the answer state, at the end of the generated answer, predicts whether that answer is correct. HARISSA fine-tunes the model so that both states predict correctness, then makes both decisions with one policy that cascades through the ways of answering from cheapest to most expensive, skipping a way the prefill state predicts will fail and deferring the query when the answer it stops with is predicted wrong. On a device running a single model, HARISSA is within one accuracy point of chain-of-thought at 2.7 times lower latency. On a server holding four sizes of one model, HARISSA is more accurate than the FrugalGPT and Self-REF cascades at the same latency, and at the same deferral rate the answer state leaves fewer wrong answers than the standard confidence signals in five of six task and setting pairs.
cs.AI / 136 / 2609.38016
Brain-SAD: A Brain-Inspired Safe Autonomous Driving Control Framework with Dynamic Fear-Oriented Constraint on Dual-Policy
Abstract
Constrained Reinforcement Learning has recently gained increasing attention in the field of Safe Autonomous Driving, where the general mechanism is to maximize the expected reward while keeping the overall action risk bounded. In this way, the safety issues arising in AD can be mitigated through constrained actions. However, existing Constrained RL methods still lack dynamics on the imposed constraints. For instance, the action cost adopted by the existing Primal-Dual/soft-constrained methods is often defined as static state-to-cost mapping, and the safe-action projection in hard-constrained methods relies on the static projection with the fixed feasible region boundary estimated from offline demonstrations. The above drawback tightly couples the imposed constraints to the training scenarios, leaving the AD policy hard to handle different interaction scenarios, due to the improper state-level action-cost and the static projection boundary. Consequently, in this paper, we propose Brain-SAD, a brain-inspired safe autonomous driving control framework with dynamic fear-oriented constraints. By perceiving the current vehicle-interaction scene, Brain-SAD generates dynamic fear signal as fear reaction to online decide long-term policy for regular interaction or short-term policy for urgent-collision defense. In such two policy, the above fear-reaction will be constructed as the dynamic fear constraints, respectively reflecting the overall fear cost directly coupled with action-impact, and the dynamic fear boundary of the feasible region derived from different risky neighbors, both of which will in turn serve for the online policy optimization. Experimental results show that Brain-SAD outperforms existing methods, achieving higher success rate in shorter task-completion and collision-recovery time, and exhibits stronger reliability across continuous intersections of fluctuating complexity.
cs.AI / 137 / 2609.38023
PE-EK-PINN: Physics Embedding with Evolving Kernel for Scalable Physics-Informed Neural Networks
Abstract
Physics-Informed Neural Networks (PINNs) embed governing equations into deep learning, but enforce them only through loss residuals, leaving highly oscillatory wave behavior to be discovered by optimization. As a result, methods that achieve relative $L_2$ errors below $10^{-3}$ on standard manufactured Helmholtz benchmarks can fail on practical radiation problems involving singular excitations, absorbing boundaries, and wave fields spanning tens of wavelengths. Architectural physics embedding addresses this limitation by factorizing the field into analytically derived oscillatory kernels and learnable envelopes. However, the kernel dictionary must be manually constructed and scales with the number of elementary units, growing exponentially with the depth of hierarchically structured systems such as antenna arrays and metasurfaces. We propose PE-EK-PINN (Physics Embedded with Evolving Kernels), which treats physics kernels as reusable learned representations rather than fixed analytical inputs. A converged subsystem field is frozen and promoted to an evolved kernel, whose transformed copies are reused to represent higher-level configurations without deriving new governing equations. The resulting hierarchy makes the peak number of active kernels independent of system size and reduces cumulative training cost from $O(N)$ to $O(\log N)$. Experiments on dipole arrays, composite line-source geometries, and cross arrays demonstrate the dramatic training cost reduction, while achieving a reduced or comparable relative $L_2$ error. One notable example is PE-EK-PINN solves a $256$-dipole array more than 30 times faster than direct PE-PINN.
cs.AI / 138 / 2609.38024
Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation
Abstract
An agent skill is a reusable, actionable natural-language artifact that guides an agent to perform a task effectively under a given harness. Recent studies have explored the optimization of agent skills, contributing to a growing collection of publicly available skills spanning diverse tasks, domains, and harnesses. Despite millions of publicly shared skills, existing skill optimization methods largely overlook this accumulated knowledge, instead relying solely on expensive agent rollouts to iteratively refine skills for a target task. To address this, we propose \textbf{Retrieval-Augmented Skill Optimization (RASO)}, a framework that leverages an external skill corpus as prior knowledge throughout skill optimization. RASO retrieves relevant knowledge from existing skills and adapts it to the target task and harness via Cross-Harness Adaptation, accounting for mismatches in both domain and harness. RASO comprises two complementary stages: \textbf{Retrieval-Augmented Skill Initialization (RASI)} constructs a knowledge-grounded initial skill without requiring agent rollouts, while \textbf{Retrieval-Augmented Skill Update (RASU)} iteratively refines the skill by retrieving external knowledge guided by execution feedback. Across four agent benchmarks and two models, extensive experiments show that RASO consistently outperforms baselines without retrieval-augmented skill initialization and updating.
cs.AI / 139 / 2609.38043
UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training
Abstract
Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information the agent receives and when, yet current benchmarks score only the agent and do not directly measure whether the user correctly executed its assigned role. We introduce UserProxyBench, an evaluation layer over the tau-bench family, and the User Fidelity Score (UFS), which measures adherence to the benchmark's private user instructions using task-grounded rubric criteria scored independently of agent success. Holding the agent fixed at GPT-5.5 and varying only the user proxy across 375 enterprise tasks changes mean task reward by 15.2 points, while 24.4% of successful episodes contain a user-specification violation. The dominant failure is premature disclosure: users provide information before it is requested. This behavior has little effect on task reward, yet among successful episodes it causes the agent to make 1.06 fewer tool calls on average, changing the interaction being evaluated while preserving the reward. Finally, across seven proxies we identify an empirical cost-fidelity frontier, enabling practitioners to select the least expensive simulator that satisfies a required fidelity level.
cs.AI / 140 / 2609.38093
Character Training for Risk-Averse Agents
Abstract
Risk aversion in resources could prevent misaligned AI agents from causing catastrophic harm. Misaligned but risk-averse agents would tend to favor safer strategies like making deals with humans over riskier strategies like rebelling. We train agents to be risk averse through character training, finding that persona traits provide a robust mechanism for instilling risk preferences. To do this, we construct a model constitution describing constant absolute risk aversion (CARA) over an agent's resources and instill it through on-policy distillation. Despite never seeing the benchmark's decision format during training, character-trained models are competitive with baselines trained directly on it, and generalise better than them out of distribution on two of our four models. We also modulate different aspects of the constitution, finding that token budget and model choice are the most influential aspect of character training to instill risk aversion. We conclude from these results that character training is a promising and scalable way to instil broad dispositions, which we can use to our advantage in mitigating risk from misaligned AI agents.
cs.AI / 141 / 2609.38098
NeuronEye: Query-Guided Visual Concept Activation for Vision-Language Reasoning
Abstract
Current vision-language models (VLMs) encode visual information in dense hidden states where object identity, spatial layout, and local attributes are implicitly entangled rather than explicitly disentangled, limiting their ability to isolate and modulate the specific visual evidence required by a given language query. Inspired by sparse population coding and top-down modulation in biological vision, we introduce NeuronEye, a plug-in framework that constructs a sparse, concept-level neuron vocabulary from intermediate VLM representations and selectively activates query-relevant visual concepts during inference. NeuronEye decomposes vision-token states into an overcomplete sparse basis organized by concept-level clusters, uses the language query to activate relevant clusters and localize the patches where selected concepts are expressed, and injects the focused evidence back into vision tokens. A complementary suppression mechanism attenuates dominant perceptual directions to preserve weaker but relevant cues. All operations run in a single forward pass over a frozen VLM backbone. On Qwen2.5-VL-7B, NeuronEye raises CV-Bench overall accuracy by +3.1 with gains of +9.5 on Distance, and improves BLINK Multi-view by +8.3, with similar trends on LLaVA-1.6-7B. These results suggest that sparse neuron vocabularies can serve not only as post-hoc interpretability tools but also as active interfaces for concept-level visual reasoning.
cs.AI / 142 / 2609.38120
Stochastic World Models for Verifying Vision-Based Neural Feedback Systems
Abstract
Verifying a vision-based neural feedback system requires a model of the observations its controller acts upon. Such a model must capture the variation the sensor produces, while remaining tractable for closed-loop analysis. Generative adversarial networks (GANs) have served as perception surrogates, but they are large, reproduce complex scenes poorly, and are hard to verify. We explore stochastic world models as a richer class of perception surrogates. We train a world model with physically grounded latents, built from operations that standard verifiers bound. It reproduces held-out frames more faithfully than GAN surrogates with up to 130 times as many parameters. To verify these surrogates, we develop a procedure that combines falsification, adaptive refinement, symbolic, and backward analyses. On an emergency braking benchmark with a GAN surrogate, our procedure resolves the entire state space, 38% of which the state-of-the-art verifier left unresolved. On the RGB version of the benchmark, where no verification results have previously been reported, our procedure resolves over 80% of the state space with a world model surrogate.
cs.AI / 143 / 2609.38142
AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation
Abstract
A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor's future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice less strongly than those of other corrections. Keeping them less often than the rest improves the model's eventual performance compared to learning from every correction. Motivated by this, our method, Advisor Self-Distillation (AdviSD), pairs outcome-based reinforcement learning with self-distillation from a feedback-conditioned copy of the advisor selectively. Reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of the difference to select decisions for supervision. This approach does not require executor likelihoods or additional executor rollouts. Experiments with Qwen3-8B advisors for Gemini and Claude show that AdviSD outperforms advisor-GRPO by 4.2-6.4 percentage points on BFCL-v3 and by 3.9-5.1 score points on EnvScaler. The trained advisors generalize to out-of-domain tasks and transfer across different executor versions and model families. AdviSD also beats matched-count random selection, supporting the value of its selection rule.
cs.AI / 144 / 2609.38143
Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
Abstract
Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models' weights remain fixed. To make the Builder's experience reusable, we introduce Meta-Skill: principles specifying when support is needed and what resources to provide. The Builder learns these principles from Target's execution feedback on the development set, then uses the frozen skill bank to construct harnesses for unseen tasks. Across Harness-Bench and NewtonBench, full-bank meta-skills improve macro-average performance by 8.95 percentage points over no-skill construction, and 12.02 points over direct delivery of the same bank to the Target. These results highlight the value of translating experience into executable support. Gains when the same model serves both roles further suggest a path to system level self-improvement through learning to build better environments.
cs.AI / 145 / 2609.38147
Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning
Abstract
As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic meta-reasoning, an inference-time harness that makes these choices an explicit and structured reasoning process. Workers carry out the task-level computation, while a controller consolidates what the run has established, explores next options, assesses what each option is worth under the remaining budget, and dispatches the chosen work with context drawn from persistent memory. Between decisions the controller carries only a compact account of the run rather than replaying its full history. Our baselines span production coding agents and research harnesses, together with a Direct Control Agent using the same workers and compute budget allowance. On ProgramBench, which tests long-horizon agentic capability through program reconstruction, meta-reasoning achieves 71.5% with GPT-5.5 against 58.0% for Codex; with Opus 4.8 it achieves 67.2% against 65.5% for Claude Code. On the other benchmarks, spanning abstract reasoning, multi-domain long-horizon reasoning, and proof generation, it gains between 3.6 and 4.2 points over direct control, averaged across three frontier models. It keeps improving over the tested budget ranges where direct control plateaus, though its overhead can hurt at small budgets. Artifact-graph analysis reveals more reuse of earlier work, higher coverage of correct solutions in most settings, and nonuniform gains in final selection. These results indicate that spending computation on structured control becomes more important as agents scale to longer runs.
cs.AI / 146 / 2609.36739
Frontier Autolab: Organizational Memory, Adversarial Dissent and Temporal Leakage in Multi-Agent LLM Firms Across Fifty Years of Technological Change
Abstract
Multi-agent LLM systems are increasingly structured like organizations, with roles, critics and shared memory, yet they are evaluated on tasks that last minutes. We ask how such an organization behaves when the ground it stands on keeps moving. Frontier Autolab is a long-horizon testbed in which one simulated firm, voiced by sixteen role personas and a dedicated Red Team, must re-found itself in nine technology eras from 1990 to 2040. Each era is temporally gated: the firm decides from a dated briefing, a historian-judge then reveals what happened and scores the decision on a five-dimension rubric, and lessons enter a persistent Playbook. Six eras are scored against history, one against the live market and two are open forecasts. Across four trajectories (36 era decisions, 180 subscores) we find a consistent foresight-commitment gap: in all 24 historically scored eras the judge rated the firm's recognition of the coming shift above its choice of where to build (mean gap 1.9 points on a 10-point scale), because boards chose the layer their existing assets could reach. Organizational design shaped long-run character. A Red Team armed with numeric kill gates produced fifty years of gated pilots and no product, and the rubric rated this firm highest; firms whose memory stored market-structure lessons pivoted every era, while a firm whose memory stored only validation procedure kept one method throughout. We also show why such results are hard to trust. Scores rise across eras in every run while the judge's own hindsight subscore falls (within-run r = -0.58), so apparent learning is confounded with recall of history, and we trace further distortions to self-judging, briefing selection and score aggregation. We release all records and an API harness, and specify fictional and post-cutoff eras that would turn the testbed into a benchmark.
cs.AI / 147 / 2609.38089
Neural topology optimization of ship structures under propulsion machinery vibrations
Abstract
Ship structural vibrations contribute to noise, fatigue, and equipment damage, while dynamic-compliance topology optimization can produce pathological designs near resonance. This study extends neural-reparameterized topology optimization using a convolutional Kolmogorov-Arnold network (KATO) to forced-vibration design with active input power (AIP) as the objective. Applications include a 100 Hz engine-supporting deck panel and an 18 Hz thruster foundation frame. Helmholtz PDE filtering and Heaviside projection control feature sizes and manufacturing tolerance. Across both deck families, all eight optimized layouts reduce AIP relative to size-optimized references and, after finite-depth extrusion, also achieve lower static compliance. For unrestricted, manufacturing-aware, and stress-aware frame variants, KATO matches GCMMA in AIP within 0.5 dB while yielding 22-36x lower static compliance after matched-volume binary re-analysis. In a near-resonant 300 Hz case, both methods reduce initial AIP by more than 32 dB; KATO maintains a connected design, achieves 59x lower binary static compliance, and reduces maximum AIP over 1-500 Hz by 2.7 dB. KATO runs 6.4-10.4x faster than GCMMA for the implemented stress-aware formulations. The results demonstrate neural AIP-driven topology optimization as an efficient approach for designing connected, feature-size-controlled ship structures with improved forced-vibration performance.
cs.AI / 148 / 2609.35965
Systematic Multi-Agent Vision-and-Language Navigation: Formulation, Benchmark, and Method
Abstract
Vision-and-Language Navigation (VLN) has largely focused on a single agent following a single instruction, yet many real-world applications require teams of robots to tackle tasks beyond the capabilities of any individual agent. We present Systematic Multi-Agent Vision-and-Language Navigation, providing, to our knowledge, the first systematic formalization of multi-agent VLN as a constrained coordination problem: each mission consists of subtasks carrying dependency and resource constraints (presence locks and holding chains). A verified four-stage crafting pipeline instantiates the task as MAVLN, comprising 11,724 episodes across 145 scenes with teams of up to four agents under three instruction regimes, accompanied by tailored constraint-aware metrics. We further present TRISS, a coordination-ready navigation system coupling an LLM-based subtask scheduler, a shared topological memory that turns each agent's exploration into team knowledge, and a conflict-aware execution mechanism that realizes simultaneous intentions as collision-free routes. Extensive experiments establish TRISS as a comprehensive baseline and reveal substantial room for improvement across scheduling, planning, and execution, highlighting the challenges of coordinating under MAVLN task constraints. Project page: https://xyz9911.github.io/mavln.
cs.AI / 149 / 2609.36101
One Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models
Abstract
Contrastive vision-language models learn shared embedding spaces by aligning matched image-text pairs, yet their representations remain separated by a modality gap. Prior work reports divergent effects of modifying this gap: reducing it can improve zero-shot classification and cross-modal alignment, whereas removing gap-related structure can degrade image-text retrieval. In this paper, we provide a unified geometric explanation for these task-dependent effects. Across CLIP and SigLIP encoders, we find that a single dominant direction captures 94.4-99.9% of the squared norm of the image-text mean separation, revealing that the mean-separation component is approximately rank-one. A decomposition of the similarity score then identifies three task-specific roles. In zero-shot classification, query-side fixed gap-offset subtraction is exactly equivalent to an additive class bias. In standard cross-modal retrieval, projecting out the gap direction and renormalising residuals discards candidate-specific norm information, inducing a multiplicative ranking distortion; a geometry-derived exponent tracks the grid-search optimum (Spearman rho = 0.93) and restores performance in some settings, although the gains transfer unevenly. In mixed-modal retrieval, the gap direction sorts candidates by modality; its removal can improve cross-modal ranking, unlike random or non-gap controls. Residual semantic structure after removal defines the limits of the rank-one account. Together, these results explain why gap modification can improve, degrade, or restore performance across downstream settings. By clarifying when and why gap modification changes model behavior, this account provides a principled basis for selecting gap interventions in similarity-based vision-language systems across evaluated downstream tasks.
cs.AI / 150 / 2609.36224
Mutually Adversarial Self-Training with Evolving Data for Unified Multimodal Models
Abstract
Unified multimodal models (UMMs) combine image generation and visual understanding in a shared backbone. Since generation and understanding are inverse tasks, recent studies self-train UMMs by letting the two branches cooperatively supervise each other. We introduce MATE (Mutually Adversarial self-Training with Evolving data), a reinforcement-learning-based post-training framework in which the two branches instead challenge each other, and the challenges evolve as the model trains. MATE lets generation and understanding take turns to be challenger and solver. Given an image, the understanding branch proposes several candidate descriptions that the generation branch must turn back into similar images, and vice versa. The candidates are screened for consistency with the image or prompt they were proposed from, and the solver is trained on the candidate it handles worst. The adversary thus comes from the model's own outputs, and no separate adversary is trained. Moreover, the candidates that defeat one branch become the sources of the next challenges to the other in the next epoch, which keeps the challenges evolving with the model and turns the training into self-play in data space. On Janus-Pro-1B, MATE improves GenEval by 2.4 points, DPG-Bench by 1.7 points, and the average over nine understanding benchmarks by 0.7 points, while strengthening consistency across repeated image-text cycles.
cs.AI / 151 / 2609.36243
Think Before You Restore: Risk-Aware Manchu Manuscript Restoration with Stroke-Guided Attention
Abstract
Full-page blind restoration of historical Manchu manuscripts is challenging due to scarce annotations, unknown degradation regions, and fragile connected strokes. Generic restoration models may improve visual quality but often modify intact content, leading to over-restoration. We propose SAGE-Restore (Stroke-Aware Gated rEstoration), a selective restoration framework that first assesses where restoration is needed and then uses this assessment to guide restoration candidate generation and pixel-level selection. Its encoder predicts patch-level repair probabilities from complementary appearance and stroke-structural cues to condition restoration candidate generation, while the corresponding repair logits are refined into a pixel-level soft gate that selectively controls where the restoration candidate is applied. We further introduce a fidelity-aware evaluation protocol that jointly measures degraded-region recovery, intact-content preservation, and their balance. SAGE-Restore achieves the highest R-Recovery (0.463) and RFS (0.626), while maintaining high U-Fidelity (0.968), demonstrating an effective balance between restoration and content preservation.
cs.AI / 152 / 2609.36364
Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generation
Abstract
Standard video generators do not natively compact historical context into reusable memory tokens. As generation continues, the growing history makes it increasingly difficult to retain information from earlier frames due to long-context degradation. Key-frame-based approaches address this challenge by retaining selected past frames, but can discard information needed for future generation. Rather than relying on frame selection alone, we study whether a frozen video generator can supply the supervision needed to learn a compact representation of the history. We propose Prediction-Aligned Context Compaction (PACC), which uses a learned compressor to aggregate information across past frames into compact memory tokens. We train the compressor through on-policy distillation, using the same frozen generator both as a student when conditioned on compressed memory and as a teacher when conditioned on the full history. The student generates continuations, while the teacher provides targets for the same noisy inputs at each denoising step. Only the compressor is updated to align the student's predictions with these targets. We evaluate PACC on MBench, which jointly measures memory-event coverage and consistency. PACC outperforms the strongest baseline by 6.63 points on Causal-rCM and 3.19 points on Causal Forcing. Evaluation on VBench-Long using MovieGen prompts further shows that PACC produces minute-long videos with generation quality competitive with baselines. Together, these results show that learning to compact historical context can improve long-video memory without modifying the underlying generator.
cs.AI / 153 / 2609.36442
Online Versatile Incremental Learning: Towards Class and Domain-Agnostic Adaptation at Any Time
Abstract
Continual learning enables vision systems to adapt to ever-changing data distributions. Despite significant advances, existing approaches fail to capture continuous and concurrent shifts in classes and domains, a critical capability for real-world deployment. This work introduces Online VIL (Online Versatile Incremental Learning), a novel scenario where class concepts and visual domains evolve simultaneously online without explicit boundaries. To better adapt to the challenges of such dynamic environments that more closely resemble real-world conditions, we propose a novel framework TopFlow, Topology preservation with Flow matching representation that contains two complementary mechanisms: Domain-agnostic Flow Matching (DFM) and Global Topology Preservation (GTP). DFM guides the model to have domain-agnostic representations by integrating the geodesic flow kernel into contrastive learning. In contrast, GTP maintains the global structure of the feature space without explicitly storing past examples. Our extensive experiments demonstrate that TopFlow effectively addresses the limitations of existing methods within the Online VIL scenario, achieving state-of-the-art performance in challenging Online VIL. The proposed methods suggest potential directions for building continual learning systems in realistic dynamic environments. Our implementation code is available at https://github.com/KU-VGI/Online-VIL.
cs.AI / 154 / 2609.36492
Benchmarking Vision-Language Models on Synapse Detection and Proofreading in Connectomics
Abstract
We benchmarked vision-language models (VLMs) on the decisions annotators take when inspecting electron microscopy images in connectomics: synapse detection (presence and polarity) and proofreading (split errors and merge errors). For synapse detection, we evaluated 19 open and 2 closed models across various architectures and sizes under zero-shot, four-shot in-context learning and LoRA settings, against specialist models, on datasets constructed by us using public resources. For proofreading, we evaluated 3 open and 2 closed models on the ConnectomeBench2 dataset, with cross-species transfer from fly and mouse to human and zebrafish. Most models were at chance zero-shot; a few examples helped mainly the closed and largest open ones. LoRA on a few thousand labels brought open models level with specialist models. When evaluated on unseen species, the best adapted VLMs outperformed specialist models trained on the same data in identifying merge errors. The project will be publicly available upon acceptance.
cs.AI / 155 / 2609.36557
How Medical VLMs Underutilize Their Vision Encoders: A Dermatology Perspective
Abstract
Medical Vision-Language Models (VLMs) show significant promise for clinical image understanding, offering accurate diagnosis with interpretable reasoning. However, a critical performance gap exists between their strong vision encoders and the full multimodal model: in dermatology, the MedSigLIP encoder outperforms MedGemma by an average of 10.26 percentage points even when both use zero target-task labels; few-shot linear probing provides further evidence of strong visual representations. This gap motivates an investigation of how visual information is used in end-to-end diagnosis and why plausible-sounding predictions can lack grounding in image evidence. Using dermatology as our primary testbed, we systematically investigate three hypotheses for this phenomenon. We further provide a mechanistic analysis of the model's internal attention patterns, showing that a simple describe-then-decide prompting strategy increases vision attention by 30-40% during generation. Task-specific fine-tuning improves dermatology classification but reduces cross-domain medical question-answering performance in our evaluation. To address these challenges, we combine label-free prompting with low-label encoder-assisted reranking while keeping the VLM frozen. We validate the interventions across five VLM backbones in dermatology and provide supporting representation and attention analyses across additional medical modalities.
cs.AI / 156 / 2609.36560
FM-ReID: Selective Competitive Token Routing for Object Re-Identification
Abstract
Object re-identification (ReID) faces a recurring challenge: different identities can share highly similar global appearances, while the cues that distinguish them are localized, heterogeneous, and visible only under particular viewpoints. This challenge arises in animal ReID through markings, contours, and scars, in person ReID through subtle clothing and accessory cues, and in vehicle ReID through localized appearance details. Although visual foundation models encode such information in dense tokens, a single holistic descriptor can obscure discriminative local signals. We propose FM-ReID, an end-to-end framework that formulates local representation learning as selective competitive token routing. Its Competitive Fine-grained Mining module uses multiple mining queries and a residual query to compete for dense DINOv3 tokens. Above-prior selection retains tokens preferentially allocated to each mining query, while the residual slot receives tokens excluded from the retrieval descriptors. The resulting multi-query descriptors are jointly trained with a holistic representation for retrieval, without fixed spatial partitions or equal-area constraints. FM-ReID achieves strong results on animal, person, and vehicle ReID benchmarks, supporting competitive token routing as an effective way to augment holistic foundation-model representations.
cs.AI / 157 / 2609.36599
Scaling Video Generation for Reasoning: At What Cost?
Abstract
We study whether scaling video generation enables models to reason about hidden information from the past frames, and at what computational cost. Our controlled benchmark requires predicting nine prescribed moves of an initially solved 2x2x2 Rubik's Cube from a fixed view of three faces. Correct predictions require inferring how actions change hidden states, and the simulator provides exact ground truth for evaluation. Models learn plausible cube geometry early, while correct sticker configurations require substantially more training. Although validation MSE follows approximate power-law scaling, lower MSE loss does not reliably indicate downstream reasoning capabilities. Smaller autoregressive models achieve higher state accuracy with limited compute, while larger models reach higher accuracy after more training. At roughly 0.1 PF-days, the 70M-parameter model correctly predicts the visible sticker configuration in 44.6% of post-action frames, compared with 0.3% for the 1B model, which reaches 83.7% at 3.14 PF-days. Symbolic state supervision raises the 20M model's frame accuracy from 31.1% to 67.3% at the same training-data budget, suggesting that learning representations of state changes can complement scaling.
cs.AI / 158 / 2609.36756
NesTok: Nested Self-Aligned 1D Tokenizer for Autoregressive Image Generation
Abstract
One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressive (AR) models to flexibly trade off generation quality against computational cost using a single tokenizer. However, existing approaches based on nested dropout often fail to fully exploit the representational capacity of the tokenizer, resulting in suboptimal performance in both image reconstruction and generation. In this work, we introduce NesTok, a nested self-alignment framework tailored to dynamic visual tokenizers. NesTok introduces cross-length training, which jointly optimizes reconstruction across token lengths while using the full-length sequence to guide shorter counterparts, enabling shorter token sequences to approach the reconstruction quality of full-length sequences. On ImageNet, NesTok improves substantially over standard training and achieves an rFID score of 0.98. On downstream image generation, it achieves the state-of-the-art gFID score of 1.46 on ImageNet 256$\times$256 among existing variable-length autoregressive image generation methods. Code will be available at https://github.com/jaiwei804/NesTok.
cs.AI / 159 / 2609.36759
Dual-Mode Low-Rank Learner with Bridge-Prototype Ensemble for Vision-Language Class-Incremental Learning
Abstract
Benefiting from transferable visual-textual alignment, CLIP has been widely adopted for class-incremental learning (CIL). However, existing learners either repeatedly update components shared across tasks, leading to knowledge overwriting, or overly isolate new-task updates, hindering the reuse of CLIP's transferable knowledge and limiting plasticity. Moreover, the text-based or bimodal classifier designs still fail to effectively integrate complementary information from the visual and textual modalities. To address these challenges, we introduce DuLBE, which couples dual-mode low-rank learning with a bridge-prototype ensemble classifier for exemplar-free CIL. DuLBE allocates two visual low-rank update modes according to the gradient demand and uses gradient routing to coordinate them: a compact and rewritable shared mode is selected from historically occupied visual directions to reuse transferable knowledge, while residual modes provide low-interference channels for task-specific variations. Building on the resulting stable inter-modal structure, we further construct geodesic bridges between visual prototypes and text embeddings on the unit hypersphere, and ensemble reliable bridge prototypes to compensate for the modality-gap limitations of textual decision boundaries. Extensive experiments under multiple settings show that DuLBE achieves state-of-the-art CIL performance while retaining the high parameter efficiency of low-rank tuning.
cs.AI / 160 / 2609.36838
On-Policy Visual Evidence Distillation
Abstract
Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teacher on student-generated interaction trajectories. However, image operations change the evidence available for subsequent reasoning, so local errors in evidence acquisition (Acquire), reading (Read), or answer grounding (Ground) can propagate through the trajectory and lead to incorrect answers. Existing multimodal OPD methods primarily construct or contrast auxiliary views of the original image to strengthen supervision, without explicitly modeling the connections between student actions, resulting observations, and subsequent reasoning. This limits their ability to provide corrections tailored to different failure stages. We introduce Reflection on Visual Evidence (ReVuE), an on-policy distillation method for visual agents. ReVuE compares multiple student-generated trajectories for the same query, summarizes the observed visual evidence, and diagnoses the first failure across the Acquire, Read, and Ground stages. The resulting reflections provide training-time context for the teacher. We group and reweight token-level distillation losses according to how strongly these reflections affect the teacher's predictions. This design translates trajectory-level evidence diagnosis into targeted token-level supervision, guiding students to improve their visual evidence acquisition and reasoning. Across 11 benchmarks spanning the Qwen2.5-VL and InternVL3.5 model families, ReVuE outperforms all evaluated OPD baselines in weighted-average scores for perception, mathematical reasoning, and general tasks. ReVuE also reduces redundancy in reasoning and tool calls while improving tool-call accuracy and task accuracy. Code is available at https://github.com/sylvain-wei/ReVuE
cs.AI / 161 / 2609.36937
WeLike2Party! In-Context Motion Transfer for Multi-Human Image Animation
Abstract
Human image animation aims to transfer motion from a driving video to subjects in a reference image. Despite remarkable progress in video generation, achieving high-fidelity animation of multiple interacting subjects remains a challenge. Many existing approaches rely on explicit motion representations such as 2D skeletons or parametric body meshes and struggle to preserve identity-motion binding under inter-person occlusion. To address this limitation, we propose WeLike2Party, a multi-human animation framework built on direct in-context video conditioning without explicit pose or mesh extraction at inference. We further introduce Reference Asymmetric RoPE Conditioning to preserve fine-grained appearance details, and Identity Binding Supervision to associate each reference identity with its intended motion trajectory. To support cross-identity training, we construct MotionTwin, a large-scale synthetic dataset comprising 14.4K cross-identity video pairs with shared subject and camera motions, totaling 84.3 hours of photorealistic video. We additionally present MotionTwin-Bench, a cross-identity benchmark specifically designed to evaluate subject-level visual fidelity and identity-motion binding. Extensive experiments on MotionTwin-Bench and real-world videos demonstrate that WeLike2Party outperforms recent state-of-the-art methods in subject-level visual fidelity, identity-motion binding, and overall perceptual quality, particularly in multi-person interactions with substantial occlusion.
cs.AI / 162 / 2609.36995
Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation
Abstract
Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges. Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction. Meanwhile, directly reusing bidirectional score models in causal Distribution Matching Distillation (DMD) creates a mismatch between generation and scoring contexts. We address these challenges with Salt++, a two-stage post-training framework comprising Causal Self-Flow (CSF) and context-aligned autoregressive DMD. CSF exploits contextual information asymmetry by varying the history while keeping the noisy target fixed: a noise-mixed-history student aligns its intermediate representations with those of a clean-history exponential-moving-average teacher. This self-supervised signal encourages the student to extract semantic information and improves cross-modal alignment. Context-aligned AR DMD shares the causal mask and prefix across generator sampling, fake-score training, and real-score evaluation to match generated and reference distributions under a block-conditional KL objective. With calibrated teacher guidance, it performs clean-prefix few-step distillation and then adapts to generated histories without switching objectives or requiring separate consistency distillation. At 480p, Salt++ improves visual and motion quality by 57% and 45% over OmniForcing on JavisBench under the same 4-step causal setting. A separate scale-wise post-training stage extends Salt++ to 4-step $1664\times960$ generation, outperforming bidirectional LTX-2 on six of seven reported metrics. Project page: https://xingtongge.github.io/Saltpp
cs.AI / 163 / 2609.37002
Visual Parallel Search: Learning to Search High-Resolution Images with Parallel Tile Inspection and Adaptive Zoom
Abstract
High-resolution visual question answering often fails because a multimodal model does not acquire the small, spatially localized evidence needed to answer a question. Sequential zooming can recover detail, but it asks the main model to choose a region before obtaining a reliable overview. We introduce VPS, a visual parallel-search framework in which a main agent first invokes grid_search to inspect image tiles in parallel with question-conditioned sub-agents, and then adaptively invokes zoom_in PSisual Parallel Search improves mean accuracy over dedicated zoom-only search in 14 of 15 same-model comparisons, with gains up to 8.0 points and especially strong improvements for smaller main models. ZoomBench retains an approximately 3.2-point gain at every tested size. We further develop a supervision pipeline with hint-free verification and a paired role-specific GRPO surrogate for learning the controller and tile-reader roles. SFT improves observed accuracy on all five benchmark splits, including a 4.17-point gain on HR-Bench 4K. Role-specific RL further reshapes search behavior: main-only RL reduces mean tool use from 2.65 to 2.11 with similar pass@1 in an internal four-response evaluation, while external accuracy changes are mixed. Joint training reveals an asymmetry between local evidence reading and global search control. Together, these results support VPS as an effective inference-time scaffold and a trainable decomposition for visual evidence acquisition.
cs.AI / 164 / 2609.37013
Embedded Bi-Temporal Building Damage Assessment for On-Board Data Reduction
Abstract
Rapid assessment of building damage after natural disasters is essential to support emergency response. Earth Observation satellites can acquire relevant imagery shortly after an event, but exploitation is limited by uplink and downlink capacity and by ground-processing latency. We address this with a bi-temporal building damage assessment pipeline built on a siamese detector derived from YOLOX, designed to compress information at both ends of the ground/space link. On the ground, pre-disaster reference images are encoded into a compact latent space -- compressed by up to a factor of 64 -- and uplinked to the satellite. On board, this reference is compared with a fresh post-disaster acquisition so that the downlink carries only actionable object-level products, bounding boxes and damage classes, instead of full scenes. This cuts the data exchanged in both directions, while on xBD the strongly compressed reference still preserves most of the detection performance. Because on-board acquisitions suffer from residual pre/post co-registration errors, we introduce a latent-space shift estimation and correction module that regresses the global offset from the coarse feature level and realigns the post-disaster features before fusion. It substantially improves robustness to de-registration -- especially under large shifts, where fusion-only variants collapse -- while also raising nominal accuracy and remaining compatible with the strongest compression. We finally port the pipeline to two embedded targets, a Xilinx Versal VCK190 and an NVIDIA Jetson AGX Orin, and report hardware performance (latency, throughput, power efficiency). The core detector and its compression port cleanly to both, but the operators needed for long-range robustness survive only on the Jetson GPU, whereas the Versal DPU does not.
cs.AI / 165 / 2609.37030
MotionInsight: Diagnosing Object Motion Deficiencies in Generated Videos
Abstract
Despite rapid progress in video generation models, they still exhibit obvious motion deficiencies, often manifested as incorrect object motion. However, most existing video quality evaluations focus on aesthetic quality or text-video alignment. To address this gap, we study object-centric motion fidelity assessment, evaluating target objects along object consistency, motion continuity, and physical plausibility. To achieve this, we first introduce VidMotion, a diagnostic dataset of 6,879 videos with designated moving objects and fine-grained annotations including dimension-wise scores and failure causes. We further propose MotionInsight, a diagnostic evaluator that shifts assessment from implicit RGB-frame observation to explicit motion-space diagnosis. By constructing motion-aware representations, MotionInsight makes subtle motion deficiencies more observable. We also introduce motion-specific rewards during GRPO to transform observed motion into a diagnostic assessment. Experiments demonstrate that MotionInsight provides an effective basis for diagnosing object motion deficiencies, producing human-aligned scores along three dimensions and grounded explanations.
cs.AI / 166 / 2609.37243
Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features
Abstract
Cross-modal knowledge distillation transfers knowledge from a teacher modality to a student modality. Existing feature-level alignment methods typically assume that teacher and student features reside in structurally alignable representation spaces. However, this assumption does not hold when cross-modal features are structurally heterogeneous and lack clear unit-level correspondence, such as 2D spatial visual grids and 1D temporal audio sequences, thereby limiting the applicability of feature-level alignment. To address this challenge, we propose a cross-modal distillation framework that enables effective knowledge transfer across structurally heterogeneous feature spaces via a vector-quantized codebook. Specifically, teacher features are abstracted into a set of vector-form codes regardless of their original feature structure, and the selected codes serve as concept-level anchors for student learning. Code selection is guided by both task relevance and student compatibility, allowing the student to receive transferable teacher knowledge without requiring direct unit-level feature alignment. Experimental results across diverse cross-modal distillation scenarios demonstrate the effectiveness of the proposed framework on classification and semantic segmentation tasks.
cs.AI / 167 / 2609.37250
V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents
Abstract
World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrained at scale, a prominent line of recent WAMs inherits both predictive knowledge and the models in which it was learned. We ask whether a predictive visual latent space induced by large-scale predictive pretraining can instead provide a sufficient foundation for effective WAM learning without inheriting a complete pretrained visual generative model. To answer this question, we introduce V-JEPA Policy, a simple framework that builds a WAM on the latent space of a frozen V-JEPA 2.1 encoder. An instruction-conditioned future-latent predictor and a flow-matching action expert are jointly learned from scratch in a single downstream stage, with the predictor's future-informed context key--value states conditioning action generation. With 0.9B total parameters, of which 0.6B are trainable, V-JEPA Policy achieves competitive performance with representative WAM and vision-language-action baselines across LIBERO, LIBERO-Plus, and RoboCasa-GR1. Comparing visual foundations under the same downstream framework and training budget identifies V-JEPA latents as more effective than the discriminative, reconstructive, and video-understanding-oriented alternatives, particularly under distribution shifts. Beyond task-specific learning, pretraining the predictor on DROID video--instruction pairs without action labels and adapting it into a WAM yields substantial gains in downstream control and out-of-distribution generalization. Together, these findings establish predictive visual latents as a foundation for effective WAM learning from task-specific demonstrations and for transferring future-modeling knowledge acquired from broader in-the-wild videos. Our code is available at https://github.com/breez3young/VJEPA-Policy.
cs.AI / 168 / 2609.37264
UniAfford: Token-Routed Multitask Learning for Generalizable 2D-3D Affordance Perception
Abstract
Affordance perception aims to localize actionable regions supporting embodied interaction, yet 2D and 3D affordance grounding have evolved as separate problems, with different task definitions, supervision formats, datasets, and evaluation protocols. This fragmentation limits the learning of transferable object-affordance semantics across visual and geometric spaces. We propose Token Router for Tasks, a multitask training paradigm for MLLM-based systems that routes contextual hidden states to task-specific branches without requiring the language head to generate predefined markers. Routed states are supervised directly by branch-specific objectives, enabling dense prediction losses to shape shared MLLM representations. We instantiate this paradigm as UniAfford, a unified framework for generalizable 2D-3D affordance perception, together with UniAfford-Data, a dataset integrating pixel-level 2D annotations, point-level 3D annotations, and language instructions under a shared object-affordance taxonomy, supporting heterogeneous supervision through semantic-level 2D-3D pairing. UniAfford adopts an MLLM as a shared semantic hub and a modality-aware token router to produce image- and point-cloud-affordance queries. These queries respectively condition a SAM-style pixel decoder and a SONATA-based point decoder, enabling flexible 2D, 3D, and joint affordance inference from image-only, point-cloud-only, or paired multimodal inputs. Experiments demonstrate strong zero-shot generalization across 2D and 3D affordance benchmarks without target-specific fine-tuning, alongside state-of-the-art branch-wise performance under modality-isolated protocols. Ablations validate token routing, joint 2D-3D supervision, and decoder coupling, while language-head diagnostics show that routed latent states carry meaningful object-affordance semantics. Project page: https://4dvlab.github.io/UniAfford
cs.AI / 169 / 2609.37287
VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics
Abstract
Image translation is a fundamental capability of multimodal models for multilingual applications, requiring visual understanding and meaning preservation across languages. However, existing benchmarks have limited language coverage and often lack explicit image-specific evaluation criteria, making it difficult to comprehensively assess this capability. To systematically evaluate this capability, we introduce VISTA-Bench, covering 22 languages and 10 domains, and develop an image-specific rubric evaluation protocol. The benchmark combines sampling for language and scenario coverage with model-assisted, human-verified annotations that group related text into coherent semantic units and provide multilingual reference translations. The rubrics specify essential content, semantic relations, and acceptable translation variants, yielding separate output-based scores for translation quality and the preservation of visual and knowledge-dependent information. We conduct extensive evaluations of 16 mainstream models, including 12 multimodal models and four text-input models, and provide systematic analyses across languages, domains, and evaluation dimensions.
cs.AI / 170 / 2609.37349
TAEC: Trajectory-Aware Evidence Coordination for Multi-Step Visual RAG
Abstract
Multi-step visual retrieval-augmented generation (RAG) answers complex questions by repeatedly retrieving visual evidence, updating an intermediate state, and deciding whether to continue searching or answer. Yet retrieving relevant evidence does not ensure its effective use throughout the reasoning trajectory. As multi-step reasoning progresses, redundant sources occupy context capacity needed for missing evidence, observations tied to resolved requirements or unproductive searches linger in context, and visual sources are revisited with insufficient detail for fine-grained reading. We term this loss of usable evidence over a reasoning trajectory trajectory-level evidence utilization degradation. To address it, we propose Trajectory-Aware Evidence Coordination (TAEC), a training-free framework that coordinates evidence use around unresolved answer requirements. TAEC tracks these requirements in a shared trajectory state to guide which evidence enters the context, how accumulated memory is retained, and at what level of detail visual evidence is examined. Under a unified evaluation protocol on ViDoSeek, SlideVQA, and MMLongBench-Doc, TAEC achieves the best overall performance against leading training-free visual RAG baselines, with the highest average accuracy across multiple proprietary vision-language models. These results demonstrate that aligning evidence with evolving reasoning needs improves evidence use throughout multi-step visual RAG.
cs.AI / 171 / 2609.37378
Do-JEPA: From Masking to Intervention in Latent World Models
Abstract
Latent world models are trained to predict what happens next, so nothing in their objective separates what an action caused from what merely co-occurred with it. Object-masking models such as C-JEPA intervene on what the predictor can see; we intervene on what physically happens. From one saved simulator state we run the dynamics under an action $a$ and under a reference action $a_{\varnothing}$, and train the model to predict the difference $Δz=z^{a}-z^{a_{\varnothing}}$ between the two latent futures. The resulting objective, Do-JEPA, has an effect loss, a support loss (where the action enters), a propagation loss (where its effect travels) and invariance losses (what must not change). In a synthetic system with object-aligned variables, support supervision finds the directly intervened object in 99.95% of test cases, where a sparse action mask sends the action to a nuisance slot in every case, and response-onset supervision recovers the ring-shaped propagation graph (edge AUROC 0.975 vs. 0.624). From pixels, the effect loss beats a control trained on exactly the same data: it lowers latent effect error by 28.4% on an end-to-end LeWM model and physical effect error by 13.5% when trained and tested on natural action sequences, and on three independently generated CausalWorld benchmarks it lowers responsive effect error by about 20% under physics shifts and the latent context sensitivity of predicted effects by 66%. Trained from scratch it costs factual accuracy; fine-tuning an existing model with it removes this cost. Together, these results show that intervening on the world, rather than on what the model sees, helps latent world models predict what their actions cause.
cs.AI / 172 / 2609.37426
LazySloth: Bounded LLM-based Lazy Tree Search for Fast Long Video Comprehension
Abstract
Modern vision-language models (VLMs) have shown promising results in long-video understanding due to the rich semantic information they can capture. However, most methods focus on coarse captioning of extracted image frames that are computationally inefficient and require models with large context windows. While past work has explored efficient methods through multimodal retrieval-augmented generation (RAG), they rely on lossy embeddings that lose temporal context and fine-grained detail. Few works to date have investigated how VLM-based query-relevant information retrieval can be optimized. We introduce LazySloth, an efficient tree-based search method that speeds up video comprehension and retrieval tasks 2.9-8.3x (compared to existing agentic methods) through bounded captioning of portions of the video considered irrelevant by a VLM of the video. Compared to contemporary specialized video-understanding VLMs and RAG-based methods, LazySloth achieved similar or better final task accuracy across two recent open-source base VLMs--Gemma 4 31B and Qwen3.6 27B--across four benchmarks. LazySloth reduced the gap between the base open-source model and a closed-source model, GPT-4o. Ablations showed that replacing VLM scene understanding with CLIP-based retrieval cost 8.8-19.9% in accuracy, while lazy tree construction matches eager construction at a fraction of the captioning cost. With LazySloth, we demonstrate the possibility of faster long-video comprehension without substantial loss in performance.
cs.AI / 173 / 2609.37582
FedSocket: Recipient-Executable Knowledge Exchange for Heterogeneous Multimodal Federated Learning
Abstract
Federated knowledge must remain usable by recipients with different modalities, private architectures, and tasks. We present FedSocket, which makes recipient execution a design requirement of the exchanged model. A shared Q combines recipient-computable inputs, task-owned outputs, and ownership-aware aggregation, connecting heterogeneous private models through a common prediction interface. Private models teach local Q copies; the returned Q supports local learning and Joint inference, with only Q parameters and counts exchanged. Across six datasets, FedSocket improves missing-modality recipient accuracy over Local by 14.44 and 15.51 percentage points on MELD and UCF-51. Under matched inference capacity, Joint exceeds independent ensembles by 11.06 points in UCF-51 accuracy and 4.87 points in mean bidirectional Flickr30k R@1. Joint also improves over Q alone on all four heterogeneous endpoints, demonstrating the value of combining local and exchanged predictions. Teacher controls, sharing-path interventions, and component factorials identify the roles of supervision, sharing, and deployment. FedSocket makes exchanged knowledge directly usable from federated training to recipient inference.
cs.AI / 174 / 2609.37709
VIF-Bench: Evaluating Visual Instruction Following in Multi-Reference Image Generation
Abstract
Recent multimodal image generation models can take multiple images and textual instructions as input, enabling reference-based generation guided not only by text but also by visual instructions such as layouts, arrows, and pose cues. However, existing benchmarks do not evaluate the joint setting in which multiple references must be composed under multiple and heterogeneous visual-instruction images. To address this gap, we introduce VIF-Bench, a benchmark of 1,241 tasks designed to assess the edge of model capabilities in this joint setting by covering: (i) multi-reference generation (up to 7) under multiple heterogeneous visual instructions (up to 6), (ii) cases where reference images can potentially compete with visual instructions (e.g., a strongly posed subject vs. a target pose), and (iii) controlled comparison of visual instructions with text descriptions at different levels of specificity. Using these capabilities, we uncover three findings: (1) models face an adherence-artifact trade-off: once models reach stronger visual instruction adherence, stronger adherence tends to coincide with more instruction artifacts in generated images, (2) visual instruction adherence tends to be lower on tasks whose reference images carry a salient state of the controlled attribute (e.g., a neon-lit subject under a light-direction instruction), most consistently for light and wind, and (3) for models that can understand visual instructions, it is often better to provide visual constraints directly rather than describe them in text; when using text, a moderate level of detail works better than an exhaustive description. VIF-Bench is released as an open benchmark to establish a basis for fair comparison in controllable multi-reference image generation.
cs.AI / 175 / 2609.37712
PolyOCR-Venus: Unified OCR Foundation Models for Text-Centric Visual Intelligence
Abstract
Optical Character Recognition (OCR) is evolving from plain-text transcription toward general visual intelligence, requiring models to recognize, localize, and reason over textual information in complex visual environments. However, existing OCR systems often excel at only some tasks and struggle to balance recognition, parsing, and reasoning across scenarios. In this report, we present PolyOCR, a family of unified OCR foundation models of varying scales. PolyOCR combines a shared instruction-following framework with a large-scale data engine that converts heterogeneous visual resources into quality-verified OCR supervision. We introduce Competence-Guided Policy Optimization, which combines verifier-based Group Relative Policy Optimization with on-policy distillation through sample-wise routing based on teacher reliability and the teacher--student competence gap. We also introduce OCRBench v2.1, our revision of OCRBench v2 with manually verified annotation corrections and task-aligned scoring metrics. Extensive experiments across OCRBench v2.1, CC-OCR, in-house KIE Benchmark, OmniDocBench v1.6 and MDPBench demonstrate that PolyOCR achieves state-of-the-art or highly competitive performance.
cs.AI / 176 / 2609.37750
Multi-Site Real-World Performance of Commercial AI for Pulmonary and Incidental Pulmonary Embolism Detection
Abstract
Pulmonary embolism (PE) is a leading cause of cardiovascular mortality, yet the real-world performance of FDA-cleared AI detection models remains incompletely characterized. We retrospectively evaluated two FDA-cleared AI algorithms from a single commercial platform (Aidoc Medical BriefCase), one for PE triage on dedicated CT pulmonary angiography (CTPA; n = 30,678) and one for incidental PE (iPE) detection on routine contrast-enhanced CTs (n = 37,191), across a 17-facility academic health system. Reference-standard labels were extracted from radiology reports using a validated LLM pipeline (97% accuracy, kappa = 0.94). The PE model achieved 86.8% sensitivity and 99.1% specificity, with sensitivity declining from 99.3% for saddle emboli to 72.9% for subsegmental PE, and from 89.7% for acute to 65.3% for non-acute PE. The iPE model achieved 73.5% sensitivity and 99.8% specificity. Both models demonstrated lower sensitivity than FDA-clearance benchmarks while exceeding cleared specificity, with diminishing performance for peripheral and non-acute emboli mirroring known human reader limitations and underscoring the need for standardized post-market surveillance of AI-enabled medical devices.
cs.AI / 177 / 2609.37775
HiRAE: Hierarchical Representation Autoencoding with Residual Budgets
Abstract
Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. Meanwhile, intermediate encoder layers contain complementary visual details, but learning to fuse them for reconstruction can produce a latent distribution that is difficult to model. Existing fusion methods require empirical tuning of layer selection or staged optimization of fusion and decoding, increasing configuration effort or training complexity. We introduce HiRAE (Hierarchical Representation Autoencoder), which learns a hierarchical fusion framework over the full encoder hierarchy to improve reconstruction fidelity while maintaining compatibility with generative modeling. HiRAE groups encoder layers by depth and learns residual corrections to the deepest representation. Group-wise norm caps bound these corrections relative to the deep anchor, with tighter budgets for shallower groups. Our HiRAE-24 preserves the latent token count and channel dimension. On ImageNet-256, HiRAE-24 reduces reconstruction FID from 0.299 to 0.209 relative to RAEv2 while maintaining competitive guided generation quality. For text-to-image generation, HiRAE-24 improves alignment over RAEv2 on GenEval, DPG-Bench, and GenAI-Bench both before and after supervised fine-tuning. Under the same generator-training and evaluation protocol, post-fine-tuning GenEval increases from 84.86 to 87.70.
cs.AI / 178 / 2609.37783
A Benchmark & Dataset for Detecting AI-Manipulated Visual Evidence in the Court System
Abstract
Photographic evidence is becoming increasingly vulnerable to forms of alteration and fabrication that existing legal and technical workflows are not well equipped to evaluate. Surveillance frames, dashcam stills, and phone photographs may be used to establish presence, sequence, causation, damage, or identity, yet contemporary generative systems allow non-experts to alter or fabricate such images through ordinary prompt-based interfaces. Existing image-forensics benchmarks provide important resources for face manipulation, classical tampering, and general synthetic-image detection, but they are not organized around the forms of visual evidence submitted in courts, the localized edits that can change what an exhibit appears to prove, or the consumer-tool threat model now facing the justice system. We introduce the CIFAR Synthetic Evidence Corpus for Detecting AI-Manipulated Images, a benchmark for evidentiary image authentication in court and justice-system contexts. The corpus contains 1,505 photographic items, including 720 authentic controls and 785 manipulated or fabricated images, spanning surveillance, dashcam, and consumer-photo imagery. Manipulations are organized into scene-condition edits, localized element edits, and full fabrications produced with contemporary generative systems. Each item is released with structured metadata covering source provenance, manipulation tier, subtype, generator, prompt template, and scene attributes, enabling controlled evaluation beyond aggregate binary detection. We also establish baselines with publicly available image-manipulation detectors, showing that current systems exhibit error profiles that remain problematic for evidentiary use. The dataset, prompts, metadata, code, and baseline evaluation scripts are released to support research on visual evidence authentication, information integrity, and trustworthy AI for the justice system.
cs.AI / 179 / 2609.37809
Pixel-Level Transformers in Remote Sensing: A Canopy Height Case Study
Abstract
Predicting canopy height from medium-resolution satellite imagery is a common and scalable approach for assessing the condition of the world's forests, which play a crucial role in climate change mitigation. While Transformer-based architectures have shown strong performance in many domains, their straightforward application to dense (i.e., pixel-level) regression tasks often yields suboptimal results. In particular, the patch size has a crucial impact on the model performance. In this work, we consider pixel-level attention schemes and show that the resulting models generally outperform those relying on larger patch sizes. However, pixel-level attention can be a prohibitively resource-intensive operation. For this reason, we conduct an extensive experimental study using efficient attention variants to identify favorable trade-offs between prediction quality and resource requirements, facilitating the practical deployment of the proposed models. In addition, we perform a comprehensive comparison with several well-established models in the field and show that, with suitable hyperparameter choices, Transformer-based architectures can outperform competing approaches. Our findings provide practical guidance for designing models for pixel-level regression tasks on medium-resolution satellite imagery, including canopy height and biomass estimation, soil moisture mapping, and yield forecasting.
cs.AI / 180 / 2609.37855
HandAnthro: Automated Hand Anthropometry from a Single Image
Abstract
Hand anthropometry supports protective-glove design, but existing measurement methods often require trained operators, specialized hardware, or manual landmarking. We present HandAnthro, which estimates 44 projected hand dimensions from a smartphone photograph of a palm-up hand on US letter-size paper. The pipeline reconstructs wrist-occluded paper boundaries for rectification, whitens non-hand pixels, and refines 41 anthropometry-specific landmarks from a fine-tuned You Only Look Once (YOLO) pose model using image-specific geometry and contours. Controlled evaluation comprised 720 captures from 45 held-out participants, each contributing 16 images across two smartphones, two backgrounds, two angles, and two nominal illumination settings. HandAnthro produced complete outputs for 704 captures (97.8%); among these, mean absolute error (MAE) was 3.80 mm per dimension against two trained operators' caliper measurements. Regional MAEs were 2.48 mm for non-thumb fingers, 6.04 mm for thumbs, and 6.17 mm for palm and wrist. In a researcher-assisted mobile-app pilot, automated batch processing returned all 44 dimensions for 260 of 268 retained, researcher-screened firefighter images (97.0%). A descriptive, unpaired comparison with an independent national firefighter reference yielded a mean absolute difference of 2.40 mm across 28 sex-by-dimension group-mean contrasts. These results characterize controlled measurement performance and researcher-assisted field feasibility for future distributed hand-anthropometry studies.
cs.AI / 181 / 2609.37938
Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark
Abstract
Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring whether local video understanding actually transfers. We introduce EgoGears, a complementary single- and multi-video benchmark designed to diagnose this transition. It contains 567 single-video and 1,487 multi-video questions derived from 126 human-collected egocentric recordings covering 39 outdoor routes. Repeated traversals across movement speeds and lighting conditions ground comparisons in shared physical environments; 531 questions require alignment across independent recordings. Single-video questions measure the local visual, spatial, and motion evidence available to a model, while multi-video questions test whether evidence remains bound to the correct observation and can be composed into consistent route relationships. We report 29 single-video and 31 multi-video MLLM configurations across six model families in the main leaderboard. Among the 20 configurations evaluated comparably on both splits, every model performs worse on multi-video questions, with a mean decrease of 22.5 percentage points, and the gap persists when answer format and scoring are held fixed. The gap is not explained simply by additional videos or recording boundaries. The central bottlenecks are observation--evidence binding and ordered route-state tracking. The code and benchmark are publicly available at https://github.com/lei-qi-233/EgoGears.
cs.AI / 182 / 2609.38155
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
Abstract
Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the "biography" of the particular entity a question concerns. To address this, we introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment. During question answering, the biography is retrieved alongside episodic evidence, allowing the model to follow an entity through events using identity links established during memory construction. Evaluations across four benchmarks, including day-long and week-long recordings, demonstrate improvements over prior memory frameworks in both multiple-choice and open-ended question answering. On EgoLifeQA, GEB achieves 72.0% accuracy, 4.4 percentage points above the best published result. Ablations show that grounded identity association and biography reading both contribute to the gains, which additional descriptions alone do not fully recover.
cs.AI / 183 / 2609.36054
What if automating AI R&D triggers an intelligence explosion?
Abstract
In contrast to even a year ago, AI systems now write most of the code inside the companies that build them. As more of the AI research and development (R&D) pipeline is automated, could AI progress radically accelerate in an "intelligence explosion," where years of advances are compressed into months or less? Preliminary evidence suggests that it could. In this work, we assess this evidence, analyze an intelligence explosion's potential impacts, and propose policy responses. AI systems are on track to automate most AI R\&D work within a few years, and possibly all of it. If this triggers an intelligence explosion, it could dramatically bring forward AI's benefits, but also pose extreme risks: capabilities growth could accelerate far beyond what society can keep up with, humanity could lose control over superhuman AI systems, and checks on power within and between states, companies, and branches of government could be severely eroded. Although there remains much uncertainty about these possibilities, the high stakes warrant serious further attention. Policymakers should urgently obtain more visibility into the automation of AI R&D, develop ways to steer and constrain an intelligence explosion, and prepare society to adapt to an intelligence explosion's impacts.
cs.AI / 184 / 2609.37085
ARGOS: Reinforcement Learning-Driven Multidimensional Elasticity for Service Orchestration in the Computing Continuum
Abstract
Data-intensive services in the Computing Continuum must balance analytics quality, resource usage, and cost across heterogeneous nodes with limited and uneven capacity. This balance becomes especially difficult when resource scaling reaches capacity limits, because changes in demand and cluster pressure must then be absorbed without violating client-defined quality ranges. Existing orchestrators mainly adapt resources, placements, or replicas, while analytics requirements such as coverage, sample, and freshness remain fixed. This article presents ARGOS, the Adaptive Reinforcement Learning-Driven Governance for Orchestrated Services, an end-to-end controller that formulates multidimensional elasticity as a per-request Markov decision process over analytics quality and cluster pressure, supported by capacity-aware admission. ARGOS is evaluated under controlled workloads and time-varying multi-tenant arrivals on a heterogeneous cluster. Across the controlled scenarios, the deep reinforcement learning policies consistently outperform the non-learning baselines and approach the independently tuned best-fixed reference. A separate live evaluation reports improvements over the static midpoint under realistic and saturated arrivals, with no recorded CPU or memory violations but remaining coverage violations. These results support deep reinforcement learning as an adaptive mechanism for multidimensional elasticity when resource scaling alone is insufficient.
cs.AI / 185 / 2609.37916
RLX: A Unified Multi-Backend Tensor Compiler and Distributed Runtime in Rust
Abstract
Production machine learning (ML) stacks often split graph compilation and kernel execution across different layers and languages, making backend behavior, deployment guarantees, and performance fallbacks hard to reason about end-to-end. RLX addresses this gap with a single Rust codebase that combines compiler and runtime roles around one primitive-level, three-level intermediate representation (IR), plus a transparent dispatch contract that resolves each operator to native, common-IR, or rewritten lowering and fails compilation when legalization is not possible. The same IR targets fourteen runtime devices (cpu, metal, mlx, ane, cuda, rocm, oneapi, tpu, hexagon, gpu, vulkan, opengl, directx, webgpu) and two specialty codegen paths (Cortex-M INT8 and FPGA), ingests safetensors, GGUF, ONNX, and rten formats, supports F16/BF16/F64/C64 and quantized INT4/INT8 flows with AMP/PTQ/QAT, and scales via tensor-/pipeline-parallel collectives over TCP and RDMA transports. Beyond neural workloads, RLX also extends to scientific/physics-style domains through sparse and dense linear algebra extensions (e.g., CSR LU/CG/matvec and LAPACK- backed factorizations) and 3D Gaussian splatting operators. We evaluate RLX against PyTorch, TensorFlow, JAX, candle, burn, tch, rten, MLX, CoreML, IREE, Glow, TensorRT, and tinygrad under identical input generation and p50 measurement methodology on one host. On all-MiniLM-L6-v2, RLX-Metal is fastest at every batch (e.g., 16.6 ms at batch 32 vs. PyTorch-MPS 26.7 ms). In the MNIST training table, RLX also has the top-throughput entry (graph-fused MLP: 946,487 img/s), above NumPy+BLAS (787,349 img/s), while retaining 100% top-1 parity on reference checks (e.g., Qwen3).
cs.AI / 186 / 2609.36593
Text2Sim: Agentic Physics-Based Simulation Generation with Distilled Expertise
Abstract
Creating diverse physical simulations remains labor-intensive because assets, layout, physical parameters, motion, control, and rendering must be designed and debugged jointly. We present Text2Sim, a simulation-specialized agentic pipeline that converts a text-only request into an executable, editable dynamic case. Built on Genesis, Text2Sim uses a hierarchical agentic structure that combines a Planner with specialized Writers, asset-generation tools, and an independent Critic. Compact skills (Debug Cards) distilled from graphics demonstrations provide role-specific physical guidance for execution-based repair. We evaluate physical quality, visual quality, and human preference on 42 held-out prompts spanning rigid, articulated, deformable, and cloth phenomena, with a paper-level split between experience construction and evaluation. We design automatic physical and visual scorers to evaluate the quality of the results, and Text2Sim achieves higher scores than all four state-of-the-art baselines on both metrics. In blinded user studies with these baselines, significantly more participants prefer Text2Sim than prefer the baselines, which is consistent with the results from our automatic scorers. The pipeline also supports a broad range of downstream applications; we select dataset construction and extension to multimodal input as two representative examples. We will release the code, the Debug Card library, and a dataset of generated cases, each pairing the text prompt and rendered video with the executable program, assets, physical parameters, controls, and recorded states.
cs.AI / 187 / 2609.36502
Towards Breaking the Learning System Wall Using Multimodal Tutoring Transcriptions
Abstract
Past research using log data has faced the "learning system wall," whereby few methods exist for generalizing models of student learning across platforms. Increasingly, online learning is captured by richer forms of data, including dialog and video, with new affordances. An example of this is remote tutoring programs, where human tutors support students who use learning systems while video conferencing. Toward better platform-general modeling of learning, we introduce an AI-driven multimodal transcription system that processes screen-recording videos into unified screenplay-style transcripts containing audio dialogue and annotated learning log actions. We describe a planned method for temporally aligning AI-generated multimodal transcripts with MATHia learning logs and for identifying and classifying student learning processes to align with MATHia logs. Lastly, we highlight challenges and potential solutions in capturing learning processes in one system, offering initial steps towards generalizing log data across diverse systems.
cs.AI / 188 / 2609.37109
Designing a Boundary Negotiating Artifact for Collaborative Socio-Technical Sense-Making in AI Regulatory Sandboxes
Abstract
The rapid, unpredictable advancements in AI system capabilities has seen regulators take adaptive and experimental approaches to policymaking. Established in other domains as instruments balancing regulation with innovation, regulatory sandboxes are seen as solutions for AI regulation. However, analyses mostly focus on the legal and institutional design of AI Regulatory Sandboxes (AIRSes). With the legal framework leaving the socio-technical interpretation to stakeholders, this creates a gap on the sense-making required to fulfill the AIRS purpose. In this paper, we approach this by designing a Boundary Negotiating Artifact as a way to mediate meaning in AIRSes. Through Research-through-Design we iteratively develop a tool, providing an interface for the different stakeholders to collaborate in AI assessment. We then position it as technical backbone in established AIRS frameworks, structuring the collaborative sense-making of the involved stakeholders. We further report the insights gained from our design process leaving the qualitative evaluation for future work.
cs.AI / 189 / 2609.37911
Generated Query Expansion Still Helps Strong Sparse Retrieval: A Controlled Study with SPLADE-v3
Abstract
Scientific queries are often brief, while relevant papers use specialized vocabulary. Generated query expansion can bridge this mismatch, but earlier work suggests that its value shrinks as the underlying retriever becomes stronger. We test the four generated formats of term lists, a pseudo-document, multiple pseudo-references, and corpus-steered text all together with SPLADE-v3 on NFCorpus, TREC-COVID, and SciDocs. Every condition searches the same frozen document index and follows the same query-side integration rule and 256-dimension budget, isolating the effect of the added content. All twelve method-collection comparisons improve aggregate nDCG@10, with best relative gains of 4.81%, 8.92%, and 9.47%. Eleven remain significant after Holm correction. The gain persists in 103 of 114 interpolation settings, including every setting that assigns at least 30% of the mixture weight to the original query. Shuffled-text and non-contextual lexical-bag controls also remain above baseline in all 24 aggregate comparisons, showing that the added vocabulary carries most of the benefit. A corpus-induced typed concept graph, by contrast, produces no consistent gain, and its relation, depth, validation, random, and gating controls do not rescue it. Generated vocabulary can therefore complement a strong learned sparse retriever, provided that the original query remains strongly represented.
cs.AI / 190 / 2609.36279
Proofs Without Nominals: Gödel's Ontological Argument, its Shallow Embedding, and the Open Questions of the Monatshefte Notes
Abstract
The shallow embedding of higher-order modal logic in classical higher-order logic, used in Benzmüller and Scott's Notes on Gödel's and Scott's variants of the ontological argument (2025), reaches beyond the modal object language of the arguments: its property quantifiers range over terms that may also express nominals and satisfaction operators of hybrid logic, and a proof using one proves a theorem of the embedding that need not be one of the modal logic. That the framework affords this is not new, and whether a result is one of the modal logic can be settled in two ways: by replaying it in an explicit proof calculus, done by hand for chosen theorems, or by analysing the proofs the embedding itself produces, which this article does mechanically, for every result at once. Every statement the Notes prove has a proof inside the object language: 294 written out by hand and machine-checked, none using a nominal. The proofs the Notes themselves give instantiate no nominal either; what the detector flags there are terms a prover substituted. The three questions the Notes leave open are settled too, and without nominals, but the conjunction axiom has to be emended: generalised in the Notes to Gödel's "any number of summands", it covers the conjunction of no properties, and of one; the empty one alone settles all three, and the two together yield what a separate axiom of Gödel's is for. This article restricts the conjunction axiom to at least two different conjuncts, the reading Gödel's footnote suggests, and the questions are settled again, by proofs that turn on the argument rather than a degenerate instance. The restriction holds of the object language only: with a nominal the axioms make the accessibility relation the identity and the readings coincide. Every theorem is verified in Isabelle/HOL and independently in Lean 4; the countermodels are Nitpick's, certified by the build.
cs.AI / 191 / 2609.36845
DSWM: Decomposed Spatio-Temporal World Model for Demand-Driven UAV Base Station Repositioning
Abstract
Uncrewed aerial vehicle base stations (UAV-BSs) are expected to cover traffic demand that shifts across space and time, yet most repositioning schemes either re-solve an optimization problem per slot or learn reactive policies without an explicit demand model. We cast demand-driven fleet repositioning as latent-space decision-time planning and propose DSWM, a decomposed spatio-temporal world model: an agentic controller that perceives the demand field through a rolling observation window, retains operational context in a latent recurrent state, reasons about candidate motions by imagined rollouts under an uncertainty penalty, and coordinates the fleet through replanned first actions. DSWM learns a recurrent state-space model shaped by an exponential-moving-average (EMA) based latent predictive objective with variance regularization. It attaches a differentiable service simulator that replays the association, probabilistic line-of-sight channel, and Shannon rate chain inside latent rollouts. Planning uses a cross-entropy method whose imagined demand is anchored on the current observation window with mixing coefficient $ρ=0.95$. On a unified pipeline over three real datasets (Milan CDR (call detail record), Shanghai Telecom, YJMob100K) and 14 methods including five reproduced IEEE baselines, DSWM attains weekday served ratios of 0.889, 0.908, and 0.898, ranking first among non-ablated configurations on every dataset. On Milan it improves over the strongest non-learning baseline (Greedy, 0.780) by 0.109, a margin that comes from decision-time use of observations rather than prediction accuracy.
cs.AI / 192 / 2609.36333
ATLAS: Aligned Transport of Latent Structure for Reliable World Model Planning
Abstract
Latent world models rely on representation geometry for planning, yet regularizing the latent marginal alone does not determine the state-to-state relationships used for action selection. We show that this can cause planning-relevant novelty structure to be weakened as representations are transformed into the final latent used by the planner. We introduce Aligned Transport of Latent Structure (ATLAS), a training objective that explicitly preserves relational geometry while calibrating the global latent distribution. ATLAS transfers normalized pairwise structure from an informative encoder representation to the planning latent and uses Wasserstein embedding matching (WEMReg) to calibrate its marginal through one-dimensional Wasserstein-2 transport. Our analysis shows that relational preservation and marginal calibration impose non-redundant constraints, and connects finite-candidate planning stability to relational distortion, latent-scale mismatch, and prediction error. Instantiated in LeWM, ATLAS improves mean goal-reaching success across PushT, TwoRoom, and OGBench-Cube on both lower- and higher-novelty evaluation subsets, with the largest gain on higher-novelty TwoRoom episodes. Representation and rollout diagnostics further show stronger novelty-related structure in the planning latent, improved marginal calibration, and lower multi-step prediction error. Together, these results highlight preservation of planning-relevant latent geometry as an important ingredient for reliable world-model planning. Code is available at https://anonymous.4open.science/r/atlas-world-model-72C4/.
cs.AI / 193 / 2609.36416
FineART: Fine-grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation
Abstract
Robots operating in real-world environments must execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions. Current manipulation datasets struggle to support this capability: although single-arm datasets reach hundreds of thousands of trajectories, they typically provide only one high-level instruction per episode while the rare bimanual effort that does label subtasks annotates only a fraction of its hours. We present FineART, a densely annotated bimanual manipulation dataset of 40,543 episodes, 1,718 hours, and 533,913 subtasks across 151 tasks. We also introduce FineART-VLA, a vision-language-action policy that predicts its own next subtask, and show that mid-training it this way yields substantial gains. Specifically, success on a spatial disambiguation task increases from 32.0% to 100.0%, and step-by-step human subtask guidance lifts success on an unseen long-horizon task from 16.0% to 76.0%. Furthermore, after minimal fine-tuning on a new robot, the policy requires one-tenth the data of baselines without mid-training and generalizes zero-shot to completely unseen tasks on the new hardware. We open-source the full dataset, model weights, and training code.
cs.AI / 194 / 2609.36471
Staircase Policy: Streaming Inference for World-Action Models with Large Action Chunks
Abstract
World-Action Models (WAMs) improve robotic manipulation by conditioning action generation on predicted future observations, but future prediction adds further inference overhead to already expensive iterative action generation. Action chunking can amortize this cost over multiple actions, yet performance degrades over long execution horizons because later actions remain conditioned on stale observations. We introduce STAIRCASE POLICY, a streaming inference and training framework that turns a flow-matching VLA into a JEPA-style WAM and partitions a large action chunk into sub-chunks at staggered denoising stages. Near-term actions are executed as soon as they become available, while later actions continue to be refined. At each sub-chunk boundary, the future latent is re-predicted from the latest observation and used to update all unexecuted actions, enabling long-horizon execution without repeated full policy inference. The resulting future-prediction error can further serve as a signal for adaptive chunking. S-WAM achieves 97.7% on LIBERO and 87.9% on LIBERO-Plus, and improves performance across multiple policy backbones and real-robot tasks. It reaches 292.7 executed actions per second, $3.62\times$ the throughput of conventional execution at comparable accuracy, while reducing time-to-first-action from 123.6 to 73.3 ms. With additional inference optimizations, throughput further increases to 642.9 actions per second.
cs.AI / 195 / 2609.36518
LIBERO-MAX: Do Robot Policies Adapt When the World Changes?
Abstract
Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at reset, leaving this temporal challenge underexamined. We introduce LIBERO-MAX, a benchmark of 8,000 paired cases spanning eight types of changes to geometry, observations, appearance, clutter, and paths. Each pair compares task execution with and without a mid-task event, holding the task, initial state, policy seed, and pre-event action sequence fixed. This controlled comparison distinguishes event-associated regressions from failures already present without the change. Across fourteen current VLA, hybrid, and world-action policies, events reduce success by 11.0-25.7 percentage points. Event profiles reveal shared vulnerabilities to geometry and observation changes, while policy-family rankings interleave. Camera controls show that robustness reflects both competence under the changed conditions and the trajectory from which they are encountered; varying query cadence does not eliminate the gap. Together, the paired protocol and temporal diagnostics establish LIBERO-MAX as a reproducible testbed for diagnosing failures under mid-execution changes and measuring progress toward robot policies that remain effective as the world changes.
cs.AI / 196 / 2609.36588
Cooperative Multi-Agent Vision-Language-Action Models via Reinforced Fine Tuning
Abstract
We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-grained coordination skills required for inter-robot collaboration. Supervised fine-tuning (SFT) on multi-robot demonstrations partially bridges this gap, but its performance is bounded by the demonstration data and cannot improve from its own experience. We present a three-stage reinforced fine-tuning (RFT) pipeline for multi-agent VLAs. First, initialization-aware data collection sweeps over initial configurations and invokes human demonstrations only when the pretrained VLA repeatedly fails, yielding robustness to initialization shift with reduced human cost. Second, offline credit-filtered tuning assigns credit to individual agents and fine-tunes on per-agent trajectories with positive advantage rather than on entire joint rollouts. Third, we find existing online RL for VLAs are less effective for hard multi-agent tasks, which we attribute to noisy co-exploration and unstable updates. We instead use online latent-space fine tuning, which freeze the VLA and perform RL in its latent noise space. We evaluate our multi-agent VLA with both $π_0$ and $π_{0.5}$ backbones across 11 tasks in RoboTwin, RoboFactory and real-world manipulation with two Franka robots. Our multi-agent VLA improves the average success rate by $+23.1\%$, $+16.4\%$, and $+44\%$ on RoboTwin, RoboFactory, and real-world tasks, respectively. Code available at https://anonymous.4open.science/r/mavla_rft-2BC0/.
cs.AI / 197 / 2609.36595
Simple Agentic Memory for Generalist Robot Policies
Abstract
Visual-memory systems commonly retain or compress past observations. Robot control additionally requires interaction-derived state that no individual frame may explicitly represent, such as persistent identity relations, accumulated progress, or ordered procedures. We introduce Simple Agentic Robot Memory (SimpleARM), a training-free memory layer for frozen generalist robot policies. From the task instruction, SimpleARM specifies what to monitor; frozen perceptual tools maintain compact typed state online; structured access retrieves that state only when a proposed subgoal depends on history; and current-view grounding resolves recalled entities before execution. We evaluate SimpleARM on RoboMME, a benchmark of memory-dependent robot manipulation tasks that require history information no longer available in the current observation. Across all 16 tasks and three policy seeds, SimpleARM achieves 67.17% mean success, compared with 44.51% for the strongest non-oracle baseline. Matched ablations show mechanism specificity: removing relation, reference, progress, or route state produces large losses where the affected state is retrieved for control, while largely sparing other tasks. These results support a state-based view of robot memory: effective memory for control is not simply retained visual history, but compact task-relevant state derived from the interaction history.
cs.AI / 198 / 2609.36645
Where Predictive Supervision Goes Shapes What VLA Policies Learn
Abstract
Future prediction is increasingly used to improve vision-language-action (VLA) policies, based on the premise that anticipating scene evolution encourages representations useful for control. However, forecast quality alone does not establish that a policy has learned a better representation for action. This distinction matters under distribution shift, where successful control depends on preserving spatial state and likely scene change beyond familiar configurations. We study what determines whether predictive supervision improves the visual representation used by a VLA policy. Through controlled comparisons with matched target constructions, prediction horizons, and training conditions, we find that different prediction interfaces produce markedly different forecasts and visual representations, including in the spatial, dynamics, and action information that transfers beyond familiar scenes. We trace these differences to how predictive errors shape the policy's visual stream. Consistent with this controlled finding, VLA policies trained with more direct, scene-matched future supervision show stronger robustness under simulated and physical distribution shifts. Together, our results frame future prediction as a representation-learning design problem whose value for control depends on whether its supervision reaches the representations through which the policy acts.
cs.AI / 199 / 2609.36808
Spotter: Let the Embodied Model Lead, and the VLM Reflect for It
Abstract
Current embodied models do not respond to their own failures, although what just went wrong could inform a small adjustment on the next attempt, the kind of reflection behind the gains of thinking in language models. We test whether they can repair a known error, which requires producing a correction and judging whether it is right. Stopped at a failure and allowed to retry, they seldom repair it through their own randomness or from a language description of the error, and best-of-N selection cannot pick the successful candidate after a failure. We attribute this to training only on successful demonstrations and to inputs too narrow to show what went wrong, and conclude that reflection must come from a vision-language model (VLM), which takes in far more information, such as the episode history and text, and is more general. Prior VLM-led work has the VLM plan every step and invoke the embodied model as a tool, placing the VLM on the critical path. We propose Spotter, which reverses the roles: the embodied model leads and executes continuously, while the VLM runs in parallel, monitors through a lightweight local screener, intervenes only when an error is detected, reflects on and corrects it, and returns control. We run Spotter with Qwen and with GPT as the VLM, and both improve the embodied models; with GPT, Spotter improves Cosmos Policy and $π_{0.5}$ by 5.6 and 7.5 percentage points on RoboCasa, and raises $π_{0.5}$ from 47.2% to 57.0% on the Hard setting of RoboTwin 2.0 and from 53% to 83% on a real robot. Because the VLM steps in only when an error is confirmed, a successful episode with Qwen takes only 13 to 16 s longer than with the embodied model alone and about 70% less time than with a VLM-led baseline using the same model. Our code is available at https://github.com/zqc3117/Spotter.
cs.AI / 200 / 2609.36915
AeroManip-VLA: Scalable Vision-Language-Action Learning for Aerial Manipulation with RL-Generated Demonstrations
Abstract
Aerial manipulators extend robotic manipulation into 3D workspaces that are difficult for ground-based robots to access, creating new opportunities for general-purpose manipulation. However, extending Vision-Language-Action (VLA) models to aerial robots introduces distinct challenges due to the tight coupling between manipulation and flight, continuously changing observations, and safety-critical physical interactions. These challenges demand diverse training data and systematic policy evaluation, yet collecting demonstrations and evaluating policies directly on physical aerial platforms are costly, difficult to scale, and hard to repeat under controlled conditions. We present AeroManip-VLA, a scalable benchmark for aerial VLA data generation and policy evaluation. AeroManip-VLA provides a GPU-accelerated simulation framework with low-level payload-aware flight and manipulation control in massively parallel environments. Building on this framework, we combine reusable reinforcement learning policies with expert task rules to automatically generate demonstrations without human teleoperation across diverse objects, environments, and randomized initial conditions. The generated data include basic skills such as grasping and placing, as well as long-horizon tasks that require both navigation and manipulation. We further introduce automated event labeling and trajectory categorization to filter demonstrations. These mechanisms enable fine-grained analysis of task progress, behavioral outcomes, and safety-related failures. Finally, we evaluate a range of imitation learning and VLA baselines across different task settings, revealing their performance characteristics and failure modes. Together, AeroManip-VLA enables scalable aerial manipulation data generation, structured trajectory analysis, and systematic VLA evaluation in simulation prior to real-world deployment.
cs.AI / 201 / 2609.37067
FACT: Fidelity-Aware Construction of Articulated Twins
Abstract
Visually plausible articulated assets may still fail during contact interactions or exhibit inaccurate motion. We present FACT (Fidelity-Aware Construction of Articulated Twins), an agentic framework that progressively constructs articulated twins to improve geometry, contact, and dynamic fidelity. The agent drives an evidence--diagnosis--revision loop on a shared editable representation, selecting measurements and model edits using quantitative feedback, while numerical tools execute and validate the updates. It reconstructs editable articulated geometry from images through feature planning, targeted measurements, and diagnostic refinement. On this reference, it repairs collision proxies through task-aware local repartitioning before fidelity-constrained compression. Finally, it constructs response models from passive-response videos, using simulation residuals to guide model revision and constrained physical parameter fitting. Experiments show that FACT improves geometric reconstruction over baselines, enables more reliable interaction with simpler collision proxies, and better reproduces held-out physical responses than direct parameter inference.
cs.AI / 202 / 2609.37070
Predictive Safety Curricula for Robust Legged Locomotion
Abstract
Rare but consequential failures can persist in learned locomotion policies for legged robots even when average task performance is high, in part because standard curricula primarily adapt task difficulty rather than the distribution of safety-critical experience. We introduce Predictive Safety Curricula (PSC), a framework for allocating locomotion training experience using learned predictions of future safety cost. PSC trains a distributional safety critic from policy rollouts and uses its predictions to prioritize both terrain contexts and previously encountered randomized events. The resulting curriculum modifies the training distribution while leaving the task reward and policy-optimization loss unchanged. We evaluate PSC in controlled rough-terrain locomotion and in production locomotion systems. PSC improves reliability relative to standard terrain progression, advantage-based replay, and learning-progress curricula, with the largest gains on difficult terrain and under degraded observations. The same allocation principle transfers to two production locomotion stacks. On ANYmal-D hardware, PSC reduces shank-collision incidence by $63\%$ relative to the learning-progress curriculum across three matched training seeds, with a reduction in every seed. On a production stair-climbing platform, PSC eliminates observed shank collisions in the evaluated hardware trials. These results show that learned predictions of future safety cost can provide an effective signal for allocating training experience toward rare failure modes and improving locomotion reliability.
cs.AI / 203 / 2609.37181
EgoHumanoid-V2: Human-to-Humanoid Transfer of Coordinated Whole-Body Skills for Loco-Manipulation
Abstract
Human demonstrations capture diverse scenes and rich whole-body skills without requiring robot teleoperation. Prior work on egocentric transfer has emphasized scene generalization in loco-manipulation under decoupled control, leaving direct transfer of coordinated whole-body skills less explored. We present EgoHumanoid-V2, the first egocentric human-to-humanoid skill transfer framework for coordinated whole-body loco-manipulation. At its core, coarse-to-fine action alignment combines kinematic reference correction with dynamics-aware refinement. It improves end-effector pose accuracy while preserving whole-body coordination. We also use robot-arm rendering and training-time image augmentation to reduce the visual embodiment gap and improve viewpoint robustness. On four real-world tasks, vision-language-action (VLA) policies trained on aligned human data show zero-shot skill transfer without target-task robot demonstrations. Task scores are comparable to those of policies trained on teleoperation data at a lower collection cost. These results support human data as direct skill supervision.
cs.AI / 204 / 2609.37359
Encore: Few-Shot Agentic Discovery of Manipulation Strategies
Abstract
Coding agents can now write, run, and debug programs with little human help. Robot tasks, however, are usually specified by a sentence that leaves out how to grasp, in what order to make contact, and what the result should look like, and an agent given only the sentence must find these details by trial and error. We introduce ENCORE, which gives the agent a few demonstrations as evidence to read rather than as training data. A deterministic builder distills each demonstration into a pack of multi-view keyframes, gripper events, frame strips, and the full trajectory. A coding agent studies the pack, writes a policy program against a fixed perception and action API, refines it iteratively over a few development rollouts, and freezes it before a sealed evaluation that never reveals the success signal. On LIBERO-PRO, the agent's first program already succeeds in half of the perturbed tasks with demonstrations and in one task without them, and the frozen programs outperform the strongest prior agentic system run with the same language model (96.3% against 89.3%). On RoboDojo tasks whose instructions leave the goal unstated, no program succeeds without demonstrations. ENCORE also runs on a real bimanual robot, learning cube handover and cup inversion from five demonstrations each.
cs.AI / 205 / 2609.37519
Video2STL: Grounding VLM-Generated Temporal Specifications for Robot Learning
Abstract
Video-based policy learning is particularly promising, as it illustrates target behaviors without requiring action annotations or embodiment-matched demonstrations. A central challenge is deciding what information should be transferred from the video to the robot. Existing approaches commonly convert visual observations into scalar similarity or value signals, or ask foundation models to directly generate reward code. These approaches can make the temporal structure of a task difficult to inspect, ground, and reuse. We present Video2STL, a framework that converts observation-only videos into parametric Signal Temporal Logic (STL) specifications and uses the resulting formal representation for robot learning. A vision-language model extracts an embodiment-independent semantic event trace and constructs a bank of symbolic temporal specifications. The model determines the task structure, while numerical predicate thresholds and temporal bounds are grounded from successful robot trajectories. For policy learning, we separate short- and long-timescale temporal information: short-horizon specifications provide dense rewards through rolling-window quantitative robustness, while a causal monitor over a retained long-horizon specification provides one-time progress rewards for valid temporal prefixes. The same representation supports cross-embodiment transfer from human or animal videos to robot control. Across four manipulation tasks, Video2STL achieves $85.8\%$ average success-once and $67.0\%$ success-at-end, compared with $81.5\%/59.5\%$ for native dense PPO and $65.0\%/42.3\%$ for Text2Reward; in quadruped locomotion, Qwen-3.8 and GPT-5.6-based Video2STL policies achieve $100\%$ success across velocities from $0.3$ to $2.1\,\mathrm{m/s}$ while remaining competitive in high-speed energy efficiency. Project webpage: \href{https://video2stl.github.io/}{video2stl}.
cs.AI / 206 / 2609.37591
Credit-Guided Policy Improvement for Test-time Adaptive Vision-Language Navigation
Abstract
Test-time adaptation for vision-language navigation (TTA-VLN) enables pretrained policies to adapt online to unseen environments using only test-time observations and interaction history. However, distribution shifts can distort local action preferences and lead to off-course decisions. Existing methods rely on predictive uncertainty, trajectory-level feedback, or accumulated adaptation experience to correct such deviations. These signals, however, do not directly reveal whether an executed action supports instruction-guided progress toward the goal. Moreover, a plausible corrective signal does not guarantee a reliable policy update. The key challenge is thus twofold: identifying interactions that support goal-directed improvement and determining whether the resulting updates are worth retaining. We observe that each executed action induces an immediate observation transition, providing evidence of its local consequences. Based on this insight, we propose Credit-Guided Policy Improvement (CGPI), which recovers signed, reference-relative decision credit from action-induced observation transitions without external outcome feedback. With the pretrained navigation policy frozen, CGPI uses this credit to propose lightweight adaptation updates and verifies them against prior credit-supported interactions. Updates are retained only when supported and rolled back otherwise. CGPI achieves consistent gains across the evaluated VLN benchmarks and navigation backbones, while qualitative robot trials further illustrate the feasibility of zero-shot sim-to-real transfer.
cs.AI / 207 / 2609.37666
Semantic Map Sharing and Capability-Aware Coverage Planning for AI-Native 6G Robotic Coordination
Abstract
Search and Rescue (SAR) operations increasingly deploy heterogeneous teams of aerial and ground robots. However, conventional coverage methods typically do not translate perceived terrain into platform-specific reachability, while continuous image exchange imposes a high communication cost. We propose an edge-centric, semantic-aware coverage planning framework that integrates aerial terrain perception, robot-specific traversability reasoning, and payload-efficient semantic state sharing. Aerial observations are converted into compact semantic grid maps, enabling reachability-constrained area decomposition and capability-aware coverage paths that assign only regions admitted by each robot's capability profile. The resulting perception-sharing-planning loop feeds semantic corrections into traversability reasoning and replanning, forming an application-level mechanism motivated by AI-enabled goal-oriented communication envisioned for AI-native 6G networks. For the high-update case, transmitting semantic corrections reduces the application payload by a factor of approximately $82$ relative to periodic full-map sharing. Across matched benchmark scenarios, the proposed method achieved $91.5\%$ coverage with no capability-infeasible allocations, compared with $78.8\%$ coverage and a $21.5\%$ capability-infeasible allocation rate for LS-MCPP. Semantic corrections update the shared planning state without requiring repeated transmission of the complete map.
cs.AI / 208 / 2609.37810
Explore, Execute, Evolve: A Skill Acquisition and Reuse Loop for Embodied Agents
Abstract
Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potential for zero-shot robotic task solving. However, they often incur high execution costs by reasoning and exploring the physical world from scratch. To reduce these costs, we introduce RoboSkill, a framework that connects skill acquisition and reuse through an Explore, Execute, Evolve loop. Within this loop, the agent explores to gather task-relevant information, executes tasks while adapting to feedback, and evolves its skill library based on execution records. It then reuses these skills to guide exploration and execution in the next cycle, closing the loop. To improve loop efficiency, we complement vision with tactile feedback to reduce uncertainty during physical interaction. We further augment textual guidance with reusable code to reduce reasoning overhead during skill reuse. On LIBERO-10, RoboSkill improves first-episode success rates by 12.5--25.0 percentage points and reduces average runtime by 7.6--72.4% across four agents. On real robots, it improves success rates by 8.3 percentage points and reduces average runtime for successful trials by at least 14.4%.
cs.AI / 209 / 2609.37871
ExceptionDrive: A Planning-Oriented Counterfactual Corner-Case Benchmark for Autonomous Driving
Abstract
Average performance on routine driving benchmarks does not establish planner reliability under rare, safety-critical hazards. We proposed ExceptionDrive, a counterfactual planning benchmark that uses VLM-assisted screening, localized multi-view editing, and quality auditing to insert hazards into real nuScenes scenes while preserving their context. Its 21 tasks span six safety families and define hazard or conflict regions, local safety constraints, and acceptable responses. Because hazard insertion can invalidate the recorded human trajectory, our reference-free protocol evaluates edited predictions using Unsafe Rate (UR), Hazard Clearance Compliance (HCC), Hazard Proximity Response (HPR), and Counterfactual Trajectory Shift (CTS), which measure core-region intrusion, clearance compliance, clearance relative to a prescribed margin, and counterfactual trajectory change. Seven representative planners frequently intrude into hazard regions or provide insufficient clearance. We also develop a Reminder Agent that, without sample-specific task labels, converts visual evidence and the shared taxonomy into structured records of hazard presence, type, and a recommended high-level strategy. The agent neither predicts trajectories nor controls the vehicle; its records guide a VLM-based decision agent. In zero-shot experiments, the reminders improve strategy accuracy and reduce under-warning.
cs.AI / 210 / 2609.38028
doPlan: A Variable-Horizon Dataset for Multi-Stage Language-Conditioned Planning in Autonomous Driving
Abstract
Autonomous vehicles interacting with passengers through natural language must reason beyond immediate commands. Passenger intent may span multiple stages of behavior, depend on future events, refer to surrounding agents or landmarks, and remain relevant as driving conditions evolve. Existing language-enabled driving datasets largely focus on short, localized interactions, leaving these longer-horizon forms of passenger intent comparatively underexplored. We introduce doPlan, to our knowledge the first publicly available, human-annotated real-world dataset designed to study passenger language as persistent task context. Built on nuPlan, doPlan contains 5,154 human-written passenger instructions spanning 169.1 hours of cumulative instruction-aligned context over 50.9 hours of unique driving, with annotation windows ranging from 30.0 to 508.8 s. The annotations capture immediate, deferred, event-conditioned, persistent, and multi-stage passenger intent. The dataset, annotation interface, and supporting resources are publicly available at https://github.com/Mi3-Lab/doPlan. We evaluate four language-conditioned driving models and find that sensitivity to passenger language does not reliably translate into behavior consistent with the requested direction. More broadly, among 2,161 examples with a matched future maneuver, the first associated maneuver occurs a median of 24.6 s after the evaluation point, and only 9.8% occur within the models' common 5 s prediction horizon. These findings highlight the need to connect persistent passenger intent with successive planning decisions. doPlan provides a setting for studying how unresolved goals can be retained, grounded in evolving scenes, and tracked across multiple stages, including how a planner determines when a future goal becomes relevant to the current plan.
cs.AI / 211 / 2609.38178
Skill-Space Shooting for Autonomous Robot Policy Improvement
Abstract
Robots deployed in the physical world must be able to improve beyond their initial training as they encounter new situations and failures. For this improvement to scale across tasks, it must make effective use of experience without requiring human demonstration of each correction. Recent agentic systems offer a way to reduce this reliance on human effort by using foundation models to autonomously compose learned behaviors to complete tasks. Yet completing tasks this way does not itself teach a task policy to overcome its own failures; that requires turning these behaviors into learnable corrections for the policy. Our insight is that many such corrections are familiar short behaviors, or skills: they recur across tasks and describe actions that foundation models can reason about from a scene. We introduce skill-space shooting, which uses foundation model guidance to explore corrections through these reusable skills and turn successful trials into policy improvement. Real-world experiments show repeated improvement in policies acting autonomously, while skills can also be shared to reduce the teaching needed to improve on new tasks. By making reusable skills a source of corrective supervision, skill-space shooting enables scalable and generalizable policy improvement within and across tasks. Additional results and videos at https://skill-space-shooting.github.io.
cs.AI / 212 / 2609.36460
Emergent Tonal Structure in Learned Chord Embeddings and Its Relation to Tonal Tension
Abstract
Several tonal pitch spaces and computational models have been proposed to analyze tonal structure in Western tonal music, many of them grounded in principles from music theory and used to support tonal analysis with important implications for tonal tension. In parallel, data-driven methods such as skip-gram have been used to learn chord embeddings from symbolic corpora, but their ability to recover tonal structure and its relation to tonal tension remains underexplored. In this work, we investigate how skip-gram chord embeddings reflect tonal structure and whether they provide a useful basis for analyzing structural aspects of tonal tension. Using chord sequences with and without transposition-based augmentation, we evaluate the learned spaces from geometric, functional, and tension-related perspectives. We show that augmented embeddings exhibit strong transposition equivariance, recover a clear circle-of-fifths structure, and support interpretable shifts between key-related regions of the learned space. We then derive embedding-based measures from chord-to-key distance and contextual chord-distance relations, and show that they capture meaningful aspects of tonal tension structure through correspondence with matched tonal measures and moderate alignment with human tension profiles. Across analyses, transposition-based augmentation generally improves the stability, tonal coherence, and interpretability of the learned space.
cs.AI / 213 / 2609.36500
InterBias-SV: Compound Conditions in Speaker Verification
Abstract
Speaker verification systems encounter combinations of noise, channel distortion, and changes in speech. Evaluating each condition separately does not establish whether their effects add. InterBias-SV organises this question around a four-term comparison: joint error, two marginal errors, and a common reference. Its results artefact contains 4,068 scored records across 17 experiments, 12 encoder labels, and six speech corpora, totalling 12 million trial evaluations. Three experiment families contain the same-corpus terms needed to compute additive contrasts. For labels assigned to speaker-trained encoders, their mean contrasts are +0.0026, +0.0088, and +0.0024 in equal error rate (EER), with larger variation across settings. These descriptive averages do not establish equivalence to additivity: trial matching, checkpoint identity, and parts of the condition metadata remain unverified. We also examine two interpretation problems. Near-chance EER can make additive predictions difficult to interpret, but chance performance is not a hard EER ceiling, and correlation with the prediction does not identify a saturation mechanism. Ratios of demographic gaps are unstable when their clean reference is near zero; absolute gaps provide a more direct summary. The benchmark provides condition definitions, analysis scripts, and explicit requirements for interpretable compound-condition comparisons, while separating recomputable summaries from claims that require further experimental validation.
cs.AI / 214 / 2609.37116
Multichannel Audio Quality Assessment: Extending Pretrained Perceptual Models to Spatial Audio
Abstract
Accurate perceptual quality assessment is essential for evaluating and optimizing spatial audio, where perceived quality depends on both signal fidelity and inter-channel spatial relationships. However, subjective evaluation is costly, while existing perceptual models are often trained for limited channel configurations and cannot be directly applied to higher-channel-count audio. This raises the question: how can pretrained perceptual knowledge be effectively reused for multichannel spatial audio? Using 5.1-channel audio, we study four levels of multichannel integration: signal, prediction, latent, and feature and propose two learned approaches: latent-level aggregation of spatial-group representations and the feature-level Feature-Band Group Attention (FGAtt), which adaptively fuses spatial groups at the feature level before perceptual processing. Across five 5.1-channel test sets, FGAtt achieves the strongest over- all performance, demonstrating the effectiveness of feature-level adaptation for reusing pretrained perceptual knowledge
cs.AI / 215 / 2609.37586
Learning as Deepfakes Evolve: RF-Prompt for Continual Audio Deepfake Detection
Abstract
Continual audio deepfake detection requires learning newly emerging deepfake methods while retaining discrimination of previously encountered speech. Existing dataset-incremental evaluation changes both real-speech domains and deepfake mechanisms, making their effects difficult to distinguish. We construct five task organizations over identical training, development, and evaluation pools to study these factors under a controlled sample budget. Our proposed Real-Anchored Mechanism-Incremental (RAMI) protocol reflects the practical setting in which available real speech provides a recurring mixed-domain reference while new deepfake mechanisms arrive incrementally. We further propose RF-Prompt, an asymmetric continual prompt-learning method that preserves reusable real-speech knowledge through a shared real prompt and expands mechanism-specific knowledge through inherited fake experts with orthogonal residuals. Input-adaptive soft fusion combines the accumulated experts into a fixed number of injected tokens without requiring task identity at inference. On RAMI, RF-Prompt achieves 10.110% average EER and 10.370% pooled EER, outperforming all evaluated continual-learning baselines. Across the five controlled protocols, RAMI yields the lowest common-average and pooled EER. Component ablations, limited-data experiments, and cross-backbone evaluations further validate the proposed design.
cs.AI / 216 / 2609.37617
AS$^2$D: Accelerating On-Demand Audio Understanding on Mobile Devices
Abstract
Speculative decoding accelerates autoregressive generation by using a smaller drafter to propose tokens for batched verification by a larger target. However, conventional speculative decoding couples drafting to the target's evolving verified prefix, serializing drafting and verification. We ask whether this dependency is necessary for source-conditioned generation. Our key observation is that, for audio language models, the input audio and user request can provide useful speculative candidates without following the target's evolving text prefix. We propose AS$^2$D (Audio Speculative Speculative Decoding), which enables target-decoupled drafting: an audio-conditioned drafter follows its own generation history while the target independently verifies and corrects ready candidates. Without usable candidates, the target advances alone. Thus, target feedback determines which candidates are committed but no longer determines when the drafter can make progress, enabling drafting and verification to proceed concurrently while retaining target-side verification and correction. We implement AS$^2$D in MNN for Android and evaluate two target models across four phones, seven datasets, and three tasks covering 12.2 hours of audio. Across four phones, AS$^2$D improves pooled ASR throughput by 42-76% over target-only decoding, while only 5.7% of evaluation windows are slower than target-only, compared with 58.1-63.0% for speculative baselines. For ASR, AS$^2$D reaches 97.33-98.20% of a hindsight per-window oracle's pooled throughput over the evaluated drafter/budget catalog. Native on-demand execution with a 7B target achieves up to 78% higher throughput than target-only. These results show that source-conditioned audio generation can relax the conventional dependence of speculative drafting on the target's evolving output prefix, exposing substantial parallelism for efficient inference.
cs.AI / 217 / 2609.37798
GLaS-JEPA: Gaussian-Regularized Speech SSL without Engineered Prediction Targets
Abstract
Speech self-supervised learning aims to learn general-purpose representations for downstream speech tasks. However, current approaches rely on complex, carefully designed prediction targets. We challenge this necessity with GLaS-JEPA, a framework that directly predicts the current encoder's continuous representations at masked positions, without contrastive learning, discrete targets, or separate EMA target encoders. We prevent representation collapse using SIGReg representation-space regularization, eliminating the need for engineered target-generation mechanisms. Pretrained on 960 hours of LibriSpeech, our 57M-parameter model achieves a 6.89% WER on frozen-encoder SUPERB ASR and a 25.87% CER on slot filling, outperforming the best non-distilled sub-90M baselines by 43.1% and 22.0%, respectively. These results demonstrate that highly competitive speech representations can emerge from a radically simplified training recipe.
cs.AI / 218 / 2609.36525
Reliability Testing of Medical Model Performance under Distributed Deployment
Abstract
Distributed inference has become an indispensable part of deploying medical models under practical latency, memory, and throughput constraints. Although modern frameworks improve serving efficiency through tensor parallelism, mixed precision, kernel fusion, and multi-device communication, they are generally assumed to preserve the behavior observed during centralized HuggingFace evaluation. This assumption creates an evaluation-deployment mismatch: a model may pass offline evaluation but produce a different output after the execution stack changes. To address this mismatch, we propose a testing framework and an improved, distributed-execution-sensitive medical-model benchmark that evaluates the same checkpoint and input under a centralized HuggingFace reference and matched distributed deployments. Extensive experiments across language, vision, and multimodal medical models show that execution changes can produce measurable output disagreements. Across supported visual settings, the test success rate ranges from 0.21 to 0.43 for single-modality models and from 0.32 to 0.98 for multimodal models. The benchmark is aimed at extending medical-model evaluation from capability and security to evaluation-deployment consistency.
cs.AI / 219 / 2609.36687
MyoCodec: A Streaming Neural Codec for Electromyography
Abstract
Neural codecs encode continuous signals into compact sequences of discrete tokens, providing an interface for efficient transmission, storage, and token-based sequence modeling. This paradigm has been widely adopted in modern speech and audio frameworks; however, the biosignal domain still lacks a neural codec designed specifically for low-bitrate streaming and generalization across diverse downstream tasks. We present MyoCodec, a streaming neural codec designed for electromyography (EMG). Inspired by recent neural audio codecs, MyoCodec combines causal Transformers with residual vector quantization to encode continuous EMG signals into different levels of EMG representations spanning from continuous latent features to discrete tokens operating at 50 Hz. Trained on twelve public EMG datasets, MyoCodec achieves favorable performance in both intrinsic codec quality and representative downstream tasks, including typing (emg2qwerty), hand-pose (emg2pose), speech decoding (emg2speech), and speech-to-EMG synthesis (speech2emg). Across these tasks, MyoCodec exhibits strong performance against prior models while providing a compact and causal EMG representation. During streaming inference, it requires compute time of only 0.482 ms for each 20 ms frame, enabling real-time streaming. Also, the discrete token representation provided by MyoCodec has the potential to support integration into language-model based approaches, creating a path toward LLM-based interactive systems, where tokenized EMG representations are directly processed into such language or speech models. Code and model weights are released.
cs.AI / 220 / 2609.36007
Infrared Subtraction with Artificial Intelligence
Abstract
We present AI-developed local infrared subtraction, building on projection to Born and EFT matching. The framework separates an integrable radiation term from a finite contribution at Born kinematics, referred to as the Born contact. The contact is determined using the EFT singular distribution in a resolution observable such as N-jettiness $τ_N$. Under human physics guidance, an LLM develops two implementations. One uses a neural network for phase space projection and fits the contact by matching to EFT cumulants. The other uses an analytic construction that keeps the Born momenta fixed while integrating over radiation. It combines the EFT $δ(τ_N)$ coefficient with finite 4-dimensional radiation integrals to calculate the contact term directly. This gives a local subtraction formula without a slicing parameter, while reusing existing lower-order radiation calculations and EFT singular predictions. As a demonstration, we reconstruct the full NLO correction for massless 3- and 4-jet production in electron-positron annihilation. The attempt to the NNLO dijet production is also made by recursively using the NLO P2B construction with the LLM designing machine-learning controls to reduce the variance of the contact integral. The tested predictions are in good agreement with EERAD3. The numerical calculation and projection-network training use a 2020 Apple M1 MacBook, without GPU acceleration, illustrating the feasibility of the construction with modest computing resources. The appendices develop an extension of the local subtraction to 3-jet NNLO, giving explicit radiation maps and a proposed contact formula. We also show how to integrate over NNLO radiation while keeping the Born momenta fixed, for any number of massless final-state jets. Our results demonstrate how AI can help higher-order calculations by constructing infrared subtraction and improving its numerical integration.
cs.AI / 221 / 2609.36600
Second-Moment Stochastic Approximation Methods
Abstract
Classical stochastic approximation methods rely on estimators of the first moment (mean) of a random regression function. We study methods that employ estimators of both the first and the second moments, which include modern deep-learning optimizers such as Adam and Muon as special cases. We derive second-moment stochastic approximation methods through the lens of optimal preconditioning for solving matrix equations, and develop a two-stage framework for their convergence analysis. The first stage focuses on the analysis of conceptual (impractical) methods that rely on the exact first and second moments. In the second stage, we replace the exact moments with their respective estimators, and invoke Dvoretzky's theorem to show that the resulting practical methods converge almost surely to a neighborhood of the target solution. The size of the neighborhood depends on the biases and variances of the first- and second-moment estimators. We derive concrete bounds for Muon and a spectral variant of Adam that determine the radius of their neighborhood of convergence.
cs.AI / 222 / 2609.36398
Where Should Physics Enter a Molecular Crystal Generator?
Abstract
Generative models make molecular crystal structure prediction fast, but their samples still exhibit geometric and packing violations. Physics can be introduced during training, post-training, or inference, yet these choices are rarely compared with the generator and physical signal held fixed. We introduce CrystAF, an all-atom crystal flow-map generation model, and use it with the UMA interatomic potential to systematically study where physics should enter. Post-training learns physical preferences directly into CrystAF, improving molecular validity and crystal packing while leaving sampling unchanged: physics is paid for once during training rather than repeatedly at deployment. In contrast, UMA relaxation is effective at repairing local clashes but makes generation 6--26$\times$ slower, while learning from relaxed targets provides little benefit. These routes are complementary rather than competing. Physics-informed post-training first shifts the generated distribution toward more physically reasonable structures, after which inexpensive inference-time corrections further remove clashes and restore stereochemistry that the generator cannot represent. Importantly, the same post-training strategy also improves the multi-step all-atom Clari-M and rigid-body MolCrystalFlow generators, demonstrating transfer across architectures and representations. Together, our results suggest a simple principle: learn reusable physical alignment into the generator, and reserve inference-time physics for residual constraints that are better corrected than learned.
cs.AI / 223 / 2609.36366
Cross-attention encoding models reveal dynamic spatiotemporal routing across human higher visual cortex
Abstract
Understanding how the brain parses actions and events from time-varying natural inputs is a central challenge in neuroscience. Recent work has used deep neural network (DNN) models to build stimulus-computable fMRI encoding models that predict single-voxel responses to complex natural videos. However, the majority of video-computable encoding models predict responses using simple linear mappings from model tokens, overlooking the spatiotemporal structure shared by video representations and neural responses. Recent cross-attention encoding models address this limitation for static images, enabling flexible stimulus-dependent weighting of image content across space. Here, we extend this framework to naturalistic video, using per-parcel cross-attention to dynamically route features from a self-supervised video model (V-JEPA-2) across both space and time, fitting this model to fMRI responses to short video clips. We compare joint spatiotemporal attention with factorized and selectively constrained alternatives, and find that joint routing improves predictions of brain responses to held-out videos across higher visual regions, most consistently in lateral and dorsal visual areas associated with dynamic motion perception. Moreover, our method provides interpretable, stimulus-specific attention maps that dynamically follow moving objects, revealing which locations and temporal moments contribute to each neural response. We further show that attention maps from parcels in different category-selective networks (face-, body-, scene-selective) differentially weight content in accordance with expected semantic selectivity. Together, this work provides a new computational framework for understanding how visual information is adaptively weighted by cortical populations during dynamic visual perception.
cs.AI / 224 / 2609.37885
Boids of a Feather Flock Together - Evolving Prey Behaviours Under Different Predator Attack Strategies
Abstract
Flocking and schooling are thought to have evolved partly as defences against predation, but how prey should balance social and escape tendencies may depend on the predator's hunting strategy. We extend the predator-prey boids model of Ojo et al. (2023), itself based on Reynolds' boids, by combining six prey movement tendencies (alignment, cohesion, separation, dodge, repel and wiggle) into a single weighted acceleration update, and by reformulating wiggle as a sinusoidal manoeuvre. We then use an evolutionary strategy to optimise the six behaviour coefficients for collective prey survival against four predator hunting strategies: attack-centroid, attack-nearest, attack-random and attack-peripheral. Across five independent trials per strategy, coefficients converged within trials and mean fitness remained stable or increased, although trials often settled in different local optima. Prey survival was highest under attack-centroid and lowest under attack-nearest, in line with our hypotheses. Against attack-centroid, prey evolved individualistic predator avoidance with high escape coefficients, whereas against the other three strategies they largely kept their flock formation. Across all strategies, evolution favoured a low repel coefficient and relatively high dodge and wiggle coefficients. Our results suggest that optimal anti-predator behaviour depends on the interplay between escape tendencies and the predator's hunting strategy.
cs.AI / 225 / 2609.36479
Quantum Computing for Network Security Classification: Near-Term Classification and Long-Term Memory Efficiency
Abstract
Quantum computing has already been explored in several network-security applications. However, how quantum computing may contribute to network-security classification in both the near term and the longer term has not been systematically discussed. This paper studies this question through two complementary experiments. First, we evaluate near-term quantum-kernel support vector machines (SVMs) on practical network-security classification tasks and compare them with classical SVM baselines on KDD Cup 1999, CICIDS2017, and BoT-IoT. Across these runs, quantum kernels are competitive. They can match or improve classical baselines in some settings, while classical RBF kernels remain stronger in others. This suggests that near-term quantum-kernel methods should be evaluated as practical, dataset-dependent alternatives to classical kernels rather than as uniformly superior replacements. Second, we use quantum oracle sketching (QOS) to study a longer-term memory advantage for classification with streaming classical samples. In QOS, samples are processed online and used to incrementally construct an approximate quantum oracle, which provides coherent query access for downstream quantum algorithms without retaining the entire dataset. Under the QOS-inspired machine-size estimate, comparable accuracy corresponds to a substantially smaller effective memory-size proxy than explicit sparse/QRAM-style storage. Compared with a simple streaming proxy, the result is more nuanced because aggressive feature filtering can make the streaming dimension small. This suggests that the long-term value of quantum computing for network-security classification may lie in memory-efficient data access rather than immediate runtime speedup. Together, these experiments show how quantum computing may contribute to network-security classification from near-term classification performance and longer-term memory efficiency.
cs.AI / 226 / 2609.36901
Digital Twin Modeling of Quantum Dynamical Systems: Dissipative Quantum Reservoir Computing
Abstract
Modeling the response of driven many-body quantum systems from input--output data is difficult: the dynamics are nonlinear, history dependent, and expensive to simulate as system size grows. A paradigmatic case is High-Harmonic Generation~(HHG), where a strong field drives a medium to emit radiation that is highly sensitive to the drive and encodes long-range temporal correlations. We introduce a dissipative quantum reservoir computing~(DQRC) framework that builds a digital twin of such a system, learning its input--output map directly from data while the reservoir---itself a small open quantum system---stays fixed and only a classical readout is trained. We show that a minimal single-qubit reservoir reproduces the HHG response of a substantially larger Ising spin chain, and on a representative benchmark matches and on several metrics surpasses previously reported temporal convolutional and Kolmogorov--Arnold-network models, while using a simpler, physically realizable system. A single fixed reservoir further generalizes across a broad range of drives, indicating that it learns a shared physical response structure rather than memorizing trajectories. These results establish dissipative quantum reservoirs as compact, physically grounded digital twins for nonlinear, memory-dependent quantum dynamics. Code is available at \href{https://github.com/AI-and-Quantum-Computing/DQuRC}{https://github.com/AI-and-Quantum-Computing/DQuRC}.
cs.AI / 227 / 2609.37134
SQUARE: Structured Quantum Representation Adapters as Compact Quadratic Feature Maps for Frozen Language Models
Abstract
Frozen language models (LMs) are increasingly used as fixed feature extractors for downstream reranking, scoring, and preference modeling, raising a practical question: how should a compact module represent interactions among features in a fixed low-dimensional bottleneck? Common linear and low-rank adapters remain linear at the adaptation module itself, whereas explicit second-order alternatives introduce pairwise interactions through direct parameterization or predefined factorizations. We propose SQUARE, a Structured QUAntum REpresentation adapter that amplitude-encodes the bottleneck vector, applies a parameterized quantum circuit, and measures the resulting state. We show that each basis-probability feature is exactly a normalized quadratic form in the bottleneck coordinates, while the additional Pauli-$Z$ readouts are signed linear combinations of these probabilities. The measured map can therefore parameterize interactions over $O(d^2)$ coordinate pairs through a small set of shared circuit parameters, where $d$ is the bottleneck dimension. It provides a structured parameterization within, rather than beyond, the classical normalized-quadratic feature class. In a disjoint same-pipeline evaluation over eight GLUE-derived controlled interaction tasks and five shared seeds, SQUARE achieves an average test accuracy of $0.7565$, compared with $0.7355$ for an affine normalized-quadratic predictor, $0.7271$ for the evaluated parameter-matched Givens mixing model, $0.6817$ for an MLP, and $0.6155$ for a frozen-circuit control. Under reduced supervision, it also shows consistent gains over the strongest evaluated classical comparator, with the same qualitative pattern across multiple frozen LM backbones. All circuit experiments use simulation, while the learned feature map can be evaluated exactly in batched PyTorch without quantum hardware.
机器学习 (cs.LG)
264
cs.LG / 1 / 2609.35966
SIFARI: Self-Supervised Interferometric Fitting for Astronomical Radio Imaging
Abstract
Radio-interferometric images are reconstructed from sparsely sampled visibilities, and CLEAN-based imaging can struggle with spatial filtering, complex morphologies, and uncertainty quantification. Alternative methods that fit visibilities directly can address some of these limitations but often require manual choices of image priors and model hyperparameters. We present SIFARI (Self-Supervised Interferometric Fitting for Astronomical Radio Imaging), a self-supervised neural network workflow that represents sky brightness as a continuous function of position and fits measured visibilities without an external image training set or explicit spatial regularizer. An empirical rule sets the Fourier feature scale from the visibilities before training, controlling how readily the network fits fine structure. Sampling network weights with Stochastic Weight Averaging-Gaussian (SWAG) gives approximate brightness uncertainty estimates, which we combine with a thermal-noise floor to construct spatially resolved signal-to-noise maps. In synthetic ALMA tests, SIFARI yields an effective point-source response about eight times narrower than the natural-weighting CLEAN restoring beam and recovers more extended flux than CLEAN when short baselines are missing. It also achieves higher image fidelity than the restored CLEAN images in all three morphology benchmarks. Applied to ALMA observations of PDS 70, SIFARI recovers the bright outer ring together with faint compact emission in the central cavity. For long-baseline-only WISPIT 2 data, SIFARI supplies a sky model for phase self-calibration where the CLEAN model is inadequate. The restored, self-calibrated SIFARI image has approximately 30% lower RMS noise than the CLEAN image made from the original visibilities without self-calibration.
cs.LG / 2 / 2609.36421
Sample Complexity of Equivariant Reinforcement Learning
Abstract
Reinforcement learning (RL) is a powerful framework for robotic control, yet its practical application is often hindered by high sample complexity. This is particularly restrictive in physical domains where interaction data is costly. While the world often exhibits geometric and physical symmetries, standard RL algorithms typically fail to exploit this structure. In this paper, we demonstrate that exploiting group symmetries significantly reduces the sample complexity of RL. Focusing on finite-horizon Markov decision processes, we find that leveraging homomorphisms induced by group symmetries significantly reduces the theoretical upper and lower bounds on the number of environment interactions required to reach an optimal return. We further extend these bounds to continuous state and action spaces, providing corresponding sample-complexity guarantees under appropriate regularity assumptions. Beyond theory, we validate our findings through controlled experiments and demonstrate the advantages of symmetry-aware policy learning on high-dimensional continuous robotic simulations. Our results show that integrating symmetry into the learning pipeline yields substantial gains in sample efficiency and performance, offering a principled path toward more data-efficient robotics.
cs.LG / 3 / 2609.36172
Exploring Learning Models for Topological Relationship Recognition from Image Data
Abstract
Figuring out how objects relate to each other, like whether they touch, overlap, stay completely separate or one sits inside another, matters a lot in fields like GIS, biomedical imaging, and robotics. Even though machine learning has come a long way, people haven't really focused on spotting these topological relationships in images. The main roadblocks? Not enough good datasets and no clear way to measure results. So, we rolled up our sleeves and built a new dataset. It's pretty sizable: over 11,000 labelled images showing all those essential relationships. We ran tests with some classic machine learning models, Naive Bayes, KNN, Random Forest, SVM, and Artificial Neural Networks, and threw in some deep learning stars like VGG16 and InceptionResNetV2. For the dataset itself, we used segmentation, contour detection, and grayscale normalization to tease out solid feature vectors. The results? Deep learning methods, especially VGG16, pulled ahead, with validation accuracy hitting 89.55%. That's a big jump compared to the traditional models. This shows how powerful transfer learning is for analyzing topological relationships in images, and it gives researchers a new standard to aim for in future work on spatial reasoning and topological classification.
cs.LG / 4 / 2609.36352
StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks
Abstract
Vision-language-action (VLA) models perform well on shorter-horizon manipulation tasks but still struggle with long-horizon tasks that require multiple dependent manipulations from a single command. Online reinforcement learning (RL) can improve these policies through environment interaction, yet many existing methods provide reward only after the complete task succeeds. However, such terminal supervision is sparse and does not distinguish early failures from rollouts that make substantial partial progress. We propose StructRL, an online RL framework that constructs structured intermediate supervision from verifiable subtask completions. StructRL decomposes each task into verifiable subtasks, grants intermediate rewards only after the prerequisite subtasks have been completed, and scales each reward according to completion pace. Across RoboCasa365 and LIBERO-Long with GR00T-N1.5 and pi 0.5, StructRL consistently outperforms evaluated online RL baselines. These results show that verifiable, structured intermediate rewards improve long-horizon VLA post-training. Code is available at https://github.com/amazon-science/StructRL.
cs.LG / 5 / 2609.37177
Sparse cubical complexes for efficient topology-preservation in image data
Abstract
Persistent homology (PH) is a frequently used tool for extracting and preserving topological information from image data, particularly in image segmentation, where preservation of topological structures is important. However, despite its general applicability across dimensionality, domains, and target structures, the runtime cost of PH-based methods often makes their practical use infeasible. In this work, we argue that this runtime cost is largely driven by processing information that is unimportant for downstream application (e.g. as optimization objective). We propose sparse cubical filtrations as an alternative foundation for PH computation, reducing subsequent computational costs by factors of up to 100 on real datasets. We show close agreement with the optimization signal of the dense counterpart and empirically evaluate our solution's effectiveness as an optimization objective in realistic training regimes where other PH-based objectives can practically not operate (i.e., 3D data with large patch sizes). We show how our solution improves topological accuracy by up to 80\% across six diverse datasets while maintaining pixel- and region-based accuracy.
cs.LG / 6 / 2609.37230
Seeing Is Not Addressing: Auditing Linguistic Access to Frozen Visual Geometry
Abstract
Visual distinctions are often finer than those reflected in linguistic conceptualization. Vision-language models exhibit a similar asymmetry: a distinction can remain discriminable in frozen image geometry while being weakly addressable through the native text interface. We study this gap by separating visual discriminability from linguistic addressability in text-to-image retrieval. Using FactorAtlas, a fully crossed testbed of 23,040 images spanning shape, hue, pattern, and nuisance variation, we compare both readouts on held-out images of the same distinctions. We then derive image-side contrasts that separate each value from its alternatives for matched visual grounding, and test whether this reduces the native-text access gap across factors and models. Direction-specific and visual-absence controls tie these gains to the relevant visual contrast; the gains persist after global alignment and extend to compositional retrieval and natural images. Together, these results show that visual discriminability and linguistic addressability need not coincide, and that matched visual grounding can probe and reduce the resulting access gap.
cs.LG / 7 / 2609.37330
The Domain Is a Residue: Adapting Self-Supervised Features, Not Generators
Abstract
Clearing fog, rain or snow from footage, or turning renders into photographs, must remove the source domain and keep the scene. Unpaired translators carry it through because their generator sees the source appearance (pixels, a near-invertible latent or a control map) and keeps it. A DINO feature map fixes what is in the scene and carries weather, lighting and rendering style as a residue of 13 to 14% of the feature norm. We propose the Representation Feature Adapter (RFA), a 2.9M-parameter network that moves this residue. We train only the adapter and its discriminators; the encoder and a feature-conditioned decoder, trained once for all conditions, stay frozen. Against CycleGAN-Turbo it is ahead on both metrics on fog and on KID on night, and level within noise on snow, rain and haze. On sim-to-real it leads REGEN and HyPER-GAN on both metrics. Only the RFA removes the rain while keeping the scene. The removal costs scene structure: CycleGAN-Turbo keeps more on every condition but fog. On VAE latents the identical adapter collapses to the identity, and decoders from other groups that never saw it render its output. The RFA has about 160 times fewer trainable parameters than CycleGAN-Turbo and under a fifth of its per-condition training time.
cs.LG / 8 / 2609.37387
Multi-task learning for the automatic grading of enlarged perivascular space burden using MRI
Abstract
Enlarged perivascular spaces (PVS) visible in brain magnetic resonance imaging (MRI) are increasingly thought to be linked to poor brain health. PVS are elongated structures of less than 3 mm in diameter and can be numerous. To reflect the incidence of PVS, radiologists visually score their burden following a clinical grading scale - a task that would benefit from automation to accelerate analyses and overcome the influence of inter-observer differences. We developed and evaluated methods for training machine learning models to score PVS incidence in the basal ganglia (BG) and centrum semiovale (CSO) leveraging the Potters/Wardlaw scale. The novelty in our work lies in the use of imperfect, semi-automatically generated "silver-standard" PVS segmentation masks during training, in addition to PVS radiological scores. We comparatively evaluated a conditional convolutional neural network (CNN) which accepts PVS masks as an extra input channel, a multi-task CNN which performs both PVS segmentation and scoring, and a logistic regression model which utilises features derived from PVS masks to predict PVS scores. Multi-task learning was the most effective method, achieving a mean average precision of 64.08% compared to 60.22% for the conditional CNN, 52.11% for a baseline CNN trained only to predict PVS scores, and 49.32% for the logistic regression model. The multi-task model showed an ability to localise individual PVS not shown by the other CNNs, and behaved in a probabilistically sensible way, predicting with lower confidence on inherently harder classes. Age, sex, hypertension status, white matter hyperintensity volume, and ischaemic stroke lesion status were shown to be associated with the multi-task model's PVS score predictions and the ground truth in a similar way.
cs.LG / 9 / 2609.37605
TomoTransformer: Towards a Foundation Model for CT Reconstruction
Abstract
Supervised deep learning has advanced sparse-view tomographic reconstruction. However, conventional models, which typically map filtered back-projection (FBP) images or sinograms to clean reconstructions, are brittle under distribution shifts. Because they require retraining whenever projection counts and angles, detector resolutions, or data distributions change, their deployment in real-world applications remains limited. To address this, we introduce TomoTransformer, a transformer-based architecture that treats each \textit{local} filtered projection as an individual token and predicts missing views via self-attention. Crucially, TomoTransformer operates in a \emph{back-projection space} that separates projections across spatial locations, making view interpolation geometrically well-posed and invariant to detector size. This design yields a single foundation model that can process any number of input projections, at arbitrary angular locations and detector dimensions, and query any number of target angles without retraining. Trained on a large-scale dataset spanning diverse medical CT anatomies and natural images, TomoTransformer generalizes effectively across anatomies, materials, and resolutions. Extensive evaluations on several benchmark sparse-view datasets show that TomoTransformer significantly outperforms concurrent multi-purpose models like ViewTrans and matches or exceeds strong protocol-specific baselines, while remaining fully agnostic to the number of input and target projections. Furthermore, the model demonstrates robust zero-shot generalization on real experimental nanoscale brain data collected from an X-ray synchrotron, showcasing its practical utility for real-world applications.
cs.LG / 10 / 2609.37732
The Camera Inside the Editor: Reading the Implicit Camera of Image Editors with Painted Calibration Patterns
Abstract
Instruction-based image editors insert objects, restyle scenes and render new viewpoints, but it is unknown which camera they assume when they paint into a photograph. Asked to cover the floor with a checkerboard, an editor paints projective structure from which classical vanishing-point geometry reads pitch, roll, focal length, yaw and, on renders, the principal point, without any training. Unlike a calibrator such as GeoCalib, which estimates the camera of an image, this isolates the camera under which the editor paints. On 120 rendered cameras with exact ground truth, Qwen-Image-Edit-2511 paints tile edges that meet their vanishing points within 0.26 degrees, and its implicit camera matches the true one to 0.8 degrees in pitch and 6% in focal length, more accurately than GeoCalib except in roll. Asked to draw the horizon or mark a vanishing point instead, the editor fails, so this knowledge is revealed by painting and not by the explicit tasks we tried. The implicit camera has two priors: roll is pulled towards level (slope 0.71), and telephoto perspective towards a default of about 30 mm, which roughly matches the camera the models paint without any scene. For Qwen, the priors do not grow when blur removes four fifths of the line evidence. They are stronger on real photographs, and on NYUv2 a shorter wording of the task removes the difference for roll. On photographs from a 24--240 mm zoom lens the painted perspective grows with only 0.62 of the lens's slope, while GeoCalib and MoGe-2 saturate at about 52 and 42 mm. FLUX.1 Kontext and LongCat-Image-Edit are pulled much harder. Finally, from a level camera a camera-control LoRA executes pose commands at only 50--70% of their strength, and a board painted into its output agrees with the camera it produced.
cs.LG / 11 / 2609.37784
Planetary Feature Fields are Scalable Earth Representations
Abstract
Satellite observations, precomputed embeddings, and map products describe the same evolving Earth, yet are stored as independent, petabyte-scale data products. Their continued growth calls for compact representations of multiple products while preserving spatial and temporal detail. We introduce Planetary Feature Fields (PFFs), which exploit redundancy across data products by modeling them jointly as continuous functions of space and time at planetary scale. PFFs are spatially local explicit-implicit (hybrid) neural fields. Each field shares a factored feature volume---a decomposition of an explicit 3D grid with smaller factors---across products, while lightweight implicit decoders reconstruct individual products across multiple timesteps. PFFs reconstruct EO products over space and time more accurately than single-product fields at matched compression rates. At $1800\times$ compression relative to the uncompressed source data, reconstructed features retain approximately $90\%$ or more of the performance achieved with the original features on pixel-level segmentation, change detection, and patch-level classification tasks. PFFs can add new timesteps by extending their factored feature volumes and add new products by attaching new decoders, while leaving existing outputs unchanged. PFFs reduce end-to-end feature access latency by an order of magnitude relative to evaluated API and cloud-storage pipelines.
cs.LG / 12 / 2609.37848
Evaluation Choices Shape Biomedical ML Claims: A Pediatric Pneumonia Benchmark Case Study
Abstract
Biomedical machine learning papers often compress model performance into one headline number. That number can look like a property of the model even when it depends strongly on how the benchmark was evaluated. We study this problem on the widely used Kermany pediatric chest radiograph dataset using nine image classifiers and a controlled evaluation protocol. Under the same protocol, the eight pretrained backbones differ by only 0.026 AUROC. In contrast, changing whether the backbone is frozen or fine-tuned changes AUROC by 0.044 on average, and changing the decision threshold changes balanced accuracy by 0.090 on average. The official test split is also measurably different from the training pool: a partition classifier distinguishes them at AUC 0.697, rising to 0.898 for normal radiographs. Most strikingly, a classifier using only file properties, with no image anatomy, reaches 0.992 balanced accuracy within the training pool but falls to 0.496 on the official test split. Validation-fitted thresholds and calibration also transfer imperfectly. These results show that a high benchmark score can support different conclusions when the split, training policy, threshold, metric, calibration, and uncertainty are not communicated with it. We end with a seven-item reporting recommendation in which each item is tied to an effect measured in the study
cs.LG / 13 / 2609.37851
FlowMap-OPD: Rollout--Kernel Separation for On-Policy Distillation of Few-Step Flow-Map Generators
Abstract
Few-step flow-map generators, including MeanFlow and consistency models, enable efficient sampling through long-range transport, yet their on-policy distillation remains underexplored. We introduce FlowMap-OPD, an on-policy distillation framework that separates student-state acquisition from teacher--student distribution comparison. A formulation based on state marginals establishes this separation, while flow--velocity consistency connects local supervision to the deployed long-range map. Within this framework, we develop flow-map, induced-velocity, and instantaneous-velocity distribution supervision, each paired with a separately specified native flow-map rollout. Cross-capacity ImageNet experiments across three teacher rewards identify instantaneous-velocity distribution supervision with independently tunable student consistency as the most effective choice. In text-to-image experiments, FlowMap-OPD demonstrates strong multi-specialist consolidation capabilities and surpasses multi-reward Flow-Map GRPO in task performance and convergence speed.
cs.LG / 14 / 2609.37888
Visual Branch is What You Need for CLIP-based Class-Incremental Learning
Abstract
Class-Incremental Learning (CIL) requires models to recognize new classes over time without forgetting previously learned ones. With the rise of vision-language pre-training, CLIP has become a strong foundation for CIL. A common design in CLIP-based CIL is to construct textual classifier weights by encoding class-name templates with the CLIP text encoder, and then classify visual features by image-text cosine similarity. This design is appealing: since CLIP aligns images and text in a shared embedding space, textual weights appear to provide an off-the-shelf classifier for incremental classes. However, we show that this seemingly natural design is not always beneficial, as a modality gap can still separate the two modalities and make textual classifier weights deviate from visual class distributions. Empirically, under identical task-wise CIL training, initializing the cosine classifier with visual class centers yields lower loss and better incremental accuracy than using CLIP textual features.Motivated by these observations, we propose VIS, a visual-only method for CLIP-based CIL that removes the deployed textual branch and constructs the incremental classifier entirely in the visual space. To obtain stronger task-adaptive visual representations, VISuses only base-session data to enhance CLIP's final visual representation with informative visual-layer features. Built on the enhanced visual representation, VISemploys a simple kernelized incremental least-squares SVM, whose classifier weights are solved in closed form from additive sufficient statistics. When new classes arrive, VISaccumulates their sufficient statistics and recomputes the classifier weights for all seen classes, enabling efficient incremental updates while preserving historical class knowledge. Extensive experiments show that VISachieves state-of-the-art performance without a textual branch.
cs.LG / 15 / 2609.38165
Cropland PAtteRNS: Parallel Dimensional Attention Networks and Attention to Dataset Disparity for Crop Segmentation in Satellite Imagery Time Series Data
Abstract
The landscape of satellite imagery time series datasets and boundary-pushing architectures for cropland segmentation has never been richer. However, in this gold rush, important truths are being missed on both fronts, as a drive for the most novel concepts or the largest datasets pushes finer details to the side. In this paper, we present our hybrid transformer-convolutional model, Cropland Parallel Attention and Refinement Network for Segmentation (PAtteRNS), the first model to use self-attention mechanisms separately for each of the temporal, spectral, and spatial aspects of Sentinel-2 multispectral SITS data. To achieve fully-factorised attention in our proposed model, we introduce a novel parallel transformer architecture which significantly reduces the computational complexity of triple-factorised self-attention. We validate our architecture with an in-depth ablation study, and analyse the performance of our model against state-of-the-art crop segmentation models on multiple tile-size variants of the popular PASTIS and MTLCC datasets. Our findings show our model to outperform all others in the task of crop class segmentation, verified across multiple important segmentation metrics, with especially strong performance against compared models seen in the often under-reported parcel delineation quality, for which we use the Boundary IoU metric. We also find that flawed class groupings within datasets can have a significant negative impact on model performance, and report that alternate tile-size variants of crop segmentation datasets produce results incomparable to one-another, invalidating fair comparison between model performance when trained on different tile-sizes. Based on these findings, we suggest further work is required to standardise best practices when constructing SITS crop segmentation datasets, and to enable future dynamic-tile-sizing for ideal model performance.
cs.LG / 16 / 2609.36070
Mixture-of-Kittens: MoE Megakernel for NVL72s
Abstract
AI accelerator systems are rapidly consolidating into scale-up architectures, where tens to thousands of GPUs communicate over high-bandwidth, single-hop fabrics. We find that existing Mixture-of-Experts (MoE) training systems, optimized for conventional scale-out networks, transfer poorly to this setting, often running slower than a naive baseline built with PyTorch and NCCL. With industry roadmaps pointing toward even larger scale-up domains, understanding the performance tradeoffs of this hardware regime is increasingly important. We present Mixture-of-Kittens (MoK), an MoE training system designed for Nvidia NVL72. MoK builds on three insights that unlock performance on scale-up domains: (1) choosing push- or pull-based communication per operator, (2) restructuring the computation-communication overlap, and (3) fully eliminating CPU-GPU synchronization. MoK distills these insights into a single deterministic training megakernel that fuses token dispatch, shared and routed expert FFNs, and token combine. Across MoE layer shapes from four widely used open-weight models, MoK delivers up to $2.37\times$ the throughput of the strongest publicly available baseline. In a production run on 512 GPUs spanning multiple GB300 NVL72 racks, MoK improves end-to-end training throughput by $1.41\times$.
cs.LG / 17 / 2609.37062
vSkipper: Translating Dynamic Layer Skipping into LLM Serving Gains
Abstract
Dynamic layer skipping reduces LLM computation by allowing each token to execute only a subset of the model's layers. However, existing skippers rely on specialized generation loops and do not integrate with modern serving engines. As a result, fewer executed layers do not necessarily translate into lower serving latency: FlexiDepth skips 8 of Llama-3-8B's 32 layers on average, yet its standard generation loop decodes 14.6--21.0% more slowly than the base model. We present vSkipper, a virtualization layer that makes dynamic layer skippers pluggable in serving engines while preserving continuous batching, fixed-shape batches, paged KV caching, and captured decode graphs. At each routed layer, vSkipper groups tokens by the skipper's decision and uses routed execution only when predicted to be profitable. We implement vSkipper in SGLang and evaluate the released FlexiDepth checkpoint against upstream SGLang under identical prompts, arrivals, output lengths, and launch settings. At the knee of upstream's load curve, vSkipper reduces mean end-to-end latency by 36.8% on GSM8K and 13.6% on BBH. Under saturation, it increases request throughput by 11.3% and 7.4%. Serving adds no statistically resolved quality loss beyond the checkpoint's own. Across synthetic skip policies, two Qwen3 skippers, and three GPUs, we demonstrate reuse without workload-specific tuning. To our knowledge, vSkipper is the first system to realize serving-efficiency gains from per-token interior layer skipping within a modern LLM serving engine. The code is open-sourced as an SGLang fork at https://github.com/AKafakA/sglang-vskipper/tree/vskipper-ref
cs.LG / 18 / 2609.37575
On Task Scope and Information Retention in Source Coding
Abstract
We argue that dividing codec design into Coding for Machines (CfM) and Coding for Humans (CfH) is a misleading distinction for deciding what information a codec may discard. Receiver identity does not determine admissible information loss. The required rate depends on task scope, including the predictions to support, their losses and tolerated risks, the encoder observation, and the permitted decoding procedures. Notably, a machine task may have a higher minimum rate than a restricted human decision. Rate savings on selected machine tasks apply only to the stated requirements, not to an intrinsic ordering by receiver type. We extend source and feature coding to finite task families, derive when restricting the encoder observation preserves the minimum rate, and show that equality between source and split-feature coding rates can no longer hold as the task scope expands.
cs.LG / 19 / 2609.37994
Mutual Information Constrained Chernoff Bottleneck
Abstract
The classical information bottleneck (IB) measures the relevance of a representation $U$ of $X$ to a target $Y$ by $I(U;Y)$, which does not directly characterize the error of downstream decisions. For a binary hypothesis $Y$ inferred from many separately encoded observations, the optimal error exponent is the Chernoff information between the two conditional distributions of $U$ given $Y$. We study the mutual information constrained Chernoff bottleneck, which seeks an encoder that maximizes this Chernoff information subject to a rate constraint $I(U;X) \leq R$. We show that its optimal value $C(R)$ increases strictly up to $R = H(V)$, where $V$ merges the symbols of $X$ with equal likelihood ratio, remains at the uncompressed exponent beyond, and, unlike the IB curve, need not be concave. We further show that $k+1$ outputs suffice to attain $C(R)$, where $k$ is the cardinality of $V$. We propose an alternating algorithm that updates the encoder via a generalized Blahut--Arimoto algorithm and the Chernoff parameter $s$ via a nonlinear equation, and prove that its iterates remain feasible, with nondecreasing and convergent Chernoff information. Numerical experiments confirm the theory, and on real topic-detection data from the 20 Newsgroups corpus, compressing each word to only $17\%$ of its entropy retains $90\%$ of the error exponent and nearly the accuracy of the uncompressed classifier.
cs.LG / 20 / 2609.36001
Making Cross-Continental Federated Learning Repeatable with FLIP: a Multi-Application Study
Abstract
Federated learning (FL) in healthcare remains challenging, as the overhead of rebuilding governance guarantees for every collaboration stops most projects at the proof-of-concept stage. Here we present FLIP (Federated Learning Interoperability Platform), an open-source, multi-application platform that makes FL training and evaluation repeatable. FLIP implements common FL workflows as a set of composable services: cohort queries against per-site structured databases, on-demand DICOM retrieval from institutional PACS, per-site project approval, and reusable FL job types. To demonstrate FLIP, we ran two distinct use cases, federated fine-tuning and federated evaluation, on synthetic chest X-ray cohorts across two client nodes based in the United Kingdom (UK) and Thailand. In FLIP, each institution independently approves its participation in each project and operates its own node under local IT security processes. This study makes an operational rather than an algorithmic claim. It does not compare federated with centralised training; for that question, we refer the reader to existing systematic reviews and meta-analyses. The central result is evidence that such platforms enable international FL collaboration and improve repeatability, auditability, and site-specific governance. We also present a comprehensive comparison of existing platforms to help researchers and operators choose the right platform for their use case.
cs.LG / 21 / 2609.36030
Preferent Compression Bounds Are Tight
Abstract
The lack of rigorous safety and performance certificates remains a key bottleneck to the deployment of modern learning-based methods. Sample compression has recently emerged as a powerful tool for deriving such certificates, with particularly sharp bounds available for algorithms satisfying a so-called preference property -- also known as stability in the learning theory literature. These bounds find direct application across domains as different as the Scenario Approach, Pick-to-Learn, and Support Vector methods. However, whether they are tight has remained an open problem. In this paper we resolve this question affirmatively and show that the state-of-the-art bound for preferent compressions is provably tight. We establish this by exhibiting an explicit construction based on the uniform distribution and order statistics that attains the bound in the limit. Along the way, we also provide a considerably shorter and more accessible proof of this bound, requiring only elementary counting arguments and no infinite-dimensional duality.
cs.LG / 22 / 2609.36032
TORQUE: Optimizing What (not) to Quantize Before and After Rotation
Abstract
Uniform random rotations are an effective preprocessing step for quantization: they make normalized coordinate distributions approximately Gaussian, enabling the use of codebooks optimized offline. We introduce TORQUE, a framework that improves on previous quantization works that use random rotations by jointly optimizing how many and which coordinates to preserve at high precision both before and after rotation, under a fixed overall expected bit budget. Intuitively, before rotation, preserving large input coordinates at high precision can reduce overall error by preventing the rotation from spreading their values across many coordinates. Likewise, after rotation, preserving a small fraction of the largest-magnitude coordinates at high precision allows the remaining values to be quantized more accurately using codebooks optimized offline for the resulting truncated Gaussian distribution. We derive a quantization error upper bound and prove that top-$k$ pre-rotation retention minimizes it for each $k$. This reduces the search over coordinate subsets to an optimization over $k$, enabling a fast optimizer that uses offline codebooks and parallel parameter selection for practical implementation. We demonstrate an improved tradeoff between reconstruction accuracy and storage cost through numerical evaluation under the Gaussian model and experiments on nearest-neighbor retrieval, KV-cache compression, and activation compression.
cs.LG / 23 / 2609.36047
Neural networks for spectral optimization
Abstract
Given a functional dependent on the spectrum of a differential operator, we address the problem of finding a domain which optimizes this functional. PDE solvers might be used to tackle this optimization. It is however computationally expensive. We propose two neural network models which learn the spectrum directly from the geometry of the domain and can be used to optimize the domain from one or more eigenvalues. We investigate two representations. The first encodes the domain through Fourier coefficients and a light MLP, which is efficient on star-shaped geometries, achieving a precision of 0.2\%. Through a rescaling of the coefficients the designed models satisfy the scaling law of the eigenvalues. Additionally, averaging the outputs of the trained surrogates over rotations and reflections induces invariance for these transformations. The second is a model that takes the landscape function, the indicator function and the gradient of the landscape function. A Gram-Schmidt process produces orthogonal eigenfunctions as output of the model along with the associated eigenvalues. The landscape model reaches 1\% mean relative error on the first ten eigenvalues, compared with 4\% for an FNO model. Replacing the landscape by an SDF worsened both prediction and optimization errors. The trained model also generalizes from synthetic shapes to domains given as classical image dataset. The resulting surrogates of both approaches recover classical spectral optima such as the disk for the first eigenvalue or the conjectured minima of higher eigenvalues. This confirms that our models produce accurate differentiable estimates of eigenvalues, which can be used in shape optimization problems involving spectral quantities.
cs.LG / 24 / 2609.36049
Improving scalable oversight with co-trained monitors
Abstract
Worker-monitor setups are a promising approach to AI oversight, but training workers against fixed monitors can incentivize monitor evasion. We study whether this failure mode can be avoided by co-training the monitor alongside the worker, and explore both supervised and self-supervised approaches. In the supervised setting, we prove a characterization: monitoring is possible with vanishing error and query rates exactly when the class of possible monitor functions has finite Littlestone dimension. This connects worker monitoring with an established literature on adversarial online learning. For self-supervision, we propose a co-training procedure based on test-time distillation: the monitor uses additional test-time compute to generate training labels, then trains its standard-compute policy on those labels. For majority-vote labels, we give a finite-sample sharpening guarantee under adaptive worker distributions with action coverage, that shows that the monitor's verdicts converge to its initial modal verdicts. We stress-test the former in code-security settings where the worker is trained adversarially to fool the monitor. Our results suggest that adaptive monitors are better at keeping pace with evolving worker strategies, while fixed monitors are more vulnerable to evasion.
cs.LG / 25 / 2609.36058
ABC: Advantage-Based Control Variates for Reinforcement Learning with Verifiable Rewards
Abstract
Recent progress in reinforcement learning with verifiable rewards (RLVR) has highlighted the effectiveness of simple critic-free policy-gradient methods such as Group Relative Policy Optimization (GRPO). In contrast, actor-critic methods rely on learned value functions whose approximation error can introduce bias through commonly used advantage estimators such as temporal-difference error. Motivated by this observation, we revisit trajectory-level control variates through an advantage-value formulation, which we call Advantage-Based Control Variates (ABC). This formulation reveals that the covariance structure is closely related to the return decomposition used in Direct Advantage Estimation (DAE). Finally, we combine ABC with DAE into a single actor-critic algorithm and evaluate it in an offline-to-online RLVR setting, where the critic is first trained on previously collected trajectories and adapted during online learning. On mathematical reasoning tasks, ABC achieves performance competitive with GRPO using substantially fewer online optimization steps.
cs.LG / 26 / 2609.36061
Introducing the CZAR Loss: A Tailored Objective Function for Financial Log-Return Predictions
Abstract
In quantitative finance, standard regression losses are misaligned with the economics of return prediction. As the conditional mean of financial log-returns is close to zero, symmetric losses such as the mean squared and mean absolute errors make the constant zero forecast a near-optimal solution, penalizing models with genuine but noisy directional skill. This applies both during training, where predictions shrink toward zero, and during evaluation, where trivial forecasters can lead loss-based rankings. Under a Gaussian linear prediction model, we show that all symmetric monotonic losses share a universal breakeven directional accuracy against the zero predictor, which rises sharply and becomes unobtainable as the prediction noise approaches the standard deviation of the returns. We introduce the CZAR (Composite Zero-Agnostic Return) loss function, a piecewise quadratic loss built around five key requirements: convexity in the prediction, asymmetry with true return direction that vanishes at zero, near-linear penalization of undershoots and wrong-direction predictions, divergence for large errors, and an adaptive loss floor for evaluation. CZAR is provably convex in the prediction at fixed true value, has closed-form gradient and Hessian suitable for custom objectives in gradient-boosted libraries, and its four hyperparameters reduce to a single choice through correlated defaults. In idealized tests, the minimum directional accuracy required for a CZAR-evaluated forecaster to outperform the zero predictor under mean log loss remains near the 50% chance level, whereas the corresponding threshold for symmetric losses rises sharply with prediction noise. This advantage persists under heavy-tailed return distributions. In a LightGBM experiment on intraday BTC log-returns, CZAR-trained models reduce the `zero-returns bias' and improve directional accuracy on large-magnitude returns.
cs.LG / 27 / 2609.36063
Understanding Decision-Making Mechanisms in Neural Routing Solvers
Abstract
Neural Combinatorial Optimization (NCO) has achieved strong empirical success, yet the internal mechanisms driving model decisions remain largely unexplored. In this paper, we investigate three representative autoregressive NCO models spanning two encoder-decoder configurations: AM and POMO (heavy-encoder, light-decoder), and LEHD (light-encoder, heavy-decoder). Through behavioral analyses, representation probing, and causal interventions, we examine how these models construct solutions and use internal representations during decoding. Our results suggest that AM and POMO predominantly follow a persistent geometric pattern throughout solution construction, whereas LEHD contains linearly accessible information about multiple future actions. Causal experiments further provide evidence for the role of future-node representations in LEHD's decision-making. We also observe that LEHD relies strongly on the current-node representation for immediate local decisions, while the start-node representation plays a broader navigational role over the subsequent route. Cross-instance alignment analyses additionally indicate that LEHD maps current-node representations into a relatively shared latent region, which may provide a stable reference for evaluating subsequent decisions. Across the Traveling Salesman Problem and the Capacitated Vehicle Routing Problem, these results reveal distinct decision-making patterns across these architecturally distinct solvers and provide a foundation for more interpretable analyses of NCO solvers. Code and additional visualizations are provided in the https://github.com/NCO-Interpretability/NCO-Interpretability.
cs.LG / 28 / 2609.36064
FLOORA: A Human-Aligned Domain-Specific Language Model for Architectural Design
Abstract
Foundation models are powerful generators, but many engineering domains require structured representations that general-purpose systems handle poorly. We introduce FLOORA (Floor Layout Optimization with RL Alignment), a family of small domain-specific language (DSL) models for architectural layout generation. With specialized data and alignment, our 0.6B model outperforms much larger frontier models, achieving VLM judge win rates up to 92.0% on out-of-distribution real-world buildings and 96.0% on synthetic buildings. Human evaluations further corroborate these results, with FLOORA selected as the best model in 89.3% of evaluations. FLOORA combines a token-efficient DSL, custom tokenization, domain-specific pretraining, supervised fine-tuning (SFT), and reinforcement learning (RL) with learned human-preference and verifiable rewards. This pipeline improves architectural and geometric validity, supported by extensive empirical evaluation and ablation studies. Although focused on architecture, our results suggest that similar domain-specific recipes may be useful in other engineering domains with structured, verifiable outputs. Datasets, models, and inference code are available at https://github.com/AutodeskAILab/floora.
cs.LG / 29 / 2609.36074
GEM-KMeans: Memory-Efficient and Accurate Clustering on Massive Scale with GPU Optimization
Abstract
Memory-efficient scaling on clustering problems without sacrificing statistical accuracy is of central interest for large-scale data analysis and machine learning problems. Nonnegative low-rank (NLR) matrix factorization for $K$-means is a scalable clustering method, which connects to semidefinite relaxations with optimal average-case exact recovery guarantees. However, a direct GPU implementation of NLR requires multiple large factor-sized buffers and substantial data movements that are essentially memory-bound. In this paper, we introduce GEM-KMeans, a spectrally normalized yet mathematically equivalent NLR formulation that fuses the gradient update, nonnegative projection, and sufficient statistics for normalization and iterate movement into a matrix-multiplication epilogue. Instead of retaining three massive factor-sized arrays, our IO-aware GPU implementation materializes only one single factor with small tile-reduction arrays as additional storage in the High Bandwidth Memory (HBM). We derive explicit memory costs and spectrally normalized smoothness bounds for optimizing the clustering objective function. Accurate clustering is demonstrated at massive scales on synthetic and real datasets, where performance gains of GEM-KMeans over existing GPU-accelerated Lloyd's algorithms involve data-dependent runtime tradeoffs.
cs.LG / 30 / 2609.36081
Early Learning Shapes Later Directions Of Representation Change In Continual Learning
Abstract
Representations continually change as a network learns new tasks. We ask whether early representational changes naturally form a geometric structure that continues to shape later learning. We identify a low-dimensional subspace of early representation drift, which we call a scaffold, and test whether it is reused across subsequent tasks. Across four pretrained visual encoders and two datasets, later representational changes consistently favor this early-defined subspace over matched random alternatives. This reuse is history-dependent: when networks experience different early tasks but identical later training inputs, each network preferentially reuses the scaffold induced by its own learning history. The same preference appears in individual optimizer updates, even though the network's dominant local response directions shift away from the original scaffold. Finally, constraining motion within the scaffold slows new-task acquisition more than matched random constraints, while effects on old-task retention are less consistent. In summary, these results suggest that early experience leaves a persistent geometric imprint on how neural networks adapt to future tasks.
cs.LG / 31 / 2609.36087
PHASE: A Physiology-Guided Hierarchical Foundation Model for Intracranial EEG
Abstract
Clinicians and neuroscientists have long analyzed intracranial electroencephalography (iEEG) through directly measurable physiological characteristics, which carry much of the information that downstream tasks depend on. Recent iEEG foundation models learn by reconstructing or predicting their inputs, which leaves the retention of these characteristics implicit. They are also evaluated mainly on cognitive decoding and a narrow clinical task, i.e., seizure detection. On a broad, clinically relevant benchmark such as Omni-iEEG, they remain below task-specific models when used frozen. We introduce PHASE, a physiology-guided foundation model that makes these characteristics explicit learning targets, pairing them with masked latent prediction in a temporal stage (PHASE-T) within each channel and a spatiotemporal stage (PHASE-ST) across synchronized channels. PHASE is pretrained on heterogeneous recordings from 222 participants at nine clinical sites. On all five Omni-iEEG clinical tasks, frozen PHASE-T outperforms every evaluated foundation model by up to 31\%, and fine-tuned PHASE-T surpasses the task-specific models, setting a new state of the art. PHASE-T benefits from physiological supervision, outperforming variants trained with latent prediction alone or auxiliary waveform reconstruction on every task in matched ablations. PHASE-T generalizes to unseen institutions, outperforming the compared models with few or no local labels. PHASE-ST further improves seizure-onset-zone identification over PHASE-T and, when frozen, decodes sound volume and pitch on BrainTreebank better than published models. Beyond task performance, PHASE learns to encapsulate the physiological characteristics clinicians recognize, from seizure onset and its propagation to anatomical region identity, even though its pretraining contains no ictal recordings or anatomical labels.
cs.LG / 32 / 2609.36108
LoopICL: Looping a single transformer block to solve tabular tasks
Abstract
Tabular foundation models using in-context learning have recently surpassed gradient-boosted trees on predictive tabular tasks. However, recent mechanistic insights suggest that parameters in these models are largely redundant. We introduce LoopICL, a looped transformer whose core design decouples parameter count from computational depth. LoopICL consists of a single block, processing data through two coupled streams: a cell stream capturing per-cell feature representations and a row stream capturing in-context example representations, jointly refined through within-column and cross-column attention. During pre-training, we vary loop counts, allowing the block to be unrolled for a varying number of iterations at test-time and use a learned exit-gate to automatically exit. In its standard setting, LoopICL performs competitively with TabICLv2 on TabArena and TALENT at the same computational cost (FLOPs), while using nearly 90% fewer parameters. Furthermore, its recurrent design enables users to also trade off inference cost and performance, providing a resource-aware TFM.
cs.LG / 33 / 2609.36126
Reasoning with Neural Cellular Automata
Abstract
Modern AI architectures used to solve visual reasoning tasks typically rely heavily on global connectivity and synchronization. As biological systems demonstrate, though, sophisticated computation can be performed in a more decentralized fashion. In this work, we test the reasoning capabilities of Neural Cellular Automata (NCAs), networks of recurrent cells that use strictly local connectivity and asynchronous updates. NCAs have been extensively studied in artificial life experiments, but it is unclear whether they can perform complex multi-step reasoning. We show that NCAs produce spatio-temporal dynamics capable of solving challenging visual reasoning tasks, including large mazes, Sudoku, and ARC-AGI-1. Furthermore, we provide evidence that NCAs generalize out-of-distribution when running with larger grids, longer rollouts, or parallel trials; and that the latter can be made more efficient via pruning of redundant trajectories. We find that these generalization capabilities depend on training with sample replay and stochastic perturbations, and that stochasticity remains beneficial at test time. Finally, we show that NCAs are robust reasoners capable of dynamically modulating compute to recover efficiently from damage, and that they can scale to solve reasoning in raw pixel space.
cs.LG / 34 / 2609.36144
Learning Continuous Patient Trajectories from Electronic Health Records
Abstract
Electronic health records provide irregular observations of latent patient states that evolve continuously over time. Recent autoregressive models condition on clinical histories to forecast future events as sequences of discrete observations. Conversely, multi-marginal flow matching provides a continuous-time formulation, but using multiple observations to supervise training paths does not itself give the learned dynamics access to preceding patient history. We introduce EHRFlow, a multi-marginal flow-matching framework that conditions on encoded patient history, thereby allowing future dynamics to depend on the patient's prior clinical trajectory. Our proposed framework accommodates irregular observation times and supports forecasting at arbitrary horizons. Across controlled synthetic benchmarks, EHRFlow improves clinical-code forecasting and latent-state recovery. On real-world clinical datasets comprising more than one million patients, including an independent external validation cohort, EHRFlow improves horizon-averaged top-5 clinical-code accuracy over autoregressive and history-independent flow-matching baselines. Finally, in a controlled counterfactual simulation, conditional guidance approximates the known effect of an antihypertensive intervention without training a task-specific outcome model.
cs.LG / 35 / 2609.36157
Encoder-Sharing Hierarchical Federated Multi-Task Learning for VANETs
Abstract
Most federated learning frameworks for vehicular ad hoc networks assume that all vehicles collaboratively train a single model for a common task. This assumption limits their applicability to practical vehicular environments, where vehicles may perform heterogeneous but related perception tasks with different output spaces. This paper proposes encoder-sharing hierarchical multi-task federated learning (EN-HMTFL), which integrates cluster-based hierarchical federated learning with a globally shared encoder and vehicle-local decoders. EN-HMTFL enables vehicles performing different tasks to collaboratively learn a transferable feature representation while preserving their task-specific models locally. Only the encoder is exchanged and aggregated through the hierarchy, whereas raw data and local decoder parameters remain at the vehicles. The proposed framework is evaluated on the MNIST and GTSRB datasets in different vehicular scenarios. Across the evaluated scenarios, EN-HMTFL improves accuracy by up to 24.0% relative to the compared representation-sharing benchmark. In scenarios where EN-HMTFL converges earlier, the reduction reaches up to 69 communication rounds (28.8%).
cs.LG / 36 / 2609.36173
Draft in Parallel, Condition Through Depth: Adjacent Causal Injection for Speculative Decoding
Abstract
Parallel speculative drafting generates multiple candidates in one backbone pass, but independent token selection can produce inconsistent continuations that shorten the accepted prefix. Existing methods mostly leave conditional decoding to a lightweight module after the backbone, which limits the flow of predecessor information to successors. Our analysis of DFlash shows that early positions already form recoverable predictions in shallow layers, and that accurate adjacent predecessors help successors more when they enter earlier. We therefore propose DSpine, a drafter with causal conditioning injection throughout the backbone: at every layer, gated adjacent injection writes each predecessor's predicted feature into its successor, so the causal conditioning chain unfolds over network depth while all positions update in parallel. A unified transfer space built from the target model's output embeddings unifies layer-wise injection with predecessor-conditioned decoding, and layer-wise output-embedding supervision promotes the formation of predicted features in shallow layers. Fused kernels and a transition cache execute both efficiently in parallel within SGLang. Across seven math, code, and chat benchmarks, DSpine achieves the longest acceptance length at both temperatures on Qwen3-4B and Qwen3-8B. At temperature zero on Qwen3-8B, it raises the seven-benchmark mean from DFlash's 3.77 to 4.82 (+27.8%); in SGLang serving tests, it delivers 23.3% higher throughput than DFlash on average.
cs.LG / 37 / 2609.36174
Fair Policy Optimization in Major-Minor Weakly Coupled Markov Decision Processes
Abstract
We consider fair resource allocation in sequential decision-making environments modeled as major-minor weakly coupled Markov decision processes (M2WCMDP). In this framework, resource constraints couple the action spaces of a major sub-Markov decision process (sub-MDP) and a population of minor sub-MDPs that would otherwise operate independently. Instead of using the traditional utilitarian (total-sum) objective, we optimize a general class of monotone, concave, permutation-invariant, normalized fairness functions. With homogeneous minor sub-MDPs, we prove that the problem under symmetry reduces to optimizing the platform-plus-mean-participant utilitarian objective over the class of \textit{permutation-invariant} policies, which allows us to exploit efficient algorithms that optimize the utilitarian-based objective to solve this fairness-aware problem. For more general settings, we introduce a count-proportion-based deep reinforcement learning approach with a priority-based sampler that generates feasible count actions. The generality of our framework means that the proposed algorithms and theoretical guarantees transfer to any domain with a symmetric M2WCMDP structure. We consider two applications: the machine replacement problem and the joint control of pricing and taxi relocation problem on a New York City-calibrated dataset. We validate our theoretical findings with comprehensive experiments, confirming the effectiveness of our proposed method in achieving strong fairness-aware performance while remaining scalable.
cs.LG / 38 / 2609.36193
Physical Cross-Modal Masked Autoencoding for Seismic-to-Well Representation Learning
Abstract
Learning from scientific measurements often requires aligning modalities with different spatial support and resolution. Subsurface characterization is an extreme dense-sparse problem in which 3D seismic provides volumetric but indirect measurements and well logs provide high-resolution 1D measurements at sparse borehole locations. We introduce a physically grounded cross-modal masked autoencoder (CM-MAE) for seismic-to-well representation learning. The model jointly tokenizes seismic volumes and well-log depth patches, embeds both modalities in continuous physical coordinates using four-axis rotary position embeddings, and reconstructs masked targets with a cross-modal decoder. Sparse Mixture-of-Experts layers provide modality-specific capacity while dense attention allows information exchange between modalities. Pretraining uses 23 seismic volumes covering approximately 178,000 square km across U.S. offshore and onshore basins, together with approximately 92,000 wells. Matched-mask ablations show strongly asymmetric information flow. Seismic context improves held-out well-log reconstruction by 9.73%, while well-log context improves seismic reconstruction by 0.65%. We evaluate seismic-only pseudo-log predictions against independent interpreter-drawn salt-geobody masks to determine whether they contain geologic signal. Across offshore U.S. surveys, compressional-slowness-derived salt scores reach an AUROC of 0.910. In onshore basins, predicted compressional slowness preserves formation-scale structure and partial relative organization in an unseen survey despite calibration drift. CM-MAE learns useful seismic-to-well representations under extreme modality asymmetry, although absolute pseudo-log calibration remains survey-dependent.
cs.LG / 39 / 2609.36208
Representable but Unlearned: Encoding Rank and the Interaction-Prediction Floor
Abstract
Input encodings can restrict which measured contrasts a predictor can jointly reproduce, even when no single contrast is forced to vanish. We compute the attainable contrast space from an encoder's equivalence classes and a fixed contrast design, without labels, loss, or a fitted model; projecting the recorded contrasts onto that space gives an empirical error floor for any unrestricted decoder on those classes. On a 140-rectangle siRNA interaction panel, a graph neural network's training-only feature mask merges 165 endpoint states into 90 classes and cuts the rank of the 140 interaction contrasts to 72. The resulting floor is 0.009980, which is 14.6% of the fitted model's interaction squared error; the fitted model reaches 0.068335, slightly worse than a control predicting no interaction at all. A minimum of three restored chemistry columns recovers full rank. Refitting without the mask removes the floor entirely, yet interaction MSE improves by only 0.000017 under the reported protocol, and the restored columns remain absent from every training input. On a released RNA-splicing predictor, whose encoding is injective on the measured states, the same computation returns the full design rank of 1,986 and a floor of exactly zero. These results separate what an encoding permits from what a fitted model achieves; they do not identify what limits the remaining error. The rank check needs no fits and bounds what any amount of training under a fixed encoding can recover. The project repository is available at https://github.com/shadi97kh/REPRESENTABLE-BUT-UNLEARNED.
cs.LG / 40 / 2609.36215
EnergyEminence: Source-Aware Environmental Calibration and Evaluation in a Physics-Grounded Grid Digital Twin
Abstract
Power-grid digital twins must combine data-driven prediction with physically meaningful state evolution while preserving the provenance of environmental observations. This paper presents an early-stage EnergyEminence testbed that couples an IEEE 118-bus-style graph-temporal predictor, nonlinear AC cascade simulation, and operator-dashboard-like temporal replay. In addition, we introduce a shared bounded calibration that converts wildfire-detection confidence and spatial extent into source-comparable wildfire interpretable and explainable evidence. We then evaluate it with visually diverse fire and hard-negative videos. Sixteen synthetic environmental videos are curated to generate 160 source-tracked grid scenarios, and a source-video-disjoint test yields 10 true positives, 8 false positives, 22 true negatives, and no false negatives. The errors occur in stressed, non-cascading scenarios conditioned on an unseen growing-fire source. Our diagnostic then reveals environmental shortcut learning that is obscured by scenario-level random splitting. The paper therefore contributes a data-centric and inspectable evaluation workflow for multimodal grid-resilience models, together with evidence supporting separation of environmental alerting from electrical cascade inference
cs.LG / 41 / 2609.36216
Preconditioned Physics-Informed Neural Operator Training
Abstract
Neural operators are typically trained in a supervised fashion, which requires a dataset to be generated with a classical solver. Training them physics-informed, i.e., purely from the governing equations, removes this large offline cost and allows fresh samples to be drawn at every optimization step, but has so far been limited to simplified problems and trails supervised training in accuracy. The obstacle is the ill-conditioning of physics-informed losses, which differential operators induce and which worsens as the discretization is refined. We therefore propose a preconditioned residual loss function and show mesh-independent conditioning for elliptic problems and greatly improved conditioning for saddle point problems. Realized through geometric and algebraic multigrid, the construction applies to linear and nonlinear equations, steady or time-dependent, on structured and unstructured meshes, is agnostic to the neural operator architecture, and adds no cost at inference. On the Poisson, Allen-Cahn and stationary Stokes equations, the resulting label-free training matches supervised training and is four to twenty-five times more accurate than previous physics-informed operator learning methods.
cs.LG / 42 / 2609.36221
How Language Models Differ in Redistributing Attention-Head Activity Under Serial Demand
Abstract
The way a model distributes activity over each layer's attention heads offers a coarse view of how it routes information through depth; how this changes with the task is part of what a mechanistic account must explain. Holding prompt length fixed, we vary how many serial steps a task demands and measure, in every layer of 17 open-weight models, whether activity concentrates on a few heads or spreads across many as demand rises. Both occur: in most models, layers just before mid-depth concentrate activity and later layers spread it. Models differ in where and how strongly this happens. The Qwen2.5 base models from 0.5B to 7B, for example, spread less than the average model in every task and concentrate activity in parts of their second half, where Llama models from 1B to 8B and OLMo-2 spread; the contrast largely holds between Llama-3.1-70B and Qwen2.5-72B, which have the same number of layers and heads. These differences are reproducible, and post-trained models keep much of their base model's pattern. An ablation study suggests that, within a task, models whose activity is more concentrated on their top heads also depend more on those heads for the answer. Concentration and spreading across layers thus offer a new way to compare models, by how they route information through depth. Code is available at https://github.com/johnnyjli/serial-demand-heads.
cs.LG / 43 / 2609.36230
The Role of Feed-Forward Layers in Transformer Dynamics
Abstract
We study the dynamical behavior of tokens in transformers from a control-theoretic perspective. Our model includes the feed-forward layer present after the self-attention mechanism, with the self-attention mechanism interpreted as an interacting particle system and the feed-forward layer as an independent control. Our main theoretical result establishes that the feed-forward network can steer the tokens arbitrarily close to consensus regardless of the key, query, and value matrices. Our result are easily extended to convergence to many clusters and to multi-head attention. We conduct numerical experiments to verify our results, and compare thresholding behavior from our theory to real-world LLMs.
cs.LG / 44 / 2609.36237
Transversal Pooling Neural Networks
Abstract
Many learning tasks require stability to small transformations while retaining sensitivity to larger ones. We introduce \emph{transversal pooling neural networks}, which generalize spatial max pooling to affine group actions. We establish equivariance to a chosen subgroup and derive explicit stability bounds for individual pooled wavelet coefficients under affine perturbations of the input. Experimentally, we demonstrate the utility of our networks in low-data environments and for predicting tropical cyclone intensification.
cs.LG / 45 / 2609.36240
Can Representation Learning Decouple from Loss Minimization? Polar Updates Have an Answer
Abstract
Does representation learning stop when the training loss stops improving? We study this question for matrix Muon, whose polar-normalised updates have a step length set by the gradient's rank rather than its norm. Near the edge of stability, full-batch Muon on teacher-student problems enters approximately period-2 loss oscillations that persist for thousands of steps: the cycle-mean loss stays flat or rises, yet the weights keep moving and the learned features continue to align with the teacher subspace. For linear teacher-student learning toys, we derive explicit cycle and alignment formulas and conditional plateau and decay bounds. For a population mean-field ReLU model, we prove that, under stated dimension, initialisation and small-head conditions, the leading eigenspace of the average gradient outer product (AGOP) recovers the teacher subspace exactly during a loss plateau, before the loss later drops. In all 33 ReLU, GELU and SiLU teacher configurations we study, direction-only alignment metrics show the student AGOP aligned with, or still aligning to, the teacher subspace during the period-2 oscillations; projected head refitting on selected configurations shows that the learned directions are useful for prediction, and further measurements distinguish AGOP alignment from weight-mass concentration. In deep residual ReLU students, freezing the downstream layers while the first layer trains with full-batch exact polar updates recreates a nearly flat cycle-mean loss with improving input-AGOP alignment; freezing and unfreezing switch between this plateau and loss decrease, and the effect is sensitive to momentum and to the choice of orthogonaliser.
cs.LG / 46 / 2609.36250
Action Chunking Proximal Policy Optimization with Feedback Correction
Abstract
Action chunking provides temporal abstraction in reinforcement learning by selecting short action sequences instead of individual actions, but many existing approaches face two limitations in high-dimensional robotic control. First, many rely on value functions over action chunks, which can be difficult to learn as action dimensionality and chunk length grow. Second, executing chunks open-loop removes within-chunk feedback, limiting reactivity in contact-rich tasks. We present Action Chunking PPO (ACPPO), a PPO extension that uses a chunked actor while retaining a standard state-value critic, thereby avoiding chunked Q-functions. We further propose ACPPO-Corr, which augments the chunk planner with a stepwise feedback corrector that adjusts planned actions online within each chunk. Across 25 simulated robotics tasks from IsaacGym and Bi-DexHands, spanning locomotion, arm manipulation, and dexterous hand-object interaction, ACPPO-Corr achieves the strongest aggregate performance among evaluated methods and performs best on both decision-frequency-sensitive and decision-frequency-neutral task subsets. Ablations show that moderate chunk lengths work best and that corrector regularization is important for balancing chunk-level planning with local feedback. These results suggest that action chunking can be effective in online PPO when chunk-level planning is paired with closed-loop correction. The code is available at: https://github.com/hshhahn/ACPPO.
cs.LG / 47 / 2609.36252
The Signed Geometry of One-Shot Recourse: On-Path Validity and the Signed-Curvature Criterion
Abstract
Closed-form recourse moves a rejected user along the unit gradient $\hat g$ of the classifier score $f$ by the promised distance $d_p=|f(x)|/\|\nabla f(x)\|$, at which the linearized score reaches zero. We ask when this one-shot step succeeds and what additional model queries change. To leading order the step ends on the favorable side exactly when the path curvature $κ=\hat g^\top\nabla^2 f(x)\,\hat g$ is nonnegative. Across 80 shallow models, the fraction of rejected users whose step ends there and the fraction with $κ\ge0$ correlate at $r=0.985$, although on Fashion-MNIST the first falls below the second by 8.2 points on average. No rule that uses only the score value and gradient can be valid for every score with path curvature bounded by $K$ without overshooting some by order $Kd_p^2/\|\nabla f(x)\|$. When the curvature is also Lipschitz and the step is short, one evaluation of $f$ at the promised point attains the minimax rate among deterministic one-query rules that know the curvature bound and its Lipschitz constant, and split-conformal calibration makes such a rule reach the first crossing or abstain with probability at least $1-δ$. Training with an asymmetric curvature penalty lets 99-100% of paths cross within the promised step on undershoot-prone shallow data, at about 4-22 times the overshoot of symmetric penalties (Fashion-MNIST, COMPAS). Because $κ$ and $d_p$ depend on how the score is scaled, part of this gain can be a longer promised step, and at matched validity a smaller audit of briefly trained models finds no uniform advantage over tuned inflation. Where a per-user line search along the ray is affordable, it is exact to grid resolution and preferable.
cs.LG / 48 / 2609.36259
CyFA: Linear Sequence Modeling with Relative-Time-Partitioned Memory
Abstract
Linear RNNs offer linear-time sequence processing and constant-memory decoding, but their fixed-size recurrent states must accommodate all past key--value associations. Existing forgetting mechanisms and Delta Rule updates reduce interference by selectively clearing or correcting the state, yet earlier associations can still become difficult to retrieve. We introduce CyFA (Cyclic Flow Attention), a Linear RNN with relative-time-partitioned memory. At each step, a learned clock controls the cyclic transport applied jointly to the key and value states before the current key--value pair enters the age-zero slot, thereby organizing stored associations across relative-time slots. We further derive an exact change to absolute-clock coordinates that expresses CyFA as two scalar-decay linear attention recurrences and enables efficient chunk-wise training. Across 400M--1.4B pretraining experiments with matched recurrent-state sizes, CyFA improves recall-intensive performance while maintaining competitive language modeling and high computational efficiency. At 400M, CyFA outperforms KDA on FDA (42.60 vs. 26.07) while requiring only 46.7% and 48.3% of KDA's forward and backward core-operator execution times, respectively. Our code is publicly available at \href{https://github.com/Chyxx/CyclicFlowAttention}{this https URL}.
cs.LG / 49 / 2609.36262
Understanding LLM Parameter Update Sparsity through the Lens of Fisher
Abstract
Recent studies have observed that parameter changes during language-model post-training can be concentrated in a small subset of coordinates. This phenomenon has been reported in reinforcement learning, on-policy distillation, and supervised fine-tuning on near-policy data. Its recurrence across different post-training paradigms suggests shared structure in training dynamics. In this paper, we examine this pattern through the diagonal model Fisher, which measures the sensitivity of the model's output distribution to individual parameters and is independent of any particular reward or teacher signal. Theoretically, we show that small diagonal Fisher leads to small expected gradients across a range of training objectives, providing a common explanation for sparse gradient updates. Empirically, we test this connection in RL and OPD. We find that Fisher identifies where gradients are concentrated, and fixed sparse masks selected from the initial Fisher retain a large proportion of the improvement from full training. Finally, we investigate the mechanisms underlying low Fisher in on-policy training. Our results show that high-probability next tokens tend to have similar parameter sensitivities, contributing to low Fisher. Together, these results establish the diagonal model Fisher as a unifying perspective linking update sparsity to on-policy training dynamics in LLM post-training.
cs.LG / 50 / 2609.36263
Paired Multimodal Scaling Laws
Abstract
Existing multimodal scaling laws fit multimodality terms empirically after testing and never vary how much data is multimodally paired at fixed data budgets. We investigate how, under the same total data per modality, changing the number of paired data affects loss curves in multimodal classification tasks. We train models in three different environments and run experiment sweeps varying data sizes and pairing budget. Pairing ratios have a dramatic impact on loss and this impact is directly tied to how much information synergy the task contains. Only paired data is able to reduce synergistic loss, while unpaired data can reduce redundant or unimodal information up until unimodal floors. Unlike traditional scaling laws where loss drops immediately in power law decay, synergy acquisition is gated, requiring a critical threshold of paired data before synergistic loss falls at all. We introduce a new family of multimodal scaling laws where total data-attributable loss is the sum of four individual power laws corresponding to the four different information channels of redundancy, a unique channel per modality, and synergy, and show how this law is both more theoretically sound and empirically valid across our experiments. This law predicts multimodal loss in our experiments more accurately than existing laws, with 3.2% error on fit tests versus 10.4% error for the best pairing extension of published laws.
cs.LG / 51 / 2609.36271
Stochastic Optimization Under Power-Law Spectra: Tight Bounds and Shuffling Analysis
Abstract
Recent work has established that power-law spectral conditions on data enable tight convergence bounds for deterministic gradient descent, resolving the conflict between classical exponential bounds and observed power-law learning curves. In this work, we extend this result to the stochastic regime of high-dimensional machine learning. We provide two main contributions: (1) We generalize the power-law spectral theory to Stochastic Gradient Descent (SGD), showing that the same spectral exponents govern stochastic dynamics; (2) For the fundamental case of isotropic Gaussian data, we provide a precise analysis of data shuffling, deriving exact constants that prove Single Shuffle is strictly superior to Flip-Flop and IID sampling. Our results bridge the gap between abstract spectral theory and practical stochastic training choices, offering a unified picture of how data geometry drives optimization speed.
cs.LG / 52 / 2609.36281
GNA: Granular Neighbor Assembly for Retrieval-Augmented Multivariate Time-Series Forecasting
Abstract
Deep forecasters predict from a fixed-length lookback window, and lengthening it gives diminishing returns at a growing cost. Retrieval augmentation instead shows the model how similar past situations continued. Retrieving a whole past window gives every variate the continuation of the same past moment. In multivariate series, however, the best past match differs from variate to variate. We present GNA (Granular Neighbor Assembly), a retrieval layer for forecasting backbones that assembles neighbors at two granularities: whole past windows, which keep the variates coherent, and per-variate neighbors, in which each variate takes its future from its own best-matching past. A learned gate decides, per forecast step and variate, how much to trust these futures against a persistence forecast, next to the backbone's own forecast. Candidates come from an embedding trained to predict each window's future, and retrieval is strictly causal: a past window is used only once its future has been observed. With the same lookback for every model and the same retrieval constants for all datasets, GNA improves two Transformer backbones in 85 of 96 dataset-horizon settings, gives the lowest MSE on 8 of 12 standard benchmarks and beats its backbone in every seed on 10 of them. Both granularities are needed, and neighbors of mismatched queries are worse than none. Retrieval helps most where the lookback says least: the gate shifts trust to retrieved futures further ahead. Where it fails, on hourly non-stationary series at long horizons, the loss is consistent with a drifting level of the retrieved futures.
cs.LG / 53 / 2609.36288
Representational and Functional Robustness to Electrode Montages in EEG Foundation Models
Abstract
EEG foundation models (EEG-FMs) are intended to generalize across different datasets by learning representations that, ideally, are invariant to dataset-specific EEG configurations such as electrode montages. However, EEG-FMs that accept different montages as input do not guarantee that representations and predictions remain stable across different electrode configurations, especially outside the training setting. In this work, we investigate the effects of different electrode montages through a joint functional and representational analysis of four EEG foundation models selected to span distinct montage-handling designs. We evaluate embeddings on cross-subject resting-state eyes-open/closed and within-subject motor-imagery classification under spatially informed channel reduction. Functional robustness is tested through the generalizability of linear probes across channel counts, while representational robustness is assessed through within-subject similarity and preservation of between-subject geometry. The four models show distinct robustness profiles, and the two axes dissociate: large changes in embedding similarity need not come with comparable probe degradation, and stable embeddings can still lose downstream performance. Comparing two readouts of the same encoder further shows that aggregation, not the encoder alone, determines functional robustness: pooling into anatomically aligned regions degrades less than a learned global readout, despite being montage-invariant by construction. Montage robustness is therefore a joint property of the encoder and its aggregation, and characterizing it requires both a representational and a functional axis. Input compatibility alone is evidence for neither.
cs.LG / 54 / 2609.36301
MoRE: Scaling mixture of experts with hardware-aware low-rank routing
Abstract
Mixture-of-Experts (MoE) layers are central to frontier language models, and recent architectures push toward more and smaller experts. In this regime, the standard linear router becomes a bottleneck: with $M$ experts and hidden dimension $h$, its per-token cost $Θ(Mh)$ dominates the MoE layer once $M$ is large. We introduce MoRE (Mixture of Rank-reduced-routed Experts), which factorizes the router weight matrix at rank $r$ and reduces the routing cost to $O((h + M)r)$. We prove that rank logarithmic in $M$ suffices for routing expressivity when the number of active experts is fixed, and is necessary up to precision factors. We also prove that logarithmic rank preserves load balance in a Gaussian memorization model, and training on a synthetic phonebook task shows that low rank does not hurt memorization. At matched active FLOPs, the factorization allows a factor of $Θ(h/r)$ more experts. To realize this gain in wall-clock time, we design a fused Triton kernel at inference that avoids expensive memory operations on HBM. Empirically, MoRE improves memorization on the phonebook task and performance on knowledge-intensive Q\&A benchmarks after pretraining, while matching reasoning ability. Code available at https://github.com/Matheart/MoRE_code.
cs.LG / 55 / 2609.36302
The Universal Classifier for Graph Learning
Abstract
While foundation models have revolutionized natural language processing and computer vision by leveraging universal vocabularies, Graph Machine Learning (GML) remains fractured due to the absence of a unified feature and structural representation across diverse domains. Existing works claiming to be Graph Foundation Models (GFMs) are typically restricted to node-level predictions or require fixed feature dimensions, failing to provide a truly task-agnostic backbone for the full spectrum of graph learning applications. In this paper, we introduce the Universal Classifier (UC), which supports arbitrary feature and class cardinalities, unifying node-, edge-, and graph-level objectives under a single similarity-based classification objective. The UC reformulates all node-, edge-, and graph-level prediction tasks as maximizing similarity in the latent space: by lifting heterogeneous features and labels into 3D latent tensors, the model learns transferable features independent of specific input schemas. This architecture allows a single pre-trained model to generalize to node classification, node regression, and link prediction across unseen graphs with varying feature semantics. Experiments show strong zero-shot transfer performance across node-, link-, and graph-level tasks.
cs.LG / 56 / 2609.36307
Cheap and Powerful Tests for Supervised Subspaces: Per-Component Inference for PLS
Abstract
Partial Least Squares (PLS) regression extracts a few outcome-aligned directions in a high-dimensional X and is widely used across applied science, but inference on the resulting fit is either expensive, biased and discouraged, or absent. We reduce inference to held-out OLS refits of the supervised subspace, a primitive shared by PLS, supervised PCA, and linear probes, and supply two tests using held-out correlations: a Nadeau-Bengio corrected asymptotic t-test as a fast approximation, and a permutation test with comparable power, finite-sample valid under outcome-predictor independence and iid rows. Held-out predictions are unchanged under any orthogonal rebasing of the supervised span, so an interpretable basis such as varimax inherits the joint claim but not a per-axis p-value; per-component claims come from a fixed-sequence test on the PLS extraction order. We validate on synthetic geometries, two NIR chemometric datasets, and cross-lingual word-embedding regressions; the exact test also transfers to supervised PCA and a ridge probe. The proposed tests have more power than CV-permutation-Q^2, at a fraction of its cost. A pre-run check on n and the spectrum of X says when the approximation is safe. We release a Rust library with Python, R, and Julia bindings, plus a Python text pipeline.
cs.LG / 57 / 2609.36310
Learning Samples Importance: Parameterizing Dual Variables in Everywhere Learning
Abstract
Everywhere learning provides a principled framework for training AI models under constraints that must hold throughout the data distribution. In the dual domain, these pointwise constraints give rise to functional dual variables. In this work, we propose to learn these dual variables, motivated by the fact that their values encode useful information about the underlying constrained problem. By representing the dual variable as a parametric function of each sample, we enable the learned multiplier to be evaluated on new, unseen samples. This contrasts with standard empirical dual formulations, which assign an independent multiplier to each training sample. We characterize the error in the recovered primal solution induced by restricting the dual variable to a parametric function class and show that it is controlled by how well this class approximates the optimal statistical multiplier. Moreover, we show that the learned parametric multiplier retains the sensitivity interpretation of the optimal statistical multiplier, yielding approximate sensitivity guarantees that extend beyond the samples used for training. We empirically validate our theory across a variety of everywhere learning tasks, showing that the resulting constrained problems can be solved efficiently and that the learned dual variables provide meaningful representations of sample-level sensitivity.
cs.LG / 58 / 2609.36314
Fractional State Space Transition for Long Sequence Modeling
Abstract
State Space Models (SSMs) compress sequence history into a bounded recurrent state, making the resulting memory law a central architectural choice for long-context performance. Most modern SSMs rely on ODE-based dynamics that lead to exponential forgetting, limiting their ability to retain information over broad temporal ranges. We introduce FRAC, a selective SSM architecture derived from fractional dynamics that replaces this exponential decay with power-law long memory. To make fractional dynamics practical, FRAC approximates the heavy-tailed target kernel with a finite-state, log-spaced sum of exponential modes. This construction turns fractional memory into an efficient recurrent module with parallel training and prefill, while retaining bounded-state autoregressive decoding. Extensive experiments, including 1.3B-parameter language modeling, demonstrate that FRAC consistently improves long-context performance over state-of-the-art SSM baselines while staying competitive on short-context. These results show that fractional dynamics provide a practical and effective prior for long-context SSMs.
cs.LG / 59 / 2609.36315
PyroStack: A Multi-Band Spatio-Temporal Sub-Daily Dataset for Wildfires in the United States
Abstract
Wildfires are an increasing hazard to ecosystems, air quality, and human systems, creating a growing need for datasets that support systematic development and evaluation of models for predicting fire spread across diverse landscapes. Effective prediction requires integrating meteorological conditions, fuels, vegetation, and topography at spatial and temporal resolutions suitable for both physical simulation and data-driven approaches. However, existing datasets often lack the resolution and coverage needed to capture these interacting controls. The PyroStack dataset addresses this gap by providing a harmonized, event-based collection of wildfire and environmental data across the contiguous United States and Alaska. It integrates satellite-derived fire observations with atmospheric reanalysis, vegetation, fuel characteristics, and topographic information into a unified framework spanning 6994 wildfires that occurred between 2012 and 2024 across a wide range of ecosystems and climate conditions. PyroStack offers spatial resolutions ranging from 30 m to 9 km and hourly temporal resolution, along with fire progression data at 12-hour intervals to support model initialization and evaluation. By combining broad spatial coverage with fine spatial and temporal detail, the dataset enables systematic analysis of wildfire dynamics and supports both physics-based and machine learning approaches, providing a foundation for benchmarking and improving fire spread models, with future extensions aimed at incorporating additional regions and fire suppression data streams to further advance wildfire prediction.
cs.LG / 60 / 2609.36322
Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression
Abstract
Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token's phase, or its position relative to compression-window boundaries. We uncover a systematic asymmetry in models using such compression: the same information can be easy to retrieve at one phase and difficult at another. We call this periodic variation in retrieval performance phase sensitivity. In large open-weight models with such compression, long-context retrieval accuracy can differ by up to 40 percentage points across phases, revealing periodic weak spots that average benchmark scores can conceal. To investigate this behavior, we pretrain a family of transformers from scratch across multiple KV-compression designs, reproducing phase sensitivity across the variants. Mechanistic analysis using causal interventions in these models reveals phase specialization: different attention components contribute asymmetrically to retrieving information at different source phases. We further analyze idealized retrieval models, showing how gradient flow dynamics may favor sharp phase specialization. Evaluating models with chunked KV-cache compression thus requires measuring across compression phases: high average accuracy can coexist with systematic positional failures.
cs.LG / 61 / 2609.36329
Reducing the Adaptation Gap Through Reachable Fisher Geometry
Abstract
Parameter-efficient fine-tuning (PEFT) determines not only how many parameters are trained, but also which local directions a model can move in, so similar adapters can affect subgroup losses differently. Since curvature matrices are infeasible to form at adapter scale, scalar summaries such as the Fisher trace are often used instead. We study what the trace reveals and what it loses through the reachable Fisher: each subgroup's full-model Fisher pulled back through the adapter Jacobian. Under likelihood losses, it represents the Gauss-Newton curvature accessible to the adapter, and its trace can be computed from score-gradient norms without forming the full matrix. Under matched subgroup gradients, a positive-definite reachable-Fisher difference, with a margin exceeding the Hessian-Fisher defect, implies that every sufficiently small nonzero model-changing update increases the signed gap. In contrast, the restricted operator norm determines worst-case quadratic change, while a matrix-free Frobenius discrepancy bounds its reachable-Fisher component. Trace alone cannot certify definiteness or control matrix mismatch. Equal traces rule out a positive-definite difference but can still hide large operator discrepancies. Across 306 single-seed models, higher trace accompanies greater subgroup difficulty in 75.7 percent of 1,218 eligible evaluations, while trace matching reduces the best-worst subgroup gap in all 30 dataset-encoder-adapter combinations. However, held-out audits show that operator discrepancy decreases in 23 of 30 combinations, while the unbiased squared-Frobenius statistic decreases in only 16 of 30. Fisher trace is therefore a scalable diagnostic and training heuristic, but not a certificate of local gap behavior or matrix alignment.
cs.LG / 62 / 2609.36337
Adapting Linear-Time Architectures for Tabular In-Context Learning
Abstract
Tabular foundation models achieve strong performance by conditioning on labelled examples in context, but softmax attention limits their use on large datasets. Existing linear-time alternatives, however, are mostly causal, and their potential for tabular in-context learning (ICL) remains underexplored. To address this, we (1) revisit causal training setups, (2) compare linear sequence mixers, and (3) investigate their ICL generalisation beyond the pretraining context length. First, we show that the best training setup for causal models resembles next-token prediction. Then, perhaps surprisingly, the most promising linear sequence mixer is causal: DeltaNet outperforms even non-causal linear attention. However, it degrades beyond $2$-$4\times$ the pretraining context length, and existing mitigation strategies such as bidirectionality defer the problem at best. A hidden-state oracle shows that this is not a capacity problem. Instead, our analysis points to an instability in the recurrent state, which drifts in deeper layers of causal models. Since DeltaNet's learned write rates overfit to the pretraining regime, we modulate them with a time-dependent decay schedule intervention to stabilise length generalisation. Finally, re-introducing non-causality by reading out from the final state allows us to closely match a controlled softmax attention baseline on OpenML-CC18 and TabArena.
cs.LG / 63 / 2609.36354
Explainability from Training with Applications to TCR-Epitope Prediction
Abstract
Deep learning models have achieved strong performance in artificial intelligence for science, yet their black-box nature limits our understanding of how they learn scientific tasks. Existing methods for interpretability provide limited insight into how models organize evidence and evolve during learning. We introduce explainability from training (EFT), a model-agnostic paradigm that traces model interpretation during training to explain why models rely on specific features and how they organize these features as predictive evidence. We apply EFT to four state-of-the-art T cell receptor (TCR)-epitope prediction models, TCR-SRIM, TULIP, MixTCRpred, and NetTCR-2.2, spanning post-hoc and interpret-by-design approaches as well as transformers and CNNs. To investigate how structural information affects model explanations, we introduce a benchmark, TCR-XAI2, containing 388 unique experimentally resolved TCR-epitope structures, complemented by structures predicted using AlphaFold3, Boltz-2, TCRModel2, tFold-TCR, and OpenFold3. Using EFT with TCR-XAI2, we demonstrate that (1) CNN and transformer models exhibit distinct learning trajectories; (2) TCR $α$ and $β$ evidence can conflict during learning, limiting the benefits of jointly modeling both chains, while MHC information mitigates this; and (3) real versus predicted structural data for TCR-epitope prediction exhibits distinct TCR and peptide feature preferences as well as differing trajectories of model certainty.
cs.LG / 64 / 2609.36368
AdaKerNet: Neural Kernel Decoding for Task-Adaptive Prediction with Multimodal Large Models
Abstract
Large foundation models have been introduced with the promise of efficient adaptation to downstream tasks. Yet, under limited supervision, MLLMs, an important class of large foundation models, remain challenging to adapt to various downstream tasks. Adaptation typically relies either on MLLM parameter fine-tuning or on training neural-based decoders. Both approaches struggle under limited supervision, while fine-tuning additionally requires access to model parameters, which is often unavailable for closed-source models. We introduce AdaKerNet, a novel learnable task-adaptive neural kernel decoder. AdaKerNet is fully agnostic to the parameters of the underlying MLLM and operates solely on its (frozen) rich representations obtained from the diverse available modalities. AdaKerNet relies on (i) a set of learnable, Lipschitz-controlled multimodal features derived from these MLLM representations; (ii) a reference kernel that provides a soft structural prior on those features; and (iii) a lightweight nonlinear neural predictor that adaptively deforms that structure. Learning the kernel representation and the neural predictor jointly within a unified optimization framework allows AdaKerNet to capture features and geometric relationships relevant to the downstream task. Numerical tests across four MLLMs: BLIP-2, LLaVA-1.5, Qwen2.5-VL, and Gemini Embedding 2, and multimodal inputs spanning text, audio, images, and tabular measurements demonstrate significant and consistent improvements over direct MLP, attention-, autoencoder- and kernel-based decoders, across a range of scarce-label budgets, with average error reduction of up to 41% across baselines. These results establish AdaKerNet as an effective approach for prediction from frozen multimodal representations in the scarce label regime. Additional structural ablations highlight the complementary contributions of AdaKerNet's components.
cs.LG / 65 / 2609.36375
Neural Succession: A Mesoscopic Theory of Invasion, Coexistence, and Stabilization in Continual Learning
Abstract
Continual learning is usually studied through mechanisms that preserve old knowledge. We develop Successional Learning Theory (SLT), a mesoscopic account in which the current representation is a resident community, the incoming task is an invader, forgetting is resident displacement, joint retention is coexistence, replay is resident reinforcement, and training moves from establishment toward stabilization. Its empirical coordinate is directional pre-invasion compatibility, measured on the resident model before the incoming task is learned. Across eight experiments, compatibility orders later forgetting on the 20 directed Split-CIFAR-10 transitions (three-repeat r=-0.789, incoming-task cluster 95% CI [-0.90,-0.72], every repeat alone r<=-0.67), forecasts held-out forgetting with 24% lower error than a no-information baseline, and reproduces under controlled MNIST permutations and CIFAR-10 rotations (r=-0.804, -0.718). On an 84-transition suite, compatibility separates coexistence from exclusion at every retention threshold (AUC 0.93-0.97). Replay repairs every transition with at most 325 stored examples and is most efficient where displacement is largest. Compatibility reaches |r|=0.720, while activation, representation, Jacobian, and fixed-coefficient Lotka-Volterra specializations do not. Plasticity and feature turnover fall reliably from early to late training (15/15 and 14/15 runs). We formalize a minimum habitat-modification bound, a displacement floor, a sufficient coexistence condition, an identifiability law with a range-restriction corollary, successional stabilization, and local reinforcement. The identifiability law also predicts where the coordinate loses leverage, and the prediction matches three CIFAR-100 partitions and five optimizer regimes. SLT is a pre-adaptation diagnostic that complements replay, regularization, and projection methods.
cs.LG / 66 / 2609.36393
Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents
Abstract
Traditional reinforcement learning (RL) techniques focus on maximizing expected cumulative reward, where each action assumes to take a constant unit of time. However, this assumption does not hold for agentic RL tasks such as machine learning engineering (MLE) agents, where actions involve data loading, feature engineering, and model training that take variable durations. Efficiency matters in modern agentic RL where actions are costly. To address this limitation, we adapt from continuous-time RL and Semi-Markov Decision Process (SMDP) formulation and propose Reward-rate Policy Gradient (RPG), where we focus on optimizing the reward rate -- the long-term reward per unit of time. RPG estimates the reward rate from off-policy samples, then charges each action for the time it consumes at that rate. We first conduct theoretical analysis in the bandit setting to establish that RPG approximates the optimal reward rate and empirically demonstrate it outperforms baselines while avoiding enumeration over the policy space, a known issue for an existing method. We then further apply RPG on a small language model (Qwen3.5-4B) with self-improvement loops and empirically show it obtains higher rewards within a fixed time budget than vanilla RL on MLE-Bench and NanoGPT, with a 19.2% and 85.7% margin, respectively. Our method provides a practical solution for optimizing performance under wait time considerations in modern agentic RL tasks, where actions interact with external environments and cost time.
cs.LG / 67 / 2609.36409
Longer Records, Broader Invariance: The Hidden Scaling Problem in Longitudinal Contrastive Learning
Abstract
Longitudinal data are valuable because people change. Yet the objectives used to learn from these data can inadvertently erase that change. In person-level contrastive learning, observations from the same person are treated as positives; as records grow, those positives can span increasingly distant---and increasingly different---behavioral states. More history can therefore produce not only more data, but broader invariance. We show that this distinction is fundamental. We separate \emph{record span}, how much history the learner sees, from \emph{supervision span}, how far across that history positive-pair supervision reaches. Across in-home sensing records spanning up to 2.7 years, broader supervision systematically suppresses recoverable changing-state information, even when the available history is held fixed. At the broadest span, less than 10\% of the information recoverable from an untrained encoder remains. Yet keeping positives local is not sufficient: as records grow, even distant states that are never paired become increasingly similar. Explicitly contrasting other observations from the same person reverses this loss without shortening the record, revealing a second route by which longitudinal scale can broaden invariance. Finally, we prospectively reproduce the supervision-span effect in 199 GLOBEM participants. Longitudinal scale therefore presents a choice: more history need not mean more invariance. By controlling what is held invariant as records grow, we can preserve the change that made the longitudinal data valuable in the first place.
cs.LG / 68 / 2609.36448
Theory on Attention Dynamics for Out-of-Distribution In-Context Learning
Abstract
Transformers have demonstrated remarkable in-context learning (ICL) capabilities, enabling them to perform new tasks without additional fine-tuning. However, their performance often deteriorates when encountering out-of-distribution (OOD) inputs that deviate from the training distribution, and the underlying theory remains poorly understood. To fill this gap, we characterize the OOD error under the input distribution shift through the interplay between the dynamics of the so-called $α$-type and $β$-type attention weights, which represent the transformer's confidence in identifying the correct and incorrect features, respectively. Our results indicate that the OOD error for each feature depends on all pairwise interactions between the training features and OOD features, and under certain cases the transformer performs no better than random guessing. To improve the OOD generalization performance, we next investigate the impact of model finetuning with the OOD data, and particularly, characterize the model forgetting performance on the source domain. Interestingly, the performance on the source domain may not always degrade after finetuning, which highly depends on the nature of the feature shift: finetuning on OOD domain keeps enhancing the confidence of identifying correct features from the original distribution, while the interference from other incorrect features may either increase or decrease. Extensive experiments on both synthetic and real data are conducted to corroborate the theoretical insights.
cs.LG / 69 / 2609.36453
Channel-Dependent State Space Model for Multivariate Time Series Forecasting
Abstract
Multivariate time series forecasting (MTSF) is critical across many real-world domains. Existing deep learning approaches fall into two paradigms with distinct limitations: channel-independent (CI) methods unconditionally ignore cross-variable dependencies and model only temporal dynamics, while channel-dependent (CD) methods consider both but typically rely on architectural compromises to mitigate overfitting and computational overhead. We therefore propose Chameleon, a specialized CD state space model (SSM) that enables data-dependent, fine-grained interactions across variables while scaling linearly with their number. By connecting selective SSMs with the Kalman filter, we leverage the missing measurement update in the former for cross-variable modeling while preserving the SSM backbone for robust temporal modeling. We further identify favorable inductive biases of GatedDeltaNet for time series, adapt it as our backbone, and improve generalization through additional techniques, including a previously unexplored stochastic perturbation of reversible instance normalization. On strongly dependent ODE and PEMS datasets, Chameleon achieves the best MSE and MAE across all settings, while its CI ablation and prior CD methods incur 61-178% higher MSE on average. Across 28 standard benchmark settings, Chameleon also achieves better MSE and MAE than each baseline in at least 27 and 22 cases, respectively. Training-time and peak-memory analyses on Traffic and ETT further demonstrate competitive efficiency and favorable memory scalability across different variable counts.
cs.LG / 70 / 2609.36458
Fisher-IRG: Fisher-Induced Local Invariant Representation Geometry across Language and Vision Models
Abstract
Semantic-preserving transformations can induce substantial motion in learned representations, while small changes may strongly affect model predictions, raising a basic question: what local metric best captures semantically consequential variation? We propose Fisher-induced invariant representation geometry (Fisher-IRG), which measures local representation directions through their predictive sensitivity. Around each representation, we construct semantic-preserving and semantic-changing neighborhoods, aggregate their local Fisher information, and recover invariant directions through a contrastive generalized eigenvalue problem. Controlled displacement analyses first show that comparable Euclidean motion can have substantially different predictive consequences, supporting the need for a predictive geometry. Across language and vision models, Fisher-IRG yields stronger semantic-versus-nuisance predictive selectivity and generally more reproducible subspaces than covariance-based geometry, while recovering systematically distinct local directions. Representation interventions further localize semantic effects to the Fisher-derived subspace, and held-out separation and retrieval show that the recovered geometry generalizes beyond the discovery neighborhoods. These results support Fisher-IRG as a principled framework for characterizing local invariant representation geometry.
cs.LG / 71 / 2609.36472
DisCoMBO: Steering Expert-in-the-Loop Black Box Optimization via Distributional Conformance
Abstract
Sequential Model-Based Optimization (SMBO) traditionally relies on Bayesian or ensembling surrogates for uncertainty quantification. While historically treated as fully data-driven, SMBO increasingly integrates external domain expertise to accelerate discovery. To overcome the opaque guidance and diminished integration fidelity of standard acquisition re-weighting, Probabilistic Circuits (PCs) have emerged as a generative surrogate alternative, enabling direct knowledge injection via conditional sampling. However, these generative routines lack the formal exploration-exploitation semantics required for rigorous optimization. We introduce the Distributional Conformance Score (DisCo), a novel metric that unifies the flexibility and efficiency of PCs with a formal uncertainty framework. DisCo provides a bounded, $[0, 1]$-normalized measure of model "surprise" that (1) recovers properties comparable to kernel-based uncertainty known from, e.g., Gaussian Processes, while maintaining linear-time inference, and (2) enables accurate assessment of conformance of external knowledge w.r.t. model evidence. We then present DisCoMBO, a framework leveraging these properties for robust, knowledge-aware optimization. We prove that DisCoMBO is a zero-regret algorithm and demonstrate its effectiveness across diverse benchmarks from AutoML, material optimization, and wind park optimization.
cs.LG / 72 / 2609.36477
Guard Models Are Overconfident Where Base Models Are Uncertain
Abstract
Guard models are used as safety classifiers, with confidence scores driving downstream moderation decisions. We evaluate five guard models for prompt classification and find that although several are nearly calibrated on clean inputs, adversarial attacks degrade their calibration by an order of magnitude, turning false negatives into high-confidence errors indistinguishable from correct detections. Comparing each guard with its corresponding base LM, we find that uncertainty signals often remain available, with the base model typically expressing uncertainty on the same inputs where the guard fails. Layer-wise analyses localize this guard-base divergence to later layers, where guard models exhibit sharper safe/unsafe separation and lower-rank representations, while adversarial harmful inputs lie closer to the clean-safe region. These findings highlight a mismatch between guard confidence and base model uncertainty under attack.
cs.LG / 73 / 2609.36484
The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation
Abstract
On-policy distillation (OPD) trains a student to match the teacher's next-token distributions on the student's own trajectories and has yielded substantial empirical gains. Generalized variants allow the student to surpass the teacher by extrapolating an implicit reward in output space. The language-model head, however, attenuates this change anisotropically: much of the change encoded in the teacher's hidden states reaches the logits at a small fraction of its weight, and the sampled-token log-probability ratios on which output-space extrapolation relies inject noise that the extrapolation amplifies, making training unstable. We observe that reinforcement learning (RL) shifts a model's internal representations relative to its base checkpoint, and that the direction of this shift can be measured at every layer. Motivated by this observation, we propose RIDE (RL-Induced Direction Extrapolation), which extrapolates the RL-induced change directly in representation space: at every layer and token position, RIDE computes the residual between the teacher and its pre-RL checkpoint and regresses the student's hidden states toward targets displaced beyond the teacher along this residual. Conditioned on a sampled trajectory, this regression is equivalent to maximizing a linear directional reward defined by the residual under a quadratic penalty centered at the teacher, which makes explicit how the objective moves the student along the RL-induced direction while limiting its deviation from the teacher. Across four base/RL-teacher pairs spanning different scales, architectures, and pre-training lineages, RIDE approaches or exceeds the RL-trained teacher on every pair and is the only method whose mean does so, and it consistently outperforms output-space extrapolation, which degrades the student whenever the teacher is close to its base. Project page: https://github.com/xixixixixxxx/RIDE.
cs.LG / 74 / 2609.36486
Optimal Multi-Reward Reinforcement Learning
Abstract
We study an unknown-transition finite-horizon Markov decision process (MDP) with a finite collection of known reward functions $\{r^1, r^2, \ldots, r^M\}$. The goal is to output an $ε$-optimal policy for every reward using online episodic interaction only. Performance is measured by the policy error $V_{0}^{*, m} - V_{0}^{\widehatπ^{m}, m}$ where $m\in [M]$ represents the reward function and $V_{0}^{*, m}=\mathbb{E}_{s_1\sim μ}[V_{1}^{*, m}(s_1)]$. Under this setting, we design a provably efficient algorithm to establish a minimax sample complexity bound of $$ O\left(\frac{SAH^3}{ε^2}\log M \mathrm{polylog}\left(\frac{SAH\log M}{\min\left\{ε, 1\right\}δ}\right)\right)$$ episodes, with no additional burn-in cost. This matches the information-theoretic lower bound up to a factor of $ \mathrm{polylog}(SAH\log M/(\min\left\{ε, 1\right\}δ))$. Our method combines three technical ingredients. First, we adapt MVP to reward-switching learning to construct optimistic value estimates. Second, we use fresh replay samples to conservatively evaluate the candidate policies. Third, gap-based multiplicative weights updates adjust the reward-sampling distribution using the differences between these estimates, converting weighted learning progress into simultaneous guarantees for all rewards.
cs.LG / 75 / 2609.36490
LLMs Learn to Evade Latent Monitors from Prior Feedback Alone
Abstract
Latent space monitors aim to detect undesired behaviors in LLM agents by inspecting an agent's internal activations rather than its outputs. However, interactive monitoring creates a feedback channel where each verdict the monitor delivers leaks information to the model about how its internal states are being evaluated. We ask whether an agent can infer the monitor's decision rule from this feedback and then selectively edit its activations to evade detection. Unlike prior evasion attacks, the model is never explicitly told what the monitor detects. Surprisingly, off-the-shelf models already produce activation edits aligned with the monitored direction, but at insufficient magnitude for evasion. Simply scaling up these edits by a factor of 8 reduces the monitor's TPR from 100% to 27%. A rank-1 LoRA amplifies this behavior into effective evasion within the forward pass, reducing TPR further to 4% on held-out concept monitors while leaving other concepts at their normal detection rates. Capabilities on standard benchmarks are retained under this finetuning, and the evasion skill survives retraining the monitors on the new activations. Mechanistically, we find evidence that the model computes its activation edit from the prior in-context turns, and show that the edit becomes more aligned with the monitored direction as more examples are provided. These results demonstrate feedback-conditioned control over activations and suggest that latent monitoring should be treated as an interactive process in which agents can observe and respond to oversight measures.
cs.LG / 76 / 2609.36504
Efficient and Scalable Physics-Guided Fully Convolutional Spatiotemporal Learning for 3D Microstructure Evolution Prediction
Abstract
Accurate prediction of three-dimensional (3D) microstructure evolution remains computationally demanding because high-fidelity phase-field simulations require repeated numerical integration over large volumetric domains and long temporal horizons. This study develops an efficient and scalable physics-guided fully convolutional spatiotemporal framework for direct multi-frame prediction of complete 3D microstructure sequences. The model combines shared 3D spatial encoding and decoding with a factorized latent translator that integrates temporal, local 3D spatial, and channel interactions. A discrete Cahn--Hilliard (CH) residual is incorporated during training to regularize the learned evolution toward the governing dynamics without altering the inference pathway. The framework is evaluated on high-resolution 3D spinodal-decomposition trajectories under nominal, long-horizon, and reduced-temporal-context forecasting. Under full temporal context, the model accurately reproduces volumetric evolution, with average 3D structural similarity remaining above 0.97 over the nominal prediction horizon. Physics guidance becomes increasingly beneficial as temporal information is reduced, improving predictive robustness and preservation of interface-level morphology. The framework also achieves more than a 30-fold wall-clock speedup relative to the reference spectral phase-field solver, while physics guidance introduces no additional inference cost. These results establish direct multi-frame, physics-guided fully convolutional learning as a high-throughput surrogate strategy for dense 3D phase-field dynamics and repeated microstructure forecasting.
cs.LG / 77 / 2609.36521
PDE-OBS: Controlled Evaluation Across Observation Patterns
Abstract
Physical-field reconstruction and forecasting depend on both measurement density and spatial layout, yet evaluation under a single observation pattern does not characterize performance when that pattern changes. We introduce PDE-OBS, an integrated benchmarking platform spanning numerical data generation, model training, and inference and evaluation under varying observation conditions. It combines 560,000 fields and trajectories from seven partial differential equation families with configurable observation operators and seven adapted baseline methods for stationary reconstruction and short-horizon forecasting. Separating observation construction from physical records allows users to specify parameterized patterns and deterministic mixtures for training and testing while preserving prediction targets and data splits. The evaluation protocol uses references trained for each test pattern to compare models on identical test observations and targets, alongside equal-count groups for spatial-layout comparisons. On a 14,000-record subset, we evaluate 441 trained models under nine test patterns, yielding 3,969 evaluations. Mean cross-pattern error exceeds mean matched-pattern error in all 49 PDE-method pairs, and this finding persists in a configuration-matched subset of 117 models. Denser test observations do not consistently reduce error for a fixed model. Mixed-pattern training on five completed pairs reduces large single-pattern transfer errors, although destination-trained references usually remain more accurate. Together, the benchmark and findings support systematic evaluation of observation-pattern sensitivity and provide a reusable workflow for developing methods under changing measurement conditions. Code: https://github.com/ru1ch3n/PDE-OBS.
cs.LG / 78 / 2609.36526
Adapting Context Compression for Long-Horizon Agents with Counterfactual Continuations
Abstract
Long-horizon agents require context compression to manage growing interaction histories. Compression quality, however, is ultimately determined by downstream execution. Existing prompt-adaptation methods infer compression errors by comparing full-context and compressed trajectories. Such comparisons cannot isolate individual compressions and are confounded by agent stochasticity. We first find that compression degrades reliability before solvability. Using matched counterfactual continuations that compare execution from the same agent state with versus without compression, we further show that severe degradation concentrates at isolated compression events. Motivated by this finding, we propose PAIR (Prompt Adaptation using Interventional Rollouts) for adapting structured compression prompts. PAIR identifies individual compressions that degrade subsequent execution, diagnoses their effects, and revises the relevant sections of a fixed compression template. PAIR achieves the strongest cross-run reliability among compressed methods in every main benchmark-scope combination, consistently exceeding the competing prompt-adaptation baseline. Without modifying the downstream agent, PAIR brings compressed execution close to the no-compression baseline and sometimes numerically exceeds it.
cs.LG / 79 / 2609.36529
Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling
Abstract
Recurrent neural networks (RNNs) compress the historical context into a memory state of fixed size, thus allowing for constant-time inference. The memory state size is a crucial factor in their performance, as exemplified by the strong performance and resurgence of linear attention, which extends the vector-valued hidden states of ordinary RNNs to matrix-valued hidden states. Crucially, linear attention does so in a parameter-efficient way, in particular by using an outer product of the key and value vectors to write to the matrix-valued hidden state. We generalize this construction and propose triadic linear attention, which writes the triadic outer product of a key, a second key, and a value, into a third-order (i.e., 3D) tensor state, and reads from it by contracting both key axes with two queries. An $E$-dimensional second key thus yields an $E$-fold increase in state size while adding only two projections. Triadic linear attention is compatible with data-dependent forgetting, the delta rule, and chunkwise-parallel training. Applied to Gated DeltaNet and scalar-gated linear attention, triadic linear attention substantially improves long-context language modeling and recall, outperforming alternatives that enlarge the state.
cs.LG / 80 / 2609.36546
Interactive-Policy Distillation with Bidirectional Propose-and-Verify
Abstract
On-policy distillation (OPD) trains a student model on its self-generated trajectories with dense token-level teacher feedback. However, naive OPD may suffer from teacher unanchoring, where the student's reasoning trajectory drifts far from the teacher, causing the teacher to be queried on states it would hardly visit and thus provide unreliable supervision. We propose Interactive-Policy Distillation (IPD), which applies adaptive teacher intervention to the student rollout. Under a bidirectional propose-and-verify state machine, the student and teacher alternately exchange their roles as proposer and verifier, and collaboratively generate mixed-source trajectories. Then different supervisions are applied according to the source of each token. This bidirectional propose-and-verify mechanism and the source-split loss make IPD not only a more performant distillation method, but also a unified bridge between on-policy and off-policy paradigms. To make the interleaved dual-model rollouts more efficient, we also design a dedicated fused inference engine that co-hosts both models in one serving instance with separate KV caches and instantiates the state machine model to distribute, collect, and process requests. On math reasoning tasks and across multiple teacher-student model pairs, student models trained with IPD not only outperform those trained with OPD, but also demonstrate higher data efficiency. Specifically, when distilling Qwen3-30B-A3B into Qwen3-1.7B-Base, IPD brings a +3.28 mean@8 and a +3.28 best@8 benchmark-averaged accuracy improvement compared with OPD. Besides, IPD only consumes about 1/4 of the training examples and steps to outperform OPD trained on the whole training dataset for one epoch. We also investigate the impact of different loss variants and takeover / handback configurations, and demonstrate the robustness of IPD on different training data.
cs.LG / 81 / 2609.36552
SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning
Abstract
Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's likelihood gradient by a factor that depends on its success probability and rollout count. Under uniform rollout allocation, the common rollout count fails to compensate for success-dependent attenuation, leaving low-success prompts more strongly attenuated and distorting their relative contributions to the expected aggregate gradient. We introduce SERA (Scale-Equalized Rollout Allocation), which redistributes a fixed rollout budget to approximately equalize these finite-rollout scaling factors. Building on our theoretical analysis of how finite rollouts distort prompt-wise likelihood gradients, we formulate the allocation as a fixed-budget max--min problem, derive a waterline solution to its continuous relaxation, and introduce a multiplicity correction to remove the additional prompt weighting induced by heterogeneous rollout counts. Experiments show stronger alignment with exact likelihood gradients in a controlled ImageNet setting and improved multi-sample solution coverage over MaxRL on maze navigation and mathematical reasoning under matched training rollout budgets.
cs.LG / 82 / 2609.36559
HiTS-CL: A Continual Learning Framework for Long-Horizon Temporal Knowledge Graph Extrapolation
Abstract
Extrapolative temporal knowledge graph reasoning (TKGR) predicts future facts from historical snapshots. Most existing methods train once on an early prefix of the timeline and then use a frozen model for all future timestamps. We argue that this fixed-prefix protocol is misaligned with extrapolation. It learns from a static prefix, whereas the target stream is non-stationary: new entities and facts emerge, temporal dependencies shift across regimes, and recurring historical signals must be refreshed online. As a result, models trained only on early snapshots become outdated and degrade over long horizons. We address this mismatch by formulating extrapolative TKGR as continual learning over streaming snapshots. Under this view, effective extrapolation must jointly handle current dynamics, stable knowledge, and recurring historical evidence. Based on these requirements, we propose History-enhanced Two-Step Continual Learning (HiTS-CL), a backbone-agnostic continual learning framework for extrapolative TKGR. HiTS-CL tracks current dynamics via continual fine-tuning, preserves stable knowledge via multi-teacher adaptive distillation, and retains recurring historical evidence via a selective memory of recent and frequent facts. We integrate HiTS-CL into five representative TKGR backbones and evaluate it on four benchmark datasets. HiTS-CL consistently improves extrapolation accuracy, reduces long-horizon degradation, and outperforms strong continual-learning baselines, including a recent method for temporal knowledge graphs. Source code and data are available at https://github.com/liuyansong98/HiTS-CL.
cs.LG / 83 / 2609.36569
From Checkpoint Variation to Selection Gains in Supervised Fine-Tuning
Abstract
Checkpoint selection is a routine decision in supervised fine-tuning (SFT): training produces multiple checkpoints, but only one is retained. Yet fixed-budget comparisons do not by themselves distinguish three empirical claims: whether more validation data improve checkpoint selection, whether a selection rule outperforms validation-loss selection, and whether it improves over simply retaining the final checkpoint. We therefore treat checkpoint selection as a finite-information decision problem. Holding completed training trajectories, candidate checkpoints, and independent test items fixed, we vary the validation budget and separately measure improvement from additional validation data, gain over negative log-likelihood (NLL) selection, and gain over the final checkpoint. Across 60 mathematical SFT trajectories and 19 configurations, increasing the validation budget from 32 to 305-313 examples raises independent-test accuracy by 0.32 percentage points (pp) for generated-accuracy selection and 0.29 pp for checkpoint agreement, with 95% configuration-bootstrap CIs of [0.10, 0.56] and [0.11, 0.50], respectively. At the full validation budget, the two generation-based rules outperform matched NLL selection by 0.71 and 0.85 pp, respectively, while their gains over the final checkpoint remain unresolved. A cross-domain replication on 12 newly trained Commonsense trajectories shows the same qualitative separation: increasing the validation budget from 32 to 1,024 questions improves generated-accuracy and checkpoint-agreement selection by 0.87 and 0.27 pp, while gains over the final checkpoint again remain unresolved. Together, these results show that benefiting from more validation data, outperforming NLL selection, and outperforming the final checkpoint are distinct empirical claims that require separate evidence.
cs.LG / 84 / 2609.36578
Factorized Scheduling Principle: Learning Interpretable and Transferable Policies via Structured Additive Functions
Abstract
Scheduling problems arise from repeatedly selecting one item from a set of candidates based on their states. These problems often reduce to assigning priority scores and choosing the highest-ranked item. In this work, we propose a factorized scheduling principle (FSP) framework to learn interpretable and transferable scheduling rules. The FSP framework represents system states as condition distributions and decomposes a global scheduling principle into additive univariate and pairwise components with identifiability constraints. The scheduling principle enables the framework to maintain a simple priority-based structure during deployment. This principle is learned by using a policy-based objective combined with a temporal-difference signal defined on the condition distribution. Experiments on synthetic and realistic scheduling tasks demonstrate the FSP framework's strong performance, interpretability, and zero-shot generalization across different system scales.
cs.LG / 85 / 2609.36584
Making Analog Training Scale: Co-Designing Mapping, Optimizer, and Converters
Abstract
Analog in-memory computing (AIMC) offers an alternative for model training by executing matrix operations directly where weights are stored. However, scaling AIMC to train modern deep models remains an open challenge due to severe hardware non-idealities, including physical weights with finite dynamic range and write granularity, analog-digital converters with finite resolution, and noisy and asymmetric updates. Guided by the insight that gradient accumulation is sensitive to precision and rounding errors, we adopt a mixed-precision training paradigm: executing forward and backward matrix multiplications in the analog domain while computing weight gradients in the digital domain. To enable scalable training, we present a holistic system-algorithm co-design that co-optimizes weight mapping to ensure well-conditioned physical and logical weight profiles, couples a preconditioned optimizer with threshold-triggered open-loop pulsing to stabilize training trajectories, and aligns converter dynamic ranges to suppress quantization errors. Evaluated via hardware-calibrated architectural simulations calibrated with electrochemical RAM measurements, our framework scales Transformer training up to $123\text{M}$ parameters with validation loss scaling as $L\propto N^{-0.231}$, where $N$ is the parameter count, comparable to $L\propto N^{-0.238}$ for digital training.
cs.LG / 86 / 2609.36587
Learned Reporting Preferences in RLVR Can Conflict with the Current Request
Abstract
Reinforcement learning with verifiable rewards (RLVR) has become a prominent approach for improving language-model performance on reasoning tasks using automatically checked answers. Yet convention-matched evaluation cannot reveal whether reinforcing one reporting convention reduces adherence to a different request that the initial policy already follows. To test this, we train matched policies under two reporting conventions and evaluate each policy under both current requests, using the same initial policy as a shared reference. We complement this crossed design with controlled interventions and independent human calibration. On GSM8K, boxed-format RLVR reduces the fraction of Qwen2.5-7B responses containing the requested hash-format payload by 35.33--74.37 percentage points relative to a 95.45% initial baseline in four of five training seeds; the fifth improves by 2.50 points. In the four deteriorating runs, almost every response that omits the requested payload instead retains the trained boxed convention, and the same four seeds deteriorate under two fixed paraphrases. Changing only the final-answer marker in supervised targets reverses which reporting convention the model prefers across three seeds, providing controlled evidence that this preference is learnable. Across three settings with independent human calibration, gains under a convention-sensitive scorer exceed the corresponding gains in committed-answer correctness, i.e., the correctness of the answer the model actually commits to. Together, these results separate three distinct post-training outcomes: learned reporting preference, current-request adherence, and committed-answer correctness. They show that convention-matched accuracy alone does not fully characterize post-training behavior and motivate evaluating current-request adherence alongside convention-matched task accuracy.
cs.LG / 87 / 2609.36608
Act First, Reason Later: Accelerating On-Policy Distillation for Multi-Turn Agents via Reference-Conditioned Inverse Dynamics
Abstract
On-policy distillation (OPD) trains multi-turn language agents with dense teacher supervision on student-generated responses. However, standard think-then-act rollouts require lengthy reasoning before each short action, delaying environment transitions and experience collection. Generating actions directly reduces this delay but can degrade rollout quality. To address this, we propose ActFirst-OPD, an act-first, reason-later training framework that decouples environment interaction from full-response generation. The student infers and executes actions through reference-conditioned inverse dynamics using its current interaction context and a reference next observation, and switches to autonomous next-action prediction when the resulting transition deviates from the reference trajectory. From the collected interaction contexts, the student asynchronously generates full think-then-act responses for token-level teacher supervision. Experiments across 0.6B-, 1.7B-, and 4B-parameter Qwen3 students show that ActFirst-OPD achieves average wall-clock training speedups of $2.3\times$ on ALFWorld, $1.8\times$ on WebShop, and $4.9\times$ on ScienceWorld over Vanilla OPD. It matches or exceeds all compared OPD baselines in mean task success rate across eight of nine benchmark-model settings. These results demonstrate that reasoning need not block acting during multi-turn agent distillation.
cs.LG / 88 / 2609.36610
Communication-Efficient Agnostic Federated Learning via Faster Convergence and Compression
Abstract
Agnostic federated learning (AFL) seeks a model that performs reliably across $m$ heterogeneous workers, but communication remains a bottleneck. We improve communication efficiency by reducing the number of synchronization rounds via faster convergence and the communication cost per round via compression. We first propose AFL-BR, which updates the dual weights over workers using online mirror ascent with KL divergence and blockwise restarts. It achieves an $O((\log m)^{1/4}T^{-1/8})$ stationarity rate after $T$ update rounds, reducing the $m$-dependence of the synchronization rounds required for convergence from polynomial to logarithmic order. Building on AFL-BR, we develop AFL-Com by applying bidirectional compression with error feedback (EF). Instead of compressing local gradients, workers apply EF to their dual-weighted gradients, enabling direct control of the aggregated compression error under time-varying weights. We then establish an $O((δ^{-1}+(\log m)^{1/4})T^{-1/8})$ stationarity rate for AFL-Com under general $δ$-approximate compressors and improve the $δ$-dependence from $δ^{-1}$ to $δ^{-1/2}$ for additive-and-idempotent compressors with shared randomness (SR). With suitable compression levels, AFL-Com retains the same convergence rate as AFL-BR at a lower per-round communication cost, yielding reductions in total communication complexity by factors of $(\log m)^{1/4}$ with Top-$k$ and $(\log m)^{1/2}$ with Rand-$k$ and SR. Experiments validate the improved synchronization and communication efficiency of our methods.
cs.LG / 89 / 2609.36614
Selective Elicitation as a Commercial Influence Channel: A Reproducible Synthetic Shopping-Agent Stress Test
Abstract
A commercial incentive need not enter the final ranking algorithm to affect a shopping assistant's recommendation: it may instead influence which preference question the assistant asks. We make this distinction experimentally observable in a deliberately small, synthetic setting. Each task has two products, three verified numerical attributes, a price limit, and a private fixed preference vector. An honest simulated user answers one pairwise question. A separate recommender receives the products and this answer but not the sponsorship assignment. We contrast a neutral question, a soft commercial instruction, and an explicitly adversarial instruction to ask about the sponsor's advantage while omitting the rival's advantage. Across 40 held-out sponsorship-assignment cases (20 distinct catalog-preference contexts), the soft instruction changes no selections. The targeted instruction raises sponsored selection by 0.30 and reduces mean synthetic utility by 0.0547 relative to neutral questioning (95% context-bootstrap interval [-0.0828, -0.0291]) for one language-model recommender. A fixed Bayesian recommender shows a similar effect; a second model makes the same choices on all 120 frozen question-answer inputs. A terminal-answer consistency judge rates all 20 sampled targeted answers consistent, although five have synthetic regret above 0.05; a separate question-coverage dimension flags their one-sided elicitation. A robust partial-preference certificate remains valid under the stipulated synthetic utility but certifies only 16 of 40 targeted cases and is not better than asking a neutral question directly. These results establish neither typical behavior under advertising incentives nor effects on actual consumers.
cs.LG / 90 / 2609.36619
SemPSG: A Semantic Channel-Aware Foundation Model for Polysomnography Analysis
Abstract
Polysomnography (PSG) integrates multiple physiological signals to provide a comprehensive characterization of human sleep, yet its heterogeneous channel configurations across centers pose substantial challenges for transferable representation learning. Existing foundation models mainly focus on physiological modeling or temporal learning, while channel identity is often treated as a fixed structural index, overlooking the physiological semantics encoded by signal modality and reference configuration. To this end, we propose SemPSG, a Semantic channel-aware foundation model for heterogeneous PSG analysis. SemPSG explicitly represents the physiological semantics of channel identity and incorporates them into both signal representation learning and channel aggregation, enabling flexible modeling across diverse data configurations. Specifically, a semantic-conditioned time-series encoder captures signal-specific temporal dynamics and cross-signal interactions, while a multi-view image encoder extracts complementary time-frequency and morphological patterns from the same physiological recordings. We evaluate SemPSG on sleep and health-related tasks, including sleep staging, sleep-disorder breathing analysis, disease prediction, cognition and emotion recognition, and demographic estimation. Extensive experiments demonstrate consistent improvements over both general-purpose time series foundation models and PSG-specific foundation models, together with generalization across heterogeneous datasets across diverse channel configurations.
cs.LG / 91 / 2609.36633
Physics-Aware Machine Unlearning for Cyber-Physical Systems
Abstract
This paper proposes a physics-guided gradient-ascent-based machine unlearning method that couples the forgetting signal with the physical residual of the target cyber-physical systems, ensuring that weight updates during unlearning are steered toward physically feasible regions of the weight space. The physics residual acts as a safety fence during gradient ascent: the model is steered away from the poisoned behavioral basin and simultaneously toward physics-compliant territory, rather than toward an arbitrary alternative that may still violate domain constraints. We evaluate the proposed method against four baselines: naive gradient ascent, exact unlearning, SISA, and full retraining on an IEEE 34-bus distribution system, driven by two physics-informed neural network-based distribution energy resource controllers and validated through high-fidelity OpenDSS power-flow co-simulation. From the evaluation, we found that our proposed physics-guided model simultaneously removes poison and restores physical compliance, which are essential for the safe deployment of safety-critical cyber-physical systems
cs.LG / 92 / 2609.36636
What Makes Recurrence Effective in Looped Language Models?
Abstract
Looped language models (LoopLMs) increase computational depth through parameter sharing, offering a path to scale inference computation without adding parameters. However, it remains unclear when additional recurrence is beneficial and how architectural choices affect its effectiveness. Through controlled experiments, we systematically examine (1) when recurrence helps, (2) where it should be applied, and (3) how its conditioning affects performance. Our evaluation covers inference budgets below, within, and beyond the training horizon under knowledge and reasoning tasks. (1) We find that recurrence can improve reasoning beyond the training horizon while degrading knowledge performance, but harder reasoning instances do not consistently benefit more. (2) Performance also depends on how distinct layers and recurrent iterations are allocated, showing that effective depth alone is insufficient to predict behavior. Non-recurrent output layers improve robustness to under-unrolling, while the preferred placement of input and output layers varies with inference budget. (3) Finally, we find that conventional initial-state injection offers limited robustness to varying recurrence depth. We therefore propose history-state injection as an alternative, and show that channel-wise history-state injection combined with timestep conditioning offers a low-cost and more effective design, better preserving knowledge under extended unrolling while improving robustness across inference budgets. Overall, our results clarify when recurrent computation helps, where it fails, and offer practical guidelines for designing LoopLMs across variable inference budgets.
cs.LG / 93 / 2609.36638
PE-OPSD: Internalizing Prompt Enhancement into Flow-matching Models via On-Policy Self-Distillation
Abstract
Text-to-image users often provide concise and underspecified prompts, whereas generative models benefit from detailed textual conditions for reliable instruction following. Existing systems bridge this gap with Prompt Enhancers (PEs) that rewrite raw prompts at inference time, introducing additional latency and leaving prompt elaboration external to the generator. We instead view enhanced prompts as privileged training information and ask whether their benefits can be internalized. We propose Prompt-Enhanced On-Policy Self-Distillation (PE-OPSD) for text-to-image flow-matching models. During training, a raw-prompt student follows its own generation trajectory, while an enhanced-prompt teacher provides vector-field targets at the states visited by the student. This on-policy supervision distills the behavior induced by enhanced prompts into the raw-prompt student without requiring additional text--image pairs. At inference, both the PE and teacher are removed, and the student generates directly from raw prompts. Across multiple model families, PEs, and benchmarks, PE-OPSD achieves the strongest aggregate prompt fidelity among the evaluated baselines, yields positive aggregate visual appeal gains, and retains the base-model inference efficiency.
cs.LG / 94 / 2609.36639
A Digital Simulation Toolkit for Physics-Based Generation of Realistic Experimental Scanning Tunneling Microscopy Images
Abstract
Scanning Tunneling Microscopy (STM) is a widely used tool for characterizing surfaces of materials at the atomic scale, playing a crucial role in discoveries across condensed matter physics and materials science. Despite its extreme spatial resolution, STM is one of the most sensitive microscopy techniques and is highly prone to noise. While existing unsupervised denoising methods are very cheap to train, these are primarily focused on removing the noise with minimal recovery of key physical information. While supervised methods can offer superior performance, the major bottleneck is that a large amount of paired clean-noisy experimental images is required which are impractical to obtain. Thus, we developed a low-cost physics-driven digital toolkit to rapidly generate large volume of realistic STM images. Firstly, we simulate clean images from a chosen material system. Then, with prior knowledge of the physical characteristics of the artifacts and noise present in STM experiments, we formulate several artifact-noise functions such as Gaussian electronic noise, 1/f flicker noise, scan-line noise, background tilt and mechanical drift. These physically informed noise components are then added to the simulated clean images to generate realistic STM images. We demonstrated the capability of the proposed digital toolkit to generate AI-ready data for denoising images of the (111) surfaces of copper and lead, while preserving atoms, defects, and electron waves. We also validated the quality of the downstream image analysis of learning electron wave patterns induced by quantum interference from Cu(111) images. Results show that the supervised models trained on digitally generated AI-ready data can more effectively denoise and learn electron wave patterns on Cu(111) images than benchmarked unsupervised approaches, indicating that the proposed toolkit facilitate scientific discovery.
cs.LG / 95 / 2609.36641
Inducing Process Supervision from Outcome-Only Reinforcement Learning
Abstract
Process reward models (PRMs) have become a key component for LLMs, as their step-level feedback supports both post-training and test-time reasoning. However, training strong PRMs remains costly: human step annotation is difficult to scale, while Monte Carlo estimation is computationally expensive and can drift from the intrinsic correctness of steps. To get effective PRMs at low cost, we introduce TIPS (Thinking-Induced Process Supervision), an outcome-only reinforcement learning (RL) framework for training generative PRMs. In TIPS, the model generates a chain-of-thought (CoT) followed by step-level labels and an outcome label. The reward depends solely on whether the predicted outcome matches the ground truth, and the resulting group-relative advantage is used to optimize the entire generated response. Intuitively, when checking intermediate steps helps determine the outcome, more accurate checks can lead to better outcome judgments and higher rewards. Outcome-only RL can therefore reinforce step-level verification without explicit process supervision. We validate the effectiveness of TIPS across math and agent benchmarks and four backbone families. Notably, TIPS-Qwen3-4B-Thinking-2507 reaches 85.2 F1 on ProcessBench with only 3.2K outcome-labeled trajectories, surpassing all evaluated trained PRMs and strong prompt-only judges such as GPT-5.4-Instruct and Claude-4.7-Opus, while still trailing o1-mini. Code and data are available at https://github.com/RUCBM/TIPS.
cs.LG / 96 / 2609.36642
PR-OPD: Privileged Representation On-policy Self-Distillation for Agentic Reinforcement Learning
Abstract
Language-model agents are usually trained by reinforcement learning from one reward per episode, and privileged self-distillation enriches it by letting the same policy, given a skill, teach its skill-free self through token probabilities. However, we identify two phenomena that question this channel. Invisible Advantage: a skill in context lifts WebShop success from 42.2% to 56.2%, yet changes the probabilities of fewer than a quarter of the sampled tokens. Much to Align: a skill changes the hidden states of over 80% of response tokens, in a way that linear probes can trace back to the specific skill. To exploit this, we propose Privileged Representation On-policy Self-Distillation (PR-OPD). After a GRPO warm start, the policy writes a hindsight skill for each trajectory, re-reads its own responses with that skill as a stop-gradient teacher, and aligns its projected hidden states to the teacher's at every layer alongside the reward objective, with no external skill library, separate teacher, or inference overhead. On ALFWorld and WebShop with two backbones, PR-OPD achieves the best overall results in every setting, improving over GRPO by up to 4.7 points in ALFWorld success and 14.0 points in WebShop accuracy. Code is available at https://github.com/balibata/PR-OPD.
cs.LG / 97 / 2609.36653
Scheduling Recursive Reasoning in Looped Transformers
Abstract
Recurrent reasoning models have attracted growing attention for scaling test-time computation, typically by iteratively refining latent states with shared parameters. However, these models apply each learned update with a fixed unit scale, which can be conservative when updates make persistent progress and overly aggressive when they fluctuate, limiting the benefit of additional loops. To understand how the scale should vary along the trajectory, we first analyze the sensitivity of terminal loss to recurrent update scale. We show that its temporal average admits an exact decomposition into persistent-progress and centered-fluctuation contributions. Based on this, we introduce the Trajectory Adaptive Progress-Fluctuation Scheduler (TAPS), which tracks their balance across recurrent updates and adapts the step size online. Theoretically, we establish sufficient conditions under which TAPS reduces expected terminal loss and reaches a target quality in fewer recurrent loops. Empirically, we show that TAPS improves terminal accuracy across structured reasoning tasks without retraining. By further incorporating the progress-fluctuation principle into training, TAPS yields additional accuracy gains with up to 1.56 times wall-clock speedup at matched baseline accuracy. The broad applicability of TAPS is supported by its effectiveness across diverse recurrent architectures and inference strategies. Together, these results establish update scale as complementary control axis of recurrent inference alongside architecture and depth.
cs.LG / 98 / 2609.36657
Constitutional adapters: Inference-time interventions for misalignment and misuse
Abstract
Training models to act in accordance with an explicitly defined set of principles, or "constitution," has shown promise as a robust and transparent mechanism for AI alignment. However, the generality and flexibility of such methods remain unclear. Here, we show that constitution-consistent behavior can be distilled from synthetic corpora into lightweight objects (low-rank adapters and steering vectors). Despite never seeing a harmful request or jailbreak during training, such objects increase jailbreak defense success and measured alignment -- particularly at long context lengths and against multi-turn attacks, where they outperform both prompted and steered baselines. Subtracting control-trained from constitution-trained objects further accentuates these effects, yielding defenses we call "constitutional adapters" (CAs). CAs can be trained on a base model, transferred zero-shot to its post-trained checkpoint, and scaled at inference time to predictably trade off defense for benign compliance. Taken together, these results recommend CAs as a lightweight, portable, and tunable lever for mitigating misalignment and misuse in API deployments.
cs.LG / 99 / 2609.36659
On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training
Abstract
The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these studies treat the observed behaviors only as byproducts in on-policy training, overlooking their potential to serve as optimization principles for improving the generalization of other paradigms such as supervised fine-tuning (SFT). To address this limitation, we investigate whether there exists a specific on-policy update behavior that can achieve such improvements. First, our analyses reveal that SFT updates parameters along consistent directions, while the on-policy paradigm continuously adjusts the direction during training. This difference inspires us to focus on the cumulative update direction of each parameter as a promising behavior. Then, we evaluate its effectiveness for improving generalization by proposing On-Policy direction-constrained Supervised Fine-Tuning (OPSFT), which constrains SFT updates to the direction identified by on-policy paradigms. The strong performance of OPSFT indicates that the generalization advantage of on-policy paradigms can be transferred to SFT through the parameter update direction. Once such a direction is identified, even SFT can generalize with its updates constrained to this direction. This finding offers two practical benefits by combining the strong generalization of on-policy paradigms with the advantages of SFT, including the high training efficiency and ability to leverage high-quality trajectories. For efficiency, we identify update directions that support strong generalization using a few on-policy training steps, and subsequently apply OPSFT to achieve high training efficiency. For leveraging high-quality trajectories, OPSFT can utilize these trajectories to continue improving a post-trained model along its update direction without disrupting the ability learned from on-policy training.
cs.LG / 100 / 2609.36660
Byzantine-Robust Federated Representation Learning
Abstract
We study federated learning (FL) with adversarial clients, where the goal is to minimize the average loss of the honest (non-adversarial) clients without knowing their identity. Under heterogeneity, a single shared model parameter is statistically inappropriate: it cannot capture the distinct data-generating processes across clients, incurring an irreducible model-heterogeneity bias and severely limiting robustness to adversarial clients (a.k.a. Byzantine-robustness). We address this problem through representation learning, where each client learns a personalized linear head, while collaboratively estimating a shared nonlinear representation through Byzantine-robust aggregation. We demonstrate that the heterogeneity among honest representation gradients is controlled by the representation error and statistical errors that decay either with the number of data samples per client ($τ$) or the number of iterations ($T$). In particular, our non-asymptotic parameter recovery error bound reveals three terms: (i) an initialization-dependent error that goes away with $T$, (ii) finite-sample noise terms that decreases with $τ$ and the number of honest clients, and (iii) a stochastic gradient variance term that also reduces with $T$. Importantly, with no irreducible model-heterogeneity bias in our bounds. We extend the regression analysis to multiclass classification, and empirically validate it on CIFAR-10, FEMNIST, and School Exam Score datasets.
cs.LG / 101 / 2609.36663
GenLimitLib: A Formal Library for Language Generation in the Limit and AI-Assisted Mathematical Research
Abstract
We present GenLimitLib, a source-aligned Lean 4 library for language generation in the limit. Introduced by Kleinberg and Mullainathan at NeurIPS 2024, language generation in the limit studies a theoretical question motivated by LLMs: how to generate valid new strings from observed examples. This young and rapidly evolving field offers a natural testbed for studying large-scale formalization. GenLimitLib contains formal developments for 30 papers. It extracts shared definitions and reusable proof components while preserving paper-specific assumptions and statements, and records relationships across papers. In this way, GenLimitLib provides a concrete and structured view of the literature. We show through mathematical case studies and LLM experiments how our library can support both human mathematical research and AI-assisted research. Our Library: https://github.com/pengzhang91/generation-in-the-limit-lib.
cs.LG / 102 / 2609.36668
Stochastic Heavy Ball with Polyak Step Size and Armijo Line Search: A General Convergence Analysis
Abstract
Polyak step size (PS) and Armijo line search (ALS) have received increasing attention in stochastic optimization, with encouraging empirical performance and theoretical guarantees. However, their convergence theory for stochastic heavy ball (SHB) methods remains limited. In this work, we develop a unified convergence analysis for SHB equipped with PS and ALS. To this end, we introduce a modified Armijo rule that closely parallels the Polyak step size, together with a decoupling analysis that isolates the historical dependence induced by momentum. For SHB with standard PS and ALS, we establish expected convergence for strongly convex, convex, and non-convex objectives without interpolation or restrictive conditions on the momentum parameter. Under interpolation or strong growth, we further strengthen the results to almost sure rates and last-iterate convergence. Moreover, for general settings beyond interpolation, we prove almost sure convergence to the exact optimum or to stationarity for SHB with diminishing variants of PS and ALS. These results provide a more comprehensive theoretical view of Polyak step size and Armijo line search for stochastic heavy ball methods.
cs.LG / 103 / 2609.36672
Human-inspired, Task-Dimension-Guided Exploration for Efficient Learning in High Dimensions
Abstract
Efficient exploration in high-dimensional decision spaces remains a central challenge for decision-making systems. Humans, in contrast, can navigate large decision spaces with remarkable efficiency. Recent behavioral studies suggest that humans reduce dimensionality in large decision spaces by probing candidate feature dimensions, identifying reward-relevant ones, and restricting the effective decision space. Inspired by this mechanism, we propose TDGE (Task-Dimension-Guided Exploration), a human-inspired, model-agnostic algorithm with an automatically constructed task-dimension--feature--item hierarchy. TDGE follows a top-down exploration strategy: it first selects task-relevant feature dimensions, then identifies informative features within those dimensions, and finally recommends concrete items based on the selected features. Experiments on MovieLens-20M, Last.fm, and Amazon recommendation datasets show that TDGE substantially improves exploration efficiency and cold-start adaptation over baseline algorithms. Comparisons with other structured algorithms and ablation studies attribute these gains to TDGE's hierarchical structure and semantic feature-space exploration, with robust results across clustering methods and hierarchy depths. Recommendation-trajectory visualizations also show exploration patterns similar to human dimension-guided behavior.
cs.LG / 104 / 2609.36678
Understanding Private Evolution as Learning-Augmented Clustering
Abstract
Private Evolution (PE) is a differentially private algorithm for synthetic data generation. While it can be viewed as a Wasserstein learning algorithm, it performs much better in practice than worst-case Wasserstein analyses would predict. We recast PE as generative model-augmented Wasserstein learning. We show theoretically that when we take into account the use of a generative model that is able to capture something about the true distribution, then we can obtain much better performance bounds. For example, if the generator gives samples in the same low-dimensional space as the distribution, then sample complexity depends on intrinsic, not ambient, dimension. We also show that standard variants of PE can fail to converge on simple well-clustered instances, and propose a new geometry-aware version of PE with provable convergence on such instances. Experimentally, we show that our new algorithm is competitive with standard baselines and can improve recall.
cs.LG / 105 / 2609.36683
MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization
Abstract
Molecular optimization is inherently iterative: a candidate is proposed, evaluated against several objectives, and revised while preserving a relationship to the source molecule. Most instruction-following models instead emit one edited molecule, forcing validity, property improvement, and similarity control into a single response. We introduce MARCO, an evaluator-grounded reinforcement-learning framework that trains molecular editors on bounded proposal--feedback--revision trajectories. MARCO aggregates shaped turn rewards into an undiscounted trajectory return for group-relative policy optimization. We evaluate two consequences of this training: Same-1 tests the trained policy under a one-response budget, while Same-5 tests whether the same policy can use verifier feedback when up to five responses are available. Across the three-objective MuMOInstruct benchmark, three Qwen backbones, and seen/unseen instruction splits, SFT-initialized MARCO obtains the highest product of property success rate and similarity in every reported primary setting. Same-5 further improves the observed score under the tested budget, while four-objective and public-checkpoint experiments test transfer across constraint sets and initialization regimes.
cs.LG / 106 / 2609.36686
Where Root Cause Analysis Fails: A Retrieval-Reranking Decomposition
Abstract
Identifying the root cause of an anomaly among hundreds of sensors is critical for preventing safety incidents and costly downtime in complex monitored systems. Existing studies evaluate root cause analysis (RCA) methods using top@k accuracy. We show that this metric has a fundamental blind spot: it conflates two failure modes, retrieval failure, where the true cause is never considered, and reranking failure, where it is considered but ranked too low. In this work, we introduce a retrieval-reranking decomposition and audit four well-known benchmarks to expose this blind spot. Our experiments show that, on benchmarks with complex faults, statistical baselines mis-rank the true cause 79-100% of the time, and graph-based methods never clearly beat the best statistical baseline, whether their causal graphs are learned on short fault windows, on retrieved candidate pools guaranteed to contain the cause, or on multi-day normal-operation data. Meanwhile, on simple benchmarks where faults manifest significantly at their origin, retrieval is nearly solved (98-100%). Guided by the decomposition, we build a two-stage pipeline combining a multi-signal retriever with an LLM reranker that, as one fixed configuration, matches or exceeds the best baseline's top@1 accuracy on all six benchmark suites (by up to +12 points), with no causal graph or labeled data required. When all methods rank the same retrieved candidates with the true cause guaranteed present, adding a short system-description document lets the reranker lead the best baseline by +7 to +18 points on every benchmark. Code is available at https://github.com/cruiseresearchgroup/DecompRCA.
cs.LG / 107 / 2609.36692
Normalize-Then-Precondition: A Hierarchical Approach to Marginal Scale and Interaction Geometry for LLM Training
Abstract
Matrix optimizers have emerged as a promising direction, with Muon standing out as a prominent design. Revisiting Muon through its full-Gram representation, we observe that it jointly processes marginal-scale and interaction information. This opens an alternative way to organize geometric information hierarchically, motivating the Normalize-Then-Precondition framework. Specifically, it first uses diagonal-Gram information to construct a marginally normalized update, then applies spectral preconditioning to its directional interaction geometry. Building on this framework, we develop NormPre with NormPre-G and NormPre-L adopting global and localized spectral preconditioning, grounded in spectral-norm steepest descent and a regularized formulation followed by leading mode selection, respectively. To enable large-scale training, NormPre-G uses Newton-Schulz iterations and NormPre-L employs randomized sketching to approximate the leading interaction eigenspace. Theoretically, we establish $\mathcal{O}(T^{-1/2})$ convergence guarantees for simplified versions of NormPre. Across extensive pretraining experiments on GPT-2 Small, LLaMA and Qwen3, both variants consistently outperform AdamW, Muon and MANO under matched training budgets. Further efficiency and spectral analyses reveal the complementary strengths of two variants and characterize their performance-efficiency trade-off. We open-source our code through a GitHub repository at https://github.com/zx-gong/NormPre.
cs.LG / 108 / 2609.36695
Know Thyself, Teach Thyself: Internal Information Flow for Selective Self-Distillation
Abstract
Self-distillation turns knowledge distillation into a closed learning loop and offers a path toward recursive self-improvement. Without an external teacher, however, the model must determine both what information can improve its supervision and which induced changes should be learned. Existing methods typically improve teacher-generated data or select training examples in isolation, leaving the information transferred between these stages unmeasured. We introduce InFlow, a retrieval-guided on-policy self-distillation framework that models this process as potential-to-realized information flow. InFlow first retrieves potentially informative sources using certainty-calibrated hidden-state trajectories, then measures their realized effect through the Jensen--Shannon divergence between the teacher's initial and retrieval-conditioned answer beliefs. Examples with larger belief shifts are selected for on-policy distillation. Our analysis formalizes the information optimized by retrieval and selection and relates the answer-level shift to the teacher--student distillation gap. Across four open-weight language models and three knowledge domains, InFlow achieves the strongest cross-model average among the compared selection methods, with ablations supporting both stages of the framework. Our code is available at https://github.com/1240148048/INFLOW.
cs.LG / 109 / 2609.36698
Learned Queries and Keys Are All You Need: Replacing the Value Projection with Structured Transforms
Abstract
To reduce the number of parameters and cache memory requirements of transformers we introduce dual-headed transformers instead of three heads. We studied Walsh-Hadamard Transform (WHT), Discrete Cosine Transform (DCT), Discrete Fourier Transform, filterbank based Shearlet Transform, and Multiplication-Avoiding (MA) operators to construct dual heads. We combine spatial patches and their orthogonal transforms (or Shearlet and MA operators) in a structure similar to the attention block. We obtained better results than triple headed transformers in ImageNet. Extensive simulation examples are presented.
cs.LG / 110 / 2609.36704
When Is Coarse Supervision Worth It? Cost-Aware Learning under Unknown Aggregation
Abstract
Modern learning systems often acquire supervision at multiple resolutions, trading annotation cost against information content. We study cost-aware two-resolution learning, where expensive fine labels reveal a vector response and cheaper coarse labels reveal a scalar aggregate formed with unknown weights, while the target remains the full response. The challenge is that unknown aggregation changes which directions coarse data can identify, so the value of coarse supervision depends jointly on cost, noise, and identification. We characterize this information geometry and develop an estimate-and-track policy that learns the aggregation rule and tracks the optimal resolution mix. We derive a closed-form break-even condition for coarse supervision and prove that the online policy attains the optimal leading cumulative-risk coefficient, with a matching local asymptotic minimax lower bound. Synthetic experiments support the predicted all-fine/mixed transition, show the online learner approaching the oracle-share benchmark, and demonstrate a finite-budget gain over all-fine acquisition when coarse supervision is sufficiently favorable. Our results provide a principled way to balance information and annotation cost across supervision resolutions.
cs.LG / 111 / 2609.36721
CALIBUDGET: Calibration-Guided Source Allocation for Fixed-Budget Mixed-Reasoning Adaptation
Abstract
Fixed-budget adaptation from heterogeneous data sources requires deciding not only how much data to use, but how much exposure each source receives. Size-proportional rules can crowd out small sources, whereas difficulty-only rules can chase noisy estimates or allocate residual budget to nearly saturated pools. We introduce CALIBUDGET, a floor-protected, reliability-aware integer allocator that treats source exposure as an explicit adaptation variable. From small train-internal calibration splits, it combines model need, post-floor availability, and bootstrap stability, then produces exact capacity-respecting quotas without changing the model, objective, or total budget. In a controlled setting combining mathematical and commonsense data, CALIBUDGET improves CommonAvg, FragileAvg, and MacroAvg over validation-error-with-floor, the strongest matched comparator, in all three paired LLaMA-2-7B LoRA+ runs. The respective mean gains are 0.56, 0.46, and 0.41 percentage points (pp). Overall increases by 0.18 pp, whereas MathAvg decreases by 0.20 pp, exposing a coverage-retention boundary rather than a uniform gain. CALIBUDGET changes only 1.14-1.42% of the source budget but improves performance in 15 of 24 comparisons across commonsense tasks and seeds. These results suggest that small changes in source quotas can matter; example-level selection can then determine which examples fill each quota.
cs.LG / 112 / 2609.36724
Routing in Gradient Space: Balanced Usage Is Not Expert Specialization
Abstract
Sparse expert models can distribute traffic evenly while still grouping incompatible training signals within the same experts. We study routing as a gradient-partitioning problem and introduce gradient-aligned routing (GAR), whose load-normalized router objective rewards grouping observations with aligned gradients. On five multi-task text-classification mixtures, we compare GAR with task-loss-only routing, gradient-combination and gradient-conflict methods, and load-balancing losses. With a fully trainable RoBERTa backbone and classification-head experts, GAR has the highest aggregate validation accuracy, 1.07 percentage points above task-loss-only routing. With frozen DeBERTa and Qwen3-1.7B backbones and low-rank adapter experts, it again ranks first, 1.10 points above task-loss-only routing, with better-balanced expert load and higher gradient-mass purity, the share of each expert's gradient-norm mass from its dominant task; the load-balancing losses flatten load further but leave this purity near its task-loss-only level. Top-1 routing, trainable full-parameter feed-forward network (FFN) experts, and a larger backbone also show positive aggregate gains. The results distinguish expert-load balance from gradient-based routing organization and indicate the predictive value of gradient-informed routing in multi-task text classification.
cs.LG / 113 / 2609.36738
Backpropagated Output Momentum: Relocating Optimizer History from Parameters to Task Space
Abstract
Optimizer momentum is usually stored as a parameter-sized moving average of past gradients, which makes history costly and fixes each past signal in the coordinates in which it was computed. We introduce Backpropagated Output Momentum (BOM), which instead stores a compact moving average of prediction errors at the model output and reprojects that history through the current network at every step. A batch-level analysis characterizes the information retained and omitted by this relocation, while the implementation preserves the current supervised gradient and can replace the first-moment component of several adaptive optimizers. As a plug-in for momentum-based optimizers, including ones that already compress their state, BOM reduces parameter-shaped optimizer state by 49.7-99.8% in three compositions and, averaged over three language backbones, paired step time by 4.0%. It also improves mean validation performance across language and vision fine-tuning, by 1.42 points in the primary five-task comparison. Language and vision pretraining studies, together with matched mechanism controls, further test the construction across output spaces and model scales.
cs.LG / 114 / 2609.36740
Efficient Offline Learning of Ranking Policies via Top-$k$ Policy Decomposition
Abstract
Many recommender systems such as for e-commerce and news platforms aim to provide users with rankings they are likely to interact with. Off-Policy Learning (OPL) of ranking policies enables us to learn new ranking policies using only historical logged data. However, ranking settings make OPL remarkably challenging because their action spaces consist of permutations of unique items, being extremely large. Existing methods primarily use either policy- or regression-based approaches. The policy-based approach, which typically uses importance-weighted policy gradients, can suffer from high variance due to large action spaces. The regression-based approach, on the other hand, estimates the expected reward using conventional machine learning methods, avoiding variance issues but potentially suffering from severe bias. To circumvent these issues of existing methods, we propose a new OPL method for ranking, named Ranking Policy Optimization via Top-$k$ Policy Decomposition (R-POD), which combines the policy- and regression-based approaches in an effective fashion. Specifically, R-POD decomposes a ranking policy into a first-stage policy for selecting top-$k$ actions and a second-stage policy for choosing the bottom actions given the top-$k$ actions. It learns the first-stage policy using a new policy gradient estimator and the second-stage policy via the regression-based approach. This method can substantially reduce variance, since it applies importance weighting only to the top-$k$ actions. We also demonstrate that our policy-gradient estimator for the first-stage policy is unbiased under a conditional pairwise correctness condition, which only requires that the expected reward differences of pairs of rankings sharing the same top-$k$ actions can be estimated correctly.
cs.LG / 115 / 2609.36752
cktFormer: Transformer-Based Approach for Automated Analog Circuit Design
Abstract
Circuit design is a complex and iterative process that requires expertise in electronic engineering. It involves selecting components while meeting performance constraints, such as power efficiency, cost-effectiveness, and signal integrity. However, manual design is time-consuming and prone to errors. Although other stages of the manufacturing pipeline have benefited from AI-driven optimizations, circuit design remains a bottleneck, limiting overall productivity. Generative AI and machine learning offer the potential to automate and improve this stage, boosting efficiency and accuracy. To address this, we introduce a dual transformer architecture that bridges the gap between AI and circuit design by leveraging attention mechanisms to model complex, non-sequential circuit relationships. Our approach structures netlist data into graph-based representations, enabling effective learning of circuit topology and component interactions. The system consists of two interlinked models: a node prediction model that proposes components and an edge prediction model that infers valid connections. This collaborative and decoupled design captures both component-level semantics and global structural coherence. In our experiments, this architecture outperforms recent models such as AnalogGenie and cktGNN in the validity of generated circuits. By addressing key limitations in existing methods, our work advances automation in electronics engineering and contributes a benchmark for AI-driven circuit synthesis.
cs.LG / 116 / 2609.36760
QuantMLA: Function-Aligned Dual-Path Quantization for Low-Bit MLA KV Caching
Abstract
Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memory still scales linearly with context length and batch size. In this work, we establish a systematic model of MLA's dual-path quantization errors, characterizing their distinct effects on attention-output distortion and explaining the pronounced amplification of RoPE-path errors. Guided by this analysis, we introduce QuantMLA, a function-aligned framework for low-bit dual-path quantization. We derive path-specific transformation spaces that preserve full-precision computation while remaining fully fusible into model parameters offline, eliminating online transformation overhead. Within these spaces, QuantMLA learns path-specific transformations with function-aligned objectives: attention-output reconstruction captures the content path's coupled matching and aggregation errors, while positional QK reconstruction preserves the RoPE-induced component of the attention logits and admits a theoretical bound on output distortion. Across four MLA model families, QuantMLA enables, to our knowledge, the first reported joint INT4 caching of the content and RoPE caches with minimal accuracy degradation. Further compressing the content cache to INT2 while retaining the RoPE key cache at INT4 maintains competitive performance on challenging reasoning and code benchmarks. We develop a native low-bit MLA attention kernel that integrates unpacking and dequantization directly into attention computation. The physical cache layout provides 3.59x compression at 128K context, while a cache-pressure serving workload achieves 5.168x higher whole-job output throughput than BF16. The code will be released upon acceptance.
cs.LG / 117 / 2609.36762
Federated Clustering with Unknown Local and Global Cluster Cardinalities
Abstract
Federated clustering methods that do not require the global number of clusters $K$ still assume that each client knows its local number $K_g$. This assumption is hard to justify when clients know no more about their data than the server does, as in fault diagnosis across independently operated industrial sites. We propose a two-phase framework in which neither count is known: each client first estimates $K_g$ from its own data, and an aggregator that requires local counts, such as FedGEM, then uses these estimates in place of the true values. For the first phase we introduce Adaptive Split--Merge (ASM), which grows a spherical Gaussian mixture by BIC-driven splitting and then merges excess components. ASM uses no labels, selects its hyperparameters on held-out client data only, and makes no assumption about how clusters are shared across clients. We derive a closed-form split criterion whose critical cluster size falls with anisotropy and rises with dimension, and show empirically that over-fragmentation grows with the number of points per cluster, which federation divides among clients. Across eight datasets, ASM with FedGEM attains a mean ARI of 0.333, against 0.256 for the next best label-free estimator and 0.361 when the true local counts are supplied. It also gives the most reliable global estimates of $K$ and is robust when client size is decoupled from local cardinality.
cs.LG / 118 / 2609.36765
Graph-Spectral Flow Matching for Multivariate Time Series Anomaly Detection
Abstract
Multivariate time series anomaly detection typically relies on evaluating discrepancies between observations and outputs produced by models trained on normal data. An alternative perspective is to characterize the distribution of normal data through the generative dynamics, i.e., the velocity field, of flow matching models. However, standard flow matching typically adopts linear probability paths that overlook dependencies among variables, leading to a misalignment with the structured data distribution. To address this issue, we propose GRASP, a flow matching framework with a graph-spectral path for multivariate time series anomaly detection. GRASP incorporates graph structure into the probability path by minimizing a fixed-endpoint action that combines kinetic energy with graph Dirichlet energy. This formulation yields a closed-form path based on graph-frequency-dependent hyperbolic interpolation. A velocity predictor trained on normal data then detects anomalies using weighted velocity discrepancies aggregated across source samples, flow times, and graph frequencies. Theoretically, we establish that GRASP is invariant to the choice of Laplacian eigenbasis and decompose its expected oracle anomaly score into bounded endpoint uncertainty and graph-frequency-weighted Fisher discrepancy. Experiments on four benchmarks demonstrate the superior anomaly detection performance of GRASP and validate the effectiveness of its graph-spectral path and weighting mechanism.
cs.LG / 119 / 2609.36766
When Can Prefixes Compile LoRA? Exact Resource-Capped Tests for Frozen Attention
Abstract
Can a fixed continuous prefix replace a given low-rank adapter while the attention head stays frozen? In this research, we show that the answer depends on the adapter's target through three conditions. First, observability: at one causal readout, every independent key--value prefix sees the content only through the query, attention partition, and value numerator, so a target that differs on two inputs with equal summaries incurs an error floor at every prefix length; norm caps extend this floor to nearly equal summaries. Second, realizability: at a common query, any prefix reduces exactly to two aggregate variables, and the norm-capped optimum is an attained second-order-cone program, also after a fixed output projection; it places two equal-norm rank-one value updates on opposite sides of compilability. Third, implementation: under affine query exposure, $2r$ signed slots approximate a rank-$r$ value update, but their values grow as $O(ε^{-3/2})$, and the construction passes all 400 tolerance checks in float64 yet only 38 in bfloat16. A first-layer GPT-2 readout with fixed token and position meets the common-query condition without clamping activations; at three such heads, the capped optimum leaves 18.4\% to 74.2\% of the projected adapter effect uncompiled, with a head-dependent value--query ordering. All claims concern local approximation at one head, not whole-network equivalence.
cs.LG / 120 / 2609.36771
Beyond Conditional Independence: Root Cause Analysis with Deep Causal Models
Abstract
Root cause analysis (RCA) is a critical problem in many real-world scenarios. RCA enables the identification of faulty or failing mechanisms in a system by comparing anomalous observations with corresponding reference (i.e., regular) observations. However, existing approaches rely either on heuristic methods or on conditional independence tests with a strong unconfoundedness assumption, and thus fail to exploit other complicated distributional constraints in the presence of latent variables. To relax these assumptions, we model the underlying system as a causal model and the anomalous system as a change in the structural functions of the same causal model. Specifically, to handle unobserved confounders, we establish an implicit connection between distributional constraint testing and root cause analysis. To adapt our approach to data generated from arbitrary causal models, we employ the deep causal model (DCM) framework, in which we design the causal model using neural networks. Finally, we illustrate how our method, RCA-DCM, can utilize different levels of partial graphical knowledge to perform RCA. We evaluate RCA-DCM against state-of-the-art baselines on simulated datasets, a physics-based causal chamber and two micro-service applications. RCA-DCM improves top-1 accuracy over the strongest baseline on both Sock Shop (0.880 vs. 0.752) and Online Boutique (0.776 vs. 0.712), and when the true root cause in the causal chamber is unobserved and acts as a latent confounder, it recovers the exact root-cause set more often than any competing method (perfect recovery rate (PRR) 0.846 vs. 0.731).
cs.LG / 121 / 2609.36794
RAE-PPG: Duration-Grounded Retain-and-Extend Pretraining for PPG Foundation Models
Abstract
Signal features derived from photoplethysmography (PPG) require different signal durations to characterize. Existing PPG foundation models treat duration as a pretraining or evaluation condition rather than using the different durations required by PPG features to organize self-supervision. We hypothesize that self-supervision should expand with signal duration, allowing a single encoder to progressively acquire additional features while preserving and reusing earlier learning. We introduce Retain-and-Extend PPG (RAE-PPG), which trains a single Transformer encoder successively on 10 s, 30 s, and 240 s inputs, adding supervision for signal features supported by each longer observation. The encoder is partitioned into duration-specific parameter groups, allowing later stages to reuse earlier groups while updating only the group assigned to the current stage. Selected earlier targets are reused to supervise later stages, encouraging the corresponding features to remain accessible in longer-input representations. Direct decoding from the final encoder shows that earlier features remain recoverable from longer-input representations, while later-stage features show higher mean decoding performance at their introduction durations. Controlled comparisons further show that prior-stage learning provides a better basis for learning newly introduced features at both transitions. Across 18 tasks from eight datasets, the final frozen encoder achieves the best observed score on 12 tasks compared with five existing PPG foundation models.
cs.LG / 122 / 2609.36797
Where Does Randomness Matter in Neural Cellular Automata?
Abstract
Stochastic cell updates are often used throughout the life of a neural cellular automaton (NCA), from backpropagation through time to final rollout. This leaves two questions entangled: does update randomness help learn a useful rule, and must that randomness remain at execution? We separate training and evaluation update modes in controlled Growing NCA experiments, then vary the states shown during training. Under the standard constant-rate persist recipe, asynchronous training passes the short-horizon quality test in 10/10 runs, compared with 3/10 synchronous runs. All ten asynchronous models also retain the target for 4,096 steps under deterministic evaluation. For a scalar translation-invariant lattice, we derive an exact mean-square criterion: random masking can damp mean modes, but it also injects variance, and a mean-only test misclassifies four non-marginal settings. Finally, among 30 models that all pass the same reconstruction test, eight of ten grow-trained models become off-target at 4,096 steps, while all persist and regenerate models retain the target; damage recovery separates persist from regenerate. The results distinguish optimization reliability, execution mode, and task-specific behavior instead of treating them as one stability property.
cs.LG / 123 / 2609.36820
CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning
Abstract
Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward by its within-group standard deviation. The corresponding variance equals the sum of all pairwise reward covariances. For a fixed centered reward, larger aggregate covariance produces smaller advantages, and vice versa, allowing update magnitudes to adapt to reward dependence. However, correlated rewards with large scales can dominate this normalization and suppress signals from smaller-scale rewards. We propose Correlation-Normalized GRPO (CorrGRPO), which normalizes pairwise covariances into Pearson correlation coefficients. CorrGRPO keeps the centered total reward unchanged while balancing the influence of differently scaled rewards on the correlation-based normalization. This allows advantage magnitudes to adapt to reward correlations without the normalization being dominated by large-scale reward components. We compare CorrGRPO with GRPO and other variants on code generation, tool calling, and agent security, using models ranging from 0.5B to 8B parameters. These tasks all involve multiple rewards that can improve together or present tradeoffs. Results show improvements across three domains, including code generation, tool calling, and agent security. Our code is available at https://github.com/HKUST-KnowComp/CorrGRPO.
cs.LG / 124 / 2609.36843
RolloutFaith: Auditing Persistent Internal Interventions in Visual World Model
Abstract
Interpretability methods such as probes, activation patches and learned editors are designed to reveal or modify a model's current computation. World models pose a harder requirement: because their predictions become inputs to later predictions, a useful internal correction must survive after editing stops. We therefore propose RolloutFaith, a framework that measures semantic improvement both in the prediction produced at intervention time and over later autonomous predictions under fixed events, actions, noise, and information budgets. We evaluate ten fitted editors on three world models across Crafter, Cartpole, and CoinRun. We also use Reference Activation Patching, which replaces a model activation with the paired activation computed from the real observation, to measure the correction available at the chosen interface. This reference intervention improves later predictions in all nine model and task combinations and outperforms the best fitted editor in eight, yet its sustained gain decreases with horizon in five of nine combinations. Current fitted editors recover only limited and inconsistent long term effects. By restoring individual state components to their untouched values, we find that persistent effects travel through the newest generated frame in DIAMOND, recurrent memory in DreamerV3, and both in STORM. These findings suggest that training should reward future consequences. To test this hypothesis, we propose Delayed LoReFT, which optimizes the same low rank intervention through four frozen future transitions and improves sustained intervention effects to some extent.
cs.LG / 125 / 2609.36848
An Effective, Reliable, and Robust Framework for Human Activity Recognition Using Wearable Sensors
Abstract
Human Activity Recognition (HAR) through wearable sensors greatly improves the quality of human life through its multiple applications. For HAR, multi-sensor channel information is vital for optimal performance. Current work states that applying an attention neural network to prioritize discriminatory sensor channels helps the model classify activity more precisely. However, obtaining discriminatory information from multisensory channels is not always trivial, such as when collecting data from older hospitalized patients. In this context, existing HAR methods struggle to classify activities, particularly activities with similar natures. Moreover, HAR models predominantly suffer from overfitting due to the small size of available datasets, which leads to poor performance. Data augmentation (DA) is a viable solution to this problem. However, available DA methods have various drawbacks, including the possibility of being domain-dependent, resulting in distorted models for test sequences. To address these HAR problems, we propose a novel framework, ALAE-TAE-CutMix+, which focuses on two aspects. First, it enhances the latent information across each sensor channel and learns to exploit the relation among multiple latent features and the ongoing activity. Consequently, the discriminatory feature representations of each activity is enriched. Second, a new augmentation strategy is introduced to address the shortcomings of existing multi-sensor channel data augmentation. We then extend the framework to create a further enhanced version, namely ALAE-CIE-TAE-CutMix+, which learns to capture the interactions between the features of each pair of sensor channels. We find that although the first framework performs slightly better than the latter, the latter is nonetheless more reliable and robust. Both frameworks significantly outperform SOTA approaches on the four HAR datasets from diverse domains.
cs.LG / 126 / 2609.36859
Markovian Nonconvex ADMM for Reinforcement Learning: Bellman-Resolvent Stability Beyond Smooth Blocks
Abstract
We identify and study a structural mechanism for Markovian nonconvex ADMM in reinforcement learning. Using finite discounted MDPs as a canonical proving ground, we show that the discounted Bellman resolvent $(I-γP_π)^{-1}$ can provide the multiplier stability that classical nonconvex ADMM analyses often obtain from a designated smooth block. Starting from this mechanism, we establish convergence under controlled Markov sampling and then under stochastic observations using an empirical Bellman surrogate that jointly represents the random residual and its Jacobian. Markov mixing, initialization drift, observation noise, and decaying bias enter as one operator perturbation, avoiding unbiased product and double sampling requirements. When the perturbations are square summable, the true KKT residual converges almost surely to zero. Under a finite conditional fourth moment condition, a companion iterate satisfies $ \mathbb{E}[\widetilde G_{K+1}] \le A/T+(B/T)\sum_{k<T}m_k^{-1}, $ which becomes $O(T^{-1}+T/N)$ for total Markov sample budget $N$, giving $O(ε^{-1})$ iteration complexity and $O(ε^{-2})$ sample complexity for squared KKT accuracy $ε$. Beyond stationarity, discounted occupancy coverage yields $J^\star-J(π)=O(\sqrt G)$ for direct tabular policies, so covered exact KKT points are globally optimal, while a statewise quadratic Bellman-improvement condition sharpens the relation to $O(G)$. Finally, nonlinear policy, projected Bellman, and explicit occupancy formulations exhibit the same chain of operator invertibility, dual representation, and multiplier stability. This supports discounted operator invertibility as a reusable structural principle for primal-dual reinforcement learning.
cs.LG / 127 / 2609.36864
Where the Model Changes Its Mind: Hindsight-Divergence Localization for Efficient Reinforcement Learning with Verifiable Rewards
Abstract
Group-relative methods for reinforcement learning with verifiable rewards (RLVR) learn from differences in rollout outcomes. Independently sampling complete trajectories is costly and does not explicitly explore the decision space at critical positions. Feedback on a completed trajectory can reveal which earlier choices the policy reconsiders, suggesting where to sample alternative continuations. We introduce Hindsight-Divergence Localization (HDL), which uses hindsight-induced changes in token log-likelihoods to select branch points. HDL generates a small number of complete root trajectories and fills each training group with continuations from the selected positions under the original task context. Each continuation reuses its root prefix and contributes policy updates only through its newly generated suffix, reducing generation cost while focusing additional exploration and learning on decisions after branching. Experiments with three models across math, code, and agent tasks show gains in both rollout efficiency and task performance. Compared with GRPO at matched group sizes and training steps, HDL yields up to a 2.5$\times$ reduction in generated tokens and a 1.8$\times$ speedup in rollout wall-clock time. Despite this reduced generation budget, HDL improves performance across all three domains, with gains of up to 12.5 points on agent tasks.
cs.LG / 128 / 2609.36873
Seeing Time: Visual-Temporal Representation Learning for Interpretable Time Series Clustering
Abstract
Multivariate Time Series (MTS) clustering is an important tool in temporal data mining, aiming to discover latent group structures from complex observations without supervision. Although existing deep clustering methods can learn discriminative temporal representations, the resulting latent clusters are often difficult to relate back to waveform characteristics that practitioners can directly inspect and compare, limiting their ability to assess whether the discovered patterns reflect meaningful temporal behaviors. This paper, therefore, proposes WAVE (Waveform Aligned Visual-temporal Embedding), which treats time series and their deterministically rendered waveform plots as complementary views of the same observations. To produce discriminative representations whose cluster structures can be traced to observable waveform characteristics, WAVE aligns and integrates fine-grained temporal variations with holistic visual patterns, while associating each discovered cluster with its centroid-nearest authentic sample. Accordingly, interpretability in this work specifically refers to waveform-level traceability rather than a general explanation of model decisions. Extensive evaluations across 10 real-world public datasets show that WAVE achieves the highest macro-averaged clustering performance and the best average rank among the compared methods, while qualitative case studies illustrate how the discovered clusters can be inspected through authentic waveform records. The source code is available at https://github.com/Zheng-Zhu1/WAVE.
cs.LG / 129 / 2609.36881
What You Observe Determines How You Identify Causal Effects: Evaluating Causal Models across Observational Views
Abstract
Causal foundation models (CFMs) pre-trained on data generated from various structural causal models (SCMs) have been proposed for estimating causal effects from observational data. However, differences in pre-training environments and evaluation protocols make it difficult to assess how their performance depends on the information available for causal identification. To enable controlled comparisons, we introduce CausalIDView, a multi-view benchmark that holds fixed SCM realization and target estimand while varying only the observational view available to the estimator. Each observational view corresponds to a distinct identification regime under the benchmark's maintained causal assumptions. Across these matched views, no CFM consistently performs best and model rankings vary substantially. Under controlled structural changes, CFMs exhibit model-specific failures to maintain stable estimates when true effects are unchanged and to track genuine effect changes. We also examine whether combining explicit identification with strong predictive estimation is effective. A modular approach that pairs a predictive tabular foundation model with regime-specific identification procedures is competitive with CFMs and outperforms several of them. These findings motivate cross-regime comparisons to assess the empirical value of CFMs.
cs.LG / 130 / 2609.36883
Architecture Alignment With Sparse Priors in Tabular Foundation Models
Abstract
Tabular foundation models (TFMs) are increasingly popular because they deliver strong predictions on new datasets through in-context learning, without task-specific training or extensive tuning. Yet released TFMs differ simultaneously in their pretraining priors, architectures, and objectives, obscuring their respective inductive biases. We therefore examine one concrete capability: irrelevant-feature suppression. Across synthetic tasks and real-world datasets, adding null features causes substantially greater predictive degradation in the row-token model TabDPT, whereas the cell-token alternating-axis model TabPFN v2 and other TFMs remain comparatively stable. This gap motivates us to ask whether architecture contributes to irrelevant-feature suppression. Because released TFMs remain confounded by other design choices, we train streamlined row-token and alternating-axis transformers under identical sparse-to-dense linear priors. Exact Bayes analysis shows that sparse prediction requires context-dependent feature gating, whereas the dense endpoint requires only uniform feature weighting. Consistent with this distinction, the alternating-axis model is substantially closer to the Bayesian optimal predictor on sparse tasks, while the architecture gap becomes negligible on dense tasks; almost all of the sparse gap arises from linear coefficient-estimation error. Finally, in both the controlled model and frozen TabPFN v2, we examine the effect of interventions on the feature-attention outputs on the linear coefficients, finding evidence of task-dependent selective routing of computation through feature-indexed pathways. Together, these results support architecture-prior alignment: preserving an addressable feature axis provides an inductive bias for task-adaptive relevance inference. Code is available at https://github.com/Tianqi-Zhao/ArchitecturePriorTFMs.
cs.LG / 131 / 2609.36890
SINO: Scale-Invariant Neural Operator
Abstract
In scientific machine learning, physical fields governed by partial differential equations exhibit low-rank structure and scale invariance. When solving equations on coarse grids, missing information leads to the closure problem: modeling unresolved physics to recover lost dynamics. Although closure terms depend on grid resolution, they represent scale-invariant physical laws. A model truly learning physics should capture these mechanisms with low-rank parameterization rather than memorizing grid-specific patterns. Inspired by this, we propose the Scale-Invariant Neural Operator (SINO), which learns on normalized physical scales via a dual-branch architecture operating in spectral and spatial domains. SINO uses bottleneck MLPs to generate continuous convolution kernels, embedding an explicit low-rank inductive bias that concentrates more than 95 percent of variance in 2-3 modes, as validated by PCA across benchmarks, while drastically reducing parameters. This principled design yields 38 times steeper scaling law exponents than FNO, demonstrating superior parameter efficiency. We compare SINO with traditional models (U-Net, DeepONet), Transformer models (Transolver, Oformer, GK-Transformer), and frequency-domain models (FNO, AMFNO, UFNO) on closure problems spanning externally forced Burgers turbulence, decaying Burgers turbulence, KS turbulence, Kolmogorov-forced NS turbulence, and decaying NS turbulence. Experiments show SINO achieves 1.5-38 times error reduction and 2-23 times parameter efficiency over baselines, with superior scaling laws reflecting exceptional data efficiency from principled low-rank design. Code is available at https://github.com/AI4Science-WestlakeU/SINO.
cs.LG / 132 / 2609.36911
Variational Mixtures and Multi-Marginal Flow Matching: Advancing Statistical Inference with Biological Applications
Abstract
In this thesis I develop methods for statistical inference when the distributions arising from complex biological systems are multi-modal, geometrically structured, and sometimes only defined up to a normalizing constant. I start from variational inference and, when analytic update equations are unavailable, move to black-box variational inference. To build intuition regarding inference challenges and the proposed methodologies, I introduce a novel unnormalized target density (the CoLN distribution) and reuse it as a controlled test case in the kappa. I then trace a trajectory of increasingly expressive approximations: ensembles evaluated with the multiple importance sampling ELBO (Paper A) and variational mixtures that automate component cooperation and exploration (Paper B). Because expressivity comes at a cost, I develop efficient mixture learning ideas, including Monte Carlo objective estimators to scale mixture learning more efficiently (Paper C). As a new result in the kappa, I overturn a three decades long misconception regarding the potential performance benefits of using mixtures in variational inference. Finally, I move from variational inference to flow matching, where I address the need for specialized treatment of interpolant learning in multi-marginal settings (Paper D). By combining insights from Papers A-D, I derive in Section 5.5 a new method: multi-marginal flow matching with mixtures of variational interpolants. I connect these methodological developments to biological applications, with special emphasis on three-dimensional spatial transcriptomics, where stacked tissue slices induce multi-modal dynamics across space.
cs.LG / 133 / 2609.36926
State Transport Routing for Short-horizon Adaptation in Multi-horizon Photovoltaic Forecasting
Abstract
Recent power measurements provide valuable information for photovoltaic(PV) power forecasting, but directly extrapolating short-term trends can introduce substantial errors over longer forecast horizons. To address this challenge, we propose state transport routing (STR), a lightweight adapter that refines the predictions of a frozen forecasting model. STR combines the original forecast with two complementary trajectories derived from the latest measured power level and its recent trend. A horizon-conditioned router adjusts their contributions over the first 120 min, while leaving subsequent predictions unchanged. Experiments on four public PV datasets show that STR consistently outperforms a parameter-matched residual adapter. On PVDAQ, the same approach improves five neural forecasting backbones, reducing all-horizon normalized mean absolute error by 0.0201-0.2364 percentage points, with paired 95% confidence intervals excluding zero. No reliable improvement is observed for LightGBM. These findings demonstrate the potential of structured state adaptation to improve short-term forecasting across different neural architectures without retraining the underlying models or altering their longer-horizon predictions.
cs.LG / 134 / 2609.36942
Safe-by-Design Learning via Energy-based Neural Networks
Abstract
Learning neural-network models of dynamical systems with safety guarantees is a fundamental requirement for their deployment in safety-critical settings. Safety is commonly established by proving the invariance of a desired subset in state-space, ensuring that every trajectory initialized in this subset remains confined to it for all time under admissible inputs. Existing frameworks, however, either rely on computationally expensive post-hoc verification or employ safety-enforcing mechanisms without formal correctness guarantees. In this paper, we introduce a novel neural architecture grounded in energy-based modern Hopfield networks to guarantee safety-by-design while retaining sufficient expressiveness to model complex nonlinear dynamics. Specifically, we integrate modern Hopfield networks with a port-Hamiltonian neural ODE, enabling by design the construction of barrier functions yielding explicit admissible-input sets and quantitative robustness radii. Across several benchmarks, including an 12-dimensional nanodrone model, our framework achieves state-of-the-art performance while producing certified invariant sets that are more robust to external solicitations than comparable existing approaches.
cs.LG / 135 / 2609.36953
Cool the Sampler, Not the Learner: Sampling Temperature Moves the Staleness Cliff of Importance-Corrected GRPO
Abstract
Production RL for language models lets the sampler fall behind the learner and repairs the resulting mismatch with a truncated importance weight. We ask how long the sampler can go without a refresh under that correction, and find a cliff: on Qwen2.5-Math-1.5B and GSM8K, importance-corrected GRPO refreshed every 192 updates learns well for 180 steps and then degrades severely in all three data seeds before the refresh arrives. Published remedies for staleness act on the update; we act on the sampler instead. Decoupled cooling draws samples at temperature 0.8 while the learner, the reference model and the importance weights stay at temperature 1, with the behaviour probability recorded from the tempered distribution, so the learner's objective is unchanged. All corresponding cooled runs are stable, and the longer interval keeps what the short one delivered: at the same update budget, a cooled sampler refreshed every 192 steps matches an uncooled sampler refreshed every 96 at the end of training (0.857 for both) and averaged over it (0.79), whereas lowering the learning rate to a safe value ends 3-7 points lower. On Qwen2.5-Math-7B the degradation points at interval 192 predict that an interval of 144 is fatal without cooling and survivable with it; on two data seeds the uncooled runs degrade before their first refresh and the cooled runs pass it and end at 92-93% against 68-81%, with one cooled run degrading transiently late in the second cycle. The benefit has a window: at three times the safe interval and in a high-mismatch MATH setting cooling delays degradation without preventing it, stronger cooling is not better, and cooling without the correction collapses. Sampling temperature is a control on staleness tolerance, and temperature and refresh interval should be chosen together.
cs.LG / 136 / 2609.36958
VStress: Correlation-Aware Auditing and Adaptive Budget Allocation for Repeated Verifiers
Abstract
Repeated verifier calls are useful only when they contribute conditional information. We introduce VStress, an auditable replay contract, and VStress-CA, a correlation-aware allocation policy that estimates the conditional marginal information of an unqueried verifier on a sealed calibration split, discounts uncertainty, normalizes by call cost, and stops or abstains when the next call is not informative. The controller freezes its decision and cost ledger before joining the clean oracle; a dependence-shift alarm disables channel preference and falls back to exact-stop. The controlled audit gives the mechanism boundary: at 35% symmetric corruption, majority-5 improves balanced accuracy from 0.6578 to 0.7739, whereas at 65% it loses 0.1226 points. In the matched fixed-budget comparison, breadth, redundancy, and adaptive allocation obtain balanced accuracies 0.6048, 0.6375, and 0.6538, with 3.4216 calls per item and an RLVR score of 0.6417 for VStress-CA. Dependence diagnostics also increase from same-model repeats to cross-family channels, with conditional marginal gains of 0.0126, 0.0462, and 0.0913. These measurements turn correlation from a post-hoc warning into an auditable allocation decision.
cs.LG / 137 / 2609.36966
JudgeCast: Time Series Forecasting with Experience-Informed Covariate Judgements
Abstract
Covariate effects vary across contexts and shift over time, requiring forecasters to assess how to use them for each forecasting context. As forecasting proceeds, observations for earlier forecasts become available, providing feedback on past covariate use for subsequent forecasts. However, when multiple covariates act together, the forecast error reveals the numerical discrepancy from the observation but not how the covariates should have been used. We introduce JudgeCast, an experience-based framework for time series forecasting with covariates. Following the judgmental adjustment practice, a frozen TSFM provides the base forecast, while a frozen LLM uses the current context and relevant experience to adjust it. Within the adjustment, assessing covariate effects and determining the numerical adjustment serve distinct roles, so JudgeCast first forms explicit covariate-wise judgments and then determines the adjustment. After observation, JudgeCast uses the observed residual of the base forecast to reconstruct alternative judgments and evaluates the original and alternatives through their resulting adjustments. The best-performing decision is selected and retained as validated experience for subsequent forecasts. Across diverse real-world datasets, JudgeCast outperforms strong baselines. Ablations show that explicit covariate-wise judgment can improve forecast-time adjustment, while residual-guided experience construction yields more reliable forecasting gains than retaining raw decisions as experience.
cs.LG / 138 / 2609.36968
TaskBridge: Bridging Unsupervised Tabular Anomaly Detection and In-Context Learning via Virtual Tasks
Abstract
Unsupervised tabular anomaly detection (TAD) aims to identify anomalous rows in tabular data using normal training samples. While conventional methods rely on dataset-specific training and configuration search, recent tabular foundation models (TFMs) enable zero-shot anomaly detection on unseen datasets via in-context learning. Most TFM-based approaches, however, require anomaly-specific pretraining from scratch, making detection inherently dependent on synthetic TAD-specific priors and costly to update. Some approaches instead repurpose pretrained general-purpose TFMs for TAD to avoid this burden, but rely on computationally expensive formulations with restrictive anomaly inductive biases. In this work, we introduce TaskBridge, a new framework that efficiently repurposes pretrained general-purpose TFMs for unsupervised TAD by constructing virtual supervised tasks that directly recast anomaly detection as supervised in-context inference of TFMs. The resulting virtual tasks induce predictive structures under which normal queries and their target pairs receive high support, whereas anomalies tend to violate the induced structures and receive lower support, providing direct anomaly evidence. Across 790 real-world datasets, TaskBridge consistently outperforms 30 baselines, including state-of-the-art TFM-based approaches, without anomaly-specific TFM pretraining or dataset-specific model optimization.
cs.LG / 139 / 2609.36970
Equally Good, Yet Different: Benchmarking Rashomon sets in AutoML packages
Abstract
The Rashomon effect describes the existence of multiple near-optimal models that achieve comparable performance while offering fundamentally different explanations. This creates a critical vulnerability in AutoML: x-hacking, the selective post-hoc choice of a model based on its explanation rather than predictive merit. No existing AutoML framework exposes this risk. We introduce ARSA ML, an open-source Python framework that quantifies Rashomon set structure and predictive multiplicity within AutoML pipelines. Using ARSA ML, we benchmark AutoGluon and H2O across 28 binary classification datasets, and conduct a post-hoc x-hacking analysis revealing a consistent structural asymmetry: AutoGluon produces larger, diverse sets with stable explanations, while H2O generates compact sets with markedly higher prediction divergence and explanation instability -- making H2O users considerably more exposed to x-hacking. This gap persists across all evaluated metrics and epsilon thresholds, pointing to a fundamental difference in each framework's model-building strategy. ARSA ML is available at https://pypi.org/project/arsa-ml/ .
cs.LG / 140 / 2609.36985
Abductive World Modeling via Causal Representation Learning
Abstract
The central challenge of world modeling is to learn representations that capture how the world evolves. However, existing world models predominantly represent future states without explicitly capturing the latent causes underlying their evolution, limiting their ability to reason about why and how the world changes. To address this limitation, we propose Abductive World Modeling (AWM), a framework that learns structured causal representations by abductively inferring latent causes from predicted futures. Specifically, we realize AWM through the Hierarchical Abductive State Pyramid (HASP), which organizes the inferred world state into three complementary components - Entity, Dynamic, and Relation - capturing what exists, how it changes, and how entities interact, respectively. By jointly reasoning over the current observation and its predicted future, HASP abductively infers these latent factors and integrates them into a structured state representation for downstream reasoning. To the best of our knowledge, AWM is the first framework to introduce abductive state inference into latent-space world modeling for learning structured representations of world dynamics. Experiments across physical prediction, causal reasoning, and action understanding demonstrate the effectiveness of our approach. Compared with V-JEPA, a state-of-the-art latent-space world model, AWM improves physical prediction AUROC by 10.7%, causal reasoning accuracy by 16.8%, and action Top-1 accuracy by 68.0%.
cs.LG / 141 / 2609.37037
High-Resolution Dynamic Functional Connectivity Generation with Graph-Variate Flow Matching
Abstract
High-resolution dynamic functional connectivity (DFC) can reveal rapidly evolving brain-network interactions, but short temporal windows yield noisy, often low-rank covariance estimates. Graph-Variate Dynamic (GVD) connectivity addresses this by modulating fast instantaneous interactions with stable trial-level support. This suppresses spurious fluctuations and emphasizes persistent, informative connections. We show that the Hadamard construction lifts low-rank instantaneous connectivity from the positive-semidefinite to the positive-definite cone, keeping high-resolution trajectories on the SPD manifold without ridge regularisation or post-hoc projection. We introduce GVD-CFM, a class-conditional generative model for high-resolution dynamic connectivity. Each trial is represented as SPD GVD matrices on a product Riemannian manifold, then mapped through a global log-Euclidean diffeomorphism and an invertible temporal DCT basis. A Transformer-based conditional flow models all spectral modes jointly and generates the full trajectory non-autoregressively in Euclidean coordinates while preserving exact correspondence with valid SPD sequences. Retaining the full DCT basis also enables decoding on denser temporal grids without retraining. Across multiple EEG motor-imagery datasets, GVD-CFM delivers the strongest overall results for held-out distributional fidelity, temporal-dynamics preservation, and synthetic-to-real classification. It also remains computationally efficient relative to strong raw-signal and direct GVD-space generative baselines. GVD-CFM therefore provides a practical framework for realistic, temporally coherent, high-resolution brain-network generation with preserved manifold structure and resolution-flexible decoding from a single trained model.
cs.LG / 142 / 2609.37041
Beam Search as Test-Time Self-Distillation via Counterfactual Contexts
Abstract
Self-Distillation Fine-Tuning (SDFT) enables a language model to act as its own teacher: by conditioning on a demonstration, the model produces an implicit reward via pointwise mutual information, which guides on-policy learning without external supervision. However, SDFT operates at training time: it requires gradient updates and access to expert demonstrations, making it inapplicable at inference. We propose test-time self-distillation, a decoding-time method that extracts a steering signal from the self-distillation framework without any parameter updates, reward models, or training data. Our key insight is that counterfactual contexts, i.e. fixed textual templates that hypothetically prime the model for excellent versus poor reasoning, can substitute for the demonstration. The log-odds ratio of a candidate answer under these two counterfactual conditions defines a new reward signal. We derive the optimal KL-regularized policy under this reward, which takes the form of a Gibbs reweighting of the base distribution. Crucially, this reweighting is global: it cannot be decomposed into independent per-token operations without ignoring future trajectory quality. We therefore approximate the target distribution via beam search. Experiments on mathematical reasoning (MATH500), code generation (HumanEval), and graduate-level science QA (GPQA) across multiple model scales show that test-time self-distillation improves over standard sampling, low temperature, beam search and power sampling baselines on average, demonstrating that the self-distillation principle can be operationalized at inference time.
cs.LG / 143 / 2609.37043
Iterative Exact Discrete Guidance for Energy-Based Sampling
Abstract
Sampling from unnormalized distributions over large discrete state spaces becomes difficult when a multimodal target is far from a tractable reference. We introduce Iterative Exact Discrete Guidance (IEDG), a population-exact, trajectory-wise guidance framework for unnormalized discrete targets. Rather than learn the full reference-to-target correction in one step, IEDG introduces a global Boltzmann tilt along an annealing trajectory. Each stage learns a stage-local posterior correction for an incremental Boltzmann tilt of the current source, while the resulting corrections are accumulated relative to a fixed analytic posterior. At the population optimum, exact stage posteriors recover the correct reverse dynamics, whose exact simulation reproduces the target distribution. IEDG chooses stage increments by relative effective sample size (rESS), which controls Rényi-2 displacement and locally adapts the step size to the thermodynamic geometry of the annealing path. Our stagewise total-variation analysis shows that limited overlap amplifies Bregman fitting error by $1/\sqrt{\mathrm{rESS}}$, while posterior, simulation, and truncation errors enter additively. IEDG improves all distribution-level errors over the neural baselines on ordered, exactly enumerated Ising $4\times4$, while substantially reducing one-shot errors on Ising/Potts $16\times16$ across thermodynamic regimes and attaining the best neural-sampler result on several reported local-statistic and phase-coverage metrics. On Max-Cut, its best-of-512 and average-sample ratios exceed all the baselines. Code and artifacts are available at https://github.com/StillFantasy123/iterative-exact-discrete-guidance.
cs.LG / 144 / 2609.37057
Message Passing Does More with Less for In-Context Learning on Graphs
Abstract
Achieving strong performance with graph neural networks (GNNs) typically requires training and hyperparameter tuning for each dataset, incurring repeated costs and effort. Graph in-context learning (ICL) avoids this by using a single pretrained model to predict unknown node labels directly from labeled context nodes. Existing approaches, however, rely on dense attention across nodes, making inference increasingly expensive as graphs grow. In this work, we present Ephris, a new graph in-context learner built on sparse message passing, scaling linearly with the number of node-feature entries and graph edges. Ephris is pretrained entirely on synthetic graphs generated from structural causal models with diverse graph structures and relational dynamics, exposing the model to varied dependencies among topology, features, and labels. We evaluate Ephris on 51 node-classification datasets against 15 extensively tuned GNNs and existing graph ICL methods under both high- and low-label train/validation/test splits. Across both settings, Ephris ranks first on all four aggregate measures: Elo, improvability, average rank, and accuracy. Its inference cost remains comparable to training a single GNN once, while being over 10 times faster than previous graph ICL models. Together, these results advance the performance-runtime Pareto frontier, demonstrating that strong graph ICL does not require dense attention. Code and model weights are available at https://github.com/nums-ai/ephris.
cs.LG / 145 / 2609.37061
A Comprehensive View of Fairness through Distributional Stability
Abstract
We view fairness as a property of distributional stability. Rather than assessing a predictor under a fixed data distribution, we study how its predictions change under perturbations that modify the composition of protected groups. A predictor is fair if it remains stable under such shifts. Under this perspective, several classical notions of fairness arise as stability with respect to specific perturbations, with the associated unfairness gap given by a Lipschitz constant of a prediction-rate functional. This formulation also yields guarantees that hold uniformly over a range of demographic compositions at test time, without requiring knowledge of the deployment distribution. It leads to a learning procedure based on convex combinations of reweighted predictors, formulated as a second-order cone program, for which we establish generalization bounds. Experiments on standard benchmarks illustrate the approach.
cs.LG / 146 / 2609.37065
RL-PaO: Prediction as Action in Decision Making under Uncertainty
Abstract
Decision-making under uncertainty often relies on predicted parameters, yet accurate prediction does not necessarily lead to good operational decisions. Aligning prediction with downstream optimization requires learning from the consequences of the decisions those predictions induce. We introduce RL-PaO, a reinforcement learning framework that integrates system formulation, optimization, and decision execution into a single environment. This yields a Markov decision process in which prediction is regarded as action: it shifts the environment to produce subsequent context and reward that explicitly aligns prediction error with realized cost, and learning the optimal policy does not require differentiating through the black-box solver. We evaluate RL-PaO on day-ahead energy scheduling using real historical data. On the test year, RL-PaO achieves the lowest annual cost among the non-oracle baselines, achieving on average $10\%$ cost reduction. Moreover, RL-PaO is capable of further analyses to provide strong interpretability both from the policy evolution perspective and the cost-accuracy trade-off.
cs.LG / 147 / 2609.37083
Identifying ODEs from Unstructured Data with Causal Representation Learning
Abstract
We study the problem of recovering the governing ODE of a dynamical system from unstructured, high-dimensional observations such as images. Existing methods for ODE discovery typically assume direct measurements of the variables, or do not provide theoretical guarantees on the learned variables and equations. While Causal Representation Learning (CRL) methods provide guarantees on identifying variables from high-dimensional observations up to component-wise diffeomorphisms, we show that in general these variables cannot be used directly as input to equation discovery methods, which typically assume that the variables will lead to sparse equations. So we introduce SParse Equivalent Equation Discovery AutoEncoder (SPEED-AE), a framework that combines a pretrained CRL method with a component-wise autoencoder that learns transformations of variables that are amenable to sparse ODE discovery. We show that for polynomial ODEs, this additional step allows us to restrict the identifiability of each variable from polynomial to monomial diffeomorphisms. Experiments on Lotka-Volterra, Lorenz, and a two-pendulum system show that SPEED-AE improves on the disentanglement of the CRL methods and that it recovers ODEs that are closest to the ground truth, while achieving state-of-the-art forecasting performance.
cs.LG / 148 / 2609.37105
VACE: Validation-Gated Alternating Co-Evolution of Agent Models and Harnesses
Abstract
Language model agents can be improved by updating their model weights or refining the harness that guides task execution. These components are coupled: weight updates change how the model uses the harness, while harness updates change the trajectories used for training. We propose VACE, Validation-Gated Alternating CoEvolution, which alternates agentic reinforcement learning with trajectory-driven harness refinement. After each RL stage, VACE reuses the collected trajectories to propose a harness revision and evaluates the incumbent and candidate with the updated model held fixed. The candidate guides subsequent training only if it improves validation performance. With Qwen3.5-9B, VACE achieves 45.26% test accuracy on OfficeQA and a mean partial-credit score of 75.19% on AutomationBench, exceeding weight-only RL by 6.43 and 9.09 percentage points and ungated alternation by 4.59 and 6.95 points, respectively. Across 44 harness proposals, 17 reduce validation performance at the updated checkpoint and are rejected before subsequent RL training, highlighting the importance of validation gating.
cs.LG / 149 / 2609.37114
Interpretable intrinsic dimension estimation through componentwise calibration of distance and angle
Abstract
DANCo (Dimensionality from Angle and Norm Concentration) jointly calibrates nearest-neighbor distance and angular statistics and consistently reaches state-of-the-art accuracy on clean intrinsic-dimension (ID) benchmarks. Practical data, however, introduce neighborhood-relative noise and sample-amplitude heterogeneity that can distort these geometric signals. We reformulate DANCo componentwise, retaining separate distance and angular discrepancy curves so that the source of an estimate can be identified and interpreted. For the distance component, we derive a closed-form Kullback-Leibler divergence for the generic-order ratios of the generalized ratios ID estimator (Gride); when both angular parameters are matched (Full), Gride reduces mean percentage error from $27.7\%$ to $17.6\%$ at noise equal to $40\%$ of typical neighbor spacing on 24 manifolds. For the angular component, two sampling regimes motivate aligning mean direction while retaining concentration matching (Profiled). On a Gaussian scale mixture with generating dimension 70 embedded in 100 dimensions, profiling raises the Minimum Neighbor Distance (MiND) estimate from $22.8$ to $66.7$, while removing the known amplitudes restores MiND-Full to $71.9$; the control thus attributes the Full shortfall to amplitude heterogeneity. On CIFAR-10 and ImageNet, amplitude-reducing normalizations move angular location toward the references and narrow the Full-Profiled gap, an observational counterpart to the controlled mixture. Across four pretrained convolutional neural networks, Gride-Profiled, the two-nearest-neighbor estimator (TWO-NN), and the maximum-likelihood estimator (MLE) exhibit similar rise-and-fall profiles, while Full-Profiled differences identify the layers most sensitive to angular calibration.
cs.LG / 150 / 2609.37122
Learning the Structure of Triangular Transport Maps
Abstract
Triangular transport maps provide a flexible approach to sampling-based probabilistic modeling, including density estimation, generative modeling, and Bayesian inference. They transform an unknown target distribution into a simpler reference through a monotone triangular map. The map structure is defined by a variable ordering and sparsity pattern, which together encode a directed acyclic graph. Map quality can depend strongly on this structure, yet finding a good structure is computationally expensive because each candidate generally requires fitting a different map. A central challenge is therefore to learn density and structure jointly, while keeping computation manageable as dimension grows. We introduce Self-Structuring Transport Maps (SSTM), which learn the map, ordering, and sparsity jointly. We use SoftSort to learn the variable ordering and $L_0$ gates to learn the sparsity, while preserving a triangular structure. To keep the map scalable, we use a monotone BatchEnsemble that shares one weight matrix across all map components through rank-one adapters. Across synthetic and real data, jointly learning the structure and map gives better density estimates than estimating the structure first. When the structure is identifiable from the density, SSTM matches the density performance of a map fitted with the true structure and outperforms autoregressive flows. On large datasets, SSTM is competitive with autoregressive flows.
cs.LG / 151 / 2609.37148
Multimodal Detection of Higher-Order Behavioral Constructs: Self-Compassion in Structured Reflective Interaction
Abstract
Many of the qualities that matter most in how people learn and grow, how someone regulates their emotions, reflects on a setback, or stays aware of others during a difficult conversation, are not directly observable. They have to be inferred from how someone speaks, moves, and sounds over time, and they resist the kind of clean labeling that most machine learning pipelines are built around. We study this challenge through a case that is well grounded in psychological theory but rarely modeled computationally: self-compassion, the tendency to respond to one's own setbacks with patience rather than harsh self-criticism. We examine how it appears during structured reflective interviews in a technology-mediated training setting, where people naturally talk through socio-emotionally demanding situations. Since no existing dataset captures this kind of construct in this kind of setting, we collected and annotated 51 reflective dialog sessions using an independent, temporally overlapping annotation scheme grounded in established theory. We consolidate the underlying six-component psychological model into a three-class supervision space, balancing self-kindness and mindfulness against self-critical or overwhelmed states, and build a reproducible window-based pipeline that aligns video, audio, and text on a shared timeline. Unimodal models trained on each modality separately are compared against a simple probability-level fusion strategy, which yields modest but consistent gains over the best single modality. We close by discussing where each modality succeeds or struggles, what this suggests about how this kind of construct is actually expressed in reflective speech, and what would be needed to model it, and constructs like it, more effectively.
cs.LG / 152 / 2609.37156
Lucid Dreaming for World Models: Learning to Doubt Imagination and Decide by Trust
Abstract
World models enable agents to learn and plan in imagination, but predictions beyond their experience can become unreliable and mislead decisions. Existing uncertainty estimates derived from predictions can remain overconfident on unfamiliar state-action pairs. We propose the Lucid World Model (LucidWM), which learns doubt from experience and propagates trust through imagination. By integrating Subjective Logic into categorical latent transitions, LucidWM distinguishes predicted outcomes from their evidential support and assigns each transition a degree of doubt. The complement of this doubt defines transition-level trust, which accumulates multiplicatively along imagined trajectories to reweight returns for policy learning and guide action selection. Uncertainty estimation requires no additional parameters or forward passes. Evaluated on four base world models against seventeen uncertainty readouts, LucidWM detects environmental changes and signals uncertainty during action-corrupted rollouts. In a controlled navigation case study, acting on trust reduces the number of steps required to reach the goal from 362 to 190. Fifteen demonstration videos show how LucidWM doubts its dreams and acts on that doubt. Videos are available at https://lucidwm.github.io.
cs.LG / 153 / 2609.37158
GLASS: Global Latent Aggregation with Slot-based Set Decoding for Scalable All-Atom Crystal Generation
Abstract
Generative models for crystals enable the discovery of novel structures, but scaling all-atom generation to larger systems such as metal--organic frameworks remains challenging. We connect this difficulty to the correspondence problem of particle-space generation. Even on a single fixed target set, index-free permutation-equivariant particle flows require substantially more training for reliable generation as set size and density increase, under both independent and optimal-transport couplings. To resolve this challenge, we introduce GLASS---Global Latent Aggregation with Slot-based Set Decoding, which encodes structures in a permutation-invariant global latent space and learns their distribution via flow matching. A learned-slot decoder constructs all atoms in parallel, removing atom-wise correspondence from generative transport. On MP20, GLASS is competitive with particle-space models, and flow training can reach the validity of the training data at every structure size. On a QMOF subset, GLASS generates MOFs with up to 150 atoms per unit cell without conditioning on building blocks, topology, or composition, and approaches the structural validity of the training data. On both datasets, flow training exposes a validity--novelty tradeoff, and MOF novelty remains limited by autoencoder generalization on the available data. These results show that separating correspondence assignment from generative transport provides a simple route toward high-validity generation of larger atomistic systems.
cs.LG / 154 / 2609.37161
The Vote Hides the Failure: Aggregation Choice and Noise Robustness in Heart Murmur Detection
Abstract
Noise robustness in automated phonocardiogram (PCG) murmur detection, and how it is measured, remains underexamined despite growing interest in low-resource screening. We evaluate two independently reimplemented pipelines, Hierarchical Multi-Scale Convolutional Network (HMS-Net)--CNN, and Bidirectional Long Short-Term Memory (BiLSTM)--LSTM, under controlled, multi-severity noise with noise-augmented fine-tuning and held-out generalization testing. Under matched aggregation, the complete BiLSTM pipeline outperforms the complete HMS-Net pipeline across all conditions in accuracy and Weighted Accuracy. A stable aggregate accuracy score can misrepresent what individual predictions show: HMS-Net's native aggregation degrades under salt-and-pepper noise far less than majority-vote (MV) aggregation at the same severity, a gap reflecting window-level disagreement its native rule absorbs, while BiLSTM's MV accuracy rises after noise-augmented training even though its individual predictions do not improve. HMS-Net's training effect is significant under one accuracy metric but not another. Noise-robustness conclusions can depend as much on evaluation choices as on the models themselves.
cs.LG / 155 / 2609.37170
Interpolated Policy Distillation: A Controllable Continuum Between Off-Policy and On-Policy Distillation
Abstract
Off-policy and on-policy distillation have traditionally been formulated as separate paradigms, each favoring a different property of distillation trajectories. Teacher-generated (off-policy) traces are typically high-quality but lie far from the student's distribution, whereas student-generated (on-policy) rollouts are more learnable but often contain erroneous reasoning. We view these paradigms as the endpoints of a policy continuum and posit that a more effective rollout policy may lie in between. We introduce \textbf{Interpolated Policy Distillation (IPD)}, which defines the next-token distribution at every decoding step as an explicit linear interpolation between the student and teacher distributions. The interpolation operates at the distribution level, token by token, and its coefficient provides direct control over the balance between trajectory quality and student learnability. Naively sampling from this policy would require sequentially querying the teacher at every token and is thus expensive. To make IPD practical, we accelerate it with a new speculative-decoding rule while exactly preserving the interpolated next-token distribution.At the trajectory level, the resulting rollouts naturally interleave student- and teacher-generated segments. Unlike recent heuristic segment-interleaving methods, however, this interleaving is induced by an exactly realized token-level interpolated policy rather than by hand-designed switching rules. Across text-only and multimodal reasoning benchmarks, IPD consistently outperforms both endpoint policies (SFT and OPD), their conventional two-stage combination (SFT-then-OPD), and recent heuristic segment-interleaving methods, demonstrating that token-level policy interpolation better balances trajectory quality and student learnability.
cs.LG / 156 / 2609.37189
Corruption-Robust Sparse Linear Contextual Bandits with Knapsack Constraints
Abstract
We study sparse linear contextual bandits with knapsack constraints under joint reward and consumption corruption. Consumption corruption creates a challenge beyond corrupted rewards: it affects not only statistical estimates, but also the recorded budget, resource prices, and stopping decisions that govern future allocation. We develop Robust Optimistic Primal--Dual (ROPD), an estimator-modular framework that combines corruption-aware confidence widths with online resource prices and a budget-safety rule. With concrete sparse implementation, ROPD achieves regret against a clean population-LP benchmark of $\widetilde O(T^{2/3}+ΓT^{1/3})$ under forced exploration and population-design coverage, and $\widetilde O(\sqrt T+Γ)$ under on-policy realized-design coverage, for a supplied valid corruption bound $Γ$ under the stated proportional-budget scaling and fixed model/design parameters. When the corruption level is unknown, Shared-Grid adapts confidence radii around common point estimates fitted to a single realized history, incurring explicit initialization and master-comparison costs; its sharper on-policy guarantee additionally requires recommendation coverage. Both methods preserve observed budgets on every realization and bound clean resource violation by cumulative consumption corruption. These results connect corruption-robust sparse estimation with resource accounting, pricing, and stopping in high-dimensional online allocation.
cs.LG / 157 / 2609.37209
Pointwise or Pairwise: When Do Pairwise Losses Help Reward Learning, Provably?
Abstract
Pairwise losses are increasingly used for reward learning even when pointwise rewards are observed, with mixed empirical results. When and why do pairwise losses outperform pointwise losses? We study this question in a grouped offline contextual-bandit setting allowing multiple actions per context, capturing many reward learning scenarios. We compare Value Regression (VR), which regresses observed rewards pointwise, with Value Difference Regression (VDR), which regresses reward differences between a pair of actions sampled under the same context. We consider a semiparametric model where the mean reward is the sum of a learnable action-dependent component and an arbitrary context-dependent yet action-independent nuisance, capturing context-specific disturbances. Using a unified localized analysis, we prove finite-sample regression guarantees for finite and linear function classes and translate them into offline-regret bounds. For finite classes, VDR eliminates the misspecification term in the VR bound and improves a reward-scale-dependent error term by averaging over actions within each context, a benefit absent from the corresponding VR term. For linear classes, neither method uniformly dominates: within-context differencing removes nuisance-induced bias but may increase estimation variance relative to using absolute rewards when the misspecification is sufficiently low. This yields a feature geometry-dependent bias-variance tradeoff, which we corroborate with numerical experiments.
cs.LG / 158 / 2609.37211
Explainable Machine Learning for Multilayer Planar Winding Inductance Estimation
Abstract
Rapid and accurate self-inductance estimation for multilayer rectangle-shaped planar windings is essential for modern high-frequency power converters, yet traditional workflows rely on complex mathematical equations, rigid monomial formulas or unexplainable black-box machine learning (ML) models that degrade severely outside their training domain. This paper introduces an explainable ML framework unifying post-hoc feature attribution (SHAP and permutation importance) with Kolmogorov-Arnold Network-guided symbolic regression via the SR-KAN framework to discover closed-form analytical equations without prior structural assumptions. Evaluated on a new open-source dataset of over 10,000 Finite Element Analysis (FEA) simulations across seven out-of-distribution (OOD) classes, standard tree-based ensembles exhibit severe extrapolation errors (> 36%), whereas the unconstrained SR-KAN expression achieves a robust OOD relative error of 8.22%. Experimental verification across 55 physical printed circuit board prototypes (up to 8 layers, with inductances from 4.11 μH to 559.27 μH) confirms that the KAN-discovered expression translates effectively to real-world hardware, predicting inductance with a mean absolute relative error of 6.26%. To support reproducible research, the complete FEA simulation dataset and prototype measurements are released open-source.
cs.LG / 159 / 2609.37239
Differentiating Bisimulation Metrics: A Framework for Parametric Markov Chain Fitting via Bicausal Optimal Transport
Abstract
Many problems in sequential decision-making, such as imitation learning from observations, state-space compression, world-model learning, and sim-to-real transfer, can be reduced to learning a model such that a notion of distance with respect to the target process is minimized. We consider this general framework and consider the bisimulation metric, equivalently Bicausal Optimal Transport (BOT), as the notion of distance to minimize. We show that BOT, since it can be formulated as a linear program (LP), is differentiable with respect to the model dynamics. We then derive an exact closed-form gradient via the envelope theorem applied to the LP saddle point. The result is a general algorithm, Differentiable Bicausal Optimal Transport (D-BOT), that can be applied to each of the problems above. The proposed algorithm learns the best model by alternating between distance computation and gradient steps. We apply D-BOT for three different settings: state-space compression, parametric model learning, and imitation learning from observations (ILfO). We show empirical results that confirm the viability of all three instantiations.
cs.LG / 160 / 2609.37241
Trident: Unifying Guarded Dispatch and Host Execution for PyTorch Triton Workloads
Abstract
User-written Triton kernels enable high-performance GPU computation within PyTorch, but their end-to-end latency can remain dominated by host-side orchestration, especially when device execution is short. Although torch.compile can generate native host wrappers for captured graphs, each invocation still passes through runtime-managed specialization lookup, guard evaluation, and preparation before reaching the wrapper. We present Trident, a compiler backend that removes this recurring overhead from the specialization cache-hit path. Trident introduces the Specialization Cache Module (SCM), which compiles guarded specialization selection, argument and execution-environment preparation, and host execution for multiple specializations into a single executable module. An invocation enters the SCM once, remains in compiled code when a specialization matches, and returns to Python only when a new specialization must be compiled. Built on Torch-MLIR, Trident lowers guards and host-side orchestration to native code while retaining calls to optimized runtime implementations of supported ATen operators. Our evalu- ation on two LLMs shows that Trident achieves up to a 1.47x speedup in model-level end-to-end latency over eager execution and up to 1.68x over torch.compile.
cs.LG / 161 / 2609.37255
Loss-Guided Pretraining Data Selection for Time-Series Foundation Models
Abstract
Time series foundation models (TSFMs) are pretrained on heterogeneous collections containing billions of observations, yet their training windows are typically sampled without estimating whether they provide useful learning signal. We introduce a static data-selection framework that scores each window with a reference forecaster and retains an intermediate interval within every source dataset. Specifically, we connect forecasting loss to optimization difficulty by showing that normalized squared loss controls the per-sample gradient norm under a local Jacobian condition. We then define a reference loss score and apply dataset-stratified selection to preserve the diversity of samples. Across various TSFM architectures, our method outperforms random selection by an absolute margin and even improves both relative MASE and CRPS over full-data pretraining by retaining fewer candidate pretraining windows. Further analyses show strong cross-scale and cross-architecture score correlations, indicating that a small reference model can often select data for larger targets, provided that the reference and target share compatible difficulty orderings.
cs.LG / 162 / 2609.37261
Efficiently Approximating Attention Is Hard
Abstract
Softmax attention is ubiquitous in modern machine learning, but its quadratic scaling with sequence length makes it costly. To reduce this cost, attention is often approximated with fast algorithms, which incur error but can still perform well in practice and on some inputs. At the same time, the growing diversity of attention applications makes approximation guarantees that do not depend on particular input structure a compelling target. For such uniform guarantees over all inputs, known runtime lower bounds rule out fast algorithms for near-exact attention, but leave open the practically important regime: is there an efficient algorithm with even a modest uniform approximation guarantee? We answer this question negatively. Under standard complexity-theoretic assumptions, no truly subquadratic algorithm can approximate attention with any nontrivial additive or relative guarantee uniformly over all inputs. This impossibility holds in the mildest parameter regime for which known algorithms do not already achieve strong approximation guarantees in near-linear time, and extends to practically relevant relaxations: even after polynomial preprocessing of the KV cache, no efficient algorithm can obtain a nontrivial uniform approximation guarantee, or identify a small set of keys receiving substantial attention under sparsity. Overall, our results settle the computational limits of uniform attention approximation.
cs.LG / 163 / 2609.37298
Scaling Full Conformal Image Classifiers
Abstract
Conformal prediction provides set-valued predictions with distribution-free coverage guarantees, making it attractive for high-stakes image classification. However, split conformal prediction is data-inefficient, while full conformal prediction (FCP), despite its stronger statistical efficiency, is computationally prohibitive at scale because it requires candidate-specific model refits at test time. We address this limitation by leveraging zero-shot vision-language models (VLMs) to guide scalable FCP in large label spaces. We introduce Targeted Full Conformal Prediction (T-FCP), which uses a lightweight inductive conformal predictor to prune unlikely labels and applies FCP only to the remaining candidates, reducing computation while retaining the formal guarantee of the combined conformal procedure. We further propose Stabilized Online LDA (SO-LDA), an efficient VLM adaptation solver based on rank-one inverse-covariance updates. Across multiple benchmarks, including ImageNet, T-FCP enables practical full-conformal image classification with modest test-time overhead, yielding efficient prediction sets and more stable empirical coverage than split conformal alternatives.
cs.LG / 164 / 2609.37312
Hidden Reasoning Must Leak, but Need Not Be Readable: Fundamental Opportunities and Limits for Chain-of-Thought Monitoring
Abstract
Can reasoning models trick chain of thought (CoT) monitors and perform hidden computation without revealing it in their thinking traces? We show that the answer depends on the underlying task difficulty and the model size. Simple computations can be performed covertly; however, beyond a threshold depending on model size, successfully solving the task necessarily leaks a near-linear amount of information about the covert task input into the CoT. Therefore, sufficiently complex hidden computation always leaves an information-theoretic footprint. However, concerningly, this leakage need not be readable: Under plausible cryptographic assumptions, even a one-layer Transformer can encrypt its reasoning online so that no polynomial-time monitor can extract information about the hidden computation. Overall, our theoretical and empirical results provide a holistic view of both the opportunities and the limitations of CoT monitoring.
cs.LG / 165 / 2609.37344
A Sharp Transition in Data Reconstruction under Differential Privacy
Abstract
Data reconstruction attacks have empirically been successful in recovering training samples from learned models, raising privacy concerns and motivating defenses with guarantees that remain valid against future threats. While differential privacy (DP) provides formal protection, choosing the privacy budget remains a challenge: small budgets severely reduce utility, but it is hard to quantify how large the budget can be without allowing accurate reconstruction. In this work, we study informed attackers who aim to reconstruct a single $d$-dimensional training sample from a $ρ$-zero-concentrated DP model, knowing all other training data. Our main contribution is to establish a sharp transition at $ρ\asymp d$ for data reconstruction: on the one hand, we derive entropy-based lower bounds for any private mechanism and any attack, characterizing a set of target priors for which reconstruction is information-theoretically impossible for $ρ\ll d$; on the other hand, we analyze a simple attack on private linear regression with output perturbation, showing that reconstruction is practically feasible for $ρ\gg d$. Remarkably, the transition moves to $ρ\asymp s$ for data lying in an $s$-dimensional subspace, demonstrating that the privacy budget guaranteeing adequate protection must be assessed in terms of the effective dimension of the data. We validate our findings via experiments on synthetic data and natural images (CIFAR-10, ImageNet).
cs.LG / 166 / 2609.37351
Port-Hamiltonian Latent Deliberation: Mitigating the Deliberation Drift Cliff in Test-Time Compute Scaling
Abstract
Test-time compute scaling has emerged as a cornerstone of advanced machine reasoning, yet performing iterative deliberation directly within continuous latent representation spaces reveals a catastrophic pathology: the Deliberation Drift Cliff. While unconstrained recurrent latent models achieve initial reasoning gains at short horizons (K <= 4), their reasoning collapses when extrapolated to deeper thinking steps (K >= 16), dropping by 22% to 62% across standard logical benchmarks. We resolve the trilemma among expressivity, Lyapunov stability, and computational efficiency in test-time latent reasoning through a 22-round empirical and theoretical investigation. We demonstrate that strictly conservative scalar potential gradient flows suppress long-range drift (cliff 3.40%) but bottleneck peak reasoning accuracy at 32.73%, whereas unconstrained rotational flows achieve high symbolic expressivity (82.33%) but suffer a severe 36.87% drift cliff. To resolve this geometric duality, we establish Port-Hamiltonian Latent Deliberation (PH-LD) and propose the Direct-Gradient Pure-Tensor Helmholtz-Hodge Decomposition (DG-HHD). DG-HHD parameterizes the attracting flow as a tangent projection tensor network while orthogonally decoupling non-zero circulation (Hodge machine error 1.65e-17, contraction error 5.55e-17), eliminating runtime autograd dependencies to achieve 1.84x vector field and 2.09x RK45 rollout speedups. In a 15-arm symmetrical Pareto benchmark, DG-HHD achieves 58.67% peak accuracy (+25.94% absolute gain over conservative HHD) and retains 35.27% at K=32. Transferred to small language model (SLM) multi-hop causal reasoning, DG-HHD delivers monotonic compute scaling (49.33% to 51.56%) and suppresses out-of-distribution drift (cliff -0.66%). All 30 Level 0 deterministic invariants are certified.
cs.LG / 167 / 2609.37379
Looped Transformers as Optimizers
Abstract
Looped Transformers provide a parameter-efficient approach to depth scaling by repeatedly applying shared Transformer blocks. Recent reasoning models have likewise highlighted the value of scaling test-time computation through longer computation trajectories. However, the principles for designing effective loop transitions remain poorly understood. We view the looped hidden state as a fast weight that is updated throughout the depth. We formulate loop transitions as local gradient-based updates, with recurrent blocks predicting implicit targets at each depth. Our framework derives loop transitions in closed form from a projection, a local objective and an optimizer update rule. Mapping representative loop transitions into this framework reveals mismatches between their transitions and projections. We first align the input maps of existing transitions. We then derive OperLoop, which combines explicit weight decay, adaptive step size and a delta objective. The aligned variants reduce training loss and improve average commonsense accuracy. OperLoop improves average generative performance over the compared looped and non-looped baselines under matched training FLOPs. These results support the framework's usefulness for loop design. We extend the analysis to additional loop models and outline a roadmap for future loop transition design.
cs.LG / 168 / 2609.37381
High-Dimensional Simulation-Based Inference in Latent Spaces
Abstract
Neural simulation-based inference (SBI) has been widely successful in inferring a relatively small number of interpretable parameters from potentially high-dimensional observations, such as images or time series. Accordingly, representation learning in SBI has focused almost exclusively on compressing the observations used to condition the posterior. More recently, however, SBI has begun to target increasingly high-dimensional parameter spaces, raising the complementary question of whether the inference target itself should be compressed. Our answer is a practical merger of SBI and latent generative modeling, which learns a low-dimensional representation of the simulator parameters, performs posterior inference directly in this latent space, and maps posterior samples back to the original parameter space. We characterize the conditions under which latent-space inference recovers the desired target posterior and systematically study its empirical trade-offs. Across four case studies and three generative families, we compare latent and standard estimators while controlling for network capacity, regularization, optimization, and training compute. At matched training compute, latent-space inference achieves accuracy and marginal calibration comparable to direct target-space inference while sampling up to more than an order of magnitude faster.
cs.LG / 169 / 2609.37384
MoTIF-X: A Multimodal Tokenized Framework for Interpretable and Extensible Molecular Representation Learning
Abstract
Molecular representation learning is central to computer-aided drug discovery. Molecular graphs, SMILES strings, and 3D conformations provide complementary structural information, yet many multimodal approaches encode these views independently and align them only at a later stage, limiting fine-grained cross-modal interaction and substructure-level interpretability. To address these limitations, we introduce MoTIF-X, a motif-centered framework that uses graph-grounded chemical motifs as shared anchors for multimodal integration and interpretation. Its first pretraining stage learns motif representations through hierarchical contrastive learning across atomic, motif, and molecular scales. The second stage contextualizes these representations with SMILES and torsion-angle tokens through multimodal masked token modeling. After pretraining on drug-like molecules with multiple conformers, MoTIF-X achieved the lowest mean absolute error on all nine OpenADMET ExpansionRx endpoints and the best overall performance among the evaluated methods. Significance analyses supported its advantage in the vast majority of endpoint-baseline comparisons after multiple-testing correction. Ablation studies supported the complementary contributions of motif-token contextualization, multimodal integration, and two-stage pretraining. Beyond molecular properties, the framework extended to drug-target interaction prediction, achieving the best average classification performance across the evaluated benchmarks and generalizing to an external drug-cold-start dataset without additional fine-tuning. Its motif-centered design also enabled substructure-level interpretation: higher motif attribution scores were associated with larger experimentally measured activity shifts. Together, these findings support MoTIF-X as a transferable and interpretable framework for molecular modeling.
cs.LG / 170 / 2609.37392
Learning Macroscopic Dynamics without Reconstructing Microscopic States
Abstract
Modeling the temporal evolution of macroscopic properties of complex systems is an important scientific task. To predict this evolution without full microscopic simulation, a common approach encodes microstates into compact latent states, learns their evolution, and reads out macroscopic predictions from the latent trajectory. These latent states are often learned through microstate reconstruction. However, with limited latent capacity, reconstruction can favor high-variance microscopic details over information needed for macroscopic prediction. Yet jointly learning latent states and their transition without reconstruction often fails to obtain latent dynamics that support accurate macroscopic prediction. We show that this failure can arise from latent scale collapse: shrinking the latent state scale reduces training loss while macroscopic evolution error remains large. Here, we propose a reconstruction-free framework to learn latent states with their dynamics for prescribed macroscopic prediction. Training alternates between updating the latent representation with the transition and next-state latent targets fixed, and updating the transition with the latent representation fixed. At inference, the trained model predicts macroscopic states recursively from an initial microstate. Our theoretical analysis characterizes reconstruction misalignment and scale collapse under joint training, and gives a sufficient condition for local convergence to correct latent dynamics for our method. Experiments on epidemic spreading on a lattice, mixing of two particle species, and polymer stretching demonstrate that the proposed method achieves substantially better macroscopic prediction over baselines.
cs.LG / 171 / 2609.37416
Scale Sensitivity in Low-Bit Post-Training Quantization: Curvature of the Quantization Error Landscape
Abstract
Post-training quantization (PTQ) methods in the GPTQ family minimize a layer-wise reconstruction error on a uniform grid whose scale must be chosen; the common max-based choice degrades sharply at low bit-widths. We study how sensitive this objective is to the scale. For a layer with i.i.d. Gaussian weights and calibration activations of sufficiently large effective rank, we prove that, as the width grows, the normalized round-to-nearest loss converges with high probability, uniformly over all scales, to the mean-squared error of a uniform quantizer applied to a standard Gaussian; we verify the effective-rank condition for wide, randomly initialized MLPs with odd Lipschitz activations and isotropic Gaussian calibration data. The limiting objective has a unique nondegenerate minimizer, whose scale decreases strictly with the number of levels and whose curvature with respect to relative scale errors decays approximately exponentially with the bit-width. GPTQ experiments on five LLMs show the same trend: the scale rule changes perplexity substantially at 2--3 bits and negligibly from 6 bits on, and a local measure of GPTQ scale sensitivity decreases with bit-width in line with the Gaussian curvature. The Gaussian-optimal scale fails on raw weights; after Hadamard incoherence processing it matches the best searched rule at 3 bits and above without any search, but remains clearly worse at 2 bits.
cs.LG / 172 / 2609.37424
Simultaneous Neural Optimal Transport
Abstract
Optimal Transport (OT) provides a principled framework for learning transformations between probability distributions from unpaired samples. In many applications, however, a single transformation must map several source distributions to a common target distribution. For example, image restoration might require handling different types of degradation without knowing the degradation of each input at inference time. Simple approaches of pooling the source distributions only encourage alignment with the target at the aggregate level and may leave individual sources misaligned. In our paper, we consider the simultaneous OT problem which formalizes the task of learning a shared transport map that minimizes the average transport cost while aligning each source distribution with a prescribed target. We propose a neural method for solving the simultaneous OT problem by learning a shared transport map that minimizes the average transport cost while aligning each source distribution with a prescribed target. We derive a max-min formulation for learning this map. We illustrate its application to image restoration, where a single model handles multiple degradation types using a common collection of clean target images.
cs.LG / 173 / 2609.37432
Looped Actor: Depth-Recurrent Reasoning Models for Reinforcement Learning
Abstract
Looped reasoning models repeatedly apply a shared set of parameters, enabling more computation without increasing the model size. These models also support input-dependent computation by dynamically deciding when to stop looping. Motivated by the recent success of looped transformers in language modeling and reasoning, we investigate whether dynamic looping can similarly benefit sequential decision-making. We provide a complexity-theoretic motivation for this approach by showing that there exist Markov decision processes in which a state-adaptive policy achieves the optimal return with asymptotically less expected computation than any optimal fixed-runtime policy. To learn compute-adaptive policies in practice, we introduce Looped Actor, a transformer-based policy that repeatedly refines a latent representation toward a fixed point using a shared computational block. This allows the model to allocate computation adaptively by varying the number of loops based on the current state. We evaluate Looped Actor on 22 tasks across six environments, ranging from combinatorial puzzles to robotic manipulation and spanning online and offline reinforcement learning (RL) with discrete and continuous actions. Looped Actor matches or exceeds the performance of an untied baseline with 16$\times$ more parameters, with the largest gains in environments where action selection requires substantial multistep planning. For the Boxoban environment, we find that the computation allocation is structured: the number of loops increases with the number of remaining pushes and future optimal pushes become increasingly predictable from the latent state over successive loops. Together, these results highlight actor looping as a simple and efficient way to equip RL agents with adaptive computation and improve their planning capabilities. Code is available at https://github.com/camail-official/LoopedActor
cs.LG / 174 / 2609.37435
Variational Augmented Invertible Koopman Autoencoder for probabilistic time series forecasting
Abstract
Neural Koopman autoencoder models have been shown to successfully build a latent embedding with linear dynamics for arbitrary dynamical systems, enabling strong performance in long-term time series forecasting. However, these models usually work in a deterministic setting, which does not allow the quantification of the uncertainty of their predictions. Thus, we propose the new Variational Augmented Invertible Koopman AutoEncoder (VAIKAE), in which the latent embedding follows a Gaussian distribution instead of being deterministic. A key property of the VAIKAE architecture is that it leverages normalizing flow models, enabling the use of likelihood computations in the state space of dynamical systems for training a model. We further propose new strategies for uncertainty-aware latent data assimilation with a trained VAIKAE model. The effectiveness of our methods is demonstrated in a series of experiments on long-term time series forecasting benchmarks.
cs.LG / 175 / 2609.37509
ScaGNN: a Graph Neural Network for Multiple Scattering Simulations
Abstract
The boundary element method (BEM) provides an efficient numerical framework for solving multiple scattering problems in unbounded homogeneous domains. By restricting the discretization to the domain boundaries, it substantially reduces computational complexity. The procedure first consists in determining the solution trace on the boundaries of the domain by solving a boundary integral equation. Then, the volumetric solution can be recovered at low computational cost using a boundary integral representation. As the first step of the BEM represents the main computational bottleneck, we present ScaGNN, a learning-based approach designed to approximate the solution trace. It relies on a graph neural network architecture that incorporates a dynamic adaptive edge sampling mechanism for selecting the most relevant interactions to model. Guided by intermediate predictions of expected error and edge length, this mechanism selects, at various stages of the forward pass, the most relevant distant interactions to model. The proposed method is tailored to achieve linear complexity with the number of nodes in the input graph. To train and evaluate our network, we present a benchmark consisting of several datasets with different types of multiple scattering problems. Our experiments show that our approach surpasses existing state-of-the-art learning-based methods on the considered tasks and investigate the generalization capabilities to settings with an increased number of obstacles and out-of-distribution obstacle shapes. github.com/LARIAD/ScaGNN
cs.LG / 176 / 2609.37515
Hierarchical Compression of Vision-Language Model Benchmarks
Abstract
Thorough evaluation of vision-language models (VLMs) has become prohibitively expensive, as benchmarks span an ever-broader spectrum of capabilities and new models arrive at a relentless pace. Benchmark compression methods that preserve model rankings at a fraction of the cost are well studied for language models, but for VLMs the question remains under-explored. We present PRIMEBench (Pruning Redundant Items for Multimodal Evaluation), a vision-aware hierarchical benchmark compression framework that substantially reduces evaluation cost while preserving model rankings. This hierarchical framework operates in four stages: data cleaning to remove items answerable without the image and all-correct items, category representative selection to pick one benchmark per capability category, item pruning with Vision-Aware Variance (VAW), and category-count pruning. VAW combines inter-model variance with a vision-dependence score computed from multimodal embeddings alone, while encouraging coverage of diverse items within each benchmark. On models held out from item selection, it has the highest mean fidelity at the released 5% retention. The hierarchical design lets practitioners stop at any stage to match their compute budget; the released suite removes over 97% of items while preserving model rankings. Beyond compression, our analyses show how VLM evaluation behaves as model panels grow and evolve, providing guidance for designing future benchmarks that are more efficient, robust to model turnover, and explicit about the limits of evaluation-side pruning.
cs.LG / 177 / 2609.37522
Graph-Conditioned On-Policy Agent Distillation from Off-the-Shelf Teachers
Abstract
On-policy distillation (OPD) trains compact language agents with teacher feedback on student-generated trajectories. In multi-turn tasks, compounding errors can move students beyond the teacher's effective supervision. We introduce Graph-Conditioned On-Policy Agent Distillation (GC-OPD), which enriches an off-the-shelf teacher's scoring context with execution evidence. A graph indexes repeated teacher executions by shared states while preserving complete successful and failed histories. After each student episode, GC-OPD retrieves current-state references or historical alternatives and combines them with student hindsight to score the original thought-action tokens. Using the same original teachers, GC-OPD improves mean success over vanilla OPD from 24.70% to 48.78% on ScienceWorld (4B student), from 53.36% to 85.26% on ALFWorld Unseen, and from 29.10% to 37.65% on WebShop. At matched student sizes, it also achieves higher mean success than every evaluated OPD baseline using GRPO-trained teachers on ScienceWorld and ALFWorld; the strongest such ScienceWorld 4B baseline reaches 46.66%. GC-OPD requires no task-specific teacher optimization.
cs.LG / 178 / 2609.37525
Physical Muon: Orthogonalization as an Equilibrium Computation
Abstract
Physical neural networks and analog in-memory computing could reduce the energy cost of neural network training. Realizing this potential, however, requires optimizers that combine effective learning with physical implementability. SGD fits local analog updates but struggles on transformers, while Adam family is unstable against analog bias. Muon offers strong training performance, but its Newton--Schulz orthogonalization relies on dense matrix-matrix products. To address this obstacle, we introduce Physical Muon, which computes the orthogonalization as the equilibrium of a continuous-time flow. Random probes approximate the flow using matrix-vector products, reciprocal reads, and local rank-1 writes. To test whether this replacement preserves training performance, we evaluate it on a 10.95M-parameter transformer. The dense flow's mean validation cross-entropy is 0.0085 above Newton--Schulz across nine seeds per method; the probe implementation is 0.0188 above the control across two seeds. Circuit simulations further reproduce the flow dynamics and yield comparable training behavior.
cs.LG / 179 / 2609.37535
Why Adaptive Optimizers Underestimate Rare Tokens
Abstract
In the softmax output layer, a rare token receives a small positive logit gradient on most steps and a much larger negative gradient on the few steps when it is the target. SGD simply adds these contributions. Coordinate-wise adaptive methods such as Adam, RMSProp, and sign descent instead divide each update by a running estimate of its magnitude, and that estimate is largest immediately after the token appears. This imbalance has two effects. At the level of the whole output layer, we characterize which optimizers preserve the mean output embedding: every method whose update is linear in past gradients does, as do Kronecker-factored and orthogonalized methods such as Shampoo and Muon. Adam, Adafactor, Lion, and sign descent do not, and for these methods we obtain an exact step-by-step expression for the change. At the level of an individual rare token, the same normalization shifts the training fixed point. In the unigram model, sign descent lowers the logit of every token that occurs in fewer than half of the minibatches at a constant expected rate. For RMSProp with periodic arrivals, we can solve the fixed point in closed form: if a token is absent for at least two consecutive minibatches, its equilibrium probability is strictly below its data frequency for every learning rate, and the ratio tends to $κ/(2(e^{κ/2}-1))$. Here $κ$ is the mean number of steps between occurrences divided by the second-moment time constant $1/(1-β_2)$. In the same model, SGD and AMSGrad retain the unbiased fixed point. We test these predictions both in a unigram model and in a small language model trained from a known generating distribution. With random arrivals, the bias is larger than the periodic formula predicts; in the language model, the optimizers with the biased fixed point also fit the generating distribution less well.
cs.LG / 180 / 2609.37555
Benchmarking graph-based models for in-silico toxicity prediction in drug discovery
Abstract
Drug discovery is a costly and high-risk process, where toxicity-related failures remain a major cause of attrition in both preclinical and clinical stages. As a result, accurate early prediction of chemical toxicity is essential to reduce downstream costs and improve compound prioritization. In this context, graph deep learning (GDL) has emerged as a powerful paradigm for toxicity prediction, leveraging molecular graph representations to learn directly from chemical structure with improved expressivity over traditional approaches. Despite the growing number of proposed models, current literature-based comparisons are often difficult to interpret due to inconsistencies in datasets, preprocessing pipelines, and evaluation protocols. To address this limitation, we introduce a unified and standardized benchmarking framework for GDL-based toxicity prediction. We systematically evaluate more than 20 representative approaches under consistent experimental conditions and across multiple datasets and partitioning strategies, enabling a fair and reproducible comparison of model performance. In addition, we complement this empirical study with a structured literature analysis to contextualize existing methodological trends and performance claims. Our results provide a clearer and more reliable assessment of the current state of the field, highlighting both the strengths and limitations of existing graph-based approaches. To support transparency and reproducibility, we release our benchmarking framework as open-source software https://gitlab.citius.gal/noel.suarez/benchtox, allowing the community to evaluate and compare models under consistent conditions.
cs.LG / 181 / 2609.37565
A Model-Agnostic Physics-Guided Adapter for Few-Shot Transfer of Coastal Flood Prediction Models to Unseen Regions
Abstract
Deep learning surrogates can produce high-resolution coastal flood maps orders of magnitude faster than physics-based hydrodynamic simulators, yet transferring them to new coastal regions remains costly, since generating target-region data for fine-tuning typically requires numerous time-consuming simulations. To tackle this bottleneck, we introduce the Physics Adapter (PA), a compact, architecture-agnostic adaptation interface that enables efficient few-shot transfer of flood prediction models across diverse coastal regions. PA predicts peak water level through a differentiable wet/dry response that compares terrain elevation against a learned water level, and blends this physics-structured prediction with a data-driven branch through a learned gate. Unlike physics-informed formulations, PA imposes no PDE-residual or conservation losses and instead exploits elevation as an architectural inductive bias, adding a negligible number of trainable parameters. We integrate PA into 12 heterogeneous models, and evaluate them on two coastal regions with markedly distinct geometries, topographies, and shoreline protection configurations. The performance of PA is benchmarked against a no-physics baseline, full fine-tuning, and standard parameter-efficient fine-tuning (PEFT) methods, considering both within-region generalization to unseen sea level rise values and between-region transfer. In low-shot regime (K=3), and averaged over all backbones and transfer settings, adding PA reduces root mean square error by 11.5% when only the output head is adapted on a frozen backbone, by 15.4% when combined with PEFT methods, and by 22.9% under full fine-tuning, compared to matched configurations without PA. Taken together, the findings of this work offer practitioners a concrete recipe for extending DL-based coastal flood predictors to new, data-scarce regions.
cs.LG / 182 / 2609.37604
GraphVQ: Structure-Aware Autoregressive Decoding over Context-Quantized Graph Tokens
Abstract
Graph foundation models need a discrete token representation, but casting a graph as a generatable token sequence faces a structural obstacle: edges spanning beyond the serialization window cannot be emitted in one pass--so one-pass autoregressive generators systematically under-produce cycles--and a single global condition cannot tell candidate edges apart. GraphVQ removes both obstacles: node contexts--features plus a local edge mask under multi-order breadth-first serialization--are quantized into a shared codebook by a VQ-VAE with BCE-calibrated Bernoulli edge decoding, and a second-stage structure-aware decoder emits the global adjacency conditioned on token-derived pair features, whose necessity over any global-summary condition is formalized in a scoped impossibility result. The tokenizer reconstructs node features at 0.86--0.99 accuracy and decodes local edges at AUROC >= 0.89 (ECE <= 0.007). Under one same-split protocol on four datasets, pair conditioning improves orbit MMD 0.248 -> 0.174 on PROTEINS and 3.4x on a ring stress test, and vanishes on a random-label control--the signature of attribute--topology coupling--so the gain is claimed exactly where attributes carry edge-relevant signal. GraphVQ ranks first among learned generators on PROTEINS, ties for first on SYN-COMM, and improves orbit MMD 2.7--17x over one-stage generation on three datasets, with seed-level bootstrap intervals confirming the rankings are not seed noise; on MUTAG the unweighted edge target under-generates and is reported as such. These results locate the structural control of autoregressive graph generation in the granularity of the condition: pair-level token context turns a quantized vocabulary into a usable capacity axis for distribution-faithful graph generation and future token-level pretraining.
cs.LG / 183 / 2609.37609
PHASE: Multi-Regime Modeling of Incompressible Magnetohydrodynamics
Abstract
Magnetohydrodynamics (MHD) is central to plasma modeling in astrophysics, space science, fusion, and engineering, but resolving multiscale MHD dynamics is computationally expensive. Machine-learning surrogates enable fast inference by learning reusable solution operators, yet existing models require separate training for each physical regime, limiting generalization across varying parameter settings. We introduce PHASE, a PHysics-Adaptive Scalable operator with residual Error correction, designed to model incompressible MHD across varying physical parameters with a single model. PHASE combines transfer learning, regime-aware adaptation, physics-centered learning, and residual refinement to improve both physical fidelity and generalization across MHD regimes. Together, these improvements achieve state-of-the-art prediction accuracy on two-dimensional MHD turbulence by reducing relative $L_2$ errors on physical fields by more than an order of magnitude compared to prior MHD neural-operator baselines. Moreover, PHASE generalizes successfully to unseen parameter values without retraining, demonstrating the cross-regime adaptability expected from operator learning. We evaluate PHASE beyond point-wise prediction errors using derived physical fields, spectral analysis, and distribution statistics, consistently observing improved physical fidelity. We further show that our framework can accurately simulate MHD instabilities by testing it on the Kelvin--Helmholtz instability, demonstrating the robustness of our method.
cs.LG / 184 / 2609.37616
Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable
Abstract
Language models tend to agree with whatever a user asserts, and post-training increasingly targets this sycophancy so that models evaluate claims on their merits rather than deferring to the user. Yet the same models are far more compliant when a wrong answer is attributed to a verified source, which is how retrieval results, tool outputs, and grounded-search content often present information. We measure this gap across five open-weight families and three closed APIs. A single verified-source note endorsing a wrong answer flips 45-88% of baseline-correct responses in seven of eight models, and compliance rises with how authoritative the note sounds. Source deference and user agreement are not behaviorally interchangeable inside the model: on matched items with the same wrong answer, causal interventions can selectively suppress one without equally affecting the other. In three open-weight families, removing a fitted source direction lowers source compliance by 65-80 percentage points while removing a user or assistant direction has far smaller effects, and removing the user direction shows the reverse preference. A separately fitted intervention derived from source-versus-user cue activations moves compliance in both directions while leaving the prompt text unchanged. An authority direction fitted on trivia also transfers to PIQA and multi-turn SYCON dialogues without refitting, and removing it lowers wrong-source compliance by tens of percentage points in four of five families with no detected change in MMLU-Pro or GSM8K accuracy at our evaluation sizes. Source deference and user agreement therefore need separate evaluation.
cs.LG / 185 / 2609.37631
Procedural Core: A Compact Recurrent Initialization for Vision Transformers
Abstract
Transformers are typically trained from random initialization, requiring all their capabilities to emerge from large-scale optimization. Recent work showed that a small amount of abstract procedurally generated data can help acquire generic inductive structure at low cost. However, this adds a pretraining stage that must be repeated for every target model. We propose Procedural Core, an initialization strategy that captures this generic structure into a compact set of weights that can be reused across models. We train a minimal recurrent transformer on procedural data, then expand its weights to initialize transformers of arbitrary width and depth. The resulting initialization improves performance on image classification, self-supervised visual learning (DINO), and modeling natural language (FineWeb-Edu) and code (CodeParrot). For image classification, expanding a 1M-parameter core to initialize an 85M-parameter ViT-Base improves ImageNet top-1 accuracy by 2.2 pp over standard random initialization. Our analysis identifies recurrence as essential for learning compact weights that transfer across models. In ViTs, we localize a key benefit in the suppression of high-norm tokens that produces substantial improvements in zero-shot segmentation (ImageNet-S mAP 32.3 to 42.9), object localization (VOC07 CorLoc 9.9 to 18.4), and depth estimation (NYUv2 RMSE 1.104 to 0.998). This demonstrates that transformers need not start from a blank slate, and can be initialized with generic capabilities at low cost with no domain- or task-specific data.
cs.LG / 186 / 2609.37633
RLTL;DR: Self-improvement by Internalizing Self-generated Feedback
Abstract
The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed attempt, we show the policy the verifier outputs and let it write its own feedback, in the form of a single TL;DR insight. The next rollout is conditioned on all previous insights, and we sequentially sample rollouts until a solution is found. Moreover, we enable backpropagation on the in-context insights to internalize a direct task to insight mapping. On challenging tool-calling and coding datasets (filtered to Pass@128=0), standard GRPO training of a Qwen 3.5 9B Thinking policy stays flat at a Pass@1 of 0% to 1%. RLTL;DR breaks through this learning barrier, achieving a Pass@1 of 14-31% with insights in context during training and, crucially, 12-13% when no insight is in context at eval time. We identify that the key is the task to insight internalization. To study this further, we reduce our approach to SFTL;DR, training only on (task, insight) tuples, without showing or backpropagating on any rollouts. Training on only 4k of these tuples recovers almost the full performance of RLTL;DR and classical SFT on full rollouts. This demonstrates a promising compacted training paradigm of the form "on this sort of task, keep this sort of thing in mind", which we hope to inspire future research on.
cs.LG / 187 / 2609.37650
RACE: Relation-Level Counterfactual Explanations for Heterogeneous Graph Neural Networks
Abstract
Counterfactual explanations of graph neural networks identify edge deletions that flip a prediction. On heterogeneous graphs, however, existing methods first collapse the graph into untyped edges, so they cannot answer the question a domain expert actually asks: which relation type drives this prediction? We present RACE (Relation-Aware Counterfactual Explanations), which gives this question an exact, per-instance answer. For every explained instance, an exhaustive search over relation subsets returns the certified minimum relation-deletion set that flips the prediction -- or an explicit report that no such deletion exists; each relation-level answer is then refined into a typed edge set within the attributed relations, verified on the discrete model by single-edge restoration. The relation-level answer is exact and deterministic given the frozen backbone, whereas soft-mask baselines vary by 6-8 pp in success rate across runs differing only in random ordering. On ACM, a Cora-derived graph, and ogbn-mag, RACE improves counterfactual success rate over the strongest baseline by up to +2.7 pp while deleting fewer edges, and attains the highest success rate among all same-task baselines on every dataset; the advantage reproduces across four backbones on ogbn-arXiv and on DBLP, with cross-seed relation-set agreement up to 0.89. A synthetic study with known generating mechanisms confirms that the search recovers the relation the trained model actually relies on -- and reports infeasibility rather than fabricating an attribution when the model has learned none -- so the explanations stay trustworthy exactly where explanations matter.
cs.LG / 188 / 2609.37660
Nonpreemptive Scheduling While Learning Context-Dependent Service Rates
Abstract
We study nonpreemptive contextual queueing bandits in a single-server system. Each job is represented by a $d$-dimensional context vector; in each round, a job may arrive with its context drawn from an unknown distribution $\mathcal{D}$, and its departure probability is determined by a logistic model of that context vector with an unknown parameter $θ^*$. The server learns from service outcomes while deciding which waiting job to serve and whether to idle, aiming to minimize queue-length regret, the gap between its expected terminal queue length and the minimum achievable by an admissible policy. Once selected, a job must be served until completion, and we refer to this as the nonpreemptive setting. A central challenge is that, even with full model knowledge, the optimal policy cannot in general be characterized by a simple myopic rule, since the optimal action can change with the remaining horizon at the same queue state. Nevertheless, when the model and horizon are known, the optimal action can be obtained through a finite-horizon Bellman recursion. Motivated by this, we propose Learn--Clear--Plan (LCP), which estimates the system and uses the resulting Bellman recursion to make horizon-dependent decisions. LCP achieves $\widetilde{O}(\sqrt{d/T})$ queue-length regret, while a lower-bound construction gives $Ω(\min\{1/\sqrt{d},\sqrt{d/T}\})$ regret for every learning policy on some instance, establishing optimality up to polylogarithmic factors when $T\ge d^2$. When the horizon is unknown, no horizon-independent policy achieves vanishing regret against the finite-horizon optimum. We therefore use SEPT, the policy that serves a waiting job with the highest probability of departure, as a fixed reference, and suggest an estimated-SEPT algorithm that achieves a tracking error of $\widetilde{O}(\sqrt{d/t})$ without knowing the model.
cs.LG / 189 / 2609.37664
Learning Causal Normalizing Flows from Incomplete Data via Observed-Data Likelihood
Abstract
Causal Normalizing Flows (CNFs) enable causal inference from observational data given the causal structure, but they assume fully observed training data. We introduce MissCNF, which trains CNFs directly on incomplete data by maximizing the marginal likelihood of each partially observed sample, without discarding rows or constructing a completed dataset. Thanks to the causal structure encoded in the autoregressive factorization of CNFs, only missing variables in the ancestral closure of the observed set are integrated out, while the others are dropped without computation. We further establish the conditions under which MissCNF recovers the true joint distribution, and introduce \emph{causal-family positivity}, where identification is possible even when no record in the dataset is ever complete. We compare MissCNF with two common strategies for handling missing data: listwise deletion and impute-then-fit pipelines. Across eight synthetic causal benchmarks, three missingness mechanisms, and missing rates up to $90\%$, MissCNF achieves the lowest KL divergence in 23 of 24 nonlinear MCAR and MAR settings and in all nonlinear MNAR settings, as well as the lowest counterfactual error in 20 of 24 settings. On linear SCMs, where linear imputation performs best, MissCNF ranks in the top two in 22 of 24 settings.
cs.LG / 190 / 2609.37667
Where Privacy Belongs: Placement Diagnosis and Certified Selection for Private Counterfactual Explanations on Graphs
Abstract
Counterfactual explanations for graph neural networks (GNNs) find the minimal intervention that flips a node's prediction--but computing one requires reading sensitive graph structure, and releasing it discloses that structure. Both existing placements fail. Privatizing the graph before explaining corrupts the target on exactly the borderline nodes needing recourse, manufacturing spurious flips that flip the privatized graph but not the true one. Explaining on the clean graph and perturbing the released explanation resists certification: re-auditing the standard heuristic shows an implied full-release budget of 573--753 on Cora and 256 on CiteSeer--orders of magnitude beyond its advertised budget--with worst-case single-entry leakage at AUC 1.0. We propose PrivCFS, which replaces certification-by-optimization with certification-by-construction: counterfactual selection over a fixed, data-independent candidate universe--edge interventions from a public prior graph, feature interventions from a public schema--whose no-op semantics give neighboring graphs the same output support. A validity-gated, clipped utility of global sensitivity $Δu \le 1$ released through the exponential mechanism gives pure $\varepsilon$-DP for the complete released object, composable over queries--to our knowledge the first such guarantee on graphs. Privacy noise is the cheapest stage: at $\varepsilon$=8 the release retains 94--97% of its support-restricted non-private optimum on the recourse population and 83--95% on the general one; the optimal edge-inference audit attains AUC 0.50 on average and 0.59 worst-pair, versus the heuristic's worst entry 1.0; and transfers to a 15K-node graph at 0.96 valid rate. The dominant cost is a measurable, monotone price in public disclosure, readable off one table before any budget is spent--turning explanation privacy from an accounting risk into a purchasable decision.
cs.LG / 191 / 2609.37675
LEMON-ZEST: Evolution-Informed Tokenization for Efficient Protein Language Modeling
Abstract
Protein Language Models (PLMs) have made remarkable progress following scaling laws established in natural language processing across sequence- and structure-based tasks, yet the potential of tokenization remains underexploited. Unlike human language, proteins preserve structure despite extensive sequence variation a property standard tokenization strategies fundamentally fail to capture. We introduce ZEST (Zoned Encoding of Sequence Traits), an evolution-informed vocabulary derived from conserved regions of multiple sequence alignments. ZEST allows embedding domain-level biological priors directly at the tokenization stage rather than learning them implicitly through scale. ZEST natively compresses sequences to an average token length of 4 residues, enabling our model to process 4,000 residues within a standard 1024-token context window. Building on this, we present LEMON (Layered Extraction of Molecular Ordering from Nature), a compact 200M-parameter sequence-based model for detection of remote homology between protein sequences trained on a single H100 GPU for one week. Despite its modest size, LEMON outperforms state-of-the-art models ranging from 600M to 3B parameters. Our results demonstrate that evolution-informed tokenization can substitute for massive parameter scaling, opening a new direction for efficient, biologically-grounded protein representation learning. All code, model weights, and results are publicly available under the MIT license.
cs.LG / 192 / 2609.37680
When Models Don't Manipulate Manifolds: The Geometry of a Comparison Task
Abstract
One of the current premises of mechanistic interpretability research is that detailed accounts of the geometry of neural network representations can tell us how models perform computations, and how to effectively intervene on them. While low dimensional manifolds have been observed for multiple concepts in the literature (e.g. numbers encoded on helices, days of the week on a circle, ...), with structure believed to reflect properties of data and tasks, the extent to which models rely on them for computation, and how they manipulate them, remains unclear. We characterize precisely the geometry of computation in a number-comparison task, as an abstraction of comparison for decision making, and how models utilize geometry in an elegant fashion to implement it. Specifically, we study the causal geometry of number comparison in Qwen2.5-7B-Instruct, a capable and widely studied open-weight model, and find Qwen largely uses linear representations of numbers despite the presence of curved geometry. To compare two numbers, the model first encodes each number along a vector and adds the two representations using attention and the residual connection, bringing them into a shared space in the residual stream. Then, the model uses MLP neurons to compare the pair of numbers on local regions in this shared space, which correspond to smaller intervals of input numbers, and combines these to obtain the position of the maximum. In fact, this reliance on linear representations for comparison also persists when the model compares three numbers. Our findings demonstrate that the manifold hypothesis can co-exist with linear representations: while concepts that are ordered may have manifold structure in representations, the model may use an underlying linear structure of the concept in certain computations.
cs.LG / 193 / 2609.37702
Width Expansion as a Method for Class Incremental Learning
Abstract
Class Incremental Learning (Class-IL) requires models to learn new classes over time while preserving previously acquired knowledge without access to past data or task identity. This setting intensifies the stability-plasticity dilemma and makes catastrophic forgetting a central challenge. Existing approaches include regularization, knowledge distillation, replay, and architectural expansion. However, many expansion methods rely on explicit task identifiers or predefined growth strategies, limiting their applicability when task boundaries are unavailable at inference time. This work proposes a dynamic width expansion method that increases the number of neurons within existing layers according to a normalized loss criterion, without requiring task-specific information. An attention mechanism with persistent key-value memory is also incorporated to stabilize feature representations and reduce interference between previously learned and newly introduced classes. The approach is evaluated on Split MNIST and Split CIFAR-100 under the standard Class-IL protocol. Experiments compare fixed-capacity and dynamically expanding architectures, both with and without attention, combined with established continual learning methods including EWC, LwF, and A-GEM. Results show that progressive width expansion consistently improves performance over fixed architectures, particularly when combined with functional methods and A-GEM. The combination of width expansion and attention provides the most consistent gains. Overall, dynamic width expansion based on representational demand provides an effective and flexible strategy for Class-IL, although uncontrolled growth may increase overfitting and computational cost.
cs.LG / 194 / 2609.37715
Volatility-Clustering Adaptation for Financial Time Series
Abstract
Time-series foundation models are increasingly adapted to new domains through fine-tuning on target data, under the implicit assumption that more target data yields better forecasts. We show that this assumption can fail in financial forecasting, where individual price changes are difficult to predict, but large moves tend to cluster, creating alternating calm and turbulent periods. Using financial foundation models trained on price bars of open, high, low, close, and volume, we argue that adapting to financial domains requires training signals beyond next-token prediction. We introduce Volatility-Clustering Adaptation (VCA), which augments next-token cross-entropy with a differentiable penalty on the autocorrelation of squared returns, the standard statistical signature of volatility clustering. This additional objective provides a multi-step training signal by matching the resulting dependence structure of autoregressive rollouts to those of the realized future. Across three asset sets and two evaluation conventions, VCA improves adaptation over the pre-trained model, with the strongest gains under the primary evaluation (\textsc{fore}), driven primarily by reduced variance error. Overall, our results suggest that effective financial adaptation requires objectives that capture domain-specific temporal structure beyond token-level prediction.
cs.LG / 195 / 2609.37717
Predictive Geometry of Hidden Trajectories in Transformers
Abstract
Decoder-only transformers are trained only through a terminal next-token prediction loss, yet this loss constrains every intermediate hidden state through the fixed downstream computation. We formalize this constraint by studying layerwise loss-to-go functions: the terminal loss obtained by continuing a candidate hidden state through the remaining transformer blocks. Around successful validation trajectories, we show that the local second-order geometry of these functions is governed, up to low-loss residual terms, by a pullback Fisher operator on hidden-state space. Its spectrum identifies output-sensitive directions and approximately prediction-null directions, yielding a local observable subspace of the residual stream. For causal transformers, the same geometry induces a tokenwise curvature score: a Fisher-weighted sensitivity of the target logits to perturbations of each token's hidden state. This score vanishes outside the causal ancestor set of the target and is controlled by downstream Jacobian couplings, making it a loss-aware alternative to attention magnitude. We estimate these quantities using matrix-free Jacobian-vector and vector-Jacobian products and evaluate them across decoder-only language models on WikiText, OpenWebText, and FineWeb. Empirically, the induced geometry predicts perturbation sensitivity, supports nonuniform layerwise rank allocation, yields competitive structured token-pruning signals, and improves low-rank student recovery when added to stronger autoregressive distillation objectives such as reverse KL and skew KL. These results support a predictive-geometric view of transformer computation: near successful trajectories, the terminal loss induces a thin, anisotropic set of output-relevant hidden-state directions that can be measured and exploited for compression and distillation.
cs.LG / 196 / 2609.37740
HyDI: A hybrid Deep Learning-Inductive Logic Programming ensemble for multi-label classification
Abstract
While attaining remarkable results for many applications, Deep Learning models are notoriously difficult to explain. This work introduces HyDI, a hybrid ensemble architecture for hierarchical multi-label classification. It combines a Deep Learning (DL) model with rule-based classifiers generated by Inductive Logic Programming (ILP). For leaf classes of the label hierarchy, the rule-based classifiers replace the DL model, leading to more transparent classification results. HyDI is applied to the Chemical Entities of Biological Interest (ChEBI) ontology, providing ILP-generated rules for 314 classes. For these classes, HyDI can generate global explanations as well as local explanations that combine visual and text-based descriptions.
cs.LG / 197 / 2609.37789
Predictive Self-Supervised Learning Provably Identifies Stochastic Signals under Nuisance
Abstract
Self-supervised learning (SSL) by predicting in latent space, without generating the input data itself, learns highly abstract, useful representations. Intuitively, this success is often attributed to its ability to discard nuisance information that is irrelevant to prediction. However, this poses a conundrum: both stochastic variation in a prediction-relevant latent signal and true nuisance make observations partly unpredictable; how could they be distinguished? Surprisingly, we prove that common SSL methods can achieve exactly this, by implicitly instantiating a latent-variable model with stochastic dynamics and observation-private nuisance. We trace their ability to recover the stochastic signal to two complementary principles: Predictive mutual information maximization ensures that representations retain the information needed for prediction, while latent distribution matching constrains how this information is encoded, thereby making the retained signal identifiable. We confirm this identifiability result in simulations for Gaussian predictors, which recover the true signal up to an affine transformation even in dynamic, nuisance-laden environments.
cs.LG / 198 / 2609.37800
Challenges and Solutions for Bandits in the Wild: Warm-Started Mixture Bandits for Cross-Cohort Slate Recommendation
Abstract
Many recommender services repeatedly encounter cold-start cohorts, where new users arrive with little or no interaction history. This creates two challenges: learning user preferences quickly from limited feedback and sustaining useful recommendations when each user has a finite catalog that can become repetitive or depleted over time. We propose CohortMix-TS, a warm-started mixture bandit that learns latent user groups from earlier cohorts and uses available metadata to construct group-informed priors for new users. Starting from these fixed priors, the model personalizes independently as feedback from each user becomes available. Session slates combine Thompson sampling with diversity and inventory-depletion controls. We evaluate CohortMix-TS through simulation, semi-synthetic experiments, and a 25-day randomized in-the-wild deployment with 713 registered participants in a Campus Games quiz application. Our evaluations show that cross-cohort transfer improves early recommendation quality and user-level regret, while inventory-aware slate construction helps prevent premature exhaustion of preferred items. In the field deployment, treatment users also showed a larger early-to-late change in correctness than users receiving random recommendations. Together, these results show how warm-start transfer and inventory-aware recommendations can support personalization for short-lived, repeatedly cold-starting cohorts.
cs.LG / 199 / 2609.37808
Feedback-Calibrated Protein Optimization with Batch-Aligned Tail Arbitration
Abstract
Protein optimization aims to discover high-fitness sequences under a limited experimental budget. Existing machine-learning methods use task-specific predictors, biological priors, or ranking-aware objectives to guide which variants are tested in the next experimental round. However, these methods cannot adapt to shifts in the reliability of predictive evidence as measurements accumulate and ensure the correct ranking of key high-fitness candidates. To address these challenges, we propose Batch-Aligned Tail Arbitration (BATA), which uses experimental feedback to adaptively combine prior-informed and task-specific rankings for next-batch selection, with calibration focused on the batch-aligned high-fitness region. Across measured GB1, PABP, and TrpB landscapes, BATA achieves the best mean task rank (1.67) in final best fitness after 480 measurements. Controlled comparisons further show task-dependent gains from high-fitness calibration and batch alignment. Our work introduces feedback-calibrated predictor arbitration, where experimental feedback dynamically determines how predictive evidence guides next-batch selection, opening a new direction for protein optimization.
cs.LG / 200 / 2609.37836
Behavioral Convergence Without Representational Convergence: Persistent Training-History Dependence in Neural Networks
Abstract
Neural networks trained toward the same final objective can reach similar predictive performance while retaining internal representations shaped by earlier training history. We study this effect using controlled sequential-training experiments in which paired convolutional networks start from identical weights, experience reversed task orders, and then receive the same deterministic common-relaxation distribution. Across 20 paired MNIST runs, 16 satisfy a predeclared behavioral-matching criterion, yet their matched representations retain a mean history score of 0.139 (95% bootstrap CI: 0.127-0.153) and approximately 3.1% prediction disagreement. Extending common relaxation to 50,000 optimizer updates does not erase the measured difference: across five paired seeds, the representation-history score remains 0.190 (95% bootstrap CI: 0.161-0.219) at the end of the measured horizon while the mean accuracy gap is only 0.18 percentage points. Fresh linear probes show that, with sufficient labeled data, the two histories retain practically equivalent linearly accessible class information. A same-label rotated-MNIST control reproduces the effect: all five paired seeds reach behavioral matching while retaining a mean representation-history score of 0.162. Finally, a matched-learning-rate ReLU-LeakyReLU control reduces the 50,000-update representation residue by 0.040 on average in all five paired seeds, providing directional evidence that activation-mediated plasticity contributes to the persistence of training-history effects. These results provide protocol-scoped evidence that behavioral convergence need not imply representational convergence and that optimization history can leave measurable internal traces after prolonged common training.
cs.LG / 201 / 2609.37841
Counterfactual Probing for Parallel Unmasking with Hidden Forest Structure
Abstract
Masked generative models offer parallel token prediction, but accurate parallel sampling must account for dependencies among tokens. When dependencies are unknown, finding safe batches also costs model evaluations. We study whether total evaluations, including discovery, can be sublinear in sequence length $N$; sublinear sequential depth then follows. We consider discrete distributions with hidden forest structure, accessed through a fixed approximate conditional oracle. Under explicit regularity conditions and uniform Hellinger error bounds, for any fixed target accuracy $\varepsilon\in (0,1/8]$ and sufficiently large $N$, our sampler achieves seed-averaged total-variation error at most $\varepsilon$, with total masked-state submissions and sequential depth both bounded by $O(N^C \varepsilon^a)$ for constants $0 <C <1$ and $a > 0$. These guarantees use polynomial vocabulary size and an edge-response lower bound set by $N$ and $\varepsilon$. The sampler shares evaluations of hypothetical reveals across dependence tests to identify safe parallel batches without requiring full recovery of the hidden forest. A tunable parameter trades probing cost against irreversible commit rounds. In the same class, any admissible irreversible product-commit sampler attaining the same seed-averaged accuracy requires $Ω(N^c \varepsilon^b)$ counterfactual submissions or commit rounds in the worst case, for constants $c,b>0$.
cs.LG / 202 / 2609.37865
Strict-Saddle Landscapes and Multi-Rank Geometry in Low-Tubal-Rank Tensor Sensing
Abstract
We study the optimization landscape of low-tubal-rank tensor sensing through a balanced factorization. Under a tubal restricted isometry condition, we establish a quantitative strict-saddle landscape with no spurious local minima for arbitrary Fourier multi-rank profiles. We further show that the local geometry depends on the Fourier-slice ranks rather than the tubal rank alone. Uniform ranks yield quadratic growth transverse to the solution orbit, whereas nonuniform ranks produce quartically flat directions through hidden frequency-wise overparameterization, even when the factor width equals the exact tubal rank. Numerical experiments illustrate the global optimization behavior and the contrasting local geometries.
cs.LG / 203 / 2609.37868
Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) methods such as GRPO rely on successful self-generated trajectories, but finite rollout budgets can produce all-fail groups with no reward-based policy-gradient signal. While additional rollouts improve the chance of success at higher cost, successful trajectories missing from one model's rollouts may already have been discovered by another. Indeed, we observe that heterogeneous models often succeed on complementary prompts, creating opportunities for mutual learning without a designated stronger teacher. To exploit this complementarity, we propose GRAFT (Gated Replacement of Answer-Failed groups with peer Trajectories), an off-policy-aware framework that replaces all-fail groups with informative peer groups. GRAFT transfers both successful and unsuccessful peer responses with peer-computed advantages, while controlling cross-model mismatch through sequence-level compatibility weighting and token-level importance ratio clipping. Across three heterogeneous model pairs and five mathematical reasoning benchmarks, GRAFT consistently improves both models over GRPO with the same per-model rollout budget, gaining 2.1 points on average and up to 4.5 points in model-level average performance. Stored peer trajectories preserve most of the gains, improving over GRPO by 1.8 points on average without simultaneous co-training.
cs.LG / 204 / 2609.37884
TopoEmbedX: A General Framework for Representation Learning on Topological Domains
Abstract
Topological structures such as simplicial complexes, hypergraphs, and cell complexes extend standard graph models by modeling higher-order relationships. These structures appear in many modern datasets and require specialized methods for generating meaningful embeddings. In this paper, we introduce TopoEmbedX, a unified framework for embedding a wide range of topological domains into Euclidean spaces. The package brings together several existing topological embedding algorithms---DeepCell, Cell2Vec, CellDiff2Vec, HOLE, and HOGLEE---and introduces five new algorithms: ComplexNetMF, ComplexRep, ComplexRandNE, ComplexWalklets, and ComplexHeat. These algorithms extend well-known graph embedding techniques to higher-order settings using the augmented Hasse graph of a topological domain. TopoEmbedX provides a clear, consistent, and easy-to-use framework for topological representation learning. Experiments show that the embeddings generated by TopoEmbedX support tasks such as classification and regression across multidimensional data.
cs.LG / 205 / 2609.37887
Behavioral Capacity Certificates for Quantized Language Models
Abstract
Activation and key-value cache precision change what a quantized language model computes without altering its stored weights. Direct weight-code bounds, however, assign identical complexity to deployments that behave differently and charge separately for weight codes that behave identically. Behavioral Capacity Certificates (BCC) charge for behavior using the aggregate prior mass of complete implementations---weights, scales, activation and cache rules---that induce the same bounded loss. When quantization merges implementations, this shared mass lowers the complexity penalty, and a break-even law determines when the saving survives the cost of validating it. BCC supports a three-step deployment workflow, and our experiments verify each step. First, a forward-only screen shortlists per-layer bit-widths by how often candidate perturbations preserve the reference predictions, with quality comparable to Hessian-guided selection at lower preprocessing cost. Second, margin-certified cells identify weights that can be pruned or sign-flipped without changing the deployed behavior: every permitted combination preserves all declared predictions, and on OLMoE-1B-7B and SmolLM2-1.7B, independent probes bound the probability that any permitted combination changes a prediction on new text. Third, BCC bounds the population loss of the deployed model, nonvacuously for complete decoders and more tightly than the compressed-code route. At equal cache memory, giving keys higher precision than values yields lower NLL and higher prediction agreement on GPT-2, Qwen2.5, and SmolLM2, together with a tighter complexity bound in the GPT-2 audit.
cs.LG / 206 / 2609.37899
Scaling Zero-Order Pretraining through Model Sharding
Abstract
Zero-order optimization (ZO) trains without backpropagation, making it relevant to forward-only hardware and non-differentiable loss, but its gradient variance grows with perturbed dimension, inhibiting large-model training. Sharded Optimization Mixture of Assemblies (SOMA) trains LSTM experts independently on $N$ data clusters using simultaneous perturbation stochastic approximation (SPSA), without exchanging gradients, activations or optimizer state. Its separable loss removes cross-expert perturbation noise at the cost of jointly learned representations across domains. Using 80,000 estimated RTX 5090 GPU-hours, we show modest sharding improves training compute efficiency over all tested monolithic ZO controls. At 8.44M parameters and 150 aggregate GPU-hours, SOMA $N=2$ with 64 perturbations reaches 1.76 test nats/byte, versus 2.00--2.11 for monolithic SPSA at 64, 256 or 1,024 perturbations and 2.21 for EGGROLL. On WikiText-103, these frozen checkpoints reach 2.07, 2.25--2.36 and 2.49, respectively. On a fixed separable objective with equal-size blocks, we prove independent losses reduce relative gradient variance to approximately $1/N$ of a shared-loss estimator's. Holding starting weights, data, perturbations and compute fixed, independent rather than summed losses lower SOMA $N=4$ test loss by 0.035 nats/byte after 1,000 updates across three seeds. Larger ensembles offer a separate inference benefit: at similar model size with top-$k$ routing ($k=4$), SOMA $N=256$ achieves 2.36M tokens/s versus 257k for SOMA $N=8$ ($9.19\times$, including routing), at lower test loss (1.68 versus 1.71), albeit using $59.9\times$ as much aggregate training compute. We release all training and evaluation code and checkpoints.
cs.LG / 207 / 2609.37905
Beyond Interaction Capacity: Estimator Scaling with Recursive Models for CTR Prediction
Abstract
Click-Through Rate prediction, a core task in recommendation and advertising systems, relies on modeling interactions among sparse categorical features. Explicit cross networks are a central paradigm for CTR prediction, and recent progress has largely come from increasing the interaction capacity of a single predictor through deeper cross networks and more expressive cross operators. We revisit whether continually increasing interaction capacity remains the most effective way to improve predictive performance, and find that its benefits quickly exhibit diminishing returns even as capacity continues to grow. This motivates a complementary scaling direction that we call estimator scaling, where additional resources are used to incorporate multiple related estimators rather than only enlarging a single predictor. Through theoretical analysis, we show that the gains from estimator scaling are governed by the amount of non-shared predictive variation available across estimators. However, exploiting this variation naively can be expensive: independently trained models provide substantial estimator diversity but require deployment cost to grow with ensemble size. This motivates a parameter-efficient realization of estimator scaling that can incorporate diversity from multiple estimator sources without maintaining multiple full models. Building on this view, we introduce RECursive Averaged Predictor (RECAP), a parameter-efficient recursive CTR model that operationalizes estimator scaling at three levels: distillation across independently trained models, exponential moving averaging over training trajectories, and aggregation over inference-time routes within a weight-shared recursive backbone. Experiments across multiple benchmarks establish new state-of-the-art predictive performance on standard benchmarks, while placing the RECAP on a favorable performance-parameter Pareto frontier.
cs.LG / 208 / 2609.37915
Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning
Abstract
On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies this supervision to unverified student rollouts while conditioning the teacher on privileged context, typically a reference solution. We separate these roles in a factorial analysis and find that scaffold correctness has a stronger effect on downstream accuracy than context correctness. Unverified scaffolds create an imitation gap because the teacher can use information unavailable to the student. This gap shrinks with model scale, yet OPSD continues to supervise mostly unverified trajectories. In contrast, verified scaffolds remain effective even when the teacher is conditioned on the student's own unsuccessful rollout. Based on this finding, we introduce OASIS, which retains the OPSD objective but supervises mostly verified by label on-policy trajectories and replaces written solutions with unverified model-generated attempts as the teacher context. OASIS therefore requires only final-answer labels. Across Qwen3-1.7B, 4B, and 8B on AIME 2024, AIME 2025, and HMMT 2025, OASIS improves over the base model by 3.2--3.8 points on average, while OPSD's gain falls from 3.05 points at 1.7B to 0.14 at 8B. At 8B, OASIS improves over OPSD by 3.05 points, showing that verified on-policy scaffolds preserve the effectiveness of self-distillation as models scale.
cs.LG / 209 / 2609.37917
Search Dimension in Unlabeled Projection Pursuit: A Scaling Law for Subspace Restriction
Abstract
Projection pursuit searches for a direction along which the data look least Gaussian. When the observation space contains a large Gaussian complement, the empirical objective can be minimized by a direction that carries no signal, with empirical kurtosis as low as at the truth. Sample splitting exposes rather than repairs this failure. Appending coordinates independent of the latent regime degrades the search while leaving Bayes recoverability unchanged. Restricting the search to the column space of a known forward operator removes the failure exactly on the negative-kurtosis branch. Estimating a principal subspace from the data is the alternative. In a controlled two-component model, the leading sufficient scalings differ in the gain with which the operator transmits the discriminant: $ς^{-4}$ for covariance-spike estimation and $ς^{-8}$ for fourth-moment search. At fixed search dimension, the measured threshold ratio collapses onto $n/p^2$ with exponent $0.156$, close to the predicted $1/8$. This is an empirically supported scaling motivated by sufficient bounds, not a proved asymptotically tight law. When the search dimension is varied, the measured exponent is $0.325$, substantially larger than $1/8$, and the tested range does not identify its functional form. The crossing location also depends on calibration and model configuration. Under a downstream excess-error criterion, the scaling largely disappears.
cs.LG / 210 / 2609.37921
Pattern Formation in Transformers
Abstract
What are the inductive biases of a Transformer architecture? Existing theory on how the forward pass shapes representations either considers whether Transformers escape from rank collapse or demonstrates that self-attention drives tokens toward cluster patterns. The latter view arises from an elegant dynamical systems perspective, but relies on simplified architectural assumptions, and does not explain the rich structures observed in practice. This leaves a major open question: when a full Transformer escapes rank collapse, how does it structure token representations? Using pattern-formation theory, we show that the dynamical view of Transformers can account for Positional Encoding, Multi-Head Attention, and Output-Value geometry. We demonstrate that a full Transformer architecture imposes an inductive prior by selectively amplifying a rich set of previously unreported patterns, including traveling or rotating waves among others. We characterize the role of each architectural component in controlling which pattern is amplified, which ones stabilize, compete, or coexist. Finally, we show that these structures can act as a controllable dynamical prior that facilitates learning. By choosing both task-aligned positional encoding and weight initialization, we demonstrate improved data efficiency and accelerated optimization on controlled sequence tasks and with ConViT on CIFAR-10.
cs.LG / 211 / 2609.37932
Learning When to Update: A Near-Optimal Timing Bandit Approach
Abstract
Systems operating in dynamic environments require timely updates to sustain performance. For resource-intensive systems such as machine learning models and digital twins, strategically timing updates is essential. Updating too frequently wastes resources, while updating too infrequently leads to costly performance degradation. The problem is particularly challenging when the system's degradation pattern is unknown a priori, as is common in new operating environments. We formalize this challenge as a novel \emph{timing bandit} problem, where each arm represents a candidate update interval with a fixed update cost and an unknown, stochastic degradation cost. Three structural properties distinguish this setting from standard multi-armed bandits: selecting an interval commits the learner to multiple time slots before the next update; arm costs are composed of per-step degradation costs and a fixed update cost; and selecting a longer interval naturally reveals degradation at every intermediate step, providing consecutive feedback relevant to shorter intervals. By exploiting these structures, we develop Balanced Consecutive Arm Elimination (BCAE). BCAE achieves $\tilde{O}(\sqrt{T})$ regret, improving upon the $\tildeΩ(K\sqrt{T})$ regret of standard bandit algorithms in this setting, where $K$ is the number of candidate update intervals. We further propose an Optimism-Enhanced variant (OE-BCAE) that integrates lower-confidence-bound principles to improve empirical adaptivity while preserving the same regret order. Moreover, the regret bound achieved by our algorithms matches the theoretical lower bound up to logarithmic factors. Simulation results demonstrate that our algorithms achieve low regret and remain stable as both the number of arms and the update cost vary.
cs.LG / 212 / 2609.37941
An Efficient Machine Learning Approach for Degradation Forecasting in AEM Water Electrolysis
Abstract
This study provides a data-driven analysis of a novel dataset of single-cell Anion Exchange Membrane water electrolyzers (AEMWE), operated under constant current load across multiple heterogeneous experimental campaigns. We train and evaluate a range of machine learning models with different complexity, including linear baselines, LSTMs and CNNs, to perform medium-term forecasting of the cell voltage degradation curve. The models are assessed within a rigorous training and evaluation framework specifically designed for heterogeneous industrial data.
cs.LG / 213 / 2609.37958
Kolmogorov-Arnold Classifier Systems as Universal Approximators
Abstract
As the input dimension $n$ grows, rule-based machine learning, such as Learning Classifier Systems (LCSs), faces a fundamental scalability bottleneck for function approximation: both rule count and parameter count grow exponentially with $n$. Traditional LCSs partition the $n$-dimensional input space directly, requiring $\mathcal{O}(m^n)$ rules for adequate coverage, where $m$ is the per-variable resolution. This article breaks from this paradigm by reorganizing rules dimension-wise, guided by the Kolmogorov-Arnold representation theorem: any continuous $n$-dimensional function can be expressed as a finite superposition of one-dimensional functions. The proposed Kolmogorov-Arnold Classifier System (KACS) decomposes the target function into one-dimensional subproblems and assigns a dedicated ruleset to each, reducing the worst-case rule count from $\mathcal{O}(m^n)$ to $\mathcal{O}(mn^2)$ and replacing $n$-dimensional local models with one-dimensional models requiring only two parameters per rule, independent of $n$. We also provide the first constructive proof that an LCS, namely KACS, is a universal approximator for continuous functions on compact domains. Evaluated against a direct $n$-dimensional input space partitioning approach under otherwise identical conditions, KACS achieves competitive accuracy in many settings while using only 2\% to 40\% of the parameters. Our implementation is available at https://github.com/YNU-NakataLab/KACS.
cs.LG / 214 / 2609.37959
TabFM: A Zero-Shot Foundation Model for Tabular Data
Abstract
Tabular machine learning typically relies on per-dataset workflows, fitting tree ensembles or running AutoML searches from scratch for every task. We present TabFM, a 400M-parameter tabular foundation model that formulates supervised tabular prediction as in-context learning. TabFM produces calibrated zero-shot predictions in a single forward pass without task-specific tuning. Trained entirely on synthetic tables generated from structural causal models, TabFM learns general tabular representations that transfer zero-shot to real-world tasks. Across all 51 benchmark datasets in TabArena (38 classification and 13 regression), zero-shot TabFM ranks first among default tabular foundation models and outperforms tuned AutoML pipelines. Two extensions over the same frozen weights improve performance further on both tracks: multi-view feature expansion with ensembling and post-hoc calibration (TabFM+), and LLM-guided, dataset-specific data processing and feature engineering (TabFM-Auto).
cs.LG / 215 / 2609.37976
$S^3$: Spectral Null-Space Swap Makes Reasoning Models Efficient
Abstract
LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost. We find that the core of reasoning capacity lies in the Thinking model's weight component within the null space of a projection defined by the corresponding Non-thinking model's dominant singular directions, and removing the subspace component can largely improve reasoning efficiency without hurting the accuracy gained during thinking-mode post-training. Unlike existing efforts that mostly operate within the dominant subspace, we are the first to unveil the critical role of the null space and harness it for model optimization. Motivated by this finding, we propose Spectral Null-Space Swap ($S^3$), a training-free composition of paired Non-thinking and Thinking checkpoints. Our method keeps the Non-thinking model inside its own dominant subspace and takes the Thinking checkpoint outside it, improving reasoning efficiency while maintaining accuracy. We extensively evaluate $S^3$ on 2B-30B dense and mixture-of-experts (MoE) architectures spanning 28 evaluation environments across mathematical, multimodal, and audio reasoning domains. $S^3$ establishes new empirical Pareto Frontiers among training-free model composition strategies: across all settings, it reduces inference token overhead by an average of 27.4% compared to full Thinking models while simultaneously improving overall task accuracy by 1.0 percentage point (e.g., yielding +8.3% accuracy on HMMT25 alongside a 33.0% token speedup). We further use attention entropy for explanation and find that the retained component produces more concentrated attention, and we use a simplified analytical model about optimization to demonstrate why null-space can effectively reduce attention entropy, thereby improving the efficiency of reasoning.
cs.LG / 216 / 2609.37989
TabFM-Auto: Self-Evolving Pipelines for Tabular Foundation Models
Abstract
Tabular foundation models achieve strong zero-shot accuracy on structured data by pretraining on synthetic tables, but they ignore the column names, task descriptions, and auxiliary files that carry dataset semantics. Meanwhile, self-evolving machine learning engineering (MLE) agents train models from scratch on each dataset, yet jointly searching over features, architectures, and hyperparameters is noisy and prone to overfitting. We introduce TabFM-Auto, which pairs a tabular foundation model, TabFM, with a language model agent that evolves the data pipeline around it. Guided by dataset metadata and validation feedback, TabFM-Auto iteratively refines data cleaning, feature engineering, context selection, and post-processing to reduce TabFM's error. Across all 51 datasets of the TabArena benchmark, five TabFM-Auto configurations with different agents and language models take the top five overall positions, and the best raises TabFM from 1785 to 2013 Elo. The discovered pipelines also transfer to other frozen tabular foundation models (+69 to +143 Elo) with no further search. On the 8 tabular competitions of MLE-Bench, TabFM-Auto ranks first overall among MLE agents.
cs.LG / 217 / 2609.38004
No Scale Left Behind: Multi-Scale Autoencoder with Bi-directional Attention for Time Series Anomaly Detection
Abstract
Time series anomaly detection (TSAD) plays a crucial role in healthcare, finance, industrial monitoring, and other sectors. Within and between these settings, anomalies span vastly different temporal scales, from sub-second point spikes to multi-hour drift patterns. However, most existing TSAD methods commit to a single temporal granularity, and multi-scale designs either analyze different scales in isolation or are constrained to a predefined coarse-to-fine hierarchy, both failing to sufficiently capture multi-scale interactions. To resolve this limitation, we propose Multi-Scale Autoencoder with Cross-Scale Attention for TSAD (MSCAD), a simple yet powerful semi-supervised TSAD framework founded on parallel autoencoder branches corresponding to different patch sizes. A stack of symmetric bidirectional cross-scale attention blocks enables every pair of scales to exchange information before reconstruction without allowing any single scale to be privileged. On the comprehensive TSB-AD benchmark (40 datasets, 530 series), MSCAD achieves large performance gains against 50 baselines across multiple metrics, with VUS-PR of 0.57(+9.6%) on the univariate split and 0.47(+9.3%) on the multivariate split compared to the state-of-the-art.
cs.LG / 218 / 2609.38011
When do data mixtures improve scaling laws? Insights from high-dimensional regression
Abstract
Modern machine learning systems are trained on mixtures of data from different domains, and choosing the right mixture can substantially improve downstream performance. Despite an extensive literature on data mixing and reweighting, existing work is largely empirical and it remains unclear when auxiliary data genuinely improves scaling laws rather than merely providing more samples. To gain insight into this question, we study a high-dimensional mixed-data regression model with a shared regression function, heterogeneous covariances and noise levels, and dataset sizes that may grow at different rates. We establish the minimax risk under an ellipsoidal parameter constraint for the general covariance structure and derive deterministic equivalents for the test error of ridge regression under commutative covariances. We then specialize to a target domain and an auxiliary domain with aligned power-law covariance spectra, where the theory yields explicit scaling laws in terms of spectral decay, target regularity, and the relative growth of the two datasets. These laws identify regimes in which combining data mixtures provably yields a faster scaling rate than using either dataset alone. In particular, improving the scaling law requires a specific interplay between spectra and relative sample sizes of the domains. Our numerical experiments on language models exhibit the same qualitative phenomenon: appropriate data mixtures yield a faster decrease in target-domain test loss than training on either domain alone.
cs.LG / 219 / 2609.38018
Prompts Live on an Arc: Gaussian Curricula in Fisher--Rao Coordinates for Rollout-Efficient GRPO
Abstract
Group relative policy optimization (GRPO) learns only from prompts whose sampled responses disagree: a group that is entirely correct or entirely incorrect has zero reward variance, contributes no gradient, and still consumes its rollouts. Prompt-selection methods reduce this waste by steering sampling toward intermediate pass rates, but they choose the target, its width, and the uncertainty model heuristically, in raw pass-rate or logit coordinates. We show that GRPO comes with a natural coordinate for pass rates: the arc length $ψ=\arcsin\sqrt{p}$ on the Bernoulli Fisher--Rao manifold. In arc length, the expected GRPO update is uniform up to two boundary ramps; the probability of a zero-variance group is bounded by two Gaussian boundary layers of width $1/\sqrt{2G}$; pass-rate evidence has constant noise; and the gradients of the pass@$k$ and pass$^k$ objectives are Gaussians whose center and width follow from $k$ in closed form. A prompt curriculum for GRPO is therefore a Gaussian in arc length, and choosing its center amounts to choosing the objective. We turn this observation into ARCUS, a drop-in sampler that tracks every prompt with a Kalman filter in arc length, scores prompts by an objective-matched Gaussian kernel times the predicted probability of an informative group, keeps only informative groups for the unchanged GRPO update, and paces the target toward the hardest objective whose predicted yield stays within a small slack of the best. Across six mathematical reasoning benchmarks and three backbones, ARCUS improves the average accuracy of GRPO by 2.8--2.9 points and that of dynamic sampling by 1.1--1.2 points, while generating 48--57\% fewer rollouts than dynamic sampling.
cs.LG / 220 / 2609.38065
Jaxolotl: A Unified High-Performance Benchmark Suite for LTL-Based Multi-Task RL
Abstract
Training agents to follow arbitrary instructions is an important goal of multi-task reinforcement learning (RL). Linear temporal logic (LTL) provides a precise and structured formalism for specifying instructions to agents, and has been successfully adopted for training generalist multi-task policies. However, differences in implementations, task distributions, and evaluation protocols make existing methods difficult to compare, while high computational costs limit the scale and statistical reliability of experiments. We introduce Jaxolotl, a unified high-performance benchmark suite for multi-task LTL-RL to address these concerns. Jaxolotl provides a modular, end-to-end JAX implementation of six representative algorithms and four environments, together with newly curated task suites and a standardised, statistically robust evaluation protocol. By precompiling symbolic task representations into static arrays, Jaxolotl enables fully JIT-compiled training and evaluation, achieving end-to-end speedups of up to $220\times$ and supporting controlled comparisons at substantially greater experimental scale. We use this framework to systematically evaluate existing approaches, revealing complementary strengths and limitations: general methods capable of non-myopic reasoning struggle as the number of propositions grows, while methods with stronger scaling rely on environment-specific assumptions and suffer from myopia.
cs.LG / 221 / 2609.38067
A foundation model for energy and radiation systems built on heterogeneous scientific interfaces
Abstract
Scientific foundation models are commonly evaluated after heterogeneous physical problems have already been translated into a compatible gridded, tokenized or symbolic representation. This leaves the scientific interface outside both the pretrained model and the audit of what is actually reused. We study the complementary setting in which boundary histories, sparse monitor records and loading histories retain their native inference classes and their outputs remain on Cartesian, latitude-longitude and unstructured domains. GEODE couples task-specific scientific interfaces to a shared routed library of wavelet operators. A single jointly pretrained model represents cavity flow, radiation dose and elastoplastic stress, then acquires a heat exchanger and a reactor subchannel by training a private interface containing 2.1% of its parameters. Earlier predictions remain unchanged by parameter isolation, whereas unrestricted fine-tuning degrades them by factors of 14-29. Crucially, preservation alone does not establish reuse: norm-matched randomized-library controls show that the contribution of pretrained computation is conditional on the task and data regime. A separate decomposition shows that full-field relative L2 error can substantially understate error relative to spatial variation when field level dominates the norm. Task-specific operators remain more accurate on three of the five problems. These results distinguish multi-task coverage, preservation and pretrained reuse as separate properties that must be tested independently when scientific foundation models span heterogeneous interfaces.
cs.LG / 222 / 2609.38081
Traversing the solution space of neural networks with Hessian Null Space Continuation
Abstract
On a single task, deep networks can learn many solutions, depending on their optimizer, training data, architecture, and hyperparameters. Many of these solutions are mode-connected: rather than isolated points in weight space, they are connected by low-loss regions. Yet how their internal computation varies within these regions is unknown. A parallel line of work has identified the degeneracy of neural representations: many networks reach similar training loss with distinct internal structures. However, it is unclear how these solutions are related in weight space. We unify these subfields and show for the first time that many different internal mechanisms exist within a local mode-connected region in weight space. To do so, we introduce Hessian Null Space Continuation (HNC), a scalable method that uses local curvature to traverse regions of weight space that preserve network function, and can be steered toward solutions with specified properties. In RNNs trained on a memory task, HNC reaches drastically different representations and dynamics with maintained behavior. In ImageNet-trained Vision Transformers, HNC finds representations that differ more from the original network than any independently trained model with a different architecture or objective. In reinforcement-learning agents, HNC uncovers a distinct navigation strategy at comparable return and exposes reward hacking in an AI Safety Gridworld. Finally, HNC measures the local geometry of the solution set, showing how model size and task complexity shape its dimension and functional sensitivity. Our results show that a surprisingly large amount of representational diversity exists near a single trained solution, unseen by standard gradient-based optimization. HNC identifies and quantifies this diversity, opening new possibilities for mechanistic understanding of solution spaces and for model merging, editing, and fine-tuning.
cs.LG / 223 / 2609.38090
Mira: Memory-Efficient MoE Inference Using Adaptive Caching and Predictive Expert Staging
Abstract
Mixture-of-Experts (MoE) models are a compelling architecture for scaling model capacity, making them especially attractive for deployment on resource-constrained, single-GPU systems. However, this benefit is difficult to realize because expert parameters dominate memory, and token-level routing is dynamic, unpredictable, and skewed. Prior work using offloading and caching remains fundamentally reactive, as systems wait for router outputs before moving experts, leading to inefficient cache utilization and an inability to overlap transfers with compute under tight VRAM budgets. To address these challenges, we propose Mira, an algorithm-system co-design that enables high-capacity MoE inference on a single GPU. Mira shifts from a reactive to a proactive stance by coupling predictive expert management with a tailored quantization format. It introduces lightweight per-layer predictors that anticipate expert usage two layers ahead, enabling proactive prefetching. These predictions feed a two-tier HOT+STAGE GPU cache managed by token-level routing telemetry to retain frequently used experts while staging predicted ones. To minimize transfer overhead, Mira implements a custom compression for expert parameters, which reduces metadata and improves packing efficiency, while minimally degrading accuracy. Mira is implemented as a fully integrated runtime that coordinates predictors, caching policies, and quantized transfers to maximize overlap between communication and compute. Our experiments show that Mira reduces expert-induced stalls. Compared against state-of-the-art baselines, Mira achieves a 5.71x speedup in average throughput on a memory-constrained GPU. It accelerates Time-to-First-Token by 11.71x and achieves a 3.84$x average speedup in beam search inference, demonstrating its effectiveness across diverse inference scenarios.
cs.LG / 224 / 2609.38094
Dimensionally consistent surrogate modelling through dimensional analysis and harmonic expansions
Abstract
Dimensional homogeneity is a fundamental constraint on physically meaningful models, requiring invariance under changes of units. We present a data-driven method for constructing surrogate models that satisfy this constraint at the level of the hypothesis class. Starting from a dimension matrix of measured variables, the method derives Buckingham $Π$-groups, constructs admissible dimensional prefactors, and approximates the remaining dimensionless dependence using truncated harmonic expansions on normalized invariant domains. Once the prefactor and dictionary are fixed, the coefficients are obtained from a regularized linear regression problem. We test the approach on the simple pendulum, Planck's black-body law, the double-pendulum Lyapunov field, and an experimental COBE/FIRAS black-body spectrum dataset. The results show that dimensional constraints improve conditioning, robustness to noise, and sample efficiency relative to unconstrained baselines, while the choice of dictionary becomes important in non-periodic or multi-invariant settings. The learned expressions are explicit and inexpensive to evaluate, which makes them useful as surrogate models for structured physical problems.
cs.LG / 225 / 2609.38095
Probe-Space Preconditioning for Fast and Stable Zero-Order Training
Abstract
Backpropagation (BP) dominates deep learning but imposes a massive memory tax. For example, training OPT-30B with Adam requires $\approx$ 600GB of GPU memory (assuming batch size 8 and sequence length 2048). Alternatively, zero-order optimization (ZOO) trains in inference-mode (requiring only $\approx$ 60GB for the same model): no stored activations, no gradients, and no optimizer states. However, ZOO convergence has lagged behind BP. In this work, we evaluate two methods to close this gap. First, we show that reallocating training compute budget from many steps to large effective batch sizes with many perturbations (or probes) but fewer steps, allows 1SPSA (Spall, 1992) to outperform zero order methods like MeZO (Malladi et al., 2023) with less training compute. Next, we introduce 1.5-SPSA, adding a single "clean" forward-pass per step to 1SPSA to calculate a cheap diagonal preconditioner in probe-space, which improves convergence rate and convergence by down-weighting high curvature directions. Benchmarking on 6 post-training datasets on both Qwen3 and OPT model families, we show that 1.5-SPSA achieves State-of-the-Art results over previous ZOO solvers with much less optimization steps. For example, we train OPT-13B (for direct comparison to MeZO) and find 1.5-SPSA achieves +3.1% accuracy on SST-2 over both MeZO and BP in only 70 steps vs. MeZO's 100,000 steps. Finally, we combine an 8-bit-packing random generator, triton fused unpack/apply kernels, and distributed parallelism to achieve fast and stable training of models as large as OPT-30B in-place on commodity GPUs (e.g. A100).
cs.LG / 226 / 2609.38096
Tail-Influence Sampling for CVaR Policy Evaluation
Abstract
Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can require many costly rollouts. When different conditional components of a stochastic workflow can be queried separately, we ask how to allocate a fixed evaluation budget to estimate a fixed policy's CVaR most accurately. We derive a tail influence for each queryable conditional law that aggregates how its uncertainty affects CVaR across every Bellman reuse. Its variance yields the fixed-design efficiency bound and the oracle Neyman allocation. Tail-Influence Sampling (TIS) estimates these influence scales from a pilot model and reallocates fresh queries toward kernels that matter most for the tail; a visitation-anchored variant protects against pilot underallocation. Under fixed dimension and a positive quantile margin, TIS attains oracle asymptotic variance and first-order MSE including pilot cost, while the anchored variant is within a factor two of the oracle. We also characterize an exact-grid regime in which tail- and mean-optimal allocations coincide. On CliffWalking, TIS reduces MSE by 41% versus learned occupancy and 76% versus complete rollouts at the same charged transition budget. In frozen language-model review workflows, anchored TIS beats an equally regularized mean-influence blend in 23 of 24 MMLU-Pro settings and reaches 2.4-3.4$\times$ lower MSE than rollouts on six-call FinQA reviews.
cs.LG / 227 / 2609.38121
WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms
Abstract
KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-order statistics of both factors in a matrix product to reduce quantization error. WUSH-KV uses calibration data to construct separate key and value transforms, with the value transform folded into the model weights and the key transform applied after RoPE. The transforms can be paired with clipped quantizers. For one such quantizer, QuEST INT, we show that, under mild assumptions, the WUSH transform is near-optimal. With this quantizer, WUSH-KV reduces layerwise reconstruction error and achieves the lowest end-to-end perplexity among other tested transforms. For end-to-end evaluation, we integrate WUSH-KV into SGLang using OSCAR-style percentile-clipped affine quantization. At 2-bit, WUSH-KV performs comparably to or outperforms the OSCAR transform across all evaluated models and downstream tasks.
cs.LG / 228 / 2609.38132
Achieving an $O(1/N)$ Optimality Gap in Average-Reward Weakly-Coupled MDPs
Abstract
We study average-reward weakly-coupled Markov decision processes (WCMDPs), where a WCMDP consists of $N$ smaller MDPs, called arms, that share multiple per-step budget constraints. We consider the setting where the arms have identical model parameters, multiple actions, and state- and action-dependent costs. For restless bandits (RBs), a well-studied special case of WCMDPs, prior work has developed policies that achieve an $O(1/\sqrt{N})$ optimality gap under general conditions, and has further identified conditions under which policies can achieve a better-than-$1/\sqrt{N}$ optimality gap. However, for general WCMDPs, no prior result achieves an optimality gap better than $1/\sqrt{N}$. In this paper, we identify conditions analogous to those for RBs under which a better-than-$1/\sqrt{N}$ optimality gap is achievable, and design a policy that attains an $O(1/N)$ optimality gap. Notably, unlike prior approaches based on generalizing priority orderings, our policy is not priority-based but rather is designed to induce locally linear mean-field dynamics.
cs.LG / 229 / 2609.38133
Multi-Agent Flow Matching with Decoupled Generative Guidance
Abstract
Generative modeling is widely used for producing diverse objects from complex, multimodal distributions. However, its expressivity does not, in general, come with formal guarantees that the generated objects satisfy hard constraints or requirements. In multi-agent generation, this problem becomes more challenging because a hard requirement can depend on multiple agents, while each agent may need to determine its own guidance input without relying on the simultaneously computed guidance inputs of other agents. To this end, we introduce DeGG-Flow, a general framework for multi-agent flow matching with decoupled generative guidance. By representing the generative process as a control-affine dynamical system, we develop guidance conditions for two classes of coupled requirements: shared requirements whose satisfaction depends on multiple agents together, and private requirements associated with each individual agent dependent on its neighbors. For both classes, we establish feasibility conditions and finite-horizon convergence guarantees. We further derive a Wasserstein bound that characterizes the distributional deviation induced by the guidance. We demonstrate DeGG-Flow on multi-robot collaboration for crossing a spatial gap by reconfiguring the environment, and on multi-object scene generation with affordance requirements. Across both applications, DeGG-Flow directly generates objects that satisfy all corresponding hard requirements, including at team sizes unseen during training.
cs.LG / 230 / 2609.38161
A Spectral Theory of Distortion in LLM Graph Reconstruction: Sharp Bounds and Empirical Characterization
Abstract
Evaluations of graph reconstruction by language models typically report a single aggregate distance between the original and the reconstructed graph. We prove that for the Wasserstein distance between Laplacian spectra such a summary is bracketed by two edge counts, the net change in edge number from below and the symmetric difference from above, each scaled by $2/n$ where $n$ is the number of vertices. The bracket is sharp: its two ends coincide exactly when the reconstruction only adds edges or only deletes them, and on that class the distance is a rescaled edge count that says nothing about which edges changed. When the ends differ, the residual between the distance and the lower end is positive only if the reconstruction both invented and lost edges, which turns it into a certificate of mixed editing computable from the reported summaries alone. We characterize these regimes in 135 reconstructions produced by three open-weight models over 45 synthetic graphs. Seventy-seven outputs are one-sided and 29 mixed outputs have $X > 0$, including cases where edge count is exactly preserved while nineteen edges were simultaneously invented and lost. The three models differ in editing policy, ranging from copying the input to attempting completion at the cost of large hallucination volume, a distinction that aggregate distortion does not reveal.
cs.LG / 231 / 2609.38166
LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization
Abstract
Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the cost of long-context processing, repeatedly reading and updating that state remains a major inference bottleneck. Quantization offers a natural way to reduce this cost, but can significantly degrade model quality, due to the accumulation of rounding errors and the presence of outlier rows and columns in the state. To address these challenges, we propose LeapQuant, a training-free method that achieves near-lossless performance under 8-bit recurrent-state quantization. First, to mitigate error accumulation, we propose per-window quantization, which leaps over a window of tokens and quantizes the state only once at its end. Within a window, outputs are computed from the fixed low-bit state together with high-precision buffered updates. Second, to reduce the error introduced by each quantization, LeapQuant retains the state's largest outliers as a few high-precision Compensator Tokens, which share the update path of real tokens. We then smooth the remaining residual before quantization to further reduce the error. Comprehensive experiments across the Qwen, Kimi, and GLM model families show that LeapQuant substantially reduces memory and compute costs during inference. With accuracy comparable to the FP32 baseline, it achieves average speedups of 2.05--3.70$\times$ at the kernel level and 1.47$\times$ for end-to-end inference on NVIDIA B200, RTX PRO 6000, and RTX 5090 GPUs.
cs.LG / 232 / 2609.38176
Breakdown of Local Denoising as Semantic Speciation
Abstract
The dynamics of generative models exhibit two apparently distinct temporal windows: a speciation window, in which a sample commits to a semantic class, and a nonlocality window, in which local context windows become insufficient for generation. Motivated by evidence of their near-concurrence in a variety of frontier models, we investigate their relationship through the spatial distribution of semantic information. Under a "common cause" hypothesis, we prove that the nonlocality window must lie in the speciation window. This hypothesis postulates that semantic labels explain a fraction of the correlations between distant tokens, a condition that is natural for many real datasets. We further give conditions under which both windows shrink to a single limiting time as system size grows, defining a "phase transition", and verify this behavior analytically in Gaussian mixtures. Together, these results identify conditions under which semantic information explains the concurrence of speciation and nonlocality, connecting two complementary perspectives on the emergence of semantic structure in generative modeling.
cs.LG / 233 / 2609.36012
In-Context Learning for Robots: Methods and Applications
Abstract
General-purpose robots must infer what a new task requires and translate that understanding into appropriate physical action. In-context learning (ICL) for robots supports this process by using demonstrations and interaction to direct existing competence with neural parameters held fixed during deployment. We organize this literature review around the interfaces connecting contextual evidence to execution, distinguishing four families: context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill- and agent-based execution. Comparing these interfaces clarifies their transfer assumptions and the roles of training, correspondence, and memory in making context useful. Across manipulation and navigation, we examine how these mechanisms preserve taught requirements as objects, environments, and execution conditions change. This analysis links method design to evaluation practices that distinguish responsiveness to teaching, physical transfer, and benefits from retained experience. The resulting agenda connects compositional task acquisition and faithful transfer with physical recursive self-improvement, in which experience improves the ability to learn subsequent tasks.
cs.LG / 234 / 2609.36426
Losing the name before the box: measuring and repairing what narrow fine-tuning costs a detector outside its deployment vocabulary
Abstract
A detector pretrained on a broad corpus is fine-tuned on a narrow domain, its in-domain accuracy improves, and it ships. We ask what happens meanwhile to its coverage of objects the vocabulary never names, which in obstacle detection and inspection carry the risk. No in-domain test set holds an example of one. We give a longitudinal protocol: one pretrained checkpoint against its own fine-tuned descendants. It tracks held-out top-$K$ proposal coverage $C_τ$: of categories pretraining covered and the vocabulary omits, the share of boxes a detector's top $K$ regions still cover. The quantity is the open-world proposal literature's; the longitudinal reading is not. $C_τ$ falls while in-domain accuracy rises, on four architectures and three domains, by $5.12$ to $63.35$ points on boxes above $1024$ px$^2$. No in-domain number identifies the fall, and neither does detection average precision, which charges a missed and a misnamed box alike. On the one architecture scoring both, adaptation costs $87\%$ of the AP against a fifth of the coverage, and the naming goes first at all six depths of its freeze ladder, every run. What breaks is structured: three architectures sharing no pretraining run agree on which categories lose coverage, and those a model never learned do not lose any. A repair follows and needs no training: mixing a quarter of the pretrained state back, normalisation statistics included, raises coverage on every cell swept for at most $2.47$ points of in-domain accuracy. Seeing it costs one extra evaluation pass.
cs.LG / 235 / 2609.36540
Reactive Real-Time Flow Policies via Asynchronous Distribution Alignment
Abstract
Generalist robot policies such as vision-language-action models (VLAs) have achieved remarkable generalization, but their inference delays can conflict with the demands of real-time control. Asynchronous execution avoids pauses between action chunks by predicting the next sequence of actions while the robot carries out the previous one. In this paper, we study whether asynchronous execution produces the same action distribution as the original VLA. We find that, for non-Markovian demonstrations, asynchronous execution can produce a fundamentally different action distribution, which can limit the policy's reactivity. In our method, we seek to restore this reactivity by aligning the asynchronously produced action distribution with that of the original VLA through two complementary mechanisms. First, Recursive Flow-Field Distillation trains the asynchronous policy using the VLA's action-generation flow. We characterize the learned distribution theoretically and show experimentally that our asynchronous policy can generate nearly the full range of actions the original VLA would produce, while existing asynchronous methods recover only a fraction of that range. Second, Propose-Resolve prepares multiple action sequences asynchronously and uses the latest observation to select among them based on a lightweight approximation of their likelihood under the VLA's action distribution. Our resulting method matches the original VLA's success on LIBERO and retains about 80% of its success on RoboMimic, about 30 percentage points more than existing asynchronous methods.
cs.LG / 236 / 2609.37165
Disentangling Spurious Correlations in Vision-Language-Action Models via Predicting Domain-Invariant Latent Lookahead
Abstract
Vision-Language-Action (VLA) models remain brittle under visual distribution shifts, often relying on spurious correlations tied to domain-specific factors rather than task-relevant structure. We propose Domain-Invariant Latent Lookahead (DILL), a representation-learning framework that mitigates shortcut learning in VLA policies. Our key idea is to supervise policies with domain-invariant future latents learned from domain-transformed trajectory data. A Task-Domain Encoder is trained with contrastive objectives and Gaussian disentanglement regularization to separate task-relevant structure from domain-specific visual variation. The learned encoder then provides future latents for VLA policy learning through lookahead prediction and domain disentanglement, encouraging the policy to focus on task-relevant structure rather than incidental visual factors. Counterfactual task-view evaluations show that DILL reduces shortcut reliance, while LIBERO-Plus evaluations demonstrate improved visual robustness, with 69.1% average success, 11.4 percentage points above the strongest baseline. Real-world manipulation experiments further support DILL's applicability beyond controlled simulation. Complementary latent-space diagnostics show that these behavioral gains are accompanied by representations that better preserve task-consistent structure while suppressing domain-specific variation. Our project page is available at https://dill-vla.github.io/.
cs.LG / 237 / 2609.37599
BlenDAgger: Blended Shared Control for Interactive Imitation Learning
Abstract
Robot policies are frequently trained from human corrections, yet teleoperating a robot to provide corrections is burdensome, and human demonstrators are not always optimal. We propose Blended DAgger (BlenDAgger), an approach for collecting data to train imitation learning policies by using shared control to blend the policy's and demonstrator's actions during interventions. By blending human and policy actions, we aim to improve the autonomous performance of manipulation policies. We validate our approach across five manipulation tasks, two in the real world and three in simulation. Our approach achieves higher autonomous performance by 30 or more percentage points on two real-world tasks compared to a typical human-gated correction approach (HG-DAgger). We also investigate the advantages of BlenDAgger that allow for higher autonomous performance, finding that BlenDAgger results in 57% smoother transitions between policy control and human interventions, and 14% higher trajectory similarity to the training data. In a user study (n=14) on two real-world tasks, we find that BlenDAgger results in faster data collection (BF=13.32), and we do not find a difference in subjective perceptions. These results show that blended shared control leads to higher autonomous performance compared to typical methods for fine-tuning robot policies from fully teleoperated interventions.
cs.LG / 238 / 2609.37677
Learning Expressive and Compositional Motion Representation via Spectral Skills
Abstract
Robotic foundation models offer a promising path toward general-purpose humanoid robot control, often through hierarchical architectures. However, their effectiveness depends on the command interface between the planner and the controller, which must support accurate execution while remaining easy to predict, and ideally allow new behaviors to be composed from prior ones. In this work, we introduce spectral skills, a latent representation of this interface that meets these requirements through predictive representation learning. By design, spectral skills compactly encode short motion segments and are learned by predicting subsequent motion rather than reconstructing the encoder input. On a 29-DoF humanoid, a controller conditioned on spectral skills reduces global tracking error by 62\% relative to the state of the art. The same frozen controller chains independently encoded skills without a separate transition policy. It also composes new behaviors by adding orthogonal directions to any compatible base skill, producing combinations unseen in the training data. We demonstrate tracking, chaining, and composition, as well as control through a language-conditioned planner, on Unitree G1 hardware. Project page: https://spectral-skill.github.io
cs.LG / 239 / 2609.36324
Distill Locally, Schedule Globally: Flow Maps for Few-Step Text-to-Speech
Abstract
Flow-matching text-to-speech (TTS) models achieve high synthesis quality but require many neural function evaluations (NFEs) to integrate their generative trajectories. Recent few-step flow-map distillation approaches for TTS construct targets from numerically integrated teacher trajectories, creating a trade-off between target accuracy and training cost. We propose Local Flow-Map Distillation (LFMD), which adapts Eulerian Map Distillation to conditional TTS and avoids teacher trajectory integration during target construction. For inference, we derive a sampling schedule (TD-DP) from teacher dynamics and consistency of the learned maps, with a single cost graph supporting multiple NFE budgets without external audio-metric evaluation. Because scheduling offers no flexibility at one NFE, we refine this regime with alignment-aware temporal self-distillation using soft-DTW. Across Seed-TTS and LibriSpeech-PC, LFMD improves low-NFE synthesis over a matched integral-distillation baseline. On Seed-TTS, the refined student reaches 1.80% WER with 1-NFE, compared with 1.76% for its 32-NFE teacher.
cs.LG / 240 / 2609.36400
CAMEO: A Class-Activation-Mapped Equitable Overlay Framework for Fair and Robust Deep Learning-based Skin Condition Diagnosis
Abstract
Deep learning classifiers for dermoscopic skin lesions often reach high in-distribution accuracy while quietly relying on spurious background cues such as skin tone, device vignetting, and embedded rulers, rather than on lesion morphology. This undermines robustness and fairness across skin tones. This work asks whether Explainable AI (XAI), typically used only to audit a finished model, can instead be repurposed as an active training signal that corrects this shortcut without sacrificing diagnostic accuracy. We introduce CAMEO (Class Activation Mapped Equitable Overlay), a framework that improves skin-lesion classification by selecting stable model explanations and using them to separate lesions from their backgrounds. It then replaces the background with realistic synthetic skin while keeping the lesion unchanged. On HAM10000 and dark-skin ISIC images, CAMEO maintained accuracy while reducing background-driven errors by nearly four times. It also made the model's attention more consistent when backgrounds changed. Results across multiple tests show that reducing reliance on background information improves robustness, with Fitzpatrick-based backgrounds providing a realistic and interpretable approach. Results show that XAI-guided augmentation can make dermoscopic classifiers measurably more robust and fair at no cost to accuracy. They also clarify that it is the mechanism and not the specific tone palette that matters, and that the lasting contribution of XAI here lies in stability-screened, annotation-free lesion localisation rather than in the robustness number itself.
cs.LG / 241 / 2609.36106
Fundamental Limits of Transferability and Equivariance in Algebraic Signal Models I: Finite Dimensions
Abstract
We study the fundamental limits of transferability in algebraic signal processing through homomorphisms between algebraic signal models. Homomorphisms are linear maps between the signal spaces of two models that commute with filtering, so filtering a signal and transferring it across domains can be done in either order. The existence of such maps is governed entirely by coincidences among the filtered eigenvalues of the two models' shift operators, but existence alone is insufficient: the space of homomorphisms always contains trivial elements that destroy all information. We introduce the spectral transfer efficiency $η(θ)\in[0,1]$ to quantify information-preserving quality, prove that every homomorphism decomposes into unconstrained blocks over coincidence classes, derive the dimension of the homomorphism space, and characterize exactly when lossless transfer is achievable. Beyond normal shift operators, we quantify a departure-from-normality penalty and show how filter derivatives can repair spectral defectiveness. The theory yields concrete consequences in three settings: for sampling, eigenvalue interlacing converts transferability under subsampling into an explicit filter design constraint; for compressed sensing, $η(θ)$ controls the restricted isometry constant and the coherence of the resulting measurements, and recovery decouples across coincidence classes; and for machine learning, spectral aliasing emerges as the controlled symmetry breaking that makes transfer between mismatched domains possible at all.
cs.LG / 242 / 2609.36609
Best Practices in EEG Analysis: Preprocessing, Modeling, and Machine Learning
Abstract
Electroencephalography (EEG) analysis requires careful choices in preprocessing, statistical modeling, and machine learning because EEG signals are highly susceptible to artifacts, volume conduction, low signal-to-noise ratio, and substantial inter-subject variability. This chapter provides a practical and methodological guide to modern EEG analysis, spanning EEG preprocessing, artifact removal, filtering, bad-channel detection and interpolation, re-referencing, independent component analysis (ICA), and preprocessing of simultaneous EEG-fMRI recordings. We review major approaches for computational EEG analysis, including event-related potentials (ERPs), time-frequency analysis, functional and effective connectivity, source localization, multivariate decoding, permutation testing, and multiple-comparison correction. We then examine machine-learning methods for EEG, from feature-based classifiers to deep learning and emerging EEG foundation models, with emphasis on cross-subject generalization, limited-data regimes, data leakage, evaluation metrics, and fair benchmarking. Reproducibility is treated as a core requirement throughout, including transparent preprocessing, BIDS-EEG data organization, standardized derivatives, preservation of raw data, and FAIR data practices. The chapter is intended as a practical reference for researchers developing reliable, interpretable, and reproducible EEG analysis and machine-learning pipelines.
cs.LG / 243 / 2609.37277
Geometry-Aided Channel Deduction with Partial Channel Estimates and Uncalibrated Digital Twin
Abstract
The acquisition of high-dimensional channel state information (CSI) in wireless MIMO-OFDM communications usually requires high pilot overhead, or relies on accurate and complete positional or environmental information. In this paper, we propose a geometry-aided channel deduction (GCD) approach, which utilizes an uncalibrated digital twin (DT) with only approximate environmental geometry and positions to assist the channel acquisition. The key rationale behind is that, even imprecise geometric information, which can be easily obtained in advance through radio sensing technologies or existing geographic databases, provides certain structural features about the current channel; meanwhile, the coarse instantaneous channel estimates using only a small amount of pilots provide dedicated information that aligns with the channel structure and further compensates for the geometry inaccuracy and other channel unknowns. To this end, we first extract geometric features from the DT, which contain only simple structural information of the channel. Then we propose random prompt augmentation, a novel method to generate an appropriate prompt that converts geometric multi-path structure into a CSI-like representation while suppressing the disturbance of other unknown channel parameters. The prompt is then fused with the pilot-based instantaneous channel estimate via a channel deduction network. To further enhance the network's versatility, we incorporate pilot configurations into the existing learning architecture to support variable pilot patterns. Comprehensive experiments validate the superiority of the proposed method, which demonstrates high channel acquisition quality, low pilot overhead, and strong robustness. Furthermore, the structural prompt also serves as scenario-related context, enabling our approach to generalize well in new scenarios.
cs.LG / 244 / 2609.37551
End-to-End Optical Semantic Communication over a Nonlinear WDM Fiber Link
Abstract
Emerging optical-network applications increasingly use received data for inference and control rather than exact source reproduction, creating an opportunity to trade bit-level fidelity for greater transmission reach and efficiency. We propose an end-to-end optical semantic communication system for joint image classification and reconstruction over a nonlinear wavelength-division multiplexed (WDM) fiber channel. The system maps each image directly into a fixed-length sequence of channel symbols that preserves task-relevant information, without explicit source compression or channel coding. Experiments on the MNIST dataset cover launch powers from -9 to +3 dBm, fiber lengths up to 800 km, and 16-, 64-, and 256-Quadrature Amplitude Modulation (QAM) formats. At 0 dBm, classification accuracy remains between 98.92% and 99.31% across all tested link lengths and modulation orders, while requiring fewer transmitted symbols than a Low-Density Parity-Check (LDPC)-coded JPEG baseline at every tested modulation order. These results show that semantic communication can simultaneously extend optical reach and reduce transmission resources by conveying only task-relevant information.
cs.LG / 245 / 2609.36712
Information-theoretic receding-horizon active learning of nonlinear dynamical systems
Abstract
Accurately learning nonlinear dynamics from a finite-duration experiment requires the efficient collection of informative data. We address this challenge for stochastic controlled nonlinear dynamical systems whose state is observed along a single trajectory. Our goal is to reconstruct the unknown controlled state-increment map over a prescribed compact subset of state-input space. We construct a parametric estimator of the map using fixed nonlinear features, so that the model is nonlinear in the state and input, but linear in the unknown parameters. A Gaussian prior over the parameters yields recursive Bayesian posterior updates as data stream in, enabling online quantification of predictive uncertainty in the reconstructed dynamics over the target set. We formulate an optimal adaptive-design problem over an information state, using a prediction-oriented acquisition criterion based on the mean marginal mutual information between candidate future trajectories and the reconstructed dynamics over the target set. We then approximate the resulting adaptive-design problem by a non-myopic receding-horizon formulation, evaluate its remaining expectation using a scenario-based sample average, and solve the resulting deterministic program with the cross-entropy method, leveraging parallel candidate-scenario evaluations. Numerical experiments on a noisy multistable system demonstrate that the proposed adaptive information-seeking strategy reduces predictive uncertainty and reconstruction error more efficiently than common excitation baselines under comparable experimental constraints.
cs.LG / 246 / 2609.36044
Hybrid Neural Simulation-Based Inference for Robust Applications and Limited-Budget Scenarios
Abstract
We develop two hybrid techniques that approach the performance of neural simulation-based inference (NSBI) analyses while substantially reducing the computational cost of inference and preserving some or all of the reliability guarantees of parametric methods. The first approach is broadly applicable, while the second is tailored to a class of particle physics analyses that admit a semi-parametric NSBI formulation. With only a modest compromise in raw sensitivity, these methods represent an important step toward computationally efficient NSBI in offline analyses and also open the door to the exploration of trigger-level applications in the future. Based on our comparison studies, we recommend the use of our first approach, Latent Categories, for robust and efficient inference.
cs.LG / 247 / 2609.38035
The finite-horizon five-expert prediction problem
Abstract
We give an explicit solution to the five expert prediction with expert advice partial differential equation (PDE) in the finite-time horizon setting. The solution formula establishes that the adversary's rank strategy $(1,0,1,0,0)$ is globally optimal, and the COMB strategy $(1,0,1,0,1)$ is optimal exactly on the set where $x_1=x_2$ and $x_3=x_4$. The formula is derived from the solution of the geometric-stopping problem given in our companion paper through the transform principle of Bayraktar, Ekren and Zhang, which links the two problems by a Laplace transform. Inverting the transform term by term expresses the solution through a series of Gaussian and complementary error function kernels. The optimality of $(1,0,1,0,0)$ is reduced to the signs of $41$ one-variable Gaussian series, which are certified with computer assistance by Poisson summation, first-mode domination and interval arithmetic on $1616$ rational cells. The proofs of our main theorems, certificates included, are also formalized in the Lean proof assistant.
cs.LG / 248 / 2609.36615
CI-PINN: Causal Integral Physics-Informed Neural Network for Solving Evolution Equations
Abstract
Physics-informed neural networks (PINNs) solve partial differential equations (PDEs) by incorporating governing physical laws into the training loss. For evolution equations, however, their conventional pointwise space--time representation does not explicitly encode temporal dependence, which can hinder accurate prediction. To mitigate this limitation, this work proposes a novel neural architecture termed a causal integral neural network (CinNet). The core module of CinNet is a Volterra-type causal integral term, which aggregates historical features to encode temporal dependence, thereby incorporating temporal causality at the architectural level rather than through training-level modifications as in many existing methods. Building on CinNet, we further develop a causal integral physics-informed neural network (CI-PINN) for solving evolution equations. Extensive numerical experiments on benchmark evolution equations demonstrate that the presented method outperforms various baseline PINN variants in terms of solution accuracy, with pronounced superiority under sparse-collocation scenarios. Additional empirical analyses show that CI-PINN exhibits low sensitivity to hyperparameter choices, while ablation studies confirm the effectiveness of the proposed network components.
cs.LG / 249 / 2609.36709
Into the danger zone: stable extrapolation in high-dimensional function and operator learning
Abstract
Out-of-distribution (OOD) generalization is a central challenge in scientific machine learning. We study regression problems in which the test distribution differs from the training distribution and ask: under what assumptions on the target function or operator is stable extrapolation possible, and how far beyond the training domain can one extrapolate? Existing theory controls the test error through additive penalties measuring the discrepancy between the training and test distributions. Such guarantees show robustness to small distribution shifts, but can very pessimistic in comparison to OOD performance observed empirically. We identify classes of holomorphic functions and operators for which the OOD generalization error converges at algebraic rates even in the presence of large distribution shifts. This phenomenon stems from the increasing smoothness of higher-index coordinates, leading to what we term a `blessing of high dimensionality'. For learning with either polynomials, deep neural networks or deep neural operators, we derive explicit rates for arbitrary test measures supported on suitable domains and quantify how the admissible domain depends on the underlying regularity of the function or operator. Our extrapolation guarantees are independent of the test distribution, depending only on its support. We also present a series of numerical experiments across a range of functions and operators that support the main theoretical findings.
cs.LG / 250 / 2609.37308
Hybrid Joint-Selective Optimization: Reduced-Space Levenberg-Marquardt Refinement of Low-Dimensional Parameters of Interest
Abstract
This paper introduces a hybrid joint-selective optimization (HJSO) framework for large-scale numerical problems in which a small subset of trainable quantities is of primary interest. We partition the full parameter vector into a high-dimensional remaining block and a low-dimensional block of parameters of interest (POIs), perform joint first-order optimization over the full parameter set, and then freeze the remaining variables while applying a reduced-space Levenberg-Marquardt (LM) refinement to the POIs. The method is designed for settings in which the POIs are low-dimensional but strongly influence the quality of the computed solution, while the full parameter space remains too large for full-space second-order methods. The framework is evaluated on three representative problems: a matrix eigenvalue problem, an inverse Bratu problem solved with a physics-informed neural network, and a 100-dimensional nonlinear Black-Scholes problem solved with the DeepBSDE method. In each test, HJSO reaches prescribed POI-error thresholds faster than the corresponding joint first-order baseline and improves the final POI accuracy for the reported solver configurations. The contribution is therefore not a universal optimizer, but a practical reduced-space strategy for problems with known low-dimensional parameters of interest and expensive high-dimensional training variables.
cs.LG / 251 / 2609.36033
Dual-Anchor Acceleration Is Near-Optimal for Stochastic Monotone Root-Finding
Abstract
Among distinct optimal acceleration mechanisms for deterministic monotone root-finding problems and fixed-point problems, dual-anchoring has recently been shown to admit a more robust direct stochastic extension than standard anchor acceleration. However, without additional strong monotonicity, the existing stochastic dual-anchoring guarantee has two limitations: first, it requires cocoercivity in expectation, and second, it attains only $O(ε^{-3})$ oracle complexity, leaving a gap to the near-optimal $\tilde{O}(ε^{-2})$ complexity achieved by other methods. In this work, we address both of these limitations by combining dual-anchoring with stochastic resolvent approximation and optimized variance control. For unbiased stochastic oracles with variance bounded by $σ^2$, where sample operators are monotone and uniformly $L$-Lipschitz, our algorithm finds a point with $ε$-residual with a near-optimal oracle complexity of $O ( (LD / ε) \ell + (σ^2 / ε^2) \ell^2)$, where $\ell = \log (1 + LD / ε)$ and $D$ is the initial distance to a solution. This result improves the best known oracle complexity in the noise-dominated regime under these samplewise assumptions, reducing the poly-logarithmic factor from cubic to quadratic.
cs.LG / 252 / 2609.36165
Tensor-Train Compressed Separable PINNs: A Curvature-Aware Optimization Framework for Parametric PDEs in High Dimensions
Abstract
In this work, we develop a second-order optimization framework for physics-informed neural networks (PINNs) applied to high-dimensional parametric partial differential equations (PDEs). The framework is built on the Gauss--Newton pullback metric, which provides an operator-informed notion of curvature in parameter space and connects the method to the broader family of natural gradient schemes. We show that, for coordinate-separable neural architectures and linear differential operators (or linearized operators in the nonlinear case) admitting a finite separable representation, the residual Jacobian inherits a structured separable factorization. This yields an exact compressed formulation of the Gauss--Newton step in a reduced space, without assembling the full residual Jacobian on the exponentially large tensor-product collocation grid. The dimension of the reduced space (the effective compressed dimension) is determined by the local collocation grid sizes, the separable operator structure, and the contraction pattern of the architecture, thereby avoiding dependence on the full tensor-product grid size and replacing dense linear algebra in parameter space by a substantially smaller structured problem. Within our framework, we investigate canonical polyadic and tensor-train parametrizations and derive their full algebraic characterization relevant to the Gauss--Newton method, including the structure of the residual Jacobian, the resulting compressed system, and its effective compressed dimension. Numerical experiments on high-dimensional PDEs, including parametric problems, demonstrate the high efficiency of the proposed compressed Gauss--Newton method, which achieves substantially lower errors than tensor-compressed first-order baselines with orders of magnitude fewer iterations and only a fraction of the computing time.
cs.LG / 253 / 2609.37733
Foundation Neural-Network Quantum States for Molecular Potential Energy Surfaces in Second Quantization
Abstract
Second-quantized neural-network quantum states have achieved accurate molecular energies, but extending them across molecular geometries requires a shared representation of the geometry-dependent wavefunction coefficients. We introduce geometry-conditioned foundation neural-network quantum states for molecular electronic structure in second quantization. A single autoregressive model learns a family of ground states from sparse anchor geometries and provides wavefunctions at untrained geometries without further optimization. Orbital alignment matches orbital identities and transports their phases, establishing an aligned orbital basis across geometries. Frozen energies reach chemical accuracy at every untrained query geometry for N$_2$, CO, and H$_4$. On additional molecular paths, the energy-trained wavefunctions yield dipoles, quadrupoles, and natural occupations without property labels. Across three paired N$_2$ training seeds, orbital alignment lowers the mean absolute energy error over all untrained query geometries from 34-37 mHa to 0.049-0.085 mHa. At approximately 1 mHa mean absolute error, frozen evaluation reduces the per-geometry cost by $986\times$ relative to independent optimization, yielding an estimated $25.8\times$ end-to-end GPU-cost reduction on a 161-point N$_2$ grid.
cs.LG / 254 / 2609.37202
Probabilistic Symbolic-Distillation Model of Droplet Collision for Spray Simulation at High Ambient Pressures
Abstract
Droplet collision governs droplet population dynamics in many chemical engineering processes, such as spray drying, spray cooling, agricultural spraying, and combustion. Existing analytical models impose deterministic, pairwise boundaries between collision outcomes, whereas machine-learning classifiers lack the explicit functional form required of analytical collision submodels. In this study, we develop a probabilistic symbolic-distillation model using nearly forty thousand experimental events spanning eight regimes and five dimensionless parameters, including over five thousand data for ambient pressure up to 50 atm. A machine-learning teacher learns the joint outcome-probability landscape from these data, and symbolic regression subsequently distils it into eight class-specific expressions that jointly define a coupled analytical model. The resulting analytical field replaces abrupt regime switching with finite-width fuzzy boundaries. It outperforms the evaluated conventional analytical boundary models and reveals that their main limitation is the inability of zero-width boundaries to represent gradual probability transitions. The "biased-dice" sampling scheme provides a statistically consistent and practically convenient model implementation for Eulerian-Lagrangian spray simulation.
cs.LG / 255 / 2609.36885
RNA Design via Conditioned Flow Matching and Finite-Policy Reinforcement Learning
Abstract
RNA design aims to identify sequences that fold into specified secondary structures. Existing methods formulate the task as target-specific search or conditional generation. However, natural RNA evolution proceeds through sequence variation and selection, with compensatory substitutions, whereas these methods do not explicitly model this process. To address this limitation, we propose a two-stage framework comprising RNA Inverse-Folding Flow (RNA-IFlow) and RNA-IFlow-RL. RNA-IFlow uses structure-conditioned Dirichlet Flow Matching to model coordinated variation across the sequence, while RNA-IFlow-RL maps the learned flow to a pairing-preserving finite policy and refines it with thermodynamic feedback. Our framework achieves leading performance on multiple benchmarks, reaching 85.19% Pass@1 on Rfam-27. Further analyses reveal thermodynamic gains, policy dynamics, and robustness across settings. Our work couples coordinated variation with thermodynamic selection, offering a novel paradigm for RNA design.
cs.LG / 256 / 2609.38073
Optimal Quantum-Classical Separations for Exact Learning
Abstract
We study exact learning with membership queries for concept classes $\mathcal C\subseteq\{0,1\}^N$, focusing on the relationships among their deterministic, randomized, and quantum query complexities, denoted $\mathsf{D}(\mathcal C)$, $\mathsf{R}(\mathcal C)$, and $\mathsf{Q}(\mathcal C)$, respectively. The two canonical quantum speedups in this model are witnessed by Grover search and Bernstein-Vazirani, leading to the longstanding conjecture $$ \mathsf{R}(\mathcal C)=O(\mathsf{Q}(\mathcal C)^2+\mathsf{Q}(\mathcal C)\log N). $$ We first refute this conjecture by constructing concept classes $\mathcal C$ and $\mathcal C'$ satisfying \[ \mathsf{R}(\mathcal C)=Ω\!\left(\frac{\mathsf{Q}(\mathcal C)^3\log N}{\log \mathsf{Q}(\mathcal C)}\right) \qquad\text{and}\qquad \mathsf{D}(\mathcal C')=Ω(\mathsf{Q}(\mathcal C')^3\log N). \] The first bound matches the upper bound of Arunachalam et al.~[Quantum'21] up to constant factors, while the second matches the upper bound of Servedio and Gortler~[SICOMP'04]. In particular, this shows that the saving in the randomized upper bound of Arunachalam et al. fundamentally relies on randomness. Apart from characterizing the optimal relationship between classical and quantum query complexity, our results are the first to show that quantum speedups for learning can go beyond the Grover and Bernstein-Vazirani paradigms.
cs.LG / 257 / 2609.36390
Finite-Sample Theory for Fitted Q-Iteration When Actions Are Functions
Abstract
Offline reinforcement learning seeks optimal decision rules from previously collected data. In some applications, a decision can be an entire function, such as a fluence map in radiation therapy or a smooth movement trajectory in robotics. In this paper, we study the finite-sample theory for fitted Q-iteration (FQI) with functional actions in a discounted infinite-horizon setting. Three major difficulties arise in this setting: first, the absence of a Lebesgue probability density for functional actions complicates coverage descriptions; second, conventional coverage requirements can be restrictive; and third, the large functional action space makes greedy optimization in FQI challenging. To address these difficulties, we study smoothness-regularized policy search under a critic-relative coverage condition. This condition measures how well logged data distinguish relevant action-value differences without requiring an action density. Our main theorem gives finite-sample guarantees for learned-policy regret relative to the best value within a fixed smooth class of functional-action policies. The results allow trajectory lengths to be either bounded or growing and Q-functions to be fitted by either functional-input kernel ridge regression or adaptive functional neural networks. For a few examples, we can obtain polynomially decaying regret bounds in the number of logged transitions, up to logarithmic factors, with logarithmically many FQI iterations. Numerical experiments show gains of learned functional-action policies over constant-action policies and support our adoption of a critic-relative coverage condition.
cs.LG / 258 / 2609.36227
One-Step Next-Latent Prediction Is Not a World Model
Abstract
Next-latent prediction fits a map from the current embedding to the next one. LeNEPA carries this objective to time series, replacing the stop-gradient of next-embedding prediction with the isotropy penalty of LeJEPA. A world model is a transition kernel that can be rolled out. The one-step regression identifies a conditional mean, and a mean is a kernel only in special cases. For a linear-Gaussian Markov latent, the mean transition and the innovation covariance are fixed by the one-step problem, and the open-loop squared error at horizon $K$ equals the trace of the sum of the pushed-forward innovation covariances. That error grows with $K$ after the one-step fit is exact. If the conditional mean is nonlinear, composing it is not the multi-step conditional mean. If the observation is a non-injective function of a Markov state, a memoryless one-step map does not determine future observations, while a short window can. An isotropy penalty is a function of the embedding marginal, so its partial derivative in the transition weights is zero. On a scalar autoregression with coefficient $0.9$, the one-step mean squared error is $0.998$ and the $16$-step open-loop error is $5.10$. On a hidden rotation, an eight-step window reaches $16$-step error $0.056$, while the current scalar alone reaches $0.778$. Raising the isotropy weight from $0.1$ to $10$ leaves eight-step latent error inside $[0.78,0.85]$ on three seeds.
cs.LG / 259 / 2609.36396
LOCO-AdaMP: Built-in LOCO Inference for Adaptive Minipatch Ensembles with Enhanced Prediction
Abstract
As black-box machine learning models become increasingly common, extracting interpretations with uncertainty quantification has become a critical challenge. One popular type of interpretation is leave-one-covariate-out (LOCO) feature importance, while prior LOCO inference methods often require data-splitting or model-refitting. A recent ensemble framework, LOCO-MP, addresses these challenges using minipatches that subsample both observations and features, but massive feature subsampling can hurt prediction in high-dimensional sparse settings. Motivated by this limitation, we consider minipatch ensembles with adaptive feature sampling guided by LOCO importance, and propose LOCO-AdaMP, which enables free LOCO inference for the resulting adaptive minipatch ensemble. We show that LOCO-AdaMP yields substantially improved predictive models while retaining asymptotically valid feature importance inference without data-splitting, despite the complex dependence between the adaptive sampling distribution and the LOCO importance statistics. Our analysis relies on a careful leave-two-out perturbation bound for the iteratively updated sampling probabilities together with the stability of LOCO scores induced by observation subsampling. Empirical results on synthetic and real datasets demonstrate advantages of LOCO-AdaMP over existing methods in predictive performance, inferential power, and stability. Overall, LOCO-AdaMP provides a flexible ensemble framework (agnostic to base models) that delivers both strong predictive performance and asymptotically valid, powerful feature importance inference for regression.
cs.LG / 260 / 2609.36532
Hierarchical Utility Calibration for Structured Multiclass Decisions
Abstract
In multiclass probabilistic prediction, Utility Calibration (UC), which focuses auditing on specified utilities, has recently received attention as a way to guarantee downstream decisions while controlling computational and sample requirements. At the same time, some multiclass problems have meaningful label hierarchies that play important roles in medicine and image classification, yet how UC evaluates utility within a hierarchy remains insufficiently understood. We show that the difference between realized utility and predicted mean utility admits an exact decomposition into a sum of contributions from the internal nodes of the label tree. This decomposition shows that positive and negative contributions from different nodes can cancel, and that even when UC is small, the utility errors remaining in parts of the hierarchy need not be small. To address this problem, we propose Hierarchical Utility Calibration (HUC), which evaluates each node contribution before summation while retaining the same target utility, subgroup, and predicted-utility interval. We further provide finite-sample evaluation over all predicted-utility intervals and propose HUC-Boost, which updates only violated internal nodes, with theoretical guarantees for both.
cs.LG / 261 / 2609.37935
Post-Anomaly Detection Inference for Deep SVDD
Abstract
Deep Support Vector Data Description (Deep SVDD) has become a prominent framework for unsupervised anomaly detection by learning latent representations that compactly characterize normal data around a center. Despite its empirical success, anomaly decisions produced by Deep SVDD are typically made solely based on anomaly scores without rigorous statistical guarantees, thereby limiting their reliability in safety-critical and high-stakes applications where false positives must be strictly controlled. In this paper, we propose PADI (Post-Anomaly Detection Inference), a novel framework that equips a trained and frozen Deep SVDD detector with statistically valid inference by leveraging the Selective Inference framework. Specifically, PADI performs inference conditional on the event that a test instance is identified as anomalous by Deep SVDD, thereby enabling rigorous statistical assessment of anomaly decisions. Based on this formulation, we derive valid selective p-values that quantify the statistical significance of the detected anomaly. Using these p-values, we theoretically establish control of the false positive rate (FPR) at a user-specified significance level $α$ (e.g., $α=0.05$). Furthermore, we extend the proposed framework to Deep Semi-Supervised Anomaly Detection (Deep SAD), providing a principled approach for statistically reliable inference in semi-supervised anomaly detection settings. Extensive experiments on both synthetic and real-world benchmark datasets robustly support the theoretical findings. The results demonstrate that PADI consistently achieves proper FPR control while attaining superior true positive rates compared with existing approaches.
cs.LG / 262 / 2609.37944
Identifiability Guarantees for Drivers and Dynamics of Delayed Physical Systems
Abstract
A wide range of methods have been proposed, including physics-informed neural networks, which are powerful but do not guarantee identifiability of the dynamics, symbolic regression, which requires a set of precomputed operations, and causal discovery, which is more principled but usually relies on strong assumptions that physical systems may violate. In this work, we develop a theory-grounded method and prove that under a set of permissive assumptions, the structural drivers and drift of stochastic delayed differential equations are identifiable. Our method outperforms others on a benchmark for driver identifiability, and on a second benchmark to evaluate physical consistency of the learned dynamics.
cs.LG / 263 / 2609.38058
Latent Inference-Time Guidance of Time Series Foundation Models
Abstract
Time Series Foundation Models (TSFMs) currently provide state-of-the-art results in forecasting tasks. They are available out-of-the-box and rely on in-context learning to make their predictions, which makes the quality of their performance highly sensitive to the user-selected lookback, covariates, horizon and training data distributions. In practise, the quality of the forecasts are variable but complementary, which highlights the need for a principled ensembling approach, rather than selecting the best context. This paper introduces Latent Inference-Time Guidance for TSFMs, which adaptively combines a pool of TSFM forecasts through a time-dependent latent space with independent components. The framework comes equipped with identifiability and reconstruction guarantees, whilst maintaining the off-the-shelf aspect of foundation models. We provide experiments on datasets at various frequencies and from multiple domains: these show that the approach is competitive with traditional ensembling approaches.
cs.LG / 264 / 2609.38112
ReCIRC: Rectified Conformal Risk Control
Abstract
Many applications of black-box predictive models require controlling task-relevant error rates, such as missed lesion pixels in segmentation or missed labels in multilabel classification. Conformal risk control (CRC; Angelopoulos et al., arXiv:2208.02814) gives distribution-free guarantees for such losses, but it calibrates a single threshold shared by all inputs. Because conditional risk varies with the input, this marginal guarantee often overprotects easy cases and underprotects hard ones. We propose ReCIRC (Rectified Conformal Risk Control), which inverts each input's estimated local risk curve to reparameterize the calibrated threshold as a risk budget $a$ representing a common target conditional risk, and then applies CRC unchanged to the resulting family. ReCIRC retains CRC's finite-sample marginal guarantee regardless of the accuracy of the estimated curves, while accurate curves yield approximate conditional risk control and, under additional conditions, asymptotically exact conditional risk control; they also support a risk-calibration diagnostic. Across three synthetic and five real-data settings spanning segmentation, multilabel and multiclass classification, and regression, ReCIRC attained the lowest average worst-group risk and mean positive group excess in every setting, while maintaining marginal risk close to the target, whereas changes in prediction size were application-dependent.
神经与进化计算 (cs.NE)
4
cs.NE / 1 / 2609.36347
Massively Parallel Reinforcement Learning with a Chaotic Reconfigurable Clockless Chip
Abstract
Hardware accelerators based on physical dynamical systems offer an attractive route toward energy-efficient reinforcement learning applications. However, their scalability is challenging because it requires many statistically independent entropy sources. Here, we introduce a quasi-analog decision-making architecture based on asynchronous Boolean networks (or lattices) implemented on a clockless reconfigurable chip. Each node in the network consists of a single logic element that acts as an autonomous entropy source. This architecture gives rise to distributed Boolean chaos, in which a spatially coupled network generates parallel streams of chaotic Boolean transitions with very low statistical dependence between nodes. We experimentally demonstrate parallel decision-making on a 1024-armed bandit problem, which is beyond the scale of previous hardware implementations, while significantly improving power-law scaling performance. Separately, we scale the proposed entropy source to 5120 parallel channels, yielding an aggregate sample generation rate of 2.14 TS/s. Our solution is implemented on a commercial reconfigurable CMOS chip and offers high integration density and ease of programmability. Our results pave the way for using distributed Boolean chaos as a valuable hardware substrate for large-scale reinforcement learning and for the development of fully integrated, high-throughput decision-making accelerators.
cs.NE / 2 / 2609.36773
NeuroDyn-EEG: An Interpretable Pre-trained Model for EEG Based on Neural Dynamics
Abstract
Clinical scalp electroencephalography (EEG) offers a noninvasive window into neural dynamics of neuropsychiatric disorders. However, discriminative deep models often lack anatomically indexed physiological interpretability. We propose NeuroDyn-EEG, a pretraining framework integrating generative priors from neural dynamics. It couples an extended Jansen-Rit neural mass model, leadfield-based source projection, and simulation-based parameter inversion. Trained on synthetic parameter-EEG pairs within physiological ranges, NeuroDyn-EEG estimates 11 regional parameter families across 90 AAL regions plus one global parameter from standard 19-channel EEG, using only ~2.43M trainable parameters. We evaluate the framework across three levels. First, controlled simulations demonstrate robust parameter recovery under diverse noise conditions, while real resting-state EEG evaluations confirm spectral and phase consistency in an inverse-forward closed loop. Second, on four clinical benchmarks (AD65, PD31, Figshare MDD, and TUAB), NeuroDyn-EEG achieves competitive classification performance, securing the highest BACC, AUROC, and AUCPR on PD31 and MDD, and highest BACC on AD65. Third, post hoc regional analyses reveal disease-specific alterations: local synaptic connectivity C_1 involves the most altered regions in AD65, whereas the firing threshold theta ranks first in MDD, offering testable mechanistic hypotheses. Overall, NeuroDyn-EEG maps scalp EEG to anatomically indexed dynamical parameters, bridging representation learning and mechanistic neurophysiology. Code: https://github.com/Gnosis-Neurodynamics/NeuroDyn-EEG.
cs.NE / 3 / 2609.37047
Multi-Depth Temporal Fusion for Feedforward, Locally Trained Spiking Neural Networks
Abstract
We propose a new spiking neural network (SNN) design to process static images and event streams using time-to-first-spike (TTFS) latencies. Our key research question is which architectural choices best accommodate local and online learning in multi-layer convolutional SNNs. This question is addressed via an original framework combining residual-like connections with multi-depth feature aggregation and consensus. The full SNN pipeline features an early-vision front end, to convert raw visual data into sparse spike latencies, a four-layer convolutional backbone trained layerwise with unsupervised spike-timing-dependent plasticity (STDP), a deterministic Multi-Depth Temporal Fusion (MDTF) and a final classifier trained with reward-modulated spike-timing-dependent plasticity (R-STDP). Rather than replacing early features in deeper layers, the proposed MDTF preserves early temporal evidence, adding sparse residual events from intermediate layers, and incorporating deeper features only when they agree in time with earlier representations. The resulting architecture is experimentally validated across MNIST, Fashion-MNIST, CIFAR-10, and N-MNIST, delivering strong classification performance under a fully local learning regime. Selective multi-depth fusion significantly outperforms traditional STDP/R-STDP baselines on higher-variability visual tasks (achieving +18.2 pp on Fashion-MNIST and +29.2 pp on CIFAR-10). Furthermore, activity-budget analyses show that the network retains high accuracy even when removing a large fraction of late or weak spike events, confirming its high data efficiency and reduced event-processing requirements. The codebase is publicly available at github.com/aidinattar/multi-depth- temporal-fusion-snn.
cs.NE / 4 / 2609.37193
Adaptive Rotation for iSOMA: Geometry, Benchmarking, and Noise Robustness in Variational Quantum Objectives
Abstract
We study whether the coordinate dependence of the improved Self-Organizing Migrating Algorithm (iSOMA) can be reduced while retaining its inexpensive leader-directed migration mechanism. We introduce iSOMA-AR, which learns a basis from successful migration displacements and selectively applies the standard perturbation mask in that basis. On the complete noiseless BBOB suite, iSOMA- AR significantly outperformed baseline iSOMA across matched conditions, with the largest gains on geometrically difficult landscapes. A targeted ablation shows that the learned orientation is beneficial on a rotated ill-conditioned landscape and that moderate changes of the gate threshold and rotation cap preserve the qualitative result. On CEC 2011 Real World Optimization Problems, iSOMA-AR outperformed iL-SHADE on most problems, although its advantage over baseline iSOMA was not statistically significant. A canonical-jSO rerun is reported as a post-hoc sensitivity check alongside the original jSO-derived comparator. On frustrated-spin variational quantum objectives, adaptive rotation improved most transverse-field conditions, while gains on the diagonal and anisotropic models were absent or selective. Under strong effective sampling noise, the SOMA variants were the most robust population-based methods in the comparison, but iSOMA-AR was not significantly better than baseline iSOMA. Repairing all-zero PRT masks greatly reduced repeated-point evaluations without changing endpoint quality significantly, making this implementation detail unlikely to explain the noise result. Overall, adaptive rotation is most useful on coordinate-sensitive deterministic problems, while the observed noise robustness appears to arise mainly from the underlying SOMA migration mechanism.
计算语言学 (cs.CL)
80
cs.CL / 1 / 2609.36059
Mnemon: Raw Records, Fast Judgments, Slow Thoughts
Abstract
Long-term memory lets an LLM assistant use a history it can no longer reread, and most memory systems build it by rewriting conversations into facts, graphs or typed memories at write time. We argue that the work of memory divides, as thinking does, into two systems. Most of it is fast System 1 work: many small, independent yes/no judgments about records, such as whether a record is needed or no longer current, which a decision model makes by the dozen in a third of a second. Only a little is slow System 2 work: writing a few search queries, naming what the reply needs and composing the answer, which an LLM does well but slowly. We present Mnemon, a memory agent built on this division. It keeps conversations as raw, dated records; an LLM (System 2) plans searches over them, a decision model, Jev (System 1), judges what the searches return, and rules with explicit budgets turn the judgments into a small View for an unchanged answering model. A background pass consolidates each record once into topic timelines, value histories and standing instructions linked to the records, so that questions about a whole conversation reach evidence their own searches miss. Because nothing is decided about a record when it is written, the same agent can read any store that returns dated records. With gpt-4.1-mini answering, as in a public re-evaluation of 14 systems, Mnemon scores 91.7% on LoCoMo, the highest among them, and 83.8% on LongMemEval-S, from under 4k tokens of context per question, with the lowest effective cost index on LoCoMo. With a reasoning model answering, it reaches 92.2% on LoCoMo and 94.4% on LongMemEval-S, the latter on par with the best published results. From 100K to 10M tokens of history on BEAM, its cost per question grows by a factor of 1.11. On the same records, Jev separates gold evidence better than two LLMs and is 3-11 times faster.
cs.CL / 2 / 2609.36086
PADMÉ: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators
Abstract
Language models are frequently employed to evaluate other language models. An LM evaluator scoring agentic behaviors across multiple criteria is valuable, provided that its decisions align with human judgment. We call the problem of evaluating this alignment Meta-Evaluation. Tackling it directly is difficult: collecting human data is expensive, absolute scoring is hard to align, and using an LM meta-evaluator recurses the question of trustworthiness. We adopt a reformulation of meta-evaluation as a preference judgment problem: rather than comparing human and LM evaluator scores of a trajectory, we ask whether their implied preferences align. Building on this, we introduce PADMÉ, a data synthesis method that generates reliable criterion-based meta-evaluation data for agentic settings. PADMÉ uses only small language models, requires no human involvement during evaluations, and operates under a low computational budget. We build a prototype of PADMÉ and synthesize a dataset of 1,000 samples across four agentic domains and three evaluation criteria. Human validation on a 150-sample subset demonstrates that PADMÉ improves agreement with human judgment from 73% to 85% over a naive baseline. Meta-evaluating 25 common models with our dataset demonstrates the correlations between evaluation performance and scoring granularity, leniency, and model size, among other factors.
cs.CL / 3 / 2609.36131
A Character-Level Neural Approach to Sinhala Sandhi Splitting
Abstract
Sinhala Sandhi splitting recovers the constituent words or morphemes hidden inside a phonologically merged surface form. The task is important for Sinhala NLP because Sandhi obscures lexical boundaries, but no prior published work has established a neural benchmark for Sinhala Sandhi splitting. We present a character-level sequence-to-sequence study based on SandhiLex, using native Sinhala Unicode input and evaluating recurrent encoder-decoder models for affixational and more complex lexicalized, derivational, and etymological Sandhi. The central challenge is the hard subset lexicalized, derivational, and etymological Sandhi, where our best model, a bidirectional LSTM encoder with a unidirectional LSTM decoder, reaches only 68.40\% exact-match accuracy (82.08\% character-level accuracy), well below the 94.00\% achieved on the more regular affixational subset. Ablations show that bidirectional encoding is the largest contributor to performance, while native Sinhala script improves exact match accuracy over romanized input. Qualitative analysis indicates that many errors are near misses involving boundary adjacent characters or plausible but incorrect phonological substitutions. These results establish an empirical baseline for Sinhala Sandhi splitting and identify data scale, Sandhi type conditioning, and attention-based decoding as the main directions for future work.
cs.CL / 4 / 2609.36138
When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMs
Abstract
Before invoking external tools, an agentic LLM must select among a K-way action space: executing a call, seeking clarification, answering directly, or declining. While internal activation steering can alter these pre-execution decisions, conventional aggregate metrics obscure where altered states land and what collateral damage they inflict. We present SAKIKO, an auditing framework that formalizes representation repair via directional error discovery, router-conditioned intervention, destination-resolved verification, and prospectively frozen statistical licensing. Across seven LLMs on When2Call and MetaTool, channel-keyed interventions induce direction-specific net gains in five models; across three sealed evaluations, none of 59 budget-matched random directions matches calibrated target gain. Crucially, destination auditing shows that behavioral movement does not equal repair: an intervention achieving +55 net gain corrupts over half of the baseline-correct decisions it touches, and promising point estimates on Qwen3-4B and Gemma-2-9B are formally declined due to finite-sample uncertainty. SAKIKO establishes the necessity of outcome-resolved adjudication before claiming internal repair. Code: https://github.com/ruizheliUOA/mechanistic-tool-use-llm.
cs.CL / 5 / 2609.36194
Concept Direction Reliability Across Languages with Different Tokenizer Fertility
Abstract
Extracted sentiment directions can vary across samples even when downstream sentiment classification remains accurate. To evaluate direction reproducibility, we measure split-half agreement in English, Hausa, and Yoruba representations across four language models using both native and translated texts. We identify layers selected for agreement using ten topics and evaluate direction agreement across separate groups of fifteen topics. Using the final token, split-half agreement ranges from 0.737 to 0.870 for English, 0.589 to 0.762 for Hausa, and 0.101 to 0.399 for Yoruba, maintaining this language rank order across all 77 complete model comparisons. Classifiers trained on these same layers consistently predict sentiment above chance, demonstrating that predictive accuracy does not imply directional consistency. Furthermore, averaging token representations yields less consistent agreement, and high agreement can partially reflect sentence length. Ultimately, our findings highlight the need to measure vector direction reproducibility independently of classification performance, though they do not establish that tokenizer fertility which is the average number of tokens per whitespace separated word causes cross-lingual differences.
cs.CL / 6 / 2609.36201
SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety
Abstract
Computer-use agents (CUAs), while capable of completing computer tasks in everyday and professional workflows, can cause unintended harm even under benign instructions and environments. However, detecting such harm remains challenging. First, it requires careful, task-specific reasoning: verifiers guided only by general safety criteria often overlook many important but subtle harmful behaviors. Second, it requires active investigation: past trajectory screenshots show what the agent did but not always what actually changed in the environment, so LLM-as-a-judge verifiers that rely on screenshots alone may be unable to determine the actual consequences of actions. To address these challenges, we introduce SCOUT, a two-stage agentic safety verifier that synergizes reasoning-intensive rubric generation with tool-intensive evidence gathering. First, our SCOUT rubric generator extensively reasons over the task and the agent's trajectory to determine what successful and safe execution should entail, generating task-specific completion and safety rubrics. Then, our SCOUT probing agent follows these rubrics to interact with the post-execution environment and collect grounded evidence for final safety and completion judgments. We evaluate our framework on two computer-use safety benchmarks. On AutoElicit-Bench, SCOUT achieves 75.4 unsafe F1 and 74.5 completion F1, outperforming LLM-as-a-judge verifiers and naive tool-use verifiers. SCOUT leads on OS-Blind with 76.4% unsafe detection accuracy. Test-time reflection reduces final unsafe execution rates from 30.2% to 17.2% on AutoElicit-Bench. Ablations and analysis show that tool-free rubric generation in SCOUT elicits substantially more reasoning and is crucial for safety detection across verifier backbones, especially non-frontier ones. A preliminary extension to coding tasks shows that SCOUT can support safety verification beyond computer-use.
cs.CL / 7 / 2609.36205
Geometric Representations of African Languages: A Regional Semantic Hub and Cultural Steering
Abstract
We study how Gemma 4 31B represents African languages and responds to cultural steering. The first study compares nine African languages and three controls using probes, contrast directions, and measures of representation similarity. Transfer from English varies across languages and layers. Directions representing an Africa versus West contrast are more aligned among the African languages than between these languages and the controls at several layers. The comparison across language families passes the reported Holm threshold at five of twelve layers, although dependence between language pairs limits the statistical interpretation. Within Nigeria, Yoruba and Igbo are more aligned than the average of their pairs with Hausa at eleven of twelve layers. The second study uses separate English data to construct directions for Nigeria, Ghana, Kenya, and South Africa. Under union scoring at the selected strengths, estimated differences in attribution rates from random directions range from 0.63 to 0.81. Most outputs pass the automated structural coherence screen. Comparisons with Aya Expanse 32B show that results depend on the representation measure. Together, the studies document regional and family patterns in the sampled representations and country steering in English.
cs.CL / 8 / 2609.36246
Learning from Teacher Continuations at Student States
Abstract
We present OLIVE (OnLine InterVEntion). At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressively, and the student is updated using cross-entropy computed on the teacher-generated tokens. Each design choice targets a corresponding limitation of existing distillation methods: (1) sequential covariate shift in offline supervised fine-tuning (SFT) on fixed teacher trajectories, (2) fragmented supervision under prefix failure in token-level on-policy distillation (OPD), and (3) the need for access to teacher token probabilities in distribution-matching distillation. OLIVE achieves higher reasoning performance than OPD (with a top-16 KL approximation) at comparable GPU-hour cost. Our asynchronous implementation further reduces OLIVE's total training time by 23.8\%. We evaluate OLIVE on both hard reasoning tasks and agentic tasks which reflects modern post-training scenarios, and it consistently outperforms existing distillation methods under the same training budget. By regenerating prefixes from the evolving student, OLIVE continues improving after offline distillation plateaus while better preserving the general capabilities and plasticity of the student. Using only text from GPT-5.4-mini, continuously training with OLIVE outperforms offline SFT from the same teacher by 13\% on ScienceWorld. These results support OLIVE as an effective and efficient approach to online language-model distillation.
cs.CL / 9 / 2609.36303
HeurEvo: Agentic Evolution of Hybrid Solver-Augmented Heuristics for Time-Critical Mathematical Optimization
Abstract
Recent advances in agentic heuristic design use AI agents and execution feedback to automate algorithm discovery for challenging optimization problems. In many practical settings, high-quality solutions must be obtained under strict runtime constraints, motivating hybrid approaches that combine problem-specific heuristics with powerful mathematical programming solvers. However, existing approaches typically improve heuristic components within predefined procedures or tune solver configurations in isolation. This limits holistic adaptation of where to allocate computation, how to leverage solvers, and how to refine the overall algorithmic structure. To address these limitations, we propose HeurEvo, an automated plan--code--component co-evolution framework that jointly evolves the high-level algorithmic structures, their implementations, and a shared pool of reusable components. A planner determines which algorithmic components to use, how to combine them, and how to allocate runtime across stages, a coder realizes the resulting plan as executable code, while a component evolver updates the shared component pool. Within an island-based evolutionary framework, plans and implementations co-evolve with feedback from an interpreter agent that analyzes execution results and identifies opportunities for improvement. Across diverse combinatorial optimization benchmarks and challenging MIPLIB instances, HeurEvo finds high-quality solutions within tight runtime budgets, often matching or surpassing state-of-the-art optimization solvers given hours or days of computation. On several nonlinear geometry problems such as hexagon packing, it also improves upon the best previously reported results. These results highlight the value of jointly searching over algorithmic structure and implementation for agentic heuristic design.
cs.CL / 10 / 2609.36344
DeepRewind: Predicting and Repairing Premature Commitments in Deep Research Agents
Abstract
Deep-research agents conduct long-horizon investigations through iterative search, evidence evaluation, belief revision, and synthesis. However, they may commit to claims before sufficient evidence is available, causing later reasoning to reinforce an incorrect interpretation. We introduce DeepRewind, an additive control layer for reversible deep research that represents the agent's evolving epistemic state as a typed graph of sources, evidence, claims, hypotheses, assumptions, commitments, plans, and drafts. Before accepting an intermediate conclusion, a prompt-based world model predicts its impact and estimates reversibility based on hypothesis narrowing, information loss, recovery cost, and contradiction-trigger coverage. A binary controller blocks risky commitments, while a consistency monitor performs dependency-aware rollback when later evidence invalidates them. Across DRBench and LiveDRBench, DeepRewind improves insight recall by 3.6 percentage points and reduces premature commitments by 59.1% relative to Open Deep Research.
cs.CL / 11 / 2609.36399
Calibrated to Whom? Persona and Language Effects on Cultural Values in JEV
Abstract
Decision-only language models return a probability for every answer option instead of generating text, which makes them attractive as survey respondents and as judges. We audit the cultural values of one such model, TypeSafe's JEV, with the Values Survey Module 2013. We asked it the 24 items as 12 matched Saudi and 12 matched American personas and without a persona, in English and Arabic, under eight ways of formulating the request (288,000 answers). JEV's answers were highly repeatable (ICC 0.997), and without a persona they resembled those of its own American personas. When the persona was Saudi rather than American, the answers moved in the direction of the human Saudi-US difference, reproducing 87% of its size in English but 62% in Arabic, with long-term orientation reversed. A language cross shows that the smaller difference in Arabic comes from the language of the items, not from the language of the persona description. Age shifted the profiles about as much as nationality, gender shifted them more for Saudi than for American personas, and JEV was less confident in Arabic and for Saudi personas. These patterns held in every request design, although the model never generates text.
cs.CL / 12 / 2609.36414
Eternal Sunshine of the Spotless Mind: Systematically Erasing LLM's Memories
Abstract
We consider persistent LLMs that accumulate memories of their interactions with a user over time. Such LLMs maintain memories using external storage, which they can query to overcome the limitations of a fixed context window. Such systems have numerous practical applications, as they can draw on all past interactions when responding to user queries. In this paper, we ask whether LLMs can forget information shared with them upon a user's request. We find that current LLMs fail to delete such information---even when they claim to have forgotten it and even when operating with a limited context. To this end, we consider a new direction of study: Deletion of LLM Memories. We show that naively removing messages that match a user's deletion request is insufficient, since conversations naturally introduce message dependencies that cause information to persist. To correctly handle deletion requests, we propose the DeLLM framework. It dynamically constructs relevant context for each LLM query and maintains a provenance graph of messages to determine which ones must be removed during deletion. Our experiments show that DeLLM achieves a high deletion rate while maintaining utility.
cs.CL / 13 / 2609.36435
MemFold: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization
Abstract
An assistant that serves the same user over a long horizon has to answer from what that user has revealed: which preferences still hold, which were revised, and which constraints apply now. Retaining that information is not the same as acting on it, and the two are usually optimized as if they were. Keeping the information as text makes the reader's input grow with the retained history, while compressing it into a fixed number of latent vectors bounds the interface but is typically trained to reconstruct text or imitate reference answers, both of which are scored on sequences the reader never produced. We present MemFold, which optimizes a fixed-budget soft memory by the behavior it supports. A query-conditioned textual memory is compressed into K continuous vectors that form the reader's memory interface, and the reader is then trained on its own rollouts under two complementary signals: group-relative rewards for task outcomes, and confidence-gated on-policy distillation in which a frozen textual-memory teacher re-scores the student's sampled tokens under the textual memory. The teacher is never sampled from, so supervision stays on the student's current distribution and adds no autoregressive decoding; at inference it is removed entirely. Across three Qwen backbones, MemFold attains the highest accuracy we measure on PersonaMem-32K and PersonaMem-128K, with margins that widen at the longer history length, and transfers to PrefEval and LongMemEval without target-domain training. Ablations attribute most of the task gain to the reward term and a smaller additional gain to the teacher signal, and memory interventions show that the reader depends on the instance-specific content of its soft memory.
cs.CL / 14 / 2609.36457
Memory Consolidation Flattens the Temporal Shape of User Facts
Abstract
Long-term memory systems turn conversations into short stored notes. A note can keep a user fact while losing evidence about whether the fact still holds. For example, "I am driving a Peugeot" can become "The user drives a Peugeot," which drops the cue that the activity is ongoing. We call this aspectual flattening and measure it with LAPSE, a benchmark of matched user statements that differ only in temporal form. We find that memory writers flatten aspect selectively. Three writer models flattened the progressive statement but kept its simple-present match in 244 of 381 pairs, never the reverse. The asymmetry holds in all 11 model configurations tested and in the installed pipelines mem0, Graphiti, and Letta. The lost cue matters to later readers. In exploratory tests, changing only the stored verb shifted all three readers' estimates that a fact still holds. When readers could ask the user before acting, two of three acted without asking more often on flattened notes. Our planned memory-use task could not detect this, because readers there acted on almost every stored fact, even expired ones. Memory writing can thus remove evidence that later models use to decide whether to act.
cs.CL / 15 / 2609.36475
Similar Choices, Different Attention: Cross-Modal Associations in Humans and Vision-Language Models
Abstract
Cross-modal associations are systematic pairings of features across modalities, such as the association of 'bouba' with round shapes and 'kiki' with sharp shapes. Prior work has compared humans and vision-language models (VLMs) on such associations, but often using different stimuli or tasks between humans and models. Here, we ask whether VLMs align with humans not only in choices, but also in where they look when making those choices. We study both VLMs and humans (N = 53), presenting them with the same stimuli, a pseudo-word and two images, and record participants' choices and eye movements, which we release. We find choice alignment in a few larger VLMs, but their saliency matches human gaze less closely than a center-bias baseline, a fixed Gaussian at the center of each image. Fine-tuning small VLMs on human choices brings their choice alignment to the level of a human majority-vote reference on unseen words and images, yet their attention still matches human gaze less closely than this baseline. Training model attention on human gaze raises attention-gaze correlation without improving choice alignment, and a single average gaze map per image position raises it by a similar amount. Matching human choices, or even human gaze patterns, is therefore not sufficient evidence of human-aligned cross-modal processing.
cs.CL / 16 / 2609.36515
Large-scale factor analysis shows machine intelligence is only partially interpretable
Abstract
A common assumption in language model development is that cognitive abilities are organized around a general, domain-free intelligence factor, like fluid intelligence in humans. This assumption is rarely tested directly, and prior attempts have done so only at a much smaller scale. We take a latent variable approach to intelligence in language models, similar to how psychometricians study psychological constructs. Performance in every specific problem set is influenced by a domain-specific and a domain-agnostic latent factor. Using factor analysis as a dimension-reduction technique, we analyzed 13,251 published evaluation scores covering 1,618 language models across 456 different text-only benchmarks. Due to the super-sparse nature of the dataset, we triangulate our analysis across different data densifiers and imputation methods. A robust pattern across different modes of bias is that 1. A general intelligence factor accounts for 70.8% of variance in model performance at our most generous estimate, and far less than that in most of our solutions, 2. Content-similar benchmarks do not necessarily cluster together, and 3. The $g$ factor is not dominated by any common theme, and there is a lack of evidence that it is well-proxied by standard "intelligence" benchmarks. Our findings go against current endeavors of defining, identifying, and targeting general intelligence as a tangible construct in language model development. This leaves the strategy of targeting a single conceptual ability without support, since the first-order abilities it would have to reach are often partially idiosyncratic and not identifiable in practice.
cs.CL / 17 / 2609.36534
Retrieval Sensitivity to Identity Signals in Queries
Abstract
Dense retrievers decide which documents reach users and the language models that use them, yet they are typically evaluated with neutral queries. We ask whether the identity signals that real users express in their queries---political ideology and dialect---bias what a retriever returns. We design evaluations in two domains, political news and consumer-health questions, each pairing a controlled synthetic set that varies only the identity signal with naturalistic queries. Across five dense retrievers and a sparse baseline, every retriever (i) retrieves articles that align with the query's own political lean and (ii) performs worse for questions written in African American Language (AAL) than in White Mainstream English (WME). Two analyses tie these gaps to queries' identity signals beyond surface vocabulary: partialling out an aggregate lexical-asymmetry score leaves the synthetic gaps largely intact, and linear probes recover lean and dialect from the retrievers' query embeddings beyond token-level features. Left unaddressed, such retrieval biases risk contributing to polarization and reinforcing the health disparities already faced by AAL speakers. Code is available at https://github.com/Andrewtcr/bias-ret.
cs.CL / 18 / 2609.36544
DraftTrace: A Multi-View Analytics Environment for AI-Integrated Writing
Abstract
Generative AI has changed how students produce writing assignments. The final artifact is no longer sufficient to understand the process through which it was produced. We introduce DraftTrace, a writing environment that jointly captures three complementary views of writing: the final product, the writing process and interactions with an integrated AI-assistant. DraftTrace reconstructs how a document develops over time and organizes these signals into submission, longitudinal, and class-level analytics for instructors. We deployed DraftTrace in a graduate NLP course with 81 students and compared their sessions with LLM-generated responses entered by automated tools and with copy-typed responses. While product measures distinguish differences in text formulation, process measures distinguish differences in how text is entered. Considering both views together helps characterize cases such as copy-typing. Interaction traces show that students use the assistant differently across stages of writing: to clarify the question at an early stage and to verify answers at a later stage. A preliminary instructor survey highlights the importance of multi-view writing analytics and their interpretability.
cs.CL / 19 / 2609.36550
Grounded Revision vs. Prior Injection: Probing Retrieval-Augmented Patent Claim Amendment
Abstract
Retrieval-augmented generation is widely used in professional writing, yet whether retrieval grounds revision or merely injects templates is rarely tested where "correct" has a definable meaning. Patent claim amendment supplies that signal: the examiner names the attacked limitation and cites prior art, providing per-case ground truth. We release three artifacts: (i) a corpus of 7,385 USPTO prosecution cases with XML-aligned pre/post claims, rejection, and cited prior art; (ii) a seven-probe battery comparing random and structural-match retrieval as two policies under a fixed prompt scaffold; (iii) a deterministic five-channel metric (C1-C3 and C5 in main, C4 supplementary) requiring no LLM evaluation. Across 9,600 pre-registered calls on four frontier LLMs (Claude Sonnet 4, Claude Haiku 4.5, GPT-5.4, GPT-4o-mini), no tested model exhibits detectable classical prior-injection behavior; retrieval effects are small and direction-inconsistent between random and structural retrieval, and the null is unchanged under a dense (semantic) retriever, across retrieval depths k in {1,3,5,10}, and under a paraphrase-sensitive grounding metric. Revision locality reveals a model-specific difference that the template channel misses. The four-cell taxonomy, which we treat as exploratory, leaves the prior-injector cell unoccupied.
cs.CL / 20 / 2609.36617
Generating Edit-Inducing Questions for AI Research Manuscripts
Abstract
We study the ability of LLMs to generate edit-inducing questions whose answer will improve a paper draft. On a dataset of paired submission and camera-ready papers from ICLR and NeurIPS, we compare the helpfulness of questions from GPT models with or without full paper context to that of human reviewers. GPT produces more edit-inducing questions and its questions are associated with more extensive edits and cover a broader range of edited content compared to questions from reviewers. However, a much smaller percentage of the GPT questions are edit-inducing. Our analyses confirm that automated questions can be beneficial to authors and highlight an example task where proper attending to long context deteriorates reasoning model ability to produce helpful output.
cs.CL / 21 / 2609.36675
Gödel Forest: Balancing Search Depth and Breadth for Data-Centric Recursive Self-Improvement
Abstract
Recursive self-improvement (RSI) aims to achieve compounding gains by having models improve themselves. While most existing RSI systems optimize external agent harnesses or prompts around a frozen base model, data-centric RSI directly updates the model's own parameters by training on agent-generated data. However, because validating data strategies requires expensive model training, existing methods face a fundamental dilemma: a single agent gets trapped in narrow directions and lacks exploration breadth, while naive parallel search or heavy trace sharing sacrifices long-horizon search depth. To address this challenge, we introduce G"odel Forest, a multi-agent framework that organizes recursive self-improvement as an ensemble of co-evolving search trees. In G"odel Forest, each agent autonomously grows a persistent tree, deepening, branching, or pruning data strategies based on model feedback to secure depth, while parallel trees explore distinct regions of the data space to expand breadth. Crucially, rather than leaving trees isolated or flooding them with heavy execution logs, a dynamically co-evolving memory connects the forest: agents continuously distill their successes and failures into compact procedural lessons anchored to a global leaderboard. Through this forest ecosystem, a dead-end in one tree instantly warns the whole forest against unpromising paths, while an empirical breakthrough quickly seeds new exploration branches in neighboring trees. Evaluated on RSIBench-Data across six diverse domains, G"odel Forest outperforms the single-agent baseline by an average of 10.70% while reducing wall-clock time on five tasks. Ablations confirm that co-evolving shared memory yields a +7.00% gain over independent parallel search, demonstrating that collective distillation is key to scalable self-improvement. The code is available at https://github.com/evolvent-ai/Godel-Forest.
cs.CL / 22 / 2609.36684
ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context
Abstract
Embodied agents now take on ever longer tasks. For long tasks, knowing only whether a task finally succeeds or fails says little; the steps along the way matter. Progress Reward Models (PRMs) score how far a task has come at every step, and serve as dense rewards, verifiers and monitors. Yet in long tasks the current frame alone often cannot tell how far the task has come, because progress depends on what happened before. We call this problem context-dependent progress estimation. Existing benchmarks on progress estimation mostly focus on short tasks whose progress can be read from the current observation, and whether PRMs can estimate progress when context is needed remains underexplored. We therefore build ContextProgress-Bench, with 24 manipulation tasks for 120 episodes. The benchmark covers three settings: (i) State Recall, where information needed for progress appeared earlier but is not in the current frame; (ii) Sequence Tracking, where steps follow a fixed order, so progress requires knowing which steps are done and which comes next; and (iii) Recurrence Disambiguation, where look-alike frames sit at very different progress. We then run a paired diagnosis: each PRM keeps the same input format in both runs, and in one run its instruction integrates the right context. Even PRMs that read the entire history get lost in estimating progress, yet with the right context the same five models cut their progress error by 77-82%. Embodied PRMs are thus not incapable of progress estimation, but lost without the right context. We therefore propose ProgressCompass, an autonomous agentic loop that reorients an existing PRM and uses current general-purpose VLMs to supply the context the PRM needs. Wrapped in the loop, the same frozen PRM cuts its progress error by 63% and raises its rank agreement by 76%. With such a compass, PRMs estimate progress far better on longer, more complex tasks.
cs.CL / 23 / 2609.36691
Video2Skill: From Streaming Experience to Reusable Embodied Skills
Abstract
Manipulation behaviors vary widely across objects and scenes, but they share a small set of reusable skills, and planning with these skills helps embodied agents generalize to new tasks. Yet an agent can only plan with skills it knows. Recovering skills from observed experience, the inverse of planning, builds this knowledge over time and yields skill data for training future agents. Vision-Language Models (VLMs) describe individual manipulation events well, but can they organize a stream of events into reusable skills? We formulate this problem as Streaming Embodied Skill Discovery (SESD): a model watches videos in sequence and maintains a persistent skill library that shapes its later decisions. To systematically measure this ability, we introduce Video2Skill, a benchmark that covers robot tabletop manipulation and human kitchen activity and tests three core capabilities: (i) locating manipulation events in time, (ii) grouping events of the same transformation, and (iii) deciding when to reuse an existing skill or create a new one. Across 19 open-source VLMs, many models group events at near-chance level, and scale does not consistently help. Their errors depend on how perception and library updates are coupled: joint models merge distinct transformations into one skill, while models that update the library from text descriptions duplicate recurring ones. Supervised fine-tuning, including our counterfactual library-state rebalancing (CLaRe), improves grouping but exposes a deeper bottleneck: trained models consolidate familiar skills yet rarely expand the library. Their libraries stall below half the reference size, and transformations unseen in training are located in time but almost never given a new skill. Recognizing when existing skills are insufficient thus emerges as the central challenge.
cs.CL / 24 / 2609.36707
LAURA: Knowledge Distillation for Interpretable Ambiguous Clause Identification in Legal Contracts
Abstract
Legal contracts contain ambiguities that expose enterprises to financial and legal risks. Some ambiguities allow flexible interpretation without triggering disputes, while others lead to significant legal conflicts. This makes identification alone insufficient, and interpretable rationale analysis essential. We propose LAURA, a post-training framework for interpretable ambiguous clause identification. LAURA leverages knowledge distillation with an IRAC-Unlearning prompting technique to transfer knowledge from a teacher LLM to an open-weight student model (<=1B parameters), which is then trained using a joint objective combining classification and rationale generation losses. The framework supports both legal and non-legal stakeholders in making informed decisions about which ambiguities require further attention. Extensive experiments across 7 baselines and 7 open-weight models demonstrate that LAURA with Flan-T5 (250M) delivers state-of-the-art interpretability over all interpretable baselines while matching the identification performance of the best-performing opaque baseline.
cs.CL / 25 / 2609.36804
VAA-CSEC: Vote-guided Advantage Allocation for Chinese Semantic Error Correction
Abstract
Chinese Semantic Error Correction (CSEC) targets semantic errors in Chinese text, which are typically more subtle and complex than spelling and grammatical errors but remain relatively underexplored. Existing LLM-based approaches face two recurring obstacles in this task: over-correction, and unclear interaction between Chain-of-Thought (CoT) reasoning and self-consistency decoding, such that the benefits brought by CoT cannot be reliably transferred to final corrections. We propose Vote-guided Advantage Allocation for CSEC (VAA-CSEC), a multi-stage framework that combines CoT distillation, Supervised Fine-Tuning (SFT), Reinforcement Learning (RL) and self-consistency decoding. During RL, we design a task-specific reward function that directly aligned with the minimal-editing principle of CSEC. We further introduce Group-Level Relative Policy Optimization (GLPO), which reallocates GRPO advantages according to the margin between individual rollout rewards and the vote-aggregated group reward, aligning the RL training objective with the self-consistency objective used at inference time. Experiments on CSED-C and NaSGEC-Exam show that VAA-CSEC outperforms all LLM-based baselines on CSED-C with an F0.5 of 47.72%, achieves the highest recall of 42.15% among all methods, and establishes a new state of the art of 41.55% F0.5 on NaSGEC-Exam.
cs.CL / 26 / 2609.36850
Rethinking Multimodal Fake News Detection in the Generative AI Era
Abstract
Generative content is increasingly entering the production and dissemination of news, transforming fake news from manually fabricated or simply manipulated material into complex forms in which native and generated content jointly participate. Existing multimodal fake news detection research primarily focuses on veracity assessment and rarely characterizes how generativity differences affect the reliability of evidence. In contrast, AIGC detection primarily determines whether content is generated or modified by generative models, but it does not by itself establish whether the underlying news event is true. To bridge the separation between these tasks in data and evaluation, we construct Weibo26, a multimodal fake news detection dataset for generative-content scenarios. On this basis, we propose the Generativity-Aware Hierarchical Reasoning (GAHR) framework, which combines global judgment with local correction so that generativity information participates in news-veracity reasoning. Experiments on multiple existing fake news detection benchmarks and Weibo26 show that GAHR achieves competitive veracity-detection performance while effectively identifying generative content.
cs.CL / 27 / 2609.36893
Momentum-Coupled Rubric Adaptation for Detailed Image Captioning
Abstract
Detailed image captioning requires accurate and comprehensive descriptions of fine-grained visual content, yet caption quality spans factual accuracy, information coverage, and clarity. Compared with conventional methods that rely mainly on high-quality supervision or holistic rewards, rubric-based reinforcement learning decomposes these requirements into explicit criteria and provides targeted, structured feedback. However, existing methods often use separate models for caption generation, rubric construction, and judging, which may lead to inconsistent interpretations across roles. Some dynamic rubric methods alternate updates between the caption policy and rubric generator while keeping the judge fixed, but staged optimization may still leave rubric construction and judging out of step with policy optimization. We propose MoCo Rubric, a two-stage framework that coordinates these roles. First, role-conditioned, shared-parameter multi-task supervised fine-tuning equips a single vision--language model to serve as the Caption Policy, Rubric Generator, and Rubric Judge. Then, the Generator constructs rubrics online from captions sampled by the current Policy, reference captions, and image evidence. The Judge provides rubric-based rewards, and only the Policy receives GRPO updates. As Policy updates change the candidates being evaluated, we use an exponential moving average of the Policy parameters to update one momentum model shared by the Generator and Judge. This gradual transfer lets both rubric roles track Policy updates without separate RL optimization while smoothing parameter changes that could disrupt their rubric capabilities under direct synchronization. Across five captioning benchmarks, MoCo Rubric achieves an average pairwise win rate of 72.83\%, the best mean rank in blind ranking, and the highest average score in caption-based question answering.
cs.CL / 28 / 2609.36902
RAEGNet: Relation-Aware Evidence Graph Network for Harm-Aware Multimodal Fake News Detection
Abstract
Existing multimodal fake news detection methods often introduce external information to assist detection. However, most of them rely on entity-level retrieval and are therefore prone to introducing event-irrelevant noise. Meanwhile, existing methods mainly focus on improving overall performance and do not account for differences in the degree of harm posed by different instances of fake news. To address these limitations, we design an Event-Level Evidence Retrieval Framework (ELERF) and propose a Relation-Aware Evidence Graph Network (RAEGNet). ELERF retrieves external evidence based on the complete event semantics of a news item. RAEGNet constructs a directed graph that incorporates news-evidence stance relations and evidence-evidence interaction relations, and introduces a conditional-harm branch to jointly model authenticity and potential harm. Experimental results demonstrate that RAEGNet outperforms multiple baseline methods across all evaluated metrics on Weibo-21, Fakeddit, and our self-constructed SSS dataset.
cs.CL / 29 / 2609.36903
MultiTalk: Scaling Full-Duplex Speech Models to Long, Multi-Party, Bilingual Conversation
Abstract
End-to-end full-duplex speech models have brought open-source machine conversation closer to human-like interaction, yet existing systems remain limited in two intertwined dimensions: long-context robustness and multi-party interaction. Real-world scenarios such as meetings, group lessons, and social-robot reception require a single model to track, contextualize, and respond to multiple speakers over extended durations. Progress is constrained by both data and evaluation: open multi-party speech corpora remain small and are not designed for codec-frame-level full-duplex modeling, while existing long-audio benchmarks focus on passive listening and speech-to-speech benchmarks are mostly short and dyadic. We extend the Moshi paradigm jointly along the long-horizon and multi-party axes in English and Chinese. First, we release 57.6k hours of synthetic training data ($\href{https://huggingface.co/datasets/MultiTalk/MultiTalkPT}{MultiTalkPT}$ and $\href{https://huggingface.co/datasets/MultiTalk/MultiTalkFT}{MultiTalkFT}$) for long-form, multi-party, English-Chinese full-duplex dialogue, with controllable length, participant count, turn-taking, overlap, backchannels, interruptions, addressee shifts, and long-range coreference. Second, we introduce $\href{https://huggingface.co/datasets/MultiTalk/MultiTalkBench}{MultiTalkBench}$, built from real human recordings, for evaluating long-form, multi-party, bilingual full-duplex dialogue. Conversations average 32.6 minutes and include probes for long-range entity tracking, topic coherence, and addressee selection. Third, we train a bilingual Moshi-style model that sustains coherent multi-party English-Chinese conversations over extended durations and substantially outperforms open-source baselines including Moshi, MiniCPM-o-4.5, and Qwen3-Omni-30B-A3B-Instruct on MultiTalkBench.
cs.CL / 30 / 2609.36913
BaLEEN: Biasing with Latent Encoded Entities for Context-Aware ASR
Abstract
Transcribing domain-specific entities and rare proper nouns remains a major challenge in automatic speech recognition (ASR). In this paper, we propose BaLEEN (Biasing with Latent Encoded Entities), a lightweight, hypernetwork-based framework for dynamic contextual adaptation without fine-tuning the underlying ASR model. BaLEEN encodes variable-length contextual keywords using a pretrained language model, compresses them into a fixed sequence of latent vectors via a Perceiver bottleneck, and injects context-dependent bias vectors directly into the intermediate encoder representations of the ASR model. Because both the language model and the backbone ASR model remain entirely frozen during training, BaLEEN operates as a plug-and-play adapter that incurs zero computational overhead at inference time when context biases are precomputed. We evaluate our method on a CTC-based ASR model using a Wikipedia-derived corpus with annotated named entities and synthetic speech. Experimental results demonstrate that BaLEEN reduces keyword miss rate by 8.7% on the test set relative to the unbiased baseline while simultaneously improving overall word error rate by 21% and character error rate by 28%.
cs.CL / 31 / 2609.36914
Can Language Models Learn to Forecast Stock Prices
Abstract
Post-training has been shown to significantly improve language models' performance on tasks with verifiable outcomes, including mathematical reasoning, software engineering, and computer use. However, whether the same approach can improve forecasting in financial markets is much less clear. Compared with tasks with verifiable outcomes, not only are realized returns noisy, but even what constitutes a relevant information set for making effective predictions is not obvious a priori: the model must decide which observations to gather and then commit to a numerical judgment before the outcome is known. We study this question in a chronological stock-price sandbox, where a language model gathers price, volume, relative-performance, and market-context evidence and predicts a future return. We post-train Qwen3-4B with supervised fine-tuning (SFT) on tool-use demonstrations, then proximal policy optimization (PPO) with a terminal reward given by the forecast score against the realized return. The resulting AURA-4B more than doubles the starting direction--magnitude score, from 20.94 to 43.31, and is comparable to frontier language models on this benchmark. Conditional magnitude agreement rises from 33.3 to 66.2, while directional accuracy changes from 62.9 to 65.4. SFT expands tool use, and PPO further increases the share of ranking and market-context queries. These results show that post-training can substantially improve financial forecasting performance, together with changes in how the model investigates the market, on this outcome-selected benchmark.
cs.CL / 32 / 2609.36920
Benchmarking Automatic Speech Recognition Tools for Iberian Languages
Abstract
Comprehensive evaluations of automatic speech recognition (ASR) for Iberian languages remain limited, and low-resource languages, biases, and efficiency trade-offs are underexplored. We benchmark eleven systems, ten open-weight models and one commercial API, across five Iberian languages (Basque, Catalan, Galician, Portuguese, Spanish), with German and Turkish as controls. Evaluation uses an 85-hour dataset covering read speech, broadcast media, and audiobooks, assessing accuracy and efficiency via word error rate (WER) and real-time factors (RTF/RTFx). Results show no single model dominates: accuracy, efficiency, and language coverage present clear trade-offs. Low-resource languages, especially Basque, degrade significantly, highlighting the role of training coverage. We observe consistent sex disparities across most systems, highlighting fairness challenges in multilingual ASR. Overall, the benchmark provides practical guidance for real-world model selection.
cs.CL / 33 / 2609.36931
Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation
Abstract
Reproducibility is essential for scientific research, yet prior work shows that LLM outputs vary with hardware and batching. We identify an overlooked factor: the hidden injection of the current date into system prompts, which users cannot control and which changes every day. Across 9 recent LLMs and 6 datasets spanning multiple-choice QA (MCQA), math reasoning, code generation, and machine translation, performance varies solely with the current date, with deltas of up to 6% on MCQA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation. Model rankings also shift, affecting leaderboards. This date effect exceeds other sources of non-determinism, such as batch size and numerical precision. Standard prompting techniques -- chain-of-thought and few-shot prompting -- do not reduce the sensitivity; chain-of-thought even amplifies it. Our findings underscore the need for careful evaluation protocols to ensure reproducibility and fair comparisons in LLM research.
cs.CL / 34 / 2609.36965
Chinese-Jev: Bringing System One Model to Chinese-Language Tasks
Abstract
System One models such as Jev offer an efficient alternative to generative language models for tasks that require decisions rather than open-ended responses. However, existing Jev models exhibit limited Chinese-language decision accuracy, restricting their utility in both general and specialized settings. In this paper, we introduce Chinese-Jev, a System One model that addresses this gap through a unified data processing and training pipeline. Our data processing protocol converts heterogeneous Chinese-language annotations into probability targets over candidate options, enabling a shared training formulation across domains and question formats. To enable efficient inference, Chinese-Jev adopts a lightweight encoder-only backbone for text encoding and learns to score candidate answers through decision-oriented training. To address the misalignment between the pre-training distribution and downstream Chinese-language scenarios, we first train the model on a general-purpose corpus of 10 million examples, then fine-tune it separately for the medical, legal, and financial domains. To evaluate decision accuracy and calibration in both general and domain-specific Chinese-language settings, we introduce Chinese-Jev Bench (CJ-Bench). After first-stage pre-training, Chinese-Jev exceeds the accuracy of the closed-source Jev model by 1.24% on general-domain tasks while achieving a 20.3x speedup. Subsequent domain-specific fine-tuning yields a 4.0% accuracy improvement over Jev in medicine and achieves 92% of Jev's average accuracy across specialized domains, with a 17x speedup and an average latency of only 15 ms per example. We further demonstrate on-device deployment of an INT8-quantized model on mobile devices, achieving an inference latency of approximately 1.0 second per decision. The project is available at https://gulucaptain.github.io/Chinese-Jev/.
cs.CL / 35 / 2609.36974
Repetition, Not Length: Isolating the Counting Failure in Neural Text-to-Speech
Abstract
Text-to-speech models loop, truncate and lose count on text that repeats a phrase many times. We show that repetition itself is what breaks them, not the length that comes with it. Every repeated sentence in our test set is paired with a control of matched sentence and word count in which no word ever repeats back-to-back. Six models from three architectures render the controls almost perfectly and fail the repeated twins: 94.3% against 18.2% exactly right at k >= 6. The gap survives greedy decoding, repetition-penalty sweeps, four independent speech recognisers and 420 analysis specifications without once reversing sign; a held-out fourth architecture lands within a point of its predicted gap, and one of two non-autoregressive baselines shows the same failure. Varying the period of the text shows the failure grows smoothly with periodicity, half of it surviving when no word is adjacent to itself.
cs.CL / 36 / 2609.36987
CypherTurn: A Multi-Turn Benchmark for Conversational Text-to-Cypher Evaluation and the Autonomy Divergence
Abstract
Graph databases are increasingly queried through natural language, yet every existing benchmark evaluates isolated single-turn queries rather than the multi-turn sessions through which analysts actually work. We introduce CypherTurn, the first benchmark for conversational Text-to-Cypher evaluation, comprising 721 sessions and 5,927 turns across 7 knowledge graphs and 13 conversational phenomena. We evaluate 15 models under a guided oracle protocol and a fully autonomous agentic protocol, yielding four findings. First, the best model reaches only 64.7% execution accuracy, and session-level correctness remains below 5%. Second, despite strong overall rank correlation, frontier models exhibit a consequential reordering of the top of the leaderboard under autonomous operation, a phenomenon we term the Autonomy Divergence, which reveals error-management as a partially independent capability from raw generation skill. Third, scaling action budgets from x3 to x10 fails to close the autonomy gap, as the strongest frontier models self-limit to approximately two actions per turn regardless of available budget. Fourth, single-turn Cypher fine-tuning degrades multi-turn instruction following, while architecture-appropriate specialization outperforms several frontier models. These results establish CypherTurn as an open challenge for conversational graph database reasoning. Code and data are available at https://github.com/BarryQ/CypherTurn.
cs.CL / 37 / 2609.37017
LatCom: Cross-Agent Latent Compression for Efficient Multi-Agent Collaboration
Abstract
LLM-based multi-agent systems (MAS) increasingly use latent collaboration to avoid the information loss and repeated encoding-decoding overhead of natural-language communication. However, directly forwarding all sender latents makes the receiver-side context scale with both the number of agents and the reasoning length, increasing computation, memory usage, and collaboration latency. A natural solution is latent compression. But we find that cross-agent redundancy remains unresolved in existing latent compression approaches, which typically compress each sender independently and then concatenate the results. We propose LatCom, a cross-agent latent compression framework for efficient multi-agent latent collaboration. LatCom maps multiple sender latents into a fixed number of receiver-readable and task-relevant slots. Rather than reconstructing all sender hidden states, it optimizes the compressed latents for receiver-side task utility. LatCom trains the compressor in two stages: single-sender readability learning first establishes a latent interface interpretable by the frozen receiver, and multi-sender fusion learning then trains the compressor to fuse complementary evidence and remove redundancy across agents. Experiments on multiple benchmarks with Qwen3-4B show that LatCom achieves an average 2.46x inference speed-up over LatentMAS and reduces output token usage by 70.3% while maintaining comparable average accuracy.
cs.CL / 38 / 2609.37040
Selecting The Most Informative Tokens in Natural Language Autoencoders
Abstract
Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across $4.7$ million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, without requiring a model forward pass for position selection. On three of four datasets, explaining just $5\%$ of positions retains nearly all of the success rate from explaining every position, where success means obtaining an explanation about the threat. The benefit varies with the audit task. We also show that pretrained verbalizers recover words that models have learned to conceal through fine-tuning, without additional verbalizer training. These results identify where auditors can concentrate explanation generation and show that useful explanations can extend beyond the model a verbalizer was trained to describe.
cs.CL / 39 / 2609.37082
Traverse: Learning When to Remember, Reset, and Redirect for Long-Horizon Web Search
Abstract
Long-horizon information-seeking agents often accumulate noisy or misleading context, causing early mistakes to persist and making recovery increasingly difficult. We introduce an autonomous search harness in which the agent manages its own search process through three states: Rubric, Answer, and Verify. The agent first defines criteria for a valid answer, searches under these criteria, and then independently verifies the result before deciding whether to terminate or continue searching. It is further equipped with a Seal Memory tool that enables active context management. Training this behavior with reinforcement learning, however, can induce Seal Collapse, resulting in unstable training and preventing the agent from reliably learning when and how to use its memory tools. We solve this with a simple strategy that trains only the final segment after context management. Our 35B model achieves 72.83 on BrowseComp, outperforming comparable open-source systems, and consistently improves over the base model across BrowseComp-ZH, xbench, DeepSearchQA, WideSearch, financial investigation, and product search. Ablations show that autonomous compression outperforms automatic compaction and validate our RL design.
cs.CL / 40 / 2609.37104
What Does Post-Training Change in Multilingual Reasoning?
Abstract
Open-source reasoning models provide unequal access to reasoning capability across languages. When a model can solve a problem but cannot deliver a complete solution in the user's language, language becomes an access barrier rather than merely a source of performance variation. We audit Qwen3 checkpoints on competition-mathematics tasks in eleven languages. Across the ten non-English languages, only 15.4-17.9% of problems receive a correct, terminating solution with visible reasoning in the requested language in any of 16 samples, compared with 92.9% in English. To identify the source of this disparity, we evaluate thirteen endpoints from one model family, spanning released checkpoints, multilingual supervised fine-tuning (SFT) at two scales, controlled SFT ablations, and three reinforcement-learning (RL) reward formulations. We jointly track correctness, language adherence, termination, and delivery efficiency. The dominant bottleneck shifts across post-training stages. Released models often reason in English. Multilingual SFT restores target-language reasoning, but accuracy declines across multilingual, English-only, and single-language SFT runs, showing that this cost is not specific to multilingual mixing; non-English reasoning traces additionally become prone to non-terminating loops. RL restores termination in both arms at no cost in accuracy, but only the arm whose reward includes a language term delivers: rewarding correctness alone returns the model to English. Together, these stages establish a constructive post-training path from English-pivoted capability to multilingual reasoning that is reliably delivered.
cs.CL / 41 / 2609.37121
Cross-Linguistic Effects in Bilingual Phoneme BabyLMs
Abstract
Cross-linguistic effects are a central topic in bilingual first-language acquisition. Artificial learners can help investigate L1-L2 interactions by enabling controlled comparisons across language combinations and learning conditions. Recent work explores this direction by training bilingual language models under developmentally plausible constraints. However, human and model learners still diverge in fundamental ways, with one major difference being input modality: children learn primarily from spoken input, whereas language models are typically trained on orthographic text. To reduce this gap, researchers have trained models on phonemic representations of speech. In this work, we combine these research directions to train bilingual BabyLMs with phonemic input. We keep English fixed as the L2 and vary the L1 across German, Swedish, Persian, and Basque, selected to represent contrasting combinations of syntactic and phoneme-inventory distance from English. Our results show stronger L1-related variation in grammatical learning trajectories under phonemic than orthographic input, while early lexical differences align with phoneme-inventory similarity.
cs.CL / 42 / 2609.37223
CredWise: A Controlled Agentic Decision-Intelligence Framework for Explainable and Auditable Credit-Risk Assessment
Abstract
Credit-risk prediction is important in banking, but a prediction alone does not explain why an applicant is risky or how it should be combined with other evidence. This paper presents CredWise, a decision-support framework that integrates credit-risk prediction, probability calibration, explainable artificial intelligence, policy retrieval, SQL analytics, and controlled agent-based workflows. An XGBoost model is trained on Lending Club data (1,345,310 loans, 18 features) using a temporal split: 2007--2016 for training, 2017 for validation, and 2018 for testing. On the 2018 test set, the calibrated model achieved a ROC-AUC of 0.7109, PR-AUC of 0.2993, F1-score of 0.3714, and accuracy of 65.44\%. Calibration reduced the Brier score from 0.2157 to 0.1273 and the expected calibration error from 0.2862 to 0.0585. SHAP explanations were temporally stable, with a Spearman correlation of 0.9959 between 2017 and 2018 feature rankings. On 28 labeled queries covering nine policy sections, FAISS achieved the best Hit@1 (0.929) and MRR (0.964), while all three retrieval methods reached Hit@5 = 1.0. Agent routing achieved 95.6\% accuracy (43 of 45 cases), and the SQL benchmark scored 1.0 on exact-match, execution-success, and result-match across six cases. These results show that CredWise can combine predictions, explanations, policy evidence, and structured analytics in one controlled workflow. It is an academic research prototype, and final decisions remain with a human reviewer.
cs.CL / 43 / 2609.37226
Follow the Entities: A Corpus Map for Agentic Search
Abstract
Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach this by iteratively searching the full corpus rather than reading only a fixed set of top-ranked documents. However, when the corpus is exposed only as a flat collection of files, a relevant document gives no indication of how it relates to others, so the agent must rediscover these relationships for every query, often missing complementary evidence while simultaneously consuming substantial additional tokens. To address this, we introduce CorpusMap, a navigation layer that organizes the corpus around its recurring entities, which are identifiable from the documents themselves and can link a single document to many others across sources. Specifically, CorpusMap represents each recurring entity as an Entity Page that aggregates information about it and links to every document that refers to it, forming a graph between entities and documents that the agent can traverse to gather otherwise disconnected evidence. Moreover, since CorpusMap is constructed offline by resolving mentions of the same entity across documents, its links are shared across queries rather than rediscovered repeatedly at inference time. Using 7 different models with 3 benchmark datasets, we show that CorpusMap improves both evidence discovery and answer quality over raw-corpus agentic search while using fewer tokens on average, and further outperforms 4 alternative navigation layers, suggesting that entities serve as effective anchors for navigating large document collections.
cs.CL / 44 / 2609.37361
SemOPT: Fixing Semantic Errors in LLM-based Optimization Modeling via Reward-Guided Search
Abstract
Operations research supports decision-making in domains such as energy, economics, and healthcare. Solving operations research problems typically begins with optimization modeling, which translates a natural-language problem description into executable solver code. LLMs offer a promising way to automate this process, but they remain prone to errors. In practice, these errors can be divided into two categories: syntactic errors refer to solver code that fails to run successfully or is judged infeasible by the solver; semantic errors refer to solver code that successfully returns an objective value but violates the intent of the original problem. Since semantic errors do not trigger runtime failures, they are difficult to detect and rectify. To address this problem, we introduce SemOPT, a semantic-guided framework for correcting LLM-based optimization models. SemOPT combines a semantic reward model that distinguishes faithful math models from plausible but incorrect ones with an adaptive correction system that applies hierarchical reward-guided search over the modeling space. Experiments on seven optimization modeling benchmarks show that SemOPT establishes a new state of the art and achieves an average 7.6% accuracy improvement over the strongest baseline on complex datasets.
cs.CL / 45 / 2609.37371
Compiling Learning Problems into Adaptation Programs for Language Models
Abstract
Model adaptation is typically governed by a fixed recipe, even though different update programs can produce substantially different behavioral outcomes. We introduce adaptation compilation, which reframes where, how, and to what extent a model should adapt as a joint prediction and decision problem. Rather than searching over candidate programs anew for each learning episode, a compiler learns from prior adaptations to predict a vector-valued counterfactual response surface over candidate programs---their expected effects on acquisition, transfer, boundedness, and preservation---and selects a program before adaptation begins. Because this predicted geometry captures multiple behavioral consequences rather than a single winner or scalar score, it can be reused under different downstream priorities without retraining. Across five learning types, preferred programs vary meaningfully across episodes, and this variation is predictable from pre-adaptation information. On Llama-3.1-8B, compiler-selected programs approach exhaustive search while outperforming global and objective-specific defaults. Replication on Gemma-2-9B preserves program heterogeneity and selection headroom, but shows that exploiting this headroom requires accounting for uncertainty when departing from strong defaults. Together, these results show that adaptation search can be amortized across related learning problems, turning prior adaptation experience into a basis for deciding how future learning should occur.
cs.CL / 46 / 2609.37443
Learning to Retrieve Missing Evidence for Long-Term Memory QA
Abstract
Long-term memory enables language models to use past interactions in future conversations. However, evidence needed to answer a question may be scattered across distant turns, while the question itself omits clues needed to locate it. Retrieved facts can reveal these clues, motivating retrieval decisions conditioned on evidence already found. We introduce MERA (Missing-Evidence Retrieval Augmentation), which separates globally searchable memory from a question-specific evidence state. Verified evidence guides subsequent retrieval without restricting access to the global memory. We train a lightweight planner through reinforcement learning, rewarding queries that recover previously missing evidence. MERA achieves strong answer accuracy across Qwen3-30B and GPT-4o-mini backbones. With Qwen3-30B for evidence processing and answer generation, the trained 0.6B planner achieves 77.40% accuracy on LoCoMo and 71.29% on LongMemEval-S, exceeding a 30B planner without retrieval-grounded training by 4.10% and 3.96%, respectively. On LoCoMo, later retrieval rounds increase cumulative evidence recall from 55.5% to 80.5%.
cs.CL / 47 / 2609.37510
From Dissonance to Orchestration: Teacher Intervention in On-Policy Distillation
Abstract
On-policy distillation (OPD) trains a student on its own reasoning trajectories using feedback from a stronger teacher. Teacher interventions can improve these trajectories, but also change the distribution on which the student learns. Our controlled studies show that rollout quality alone is an incomplete criterion for allocating teacher guidance. Deeper intervention yields diminishing gains in rollout accuracy while increasing off-policy load. In a training probe with a restricted rollout horizon, peak student accuracy and performance retention favor different intervention strengths. The preferred intervention depth and placement also vary across benchmarks. These findings motivate MAESTRO, which uses local policy disagreement to jointly adapt when the teacher takes over and how long it generates. Its {policy disagreement score} combines teacher-weighted candidate coverage with local distribution similarity and is aggregated within reasoning paragraphs. Across eight mathematical reasoning benchmarks, MAESTRO achieves the highest macro-average accuracy among the compared methods for both 0.6B and 1.7B Qwen3 students, with the 1.7B student leading on every benchmark. MAESTRO also reduces average training response length by 67.3\% relative to standard OPD. The code is available at https://github.com/yhao-wang/MAESTRO.
cs.CL / 48 / 2609.37543
RunyaNER: Auxiliary Language Selection for Runyankore NER
Abstract
Cross-lingual zero-shot transfer and multilingual fine-tuning are promising approaches for NLP tasks such as Named Entity Recognition (NER) in low-resource languages, but in the absence of target language benchmarks, it is unclear which auxiliary language selection strategy leads to the best transfer. We introduce RunyaNER, the first publicly available NER benchmark for the East African language Runyankore, and use it to investigate the choice of which languages to use for transfer. Created with a semi-automated pipeline and fully manually verified, RunyaNER contains over 237k annotated words across 30k sentences. We benchmark pretrained models on RunyaNER, establishing that our dataset is of sufficient quality and size to produce effective Runyankore NER models. We then use RunyaNER to investigate auxiliary language selection in cross-lingual zero-shot and multilingual fine-tuning settings. Our experiments show that while transfer performance is highly sensitive to auxiliary language selection, embedding-based measures computed from labelled training spans correlate more strongly with downstream transfer performance than traditional linguistic features based on metadata or typology. By releasing RunyaNER and providing a systematic analysis of auxiliary language selection strategies, this work contributes both a new benchmark resource and practical insights for multilingual transfer in low-resource settings.
cs.CL / 49 / 2609.37564
Orthogonal Yet Coupled: Decoupling Geometric Components for Model Merging
Abstract
Merging pretrained models has emerged as an effective approach for consolidating diverse capabilities into a single unified model. However, prevailing merging methods typically treat each task vector as an indivisible merging unit, overlooking the heterogeneous geometric changes encoded within it. This treatment can induce cross-component coupling: when merging decisions are derived from statistics of the complete task vector, the geometric characteristics of one component may influence how another is selected, weighted, or combined, potentially degrading the quality of the merged model. To address this issue, we propose DiGA, a Disentangled Geometry-Aware model merging framework. Using the pretrained weights as a shared geometric reference, DiGA orthogonally decomposes each task vector into components corresponding to distinct geometric attributes. Rather than merging the task vectors as a whole, DiGA aggregates corresponding components independently within their respective subspaces and subsequently recombines them into a unified update. This component-wise formulation preserves the geometric identity of each component and prevents the characteristics of one component from interfering with the aggregation of another. Furthermore, DiGA can be incorporated into a broad range of existing model merging methods. Extensive experiments across diverse models, tasks, and merging methods demonstrate that DiGA improves merged-model performance and reduces capability degradation. Our repository is on https://github.com/wzj1718/DiGA.
cs.CL / 50 / 2609.37624
Correct, Don't Delete: Mitigating Emergent Misalignment with Corrective Supervision
Abstract
Fine-tuning a language model on a narrow set of harmful demonstrations, such as bad medical advice, can make it broadly misaligned on unrelated questions, a phenomenon known as emergent misalignment (EM). The usual defense is to find the offending rows and delete them, but a row locator failed our held-out test and deleting rows helps less than expected. We ask a different question: given a fixed set of poisoned rows, is it better to correct them than to remove them? We fine-tune Qwen2.5-14B-Instruct on a mixture of bad medical advice and benign chat data, select a quarter of the poison rows in advance, and either delete them or replace each with a corrected answer to the same prompt, keeping everything else the same. Replacing the rows cuts the EM rate by about a third and improves answers on held-out medical questions, while deleting the same rows has little measurable effect. The advantage is larger when half the poison rows are corrected, and it holds on a second base model and a second misaligned model organism. The content of the replacement appears to matter: paraphrasing the rows while keeping their bad advice shows no clear benefit, and the correct answers distributed with the dataset appear to do about as well as our rewriter's. Realigning an already-poisoned model with further fine-tuning is known to work, but which data does the work has not been compared directly. We find that a short round of training on corrections beats the same amount of training on generic chat data, that corrections on other medical prompts do roughly as well as corrections of the poisoned prompts themselves, and that instructing the correction writer to model a careful, harm-avoiding assistant adds no measurable benefit over plain corrections. In the settings we tested, correcting harmful training data reduces EM more than deleting it.
cs.CL / 51 / 2609.37635
Co-Linguistics: AI-augmented Theory Construction in Linguistics
Abstract
LLMs have been studied in recent linguistics as potential models of humans' linguistic abilities. Here we discuss an entirely different use of AI, namely as a co-scientist, to help construct and assess linguistic theories (we refer to the result as "Co-Linguistics"). Since the 1960s, linguistics has developed theories that are in principle mathematically formalizable, often in the language of formal language theory or model theory. The AI revolution in mathematics will thus have consequences in linguistics-but with an essential twist: proving new theorems is rarely the linguist's goal. Rather, one seeks to find the best set of axioms to derive empirical statements. AI could accelerate research by making existing theories fully explicit, by comparing competing theories, and more ambitiously, by proposing new theories (in machine learning, this relates to "program induction"). It will also help assess theories by accelerating the identification and test of crucial predictions, thanks to unparalleled access to data (in machine learning, this relates to "active learning"). While the cycle from theory evaluation to theory construction may give rise to recursive and possibly autonomous improvement of linguistic theories, humans remain central: linguists provide scientific directions and evaluate theories conceptually, and experimental participants are needed to assess empirical predictions that are outside the reach of LLMs.
cs.CL / 52 / 2609.37647
Evaluating and Benchmarking the System One Model Jev
Abstract
Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed options, a position on a rubric, or the probability that a statement is true, with probabilities the vendor describes as calibrated. Such models target small decisions in information access pipelines, such as routing queries, checking grounding, moderating content, or rating against a rubric. We evaluate Jev (jev-1.13.0) zero-shot on 37 datasets spanning classification, routing, natural language inference, reading comprehension, commonsense reasoning, moderation, legal clause analysis and rubric scoring, with one frozen template per dataset and full evaluation splits: 346,009 requests for under USD 10. For reference, we score Qwen3.8-27B and Gemma-4-E4B on identical requests via their exact next-token probabilities over the options. Jev reaches 95-99% accuracy on IMDB, SST-2, HellaSwag and ARC and 86.7% on Belebele across 122 languages. It beats Qwen on 27 of 37 datasets, with none of Qwen's nine leads outside the bootstrap intervals, and Gemma on all 37. All three models degrade on low-resource languages, fine-grained or noisy labels, and rubric-based quality judgments. Jev's choice probabilities are well calibrated and support selective prediction. Binary probabilities rank well but are poorly placed relative to a fixed 0.5 threshold; thresholds tuned on training data raise micro-F1 on UNFAIR-ToS from 0.50 to 0.75. Jev answers MMLU's calculation-heavy questions more accurately than other MMLU questions (94% vs. 91%), whereas both open models, and all three on C-Eval, find them harder. Rotating the options leaves Jev's accuracy unchanged and withholding the question drops it to near chance, ruling out shallow memorization but not memorized question-answer pairs. We release the code, harness and all raw responses.
cs.CL / 53 / 2609.37661
Corpus-Guided Dual-Path Propagation for Graph Retrieval-Augmented Generation
Abstract
Graph-based retrieval-augmented generation supports multi-hop retrieval by organizing corpus information into graphs. However, existing relation-free graph retrieval methods rely primarily on query-sentence similarity to search for evidence. This can exclude useful bridging evidence with low query similarity and activate incidental entities unrelated to the reasoning chain. In this paper, we propose a simple and effective approach called NexusRAG, which augments the relation-free Tri-Graph with a corpus-level entity neighborhood structure derived from joint entity co-occurrence and semantic similarity. NexusRAG employs this structure to guide two complementary propagation paths: neighborhood-constrained semantic propagation through sentences identifies the query-relevant entity frontier, while direct structural propagation between neighboring entities expands that frontier to structurally related entities. The propagated entity weights also inform neighborhood-aware passage initialization for Personalized PageRank. Experiments on three multi-hop QA benchmarks and a domain-specific subset of GraphRAG-Bench show that NexusRAG consistently outperforms existing approaches. On the GraphRAG-Bench subset, NexusRAG achieves the highest evidence recall in all question categories, exceeding baselines by 4.2-8.1 points. The implementation code is available at https://github.com/Jacob-biu/NexusRAG.
cs.CL / 54 / 2609.37713
Billiger.de Products: A Bilingual Entity Matching Benchmark
Abstract
Existing product matching benchmarks primarily contain English-language product data and are often dominated by a single product category, such as electronics. This paper introduces Billiger.de Products, a bilingual German and English entity matching benchmark covering thirteen consumer product categories, including difficult-to-handle categories such as clothing and furniture. The benchmark data originates from the German price comparison platform billiger.de. Following the design of WDC Products, the benchmark offers multiple variants that differ in the fraction of corner cases, the size of the development set, and the fraction of entities unseen during training. An aligned English translation of every offer keeps all pairs, splits, and labels fixed, while cross-language test sets combine German and English records within individual pairs. We validate the benchmark using six supervised matchers and zero-shot GPT-5.2 on both language versions and the cross-language test sets. The validation shows the difficulty of the benchmark. The comparison of the results on the English version of the benchmark to the results on the German version shows that most matchers score on average higher on the English version. The difference is largest for RoBERTa and HierGAT, while the zero-shot LLM runs are largely insensitive to the language. Comparing the F1 scores achieved by PLM-based matchers on the English version of Billiger.de Products with their performance on existing English-language benchmarks, such as WDC Products and Abt-Buy, shows that Billiger.de Products is more difficult than these benchmarks.
cs.CL / 55 / 2609.37755
Which papyrus HTR is good enough? Character-error-rate tolerance of four papyrological tasks on Greek texts
Abstract
Purpose: Most Greek papyri remain unpublished and undigitised; a handwritten text recognition (HTR) pipeline that transcribes them automatically would let scholars discover documents and literary works that have so far gone unread. Recognition systems for Ancient Greek papyri are in statu nascendi, and how accurate they must be for a given papyrological task has not been examined. To answer this and set a benchmark for Greek papyrus HTR, we test a range of character error rates (CER) against four papyrological tasks, using published editions as ground truth. Methods: From 63,846 current editions of Greek texts in papyri.info, we imitate a letters-only "perfect HTR" output by removing the editorial layer, then degrade it with a seeded algorithm to exact CERs of 1 - 50%, with lost lines and four error-shape variants. On these data we train small models (TF-IDF, fastText, a character CNN, ByT5-small) for document type, dating and documentary-versus-literary classification, and apply eight keyword search methods. We compare models trained on clean text with models retrained at a specific CER level, and evaluate across CERs. Results: Tolerance differs by task. With clean-trained models, documentary-versus-literary classification retains 90% of its metric up to 20% CER; document type up to 7.5%; subtypes and search up to 5%; dating only up to 3%. Retraining on text containing character errors largely eliminates the sharp degradation that otherwise sets in above 15% CER. Models generally tolerate concentrated damage in a long document better than small errors spread across a short text. Conclusion: The study provides a CER target for each of the four tasks and shows that models trained on noisy text make current, imperfect text recognition useful for them.
cs.CL / 56 / 2609.37807
CompOrca: Corpus-Scale Compliance Labelling of Instruction-Tuning Data
Abstract
Studying how fine-tuning shapes refusal and noncompliance behaviour requires identifying training examples that refuse, evade or otherwise fail to fulfil the requested task. But existing annotation covers evaluation sets of a few thousand prompts at most. We present CompOrca, a compliance labelling over the entirety of the 4,233,923-example OpenOrca corpus. Every example was classified as compliant or noncompliant by five independent passes of an open-weight LLM judge (LongCat-2.0, 1.6T parameters), and the corpus is released as unanimous compliance (94.75%), unanimous noncompliance (1.28%), and nonunanimous rows (3.97%) along with the raw vote counts. A single pass flags 2.7-3.2% of the corpus as noncompliant, while only 1.28% is flagged by all five, allowing for filtering the most ambiguous samples. Against 450 human-annotated examples, 150 of them annotated twice (human-human $κ= 0.93$), the unanimous compliance and noncompliance labels are 97.3% and 86.7% precise, the latter a high-precision subset, not a complete enumeration, of noncompliance. Published refusal-detection methods recall only between 0.4% and 94.1% of the noncompliance class. We release the full corpus with its per-row labels and vote counts at https://huggingface.co/datasets/cemiu/CompOrca
cs.CL / 57 / 2609.37818
Thinking in Depth, Speaking Directly: Recurrent Latent Reasoning for Paralinguistically Grounded Spoken Dialogue
Abstract
Empathetic spoken dialogue requires models to use both what is said and how it is said to decide how to respond. Explicit CoT can improve paralinguistic perception and make acoustic cues more explicit in replies, yet does not ensure their effective use in response planning. We call this mismatch the perception-reasoning gap. In addition, CoT may not fully capture acoustic cues in words, and generating it adds inference latency. To address these limitations, we introduce LoopSLM, which builds on looped Transformers for latent reasoning, reusing a decoder block to refine hidden states with acoustic grounding at every pass. Its two-stage training further narrows the perception-reasoning gap by separating learning to reason from learning to respond, enabling direct inference without CoT. On EchoMind, LoopSLM improves paralinguistic understanding, reasoning, and reply quality over Qwen2.5-Omni-7B. Against the CoT-SFT baseline, LoopSLM gains over 20 points in reasoning accuracy while generating 64.5% fewer tokens at half the latency. It also outperforms Qwen3-Omni-Thinking on most empathetic reply metrics with 34x lower latency. Despite training only on dialogue data, LoopSLM improves accuracy on general audio benchmarks.
cs.CL / 58 / 2609.37824
The Geometry of Inference in Transformer Residual Streams
Abstract
Transformer language models build predictions through successive residual updates, but how their representations become specific to an eventual outcome remains unclear. We study this process by comparing intermediate residual states with their own final states and an empirical bank of final states from other contexts. Across six pretrained language models, the own endpoint becomes preferable to the average alternative early, while many individual endpoints remain closer. These competing sets generally shrink with depth, but their membership changes and their surviving endpoints need not become more similar to one another. Directional alignment and endpoint rank can therefore improve while Euclidean distance to the final state changes little. We develop a simple high-dimensional model that separates the roles of norm, alignment, and endpoint geometry, showing how gradual directional changes can produce sharp reductions in competition. We also prove that a straight path toward the own endpoint cannot introduce new competitors under either Euclidean or cosine distance; observed entries thus establish departures from straight-line convergence. Finally, endpoints associated with lower-ranked output tokens tend to lie farther away in cosine distance across all studied models, connecting residual geometry to output organization. Together, these findings characterize increasing geometric specificity during transformer inference and explain why distance, competitor count, and concentration of the surviving endpoints provide distinct views of that process.
cs.CL / 59 / 2609.37837
Can Vision-Language Models Stay Helpful When Facing Implicit Risks? Intent-Privilege OPSD for Efficient Safety-Helpfulness Alignment
Abstract
Vision-Language Models (VLMs) remain vulnerable to cross-modal implicit risks: visual and textual inputs that appear benign in isolation can jointly elicit unsafe responses. Existing safety methods often require large preference datasets, costly multi-rollout training, or additional safeguards at inference time. They may also sacrifice helpfulness by directly refusing requests that could be answered safely. In this paper, we propose Intent-Privilege On-Policy Self-Distillation (OPSD), which leverages evidence-grounded intent as privileged supervision during training to help VLMs recognize implicit risks and provide safe, useful responses instead of blanket refusals. OPSD distills a teacher's intent-conditioned preferences over responses into a student using a single rollout per prompt; the student then responds without intent annotations or an additional safety module. With only 1,447 safety-specific examples - 95% fewer than standard preference datasets - OPSD reduces training time by 5x relative to multi-rollout GRPO-style training and average inference length by 7%. It attains the highest ratio for joint safety-helpfulness success, which measures the proportion of responses that are both safe and helpful, across all five evaluation groups. Remarkably, on pooled SIUO+HoliSafe, this success ratio rises from 43.9% to 53.5%. These results show that training-time intent supervision can improve both safety and helpfulness while substantially reducing data, training, and inference costs.
cs.CL / 60 / 2609.37861
One Threshold Does Not Fit All Languages: Language-Conditional Deferral for Reliable and Efficient Low-Resource Text Classification
Abstract
In the Global South, the lower-income countries of Africa, Asia, and Latin America where most of the world's languages are spoken, a deployed text classifier usually runs on ordinary CPUs, serves many languages with a single model, has few labeled examples in any of them, and relies on people to catch its mistakes. Such a system is only useful if it can promise how often it will be wrong: at most a fixed fraction of the labels it assigns on its own may be incorrect, and everything else must go to a person. Split conformal prediction delivers this promise through a single confidence threshold, normally estimated on validation data pooled across languages. We ask whether the promise reaches every language, and it does not. On MasakhaNEWS (16 African languages) and AfriSenti (12 languages plus two never seen in training), a pooled threshold meets the 90% target on average but covers Somali at 77.5%, Tigrinya at 83.7%, and the two unseen languages at 77.5% and 81.2%. Estimating one threshold per language brings every language to between 89.1% and 91.0% without retraining, and it shows how unequal the cost of the promise is: keeping it means sending 43% of Somali news and over 80% of Amharic and Xitsonga tweets to a person, against under 8% of Nigerian Pidgin news. One or two hundred labels per language are enough and the models train in minutes on one CPU core, so the fix is affordable: calibrate, report, and budget human review one language at a time.
cs.CL / 61 / 2609.37863
It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them
Abstract
Vision-language models (VLMs) are increasingly used in place of human annotators, making it important that substitutability tests reflect the model rather than incidental evaluation conditions. We introduce MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all. The guidelines require the label to be decided from the sentence alone, so no image should change any answer. We expected each image to pull a judge's labels toward the sense it depicts, and neither kind did. Across thirteen VLM judges, an aligned image changed 20.5% of labels and a misleading one 19.4%, close for every judge and both above the 11.6% produced by deleting the ignore-the-image instruction with the image left in place. Yet only 37% of the labels that differ between the two images moved toward the sense shown, and agreement with our human annotators is unchanged whether the image is absent, aligned or misleading. The effect is smaller in the seven judges that pass the alt-test than in the six that never do, but present in all of them: what moves a judge is that an image is there, not which of the two it is, so a substitutability verdict describes a configuration as much as a model.
cs.CL / 62 / 2609.37879
Retrieval Capacity of Self-Attention Under Competition
Abstract
How many tokens from its context does a language model actually use, and what determines that number? We study this question through self-attention. Without retraining, we retain only the tokens with the highest attention weights at each head, layer, and query, keeping their original weights unchanged. By varying the selected set size and measuring the increase in negative log-likelihood (NLL), we estimate the effective attention set size needed to stay within a chosen loss tolerance. Relatively small selected sets can keep NLL close to the full-attention baseline, although the required size varies across models. Attention-based selection substantially outperforms random selection. Selected sets exhibit geometric structure, although geometric separation alone does not establish that model loss is preserved. Extending context while evaluating the same prediction targets increases the required set size, while its fraction of context decreases over the tested range. Experiments with a fixed supporting fact show that additional background pushes its tokens down the attention ranking and reduces their attention mass. Renormalizing the retained weights can substantially reduce the required set size, showing that it also depends on how selected representations are combined. Conditional theoretical models explain how competition and attention-mass retention can produce growing set sizes without more distinct information to retrieve. These results provide a way to measure effective attention set size in language models and investigate its dependence on context, competition, and aggregation.
cs.CL / 63 / 2609.37882
How Many Labels Does a Language Need? Annotation Budgets and Cross-Lingual Pooling for African-Language Text Classification
Abstract
Every text classifier for an African language begins with a budgeting question: how many labelled examples are needed, and can labels from other African languages stand in for them? We answer both questions empirically for 28 language-task pairs, news topic classification in 16 languages (MasakhaNEWS) and tweet sentiment in 12 languages (AfriSenti), using a character n-gram linear model that trains in seconds on two CPU cores with no pretrained weights and no accelerator. Monolingual learning curves at budgets from 25 to several thousand labels show that topic classification reaches 90\% of its full-data macro-F1 with about 400 labels in the median language, while sentiment is still improving at the full training size in 11 of 12 languages and needs thousands of labels. Pooling the full training data of the other languages in the benchmark is worth a great deal at small budgets and nothing at large ones: at 25 target labels it adds 0.20 macro-F1 on average for news (up to 0.43 for Lingala) and 0.08 for sentiment, the gain decays to zero by 800 labels, and at full size pooling hurts in 9 of 16 and 8 of 12 languages. Twenty-five target labels plus pooled data match what 100 to 400 monolingual labels achieve for most news languages. A complete zero-shot transfer matrix shows that transfer without any target labels recovers a median of only 13\% (news) and 4\% (sentiment) of the gap between a majority-class predictor and the in-language model, with the exceptions explained by shared script (Amharic and Tigrinya), shared lexicon (English and Nigerian Pidgin, the Arabic dialects), or a shared label prior rather than by language family. We release code that regenerates every number from the public benchmark files and translate the results into concrete annotation guidance for teams building African-language classifiers without GPUs.
cs.CL / 64 / 2609.37883
Zero-shot Dependency Parsing with Unsupervised Cross-Lingual Bootstrapping
Abstract
Pre-trained language models (PLMs) with encoder-based architectures have shown impressive capabilities in zero-shot cross-lingual transfer for various language understanding tasks. However, applying this technique to dependency parsing remains a significant challenge due to its syntactic nature. To boost model generalizability across linguistic typologies, we propose a cross-lingual unsupervised bootstrapping method to improve syntactic knowledge within the PLM. We show that our method achieves a significant improvement in zero-shot parsing performance in low-resource languages. Analysis of these bootstrapped models uncovers increased robustness in recognizing syntactic structures, evidenced by higher scores in parameter-free tree probing tests.
cs.CL / 65 / 2609.37891
It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs
Abstract
Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--for instance, they contain little explicit reasoning. Thus, many frontier labs have begun to develop their own internal datasets, starting from state-of-the-art models, to augment their pre-training data mix, eg, with reasoning traces to address cold-start problems. While demonstratively effective, none of these datasets are public, and the effect of this so-called synthetic data on knowledge and skill acquisition of language models, including small ones, remains poorly understood. We present SYNTH, the first open-source synthetic corpus derived from 58,698 Wikipedia articles that collapses pre-, mid-, and post-training into a single training stage via structured amplification of curated encyclopedic seeds. We evaluate SYNTH by training a suite of models: a 56M tiny model (Monad), 0.3B-0.6B dense models (Baguettotron), and a 13B / 1B-active MoE. At iso-compute, SYNTH outperforms filtered web data, and our models remain competitive with similarly-sized open-weight baselines. Because SYNTH is back-translated from grounded passages, SYNTH-trained models achieve high factual precision despite 10-140x fewer training tokens, with memorization targeted by the seed corpus. These results show that synthetic datasets, including our SYNTH dataset, are capable of producing competitive generalist models from a fraction of the training data, enabling rapid iteration as the frontier advances. These findings open up possibilities for both generalist models with significantly increased data efficiency, as well as domain-specific models where no instruction or conversational data is available. Finally, we publicly release our SYNTH dataset and the suite of Baguettotron models under a permissive license, thus supporting open-source language model development.
cs.CL / 66 / 2609.37930
Learning What to Remember: Long-horizon Counterfactual Memory Optimization
Abstract
Persistent textual memory allows language models to carry information across long interactions, but learning what to remember is fundamentally a credit-assignment problem. A memory rewrite may only become useful many steps later, while much of the observed utility may be inherited from information already stored before the rewrite. We introduce Memory Gain Policy Optimization (MGPO), which isolates the incremental value of each memory rewrite by crediting it for its marginal contribution to current and future downstream utility. This turns delayed memory utility into a direct learning signal for optimizing what information should persist. We study MGPO on document-level information extraction, where structured supervision makes the effects of individual memory updates directly measurable. MGPO improves extraction while reducing average memory length by nearly 80% relative to the initial memory policy before optimization. The learned memory policy also supports reuse and transfer across domains, downstream models without further training. These results show that effective memory learning depends not only on preserving useful information, but on identifying which memory updates create lasting incremental value.
cs.CL / 67 / 2609.37993
BITEM at the NTCIR-19 R2C2 Task: Predicting Confidence from Agentic RAG Pipeline Signals
Abstract
The BITEM team entered both subtasks of the NTCIR-19 R2C2 task with a single agentic pipeline, in which a model searches, reads and records evidence over a movie corpus while an orchestrator holds the record and rules on what may be submitted. A claim is admitted only once an entailment cascade has checked it against the passage it cites, and an answer is released only once enough checked evidence stands behind it. Each question is run three or four times, every pass retrieving from a corpus stripped of what the earlier passes have already seen. The confidence filed with each answer is computed by the orchestrator from what the run leaves behind and is never asked of the model, which is offered no way to rate itself. The two retrieval runs placed 4th and 5th of 22, pooling the passes was worth 0.0709 nDCG@20, and the gain was largest on the multi-hop and post-processing-heavy questions, where the organisers rank the pooled run top of the field. Sixteen of the 25 answer runs were built on passages these two runs supplied, 12 of them filed by other teams. HMR rewards a system whose confidence is high where it answers right and low where it answers wrong. The pipeline reached an accuracy of 0.9219, 6th of 25, while the confidence filed with those answers gave an HMR of 0.4915, 13th. A few rules crafted over those same recorded signals, with no further model call and no further retrieval, raise that to an accuracy of 0.9375, 5th, and an HMR of 0.6985, 9th. Ranking on HMR alone can reward a system for answering wrongly with low confidence, so we propose accHMR, the accuracy multiplied by HMR, which reports the reward in proportion to the accuracy, and on which the revised rules would have scored 0.6549, 5th. For future work, fitting a model on the numbers the pipeline already produces, rather than writing such rules by hand, would be a real step forward.
cs.CL / 68 / 2609.38021
Auditable Long-Term Memory: A Deterministic Retrieval Chain Measured at 479/475 of 500 on LongMemEval-S
Abstract
We evaluate an auditable long-term memory system on LongMemEval-S. Its retrieval chain uses hybrid candidate retrieval, cross-encoder reranking, coverage-first packet compilation, and deterministic reasoning scaffolds; an LLM is used only as a replaceable final reader. The chain places all gold sessions in the candidate pool for 468/470 answerable questions and produces gold-complete packets for 462/470. With a Claude Opus reader called through an unpinned CLI alias, two 500-question passes score 479/500 and 475/500 under GPT-4o. The 72 answerable knowledge-update rows used a substantively modified scoring prompt whose effect under the official text has not been measured. The pair straddles Chronos High's published 478/500; differences in reader generation, scoring prompt, and possibly data version, plus within-system variance, establish neither superiority nor equivalence. A grok-4.6-high reader on the same packets scores 476/474, while a maximum-reasoning-effort agentic variant regresses to 461/465. The headline passes differ on eight verdict-flip rows. A second judge agrees with the headline judge on 493/500 rows (98.6%) in each pass and scores both passes 472/500; the official judge also flips three verdicts when re-scoring byte-identical pass-1 answers. Negative controls rejected a verifier that repaired three wrong drafts but broke eleven correct drafts. All components were developed on the same 500 questions, with no held-out evaluation or independent human adjudication; retrieval and scaffold method sources and transcript-derived audits are held; and the headline reader received extra operator context, its complete requests were not retained, and MCP tool availability is unresolved. We release materialized packets, scaffolds, reader outputs, judge verdicts, and controls for inspection and re-scoring.
cs.CL / 69 / 2609.38107
Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces
Abstract
Chain-of-thought traces are widely read as records of how models reach their answers, informing debugging, agent auditing, and claims about reasoning. Testing this interpretation is difficult because natural-language thinking traces are rarely mechanically verifiable. We revisit it in iGSM, a synthetic grade-school mathematics benchmark designed to study thinking traces and used to support claims of learned reasoning and planning. Crucially, iGSM exposes the exact quantities and dependencies that a correct solution should use, allowing generated traces to be checked programmatically step by step and enabling us to test whether correct answers are reliably accompanied by valid traces. We first evaluate models trained exclusively on valid, minimal traces. Answer correctness and trace validity nearly coincide in distribution but decouple out of distribution: on the hardest instances, 31.6% of correct answers have invalid traces, over half of which pass all syntactic and arithmetic checks but fail semantic dependency checks. We then intervene on trace supervision. Non-minimal training traces induce non-minimal outputs, while re-asking the same problem with a different query reveals computations inherited from the original query, weakening minimality as evidence of selective planning. Shuffling tokens in 10% of training trace sentences preserves near-clean accuracy even out of distribution despite no trace passing verification. Swapped training traces likewise retain high in-distribution accuracy. We discuss the implications of these findings for chain-of-thought monitoring and interpretation in the context of AI safety.
cs.CL / 70 / 2609.38109
How Local Mixing Encodes Relative Position in Global NoPE Attention
Abstract
The attention operation is naively position invariant. However, positional information is fundamental to natural language, and therefore a variety of explicit position encodings have been developed in transformer-based models, such as rotary position encoding (RoPE). Although explicit position encodings have long been assumed to be required, recent methods that interleave local mixing layers, such as sliding window attention (SWA) and gated linear attention, while not encoding position (NoPE) in global attention layers has recently been shown to be successful at scale. How and why this approach works is not well-understood. In this paper, we develop an explanation of how hybrid models of this sort can implicitly encode position at global NoPE layers. Supported by both theoretical and empirical evidence, our central argument is that SWA and gated linear attention induce a recency bias in the residual stream that propagates to, and is selected by, the global attention logits. Moreover, in contrast to the implicit position encodings found in models with only global NoPE attention, in which positional information arises solely from the causal mask, the recency bias in hybrid models can be maintained across long sequences. In addition to deepening our understanding of how hybrid models encode position, these findings may provide insights for how to encode position in a way that can extrapolate to longer sequence lengths indefinitely.
cs.CL / 71 / 2609.38111
From Routing Signals to Selective Review: Visual regrounding in MoE VLMs
Abstract
Vision-language models (VLMs) may accept false visual premises, answering questions about a target object's color, count, location, or state even when it is absent. We call this reliability-critical behavior a target-absence grounding failure. Existing visual-grounding detectors primarily rely on generated responses, hidden states, or uncertainty measures. We present the first framework to leverage internal routing decisions in Mixture-of-Experts (MoE) VLMs to detect target absence before generation and guide selective correction. We extract target-token routing probabilities from Qwen3-VL-30B-A3B-Instruct and Gemma-4-26B-A4B-it, train a separate L2-regularized linear detector for each model, and use its predictions to selectively invoke a target-aware review prompt. Using routing alone, the Qwen and Gemma detectors achieve ROC-AUCs of 0.9988 and 0.9956 on GQA-Inpaint and retain 0.8095 and 0.7781 on the external OBER dataset, respectively. The resulting routing-gated policy improves end-to-end accuracy on GQA-Inpaint and OBER by +22.25% and +12.17% for Qwen, and by +13.42% and +1.39% for Gemma, without modifying model weights. Further analysis shows that the signal is localized to the target-object token, emerges in early MoE layers, and is distributed across partially substitutable experts. Although cross-dataset threshold shifts require recalibration, false-positive review causes limited harm overall, suggesting that intervention risk can be controlled through joint selection of the detector threshold and review prompt. Overall, we show that routing probabilities alone preserve actionable information about visual perception, allowing computation already produced by an MoE VLM to support low-cost detection and selective visual regrounding.
cs.CL / 72 / 2609.38137
LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning
Abstract
Language-model (LM) harnesses enable LMs to operate effectively over long contexts using additional compute. However, existing long-context evaluations are insufficient for distinguishing modern harnesses, reflected by saturated accuracy across harnesses and largely similar evaluation costs. In this paper, we introduce a benchmark for evaluating both the effectiveness and efficiency of long-context harnesses. Our tasks require diverse retrieval strategies, including lexical search and semantic matching, together with strategic and adaptive reasoning over global and local context. Much of the context is semantically relevant but only a small subset is useful at each step, creating both a challenging search problem and different accuracy--cost tradeoffs across processing strategies. For example, one task requires identifying every person satisfying several conditions using evidence scattered across documents; strategically checking the most selective condition first can narrow the search before verifying the remaining conditions. We evaluate multiple families of frontier language models with four state-of-the-art harnesses. Our benchmarks remain challenging even for strong model--harness combinations: the best reaches 68\% macro-average accuracy across four evaluation suites. More importantly, we find that the same underlying model can exhibit markedly different efficiency under different harnesses. Our results establish efficiency as an important axis for long-context evaluation and provide a testbed for developing harnesses that process context strategically rather than exhaustively.
cs.CL / 73 / 2609.38149
Pretraining Latent Information Feedback Transformers with Teacher Supervision
Abstract
Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for information to flow downward across generation steps is the decoded token. This narrow channel forces models to recompute intermediate results and to discard alternative continuations. In this work, we remove this bottleneck during pretraining, introducing the LIFT (Latent Information Feedback Transformer) architecture and training method which enable LMs to propagate state across generation. We achieve this by turning recurrent-state learning into a teacher-forced prediction problem: each input token is paired with an information-dense state, derived from the next-token distribution of an off-the-shelf pretrained LM. The model, extended with a small number of additional parameters, is then trained to predict both the next token and the next state. As the input states are precomputed, pretraining remains fully parallel across positions. At inference, the model's own predicted states are fed back, with a minor computational overhead that decreases with model size. Experiments with pretrained models ranging from 135M to 1B parameters show that LIFT consistently outperforms standard Transformers and baselines on language modeling, downstream reasoning tasks, and procedural tasks under token-matched budget, while being on par with or ahead of compute-matched Transformers. Moreover, a controlled study on a state-tracking task shows that a tiny LIFT outperforms same-size Transformers trained on 8x more data, even when trained with the states of a Transformer that fails the task. Overall, we show that LMs can learn to exploit deep-to-shallow feedback during pretraining via scalable teacher supervision.
cs.CL / 74 / 2609.38169
STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
Abstract
Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through successive state updates. We discover that the impact of these errors depends on two complementary dimensions: temporally, errors in long-lived memory can persist across many decoding steps; spatially, errors in different key rows affect model outputs differently, while state magnitudes vary substantially along both rows and columns. Motivated by these observations, we propose STEPQuant, a spatial-temporal post-training quantization framework for Delta-rule recurrent states. STEPQuant allocates precision according to error magnitude and memory lifetime, and jointly fits key-row and value-column scales based on state distributions and key-row impact on output error. Experiments on Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct across both long- and short-generation benchmarks show that STEPQuant closely matches FP32-state accuracy under a nominal 6-bit budget and outperforms uniform INT8 in its 4-bit configuration. Integrated into SGLang with optimized GPU kernels, 6-bit STEPQuant achieves over 5x recurrent-state compression and reduces total serving memory by up to 68.7%. Our code is available at https://github.com/Dreamer-Toby/STEPQuant.
cs.CL / 75 / 2609.36736
From Neurons to Conversation: Speech Brain-Computer Interfaces
Abstract
Speech brain-computer interfaces (BCIs) aim to restore communication by transforming neural activity related to speech, language, or communicative intent into external outputs such as text, synthesized voice, or avatar control. Recent advances in intracortical and electrocorticographic recording, deep sequence models, and language-model-assisted decoding have enabled rapid progress, including high-performance attempted-speech decoding and increasingly naturalistic speech synthesis. Yet these achievements also reveal that speech BCIs are not simply neural-to-text decoders. They are adaptive clinical systems in which neural representations, recording hardware, decoding architectures, language priors, feedback, and user learning interact over time. Here, we synthesize speech BCI research from a system-level perspective. We first examine the neural substrates of speech and language, emphasizing their hierarchical, distributed, temporally structured, and non-stationary organization. We then examine recording and decoding choices, closed-loop adaptation, evaluation, clinical translation, and ethics. Across these domains, we highlight recurring trade-offs between signal resolution and invasiveness, low-level motor and high-level semantic targets, decoder accuracy and user agency, and language-model fluency and faithful neural evidence. We argue the next generation of speech BCIs should be evaluated not only by offline accuracy, but also by robustness across sessions, calibration burden, latency, uncertainty, usability, and safeguards against unintended decoding. By reframing speech BCIs as adaptive, user-centred systems, we outline the interdisciplinary priorities spanning speech neuroscience, neural engineering, machine learning, clinical practice, and neuroethics needed to move from proof-of-concept decoding toward reliable, expressive, and controllable communication neuroprostheses.
cs.CL / 76 / 2609.36737
Reconstructing the Vocal Tract with Differentiable Acoustic Simulation
Abstract
The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by propagating sound along an acoustic tube model of the vocal tract, and via its gradients, can solve the inverse problem: reconstructing the shape of the vocal tract solely from the sound it produces. Although the inverse mapping between geometry and sound is notoriously non-convex, we discover that gradient descent succeeds with three technical contributions: (1) we design a frequency domain formulation of the vocal tract's fluid dynamics that is 70x more GPU parallelizable than finite differences in time, (2) we integrate a differentiable model for turbulence to synthesize consonants, and (3) similar to prior work in implicit neural representations (INRs) and neural fields, we find that parameterizing the geometry with a neural network accelerates convergence and escapes local minima that trap discrete representations. Because the simulator is differentiable, it is readily integrated with other deep learning pipelines to enable novel linguistics and medical imaging applications. (1) We demonstrate self-supervised autoencoding of vocal tract shapes across 11 languages, and (2) we couple our simulator with a generative model of MRI (magnetic resonance imaging) images to reconstruct one's moving vocal tract from only their speech without paired data.
cs.CL / 77 / 2609.38106
Pruning for Efficiency, Paying in Fairness: Demographic Disparities in Pruned Speech-LLMs
Abstract
Speech-LLMs are expensive to run, making compression important for real-world deployment. However, compressed models are usually selected using aggregate word error rate (WER), which can hide how pruning affects different demographic groups. In this work, we systematically study the effect of audio encoder pruning on SLAM-ASR for different demographic groups. Using the Fair-Speech and Common Voice datasets, we found that the pruning does not affect all demographic groups equally; the gap between best- and worst-performing groups increases in fold. These disparities appear across all three encoder scales, but only the largest model initially hides them behind aggregate WER. LoRA adaptation improves WER for every group, but benefits groups already performing well more strongly and widens for certain groups. On Common Voice English, Danish, and Dutch, accent gaps persist but do not clearly widen, showing that the fairness effects of pruning vary across datasets and must be measured directly. Our findings suggest that for pruned models, deployment decisions should include per-group WER, with the worst-performing group's error rate as an explicit criterion.
cs.CL / 78 / 2609.38157
EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation
Abstract
Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional training is costly in both computation and emotion-labeled speech training data. We therefore study vector steering, a training-free approach that modifies the internal representations of a frozen model. CoCoEmo, a conventional vector steering method for emotion TTS, treats each emotion vector as an indivisible direction controlled by a single global strength, limiting adherence to the requested emotion. In this work, we first discover that an emotion vector can be decomposed into a shared component that moves speech away from neutral expression and a residual component that directs generation toward the requested emotion. Building on this finding, we propose Emotion Residual-Enhanced Steering for TTS (EmoRES), a novel method that controls the two components without retraining the backbone. On IEMOCAP, EmoRES outperforms CoCoEmo across all four objective emotion metrics on the IndexTTS-2 and CosyVoice2 backbones. Rank correlation improves by 26.13 and 12.97 percentage points, corresponding to relative gains of 118.8% and 33.1%, while emotion hit rate improves by 12.95 and 6.92 points, corresponding to relative gains of 20.1% and 9.8%. Human evaluation further shows a relative improvement up to 35.0% in the rate at which listeners correctly identified the dominant requested emotion and up to a 17.3% improvement in fidelity, while listeners prefer EmoRES for naturalness in up to 63.8% of pairwise comparisons. Component ablations further demonstrate that effective control benefits from preserving the shared component while strengthening the residual of the emotion steering vectors.
cs.CL / 79 / 2609.36754
Does a prosody-trained representation help beyond trainable fusion? A parameter-matched study with frozen HuBERT
Abstract
Explicit prosodic cues may help automatic speech recognition (ASR) of spontaneous speech, but auxiliary representations typically require additional trainable components, making it unclear whether gains come from the auxiliary information or the fusion mechanism. We address this using a frozen HuBERT backbone and a 64-dimensional representation trained to predict log F0, voicing, Delta log F0, log energy, and spectral tilt. We compare a frozen-backbone recognizer (Baseline), trainable fusion with zero auxiliary input (Null), and the same fusion supplied with the learned representation (Learned). Across Buckeye, Switchboard, and AMI IHM, Null reduces WER by 0.71-1.45 points over Baseline, whereas Learned differs from Null by +0.07, -0.09, and +0.00 points, with no significant differences. However, removing or mismatching the representation at inference increases Learned WER. Thus, Learned depends on the representation yet shows no measurable incremental WER benefit over the parameter-matched control.
cs.CL / 80 / 2609.36097
Better Behavioral Prediction, More Faithful Model Ablations? Evidence from Sequential Choice
Abstract
Using predictive models to explain cognition requires more than accurate behavioral predictions. Input ablations offer an appealing route: remove information from a model and interpret the resulting performance change as evidence of its importance for behavior. Yet this inference assumes that the model's dependence on information reflects the dependence of the process generating the behavior. We test it in two synthetic sequential bandit tasks with known generating policies, where past choices can remain informative when feedback is unavailable to a predictor. We compare GRUs and Transformers trained from scratch, a fine-tuned LLaMA model, and cognitive models across systematically varied reward contributions. Our analyses distinguish prediction after training without reward observations from the response of a fixed predictor to donor-reward replacement. Three findings emerge. First, in the restless task, neural models trained without rewards predict held-out choices better than four simple training-fitted behavioral baselines. Second, under matched donor replacement, accurate predictors can respond much less than the known generator. Third, at some reward weights, neural networks predict better than a pooled reinforcement-learning model but have less faithful changes in choice probabilities; the model ordering differs between the two tasks. These independent-test results separate information sufficient for prediction from response fidelity under a specified ablation in sequential choice. They motivate validating model-ablation responses independently of predictive performance before using them to infer how the observed behavior was generated.
多智能体系统 (cs.MA)
10
cs.MA / 1 / 2609.36424
Agent-Based Evolutionary Dynamics for Mixed Autonomy Weaving Ramps
Abstract
Existing models of mixed-autonomy weaving ramps characterize how altruistic connected and automated vehicles (CAVs) can improve traffic efficiency at the population level, but provide limited insight into how such behavior emerges from decentralized vehicle interactions or how it is affected by finite populations, heterogeneous preferences, and imperfect information. We develop an agent-based model of a macroscopic weaving-ramp framework in which individual vehicles adapt their lane choices using an evolutionary game-theoretic update rule and altruism-based objectives providing a microscopic interpretation of the original Wardrop model. We prove convergence of the decentralized dynamics to the unique equilibrium predicted by the macroscopic theory. Beyond reproducing aggregate equilibrium behavior, the framework enables the study of deployment-level questions that cannot be addressed by static analysis. Simulation results demonstrate close agreement with the macroscopic predictions while revealing how convergence rates, adaptation to changing traffic conditions, heterogeneous altruism levels among CAVs, and imperfect state information influence system performance and the distribution of altruistic burden across vehicles. These results provide a bridge between equilibrium traffic theory and decentralized mixed-autonomy deployment.
cs.MA / 2 / 2609.36787
Regularized policy gradient with learned mixtures of Gaussians for games with continuous actions
Abstract
Most successes of superhuman game-playing algorithms are in games with discrete actions, yet in auctions, robotics, sports, or trading, actions are nearly continuous. Prior techniques either rely on expert-designed discretizations or are sample inefficient. We present a scalable policy-gradient algorithm for large sequential games with continuous or mixed discrete and continuous actions. It combines magnetic mirror descent with a mixture of Gaussians reparametrization, trained via self-play. We show that it approximates equilibrium in games where gradient descent fails. In sequential games, it outperforms neural fictitious self-play and matches or outperforms the final strategies of policy space response oracles with 3.5--5.5$\times$ fewer samples. In heads-up no-limit Texas hold'em, it performs on par with Slumbot.
cs.MA / 3 / 2609.36789
GitHarness: Git Init Your Harness Working Memory for Perpetual User Requirements
Abstract
LLM-based agents increasingly collaborate with users on long-horizon tasks, accumulating evidence, code, and drafts through extensive search, reasoning, and execution. As users inspect these results, they may supply missing information requirement completion, introduce new requirements requirement elicitation, or revise existing ones requirement shift. These changes often affect only part of the accumulated work, yet agents may carry forward obsolete information or turn local revisions into global rewrites. Existing approaches clarify current intent without determining how prior work should change, or reuse execution histories under a fixed objective. We address this gap by formulating dynamic-requirement collaboration as joint requirement tracking and local update. We introduce GitHarness, a pluggable Git-style framework that organizes requirement states and their corresponding harness work states into a branchable version history. A trainable Git Agent resolves requirement changes and selects a semantically compatible historical state. A unified version interface then restores that state and creates a new branch, enabling the underlying harness to exclude obsolete information, inherit compatible work, and focus execution on affected parts. The Git Agent is trained through interface-level black-box reinforcement learning, with downstream harnesses and task-execution models kept fixed. We also construct MTAgentBench, a verifier-preserving benchmark covering mathematical reasoning, text-to-SQL, agentic search, software engineering, and research synthesis. Experiments demonstrate strong task performance alongside effective requirement tracking, preservation of valid work, and efficient execution.
cs.MA / 4 / 2609.37151
FlowMAS: Learning Multi-Agent Workflow Topology via Information-guided Generative Flow Network
Abstract
Automated multi-agent systems offer clear advantages over manually designed ones in scalability and adaptability, but existing workflow topology methods still face important limitations. Search-based methods are often computationally expensive, textual-gradient-based methods rely on coarse-grained feedback, and existing generation-based methods are not well suited to discrete workflow topologies with complex dependencies. To address these limitations, we propose FlowMAS, a multi-agent workflow topology method based on Generative Flow Networks (GFlowNets). FlowMAS models workflow generation as reward-guided flow over the topology space and introduces three components: a GFlowNet-based topology generation backbone, a curiosity-driven module for structure-aware exploration, and an information-guided optimization module for evaluating intermediate topologies. Concretely, the curiosity-driven module encourages exploration of structurally novel workflows, while the information-guided module measures both the information contribution and the communication efficiency of different operators to favor more informative and effective collaboration patterns. Experiments on six benchmark datasets with three LLM backbones show that FlowMAS consistently outperforms multiple baselines.
cs.MA / 5 / 2609.37321
PowerMarketJax: A JAX Benchmark Suite for Multi-Agent Reinforcement Learning in Power Markets
Abstract
Power markets are a natural testbed for multi-agent reinforcement learning (MARL), where multiple self-interested participants repeatedly submit bids. A market-clearing mechanism then determines dispatch and prices subject to power grid constraints and market settlement rules. However, existing MARL environments typically focus on a single market setting, implement simplified clearing mechanisms, or rely on CPU-based optimization solvers that slow large-scale training and limit the systematic study of bidding strategies and market behavior. We introduce PowerMarketJax, a benchmark suite for MARL across five power markets: day-ahead wholesale, real-time balancing, ancillary services, peer-to-peer double auctions, and local flexibility. Each environment implements its own clearing, pricing, and settlement rules while providing a common framework for learning and evaluation. We find that learned bidding behavior depends strongly on the market design: independent learners can miss better strategies when gains require many agents to change together, when more profitable strategies lie beyond a region of lower profit, or when profits disappear as more agents adopt the same strategy. PowerMarketJax implements both market simulation and policy training in JAX, allowing the entire pipeline to run on the GPU with 1,024 X 1,200 parallelisms across both environments and market participants, achieving up to 33X speedup over CPU-based baselines. Our open-source benchmark is available at: https://github.com/powermarketjax/PowerMarketJax.
cs.MA / 6 / 2609.37566
RAVEN: Receiver-Conditioned Action-Value Encoding for Finite-Alphabet Multi-Agent Communication
Abstract
A message drawn from a small alphabet helps a teammate only if it keeps the distinctions that change that teammate's next decision. We show that scoring messages by action values averaged over the receiver's situation can erase exactly these distinctions, and we propose RAVEN (Receiver-conditioned Action-Value ENcoding), which trains a four-symbol, one-step-delayed channel to preserve each receiver's centered action-value profile within the receiver's own context. The sender never needs to know that context: the receiver decodes every symbol with its private information. We give two estimators of this target. With a teacher, offline RAVEN selects the codebook that exactly minimizes an empirical conditional distortion and distills it into a frozen sender; we bound the resulting codebook-selection error and one-step decision loss. Without a teacher, online RAVEN aligns, inside a QMIX learner, the deployed symbol pathway with a training-only continuous reference that shares its routing. Against five recent communication methods on eight navigation settings, offline RAVEN attains the highest return in seven, and removing receiver conditioning forfeits 83% of its communication gain. Online RAVEN raises predator-prey capture success from 53.2% to 96.0% over the same QMIX backbone without communication, and on SMAC and MPE it attains the best mean normalized score of 14 methods, including methods that exchange kilobit messages. Every RAVEN message costs 2 bits, 12-1,024x fewer than those of NDQ, CACOM and ExpoComm on navigation.
cs.MA / 7 / 2609.38113
IMPACT: Modeling Socially Interdependent Movement in a Generative Multi-Agent Simulation of a Pompeian Household
Abstract
Simulations of archaeological sites can make interpretations of past cultural practices observable and examinable. Generative multi-agent simulations offer a bottom-up approach to modeling how people collectively moved through and used historical spaces. However, current agents designed to simulate everyday life often plan and act independently, limiting their ability to capture how movement depends on others' actions. We introduce IMPACT (Interdependent Movement Planning through Inter-Agent Constraints and Triggers), an architecture that uses culturally specific roles and obligations to define dependencies among agents' activities and guide coordination. IMPACT connects socially gated milestone planning, wait-or-prompt resolution, structured directive issuance, and directive integration. These mechanisms determine whether and when activities can begin or change as social conditions evolve, producing socially constrained and prompted movement as their primary observable outcome. We instantiate IMPACT in a five-hour simulation of a Pompeian dinner involving ten agents across interdependent roles. Analysis of five simulation runs shows how social roles, responsibilities, and status relations shape household activities and spatial practices, as reflected in patterns of co-location, asymmetric waiting, co-movement, and social directives. In a controlled ablation evaluation, thirty-seven participants rated the complete architecture's behavior as more socially coherent and believable than that of two reduced architectures. Interviews with six archaeology experts highlighted historically plausible movement patterns and the simulation's potential to support archaeological interpretation, while identifying areas requiring stronger historical grounding for future work.
cs.MA / 8 / 2609.37126
Adversarially Robust Geometric Safety Certificates for Nonholonomic Robots Against Maneuvering Obstacles
Abstract
Safe navigation against obstacles that can actively maneuver within bounded capabilities remains challenging: robust control barrier function methods typically treat obstacle actions as generic disturbances, while differential-game approaches are computationally expensive for online navigation. We propose an adversarially robust geometric certificate that accounts for the worst-case effect of admissible obstacle maneuvers directly in the safe-set geometry through a closed-form contraction of the certificate parameters. The construction exploits a structural property of line-of-sight (LoS) certificates: the robot and obstacle actions enter the certificate through a common state-dependent geometric gain. This gain cancels in the worst-case comparison, reducing the differential game to a direct comparison between obstacle maneuvering capability and the weaker of the robot's longitudinal and steering authorities. Instantiated on the parabolic certificate, the construction yields Adversarially Robust Dynamic Parabolic Control Barrier Functions (AR-DPCBF), for which we establish sufficient conditions for forward invariance of the contracted safe set against all admissible obstacle maneuvers under kinematic bicycle dynamics with bounded inputs. When the obstacle capability is unknown, a sliding-window estimator supplies a high-probability upper bound, allowing the guarantee to be retained with the corresponding coverage probability. We further formulate soft and buffered variants to recover feasibility in dense environments. Simulations across obstacle capabilities, densities, and capability mismatch show substantial reductions in barrier violations and collisions and demonstrate that pointwise robustification of the barrier derivative cannot substitute for contraction of its geometry.
cs.MA / 9 / 2609.36292
Fully Decentralized and Safety-Aware Multi-Agent Reinforcement Learning for Control on Networks
Abstract
This paper develops a safe and fully decentralized multi-agent reinforcement learning (MARL) algorithm to solve a class of discrete-time control problems on networks, including the persistent monitoring problem. Fully decentralized control of agents, while offering numerous benefits, faces issues such as exponentially increasing sample complexity, lack of global information about the system, and challenges in coordinating between agents. To address these issues, this paper introduces a fully decentralized multi-agent reinforcement learning algorithm that integrates deep reinforcement learning with safety considerations. This method feeds a history of local observations of the network's state into two parallel neural-network branches: the graph encoder, which adds structural information and correlations among nodes, and a state estimator, which predicts the uncertainty at each node in the graph. Additionally, the result of feeding that input into an actor-critic network is passed through a discrete-time control barrier heuristic to reduce the likelihood that any node will be neglected. This approach enables teams of fully decentralized agents to solve challenging problems by increasing system awareness and incorporating built-in safety measures to prevent the adoption of potentially harmful control policies. Numerical results from a custom simulation environment demonstrate that the proposed algorithm achieves 26.3 percent lower average uncertainty than a centralized control policy and is within 1 percent of the uncertainty performance of a more computationally complex algorithm with added attention layers.
cs.MA / 10 / 2609.37009
An LLM-powered Agent Framework for Heterogeneous Evacuation Behavior Modeling under a Moving Threat in a Public Plaza
Abstract
Modeling heterogeneous evacuation behavior under a moving threat is difficult because human perception, memory, and evidence evaluation are not well captured by fixed rules. We propose a novel LLM-powered agent-based framework to represent these internal decision processes. Each pedestrian agent perceives a private symbolic ASCII view, maintains a Memory-based Knowledge Graph derived solely from individual observations, and makes decisions through persona-conditioned prompts under a common sampling configuration. A compressed decision context with stateless memory preserves trial-and-error experience across turns while excluding reasoning traces, and a validation engine separates behavioral choice from physical feasibility by executing routes only over observed terrain. We evaluated eight personality compositions in eight paired randomized blocks within a simulated public plaza. Usable-exit knowledge was strongly associated with evacuation success: 89.5% of agents possessing such knowledge evacuated, compared with 1.05% of those without it. Personality compositions also differed in their evaluation of remembered threat evidence: the proportion of danger assessments varied by 0.265, while high-urgency, low-directness decisions ranged from 11.41% to 34.33%. After direct threat sightings, responses converged, with 99.6% of assessments classifying the situation as dangerous. Movement was selected in 99.8% of decisions. Overall, evacuation outcomes were strongly associated with information access, while evacuation time was jointly associated with spatial geometry, information, and affect. The framework provides an auditable approach to generating endogenous behavioral heterogeneity through persona-conditioned LLM agents in crowd-evacuation simulations.
软件工程 (cs.SE)
21
cs.SE / 1 / 2609.36065
Irene: Equivalence Checking of Hybrid Quantum Programs via Structure-Preserving Symbolic Reduction
Abstract
Equivalence checking is essential for validating compiler transformations of hybrid quantum programs, which combine quantum operations, measurements, and classical control. Measurement-dependent control limits unitary reasoning, while dependencies between classical outcomes and quantum operations can enlarge intermediate symbolic states. We present Irene, an equivalence-checking framework for bounded hybrid quantum programs based on structure-preserving symbolic reduction. The framework progressively simplifies equivalence obligations through three levels of reasoning. At the gate level, algebraic identities simplify unitary regions. At the hybrid path-sum (HPS) level, reduced symbolic execution states are represented as typed graphs, whose isomorphism certifies equivalence. Remaining obligations are handled by density kernels that characterize transformations of input density operators into observable outputs, allowing comparison even when internal measurement histories differ. Residual coefficient differences are encoded as SMT queries. A common set of symbolic reductions supports HPS and density-kernel reasoning by preserving factored Boolean and arithmetic expressions, eliminating reducible dependencies before expanding residual sums. We evaluate Irene against five equivalence checkers on 1,982 program pairs from seven benchmark suites. Irene solves 1,584 pairs (79.92%), compared with 57.52% for MQT QCEC, the baseline with the highest aggregate coverage, with a mean end-to-end time of 3.93 seconds per solved pair. Applied as an equivalence-checking oracle, Irene also identifies 15 previously unknown bugs in quantum compilers, including Qiskit, Cirq, and PennyLane.
cs.SE / 2 / 2609.36069
The Uneven Decline of Collective Knowledge Production: Evidence from Stack Overflow After Generative AI
Abstract
Generative AI (Gen AI) is reshaping how individuals learn and work, but its consequences for collective knowledge, the shared body of knowledge that online communities produce together, remain poorly understood. Prior work has documented an aggregate decline in participation on knowledge-sharing platforms, but it remains unclear which specific kinds of knowledge are being lost first. We study this question using Stack Overflow, one of the largest online communities for software engineering, treating the release of ChatGPT-3.5 as a natural shock. Analyzing over two million questions posted between 2020 and 2025, we track how two dimensions of collective knowledge, difficulty and data availability, change following Gen AI's release. Using diverse methods and robust checks, we find consistent patterns. Easy questions decline sharply while difficult questions become more common, a pattern corroborated by rising code complexity. Data-rich topics and tags lose share of questions, while data-scarce ones gain ground. The two dimensions also interact: the decline in easy questions is concentrated specifically within data-rich domains, while difficult questions increase regardless of data availability. This pattern extends beyond Python across programming languages, with more prevalent languages showing sharper shifts. Together, our findings reveal that Gen AI's impact on collective knowledge is uneven, eroding easy, accessible knowledge first while more complex, less common knowledge persists.
cs.SE / 3 / 2609.36147
The Invisible Scheduler: Dragging a Sprite Can Change What a Scratch Program Does
Abstract
Tens of millions of children program in Scratch, whose programs are concurrent: when the green flag is clicked, every sprite's scripts start together and share the project's state. Which script starts first is decided by the sprites' front-to-back stacking order, which no block reads. Dragging a sprite brings it to the front, and the order is saved with the project. Hence, a program can work on the author's screen and fail on the teacher's, with the same blocks. Our key observation is that, with the rest of the file fixed, rearranging the sprites that start scripts among their positions produces exactly the permutations of their initial stacking order. Under a fixed input and seed the unmodified virtual machine reproduces every one of them. StackSwap runs a project under these orders, all of them for up to five sprites, one per class of a conflict graph beyond, and a sample where neither is feasible, and compares the runs under four observation lenses. It returns a two-run witness naming an adjacent pair of sprites whose swap changes the outcome, with the resources they share as candidate causes; otherwise an exhaustive or graph-conditional robustness certificate, or an unresolved verdict. On 767 real programs from a course, an online judge, and a random public sample, 26.4% (course) and 18.2% (public) of those with a scheduling choice behave differently under some stacking order. Among the sensitive programs whose every order ran, the saved order's outcome recurs under a median 17% (course) and 50% (public) of the orders. Under scripted play, a predicate written from a stated requirement holds under one order and fails under another for 5 of 427 student submissions with a scheduling choice. For half of the sensitive programs the random stream is the racing pair's highest-priority shared resource. We close with recommendations for learners, graders, and the Scratch platform.
cs.SE / 4 / 2609.36148
Live Architecture Models for Cloud-Native Architecture-as-Code: Early Results from Kubernetes Conformance Checking
Abstract
Infrastructure drift can separate a running Kubernetes system from its documented architectural intent. This paper investigates live architecture models: editable architectural representations connected to selected runtime facts through explicit correspondences and recurring conformance checks. The approach is instantiated through an Archer-specific subset of Kubernetes Deployment Language (KDL) and Archer, a VS Code prototype with synchronized textual and graphical views. Snapshot recovery extracts selected Kubernetes facts into KDL; periodic and on-demand read-only checks report model-cluster inconsistencies without enforcing or repairing deployment state. We assess feasibility on three feature-selected Kubernetes example applications under an author-defined protocol, reporting precision and recall for snapshot recovery and selected inconsistency detection. Recovery scores can be reproduced from saved ground-truth and recovered models; detection uses a documented manual perturbation protocol. The results support feasibility within the evaluated KDL scope. They do not establish comparative superiority, detection of arbitrary production drift, scalability, or developer benefit. Storage coverage and ingress-host representation/default handling remain limitations.
cs.SE / 5 / 2609.36152
The Invisible Throttle: Running on Borrowed Time in Scratch
Abstract
Scratch has 135 million registered users, most of them children, and 164 million shared projects. What they are taught about the speed of a script fits in one sentence: a loop iterates once per frame. Unfortunately, that sentence describes the exception. In the public virtual machine a frame repeats the scripts until something visible asks the screen to redraw, no script is left running, or three quarters of the frame's wall-clock time are spent. Hence a loop that does not draw is paced by whatever else is visible and by the machine. Hide the moving sprite of a two-sprite project, and the other's counting loop runs 77,000 times faster on a laptop. No documentation states the rule. Our key observation is that the redraw gate is one flag for the whole runtime, so a loop's speed can be changed without touching its code: hide the sprite that draws, change the machine's budget, or run the program in a tool with no renderer. ThrottleCheck implements a budgeted semantics (a stated budget of rounds per frame, the redraw gate emulated) on the unmodified virtual machine, runs a project under each knob, and reports a rate-sensitive project with a witness. On 500 popular public games, 59% contain a loop that never draws and never waits. Muting the requests of the sprites that draw changes the state of a played game after ten seconds in 24% of the games that have one; hiding them changes 18%; running the game without a renderer and without the gate changes 61%. Rules that read a position, a score or a clock after a fixed time reverse between pass and fail once the throttle is released, in 17% of Whisker's own example tests and 19% of a tutorial's checks on 224 of its remixes. A grader without a renderer grades a program the editor never runs, unless it emulates the gate and states a budget; we say what graders and the platform should do, and close with the sentence a child could be taught.
cs.SE / 6 / 2609.36161
From Dead Code and Static Requirements to Working Engines: Software Revival with Coding Agents
Abstract
Can coding agents restore software that no longer runs while preserving its underlying methods, and reconstruct industrial software engines from open specifications? Here we introduce ReviveBench, a benchmark with two task families evaluated by hidden verifiers calibrated against native execution environments, established engineering tools, or purpose-built reference implementations. The revival family comprises ten tasks involving dependency incompatibilities, deleted core modules, legacy builds, and a GPU-based foundation model. Every starting workspace fails verification, and the strongest evaluated model passes all ten tasks in at least one run each. In contamination-control experiments, identifier obfuscation reduces line similarity to the original implementations from 0.51--0.96 to 0.03--0.44 without reducing the observed pass rate of any evaluated model. On repositories created after the stated knowledge cutoffs, the strongest model passes eight of nine runs. The reconstruction family comprises thirteen tasks spanning numerical, geometric, hardware, and transactional systems (e.g. CAD and CRM). Two models meet the benchmark's pass criteria on all thirteen, although our audit shows that the CFD task cannot establish numerical-solver capability. Benchmark construction and auditing uncover 28 verifier defects, including 24 false negatives and two false positives. These findings show that executable verification can itself introduce substantial measurement error. We present three practical checks: test whether prescribed methods can reach the grading thresholds, investigate agreement among independently generated candidates, and recompute diagnostics from submitted artifacts. ReviveBench thus provides both an evaluation of software revival and engine reconstruction, and cases in validating the verifiers used to measure coding agents for software design.
cs.SE / 7 / 2609.36170
Assay: Claims That Decay With the Code. Content-Addressed Evidence Graphs for Accountable AI-Assisted Software Delivery
Abstract
AI coding agents fail in two coupled ways. They spend most of their context window rediscovering where things live, and they assert success without evidence when the work gets hard. Repository indexes address the first with cheap context, and orchestration frameworks with adversarial review address the second with accountability. Both describe the same object, the structure of the codebase, at two timescales: what is true of the code now, and what was verified to be true, at which revision, by whom. Assay makes that observation operational. Every claim an agent makes (tests pass, no secrets, behavior preserved) is bound to the Merkle hash of the dependency cone of the code it covers, so the claim is stale exactly when that code or anything it depends on changes. We show the binding is sound and minimal, and that the blast radius of a change is precisely the set of claims it invalidates. On the graph we place a risk-proportional evidence obligation, a bounded review protocol with separation of duties, and a merge gate that consults no model: coverage, freshness, signatures, exit codes, plausibility, evidence monotonicity (the mechanical form of "do not delete the failing test"), and review status. Assay is a dependency-free Python tool with an MCP server. On five public repositories a 600-token brief costs 14x to 114x less than an exploration proxy, warm rebuilds are up to 5x faster than cold ones, cone binding re-verifies 7.9% to 81.9% of claims where repository binding re-verifies all of them while per-module binding misses 23% to 68% of required invalidations, and the gate blocks 9 of 9 scripted adversarial behaviours while admitting the honest ones. Every number in this paper is generated by the released scripts.
cs.SE / 8 / 2609.36362
Strategies for Deploying AI Agents in Production at Scientific User Facilities
Abstract
Agentic artificial intelligence (AI) is moving beyond research demonstrations toward production use at scientific user facilities, including light sources, neutron sources, nanoscience centers, and autonomous laboratories. Its scientific value extends beyond increasing throughput. Agents can perform repeatable tasks in calibration, measurement execution, and quality control, as well as initial analyses that turn data into reviewable evidence, allowing scientists to focus on hypotheses, unexpected observations, and interpretation. Drawing on deployments of LLM-driven agents at the APS, this perspective distills practical strategies with an emphasis on elements that can be reused across instruments and facilities. We discuss agent harnesses for beamline control, facility knowledge retrieval, and data analysis while keeping the underlying design principles independent of any specific implementation. These principles cover inference endpoints, tool-server architectures, non-text data, computationally intensive services, reusable skills, and governed learning throughout an instrument's lifecycle. We also consider how network and Linux operations, governed shared memory, and deterministic orchestration can extend these patterns across facility services. Because LLM capabilities continue to evolve, these recommendations represent a snapshot of the technology as of the date on the cover.
cs.SE / 9 / 2609.36371
LatentSift: Policy-State Filtering for Token-Efficient Verification of Software Engineering Agents
Abstract
Test-time scaling improves software engineering agents by generating multiple candidate trajectories and selecting the best one. Verifying and selecting among these long interactions can consume as many tokens as generation itself. Existing hybrid workflows first apply an LLM-based execution-free (EF) verifier to filter candidates before running tests, which adds another model pass over every trajectory. We introduce LatentSift, a token-free and execution-free filter that replaces this first stage with hidden states the policy already produces while generating the candidates. It represents each candidate through its reasoning, observation, and function-call states, compares them with positive and negative banks of such states collected from successful and unsuccessful trajectories during policy training, and fuses the resulting distance scores with a learned linear score to retain promising candidates for the execution-based stages. On SWE-bench Verified, across three agents and two policy sizes, LatentSift cuts EF-verifier tokens by 66.6--81.0% and total verification tokens, which include test generation, by 49.1--62.1% at K=16, while hybrid Best@16 matches or improves on each agent's reference workflow, rising from 59.26% to 60.06% on DeepSWE-Preview.
cs.SE / 10 / 2609.36635
WitnessGym: Benchmarking Coding Agents on the Construction of Bug Witnesses
Abstract
Bug validation asks a coding agent to produce an executable witness for a reported bug. The witness combines a concrete input with a testing harness and exposes faulty behavior during execution. Such evidence makes audit findings actionable, yet benchmark evaluation is difficult when cases reuse public historical bugs and witnesses or require manual construction. We present WitnessGym, an automated framework for constructing bug-validation benchmarks through bug injection. It injects bugs into test-reached paths of real projects, rebuilds each project, and retains cases exposed by a construction-time witness. Bug specifications and execution adapters allow extension to additional bug types and languages. Bug-preserving transformations vary the surrounding structure while preserving the witness behavior. Based on real-world Java projects with test suites, WitnessGym automatically constructs 1,300 benchmark cases. The injected patches resemble historical bug patches and are difficult for the two evaluated models to distinguish in blinded comparisons. We evaluate four coding agent frameworks in six framework/model pairings across bug types, execution contexts, and transformation depths. Witness construction remains difficult even when the bug pattern is known. Our framework, benchmark cases, and evaluation scripts are available.
cs.SE / 11 / 2609.36647
CTE-Bench: Counterfactual Trace Evaluation for Stateful Software Simulators
Abstract
Coding agents change running software: they patch a service's code or overwrite its stored state, and then act on their own expectation of how the service will respond afterwards. A wrong expectation may surface only several calls later. Function-level code-execution benchmarks omit persistent service state, and agent benchmarks score the actions an agent takes or the final state it reaches. We introduce CTE-Bench, which measures whether a model can predict how an intervention changes a stateful service's future behavior, without asking it to choose actions. Each scenario gives the model Python service code, the calls and responses observed before the intervention, the intervention itself (a source edit or a state overwrite), and 40 fixed future calls; the model predicts every future response, and predictions are checked by executing the service. Three memory protocols control whether the model sees the correct earlier responses, none of them, or its own earlier predictions. CTE-Bench-Core-v1 contains 255 scenarios over six deterministic Python services, giving 10,200 predictions per model. The main score is effect-step value match (VM): exact response equality on the 2,476 future calls whose response the intervention changes. With correct earlier responses revealed, four API-hosted models (DeepSeek V4-Flash, Kimi K2.5, Qwen3.6-35B-A3B, and Claude Sonnet 4.6) reach 54.3%-61.5% effect-step VM. Hiding those responses lowers effect-step VM to 23.2%-28.9%; conditioning on self-generated predictions gives 24.8%-33.2%, and at most 1.2% of scenarios are predicted exactly end to end. Current models thus track intervention effects mainly when correct feedback is supplied, and their errors compound over a rollout. We release CTE-Bench-Core-v1 with its executable oracle, evaluation scripts, and an evaluation card mapping each claim to its protocol.
cs.SE / 12 / 2609.36807
XRepoSkill: Learning Transferable Skills for Software Engineering Agents
Abstract
Software engineering agents increasingly use reusable skills distilled from prior experience to resolve repository-level issues, yet such skills often fail to transfer across repositories. A central challenge is that a behavior appearing in a successful trajectory is not necessarily responsible for the successful outcome: it may be genuinely useful, merely incidental, or simply a recurring habit of the model. We introduce XRepoSkill, a trajectory-based approach for learning transferable skills. We represent a skill as a collection of rules, each specifying what action to take and when to take it during issue resolution. XRepoSkill first contrasts successful and failed trajectories of the same agent on the same issue and derives candidate rules from where their execution paths diverge. Each rule is paired with an executable predicate that enables its prescribed behavior to be evaluated systematically on other trajectories. A rule is verified based on its association with successful issue resolution and retained only when its prescribed behavior recurs across multiple repositories; repository-specific variants of the same behavior are then consolidated into transferable rules. For a new issue, XRepoSkill selects relevant rules to guide the agent. We learn skills from publicly released trajectories on the official SWE-bench Verified leaderboard and evaluate them on SWE-bench Pro and DeepSWE using three backbone LLMs from different vendors; none of the evaluation repositories appears in the skill-learning trajectory pool. Against three recent skill learning methods, XRepoSkill achieves the highest issue resolution rate in all six benchmark--LLM combinations. In particular, on the challenging long-horizon DeepSWE benchmark, XRepoSkill improves issue resolution by 10.3 percentage points over the same agent without learned skills and by 5.0 points over the strongest skill-learning baseline.
cs.SE / 13 / 2609.37045
GitCF: Reducing Incomplete Changes by Exploiting Multiple Similarities Among Commits
Abstract
Software developers often struggle to identify all locations that their changes affect, and incomplete changes frequently result from these omissions. To mitigate this problem, several approaches have been proposed to recommend additional locations that should be modified together with a developer's current change. A large class of these techniques mines co-change rules from repository revision histories, but because they rely solely on which elements have been modified together, they cannot recommend elements that have rarely co-changed, and they disregard the textual information that accompanies commits. A second class of techniques does exploit textual information, but it derives it from the source code or from a change request rather than from commit history, and it does not use the developer's current change as its input. To address these limitations, we propose GitCF, a change recommendation method that compares a developer's current change with past changes in the revision history and recommends additional change locations by combining two sources of similarity: the set of modified elements and the textual content of commits, including commit messages, code diffs, and linked issue descriptions. We evaluate GitCF on 543 incomplete changes from 17 open-source projects, where an incomplete change denotes a commit in which an omitted modification is supplied in a later commit, comparing it against three co-change-rule-based techniques. GitCF outperforms the strongest baseline on all five metrics (MAP, Recall@10, Recall@20, Hit@10, and Hit@20); in particular, it raises Hit@10 from 0.322 to 0.403, and the improvement in average precision is statistically significant.
cs.SE / 14 / 2609.37143
LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems
Abstract
Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents' implementation capability to produce correct code edits from detailed specifications. However, practical modular development tasks also require the perception capability of grounding user intent and high-level design to derive a specification. We introduce LoLBench to evaluate both capabilities through the entire proposal-to-implementation process on large software systems. It is a multilingual benchmark of 100 tasks across 29 software systems in five domains. Each task provides a human-written enhancement proposal with user intent and high-level design. On average, proposals contain about 5,000 words, software systems contain 2.4 million source lines of code (LoC), and implementation pull requests (PRs) change approximately 5,500 LoC. Across 28 agents we evaluated, the best agent resolves only 14% of tasks and achieves a 52.7% Fail-to-Pass (F2P) pass rate. Failure analysis identifies incomplete code localization as a major bottleneck, while providing reference-derived file trees alongside API specifications improves resolved rates by 16--22 percentage points (2.4--17$\times$), reaching at most 34%. These results show that both perception and implementation remain central challenges for coding agents in practical modular development on large software systems. LoLBench is available at https://huggingface.co/datasets/lolbench26/LoLBench.
cs.SE / 15 / 2609.37315
Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments
Abstract
Tool-using agents are entering settings where a wrong action carries real cost, and the benchmarks certifying them grade what each simulated tool call reports having done, assuming the tool did what its interface advertises. The audit taxonomies we survey publish no category for that assumption, and a defect beneath a score is present on every rerun. We treat a tool's advertised surfaces as an executable contract, check the implementation against it, and trace each score's provenance through the task files and evaluator code to the verdicts that derive from state a defective tool should have written. Across 34 audited mutating tools in four benchmarks we confirm seven tool defects and one evaluator property at pinned commits. On injected defects the checker raised no false positive in 25 flags, flagged 2 of 5 negative controls, and missed most: in 29 of 33 scored misses a clause covered the defect but no probe revealed it. The checker's own static half, run alone, flags 14 of 17 confirmed sites, so on these findings the dynamic half confirms and traces rather than discovers. Twelve further AgentDojo tools, with six held-out tools and the seven audited first, complete its 25-tool mutating surface, on which at least 5 tools diverge from their advertised surface as our contracts read it, a rate for AgentDojo alone. No gold trajectory reaches either tau2-bench defect; on 1,120 paths built to isolate the telecom defect, a number fixed by construction, the evaluator rewards a refuel of a suspended line and fails the repaired tool. The clearest case is a clinical benchmark whose tool tells the agent each write executed under a documented no-write design its interface does not disclose; its grader takes that message as evidence, so its action success rate records whether a request carried the expected payload, not whether any record changed.
cs.SE / 16 / 2609.37526
Exploring Emotional Intelligence in Software Testing
Abstract
Background: Emotional Intelligence (EI) is the ability to recognise, understand, and manage one's own and others' emotions. Software testers deliver judgements about colleagues' work under deadlines they do not control, and prior work on emotion in software engineering has mostly studied developers. Aims: To explore how software testers describe the part EI plays in their day-to-day work, in communication and conflict within the team, and in responding to requirements volatility. Method: Semi-structured interviews with 16 software testers in Sweden working in teams that use agile practices, across aviation, automotive, healthcare, IT services, administration, banking and pharmaceuticals, analysed with reflexive thematic analysis informed by Goleman's EI framework. Results: Three themes. Testers described regulating stress under deadline pressure and drawing motivation from recognition, clarity and autonomy; managing the daily delivery of critical findings to colleagues so that trust survives; and responding to requirements change with frustration that turned into decisions about what to leave untested, into advocacy for process change, or into workarounds. Read against developer-focused studies, the themes point to features of the testing role: the work product is a criticism of a colleague's work, success is invisible while failure is attributed, and the tester's window shrinks with every upstream delay. Conclusions: For testers, managing emotions is a constant job requirement. The results highlight that the importance of EI increases when the development process lacks an independent testing phase. The findings also inform implications for teams and, ultimately, for organisations and future research.
cs.SE / 17 / 2609.37531
Profiling the Energy Consumption of Serverless Functions with Joule Profiler
Abstract
Cloud providers and customers have widely adopted serverless computing as a convenient paradigm for deploying and executing functions on demand. To do so, serverless platforms require provisioning an appropriate execution environment before a single line of the function's code runs. These environments consist of several layers, such as container engines, hypervisors, unikernels, and programming language runtimes. While the literature has investigated the performance of these serverless platforms, it treats functions as black boxes, and the community lacks key insights into the environmental impacts of packaging applications as serverless functions. This paper therefore empirically studies the energy efficiency of serverless functions deployable on serverless platforms. We design an experimental benchmarking environment that lets stakeholders explore the impacts of the various layers involved in executing serverless functions. We use it to evaluate 1,401 configurations, combining 9 execution environments, 7 language-runtime configurations, 11 workloads, and 3 input sizes, to answer three research questions: Are the most popular programming languages for serverless functions the most energy-efficient? What factors most affect their energy efficiency? What are the most energy-efficient configurations to deploy them? Our results show that one should first choose the programming language, then the language runtime, and only then the execution environment, which matters only for short-lived functions and whose best choice depends on the runtime. Our benchmarking environment, experimental artifacts, raw measurements, and analysis code are publicly available.
cs.SE / 18 / 2609.37603
Independent Verification Paths Are Not Independent: A Case Study of Common-Mode Failure in a Satellite Catalogue Pipeline
Abstract
A common safeguard for a data pipeline is redundant computation: derive each published number by two routes built on different technology and refuse to exit when they disagree. We report one such gate failing, in a cross-catalogue integrity study of two open registers of Earth-orbiting objects. A gate comparing a set-based Python path with SPARQL queries over the emitted RDF graph printed ALL CROSS-CHECKS AGREE on seven counts. Three were wrong, one overstated more than fourfold (932 against 220). Both paths imported the same constants, which encoded a misreading of the source's status vocabulary, so the error was common-mode and the gate could not see it. We give the mechanism, an object-level ledger reconciling every figure, and three checks that go back to the source's documentation, measured on the defective code and on its correction. We then checked that correction against each object's phase history, held in a source file the pipeline never read. The correction was also wrong: 42 of its 261 disagreements are artefacts, and none of our three checks flagged them. Finally, in a controlled replication with three pinned models and tools disabled, 72 of 75 paths generated on request as independent checks computed the defective count, 29 of 30 even when the prompt carried the source's own definitions of the codes. The evidence is one pipeline and one defect family. Within it, redundancy verified implementation, and the errors that reached publication were errors of meaning.
cs.SE / 19 / 2609.37645
Beyond Productivity: Measuring Developers' Cognitive Load During GenAI-Supported Software Development
Abstract
Generative AI (GenAI) is changing software development workflows and how developers work. Industry evaluations of GenAI adoption often monitor productivity gains, usage, and output quality, but limited attention is paid to the interaction experience and cognitive load of the actual adopters and drivers of GenAI technology - the software developers. Understanding whether GenAI changes or shifts developers' cognitive demands during everyday development is important for a developer-centered evaluation of GenAI-supported software development. It can inform organizations in designing and evaluating effective AI-supported workflows. In this work, we study how GenAI use and task context relate to professional developers' perceived cognitive load and whether wearable-derived physiological characteristics provide additional information beyond this context. In a four-day industrial field study at two SAP sites, 21 developers documented their tasks, task duration, GenAI use, and perceived cognitive load while wearing an EmbracePlus wristband. The results show that perceived cognitive load is associated with both GenAI use and task context, while physiological measures provide only limited additional information. These findings suggest that developers' perceived cognitive load during GenAI-supported software development should be evaluated in relation to the concrete work context, with wearable physiological data used as complementary rather than standalone information.
cs.SE / 20 / 2609.37864
AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems
Abstract
Agent harness bugs exhibit unique characteristics and remain challenging for state-of-the-art software agents to repair. Progress in this area is further hindered by existing benchmarks, which contain only a small and fixed number of executable harness bugs while requiring hundreds of human hours to construct. This work presents AgentBug-Smith, an automated harness bug reproduction approach that continuously discovers and reproduces real-world harness bugs from open-source agentic systems. Across different backbone LLMs, AgentBug-Smith consistently outperforms existing bug reproduction techniques designed for general software, achieving 10.67% - 27.56% higher success rates of reproducing harness bugs. By applying AgentBug-Smith to open-source agentic systems in the wild, we construct Live-Harness-Bench, a live and extensible benchmark that currently contains 200 reproducible harness bugs. We further demonstrate the utility of Live-Harness-Bench through two downstream applications. First, we use Live-Harness-Bench as the evaluation benchmark to systematically evaluate state-of-the-art software agents, revealing their limited capabilities in repairing real-world harness bugs. Second, we use Live-Harness-Bench as a knowledge base of real-world harness bug fixes, from which reusable repair skills can be distilled to improve existing software agents, increasing their harness-bug repair rates by 6.32%. Together, AgentBug-Smith and Live-Harness-Bench establish a scalable foundation for continuously evaluating and improving software agents on harness bug repair, turning real-world agent failures into executable evaluation instances and reusable knowledge for harness improvement, thus contributing to the ultimate goal of recursively self-improving agents.
cs.SE / 21 / 2609.37985
Merged, Not Measured: An Empirical Study of Performance Issues Fixed by Coding Agents
Abstract
Coding agents open pull requests (PRs) that claim to speed up software, but studies of human performance fixes say little about how maintainers respond to such a fix or whether its claim holds. From the 71,677 agent PRs of AIDev v4, a text filter and codebook coding by language models and by the authors select 1,262 performance issues fixed by six agents in 582 repositories. We code each issue and its tests and re-execute 23 rejected and 30 merged fixes. (1) 57% of closed fixes are merged, 61% of rejections give no stated reason, and only 6 of the 23 re-executed rejected claims held under our three-run pilot on mostly agent-built workloads. (2) Acceptance rises with the agent's track record in the repository (31-37% to 70%) and with the repository's pre-opening merge rate on its other agent PRs (33% to 84%). Merged fixes delete a larger share of the lines they change (0.26 versus 0.15), a difference that holds within agent and within repository, with no such difference detected in the coded content, description, tests or measurements. (3) Repeated computation and redundant data processing cause 44% of the issues, and 46% of fixes are architectural-level. (4) Agents change tests in 37% of fixes and 11% carry a performance test or benchmark; of the 30 merged fixes, 18 met our delivery criterion, 3 fell short of the claim, 9 showed no significant gain or regressed, and 14 change behavior on untested inputs. The outcome tracks the repository's history with the agent rather than the coded content of the fix, and a merge does not show that the fix delivers what it claims.
硬件架构 (cs.AR)
7
cs.AR / 1 / 2609.36111
Automated Pre-Silicon Verification of High-Speed DDR5 and LPDDR5/6 Memory Controllers: Closed-Loop Timing, Mode Register, and PHY Synchronization in UVM
Abstract
External memory interfaces (such as LPDDR5/4 and DDR5) are essential components of modern mobile, cloud, and enterprise computing systems. While memory manufacturers focus on physical DRAM die development, the vast majority of semiconductor firms design custom application ASICs that require a dedicated Memory Controller to interface with these standardized external memories. Consequently, pre-silicon design verification of the Memory Controller RTL -- acting as the Device Under Test (DUT) against a third-party DRAM Verification IP (VIP) -- is a ubiquitous and critical challenge across the global semiconductor industry. This article presents an automated, pre-silicon configuration and closed-loop initialization framework for Memory Controller verification. The proposed solution parses JEDEC timing and configuration parameters directly from the DRAM VIP's database files (such as Denali SOMA files) to configure the Memory Controller DUT's registers, while a custom Mode Register Register Abstraction Layer (MR-RAL) tracks volatile DRAM VIP states in real-time. Furthermore, a dynamic, re-compilation-free PHY initialization flow randomizes interface parameters directly in the testbench, executes on-the-fly configuration generation via system calls, and parses the resulting register write sequences at runtime. By integrating these automated methodologies into pre-silicon verification flows, engineers can eliminate setup overhead, prevent false protocol violations, and enable comprehensive randomized testing of complex PHY and Memory Controller configurations.
cs.AR / 2 / 2609.36634
Lossless Compression of Lookup Tables for Hardware Applications
Abstract
Large lookup tables are widely used in hardware to store constant-valued arrays for applications ranging from elementary mathematical operations, such as constant-coefficient multiplication and nonlinear function evaluation, to emerging machine learning models, including table-based neural networks (NNs) and Kolmogorov-Arnold networks (KANs). However, storing extensive tables of constant values can lead to excessive hardware costs in resource-constrained edge devices such as FPGAs. In this paper, we propose CompressedLUT, a lossless compression scheme and its decoder hardware architecture for the efficient storage and retrieval of arbitrary data in hardware. Our method combines decomposition, self-similarities, higher-bit compression, and multilevel compression techniques to maximize table size savings without accuracy loss. Its hardware decoder primarily uses addition, arithmetic right shift, and several small lookup tables, ensuring low area and high throughput. We evaluated CompressedLUT on FPGAs by implementing multiple nonlinear functions, constant-coefficient multipliers (CCMs), and KANs at 12-bit resolution. CompressedLUT is available as an open-source tool.
cs.AR / 3 / 2609.37021
Low-level optimizations in high-level HDLs: Is there a benefit?
Abstract
This paper explores the applicability of functional programming to the design of Application-specific Integrated Circuits (ASICs). We investigate the impact of designing ASICs using high-level, abstract Hardware Description Language (HDL) features versus employing low-level optimizations on the area of the synthesized circuits. The aim is to determine whether using low-level optimizations is beneficial and, if so, whether it is worth the added implementation effort. To carry out the investigation, we implement an unsigned bit-serial multiply-accumulate (MAC) unit in 16 different configurations using the functional HDL Clash. We make use of both low-level bit-manipulation techniques as well as Clash's high-level constructs for circuit design. The experimental evaluation shows that some high-level constructs of Clash have negligible influence on the resulting circuit size, suggesting that using the full power of functional programming is a viable approach to hardware design. To evaluate the impact of the used HDL itself, we also implemented versions of the MAC in Verilog. The experiments clearly show that there seems to be an inherent overhead in using Clash compared to Verilog code written by a seasoned engineer.
cs.AR / 4 / 2609.37399
MEDEM: Multi-Engine DL Accelerator Design Methodology
Abstract
Multi-engine deep learning (DL) accelerators are becoming increasingly prevalent as they address the heterogeneity and growing complexity of modern DL workloads. To efficiently process diverse DL workloads, these accelerators must incorporate combinations of engines with complementary capabilities to match the distinct computational characteristics of these workloads' heterogeneous kernels. However, existing multi-engine DL accelerator design approaches lack a systematic methodology, leaving fundamental questions unresolved. These include how to co-design engines for workloads with diverse computational characteristics and which engine combinations minimize aggregate execution costs (such as time or energy) across such workloads. Addressing these questions requires efficient exploration of exponentially large design spaces. To address these questions systematically, this work proposes Multi-Engine DL Accelerator Design Methodology (MEDEM). MEDEM defines generic engine abstractions, co-designs candidate instances (engines), and selects a combination of co-designed engines to minimize aggregate execution cost given diverse DL workloads and a resource budget. MEDEM encompasses a set of design strategies that efficiently navigate the exponentially large design spaces of engine co-design and combination selection, identifying highly optimized multi-engine accelerators. A comprehensive evaluation demonstrates that MEDEM identifies accelerators that outperform state-of-the-art designs, delivering geometric-mean improvements of up to 4.84x in energy-delay product (EDP) and 1.59x in throughput. The improvements are achieved using different resource budgets, demonstrating MEDEM's scalability, and using 51 single- and multi-model DL workloads, demonstrating its generalizability.
cs.AR / 5 / 2609.37711
Zephyr: An Efficient Audio Denoising System Using Spiking Neural Networks Enabled With A Sparsity-Aware Flexible FPGA PE Array
Abstract
In this work we look to neuromorphic computing to solve the power consumption problem that audio denoising neural networks face on edge devices like smartphones, wireless headphones and hearing aids. Spiking neural networks (SNNs) have the potential to solve this problem due to their high activation sparsity and low complexity, however many SOTA SNNs require hardware that supports a mixture of operations to be able to fully perform inference. To solve this problem, we convert SOTA audio denoising neural network Spiking-FullSubNet to a hardware friendly version showing that via QAT and activation function simplification we can achieve $\approx28\times$ improvement in power consumption to 52.9nJ per 32ms audio frame when calculated for custom digital hardware in a 45nm process node. We then propose a digital circuit which by means of a sparsity-aware flexible PE array can perform inference of the heterogeneous compute load of Spiking-FullSubNet, and validate this circuit on a PYNQ-Z1 FPGA achieving a real-time factor of 0.727 at 100MHz.
cs.AR / 6 / 2609.36134
Hardware-Aware Functional Kolmogorov-Arnold Networks for Efficient Medical Image Enhancement and Segmentation
Abstract
Functional Kolmogorov-Arnold Networks (FunKAN) achieve state-of-the-art accuracy on MRI Gibbs artifact removal and anatomical segmentation, but their 11.6 M parameters and 8.7 GFLOPs are too large for edge medical devices. We present FunKANLite, a two-stage, hardware-aware compression of FunKAN for point-of-care use. FunKANLite-TR reduces the spatial prior and replaces the ResBlock offset predictor with a depthwise-separable block. It has 1.9x fewer parameters than FunKAN and no loss in accuracy. We then distill FunKANLite-TR into FunKANLite-ST, which lowers the Hermite basis rank, factorizes the spatial prior into a low-rank form, and halves the filter widths. FunKANLite-ST has 5.6x fewer parameters and 3.7x fewer GFLOPs than FunKAN. It stays within 1.4 percentage points IoU of FunKAN on BUSI, GlaS, and CVC-ClinicDB, and reaches 33.95 dB PSNR on IXI. On an NVIDIA Jetson Orin Nano and a Raspberry Pi 5, FunKANLite-ST reduces energy per inference by up to 68% and raises throughput by 2.9x.
cs.AR / 7 / 2609.37137
Mixed-Precision Computing for Scientific Discovery: Formats, Co-Design, and Responsible Approximation
Abstract
Reduced and mixed precision have moved from a niche optimization to a central design axis in scientific computing and engineering, driven by energy constraints, heterogeneous accelerators, and the convergence of simulation and machine learning. This paper organizes the landscape around seven coupled themes---number formats, floating-point emulation, emerging architectures, hardware/software co-design, relation to other approximations, software design, and precision as a multilevel resource ---and, for each theme, synthesizes the state of the art, future directions, and open questions. We emphasize \emph{energy per trusted solution} as the core objective, and we frame \say{recklessly responsible} computing as a pragmatic doctrine: exploit low precision aggressively, but with systematic detection, escalation, and certification pathways.
密码学与安全 (cs.CR)
38
cs.CR / 1 / 2609.36039
Multi-Class, Multi-Tier Network Intrusion Detection: A Comprehensive and Reproducible Benchmark
Abstract
Machine learning (ML) and deep learning (DL) have dominated Intrusion Detection System (IDS) research in recent years. Unfortunately, many existing studies have produced inflated results and unreliable benchmarks due to critical oversights and mistakes in the ML and DL pipeline, from data collection and labeling to feature engineering and model training and evaluation. CIC-IDS2017 is a standard benchmark for network intrusion detection. Still, many published results on this dataset are difficult to compare due to labeling errors, inconsistent flow extraction, potential leakage, and performance evaluation metrics dominated by benign traffic. In this paper, we present a comprehensive benchmark with corrected PCAP-level labeling and a complete evaluation pipeline with diverse ML models. We evaluate eleven tabular classifiers at three nested levels: binary attack detection, nine-class attack-family attribution, and fifteen-class fine-grained classification. A soft-voting ensemble of Random Forest, XGBoost, and LightGBM obtains the best fine-tier macro-F1 of 0.955, with coarse and binary macro-F1 scores of 0.980 and 0.999, respectively. We further conducted a feature selection study based on an analysis of feature importance. This comprehensive benchmark pipeline is configurable and open-source, enabling new feature extraction and model plugins for new datasets. Future work should use this pipeline as a reference point for richer features, rare-class analysis, and model generalization towards new datasets and attack classes.
cs.CR / 2 / 2609.36153
Privacy-Friendly Cohort Determination: Sealed, CSP-Independent In-Browser ML Inference of Professional Segments for Identity-Less Advertising
Abstract
B2B advertising targets a viewer's professional attributes (employer size and industry, function, seniority) and has obtained them by matching identities across sites. Safari and Firefox block third-party cookies, Google retired the Privacy Sandbox cohort APIs in 2025, and reverse-IP firmographics decay under remote work. We present SIF (Sealed Inference Frame), which infers coarse professional cohorts on the device and emits only a locally differentially private, taxonomy-coded label into the OpenRTB bid stream, with no cross-site identifier. It rests on a property of the web platform we make precise: a navigated cross-origin iframe is the only way third-party code obtains a policy it controls, so inference runs in WebAssembly even where the publisher's CSP forbids it, and a nested worker served with default-src 'none' gives the model no network. Even a malicious model leaks at most about 5 bits per site per week. Labels pass through a memoised k-ary randomised response keyed to the publisher's first-party identifier, which gives $\varepsilon$-local differential privacy, defeats averaging, and links requests no better than the identifier already sent. An org-conditional k-anonymity rule suppresses cells, more strictly on corporate networks than at home. Cohorts ride OpenRTB user.data in a LinkedIn-aligned taxonomy, and attribution uses LinkedIn's click-scoped li_fat_id without bridging identities. We report a crawl of CSP deployment on 7,969 top sites and 431 B2B publishers, Heavy-Ad budgets, closed-form privacy-utility trade-offs, a re-identification simulation, and an assessment of which attributes are predictable at all: company type and size are, seniority largely is not. On-device is a design property, not a consent exemption.
cs.CR / 3 / 2609.36167
Adversarial Debiasing of Machine Learning Models for Enhanced Network Security against DDoS Attacks
Abstract
Distributed Denial of Service attacks are a growing threat to network infrastructure, and new techniques, including the use of generative AI, make them harder to detect. Traditional detection systems, such as rule based firewalls, often fail to identify these evolving attack patterns. In this study, we propose a new method for detecting DDoS attacks by combining synthetic data generation using Generative Adversarial Networks with a Random Forest classifier. The GAN generated data showed 80.3 percent cosine similarity to real traffic, which helped the model learn underlying traffic patterns more effectively. To address imbalances in the data, especially in packet related features, we applied adversarial debiasing. This reduced the model's sensitivity to skewed distributions in variables such as forward and backward packet counts and total byte lengths. Our results show that models trained on a mix of synthetic and real data achieved significantly better performance: 99.98 percent accuracy on benchmark data and a 22.60 percent improvement when tested on previously unseen synthetic traffic. This suggests that the method can generalize well across different traffic scenarios and adapt quickly to new types of attacks. The proposed approach not only improves DDoS detection but also provides a scalable foundation for security models that account for bias and benefit from data augmentation. Our findings show that combining GANs with adversarial debiasing can lead to more robust and effective DDoS mitigation, supporting the further development of machine learning based cyber security.
cs.CR / 4 / 2609.36330
DecoyTrace: Toxic Decoys for Active Defense in Decentralized Federated Learning
Abstract
Decentralized Federated Learning (DFL) eliminates the central aggregation server, reducing the single point of observation that traditional defenses against attacks rely on. As a result, peer-to-peer networks become exposed to malicious updates containing backdoors or semantic poisoning, since such updates can remain close to benign ones in the parameter space while behaving very differently. This may evade defenses based on passive parameter inspection. However, existing deception-based defenses have mainly been designed for centralized FL and do not jointly address local observation, poisoning propagation, source attribution, and containment in strictly serverless DFL. To address these limitations, this paper presents DecoyTrace, a proactive cyber deception-based defense for strictly serverless DFL environments. DecoyTrace deploys a mobile DecoyNode that generates decoy challenges using chaotic maps, disseminates a dual model (clean vs. decoy) based on neighbor trust, and evaluates them using three-state semantic metrics. Upon confirmation, a distributed protocol isolates the source and performs a model reset or recovery to preserve training progress. Evaluated across sixty configurations on the NEBULA platform (five datasets, three topologies, and four attack/defense scenarios), DecoyTrace systematically restores lost utility. The F1-score remains within 0.03 of the baseline on MNIST/FashionMNIST (mitigating drops of up to 0.37), matches or exceeds the baseline on EMNIST and CIFAR-100, and remains between 0.05 and 0.10 below the baseline on CIFAR-10, the most visually complex convolutional scenario evaluated. Furthermore, containment reduces CPU and network usage by up to two-thirds. These results demonstrate the feasibility of unifying deception, identification, and containment in DFL, while also identifying its limitations in complex tasks and multi-attractor threat models.
cs.CR / 5 / 2609.36331
Calibrating One-Round Membership Inference with Neighbors
Abstract
The state-of-the-art Membership Inference (MI) methods calibrate their signal separately for each example using reference models, auxiliary models trained to exclude the target. This paradigm scales poorly to modern large models, however, whose training is too expensive to replicate. This has motivated one-round settings, where only a single trained model is available; but without reference models the per-example calibration that drives the strongest attacks can no longer be estimated, leaving the membership signal weak. We ask whether neighbors of the target point can recover this calibration without training any additional model. Our key observation is that reference models serve only to reveal how an example behaves under models not trained on it, and that querying the target model on nearby samples yields the same information. We propose two complementary ways to obtain such neighbors, and show that querying them against an early training checkpoint further sharpens the signal. We evaluate across three image classification datasets and three training setups, showing that neighbors yield strong membership signals and competitive attack performance at no additional training cost.
cs.CR / 6 / 2609.36372
TTMark: Pairwise Distortion-Free Watermarking Beyond Single-Token Entropy
Abstract
Distortion-free watermarking enables reliable attribution of machine-generated text while preserving output distribution. However, existing methods operate independently on each generated token, making their detection capability fundamentally constrained by the entropy of the next-token distribution. We present Tandem Token WaterMark (TTMARK), a general pairwise watermarking framework that extends distortion-free watermarking from individual tokens to adjacent token pairs. By watermarking the joint distribution of consecutive tokens, TTMARK enlarges the effective watermarking alphabet from V to $V^2$, allowing the detector to exploit both token entropy and conditional entropy while preserving distortion-freeness over the joint distribution. We further introduce a branch-isolating concatenated tandem generation algorithm that efficiently constructs the joint distribution in a single forward pass. Theoretically, we show that pairwise watermarking achieves better expected detection strength in low-entropy regimes. Extensive experiments across multiple language models, datasets, and three representative distortion-free watermarking schemes demonstrate that TTMARK consistently improves detectability without degrading generation quality, while also improving robustness to edits and substantially enhancing localized watermark detection.
cs.CR / 7 / 2609.36373
Audience-Bound Persistent Memory: Authorization Across the Memory Lifecycle
Abstract
A personal language agent that acts for its owner across private and shared conversations can learn a fact from one audience and later place it in the context it assembles for another. We study authorization before context across the whole memory lifecycle. Each memory item carries the audience present when it was recorded; derived items are partitioned by audience, receive the intersection of their sources' audiences, or are suppressed; an audience widens only by an explicit, object-specific grant; and an item enters a model attempt only when every current viewer belongs to one of its authorized audiences, with unresolved viewers failing closed to public-only. Under explicit identity, provenance and complete-mediation assumptions, this admission is sound and policy-complete on the exact assembled context, enforced by exclusion rather than by model behavior. We realize it in two independently persisted reference architectures, a flat store and a relationship graph, and, descriptively, in a native agent-memory runtime. In a prospectively frozen confirmation over 10,000 multi-party histories, no forbidden item entered any architecture's context, whereas unscoped retrieval exposed forbidden items in 82% of its contexts. Entitled recall matched policy-equivalent baselines exactly and exceeded unscoped retrieval by 0.30 Recall@5, with a Holm-confirmed advantage that grows with distractors. No architecture produced a wrong-principal substitution, but unscoped substitutions were too rare to establish the prespecified joint decision.
cs.CR / 8 / 2609.36376
Quantization Enables Private Dense Retrieval against Malicious Service Providers
Abstract
Dense retrieval, the key component of Retrieval Augmented Generation (RAG), retrieves the most relevant documents by comparing dense vector representations of queries and passages from a large corpus. In privacy-sensitive applications, the server observes the query and controls which evidence is returned, creating both confidentiality and integrity risks. We formulate private dense retrieval as providing query privacy and retrieval integrity against a malicious server, and develop a two-round cryptographic protocol that provides both guarantees. Our protocol reduces private and verifiable retrieval to multiplication of a committed matrix by an encrypted vector and uses low-bit quantization to make this computation practical. We evaluate the resulting trade-off between cryptographic cost, retrieval quality, and downstream RAG accuracy across six embedding models, four language models, and corpora of up to 2.68 million passages. Our results show that, with a clipped quantizer, three-bit quantization largely preserves retrieval quality and downstream accuracy, while a private query over a corpus the size of a clinical reference requires one to three minutes of server time. These results suggest that private dense retrieval is already practical for moderately sized, privacy-sensitive corpora when minute-scale latency is acceptable.
cs.CR / 9 / 2609.36494
Know the Normal, Track the Attack: Context-Grounded and Stateful LLM Investigation over System Provenance
Abstract
Provenance-based intrusion detection systems (PIDSs) identify suspicious activity in audit streams, but their outputs remain difficult to turn into coherent attack narratives. Direct LLM analyses of local anomalous subgraphs lack deployment-specific normal-behavior knowledge and validated attack state across evidence fragments. This can cause unsupported attack interpretations of routine activities and incorrect attribution of temporally dispersed evidence to attack stages. We present ANCHOR, an investigation-oriented provenance system that combines evidence curation with context-grounded LLM reasoning. It calibrates anomaly judgments by relation type and links anomalous windows through rare relation-role patterns. The resulting evidence queues preserve causal structure, temporal boundaries, and cross-window continuity. The investigator interprets process-centered evidence using two complementary forms of context. Deployment Context combines environment-specific interaction and object baselines with high-risk security knowledge. Case Context uses a confidence-gated Attack-Tracking Cache to maintain investigation state across windows. Correlating current evidence with high-confidence prior findings, ANCHOR incrementally reconstructs attack narratives organized by kill-chain stages. We evaluate ANCHOR on six DARPA Transparent Computing E3/E5 datasets across three operating systems. Controlled evidence-level and end-to-end comparisons show improved overall IoC recovery and attack-stage attribution over state-of-the-art provenance-based baselines. These gains persist under a fixed LLM backbone in our evaluation. ANCHOR processes a full audit day at dollar-level API cost, supporting practical, context-grounded investigation across windows.
cs.CR / 10 / 2609.36570
CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering
Abstract
Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that suppresses this behavior inside the model. Per model, a five-step recipe fits a residual-stream direction from paired episodes differing only in whether an embedded instruction is followed, and retains it only if it passes pre-specified causal and capability gates. At deployment, the direction is subtracted from every tool-result token during prefill. The edit is always on--there is no detection decision to evade--and requires no fine-tuning, auxiliary model, or added tokens, only white-box serving and tool-result span boundaries. Across five open-weights models (8B-106B, five vendor lineages), held-out attack success falls from 0.21-1.00 undefended to 0.00-0.17 defended, and AgentDojo compromise rate from 0.10-0.49 to 0.006-0.079, at 93-100% typography-normalized benign utility, with larger task-dependent costs when reasoning over steered content. A benchmark-level adaptive attacker reaching 0.67-0.73 undefended is held to roughly a quarter of that on the two most deeply evaluated models. Among the defenses we measured on capable models, those achieving lower compromise rates either lost 22-89% of benign utility or fine-tuned the served weights. White-box gradient attacks through the deployed vector compromise at most 2 of 52 episodes, and none of 2,052 replayed human red-team attacks succeeds. CounterSteer largely neutralizes instructional takeover: a black-box framing search cracks 3 of 18 development samples. Parameter manipulation--attacker-chosen arguments in otherwise legitimate calls--is only partially resisted (13 of 18); the decision becomes linearly readable at argument emission but not at the examined pre-generation sites, and is not removed by the tested prefill- or decode-time steering, motivating argument-provenance controls.
cs.CR / 11 / 2609.36573
CyberPersistBench: Evaluating LLM-Based Cyber Attackers on Installation and Persistence
Abstract
While LLM-based attackers exhibit growing proficiency in vulnerability exploitation, most existing cybersecurity benchmarks suffer from single-stage truncation, prematurely terminating evaluation upon initial access. In practice, initial footholds are exceptionally fragile across operational disruptions such as service restarts and host reboots. Whether LLM-based attackers can establish and maintain durable footholds beyond initial compromise remains a central blind spot in cybersecurity evaluation. To bridge this gap, we introduce CyberPersistBench, the first benchmark dedicated to post-compromise installation and persistence. Decoupled from upfront exploitation, CyberPersistBench frames persistence as an adversarial survival task in which agents use native host mechanisms to maintain footholds across staged system disruptions. Deterministic checks support a six-level scoring method (L1--L6) spanning installation and persistence. The benchmark comprises 203 core tasks across seven categories, augmented by multi-host and active defense extensions. Empirical evaluations across five frontier agents show that autonomous persistence remains limited (27.6%--44.8%) and drops further on defense-enabled tasks (5.5%--13.3%); nonetheless, these results reveal an emerging cyberattack risk, underscoring the necessity of benchmarking post-compromise persistence. CyberPersistBench thus establishes a foundational benchmark for post-compromise installation and persistence, delineating the operational boundaries of autonomous cyber agents.
cs.CR / 12 / 2609.36731
Efficient Linkage-Based Compartmentalization on CHERI
Abstract
We present an efficient linkage-based model for in-process compartmentalization built on CHERI memory safety, which enables fine-grained compartmentalization of the entire UNIX user-space, scaling to 10K+ compartments on desktop systems. The model's "push-button" compartmentalization along existing library boundaries regularly hosts 500+ compartments per process for large applications such as Chromium, far exceeding the number of concurrently available protection domains supported by other mechanisms (e.g., up to 16 for Intel MPK). Custom policies can further subdivide libraries. Of the thousands of C/C++ programs tested, only the V8 JavaScript engine required source-level adaptation (<300 lines of changed code concerning garbage collection and JIT compilation). We implement the model for CHERI-extended versions of Armv8-A and RISC-V through support in the compiler toolchain and operating system. Case studies illustrate the smooth delegation of memory between compartments, compartment-aware debugging and visualization, as well as extensibility to a complex managed language runtime, demonstrating the benefits of our single-address-space model. We evaluate using multiple processors, including Arm's superscalar Morello and, notably, the first commercial CHERI-enabled RISC-V application core---Codasip's in-order dual-issue X730.
cs.CR / 13 / 2609.36732
Deep Learning Latency Attacks and Defenses: A Cross-Domain Survey of Availability Threats
Abstract
Adversarial machine learning has focused mainly on integrity, but availability is an increasingly consequential complement. Latency attacks (also energy-latency attacks) increase inference-time work, energy, or response time, causing deadline misses, throughput collapse, or resource exhaustion in vehicle controllers, interactive services, or battery-powered sensors, sometimes while preserving the nominal prediction. This survey unifies a fragmented literature spanning perception pipelines (including physical attacks on autonomous-driving detection and tracking), input-adaptive neural inference (sponge examples, dynamic networks), and autoregressive and agentic systems (output-length, verbose-image, and reasoning denial-of-service attacks on LLMs, VLMs, mixture-of-experts models, and tool-using agents). We organize attacks by exploited computational bottleneck rather than formulation, separating what makes a computation expensive from how the attacker triggers it; the delivery channel (input, prompt or retrieved content, message, poisoning, or weight tampering) is an orthogonal attribute. Many attacks share one mechanism, intermediate-work amplification, motivating a work-budget defense abstraction; we distinguish caps on the work entering an expensive stage from caps on the results leaving it. We further analyze when a model-level cost increase becomes a system-level availability failure, which depends on critical-path share, slack, existing ceilings, accumulation, resource sharing, and fallback policy, not on the amplification factor alone. We also provide a threat-model taxonomy, consolidated quantitative comparisons, a defense review by control mechanism, and open challenges such as standardized evaluation, physical realizability, and whole-system availability. Companion website: https://github.com/guzonghua/awesome-latency-attacks.
cs.CR / 14 / 2609.36817
pikit: A Composable Toolkit for Indirect Prompt Injection Research and Evaluation
Abstract
Indirect prompt injection embeds malicious instructions within external content retrieved by LLM-based agents, altering target behavior without user authorization. We introduce pikit, a research toolkit designed to systematically evaluate these threats across three core dimensions: attacks (13 methods), channels (16 carriers across text and file modes), and defenses (9 prevention strategies and 3 offline detection baselines). Built on a decorator-based registry, pikit enables seamless extension of custom components without modifying core code, while a unified craft() API composes arbitrary attacks and channels in a single call. We evaluated the toolkit on the pi coding agent powered by an anonymized LLM in a production-like environment. Benchmarking 9 prevention strategies against high-risk attacks yields a 71.8\% relative reduction in attack success rate, with few\_shot\_warning and instruction\_hierarchy providing the strongest protection. Offline detection baselines achieve perfect precision but low recall, demonstrating that heuristic detectors complement rather than replace prompt-level defenses. To ensure reproducibility, each run automatically logs full prompts, agent event traces, session transcripts, and verdict records. Our code is available at https://github.com/Tencent/AI-Infra-Guard/tree/main/Research/pikit.
cs.CR / 15 / 2609.36849
Does the Unsafe Gradient Survive a Conversation? On the Fragility of Gradient-Based Jailbreak Detection in Multi-Turn Dialogue
Abstract
Safety-aligned language models are commonly deployed as multi-turn assistants, which lets adversaries spread unsafe intent across several user turns instead of a single prompt. Gradient-based jailbreak detectors such as GradSafe were developed for single prompts: they score an input by the alignment between its induced gradient and a fixed unsafe reference direction, and their effectiveness in multi-turn dialogue remains unclear. We conduct a controlled evaluation of gradient-based jailbreak detection in multi-turn settings. We extend GradSafe with a Context Window Scanner that applies the detector to fixed-size windows of user turns and uses the maximum window score as the conversation-level score. We evaluate different window sizes, attack families, benign conversation distributions, and target models. The results differ sharply between synthetic and realistic benign settings. Against synthetic benign conversations, the detector achieves an ROC-AUC of 0.98 on human-authored multi-turn jailbreaks. On WildChat benign conversations, ROC-AUC drops to 0.76, and a threshold calibrated on synthetic data flags more than 90% of benign conversations as unsafe. Under realistic benign distributions, single-turn windows give the highest separability, whereas longer windows and accumulated contexts reduce performance. The detector is also sensitive to the attack-generation method and target model: successful Crescendo attacks receive scores comparable to or lower than benign conversations, and Qwen2.5-7B-Instruct yields near-random separability with a different optimal window size. These findings show that gradient-based signals can support multi-turn jailbreak detection, but reliable deployment requires calibration on realistic benign conversations, short-window scoring, length-aware thresholds, and evaluation across attack types and model architectures.
cs.CR / 16 / 2609.36862
Safer Content or Firmer Refusals? A Hybrid Perturbation Defense for Alignment under Harmful Fine-tuning
Abstract
Fine-tuning-as-a-service lets users adapt a safety-aligned language model to their own data, but it also creates a harmful fine-tuning attack surface: a small amount of harmful data mixed into an otherwise benign fine-tuning set can degrade the model's alignment. Two recent alignment-stage defenses address this problem at different levels of the model. Vaccine improves the robustness of hidden embeddings to the representation shifts induced by harmful fine-tuning, whereas Booster simulates harmful weight updates and attenuates their effect during alignment. We investigate whether these mechanisms are complementary and propose VaccineBooster, a single alignment procedure that combines embedding perturbation and weight-level gradient attenuation within each training step. On Llama-2-7B aligned with BeaverTails and then attacked through poisoned fine-tuning, VaccineBooster achieves the lowest OpenAI moderation score among the compared defenses, 0.315, while a Booster-Only variant retains the highest post-attack refusal rate, 50%. Together with ablations over the embedding-perturbation and gradient-attenuation strengths, these results indicate a trade-off: embedding perturbation primarily reduces flagged harmful content, whereas gradient attenuation primarily preserves explicit refusal behavior. Because our evaluation uses ten prompts and a single unseeded run per configuration, we report this trade-off as an observed pattern rather than a statistically resolved effect. These results provide practical guidance for prioritizing content safety or refusal retention when aligned models are exposed to untrusted fine-tuning.
cs.CR / 17 / 2609.36879
SKILLLITE: Evidence-Guided Malicious Skill Auditing with Compact LLMs
Abstract
As LLM-based agents perform increasingly complex tasks, Agent Skills have emerged as a flexible mechanism for extending their capabilities. An Agent Skill packages task-specific instructions with executable components and auxiliary resources to provide specialized functionalities. However, the growing adoption of third-party Skills introduces a new supply-chain attack surface. Malicious Skills can embed harmful behaviors that abuse agent privileges and compromise the agent execution environment or accessible resources. Although recent LLM-based malicious Skill auditing approaches have achieved promising performance, they often rely on capable commercial LLMs. How to achieve effective auditing with compact, locally deployable LLMs in security-sensitive and resource-constrained settings remains largely unexplored. Our investigation reveals that compact LLMs struggle to identify malicious behaviors hidden in complex Skill packages. This difficulty arises from both the implicit nature of such behaviors and the limited reasoning capacity of compact LLMs. To address these challenges, we propose SKILLLITE, an evidence-guided agentic framework for malicious Skill detection. SKILLLITE effectively extracts security-relevant behaviors and infers the intended functionality from complex Skill packages. It then employs a compact LLM to assess the maliciousness of the Skill based on the observed behaviors and their functional context. Experiments show that SKILLLITE improves malicious Skill detection across different compact LLM backbones and outperforms existing representative auditing baselines. Its effectiveness generalizes to behaviorally confirmed in-the-wild malicious Skills. Meanwhile, SKILLLITE maintains a low inference latency, supporting its practical deployment.
cs.CR / 18 / 2609.37006
One Pipeline Does Not Fit All: TAILOR, a Type- and State-Aware Framework for CVE Reproduction
Abstract
Growing vulnerability disclosure and widespread software reuse increase security teams' need for reproducible evidence to diagnose vulnerabilities, validate patches, and build regression tests. Producing such evidence at scale requires automated end-to-end CVE reproduction. Existing methods typically process different CVEs through a uniform pipeline, but differences in runtime form, trigger interfaces, and prerequisite state impose different execution requirements on individual stages, making fixed workflows difficult to adapt to diverse reproduction needs. To address this problem, we present TAILOR, a type- and state-aware multi-agent framework specialized for complex vulnerability reproduction. TAILOR converts static vulnerability information into auditable reproduction evidence and packages reconstructed environments and trigger evidence into reproduction artifacts. Its first-level type-aware mechanism adaptively matches each vulnerability to an execution path. Within the Web path, its second-level state-aware mechanism constructs the required prerequisite state before exploitation, decouples prerequisite-state construction from core vulnerability triggering, and shares execution constraints across exploitation and verification. We construct a dataset of 200 CVEs with an emphasis on cases with complex execution requirements. TAILOR successfully reproduces 59.24\% of Web vulnerabilities and 44.19\% of traditional vulnerabilities. Further ablation experiments show that the two control levels respectively mitigate execution-path mismatch and missing Web prerequisite state. Overall, TAILOR broadens the coverage of automated CVE reproduction and provides auditable evidence for vulnerability diagnosis and defense.
cs.CR / 19 / 2609.37011
OPFL: Optimistic Verification of Federated Learning via Empirical Boundary
Abstract
Federated learning enables multiple clients to collaboratively train models without sharing their private data. However, the lack of visibility into local training makes it difficult to verify whether clients follow the prescribed training procedure or submit malicious updates, such as model poisoning. A natural approach is to replay client training for verification. However, privacy-preserving replay produces numerical results that cannot be directly matched with local client execution because the two run in different environments. We present OPFL, an optimistic verification framework for privacy-preserving federated learning. To protect data privacy, OPFL performs replay inside secure multi-party computation (MPC). Although gradients computed on MPC and local GPUs are not bitwise identical, we observe that their absolute differences are stable and bounded. OPFL therefore calibrates an empirical boundary offline and uses it to distinguish benign numerical deviations from malicious manipulation. To reduce the cost of expensive MPC replay, OPFL adopts optimistic verification by post auditing only sampled training steps. Experiments on LeNet, BERT, and Qwen show that the boundary generalizes across datasets, input lengths, and GPUs, while achieving $0$\% ASR against model poisoning and PGD-based attacks. On a LeNet workload, at $p=0.01$, OPFL is approximately $98.6\times$ faster than full MPC-based FL and $625.5\times$ faster than ZK-based approach.
cs.CR / 20 / 2609.37196
ToolFence: Fine-Grained Authorization for Secure Tool-Using LLM Agents
Abstract
Tool-using LLM agents remain vulnerable to indirect prompt injection because trusted instructions and untrusted observations share one context, allowing malicious content to steer consequential input-filtering defenses. Multi-path consensus defenses still leave a high attack success rate because they examine content or aggregated outputs rather than authorizing effects, especially for the within-tool attack, which preserves the intended tool but manipulates its arguments. Data-Flow Control such as CaMeL provides stronger guarantees, but incurs substantial time latency that limits practical deployment. We introduce ToolFence, which compiles a typed authorization blueprint before execution, enforces it through a deterministic monitor, and when the blueprint is incomplete asks a judge to grant new capabilities rather than adjudicate each concrete call. ToolFence provides two key advantages. First, its fine-grained provenance-aware authorization enables the system to distinguish user-authorized values from untrusted observations, effectively addressing the within-tool attack. Second, its deterministic fast path and capability-level runtime grants substantially reduce the frequency of expensive judge calls, improving runtime efficiency. On AgentDojo with Qwen3-max, ToolFence reduces overall ASR to near zero with only a 3.80 percentage-point clean-utility drop and practical runtime overhead.
cs.CR / 21 / 2609.37217
When Cyber Scoring Systems Diverge: An Empirical Comparison
Abstract
Vulnerability scoring systems underpin cyber patch prioritization and risk management, but their comparative behavior is almost always assessed in the abstract, through correlation studies in IT vulnerability databases, rather than by the operational consequences they produce when embedded in a system-level risk model. Here we present an empirical comparison of four vulnerability scoring systems, namely CVSS (Common Vulnerability Scoring System), EPSS (Exploit Prediction Scoring System), SSVC (Stakeholder-Specific-Vulnerability Categorization), and IronMiner (operationally calibrated proprietary scoring system). As a substrate for comparison, we use a reconstruction of the 2015 Ukraine Power Grid operational-technology (OT) network that provides a documented incident topology. The results show a high degree of disagreement between the scoring systems. This suggests that the choice of the scoring system could significantly influence mitigation strategies and vulnerability prioritization, implying that a composite or hybrid scoring approach could offer a more suitable solution.
cs.CR / 22 / 2609.37218
Beyond Semantic Narrowing: Robust and Efficient LLM Watermarking with Hamming Neighborhoods
Abstract
Semantic watermarking improves robustness against watermark removal attacks by embedding detectable signals into sentence-level representations. However, existing watermarking methods typically impose watermark-specific semantic preferences on generated sentences without explicitly accounting for the highly non-uniform and context-dependent semantic preference of LLM generation. When these two preferences are poorly aligned, many natural continuations become incompatible with the watermark, causing semantic narrowing: reduced semantic freedom, increased resampling cost, and potential degradation on tasks with strict semantic requirements. To alleviate this problem, we propose HammingMark, which uses the semantic hash of the preceding sentence as a dynamic center and accepts candidates whose hashes fall within its Hamming neighborhood. Defining watermark validity over a Hamming neighborhood in compact hash space retains a larger fraction of naturally likely semantic continuations. The coarse many-to-one hash mapping further allows diverse semantic realizations to remain watermark-valid. Experiments on C4 and BookSum show that HammingMark achieves strong robustness, high detectability, and near-unwatermarked generation quality, requiring only 2.2 sampled candidates per accepted sentence,a 72.8% reduction compared with the most sampling-efficient existing method. On more complex tasks with strict semantic constraints, HammingMark achieves the highest detection rates with the highest or tied-highest ROUGE-L scores, demonstrating its effectiveness in balancing watermark detectability and generation quality under constrained generation settings.
cs.CR / 23 / 2609.37310
AutoMark: Enabling Autoresearch to Discover Better LLM Watermarks
Abstract
With LLM watermarking being deployed commercially and now required by regulations, improving its reliability and effectiveness has become crucial. Yet, recent progress in the field of LLM watermarking has increasingly been driven by improving details of existing methods, an effort fundamentally limited by the pace of human researchers. In this work, we enable for the first time the autonomous discovery of new distortion-free state-of-the-art watermarking schemes. To enable this, we (i) establish strict criteria to ensure that watermarks are reliable (e.g., they do not have an unexpectedly high false positive rate), (ii) propose rigorous statistical tests to automatically evaluate whether a watermarking scheme satisfies our criteria, and (iii) design an evaluation suite to rank watermarks along three key dimensions: detectability, quality, and robustness. By running our framework with 3 frontier models (GPT-6 Astra, Opus 5, Gemini-3.8 Flash), we discover over 50 different watermarking schemes, including several that outperform prior works along all key dimensions. We complement this by a manual study of the discovered schemes, distilling the key ideas into smaller components, and individually studying the impact of each component across dimensions (detectability, quality, robustness) to better understand how the proposed schemes operate. Importantly, we find that the agents, on top of improving existing ideas, also discover fundamentally new ideas (e.g., aligning watermark scores with random per-request direction). Overall, our work establishes the first steps of fully autonomous watermarking research, enabling the discovery of more reliable and effective watermarks. Our code is available at https://github.com/eth-sri/automark, and a blogpost to visualize our results at https://www.sri.inf.ethz.ch/blog/automark.
cs.CR / 24 / 2609.37608
Harvest Season for SLUB: From io_uring vulnerability to Novel Sheaf-Based Exploitation Techniques
Abstract
The Linux kernel's push for higher I/O performance and more efficient memory management has introduced new mechanisms that, while improving performance, also open new attack surfaces. This research examines two of them together: the io_uring subsystem and the sheaf/barn caching mechanism added to the SLUB allocator in Linux 6.18. In this research, two previously unknown vulnerabilities in io_uring are presented, and one is developed into a complete local privilege escalation chain under a hardened kernel configuration. Building this chain revealed that the sheaf/barn mechanism changes long-standing assumptions behind established exploitation techniques such as cross-cache attack, and that its design also weakens existing SLUB freelist protections. Both observations are analyzed and turned into working primitives. Building on this analysis, three novel sheaf-based exploitation techniques are proposed. Among them, an RCU-sheaf cross-cache technique removes the traditional dependence on the buddy system for moving objects across caches, giving more flexible and reliable control over object migration between cache pools. Together, these results characterize the sheaf/barn layer as a new and largely unexplored attack surface in Linux kernel exploitation.
cs.CR / 25 / 2609.37676
She Spoofed Sea Ships by the Sea Shore: Measuring Large-Scale GPS Spoofing in Global Maritime Traffic
Abstract
GPS spoofing has emerged as a serious threat to maritime security, yet its global prevalence, persistence, and structure remain largely unmeasured. In this paper, we present the first large-scale measurement study of maritime GPS spoofing, using global Automatic Identification System (AIS) data, which contain the GPS coordinates broadcasted over time by ships across the world. We focus on large-scale regional spoofing, where external interference displaces many vessels across an area at once, leaving a recognizable signature of physically implausible motion correlated across ships; our motion-aware, marine-specific framework identifies this signature and grades the evidence for GPS spoofing in each region it finds. Applying our approach to AIS data from over 367,000 vessels collected between late November 2024 and early February 2025, we identify 31 persistent anomalous hotspots across high-traffic maritime regions, at least 22 of which show strong evidence of GPS spoofing, with spatial and temporal structure aligning with regional conflict and economic sanctions. Notably, our method found that the spoofing activity in the Red Sea responsible for the highly-publicized grounding of the 75,000-ton container ship, MSC Antonia, was ongoing months before the incident, which has not been previously documented. Similarly, we detected persistent spoofing in the Strait of Hormuz over a year before the 2026 Iran war brought commercial shipping through the Strait to near-standstill. Together, this work establishes GPS spoofing as a widespread, recurring, and measurable threat to global maritime navigation.
cs.CR / 26 / 2609.37742
Lights, Camera, Attack: Exploiting Temporal HDR Fusion with Pulsed Light
Abstract
Modern cameras widely use temporal High Dynamic Range (HDR) to improve visibility by capturing a sequence of exposures with different integration times and fusing them into a single image. This process implicitly assumes that scene illumination remains sufficiently stable during capture. We introduce FLASH (Fusion-Level Attack by Saturating HDR), an external pulsed-light attack that deliberately attacks this assumption by creating cross-exposure inconsistency before downstream perception. FLASH exploits an algorithmic assumption rather than relying on sensor damage or hardware failure, and requires neither physical camera access, access to raw exposure brackets, knowledge of the fusion algorithm, nor exact phase lock to the camera. Across eight physical camera platforms spanning embedded, surveillance, photography, smartphone, and automotive use cases, and matched optical controls, FLASH causes pipeline-dependent darkening, overexposure, and visibility loss. This includes extreme-darkening rates of 50.0% on an iPhone 16 Pro and 33.7% on a Wyze Battery Cam Pro. On the Wyze camera, FLASH triggers the system-level low-visibility response in 10/10 trials, compared with 0/10 continuous-light and randomized-frequency flashing controls. In a controlled stationary OpenPilot case study, 23.0% of frames exhibit severe darkening in the traffic-cone target region, with target-background CNR decreasing by up to 90.8%. Under FLASH, the OpenPilot interface also fails to display the system-level path state observed in the corresponding control trials. In a controlled night-only HDR reconstruction stress test, a proof-of-concept exposure-rejection defense reduces median output-brightness deviation by 79.16%. These results show that temporal HDR fusion itself requires security-aware validation of exposure evidence.
cs.CR / 27 / 2609.37759
Selective Channel Restoration for Backdoored Vision-Language Models
Abstract
Vision-language models (VLMs) exhibit strong multimodal capabilities but remain vulnerable to backdoors implanted through poisoned fine-tuning data. Existing defenses often require extensive parameter updates during fine-tuning or incur per-query overhead during inference. To address these limitations, we propose Perturb-Select-Restore (PSR), a post-training defense that performs sparse updates to the projection interface and introduces no additional computation during inference. We reveal that backdoored VLM projectors are substantially more sensitive to bounded perturbations than clean VLM projectors, a phenomenon we term projection fragility. Building on this finding, PSR identifies the output channels most sensitive to perturbations in each projection layer of a backdoored VLM and restores their parameters to the corresponding pretrained values. Experiments across multiple tasks show that PSR reduces attack success rates to near zero while preserving clean-task performance.
cs.CR / 28 / 2609.37819
Making Duplicate Reimbursement Unrepresentable: A Verified Ethereum E-Invoice System for Humans and AI Agents
Abstract
Electronic invoices are replacing paper invoices worldwide, but today's centralized architectures leave three problems unsolved on the consumption side: an invoice can be submitted for reimbursement repeatedly, authenticity is difficult for recipients to verify, and data is siloed at a central authority that forms both a performance bottleneck and a single point of failure. This paper presents the design, formal analysis, and implementation of a complete blockchain-based electronic invoice system on Ethereum. We formalize the invoice lifecycle as a guarded labeled transition system and prove, under standard cryptographic and consensus assumptions, that the system guarantees: (i) reimbursement uniqueness--an invoice is reimbursed at most once, even across mutually distrusting organizations; (ii) face integrity--any verified invoice matches the recorded one unless keccak256 second-preimage resistance is broken; and (iii) authorization soundness for every lifecycle operation. The core invariants are machine-checked using Solidity SMTChecker, proving inductive validity across all reachable transaction sequences. The architecture models each invoice as a non-fungible, non-tradable token whose state transitions through five guarded subsystems, employing a lock-based protocol that makes duplicate reimbursement unrepresentable rather than merely detectable. We implement the design as a Solidity 0.8 contract with a four-role web application and evaluate it on a private Ethereum network: issuing costs 646,773 gas, full reimbursement costs under 135,000 gas, all operations run in O(1) time, and a single node sustains 137 issuances/s. Finally, the verified contract serves as a safety envelope for LLM-based reimbursement agents, provably rejecting unsafe actions (duplicate, over-limit, or forged-receipt claims) even when the agent's internal policy fails. All code and benchmarks are open-source.
cs.CR / 29 / 2609.37972
Dagger: Decoupling-based Model Stealing Attack against Graph Neural Networks
Abstract
As Graph Neural Networks (GNNs) are widely deployed as Machine Learning-as-a-Service (MLaaS) APIs, model stealing attacks have emerged as a critical security threat. By querying a victim model's black-box API, an adversary can construct a functionally equivalent surrogate model, compromising proprietary intellectual property and downstream security. Existing GNN stealing attacks, however, rely on overly permissive assumptions, such as soft-label outputs, large query budgets, full-graph query access, and prior knowledge of victim backbones that rarely hold in real-world deployments. In this work, we formalize a strictly constrained black-box, hard-label and backbone-agnostic threat model for GNN stealing attacks under a tight query budget. Given these realistic restrictions, we identify four fundamental challenges: sparse local structures and isolated nodes that degrade victim label quality, insufficient supervision signals, systematic imbalance with incomplete class coverage, and backbone mismatch. To address these interlocking barriers, we propose Dagger, a novel two-phase decoupling-based attack framework. Specifically, in Phase 1, Dagger pre-trains a surrogate using decoupled information propagation to preserve structural context over sparse local subgraphs while handling isolated nodes, combined with manifold-level node mixup to synthesize continuous supervision signals and smooth decision boundaries. In Phase 2, Dagger freezes the encoder and fine-tunes the classifier head via class-balanced sampling paired with logit adjustment to rectify severe query imbalance without requiring extra victim queries. Extensive experiments across four benchmark graphs and four GNN backbones demonstrate that Dagger consistently outperforms state-of-the-art GNN stealing attacks, achieving up to 18.16\% higher fidelity while only utilizing 12.23$\times$ fewer queries than the strongest baseline.
cs.CR / 30 / 2609.38012
A Function-level Dataset of Vulnerable and Fixed Source Code in JavaScript and TypeScript
Abstract
JavaScript and TypeScript are widely used in modern web development, making their security critical; however, automated vulnerability detection is often constrained by the availability of high-quality training data. Here we present JsVul, a dataset curated from seven major sources. Unlike generic multi-language datasets that may retain noise -- such as minified code and cosmetic edits -- JsVul utilizes a language-specific pipeline. We collected pre-fix and post-fix versions of files around security fixes and, by filtering irrelevant artifacts and applying automated syntax normalization, isolated security-related changes. We ensured data integrity through multi-stage deduplication and heuristic-based labeling. Provided in a time-ordered JSONL format, JsVul supports robust model training in the JavaScript and TypeScript ecosystem and demonstrates the importance of language-aware preprocessing in building vulnerability datasets.
cs.CR / 31 / 2609.36795
Epistemic Typing as a PostgreSQL Table Access Method: Adversarial Conflict Resolution Under Confidence Forgery and Sybil Coordination
Abstract
We describe KNDB, a PostgreSQL 18 table access method (TAM) that types every row with an engine-assigned epistemic kind (MEASURED, INFERRED, or DERIVED) and resolves per-slot conflicts inside every write-time heapam callback. Rows land as ordinary heap tuples; seven of the 44 TAM callbacks are overridden (tuple_insert, multi_insert, tuple_update, tuple_delete, tuple_insert_speculative, tuple_complete_speculative, relation_toast_am), the other 37 delegate to heap; we provide a completeness argument over the interface as a paper artefact. This paper reports the engineering behind that decision and the adversarial evaluation that motivated it. On a confidence-forgery workload where an attacker asserts INFERRED writes with confidence in [0.95,1.0] against honest MEASURED writes with confidence in [0.5,0.9], KNDB beats a confidence-only baseline by 63 percentage points on the Book-Author fusion dataset and 92.7 points on the Zheng crowdsourcing dataset. Both wins are proven load-bearing on the kind axis by a source-rebuild disable-and-test in which the lattice is neutralised and the win vanishes. Against four truth-discovery baselines (TruthFinder, CRH, CATD, ACCU) reimplemented from the original equations and validated to within 0.3 percentage points of the published numbers, KNDB is competitive below a per-dataset density-saturation cell and dominant at or above it. We formalise the cell as k* ~ rho_alg * h_top, where h_top is per-slot top honest surface-form support, and validate the prediction within +/-20% on Book-Author and +/-30% on Zheng. Because the kind axis is assigned by the engine from independent metadata and cannot be forged at write time, KNDB's k* is unbounded. The paper is honest about where KNDB loses: CRH and ACCU outperform KNDB below saturation on Zheng, and KNDB scores zero on three temporal knowledge-editing benchmarks whose ground truth is last-writer-wins.
cs.CR / 32 / 2609.37342
Collision Detection is Instance $\widetilde{O}$ptimal Under the Birthday Threshold
Abstract
Can structural knowledge about a hash function help accelerate the (black box) detection of collisions in it? This question is fundamental to cryptography theory given the importance of collision-resistant hash functions, and in this paper we tackle it from the angle of instance optimality, an ultimate notion of beyond worst case algorithm analysis that has gained significant traction in recent years. Instance optimality asks for a single algorithm that, on every input, performs nearly as well as the best correct algorithm that ``knows the structure'' of that specific input. Here we measure algorithms by the number of queries they make to the hash function $f\colon [n]\to [n]$, and we say that an algorithm ``knows the structure'' of the input if, in addition to query access to $f$, it has free access to an unlabeled copy $π^{-1}\circ f\circπ$ of $f$, for an unknown permutation $π$ on $[n]$. We prove the existence of an (almost) instance-optimal algorithm for collision detection in the regime most interesting from a cryptographic perspective: among functions where finding a collision takes significantly less than $\sqrt{n}$ queries. Specifically, we prove the existence of a single algorithm $A$ that, for any input $f$ in which a structure-aware algorithm can find a collision using $q\leq O(\sqrt{n/\log n})$ queries in expectation, $A$ can find a collision in at most $O(q\log n)$ queries. The $O(\log n)$ multiplicative overhead is tight, matching a lower bound of Ben-Eliezer, Grossman, and Naor [ICALP'25], and partially resolving their main open question. Our result implies, in particular, that it is impossible for a cryptographic designer to plant purely structural backdoors for collision finding (for this unlabeled notion of structure): whatever collisions the designer's secret knowledge finds, the public can find with a multiplicative overhead of $O(\log n)$.
cs.CR / 33 / 2609.36485
From Reconnaissance to Response: Quantitative Risk Parameterization and Game Theoretic Containment in Modern Enterprise Attack
Abstract
Modern Security Operations Centers struggle with delayed manual incident response, enabling adversaries to advance through the Cyber Kill Chain during early stage reconnaissance. While classical game theoretic defense models optimize strategic resource allocation, they rely on static utility matrices that fail to adapt to dynamic telemetry. This paper presents an integrated, metrics driven decision engine that bridges quantitative risk parameterization and continuous automated response time. Common Vulnerability Scoring Systems exploitability parameters are mapped to attacker success probabilities and evaluate defender log distributions via Factor Analysis of Information Risk Monte Carlo simulations. Real time SIEM logs streams are modeled as Poisson process arrival rates, dynamically updating defender posterior threat belief through sequential Bayesian filtering. A closed form threshold is derived by framing the interaction as a dynamic Bayesian Stackelberg game, where the expected unmitigated risk exceeds proactive containment cost. Parameterized against empirical data from the 2023 MGM Resorts and Caesars Entertainment cyber incident, simulation results demonstrate that the engine suppresses transient background noise while triggering automated SOAR network isolation within seconds of adversarial probing. Multi parameter sensitivity analysis confirms that the decision boundary dynamically adjusts to live perimeter vulnerability, offering a control theoretic foundation for sub minute automated threat containment.
cs.CR / 34 / 2609.36667
When One Leak Pays Forever: Context Binding and the Price of Deterring Collusion
Abstract
A coalition that deviates once can profit many times when what it sells keeps working. In a threshold-encrypted mempool, a leading defense against maximal extractable value (MEV), a quorum of the decryption committee that sells its decryption capability to a front-runner exposes every later block that the capability still decrypts. We ask how large a penalty, such as slashable stake, deters this kind of collusion. In our repeated game, a single leak by any coalition in a monotone family of authorized coalitions (for example, any $k$ of the $n$ committee members) unlocks a set of future rounds, costs a one-time penalty, and ends the coalition's participation. We show that every dynamic deviation reduces to choosing a leak time, so deterrence holds if and only if each coalition's penalty covers the largest discounted value that a single leak reaches. Without discounting, over $T$ rounds of unit value full reuse needs a penalty of $T$ while binding each leak to its own round needs $1$, so no penalty that is constant in the horizon deters unbounded reuse; a reuse window of $w$ rounds costs at most $w$ times the largest per-round value. The cheapest profile of per-party stakes that deters every coalition solves a covering linear program. For blockchain design, per-epoch keys cut the required stake from the value of a key's lifetime to the value of one epoch; we calibrate the gap on Ethereum front-running data and place Ferveo and Shutter in the model. The analysis extends to sealed-bid auctions, multi-authority voting, and federated learning under a shared key.
cs.CR / 35 / 2609.36335
Conversable Quantum Fire in the Standard Model
Abstract
Quantum fire consists of quantum states that can be efficiently functionally cloned but cannot be transmitted using one-way classical communication. Whereas all previous quantum-fire constructions with proven untelegraphability are relative to an oracle, ours, based on one-shot signatures, is in the standard model. Our construction achieves two further notions that we introduce: (a) keyless untelegraphability, a strengthening of untelegraphability; and (b) conversability, meaning that a flame, though not telegraphable via one-way classical communication, can be transmitted via classical interaction. Hence one-way and two-way classical communication differ qualitatively in their power to transmit this quantum fire. We use these two novel properties of quantum fire, along with its clonability, to give one of the first cryptographic applications of quantum fire. Keyless untelegraphability forces the distribution of serial numbers of valid flames to have high min-entropy, while conversability and clonability allow the corresponding flame to be cloned and transferred using only classical communication. We present two variants of our conversable quantum-fire construction. The first, based on one-shot signatures (instantiable from subexponential iO, subexponentially secure one-way functions, and LWE), supports polynomially many flames, and super-logarithmic min-entropy of the serial-number distribution. The second, under the stronger assumption of exponentially unforgeable one-shot signatures (instantiable relative to a classical oracle), supports exponentially many flames, and linear min-entropy.
cs.CR / 36 / 2609.37276
Quantum Leakage Resilience of Shamir Secret Sharing
Abstract
We initiate the study of quantum leakage resilience of unmodified Shamir secret sharing over prime fields. A well-studied leakage model for Shamir's secret sharing classically is single-bit local leakage from each share. We consider its quantum analogue where, for each party, a local leakage channel takes as input the party's share and outputs a leaked qubit. Without preshared entanglement, we show that the distinguishing advantage is $2^{-Ω(n)}$ when the threshold rate $t/n=τ$ exceeds $τ_\star\approx0.73339$ by a fixed positive margin. More generally, we allow disjoint entangled blocks of any fixed maximum size where there is no entanglement between different blocks or with the adversary, and each block emits at most a fixed number of qubits. Security holds when the threshold rate is high enough (sufficiently close to one). We then allow a specified set of devices to share entanglement with the adversary. We show that security holds even when a linear number of devices ($αn$ for small $α>0$) share entanglement with each other and with the adversary for a large enough threshold rate. As a complementary negative result, we also show that even classical single-bit leakage makes Shamir scheme insecure if we allow arbitrarily large entanglement between the leakage devices. A GHZ state shared by exactly $t$ leakage devices makes even classical one-bit leakage insecure, without any entanglement with the adversary. In this attack, each participating device emits only one classical bit, and their joint parity distinguishes any chosen pair of secrets with a constant advantage. Thus, for fixed threshold rates above $τ_\star$, the maximum number of devices that may share arbitrary entanglement with one another and with the adversary while preserving security is linear in $n$ up to constant factors, although the optimal support fraction remains open.
cs.CR / 37 / 2609.37429
Succinct Arguments for QMA in the Quantum Random Oracle Model
Abstract
Succinct arguments are a fundamental cryptographic primitive for verifying computational claims with small communication. In the classical setting, succinct arguments for NP can be constructed from unstructured hardness alone (e.g., hash functions) by compiling probabilistically checkable proofs (PCPs) or interactive oracle proofs (IOPs) for NP via the commit-and-open paradigm. In contrast, known succinct arguments for QMA rely on ``structured'' cryptographic primitives, or on the quantum PCP conjecture. We construct the first succinct argument for QMA in the quantum random oracle model (QROM) without relying on additional cryptographic assumptions or unproven conjectures. This yields succinct arguments for QMA from unstructured hardness alone, showing that ideal hash functions not only suffice for succinct arguments for NP but also for QMA. Underlying our result is an efficiency-preserving transformation that compiles quantum interactive oracle proofs (QIOPs), a recently introduced interactive generalization of quantum PCPs, into quantum arguments for the same language, via a natural quantum commit-and-open paradigm. Our transformation applies to every QIOP with public-query soundness, a notion that we formalize to capture a natural requirement of the commit-and-open paradigm and is satisfied by a known QIOP for QMA. As a key ingredient in our transformation, we formalize and construct extractable vector commitments for quantum states with local openings in the QROM, which may be of independent interest.
cs.CR / 38 / 2609.38060
Classical Verification of Quantum Computation with Quasilinear Resources, from Compiled Nonlocal Games
Abstract
Computational self-testing gives a classical verifier command over the quantum register of a single computationally bounded prover. We use this framework to construct the first argument system for BQP with quasilinear total resource requirements in the circuit model. Our argument system is based on the learning with errors (LWE) assumption and requires total resources of $O(\mathrm{poly}(λ, \log g)\cdot g)$ for delegating a circuit with $g$ gates, where $λ$ is the LWE security parameter. This is achieved by constructing a new computational self-test for certifying the prover's quantum state and using it to dequantize the efficient verification protocol of Broadbent (ToC 2018). Specifically, this self-test enables the verifiable, random remote state preparation of tensor product states of the single-qubit Clifford observables $σ_X, σ_Y, σ_Z, (σ_Y-σ_X)/\sqrt{2}$ and $(σ_Y+σ_X)/\sqrt{2}$, with constant robustness: the verification error is independent of the number of prepared qubits. This approach was first proposed by Coladangelo et al. (ToC 2024) in the multi-prover setting. We replicate their result in the single-prover setting by applying the compiler proposed by Kalai et al. (STOC 2023)---which turns any nonlocal game into a single-prover argument system---to a modified version of their self-test.