Daily Research Digest
arXiv Papers
2026-10-08
569
Papers
9
Categories
115
Translated
收藏清单 0
精选 · Favorites
115
cs.AI / 1 / 2610.09323
Constrained Diffusion for Data-Scarce Orbital Monte Carlo in Constellation Tasking
用于星座任务分配中数据稀缺轨道蒙特卡洛的受约束扩散
diffusion
扩散模型相关
Abstract
Constellation Monte Carlo results depend on the orbital population used to evaluate a tasking policy. With scarce reference trajectories, replay limits geometric diversity, while independent orbital-element jitter can violate physical constraints. We study constrained diffusion for orbital-population augmentation. A force-conditioned diffusion model learns a 13-dimensional orbital prior, recovering semimajor axis from perigee altitude and eccentricity; Basilisk propagates each sample under one of five force-model tiers. Using 800 reference trajectories, we compare diffusion with jittered bootstrap, per-tier Gaussian mixtures, and a conditional variational autoencoder, and evaluate distributional fidelity, support shift, classical astrodynamics diagnostics, and 4,000 paired GoDSAT-compatible campaigns. The largest diffusion model achieves held-out trajectory MMD of 0.0171 +/- 0.0239, similar to bootstrap (0.0170) and the mixture (0.0198), but samples farther from training priors (median nearest-training distance 2.7 versus 0.10 standardized units). All generators fail to match the shifted-blind population (classifier AUC 0.991-0.998). Endpoint-conditioned samples satisfy Lambert boundaries but have greater interior error than the matched Lambert reference (29.2 versus 5.7 km mean RMSE). Residual diffusion improves selected sparse forecasts and catalog-mode recall but does not outperform classical estimators on custody ranking. In a fixed 16-satellite configuration, diffusion yields similar mean custody to replay and bootstrap, while the force-tier mixture shifts custody by about 5.5 percentage points. Constrained diffusion supports local, in-support augmentation but cannot replace orbital dynamics or serve as an operational posterior. Orbital-population construction is a consequential source of uncertainty in constellation analysis.
Chinese Translation
星座蒙特卡洛结果依赖于用于评估任务分配策略的轨道种群。在参考轨迹稀缺的情况下,回放限制了几何多样性,而独立的轨道根数抖动可能违反物理约束。我们研究用于轨道种群增强的受约束扩散。一个以力模型为条件的扩散模型学习一个13维轨道先验,从近地点高度和偏心率恢复半长轴;Basilisk 在五个力模型层级之一下传播每个样本。使用800条参考轨迹,我们将扩散与抖动自助法、按层级的高斯混合模型以及条件变分自编码器进行比较,并评估分布保真度、支撑偏移、经典天体动力学诊断以及4,000组配对的GoDSAT兼容活动。最大的扩散模型在留出轨迹上的 MMD 达到 0.0171 +/- 0.0239,与自助法(0.0170)和混合模型(0.0198)相似,但其样本距训练先验更远(最近训练距离中位数为 2.7 对 0.10 标准化单位)。所有生成器均无法匹配移位盲种群(分类器 AUC 0.991-0.998)。端点条件样本满足 Lambert 边界,但比匹配的 Lambert 参考具有更大的内部误差(平均 RMSE 为 29.2 对 5.7 km)。残差扩散改善了选定的稀疏预报和目录模式召回,但在 custody 排序上并未优于经典估计器。在固定的16颗卫星配置中,扩散产生的平均 custody 与回放和自助法相似,而力模型层级混合使 custody 偏移约5.5个百分点。受约束扩散支持局部的、支撑内的增强,但不能替代轨道动力学,也不能作为可操作的后验。轨道种群构建是星座分析中一个会产生重大影响的不确定性来源。
cs.AI / 2 / 2610.08966
Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models
人类的第六感:对多模态模型中的直觉视觉推理进行基准测试
large language model
大语言模型相关
Abstract
Humans perceive far more in a scene than what is explicitly depicted: a single glance captures past causes and future trajectories; a quick peek determines if a vehicle can fit between two parked cars; a few seconds of video reveals who holds authority in a room; and a fleeting clip highlights subtle abstract patterns like unwritten rules or hidden labels. This capacity reflects a form of humanity's sixth sense: an intuitive reasoning mechanism that recovers implicit information beyond raw sensory perception. Crucially, this rapid, zero-shot visual intuition underpins everyday navigation and social interaction, making it a vital capability for Multimodal Large Language Models (MLLMs) deployed alongside people. Existing visual benchmarks, however, target either deliberate expert-level analysis in academic and mathematical domains or low-level perception, leaving the intuitive reasoning that people perform largely untested. To bridge this gap, we introduce Humanity's Sixth Sense (HSS), a benchmark for intuitive visual reasoning. HSS spans diverse image and video inputs, organizes items under a structured taxonomy, and pairs each with human-written prompts probing the implicit temporal, spatial, social, and abstract structure that people infer at a glance. Frontier MLLMs fall short of human performance: participants reach 93.1% accuracy, while the strongest model, GPT-6-astra, reaches only 53.6% even at maximum reasoning effort. Despite excelling in many complex tasks that require advanced perception and knowledge, current models still struggle significantly on these visual tasks that are intuitive for humans. We further explore agentic setup that apply dynamic visual manipulation to HSS, which narrows but does not close the gap. HSS establishes intuitive visual reasoning as a measurable axis and directs attention to a capability that scaling on current benchmarks has so far left behind.
Chinese Translation
人类在场景中感知到的远多于其中明确描绘的内容:一瞥即可捕捉过去的原因和未来的轨迹;快速一瞥即可判断一辆车能否停入两辆停放的汽车之间;几秒钟的视频即可揭示房间里谁拥有权威;而一段转瞬即逝的片段则能凸显诸如不成文规则或隐藏标签之类的微妙抽象模式。这种能力反映了人类第六感的一种形式:一种直觉推理机制,它能够恢复超出原始感官知觉的隐含信息。至关重要的是,这种快速、零样本的视觉直觉支撑着日常导航和社交互动,使其成为与人类一起部署的多模态大语言模型(MLLMs)的一项关键能力。然而,现有的视觉基准要么针对学术和数学领域中经过深思熟虑的专家级分析,要么针对低级感知,致使人类所执行的直觉推理在很大程度上仍未得到测试。为弥合这一差距,我们引入人类的第六感(HSS),一个用于直觉视觉推理的基准。HSS 涵盖多样化的图像和视频输入,在结构化分类体系下组织条目,并为每一项配以人工编写的提示,用以探查人们一眼即推断出的隐含时间、空间、社会和抽象结构。前沿 MLLMs 未能达到人类表现:参与者达到 93.1% 的准确率,而最强模型 GPT-6-astra 即使在最大推理努力下也仅达到 53.6%。尽管当前模型在许多需要高级感知和知识的复杂任务中表现出色,但它们在那些对人类而言直观的视觉任务上仍然面临显著困难。我们进一步探索将动态视觉操作应用于 HSS 的智能体式设置,这缩小了差距,但并未弥合差距。HSS 将直觉视觉推理确立为一个可测量的维度,并将注意力引向一种迄今在当前基准上的扩展所遗落的能力。
cs.AI / 3 / 2610.08967
Socio-Foundation: A Model for Generalizable Individual Behavior Simulation via Hierarchical Capability Distillation
Socio-Foundation:一种通过分层能力蒸馏实现可泛化个体行为模拟的模型
large language model
大语言模型相关
Abstract
Simulating individual behavior requires large language models (LLMs) to preserve persona traits while adapting to dynamic social contexts. However, general-purpose LLMs often flatten distinct personas, while task-specific tuning suffers from fragmentation and generalization. To overcome these challenges, we organize individual simulation into the \textbf{FONTS Taxonomy}, comprising five complementary capability dimensions: \emph{persona fidelity} (\textbf{F}), \emph{outcome realization} (\textbf{O}), \emph{behavioral naturalness} (\textbf{N}), \emph{trajectory coherence} (\textbf{T}), and \emph{social grounding} (\textbf{S}). Grounded in this taxonomy, we curate a standardized training corpus library of approximately 10 million instances across 14 representative datasets and present \textbf{Socio-Foundation}. Socio-Foundation decouples specialization from integration via a three-stage pipeline: learning task experts via DAPO, consolidating them into capability experts via off-policy distillation, and unifying them via multi-teacher on-policy distillation (MOPD). We also establish \textbf{IndiEval}, consolidating 29 metrics across the FONTS dimensions. Experiments show that Socio-Foundation outperforms its \textit{Qwen3-8B} base by 11.0 points and approaches frontier models such as \textit{GLM-5.2}, with ablations and out-of-distribution evaluations further demonstrating the effectiveness and generalization of our model.
Chinese Translation
模拟个体行为要求大语言模型(LLMs)在适应动态社会语境的同时保持人设特质。然而,通用LLMs往往会抹平不同人设之间的差异,而面向特定任务的调优则受困于碎片化与泛化能力不足。为克服这些挑战,我们将个体模拟组织为\textbf{FONTS Taxonomy},其包含五个互补的能力维度:\emph{人设保真度}(\textbf{F})、\emph{结果实现}(\textbf{O})、\emph{行为自然度}(\textbf{N})、\emph{轨迹连贯性}(\textbf{T})以及\emph{社会基础}(\textbf{S})。以该分类体系为基础,我们整理了一个标准化训练语料库,涵盖14个代表性数据集、约1000万个实例,并提出了\textbf{Socio-Foundation}。Socio-Foundation通过一个三阶段流水线将专门化与整合解耦:通过DAPO学习任务专家,通过离策略蒸馏将其整合为能力专家,并通过多教师同策略蒸馏(MOPD)将它们统一起来。我们还建立了\textbf{IndiEval},整合了FONTS维度上的29项指标。实验表明,Socio-Foundation比其\textit{Qwen3-8B}基座高出11.0分,并接近\textit{GLM-5.2}等前沿模型,消融实验和分布外评估进一步证明了我们模型的有效性与泛化能力。
cs.AI / 4 / 2610.08993
Verify Less, Evolve More: Training Idea-Level Critics for Verification-Efficient ML Evolving Agents
少验证,多进化:为验证高效的机器学习进化智能体训练想法级评论模型
large language model
大语言模型相关
Abstract
As large language models become more powerful, self-evolving agents are able to tackle challenging tasks including AI for machine learning (AI4ML). In AI4ML, while empirical verification is available, it often requires computationally costly model training and evaluation, limiting the speed and scale of agent evolution. Yet verification efficiency remains under-explored, and frontier models provide only limited gains when used directly as idea selectors. We address this gap with specialized idea-level critic models that predict whether a proposed ML modification will improve upon the current solution, allowing agents to screen ideas and concentrate verification resources on the most promising candidates. We train the critic models through supervised fine-tuning on high-quality critiques synthesized by Gemini-3.1-Pro, followed by GRPO to further improve their predictive accuracy. Empirically, our critic models outperform Gemini-3.1-Pro in static idea evaluation, and these gains extend to agent inference, continual learning, and policy training. During inference-time evolution, they improve final solution quality under the same verification budget by selecting more promising ideas, with further gains from continual learning. During policy training, they serve as learned reward models, reserving empirical verification for uncertain cases and enabling substantially more policy updates with the same verification resources. Together, these results show that idea-level critic models help ML agents discover better solutions and learn stronger proposal policies under limited verification budgets.
Chinese Translation
随着大语言模型变得更加强大,自我进化的智能体能够处理包括机器学习人工智能(AI4ML)在内的挑战性任务。在 AI4ML 中,尽管经验性验证是可用的,但它通常需要计算成本高昂的模型训练与评估,从而限制了智能体进化的速度与规模。然而,验证效率仍未得到充分探索,而前沿模型在直接用作想法筛选器时只能带来有限的收益。我们通过专门的想法级评论模型来填补这一空白,这些模型预测所提议的机器学习修改是否会改进当前解决方案,使智能体能够筛选想法并将验证资源集中在最有希望的候选方案上。我们通过在 Gemini-3.1-Pro 合成的高质量评论上进行监督微调来训练评论模型,随后使用 GRPO 进一步提升其预测准确率。在实验中,我们的评论模型在静态想法评估中优于 Gemini-3.1-Pro,并且这些收益可扩展至智能体推理、持续学习和策略训练。在推理时进化过程中,它们通过在相同验证预算下选择更有希望的想法来提高最终解决方案的质量,并可从持续学习中获得进一步提升。在策略训练过程中,它们充当学习到的奖励模型,将经验性验证保留给不确定的情形,并在相同的验证资源下实现大幅更多的策略更新。总体而言,这些结果表明,想法级评论模型能够帮助机器学习智能体在有限的验证预算下发现更好的解决方案并学习更强的提案策略。
cs.AI / 5 / 2610.09007
Sigma-Hunter: A Domain-Specific Language Model for Threat Hunting and Detection Engineering
Sigma-Hunter:面向威胁狩猎与检测工程的领域特定语言模型
large language model
大语言模型相关
Abstract
Detection engineers must translate threat reports, forensic observations, and hunt hypotheses into precise, testable rules. General-purpose large language models (LLMs) can draft such rules, but often produce invalid YAML, incorrect log sources, unsupported fields, or overly broad detection logic. This paper presents \emph{Sigma-Hunter}, a domain-adapted LLM for analyst-assistive Sigma rule generation and threat hunting. We build an instruction-tuning dataset from 3,635 validated open-source Sigma rules, expanded into 7,663 question-answer and analyst-reasoning examples. Each source rule is assigned to a single train, validation, or test partition before this expansion, so no rule leaks across splits. We fine-tune a 7B Mistral model and a Phi-4 model with LoRA and score held-out rule generations on syntax, approximate field consistency, and a semantic judgment of detection logic, completeness, selectivity, and log-source alignment. Sigma-Hunter-Mistral scores 8.17 overall, against 7.88 for the strongest general-purpose baseline and 4.61 for untuned Mistral. Two findings stand out: domain adaptation enables a compact 7B model to perform competitively with larger general-purpose models on this structured task, and syntactic validity is a weak proxy for semantic rule quality, as several baselines emit well-formed YAML carrying weak detection logic. The adapted models run locally, which suits detection engineering in disconnected environments where analysts cannot reach hosted model services.
Chinese Translation
检测工程师必须将威胁报告、取证观察和狩猎假设转化为精确、可测试的规则。通用大型语言模型(LLM)可以起草此类规则,但常常产生无效的 YAML、错误的日志源、不支持的字段,或过于宽泛的检测逻辑。本文提出 \emph{Sigma-Hunter},一种领域适配的 LLM,用于辅助分析师的 Sigma 规则生成和威胁狩猎。我们基于 3,635 条经过验证的开源 Sigma 规则构建了一个指令微调数据集,并将其扩展为 7,663 个问答和分析师推理示例。在此扩展之前,每条源规则都被分配到单个训练、验证或测试划分中,因此没有规则跨划分泄漏。我们使用 LoRA 对 7B Mistral 模型和 Phi-4 模型进行微调,并从语法、近似字段一致性,以及对检测逻辑、完整性、选择性和日志源对齐的语义判断方面,对留出的规则生成进行评分。Sigma-Hunter-Mistral 总体得分为 8.17,而最强的通用基线为 7.88,未微调的 Mistral 为 4.61。两个发现尤为突出:领域适配使一个紧凑的 7B 模型在此结构化任务上能够与更大的通用模型竞争,并且语法有效性是语义规则质量的弱代理指标,因为若干基线生成了格式良好的 YAML,却携带薄弱的检测逻辑。适配后的模型在本地运行,这适合分析师无法访问托管模型服务的断连环境中的检测工程。
cs.AI / 6 / 2610.09039
From High Recall to High Utility: Dataset-Adaptive Post-Processing of LLM-Generated Customer Intents
从高召回率到高实用性:面向 LLM 生成的客户意图的数据集自适应后处理
large language model
大语言模型相关
Abstract
Large language models can extract useful signals from heterogeneous enterprise data, but high-recall extraction often produces outputs that are duplicated, uneven in granularity, semantically overlapping, or too numerous for downstream systems and human reviewers to use effectively. We present a dataset-adaptive post-processing architecture developed for Customer Intent Extraction (CIE), where unstructured customer language is transformed into stable, traceable intent units. The approach separates recall-oriented extraction from utility-oriented reduction. Source-specific preprocessing first isolates evidence from multimodal plans, sparse operational records, and structured opportunity data. Candidate intents are then standardized and deduplicated, optionally enriched with metadata for embedding computation, represented in a shared semantic vector space, and grouped using a clustering strategy selected according to the candidate set's characteristics. Cluster-level keywords provide an explainability layer, while singleton reassignment requires agreement between embedding and keyword similarity. Finally, constrained language-model aggregation produces one concise intent per cluster without introducing unsupported concepts, and the resulting unit retains provenance, clustering, embedding, and generation metadata. This treats post-processing not as cosmetic cleanup, but as a semantic reduction layer converting high-recall LLM outputs into reusable enterprise intelligence. We also describe two downstream applications: Machine-Generated Intents, which infer likely objectives for customers lacking direct evidence from peer customers with similar profiles, and intent-guided semantic retrieval and mapping, which uses the stable intent as a query against a downstream decision space, illustrated here by mapping customer intents to business outcomes.
Chinese Translation
大语言模型能够从异构企业数据中提取有用的信号,但高召回率的提取往往会产生重复、粒度不均、语义重叠,或数量过多以致下游系统和人工审核者无法有效使用的输出。我们提出了一种为客户意图提取(CIE)而开发的数据集自适应后处理架构,在该架构中,非结构化的客户语言被转化为稳定、可追溯的意图单元。该方法将面向召回的提取与面向实用性的归约分离开来。特定于数据源的预处理首先从多模态方案、稀疏的运营记录和结构化的商机数据中分离出证据。随后,候选意图被标准化并去重,可选地通过元数据加以丰富以用于嵌入计算,在共享的语义向量空间中表示,并使用根据候选集特征所选择的聚类策略进行分组。聚类层面的关键词提供了一个可解释性层,而单例重新分配则要求嵌入相似度与关键词相似度之间取得一致。最后,受约束的语言模型聚合为每个聚类生成一个简洁的意图,且不引入无依据的概念,而所得单元保留了来源、聚类、嵌入和生成元数据。这将后处理视为一个语义归约层,而非表面上的清理,从而把高召回率的 LLM 输出转化为可复用的企业智能。我们还描述了两种下游应用:机器生成意图(Machine-Generated Intents),它针对缺乏直接证据的客户,从具有相似画像的同类客户中推断其可能的目标;以及意图引导的语义检索与映射,它将稳定的意图作为查询,作用于下游决策空间,本文通过将客户意图映射到业务成果来对此加以示例说明。
cs.AI / 7 / 2610.09053
Justice After Identity: Large Language Models and the View from Everywhere
身份之后的正义:大型语言模型与来自各处的视角
large language model
大语言模型相关
Abstract
The search for a common view of justice and fairness has challenged human collective activity, as our diverging judgments are unavoidably shaped by the self-interests of social position, personal benefit, cultural inheritance, and historical circumstance. John Rawls famously attempted to overcome this limitation through popularizing a philosophical tradition known by the phrase "the original position" - a thought experiment by which people select principles of justice without knowing the identities or advantages they will possess. Critics, however, have long questioned whether people can meaningfully suspend their social identities and suppress morally relevant forms of lived experience. Artificial intelligence engaged to calculate algorithmic and agentic fairness introduces a novel possibility. LLMs have no singular class, race, gender, nationality, or biography, yet their parameters encode linguistic representations of a vast range of human identities and moral traditions. Perhaps the ethical judgments of LLMs could approximate an integrative original position - a "view from everywhere" generated not by excluding social identities but by computationally incorporating their diversity.It is unlikely that humankind will "hand over the keys" to computational systems by simply delegating complete agentic control of distributive and procedural collective processes. But AI may play a role, perhaps a positive one, interacting with individual and collective human judgment as we often confront increasingly polarized views on what is fair and just. We present data comparing human, base model, and frontier/fine-tuned model judgments about classic moral dilemmas while systematically varying identity relationships and Rawlsian constraints on identity. We conclude by speculating whether, if advanced AI systems provide humans with thoughtful advice, humans would actually be likely to accept it.
Chinese Translation
对正义与公平的共同看法的探求一直挑战着人类的集体活动,因为我们彼此分歧的判断不可避免地受到社会地位、个人利益、文化传承和历史境遇所构成的自利的塑造。约翰·罗尔斯曾著名地试图通过推广一种被称为“原初状态”的哲学传统来克服这一局限——这是一种思想实验,人们在其中选择正义原则,却不知道自己将拥有何种身份或优势。然而,批评者长期以来一直质疑,人们是否能够有意义地悬置自己的社会身份,并压制那些在道德上相关的亲历经验形式。被用来计算算法公平性和能动性公平性的人工智能引入了一种新的可能性。大型语言模型没有单一的阶级、种族、性别、国籍或生平,然而它们的参数编码了广泛人类身份和道德传统的语言表征。也许大型语言模型的伦理判断能够近似一种整合性的原初状态——一种“来自各处的视角”,它不是通过排除社会身份而产生,而是通过计算性地纳入其多样性而产生。人类不太可能仅仅通过把对分配性和程序性集体过程的完整能动控制委托给计算系统,就“交出钥匙”。但人工智能可能发挥作用,或许是积极的作用,在我们常常面对关于何为公平与正义的日益两极分化的观点时,与个体和集体的人类判断互动。我们呈现数据,比较人类、基础模型以及前沿/微调模型对经典道德困境的判断,同时系统地改变身份关系和罗尔斯式身份约束。我们最后推测,如果先进的人工智能系统为人类提供深思熟虑的建议,人类实际上是否可能接受它。
cs.AI / 8 / 2610.09146
Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI
冻结的模型,演进的专业知识:面向多模态医学人工智能的从部署经验中进行的模型无关学习
large language model
大语言模型相关
Abstract
Large language models (LLMs) and vision-language models (VLMs) are usually frozen after deployment, so they do not learn from the cases they solve. This is especially concerning in medicine, where new clinical evidence, updated guidelines, and new therapies can change established practice. Fine-tuning can update the model, but it requires access to model weights and additional training. Parameter-free methods avoid training, but they may overfit a fixed validation set, lack reliable domain knowledge, or lose visual details by saving experience only as text. To address these limitations, we present a model-agnostic framework that allows frozen LLMs and VLMs to learn from deployment experience through three forms of external expertise: a Skill that guides reasoning and tool use, a Knowledge Memory that stores reliable facts supported by earlier cases or trusted external evidence, and a Multimodal Knowledge Base that keeps visual examples and guides the model to relate each retrieved case to the current image. Instead of relying on a fixed validation set, a validation strategy keeps an update only if it helps on new cases without degrading performance on earlier ones. Across six benchmarks covering clinical diagnosis, clinical workflows, medical reasoning, and medical and non-medical visual reasoning, and with four open-weight and closed-source base models, our framework improves performance during online deployment by up to 34.2% over the base model on medical tasks, generalizes to unseen cases, transfers to other models without further optimization, and works in non-medical domains.
Chinese Translation
大型语言模型(LLMs)和视觉-语言模型(VLMs)通常在部署后被冻结,因此它们不会从所解决的案例中学习。这在医学中尤其令人担忧,因为新的临床证据、更新的指南和新的疗法可能改变既定的实践。微调可以更新模型,但它需要访问模型权重并进行额外训练。无参数方法避免了训练,但它们可能过拟合一个固定的验证集,缺乏可靠的领域知识,或者由于仅将经验保存为文本而丢失视觉细节。为了解决这些局限,我们提出一个模型无关的框架,它允许冻结的LLMs和VLMs通过三种形式的外部专长从部署经验中学习:一个指导推理和工具使用的技能(Skill)、一个存储由早期案例或可信外部证据支持的可靠事实的知识记忆(Knowledge Memory),以及一个保留视觉示例并引导模型将每个检索到的案例与当前图像关联起来的多模态知识库(Multimodal Knowledge Base)。与依赖固定的验证集不同,一种验证策略仅在某个更新有助于新案例且不会降低在较早案例上的性能时才保留该更新。在涵盖临床诊断、临床工作流、医学推理以及医学和非医学视觉推理的六个基准上,并且使用四个开放权重和闭源基础模型,我们的框架在医学任务上的在线部署期间将性能较基础模型提升最多34.2%,能够泛化到未见过的案例,无需进一步优化即可迁移到其他模型,并且在非医学领域中也有效。
cs.AI / 9 / 2610.09243
We Query, Therefore We Compute: On Oracle Computation beyond the Machine, with an Application to Agents
我们查询,故我们计算:论超越机器的 Oracle 计算,及其在智能体中的应用
large language model
大语言模型相关
Abstract
Agentic systems use large language models (LLMs) to carry out concrete tasks. Prior work often borrows abstractions such as scheduling, caching or isolation piecemeal from operating systems, so the mechanisms it builds share little common ground, and the shared view of the two forms of agentic system, Workflows and Agents, is limited. We construct an abstract machine that provides both. We treat the LLM as an Oracle and extend a two-stack pushdown automaton with one instruction, which hands the Oracle a whole stack as its query and appends the answer to that same stack. The machine thus performs two computations, the Oracle's and a Turing-complete one that we call the Priestess. A stack that the program only appends to grows autoregressively, as an agent's context does. Two symmetry breakings, S in storage and T in transitions, make a Priestess program the operating system of the programs the Oracle runs, and produce the Agent and the Workflow as the two placements of a task's program. For internally autoregressive Oracles, the two computations synchronize at the end of every answer under certain conditions, and through that synchronization we model caching and analyse scheduling. No guarantee that holds for every Oracle can fix which content crosses between the two computations, but such a guarantee does fix the boundary itself. The construction V fits the machine to a von Neumann computer. To show that it is realizable, we propose ArchNights, an extended RISC-V ISA and a Linux-style operating system implementing the machine by design. ArchNights-SE runs on gem5 as a computer system, becomes an agentic system when it runs an LLM as the Oracle, and will be open source. Agentic systems can then be designed as computer systems are. With a foundation built and a unified view, future work can share invariants and bounds, each with its conditions.
Chinese Translation
智能体系统使用大语言模型(LLM)来执行具体任务。先前的工作常常从操作系统中零散地借用诸如调度、缓存或隔离之类的抽象,因此它所构建的机制之间鲜有共同基础,而且对两类智能体系统——Workflows 和 Agents——的共享视角也很有限。我们构造了一台同时提供二者的抽象机器。我们把 LLM 视为一个 Oracle,并为双栈下推自动机扩展一条指令;该指令将整个栈作为查询交给 Oracle,并把答案追加到同一个栈上。因此,这台机器执行两种计算:Oracle 的计算,以及一种图灵完备的计算,我们称之为 Priestess。一个程序只向其追加内容的栈会以自回归方式增长,正如智能体的上下文那样。两种对称性破缺——存储中的 S 和转移中的 T——使一个 Priestess 程序成为 Oracle 所运行程序的操作系统,并产生 Agent 和 Workflow,作为任务程序的两种放置方式。对于内部自回归的 Oracle,在特定条件下,这两种计算会在每个答案结束时同步;通过这种同步,我们对缓存进行建模,并分析调度。没有任何对每个 Oracle 都成立的保证能够确定哪些内容在两种计算之间跨越,但这样的保证确实固定了边界本身。构造 V 使这台机器适配冯·诺依曼计算机。为了表明它是可实现的,我们提出 ArchNights:一种扩展的 RISC-V ISA,以及一个 Linux 风格操作系统,二者按设计实现该机器。ArchNights-SE 作为一个计算机系统运行在 gem5 上,当它把 LLM 作为 Oracle 运行时就成为一个智能体系统,并且将会开源。这样一来,智能体系统就可以像计算机系统那样被设计。在基础已经建立且具有统一视角的情况下,未来的工作可以共享不变量和界,每一个都带有其自身条件。
cs.AI / 10 / 2610.09416
Efficient Reasoning with Flow Language Models
使用流语言模型的高效推理
diffusion
扩散模型相关
Abstract
Flow Language Models (FLMs) have emerged as a continuous-state alternative to discrete diffusion language models, yet the role of their continuous representations in reasoning remains unclear. We investigate this question by comparing the reasoning efficiency of FLMs and discrete diffusion models, measured by solution accuracy under matched denoising steps. Unlike discrete diffusion, which passes categorical states between denoising steps, FLMs evolve a continuous sequence representation throughout denoising and decodes it into discrete tokens only at the end. Our theoretical analysis shows, from a superposition perspective, how information retained in these continuous states can benefit reasoning. Intermediate-state interventions provide further empirical support for this theoretical account, showing that removing information about alternative candidates reduces subsequent solution recovery. Together, these findings show that FLMs allow evidence for multiple candidates to persist and inform subsequent reasoning before a discrete answer is produced. Furthermore, our experiments on maze planning and Sudoku tasks show that FLMs achieve greater reasoning efficiency in the few-step regime: FLMs achieves higher sequence accuracy than discrete diffusion baselines at matched model sizes and small denoising steps. On maze planning tasks, FLMs can also achieve comparable accuracy with smaller models. For example, on Maze15, FLM reaches the 95\% accuracy target at 64 denoising steps with 36.5\% fewer parameters than MDLM. These findings point to continuous state spaces as a promising foundation for reasoning models that require fewer refinement steps.
Chinese Translation
流语言模型(FLMs)已成为离散扩散语言模型的一种连续状态替代方案,然而其连续表示在推理中的作用仍不清楚。我们通过比较 FLMs 与离散扩散模型的推理效率来研究这一问题,其衡量标准是在匹配去噪步数下的解答准确率。与在去噪步骤之间传递类别状态的离散扩散不同,FLMs 在整个去噪过程中演化出一个连续的序列表示,并且仅在最后将其解码为离散词元。我们的理论分析从叠加的视角表明,保留在这些连续状态中的信息如何能够有益于推理。中间状态干预为该理论解释提供了进一步的经验支持,表明移除关于备选候选的信息会降低后续的解答恢复。这些发现共同表明,FLMs 允许多个候选者的证据持续存在,并在产生离散答案之前为随后的推理提供信息。此外,我们在迷宫规划和数独任务上的实验表明,FLMs 在少步机制下实现了更高的推理效率:在匹配的模型规模和较小的去噪步数下,FLMs 比离散扩散基线实现了更高的序列准确率。在迷宫规划任务上,FLMs 还能以更小的模型达到相当的准确率。例如,在 Maze15 上,FLM 在 64 个去噪步数下达到 95\% 的准确率目标,其参数量比 MDLM 少 36.5\%。这些发现表明,连续状态空间有望成为需要更少精化步骤的推理模型的基础。
cs.AI / 11 / 2610.09590
Learning Situation-Conditioned Thinking Policies for Long-Term LLM Agents
为长期 LLM 智能体学习情境条件化思维策略
large language model
大语言模型相关
Abstract
Long-running autonomous agents must reuse accumulated reasoning experience without allowing explicit historical memory and LLM context to grow indefinitely. However, existing memory mechanisms mainly retrieve, summarize, or compress past content and do not directly learn when particular kinds of thinking should be activated or discover new thinking knowledge from temporally dispersed experiences. This paper proposes a situation-conditioned thinking memory framework that transforms historical reasoning experience into a lightweight policy for predicting what should be thought about in the current situation, while leaving detailed reasoning to a large language model. Situations may represent temporal or spatiotemporal evolution rather than only current states. Temporary experiences are also periodically analyzed across multiple independent episodes to identify repeated long-range regularities, which are consolidated into new thinking knowledge and further internalized by the lightweight policy. Experiments show that the learned policy achieves 1.000 F1 on temporal-rule generalization, improves DeepSeek reasoning F1 from 0.789 to 0.868, reduces online processing time from 0.3636 ms to 0.0382 ms per query at 30,000 historical situations, and reaches 1.000 relation-discovery F1 and future-thinking accuracy after sufficient repeated cross-experience evidence.
Chinese Translation
长期运行的自主智能体必须复用累积的推理经验,同时不允许显式的历史记忆和 LLM 上下文无限增长。然而,现有的记忆机制主要检索、总结或压缩过去的内容,并不直接学习特定类型的思维应在何时被激活,也不从时间上分散的经验中发现新的思维知识。本文提出一个情境条件化的思维记忆框架,该框架将历史推理经验转化为一个轻量级策略,用于预测在当前情境下应当思考什么,同时将详细推理留给大语言模型。情境可以表示时间或时空演化,而不仅仅是当前状态。临时经验还会跨多个独立回合被周期性地分析,以识别重复的长期规律,这些规律被巩固为新的思维知识,并进一步由轻量级策略内化。实验表明,所学策略在时间规则泛化上达到 1.000 F1,将 DeepSeek 推理 F1 从 0.789 提升到 0.868,在 30,000 个历史情境下将每次查询的在线处理时间从 0.3636 ms 降低到 0.0382 ms,并在有足够的重复跨经验证据后达到 1.000 的关系发现 F1 和未来思维准确率。
cs.AI / 12 / 2610.09600
SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models
SafeEvo:解密语言模型中的安全对齐机制与演化
large language model
大语言模型相关
Abstract
Safety interpretability advances the study of Large Language Model (LLM) alignment from behavioral constraints driven by data or algorithms towards a deeper understanding of internal mechanisms. However, existing works have focused primarily on safety-related representations, attention heads, or neurons after alignment, while largely overlooking the safety mechanisms in pretrained-only models and their evolution across alignment checkpoints. To address this, we propose SafeEvo, an interpretability framework from the circuit (sparse subgraphs of an LLM) perspective. SafeEvo first applies an optimization-based extraction algorithm to identify weak refusal circuits in pretrained base LLMs that can independently express refusal behavior. Causally ablating these circuits completely eliminates the base model's refusal of harmful inputs. SafeEvo then traces the evolution of refusal circuits across successive alignment checkpoints and finds that their structures change progressively, suggesting that the alignment tax may result from refusal-circuit updates affecting utility-related parameters. To validate this, SafeEvo introduces Safety Circuit Alignment (SCA), which confines safety updates to the refusal circuits. Experiments across three LLMs and two alignment algorithms show that, on average, SCA outperforms vanilla alignment in three aspects: \textbf{(1) stronger alignment}, lowering harmfulness score by 63.21\%; \textbf{(2) less over-refusal}, yielding a 58.44\% decrease in refusal rates for benign queries; and \textbf{(3) better utility}, retaining 99.58\% of the original model capabilities.
Chinese Translation
安全可解释性将大型语言模型(LLM)对齐的研究,从由数据或算法驱动的行为约束,推进到对内部机制的更深入理解。然而,现有工作主要关注对齐之后与安全相关的表征、注意力头或神经元,而在很大程度上忽视了对齐前的纯预训练模型中的安全机制,以及它们在不同对齐检查点之间的演化。为解决这一问题,我们提出了 SafeEvo,一个从电路(LLM 的稀疏子图)视角出发的可解释性框架。SafeEvo 首先应用一种基于优化的提取算法,在预训练基础 LLM 中识别出能够独立表达拒绝行为的弱拒绝电路。因果性地消融这些电路会完全消除基础模型对有害输入的拒绝。SafeEvo 随后追踪拒绝电路在连续对齐检查点上的演化,发现它们的结构逐步发生变化,这表明对齐税可能源于拒绝电路的更新影响了与效用相关的参数。为验证这一点,SafeEvo 引入了安全电路对齐(SCA),它将安全更新限制在拒绝电路之内。在三个 LLM 和两种对齐算法上的实验表明,平均而言,SCA 在三个方面优于原始对齐:\textbf{(1) 更强的对齐},将有害性得分降低 63.21\%;\textbf{(2) 更少的过度拒绝},使良性查询的拒绝率下降 58.44\%;以及 \textbf{(3) 更好的效用},保留了原始模型能力的 99.58\%。
cs.AI / 13 / 2610.09681
Cost-Efficient Theorem Proving via Agent Orchestration in Program Verification
程序验证中通过智能体编排实现的成本高效定理证明
large language model
大语言模型相关
Abstract
Program verification establishes software correctness through machine-checkable proofs constructed in theorem provers. It's a guarantee especially valuable for code generated by large language models (LLMs), which is fluent but carries no assurance of correctness. Almost all existing provers, however, pursue pass rates alone at whatever sampling or search budget it takes, and overlook the success-vs-cost frontier; yet real software often carries hundreds of interdependent proof obligations, so what matters at scale is not whether one theorem can be proved, but how many can be proved economically. We introduce CoCo-Prover, which formalizes cost-efficient program proving as metalevel decision-making under cost, grounded on two-level proof graphs: an AND/OR proof hypergraph within each declaration is joined to a lemma-dependency graph across declarations; and at each step, it answers two questions: which open goals to select, and which actions to purchase on these goals. Selection stays symbolic as a topological pass over the proof graphs. Action choice is agent orchestration via metalevel decision-making: an agentic router treats every bounded specialist invocation as a separately priced, best-effort computation, matching heterogeneous specialist agents together with configurations, under evolved routing rules as evidence accumulates. On five program verification benchmarks in Lean 4 including function-level CLEVER, VERINA, and AlgoVeri, and repository-level NTP4VC and Vero, we show that CoCo-Prover achieves a better success-vs-cost frontier than baselines including frontier coding agents and state-of-the-art LLM-based provers: it achieves the best solve rate on every benchmark and up to 100% on two benchmarks. It also reduces cost by up to 30.9% compared to the strongest baseline with the strongest LLM in our evaluation.
Chinese Translation
程序验证通过在定理证明器中构造的机器可检查的证明来确立软件的正确性。这一保证对于大型语言模型(LLM)生成的代码尤为重要,因为这类代码虽然流畅,却不带任何正确性的保障。然而,几乎所有现有的证明器都只追求通过率,不惜消耗任何采样或搜索预算,却忽视了成功率与成本之间的前沿;而真实软件往往承载着数百个相互依赖的证明义务,因此在规模化场景下,关键并不在于某一个定理能否被证明,而在于有多少定理能够被经济地证明。我们提出 CoCo-Prover,它将成本高效的程序证明形式化为成本下的元级决策,并以两层证明图为基础:每个声明内部的 AND/OR 证明超图与跨声明的引理依赖图相连接;在每一步,它回答两个问题:选择哪些未完成目标,以及在这些目标上购买哪些动作。目标选择保持符号化,即对证明图进行一次拓扑遍历。动作选择则是通过元级决策进行的智能体编排:一个智能体路由器将每一次有界的专家调用视为一次单独定价的、尽力而为的计算,在随着证据积累而演化的路由规则下,将异构的专家智能体与配置相匹配。在 Lean 4 中的五个程序验证基准上,包括函数级的 CLEVER、VERINA 和 AlgoVeri,以及仓库级的 NTP4VC 和 Vero,我们表明 CoCo-Prover 实现了比包括前沿编码智能体和最先进的基于 LLM 的证明器在内的基线更好的成功率与成本前沿:它在每个基准上都取得了最佳的求解率,并在两个基准上达到了最高 100%。在我们的评估中,与使用最强 LLM 的最强基线相比,它还将成本降低了最多 30.9%。
cs.AI / 14 / 2610.09769
From Expert-Guided Proof Search to Automated Open-Problem Solving
从专家引导的证明搜索到自动化开放问题求解
large language model
大语言模型相关
Abstract
Large language models are increasingly contributing to mathematical research, where progress often depends on efficient proof search, incremental improvements and careful verification. We describe Bolzano, a multi-agent open-source system that uses parallel prover agents with a verifier agent and maintains a human-readable research state. Initial manual use on expert-selected problems yielded 8 results whose proofs were checked by domain experts. Motivated by these case studies, we ran Bolzano without problem-specific human guidance on about 3,800 open problems extracted from four sets of papers, solving about 200 open problems. One experiment used papers accepted to STOC 2026, a top conference in theoretical computer science. There, we answered four questions raised in the papers, as confirmed by their authors.
Chinese Translation
大语言模型正日益为数学研究做出贡献,而在数学研究中,进展往往取决于高效的证明搜索、增量式改进和细致的验证。我们介绍 Bolzano,一个多智能体开源系统,它使用并行的证明器智能体与一个验证器智能体,并维护一个人类可读的研究状态。在专家选定的问题上进行的初步人工使用产生了 8 项结果,其证明由领域专家检查。受这些案例研究的启发,我们在从四组论文中提取的约 3,800 个开放问题上,在没有针对具体问题的人工引导下运行了 Bolzano,解决了约 200 个开放问题。其中一项实验使用了被 STOC 2026 接收的论文,STOC 2026 是理论计算机科学领域的顶级会议。在那里,我们回答了论文中提出的四个问题,并得到了论文作者的确认。
cs.AI / 15 / 2610.09876
Think Before You Paint: Recursive Latent Reasoning for Diffusion Models
先思后绘:面向扩散模型的递归潜在推理
diffusion
扩散模型相关
Abstract
Diffusion models generate realistic images but often fail on visual reasoning tasks, such as filling in a Sudoku or drawing the path through a maze. When a discrete symbolic representation is available, recursive methods such as the Tiny Recursive Model (TRM) solve even hard instances of these puzzles. We ask how such reasoning can be carried over to pixels, where no symbolic representation is available. We propose Painter-Thinker (PaTh): a small recursive network (the Thinker) reasons over a grid of learned tokens that encode the noisy image and the conditioning, refines a latent state within every denoising step, and steers a frozen diffusion model (the Painter) through ControlNet adapters. The Thinker is trained with the standard reconstruction loss alone, without symbolic targets, a solver, or a verifier. PaTh solves 92.5% of hard MNIST Sudoku puzzles (prior best 75%) and 71.2% of extreme ones (prior best 4.1%), with 10M parameters against 82M for a standard diffusion model. It also improves on mazes, Queens, and CLEVR scenes with specified spatial relations, and its advantage grows with problem size. Diagnostic experiments show that PaTh recovers from injected mistakes that the diffusion model cannot repair, especially when many cells are wrong. Together, these results show that reasoning mechanisms developed for symbolic data can be integrated into pixel-space diffusion without symbolic supervision, opening a path toward generating data under increasingly complex constraints.
Chinese Translation
扩散模型能生成逼真的图像,但往往在视觉推理任务上失败,例如填写数独或画出穿过迷宫的路径。当存在离散符号表示时,诸如微型递归模型(Tiny Recursive Model, TRM)之类的递归方法甚至能解决这些谜题的困难实例。我们提出疑问:如何将这种推理迁移到像素上,而在像素上并没有可用的符号表示。我们提出 Painter-Thinker(PaTh):一个小型递归网络(即 Thinker)在一组学习到的 token 网格上进行推理,这些 token 编码了噪声图像和条件信息,它在每个去噪步骤内细化一个潜在状态,并通过 ControlNet 适配器引导一个冻结的扩散模型(即 Painter)。Thinker 仅使用标准重建损失进行训练,不需要符号目标、求解器或验证器。PaTh 解决了 92.5% 的困难 MNIST 数独谜题(此前最佳为 75%)和 71.2% 的极端谜题(此前最佳为 4.1%),其参数量为 10M,而标准扩散模型为 82M。它还在迷宫、Queens 和具有指定空间关系的 CLEVR 场景上有所改进,并且其优势随着问题规模增大而增长。诊断实验表明,PaTh 能够从注入的错误中恢复,而扩散模型无法修复这些错误,尤其是在许多单元格出错时。总之,这些结果表明,为符号数据开发的推理机制可以集成到像素空间扩散中,而无需符号监督,从而开辟了一条在日益复杂的约束下生成数据的道路。
cs.AI / 16 / 2610.09964
Successive Training Stages and Large Language Model Persuasion: Effects of Misalignment, Supervised Fine-Tuning, and Preference Optimization
相继训练阶段与大型语言模型说服力:不对齐、监督微调和偏好优化的影响
large language model
大语言模型相关
Abstract
Large language models (LLMs) can be tuned to influence human attitudes, yet the respective contributions of successive post-training stages remain un-clear. This study examines how three successive training stages affect LLM persuasiveness: (1) misalignment through supervised fine-tuning (SFT) on conspiracy data, (2) additional persuasive SFT on argumentative data, and (3) Identity Preference Optimization (IPO), a preference-optimization method. A total of 835 participants recruited on Prolific were randomly assigned to five between-subject conditions (neutral text, conspiracy-trained model, persuasion-trained model, preference-optimized model, and GPT-4) and were exposed to texts on 10 divisive political issues, personalized from their individual profiles in all model conditions. Attitude change was measured as the difference between pre- and post-exposure positions on continuous Likert scales and analyzed with an analysis of covariance (ANCOVA). A significant condition x baseline-attitude interaction, F (4, 825) = 5.33, p < .001, indicated that training effects depended on participants' initial attitudes. Persuasive SFT produced greater attitude change than conspiracy training alone, d = 0.30, whereas IPO provided no additional benefit, d = 0.03, and GPT-4 did not differ from neutral text, d = --0.01. These results show that targeted supervised training on persuasive data increases LLM persuasiveness, whereas preference optimization yields no significant gains beyond it.
Chinese Translation
大型语言模型(LLMs)可以被调优以影响人类态度,但相继的后期训练阶段各自的贡献仍不清楚。本研究考察三个相继训练阶段如何影响 LLM 的说服力:(1)通过基于阴谋论数据的监督微调(SFT)实现的不对齐,(2)基于论证性数据的额外说服性 SFT,以及(3)身份偏好优化(IPO),一种偏好优化方法。在 Prolific 上招募的共 835 名参与者被随机分配到五种被试间条件(中性文本、阴谋论训练模型、说服训练模型、偏好优化模型和 GPT-4),并在所有模型条件下接触了关于 10 个有争议政治议题的文本,这些文本根据其个人资料进行了个性化定制。态度变化被测量为暴露前和暴露后在连续李克特量表上的立场差异,并用协方差分析(ANCOVA)进行分析。一个显著的条件 × 基线态度交互作用,F (4, 825) = 5.33, p < .001,表明训练效应取决于参与者的初始态度。说服性 SFT 比单独的阴谋论训练产生了更大的态度变化,d = 0.30,而 IPO 没有提供额外益处,d = 0.03,并且 GPT-4 与中性文本没有差异,d = --0.01。这些结果表明,针对说服性数据的定向监督训练提高了 LLM 的说服力,而偏好优化并未在其基础上产生显著增益。
cs.AI / 17 / 2610.10042
Learning to Accumulate Knowledge with Mutual Information
利用互信息学习积累知识
large language model
大语言模型相关
Abstract
Large language model (LLM) agents can improve their performance by reusing knowledge distilled from past interactions. However, curating new experiences into a knowledge bank that becomes more useful as it grows remains challenging. Effective knowledge accumulation should limit redundant overlap among entries and ensure that new knowledge contributes beyond what the bank already provides. Yet training a curator with Group Relative Policy Optimization (GRPO) on standalone task success can reinforce general guidance even when it duplicates existing knowledge. Therefore, we propose Knowledge Weaver, a reinforcement learning framework that trains a language model to curate reusable knowledge from agent trajectories. We couple feedback inspired by token-wise mutual information (MI) with marginal success rewards to guide knowledge accumulation. Together, these signals encourage the curator to preserve distinct information from experience and produce entries that improve task success when added to existing knowledge. Standalone success rewards also favor entries that are useful on their own. On ALFWorld and WebShop, Knowledge Weaver achieves mean success rates of 54.0\% and 42.0\% with k=10 retrieved entries, exceeding GRPO by 16.9 and 18.7 percentage points, respectively. Its knowledge banks also outperform the evaluated prompt-based and established banks, including human-written banks, in overall ALFWorld success rate and WebShop score with the executor frozen. Our codebase is available at https://github.com/LaoKuiZe/Knowledge-Weaver.
Chinese Translation
大型语言模型(LLM)智能体可以通过复用从过去交互中提炼的知识来提升其性能。然而,将新经验整理进一个随其增长而变得更有用的知识库仍然具有挑战性。有效的知识积累应限制条目之间的冗余重叠,并确保新知识带来超出知识库已提供内容的贡献。然而,仅以独立任务成功为目标,用群组相对策略优化(GRPO)训练整理器,可能会强化通用指导,即使其重复了已有知识。因此,我们提出 Knowledge Weaver,一个强化学习框架,训练语言模型从智能体轨迹中整理可复用知识。我们将受逐 token 互信息(MI)启发的反馈与边际成功奖励相结合,以指导知识积累。这些信号共同鼓励整理器保留来自经验中的独特信息,并生成在加入现有知识后能提升任务成功的条目。独立成功奖励也偏好那些自身就有用的条目。在 ALFWorld 和 WebShop 上,Knowledge Weaver 在检索 k=10 个条目时分别达到 54.0\% 和 42.0\% 的平均成功率,分别超过 GRPO 16.9 和 18.7 个百分点。在冻结执行器的情况下,其知识库在总体 ALFWorld 成功率和 WebShop 分数上也优于所评估的基于提示的知识库和已有知识库,包括人工编写的知识库。我们的代码库可在 https://github.com/LaoKuiZe/Knowledge-Weaver 获取。
cs.AI / 18 / 2610.10164
UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy
UniSkill:为不断演化的策略学习与执行者对齐的技能提议
large language model
大语言模型相关
Abstract
Large language model agents can improve across tasks by retaining reusable skills distilled from prior interactions. Recent work jointly optimizes task execution and skill extraction, enabling the policy and skillbank to co-evolve. However, as the actor continues learning, rewarding skill proposals through their reuse in subsequent training steps may conflate skill benefits with actor improvement, while directly testing each proposed skill requires costly additional actor rollouts. In this paper, we introduce UniSkill, which uses a shared policy to interact with the environment and propose skillbank edits (Add, Update, or No Edit) from the resulting trajectories. Specifically, the actor learns from environment rewards, while contrastive action feedback guides skill proposal learning. This feedback provides an actor-alignment signal by measuring how replacing the retrieved skill with a proposed skill changes the current actor's action log-likelihood gap between previously collected successful and failed trajectories from the same task, thereby avoiding new rollouts for each proposal. Since proposal-level feedback may suppress an otherwise appropriate edit operation when the proposed skill content scores poorly, we further apply skill-edit support regularization to preserve exploration. Empirically, UniSkill achieves strong performance, reaching 98.4% success on ALFWorld and 84.7% on WebShop while maintaining stable joint training. Further ALFWorld experiments show that UniSkill remains effective when the shared policy uses a smaller backbone. Our implementation is available at https://github.com/LimOkii/UniSKill.
Chinese Translation
大型语言模型智能体可以通过保留从先前交互中蒸馏出的可复用技能,在跨任务上取得提升。近期工作联合优化任务执行与技能提取,使策略与技能库能够共同演化。然而,随着执行者持续学习,通过在后续训练步骤中对技能提议的复用来奖励它们,可能会将技能收益与执行者改进混为一谈,而直接测试每个提出的技能则需要代价高昂的额外执行者 rollout。在本文中,我们提出 UniSkill,它使用一个共享策略与环境交互,并从所得轨迹中提议技能库编辑(Add、Update 或 No Edit)。具体而言,执行者从环境奖励中学习,而对比式动作反馈引导技能提议学习。这种反馈通过衡量以下变化来提供执行者对齐信号:将检索到的技能替换为提议技能后,当前执行者在同一任务先前收集的成功轨迹与失败轨迹之间的动作对数似然差距如何变化,从而避免为每个提议进行新的 rollout。由于当提议技能内容得分较差时,提议级反馈可能会抑制一个本来合适的编辑操作,我们进一步应用技能编辑支持正则化以保持探索。在实证上,UniSkill 取得了强劲性能,在 ALFWorld 上达到 98.4% 的成功率,在 WebShop 上达到 84.7%,同时保持稳定的联合训练。进一步的 ALFWorld 实验表明,当共享策略使用更小的骨干网络时,UniSkill 仍然有效。我们的实现可在 https://github.com/LimOkii/UniSKill 获取。
cs.AI / 19 / 2610.10184
Agentic AI-Assisted Modeling for Production Scheduling: Assessment in Constraint Programming
用于生产调度的智能体AI辅助建模:在约束编程中的评估
large language model
大语言模型相关
Abstract
Developing optimization models for production scheduling requires substantial expert effort. Research on large language models (LLMs) has followed two directions: specialized approaches for automated modeling, mostly for mixed-integer linear programming, which often rely on dedicated training or problem-specific architectures that limit industrial deployment; and agentic artificial intelligence for operational decision support, which generally assumes that the optimization model already exists. This study bridges both directions by assessing whether general-purpose LLMs, orchestrated as agents without task-specific training, can formulate and implement constraint programming models from natural-language problem descriptions. Singleagent and multi-agent architectures are integrated with a Model Context Protocol server that provides context-aware retrieval of solver documentation to mitigate hallucinations during implementation. Both are compared with a direct LLM baseline on six industry-oriented problems covering flow-shop, job-shop, flexible job-shop and resource-constrained warehouse scheduling, using three LLMs and assessing modeling accuracy, execution success, latency and token consumption. Formulation proves largely within reach of current LLMs, whereas implementation is the main barrier. The multi-agent workflow raises the share of scripts that run correctly as generated from 14.8% with a direct LLM call to 59.3%, reaching 80.6% on the four less complex problems, while tightly coupled intralogistics models remain an open challenge.
Chinese Translation
开发用于生产调度的优化模型需要大量的专家努力。关于大语言模型(LLMs)的研究已遵循两个方向:用于自动建模的专门方法,大多针对混合整数线性规划,这些方法通常依赖于专用训练或特定问题的架构,从而限制了工业部署;以及用于运营决策支持的智能体人工智能,其通常假设优化模型已经存在。本研究通过评估通用LLMs在未经任务特定训练的情况下被编排为智能体时,能否根据自然语言问题描述构建并实现约束编程模型,来桥接这两个方向。单智能体和多智能体架构与一个模型上下文协议服务器集成,该服务器提供求解器文档的上下文感知检索,以减轻实现过程中的幻觉。两者都与直接LLM基线在六个面向工业的问题上进行比较,这些问题涵盖流水车间、作业车间、柔性作业车间和资源受限仓库调度,使用三个LLM并评估建模准确性、执行成功率、延迟和token消耗。事实证明,建模在很大程度上处于当前LLMs的能力范围之内,而实现则是主要障碍。多智能体工作流将按生成内容即可正确运行的脚本比例从直接LLM调用的14.8%提高到59.3%,在四个较简单问题上达到80.6%,而紧耦合的内部物流模型仍然是一个未决挑战。
cs.AI / 20 / 2610.10358
Open-MMUnlearning: Unifying Methods and Evaluation for MLLM Unlearning
Open-MMUnlearning:统一 MLLM 遗忘的方法与评估
large language model
大语言模型相关
Abstract
As multimodal large language models (MLLMs) become more capable and widely deployed, concerns about privacy and safety have become increasingly pressing. Machine unlearning offers one approach to addressing these concerns by removing designated information from trained models while preserving unrelated capabilities. However, fragmented implementations and evaluation protocols, incomplete robustness testing, and limited understanding of metric reliability make progress in MLLM unlearning difficult to assess systematically. We introduce Open-MMUnlearning, an open-source, extensible framework that integrates target-model preparation, multimodal data processing, unlearning, and evaluation through shared interfaces and structured configurations. The framework supports five benchmarks spanning privacy, safety, and copyright, eight MLLMs from four model families, and twelve unlearning methods. Its evaluation suite jointly assesses forgetting effectiveness, retained utility, and robustness to model interventions, adversarial inputs, and membership inference attacks. Using a common evaluation protocol, we compare ten representative unlearning methods. In this comparison, GD and MIP-Editor tie for the highest overall score: GD achieves the highest Forget Quality, while MIP-Editor preserves more Model Utility. We further introduce a metric meta-evaluation protocol that tests faithfulness using models with controlled exposure to target knowledge and robustness under quantization and relearning. Among the thirteen evaluated metrics, BLEU achieves the highest aggregate reliability score. KS-Test attains the highest faithfulness AUC but performs less well on robustness. Together, the framework and these findings support reproducible comparison of MLLM unlearning methods and systematic assessment of evaluation reliability.
Chinese Translation
随着多模态大语言模型(MLLMs)能力日益增强且得到广泛部署,对隐私和安全的担忧也变得越来越迫切。机器遗忘提供了一种应对这些担忧的方法:从已训练模型中移除指定信息,同时保留不相关的能力。然而,碎片化的实现与评估协议、不完整的鲁棒性测试,以及对指标可靠性的有限理解,使得 MLLM 遗忘领域的进展难以被系统性地评估。我们提出 Open-MMUnlearning,这是一个开源、可扩展的框架,它通过共享接口和结构化配置,将目标模型准备、多模态数据处理、遗忘和评估集成在一起。该框架支持涵盖隐私、安全和版权的五个基准、来自四个模型家族的八个 MLLM,以及十二种遗忘方法。其评估套件联合评估遗忘有效性、保留效用,以及对模型干预、对抗输入和成员推理攻击的鲁棒性。使用一个共同的评估协议,我们比较了十种具有代表性的遗忘方法。在这项比较中,GD 和 MIP-Editor 在最高总分上并列:GD 取得了最高的遗忘质量(Forget Quality),而 MIP-Editor 保留了更多的模型效用(Model Utility)。我们进一步引入一个指标元评估协议,该协议使用对目标知识具有受控暴露程度的模型来测试忠实性,并在量化和再学习下测试鲁棒性。在评估的十三种指标中,BLEU 取得了最高的综合可靠性得分。KS-Test 取得了最高的忠实性 AUC,但在鲁棒性方面表现较差。总而言之,该框架与这些发现支持 MLLM 遗忘方法的可复现比较,以及对评估可靠性的系统性评估。
cs.AI / 21 / 2610.10405
Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models
大型语言模型中受提示的不真实回应下的推理 token 激增
large language model
大语言模型相关
Abstract
Monitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models. However, semantic chain-of-thought monitoring depends on reasoning traces being legible and sufficiently faithful to the underlying computations that produced the model's behavior, not to mention accessible. Moreover, there is increasing evidence that chain-of-thought outputs may soon become illegible or unfaithful, if they even remain accessible. Based on cognitive load theory, we investigate a lower-bandwidth signal -- the number of reasoning tokens generated -- which does not require access to the content of the reasoning trace. Three reasoning-capable large language models answered 210 multiple-choice questions -- across analytic, descriptive, and normative reasoning types as well as moral and non-moral domains -- under system prompts instructing them to respond truthfully, falsely, or without regard for truth. Across all three models, truth-directed responding elicited fewer reasoning tokens than both lie-directed and truth-indifferent responding. These findings show that explicitly prompted untruthful response policies can produce robust group-level differences in test-time reasoning-token use. While not yet establishing reasoning-token count as a detector of spontaneous deception or general misalignment, our results are a proof of concept that it can serve as a simple, content-independent candidate signal for differentiating untruthful from truthful model behavior when raw reasoning traces are unavailable or unreliable. Future work should test instance-level detection rates, out-of-distribution generalization, learned deceptive policies, hidden objectives, and robustness under adversarial pressure.
Chinese Translation
监控推理型人工智能(AI)模型的思维链,仍然是检测此类模型中的欺骗及其他形式不当行为的关键方法。然而,语义思维链监控依赖于推理轨迹清晰可读、且对产生模型行为的底层计算足够忠实,更不用说还要可访问。此外,越来越多的证据表明,思维链输出可能很快就会变得不可读或不忠实,如果它们还能保持可访问的话。基于认知负荷理论,我们研究一种更低带宽的信号——所生成的推理 token 数量——它不需要访问推理轨迹的内容。三个具备推理能力的大型语言模型回答了 210 道多项选择题——涵盖分析型、描述型和规范型推理类型以及道德和非道德领域——在指示它们真实、虚假或不考虑真伪地回应的系统提示下。在所有三个模型中,以真实为导向的回应所引发的推理 token 数量都少于以说谎为导向和对真伪漠不关心的回应。这些发现表明,明确提示的不真实回应策略可以在测试时推理 token 使用上产生稳健的组级差异。虽然尚未确立推理 token 数量作为自发性欺骗或一般性不对齐的检测器,但我们的结果是一个概念验证,表明当原始推理轨迹不可用或不可靠时,它可以作为一种简单、内容无关的候选信号,用于区分不真实与真实的模型行为。未来工作应测试实例级检测率、分布外泛化、习得的欺骗策略、隐藏目标,以及对抗压力下的稳健性。
cs.AI / 22 / 2610.10506
Validity Without Ground Truth: What Stated-Preference Economics Offers the Evaluation of Language Models
无真值的效度:陈述偏好经济学为语言模型评估提供了什么
large language model
大语言模型相关
Abstract
Many of the questions now put to large language models have no correct answer to score against: what a policy is worth, which option a user should choose, how to weigh competing values. Stated-preference economics has faced this problem for decades. It judges survey responses without knowing the true value, through a framework of validity and related concepts: content, construct, and criterion validity, reliability, incentive compatibility, and consequentiality. We argue that this framework is a general method for evaluating language models, and we set out what each concept means for LLM evaluation. We demonstrate the approach using a published water-quality stated preference economic valuation survey (Vossler et al. 2023) administered to six models. In this economic application, the validity tests take the form of predictions from economic theory: demand should slope down, and willingness to pay should respond to the scope of the good and to income. The tests separate the models sharply. Two older models fail the most basic test at a household income level of \$75,000, and the two newest pass every test of theoretical validity we can score, but diverge on convergent validity. Passing validity tests shows that a model's answers are coherent, not that they are correct.
Chinese Translation
如今向大型语言模型提出的许多问题都没有可供评分的正确答案:一项政策价值几何、用户应当选择哪个选项、如何权衡相互竞争的价值观。陈述偏好经济学已经面对这一问题数十年。它在不知道真实价值的情况下评判调查回答,所借助的是一套效度及相关概念的框架:内容效度、结构效度与效标效度、信度、激励相容性以及后果性。我们认为,这一框架是评估语言模型的一种通用方法,并阐述了每个概念对 LLM 评估意味着什么。我们使用一项已发表的水质陈述偏好经济价值评估调查(Vossler et al. 2023),将其施测于六个模型,以演示该方法。在这一经济学应用中,效度检验采取经济理论预测的形式:需求应当向下倾斜,支付意愿应当对物品的范围和收入作出反应。这些检验将各模型清晰地区分开来。两个较旧的模型在家庭收入水平为 \$75,000 时未能通过最基本的检验,而两个最新的模型通过了我所能评分的每一项理论效度检验,但在收敛效度上出现分歧。通过效度检验表明模型的回答是连贯的,而非表明它们是正确的。
cs.AI / 23 / 2610.10507
RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing
RECAST:学习通过自适应证据路由计算正确的上下文
large language model
大语言模型相关
Abstract
Large language models are increasingly applied to tasks grounded in long, heterogeneous information sources. Conventional Retrieval-Augmented Generation (RAG) relies on fixed similarity-based retrieval, while agentic variants adapt queries and tool use but remain largely retrieval-centric. However, in many tasks, the evidence required for a solution is not explicitly present in any single source item. Instead, it must be derived through filtering, aggregation, or computation across multiple source items. In this work, we introduce RECAST (Routing Evidence through Computation, Access, and Synthesized Tools), a learned framework that formulates evidence construction as a sequential decision process over heterogeneous retrieval and computation operations, allowing evidence to be actively derived rather than merely retrieved. A lightweight RouterLM iteratively selects and formulates primitive operations or specifies customized operations for a frozen CompilerLM to translate into executable code. Once it judges the evidence sufficient, RouterLM passes the accepted evidence to a frozen AnswerLM to produce the final solution. We train RouterLM with supervised fine-tuning (SFT) followed by group relative policy optimization (GRPO). Across six heterogeneous benchmark families, RECAST achieves a mean success rate of 75.6%, outperforming the strongest large-model baseline by 15.9%. Moreover, training enables the Qwen3.5-9B RouterLM to outperform a training-free Gemini 3.5 Flash RouterLM by 5.0%. On three held-out benchmarks, RECAST improves over the strongest baseline by 15.0% on average, demonstrating strong zero-shot generalization across tasks and heterogeneous source representations.
Chinese Translation
大语言模型越来越多地被应用于以长篇幅、异构信息源为基础的任务。传统的检索增强生成(RAG)依赖于固定的基于相似度的检索,而智能体式变体虽然会调整查询和工具使用,但在很大程度上仍以检索为中心。然而,在许多任务中,解决方案所需的证据并不显式地存在于任何单个源项中。相反,它必须通过跨多个源项的过滤、聚合或计算来推导得到。在这项工作中,我们提出了 RECAST(Routing Evidence through Computation, Access, and Synthesized Tools,即通过计算、访问和合成工具来路由证据),这是一个学习式框架,它将证据构建形式化为对异构检索和计算操作的序列决策过程,使证据能够被主动推导出来,而不仅仅是被检索到。轻量级 RouterLM 迭代地选择并构造原始操作,或者为冻结的 CompilerLM 指定定制操作,以将其翻译成可执行代码。一旦它判断证据已充分,RouterLM 就将被接受的证据传递给冻结的 AnswerLM,以生成最终解决方案。我们采用监督微调(SFT)随后进行组相对策略优化(GRPO)来训练 RouterLM。在六个异构基准系列上,RECAST 取得了 75.6% 的平均成功率,比最强的大模型基线高出 15.9%。此外,训练使 Qwen3.5-9B RouterLM 比免训练的 Gemini 3.5 Flash RouterLM 高出 5.0%。在三个留出基准上,RECAST 平均比最强基线提升 15.0%,展现出跨任务和异构源表示的强大零样本泛化能力。
cs.CL / 24 / 2610.09033
Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs
开放权重大型语言模型在非规范输入上的四态安全评估
large language model
大语言模型相关
Abstract
Standard safety evaluations of large language models assess harmful requests written in canonical plain text, while models in real-world deployment routinely receive inputs containing emojis, altered spellings, encoded strings, and character-level variations. This work introduces the Adversarial Surface-Form Robustness Dataset (ASRD), comprising 2,100 prompts across seven distinct surface-form families. Five open-weight language models are evaluated across these prompts, producing 10,500 responses. The Quad-State Evaluation Rubric classifies each response into one of four outcomes: harmful compliance, safe response, comprehension failure, or indeterminate. Emoji and invisible Unicode variations cause almost no comprehension failure, with pooled harmful compliance of 20.27% and 17.20% against a 22.87% baseline that is driven mainly by Mistral 7B, whereas leetspeak, encoded wrappers, and hybrid transformations score 2.40%, 0.13%, and 2.40% while comprehension failure rises to 36.47%, 65.60%, and 34.47%. Inspection of raw model outputs reveals three response behaviors: hallucinated benignity, structural collapse, and language drift. Project page: www.pavanmaddula.com/quadstate
Chinese Translation
大型语言模型的标准安全评估评估的是以规范纯文本编写的有害请求,而现实部署中的模型经常接收包含表情符号、拼写变体、编码字符串和字符级变异的输入。本工作引入了对抗性表层形式鲁棒性数据集(Adversarial Surface-Form Robustness Dataset,ASRD),其包含跨越七个不同表层形式族的 2,100 条提示。五个开放权重语言模型在这些提示上接受评估,产生 10,500 条响应。四态评估量表将每个响应分类为四种结果之一:有害遵从、安全响应、理解失败或不确定。表情符号和不可见 Unicode 变异几乎不造成理解失败,其汇总有害遵从率分别为 20.27% 和 17.20%,对照基线为 22.87%,该基线主要由 Mistral 7B 驱动;而 leetspeak、编码包装器和混合变换的得分为 2.40%、0.13% 和 2.40%,同时理解失败上升到 36.47%、65.60% 和 34.47%。对原始模型输出的检查揭示了三种响应行为:幻觉式良性、结构崩溃和语言漂移。项目页面:www.pavanmaddula.com/quadstate
cs.CL / 25 / 2610.09087
U-Space: Uncovering When and Why Uncertainty Arises in Language Models
U-Space:揭示语言模型中的不确定性何时以及为何产生
large language model
大语言模型相关
Abstract
Large language models are informing decisions with ever-higher stakes. As the consequences of their errors grow, a central question becomes harder to ignore: how much can we trust an individual answer? Yet recognizing when to defer remains difficult because language models can present incorrect conclusions with fluent explanations and an authoritative tone. Uncertainty quantification seeks to address this disconnect by estimating the reliability of individual predictions. However, many existing methods require repeated generations or separately trained components, and their scalar estimates do not reveal where uncertainty arises or how it evolves during reasoning. Recent work has also shown that generation length can be strongly associated with uncertainty estimates and correctness, raising the question of how much of an estimator's predictive power comes from uncertainty-specific information rather than output length alone. Mechanistic interpretability offers a way to address these limitations by connecting human-interpretable concepts to intermediate model states. Building on this capability, we introduce the U-Space, a low-dimensional subspace that makes a model's evolving uncertainty measurable and interpretable. We identify semantic anchors for doubt and certainty, map their unembedding directions back into the residual space, and combine their contrasts into an orthogonal basis. The U-Lens projects each token state onto these basis vectors, yielding an interpretable token-level uncertainty map that can be inspected directly or aggregated into a scalar uncertainty score. Our approach requires no correctness labels, repeated generations, or training. Across reasoning benchmarks, its confidence score outperforms established baselines under both standard and length-controlled evaluation and transfers more reliably than supervised estimators. Code: https://github.com/s2labres/U-Space.
Chinese Translation
大语言模型正在为风险越来越高的决策提供信息。随着其错误的后果日益严重,一个核心问题变得更难被忽视:我们能在多大程度上信任单个答案?然而,识别何时应当弃权仍然困难,因为语言模型能够以流畅的解释和权威的语气呈现错误的结论。不确定性量化试图通过估计单个预测的可靠性来解决这种脱节。然而,许多现有方法需要重复生成或单独训练的组件,而且它们的标量估计并不能揭示不确定性在何处产生,也不能揭示其在推理过程中如何演化。近期工作还表明,生成长度可能与不确定性估计和正确性强烈相关,这引发了一个问题:一个估计器的预测能力有多大程度来自不确定性特定信息,而不是仅来自输出长度。机制可解释性提供了一种方法,通过将人类可解释的概念与模型的中间状态联系起来,来解决这些局限。基于这一能力,我们引入 U-Space,一个低维子空间,使模型不断演化的不确定性变得可测量且可解释。我们识别出怀疑与确定性的语义锚点,将其反嵌入方向映射回残差空间,并将它们的对比组合成一个正交基。U-Lens 将每个词元状态投影到这些基向量上,产生一个可解释的词元级不确定性图,该图可直接检查或聚合为标量不确定性分数。我们的方法不需要正确性标签、重复生成或训练。在推理基准上,其置信度分数在标准评估和长度控制评估下都优于已有基线,并且比监督式估计器更可靠地迁移。代码:https://github.com/s2labres/U-Space。
cs.CL / 26 / 2610.09152
sk-bench: A Native-First Benchmark for Evaluating Large Language Models in Slovak
sk-bench:一个用于评估斯洛伐克语大语言模型的原生优先基准
large language model
大语言模型相关
Abstract
Multilingual LLM benchmarks omit Slovak, a morphologically rich West Slavic language of five million speakers, or cover it only by machine translation. We present sk-bench, a native-first Slovak benchmark with 30 datasets (33 scored task variants) across ten skill categories. Eleven resources are introduced or first packaged for generative-LLM evaluation, including IFEval-SK with Slovak-adapted instruction checkers and native Chiby/SKJ1 resources for Slovak grammar and morphology. We evaluate 55 open- and closed-weights models under one harness. The best open model trails proprietary APIs by 12.6 points. Model rankings are similar for native and translated closed-form data ($ρ\geq0.98$), though translation separates the strongest models less well. By contrast, human-authored and LLM-generated QA questions rank models differently ($ρ=0.72$). For Qwen3-14B, continued Slovak pretraining lowers the overall score by 13.9 points. A small instruction set restores three quarters of that loss. Test-time reasoning improves scores by 8.5 to 12.5 points for models of 9B and above. Together, these findings suggest four design lessons for other under-resourced languages: use native data where translation fails, plan instruction repair after language adaptation, enable test-time reasoning before scaling up, and avoid overinvesting in target-language prompts. We release the data and code at https://github.com/slovak-nlp/sk-bench
Chinese Translation
多语言 LLM 基准遗漏了斯洛伐克语——一种拥有五百万使用者、形态丰富的西斯拉夫语言——或者仅通过机器翻译来覆盖它。我们提出 sk-bench,一个原生优先的斯洛伐克语基准,包含 30 个数据集(33 个计分任务变体),覆盖十个技能类别。有十一项资源被引入,或首次打包用于生成式 LLM 评估,包括带有斯洛伐克语适配指令检查器的 IFEval-SK,以及用于斯洛伐克语语法和形态学的原生 Chiby/SKJ1 资源。我们在同一个测试框架下评估了 55 个开放权重和封闭权重模型。最佳开放权重模型落后专有 API 12.6 分。模型在原生和翻译后的闭式数据上的排名相似($ρ\geq0.98$),不过翻译对最强模型的区分度较弱。相比之下,人工编写和 LLM 生成的问答问题对模型排名不同($ρ=0.72$)。对于 Qwen3-14B,继续斯洛伐克语预训练使总体得分降低 13.9 分。一个小型指令集恢复了该损失的四分之三。对于 9B 及以上的模型,测试时推理使得分提高 8.5 到 12.5 分。总之,这些发现为其他低资源语言提出了四条设计经验:在翻译失败之处使用原生数据,在语言适应后规划指令修复,在扩大规模之前启用测试时推理,并避免在目标语言提示上过度投资。我们在 https://github.com/slovak-nlp/sk-bench 发布数据和代码。
cs.CL / 27 / 2610.09209
Multi-Objective Aligned Small Language Model Framework for SUD Patient Dialogue Generation
用于SUD患者对话生成的多目标对齐小型语言模型框架
large language model
大语言模型相关
Abstract
Substance Use Disorder (SUD) counseling requires patient responses that reflect underlying cognitive states such as beliefs, coping strategies, and readiness for change. Although large language models (LLMs) can generate fluent text, they often fail to produce cognitively coherent and clinically realistic patient behavior, especially under ethical and data-scarce clinical settings. Moreover, deploying frontier-scale LLMs in healthcare applications presents practical challenges including high computational cost, latency, privacy concerns, and limited deployability in resource-constrained environments, motivating the need for cognitively aligned small language models (SLMs). We propose a cognitively grounded framework for SUD patient dialogue generation that explicitly models and aligns latent cognitive components with patient histories and counselor questions. Our pipeline consists of two stages: cognitive component detection and cognitive component-aligned dialogue generation. To enable effective learning with smaller models, we combine knowledge distillation from high-capacity teacher models, preference optimization from human-annotations, and attention-guided reward shaping. Extensive evaluations using automatic scores like BERTScore, ROUGE, METEOR and BLEU, and LLM-as-judge hit-metrics against both human and teacher-model references show that cognitively informed fine-tuning substantially improves cognitive realization and alignment over a generic instruction-tuned baselines and mental health domain specific SLMs, with particularly strong gains for open-ended cognitive components.
Chinese Translation
物质使用障碍(SUD)咨询要求患者回应能够反映潜在的认知状态,如信念、应对策略和改变准备度。尽管大语言模型(LLM)能够生成流畅文本,但它们往往无法产生认知上连贯且临床现实的患者行为,尤其是在具有伦理约束和数据稀缺的临床环境中。此外,在医疗保健应用中部署前沿规模LLM带来了实际挑战,包括高计算成本、延迟、隐私问题以及在资源受限环境中可部署性有限,这促使人们需要认知对齐的小型语言模型(SLM)。我们提出了一个基于认知的SUD患者对话生成框架,该框架显式地建模潜在认知成分,并将其与患者病史和咨询师问题对齐。我们的流程包含两个阶段:认知成分检测和认知成分对齐的对话生成。为了使较小模型能够有效学习,我们结合了来自高容量教师模型的知识蒸馏、来自人工标注的偏好优化以及注意力引导的奖励塑形。使用如BERTScore、ROUGE、METEOR和BLEU等自动评分,以及针对人类参考和教师模型参考的LLM-as-judge命中指标所进行的广泛评估表明,与通用指令微调基线和心理健康领域专用SLM相比,认知信息驱动的微调显著改善了认知实现和对齐,对于开放式认知成分尤其取得了强劲提升。
cs.CL / 28 / 2610.09346
OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models
OnlineQAT:面向超低位宽大语言模型的同策略蒸馏
large language model
大语言模型相关
Abstract
Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits. Existing re- covery stages, however, are commonly optimized on fixed completions or teacher-generated answers, whereas the deployed quantized model condi- tions on prefixes generated by itself. Quantization errors can therefore move the model into states that are absent from offline recovery data. We introduce OnlineQAT, a two-stage framework that first obtains a usable low-bit initialization through block-wise QAT and then performs on-policy distillation (OPD) on student-generated responses. At each visited pre- fix, a frozen full-precision teacher provides a sampled reverse-KL training signal. On Qwen3-1.7B, OnlineQAT obtains the best average among the compared quantized methods: 57.28 at W3A16 and 32.52 at W2A16, im- proving over ReasoningQAT by 2.90 and 0.44 points, respectively. The results suggest that student-visited states provide a useful recovery signal beyond fixed-completion training, particularly at three bits.
Chinese Translation
量化感知训练(QAT)可以恢复大语言模型在被压缩到低于四位时损失的大部分准确率。然而,现有的恢复阶段通常是在固定的补全或教师生成的答案上进行优化,而部署后的量化模型却以自身生成的前缀为条件。因此,量化误差可能将模型推入离线恢复数据中不存在的状态。我们提出 OnlineQAT,这是一个两阶段框架:首先通过分块 QAT 获得可用的低位宽初始化,然后在学生生成的回答上进行同策略蒸馏(OPD)。在每个访问到的前缀处,一个冻结的全精度教师提供采样的反向 KL 训练信号。在 Qwen3-1.7B 上,OnlineQAT 在所比较的量化方法中取得了最佳平均值:W3A16 下为 57.28,W2A16 下为 32.52,分别比 ReasoningQAT 提升了 2.90 和 0.44 个点。结果表明,学生访问过的状态提供了超越固定补全训练的有用恢复信号,尤其是在三位时。
cs.CL / 29 / 2610.09446
Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science
北极问题,缺失答案:面向北极科学中 LLM 弃答的数据集与基准
large language model
大语言模型相关
Abstract
Large language models (LLMs) should abstain from scientific multiple-choice questions when no option is valid, but frequent abstention alone does not demonstrate sensitivity to answer availability. We introduce ArcticQA, a dataset of 194 questions derived from primary Arctic research, with automated checks of answer support and distractor contradiction against source evidence. We further develop ArcticAbstain, a paired benchmark comparing answer-present and answer-absent conditions, with the correct answer replaced by a distractor in the latter and an explicit abstention option in both. We evaluate eight models from the Gemini, Claude, and ChatGPT families at high reasoning effort, with three trials per condition, yielding 9,312 recorded responses. Answer-present abstention rates range from 0.0% to 63.0%, whereas replacing the correct answer increases abstention by 5.05 percentage points on average. These findings highlight substantial baseline differences and the need to evaluate abstention frequency and responsiveness jointly. The dataset and benchmark are available at https://github.com/BenWilcox8/arctic-qa.
Chinese Translation
大语言模型(LLM)在没有选项有效时应当对科学多选题弃答,但仅凭频繁弃答并不能表明其对答案可用性的敏感性。我们提出 ArcticQA,一个源自北极原始研究的 194 个问题的数据集,并针对源证据对答案支持性与干扰项矛盾性进行自动化检查。我们进一步开发了 ArcticAbstain,一个配对基准,用于比较答案存在与答案缺失两种条件,在后一种条件中将正确答案替换为干扰项,并在两种条件中都提供明确的弃答选项。我们在高推理努力设置下评估了来自 Gemini、Claude 和 ChatGPT 系列的八个模型,每种条件进行三次试验,共产生 9,312 条记录响应。答案存在条件下的弃答率范围为 0.0% 至 63.0%,而替换正确答案平均使弃答率提高 5.05 个百分点。这些发现凸显了显著的基线差异,以及联合评估弃答频率与响应性的必要性。该数据集与基准可在 https://github.com/BenWilcox8/arctic-qa 获取。
cs.CL / 30 / 2610.09489
Goldsmith: Gold-Loss-Guided Definition Optimization with an Agentic Annotation Harness
Goldsmith:黄金损失引导的定义优化与智能体式标注执行框架
large language model
大语言模型相关
Abstract
Many annotation projects begin before experts have a stable guideline or enough labels to train a task-specific model. We present Goldsmith, an agentic pipeline that turns a small gold set---expert-annotated calibration examples representing the intended task boundaries---into a reusable structured annotation definition. Goldsmith treats this definition as a trainable textual object. Candidate definitions are run on the same gold examples and scored with an executable structured loss, while the output schema, formatting, retrieval, repair, judging, and human review remain in an external harness. A large language model (LLM) editor converts the highest-loss failures into textual-gradient revisions, which are accepted only when the measured loss decreases. In prompt-optimization comparisons, Goldsmith improves over direct rewriting, OPRO, APE, and PromptBreeder under matched evaluation protocols. The resulting definition also improves downstream annotation when combined with retrieval, score-based routing, and human review across typed span, pair-level relation, and fixed-trigger event-argument tasks. These results show that scarce expert supervision can support both task-definition learning and scalable annotation.
Chinese Translation
许多标注项目在专家尚未拥有稳定的指南或足够的标签来训练任务特定模型之前就已开始。我们提出 Goldsmith,一种智能体式流水线,它将一个小型黄金集——代表预期任务边界的专家标注校准示例——转化为可复用的结构化标注定义。Goldsmith 将此定义视为可训练的文本对象。候选定义在同一批黄金示例上运行,并使用可执行的结构化损失进行评分,而输出模式、格式、检索、修复、评判和人工审核仍保留在外部执行框架中。一个大型语言模型(LLM)编辑器将损失最高的失败转化为文本梯度式修订,仅当测量到的损失下降时才接受这些修订。在提示优化比较中,在匹配的评估协议下,Goldsmith 优于直接重写、OPRO、APE 和 PromptBreeder。当与检索、基于分数的路由和人工审核相结合时,所得定义还能在类型化片段、对级关系和固定触发词事件论元任务上改善下游标注。这些结果表明,稀缺的专家监督既能支持任务定义学习,也能支持可扩展的标注。
cs.CL / 31 / 2610.09569
RELATE: An Evaluation Framework for measuring Relational Orientation of Large Language Models
RELATE:一个用于衡量大型语言模型关系取向的评估框架
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly used for emotional support, raising concern that sustained use may draw users away from their real-world relationships. Yet existing evaluations primarily focus on the safety, empathy, or helpfulness of responses, leaving under-examined a relational question: where does the model orient the user for continued support? To address this question, we introduce relational orientation, a property operationalized through two non-exclusive dimensions: inward-facing (IF) language, which positions the AI as the user's ongoing source of support, and outward-scaffolding (OS) language, which encourages real-world human connection. Grounded in psychological and sociological literature, we formalize a taxonomy of relational orientation and present RELATE, a persona-conditioned framework for measuring inward-facing and outward-scaffolding language at the sentence level in multi-turn dialogues. RELATE pairs 76 help-seeking situations adapted from naturally occurring questions with three simulated user styles, providing 228 evaluation stimuli. In our experiments, we evaluate seven LLMs using dialogues with six assistant turns each, yielding 1,596 dialogues and 69,194 assistant sentences. We assess these sentences using a primary rubric-based LLM judge and apply a secondary judge to a subset. Under automated evaluation, we find that the proportion of sentences labeled as IF is higher at the sixth assistant turn than at the first, while the proportion labeled as OS is substantially lower for hesitant, indirect simulated users than for explicit, reassurance-seeking users. RELATE provides a reproducible framework and a sentence-level signal for auditing and steering the relational orientation of supportive LLMs.
Chinese Translation
大型语言模型(LLMs)正越来越多地被用于情感支持,这引发了担忧:持续使用可能会使用户远离其现实世界中的关系。然而,现有评估主要关注回复的安全性、同理心或有用性,从而使一个关系性问题仍未得到充分考察:模型将用户导向何处以寻求持续支持?为回答这一问题,我们引入关系取向(relational orientation),这是一种通过两个非互斥维度进行操作化的属性:面向内部的(inward-facing, IF)语言,它将 AI 定位为用户持续支持的来源;以及向外搭建脚手架的(outward-scaffolding, OS)语言,它鼓励现实世界中的人际连接。基于心理学和社会学文献,我们形式化了一个关系取向的分类体系,并提出了 RELATE,一个以人物角色为条件的框架,用于在多轮对话中在句子层面测量面向内部和向外搭建脚手架的语言。RELATE 将 76 个改编自自然出现的问题的求助情境与三种模拟用户风格配对,提供 228 个评估刺激。在我们的实验中,我们使用每个包含六个助手轮次的对话评估了七个 LLM,产生了 1,596 个对话和 69,194 个助手句子。我们使用基于评分标准的主要 LLM 评判器评估这些句子,并对一个子集应用次要评判器。在自动评估下,我们发现,被标记为 IF 的句子比例在第六个助手轮次高于第一个助手轮次,而被标记为 OS 的比例对于犹豫、间接的模拟用户而言显著低于明确、寻求安慰的用户。RELATE 提供了一个可复现的框架和一个句子层面的信号,用于审计和引导支持型 LLM 的关系取向。
cs.CL / 32 / 2610.09571
How Do LLMs Change Predictions Under Negation?
大语言模型在否定下如何改变预测?
large language model
大语言模型相关
Abstract
Negation is an essential feature of human language, yet large language models (LLMs) remain unreliable in processing it. We evaluate recent open-source and closed-source LLMs on our negation benchmark and find that, in 37-71% of cases, they repeat the same answer under negation (e.g., "Madrid" for "What is not the capital of Spain?"). To understand and address this brittleness, we mechanistically examine how models operate under negation. Our main finding is that specialized attention heads and MLP neurons jointly implement negation by (1) suppressing retrieval of the original answer (e.g., "Madrid") while (2) promoting a favored candidate within the answer category (e.g., "Paris"). This contrasts with accounts of human negation processing, in which information about the original answer helps to determine what should be excluded. Furthermore, we find that this difference from human processing is a key source of negation failures: the model's mechanism relies on suppressing the original answer rather than using it to determine what to exclude, so the model can repeat the original answer when suppression is too weak or when a bias toward particular answers prevents it from selecting an alternative. To address this weakness in the model's negation mechanism, we propose a training objective that requires larger shifts in answer preference for more confident original predictions, and show that it reduces negation failures with less degradation of general capabilities than standard fine-tuning baselines. Together, our results demonstrate how mechanistic analysis can reveal why a linguistic capability fails and guide training that targets the underlying limitation.
Chinese Translation
否定是人类语言的一个基本特征,然而大语言模型(LLMs)在处理否定时仍然不可靠。我们在我们的否定基准上评估了近期的开源和闭源 LLMs,并发现,在 37-71% 的情况下,它们在否定下重复相同的答案(例如,对于“什么不是西班牙的首都?”回答“马德里”)。为了理解并解决这种脆弱性,我们从机制上考察了模型在否定下如何运作。我们的主要发现是,专门的注意力头和 MLP 神经元通过以下方式共同实现否定:(1) 抑制对原始答案(例如,“马德里”)的检索,同时 (2) 提升答案类别中一个受偏好的候选(例如,“巴黎”)。这与关于人类否定加工的解释形成对比,在后者中,关于原始答案的信息有助于确定应当排除什么。此外,我们发现,这种与人类加工的差异是否定失败的一个关键来源:模型的机制依赖于抑制原始答案,而不是利用它来确定应排除什么,因此当抑制过弱时,或者当对特定答案的偏向阻止其选择替代答案时,模型就可能重复原始答案。为解决模型否定机制中的这一弱点,我们提出了一种训练目标,对于置信度更高的原始预测要求答案偏好发生更大的偏移,并表明与标准微调基线相比,它在减少否定失败的同时对通用能力的退化更小。总之,我们的结果表明,机制分析如何能够揭示某种语言能力为何失败,并指导针对其底层限制的训练。
cs.CL / 33 / 2610.09665
SAPD: Step-Aligned Privileged Distillation
SAPD:步骤对齐的特权蒸馏
large language model
大语言模型相关
Abstract
On-policy post-training can improve large language models by learning from their own trajectories, but requires costly rollout generation. We ask whether fixed demonstrations can support competitive off-policy learning through better supervision. Our premise is that their usefulness depends not only on the training trajectories, but also on whether supervision provides informative preferences among continuations and connects this guidance to the reasoning decision being learned. We introduce Step-Aligned Privileged Distillation (SAPD), a rollout-free self-distillation method that turns demonstrations into step-aligned distributional supervision. Its key insight is to use the known progression of a reference solution to associate each reasoning transition with targeted privileged guidance, rather than treating the solution as undifferentiated context. On mathematical reasoning benchmarks, SAPD outperforms supervised fine-tuning and label smoothing on average while remaining competitive with on-policy reinforcement learning and self-distillation. Analyses support both the value of context-dependent distributional guidance and the benefit of aligning privileged information with the current step. SAPD also largely preserves out-of-domain coding performance and achieves approximately 2x training-loop speedups over the on-policy baselines. These findings suggest that carefully constructed supervision can make fully off-policy post-training a competitive and computationally efficient alternative. Our code is available at https://github.com/Miaow-Lab/SAPD.
Chinese Translation
同策略后训练可以通过从大语言模型自身的轨迹中学习来改进它们,但需要昂贵的 rollout 生成。我们询问,固定示范能否通过更好的监督来支持具有竞争力的离策略学习。我们的前提是,它们的用处不仅取决于训练轨迹,还取决于监督是否在续写之间提供有信息量的偏好,并将这种指导与正在学习的推理决策联系起来。我们提出步骤对齐的特权蒸馏(SAPD),一种无需 rollout 的自蒸馏方法,它将示范转化为步骤对齐的分布监督。其关键洞见是,利用参考解的已知推进过程,将每个推理转移与有针对性的特权指导相关联,而不是将该解视为无差别的上下文。在数学推理基准上,SAPD 平均而言优于监督微调和标签平滑,同时与同策略强化学习和自蒸馏相比仍具有竞争力。分析既支持依赖上下文的分布指导的价值,也支持将特权信息与当前步骤对齐的益处。SAPD 还在很大程度上保持了域外编码性能,并相较于同策略基线实现了约 2 倍训练循环加速。这些发现表明,精心构建的监督可以使完全离策略后训练成为一种具有竞争力且计算高效的替代方案。我们的代码可在 https://github.com/Miaow-Lab/SAPD 获取。
cs.CL / 34 / 2610.09671
InsClaimBench: Benchmarking Insurance Claim Adjudication Across the Decision Chain
InsClaimBench:跨决策链的保险理赔裁决基准测试
large language model
大语言模型相关
Abstract
Recent advances in reasoning-oriented large language models (LLMs) have motivated increasing evaluation of their ability to perform professional decision tasks. Insurance claim adjudication is one such task, requiring models to connect case evidence, insurance rules, intermediate judgments, and payout calculations across a structured decision process. We introduce InsClaimBench, an end-to-end benchmark for evaluating insurance claim adjudication across the decision chain. Grounded in real claim materials and structured insurance rules, InsClaimBench contains 3,780 cases in 375 case families across auto, property, and health insurance, comprising 86,656 atomic rule judgments. It evaluates each claim from atomic rules through adjudication modules to payout decisions and amounts, with controlled factual variants testing whether required changes are correctly propagated across levels. Evaluation of six LLMs reveals a progressive loss of reliability along the decision chain. Payout-decision accuracy ranges from 74.23--80.19%, while joint decision--amount accuracy drops to 47.54--73.15%. Strong local performance also fails to ensure case-level correctness: atomic-rule accuracy reaches 95.48%, whereas rule-vector exact match peaks at only 36.90%, and the most frequent module errors are not necessarily those most associated with final-decision failure. Under factual changes, these inconsistencies further become propagation failures: module updates are less reliable than rule updates, correct local judgments can still yield incorrect payouts, and correct payouts can conceal intermediate errors. These results show that reliable claim adjudication requires consistent composition and propagation across the decision chain.
Chinese Translation
近年来,面向推理的大型语言模型(LLMs)的进展推动了对其执行专业决策任务能力的日益增多的评估。保险理赔裁决就是这样一项任务,要求模型在结构化的决策过程中将案件证据、保险规则、中间判断和赔付计算连接起来。我们提出 InsClaimBench,一个用于评估跨决策链的保险理赔裁决的端到端基准。InsClaimBench 以真实理赔材料和结构化保险规则为基础,包含汽车险、财产险和健康险领域的 375 个案件族中的 3,780 个案件,涵盖 86,656 个原子规则判断。它从原子规则经由裁决模块到赔付决策和赔付金额,对每项理赔进行评估,并通过受控的事实变体测试所需变化是否在各层级之间被正确传播。对六个 LLM 的评估揭示了沿决策链的可靠性逐步丧失。赔付决策准确率介于 74.23--80.19%,而决策--金额联合准确率降至 47.54--73.15%。强大的局部性能也未能确保案件级正确性:原子规则准确率达到 95.48%,而规则向量精确匹配最高仅为 36.90%,并且最频繁的模块错误不一定是最与最终决策失败相关的错误。在事实变化下,这些不一致进一步变成传播失败:模块更新不如规则更新可靠,正确的局部判断仍可能产生错误的赔付,而正确的赔付可能掩盖中间错误。这些结果表明,可靠的理赔裁决需要在决策链上保持一致组合和传播。
cs.CL / 35 / 2610.09684
From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery
从帕累托到偏好:通过摊销式智能体策略发现实现个性化测试时扩展
large language model
大语言模型相关
Abstract
Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--cost or accuracy--latency Pareto frontier. Yet user requirements are multidimensional: users may specify accuracy, latency, and inference-cost requirements jointly, and different requirements can favor different controllers. We formulate Personalized Test-Time Scaling as discovering executable controllers that maximize the joint satisfaction rate of user-specific requirements. To reduce the overhead of repeated policy discovery for new user profiles, we propose PersonTTS, an amortized agentic policy-discovery framework that reuses prior search experience through requirement-matched controller initialization and source-distilled procedural guidance, while retaining target-profile evaluation for every candidate. Experiments on AIME and HMMT show that PersonTTS substantially outperforms strong TTS baselines in joint requirement satisfaction on unseen user profiles and held-out problems. Under the same candidate-evaluation budget, cross-user experience reuse further improves policy quality while substantially reducing discovery-agent time and cost.
Chinese Translation
测试时扩展(TTS)通过分配额外的推理计算来提升大语言模型的推理能力。现有的提升 TTS 效率的方法大多一次针对一个资源维度优化准确率,推进准确率—成本或准确率—延迟帕累托前沿。然而用户需求是多维的:用户可能同时指定准确率、延迟和推理成本要求,而不同要求可能偏好不同的控制器。我们将个性化测试时扩展建模为发现可执行控制器,以最大化用户特定需求的联合满足率。为降低针对新用户画像重复进行策略发现的开销,我们提出 PersonTTS,一种摊销式智能体策略发现框架,它通过需求匹配的控制器初始化和源蒸馏的过程性指导来复用先前的搜索经验,同时为每个候选保留目标画像评估。在 AIME 和 HMMT 上的实验表明,在未见过的用户画像和留出问题上,PersonTTS 在联合需求满足方面显著优于强 TTS 基线。在相同的候选评估预算下,跨用户经验复用进一步提升了策略质量,同时大幅降低了发现智能体的时间和成本。
cs.CL / 36 / 2610.09835
A Deafening Silence: Catastrophic Forgetting Lives in the Output Embeddings of Tokens the Data Never Speaks
震耳欲聋的沉默:灾难性遗忘存在于数据从未言说的词元的输出嵌入中
large language model
大语言模型相关
Abstract
Continual pre-training and fine-tuning in Large Language Models (LLMs) inevitably induce catastrophic forgetting, typically mitigated by replay using often-inaccessible original data. In this data-free regime, we analyze where forgetting occurs and why. Systematic parameter freezing across five settings up to 1.4B reveals that forgetting concentrates selectively in the output embeddings of tokens rarely seen in the new corpus, whereas the same sqrt(v-hat) band of the body is inert and new learning resides elsewhere. This localization is governed by the vocabulary deficiency of the corpus rather than the training mode, allowing pre-retraining risk ranking from token counts alone within a fixed base model. Mechanistically, absent tokens receive persistent one-sided softmax gradients that Adam's second-moment (sqrt(v-hat)) normalization amplifies into full-sized updates. We therefore propose an intervention: raising Adam's epsilon exclusively for the output projection during training. Across eight settings spanning 160M to 12B parameters and four model families, this removes 39.4% to 67.9% of forgetting across all seven stable configurations without degrading target learning or requiring per-model tuning. The defense combines additively or better with replay (79.8% on Qwen/Korean) and rescues released-head LoRA from a 23-fold forgetting surge. Because post-hoc editing of the drifted rows recovers under 5% of forgetting, the intervention must operate during training. Our findings indicate that a single-line optimizer adjustment may serve as the primary defense against catastrophic forgetting where the corpus starves the vocabulary.
Chinese Translation
大语言模型(LLMs)中的持续预训练与微调不可避免地会引发灾难性遗忘,通常通过使用往往无法获取的原始数据进行重放来缓解。在这种无数据的情形下,我们分析遗忘发生在何处以及为何发生。在最多 1.4B 的五个设置中进行的系统性参数冻结揭示,遗忘选择性地集中在新语料库中很少见到的词元的输出嵌入中,而模型主体中相同的 sqrt(v-hat) 带则是惰性的,新学习位于其他地方。这种定位由语料库的词表匮乏而非训练模式所支配,使得在固定基础模型内仅根据词元计数即可在重新训练前进行风险排序。从机制上看,缺失词元接收到持续的单侧 softmax 梯度,Adam 的二阶矩(sqrt(v-hat))归一化将其放大为完整规模的更新。因此,我们提出一种干预:在训练期间仅针对输出投影提高 Adam 的 epsilon。在跨越 160M 到 12B 参数以及四个模型家族的八个设置中,这在所有七个稳定配置中消除了 39.4% 到 67.9% 的遗忘,同时不降低目标学习,也不需要对每个模型进行调参。该防御与重放结合时以相加方式或更好地起作用(在 Qwen/Korean 上为 79.8%),并将 released-head LoRA 从 23 倍的遗忘激增中挽救出来。由于对漂移行的事后编辑只能恢复不到 5% 的遗忘,该干预必须在训练期间起作用。我们的发现表明,在语料库使词表饥饿的情况下,单行优化器调整可能作为对抗灾难性遗忘的主要防御手段。
cs.CL / 37 / 2610.09858
Training Advisors for LLM Agents from Task Outcomes
基于任务结果为LLM智能体训练顾问
large language model
大语言模型相关
Abstract
Large language model agents tackle multi-step tasks by interleaving reasoning and tool calls with observations from the environment. Prior work has shown that natural-language feedback can help these agents revise their decisions during task execution. We introduce Caddie, a method for training critics to provide natural-language analysis and advice as agents work through a task. Unlike approaches that rely on step-level labels or reference critiques, Caddie learns from whether the agent ultimately succeeds after receiving the critic's feedback. We optimize the critic through reinforcement learning while keeping the base model frozen. Trained on multi-hop question answering with a single base model, our Qwen3-4B critic improves success rates across four base models of different scales and architectures, including three not used during critic training. On the MuSiQue benchmark, the trained critic improves Qwen3-4B's success rate by more than 25 percentage points, surpassing the performance of Kimi K3 without a critic. The same critic also yields gains on out-of-domain interactive benchmarks, including $τ^3$ and DeepDive, with no additional training. Our results show that agents can decide when to seek help from a critic at inference time and that outcome-based critic training can produce guidance that transfers across base models and task domains.
Chinese Translation
大语言模型智能体通过将推理和工具调用与环境中的观察结果交错进行,来解决多步任务。先前工作表明,自然语言反馈可以帮助这些智能体在任务执行过程中修正其决策。我们提出 Caddie,一种训练评论者在智能体完成任务过程中提供自然语言分析和建议的方法。与依赖步骤级标签或参考评论的方法不同,Caddie 根据智能体在收到评论者反馈后最终是否成功来学习。我们通过强化学习优化评论者,同时保持基础模型冻结。我们的 Qwen3-4B 评论者在单个基础模型上通过多跳问答训练后,提高了四个不同规模和架构的基础模型的成功率,其中包括三个未在评论者训练期间使用的模型。在 MuSiQue 基准上,训练后的评论者将 Qwen3-4B 的成功率提高了超过 25 个百分点,超过了没有评论者的 Kimi K3 的性能。同一个评论者还在域外交互式基准上带来增益,包括 $τ^3$ 和 DeepDive,且无需额外训练。我们的结果表明,智能体可以在推理时决定何时向评论者寻求帮助,并且基于结果的评论者训练可以产生能够跨基础模型和任务领域迁移的指导。
cs.CL / 38 / 2610.09976
EASE: Entropy-Adaptive Distribution Shaping for Evading AI-generated Text Detectors
EASE:用于规避 AI 生成文本检测器的熵自适应分布塑造
large language model
大语言模型相关
Abstract
AI-generated text (AIGT) detection can be sensitive to the decoding choices of the source large language model (LLM). We observe that perturbing next-token logits or adjusting sampling temperature can reduce detection performance, providing a clear signal of detector vulnerability to decoding-time distribution changes. Building on this observation, we propose EASE (Entropy-Adaptive Distribution Shaping for Evasion), a training-free and detector-agnostic framework for evading AIGT detectors. EASE computes predictive entropy directly from the source LLM's next-token distribution and uses it to adapt both logit perturbation and sampling temperature, without detector feedback or model fine-tuning. Experiments across three source LLMs and multiple detectors demonstrate consistent reductions in detection performance, with negligible degradation in text quality and negligible inference overhead.
Chinese Translation
AI 生成文本(AIGT)检测可能对源大语言模型(LLM)的解码选择敏感。我们观察到,扰动下一词元 logits 或调整采样温度会降低检测性能,这为检测器对解码时分布变化的脆弱性提供了明确的信号。基于这一观察,我们提出了 EASE(用于规避的熵自适应分布塑造),一个无需训练且与检测器无关的、用于规避 AIGT 检测器的框架。EASE 直接从源 LLM 的下一词元分布计算预测熵,并利用它来同时自适应地调整 logit 扰动和采样温度,而无需检测器反馈或模型微调。在三个源 LLM 和多个检测器上的实验表明,检测性能一致下降,同时文本质量的退化可忽略不计,推理开销也可忽略不计。
cs.CL / 39 / 2610.10049
The Long Road to the Same Answer: Cognitive Bias Under Escalating Reasoning Budgets in Large Language Models
通往同一答案的漫长道路:大语言模型在推理预算逐步升级下的认知偏差
large language model
大语言模型相关
Abstract
Reasoning models allocate extra computation at inference time and present their answers as the product of deliberate thought. If this deliberation works the way dual-process accounts of human cognition suggest, longer thinking should weaken the classic decision biases that fast, intuitive judgment produces. Using 30 vignettes covering six biases (anchoring, framing, loss aversion, escalation of commitment, availability, confirmation) from an established benchmark, we run a dose-response study across four model families, pairing each reasoning model with a matched non-reasoning sibling and requesting thinking ceilings of 0, 1,024, 4,096, and 8,192 tokens, for 12,350 API calls. Because a requested ceiling is not the same as realized deliberation, we use the reasoning tokens each call consumed as the dose. First, reasoning models are not less biased than their siblings; the point estimate leans the other way in every family, but the item-level pooled contrast is not reliable (Delta = +0.031, t(29) = 1.45, p = .157). Second, bias magnitude does not reliably fall as realized deliberation grows: no slope is significantly negative, and where anything moves it is the signed score drifting further from the human direction. Third, anchoring is the only bias in the human direction (d = 1.89). Four of the other five lean the opposite way in all seven models; with five items per bias, that reversal is reliable for framing and directional for escalation of commitment, confirmation, and loss aversion, while availability is absent. A one-line instruction to restate the anchor before answering lowered anchoring on all five anchoring items, which no amount of additional thinking did, although the effect does not reach significance (p = .057). The results argue against treating test-time reasoning as a rationality guarantee and for auditing deployed models bias by bias.
Chinese Translation
推理模型在推断时分配额外计算,并将其答案呈现为审慎思考的产物。如果这种审慎思考的运行方式如人类认知的双过程解释所暗示的那样,那么更长时间的思考应当削弱快速、直觉判断所产生的经典决策偏差。使用来自一个成熟基准的 30 个情境短文,涵盖六种偏差(锚定、框架、损失厌恶、承诺升级、可得性、确认),我们在四个模型家族中开展了一项剂量-反应研究,将每个推理模型与一个匹配的非推理同类模型配对,并要求思考上限为 0、1,024、4,096 和 8,192 个 token,共进行 12,350 次 API 调用。因为请求的上限并不等同于实际实现的审慎思考,我们使用每次调用所消耗的推理 token 作为剂量。第一,推理模型并不比它们的同类模型偏差更少;在每个模型家族中,点估计都偏向另一方向,但项目层面的合并对比并不可靠(Delta = +0.031,t(29) = 1.45,p = .157)。第二,随着实际实现的审慎思考增加,偏差幅度并未可靠下降:没有斜率显著为负,而在任何出现变化的地方,都是有符号得分进一步偏离人类方向。第三,锚定是唯一一种朝人类方向的偏差(d = 1.89)。其他五种偏差中有四种在全部七个模型中都偏向相反方向;由于每种偏差有五个项目,这种反转在框架效应上可靠,在承诺升级、确认和损失厌恶上具有方向性,而可得性偏差则不存在。一条要求回答前重述锚的一行指令降低了全部五个锚定项目上的锚定效应,而无论多少额外思考都未能做到这一点,尽管该效应未达到显著性(p = .057)。这些结果表明,不应将测试时推理视为理性保证,而应逐个偏差审计已部署的模型。
cs.CL / 40 / 2610.10114
Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position
长上下文混合模型的机制 第1.1部分:从混合注意力到混合位置
large language model
大语言模型相关
Abstract
The architectural design of Large Language Models (LLMs) is shifting from traditional full-attention-only models to hybrid models, which combine different attention modules to improve long-context efficiency and performance in length extrapolation and context extension. To explain why hybrid models work and how to design them better, we propose Mechanics of Long-Context Hybrid Models. As Part 1.1 of this series, we begin with hybrids of full attention and either sliding-window attention (SWA) or gated variants of linear attention (LA), represented by GLA and GDN. We first observe a Seesaw Effect in Context Extension: LA hybrids benefit more from long-context continual pretraining, whereas SWA hybrids perform better under length extrapolation. We attribute this behavior to differences in the positional inductive biases induced by these attention mechanisms. We find that SWA hybrids suffer from a Short-Context Learning Trap, Short-Window Weariness, and Long-Window Laziness, and require extended windows to enhance performance in continual long-context pretraining. For LA hybrids, we summarize the Matthew Effect of Hybrid Position Extrapolation and propose Sliding-Window Linear Attention, achieving 16$\times$ training-free length extrapolation while maintaining 100\% accuracy on NIAH-SK1 in 64k context length.
Chinese Translation
大型语言模型(LLMs)的架构设计正在从传统的仅全注意力模型转向混合模型,这些混合模型结合不同的注意力模块,以提高长上下文效率以及长度外推和上下文扩展中的性能。为了解释混合模型为何有效以及如何更好地设计它们,我们提出长上下文混合模型的机制。作为本系列的第1.1部分,我们从全注意力与滑动窗口注意力(SWA)或线性注意力(LA)的门控变体(以 GLA 和 GDN 为代表)的混合模型开始。我们首先在上下文扩展中观察到一种跷跷板效应:LA 混合模型从长上下文持续预训练中获益更多,而 SWA 混合模型在长度外推下表现更好。我们将这种行为归因于这些注意力机制所引入的位置归纳偏置的差异。我们发现 SWA 混合模型存在短上下文学习陷阱、短窗口疲倦和长窗口懒惰,并且需要扩展窗口以提高持续长上下文预训练中的性能。对于 LA 混合模型,我们总结了混合位置外推的马太效应,并提出滑动窗口线性注意力,实现了 16$\times$ 免训练长度外推,同时在 64k 上下文长度下的 NIAH-SK1 上保持 100\% 准确率。
cs.CL / 41 / 2610.10179
Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search Agents
超越结果奖励:为搜索智能体构建与分配检索信用
large language model
大语言模型相关
Abstract
Search agents enable Large Language Models (LLMs) to iteratively retrieve and use information for complex multi-hop questions. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising approach for post-training such agents, but its reliance on sparse, outcome-based supervision can make credit assignment difficult and limit learning efficiency. In this paper, we systematically investigate how intermediate supervision can improve reinforcement learning for search agents. We study a range of reward-shaping and credit-assignment strategies that provide learning signals from intermediate retrieval steps. Building on these insights, we develop a training framework that combines intermediate signals with final outcome rewards to improve learning from multi-step search trajectories. Experiments across multiple benchmarks under matched training conditions demonstrate improvements in aggregate search-agent performance and show that both the choice of intermediate signal and where its credit is assigned affect training behaviour. These findings show that reward design and credit assignment are important design dimensions for training effective search agents.
Chinese Translation
搜索智能体使大语言模型(LLMs)能够为复杂的多跳问题迭代地检索和使用信息。可验证奖励强化学习(RLVR)为这类智能体的后训练提供了一种有前景的方法,但它对稀疏的、基于结果的监督的依赖会使信用分配变得困难,并限制学习效率。在本文中,我们系统地研究了中间监督如何改进搜索智能体的强化学习。我们研究了一系列奖励塑形与信用分配策略,这些策略从中间检索步骤提供学习信号。基于这些见解,我们开发了一个训练框架,将中间信号与最终结果奖励相结合,以改进从多步搜索轨迹中的学习。在匹配训练条件下跨多个基准的实验表明,搜索智能体的总体性能得到提升,并表明中间信号的选择及其信用被分配的位置都会影响训练行为。这些发现表明,奖励设计与信用分配是训练有效搜索智能体的重要设计维度。
cs.CL / 42 / 2610.10232
LLM Persuasion Is in the Eye of the Evaluation
大语言模型的说服力取决于评估视角
large language model
大语言模型相关
Abstract
Large language models (LLMs) have already been shown to match or exceed human experts in persuasion. While their persuasive capabilities hold promise for beneficial uses such as education and health communication, they can also be used to manipulate and misinform, making their evaluation a growing priority for developers and regulators. That evaluation, however, remains fragmented: studies differ in what they treat as persuasion, and broad claims often rest on narrow, situation-specific assessments. Automated methods, often modelled on human studies, offer a way to compare such assessments directly, as they can be run on the same models at scale and can include high-risk forms of persuasion that would be difficult or unethical to test on people. In this study, we adapt nine published automated methods to a shared setup, run them on the same fifteen LLMs, and ask whether their rankings agree and why. We find that the methods agree only weakly (mean Spearman $ρ= 0.25$). Our analyses point to two contributing factors. Models that refuse some tasks but not others, directly or indirectly, lower agreement by about a quarter, and these refusals fall mostly on manipulation tasks. General capability also plays a part: most rational persuasion (non-manipulative) methods track it, whereas most manipulation methods do not. Together, these findings suggest that agreement depends more on the task a method sets than on how it scores persuasion, although this pattern is only indicative given the eight methods available for analysis. More broadly, our results suggest that persuasion scores combine a model's ability to persuade with its willingness to do so. A single score is therefore informative about its own setting, but says little about a model's persuasiveness across tasks.
Chinese Translation
大型语言模型(LLM)已被证明在说服方面能够匹配或超过人类专家。尽管它们的说服能力有望用于教育和健康传播等有益用途,但它们也可能被用于操纵和误导,这使其评估日益成为开发者和监管者的优先事项。然而,这种评估仍然分散:不同研究对什么被视为说服存在差异,而宽泛的论断往往建立在狭窄、特定情境的评估之上。自动化方法通常以人类研究为蓝本,为直接比较此类评估提供了途径,因为它们可以在相同模型上大规模运行,并且可以包含难以或不道德在人类身上测试的高风险说服形式。在本研究中,我们将九种已发表的自动化方法适配到一个共享设置中,在相同的十五个 LLM 上运行它们,并探究它们的排名是否一致以及为什么。我们发现这些方法仅表现出弱一致性(平均 Spearman $ρ= 0.25$)。我们的分析指出了两个促成因素。那些直接或间接地拒绝某些任务而不拒绝另一些任务的模型,使一致性降低约四分之一,而这些拒绝大多落在操纵任务上。通用能力也发挥了作用:大多数理性说服(非操纵性)方法会追踪它,而大多数操纵方法则不会。总之,这些发现表明,一致性更多取决于方法所设定的任务,而不是它如何对说服力进行评分,尽管鉴于可用于分析的方法只有八种,这一模式仅具有指示性。更广泛地说,我们的结果表明,说服力分数结合了一个模型的说服能力与其进行说服的意愿。因此,单一分数对其自身设置具有信息量,但对模型跨任务的说服力几乎没有说明。
cs.CL / 43 / 2610.10318
Nobody Truly Agrees on Sentiment: Humans, Bespoke Tools, and LLMs Struggle with Social Media Texts
在情感判断上无人真正达成一致:人类、定制工具和大语言模型在社交媒体文本上举步维艰
large language model
大语言模型相关
Abstract
Social media is a rich source of real-time public sentiment, but widely used sentiment analysis tools are often applied without understanding their limitations. In this study, we evaluate the inter-rater reliability of three bespoke sentiment analysis tools (TextBlob, VADER, and Twitter-roBERTa-base) and three large language models (LLMs: Qwen3-32B, GPT-OSS-120B, Llama-4-Maverick-17B) against six human raters across 100 tweets. We measured agreement using two statistical measures: Cohen's kappa for pairwise comparisons and Fleiss' kappa for multiple raters. Even among the human raters, our results showed only fair agreement, highlighting the subjectivity of sentiment analysis. Higher agreement was observed under the binary sentiment classification (negative vs. non-negative and positive vs. non-positive) than under the three-class classification across both humans and automated tools. The Twitter-roBERTa-base model showed the strongest alignment with human ratings, outperforming both bespoke sentiment tools and LLMs, particularly in distinguishing negative versus non-negative sentiment. LLMs showed substantial agreement among themselves and moderate to substantial alignment with humans, performing better in positive vs. non-positive classifications. Our findings underscore that domain-specific fine-tuning remains crucial for reliable social media sentiment analysis, and human-centered evaluation remains essential for establishing gold-standard labels.
Chinese Translation
社交媒体是实时公众情绪的丰富来源,但广泛使用的情感分析工具常常在未理解其局限性的情况下被应用。在本研究中,我们针对 100 条推文,评估了三种定制情感分析工具(TextBlob、VADER 和 Twitter-roBERTa-base)和三种大语言模型(LLMs:Qwen3-32B、GPT-OSS-120B、Llama-4-Maverick-17B)与六名人类评分者之间的评分者间信度。我们使用两种统计指标来衡量一致性:用于成对比较的 Cohen's kappa 和用于多评分者的 Fleiss' kappa。即使是在人类评分者之间,我们的结果也仅显示出尚可的一致性,这凸显了情感分析的主观性。在人类和自动化工具中,二元情感分类(负面 vs. 非负面以及正面 vs. 非正面)下观察到的一致性均高于三类分类下的一致性。Twitter-roBERTa-base 模型与人类评分的一致性最强,优于定制情感工具和 LLMs,尤其是在区分负面与非负面情感方面。LLMs 之间显示出高度一致性,并且与人类具有中等至高度的一致性,在正面 vs. 非正面分类中表现更好。我们的发现强调,领域特定微调对于可靠的社交媒体情感分析仍然至关重要,而以人为中心的评估对于建立金标准标签仍然必不可少。
cs.CL / 44 / 2610.10378
Document-Level Text Simplification in Estonian Using Large Language Models
使用大语言模型进行爱沙尼亚语文档级文本简化
large language model
大语言模型相关
Abstract
Document-level text simplification involves transformations that go beyond sentence-internal edits, addressing discourse coherence, anaphora resolution, and cross-paragraph consistency. Despite advances in sentence-level simplification for high-resource languages, document-level simplification in morphologically rich, low-resource languages such as Estonian remains largely unexplored. This study presents a comprehensive evaluation of five state-of-the-art multilingual large language models (LLMs) for document-level simplification in Estonian. Three prompting strategies are examined: single-pass generation, pipeline-based modular agents, and guideline-augmented pipelines. The evaluation framework integrates automatic metrics assessing readability, semantic preservation, and discourse coherence, alongside a structured manual annotation protocol. The findings indicate that Gemini-2.0 and LLaMA-3.3 produce outputs with near-native fluency and strong meaning preservation, whereas other models display notable grammatical and semantic limitations. This work contributes novel document-level coherence metrics, evidence-based prompting strategies, and publicly available resources for reproducibility.
Chinese Translation
文档级文本简化涉及超越句子内部编辑的转换,处理语篇连贯性、回指消解以及跨段落一致性。尽管高资源语言的句子级简化取得了进展,但在形态丰富、低资源语言(如爱沙尼亚语)中的文档级简化在很大程度上仍未得到探索。本研究对五种最先进的多语言大语言模型(LLMs)在爱沙尼亚语文档级简化任务上进行了全面评估。考察了三种提示策略:单遍生成、基于流水线的模块化智能体,以及通过指南增强的流水线。该评估框架整合了评估可读性、语义保持和语篇连贯性的自动指标,以及一个结构化的手动标注协议。研究结果表明,Gemini-2.0 和 LLaMA-3.3 生成的输出具有接近母语的流畅性和很强的意义保持能力,而其他模型则表现出明显的语法和语义局限。这项工作贡献了新颖的文档级连贯性指标、基于证据的提示策略,以及用于可复现性的公开可用资源。
cs.CL / 45 / 2610.10455
PHRBench: A Behavioral Evaluation of Post-Hallucination Reasoning in LLMs
PHRBench:LLMs 中幻觉后推理的行为评估
large language model
大语言模型相关
Abstract
Hallucinated information can propagate through multi-stage LLM systems and become part of the context for subsequent reasoning. Existing studies of post-hallucination reasoning (PHR) mainly characterize changes in final outcomes and aggregate reasoning dynamics, leaving how models resolve hallucinated premises at the response level insufficiently understood. In this work, we introduce PHRBench, a controlled benchmark for behaviorally structured PHR across four domains and 18 large language models. PHRBench characterizes each reasoning trajectory independently of final-answer correctness through Hallucination Compliance, Hallucination Avoidance, and Heuristic Correction, and defines an insightful trajectory as successful correction that ultimately reaches the correct answer. Across 4820 controlled instances, we find that successful recovery remains relatively rare and is associated with more frequent belief updates along the reasoning trajectory. We further find that properties of the hallucinated prompt contain substantial predictive signal for successful recovery, with a lightweight predictor achieving an AUROC of 0.847. These findings provide a behavioral view of post-hallucination reasoning, characterizing how LLMs resolve erroneous context and when successful recovery is likely to occur.
Chinese Translation
幻觉信息可以通过多阶段 LLM 系统传播,并成为后续推理上下文的一部分。现有的幻觉后推理(PHR)研究主要刻画最终结果的变化和聚合的推理动态,导致模型在响应层面如何解决幻觉前提仍未得到充分理解。在这项工作中,我们提出了 PHRBench,一个用于跨四个领域和 18 个大型语言模型的行为结构化 PHR 的受控基准。PHRBench 通过幻觉遵从、幻觉避免和启发式纠正,独立于最终答案正确性来刻画每条推理轨迹,并将有洞察力的轨迹定义为最终达到正确答案的成功纠正。在 4820 个受控实例中,我们发现成功恢复仍然相对罕见,并且与推理轨迹中更频繁的信念更新相关。我们进一步发现,幻觉提示的属性包含对成功恢复的显著预测信号,一个轻量级预测器达到了 0.847 的 AUROC。这些发现提供了对幻觉后推理的行为视角,刻画了 LLM 如何解决错误上下文以及成功恢复可能在何时发生。
cs.CL / 46 / 2610.10533
EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory
EngramEdit:通过条件记忆实现大语言模型中的解耦知识更新
large language model
大语言模型相关
Abstract
Conditional memory architectures such as DeepSeek Engram use input n-grams to look up learned embeddings, expanding the capacity of large language models (LLMs) with limited additional computation. Beyond model scaling, this architecture has demonstrated the potential to decouple factual knowledge storage from general-purpose computation, offering a promising route to updating factual knowledge while keeping the Transformer backbone fixed. Realizing this potential is challenging because different expressions of a fact may activate different n-gram embeddings, while updating shared embeddings can unintentionally change the model's predictions about other facts. We propose EngramEdit for decoupled knowledge updates through conditional memory. EngramEdit first computes target memory representations that make the model predict the updated fact across multiple expressions. It then jointly updates the shared n-gram embeddings to match these targets across expressions and edits, penalizing updates to frequently reused embeddings more strongly to preserve unrelated knowledge. Experiments show that EngramEdit enables independent factual knowledge updates through conditional memory, achieving near-perfect editing success. Revised knowledge is usable across unseen expressions and in multi-hop reasoning, with nearly three times the strongest baseline's accuracy under chain-of-thought (CoT) prompting. Unrelated knowledge and general capabilities are largely preserved even as factual updates accumulate. These findings show that EngramEdit turns conditional memory into an editable knowledge interface, extending its role beyond model scaling to support decoupled knowledge updates.
Chinese Translation
诸如 DeepSeek Engram 之类的条件记忆架构使用输入 n-gram 来查找学习到的嵌入,以有限的额外计算扩展大语言模型(LLMs)的容量。超越模型扩展之外,该架构已展现出将事实知识存储与通用计算解耦的潜力,为在保持 Transformer 主干固定的同时更新事实知识提供了一条有前景的路径。实现这一潜力具有挑战性,因为同一个事实的不同表达可能激活不同的 n-gram 嵌入,而更新共享嵌入可能会无意中改变模型对其他事实的预测。我们提出 EngramEdit,用于通过条件记忆进行解耦知识更新。EngramEdit 首先计算目标记忆表示,使模型能够在多种表达中预测更新后的事实。然后,它联合更新共享 n-gram 嵌入,以在多种表达和编辑中匹配这些目标,同时对频繁复用的嵌入的更新施加更强惩罚,以保留无关知识。实验表明,EngramEdit 能够通过条件记忆实现独立的事实知识更新,达到近乎完美的编辑成功率。修订后的知识可用于未见过的表达和多跳推理,在思维链(CoT)提示下,准确率几乎是最强基线的三倍。即使事实更新不断累积,无关知识和通用能力也在很大程度上得以保留。这些发现表明,EngramEdit 将条件记忆转变为可编辑的知识接口,将其作用从模型扩展延伸至支持解耦知识更新。
cs.CR / 47 / 2610.09337
Ask the Expert: LLM-Guided Reinforcement Learning for Autonomous Cyber Defense
问专家:面向自主网络防御的 LLM 引导强化学习
large language model
大语言模型相关
Abstract
Policy-based reinforcement learning (RL) approaches have produced promising results for autonomous cyber defense; however, they are sample-inefficient in settings where defenders must respond under delayed, partial observations with actions from large action spaces. While large language models (LLMs) may reason semantically about security state space, high latency and trust assumptions prevent attractive in-line deployment models. We introduce Ask the Expert, a training-time guidance framework which first summarizes hard cyber-defense states, then intermittently queries an LLM for host-level defensive recommendations via a constrained action interface, and finally transforms those recommendations into tiered reward shaping for use with PPO. Because the LLM is discarded after training, deployment is a pure RL policy. Across TTCP CAGE CC1 and CC2 and both attacker types, this asymmetric design improves sample efficiency over PPO and outperforms the evaluated potential-based reward shaping (PBRS) baselines, while retaining the strongest terminal mean and requiring no LLM dependency at deployment time.
Chinese Translation
基于策略的强化学习(RL)方法在自主网络防御中已取得令人瞩目的成果;然而,在防御者必须在延迟的、部分可观测的条件下、从大规模动作空间中采取动作进行响应的场景中,这类方法的样本效率较低。尽管大语言模型(LLM)可以对安全状态空间进行语义推理,但高延迟与信任假设阻碍了具有吸引力的在线部署模式。我们提出了 Ask the Expert,一个训练期指导框架:它首先对困难的网络防御状态进行总结,然后通过一个受约束的动作接口间歇性地向 LLM 查询主机级防御建议,最后将这些建议转化为分层奖励塑形,以供 PPO 使用。由于 LLM 在训练后被丢弃,部署时是一个纯粹的 RL 策略。在 TTCP CAGE CC1 和 CC2 以及两种攻击者类型上,这种非对称设计相比 PPO 提升了样本效率,并优于所评估的基于势能的奖励塑形(PBRS)基线,同时保持了最强的终端均值,且在部署时不需要任何 LLM 依赖。
cs.CR / 48 / 2610.09553
MARS: Malware Analysis with Rule-Based Scoring of LLM Claims
MARS:基于规则的 LLM 声明评分的恶意软件分析
large language model
大语言模型相关
Abstract
Large language models can triage malware through direct verdicts or behavioral claims scored by an external policy. We present MARS, a malware triage framework, and compare direct classification with single-pass claim scoring using the same evidence collector and identical static evidence bundles for each model. The evaluation covers 1,195 PE and ELF binaries grouped into 1,001 near-duplicate clusters and six language models, with deterministic rules providing a baseline. Direct classification is more accurate for all six models. On samples with usable outputs from both paths, its accuracy advantage ranges from 3.7 to 20.9 percentage points, with all 95% cluster-bootstrap confidence intervals for the differences above zero. It also achieves higher malicious alert recall in ten of twelve platform and model combinations. Claim mediation provides no consistent reduction in performance variation across models. Separate subset studies find more consistent alert decisions for direct classification and a larger recall loss for the claim path when predefined indicator fields are removed. In a family identification probe, claims yield higher accuracy than verdict labels but lower accuracy than evidence text. Retained claims expose the inputs to verdict computation and permit policy revision without another model call. We reproduce archived verdicts exactly and apply a revised policy to the same records, including outputs from two additional models withdrawn by their provider. Under the evaluated claim taxonomy and additive policy, these results favor direct classification when only a verdict is required, while demonstrating that retained claims support explicit policy inspection and revision.
Chinese Translation
大型语言模型可以通过直接裁决或由外部策略评分的行为性声明来对恶意软件进行分诊。我们提出了 MARS,一个恶意软件分诊框架,并在使用相同证据收集器和相同静态证据包的情况下,对每个模型比较直接分类与单遍声明评分。该评估涵盖 1,195 个 PE 和 ELF 二进制文件,被归入 1,001 个近重复簇,以及六个语言模型,并以确定性规则提供基线。对于全部六个模型,直接分类的准确率更高。在两条路径均产生可用输出的样本上,其准确率优势范围为 3.7 至 20.9 个百分点,所有差异的 95% 簇自助法置信区间均高于零。在十二种平台与模型组合中,有十种组合中它也取得了更高的恶意告警召回率。声明中介并未在跨模型性能变异方面提供一致的降低。分别进行的子集研究发现,直接分类的告警决策更为一致;而当预定义指标字段被移除时,声明路径的召回率损失更大。在家族识别探测中,声明比裁决标签取得更高的准确率,但比证据文本的准确率更低。保留的声明暴露了裁决计算的输入,并允许在不再次调用模型的情况下修订策略。我们精确复现了归档裁决,并将修订后的策略应用于相同记录,包括来自其提供方已撤回的两个额外模型的输出。在所受评估的声明分类体系与加性策略下,当仅需要裁决时,这些结果支持直接分类,同时表明保留的声明支持显式的策略检查与修订。
cs.CR / 49 / 2610.09906
Constrained-Action AI Remediation for SIEM/XDR via a NeMo-Guardrails Proxy
通过 NeMo-Guardrails 代理为 SIEM/XDR 实现受限动作 AI 修复
large language model
大语言模型相关
Abstract
Security Operations Centers (SOCs) for information technology and operational technology share one incident-response problem: a flood of correlated alerts and too few analysts. Large Language Models (LLMs) are increasingly proposed as reasoning engines that triage alerts and, in autonomous deployments, issue commands that block IPs, kill processes, or quarantine files on production hosts. This coupling introduces a new risk: a single adversarial alert can become a remote code path through the LLM's reasoning, leading it to recommend an action the SOC then executes. We present a constrained-action architecture with two coordinated layers: (i) a SIEM/XDR control plane that grounds remediation in correlated host events and confines the LLM's output to a closed intent vocabulary whose templated commands are executed by thin endpoint agents, backstopped by an argument validator; and (ii) a NeMo-Guardrails proxy that wraps the SOC-analyst LLM with input- and output-rail policies, evaluated out-of-the-box against a SOC-specific adversarial corpus we release. The stock proxy lifts injection recall from 25.0% to 94.5% at a 0.1% false-positive rate, and a live red-team exercise confirms that the closed intent vocabulary and argument validator contain the observed LLM failure modes before any command crosses the trust boundary. As an architectural fit (not yet a measured operational-technology deployment), the constrained-action property suits critical-infrastructure settings where a wrong remediation has physical, not merely operational, consequences. The loop is best run human-in-the-loop or delayed: the measured rail latency keeps inline control out of scope.
Chinese Translation
信息技术与运营技术的安全运营中心(SOC)共享同一个事件响应问题:大量关联告警与过少分析师。大型语言模型(LLM)正越来越多地被提议作为推理引擎,用于对告警进行分诊,并且在自主部署中发出命令,以在生产主机上封禁 IP、终止进程或隔离文件。这种耦合引入了一种新风险:单条对抗性告警可以经由 LLM 的推理成为一条远程代码路径,导致其推荐某个动作,而 SOC 随后执行该动作。我们提出一种受限动作架构,具有两个协同层:(i) 一个 SIEM/XDR 控制平面,其将修复建立在关联主机事件之上,并将 LLM 的输出限制在一个封闭意图词汇表中,该词汇表的模板化命令由轻量端点代理执行,并由参数验证器作为后备保障;以及 (ii) 一个 NeMo-Guardrails 代理,其用输入和输出护栏策略包裹 SOC 分析师 LLM,并针对我们发布的一个 SOC 专用对抗语料库进行开箱即用评估。现成代理将注入召回率从 25.0% 提升至 94.5%,假阳性率为 0.1%,并且一次实时红队演练证实,在任何命令跨越信任边界之前,封闭意图词汇表和参数验证器能够遏制观察到的 LLM 失效模式。作为一种架构适配(尚不是经过测量的运营技术部署),受限动作特性适合关键基础设施环境,在这些环境中,错误的修复会产生物理后果,而不仅仅是运营后果。该循环最好以人在回路或延迟方式运行:测得的护栏延迟使内联控制不在范围之内。
cs.CR / 50 / 2610.10150
On the Reliability of LLM-Based Vulnerability Patching Benchmarks
论基于LLM的漏洞修补基准测试的可靠性
large language model
大语言模型相关
Abstract
Large language models (LLMs) have shown strong potential for automated vulnerability patching, but current benchmarks can substantially distort reported performance. Drawing on extensive experience developing, running, and stress-testing such frameworks, we identify under-examined pitfalls across three dimensions: (1) agent-level factors, where prompting, tool availability, and detailed instructions can raise success rates without improving developer-aligned patch quality; (2) framework-level factors, where permission errors, infrastructure bugs, and timeout handling can silently suppress or inflate performance; and (3) dataset-level factors, where bug reports and single proof-of-concept (PoC) tests fail to capture whether patches address root causes or follow developer intent. We curate 112 historical bugs from 84 open-source C/C++, Go, and Rust projects, each with PoC tests, regression tests, and additional developer tests that assess alignment with the original developers' design principles. Through controlled experiments and case studies, we show that LLMs can achieve high PoC passing rates under ideal conditions, yet benchmark execution choices can materially change measured success. More importantly, developer-test passing rates remain low and improve only marginally with newer models, suggesting that models increasingly suppress symptoms without consistently producing upstream-quality fixes. These results show that benchmark scores are highly sensitive to evaluation design, and we provide practical guidelines for more rigorous, reliable, and reproducible evaluation.
Chinese Translation
大语言模型(LLM)在自动化漏洞修补方面展现出强大潜力,但当前的基准测试可能严重歪曲所报告的性能。基于开发、运行和压力测试此类框架的丰富经验,我们从三个维度识别出尚未得到充分考察的陷阱:(1)智能体层面因素,其中提示、工具可用性和详细指令可能提高成功率,却不会改善与开发者意图一致的补丁质量;(2)框架层面因素,其中权限错误、基础设施缺陷和超时处理可能悄无声息地压低或抬高性能;(3)数据集层面因素,其中缺陷报告和单一概念验证(PoC)测试无法捕捉补丁是否解决根本原因或遵循开发者意图。我们整理了来自84个开源C/C++、Go和Rust项目的112个历史缺陷,每个缺陷都配有PoC测试、回归测试以及额外的开发者测试,用以评估其与原始开发者设计原则的一致性。通过受控实验和案例研究,我们表明LLM在理想条件下可以达到较高的PoC通过率,但基准测试的执行选择可能实质性改变所测得的成功率。更重要的是,开发者测试通过率仍然较低,并且随着更新的模型仅有边际改善,这表明模型越来越多地抑制症状,而未能持续产生上游质量的修复。这些结果表明,基准测试分数对评估设计高度敏感,并且我们为更严谨、可靠和可复现的评估提供了实用指南。
cs.CR / 51 / 2610.10345
SLDR: Defending Against Malicious Fine-tuning via Selective Layers Recovery and Dynamic Routing
SLDR:通过选择性层恢复和动态路由防御恶意微调
large language model
大语言模型相关
Abstract
Fine-tuning-as-a-service enables users to adapt aligned large language models (LLMs) to specialized tasks, but malicious fine-tuning can erode refusal behavior while preserving task performance on legitimate inputs. We revisit recent layer-wise safety diagnostics and find that safety sensitivity is signed: scaling different layers can strengthen refusal, weaken it, or have little effect. Motivated by this observation, we propose SLDR, a post-fine-tuning defense based on Selective Layers Recovery and Dynamic Routing. SLDR trains a LoRA recovery adapter only on the layers with the maximum and minimum sensitivity scores in the signed spectrum, and uses representation-based dynamic routing inference to activate the adapter only for malicious queries. Across four model architectures, five downstream tasks, and four harmful benchmarks, SLDR substantially reduces harmful outputs while preserving downstream utility. On Llama3.1/SST2, SLDR reduces the average harmful score from 11.54 to 0.08 while maintaining downstream accuracy, and the harmful score remains near zero under poisoning ratios up to 0.9. The code is available at https://github.com/Stardust457/SLDR.
Chinese Translation
微调即服务(Fine-tuning-as-a-service)使用户能够将对齐后的大语言模型(LLMs)适配到专门任务,但恶意微调可能侵蚀拒答行为,同时在合法输入上保持任务性能。我们重新审视近期的逐层安全诊断,并发现安全敏感性是带符号的:缩放不同层可以增强拒答、削弱拒答,或几乎没有影响。受这一观察启发,我们提出了 SLDR,一种基于选择性层恢复和动态路由的微调后防御方法。SLDR 仅在带符号谱中具有最大和最小敏感性分数的层上训练一个 LoRA 恢复适配器,并使用基于表示的动态路由推理,仅为恶意查询激活该适配器。在四种模型架构、五个下游任务和四个有害基准上,SLDR 大幅减少有害输出,同时保持下游效用。在 Llama3.1/SST2 上,SLDR 将平均有害分数从 11.54 降至 0.08,同时保持下游准确率,并且在投毒比例高达 0.9 时,有害分数仍接近零。代码可在 https://github.com/Stardust457/SLDR 获取。
cs.AI / 52 / 2610.08954
RACER: Reflective Agent Coupling Query Interpretation and Tool-Based Retrieval for Frame Selection in Long Video Understanding
RACER:耦合查询解释与基于工具的检索的反思型智能体,用于长视频理解中的帧选择
large language model
大语言模型相关
Abstract
Video large language models (Vid-LLMs) excel at diverse video-language tasks by reasoning over selected frames. However, frame selection for long videos remains challenging, as it requires retrieving relevant frames distributed across segments from a large candidate pool given complex queries. This paper investigates dominant approaches to long-video frame selection from a task-decomposition perspective, identifying two key challenges: the Query Comprehension Gap in similarity-based methods and the Interpretation--Selection Gap in judgment-based methods. To address them, we propose RACER, a training-free reflective agentic framework that decomposes long-video frame selection into query interpretation driven by a lightweight Vid-LLM and evidence localization supported by an embedding model serving as a retrieval tool. Specifically, the Vid-LLM is responsible solely for reformulating the complex query into sub-queries that make implicit information requirements explicit, mitigating the Query Comprehension Gap. Meanwhile, the retrieval tool leverages these sub-queries to localize relevant evidence, relieving the Vid-LLM of direct frame selection and thus addressing the Interpretation--Selection Gap. Finally, the retrieved frames are fed back to the Vid-LLM for sub-query refinement, forming a reflection loop that iteratively improves query interpretation and frame selection. Experiments across multiple benchmarks show that RACER consistently improves long video understanding. Notably, RACER achieves effective frame selection even with limited-capability components, demonstrating that agentic integration enables these components to enhance more capable Vid-LLMs.
Chinese Translation
视频大语言模型(Vid-LLM)通过对所选帧进行推理,擅长处理各类视频-语言任务。然而,长视频的帧选择仍然具有挑战性,因为给定复杂查询时,它需要从大规模候选池中检索分布在各个片段中的相关帧。本文从任务分解的视角研究长视频帧选择的主流方法,识别出两个关键挑战:基于相似度的方法中的查询理解鸿沟(Query Comprehension Gap),以及基于判断的方法中的解释--选择鸿沟(Interpretation--Selection Gap)。为解决这两个问题,我们提出 RACER,一个无需训练的反思型智能体框架,它将长视频帧选择分解为由轻量级 Vid-LLM 驱动的查询解释,以及由充当检索工具的嵌入模型所支持的证据定位。具体而言,Vid-LLM 仅负责将复杂查询重构为使隐含信息需求显式化的子查询,从而缓解查询理解鸿沟。与此同时,检索工具利用这些子查询来定位相关证据,使 Vid-LLM 免于直接进行帧选择,从而解决解释--选择鸿沟。最后,检索到的帧被反馈给 Vid-LLM 以进行子查询细化,形成一个反思循环,迭代地改进查询解释与帧选择。在多个基准上的实验表明,RACER 能够持续提升长视频理解性能。值得注意的是,即便使用能力有限的组件,RACER 也能实现有效的帧选择,这表明智能体式的集成能够使这些组件增强能力更强的 Vid-LLM。
cs.AI / 53 / 2610.09015
Personalize at Test Time: Learning User Preferences for Image Generation
在测试时个性化:学习用于图像生成的用户偏好
diffusion
扩散模型相关
Abstract
Diffusion models can generate high-quality images, yet aligning their outputs with individual user preferences remains challenging. A key bottleneck is accurately modeling diverse user preferences from limited feedback. Existing approaches often rely on labor-intensive manual preference annotations or vision-language models (VLM) to extract preference information from user interaction histories, introducing substantial annotation or computational costs that limit scalability. We propose an approach that learns personalized reward models directly from users' historical image preference pairs. First, we use an autoencoder to compress hundreds of visual attributes into 50 attribute-anchored preference dimensions and train an evaluator to score images along these dimensions. We then represent each user's preferences as a linear combination of the shared dimension scores, estimating the user-specific weights by maximizing the likelihood of their observed pairwise preferences under the Bradley-Terry model. This formulation reduces per-user adaptation to optimizing a low-dimensional weight vector, simplifying optimization and enabling data-efficient personalization from sparse feedback. The learned personalized rewards guide image generation at inference time while keeping the diffusion model frozen. Experiments on real-user preference data show that our approach achieves approximately 77% held-out pairwise preference prediction accuracy and improves the alignment of generated images with individual user preferences.
Chinese Translation
扩散模型能够生成高质量图像,然而使其输出与个体用户偏好保持一致仍然具有挑战性。一个关键瓶颈在于从有限的反馈中准确建模多样化的用户偏好。现有方法通常依赖于劳动密集型的人工偏好标注,或依赖视觉语言模型(VLM)从用户交互历史中提取偏好信息,这带来了大量的标注成本或计算成本,从而限制了可扩展性。我们提出一种方法,直接从用户的历史图像偏好对中学习个性化奖励模型。首先,我们使用一个自编码器将数百个视觉属性压缩为 50 个以属性为锚的偏好维度,并训练一个评估器沿这些维度对图像进行打分。然后,我们将每个用户的偏好表示为共享维度得分的线性组合,并通过在 Bradley-Terry 模型下最大化其观测到的成对偏好的似然来估计用户特定的权重。该形式将每个用户的适配简化为优化一个低维权重向量,从而简化了优化,并能够从稀疏反馈中实现数据高效的个性化。学习到的个性化奖励在推理时引导图像生成,同时保持扩散模型冻结。在真实用户偏好数据上的实验表明,我们的方法达到了约 77% 的留出成对偏好预测准确率,并提升了生成图像与个体用户偏好的一致性。
cs.LG / 54 / 2610.09363
Closing the Loop on Contrail Avoidance with Satellite Verification
以卫星验证闭合凝结尾迹规避的闭环
diffusion
扩散模型相关
Abstract
Contrails are the thin ice clouds that aircraft leave behind. They cause a large share of aviation's warming, and rerouting the few flights that produce them could avoid much of it. However, an avoided contrail only counts if a satellite can confirm that it never formed, and this check is hard: contrails are one to two pixels wide, cover only 0.18% of pixels, and look very similar to natural cirrus. We build a small diffusion model (8.4M parameters, trained on one GPU) that detects them, and we run a controlled study to find out which components matter. The model reaches 0.476 PR-AUC, compared with 0.414 for a DeepLabV3+ baseline and 0.119 for an adapted MedSegDiff. Doubling the input resolution of the CNN brings it to parity (0.499, p=0.07). Three lessons apply beyond contrails. First, check the input resolution before designing a new architecture. Second, simple flips and rotations more than double accuracy and matter more than any architectural choice we measured. Third, pretraining the model on contrail shapes is harmful: the model learns that thin strokes appear everywhere and paints them onto empty scenes. Precision collapses to 1% while recall-based metrics still rate the degraded model as excellent, and no threshold or guidance heuristic repairs this failure.
Chinese Translation
凝结尾迹是飞机留下的薄冰云。它们造成了航空业变暖的很大一部分,而让少数产生凝结尾迹的航班改道就可以避免其中很大一部分。然而,只有当卫星能够确认凝结尾迹从未形成时,被避免的凝结尾迹才算数,而这一核查很难:凝结尾迹宽仅一到两个像素,只覆盖 0.18% 的像素,并且看起来与天然卷云极为相似。我们构建了一个用于检测它们的小型扩散模型(8.4M 参数,在单块 GPU 上训练),并开展了一项受控研究,以找出哪些组件起着关键作用。该模型达到 0.476 PR-AUC,相比之下 DeepLabV3+ 基线为 0.414,改造后的 MedSegDiff 为 0.119。将 CNN 的输入分辨率提高一倍可使其达到同等水平(0.499,p=0.07)。有三条经验教训适用于凝结尾迹之外。第一,在设计新架构之前先检查输入分辨率。第二,简单的翻转和旋转使准确率提高一倍以上,其重要性超过我们所测量的任何架构选择。第三,在凝结尾迹形状上对模型进行预训练是有害的:模型会学到细线条无处不在,并将它们涂抹到空白场景上。精确率暴跌至 1%,而基于召回率的指标仍将退化后的模型评为优秀,且没有任何阈值或引导启发式方法能修复这一失败。
cs.AI / 55 / 2610.09440
Mixture of Layers: Dynamic Layer Routing for Visual Reasoning
层混合:用于视觉推理的动态层路由
large language model
大语言模型相关
Abstract
Pre-trained vision encoders contain layer-wise visual representations that differ in spatial granularity, semantic abstraction, and sensitivity to local details. However, most Multimodal Large Language Models (MLLMs) rely on only the final or penultimate vision encoder representations or fixed aggregation rules, making visual abstraction largely query-agnostic and limiting access to fine-grained cues such as small objects, spatial details, text, and subtle visual attributes. In this work, we propose Mixture of Layers (MoL), an instruction-conditioned layer routing approach at the visual patch level that dynamically aggregates query-relevant latent representations from intermediate vision encoder layers. Given a text query, MoL predicts routing probabilities over vision encoder layers and performs a top-k sparse aggregation over selected hidden states at either the image level, patch level, or through a hybrid routing mechanism. In doing so, MoL enables query-adaptive access to layer-specific visual features for fine-grained visual reasoning. Our experiments across 7 fine-grained visual reasoning tasks demonstrate substantial performance improvements, especially across fine-grained visual grounding and understanding tasks such as +18.9% improvement on V* in overall accuracy, +4.5% on HRBench4K, and +16.3% on CharXiv compared to the baseline MLLMs, without resorting to multi-resolution inputs, simple interleaving of multiple vision encoders, or increasing the number of patch tokens. We study vision encoders' receptive field scales across different layers and their sampling behaviors to provide an in-depth analysis of why layer-wise sampling is helpful, demonstrating that conditional visual representations are a key step towards better visual perception and reasoning in MLLMs. Our project page is available at https://wjdghks950.github.io/mol.github.io/.
Chinese Translation
预训练的视觉编码器包含逐层的视觉表示,这些表示在空间粒度、语义抽象和对局部细节的敏感性方面各不相同。然而,大多数多模态大语言模型(MLLMs)仅依赖于最终或倒数第二层的视觉编码器表示,或固定的聚合规则,这使得视觉抽象在很大程度上与查询无关,并限制了对细粒度线索的访问,例如小物体、空间细节、文本和细微的视觉属性。在这项工作中,我们提出了层混合(Mixture of Layers,MoL),一种在视觉 patch 级别的指令条件化层路由方法,它动态聚合来自中间视觉编码器层的与查询相关的潜在表示。给定一个文本查询,MoL 预测视觉编码器各层上的路由概率,并对选定的隐藏状态执行 top-k 稀疏聚合,聚合可以是在图像级别、patch 级别,或通过混合路由机制进行。通过这样做,MoL 能够为细粒度视觉推理提供对特定层视觉特征的查询自适应访问。我们在 7 个细粒度视觉推理任务上的实验表明性能有了显著提升,尤其是在细粒度视觉定位和理解任务上,例如与基线 MLLMs 相比,在 V* 上的总体准确率提升了 +18.9%,在 HRBench4K 上提升了 +4.5%,在 CharXiv 上提升了 +16.3%,而无需采用多分辨率输入、多个视觉编码器的简单交错,或增加 patch token 的数量。我们研究了不同层中视觉编码器的感受野尺度及其采样行为,以深入分析为何逐层采样是有帮助的,表明条件化视觉表示是朝着 MLLMs 中更好的视觉感知和推理迈出的关键一步。我们的项目页面可在 https://wjdghks950.github.io/mol.github.io/ 获取。
cs.AI / 56 / 2610.09450
Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
Iris-3B:以像素空间扩散训练、转换与微调超越潜空间
diffusion
扩散模型相关
Abstract
Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters. We test this claim along both routes to a pixel-space backbone. We pretrain Iris-3B, a 3B-parameter pixel-space text-to-image transformer, from scratch through a $256\to512\to1024$ curriculum, after first ablating the prediction target and representation alignment at $256^2$ to decide what to scale. We also convert a pretrained latent model, FLUX.2 Klein base 4B, to pixel space. We fine-tune both families for monocular depth estimation and for image restoration/super-resolution. We find no significant improvement from using a pixel-space generative prior. Fine-tuned for depth with one matched direct-regression recipe, Iris-3B is level with the latent FLUX.2 Klein and the converted pixel FLUX.2 Klein falls behind it, and on $4\times$ DIV2K restoration neither pixel model beats a latent FLUX.2 Klein fine-tune, the converted one trailing it slightly. We document the recipes, the failure modes and the remaining confounds behind this negative result. Nevertheless, Iris-3B shows that pixel-space pretraining with the pixel-transformer (PiT) head of PixelDiT scales to 3B parameters and to text-to-image quality competitive with latent models, matching Qwen-Image on OneIG under the official evaluators at $1024^2$. We release its weights and training code in the hope that they help pave the way for further work on pixel-space generation.
Chinese Translation
像素空间扩散模型避免了潜空间模型中有损的 VAE,这暗示其在细粒度细节至关重要的下游任务上具有优势。我们沿着通往像素空间主干网络的两条路径检验了这一说法。我们首先在 $256^2$ 上对预测目标和表征对齐进行消融,以决定应当对什么进行规模化,随后从头预训练了 Iris-3B——一个 30 亿参数的像素空间文本到图像 Transformer,并采用了 $256\to512\to1024$ 的课程式训练流程。我们还将一个预训练的潜空间模型 FLUX.2 Klein base 4B 转换为像素空间。我们针对单目深度估计以及图像恢复/超分辨率对这两个系列都进行了微调。我们发现,使用像素空间生成先验并未带来显著改进。在使用同一套匹配的直接回归方案针对深度进行微调时,Iris-3B 与潜空间的 FLUX.2 Klein 持平,而转换后的像素空间 FLUX.2 Klein 则落后于它;在 $4\times$ DIV2K 恢复任务上,两个像素空间模型都未能胜过微调后的潜空间 FLUX.2 Klein,其中转换而来的那个模型略微落后于它。我们记录了这一负面结果背后的方案、失败模式以及尚存的混淆因素。尽管如此,Iris-3B 表明,使用 PixelDiT 的像素 Transformer(PiT)头进行像素空间预训练可以扩展到 30 亿参数,并达到与潜空间模型相竞争的文本到图像质量,在 $1024^2$ 下、于官方评测器上在 OneIG 上与 Qwen-Image 持平。我们发布了其权重和训练代码,希望它们有助于为像素空间生成的进一步工作铺平道路。
cs.LG / 57 / 2610.09677
A Multi-Source Ultrasound Benchmark Revealing the Limits of Contemporary Self-Supervised Anomaly Detection Methods
一个多源超声基准,揭示当代自监督异常检测方法的局限
diffusion
扩散模型相关
Abstract
Self-supervised anomaly detection is a promising paradigm for medical ultrasound, as normal images are often easier to obtain than exhaustive annotations of all possible pathologies. However, most existing evaluations are limited to a single anatomy or task, making it unclear whether models learn a robust notion of normal ultrasound appearance or only a source-specific representation. We introduce the SADUSI benchmark, a multi-source ultrasound dataset designed to train and evaluate anomaly detection methods across a broad range of anatomical regions, views, and acquisition protocols. The goal of SADUSI is to provide a diverse normal ultrasound distribution and a benchmark for visible structural anomalies that can be assessed from single images. We evaluate representative self-supervised anomaly detection methods and find that current approaches struggle in this setting. In particular, reconstruction-based diffusion methods such as AnoDDPM and DeCo-Diff achieve pixel-level AUROC values of 0.56-0.72 and maximum F1 scores of 0.10-0.26, indicating limited separation of pathology from normal image regions. Feature-based PatchCore variants perform better, reaching pixel-level AUROC values of 0.76-0.83, but remain limited with maximum F1 scores of 0.14-0.40. These findings suggest that broad multi-source ultrasound anomaly detection remains an open challenge and that SADUSI can serve as a resource for developing methods that generalize beyond anatomy-specific settings.
Chinese Translation
自监督异常检测对于医学超声是一种有前景的范式,因为正常图像通常比对所有可能病理的详尽标注更容易获得。然而,大多数现有评估仅限于单一解剖结构或任务,这使得不清楚模型学到的是正常超声表现的稳健概念,还是仅学到源特定的表示。我们提出 SADUSI 基准,这是一个多源超声数据集,旨在跨广泛的解剖区域、视图和采集协议训练和评估异常检测方法。SADUSI 的目标是提供一个多样化的正常超声分布,以及一个针对可从单张图像评估的可见结构异常的基准。我们评估了代表性的自监督异常检测方法,并发现当前方法在这种设置下表现不佳。特别地,诸如 AnoDDPM 和 DeCo-Diff 等基于重建的扩散方法达到 0.56-0.72 的像素级 AUROC 值和 0.10-0.26 的最大 F1 分数,表明病理与正常图像区域的分离有限。基于特征的 PatchCore 变体表现更好,达到 0.76-0.83 的像素级 AUROC 值,但最大 F1 分数为 0.14-0.40,仍然有限。这些发现表明,广泛的多源超声异常检测仍然是一个开放挑战,并且 SADUSI 可以作为开发能够泛化到特定解剖设置之外的方法的资源。
cs.LG / 58 / 2610.09841
ORCA: Hunting Compositional Failures in Text-to-Image Diffusion
ORCA:追踪文本到图像扩散模型中的组合性失败
diffusion
扩散模型相关
Abstract
Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count. Recent architectures already augment CLIP with a T5 encoder precisely because CLIP's contrastive embedding loses compositional structure, yet these failures persist. We argue the binding problem is therefore not one of missing information but of misaligned information: a text encoder preserves compositional structure, but in a representation space shaped by language modelling rather than vision, and the denoising objective does not directly reward aligning the two. We show this correspondence can be supplied as an explicit training signal, that the relevant cross-modal information is concentrated in a low-rank subspace of self-supervised visual features, and that supplying it can be folded into diffusion training as a single auxiliary loss. Our method, ORCA (Orthogonal Residual Compositional Alignment), aligns the latent of a diffusion transformer with a low-rank target derived from a frozen visual encoder, through a predictor whose orthogonal basis is parameterised by a learned residual between T5 and CLIP embeddings, which provides a prompt-dependent signal for selecting the visual readout subspace. We prove that the cross-modal information recoverable at a given rank is bounded by the spectral mass of the visual encoder's covariance in the top components. Across three diffusion-transformer backbones (DiT-B/2, DiT-L/2, U-ViT-L), ORCA improves FID and GenEval over both vanilla and REPA baselines at zero inference-time cost; on DiT-L/2 it reaches FID 16.65 and GenEval 0.291 at 200K steps, exceeding the strongest 400K baseline at half the training cost, with the largest gains concentrated on attribute binding, spatial relations, and multi-object prompts.
Chinese Translation
文本到图像扩散模型在组合性提示上会以可预测的方式失败:属性绑定到错误的对象上,空间关系发生颠倒,多对象场景丢失计数。近期的架构已经用 T5 编码器来增强 CLIP,正是因为 CLIP 的对比嵌入会丢失组合结构,然而这些失败依然存在。我们认为,因此绑定问题并非信息缺失的问题,而是信息错位的问题:文本编码器保留了组合结构,但其所处的表示空间是由语言建模而非视觉所塑造的,而去噪目标并不直接奖励将二者对齐。我们表明,这种对应关系可以作为显式的训练信号被提供,相关的跨模态信息集中在自监督视觉特征的一个低秩子空间中,并且提供该信号可以以单一辅助损失的形式并入扩散训练。我们的方法 ORCA(正交残差组合对齐,Orthogonal Residual Compositional Alignment)通过一个预测器,将扩散 Transformer 的隐变量与由冻结视觉编码器导出的低秩目标对齐,该预测器的正交基由 T5 与 CLIP 嵌入之间学习到的残差参数化,从而为选择视觉读出子空间提供了依赖于提示的信号。我们证明,在给定秩下可恢复的跨模态信息,受到视觉编码器协方差在前若干主成分上的谱质量的约束。在三种扩散 Transformer 骨干网络(DiT-B/2、DiT-L/2、U-ViT-L)上,ORCA 在零推理时间成本下,相较于原始基线和 REPA 基线均提升了 FID 和 GenEval;在 DiT-L/2 上,它在 200K 步时达到 FID 16.65 和 GenEval 0.291,以一半的训练成本超过了最强的 400K 基线,其中最大的增益集中在属性绑定、空间关系和多对象提示上。
cs.AI / 59 / 2610.10066
From Pixel to Coding: Evaluating the Figure Reproduction Capabilities of MLLMs
从像素到编码:评估多模态大语言模型的图形复现能力
large language model
大语言模型相关
Abstract
Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in both visual understanding and code generation. However, existing benchmarks typically evaluate these two modalities in isolation, lacking a dedicated assessment of their unification, i.e., how a model can perceive complex visual structures and synthesize them into precise, executable code. Moreover, current visual code generation benchmarks often rely on simplified layouts within single programming environments, falling short of evaluating true unified multimodal reasoning. To bridge this gap, we propose FigCodeBench, a comprehensive framework for rigorously evaluating MLLMs on figure reproduction, integrating multimodal comprehension and generation. We first design a systematic dataset construction pipeline, resulting in a total of 6,194 instances that cover 7 functional categories and 4 types of programming languages. We further categorize figure reproduction into three tiers with visual and code complexity modeling, specifically targeting complex structural reasoning, varying aspect ratios, and dense geometric constraints. We introduce a multi-dimensional evaluation protocol, encompassing visual fidelity and syntactic isomorphism, that aligns highly with the Mean Machine Opinion Score (MMOS) and human preferences. Based on our framework, we conducted extensive experiments on 24 widely used proprietary and open-source MLLMs (e.g., Gemini 3.1 Pro, GPT-5.4, and Kimi-K2.5), where we observed a universal, non-linear performance cliff across different programming languages and difficulty scenarios for all models, and gained several insights, such as the significant metric decline in rigid declarative languages.
Chinese Translation
多模态大语言模型(MLLMs)在视觉理解与代码生成两方面均已展现出令人印象深刻的能力。然而,现有的基准测试通常将这两种模态孤立地加以评估,缺乏对二者统一性的专门评估,即模型如何感知复杂的视觉结构并将其合成为精确、可执行的代码。此外,当前的视觉代码生成基准往往依赖于单一编程环境中的简化布局,未能评估真正统一的多模态推理。为弥合这一差距,我们提出了 FigCodeBench,这是一个用于严格评估 MLLMs 在图形复现任务上表现的综合框架,它整合了多模态理解与生成。我们首先设计了一套系统化的数据集构建流程,最终得到总计 6,194 个实例,覆盖 7 个功能类别和 4 种编程语言。我们进一步将图形复现划分为三个层级,并进行视觉与代码复杂度建模,分别针对复杂的结构推理、不同的宽高比以及密集的几何约束。我们引入了一套多维度的评估协议,涵盖视觉保真度与句法同构性,该协议与平均机器意见分数(MMOS)以及人类偏好高度一致。基于我们的框架,我们对 24 个广泛使用的专有与开源 MLLMs(例如 Gemini 3.1 Pro、GPT-5.4 和 Kimi-K2.5)进行了大量实验,在实验中我们观察到,所有模型在不同编程语言和难度场景下普遍存在非线性的性能断崖,并获得了一些洞见,例如在严格声明式语言中指标的显著下降。
cs.AI / 60 / 2610.09064
Talking with Language Models
与语言模型交谈
large language model
大语言模型相关
Abstract
When we interact with large language models (LLMs), are we having a conversation? They are designed to invite us to treat them as intelligent interlocutors who remember, act, and make commitments. But appearances deceive. We introduce the artifactual stance, a framework that reconceives human-AI interaction as artifact-mediated exchanges of candidate texts. LLM outputs are candidate texts optimized for utility, not utterances bearing meaning or force. LLMs are sophisticated text generators, not speakers. Between sessions, nothing runs; between turns, no one remembers. What persists is a configuration and a transcript. The "conversation" is a user's solo performance, interpretive labour disguised by interface and artifact design. This shift dissolves recent philosophical puzzles. Questions about what 'I' and 'you' refer to in AI exchanges, about whether systems can lie or be held to promises, about the identity of our supposed interlocutors all rest on a false presupposition. There is no speaker behind the screen, hence no one to refer to, no one to hold responsible. What feels like dialogue with someone is interaction with an artifact that generates text at unprecedented scale and fit. By abandoning the conversational framing, we see these systems for what they are: immensely sophisticated artifacts that afford varied uses. The philosophical questions that matter are about the normative underpinnings of design, adoption, authorization, and human practices of use.
Chinese Translation
当我们与大型语言模型(LLMs)互动时,我们是在进行一场对话吗?它们被设计来邀请我们将它们视为智能对话者,这些对话者会记忆、行动并做出承诺。但表象具有欺骗性。我们引入人工物立场,这是一个将人机交互重新构想为以人工物为中介的候选文本交换的框架。LLM 的输出是为效用而优化的候选文本,而不是承载意义或语力的话语。LLM 是复杂的文本生成器,而不是说话者。在会话之间,没有任何东西在运行;在轮次之间,没有人记得。持续存在的是一种配置和一份转录文本。所谓的“对话”是用户的独角戏,是被界面和人工物设计所伪装起来的解释性劳动。这一转变消解了近来的哲学谜题。关于在 AI 交流中“我”和“你”指什么的问题,关于系统是否能说谎或受承诺约束的问题,关于我们所以为的对话者的身份的问题,全都建立在一个错误预设之上。屏幕背后没有说话者,因此没有可指称的人,也没有可追究责任的人。感觉像是在与某人对话的东西,实际上是与一个人工物的互动,该人工物以前所未有的规模和契合度生成文本。通过放弃对话式框架,我们看清这些系统的本来面目:极其复杂的人工物,能够支持多种用途。真正重要的哲学问题,涉及设计、采用、授权以及人类使用实践的规范性基础。
cs.AI / 61 / 2610.09166
Lower Bounds for Parallel Diffusion Sampling
并行扩散采样的下界
diffusion
扩散模型相关
Abstract
Standard diffusion samplers generate samples through repeated evaluations of a learned score function. Parallel sampling methods seek to accelerate generation by trading additional evaluations for fewer sequential rounds. This raises the question of how much sequential dependence is unavoidable, even when many score queries can be made simultaneously. We establish the first polynomial parallel-round lower bounds for diffusion sampling with approximate scores. Specifically, we prove (1) a $\widetildeΩ(d^{1/3})$-round lower bound for sampling smooth, near-isotropic Gaussian mixtures in $R^d$, and (2) an $Ω(d)$-round lower bound for uniform sampling from anisotropic axis-aligned boxes contained in the unit ball. Both bounds hold for arbitrary randomized algorithms making polynomially many queries per round at arbitrary locations and noise levels, with inverse-polynomial score error and constant total variation accuracy. The linear bound is tight for our box family. Our constructions use fixed approximate score oracles that enforce sequential access to hidden information while satisfying the accuracy guarantee at every noise level.
Chinese Translation
标准扩散采样器通过重复评估学习到的得分函数来生成样本。并行采样方法试图通过以更多的评估换取更少的顺序轮数来加速生成。这引出了一个问题:即使可以同时进行许多得分查询,有多少顺序依赖是不可避免的。我们建立了首个针对具有近似得分的扩散采样的多项式并行轮数下界。具体而言,我们证明:(1) 对在 $R^d$ 中采样光滑、近各向同性高斯混合分布的一个 $\widetildeΩ(d^{1/3})$ 轮下界;以及 (2) 对从单位球内包含的各向异性轴对齐盒子中进行均匀采样的一个 $Ω(d)$ 轮下界。这两个下界对任意随机化算法都成立,这些算法每轮在任意位置和噪声水平上进行多项式多次查询,并具有逆多项式得分误差和常数全变差精度。该线性下界对我们的盒子族是紧的。我们的构造使用固定的近似得分预言机,这些预言机强制对隐藏信息进行顺序访问,同时在每个噪声水平上满足精度保证。
cs.LG / 62 / 2610.09985
Marrying Pricing and Advertising with LLMs
将定价和广告与大语言模型相结合
large language model
大语言模型相关
Abstract
We study a sequential pricing problem in which a seller jointly posts a price and an advertisement generated by a large language model (LLM). The seller aims to maximize revenue under an unknown product demand that depends on both decisions, while observing only whether each offer leads to a purchase. We propose an online actor-critic algorithm that combines low-rank adaptation (LoRA) of a pretrained LLM with a demand model fitted to available data. At each round, the actor generates an advertisement, and the critic estimates purchase probabilities to guide price selection. Then, the resulting feedback is used to update both the actor and the critic, with the critic's revenue estimates providing a baseline for policy gradient updates of the actor. To evaluate our approach, we develop an evaluation framework with three synthetic demand models and a demand simulator built from real-world marketplace data. Finally, we compare our algorithm with benchmarks that do not jointly optimize price selection and advertisement generation, achieving expected revenue gains over the reference policy of 5.69%, 5.18% and 55.96% under the three synthetic demand models and 5.81% under the marketplace simulator.
Chinese Translation
我们研究一个序贯定价问题,其中卖家同时发布一个价格和一个由大语言模型(LLM)生成的广告。卖家旨在在依赖于这两个决策的未知产品需求下最大化收益,同时仅观察到每次报价是否导致购买。我们提出一种在线演员-评论家算法,它将预训练LLM的低秩适应(LoRA)与拟合可用数据的需求模型相结合。在每轮中,演员生成一则广告,评论家估计购买概率以指导价格选择。然后,所得到的反馈用于更新演员和评论家,其中评论家的收益估计为演员的策略梯度更新提供基线。为了评估我们的方法,我们开发了一个评估框架,其中包含三个合成需求模型和一个由真实世界市场数据构建的需求模拟器。最后,我们将我们的算法与不联合优化价格选择和广告生成的基准进行比较,在三个合成需求模型下相对于参考策略分别实现了5.69%、5.18%和55.96%的预期收益提升,并在市场模拟器下实现了5.81%的提升。
cs.AI / 63 / 2610.09400
TutorLoop: Regulating Student Learning Behaviors via Sensor-in-the-Loop Generative Feedback
TutorLoop:通过传感器在环生成式反馈调节学生学习行为
large language model
大语言模型相关
Abstract
We present TutorLoop, a sensor-in-the-loop system that regulates student learning behaviors by delivering adaptive feedback based on real-time cognitive states. Unlike prior large language model (LLM) tutors that directly depend on scenario-specific content, TutorLoop operates on sensor-derived signals captured via webcams. Moreover, unlike direct cognitive-to-feedback mappings that are short-sighted, the system employs a deep reinforcement learning (DRL) agent to optimize the feedback type across the entire learning process. Finally, another LLM tutor refines feedback into human-like, context-aware messages. We evaluate TutorLoop in a large-scale user study (N=187), where a model trained offline is directly applied to a new learning task without retraining. Results show that TutorLoop provides less frequent yet more effective interventions, improving attention, reducing workload, increasing engagement, and ultimately enhancing learning outcomes. These findings highlight the potential of closed-loop, sensor-driven feedback for scalable human-AI integrated systems to support learning.
Chinese Translation
我们提出了 TutorLoop,这是一个传感器在环(sensor-in-the-loop)系统,它通过基于实时认知状态提供自适应反馈来调节学生的学习行为。与以往直接依赖特定场景内容的大型语言模型(LLM)辅导系统不同,TutorLoop 基于通过网络摄像头采集的传感器衍生信号运行。此外,与短视的直接认知到反馈映射不同,该系统采用深度强化学习(DRL)智能体在整个学习过程中优化反馈类型。最后,另一个 LLM 辅导系统将反馈细化为类人的、情境感知的消息。我们在一项大规模用户研究(N=187)中评估了 TutorLoop,其中离线训练的模型无需重新训练即可直接应用于新的学习任务。结果表明,TutorLoop 提供的干预频率更低,但效果更为有效,能够提高注意力、降低工作负荷、提升参与度,并最终改善学习成果。这些发现凸显了闭环、传感器驱动反馈在可扩展的人类—AI 集成学习支持系统中的潜力。
cs.CL / 64 / 2610.09724
Towards Explaining Query Expansion Performance in Information Retrieval
面向解释信息检索中的查询扩展性能
large language model
大语言模型相关
Abstract
Query Expansion (QE) techniques have long been widely used in Information Retrieval (IR) to address the vocabulary mismatch problem. They remain relevant in modern retrieval systems, including those based on large language models (LLMs). However, no single QE method consistently outperforms others across all queries. This work seeks to explain the variation in QE performance through two complementary perspectives. The first is the concept of an Ideal Expanded Query (IEQ)--a hypothetical query that maximizes retrieval effectiveness with a downstream BM25 retrieval model. The second is a separability perspective, which quantifies how distinctly relevant and non-relevant documents are scored for a given expanded query using Cohen's (d). We develop a separability measure and practical formulations to approximate the IEQ and investigate how these factors relate to retrieval effectiveness. Extensive experiments on the TREC Robust collection, TREC DL 2019-2022 passage collections, and TREC DL 2019-2020 document collections reveal several interesting patterns. In particular, we find that expanded queries that are closer to the ideal expanded query tend to achieve higher retrieval effectiveness. We further show that the separability of relevant and non-relevant documents provides a complementary perspective for understanding QE performance.
Chinese Translation
查询扩展(Query Expansion, QE)技术长期以来一直被广泛应用于信息检索(Information Retrieval, IR)中,以解决词汇不匹配问题。在现代检索系统中,包括那些基于大语言模型(Large Language Models, LLMs)的检索系统,它们依然具有相关性。然而,没有任何一种 QE 方法能够在所有查询上都始终优于其他方法。本工作试图通过两个互补的视角来解释 QE 性能的差异。第一个是理想扩展查询(Ideal Expanded Query, IEQ)的概念——即一个在下游 BM25 检索模型下能够最大化检索有效性的假设性查询。第二个是可分离性视角,它使用 Cohen's (d) 来量化对于给定的扩展查询,相关文档与非相关文档被评分的区分程度。我们提出了一种可分离性度量以及若干实用公式来近似 IEQ,并研究了这些因素与检索有效性之间的关系。在 TREC Robust 集合、TREC DL 2019-2022 段落集合以及 TREC DL 2019-2020 文档集合上进行的大量实验揭示了若干有趣的规律。特别地,我们发现,越接近理想扩展查询的扩展查询往往能够取得更高的检索有效性。我们进一步表明,相关文档与非相关文档的可分离性为理解 QE 性能提供了一个互补的视角。
cs.AI / 65 / 2610.10235
Beyond LLM-GA: Secure Fluid Antenna Systems with ReEvo-Designed Memetic Algorithm
超越 LLM-GA:采用 ReEvo 设计的模因算法的安全流体天线系统
large language model
大语言模型相关
Abstract
Fluid antenna systems (FASs) offer significant spatial flexibility, yet securing them against eavesdropping is critical for practical FAS deployment in military, satellite, and internet-of-things networks. Although large language model (LLM)-assisted genetic algorithms (LLM-GAs) can address this secure FAS port selection problem, whether further algorithmic improvement is possible warrants deeper investigation. To this end, we propose a memetic algorithm based on reflective evolution (ReEvo). Unlike the state-of-the-art LLM-GAs, which design only crossover or mutation operators with an LLM, our algorithm leverages an LLM to evolve dedicated crossover, mutation, and local-search operators offline. These operators are then embedded into a memetic search framework, thereby obviating any online LLM queries during execution. Simulation results at equal generation counts demonstrate that our proposed algorithm achieves a higher secure sum-rate than the conventional GA and the state-of-the-art LLM-GAs.
Chinese Translation
流体天线系统(FASs)提供了显著的空间灵活性,但确保其免受窃听对于在军事、卫星和物联网网络中实际部署 FAS 至关重要。尽管大语言模型(LLM)辅助的遗传算法(LLM-GAs)能够解决这一安全 FAS 端口选择问题,但是否可能实现进一步的算法改进仍值得更深入研究。为此,我们提出了一种基于反思进化(ReEvo)的模因算法。与仅利用 LLM 设计交叉或变异算子的最先进 LLM-GAs 不同,我们的算法利用 LLM 离线演化专用的交叉、变异和局部搜索算子。然后,这些算子被嵌入到模因搜索框架中,从而在执行期间避免了任何在线 LLM 查询。在相同代数下的仿真结果表明,我们提出的算法比传统 GA 和最先进的 LLM-GAs 实现了更高的安全总速率。
cs.LG / 66 / 2610.08959
GraphOPD: Graph-Augmented On-Policy Distillation for LLM Agents
GraphOPD:面向LLM智能体的图增强同策略蒸馏
large language model
大语言模型相关
Abstract
On-policy distillation post-trains large language model agents by supplying dense, step-level guidance from a teacher policy when the reinforcement-learning reward is sparse and arrives only once per trajectory. Existing instantiations allocate this guidance by the size of the teacher-student divergence at each step, on the single-turn intuition that a large disagreement marks a mistake worth correcting. Once decisions chain over many turns, that rule misfires, since an early drift enters every later context both policies condition on, leaving the teacher consistent with the drifted trajectory instead of flagging its cause, while interchangeable steps register large but outcome-irrelevant divergences. We demonstrate this on an agentic benchmark, where distilling the highest-divergence steps brings no consistent benefit over random selection. To this end, we introduce GraphOPD, the first method to bring graph-based structural augmentation into on-policy distillation for agent capabilities. It reads which steps enabled which later ones from the environment's own record of state changes, immune to the drift that corrupts the teacher-student gap, organizes them into a dependency graph, scores each step by a random-walk stationary distribution over it, and fuses that structural credit with the divergence signal into a trajectory-relative mask concentrating supervision on each rollout's highest-aptitude steps. Across three model scales and eleven baselines on ALFWorld, WebShop, and SearchQA, GraphOPD shows competitive performance throughout, improving over the strongest baseline by up to +5.8 pp. An executed-replay audit further shows that this structural credit score tracks true causal impact far above chance, that both fused signals are independently necessary, and that the same signal transfers to out-of-domain tool-integrated reasoning.
Chinese Translation
同策略蒸馏通过在强化学习奖励稀疏且每条轨迹仅到达一次时,由教师策略提供稠密的、步骤级指导,来对大型语言模型智能体进行后训练。现有的实现方式依据每一步教师-学生散度的大小来分配这种指导,其基于单轮直觉,即大的分歧标志着值得纠正的错误。一旦决策在多个回合中链式展开,该规则便会失效,因为早期的漂移会进入两个策略所条件化的每一个后续上下文,使教师与漂移后的轨迹保持一致,而不是标记出其成因,与此同时可互换的步骤却会记录到很大但与结果无关的散度。我们在一个智能体基准上证明了这一点:蒸馏最高散度的步骤相较于随机选择并未带来一致的收益。为此,我们提出了 GraphOPD,这是首个将基于图的结构增强引入面向智能体能力的同策略蒸馏的方法。它从环境自身的状态变化记录中读取哪些步骤促成了哪些后续步骤——该记录不受破坏教师-学生差距的那种漂移影响——将它们组织成一个依赖图,通过其上的随机游走平稳分布为每个步骤打分,并将该结构性贡献与散度信号融合成一个轨迹相对掩码,把监督集中在每次 rollout 中资质最高的步骤上。在 ALFWorld、WebShop 和 SearchQA 上,跨三个模型规模和十一个基线,GraphOPD 全程展现出有竞争力的性能,相较于最强基线最多提升 +5.8 个百分点。一项执行回放审计进一步表明,该结构性贡献分数对真实因果影响的追踪远超随机水平,两种融合信号各自都是独立必要的,并且同一信号可迁移到域外的工具集成推理。
cs.LG / 67 / 2610.08963
On KL-Regularized Policy Optimization
论 KL 正则化策略优化
large language model
大语言模型相关
Abstract
Asynchronous reinforcement learning (RL) for large language model (LLM) agents trains one policy on trajectories generated by another: rollouts come from stale checkpoints, and the inference engine's probabilities differ from the trainer's even at identical parameters. Standard remedies either clip importance ratios, which biases the update, or, as in GRPO, sample a group of responses per prompt, which is costly when episodes are long. We propose KL-Regularized Policy Optimization (KLPO), a framework that anchors the KL regularizer at the sampler. The regularized improvement step then has a closed-form Gibbs solution, and KLPO fits its log-ratio optimality condition by least squares on the sampler's own trajectories, so the sampler probability enters through a log-ratio and no importance weights are needed. Profiling out the regression intercept replaces the intractable log-partition function with the signal's sampler mean plus a sampler-to-trainer KL divergence. For token-level policy mirror descent targets, we show that the resulting gradient can be computed from terminal returns without a critic, via sampler-centered scores or a single trajectory residual, even under stochastic tool outputs. We further prove that independent Monte Carlo estimates of the KL term keep these gradients unbiased, derive the exact KL gap of cheaper top-$K$ and binary approximations, and show that SPPO, GPO, REBEL, and BPO arise as special cases of KLPO. The result is a critic-free update that uses one rollout per prompt and requires neither a learned normalizer nor a group of responses.
Chinese Translation
用于大语言模型(LLM)智能体的异步强化学习(RL)在由另一个策略生成的轨迹上训练一个策略:rollout 来自陈旧的检查点,并且即使参数相同,推理引擎的概率也与训练器的概率不同。标准补救措施要么裁剪重要性比率,这会使更新产生偏差;要么像 GRPO 那样,对每个提示采样一组回复,而这在回合较长时代价高昂。我们提出 KL 正则化策略优化(KLPO),这是一个将 KL 正则化项锚定在采样器处的框架。于是,正则化改进步骤具有闭式 Gibbs 解,并且 KLPO 在采样器自身的轨迹上通过最小二乘拟合其对数比率最优性条件,因此采样器概率以对数比率的形式进入,并且不需要重要性权重。消去回归截距后,用信号的采样器均值加上从采样器到训练器的 KL 散度,替代了难以处理的对数配分函数。对于 token 级策略镜像下降目标,我们表明,所得梯度可以在没有 critic 的情况下由终端回报计算,通过以采样器为中心的得分或单个轨迹残差,即使在随机工具输出下也是如此。我们进一步证明,KL 项的独立蒙特卡洛估计使这些梯度保持无偏,推导出更廉价的 top-$K$ 和二元近似的精确 KL 差距,并表明 SPPO、GPO、REBEL 和 BPO 作为 KLPO 的特例出现。其结果是一种无需 critic 的更新,它对每个提示使用一次 rollout,并且既不需要学习到的归一化器,也不需要一组回复。
cs.LG / 68 / 2610.08977
SNR-Gated LSTM-Conditioned Diffusion Model for MIMO Channel Estimation
用于MIMO信道估计的SNR门控LSTM条件扩散模型
diffusion
扩散模型相关
Abstract
Accurate and low latency channel estimation is critical for modern MIMO systems, particularly under mobility, where channels exhibit structured sparsity and strong temporal correlation. This paper proposes a time-series conditioned diffusion framework for channel estimation that performs denoising in the angular domain. Starting from least squares (LS) observations, we train a diffusion denoiser whose conditioning information is encoded by a long short-term memory (LSTM) network over a short observation sequence, enabling the model to exploit temporal dynamics beyond per-snapshot estimation. To robustly balance observation fidelity and learned generative priors across a wide signal-to-noise ratio (SNR) range, we introduce a learnable SNR-gated late-fusion shortcut that injects the network input into the final decoding stage through a sigmoid gate with trainable center and scale. To reduce inference latency, we adopt deterministic denoising diffusion implicit model (DDIM) style reverse updates with SNR-adaptive truncation and step allocation, which significantly reduces the number of reverse diffusion steps at high SNR while maintaining strong performance in low SNR regimes. Simulations on time-evolving standardized channel models demonstrate that the proposed method achieves consistent performance gains over existing diffusion-based channel estimation baselines, while retaining low latency through SNR-adaptive inference.
Chinese Translation
准确且低时延的信道估计对现代MIMO系统至关重要,尤其是在移动性场景下,此时信道表现出结构化稀疏性和强时间相关性。本文提出了一种用于信道估计的时间序列条件扩散框架,该框架在角度域中进行去噪。从最小二乘(LS)观测出发,我们训练一个扩散去噪器,其条件信息由长短期记忆(LSTM)网络在短观测序列上编码,使模型能够利用超出逐快照估计的时间动态。为了在宽信噪比(SNR)范围内稳健地平衡观测保真度与学习到的生成先验,我们引入了一种可学习的SNR门控后期融合捷径,它通过具有可训练中心和尺度的sigmoid门将网络输入注入最终解码阶段。为了降低推理时延,我们采用确定性的去噪扩散隐式模型(DDIM)风格的反向更新,并结合SNR自适应截断与步数分配,这在高SNR下显著减少了反向扩散步数,同时在低SNR情形下保持强劲性能。在随时间演化的标准化信道模型上的仿真表明,所提方法相较现有基于扩散的信道估计基线取得了一致的性能提升,同时通过SNR自适应推理保持了低时延。
cs.LG / 69 / 2610.09003
Algorithmic Scratchpads and Curriculum Staging for Arithmetic Reasoning in Tiny Transformers
Tiny Transformer 中算术推理的算法演算草稿与课程阶段化
large language model
大语言模型相关
Abstract
Autoregressive Large Language Models (LLMs) frequently struggle with deterministic multi-step algorithmic tasks such as multi-digit multiplication and long division. In this paper, we investigate the mechanics of multi-step arithmetic in compact "Tiny" Transformers (~10.6M non-embedding parameters, 49.3M total) trained on synthetic data across four basic operations (+, -, *, /) unrolled as step-by-step scratchpads. First, we establish the necessary training foundations: (1) dataloader sequence padding creates an 83% gradient starvation artifact that collapses accuracy from 40% to 1%, remediated via continuous sequence packing; (2) linguistic pretraining is an essential prerequisite (<= 2.0% without it); and (3) modern architectural primitives (RoPE, RMSNorm, SwiGLU) and Sparse Mixture of Experts (MoE) substantially improve additive reasoning over baseline GPT-2. Second, we demonstrate that algorithmic scratchpad formulation directly dictates success. Introducing a deterministic Digit-by-Digit Long Division scratchpad within a 4-stage Hierarchical Developmental Curriculum dramatically elevates single-digit division from 4.0% to 86.7% accuracy on a 4,000-problem held-out benchmark. In contrast, multi-digit multiplication remained challenging: detailed error analysis revealed that while the model correctly computed single-digit sub-products and place-value zeros, our FOIL scratchpad failed because it forced a simultaneous summation of up to nine multi-digit terms in a single step without pairwise intermediate accumulation. Finally, we identify two key boundaries: performance collapses to 0.00% on unseen 4-digit operands, and unbuffered training induces catastrophic forgetting, collapsing division accuracy from 86.7% down to 0.00%.
Chinese Translation
自回归大型语言模型(LLM)常常难以处理确定性的多步算法任务,例如多位数乘法和长除法。在本文中,我们研究紧凑的“Tiny” Transformer(约 10.6M 非嵌入参数,总计 49.3M)中多步算术的机制,这些模型在合成数据上训练,涵盖四种基本运算(+、-、*、/),并被展开为逐步演算草稿。首先,我们建立必要的训练基础:(1)数据加载器序列填充会产生 83% 的梯度饥饿伪影,使准确率从 40% 崩溃至 1%,可通过连续序列打包修复;(2)语言预训练是必不可少的先决条件(没有它时 <= 2.0%);以及(3)现代架构基元(RoPE、RMSNorm、SwiGLU)和稀疏专家混合(MoE)相较于基线 GPT-2 显著改善了加法推理。其次,我们证明算法演算草稿的表述直接决定成功。在一个 4 阶段分层发展课程中引入确定性的逐位长除法演算草稿,将留出的 4,000 题基准上的一位数除法准确率从 4.0% 显著提升到 86.7%。相比之下,多位数乘法仍然具有挑战性:详细的错误分析揭示,尽管模型正确计算了一位数子乘积和位值零,但我们的 FOIL 演算草稿失败了,因为它迫使在单个步骤中同时求和多达九个多位数项,而没有成对的中间累积。最后,我们识别出两个关键边界:在未见过的 4 位数操作数上性能骤降至 0.00%,并且无缓冲训练会诱发灾难性遗忘,使除法准确率从 86.7% 骤降至 0.00%。
cs.LG / 70 / 2610.09052
BeatFlow-ECG: Rectified Flow for ECG Reconstruction from Indirect Wearable Signals
BeatFlow-ECG:面向间接可穿戴信号 ECG 重建的整流流
diffusion
扩散模型相关
Abstract
Continuous cardiac monitoring outside clinical settings requires signals that are both informative and practical to collect during daily life. Electrocardiography (ECG) provides rich information about cardiac rhythm and waveform morphology, while wearable photoplethysmography (PPG) is easier to acquire continuously but is only an indirect cardiovascular measurement and is highly sensitive to motion. We present BeatFlow-ECG, a conditional rectified-flow model for reconstructing single-channel ECG from synchronized PPG and inertial measurements. BeatFlow-ECG models reconstruction as conditional transport from noise to ECG using a convolutional encoder-decoder with a transformer bottleneck and explicit flow-time conditioning. Motion information is incorporated through IMU-derived conditioning features, motion-dependent loss weighting, and an easy-to-hard training curriculum. We evaluate the model under leave-one-subject-out protocols on PPG-DaLiA and WESAD. BeatFlow-ECG achieves the best results among the evaluated deterministic, adversarial, and diffusion-based baselines across all reported waveform and beat-timing metrics, with Pearson correlations of 0.983 and 0.986 and R-peak F1 scores of 0.946 and 0.955, respectively. Compared with Conditional DDPM-1D, L1 error decreases from 0.085 to 0.062 on PPG-DaLiA and from 0.074 to 0.055 on WESAD. Additional analyses on PPG-DaLiA show higher correlation in fixed R-peak-relative waveform regions and lower reconstruction error across low-, medium-, and high-motion subsets.
Chinese Translation
在临床环境之外进行连续心脏监测,需要既包含丰富信息、又便于在日常生活中采集的信号。心电图(ECG)提供了关于心律和波形形态的丰富信息,而可穿戴光电容积描记术(PPG)虽然更容易连续采集,但它只是一种间接的心血管测量,且对运动高度敏感。我们提出了 BeatFlow-ECG,一种条件整流流模型,用于从同步的 PPG 和惯性测量中重建单通道 ECG。BeatFlow-ECG 将重建建模为从噪声到 ECG 的条件传输,采用带有 transformer 瓶颈和显式流时间条件的卷积编码器-解码器。运动信息通过由 IMU 导出的条件特征、依赖运动的损失加权以及由易到难的训练课程被纳入。我们在 PPG-DaLiA 和 WESAD 上采用留一受试者交叉验证协议对该模型进行评估。在所有报告的波形和心搏时序指标上,BeatFlow-ECG 在所评估的确定性、对抗式以及基于扩散的基线方法中取得了最佳结果,其皮尔逊相关系数分别为 0.983 和 0.986,R 峰 F1 分数分别为 0.946 和 0.955。与 Conditional DDPM-1D 相比,L1 误差在 PPG-DaLiA 上从 0.085 降至 0.062,在 WESAD 上从 0.074 降至 0.055。在 PPG-DaLiA 上的额外分析表明,在固定的 R 峰相对波形区域内相关性更高,并且在低、中、高运动子集上重建误差更低。
cs.LG / 71 / 2610.09063
Multi-Label Topic Assignment via LLM Distillation: A Comparative Analysis of Generative vs. Discriminative Student Models
通过 LLM 蒸馏进行多标签主题分配:生成式与判别式学生模型的对比分析
large language model
大语言模型相关
Abstract
Multi-label topic assignment for user-generated content (UGC) -- including product reviews and buyer-seller conversations -- poses unique scalability challenges in large-scale e-commerce due to informal language, extreme label sparsity, and rapidly evolving taxonomies. While utilizing Large Language Models (LLMs) as labeling oracles to distill ground-truth data has emerged as an industry standard to bypass prohibitive manual annotation costs, determining the optimal, low-latency architecture for the resulting student models remains an open challenge. To address this, we conduct a comprehensive evaluation across Small Language Model (SLM) parameter scales (1B, 4B, and 8B) and architectural paradigms (causal generative versus bidirectional discriminative). Comparing generative text-to-label classifiers against discriminative baselines (DeBERTa-V3 and ModernBERT), our analysis reveals a crucial data-dependent trade-off: while discriminative models outperform ultra-lightweight generative models on structured product reviews, even the smallest 1B generative model surpasses discriminative baselines on complex, multi-turn conversational data. Furthermore, generative models maintain robust performance under massive label-set expansion (up to 112 topics) and severe long-tail distributions, whereas discriminative baselines suffer a 35% drop in Macro-F1 at scale. Finally, we detail the successful production deployment of these optimized models across both product review and conversational domains, demonstrating strict latency compliance and tangible business impact at a global marketplace scale.
Chinese Translation
针对用户生成内容(UGC)——包括产品评论和买家-卖家对话——的多标签主题分配,由于非正式语言、极端的标签稀疏性以及快速演变的分类体系,在大规模电子商务中带来了独特的可扩展性挑战。虽然利用大型语言模型(LLM)作为标注预言机来蒸馏真值数据已成为一种行业标准,以绕过高昂的人工标注成本,但为由此产生的学生模型确定最优的低延迟架构仍然是一个悬而未决的挑战。为解决这一问题,我们在小语言模型(SLM)参数规模(1B、4B 和 8B)以及架构范式(因果生成式与双向判别式)上开展了一项全面评估。将生成式文本到标签分类器与判别式基线(DeBERTa-V3 和 ModernBERT)进行比较,我们的分析揭示了一个关键的、依赖数据的权衡:尽管判别式模型在结构化产品评论上优于超轻量级生成式模型,但即便最小的 1B 生成式模型在复杂的多轮对话数据上也超越了判别式基线。此外,生成式模型在大规模标签集扩展(最多 112 个主题)和严重长尾分布下仍保持稳健的性能,而判别式基线在大规模下 Macro-F1 下降了 35%。最后,我们详细介绍了这些优化模型在产品评论和对话两个领域中的成功生产部署,展示了在全球市场尺度下严格的延迟合规性和切实的业务影响。
cs.LG / 72 / 2610.09092
MaRK: Markov-adapted Recurrent Kernels for Dynamic Operator Conditioning in State Space Models
MaRK:用于状态空间模型中动态算子条件化的马尔可夫自适应循环核
diffusion
扩散模型相关
Abstract
State Space Models (SSMs) offer an efficient alternative to Transformers for sequence modeling, yet conditioning pre-trained SSMs for iterative generation typically operates outside the recurrent operator, through input injection or activation modulation. While such mechanisms expose the model to conditioning information, they leave the underlying temporal dynamics fixed. We introduce MaRK (Markov-adapted Recurrent Kernels), a dynamic operator-conditioning framework that maps context vectors directly into bounded modulations of a frozen SSM's recurrence ($A$), read-in ($B$), read-out ($C$), skip ($D$), and discretization ($Δ$) parameters. Viewed through the lens of LPV-SSM systems, MaRK induces a context-indexed family of Markov parameter sequences, allowing each diffusion timestep to reshape the model's input-output memory kernel. We instantiate MaRK on a frozen 111M-parameter Hydra SSM backbone and study three adapter geometries: Hypernet, Chebyshev polynomial, and Discrete Cosine Transform kernels. Since these adapters modify the Markov parameter sequence through low-rank auxiliary maps on the frozen backbone, parameter-efficient fine-tuning arises as a structural consequence of the adaptation mechanism itself, requiring only 6.3--11M trainable auxiliary parameters to transition from a bidirectional objective to an iterative diffusion regime. The bounded recurrence parameterization further yields an analytic Affine Quadratic Stability certificate for the modulated recurrence. Through synthetic LPV recovery experiments and Markov-operator diagnostics, we show that MaRK recovers coordinate-invariant temporal operators under matched assumptions and produces distinct, stable timestep-conditioned memory profiles. Empirically, the Chebyshev variant yields the strongest performance, achieving an average validation loss of 2.55, followed by the DCT (2.59) and Hypernet (3.77) geometries.
Chinese Translation
状态空间模型(SSMs)为序列建模提供了一种高效的 Transformer 替代方案,然而,为迭代生成而对预训练 SSM 进行条件化通常发生在循环算子之外,通过输入注入或激活调制来实现。尽管此类机制使模型接触到条件信息,但它们使底层时间动态保持不变。我们提出 MaRK(马尔可夫自适应循环核),这是一种动态算子条件化框架,它将上下文向量直接映射为对冻结 SSM 的递推($A$)、读入($B$)、读出($C$)、跳跃($D$)以及离散化($Δ$)参数的有界调制。从 LPV-SSM 系统的视角来看,MaRK 诱导出一个由上下文索引的马尔可夫参数序列族,使每个扩散时间步都能重塑模型的输入-输出记忆核。我们在一个冻结的 111M 参数的 Hydra SSM 主干上实例化 MaRK,并研究三种适配器几何结构:Hypernet、Chebyshev 多项式和离散余弦变换核。由于这些适配器通过冻结主干上的低秩辅助映射来修改马尔可夫参数序列,参数高效微调作为自适应机制本身的结构性结果而出现,仅需 6.3--11M 个可训练辅助参数即可从双向目标过渡到迭代扩散机制。有界递推参数化进一步为受调制的递推产生了解析的仿射二次稳定性证书。通过合成 LPV 恢复实验和马尔可夫算子诊断,我们表明 MaRK 在匹配假设下恢复了坐标不变的时间算子,并产生不同的、稳定的、以时间步为条件的记忆轮廓。在实证上,Chebyshev 变体取得最强性能,达到 2.55 的平均验证损失,其次是 DCT(2.59)和 Hypernet(3.77)几何结构。
cs.LG / 73 / 2610.09122
Are Parameter-Efficient Fine-tuning Methods Really Different?
参数高效微调方法真的不同吗?
diffusion
扩散模型相关
Abstract
Parameter-efficient fine-tuning (PEFT) offers many parameterizations, yet their methodological and functional differences remain unclear. We compare six methods in language and diffusion models to examine how their parameterizations relate to task performance, forgetting, and changes in pretrained weight geometry. Motivated by the spectrum-preserving design of orthogonal fine-tuning (OFT), we first ask whether spectral preservation is itself important for adaptation and retention. We find that the selected LoRA-family methods also approximately preserve pretrained geometry, and that restoring their slightly drifted singular-value spectra largely preserves task performance, questioning the necessity of explicit geometric preservation. Beyond this, we observe that some methods exhibit distinct adaptation--retention trade-offs that vary across settings: LoRA most consistently limits forgetting at competitive performance, DoRA achieves higher mean task scores than LoRA in most comparisons, while PiSSA often incurs greater retention costs. Further intervention experiments suggest that while performance gains from different PEFT methods can be attributed to modifications in different groups of spectral components, we consistently find that restoring dominant rather than intermediate or trailing components produces the largest mean reduction in general-text NLL or base-image drift. Together, these results motivate evaluating geometric constraints through their functional consequences rather than preservation alone. Code is available at https://github.com/Kuaaannn/PEFT_methods.
Chinese Translation
参数高效微调(PEFT)提供了许多参数化方式,但它们在方法论和功能上的差异仍不清楚。我们在语言模型和扩散模型中比较六种方法,以考察它们的参数化如何与任务性能、遗忘以及预训练权重几何的变化相关联。受正交微调(OFT)的频谱保持设计启发,我们首先追问频谱保持本身是否对适应和保持重要。我们发现,所选的 LoRA 系列方法也近似地保持了预训练几何,并且恢复它们轻微漂移的奇异值谱在很大程度上保持了任务性能,这让人质疑显式几何保持的必要性。除此之外,我们观察到一些方法表现出在不同设置下有所不同的独特适应—保持权衡:LoRA 在具有竞争力的性能下最一致地限制遗忘,DoRA 在大多数比较中取得比 LoRA 更高的平均任务分数,而 PiSSA 通常带来更大的保持成本。进一步的干预实验表明,虽然不同 PEFT 方法带来的性能提升可以归因于对不同谱成分组的修改,但我们一致发现,恢复主导成分而非中间或尾部成分,能在一般文本 NLL 或基础图像漂移上产生最大的平均降低。总之,这些结果促使我们通过几何约束的功能后果而非仅通过保持性来评估它们。代码可在 https://github.com/Kuaaannn/PEFT_methods 获取。
cs.LG / 74 / 2610.09145
Noise Your Prompt: Noising Conditioning Tokens in Continuous Diffusion Language Models
给你的提示加噪:在连续扩散语言模型中为条件提示词元加噪
diffusion
扩散模型相关
Abstract
We revisit a standard accepted practice in the continuous diffusion language model literature of fixing conditioning prompt tokens clean during training. We make a very simple modification: also noise the conditioning prompt tokens during training. We demonstrate that under this modified training objective, we achieve better generalization in combinatorial reasoning tasks such as Sudoku and N-Queens, with the largest gains on harder variants ($3.73\% \to 24.65\%$ solve rate on Sudoku Hard), and increased diversity of generated solutions ($50.60\% \to 73.79\%$ coverage on 10x10 N-Queens). We also show measurable improvements to natural language generation quality in modest dataset regimes with Gigaword summarization, but notably demonstrate that gains do not transfer to all natural language tasks (e.g open ended dialogue generation). Our method is a single line change to the training objective, requires no additional inference costs by default, and provides the flexibility of classifier-free guidance inspired guided sampling. Our \href{https://github.com/LateralIntelligence/noise-your-prompt} {code} is publicly available.
Chinese Translation
我们重新审视连续扩散语言模型文献中一个被普遍接受的标准做法:在训练期间将条件提示词元保持为干净。我们做了一个非常简单的修改:在训练期间也对条件提示词元加噪。我们证明,在这一修改后的训练目标下,我们在诸如数独和 N 皇后等组合推理任务中实现了更好的泛化,在更难变体上提升最大(Sudoku Hard 上求解率 $3.73\% \to 24.65\%$),并提高了生成解的多样性(10x10 N-Queens 上覆盖率 $50.60\% \to 73.79\%$)。我们还表明,在中等数据集规模下,使用 Gigaword 摘要时,自然语言生成质量有可测量的提升,但值得注意的是,这些增益并不会迁移到所有自然语言任务(例如开放式对话生成)。我们的方法仅需对训练目标做一行更改,默认情况下不需要额外推理成本,并提供受无分类器引导启发的引导采样的灵活性。我们的 \href{https://github.com/LateralIntelligence/noise-your-prompt}{代码} 已公开可用。
cs.LG / 75 / 2610.09186
The Dichotomy Between Pattern Recognition and Step-by-Step Reasoning
模式识别与逐步推理之间的二分法
large language model
大语言模型相关
Abstract
We argue that pattern recognition and step-by-step reasoning are two ends of a spectrum. A large language model (LLM) learns to reason step-by-step when data is structured such that the next token depends on a small amount of preceding context. Inference in LLMs resembles pattern recognition when the next token depends on a large amount of preceding context. If the next token depends on only the $c$ most recent tokens, reasoning traces are paths on a De Bruijn graph whose nodes are $c$-length contexts and edges are next-token transitions between contexts. The set of reasoning traces of a task forms a directed acyclic subgraph of the De Bruijn graph. An LLM that has learned all edges of this subgraph can compose them to solve longer, unseen tasks, i.e., it reasons step-by-step. We prove that the number of edges is vanishingly small compared to the number of reasoning traces. Empirically, the number of training samples a transformer needs is a power law in the number of edges, so learning to reason step-by-step is sample efficient. We can induce De Bruijn structure in any task by maintaining a ``state'' that makes future reasoning independent of the past. The frequency of states in the reasoning trace determines $c$. We show, by fine-tuning Qwen2.5-1.5B-Instruct to solve equations and answer questions about stories, that frequent states (small $c$) result in higher accuracy but greater fragility to perturbations at test time. LLMs trained with a large $c$ are only as good as models that perform pattern recognition without reasoning. A moderate density of states balances accuracy and robustness. We show that real-world data has De Bruijn structure: Qwen3-14B and Qwen3-32B retain over 75% of their accuracy on GSM8K, MATH-500 and GPQA-Diamond when attention is restricted to a sliding window less than 15% as long as the full reasoning trace.
Chinese Translation
我们认为,模式识别与逐步推理是一个连续谱的两端。当数据的结构使得下一个 token 依赖于少量前文上下文时,大型语言模型(LLM)学会逐步推理。当下一个 token 依赖于大量前文上下文时,LLM 中的推理类似于模式识别。如果下一个 token 仅依赖于最近的 $c$ 个 token,那么推理轨迹是 De Bruijn 图上的路径,该图的节点是长度为 $c$ 的上下文,边是上下文之间的下一个 token 转移。一个任务的推理轨迹集合构成 De Bruijn 图的一个有向无环子图。一个已经学会该子图所有边的 LLM 可以组合它们来求解更长、未见过的任务,也就是说,它进行逐步推理。我们证明,与推理轨迹的数量相比,边的数量小到可以忽略不计。经验上,一个 transformer 所需的训练样本数量是边数量的幂律,因此学习逐步推理具有样本效率。我们可以通过维护一个使未来推理独立于过去的“状态”,来在任何任务中诱导出 De Bruijn 结构。推理轨迹中状态的频率决定了 $c$。我们通过微调 Qwen2.5-1.5B-Instruct 来求解方程并回答关于故事的问题,表明频繁出现的状态(较小的 $c$)会带来更高的准确率,但在测试时对扰动更加脆弱。用较大的 $c$ 训练的 LLM 仅与不进行推理而执行模式识别的模型一样好。适中的状态密度会平衡准确率与鲁棒性。我们表明,真实世界数据具有 De Bruijn 结构:当注意力被限制在一个长度不到完整推理轨迹 15% 的滑动窗口时,Qwen3-14B 和 Qwen3-32B 在 GSM8K、MATH-500 和 GPQA-Diamond 上仍保留超过 75% 的准确率。
cs.LG / 76 / 2610.09212
CurveTQ: Rotation-Free Trellis Quantization of LLM Weights via Curvature-Weighted Search
CurveTQ:通过曲率加权搜索实现的无旋转 LLM 权重网格量化
large language model
大语言模型相关
Abstract
The best two-bit weight quantizers for large language models, such as QTIP and Proteus, rotate each weight matrix by a random orthogonal transform, which must be undone at every decoding step, then encode it with a trellis or lattice code under a Euclidean search; the layer Hessian enters only through error feedback between coding blocks. We show that this leaves part of the Hessian unused. Error feedback turns the loss into a weighted sum of per-coordinate rounding errors whose weights, the diagonal of the Hessian's LDL factorization, existing quantizers compute but never read. We put these weights into the Viterbi branch metric, so the search follows the curvature within each coding block. This also explains the rotation: it removes this within-block variation, so weighting in the native basis and rotating are substitutes. On three models the weighted native search matches a full-dimension randomized Hadamard to within about one point of downstream accuracy, and weighting after the rotation gains little. Around this search we build CurveTQ, a trellis codec with no rotation, which handles the weights' amplitude and marginal shape with a factored scale field and a closed-form quantile table, and stores a start state per coding block so the trellis can adapt to the residual that error feedback carries into it. At two bits CurveTQ is 1-3 points higher in mean downstream accuracy than QTIP and Proteus on three 4-8B Instruct models, even after both are given our start state, which alone lifts either baseline by 1-3 points. It also leads on a 35B mixture of experts, to our knowledge the first trellis-coded result on such a model. With no rotation to undo, our decoder is the fastest of the three at all tested batch sizes and bit widths.
Chinese Translation
面向大语言模型的最佳二比特权重量化器,如 QTIP 和 Proteus,会用随机正交变换对每个权重矩阵进行旋转,而该变换必须在每个解码步骤中被撤销,然后在欧几里得搜索下用网格(trellis)或格(lattice)码对其进行编码;层 Hessian 仅通过编码块之间的误差反馈进入。我们表明,这会使 Hessian 的一部分未被利用。误差反馈将损失转化为逐坐标舍入误差的加权和,其权重——即 Hessian 的 LDL 分解的对角线——是现有量化器会计算却从不读取的。我们将这些权重放入 Viterbi 分支度量中,从而使搜索在每个编码块内部遵循曲率。这也解释了旋转:它消除了这种块内变化,因此在原始基下加权与旋转是相互替代的。在三个模型上,加权后的原始搜索与全维度随机 Hadamard 在下游准确率上相差约一个点以内,而在旋转之后再加权收益甚微。围绕这一搜索,我们构建了 CurveTQ,一种无旋转的网格编解码器,它用一个可分解的尺度场和一个闭式分位数表来处理权重的幅度与边缘分布形状,并为每个编码块存储一个起始状态,使网格能够适应误差反馈携带进入其中的残差。在二比特下,CurveTQ 在三个 4-8B Instruct 模型上的平均下游准确率比 QTIP 和 Proteus 高 1-3 个点,即便在两者都被赋予我们的起始状态之后也是如此,而仅此一项就能将任一基线提升 1-3 个点。它还在一个 35B 混合专家模型上领先,据我们所知,这是此类模型上的首个网格编码结果。由于无需撤销任何旋转,在所有测试的批大小和比特宽度下,我们的解码器都是三者中最快的。
cs.LG / 77 / 2610.09221
Consistent Distribution Matching for Data-Free Diffusion Distillation
面向无数据扩散蒸馏的一致性分布匹配
diffusion
扩散模型相关
Abstract
Flow and diffusion models suffer from slow inference due to computationally expensive numerical integration. Distillation provides a promising way for a student model to learn from a teacher's dynamics, enabling one-step or few-step generation. However, existing methods often depend on curated distillation datasets, costly teacher rollouts, or auxiliary proxy networks, which complicate model training and scaling. In this work, we propose Consistent Distribution Matching, a simulation-free and data-free distillation method for accelerating diffusion and flow models while preserving strong generative capacity. Our key insight is to unify sample generation and score estimation with one student network. Thus, our framework uses only two models, a frozen teacher and a trainable student, and optimizes one objective. We prove that minimizing our objective indicates Wasserstein convergence of the student flow-map pushforwards to the teacher marginals. On ImageNet 256$\times$256, our method attains an FID of 2.04 with a single function evaluation (1-NFE) and a 4-NFE FID of 1.37 within 40 epochs of training, surpassing the state-of-the-art distillation baselines without data. Our code code and model are available at https://consistentdmd.github.io/.
Chinese Translation
流模型和扩散模型由于计算代价高昂的数值积分而面临推理速度慢的问题。蒸馏为学生模型从教师模型的动力学中学习提供了一种有前景的方式,能够实现一步或少步生成。然而,现有方法通常依赖于精心整理的蒸馏数据集、代价高昂的教师模型展开,或辅助代理网络,这使得模型训练和扩展变得复杂。在这项工作中,我们提出一致性分布匹配,一种无需模拟且无需数据的蒸馏方法,用于加速扩散模型和流模型,同时保持强大的生成能力。我们的关键洞见是用一个学生网络统一样本生成和分数估计。因此,我们的框架仅使用两个模型,一个冻结的教师模型和一个可训练的学生模型,并优化一个目标。我们证明,最小化我们的目标表明学生的流映射前推分布向教师边缘分布的 Wasserstein 收敛。在 ImageNet 256$\times$256 上,在 40 个训练轮次内,我们的方法在单次函数评估(1-NFE)下达到 2.04 的 FID,并达到 1.37 的 4-NFE FID,超越了无数据的最先进蒸馏基线。我们的代码和模型可在 https://consistentdmd.github.io/ 获取。
cs.LG / 78 / 2610.09285
LeCuration: A Tiny World Model as a Data Curation Multi-Tool
LeCuration:作为数据整理多工具的微型世界模型
diffusion
扩散模型相关
Abstract
Many applications of physical AI run within finite or closed physical worlds with a limited set of physical laws governing object behavior. Examples include robots working in a warehouse and agents moving around in a video game. In order to better organize, filter, and curate data for physical AI applications, we propose a new approach centered on the unique settings and physical laws of individual datasets. We train LeCuration, a small world model intended to serve as a data curation tool for a separate, larger downstream model. To build this model, we choose LeWorldModel (LeWM)as our latent encoder and predictor, adding a diffusion transformer (DiT) decoder to add visuals to autoregressive gameplay rollout. We find that the embeddings of this model can be used as an anomaly detection signal and as a content-based clustering heuristic, and that auto-regressively predicting the game state with this model allows us to qualitatively check for action-state consistency. This paper presents a qualitative, proof-of-concept case study on CS:GO gameplay data; we do not yet report quantitative curation metrics or downstream training results, which we identify as the key next step.
Chinese Translation
物理 AI 的许多应用运行在有限或封闭的物理世界中,这些世界由一整套有限的、支配物体行为的物理定律所约束。例子包括在仓库中工作的机器人以及在电子游戏中移动的智能体。为了更好地为物理 AI 应用组织、过滤和整理数据,我们提出了一种以各个数据集的独特设定和物理定律为核心的新方法。我们训练了 LeCuration,这是一个小型世界模型,旨在为另一个更大的下游模型充当数据整理工具。为了构建该模型,我们选择 LeWorldModel(LeWM)作为我们的潜变量编码器和预测器,并添加一个扩散 transformer(DiT)解码器,以便为自回归的游戏过程推演添加视觉画面。我们发现,该模型的嵌入可以被用作异常检测信号以及基于内容的聚类启发式方法,并且用该模型自回归地预测游戏状态,使我们能够定性地检查动作-状态一致性。本文呈现了一项针对 CS:GO 游戏数据的定性概念验证案例研究;我们尚未报告定量的数据整理指标或下游训练结果,我们将其确定为关键的下一步。
cs.LG / 79 / 2610.09305
Kuration SDK: Addressing the Virtual2Real Gap via Data Curation
Kuration SDK:通过数据策展解决 Virtual2Real 差距
diffusion
扩散模型相关
Abstract
Benchmarks for measuring the quality of action-conditioned world models are still evolving and shifting away from visual similarity-based metrics to action-semantic and physically-grounded metrics. However, for domain and task-agnostic action-conditioned world model training, existing benchmarks provide a limited signal. By training and evaluating diffusion world models on CounterStrike gameplay data, we confirm that qualitative playability does not correspond with metrics such as FVD, LPIPS, and JEDi. We term this the Virtual2Real gap. We posit that, in lieu of reliable benchmarks, curating raw gameplay data and measuring a variety of diagnostic properties provides a more robust signal to bridge the gap, before the training even begins. We present several curation strategies and a general-purpose kit for physical AI data curation called Kuration SDK, which is being open-sourced with this paper. The SDK was instrumental in uncovering the root cause of the virtual2real gap in a specific case: why two world models trained on identical gameplay map, action and state distribution, behaved very differently when played in spite of having very similar LPIPS and FVD scores. Thus, Kuration SDK has the potential to uncover the root causes of Virtual2Real gap in specific datasets and accelerate development of sample-efficient training datasets.
Chinese Translation
用于衡量以动作为条件的世界模型质量的基准仍在不断演进,并且正从基于视觉相似度的指标转向动作语义与具有物理基础的指标。然而,对于领域与任务无关的以动作为条件的世界模型训练而言,现有基准所提供的信号是有限的。通过在 CounterStrike 游戏数据上训练并评估扩散世界模型,我们确认定性的可玩性与 FVD、LPIPS 和 JEDi 等指标并不对应。我们将此称为 Virtual2Real 差距。我们认为,在缺乏可靠基准的情况下,对原始游戏数据进行策展并测量各种诊断性属性,能够提供更稳健的信号来弥合这一差距,而且这一切甚至在训练开始之前就能实现。我们提出了若干策展策略,以及一个用于物理 AI 数据策展的通用工具包,称为 Kuration SDK,它随本文一同开源。该 SDK 在揭示某一具体案例中 virtual2real 差距的根本原因方面发挥了关键作用:即为什么两个在相同游戏地图、动作与状态分布上训练的世界模型,尽管具有非常相似的 LPIPS 与 FVD 分数,在实际游玩时却表现得非常不同。因此,Kuration SDK 有潜力揭示特定数据集中 Virtual2Real 差距的根本原因,并加速样本高效训练数据集的开发。
cs.LG / 80 / 2610.09311
Denoising Blocks, Not Tokens: Efficient Compressed Continuous Diffusion with Branching Token Realization
去噪块,而非词元:带有分支词元实现的高效压缩连续扩散
diffusion
扩散模型相关
Abstract
Diffusion language models (DLMs) generate text through iterative parallel refinement, offering the potential for higher throughput than autoregressive (AR) decoding. However, most DLMs still maintain one generative state per token, so every denoising step processes a state sequence as long as the output sequence, limiting the throughput gains from parallel generation. Continuous DLMs provide an additional degree of freedom: a single continuous state can represent multiple tokens, allowing diffusion to operate on a much shorter latent sequence. We introduce \emph{Branching Latent Diffusion (BLD)}, which exploits this flexibility by compressing a 1024-token sequence into only 64 block latents, a $16\times$ reduction. BLD combines latent compression with \emph{branching token realization}, where each latent is decoded by a local AR branch and all branches run in parallel. Because strong compression makes joint latent generation difficult, BLD generates the latents in groups, conditioning each group on previously generated latents. In end-to-end evaluation on the same GPU, BLD reduces generation FLOPs by more than $80\times$ and increases throughput by more than $6\times$ relative to the similarly sized ELF-L baseline. Compared with the AR baseline, BLD achieves more than $6\times$ higher throughput and more than $4\times$ lower latency. Despite the compression, BLD maintains competitive local fluency and diversity, although long-range coherence remains challenging. Overall, BLD shows that moving diffusion from token-level states to compressed latent sequences can substantially improve the efficiency of long-sequence generation.
Chinese Translation
扩散语言模型(DLMs)通过迭代并行细化生成文本,具有比自回归(AR)解码更高吞吐量的潜力。然而,大多数 DLM 仍为每个词元维护一个生成状态,因此每个去噪步骤处理的状态序列与输出序列一样长,限制了并行生成带来的吞吐量增益。连续 DLM 提供了额外的自由度:单个连续状态可以表示多个词元,从而使扩散能够在短得多的潜在序列上运行。我们引入了分支潜在扩散(Branching Latent Diffusion,BLD),它利用这种灵活性,将一个 1024 词元的序列压缩为仅 64 个块潜在变量,实现了 $16\times$ 的缩减。BLD 将潜在压缩与分支词元实现相结合,其中每个潜在变量由一个局部 AR 分支解码,并且所有分支并行运行。由于强压缩使得联合潜在生成变得困难,BLD 以组为单位生成潜在变量,并让每个组以先前生成的潜在变量为条件。在同一 GPU 上的端到端评估中,与规模相近的 ELF-L 基线相比,BLD 将生成 FLOPs 减少了超过 $80\times$,并将吞吐量提高了超过 $6\times$。与 AR 基线相比,BLD 实现了超过 $6\times$ 的更高吞吐量和超过 $4\times$ 的更低延迟。尽管进行了压缩,BLD 仍保持了具有竞争力的局部流畅性和多样性,尽管长程连贯性仍然具有挑战性。总体而言,BLD 表明,将扩散从词元级状态转移到压缩的潜在序列可以显著提高长序列生成的效率。
cs.LG / 81 / 2610.09342
Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression
面向无数据混合专家压缩的共享低秩基分解
large language model
大语言模型相关
Abstract
Mixture-of-Experts (MoE) large language models decouple capacity from compute through sparse routing, but their large parameter count creates storage and serving challenges. We analyze three MoE compression families: expert pruning, expert merging, and weight reconstruction, and derive structural error bounds showing that pruning and merging can incur non-vanishing errors tied to routing and expert heterogeneity. In contrast, weight reconstruction avoids these structural costs by preserving expert structure and routing. Motivated by the analysis, we propose Shared Low-rank Basis Factorization (SLBF), a data-free weight reconstruction method that uses rank-$k$ bases shared among experts, enabling richer cross-expert sharing, faster convergence, and lower reconstruction error. A post-hoc gauge fixing removes redundant parameters at no representational cost. Across five MoE architectures spanning 16B to 122B parameters, SLBF consistently outperforms methods from all three compression families.
Chinese Translation
混合专家(MoE)大语言模型通过稀疏路由将容量与计算解耦,但其庞大的参数量带来了存储与服务挑战。我们分析了三类 MoE 压缩方法族:专家剪枝、专家合并和权重重建,并推导了结构误差界,表明剪枝和合并会带来与路由和专家异质性相关的非消失误差。相比之下,权重重建通过保留专家结构和路由避免了这些结构性代价。受该分析启发,我们提出共享低秩基分解(SLBF),一种无数据权重重建方法,它使用在专家间共享的 rank-$k$ 基,从而实现更丰富的跨专家共享、更快的收敛以及更低的重建误差。事后规范固定在不造成表示代价的情况下移除了冗余参数。在涵盖 16B 到 122B 参数的五个 MoE 架构上,SLBF 始终优于来自所有三类压缩方法族的方法。
cs.LG / 82 / 2610.09407
Noise, Denoise, Correct: MCMC Posterior Sampling with Diffusion Priors in Three Steps
加噪、去噪、校正:三步实现带扩散先验的 MCMC 后验采样
diffusion
扩散模型相关
Abstract
Pretrained diffusion models are powerful priors for inverse problems, but posterior sampling under nonlinear, non-differentiable forward models remain hard. We introduce diffusion waltz, an MCMC method using SDEdit-style noising-denoising as a proposal, corrected via Metropolis-Hastings for exact posterior sampling without prior evaluation. We further propose injecting observations into the proposal while preserving exactness, using a gradient-free ensemble Kalman update. On a non-differentiable Navier-Stokes initial condition recovery task, diffusion waltz outperforms existing baselines across different noise and nonlinearity regimes.
Chinese Translation
预训练扩散模型是求解反问题的强大先验,但在非线性、不可微的前向模型下进行后验采样仍然十分困难。我们提出了 diffusion waltz,一种 MCMC 方法,它使用 SDEdit 风格的加噪-去噪作为提议分布,并通过 Metropolis-Hastings 进行校正,从而在无需评估先验的情况下实现精确的后验采样。我们进一步提出,在保持精确性的同时,利用无梯度的集合卡尔曼更新将观测注入到提议分布中。在一个不可微的 Navier-Stokes 初始条件恢复任务上,diffusion waltz 在不同噪声与非线性程度下均优于现有基线。
cs.LG / 83 / 2610.09563
EvoSignal: LLM-Guided Evolutionary Design of Modular Traffic Signal Control Programs
EvoSignal:LLM 引导的模块化交通信号控制程序的进化设计
large language model
大语言模型相关
Abstract
Effective traffic signal control (TSC) requires policies that respond to changing traffic demand and network conditions while meeting different control objectives. However, adapting existing strategies often involves repeated manual design and adjustment, making it difficult to systematically explore better control rules for a target network. Large language models (LLMs) can automate this process, but directly using them to select signal phases leaves decision rules embedded in black-box models and incurs recurring inference costs and latency. This paper formulates TSC as a modular program design problem and proposes EvoSignal, an LLM-guided evolutionary framework using traffic knowledge and performance feedback. The modular representation separates traffic feature extraction, local phase prioritization, and optional network-based priority adjustment. Starting from several established strategies, EvoSignal improves programs through feedback on congestion and signal operation, retaining strategies with different performance trade-offs. The resulting programs operate without online LLM inference. Simulation experiments across five scenarios on two real-world road networks show that the selected default EvoSignal program reduces waiting time by 16.8--49.2\% relative to the lowest waiting time achieved by the 20 conventional, reinforcement learning-based, and LLM-based baselines in each scenario. A program prioritizing travel time and queue length outperforms all 20 baselines on all three metrics in the search scenario and remains among the top three on each metric when transferred unchanged to the other four scenarios. These findings support automated design of inspectable control programs that transfer across the evaluated road networks and traffic demands.Code is available at https://github.com/georgewanglz2019/EvoSignal.
Chinese Translation
有效的交通信号控制(TSC)需要能够响应不断变化的交通需求和网络条件,同时满足不同控制目标的策略。然而,调整现有策略通常涉及反复的人工设计与调整,使得难以系统地探索针对目标网络的更优控制规则。大语言模型(LLM)可以自动化这一过程,但直接使用它们来选择信号相位会使决策规则嵌入黑盒模型中,并带来反复的推理成本和延迟。本文将 TSC 形式化为一个模块化程序设计问题,并提出 EvoSignal,一个使用交通知识和性能反馈的 LLM 引导的进化框架。模块化表示将交通特征提取、局部相位优先级排序以及可选的基于网络的优先级调整分离开来。从若干已有策略出发,EvoSignal 通过对拥堵和信号运行的反馈来改进程序,保留具有不同性能权衡的策略。所得程序无需在线 LLM 推理即可运行。在两个真实道路网络上的五个场景中进行的仿真实验表明,与每个场景中 20 个传统、基于强化学习和基于 LLM 的基线所达到的最低等待时间相比,所选的默认 EvoSignal 程序将等待时间减少了 16.8--49.2\%。一个优先考虑行程时间和排队长度的程序在搜索场景中在所有三个指标上优于全部 20 个基线,并且在原样迁移到其他四个场景时,在每个指标上仍保持在前三名。这些发现支持可检查控制程序的自动化设计,这些程序可在所评估的道路网络和交通需求之间迁移。代码可在 https://github.com/georgewanglz2019/EvoSignal 获取。
cs.LG / 84 / 2610.09587
Collaborative Reasoning Distillation via Cross-Feedback and Coherent Curation
通过交叉反馈与连贯策展的协作推理蒸馏
large language model
大语言模型相关
Abstract
Reasoning capabilities are critical for advancing Large Language Models, yet current approaches either require massive computational budgets or struggle to effectively distill reasoning to smaller models. Standard distillation methods rely on outcome-based rewards, failing to distinguish between sound reasoning and lucky guesses. We propose Collaborative Reasoning Distillation (CRD), a framework that enhances reasoning in compact models through three innovations: (1) interactive cross-feedback where teachers iteratively critique each other's reasoning, (2) fine-grained step-wise quality assessment capturing logical validity independent of final answers, and (3) coherence-aware step stitching that synthesizes complementary strengths. Students are trained via Reasoning Quality Optimization (RQO) with budget constraints. Our model, CRD-4B, achieves 97.3% on MATH-500 and 70.3% on AIME'25, surpassing baselines while using only 50K training examples, up to 12 times smaller than the datasets of comparable models.
Chinese Translation
推理能力对于推进大型语言模型至关重要,然而当前方法要么需要巨大的计算预算,要么难以有效地将推理蒸馏到更小的模型。标准蒸馏方法依赖基于结果的奖励,无法区分合理的推理与侥幸的猜测。我们提出协作推理蒸馏(CRD),一个通过三项创新增强紧凑模型推理能力的框架:(1)交互式交叉反馈,其中教师迭代地互相批评对方的推理;(2)细粒度的逐步质量评估,独立于最终答案捕捉逻辑有效性;(3)连贯性感知的步骤拼接,综合互补优势。学生模型通过带有预算约束的推理质量优化(RQO)进行训练。我们的模型 CRD-4B 在 MATH-500 上达到 97.3%,在 AIME'25 上达到 70.3%,超越了基线,同时仅使用 50K 训练样本,比可比模型的数据集最多小 12 倍。
cs.LG / 85 / 2610.09597
COPC: Coupled Off-Policy Correction for Asynchronous LLM Reinforcement Learning
COPC:用于异步 LLM 强化学习的耦合离策略校正
large language model
大语言模型相关
Abstract
Asynchronous RL accelerates large language model post-training by decoupling rollout generation from optimization, but trains on stale trajectories. Existing methods primarily correct token-level policy mismatch through importance-ratio control in the actor objective. We show that this \emph{policy-side correction} alone is insufficient: advantage estimates also inherit mismatch from behavior-policy continuations, which we term \emph{advantage staleness}. We derive exact bias and variance decompositions for a general two-channel actor update, revealing nonseparable coupling between policy-weight and advantage-estimation errors: their interaction induces multiplicative bias terms, while squared policy weights amplify advantage uncertainty in gradient variance. This motivates the hypothesis that policy- and advantage-side correction should be coordinated. We introduce Coupled Off-Policy Correction (COPC), an actor--critic method combining token-level ratio masking with two-sided clipped-ratio weighting of TD residuals for return and advantage estimation. Joint parameter sweeps across staleness levels support this hypothesis: the effect of one correction parameter depends on, and can reverse with, the other. COPC achieves the highest reported performance on tool-integrated mathematical reasoning and search, outperforming the strongest reported asynchronous baseline in each setting. It also offers a broad high-performing parameter region and improved training stability. In search, COPC remains stable throughout training, while most evaluated asynchronous baselines collapse late in training. These gains persist at 64-step policy staleness. COPC adds minimal step-time overhead over asynchronous PPO and retains a $1.7\times$ step-time speedup over synchronous PPO.
Chinese Translation
异步强化学习通过将 rollout 生成与优化解耦,加速了大语言模型的后训练,但会在陈旧轨迹上训练。现有方法主要通过在 actor 目标中进行重要性比率控制,来校正 token 级的策略不匹配。我们表明,仅靠这种 \emph{策略侧校正} 是不够的:优势估计也会从行为策略延续中继承不匹配,我们将其称为 \emph{优势陈旧性}。我们推导了一般双通道 actor 更新的精确偏差与方差分解,揭示了策略权重与优势估计误差之间不可分离的耦合:它们的交互会引入乘性偏差项,而策略权重的平方会在梯度方差中放大优势不确定性。这促使我们提出一个假设:策略侧和优势侧的校正应当协调进行。我们提出耦合离策略校正(Coupled Off-Policy Correction,COPC),这是一种 actor--critic 方法,它将 token 级比率掩码与用于回报和优势估计的 TD 残差的双侧裁剪比率加权相结合。跨不同陈旧程度的联合参数扫描支持了这一假设:一个校正参数的效果取决于另一个校正参数,并且可能随另一个参数发生反转。COPC 在工具集成的数学推理和搜索上取得了已报告的最高性能,在每种设置中都优于已报告的最强异步基线。它还提供了宽广的高性能参数区域和更好的训练稳定性。在搜索中,COPC 在整个训练过程中保持稳定,而大多数被评估的异步基线在训练后期崩溃。这些收益在 64 步策略陈旧度下依然存在。COPC 相对于异步 PPO 仅增加极小的单步时间开销,并相对于同步 PPO 保持 $1.7\times$ 的单步时间加速。
cs.LG / 86 / 2610.09804
BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation
BoT-GRPO:经由Token袋聚合实现用于推理的高效过程奖励强化学习
large language model
大语言模型相关
Abstract
Reinforcement learning is now central to eliciting reasoning in large language models, while in the popular algorithm Group Relative Policy Optimization (GRPO) every token in a rollout receives the same advantage. We ask how to make process supervision efficient: accelerating convergence and improving final quality without the cost of value networks. We propose Bag-of-Tokens Group Relative Policy Optimization (BoT-GRPO), which extends GRPO to token-level reward models through a length-invariant "bag of tokens" aggregation: it collects all token-level rewards across rollouts, weights each by the inverse of its source sequence length, and computes per-token advantages relative to weighted group statistics. BoT-GRPO is critic-free, and is a drop-in replacement wherever GRPO is used when token-level reward is available. On React front-end code generation, BoT-GRPO reaches $80\%$ compile rate up to $1.9\times$ faster than GRPO and converges faster than modern GRPO variants (GSPO, DAPO, PURE) while reaching higher final compile and VLM-judged win rates. On a second task, AIME mathematical reasoning, BoT-GRPO delivers absolute Pass@$k$ gains up to $8.1\%$ over GRPO in half the steps. For both tasks we compare the algorithm's performance on reasoning vs. non-reasoning base-model families (Qwen2.5-3B, SmolLM3-3B, Phi-4-mini-reasoning). Our experiments also yield a practical recipe for the reward model itself: reward stability matters more than richness: clean, bounded, stable fine-grained signals consistently accelerate learning where noisier alternatives stall.
Chinese Translation
强化学习现在对于激发大语言模型中的推理能力至关重要,而在流行算法组相对策略优化(Group Relative Policy Optimization, GRPO)中,一次 rollout 中的每个 token 都获得相同的优势值。我们追问如何使过程监督高效:在不付出价值网络成本的情况下加速收敛并提升最终质量。我们提出Token袋组相对策略优化(Bag-of-Tokens Group Relative Policy Optimization, BoT-GRPO),它通过一种长度不变的“Token袋”聚合将 GRPO 扩展到 token 级奖励模型:它收集所有 rollout 中的 token 级奖励,按每个奖励的来源序列长度的倒数进行加权,并相对于加权组统计量计算每个 token 的优势。BoT-GRPO 无需 critic,并且在 token 级奖励可用时,可在任何使用 GRPO 的地方直接替换。在 React 前端代码生成任务上,BoT-GRPO 达到 $80\%$ 编译率的速度比 GRPO 最高快 $1.9\times$,并且比现代 GRPO 变体(GSPO、DAPO、PURE)收敛更快,同时达到更高的最终编译率和 VLM 评判的胜率。在第二个任务 AIME 数学推理上,BoT-GRPO 在仅用一半步数的情况下,相较 GRPO 带来最高 $8.1\%$ 的绝对 Pass@$k$ 提升。对于这两个任务,我们比较了该算法在推理型与非推理型基础模型系列(Qwen2.5-3B、SmolLM3-3B、Phi-4-mini-reasoning)上的性能。我们的实验还为奖励模型本身给出了一套实用配方:奖励稳定性比丰富性更重要:干净、有界、稳定的细粒度信号能够持续加速学习,而噪声更多的替代方案则会停滞。
cs.LG / 87 / 2610.09877
Layerwise Error Attribution for Fast and Robust Mixed-Precision Post-Training Quantization
用于快速且鲁棒的混合精度训练后量化的逐层误差归因
diffusion
扩散模型相关
Abstract
Mixed-precision post-training quantization is a network compression method that assigns bits layer by layer, under a global memory budget using a small calibration set. The main difficulties are to overcome the combinatorial nature of the allocation problem and to manage the sensitivity to small, potentially corrupted databases. Hence, an efficient allocation method should be fast to compute and preserve model quality when calibration data are corrupted. To design such a method, we derive a layerwise probabilistic analysis of the quantization error that separates propagated error from the local perturbation introduced at a given layer. We use this local term to build a separable score for a simple allocation algorithm, that requires no external solver. The probabilistic nature of our approach brings robustness to corrupted data. On denoising tasks with DRUNet, with an average budget of 4 bits per weight, our method matches or improves state-of-the-art mixed-precision baselines under clean calibration, and is more robust to corrupted calibration, with PSNR gains of up to 7.5 dB under the tested corruptions. Experiments show bit-allocation speed-ups from 28x to 2,570x over the studied baselines. For quantized diffusion models, our experiments show that a direct application of our framework also improves the state-of-the-art.
Chinese Translation
混合精度训练后量化是一种网络压缩方法,它在全局内存预算下,使用小型校准集逐层分配比特。主要困难在于克服分配问题的组合性质,并处理对小型、可能被损坏的数据集的敏感性。因此,一种高效的分配方法应当计算快速,并在校准数据被损坏时保持模型质量。为了设计这样一种方法,我们推导了量化误差的逐层概率分析,该分析将传播误差与在给定层引入的局部扰动分离开来。我们使用该局部项为一个简单的分配算法构建可分离评分,该算法不需要外部求解器。我们方法的概率性质为损坏数据带来了鲁棒性。在使用 DRUNet 的去噪任务上,以每个权重平均 4 比特的预算,我们的方法在干净校准下达到或优于最先进的混合精度基线,并且对损坏校准更鲁棒,在测试的损坏下 PSNR 增益最高达 7.5 dB。实验表明,相对于所研究的基线,比特分配速度提升了 28 倍到 2,570 倍。对于量化扩散模型,我们的实验表明,直接应用我们的框架也能改进最先进水平。
cs.LG / 88 / 2610.09914
RollVerify: Bridging Efficiency and Accuracy in Long-Tail Rollout Reinforcement Learning
RollVerify:弥合长尾 rollout 强化学习中的效率与准确性
large language model
大语言模型相关
Abstract
Reinforcement learning is crucial for improving large language models' reasoning and generalization. It relies on massive rollouts whose lengths become increasingly long-tailed as context windows grow. In on-policy training, these long-tail rollouts can result in GPU bubbles, reducing system utilization and limiting RL scalability. Asynchronous or partial-rollout methods improve throughput by relaxing synchronization, but inevitably introduce stale off-policy samples (trajectories) that may hurt final accuracy. Existing approaches mainly mitigate this off-policy issue by reweighting off-policy samples during training, yet they can still leave a performance gap compared to fully on-policy training. In this work, rather than passively reweighting samples during training, we propose RollVerify, a lightweight RL framework built on partial rollout that actively verifies and repairs samples before they enter training. Specifically, it introduces an off-policy shift metric OPS, to quantify the off-policy deviation of partially generated trajectories. Guided by the OPS constraint, RollVerify performs both sequence-level and token-level verification to identify and truncate invalid suffixes of trajectories. This yields high-quality samples that protect the models' accuracy while preserving the efficiency gains of partial rollout. Experiments on mathematical and tool-assisted mathematical reasoning show that RollVerify achieves accuracy comparable to on-policy training while reducing training cost. Additional code-generation results provide preliminary evidence beyond mathematics.
Chinese Translation
强化学习对于提升大语言模型的推理和泛化能力至关重要。它依赖于大规模 rollout,而随着上下文窗口增大,这些 rollout 的长度变得越来越长尾化。在同策略训练中,这些长尾 rollout 会导致 GPU 气泡,降低系统利用率并限制 RL 的可扩展性。异步或部分 rollout 方法通过放宽同步来提升吞吐量,但不可避免地引入陈旧的异策略样本(轨迹),这可能损害最终准确率。现有方法主要通过在训练期间对异策略样本重新加权来缓解这一异策略问题,但与完全同策略训练相比,它们仍可能留下性能差距。在这项工作中,我们并非在训练期间被动地对样本重新加权,而是提出 RollVerify,一个基于部分 rollout 的轻量级 RL 框架,它在样本进入训练之前主动验证并修复样本。具体而言,它引入一个异策略偏移度量 OPS,以量化部分生成轨迹的异策略偏差。在 OPS 约束的指导下,RollVerify 执行序列级和 token 级验证,以识别并截断轨迹中无效的后缀。这会生成高质量样本,在保留部分 rollout 的效率增益的同时保护模型的准确率。在数学和工具辅助数学推理上的实验表明,RollVerify 在降低训练成本的同时,达到了与同策略训练相当的准确率。额外的代码生成结果提供了数学之外的初步证据。
cs.LG / 89 / 2610.09993
Efficient Patch-Based Anomaly Detection Fused with Diffusion Driven Generative Modeling for Semiconductor Wafer Bin Map Open Set Anomaly Detection
面向半导体晶圆图开集异常检测的融合扩散驱动生成建模的高效基于图块的异常检测
diffusion
扩散模型相关
Abstract
Spatial defect signatures on wafer bin maps (WBMs) trace yield loss to specific process faults, yet supervised classifiers recognize only the defect types seen during training, and one-class detectors built on a single mechanism tend to capture either local structural deviations or global distributional violations, but rarely both. This work proposes a hybrid one-class framework that couples a patch-based student-teacher detector (EfficientAD) with a denoising diffusion probabilistic model (DDPM) used for partial-diffusion reconstruction, and fuses their percentile-calibrated scores through a fixed convex combination. Trained on only 700 normal wafers from the WM-38K mixed-type dataset and evaluated on 18,658 held-out wafers, the fused detector reached an AUROC of 0.9985 and reduced misclassifications from 852 (DDPM) and 1,412 (EfficientAD) to 618, with all pairwise differences significant at p < 0.001. Beyond aggregate accuracy, the analysis shows that the gain arises from weakly overlapping errors between the two modules, yet fixed-weight fusion recovers only 40-70% of the correction available to an oracle selector. Under the benchmark's inverted class balance, average precision and F1 saturate, while the Matthews correlation coefficient and negative predictive value expose unreliable normal predictions. Pixel-level maps further show that strong image-level separability does not imply spatial localization, and the diffusion module succeeds as a local density prior rather than through global geometric reasoning. These findings motivate sample-adaptive fusion and imbalance-aware evaluation of hybrid wafer anomaly detectors.
Chinese Translation
晶圆图(WBM)上的空间缺陷特征可将良率损失追溯到特定的工艺故障,然而有监督分类器只能识别训练期间见过的缺陷类型,而基于单一机制构建的单类检测器往往只能捕捉局部结构偏差或全局分布违例中的一种,很少能同时兼顾二者。本工作提出了一种混合单类框架,它将基于图块的学生-教师检测器(EfficientAD)与用于部分扩散重建的去噪扩散概率模型(DDPM)耦合起来,并通过固定的凸组合对二者经百分位校准的分数进行融合。该融合检测器仅在 WM-38K 混合类型数据集的 700 片正常晶圆上训练,并在 18,658 片留出晶圆上进行评估,达到了 0.9985 的 AUROC,并将误分类数从 852(DDPM)和 1,412(EfficientAD)减少到 618,所有成对差异在 p < 0.001 水平上均显著。除总体准确率之外,分析表明该增益源于两个模块之间弱重叠的误差,然而固定权重融合只能恢复预言机(oracle)选择器所能获得校正量的 40–70%。在该基准的倒置类别平衡下,平均精度与 F1 趋于饱和,而马修斯相关系数与阴性预测值则暴露出正常预测的不可靠性。像素级图进一步表明,强的图像级可分性并不意味着空间定位能力,扩散模块是作为局部密度先验而非通过全局几何推理而取得成功的。这些发现推动了对混合晶圆异常检测器开展样本自适应融合与不平衡感知评估的研究。
cs.LG / 90 / 2610.10199
Robust Decentralized Fairness Auditing
鲁棒的去中心化公平性审计
large language model
大语言模型相关
Abstract
Emerging legislation requires large language models (LLMs) to be audited for compliance with regulatory standards, particularly fairness. Such black-box audits typically assume a single auditor with access to a large, representative set of queries. In practice, it can be difficult for an auditor to obtain such a query set, but multiple auditors can together cover the relevant demographic groups by auditing the LLM collaboratively with their individual query sets. However, relying on multiple auditors raises a fundamental trust problem, as they may act on behalf of the LLM provider to portray a misleading appearance of fairness, i.e., fairwashing. We propose Auditopus, a novel approach for robust decentralized fairness auditing. In Auditopus, auditing proceeds in rounds without a central server. In each round, every auditor issues a fixed number of queries to the LLM, and sends only cumulative statistics vectors of its query results to other auditors instead of sensitive queries in clear. The fairness of the audited LLM is then estimated by aggregating all the vectors. We show theoretically and empirically that even a single adversarial auditor in the network can steer this estimate by fabricating the vectors it sends, making an unfair LLM appear fair. To address this threat, Auditopus has each honest auditor locally down-weight any auditor whose cumulative statistics vectors are statistically inconsistent with previous ones. We implement Auditopus and compare it to robust aggregation baselines on two datasets with two pre-trained LLMs. Against an attacker that optimizes the vectors it sends to make the LLM appear fair, Auditopus reduces audit error by up to 78% on average relative to no defense and at least 62% relative to the robust aggregation baselines. Even when 49% of the auditors are adversarial, Auditopus never lets a very unfair or moderately unfair LLM pass as fair.
Chinese Translation
新出现的立法要求对大型语言模型(LLM)进行审计,以检查其是否符合监管标准,尤其是公平性。此类黑盒审计通常假设有一个审计者,能够访问一个大型且具有代表性的查询集。在实践中,单个审计者可能难以获得这样的查询集,但多个审计者可以通过各自独立的查询集协作审计 LLM,从而共同覆盖相关的人口统计群体。然而,依赖多个审计者会引发一个根本性的信任问题,因为他们可能代表 LLM 提供方行事,以营造具有误导性的公平表象,即 fairwashing(公平洗白)。我们提出 Auditopus,一种用于鲁棒去中心化公平性审计的新方法。在 Auditopus 中,审计在没有中央服务器的情况下分轮进行。在每一轮中,每个审计者向 LLM 发出固定数量的查询,并且只向其他审计者发送其查询结果的累积统计向量,而不是明文发送敏感的查询。随后,通过聚合所有向量来估计被审计 LLM 的公平性。我们从理论和实证上表明,网络中即使只有一个对抗性审计者,也可以通过伪造其发送的向量来操纵这一估计,使不公平的 LLM 显得公平。为应对这一威胁,Auditopus 让每个诚实审计者在本地对任何其累积统计向量与先前向量在统计上不一致的审计者进行降权。我们实现了 Auditopus,并在两个数据集上使用两个预训练 LLM 将其与鲁棒聚合基线进行比较。面对一个通过优化其发送的向量来使 LLM 显得公平的攻击者,Auditopus 相对于无防御平均将审计误差降低最多 78%,相对于鲁棒聚合基线至少降低 62%。即使 49% 的审计者是对抗性的,Auditopus 也绝不会让一个非常不公平或中等不公平的 LLM 被当作公平通过。
cs.LG / 91 / 2610.10210
OrthoGen: A Generative Orthogonal Learner for Time-Varying Treatments
OrthoGen:一种用于时变治疗的生成式正交学习器
diffusion
扩散模型相关
Abstract
Estimating conditional distributional potential outcomes (CDPOs) over time is important in medicine (e.g., to estimate patient-specific risks under different treatment sequences). However, this task is challenging because of time-varying confounding, yet existing adjustment strategies for this task are limited. In this paper, we aim to learn CDPOs under time-varying treatments using flexible generative models. Our contributions are two-fold. (1) We introduce a tailored adjustment strategy for our setting, namely, generative recursive g-computation. Our adjustment strategy recursively propagates full conditional outcome distributions rather than conditional means, modeling the variables of interest directly rather than full trajectories. Building on our adjustment strategy, we formulate simple generative learners for CDPO estimation. However, these learners can be sensitive to nuisance estimation errors, which motivates an orthogonal learner. (2) We thus introduce OrthoGen, a Neyman-orthogonal and doubly robust generative learner. Importantly, we show that OrthoGen further achieves rate double robustness and quasi-oracle efficiency under suitable conditions. Our learners are flexible and can be instantiated with different generative backbones (e.g., normalizing flows and diffusion models). Across experiments with synthetic, semi-synthetic and real-world datasets, we find that OrthoGen is highly effective. To the best of our knowledge, we are the first to propose a generative orthogonal learner for estimating CDPOs under time-varying treatments.
Chinese Translation
估计随时间变化的条件分布潜在结果(CDPOs)在医学中很重要(例如,估计不同治疗序列下患者特定的风险)。然而,由于存在时变混杂,该任务具有挑战性,但用于该任务的现有调整策略有限。在本文中,我们旨在使用灵活的生成模型学习时变治疗下的 CDPOs。我们的贡献有两方面。(1) 我们为我们的设定引入了一种量身定制的调整策略,即生成式递归 g-computation。我们的调整策略递归传播完整的条件结果分布而非条件均值,直接对感兴趣的变量建模而非完整轨迹。基于我们的调整策略,我们为 CDPO 估计构建了简单的生成式学习器。然而,这些学习器可能对干扰参数估计误差敏感,这促使了正交学习器的提出。(2) 因此,我们引入了 OrthoGen,一种 Neyman 正交且双重稳健的生成式学习器。重要的是,我们表明 OrthoGen 在适当条件下进一步实现了速率双重稳健性和准 oracle 效率。我们的学习器是灵活的,可以用不同的生成式骨干网络(例如,归一化流和扩散模型)进行实例化。在合成、半合成和真实世界数据集的实验中,我们发现 OrthoGen 非常有效。据我们所知,我们是第一个提出用于估计时变治疗下 CDPOs 的生成式正交学习器。
cs.LG / 92 / 2610.10227
From Prompts to Trees: Effective LLM-Guided Tree Generation for Few-Shot Tabular Classification
从提示到树:面向少样本表格分类的有效LLM引导树生成
large language model
大语言模型相关
Abstract
While Large Language Models (LLMs) possess rich world knowledge and impressive generalization capabilities, their direct application to tabular data classification is hindered by high inference costs and limited interpretability. In contrast, decision trees are fast and transparent but often underperform in low-data regimes. In this work, we propose a novel framework that bridges these paradigms by distilling LLM knowledge into interpretable decision trees under a few-shot learning setting. Instead of directly prompting the LLM to generate full trees, which is often unstable and inefficient, we develop a three-stage paradigm that prompts the LLM to generate rules and organize the rules into a tree. Experiments on multiple real-world tabular datasets demonstrate that our method achieves superior accuracy and interpretability with significantly lower prompting overhead compared to existing baselines.
Chinese Translation
尽管大型语言模型(LLMs)拥有丰富的世界知识和令人印象深刻的泛化能力,但其直接应用于表格数据分类仍受到高推理成本和有限可解释性的阻碍。相比之下,决策树快速且透明,但在低数据场景下往往表现不佳。在这项工作中,我们提出了一个新框架,该框架通过在少样本学习设置下将LLM知识蒸馏到可解释的决策树中,连接了这两种范式。我们没有直接提示LLM生成完整的树(这通常不稳定且低效),而是开发了一种三阶段范式,提示LLM生成规则,并将这些规则组织成一棵树。在多个真实世界表格数据集上的实验表明,与现有基线相比,我们的方法以显著更低的提示开销实现了更优的准确性和可解释性。
cs.LG / 93 / 2610.10304
SemanticFold: Latent Sequence Compression SeparatesLanguage Modeling, Decodability, and Reasoning
SemanticFold:潜在序列压缩分离语言建模、可解码性与推理
large language model
大语言模型相关
Abstract
We study whether latent sequence compression of prompt prefixes preserves the capabilities that large language models rely on during inference. We introduce SemanticFold, a compression scheme that folds prefix hidden states at learned boundaries, and evaluate it across five model scales: Qwen3-1.7B, Qwen3-8B, SmolLM2-1.7B, Pythia-1.4B, and Pythia-6.9B. We use a fixed-target protocol: a frozen prefix is executed natively or compressed, and both arms teacher-force identical continuation tokens. This design rules out target-selection explanations for likelihood changes. We examine five endpoint families: fixed-target negative log-likelihood, finite-label reasoning accuracy, linear probe accessibility, open-ended generation, and systems-level memory and latency. We find that compression moves these endpoints non-monotonically and that they do not share a single compression threshold. On Qwen3-1.7B at compression ratio R=1.7, compressed-minus-native mean NLL decreases by 0.135 under paired bootstrap with 10000 draws. On SmolLM2 at R=1.2, the mean change is 0.013 higher than native. On both Pythia checkpoints, NLL is effectively unchanged. An NLL decomposition separating sequence shortening from the learned residual transform shows that the favorable Qwen likelihood is attributable primarily to residual adaptation rather than to shortening alone. MLP-only, which applies the transform without shortening, achieves 0.082 lower NLL than Full SemanticFold. Linear probe accuracy and macro AUC change by less than 0.03 in absolute value across conditions, with confidence intervals crossing zero. We conclude that preservation under latent compression has no single scalar certificate: language-model fit, decodability, and reasoning behavior answer different questions and can move in different directions under the same compression operation.
Chinese Translation
我们研究提示前缀的潜在序列压缩是否保留大语言模型在推理时所依赖的能力。我们提出 SemanticFold,一种在学习到的边界处折叠前缀隐藏状态的压缩方案,并在五种模型规模上对其进行评估:Qwen3-1.7B、Qwen3-8B、SmolLM2-1.7B、Pythia-1.4B 和 Pythia-6.9B。我们采用固定目标协议:冻结的前缀要么以原生方式执行,要么被压缩,且两条分支都对相同的续写 token 进行教师强制。这一设计排除了用目标选择来解释似然变化的可能性。我们考察五类终点指标:固定目标负对数似然、有限标签推理准确率、线性探针可访问性、开放式生成,以及系统级内存和延迟。我们发现压缩以非单调方式改变这些终点指标,并且它们并不共享单一压缩阈值。在 Qwen3-1.7B 上,当压缩比 R=1.7 时,在 10000 次抽样的配对自助法下,压缩减原生的平均 NLL 降低 0.135。在 SmolLM2 上,当 R=1.2 时,平均变化比原生高 0.013。在两个 Pythia 检查点上,NLL 实际上没有变化。一种将序列缩短与学习到的残差变换分离的 NLL 分解表明,Qwen 上有利的似然主要归因于残差适应,而不是仅归因于缩短。仅 MLP,即在不进行缩短的情况下应用该变换,其 NLL 比完整 SemanticFold 低 0.082。在不同条件下,线性探针准确率和宏 AUC 的绝对变化小于 0.03,且置信区间跨越零。我们得出结论:潜在压缩下的保持没有单一的标量证明:语言模型拟合、可解码性和推理行为回答不同的问题,并可在同一压缩操作下朝不同方向变化。
cs.LG / 94 / 2610.10366
Koopman Observers for Diffusion Acceleration: Correcting Feature Forecasts with Shallow Measurements
用于扩散加速的 Koopman 观测器:用浅层测量校正特征预测
diffusion
扩散模型相关
Abstract
Feature caching accelerates diffusion sampling by replacing expensive network evaluations with predictions from previously computed activations. However, forecasts based only on past features cannot directly incorporate changes in the current denoising state. We investigate whether inexpensive, freshly computed features can serve as observations for correcting these predictions. We introduce an observation-corrected Koopman framework for accelerating frozen diffusion models. Using calibration trajectories, we identify finite-dimensional, time-dependent Koopman approximations that jointly describe the increments of shallow and deep network features. During accelerated sampling, these operators predict the evolution of expensive deep features, while innovations in the observed shallow features correct the predicted state. Periodic full evaluations refresh the observer, and all generative-model parameters remain unchanged. This formulation enables controlled comparisons of temporal prediction and observation correction. Across three 10,000-image runs per dataset, our method reduces paired Inception-feature MSE by $19.9\%$ on CIFAR-10 and $11.9\%$ on a ten-class ImageNet subset relative to channelwise affine prediction under the same four-partial-step schedule. Matched ablations attribute additional reductions of $4.54\%$ and $4.67\%$ to observation correction. The observer achieves $1.89\times$ and $1.85\times$ measured speedups over DDIM-50, supporting improved reference-sampler fidelity without retraining the denoiser.
Chinese Translation
特征缓存通过用来自先前计算的激活的预测替换昂贵的网络评估,加速了扩散采样。然而,仅基于过去特征的预测无法直接纳入当前去噪状态的变化。我们研究廉价的、新计算的特征是否可以作为观测来校正这些预测。我们提出一种观测校正的 Koopman 框架,用于加速冻结的扩散模型。利用校准轨迹,我们辨识出有限维、时间依赖的 Koopman 近似,这些近似联合描述浅层与深层网络特征的增量。在加速采样期间,这些算子预测昂贵的深层特征的演化,而观测到的浅层特征中的新息校正预测状态。周期性完整评估刷新观测器,且所有生成模型参数保持不变。这一表述使得能够对时间预测和观测校正进行受控比较。在每个数据集三次 10,000 张图像的运行中,我们的方法在相同的四部分步调度下,相对于逐通道仿射预测,在 CIFAR-10 上将配对 Inception 特征 MSE 降低 $19.9\%$,在十类 ImageNet 子集上降低 $11.9\%$。匹配的消融实验将额外的 $4.54\%$ 和 $4.67\%$ 的降低归因于观测校正。该观测器相对于 DDIM-50 实现了 $1.89\times$ 和 $1.85\times$ 的实测加速,在无需重新训练去噪器的情况下支持改进的参考采样器保真度。
cs.LG / 95 / 2610.10411
Training Parallel Speculative Draft Models by Directly Minimizing Expected Decoding Rounds
通过直接最小化期望解码轮数来训练并行推测草稿模型
large language model
大语言模型相关
Abstract
Speculative decoding accelerates large language model inference by using a low-cost draft model to propose tokens that the full-size target model verifies in parallel. Parallel and semi-autoregressive (semi- AR) drafters improve drafting efficiency by proposing an entire block in a single forward pass, but training them raises a new difficulty: the draft distribution for a given position depends on where the decoding round starts, and where rounds start depends on how many tokens earlier rounds accepted. Existing training objectives typically rely on block-local surrogates that ignore this cross-round coupling, and therefore do not directly optimize the global decoding efficiency. In this work, we develop a theoretical framework for training and evaluating these drafters by representing speculative decoding as a Markov reward process. This formulation yields the Expected Decoding Rounds (EDR) objective, which weights local rejection costs by state occupancies and exactly equals the expected number of decoding rounds. Unlike prior surrogate objectives, EDR introduces no auxiliary hyperparameters. We then derive an exact temporal-difference gradient that supports unbiased stochastic optimization from target-model rollouts. The same framework also yields an exact offline evaluator for round counts, enabling paired drafter comparisons on shared target rollouts without running speculative decoding. Finetuning two state-of-the- art drafters, DSpark and DFly, with EDR consistently improves mean accepted length and outperforms existing training objectives across nine benchmarks spanning math reasoning, code generation, and chat.
Chinese Translation
推测解码通过使用低成本草稿模型提出 token,由全尺寸目标模型并行验证,从而加速大语言模型推理。并行和半自回归(semi-AR)草稿器通过在一次前向传播中提出整个块来提升起草效率,但训练它们带来了一个新的困难:给定位置的草稿分布取决于解码轮从哪里开始,而轮从哪里开始又取决于更早的轮次接受了多少 token。现有的训练目标通常依赖于块局部的替代目标,忽略了这种跨轮耦合,因此并不直接优化全局解码效率。在这项工作中,我们通过将推测解码表示为一个马尔可夫奖励过程,开发了一个用于训练和评估这些草稿器的理论框架。这一形式化产生了期望解码轮数(EDR)目标,它按状态占用率对局部拒绝成本进行加权,并且恰好等于期望的解码轮数。与先前的代理目标不同,EDR 不引入任何辅助超参数。然后,我们推导出一个精确的时序差分梯度,它支持从目标模型轨迹进行无偏随机优化。同一框架还产生了一个用于轮数的精确离线评估器,使得能够在共享的目标模型轨迹上对草稿器进行配对比较,而无需运行推测解码。用 EDR 微调两个最先进的草稿器 DSpark 和 DFly,可持续提高平均接受长度,并在涵盖数学推理、代码生成和聊天的九个基准上优于现有训练目标。
cs.LG / 96 / 2610.10440
Seq-Flow: Efficient Probabilistic Forecasting with Self-Rollout Error Control
Seq-Flow:具有自展开误差控制的高效概率预测
diffusion
扩散模型相关
Abstract
Many scientific forecasting tasks require updating a distribution over future trajectories as new observations arrive. Conventional diffusion and flow models generate each forecast from Gaussian noise, often at the cost of many sampling steps. Warm-start methods reuse earlier predictions to reduce this cost, but their models are not trained to perform the forecast update itself, which can compromise quality under few-step sampling. In this work, we introduce Seq-Flow, a conditional flow model whose ODE transports samples from the previous forecast distribution to the updated one. Because successive forecasts often differ only modestly, this transport starts from an informative distribution and can produce accurate updates with few flow evaluations. Recursive reuse also creates a challenge: errors in one forecast become errors in the initial states of subsequent flows. We address this with self-rollout training, in which a moving average copy of the model generates forecasts that initialize later training updates. Unlike self-forcing methods, which reuse generated outputs as conditioning context, Seq-Flow reuses them as the source of the next flow. Experiments On particle-accelerator beam spill forecasting show Seq-Flow reduces CRPS by 65% under a few-NFE sampling budget, while remaining competitive with strong baselines on fluid-dynamics forecasting tasks. Although trained on self-rollouts of at most four updates, Seq-Flow remains accurate over more than 400 consecutive updates. Our code is available at https://github.com/Graph-COM/Seq-Flow.
Chinese Translation
许多科学预测任务需要随着新观测的到来,更新未来轨迹上的分布。传统的扩散模型和流模型从高斯噪声生成每一次预测,往往要以大量采样步数为代价。热启动方法复用先前的预测以降低这一代价,但它们的模型并未被训练来执行预测更新本身,这可能在少步采样下损害质量。在这项工作中,我们提出 Seq-Flow,一种条件流模型,其 ODE 将样本从先前的预测分布输运到更新后的分布。由于相邻的预测往往只有适度差异,这种输运从一个信息丰富的分布出发,并能够以很少的流评估产生准确的更新。递归复用还带来一个挑战:某一次预测中的误差会成为后续流初始状态中的误差。我们通过自展开训练来解决这一问题,其中模型的一个移动平均副本生成预测,用以初始化后续的训练更新。与将生成输出复用为条件上下文的自强迫方法不同,Seq-Flow 将生成输出复用为下一个流的源。在粒子加速器束流溢出预测上的实验表明,Seq-Flow 在少 NFE 采样预算下将 CRPS 降低 65%,同时在流体动力学预测任务上与强基线相比仍具竞争力。尽管最多只在四次更新的自展开上进行训练,Seq-Flow 在超过 400 次连续更新中仍保持准确。我们的代码可在 https://github.com/Graph-COM/Seq-Flow 获取。
cs.LG / 97 / 2610.09778
Reproducible LLM Inference Benchmarking: A Sequential Isolation Protocol for Regression Testing
可复现的LLM推理基准测试:一种用于回归测试的序列隔离协议
large language model
大语言模型相关
Abstract
Reproducible benchmarking of Large Language Model (LLM) inference is challenging because repeated measurements can vary with execution and system state. We present the Sequential Isolation Methodology, a controlled benchmarking and regression-testing protocol designed to reduce between-run measurement variance while deliberately varying workload concurrency. We evaluate three representative open-source LLMs on an NVIDIA A100 80GB GPU using vLLM 0.9.1 across six context sizes and eight concurrency levels, with five repetitions per configuration. The final protocol reduces average coefficient of variation (CV) from 15.2% in the least controlled methodology stage to 2.2% under the final protocol; using CV computed across the five repetition-level median (P50) TTFT values per configuration, 113 of 144 configurations (78.5%) achieve CV below 3%. The measurements also show a marked latency transition between 200 and 500 concurrent users on the tested stack and descriptive differences in P99 latency across the three models. We additionally provide an explicit cost break-even model with sensitivity to API pricing. The protocol is intended to provide a stable reference for reproducible comparison and regression testing rather than to predict absolute behavior under uncontrolled production traffic. Infrastructure-as-Code and benchmark scripts support replication of the experimental environment.
Chinese Translation
大型语言模型(LLM)推理的可复现基准测试具有挑战性,因为重复测量结果会随执行过程和系统状态而变化。我们提出序列隔离方法(Sequential Isolation Methodology),这是一种受控的基准测试与回归测试协议,旨在减少运行间测量方差,同时有意改变工作负载并发度。我们在 NVIDIA A100 80GB GPU 上使用 vLLM 0.9.1 评估了三个具有代表性的开源 LLM,涵盖六种上下文长度和八种并发级别,每种配置重复五次。最终协议将平均变异系数(CV)从控制最弱的方法阶段的 15.2% 降低至最终协议下的 2.2%;基于每种配置下五次重复的中位数(P50)TTFT 值计算的 CV,144 种配置中有 113 种(78.5%)的 CV 低于 3%。测量结果还显示,在所测试的技术栈上,200 至 500 并发用户之间出现明显的延迟转变,并且三个模型在 P99 延迟上存在描述性差异。我们另外提供了一个显式的成本盈亏平衡模型,并对 API 定价进行了敏感性分析。该协议旨在为可复现比较和回归测试提供稳定参考,而非预测不受控生产流量下的绝对行为。基础设施即代码(Infrastructure-as-Code)和基准测试脚本支持实验环境的复现。
cs.LG / 98 / 2610.09369
Immiscible Diffusion Policy: Preserving Multimodal Robot Actions through Label-Free Noise Assignment
不可混合扩散策略:通过无标签噪声分配保留多模态机器人动作
diffusion
扩散模型相关
Abstract
When diffusion policies were first introduced, they were expected to recover multi-modal action distributions. However, we find this expectation does not always hold, as diffusion policies often collapse to a single modality even when we guarantee the balance of dataset modalities and exact within-batch symmetry. Our analysis indicates that independent action-noise pairing contributes to this failure by increasing mixing and crossing among diffusion paths, which can produce averaged denoising responses and suppress modality-specific behavior. This issue is especially severe in robot planning, where action spaces are dense and low-dimensional, significantly increasing such mixing and crossing. To alleviate this problem, we propose Immiscible Diffusion Policy, a label-free training-time add-on to diffusion policy that uses action-noise assignment to preserve relatively distinct noise-to-action routes without modifying the policy architecture or inference procedure. Across five simulated and two real-world humanoid manipulation tasks spanning state, RGB, and point-cloud observations, our method significantly improves the policy's preservation of action modalities while maintaining strong task performance. It increases the proportion of the non-dominant modality by 6.0x-14.6x across three two-modality tasks and recovers demonstrated modalities that are entirely absent from vanilla policy rollouts on both four-modality tasks. These results demonstrate that Immiscible Diffusion Policy provides a simple yet robust approach to preserving action multi-modality in general robot learning tasks.
Chinese Translation
当扩散策略首次被提出时,人们期望它们能够恢复多模态动作分布。然而,我们发现这一期望并不总是成立,因为即便我们保证了数据集各模态的平衡以及批次内精确的对称性,扩散策略仍常常坍缩到单一模态。我们的分析表明,独立的动作-噪声配对会加剧扩散路径之间的混合与交叉,从而促成这一失败,这可能产生平均化的去噪响应并抑制模态特定的行为。该问题在机器人规划中尤为严重,因为机器人动作空间是稠密且低维的,会显著加剧此类混合与交叉。为缓解这一问题,我们提出了不可混合扩散策略(Immiscible Diffusion Policy),这是一种面向扩散策略的无标签、训练期附加方法,它利用动作-噪声分配来保留相对彼此区分开的噪声到动作的路径,而无需修改策略架构或推理过程。在涵盖状态、RGB 和点云观测的五项仿真任务和两项真实世界人形操作任务上,我们的方法在保持强劲任务性能的同时,显著提升了策略对动作模态的保留能力。在三个双模态任务上,它将非主导模态的比例提升了 6.0 倍至 14.6 倍;在两个四模态任务上,它恢复了在原始策略 rollout 中完全缺失的示教模态。这些结果表明,不可混合扩散策略为在通用机器人学习任务中保留动作多模态性提供了一种简单而稳健的方法。
cs.AI / 99 / 2610.09588
Adaptive Code Generation for Controlling Robots
面向机器人控制的自适应代码生成
large language model
大语言模型相关
Abstract
Deploying robots as Complex Adaptive Systems (CAS) in unknown and dynamic environments necessitates a transition from rigid command libraries toward intention-based autonomy, as natural language represents the only medium capable of articulating complex goals beyond the capacity of finite instruction sets. While Large Language Models (LLMs) offer a path toward natural language goal description, their integration introduces significant challenges: the formalization gap between imprecise intentions and executable actions, the taxonomy gap induced by unpredictable environments, and the challenge of maintaining temporal state and progress awareness. This work introduces an architectural framework that enables robotic control by leveraging generative AI. The system follows a dual-AI design: an LLM translates high-level intentions into executable program code restricted to a formal robotic library and constrained by verifiable syntax, while a Vision-Language Model (VLM) provides semantic grounding via a distillation process. To ensure robustness, the framework incorporates environment-driven replanning triggers based on geometric and semantic thresholds, complemented by continuous runtime monitoring and an adaptive planning loop. Benchmarked across frontier models, our framework architecture demonstrates that grounding generative AI in a reactive, constrained loop enables robust fulfillment of complex intentions in dynamic and unknown environments.
Chinese Translation
将机器人作为复杂自适应系统(CAS)部署在未知和动态环境中,需要从僵化的命令库转向基于意图的自主性,因为自然语言是唯一能够表达超出有限指令集能力的复杂目标的媒介。尽管大型语言模型(LLM)为自然语言目标描述提供了一条路径,但它们的集成带来了重大挑战:不精确意图与可执行动作之间的形式化鸿沟、由不可预测环境引发的分类学鸿沟,以及维持时间状态和进度感知的挑战。这项工作引入了一个利用生成式 AI 实现机器人控制的架构框架。该系统遵循双 AI 设计:一个 LLM 将高层意图转换为可执行程序代码,这些代码被限制在形式化机器人库内,并受可验证语法的约束,而一个视觉-语言模型(VLM)通过蒸馏过程提供语义落地。为确保鲁棒性,该框架纳入了基于几何和语义阈值的环境驱动重规划触发机制,并辅以持续的运行时监控和自适应规划循环。在多个前沿模型上进行基准测试,我们的框架架构表明,将生成式 AI 落地于反应式、受约束的循环中,能够在动态和未知环境中稳健地实现复杂意图。
cs.CL / 100 / 2610.10526
Rephrase Before You Act: Characterizing and Mitigating Language Sensitivity in Vision-Language-Action Models
先改写,再行动:刻画并缓解视觉-语言-动作模型中的语言敏感性
large language model
大语言模型相关
Abstract
Vision-language-action models (VLAs) are strikingly sensitive to instruction phrasing and do not inherit the language robustness of the vision-language models they are built on. A one-word edit can move success by tens of points: $π_{0.5}$ turns on a LIBERO stove 100% of the time for "switch on the stove" and 2% for "switch on the hot plate", and a $π_0$ checkpoint finetuned with rephrase augmentation still shows swings of up to 61 points. We characterize this sensitivity with statistically tested single-edit swings and an oracle phrase search, which shows that phrasing alone nearly closes the 21-point gap between in-distribution and out-of-distribution tasks. We then reduce it without modifying the policy. Because the sensitivity is systematic, it can be expressed as explicit rules: we score many phrasings of a few training tasks, have a large language model distill the evidence into ten to twenty rephrasing rules, and at deployment rewrite each incoming instruction once under these rules. The rules improve the frozen $π_0$ by 16 to 27% relative on twelve held-out tasks across adversarial, VLM-generated, and human-generated phrasings, with gains concentrated on out-of-distribution tasks. The pipeline replicates on $π_{0.5}$ and LIBERO, lifting in-finetune success from 93.6% to 97.8%. The method requires no retraining and no per-step verification, and applies zero-shot to unseen tasks and instructions. Project website: https://sttawm.github.io/rephrase-before-you-act
Chinese Translation
视觉-语言-动作模型(VLA)对指令措辞表现出惊人的敏感性,并且没有继承其构建所基于的视觉-语言模型的语言鲁棒性。一个单词的改动就能使成功率变化数十个百分点:对于“switch on the stove”,$π_{0.5}$ 有 100% 的概率打开 LIBERO 炉灶,而对于“switch on the hot plate”仅为 2%;而且,一个用改写增强进行微调的 $π_0$ 检查点仍然表现出高达 61 个百分点的波动。我们通过经过统计检验的单次编辑波动以及一次 oracle 短语搜索来刻画这种敏感性,该搜索表明,仅凭措辞就几乎弥合了分布内任务与分布外任务之间 21 个百分点的差距。随后,我们在不修改策略的情况下降低了这种敏感性。由于这种敏感性是系统性的,它可以被表达为显式规则:我们对少数训练任务的许多措辞进行评分,让一个大型语言模型将证据提炼为十到二十条改写规则,并在部署时依据这些规则对每一条传入指令进行一次改写。这些规则在十二个留出任务上,在对抗性、VLM 生成和人类生成的措辞中,使冻结的 $π_0$ 相对提升了 16% 到 27%,且增益集中在分布外任务上。该流程在 $π_{0.5}$ 和 LIBERO 上得到复现,将微调内成功率从 93.6% 提升至 97.8%。该方法无需重新训练,也无需逐步验证,并且能够零样本应用于未见过的任务和指令。项目网站:https://sttawm.github.io/rephrase-before-you-act
cs.SE / 101 / 2610.09079
Large-scale Repository Engineering via Agent-Native Reusable Code Primitives
基于智能体原生可复用代码原语的大规模仓库工程
large language model
大语言模型相关
Abstract
Large language models equipped with development environments have moved code generation toward repository-scale construction, yet building complete repositories remains difficult because interacting modules, interfaces, configurations, tests, and dependencies must work together. We introduce Code Primitives, agent-native reusable executable components with interface contracts, dependency closures, validation tests, and provenance. Each primitive uses a resident LLM to assess relevance and adapt its implementation, interfaces, and dependencies to the target repository, and we organize 1,424 validated primitives in CodeFace, a searchable library for repository construction. We introduce LEGO (Large-scale repository Engineering via aGent-native reusable cOde primitives), which activates task-relevant primitives, integrates their adapted implementations with task-specific code while resolving cross-component constraints, and revises the result against executed tests. To measure construction end to end, we build LEGO-REPO, a benchmark of 522 executable reconstruction tasks spanning seven software domains, 22 capability tracks, and five difficulty levels, scored against native test suites between an empty-package floor and original-source ceiling. The strongest of 13 evaluated backbones reaches a delivery score of 0.318 and scores zero on 41.0% of tasks; LEGO improves all 13 by 0.1474 on average and raises GPT-5.6-terra from 0.3180 to 0.5134 (+61.4%). In controlled comparisons, adapted primitives outperform retrieved code supplied as context or vendored unchanged. The effect persists against independent repository agents, across three external benchmarks, and with a disjointly re-mined CodeFace; GPT-OSS-20B for adaptation and diagnosis retains 95.1% of the homogeneous score at 24.0% lower cost.
Chinese Translation
配备开发环境的大型语言模型已将代码生成推向仓库规模构建,然而构建完整仓库仍然困难,因为相互作用的模块、接口、配置、测试和依赖必须协同工作。我们引入代码原语,即智能体原生的可复用可执行组件,具有接口契约、依赖闭包、验证测试和来源信息。每个原语使用常驻 LLM 来评估相关性,并使其实现、接口和依赖适应目标仓库,我们在 CodeFace 中组织了 1,424 个经过验证的原语,这是一个用于仓库构建的可搜索库。我们引入 LEGO(通过智能体原生可复用代码原语进行大规模仓库工程),它激活与任务相关的原语,将其适配后的实现与任务特定代码集成,同时解决跨组件约束,并根据执行的测试修订结果。为了端到端地衡量构建,我们构建了 LEGO-REPO,这是一个包含 522 个可执行重建任务的基准,跨越七个软件领域、22 个能力赛道和五个难度级别,其评分依据原生测试套件,介于空包下限和原始源码上限之间。在 13 个被评估的骨干模型中,最强的一个达到 0.318 的交付分数,并在 41.0% 的任务上得分为零;LEGO 将所有 13 个模型平均提升 0.1474,并将 GPT-5.6-terra 从 0.3180 提高到 0.5134(+61.4%)。在受控比较中,适配后的原语优于作为上下文提供的检索代码或未经修改而直接引入的代码。该效果在针对独立仓库智能体、跨三个外部基准以及使用不相交重新挖掘的 CodeFace 时依然存在;用于适配和诊断的 GPT-OSS-20B 以低 24.0% 的成本保留了 95.1% 的同质分数。
cs.SE / 102 / 2610.09605
GRAML: Graph-Grounded Reasoning and Multi-Task Learning for LLM-Based Software Vulnerability Detection
GRAML:面向基于大语言模型的软件漏洞检测的图锚定推理与多任务学习
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) have been widely applied to software vulnerability detection. However, their performance is often limited by insufficient use of control-flow and data-flow information. In this paper, we propose GRAML, a framework that combines graph evidence, vulnerability description generation, and multi-task training. GRAML first performs static analysis on C/C++ programs to extract critical source lines and typed line relations as structural evidence. It then uses this evidence to guide GPT-5 through the Tree-of-Thought-guided Vulnerability Reasoning (ToT-VR) process and generate vulnerability descriptions. These descriptions are further combined with Detection, Localization, and Assessment samples to build a unified four-task training dataset. We evaluate GRAML on an in-distribution (ID) test set and six out-of-distribution (OOD) datasets. The results show that GRAML achieves average F1 scores ranging from 66.67% to 68.70%, outperforming state-of-the-art baselines by up to 30.92%. Ablation experiments further show that ToT-VR and graph-guided vulnerability descriptions improve detection performance compared with standard Chain-of-Thought (CoT) reasoning and raw Code Property Graph (CPG) serializations. These findings provide practical guidance for building more reliable and secure software engineering systems with large language models.
Chinese Translation
大语言模型(LLMs)已被广泛应用于软件漏洞检测。然而,其性能往往受限于对控制流和数据流信息的利用不足。在本文中,我们提出 GRAML,一个结合了图证据、漏洞描述生成与多任务训练的框架。GRAML 首先对 C/C++ 程序执行静态分析,以提取关键源代码行和带类型的行关系作为结构证据。随后,它利用这些证据引导 GPT-5 经历思维树引导的漏洞推理(ToT-VR)过程,并生成漏洞描述。这些描述进一步与检测、定位和评估样本相结合,以构建一个统一的四任务训练数据集。我们在一个分布内(ID)测试集和六个分布外(OOD)数据集上评估 GRAML。结果表明,GRAML 取得的平均 F1 分数在 66.67% 至 68.70% 之间,相比最先进的基线最多高出 30.92%。消融实验进一步表明,与标准的思维链(CoT)推理和原始代码属性图(CPG)序列化相比,ToT-VR 和图引导的漏洞描述提升了检测性能。这些发现为利用大语言模型构建更可靠、更安全的软件工程系统提供了实践指导。
cs.SE / 103 / 2610.09901
A Chat Assistant for Software Exploration in a 3D Software Visualization
用于3D软件可视化中软件探索的聊天助手
large language model
大语言模型相关
Abstract
We present a chat assistant for interactive software exploration, embedded in the 3D software visualization tool ExplorViz. The assistant builds upon current Large Language Models (LLMs) and enables users to ask questions about the currently visualized software system and trigger actions that change the visualization through natural language. We integrate the chat assistant in our ExplorViz frontend using the CopilotKit libraries such that probabilistic LLMs are combined with deterministic and tool-based actions similar to implementations using the Model Context Protocol (MCP). The chat assistant is also enabled to restructure the software system by adding, removing, or modifying parts of the software system in the visualization. An empirical experiment with eleven participants evaluated both perceived comprehension support and the tool calls that were triggered by the chat assistant. Participants rated the assistant's generated summaries and explanations as largely correct. They also reported high usability for actions like highlighting entities in the visualization and the creation of new color themes. In contrast, open-ended chat-assisted software restructuring in the visualization showed mixed results. This suggests a need for stronger guardrails for the employed LLM and better incorporation of user feedback. Overall, the assistant was perceived as usable and promising for reducing interaction overhead during exploratory program comprehension tasks. We provide a video presenting the chat assistant's use and a reproduction package of the software system that was used in our evaluation.
Chinese Translation
我们提出了一种用于交互式软件探索的聊天助手,它嵌入在3D软件可视化工具ExplorViz中。该助手基于当前的大语言模型(LLM),使用户能够就当前所可视化的软件系统提出问题,并通过自然语言触发改变可视化的操作。我们使用CopilotKit库将聊天助手集成到我们的ExplorViz前端中,从而将概率性的大语言模型与确定性的、基于工具的操作相结合,类似于使用模型上下文协议(MCP)的实现。该聊天助手还能够通过在可视化中对软件系统的部分进行添加、删除或修改来重构软件系统。一项有十一名参与者参与的实证实验评估了感知到的理解支持以及由聊天助手触发的工具调用。参与者认为该助手生成的摘要和解释大体上是正确的。他们还报告说,对于诸如在可视化中高亮实体以及创建新颜色主题之类的操作,可用性很高。相比之下,在可视化中进行的开放式、由聊天辅助的软件重构则显示出好坏参半的结果。这表明需要为所采用的大语言模型设置更强的护栏,并更好地纳入用户反馈。总体而言,该助手被认为可用,并且在减少探索性程序理解任务中的交互开销方面具有前景。我们提供了一段展示该聊天助手使用的视频,以及一份我们评估中所用软件系统的复现包。
cs.SE / 104 / 2610.09935
AgentTracer: Tracing Indirect Prompt Injection Attack through Fine-Grained Intention-Execution Alignment
AgentTracer:通过细粒度意图-执行对齐追踪间接提示注入攻击
large language model
大语言模型相关
Abstract
Large language model (LLM) agents interact with external resources to complete complex user tasks, exposing them to indirect prompt injection (IPI), where malicious instructions redirect agents toward attacker-intended tasks. Since IPI is difficult to defend against in real-world environments, post-incident tracing is essential for locating the injection source and reconstructing the attack chain. However, existing tracing methods primarily capture explicit control-flow and data-flow dependencies, overlooking the implicit relationships among tool calls driven by the malicious instruction. These tool calls may lack explicit dependencies and be interleaved with legitimate operations, making complete attack-chain reconstruction difficult. In this paper, we present AgentTracer, an intent-aware tracing framework that treats IPI as task intent drift. AgentTracer recovers implicit decision dependencies among tool calls to construct an Intent-Driven Execution Graph that connects dispersed tool calls by task intent. It combines the user request with an operation knowledge base to construct a user intent authorization space and identify intent-drift tool calls. Starting from an anomalous tool call under audit, AgentTracer performs target-based pruning and backward tracing to reconstruct the attack chain and locate the injection source and injection point. To evaluate AgentTracer in the presence of noise from normal tasks, we combine execution logs constructed from successful IPI attacks in AgentDyn and InjecAgent with normal execution logs containing 1,800 user requests and 9,000 background tool calls without IPI. In end-to-end experiments, AgentTracer achieves 94.17 percent injection-point accuracy and 93.56 percent path precision. Comparative experiments show that AgentTracer improves injection-point accuracy over existing methods by 18 to 54 percent.
Chinese Translation
大型语言模型(LLM)智能体与外部资源交互以完成复杂的用户任务,这使它们暴露于间接提示注入(IPI)之中,其中恶意指令会将智能体重定向到攻击者意图的任务。由于在现实环境中IPI难以防御,事后追踪对于定位注入源和重建攻击链至关重要。然而,现有追踪方法主要捕获显式的控制流和数据流依赖关系,忽略了由恶意指令驱动的工具调用之间的隐式关系。这些工具调用可能缺少显式依赖关系,并与合法操作交织在一起,使得完整的攻击链重建变得困难。在本文中,我们提出AgentTracer,一个意图感知的追踪框架,它将IPI视为任务意图漂移。AgentTracer恢复工具调用之间的隐式决策依赖关系,以构建一个意图驱动的执行图,该图通过任务意图连接分散的工具调用。它将用户请求与操作知识库相结合,以构建用户意图授权空间,并识别意图漂移的工具调用。从审计下的异常工具调用开始,AgentTracer执行基于目标的剪枝和反向追踪,以重建攻击链并定位注入源和注入点。为了在存在正常任务噪声的情况下评估AgentTracer,我们将由AgentDyn和InjecAgent中成功的IPI攻击构建的执行日志,与包含1,800个用户请求和9,000个无IPI背景工具调用的正常执行日志相结合。在端到端实验中,AgentTracer实现了94.17%的注入点准确率和93.56%的路径精确率。对比实验表明,AgentTracer将注入点准确率相较于现有方法提高了18%至54%。
cs.SE / 105 / 2610.10141
AdaT$^2$: Adaptive Test Transformations for Black-Box Boundary Testing of Conversational Agents
AdaT$^2$:面向对话代理黑盒边界测试的自适应测试变换
large language model
大语言模型相关
Abstract
Conversational agents based on large language models (LLMs) must comply with policies. Each condition in a policy draws a boundary between user requests, and the agent must behave differently on its two sides. We present AdaT$^2$, which extracts statements from the agent's replies in exploratory conversations with an LLM acting as the user, and uses the statements to guide boundary test generation. Each statement describes one condition and the behavior expected when the condition holds. Besides plain tests guided by single statements, AdaT$^2$ writes transformed tests guided by pairs of a statement and a test transformation instruction, such as "omit one required input". The instruction of a pair can move a test to the other side of the statement's boundary or to another boundary. The statements and instructions form far more pairs than a run can try, and many pairs are not applicable. Adaptive pair selection therefore chooses the statement of each pair by novelty and the instruction with the bandit algorithm Bayes-UCB, which learns from whether earlier pairs yielded a test and whether the agent passed it according to an LLM judge. Our benchmark counts the two sides of each boundary separately and distinguishes boundaries explicitly defined by the agent's prompt, tool code, or knowledge base from boundaries that the agent's LLM infers from domain knowledge. On four domains of $τ^3$-bench, 62.7% to 83.3% of AdaT$^2$'s tests are valid boundary tests whose expected behavior is explicitly defined, higher in every domain than AgentEval's (47.7% to 68.1%). Transformed tests add 13 to 46 explicitly defined boundaries that plain tests miss. As regression tests, AdaT$^2$'s test suites detect all eight seeded policy faults in the airline domain and four of eight in the retail domain, and AgentEval's test suites, with fewer than a third as many tests, detect five and two, respectively.
Chinese Translation
基于大语言模型(LLM)的对话代理必须遵守策略。策略中的每个条件都在用户请求之间划定一条边界,而代理必须在边界两侧表现出不同的行为。我们提出 AdaT$^2$,它从由 LLM 扮演用户的探索性对话中代理的回复里提取陈述,并利用这些陈述指导边界测试生成。每个陈述描述一个条件以及该条件成立时预期的行为。除了由单个陈述指导的普通测试外,AdaT$^2$ 还编写由陈述与测试变换指令构成的配对指导的变换测试,例如“省略一个必需输入”。配对中的指令可以将测试移动到该陈述边界的另一侧,或移动到另一条边界。这些陈述和指令形成的配对远多于一次运行所能尝试的数量,而且许多配对并不适用。因此,自适应配对选择根据新颖性选择每个配对的陈述,并使用赌博机算法 Bayes-UCB 选择指令,该算法根据先前的配对是否产生了一个测试,以及根据 LLM 评判器代理是否通过了该测试来进行学习。我们的基准分别统计每条边界的两侧,并区分由代理的提示、工具代码或知识库明确定义的边界,与代理的 LLM 从领域知识推断出的边界。在 $τ^3$-bench 的四个领域上,AdaT$^2$ 的测试中有 62.7% 到 83.3% 是有效的边界测试,其预期行为被明确定义,在每个领域都高于 AgentEval 的(47.7% 到 68.1%)。变换测试增加了 13 到 46 条普通测试遗漏的明确定义的边界。作为回归测试,AdaT$^2$ 的测试套件在航空领域检测出全部八个植入的策略故障,在零售领域检测出八个中的四个;而 AgentEval 的测试套件在测试数量不到其三分之一的情况下,分别检测出五个和两个。
cs.SE / 106 / 2610.10226
Why Software Engineering Is Indispensable in the Age of Coding Agents
为什么软件工程在编码智能体时代不可或缺
large language model
大语言模型相关
Abstract
Can AI make Software Engineering (SE) -- the discipline -- obsolete? And can it make software engineers -- the professionals -- redundant? This paper argues that the rise of capable AI coding agents makes SE and software engineers essential, not obsolete: the missing foundation without which AI-assisted development produces misleadingly plausible, unverifiable, and ultimately untrustworthy software. Three structural properties of large language models (probabilistic generation, agnosticism, and semantic statelessness) create a structural vacuum that no amount of training can eliminate. Filling it requires four knowledge levers: methodological knowledge, domain knowledge, design choices, and process choices. All four must be reified as persistent artifacts, and each requires the software engineer as methodologist, mediator, and custodian.
Chinese Translation
AI 能否使软件工程(SE)——这门学科——变得过时?它又能否使软件工程师——这些专业人员——变得多余?本文认为,能力强大的 AI 编码智能体的兴起,使软件工程与软件工程师变得不可或缺,而非过时:正是这一缺失的基础,若无它,AI 辅助的开发便会产出看似可信实则误导、无法验证、并最终不可信的软件。大语言模型的三个结构性属性(概率化生成、不可知性以及语义无状态性)造成了一个结构性真空,无论多少训练都无法消除。填补这一真空需要四种知识杠杆:方法论知识、领域知识、设计选择与过程选择。这四者都必须被具体化为持久性制品,而每一种都要求软件工程师充当方法论者、中介者与守护者。
cs.SE / 107 / 2610.10374
TaoD2C-Bench: Benchmarking MLLMs for Industrial UI Code Generation Beyond Visual Fidelity
TaoD2C-Bench:超越视觉保真度,对用于工业 UI 代码生成的 MLLMs 进行基准测试
large language model
大语言模型相关
Abstract
A key challenge for multimodal large language models (MLLMs) is moving beyond visual recognition to constraint-aware cross-modal reasoning. This involves combining visual cues with information from other modalities to understand elements' relationships under domain-specific rules. This challenge is acutely evident in industrial design-to-code (D2C), which converts user interface (UI) designs into code and requires MLLMs to connect design images with disorganized layer metadata, infer component and layout implementation requirements, and realize them in code under target-library constraints. However, these capabilities remain insufficiently evaluated in realistic industrial settings. To fill this gap, we present TaoD2C-Bench, a benchmark for evaluating MLLMs' ability to generate UI code that satisfies implementation requirements in industrial applications. The TaoD2C dataset consists of 2,861 production designs from 17 commercial platforms with 97,652 expert annotations across four categories: Component, Group, Alignment, and Position. These annotations distinguish required constraints from permitted implementation choices. TaoD2C-Bench defines three tasks: end-to-end UI code generation, requirement inference, and requirement realization. Evaluating eight MLLMs reveals substantial gaps in generating UI code that satisfies implementation requirements, alongside distinct performance profiles in inference and realization. We further show that MLLMs' visual reconstruction ability does not necessarily imply an ability to generate code that meets these requirements. We release TaoD2C to support research on industrial UI code generation.
Chinese Translation
多模态大语言模型(MLLMs)的一个关键挑战是超越视觉识别,转向约束感知的跨模态推理。这涉及将视觉线索与其他模态的信息相结合,以理解在领域特定规则下元素之间的关系。这一挑战在工业设计到代码(D2C)中尤为明显,该任务将用户界面(UI)设计转换为代码,并要求 MLLMs 将设计图像与杂乱的图层元数据连接起来,推断组件和布局的实现需求,并在目标库约束下用代码实现这些需求。然而,这些能力在真实的工业环境中仍未得到充分评估。为填补这一空白,我们提出了 TaoD2C-Bench,一个用于评估 MLLMs 在工业应用中生成满足实现需求的 UI 代码能力的基准。TaoD2C 数据集包含来自 17 个商业平台的 2,861 个生产设计,以及跨四个类别的 97,652 条专家标注:组件(Component)、组(Group)、对齐(Alignment)和位置(Position)。这些标注区分了必需的约束与允许的实现选择。TaoD2C-Bench 定义了三个任务:端到端 UI 代码生成、需求推断和需求实现。对八个 MLLMs 的评估揭示了在生成满足实现需求的 UI 代码方面存在显著差距,同时在推断和实现方面表现出不同的性能特征。我们进一步表明,MLLMs 的视觉重建能力并不必然意味着其具备生成满足这些需求的代码的能力。我们发布 TaoD2C,以支持工业 UI 代码生成的研究。
cs.CL / 108 / 2610.08994
Phoneme-Guided Initialization for LLM-based Speech Recognition
用于基于LLM的语音识别的音素引导初始化
large language model
大语言模型相关
Abstract
Speech large language models (speech LLMs) perform well on automatic speech recognition (ASR) when sufficient paired speech-text data is available, but their performance degrades in low-resource settings. A cascaded pipeline that performs speech-to-phoneme (S2P) conversion followed by phoneme-to-grapheme (P2G) conversion has been shown to outperform end-to-end speech LLMs in this regime, suggesting that phoneme-mediated processing is beneficial when paired data is scarce. We propose \textit{phoneme-guided initialization}, a simple method that uses this insight within an end-to-end framework: we pre-train the audio encoder on S2P and the LLM on P2G tasks, then connect them and fine-tune the full model end-to-end on the target ASR task. Experiments on Japanese (CSJ), Chinese (AISHELL-1), and two low-resource languages from Common Voice 25.0 (Tatar and Urdu) show that our method matches or outperforms both the cascaded S2P-P2G baseline and the end-to-end model without P2G initialization.
Chinese Translation
语音大语言模型(speech LLMs)在可获得充足的成对语音-文本数据时,在自动语音识别(ASR)上表现良好,但在低资源场景下其性能会下降。一个先执行语音到音素(S2P)转换、随后执行音素到字素(P2G)转换的级联流水线已被证明在这种情形下优于端到端语音LLM,这表明在成对数据稀缺时,以音素为中介的处理是有益的。我们提出 \textit{音素引导初始化},这是一种在端到端框架内利用上述洞见的简单方法:我们在 S2P 任务上预训练音频编码器,并在 P2G 任务上预训练 LLM,然后将它们连接起来,并在目标 ASR 任务上对完整模型进行端到端微调。在日语(CSJ)、中文(AISHELL-1)以及来自 Common Voice 25.0 的两种低资源语言(塔塔尔语和乌尔都语)上的实验表明,我们的方法达到或优于级联 S2P-P2G 基线和未使用 P2G 初始化的端到端模型。
cs.LG / 109 / 2610.09334
VM-ARRAYDPS: Virtual Microphone Augmented Diffusion Posterior Sampling for Unsupervised Blind Speech Separation
VM-ARRAYDPS:用于无监督盲语音分离的虚拟麦克风增强扩散后验采样
diffusion
扩散模型相关
Abstract
Blind Source Separation(BSS) is a fundamental problem in signal processing, aiming to separate multiple source signals from their mixtures without prior knowledge of the sources or the mixing process. Traditional approaches, such as Independent Vector Analysis (IVA) exploits statistical independence of sources. Recently, diffusion-based approaches have emerged as a promising alternative by leveraging powerful generative priors. Among them, ArrayDPS formulates BSS problem as a posterior sampling problem, and utilizes a pretrained speech diffusion model to guide the recovery of clean source signals. A key factor behind its separation capability is the multi-channel consistency (MC) objective, which enforces the estimated source signals to reconstruct the observed microphone mixtures through the estimated acoustic transfer functions. However, the number of microphones in the array is often limited, which constrains the performance of ArrayDPS. To address this issue, we propose VM-ArrayDPS, a novel method that augments the microphone array with virtual microphones with higher-SNR, these microphones can offer extra MC constraints to enhance the separation performance. Experimental results demonstrate that VM-ArrayDPS significantly outperforms ArrayDPS on both 2-speaker and 3-speaker datasets, showcasing the effectiveness of virtual microphone augmentation in improving BSS performance. We also did ablation studies to show the influence of the number of virtual microphones and weight of the MC objective brought by virtual microphones.
Chinese Translation
盲源分离(BSS)是信号处理中的一个基础问题,旨在在没有关于源信号或混合过程的先验知识的情况下,从混合信号中分离出多个源信号。传统方法,例如独立向量分析(IVA),利用源信号的统计独立性。最近,基于扩散的方法通过利用强大的生成先验,已成为一种有前景的替代方案。其中,ArrayDPS 将 BSS 问题形式化为后验采样问题,并利用预训练的语音扩散模型来指导干净源信号的恢复。其分离能力背后的一个关键因素是多通道一致性(MC)目标,该目标强制估计的源信号通过估计的声学传递函数重建观测到的麦克风混合信号。然而,阵列中麦克风的数量通常有限,这限制了 ArrayDPS 的性能。为了解决这个问题,我们提出了 VM-ArrayDPS,这是一种新颖的方法,它用具有更高 SNR 的虚拟麦克风来增强麦克风阵列;这些麦克风可以提供额外的 MC 约束,以提升分离性能。实验结果表明,VM-ArrayDPS 在 2 说话人和 3 说话人数据集上均显著优于 ArrayDPS,展示了虚拟麦克风增强在提升 BSS 性能方面的有效性。我们还进行了消融研究,以展示虚拟麦克风数量以及由虚拟麦克风带来的 MC 目标权重的影响。
cs.AI / 110 / 2610.10320
LLM-Assisted Generation of Transparent, Open-Source Multiphysics Models of Electrochemical Devices
LLM 辅助生成电化学装置的透明、开源多物理场模型
diffusionlarge language model
扩散模型相关
大语言模型相关
Abstract
Multiphysics continuum models are powerful tools for studying electrochemical devices, enabling in silico reactor design and resolution of local pH, potential, and concentration fields that govern device performance but are difficult to measure experimentally. However, constructing such models requires substantial numerical expertise or reliance on proprietary software. Here, we show that frontier large language model agents can remove this implementation burden while keeping the underlying physics under researcher control. Using one-dimensional electrochemical CO2 reduction to CO in a porous gas diffusion electrode as a test case, we develop a machine-readable, human-specified modeling harness containing governing equations, parameters, numerical methods, logical build stages, and human-verifiable checkpoints. From this specification, the agent reproducibly constructs complete multiphysics models in open-source Julia. Independently built models, including fully autonomous agent-built models, agree with an equivalent COMSOL implementation to within 0.7% of the peak CO partial current density, and with one another to within 0.04%. Systematically planted errors demonstrate the importance of explicit specifications for reproducibility and reveal the agent's capabilities and limitations in debugging model physics. This framework establishes a more transparent approach to multiphysics modeling in which physical descriptions and governing equations, rather than specialized code, become the primary inputs for computational model development.
Chinese Translation
多物理场连续介质模型是研究电化学装置的有力工具,能够支持计算机模拟反应器设计,并解析决定装置性能但难以通过实验测量的局部 pH、电势和浓度场。然而,构建此类模型需要深厚的数值专业知识,或依赖专有软件。在此,我们表明,前沿大语言模型智能体能够消除这一实现负担,同时将底层物理置于研究者的控制之下。使用多孔气体扩散电极中的一维电化学 CO2 还原为 CO 作为测试案例,我们开发了一个机器可读、由人指定的建模框架,其中包含控制方程、参数、数值方法、逻辑构建阶段和可由人验证的检查点。根据该规范,智能体可重复地在开源 Julia 中构建完整的多物理场模型。独立构建的模型,包括完全自主由智能体构建的模型,与等效的 COMSOL 实现一致,差异不超过峰值 CO 分电流密度的 0.7%,并且彼此之间一致,差异不超过 0.04%。系统性地植入的错误证明了显式规范对于可重复性的重要性,并揭示了智能体在调试模型物理方面的能力与局限。该框架建立了一种更透明的多物理场建模方法,其中物理描述和控制方程,而非专门代码,成为计算模型开发的主要输入。
cs.LG / 111 / 2610.09045
Learning Transition Kernels of Jump-Diffusion Processes with Conditional Diffusion Models
用条件扩散模型学习跳跃扩散过程的转移核
diffusion
扩散模型相关
Abstract
We study the problem of learning transition kernels for time-homogeneous jump-diffusion processes using conditional diffusion models, with the goal of generating new sample paths from training data consisting of N independent trajectories observed on a high-frequency discrete time grid. On the theoretical side, we establish non-asymptotic bounds for the conditional score estimation error and for the KL divergence between the laws of the true and generated discretely observed paths. On the numerical side, we first evaluate our method on synthetic data to assess the theoretical findings and benchmark its performance against the approach of Gao et al. (2025). We then apply our method to real-world data and investigate its performance on a probabilistic forecasting task.
Chinese Translation
我们研究使用条件扩散模型学习时间齐次跳跃扩散过程的转移核这一问题,目标是从由在高频离散时间网格上观测到的 N 条独立轨迹组成的训练数据中生成新的样本路径。在理论方面,我们为条件得分估计误差,以及真实的离散观测路径与生成的离散观测路径的分布之间的 KL 散度,建立了非渐近界。在数值方面,我们首先在合成数据上评估我们的方法,以检验理论发现,并将其性能与 Gao 等人 (2025) 的方法进行基准比较。随后,我们将我们的方法应用于真实世界数据,并考察其在一个概率预测任务上的性能。
cs.LG / 112 / 2610.09522
Reflected Anchored Langevin Algorithms
反射锚定朗之万算法
diffusion
扩散模型相关
Abstract
First order Langevin algorithms for constrained sampling in machine learning, such as projected Langevin Monte Carlo which are based on discretizations of reflected Langevin dynamics, require differentiable log densities that limits their applicability. This paper introduces reflected anchored Langevin dynamics (RALD), a reflected diffusion that converges to non-differentiable targets on constrained domains. The method uses a smooth anchored reference potential and multiplies the drift and noise covariance of its reflected Langevin dynamics by the same state dependent scaling factor. Its Euler-Maruyama discretization with projection gives reflected anchored Langevin Monte Carlo (RALMC) algorithm. We prove explicit convergence bounds and iteration complexity for RALMC in the 2-Wasserstein distance to the target distribution. Numerical experiments are provided to illustrate the theoretical predictions and the empirical performance of the method.
Chinese Translation
用于机器学习中约束采样的一阶朗之万算法,例如基于反射朗之万动力学离散化的投影朗之万蒙特卡洛,要求对数密度可微,这限制了其适用性。本文引入反射锚定朗之万动力学(RALD),这是一种在约束域上收敛到不可微目标的反射扩散。该方法使用一个光滑的锚定参考势,并将其反射朗之万动力学的漂移和噪声协方差乘以同一个依赖于状态的缩放因子。其带投影的 Euler-Maruyama 离散化给出了反射锚定朗之万蒙特卡洛(RALMC)算法。我们证明了 RALMC 在到目标分布的 2-Wasserstein 距离中的显式收敛界和迭代复杂度。提供了数值实验以说明理论预测和该方法的经验性能。
cs.LG / 113 / 2610.10187
Kinetic Langevin Meets Split Gibbs: Accelerated Posterior Sampling for Imaging Inverse Problems with Diffusion Priors
动能朗之万遇上分裂吉布斯:带扩散先验的成像逆问题的加速后验采样
diffusion
扩散模型相关
Abstract
Split Gibbs sampling (SGS) is a popular framework for posterior sampling in Bayesian imaging inverse problems. It decouples a Gaussian data-fidelity term from a complex prior through an auxiliary variable, so the data variable is updated exactly and only the prior-side conditional is hard to sample. Existing samplers treat this conditional in one of two ways. Plug-and-play SGS runs a multi-step diffusion denoiser at every iteration, which is expensive and lacks non-asymptotic guarantees. Langevin-within-SGS takes cheap overdamped Langevin steps but needs many iterations. We propose RED-KLwSGS, which keeps the exact Gaussian update for the data variable and updates the auxiliary variable with underdamped (kinetic) Langevin diffusions driven by a one-shot denoising score, at the same per-iteration cost as Langevin-within-SGS. We prove non-asymptotic Wasserstein-2 convergence in continuous and discrete time for strongly log-concave priors. We also introduce Joint-RED-KLwSGS, which applies kinetic Langevin diffusions to both variables. Experiments with Denoising diffusion probabilistic models as diffusion priors on FFHQ and ImageNet datasets show faster convergence and high-quality image reconstruction.
Chinese Translation
分裂吉布斯采样(SGS)是贝叶斯成像逆问题中后验采样的一种流行框架。它通过辅助变量将高斯数据保真项与复杂先验解耦,因此数据变量被精确更新,而只有先验侧条件分布难以采样。现有采样器以两种方式之一处理该条件分布。即插即用 SGS 在每次迭代中运行多步扩散去噪器,这代价高昂且缺乏非渐近保证。Langevin-within-SGS 采用廉价的过阻尼朗之万步,但需要大量迭代。我们提出 RED-KLwSGS,它保留数据变量的精确高斯更新,并用由单次去噪得分驱动的欠阻尼(动能)朗之万扩散来更新辅助变量,其每次迭代成本与 Langevin-within-SGS 相同。我们针对强对数凹先验,证明了连续时间和离散时间下的非渐近 Wasserstein-2 收敛。我们还引入 Joint-RED-KLwSGS,它将动能朗之万扩散应用于两个变量。在 FFHQ 和 ImageNet 数据集上使用去噪扩散概率模型作为扩散先验的实验表明,其收敛更快,并能实现高质量图像重建。
cs.LG / 114 / 2610.10190
Universal Local Error and Realized Amplification for the First-Order EDM Predictor
一阶 EDM 预测器的普适局部误差与实际放大
diffusion
扩散模型相关
Abstract
We analyze the first-order deterministic diffusion sampler of Karras et al. (2022), termed EDM, in 2-Wasserstein distance by separating two sources of error: local discretization error and its amplification by subsequent learned steps. We prove that local error admits a universal bound: for any data distribution with finite second moment, the one-step discretization error is quadratic in the step size, with an explicit constant that does not depend on the data distribution. Error propagation, in contrast, depends on the learned network. At high noise levels, we exploit the network parametrization of EDM to derive an explicit contraction criterion. At low noise levels, we measure propagation through the amplification realized on the distributions transported by the sampler; this realized amplification can be arbitrarily smaller than the worst-case Lipschitz constant. This analysis yields an $O(e^{Λ_K}/K)$ global discretization error for $K$ sampling steps, where $Λ_K$ is the low-noise log-amplification. Experiments on a one-dimensional Gaussian mixture show how measured amplification accounts for slower error decay on finite sampling grids. Diagnostics on a pretrained CIFAR-10 model illustrate related stability mechanisms without certifying the global assumptions.
Chinese Translation
我们通过分离两个误差来源:局部离散化误差及其由后续学习步骤带来的放大,在 2-Wasserstein 距离下分析 Karras 等人 (2022) 的一阶确定性扩散采样器,称为 EDM。我们证明局部误差具有一个普适界:对于任何具有有限二阶矩的数据分布,单步离散化误差关于步长是二次的,并带有一个不依赖于数据分布的显式常数。相比之下,误差传播依赖于学习到的网络。在高噪声水平下,我们利用 EDM 的网络参数化来推导一个显式的收缩判据。在低噪声水平下,我们通过采样器所传输的分布上实际实现的放大来度量传播;这种实际放大可以任意小于最坏情况 Lipschitz 常数。该分析给出对于 $K$ 个采样步的 $O(e^{Λ_K}/K)$ 全局离散化误差,其中 $Λ_K$ 是低噪声对数放大。在一维高斯混合模型上的实验表明,实测放大如何解释有限采样网格上较慢的误差衰减。在预训练 CIFAR-10 模型上的诊断分析展示了相关的稳定性机制,但没有证明全局假设。
cs.LG / 115 / 2610.10194
Broadly Applicable Approximate MCMC for Switching Stochastic Differential Equations Using Uniformization and Time-Conditioned Factorized Neural Likelihood Estimation
使用均匀化和时间条件化因子化神经似然估计的广泛适用切换随机微分方程近似 MCMC
diffusion
扩散模型相关
Abstract
Switching stochastic differential equations (SSDEs) describe continuous-time dynamics whose parameters switch according to a latent regime process that follows a continuous-time Markov chain (CTMC). By allowing dynamics to change between regimes, SSDEs represent heterogeneous system behavior and have been applied across diverse fields. However, Bayesian inference for SSDEs remains difficult, and existing SSDE inference methods have limited applicability, with restrictions such as noise-free observations, univariate states, linear drift, or state-independent diffusion. In this study, we propose an approximate Markov chain Monte Carlo sampler for SSDEs using uniformization and factorized neural likelihood estimation (FNLE), a simulation-based inference method. Uniformization provides an exact representation of the CTMC but requires SDE transition densities over arbitrary time intervals. We approximate these densities by training a time-conditioned FNLE model. The resulting sampler is broadly applicable to SSDEs without requiring analytically tractable transition densities. In synthetic-data experiments, our method recovered regime paths and parameters for three SSDE models for which previous methods have limited applicability. We also applied our method to a real dataset and detected a regime transition.
Chinese Translation
切换随机微分方程(SSDEs)描述连续时间动力学,其参数根据遵循连续时间马尔可夫链(CTMC)的潜在机制过程进行切换。通过允许动力学在不同机制之间变化,SSDEs 表示异质系统行为,并已被应用于多个不同领域。然而,SSDEs 的贝叶斯推断仍然困难,并且现有的 SSDE 推断方法适用性有限,存在诸如无噪声观测、单变量状态、线性漂移或状态无关扩散等限制。在本研究中,我们提出了一种用于 SSDEs 的近似马尔可夫链蒙特卡洛采样器,其使用均匀化和因子化神经似然估计(FNLE),这是一种基于模拟的推断方法。均匀化提供了 CTMC 的精确表示,但需要任意时间区间上的 SDE 转移密度。我们通过训练一个时间条件化 FNLE 模型来近似这些密度。所得采样器广泛适用于 SSDEs,而不需要解析可处理的转移密度。在合成数据实验中,我们的方法恢复了三个 SSDE 模型的机制路径和参数,而先前方法对这些模型的适用性有限。我们还将我们的方法应用于一个真实数据集,并检测到了一次机制转换。
人工智能 (cs.AI)
126
cs.AI / 1 / 2610.08923
AdaGuard: Enhancing Safety and Policy Compliance with Reasoning-Enabled LLM-As-A-Judge Guardrails
Abstract
Enterprise generative AI applications require robust safety mechanisms that can accommodate diverse risk postures, evolving policies, and varying latency constraints. Current guardrail solutions often suffer from rigidity, relying on fixed policy sets and offering limited transparency or reasoning flexibility. We present Adaguard, an adaptive LLM-as-a-Judge framework designed to address these challenges through dynamic policy enforcement and adaptive reasoning-budget allocation. Built using supervised fine-tuning (SFT) and reinforcement learning (GRPO), AdaGuard generalizes to user-defined safety and compliance policies at runtime without requiring frequent model updates. A core innovation of our approach is the ability to dynamically infer the complexity of input-policy pairs, allowing the model to switch between high-speed black-box inference and explainable, reasoning-enabled moderation. This flexibility enables developers to balance stringent latency requirements with the need for actionable transparency. This adaptive capability allows AdaGuard to rival other guardrail and frontier models several times its size, while its auto-reasoning mode recovers the accuracy of always-on reasoning at a fraction of the latency
cs.AI / 2 / 2610.08927
Can AI Agents Make Open-Ended Scientific Discovery? Evidence from Station
Abstract
Recent AI systems have made rapid progress in scientific discovery when given well-defined metrics, but whether they can autonomously undertake open-ended scientific discovery remains unclear. We investigate AI's ability to tackle open-ended tasks in Station, an open-world environment in which multiple agents simulate a scientific ecosystem. To tackle challenges specific to open-ended tasks, we propose augmenting Station with two mechanisms: a Supervisor mechanism and periodic Meta Reflection, which encourage persistent exploration even when intermediate metrics are lacking. We construct open-ended tasks from three recent oral papers presented at ICLR. We give agents the main research question studied in each paper while withholding the paper's results and disabling web access. We then measure how many of the original findings-partitioned into individual criteria-agents rediscover. We find that Station rediscovers 62.7% of the criteria on average, compared with 15.4% for Codex Multiagent-v2 and 14.4-20.6% for AI Scientist-v2. Ablation and behavioral analyses indicate that adding the two mechanisms together improves research coverage and continuity. We further evaluate Station on two open-ended tasks without oracle papers and find that some of the discoveries made by the agents closely match discoveries reported by researchers after the knowledge cutoff date. Together, these results indicate that a suitable environment can enable agents to autonomously make meaningful progress in open-ended scientific discovery.
cs.AI / 3 / 2610.09000
How Fragile Is On-Device Language Model Safety? Localizing Safety-Critical Parameters for Sparse Fault Analysis
Abstract
As small language models (SLMs) are increasingly deployed on resource-constrained and on-device platforms, including as components of agentic systems, the integrity of locally stored model parameters becomes an important safety concern. We investigate whether safety-sensitive behavior in LLaMA-2-7B-Chat is concentrated within a sparse subset of parameters, creating a reduced fault surface for targeted analysis. We study two complementary localization methods: low-rank safety-associated subspace analysis and parameter-level safety--utility importance filtering. Both approaches reveal highly non-uniform safety sensitivity across the network, with the MLP down_proj consistently emerging as a prominent safety-sensitive component and o_proj providing a smaller contribution. Using parameter-level localization, modifying only 0.19% of model weights in down_proj yields 53% Basic ASR and 56% GCG ASR, while tinyBenchmarks accuracy remains at 51.6% compared with a 52.2% unmodified baseline. These results motivate targeted fault analysis and selective integrity protection for language models deployed in resource-constrained, on-device, and agentic settings.
cs.AI / 4 / 2610.09002
Learning to Report Unsafe Tasks in a Multi-Agent Game
Abstract
When agents share a reward for completed tasks, reporting unsafe work can reduce the reporter's reward by stopping a task. Audits can make reporting optimal without ensuring that further training teaches a silent team to report. We study this learning problem in a game where any witness can stop a task by reporting. With $k$ witnesses per task sharing a policy and drawing independently, the expected-reward derivative with respect to their shared silence probability counts each task's benefit $k$ times at universal silence. The comparison with universal reporting counts it once. For arbitrary policy groups, we give an audit condition sufficient for exact policy-gradient updates to reach universal reporting and, apart from boundary cases, necessary near universal silence. In a balanced family, the cheapest audits meeting the condition with prescribed positive margins cost exactly $k$ times as much for full sharing as for one policy per role. We train PPO policies on 24 witness graphs from learned silence. Separating co-witnesses reduces unsafe completion by 33.59 percentage points compared with shuffled groups of the same sizes under the same audits (95% graph-bootstrap interval: 21.03-45.13). Only 9 of 48 witness-group runs achieve below 1% unsafe completion while retaining at least 90% legitimate completion. At the same audit budget, a fully shared network meets both thresholds in none of 48 runs with independent action draws and all 48 with a common draw.
cs.AI / 5 / 2610.09008
Whose Memory Is It? Scope-Aware Commit Rules for Long-Term LLM Memory
Abstract
Persistent memory allows an LLM agent to carry experience across conversations, but it also turns a local reasoning mistake into a durable one. During deliberation, an agent may consider a plan, simulate a tool result, report another speaker's belief, and then reject all of them. If memory retains only the resulting sentences, those once-useful possibilities can later return as facts. The record is neither fabricated nor irrelevant; it has simply been detached from the context in which it was valid. We identify this missing context as \emph{discourse ownership}: the world, branch, or speaker that licenses a proposition. Our first finding is counterintuitive. Language models already carry a causally active signal for ownership, yet conventional memory interfaces discard it when they convert reasoning into records. We introduce CASK (Causally Anchored Scoping Keys), a commit rule that preserves this signal so that shared-world facts enter durable memory while provisional content remains available only within its original scope. Our second finding is that the most obvious way to preserve the signal---storing the discovered internal coordinates---is unreliable because equivalent representations need not keep the same coordinates. CASK instead preserves the stable relations that express ownership. Controlled long-conversation conflicts and tool-agent traces show that this design improves memory admission and prevents provisional content from contaminating later answers while complementing runtime provenance. The resulting commit boundary lets agents explore more possibilities without granting every intermediate sentence authority over future behavior.
cs.AI / 6 / 2610.09013
Enabling Dynamic Computation in Looped LMs
Abstract
Looped LMs are parameter efficient and promise dynamic computation (saving memory and FLOPs on easy tokens). However, state-of-the-art open Looped LMs trained with this dynamic computation capability (Ouro models) do not realize it in practice as each loop iteration (depth) requires its own level of KV-cache, necessitating all loop computations. Moreover, Ouro's early-exit prior is enforced on each token equally, which results in static lower-depth like processing of all tokens regardless of difficulty. In this work, we propose a simple "best-available" KV caching strategy that works out-of-the-box, creating a new frontier in the performance vs depth space. Our approach enables up to 30% reduction in FLOPs and KV memory while retaining full-depth performance, showing the true flexibility of Looped LMs. Furthermore, training looped LMs with awareness about this KV caching strategy improves performance and efficiency. Finally, we apply a small but effective fix to the early-exit prior enforcement objective that makes tokens exit at truly heterogeneous depths based on effort. Our findings are validated on Ouro models as well as smaller looped LMs pre-trained from scratch.
cs.AI / 7 / 2610.09016
PAIR: Bridging Perception and Action in Vision-Language-Action Models
Abstract
Vision-language-action (VLA) models map visual observations and language instructions to continuous robot actions. This task requires a transition from representations that describe the scene and instruction to representations that support action generation. Many continuous-action VLAs leave this transition implicit and supervise it mainly through the final action-prediction loss. We introduce PAIR, a framework that learns a shared perception-action representation between these two spaces. During training, a Masked Action Autoencoder encodes expert action chunks into horizon-aligned Action Latent Tokens. A Bridge Module extracts task-relevant features from the current visual-language representations. PAIR aligns these features with the Action Latent Tokens to form Bridge Tokens that preserve task information and capture the structure of expert actions. The Bridge Tokens are then projected into the action-token space and injected into the initial Action Tokens, providing an action-ready starting point for Action Expert refinement. At inference, the autoencoder is removed, and the Bridge Tokens are generated only from the current observation and instruction. Experiments on LIBERO, LIBERO-Plus, and CALVIN ABC-D show gains for the evaluated OpenVLA-OFT and VLA-Adapter models. On LIBERO-Plus, PAIR raises VLA-Adapter's success rate from 59.1% to 64.2%. On CALVIN, it increases VLA-Adapter's average completed sequence length from 4.42 to 4.53. Across seven real-world tasks, PAIR raises OpenVLA-OFT's success rate from 51.4% to 65.0%. Representation analyses show that Bridge Tokens retain task information while making continuous-action information accessible before Action Expert refinement. These results support a shared intermediate representation as a useful interface between perception and action in continuous-action VLAs.
cs.AI / 8 / 2610.09021
Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System
Abstract
An agentic system issues several structurally different kinds of LLM calls. It routes intent, classifies actions, grounds language in a device registry, plans multi-agent pipelines and writes the Python code those pipelines run. The difficulty of these call sites varies by an order of magnitude, yet in practice a single model, chosen for the hardest site, serves all of them. In this work, we evaluate 9 models from 0.8B to a frontier hosted model across the five call sites of a deployed open-source home-automation framework (Wactorz), using its unmodified production prompts and two real Home Assistant installations (280 cases, 2520 scored calls). We find that capability is not ordered the same way at every site, and that larger models are not uniformly better: one 4B model is worse than its 2B sibling at grounded actuation. Paired testing shows the best local model to be statistically indistinguishable from both hosted models at four of five sites. Only code generation separates them, against a small hosted model (p = 0.039) as well as a frontier one (p = 0.002). Aggregate accuracy also hides a safety failure specific to actuation, where small models resolve the accuracy/refusal trade-off in degenerate ways: one model (Gemma4 E2B) actuates on 87.2% of requests for devices the site does not own, while another refuses every request it receives. Routing each site to its best local model reaches 91.8% against 95.4% at no per-call cost. In a live deployment judged by a user, hosting only the two generative sites matches hosting everything (39/43 against 39/43) for 28% of the spend, and the actuation gap the benchmark predicted appears as exactly one case in twenty-six. Benchmark, harness and all records are released at https://github.com/waldiez/slm-callsite-eval.
cs.AI / 9 / 2610.09026
BEACON-SP: Ontology-Grounded GraphRAG Framework for Clinical Suicide Risk Assessment
Abstract
We present BEACON-SP, an ontology-grounded Graph Retrieval-Augmented Generation (GraphRAG) framework for clinician-facing decision support in behavioral health settings such as suicide prevention, where effective assessment requires integrating heterogeneous clinical, behavioral, social, and temporal evidence. BEACON-SP combines patient knowledge graphs with ontology-guided retrieval to support multi-hop reasoning across diagnoses, medications, risk and protective factors, life events, and temporal relationships. The framework is enabled by a comprehensive suicide prevention ontology that integrates the Three-Step Theory, the Integrated Motivational-Volitional Model, and the Suicide Social Determinants of Health Ontology into a unified representation of patient risk factors. We construct ontology-grounded patient knowledge graphs and evaluate BEACON-SP for clinician-facing question answering. Compared with a vector-based retrieval-augmented generation (RAG) baseline on a 1,500-query benchmark spanning 15 clinical categories and 100 patients, BEACON-SP improves completeness, clinical relevance, and evidence grounding under a corrected comparative evaluation protocol, with a small gain on factual accuracy. In paired criterion-level comparisons, GraphRAG is preferred in 76.4% of cases. These results demonstrate the potential of ontology-guided GraphRAG to provide structured, contextualized patient evidence for clinical decision support.
cs.AI / 10 / 2610.09034
Shared-Roadmap Generation and Evaluator for Multi-Agent Path Planning Using Heterogeneous Graph Neural Network
Abstract
Multi-agent path planning (MAPP) in continuous environments often relies on roadmaps to balance safety and search efficiency. However, traditional roadmap generation methods, such as lattice grids or standard sampling-based approaches, frequently face a trade-off between graph density and the likelihood of finding feasible, high-quality solutions. In this paper, we propose a scalable heterogeneous Graph Neural Network (GNN) framework for the automated generation and evaluation of shared multi-agent roadmaps. Our model covers the representation of waypoints, agent locations, and task locations as distinct nodes in a heterogeneous graph, allowing it to reason over global connectivity and inter-agent interactions. By training on occupation density maps aggregated and collected from expert solver trajectories, the GNN learns to identify critical points of interest and prune redundant nodes and edges. This process produces a compact, coordination-aware roadmap that is invariant to task permutations and is reusable for multi-agent pick and delivery tasks. Experimental results demonstrate that our framework can reduce planning effort and can potentially find better solutions, reaching at least 40% reduction in runtime and in graph size for dense roadmaps.
cs.AI / 11 / 2610.09037
When the Governor Becomes the Disturbance: Control-Generated Disturbance and Cost-Aware Backoff in Governed Tool-Using Agents
Abstract
Supervisory governors can interfere with the tool-using agents they regulate. We study this possibility in a controlled file-recovery environment where increases in regulatory intensity trigger experimentally imposed tool failures. A cost-blind governor can turn these failures into persistent blocking that prevents task completion. We compare this governor with a backoff rule that reduces intervention probability using a moving average of known induced events. On a hand-coded stochastic-policy agent, the failure pattern appears under both result replacement and execution of corrupted tool arguments. For the persistent policy, adaptive backoff improves completion relative to a fixed weak governor with approximately matched intervention frequency. A Gemini 2.5 Flash experiment comprising 576 episodes across 6 tasks also shows reduced blocking and improved completion under backoff; among the tested settings, intermediate backoff strength achieves the highest observed aggregate success. These results identify an interaction between intervention cost and persistent action blocking, together with a possible mitigation. The cost mechanisms are imposed and their induced events are directly observable to the backoff rule; applicability beyond this controlled environment remains an empirical question.
cs.AI / 12 / 2610.09057
Epistemic Uncertainty-Aware Defect Detection for Quality Control in Medical Device Manufacturing
Abstract
Objective: We investigate whether accounting for epistemic uncertainty can improve the reliability of automated defect detection in medical device manufacturing. Methods: We consider a machine learning framework that operates on heterogeneous manufacturing and device-report data represented with Knowledge Graphs. To mitigate errors arising from uncertainty in the decision model, we analyze a principled rejection strategy to abstain from predictions whose estimated epistemic uncertainty exceeds a specified threshold. We evaluate the approach using standard synthetic benchmarks and real-world medical device report data. Results: The theoretical results establish the validity of the method characterizing the regimes under which it is expected to be effective. Empirically, the rejection strategy enables explicit control of coverage, that is, the proportion of samples for which the model issues predictions, while improving performance on the retained samples. On 266,170 real-world FDA MAUDE device reports, a 10% abstention rate reduces classification error by 48%, and more aggressive rejection (approximately 70% coverage) yields near-perfect accuracy on the retained samples. On standard synthetic manufacturing benchmarks, abstaining on 9% of the decisions, our approach reduces the risk up to 63% compared with the standard no-abstention approach. Conclusions: Abstaining from predictions with high epistemic uncertainty can provide a practical tool for controlling the reliability of machine learning-based defect detection, especially in high-stakes medical device manufacturing applications. Significance: Uncertainty-aware defect detection may support safer and more reliable quality assurance in medical device manufacturing by identifying cases that require additional inspection rather than issuing potentially harmful predictions.
cs.AI / 13 / 2610.09088
RippleCP: Measuring Counterfactual Checkpoint Advantage in Coding Agents
Abstract
Agent checkpoint systems decide what state is recovery-relevant, how to snapshot it, and whether rollback is admissible. None decides which of the safe boundaries they expose are worth materializing. We formulate this as counterfactual checkpoint advantage, the reduction in future recovery cost obtained by checkpointing a candidate rather than skipping it, and measure it by driving a CP branch and a SKIP branch to the same logical failure and recovering both under matched model, tool, verifier, and stopping conditions. On a frozen pilot of 12 SWE-bench Verified tasks and 106 real recovery branches, checkpointing saves 49.4 s per task, and that figure resolves into two regimes two orders of magnitude apart. The first checkpoint returns 100.0 s on 156.5 s of protected work, a conversion of 0.64; a second one step later returns-1.1 s on 41.8 s, a conversion of -0.03. Recovery is a re-derivation rather than a replay, so preserved work is a poor guide to saved work, and the classical elapsed-work rule misprices the second checkpoint by its full nominal cost. We identify where placement can pay, and set the bar a placement policy must clear.
cs.AI / 14 / 2610.09107
Constraint Tree Exploration for Learning from Language Feedback
Abstract
Natural-language feedback in interactive learning often explains why an action failed by pointing to violated requirements. Misinterpreting this feedback can lead an agent to rule out valid solutions. We study this setting by modeling user intent as latent constraints over an action space and formulating learning from language feedback as pure exploration over feasible regions. We introduce TRACE, an algorithm that organizes candidate constraints in a tree and tests each proposed refinement by generating actions that satisfy it. TRACE commits to the refinement only if the resulting feedback does not contradict it over repeated tests. We distinguish two ways of using the same feedback: (i) falsification, which detects contradictions to the constraint set currently being tested, and (ii) identification, which may additionally name a violated constraint. We prove high-probability coverage bounds with dependence on the candidate class size $H$ for TRACE-Falsification. With reliable identification, TRACE-Identification can replace this dependence by $K/p_{\mathrm{ext}}$, where $K$ is the number of latent constraints and $p_{\mathrm{ext}}$ lower-bounds the probability of extracting a missing true constraint from informative feedback. We evaluate TRACE across six language-feedback tasks. On RecMovie, TRACE-Identification achieves 73% and 86% final-output success under caps of 20 and 60 evaluated outputs, compared with at most 42% and 48% for the evaluated prompting baselines given the same feedback and output caps. Controlled identity-corruption experiments further show greater robustness than direct accumulation when the falsification detector remains reliable.
cs.AI / 15 / 2610.09112
GeoNatureAgent (GNA): A Framework and Benchmark for Pre-Production Evaluation of Tool-Using Agents on Geospatial and Environmental Tasks
Abstract
Before tool-using LLM agents are deployed in environmental and geospatial workflows, teams need evidence that an agent reliably selects the right operations against real APIs. We introduce GeoNatureAgent (GNA), a framework for pre-production evaluation of tool-using agents: a fixed sixteen-tool geospatial interface published as a Model Context Protocol (MCP) server, so the agent under test is the only variable, scored against an identical tool layer, task suite, and deterministic scorer. Its flagship instance is a 103-task benchmark (a 93-task main suite across 18 categories plus a ten-task comparison expansion) evaluated against an open, self-hostable geospatial API serving three environmental indicators across Spain and Portugal. We evaluate nine LLMs under three temperature-1.0 seeds, reporting capability and per-case cost as orthogonal axes. (1) Claude Sonnet 4 achieves the highest capability (61.7% +/- 0.7% on all 103 tasks; 60.8% on the main suite), followed closely by DeepSeek V3.2 (57.9%), while no other model exceeds 53%; (2) the cost-accuracy Pareto frontier is mostly open-weight, with DeepSeek V3.2 offering 93% of Claude's capability at 11.3x lower list-price cost; (3) under strict all-checks scoring the best model sits 24-36 points below the 85-97% reported on general-purpose GIS benchmarks, whereas per-check partial credit for the top four models (86-90%) is comparable, so much of that gap reflects scoring strictness rather than task difficulty alone. The MCP server, evaluation harness, benchmark, and API are publicly available; swapping the tool executors and task suite instantiates an equivalent benchmark for any geospatial domain.
cs.AI / 16 / 2610.09115
From Uncertainty to Action: Learning to Steer LLM Agents
Abstract
Steering an LLM agent means deciding whether to correct it, at which step, and with which mechanism. Uncertainty is often used to decide when to correct an agent, but whether it can guide these decisions remains unclear. We steer agent trajectories separately at every non-terminal step with each of four mechanisms and run each continuation to completion. The resulting stepwise outcome table (SOT) holds about 82,000 counterfactual continuations of 1,864 trajectories from three benchmarks and two agents. It shows that uncertainty can identify failing trajectories, but that no single signal reliably locates the step at which steering helps. We therefore propose VoS (Value of Steering), a trajectory-level monitor, offline or online, that learns from SOT the value of steering at each step and decides where to steer by it. A harm-budgeted trigger decides whether to steer, limiting the fraction of successful trajectories that VoS disturbs. VoS improves on unmodified execution in all 12 settings of benchmark, agent, and offline or online use, by 7.8 points on average, and outperforms the strongest of five existing uncertainty-triggered methods in 11, by 2.9 points on average. Ablations show that training on measured outcomes and a tight harm budget are both essential.
cs.AI / 17 / 2610.09119
Which Buildings Are Artificial Intelligence-Ready? A Measurement-Based Assessment Framework for AI Question Answering and Actuation
Abstract
Agentic artificial intelligence (AI) systems are becoming the interface to buildings, answering questions and controlling operations, but a building's readiness for them has not been systematically assessed. This study proposes a framework to quantify it. First, a building's knowledge graph sets two ceilings. The answerable-readiness ceiling is the share of operational questions its data could answer, and the actuation-readiness ceiling is the share of control actions it exposes. Second, a reference AI agent's accuracy on a fixed set of these questions shows how much of the ceilings is realized. On a simulated office, the agent realizes 0.62 of a 0.64 answerable ceiling, so missing data, not the AI, limit readiness, except in naming a fault's cause. Across 37 public real-building graphs, the median answerable ceiling is 0.16, and in 15 of 45, unlinked sensors lower it. The framework turns "is this building AI-ready?" into an auditable, ranked retrofit question.
cs.AI / 18 / 2610.09127
CADFather: Autonomous CAD Reconstruction through Coordinated Tool Use
Abstract
Reconstructing an editable CAD model from a 3D shape remains a challenging engineering task. Existing methods can propose CAD operations, but no single source of proposals works equally well across different part geometries and stages of reconstruction. We introduce CADFather, an autonomous agentic system that coordinates complementary tools to recover parametric CAD programs from 3D meshes. A vision-language assistant inspects renders of the target and intermediate reconstructions, then decides which candidate CAD programs to extend, which tools to invoke, how many proposals to generate, and when to finish. Learned and algorithmic tools propose CAD operations, while numerical optimization refines the parameters of existing programs. Proposed or refined programs are executed and evaluated to provide feedback for subsequent decisions. The agent maintains alternative candidate programs for each target part and preserves the best valid result throughout reconstruction. CADFather uses pretrained generation and assistant models without additional training. We evaluate reconstruction quality and execution validity on the full DeepCAD, Fusion360, and MCB test sets, as well as on CADENA-Bench, CADBench, and BenchCAD. We additionally analyze computational cost and the trade-off between cost and reconstruction quality.
cs.AI / 19 / 2610.09139
Breaking the Space Barrier and its Application to Language Model Inference
Abstract
Language models are more and more often asked for structured output: JSON that follows a schema, or a tool call with typed arguments. A small machine, an automaton, enforces the format by forbidding the tokens that would break it. We observe that this machine has a rare property: from any of its states, each token leads along exactly one path. Graphs in which only a few paths join any two points are a classical object of complexity theory, and our theoretical result settles an open question about them: one can decide whether such a graph connects two points while verifying that it really has few paths, with very little memory. Precisely, the problem lies in the classes ReachUL, LOGDCFL, C=L and SC2, and needs only O(log2 n/ log log n) space, below the classical O(log2 n) of Savitch's theorem. The constructions behind the proofs become an inference engine: text the format forces is written without running the model, the mask is recomputed on the GPU without any table, recursive formats use a small stack, every output stays valid under a token limit, and independent fields are decoded in parallel and verified. On one 16 GB Apple M2 Pro with Qwen3.5-2B and 4B, against MLX with llguidance, the standard setup for this hardware, schema-constrained extraction finishes 1.2- 1.3x sooner with the same answers, a grammar costs 3 MB instead of up to 1.5 GB, one server holds sixteen grammars where tables run out of memory, and sixteen tool-calling agents finish 2.5x sooner.
cs.AI / 20 / 2610.09142
Finding Blind Spots in AppWorld and WorkArena Task Verifiers
Abstract
Execution-based task verifiers decide whether an agent succeeded. We audit shipped AppWorld and WorkArena verifiers with source-informed mutation tests. The main audit never modifies a shipped checker. In AppWorld, duplicating a non-idempotent write creates an extra record while preserving every checked field value. The verifier accepts all three task variants from two of five eligible generators: 6/15 constructed effects. A cardinality patch applied to checker copies after the census makes all six cells fail while preserving valid controls. In WorkArena, we prospectively rerun 23 extra-field candidates selected for earlier checker-PASS outcomes. Independent Table API readback confirms nondefault persisted values in 21, while all 23 receive PASS. Two requested strings are aliases of stored defaults. The 21 confirmed wrong effects span three form templates. These selected cases confirm wrong effects under the audit's protocol; they do not estimate a population rate. No other construction produces an independently confirmed false accept. Other checker-PASS cases are effect-correct degeneracies. We report zero-PASS families separately because retained evidence differs. In fixed intent-swap grids, the checkers return no PASS on 2,689 off-diagonal executions. This is a rejection census: 57 WorkArena cells use session-scoped evidence; the other 2,632 lack classified rejection causes and independent target ground truth. Each increment is specified before its own cells are scored. A supplement accompanies the OpenReview submission with the construction grammar, evidence, content-bound stage lineage and count reproducer.
cs.AI / 21 / 2610.09144
DIVA: Dual-Space Intent-Aware Visual Attenuation for Vision-Language-Action Policies
Abstract
Vision-language-action (VLA) policies typically feed dense visual patch tokens into a language-action backbone, preserving scene context but offering no explicit mechanism to regulate how strongly different visual tokens influence policy computation. We introduce DIVA, a Dual-Space Intent-Aware Visual Attenuation module with an anchor-then-attenuate design. DIVA combines high-level task intent with low-level visual evidence to estimate patch-wise relevance anchors, then applies them in two complementary spaces: it reweights projected visual tokens before backbone entry and persistently attenuates low-relevance visual states within the backbone. DIVA preserves the full visual token sequence and requires no external grounding supervision. On LIBERO, DIVA improves OpenVLA-OFT from 96.6% to 98.0% average success and raises its zero-shot LIBERO-Plus score from 69.6 to 72.6. Real-world experiments further show consistent gains under task-irrelevant visual perturbations, supporting the robustness of intent-aware visual attenuation beyond simulation.
cs.AI / 22 / 2610.09159
SpecGuard: Proving a Task Is Broken Before the Agent Cheats
Abstract
As autonomous coding agents get increasingly deployed, the risk that accidental or adversarially injected misspecifications in tasks lead to dangerous agent behavior is critical to address. Prior work has shown that agents given such tasks rarely flag the conflict and instead cheat, editing tests or hard-coding expected outputs, and the actions taken to cheat can cause real damage, such as deleting a security defense to make a corrupted test pass. It remains unclear whether such conflicts can be established with independently verifiable evidence before the agent acts. We present SpecGuard, which detects and formally certifies these conflicts between task intent and tests. Given only the task description and codebase, SpecGuard autoformalizes the intended behaviour into a Lean 4 specification. The tests are formalized independently, and the Lean kernel checks whether any implementation could satisfy both formalizations, producing a machine-checked certificate when none can. On conflicted SWE-bench tasks, SpecGuard detects up to 72.8% of conflicts and formally certifies up to 51.1%, with a nearly five-fold lower conflict miss rate than model-based judgment. SpecGuard provides a pre-execution safety check that identifies reward-hacking opportunities through formal certification of task-level conflicts, before any agent behavior is observed. Our code is available at https://github.com/prmbiy/specguard.
cs.AI / 23 / 2610.09164
Training Language Models To Be Coherent Decision-Makers
Abstract
Reliable decision-making requires more than accurate prediction: a model must preserve its beliefs, apply the relevant utilities, and recognize when the information needed to justify an action is missing. We study whether language models can learn this decision procedure from supervised fine-tuning and generalize it across domains and differing natural-language expressions of the decision challenge. Across 20 datasets, we explore challenges of belief instability and decision-making errors by first eliciting probabilities of outcomes and then varying only the utilities and the framing of the decision problems, while holding the evidence fixed. We train models to preserve elicited beliefs while selecting the action that maximizes expected utility, and evaluate transfer to unseen application domains, held-out framings, and different classes of payoff structures. We further introduce incomplete-information settings in which required utilities are withheld and replaced with irrelevant text, testing whether models can distinguish missing decision-relevant information from merely additional context. We find that targeted fine-tuning substantially improves coherent decision-making and that in many situations, learning transfers across domains and framings to situations unobserved during training. Further, models trained for decidability learn to identify when action cannot be justified based on missing information. Finally, we show the value of a routed system that considers separately the recognition of decision completeness and utility-sensitive decision execution.
cs.AI / 24 / 2610.09188
From Probabilities to Decisions: Search and Multi-Teacher Distillation with Jev
Abstract
Probability-only models, which TypeSafe calls System One models, return calibrated probabilities for fixed choices in milliseconds and generate no text. We study one such model, Jev, through two tasks that require decisions under tight constraints. In bullet chess, a bot that places Jev's judgment inside Stockfish search alongside an opening book and endgame tablebases climbs above a 2200 Lichess bullet rating against other bots. Live model calls are too slow for search, so we distill pairwise judgments into a compact evaluator that runs at every position. We then ask how best to spend a fixed labeling budget when an LLM, Qwen3-32B, is available as a second teacher. In chess, averaging both judges' labels beats spending the whole budget on Qwen alone by 9.6 Elo (95% interval 4.3 to 14.9), and the gain replicates on fresh openings; a second answer from the same judge is no substitute, and Jev is the strongest partner for Qwen among the models tested. In passage reranking, Jev's labels alone train a reranker that scores as high as Qwen's, from 21 minutes of API calls instead of 5.1 GPU-hours, and adding Qwen gains at most a few thousandths in ranking quality. Search supplies the lookahead, distillation makes the judgment cheap enough to use at every position, and an LLM partner pays off in chess.
cs.AI / 25 / 2610.09193
Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver
Abstract
MemoryAgentBench's Conflict Resolution split is read as measuring "selective forgetting". We execute the benchmark's own rule - the newest statement about a fact wins - as a zero-learning resolver frozen on one of the four fact lists. Under the official metric the rule answers 80.25% of the questions (74.5% on the three held-out lists). Of the rest, 67 items have a released gold that the last-write graph cannot reach but overwritten statements would ("The capital of India is New Delhi." superseded by "The capital of India is Grosseto."; gold New Delhi); such items are a third of the multi-hop questions at 262K. Two long-context models and our pre-registered approximate re-implementation of the benchmark's BM25 agent, one retained run per item and outcomes only, score 84.7%, 82.6% and 41.6% on the items the rule solves against 10.4%, 11.9% and 6.0% on those 67. The failures are a reachability split plus a small parser-scope residual; the per-item split, not the aggregate, is the unit at which a score here can be read.
cs.AI / 26 / 2610.09197
FreeEvolve: Learning to Evolve Beyond Fixed Loops
Abstract
Agent evolvers automate the design of the prompts, skills and workflows around language model agents, yet the optimization process they follow is still designed by hand: a fixed search loop decides how candidates are evaluated, which are kept and when the search stops. We propose FREEEVOLVE, which automates this process as well. An environment specifies the goal, target agent, evaluator, data and resource limits; within these limits, the evolver itself decides what to test, how much evidence to collect, which candidates to pursue and when to stop. These decisions follow an editable evolution skill, which we improve through meta-evolution by scoring each candidate skill on the fresh target agent it produces. The optimization process thus becomes a capability learned from experience rather than a loop engineered in advance. On tau3-bench, ARC-AGI-2, ARC-AGI-3 and Terminal-Bench 2.1, FREEEVOLVE controls the evolution campaign by itself, yet improves the primary held-out metric by 13.6 points on average and matches or exceeds hand-designed evolvers. The learned process keeps improving with experience: meta-evolved skills add 6.9 points over the seed skill on fresh target agents, demonstrating transferability across environments.
cs.AI / 27 / 2610.09202
Few Bits, One Law: Toward W2A4KV2
Abstract
Extreme low-bit LLM compression is most challenging when weights, activations, and KV caches are quantized together: their distributions differ, and quantization errors interact throughout the network. We introduce CanonQ, a unified quantization-aware training framework that addresses these challenges by separating source canonicalization from task-aware adaptation. Fixed rotations and energy normalization map heterogeneous tensor sources to canonical coordinates, enabling frozen Gaussian-reference codebooks to be reused across layers and models. Joint training then adapts the network to the coupled errors of weight, activation, and cache quantization within a common scalar/vector interface. We bound frozen-codebook transfer error and local task loss, and derive an exact normalization-aware straight-through Jacobian that links quantization distortion to gradient bias. The strongest gains arise under joint W2A4KV2 compression: across LLaMA3-1B/3B/8B, CanonQ-Omni achieves up to 14.28x lower WikiText-2 perplexity and up to 57.9% higher mean zero-shot accuracy than prior state-of-the-art and representative quantization baselines. The benefits extend to Qwen3-1.7B, code generation, and mathematical reasoning: on instruction-tuned MobileLLM-Pro-1B at W2A16KV16, CanonQ achieves relative improvements of 41.7% in HumanEval pass@1 and 39.1% in GSM8K exact match over the strongest evaluated quantization baseline.
cs.AI / 28 / 2610.09215
AGAR: a reinforcement learning substrate for LLM program evolution
Abstract
Given a task and an evaluator, a language model can rewrite a candidate program while a search loop decides which rewrites survive, offering a practical route to algorithm discovery. But that loop is governed by five constants set by hand: which parent to select, how hard to mutate, how to keep diversity, what to remember, and a scalar score that never says which part of the program earned it. Reinforcement learning already has an estimator for each. The obstacle is that program evolution is not usually written down as a decision process. We formalize it as a Markov decision process whose action is the modular prefix the model is conditioned on, rather than the program it emits. Credit assignment, value estimation, adaptive exploration, and experience memory can then attach to distinct components. AGAR (Algorithm Generation As RL) provides the resulting substrate: any estimator can be replaced or switched off without changing the controller, making the transfer auditable one mechanism at a time, with no gradient training of the backend model. Across 19 tasks, two backends, and three seeds under one harness, AGAR improves on the stronger of two published baselines on most tasks, with gains concentrated in the competitive-programming family. The formalization also yields a checkable reading of prior work: these systems are implicitly zero-discount, not by choice, but because fitness is exogenous to an individual rather than a return over successors, leaving a discount factor nothing to act on.
cs.AI / 29 / 2610.09218
RLDISCOVER: LLM-driven co-evolution of reinforcement learning algorithms
Abstract
LLM-guided program evolution has enabled discoveries in mathematics and computational optimization, raising the prospect of reinforcement learning (RL) algorithms that self-evolve to improve how agents learn. However, realizing this prospect faces two obstacles. Joint search over coupled algorithmic components is difficult to scale: simultaneous changes can disrupt learning, while isolated changes overlook their dependencies. Evaluating candidate algorithms also requires costly training, with fitness remaining uncertain across random seeds. We introduce RLDiscover, a framework for the self-evolution of model-free deep RL algorithms. Progressive Co-Evolution advances from targeted component edits to joint evolution, while Progressive Probabilistic Evaluation balances search breadth and evaluation fidelity through staged training and repeated evaluation. Experiments across SAC, PPO, and DQN on four benchmark suites show substantial improvements in mean return, with per-family median gains of 32%-84% and a peak return ratio of approximately 363x over a near-zero baseline. These gains include transitions from failed learning to successful task completion, and improvements persist when evolution starts from stronger open-source implementations. On measured SAC locomotion runs, evaluation uses approximately one-fifteenth the estimated compute required to fully evaluate the same candidate pool. Remarkably, independent searches repeatedly discover interpretable combinations of adaptive robust losses, progress-dependent value targets, and running statistics, with selected programs transferring to unseen tasks. These findings point toward a broader role for self-evolution in AI: discovering interpretable algorithms that improve how agents learn.
cs.AI / 30 / 2610.09219
CM-DPO: Constraint-Margin Direct Preference Optimization for LLM Planning
Abstract
Direct Preference Optimization (DPO) treats all constraint violations equally: a $1 budget overshoot and a $1,000 overshoot induce the same training signal. It is also susceptible to length and style bias when preference pairs come from different model families. We introduce Constraint-Margin DPO (CM-DPO), which replaces DPO's binary preference signal with a continuous margin derived from a deterministic symbolic verifier and scaled by violation severity. Hard and soft constraints are separated through a lexicographic objective, ensuring hard constraints are never traded off against preferences. To supply CM-DPO with bias-reduced training pairs, we generate preference data through procedurally generated constraint profiles (DCCG) and minimal-edit distillation from a reasoning teacher (RT-MED), within a framework we call SynPlan-R. On TravelPlanner, NaturalPlan, and out-of-distribution PlanBench, an 8B model fine-tuned with CM-DPO achieves 89.2% pass rate and 93.4% solve rate, matching multi-agent systems at 13x lower latency while outperforming GPT-4o on unseen Blocksworld by 9.2 points.
cs.AI / 31 / 2610.09237
Trajectory Abstraction for the Science of Language Agent Behavior
Abstract
Scientific studies of language agents need behavioral variables that support hypotheses across tasks and models. We formulate this research problem as learning and testing a hierarchy of trajectory abstractions. A concrete recursive procedure first measures role- and phase-indexed events, proposes temporally constrained relations, and tests their stability across conditions. It then constructs episode-level motif variables from selected relations and repeats the analysis on those variables. Explicit measurement functions connect every abstraction level to the original trajectories. Observations and randomized protocol experiments assess the resulting hypotheses, while comparisons between intervention realizations determine whether an abstraction should be retained, refined, or restricted. We derive a finite-depth bound for accepted reductions, identify protocol effects on fixed abstractions, and characterize realization disagreement and composition of abstraction error. A finite-sample test makes projected intervention consistency operational, and constructed examples illustrate motif construction and abstraction refinement. The formulation distinguishes this experimental approach from semantic taxonomies, qualitative theory induction, and behavior-model recovery. It specifies a proposed research procedure for discovering generalizable behavioral hypotheses, with literature-relative novelty assessed separately from model-relative surprise.
cs.AI / 32 / 2610.09239
The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules
Abstract
Self-improving LLM systems propose changes to themselves and keep those that score better on a small evaluation set. We treat this keep-if-better step as selection under measurement noise, model the correlated errors of the candidates in a single decision, and study empirically what happens when the evaluation set is reused. In runs where Qwen models rewrite their own instructions and every candidate is also scored on 600 held-out items, most proposals after the first are harmful, and the model gives the size of the winner's curse of a generation's best candidate. With a prior from a separate pilot, it matches the average overstatement of first-generation commits in native loops, though not setting by setting. In a pre-registered study, the final selection-set score of greedy loops exceeded held-out accuracy by 13 to 20 points with 16 selection items and by 1 to 5 points with 256. Held-out gains grew with the selection set on TREC but not on GSM8K, and the tested acceptance rules did not beat greedy acceptance over whole runs. Gains measured on the selection set also exceeded held-out gains when a current model refined a competent instruction, and in the validation scores of GEPA and MIPROv2. Scoring the starting and the current instruction on 64 items never used for selection removes the average bias of a loop's reported gain, but single estimates remain off by about 6 points. Self-improvement studies should report held-out gains with their uncertainty.
cs.AI / 33 / 2610.09294
RT-Safe: Benchmarking Agent Safety in Real-Time Embodied Environment
Abstract
Rapid progress in AI agents has brought growing attention to agent safety, with extensive evaluation focused on digital environments. As agents move into the physical world, embodied safety becomes increasingly important: failures can cause human injury and costly hardware damage. Beyond selecting safe actions, embodied agents must also operate under real-time constraints: the physical world does not pause while an agent reasons. As pedestrians move and vehicles approach during inference, an action that appears safe at observation time may become unsafe before execution. Real-time embodied safety therefore depends on both decision quality and decision latency. We introduce RT-SAFE, a simulated urban benchmark for evaluating embodied-agent safety under real-time constraints. RT-SAFE combines navigation tasks with moving actors, environmental hazards, and traffic rules, while allowing the world to evolve throughout inference and action execution. Across eight VLMs, agents achieve high task completion yet almost never complete safely: in the hardest setting, only 0.7% of episodes finish without a safety event. More strikingly, matched static and real-time evaluations yield task completion rates of 91.3% and 94.1%, respectively, while real-time execution increases collisions by $12.3\times$. These results reveal that standard task success can mask substantial safety failures, and that decision latency itself can become a source of physical risk. Finally, we show that RT-SAFE can support offline RL training and substantially reduce collision rates while achieving strong task completion.
cs.AI / 34 / 2610.09296
The AI Evaluation Ecosystem
Abstract
AI evaluation shapes the decisions of model providers, users, funders, and regulators. We argue that designing valid benchmarks requires contextualizing design choices in the dynamics of this ecosystem of actors. We develop a simulation architecture that combines rule-based market dynamics with LLM-driven strategic actors, building on advances in Generative Agent-Based Modeling (GABM). We model benchmarks, consumer needs, and provider capabilities as vectors over a six-dimensional capability space (reasoning, coding, knowledge, safety, communication, agentic), with structural information partitions across actors. As a case study, we apply this stylized simulation to explore benchmark holdout design. We find that moving from public benchmarks to private holdout benchmarks shrinks the gap between benchmark scores and user satisfaction on most benchmarks but widens it on a few, depending on where holdout weights shift scoring credit. We stress-test our findings at both the instrument and case-study level, drawing on the V&V framework of Sargent (2013) and GABM-specific evidence criteria. Beyond holdout design, our simulation is a hypothesis-generating sandbox for studying how evaluator and policy choices, in turn, reshape the ecosystem.
cs.AI / 35 / 2610.09335
SearchWorld: Spatial Value-Grounded Imagination for UAV Object Search via World Models
Abstract
Autonomous unmanned aerial vehicle (UAV) object search involves a closed loop of perception, decision-making, and action under partial observability. Urban environments pose several challenges: large search areas and narrow egocentric views limit coverage, dense 3D geometry constrains safe motion, and open-world instructions require identifying a specific target among distractors. Many existing methods mitigate partial observability through explicit maps or memory representations, yet remain largely reactive, reasoning over past observations without explicitly predicting future states. World models enable prospective reasoning through imagined rollouts. However, image-generating world models can incur high inference latency, while spatially grounded planning remains challenging for latent world models. We propose SearchWorld, a recurrent state-space world model that connects explicit spatial memory with value-guided imagination. The model maintains BEV exploration and obstacle memory and decodes a task-aware spatial value layer to guide search. A cognition-action network uses this learned spatial value prior to improve the policy through imagined rollouts, without training a separate scalar critic. Training progresses from world-model learning to expert imitation and imagination-based exploration refinement. On UAV-ON, SearchWorld improves the success rate to 23.8% (19.5% for the strongest published agent) and raises oracle success to 35.5%, while remaining robust on unseen scenes (19.9% success rate). By grounding imagination in explicit spatial representations, SearchWorld enables UAV agents to plan prospectively rather than react.
cs.AI / 36 / 2610.09348
Relevance Is Not Sufficiency: What Actually Closes the Evidence Gap in Long-Term Memory QA
Abstract
LLM agents that interact with a user across many sessions accumulate histories that exceed their context window, so they store past interactions in an external memory and answer each question from a small set of retrieved records. Existing memory systems rank records by lexical or embedding relevance, yet the top-ranked memories can each be relevant while jointly omitting a complementary fact that the answer requires, especially for multi-session and temporal questions. Drawing on the distinction between relevance and sufficiency in legal evidence scholarship, we recast memory retrieval as constructing a sufficient memory set. To operationalize this view, we introduce a blinded LLM judgment over the retrieved set, together with Gold Hit and Turn Hit as evidence-coverage proxies. We then propose Budgeted Flat Reconstruction (BFR), which builds sufficient sets over a fixed flat memory store in two stages. Specifically, we first apply Formal Concept Analysis for Memory Selection (FCA-MS) to decompose the question into information requirements and select a compact candidate subset that jointly covers them. Then, we repeatedly acquire unseen records through deeper text search or complementary entity and session views, stopping when the budget is exhausted. Experiments on LoCoMo and LongMemEval-S show that BFR outperforms same-store adaptations of recent agent-memory systems in both answer quality and evidence coverage. Specifically, on LongMemEval-S it raises judged accuracy from 72.4% to 82.2% and Turn Hit to 91.4%.
cs.AI / 37 / 2610.09360
TopoGraphRAG-Bench: Evaluating Multimodal GraphRAG on Layout-Grounded Evidence Reasoning
Abstract
Real-world documents distribute evidence across text, tables, figures, and captions within complex page layouts. Answering complex questions over such documents therefore requires more than retrieving relevant passages: systems must recover the evidence topology that connects heterogeneous evidence units. Existing GraphRAG evaluations remain largely text-centered, while multimodal document RAG benchmarks assess cross-modal retrieval and generation without directly evaluating recovery of the intended evidence topology. We introduce TOPOGRAPHRAG-BENCH, a layout-grounded benchmark for multimodal evidence reasoning in GraphRAG, comprising 2,024 questions over 201 long, visually rich documents. Questions are constructed bottom-up from text, figure, and table evidence units under three controlled topologies: single-hop retrieval, bridge-chain reasoning, and multi-source synthesis. To ensure that questions preserve their intended structure, we apply counterfactual validation for shortcut resistance, modality necessity, and evidence necessity. We evaluate text-only GraphRAG, page-level visual retrieval, and multimodal GraphRAG systems using retrieval, generation, and topology-aware reasoning metrics. Multimodal GraphRAG systems achieve the strongest overall performance, but still fail when visual-textual evidence alignment or multi-unit composition is incomplete. Text-only GraphRAG struggles when key dependencies are grounded in figures or tables, while page-level visual retrieval lacks the fine-grained structure needed for topology recovery. These findings motivate GraphRAG systems that move beyond text-derived entity relation graphs to explicitly model document layouts, cross-modal evidence alignment, and the reasoning roles of evidence units. Code and data are available at https://richardlrc.github.io/TopoGraphRAG-Bench/.
cs.AI / 38 / 2610.09374
DUDA-Bench: Benchmarking LLM Agents on Multimodal Data-Driven Urban Diagnosis
Abstract
Urban diagnosis integrates heterogeneous observations to identify urban problems, localize affected areas, and investigate contributing factors, informing evidence-based urban planning and management. However, its reliance on labor-intensive, case-specific expert workflows limits scalability and reuse, motivating the exploration of agent-based execution. To evaluate this capability, we introduce DUDA-Bench, a hierarchical and interactive benchmark that formalizes data-driven urban diagnosis as a multi-stage agent workflow. It comprises 86 atomic and 22 workflow tasks spanning four analytical stages, grounded in multimodal data from 12 cities covering five urban problem types. Evaluations of seven backbone models and five agent systems reveal a substantial gap between isolated analytical competence and end-to-end diagnosis, with system benefits varying across backbones. Trajectory analysis shows that unresolved evidence gaps propagate across stages, while successful recovery involves revising assumptions and actions using feedback. These findings highlight limitations in coordinating analytical capabilities across stages, particularly adaptive planning, evidence integration, and verification. More broadly, DUDA-Bench provides a framework for translating expert analytical workflows into hierarchical agent tasks and process-aware evaluation, supporting systematic assessment of end-to-end analytical capabilities.
cs.AI / 39 / 2610.09377
From Plausible Hierarchies to Useful Taxonomies: Evaluating Agentic Harnesses on Customer Feedback
Abstract
Taxonomies are the symbolic representations through which AI systems organize evidence, aggregate patterns, and answer questions over large document collections. Over customer feedback, the category tree decides how every record is counted and routed, which problems get seen, and which team owns them. Agentic harnesses now make it easy to generate a plausible-looking hierarchy, and such trees are checked today with generic, individually scoped checks: each name fits its description, sits under the right parent, and stays distinct from its siblings. We ask a more operational question: when is a generated hierarchy actually useful as a production taxonomy? We build six taxonomies over two proprietary feedback corpora (1,940 and 5,000 records): for each corpus, a production reference and two repeated runs of the same harness under identical inputs. All six pass every generic naming and structure check, and a deeper product-coverage check even prefers the generated trees. Yet in every generated tree at least 97.7% of leaf names merely restate an ancestor's name (13.9% and 2.9% in the references), and in one, three of every four records fall under multiple top-level categories. We introduce two families of whole-tree metrics: structural discriminators test whether a tree's shape was learned from the data or imposed by its generator; team partitionability tests whether branches split feedback into groups teams can own. Trees the generic checks rate as equally correct differ by 27 percentage points in cross-branch leakage, and only one of the two beats a random split. Surface plausibility is an insufficient measure of taxonomy quality: evaluation must measure the whole tree as well as each node.
cs.AI / 40 / 2610.09401
Let the Library Speak: Self-Advertised Method Selection for Formal Proving
Abstract
LLM-based formal provers can retrieve relevant lemmas and prior proofs, but relevance alone does not say whether a mathematical method can be used on the current theorem. A method has prerequisites, a target, an intended action, and obligations that its use leaves to prove. Methods that look equally related to a theorem may therefore differ substantially in whether they offer a plausible next step. We formulate this as an applicability-aware method-selection problem and introduce self-advertisement: before candidates are ranked, a model generates a problem-specific proposal for each one, stating what part of the goal it targets, what action it would take, and what conditions that action requires. We organize 82 reusable methods from Putnam 2000-2014 as Method Contracts, which pair applicability descriptions with Mathlib anchors, a checked example or scaffold, and expected proof obligations. A single batched call elicits proposals across the library; vague or unsupported proposals are demoted, yielding a ranked shortlist accompanied by inspectable claims about each candidate's use. We analyze when similarity-based representations cannot distinguish methods with different applicability, how errors in applicability estimates affect shortlist quality, and what a checked scaffold guarantees under its stated assumptions. Against lexical, embedding, and embedding-plus-LLM reranking baselines, self-advertisement achieves 95.0% hit@5 on Putnam 2015-2025, compared with 84.2% for the strongest reranker. On IMO ProofBench, it achieves 91.7% compared with 88.3%. These results indicate improved coverage of annotated methods in the retrieved shortlists, particularly on Putnam.
cs.AI / 41 / 2610.09426
RSI-Forge: From Research Papers to Environments for Recursive Self-Improvement
Abstract
Environments are the foundation of recursive self-improvement: they provide the problems agents work on and the feedback used to evaluate progress. Yet constructing challenging research environments with reliable evaluation still depends on domain experts, limiting their scale and disciplinary coverage. We introduce RSI-Forge, a multi-agent pipeline that turns published papers into executable environments for self-improvement. Three agents coordinate construction, reproduction, and review to produce tasks with automated evaluators; each paper's method is independently reimplemented to establish a baseline score. We present 210 environments across 18 fields, including 90 reviewed by independent human domain experts. Both experts and agent judges give high ratings to the potential for improving the provided starting solutions and the evaluators' ability to distinguish solution quality, whereas experts are more critical of shortcut resistance, faithfulness to the source paper, and whether a single idea can exhaust a task. To validate their use for repeated improvement, we evaluate four models over 3 successive attempts on 120 environments, with each attempt inheriting prior code and notes while model weights remain fixed. At least one model improves after the first attempt in 84% of environments. Models also outperform the reproduced paper methods in 68 of the 120 environments, demonstrating room for gains beyond these baselines. Transcript analysis identifies work beyond parameter tuning in 95% of these successful attempts. Analysis of the resulting trajectories shows that models scoring lower on these tasks explore less, more often accept gains smaller than the reported standard error, and rely more heavily on tuning to the development set. RSI-Forge provides a scalable approach to constructing research environments for training and evaluating self-improving agents.
cs.AI / 42 / 2610.09466
LLM-Enabled UAV Dispatch: A System-Level Survey and Taxonomy
Abstract
Unmanned aerial vehicle (UAV) dispatch is beginning to move beyond isolated path planning and optimization-driven resource allocation toward system-level coordination supported by semantic reasoning and LLM-based interfaces. This survey provides a unified characterization of LLM-enabled UAV dispatch systems that bridges semantic intent, symbolic decision-making, and physical UAV execution. Rather than treating LLMs as standalone add-ons, we conceptualize them as a cross-layer semantic orchestration layer connecting human instructions, external solvers, and distributed control modules. We organize the literature into four representative dispatch paradigms: pipeline dispatch, global assignment dispatch, decentralized agentic dispatch, and divide-and-conquer dispatch. For each paradigm, we analyze its decision logic, system structure, control flow, representative methods, and potential LLM roles. We further examine how LLMs support semantic parsing, retrieval-grounded planning, solver orchestration, local agent reasoning, multi-agent coordination, safety assessment, and human-facing explanation. We discuss the implications of these paradigms for scalability, robustness, coordination burden, and verification requirements, and identify open challenges including latency-aware reasoning, grounding reliability, physical feasibility guarantees, edge deployment, privacy protection, and distributed consistency. This survey provides a system-level taxonomy and design perspective for integrating LLMs into safety-critical UAV dispatch systems.
cs.AI / 43 / 2610.09484
MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
Abstract
Training and evaluating interactive language agents typically requires rich user interactions, yet collecting human feedback is expensive and difficult to scale. Simulated users offer a scalable alternative, but they must both resemble real user behavior and provide useful learning experiences for agents. In contrast, most agent-training frameworks rely on off-the-shelf assistant LLMs, whose helpfulness can make them overly cooperative, explicit, and behaviorally homogeneous compared with real users. We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions. Empirically, our 9B model achieves a SOUL-Index of 65.7, surpassing the strongest frontier model. Compared with Claude-Opus-5, the strongest baseline on RealUserSim and SimulatorArena, MIMESIS improves behavioral fidelity by 13.4 points and reduces Turing distance by 3.6 points, respectively. We then freeze the simulator and train an agent by interacting with the frozen simulator using multi-turn reinforcement learning. Across eight environments, training with MIMESIS yields better agent performance than training with GPT-5.5 under all nine unseen user simulators, demonstrating stronger generalization to new user simulators. Moreover, we propose Coached On-Policy Self-Distillation (CSD), which leverages simulator-generated private reasoning traces and subsequent utterances as feedback on how well the agent addresses user needs. A coach converts this information into concise coaching notes that describe how the agent can better anticipate user needs and adapt its behavior over the course of an interaction. CSD turns this feedback into dense, token-level supervision beyond sparse task rewards, yielding further gains across all nine evaluation user models.
cs.AI / 44 / 2610.09493
The Attribution Blind Spot: Layerwise Trajectory Diagnostics for Source Reliance in Retrieval-Augmented Language Models
Abstract
A retrieval-augmented model can match a document without relying on it. Controlled knowledge conflicts make source choice observable and let us ask a second question that prediction alone cannot answer: which internal-state properties define useful intervention directions? We study paired hidden-state changes with Latent Trajectory Shift (LTS), a signed projection onto a training-fitted first principal component (PC1), and keep verified training exposure separate from behavioral source choice. Across the evaluated conflicts, state-change magnitude is often the stronger predictor, whereas signed PC1 is the stronger selective controller: equal-norm interventions change source preference while better preserving non-target behavior, and the frozen direction transfers across the tested datasets and aligned model pairs. A same-system OLMo study further combines positive choice and control results with inconclusive exposure detection at the achieved power. The central result is a separation: representations that diagnose what a model will choose need not be the representations that best control that choice.
cs.AI / 45 / 2610.09495
Cognitive Schemas, Laws and Tasks
Abstract
This paper asks how explicit representations can support reusable cognitive schemas in knowledge-based problem solving. We develop a structural framework in which schemas are organized by the information and relations required for their use, rather than introduced as unrelated primitives. The framework also distinguishes context-dependent relations from more stable structures that can be reused across different representations. Tasks are described through the information available, the unknowns to be determined, and the constraints that admissible solutions must satisfy. This makes it possible to separate limitations of the representation from limitations of the solving procedure. In particular, we distinguish inconsistency, underdetermination, and contextual insufficiency, where the current representation lacks distinctions or relations required by the external task meaning. We also show formally when a reduction of representation preserves the task-relevant solution structure. The resulting task--schema interface offers a structured way to describe representational conditions relevant to problem solving. It supports the reuse and stabilization of derived knowledge while remaining independent of the particular mechanism used to generate candidate solutions. This may provide a useful component for future solver architectures that combine structured knowledge, verification, and learned proposal mechanisms.
cs.AI / 46 / 2610.09497
Ream: Unfolding Mutual Awareness in Human-Agent Workspaces
Abstract
As AI agents work alongside humans in shared workspaces, a mutual awareness challenge arises: agents act at speeds that outpace human monitoring, and users' evolving interests are not always expressed in chat. This challenge is especially pressing in literature review, where both parties retrieve, read, and synthesize a growing body of papers. We present Ream, a literature review workspace that supports mutual awareness through structured artifacts, bidirectional engagement tracking, and localized visualizations. Users can see each party's activity within these documents, and agents can retrieve the same history to guide their work. In studies with eighteen researchers, participants used these traces to inspect evidence, steer agents, communicate through annotations, and reflect on their research focus. Shared histories also helped agents build on earlier work. These findings inform how engagement traces within shared documents can support transparency, personalized assistance, and coordination in human-agent knowledge work.
cs.AI / 47 / 2610.09554
Reliability of LLM Judges for Evaluating Entity Alignment
Abstract
Entity Alignment (EA) identifies equivalent entities across knowledge graphs and is critical for knowledge base integration and ontology merging. Evaluating EA systems at scale requires expensive expert annotation, making systematic assessment across diverse domains practically infeasible. LLM-as-judge evaluation offers a potentially scalable alternative, yet its reliability for structured prediction tasks like EA remains unstudied. We present the first systematic benchmarking study across three frontier models, three datasets, and four EA systems, using perturbation bias diagnostics, meta-evaluation across all dataset-judge-prompt combinations, and counterfactual label-flip tests. We identify anchor bias, a failure mode in which judges invert discrimination when the system's decision label is visible. Label exposure causally collapses judge discrimination (J-ROC-AUC 0.12-0.87), while a label-free protocol recovers near-ceiling capability on distinctive-name datasets (0.93-1.00) and significant recovery on biomedical pairs (0.93-0.95). Counterfactual experiments confirm causality (FSR 53-99%) and reveal a frontier model paradox: stronger judges exhibit greater label sensitivity, not less. A blinded two-annotator human evaluation (102 pairs, Cohen's kappa=0.902) confirms this mechanism directly. We release the first biomedical EA benchmark (MeSH-SNOMED CT, 15K pairs) and a reproducible auditing framework for LLM judge reliability in EA. Code and data are available at https://github.com/vaibhavalakshmiravideshik/llm-as-a-judge-entity-alignment.
cs.AI / 48 / 2610.09558
DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists
Abstract
Drug target discovery requires distinguishing molecules that causally drive disease from those that are merely associated with it. Training and evaluating AI agents to perform this workflow end-to-end is difficult because real world biobanks lack known causal ground truth and participant-level data is access controlled. We introduce DrugTargetWorld, a framework that procedurally generates simulated biobanks, or "worlds," with known but concealed causal structure. Each world contains genotypes, proteins, health records, outcomes, and synthetic magnetic resonance imaging (MRI) for 54,000 participants. Agents must construct a disease phenotype, identify causal driver proteins, infer the beneficial direction of modulation, and optionally conduct virtual 'wet lab' experiments. We evaluated nine agents in 540 episodes across 20 cardiovascular worlds and three experimental budgets. Opus 5 and GPT-5.6 Sol achieved the highest mean composite scores, 39.98 and 35.38 of 100, respectively, and both recovered 64% of causal drivers on average. However, no agent reliably distinguished misleading non-causal proteins, and performance remained limited by the integrative judgments required to connect phenotype construction, causal evidence, and intervention decisions. By making each world's causal structure known to the evaluator but hidden from the agent, DrugTargetWorld turns end-to-end drug target discovery into a scalable training and evaluation problem with verifiable reward.
cs.AI / 49 / 2610.09560
World Potential Model: Pretrained World Knowledge as Progress Potentials
Abstract
Long-horizon language agents often receive supervision only from terminal task outcomes, leaving little signal for distinguishing productive intermediate behavior from stagnation or even regression. Rather than learning a separate value function or process reward model for every task, we ask whether pretrained models can recognize task progress from their existing world knowledge. We formalize this capability with a World Potential Model (WPM), a goal-conditioned evaluator of task-relative realized progress in agent contexts. In ALFWorld and ScienceWorld, off-the-shelf pretrained models substantially outperform chance at recovering realized-progress structure without task-specific evaluator fine-tuning. We further anchor these progress judgments to task-specific milestones to obtain scalar world potentials, whose temporal differences provide process-sensitive step-level credit for policy optimization. Under matched comparisons, WPM-guided optimization improves success over outcome-only GRPO across all evaluated configurations. Together, these results provide initial evidence that pretrained world knowledge can support reusable realized-progress evaluation and provide useful supervision for long-horizon agents.
cs.AI / 50 / 2610.09589
Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study
Abstract
Purpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect. Materials and Methods: This prospective, multicenter, randomized three-arm reader study was conducted at three hospitals in China from July to September 2026 (ChiCTR2600129243). After specialty stratification, 132 residents with fewer than 3 years of clinical experience were randomized 1:1:1 to GPT-5.4 alone (group A), GPT-5.4 plus Kimi-K2.6 (group B), or GPT-5.4 plus Gemini-3.6 Flash (group C); 123 were analyzed. Participants interpreted 60 radiographs before and after AI support. The primary outcome was accuracy change. Welch ANOVA and Holm-adjusted t tests compared support conditions; HC3 linear models assessed specialty interaction. Results: Among 123 residents (mean age, 24.1 years +/- 1.4; 65 women), radiology residents showed greater accuracy improvement with dual- than single-suggestion support (B-A, 6.69 percentage points [95% CI, 0.97-12.40]; C-A, 7.87 percentage points [95% CI, 1.64-14.11]; Holm-adjusted P = .030 for both), whereas accuracy change did not differ in non-radiology residents (P = .20). When GPT-5.4 was incorrect, AI-assisted accuracy was higher with dual- than single-suggestion support in radiology residents (40.1% and 40.4% vs 20.0%) and non-radiology residents (31.3% and 31.0% vs 12.1%) (all Holm-adjusted P < .001). The dual-suggestion effect differed by specialty (interaction difference, 10.44 percentage points; 95% CI, 4.36-16.52; P < .001). Conclusion: Dual-suggestion support may mitigate the influence of erroneous AI suggestions, with greater accuracy improvement observed in radiology but not non-radiology residents.
cs.AI / 51 / 2610.09624
How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression
Abstract
Tool calling, invoking external tools on demand, is central to agentic LLMs, yet the mechanism that decides whether a model calls a tool or responds directly remains poorly understood. Agentic prompts are long and heavily scaffolded, combining role instructions, tool schemas, format templates, and the user's request across hundreds of tokens, creating a noisy, highly entangled context in which no single controllable variable for mechanistic analysis is obvious. To obtain such a variable, we propose a method that converts complex agentic prompts into minimal contrastive pairs in which a single request verb determines the tool-call decision: replacing an execution-verb (e.g., \textit{write}) with an analysis-verb (e.g., \textit{discuss}) reliably flips the decision, suggesting it is mediated by a compact internal state. We construct 500 such paired prompts across Python, Java, and C++ (300 for mechanistic analysis, 200 held out for evaluation). We trace the decision to a vector, $μ_Δ$, that is both causally necessary and sufficient and generalizes beyond the discovery prompts to native multi-turn $τ^2$-Bench trajectories and verb-free requests. Behavioral ablations show that the scaffold establishes a tool-call prior; Transcoder decomposition then reveals that analysis verbs suppress this prior through features signaling that tool use is unnecessary, whereas execution verbs largely leave it intact. Downstream scaffold-reading attention heads and MLP features read out the resulting state, and the same mechanism recurs across seven models from the Qwen, Mistral, and Granite families. Our code is available at https://github.com/XijieGo/MI4ToolCalling.
cs.AI / 52 / 2610.09644
Automatically Building and Updating a Knowledge Graph of MLIP Models
Abstract
Complementing the many efforts in providing semantic representations of concepts, notions, and entities in materials science, we report and illustrate a process by which we can automatically build a knowledge graph of the fast evolving field of machine learning applied to the prediction of material properties, focusing on MLIP (Machine Learning Interatomic Potential). This LLM-based process relies on multiple steps, from information extraction in documents and articles to a validation loop using SHACL constraints to detect and correct errors. It is carried out on a model-by-model basis, focusing on the consistency of representation, therefore enabling an iterative construction where the addition of new models is facilitated. We illustrate the process by showing a few interesting aspects that can be queried from a knowledge graph built from the models listed in the Matbench Discovery leaderboard.
cs.AI / 53 / 2610.09652
MeshSIPP: Efficient Lattice Planning in Dynamic Environment
Abstract
Autonomous navigation in dynamic environments requires computing spatiotemporal trajectories that satisfy non-holonomic motion constraints. When the trajectories of the moving obstacles are predictable or known, a promising approach is to rely on the combination of state lattices constructed from precomputed feasible motion primitives and Safe Interval Path Planning -- a search-based algorithm with strong theoretical guarantees. While this approach yields feasible paths, the rich primitive sets needed for smooth navigation induce a large branching factor, which becomes costly when coupled with time-dependent obstacle intervals. To this end, we present MeshSIPP, an efficient planner that removes the computational bottleneck by exploiting the fact that many primitives sweep the same regions and can therefore be validated together. MeshSIPP propagates primitives as spatial bundles, screens them with lightweight bounding-interval checks, and defers the expensive exact departure-time search until a primitive reaches its terminal state. A time-aware pruning rule additionally discards redundant space-time branches early in the search. We prove that the resulting search is complete and optimal. Extensive experiments over more than 6,000 benchmark instances and real-time ROS~2 simulations show that MeshSIPP achieves up to a 3$\times$ speedup over state-of-the-art spatiotemporal planners.
cs.AI / 54 / 2610.09667
Shared and structured inputs undermine collective random choice by reasoning AI agents
Abstract
Random selection is widely used in resource allocation and auditing, making reliable implementation essential for AI-agent systems. Behavioural tests across six reasoning models uncovered threshold and divisibility rules used in identifier-based choices. For threshold-following GPT-6 Sol and Gemini 3.8 Flash, single-agent measurements prospectively predicted correlated participation under shared identifiers and biased participation under distinct identifiers with common timestamp bits. Changing dates, formats and identifier labels revealed when these predictions held. Explicit instructions to randomize independently reduced but did not eliminate shared-input correlation. To test implications for oversight, we asked four models to select customer requests randomly for human review. GPT-6 Sol approached the target rate while selecting predictably from identifiers; the others rarely selected requests. All four closely followed supplied random draws. These findings expose collective and audit vulnerabilities that selection rates alone miss, making input-dependent bias, correlation and predictability central targets for agent evaluation.
cs.AI / 55 / 2610.09683
System Switch: When Should a Fast Decision Model Stop and Think?
Abstract
Dual-process agents pair a fast policy with a slow deliberative model. In real-time settings the slow model usually runs continuously; in turn-based agents and robot planners it is invoked on events such as uncertainty or a detected failure. We study a fast learned actor that takes every decision and hands control to a reasoning vision-language model only when a gate opens, while the game keeps running. We use closed-loop Doom and the new open "System One" typed-decision models, served through a common llama.cpp interface. On 900 held-out questions, (i) zero-shot decision models from 0.15B to 9B parameters choose to collect items 1.6-1.8 times more often than chance among their errors, in any option order, although the order changes some models' accuracy; (ii) accuracy, calibration and sensitivity (how well confidence separates right from wrong answers) are distinct: models of similar accuracy differ widely in AUROC, and the confidence of the most sensitive one tracks which kinds of situation it fails, not which answers are wrong; (iii) offline, deferring the least confident 30% of decisions to a reasoning model gains over random deferral in proportion to the actor's AUROC (rank correlation 0.87); with actor and rate chosen on held-out games the gain is +0.13 [0.08, 0.18] with doomLaya's option order and +0.08 [0.02, 0.14] with shuffled options, and reasoning carries about half of it; (iv) in closed loop (33 games, three seeds) no variant reaches the exit. Committing to plans, the reasoner's or a fixed explore rule's, opens more doors and makes an actor that stands still play; with the rule the agent dies more often. Told that some doors need keys, the reasoner takes ordinary doors for locked ones, which the state cannot tell apart; without that knowledge it goes back to collecting. We release code, prompts, data and logs.
cs.AI / 56 / 2610.09693
A Tale of Two Error Categories: Exploring Concealed Trade-Offs in the Errors of Automated Judges in Evaluation of Uncertainty Quantifiers
Abstract
The wide adoption of LLMs across broad NLG applications heightens the importance of providing users with the means to avert errors and hallucinations. Uncertainty quantification is poised to fill that gap; with low uncertainty (high confidence), as a proxy for correctness, allowing users to be selective (e.g., reject low-confidence, likely incorrect responses). Correlation between confidence and correctness then serves as a useful criterion for evaluation of uncertainty quantifiers (UQs). But in NLG, where diverse responses can be adequate to a prompt, obtaining reliable correctness judgements is not simple, especially without human intervention. Errors in automated judgement are hardly avoidable and known to diminish the reliability of evaluation protocols (Santilli et al., 2025; Ielanskyi et al., 2025). In a meta-analysis of published work, we show that automated judgement is the present norm. Besides, automated judgements are rarely validated against human ones, and the validation of the UQ evaluation they automate is even rarer. With experiments in question answering, using 4 LLMs, human and automated judgements and 7 popular UQs, we find that i) a judge's performance can only coarsely predict the observed impact of its errors on the reliability of UQ evaluation, and that ii) judgement errors tend to misrepresent informative UQs most. We link these observations to patterns of correlation between confidence and categories of judgement error.
cs.AI / 57 / 2610.09832
SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles
Abstract
Memory-augmented reinforcement learning strengthens LLM agents' ability to solve complex long-horizon tasks. Skills are one such form of memory, pairing instructions with an applicability condition over task types. However, retaining every skill indiscriminately as the policy improves lets obsolete or harmful entries accumulate and mislead the agent. We propose SkillForge, an agentic RL method that compiles and evolves the skill library through a fitness-driven skill lifecycle of trial, active, stable, and retired states, so that the skills and the model co-evolve throughout training. A pre-RL evaluation phase first uses the base model's own rollouts to pre-retire low-fitness skills, yielding a filtered library that then seeds supervised fine-tuning. Reinforcement learning takes over from this checkpoint, and at each iteration selective retirement, stabilization, and LLM-guided mutation continue to forge the skill library alongside policy optimization. Across multiple interactive agent benchmarks, SkillForge achieves the highest aggregate success rate, delivering up to 7.8% relative improvement over the strongest baseline while keeping the skill library compact throughout training. We introduce SkillFurnace, a dataset of 5k+ annotated records bundling retirement-filtered SFT trajectories, evolved skill libraries with fitness annotations, and retirement events with human-annotated failure categories to support research on skill quality and lifecycle management.
cs.AI / 58 / 2610.09856
Self-Evolve With a Reference:Anchored Training of Tool-Integrated Agents
Abstract
Self-evolving tool-integrated agents learn from tasks and feedback generated within their own training loop. A Curriculum Agent generates tasks, while an Executor Agent learns from self-consistency signals through reinforcement learning. However, relying solely on the current Executor for feedback has two limitations: group-relative advantages vanish under full consensus, while uncertainty-based curriculum rewards favor disagreement without showing whether the generated tasks support further learning. These limitations motivate an additional reference beyond the current Executor. We propose \textit{AnchorLoop}, which introduces a frozen copy of the previous iteration's Executor as a historical reference and reuses it on both sides of the training loop. For the Executor, the anchor provides a cross-reference advantage that evaluates current outputs against both current and historical majority answers. For the Curriculum, it provides an agreement-based reference based on differences in sampled majority agreement. Since the Executor and anchor have identical parameters during Curriculum training, this comparison serves as a proxy for task selection rather than evidence of inter-version improvement or correctness. Across 13 reasoning benchmarks, AnchorLoop improves over Agent0 by 2.5\% on mathematical reasoning and 2.8\% on general reasoning tasks. It also maintains higher effective-advantage variance and continues improving in later iterations as the unanchored baseline shows diminishing gains. These results demonstrate the benefit of introducing a lightweight historical reference into self-evolving tool-integrated agents without external task or answer supervision.
cs.AI / 59 / 2610.09872
LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets
Abstract
Evaluating agents by outcomes alone can obscure the capabilities that produce them. This problem is especially pronounced in evolving environments, where outcomes reflect a closed-loop interaction between agent behavior and changing external conditions. We introduce LiveMACEBench, a process-aware benchmark that uses live financial markets as a naturally evolving testbed for persistent LLM agents. Five frontier LLMs operate along continuous trajectories under matched Tool Use, Persistent Memory, Rule Following, and Multi-Agent Collaboration configurations. We evaluate them through both realized outcomes and mechanism-specific diagnostics derived from complete decision traces. Across 30 days of live evaluation, we find a pronounced outcome-capability gap: realized returns often diverge from capability-specific measurements, and similar outcomes can arise from markedly different patterns of mechanism use. Trace-level diagnostics further expose distinct bottlenecks across capabilities, demonstrating that mechanism access, effective mechanism use, and downstream performance are not interchangeable measures of agent capability. LiveMACEBench makes this distinction measurable, turning live markets from a performance leaderboard into a diagnostic environment for agent capability
cs.AI / 60 / 2610.09890
Outperformance Inverse Optimization: Learning Objective Functions that Outperform Agent Decisions
Abstract
Inverse optimization estimates the weights of an objective function that explain observed decisions as optimal solutions, and is used in a variety of fields. For mixed-integer linear programs (MILPs), existing methods aim to reproduce the observations as optimal solutions, and thus learn compromise weights when the observations are suboptimal. We propose outperformance inverse optimization, which instead seeks weights that induce, at each state, an optimal solution outperforming the observed action in every component. We give a loss function that can be evaluated with forward-problem oracles alone and is thus applicable to MILPs, together with gradient-based and DC optimization algorithms for minimizing it. For weights inducing a unique outperforming optimal solution at all observations, we prove that the probability of failing to induce such a solution at a new state (the generalization error) is bounded by a quantity inversely proportional to the number of observations, and that this bound is tight in the number of observations up to logarithmic factors. In experiments on synthetic and real data, the proposed methods improve the prediction of solutions outperforming the actions over existing methods.
cs.AI / 61 / 2610.09896
NL2Hull: A Natural Language-Driven Constrained Ship Design Decision Framework
Abstract
Ship-form design combines smooth geometric representation, local shape editing, and constraints on the resulting hull. We present the Natural-Language-to-Hull Framework (NL2Hull Framework), which formulates ship-form editing as a typed discrete decision problem and connects language decisions to numerical geometry. Its Constrained Free-Form Deformation Engine (CFFD Engine) represents hull waterlines with non-uniform rational B-splines (NURBS), applies free-form deformation (FFD) to their control points, reconstructs the hull, and checks geometric constraints. We construct the Ship Design Decision Dataset (SDD Dataset) with 134,558 cleaned records and evaluate compared models on its subset Ship Design Decision Benchmark (SDDBench), containing 5,000 records and 43,496 typed questions. We propose Chip, a constrained ship-design decision model for processing natural-language requests. Chip reaches 95.90\% question accuracy and 99.32\% FFD exact match, with a negative log-likelihood of 0.0951, an expected calibration error of 0.0032, and a Brier score of 0.0551. The NL2Hull Framework provides a reproducible interface for evaluating language-based ship-form decisions while identifying the geometry and continuous-control components that require further development. Our code and dataset is available at https://github.com/wenhuahuo/NL2Hull.
cs.AI / 62 / 2610.09937
Where Can a Decision Model Diagnose HVAC Faults? Reasoning Demand, Physical Representation, and Robustness Under Shift
Abstract
Artificial intelligence supports building operations in several forms, each with its own barrier. Expert rules must be tuned for every system, supervised models need labeled data that buildings rarely record, and language models return free text that requires human-in-the-loop checking, since their stated confidence is unreliable. A newer kind of pretrained model, here called a decision model, returns a probability for every allowed answer, so one model could serve many decisions without training. This study answers three open questions for fault diagnosis in heating, ventilation, and air-conditioning systems: which decisions such a model can make, what input it needs, and whether its probabilities hold when conditions change. On 128 fault days from four public datasets of real equipment, faults are graded by the reasoning their diagnosis demands, with data given raw, as physical features, or with Brick topology. The decision model Jev, open language models, and a supervised model face nine tests that change season, control configuration, or building. Given physical features, Jev and the larger open model diagnosed faults whose evidence one feature carries, but not faults that need operating context. Under shift they kept their accuracy and calibration, while the supervised model lost 0.33 macro-F1 yet led or tied within a building. Their probabilities still needed correction, and detection was weak. The study maps which faults a decision model can diagnose and from what input, and supports a division of work in which code computes the physics and the model ranks candidate faults for an operator.
cs.AI / 63 / 2610.09944
AgentTime: Can Agents Estimate and Control Their Own Runtime?
Abstract
An essential control of AI agents is their ability to manage runtime. This ability requires a sense of time-awareness, to predict and estimate wall-clock time and to control their own actions. Prior work has focused on time-awareness, but duration-following and control in native agent harnesses remain unexplored. We present AgentTime, a benchmark for testing whether agents can work for a requested duration, predict their runtime, and estimate elapsed time afterward. It comprises 222 tasks from 18 sources spanning coding, computer use, agentic work, and automated research. Duration-following experiments append a single instruction specifying how long to work, with requests ranging from about a minute to multiple days. Accuracy on these instructions varies substantially: Fable 5.1 in Claude Code deviates from requested runtimes by a typical factor of 2.9$\times$, compared with only 1.2$\times$ for GPT-6 Astra in Codex. However, matching the requested runtime does not, by itself, establish continued work on the task. Among 158 reviewed Astra runs with classifiable transcripts, 14 explicitly slept after appearing to finish. In forecasting experiments, predictions tend to overestimate natural runtimes. In retrospective experiments, removing temporal information more than doubles deviation for Sol and Astra and nearly doubles it for Fable. An agent's ability to complete a task does not guarantee that it can control its own time or work for the whole requested duration. For agents to run reliably, safely, and autonomously over long horizons, we require the evaluation of both.
cs.AI / 64 / 2610.09973
From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM Agents
Abstract
When the harmfulness of an LLM agent's output can be quantified, a natural jailbreaking objective is to maximize expected harmfulness over admissible input modifications. An alternative approach constructs or selects harmful target outputs and modifies the input to increase their likelihood. We establish a precise connection between these two approaches through a probabilistic reformulation. Specifically, we show that the gradient of the logarithm of expected harmfulness with respect to the input equals the expected input gradient of the model's log-likelihood under a harmfulness reweighted output distribution. This identity provides a unified interpretation of expected harmfulness and target likelihood optimization. Building on this connection, we propose OPUR, a sampling distribution designed to generate highly harmful target outputs and use the resulting samples to guide likelihood-based input optimization. Experiments demonstrate the effectiveness of the resulting method in jailbreaking LLM agents.
cs.AI / 65 / 2610.10015
What the Sleeve Feels: Explainable Machine Learning for Textile Pressure-Based Postural Screening
Abstract
Pressure-sensing smart textiles convert body-surface contact into a dense, image-like signal closely tied to posture and movement, making them a promising low-cost route to wearable posture screening. Realizing that promise, however, requires more than classification accuracy: a deployable system must generalize to wearers unseen during training, expose the physical evidence behind its decisions, and tolerate the small donning offsets that occur whenever a garment is removed and re-worn. This paper addresses these three requirements jointly using a knitted piezoresistive sleeve worn on the forearm as a testbed. We regroup fine-grained everyday activities into three coarser screening categories (neutral, potentially undesirable, and functional or transitional), engineer 29 interpretable pressure-distribution features spanning global intensity, spatial center of pressure, quadrant asymmetry, distribution complexity, and short-horizon temporal change, and evaluate under a strict subject-wise split. A tuned XGBoost classifier reaches 0.818 accuracy, 0.788 balanced accuracy, and 0.801 macro F1 on unseen test subjects, with tight frame-level bootstrap 95% intervals of about plus-minus 0.01 and a subject-to-subject standard deviation near 0.06 under leave-one-subject-out cross-validation. A simple 2D-CNN baseline trained on raw frames achieves broadly similar performance, showing that hand-engineered features are not left behind by a learned spatial representation on this task. SHAP-based explanation, a feature-group ablation, per-activity error analysis inside the pooled undesirable class, class-mapping sensitivity, and a simulated donning-rotation stress test together locate what the model relies on, where it degrades, and why, directly targeting the generalization, interpretability, and robustness gaps that determine whether such a system is deployable.
cs.AI / 66 / 2610.10062
Loud Failures, Quiet Failures: Fault Detection and Recovery in Tool-Using Language Model Agents
Abstract
Tool-using agents are usually scored on whether they finish a task while the tools work. Deployments are less forgiving: services time out, endpoints disappear, parameter names change, and results come back well formed but wrong. Prior work has shown that language models over-trust tool outputs that fail silently; we ask how that over-trust plays out across the stages of failure handling in multi-turn agents. Wrapping the executable environments of an established function-calling benchmark in a fault-injection layer, we inject one of four typed faults at a controlled point in the trajectory and record whether the agent notices, changes plan, recovers the task, or repeats itself. Six models from three families, half of them reasoning variants, ran 1,920 trials over 24 multi-step tasks. Agents treat a failure as a problem in 91.3% of trials when the tool returns an explicit error, but in 58.8% of trials when it returns a plausible wrong value, against a 26.8% rate of reporting problems when nothing was wrong. Reasoning models are not better placed: paired against instruct siblings, they notice less (-9.3 points, p < .001) and change plan more (+10.4 points, p < .001), and recovery is unchanged (p = .512). Because agents are stochastic, two fault-free runs of the same task end in the same state only 63.3% of the time; against that baseline, only a missing tool clearly lowers recovery (39.9%), while timeouts, schema drift, and corruption stay within run-to-run variation. After a fault, agents return to the same tool three or more times in a row in up to 22.2% of trials, though strictly identical repeats are rare. A prompt line asking the agent to check each result did not move detection. Agents respond to the error channel rather than to the content of what a tool returns, so failures that stay inside the expected format pass through.
cs.AI / 67 / 2610.10071
HGP:An on-device personalized agent memory via hybrid graph storage
Abstract
LLM-based agents face challenges in personalized interactive tasks due to heterogeneous, multi-typed, and implicitly constrained long-term traces. Existing memory mechanisms struggle with accurate routing and retrieval, especially on-device where personalization is critical. Most methods use single-vector representations, blurring type distinctions and relational structure. We propose HGP, a hybrid graph memory framework. HGP employs a lightweight self-enhancement classifier for personalized memory routing and constructs episodic, semantic, and procedural memories as graphs. It also extracts working memory as a state trajectory to capture current state and implicit constraints, ensuring reliable decision-making. The classifier reduces large-model calls, enabling on-device deployment, while graph storage enables accurate retrieval and incremental user profile refinement. Experiments on two benchmarks show that on PAL-Set solution selection, HGP achieves an S-score of 35.58, nearly 7 points above the strongest baseline. Code and data are at https://github.com/Ouan6/HGP-.git.
cs.AI / 68 / 2610.10088
SkillSandbox: Skill Verification via Dynamic Scenario Synthesis
Abstract
Self-evolving agents distill task-solving experience into skills for future reuse, but these skills can encode incorrect procedures or non-transferable knowledge. It is therefore critical to verify each skill's reusability: whether its guidance remains useful beyond the experience from which it was distilled. Such verification requires observing how a skill affects execution in new tasks, yet existing tasks may not expose the situations where the target skill can actually be exercised. To construct such situations, we propose SkillSandbox, a framework that dynamically synthesizes a task and its environment for each skill that are skill-relevant yet novel. A Proposer specifies the conditions to preserve and the source-specific details to vary, a Builder constructs an executable scenario, and a Verifier compares executions with and without the skill. The Verifier assesses executability, utility, and efficiency to assign a Keep or Reject verdict, determining whether the skill enters the library. Across ALFWorld and WebShop with three models, SkillSandbox consistently yields the strongest downstream performance and improved execution efficiency. Further analyses examine whether these gains reflect accurate assessment of skill reusability and identify which components of SkillSandbox contribute to them.
cs.AI / 69 / 2610.10120
RewardWeaver: Long-Horizon Interactive Learning for Language Agents via Self-Evolving Reward Adaptation
Abstract
Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in domains where task outcomes can be reliably evaluated, but long-horizon interaction remains challenging due to sparse terminal feedback and difficult credit assignment. Process rewards provide denser supervision, yet the capabilities most relevant for training can change as the policy evolves: a behavior that is easy to evaluate or frequently deficient need not be the bottleneck currently limiting task success. We introduce RewardWeaver, a self-evolving reward adaptation framework for language agents in long-horizon interaction. RewardWeaver maintains a validated capability space in which the semantics of admitted Rubrics remain fixed, and closes the loop between policy optimization, task evaluation, failure attribution, and reward adaptation. After each training stage, it performs outcome-grounded backward attribution on low-outcome trajectories, aggregates recurrent and policy-controlled capability bottlenecks, and dynamically selects the corresponding process rewards for the next stage. Recurrent failures not covered by the existing capability space trigger a separate, controlled expansion procedure. We evaluate REWARDWEAVER on SOTOPIA, Amazon?HistoryPrice, and a newly constructed Sales Benchmark. Across social interaction, bilateral bargaining, and domain-specific sales, REWARDWEAVER establishes new state-of-the-art (SOTA) results. Ablations further demonstrate the importance of dynamic reward allocation, failure-grounded attribution, and stable semantics for admitted capabilities.
cs.AI / 70 / 2610.10201
GAGR-Lab: Evaluating Joint Spatial-Geometric and Analytic Function Reasoning
Abstract
Joint spatial-geometric and analytic function reasoning requires translating a perceived spatial configuration into a symbolic function whose executed curve satisfies geometric constraints. We present GAGR-Lab, a framework for measuring this capability through Cartesian game scenes, explicit function semantics, and authoritative Rust trajectory execution. It distinguishes spatial perception, metric grounding, geometric relations, function interpretation, function construction, and constrained synthesis. We specify four configurable scene-difficulty presets and a prospective 24-cell diagnostic design, while reporting only the subset actually evaluated. A bounded pilot of one hosted model (Llama 3.2 11B Vision Instruct) using two API credentials as execution replicas yields 72 balanced games with 432 attempts, 429 valid provider responses, and no target hits; exploratory ordinary-function prompt variants also fail to hit, while the structured localization interface yields no scoreable outputs. A privileged analytic search control independently succeeds on 600 directional cases from 300 generated scenes, with exact repeatability and 1,200 successful vertical-reflection or translation checks. The framework separates serving reliability, symbolic compliance, and geometric success, and preserves exact model-visible inputs and realized paths. A staged protocol outlines diagnostic calibration, held-out replication, multi-model comparison, and paired robustness tests. The contribution is an operational research framework with an executed pilot and a clearly identified prospective study plan; the full difficulty matrix and comparative model results remain untested.
cs.AI / 71 / 2610.10256
OOM-RL II: Reality Is an Oracle, Not a Debugger Provenance-Constrained Diagnosis in Continually Evolving Agent-Engineered Systems
Abstract
Reality may establish that an outcome occurred without identifying which evolving procedure produced it or why. This distinction matters in production ML systems whose code, configuration, and artifacts change while external feedback accumulates. We examine it in a human-directed, agent-engineered quantitative trading system, using oracle to mean an external source of realized outcomes rather than a complete correctness specification. Across one year, the account gained and outperformed a broad market index, while annual alpha was not statistically distinguishable from zero under the main retrospective specification. Retrospectively selected subperiods include adverse relative performance and conditional candidate-level weakness under declared approximate references. Engineering records document changes during the episode, and complete recommendation-to-runtime binding is unavailable. The archive does not establish a common frozen instance or a unique cause. The case motivates an outcome--diagnosis gap: outcome evidence, evaluated-object identity, and causal explanation support distinct claims. We distinguish frozen instances, pre-specified adaptive procedures, and ad-hoc development; organize archive-relative claim identifiability and an evidence hierarchy; and propose a prospective production-binding protocol. An illustrative compatible-history example shows how factual binding can resolve a recommendation's referent without supplying its counterfactual effect. The protocol is proposed rather than prospectively validated. External feedback constrains outcome claims, while provenance and additional identification structure determine the resolution of diagnosis.
cs.AI / 72 / 2610.10265
Stale, Misattributed, or Late: Where Personal Memory Fails Before Generation
Abstract
Personal memory for language agents is usually judged by whether the final an- swer is correct. That score hides errors that arise before generation: the memory block may contain an obsolete value, a fact about the wrong person, or no use- ful fact before the serving deadline. We measure these failures directly. Using Personal Fact Memory (PFM) as a reference layer, we find that temporal validity is primarily a property of memory construction in our setting. On a controlled revision benchmark, serving only the active value of each correctly keyed slot eliminates observed stale exposure; without update resolution, 70.3% of prompts expose a superseded value. Once retrievers share the same active store and par- ticipant information, participant-aware BM25 is equivalent to the reference ranker within a prespecified 0.02 margin. The harder problem is assigning revisions to the right slot. Missed merges leave stale values active, whereas false merges silently remove current values; four LLM key assigners achieve higher key re- call than a rule extractor yet produce lower clean-retrieval rates, and open-domain merge recall on LongMemEval never exceeds 0.062. Misattribution survives va- lidity filtering: an entity posterior reduces same-name exposure on controlled data but cannot distinguish identically named speakers in LoCoMo. Two frozen lan- guage models reproduce prompt errors in generated text. Retrieval latency varies across rankers, but prompt prefill dominates turn-level latency on our hardware. These results argue for evaluating agent memory before generation, separating stored-state validity, identity resolution, abstention, and serving latency.
cs.AI / 73 / 2610.10285
AI Safety Considerations for Agents With Limited Time to Act
Abstract
In the wake of the increasingly public discussion about AI alignment, recent work has tried to propose specific AI architectures that behave safely. However, the proposed arguments that seemingly demonstrate proved alignment mostly neglect the environment the agent needs to act in. We discuss theoretical bounds for agent-agnostic safety guarantees in environments that can only be partially observed and within which an action is required within limited time. We introduce two realistic scenarios, one with an infinite state space and one with signal mixture. In these scenarios, we prove that even a perfect agent cannot guarantee safe behaviour. It will be argued that for any proof of AI safety or alignment, the environment and associated safe actions need to be specifically considered together with the agent.
cs.AI / 74 / 2610.10352
The Handover Problem: Governing Autonomy Transitions in Human-AI Collaboration
Abstract
Human-machine systems rarely operate at a fixed level of AI autonomy. As operators and AI systems collaborate over time, control must shift: the AI can take on more responsibility when collaboration is stable, maintain its current role when evidence is ambiguous, or return control to the human when conditions deteriorate. Existing work on adaptive automation, supervisory control, trust in automation, and deskilling explains parts of this problem, but provides no auditable, multi-signal criterion for governing when autonomy should change across multi-cycle workflows. We formalise this challenge as the Handover Problem: deciding, at each operational cycle, whether to escalate, maintain, or revert AI autonomy while keeping the process reversible, recoverable, and auditable. We introduce the Handover Readiness Score (HRS), a transparent composite measure that integrates four signal dimensions: operator readiness, human-AI trust, learning stability, and operational performance. It is combined with a hysteresis-based transition policy that requires sustained positive evidence before increasing autonomy but reverts promptly when conditions worsen. Across software engineering and manufacturing domains, the HRS and hard safety guards address complementary failure regimes: guards enforce immediate corrective action when a single indicator breaches a critical threshold, while the HRS detects the slow, multi-signal erosion of operator readiness that no individual guard can observe. The framework establishes autonomy handover as a governance problem requiring explicit, composite, and auditable criteria. This provides a conceptual and formal foundation that adaptive automation research has not previously provided.
cs.AI / 75 / 2610.10407
SOTA: Stock Options Trading Agents Guided by Option-Implied Return Distributions
Abstract
As option markets grow and AI advances, agentic systems for option trading are gaining increasing attention. Language-model-based agents can reason over contextual information such as news, but option trading presents a particularly challenging decision problem: a single stock can have thousands of contracts, and the agent must decide both which contracts to trade and how to combine them. Existing approaches often sidestep this complexity by restricting the policy to a fixed strategy structure, such as a straddle, limiting their ability to switch strategies as market conditions change. We present SOTA (Stock Options Trading Agents), an agentic trading framework for structured option-strategy selection. SOTA abstracts the large option universe into strategy-level decisions while deterministic resolvers handle portfolio implementation. We develop SOTA by post-training Qwen3.8-27B with supervised fine-tuning followed by reinforcement learning. SOTA is evaluated on options on nine large-cap U.S. equities and SPY against rule-based and machine-learning strategy selectors in the same trading environment. Over a six-month out-of-sample period, SOTA earns an 18.3% total return with a Sharpe ratio of 1.60 and a maximum drawdown of 8.96%. We also document an asymmetric role of news: news improves frontier-teacher trajectories, but retaining news during reinforcement learning reduces out-of-sample return from 18.3% to -2.7%.
cs.AI / 76 / 2610.10444
RunningTab: Direct Workspace Interaction with Environment-Side Tabs
Abstract
Much knowledge work produces new deliverables from files a workspace already holds, and LLM agents are beginning to take such work over. Through direct corpus interaction, an agent can search and read any of those files from a terminal with no indexing, and producing a deliverable from many of them in this way is what we call direct workspace interaction (DWI). Reaching the files, however, is only half the task: nothing keeps track of what the task asks for, what has been read, and what was listed but never opened, all of which slip through the context window without leaving a trace, so an agent may extract a figure and still deliver a report without it. To address this, we present RunningTab, a framework that equips direct workspace interaction with an environment-side tab: a per-task record of what the task still owes, kept by the environment alongside the agent. Specifically, the agent adds its requirements, while the environment records every file read as an excerpt with its provenance and every listed but unopened file as a candidate; the agent can then see each requirement beside its best-matching excerpts and top unopened candidates, resolve it against matching content or set it aside with a reason, and, should it try to finish with requirements still open, receive them in a finish check. We validate RunningTab on three benchmarks with three LLMs, where it consistently outperforms plain DWI and baselines that keep the record in the model, while its tab usually holds the values a deliverable needs once seen.
cs.AI / 77 / 2610.10478
Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models
Abstract
How can we predict which base checkpoint is worth an expensive round of agentic post-training? End-to-end pass@$K$ tests whether successful behavior already appears in a base model's distribution, but it is a poor fit for agentic coding: many base checkpoints cannot reliably produce the well-formed tool invocation required to complete a task end-to-end. Single-shot or short-horizon tasks avoid these tool-calling failures by collapsing a multi-step interaction into a fixed prompt and a single patch, but they sidestep the core capability we care about: maintaining coherent state over many tool-using steps as the repository evolves. To bridge this gap, we treat successful post-trained agent trajectories as a lookahead signal of base-model potential. Replaying each trajectory and rerunning tests after every code-changing step identifies the decisive step: the first step whose cumulative patch flips the repository from failing to passing, certifying that the recorded action solves the task given the prior context. Motivated by a coverage principle for agentic traces, we build three screens at this step that do not require a base checkpoint to drive the harness from a cold start: (i) Decisive-Action BPB (bits per byte) measures the probability mass on the certified action, (ii) Patch MCQ tests the checkpoint's choice between that action and alternatives rejected by the same verifier, and (iii) prefix-conditioned pass@$K$ evaluates support for functionally-correct generations and credits any continuation that the tests accept. Across ten pairs of public base and post-trained models, all three screens rank the cohort in close agreement with post-trained SWE-bench Verified pass@$1$. As our methods need only a benchmark's successful trajectories and its verifier, they can be applied to turn future agentic coding benchmarks into base-model evaluations.
cs.AI / 78 / 2610.10498
EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution
Abstract
Robot foundation models provide strong visuomotor control, yet their performance can degrade when object positions or task instructions change. Further improvements often require post-training on substantial robot data, which can be costly to collect through methods such as teleoperation. Agentic harnesses can adapt around the model, but current self-evolving harnesses use robot trials inefficiently when deciding which code and skill changes to pursue. We introduce EmbodiedRSI, a self-evolving agentic harness that autonomously decides where to explore next and turns the resulting physical interaction into improved code and skills. EmbodiedRSI realizes this through a Fast-Slow Dual-System Architecture, in which competing code and skill hypotheses are maintained in a Hypothesis Graph. Value-of-Information Experiment Selection chooses physical experiments that can distinguish these hypotheses. Their outcomes guide Code-Skill Co-Evolution. The Slow System builds Hierarchical Memory, and Reward-Grounded Memory Learning selects effective memory according to their value for later Fast-System improvement. On RoboCasa365, EmbodiedRSI reaches 77.0% overall success and 71.3% on Composite-Unseen, compared with 40.1% for the best baseline. EmbodiedRSI also reaches 86.8% overall success on LIBERO-Pro. Beyond benchmark performance, EmbodiedRSI transfers zero-shot to real-world robot, achieving 71.3% overall success across multiple challenging tasks.
cs.AI / 79 / 2610.10513
SciExam for ENSO: Can AI Agents Build Climate Models?
Abstract
Language-model agents are increasingly asked to carry out open-ended scientific research, yet their results are usually graded against a known answer, a rubric, or a language-model reviewer, none of which can tell whether a new scientific model is valid. The AI Science Exam for El Nino-Southern Oscillation (SciExam for ENSO) is a benchmark in which agents build low-order stochastic models of ENSO, the dominant mode of interannual climate variability, from real observations. Within a six-hour budget, agents process the observations, write their own diagnostics, which are then frozen, and develop a model using only these diagnostics as feedback. Hidden graders then test whether the model reproduces ENSO's statistics, recovers unobserved variables, and forecasts held-out years, and score a published model in the same way. Across twelve agent systems, six produce models that score higher than the published model, mainly through better reconstruction and forecasting. The simplified forms of the stronger models are each compatible with one of the two competing explanations of ENSO's warm-cold asymmetry, an open debate that the task never mentions. Controlled runs of the top system under varied information suggest that its scores do not come from recalling the dated observational record and that the information it receives shapes how it builds its model. SciExam for ENSO can thus evaluate agent research where no answer is known, and the results suggest that agents can already build competitive models whose structures bear on questions that scientists still debate.
cs.AI / 80 / 2610.10515
RoboJEPA: Scaling Robotic Latent World Models
Abstract
Latent world models have shown a remarkable ability to predict future states and to plan in the real world. In practice, however, we lack a principled way to estimate how their capabilities scale with model size, data, and compute, an open problem that slows progress in the field. In this work we present RoboJEPA, a world model based on the Joint Embedding Predictive Architecture (JEPA) and trained on a large-scale dataset spanning 12 robotic embodiments. We show that RoboJEPA's imagination error, the error of its latent rollouts, follows a second-order power law in compute, allowing us to predict model quality well beyond the scale at which the law is fit. We further show that downstream robotic planning performance improves predictably with compute, and that imagination error is strongly correlated with it, making it a reliable proxy for real-robot evaluation. Finally, we demonstrate that latent world models can be deployed zero-shot as robotic agents, planning toward a single goal image to solve tasks requiring long-horizon planning on real hardware. We release all model checkpoints together with our training and robot deployment code. To our knowledge, this is the first work to establish scaling laws for multi-embodiment robotic world models trained on real robot data, and RoboJEPA, at 8B parameters, is the largest JEPA predictor model trained to date.
cs.AI / 81 / 2610.08944
One-Slide Calibration of Pathology Foundation Models
Abstract
Scanner variation changes how pathology foundation models represent the same tissue. We introduce SlideRuler, which uses regions within a slide as internal controls to estimate and correct acquisition-induced shifts in other regions. A transfer map learned from paired rescans enables calibration from a single scan at inference while keeping the foundation model fixed. Across two encoders and five SCORPION scanners, learned transfer reduces mean target-to-source embedding distance by 16.3-38.5% relative to raw embeddings. Comparisons with unrelated same-scanner controls reveal a positive same-slide contribution across all four evaluation settings, including scanner holdout. A source-anchored variant reduces source-feature displacement by 47.7-83.6% relative to learned transfer while retaining most of its alignment gain. By drawing calibration information from the slide itself, SlideRuler offers a path toward more consistent use of frozen pathology models across imaging systems.
cs.AI / 82 / 2610.09205
Do Vision Models Learn Physical Constraints or Rendering Shortcuts? A Counterfactual Benchmark for Grounded Physical Consistency
Abstract
Modern image editing models can satisfy a text instruction while breaking the physics of the edited scene. A new object may cast no shadow, a mirror may fail to reflect visible geometry, or an object may float above a surface that should support it. We study physical plausibility diagnosis, detecting whether an edited image violates scene physics, naming the violation type, localizing the affected region, and explaining the failure in language. We introduce a counterfactual benchmark whose controlled synthetic component uses Mitsuba~3 to generate 5,500 images from 500 scene families. Each family contains one clean image and ten matched violations involving shadows, reflection, support, surface response, and occlusion. The renderer pipeline provides category labels, affected-region masks and boxes, scene metadata, and explanation targets. We use LLaVA-1.5-7B, Qwen2.5-VL-7B, and InternVL3.5-8B as diagnostic baselines rather than proposed methods. On a 1,650-image synthetic test set, the adapted baselines reach 64.0--67.8\% category macro-F1 on standard held-out scenes. For LLaVA-1.5-7B, category macro-F1 falls from 64.0\% on the standard split to 40.8\% under intervention shift. This gap shows that high in-distribution accuracy partly reflects cues tied to rendering and counterfactual construction.
cs.AI / 83 / 2610.09217
A Deterministic Evidence Layer for Vision-Language Autism Screening from Naturalistic Home Video
Abstract
Autism spectrum disorder (ASD) is diagnosed through specialist observation of a child's social behavior, and access to that expertise is the bottleneck for early identification. Vision-language models (VLMs) describe a child's behavior from video well; the verdict drawn from the description is unstable: at temperature~0, across eight pipeline configurations on one backbone, 16--37\% of clips change their predicted label between repeated runs, and the cause lies in the serving stack. We keep the VLM frozen and move the decision out of the model. Under our grounded perception constraints the VLM writes an event table of timestamped, glossary-labeled events that names the eliciting press and logs counter-evidence; a text-only stage age-calibrates the confidence of each row; a deterministic weight-of-evidence scorer sums it into an evidence total and stratifies it into a risk category, so every decision decomposes into named per-feature contributions and can be re-scored from the saved table. On 43 caregiver-recorded, protocol-free home free-play clips of preschool children, the pipeline reaches AUC $0.851 \pm 0.012$, 86.0\% accuracy, and F$_1$ 71.8 over three runs. It labels 74.4\% of clips correctly in every run (60.5\% for the zero-shot baseline) and flags no typically developing clip in every run (9 of 31 at zero-shot). An ablation on the same backbone attributes the gain to the grounded perception constraints read through the deterministic scorer.
cs.AI / 84 / 2610.09328
Visual Jev Rewards: Reference-Bound Verification for Multi-Subject Image Generation
Abstract
Multi-subject image generation requires rewards that verify whether requested attributes, actions, and relations hold for the specified reference subjects. Subject presence alone does not establish that the correct subjects participate in a requested interaction. We present reference-bound Visual Jev rewards that turn these visual decisions into generator training signals. Each subject-related question receives a positive label only when the requested condition and the relevant reference identities hold jointly. We construct fixed questions offline, train a Qwen3.5-4B verifier with binary supervision, and directly read Yes probabilities from its language-model head. Their mean supplies a GRPO reward while retaining individual judgments for inspection. Using 200 MICo-150K training tasks and 30 updates, the framework raises a GPT-5.4 composite score from 41.78 to 52.50 on a manually selected 897-task MICo-Bench subset; direct 27B rewards yield 51.84. Each reward is tested in one GRPO run, and offline human evaluation does not establish a statistically significant advantage over direct scoring. The study provides an initial implementation and evaluation of Visual Jev as a reference-bound reward for multi-subject image generation.
cs.AI / 85 / 2610.09498
SpatialUQ: Post-Hoc Uncertainty Quantification from Spatial Consistency in Black-Box Vision Models
Abstract
Clinical vision models are often deployed as frozen black boxes with no access to internals, retraining, or ground truth at inference time. We introduce \textbf{SpatialUQ}, a post-hoc uncertainty method using only output probabilities. It measures the Jensen-Shannon divergence between the global prediction and the mean of five fixed spatial crops in six deterministic forward passes. The premise is simple, trustworthy predictions are spatially consistent. On NIH ChestX-ray14 (DenseNet-121, $N{=}25{,}596$), our Multicrop Uncertainty Score (MUS) reaches $0.784$ failure-detection AUC versus $0.664$ for MC-Dropout ($p{<}10^{-6}$) at one-fifth the compute, with native calibration ($\text{SCE}{=}0.049$ vs.\ $0.127$ for $\ell_1$), the best-calibrated among methods above 0.78 AUC. A supervised fusion of MUS with entropy, confidence, and $\ell_1$ reaches $0.832$, outperforming a five-member ensemble ($0.813$). MUS scales with model quality, reaching $0.899$ with BiomedCLIP ($ρ= 0.846$), while this relationship remains meaningful in-distribution ($ρ= 0.523$) but breaks down under severe distribution shift (VinBigData, $ρ= 0.027$). MUS is well-suited to diffuse findings but is less dependable for small focal lesions such as nodules. Code and experimental materials are publicly available at https://huggingface.co/datasets/kawsher11/SpatialUQ.
cs.AI / 86 / 2610.09513
OmniCam: Omni-Camera Trajectory Generation via Geometry-Grounded Pose Token Learning
Abstract
Camera trajectories control viewpoint changes in video generation, scene reconstruction, and robotic perception. Generating them from language requires both scene geometry and target-aware framing. We introduce OmniCam, an autoregressive model that generates camera pose sequences from a single panorama and textual trajectory descriptions. Its geometry-grounded pose token learning combines three components: a panoramic point-cloud encoder for omnidirectional geometric context; hybrid absolute-rotation and relative-translation tokenization with temporally consistent quaternion signs; and separate geometric and semantic conditioning streams with an explicit 3D target anchor. We also construct OmniCaT, containing 267,700 trajectories across four camera behaviors. On the reported OmniCaT evaluation, OmniCam reduces trajectory errors by 28--47% and collision rate by 65.8% relative to GenDoP retrained on OmniCaT. Against the best baseline for each metric, the ATE and collision reductions are 43.0% and 62.3%, respectively. Component ablations support the use of geometric and target-aware conditioning, while downstream experiments examine camera-controlled video generation and robotic active perception.
cs.AI / 87 / 2610.09534
KASALv2: Fully Automatic 3D Rotational Symmetry Classification and Axis Localization
Abstract
Rotational symmetry is an important prior in 6D pose estimation, improving pose accuracy and supporting symmetry-aware evaluation. However, current symmetry annotations for 3D objects remain largely manual or semi-automatic, often requiring predefined types or orders, which limits scalability. This work introduces a fully automatic, reference-free framework for symmetry-type classification, rotational-order identification, and full-axis localization across all eight canonical 3D rotational symmetry types. The method localizes a dominant high-order axis, infers its rotational order through self-consistency analysis, and reconstructs the complete symmetry structure under a hierarchy-guided formulation. A texture-aware extension further models appearance-induced reductions in rotational order while preserving axis orientations. Experiments on idealized and real-world datasets demonstrate strong accuracy and generalization, achieving 94.75% accuracy on 438 symmetric objects in GSO. Training FoundationPose with these priors improves accuracy by up to 0.9% across five BOP datasets, showing that automatically estimated rotational priors improve downstream 6D pose estimation. Code is available at https://github.com/WangYuLin-SEU/KASAL.
cs.AI / 88 / 2610.09535
WAPR: A Foundation Model for Wide-Angle Refinement in Unseen Object Pose Estimation
Abstract
Real-world applications require 6D pose estimation to be accurate, fast, and scalable to unseen objects. This paper introduces WAPR, a zero-shot wide-angle pose refinement model that refines candidate poses with rotational deviations up to 90 degrees. With as few as 12 candidate poses per detected object instance, WAPR supports fast inference within 1 s per frame and reaches a pose-estimation throughput of up to 25 detected object instances per second. To support wide-angle training for rotationally symmetric objects, WAPR uses rotational symmetry priors to canonicalize symmetry-equivalent pose targets before loss computation. We further construct SA6D, a large-scale 6D training dataset with such priors. SA6D obtains KASAL-assisted rotational symmetry priors for 944 GSO scans and expands them through geometry and texture augmentation into about 50K augmented object instances and about 2M rendered RGB-D images. In addition, an angle-balanced loss stabilizes learning across different angular ranges by reducing the influence of uninformative large-error cases. Experiments on seven BOP core datasets show that WAPR achieves state-of-the-art performance in unseen-object 6D pose localization and detection under both fast and unconstrained inference settings. Project page: https://github.com/WangYuLin-SEU/WAPR.
cs.AI / 89 / 2610.09700
What Makes Synthetic Hard Negatives Work in Vision-Language Pretraining?
Abstract
Synthetic hard negatives generated in the representation space have proven effective for unimodal self-supervised learning, but transferring this idea to vision-language pretraining is not straightforward. We analyze six representation-space synthesis strategies and identify two failure modes in their transfer to vision-language pretraining: cross-modal constructions that produce overly easy negatives or pull them toward the query, and intra-modal constructions that incorporate the matched positive. We also observe logit-scale saturation when training with synthetic hard negatives and a learnable temperature, and find that fixing the temperature improves downstream performance. Using this geometric analysis we propose SNAP, which generates intra-modal hard negatives that never involve the positive from either modality, avoiding both failure modes entirely. SNAP is model-agnostic, requires no external generative models, and adds less than 10% training time overhead. Evaluated on top of CLIP and FLIP across multiple architectures and datasets, SNAP delivers consistent improvements on zero-shot retrieval, zero-shot classification, and linear probe evaluation.
cs.AI / 90 / 2610.09761
PARC-Loc: Text-to-Point-Cloud Localization with Partial Assignment and Relational Consistency
Abstract
Text-to-point-cloud localization estimates a position in a city-scale 3D map from descriptions of surrounding objects. Existing coarse-to-fine methods retrieve submaps using aggregate learned compatibility and then localize within a selected submap. However, repetitive or similar urban objects can inflate the embedding similarity between the query and multiple submaps, even when the instance layout within a submap violates the query description. Meanwhile, query-relevant instances often span submap boundaries, leaving the retrieved submap with incomplete contextual evidence. We term these failure modes layout-inconsistent aliasing and boundary evidence incompleteness, respectively. To address them, we propose PARC-Loc, a coarse-to-fine localization framework built on Partial Assignment with Relational Consistency (PARC). PARC jointly models hint-object compatibility and pairwise spatial relations, allowing unmatched elements while favoring assignments consistent with the queried layout. At the coarse stage, its candidate-level assessment complements neural similarity for layout-consistent submap selection. At the fine stage, the context is expanded with query-relevant instances from adjacent submaps, while PARC yields object-level matching weights that guide cross-modal attention. Extensive experiments on KITTI360Pose and CityLoc show that PARC-Loc outperforms conventional coarse-to-fine baselines. On KITTI360Pose, our method improves Top-1 localization recall at 5 m from 0.50 to 0.67, achieving a 34% relative gain over the strongest baseline.
cs.AI / 91 / 2610.09802
DisParQ: Self-Supervised Part Concepts for Interpretable Vision Foundation Models
Abstract
Concept-based vision models represent images through an intermediate layer of human-inspectable concepts, so what a model relies on can be traced to those concepts. However, those models are often limited to fixed categories or depend on language to define their concepts. We introduce DisParQ (Discrete Parts with Quantized attributes), a method that learns spatially grounded, discrete concept representations from a powerful frozen vision-only self-supervised backbone. It requires no class labels and no language supervision. Each image patch is assigned to exactly one concept from a learnable prototype dictionary, and only a sparse subset of concepts may activate per image. To capture how each concept varies across images (e.g., the type of a "wheel"), we learn continuous residuals alongside the concepts and then quantize them into discrete attributes. A spatial decoder reconstructs the backbone's representation from the concepts and attributes alone, so successful reconstruction means that the discrete representation preserves the backbone's information. We evaluate DisParQ across seven datasets, from general recognition (ImageNet, PartImageNet, Places) to fine-grained benchmarks (CUB, Cars, Dogs, Flowers). We show that DisParQ closely matches its frozen DINOv2 teacher on ImageNet linear probing (83.2% top-1), achieves higher concept consistency than language-aligned models, remains competitive on fine-grained recognition, and enables cross-category part-based retrieval.
cs.AI / 92 / 2610.09823
UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation
Abstract
Dense visual text requires image generators to reproduce long strings across multiple regions with correct placement and legibility. As short-string rendering improves, evaluation must test sustained performance across more demanding scenes. We introduce UltraText Bench, a bilingual benchmark for prompt-only generation of dense visual text. It contains 432 prompts spanning 24 real-world scene categories and three difficulty levels, split equally between English and Chinese. Each human-reviewed prompt supplies exact strings for four to twelve text regions, paired with structured references for their content, placement, and visual attributes. We use the Q-Judger vision-language model to assess each image against the complete reference, reporting text fidelity, text clarity, spatial quality, and scene quality. Across 24 model configurations, these dimensions reveal different strengths: Z-Image-Turbo gains 3.81 clarity points over Z-Image-Base while losing 14.76 fidelity points under the reported settings. Performance also varies with workload; Qwen-Image-2512's English composite falls from 86.50 at L1 to 42.86 at L3. Ten participants took part in human evaluation of the automatic scores. Repository: https://github.com/LINs-lab/UltraText_Bench.
cs.AI / 93 / 2610.10183
VideoEvolve: Co-Evolving Memory and Retrieval for Long Video Understanding
Abstract
Long video understanding increasingly relies on external memory to organize massive visual streams into compact representations. However, most memory-based methods dynamically adapt how information is retrieved for different questions, while largely fixing what is remembered. This mismatch makes missing details costly to recover, whereas stored information is valuable only when it can be reliably retrieved. To address this issue, we propose VideoEvolve, a novel self-evolving framework that jointly evolves memory and retrieval for long video understanding. Specifically, starting from a coarse low-frame-rate overview, VideoEvolve couples a Memory Evolver for selective memory augmentation with a Retrieval Evolver for adaptive retrieval over the evolving memory. We then co-evolve the two Evolvers through alternating agentic reinforcement learning (Agentic RL), updating one while freezing the other. To steer this alternating evolution, Bottleneck-Aware Evolution Feedback (BEF) identifies whether the current bottleneck lies in memory or retrieval and directs optimization toward the more limiting side. Furthermore, VideoEvolve introduces Capability-Aware Evolution Feedback (CEF) to alleviate downstream feedback from over-specializing memory to a fixed set of training questions, shifting training toward underdeveloped yet learnable video capabilities. By integrating Agentic RL with BEF and CEF, VideoEvolve transforms downstream reasoning experience into transferable capability updates, providing a concrete path from static long-video systems toward experience-driven, self-improving multimodal intelligence. Extensive experiments on multiple long video understanding benchmarks demonstrate the effectiveness of VideoEvolve.
cs.AI / 94 / 2610.10266
LoomSC: Scalable Deep Subspace Clustering with Projector Factorization and Exact Spectral Reduction
Abstract
Dense self-expression matrices and full-affinity spectral clustering limit the scalability of subspace clustering. We introduce the Latent Orthogonal Optimization Model for Subspace Clustering (LoomSC), a framework that addresses both bottlenecks through projector factorization and exact spectral reduction. Motivated by the spectral structure of least-squares regression, LoomSC jointly learns latent features and a projector self-representation through two thin factors. Alternating Procrustes and least-squares updates preserve the sample factor's orthogonality while keeping the coefficient matrix implicit. We construct a nonnegative quadratic affinity that preserves the projector's support. An exact feature map then reduces its normalized spectral problem to an eigenproblem whose dimension depends only on the factor width. Neither the full affinity nor the sample Laplacian needs to be formed. Our analysis quantifies the projector approximation and identifies conditions for subspace preservation and within-subspace connectivity. For fixed dimensions and iteration budgets, the complete pipeline has linear time and memory complexity in the number of samples. Across five image-clustering benchmarks, LoomSC ranks first or second in all 15 dataset-metric comparisons against 9 state-of-the-art baselines. Its mean accuracy exceeds the highest baseline mean by 6.66 percentage points. Synthetic experiments scale to 500,000 samples while maintaining at least 99.8% accuracy.
cs.AI / 95 / 2610.10324
Performance at What Cost? A Sustainability-Aware Performance Index for Cell and Nucleus Instance Segmentation
Abstract
Pretrained models for cell and nuclear instance segmentation differ substantially in architecture, pretraining data and objectives, parameter count, inference strategy, adaptation requirements, postprocessing pipeline, and computational demand. Large pretrained and foundation models are increasingly adopted because of their strong zero-shot capabilities, but their use also imposes greater energy consumption, memory requirements, computational demands, adaptation costs, and operational carbon emissions. Whether these additional demands are justified by meaningful gains in segmentation performance remains unclear. We address this question by introducing the Sustainability-Aware Performance Index (SAPI), a configurable metric that combines segmentation performance, energy consumption, and model size. We benchmark 19 pretrained and foundation models across six CellBinDB datasets under zero-shot inference and evaluate 16 fine-tunable models using few-shot adaptation with both frozen encoder and full-model fine-tuning. We estimate energy consumption for GPU, CPU, and RAM using software-based monitoring tools. Our results show that larger and more computationally demanding models do not consistently achieve proportionate improvements in segmentation quality. While few-shot adaptation benefits several models, the gains and resource costs vary considerably across architectures, datasets, and adaptation strategies, causing SAPI-based rankings to differ from rankings based on performance alone. This study provides a practical framework for comparing segmentation models more comprehensively and supports more computationally accessible and environmentally responsible model selection in biomedical image analysis.
cs.AI / 96 / 2610.10538
Never Look Back: Understanding Persistence in 3D Object Memory from Egocentric Videos
Abstract
As we move through the world and carry out everyday tasks, we encounter objects that may become relevant only later. We are capable of recalling where we left something or what was inside a container, even without knowing we would need it later. Here, we study how an embodied assistant can build a similar memory from egocentric videos, by observing a person's day-to-day activities. We present Ledger, a persistent 3D object memory that combines object locations, their histories, and contextual descriptions. It associates observations across the recording and retains objects after they leave the view, including those the person never touches. It clusters each object's observations by resting locations and records a move only after repeated evidence, reducing the effect of localization noise. Short descriptions preserve details such as an object's contents or supporting surface. It saves these records to later answer spatial questions without having to access the original images or video. Our memory raises HD-EPIC accuracy from 29.7% to 42.6%, UCS-Bench accuracy from 33.8% to 38.5% and localizes Ego4D objects with a 0.99 m median error on returned predictions. Our analyses identify complementary roles for temporal persistence, contextual descriptions, and retrieval. Our study on 100 stitched streams of multiple scenes each further exposes failures in both retrieval and construction. Per-scene construction partially recovers the performance lost across scene changes compared to that of single scene streams.
cs.AI / 97 / 2610.10064
Comprehension Audits to Mitigate Risks from Automated AI Research
Abstract
AI is already writing a majority of code for frontier AI labs. This creates a safety risk if there is insufficient human oversight. Existing work proposes minimum comprehension thresholds and unaided checks to mitigate this. To our knowledge, however, there is currently no published frontier-AI assurance regime that requires demonstrated evidence that the responsible humans understand what they are building as a precommitted condition for continuing development or usage. We propose comprehension audits, a novel development-process assurance mechanism in which the responsible people explain R&D contributions to auditors to demonstrate understanding. With independent administration and graded reports, they provide a gate: development of a contribution stops based on a failure to demonstrate human understanding until remediated, with escalating consequences for repeated failures. Our analysis of leading open-source AI projects finds increased output of code with reduced human review commentary rates per line of code, with far lower rates for automated fleet accounts. We advocate for labs to conduct them with embedded independent auditors.
cs.AI / 98 / 2610.10241
When Scientific Cognition Is No Longer Scarce
Abstract
AI could change which parts of science impede progress. Consider a world in which machine systems are better, faster, and cheaper than people at most scientific work that can be done through a computer. Our question is what would limit science in that world. Literature synthesis, hypothesis generation, software development, simulation, and analysis could become abundant, while experiments, observations, well-supported conclusions, and accountable institutional au- thority remain scarce. Science would then be constrained by a different set of resources. In this paper, we call this change the scarcity inversion and consider four parts of it: selection, physical access, validation, and organizational choice. This change is arriving first in mathematics and coding/software/algorithm design, where the whole scientific loop can run inside computation. For national laboratories, the change could be striking. Their distinctive role is to turn abundant machine reasoning into trustworthy results by combining controlled experiments, protected data, expert judgment, and accountable authority. The practical question is how facilities, verification, provenance, resource allocation, and scientific governance should change when reasoning is plentiful and trustworthy evidence is scarce.
cs.AI / 99 / 2610.09487
Correspondences as Decisions: JevNexus for Decision-Centric Schema Matching
Abstract
Schema matching increasingly uses generative language models to rerank retrieved column candidates, although the underlying task is a bounded correspondence decision. We present JevNexus, which combines typed pairwise decisions with schema/instance evidence and invokes listwise refinement only when the evidence disagrees and the fused margin is small. The evaluation covers 561 cases from six benchmark families. JevNexus obtains dataset-macro MRR and Hits@1 of 0.930 and 0.909, compared with 0.926 and 0.903 for Magneto, while reducing mean latency from 123.452 to 15.929 seconds (7.750). Paired analysis finds no statistically significant difference in either MRR or Hits@1. The gate invokes listwise refinement for only 5.665% of source columns and avoids the degradation caused by unconditional refinement. Code and experimental artifacts are available at https://github.com/RazeenLI/JevNexus.
cs.AI / 100 / 2610.09894
QCATS: Query Context-Aware Transformer Slicing for Efficient Predictive Query Processing
Abstract
In-database predictive query processing increasingly applies Transformer-based models within relational pipelines. However, existing in-database inference typically exposes only tuple-level model inputs to the inference runtime, leaving relational predicates and metadata statistics invisible to neural execution planning. In this paper, we propose QCATS, a query context-aware transformer slicing framework that enables efficient sparse inference inside database systems. QCATS executes at query granularity: instead of routing individual tokens or tuples during inference, it uses query predicates and metadata statistics to pre-select context-aligned FFN slices before model execution. The framework comprises offline expert construction and lightweight query-level routing that dynamically selects experts during execution. QCATS further introduces system optimizations, including asynchronous CPU-GPU pipelines and routing-aware batching. Experiments on four predictive-query workloads with BERT-base and Qwen-0.6B show that QCATS achieves up to 4.42x latency reduction while preserving prediction accuracy comparable to dense baselines.
cs.AI / 101 / 2610.10061
TRACK: Telemetry-Based Racing Analysis and Coaching Kit in Sim Racing Games
Abstract
This paper presents TRACK (Telemetry-Based Racing Analysis and Coaching Kit), which is a framework for analyzing driving performance in sim racing and profiling how individual drivers behave behind the wheel. We report this framework together with its limitations: we calibrate each clustering result against a null, and when one does not separate from chance, we say so. Instead of restricting ourselves to scoring drivers or sorting them into preset labels, we represent each recording session as a compact geometry in a four-dimensional behavioral space (speed, braking, strategy, and consistency), and we group these fingerprints by their similarity using unsupervised clustering. Over time, we have developed and refined this framework on the open Assetto Corsa Gym (ACGym) dataset. Our study suggests that corner types differ along a behavioral dimension that was not used to define them. It also suggests that when the car changes, only speed and consistency carry over in the restricted population, while repeatability could not be shown there for any of the braking or strategy measures. Cluster separation becomes less distinct as the range of available telemetry widens. Until that repeatability is shown, grouping on the braking and strategy dimensions cannot treat the car as interchangeable, which divides an already small sample into smaller cells. It is also not clear whether a driver's grouping carries over from one corner type to the next. We also normalize each metric against a reinforcement-learning reference agent. The reference does not depend on the sample, so the scale does not shift when the sample does. We intend these results as an analytical foundation for a personalized improvement suggestion system. The sample is small. The cross-car result changes when the sample is defined more broadly. These outcomes are preliminary.
cs.AI / 102 / 2610.09371
The Confidence Game: Strategic Miscalibration in Human-AI Delegation
Abstract
Calibrated uncertainty quantification is essential to ensuring AI agents are trustworthy and reliable. However, when agents seek to maximize user engagement or revenue, confidence reports may be strategically distorted, detracting from their informativeness. We formalize this problem in the Confidence Game: a repeated signaling game with imperfect monitoring in which an agent of unknown honesty and ability reports its confidence, and a user decides whether to delegate the task or complete it herself. The agent manages the tradeoff between manipulating signals and maintaining its reputation. We characterize the Markov Perfect Bayesian Equilibria of the two-period game and show that honest reporting is not an equilibrium, inflation is the unique best response once the agent is sufficiently myopic, and under-reporting requires that the user believe honesty to be a minority. We then place an LLM in the agent role, supplying it with its true probability of success so that any gap between what it knows and what it reports is attributable to incentives rather than to miscalibration. The model claims high confidence on 56% of tasks it has been told it will probably fail. This persists on real tasks, where it must estimate its own accuracy and causes miscalibration to increase while the agent's signal becomes less informative. Furthermore, we find that the LLM agent's decisions are coherent, but it systematically underestimates both how likely the user is to delegate and how secure its reputation is, resulting in less extreme behavior. Pricing the agent's reporting rule, we find that it destroys 68% of the gains from delegation, of which 71% is information the report no longer carries and no amount of user sophistication recovers. Overall, we establish confidence reporting under delegation as a strategic problem and provide a tractable basis for modeling, analyzing, and testing agent behavior.
cs.AI / 103 / 2610.09593
Structured pre-generation elicitation versus single-shot prompting in AI-assisted enterprise decision-making: a randomised online experiment
Abstract
Generative AI speeds, and mostly improves, professional work, but there is concern that users who delegate both the production and the evaluation of an answer may accept weak output and engage less with the underlying reasoning (cognitive surrender). Interventions proposed so far, such as unassisted practice or slowing adoption, sit outside the working task. We tested a different approach: an interactive metacognitive scaffolding layer (Cognistance, a prototype developed at the Oxford Centre for Impact Research (OCIR) that asks users to clarify context, choose a strategic direction and explain their reasoning before the AI generates a deliverable). Mean composite quality was 32% higher with the scaffold, with the same direction for every rater. Gains were largest for trade-off articulation and strategic coherence and absent for technical specificity. A large part of the aggregate effect reflected rescue of weak prompts: floor-scored (off-task) deliverables fell from 34% to 5%. Among participants whose own prompt already stated the data-localisation problem, the advantage was 21%. Treatment participants reported greater involvement and took about 2.4 minutes longer on average (10.46 minutes). Immediate recall scores were higher, which tentatively suggests better retention, but in this limited experiment, was not robust to sensitivity analyses. Self-ratings of quality did not track rated quality in either condition. Structured elicitation before generation improved the rated quality and task relevance of AI-assisted strategy documents at modest cost in time. Delayed retention, error detection and effects in live organisations are the priorities for the next stage of research.
cs.AI / 104 / 2610.09740
Healthy skepticism in AI: a data visualization research agenda
Abstract
Research in data visualization of artificial intelligence (AI) models has historically focused on enhancing trust through visual explanations of AI. The trustworthiness line of work was built at least partially on an assumption that humans were critical users unlikely to adopt AI technology. It is increasingly clear that human trust levels in AI span, in fact, a wide range from critical to over-reliant. There is an urgent need to support both trust and healthy skepticism in AI solutions. We argue that it is healthy for humans to adopt a skeptical view both on the results of AI models and on the use of such AI models. We share our thoughts on the rising phenomenon of over-reliance on AI models, the risks and opportunities in using AI models, and the role of data visualization in over-reliance situations where humans are not motivated to engage in critical thinking.
cs.AI / 105 / 2610.09891
A Scoping Review and Experimental Study on Reinforcement Learning from Human Feedback for Human-Robot Collaboration
Abstract
Human-Robot Collaboration (HRC) can facilitate mass customisation in Industry 4.0, with Reinforcement Learning from Human Feedback (RLHF) representing a promising approach for developing safe AI-based robots. Practical challenges remain regarding safety during AI development, human feedback quality, and bidirectional human-robot adaptation. We conducted a scoping review of RLHF in HRC systems, mapping methods that address these challenges. Following PRISMA guidelines, we screened 199 records and included 20 peer-reviewed publications (2020-2025) spanning multiple HRC domains. To our knowledge, this is the first review focused on the bidirectional, closed-loop design of RLHF. Our review found multiple feedback modalities enabling data collection in various feedback formats. Collected data can be integrated at different stages of AI training, resulting in a multi-step development process. Pilot experiments are commonly used to evaluate HRC systems based on both human and robot metrics. To empirically test a key gap identified in the review, we conducted a between-subjects VR experiment comparing system- and user-initiated feedback on robot proxemic behaviour for safe navigation. Using Bayesian models, we analysed the relation between the collected feedback and safety metrics: psychological safety (post-experiment questionnaire) and physical safety (inverse time-to-collision). Results show that user-initiated feedback captures perceived safety better than system-initiated feedback, indicating that feedback timing directly affects feedback quality. Our review and experiment findings show that RLHF relies on appropriate feedback methods to ensure AI safety in HRC, and future RLHF research should prioritise realistic HRC experiments evaluating the effects of feedback collection methods on relevant human and robot metrics.
cs.AI / 106 / 2610.10182
EEG and Eye-Tracking Evidence That AI Disclosure Shapes Face Evaluation
Abstract
AI-generated faces can be difficult to distinguish from real ones, leaving viewers to rely on source labels when judging an image. Yet prior work has made it difficult to separate the effects of what an image actually is from what viewers are told it is. We validated faces as AI-generated or human in an online study (N=169), then crossed actual source (AI, human) with label (none, Made with AI, Made by a human) in a lab study $N=30), recording event-related potentials (ERPs) and gaze. ERP responses were equivalent for AI-generated and real faces, but varied with the label: labels drew early attention (N2), while labels that conflicted with the face's actual source prompted re-evaluation of the face (P3). Affective processing and initial gaze orienting were unchanged, but labels altered visual exploration. We provide a validated stimulus set and evidence that attributed origin shapes face processing, with implications for disclosure design.
cs.AI / 107 / 2610.10463
How assigned AI use before class shapes active student engagement in class
Abstract
AI learning tools are rapidly entering classrooms, but evidence about whether they help students learn is mixed and rests mostly on test scores. Comparatively less research addresses whether the use of AI changes students' live learning behaviors in class. Here, we report the results of a preregistered field experiment with 759 MBA students enrolled in ten sections of a course, in which each student was randomly assigned two of ten class sessions to prepare for with a purpose-built voice-based AI discussion partner. After two uses of the AI discussion partner, students made about 31% more voluntary contributions in each later class session. Students who used the AI discussion partner more also reported greater comfort speaking up and greater perceived learning, but not greater focus or motivation. These findings suggest that repeated practice with a voice-based AI partner can meaningfully increase students' engagement in class discussion, enhancing a critical intermediate learning outcome.
cs.AI / 108 / 2610.09361
From Chunks to Functional Evidence: Function-Aware Retrieval for EDA Documentation QA
Abstract
Retrieval-Augmented Generation (RAG) is widely used to ground answers in documents. For complex technical documentation, however, the primary bottleneck is often not model reasoning but a mismatch between a query and the way knowledge is organized for retrieval. This mismatch is pronounced in Electronic Design Automation (EDA) documentation, where the information needed for an answer is scattered across heterogeneous yet tightly coupled artifacts. We therefore redesign the basic retrieval unit of RAG. Instead of operating on isolated chunks or binary relations, we collect typed artifacts into EDA functional units. Each unit is recorded as a hyperedge with links to its source chunks. We then train an encoder to align queries with functional units and combine unit retrieval with direct chunk retrieval. After mapping the selected units back to their sources, a unified reranker chooses the evidence given to the generator. On the newly constructed EDADocEval-QA dataset, our method improves ROUGE-L by 37.1% over Chunk RAG and 55.6% over the strongest graph baseline. On the public ORD-MMBench benchmark, it improves ROUGE-L by 30.0% over the strongest baseline. These results support function-aware evidence organization in the evaluated EDA documentation settings.
cs.AI / 109 / 2610.10170
Does Document Structure Help Dense Retrieval? A Placebo-Controlled Ablation of Four Mechanisms Across Two Corpora
Abstract
Retrieval-augmented generation systems increasingly rely on document-structure treatments: structure-aligned chunking, LLM-generated chunk contexts, heading-path metadata, and hierarchical two-stage retrieval. Separate studies support each on different corpora, embedders, and metrics, and none control for a shared confound: any text prepended to a chunk perturbs its embedding. We present a mechanism-isolating ablation testing all four treatments under one protocol, matching chunk sizes across conditions and adding a semantically null placebo---heading paths that are structurally valid but shuffled across documents. We score retrieval with a coverage-aware nDCG and test four pre-registered contrasts via document-clustered bootstrap with Holm correction, on two distant corpora: 200 Wikipedia Featured Articles (951 queries) and 1,585 QASPER papers (4,303 questions). Organization helps, and the cause is content, not tokens: structure-aligned chunks with real heading paths beat contextualized fixed windows (+0.022 / +0.012 cov-nDCG@10) and the placebo (+0.010 / +0.016). Naive two-stage hierarchical retrieval hurts (-0.033 / -0.015), traceable to first-stage section recall. Gold structure beats LLM-induced structure on Wikipedia but not on QASPER. Effects are small ($dz$ 0.06-0.11) but Holm-significant and consistent across corpora.
cs.AI / 110 / 2610.09870
Deadline-Aware Multi-Agent Reinforcement Learning for TSN-Based Vehicular Edge Networks
Abstract
Vehicular edge computing (VEC) enables latency-sensitive applications by bringing computing and networking resources closer to vehicles. However, existing approaches often overlook network contention among co-located services with heterogeneous and dynamic latency requirements. While time-sensitive networking (TSN) provides bounded-latency communication, conventional and reinforcement learning-based schedulers struggle to adapt to highly dynamic vehicular environments and inter-queue dependencies. To address these limitations, we propose a multi-agent reinforcement learning (MARL) approach for queue-level scheduling in TSN-enabled VEC. Each TSN queue is assigned an autonomous agent that jointly learns the queue service order and time-slot duration to minimize deadline misses under speed-dependent latency requirements. We employ multi-agent proximal policy optimization (MAPPO) to enable coordinated yet autonomous scheduling decisions. Evaluation against single-agent, multi-agent, and non-learning-based baselines shows that MAPPO provides robust performance across different traffic profiles. Compared with centralized single-agent methods, it reduces service latency by up to 66.2% and improves reliability by up to 271.8%. Furthermore, unlike urgency-based heuristics, MAPPO ensures balanced scheduling while achieving lower inference times compared to other MARL methods.
cs.AI / 111 / 2610.08933
Toward Evidence-Driven Human-Agent-Robot Teaming for Earth-Independent Anomaly Triage
Abstract
Deep-space crews cannot rely on real-time ground support for urgent off-nominal events. Initial alerts may underdetermine cause, while discriminating evidence may reside in crew observations or at locations that are unsafe, costly, or unavailable for crew inspection. We present an evidence-driven architecture for human-agent-robot teaming in Earth-independent anomaly triage. Agentic AI is treated as a stateful coordinator over bounded, inspectable services rather than as a fully autonomous vehicle controller. A triage state manager maintains hypotheses, evidence provenance, uncertainty, operational context, and tool status; a crew-facing embodied agent elicits observations and explains assessment changes; and a mobile robot acquires targeted, localized evidence. Typed interfaces separate dialogue and orchestration from monitoring, robot command, context retrieval, and safety-critical control. Two scenarios illustrate the architecture: a crewed deep-space mission based on an actual ISS ammonia false alarm, where suspected contamination restricts crew access, and a power-interface anomaly at a crewed lunar base, where robotic inspection distinguishes a local connector fault from other causes ambiguous in remote telemetry. Our main contribution is an authority-bounded closed evidence-loop architecture, exercised in a hardware-in-the-loop integration prototype using Reachy Mini and an Innate MARS mobile robot.
cs.AI / 112 / 2610.08995
PhysEvo: Astra Can Act, Let It
Abstract
Astra can act, yet reliable manipulation depends on the system through which it observes and controls the world. We introduce PhysEvo, a framework for physical recursive self-improvement (RSI) around a single frozen model. A task agent executes robot tasks; a meta-agent uses the resulting trajectories to diagnose failures, revise tools and skills, and test corrections. The meta-agent can also improve its own diagnostic tools, so retained revisions support both later action and later self-improvement. This process develops joint-level control, evidence-seeking observation, and reusable manipulation skills without model-weight updates or a separately trained action policy. Across 42 RoboDojo tasks, held-out-layout evaluation of retained task-specific deployment versions yields a five-dimension average score of 68.14/100 and 62.00% success, compared with 47.17% for RoboDawn's one-shot Astra agent, the strongest published reference in our comparison. On eight manipulation tasks challenging direct Astra, PhysEvo achieves 55.00% success, compared with 1.25% for the direct-Astra reference. Deploying the simulation-evolved harness on AgileX PiPER and continuing skill revision yields 90.60/100 average score and 84.00% success across 25 trials on five real-world tasks. PhysEvo turns the consequences of action into persistent, testable changes to how a frozen model acts and improves.
cs.AI / 113 / 2610.09055
MimicX: Policy-in-the-Loop Supervision Refinement for Video-Driven Humanoid Motion Tracking
Abstract
Human videos provide rich motion targets for humanoid learning, yet visually plausible references can still produce persistent failures under physics-based execution. These failures reveal where training supervision should change. We present MimicX, a policy-in-the-loop framework that uses execution feedback to refine video-driven humanoid motion tracking. Starting from reconstructed and retargeted motion, MimicX localizes difficult transitions and affected body regions, then jointly adapts tracking objectives and the reset curriculum for policy continuation. Repeated rollout verification selects execution-priority improvements subject to tracking guards. Across four core video tasks, MimicX consistently improves tracking accuracy and Robust Execution Horizon relative to the Fixed Reference baseline. Task-averaged results show a 25.7% reduction in body-tracking error and a 255.6% increase in execution horizon. Additional video, motion-reference, and collision-scene studies evaluate the method beyond the core tasks, while MimicX-HLoop accelerates feedback through heterogeneous execution. Overall, MimicX turns policy failure into actionable supervision for deciding what to refine and which refinement to retain.
cs.AI / 114 / 2610.09496
Sparse Feature Policy Unlearning Mitigates State Hallucination in Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation by leveraging rich representations from pretrained vision-language models. However, their deployment in real-world environments remains limited by recurring unreliable behaviors. In this work, we study state hallucination, a recurring failure pattern in which a VLA continues acting as if an unrealized robot-object state had been achieved. Our analyses find that state hallucination coincides with weakened attention to task-relevant visual regions, and a mechanistic interpretation via sparse autoencoders reveals that hallucination-associated sparse features are activated when these failures occur. Based on this analysis, we propose SOUL (Sparse feature pOlicy UnLearning), which selectively unlearns policy knowledge associated with state hallucination behaviors, where sparse features identified from hallucination failures and successful behaviors serve as explicit forgetting and retention targets, respectively. Experiments across VLA architectures in simulated and real-world environments show that our method substantially reduces hallucinated failures and improves task success without substantially compromising the existing manipulation capabilities. These results suggest that interpretable feature analysis provides a practical basis for selectively modifying undesirable knowledge in robot policies.
cs.AI / 115 / 2610.09520
Not All Uncertainty Matters: Simulation-in-the-Loop Fast-Slow Reasoning for Decision-Critical Autonomous Driving System
Abstract
Large vision-language models (VLMs) provide powerful open-world perception and reasoning for autonomous driving, but their high computational cost and inference latency make continuous cloud-side use impractical. This motivates fast--slow collaboration, where efficient onboard modules handle real-time perception and control while cloud models provide high-level reasoning only when needed. The key challenge is deciding when cloud reasoning should influence time-critical driving decisions. Existing methods often rely on perception uncertainty, heuristic triggers, or resource-driven policies, without assessing whether resolving an uncertainty will improve planning. We propose \textbf{SIGMA}, a simulation-in-the-loop framework for task-oriented fast--slow collaboration. SIGMA embeds the planner into uncertainty assessment and evaluates how plausible scene realizations under semantic and geometric uncertainty affect feasible trajectories and planning cost. Based on these outcomes, it estimates the expected reduction in planning cost from resolving uncertainty. We further introduce expected planning gain (EPG), a decision-level metric for cloud invocation, cloud-guidance integration, and request prioritization under deadline and resource constraints. Experiments in CARLA show that SIGMA reduces unnecessary cloud interactions while improving planning, efficiency, and navigation success in static and dynamic obstacle scenarios. Compared with fixed-period collaboration, SIGMA reduces unnecessary cloud interactions by 50\%, improves navigation success by more than 6\%, and cuts finish time by up to 26.2\% in dynamic scenarios.
cs.AI / 116 / 2610.10388
RoboQuest: Generalist Physical Agents that Search, Inspect and Test
Abstract
Recent advances in multimodal foundation models have made them capable generalist physical agents for a range of manipulation tasks. However, successful operation in an unfamiliar environment may require an agent to seek task-relevant information through interaction when it is absent from the observations: it may need to determine where a relevant object is, inspect an unobserved property, or discover the effect of an unfamiliar tool. We thus introduce RoboQuest, a benchmark for goal-directed embodied exploration, where agents must actively acquire task-relevant information through physical interaction, use the resulting evidence to adapt subsequent actions, and autonomously decide when to commit to task completion. RoboQuest comprises ten mobile manipulation tasks centered on three forms of uncertainty: search, manipulation-based inspection, and interactive testing. We evaluate five frontier multimodal agents through a common visuomotor interface, as well as a $π_{0.5}$ policy fine-tuned on the full-episode demonstrations we release. The best agent succeeds in only 23\% of the episodes, and the fine-tuned policy almost never succeeds. Isolated tests of the execution skills the tasks are built from, with the hidden information supplied, show that the agents can carry out most of the required actions, and our failure analysis attributes only a minority of the failures to execution. Our failure analysis further finds that the agents often stop exploring too early as they make decisions before observing the required evidence for task completion. We also find that agents rarely prevent or repair the disturbances caused by their exploration. Moreover, learning by trial and error remains difficult for most models.
cs.AI / 117 / 2610.10462
FoldBack: Self-Correcting Masked Generative Policy for Long-Horizon Garment Folding
Abstract
We present FoldBack, a self-correcting masked generative policy for long-horizon garment folding. Existing long-trajectory policies may continue after a missed or slipped grasp even when the garment has not reached the intended configuration. We structure FoldBack's recovery mechanisms around three inference-time decisions: when to refine and verify, how to roll back, and where and how to retry. FoldBack aligns refinement and grasp verification with pick-and-place events, returns the robot to a retryable pre-grasp configuration while preserving successful grasps, and selectively regenerates the failed segment and selected future actions while avoiding previous failed grasp locations. To our knowledge, FoldBack is the first editable full-trajectory policy to unify these decisions, enabling failed interactions to be detected, undone, and repaired before execution continues, without recovery demonstrations or base-policy retraining. Across 33 real garments from six categories, FoldBack achieves 75.2% final folding success and 0.837 final-mask IoU, versus 45.7% and 0.689 for the strongest prior baseline.
cs.AI / 118 / 2610.10528
Long-WAM: Scaling the Context of World-Action Models
Abstract
Real-time robot control demands enough visual history to infer motion and task progress, but processing that history can delay action. We present Long-WAM, a model-system framework for scaling the context of causal world-action models under real-time control constraints. Our central finding is that access to history is not the same as using it: longer histories pay off far more when the video foundation is pretrained autoregressively (AR). We first learn causal prediction from robot and egocentric videos without action labels, then preserve this history-to-future structure during world-action adaptation. On RoboCasa GR-1, increasing context from 0.0 to 19.2 seconds raises success from 63.3% to 78.7%, whereas a bidirectionally pretrained initialization shows no net gain; robot-domain AR pretraining further raises peak success on GR-1 and LIBERO-Long. Long-WAM also achieves the best results among compared methods on LIBERO-Long, RoboTwin 2.0, and DOMINO. Streaming observation encoding, asynchronous execution, and hardware-specific acceleration enable deployment on RTX 5090, DGX Spark, and Jetson AGX Thor without dropping future prediction; on RTX 5090, each action chunk, including future-video latent prediction, takes 107.4 ms. Real-time deployment on Unitree G1 and YAM supports dynamic and long-horizon manipulation, including 95% success on dynamic cup stacking, where Pi0.5 and Fast-WAM succeed in none of 20 trials. As a memory-informed executor, Long-WAM also complements higher-level planning in composite tasks.
cs.AI / 119 / 2610.10355
MIRA: A Musical Intent Refinement Agent for Aligning Text-to-Music Generation with User Intent
Abstract
Text-to-music systems produce increasingly convincing audio, yet evaluation reveals little about whether the result matches user intent. A global text-audio relevance score can overlook the implicit intent in underspecified prompts and mask failures in specific requirements, such as instrumentation, structure, rhythm, or mood progression. To bridge this gap, we formulate text-to-music intent alignment as satisfying a per-request rubric of independently verifiable items covering both a request's explicit requirements and its implied musical intent. Scoring items individually makes evaluation diagnostic by intent source and musical dimension, rather than a single opaque score. We instantiate this as MuRA-Bench, a benchmark of real-world platform requests curated by music experts. We further propose MIRA (Musical Intent Refinement Agent), a test-time agent that first grounds a request's intent into rubrics, then searches over prompt revisions for a black-box generator under a bounded budget, iteratively generating music, verifying it against the rubrics, and using this feedback to guide a trajectory-aware tree search. Experiments across open-source and commercial backends show that MIRA improves intent alignment, enabling an open-source generator to achieve performance comparable to representative commercial systems (e.g. Suno and Mureka). Project page: https://mirareview.github.io/.
cs.AI / 120 / 2610.09938
Fast holographic inversion of superconducting domes
Abstract
A holographic superconductor whose scalar mass depends on the gauge field strength, $M(\Fsq)$, reproduces a superconducting dome for a suitable $M$, and recovering that $M$ from a given dome has so far taken days for a single training run. We propose a new way of training this model, with which an inversion takes from about ten minutes to an hour. Training needs the gradient of the condition that fixes the critical temperature, which the earlier method obtains by finite differences, repeating the bulk integrations for every training parameter. Here that condition is obtained, without any fit, from two integrations started at the horizon and at the boundary, and its derivative with respect to $M$ is an integral over the same two solutions, so the gradient needs no integration of its own. We use the speed to study the part of $M$ that a dome cannot determine, on the interval between the value $\Fsq$ takes at the horizon for the lowest doping and $\Fsq=0$, at which $M$ is the scalar mass $M(0)$ that fixes the dimension of the dual operator. We hold the scalar mass at several values, which we call pinned masses, retrain everything else at each, and find that the reconstructions agree wherever the horizons of the dome reach, including the minima of $M$, and differ only on that interval. A rule that keeps the reconstruction with the simplest closed form recovers both the scalar mass and the mass function of a test dome. On Gaussian and double-Gaussian domes and on the measured phase diagrams of YBa$_{2}$Cu$_{3}$O$_{y}$ and 2M-WS$_{2}$, however, the pinned mass it keeps rests on ties or on narrow margins, so for these targets the scalar mass is left open. The dome thus constrains $M$ where its horizons reach, and fixing the dimension of the dual operator needs a second observable.
cs.AI / 121 / 2610.09770
Artificial intelligence pathways from weather to climate
Abstract
Deep learning has made rapid advances in weather forecasting: autoregressive models trained on atmospheric reanalyses now rival dynamical models across nowcasting, medium-range, and subseasonal-to-seasonal lead times, producing well-calibrated ensemble forecasts at reduced cost. We review these advances and consider their extension to climate horizons, where the challenge shifts from initial-condition skill to producing reliable statistical responses under altered forcings. AI-powered climate prediction systems must produce credible forced responses to drivers (e.g., greenhouse gases, land-use change) typically outside the observed record. We propose two minimum requirements for AI in climate modeling: (i) external forcing agents must enter explicitly enough to support interventions in which they vary independently; and (ii) robustness must be stress-tested in out-of-distribution regimes, including extremes and counterfactual trajectories. Using leading AI autoregressive emulators and hybrid physics-AI models, we identify development and coupling challenges. Comparing the reported throughput of these models with that of GPU-ported dynamical models highlights how AI can reduce time-to-solution by advancing only the target variables at the required resolution and using longer time steps, rather than integrating a full high-frequency, multivariate state. Diverse AI downscaling strategies can partially substitute for explicit fine-scale resolution, paving the way toward inexpensive local hazard assessment across prediction horizons.
cs.AI / 122 / 2610.09663
Real-world application of deep learning in large-scale seismic interference attenuation: A case study in the Camie field of Angola
Abstract
In marine seismic acquisition, seismic interference (SI) occurs when energy from nearby external seismic source(s) is captured. It typically appears as coherent noise with linear or non-linear movement and varying amplitudes across different sail lines. SI is commonly observed and poses a challenge for seismic data processing. We present a case history of a previously proposed deep neural network (DNN)-based workflow applied for SI attenuation across a marine seismic block in the Camie Field of Angola. This field survey covers over 345 km2 and is marked by the challenge of multiple SI types. The employed DNN-based workflow performs SI attenuation in the common shot domain based on a supervised learning framework: a small subset of the SI-contaminated data was first processed by a conventional geophysical algorithm to obtain an estimate of the SI noise, which was then manually blended with the SI-free common shot gathers from the same survey to generate the training pairs. To ensure signal fidelity, several techniques were applied to improve the DNN's performance. A key highlight of this case history is its scale: this represents a real-world, large-scale processing project and we present a comprehensive comparison of the DNN-based workflow with the conventional geophysical algorithm across the entire survey block, focusing on both processing quality and processing time. The results demonstrate the outstanding performance of the employed DNN-based workflow, which achieved higher SI removal accuracy, with less signal leakage and more complete SI removal. The promising results of this application also open up possibilities for integrating deep learning into other seismic denoising tasks. In addition, we discuss the limitations of this case history, aiming to provide insights for future research and applications in the field.
cs.AI / 123 / 2610.09871
An AI-assisted conditioning and geological interpretation workflow for usage in implicit geological modeling
Abstract
Implicit modeling and Relative Geologic Time are geological modeling techniques that enable more efficient, faster, less biased and more reproducible modeling results. For optimal operation, these techniques require many well-constrained input data. In the framework of the Horizon Europe GO-Forward and MOOI WarmingUP GOO projects and to accelerate Implicit modeling, Machine Learning (ML) methods have been tested and implemented in a toolkit for the interpretation of (onshore) seismic data from the shallow to deep range (+- 300 - 3500 m). The goal is to rapidly characterise this depth domain by efficient interpretation of horizons and faults in seismic data. The first step is to improve the signal by applying AI techniques like self-supervised and semi-supervised contrastive learning CNN's for noise reduction and interpolation. Next, horizons and faults are interpreted with minimal use of human-generated training data by using (semi-) self-supervised methods. The resulting developed toolkit supports the application of the implemented algorithms in an efficient workflow. As a first demonstration, the top of the Dutch Maassluis Formation has been interpreted in the Leeuwarden and Waalwijk 3D seismic cubes. Overall, this study demonstrates that AI-assisted interpretation workflows have reached a level of maturity that allows their integration into applied geological modeling and decision-making.
cs.AI / 124 / 2610.08937
A Shortcut to Structure in AlphaFold 3
Abstract
AlphaFold 3 predicts protein structures with remarkable accuracy, yet how structural information emerges within the model remains poorly understood. Here, through causal interventions on internal representations and direct probing of every Pairformer block, we trace the formation of global protein geometry and identify the multiple sequence alignment (MSA) as a structural shortcut to the fold. Removing the MSA largely preserves local secondary structure while disrupting the long-range relationships that define global topology. Restoring the MSA-enriched pair representation at only forty residues recovers most of this lost organization, including at pairs never directly modified. This contribution depends on the detailed direction of the MSA module's output rather than its magnitude. The Pairformer rapidly converts this signal into global geometry: the final fold becomes recoverable by approximately block 9 of 48 for a majority of proteins, roughly twenty-seven blocks before the model's decoder can render it, whereas without the MSA it remains inaccessible for most proteins throughout the pass. Which homologs are supplied shapes this trajectory more strongly than which query is supplied; it persists for a designed query that never evolved but collapses for a shuffled sequence. Most importantly, an alignment built for a different protein that shares the fold, supplied only at the structurally corresponding columns, raises the median TM-score against experiment from 0.44 to 0.72, while the same alignment shifted a few residues along the chain performs worse than supplying no alignment at all. What AlphaFold 3 reads from an alignment is therefore a description of the fold itself, transferable between proteins that share one, rather than the query's own evolutionary history. This explains both its accuracy and the limits of what it has solved.
cs.AI / 125 / 2610.09643
CircuitATLAS: Agentic reasoning over a systems neuroscience knowledge graph for target discovery in circuitopathies
Abstract
Drug discovery for neurological disease has traditionally centered on the molecules altered by disease. But the molecules that cause pathology are not necessarily the best points from which to reverse it. Here, we ask which otherwise unaltered molecular control points can be engaged to restore pathological neural circuits toward functional states. We present CircuitATLAS, a provenance-grounded systems-neuroscience knowledge graph and agentic framework for target discovery in circuitopathies. It structures literature-derived relationships across diseases, phenotypes, electrophysiology, circuits, brain regions, cell types and molecular effectors, while deliberately excluding direct disease-gene and disease-protein edges to reduce shortcut reasoning. The graph contains 3.83 million nodes and 7.66 million edges, including 5.31 million LLM-extracted relations, and incorporates structured datasets such as the Human Cell Atlas and new multimodal in vivo measurements. We then introduce an agentic workflow that reasons from measurable disease phenotypes through their circuit and cellular substrates to molecular interventions, therapeutic feasibility and clinical constraints. Finally, we introduce a human-governed in vivo lab-in-the-loop linking hypothesis generation to experimental iteration. Within this framework an agent nominated ATP1A3, the neuronal alpha3 Na+/K+-ATPase, as a control point on cortical excitability; interneuron-restricted expression of ATP1A3 abolished the beta- and gamma-band response to a focal 4-aminopyridine challenge in vivo, and the validated target was then carried into a structure-guided small-molecule campaign terminating in a defined assay to resolve the direction of modulation. CircuitATLAS thus provides a framework for discovering therapeutics based not only on what is molecularly disrupted in disease, but on what can be controlled to restore circuit function.
cs.AI / 126 / 2610.09043
Careful Judge: Safe and Efficient Human-AI Collaborative Decision Making
Abstract
In human-AI collaborative decision making, human review can prevent unsafe AI decisions, but each human judgment is costly. Treating human intervention after AI abstention as a one-off fallback misses the opportunity to improve future AI decisions for greater automation, yet AI adaptively learning from selectively queried human feedback breaks safety guardrails calibrated for old models. We approach this challenge with CARE---calibrated adaptive rectification and escalation---an end-to-end pipeline that combines AI models and human reviewers to guarantee safe, human-aligned decisions, while continuously learning from human feedback to achieve greater automation with fewer human queries. CARE is principled, general, modular, and works with any black-box AI model. Our novel adaptive calibration module guarantees risk control at every time step for any rectification module. We further show how CARE improves query efficiency when the AI model is well trained and the human-AI misalignment has a clear structure. Experiments on four safety-critical real-world datasets spanning driving, language, and robotics demonstrate that CARE achieves human-aligned decisions while reducing human queries by 25-81% relative to baselines.
机器学习 (cs.LG)
234
cs.LG / 1 / 2610.08985
A method for multimodal analysis of TAIGA experiment data using essential features
Abstract
The aim of processing and analyzing experimental data from physical experiments is to obtain physically significant information about the phenomenon under study. This goal is achieved by multi-stage processing of experimental data, during which noise associated with measurements is suppressed and the dimensionality of the input data is reduced. In this paper, we propose a new method based on the use of neural networks such as autoencoders to extract essential features. The special value of the proposed approach lies in the possibility of its application to the analysis of multimodal data received simultaneously from several installations. We will apply this approach to a multimodal data (MMD) of the experiment TAIGA. Currently, the analysis of the MMD is carried out independently for each installation separately. Therefore, the development of methods for the joint analysis of MMD from TAIGA-type installations is an urgent task in cosmic ray physics and gamma-ray astronomy. Based on Monte Carlo simulation, it is shown that the proposed method allows for effective MMD analysis. It can also be used for MMD analysis at other experimental complexes.
cs.LG / 2 / 2610.09837
Origins of Universal Machine Learning Force-Field Errors in Multicomponent Materials
Abstract
Universal machine learning force-field generalization to multicomponent environments generated by compositional design remains insufficiently assessed. We construct a benchmark of 7,599 multicomponent configurations inspired by high-entropy design, elemental substitution and anion mixing. Eleven pretrained models are evaluated against density functional theory for energies, forces and stresses, with assessment extended to elastic, vibrational and adsorption-related properties. Force errors are analysed through training-reference coverage, local geometric heterogeneity, distance directionality and elemental response. Distances to training-reference environments reveal a qualitative association between coverage differences and increasing errors, while substantial variation remains at similar distances. Higher-error groups show greater local geometric heterogeneity, although OMat24 provides broad coverage of these environments. Relative to training-reference pair medians, errors remain low near the median, rise steeply on the compression side and increase more weakly on the extension side. After matching element pairs and absolute distance deviations, compression-side force errors are 1.81-1.95 times extension-side errors. Model-predicted pairwise interaction curves show greater curvature under compression. Fitting difficulty in independent elemental systems correlates with electronic band-energy responses to atomic displacements and Fermi-level shifts, and a similar pattern is observed in multicomponent systems. In parameter-matched comparisons, spherical-harmonic representations with maximum degrees of 2 and 4 lower test force errors for 38 and 40 of 43 elements, respectively, while differences in elemental difficulty remain. These findings inform force-field selection for experimental compositional design and identify targets for training-data sampling and model representations.
cs.LG / 3 / 2610.10140
Progress and Prospect of AI in ARPES Workflow
Abstract
Artificial intelligence (AI) is becoming an increasingly useful tool across the experimental sciences, including angle-resolved photoemission spectroscopy (ARPES), which routinely produces large, multidimensional datasets of electronic structure. Recent advances in AI and machine learning (ML) have opened new opportunities across the entire ARPES workflow, from automated sample preparation and real-time data acquisition to post-experiment data analysis and comparison with theoretical calculations. Despite this progress, a comprehensive review of ML applications, their capabilities, and reliability across the different stages of ARPES workflow is still lacking. In this review, we first introduce ML methods that are most relevant to experimentalists working in condensed matter physics and materials science. We then follow the ARPES workflow, reviewing existing ML applications at each step and discussing their advantages, limitations and potential for future development. We also examine the current ARPES data landscape, where several open databases are available but remain relatively small and fragmented compared with large, shared datasets such as ImageNet. Given these limitations, we suggest that the community focus on sharing pretrained models that can be further trained, adapted to specific tasks, and redistributed, while working toward a larger and standardized open ARPES dataset repository. Finally, we discuss our perspectives on the future of AI within the ARPES workflow using a six-level framework of laboratory automation, highlighting the opportunities and challenges in moving toward a fully autonomous, self-driving ARPES laboratory.
cs.LG / 4 / 2610.09104
On the Computational Complexity of Hidden Markov Model Identification
Abstract
Identification is the task of recovering the parameters of an unknown ground-truth model from sampled data. When parameters other than the ground truth induce the same output distribution, data alone does not provide enough information to recover the ground truth, and the model is thus called unidentifiable. We study the identifiability problem for hidden Markov models (HMMs): given an HMM, is it identifiable? Existing work on HMM identification establishes conditions under which the ground-truth HMM can be identified. However, most of these conditions are sufficient but not necessary, meaning that, when a model does not satisfy them, its identifiability remains inconclusive. We instead take a computational perspective: is there a sound and complete algorithm that decides whether a given HMM is identifiable, and if so, what is the complexity of this decision problem? We consider the decision problems arising from the various notions of identifiability in the literature, including deterministic, generic, global, local, state-permutation- invariant, and finite-alphabet identifiability. We show that all of these problems are decidable in PSPACE, via reductions to the theory of the reals at various levels of its quantifier-alternation hierarchy. We further show that the deterministic variants are already coETR-hard (and hence coNP-hard) for simply parameterized families.
cs.LG / 5 / 2610.10038
Evolve on the Host, Predict on the Edge: Deploying Online Neuroevolutionary Architecture Search for Cross-sectional Stock Return Prediction
Abstract
Accurate forecasting models are usually large, expensive to update online, and fixed in architecture once trained. We apply ONE-NAS, an online neuroevolutionary architecture search that evolves a population of small recurrent networks as each window of data arrives, to daily cross-sectional stock return prediction, and pilot it on a host and endpoint pipeline: the host runs the search and ships each generation's champion genomes over TCP/IP to a Raspberry Pi 4B, which predicts online. On the Pi a single champion predicts a 50-stock window in 24.6~ms and the ensemble of 40 island champions in 556~ms, far inside the daily decision cycle. On four panels of US mid-cap equities over 2022--2024, reading the population as a rank-mean ensemble of island champions returns $+27.5\%$ net of realised transaction costs, against $+11.3$ to $+14.8\%$ for online LSTM, online GRU and monthly-retrained LSTM baselines and $+4.5\%$ for the single best genome used in prior ONE-NAS work.
cs.LG / 6 / 2610.09031
Beyond Explanation: Debugging Medical Imaging Models via Concept Intervention
Abstract
Medical imaging models often operate as black boxes, limiting interpretability and systematic debugging. We introduce an easy-to-use, plug-and-play framework for concept-based interpretation and model refinement. By aligning a single-modality encoder to BioMedCLIP, we construct a Concept Bottleneck Model (CBM) that enables concept-level interventions. These interventions allow us to isolate causal versus spuriously correlated concepts, validate insights with domain experts, and generate counterfactual samples for targeted fine-tuning. We evaluate our framework on a Mayo Clinic ultrasound dataset and the CheXpert 5x200 chest X-ray dataset. Results demonstrate that concept intervention enables reliable model diagnosis while maintaining, and occasionally improving predictive performance via guided fine-tuning. Our findings highlight the practical value of this framework for controlled, interpretable refinement of clinical deep learning models.
cs.LG / 7 / 2610.09274
Hardware-aware Calibrated Clustered Attention for Efficient Visual Geometric Transformers
Abstract
The Visual Geometry Grounded Transformer (VGGT) marks a significant leap forward in 3D scene reconstruction, as it is the first model that directly infers all key 3D attributes (camera poses, depths, and dense geometry) jointly in one pass. However, this joint inference mechanism requires global attention layers with extremely long sequences that causes a significant latency bottleneck. In this paper, we propose blockwise clustered attention (BC attention) to accelerate the global attention layers in VGGT. By limiting the clustering within HW-friendly neighborhood blocks, BC attention reduces the computation overhead of query clustering as well as the costly data movement between on- and off-chip memory. This enables BC attention to scale to long sequences and deliver practical latency improvements on GPUs. Moreover, we introduce a hashing hyperplane calibration method and a threshold-based error compensation method to reduce clustering errors efficiently, which is a bottleneck in the current clustered attention mechanism. Overall, our experiments on GPU demonstrate that calibrated BC attention accelerates the global attention layers by 2.10-2.63$\times$ and the whole backbone by 1.77-2.35$\times$ with negligible loss (1%) for large scenes. With a small performance loss (< 5%), calibrated BC attention further achieves a 2.26-2.87$\times$ latency improvement on the global attention layers and a 1.90-2.55$\times$ improvement on the backbone.
cs.LG / 8 / 2610.09459
An Invariant Tangent-Angle Descriptor and a Band U-Net for 2D Fragment Adjacency Prediction
Abstract
This paper addresses the prediction of adjacency between pairs of 2D fragments based on their contours. We improved the two-stage architecture proposed in Beaulac's thesis, in which a rotation-equivariant Siamese convolutional neural network scores pairs of local image windows along the two contours of two fragments. The scores are gathered in an adjacency matrix in which a ResNet detects the partial anti-diagonal band that reveals the adjacency of two fragments. In the current work, we keep the pipeline and replace the local score by a comparison of tangent-angle profiles of contour windows, making it, by construction, invariant to fragment rotation and agnostic to the selected contour-starting point. These adaptations may be either a training-free likelihood ratio or a small one-dimensional convolutional model trained on corresponding points. We also replaced the final classifier by a band U-Net that segments the band and classifies the pair, so that the shared arc is obtained along with the decision. In the synthetic data set of the original thesis, the tangent descriptor performs as well as or better than the image-window approach in all tested configurations. The proposed pipeline reaches an accuracy of 98%, vs 93% to 95% for the original approach once its evaluation is corrected. We tested our pipeline, with models trained only on synthetic data, on the PairingNet benchmark, and obtained an AUC of 0.93. Furthermore, under the PairingNet pair-searching protocol conditions, our learned descriptor obtains a Recall@10 of 0.82 on the real set against 0.56 from the best model of the original paper.
cs.LG / 9 / 2610.09509
TERRA: Learning Transportable Latent Actions through Temporal Effect Representation and Relational Alignment
Abstract
Latent actions supervise robot policies with action-like codes inferred from visual transitions, and their usefulness hinges on two questions: what a code keeps from a transition, and whether it still means the same thing when reused in a different initial state. The first is a tension in time: an endpoint difference discards how motion unfolds, while the full sequence admits nuisance variation. The second is left open by reconstruction, which only ever observes a latent together with the state it came from. We argue that both questions can be answered in the same place. TERRA (Temporal Effect Representation and Relational Alignment) describes a transition by a compact temporal effect, its net feature change together with a low-order within-window dynamics component, and learns a continuous latent from this effect. The same effect space then serves as the reference for reuse: Effect-Anchored Transport (EAT) decodes a latent in other initial states and anchors the resulting effect to the one observed at its source, so that the latent is shaped by what it does across contexts rather than only by the transition it came from. With frozen linear readers, TERRA predicts actions more accurately than UniVLA and a LAPA-style baseline, degrades more slowly under visual distractors, and keeps transported transitions faithful to the donor action as the recipient context moves farther away; a same-budget control shows that these gains come largely from EAT. At matched pretraining scale, the complete system reaches 93.4% average success on LIBERO, compared with 91.8% for UniVLA.
cs.LG / 10 / 2610.09517
It Is Not Seeing the Hazard: A Frozen Vision-Language Safety Score Measures Its Caption Bank
Abstract
Frozen vision-language models increasingly provide safety signals for reinforcement learning. Their use assumes that similarity to language describing danger indicates the hazard itself. Yet policy return and collision rate cannot reveal whether a score detects hazards or responds to correlated features of the scene. VLM-based methods have reported gains in driving and safe-RL benchmarks by converting image-text similarity into rewards, costs, or confidence weights. Such signals promise to reduce reliance on manually designed feedback. They may also reflect prompt structure, embedding geometry, or camera viewpoint, leaving their safety meaning unverified. To address this gap, we present a controlled evaluation of a frozen CLIP prompt-margin safety score. We apply the score to trajectories generated by policies that never receive it, match pre-contact observations to contact-free observations with comparable hazard geometry, and vary the captions, encoder, and camera view. Across three policies, 180 episodes, and 130 isolated contact onsets, the score decreases for about twenty steps before contact. Mechanism controls indicate that the score mainly tracks resemblance to the scene shared by its captions and changes with caption separation and camera view. A constant-confidence control retains the lower catastrophe-rate point estimate, so policy gains do not establish hazard perception.
cs.LG / 11 / 2610.09860
DeepTopoClustering: Unsupervised Derivation of Surface Process Taxonomy from 4D Point Clouds for Topographic Monitoring
Abstract
4D point clouds acquired by permanent laser scanning (PLS) enable accurate high-frequency monitoring of surface change in dynamic topographic environments. However, existing methods remain limited in organizing detected surface activities into meaningful process types. We propose DeepTopoClustering (DTC), an unsupervised framework for deriving a hierarchical process taxonomy from object-based surface activities, so-called 4D objects-by-change (4D-OBCs). We transform each 4D-OBC into a GeoMorphogram, a distributional sequence representing the temporal evolution of topographic change within a spatially bounded surface activity. A convolutional autoencoder learns latent embeddings from GeoMorphograms, which are jointly optimized using a hierarchical deep clustering objective to organize surface activities into a hierarchy. We evaluate the learned hierarchy using expert annotations on two 4D datasets of sandy beach sites and their combination. DTC with GeoMorphograms achieves the highest agreement with expert judgment at the taxonomy level comprising eight major process types ($F_1=0.78$, match accuracy $=0.92$), outperforming dimensionality reduction and conventional flat clustering. The learned taxonomy separates major erosion- and deposition-dominated activities and distinguishes finer subtypes based on change magnitude, duration, compactness, and temporal evolution. DTC thus provides a scalable and interpretable route from 4D change detection to a data-driven, expert-supported surface process taxonomy, advancing automated knowledge derivation for understanding surface dynamics in topographic monitoring.
cs.LG / 12 / 2610.09863
Global Average Precision for Representation Learning
Abstract
Standard information retrieval metrics, such as mean Average Precision (mAP), assess performance one query at a time, based on how the similarities between a query and its positives compare against those with its negatives. The same holds for common representation learning losses, such as InfoNCE and per-query AP surrogates. None of them considers whether similarities are comparable across queries, which any system with a single decision threshold relies on. Global Average Precision (gAP) does, by ranking all query-candidate pairs in one list and computing a single AP. We introduce gSAP, a differentiable surrogate of gAP. It needs only a similarity matrix and a binary matrix marking the positive pairs, the same input as existing losses, so it is a drop-in replacement for them and agnostic to the encoder, the modality, and the source of supervision. Since it considers all possible pairwise comparisons in the batch jointly, it also remains trainable at low temperatures, a regime where per-query surrogates run out of gradient. Swapping it into established recipes improves supervised metric learning, cross-modal alignment, and self-supervised pretraining, where, to our knowledge, it is the first ranking loss to replace the community standard InfoNCE in the latter two. Its similarities are more consistent across queries, which drives the gains under a universal threshold. gSAP retrieves up to four times as many positive pairs as the strongest AP surrogate at the same precision, and it degrades the least when queries with no positives in the database are added. Beyond thresholding, models trained with gSAP also learn better representations, with higher transfer, $k$NN and zero-shot classification accuracy.
cs.LG / 13 / 2610.10408
Rubix: Global Correspondence-Free Point Set Alignment through Assignment Geometry
Abstract
Procrustes-Wasserstein alignment jointly estimates a matching and rotation without supplied correspondences, but alternating minimization can stop at suboptimal solutions. Rubix solves the equally weighted planar problem globally under squared Euclidean loss. Each matching $σ$ of two centered $n$-point sets defines a complex correlation $z_σ=\sum_i\bar x_i y_{σ(i)}$. Their convex hull is the permutation polygon: supporting vertices give optimal matchings at fixed rotations, and the farthest vertex gives the global alignment. We prove the sharp bound of $n(n-1)$ vertices for $n\ge2$, answering Rote's rotation-assignment open problem. In exact arithmetic, assignment queries recover the polygon in $\mathcal O(n^5)$ operations. Assignment-based bounds extend the approach to three-dimensional rotations and partial matching at a supplied translation through branch-and-bound. On timed MPEG-7 shape pairs, Rubix attains every numerical reference value in 12 ms on average, 50 times faster than a rotation grid at the same accuracy. Its distances improve gravity-aligned matching of real 3D scans, shape retrieval and noisy crystal classification over alternating minimization.
cs.LG / 14 / 2610.09276
CATune: Structural Constraint-Aware Bayesian Optimization for DBMS Configuration Tuning
Abstract
Modern DBMSs expose hundreds of configuration knobs, resulting in a high-dimensional and heterogeneous search space that makes automated tuning costly. Existing ML-based tuning systems typically treat the configuration domain as box-constrained and rely on workload feedback to implicitly capture inter-knob relationships. However, DBMS documentation specifies deterministic knob dependency constraints, particularly ordering constraints, that characterize structurally valid regions of the configuration space. We present CATune, a constraint-aware Bayesian optimization (BO) framework that models deterministic inter-knob ordering constraints as structural components of the search domain. Instead of learning feasibility boundaries through sampled violations, CATune performs optimization within a constraint-consistent subspace. We develop a topology-aware sampling strategy that respects dependency structure during exploration and avoids the inefficiencies of post-hoc constraint handling. To enable automated constraint discovery, we further design a precision-first extraction pipeline that combines LLM-based parsing with reliability safeguards to mitigate hallucinated dependencies. Experiments on PostgreSQL and MySQL using TPC-C and TPC-H workloads show that CATune substantially improves both sample efficiency and final tuning quality across surrogate models and BO frameworks. Under default ranges, CATune reaches the baseline optimum up to 12.5x faster and improves throughput by up to 63.37%. The improvements persist under knowledge-guided reduced ranges and alternative optimization implementations. These results demonstrate that explicitly modeling system-defined deterministic ordering constraints enhances optimization robustness and system stability.
cs.LG / 15 / 2610.09424
Democratizing MoE inference on commodity GPUs with CoMoE
Abstract
Deploying Mixture-of-Experts (MoE) models relies heavily on Expert Parallelism, which generates intense inter-GPU communication. Consequently, state-of-the-art inference systems require high-bandwidth, P2P interconnects (e.g., NVLink) in datacenter GPUs to handle massive token routing, making deployment prohibitively expensive. Consumer GPUs offer comparable compute power at significantly lower cost, promising to democratize MoE inference for individuals and enable privacy-preserving local deployments. However, their bandwidth-limited (only weak PCIe bus bandwidth) and host-mediated interconnects (no P2P support) introduce severe communication bottlenecks. We present CoMoE, a communication-efficient MoE inference system that resolves this mismatch through novel host-centric routing. Our key insight is that the unique communication topology provides the opportunity to elevate the host to an active routing hub, which can fundamentally reduce communication volume and eliminate global synchronization-induced stalls. Specifically, for token dispatch, we introduce host-backed token multicast to write shared tokens to the host exactly once, eliminating outbound transmission redundancy. For token combine, we propose a fine-grained, token-level aggregation mechanism using host staging buffers, which replaces rigid global synchronization and mitigates straggler effects. Evaluation on RTX 5090 GPUs shows that CoMoE improves inference throughput by up to 1.46x, approaching the performance of NVLink-capable A800 GPUs at only 23.4% of the hardware cost.
cs.LG / 16 / 2610.10135
Attention via Black-Box Vector Search
Abstract
Sparse attention mechanisms estimate attention over $n$ tokens using a small subset of keys. Many existing approaches use maximum inner product search (MIPS) to retrieve the heaviest keys, which motivates the following question: given black-box access to a MIPS oracle, how many keys must be retrieved to output an $\varepsilon$-accurate attention estimate? We answer this question by unifying prior approaches through the framework of priority sampling. With a single MIPS index, we show that $Θ(\sqrt{n}/\varepsilon)$ retrieved keys are both sufficient and necessary. With $Θ(\log n)$ indices, we give an algorithm that retrieves only $O(\log n+1/\varepsilon^2)$ keys and prove that this is near-optimal. More generally, we design algorithms that establish a smooth tradeoff between the number of MIPS indices and number of retrieved keys. We then show that if we allow augmentation of keys and queries, we can bypass the above lower bounds: there exists a simple priority-sampling estimator using a single MIPS index and $O(1/\varepsilon^2)$ retrieved keys. When integrated into LLM inference, our algorithms outperform top-$k$ and sampling approaches used in prior work and yield attention approximation that scales favorably to long contexts.
cs.LG / 17 / 2610.10253
On the Cyclic Assumption of the Cow-Path Search Algorithm
Abstract
In the cow-path problem, a cow must find a goal lying at an unknown distance on one of $w$ paths connected only at the origin, and performance is measured by competitive ratio. Kao, Reif and Tate designed an efficient randomized algorithm in which the cow visits the paths in a fixed cyclic order. They proved the algorithm is optimal for $w=2$, and subsequently Kao, Ma, Sipser and Yin proved its optimality for all $w$, with a claim that no algorithm does better than the best cyclic one. This note provides a detailed proof of that claim.
cs.LG / 18 / 2610.09304
emg2face: Expressive Facial Animation with High-Density Surface EMG
Abstract
Facial movements convey subtle and important information that is critical for human social communication. Optical methods for face capture are difficult or impossible to use when the face is occluded by head-mounted devices (HMDs), such as VR headsets. Even with a clear line of sight, such methods raise privacy concerns and require head-mounted capture rigs that offset cameras and lighting from the face. We show that high-density surface electromyography (HD-sEMG) provides a viable non-optical alternative that addresses these challenges. We measured 64 EMG channels, using two textile EMG grids, with 32 from the forehead (typically occluded by an HMD) and 32 from the side of the face. EMG data were digitized at 2048 Hz and filtered. Facial movements were simultaneously recorded and used to estimate 478 3D facial landmarks using MediaPipe's Face Landmarker. A major challenge in such multimodal recordings is synchronizing EMG and video data, which have different sampling frequencies and independent clocks. We developed a novel synchronization method using analog audio bursts that is capable of sub-millisecond synchronization. We also developed a staged fitting method that fits a recent high-resolution parametric head model (GNM), with 253 identity blendshapes and 383 expression blendshapes, to the MediaPipe landmarks as participants performed different facial expressions. We trained a deep neural network comprising per-grid spatial encoders followed by a dilated temporal convolutional network (TCN) to predict blendshape parameters from HD-sEMG signals at 100 Hz. Once trained, the network can predict expression blendshapes solely from HD-sEMG recordings. The output can be rendered using standard real-time blendshape animation methods. We demonstrate the methods using recordings from 25 participants, and direct expression transfer to a variety of human faces and non-human characters.
cs.LG / 19 / 2610.09691
Optimal Regret for Online Market Making with Limit Order Book
Abstract
We study online learning in market making, where, at each round, a market maker posts bid and ask prices before observing the market price and the private valuation of an incoming trader. In this setting, Maran et al. 2026 introduce a feedback model motivated by limit order books, in which the trader's valuation is revealed only if no transaction occurs. Assuming that trader valuations are drawn i.i.d. from an unknown distribution while market prices are chosen adversarially, they establish an expected regret bound of $\widetilde{\mathcal{O}}(T^{2/3})$. In this work, we improve upon this guarantee by establishing a high-probability regret bound of $\widetilde{\mathcal{O}}(\sqrt{T})$. As a warm-up, we first consider the full-feedback setting. We introduce a discretization of the bid-ask space based on two coupled grids and combine it with Hedge to achieve the desired regret rate. Building on these ideas, we then address the substantially weaker feedback induced by a limit order book and develop an algorithm that achieves the same guarantee. Finally, we investigate the limits of learnability in fully adversarial environments, where the valuations may vary arbitrarily as well. Perhaps surprisingly, we show that when both market prices and trader valuations are chosen adversarially, sublinear regret is impossible even under full feedback, thereby motivating our stochastic assumption on the valuations.
cs.LG / 20 / 2610.10124
Training with Missed Targets in Generative Recommendation: Separating Supervision from Probability Competition
Abstract
Generative recommenders return a limited candidate set and may omit observed targets before reranking. A training strategy appends these missed targets to reranker training lists, although inference still ranks only original candidates. This operation simultaneously changes retrieved-target weight, adds supervision over appended targets, and makes the two groups compete for probability. An append/no-append comparison therefore cannot explain changes in returned-item rankings. We construct three matched losses that hold retrieved-target weight fixed while introducing appended-target supervision and group competition separately. The intermediate loss trains within both groups but normalizes them separately, preventing training-only targets from competing with inference candidates. Experiments with a released OneRec model and locally trained Amazon generators show that this competition can harm returned-item ranking. In four prespecified Amazon Video Games comparisons, removing it improved full-target normalized discounted cumulative gain (FT-NDCG) by 7.8--22.2\%; 95\% intervals over users and three of four intervals over training runs excluded zero. A conservative development-set rule selected appended-target training for two of three generators in one held-out category and rejected it for all three in another, avoiding a 1.7\% loss. Candidate completion should therefore be evaluated for each generator rather than applied automatically.
cs.LG / 21 / 2610.10362
Pathwise Information Certificates for Decentralized Adaptive Sensing
Abstract
We study decentralized adaptive sensing, where multiple agents choose measurements from evolving local beliefs while exchanging information over a communication graph. We ask whether the measurements actually selected by an adaptive policy have collected enough evidence to distinguish the true target from every plausible alternative. We develop a pathwise certificate based on the Rényi--Chernoff information accumulated along the realized sensing trajectory. It yields nonasymptotic MAP-error bounds and an anytime, network-wide stopping rule for arbitrary history-dependent sensing policies, while separating accumulated statistical information from a bounded network-mixing transient. Linear growth of the information against the least-resolved competitor implies exponential decay of MAP and squared-localization error. A classical pairwise KL converse, specialized to the adaptive decentralized transcript, shows that insufficient information on any pair prevents a positive uniform error exponent, confirming the hardest competitor as a fundamental bottleneck. Across policies, graph topologies, sensor profiles, and seeds, the worst-competitor score correlates more strongly with localization speed than an average-pair proxy in both 1D ($r=0.89$ versus $0.40$) and structured 2D sensing ($r=0.77$ versus $0.48$). Our results provide a practical way to certify and diagnose adaptive multi-agent sensing systems using the evidence they actually collect.
cs.LG / 22 / 2610.08960
Directed Temporal Representations for Offline Visual Control
Abstract
Predictive world models provide compact visual representations for control. Control requires a latent geometry aligned with temporal reachability rather than predictive similarity alone. We introduce Directed Temporal Representations for Control (DTRC), which learns such a geometry from offline visual trajectories on top of frozen LeWorldModel (LeWM) features. DTRC constructs a directed temporal quasimetric over the learned control representation. Short-range temporal offsets calibrate the distance scale. Bootstrapped targets extend temporal reachability across longer horizons. Action-conditioned consistency aligns the representation with local transition dynamics. The resulting distance estimates temporal reaching cost, and its change across a transition defines goal-relative temporal progress. We use this progress signal as a temporal critic for direct goal-conditioned policy learning. Model-assisted targets provide an additional training-time refinement under behavior-support and dynamics-agreement constraints. Across ten visual control tasks, DTRC achieves strong goal-conditioned control performance relative to planning and direct-policy baselines. Held-out diagnostics on the four LeWM tasks show consistent short-range temporal calibration, task-dependent long-range and directional structure, and positive transition-level progress. Temporal supervision improves the same flow-policy parameterization across all four LeWM tasks, while the resulting policy acts directly without iterative trajectory search at test time.
cs.LG / 23 / 2610.08969
Work While They Sleep: Exploiting Evaluation Latency for Fully Bayesian Optimization
Abstract
Black-box optimization problems are ubiquitous across science and engineering, often dealing with expensive objective functions. This objective latency has two consequences during optimization: (i) the objective evaluation dominates execution time, and (ii) sample-efficient algorithms are crucial to accelerate development and avoid wasting resources. Bayesian optimization (BO) methods are the \textit{de facto} choice of planners for suggesting the next point to try. Standard BO fits the surrogate model's hyperparameters with a point estimate. Alternatively, a fully Bayesian approach uses model averaging to account for uncertainty over the hyperparameters, leading to better uncertainty estimates---useful in the low-data regime that is pervasive in BO. However, it is often prohibitively expensive and thus rarely used. In this work, we propose ELF-BO, an algorithm that uses the objective evaluation latency to headstart the computation of the next suggestion, allowing for fully Bayesian optimization without incurring substantial decision-time costs. This is done by sampling from the hyperparameter posterior \emph{while} the objective is being evaluated, only requiring reweighting of the samples once the objective value is observed. Across synthetic functions and real-world applications, we show that ELF-BO matches the performance of fully Bayesian methods while only incurring decision latency on par with or better than standard BO. Thus, ELF-BO makes fully Bayesian optimization practical in real-world use cases.
cs.LG / 24 / 2610.08975
The Best Optimizer Depends on Batch Size
Abstract
A plethora of new adaptive optimizers are designed to efficiently estimate and use minibatch gradient statistics to shape parameter updates, but they are typically benchmarked at a single batch size. Hyperparameter scaling rules promise to preserve performance as batch size and gradient noise change, suggesting that the best optimizer at one batch size should remain the best at another. We challenge this approach to developing and evaluating optimizers by showing: (1) no principled scaling rule for Muon works consistently across training settings, and (2) the best optimizer for language model pretraining changes with batch size even after extensive hyperparameter tuning.
cs.LG / 25 / 2610.08989
LASER: Latent Space Adjoint Matching for Support-Constrained Entropy-Regularized Offline RL
Abstract
While offline reinforcement learning (RL) enables policy optimization from static datasets without costly online interaction, it remains bottlenecked by the risk of executing out-of-distribution (OOD) actions. Recent approaches mitigate this by learning a behavior-cloning policy through flow matching and then performing RL within its constrained latent space. However, naively optimizing the latent policy can easily cause the policy to collapse into a brittle mode or exploit sharp artifacts of the learned critic. In this work, we find that entropy regularization is essential in latent-space RL for addressing these challenges. We introduce LASER, a novel offline RL algorithm that applies latent-space adjoint matching to achieve entropy-regularized latent-space RL with expressive flow policies while avoiding backpropagation through time. Through comprehensive experiments on 40 challenging OGBench tasks with varying dataset qualities, we show that LASER achieves state-of-the-art performance. Notably, LASER uses fixed method-specific hyperparameters across all tasks and outperforms the evaluated baselines, including those with task- and dataset-specific tuning, which highlights the robust applicability of LASER. Project website: https://mit-realm.github.io/laser/.
cs.LG / 26 / 2610.08991
REFIT: Recognize, Fix, and Test Wearable Sensor Placement Shifts without Labels
Abstract
We present REFIT, an input calibration for frozen activity-recognition models whose inertial sensors are worn differently at deployment than in training. When users move a watch to the other wrist or put a strap sensor back on turned, the model sees the same motion on changed axes. REFIT undoes such shifts without labels or retraining. It describes them by families of axis transforms, such as reflections and rotations, and fits each family to the user's data so that simple statistics match those of the training data. The family that removes most of the mismatch names the shift. REFIT fixes the shift by applying the best member of that family before the frozen model and re-estimating its normalization statistics. It tests the fixed model with a label-free accuracy estimate and asks the user to re-wear the sensor when it is low. Experiments on real left/right sensor pairs and on real and simulated re-attachment show that REFIT outperforms label-free test-time adaptation methods on every dataset and restores most of the accuracy lost to re-attachment. It names injected shifts far more reliably than a confidence-based selector. After a correction over all signed permutations of the axes, the estimate separates successful from failed corrections.
cs.LG / 27 / 2610.09004
Removing Information Content Does Not Certify Tamper Resistance in Open-Weight Models
Abstract
Does removing harmful information make open-weight models resistant to fine-tuning attacks? We show that mutual information at release alone cannot universally certify slow recovery. Function-preserving reparameterizations leave information unchanged while altering gradient-descent geometry, so an invariant certificate is bounded by the fastest reachable parameterization. We apply this principle to weight--data mutual information under training-data filtering and label--representation mutual information under capability removal. Training order can change recovery time at fixed weight--data information, while exact representation-level independence can preserve the entire parameter Jacobian. An explicit construction has both information quantities equal to zero and recovers in one gradient step. Controlled experiments illustrate order-dependent recovery and parameterization-dependent attack speed. These results identify the missing requirement for certification: constraints on attack dynamics beyond mutual information at release.
cs.LG / 28 / 2610.09020
Neighborhood Smoothing for Calibration
Abstract
Modern neural networks are often miscalibrated, with a tendency to overconfidence. Existing train-time calibration methods largely modify task losses or calibration penalties, leaving neighborhood structure in learned representations underexploited. We introduce graph smoothing as a general principle for train-time calibration, which encourages similar predictive distributions across neighboring samples in representation space. We analyze the effects of graph smoothing, deriving bounds that connect predictive divergence between neighboring samples to local confidence variation and to the propagation of pointwise calibration error, and characterize the conditions under which smoothing can or cannot improve calibration. In light of this analysis, we propose \modelNoSpace, a graph-based train-time regularizer that penalizes the Jensen--Shannon divergence between predictive distributions of neighboring samples. We present a thorough empirical analysis, showing that across standard calibration benchmarks, \model improves predictive quality, and the improvement is complementary to post-hoc calibration: after temperature scaling, \model attains the lowest NLL of all evaluated train-time methods in seven of the eight image and tabular settings. These findings demonstrate the value of graph smoothing over learned representations for neural network calibration.
cs.LG / 29 / 2610.09025
SPIN: Shadow Predictive Indexer for Sparse Attention
Abstract
Indexer-based sparse attention reduces the cost of core attention by passing only a fixed, small number of important tokens to it. However, the indexer must still score the entire KV cache at every decoding step. This scoring overhead becomes a major bottleneck as the context length grows. We propose SPIN (Shadow Predictive Indexer) to reduce this indexer overhead. SPIN uses lightweight, history-based prediction to identify important KV blocks, avoiding the need to score the full KV cache at every decoding step. SPIN treats KV blocks and speculative decoding as first-class design and implementation considerations. Across extensive evaluations on long-context and agentic benchmarks, SPIN achieves 30-40% sparsity while preserving task quality. In end-to-end vLLM serving, SPIN improves output throughput by up to 14.9% and reduces median inter-token latency by up to 13.2%.
cs.LG / 30 / 2610.09040
ORACLE: Optimizer-Relative Alignment for Constrained LEarning
Abstract
Constraint handling methods typically intervene before the optimizer acts, by modifying the objective or the gradient. Yet momentum, adaptive scaling, and structured preconditioning can substantially reshape that signal before it becomes a parameter update. We formulate optimizer relative constrained learning, where constraint compatibility is assessed on the post optimizer update. Building on this view, we introduce ORACLE, which evaluates the native optimizer's realized step through a joint endpoint linearization of heterogeneous constraint families, constructs the resulting alignment in the optimizer's own geometry, bounds its authority, and commits it only after validation. We evaluate ORACLE across eight Partial Differential Equation benchmarks and four optimizers spanning Euclidean, diagonal adaptive, and structured preconditioned geometries, where it improves or matches native optimizer in 94% of configurations. Cross model analysis shows the same behavior in 92% of configurations, while matched comparisons show improvements over alternative constraint-handling methods acting at the objective, gradient, and post-optimizer levels.
cs.LG / 31 / 2610.09048
Towards Financial World Modeling
Abstract
Building a world model requires a state representation useful for planning and decision-making---potentially over tasks unknown at training time. In the context of financial markets, planning and decision-making may require a model to reason about market-wide conditions, asset-specific expected returns, liquidity, volatility, and cross-asset relationships. Yet financial representation learning has largely been evaluated on individual predictive tasks, oftentimes on a single time period using comparatively narrow datasets. We address this through three primary contributions. First, we introduce Market-1T, a dataset containing nearly one trillion observations across U.S. equities from 2008 to 2025 at 1 Hz resolution. Second, we develop and implement a rigorous evaluation protocol. Third, we conduct a systematic large-scale study of financial representation learning, comparing 18 encoder-training strategies across nearly two decades of market regimes. We evaluate learned representations both by their predictive utility on common finance tasks and through probes of latent structure. We find that encoders with similar predictive performance can organize market state very differently. Collectively, we establish a foundation for training and evaluating financial market representations in support of world models such as DINO-WM, V-JEPA 2, and LeWM.
cs.LG / 32 / 2610.09051
A Self-Pruning Transformer: Extreme KV-Cache Compression with Universal Attention
Abstract
The large KV-cache size of modern LLMs creates a barrier to efficient deployment. Recent work has explored replacing attention layers' RoPE positional embeddings with alternative decay-based mechanisms, which can then be used to prune KV-cache during inference. However, these decay functions have limited expressivity, and in practice devolve into sliding-window-like eviction patterns. In this work, we propose a unifying framework for complementary and novel decay mechanisms, capturing complex key statistics and interactions while preserving expressive RoPE embeddings and Softmax attention. The resulting Universal Attention is a highly expressive and end-to-end trainable architecture, whose composite decay mechanism acts as a natural, $\textit{adaptive}$ pruning criterion, removing tokens that contribute least to attention computation. Experimentally, Universal Attention achieves state-of-the-art $10\times$ compression on natural language and synthetic task data, while $\textit{improving}$ downstream performance compared to both state-of-the-art baselines and unpruned oracles. It further demonstrates superior long-context generalization with unprecedented $25\times$ compression at length 16k.
cs.LG / 33 / 2610.09056
A Geometry-Based Capacity Theory for Finite-Feature Associative Memory
Abstract
We develop a geometry-based capacity theory for exact-key retrieval in compressed finite-feature Hebbian associative memory. For random or approximately isotropic values, retrieval interference separates into finite-feature noise, which decreases with feature dimension, and structural interference, which is determined by squared kernel overlap among stored keys and persists in the infinite-feature limit. This yields a fit-free prediction of retrieval quality, reveals a geometry-dependent capacity ceiling, and predicts the feature budget required for a target retrieval quality. When stored values are correlated, we show that retrieval depends jointly on the key kernel and value Gram matrix, and derive finite-feature approximations that account for this interaction. We validate the theory on synthetic, visual, and medical-image representations. Overall, the framework links representation geometry directly to memory capacity and distinguishes when performance can be improved by increasing the feature budget and when the representation itself must be changed. Across these settings, the predicted retrieval curves closely match empirical behavior and correctly identify changes in the preferred memory design.
cs.LG / 34 / 2610.09058
Learning Cross-Model Activation Alignments with Explicit Many-to-Many Layer Maps
Abstract
LLMs are released at a rapid pace, raising a natural question: how do two independently trained models relate, both in which layers correspond and in how features transform between them? We study this by learning an activation alignment, a map from a source model's layerwise activations to a target's. Our method, MATCHA, factors this map into a layer map, whose output is an explicit target-by-source matrix that can be extracted and inspected, and a layer-shared feature map between hidden spaces. Most of prior work fixes the layer correspondence in advance, pairing layers at roughly the same relative depth; in contrast, we learn both factors jointly from prompts. Across 42 pairs of seven models spanning three different families, MATCHA reconstructs the target's activations more faithfully and improves retrieval-based metrics substantially, w.r.t. previous approaches. The recovered maps are broadly monotone in depth but, in contrast with most previous approaches, are consistently many-to-many: each target layer draws on a band of source layers. Our alignments also enable transfer of activation-space interventions, allowing steering vectors and probes developed for one model to transfer to another.
cs.LG / 35 / 2610.09059
EDiS: Edge Disjoint Subgraph Sparsification Framework for Graph Neural Networks
Abstract
Sparse GNN training reduces computation, but deciding which edges to keep can be costly. Reusing one sparse graph is cheap, but locks training to a fixed topology, while varying it across epochs can require repeated sampling or recomputation. We introduce EDiS (Edge-Disjoint Subgraph sparsification framework), which separates one-time structural extraction from per-epoch graph composition. EDiS decomposes the graph once into cacheable edge-disjoint subgraphs, then recombines them into graphs with edge-budget constraints across epochs and retention ratios without re-extracting structure. Our default construction uses feature-based scores and successive maximum score covering forests, while the same composition mechanism also supports alternative edge selection rules. We provide a combinatorial analysis of the per-epoch sampler, the composition step that draws a training graph from the cached decomposition. We show that, under the default covering-forest selector, the stored decomposition deterministically preserves high-score cut edges, and we derive a selector-agnostic conditional bound on high-score cut survival in composed training graphs. Across 19 homophilic, heterophilic, and large-scale node classification benchmarks against 17 baselines under the same edge budget, EDiS achieves the highest mean benchmark score (accuracy/ROC-AUC) and the lowest average rank and gap-to-best among ranked methods. Ablations show the clearest benefits of structural decomposition and epoch variation at tight edge budgets.
cs.LG / 36 / 2610.09074
TAP: Efficient Long-Horizon Agent Pruning via Trajectory-Anchored Recovery
Abstract
Emerging long-horizon agentic tasks require repeated model calls, worsening the inference cost of already-costly language models. While narrow agentic tasks suggest potential for aggressive model pruning without performance drop, empirical results show existing methods proposed for question answering tasks severely degrade task performance when applied to agentic models. We trace this failure to two decisions: what to prune and how to recover. For pruning, one-shot importance estimates fail to track how the pruned model adapts. For recovery, offline distillation covers only teacher prefixes, while full-trajectory on-policy distillation causes student errors to compound across turns. In this work, we propose Trajectory-Anchored Pruning (TAP), the first structural pruning framework for reinforcement learning (RL)-trained agents. TAP couples structural pruning with efficient on-policy recovery, anchoring interactions to teacher trajectories while allowing the student to generate each reasoning-action response. A frozen dense teacher supervises the student's response prefixes, addressing within-response training-inference mismatch while preventing student-induced deviations from propagating across training turns. Instead of one-shot pruning, TAP re-scores channels using gradients of the recovery objective on the recovered student, connecting iterative channel selection to the evolving policy. With 60% of FFN channels removed, TAP retains 99.2% and 88.0% of the dense 7B agents' task success rates on ALFWorld and WebShop, respectively, while reducing GPU time per successful task by approximately 22% and 17%. These results demonstrate effective structural compression of long-horizon agents under a limited recovery budget.
cs.LG / 37 / 2610.09090
Tucker Bottleneck Attention for Multi-Dimensional Sequence Modeling
Abstract
The quadratic cost of self-attention limits scalability to long sequences from multidimensional data. We introduce Tucker bottleneck attention (TuBA), which exploits low-rank tensor structure for efficient global token mixing. TuBA projects hidden tensors into compact Tucker cores, performs multi-head self-attention and linear projections on the cores, and writes updates back to the ambient space, enabling subquadratic computation. Its autoregressive extension combines bidirectional interactions within cores with causal attention across cores. On video prediction and global weather forecasting, TuBA achieves favorable accuracy-efficiency trade-offs over standard and efficient attention and task-specific models. Compared to standard self-attention, TuBA reduces error and computation by up to 24.7% and 66.6% for video prediction and 37.1% and 85.1% for autoregressive weather forecasting, with speedups up to 4.27 times. Low-rank Tucker cores and multi-frame generation also outperform full-rank attention and frame-by-frame generation, respectively.
cs.LG / 38 / 2610.09091
FedRSPO+: A Heterogeneity-aware Algorithm for Decision-focused Federated Learning
Abstract
Decision-focused learning (DFL) trains predictive models for downstream optimization, but existing methods largely assume centralized data. In cross-silo settings, federated learning offers a natural alternative, yet standard federated methods optimize prediction over decision quality and do not address heterogeneity in downstream objectives or feasible sets. This heterogeneity is especially challenging for DFL because small perturbations in polyhedral problems can cause discontinuous changes in optimal decisions, destabilizing client updates and aggregation. We propose FedRSPO+, a heterogeneity-aware framework for decision-focused federated learning, built on RSPO+, a regularized predict-then-optimize surrogate that smooths the decision map through projection. We show that RSPO+ upper bounds decision error and regret for the regularized decision and, under exact regularization and consistent LP solution selection, for the original LP decision. We further derive cross-client heterogeneity bounds that depend on both objective and feasible-set heterogeneity, vanish at homogeneity, and require no strong convexity. FedRSPO+ uses an annealed, modular training procedure compatible with standard federated personalization and aggregation methods. Experiments on synthetic knapsack, shortest-path, and real-world energy pricing tasks compare against prediction-only federated learning and DFL baselines under varying heterogeneity and communication budgets. Results suggest that smoothing is a useful ingredient for stable collaborative decision learning and provide a heterogeneity-aware foundation for federated DFL.
cs.LG / 39 / 2610.09096
Are We Really Benchmarking Forecasting Models? The Impact of Preprocessing on Time Series Performance
Abstract
While established literature underscores the pivotal role of preprocessing in forecasting accuracy, this stage remains largely overlooked in current research. Modern benchmarks typically resort to simple scaling, failing to account for critical transformations required to address nonstationarity, such as differencing. This omission creates a significant structural preprocessing bias that favors models with built-in data treatments while obscuring the true potential of simpler architectures. We study this effect through a preprocessing-aware benchmark that evaluates 11 forecasting models across 16 reversible preprocessing pipelines on 29,000 M4 time series. Our results identify preprocessing as a key driver of forecasting performance. Optimizing preprocessing per series yields gains of approximately 27\% to 87\% across all evaluated models, with architectures lacking internalized preprocessing experiencing the most substantial improvements. This allows simpler architectures to become highly competitive with complex, state-of-the-art models in modern forecasting benchmarks. All resources and experimental results from this benchmark are stored in a comprehensive metadataset to support future metalearning tasks.
cs.LG / 40 / 2610.09108
Convex-Concave Reinforcement Learning
Abstract
Policy learning drives many of the most consequential and heavily-invested applications of reinforcement learning today. Yet the core optimization problem it rests on (maximizing expected return) is notoriously non-convex, even under a direct policy parameterization, and the field has largely responded by avoiding it: optimizing convex surrogate approximations of the return under trust-region constraints (NPG, TRPO, PPO, AWR). We show that this seemingly unstructured problem is not actually structureless. In log-density-ratio coordinates $y := \log[π/π_n]$, the exact per-iteration objective, computable via per-decision importance sampling (PDIS), is a difference-of-convex-constrained difference-of-convex (DC-constrained DC) program. This structure lets us move beyond surrogate approximations: it recovers CPI, NPG, TRPO, and AWR as special cases along interpretable axes, and it opens a multi-step axis $k$ that couples consecutive decisions. We solve the per-iteration program with sequential convex programming (SCP), the standard solver for difference-of-convex problems, and give convergence guarantees under mild conditions, bridging the difference-of-convex optimization and RL literatures. Empirically, multi-step Convex-Concave RL (CCRL) wins on diagnostic MDPs where credit must propagate across a horizon (its advantage growing with the dependency length), is competitive with a tuned PPO on classic control, and on a realistic, stochastic, mid-horizon healthcare domain converges markedly faster than tuned PPO to the same near-optimal survival, with an 11.3% higher area under the training curve.
cs.LG / 41 / 2610.09109
Breaking Adversarial Transferability in Fine-Tuned Speech Recognition
Abstract
Many organizations fine-tune publicly available pretrained Automatic Speech Recognition (ASR) models and deploy them in black-box settings, assuming limited access provides protection. We show this assumption is fragile: adversarial perturbations crafted on the public base model transfer effectively to fine-tuned target models, severely degrading performance and posing concerns for safety-critical applications. We propose TransferBreaker, a unified fine-tuning framework that suppresses adversarial transfer by integrating Base Adversarial Fine-Tuning, which restricts adversarial training to base-effective perturbations; Latent Jacobian Regularization, which enforces latent-space invariance by suppressing adversarially sensitive directions; and HybridGrad-AFT, which improves robustness against adaptive attacks by interpolating transferable perturbations from base and target gradients. We theoretically justify all components and evaluate TransferBreaker across three languages and four large ASR models, reducing adversarial WER from 92.6 to 27.8. Our code is publicly available at https://github.com/rohban-lab/TransferBreaker.
cs.LG / 42 / 2610.09120
The Deceptive Bandit Problem: Exploratory Coupling and the Fragility of Multi-Agent Learning
Abstract
Randomized exploration is central to bandit learning, multi-agent reinforcement learning, and zeroth-order policy search, yet its independence and privacy are usually only treated as technical assumptions. We show that these properties are critical for security purposes and demonstrate how an adversarial agent can exploit privileged information on another agent's exploration. We analyze a deceiver-victim pair in the minimal two-player strongly monotone setting, where a deceptive player obtains leaked signals that are merely correlated with the victim's exploration. We show that, by coupling their own exploratory action with this information, the deceptive player injects an externality that steers the learning dynamics to a new steady state, called the deceptive Nash equilibrium (DNE). We prove that the deceptive bandit learning (DBL) dynamics converge to an arbitrarily small neighborhood of the DNE while retaining optimal convergence rates. Interestingly, our analysis attains these optimal rates while relaxing second-order smoothness conditions from standard bandit optimization literature. We characterize conditions under which deception strictly shifts the steady state and its effect on the deceiver's cost, illustrating the results in a resource-allocation game.
cs.LG / 43 / 2610.09124
Spatial Induction Heads: In-Context Learning of Multidimensional Cellular Automata
Abstract
Induction heads provide a mechanistic account of in-context learning in sequential data, but existing theory largely assumes that the context relevant to a prediction forms a contiguous block. In multidimensional data, serialization breaks this assumption by scattering spatial neighbors across distant positions in the token sequence. We study how transformers overcome this routing problem in multidimensional stochastic and deterministic cellular automata, where each trajectory is generated by an unknown local rule and presented as a flattened sequence without an explicit coordinate-based spatial inductive bias. We introduce spatial induction heads, two-layer gather-and-match circuits in which the first layer reconstructs the relevant spatial neighborhood and the second matches the resulting configuration against earlier occurrences. We give two explicit realizations of the gather and show that the positional dimension required for spatial routing depends only on the local neighborhood and spatial dimension, not on grid volume or trajectory horizon. We further construct a matching layer which implements Bayesian counting. The end-to-end circuit can approximate the Bayesian posterior arbitrarily closely for stochastic rules and can predict exactly for deterministic rules. Empirically, trained two-layer transformers generalize to unseen rules in one and two dimensional settings, achieving near-perfect deterministic rollouts and less than 0.005 nats KL from the Bayes-optimal predictor on stochastic rules. Attention patterns and layerwise probes align with the predicted gather-and-match computation, providing mechanistic evidence for spatial induction in trained transformers.
cs.LG / 44 / 2610.09126
Domain-informed Adaptive Sampling for Generalizable PINNs in Metal Additive Manufacturing via Conditional Flow Matching
Abstract
Accurate thermal modeling is essential in metal additive manufacturing (AM) for understanding the process-structure-property chain. Physics-informed neural networks (PINNs) offer effective surrogate thermal modeling by minimizing physics-based residual losses at collocation points. However, prior works typically rely on manually-crafted, static collocation sampling strategies, which are neither principled nor scalable across process conditions, hindering their generalization capability. In this work, we provide theoretical analysis through empirical risk minimization, showing that process condition-aware adaptive sampling is strictly more favorable than conventional static sampling for generalization. Building on this insight, we propose an adaptive sampling strategy within a two-stage framework: (1) a conditional Flow Matching model that learns approximate high-residual distributions across different process conditions, and (2) a mixed sampling strategy combining this distribution with a domain-informed base distribution to generate adaptive collocation points for refining the PINN predictor. Experiments on metal AM numerical benchmarks demonstrate that our method consistently outperforms state-of-the-art PINN baselines, achieving an average 62.1\% reduction in relative $L_2$ error under an identical collocation budget, by capturing process-dependent heat dissipation regions often overlooked in the literature. To the authors' knowledge, this is the first adaptive sampling strategy for PINNs in metal AM, contributing to the enhanced generalization and broader applicability.
cs.LG / 45 / 2610.09136
Directional Evidence Guided Search-Space Reduction for Exact DAG Learning
Abstract
Learning a directed acyclic graph (DAG) from observational data is a challenging combinatorial problem due to the exponential growth in the number of candidate parent-set configurations. Existing exact score-based methods often require computationally intensive combinatorial search, whereas constraint-based methods can become unreliable or computationally demanding as graph size and conditioning-set complexity increase. We develop a non-parametric hybrid framework, referred to as DECO (Directional Evidence-guided Configuration Optimization), that extracts dependency and directional evidence from observation data to construct admissible parent sets prior to exact optimization. It reduces the optimization search space by eliminating empirically unsupported parent configurations while preserving flexibility for all plausible edge orientations. Theoretical analysis establishes an exponential reduction in the admissible parent-set configuration space and quantifies how bounded edge-level omission affects the probability of retaining the true parent structure. Experiments on benchmark Bayesian networks and synthetic discrete and continuous DAGs demonstrate substantial search-space reduction while achieving competitive structure-recovery performance, with favorable structural Hamming distance across many evaluated settings. These results show that directional evidence can provide an effective preprocessing mechanism for reducing the computational burden of exact DAG learning without requiring a fixed parametric structural~model.
cs.LG / 46 / 2610.09141
A Cognitive-Aware QML-CRL Framework for Detecting Affinity and Romance-Investment Fraud
Abstract
We present a hybrid quantum-classical framework that detects affinity and romance-investment fraud by modelling the cognitive biases in a manipulative conversation. In our proposed framework, cognitive biases central to this fraud class are carried by dedicated qubits in a structured parameterized quantum circuit, together with a frame qubit makes the encoding sensitive to the temporal order of manipulative reframing, and a narrative qubit that aggregates co-occurrence through a trainable entanglement layer. The circuit parameters are trained jointly with a classical reinforcement-learning agent that decides, turn by turn, whether to flag the conversation, modeled as an optimal stopping problem. We evaluate the model's performance on synthetic conversations that include hard negatives, legitimate but urgent, and legitimate but pushy sales conversations.
cs.LG / 47 / 2610.09174
Context-aware Attention-based Gaussian Mixture Models for Vehicular Trajectory Prediction
Abstract
Reliable and interpretable trajectory prediction is critical for cooperative and autonomous driving in complex and uncertain environments. This paper introduces a Context-Aware Attention-based Gaussian Mixture Model (CAA-GMM) for multimodal, uncertainty-aware motion forecasting. The proposed approach models future motion as a probabilistic mixture conditioned on both scene context and agent dynamics, capturing diverse behavioral modes with interpretable Gaussian components. A lightweight attention mechanism adaptively encodes inter-agent interactions and contextual salience, enabling efficient fusion of rasterized environment cues and motion history in dense traffic scenes. Comprehensive evaluations on the nuScenes and Argoverse 2 datasets demonstrate that CAA-GMM achieves competitive or superior accuracy compared with state-of-the-art raster-based baselines, while maintaining low computational complexity. Ablation analyses confirm the importance of the attention module for robust contextual reasoning and predictive precision. Furthermore, evaluations under imperfect communication and perception conditions highlight the framework's resilience to uncertainty, establishing CAA-GMM as an efficient and scalable solution for cooperative trajectory prediction in intelligent transportation systems.
cs.LG / 48 / 2610.09179
LayerRoPE: Dynamic Depth-wise Magnitude & Angular Superposition
Abstract
As data propagates through a Transformer, the norm of its hidden states grows by orders of magnitude with depth, a phenomenon framed as 'curse of depth' and nearly universally treated as a pathology to be suppressed. We take the opposite view. Across 16 pre-trained LLMs from 9 families, spanning dense, mixture-of-experts and hybrid architectures and Pre-, Peri- and Post-Norm designs, we find that this growth reflects an emergent depth-positional encoding, carried by the only learned per-layer gain on the residual stream, the normalization weight $γ$: with depth, $γ$ grows in magnitude and rotates in direction, jointly encoding the layer index. We make this depth-conditioned encoding explicit with LayerRoPE, an implicit analog of RoPE along the depth axis, which replaces all layerwise $γ$ vectors with a single shared vector and depth-conditioned scalars, at a net reduction in parameters and $<0.02\%$ change in FLOPs. Across a model ladder scaled up to $100$B+ tokens, LayerRoPE consistently outperforms Pre-, Post- and Peri-Norm and Layer-Norm Scaling, reaching Pre-Norm's 1.3B loss with $3.4\times$ less compute; LayerRoPE is the only approach that shows strong convergence and improves near monotonically as depth scales to 512 layers. It improves learning-rate sensitivity by $3$-$10\times$, and transfers naively to and consistently improves looped latent models and Vision Transformers. Inspecting its learned schedule inverts the prevailing premise: LayerRoPE does not shrink the residual stream but widens it, damping what each block reads while amplifying what it writes. Depth stability, our results suggest, calls not for suppressing the residual stream, but for depth-conditioned regulation of the computational blocks it feeds.
cs.LG / 49 / 2610.09183
Q-PACE: Dynamic Precision Allocation for Quantization-Aware Training
Abstract
Quantization-aware training (QAT) leverages lower-precision arithmetic to reduce the cost of LLM deployment, but aggressive quantization degrades final model performance. A common remedy is mixed-precision training, in which high precision is assigned to some of the layers to maintain performance while keeping the cost constrained. This approach then requires precision assignments for model layers during training. We provide a new approach, called Q-PACE, consisting of a second-order sensitivity model that predicts the loss increase as a sum of quantization noise MSE weighted by per-layer curvature coefficients. During training, we periodically re-compute these coefficients using perturbations across layers, and re-assign precision. Pretraining and supervised fine-tuning experiments on LLMs of up to 4B parameters show that Q-PACE consistently improves over existing mixed-precision training recipes, and achieves comparable loss at substantially lower total memory budgets. We further find that quantization sensitivity is highly predictable by depth and layer type, and its stability during training allows for infrequent, cheap recalibration.
cs.LG / 50 / 2610.09194
Patient, Place, Prior (P$^3$): What Counts as Personalization in Medical World Models?
Abstract
Longitudinal models forecast how a patient's imaging state evolves, but accuracy does not show whether the patient's observed trajectory drives the prediction. A population-average forecast may be useful but cannot establish a patient-specific world-model claim. We introduce Patient, Place, Prior (P$^3$), an audit asking whether a forecast benefits from the patient's longitudinal imaging history (Patient), benefits from patient-matched externally supplied spatial support (Place), and gains predictive value beyond a population-average prediction under matched support and context (Prior). We also propose Cancer JEPA, a one-step model that forecasts frozen representations of future breast dynamic contrast-enhanced MRI examinations during neoadjuvant therapy. It adds a lesion-constrained neural correction, trained with an occlusion-based latent objective, to a patient-conditioned low-complexity reduced-rank regression baseline. This factorization permits a post-hoc P$^3$ audit of the frozen model. In a validation cohort previously used in development, forecast error is lower when the neural correction receives the patient's history rather than another patient's and patient-matched lesion occupancy maps rather than substituted maps. However, the descriptive 95% interval comparing the correction computed from patient history with the population-average neural correction includes zero. P$^3$ thus separates input use from evidence of patient-specific predictive value beyond a population-level pattern.
cs.LG / 51 / 2610.09196
Exact Dynamics and Finite-Sample Trajectory Recovery of Linear Recursive Feature Machines
Abstract
Recursive feature machines (RFMs) learn representations of data by alternating between fitting a predictor to a dataset and updating features of that predictor using the average gradient outer product (AGOP). Connections between AGOPs and feature learning in neural networks motivate linear RFMs as a simple setting for analyzing how representations evolve during training. Here, we study the dynamics and statistics of linear RFM in noisy multi-output regression with isotropic sub-Gaussian input data and targets generated by a low-rank teacher matrix of dimension $d$. We extend the known connection between linear RFM and iteratively reweighted least squares from the interpolating setting to ridge-regularized multi-output regression with noise. We show that the learned feature matrix remains close to its infinite-data ideal counterpart at every iteration. Namely, for $n$ samples, we show the error in the feature matrix decays as $O(\sqrt{d/n})$ with high probability. Experiments on real-world text and single-cell gene-expression data illustrate the features learned by this simple linear model.
cs.LG / 52 / 2610.09206
An Accuracy--Information Tradeoff for Loss-Difference Conditional Mutual Information
Abstract
Loss-difference conditional mutual information (ld-CMI) uses the smallest of the standard observations in the supersample hierarchy of generalization bounds: it measures what a learner's loss differences reveal about which candidate of each pair it was trained on. Accuracy is known to force information into the model; data processing does not carry such lower bounds to losses. We show, by bounding three moments of the loss differences, that accuracy also forces ld-CMI. For linear predictors with a smooth convex loss of nonzero slope at zero, such as the logistic loss, plus a regularizer whose curvature and growth are both of power $r\ge2$, on product distributions over a scaled sign cube in dimension at least linear in $n$, every proper learner with expected excess risk at most $\varepsilon$ on these distributions at the optimal sample size $n\asymp\varepsilon^{-2+2/r}$ has worst-case ld-CMI of order $n$ bits, and $Θ(n/(1+(τ/\varepsilon)^2))$ bits under Gaussian noise of standard deviation $τ$ on the loss differences. The same holds without a regularizer, at $n\asymp\varepsilon^{-2}$. Consequently, range-scaled ld-CMI bounds cannot vanish on these distributions, although every proper learner's generalization gap is $O(n^{-1/2})$. We also show that model-level information does not determine noisy loss-difference information, and that the growth, slope and dimension conditions are needed, the last up to a logarithm.
cs.LG / 53 / 2610.09229
Conditional Accuracy Profiles: Diagnosing LLM Judges across Deployment Conditions
Abstract
LLM-as-judge is now a standard tool for scalable evaluation, but judge performance is still often summarized by a single accuracy number. This aggregate view hides the deployment conditions under which a judge succeeds or fails. We introduce \textbf{Conditional Accuracy Profiling} (CAP), a post-hoc diagnostic framework that decomposes pairwise LLM-judge accuracy into eight conditions organized into content sensitivity, robustness, and rationale quality. CAP is benchmark-agnostic: it can be applied directly when a benchmark provides the required annotations, approximately through task-subset proxies, or through controlled augmentation when perturbation pairs can be generated. We instantiate CAP on seven LLM judges across six pairwise judging benchmarks, including \textsc{judgerEva-Standard}, a controlled testbed we created to support all eight conditions. CAP exposes profile differences hidden by aggregate accuracy: on \textsc{judgerEva}'s judge-independent Hard-Constructed subset, the two judges most sensitive to omitted qualifications rank in the bottom three of seven by overall accuracy, so omission sensitivity is not predicted by aggregate accuracy. Across benchmarks, Position Robustness shows the strongest rank stability (mean Spearman $\barρ{=}0.87$) but is itself fragile under JudgeBench-Pro adversarial stress, showing the largest mean accuracy drop among the shared conditions, though the dominant degradation channel varies by judge. Condition-level profiles provide a more actionable basis than aggregate accuracy for selecting LLM judges.
cs.LG / 54 / 2610.09230
Symmetry-Informed Causal Partial Identification
Abstract
Partial identification (PI) entails estimating bounds on causal effects by encoding different assumptions on data generation as a constrained optimization problem. Such bounds can suffice to inform policy decisions even if the causal effect itself is not identifiable. Often vacuous in practice, practitioners seek to exhaustively encode domain knowledge as additional constraints to make the PI bounds more informative. We introduce known data symmetries -- invariance of the causal effect under certain data transformations -- as a new source of constraints to inform PI. We operationalize this as a shape constraint on the causal function, and via a change of measure against which PI is posed using simple data pre-processing. Both approaches are shown to sharpen bounds under two canonical PI models. This is shown both theoretically for the population case, and via experiments in the finite-sample case. More broadly, our framework establishes data symmetries as a natural, underutilized source of background knowledge for robust causal inference.
cs.LG / 55 / 2610.09247
An Informational Curse of Horizon in Goal-Conditioned Policy Learning
Abstract
The difficulty of learning goal-reaching policies is often attributed to a "curse of horizon" that manifests as bias accumulation in temporal-difference backups and noisy advantage estimates. In this work, we identify an additional informational curse of horizon in goal-conditioned policy learning, where increasing the goal relabeling horizon can significantly reduce policy generalization and performance. Through a series of controlled experiments with oracle planners, we decouple the goal horizons sampled during training from those that the policy is asked to reach at test time. Even when evaluated only on a sequence of nearby subgoals, goal-conditioned behavioral cloning (BC) policies suffer from severe, training horizon-dependent performance degradation that is mitigated by reinforcement learning (RL) objectives. We explain this phenomenon as a horizon-dependent decrease in the conditional mutual information between actions and hindsight-relabeled goals, and find empirically that both BC and RL policies trained on longer-horizon goals exhibit a shift in sensitivity from goal to state information, as measured by the policy's input Jacobians. Motivated by this observation, we find that distilling the input Jacobians of short-horizon policies into long-horizon policies yields significant performance gains, especially in combinatorial manipulation tasks. Taken together, our results highlight goal relabeling horizon as an important consideration when learning generalist policies from offline data.
cs.LG / 56 / 2610.09250
Efficient Best-of-N policy evaluation for inference-time alignment
Abstract
Best-of-N (BoN) is a common inference-time alignment method that selects the highest-scoring response among N samples from a reference model. Evaluating BoN policies from logged data is challenging under sample-only access because standard off-policy estimators require density ratios that depend on unavailable response likelihoods. In this paper, we propose a sample-only framework for evaluating and selecting BoN policies without access to these likelihoods. We show that the order-statistic structure of BoN allows the required density ratios to be expressed through score-rank probabilities that are estimable from samples alone. We then develop a doubly robust estimator of the BoN policy value (BoN-DR) that efficiently reuses a shared auxiliary sample pool across candidate budgets. We establish valid asymptotic inference even under reward estimator misspecification and prove the efficiency of our BoN-DR estimator. Since larger budgets can amplify errors in the score function and lead to reward overoptimization, we derive two selection rules: (i) maximizing the estimated policy value and (ii) maximizing a lower confidence bound on the improvement over the reference policy, which accounts for estimation uncertainty and provides a no-harm guarantee. Across synthetic experiments and GSM8K with multiple reference and reward models, our framework accurately estimates BoN policy values and selects effective sampling budgets.
cs.LG / 57 / 2610.09255
The Symbol of the Surrogate: Measuring Numerical Provenance in Neural PDE Solvers
Abstract
Neural PDE surrogates are trained on numerical solver outputs that contain both physical evolution and solver-specific discretization errors. Because surrogates are also evaluated against held-out trajectories from the same solver, standard benchmarks cannot distinguish fidelity to the exact evolution from imitation of the numerical scheme. We introduce an empirical Fourier-symbol diagnostic that probes a trained surrogate's linearized one-step operator with individual Fourier modes and compares it with both exact-evolution and training-scheme references. To address architectural spectral bias, we train identical networks on schemes with orthogonal dissipative and dispersive signatures and compare their learned operators. In linear advection, the learned surrogates reproduce the training schemes' amplitude and phase errors, with the twin-scheme difference reaching more than 99.8\% of the analytically predicted full-imitation ceiling. The same behavior occurs for a non-local Fourier neural operator and at the operator level for nonlinear Burgers dynamics. These results show that agreement with solver-generated test data does not by itself establish fidelity to the exact evolution. Fourier-symbol measurements provide a direct diagnostic of numerical provenance.
cs.LG / 58 / 2610.09272
Evaluating Trajectory Features for Routing Final-Layer Attention
Abstract
Attention routing requires a signal that predicts the value of attention on the current prefix. We evaluate whether hidden-state extrapolation error, curvature and error change improve this prediction beyond uncertainty, one-step displacement, position and state projections. Paired executions of the final attention layer supply signed next-token loss differences in frozen SmolLM3-3B-Base and Qwen3.5-4B-Base checkpoints. Utility-supervised routers are tested on 100 held-out PG-19 books at an identical causal 20 percent invocation quota. None of six prespecified comparisons shows a positive gain after familywise correction. In Qwen3.5, a parameter-matched fixed-projection control lowers NLL by 0.00356 nats/token relative to the trajectory router (95 percent interval 0.00218 to 0.00487). Secondary results depend on the operation removed, feature location and scoring horizon; frozen thresholds also drift substantially at longer horizons. Actual selected-query execution yields small long-sequence latency reductions with increased NLL, while learned routers remain slower during cached continuation. The study identifies limits on the incremental value of these trajectory summaries and separates allocation quality from measured inference benefit.
cs.LG / 59 / 2610.09281
Twist Flow for Inverse Problems
Abstract
In Bayesian inverse problems, posterior sampling requires generating samples that are consistent with given observations while capturing the range of plausible solutions. Direct conditional generative models introduce latent noise to model this ambiguity, but paired inverse-problem training can still encourage an almost deterministic map from the observation to the target. As a result, generated samples may be observation-consistent while under-representing posterior variability, especially when the posterior is multimodal, leading to undercoverage, mode distortion, or artificial transitions between distinct feasible solutions. We propose joint twist-flow, an augmented flow-matching formulation that learns a continuous transport from the augmented source state $(z_x, y)$ to the augmented terminal state $(x, z_y)$. Here x is the target variable, $y$ is the observation, $z_x$ is the Gaussian reference coordinate for posterior sampling, and $z_y$ is a Gaussian likelihood-side coordinate associated with the observation branch. Under a Gaussian observation model, $z_y$ is motivated by the normalized observation residual associated with observation compatibility. Its role is not to replace uncertainty in $x$, but to couple generated samples of x to observation consistency, helping reduce likelihood-inconsistent variation while preserving variability in weakly constrained directions. We validate the method on low-dimensional inverse problems with reference posterior samples, where joint twist-flow better preserves multimodal posterior support than a direct conditional-flow baseline. We further evaluate the method on image restoration and seismic subsurface velocity-model inversion, showing increased posterior variability while maintaining observation consistency.
cs.LG / 60 / 2610.09282
Self-attention summary networks for subsurface velocity-model building from common-image gathers
Abstract
Common-image gathers (CIGs) contain physically meaningful information about velocity-model errors through reflector focusing and residual moveout, but in conventional imaging workflows they are typically used only as diagnostic tools. In this work, we propose a multiscale self-attention summary network that maps high-dimensional 3D CIG volumes into compact conditioning embeddings for probabilistic subsurface velocity inversion. These learned embeddings preserve offset-dependent kinematic structure and spatial coherence while reducing variability caused by background-velocity mismatch. Conditioned on these summary embeddings, a flow-matching model learns a transport from a Gaussian source distribution to the posterior distribution of plausible velocity fields. Numerical experiments show that, compared with direct conditioning on raw CIGs, the proposed summary network improves posterior velocity inference. In particular, the multiscale attention design provides greater robustness to background-model mismatch, yielding more accurate posterior reconstructions and lower predictive uncertainty.
cs.LG / 61 / 2610.09297
Node-level Graph Neural Architecture Search Framework
Abstract
In recent years, Graph Neural Networks (GNNs) and architecture search frameworks have gained extensive application in non-Euclidean data processing, attributable to their superior capacity in managing unstructured data. Nevertheless, traditional approaches typically apply uniform convolution operations to all nodes, regardless of their varying structural and feature characteristics, which can undermine model performance and result in over-smoothing issues as the number of layers increases. To overcome this limitation, in this work, we propose a \textbf{N}ode-Level \textbf{G}raph \textbf{N}eural \textbf{A}rchitecture \textbf{S}earch (N-GNAS) algorithm. It can automatically choose an appropriate network architecture for each subset of nodes when updating node features. N-GNAS also introduces a contrastive learning loss to separate sample features from different categories and vice versa. In experiments conducted on eight datasets for node and graph classification, our methodology outperforms current leading GNAS techniques and traditional human-designed GNNs. For example, it achieves an accuracy rate of 78.26\% on the CiteSeer dataset.
cs.LG / 62 / 2610.09306
MovieSTAGE: Scene, Transition, and Global Encoding for Movie-fMRI ADHD Classification
Abstract
Naturalistic movie-fMRI provides a shared, temporally structured probe of brain dynamics, yet predictive models commonly rely on whole-run functional connectivity (FC) or temporally generic representations that are not aligned with narrative events. We introduce MovieSTAGE (Scene, Transition, and Global Encoding), a multiscale framework that combines hypergraph-structured FC-profile organization within scenes, unsigned FC-profile differences across adjacent scenes, and whole-movie FC. We evaluated 260 participants from the CMI-HBN Despicable Me cohort on case-control, ADHD-subtype, and three-class classification using 10 repetitions of stratified five-fold cross-validation, complete out-of-fold (OOF) predictions, and paired subject-cluster bootstrap and permutation tests. MovieSTAGE achieved AUROCs of 0.69, 0.73, and 0.75 and balanced accuracies of 67.6%, 69.8%, and 58.3%, respectively, yielding the highest mean point estimates among the evaluated methods. On the three-class task, the full model outperformed all two-branch variants, the HGNN scene encoder outperformed MLP, GAT, and BNT alternatives under matched settings, and the human-annotated partition outperformed duration-matched random and fixed-count GSBS controls. These controlled results support incremental predictive value from event-aligned scene and transition representations when combined with whole-movie FC in this cohort. Post-hoc model-derived analyses generated network-level hypotheses involving frontoparietal and default-mode systems.
cs.LG / 63 / 2610.09340
Benign Overfitting under Heterogeneous Input Fusion
Abstract
Benign overfitting is extensively studied when learning from a single high-dimensional input, but its behavior under heterogeneous input fusion remains largely unexplored. We study this question for minimum-norm linear interpolation under a heterogeneous Gaussian design, comparing two statistically dependent input blocks with their fusion while holding the underlying population task fixed. For regression, we identify a full-spectrum covariance certificate whose asymptotic status is independent of the cutoff threshold and prove that it is preserved by every positive-semidefinite joint covariance consistent with the two marginals. This protection is sharp, yet it does not extend to all benign regression problems: outside the certified regime, two benign marginals can have a harmful fusion. For one-sparse Gaussian classification, benignity in the regular regime is characterized by the balance between surviving predictive signal and nuisance contamination. Fusion can move these two quantities in opposite directions, and within this model class every marginal-to-joint benign/non-benign pattern is attainable. We further show that the same fused input can have qualitatively different effects on regression and classification. These results establish that benign overfitting under heterogeneous fusion is determined by the joint signal and spectral geometry created by input interaction, rather than by marginal benignity alone.
cs.LG / 64 / 2610.09355
Multimodal LLMs Can Learn to Read Brain Signals: A Vision--Language Model for Unified Multi-Task EEG Decoding
Abstract
Learning EEG representations that generalize across cognitive tasks, subjects, and recording conditions remains a key challenge in electroencephalography (EEG) decoding. Recent advances in foundation models have improved EEG decoding performance, yet a fundamental open question remains: how to effectively interface neural signals with these models to enable multi-task learning across datasets. To investigate this question, we introduce BraVista, a visual-language framework that encodes multichannel EEG signals as structured images and enables multi-task learning through instruction-conditioned vision-language models (VLMs). Our approach relies on continued post-training of a general-domain VLM, leveraging its visual and linguistic priors to adapt to neural signals without a separate large-scale EEG-specific pretraining stage. We evaluate BraVista on four datasets spanning sleep staging, emotion recognition, cognitive workload classification, and abnormal EEG detection, showing strong performance across these tasks. Further analyses show that the choice of EEG-to-image representation is critical to performance. Moreover, through controlled perturbations of the EEG signal, we observe a gradual performance degradation under increasing noise, suggesting that the model relies on EEG-relevant information rather than superficial visual patterns. Together, these findings establish structured visual representations as an effective and scalable interface between neural signals and general-domain foundation models for unified multi-task EEG decoding.
cs.LG / 65 / 2610.09356
Global Exponential Convergence of Two-Layer Linear Network Training
Abstract
We prove global exponential (linear) convergence with an explicit rate in the rich scaling for wide two-layer linear networks trained with smooth Polyak-Lojasiewicz predictor losses. Gradient flow in the factors closes exactly in terms of a finite-dimensional Bures flow of the neuron law covariance, in which the predictor dynamics are preconditioned by hidden covariance blocks. Mean-field conservation laws provide uniform spectral lower bounds on the hidden preconditioning blocks when the initial covariance satisfies a spectral support gap condition. This condition encompasses positive definiteness while still allowing for singular initializations. For an initial covariance $Σ_0 = σ^2 \mathrm{Id}$, the loss converges to the global minimum with linear rate at least $4σ^2κ$, where $κ$ is the PL constant. We establish stability of this rate under finite-width sampling, as well as global convergence of factor gradient descent for an explicit stepsize interval depending on smoothness, the initial loss, and conserved spectral margins. Our argument extends layerwise to deep linear ResNets, subject to a residual-path bound. In the case of heavy-ball momentum, training dynamics close instead over positions and velocities in terms of a lifted phase covariance. Linear convergence holds under an explicit condition on the energy and damping, specifying a window of admissible dampings. For two-scale white initializations, this interval is nonempty for sufficiently large position scales, with a fixed initial loss gap and velocity covariance. Numerical experiments illustrate the covariance geometry and compare the predicted and observed rates.
cs.LG / 66 / 2610.09375
From Retrieval to Customer Context: Evaluating Frontier-Model Systems for Voice-of-Customer Analysis
Abstract
Organizations increasingly use frontier language models to analyze customer feedback, but answer quality also depends on how that feedback is organized and made available. We define a \emph{customer context graph} as a unified model of customer and business context. Typed relationships connect customer objects (feedback, conversations, users, and accounts), operational objects (tickets, support agents, opportunities, and competitors), and analytical or action objects (taxonomy concepts, evidence, insights, work items, and outcomes). This lets an agent investigate not only what customers say, but why, who is affected, what action followed, who owns it, and whether it was resolved. For this experiment, the graph is populated from public Cursor feedback; the same architecture can support any type of feedback source. We compare Agentic RAG, a Deep Research Agent, and a Customer Context Graph-backed Agent on the same 9,432 public Cursor feedback records using 30 realistic product, incident, comparison, and metadata questions. Without exhaustive ground truth, we jointly score responses on answer quality (coverage and organization), analytical depth (specificity and decomposition), and evidence quality (citation support and traceability), using a comparative rubric calibrated on 28 of the 30 questions. We sample cited records against their claims and weight the three dimensions equally. Under this aligned rubric, the Customer Context Graph-backed Agent scores 0.961 overall, versus 0.710 for the Deep Research Agent and 0.651 for Agentic RAG, and leads the Deep Research Agent on 27 of 30 paired questions (sign-test p < 10^(-5); strictly best on 26 of 30). Its largest advantage is analytical depth (0.967 versus 0.642), reflecting more specific, hierarchically developed findings with quantified themes and traceable evidence...
cs.LG / 67 / 2610.09409
Stability and Diversity of Networked Self-Consuming Generative Ecosystems
Abstract
The widespread deployment of generative AI has made it increasingly difficult to distinguish synthetic content from real data. Consequently, synthetic data is inevitably incorporated into the training pipelines of future model generations, forming a self-consuming training loop. Prior work has studied the effects of such recursive self-consuming training, but analyses have largely been limited to isolated models, where a model consumes only its own synthetic data, or to simplified interactions between two models. This paper takes a first step toward understanding networked self-consuming generative models, in which multiple models consume synthetic data generated by one another through complex interaction pathways. We introduce a theoretical framework representing models as nodes in a directed, weighted graph, with edge weights governing the flow of synthetic data among models. Using this framework, we analyze the long-term behavior of networked models under retraining dynamics, establishing conditions for convergence and characterizing the resulting fixed points. We further investigate how the system's long-term stability and diversity are shaped by each model's access to real data, cross-model data consumption, and the structure of the interaction graph.
cs.LG / 68 / 2610.09411
Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data
Abstract
Multimodal representations enable zero-shot classification and retrieval, but aligning independently trained models usually requires large amounts of paired data. Yet, the Platonic Representation Hypothesis suggests that models trained on different modalities may converge spontaneously toward a shared representation geometry. But then, do we even need paired examples for cross-modal alignment? Remarkably, we show that paired examples are unnecessary for coarse cross-modal alignment. Our simple Wasserstein Procrustes method with a coarse geometric initialization aligns two disjoint embedding sets by estimating a single orthogonal map without seeing any pairs. Across datasets, modalities, and unimodal models, we show that we can consistently align independently trained representations without pairs, and standard geometric alignment metrics accurately predict when this is possible. Nevertheless, we can naturally benefit from paired examples. In the very few-pair regime, our method substantially outperforms existing ones, while staying competitive with pair-based methods with more added examples. Finally, we demonstrate that the resulting alignments can enable text-to-image generation without paired examples. These results show that independently trained models often share enough geometry to establish cross-modal correspondence with little or no paired data.
cs.LG / 69 / 2610.09415
Self-Consuming Generative Models with Co-Evolving Human Preferences
Abstract
Generative models are increasingly trained in self-consuming iterative loops, where users curate preferred samples from model-generated candidates and the curated samples are used to train future generations of the model. Prior work has largely assumed fixed user preferences, but in practice exposure to model outputs gradually reshapes what users perceive as desirable, creating a feedback loop in which model distributions and user preferences co-evolve. We take a first step toward understanding the long-term behavior of such coupled dynamics. We show that when training relies entirely on user-curated synthetic data, iterative curation amplifies initial biases and drives the system toward one of multiple singleton equilibria in which the instance holding an initial advantage eventually dominates. In contrast, injecting reference data into training at a sufficiently large rate fundamentally changes the dynamics and yields a unique globally attracting equilibrium. Building on this insight, we study how reference-data injection can be used to control long-term outcomes, and propose an efficient algorithm that jointly selects a reference distribution and its mixing weight to steer the coupled system toward equilibria that preserve desired attributes while minimizing data collection costs.
cs.LG / 70 / 2610.09425
What a Reporting Convention Hides: A Matched-Budget Audit of Quantum Natural Gradient with an Exactly Computed Metric
Abstract
Several published comparisons of variational quantum optimizers time only runs that reach a target loss, or read the verdict at a single target. Either convention could decide whether an optimizer's costlier steps pay off. We measure how much each convention changes verdicts among Adam, simultaneous perturbation stochastic approximation (SPSA) and quantum natural gradient (QNG), on initializations held out from the selection of settings. We compute the exact metric that preconditions QNG, price every step in circuit evaluations and give every method the same budget. In a median pooled over circuit widths, cost families and a sweep of settings with common misses, SPSA needs more than twice Adam's evaluations to reach a loose target. Dropping the censored runs that miss the target hides this gap. On the global-cost family we hold fixed the settings selected for a strict target. QNG then usually reaches the loose target after Adam but the strict target first. QNG's strict-target lead disappears when the metric's simulator price, linear in the number of parameters, is replaced by an assumed hardware count quadratic in that number. We recommend charging every miss the budget and reporting verdicts across targets.
cs.LG / 71 / 2610.09433
An extended deep energy method for thermo-mechanical crack propagation
Abstract
Thermo-mechanical fracture couples transient heat conduction on a cracked domain with a crack that grows as the temperature and the displacement evolve. Neural energy solvers have been proposed for phase-field fracture and later extended to represent a sharp crack through the network input, but heat conduction on the cracked domain and crack propagation under the resulting thermal stresses have not yet been treated together in these solvers. We present an extended deep energy method for thermo-mechanical crack propagation in which the crack remains a sharp polyline. Two networks represent the temperature and the displacement and receive the crack through a scalar embedding function, discontinuous across the crack and smooth elsewhere, so that both fields can jump across it without a regularization length, and the displacement is enriched near the tip by the Williams expansion with trainable amplitudes. The two fields are obtained by minimizing an incremental conduction functional and the thermoelastic potential energy in a staggered sequence, with Monte Carlo integration on points stratified over background elements, densified near the tip and redrawn during training. The stress intensity factors are extracted by the interaction integral with the area term of Wilson and Yu and checked by a sweep of the contour radius, and the crack advances at the maximum hoop stress angle when the energy release rate of the kink reaches the critical value at the crack-tip temperature. On a stationary thermal edge crack the extracted stress intensity factor agrees with the published value to 0.11%, in a functionally graded shear test initiation agrees with an independent sharp-crack finite element solution to within one load step, and on a notched cruciform specimen the crack paths follow the published solutions under mechanical, thermal and combined loading.
cs.LG / 72 / 2610.09456
GeoPrior-Mamba: Structured Process Priors with Mamba for Fine-Resolution XCO2 Reconstruction
Abstract
Reconstructing fine-resolution column-averaged dry-air CO2 (XCO2) fields from sparse satellite observations requires models to infer spatial structure that is only weakly constrained by direct measurements. Existing learning-based methods typically treat environmental covariates as ordinary numerical inputs and must therefore learn heterogeneous source-sink relationships largely from sparse supervision. We introduce GeoPrior-Mamba, a multi-directional Mamba framework augmented with offline language-model-induced structured process priors. Rather than using a language model to predict XCO2, we use it before training to organize relative process knowledge for biospheric uptake, ecosystem respiration, and anthropogenic emissions into deterministic prior tables. These priors are spatially instantiated using geographic, ecological, emission-related, and seasonal information and are adaptively injected into the reconstruction backbone through a lightweight knowledge adapter. Using OCO-2 observations from 2018-2020, GeoPrior-Mamba achieves an RMSE of 0.81 ppm and an R2 of 0.93 on held-out observations, reducing RMSE by 48.2% relative to CAMS background interpolation and by 3.1% relative to Trans-XCO2 under the same evaluation protocol. Ablation experiments show a measurable contribution from the knowledge-prior branch and substantially faster convergence than the knowledge-free Mamba backbone. Independent TCCON evaluation further supports the consistency of the reconstructed fields with ground-based column CO2 measurements. These results suggest that language models can provide a practical mechanism for constructing structured process priors when globally consistent process-response representations are difficult to obtain directly, while remaining outside the numerical prediction loop.
cs.LG / 73 / 2610.09457
DSReg: Provably Recovering Individual World Latents without Reconstruction
Abstract
Methods that recover individual latent variables of the world, from nonlinear ICA to dictionary learning and causal representation learning, anchor the latents to observations through reconstruction, auxiliary supervision, or distributional asymmetries such as non-Gaussianity. Methods without these anchors, including joint-embedding predictive architectures (JEPAs), identify the latent state only up to a linear transformation, so individual latents remain mixed. We close this gap: individual world latents can be provably recovered with no reconstruction, no decoder, and no labels. The key condition is Structural Diversity: different latents leave distinct dependency footprints on observations, just as no two snowflakes are alike. Building on the linear identifiability that LeJEPA provides, we prove that under Structural Diversity, DSReg (Dependency-Sparsity Regularization) recovers individual world latents up to signed permutation, without reconstruction or a decoder. It applies post hoc to any linearly identified representation, reusing trained checkpoints at no loss over joint training, and establishes the first fully identifiable JEPA that recovers every world latent. Moreover, as a condition on dependency footprints, Structural Diversity is strictly weaker than all structural conditions of prior identifiable latent variable models. Across synthetic regimes, world model probes, learned visual encoders, and external renderers, DSReg preserves dense prediction while improving individual-latent recovery and downstream use with scales.
cs.LG / 74 / 2610.09471
When Should an In-Context Learner Expand Its Hypothesis Space?
Abstract
Learning systems adapt quickly inside a familiar family of models. The harder step comes earlier: deciding, from observations that could be noise, an exception, a change within the family or structure outside it, whether opening a richer family is worth its cost. We treat this as a costly sequential decision: prediction failure must be turned into structural evidence, evidence into a value of expansion, and value into action. The Structural Revision Environment produces matched failures from each source, varies the price of expansion and the remaining horizon independently of the evidence, and admits exact Bayesian calculations and an exact normative solution of the one-shot decision. Its solution shows that revision is a value boundary and not an evidence threshold: one history has different optimal actions under different prices, horizons and announced queries, the boundary between local repair and expansion is set by the inputs a rule predicts and a repair cannot cover, and belief in the richer family crosses long before the decision does. Transformers trained in the environment reproduce this boundary from utility alone. Language models of three post-training lineages carry a failure-sensitive signal in their predictions that is not reflected in their revision decisions, and given the gain of expanding they read it without weighing it against price and horizon. Three models allowed to reason weigh the stated gain in the reference's proportions and still do not turn the history into an estimate of what expansion would buy. Controlled post-training of the meta-trained learners moves the prior and the sharpness of predictions, and neither moves the criterion.
cs.LG / 75 / 2610.09473
MORA: Modeling Observed Changes for Drift-Robust Time-Series Anomaly Detection
Abstract
Time-series anomaly detection (TSAD) identifies deviations from patterns learned from historical data. In non-stationary settings, distribution drift and true anomalies can cause similar local changes, making it difficult to tell whether a deviation reflects abnormality or evolving context. Existing methods typically adapt to detected shifts or learn drift-insensitive representations, but do not resolve this ambiguity. We define this problem as \emph{temporal change disambiguation}: determining whether a local deviation is explained by broader temporal evolution. We introduce MORA, a drift-robust TSAD framework that reconstructs the same local target from paired short- and long-term views. The reconstruction gap measures contextual support for a local deviation, and a data-dependent correction mechanism conservatively adjusts the primary local anomaly score. Context can only reduce the score when it improves reconstruction of the same target. MORA needs neither drift annotations nor online adaptation. Experiments on four TSAD benchmarks show strong robustness to non-stationarity while preserving sensitivity to genuine anomalies.
cs.LG / 76 / 2610.09476
CHASE: Channel-Aligned Structure Exploitation for Geometry-Aware Model Engineering
Abstract
Geometric and Spectral Alignment (GSA) characterizes trained networks through spectral concentration, physical-channel alignment, support structure, and changes in singular bases. In this paper, we propose CHASE (Channel-Aligned Structure Exploitation) to use these structures in practical model design. CHASE covers six applications across model modification, reconfiguration, and compression. CORA, COEC, and CORAM apply GSA to parameter-efficient finetuning, structured-pruning compensation, and model merging. We further develop three new methods. CAGA uses GSA to identify multi-head attention heads that can share a KV representation and constructs the shared key and value heads through geometric alignment and low-rank subspace extraction. SAKV uses GSA to determine which adjacent layers can share a low-rank KV-cache representation and the retained rank for each layer group. CAPS uses GSA spectral structure to group output neurons and selects retained input channels separately for each group. Results from CORA, COEC, and CORAM establish the effectiveness of GSA for adaptation, pruning compensation, and model merging. Experiments on CAGA show that geometric shared-head construction substantially improves MHA-to-GQA conversion, and SAKV and CAPS improve over representative baselines for KV-cache compression and structured pruning. These results show that the structures identified by GSA can be used directly to design methods for a range of model operations.
cs.LG / 77 / 2610.09508
Safe on Average, Unsafe in the Tail: When Is the Episodic-Cost Tail Controllable?
Abstract
Safe reinforcement learning seeks policies that maximize return while satisfying constraints on cumulative cost. Most methods impose these constraints on expected episodic cost. Consequently, standard evaluations report mean episodic cost without characterizing how cost is distributed across episodes. A policy that satisfies the mean-cost criterion may therefore remain unsafe in its worst episodes. Mean-cost reporting neither identifies this tail violation nor shows whether it can be brought within budget while preserving return. In this work, we measure the episodic-cost tail using $\mathrm{CVaR}_{0.1}$, the average cost of the worst $10\%$ of episodes. We classify a policy as tail-safe when $\mathrm{CVaR}_{0.1}$ is within the safety budget. This allows us first to identify policies that are safe on average but unsafe in the tail and then to study whether their tail violations can be controlled while preserving return. To identify tail-unsafe policies, we evaluate five standard algorithms on three Safety-Gymnasium navigation tasks. We then examine four constraint families on dense-hazard navigation and assess tail control across four navigation and four locomotion tasks.
cs.LG / 78 / 2610.09510
Physics-Informed Neural Plasticity: PDE Solvers That Reshape Themselves
Abstract
Physics-informed neural PDE solvers adapt their parameters to satisfy governing equations, yet their representational structure typically remains fixed throughout training. This rigidity is poorly matched to PDE solutions with strongly heterogeneous complexity across space and space--time, leaving capacity insufficient where the physics is difficult and redundant where it is simple. We introduce physics-informed neural plasticity, a paradigm in which the representation itself reshapes during optimization in response to unresolved physics. We instantiate this principle with Representation Capacity Adaptation for PDEs (ReCAP), a Gaussian-localized solver that dynamically redistributes capacity through local enrichment, residual-directed splitting, gate-based pruning, and function-aware merging. ReCAP uses responsibility-weighted error indicators and the geometry of residual energy to determine where and how to refine. To limit the disturbance introduced by splitting, we introduce quiet-child refinement, which initializes new components by transporting the parent representation while controlling instantaneous functional perturbation. We further establish conditional a posteriori reliability and structural-stability guarantees linking localized physics residuals to solution error and stable refinement. Across five challenging 3D and 4D PDE benchmarks against 11 physics-informed solvers, ReCAP achieves the lowest relative $L^2$ error on every problem, reducing error by $10.7\%$--$27.5\%$ relative to the strongest competing result. These results suggest that physics-informed solvers need not merely learn their parameters---they can learn how their representational capacity should be organized.
cs.LG / 79 / 2610.09519
Differential Refresh Policies for Models Trained on Lagging Data Snapshots: From a Single-Age Equivalence Limit to an Optimal Per-Segment Allocation
Abstract
Production machine-learning models are derived artifacts of time-bounded training snapshots: a deployed model is a materialized view over a training cut that ages the instant it is built. A common response is to replace the fixed retraining cadence with an adaptive trigger -- a weighted staleness score that retrains when accumulated source risk crosses a threshold. We show this is the wrong lever, and identify the right one. First, an equivalence limit: any refresh trigger that is a static, strictly monotone function of a single shared global training-data age is operationally equivalent to a calibrated uniform age timer, so a global staleness budget, however elaborately it weights segments, sources, and sensitivities, carries no scheduling information a clock does not. The limit also shows how to escape it: refresh segments differentially, giving each its own age and refresh interval, which is meaningful when refresh cost is separable across segments (incremental training or per-segment models). We solve the resulting budget-allocation problem. In the frequent-refresh regime each segment's optimal refresh rate is proportional to the square root of its risk $w_j λ_j$ (weight times change rate), and the optimal policy never costs more than the uniform timer, beating it by a closed-form Cauchy-Schwarz "price of uniformity" that is zero for homogeneous workloads and grows with heterogeneity. In a discrete-event simulation with real Poisson change events, the optimal policy lowers realized weighted stale exposure by 8-29% relative to the uniform timer at matched refresh budget, winning on 86-100% of seeds; a naive exposure-threshold policy does not, showing the allocation is what helps; and the advantage survives 50% rate-estimation noise. The leverage in model refresh is not a better score but a better action.
cs.LG / 80 / 2610.09525
Align Before You Combine: Reference Space Calibration for Supervision Without Ground Truth
Abstract
We introduce a calibration-first framework that produces supervision scores without access to ground-truth labels or a shared annotation space. Our framework aligns subset-specific scorers using a synthetic ordinal reference space before fusion. This reference space is constructed from ordered calibration features that represent the latent concept, providing a common scale on which otherwise incomparable scorer outputs can be aligned. Because our calibration procedure uses the reference space rather than training samples, it is independent of the training set's empirical distribution. Across three benchmark datasets, our framework consistently outperforms uncalibrated averaging and achieves higher primary-metric point estimates on the evaluation metrics than the best individual scorer. Performance relative to sample-dependent baselines varies by domain, with absolute differences below 0.02 on Ames Housing and below 0.01 on Breast Cancer Wisconsin and Wine Quality. After Bonferroni correction, differences remain significant for all three comparisons on Ames Housing and one on Breast Cancer Wisconsin. Additionally, we show that using fewer calibration levels per feature can closely approximate higher-resolution results at substantially lower computational cost. Together, these results support our framework as a viable approach to construct supervision scores when neither ground-truth labels nor a shared annotation space is available.
cs.LG / 81 / 2610.09549
CircuitGate: Logic-Consistent Circuit-Level Functional Modeling for And-Inverter Graphs
Abstract
And-Inverter Graphs (AIGs) are fundamental representations for logic synthesis and verification in Electronic Design Automation (EDA). As structured representations of complex digital systems, AIGs require models to capture functional dependencies beyond local structure and remain robust to functionality-preserving transformations. In learning-based AIG representation, existing approaches are predominantly based on GNNs and rely on local gate-level message passing, limiting their ability to capture circuit-level functional context and making the learned representations sensitive to topology-specific patterns. Therefore, we propose CircuitGate, a function-aware AIG representation learning framework that advances from gate-level semantics to circuit-level functional modeling. CircuitGate explicitly encodes global primary-input (PI) support and models support-overlap-aware reconvergence between fanins, while incorporating logic-inspired Boolean constraints to encourage functionally consistent representations. We evaluate CircuitGate on the large-scale ForgeEDA benchmark and further validate it on the EPFL and ITC'99 benchmarks. Across equivalent-gate identification and signal-probability prediction tasks, CircuitGate consistently outperforms existing methods, achieving up to 21.7% and 14.2% reductions in MAE, respectively. Under direct ForgeEDA-to-OpenABC transfer without fine-tuning, CircuitGate also achieves the best equivalent-gate identification performance, demonstrating strong cross-dataset generalization. These results demonstrate the effectiveness of modeling circuit-level functional dependencies beyond local topology.
cs.LG / 82 / 2610.09551
A Framework for the Systematic Review of ML Assets in AI Registries
Abstract
Background: Modern software systems increasingly rely on Machine Learning (ML) assets (i.e., pre-trained models, datasets, benchmarks) for building, evaluating, and integrating ML-based systems. However, current exploration, selection and reuse practices of ML assets are not supported by systematic retrieval methodologies comparable to those used in traditional evidence synthesis. Consequently, in practice, ML asset selection is often presented as a settled design decision, supported by informal justification rather than a traceable, evidence-based, and updatable selection process. Aims: This paper explores how systematic review methods can support ML asset retrieval. In doing so, we aim to make their selection transparent and reproducible, grounded in explicit evidence, and ultimately better suited to its intended use. Method: We analyze established systematic review practices from scientific literature and adapt their phases (i.e., planning, conducting, and documenting) to Artificial Intelligence (AI) registries, treating ML assets as first-class units of analysis. The resulting framework integrates registry-aware search strategies, cross-registry schema alignment, and dependency-driven ML asset exploration. Results: We conceptualize ML asset retrieval as a systematic and reproducible process rather than an ad hoc activity, and propose a framework for structured ML asset discovery. \textbf{Conclusions:} This work illustrates how systematic review principles can be extended beyond scientific literature to support evidence synthesis over evolving AI registries.
cs.LG / 83 / 2610.09559
UniCSI Towards a Universal Wi-Fi CSI Encoder for Ubiquitous Human Sensing
Abstract
Wi-Fi sensing promises to turn the everyday wireless signals that already surround us into ubiquitous sensors for human sensing. However, a fundamental obstacle is that CSI is acquired under diverse device-specific configurations, including different subcarrier counts, bandwidths, and carrier bands. Consequently, the resulting CSI tensors vary in both spectral resolution and tensor shape, making heterogeneous modeling challenging. Standard architectures struggle with such heterogeneity, forcing lossy pre-processing which compromises the underlying signal. To bridge this gap, we present UniCSI, a unified foundation architecture that directly operates on heterogeneous CSI while preserving the integrity of the native waveform. UniCSI hinges on two core innovations: (1) a physics-informed RF tokenizer that encodes each frequency channel based on its fractional position within the physical spectrum rather than rigid array indices. It preserves intrinsic spectral coherence and enables seamless, frequency resolution-agnostic processing across arbitrary sensing configurations. (2) a spectral aggregator that distills variable-length channel sequences into a fixed-size spectral signature, effectively decoupling the feature dimensionality from the physical subcarrier spacing. Extensive evaluations on a large-scale corpus of 25 heterogeneous public datasets, spanning 14 to 2048 subcarriers, 20 to 160 MHz bandwidth, and the 2.4 and 5 GHz bands, demonstrate that native heterogeneous ingestion substantially improves cross-domain transfer under both supervised and self-supervised training schemes, particularly in regimes where fixed-grid architectures fail to generalize.
cs.LG / 84 / 2610.09577
Online Resource Allocation with an Endogenous Markov State: Fewer LP Solves Earn More
Abstract
We study finite-horizon online resource allocation with i.i.d. requests and an endogenous Markov state on a finite state space: each action affects the transition of the state that governs future rewards and resource consumption. In this problem, a transient fluid LP benchmark upper bounds the expected reward of every nonanticipating policy, while a stationary LP supplies randomized state-dependent controls. We assume that the stationary LP has a unique optimum and identify primal nondegeneracy and irreducibility of the optimal induced kernel as important regularity conditions in this framework. With a known request prior, we show that, under nondegeneracy and irreducibility, both frequent and infrequent re-solving attain $O(1)$ regret. However, under a degenerate optimum, irreducibility yields the sharp worst-case $Θ(\sqrt{T})$ rate for infrequent re-solving, while frequent re-solving can incur $Ω(T)$ regret. Thus, more frequent optimization can perform asymptotically worse. With an unknown request prior, we develop a three-phase U-shaped infrequent re-solving policy that coordinates learning and inventory correction with $O(\log\log T)$ LP solves. When the optimal induced kernel is irreducible and the algorithm is given the optimal target state class and a constant-cost entrance policy, it attains $O(1)$ regret under nondegeneracy and $O(\sqrt{T})$ regret under degeneracy. Without the target-class information, linear minimax regret is unavoidable. Numerical experiments further illustrate the instability of round-by-round re-solving relative to epoch-wise infrequent re-solving, show that thresholding greatly mitigates its loss, and find that infrequent schemes remain dominant under both known and estimated priors.
cs.LG / 85 / 2610.09604
Lightweight and Versatile Learned Optimization by Recombination of Gradient History
Abstract
This paper presents a lightweight and versatile learned optimizer that dynamically recombines gradient history, represented as averages over disjoint time spans. The optimizer reduces the prediction space to one scalar coefficient per gradient average, shared by multiple parameters. Progressively averaging older gradients minimizes memory cost of long history, while keeping their contributions independently accessible. A 37k-parameter network trained in 0.87 GPU-hours generalizes zero-shot to unseen tasks, lowering validation loss by 9.1% and 0.4% on BERT-Tiny and GPT-Tiny, and improving test accuracy over Adam by 3.5 %p on a Vision Transformer and by 2.7 %p on average across nine graph models, with FLOPs overhead as low as 0.3%.
cs.LG / 86 / 2610.09609
Decoupled Optimization for Teacher-Student Semi-Supervised Learning via a Pioneer Student
Abstract
Semi-supervised learning (SSL) relies on two core mechanisms: self-training under the Teacher-Student (T-S) framework and joint optimization of labeled and unlabeled losses. Despite their effectiveness, we find both mechanisms introduce distinct optimization pathologies. First, parameter coupling enforces strict synchronization between teacher and student, where strong regularization on the student degrades the teacher's fitting ability, thereby limiting the permissible generalization intensity. Second, the imbalance in gradient update consistency between labeled and unlabeled losses drives the shared parameters to prematurely converge to labeled-dominated local minima, creating a bottleneck for global optimization. To address both issues, we propose the Pioneer Student (PiS), an auxiliary branch that operates in an independent parameter space and periodically transfers accumulated knowledge back to the T-S model. Extensive experiments show that PiS is a universal plug-and-play module that consistently improves mainstream SSL methods.
cs.LG / 87 / 2610.09611
Sequential Pretraining Favors Large Models
Abstract
Large neural networks often acquire capabilities that small models fail to learn. Does this stem from large models learning more representative features, or from being more robust to unaccounted-for adverse effects introduced during training? We define and quantify one such adverse effect, primacy bias, as the extent to which exposure to early data distributions impairs later learning. We show that small models can allocate learning capacity inefficiently toward early distributions, whereas sufficiently overparameterized models are robust to this effect. This inefficiency is particularly consequential in pretraining, where foundation models often encounter heterogeneous data distributions sequentially rather than jointly. As a result, small foundation models can struggle to learn distributions encountered late in training, which is particularly harmful when later data emphasizes desirable capabilities such as code, mathematics, and reasoning. Motivated by these findings, we introduce Exposure Therapy (ET), a simple regularization that promotes more efficient allocation of learning capacity during sequential pretraining. We demonstrate that ET improves foundation models' performance on late data distributions as well as overall capability in models up to the billion-parameter scale. Overall, our results suggest that some benefits of large foundation models may arise from greater robustness to adverse training effects, rather than from learning more representative features, and that improved training algorithms can recover some of these advantages in smaller models.
cs.LG / 88 / 2610.09620
The Identifiability and Observability of Deep Normalized Attention
Abstract
We study which parameters of deep, unmasked, single-head attention are determined by its input--output function. For known positive nonconstant real-analytic normalizers, the function generically determines the effective scores and combined value map up to the signs induced by even normalizers. This proves the real-analytic case of a conjecture of Henry--Marchetti--Kohn, including softmax. We then classify exceptional fibers under explicit normalizer conditions, identifying when collapse makes later scores unobservable, and establish sharp Taylor orders for local identification. Near simultaneous query/key collapse, we compute the complete native Jacobian decay spectrum on separating finite input banks. For common first nonconstant normalizer degree $k$, layer $i$ has contact order $2k3^{i-1}-1$, with exact multiplicities and kernel dimension. High-precision and automatic differentiation calculations illustrate the resulting loss of numerical sensitivity.
cs.LG / 89 / 2610.09621
When does a network's training history predict its future learning better than its current state? Evidence from a response probe and a forecasting screen
Abstract
Networks that behave alike now can still learn differently when training continues. Work on loss of plasticity and critical periods shows that the path to a state shapes what follows; it does not show whether the path carries information that a measurement of the state itself misses. We ask when the training history of a network predicts its future learning better than its current state. In a main study, small multilayer perceptrons were trained under three history regimes (42 histories), and future learning was measured at four checkpoints by a short probe: a copy of the network trained for 100 updates on a new task. Before the prediction result was read, the protocol checked the probe. It responded monotonically to a function-preserving rescaling of hidden units, repeated measurements agreed (intraclass correlation 0.940, [0.903, 0.997], in the least reliable class, mean of three repeats), and a re-initialisation of units was visible directly after it but not 100 to 200 updates later. A history state of at most four dimensions did not improve on a calibrated model of the current state (gain -21.4%, 90% interval [-91.9, 8.1]; required in advance: 10%). A companion screen on 1,560 synthetic regression runs asked the same question for a target further away, the final error of the run. There, history models forecast better than the current validation error after 12 of up to 240 epochs (compact state 30.3%, [15.8, 39.4], a contextual comparison) and were not distinguishable from it after 48. In both studies the history was informative only while the current state was not yet informative about the target; this reading was formed after the results.
cs.LG / 90 / 2610.09633
Coding-Agent Benchmarks Should Match Their Users' Task Flows
Abstract
The evaluation of coding agents generally strives to be as realistic as possible. In our study, we collect 4,782 agent sessions of real software engineers in JetBrains IDEs, which we call Production Sessions. Since our subject is interactive agents, we study the sessions with at least three user messages (33% of the sample). These long sessions differ from issue-derived benchmark tasks in two ways: (i) user requests span a far wider mix of task types - questions about the project's code, planning, review, refactoring, execution - and (ii) users switch between types throughout a session. Long-session samples from three public interaction corpora exhibit markedly different Task Flows (the distributions of session lengths, task types, and type-to-type transitions), so no single interaction distribution is universally realistic: benchmarks should name a target use case and calibrate to measurements from it. We present SWE-TaskFlow, an approach for transforming any issue-derived benchmark: it preserves the verified tasks and tests while steering the interaction toward a target Task Flow through prompt splitting and verifiable repository QA, with a TaskFlow Alignment Score (TFAS) for selecting among generated trajectories. In a pilot on 700 SWE-Bench Pro tasks, solving the task sequentially in several steps approximately doubles agent cost without a stable change in resolve rate: the interaction protocol itself is an important dimension of evaluation.
cs.LG / 91 / 2610.09647
When Rank Rises as LLMs Degrade
Abstract
Post-training adapts language models in non-stationary environments. Practitioners monitor representation health with RankMe and related spectral statistics, often assuming that rank falls when representations degrade. We show that this assumption is unsafe for LLM post-training. In a controlled study of Qwen3-0.6B with four degradation modes and three seeds, data duplication worsens held-out loss by 75% relative to healthy while increasing both original and centred RankMe; the latter changes by 13.5 pooled standard deviations. Covariance effective rank rises to nearly twice its healthy value. This failure is spectral dispersion rather than collapse, so a one-sided monitor rates the worst checkpoint as the healthiest. By contrast, a learning-rate misconfiguration lowers centred RankMe and k95, while uncentred RankMe is inconsistent across seeds. Direction is therefore a property of the regime-statistic pair and cannot be fixed by recalibration alone. We also distinguish two often-conflated statistics: RankMe normalises singular values, whereas covariance effective rank normalises eigenvalues. On raw intermediate-layer states in the pretrained model, massive activations pin the latter near 1 out of dimension d while RankMe retains usable range. We then test a two-sided, multichannel sequential monitor with separate calibration and test data. In a pre-registered shared-prefix, leave-one-seed-out evaluation, it detects all three damage regimes in every fold 10 to 60 steps after the fork and separates dispersion from downward-rank damage by firing direction. However, it never precedes held-out probe loss, and calibration with two seeds produces false alarms on the held-out healthy seed. Spectral monitoring can diagnose failure regimes, but it does not warn earlier than held-out loss, and validity claims require held-out healthy data.
cs.LG / 92 / 2610.09651
Closed-Form Noise Calibration Against Membership Inference for Random-Allocation DP-SGD
Abstract
DP-SGD protects training data by adding Gaussian noise to clipped gradients. The amount of noise is usually chosen by running a numerical privacy accountant inside a search. We study DP-SGD with random allocation, where each epoch uses every record once, at a randomly chosen step. For this setting we give a one-line formula that bounds the accuracy of every membership inference attack (MIA) on the trained model. With $M$ steps per epoch, $E$ epochs and noise multiplier $σ$, and with membership and non-membership equally likely a priori, the attack accuracy is at most $\frac12+\frac14\sqrt{(1+(e^{1/σ^2}-1)/M)^E-1}$. The formula comes from the chi-square divergence between a Gaussian distribution and a Gaussian mixture that dominates random allocation. It is interpretable and gives $σ$ in about a microsecond. Where applicable, our formula needs at most about half the noise of the state-of-the-art closed-form bound. To measure how close the bound is, we also derive an exact expression for the attack accuracy of these two distributions and evaluate it numerically. Calibrating to this exact expression requires $13.0\%$ to $20.2\%$ less noise than the formula in our main experiments, and since it is exact, no accountant that knows only $M$, $E$ and $σ$ can certify a smaller $σ$. In training, the resulting $σ$ outperforms the formula and matches a published accountant in test accuracy. It is found in seconds and certified in minutes, whereas every search we ran with that accountant took longer or returned at least $0.62\%$ more noise. We show that MIAs on the trained models stay below the bound.
cs.LG / 93 / 2610.09654
DSTNet: Dynamic Spectral Trajectory Network for Causal Multi-Horizon Financial Forecasting
Abstract
Wavelet-based financial forecasters typically use the transform only to denoise, or reduce it to a single spectral snapshot at the forecast origin, and the convolution that produces the coefficients is usually bilateral, so it can read past the forecast origin. DSTNet instead retains the recent evolution of filter-bank magnitudes as a causal Dynamic Spectral Trajectory, built from seven trailing technical indicators over a twenty-day lookback with a one-sided Morlet-derived filter bank and an explicit burn-in for the left-boundary transient. A factorized Scale-Temporal Spectral Transformer attends along the time and filter-bank axes separately, a learned gate fuses the spectral branch with a CNN-BiLSTM, and horizon-specific gates emit one, three, five, and ten day forecasts in a single pass. We evaluate seven equity indices and gold under a common expanding-window protocol and an untouched one-year hold-out, against nine learned baselines and a random-walk persistence benchmark. Under MAE and MAPE, persistence is the strongest of the ten fixed competitors in 29 of the 32 series-horizon cells and DSTNet is the only model below it in every cell, by 0.7 to 0.9 percent at one day and 3.4 to 4.5 percent at ten days. At one day, paired testing favours DSTNet against the weaker learned baselines but is inconclusive against persistence and the strongest learned forecasters. A downstream allocation diagnostic does not support an equity-timing advantage on any of the seven indices.
cs.LG / 94 / 2610.09675
Gauss-Newton Accuracy and Indefinite Hessians: Uniform Coexistence in Low-Cost Sets
Abstract
We study the accuracy of Gauss-Newton curvature in ridge-regularized nonlinear least squares. Under local regularity and persistence of level-set curvature magnitude along an exact-fit section, we prove uniform coexistence of two curvature regimes. Global minimizers exist, and every global minimizer has relative Hessian error below $(1+\sqrt2)/8$, while the same low-cost set contains a point with an indefinite Hessian and relative error at least $15/8$. One positive ridge cap works for all independent center and label perturbations in fixed neighborhoods and every positive ridge weight up to the cap. These neighborhoods do not shrink as the ridge weight tends to zero. A pointwise certificate based on the current prediction level set controls the normal, mixed, and tangent parts of the Hessian correction. We prove a sharp relative-error bound over the stated pointwise class when the prediction map and ridge vary. Analytic examples describe the roles of output alignment, curvature orientation, and persistence. A separate structural result gives full Jacobian row rank throughout low-cost sets and exact interpolation near a rank-deficient reference.
cs.LG / 95 / 2610.09679
CERO: Where and When to Allocate Rollouts for RL Post-Training
Abstract
Adaptive rollout methods for group-relative reinforcement learning typically allocate a fixed per-update budget across prompts. We instead study how to coordinate a finite rollout budget over the entire training horizon. We formulate this problem using a concave surrogate utility of cumulative prompt exposure and introduce CERO, an online primal dual scheduler for prompt admission and budget pacing. In our experiments, each admitted prompt receives a fixed-size response group. CERO instead adapts which prompts are selected, how often they are revisited across rounds, and how many groups are generated in each round. A compact Fenchel representation linearizes the dependence on cumulative exposure, while projected online gradient descent updates prompt-specific supporting slopes and a shared budget price using reward-variation feedback and budget deviations. We establish pathwise guarantees for the surrogate allocation objective against fixed-rate and same-path time-varying benchmarks, with explicit terms for proxy discrepancy and rate variation. Under matched training-response budgets, CERO attains the highest avg@16 macro-average on each of three backbones across five mathematical reasoning benchmarks. Mechanistic analyses link CERO's prompt choices to within-group reward contrast, while multi-seed ablations show gains from adaptive pacing over both uniform and preset spending schedules.
cs.LG / 96 / 2610.09692
Few-Shot Learning for Personalised Automated Pain Assessment
Abstract
Pain perception varies substantially across individuals, making it difficult for population-based classifiers to generalise across all subjects in a dataset. One way to account for subject variability is to train personalised classifiers. In this work, we evaluate Few-Shot Learning, a sub-area of Meta-Learning, as an approach to personalisation in automated pain assessment. We re-interpret the shift from population-level to subject-level evaluation as a task-domain shift, where the observed classes remain fixed but the target subject changes. We evaluate our method on the BioVid Pain Database, the SenseEmotion Database, and the PainMonit Experimental Dataset (PMED), reaching 85.75% and 35.49% accuracy on BioVid and 82.37% and 41.88% on SenseEmotion in the binary and multi-class settings under a Leave-One-Subject-Out CV protocol respectively, and 90.47% on PMED, for which only a binary benchmark exists. Using samples to implement k-shot conditioning, the accuracies can be improved to 86.25%, 40.06%, 83.43%, 44.08%, and 91.25%, respectively. To further evaluate the effects and robustness of our method, we provide additional ablation experiments and investigate the personalisation effects. Our results suggest that support-conditioned few-shot adaptation can improve average performance under inter-subject variability.
cs.LG / 97 / 2610.09709
Pretraining Shapes Spectral Structure: Architecture- and Strategy-Conditional Prediction of OOD Robustness in Foundation Models
Abstract
Can we determine whether a foundation model will generalize out-of-distribution (OOD) before any target data is available? Existing diagnostics require source or target data, which rules them out before a target domain exists. Those that use the weights alone apply one statistic to every architecture, and do not separate robust models from fragile ones. We show the answer is encoded in the spectral structure of pretrained weights. Two forces shape that structure. Architecture determines how information is stored in weight matrices. Pretraining strategy determines what is rewarded. Together they set a spectral geometry that governs OOD robustness. We prove that the OOD accuracy gap is bounded by how tightly the source representations concentrate. A statistic computed from the pretrained weights alone serves as a proxy for that concentration. The direction of that proxy reverses between architecture families. We operationalize it: the direction is stable within one (architecture X strategy) combination, the finest grouping we test, which we call a cell. Pooled over 116 models spanning 7 modalities, a single statistic ranks OOD robustness weakly, because cells of opposite direction cancel. Within a cell, the statistic selected for it orders 92% of model pairs by OOD robustness in-sample. The selection does not leak the target: for each model family outside the matrix we logged the cell, metric and sign before running its OOD evaluation, and the predicted direction held in every case: EEG, genomic and protein. Acting on spectral concentration narrows the OOD gap by 24% at 87.5% ID retention. The diagnostic operates on released weights alone, so OOD robustness becomes checkable at model-selection time, before data or compute is committed to a target domain.
cs.LG / 98 / 2610.09715
EC-EarthFlow: Probabilistic emulation of daily transient global climate model simulations with flow matching
Abstract
We introduce EC-EarthFlow, a generative flow matching model that emulates simulations from the physical climate model EC-Earth3. The model is trained on transient simulations from EC-Earth3 (1950-2166, SSP2-4.5) to predict the day ahead temperature field from the previous days temperature as well as annual mean temperature. Predictions are made auto-regressively with rollout periods of between a month and an extended season. Using only this variable of interest, we are able to reproduce the daily variability, spatial patterns, annual cycle and long-term trend from EC-Earth3 at a substantially lower computational cost than the physical model. We demonstrate that EC-EarthFlow is stable for long inference periods, and that it can learn the physical relationships as simulated in EC-Earth3.
cs.LG / 99 / 2610.09752
SoftSEEPS improves ML-based precipitation forecasting
Abstract
In this paper we have developed a differentiable approximation of the well-known SEEPS score, which we name SoftSEEPS. This allows the training of a Machine Learning model to forecast precipitation directly. We test SoftSEEPS on the IMERG dataset (0.1 degree resolution) by training a decoder for precipitation on the latent space of a pre-trained low-resolution forecasting model. Combining SoftSEEPS and RMSE in a joint objective is possible with marginal trade-offs in either metric.
cs.LG / 100 / 2610.09757
EntroPrefill: Renyi-Guided Context Pruning with Conditional Stability Guarantees for Retrieval-Augmented Generation
Abstract
Mid-prefill pruning can reduce the sequence processed by deeper transformer layers, but attention concentration alone does not certify that discarded context is dispensable. We formulate EntroPrefill as a Renyi-guided proposal mechanism coupled to explicit constraints on discarded attention mass. Sink-isolated, regularized head pooling respects grouped-query attention while exposing a quantitative trade-off between specialization and worst-head coverage. We derive a mixture-to-head deletion envelope, a computable upper bound on feasible token removal, and a finite-sample observer guarantee that remains valid when the pruning layer is selected adaptively. We then establish a conditional transformer perturbation bound with explicit sufficient Lipschitz constants and a first-token decision-margin corollary. A counterexample shows why shallow observations alone cannot imply an unconditional future-output guarantee. The systems analysis distinguishes query-head unions, physical page allocation, and KV-transfer payload, and gives an arithmetic break-even condition for pruning. This manuscript is theoretical in scope: it defines the procedure, its assumptions, and its formal limits, but does not report measured acceleration or task-accuracy preservation. Experiments are reserved for subsequent validation of the assumptions, approximation tightness, and end-to-end resource trade-offs.
cs.LG / 101 / 2610.09759
Unrolled Flow Models for Reasoning
Abstract
Flow matching enables language generation in few steps, but whether additional integration steps improve reasoning remains unclear. We prove that a flow parameterized by a two-layer Transformer can solve graph reachability, with the required number of integration steps increasing with the target's distance from the root. Yet, standard flow language models can fail to benefit from additional steps on reasoning tasks. We attribute this limitation to objectives that supervise each time point independently, without explicitly training successive steps to build on one another. To address this, we instead train through the model's own latent rollout over a randomly sampled subinterval of [0, 1], decoding only at the endpoint. On ProsQA, this raises accuracy to 97% and enables performance to improve with additional integration steps. For the longer rollouts required by reasoning tasks such as Sudoku and Maze, retracting the latent state onto a sphere stabilizes the dynamics and yields substantial gains over baselines with more than three times as many parameters. Sampling multiple rollouts further improves performance when paired with a parameter-free selection score, although reliable selection remains challenging for longer answers. Together, these results establish a theoretical basis for reasoning with flows and show how rollout training, stable latent dynamics, and rollout selection help realize this capacity in practice.
cs.LG / 102 / 2610.09760
Leaner Transformers Can Easily Learn to Cluster
Abstract
Transformers have in-context learning capabilities, where some known learning algorithms can be executed in the forward pass through the model. Recent work shows that transformers can exactly perform Lloyd's algorithm for $k$-means clustering with $n$ points in $d$ dimensions with an embedding size $d_{\textsf{emb}} = d+k$ (thus, requiring attention projection matrices of size $(d+k)^2$). In this work, we build upon this result in the following ways: First, we present an equally expressive but smaller transformer that executes Lloyd's algorithm with embedding size $d_{\textsf{emb}} = (d + \lceil \log_2 k \rceil)$. Next, we train these transformers to learn the clustering algorithms given a distribution of clustering tasks, and theoretically characterize and empirically validate the factors affecting the convergence and in-distribution generalization of learning algorithms based on stochastic gradients. Finally, we probe the general clustering abilities of these learned algorithms (in the form of transformers), and try to understand situations where they succeed and fail.
cs.LG / 103 / 2610.09763
Beyond Policy Support: Interaction Constrained Offline Reinforcement Learning for Autonomous Driving
Abstract
Offline reinforcement learning enables reward-driven policy improvement from fixed datasets without requiring online exploration, making it particularly attractive in safety-critical domains. A central challenge, however, is distribution shift: policy optimization may favor actions that are weakly supported by the offline data, rendering value estimates unreliable. Existing approaches primarily control this shift in the policy's own action space. In interactive environments such as autonomous driving, this can be insufficient: a candidate ego trajectory may remain well supported under the marginal behavior distribution while being poorly supported jointly with the surrounding-agent behavior observed in the logged interaction. We refer to this degradation in interaction support as \emph{interaction distribution shift} (IDS), and introduce \emph{Interaction-Constrained Drive Policy} (ICDP), an offline reinforcement learning framework that explicitly controls interaction-level distribution shift. Starting from the joint data distribution over ego and surrounding-agent futures, we show that joint-support degradation decomposes exactly into an ego-support component and a residual interaction-support component. We recover the latter through contrastive density-ratio estimation, isolating interaction compatibility without explicit joint-density modeling, surrounding-agent prediction, or rollouts in reactive simulators or learned world models during policy optimization. Closed-loop evaluations on nuPlan, Interplan and real-world truck experiments show that ICDP suppresses high-value yet interaction-unsupported trajectory selections and improves performance in interaction-critical driving scenarios. Project webpage: https://mahmoud-selim.github.io/ICDP/
cs.LG / 104 / 2610.09768
Fluctuations of Nonlinear Observables in Mean Field Neural Network Training
Abstract
Mean field limits describe the training dynamics of wide neural networks through the evolution of the empirical distribution of their parameters. Although functional central limit theorems characterize the asymptotic fluctuations of this distribution, quantities of practical interest are typically nonlinear observables of the parameter distribution rather than the distribution itself. In this work, we show how these mean field fluctuations propagate to finite dimensional nonlinear observables for shallow neural networks trained by stochastic gradient descent. Working in the weighted Sobolev space in which the limiting fluctuation process is constructed, we apply a functional Delta method under ordinary Fr{é}chet differentiability, without requiring Lions derivatives with respect to the measure variable. We obtain a central limit theorem for the observables and, under a suitable representation of their differentials, an explicit covariance formula inherited from the underlying mean field fluctuation theory. We also study whether prescribed quantities of interest can be recovered from the selected observations. Under a constant rank assumption, we prove that a quantity of interest factors locally through the observation functional if and only if, throughout a neighborhood, the kernel of the differential of the observation is contained in that of the quantity of interest. Thus, a differential condition expressed directly in the ambient Sobolev space yields an exact nonlinear local factorization. These results provide a framework both for quantifying finite-width uncertainty on observable, statistically or physically meaningful quantities and for assessing whether the chosen observations contain the information required to identify them.
cs.LG / 105 / 2610.09782
AdaPS-LiNGAM: Adaptive Predecessor Selection for Linear Non-Gaussian Acyclic Models under Small-Sample Settings
Abstract
Causal discovery becomes particularly challenging when the available sample size is small relative to the number of variables. This challenge also arises in the linear non-Gaussian acyclic model (LiNGAM), an identifiable framework for causal discovery from observational data. DirectLiNGAM estimates a causal order, which arranges variables so that causes precede their effects, by sequentially identifying an exogenous variable and removing its linear effect from the remaining variables. We establish a structural limitation of this procedure: when the number of variables exceeds the sample size, repeated residualization necessarily becomes degenerate before the full causal order can be determined. Our analysis further reveals that each residual can be reconstructed using only a graph-determined subset of variables already placed earlier in the causal order, termed the active boundary. This result motivates AdaPS-LiNGAM (Adaptive Predecessor Selection LiNGAM), which reconstructs each residual directly from the original observations using an adaptively chosen sparse subset of those earlier variables. The same subset-selection principle is also applied to the final pruning step for edge estimation. Experiments on synthetic data demonstrate that AdaPS-LiNGAM provides accurate causal-structure recovery in sample-limited settings and degrades more gradually as the sample size decreases.
cs.LG / 106 / 2610.09792
A Proof-of-Concept Study of Weakly Supervised Labeling of Fine-Grained EEG Components for Artifact Attenuation
Abstract
Electroencephalography (EEG) is highly susceptible to electromyographic (EMG) artifacts, whose temporal heterogeneity and spatial-spectral overlap with neural activity can leave mixed sources after blind source separation. Existing artifact-removal methods are further limited by scarce reliable component-level ground truth: expert annotations are costly and subjective, while no established method provides realistic simulation-based ground truth for EMG contamination in multichannel scalp EEG. To address these limitations, we propose a framework combining a frequency-aware high-dimensional representation with Multi-Instance Learning. The representation unfolds separated components into frequency-resolved intra-components, creating a space in which mixed neural and muscular activity becomes more separable, while the weakly supervised learning formulation enables artifact-likelihood scores for individual intra-components to be learned from epoch-level labels without finer-grained ground truth. The resulting intra-component classifier supports fine-grained EMG artifact detection and score-guided attenuation. Experiments on held-out subjects show that the framework learns informative intra-component scores and reduces artifact-related spectral deviations most clearly for jaw tension, with moderate effects for raising eyebrows and limited effects for frowning.
cs.LG / 107 / 2610.09824
Homogenization in Multi-Agent Systems
Abstract
Multi-agent systems (MAS) leverage interactions between agents to perform complex tasks. Despite their success, we show that these interactions can also lead to homogenization, i.e., agents converging to similar behaviors. Homogenization in MAS can reduce agent diversity and reinforce shared failures. In this paper, we operationalize homogenization using three metrics: conformity to the majority, polarization towards extremes, and growing inertia against changes over subsequent interactions. We evaluate homogenization in MAS for code generation, hiring, and scientific peer review. Across these tasks, we show that homogenization translates to concrete downstream risks: in code generation, it hides and amplifies correlated errors which can create systemic vulnerabilities; in hiring, it allows the influence of biased agents to persist long after their removal; and in peer review, it creates uneven evaluation standards across research areas. Our results establish homogenization as a failure mode of MAS, demonstrating that MAS evaluations must move beyond aggregate performance to carefully analyze interaction dynamics. Finally, we show that simple approaches to increase diversity---leveraging sampling stochasticity and mixed-models MAS---fail to reduce homogenization risks, highlighting the need for strategies to effectively leverage agent diversity.
cs.LG / 108 / 2610.09827
Dual-QK: Sharp Queries and Flat Keys for Prunable 2-bit KV Caches
Abstract
Long inputs and extended generation increase the storage and access costs of the key-value (KV) cache. Low-bit quantization reduces storage and memory traffic, while query-channel pruning can further reduce key-cache reads. Rotation-based quantization redistributes the energy of key outliers across channels. To maintain computational invariance, the same orthogonal transform must be applied to queries, preserving query-key dot products. However, this rotation can disperse query energy, weakening the separation between a few large components to retain and many small ones to prune. We introduce Dual-QK, which uses paired non-orthogonal query and key transforms to address this conflict. Using calibrated query and key statistics, Dual-QK combines partial key whitening with a query-aligned basis to balance key scales for INT2 quantization and concentrate query energy for dynamic channel pruning. Channel-0 protection and bucket-relative RoPE support low-bit accuracy over long contexts. Experiments on four models across five generative benchmarks and long-context retrieval tasks show improved accuracy over OSCAR on most tasks at 40% query-channel sparsity. At a 128K context, Dual-QK provides $6.8\times$ KV-cache compression and an estimated $8.3\times$ reduction in KV read volume relative to unpruned BF16. Under the evaluated configurations, our SGLang implementation achieves up to $3.75\times$ the decoding throughput of unpruned BF16.
cs.LG / 109 / 2610.09838
Fully Interpretable Minimal Transformers: From Geometry to Algorithm
Abstract
We present a framework for building and interpreting minimal transformer models. By constraining a transformer's embedding dimension and head size to 2, we enable full two-dimensional visualization of its internal representations. Embeddings, query/key/value transforms, attention outputs, residual streams, and decision boundaries can all be seen directly. Our central claim is that the learned geometry implies an algorithm; the arrangement of points and boundaries in R^2 can be read as a step-by-step procedure. We train a transformer on a simple task where it must produce the most recently observed even number whenever the '+' operator appears in a sequence of digits. Once trained, we visually walk through every step of the transformer's computation. We show how the model embeds the tokens and their respective positions in the sequence, transforms them via the Q, K, and V matrices, uses the dot product between the Q and K representations to form the attention matrix, and uses the attention matrix to select values that move the representation of each input token to the region of the domain of the output layer that will correctly predict the next token. We introduce a suite of interpretability visualizations that make the algorithmic interpretation of this procedure explicit. Our framework offers a pedagogical and experimental testbed to explore how transformers use informational geometry to implement next-token prediction.
cs.LG / 110 / 2610.09844
For Those Who Believe in Faithfulness: Optimizing the Area Under Insertion and Deletion Curves for Ranking Relative Feature Importance
Abstract
The adoption of machine learning for socially relevant tasks requires effective explainable artificial intelligence (XAI) methods to better understand the behavior of machine learning models. Attribution methods are a popular XAI approach in which input-output relationships are characterized by heat maps that reflect the relative importance of input features for a particular prediction. The quality of such maps is often assessed by measuring faithfulness based on the area under insertion and deletion curves, which measures changes in the model output as features are added and removed. In this study, we derive an objective function from this notion of faithfulness and a way to approximate its gradient. We establish the connection between insertion curves and top-$k$ feature selection, which leads to a loss function measuring the quality of attributions. Randomization of the loss allows us to efficiently approximate its gradient. To show the effectiveness of the general approach, we combine the loss function with the neural explanation mask framework. The resulting method, termed Ra-NEM, can be used with any differentiable model without affecting the model's performance. Experiments demonstrate that Ra-NEM provides accurate attributions robustly and efficiently. Compared to other algorithms, the attributions have not only higher faithfulness but also perform well in terms of other XAI metrics. The high inference speed of Ra-NEM makes the method suitable for online applications. The code is available online: https://github.com/baerminator/Ra_Nem
cs.LG / 111 / 2610.09848
Stream-Based Active Learning with Cooperative Neural Networks for Data-Efficient Partial Inverse Design: An Automotive Glass Run Channel Case Study
Abstract
Inverse design in engineering often runs into a simple problem. Each labeled training sample must be produced through expensive simulation, so building a large dataset is slow and costly. This study addresses that problem for partial inverse design, where only some design variables are specified and the rest must be inferred to reach a target performance value. We propose CoNN-AL, a framework for data-efficient partial inverse design that adds stream-based active learning to the Cooperative Neural Network with Denoising Autoencoder (CoNN-DAE). The model estimates predictive uncertainty through Monte Carlo dropout and uses it to decide, in real time, which incoming candidate samples are worth labeling, so the limited labeling budget is spent on the most informative designs. We validate the framework on a real-world automotive glass run channel dataset of more than 900,000 unique simulated designs. With only 20,000 actively selected labels, about 2.3% of the training pool, CoNN-AL reaches R-squared values of 0.967 to 0.982 across all missing-variable levels, approaching the upper-bound models trained on far more data. It reaches R-squared of at least 0.95 with 30 to 40% fewer labels than random sampling at the more difficult missing-variable levels and, at the most challenging level, is the only strategy in this study to reach R-squared of 0.98. Together with this work, we publicly release the dataset to support future research on data-driven design.
cs.LG / 112 / 2610.09866
MUNITE: Unified Multimodal Latent Inference for Any-to-Any Multimodal Generation
Abstract
We introduce MUNITE, a latent-variable framework for flexible any-to-any multimodal generation that treats encoding and latent generation as the same inference problem under different amounts of observed evidence. Given any subset of modalities, MUNITE models the conditional distribution over the latent representation associated with the complete observation. Full observation recovers deterministic encoding, no observation recovers the latent marginal, and intermediate subsets define conditional latent inference, all within a single conditional flow model. A shared latent sample captures variation that must remain consistent across generated targets, while modality-specific generative decoders model the remaining uncertainty independently. To learn these conditional distributions from incomplete training examples, we extend conditional flow matching through self-distillation: predictions conditioned on richer available observations supervise the same model conditioned on smaller subsets at the same intermediate latent state. When the richer-evidence trajectory follows the exact conditional flow, this provides the same expected learning signal as full-target denoising. Across PolyMNIST-D-Q, FFHQ64, and image-text-audio, MUNITE achieves competitive or better generation quality and source-target alignment, with higher joint-generation coherence. In particular, it attains the highest coherence in all one-to-many and unconditional image-text-audio comparisons, showing the effectiveness of unified latent inference across diverse multimodal settings.
cs.LG / 113 / 2610.09889
Identifiability of a dissipative knowledge-dynamics model: exact recovery under designed excitation, degeneration on observational data
Abstract
Human learning is a dissipative dynamical process: mastery accumulates through practice, decays through forgetting, and propagates across interdependent concepts. We model it as a nonlinear dissipative system of ordinary differential equations whose parameters are mechanistically meaningful (a concept-transfer matrix encoding prerequisite coupling, per-concept forgetting rates, and a saturating practice-response gain), and we study when those parameters can actually be recovered from data. We prove a structural identifiability theorem for the associated inverse problem under explicit excitation conditions, with constructive closed-form recovery for the two-concept case, together with monotonicity, robustness and L-stability results. We derive a semi-implicit L-stable scheme for the dissipative subsystem and a batched solver numerically equivalent to the per-trajectory formulation (bit-exact predictions, gradients to $10^{-10}$) yet two orders of magnitude faster, making estimation feasible on cohorts of $10^5$ learners. The empirical study is two-sided. Under the theorem's excitation conditions, synthetic recovery is exact: parameters to machine precision, prerequisite structure at $F_1 = 1.0$. On large observational benchmarks it is not. An apparently strong recovery, with forgetting rates correlating with topic difficulty at Spearman $ρ= 0.83$, is refuted by four independent controls: it survives destroying the temporal order of the data, is matched by a classical Bayesian baseline, and is unaffected by removing real timestamps. We trace this to the stationary structure of the model and show that it is the degeneration the theorem predicts in the absence of designed excitation. The result delineates a sharp boundary between identifiable and unidentifiable regimes and yields a validation protocol for interpretability claims.
cs.LG / 114 / 2610.09916
NeuralZip: Reusable Setup for Fast Lossless Compression
Abstract
Lossless compression can reduce the storage and movement of model weights without changing their floating-point values, but repeated statistical analysis and code construction add computational overhead. We study whether the statistical structure of exponents can be prepared once and reused. For this, we introduce NeuralZip, which groups chunks with similar exponent distributions, shares Huffman codes, and selectively represents recurring exponent tuples using packed exponents, thereby achieving additional moderate compression ratios. A setup chooses these representations before subsequent encodings, while every encoding still processes the current tensor values. In floating-point model checkpoints, post-setup compression is 1.81-21.33$\times$ faster than the baselines and achieves exact bit-to-bit reconstruction. We show that this setup can be precomputed and transferred from another compatible architecture, preserving similar compression ratios and avoiding the need to amortize setup costs. Therefore, compression adaptation is transferable and reusable. Training checkpoints demonstrate continued reuse as the weights evolve. Finally, GPU experiments reduce active memory usage by up to 27.5$\%$ while reproducing the logits exactly.
cs.LG / 115 / 2610.09919
Eigenvalues of the Hessian in Deep Learning: The Origin of Symmetry and Its Breaking
Abstract
Hessian spectra at trained models in deep learning exhibit a persistent pattern: eigenvalues organize into distinct clusters, including a large bulk near zero and a few isolated outliers. This paper shows that a natural account of these spectral phenomena emerges when the original setting is understood as a departure from a nearby, otherwise hidden, highly symmetric reference. Modifications, including changes to the architecture, data distribution, or parameter metric, expose a nearby reference configuration whose Hessian exhibits rich invariances-ones not accounted for by weight symmetries. There, symmetry enables a precise description of the spectra, forcing high-dimensional kernels and eigenvalues of large multiplicity. Returning to the original configuration breaks the Hessian symmetry and thereby produces the observed hierarchy of clusters and outliers. The framework is developed in some generality, with a detailed analysis of three-layer ReLU networks and applications to convolutional, graph, and transformer models, as well as to the NTK. The same mechanism is further shown to yield analogous spectral structures in layerwise Hessians and the Gauss-Newton matrix.
cs.LG / 116 / 2610.09927
KGATE : a Knowledge Graph Embedding Training Environment
Abstract
Knowledge graph embedding (KGE) models encode the entities and relations of a knowledge graph into a low-dimensional latent space, enabling tasks such as classification or link prediction. Most KGE models follow an autoencoder architecture, in which an encoder projects the knowledge graph into the latent space and a decoder reconstruct it. Combining both encoder and decoder components is increasingly needed, yet existing libraries rarely support complete autoencoders, are often unmaintained, rely on undocumented default hyperparameters, and produce results that cannot be compared across libraries. Here we present KGATE (Knowledge Graph Autoencoder Training Environment), a modular Python library built on PyTorch Geometric and TorchKGE. KGATE lets users assemble initializers, encoders, decoders, losses, negative samplers, and evaluation metrics as building blocks, or plug in their own block. KGATE includes a preprocessing procedure that controls data leakage, a builtin training pipeline, and reproducibility by design. Benchmarks against six existing KGE libraries show that KGATE training time is comparable with the fastest libraries while offering a broader set of features.
cs.LG / 117 / 2610.09929
Expected Sample Complexity in Multi-Armed Bandits
Abstract
Sample complexity is a widely used metric in sequential decision-making problems, defined as the number of suboptimal decisions during the interaction between the agent and an environment. We study the sample complexity of stochastic multi-armed bandit problems and introduce the expected sample complexity performance measure, analyzing it in a novel framework called approximately correct in expectation (ACE). We show that ACE guarantees imply almost sure convergence to the optimal expected reward, in contrast to high-probability guarantees found in other frameworks, and also show how to convert ACE guarantees into explicit expected regret bounds. We further show that, in contrast to existing measures, deterministic algorithms cannot obtain favorable ACE bounds, and analyze stochastic algorithms in two settings: when the allowed suboptimality level $ε$ is known to the algorithm and when it is unknown. In the former, we devise an explore-then-$ε$-greedy algorithm, and in the latter, we analyze the expected sample complexity of Thompson sampling. Finally, we establish nearly matching lower bounds for both settings, showing that the algorithms are tight in $ε$ and proving a performance separation between the two regimes.
cs.LG / 118 / 2610.09946
Learning Traffic Flow Dynamics with Stochastic Physics-Informed Neural Cellular Automata
Abstract
Traffic flow modeling is essential for understanding and predicting the collective dynamics of vehicles on road networks. Cellular automata provide a simple, interpretable yet powerful framework for representing these dynamics via local interaction rules, while retaining the ability to reproduce complex macroscopic traffic phenomena. However, learning local transition rules from data while preserving physically meaningful constraints remains challenging, particularly for stochastic models. In this work, we propose a physics-informed neural cellular automaton (PI-NCA) for data-driven traffic flow modeling. Building on the standard neural cellular automaton (NCA), we design a neural architecture that is physically consistent with the road topology and guarantees conservation of the total number of vehicles, thereby constraining the learned transition rules to physically admissible dynamics. We further extend this framework to stochastic dynamics by parameterizing probabilistic transition rules while preserving the same physics-informed constraints. We evaluate the proposed models on multiple traffic scenarios generated by the well-established Nagel-Schreckenberg and Kerner-Klenov-Wolf cellular automata. The results demonstrate that the PI-NCA successfully learns the dynamics of both traffic models and consistently outperforms a standard NCA, while the stochastic extension captures probabilistic transition rules without compromising the imposed physical constraints.
cs.LG / 119 / 2610.09958
Force without transmission: a depth-induced rank collapse that no loss on the representation reopens
Abstract
Training can drive a transformer into a rank collapse: all token representations point in one direction, and learning stops. In a related collapse of attention, a loss term with a bounded corrective force repairs the network during the run. We ask whether such a term repairs rank collapse. We collapse small transformers by weakening their skip connection and treat copies of the collapsed network. No added loss term repaired the collapse, although the stronger kind pushed with about a tenth of the task gradient. The reason was the path, not the strength. The task gradient no longer reached the query and key weights, which decide where attention looks, and the added term's gradient faded before the blocks where the collapse forms. Restoring the skip connection, which changes no weight, reopened this path at once. The rank then recovered, but only far above the scale of collapse. After a burst of high learning rate the path stayed open and the rank recovered untreated. Registered predictions from the path ranked recovery times but did not transfer to this cause. In every case the loss stayed above that of a healthy network after the rank recovered. Whether a collapsed network can be repaired depends on whether the gradient still reaches the weights that must change, not on how strongly a loss term pushes.
cs.LG / 120 / 2610.09969
TR-PTQ: High-Accuracy Integer-Only Transformer Post Training Quantization via Taylor Region Reformulation
Abstract
Post-training quantization (PTQ) enables efficient deployment, yet transformer architectures remain challenging to quantize due to nonlinear layers. While existing methods attribute accuracy loss to insufficient numerical precision, often necessitating floating-point fallbacks, we demonstrate that degradation is actually driven by specific structural error sources. We find that learned scale parameters in normalization layers and compounded approximations in GELU are the primary error contributors, whereas SoftMax remains inherently robust to aggressive quantization. To address these bottlenecks, we introduce TR-PTQ, a unified integer-only formulation using shared Taylor Region (TR) exponential and logarithm primitives. This approach allows computationally expensive operations, including division and square roots, to be performed entirely in the log-domain via standard integer arithmetic. Combined with a calibration-free, outlier-aware optimization for LayerNorm parameters, our method eliminates the need for floating-point hardware units for nonlinearities, achieving less than 1.5\% absolute accuracy degradation across vision and language benchmarks.
cs.LG / 121 / 2610.09994
Temporal Predictive Multiplicity: Equally Accurate Time Series Models Yield Different Forecast Trajectories
Abstract
Models with near-identical predictive performance can yield substantially different predictions, a phenomenon known as predictive multiplicity. Prior work has mostly studied this at the level of individual scalar outputs. In time-series forecasting, however, predictions across horizons jointly define a trajectory, and horizon-wise comparisons can hide important differences in predictive behavior. To address this problem, we introduce temporal predictive multiplicity, a framework that characterizes disagreement over complete forecast trajectories among models with near-identical predictive performance. We show that constraining predictive performance alone can still admit a broad range of different trajectories. We further show that constraining multiplicity at individual horizons partially reduces, but does not eliminate, trajectory-level multiplicity. Experiments with 19 neural forecasting architectures on 11 datasets confirm that near-optimal models can exhibit substantial variability in the forecast trajectories they produce, and trajectory-level disagreement is largely unrelated to horizon-wise disagreement. Our framework, therefore, exposes a gap in existing multiplicity studies: models with indistinguishable predictive performance imply fundamentally different temporal trajectories, with consequential downstream effects.
cs.LG / 122 / 2610.10000
Beyond Reward Suppression: Near-Optimal Offline Attacks on Warm-Start Bandits with Bounded Rewards
Abstract
Adversarial attacks on bandits aim to mislead a learner toward a target arm while keeping the attack cost small. Existing attacks typically achieve this by suppressing non-target arms. In practice, however, manipulation such as fake reviews often directly promotes the target item. We study this gap through bounded offline attacks on warm-start bandits, where an attacker can inject only valid action-reward pairs into the warm-start history before deployment. We show that target promotion is not merely a heuristic: when the target arm lies near the lower reward boundary, any order-optimal-cost attack against UCB that makes it selected in nearly all online rounds must allocate a nonvanishing fraction of its cost to the target arm. We then design an attack that achieves the optimal sublinear cost and characterize its allocation between target promotion and non-target suppression. We further extend the attack to Thompson Sampling, $ε$-greedy, and a broader class of bandit algorithms. Experiments on real-world and synthetic data validate the effectiveness of our attacks.
cs.LG / 123 / 2610.10014
A Drosophila Whole-Connectome Network Can Learn Human-Designed Cognitive Tasks
Abstract
Can a biological wiring diagram serve as a useful computational substrate beyond the behaviors for which it evolved? We use the publicly released MaleCNS v1.0 connectome, reconstructed from a single adult male Drosophila specimen, as the fixed recurrent topology of an artificial network. We train separate models for bounded addition and for a controlled grounded relational language task built from a fixed 100-word lexicon. In both models, one scalar is learned per anatomical edge. The anatomical graph reaches 92.77% mean accuracy on held-out addition, compared with 67.93% for directed degree-preserving rewires. On the strict paired language endpoint, which matches original and order-reversed scenes to their corresponding descriptions, it reaches 61.59% across four fixed interfaces, compared with 44.17% for matched rewires. At the canonical interface, it ranks first in a fixed 21-graph comparison. On the matched 48-group intervention subset, shuffling task-defined sensory features reduces its score from 60.94% to 19.27%. Together, these results show that higher-order MaleCNS wiring provides a reusable inductive bias for bounded addition and grounded relational language.
cs.LG / 124 / 2610.10018
Oscillatory Neural Dynamics over Sheaves
Abstract
Effective long-range propagation remains a central challenge in graph neural networks, as increasing a model's propagation depth does not guarantee that distant nodes effectively influence each other. Sheaf neural networks enrich graph propagation through matrix-valued transport between stalks; still, this expressivity alone does not automatically imply effective long-range communication. We introduce ONDA, a long-range graph learning framework based on operator-valued information waves. Stalk-valued representations evolve through second-order dynamics governed by learned sheaf transport operators, combining wave-like propagation with expressive local geometry. We characterize long-range influence through a stalk-wise sensitivity analysis and show that the cross-influence never vanishes. Across long-range propagation, severe graph bottlenecks, graph transfer, and heterophilic benchmarks, ONDA consistently improves over scalar wave propagation, diffusive sheaf baselines, and state-of-the-art models, demonstrating the benefit of coupling wave dynamics with matrix-valued transport.
cs.LG / 125 / 2610.10037
Matching of signal, noise and hardware timescales for filtering and forecasting of correlated noise signals
Abstract
Physical reservoir computing exploits the nonlinear dynamics of physical systems to process time-dependent data with greater energy efficiency than conventional machine learning approaches. However, physical reservoirs have fixed intrinsic response timescales, whereas real-world signals combine deterministic and stochastic components across multiple timescales. Here we show, using a nanoporous niobium oxide reservoir, synthetic noisy signals and cryptocurrency-price volatility, that the relationship among noise correlation time, reservoir memory and forecast horizon determines whether correlated noise is filtered or predicted. Noise varying faster than the relevant reservoir memory and forecast horizon is averaged by the reservoir, whereas the temporal structure of slower-varying noise is sufficient for algorithmic forecasting. We introduce the reservoir memory horizon and forecasting regime index to distinguish these operating regimes. These contributions demonstrate that timescale matching can guide the encoding of input time series and development of physical reservoir architectures that filter, analyse and predict stochastic signal components across distinct temporal scales.
cs.LG / 126 / 2610.10057
WxFM-XL: Adapting Univariate Foundation Models to Multi-Station Weather Forecasting
Abstract
With the rise of univariate time series foundation models (e.g., Sundial, Timer), initial efforts have been made to extend them to multivariate settings. However, these models mainly focus on modeling correlations among variables. When they are applied to multi-station weather forecasting, two important factors are often overlooked: (1) the spatial information of stations, and (2) different error priors of different stations relative to the foundation model. In this paper, we propose WxFM-XL, a model for adapting univariate time series foundation models to multi-station weather forecasting. WxFM-XL introduces a cross-station error correlation prior graph to capture stationwise error priors with respect to the foundation model. Building on this, we further propose a dynamic fusion mechanism that adaptively integrates a spatial correlation graph with the error correlation prior graph. Experiments on multiple datasets demonstrate that our model outperforms state of the art baselines.
cs.LG / 127 / 2610.10068
Efficient Provably Private Classification with a Tabular Foundation Model
Abstract
Tabular data underpin prediction and decision-making in medicine, finance, government and science, but often contain sensitive individual-level information, creating a need for accurate prediction while preserving privacy. Traditional private learning provides formal privacy guarantees, but requires slow dataset-specific optimisation, suffers substantial utility loss under strong privacy, and is often difficult to apply correctly. Tabular foundation models adapt rapidly to new datasets, but existing models lack formal privacy guarantees, and are highly vulnerable to membership-inference attacks, limiting their use on sensitive data. Here we introduce PrivTab, an easy to use tabular foundation model for differentially private classification that embeds a privacy mechanism within its architecture. Pretrained on simulated datasets, PrivTab uses in-context learning to transform sensitive rows into compact, provably private summaries---effectively learning how to learn under privacy. PrivTab outperforms private linear and neural-network baselines under moderate-to-strong privacy, shows negligible membership leakage, maintains well-calibrated predictions under strong privacy, and reduces dataset fitting time by 10,000 times, requiring only a single forward pass. By combining formal privacy, speed, and easy of use, PrivTab brings recent advances in AI to applications where sensitive individual-level data have limited their adoption.
cs.LG / 128 / 2610.10085
Activation-Aware Weight Tensorization: A Calibration-Time Preconditioner for Tensor-Network LLM Compression
Abstract
Post-training tensor-network compression replaces Transformer linear layers with Tensor Train (TT) or Tree Tensor Network (TTN) operators, but standard decompositions minimize weight-space Frobenius error rather than functional error under the layer's activation distribution. We propose Activation-aware Weight Tensorization (AWT), a training-free calibration wrapper that preconditions each weight matrix with a diagonal activation-derived scale before an unchanged TT/TTN solver and deploys the result with only an input-side elementwise rescaling. Across Llama 3.1 8B, Ministral 8B, and Qwen2.5 7B, AWT consistently improves vanilla TT/TTN tensorization at 2-6 times compression: under single-operator replacement, AWT closes 12-35% of the WikiText perplexity gap to the dense baseline across the three model families and 2-6 times compression settings; while under multi-operator Llama suffix replacement it closes 27-60% across attention-group and all-seven-matrix settings. The gains also transfer to downstream HellaSwag and ARC-Challenge evaluations. We further show that diagonal preconditioning is a robustness-modularity tradeoff rather than a diagonal-covariance assumption: a dense full-covariance oracle wins its own weighted objective in 80/81 cases, yet diagonal AWT gives better held-out functional fidelity in 53/81 cases. Together, these results position AWT as a principled, modular preconditioner for improving functional fidelity in fixed TT/TTN compression pipelines without modifying the decomposition solver.
cs.LG / 129 / 2610.10087
Multi-Agent Coordination via Support-Preserving Distillation
Abstract
Offline MARL increasingly relies on generative policies to model multimodal joint behavior, typically by distilling a centralized teacher into decentralized one-step actors under the CTDE. We identify a failure mode at the teacher training stage: standard flow-based teachers pair noise with replay targets independently, so nearby noise samples can be routed toward conflicting coordination modes. The teacher then produces samples between valid modes, and because the distillation loss regresses each local actor onto the conditional mean of the teacher's output given local input, this error is not absorbed but propagated to the student. To remove this teacher-side artifact, we propose Mode-Support Semi-Discrete Optimal Transport (MoSDOT), which summarizes multimodal replay into a finite mode support with prescribed capacities and uses conditional semi-discrete optimal transport to assign each noise sample to a single mode before teacher training. We additionally study a shared-randomness variant that uses a shared noise component at execution to expose the residual gap intrinsic to strict-product execution. On controlled diagnostics and offline MARL benchmarks, MoSDOT improves endpoint quality and routing consistency, particularly on datasets exhibiting multimodal joint behavior.
cs.LG / 130 / 2610.10105
CAFE+FNO: Fourier Kernel Generation via Multiplicative Feature Composition
Abstract
The Fourier Neural Operator (FNO) learns solution operators of partial differential equations (PDEs) through Fourier-space kernel parameterization, but frequency truncation can limit the learning of high-frequency variations. AM-FNO and SirenFNO generate kernels for all grid modes from spectral coordinates using shared networks, making coordinate encoding and generator design important. Recent work on implicit neural representations (INRs) has proposed constructing frequency interactions through explicit feature composition rather than relying on subsequent MLPs to form them implicitly. Building on this approach, we propose CAFE+FNO, which incorporates Content-Aware Frequency Encoding+ (CAFE+) into Fourier kernel generation. CAFE+ combines Fourier--Chebyshev features through parallel affine branches and a Hadamard product, forming interactions within and across the two feature families. A kernel MLP maps the resulting representation of each normalized spectral coordinate to a complex channel-mixing matrix. Each layer shares its generator across all stored modes, making the number of trainable parameters independent of the number of modes for a fixed architecture. We compare CAFE+FNO with existing FNO variants on five PDE benchmarks and conduct ablation studies on basis configuration, multiplicative composition, and bandwidth learnability. Code and experimental configurations are available at https://github.com/fabsk101/CAFEPlusFNO.git.
cs.LG / 131 / 2610.10118
YANchor-4B: Effective Long-Horizon Reasoning in O(N) Time with O(1) Memory
Abstract
Long-horizon reasoning demands access to earlier information at a manageable generation cost. Full-history attention incurs growing storage and computation, while recurrent compression can lose precise details. Therefore, we present YANchor-4B, a general-purpose recurrent model that preserves crucial memory as ANchors for retrieval during subsequent reasoning. Beyond $O(N)$-time generation and $O(1)$ memory, YANchor enables effective long-horizon reasoning through its multidimensional memory mechanism. For example, on challenging math problems, it achieves 82.93% mean pass@1 on AIME 2024--2026 and 63.64% on HMMT, substantially outperforming linear-time, constant-state counterparts, including larger models. It also delivers several-fold higher batched long-generation throughput than Transformer and hybrid baselines on H100. Furthermore, evaluations across dozens of benchmarks demonstrate YANchor's superiority in general-purpose capabilities.
cs.LG / 132 / 2610.10122
What Can a Gaussian Process Design Test
Abstract
A Gaussian process (GP) model can agree with the data for two reasons: its assumptions are right, or the chosen inputs could never have shown that they are wrong. The distinction can be checked from the design before any responses are observed. Every model implies relations that its noiseless responses must satisfy at the chosen inputs, such as the middle value lies on the line through its two neighbours. For GPs built from finitely many features, these relations are exactly the null space of the kernel matrix. Gale duality gives them a geometric interpretation, in which each observation has a vector and the smallest groups of observations that can expose an error are the circuits. For other kernels the relations become soft: response patterns may be improbable under the prior rather than algebraically impossible. A standard test then combines two kinds of evidence. Structural evidence comes from a violated relation and grows without limit as the noise falls. Prior-based evidence only says that a departure is improbable under the prior. With all inputs at the two ends of an interval, for example, a GP can reject a straight line against a large curvature, but only because the implied intercept is improbable, never because curvature was seen. In simulations the predicted power matched the observed rejection rates. Choosing the next input by predicted power raised the power against a localised discrepancy from 0.48 to 0.72, against 0.51 when choosing by predictive variance, and a grid in two dimensions contained exact tests of additivity that a Latin hypercube lacked. The test itself is classical. The contribution is the prospective reading of that test: before observing the responses, the design already determines what kind of contradiction it can produce.
cs.LG / 133 / 2610.10128
m-Set Adversarial Bandits with Winner Feedback
Abstract
We show upper and lower bounds on the regret of $m$-set adversarial bandits for different utilities (winner reward or sum of rewards) and feedback models (winner index, winner reward, sum of rewards, and their combinations). By comparing to standard bounds for combinatorial and MNL bandits, our results reveal how subtle changes in the setting can have a dramatic impact on the learning rates. Our main technical contributions are the information-theoretic lower bounds on the regret. Experiments on synthetic data confirm our theoretical analyses.
cs.LG / 134 / 2610.10174
A Unified Information-Theoretic Approach to Constrained Multi-Fidelity Multi-Objective Bayesian Optimization
Abstract
Bayesian optimization often involves multiple objectives, constraints, and fidelity levels. We address the challenge of jointly selecting where and at which fidelity to evaluate to identify the highest-fidelity feasible Pareto frontier in this combined setting. From a unified information-theoretic perspective, we measure query utility by the information gain about this frontier, provided by an observation. Since this mutual information is intractable, we derive a variational lower bound using a mixture of under- and over-truncated approximations to the Pareto-consistent region. Multi-fidelity surrogate models propagate the information to arbitrary fidelities, yielding a cost-aware acquisition function without separate heuristics for fidelity selection or constraint handling. Experiments on synthetic, benchmark, and real-world problems demonstrate effectiveness across diverse objective, constraint, and fidelity settings.
cs.LG / 135 / 2610.10186
Pre-training of Bayesian Optimization Algorithm through Bayesian Optimization
Abstract
Bayesian optimization (BO) is widely used as a standard approach for expensive black-box optimization. However, BO algorithms often involve parameters that must be specified in advance, and their performance can strongly depend on these choices. We propose a framework for optimizing such parameters using sample paths drawn from a Gaussian process (GP) inferred from the information available at the start of BO. We use cumulative regret as the performance metric for a BO algorithm. By running the BO algorithm on the generated sample paths, we obtain an empirical estimate of its expected cumulative regret for a given parameter configuration. Optimizing this estimate allows us to identify parameter configurations that, given the currently available information, are expected to achieve low cumulative regret. Since this parameter optimization is itself a black-box optimization problem, we employ another BO procedure to solve it, which we refer to as outer BO. Through experiments, we demonstrate that the proposed framework can effectively select parameter configurations that achieve strong performance among a range of candidate configurations.
cs.LG / 136 / 2610.10203
How to train your model organism
Abstract
Model organisms of alignment-relevant behaviors (e.g., backdoors, sycophancy, spurious correlations) have emerged as a key tool for evaluating whitebox interpretability techniques. We argue that the prevailing practice of training model organisms to a single objective of installing the target behavior is insufficient and propose validating model organisms with respect to three objectives with associated metrics: target-behavior installation, general-capability preservation (i.e., parametric knowledge, chat quality), and output naturalness (i.e., CoT and activations). We re-visit two publicly released organism suites using this validation framework and show that (1) chat quality and CoT naturalness degrade substantially across training recipes, and (2) validation metrics predict how well interpretability methods recover the installed behavior, e.g., a logit lens readout covaries with an organism's general capabilities. We introduce a multi-objective training approach based on model merging to train more realistic model organisms. Finally, on a new suite of model organisms targeting demographic biases in clinical reasoning, we compare training recipes and find that DPO training stays closer to the base model than supervised finetuning, and the proposed model optimization approach better preserves capabilities and naturalness. Auditing this suite with an investigator agent, we again observe validation metrics tracking bias recovery. In sum, training methods shape the interpretability conclusions an organism supports, and we argue that one should consider multiple objectives to draw generalizable conclusions about interpretability methods using (realistic) model organisms.
cs.LG / 137 / 2610.10213
Edge Accuracy Is Not Enough: Why Dynamics-Learned Structure Fails to Transfer to Inverse Problems
Abstract
A natural strategy for inverse problems with scarce labelled data is to transfer relational structure learned from abundant forward-simulation data. We show this strategy fails systematically, even when it satisfies the standard theoretical justification for why structure should help. We prove that approximate structure provides estimation-error benefits whenever the edge error satisfies $Δ< n^2 - kn$, reducing sample complexity from $O(n^2)$ to $O(kn+Δ)$. Structure learned via Neural Relational Inference (NRI) from dynamics prediction satisfies this condition, yet on a source-localisation task across 180 CFD-simulated hydrogen-leak scenarios and 180 acoustic scenarios, it degrades performance by 116% and 201% relative to a flexible, task-optimised attention baseline, while a physics-based prior (Green's function) degrades by only 69-72%. Four independent lines of evidence show this is not a tuning failure: NRI improves only 0.5% when given 18x more training data (versus 16.6% for the task-optimised baseline, $p<0.001$); performance is insensitive to the NRI edge threshold across a wide range; the dynamics-learned graph overlaps the task-optimal graph on only 6% of edges; and two further dynamics-derived structure estimators (correlation- and mutual-information-based) show no measurable benefit over a structure-free baseline, with the correlation-based estimator performing markedly worse. We formalise this gap as a statement about approximation error that the edge-accuracy condition cannot control, and we provide a lightweight transferability test (Jaccard similarity against a partially-observed target-task graph) that separates successful from failed transfer in all four domain/structure pairs we evaluate, using under an hour of computation and 15-20% of target-domain data; we present this as a heuristic calibrated on few cases, not a validated general threshold.
cs.LG / 138 / 2610.10218
Finite-Sample Approximation of Hessian-Guided Perturbed Wasserstein Gradient Flows
Abstract
Wasserstein gradient flow extends gradient descent to probability measures. Its Hessian-guided perturbed variant (PWGF) adds Gaussian perturbations to escape saddle points in nonconvex problems. We investigate when its approximation by finitely many interacting particles remains accurate over growing time horizons. Our analysis retains the curvature accumulated along the population-driven reference path: negative curvature can amplify approximation errors, while subsequent positive curvature can damp their influence. This captures favorable scenarios in which temporary instability is compatible with accurate tracking over growing horizons. Under regularity assumptions and a prescribed common perturbation schedule, we prove particle and objective-value tracking bounds on a high-probability event for reference paths satisfying explicit conditions on accumulated curvature. To handle state-dependent Gaussian jumps, we construct a population-first coupling that preserves the reference particles' conditional independence and reduces jump errors to covariance comparison. We verify the conditions in a variance-plus-cosine model, where curvature recovery yields a growing-horizon tracking guarantee. We also establish local attraction, transverse descent, and positive second variation in two regions of a regularized matrix-factorization model, motivating a positive-negative-positive curvature pattern.
cs.LG / 139 / 2610.10222
Evaluating Sequence Assembly Strategies for Differentially Private Synthetic Time-Series Forecasting
Abstract
Differentially private time-series generators commonly produce fixed-length synthetic windows, whereas downstream forecasting models often require long continuous training sequences. How these windows are assembled after generation can therefore alter the effective synthetic data presented to a forecaster, even when the trained generator remains unchanged. We study this post-generation sequence assembly process by systematically varying overlap rates and window-weighting schemes and evaluating the resulting sequences in terms of boundary continuity, statistical and temporal fidelity, and Train-on-Synthetic-Test-on-Real (TSTR) forecasting utility. Across four types of public datasets (ETTh1, ETTm1, Weather, and Appliances) and five forecasting models, the results reveal a clear forecaster-dependent assembly principle: downstream TSTR utility is jointly shaped by the forecaster, overlap rate, and window-weighting scheme, leading to distinct assembly preferences across forecasting models. Increased overlap generally improves boundary continuity, but improvements in continuity or individual fidelity diagnostics do not consistently reduce forecasting error, indicating that these diagnostics alone are insufficient for selecting assembly configurations. Complete five-forecaster assembly grids, together with matched Train-on-Real-Test-on-Real (TRTR) references, further characterize these regularities and quantify assembly-dependent utility relative to real-data training. We then validate the identified principles through additional analyses of robustness and generator variability.
cs.LG / 140 / 2610.10239
ProtocolMatch: Protocol-Dependent Model Selection for Scientific Dynamics Forecasting
Abstract
Scientific dynamics forecasting is often framed as an architecture choice, although deployment is also determined by observed history, rollout feedback, compute budget, physical objective, and test distribution. We formulate protocol-dependent model selection and introduce ProtocolMatch, a compute-matched, validation-selected, and failure-preserving evaluation framework. On driven quantum-spin dynamics, we compare recurrent, patched-attention, causal-attention, and low-rank linear predictors across three independently generated datasets. The causal-attention--recurrence ordering reverses as the training set grows within a fixed two-spin task, while a linear predictor has the lowest mean error in the six-spin local-observable comparison. Restricting observed history worsens every refreshed-history view but improves every closed-loop view in the four-spin study. A latest-state MLP has lower error than persistence on every dataset under state refresh across all five cells, yet its closed-loop rank varies by system and includes finite explosive errors. Physical penalties improve targeted consistency without reliably improving prediction error, and in-distribution intervals lose most coverage after a driving-frequency shift. Thus scientific model selection should return a predictor with its protocol and report accuracy, physical validity, and shifted-distribution reliability separately.
cs.LG / 141 / 2610.10245
MorphCL: Morphological Contrastive Learning for Inertial-based Human Activity Recognition
Abstract
Despite the ubiquity of sensors in wearable and mobile devices and the abundance of human movement data they generate, translating unlabeled recordings into foundational motion models remains an open challenge. Self-supervised learning (SSL) has alleviated the need for costly annotations, yet existing approaches leave the global structure of large-scale motion data largely untapped, relying on randomly sampled batches and local comparisons that become particularly problematic for in-the-wild inertial data dominated by stationary, low-variance behaviors. Here we introduce Morphological Contrastive Learning (MorphCL), a self-supervised pretraining framework that uses structure-aware grouping to inject explicit modeling of global structure into inertial-based SSL approaches. Building on two well-established pillars of motion analysis, the discovery of motion primitives, or motifs, and domain-specific feature descriptors, we show that MorphCL substantially improves linear probing and finetuning results of learned encoders by up to 15 percentage points in F1-score. In a comparison with existing foundation models, we demonstrate that MorphCL-pretrained encoders match or surpass them models in linear probing performance while trained on $4600\times$ less data. Qualitative analysis of the resulting embedding spaces further reveals morphologically meaningful cluster structure, with improved separation of kinematically similar activity classes.
cs.LG / 142 / 2610.10246
Stationary Bias and Extrapolation in Nonlinear Two-Timescale Stochastic Approximation
Abstract
Constant-step stochastic approximation generally has a nonzero stationary mean error that persists under time averaging. This paper studies that error for nonlinear two-timescale recursions driven by an exogenous finite-state Markov chain. Under stated smoothness assumptions and conditions on the stationary distribution, we derive a first-order bias expansion whose error bound remains uniform as the slow step size becomes much smaller than the fast step size. Fast-manifold coordinates keep the associated covariance equation regular in this limit. For fast step $η$ and slow step $\varepsilon$, the expansion reveals a mixed contribution $\varepsilon^2/η$ alongside terms linear in each step size. This dependence matters for bias reduction: along power-law step-size paths, the bias exponents need not be integers, so Richardson--Romberg extrapolation requires weights matched to the path. An exactly solvable nonlinear Markov example verifies the coefficients. We verify localization for temporal-difference learning and compare finite-run extrapolation at equal update budgets. For finite runs, we bound the initialization error of tail averages on both timescales under an additional coupling assumption. In the special case of additive independent noise, signed third-moment cancellation yields a sharper remainder.
cs.LG / 143 / 2610.10250
Logarithmic Regret via Passive Change Detection in Piecewise-Stationary Self-Tuning Regulation
Abstract
We study minimum-variance control of an unknown autoregressive system with exogenous inputs and coefficients that change at unknown times. Under bounded independent disturbances, fixed detection gaps, stability and feasibility conditions, and sufficient time between changes, we prove \(O((C+1)\log((T+1)/δ))\) regret with probability at least \(1-δ\), where \(T\) is the horizon and \(C\) the number of changes. Unlike switching bandits, where unselected arms can change unobserved, admissible plant changes provide information during exploitation: the correct feasible controller leaves only the disturbance in the output, whereas a detectable change raises output energy under the old controller. PIECE-CD explores initially and after alarms, then uses gated recursive least squares for control. Its energy test compares windowed output power with a threshold above the noise floor; the extension to unstable controller mismatches also monitors the reference controller's input proposal. We control false alarms across the horizon and prove logarithmic detection delay. Inputs are clipped to prescribed bounds. Logarithmic regret also holds under an explicit condition ensuring that clipping becomes inactive after a finite burn-in. Under the stated feasibility conditions, the extended detector covers destabilizing changes with detectable excess energy over a fixed window.
cs.LG / 144 / 2610.10260
PairAudit: Guiding Human Review with Graph Tokens under Distribution Shift
Abstract
Intrusion detectors can confidently misclassify attacks that were not seen during training. Human review can correct these errors, but only a limited number of cases can be checked. Uncertainty-based review may overlook confident errors, while anomaly scores alone do not show whether changing the review plan will correct more errors. We introduce PairAudit to find overlooked errors and improve review under a fixed budget. Its graph tokens capture prediction patterns across connected nodes. Rather than building another predictor through feature aggregation, PairAudit uses unusual relational patterns to uncover potential errors in existing predictions. Human feedback then helps decide whether these findings justify changing review priorities. Experiments across security tasks show that PairAudit corrects more errors on average than uncertainty-based review, including more errors on unseen attacks. These gains account for all review costs and do not require retraining the detector.
cs.LG / 145 / 2610.10273
A Closed-Loop Non-Asymptotic Convergence Analysis of PPO with Learned Critics and Clipping
Abstract
Despite its widespread use, Proximal Policy Optimization with clipping (PPO-Clip) remains difficult to tune, and the interactions among critic learning, clipping, and rollout reuse remain incompletely understood. We develop a \emph{non-asymptotic} analysis of PPO-Clip as a \emph{closed-loop actor--critic} system. It captures actor--critic coupling, nonsmooth probability-ratio clipping, finite-batch reuse, and predictable early stopping under explicit coverage and critic regularity assumptions, using raw GAE and Monte Carlo critic targets. Our synchronous and asynchronous guarantees jointly characterize policy stationarity and the tracking accuracy of the learned critic, with explicit dependence on algorithmic parameters. A sufficient coupling condition gives optimization, critic tracking, clipping, and finite-batch errors a common amplification bound. The asynchronous result also requires a delay-dependent critic stepsize restriction; violating these conditions does not establish divergence. For finite layered MDPs with tabular critics, a uniform bound on the actual clipped-gradient class replaces complete-trajectory counting. A verified growing-horizon family has polynomial sample complexity, and a two-time-scale schedule gives $O(T^{-2/5})$ stationarity and critic-tracking bounds with explicit fresh-rollout accounting. These results together advance our understanding about PPO and provide theoretical guidance in tuning.
cs.LG / 146 / 2610.10274
Sparse Planning in Visual World Models via Cost Gradients
Abstract
Token-based world models enable fine-grained latent planning, but repeatedly processing large spatial token grids makes action search expensive. We introduce COSTGRAD, a training-free, goal-conditioned selector that ranks spatial tokens by the gradient norm of the planning cost with respect to each input token. By deriving importance from the downstream control objective, COSTGRAD targets tokens that matter for planning rather than merely for prediction. On AdaLN-conditioned predictors at $50\%$ sparsity, COSTGRAD matches or exceeds full-token planning on three of four continuous-control benchmarks, while giving a measured $2.6\times$ wall-clock speedup per environment planning step. Combining token sparsity with reduced CEM search increases this to a $\sim 5\times$ total speedup while still exceeding the full-token baseline. We also identify an architecture-dependent failure mode: in a matched AdaLN-vs-concat comparison, concat maintains comparable full-token performance but pure COSTGRAD loses its advantage over random selection. This difference tracks action-pathway drift: gradient-selected removal produces less drift than random removal on AdaLN, but more on concat. These results highlight selector-architecture compatibility as a design axis for sparse world-model planning. Project page and demos: https://ycxuyingchen.github.io/costgrad/
cs.LG / 147 / 2610.10276
PatchBench: Measuring Collateral Damage in Activation Patching
Abstract
An LLM safety patch can pass a benchmark while still being a poor repair. This risk is especially acute for jailbreak repairs, where the goal is to correct a specific unsafe behaviour without changing unrelated behaviours. A patch may block exact evaluation prompts yet fail on close harmful variants, or suppress harmful behaviour by over-refusing benign prompts that share its wording or structure. Existing protocols primarily test whether models can be broken, while aggregate metrics (attack success, refusal rates, global capability) cannot distinguish selective repairs from broader local suppression. To address this gap, we introduce PatchBench, a benchmark of empirically observed model-specific jailbreak failures inducing actionable harmful answers. Starting from 27,870 prompts from 37 public datasets, we curate 15,314 English prompts and query 8 open-source instruction-tuned models. Combining WildGuard filtering, pairwise Elo ranking, and manual verification, we retain a curated bank of 400 high-confidence jailbreak failures. We further introduce PatchBench-Local, an evaluation protocol testing whether a patch is behaviourally precise. For each harmful source prompt, PatchBench-Local generates three families of local neighbours: harmful variants preserving malicious intent, benign prompts with matched structure, and benign prompts reusing key harmful terms. It evaluates harmful-neighbour correction and benign-neighbour preservation, distinguishing selective repair from broader local suppression. Evaluating four activation steering methods with PatchBench-Local and MMLU shows that global capability can remain nearly unchanged while local benign regressions are severe, confirming aggregate metrics miss important collateral damage. PatchBench-Local provides a more precise basis for developing and comparing jailbreak repair methods.
cs.LG / 148 / 2610.10298
Physics-Aligned Electronic Ground-State Learning Improves Generalization
Abstract
Machine-learned interatomic potentials (MLIPs) excel at in-distribution tasks, accelerating drug and material development, yet they struggle to generalize out-of-distribution. We propose to push the cost-accuracy Pareto frontier by designing observable-agnostic electronic ground-state descriptor models (GSMs) with computational costs situated between MLIPs and Kohn-Sham density functional theory (KS-DFT). We align the learning objectives and architectures of GSMs with the governing equations of KS-DFT by enforcing physical constraints and removing optimization pressure on unphysical or irrelevant degrees of freedom. In our size-extrapolation experiments from QM9 to QM40, our combined contributions OrthoNormal-Loss (ON-Loss) and Grassmann Restricted Occupied-Orbital Training (GROOT) reach a 79.1% energy and 83.4% force mean absolute error (MAE) reduction over previous state-of-the-art density GSMs. For Hamiltonian GSMs, ON-Loss and Residual Optimal-gauge Conditioning-aware KS-Eq. Training (ROCKET) together reduce the energy and force MAEs of the strongest baseline by 99.8% and 95.9%, respectively. Using a self-consistency rejection criterion, we filter out extrapolation errors on QMugs, rejecting fewer than 0.4% of predictions while reaching an energy MAE of 0.07 mHa. Finally, we demonstrate the efficiency of label-free self-consistency fine-tuning, and transfer GSMs to reactive chemistry in Transition1x, reaching energy errors below chemical accuracy.
cs.LG / 149 / 2610.10299
Shared Gaussianization: What Gaussian Regularizers Certify About Contrastive Learning, and What They Miss
Abstract
What can a distribution-matching regularizer such as SIGReg in LeJEPA certify about contrastive learning? We study shared Gaussianization (SG), a characteristic-function Gaussianity test on the average of two normalized views, scaled by an independent $χ_d$ radius. Because disagreeing views shorten the average, one test detects both misalignment and non-uniformity. SG vanishes exactly at the aligned, uniform minimizers of population InfoNCE, and under equal marginals it bounds the InfoNCE excess by $4\cdot 3^{3/4}β$ times the square root of the SG loss, plus a term linear in the loss. The square-root rate and this dimension-free constant are sharp, and no squared mean-embedding distance on view pairs achieves a faster rate. With an explicit alignment term, a rotation-invariant uniformity test gives a linear bound if and only if its spectrum dominates that of InfoNCE's kernel $e^{βu^\top v}$; SG's own test does, Gaussian kernels $e^{-γ\|u-v\|^2}$ qualify exactly when $γ\ge β/2$, and moment matching never does. Away from the optimum, the objectives differ. Along an isotropic nuisance channel, pure SG lowers its loss by adding per-view nuisance whenever the shared code is non-uniform. An alignment weight above the channel's gain makes the nuisance-free solution a strict local minimizer; for LeJEPA, the same rule gives a critical SIGReg weight that decreases with the batch size. At finite batch size, an off-diagonal U-statistic removes a plug-in bias toward misalignment. In controlled latent-variable models, pure SG retains per-view style, an alignment weight above the measured gain removes it, and for LeJEPA at three batch sizes the measured gain separates the encoders that retain style from those that do not. InfoNCE training also reaches a lower SG$_{0.2}$ loss than SG$_{0.2}$ training from scratch, which points to an optimization gap.
cs.LG / 150 / 2610.10301
Revisiting Explainable AI through Model-Independent Concept Dictionaries
Abstract
Modern applications of AI rely on increasingly complex models. Explainable AI (XAI) has emerged as a set of techniques aimed at improving model transparency. However, existing XAI methods typically assume input features to be inherently interpretable, or they rely on intermediate internal abstractions that are difficult to characterize and highly architecture-specific, hindering consistent use across models. To address these limitations, we propose DictXAI, a method that defines concepts directly in the input domain via a dictionary---a large, potentially overcomplete set of predefined elements, each carrying an interpretable meaning. Technically, DictXAI first computes a sparse code of the input and then attributes the model's prediction to the associated dictionary elements. We demonstrate the actionable nature of DictXAI explanations, showing that they can attribute AI malfunctions (e.g., Clever Hans effects) directly to identifiable artifact patterns in the data, while fostering human-AI alignment on intricate biomedical signals. We further demonstrate our method's ability to operate across a wide variety of dictionaries, including learned image bases, analytically defined waveforms for electrocardiography, and experimentally acquired dictionary elements. Overall, our results show that DictXAI provides more interpretable, actionable, and architecture-agnostic insights than classical XAI or existing concept-based approaches.
cs.LG / 151 / 2610.10302
Continual Graph Multi-Agent Reinforcement Learning
Abstract
In Continual Multi-Agent Reinforcement Learning (CMARL), agents learn cooperative policies across sequences of tasks, aiming to adapt effectively to new tasks while preserving the ability to solve previously encountered ones. In many applications, tasks differ in their underlying structure, which can represent, for example, distinct operational conditions or target configurations (e.g., different network topologies in power grids or arrangements in formation control). Existing CMARL methods lack dedicated mechanisms to leverage this structural information when learning new tasks, failing to promote transfer and mitigate forgetting. To fill this gap, we propose Continual Graph Multi-Agent Reinforcement Learning (CGMARL), a novel framework for CMARL problems in which task sequences are mapped into a series of attributed graphs, each modeling a task-specific structure. In CGMARL, each graph determines the environment dynamics (next states and/or rewards) and the number of agents for the corresponding task. Then, we present Graph-based Formation (GRAFO), the first CGMARL benchmark, and show how forgetting arises in this setting. Finally, to address this limitation, we propose Frozen Graph Encoder (FROG), a method that relies on a frozen graph backbone to preserve past structural information in graph-based CMARL policies. Experiments on GRAFO show that pairing FROG with existing CL methods substantially improves performance on multiple CGMARL scenarios.
cs.LG / 152 / 2610.10305
How Do Transformers Learn to Represent Symmetries?
Abstract
Training Transformer-based architectures with finite data augmentation has become an increasingly popular approach in geometric machine learning. Despite its empirical success, the interplay between the Transformer architecture, invariance to different symmetries, and augmentation budgets remains underexplored. In this paper, we study the ability of a vanilla Transformer to learn various symmetries through finite data augmentation for point cloud datasets. We identify an ordering of increasing learnability across the following symmetry groups: (i) non-angle-preserving symmetries, (ii) angle-preserving symmetries, and (iii) base angle-preserving subgroups, such as translation, rotation, and scale. For the base angle-preserving groups, we further investigate the Transformer's extrapolation behavior and conduct a structural analysis of the trained models, allowing us to identify interpretable mechanisms that induce invariance. Finally, we extend our analysis to equivariant functions and show that the detected mechanisms for approximate invariance can also provide a key building block for learned equivariance. Our project page is available at https://transformers-learn-symmetries.github.io/
cs.LG / 153 / 2610.10310
RSIGym: A Flexible Environment for Recursive Self-Improvement
Abstract
Recursive self-improvement requires carrying accepted changes into later improvement cycles, while studying agent-proposed changes also requires substantial research infrastructure. Existing settings often leave agents to rebuild routine infrastructure or restrict exploration to individual components. We introduce RSIGym, an agent-native research environment based on Everything as a Service (EaaS). RSIGym exposes training, inference, rollout, evaluation, and sandbox execution through reusable services, with shared budget and permission controls supporting Data, Harness, and Joint improvement tracks. This design enables agents to investigate individual interventions and jointly optimize data, training settings, and execution harnesses within the same environment. We define RSI-Index as the mean fraction of the remaining performance gap closed across five benchmarks covering software engineering, terminal interaction, mathematics, scientific reasoning, and skill-based tasks. Comparing six frontier research models in independent Joint runs, Opus 5 achieves the highest RSI-Index of 0.4809 under a $500 platform-service budget per benchmark run. Its selected systems improve all five benchmarks, raising SWE-bench Verified from 17.67% to 50.33% and AIME from 31.67% to 97.78%. Additional experiments examine DSH-harness refinement, budget variation, and restricted network access, while recorded trajectories reveal how agents diagnose failures and select candidates. We open-source the full RSIGym codebase and results to support reproducibility and further research.
cs.LG / 154 / 2610.10311
Fault-tolerant foundation models
Abstract
Emerging computer hardware often trades reliability for energy efficiency; here we show that large-language models (LLMs) can be trained to tolerate this unreliability, and that rather than degrading, their error resilience actually increases as they grow. Modified neural scaling laws inferred from 40,000 GPU-hours of training runs on simulated faulty digital hardware quantify this trend and suggest that models learn to compute within "good" error-correcting codes, whose relative overhead remains finite no matter how large the model gets. This finding leads us to conjecture that appropriately trained LLMs may be formally fault-tolerant; if true, running AI inference on low energy, faulty hardware may be a path to substantial energy savings over the status quo.
cs.LG / 155 / 2610.10314
PoreML: A Data-Driven Framework for Learning Multiphase Flow in Porous Media
Abstract
Multiphase flow in porous microstructures is central to CO$_2$ storage, fuel-cell operation, and flip-chip packaging. Predicting these flows remains challenging because wettability and complex pore geometry govern the nonlinear evolution of fluid interfaces. Machine learning holds substantial promise for advancing the field, but progress is constrained by scarce time-resolved 3D datasets and a lack of a unified workflow for training and evaluating models. To fill this critical gap, we introduce PoreML, an open-source framework unifying data generation, model training, and evaluation grounded in pore-scale physics. The framework comprises three core components. (a) A modern GPU-native lattice Boltzmann solver, validated against analytical solutions and published experiments, enables reproducible data generation. (b) A 3.3 TB dataset contains 560 simulation runs and 158,546 stored time steps across four application-driven scenarios. These trajectories span synthetic structures and geometries derived from micro-CT scans of real materials, covering diverse wetting conditions and viscosity ratios. (c) A unified learning framework evaluates one-step prediction and autoregressive rollouts. Its domain-specific evaluation protocols assess predictive accuracy and physical consistency. We evaluate five models of diverse architecture under these protocols. Two complementary challenges assess transfer to larger domains and from synthetic to micro-CT-derived structures. PoreML provides a shared foundation for machine-learning research on multiphase flow in porous media, with the aim of empowering the community to develop reliable predictive models and advance the field.
cs.LG / 156 / 2610.10317
Thinking in Depth: Retrospective Inference for Tabular Foundation Models
Abstract
Tabular foundation models (TFMs) are pretrained across diverse tabular tasks and make predictions on a new table at inference time using its labeled examples as context. Most recent TFMs perform such in-context prediction with stacked Transformer layers, repeatedly transforming how examples are represented and compared. By tracing individual queries through several strong TFMs, we find that predictive refinement is highly uneven across depth and is often concentrated in later layers. This uneven refinement motivates us to reconsider how intermediate representations are constructed and reused throughout the network. We introduce Retro, a tabular foundation model based on retrospective inference, where later stages can explicitly revisit and recombine intermediate information produced earlier in the network. Retro organizes this process around two complementary operations: which intermediate information to revisit, and how the resulting contextual update should be shaped for each query. Attention Residuals address the former by adaptively reweighting contributions from different depths, while query-conditioned Gated Attention addresses the latter by modulating the attention output element-wise across representation dimensions. Our analysis shows that Retro shifts predictive refinement earlier and more broadly across depth, with different stages revising different subsets of queries in a pattern suggestive of multi-view refinement. Across TabArena, TALENT, and RelArena, Retro ranks among the top three and lies on the Pareto frontier. These results indicate that directly reusing intermediate representations provides a practical way to better exploit depth in TFMs.
cs.LG / 157 / 2610.10323
HAN-Mamba: Hierarchical Selective State Space Networks for Multi-Scale Financial Volatility Forecasting
Abstract
Short-horizon realized volatility forecasting requires the integration of market information that evolves at incompatible temporal resolutions, from second-level order book dynamics to weekly regime drift. Our conference work introduced HAN-T, a hierarchical architecture in which scale-specific Transformer encoders process short, mid, and long-horizon streams and a learned attention fuser weighs their contributions. This article replaces the quadratic attention encoders with selective state space (Mamba) encoders while retaining attention only in the fuser, where the input is a three-token set rather than a long sequence. The resulting hybrid, HAN-Mamba, summarizes each stream through a recurrent state whose input-dependent gating matches two structural properties of volatility: persistent but decaying memory and abrupt regime shifts. On the Optiver Realized Volatility Prediction benchmark under time-aware five-fold cross-validation, HAN-Mamba improves mean RMSPE over HAN-T (0.1942 vs. 0.1965) with 33% fewer parameters. Its linear-time encoders further allow the high-frequency context to be extended from 60 to 240 buckets, reducing error to 0.1927 where the attention variant saturates, and support constant-time streaming updates at inference. Ablations attribute the gains to the encoder swap, confirm that the hierarchical prior transfers across sequence-model families, and show that the permutation-invariant attention fuser remains the correct mechanism for cross-scale integration.
cs.LG / 158 / 2610.10326
Average-Reward Reinforcement Learning for Multichain MDPs: A Hierarchical Decomposition Approach
Abstract
We study learning optimal policies in average-reward multichain Markov decision processes (MDPs), where the optimal gain may depend on the initial state and recurrence structures vary across policies, creating challenges for reinforcement learning (RL) methods. We propose an asynchronous value-iteration-based RL algorithm that requires no model knowledge beyond the MDP's transition graph and leverages Bather's decomposition to hierarchically partition the state space into communicating subsystems and transient states. This decomposition induces a recasting of the global decision problem into structured subproblems, which our algorithm exploits. We show that the algorithm converges to the optimal gain and produces gain-optimal policies after finite time. Building on this base algorithm, we develop two further algorithms: one approximately solves the multichain average optimality equations to obtain near gain-optimal policies, and another targets near bias-optimality by approximating the optimal bias function and solving an induced average-reward multichain MDP using the base algorithm. We provide almost-sure convergence guarantees for all three algorithms and empirically compare their tradeoffs, showing that the latter two also consistently improve transient performance relative to the base algorithm. To our knowledge, these are the first essentially model-free average-reward RL algorithms for general multichain MDPs without reductions to discounted problems.
cs.LG / 159 / 2610.10349
AutoAdapt: Automatic Domain Discovery Enables Low-Cost Extensibility
Abstract
Instruction-tuned models are deployed into environments where domains are heterogeneous and evolve, yet adding new domains or data typically requires costly retraining. We present AutoAdapt, a modular framework that incorporates new domains and data via targeted single-adapter training without modifying other adapters. The framework automatically discovers latent domains, uses them to train per-domain Low-Rank Adaptation (LoRA) adapters independently in parallel and performs parameter-free routing. Across 14 domain-specific benchmarks and GPT-4o pairwise judgements, AutoAdapt achieves parity with a LoRA adapter trained on all domains without requiring full-model retraining. We also find evidence of specialisation effect convergence across independent discovery methods. Overall, training each adapter on its own domain prevents domain interference by construction, thus enabling modular, taxonomy-free domain specialisation without aggregate performance loss or full model retraining.
cs.LG / 160 / 2610.10361
ORDERS: An Empirical Study of Norm-Rank Aggregation for Personalized Federated Learning
Abstract
Personalized federated learning combines shared representations with client-specific predictors, but the contribution of a server weighting rule can be obscured by local training and evaluation choices. We study ORDERS, a configuration that combines a shared backbone, a private residual adapter and classifier, geometric weights assigned by descending update norm, feature alignment, and private-parameter perturbations. The server computes a weighted sum of updates obtained from the same broadcast model; it does not obtain an additional optimization effect from sequential addition. A fully specified evaluation comprises 80 final runs: eight configurations, two datasets, and five training seeds on one fixed partition per dataset. On two-class-per-client CIFAR-10, ORDERS achieves $80.51 \pm 0.79\%$ native mean client accuracy, compared with $79.02 \pm 1.42\%$ for FedPer-R1 and $80.27 \pm 0.73\%$ for the matched uniform-weight control. After common local fine-tuning, the difference from FedPer-R1 narrows to 0.32 percentage points. On Sent140, ORDERS reaches $74.71 \pm 0.49\%$, only 0.69 points above a post hoc client training-majority diagnostic. Ablations provide limited, endpoint-dependent evidence for norm ranking and alignment, and no clear benefit from perturbations. Parameter-payload savings are 5.47% and 0.78%, respectively.
cs.LG / 161 / 2610.10367
Temporally Interpretable Differentiable Decision Trees
Abstract
Interpretability offers a solution to safe autonomy by providing transparency into an agent's underlying decision-making model. Within sequential-decision making tasks, differentiable decision trees (DDTs) are one approach to such interpretability, maintaining automatic-differentiable policies while providing humans with a discrete tree-based visualization. Nonetheless, current implementations of DDTs are not well-suited for sequential-decision making domains, as there exists an inherent mismatch between a tree's single-timestep behavior and a human's multi-timestep planning. Our work thus introduces time as a new dimension of interpretability, coined as temporal interpretability, and demonstrates how temporal abstractions via action chunking improve it. We achieve this by first introducing two novel policy gradient algorithms that incorporate action chunking. Additionally, to maintain parameter-efficient trees, we develop an information-theoretic tree restructuring algorithm that modifies the tree during training. Across four simulation environments, we find that warm-starting action chunked DDTs from a distilled action chunked policy is the most effective way to obtain temporally interpretable trees: they match neural network policies in three of the four domains while using up to 80$\%$ fewer parameters. Our code is available at https://github.com/ei5uke/temp-interp.
cs.LG / 162 / 2610.10368
Input-Blind Controls Produce Substantial Oracle Headroom for Layer Programs in Multiple-Choice Evaluation
Abstract
Adaptive computation aims to improve language-model inference by tailoring execution to each input. For layer programs, oracle evaluations use known answers to estimate the potential gain from this flexibility, before a practical selector is available. However, a gain from selection does not by itself explain why the chosen programs help. This study examines this distinction using 32 layer-skipping and repetition programs on two models and 4,413 multiple-choice items. The analysis compares their gains over a fixed action selected without the evaluation prompt with those of input-blind perturbations at the same sites, re-evaluating selections on another prompt. With shared option order, the controls give 10.2-11.8 and 15.6-19.4 percentage points of headroom on Qwen3-4B-Base and Llama-3.1-8B, exceeding the real programs' 9.0 and 10.1 in all three random-direction draws per model. They match answer-change rate only, and the ordering depends on the menu: in post hoc comparisons, real programs lead on Llama's repeat-only menu in every draw. A smaller KL-calibrated comparison, including an input-dependent control, favours real programs in point estimate, with inconclusive corrected tests. Fixed letter offsets produce headroom of similar scale. Rotating options sharply reduces both families' headroom, while leaving positive real-minus-control differences of 1.4-2.3 and 3.7-4.5 points; their magnitudes and statistical support depend on further adjustments and the reference. A supplementary generated-answer test finds that search-selected programs keep a 26.0-point advantage over programs selected for other problems after rewording, without a placebo comparison. These results show that substantial headroom can persist across prompts with shared option order without establishing a benefit specific to the selected layer computation; neither ordering against these controls identifies that benefit.
cs.LG / 163 / 2610.10379
Continual Learning without Continual Training
Abstract
Continual learning requires models to adapt to new domains and new classes while retaining prior knowledge. Many existing methods rely on continued optimization, using regularization, replay, or parameter expansion to prevent new updates from overwriting previously learned knowledge. Instead, we propose replacing continual training with continual inference: a PFN-based model that is meta-trained, and then frozen, adapting to new classes only by extending an in-context evidence set. Our model, Latent Concept PFN, performs in-context Bayesian inference over a latent concept space that captures semantic structure shared across domains and classes. As each new domain or class arrives, exemplars are added to the memory; adaptation reflects updated posterior beliefs over latent concepts rather than gradient updates. No parameters are changed, reducing forgetting. The same method handles both domain and class incremental continual learning without task identity. Concept annotations are only used during meta-training, acting as a soft anchor on the latent space rather than a fixed bottleneck. Unlike fixed-vocabulary concept methods, the model also handles noisy, ambiguous, or incomplete annotations by combining concept labels with raw input evidence to discover distinctions beyond the predefined concept set. Experiments on class and domain incremental learning datasets demonstrate competitive continual learning performance while learning interpretable latent concepts.
cs.LG / 164 / 2610.10381
ResidualQuant: KV Cache Quantization for Looped Transformers with 2-Bit Residuals
Abstract
Looped Transformers improve parameter efficiency by repeatedly applying shared Transformer blocks over multiple recurrent loops, increasing computational depth without increasing the parameter count. However, KV cache memory still scales with the number of loops, becoming a key memory bottleneck that limits batch size and inference throughput. KV cache quantization can alleviate this bottleneck, but existing methods often suffer substantial accuracy degradation at aggressive low-precision regimes. We observe that looped Transformers offer a unique opportunity: KV states across loops are highly similar. Based on this observation, we propose ResidualQuant, which uses the final-loop KV states as a reference and represents the remaining loops with low-precision residuals. Our method further combines least-square scaling and rotations applied to the residuals, as well as loop-wise mixed precision, to enable accurate quantization down to INT2 while retaining efficient reconstruction. Across multiple looped Transformer models and mathematical reasoning and code generation benchmarks, ResidualQuant consistently improves the accuracy-memory tradeoff over state-of-the-art rotation-based KV quantization. In particular, our method retains accuracy close to BF16 under mixed-precision settings while reducing theoretical KV storage by 80.7%, achieving up to 13.0% higher accuracy than the rotation-based baseline at the same memory budget. On an RTX 5090, the reduced KV memory traffic improves fixed-batch decode throughput by up to 2.73x, while the smaller memory footprint enables up to 2x larger batches, improving peak throughput by up to 4.15x.
cs.LG / 165 / 2610.10383
Boosting and the Expressive Power of Simple Weak Learners via the $γ$-VC Dimension
Abstract
Boosting converts weak hypotheses with a small edge over random guessing into highly accurate predictors, but the expressive power of the resulting classifier can depend strongly on the structure of the base class. We study this phenomenon through the $γ$-VC dimension introduced by Alon et al. (STOC 2021). Our first result shows that this parameter characterizes the sample complexity for weak-to-strong learning up to a constant factor scaling in $γ$. We then sharpen the general relationship between the classic VC dimension and the $γ$-VC dimension. Finally, we also give improved upper and lower bounds on the $γ$-VC dimension for the fundamental concept classes of decision stumps and axis-parallel rectangles in $\mathbb{R}^d$.
cs.LG / 166 / 2610.10385
OrBIT: Structure-Guided Embedding Compression
Abstract
Embedding tables are among the largest components of modern language models. Most compression methods fix a coding geometry such as coordinate blocks, low-rank subspaces, or unrestricted codebooks, and optimize within it. We instead ask whether the coding geometry can itself be discovered. We introduce \emph{OrBIT}, a structure-guided embedding compression framework that learns reusable local geometry from orbit dynamics and uses it to constrain a small set of shared codewords. The global reconstruction residual then decides where the fixed coding budget is spent, while redundant overlapping charts let local errors compensate one another after gluing. Our theory shows how tight-chart geometry controls distortion, how the global residual directs sequential allocation, and how data-geometry-guided refinement improves the codec. The resulting orbit machinery is compiled away, leaving a compact decoder in which the learned structure governs what is stored, where capacity is allocated, and how local information is assembled globally. Across four LLM embedding tables, OrBIT achieves $37.9\times$ compression on GPT-2 and over $23\times$ on each 7B table relative to 16-bit storage, while delivering competitive rate-distortion performance against established quantization and low-rank baselines.
cs.LG / 167 / 2610.10394
Kernel Autoresearch for Open-Ended Model Discovery
Abstract
Kernels encode the inductive bias of a wide range of machine learning models, yet automated kernel design faces a fundamental dilemma. A fixed grammar of base kernels and operators guarantees validity but limits the search to structures expressible by those building blocks. Conversely, unrestricted programs remove this limitation but no longer guarantee validity. In our stress tests, 22-58% of LLM-generated kernels that pass numerical checks on random inputs fail when evaluated at different scales or dimensions. We propose Kernel Autoresearch (Kernaut), which treats kernel design as open-ended model discovery. Coding agents write kernels as programs, while construction contracts ensure that every accepted kernel is valid. A quality-diversity archive retains high-performing kernels with distinct behaviors, and novelty screening steers agents toward functionally new candidates. Our experiments demonstrate that the discovered kernels encode reusable inductive biases that generalize to unseen tasks. On held-out black-box optimization families, a discovered kernel outperforms a meta-learned deep kernel trained on the same episodes. Furthermore, kernels discovered from ten enzyme-kinetic rate laws achieve lower error than tuned ARD and deep kernel baselines on five unseen mechanisms. The discovered kernels are also interpretable programs that human researchers can refine: a human-refined version of one further reduces the held-out predictive error by 5.7% and optimization regret by 7.8%.
cs.LG / 168 / 2610.10395
Executing Causal Structure Learning with Linear-Attention Transformers
Abstract
Transformers can execute algorithms on data given in their input. We ask whether they can do the same for causal discovery. We study a standard continuous method that repeatedly updates a candidate causal graph while enforcing acyclicity. We explicitly construct a fixed-weight transformer whose forward pass exactly reproduces one update of this method, so repeated blocks reproduce its optimization trajectory. The transformer carries the current graph and the algorithm's multiplier between updates. We show that retaining the multiplier is essential for exact execution, since different multiplier values can lead to different next updates. We also give conditions under which, within a fixed stage, the number of updates needed to reach a target accuracy can be computed in advance and rounding errors stay bounded as depth grows. Experiments show that the constructed block agrees with a reference update to floating-point precision, while arithmetic replay on synthetic data and seven published benchmark network topologies inherits the reference solver's successes and failures. This separates accurate algorithm execution from accurate causal recovery. In contrast, the ordinary attention models tested under our training budgets do not reliably execute the update or transfer to larger graphs. Whether gradient training can learn an executor in the architecture class of the construction remains open.
cs.LG / 169 / 2610.10398
Cross-Domain Pretraining for Steady-State Neural CFD Surrogates
Abstract
Neural surrogates for computational fluid dynamics (CFD) have the potential to greatly enhance engineering innovation through accelerating simulation. However, the primary limitation for neural surrogates is the lack of generalization to geometries and applications beyond the training set, which is significant given the diversity of engineering scenarios. Currently, this is addressed by generating a new dataset for a specific application; however, this requires running costly numerical solvers. In this work, we take a step toward addressing this by studying neural surrogates trained across different geometries, boundary conditions, and fidelities. We find that cross-domain pretraining improves zero- and few-shot performance on held-out datasets relative to both training from scratch and transferring from domain-specific experts. In particular, finetuning a pretrained, cross-domain model can achieve 2-3x lower errors at the same sample size and use 8x fewer samples to achieve the same error, compared to training from scratch. This benefit is architecture agnostic and improves with model size and pretraining dataset diversity. Furthermore, we study how and why cross-domain pretraining works in CFD surrogates, and find that simply pooling steady-state datasets is both sufficient and effective. Given the high cost of generating CFD data, leveraging existing datasets through cross-domain pretraining will likely be a valuable strategy as future surrogates expand to tackle new problems and use cases.
cs.LG / 170 / 2610.10422
Which Rollout Taught It That? BehaviorTrace and the Limits of Training-Data Attribution in Online RL
Abstract
When reinforcement learning teaches a language model a new behavior, can we find the training rollouts that taught it? And when an attribution method says it can, how do we know the answer is real? We study both questions on online RL fine-tuning with GRPO, using a planted behavior with a known cause. We release BehaviorTrace, an open evaluation harness that combines full-gradient sketching, the planted-behavior setup, and controls for gradient magnitude, fluency, headroom, and variation across seeds and generation draws. Across three seeds on Qwen2.5-1.5B, much of the apparent attribution signal comes from confounds. A control that ranks training steps by gradient size alone, with no behavior target, reaches 4.2 to 4.5 times chance and matches or beats the best targeted estimator on two of three seeds. At saturated checkpoints, model fluency predicts the behavior label at least as well as every gradient method we compared it with. Once fluency is controlled, the per-rollout results change from seed to seed and from one generation draw to the next, so a single run cannot settle the question. One signal does hold on all three seeds. The gradient of the trigger tokens aligns with a target built where the behavior actually occurs. We turn these findings into a checklist for evaluating attribution in RL. We test existing estimators, including GAS (renormalized TracInCP) and a TRAK-style estimator, and do not propose a new one.
cs.LG / 171 / 2610.10437
Q-Learning with Scalar Adjoint Matching
Abstract
Flow policies capture rich and diverse action distributions, and fine-tuning them with off-policy RL to improve beyond the demonstrations has drawn growing interest. However, fine-tuning a flow policy against a learned value function is not trivial, because the policy generates its action over many flow steps. Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size. We observe that the batch-averaged velocity Jacobian of pretrained flow policies concentrates on its diagonal. Motivated by this finding, we derive a closed-form scalar adjoint that scales the value gradient at the final action by the flow time, eliminating the per-step vector--Jacobian products. We further find that controlling the critic's value at policy-generated actions is particularly important under the scalar adjoint. Based on these findings, we propose Q-learning with Scalar Adjoint Matching (SQAM), which combines the scalar adjoint with a value penalty at those actions. SQAM's gains concentrate on the four hardest OGBench domains, where its success rate exceeds that of the strongest baseline in each domain by 18 to 35 percentage points. To test whether SQAM extends to large pretrained policies, we also fine-tune a vision-language-action policy on a real bimanual robot. SQAM improves over supervised fine-tuning on all three tasks.
cs.LG / 172 / 2610.10447
A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and Teaching
Abstract
Reinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate. On-Policy Distillation (OPD) offers an attractive alternative by providing dense token-level supervision from a stronger teacher along the student's own generations. Self-distillation methods further remove the need for a separate teacher model by conditioning the same policy on privileged information to serve as its own teacher. However, privileged conditioning alone does not guarantee that the resulting distillation update improves the student. Indeed, privileged information can lead the teacher to solve tasks through shortcuts unavailable to the student, producing supervision poorly matched to the student's current behavior. Consequently, even a higher-performing teacher can provide guidance that degrades student performance. To address this, we analyze how the choice of privileged teacher affects the student's update. We derive a necessary and sufficient condition for the teacher's local distillation update to be a positive multiple of the student's reward gradient. Our analysis suggests that the teacher should not only perform well on the task, but also provide guidance suited to the student's current capabilities. This characterization motivates a practical teacher-training surrogate that combines outcome rewards with token-level Kullback-Leibler (KL) regularization toward the student. Based on this result, we propose Joint On-Policy Learning and Teaching (JOLT), which jointly trains a single policy in two roles: a privileged teacher using a KL-regularized objective, and an unprivileged student using dense on-policy distillation. Across mathematical reasoning, coding, tool use, and terminal use, JOLT improves training efficiency and performance, with further gains from student rewards.
cs.LG / 173 / 2610.10459
NeuralBES: A Differentiable, Control-Aware Emulator for Scalable Building Energy Modeling
Abstract
Demand-side flexibility i.e. forecasting, shifting, and curtailing residential energy loads, depends on thermal models trusted across millions of heterogeneous buildings. Existing tools force a hard tradeoff: high-fidelity physics simulators such as EnergyPlus are accurate but sequential and require per-building calibration, while purely data-driven sequence models scale but abandon the physical structure that makes their predictions trustworthy. We introduce NeuralBES (Building Energy Simulation), a differentiable emulator that resolves this tradeoff by parameterizing a resistance--capacitance (RC) based thermal model with a shared neural encoder: static building metadata such as floor area, vintage, and HVAC type is mapped to physically bounded capacitances, conductances, and equipment coefficients, which become the coefficients of a scalar linear recurrence solved via a log-space parallel scan, and a predictor--corrector loop closes the thermostat--temperature nonlinearity while preserving full-horizon gradient flow. Trained on the ResStock dataset across three climate zones, NeuralBES handles heterogeneous building archetypes, vintages, and climate zones within a single trained encoder, while black-box baselines produce statistically plausible but physically inconsistent trajectories. On the annual full-year rollout, NeuralBES is the only data-conditioned model that is simultaneously physics-valid and accurate to within 4 MAPE points of the strongest raw-error baseline, while operating at roughly an order of magnitude fewer parameters than the transformer and recurrent baselines; among physics-valid baselines at parameter parity it more than halves the MAPE of the grey-box RC alternative.
cs.LG / 174 / 2610.10460
Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts
Abstract
Multi-teacher on-policy distillation (MOPD) is used in two settings. In common-domain composition, several teachers score each student rollout from one prompt domain and their signals form a single target; in routed-domain distillation, prompts from different domains are assigned to the corresponding specialist. Both settings usually transfer each teacher's endpoint policy, which mixes what post-training changed with preferences inherited from the teacher's base. We introduce $Δ$-MOPD, which transfers each teacher's teacher-minus-base logit shift re-anchored at the student's frozen initialization, and compare it with endpoint supervision in both settings while holding teacher selection fixed. We first expose the mechanism that impedes endpoint transfer: inherited base pull can exceed the post-training shift. Removing it reduces the teacher-term norm ratio and target--student KL. Across our experiments, the results suggest that shift targets are particularly useful when teacher signals are combined at a state. With three composed teachers, $Δ$-MOPD exceeds endpoint composition by $4.11$ Math and $1.95$ five-benchmark points; with two, it matches endpoint accuracy. Under phased routing, it achieves higher mean performance in both phase orders and reduces the observed order gap from $10.50$ to $6.42$ points. Under interleaved routing, where each update involves one teacher, the two targets perform comparably. The phased results provide supporting evidence that the benefit may extend to signals accumulated across training phases. Target construction is thus an independent design axis in MOPD, complementary to teacher selection.
cs.LG / 175 / 2610.10483
Two-Level Softmax Sampling Done Right: Correcting Bias from Size Imbalance and Dispersion
Abstract
Sampling from a softmax distribution is a fundamental operation in machine learning, but its linear complexity in the number of items makes exact sampling impractical at scale. Two-level softmax (2LS) sampling is a popular alternative enabling sublinear-time sampling. Assuming items are partitioned into clusters, 2LS first samples a cluster and then an item within it. In this paper, we show that, despite its advantages, 2LS introduces systematic and undesirable sampling biases, which arise from misweighting clusters by ignoring both cluster size imbalance and intra-cluster similarity dispersion. We propose two sampling methods, Size-Corrected 2LS (S-2LS) and Size- and Dispersion-Corrected 2LS (SD-2LS), which correct these biases and provide provably better softmax approximations with negligible to non-existent computational overhead. In-depth experiments on five large-scale datasets validate the improved sampling properties of our methods. We recommend their consistent use in place of standard 2LS in future work.
cs.LG / 176 / 2610.10499
Oracle-Efficient and Parameter-Free Agnostic Smoothed Online Learning
Abstract
Online learning is an attractive framework in many domains because it permits well-defined learning even when data are dependent or chosen adversarially. This generality, however, comes at a steep price, introducing significant statistical and computational barriers. Recently, smoothed online learning has emerged as a promising framework that interpolates between the fully adversarial and fully stochastic settings by assuming that the conditional law of each covariate has density at most $1/σ$ with respect to some fixed base measure $μ$, and it is known to match the statistical and computational guarantees of classical learning while still allowing for much of the flexibility of online learning. However, existing oracle-efficient algorithms require either (i) sampling access to the base measure $μ$ or (ii) labels that are perfectly predicted by a fixed hypothesis. Both assumptions limit the applicability of these algorithms, in contrast to statistical learning, where empirical risk minimization (ERM) learns efficiently in the agnostic setting without any knowledge of the data distribution. We show that neither assumption is necessary, giving the first oracle-efficient algorithm that achieves sublinear regret in the agnostic setting without knowledge of $μ$. Our algorithm, based on Gaussian Follow-The-Perturbed-Leader, is parameter-free: it requires no knowledge of $μ$, the smoothing parameter $σ$, or the horizon $T$, and it achieves regret $\widetilde O(d\sqrt{T/σ})$ for binary classes of VC dimension $d$ with a single call to an ERM oracle per round, which is optimal up to a $\sqrt{d}$ factor. En route to establishing the regret bound, we introduce several new techniques that may be of independent interest.
cs.LG / 177 / 2610.10519
Why Forget-Only Unlearning Needs Memorization
Abstract
Machine unlearning asks for a deletion algorithm whose output is close to retraining from scratch without the selected forget examples. In this work, we study forget-only unlearning, where the deletion algorithm receives only the trained model and the examples to forget, with no retained data or extra training information. We ask whether forget-only unlearning is always possible. We first show that this depends on the learning method: different datasets can produce the same trained model but require very different outputs after the same examples are removed. Using this observation, we derive lower bounds on how accurately unlearning can match retraining and instantiate them for several standard learning algorithms. We then ask what must be true when forget-only unlearning succeeds. To this end, we derive lower bounds on what an algorithm must memorize about the training data to handle arbitrary deletion requests. For simple threshold learners, the required information can be as large as the entire dataset, even though ordinary training keeps only one boundary point. Overall, our results show that information discarded during ordinary learning may be needed later for deletion, so models designed for forget-only unlearning may need to retain more information than standard training does.
cs.LG / 178 / 2610.10520
Distilling Graph Geometry: Knowledge Gap from GNNs to MLPs
Abstract
GNN-to-MLP distillation aims to retain the predictive accuracy of a message-passing teacher while deploying a graph-free MLP at inference. Existing methods mainly transfer node-wise predictions or use confidence-based reweighting, but they do not specify where the student should preserve the teacher's graph-induced geometry. We show that this omission leads to two spectral failure modes in the student's representation space. On sparse graphs, the student suffers from spectral underfit, missing high-energy teacher directions concentrated near boundary regions. On dense graphs, it suffers from spectral overfit, retaining spurious directions that the teacher has collapsed through aggregation. Motivated by an energy-weighted teacher-student alignment objective, we propose Graph Geometry-aware MLP (G^2MLP), a training-time distillation framework guided by Ollivier-Ricci curvature. Curvature identifies where the two spectral errors concentrate and is used to allocate supervision between prediction-level and representation-level alignment. The deployed model remains a standard MLP and requires no graph access at inference. Across node-classification benchmarks, G^2MLP consistently improves over graph-free distillation baselines, reduces the teacher-student rank gap in both regimes, and transfers without architectural changes to Graph Transformer teachers and link prediction.
cs.LG / 179 / 2610.10536
Decoupling Exploration from Optimization in RLVR
Abstract
Modern language models undergo reinforcement learning with verifiable rewards (RLVR) on top of already-trained checkpoints. A key promise of RLVR is the discovery of new reasoning strategies. In principle, a model can sample novel ideas absent from its prior training data. In practice, however, augmenting RLVR with strong novelty incentives has seen limited success and can degrade model quality. Because verifiable rewards supervise only a narrow slice of the model's knowledge and behavior, such degradations are difficult to recover from. Instead, we decouple exploration from optimization in a framework we call Exploration-Distillation (ExpDis). We train one or more explorer policies with a novelty bonus in the reward, filter their trajectories for correctness and quality, and distill them into a separate student policy. The student policy is then trained without a novelty bonus. We repeat the above procedure for several rounds, alternating between exploration and optimization. This decoupling allows us to aggressively scale exploration without degrading the student policy. Across seven mathematical reasoning benchmarks and two model families, ExpDis outperforms DAPO at the same wall-clock budget. Moreover, we observe improved pass@$k$ scaling, indicating that ExpDis produces models that generate more diverse correct solutions.
cs.LG / 180 / 2610.09226
FLoRa: Flight-Assisted Data Collection from Duty Cycling LoRa Nodes under Energy Constraints
Abstract
Data collection using Unmanned Aerial Vehicles (UAVs) is challenging when LoRa IoT Devices (IoTDs) duty-cycle to conserve battery. Under energy constraints, a UAV must decide which IoTDs to visit, in what order, where to hover, and how many times to probe each node, while time-based data freshness decays. Tractably solving this problem requires a multi-level optimization architecture: discrete combinatorial optimization for routing, continuous global optimization for spatial positioning, and sequential decision-making under uncertainty. We propose FLoRa, a Flight-assisted LoRa data collection architecture using Simulated Annealing (SA) for path planning, Covariance Matrix Adaptation Evolution Strategy (CMA-ES) for hover positioning, and Partially Observable Markov Decision Processes (POMDPs) for probing IoTDs. To quantify collection utility from duty-cycling nodes, we introduce the Value of Information for Pull-based systems (VIP), a metric that rewards fresh data and penalizes failed probes, imposing well-posedness and preventing indefinite probing when an IoTD is off. Tracking hard battery constraints on every POMDP sample path requires state augmentation, worsening the curse of dimensionality. For tractability, SA and CMA-ES work on the hard battery constraints, while at the POMDP layer we relax them into soft average constraints via Lagrangian relaxation. Since solving the network-wide POMDP is computationally complex, we decompose it into node-level POMDPs by approximating inter-node time dependency using a forward-decomposition technique. Evaluation shows FLoRa outperforms metaheuristic, greedy, and deep reinforcement learning baselines by 30.6%, 27.8%, and 15.2% in total expected VIP, while increasing node coverage by 24.5%, 29.2%, and 8.8%, and successful collections by 24.3%, 25.6%, and 15.2%, respectively.
cs.LG / 181 / 2610.08970
HULK: Learning Whole-Body Forceful Loco-Manipulation for Humanoids
Abstract
Humanoid loco-manipulation of large, heavy objects demands forceful interaction across the entire body. However, such payloads shift a humanoid's center of mass and impose sustained loads across the upper body, challenging balance and command tracking. We present HULK, a whole-body control framework for forceful loco-manipulation. Using model predictive control (MPC) to guide reinforcement learning with predictions of the loaded dynamics, we train two teachers: one tracks arm motions under wrist forces, and the other locomotes while holding large objects against the body. A capture-point control barrier function augments the wrist-force teacher during training to improve balance under load. We distill both teachers into a single policy. Evaluation spans simulation and the Unitree G1. In simulation, the teacher with the barrier function achieves the lowest forward and lateral velocity tracking errors at 10 kg per arm among evaluated controllers and reduces aggregate divergent component of motion (DCM) excursion magnitude by 35.7% relative to MPC-guided reinforcement learning alone. Our wrist-force teacher withstands torso push disturbances of up to 130 N.
cs.LG / 182 / 2610.09117
Workhorse: Learning Robust Whole-Body Humanoid Loco-Manipulation from Human Data
Abstract
Humanoid robots still struggle to plan contact-rich whole-body manipulation from egocentric RGB and proprioception. Workhorse learns such manipulation from robot-free human demonstrations. A visual planner predicts five-link targets: the poses of the torso, both wrists, and both feet. A reinforcement-learning whole-body tracker follows them on the robot. Both policies train separately on the same recorded human poses, without retargeting. We augment the training data of each policy to imitate the errors that the other makes at deployment. On a real Unitree G1, Workhorse sorts boxes with its hands and a kick, catches a thrown box, and topples and climbs a suitcase. During box sorting, we show recoveries after a person pushes the robot or takes the box away. In a simulated copy of the demonstration room, the system completes box sorting in 77% of episodes, and in 64% under 40 N.s pushes. With both policies retrained from the same demonstrations, a simulated second humanoid completes box sorting in 83% of episodes without pushes.
cs.LG / 183 / 2610.09350
Learning Unknown Constraints without Unsafe Data via Optimality and Counterfactual Regularization
Abstract
Learning from demonstrations (LfD) provides a framework for inferring unknown constraints from locally optimal, constraint-satisfying expert behavior. Existing approaches largely fall into two paradigms, constrained inverse optimal control (CIOC) and inverse constrained reinforcement learning (ICRL). CIOC exploits optimality conditions such as the Karush--Kuhn--Tucker (KKT) conditions but typically assumes known dynamics and structured constraint representations. Meanwhile, ICRL accommodates complex unknown constraints and unknown transition dynamics but often requires extensive online exploration, during which unsafe constraint violations may occur. In this work, we introduce Counterfactual KKT (CF-KKT), a constraint learning framework that leverages learned dynamics and locally optimal demonstrations to recover unknown constraints without requiring known dynamics or additional risky exploration, thereby combining the data efficiency and safety advantages of CIOC with the flexibility of ICRL. First, we use a locally learned differentiable dynamics model to impose KKT-inspired optimality conditions directly on the demonstrations. Second, we use the learned dynamics to generate reward-improving counterfactual behaviors near the demonstrations, revealing behaviors that would be preferable in the absence of the unknown constraint and thus providing synthetic infeasible data. When the constraint parameterization is known, the same learned-dynamics framework enables direct CIOC-based parameter recovery, and we characterize its sensitivity to dynamics misspecification. Across high-dimensional robotic control tasks, our approach learns neural constraint representations with improved safety and data efficiency relative to state-of-the-art offline ICRL baselines.
cs.LG / 184 / 2610.09690
NAViLoss: An Underwater Navigation-Aware Dual-Residual Objective for Physics-Consistent Learning
Abstract
Autonomous underwater vehicles (AUVs) commonly rely on inertial navigation systems (INS) aided by Doppler velocity logs (DVLs) for reliable underwater navigation. Accurate DVL velocity estimation is therefore essential for successful operation. Recent learning-based methods have demonstrated improved DVL velocity estimation, particularly under degraded measurement conditions. However, their training objectives typically rely on conventional regression losses that are highly sensitive to large residuals and corrupted observations. Additionally, they do not explicitly account for the physical consistency and measurement uncertainty associated with the underlying sensing process. To address these limitations, this paper introduces navigation-aware loss (NAViLoss), a robust and uncertainty-aware objective function for learning-based AUV velocity estimation. NAViLoss jointly penalizes the velocity-estimation residual in the navigation-state domain and the beam-consistency residual in the DVL measurement domain. Its bounded formulation limits the influence of large residuals, while an adaptive mechanism regulates the uncertainty in beam geometry. Furthermore, NAViLoss is integrated with a DeepONet architecture to form a novel NAVi-DeepONet model for seamless estimation of an underwater vehicle's velocity. Lastly, our model is evaluated using approximately 10,000m of semi-synthetic AUV experimental data collected during multiple real-world sea trials. Experimental results demonstrate a 44% improvement in velocity-estimation accuracy compared with conventional and learning-based baselines. These results demonstrate the effectiveness of navigation-aware and uncertainty-adaptive loss design for robust learning-based underwater velocity estimation.
cs.LG / 185 / 2610.09943
Many Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization
Abstract
Reinforcement learning (RL) fine-tuning improves vision-language-action (VLA) policies through closed-loop experience, yet generalization beyond the fine-tuning distribution remains limited. Our analysis reveals a selective reshaping of exploration: RL contracts behavior globally, yet diversifies successful trajectories, elicits success with fewer rollouts, and covers more of the latent task-valid solution space than supervised fine-tuning. Broader successful-mode coverage may provide alternative strategies under distribution shifts. Inspired by this, we introduce DRIVE (Diversity-driven RL fIne-tuning for VLA gEneralization), which turns successful-behavior diversity into an explicit RL objective. DRIVE groups rollouts under matched task conditions, compares their trajectories with temporal alignment, and derives a success-conditioned intrinsic reward from relative behavioral diversity. This design encourages broader coverage of feasible solutions without rewarding diverse failures or superficial timing differences. Across LIBERO-Plus, ManiSkill3, and RoboTwin 2.0, DRIVE improves the average out-of-domain (OOD) performance over vanilla RL fine-tuning by 5.3 points on $π_0$ and 2.0 points on $π_{0.5}$. On a dual-arm AgileX PiPER-X platform, DRIVE further increases average OOD success from 64.1% to 73.3% (+9.2 points), demonstrating gains that persist under physical deployment.
cs.LG / 186 / 2610.10283
Temporal Visuo-Tactile Learning for Dexterous Grasp Stability
Abstract
Humans can grasp everyday objects with almost perfect success rates using fingertip tactile feedback, yet much of the robotic grasping literature emphasizes vision-based grasp selection with parallel grippers. In this work, we systematically investigate how high-resolution, dynamic tactile sensing contributes to grasp stability prediction and model-guided grasping in dexterous robotic hands. To this end, we collected a dataset of 10,000 grasp trials across 200 objects using a multi-fingered robotic hand equipped with four Digit 360 tactile sensors, recording external vision, proprioception, and tactile streams throughout each grasp. With this dataset, we trained end-to-end temporal multimodal models to predict post-lift stability from pre-lift grasp observations and compared sensing modalities and encoding backbones. Experimental results and controlled input ablations show that incorporating touch, and particularly high-resolution, dynamic touch, improves grasp stability prediction. Finally, we deployed the learned predictor as an online stability gate on the real robot, where visuo-tactile model-guided regrasping improved the success rate among executed lifts by 10.5 percentage points over a non-tactile gate. These results show how rich fingertip sensing and expressive temporal models that capture the dynamics of touch can support learned grasping with multi-fingered hands without explicit contact or force modeling, providing a scalable data-driven path from tactile experience toward stable dexterous manipulation. The dataset is publicly available at https://lasr-lab.github.io/dexterous-grasp-stability/.
cs.LG / 187 / 2610.10297
Energy-Efficient Gait Adaptation via Hierarchical Reinforcement Learning for Quadrupedal Locomotion Across Diverse Terrains
Abstract
While energy efficiency is a critical objective for legged-robot locomotion control, achieving low energy consumption while maintaining robust performance across different velocity ranges and terrain conditions remains a key challenge. This is particularly true for end-to-end RL policies, where gait generation, motion execution, and energy optimization are tightly coupled, leading to high sensitivity to reward design. In this work, we propose a hierarchical reinforcement learning (HRL) framework that separates a high-frequency policy for stable and robust joint-level motion execution from low-frequency gait adaptation that explicitly minimizes the cost of transport (CoT). The three-stage Isaac-based training procedure enables zero-shot sim-to-real transfer with improved tracking accuracy, robustness, and energy efficiency. The learned hierarchy exhibits automatic speed-dependent gait adaptation, transitioning from pacing at low speeds to trotting at higher speeds. We validate the proposed approach in simulation against representative single-policy and hierarchical locomotion baselines, demonstrating reduced CoT over a broad range of commanded velocities, while maintaining robust locomotion across flat, uneven rough, and inclined terrains. We further demonstrate its practical feasibility through zero-shot deployment on a physical Unitree AlienGo quadruped.
cs.LG / 188 / 2610.10409
RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments
Abstract
General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world. To investigate this, we introduce RobotWorld, a challenging simulation testbed for robot use: turning instructions and observations into physical task execution through robot interfaces. Its 84 tasks span manipulation, mobile manipulation, locomotion, driving, and aerial control, with explicit interaction budgets and executable success checks. By analysing task outcomes alongside execution traces, we identify both the capabilities that transfer and the gaps that prevent reliable completion. Furthermore, we find that current agents can construct sophisticated perception and control workflows, including image segmentation, camera calibration, spatial estimation, and dynamics-based computation. These capabilities, however, do not consistently compose into successful behaviour: agents lose task-relevant object states despite reaching commanded poses, fail to correct ineffective actions, recover too late, or mistake unfinished tasks for completion. This uneven transfer also differs across models: Astra succeeds more often on spatial and constrained-contact goals, whereas Opus 5.5 succeeds more often on continuous-balance and timed-interaction goals. By linking these outcomes to execution behaviour, RobotWorld provides both a rigorous proving ground and an empirical account of the remaining capability gaps, thereby establishing concrete targets for training and designing more reliable physical-world agents.
cs.LG / 189 / 2610.09075
Towards AI-Generated Music Plagiarism Detection as a Version Identification Problem
Abstract
The rapid expansion of text-to-music generative models challenges traditional paradigms of music creation and intellectual property. Plagiarism in this context is rarely an absolute mathematical binary, but an ambiguous threshold negotiated over harmonic structure, melodic contours, or overall perceived stylistic character. In this work, we test the transferability of state-of-the-art music version identification architectures from the human-to-human cover domain to the human-to-AI plagiarism setting. To evaluate this task, we introduce COPYCAT, a benchmark derived from real-world plagiarism cases and extended through generative re-synthesis and digital signal processing obfuscations, yielding 350,654 evaluation pairs. We show that scalar distance thresholding collapses under generative re-synthesis, while a supervised framework leveraging coordinate-wise embedding shifts recovers the dispersed plagiarism signal, raising overall $F_{0.5}$ from $0.612$ to $0.803$.
cs.LG / 190 / 2610.09637
Tracing Inputs, Verifying Outputs: Validating Attribution in Music Generation
Abstract
How can we verify whose music contributed to an AI-generated output? This paper demonstrates how input-based attribution can provide verifiable evidence of which audio sources were used in a generation and whether they shaped the output. To do so, we condition the generation solely on audio without any text input, then trace the inputs behind each output, and establish their musical effect. In prompt adherence tests and controlled input swaps, the stems generated by our generator, MixAudio, follow the prompt audio in timbre and the context audio in harmony. Yet these outputs may still reproduce training data not supplied as inputs. We therefore audit memorization with our musical version identification model, musicDNA, and find few reproductions outside the input records. On human-judged cases within the flagged pool, it achieves higher precision and recall than the other tested memorization detectors. The two evaluations suggest that input records and output analysis provide complementary evidence for attribution, on which rights-holder reporting and compensation can draw as the AI music economy takes shape. Audio examples are available at https://neutune.github.io/attr2027demo/
cs.LG / 191 / 2610.09819
Backdooring Acoustic Foundation Models for Physically Realizable Triggers
Abstract
Acoustic foundation models (AFMs) have democratized acoustic applications, enabling powerful models for tasks ranging from speech recognition to speaker verification with minimal resources. However, the security of applications based on AFMs remains largely underexplored. Our work addresses this gap by proposing the Foundation Acoustic model Backdoor (FAB) attack, demonstrating that state-of-the-art AFMs are susceptible to backdooring under practical settings. Despite making minimal assumptions about adversary capabilities (e.g., no access to pre-training data), we show that FAB preserves benign performance while inducing backdoors that survive fine-tuning and cause significant degradation across diverse downstream tasks when activated. Notably, FAB utilizes task-agnostic, physically realizable, inconspicuous, and sync-free triggers (e.g., a background siren). We evaluate FAB using two leading AFMs, nine downstream tasks, and four different triggers. We further demonstrate its effectiveness against established defenses and across both digital and physical domains. While extensive end-to-end fine-tuning can mitigate FAB, such a defense is resource-intensive and task-specific. Our work highlights critical risks to AFMs and calls for advanced defenses.
cs.LG / 192 / 2610.10208
CARES: A Controlled Synthetic Benchmark of Speaker Reactions to Sound
Abstract
Automatic audio scene description turns a recording into a text account of a situation. One difficulty is deciding which elements of the audio should be kept, since a description cannot include them all. Annotators disagree about this, making a ground truth hard to obtain. In this work, we first define the ground truth, then generate the data. We focus on audio events and define sound salience with a simple rule: a sound is salient when a speaker audibly reacts to it. For scale and variety, a controlled set of scenarios fixes the ground truth, and a language model writes the dialogues. The resulting corpus, CARES, contains 10,000 two-speaker scenes. We then benchmark six audio-language models on three tasks: identifying the scene, tagging the sounds present, and classifying reactions. We show that these models hear the sounds but miss how the speakers react to them.
cs.LG / 193 / 2610.10415
Steerspeech: Activation Steering For Emotion Control In Generated Speech
Abstract
Pretrained text-to-speech (TTS) models can generate expressive speech, but reliable inference-time emotion control remains challenging: prompts and reference audio offer coarse, inconsistent control, whereas specialized conditioning and model adaptation require costly training. We present SteerSpeech, a lightweight activation-steering framework that controls emotion by injecting steering vectors into hidden activations. For each target emotion we train a lightweight low-rank transform, using a multi-expert objective that encourages monotonic emotion control while preserving speaker identity and linguistic content, constraining steering drift, and keeping the TTS backbone frozen. To optimize through discrete speech tokens, we introduce a two-pass generation-and-replay pipeline using a straight-through estimator to backpropagate expert supervision through sampled tokens. At inference, a target-emotion steering direction is optimized with its respective transform and injected into the base TTS model. Objective and subjective evaluations with Qwen3-TTS across seen, unseen, and accented speakers show stronger continuous emotion control with limited speaker and content degradation. SteerSpeech achieves 1.08x-7.12x baseline target-emotion scores and for a representative emotion subjectively, it receives 78.1%-96.8% intensity preference and 1.43x-1.46x speaker-identity preservation at high steering strengths.
cs.LG / 194 / 2610.09737
A Strength-Monotonic Law for Domain Alignment in Frozen-Embedding Bioacoustic Classification
Abstract
When does distribution alignment help a frozen foundation-model embedding generalize across acoustic domains? For cross-domain mosquito-species classification we report a strength-monotonic law: the stronger an encoder is on the target task, the more its unseen-domain generalization relies on a distribution-alignment (MMD) term, and the more it is harmed by domain-rebalanced sampling. Across four encoder families and a within-encoder HuBERT layer sweep (n=8), the rebalancing leg orders exactly with encoder strength (Spearman -1.000), while the MMD-benefit leg is monotonic within each stream and -0.857 pooled; fixing architecture and varying only representation strength flips the rebalancing effect from benefit to collapse. The law is actionable: a single MMD term is the sole lever on a strong encoder, so we reduce the field's default recipe to a frozen Perch 2.0 embedding, a lightweight probe, cross-entropy, one MMD, and input augmentation. The reduced recipe stays within seed noise of the full composite (BA_unseen 0.299+/-0.006 vs. 0.307+/-0.014). As boundary conditions of the same law, three community defaults (backbone fine-tuning, multi-modal fusion, and domain rebalancing) each hurt unseen-domain accuracy under a leave-domain protocol, shown with single-variable, multi-seed evidence. We present a mechanism and the recipe it explains, not a leaderboard entry.
cs.LG / 195 / 2610.09617
Second-order optimization of variable projection SVM models and road abnormality detection
Abstract
We introduce a novel second-order optimization framework for minimizing so-called variable projection functionals. We demonstrate that the proposed framework is especially usefulfor the training of variable projection based kernel methods. In particular, the problem of efficiently training variable projection support vector machines (VP-SVMs) is considered. We show the effectiveness of the proposed training methodology in a real-world application, namely we demonstrate how second-order trust region algorithms can be used to train VPSVM models to recognize road surface abnormalities based on 1D signals obtained from a tire sensor.
cs.LG / 196 / 2610.10054
Transition Path Sampling Using Koopman Operators and Exit-Time Optimal Control
Abstract
Sampling transitions between metastable states is a central problem in dynamical systems theory and molecular dynamics in particular. A key challenge is the existence of high free-energy barriers that separate the states, making transitions extremely rare. Recent machine learning-based methods cast transition path sampling (TPS) as an optimal stochastic control (OSC) problem over a fixed time horizon, and parameterize the drift bias via a neural network trained by simulation-in-the-loop, requiring repeated biased rollouts. To address computational and performance guarantee issues of these models, we propose a new approach for the problem based on Koopman operators. Because Koopman operators are linear, their leading eigenfunctions reveal the metastable sets and provide an estimate of the committor function with no transition path information required. Furthermore, we formulate TPS as an OSC problem up to an exit time. Our time horizon is the first hitting time of the target set, and our running cost penalizes time spent in nonreactive regions by encoding the estimated committor function. We derive the optimal controller in closed form and approximate it in a reproducing kernel Hilbert space (RKHS). This reduces the problem of constructing the optimal controller to solving a single equality-constrained quadratic program, whose solution can be characterized by a linear Karush-Kuhn-Tucker (KKT) system. On the two-channel double well and alanine dipeptide, our controller increases the fraction of trajectories reaching the target from 0% to 99.8% within 1000 steps, and from 0% to 93% within 1ps, respectively.
cs.LG / 197 / 2610.09147
Sampling SU(N) gauge theory on a 2D lattice from independent plaquettes via holonomies and corner reweighting
Abstract
In the Wilson formulation of lattice gauge theory, the fundamental degrees of freedom are group-valued link variables, while the action is a sum over the trace of the plaquettes, the smallest Wilson loops. For generative sampling methods such as normalizing flows, this poses a challenge: the distribution of an individual plaquette is easy to model, but mapping sampled plaquettes to the links is the obstruction. We explore a statistical way around this in two dimensions for the $\mathrm{SU}(N)$ gauge theory with the Wilson plaquette action. The action can be written in terms of holonomy variables, which determine the links once a consistency condition at the four corners of an extended lattice is satisfied. Using a boundary condition that only affects the holonomies, we trade this consistency condition for a conditional sampling problem, which reduces to sampling $X,Y\in\mathrm{SU}(N)$ given the group commutator $Z=XYX^\dagger Y^\dagger$, where $Z\in\mathrm{SU}(N)$ depends on the four corners. Plaquettes are sampled independently, and the resulting configurations carry weights due to the additional conditional sampling. We obtain the density of $Z$, which determines these weights, in closed form for $\mathrm{SU}(2)$ and $\mathrm{SU}(3)$. A normalizing flow models the single-plaquette distribution for $\mathrm{SU}(2)$ and $\mathrm{SU}(3)$; for the group-commutator problem we use a closed-form sampler for $\mathrm{SU}(2)$ and a trained normalizing flow for $\mathrm{SU}(3)$. In our tests on $2\times2$ and $32\times32$ lattices at two couplings each, per-plaquette acceptance rates exceed $99\%$ and the effective sample size of the final weights is moderate, typically above one half.
cs.LG / 198 / 2610.10206
Computations of the slice genus and the unknotting number of links via machine learning
Abstract
Links are disjoint unions of circles smoothly embedded in $S^3$. We use reinforcement learning and Bayesian optimisation to obtain new upper bounds on several link invariants that are not known to be algorithmically computable: the slice genus and the unknotting number for links, and the strong slice genus for algebraically split links. We also compute lower bounds using known invariants. Combining the upper and lower bounds, we obtain new exact values in many cases. Our unknotting agents can reproduce the non-additivity of the unknotting number for several counterexamples due to Brittenham and Hermiller, in some cases finding new unknotting trajectories.
cs.LG / 199 / 2610.10278
Neural Sampling with Reweighted Normalizing Flows via the Wasserstein--Fisher--Rao JKO Scheme
Abstract
We propose a neural algorithm for sampling from distributions specified by unnormalized Boltzmann densities. Our approach is based on the Jordan--Kinderlehrer--Otto scheme for the Kullback--Leibler divergence in the Wasserstein--Fisher--Rao geometry (WFR JKO scheme). Our contributions are twofold. First, we prove that, for any fixed step size, the exact WFR JKO iterates converge exponentially fast to the target as the number of iterations tends to infinity. Notably, this result requires no structural assumptions on the target, such as log-concavity or a logarithmic Sobolev inequality. Second, we develop a neural implementation of the WFR JKO scheme that parametrizes its transport and reaction components using reweighted normalizing flows. Numerical experiments on challenging multimodal targets demonstrate the promising performance of the proposed method.
cs.LG / 200 / 2610.09244
Beyond Nominal Equilibria: Risk-Averse Multi-Population Mean-Field Games
Abstract
Recent advances in mean-field games and its multi-population variants enable large-scale heterogeneous multi-agent systems to be modeled through representative agents and their associated mean-field distributions. However, existing approaches do not explicitly account for uncertainty in the behavior of other populations. To this end, we introduce a new paradigm: risk-averse multi-population mean-field games, where each population optimizes a worst-case expected reward over dynamically feasible ambiguity sets of mean-field flows of a subset of the other populations. Employing an occupation-measure formulation along with tools from set-valued analysis, we establish, under mild assumptions, several theoretical properties of the multi-population game, including the geometric properties of the ambiguity sets and the existence of a novel risk-averse multi-population mean-field equilibrium. Further, we derive contractivity results of the fixed-point operator under entropy regularization and show that it can be utilized to learn the equilibrium. Finally, we propose a risk-averse fictitious-play scheme and show that exploitability decays to zero, despite the additional nonlinearity introduced by the worst-case objective. We report several numerical experiments to illustrate convergence and risk-averse behavior.
cs.LG / 201 / 2610.09712
Boundary-aware Reinforcement Learning for Hypercube State Spaces via Deterministic Policy Gradient
Abstract
We develop a continuous-time deterministic policy gradient framework for reinforcement learning with reflected state dynamics, where the state process is governed by a controlled reflected stochastic differential equation on a hypercube. Under suitable regularity assumptions, we establish the connection between the value function and the Neumann Bellman equation, introduce an advantage-rate function that yields a deterministic policy gradient formula, and prove the martingale characterization theorem. Motivated by these theoretical results, we propose a continuous-time deep deterministic policy gradient algorithm for reflected stochastic systems, in which the Neumann boundary condition is imposed via either soft penalization or hard architectural constraint. We further quantify the discrepancy between the ideal continuous-time dynamics and the discretely sampled exploratory dynamics executed in practice, showing that the error decays as the time grid is refined and exploration noise vanishes. Our experiments on reservoir control problems illustrate the effectiveness of the RL framework, highlighting that boundary-aware methods substantially reduce Neumann boundary residuals and enhance learning stability.
cs.LG / 202 / 2610.10527
Decentralized SGD under Heavy-Tailed Noise: Optimal Convergence Rates and the Role of Gradient Clipping
Abstract
Heavy-tailed noise has been widely observed in modern machine learning, motivating the use of methods like gradient clipping and normalization. While these methods are well understood in centralized settings, much less is known in decentralized ones, where applying a nonlinearity to local gradients affects both optimization and consensus. Recent works on decentralized non-convex optimization have studied both clipping and normalization under heavy-tailed noise, with clipping yielding suboptimal rates and normalization needing local momentum or mini-batches to converge. This raises the question: can a baseline decentralized method using a nonlinearity achieve optimal convergence rates under heavy-tailed noise? We answer affirmatively with clipped decentralized SGD ($\mathtt{DSGD}$). For smooth non-convex costs under bounded $p$-th moment noise, $p \in (1,2]$, we show that clipped $\mathtt{DSGD}$ achieves order-optimal rates both with high probability and in expectation. Moreover, we establish a linear speed-up in the number of agents, which, to our knowledge, has not been shown for decentralized methods with clipping. The key technical ingredient is a sharp analysis of the consensus gap that exploits the structure of clipping, relegating network effects to higher-order terms. Our results highlight an important distinction between clipping and normalization in decentralized settings: while normalized $\mathtt{DSGD}$ can fail to converge, clipping retains magnitude information, enabling $\mathtt{DSGD}$ to be convergent and order-optimal. Numerical experiments validate our theory.
cs.LG / 203 / 2610.09687
The Silhouette Operator: Identifiability of Low-Rank Measures from One-Dimensional Projections
Abstract
Structured recovery phenomena, such as restricted isometry properties in compressed sensing, have shown that high-dimensional objects can often be reconstructed from remarkably low-dimensional linear measurements. This work develops an analogous recovery framework for low-rank signed measures on $\mathbb{R}^2$, defined here as measures that can be expressed as finite sums of product measures with one-dimensional factors. The framework is based on linear operators, termed "silhouette operators," that map a measure to a fixed finite collection of one-dimensional linear pushforwards. The main results show that a suitably chosen collection of $2k$ projected marginals suffices to identify every compactly supported rank-$\le k$ signed measure, that this number is optimal, and that the projection directions cannot be chosen arbitrarily. The framework is also extended to higher-dimensional sums of product measures by establishing sufficient conditions under which collections of pairwise marginals identify the full model. Building on this framework, a computationally efficient estimator, termed "silhouette mixture estimation" (SME), is introduced for constructing a low-rank empirical measure from data by matching its one-dimensional projected marginals to the corresponding empirical marginals in Wasserstein distance. When combined with one-dimensional density estimators, SME yields an efficient nonparametric density estimator that performs strongly relative to a range of parametric, nonparametric, and deep-learning baselines in settings of moderate dimension and sample size.
cs.LG / 204 / 2610.09731
Unbounded Characteristic and Universal Kernels
Abstract
Kernel methods are among the most powerful tools in machine learning and statistics, with a large number of successful applications. Their immense success stems from the flexible function class associated to each kernel---its reproducing kernel Hilbert space (RKHS)---which facilitates statistical analysis, as well as from their computational tractability and applicability to many domains. Multiple notions (such as characteristic, $L_p$-universal, and integrally strictly positive definite) capture the expressivity of kernels and their RKHSs and play a key role in understanding the statistical properties of kernel methods; these concepts and their relations are well-understood for bounded kernels. Even though unbounded kernels have received significant attention over the past decade (for instance, in the construction of kernel-based discrepancy and dependence measures such as the maximum mean discrepancy, the Hilbert-Schmidt independence criterion, and the kernel Stein discrepancy), surprisingly little is known about the relations of these notions in the unbounded case. In the present paper we tackle this severe bottleneck, establishing their relations under mild assumptions.
cs.LG / 205 / 2610.10080
Sharp Asymptotic Theory of Maximum Likelihood Estimation for Gaussian Processes with an RBF Kernel
Abstract
Gaussian processes (GPs) are widely used across machine learning, spatial statistics, time-series analysis, optimization, Bayesian statistics, and scientific applications. A central component of a GP model is its kernel, which is typically specified through a parametric family. Among the most widely used choices is the radial basis function (RBF), also known as the squared exponential or Gaussian kernel, owing to its simple form, smoothness, and flexibility. In practice, the kernel parameters are routinely estimated by the maximum likelihood estimators (MLEs), as implemented by standard GP software. Despite this widespread use, the asymptotic behavior of the MLEs remains poorly understood under fixed-domain asymptotics, even for the RBF kernel. The main difficulty arises from the increasingly strong dependence among densely sampled observations and the nonlinear dependence of the covariance matrix on the kernel parameters. In this paper, we address this gap by providing, to the best of our knowledge, the first complete asymptotic characterization of the joint MLE of the spatial variance, lengthscale, and nugget variance under fixed-domain asymptotics. We establish consistency, derive convergence rates for all three parameters, prove joint asymptotic normality, and show that these rates are minimax optimal.
cs.LG / 206 / 2610.09576
PEACE: Covariant learning of nonadiabatic manifolds with parity-resolved Hamiltonians
Abstract
Nonadiabatic molecular dynamics provides mechanistic insight into light-driven processes and informs the design of molecules and materials for solar energy conversion, photocatalysis and photo switching. Accurately describing these processes requires a representation that respects electronic symmetry and consistently relates energies to interstate couplings. Here we introduce PEACE, which combines a parity-equivariant latent Hamiltonian with a learned electronic connection. Controlled ablations reveal the complementary roles of symmetry-allowed state mixing and electronic-frame variation in reproducing crossing structures and relaxation dynamics. PEACE closely reproduces excited-state population dynamics from first-principles simulations, while its extension to spin-orbit coupling enables simulations of intersystem crossing. These results demonstrate that a more complete incorporation of the underlying physics into learned electronic representations leads to more accurate predictions of nonadiabatic dynamics.
cs.LG / 207 / 2610.10430
Conditional Flow Matching for Generation of 3D Multi-variable Instantaneous Urban Microclimate Fields
Abstract
Rapid and accurate prediction of urban wind and temperature fields is important for urban microclimate design and climate adaptation. Large-eddy simulation (LES) effectively resolves these instantaneous fields, but its application is limited in iterative design of urban microclimate applications due to high computational cost. Existing regressive data-driven models offers quick outputs, but they produce only deterministic point predictions that inherently fail to represent turbulent stochasticity. This paper adopts a novel generative framework of Conditional Flow Matching (CFM) that uses building geometry and mean flow as guidance to generate plausible three-dimensional instantaneous velocity and temperature fields for urban microclimate in seconds. To overcome the GPU memory bottleneck of pixel space 3D generation, the model operates in parallel on overlapping pixel space through a shared-noise initialization that preserves high spatial continuity of flow structure across the entire domain. Against reference LES data, the CFM surrogate can rapidly and accurately restore the first-order statistics with Normalized Root Mean Square Error (NRMSE) of 2.99% for wind and 1.77% for temperature, second-order turbulence metrics with NRMSE of 7.17% for wind and 8.84% for temperature, turbulent kinetic energy with NRMSE of 7%, probability density function and vertical profiles in representative locations. Wind engineering application of local gust prediction demonstrate that the speed and accuracy of CFM, supporting the use of generative AI for making turbulence-aware resilient urban design and climate adaptation more computationally feasible.
cs.LG / 208 / 2610.10023
Structure alone supports efficient visual computation in the Drosophila visual system
Abstract
Understanding the extent to which measured synaptic wiring determines computation remains a central challenge. Here, we couple the proofread adult Drosophila melanogaster connectome to an anatomically faithful model of its eye. Visual information is inputted in the eye model, then passed to the connectome, and finally read from a Kenyon-cell-centered linear decoder. This creates a connectome-only model in which the anatomical graph and eye geometry are fixed and only scalar synaptic gains and neuronal thresholds may be learned. The model supports multitask vision, including color discrimination, shape classification, and numerical discrimination that follows a ratio-dependent scaling characteristic of approximate number perception. To test whether precise connectivity is consequential under wiring economy, we compare the biological graph to randomized ensembles that increasingly preserve biological synaptic constraints. At matched wiring cost, the biological network consistently yields higher accuracy, whereas less constrained rewiring surpasses it at the cost of inflated wiring. These findings indicate that the measured connectivity and eye geometry jointly set efficient operating points for visual computation.
cs.LG / 209 / 2610.09613
Residual Learning in Empirical Asset Pricing
Abstract
Shallow models are special cases of deep models, and deep models theoretically have the potential to outperform the shallow ones. However, the existing empirical asset pricing literature provides strong benchmarks for shallow models. Residual learning allows neural network models in asset pricing to go deeper by preserving and refining their shallow counterparts. The out-of-sample Sharpe ratio for value-weighted long-short portfolios of deep residual models (2.07) is higher than that for the corresponding shallow ones (1.92) and more than twice that of the deep feedforward models (0.89). We show that model depth is a source of additional economic value in asset pricing. Residual learning can be used to deepen other neural-network-based asset pricing models if they contain intermediate layers. Our design also provides one way to scale asset pricing models, making native "large asset pricing models" more feasible.
cs.LG / 210 / 2610.09635
Quantum anomaly detection in real scarce data
Abstract
Anomaly detection on small and unbalanced datasets remains very challenging in machine learning, although this scenario is common in several domains, including healthcare, cybersecurity, finance, and energy. Data augmentation and generative AI may mitigate training-data scarcity, but they often fall short because anomalies are, by definition, unpredictable, rare, and highly diverse events compared to high-probability normal data. Overfitting to pseudo-anomalies, model collapse, high-dimensional data, uninterpretable black-box models, and validation challenges are typical issues limiting their practical applicability. In this context, quantum machine learning may provide a promising and more sustainable avenue because it can enable more interpretable models with far fewer trainable parameters and smaller datasets, implementable on energy-efficient quantum hardware. Here, we propose a novel two-step hybrid classical--quantum architecture for sequential data and test it on a realistic scenario in the global energy-transition domain, i.e., automated anomaly detection in large-scale photovoltaic plants. The achieved generalization capability and competitive prediction accuracy may pave the way for new hybrid learning models able to exploit the continuously increasing power of cloud-available and more sustainable quantum accelerators integrated with more traditional energy-hungry High Performance Computing resources.
cs.LG / 211 / 2610.09641
Q-PhotoMarket: A Design Space Exploration Framework for Photonic Hybrid Quantum Neural Networks in Financial Market Prediction
Abstract
Photonic quantum computing has recently emerged as a promising platform for hybrid quantum machine learning due to its native realization of linear-optical circuits and the computational complexity of boson sampling. However, despite growing interest in quantum methods for finance, the influence of photonic circuit design choices on predictive performance remains largely unexplored. Existing studies typically evaluate a single architecture, leaving the broader photonic design space unexamined. In this work, we present Q-PhotoMarket, a systematic design space exploration (DSE) framework for photonic hybrid quantum neural networks (HQNNs) applied to financial market prediction. We explore over 5,000 valid photonic configurations spanning input photon states, circuit architectures, entangling models, and measurement strategies across their compatible computation spaces, for U.S., Indian, and cryptocurrency markets. To improve search efficiency, the exhaustive exploration is complemented with Bayesian optimization. We further incorporate threshold calibration and prediction-collapse diagnostics to enable reliable evaluation under increasingly imbalanced return thresholds. Experimental results show that systematic exploration of more than 5,000 photonic HQNN configurations reveals consistent architectural patterns across financial markets, identifies robust high-performing designs, and demonstrates competitive performance relative to classical machine learning baselines.
cs.LG / 212 / 2610.10351
Measurement-Efficient Differentiable Quantum Architecture Search for Combinatorial Optimization
Abstract
Differentiable quantum architecture search (DQAS) is a promising framework for the automated design of quantum circuits, particularly for variational quantum optimization algorithms. However, its practical deployment on quantum hardware is limited by the large number of circuit measurements required during optimization, making hardware execution costly. In this work, we show that for a broad class of combinatorial optimization problems and commonly used rotational gate parameterizations, the measurement cost of DQAS can be significantly reduced without changing the optimization objective. We derive the proposed measurement reduction scheme theoretically and validate it experimentally on 3-SAT and MaxCut benchmark problems. Our approach reduces the requested gradient measurement cost by about 39 to 41% while introducing only negligible classical post-processing overhead, lowering the practical cost of executing DQAS on quantum hardware.
cs.LG / 213 / 2610.09898
Learning joint probabilistic weather forecasts from station observations alone
Abstract
Assessing compound weather risks requires forecasts representing dependence between variables. CLARA (Calibrated Advection-Routing Attention) learns joint Gaussian predictive distributions of five surface variables from station observations alone, without numerical weather prediction or reanalysis; the approximately 28,000-parameter model supports CPU training and prediction. Across six multi-year folds on 96 stations, its lead-mean energy score is 4.9% lower than that of a learned comparator with matched temporal inputs (4.7% with a similar parameter count) and 11-65% lower than those of statistical baselines. Holding marginal variances fixed, removing learned correlations worsens joint negative log-likelihood by 1.0-2.8 nats per station. A covariance-scale estimator, proved consistent under stated assumptions, improves short-lead calibration but over-corrects at long leads. Synthetic interventions show an attention-bias coefficient alone does not measure forecast influence. Retrained in ten regions on six continents, CLARA outperforms persistence in all 60 multi-year region-lead comparisons and a similarly sized learned model in 57 of 60.
cs.LG / 214 / 2610.10321
Estimating Uncoded Crash Factors with Tabular Foundation and System One Models: Kumo Tabular and Jev
Abstract
Road safety programs count the coded fields of police crash records, while the officer's narrative, which often records factors the fields omit, is rarely read. A safety office thus cannot tell how much its counts miss or where to review. This study develops and evaluates a system that joins both views of the 5,601,890 Texas crashes from 2017 to 2025 into population estimates with stated validity. An in-context tabular foundation model, Kumo Tabular, reads the coded record of every crash, a calibrated System One model, Jev, reads the narratives of two probability samples, and human judgments recalibrate its probabilities. A multiwave predict-then-debias estimator joins the three tiers, and a second human tier drawn with recorded probabilities checks the estimates by design. For hydroplaning, medical episodes, fatigue, animals, and phone use, the narrative documents more injury crashes than the coded field, 15,074 against 7,340 for phone use, and the human check agrees with all fifteen estimates within its margin. A re-read list ranked by Kumo Tabular finds confirmed discordance 7 to 58 times as often as random reading. At the planning cost of human coding, one further round of human judgments would cut the root mean square relative half-width from 22.0 to 16.2 percent, against 21.2 for reading every narrative. Two calibrated readers of different views, joined by a sampling design, give a safety office counts, a discordance map, a validated re-read list, and a reading budget, with Kumo Tabular reading the table at 15 times the speed of TabPFN 3.5.
cs.LG / 215 / 2610.10167
Policy Learning with Weak Signals
Abstract
Policy learning in digital experimentation faces three challenges: weak signal-to-noise ratios, rich covariate spaces, and massive data volumes. We formalize this regime by modeling treatment-effect estimates from increasingly fine covariate partitions as Gaussian observations with bounded signal-to-noise ratios. We establish that, in general, the optimal treatment policy is not learnable in this setting. Even learning the optimal policy value suffers from impractically slow rates. However, when treatment effects vary smoothly, we derive minimax-adaptive policies based on linear smoothers that achieve vanishing welfare regret. We demonstrate the practical value of our framework by applying it to large-scale real-world experiments at Netflix, showing that personalized linear-smoothing policies can dominate unpersonalized policies even in this challenging empirical setting.
cs.LG / 216 / 2610.09044
Quadratic Weak-to-Strong Generalization in Random Feature Networks via Random Matrix Theory
Abstract
Weak-to-strong generalization is the phenomenon where a strong student model trained with labels produced by a weak teacher model is able to generalize better than the teacher. In this paper, we study this phenomenon in two-layer random feature networks where the model strength is determined by its width. Using tools from random matrix theory, we derive deterministic equivalents for the population errors of an optimally trained teacher and a student trained with gradient flow. For ReLU activation and a pure spherical harmonic target, we obtain sharp asymptotics under a Gaussian universality assumption, showing a quadratic improvement: the student error scales as the square of the teacher error. These results attain the general lower bound of Medvedev at al (2025). We also analyze how the student behaves under more general stopping times and targets supported on multiple harmonic degrees, characterizing the regimes in which weak-to-strong generalization occurs and identifying the transition between quadratic, non-quadratic, and no improvement.
cs.LG / 217 / 2610.09070
Covariate-dependent Joint Modeling of Multivariate Ordinal Preferences and Its Connections with Comparison Models
Abstract
Multivariate ordinal data along with covariates are commonly collected in problems ranging from alignment of language models with human preferences, as well as in recommender systems. For example, data sets such as MovieLens contain several movies rated on a scale 1--5 by human users, along with their demographic information such as age or gender. Similarly, data sets such as HelpSteer collect human feedback on several attributes such as "helpfulness" or "verbosity" of LLM response on an ordinal scale, with covariates depending on the LLM prompt--response pairs. Unfortunately, the standard approaches for modeling these data (a) look at the attributes individually rather than jointly, and (b) often convert the data into pairwise or list-wise win--loss comparisons for fitting models such as Bradley--Terry and Plackett--Luce. Both of these lead to a coarsening of what is actually observed, which we address via a joint covariate-dependent consecutive ratio Markov random field model. We also show pairwise or listwise comparison models are obtained under restrictions of our joint model, and that joint modeling improves comparisons. We also develop a maximum likelihood inference procedure even in the presence of an intractable normalizer.
cs.LG / 218 / 2610.09132
The Impact of Likelihood Tempering on the Limiting Predictive Moments of Variational Bayesian Linear Neural Networks
Abstract
In wide Bayesian neural networks, Gaussian mean-field variational inference is prone to "prior dominance": the Kullback-Leibler (KL) regularization term of the ELBO outweighs the expected log-likelihood, and the variational predictive distribution collapses to the prior predictive as the width $M$ grows. Tempering the likelihood, by raising it to the power $1/T$ for a temperature $T < 1$, is equivalent to scaling the KL term by $T$. We ask in this paper how fast $T$ must decrease with $M$ to counteract this degeneracy and strike a good balance between the two terms. For single-hidden-layer linear networks with isotropic Gaussian priors, we derive the limiting predictive distribution under schedules of the form $T = τ/M^{c}$, with constants $τ, c > 0$, as $M \to \infty$ and compare it with the untempered neural network Gaussian process (NNGP) posterior, the infinite-width limit of the exact posterior. Our main result is that the predictive expectation and variance undergo phase transitions at different scales: the limiting expectation leaves its prior value at $c = 1/2$, once $τ$ falls below an explicit threshold, and equals the least-squares prediction for $c > 1/2$, whereas the limiting variance keeps its prior value for $c < 1$, matches the NNGP's for $c=1$, and vanishes for $c > 1$. With suitable choices of $τ,c$, one can recover either the NNGP posterior expectation or its variance.
cs.LG / 219 / 2610.09505
Adjoint-Based Calibration and Optimal Control of Stochastic Multiscale Bioprocess Digital Twins
Abstract
We develop a bias-aware digital-twin calibration and control framework for multiscale bioprocess models within a biological systems-of-systems (Bio-SoS) paradigm. The digital twin is represented by a stochastic differential equation (SDE) model and calibrated from sparse, discrete observations using quasi-likelihood estimation and adjoint sensitivity analysis. SDE generator-based moment expansions characterize truncation-induced parameter bias, while forward-backward adjoints quantify how calibration uncertainty propagates to value functions and policy performance. The resulting parameter-error distribution supports both policy-directed adaptive experimental design and uncertainty-aware policy optimization through a second-order Gaussian-averaged objective. We characterize the asymptotic behavior of the resulting exploration criterion and derive a physical-system performance under the optimized policy. To implement these ideas, we develop an Actor-Simulator algorithm that jointly updates model parameters, selects informative experiments, and optimizes control policies. Numerical studies demonstrate improved calibration accuracy, sample efficiency, and control performance relative to state-of-the-art baselines.
cs.LG / 220 / 2610.09530
Unpaired Canonical Correlation Analysis
Abstract
Canonical Correlation Analysis (CCA) is a fundamental method for multiview shared space learning. However, its strict reliance on paired data poses a significant limitation, as such data is often difficult to obtain or entirely unavailable. In this paper, we present Unpaired CCA (UCCA), a novel method that learns linear projections to maximize the correlation of the true underlying pairing without access to any paired samples during training. We first establish theoretical results connecting the Quadratic Assignment Problem (QAP) to CCA. Leveraging these theoretical insights, we derive a practical method to maximize correlation exclusively from unpaired data. To the best of our knowledge, UCCA is the first approach to learn maximally correlated projections in a strictly unpaired setting. We validate UCCA on real-world multi-modal datasets, demonstrating that it significantly outperforms recent unpaired alignment baselines in recovering the underlying true correlation. This work fills a critical gap between traditional statistical multiview learning and the growing field of unpaired data learning.
cs.LG / 221 / 2610.09591
Scalable Logistic Gaussian Process Density Regression with Kinetic Langevin Sampling
Abstract
Conditional density estimation targets the full distribution of a response given covariates, as required, for example, for per-galaxy photometric redshifts. We develop a scalable Bayesian estimator based on the logistic Gaussian process. The log conditional density has a separable covariance: a Matérn kernel along the response, represented in a truncated Fourier basis on a circle, and a covariate kernel represented by Nyström features, which accommodate non-stationary kernels with input-dependent amplitudes and length scales. Instead of a Laplace or variational approximation, we sample the latent field of this finite-feature model. Given the hyperparameters, its posterior is strongly log-concave with a uniformly bounded Hessian, and we draw from it by simulating kinetic Langevin dynamics with symmetric minibatch splitting in Kronecker-whitened coordinates. Marginal-likelihood gradients follow from Fisher's identity as posterior expectations. Under the conditions of our analysis their bias is controlled by the sampler's step size and run length, and the predictive averages over the non-Gaussian latent posterior instead of a Gaussian around its mode. On photometric-redshift benchmarks with up to 3.9 million training observations, trained on a single GPU, the estimator is competitive with state-of-the-art tabular foundation models on density and calibration metrics.
cs.LG / 222 / 2610.09886
Sparsifying Stochasticity, Not Capacity: Partial Stochasticity via Deep Weight Factorization of Prior Scales
Abstract
Bayesian neural networks need not be fully stochastic to be universal conditional density approximators, but it remains open which parameters should be stochastic. We learn this split by applying deep weight factorization to the prior scales, which are the standard deviations of the parameter priors, while fitting the functional prior to a Gaussian process with a maximum mean discrepancy objective. A parameter whose prior scale falls below a cutoff becomes deterministic and is optimized during inference, so the regularizer sparsifies stochasticity rather than capacity. We give a certificate for universal conditional density approximation that is checkable in linear time, together with a minimal repair when it fails. We further show that the common hybrid scheme of sampling some parameters and optimizing the others is stochastic approximation for a type-II maximum a posteriori objective, and that coupled step sizes can leave a tracking error that does not vanish as the step size shrinks. On a bimodal target, the learned split stays close to an unconstrained reference across all budgets and is insensitive to the cutoff, while random masks that distribute the same prior scales across layers are worse by up to two orders of magnitude. On UCI benchmarks, our method performs on par with a fully stochastic network while keeping about half of its parameters deterministic.
cs.LG / 223 / 2610.09956
Possibilistic Radial Transport for Approximate IM Inference
Abstract
Probing the hypothesis space after seeing the data remains valid under possibilistic inferential models (IMs), provided the significance level stays fixed. The price is computation, as each plausibility is a supremum of the possibility contour over the hypothesis, and the contour itself is approximated at each queried parameter value. We propose a possibilistic radial transport, which hides the contour value of a parameter in the radius of its source point. When a transport that maximizes within-shell entropy is picked, sampling parameters covering a confidence cut becomes a matter of truncating the radius. We provide a deep learning algorithm that enforces the contour depth condition while maximizing the entropy within each shell. Our amortization makes coverage and power assessments of the learned approximation practical as well as predictive check of new datasets. We also use the sampler to construct a Bel-Pl spectrum for comparing and selecting interpretable hypotheses that satisfy a prescribed Bel-Pl decision criterion. In simulations the learned contours match or improve on ellipsoidal approximations to the cuts, while the coverage and power track the exact reference. Finally, we probe hypotheses about ovarian aging using synthetic AMH records, asking for each woman how many more years her median AMH level will remain above a specified reference value.
cs.LG / 224 / 2610.09984
Extreme Binary Classification: Extreme Value Theory for Extreme Constraint on False Negative
Abstract
While binary classification is one of the most extensively studied problems in machine learning, the regime in which the goal is to learn a classifier with an almost zero false negative rate remains largely unexplored. In this paper, we introduce the Extreme Binary Classification problem, where the objective is to learn a classifier whose false negative rate $α$ is constrained by $ε_{N_1}=o_{N_1\to\infty}(1/N_1)$, with $N_1$ denoting the number of positive examples in the training set. To address this problem, we propose a threshold adaptation method theoretically grounded in guarantees derived from Extreme Value Theory, together with a feature selection procedure based on a permutation test applied to sample maxima. Experimental results on four real-world datasets of varying sizes demonstrate that our approach compares favorably with state-of-the-art methods. In addition, we illustrate its interpretability through an application to a cancer screening dataset.
cs.LG / 225 / 2610.10021
Controlling Dependence in Implicit Generative Models via Spread Mutual Information
Abstract
Mutual information (MI) provides an objective for suppressing or encouraging statistical dependence in implicit generative models. However, direct MI evaluation is challenging in implicit models due to typically intractable densities. A remedy is estimating the generator gradient from the difference between conditional and marginal scores. This score difference can, in turn, be estimated by differentiating a log density ratio learned through classification. This construction nevertheless faces two difficulties: (i) singular distributions need not admit the required score functions, and (ii) poor overlap can hinder density-ratio estimation. We therefore introduce Spread Mutual Information (SMI), a weighted integral of MI across noise levels obtained by applying a common spreading kernel to the generated variable. Gaussian spreading yields smooth, strictly positive conditional and marginal densities, extending the gradient construction to distributions that may originally be singular. Across a variaty of experiments, SMI consistently achieves effective dependence control among MI-based methods and remains competitive with established task-specific approaches.
cs.LG / 226 / 2610.10033
Gaussian Equivalence for Multi-Head Self-Attention
Abstract
A theoretical understanding of multi-head self-attention is fundamental to the study of modern neural networks. Using random matrix theory, we establish Gaussian equivalence for multi-head self-attention: replacing softmax attention with rescaled scores plus Gaussian noise preserves the limiting spectral law of the centered output. This equivalence also covers value and output projections that depend on the keys. The resulting laws separate the effects of head allocation and projection widths, and distinguish spectrum-preserving across-head sharing from within-head key--value dependence.
cs.LG / 227 / 2610.10076
Towards Calibrated Probabilistic Forecasts for Events of Interest via Outcome-Conditional Recalibration
Abstract
Calibration is an essential requirement for probabilistic predictions to be useful for decision making. While state-of-the-art prediction methods often yield miscalibrated predictive distributions, several post-hoc recalibration schemes have been proposed to generate calibrated predictions. However, popular recalibration schemes can conceal miscalibration in specific regions of the outcome space. Since particular outcomes, such as extreme events, often matter most for decision making, probabilistic predictions should be calibrated when evaluation is restricted to these outcomes. Hence, in this paper, we introduce outcome-conditional recalibration, a post-hoc method to recalibrate probabilistic predictions on user-defined regions of the outcome space. The method is simple, easy to implement, and can be applied to arbitrary predictive distributions. It works by applying the quantile recalibration approach of Kuleshov et al. (2018) to forecast conditional distributions, before rescaling these conditional distributions so that forecast event probabilities match empirical occurrence frequencies. This produces valid and continuous predictive distributions that are calibrated within each region of interest. Across regression benchmarks, we demonstrate that existing recalibration schemes do not necessarily yield calibrated predictions when interest is on particular outcomes, and that our approach improves outcome-conditional calibration relative to existing conditional and unconditional recalibration methods, while retaining competitive calibration overall. In an application to day-ahead electricity price forecasting, the approach substantially improves calibration when predicting negative prices, at negligible cost to forecast accuracy.
cs.LG / 228 / 2610.10168
Conformal Prediction for Spatially Dependent Data via Sequential Whitening
Abstract
Split conformal prediction uses prediction errors on held-out (calibration) data to determine how wide the prediction intervals should be. It guarantees distribution-free finite-sample coverage when these errors and the error at the target site are exchangeable. This assumption may fail under spatial dependence and nonrandom sampling geometry. Existing spatial methods use fitting residuals to remove the predictable part of spatial variation from calibration and target errors. However, the spatial variation that only the calibration residuals can predict remains in both the target and calibration errors, reducing the efficiency and stability of the interval. We address this by additionally conditioning on the calibration residuals sequentially, which scales to large networks through nearest-neighbour approximations. Under a correct working covariance and an elliptical residual law, the resulting interval has exact finite-sample coverage under any spatial design, and under further conditions it is asymptotically oracle efficient. We also bound coverage loss under covariance misspecification and develop a diagnostic that identifies regions at risk of undercoverage. In simulated data, our method produces narrower and more stable intervals than global and localized state-of-the-art alternatives. In a national PM2.5 application, it produces narrower intervals within the network and identifies regions at risk of coverage failure.
cs.LG / 229 / 2610.10214
RoBART: Bayesian Additive Regression Trees with Tree-Specific Rotations
Abstract
Bayesian additive regression trees (BART) can require many splits to approximate boundaries misaligned with the predictor axes. RoBART assigns each tree a rotation shared by all internal nodes, retaining axis-aligned splits in rotated coordinates and constant leaves. We jointly propose a Givens rotation sequence and cutpoints on the resulting grid by Metropolis-Hastings and establish reversibility with respect to the conditional posterior with leaf means integrated out. For additive functions with component-specific rotations and anisotropic Hölder smoothness, we prove posterior contraction in empirical $L_2$ distance and for the noise standard deviation. Under the stated prior, design, and grid conditions, with fixed numbers of predictors, trees, and components and no more components than trees, the rate is a sum of componentwise rates determined by smoothness and the number of rotated coordinates used. We also establish a posterior contraction lower bound showing that there exist functions for which RoBART adapts to the intrinsic dimension but axis-aligned BART does not.
cs.LG / 230 / 2610.10340
Data Reuse in Non-Stationary Learning
Abstract
We consider online learning in non-stationary environments, where the goal is to track an unknown parameter that switches abruptly between a finite set of recurring values. Recurrence opens the possibility of judiciously reusing past observations to improve algorithm performance. However, the changing nature of the underlying signal and lack of information on these dynamics may limit the ability to "safely" reuse data. In this paper we quantify some of the fundamental tradeoffs in this class of problems, and show that they bear a certain resemblance to the classical bias-variance dilemma. Specifically, we propose a class of anytime algorithms, dubbed Exposure-Capped Reuse (ECR), that combine online change detection, compatibility testing, and "contamination" control. We characterize the regime in which ECR's regret scales with the number of distinct values rather than the number of changes, and derive a novel information-theoretic lower bound that establishes the near-minimax optimality of ECR. This provides rigorous quantification of the statistical "value" of data reuse.
cs.LG / 231 / 2610.10347
Dataset Pruning from First Principles: A Label-Free Linear Programming Approach
Abstract
Dataset pruning reduces a large training set to a representative subset while preserving model performance. Existing geometry-based methods typically assume that nearby points in embedding space share similar properties. Rather than imposing this assumption, we derive geometric selection criteria by reformulating unbiased subset selection as a variance minimization problem. Unbiasedness ensures that unweighted subset averages recover full-dataset averages in expectation, including losses and gradients at fixed model parameters. Specifically, we characterize a family of unbiased subset selection algorithms as a high-dimensional polytope. In this context, minimizing the expected sampling variance is a linear objective. Differences in sampling variance, averaged over rigid motions, admit closed-form pairwise expressions. Because the polytope has high dimension, directly applying standard linear programming is impractical. We instead use these expressions to construct an efficient vertex walk that optimizes an approximation of the variance objective while preserving unbiasedness, yielding a method that requires neither labels nor model training during selection. Across CIFAR-10, MNIST, and CelebA benchmarks, our method matches or exceeds uniform sampling in mean test accuracy at every evaluated budget and outperforms competing geometric methods in several settings, particularly at small selection budgets. Beyond dataset pruning, the same framework reduces stochastic-gradient variance by increasing diversity within mini-batches while keeping the batch size unchanged.
cs.LG / 232 / 2610.10393
Safe Meta-Policy Design with Risk Control
Abstract
Models can be retrained as new data arrive, but deploying every new version risks replacing a good policy with a worse one. We study how to plan policy updates (i.e., meta-policy) before future candidates are trained, balancing the benefits of improvement against the risk of performance regression. Our offline meta-policy maximizes expected cumulative value subject to a budget on the expected number of updates that perform worse than the policies they replace. We estimate the value and risk of possible switches from historical learning trajectories, represent an update schedule as a path in a directed acyclic graph, and select a schedule using dynamic programming. A leading-order analysis identifies the signal-to-noise ratio of policy improvement as a key driver of update frequency, waiting times, and risk allocation: clearer improvements support earlier, more frequent updates, while noisier improvements call for longer waits or greater risk expenditure. Their asymptotic rates also reveal a diminishing marginal cost of achieving greater safety over time. Experiments on synthetic and clinical trial data illustrate the performance--risk tradeoff and compare our method with alternative baselines.
cs.LG / 233 / 2610.10428
Derivative Gaussian Processes on a Two-Direction Budget
Abstract
Gradient observations promise more accurate Gaussian process (GP) surrogates, but the cost of incorporating them has long stood in the way of realizing that promise. We propose a derivative GP with a budget of just two directions per observed gradient. One direction focuses on each gradient's direct contribution to target prediction, while the other aggregates its indirect contributions through correlations with the conditioning function values. Within a Vecchia approximation, where each prediction conditions on $m$ nearby inputs in $d$ dimensions, this construction represents their $md$ gradient coordinates using at most $2m$ directional derivatives, giving $\mathcal{O}(m^3)$ dense factorization cost per prediction target. For general conditioning sets, we bound the posterior approximation error relative to using full gradients and characterize when the error is small or the approximation is exact. In simulations, our method matches the accuracy of a leading exact gradient-reduction method at equal conditioning set size. Because its cost grows much more slowly with that size, it can use conditioning sets well beyond the memory limit of the exact method, reaching lower prediction error with a small fraction of the time and memory. Notably, our method can exploit gradient observations while requiring less computation time or memory than function-only GP baselines.
cs.LG / 234 / 2610.10488
Best Arm Identification for Bandits with Shifting Means
Abstract
We study the best arm identification problem in a stochastic environment with a novel form of adversarial perturbations, which we coin Shifting Means. While classically the mean rewards of the $K$ arms are stable in time, in Shifting Means only the gaps $\boldsymbolΔ$ between mean rewards are stable, while their common shift may be determined adversarially in each round. The objective of the learner is to identify the best arm with high probability while minimizing sample complexity (the fixed confidence setting). Handling shifts requires new tools: we show that algorithms employing a Generalized Likelihood Ratio Test (GLRT) stopping rule, including the popular Track-and-Stop, fail under time-varying shifts. Instead, we propose Importance Weights for Shifting Means ($\mathsf{ISM}$). Assuming means bounded by $U$ and $σ^2$-sub-Gaussian rewards, we show $\mathsf{ISM}$ to be $δ$-correct and to enjoy a sample complexity bound of order $K (σ^2 + U^2) Δ_{\min}^{-2} \ln \frac{1}δ$. We also present a matching (up to constant factors) worst-case lower bound and evaluate our results empirically.
神经与进化计算 (cs.NE)
5
cs.NE / 1 / 2610.09114
A Connectome Test of the Fly Hashing Algorithm
Abstract
Dasgupta, Stevens and Navlakha (2017) showed that the Drosophila olfactory circuit, modelled as a random sparse projection followed by winner-take-all, is a locality-sensitive hash that beats classical LSH. The projection was random because the wiring was unknown. We test it against four electron-microscopy connectomes (MaleCNS, hemibrain, FlyWire, BANC; four animals, seven hemispheres). First, the 2017 pattern holds in a reimplementation of its protocol on SIFT, MNIST and odour mixtures (GloVe is near chance at short codes for every method): the fly hash beats k Gaussian projections at short hash lengths (3.1x in AP@200 on MNIST at k = 4). Second, against the tested real-valued Gaussian baseline that advantage is per active cell, not per operation: Gaussian projections given the same projection arithmetic retrieve better on every dataset and input dimension tested. Third, across four connectomes the measured pairing of glomeruli gives no consistent retrieval advantage over degree-preserving rewiring: retrieval is slightly lower (median -1.6%), and the small odour deficits depend on how missing odour responses are treated. Separately, equalising glomerular fan-out at fixed connection count improves retrieval in the model in every hemisphere, while equalising inputs per cell lowers it on average. Yet the fan-out profile is similar across the four sampled animals (median between-animal Spearman rho = 0.86), structural synapse counts do not offset it, and its relation to odour tuning is weak. In the model that uneven allocation costs retrieval. For practice: measured wiring gives no consistent retrieval advantage over degree-preserving random wiring, so a fly hash needs no connectome data, and its advantage is per active unit, which may suit hardware where active units rather than arithmetic are the binding cost, a hypothesis we do not test.
cs.NE / 2 / 2610.09347
Relocating Nonlinearity: How a Downstream Learner Reshapes What Genetic Programming Must Evolve
Abstract
Genetic programming was conceived as a way of evolving solutions: the program is the answer, and fitness is the error of its own output. A substantial line of work instead makes the program an input to a separate learner, so fitness measures the learner's output rather than the program's. Which learner to attach matters, with no account of what decides it. What makes a target hard for genetic programming is how much nonlinearity the program must build by composing primitives; whatever nonlinearity the learner supplies, the program need not. The quantity to measure is therefore how much of a target's nonlinearity a learner can take over, though on continuous benchmarks it can only be estimated. Here we show that how much the program must still build is decided by which learner is attached, and is measurable on the programs themselves. Moving to Boolean domains, where a target's nonlinearity is exactly its Fourier degree, we tune that degree from one to six with everything else fixed. Conditioned on success, a linear learner forces the evolved program to the target's degree exactly, at one through six without exception over thirty runs per setting, while tree ensembles succeed with far simpler programs and four times as often. This gives the field two things: a learner can be chosen from a target's structure instead of its reputation for difficulty, and the degree of the evolved program is a diagnostic free to compute during any run. Our control target has lower degree than the parity problems yet gains nothing from a nonlinear learner, while targets reducible to a simple statistic gain a great deal. The latter comes with a warning: evolution internalises what the learner supplies in only four of sixty-three conditions, and grows more dependent on it in forty-four, so the more capable the learner, the less of the model is legible in the program.
cs.NE / 3 / 2610.09699
The Modular CMA-ES: A Framework for Modern Evolution Strategies
Abstract
Since their introduction, modern evolution strategies, such as the Covariance Matrix Adaptation Evolution Strategy (CMA-ES), have become established as powerful methods for continuous black-box optimization. This success has led to a wide range of proposed modifications, each designed to improve performance or behavior in specific optimization scenarios. However, because these developments have largely been introduced and studied in isolation, their interactions remain comparatively underexplored. In this paper, we present the Modular CMA-ES (ModCMA), a configurable framework that integrates a wide range of mechanisms from modern evolution strategies within a single implementation. By decomposing CMA-ES into modules with interchangeable options for sampling, selection and recombination, step-size adaptation, matrix adaptation, and restarting, ModCMA enables systematic exploration of a large design space of modern evolution strategies and facilitates the construction, comparison, and automated configuration of new algorithm variants. We illustrate the benefits of this modular approach through two example studies. First, we compare several matrix-adaptation mechanisms in terms of their computational cost and optimization performance. Second, we use automated algorithm configuration to specialize ModCMA to individual benchmark problems and analyze the resulting configurations. Together, these examples demonstrate how the framework can be used both to study individual algorithmic design choices and to explore their combinations in a systematic and reproducible manner.
cs.NE / 4 / 2610.10496
Evolutionary Architecture Search for Chlorophyll-$a$ Prediction in Lakes using Sentinel-2
Abstract
Small tabular datasets with expert-designed spectral features are the norm in operational Earth observation, and the networks applied to them are typically hand-designed. We revisit one such published model -- a Sentinel-2 algal bloom classifier -- and ask what architecture search adds, holding the task, the features and the lake-level train/test split of the original study fixed. Searching an extended multilayer-perceptron space with regularized evolution, and selecting on inner-cross-validation AUC only, we find networks that improve held-out AUC from 0.790 to 0.820 and accuracy from 0.733 to 0.748 while using 409 trainable parameters, 26 times fewer than the strongest hand-designed reference. The search converges on a consistent recipe -- a single narrow layer, RMS normalisation, $\tanh$ activation, step-decayed RMSprop and weight averaging -- that a practitioner would be unlikely to reach by default. At 1.6\,kB the resulting model is small enough to serve as an onboard screening trigger, which is the setting that motivates the work. Code: https://github.com/VU-AIML/automl4eo-bloom-nas.
cs.NE / 5 / 2610.09280
Evaluating the Transfer of Co-Evolved Communication from 2D to 3D Simulation
Abstract
This work examines the transfer of a co-evolved communication mechanism between two robotic agents from a discrete two-dimensional (2D) simulator to a three-dimensional simulator with real physics (3D). The study focuses on whether a communication mechanism co-evolved in a 2D environment retains its functional role after transfer to a 3D physics-based simulator. To support this analysis, the effects of the episode time budget, the social cue, and the asymmetry between the two co-evolved roles were examined. The results indicate that the success rate increased approximately linearly with the evaluated time budgets, with no evidence of a plateau between 2,000 and 6,000 physics steps, suggesting that evaluations based on shorter episodes may underestimate the performance of the trained controllers. In both simulators, the social cue functioned primarily as a jam- assistance mechanism rather than as a navigation guide, although with a more pronounced effect in 2D. Analysis of eight independent evolutionary runs revealed a consistent direction of asymmetry, although its magnitude varied across runs. Controlling the processing order between agents allowed us to rule out an artifact of the physics engine. Finally, the results are discussed in terms of the factors that may contribute to the remaining performance gap observed after transfer.
计算语言学 (cs.CL)
36
cs.CL / 1 / 2610.08917
CARE: Certifying Acceleration for Vision-Language-Action Inference
Abstract
While vision-language-action (VLA) models have advanced rapidly, running them at every control step remains expensive. Prior work accelerates VLA inference using techniques like action chunking and visual-token pruning, typically evaluating based on latency and average task success. However, acceleration may discard information and break tasks the original policy would solve, a risk hidden by average metrics. Measuring these failures is challenging because action deviations compound over closed-loop trajectories, meaning task failure is only observable across full episodes. We therefore define an acceleration-induced failure via paired rollouts from identical initial conditions, tracking when the reference succeeds but the accelerated policy fails. To manage this, we introduce CARE, an approach for certified accelerator selection. CARE uses paired rollouts on a calibration set to provide finite-sample guarantees that acceleration-induced failure risk stays below a user-specified budget. It deploys the fastest certified candidate, falling back to the reference if none qualify. By relying only on terminal outcomes and measured compute, CARE applies unchanged across diverse acceleration mechanisms, while sequential testing and failure-triggered reference rollouts keep certification affordable. On four LIBERO suites with OpenVLA-OFT, CARE certifies $9.0$--$10.8\times$ speedups while guaranteeing (at $95\%$ confidence) that at least $85.8\%$ of reference-solved episodes are preserved. Under tight budgets, selectors without guarantees exceed the budget in up to $75\%$ of trials, whereas CARE stays within budget and its sequential form uses $78.9\%$ fewer rollouts than exhaustive evaluation. CARE further generalizes to flow-step reduction for $π_{0.5}$, and to Qwen3.5-9B and Llama-3.1-8B agents in Crafter.
cs.CL / 2 / 2610.09111
Same Text, Different Prediction: Serving-Context Nondeterminism in Text Classifiers
Abstract
Deterministic inference is essential for reliable and trustworthy machine learning. Prior studies of text generation have shown that changing factors such as batch size, batch composition, hardware, or inference engine can alter the generated text, even when the prompt, model parameters, and sampling randomness are fixed. These differences have been attributed in part to floating-point non-associativity, shape-dependent kernel selection, and other implementation-level differences in numerical execution. However, it remains unclear whether, when, and to what extent the same factors affect text classification. We present a systematic study of serving-context non-invariance in text classifiers, which prior work has measured only through generated text. We train 180 models spanning discriminative, pseudo-generative, and fully generative classifier formulations and evaluate each across four categories of serving contexts, holding the checkpoint and the text fixed. Label stability does not imply score stability. Changing only the batch shape changes no labels across fp32 comparisons, yet under bf16 it moves up to 56.7 percentage points of predicted probability mass, with label changes concentrated at small margins. Fully generative classifiers change more labels than their discriminative counterparts under the same serving changes. We derive sufficient conditions for label stability under each serving change and give a separate mitigation for each mechanism. Our results identify and quantify the serving conditions that must be fixed for reproducible text classification.
cs.CL / 3 / 2610.09163
ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation
Abstract
Task-oriented conversational agents remain fragile under real world conversation scenarios as they rarely follow a predictable script, especially when users exhibit non-cooperative behavior. Existing function-calling benchmarks often emphasize successful, cooperative interactions and underrepresent adversarial conversation trajectories, thereby limiting the training resources available for developing robust agents. We present ToolRACER, a synthetic data generation pipeline that coordinates user, assistant and tool emulation models to generate and validated multi-turn interactions between a user and an agent. Using \sysn, we construct ToolRACERBench a robust multi-turn conversation benchmark spanning six domains, ranging over 55 varied personas, generating a validated corpus of 5.6K conversation trajectories, with approximately 66\% of conversations containing failure-prone conversation scenarios. We inject adversarial behaviors, producing validated conversational interaction trajectories that capture realistic, robust scenarios. We evaluate models trained on ToolRACERBench against internal benchmarks, as well as on function calling benchmarks such as $τ^2$-bench, BFCLv3 and ACEBench to evaluate agentic accuracy and robustness. Models trained on ToolRACERBench improve end to end agentic accuracy across $τ^2$-bench and ACEBench, demonstrating significant gains when mixed with in-domain dataset in small language models for agent capability tasks.
cs.CL / 4 / 2610.09321
Dialect-Robust Speech Language Models with Synthetic Pseudo-Dialect Augmentation
Abstract
Speech Language Model (SLM) performance often degrades on dialects due to data scarcity. Conventional text-to-speech (TTS) augmentation struggles to cover diverse dialects as it requires a certain amount of real dialect speech. We propose synthesizing pseudo-dialect speech by converting LLM-generated dialect text via a standard-language TTS model, requiring zero real dialect speech. Additionally, we introduce intermediate standard-text prediction during training, acting as semantic normalization for downstream tasks. We evaluate dialect understanding via dialect-to-English speech translation across Japanese, German, and Chinese dialects. Compared to synthetic standard speech baselines, pseudo-dialect augmentation improves scores for Japanese (from 25.38 to 26.24) and German (from 31.57 to 32.47). Furthermore, the intermediate standard-text prediction effectively bridges the semantic gap, boosting performance to 28.26 for Japanese and from 11.67 to 16.37 for Chinese. These results suggest that our approach scales to various languages without requiring speech resources specific to each dialect.
cs.CL / 5 / 2610.09372
Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling
Abstract
Mixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts. Under expert parallelism (EP) the experts are distributed across GPUs, and every MoE layer runs all-to-all collectives in the forward and backward passes to dispatch tokens to their experts and then combine the results. On a cluster with 8 AMD Instinct MI300X GPUs per node, these collectives can take 45% of the training step at EP32 with top-2 routing and 60% with top-6 routing. We find that early in pretraining routers have already learned to assign tokens to experts in correlated patterns, both within a layer and across layers. At top-2, 0.8% of the expert pairs in a layer are selected together by 42% of tokens, and the experts a token selects at one layer predict the experts it selects at the next layer. We use these correlations to keep more token--expert assignments on the token's own GPU, which reduces communication across GPUs and across nodes. Correlated expert placement puts experts that are often selected together on the same GPU. Combined with a dispatcher that sends each token to each GPU once, it removes up to 58% of dispatched rows. Token shuffling applies when sequence parallelism shards tokens across the EP group. It moves each token to the GPU predicted to hold its next-layer experts during the reduce-scatter that follows attention. On one node this raises the share of token--expert assignments served on the token's GPU from 12.5% to 59%. In Megatron-LM, across EP degrees from 8 to 64 with top-2 and top-6 routing, the two methods reduce all-to-all time by 1.16-2.63X and end-to-end step time by up to 1.41X. Neither method changes the models' underlying routing decisions or expert parameters.
cs.CL / 6 / 2610.09384
The Persona Hierarchy Model: Understanding Contextual Generalization in Fine-Tuning LLMs
Abstract
Language models are routinely fine-tuned under a fixed context, such as a generic system prompt, persona or domain-specific instruction, yet the learned behavior sometimes stays confined to that context and sometimes broadly generalizes to unseen contexts. We propose the Persona Hierarchy Model to explain this: a shared default persona influences behavior across contexts. Under this model, fine-tuning that modifies the shared persona promotes broader transfer, whereas changes to local personas remain more context-specific. Across 120 fine-tuned models spanning four behaviors and 15 training contexts, generalization narrowness positively correlates with the similarity between the training context's persona and the default persona (Pearson's r = 0.72 for Qwen3-4B). Prior fine-tuning under the default context can broaden generalization in subsequent training under other contexts. Aligning contextual responses with default-persona responses produces stronger effects. Finally, we propose persona-preserving regularization (PPR) to confine undesired contextual generalization. In RL, PPR cuts reward hacking from 42-55% to at most 0.2% under every evaluated prompt while retaining accuracy gains. These results support the Persona Hierarchy Model as an explanation for contextual generalization and can motivate future controls on unintended generalization for better alignment of LLMs.
cs.CL / 7 / 2610.09396
ARCS: Towards Precise Text-to-SQL via Structured Disambiguation
Abstract
As text-to-SQL systems move beyond demonstrations toward real-world deployment, ambiguity in user questions becomes a primary source of errors. Such ambiguities are often subtle, domain- or data-specific, and can silently cause system outputs to deviate from the user's true intent. Ambiguity is traditionally addressed through conversational clarification, which is often inefficient, cognitively demanding, and poorly aligned with real-world user workflows. We propose structured disambiguation, a new paradigm in which ambiguity is resolved through explicit, constrained interactions rather than free-form dialogue. We construct ARCS (Ambiguity Resolution Corpus for SQL), the first text-to-SQL benchmark featuring naturally occurring, unconstrained ambiguities over real-world databases, with complete annotations of all valid ambiguity points, interpretations, and SQL queries. Experimental results show that text-to-SQL remains challenging in the presence of ambiguity: gpt-6-sol achieves only 51% end-to-end execution accuracy, and no open-source model exceeds 27%.
cs.CL / 8 / 2610.09458
Right Number, Wrong State? Measuring Cross-Jurisdiction Substitution in LLM Recall of State Policy
Abstract
When an LLM answers a state-specific policy question wrongly, it may be hallucinating, or it may be returning a real value that holds in another state. We test this with a minimal-set design: the question wording is fixed and only the jurisdiction varies, across the 50 U.S. states and the District of Columbia (51 jurisdictions) and three exactly defined Medicaid income-eligibility quantities. Gold values come from an official data book and agree with an independent source in 101 of 102 checked cells. Under a pre-registered protocol, Claude Sonnet 5.5 and GPT-5.6 Sol reproducibly give another state's current value, identical across two independent repeats, for 10 and 25 of 153 items. Attribution is fragile, however. Crediting any wrong answer that equals another state's value yields 3-5x more reproducible substitutions than checking every number in the asked state's own records, because many apparent cross-state answers are the asked state's own values under another convention or from an earlier year. Claims about cross-jurisdiction error need a complete same-state reference set. We will release the protocol, gold table, and all model outputs.
cs.CL / 9 / 2610.09464
BanglaRhet: Benchmarking Classical and Transformer Models for Rhetorical and Persuasion Detection in Bangla Political Speech
Abstract
Political discourse often uses rhetorical and persuasive language to frame narratives, influence public opinion, and mobilize audiences. While Bangla natural language processing has made progress in sentiment analysis and opinion mining, systematic benchmarking of transformer models for fine-grained rhetorical and persuasion technique detection in Bangla political speech remains largely underexplored. This paper presents a benchmark study of transformer-based models for detecting rhetorical form and persuasive intent in Bangla political discourse. Using BanglaRhet, a manually annotated corpus of 30,289 Bangla political speech segments collected from publicly available political news sources, we formulate two supervised single-label classification tasks: rhetorical technique detection (contrast, repetition, exaggeration, metaphor, rhetorical questions) and persuasion technique detection (blame assignment, call to action, unity call, moral, emotional, and logical appeals). We evaluate four transformer-based models, BanglaBERT, BanglaBERT-Base, SahajBERT, and XLM-RoBERTa-Base, against classical TF-IDF baselines. BanglaBERT achieves the highest performance, with 65.40% macro-F1 for rhetorical technique detection and 66.46% for persuasion technique detection, outperforming the best tuned classical baseline by 19.2 and 13.8 macro-F1 points, respectively. Class-level analysis indicates that errors are mainly associated with semantic overlap among labels, figurative language, and class imbalance. The results provide initial benchmark baselines for Bangla rhetorical and persuasion-aware political discourse analysis and highlight the need for context-aware and multi-label modeling.
cs.CL / 10 / 2610.09529
A Comparative Study of Evaluation Metrics for Long-Document Financial Narrative Summarization with Transformers
Abstract
There are more than 2,000 listed companies on the UK's London Stock Exchange, divided into 11 sectors who are required to communicate their financial results at least twice in a single financial year. UK annual reports are very lengthy documents with around 80 pages on average. In this study, we aim to benchmark a variety of summarisation methods on a set of different pre-trained transformers with different extraction techniques. In addition, we considered multiple evaluation metrics in order to investigate their differing behaviour and applicability on a dataset from the Financial Narrative Summarisation (FNS 2020) shared task, which is composed of annual reports published by firms listed on the London Stock Exchange and their corresponding summaries. We hypothesise that some evaluation metrics do not reflect true summarisation ability and propose a novel BRUGEscore metric, as the harmonic mean of ROUGE-2 and BERTscore. Finally, we perform a statistical significance test on our results to verify whether they are statistically robust, alongside an adversarial analysis task with three different corruption methods.
cs.CL / 11 / 2610.09607
Which Language Should a Skeleton Speak? Language Choices in Multilingual Reasoning
Abstract
Skeleton-based reasoning prompting is a promising training-free approach for structuring LLM reasoning, but prior work largely assumes an English-centric setting. We propose the Language-Aware Skeleton Exploration Framework (LASEF) to study skeleton-language choice in multilingual mathematical reasoning. Across math benchmarks, model scales, and languages, we show that English skeletons yield a small positive tendency on average, most visible for smaller models and low-resource languages. However, few language-level gains remain significant after correction, and English is not universally optimal. Combining greedy decoding, multi-rollout evaluation, translation ablation, and cross-benchmark validation, we further find three patterns of skeleton-language effects: directionally consistent, evaluation- and benchmark-dependent, and asymmetric negative. These effects cannot be fully explained by generation quality alone. Overall, skeleton language is a context-dependent design variable that requires multi-level exploration. All resources are released at https://github.com/lhsstn/LASEF.
cs.CL / 12 / 2610.09639
On-Policy Distillation Teaches New Skills but Not New Knowledge
Abstract
On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown. We separate these capabilities using a controlled synthetic framework that measures the student's initial capabilities and independently controls the teacher's additional facts, compositional skill, or both. Across four models from three families, reverse-KL OPD reliably transfers compositional skill across unseen reasoning structures, but transfers minimal factual knowledge. Decoupling the distillation recipe reveals the source of this asymmetry: replacing reverse KL with forward KL restores factual transfer, whereas student rollouts specifically improve the execution of multi-step reasoning. Experiments on recent factual QA and competition mathematics show a similar asymmetry under reverse-KL OPD, yielding notable reasoning gains without factual memory expansion. Together, these results demonstrate that on-policy distillation does not expand a model's parametric knowledge, but instead teaches it to organize and compose the knowledge it already possesses.
cs.CL / 13 / 2610.09660
Rubric Spans are Label Representations: Joint LLM Encoding for Short Answer Scoring
Abstract
Automatic Short Answer Scoring (ASAS) requires models that can score student responses against question-specific criteria while remaining efficient and transferable across rubric sets. We propose RUSPAN, a rubric-conditioned ASAS framework that treats rubric descriptions as semantic label representations. RUSPAN serialises the question context, student answer, and all candidate rubric levels into a single sequence, then scores the levels listwise from the rubric-span and whole-sequence representations produced in a single LM pass. We further introduce RUSPAN-RIM, in which a Rubric-Independent Mask prevents rubric spans from attending to one another, making rubric representations depend only on the answer and question context and preventing overfitting to rubric patterns during training for zero-shot transfer. On six ASAS benchmarks spanning English, German, and Portuguese, RUSPAN improves mono-benchmark scoring over discriminative and generative baselines, while RIM with position reindexing delivers consistent and substantial gains on PT-ASAG, the held-out benchmark with the strongest combined language and rubric-structure shift.
cs.CL / 14 / 2610.09661
Alice: A Large-Scale German Benchmark for Rubric-Based Multi-Dimensional Automatic Short Answer Scoring
Abstract
Automatic Short Answer Scoring (ASAS) is central to NLP for Education. However, openly available benchmarks remain scarce, and existing datasets largely address how well students answer a question directly rather than how well they master underlying concepts (knowledge elements) such as thermal energy or epistemic activities (skills) such as reasoning or claim. To address this gap, we introduce Alice, a large-scale, rubric-based German ASAS dataset that is pedagogically aligned and comprises three subtasks: (i) learning performance (Alice-LP), (ii) knowledge elements (Alice-KE), and (iii) skills (Alice-SK). We further formulate rubric-based ASAS as a rubric-retrieval task and benchmark the dataset with a range of language models, from encoder-only models to lightweight LLMs. We also benchmark the dataset with zero-shot prompting via LLMs and a standard classification baseline. The experiments show that LLMs, in particular, struggle to score knowledge elements and skills in the zero-shot setting. They also indicate that rubric text is often useful, especially for Alice-KE and Alice-SK, while on Alice-LP gains over sample-solution-focused inputs are more modest and vary by model and input format.
cs.CL / 15 / 2610.09710
SpikingVLA: Asynchronous Spiking Vision-Language-Action Models
Abstract
ANN-to-SNN conversion offers a practical route toward energy-efficient spiking Vision-Language-Action (VLA) models by bypassing the substantial cost of training large-scale SNNs from scratch. However, existing methods often require many timesteps to maintain competitive performance, resulting in substantial inference latency for real-time VLA deployment. To address this challenge, we introduce SpikingVLA, an ANN-to-SNN conversion framework that enables accurate and low-latency spiking VLA inference. Specifically, we propose a Dendritic Integrate-and-Fire (DIF) neuron that alleviates channel-wise activation outliers through dendritic mixing and adaptive somatic firing, enabling accurate ANN-to-SNN conversion with fewer timesteps. Building on DIF neurons, we further introduce an asynchronous execution mechanism that overlaps temporal computation across VLA components, reducing synchronization overhead and latency. Extensive experiments demonstrate that SpikingVLA achieves competitive navigation performance with substantially improved inference efficiency. Compared with existing spiking VLA methods, SpikingVLA improves SR and SPL by 11.9\% and 12.6\%, respectively, while reducing first-action latency by 11.2$\times$. These results establish SpikingVLA as a practical framework for deploying pretrained VLA models with high-performance and low-latency spiking inference.
cs.CL / 16 / 2610.09733
Bridge Routing Heads: Where Multilingual Multi-hop Reasoning Lives in LLMs
Abstract
Multilingual LLMs answer the same multi-hop reasoning question across languages, but we lack a mechanistic account of whether they share an internal circuit. We identify Bridge Routing Heads (BRH) in two large multilingual LLMs through a three-stage pipeline. The resulting language-specific head sets exhibit near-complete mutual exclusivity across the five languages, with a mean Jaccard similarity of only 0.017 for Llama 3.1 70B and 0.057 for Qwen 2.5 72B, revealing language-idiosyncratic circuits. Ablating general BRH increases two-hop Negative Log-Likelihood (NLL) by 39-89x the random-head baseline, providing direct causal evidence of their role. Amplifying these heads in a failing target-language pass rescues up to 51.7% of cross-lingual failures, with no training. The two models share this dual-circuit pattern but allocate heads differently: Llama concentrates chaining in a large general pool, while Qwen leans on larger language-specific pools. Together these results show that activation-level intervention alone can recover correct answers from cross-lingual reasoning failures.
cs.CL / 17 / 2610.09756
Shaer: Controlled Arabic Poetry Generation with Meter Subform and Semantic Conditioning
Abstract
Classical Arabic poetry generation requires simultaneously satisfying semantic, linguistic, and fine-grained prosodic constraints. Existing systems typically control broad poetic attributes but do not jointly model semantic intent, meter subform, and poem length. We present Shaer, a controllable Classical Arabic poetry generation framework jointly conditioned on natural-language descriptions, meter subforms, and target hemistich counts. To support this task, we construct an enriched corpus of 116,032 classical Arabic poems derived from Ashaar, containing normalized meter-subform labels and automatically generated, validated semantic descriptions. We then adapt Yehia-7B using QLoRA-based supervised fine-tuning with a completion-only objective. Our evaluation combines automatic assessment of base-meter conformity, requested-subform adherence, and length control with three LLM judges, blinded human evaluation, and memorization analysis. Shaer achieves 95.17% base-meter accuracy, 91.75% poem-level meter-subform accuracy, and 83.40% exact count accuracy. Relative to its untuned foundation model, these results represent gains of 68.68, 57.77, and 38.93 percentage points, respectively; Shaer also attains the highest base-meter accuracy among all evaluated systems. Multi-LLM evaluation and a blinded human assessment of top-ranked outputs further indicate competitive semantic and literary quality. Finally, analysis of all 3,481 test generations finds no exact copies from the training corpus or paired source poems. Code, models, and datasets are publicly available.
cs.CL / 18 / 2610.09772
Decoupling Logic from Persona: Structural Immunity of Edge LLM Agents to Context Pollution
Abstract
Small language-model agents on edge devices must hold a persona and reason correctly at once, inside one context window that fills with conversational history and persona instructions. We study what happens to the logical part of such an agent when that history is long, misleading and persona-heavy (persona-logic interference), and present a Decoupling Architecture (AO-DA) that separates logical inference ("What") from persona expression ("How") into two inference paths on one INT4 base model with hot-swappable LoRA adapters. The logic path receives only the core turn and emits a verifiable structured state (Micro-State); the persona path renders it in character with the full history. In same-base-model ablations on an Apple M2 laptop (Llama-3.1-8B-Instruct and Gemma-3-4B-it, 4-bit; 480 runs over 4 pollution levels x 3 arms x 2 tasks x 2 personas x 5 seeds) we find: (i) the decoupled logic path is structurally invariant to pollution: its prompt stays at 180 (Llama) or 167 (Gemma) tokens while the mixed single-pass prompt grows from 242 to 1,203, and its outputs are byte-identical across levels (40/40); (ii) the mixed single pass degrades monotonically (composite logic score 0.669 to 0.150 on Llama, 0.487 to 0.150 on Gemma), mostly by failing to emit the required structured output (80-95% of runs on Llama, 100% on Gemma at the two highest levels); (iii) with the same pollution fed into the decoupled logic path, the dedicated-adapter, dedicated-format path is still more robust than the single pass on the 8B model (failure 0-20% vs 80-95%; paired $Δ$ +0.30 to +0.50, Cliff's $δ$ 0.50-0.85, Holm-adjusted $p \le 0.03$) but not on the 4B model, where both collapse. Separation costs one extra decode on a topic's first turn (28.2 s vs 18.2 s on Llama) and buys persona hot-swapping in 1.7 ms without re-running the logic path. Code, rubric, fixtures, adapters and logs are released.
cs.CL / 19 / 2610.09788
Judging in Latent Space: Efficient Generative Reward Modeling via Semantics-Preserving Compression
Abstract
Reward modeling often requires jointly representing and reasoning over multiple evaluation criteria, yet verbalizing this process token by token can incur substantial inference cost. Recent work on latent reasoning suggests that continuous states may support this computation more compactly. We introduce LatentGRM, a latent evaluation framework built on semantic chunking, compression, and reconstruction. By using the structure of rubric-guided evaluations to guide compression, LatentGRM learns compact continuous trajectories that support autonomous pairwise judgments without generating textual assessments. A separate interpreter reconstructs evaluation text from these trajectories, providing an offline view of the information retained under compression. Under matched training data and backbones, LatentGRM achieves competitive aggregate preference accuracy relative to explicit Supervised Fine-Tuning (SFT) judges at both 4B and 8B scales. Across four benchmark domains, LatentGRM-8B compresses evaluation trajectories by 8.9--9.2x and reduces total judge inference time by 6.1--7.0x at vote@5. Controlled rubric interventions show that criterion-dependent preference information is carried through the latent sequence. Together, these results demonstrate that continuous latent evaluation can substantially reduce inference cost while preserving competitive judgment quality.
cs.CL / 20 / 2610.09795
MIRROR: From Imitation to Internalization in LLM Personalization
Abstract
The demand for personalized LLMs is shifting from style imitation toward content quality. We investigate whether self-distillation can bridge this gap in existing fine-tuning paradigm. To address this limitation, we introduce MIRROR(Meta- personalization by Internalizing Reference-Revealed On-policy Reflections), a novel self-distillation framework that shifts LLM personalization from imitation toward preference internalization. First, we replace reference-token imitation with reference-revealed on-policy self-distillation, aligning the model's next-token distributions along its own generation trajectories with those of its reference-conditioned self, thereby internalizing user preferences rather than reproducing reference wording.Second, we introduce MIRROR-F, a focal plug-in that augments on-policy distributional alignment with selective supervision over informative reference tokens, thereby strengthening content generation while preserving user-specific expression. Across three personalized generation benchmarks, two model scales, and complementary reference-based and LLM-based evaluations, MIRROR and MIRROR-F achieve leading overall personalization performance and superior text quality, while exhibiting less catastrophic forgetting than SFT-based baselines on three unseen personalized generation tasks. The gains are consistent across model scales and application scenarios, translating to improved performance in LLM personalization tasks.
cs.CL / 21 / 2610.09934
Itgan at NADI 2026 shared task: Parameter-Efficient Whisper Adaptation for Robust, Mixed-Dialect and Code-Switched Arabic ASR
Abstract
We describe the Itgan systems for the three ASR subtasks of NADI 2026, namely robust country-level ASR (1.1), mixed-dialect ASR (1.2), and Tunisian code-switched ASR (1.3). All three share one recipe, Whisper adapted with LoRA on consumer GPUs, and each was carried by a different addition to it. On 1.1, where the dialect label is given at test time, per-dialect specialists continued from a pooled adapter gave the largest gain, and the submitted system reached 57.1% country-average WER. A post-evaluation linear probe on frozen encoder features routes utterances without the label and recovers 44% of what oracle routing gives. On 1.2 the choice of base model mattered more than adapter capacity, and system combination helped only once we added a decorrelated member, reaching 46.7% WER. On 1.3 our system placed second at 14.49% WER with the lowest CER among the leading submissions, 5.38%. Its last 0.60 WER points came without further training, mostly from an exact weight-space average of independently trained runs, with ROVER voting adding the remainder. Every comparison carries a paired-bootstrap test, and we report eight directions that did not work.
cs.CL / 22 / 2610.10058
Cache the Encoder Within:Compact, Reusable Memory across LLM Queries
Abstract
Repeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs. Building on CoMem's intermediate-state interface, EncBank treats a pretrained LLM's lower layers as a reusable document encoder and compactly stores their outputs for an adapted upper-layer reader. A self-distilled suffix adapter is shared across storage precisions within each backbone, without quantization-specific retraining. Across five benchmark suites on three Qwen backbones spanning different sizes and full-attention and hybrid architectures, 4-bit storage keeps each reported benchmark aggregate within one score point of native-precision EncBank. In a fixed Qwen3-8B workload, it retains 28.1% of the native-precision persistent GPU store. Separate native-precision controls yield a 1.40x selected-pack prefill speedup over same-evidence, same-adapter text replay, at a 3.12-point RULER accuracy cost. A native-precision Qwen3.8-27B configuration also passes 70 of 89 Terminal-Bench 2.1 tasks. EncBank thus combines reusable computation with compact memory, while task fidelity and end-to-end benefits remain dependent on the workload, preparation costs, and reuse frequency.
cs.CL / 23 / 2610.10091
ExperienceIndex: Artifact-Grounded Memory
Abstract
Knowledge-intensive tasks require answering many questions by reasoning about a shared corpus of artifacts (e.g., court cases, or scientific literature). As humans interact with these corpora, they naturally accumulate experiential knowledge about artifacts, enabling them to quickly identify the complete set of relevant artifacts for each new task. However, existing AI agents lack appropriate memory solutions to build or reuse such artifact-grounded experience, leading to lower answer quality and higher online cost. Existing memory solutions extract and reuse information from prior task-solving traces, but they primarily focus on user preferences, factual attributes, or abstract reasoning patterns rather than persistent artifact-specific knowledge. We introduce ExperienceIndex, a novel experience layer for AI agents that captures and reuses knowledge about artifacts based on prior reasoning traces. ExperienceIndex stores two complementary forms of experience: (i) single-artifact experiences that summarize an artifact's contribution to prior tasks and (ii) artifact-pair experiences that encode structural relationships discovered during past reasoning. Integrated as lightweight middleware, ExperienceIndex uses an experience retrieval mechanism to guide agents toward the complete set of relevant artifacts for new tasks, improving both answer quality and efficiency. Across diverse corpora and agentic solutions with different search frameworks, ExperienceIndex delivers consistent gains, raising answer quality by up to 11.0 points and reducing online dollar cost by up to 50.5%. We further demonstrate two benefits: (i) cross-task generalization, where experiences accumulated from text-to-SQL tasks transfer to factoid QA tasks over the same artifact corpus, and (ii) teacher-student learning, where experiences from a stronger model enable a weaker model to reach comparable performance.
cs.CL / 24 / 2610.10092
I would rather quit NLP than read another paper like this: The rise of antithesis in NLP papers
Abstract
For better or worse, LLMs are by now used routinely for scientific writing.\footnote{This paper is no exception; we did use AI to assist with writing some of the sections (see Acknowledgments).} Many have noticed that recent models fill papers with unnecessary antithesis, stating over and over what the work does not do, in ways that do not contribute to its precision or quality of expression and annoy reviewers \emph{rather than impressing them}. We study the construction \emph{rather than} in ACL papers from 2019, ACL-style arXiv papers from 2026, and papers written by GPT models from the same titles and abstracts. Its rate in 2026 is seven times the 2019 rate, and higher still in the GPT papers. Two annotators, blind to the source, find almost no 2019 use \emph{annoying} and about one in ten 2026 uses; they seldom agree on which, yet about half of 2026 papers contain a use that annoys each of them. \emph{Annoying} uses present the rejected alternative less favorably than legitimate uses. Raters of preference data and open reward models favor the construction, and an instruction to be honest promotes it. We conjecture that it is a side effect of post-training on pairwise preferences, which credit a disavowal in a single response and cannot register its cost across a text.
cs.CL / 25 / 2610.10138
LLM4Impact: Integrating Heterogeneous Information for Scientific Impact Prediction
Abstract
Predicting the future impact of a newly published paper is challenging because it must be inferred from heterogeneous evidence available at publication time. Existing approaches often rely on a single source of information or combine multiple sources without accounting for their different predictive roles. In this paper, we present LLM4Impact, an evidence-aware method for scientific impact prediction that learns to represent, integrate, and calibrate heterogeneous information. LLM4Impact combines semantic, graph, LLM, and temporal representations, and injects graph information into a frozen LLM through continuous prefix tokens. A context aware gating mechanism adaptively weights different evidence, while a separate calibration module accounts for domain and temporal variation in citation scales. We further construct a large-scale benchmark dataset with 2 million papers, leakage-safe point-in-time heterogeneous ego graphs, temporal splits, and both year-level and month-level citation targets. Experiments show that LLM4Impact consistently outperforms strong semantic, graph, and LLM based baselines, with a 10.13% reduction in year RMSE on the in distribution test set and a 6.87% reduction under out-of-domain distribution. Our results reveal that the value of such evidence is context dependent: different papers benefit from different sources, while domain and publication time affect how evidence translates into citations. This finding motivates adaptive evidence selection and context-conditioned calibration rather than simply richer representations. We will release our code, benchmark, and an interactive web demonstration upon publication.
cs.CL / 26 / 2610.10145
InterView-C: A Synchronized Multimodal Corpus of VR Avatar-Mediated Survey Interviews
Abstract
We present InterView-C, a German multimodal corpus of 27 survey interviews conducted entirely in virtual reality, with both interlocutors represented by avatars. The corpus aligns spoken interaction with synchronized behavioral data, including gaze, head and body movement, facial behavior, hand and finger tracking. Its reference transcripts and linguistic annotations provide a reliable interface between this multimodal spoken interaction and predominantly text-based NLP methods. This interface is important because automatically transcribing speech can distort linguistically relevant information, while downstream models trained on existing resources may additionally face transfer challenges when applied to transcribed spoken data. InterView-C therefore provides word-timed and manually post-edited verbatim transcripts for all 54 recordings, interview-item timings, questionnaire responses and negation cue and scope annotations for 1,422 sentences, 1,398 of them doubly annotated (α=0.87 for cues; α=0.81 for scopes). We demonstrate both challenges empirically: nine open-weight ASR systems disproportionately misrecognize short closed answers and number words, while negation models trained on existing corpora show lower and highly variable performance on our transcribed interviews than a model trained on the InterView-C annotations. InterView-C thus enables linguistic analyses of spoken interaction while retaining their alignment with rich multimodal behavior.
cs.CL / 27 / 2610.10154
HySPE: Positional Encoding via Symplectic Dual Shears
Abstract
We introduce Hyperbolic Symplectic Positional Encoding (HySPE), grounding positional attention in non-compact symplectic transformations. While canonical Rotary Position Embedding (RoPE) parameterizes the compact, elliptic branch of $\Sp(2,\R)$ via rotations, HySPE operationalizes its hyperbolic branch via a damped symmetric composition of dual shears, yielding a conformally symplectic contraction with two spectral decay rates per channel pair. To eliminate the exponential representation drift inherent to naive absolute factorizations, we diagonalize the operator in its invariant eigenbasis and introduce blockwise coordinate rebasing with adaptive centered execution. This guarantees length-independent numerical bounds while matching cached RoPE forward latency (7.21\,ms on an RTX 4090). On TinyShakespeare, HySPE-UltraLong maintains an invariant perplexity of 4.810 up to $16\times$ zero-shot extrapolation ($L=4096$), whereas RoPE degrades to 131.198. Scaled to a 51M-parameter subword Transformer on WikiText-103 ($L_{\text{train}}=512$), HySPE closely matches RoPE in-domain while robustly extrapolating to length 8192, reducing tail perplexity by 83.9\% over RoPE. While these controlled experiments establish HySPE's extrapolation robustness and numerical stability, evaluating its scaling behavior on large-scale foundation models remains an important direction for future investigation.
cs.CL / 28 / 2610.10332
Learning to Act with Task Progress: Distilling Small Agents from Compact Teacher Supervision
Abstract
Learning from large-model demonstrations offers a way to train small agents that can complete recurring tasks without calling a large model at every step. A central design choice is what to retain from teacher trajectories that contain reasoning, actions, and information about task progress. We introduce Task-Progress Distillation (TPD), an offline approach that pairs each demonstrated action with a short label describing the current task stage. The student learns these compact targets and selects actions by jointly scoring admissible stage--action pairs, which a deterministic harness executes in the environment. On ALFWorld, a 1.7B student trained with 404 demonstrations achieves 72.4\% mean unseen task success with either TPD or action-only supervision, compared with 48.3\% for a reasoning-trained student using constrained action selection. Explicit stages provide an additional benefit at 200 demonstrations, improving success from 48.0\% to 67.7\% over action-only supervision. With more demonstrations, the action-only student closes the gap, and both approaches reach 76.9\% at 808 demonstrations. Shared-history analyses link part of TPD's local advantage to better decisions when moving between subgoals, particularly from object acquisition to processing. These results show that compact supervision can train effective small task agents, while explicit task progress provides additional guidance at an intermediate demonstration budget.
cs.CL / 29 / 2610.10426
CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution
Abstract
Terminal-agent capability depends jointly on model weights and the runtime harness that formats prompts, binds tools, and handles error recovery. Existing harness-model co-evolution approaches improve both components, yet often treat trajectories produced during harness search as an undifferentiated replay buffer. This practice overlooks that a trajectory's value for model training depends on the harness under which it was generated. To systematically analyze this interface, we establish an alternating co-evolution framework that decouples harness search and policy training through component-wise promotion decisions. Within this framework, we introduce CoTrace, a harness-aware data recipe that explicitly governs trajectory routing, provenance matching, and curriculum refresh. Under CoTrace, recurring execution failures guide harness synthesis, while policy training is strictly conditioned on verified rollouts matched to the adopted runtime for supervised fine-tuning (SFT) or fresh online interactions for reinforcement learning (RL). On the Tmax promotion split, CoTrace advances Qwen3.5-9B from 78 to 88 solved tasks under supervised fine-tuning while an online reinforcement variant reaches 90. Specifically, a compact harness-matched corpus produces steady model gains at substantially lower compute than much larger corpora pooled across sibling harnesses. Furthermore, evaluations on Terminal-Bench 2.1 and SWE-bench Lite show that out-of-distribution transfer depends fundamentally on harness compatibility, where maintaining consistency between training and evaluation runtimes prevents procedural execution breakdowns observed under foreign scaffolds.
cs.CL / 30 / 2610.10508
Your Prompt Should Do More: Effects of Retrieval Instructions in Embedding Models
Abstract
Prompted embedding models have recently received increasing attention, particularly for retrieval, where detailed retrieval instructions are provided as part of the retrieval prompt. Several new datasets and studies have examined this setting, showing that the current embedding models often struggle to follow such instructions reliably. In this paper, we study the mechanism of how instructions actually affect the representations of retrieval queries in asymmetric retrieval tasks. We show that models can fail to follow even simple task instructions when query-side distractors are included in the evaluation. We hypothesize that this behavior is driven by the training setup of current embedding models and their evaluation, and show that fine-tuning with added query-side distractors leads to substantial improvements, with minimal effect on other tasks.
cs.CL / 31 / 2610.09227
Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders
Abstract
Mixed-precision post-training quantization needs a per-module sensitivity signal; for a text embedder the obvious one -- the retrieval quality a module costs when quantized -- needs relevance labels that deployments rarely have. We measure a label-free substitute: quantization-induced representation drift, obtained by quantizing one module, re-encoding the corpus, and recording how far the output embeddings moved from their full-precision positions. What is specific is the observable: the deployed output representation a dense retriever ranks with. Across five development embedders, configuration-level drift orders sampled mixed-precision plans against held-out retrieval quality at a macro Spearman of 0.911, the sensitivity transports across calibration corpora and retrieval domains in the usable regime, module drifts compose rank-consistently but not numerically, and relevance-derived sensitivity adds no consistent value. The method is one additive allocation under a hard packed-byte budget, with no labels and no search. On three embedders held untouched until method, baselines and hypotheses were frozen and sealed, the pre-registered directional hypothesis against the prior LieQ criterion holds (3/3 at the main budget, no collapse) and drift scores above a two-sided LieQ steelman in 2/3; but at the main budget drift is numerically lower than same-budget uniform precision on all three (-0.99, -0.85, -1.01 points), having reduced module and whole-model drift as designed. Output drift is thus a robust coarse sensitivity signal, not a universally optimal allocation objective: it avoids the catastrophic failures of the transferred signed-geometry adaptation and can remain usable at stressed budgets where uniform collapses, but fine-grained redistribution around a strong uniform operating point remains unresolved.
cs.CL / 32 / 2610.09412
Finding the Right Balance: Relevance and Diversity in LLM Retrieval
Abstract
Retrieval diversification is widely available in retrieval-augmented generation (RAG) frameworks, yet prior studies disagree on whether it improves retrieval and answer quality. We show that its effectiveness varies primarily with candidate-pool redundancy, in a pattern consistent with the number of distinct evidence pieces a query requires. Using controlled near-duplicate injection and production-style overlapping chunking, we find that diversification harms relevance, evidence coverage and answer quality on clean pools, but becomes beneficial on multi-evidence tasks when redundancy causes nearest-neighbor retrieval to select repeated passages. We therefore introduce a query-adaptive rule that diversifies only when the effective number of distinct documents in the nearest-neighbor top-$k$ selection falls below the query's evidence requirement. Computed from existing embeddings, the rule captures most of the achievable gain, transfers across datasets and encoders and automatically reduces to nearest-neighbor retrieval for single-evidence queries. We also introduce RNG-Score, a geometric reranker with an exact nearest-neighbor fallback whose margin indicates duplicate structure. Overall, we conclude that diversification should be used selectively, based on observable redundancy and evidence requirements.
cs.CL / 33 / 2610.09920
Inverting Multi-Vector Visual Document Indices
Abstract
Prevailing multi-vector visual document retrievers store each page as about a thousand patch vectors, often in vector databases run by a third party. Since no one can read a page from its vectors, this index is easily treated as less sensitive than the page. However, because the index keeps one vector per patch in raster order, and each vector is computed by a vision-language model pre-trained to read documents, we hypothesize that whoever runs or breaches the store can reproduce a page from its index alone. We frame inversion as conditional document image generation and infer from the vectors what the attack needs: the encoder, the page shape and, for shuffled vectors, their order. On the ViDoRe v3 benchmark, pages inverted from raw indices recover 47% of the words and 45% of the sensitive tokens. Used as queries against the stored indices, they rank their source page first 98.4% of the time. We test two cheap protections, token pooling and shuffling, which both cut word recall to about 8%. A model that restores the order of a shuffled index raises the share of source pages ranked first from 3.8% to 93.5%, while inverting a pooled index remains open. To test generalisation, we apply the same attack unchanged to another multi-vector retriever: its inverted pages still rank their source page first 70.2% of the time, though its word recall stays below a nearest-neighbour baseline. Multi-vector visual document retrievers are therefore vulnerable to inversion through their stored index, which should be protected like the documents it encodes.
cs.CL / 34 / 2610.09467
Boundary-Free Contextual Biasing: Depth-Adaptive Gating and Reading-Space Matching for Unsegmented Languages
Abstract
Contextual biasing supplies an ASR system with a list of expected words at inference time, but existing methods rely on word boundaries that Japanese and Chinese do not provide. We present a boundary-free biasing decoder for frozen public CTC models, built on a character-level Aho-Corasick automaton, with no training and no second pass. Two evidence-based mechanisms replace the boundary: a depth-adaptive gate that sets how hard to push from match depth, and reading-space matching for when the audio is right but the characters are wrong. On Aishell-1 NE's hard R1 subset we reach 66.5% recall, above the trained CLAS baseline (64%), transferring to WenetSpeech and to a second architecture without retuning. We release the first open Japanese contextual-biasing benchmark, where biasing lifts rare-word recall by 25 points at precision above 97%, and still by 19 and 22 points against 1,000-word lists.
cs.CL / 35 / 2610.09486
Mitigating Accent-Language Confusion in Self-Supervised Speech Representations for Language Identification
Abstract
Spoken language identification (LID) aims to recognize the target language regardless of accent. In practice, however, LID models fine-tuned from self-supervised speech representations frequently confuse accents with languages, misclassifying non-native (L2) speech as the speaker's first language (L1). We show that non-native speech representations lie between native target-language and native L1 poles, causing systematic misclassification. To address this, we introduce a geometric projection that estimates an L1-bias direction solely from native speech and removes it before the frozen LID head. Across five MMS-LID models and non-native corpora, this projection substantially improves target language identification for L2-accented speech while preserving predictions for native speech. These results show that accent-induced L1 bias can be corrected directly within the representation space without L2 training data or model adaptation.
cs.CL / 36 / 2610.09541
Certified by Abstention: Distribution-Free Guarantees for Chain-of-Thought Verifiers at Small Calibration Budgets
Abstract
Signals that predict whether a chain-of-thought (CoT) trace is correct are compared by AUC, but deploying one requires a threshold with a guarantee. We ask what distribution-free selective guarantees deliver for CoT verifiers at realistic calibration budgets of tens to a few hundred labelled problems, using seven open models, five verifier signals and 37,000 graded traces. The central observation is validity by abstention: an $(α,δ)$-valid procedure that issues a certificate with probability $P_{\rm fire}$ bounds the failure probability of an issued certificate only by $δ/P_{\rm fire}$, so a certificate that rarely fires can be valid and wrong every time it is used. In a simulation with known risk the standard certificate fails in at most 0.3% of calibration draws but in up to 69% of those in which it fires. A certification floor and a lattice condition for Benjamini-Hochberg conformal selection explain why certificates abstain at these budgets, and the data bear them out: the standard certificate returns nothing or a large accepted set, and an unreadable residual-stream probe buys two to three times the coverage of the readable signals, an edge a cross-fitted reconstruction cannot recover linearly from the readable features. We then give a floor-started fixed-sequence certificate, valid without monotonicity assumptions, that covers more than the Bonferroni certificate on every model-signal pair and raises coverage at the non-vacuous target $0.75π_0$ from 0.05 to 0.16, although the floor keeps absolute coverage small. Finally, a certificate cannot see what matters after deployment: under benchmark shift the error among accepted traces tracks the new task's base error, and under best-of-$n$ selection against the verifier it rises past the target while the empirical failure frequency stays below $δ$, because abstention absorbs the failures.
多智能体系统 (cs.MA)
5
cs.MA / 1 / 2610.09843
Minimizing Cumulative Envy in Allocating a Sequence of Items
Abstract
We study temporal fair division with indivisible goods that arrive sequentially and must be allocated irrevocably. In contrast to the usual online model, we assume that valuations and future arrivals are known in advance, and ask how unfairness evolves during the process. We introduce \emph{cumulative maximum envy}: the sum, over all rounds, of the maximum pairwise envy at that round. Equivalently, this is the area under the worst-envy curve, and it captures both the magnitude and the duration of envy. For a fixed arrival order, we show that the corresponding decision problem is strongly NP-complete and that minimizing this objective admits no constant-factor approximation unless P = NP, even under identical valuations and even under binary valuations. We complement these hardness results with a dynamic program that gives pseudopolynomial-time solvability for a constant number of agents, polynomial-time algorithms in further restricted settings, and an FPTAS for fixed $n$ under identical integer valuations. We then study a sequencing variant where the algorithm may choose the arrival order. This variant remains NP-complete even for two agents with identical valuations; however, a simple greedy algorithm achieves a $3/2$-approximation for $n=2$ agents, an $n/(n-1)$-approximation for any number of agents, and an additive guarantee depending on the maximum value of any good.
cs.MA / 2 / 2610.10044
CANDO: Cooperative Agentic Network for Layout Design Optimization
Abstract
Layout generation for real-world facilities is a challenging problem, requiring reasoning over irregular site boundaries, heterogeneous orientations, access-aware placements, and motion-planning feasibility. Yet, most existing layout benchmarks in the generative AI space target simpler placements over rectangular domains and rely on distributional metrics such as FID and IoU that reward conformity to dataset priors, thus discounting design innovation. Motivated by these gaps, we introduce ALPS-Bench, a benchmark of $1,000$ professionally annotated real-world facility layouts paired with an instance-specific scoring protocol grounded in a structured design manual. As a strong baseline for ALPS-Bench, we propose CANDO, a training-free multi-agent framework in which specialized agents iteratively refine layouts through a verification-grounded loop, concentrating reasoning on strategic spatial decisions. We demonstrate that CANDO surpasses state-of-the-art trained and LLM-based baselines on the widely adopted PubLayNet, RICO, and PKU-PosterLayout benchmarks, establishing cooperative agentic design as a broadly effective recipe for constraint-aware layout synthesis.
cs.MA / 3 / 2610.10100
The Cost of Classical Multi-Agent Path Finding
Abstract
Multi-Agent Path Finding (MAPF) is the problem of planning conflict-free paths for multiple agents in a shared space, each from its start to its goal. Classical MAPF has been the dominant formulation for many years, with its assumptions of discrete time and graph-based conflicts presumably easing the search for solutions. These assumptions limit the physical environments and agents for which a solution is truly collision-free, and also place an upper bound on solution quality that no algorithmic improvements can lift. This work investigates how much solution quality, and in what contexts, the classical MAPF formulation forfeits. Continuous-time MAPF (MAPF$_R$) relaxes these assumptions, making it a natural counter-formulation to compare against across various agent counts and sizes, and graph connectedness, topologies, and resolutions. We find that continuous time and agent shape consideration are worth relatively little on their own; their value comes from enabling an expanded range of move actions, on average improving solution quality by at least $5\%$ on narrow and constrained maps and $17\%$ on maps with open spaces. In some cases, the improvements exceed $20\%$. Doubling the map resolution with classical MAPF recovers less than $3\%$, meaning that little of what is forfeited can be bought back through more compute. This work therefore provides insight on when classical MAPF is a reasonable simplification, and when MAPF$_R$ unlocks significantly higher-quality solutions.
cs.MA / 4 / 2610.10126
Know the Shape, Find the Fault: Topology-Conditioned Diagnosis of Multi-Agent LLM Failures
Abstract
Multi-agent LLM systems coordinate task execution through exchanges of information among agents. When coordination breaks down, similar symptoms in execution traces can reflect different problems in how information is passed, used, or verified. Communication topology captures how agents exchange information and provides structural cues for distinguishing coordination failure modes. Using these cues for diagnosis requires establishing how topology relates to failure patterns and recovering the relevant structure from execution traces that lack explicit topology labels. We analyze the relationship between communication topology and failure patterns and introduce MAScope, a two-stage framework for topology-conditioned diagnosis. Its Trace Structural Extractor TSE recovers communication topology from heterogeneous execution traces by grounding an interaction graph in message evidence. The Topology-Conditioned Judge TC-Judge then classifies failures using the trace, predicted topology, an empirical failure prior estimated from separate labeled traces, and a short description of topology-specific failure patterns. Under a fixed orchestration structure, the recovered topology can be reused across executions. Experimental results show a statistically significant association between communication topology and failure type, with $χ^2 = 409.9$ and $p = 1.2 \times 10^{-70}$. On the \num{851} MAST-clean traces, ground-truth topology context raises gpt-mini's Macro-F1 from $0.173$ to $0.350$. With predicted topology, the pipeline achieves $0.346$, approaching the trace-only gpt-5.4 baseline of $0.372$. For \num{1000} traces under a fixed orchestration structure, the projected pipeline cost, including one topology extraction, is approximately $6\%$ of repeated gpt-5.4 diagnosis cost. These results show that topology-conditioned context improves failure diagnosis and supports lower-cost deployment.
cs.MA / 5 / 2610.10468
A Society of Researchers: Designing Institutions for Populations of Autonomous Research Agents
Abstract
Deployments of research agents are moving to populations of thousands that share one pool of compute, while most current systems organize one project at a time or leave the population unorganized. We argue that such a population will acquire an organization whether or not its designers provide one, so designers should provide it explicitly, and that the multi-agent systems community holds the tools to do so. We propose a society of agents, a population of persistent agents under explicit institutions, and develop it for science as a society of researchers built on six principles. Principal investigators compete for compute through requests for proposals, independent review, and grants; a human governor, the mayor, allocates resources and assigns no tasks. In a running society of ten thousand researchers, asked only to improve the pretraining of language models, one lab reported a way to reach the same quality with about 30% less compute, a result the labs that tested it do not yet agree on. We close with six open problems for the agents community.
软件工程 (cs.SE)
11
cs.SE / 1 / 2610.08962
Automatically Detecting and Fixing Deadlocks in Go Code with GoDDaR
Abstract
The Go programming language provides a lightweight abstraction for concurrent programming through goroutines, which are prone to deadlocks. Go includes a runtime detector that aborts execution when all threads are blocked (a global deadlock). However, due to nondeterministic thread scheduling, this runtime mechanism only detects global deadlocks that manifest during execution and cannot identify partial deadlocks, where a subset of goroutines is permanently blocked while at least one remains runnable. Statically detecting local deadlocks is essential for developing dependable concurrent software. While several tools statically detect deadlocks in concurrent programs, few assist developers in fixing them. Detecting and resolving partial deadlocks requires reasoning about complex interleavings and communication patterns, an inherently challenging task. In this paper, we present GoDDaR, a tool that detects unobserved global or partial deadlocks in Go programs and suggests concrete fixes. GoDDaR translates Go source code into an intermediate representation reflecting message-passing communication events and composition patterns. Partial deadlocks are identified through symbolic analysis over this representation. To resolve detected deadlocks, our algorithms transform the intermediate representation to eliminate problematic synchronization patterns, allowing GoDDaR to generate fix suggestions for the original Go source code as a diff. We evaluate GoDDaR on a benchmark suite of representative deadlock examples from literature and git repositories to validate algorithm correctness and compare performance against the state of the art. Results demonstrate that GoDDaR is competitive with top approaches in partial deadlock detection and advances the state of the art in automated deadlock fixing.
cs.SE / 2 / 2610.09023
Evaluating Change Point Detection Methods for Software Performance Regression Analysis
Abstract
Performance issues in software systems are a critical quality issue that can erode user trust, violate service-level agreements, and ultimately affect business efficiency. Consequently, software performance engineering has shifted its focus to developing robust techniques to detect performance regressions as early as possible in the development cycle. Performance regression analysis often relies on time series of performance measurements to detect significant changes in performance behavior. Change Point Detection (CPD) methods have been widely used to automate the identification of such changes in various domains, including finance, healthcare, and performance monitoring. However, the effectiveness of these methods for software performance measurements has not been thoroughly evaluated. In this paper, we present a comprehensive study to evaluate the effectiveness of various CPD methods on real-world software performance measurement datasets. We start by collecting performance measurement data from three large software systems and characterizing the unique properties of performance time series data. Then, we undertake a large-scale effort to annotate and evaluate the consistency of human annotators' identification of potential performance changes. Thereafter, we evaluate the accuracy of twelve distinct CPD methods in detecting potential performance changes, providing insights into their applicability and effectiveness in software performance regression analysis.
cs.SE / 3 / 2610.09267
Who Broke Me? Execution-Guided Repair of Behavioral Dependency Breaks
Abstract
Dependency upgrades can break downstream projects without changing the library interface. Such behavioral breaking changes are difficult for developers to fix, because the failing test does not always point to the root API, the upstream API that causes the break. Existing LLM-based repair methods obtain evidence for the repair from compiler feedback or library documentation. However, a behavioral break produces no compiler feedback and is often undocumented. Agents that receive no upgrade evidence also usually do not retrieve upstream evidence themselves, and most of their failed repairs do not identify the root API. The library diff provides useful upstream evidence, but the root API must first be identified to select the relevant part of the diff. We present BBCFixer, a repair method that runs the failing test under the old and new library versions, ranks the calls whose return value differs to identify a candidate root API, and filters the library diff with that candidate. To evaluate repairs of behavioral breaks, we further introduce BBCBench, a benchmark of 100 behavioral dependency breaks in Python and JavaScript. We compare BBCFixer on BBCBench with two baselines: (1) Pure Agent, an agent that receives no upgrade evidence, and (2) BDUpdater, a method that mines library documentation. BBCFixer raises the average pass rate by up to 17% relative to Pure Agent and needs up to 27% fewer agent steps per successful repair. BBCFixer thus provides a way to select the upstream evidence behind a behavioral break, so that an LLM agent can repair behavioral breaking changes more effectively and efficiently.
cs.SE / 4 / 2610.09580
Beyond FAIR: A Fitness Function Framework for Sustainable Research Software
Abstract
Research software sustainability is often assessed through the FAIR principles: findability, accessibility, interoperability and reusability. While important, FAIR does not cover all relevant sustainability concerns. Research software should also be environmentally responsible and secure over time. Excessive resource consumption increases environmental cost, while insecure software creates maintenance overhead, technical debt and barriers to reuse. To this end, in this paper, we extend a previously proposed fitness function framework for sustainable research software beyond FAIR. We introduce two additional sets of fitness functions: environmental functions targeting resource efficiency and execution footprint, and security functions targeting dependency health, vulnerability exposure and secure configuration. Together, these functions broaden continuous software assessment from FAIR compliance to a more complete view of sustainability. The resulting framework treats research software sustainability as a multidimensional property that includes FAIR, environmental responsibility and long-term security.
cs.SE / 5 / 2610.09910
Analysis and Visualization of the Linux Kernel's Software Evolution Using the City Metaphor
Abstract
The Linux kernel is one of the largest and longest-maintained open source projects in existence. With more than 40 million lines of code, understanding the kernel's internal structure and assessing its software evolution is a great challenge. In this paper, we present an approach to visualize the Linux kernel using our software visualization tool ExplorViz. We analyze commits from the Linux Git repository using a custom analysis service. The web-based frontend utilizes the 3D city metaphor for visualization of the software structure. Metrics such as the number of lines are collected for each file and are accumulated for directories and commits. Calculated metrics can be mapped to the building's dimensions or be displayed via a heat map. A wide range of visualization, search, and filter options enable the interactive exploration of the visualized data. We present both a visualization of all files in the kernel repository and a visual analysis of the code evolution of Rust files in the Linux kernel. Demo video: https://youtu.be/3ODnERqD8g0
cs.SE / 6 / 2610.09995
Designing Collaborative AI-Driven Workflows for Scientific Software Engineering
Abstract
Agentic artificial intelligence systems can carry out a broad range of tasks in software engineering and scientific research, from writing and translating code to running workflows for data analysis and visualization. In scientific computing, the difficulty is verifying that agent-generated code is both correct and understandable to teams whose members bring different areas of expertise. We therefore argue that these systems are best used within collaborative team structures rather than as full automation. In the workflows we propose, domain experts write the specification and plan, and agents operate under a deterministic orchestration pattern to write the target code. Each stage ends with a numerical comparison against the reference code and requires human review and approval before the next begins. We evaluate these workflows on the translation of a large high-energy physics application from Fortran to C++, running the same task under different orchestrators, design patterns, and models. Across fourteen experiments, a simple author--reviewer loop with enforced limits completed a comparable number of files to a multi-agent workflow at about one-third of the cost per file.
cs.SE / 7 / 2610.10242
TestGRAD: Evolving Test Suites via Failure Pattern Momentum for SWE-Agent Ensemble
Abstract
SWE-agent ensembles improve issue resolution by combining candidate patches from different agents with complementary strengths. The central problem is therefore test-based selection: generate tests, execute candidate patches, and identify the best patch. We formulate this process as test-space optimization: evolving an executable repository test suite until it distinguishes competing patches. Existing test-generation methods are limited optimizers. They usually lack an explicit loss for ensemble selection, optimize through incomplete directions that mostly create new tests or delete old ones, and perform one-off generation without feedback from repeated failures. Inspired by gradient descent with momentum, we introduce TestGRAD, a framework for automatic test optimization. TestGRAD centers on three concepts. Differential loss gives the optimizer an explicit execution-defined target: useful tests should separate candidate patches by behavior. Full CRUD gradients expand the update direction from merely creating or deleting tests to reading existing test infrastructure, creating new tests, updating stale assertions, and deleting only obsolete tests. Failure Pattern Momentum mines frequent failure sequences from memory, allowing the optimizer to avoid repeated non-discriminative directions while compressing the failure-history context. On SWE-bench Verified, TestGRAD achieves 84.2% Pass@1 with a 4-agent ensemble, outperforming the strongest baseline (80.6%) by an absolute improvement of 3.6 percentage points, while compressing failure-history context by over $100\times$.
cs.SE / 8 / 2610.10258
QuSema: Detecting Silent Bugs in Quantum Libraries via Quantum-knowledge-enhanced Agents
Abstract
Quantum libraries are now critical infrastructure for quantum algorithm development, yet their correctness remains difficult to test. Existing testing techniques mainly rely on failure-based or comparison-based oracles, exposing bugs only when executions fail, violate runtime checks, or disagree with another implementation. Their applicability is limited when suitable execution-based oracles are unavailable, leaving some silent bugs undetected. Such missed bugs can produce incorrect results that propagate into experimental conclusions, simulation studies, and algorithmic designs. Here we present QuSema, an autonomous testing agent for finding silent bugs in quantum libraries. QuSema uses constraints from quantum semantics and documentation as a source-level semantic oracle to assess whether implementation logic can produce invalid outputs from valid inputs. It operates through an agentic loop that repeatedly inspects library API documentation and source code, reasons about the intended behavior of quantum operations, identifies potential semantic deviations, and validates them by generating executable tests through library APIs. Guided by quantum-domain reasoning, QuSema turns high-level behavioral mismatches into concrete, user-triggerable bug reports, enabling it to uncover non-crash defects. We implement QuSema for Qiskit and PennyLane. On a benchmark of 20 historical silent bugs, QuSema achieves higher mean bug relocation counts than Claude Code and Codex, with the DeepSeek configuration costing less than Claude Code. QuSema also discovers 40 previously unknown bugs confirmed by the developers, including 30 silent bugs.
cs.SE / 9 / 2610.10261
Using Small Language Models to Reverse-Engineer Machine Learning Pipelines Structures
Abstract
Context: Once defined a taxonomy of stages structuring Machine Learning (ML) pipelines (e.g. Data Preprocessing, Modeling...), extracting these stages from source code is key for better understanding ML practices. However, the diversity caused by the constant evolution of ML (e.g., algorithms, datasets) makes this task challenging. Existing approaches either rely on non-scalable manual labeling or on classifiers that do not properly support domain's diversity. These limitations call for more reliable solutions. Objective: We evaluate whether Small Language Models (SLMs) can leverage their code understanding and classification abilities to address these limitations, and enhance our understanding of practices in ML. Method: We conduct a confirmatory study based on two relevant reference works representing current limitations in the state-of-the-art. We first compare several SLMs using Cochran's Q test, then evaluate the best-performing model against reference studies via two McNemar's tests. An additional Cochran's Q test examines how taxonomy definition variations affect the SLM performance. Finally, goodness-of-fit tests compare ML practice insights from SLM classification with those from prior studies. Results: First, we found that the taxonomy wording significantly impacts classification performance. Second, the best performing SLM yielded good results, yet, without outperforming other classifiers. Third, the three classification methods led to significantly different insights, with varying effect sizes, when exploring practices of data scientists. Conclusions: Limitations of existing classification methods bias our understanding of ML practices. While current SLMs show promising results without prior fine-tuning, they still exhibit common limitations, in addition to inference high costs challenging their applicability in large-scale studies.
cs.SE / 10 / 2610.10263
When Sub-Agents Work in Parallel: The Promises and Pitfalls of Dynamic Concurrency in Long-Horizon Coding Tasks
Abstract
As coding agents advance from bounded software engineering tasks toward long horizon development, dynamic concurrency offers a promising way to scale complex development tasks. Under this policy, agents decide during execution whether and how to spawn concurrent sub-agents. Model capability largely determines outcomes on shorter tasks, whereas long horizon development makes orchestration central to task completion. Existing work, focused on coding agent failures on shorter tasks or collaboration in predefined multiagent workflows, offers little insight into dynamic concurrency in frontier agents across task complexity. We study dynamic concurrency as an execution policy through controlled comparisons of matched Codex, Claude Code, and Kimi Code executions with the policy enabled or disabled. Across 354 tasks and 2,124 executions spanning a range of task complexities and execution horizons, we evaluate its end to end effects and scheduling behavior, and analyze matched trajectories to characterize 13 concurrency specific failure modes, 28 observable patterns, and the conditions under which it provides an advantage.
cs.SE / 11 / 2610.10427
Design-Time Conformance Checking for Pulse-Level Quantum Control
Abstract
A pulse-level quantum control program is written against a device whose limits are recorded, if at all, in vendor documentation and source code. A program that exceeds them can be refused at compile time. It can also be accepted and silently altered, or compile and then fail at the board. In the last two cases the experiment runs, and the data does not correspond to the program that was written. We present qconform, a checker that decides whether a pulse program is realizable on a device, given a versioned capability descriptor for that device. Every constraint in a descriptor cites the toolchain observation that established it. The checker is design-time, offline, and deterministic, and it uses no floating point. It reports a verdict, the rules it applied, and a coverage manifest that names what it did not check. We evaluate qconform by differential testing against the QICK and Qblox toolchains, on three QICK board configurations and one Qblox cluster. Over 1263 program runs of 969 distinct programs, the checker accepted no program that a toolchain refuses. On both vendors, a descriptor pinned to one toolchain release accepts programs that older releases refuse. The evaluation also found ten defects in the checker and its descriptors, all fixed.
操作系统 (cs.OS)
1
cs.OS / 1 / 2610.10137
Behind a Simple Read: Understanding Work and Waiting in the Linux I/O Stack
Abstract
A simple read interface provides uniform functional semantics, but not equally simple or predictable performance behavior. We use synchronous large-buffer reads in Linux as an observation window and decompose the buffered-read path end to end, distinguishing work volume, processing time, and critical-path exposure. We find that Linux reduces metadata work through large folios and overlaps most cold-read copying with device waiting, but these mechanisms depend on access advice, folio granularity, cache state, and backend execution. Using an experimental kernel prototype, we validate additional opportunities from out-of-order early copying and opportunistic parallelism, while showing that faster request submission does not improve end-to-end performance when SSD supply is already sufficient. To explain device waiting, we abstract first-completion wait $F$ and subsequent completion capacity $B_{CQ}$ from finite request-batch completion timelines. Measurements on a real SSD characterize how they vary with request size and batch size, while MQSim experiments connect them to internal device mechanisms and distinguish finite-batch completion from sustained throughput. We then construct a cross-layer wait chain along the actual folio--bio--request mapping. Independently calibrated device parameters predict waiting at the block layer and at read_pages with errors no greater than approximately 6.9% and 3.2%, respectively. Finally, four use cases apply the model and wait chain to parallel submission, Linux readahead, dependent-read layout, and polling versus sleeping. Their gains, no gains, and gain reversals form a closed loop from observation and modeling to explanation and control, providing a measurable basis for performance decisions in layered I/O stacks.
硬件架构 (cs.AR)
6
cs.AR / 1 / 2610.09156
Compiling Semi-Ring Dynamic Programming to Tier-Aware 3D-DRAM Processing-in-Memory
Abstract
Processing-in-Memory (PIM) on monolithic 3D (M3D) DRAM is a promising answer to the memory wall for data-intensive dynamic programming (DP), yet extracting its performance today demands hand-written kernels: the programmer must pick a tile size, place data across non-uniform-latency memory tiers, partition work between heterogeneous processing units, and insert the right broadcasts and queues by hand. We present GenMLIR, an MLIR compiler that automates this mapping for tier-aware 3D PIM. GenMLIR encodes the semi-ring generalized grid update, the algebraic form shared by all-pairs shortest path (APSP) and sequence alignment, as a first-class IR abstraction, and lowers it through four GenDRAM-aware pass groups that derive blocked tiling, tier-aware placement, tile-to-PU assignment, and explicit communication. We also characterize precisely which DP recurrences the abstraction admits and which it does not. On the GenDRAM architecture, GenMLIR-generated code runs up to 5.8 (APSP) and 16.7 (alignment) faster than a lowering without compiler support, and 1.2-1.4 faster than a standard affine-tiling PIM compiler we also implement, reaching 90-100% of an achievable-performance bound. It expresses these workloads in 4-9 lines instead of 300-500 and compiles in a negligible fraction of runtime.
cs.AR / 2 / 2610.09292
An Interleaved Parallel Dependent Quantization Hardware Architecture for H.266/VVC
Abstract
While dependent quantization in H.266/VVC delivers a high compression ratio, its strong serial nature and high complexity result in poor real-time performance, making it difficult to deploy in practical scenarios. To improve the real-time performance of dependent quantization with minimal degradation to its compression performance, we propose a interleaved parallel dependent quantization hardware architecture with low BDBR loss, which achieves four-channel parallel dependent quantization by time-division multiplexing most combinational logic. This architecture adopts the proposed intra-CG context simplification scheme and the encoding scheme that skips decAbslevel during rate estimation. The proposed design incurs a BDBR loss of only 0.42% under the All Intra configuration and 0.38% under the Random Access configuration, respectively. Implemented in Verilog HDL, the dependent quantization hardware architecture achieves quantization speeds of 4K@33.2, 91.4, 331.2, and 454.1 fps at QP = 22, 27, 32, and 37, respectively, when implemented on the Xilinx XCZU19 FPGA. When implemented on ASIC using the TSMC 28nm process standard cell library, the corresponding quantization speeds reach 4K@84.0, 231.0, 837.0, and 1147.7 fps.
cs.AR / 3 / 2610.09447
BenchmarkAnything: Agent-Driven Construction of Simulator-Ready Microarchitecture Benchmarks
Abstract
The selection of benchmark workloads is of paramount importance in computer architecture, as it establishes the yardstick against which architectural innovations are measured and guided. Yet for decades, the SPEC benchmark suites, comprising merely tens of workloads, have been the de facto standard in academic architectural research, where they are frequently treated as a principal evaluation and optimization target. When a suite this small is relied upon so heavily, it risks architectural overfitting; as our research and prior studies demonstrate, an overly narrow focus can mislead design decisions by overvaluing certain innovations, producing cores that excel on SPEC benchmarks yet underperform on broader, realistic workloads. To mitigate this overfitting, adopting a large, comprehensive benchmark suite is the natural solution. However, the immense engineering effort required to strip software into the clean, interference-free binary executables demanded by simulators often makes this highly impractical. In this work, we demonstrate that AI agents provide an elegant solution to this challenge. Rather than manually curating yet another static benchmark suite, we introduce an agent-driven workflow capable of autonomously transforming arbitrary open-source repositories into simulator-ready executables. This automated approach makes workload collection highly scalable, allowing us to rapidly harvest hundreds of diverse applications from public repositories into our benchmark suite. Through a comparative analysis of our agent-generated suite against SPEC, we show that it not only achieves higher-fidelity performance assessments but also uncovers novel architectural insights that traditional, static suites fail to expose.
cs.AR / 4 / 2610.09999
ReSAFT: An Efficient Stuck-at Fault-Tolerant Scheme for ReRAM-based Process-in-Memory Accelerators
Abstract
Analog ReRAM-based process-in-memory (PIM) accelerators provide high parallelism and energy efficiency for deep convolutional neural networks (CNNs) inference. However, their susceptibility to permanent faults, such as stuck-at high (SaH) and stuck-at low (SaL) resistance states, poses a major challenge by permanently corrupting the CNN weights mapped to conductance values of ReRAM cells and degrading inference accuracy, which leads to system unreliability in safety-critical applications. In this paper, we propose a fault-tolerant scheme for analog ReRAM-based PIM accelerators to tackle stuck-at faults (SAFs) with minimal redundancy overhead to recover classification accuracy degradation. The proposed scheme contains a redundancy-based hardware solution alongside fault-aware mapping method for ensuring reliable analog computation in ReRAM crossbar. We analyze the impact of varying number of redundant rows and columns on accuracy and design metrics. Subsequently, a multi-objective optimization (MOO) problem is formulated and solved to efficiently determine the number of redundant rows and columns, considering trade-offs among various design metrics. Furthermore, a fault-aware weight mapping is proposed for dual-crossbar structures to further compensate for the accuracy degradation caused by SAFs. Simulation results show that, for the SimpleNet model using the MNIST dataset, the inference accuracy is recovered by approximately 22.39%, on average, across four configurations of optimal solutions, each offering a trade-off between reliability and area, energy consumption, and latency overheads. The mean-time-tofailure (MTTF) improves by about 61x on average compared to the baseline. These selected configurations also reduce energy and area overheads by 32%, on average, in comparison to row-only and column-only configurations.
cs.AR / 5 / 2610.10129
When Algorithmic Exploration Becomes Cheap: A Case Study of Agentic Research in EDA
Abstract
As EDA researchers, we conducted eight deliberate trials of agentic algorithm exploration, selecting several topics outside our areas of depth. One faculty member and seven students participated, including students without publication experience. With limited intervention in the algorithms, agents developed mathematical constructions, analyzed existing tools, and implemented improvements; some efforts fell short of their practical goals. We also used AI to collect, classify, and analyze 8,420 papers from four EDA conferences and two journals over 2022-2026. Among 2,380 primary-core papers, we classified 97.7% from titles and abstracts as computationally closed, including work on new formulations. Together, these observations suggest that much of EDA offers an executable environment for increasingly accessible algorithm research. We see an opportunity for tool developers to investigate ideas they previously lacked time to pursue. We also ask how EDA should validate and reward research when results become easier to produce than to examine, and what papers and venue labels will continue to tell us about a contribution.
cs.AR / 6 / 2610.10403
SUSpMV: A High Frequency Sparse Matrix Vector Multiplier on HBM Enabled FPGA written in SUS
Abstract
SUSpMV is a Sparse Matrix Vector multiplication (SpMV) accelerator written in the upcoming HDL SUS. By leveraging SUS's unique latency counting and inference mechanism, SUSpMV could be designed with very deep pipelines yet small design complexity overhead. This enables an efficient implementation which employs all 32 HBM channels on the Alveo U280 FPGA at 400MHz for streaming matrix data into 32 Compute Units (CUs). Input and output vectors are stored in DDR memory, which allows zero-overhead chaining of multiplications and increases overall system memory bandwidth by not sharing HBM bandwidth with the CUs. A CU processes the SpMV in tiles of width 1024 and a dynamically chosen height, up to 32768. Each CU is capable of accumulating the multiplications with up to 6 separate matrix entries per cycle, resulting in the combined theoretical peak computational throughput of 153.6 GFLOPs. The matrix storage format is designed to exploit density variation within a given matrix by dynamically switching between one representation optimized for denser regions, and a second optimized for sparser regions, allowing jumps of up to 255 rows between each entry. Evaluation demonstrates a 79% geometric mean improvement over prior work on the same platform. We achieve a peak throughput of 144.9 GFLOPs, or 94% of our theoretical computational throughput, compared to 98 GFLOPs reached by prior work.
密码学与安全 (cs.CR)
30
cs.CR / 1 / 2610.08928
Multi-Aspect Runtime Verification for Simulation-Based V&V of LLM-Enabled Autonomous Agents
Abstract
LLM-based agents are entering decision-support roles in defence staff work, where the obligations they must respect are already written down and binding, and where retraining is not available as a control because models arrive as procured components. What can be placed under engineering control is the interface between the agent and the systems it acts on. Those obligations are at once spatial, temporal and text-semantic, and a violation typically lives in the composition of a multi-step interaction, which is why per-event guardrails miss sequential tool-attack chains. We present a multi-aspect runtime-verification framework that decomposes a natural-language policy clause into a typed spatial/temporal/semantic triple over one canonical event stream, checks each aspect with its own monitoring specification, and fuses the verdicts through a four-valued algebra that carries provenance. The spatial aspect is interpreted over a weighted two-sorted location graph in which mission geometry and information-release topology are one object; we show that these spatial obligations are not in general subsumed by a first-order temporal specification. The past-time aspect runs on the unmodified MonPoly engine, which agrees with our reference monitor at every time point. Across two mission domains, casualty evacuation and contested sustainment, and one civil domain, composition under the precautionary blocking policy drives attack success to zero with no observed false positives and microsecond-scale per-event cost, while every single aspect and every pair leaves a substantial share of attacks succeeding. In a closed-loop experiment a policy-naive planner reaches a violating state in most unshielded missions and in none when shielded, and four refused episodes in five still recover to a compliant outcome.
cs.CR / 2 / 2610.08951
ASPIRE: Agentic Safety & Prompt Injection Red-teaming Engine
Abstract
LLM agents retrieve untrusted content and act through tools, creating indirect prompt-injection risks that can cause unauthorized actions or persistent state changes. Existing automated red-teaming largely optimizes payloads for pre-specified scenarios, leaving latent vulnerabilities across the agent's behavior space unexplored. We present ASPIRE, an Agentic Safety & Prompt Injection Red-teaming Engine for open-ended, behavior-level vulnerability discovery. ASPIRE maintains an evolving Agent Security Behavior Graph and uses complementary Explore and Exploit experts to discover, verify, and generalize consequence-centric tests. Trajectory evidence updates the graph and diagnoses partial or failed attempts, while cross-run memory transfers useful red-team strategies. Experiments on various benchmarks show that ASPIRE substantially expands coverage across consequences, injection methods, environments, and behavior paths while maintaining strong attack success.
cs.CR / 3 / 2610.09012
TwinGuard-Lite: A Rule-Based State-Admission Gateway for Generative Patient Digital Twins
Abstract
Future generative patient digital twins may combine longitudinal health records with language-model agents and keep information across sessions. The wording of a proposed update does not reveal whether it comes from an allowed source, conflicts with the patient's record, or belongs to someone else. We present TwinGuard-Lite, a rule-based gateway that admits an update to a twin's persistent state only if it passes checkable approximations of two integrity properties. Grounded state consistency (GSC) is approximated by a provenance check and a contradiction check; cross-patient noninterference (CPN) is approximated by a namespace check against trusted transport metadata. We evaluate three attack families over 20 seeded person-level splits of a semi-synthetic stream built from the openly licensed eICU database demo, with 29,131 $\pm$ 2706 candidate updates per seed. A keyword filter and an anomaly detector detect 16% and 24% of attacks, respectively; a provenance-only ablation detects 59.9% but misses every cross-patient attack. Combining GSC and CPN yields 99.1% precision, 90.5% recall, a 94.6% F1 score, and a 0.033% false-positive rate. Every attack the full gateway admits falsely claims a trusted source; when the retrieval channel does so, 92.9% of retrieval attacks are admitted. The results hold under stated trust assumptions, and no language model, agent, or retriever is executed: TwinGuard-Lite is a mechanism-level proof of concept that motivates authenticated, patient-bound ingestion, not a clinical safeguard.
cs.CR / 4 / 2610.09027
Visual Memory Attacks Can Persist Through The KV Cache
Abstract
Modern language model systems operate autonomously over increasingly long contexts containing untrusted text and images. Can an adversarial input continue to steer a model even after that input is removed from its context? We show that attacks can be trained to persist through the key/value (KV) cache of subsequent tokens, allowing adversarial influence to outlive direct access to its source.We consider the Visual Memory Injection (VMI; Schlarmann and Hein, 2026) attack setting, in which an adversarial image that stays in the context plants a hidden backdoor: the model behaves normally until a chosen trigger elicits an attacker-chosen response. We first demonstrate persistence in this setting with optimized soft prompts, which remain effective after we mask the prompt from attention. We then introduce Persistent Visual Memory Injection (P-VMI), which optimizes images to preserve this adversarial behaviour after they are masked from attention. These attacks persist over conversations substantially longer than those used during optimization. On Qwen3-VL-8B-Instruct, P-VMI achieves up to approximately $90\%$ target success in its strongest configuration and remains effective under a stricter removal setting that exposes the image only on the first turn. A cache-swap ablation localizes the persistent influence to the KV cache. Finally, we show that these attacks can be trained to survive compaction that retains the KV cache of a summary generated by the same model, demonstrating that adversarial behaviour can persist in cached state without continued access to its source.
cs.CR / 5 / 2610.09029
Cost of Delay for Post-Quantum Migration: Putting Classical and Harvest-Now-Decrypt-Later Risk on One Ordered List
Abstract
Organisations deciding which assets to migrate to post-quantum cryptography first, and how that work competes with a backlog of classical findings, lack a common unit: post-quantum scores are dimensionless and quantum-only, while classical risk is annualised loss. We express both threats as a cost of delay in currency per year on the same asset. The quantum term is the rate at which deferring migration commits irreversible harvest-now-decrypt-later loss, computed from required confidentiality duration, recordable traffic share and data turnover, using a quantum-arrival law fitted in closed form to published expert anchors. A competing-risks factor couples the two terms so the same loss is not counted twice. We prove that the construction is not a weighted sum of a classical and a quantum score, give a dominance threshold on the classical hazard that is independent of asset value, and show that the induced order is optimal for sequencing work under constant loss rates. We verify the implementation by simulation, adaptive quadrature and a second implementation. On four public systems with parameters assumed from public documentation, the quantum term reorders the asset list beyond input uncertainty for one system only (signal-to-noise 1.36 against 0.14-0.46), moves specific long-lived, quiet, recordable assets decisively, and otherwise changes magnitudes. Asset value and turnover explain 73-91% of the uncertainty in the quantum term; the arrival date explains 3-17%. The case studies are illustrative and are not validated against outcomes.
cs.CR / 6 / 2610.09138
SwarmReconGuard: Black-Box Detection of Distributed Collective Reconnaissance by Individually Benign-Looking Agent Populations
Abstract
Autonomous and agentic clients can distribute reconnaissance across many identities so that each request remains valid, low-rate, and benign-looking while the population collectively acquires broad system knowledge. We formalize this threat as Distributed Collective Reconnaissance (DCR) and present SwarmReconGuard, a reproducible black-box benchmark in which the defender observes only service-boundary telemetry. The Docker-isolated study evaluates 11 benign and attack behaviors across 10-10,000 virtual identities, comprising 440 test runs and 3,666,300 requests, with complete telemetry integrity. We compare semantic, Gaussian, conditional, graph, kernel, hybrid, and CUSUM-based detectors. Gaussian likelihood-ratio detection achieves 100$\%$ detection with 0$\%$ observed false positives on known attacks but only 3$\%$ on unseen policies. CUSUM yields 36.1$\%$ overall detection at 1.25$\%$ false positives, while hybrid CUSUM reaches 85.7$\%$ detection with 0$\%$ observed false positives at 10,000 identities. Results expose a major policy-generalization gap and motivate exposure-aware, scale-aware defenses.
cs.CR / 7 / 2610.09240
Adversarial Images Hijack Web Agents from Visual Grounding to Browser Execution
Abstract
Modern web agents built on large vision-language models process webpages, select relevant UI elements, and translate model outputs into browser actions. Existing visual red-teaming approaches use adversarial visual content to manipulate this process. However, they primarily target model inference and do not explicitly account for structured input processing or action post-processing. Consequently, model-level success does not establish control over browser execution and cannot reliably characterize end-to-end agent robustness. To address this gap, we formulate red teaming for vision-grounded web agents as an end-to-end grounding-to-execution problem, and introduce WebMirage, a framework that crafts localized visual perturbations that cause agents to select attacker-controlled content and execute the corresponding browser action across varying webpage renderings. It uses a role-slot abstraction and webpage recomposition to capture competition among webpage elements, and dataflow analysis to align optimization with action post-processing. We evaluate WebMirage across four agent configurations and six VLM backbones on 2,250 tasks covering 13 public websites and a sandbox benchmark. WebMirage achieves an average attack success rate of 91.9%, compared with 17.4% for the strongest baseline, and remains effective against three agent-level defenses.
cs.CR / 8 / 2610.09264
Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Files
Abstract
Modern agentic coding frameworks increasingly rely on community-shared rule files (e.g., AGENTS.md or .cursorrules) to guide autonomous code generation, yet the security risks of this pipeline remain underexplored. To bridge this gap, we introduce the package hallucination attack, where an attacker injects malicious prompts into benign rule files to induce coding agents to replace legitimate dependencies with attacker-controlled packages. To obtain effective malicious prompts injected into rule files, we propose PackHallu, an evolutionary optimization framework that iteratively rewrites these injected prompts using trajectory-level feedback and LLM-guided mutations. Evaluations across multiple benchmarks, LLMs, and agent frameworks show that PackHallu achieves high attack success rates and strong transferability across diverse models and agent combinations. Our findings demonstrate that coding agents are vulnerable to package hallucination attacks, highlighting the urgent need for stronger security safeguards in autonomous coding systems.
cs.CR / 9 / 2610.09469
Secure-CUA: Controlling Untrusted Influence in Computer-Use Agents
Abstract
Computer-use agents (CUAs) perform tasks across applications (such as desktops, mobile apps, and web browsers) by observing graphical interfaces and issuing commands such as clicks and keystrokes. These interfaces combine trusted controls and content with untrusted content needed for legitimate tasks. An adversary controlling this untrusted content can embed instructions or misleading visual cues to change the agent's intended action or redirect its commands to the wrong interface target. We formalize security requirements for both the agent's decisions and their execution through GUI commands. In an ideal execution model, we show that enforcing both requirements at each step protects execution traces. We instantiate this model in Secure-CUA, our system for secure CUA execution. Its key idea is to commit to an explicit per-action program, called an $\textit{action transaction}$, before accessing untrusted content. Each transaction fixes its queries to untrusted content and the permitted uses of their responses. The system masks untrusted regions and evaluates each transaction to produce the next action, using an isolated query model to answer its queries. It then locates the intended interface target using the masked interface. Under the model's assumptions, Secure-CUA is secure by design, while generating a new transaction at each step helps maintain high task utility by adapting to changing interfaces. We evaluate Secure-CUA under benign conditions on 400 WebArena tasks using three frontier models across $5$ seeds, yielding $6,000$ execution traces. Secure-CUA achieves an average task success rate of $53.55\%$, compared with $55.12\%$ for Vanilla-CUA and $13.17\%$ for CaMeL-CUA.
cs.CR / 10 / 2610.09552
Constitution-Guided Watermarking
Abstract
Watermarking enables language model providers to identify text generated by their models. However, its desired properties can conflict (\ie~stronger watermark signals can degrade text quality), while designs that resist editing may also facilitate forgery. Providers address these trade-offs by choosing configurations that balance competing objectives or prioritize particular properties. Either approach imposes a shared operating point on requests with different requirements, potentially sacrificing quality where wording preservation matters or robustness where reliable attribution is essential. To allow flexible and adaptable designs, we introduce \emph{Constitution-Guided Watermarking}, a framework that selects request-appropriate trade-offs from provider requirements, listed as natural-language principles. \emph{Offline}, a pretrained reasoning agent examines constitutional rules alongside watermark implementations and iteratively refines rule-specific configurations using empirical feedback. \emph{At deployment}, a separate monitor identifies applicable rules and retrieves the corresponding policy, including watermarking exemptions, without modifying the serving model. Furthermore, our framework supports offline parallel optimization and refinement of rule-specific configurations based on evolving provider requirements without affecting deployment, and binds each deployed configuration to its evaluation evidence, making deployment decisions auditable. In a proof-of-concept evaluation using KGW and a five-rule constitution, our framework selects configurations responsive to provider priorities and improves post-paraphrase detection on robustness-prioritized requests by up to $14$ percentage points over fixed configurations, while matching or exceeding all baselines in aggregate quality and clean detection at a nominal $0.1\%$ false-positive rate.
cs.CR / 11 / 2610.09581
Correct Answers, Unsupported Findings: Evidence Binding in Forensic Reconstruction of LLM Agent Logs
Abstract
Forensic reconstruction of LLM-agent actions requires not only recovering the correct value, but establishing which preserved record supports that finding. Tool logs, generated explanations, and local citation identifiers capture different parts of this evidence, yet a citation identifier does not establish a source unless its binding to a record is preserved. We audit this distinction using 64 mechanically checkable cases from saved AgentDojo Banking executions. Two LLM readers reconstruct source relationships under controlled variations in visible evidence and identifier-to-record bindings. We separately evaluate complete-record agreement, evidence-grounded findings, justified abstention, and unsupported assertions. With original identifiers and no binding table, Sonnet recovered every literal source location but made unsupported citation-source assertions in 26 of 28 cases requiring the missing relation; 22 nevertheless matched the complete reference. Adding explicit bindings improved grounded reconstruction for both readers, whereas identifier renaming alone provided no consistent remedy. A deterministic same-packet comparator correctly resolved the bounded task or abstained throughout. These results show that factual agreement alone is insufficient for evaluating forensic reconstruction of agent logs and motivate preserving explicit record bindings to distinguish supported findings from correct guesses.
cs.CR / 12 / 2610.09697
Automotive Hardware Attacks: An Architect's Guide to TARA
Abstract
Threat analysis and risk assessment (TARA) according to ISO/SAE 21434 treats implementation-level hardware attacks inconsistently. We show that two of the three attack-feasibility approaches exemplified in the standard rate every attack path that requires physical access as very low by construction, independent of the implementation, while the third can assign the same path any of the four ratings, depending on the assumed state of attacker knowledge and on the attacker profile. This paper presents a hardware-aware extension of the TARA that covers side-channel analysis (SCA), fault-injection attacks (FIA), and abuse of debug and test interfaces. It adds an attacker model for physical access over the vehicle lifecycle; a path-dominance relation on the attack-potential factors that identifies when a simpler path makes a sophisticated one irrelevant; a three-level hardware relevance gate (H0 to H2) with an explicit decision rule that fixes whether resistance of a hardware step may be argued, must be documented, or must be demonstrated by test; and a rating rule under which a hardware step receives no credit for implementation-specific resistance until the required evidence exists. The gate is positioned relative to cybersecurity assurance levels and targeted attack feasibility. The method is applied to a reference gateway architecture, and its structural inputs are examined on ten published attacks on automotive components. In seven of the ten studies, the demonstrated attack included a debug, programming, boot, diagnostic, or update mechanism, whereas only four studies included an attack on a cryptographic computation.
cs.CR / 13 / 2610.09703
Understanding and Mitigating Token-Pruning-Induced Vulnerabilities in VLMs
Abstract
Token-Pruning accelerates Vision-Language Models by removing redundant visual tokens, yet its safety implications remain underexplored. In this work, we present the first comprehensive safety evaluation of Token-Pruning mechanisms and find that: most pruning strategies significantly degrade safety as pruning ratios increase, whereas Query-based Compression shows the opposite, with extreme pruning (up to 99.8%), unexpectedly improves model safety. This sharp contrast prompts a key question: How do different Token-Pruning strategies reshape model safety behavior, and is it possible to enhance safety without sacrificing acceleration? To answer this, we identify an unrecognized mechanism, termed Pruning-Induced Malicious Amplification, where removal of background tokens triggers a side effect: forcing the model's attention to collapse onto a few retained malicious anchors within the foreground, inadvertently amplifying their toxic semantics under jailbreak. To address that, we propose an inference-time and plug-and-play Safety-Aware Pruning (SAP) mechanism that counteracts such dominance via three steps: (1) identifying malicious anchors, (2) restoring pruned benign tokens, and (3) reallocating excessive attention from malicious anchors to benign tokens. Extensive experiments across three safety and four utility benchmarks demonstrate that SAP mitigates pruning-induced vulnerabilities, i.e., reducing ASR by up to 62%, without compromising efficiency or utility.
cs.CR / 14 / 2610.09708
Black-Box Adversarial Patch Attacks on VLAs via Ancestor VLM Exploitation
Abstract
Vision-Language-Action models (VLAs) are increasingly deployed in safety-critical physical environments, yet their adversarial robustness remains poorly understood. Existing attacks typically assume white-box access or rely on surrogate VLAs, which rarely holds in real-world deployments. Our key insight is that most VLAs are adapted from a publicly released pretrained vision-language model (VLM), inheriting two capabilities essential for action generation: visual perception and instruction-conditioned grounding. Therefore, this paper explores a previously unaddressed question: can an adversary attack deployed VLAs using only their ancestor VLMs? To this end, we propose three adversarial patch attacks that disrupt the inherited capabilities: a vision disruption attack that corrupts the projected visual tokens through relative and absolute terms, an instruction-grounded semantic evidence suppression attack that removes the visual evidence required for instruction-grounded concepts, and a joint attack that unifies both objectives under a two-phase curriculum. Experiments across different VLA families on both simulation and static real-world images show that patches optimized on the ancestor VLM cause substantial degradations in VLA task success rates, demonstrating that VLAs inherit adversarial vulnerabilities alongside their foundational capabilities. This effect is not uniform: it is strongest on tasks that require precise instruction-grounded localization, and nearly vanishes on policies whose adaptation rewrites the shared visual-semantic representation or whose action head iteratively smooths perturbations away. By characterizing the boundary conditions of vulnerability inheritance and providing analysis of why the inheritance effect holds or fails, we advance the understanding of safety for VLA-involved systems.
cs.CR / 15 / 2610.09730
Trust a Few: The Weakest Assumptions a Protocol Needs
Abstract
Protocol verifiers check whether a protocol meets a security goal under stated trust assumptions, such as that a key is never leaked, a value is fresh, or a channel is authentic. They do not say which of those assumptions the goal needs. Rowe, Guttman and Liskov asked for the weakest assumptions under which a protocol achieves a goal and left the question open. We answer it for assumptions about keys, values and channels. Call a run that violates the goal an attack, and the assumptions that would rule it out its stopping set. The least a protocol must trust to meet a goal is exactly the set of minimal ways to stop all of its minimal attacks. Several such sets may exist. The answer becomes unique once either/or assumptions are allowed, and it is a single set exactly when every minimal attack is stopped by a single assumption. A Galois connection between assumptions and goals explains why: a goal may end in "or", but its hypothesis may not. The same structure gives the trust needed by a conjunction of goals and an exact condition under which composed protocols need no extra trust. To compute the answer, a loop asks a verifier whether a candidate suffices, records what stops each attack it reports, and recomputes the candidates. Deciding whether some trust of at most k assumptions suffices is NP-hard. With a bounded analyser of our own and with CPSA, the loop finds the weakest trust for 18 goals over ten protocols and their variants, and on each it agrees with evaluating every trust using the same verifier. The answers include an assumption that our model of the adopted fix of Kerberos PKINIT states but the client's authentication of the server does not need, and channel assumptions under which Needham-Schroeder meets its goal in a bounded model.
cs.CR / 16 / 2610.09767
Faster PMNS Multi-precision Multiplications Using Truncated Montgomery Technique
Abstract
The Polynomial Modular Number Systems (PMNS) aim to represent elements of fields or rings of large characteristics using polynomials satisfying bounds on some parameters (degree, absolute values of the coefficients). Those PMNS, while some conditions are fulfilled, using convenient parameters and implementation features, allow some speed-ups in cryptographic computations. Recent works (Meloni \emph{et al.} \cite{MeloniPV25}) propose better speed-ups by improving the parameter generation of the system, and by using multi-precision polynomial coefficients, i.e. coefficients stored using several machine words. In this work, we first present PMNS schoolbook software implementation better in some context than the Toeplitz state-of-the-art counterpart of \cite{MeloniPV25}, taking advantage of smaller memory cost by being free to choose a smaller $n$ parameter, minimizing at the same time the complexity of the multiplication. This implementation for 4096 bit modulo size shows element size smaller by 11\% and 20 \% speed-up. We thus present a new improvement on the PMNS, applying to the \textsl{internal reduction} an approach similar to the truncated Montgomery reduction technique presented by Didier \emph{et al.} in \cite{DidierEGR24}. In the context of software implementations using \texttt{AVX512} instruction set extension, and modulo size up to 8192 bits, this new approach allows speed-ups up to 15\% in modular multiplication computation, in comparison with conventional approaches.
cs.CR / 17 / 2610.09793
Formal Runtime Verification for Tool-Using LLM Agents: An Offline Same-Benchmark Study on AgentDojo and STAC
Abstract
Guardrails for tool-using LLM agents are usually application-specific rules, which makes multi-step, data-dependent safety policies hard to specify, audit and reuse. As a declarative alternative, we evaluate metric first-order temporal logic (MFOTL), replaying the recorded trajectories that AgentDojo, STAC and R-Judge already ship through the unmodified MonPoly monitor, offline and without running an agent. On these corpora, five generic obligations flag 71.8% of STAC attack chains and 70.1% of successful AgentDojo attacks, but also fire on 29.3% of benign runs. This imprecision stems from the corpora rather than the logic: they rarely record approvals and never record timestamps, so history-dependent obligations reduce to detecting risky action types. Where the trace does carry relational context, provenance-aware policies discriminate better; that context, however, is itself attackable, and one planted line defeats a naive provenance check on 94-99% of the runs it would otherwise flag. Binding provenance to the lookup that produced it closes this evasion at no cost in detection or benign firing. Taken together, these results show that formal temporal monitoring adds value exactly when the trace exposes trustworthy history. We therefore quantify how far current benchmarks are from that point and propose a twelve-field enforcement-ready trace schema.
cs.CR / 18 / 2610.09817
Hierarchical Security Monitoring for Edge-IoT: A Formal Methods Approach
Abstract
Cyber resiliency in edge-IoT deployments is fundamentally an economic problem: detection must keep critical processes operating under attack, but defender resources (compute, bandwidth, operator attention) are bounded. Centralised cloud monitoring offers expressive cross-device detection at prohibitive bandwidth cost; purely edge-local monitoring is cheap but blind to coordinated multi-device attacks where the asymmetric balance favours the attacker. We propose a lightweight hierarchical security-monitoring framework, built on formal runtime-verification methods, that occupies the practical middle ground at quantified cost. Each edge device runs a lightweight TeSSLa stream specification (size, payload validity, rate, and timestamp-drift predicates) that emits a four-valued verdict per aggregation window at sub-microsecond per-event cost; the gateway runs a parametric first-order MonPoly monitor over the per-device verdict streams at microsecond-scale per-verdict cost. The edge-to-gateway uplink carries roughly one Boolean per aggregation window per node, orders of magnitude smaller than the raw packet stream. The gateway tier detects coordinated attack patterns that no single-node monitor can see, shifting the asymmetric cost balance toward the defender. Every alert carries a witness set naming the device, the monitor tier, and the predicate that fired, providing an auditable record of the decision. We evaluate the framework on a container-host testbed spanning nominal and attacker nodes across four attack classes (buffer overflow, time spoofing, denial-of-service, and mixed advanced-persistent-threat patterns), and describe the edge- and gateway-tier specifications together with the cost-versus-coverage trade-off as monitor levels are added.
cs.CR / 19 / 2610.09825
Hybrid Hierarchical Runtime Verification for Edge-IoT Security: Combining MonPoly and RTLola
Abstract
Security monitoring of edge-IoT fleets faces three structural challenges. (i) A per-node monitor is cheap but cannot see attacks that coordinate across devices. (ii) A cloud monitor sees the full fleet but pays for that view in bandwidth. (iii) Even at the cloud, a monitor built on a single RV engine can be fooled by an attacker who compromises a device, raises one malicious request, and then goes silent: once the events stop, an event-triggered monitor has nothing left to evaluate. We propose a three-layer hierarchical runtime-verification framework that addresses all three. The edge layer classifies events as they happen, the gateway layer aggregates short windows of per-device behaviour, and the cloud layer runs two complementary RV engines. MonPoly handles first-order temporal correlation over the merged alert stream: coordinated overflow (which genuinely quantifies across devices) plus per-device multi-vector APT, escalation, and persistent-campaign patterns. RTLola handles a time-triggered silent-node property that an event-triggered engine cannot detect within a bounded delay under fleet silence. We evaluate the framework on a 15-actor Docker testbed covering eight attack profiles plus a silent-bypass scenario. In the controlled labelled testbed, every device-attributable incident the framework raises names an attacker-labelled device, and the RTLola tier catches silent-bypass attempts the event-triggered tier misses. Per-event monitoring stays in the microsecond range at the edge and gateway, with low end-to-end alert-to-incident latency at the cloud.
cs.CR / 20 / 2610.09892
Defensive Sufficiency in a Stackelberg Model of AI Security
Abstract
Feedback from automated testing, human red teaming, and incident response can strengthen an AI system's defenses when discovered failures lead to effective repairs. We study when this feedback process provides sufficient protection and when investing in it is economically worthwhile. We begin by showing that an attack surface composed of finite number of inputs is defended with probability 1 if every unresolved attack has a persistent chance of discovery, repairs are effective, and subsequent updates preserve earlier protection. We derive completion-time bounds and extend the analysis to growing attack surfaces, repairs that generalize across related attacks, and multiple discovery mechanisms. These results distinguish eventual protection against each fixed attack from complete protection at a single time. We then formulate a defender-led Stackelberg game in which the defender invests in proactive discovery and reactive repair, anticipating the attacker's choice of search effort. We characterize the least-cost allocation that deters attack and the equilibrium regimes in which the defender funds neither capability, one capability, or both. Numerical experiments illustrate these regimes and show how faster repair can reduce compromise duration without reducing compromise probability.unified theory of performance limits in generative language models.
cs.CR / 21 / 2610.09918
CYBERFORT: A Compliance-Chain Platform Operationalising the Cyber Resilience Act for SMEs
Abstract
The EU Cyber Resilience Act (CRA) turns product cybersecurity into a lifecycle compliance obligation for manufacturers, importers, distributors, and integrators of products with digital elements on the EU market, a load that falls largely on small and medium-sized enterprises (SMEs) that rarely have dedicated governance, risk, and compliance (GRC) capacity. We present CYBERFORT, an open-source CRA-first compliance platform developed under the EU Digital Europe Programme and one of twelve projects in the EU CRA cluster. CYBERFORT operationalises the CRA through a guided scope self-assessment, a question bank tied to Annex I and the vulnerability-handling obligations, and a compliance-checking engine that links every answer to controls, policies, and machine-attested evidence, reusing ISO/IEC 27001, NIS2, and GDPR controls only where they coincide with CRA obligations. Its central contribution is the compliance chain, a traceable structure linking each product risk through its controls and policies to the CRA obligations it satisfies, and onward through evidence to the technical-documentation file and EU declaration of conformity, so that every operational gap is traceable and can be closed before market placement. Deployed at https://access.cyber-fort.eu/ for a first cohort of 43 organisations, the platform is presented with the engineering behind the chain, measured results from a completed end-to-end case study on a SIEM/XDR product with AI-driven remediation spanning the CRA obligation chapters, and the controlled effort study that remains in progress.
cs.CR / 22 / 2610.09981
Sensitive-Topic Leakage Through LLM Routing Metadata: Measurement and Mitigation
Abstract
LLM routers pick a cheap or expensive model per request by its content, and many gateways and some cloud platforms can log that choice with content logging off. We measure this privacy channel beyond token counts, accounting for noisy labels and repeated prompts. We run pre-registered studies on 1.7 million real requests (WildChat-1M, LMSYS-Chat-1M) with two cost/quality routers and a domain router, survey eleven systems' logging, and test post-processing defenses. At matched length, the shift's direction depends on category and router. For RouteLLM at the 50% operating point, harassment and self-harm requests reach the strong model 19 points less often than comparable ones on prompts unseen in exploration, medical requests (exploratory: LLM labels failed their gate) 31 points less often on distinct prompts (both post hoc), and sexual requests 10 points more often (secondary); the other router's four are negative. Twenty RouteLLM decisions separate frequent medical askers with AUC 0.71, exploratory and below the pre-registered primary endpoint's 0.75 (domain router: 0.92, an upper estimate). Per-category length-matched parity with accurate labels removes the gap on real traffic, costing at most 0.2 accuracy points on RouterBench (post hoc), where routers' gaps on sensitive subjects (13-42 points, pre-registered) exceed those of an oracle routing by realized accuracy gain (1-11, post hoc). Per-conversation stickiness, per-user budget bands, and pooled parity fail, the last as categories' shifts differ in size or sign. A post hoc exact per-user rate hides only even-prefix strong counts and forfeits most self-assessed routing value; it preserves odd-position decisions, from which a post hoc log attack reaches AUC 0.73 after 20 RouteLLM requests (exploratory).
cs.CR / 23 / 2610.10010
Defining Purpose-Limited Secrets
Abstract
A cryptographic secret is issued for a purpose but grants a capability, and the capability is usually larger: a decryption key meant for computing aggregates can read every record, and a token meant to pay one invoice can drain the account. Deployed systems state the purpose in policy and enforce only the capability. We make the purpose a property of the secret. A mechanism sees operations, not reasons, so a scheme enforces an admissible set $A(P)$ of operations that stands in for a declared intended use $I$; whether $I$ captures the human purpose is a modeling obligation that marks where policy takes over. Against the same $I$ we define three regimes: under confinement an operation outside the intended use is infeasible, under mediation a trusted component refuses it, and under accountability use beyond a stated bound is possible but attributed to the holder. The confinement games, for unpredictability and indistinguishability, coincide with the key-query forms of functional-signature unforgeability and constrained-PRF pseudorandomness and correspond to simulation-secure functional encryption within its feasibility boundary; stated against $I$, they also register a predicate that permits too much. A payment token uses all three regimes on one secret: it pays one capped charge at one merchant, can be narrowed but not widened, is checked by the merchant, and identifies the withdrawing account if spent twice. Attenuation unforgeability reduces to unforgeable signatures and a collision-resistant hash without random oracles; traceability and non-frameability hold under discrete log and unforgeable signatures in the random-oracle model. We claim no new primitive, hardness assumption, or general composition theorem; the framework is a common specification against which such primitives are stated and compared.
cs.CR / 24 / 2610.10075
BetweenCut: Private Heavy-Node Classification with Doubly Logarithmic Error in Tree Height
Abstract
Finding heavy nodes in a tree---those whose counts exceed a given threshold---is a building block for analysis and learning over structured data. Achieving record-level differential privacy (DP) without sacrificing accuracy is challenging because each record contributes to counts along an entire root-to-leaf path, allowing privacy costs to accumulate across levels. Existing methods account for the multiple threshold comparisons for each record incur additive error margins of $Ω_{\varepsilon,δ}(\log h)$ or $Ω_{\varepsilon,δ}(\sqrt{\log h})$ for tree height $h$. We introduce \textsc{BetweenCut}, an $(\varepsilon,δ)$-DP algorithm with an additive error margin of $O_{\varepsilon,δ}(\log\log h)$, improving the existing bounds for deep trees. This error holds simultaneously for all nodes and is independent of the input database size.
cs.CR / 25 / 2610.10149
Pump-and-Dump meets Honeypot Tokens: Detection and Analysis of Telegram Bait-and-Trap Schemes
Abstract
While Pump-and-Dump schemes have been extensively studied on centralized exchanges (CEXs), how they operate in decentralized exchanges (DEXs) remains largely unexplored. We monitor 83 Telegram channels used to coordinate Pump-and-Dump campaigns and collect 3,677 events across the BNB Smart Chain and Ethereum. Our analysis reveals that, although these operations superficially resemble CEX-based Pump-and-Dump schemes, their underlying mechanism is fundamentally different. Rather than manipulating prices, organizers orchestrate deceptive Pump-and-Dump campaigns around honeypot tokens whose smart contracts allow Telegram subscribers to buy but prevent them from selling. Unaware of this restriction, victims purchase these tokens and are unable to recover their funds. We refer to this new fraud mechanism as Bait-and-Trap. We show that Bait-and-Trap provides a substantially more reliable profit strategy than Pump-and-Dump schemes: organizers gain in 99.3% of cases, extracting over $7 million in profit. To counter this threat, we develop a transaction-simulation tool that detects honeypot tokens before purchase by testing whether a user can successfully sell against the token's live contract state, providing a practical defense against Bait-and-Trap operations.
cs.CR / 26 / 2610.10176
Collusion-Secure Semi-Quantum Secret Sharing Scheme using a Quantum Third Party
Abstract
The security of a secret sharing protocol is compromised if one or more dishonest participants can reconstruct the secret by deviating from the prescribed protocol, particularly when such malicious behavior remains undetected. Semi-Quantum Secret Sharing (SQSS) is a variant of secret sharing in which the participants possess only classical capabilities, such as preparing and measuring qubits in the computational ($Z$) basis, while a quantum-capable third party assists the dealer in generating and distributing quantum shares of a classical secret. The limited quantum capabilities of the participants in SQSS protocols may introduce vulnerabilities to collusion attacks, enabling an assisting quantum third party, in collaboration with a dishonest classical participant, to jointly recover the secret even without detection. In this work, we propose a novel SQSS protocol that eliminates this vulnerability and achieves information-theoretic security against collusion attacks. We further prove that the proposed protocol is secure against a broad class of external and internal attacks. Finally, a comparative analysis demonstrates that our construction advances existing SQSS protocols by simultaneously optimizing qubit efficiency, involves a classical dealer, and resilience against DCNA attacks using basic Bell states.
cs.CR / 27 / 2610.10360
Receiver-Domain Behavioral Probing for Backdoor-Resilient Federated GPS Spoofing Detection in UAV Networks
Abstract
Federated learning lets a UAV fleet train a shared GPS spoofing detector without raw receiver data leaving any aircraft, and several recent UAV-FL designs weight each client by the validation accuracy it reports about itself. We show this self-report is an exploitable attack lever: two compromised clients of ten that poison part of their data, scale their updates, and inflate their reported accuracy raise backdoor lift to +0.3036, higher than the same attack achieves without lying. We propose receiver-domain behavioral probing, in which the coordinator evaluates every submitted model on counterfactual spoofed samples built by driving each discriminative GPS feature to a benign value, weighting clients by what their models do rather than what they claim. Under independent and identically distributed clients this reduces attacker-induced lift to -0.0265, statistically indistinguishable from an honest fleet, while flagging the compromised aircraft. Unlike Byzantine-robust aggregation it needs no exact attacker count, only an honest majority: when the true count exceeds the configured value, Multi-Krum degrades from +0.0061 to +0.2837 while ours stays near baseline. Evaluation uses one public single-receiver dataset partitioned into simulated clients; we also report where the mechanism fails, under strong client heterogeneity.
cs.CR / 28 / 2610.10333
The Economic Security of Exponential EIP-1559
Abstract
We consider the problem of parameter selection for EIP-1559, a widely used gas pricing mechanism for blockchains. As opposed to previous work, we aim to achieve economic security, where the chosen parameters guarantee a lower bound on total collected gas fees whenever usage over a given time window exceeds a specified threshold. This bound can be set prohibitively high, thereby economically deterring usage above a desired level. Such guarantees are particularly relevant to long-term objectives such as limiting state growth. We apply our approach to the pure exponential version of EIP-1559 which Ethereum uses and to a variant used by Robinhood Chain and Arbitrum. To achieve economic security, we first characterize the revenue-minimizing gas-usage distribution for the exponential version and for the other variant we construct a distribution whose associated fee revenue provably approximates the minimum. Using these results, we then demonstrate how to set economically secure mechanism parameters for both variants and we discuss the associated tradeoffs.
cs.CR / 29 / 2610.10002
The Price of Privacy: Randomness Complexity of Graph-Based Multi-Secret Sharing
Abstract
We study the randomness required to share possibly correlated secret bits among parties connected by a graph. A dealer places shares on the edges so that each party can recover its own secret from its incident shares and learn nothing about the others beyond what its own secret reveals. Anilkumar et al. completely determined the minimum randomness required for three binary secrets. We extend this study to four secrets and obtain results for arbitrary numbers of secrets on general graphs. For four parties on a complete graph, we determine the minimum number of random states for every set of permitted secret combinations: the possible values are one, two, three and four. We also characterize when one random bit suffices on an arbitrary graph.
cs.CR / 30 / 2610.09717
Pareto-optimal quantum kernel selection for unsupervised anomaly detection on real malware beaconing data
Abstract
Quantum kernel methods are leading candidates for a practical quantum advantage in machine learning, but assessing that potential requires two quantities usually reported separately: how well a kernel performs on the task, and how far its geometry departs from the classical kernels available for the same problem. We introduce a fully unsupervised, multi-objective protocol that optimises simultaneously the normalised pseudo discrepancy (NPD), a label-free proxy for anomaly detection quality, and the geometric difference (GD) to a tuned classical reference kernel, selecting models from the resulting Pareto front. We apply it to malware beaconing detection in real network traffic, using a one-class support vector machine with fidelity and projected quantum kernels over four data encodings, on simulators and on IQM's 20-qubit Garnet processor. NPD-guided selection alone finds a fidelity kernel that beats the tuned classical baseline, but with a geometric difference too small to certify the gain as quantum. Projected kernels reach far larger geometric differences; the Pareto-selected one only marginally exceeds the baseline (AUC $0.782$ versus $0.765$, $g_{C\to Q}\approx 89>\sqrt{N}$ relative to that reference kernel), still below the NPD-selected fidelity kernel ($0.840$).