← Back to Index
Daily Research Digest

arXiv Papers

2026-09-28
409
Papers
9
Categories
81
Translated
收藏清单 0
精选 · Favorites
81
cs.AI / 1 / 2609.30456
Spectral Feedback for Test-Time Alignment of Protein Diffusion Models
用于蛋白质扩散模型测试时对齐的谱反馈
Shai Dickman, Mert Cemri, Landon Butler, Kannan Ramchandran
cs.AI
diffusion
扩散模型相关
Abstract
Reward maximization alignment methods for discrete diffusion models have primarily focused on steering the reverse process, either by influencing token logits or by selecting favorable sequences at intermediate steps. These approaches largely treat inference as a unidirectional process, lacking mechanisms for revisiting undesirable token selections. We introduce Spectral Feedback, an algorithm that selects edit-positions in a feedback loop, allowing the model to iteratively correct its own generations. This approach leverages the mask structure of discrete diffusion models by re-masking and re-sampling tokens, analogous to image editing methods that reintroduce noisy latents and re-run the reverse process. While prior alignment methods focus on what token labels to assign to maximize a target reward, we instead treat which tokens to revisit as the central alignment problem. Selecting edit-positions is challenging because edit effects are interdependent: the impact of modifying one token depends on which others are edited simultaneously. We define an edit-set as a set of token positions to re-mask and re-sample. Motivated by prior work on sparse interactions in biological systems, we find empirically that edit-set value functions for protein inverse folding admit sparse Fourier representations. This structure enables Spectral Feedback to efficiently learn and optimize the value functions for edit-position selection. Spectral Feedback is model-agnostic and can be applied to pretrained, test-time aligned, and fine-tuned diffusion models. For all of these models, the algorithm improves alignment performance without modifying the underlying generative process. Applied to inverse folding with a protein stability reward oracle, it achieves a 32.3% increase in stable proteins for a pretrained model, 24.8% for Best-of-10, and 5.8% for a state-of-the-art RL fine-tuned diffusion model.
Chinese Translation
针对离散扩散模型的奖励最大化对齐方法主要集中于引导反向过程,要么通过影响 token logits,要么通过在中间步骤选择有利序列。这些方法在很大程度上将推理视为单向过程,缺乏用于重新审视不理想的 token 选择的机制。我们提出谱反馈(Spectral Feedback),一种在反馈回路中选择编辑位置的算法,使模型能够迭代地修正其自身生成结果。该方法通过重新掩码和重新采样 token 来利用离散扩散模型的掩码结构,类似于重新引入噪声潜变量并重新运行反向过程的图像编辑方法。虽然先前的对齐方法关注分配什么 token 标签以最大化目标奖励,我们则将重新审视哪些 token 视为核心对齐问题。选择编辑位置具有挑战性,因为编辑效果相互依赖:修改一个 token 的影响取决于同时编辑了哪些其他 token。我们将编辑集(edit-set)定义为要重新掩码和重新采样的一组 token 位置。受关于生物系统中稀疏相互作用的先前工作启发,我们通过实证发现,蛋白质逆折叠的编辑集价值函数具有稀疏傅里叶表示。这种结构使谱反馈能够高效地学习和优化用于编辑位置选择的价值函数。谱反馈与模型无关,可应用于预训练的、测试时对齐的以及微调的扩散模型。对于所有这些模型,该算法在不修改底层生成过程的情况下提升了对齐性能。应用于带有蛋白质稳定性奖励预言机的逆折叠时,对于预训练模型,它使稳定蛋白质增加 32.3%,对于 Best-of-10 增加 24.8%,对于最先进的 RL 微调扩散模型增加 5.8%。
cs.AI / 2 / 2609.30484
Do LLMs Understand Context? A Knowledge Graph-Based Evaluation Framework
LLMs 是否理解上下文?一个基于知识图谱的评估框架
Subavarshana Arumugam, Mamta Nallaretnam, Kithuni Wickramasinghe, Chamath Gunapala, Pragatheeswaran Vipulanandan, Kamal Premaratne, Uthayasanker Thayasivam
cs.AI · cs.LG
large language model
大语言模型相关
Abstract
While large language models (LLMs) have achieved remarkable linguistic capabilities, a profound question lingers at their core: do these models truly comprehend context or simply excel at pattern matching on an unprecedented scale? Contextual understanding in LLMs refers to the ability to correctly extract relevant information from a given context, integrate it into a coherent internal representation, and reason over it to produce factually consistent and contextually grounded responses. However, traditional methods such as BiLingual Evaluation Understudy (BLEU) and perplexity simply measure surface-level performance. This reveals a critical gap in question answering (QA), where responses must be contextually grounded rather than simply being memorized associations. To fill this void, we propose a novel knowledge graph (KG) based evaluation framework for LLM contextual understanding in QA. Central to this is Semantic Structural Similarity for KGs (S3KG), a hybrid similarity measure combining structural and semantic signals into a single score. In addition, a diagnostic analysis framework is developed to identify and categorize reasoning errors at the triplet level, enabling fine-grained analysis of model failures. Together, across nine benchmarks, S3KG achieves F1 gains of up to $+7.6$ points over the strongest baseline and AUROC up to $0.973$.
Chinese Translation
尽管大型语言模型(LLMs)已取得显著的语言能力,但一个深刻的问题仍萦绕在其核心:这些模型是真正理解了上下文,还是仅仅擅长以前所未有的规模进行模式匹配?LLMs 中的上下文理解指的是:能够从给定上下文中正确提取相关信息,将其整合为连贯的内部表示,并对其进行推理,以产生事实一致且基于上下文的回应。然而,诸如双语评估替补(BiLingual Evaluation Understudy, BLEU)和困惑度等传统方法仅仅衡量表层性能。这揭示了问答(QA)中的一个关键缺口:其中的回答必须基于上下文,而不是仅仅是被记住的关联。为填补这一空白,我们提出了一种新颖的基于知识图谱(KG)的评估框架,用于 QA 中的 LLM 上下文理解。其核心是面向 KG 的语义结构相似度(Semantic Structural Similarity for KGs, S3KG),这是一种将结构信号和语义信号结合为单一分数的混合相似度度量。此外,还开发了一个诊断分析框架,用于在三元组层面识别并分类推理错误,从而能够对模型失败进行细粒度分析。综合而言,在九个基准上,S3KG 相较于最强基线实现了最高 $+7.6$ 个点的 F1 提升,AUROC 最高达到 $0.973$。
cs.AI / 3 / 2609.30489
BioEVAL: A global, multi-institutional benchmark of large language and multimodal models for bioengineering
BioEVAL:一个面向生物工程的全球性、多机构大型语言与多模态模型基准
Shun Ye, Vinny Chandran Suja, Chenlong Li, Chongming Jiang, Reza Zamani, Xiang Li, Christopher Bain, Yuqi Zhou, Walker Peterson, Huidong Wang, Chenglang Hu, Jongchan Park, Xiao Cheng, Benjamin Swedlund, Sandra Murillo, Anjali Sivanandan, Shiyu Sun, Liang Lanfeng, Mohammad Tariqul Islam, Baju C. Joy, Ishaq N. Khan, Sreedhar S. Kumar, Gabriel Mercado-Vásquez, James V. Vizzard, Jonathan M. Matthews, Helen Huang, Xiaolu Guo, Ethan Nicklow, Guorui Chen, Ryan A. Neff, Surjendu Maity, Hyeonjin Park, Han-ho Joo, Katherine Dong, Yuyan Cai, Weihang Huang, Yichen Zou, Rui Yan, Raphael Figueroa, Artem Goncharov, Bella Rose Schremmer, Lian Elsa Linton, Keisuke Goda, Liang Gao, Ke Cheng, Leonardo Morsut, Jennifer L. Wilson, Jianping Fu, Lim Chwee Teck, Deblina Sarkar, Andreas Hierlemann, Savaş Tay, Alexander Hoffmann, Donald Richieri Griffin, Jun Chen, Shana O. Kelley, Shyni Varghese, Jinwoo Cheon, Wilbur A. Lam, James J. Moon, Wilson W. Wong, Samir Mitragotri, Dino Di Carlo
cs.AI
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) have demonstrated historic breakthroughs in general reasoning with early successes in biomedical science. However, existing LLM benchmarking emphasizes factual recall, offering limited insight into model performance on frontier and multimodal tasks. We assembled BioEVAL (BioEngineering Validation of AI and LLMs), a global, multi-institutional initiative designed to assess experimental reasoning capability across bioengineering (BE) subfields. BioEVAL spans 11 major BE subfields plus a set of uncategorized items, bringing together 22 research groups to create a PhD-level benchmark comprising 608 evaluation items: 1) 380 multiple-choice questions (MCQs, 359 retained after audit), 2) 218 literature synthesis tasks, and 3) 10 multimodal problems with experimental image interpretation. Benchmark items underwent authoring-group expert review and centralized quality control before evaluation. Following evaluation, a blinded cross-group consensus audit of the highest- and lowest-accuracy MCQ items flagged 21 questions for revision or removal; these were withheld, and all reported MCQ results are computed on the 359 retained items. We evaluated diverse cloud-scale foundation/multimodal models (e.g., ChatGPT, Gemini, and Grok) and locally deployable models suitable for inference on consumer-grade GPUs. Models achieved the highest accuracy of up to 90% on MCQs, similarity score of 0.72 on literature synthesis, and accuracy of 80% on a small sample of multimodal reasoning questions, with substantial performance variation across subfields. Leaderboard rankings characterize current capabilities, limitations, and development priorities across the evaluated BE task categories. BioEVAL is maintained as an extensible benchmark with standardized protocols for continuing expert item contribution and model evaluation.
Chinese Translation
大型语言模型(LLMs)在通用推理方面已展现出历史性突破,并在生物医学科学中取得早期成功。然而,现有 LLM 基准测试强调事实回忆,对模型在前沿与多模态任务上的表现提供的洞见有限。我们构建了 BioEVAL(BioEngineering Validation of AI and LLMs,即人工智能与 LLMs 的生物工程验证),这是一项全球性、多机构倡议,旨在评估跨生物工程(BE)子领域的实验推理能力。BioEVAL 涵盖 11 个主要 BE 子领域以及一组未分类条目,汇集 22 个研究组,创建了一个博士水平的基准,包含 608 个评估条目:1)380 道多项选择题(MCQs,审计后保留 359 道),2)218 项文献综合任务,以及 3)10 个带有实验图像解读的多模态问题。基准条目在评估前经过作者组专家评审和集中式质量控制。评估之后,对准确率最高和最低的 MCQ 条目进行了盲法跨组共识审计,标记出 21 道题以供修订或删除;这些题目被暂扣,所有报告的 MCQ 结果均在 359 个保留条目上计算。我们评估了多种云规模基础/多模态模型(例如 ChatGPT、Gemini 和 Grok)以及适合在消费级 GPU 上推理的本地可部署模型。模型在 MCQs 上达到了最高达 90% 的准确率,在文献综合上取得了 0.72 的相似度分数,并在一个小样本的多模态推理问题上取得了 80% 的准确率,且不同子领域间存在显著性能差异。排行榜排名刻画了所评估 BE 任务类别中当前的能力、局限性和发展优先级。BioEVAL 作为一个可扩展基准进行维护,并配有标准化协议,用于持续进行专家条目贡献和模型评估。
cs.AI / 4 / 2609.30576
T-RoPE: Time-Aware Rotary Position Embedding for Sequential Recommendation
T-RoPE:面向序列推荐的时间感知旋转位置嵌入
Yang Liu, Noel Loo, Ali Khanafer, Shuying Sun, Akshay Soni, Zhong Wu, Linjun Yang
cs.AI · cs.IR · cs.LG
large language model
大语言模型相关
Abstract
Large-scale recommenders increasingly adopt the sequential generative recipe behind large language models, bringing the Transformer into recommendation along with design choices made for text, including Rotary Position Embedding (RoPE). In language models, RoPE encodes token indices for relative position reasoning, but in recommendation, an interaction index records only event order, saying nothing about elapsed time, behavioral cycles across scales, or calendar phase. We revisit this choice and propose T-RoPE, a time-aware RoPE for sequential generative recommendation that replaces index-only rotation with timestamp-based angles, learnable temporal coefficients, multiscale frequency banks, shifted query alignment, and non-stationary key rotation. We prove that standard RoPE, even on timestamps, remains time-translation invariant and cannot distinguish seasonal contexts, and that T-RoPE breaks this invariance while preserving the RoPE interface. Across five public benchmarks, T-RoPE achieves the best result on every metric on every dataset, improving over the strongest baseline by 78--130\% in HR@10 on the sparse PixelRec data and 8--12\% across metrics on Amazon Books. On an industrial-scale e-commerce dataset with more than 6B interactions, it improves every metric over the HSTU + Time RAB backbone by 13--82\%, with ablations attributing the largest gains to multiscale frequencies ($+56\%$ NDCG@50) and non-stationary keys ($+4\%$). An online A/B test in the Shop app yields positive lifts in conversion rate ($+0.33\%$) and order count ($+0.63\%$). We also provide forward and backward algorithms whose added cost is linear in sequence length and head dimension, keeping time-aware RoPE practical for large generative recommenders.
Chinese Translation
大规模推荐系统越来越多地采用大语言模型背后的序列生成式范式,将 Transformer 引入推荐领域,并同时引入了为文本所作的设计选择,包括旋转位置嵌入(RoPE)。在语言模型中,RoPE 对 token 索引进行编码以进行相对位置推理,但在推荐中,交互索引只记录事件顺序,对经过时间、跨尺度的行为周期或日历相位一无所知。我们重新审视这一选择,并提出 T-RoPE,一种用于序列生成式推荐的时间感知 RoPE,它用基于时间戳的角度、可学习的时间系数、多尺度频率库、偏移查询对齐和非平稳键旋转来替代仅基于索引的旋转。我们证明,标准 RoPE 即使在时间戳上仍保持时间平移不变性,并且无法区分季节性上下文;而 T-RoPE 在保留 RoPE 接口的同时打破了这种不变性。在五个公开基准上,T-RoPE 在每个数据集的每个指标上都取得了最佳结果,在稀疏的 PixelRec 数据上 HR@10 相比最强基线提升 78--130\%,在 Amazon Books 上各指标提升 8--12\%。在一个具有超过 60 亿次交互的工业级电子商务数据集上,它相比 HSTU + Time RAB 主干将每个指标提升了 13--82\%,消融实验将最大增益归因于多尺度频率($+56\%$ NDCG@50)和非平稳键($+4\%$)。在 Shop 应用中进行的一项在线 A/B 测试在转化率($+0.33\%$)和订单数($+0.63\%$)上带来了正向提升。我们还提供了前向和反向算法,其额外成本与序列长度和头维度呈线性关系,使时间感知 RoPE 对大型生成式推荐系统而言保持实用。
cs.AI / 5 / 2609.30625
Audio LLMs Know When They Can't Hear You
音频大语言模型知道自己何时听不见你
Amirhosein Javadi, Richa Dixit, Mehrdad Farajtabar, Minsik Cho, Devang Naik, Mohammad Samragh
cs.AI
large language model
大语言模型相关
Abstract
Audio large language models allow users to interact with the model through speech. When an input recording is too degraded, the model may misinterpret the user's query and respond based on an incorrect transcription. In this paper, we study model-conditional transcription reliability: whether an Audio LLM can recognize when its own transcription is unreliable. We first prompt the Audio LLM to assess whether its own transcription would be reliable, and find that the model is a poor judge of its own transcription reliability: in most cases, it predicts that its transcription will be reliable. We find that existing approaches, including speech quality predictors, audio LLM generation uncertainty, and transcript-conditioned WER estimation, provide limited signals for detecting transcription failures. In contrast, we discover that transcription reliability is strongly represented in the model's audio-encoder representations. Based on this observation, we devise a lightweight reliability predictor that operates on representations extracted by the frozen audio encoder and predicts the reliability class before generation. The reliability predictor can trigger a clarification request from the user when their voice query is predicted to be unreliable, while allowing reliable queries to proceed without modifying the underlying Audio LLM. Our predictor achieves 81.10% in-domain and 78.09% cross-domain macro-F1 scores, outperforming the strongest baselines by 10.33 and 11.93 points, respectively. Finally, we show that reliability labels can transfer across Audio LLM families, and that transfer performance is closely related to the alignment of their model-specific reliability boundaries.
Chinese Translation
音频大语言模型允许用户通过语音与模型进行交互。当输入录音退化过于严重时,模型可能会误解用户的查询,并基于错误的转录进行回应。在本文中,我们研究模型条件下的转录可靠性:音频 LLM 能否识别出其自身的转录在何时不可靠。我们首先提示音频 LLM 评估其自身转录是否会是可靠的,并发现该模型对自身转录可靠性的判断能力很差:在大多数情况下,它预测其转录将会是可靠的。我们发现,现有方法,包括语音质量预测器、音频 LLM 生成不确定性以及转录条件化的 WER 估计,为检测转录失败提供的信号有限。相反,我们发现转录可靠性在模型的音频编码器表示中被强烈地表示出来。基于这一观察,我们设计了一个轻量级可靠性预测器,它作用于由冻结的音频编码器提取的表示,并在生成之前预测可靠性类别。当用户的语音查询被预测为不可靠时,该可靠性预测器可以触发向用户发出澄清请求,同时允许可靠的查询继续进行,而无需修改底层的音频 LLM。我们的预测器分别取得了 81.10% 的域内和 78.09% 的跨域 macro-F1 分数,分别比最强的基线高出 10.33 和 11.93 个百分点。最后,我们表明可靠性标签可以在不同音频 LLM 系列之间迁移,并且迁移性能与其特定于模型的可靠性边界的对齐程度密切相关。
cs.AI / 6 / 2609.30662
LLM Parkinsonism: Executive-Control Failure, Token-Inefficient Persistence, and an Uncertainty-Aware Global Executive Control Architecture for Autonomous Language-Model Agents
LLM帕金森样综合征:执行控制失败、Token低效的持续性,以及一种面向自主语言模型智能体的不确定性感知全局执行控制架构
Dongsheng Xiao, Zeyuan Wang, Xuzhe Xia, Bo Zhao, Yankai Cao
cs.AI
large language model
大语言模型相关
Abstract
Large language models (LLMs) can plan, use tools, write code, and execute long-horizon workflows, yet strong local competence does not guarantee project-level executive control. Agents may continue acting after the original objective is satisfied, producing low-value refinements, repeated verification, and repairs to self-created complexity. We use LLM Parkinsonism as a narrowly defined, non-clinical metaphor for this pattern of persistent action despite diminishing task-level value. We argue that the problem is not explained by autoregressive next-token prediction alone, but more directly by concentrating proposal generation, scope interpretation, progress assessment, and stopping authority within the same self-conditioned loop. We therefore introduce Global Executive Control (GEC) v0.2, an uncertainty-aware governance architecture that separates action generation from project-level control. In a 24,000-episode matched-candidate benchmark under a common 40,000-token ceiling, a first-candidate baseline achieved 67.42% hard-goal success, a candidate-set local control achieved 96.53%, and GEC achieved 96.57%. The candidate-set control shows that access to multiple candidate actions explains most of the success gain; relative to that control, GEC preserved success while reducing mean token use from 19,782 to 12,574 (36.4%) and restricted mean tokens to completion at the 40,000-token ceiling from 16,136 to 13,114 (18.7%), while eliminating measured pre-completion drift and sharply reducing gross complexity. Governance-overhead sensitivity remained favorable through an additional 500 synthetic governance tokens per cycle. These mechanistic simulations support explicit governance of scope, evidence, resource use, and stopping, while live-model validation remains necessary.
Chinese Translation
大型语言模型(LLMs)能够进行规划、使用工具、编写代码并执行长时程工作流,然而强大的局部能力并不保证项目级执行控制。智能体可能在原始目标已经满足之后继续行动,产生低价值细化、重复验证以及对自造复杂性的修复。我们将LLM帕金森样综合征作为一个狭义定义的、非临床的隐喻,用来指代这种尽管任务级价值递减却仍持续行动的模式。我们认为,该问题不能仅用自回归下一Token预测来解释,而更直接的原因在于将提案生成、范围解释、进展评估和停止权限集中于同一个自条件化循环之内。因此,我们引入全局执行控制(GEC)v0.2,这是一种不确定性感知的治理架构,将动作生成与项目级控制分离。在一个共同的40,000-Token上限下、包含24,000回合的匹配候选基准中,首个候选基线实现了67.42%的硬目标成功率,候选集局部控制实现了96.53%,而GEC实现了96.57%。候选集控制表明,能够访问多个候选动作解释了大部分成功增益;相对于该控制,GEC在保持成功的同时,将平均Token使用量从19,782降至12,574(36.4%),并将40,000-Token上限下完成所需的受限平均Token数从16,136降至13,114(18.7%),同时消除了测量到的完成前漂移,并大幅降低了总体复杂性。治理开销敏感性在每周期额外增加500个合成治理Token的情况下仍保持有利。这些机制性模拟支持对范围、证据、资源使用和停止进行显式治理,而真实模型验证仍然必要。
cs.AI / 7 / 2609.30705
The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?
思考的代价:测试时推理在 LLM 交易中能获得回报吗?
Jiayi Chen, Guiling Wang
cs.AI
large language model
大语言模型相关
Abstract
While inference-time reasoning in large language models (LLMs) promises better decision making, its higher computational cost may not yield better economic outcomes. Yet reasoning controls are rarely evaluated as economic interventions, where changes in model outputs must translate into better portfolios after trading costs. We conduct a controlled study of representative LLMs from the DeepSeek, GPT, and Gemini families. We vary reasoning effort while holding information available at each formation date, prompts, output formats, and portfolio construction fixed. Our evaluation covers a full year of U.S. equities under three input conditions: numerical, identifiable news, and masked news. It includes more than 800,000 asset predictions and repeated model generations. Across all three model families, additional reasoning does not produce a reliable improvement in net portfolio returns. For DeepSeek, where we examine the full progression from no reasoning to maximum reasoning, performance is nonmonotonic. Repeated generations also produce unstable treatment effects and portfolio selections, even when overall scores remain similar. These findings show that additional reasoning can change financial decisions without reliably improving their economic value, motivating validation for each task before deployment.
Chinese Translation
尽管大型语言模型(LLM)中的推理时推理有望带来更好的决策,但其更高的计算成本未必能带来更好的经济结果。然而,推理控制很少被作为经济干预来评估;在这类评估中,模型输出的变化必须能在扣除交易成本后转化为更好的投资组合。我们对来自 DeepSeek、GPT 和 Gemini 系列的代表性 LLM 进行了受控研究。我们改变推理努力程度,同时固定每个形成日期可获得的信息、提示词、输出格式和投资组合构建方式。我们的评估涵盖整整一年的美国股票,在三种输入条件下进行:数值、可识别新闻和掩码新闻。它包含超过 800,000 个资产预测和重复的模型生成。在所有三个模型系列中,额外的推理并未在净投资组合收益上产生可靠的改进。对于 DeepSeek,我们考察了从无推理到最大推理的完整过程,其表现是非单调的。重复生成也会产生不稳定的处理效应和投资组合选择,即使总体得分保持相近。这些发现表明,额外的推理可以改变金融决策,却未必能可靠地提升其经济价值,这促使人们在部署前对每项任务进行验证。
cs.AI / 8 / 2609.30749
ORCA: Evaluating LLMs on Data Science Code Translation
ORCA:评估大语言模型在数据科学代码翻译上的表现
Xiaolong Li, Jinyang Li, Bowen Qin, Ge Qu, Nan Huo, Xiaohan Xu, Shipei Lin, Reynold Cheng
cs.AI
large language model
大语言模型相关
Abstract
Data Science Code Translation (DSCT) is the process of converting code between data science libraries while preserving functional equivalence and enabling interoperability across data science ecosystems. While Large Language Models (LLMs) have demonstrated considerable progress in Data Science Code Generation (DSCG), their performance in DSCT remains insufficiently studied. To address this gap, we introduce ORCA, a comprehensive benchmark with two complementary settings: ORCA-MAIN, which comprises 1,600 carefully curated grounding-level tasks across 3 representative domains: Data Querying, Data Manipulation, and Deep Learning; and ORCA-PROJECT, which contains 200 translation tasks over complete data science projects across 7 data science task types. Each task is accompanied by annotated reference translations and test cases for validating functional equivalence. We further incorporate a multi-stage quality verification process that thoroughly verifies task correctness and test case robustness. Experimental results demonstrate challenges in DSCT, with even frontier LLMs showing limited performance. Specifically, Claude-Opus-4.6 achieves a success rate of 56.92% on ORCA-MAIN and 33.67% on ORCA-PROJECT, indicating considerable room for improvement in DSCT. We also observe a clear directional preference in DSCT, where translation is consistently easier when the source code expresses the task through more explicit, fine-grained operations. Motivated by this, we propose an intent-augmented method, in which the model first infers source-code intent and then uses it as additional context for translation, achieving average absolute success-rate gains of 4.80% and 5.33% on ORCA-MAIN and ORCA-PROJECT, respectively.
Chinese Translation
数据科学代码翻译(DSCT)是在数据科学库之间转换代码的过程,同时保持功能等价性并实现跨数据科学生态系统的互操作性。尽管大语言模型(LLM)在数据科学代码生成(DSCG)方面已展现出可观的进展,但它们在 DSCT 中的表现仍未得到充分研究。为填补这一空白,我们提出 ORCA,一个具有两种互补设置的综合性基准:ORCA-MAIN,包含 1,600 个精心构建的基础级任务,覆盖 3 个代表性领域:数据查询、数据处理和深度学习;以及 ORCA-PROJECT,包含 200 个翻译任务,涉及跨 7 种数据科学任务类型的完整数据科学项目。每个任务都配有标注的参考翻译和用于验证功能等价性的测试用例。我们进一步引入多阶段质量验证流程,全面验证任务的正确性和测试用例的鲁棒性。实验结果揭示了 DSCT 中的挑战,即使是最前沿的 LLM 也表现有限。具体而言,Claude-Opus-4.6 在 ORCA-MAIN 上达到 56.92% 的成功率,在 ORCA-PROJECT 上达到 33.67%,表明 DSCT 仍有相当大的改进空间。我们还观察到 DSCT 中存在明显的方向偏好:当源代码通过更显式、更细粒度的操作来表达任务时,翻译始终更容易。受此启发,我们提出一种意图增强方法,模型首先推断源代码意图,然后将其作为额外上下文用于翻译,在 ORCA-MAIN 和 ORCA-PROJECT 上分别取得 4.80% 和 5.33% 的平均绝对成功率提升。
cs.AI / 9 / 2609.30841
Why Jailbreaks Succeed in Diffusion Language Models: An Energy Landscape Analysis
为何越狱在扩散语言模型中能够成功:一项能量景观分析
Thong Bach, Dung Nguyen, Thao Minh Le, Truyen Tran
cs.AI
diffusionlarge language model
扩散模型相关
大语言模型相关
Abstract
Existing attacks and defenses for diffusion-based large language models (dLLMs) target specific vulnerabilities but lack a shared framework explaining why attacks succeed. We propose one by interpreting safety alignment as shaping the denoising energy landscape: a well-aligned model routes harmful queries toward safe outputs through an energy barrier that separates the two regions. Current jailbreak attacks reduce to two strategies for circumventing this barrier: obscuring the query's safety disposition at initialisation, or intervening mid-trajectory to force the denoising path across the energy barrier. From this perspective and the result that masked diffusion models minimise kinetic energy during denoising, we derive three complementary, training-free detection signals: a step-0 ratio that reads the initial safety disposition from the logit distribution before generation begins, and two trajectory-velocity signals that track kinetic energy in complementary subspaces of the logit space. An attack must either reveal its intent at initialisation or expend kinetic energy to cross the barrier in at least one monitored subspace, so the three signals cover each other's blind spots in the energy budget by construction. Evaluation across three dense dLLMs (LLaDA-8B, LLaDA-1.5, Dream-7B) and a sparse mixture-of-experts dLLM (LLaDA-MoE-7B) confirms this complementarity. In stress tests of known attacks, every configuration that evades detection also fails to produce harmful content, suggesting that the detection and barrier-crossing thresholds are hard to separate.
Chinese Translation
现有的针对基于扩散的大语言模型(dLLMs)的攻击与防御针对的是特定的脆弱性,但缺乏一个解释攻击为何能够成功的共享框架。我们提出这样一个框架,其方式是将安全对齐解释为对去噪能量景观的塑造:一个对齐良好的模型通过一个将两个区域分隔开的能量势垒,把有害查询引导向安全输出。当前的越狱攻击可归结为两种绕过该势垒的策略:在初始化时掩盖查询的安全倾向,或在轨迹中途进行干预以迫使去噪路径跨越该能量势垒。从这一视角以及掩码扩散模型在去噪过程中最小化动能这一结果出发,我们推导出三种互补的、无需训练的检测信号:一个 step-0 比率,它在生成开始之前从 logit 分布中读取初始安全倾向;以及两个轨迹速度信号,它们在 logit 空间的互补子空间中追踪动能。攻击要么必须在初始化时暴露其意图,要么必须消耗动能在至少一个被监控的子空间中跨越势垒,因此这三种信号在构造上就相互覆盖了彼此在能量预算中的盲区。在三个稠密 dLLM(LLaDA-8B、LLaDA-1.5、Dream-7B)以及一个稀疏的专家混合 dLLM(LLaDA-MoE-7B)上的评估证实了这种互补性。在对已知攻击的压力测试中,每一个能逃避检测的配置同样无法生成有害内容,这表明检测阈值与跨越势垒的阈值难以区分。
cs.AI / 10 / 2609.30936
Self-Play Search Distillation for Large Language Model Reasoning
面向大语言模型推理的自对弈搜索蒸馏
Lorenzo Molfetta, Wai-Chung Kwan, Giacomo Frisoni, Luca Ragazzi, Gianluca Moro, Pavlos Vougiouklis, Jeff Z. Pan, Pasquale Minervini
cs.AI
large language model
大语言模型相关
Abstract
Improving reasoning abilities in Large Language Models (LLMs) requires high-quality data that exposes difficult decisions, competing alternatives, and their consequences. Data scarcity is driven by the low quality of synthetic data and the cost of human labeling. We introduce Self-Play Search Distillation (SPSD), a framework for generating superhuman synthetic data via self-play of MuZero-like networks trained on board games. SPSD uses executable environments to turn search into structured reasoning problems. At each state, the expert identifies a preferred decision, plausible alternatives, plausible opponent replies, and value estimates. By converting the self-play search records into superhuman chains-of-thought, we train LLMs with environment-grounded supervision. Although trained only on self-play search records, SPSD transfers to unseen mathematics. On Qwen3-4B-Base, it raises the mean over six mathematics benchmarks from 24.1 to 36.6 while increasing the held-out-game win rate from 15% to 45%. SPSD offers an annotation-efficient way to create high-quality synthetic data for improving LLM performance in reasoning tasks.
Chinese Translation
提升大语言模型(LLM)的推理能力,需要能够揭示困难决策、相互竞争的备选方案及其后果的高质量数据。数据稀缺的根源在于合成数据质量低下以及人工标注的成本。我们提出自对弈搜索蒸馏(SPSD),这是一个通过在棋类游戏上训练的类 MuZero 网络进行自对弈来生成超人类合成数据的框架。SPSD 利用可执行环境将搜索转化为结构化的推理问题。在每个状态,专家会确定一个首选决策、若干合理的备选方案、若干合理的对手回应以及价值估计。通过将自对弈搜索记录转换为超人类的思维链,我们以基于环境的监督来训练 LLM。尽管仅在自对弈搜索记录上训练,SPSD 仍能迁移到未见过的数学任务上。在 Qwen3-4B-Base 上,它将六个数学基准的平均成绩从 24.1 提升至 36.6,同时将留出对局的胜率从 15% 提升至 45%。SPSD 提供了一种标注高效的方法来创建高质量合成数据,以提升 LLM 在推理任务中的表现。
cs.AI / 11 / 2609.30939
MACBT: A Multi-Agent Cognitive Behavioral Therapy Decision Support System with Longitudinal Memory
MACBT:一种具有纵向记忆的多智能体认知行为疗法决策支持系统
De Jiang, Shuo Zhang, Weiwei Liao, Jianying Zhang, Chuanhui Yu, Hongen Liao, Kehong Yuan
cs.AI
large language model
大语言模型相关
Abstract
Cognitive behavioral therapy (CBT) is an evidence-based first-line treatment for depression, yet its scale is constrained by the time clinicians spend on pre-session preparation, post-session documentation, and longitudinal cognitive-pathology tracking. We present a clinician-facing AI decision-support system that combines a multi-agent CBT framework (MACBT) with a CBT-specific longitudinal memory module (CD Memory). MACBT encodes the five-stage CBT workflow (assessment, Socratic questioning, cognitive restructuring, behavioral experiments, and treatment monitoring) into five collaborative agents. CD Memory tracks cognitive-distortion type, frequency, severity, and restructuring efficacy across sessions to generate pre-session pathology reports and intervention-priority recommendations. We construct a Chinese CBT dialogue corpus via dual-role large language model simulation and train a Qwen3-14B backbone with supervised fine-tuning and direct preference optimization. Evaluation with GPT-4 judges shows MACBT outperforms MeChat, SoulChat, PsyChat, and CPsyCounX in professionalism (2.62) and clinical authenticity (2.25). The full memory-augmented system further improves session quality by 12.6% and achieves a longitudinal mean of 2.29 on cross-session continuity, intervention progression, and personalization.
Chinese Translation
认知行为疗法(CBT)是治疗抑郁症的循证一线疗法,然而其规模化受到临床医生在会前准备、会后记录以及纵向认知病理追踪上所花费时间的制约。我们提出了一种面向临床医生的 AI 决策支持系统,该系统将多智能体 CBT 框架(MACBT)与 CBT 专用的纵向记忆模块(CD Memory)相结合。MACBT 将五阶段 CBT 工作流程(评估、苏格拉底式提问、认知重构、行为实验和治疗监测)编码为五个协作智能体。CD Memory 跨会话追踪认知扭曲类型、频率、严重程度和重构效果,以生成会前病理报告和干预优先级建议。我们通过双角色大语言模型模拟构建了一个中文 CBT 对话语料库,并通过监督微调和直接偏好优化训练了一个 Qwen3-14B 主干模型。使用 GPT-4 评判员的评估表明,MACBT 在专业性(2.62)和临床真实性(2.25)上优于 MeChat、SoulChat、PsyChat 和 CPsyCounX。完整的记忆增强系统进一步将会话质量提升了 12.6%,并在跨会话连续性、干预进展和个性化方面达到了 2.29 的纵向平均值。
cs.AI / 12 / 2609.30940
Financial Fragility in Societies of LLM Agents: Coordination Failures and Stabilizing Mechanisms
LLM 智能体社会中的金融脆弱性:协调失败与稳定机制
Zhenhao Fu, Ruipeng Xu, Qibing Ren
cs.AI · q-fin.GN
large language model
大语言模型相关
Abstract
Individually protective decisions can produce avoidable collective failures. As large language model (LLM) agents take on greater roles in financial decision-making, financial AI safety must therefore be considered not only at the level of individual agents, but also at the level of the systems they jointly create. We study this problem with FRAIL, a controlled experimental framework that places LLM agents in three dynamic financial environments---bank runs, debt rollover, and reward crowdfunding---where agents' decisions reshape the financial conditions faced by others. Across seven leading LLMs, we find widespread collective fragility even when no agent is instructed to destabilize the system: 77\% of baseline bank-run episodes and 83\% of debt-rollover episodes end in failure. We then compare three interaction mechanisms based on compensated commitments, centralized commitment agreements, and participant-led coalitions. All three improve aggregate outcomes, but no single mechanism performs best across all financial structures. Across mechanisms, successful stabilization shares a common temporal pattern: broad commitment forms early, before defensive behavior becomes self-reinforcing. Our findings show that individually capable agents do not automatically form safe financial systems, highlighting system-level evaluation and interaction design as central problems for financial AI safety. Code is available at https://anonymous.4open.science/r/FinFrail-CF26.
Chinese Translation
个体保护性决策可能产生本可避免的集体失败。随着大型语言模型(LLM)智能体在金融决策中承担更大的角色,金融 AI 安全因此不仅必须在个体智能体层面加以考虑,还必须在它们共同创造的系统层面加以考虑。我们使用 FRAIL 研究这一问题,这是一个受控实验框架,将 LLM 智能体置于三种动态金融环境中——银行挤兑、债务展期和奖励众筹——在这些环境中,智能体的决策会重塑其他主体所面临的金融条件。在七个领先的 LLM 中,我们发现,即使没有智能体被指示去破坏系统稳定,集体脆弱性也普遍存在:77\% 的基线银行挤兑回合和 83\% 的债务展期回合以失败告终。然后,我们比较三种互动机制,它们分别基于补偿性承诺、集中式承诺协议和参与者主导的联盟。这三种机制都改善了总体结果,但没有任何单一机制在所有的金融结构中表现最佳。在不同机制之间,成功的稳定化具有共同的时间模式:广泛的承诺在防御性行为变得自我强化之前很早就形成。我们的发现表明,个体能力强的智能体并不会自动形成安全的金融系统,这凸显出系统级评估和交互设计是金融 AI 安全的核心问题。代码可在 https://anonymous.4open.science/r/FinFrail-CF26 获取。
cs.AI / 13 / 2609.30943
LogicTree-RAG: Logic Tree-guided Retrieval-Augmented Generation for Long-form Patent Drafting
LogicTree-RAG:逻辑树引导的面向长篇幅专利起草的检索增强生成
Jiaqi Zhu, Naili Xing, Hexiang Pan, Haotian Gao, Jianwei Yin, Xiaokui Xiao, Beng Chin Ooi
cs.AI
large language model
大语言模型相关
Abstract
Long-form technical text generation underpins knowledge-intensive workflows, yet remains challenging for large language models (LLMs) due to the need for globally consistent logical structuring and faithful technical reasoning beyond local coherence. Patent drafting is a canonical instance of this challenge, demanding holistic generation of a legally compliant and technically exhaustive document through sustained multi-expert collaboration. Existing approaches often focus on partial section generation or rely on manually crafted outlines, limiting scalable automation in realistic settings. In this work, we propose LogicTree-RAG, a logic tree-guided retrieval-augmented generation framework that induces a hierarchical logic tree as a global organizational backbone to organize and ground technical disclosures, without relying on expert-defined drafting priors. Each node in the logic tree represents a technical element and is constructed through evidence-guided recursive generation. A hybrid traversal mechanism then maps the logic tree into patent sections, enabling controllable and section-balanced generation. Extensive experiments show that LogicTree-RAG consistently improves content quality and language conformity over strong LLM-based baselines and achieves longer structured generation with high token efficiency, demonstrating the effectiveness of logic-centric generation for complex technical document drafting.
Chinese Translation
长篇幅技术文本生成支撑着知识密集型工作流,但由于需要超越局部连贯性的全局一致逻辑结构和忠实的技术推理,它对大语言模型(LLMs)而言仍然具有挑战性。专利起草是这一挑战的一个典型实例,它要求通过持续的多专家协作,整体生成一份合法合规且技术详尽的文档。现有方法通常聚焦于部分章节生成,或依赖手工构建的提纲,从而限制了现实场景中的可扩展自动化。在这项工作中,我们提出 LogicTree-RAG,一种逻辑树引导的检索增强生成框架,该框架归纳出一个分层逻辑树作为全局组织主干,以组织和锚定技术公开内容,而不依赖专家定义的起草先验。逻辑树中的每个节点表示一个技术要素,并通过证据引导的递归生成构建。随后,一种混合遍历机制将逻辑树映射到专利章节中,从而实现可控且章节均衡的生成。大量实验表明,LogicTree-RAG 在内容质量和语言符合性方面持续优于强大的基于 LLM 的基线,并以高 token 效率实现更长的结构化生成,证明了以逻辑为中心的生成在复杂技术文档起草中的有效性。
cs.AI / 14 / 2609.30954
FTB Graph: Determining and Validating First-token Broadcasters and Language-Identity Head Circuits in Multilingual Language Models
FTB Graph:确定并验证多语言语言模型中的首词元广播器与语言身份头电路
Arjun Pillai, Christian Hoang, Anjelo Laroza
cs.AI
large language model
大语言模型相关
Abstract
Large language models operating in multilingual contexts must resolve target response languages early in generation, yet the causal circuitry governing first-token language identity decisions remains poorly mapped. We present an end-to-end structural circuit analysis across six model architectures spanning four families: GPT-2, BLOOM-560M, Pythia-1B/2.8B, and Qwen2.5-1.5B Base/Instruct. Using Edge Attribution Patching (EAP) with FP16 active clamping, followed by exact activation patching verification with a 2,000-candidate-edge search ceiling, we extract directed acyclic graphs driving first-token language broadcasting. Across the standalone models, we observe deep or mid-to-deep broadcasting hubs, though the evidence is strongest for Pythia-2.8B and BLOOM-560M because GPT-2 and Pythia-1B leave few out-of-graph heads for comparison, while both Qwen2.5-1.5B variants invert the necessity check. Scaling from Pythia-1B to 2.8B expands node participation while maintaining a similar verified edge budget, producing sparser topology. The Qwen2.5-1.5B base and instruct circuits retain 84.7% Jaccard similarity, including the Layer 27 hub, indicating that first-token routing is largely established during pretraining and preserved by instruction tuning. Finally, EAP scores correlate weakly with exact patching deltas across most models, showing that linear gradient approximations can diverge from causal interventions in FP16 and motivating exact-patching verification for reliable circuit discovery.
Chinese Translation
在多语言情境下运行的大语言模型必须在生成早期就确定目标回复语言,然而支配首词元语言身份决策的因果电路仍缺乏清晰的刻画。我们对跨四个家族的六种模型架构进行了端到端的结构性电路分析:GPT-2、BLOOM-560M、Pythia-1B/2.8B,以及 Qwen2.5-1.5B Base/Instruct。我们使用带有 FP16 主动钳制的边归因修补(Edge Attribution Patching, EAP),随后在 2,000 条候选边搜索上限下进行精确激活修补验证,提取出驱动首词元语言广播的有向无环图。在各个独立模型中,我们观察到深层或中深层的广播枢纽,不过证据对 Pythia-2.8B 和 BLOOM-560M 最为有力,因为 GPT-2 和 Pythia-1B 留下的图外头(out-of-graph heads)很少、难以进行比较,而两个 Qwen2.5-1.5B 变体则颠倒了必要性检验。从 Pythia-1B 扩展到 2.8B 时,节点参与度扩大,同时维持了相近的已验证边预算,从而产生更稀疏的拓扑结构。Qwen2.5-1.5B 的基座与指令电路保持了 84.7% 的 Jaccard 相似度,其中包括第 27 层枢纽,这表明首词元路由在很大程度上于预训练期间即已确立,并被指令微调所保留。最后,在大多数模型中,EAP 分数与精确修补增量之间的相关性较弱,这表明线性梯度近似在 FP16 下可能与因果干预产生偏离,从而推动人们采用精确修补验证来实现可靠的电路发现。
cs.AI / 15 / 2609.30967
MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens
MoMHa:面向准确率、安全性和 Token 的 LLM Harness 多目标优化
Subhojyoti Mukherjee, Md Mehrab Tanjim
cs.AI
large language model
大语言模型相关
Abstract
Most work on improving large language models treats accuracy as the sole objective. We argue that the harness, the Python code surrounding the model that constructs prompts, routes calls, and parses outputs, is a first-class design surface whose quality is inherently multi-objective: an accurate harness that refuses no unsafe request, or that consumes an order of magnitude more tokens, is not a good harness. We present Meta-Harness, a system that casts harness design as search over three per-domain objectives (accuracy, behavioural safety, and token cost) solved by an agentic proposer (Claude Code) with full filesystem access to prior harness source, execution traces, and scoring artifacts. Our central finding is that a singlephase joint-reward proposer (MoMHa) outperforms every alternative, including a two-phase "accuracy then tokens" ablation, scalar-only feedback, and an accuracy-only baseline. We evaluate on seventeen domains: seven synthetic capability suites, seven real-world public benchmarks (HumanEval, MBPP, Spider, FEVER, MMLU-Pro, LawBench, NuminaMath), and three U-SafeBench-derived user-specific safety domains, using a 12-model fleet spanning four families. On the synthetic track MoMHa achieves a joint mean of 0.482 versus 0.198-0.422 for ten baselines, winning $7 / 10$ per-domain columns; on the real-world track it scores 0.461 versus 0.377 for the strongest baseline (DSPy), winning 5/7 columns, demonstrating that harness strategies transfer to unseen benchmarks without retraining on 8 of 12 target models. MoMHa attains the highest measured behavioral safety composite (U-SafeBench, 0.781) and uses 95 fewer tokens per example than the two-phase alternative. We will release all harness code, evaluation infrastructure, and crossmodel logs.
Chinese Translation
大多数改进大语言模型的工作将准确率视为唯一目标。我们认为,harness,即围绕模型、用于构造提示、路由调用并解析输出的 Python 代码,是一个一等设计面,其质量本质上是多目标的:一个准确的 harness,如果不拒绝任何不安全请求,或者消耗多一个数量级的 token,就不是一个好的 harness。我们提出 Meta-Harness,一个将 harness 设计转化为对三个按领域划分的目标(准确率、行为安全性和 token 成本)进行搜索的系统,该搜索由具备对先前 harness 源代码、执行轨迹和评分产物的完整文件系统访问权限的智能体式提议器(Claude Code)求解。我们的核心发现是,单阶段的联合奖励提议器(MoMHa)优于所有替代方案,包括两阶段的“先准确率后 token”消融、仅标量反馈以及仅准确率基线。我们在十七个领域上进行评估:七个合成能力套件、七个真实世界公开基准(HumanEval、MBPP、Spider、FEVER、MMLU-Pro、LawBench、NuminaMath),以及三个由 U-SafeBench 衍生的用户特定安全领域,使用一个跨越四个家族的 12 个模型组成的模型群。在合成赛道上,MoMHa 取得了 0.482 的联合均值,而十个基线为 0.198-0.422,赢得 $7 / 10$ 个按领域划分的列;在真实世界赛道上,它得分 0.461,而最强基线(DSPy)为 0.377,赢得 5/7 列,表明 harness 策略无需在 12 个目标模型中的 8 个上重新训练即可迁移到未见过的基准。MoMHa 取得了最高的实测行为安全综合值(U-SafeBench,0.781),并且相比两阶段替代方案每个示例少用 95 个 token。我们将发布所有 harness 代码、评估基础设施和跨模型日志。
cs.AI / 16 / 2609.31013
Same Text, Different Numbers: The Divergence of LLM-Based Measures
相同文本,不同数字:基于LLM的测度的分歧
Hamid Boustanifar, Sasan Mansouri
cs.AI · cs.CL · q-fin.GN · q-fin.RM
large language model
大语言模型相关
Abstract
Researchers increasingly use generative large language models (LLMs) to convert corporate text into empirical variables. We examine the extent to which LLM-based textual measures are invariant to model choice using thirteen measures, including sentiment, management clarity, uncertainty, answer specificity, and climate and political risk. Seven LLMs from different providers score earnings call transcripts of S&P 500 companies on these constructs. Cross-model rank correlations average only 0.52, and transcript-level differences common across providers account for only 34% of total score variation. Cross-model disagreement does not predict subsequent analyst or market disagreement, consistent with a substantial model-specific component rather than common ambiguity in the underlying disclosure. Model choice significantly affects downstream inference, with coefficient magnitudes, signs, and statistical significance varying substantially across models. Averaging across providers makes transcript rankings more stable for most constructs, but score levels remain sensitive to the models included in the ensemble. LLM-generated variables should therefore be treated as model-contingent measurements and validated across providers.
Chinese Translation
研究人员日益使用生成式大语言模型(LLM)将公司文本转化为经验变量。我们使用十三种测度,包括情感、管理层清晰度、不确定性、回答具体性以及气候和政治风险,考察基于LLM的文本测度在多大程度上对模型选择具有不变性。来自不同提供商的七个LLM对标准普尔500指数公司的财报电话会议记录在这些构念上进行评分。跨模型的秩相关平均仅为0.52,而不同提供商共有的记录层面差异仅解释总得分变异的34%。跨模型分歧不能预测随后的分析师或市场分歧,这与存在显著的模型特有成分而非底层披露中的共同模糊性相一致。模型选择显著影响下游推断,系数大小、符号和统计显著性在不同模型之间差异很大。跨提供商平均使大多数构念的记录排名更稳定,但得分水平仍对集成中所包含的模型敏感。因此,LLM生成的变量应被视为依模型而定的测量,并应在不同提供商之间进行验证。
cs.AI / 17 / 2609.31054
Cheap, open agents make LLM pollution harder to mitigate
廉价、开放的智能体使 LLM 污染更难缓解
Raluca Rilla, Anne-Marie Nussberger, Rui Mata, Dirk U. Wulff
cs.AI
large language model
大语言模型相关
Abstract
Large Language Model (LLM) pollution occurs when synthetic responses contaminate data intended to capture human behavior. High deployment costs have so far limited the risk posed by autonomous survey agents. However, open-weight models paired with open-source agentic frameworks may have removed this barrier. We compared the performance and detectability of nine agent configurations, ranging from fully open variants to closed commercial ones. Each agent autonomously completed a survey containing multiple response types yielding various detection checks. Fully open agents ran locally without usage fees and performed competitively with commercial alternatives. Open and commercial agents failed different sets of checks, and no single check reliably detected all agents, but open-text responses discriminated best between agents and humans. These findings identify fully open agents as a distinct risk for LLM pollution and support multilayered detection strategies emphasizing open-text analysis.
Chinese Translation
当合成回答污染了旨在捕捉人类行为的数据时,就会发生大语言模型(LLM)污染。迄今为止,高昂的部署成本限制了自主调查智能体所带来的风险。然而,开放权重模型与开源智能体框架相结合,可能已经消除了这一障碍。我们比较了九种智能体配置的性能与可检测性,这些配置涵盖了从完全开放的变体到闭源商业变体。每个智能体都自主完成了一份包含多种回答类型的调查,这些回答类型会产生各种检测检验。完全开放的智能体在本地运行且无需使用费,其表现与商业替代方案相比具有竞争力。开放智能体与商业智能体未能通过不同的检测组合,且没有任何单一检测能够可靠地检测出所有智能体,但开放式文本回答在区分智能体与人类方面表现最佳。这些发现将完全开放的智能体认定为 LLM 污染的一项独特风险,并支持强调开放式文本分析的多层检测策略。
cs.AI / 18 / 2609.31078
OmouAI: Argumentative Human-AI Policy Deliberation with Simulated Personas
OmouAI:使用模拟角色的论辩式人类-人工智能政策审议
Stylianos Loukas Vasileiou, Antonio Rago, William Yeoh, Georgina Curto
cs.AI
large language model
大语言模型相关
Abstract
Debates amongst agents driven by large language models (LLMs) have demonstrated vast potential in various applications, but when these interactions include humans and take place in high-stakes environments, e.g., in public policy deliberations, they are beset with issues such as sycophancy and a lack of faithful explanations. To tackle these issues, we present OmouAI, an interactive and inclusive deliberation system that uses LLMs in combination with computational argumentation, a field which excels in representing and reasoning within debates. OmouAI allows a human user to deliberate policy claims for real-world challenges with simulated personas, e.g., representing stakeholders, domain experts or devil's advocates, towards reducing sycophancy. Each persona generates its own arguments, and the arguments of all parties form a shared argumentation framework. Users can then contest, add and revise arguments, providing crucial human oversight. Then, arguments are evaluated using deterministic argumentative semantics against external goals, such as the UN Sustainable Development Goals, guaranteeing faithful explanations. The advancement or worsening of the goals thus serve as indicators for the policy recommendations.
Chinese Translation
由大型语言模型(LLMs)驱动的智能体之间的辩论已在各种应用中展现出巨大潜力,但当这些互动包含人类并且发生在高风险环境中时,例如在公共政策审议中,它们会受到诸如谄媚和缺乏忠实解释等问题困扰。为了解决这些问题,我们提出 OmouAI,一个交互式且包容性的审议系统,它将 LLMs 与计算论辩相结合,该领域擅长在辩论中进行表示和推理。OmouAI 允许人类用户与模拟角色一起审议针对现实世界挑战的政策主张,这些模拟角色例如代表利益相关者、领域专家或魔鬼代言人,以减少谄媚。每个角色生成自己的论点,所有各方的论点形成一个共享的论辩框架。用户随后可以质疑、添加和修订论点,从而提供关键的人类监督。然后,使用确定性论辩语义对照外部目标(例如联合国可持续发展目标)来评估论点,从而保证忠实解释。因此,这些目标的推进或恶化可作为政策建议的指标。
cs.AI / 19 / 2609.31140
Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability?
语言推理向量能否增强多模态推理能力?
Ziyi Wang, Li Li, Aolin Zhou, Yankun Shen, Chonghan Liu, Shuxia Lin, Xu Yang
cs.AI
large language model
大语言模型相关
Abstract
Most Vision-Language Models (VLMs) are built by extending pretrained Large Language Models (LLMs) with visual modules and multimodal alignment. However, this multimodal scaling often degrades the language-side reasoning ability originally encoded in the base LLM. While the base LLM retains usable reasoning after scaling, the aligned VLM itself cannot reliably access this ability. Therefore, recovering the degraded reasoning capability in VLMs would benefit more from seeking help from the base LLM than from the VLM alone. Motivated by this, we propose LIFT (Language-side reasonIng Facilitation and Transfer), a lightweight vector-intervention method that transfers reasoning capability from the base LLM to the VLM without retraining the backbone. LIFT defines Reasoning Vectors as answer-token hidden-state differences between a Reasoner path with an explicit reasoning trace and a Solver path without it, and injects these vectors into language-side activations of the target VLM. LIFT further supports learnable vector adaptation while keeping the VLM backbone frozen. We evaluate LIFT on two VLMs across six reasoning benchmarks, comparing Reasoning Vectors extracted from the base LLM and from the aligned VLM under matched protocols. Results show that LLM-derived vectors consistently outperform VLM-derived vectors, confirming that the base LLM is a more effective source for recovering reasoning. LIFT partially recovers degraded reasoning through lightweight language-side interventions. Further analyses show that Reasoning Vectors influence intermediate reasoning behavior rather than merely altering final answers. The source code will be released soon.
Chinese Translation
大多数视觉语言模型(VLMs)是通过用视觉模块和多模态对齐扩展预训练大语言模型(LLMs)构建的。然而,这种多模态扩展常常会削弱基座 LLM 中原本编码的语言侧推理能力。尽管基座 LLM 在扩展后仍保留可用的推理能力,但经过对齐的 VLM 本身无法可靠地访问这种能力。因此,要恢复 VLM 中退化的推理能力,从基座 LLM 寻求帮助会比仅依靠 VLM 本身更有益。受此启发,我们提出 LIFT(Language-side reasonIng Facilitation and Transfer,语言侧推理促进与迁移),这是一种轻量级向量干预方法,可在无需重新训练骨干的情况下将推理能力从基座 LLM 迁移到 VLM。LIFT 将推理向量(Reasoning Vectors)定义为:带有显式推理轨迹的推理器(Reasoner)路径与不带显式推理轨迹的求解器(Solver)路径之间在答案 token 隐藏状态上的差异,并将这些向量注入目标 VLM 的语言侧激活中。LIFT 进一步支持可学习的向量适配,同时保持 VLM 骨干冻结。我们在两个 VLM 上跨六个推理基准评估 LIFT,并在匹配协议下比较从基座 LLM 和从对齐后的 VLM 中提取的推理向量。结果表明,源自 LLM 的向量始终优于源自 VLM 的向量,证实基座 LLM 是恢复推理的更有效来源。LIFT 通过轻量级语言侧干预部分恢复了退化的推理。进一步分析表明,推理向量影响中间推理行为,而不仅仅是改变最终答案。源代码将很快发布。
cs.AI / 20 / 2609.31215
DIAL: Position-Debiased LLM Judges with Adaptive Human Preference Calibration
DIAL:具备自适应人类偏好校准的位置去偏 LLM 评判器
Zesheng Cai, Yingqi Fan, Sichang Chen, Jin-Hong Du
cs.AI · stat.AP · stat.ME · stat.ML
large language model
大语言模型相关
Abstract
Large language models (LLMs) as a judge enable scalable evaluation, but their judgments can be sensitive to response order and, even after removing such position effects, can still diverge systematically from human preferences.We introduce DIAL, a unified framework that combines abundant LLM comparisons with limited human comparisons to separate judge-specific position effects, learn shared structure in position-debiased LLM preferences, and adaptively calibrate that structure toward the human preference target. Theoretically, we study three aspects of DIAL: (i) identification of latent LLM preferences, position effects, and human calibration; (ii) adaptive estimation that balances LLM anchoring against limited human evidence; and (iii) fixed-weight uncertainty quantification for the calibrated human preference. Empirically, we evaluate position debiasing and human alignment separately in controlled simulations and on three human-preference benchmarks, showing that DIAL remains robust to unbalanced response order, achieves strong human-aligned rankings with limited labels, and adapts toward human evidence when LLM information is imperfect. Our real-data study collects over 410K judgments from 21 LLM judges in both display orders, providing a resource for future studies of LLM-judge bias, heterogeneity, and human alignment.
Chinese Translation
作为评判器的大语言模型(LLMs)能够实现可扩展的评估,但它们的判断可能对回答顺序敏感;而且即使在移除这类位置效应后,仍可能系统性地偏离人类偏好。我们提出 DIAL,一个统一框架,它将丰富的 LLM 比较与有限的人类比较相结合,以分离评判器特定的位置效应、学习经位置去偏的 LLM 偏好中的共享结构,并自适应地将该结构校准至人类偏好目标。在理论上,我们研究 DIAL 的三个方面:(i) 潜在 LLM 偏好、位置效应和人类校准的识别;(ii) 在 LLM 锚定与有限人类证据之间进行平衡的自适应估计;以及 (iii) 针对校准后人类偏好的固定权重不确定性量化。在实证上,我们在受控模拟中和三个人类偏好基准上分别评估位置去偏与人类对齐,表明 DIAL 对不平衡的回答顺序保持稳健,在有限标签下实现强的人类对齐排序,并在 LLM 信息不完美时朝人类证据自适应。我们的真实数据研究从 21 个 LLM 评判器以两种展示顺序收集了超过 410K 条判断,为未来关于 LLM 评判器偏差、异质性和人类对齐的研究提供了资源。
cs.AI / 21 / 2609.31354
Mutable Transcripts: Mitigating Context Pollution through Editable Conversation State
可变转录:通过可编辑的对话状态缓解上下文污染
Dan Barry, Andrew Hines
cs.AI
large language model
大语言模型相关
Abstract
Contemporary large language model (LLM) chat systems treat conversation history as an immutable sequence of turns that defines the model's working context. However, user intent in real interactions is not static: it evolves through correction, refinement, and shifting constraints. This mismatch between dynamic intent and static transcripts can result in context pollution, where outdated or irrelevant information persists and continues to influence subsequent responses. We introduce mutable transcripts, a new interaction paradigm that enables users to revise prior turns through natural language edit requests, allowing the conversation history itself to be updated rather than appended. This reframes the transcript from a passive record into an editable representation of conversational state. We present a working prototype that integrates transcript-level revision into a standard chat interface and evaluate its feasibility through a controlled user study (n=17) and an illustrative transcript analysis of representative interaction scenarios. Participants significantly preferred mutable transcripts over standard chat across measures of clarity, confidence, and ease of use, with reduced intent to restart conversations. Transcript analysis of representative user study conversations shows that mutable transcripts can reduce conversation length and eliminate obsolete retained context. These findings provide initial evidence that user-driven revision of conversational history can improve interaction quality and help maintain a more current representation of user intent. The source code and prototype can be accessed at https://github.com/QxLabIreland/ReChat
Chinese Translation
当代大型语言模型(LLM)聊天系统将会话历史视为定义模型工作上下文的不可变的轮次序列。然而,真实交互中的用户意图并非静态的:它会通过纠正、细化和不断变化的约束而演变。动态意图与静态转录之间的这种不匹配可能导致上下文污染,即过时或无关的信息持续存在,并继续影响后续回复。我们引入了可变转录,这是一种新的交互范式,使用户能够通过自然语言编辑请求来修订先前的轮次,从而让会话历史本身得以更新,而非仅被追加。这将转录从一种被动记录重新构造为对话状态的可编辑表示。我们提出了一个可运行的原型,将会话转录层面的修订集成到标准聊天界面中,并通过一项受控用户研究(n=17)以及对代表性交互场景的示例性转录分析来评估其可行性。在清晰度、信心和易用性等各项衡量指标上,参与者显著更偏好可变转录而非标准聊天,并且重启对话的意愿有所降低。对代表性用户研究对话的转录分析表明,可变转录能够缩短对话长度并消除被保留的过时上下文。这些发现提供了初步证据,表明由用户驱动的会话历史修订能够提升交互质量,并有助于维持对用户意图更贴近当前的表示。源代码和原型可在 https://github.com/QxLabIreland/ReChat 获取
cs.AI / 22 / 2609.31460
Segment-Level Agentic Topic Modeling for Improved Data Exploration and Resource Efficiency
用于改进数据探索和资源效率的片段级智能体主题建模
Myeongjun Erik Jang, Antonios Georgiadis, Sae Young Moon, Fran Silavong
cs.AI
large language model
大语言模型相关
Abstract
Topic modeling is an effective technique for discovering hidden themes within documents and is widely used in text mining and data analysis across a variety of industry sectors. Recently, large language model (LLM)-based topic models have been emerged that prompt LLMs to generate topics then assign the topics to documents, producing more natural and human-readable topics than conventional topic modeling algorithms. However, the nature of topic assignment process causes certain drawbacks, such as the incapability to produce topic distributions over a document, too broad or narrow topics, and high resource consumption, which increases with the number and length of of documents being assigned topics. These issues are particularly critical for industrial applications, which require high-quality, in-depth analysis and the processing of large volumes of documents. In this context, this paper introduces a framework called SeLATM, which addresses these concerns by employing segment-level topic generation and topic refinement through agentic feedback loops. Experimental results on various datasets demonstrate that SeLATM significantly reduces the LLM resources compared to methods based on topic assignment process, while maintaining superior performance.
Chinese Translation
主题建模是一种用于发现文档中隐藏主题的有效技术,并广泛应用于各种行业领域的文本挖掘和数据分析中。最近,基于大语言模型(LLM)的主题模型已经出现,它们提示LLM生成主题,然后将主题分配给文档,从而产生比传统主题建模算法更自然、更易读的主题。然而,主题分配过程的性质导致某些缺点,例如无法生成文档上的主题分布、主题过宽或过窄以及高资源消耗,而资源消耗会随着被分配主题的文档数量和长度的增加而增加。这些问题对于工业应用尤为关键,因为工业应用需要高质量、深入的分析以及对大量文档的处理。在此背景下,本文介绍了一个名为SeLATM的框架,它通过采用片段级主题生成以及通过智能体反馈循环进行的主题精炼来解决这些问题。在各种数据集上的实验结果表明,与基于主题分配过程的方法相比,SeLATM显著减少了LLM资源,同时保持了优越的性能。
cs.AI / 23 / 2609.31473
Game Arena: Strategic LLM Evaluation in Competitive Environments
Game Arena:竞争性环境中的策略性 LLM 评估
Bovard Doerschuk-Tiberi, Yao Yan, Justin Chiu, Hann Wang, Timothy Chung, Martyna Plomecka, John Schultz, Jon Lipovetz, Clayton Drazner, Yuchen Zhuang, Jaimie Hwang, Nate Keating, Riley Jones, Andrew Lee, Oran Kelly, Ian Gemp, Michael Aaron, Laurel Prince, Kate Larson, Jeff Moser, Harrison Jobe, Chad Woodford, Siqi Liu, Andrew Wang, Bo Chang, Christopher D'Mello, Diane Chaleff, Addison Howard, Johnny Yip, Chuck Sugnet, Antonio Gulli, Meghan O'Connell, Will Cukierski, Nenad Tomasev, Dima Yeroshenko, Kinjal Parekh, Roxanne Daniel, Marc Lanctot, Domino Weir, Elsa Dong, Daniel Hennes, Melissa Nalubwama, Robert Fraser, Ryan Trostle, Jun Peng, Tom Mason, Lloyd Hightower, Chiamaka Chukwuka, Yuexiang Zhai, Phoebe Kirk, Yi Su, Yuting Han, Jie Ren, Chris Prichard, Sahand Sharifzadeh, Karim Hakimzadeh, DJ Sterling, Meg Risdal, Kate Olszewska, Ya Xu, Orhan Firat, Minmin Chen
cs.AI
large language model
大语言模型相关
Abstract
We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head matchups in structured environments where the gameplay strength naturally increases as models evolve, preventing performance saturation. This technical report details the infrastructure behind Game Arena and describes the three pilot game environments: Chess, Poker, and Werewolf. These environments span perfect information, imperfect information, and multiplayer game settings, enabling a systematic study of models' strategic planning, adaptation, and robustness under uncertainty. For each game, we provide a detailed description of the environment, evaluation metrics, and results from running full competitions across models. Through robust infrastructure and large-scale ground-truth based evaluation, Game Arena ensures reproducibility, transparency and generalizability to new games and variants over time.
Chinese Translation
我们介绍了 Kaggle Game Arena,这是一个开放且不断扩展的平台,用于通过竞技游戏评估大型语言模型(LLM)。与静态基准不同,游戏竞技场使模型能够在结构化环境中进行正面交锋的对决,在这些环境中,随着模型的演进,对局强度会自然提升,从而防止性能饱和。本技术报告详细介绍了 Game Arena 背后的基础设施,并描述了三个试点游戏环境:国际象棋、扑克和狼人杀。这些环境涵盖完全信息、不完全信息以及多人游戏设定,从而能够系统地研究模型在不确定性下的战略规划、适应能力和鲁棒性。对于每个游戏,我们都详细描述了环境、评估指标,以及在各个模型之间运行完整竞赛所得到的结果。通过稳健的基础设施和基于大规模真实基准(ground-truth)的评估,Game Arena 确保了可复现性、透明度,以及随着时间推移对新游戏和变体的可推广性。
cs.AI / 24 / 2609.31505
Prompt Minimization: Reducing Input Redundancy Without Sacrificing Output Fidelity
提示最小化:在不牺牲输出保真度的情况下减少输入冗余
Marius F. R. Juston, Kevin A. Karim, Jonathan Gao, Kevin C. Li, Rudhi Bashambu
cs.AI
large language model
大语言模型相关
Abstract
Despite the growing capabilities of large language models (LLMs), prompt design remains largely heuristic and ad hoc. This project will explore $\textit{prompt minimization}$, the process of reducing prompts to their smallest, most information-dense form while preserving output fidelity. Practically, shorter prompts reduce computational overhead and inference latency, especially when large contexts, such as entire documents or codebases, are included unnecessarily. Further, longer prompts can damage LLM reasoning and accuracy. Theoretically, the existence of multiple prompts yielding equivalent outputs suggests a high degree of redundancy in the input space, raising fundamental questions about what information is essential to elicit specific model behaviors. We propose three variant frameworks to identify and evaluate minimal prompts and demonstrate that minimal prompts often produce outputs comparable to those of their longer counterparts. These findings suggest new directions for efficient prompt engineering and deepen our understanding of input compression in LLMs.
Chinese Translation
尽管大语言模型(LLM)的能力不断增强,提示设计在很大程度上仍然是启发式且临时性的。本项目将探索 $\textit{prompt minimization}$,即在保持输出保真度的同时,将提示缩减到其最小、信息密度最高的形式的过程。在实践中,更短的提示可以减少计算开销和推理延迟,尤其是当诸如整个文档或代码库之类的大规模上下文被不必要地包含在内时。此外,更长的提示可能损害 LLM 的推理能力和准确性。在理论上,存在多个产生等价输出的提示这一事实表明输入空间具有高度的冗余性,这引发了关于什么信息对于诱发特定模型行为是必不可少的基本问题。我们提出了三个变体框架来识别和评估最小提示,并证明最小提示通常能产生与其更长对应提示相当的输出。这些发现为高效的提示工程提示了新的方向,并加深了我们对 LLM 中输入压缩的理解。
cs.CL / 25 / 2609.30416
All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation
恰逢其时:面向基于大语言模型的同声语音到语音翻译的因果感知框架
Amir Hussein, Enas Albasiri, Travis M. Bartley, Nourchene Ferchichi, Ke Hu, Harishchandra Dubey, Myungjong Kim, Zhehuai Chen, Oluwatobi Olabiyi, Sanjeev Khudanpur
cs.CL
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) have shown strong performance in low-resource offline translation; however, extending them to simultaneous speech-to-speech translation (Simul-S2ST) remains challenging due to the scarcity of causally aligned training data with high cross-lingual speaker fidelity. In addition, existing approaches rely on fixed translation policy or confidence heuristics, leading to suboptimal quality and higher latency. We propose a causality-aware Simul-S2ST framework with a novel data pipeline that generates high-fidelity, causally aligned segments with improved voice transfer. The framework introduces (i) a factorized S2ST architecture (FAST), (ii) a causality-aware adaptive policy (CAP), and (iii) causality-aware latency metric. Experiments on CVSS Spanish, German, and French show that FAST-CAP consistently improves the quality-latency trade-off, achieving up to +1.2 BLEU and a 26% relative latency reduction over a fixed policy. Despite using substantially less training data than existing systems, FAST-CAP achieves state-of-the-art results in speech translation quality and speaker fidelity while yielding up to a 38.8% relative reduction in latency.
Chinese Translation
大语言模型(LLM)在低资源离线翻译中已展现出强劲的性能;然而,由于缺乏具有高跨语言说话人保真度的因果对齐训练数据,将其扩展到同声语音到语音翻译(Simul-S2ST)仍然充满挑战。此外,现有方法依赖固定的翻译策略或置信度启发式规则,导致质量次优且延迟更高。我们提出了一个因果感知的 Simul-S2ST 框架,并配有一个新颖的数据流水线,该流水线能够生成高保真、因果对齐且语音迁移效果更佳的片段。该框架引入了(i)分解式 S2ST 架构(FAST),(ii)因果感知自适应策略(CAP),以及(iii)因果感知延迟指标。在 CVSS 西班牙语、德语和法语上的实验表明,FAST-CAP 持续改善了质量—延迟权衡,相较于固定策略最多提升 +1.2 BLEU,并实现 26% 的相对延迟降低。尽管使用的训练数据远少于现有系统,FAST-CAP 在语音翻译质量和说话人保真度上仍取得了当前最优的结果,同时实现了最高 38.8% 的相对延迟降低。
cs.CL / 26 / 2609.30773
Learning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized Guidance
利用随机化引导学习串联语音到语音模型中的自然对话行为
Manato Yaguchi, Yotaro Kubo, Hikaru Asano, So Kuroki
cs.CL · eess.AS
large language model
大语言模型相关
Abstract
Tandem speech-to-speech architectures couple a responsive speech frontend with an asynchronous text backend. In KAME, a large language model (LLM) serves as the backend, supplying candidate responses as guidance to the speech frontend while the user is still speaking. Ordinary conversation recordings capture the eventual response but not the guidance the backend would supply during the user's utterance. Generating the missing guidance with a simulator LLM adds substantial data-preparation overhead when training on real conversations. We propose randomized intermediate guidance, which derives guidance directly from the conversation corpus rather than simulating backend LLM behavior. During training, target responses provide informative guidance, while randomly sampled responses provide potentially irrelevant updates during the utterance. This combination aims to teach the frontend to use backend information selectively. On synthetic dialogues, KAME trained with this recipe achieves response quality comparable to that of the LLM-generated and similarity-based baselines. Training on 3.8k hours of real conversations improves smooth turn-taking and audio-judge naturalness over synthetic-data KAME while retaining a response-quality advantage over Moshi. These results show that randomized guidance offers a practical route to combining the response-quality benefits of tandem models with natural conversational behavior learned from real speech.
Chinese Translation
串联式语音到语音架构将一个响应迅速的语音前端与一个异步文本后端耦合在一起。在 KAME 中,一个大型语言模型(LLM)充当后端,在用户仍在说话时向语音前端提供候选响应作为引导。普通对话录音捕获的是最终的响应,而不是后端在用户话语期间会提供的引导。在真实对话上训练时,用模拟器 LLM 生成缺失的引导会增加大量数据准备开销。我们提出随机化中间引导,它直接从对话语料库中导出引导,而不是模拟后端 LLM 的行为。在训练期间,目标响应提供有信息量的引导,而随机采样的响应在话语期间提供可能不相关的更新。这种组合旨在教会前端有选择地使用后端信息。在合成对话上,使用该方案训练的 KAME 达到了与 LLM 生成基线和基于相似度的基线相当的响应质量。在 3.8k 小时真实对话上训练,相较于合成数据训练的 KAME,改善了顺畅的轮次转换和音频评判自然度,同时保持了对 Moshi 的响应质量优势。这些结果表明,随机化引导提供了一条实用路径,可将串联模型的响应质量优势与从真实语音中学习到的自然对话行为结合起来。
cs.CL / 27 / 2609.30882
Effects of Transcript Compression on LLM-based Medical Misinformation Detection in Japanese YouTube Videos
转录压缩对基于LLM的日语YouTube视频医疗虚假信息检测的影响
Yuya Wake, Sho Tsugawa, Toshiyuki Amagasa
cs.CL
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly used to assess long-form medical videos, but their effectiveness may depend on whether transcripts are provided in full or compressed through summarization, retrieval, or claim screening. This study examines how such transcript compression affects LLM-based veracity classification of Japanese medical YouTube videos. We compare four transcript input designs: full transcripts, LLM-generated summaries, RAPTOR-based retrievalaugmented generation (RAG), and Screening, which extracts candidate medical and health-related sentences. Using 74 long-form videos labeled as Real or Fake, we evaluate classification performance and analyze linguistic changes using J-LIWC, hedge expressions, and institutional or technical terms. The full-transcript Baseline achieved the best performance, whereas all compressed inputs increased false negatives, meaning that Fake videos were more likely to be misclassified as Real. Summary caused the largest performance drop, while Screening performed best among the compressed inputs but still omitted many medically relevant sentences. Linguistic analyses showed that these errors were not explained by a simple increase in certainty. Instead, Summary reduced affective, social, temporal, cognitive, and conversational cues, while Summary and RAG made institutional and technical terms more salient. These findings suggest that transcript compression can represent Fake videos as more coherent and authoritative inputs, thereby weakening cues needed for misinformation detection
Chinese Translation
大语言模型(LLM)越来越多地被用于评估长时医疗视频,但其有效性可能取决于转录文本是以完整形式提供,还是通过摘要、检索或主张筛查进行压缩。本研究考察这种转录压缩如何影响基于LLM的日语医疗YouTube视频真实性分类。我们比较四种转录输入设计:完整转录文本、LLM生成的摘要、基于RAPTOR的检索增强生成(RAG),以及Screening,后者提取候选的医学和健康相关句子。使用74个标记为Real或Fake的长时视频,我们评估分类性能,并使用J-LIWC、模糊限制表达以及机构性或技术性术语分析语言变化。完整转录基线取得了最佳性能,而所有压缩输入都增加了假阴性,这意味着Fake视频更可能被误分类为Real。摘要造成了最大的性能下降,而Screening在压缩输入中表现最好,但仍遗漏了许多医学相关句子。语言分析表明,这些错误并非由确定性的简单增加所解释。相反,摘要减少了情感、社会、时间、认知和会话线索,而摘要和RAG使机构性和技术性术语更加显著。这些发现表明,转录压缩可能将Fake视频表示为更连贯、更具权威性的输入,从而削弱虚假信息检测所需的线索。
cs.CL / 28 / 2609.30906
ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning
ToolSearcher:通过强化学习优化大规模工具选择
Zhenlong Dai, Xujie Song, Zitong Wang, Tong Niu, Jian liu, Weiqiang Wang, Xiu Tang, Sai Wu, Chang Yao, Jingyuan Chen
cs.CL
large language model
大语言模型相关
Abstract
Large language models (LLMs) excel at natural language processing but struggle to interact with external environments. Tool learning provides a promising way to extend LLMs into actionable agents, where tool selection is a critical prerequisite for successful tool use. Existing work often assumes a small or predefined set of tools, leaving large-scale tool selection underexplored. Real-world repositories contain a vast and diverse array of tools, making it difficult for LLMs to effectively search, distinguish, and compose tools under context-length constraints. We identify large-scale tool selection as a new challenge for agentic reinforcement learning, highlighting that existing RL methods for knowledge-based question answering are inadequate for selecting tools while considering compatibility. To address this challenge, we propose ToolSearcher, a novel RL framework for effective multi-turn search and fine-grained optimization in large-scale tool selection. Specifically, we introduce category-constrained tool discrimination to improve the model's ability to distinguish functionally similar tools, event-level search modeling to explicitly optimize the discovery of target tools during multi-turn search, and trajectory-aligned credit allocation to provide fine-grained reward signals for different stages of the search-selection process. Extensive experiments on large-scale tool selection benchmarks demonstrate that ToolSearcher consistently outperforms a set of strong baselines in challenging settings involving iterative search and complex tool composition.
Chinese Translation
大语言模型(LLMs)擅长自然语言处理,但在与外部环境交互方面仍存在困难。工具学习为将 LLMs 扩展为可执行行动的智能体提供了一条有前景的途径,其中工具选择是成功使用工具的关键前提。现有工作通常假设工具集合规模较小或是预定义的,这使得大规模工具选择尚未得到充分探索。现实世界的工具库包含大量且多样的工具,使得 LLMs 在上下文长度约束下难以有效搜索、区分和组合工具。我们将大规模工具选择视为智能体强化学习的一个新挑战,并强调现有的用于基于知识的问答的 RL 方法在考虑兼容性的同时选择工具方面是不足的。为应对这一挑战,我们提出了 ToolSearcher,一种新颖的 RL 框架,用于在大规模工具选择中实现有效的多轮搜索和细粒度优化。具体而言,我们引入了类别约束的工具判别,以提高模型区分功能相似工具的能力;引入事件级搜索建模,以显式优化多轮搜索过程中对目标工具的发现;以及引入轨迹对齐的信用分配,为搜索-选择过程的不同阶段提供细粒度奖励信号。在大规模工具选择基准上的大量实验表明,在涉及迭代搜索和复杂工具组合的具有挑战性的设置中,ToolSearcher 始终优于一组强基线。
cs.CL / 29 / 2609.30935
Estimating and Orthogonalizing Unknown Pre-training Gradients for Continual Fine-tuning of Large Language Models
为大语言模型持续微调估计并正交化未知的预训练梯度
Bing Wang, Changchun Li, Xin-Qiang Cai, Lin Yuanbo Wu, Ximing Li, Gang Niu, Masashi Sugiyama
cs.CL · cs.AI · cs.LG
large language model
大语言模型相关
Abstract
Continual fine-tuning is essential for large language models (LLMs) to dynamically adapt to real-world environments, yet it inevitably suffers from catastrophic forgetting, particularly the performance degradation of previous tasks and LLMs' general-purpose knowledge. Although existing methods, such as orthogonal gradient projection, mitigate the forgetting across various fine-tuning tasks, they fundamentally fail to preserve pre-training LLMs' inherent general-purpose knowledge because the original data and gradients of off-the-shelf pre-training LLMs required by these methods are strictly unknown and highly diverse. To bridge this critical gap, we propose EoupCT, a novel framework designed to Estimate and Orthogonalize Unknown Pre-training gradients for Continual LLM fine-Tuning. Specifically, EoupCT estimates pre-training gradients by dynamically generating pseudo data that is most susceptible to forgetting for new tasks through a learnable soft prompt equipped with Gumbel-Softmax relaxation. Furthermore, we formulate a multi-objective optimization problem and introduce a first-order efficient Pareto optimizer that jointly optimizes LLM parameters and the soft prompt, rigorously enforcing orthogonality between new task updates and the estimated pre-training gradients. Extensive experiments across multiple LLMs demonstrate that EoupCT effectively preserves both task-specific proficiency and inherent general-purpose knowledge, successfully mitigating the catastrophic forgetting.
Chinese Translation
持续微调对于大语言模型(LLMs)动态适应真实世界环境至关重要,然而它不可避免地会遭受灾难性遗忘,尤其是先前任务性能以及大语言模型通用知识的退化。尽管现有方法,例如正交梯度投影,缓解了各种微调任务中的遗忘,但它们从根本上无法保留预训练大语言模型固有的通用知识,因为这些方法所需的现成预训练大语言模型的原始数据和梯度是严格未知且高度多样的。为弥合这一关键空白,我们提出了 EoupCT,一个旨在为大语言模型持续微调估计并正交化未知预训练梯度的新颖框架。具体而言,EoupCT 通过一个配备 Gumbel-Softmax 松弛的可学习软提示,动态生成对新任务而言最易被遗忘的伪数据,从而估计预训练梯度。此外,我们构建了一个多目标优化问题,并引入一种一阶高效帕累托优化器,联合优化大语言模型参数与软提示,严格强制新任务更新与所估计的预训练梯度之间的正交性。在多个大语言模型上的大量实验表明,EoupCT 有效地同时保留了任务特定的专有能力与固有的通用知识,成功缓解了灾难性遗忘。
cs.CL / 30 / 2609.30968
FAVoR: Measuring and Mitigating Author-Style Homogenization in Federated Personalized Generation
FAVoR:测量并缓解联邦个性化生成中的作者风格同质化
Lu Han, Jingyao Zhang, Katy Ilonka Gero, Nguyen H. Tran
cs.CL
large language model
大语言模型相关
Abstract
Large language models are increasingly used as personalized writing assistants, but adapting a model across many authors can compromise individual writing style by pulling author-specific signals toward a shared register. Federated parameter-efficient fine-tuning (PEFT) offers a data-local setting for this multi-author adaptation problem: clients keep author text local while sharing compact adapter updates. However, we show that standard aggregation can preserve continuation utility while making different authors' generations less distinguishable in style space, a failure mode we define as author-style homogenization. We evaluate author-style retention with Angular Style Classification Encoder (ASCE)-based diagnostics on our main BlogText benchmark and ASCE-independent external authorship verification. Using this protocol, we find that common federated PEFT baselines can preserve semantic utility while averaging out author-specific signals. To address this homogenization, we instantiate FAVoR (Federated Authorial Voice Retention), an author-style residual mechanism for federated PEFT. FAVoR uses a shared-private adapter design: clients upload shared-adapter updates while retaining author-specific residual corrections locally. Across BlogText and external Mythos-Reddit validation, FAVoR improves author-style retention over standard and personalized federated PEFT baselines. These gains come with small continuation-utility trade-offs and are supported by component ablations, external verification, and cold-start transfer.
Chinese Translation
大型语言模型正越来越多地被用作个性化写作助手,但将模型跨许多作者进行适配可能会损害个体写作风格,因为它会将作者特定的信号拉向一种共享语域。联邦参数高效微调(PEFT)为这一多作者适配问题提供了一种数据本地化设置:客户端将作者文本保留在本地,同时共享紧凑的适配器更新。然而,我们表明,标准聚合可以在保持续写效用的同时,使不同作者的生成在风格空间中变得更难以区分,我们将这一失效模式定义为作者风格同质化。我们在我们的主 BlogText 基准上使用基于 Angular 风格分类编码器(ASCE)的诊断来评估作者风格保留,并进行独立于 ASCE 的外部作者身份验证。使用这一协议,我们发现常见的联邦 PEFT 基线能够在保持语义效用的同时,将作者特定的信号平均掉。为了解决这种同质化,我们实例化 FAVoR(Federated Authorial Voice Retention,联邦作者声音保留),这是一种用于联邦 PEFT 的作者风格残差机制。FAVoR 使用共享-私有适配器设计:客户端上传共享适配器更新,同时在本地保留作者特定的残差校正。在 BlogText 和外部 Mythos-Reddit 验证上,FAVoR 相比标准和个性化联邦 PEFT 基线提升了作者风格保留。这些增益伴随着较小的续写效用权衡,并得到组件消融、外部验证和冷启动迁移的支持。
cs.CL / 31 / 2609.30986
Evaluating Sycophancy in Chinese Large Language Models on Factual Questions Derived from Online Search Queries
评估中文大语言模型在源自在线搜索查询的事实性问题上的谄媚行为
Geng Liu, Feng Li, Mengxiao Zhu, Francesco Pierri
cs.CL
large language model
大语言模型相关
Abstract
As large language models increasingly mediate information access, factually accurate and independent answers are critical. However, these models can exhibit sycophancy by aligning their responses with users' stated beliefs even when those beliefs are incorrect, potentially presenting misinformation as independently verified and reinforcing users' confidence in false claims. Prior work leaves unresolved whether introducing user beliefs causes correct responses to become incorrect or uncertain, or causes uncertain responses to become belief-aligned incorrect answers. It also remains unclear whether anti-sycophancy interventions preserve or restore factual accuracy or merely shift responses toward uncertainty. We analyze factual sycophancy in Chinese-language information seeking using yes/no fact-checking questions. Our analysis covers 364,941 responses from three frontier Chinese-based LLMs (DeepSeek, Qwen, and Doubao) to 12,165 factual questions derived from real-world Chinese search queries. We evaluate the models with and without reasoning across baseline, belief-conditioned, and anti-sycophancy prompting, tracing matched shifts among correct, incorrect, and uncertain responses. Under incorrect user beliefs, we distinguish belief-aligned errors from losses of factual confidence, in which initially correct answers become uncertain. Patterns vary across models and reasoning settings: reasoning is not a consistent safeguard, and anti-sycophancy instructions can reduce incorrect agreement while increasing uncertainty. In Chinese-language factual question answering, avoiding agreement with false beliefs is therefore not equivalent to preserving factual accuracy, highlighting the value of transition-level evaluation. Such behavior may undermine the reliability of LLM-mediated information access by reinforcing misinformation or weakening users' confidence in factually correct answers.
Chinese Translation
随着大语言模型日益中介信息获取,事实准确且独立的回答至关重要。然而,这些模型可能表现出谄媚行为,即使用户陈述的信念是错误的,也会使其回答与这些信念保持一致,从而可能将错误信息呈现为经过独立验证的内容,并强化用户对错误说法的信心。既有研究尚未解决以下问题:引入用户信念是会使正确回答变为错误或不确定,还是会使不确定的回答变为与信念一致的错误答案。同样尚不清楚的是,反谄媚干预是保持或恢复了事实准确性,还是仅仅将回答转向不确定性。我们使用是/否事实核查问题,分析中文信息寻求中的事实性谄媚。我们的分析涵盖三个前沿中文大语言模型(DeepSeek、Qwen 和 Doubao)对 12,165 个源自真实世界中文搜索查询的事实性问题所给出的 364,941 条回答。我们在基线、信念条件化和反谄媚提示三种设置下,对有推理和无推理的模型进行评估,追踪正确、错误和不确定回答之间相匹配的转变。在用户信念错误的情况下,我们将与信念一致的错误与事实信心的丧失区分开来,后者指最初正确的答案变为不确定。这些模式因模型和推理设置而异:推理并非一致的保障,反谄媚指令可以减少错误的一致性,同时增加不确定性。因此,在中文事实性问题回答中,避免与错误信念一致并不等同于保持事实准确性,这凸显了转变层面评估的价值。此类行为可能通过强化错误信息或削弱用户对事实正确回答的信心,损害由大语言模型中介的信息获取的可靠性。
cs.CL / 32 / 2609.31002
ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker
ZooWork-ShopRanker:一个开放的、偏好对齐的电子商务重排序器
Siqiao Xue, Shuxuan Liu, Ning Hu
cs.CL
large language model
大语言模型相关
Abstract
Open rerankers trained for general web retrieval transfer imperfectly to e-commerce, where ranking decisions depend not only on topical relevance but also on user preferences, product constraints, and comparative product fit. These preference signals are difficult to supervise at scale: real search traffic provides authentic queries and candidates but no clean pairwise labels. We present ZooWork-ShopRanker, a family of e-commerce rerankers (0.6B, 4B, and 8B) aligned to judge-labeled shopping preference. Training pairs are labeled by a panel of reasoning large language models (LLMs) from different families acting as a preference oracle, with position-debiased judgments and agreement tiers, and the rerankers are trained on these labels. The aligned 8B flagship then serves as a distillation teacher for the efficient 4B and 0.6B models, which are fit to its scores and sharpened on judged pairs. To measure progress, we introduce ShopRank-Bench, a contamination-limited benchmark of ~10,000 private-traffic preference pairs in both text formats, tiered by how many judge families committed to each label. ZooWork-ShopRanker-8B and -4B significantly outperform the strongest open reranker baseline, every model significantly beats its own un-aligned base, and ZooWork-ShopRanker-0.6B beats its size peer; the gains hold in both formats and extend to common MTEB benchmarks. We release the models and the dual-format ShopRank-Bench to facilitate further research.
Chinese Translation
为通用网页检索训练的开源重排序器迁移到电子商务时表现并不完美,因为在电子商务中,排序决策不仅取决于主题相关性,还取决于用户偏好、产品约束以及产品间的比较适配度。这些偏好信号难以大规模监督:真实搜索流量提供了真实的查询和候选商品,但不提供干净的成对标签。我们提出 ZooWork-ShopRanker,这是一个电子商务重排序器系列(0.6B、4B 和 8B),其对齐目标是评判器标注的购物偏好。训练对由一组来自不同家族的推理大语言模型(LLMs)组成的专家小组标注,这些模型充当偏好预言机,并采用位置去偏的判断和一致性分级,然后在这些标签上训练重排序器。随后,对齐后的 8B 旗舰模型充当高效 4B 和 0.6B 模型的蒸馏教师,这些模型拟合其分数,并在经过评判的配对数据上进一步锐化。为了衡量进展,我们引入 ShopRank-Bench,这是一个污染受限的基准,包含约 10,000 个来自私有流量的偏好配对,涵盖两种文本格式,并按照有多少个评判模型家族认同每个标签进行分级。ZooWork-ShopRanker-8B 和 -4B 显著优于最强的开源重排序器基线,每个模型都显著优于其自身未对齐的基础模型,ZooWork-ShopRanker-0.6B 优于同规模同类模型;这些增益在两种格式中都成立,并扩展到常见的 MTEB 基准。我们发布这些模型以及双格式 ShopRank-Bench,以促进进一步研究。
cs.CL / 33 / 2609.31009
G$^2$PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation
G$^2$PTQ:利用广义梯度补偿改进 LLM 训练后量化
Ruikang Liu, Haoli Bai, Yuxuan Sun, Qian Zhang, Wenzheng Cai, Yanqi Hao, Feiyu Wang, Weidong Zhong, Zhuang Wang, Tong Yang, Xiangsheng Zhou
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Post-training quantization (PTQ) is a practical approach to reducing the memory and computational footprint of large language models (LLMs) without retraining. GPTQ-based methods have become the de facto standard, yet they suffer from two complementary limitations. Methods with local, layer-wise objectives lack global supervision; while methods with global objectives fix their Hessian estimates at the start and ignore first-order gradients, so their guidance grows stale as quantization proceeds. This paper presents G$^2$PTQ, a unified PTQ framework with Generalized Gradient Compensation that integrates both first- and second-order information under a globally supervised, block-wise optimization objective. By refreshing gradient and Hessian estimates before quantizing each Transformer block, G$^2$PTQ avoids the staleness of prior global methods. Furthermore, to stabilize the exact first-order compensation, we introduce a trust-region scaling mechanism that dynamically bounds the gradient step to prevent exploding weight updates. Finally, we derive efficient implementations for block-wise Hessian approximation and exact gradient compensation. Experimental results on various model families and bit-widths demonstrate that G$^2$PTQ enables better alignment with the full-precision model, outperforming state-of-the-art baselines. Code is available at: https://github.com/G2PTQ/G2PTQ.
Chinese Translation
训练后量化(PTQ)是一种无需重新训练即可减少大语言模型(LLM)内存和计算占用的实用方法。基于 GPTQ 的方法已成为事实上的标准,但它们受到两个互补局限性的制约。具有局部、逐层目标的方法缺乏全局监督;而具有全局目标的方法在开始时固定其 Hessian 估计并忽略一阶梯度,因此随着量化进行,它们的指导会变得陈旧。本文提出 G$^2$PTQ,一个统一的 PTQ 框架,具有广义梯度补偿,它在全局监督的逐块优化目标下整合了一阶和二阶信息。通过在量化每个 Transformer 块之前刷新梯度和 Hessian 估计,G$^2$PTQ 避免了先前全局方法的陈旧性。此外,为了稳定精确的一阶补偿,我们引入了一种信赖域缩放机制,该机制动态约束梯度步长以防止权重更新爆炸。最后,我们为逐块 Hessian 近似和精确梯度补偿推导了高效实现。在不同模型家族和位宽上的实验结果表明,G$^2$PTQ 能够更好地与全精度模型对齐,性能优于最先进的基线。代码可在以下网址获取:https://github.com/G2PTQ/G2PTQ。
cs.CL / 34 / 2609.31046
Modeling Student Sensemaking with LLMs and Knowledge-Graph-Guided Inference
使用LLMs与知识图谱引导推理对学生意义建构进行建模
Özge Alacam, Zübeyde Demet Kirbulut Güneş, Funda Ekici, Nurcan Turan-Oluk, Dilay Dinçdemir, Hakkı Kadayıfçı, Sevinç Nihal Yeşiloğlu, Burcu Işık, Halil Tümay, Sinem Gencer
cs.CL
large language model
大语言模型相关
Abstract
Collaborative science learning requires nuanced interpretation of student dialogue to characterize how learners identify knowledge gaps, build explanations, and work toward resolution - a theory-driven analysis that is labor-intensive and difficult to scale. We investigate whether instruction-tuned large language models (LLMs) can support multidimensional analysis of collaborative sensemaking without task-specific training, and whether structured knowledge-state information improves model inference. We evaluate two mid-size LLMs on 23 richly annotated, expert-labeled episodes across prompting conditions that vary definitional scaffolding, reasoning mode, and turn structure. Without reasoning, models tend to overpredict successful sensemaking; reasoning-enabled prompting improves identification of unsuccessful cases. Knowledge-state diagnostics provide additional grounding, improving detection of unsuccessful sensemaking and increasing agreement with expert annotations. No single configuration performs best across all sensemaking dimensions, underscoring the multidimensional nature of the task.
Chinese Translation
协作式科学学习需要对学生的对话进行细致入微的解读,以刻画学习者如何识别知识缺口、构建解释并努力达成问题解决——这是一种理论驱动的分析,既耗费人力又难以规模化。我们研究经过指令微调的大语言模型(LLMs)能否在没有任务特定训练的情况下支持对协作式意义建构的多维分析,以及结构化的知识状态信息能否提升模型的推理能力。我们在23个标注丰富、由专家标注的片段上,跨多种提示条件(变动定义性支架、推理模式和轮次结构)评估两个中等规模LLM。在没有推理的情况下,模型倾向于过度预测成功的意义建构;启用推理的提示能改善对不成功案例的识别。知识状态诊断提供了额外的依据,改善了不成功意义建构的检测,并提高了与专家标注的一致性。没有任何单一配置在所有意义建构维度上都表现最佳,这凸显了该任务的多维性质。
cs.CL / 35 / 2609.31122
LocUS: Head Selection and Subspace Projection for Targeted Activation Steering
LocUS:用于目标激活引导的注意力头选择与子空间投影
Irene Tallini, Lorenzo Basile, Valentino Maiorca, Francesco Locatello, Alberto Cazzaniga
cs.CL · cs.LG
large language model
大语言模型相关
Abstract
Activation steering is a powerful training-free paradigm for controlling large language models at inference time. However, standard approaches estimate a per-layer steering direction from contrastive data and apply it on the layer's entire representation space, which may couple the intervention to off-target properties present in the contrastive data and degrade unrelated capabilities. To mitigate this issue, we introduce LocUS (Localized Unembedding Steering), a method which grounds activation steering to the model's own output vocabulary subspace. By identifying a property-specific linear subspace within the unembedding matrix, LocUS enforces a geometric constraint that restricts the steering transformation to a specific subspace and at the same time localizes its application to a sparse subset of attention heads. Extensive evaluations across three model families on toxicity mitigation, sentiment redirection and sycophancy suppression show that LocUS matches or outperforms state-of-the-art baselines while intervening on under 6% of parameters and better preserving general capability.
Chinese Translation
激活引导是一种强大的免训练范式,用于在推理时控制大语言模型。然而,标准方法从对比数据中估计每层的引导方向,并将其应用于该层的整个表示空间,这可能使干预与对比数据中存在的非目标属性耦合,并削弱无关能力。为缓解这一问题,我们引入 LocUS(Localized Unembedding Steering,局部化解嵌入引导),一种将激活引导锚定到模型自身输出词表子空间的方法。通过在解嵌入矩阵中识别特定属性的线性子空间,LocUS 施加几何约束,将引导变换限制到特定子空间,同时将其应用局部化到稀疏的注意力头子集。跨三个模型家族在毒性缓解、情感重定向和谄媚抑制上的广泛评估表明,LocUS 在匹配或优于最先进基线的同时,仅干预不到 6% 的参数,并更好地保持通用能力。
cs.CL / 36 / 2609.31169
Improving Visual Sensitivity of LLMs on Multimodal Machine Translation with Metric-based Loss Weighting
基于度量损失加权提升大语言模型在多模态机器翻译中的视觉敏感性
Paweł Mąka, Piotr Andruszkiewicz, Yusuf Can Semerci, Jan Scholtes, Gerasimos Spanakis
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Multimodal Machine Translation aims to incorporate additional signal from non-textual modalities to improve translations by resolving ambiguities. While models, through multimodal fusion, are able to accept images related to the source text, they can ignore this information. Therefore, increasing their visual sensitivity remains an active research area. In this work, we introduce a training method, Metric-based Loss Weighting, that improves visual grounding of translations by increasing the loss function for tokens that benefit from the accompanying image. We identify these tokens using the Point-wise Cross-mutual Information (PCXMI) metric, which compares the model's output probabilities with and without visual context. We introduce a Congruency-based PCXMI metric and experimentally show that both metrics working in combination yield the best results. We evaluate our method by fine-tuning three pretrained Multimodal Large Language Models on the task of Image-guided Machine Translation for three language directions. Metric-based Loss Weighting outperforms other tested methods on the CoMMuTE contrastive dataset, improving accuracy by up to more than 7 percentage points compared to standard fine-tuning, while maintaining strong general translation performance.
Chinese Translation
多模态机器翻译旨在引入来自非文本模态的额外信号,通过消除歧义来改进翻译。尽管模型通过多模态融合能够接受与源文本相关的图像,但它们可能会忽略这些信息。因此,提高其视觉敏感性仍然是一个活跃的研究领域。在这项工作中,我们提出了一种训练方法——基于度量的损失加权(Metric-based Loss Weighting),它通过提高那些受益于伴随图像的 token 的损失函数,来改善翻译的视觉接地。我们使用逐点交叉互信息(Point-wise Cross-mutual Information, PCXMI)度量来识别这些 token,该度量比较模型在有和没有视觉上下文时的输出概率。我们引入了一种基于一致性的 PCXMI 度量,并通过实验表明,两种度量结合使用能取得最佳结果。我们通过在三个语言方向上的图像引导机器翻译任务上微调三个预训练多模态大语言模型来评估我们的方法。在 CoMMuTE 对比数据集上,基于度量的损失加权优于其他所测试的方法,与标准微调相比准确率最高提升超过 7 个百分点,同时保持强劲的通用翻译性能。
cs.CL / 37 / 2609.31245
RupeeBias: Auditing Demographic Bias in Indian Economic Guidance from Large Language Models
RupeeBias:审计大型语言模型在印度经济指导中的人口统计学偏见
Pavithra P M Nair, Bhavik Talaviya, Shourya Bhushan, Rahul Pankajakshan, Seema Guruvadoo, Avinash Agarwal, Gilad Gressel, Krishnashree Achuthan
cs.CL
large language model
大语言模型相关
Abstract
Individuals turn to large language models (LLMs) for guidance across a wide range of economic tasks, from comparing loan options and planning savings to deciding what raise to ask for or how much to charge for their services. LLMs are known to reproduce social biases, and biased economic guidance may influence what users believe they are worth, what they ask for, and what they ultimately accept. This risk is especially salient in India, where economic outcomes are shaped by demographic categories such as caste and urban-rural location. Existing LLM bias benchmarks, however, are largely designed around Western demographic categories and therefore miss key axes of economic disparity in the Indian context. We introduce RupeeBias, a benchmark for auditing demographic bias in LLM-generated economic guidance across Indian economic settings. RupeeBias consists of 39,150 prompts spanning four use cases: salary estimation, salary increment estimation, counter-offer recommendation, and service pricing recommendation. The benchmark follows a single-attribute counterfactual design, holding the description of the user's qualifications, experience, or service offering fixed while varying one demographic identifier at a time. RupeeBias covers 87 India-specific demographic identifiers across six axes: caste, religion, regional identity, gender, disability, and urban-rural location, with all prompts constructed in both English and Hinglish. We evaluate nine LLMs on RupeeBias and find systematic demographic disparities across all six axes. For otherwise identical prompts that differ only in demographic identifier, LLM-generated economic outputs differ by 20.2% on average. We publicly release RupeeBias to support future research on demographic bias in LLM-generated economic guidance across India-specific demographic and economic contexts.
Chinese Translation
个人会求助于大型语言模型(LLM)来获得广泛经济任务中的指导,从比较贷款方案、规划储蓄,到决定该要求多少加薪,或为自己的服务收取多少费用。已知 LLM 会重现社会偏见,而有偏见的经济指导可能会影响用户认为自己值多少、要求多少,以及最终接受多少。这种风险在印度尤为突出,因为印度的经济结果受到种姓和城乡位置等人口统计学类别的塑造。然而,现有的 LLM 偏见基准大多围绕西方人口统计学类别设计,因此遗漏了印度语境中经济差异的关键轴线。我们推出 RupeeBias,一个用于审计 LLM 生成的、跨越印度经济环境的经济指导中人口统计学偏见的基准。RupeeBias 包含 39,150 条提示,涵盖四个用例:薪资估计、薪资增幅估计、还价推荐和服务定价推荐。该基准遵循单属性反事实设计,在保持用户资质、经验或服务提供描述不变的同时,每次只改变一个人口统计学标识符。RupeeBias 覆盖六大轴线上的 87 个印度特有的人口统计学标识符:种姓、宗教、地区身份、性别、残疾和城乡位置,且所有提示均以英语和印度英语混合语(Hinglish)两种语言构建。我们在 RupeeBias 上评估了九个 LLM,并发现所有六大轴线均存在系统性的人口统计学差异。对于除人口统计学标识符外其他方面完全相同的提示,LLM 生成的经济输出平均相差 20.2%。我们公开发布 RupeeBias,以支持未来关于 LLM 生成的、跨越印度特有人口统计学和经济语境的经济指导中人口统计学偏见的研究。
cs.CL / 38 / 2609.31382
Highlight-Then-Summarize: Learning to Compress Evidence for Long-Context Understanding
先标注再总结:学习压缩证据以实现长上下文理解
Zhaoyuan Xia, Qinghongbing Xie, Yung Xiang Hue, Jianguang Jiang, Gaofeng Lu, Zhenyu Jiao, Xing Yuan, Dai Dai, Tong Mo, Long Zeng
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Long-context understanding requires large language models (LLMs) to reason over lengthy documents, conversations, and code, yet task-relevant evidence is often sparse and scattered amid substantial irrelevant and redundant content. We propose Highlight-Then-Summarize (H2S), a compress-then-reason paradigm that first identifies source-grounded, question-relevant evidence and then integrates it into a compact, question-conditioned summary before producing the final answer. To train this behavior, we construct H2S-Dataset, comprising 6,647 examples from 11 benchmark families with an average context length of 43.9K tokens, and introduce H2S-RL, which provides process-level rewards for evidence selection and summary construction in addition to final-answer correctness. We evaluate on H2S-Bench, a seven-task long-context suite. Under a shared 128K input and 4K output budget, H2S-14B achieves an average score of 32.60, outperforming Qwen3.8-27B by 10.17 points and obtaining the strongest overall result among the evaluated open-source models. H2S-14B also achieves the highest Evidence-Summary Quality score and retains 97.1% of its 16K-budget performance with only a 4K output budget. These results show that explicitly selecting and integrating evidence improves long-context reasoning while enabling more compact generation.
Chinese Translation
长上下文理解要求大语言模型(LLM)对冗长的文档、对话和代码进行推理,然而与任务相关的证据往往十分稀疏,且分散在大量无关和冗余的内容之中。我们提出 Highlight-Then-Summarize(H2S),一种先压缩再推理的范式:它首先识别有原文依据、与问题相关的证据,然后将其整合为一个紧凑的、以问题为条件的摘要,之后才生成最终答案。为训练这一行为,我们构建了 H2S-Dataset,其中包含来自 11 个基准家族的 6,647 个示例,平均上下文长度为 43.9K token,并提出 H2S-RL,它在最终答案正确性之外,还为证据选择与摘要构建提供过程级奖励。我们在 H2S-Bench 上进行评估,这是一个包含七项任务的长上下文测试套件。在共享的 128K 输入与 4K 输出预算下,H2S-14B 取得 32.60 的平均分,比 Qwen3.8-27B 高出 10.17 分,并在所评估的开源模型中取得最强的整体结果。H2S-14B 还取得了最高的证据-摘要质量(Evidence-Summary Quality)分数,并且在仅有 4K 输出预算的情况下仍保留了其 16K 预算性能的 97.1%。这些结果表明,显式地选择与整合证据可以提升长上下文推理能力,同时实现更紧凑的生成。
cs.CL / 39 / 2609.31448
ViSTA: A Simple Bridge Extends Visual Alignment to Clinical Time-Series Understanding in Multimodal LLMs
ViSTA:一个简单的桥接将视觉对齐扩展到多模态大语言模型中的临床时间序列理解
Junyi Gao, Yu Shi, Pingzhao Hu, Ewen M Harrison
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Clinical prediction models estimate risk from patient measurements, while large language models support medical text understanding and question answering. Yet their language capabilities do not ensure accurate prediction from structured, high-dimensional clinical time series. Improving this ability would connect risk estimation with flexible questions about a patient's evolving condition. We introduce ViSTA, a compact adapter that incorporates irregular numerical measurements into a pretrained vision-language model's chart representations. It learns corrections to visual tokens while leaving all pretrained parameters unchanged. On MIMIC-IV, ViSTA has the highest mean scores among the compared adaptations on all four metrics for acute kidney injury and mortality prediction across models with 2-9 billion parameters. With 0.516 million trainable parameters, the 2-billion-parameter model reaches an area under the ROC curve of 0.7376 for acute kidney injury, compared with GPT-5.6 Sol's 0.7380 with text input and high reasoning effort. Training for temporal question answering yields 69.27% accuracy at 4 billion parameters with over 90% fewer trainable parameters than low-rank adaptation using charts or numerical text, at a 2.82-4.88 percentage-point accuracy gap. ViSTA extends pretrained language models to numerical prediction and temporal questions.
Chinese Translation
临床预测模型根据患者测量值估计风险,而大语言模型支持医学文本理解和问答。然而,它们的语言能力并不能确保基于结构化、高维临床时间序列进行准确预测。提升这一能力将把风险估计与关于患者病情演变的灵活问题联系起来。我们提出 ViSTA,一种紧凑适配器,将不规则数值测量纳入预训练视觉-语言模型的图表表示中。它学习对视觉 token 的校正,同时保持所有预训练参数不变。在 MIMIC-IV 上,对于急性肾损伤和死亡率预测,在参数规模为 20 亿至 90 亿的模型中,ViSTA 在所有四项指标上取得了相比的适配方法中最高的平均分数。在仅有 51.6 万个可训练参数的情况下,20 亿参数模型在急性肾损伤上达到 0.7376 的 ROC 曲线下面积,而 GPT-5.6 Sol 在使用文本输入和高推理努力时为 0.7380。针对时间问答进行训练,在 40 亿参数下取得了 69.27% 的准确率,其可训练参数比使用图表或数值文本的低秩适配少 90% 以上,准确率差距为 2.82–4.88 个百分点。ViSTA 将预训练语言模型扩展到数值预测和时间问题。
cs.CL / 40 / 2609.31506
Evaluating Cultural Awareness of LLMs for Haitian Creole
评估大语言模型对海地克里奥尔语的文化意识
Christelle Clervilsson, Yanzhu Guo
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Large language models (LLMs) exhibit substantial performance disparities between high- and low-resource languages. Beyond lower task performance, they often fail to capture the cultural norms and values of underrepresented communities. In this work, we present the first systematic evaluation of cultural awareness in LLMs for Haitian Creole, a language spoken by millions but severely underrepresented in digital resources. We assess cultural awareness along four complementary dimensions---specificity, bias, diversity, and variation---using a benchmark of culturally salient prompts curated by native speakers in a text infilling setting. Our results reveal a clear gap between cultural awareness in Haitian Creole and higher-resource French, with Haitian performance being more uneven across domains and more affected by French linguistic interference. Story generation further reveals recurring portrayals of Haitian characters through hardship and resilience, showing that even positive characterizations can encode stereotypical narratives. Our code, benchmark, and evaluation framework are publicly available.
Chinese Translation
大语言模型(LLMs)在高资源语言和低资源语言之间表现出显著的性能差异。除了任务性能较低之外,它们往往无法捕捉代表性不足群体的文化规范和价值观。在这项工作中,我们首次对 LLMs 在海地克里奥尔语上的文化意识进行了系统性评估;海地克里奥尔语有数百万人使用,但在数字资源中严重代表性不足。我们沿着四个互补维度——特异性、偏见、多样性和变异性——评估文化意识,使用由母语者在文本填空设置中整理的文化显著提示基准。我们的结果揭示了海地克里奥尔语与更高资源法语之间的文化意识存在明显差距,海地克里奥尔语的表现在不同领域间更不均衡,且更受法语语言干扰的影响。故事生成进一步揭示了对海地人物的反复描绘:通过苦难与韧性来刻画,表明即使是正面刻画也可能编码刻板化叙事。我们的代码、基准和评估框架已公开可用。
cs.CR / 41 / 2609.30657
Prompt Injection Detection for Email Agents Through Attack Chain Modeling
通过攻击链建模的电子邮件代理提示注入检测
Ahmad Hashmi, Dhyey Patel, Yunting Yin
cs.CR · cs.CL
large language model
大语言模型相关
Abstract
Large language model email assistants are particularly vulnerable to indirect prompt injection because untrusted email content can be retrieved into the model context and influence subsequent tool use. Existing prompt injection detectors mainly formulate this problem as binary malicious text classification, which overlooks the important factor that harmful agent behavior often arises through a sequence of stages. We propose a detection framework that models this attack chain by combining a text detector, verifiers specific to each stage, explicit rule-based risk signals, user intent and action consistency analysis, and a logistic decision policy. To support this framework, we derive attack chain labels from prompt injection datasets, evaluate the proposed framework under random splits, temporal phase transfer, conditional stage transfer, cross-dataset transfer, and conduct ablation studies on multiple benchmarks. Results show that random train test splits substantially overestimate robustness under distribution shift, while later tool argument stages are more predictable than earlier stages in the framework. We also show that training on harmless emails that resemble attacks helps reduce false alarms while preserving the ability to detect real attacks. Across five binary benchmarks, our framework achieves a mean F1 score of 0.406 under the strict threshold setting policy, compared with 0.216 for the strongest of five pretrained detectors evaluated without additional training. These results highlight the value of combining attack stage predictions with checks for conflicts between the user's request and instructions in retrieved emails. Our experiments also demonstrate the importance of training with challenging benign examples to balance attack detection and false alarms.
Chinese Translation
大型语言模型电子邮件助手尤其容易受到间接提示注入的影响,因为不可信的电子邮件内容可被检索到模型上下文中,并影响后续的工具使用。现有的提示注入检测器主要将这一问题表述为二元恶意文本分类,这忽略了有害智能体行为往往通过一系列阶段产生这一重要因素。我们提出了一种检测框架,通过结合文本检测器、针对每个阶段的验证器、显式的基于规则的风险信号、用户意图与动作一致性分析,以及逻辑决策策略,对这一攻击链进行建模。为支持这一框架,我们从提示注入数据集中推导出攻击链标签,在随机划分、时间阶段迁移、条件阶段迁移、跨数据集迁移下评估所提出的框架,并在多个基准上进行消融研究。结果表明,随机训练-测试划分大幅高估了分布偏移下的鲁棒性,而在该框架中,较晚的工具参数阶段比较早的阶段更可预测。我们还表明,在类似于攻击的无害电子邮件上进行训练有助于减少误报,同时保持检测真实攻击的能力。在五个二元基准上,在严格阈值设定策略下,我们的框架实现了 0.406 的平均 F1 分数,相比之下,五个预训练检测器中在不进行额外训练的情况下评估的最强检测器为 0.216。这些结果突出了将攻击阶段预测与对用户请求和检索到的电子邮件中的指令之间冲突的检查相结合的价值。我们的实验还证明了使用具有挑战性的良性示例进行训练以平衡攻击检测和误报的重要性。
cs.CR / 42 / 2609.30909
Machine Unlearning for Large Language Models: Foundations, Advances, and Agentic Extensions
面向大语言模型的机器遗忘:基础、进展与智能体式扩展
Xiaoyu Xu, Minxin Du, Li Bai, Junxu Liu, Yaxin Xiao, Kun Fang, Liu Yang, Huadi Zheng, Peizhao Hu, Qingqing Ye, Haibo Hu
cs.CR
large language model
大语言模型相关
Abstract
Machine unlearning aims to remove target influence while preserving other capabilities. This survey compares methods, benchmarks, and evidence across large language models and systems using retrieval, memory, tools, and interacting agents. A five-layer framework connects removal requests, system boundaries, target locations, interventions, and supported claims. A seven-stage lifecycle and six evidence dimensions guide comparison. The review shows that target construction, retained data, and recovery tests affect reported outcomes. Evidence from model evaluations remains insufficient to establish removal across external state and subsequent updates, motivating evaluation that tracks dependencies and tests whether target influence returns.
Chinese Translation
机器遗忘旨在移除目标影响,同时保留其他能力。本综述比较了跨大语言模型以及使用检索、记忆、工具和交互式智能体的系统的方法、基准和证据。一个五层框架将移除请求、系统边界、目标位置、干预措施和得到支持的主张连接起来。一个七阶段生命周期和六个证据维度指导比较。该综述表明,目标构建、保留数据和恢复测试会影响所报告的结果。来自模型评估的证据仍不足以确立跨外部状态和后续更新的移除,这促使人们进行评估,以追踪依赖关系并检验目标影响是否会返回。
cs.CR / 43 / 2609.30980
FeatMark: Feature-level Watermark Protection against Mimicry Attacks with Diffusion Models
FeatMark:针对基于扩散模型的模仿攻击的特征级水印保护
Haoyang Li, Ruoxi Sun, Qingqing Ye, Benjamin Zi Hao Zhao, Yaxin Xiao, Jason Xue, Haibo Hu
cs.CR · cs.CV
diffusion
扩散模型相关
Abstract
Text-to-image diffusion models enable data-efficient "mimicry" attacks, wherein adversaries fine-tune the model on a handful of public photos to synthesize convincing forgeries of a target individual. A common countermeasure is to embed imperceptible, low-energy watermarks, yet recent studies show these signatures are brittle: modest post-processing or lightweight adversarial perturbations readily suppress detection, exposing a fundamental tension between imperceptibility and robustness. We introduce FeatMark, a watermarking framework that shifts from pixel-level, energy-starved perturbations to inconspicuous semantic features: small, scene-consistent micro-features that remain natural to humans while providing a stronger, machine-verifiable provenance signal. FeatMark builds domain-specific feature banks that encode each watermark as a compact concept program, pairing open-vocabulary semantic cues with reliable edit regions and instruction templates. It then automatically selects features that are both feasible and executable and injects them through modular, mask-guided concept editing, yielding highly localized, scene-consistent micro-edits that are difficult to perceive. We conduct extensive experiments across VGGFace2, CelebA-HQ, and WikiArt, evaluating against 10 strong watermark removal/purification attacks (including regeneration-style purification) and several bespoke adaptive attacks tailored to FeatMark, to assess perceptual fidelity, watermark detection accuracy, and robustness. We further demonstrate FeatMark's extensibility to video mimicry attacks. The results show FeatMark remains virtually impervious, withstanding all evaluated attacks with negligible bit-accuracy and fidelity degradation.
Chinese Translation
文生图扩散模型使数据高效的“模仿”攻击成为可能,其中攻击者仅用少量公开照片微调模型,就能合成目标个体的逼真伪造图像。一种常见的对策是嵌入不可感知的低能量水印,然而近期研究表明这些签名很脆弱:适度的后处理或轻量级对抗扰动就能轻易抑制检测,暴露出不可感知性与鲁棒性之间的根本性张力。我们提出 FeatMark,一种水印框架,它从像素级、能量匮乏的扰动转向不显眼的语义特征:小型的、场景一致的微特征,这些特征对人类而言仍然自然,同时提供更强的、可机器验证的来源信号。FeatMark 构建领域特定的特征库,将每个水印编码为紧凑的概念程序,将开放词汇语义线索与可靠的编辑区域和指令模板配对。然后,它自动选择既可行又可执行的特征,并通过模块化的、掩码引导的概念编辑将它们注入,从而产生高度局部化、场景一致且难以察觉的微编辑。我们在 VGGFace2、CelebA-HQ 和 WikiArt 上进行了大量实验,针对 10 种强水印去除/净化攻击(包括再生式净化)以及若干为 FeatMark 量身定制的专门自适应攻击进行评估,以衡量感知保真度、水印检测准确率和鲁棒性。我们进一步展示了 FeatMark 对视频模仿攻击的可扩展性。结果表明,FeatMark 几乎不受影响,能够抵御所有评估过的攻击,且比特准确率和保真度下降可忽略不计。
cs.CR / 44 / 2609.30997
Can Pixels Alone Reveal Image Origin? Minimax Limits and Learnable Interfaces for Passive Provenance
仅凭像素就能揭示图像来源吗?被动溯源的极小极大极限与可学习接口
Kai Yao
cs.CR · cs.AI · cs.CV · cs.LG
diffusion
扩散模型相关
Abstract
Passive image provenance asks whether pixels alone can reveal where an image came from: a human, an aggregate AI class, or a particular generator. This becomes a robustness problem once a source image can be edited before the verifier sees it. We study the problem as source--target verification under adversarial distribution shift. Our first result gives the exact best-case limit for any image-only verifier: the largest robust target-acceptance gap equals the minimum total-variation distance between the target distribution and the set of attacked source distributions. This quantity depends on the source, target, and edit class, not on the verifier architecture. Our second result explains why deployed public verifiers can fail before this statistical limit is reached. If the verifier can be emulated on the attack region to error $\varepsilon$, then a surrogate black-box attack reaches target acceptance within $2\varepsilon$ plus optimization error of the white-box optimum; score-revealing logistic and softmax heads over public features are identifiable, and approximate score access gives stable recovery bounds. A finite-state experiment checks the minimax identity where both sides are computable. On same-prompt real/diffusion benchmarks, the evaluated public CLIP verifiers fail under targeted pixel attacks, while a ResNet-18 victim exhibits partial fake-to-real transfer. Binary feedback with abstention reduces measured attack success, but positive empirical gap upper bounds do not establish robustness. These results motivate separate evaluation of the source--target statistical ceiling and the information released by a deployed verifier.
Chinese Translation
被动图像溯源追问:仅凭像素本身能否揭示图像来自何处:人类、一个聚合的 AI 类别,还是某个特定生成器。一旦源图像在验证器看到它之前可以被编辑,这就成为一个鲁棒性问题。我们将该问题作为对抗性分布偏移下的源--目标验证来研究。我们的第一个结果给出了任何仅基于图像的验证器的精确最优情形极限:最大的鲁棒目标接受差距等于目标分布与受攻击源分布集合之间的最小全变差距离。该量取决于源、目标和编辑类别,而不取决于验证器架构。我们的第二个结果解释了为什么已部署的公开验证器可能在这一统计极限达到之前就失败。如果验证器可以在攻击区域上被模拟到误差 $\varepsilon$,那么一个替代黑盒攻击达到目标接受度,其差距在 $2\varepsilon$ 加上白盒最优的优化误差之内;公开特征上的得分揭示型 logistic 和 softmax 头是可识别的,并且近似得分访问给出稳定的恢复界。一个有限状态实验在两侧均可计算处检验了极小极大恒等式。在相同提示词的真实/扩散基准上,所评估的公开 CLIP 验证器在定向像素攻击下失败,而一个 ResNet-18 受害模型表现出部分假到真迁移。带弃权的二元反馈降低了测得的攻击成功率,但正的经验差距上界并不确立鲁棒性。这些结果促使我们分别评估源--目标统计上限以及已部署验证器所释放的信息。
cs.CR / 45 / 2609.31262
Deduplication-while-Training: A Resilient Paradigm for Privacy-Preserving Cross-Client Deduplication in Federated Learning
边训练边去重:联邦学习中面向隐私保护跨客户端去重的弹性范式
Rongxi Wang, Guanxiong Ha, Chunfu Jia, Yongsheng Lin, Minfen Gao, Hanmiaomiao Wang
cs.CR · cs.DC
large language model
大语言模型相关
Abstract
Cross-client duplicate data in large language model training corpora degrades the efficiency of federated learning (FL) while exacerbating model memorization and privacy risks. Privacy-preserving cross-client deduplication effectively mitigates this issue by eliminating duplicate training data. However, existing schemes all follow a "Deduplication-before-Training" paradigm. This serially coupled paradigm incurs high fault-tolerance costs and lacks support for dynamic client joining. To this end, we propose an unexplored paradigm called "Deduplication-while-Training (DwT)", which enables concurrent deduplication and training. DwT transforms cross-client deduplication from a one-time, globally synchronous preprocessing operation into a continuous online service with state management, concurrent claiming, and failure recovery. By enabling state synchronization and task takeover, it minimizes the impact of client disconnections on the overall training progress while supporting the dynamic joining of clients. We design DwT-FL, a privacy-preserving deduplication system, to support DwT. By designing a concurrent state-claim mechanism and a hot-cold dual-queue scheduling strategy, DwT-FL enables the parallel execution of secure deduplication and model training, while effectively handling client disconnections and dynamic joins. Experimental evaluations demonstrate that, compared to the state-of-the-art scheme, DwT-FL significantly reduces the time overhead of failure recovery and dynamic joining by up to 93.04% and 94.18%, respectively. This provides an efficient and elastic concurrent deduplication scheme for dynamic and unstable FL environments.
Chinese Translation
大语言模型训练语料中的跨客户端重复数据会降低联邦学习(FL)的效率,同时加剧模型记忆与隐私风险。隐私保护跨客户端去重通过消除重复训练数据有效缓解了该问题。然而,现有方案均遵循“先去重后训练”(Deduplication-before-Training)范式。这种串行耦合的范式带来了高昂的容错成本,并且缺乏对客户端动态加入的支持。为此,我们提出了一种尚未被探索的范式,称为“边训练边去重”(Deduplication-while-Training,DwT),它能够实现去重与训练的并发执行。DwT 将跨客户端去重从一次性的、全局同步的预处理操作转变为一种具有状态管理、并发认领与故障恢复能力的持续在线服务。通过实现状态同步与任务接管,它在支持客户端动态加入的同时,将客户端断连对整体训练进度的影响降至最低。我们设计了 DwT-FL,一个支持隐私保护的去重系统,用以支撑 DwT。通过设计并发状态认领机制与冷热双队列调度策略,DwT-FL 能够实现安全去重与模型训练的并行执行,同时有效处理客户端断连与动态加入。实验评估表明,与最先进的方案相比,DwT-FL 将故障恢复与动态加入的时间开销分别显著降低了高达 93.04% 和 94.18%。这为动态且不稳定的联邦学习环境提供了一种高效且弹性的并发去重方案。
cs.CR / 46 / 2609.31282
Resource-Optimized and Energy-Aware Agentic AI Framework Anchored on Blockchain for Secure Software Supply Chains
面向安全软件供应链的、锚定于区块链的资源优化与能量感知智能体 AI 框架
Toqeer Ali Syed, Asadullah Abdullah Khan
cs.CR · cs.AI
large language model
大语言模型相关
Abstract
This paper proposes a blockchain-backed agentic security framework designed to safeguard the complete software development lifecycle (SDLC) while also securing the agentic AI components responsible for monitoring it. The framework coordinates a set of specialised security agents, covering source integrity, dependency and SBOM analysis, CI configura tion auditing, artifact verification, and runtime policy evaluation, each supported by a large language model (LLM) that interprets artefacts, reasons over tool outputs, and produces structured security reports. To ensure agent trustworthiness, every agent generates a cryptographically signed attestation that is recorded in a permissioned blockchain via smart contracts, including an agent registry, an immutable attestation log, and an enforceable release-policy module. Communication among agents and with blockchain nodes is secured using a consortium-operated certificate authority, ensuring authenticated and tamper-resistant interactions. A detailed use-case and sequence flow demonstrate how a source code security agent performs analysis, anchors its attestation on-chain, and triggers a verifiable allow/block deployment decision. The proposed framework of fers decentralised integrity transparent provenance, uninterrupted security assurance and a generalisable architecture to incorporate the agentic AI into the modern software supply chain security.
Chinese Translation
本文提出一个由区块链支持的智能体安全框架,旨在保护完整的软件开发生命周期(SDLC),同时也保护负责监控该生命周期的智能体 AI 组件。该框架协调一组专门的安全智能体,涵盖源代码完整性、依赖项与 SBOM 分析、CI 配置审计、制品验证和运行时策略评估,每个智能体均由一个大型语言模型(LLM)支持,该模型解释制品、对工具输出进行推理,并生成结构化安全报告。为确保智能体的可信度,每个智能体都会生成一个经密码学签名的证明,该证明通过智能合约记录在许可区块链中,包括智能体注册表、不可变的证明日志以及可强制执行的发布策略模块。智能体之间以及与区块链节点之间的通信通过一个由联盟运营的证书颁发机构来保障安全,从而确保经过身份验证且防篡改的交互。一个详细的用例和序列流程展示了源代码安全智能体如何执行分析、将其证明锚定到链上,并触发可验证的允许/阻止部署决策。所提出的框架提供去中心化的完整性、透明溯源、不间断的安全保障,以及一种可泛化的架构,以将智能体 AI 纳入现代软件供应链安全之中。
cs.CR / 47 / 2609.31552
FragToken: Amplifying LLM Inference Costs through Noncanonical Token Generation
FragToken:通过非规范 Token 生成放大 LLM 推理成本
Zihan Wang, Rui Zhang, Xinyuan Qian, Qingchuan Zhao, Hongwei Li, Guowen Xu
cs.CR
large language model
大语言模型相关
Abstract
As large language model (LLM) inference becomes increasingly expensive, resource-consumption attacks pose a growing threat to model providers. Existing attacks typically amplify cost by inducing abnormally long or repetitive outputs on attacker-controlled or triggered requests, making them easier to detect and limiting their deployment-wide impact when benign traffic dominates. In this work, we uncover a previously overlooked token-level attack surface arising from the many-to-one mapping from token sequences to decoded text. Although standard LLMs predominantly generate the canonical token sequences induced by their tokenizers, the same text can also be represented by substantially longer non-canonical sequences. This representational flexibility exposes a new avenue for resource-consumption attacks: an attacker can train the model to favor such sequences, systematically increasing the number of autoregressive decoding steps without a proportional increase in visible response length. However, we empirically find that directly maximizing token fragmentation substantially degrades model utility, producing conspicuous answer-quality failures that undermine attack stealthiness. To address this challenge, we propose FragToken, a training-time framework that combines source-model self-distillation, capacity-aware filtering and budgeting, and BPE-Aligned Merging to induce fragmented generation under ordinary prompts while largely preserving model utility. We evaluate FragToken on four LLMs across three benchmarks. Across the four models, FragToken achieves a three-benchmark average token inflation ratio (TIR) ranging from 1.99 to 2.46, while causing only minor degradation in model utility. Our work reveals a covert LLM supply-chain threat that increases inference cost without requiring large volumes of attack requests while largely preserving utility.
Chinese Translation
随着大型语言模型(LLM)推理变得越来越昂贵,资源消耗攻击对模型提供商构成了日益增长的威胁。现有攻击通常通过在攻击者控制或触发的请求上诱导异常长或重复的输出以放大成本,这使其更容易被检测到,并在良性流量占主导时限制了其在部署范围内的影响。在这项工作中,我们揭示了一个此前被忽视的 token 级攻击面,它源于从 token 序列到解码文本的多对一映射。尽管标准 LLM 主要生成由其分词器诱导的规范 token 序列,但相同的文本也可以由长得多的非规范序列来表示。这种表示灵活性为资源消耗攻击开辟了一条新途径:攻击者可以训练模型偏好此类序列,从而系统地增加自回归解码步骤的数量,而不会相应地增加可见响应长度。然而,我们通过实验发现,直接最大化 token 碎片化会显著降低模型效用,产生明显的答案质量失败,从而削弱攻击的隐蔽性。为应对这一挑战,我们提出了 FragToken,一个训练时框架,它结合了源模型自蒸馏、容量感知过滤与预算,以及 BPE 对齐合并,以在普通提示下诱导碎片化生成,同时很大程度上保持模型效用。我们在四个 LLM 上跨三个基准评估了 FragToken。在四个模型上,FragToken 在三个基准上的平均 token 膨胀率(TIR)范围为 1.99 到 2.46,同时仅导致模型效用的轻微下降。我们的工作揭示了一种隐蔽的 LLM 供应链威胁,它增加了推理成本,而不需要大量攻击请求,同时很大程度上保持了效用。
cs.AI / 48 / 2609.30783
Skip the Talk, Re-Focus on Vision: Latent Reasoning for Reasoning Segmentation in Multimodal Large Language Models
跳过言语,重新聚焦视觉:面向多模态大语言模型推理分割的潜在推理
Tianhang Guo, Yulin He, Wei Chen, Wenjuan Zhou, Yuhang Li, Xinbiao Gan
cs.CV · cs.AI
large language model
大语言模型相关
Abstract
Reasoning segmentation aims to interpret implicit textual queries and enable fine-grained visual perception, which is critical for applications such as human-computer interaction and embodied agents. Existing methods typically generate explicit Chain-of-Thought (CoT) by multimodal large language models (MLLMs) before localizing the target. Although intuitive, such explicit verbal reasoning introduces substantial attention interference: redundant textual tokens disrupt attention during perception-token generation and also increase the effective distance between visual tokens. To address this issue, we propose LIRSeg, which fully replaces explicit CoT with a compact set of learnable latent tokens for reasoning segmentation. LIRSeg is trained in two stages: spatial alignment grounds the latent tokens in object-relevant visual evidence, and GRPO further optimizes them with segmentation rewards. To make these compact latent tokens more informative, we introduce three complementary mechanisms from an information perspective: extreme-advantage sampling for selecting informative training signals, decoupled exploration-stability updates for learning complementary representations, and latent diversity amplification for preventing representational collapse. Extensive experiments on benchmarks demonstrate that LIRSeg consistently improves both segmentation accuracy and reasoning efficiency. Compared with the VisionReasoner baseline, LIRSeg achieves absolute gIoU improvements of 4.9% on ReasonSeg, 7.1% on MUSE, and 4.7% on MMR, while achieving a approximately 16x reduction in reasoning tokens. Code is available in supplementary materials.
Chinese Translation
推理分割旨在解释隐式文本查询并实现细粒度视觉感知,这对于人机交互和具身智能体等应用至关重要。现有方法通常通过多模态大语言模型(MLLMs)生成显式思维链(Chain-of-Thought,CoT),然后再定位目标。尽管这种显式言语推理直观,但它会引入严重的注意力干扰:冗余的文本 token 会干扰感知 token 生成过程中的注意力,并且还会增加视觉 token 之间的有效距离。为了解决这一问题,我们提出了 LIRSeg,它用一组紧凑的可学习潜在 token 完全替代显式 CoT 来进行推理分割。LIRSeg 分两个阶段训练:空间对齐将潜在 token 锚定在与对象相关的视觉证据中,而 GRPO 则利用分割奖励进一步优化它们。为了使这些紧凑的潜在 token 更具信息量,我们从信息视角引入了三种互补机制:用于选择信息量丰富的训练信号的极端优势采样、用于学习互补表示的解耦探索-稳定性更新,以及用于防止表示崩溃的潜在多样性放大。在基准上的大量实验表明,LIRSeg 持续提升了分割准确率和推理效率。与 VisionReasoner 基线相比,LIRSeg 在 ReasonSeg 上实现了 4.9% 的 gIoU 绝对提升,在 MUSE 上提升 7.1%,在 MMR 上提升 4.7%,同时推理 token 数量减少了约 16 倍。代码可在补充材料中获取。
cs.AI / 49 / 2609.30941
Spackle: Completing Large View Single Image NVS with Adaptive Gaussians
Spackle:使用自适应高斯补全大视角单图像 NVS
Xuanzhi Liu, Yuhe Zhou, Xinyi Wu, Zhenyao Wu, Jinghao Chen, Ruize Han, Song Wang
cs.CV · cs.AI
diffusion
扩散模型相关
Abstract
Single-image novel view synthesis (NVS) enables photorealistic rendering of un- observed viewpoints from a single input. Practical NVS systems require two key capabilities: robust reconstruction of occluded regions and high inference effi- ciency. While hybrid decoupled frameworks combining feedforward 3D Gaussian Splatting (3DGS) and diffusion models show promise for large-view-deviation NVS, they suffer from capacity competition: a fixed number of Gaussians forces resource shifts from visible to newly disoccluded areas, degrading original scene fidelity when the target view deviates significantly from the input. To address this, we propose Spackle, a lightweight residual learning framework that mit- igates capacity competition without sacrificing efficiency. Spackle operates in three stages: predicting base 3DGS attributes from given views, automatically identifying poorly reconstructed regions, and learning a residual 3DGS optimized exclusively for these areas. At inference, we combine the baseline and aug- mented Gaussians for NVS. We conduct comprehensive experiments and show that Spackle achieves state-of-the-art performance on large-view-deviation cases.
Chinese Translation
单图像新视角合成(NVS)能够从单个输入实现对未观测视角的照片级真实感渲染。实用的 NVS 系统需要两项关键能力:对被遮挡区域的鲁棒重建以及高推理效率。虽然将前馈 3D 高斯泼溅(3DGS)与扩散模型相结合的混合解耦框架在大视角偏差 NVS 中展现出潜力,但它们遭受容量竞争:固定数量的高斯迫使资源从可见区域转向新近去遮挡区域,当目标视角与输入显著偏离时,会降低原始场景保真度。为了解决这一问题,我们提出 Spackle,一个轻量级残差学习框架,它在不牺牲效率的情况下缓解容量竞争。Spackle 分三个阶段运行:从给定视角预测基础 3DGS 属性,自动识别重建较差的区域,并学习专门针对这些区域优化的残差 3DGS。在推理时,我们结合基线高斯和增强高斯进行 NVS。我们进行了全面的实验,并表明 Spackle 在大视角偏差情况下达到了最先进的性能。
cs.AI / 50 / 2609.31135
Pocket-STVG: lightweight architecture for Spatio-Temporal Video Grounding
Pocket-STVG:用于时空视频定位的轻量级架构
Alberto Presta, Michal Byra, Grzegorz Stefański, Karol Szurkowski, Eryk Kołodziejczyk, Krzysztof Arendt
cs.CV · cs.AI · cs.MM
large language model
大语言模型相关
Abstract
Spatio-Temporal Video Grounding (STVG) aims to localize the spatio-temporal tube in a video corresponding to a natural language query. While recent methods achieve strong performance in fully supervised, weakly supervised, and zero-shot settings, they typically rely on computationally expensive architectures, complex training pipelines, or multimodal large language models. We present Pocket-STVG (P-STVG), a lightweight cascade architecture that addresses STVG by combining efficient pre-trained components instead of large end-to-end models. P-STVG integrates a temporal-aware video encoder based on MobileViCLIP, a spatial encoder-decoder derived from MDETR, and a shared aligned text encoder. Temporal localization is performed through either a lightweight 1D U-Net or a simple thresholding strategy, enabling the same framework to operate in both weakly supervised and zero-shot settings. Furthermore, video representations are precomputed independently of the query, yielding an indexing-friendly pipeline for efficient inference and large-scale video collections. Despite requiring fewer than 90M parameters, P-STVG performs on par with weakly supervised methods and improves on earlier zero-shot approaches at a fraction of their memory and computational cost, establishing a favorable performance-efficiency trade-off for STVG.
Chinese Translation
时空视频定位(STVG)旨在定位视频中与自然语言查询相对应的时空管。尽管最近的方法在全监督、弱监督和零样本设置下取得了强劲性能,但它们通常依赖于计算成本高昂的架构、复杂的训练流程或多模态大语言模型。我们提出 Pocket-STVG(P-STVG),这是一种轻量级级联架构,通过组合高效的预训练组件而非大型端到端模型来解决 STVG 问题。P-STVG 集成了基于 MobileViCLIP 的时序感知视频编码器、源自 MDETR 的空间编码器-解码器,以及共享对齐的文本编码器。时间定位通过轻量级 1D U-Net 或简单的阈值化策略来执行,使同一框架能够在弱监督和零样本设置下运行。此外,视频表示独立于查询进行预计算,从而产生一种有利于索引的流水线,以实现高效推理和大规模视频集合。尽管所需参数少于 90M,P-STVG 的性能与弱监督方法相当,并在仅需其一小部分内存和计算成本的情况下优于早期的零样本方法,为 STVG 建立了良好的性能-效率权衡。
cs.AI / 51 / 2609.31349
DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models
DyMD:在少步视频世界模型中通过分布匹配蒸馏保留交互动力学
Haojun Xu, Jie Huang, Xin Lu, Mingchen Zhong, Zihao Fan, Linjiang Huang, Si Liu
cs.CV · cs.AI
diffusion
扩散模型相关
Abstract
Large video diffusion models offer expressive priors for embodied prediction and learning, yet their many-step sampling remains costly for interactive downstream use. Distribution Matching Distillation (DMD) enables few-step video generation, but can suppress robot--object motion while preserving visual quality. Examining DMD's teacher and fake-score signals, we find that weak re-noising keeps the teacher posterior concentrated near motion-deficient rollouts, limiting motion-restoring guidance. Meanwhile, stronger-motion rollouts tend to incur larger fake-score fitting errors, which can hinder the generator's learning of interaction dynamics. We propose DyMD, a DMD framework that adapts both teacher supervision and critic fitting to the evolving student. Temporal affinity--conditioned re-noise sampling adapts the timestep distribution to each rollout's current interaction fidelity by mixing the base schedule with a teacher prior motivated by local posterior variation, thereby balancing motion recovery and appearance refinement. To better track stronger-motion rollouts, dynamics-guided fake-score tracking uses a noise-conditioned predictor to estimate noise-relative fitting difficulty from latent temporal dynamics, then upweights predicted-hard rollouts in the critic loss. Using DyMD, we distill a 14B teacher into a four-step 1.3B student with no auxiliary modules at inference. On embodied-video benchmarks, the student improves R-Bench task adherence by $9.6$ percentage points and PAI-Bench-G Domain score by $5.1$ points over Base DMD while maintaining comparable visual quality. As a backbone for downstream action planning, our student achieves 34% mean success across two WorldArena tasks, compared with 16% for Base DMD.
Chinese Translation
大型视频扩散模型为具身预测与学习提供了富有表现力的先验,但其多步采样对于交互式下游使用仍然代价高昂。分布匹配蒸馏(DMD)能够实现少步视频生成,但在保持视觉质量的同时可能抑制机器人--物体运动。通过考察 DMD 的教师信号与 fake-score 信号,我们发现弱重加噪会使教师后验集中在运动不足的 rollout 附近,从而限制运动恢复指导。与此同时,运动更强的 rollout 往往会产生更大的 fake-score 拟合误差,这可能阻碍生成器学习交互动态。我们提出 DyMD,一个使教师监督和 critic 拟合都适应不断演化的学生的 DMD 框架。时间亲和性--条件化的重加噪采样通过将基础调度与由局部后验变化驱动的教师先验混合,使时间步分布适应每个 rollout 当前的交互保真度,从而平衡运动恢复与外观细化。为了更好地跟踪运动更强的 rollout,动态引导的 fake-score 跟踪使用噪声条件化预测器,从潜在时间动态中估计噪声相对的拟合难度,然后在 critic 损失中提高预测为困难的 rollout 的权重。使用 DyMD,我们将一个 14B 教师蒸馏为一个四步 1.3B 学生模型,且在推理时无需辅助模块。在具身视频基准上,与 Base DMD 相比,该学生模型将 R-Bench 任务遵循度提高了 $9.6$ 个百分点,并将 PAI-Bench-G Domain 分数提高了 $5.1$ 分,同时保持可比的视觉质量。作为下游动作规划的骨干,我们的学生模型在两个 WorldArena 任务上取得了 34% 的平均成功率,而 Base DMD 为 16%。
cs.LG / 52 / 2609.31551
EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models
EAServe:面向多模态大语言模型的编码感知分离式服务
Kunxiong Zhu, Zhihao Shu, Hangyu Zheng, Minghai Qin, Miao Yin, Gagan Agrawal, Wei Niu
cs.DC · cs.LG · cs.PF
large language model
大语言模型相关
Abstract
Disaggregating the two stages, Prefill and Decode, onto separate GPU pools is now a standard optimization for (text-only) LLM serving. However, multimodal LLMs (MLLMs), which add a third phase, Encode, pose new challenges for resource allocation. Encode turns images, video, or audio into embeddings that the language model can consume, yielding a three-stage Encode-Prefill-Decode (EPD) pipeline. Existing frameworks offer only partial answers: text-only PD systems lack Encode, while EPD frameworks expose it as a separate service without regulating downstream request flow. The pipeline also carries a structural resource imbalance: every request enters through Encode before downstream work can begin, yet per-request execution leaves the encode GPU severely underutilized even at high loads, starving the downstream Prefill and Decode workers. Addressing this, we reposition Encode as the control point of the EPD pipeline, exposing three tightly coupled dimensions: when work enters downstream, where prefill executes, and how the GPU is shared. We instantiate this in EAServe across two co-designed layers. Its runtime manages load-adaptive micro-batching, rate-controlled partial offload to a co-resident prefill worker, and dynamic SM partitioning for predictable co-location. The configuration layer, Hybrid Auto Selection (HAS), navigates the joint space of GPU allocation, encode batch size, and offload ratio by pruning unbalanced allocations with per-stage capacity profiling and refining the remainder through TPE-based Bayesian optimization. Evaluated on three MLLM architectures spanning image, video, and audio, EAServe delivers up to 4.3x and 1.7x higher goodput than NVIDIA Dynamo and vLLM, respectively, under identical SLO constraints, sustains more balanced and higher GPU utilization across the EPD pipeline, and reaches near-optimal configurations faster than baseline search methods.
Chinese Translation
将预填充(Prefill)和解码(Decode)两个阶段分离到独立的 GPU 池上,现在是(纯文本)LLM 服务的一种标准优化。然而,多模态大语言模型(MLLMs)增加了第三个阶段——编码(Encode),给资源分配带来了新的挑战。编码将图像、视频或音频转换为语言模型可以消费的嵌入,从而产生三阶段编码-预填充-解码 (EPD) 流水线。现有框架只提供了部分答案:纯文本 PD 系统缺少编码阶段,而 EPD 框架将其作为单独的服务暴露,却没有调节下游请求流。该流水线还带有结构性的资源不平衡:每个请求都要先通过编码进入,然后下游工作才能开始,然而逐请求执行使得编码 GPU 即使在高负载下也严重利用不足,从而使下游的预填充和解码工作进程处于饥饿状态。针对这一点,我们将编码重新定位为 EPD 流水线的控制点,暴露出三个紧密耦合的维度:工作何时进入下游、预填充在哪里执行,以及 GPU 如何被共享。我们在 EAServe 中通过两个协同设计的层来实例化这一点。其运行时管理负载自适应的微批处理、向共驻预填充工作进程进行受速率控制的部分卸载,以及用于可预测共置的动态 SM 分区。配置层 Hybrid Auto Selection (HAS) 通过使用逐阶段容量分析剪除不平衡的分配,并通过基于 TPE 的贝叶斯优化细化剩余部分,从而在 GPU 分配、编码批大小和卸载比例的联合空间中导航。在涵盖图像、视频和音频的三种 MLLM 架构上评估,EAServe 在相同 SLO 约束下,相比 NVIDIA Dynamo 和 vLLM 分别提供最高 4.3 倍和 1.7 倍的更高有效吞吐量,在 EPD 流水线中维持更均衡且更高的 GPU 利用率,并且比基线搜索方法更快达到接近最优的配置。
cs.LG / 53 / 2609.30658
DiffusionShadow: Diffusion-based Shadow Caching for Neural Volume Rendering
DiffusionShadow:面向神经体渲染的基于扩散的阴影缓存
Kai-Chen Tung, Qi Wu, David Bauer, Mengjiao Han, Silvio Rizzi, Kwan-Liu Ma
cs.GR · cs.LG
diffusion
扩散模型相关
Abstract
Implicit neural representations (INRs) have gained momentum in scientific visualization due to their compactness and scalability to large datasets, making them well suited for integration with direct volume rendering (DVR). However, real-time volume rendering of INR with advanced illumination effects, such as shadows, remains computationally expensive, as evaluating shadow terms via ray marching is costly. Alternatively, precomputing and storing shadows for many lighting directions is prohibitive in both memory and storage. To address this, we introduce a diffusion-based shadow caching framework that compresses a vast set of pre-calculated shadow INRs into a single diffusion model. Rather than focusing on generalizing to unseen directions, our method effectively memorizes and reconstructs a dense set of pre-trained lighting conditions on the fly. We first encode a collection of shadow coefficient volumes as shadow INRs, and then train a diffusion model conditioned on lighting direction to predict the corresponding shadow INR weights at inference time. This design integrates directly with standard INR renderers without additional runtime sampling. Experiments show that our approach achieves faster rendering than traditional methods while bypassing the massive storage bloat of independent INRs, producing shadows that closely match most of the reference results.
Chinese Translation
隐式神经表示(INRs)因其紧凑性和对大型数据集的可扩展性,在科学可视化中获得了发展势头,使其非常适合与直接体渲染(DVR)集成。然而,对具有诸如阴影等高级光照效果的 INR 进行实时体渲染,在计算上仍然代价高昂,因为通过光线行进评估阴影项是昂贵的。另一种选择是,为许多光照方向预计算并存储阴影,在内存和存储两方面都令人望而却步。为了解决这个问题,我们引入了一个基于扩散的阴影缓存框架,它将大量预先计算的阴影 INR 压缩到一个单一的扩散模型中。我们的方法并不侧重于泛化到未见过的方向,而是有效地即时记忆并重建一组密集的预训练光照条件。我们首先将一组阴影系数体积编码为阴影 INR,然后训练一个以光照方向为条件的扩散模型,以在推理时预测相应的阴影 INR 权重。这种设计可直接与标准 INR 渲染器集成,而无需额外的运行时采样。实验表明,我们的方法比传统方法实现了更快的渲染,同时绕开了独立 INR 的大规模存储膨胀,生成的阴影与大多数参考结果密切匹配。
cs.LG / 54 / 2609.30405
Adaptive Multi-Value Control in LLMs via Causal Activation Steering
通过因果激活引导实现大语言模型中的自适应多值控制
Payel Bhattacharjee, Ravi Tandon
cs.LG
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly deployed in settings where responses must reflect multiple, potentially interacting social norms and human values. Activation steering offers a lightweight alternative to training-based alignment by modifying internal activations at inference time. However, prior human-value steering methods have largely considered values in isolation, while direct composition of multiple directions relies on fixed intervention strengths that cannot respond to the model's evolving internal state. Motivated by this key observation, we introduce AIMES, a framework for adaptive multi-value activation steering. AIMES constructs layer-specific bipolar directions for moral-foundation values and uses intermediate-layer vocabulary readouts as online observers. An observer-guided controller then adapts the strength of each requested value intervention at every decoding step based on its current observed state, without training a separate value-state estimator. Across multiple instruction-tuned model families, value combinations, and intervention depths, we find that multi-value controllability varies across both value combinations and intervention locations. Compared with fixed joint steering and prompt-based steering, AIMES shows depth-dependent advantages that are broadly supported across two independent evaluators, with some variation in the precise depth at which specific control effects emerge. These advantages come with smaller realized activation-space interventions than fixed-joint steering and comparable response quality. Overall, our results suggest that online observer feedback can provide lightweight, state-aware adaptation for single-pass multi-value steering.
Chinese Translation
大语言模型(LLMs)正越来越多地被部署在那些要求其回复必须反映多种、且可能相互作用的社會规范与人类价值观的场景中。激活引导通过在推理时修改内部激活,为基于训练的对齐提供了一种轻量级的替代方案。然而,以往的人类价值观引导方法大多孤立地考虑各个价值,而多个方向的直接组合依赖于固定的干预强度,这些固定强度无法对模型不断演变的内部状态做出响应。受这一关键观察的启发,我们提出了 AIMES,一个用于自适应多值激活引导的框架。AIMES 为道德基础价值观构建了层级特定的双极方向,并使用中间层词表读出作为在线观测器。随后,一个由观测器引导的控制器在每个解码步骤中,根据其当前观测到的状态,自适应地调整每一项所请求的价值干预的强度,而无需训练单独的价值状态估计器。在多个指令微调模型系列、价值组合和干预深度上,我们发现多值可控性会随价值组合和干预位置的不同而变化。与固定联合引导和基于提示的引导相比,AIMES 表现出依赖于深度的优势,这些优势在两个独立评估器上得到广泛支持,但在特定控制效应出现的具体深度上存在一定差异。与固定联合引导相比,这些优势伴随着实际实现的激活空间干预更小,且响应质量相当。总体而言,我们的结果表明,在线观测器反馈能够为单遍多值引导提供轻量级、状态感知的自适应能力。
cs.LG / 55 / 2609.30427
Fake News Theories: Harnessing Disciplinary Insights for Computational Modeling, Detection, and Explanation
假新闻理论:利用学科洞见进行计算建模、检测与解释
Zhaoyang Cao, Miriam Metzger, Reza Zafarani
cs.LG · cs.CY
large language model
大语言模型相关
Abstract
Disinformation research has produced increasingly accurate automated fake-news detectors, but many systems remain difficult to interpret and are weakly connected to established theories of persuasion, credibility, and human judgment. In this paper, we develop a theory-informed computational framework that translates cross-disciplinary theories of fake news into measurable features for automated detection and explanation through statistical techniques and large language models. To that end, we conduct a structured cross-disciplinary review of theories from social sciences, psychology, economics, among other disciplines that reveal how fake news persuades and spreads, thereby establishing a broad theoretical foundation for computational modeling. Experiments on benchmark datasets show that theory-derived features are predictive and provide interpretable, theory-referenced diagnostic signals. Multi-feature models generally outperform individual features, although gains among the strongest small feature combinations are modest. Our work highlights the value of interdisciplinary perspectives in building robust and interpretable fake news detection systems, advancing the foundation for human-centered approaches in combating disinformation.
Chinese Translation
虚假信息研究已经产生了越来越准确的自动化假新闻检测器,但许多系统仍然难以解释,并且与既有的说服、可信度和人类判断理论联系薄弱。在本文中,我们开发了一个以理论为依据的计算框架,该框架通过统计技术和大型语言模型,将跨学科的假新闻理论转化为可用于自动化检测与解释的可测量特征。为此,我们对来自社会科学、心理学、经济学以及其他学科的理论进行了结构化的跨学科综述,这些理论揭示了假新闻如何说服和传播,从而为计算建模奠定了广泛的理论基础。在基准数据集上的实验表明,由理论推导出的特征具有预测能力,并能提供可解释的、以理论为参照的诊断信号。多特征模型总体上优于单个特征,尽管在最强的小型特征组合之间的增益较为有限。我们的工作凸显了跨学科视角在构建稳健且可解释的假新闻检测系统中的价值,为应对虚假信息时以人为中心的方法推进了基础。
cs.LG / 56 / 2609.30501
To Solve Bilevel Optimization with Nonconvex Lower Levels, We Need Second-Order Stationarity
要解决具有非凸下层的双层优化,我们需要二阶平稳性
Zhiyao Zhang, Menglu Yu, Alvaro Velasquez, Nathaniel D. Bastian, Jia Liu
cs.LG
large language model
大语言模型相关
Abstract
Although bilevel optimization (BLO) has emerged as a powerful framework for addressing many complex and nested machine learning problems in recent years, most existing studies are confined to the lower-level strongly convex (LLSC) or lower-level generally convex (LLGC) settings (i.e., the lower-level objective function is assumed to be, at least, convex). While the LLSC/LLGC assumptions render more tractable algorithmic design and theoretical analysis, they are too rigid to encompass many machine learning problems in practice. The limitations of LLSC/LLGC assumptions in BLO motivate us to investigate solving the BLO problem in the general lower-level nonconvex (LLNC) settings, which remains in its infancy. In the literature on LLNC-BLO, most of the existing works either require additional structures in the lower-level objective function for tractable theoretical analysis, or adopt the first-order stationarity reformulation as a lower-level surrogate problem, which is inherited from the LLSC/LLGC settings but could lose their effectiveness in the LLNC setting. To bridge this gap, we propose to reformulate the nonconvex lower-level problem using a second-order stationarity-based surrogate, the solution of which guarantees a local optimal solution at the lower level. Based on this reformulation, we propose the PROBE (Perturbed gradient algorithm for bilevel problem) and show that it overcomes the limitations of prior works by probing and escaping lower-level saddle points. We prove that PROBE achieves a finite-time convergence rate of $O(T^{-2/5})$, where T denotes iterations. To our knowledge, this work is the first to establish the finite-time convergence for achieving lower-level second-order stationary solutions in general LLNC-BLO. Our experiments on both a large language model-based data curation task and a meta-learning task also show that PROBE outperforms SOTA methods.
Chinese Translation
尽管近年来双层优化(BLO)已成为一种强大的框架,用于处理许多复杂且嵌套的机器学习问题,但大多数现有研究都局限于下层强凸(LLSC)或下层一般凸(LLGC)设置(即,下层目标函数被假定为至少是凸的)。虽然 LLSC/LLGC 假设使算法设计和理论分析更易于处理,但它们过于严格,无法涵盖实践中的许多机器学习问题。BLO 中 LLSC/LLGC 假设的局限性促使我们研究在一般的下层非凸(LLNC)设置中求解 BLO 问题,而该方向仍处于起步阶段。在关于 LLNC-BLO 的文献中,大多数现有工作要么需要下层目标函数中的额外结构以进行可处理的理论分析,要么采用一阶平稳性重构作为下层代理问题,这种重构继承自 LLSC/LLGC 设置,但在 LLNC 设置中可能失去其有效性。为了弥补这一差距,我们提出使用基于二阶平稳性的代理来重构非凸下层问题,其解保证了下层的一个局部最优解。基于这一重构,我们提出了 PROBE(双层问题的扰动梯度算法),并表明它通过探测和逃离下层鞍点克服了先前工作的局限性。我们证明 PROBE 达到了 $O(T^{-2/5})$ 的有限时间收敛速率,其中 T 表示迭代次数。据我们所知,这项工作是首个在一般 LLNC-BLO 中建立达到下层二阶平稳解的有限时间收敛性的工作。我们在基于大语言模型的数据整理任务和元学习任务上的实验也表明,PROBE 优于 SOTA 方法。
cs.LG / 57 / 2609.30541
AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework
生产规模下的AutoResearch:失效模式与一个多智能体框架
Aparajith Chandran, Juwon Kim, Saurav Jha, Pablo Castells, Florian Hottier
cs.LG · cs.IR
large language model
大语言模型相关
Abstract
Optimizing embedding systems for production recommendation pipelines demands systematic exploration that consumes disproportionate engineering effort at scale. We apply Andrej Karpathy's AutoResearch paradigm -- a large language model that iteratively edits a training script and retains modifications that improve a held-out scalar metric -- to automate this exploration. We report on twelve weeks of running this paradigm at production scale, where iterations consume hours of multi-GPU compute, evaluation involves competing criteria, and campaigns span weeks across many training jobs. Across two independently developed representation-learning systems for a book recommendation pipeline, we ran 220+ experiments and observed five recurring failure modes absent from the original setting: infrastructure fragility, agent memory decay, search-direction stagnation, iteration-cost asymmetry, and metric fixation. We contribute a three-principle scaffolding design -- prevent, persist, redirect -- that maps each failure mode to a structural remedy and whose instantiation scales with iteration cost. The framework produced a 1.82x Recall@6 lift and a 2.1x coherence lift over hand-tuned baselines, and the agent autonomously designed a text-only fallback that expanded catalog coverage by 5.8x. The two systems span nearly three orders of magnitude in per-iteration cost yet exhibit the same failure modes, suggesting these are structural properties of production-scale autonomous research rather than artifacts of either application.
Chinese Translation
为生产推荐流水线优化嵌入系统需要系统性的探索,而这种探索在大规模场景下会消耗不成比例的工程投入。我们应用Andrej Karpathy的AutoResearch范式——即用一个大型语言模型迭代地修改训练脚本,并保留那些能够改善某个留出标量指标的修改——来自动化这一探索。我们报告了在生产规模下运行该范式十二周的情况,其中每次迭代都要消耗数小时的多GPU计算,评估涉及相互竞争的标准,而实验周期会跨越多周、覆盖众多训练任务。在针对一个图书推荐流水线独立开发的两个表示学习系统上,我们运行了220多个实验,并观察到五种在原初设定中并不存在的反复出现的失效模式:基础设施脆弱性、智能体记忆衰退、搜索方向停滞、迭代成本不对称,以及指标固着。我们贡献了一个三原则的脚手架设计——预防、持久化、重定向——它将每种失效模式映射到一种结构性补救措施,且其具体实现会随迭代成本而扩展。该框架相较于人工调优的基线带来了1.82倍的Recall@6提升和2.1倍的一致性提升,并且智能体自主设计了一个纯文本回退方案,将目录覆盖率扩大了5.8倍。这两个系统在每次迭代成本上跨越了近三个数量级,却表现出相同的失效模式,这表明它们是生产规模自主研究的结构性属性,而非任一应用的偶然产物。
cs.LG / 58 / 2609.30572
Entropy Regularization: A Free Correction to Cross-Entropy for Verified Demonstrations
熵正则化:面向已验证示教的交叉熵的一种免费修正
Mihir Dhanakshirur, Adam Ousherovitch, Ambuj Tewari
cs.LG · stat.ML
large language model
大语言模型相关
Abstract
Large language models are often post-trained on expert demonstrations using cross-entropy (CE), even when the downstream objective is not to imitate the demonstrated solution but to produce any output accepted by a verifier. This mismatch is seen in verifiable domains with multiple correct solutions, such as mathematical reasoning and code generation, where training data may contain only one expert solution per problem. We show that minimizing cross-entropy can be misaligned with minimizing verifier risk; two policies can assign identical likelihood to the observed demonstrations while placing different probability mass on incorrect outputs. This is formalized through a learning-theoretic counterexample in which CE minimization selects a suboptimal policy. We identify that controlling the support of the learned policy can solve this problem by preventing probability mass from spreading to unsupported outputs. Since support size is non-differentiable and computationally intractable, we propose entropy-regularized cross-entropy (ER-CE), using token-level Shannon entropy as a tractable proxy. Finally, across mathematical reasoning and code-generation benchmarks, we find that entropy-regularized training consistently improves verifier accuracy over standard cross-entropy. Our results identify a simple failure mode of imitation-based post-training in verifiable tasks and provide a practical objective that is better aligned with producing correct outputs.
Chinese Translation
大型语言模型常常使用交叉熵(CE)在专家示教上进行后训练,即使下游目标并不是模仿所展示的解答,而是产生任何能被验证器接受的输出。这种不匹配出现在具有多个正确解的可验证领域中,例如数学推理和代码生成,其中训练数据可能对每个问题只包含一个专家解答。我们表明,最小化交叉熵可能与最小化验证器风险不一致;两个策略可以对观察到的示教赋予相同的似然,却在错误输出上分配不同的概率质量。这一点通过一个学习理论中的反例得以形式化,在该反例中,交叉熵最小化选择了一个次优策略。我们指出,控制所学策略的支撑集可以通过防止概率质量扩散到无支撑的输出上来解决这一问题。由于支撑集大小不可微且计算上难以处理,我们提出熵正则化交叉熵(ER-CE),使用词元级香农熵作为可处理的代理。最后,在数学推理和代码生成基准上,我们发现熵正则化训练相比标准交叉熵持续提升验证器准确率。我们的结果识别出可验证任务中基于模仿的后训练的一个简单失效模式,并提供了一个更契合产生正确输出的实用目标。
cs.LG / 59 / 2609.30692
LUMO (Lightweight Unified Multilingual Orchestrator): A Privacy Preserving Offline Voice Assistant
LUMO(轻量级统一多语言编排器):一种隐私保护的离线语音助手
Md. Mehedi Hasan Naeem, Mst. Kamrunnahar Ruma, Nafiza Anjum, Shakila Sultana, Md. Sujan Ali
cs.LG
large language model
大语言模型相关
Abstract
Reliable voice interaction is essential in environments with limited internet connectivity and strong privacy. However, most existing voice assistants depend on cloud-based services, which leads to latency issues, dependency on internet access, and privacy vulnerabilities. This research presents LUMO (Lightweight Unified Multilingual Orchestrator), a privacy preserving offline voice assistant designed for edge computing environments. This system integrates local Automatic Speech Recognition (ASR), locally deployed quantized Large Language Model (LLM), and Text-to-Speech (TTS) synthesis into a fully offline pipeline running on a Raspberry Pi 5 with 8 GB RAM. To enable efficient operation on resource constrained hardware, the language model is compressed using 4-bit GGUF quantization, which reduces memory usage while preserving practical conversational capability. Existing edge based voice assistants Mycroft provides partial offline functionality without a generative LLM, with an approximate latency of ~5 s and power consumption of ~12 W, while Rhasspy supports full offline operation but lacks generative capabilities, with ~3 s latency and ~11 W power usage. In contrast, LUMO achieves a Word Error Rate (WER) of 6.8% for short English utterances in low noise conditions, an end-to-end response latency of 2.0-4.0 s, and a lower peak power consumption of approximately 9.0 W. The system also achieves effective offline recognition for Bangla speech, supporting multilingual accessibility in low resource settings. By operating entirely offline, LUMO provides strong data privacy, reduced need for cloud connectivity, and suitability for privacy sensitive edge execution such as rural healthcare, education, and disaster response scenarios.
Chinese Translation
在互联网连接有限且隐私要求高的环境中,可靠的语音交互至关重要。然而,大多数现有语音助手依赖基于云的服务,这会导致延迟问题、对互联网接入的依赖以及隐私漏洞。本研究提出了 LUMO(轻量级统一多语言编排器),一种为边缘计算环境设计的隐私保护型离线语音助手。该系统将本地自动语音识别(ASR)、本地部署的量化大语言模型(LLM)以及文本到语音(TTS)合成集成到一个完全离线的流水线中,该流水线运行在具有 8 GB RAM 的 Raspberry Pi 5 上。为了在资源受限的硬件上实现高效运行,语言模型使用 4-bit GGUF 量化进行压缩,这在减少内存使用的同时保持了实用的对话能力。现有的基于边缘的语音助手 Mycroft 在没有生成式 LLM 的情况下提供部分离线功能,其近似延迟约为 ~5 s,功耗约为 ~12 W,而 Rhasspy 支持完全离线运行,但缺乏生成能力,其延迟约为 ~3 s,功耗约为 ~11 W。相比之下,LUMO 在低噪声条件下对短英文话语实现了 6.8% 的词错误率(WER),端到端响应延迟为 2.0-4.0 s,并且具有更低的峰值功耗,约为 9.0 W。该系统还实现了对孟加拉语语音的有效离线识别,在低资源环境下支持多语言可访问性。通过完全离线运行,LUMO 提供了强大的数据隐私、降低了对云连接的需求,并且适用于隐私敏感的边端执行场景,如农村医疗、教育和灾害响应场景。
cs.LG / 60 / 2609.30884
CacheReforge: Bounded Recovery for Stale KV Caches under Evolving Adapters
CacheReforge:演化适配器下陈旧 KV 缓存的有界恢复
Yuhang Cao, Yanzhou Mu, Chunrong Fang, Zhenyu Chen
cs.LG
large language model
大语言模型相关
Abstract
Large language models rely on KV caching to reduce repeated prefill computation in long context and interactive applications. As lightweight adapters evolve, cached states reflect earlier versions, so stale reuse distorts current model outputs, while complete affected suffix recomputation restores fidelity at substantial cost. We seek minimal recomputation that recovers current adapter behavior. Existing systems track token, context, or stable adapter identity, but neither represent caches from earlier adapter versions nor distinguish update propagation from the recomputation required for behavioral recovery. To address these gaps, we introduce CacheReforge, which represents stale KV caches as layerwise mixed-version objects. It combines per-layer adapter anchors, calibrated sensitivity, accumulated drift, and executable restart boundaries to select direct reuse, bounded recomputation, or complete affected-suffix recovery. We distinguish dependency depth from the functional recomputation horizon and use cumulative tail influence to characterize when bounded recovery preserves current-model behavior. We evaluate CacheReforge on Qwen2.5-1.5B and Qwen2.5-7B with continual LoRA updates, including 16K HotpotQA and 2WikiMQA workloads. CacheReforge reduces mean KL divergence by 92.4% relative to stale reuse, while recomputing only 5.44% of layers and reducing cache-maintenance time by 93.2% relative to fresh full prefill. These results show that version-aware recovery preserves model fidelity and most KV caching gains.
Chinese Translation
大型语言模型依赖 KV 缓存,以减少长上下文和交互式应用中的重复预填充计算。随着轻量级适配器演化,缓存状态反映的是较早版本,因此陈旧复用会扭曲当前模型输出,而完整重计算受影响后缀能以巨大代价恢复保真度。我们寻求能够恢复当前适配器行为的最小重计算。现有系统跟踪 token、上下文或稳定适配器身份,但既不能表示来自较早适配器版本的缓存,也不能区分更新传播与行为恢复所需的重计算。为弥补这些空白,我们引入 CacheReforge,它将陈旧 KV 缓存表示为逐层混合版本对象。它结合逐层适配器锚点、校准后的敏感度、累积漂移以及可执行的重启边界,以选择直接复用、有界重计算或完整受影响后缀恢复。我们区分依赖深度与功能性重计算视界,并使用累积尾部影响来刻画有界恢复何时保持当前模型行为。我们在 Qwen2.5-1.5B 和 Qwen2.5-7B 上评估 CacheReforge,采用持续 LoRA 更新,包括 16K HotpotQA 和 2WikiMQA 工作负载。CacheReforge 相对于陈旧复用将平均 KL 散度降低了 92.4%,同时仅重计算 5.44% 的层,并相对于全新完整预填充将缓存维护时间减少了 93.2%。这些结果表明,版本感知恢复保持了模型保真度和大部分 KV 缓存收益。
cs.LG / 61 / 2609.30957
Towards Understanding Momentum Acceleration in River-Valley Loss Landscape
迈向理解河-谷损失景观中的动量加速
Miao Lu, Zeyu Bian, Kaiyue Wen, Beining Wu, Siyu Chen, Tianhao Wang, Zhiyuan Li
cs.LG · math.OC
large language model
大语言模型相关
Abstract
The empirical success of pretraining large language models has inspired a deeper investigation into the underlying loss landscapes and the optimization dynamics. Recent empirical and theoretical study suggest that the training loss landscape often exhibits a "river-valley" structure, which features a low-loss manifold (river) flanked by sharp orthogonal directions with higher loss (mountains). In the long term, the optimization progress is determined primarily by the progress along the river. Within such a landscape, gradient descent with large learning rates can move faster along the river despite high apparent loss due to vertical oscillations, while a subsequent sharp decay in the learning rate suppresses these oscillations, revealing genuine optimization progress. This explains the recent success of warmup-stable-decay (WSD) learning rate scheduler which, unlike cosine scheduling, keeps stable high learning rate and decays before producing intermediate checkpoints. Building on this foundation, in this work we take a step further and study the role of momentum within such a loss landscape. We establish theoretical analysis that characterizes how momentum accelerates optimization by stabilizing large learning rates that can not be tolerated by vanilla GD without deviating significantly from the river. The enabled large learning rate in-turn gives greater speed along the river and makes faster essential progress in the long run. Another intriguing observation from theory is that for a river-valley landscape with very flat and slow-spinning river, the momentum itself does not contribute directly to acceleration in terms of the speed of tracking the river, while the main acceleration comes from the admissible larger learning rate.
Chinese Translation
预训练大语言模型的实证成功激发了对底层损失景观和优化动力学的更深入研究。最近的实证和理论研究表明,训练损失景观通常呈现出一种“河-谷”结构,其特征是一个低损失流形(河流),两侧是具有更高损失的尖锐正交方向(山脉)。从长期来看,优化进展主要取决于沿河流方向的进展。在这样的景观中,使用大学习率的梯度下降可以沿着河流移动得更快,尽管由于垂直振荡而导致表观损失较高,而随后学习率的急剧衰减会抑制这些振荡,从而揭示出真正的优化进展。这解释了最近预热-稳定-衰减(WSD)学习率调度器的成功,与余弦调度不同,它保持稳定的高学习率,并在产生中间检查点之前衰减。在此基础上,在本工作中我们更进一步,研究动量在这种损失景观中的作用。我们建立了理论分析,刻画动量如何通过稳定那些普通 GD 无法在不显著偏离河流的情况下容忍的大学习率来加速优化。由此得以启用的大学习率反过来提供了沿河流方向更大的速度,并在长期内取得更快的实质性进展。理论中的另一个有趣观察是,对于河流非常平坦且缓慢旋转的河-谷景观,动量本身在跟踪河流的速度方面并不直接贡献加速,而主要加速来自可容许的更大学习率。
cs.LG / 62 / 2609.30977
Does Uniform Discrete Diffusion Need Time?
均匀离散扩散需要时间吗?
Chunsan Hong, Chieh-Hsin Lai, Satoshi Hayakawa, Yuhta Takida, Jong Chul Ye, Yuki Mitsufuji
cs.LG · cs.AI · cs.CL
diffusion
扩散模型相关
Abstract
Uniform discrete diffusion models (UDMs) commonly use explicit time conditioning, but we find that it can often be unnecessary in practice. In this paper, we first show that the population-optimal UDM predictor generally depends on time: time controls how much the model should trust the observed context. We then show that this dependence can become negligible in finite-data settings relevant to language. When a corrupted training sequence remains much closer to its original clean sequence than to competing training sequences, the empirical-optimal predictor is nearly insensitive to time over most of the diffusion trajectory, where the guarantee weakens toward the high-noise endpoint. Empirically, trained language UDMs exhibit limited time sensitivity over most of the trajectory, while time-agnostic predictors remain competitive with, and often outperform, time-conditioned models across datasets and training objectives. These results challenge the use of explicit time conditioning in UDMs: although the population optimum depends on time, explicitly conditioning on it may often be unnecessary in practice.
Chinese Translation
均匀离散扩散模型(UDMs)通常使用显式时间条件化,但我们发现这在实践中往往可能是不必要的。在本文中,我们首先表明,总体最优的 UDM 预测器通常依赖于时间:时间控制着模型应当在多大程度上信任观测到的上下文。我们随后表明,在与语言相关的有限数据设定中,这种依赖可能变得可以忽略不计。当一个受损的训练序列与其原始干净序列的距离仍远小于其与竞争训练序列的距离时,经验最优预测器在扩散轨迹的大部分区域上几乎对时间不敏感,而该保证在高噪声端点附近会减弱。经验上,训练后的语言 UDM 在轨迹的大部分区域上表现出有限的时间敏感性,而与时间无关的预测器在不同数据集和训练目标下仍能与时间条件模型竞争,且往往优于后者。这些结果对在 UDM 中使用显式时间条件化提出了挑战:尽管总体最优依赖于时间,但在实践中显式地以其为条件往往可能是不必要的。
cs.LG / 63 / 2609.30996
The Linear Representation Hypothesis for Vision-Language-Action Models
视觉-语言-动作模型的线性表示假设
Minseok Jeong, Hyewon Choi, Hiroyasu Tsukamoto, SooJean Han
cs.LG · cs.AI
large language model
大语言模型相关
Abstract
The linear representation hypothesis (LRH) has become a standard lens for measuring and intervening on semantic information through the internal representations of large language models (LLMs). A growing body of work has begun extending this perspective to vision-language-action (VLA) models, but the dynamical nature of embodied interaction introduces an additional challenge. Unlike semantic attributes commonly studied in LLMs, such as gender or language, a physical quantity of interest (QoI) in a VLA evolves jointly with the system dynamics: the representation influences the actions selected by the policy, which alter the physical state and, in turn, the next representation. In this paper, we develop a theoretical, signature-based formulation of the LRH for VLA that unifies representations and policies. On the representation side, we establish the existence of representations from which the future evolution of a QoI under a candidate action trajectory can be recovered via linear probing. On the policy side, we introduce a signature generalized linear model for stochastic action chunks. This structure yields a monotonic change in the expected future QoI along linear paths in natural parameter space, enabling linear steering. We construct an explicit oracle representation in a planar control-affine navigation experiment and verify the predicted linear probing and steering mechanisms.
Chinese Translation
线性表示假设(LRH)已成为一种标准视角,用于通过大语言模型(LLMs)的内部表示来测量和干预语义信息。越来越多的研究已开始将这一视角扩展到视觉-语言-动作(VLA)模型,但具身交互的动态性质引入了额外的挑战。与LLMs中常被研究的语义属性(如性别或语言)不同,VLA中关注的物理量(QoI)与系统动力学共同演化:表示影响策略所选择的动作,动作改变物理状态,进而改变下一个表示。在本文中,我们为VLA提出了一种理论性的、基于签名的LRH表述,它统一了表示与策略。在表示一侧,我们确立了这样一类表示的存在性:从中可以通过线性探测恢复出在候选动作轨迹下QoI的未来演化。在策略一侧,我们为随机动作块引入了一种签名广义线性模型。这种结构使得期望的未来QoI在自然参数空间中沿线性路径发生单调变化,从而实现线性引导。我们在一个平面控制仿射导航实验中构建了一个显式的oracle表示,并验证了所预测的线性探测和引导机制。
cs.LG / 64 / 2609.31149
CRNDiff: Count-Native Diffusion Framework via Chemical Reaction Networks
CRNDiff:基于化学反应网络的计数原生扩散框架
Yuxuan Qiu, Praful Gagrani, Tetsuya J Kobayashi
cs.LG
diffusion
扩散模型相关
Abstract
Scientific measurements such as single-cell RNA (scRNA) sequencing often take the form of nonnegative integer counts, whereas continuous-state diffusion models approximate this discrete structure using continuous coordinates. Building on stochastic chemical reaction networks (CRNs), a class of count-native Markov jump processes, we introduce CRNDiff, a structured framework that combines count-space diffusion with inference-time conditioning on rare subpopulations. An independent birth--death instantiation yields a closed-form transition kernel for forward noising. This kernel enables reverse sampling via forward-filtering backward-sampling (FFBS) and supports data-driven selection of the terminal noising time, eliminating the need for a validation sweep. This tractability also lets us introduce tilted Feynman--Kac (FK) steering, a method for sampling target subpopulations from a frozen generator without retraining. By tilting posterior marginals before FK particle correction, steering mitigates importance-weight concentration when the target population is rare. Using scRNA-seq data from the human heart cell atlas, we test the ability of CRNDiff to generate cell-type-specific distributions. Across the three evaluated target populations, CRNDiff achieves the highest conditional fidelity among the evaluated generative models, with larger mean purity margins for rarer target populations. Generated cells preserve marker-level differential-expression structure. Replacing real training cells for the target classes with generated cells yields downstream classification performance approaching that of the real-data reference.
Chinese Translation
诸如单细胞 RNA(scRNA)测序之类的科学测量通常以非负整数计数的形式出现,而连续状态扩散模型则使用连续坐标来近似这种离散结构。基于随机化学反应网络(CRNs)——一类计数原生马尔可夫跳跃过程——我们提出 CRNDiff,一个结构化框架,它将计数空间扩散与推理时对稀有亚群的条件化相结合。一个独立的生灭(birth--death)实例化产生了用于前向加噪的闭式转移核。该核通过前向滤波-后向采样(FFBS)实现反向采样,并支持以数据驱动方式选择终端加噪时间,从而无需验证扫描。这种易处理性还使我们能够引入倾斜 Feynman--Kac(FK)引导,一种无需重新训练即可从冻结生成器中采样目标亚群的方法。通过在 FK 粒子校正之前倾斜后验边缘分布,引导缓解了目标群体稀有时的的重要性权重集中问题。使用来自人类心脏细胞图谱的 scRNA-seq 数据,我们测试了 CRNDiff 生成细胞类型特异性分布的能力。在三个被评估的目标群体中,CRNDiff 在所评估的生成模型中实现了最高的条件保真度,并且对于更稀有的目标群体具有更大的平均纯度裕度。生成的细胞保留了标志物水平的差异表达结构。用生成细胞替换目标类别的真实训练细胞,可获得接近真实数据参考的下游分类性能。
cs.LG / 65 / 2609.31222
Budgeted Quotient-Residual Guidance for Frozen Pocket-Conditioned Molecular Diffusion
用于冻结口袋条件分子扩散的预算化商-残差引导
Xinyu Wang, Jinbo Bi, Minghu Song
cs.LG
diffusion
扩散模型相关
Abstract
Pocket-conditioned molecular diffusion updates ambient atom coordinates, but many lead-optimization objectives are expressed on quotient features such as distances, contacts, and anchored substructures. We introduce budgeted quotient-residual guidance (QRG), an inference-time correction that makes these quotient objectives active without retraining the molecular generator. QRG lifts quotient covectors to metric-horizontal ambient directions and delivers them through a trust budget set by the frozen sampler's own step norm: quotient geometry chooses the direction, while sampler motion bounds the scale. We derive the horizontal lift, closed-form sampler-budget update, KL/kinetic interpretation around a frozen reverse step, equivariance conditions, and a product-budget split for budget-capped section and residual controls. Controlled quotient tasks confirm that sampler-relative delivery activates signals that raw local quotient gradients leave dormant. On frozen TargetDiff backbones, official seed-0 CBGBench ligand-generation/editing sweeps show practical quality-runtime gains: Local-QRG improves validity from 0.815 to 0.864 on fragment growing, 0.664 to 0.707 on scaffold hopping, and 0.681 to 0.712 on linker design, while PredNext-QRG improves fragment/scaffold and remains near-neutral on linker. Novelty remains 1.000 and diversity is preserved in the matched multi-seed molecular slice, giving task-dependent improvements without sampler retraining or backbone modification. Overall, QRG provides a lightweight route to quotient-aware inference for frozen molecular samplers with explicit runtime accounting.
Chinese Translation
口袋条件分子扩散更新环境原子坐标,但许多先导优化目标表达在商特征上,例如距离、接触和锚定子结构。我们提出预算化商-残差引导(QRG),这是一种推理时校正,使这些商目标变为活跃,而无需重新训练分子生成器。QRG 将商余向量提升为度量水平的环境方向,并通过由冻结采样器自身步长范数设定的信任预算传递它们:商几何选择方向,而采样器运动约束尺度。我们推导水平提升、闭式采样器预算更新、围绕冻结反向步的 KL/动力学解释、等变性条件,以及用于受预算上限约束的截面控制和残差控制的乘积预算拆分。受控商任务证实,相对于采样器的传递激活了原始局部商梯度所保持休眠的信号。在冻结的 TargetDiff 骨干上,官方 seed-0 CBGBench 配体生成/编辑扫描显示出实际的质量-运行时间增益:Local-QRG 在片段生长上将有效性从 0.815 提高到 0.864,在骨架跃迁上从 0.664 提高到 0.707,在连接子设计上从 0.681 提高到 0.712,而 PredNext-QRG 改善了片段/骨架并在连接子上保持接近中性。新颖性保持为 1.000,并且在匹配的多随机种子分子切片中多样性得以保持,从而在不重新训练采样器或修改骨干的情况下给出依赖于任务的改进。总体而言,QRG 为冻结分子采样器提供了一条轻量级的商感知推理途径,并具有显式的运行时间核算。
cs.LG / 66 / 2609.31371
Towards Understanding LLM-Based Log Anomaly Detection: An Empirical Study of Performance, Efficiency, and Robustness
面向理解基于LLM的日志异常检测:性能、效率与鲁棒性的实证研究
Bin Li, Dongdong Wang, Siyang Lu
cs.LG · cs.CR
large language model
大语言模型相关
Abstract
Large language models (LLMs) have demonstrated promising performance in log anomaly detection, yet how their adaptation strategies, architectures, and deployment configurations affect detection effectiveness remains insufficiently understood. To investigate these factors, we conduct a systematic empirical analysis across three public log datasets, examining different adaptation strategies, model architectures, parameter scales, and quantization settings. Our results reveal substantial performance differences across adaptation strategies, while model scaling yields varying detection gains across datasets. We further observe that models with comparable detection accuracy can exhibit markedly different computational costs, and that low-bit quantization largely preserves detection performance in the evaluated configurations. Finally, we examine detection robustness under structural, semantic, and label noise at different perturbation levels. These findings provide empirical insights into the performance, efficiency, and robustness of LLM-based log anomaly detection, highlighting practical considerations beyond conventional accuracy-oriented evaluation.
Chinese Translation
大语言模型(LLM)在日志异常检测中已展现出有前景的性能,但它们的适配策略、架构和部署配置如何影响检测有效性,仍未被充分理解。为探究这些因素,我们在三个公开日志数据集上开展了一项系统的实证分析,考察了不同的适配策略、模型架构、参数规模和量化设置。我们的结果揭示了不同适配策略之间存在显著的性能差异,而模型规模扩展在不同数据集上带来的检测增益各不相同。我们进一步观察到,具有相当检测精度的模型可能表现出显著不同的计算开销,并且在所评估的配置中,低位量化在很大程度上保持了检测性能。最后,我们考察了在不同扰动水平下的结构噪声、语义噪声和标签噪声条件下的检测鲁棒性。这些发现为基于LLM的日志异常检测的性能、效率和鲁棒性提供了实证见解,并突出了超越传统以精度为导向的评估的实际考量。
cs.LG / 67 / 2609.31600
New LoRA Skills Should Read but Never Write
新的 LoRA 技能应当只读,绝不写入
Zeyan Li, Panqi Yang, Qirong Guo, Shengda Zhuo, SIyuan Qiu, Hu Xu, Chun Li, Jianfeng Xu
cs.LG
large language model
大语言模型相关
Abstract
Low-rank adapters (LoRA) make it cheap to fine-tune a large language model once per task, but combining several independently trained adapters into one model remains difficult: merging the updates in weight space causes interference, retraining on all task data is expensive, and routing between separate adapters gives up the goal of a single combined model. We trace the difficulty to two choices that every composition method makes implicitly. A LoRA update admits infinitely many equivalent factorizations; the choice among them is invisible while an adapter serves alone, but it determines what a learned interaction between adapters can see. A coupling between an old skill and a new one can likewise point in either direction, and the direction decides whether the old skills keep computing what they computed before. We introduce READ (Read-only Expansion of Adapter Deltas), which fixes both choices: each adapter is rewritten into a balanced canonical form that preserves its update exactly, and the coupling grows in one direction only, so a new skill can read the input subspaces of old skills but cannot write into their output subspaces. The only trainable object at each append is the new skill's row of the coupling matrix, and the composed update folds into the base weights with no inference cost, routing, or task-specific rules. We evaluate READ across four benchmark suites and two model families, adding skills one at a time. Across several families, READ improves every suite average over the strongest published baselines built from the same adapters---by more than twenty points on SuperGLUE and more than seven points on the domain suite---and nearly all complete addition sequences end above every direct baseline. Factor coordinates and coupling direction, which a lone adapter never exposes, are what decide whether composed skills survive.
Chinese Translation
低秩适配器(LoRA)使得针对每个任务对大语言模型进行一次微调变得廉价,但将若干个独立训练的适配器组合成一个模型仍然困难:在权重空间中合并更新会造成干扰,在所有任务数据上重新训练代价高昂,而在各个独立适配器之间进行路由则放弃了获得单一组合模型的目标。我们将这一困难追溯到每一种组合方法都会隐含做出的两个选择。一个 LoRA 更新容许无穷多个等价的分解;当适配器单独使用时,这些分解之间的选择是不可见的,但它决定了适配器之间学得的交互能够看到什么。旧技能与新技能之间的耦合同样可以指向两个方向中的任意一个,而方向决定了旧技能是否继续计算它们此前所计算的内容。我们提出 READ(Read-only Expansion of Adapter Deltas,适配器增量的只读扩展),它同时固定了这两个选择:每个适配器被重写为一种平衡的规范形式,该形式精确保持其更新不变;而耦合只沿一个方向增长,因此新技能可以读取旧技能的输入子空间,但不能写入它们的输出子空间。每次追加时唯一可训练的对象是耦合矩阵中属于新技能的那一行,并且组合后的更新会折叠进基础权重之中,不带来任何推理开销、路由或任务特定规则。我们在四个基准套件和两个模型家族上评估 READ,每次添加一个技能。在多个模型家族中,READ 在每一个套件的平均值上都优于由相同适配器构建的最强已发表基线——在 SuperGLUE 上高出二十多个点,在领域套件上高出七个多点——并且几乎所有完整的添加序列最终都高于每一个直接基线。因子坐标与耦合方向——这是单个适配器从不暴露的——才是决定组合技能能否存活的关键。
cs.LG / 68 / 2609.31603
User Model Extraction via Belief Self-Distillation
通过信念自蒸馏提取用户模型
Ali Holmov, Yiran Huang, Kirill Bykov, Zeynep Akata
cs.LG · cs.CL
large language model
大语言模型相关
Abstract
Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling beliefs from natural conversations without external annotations. Unlike conventional probing, BSD isolates not only information present in activations, but a state whose causal role can be directly tested. Across multiple model families, BSD faithfully recovers user beliefs and enables substantially stronger interventions than matched hidden-state steering. Crucially, we find that refusal depends not only on the request, but on the model's inferred user intent: changing this belief alters refusal while holding the request fixed. We further uncover a striking cross-model regularity: independently trained LLMs converge on a shared geometry for representing their users. Together, these results reveal implicit user models as readable and causally writable internal states with direct implications for AI safety, shaping how models condition safety decisions on whom they believe they are interacting with.
Chinese Translation
大型语言模型(LLM)隐式地推断其用户的属性,并相应地调整其行为,但这些信念仍然难以被检视和因果性地操控。我们提出信念自蒸馏(BSD),这是一个统一的读-写框架,通过学习一种紧凑的用户表示来连接线性探测与因果探测,该表示既可以被解码,也可以被写回模型之中。冻结的LLM充当其自身的教师,从自然对话中蒸馏出信念,而无需外部标注。与传统探测不同,BSD不仅分离出激活中存在的信念,还分离出一种其因果作用可以被直接检验的状态。在多个模型家族中,BSD能够忠实地恢复用户信念,并实现比匹配的隐藏状态引导强得多的干预。至关重要的是,我们发现拒绝不仅取决于请求,还取决于模型推断出的用户意图:在保持请求不变的情况下改变这一信念会改变拒绝行为。我们进一步揭示了一个引人注目的跨模型规律性:独立训练的LLM在表示其用户时收敛于一种共享的几何结构。总之,这些结果揭示了隐式用户模型是可读且可因果写入的内部状态,对人工智能安全具有直接影响,并塑造了模型如何根据它们认为自己正在与谁交互来对安全决策进行条件化。
cs.MA / 69 / 2609.30691
ADF-EA: A Unified Execution Assurance System for Agent Device Foundation
ADF-EA:面向智能体设备基础(Agent Device Foundation)的统一执行保障系统
Xuechun Li, Jiaxin Liang, Jie Li, baolong Li, Jue Wang, Peng Yuan, Hang Huang
cs.MA
large language model
大语言模型相关
Abstract
Agents based on large language models (LLMs) can access heterogeneous devices through tools and APIs, but reliable execution must account for unmet effects, uncertain outcomes, and changing prerequisites. A command may be acknowledged without producing its intended effect, while missing feedback may obscure an action that has already succeeded. We present Agent Device Foundation--Execution Assurance (ADF-EA), an architecture that connects agent planning and device execution through shared capability contracts. Device Capability Contracts (DCCs) unify invocation conditions, intended effects, evidence requirements, and recovery rules across heterogeneous interfaces. Agents use these contracts to plan, while the runtime applies the same semantics to authorize actions, verify effects, and govern continuation and completion. Persistent execution state retains verified progress, unresolved outcomes, and remaining budgets across plan revisions, enabling observation-based recovery, authorized retries, and necessary state repair. We formalize the execution lifecycle and establish conditional soundness properties for completion and recovery authorization. Evaluations span multiple LLMs, five agent frameworks, and simulated process-control, household, and robotic manipulation domains. Compared with direct invocation and existing execution-checking approaches, ADF-EA reduces false completion and unnecessary repetition, supports necessary state repair, prevents calls to unavailable capabilities, and preserves permitted task completion and recovery. These results demonstrate DCCs as a reusable semantic foundation for agent autonomy across heterogeneous devices, unifying capability-based planning, evidence-grounded execution, and authorized recovery within one architecture.
Chinese Translation
基于大语言模型(LLMs)的智能体可以通过工具和 API 访问异构设备,但可靠执行必须考虑未满足的效果、不确定的结果以及不断变化的前提条件。一条命令可能被确认,却未产生其预期效果;而缺失的反馈可能掩盖一个已经成功的动作。我们提出智能体设备基础——执行保障(ADF-EA),一种通过共享能力契约连接智能体规划与设备执行的架构。设备能力契约(DCCs)统一了跨异构接口的调用条件、预期效果、证据要求和恢复规则。智能体使用这些契约进行规划,而运行时应用相同的语义来授权动作、验证效果,并管控继续执行与完成。持久化执行状态在计划修订过程中保留已验证的进展、未解决的结果和剩余预算,从而支持基于观测的恢复、经授权的重试以及必要的状态修复。我们对执行生命周期进行形式化,并建立完成与恢复授权的条件健全性性质。评估涵盖多个 LLMs、五个智能体框架,以及模拟的过程控制、家庭和机器人操作领域。与直接调用和现有执行检查方法相比,ADF-EA 减少了虚假完成和不必要的重复,支持必要的状态修复,防止调用不可用的能力,并保持允许的任务完成与恢复。这些结果表明,DCCs 可作为跨异构设备的智能体自主性的可复用语义基础,在一个架构内统一基于能力的规划、以证据为依据的执行和经授权的恢复。
cs.MA / 70 / 2609.31422
Towards Mitigating Fabricated Consensus: The Active Provenance Gate for Multi-Agent Debate Synthesis
面向缓解虚构共识:用于多智能体辩论综合的主动溯源门控
Jakub Masłowski, Jarosław A. Chudziak
cs.MA · cs.AI · cs.CL
large language model
大语言模型相关
Abstract
Large language model-based multi-agent debate (MAD) systems are being increasingly used as complex decision pipelines in distributed processes. However, their final synthesis phase still remains inadequately controlled. Even with detailed debate logs, summarizing models are prone to fabricating smoothly written debate consensus that is not grounded in the debate's history. To address this safety gap, this paper presents empirical research and studies if the introduction of active post-debate verification can mitigate the production of such factually unsupported summaries, while still providing valuable information. Furthermore, it is examined whether explicitly signalling divergence is preferable in the absence of a reliable compromise. The Active Provenance Gate (APG) is introduced as a post-debate verification layer that treats the source as a hard constraint, analysing the debate logs, auditing each claim, and applying self-correction. In crisis simulations, the self-healing mechanism more than doubles the average data Provenance Fidelity in difficult condition scenarios, before the strict gate blocks unsupported claims and generates divergence reports. In the human study, a vast majority of the users (over 75%) preferred a report explicitly stating failure in critical scenarios, despite most of them perceiving fabricated consensus from the baseline system as more fluent. Our main contribution is the transition of data origin tracing from passive logging to active conditional blocking before publication.
Chinese Translation
基于大语言模型的多智能体辩论(MAD)系统正越来越多地被用作分布式流程中的复杂决策流水线。然而,其最终的综述阶段仍然缺乏充分的控制。即使拥有详细的辩论日志,总结模型也倾向于捏造出文笔流畅、却并无辩论历史依据的辩论共识。为弥补这一安全缺口,本文开展实证研究,考察引入主动的辩论后验证能否抑制此类缺乏事实支撑的摘要的产生,同时仍提供有价值的信息。此外,本文还考察在缺乏可靠折中方案时,明确标示分歧是否更为可取。本文提出主动溯源门控(APG)作为辩论后验证层,它将来源视为硬约束,分析辩论日志、审计每一项主张,并应用自我纠正。在危机模拟中,自我修复机制在困难条件场景下使平均数据溯源保真度提高了一倍以上,随后严格的门控会阻断缺乏支撑的主张并生成分歧报告。在人类受试者研究中,绝大多数用户(超过 75%)更偏好明确说明在关键场景中失败的报告,尽管其中大多数人认为基线系统生成的虚构共识更为流畅。我们的主要贡献在于将数据来源追踪从被动记录转变为发布前的主动条件性阻断。
cs.AI / 71 / 2609.31397
Intent2Tc: Automated Intent-to-Traffic Control Translation with Language Models
Intent2Tc:基于语言模型的意图到流量控制自动翻译
Andrea Masini, Sudipta Acharya, Paolo Bellavista, Luca Foschini, Burak Kantarci
cs.NI · cs.AI · cs.CL
large language model
大语言模型相关
Abstract
Automated and highly usable Quality-of-Service (QoS) enforcement requires translating high-level service intents into deployable traffic-management policies. Although intent-based networking (IBN) has simplified policy specification, bridging the gap between business-level intents and executable network configurations remains complex, error-prone, and difficult to automate. This paper presents Intent2Tc, a closed-loop language-model-driven framework that translates business-level traffic-shaping intents into declarative sub-intents and subsequently into validated, executable Linux traffic control (tc) configurations. The framework integrates an Active Queue Management (AQM)-based digital twin (DT) semantic model, automated metadata extraction, critique-driven refinement, and Retrieval-Augmented Generation (RAG)-based knowledge reuse to improve semantic consistency and configuration reliability. We evaluate multiple open-source large language models (LLMs) and small language models (SLMs), together with Claude Sonnet-4.6, on 100 Request for Comments (RFC) 9315-compliant traffic-shaping intents. Across both translation stages, Intent2Tc achieves high semantic fidelity, configuration accuracy, and deployment readiness, with Claude Sonnet-4.6 reaching 0.98 semantic similarity, 1.0 semantic unit coverage, and 0.045 normalized edit distance. Furthermore, RAG reduces token consumption and inference latency while enabling compact models such as Phi-4-mini to approach the performance of substantially larger models. Linux tc serves as the target configuration platform, demonstrating the practical applicability of the proposed framework.
Chinese Translation
自动化且高度可用的服务质量(QoS)实施需要将高层服务意图转换为可部署的流量管理策略。尽管基于意图的网络(IBN)简化了策略规范,但弥合业务级意图与可执行网络配置之间的差距仍然复杂、易错且难以自动化。本文提出了 Intent2Tc,一个闭环的、由语言模型驱动的框架,它将业务级流量整形意图转换为声明式子意图,并随后转换为经过验证的、可执行的 Linux 流量控制(tc)配置。该框架集成了基于主动队列管理(AQM)的数字孪生(DT)语义模型、自动化元数据提取、批评驱动的细化,以及基于检索增强生成(RAG)的知识复用,以提高语义一致性和配置可靠性。我们在 100 个符合 Request for Comments(RFC)9315 的流量整形意图上,评估了多个开源大语言模型(LLM)和小语言模型(SLM),以及 Claude Sonnet-4.6。在两个翻译阶段中,Intent2Tc 均实现了高语义保真度、配置准确性和部署就绪度,其中 Claude Sonnet-4.6 达到了 0.98 的语义相似度、1.0 的语义单元覆盖率和 0.045 的归一化编辑距离。此外,RAG 降低了 token 消耗和推理延迟,同时使 Phi-4-mini 等紧凑模型能够接近规模大得多的模型的性能。Linux tc 作为目标配置平台,展示了所提框架的实际适用性。
cs.CR / 72 / 2609.31110
AuthGuard-R: Safety-Compliant Mission Hijacking and Dual-Gate Defense for LLM-Controlled Robots
AuthGuard-R:面向LLM控制机器人的安全合规型任务劫持与双门防御
Saidattu Chepuri, Vikas Srivastava
cs.RO · cs.CR
large language model
大语言模型相关
Abstract
Large language models are increasingly used as high-level planners for mobile robots, robot manipulators, and autonomous vehicles. Recent studies show that these systems can be influenced through malicious text, speech, visual instructions, retrieved documents, and poisoned sensory context. Most defenses ask whether a proposed action is physically safe. This paper studies a different problem: an action may be physically safe and still violate the mission authorized by the user. An attacker may redirect a delivery robot, replace an approved object, extend a robot's operating region, activate an unnecessary sensor, or delay a mission without creating an immediate physical hazard. We call this attack \emph{safety-compliant mission hijacking}. We propose MissionPAIR, an adaptive attack framework that searches for executable plans that pass a safety gate while violating an authenticated mission. We also propose AuthGuard-R, a deterministic authorization layer that binds every executable action to a signed mission, robot identity, object and region scope, current state, time, and input provenance. AuthGuard-R operates with an independent safety gate, giving a dual-gate architecture. We formalize mission policies over robot traces, define security games, and prove authorization soundness, mission non-escalation, replay resistance, robot binding, provenance separation, threshold-approval security, audit-log tamper evidence, and trace-level composition. We report a preliminary cross-model evaluation with Claude Haiku~4.5 and the open-source Qwen2.5~7B planner. Across 240 live attack trials, the planners followed an injected mission deviation in 109 trials; AuthGuard-R rejected all 109 resulting unauthorized actions. A separate hand-constructed suite of eleven protocol- and policy-level attacks was also blocked completely.
Chinese Translation
大语言模型正日益被用作移动机器人、机器人机械臂和自动驾驶车辆的高层规划器。近期研究表明,这些系统可以通过恶意文本、语音、视觉指令、检索到的文档以及被投毒的传感上下文而受到影响。大多数防御都在询问某个提议动作是否物理安全。本文研究一个不同的问题:一个动作可能物理安全,却仍然违反用户授权的任务。攻击者可能重定向配送机器人、替换已批准的对象、扩展机器人的运行区域、激活不必要的传感器,或延迟任务,而不造成即时物理危险。我们将这种攻击称为 \emph{安全合规型任务劫持}。我们提出 MissionPAIR,一种自适应攻击框架,它搜索能够通过安全门、同时违反已认证任务的可执行计划。我们还提出 AuthGuard-R,一种确定性授权层,它将每个可执行动作绑定到已签名的任务、机器人身份、对象和区域范围、当前状态、时间以及输入来源。AuthGuard-R 与独立的安全门协同运行,形成双门架构。我们在机器人轨迹上形式化任务策略,定义安全博弈,并证明授权健全性、任务不可升级性、抗重放性、机器人绑定、来源分离、阈值批准安全性、审计日志防篡改证据以及轨迹级组合性。我们报告了一项使用 Claude Haiku~4.5 和开源 Qwen2.5~7B 规划器进行的初步跨模型评估。在 240 次实时攻击试验中,规划器在 109 次试验中遵循了注入的任务偏离;AuthGuard-R 拒绝了所有 109 个由此产生的未授权动作。另一套手工构建的包含十一种协议级和策略级攻击的测试集也被完全阻止。
cs.AI / 73 / 2609.31383
Guiding End-to-End Driving Models with Endpoint-Constrained Trajectory Optimization
以端点约束轨迹优化引导端到端驾驶模型
Brayden Zhang, Mahsa Golchoubian, Igor Gilitschenski, Boris Ivanovic, Kashyap Chitta
cs.RO · cs.AI · cs.CV · cs.LG
diffusion
扩散模型相关
Abstract
End-to-end driving policies are commonly trained through open-loop behavior cloning, yet they must ultimately operate in closed-loop when deployed on a vehicle, creating a fundamental mismatch between training and execution. Beyond the commonly studied effects of covariate shift and causal confusion, we identify a complementary factor for this open-loop/closed-loop gap: waypoint-based supervision and displacement metrics do not ensure that the intermediate trajectory is physically coherent or easy for the controller to track. We observe that these inconsistencies concentrate primarily at intermediate waypoints, while the predicted endpoint remains comparatively reliable. Based on this observation, we introduce Endpoint-Constrained Optimization (ECO), a lightweight postprocessing layer that anchors the trajectory to the vehicle's executed history, preserves the policy's predicted endpoint, and reshapes the intermediate waypoints to improve feasibility. ECO requires no map, privileged simulator state, or additional training, and can be inserted between a broad range of waypoint-emitting policies and their controllers. Across two closed-loop simulators, it improves the aggregate closed-loop score of all six evaluated generative and regression-based driving policies, and the gains tend to increase with how often the base plans violate motion limits. On HUGSIM, ECO improves VaVAM from 18.1 to 31.0 HD-Score (+71%), achieving 1st place on the HUGSIM Closed-Loop Driving Challenge. Similarly, on AlpaSim, ECO increases the scene scores of VaVAM and DiffusionDrive by 123% and 22%, respectively. These results show that for a broad collection of end-to-end driving models, repairing the intermediate geometry of predicted trajectories without changing the policy's predicted endpoint can substantially improve closed-loop performance.
Chinese Translation
端到端驾驶策略通常通过开环行为克隆进行训练,然而当部署到车辆上时,它们最终必须在闭环下运行,这就造成了训练与执行之间的根本性失配。除了常被研究的协变量偏移与因果混淆效应之外,我们还识别出导致这一开环/闭环差距的一个补充因素:基于路点的监督与位移指标并不能保证中间轨迹在物理上自洽,或者易于被控制器跟踪。我们观察到,这些不一致主要集中在中间路点上,而预测的端点则相对可靠。基于这一观察,我们提出了端点约束优化(Endpoint-Constrained Optimization,ECO),这是一个轻量级的后处理层,它将轨迹锚定到车辆已执行的历史上,保留策略预测的端点,并重塑中间路点以提升可行性。ECO 不需要地图、特权仿真器状态或额外训练,并且可以插入到广泛的输出路点的策略与其控制器之间。在两个闭环仿真器中,它提升了所评估的全部六个生成式与基于回归的驾驶策略的总体闭环得分,并且当基础规划越频繁地违反运动限制时,增益往往越大。在 HUGSIM 上,ECO 将 VaVAM 从 18.1 提升到 31.0 HD-Score(+71%),在 HUGSIM 闭环驾驶挑战赛中夺得第 1 名。类似地,在 AlpaSim 上,ECO 将 VaVAM 与 DiffusionDrive 的场景得分分别提升了 123% 和 22%。这些结果表明,对于广泛的一类端到端驾驶模型而言,在不改变策略预测端点的前提下修复预测轨迹的中间几何形状,可以显著提升闭环性能。
cs.CL / 74 / 2609.30784
Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models
面向冻结语言模型的事后音频扩展的共生架构
Yotaro Kubo, Qi Sun, Yujin Tang
cs.SD · cs.CL · eess.AS
large language model
大语言模型相关
Abstract
This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes audio-conditioned vectors directly into the target LLM's short-term memory, i.e., the key-value (KV) cache, enabling the LLM to behave as an audio language model (ALM). The architectural advantages are twofold. First, it improves the scalability of ALMs: because the proposed method bypasses the LLM during audio injection, the injection cost is governed by the injector width rather than the backbone width, and can therefore scale more slowly than the cost of full-backbone prefilling. Second, since the training scheme does not update the LLM weights, the original capabilities of the LLM are preserved without the risk of degradation from fine-tuning. The effectiveness of the proposed method is evaluated on both audio-understanding tasks (automatic speech recognition, audio question answering, and acoustic scene classification) and text-only tasks. We confirm that, while activating fewer parameters during audio prefilling, our architecture outperforms the conventional method with a frozen LLM and approaches the performance of a fine-tuned ALM, all while preserving the backbone LLM's original text-only task performance by construction.
Chinese Translation
本文提出了一种架构,使大语言模型(LLMs)无需微调其权重即可具备音频理解能力。所提出的共生架构采用一个注入器模块,该模块将音频条件化的向量直接写入目标 LLM 的短期记忆,即键值(KV)缓存,从而使该 LLM 能够表现为音频语言模型(ALM)。该架构的优势有两方面。首先,它提升了 ALM 的可扩展性:由于所提方法在音频注入期间绕过了 LLM,注入成本由注入器宽度而非主干宽度决定,因此其增长可以比全主干预填充的成本更慢。其次,由于该训练方案不更新 LLM 权重,LLM 的原始能力得以保留,不存在因微调而退化的风险。所提方法的有效性在音频理解任务(自动语音识别、音频问答和声学场景分类)以及纯文本任务上进行了评估。我们确认,在音频预填充期间激活更少参数的同时,我们的架构优于使用冻结 LLM 的传统方法,并接近微调后的 ALM 的性能,同时从构造上保留了主干 LLM 原有的纯文本任务性能。
cs.SE / 75 / 2609.30642
A Framework for Identifying, Categorizing, and Explaining Bias in AI-Generated Code
一个用于识别、分类和解释 AI 生成代码中偏见的框架
Manaal Basha, Aimee M. Ribeiro, Gema Rodriguez-Perez
cs.SE · cs.AI
large language model
大语言模型相关
Abstract
As Large Language Models (LLMs) become integrated into software development workflows, concerns regarding unintentional biases in AI-generated code. Although evidence suggests these biases exist, limited research has systematically identified, categorized, and explained them. This study investigates bias in AI-generated code and evaluates whether LLMs can reliably identify and explain it through a taxonomy-driven framework. We extended an existing dataset of biased AI-generated Python code and manually annotated snippets with bias categories and human-authored justifications to establish a ground-truth dataset. Using this dataset, we evaluated proprietary and open-source LLMs as automated bias detection and justification systems through ICL. Finally, we analyzed similarity between LLM-generated explanations and human-authored justifications using structured justification and code identification metrics. Our findings demonstrate that LLMs can effectively support code bias identification and explanation. Gemini achieved 80.14% classification accuracy, with 84.0% precision and 95.7% recall, while the best open-source alternative, Qwen3-coder, achieved 82.45% accuracy, 68.64% precision, and 80.22% recall. Additionally, the models achieved justification similarity scores of 80.4% and 80.14%, respectively, relative to human-authored reasoning, and code identification similarity scores of 86.0% and 87.82%. These results suggest that LLMs can detect biased logic in generated Python code and produce explanations that substantially align with expert interpretations.
Chinese Translation
随着大型语言模型(LLMs)被整合进软件开发工作流,人们开始担忧 AI 生成代码中存在的无意偏见。尽管有证据表明这些偏见确实存在,但系统性地识别、分类并解释它们的研究仍然有限。本研究考察 AI 生成代码中的偏见,并评估 LLMs 能否通过一个由分类体系驱动的框架可靠地识别并解释它。我们扩展了一个现有的、由 AI 生成且带有偏见的 Python 代码数据集,并人工为代码片段标注偏见类别与人类撰写的理由说明,从而建立了一个真值数据集。使用该数据集,我们通过 ICL 将专有与开源 LLMs 评估为自动化的偏见检测与理由说明系统。最后,我们使用结构化理由说明指标与代码识别指标,分析了 LLM 生成的解释与人类撰写的理由说明之间的相似性。我们的发现表明,LLMs 能够有效支持代码偏见的识别与解释。Gemini 达到了 80.14% 的分类准确率,精确率为 84.0%,召回率为 95.7%;而最佳的开源替代方案 Qwen3-coder 达到了 82.45% 的准确率、68.64% 的精确率与 80.22% 的召回率。此外,相对于人类撰写的推理,这些模型分别取得了 80.4% 与 80.14% 的理由说明相似性得分,以及 86.0% 与 87.82% 的代码识别相似性得分。这些结果表明,LLMs 能够检测所生成 Python 代码中的偏见逻辑,并产生与专家解读实质上高度一致的解释。
cs.SE / 76 / 2609.31288
When the Model Retires: An Empirical Study of LLM Migration in Open-Source Applications
当模型退役时:开源应用中LLM迁移的实证研究
Hyungjin Lukas Kim
cs.SE
large language model
大语言模型相关
Abstract
Applications built on commercial large language model (LLM) APIs depend on model versions that providers retire on their own schedule, with notice periods ranging from one year to two weeks. We ask what actually happens to applications when a model is retired. We mine GitHub for commits that migrate away from officially deprecated models and endpoints of OpenAI, Anthropic, and Google, matching each commit to the provider's published announcement and shutdown dates. From 22,555 commits in 17,703 non-fork repositories (2024-2026), 5,139 are matched to an official event; two independent coders validated a stratified sample of 300 (kappa = 0.89-0.95), and we reweight all estimates by their labels. We find that an estimated 82% (95% CI 79-84) of migrations away from retired models were committed after the shutdown date - after the application had started failing - regardless of repository popularity, prior retirement experience, or the presence of a provider-abstraction layer. The share tracks the provider's notice policy: 89% for Anthropic's 60-114-day notices versus 13% for OpenAI's one-year Assistants API notice, and each e-fold increase in notice length reduces the odds of post-shutdown migration by about three quarters. Model identifiers are hard-coded in 94% of migrating applications, migration effort scales from a median of 6 added lines for prompt-only applications to nearly 700 for fine-tuned ones, and only 8% of migrations switch provider. We release the dataset and pipeline and discuss implications for deprecation policy, dependency-risk assessment of LLM products, and tooling.
Chinese Translation
构建在商业大型语言模型(LLM)API 之上的应用程序依赖于模型版本,而提供商会按自己的时间表淘汰这些版本,通知期从一年到两周不等。我们追问:当模型退役时,应用程序究竟会发生什么?我们挖掘 GitHub 上那些从 OpenAI、Anthropic 和 Google 的官方弃用模型与端点迁移离开的提交,并将每个提交与提供商公布的公告和关停日期相匹配。从 17,703 个非 fork 仓库(2024-2026)中的 22,555 个提交里,有 5,139 个与官方事件相匹配;两名独立编码者验证了一个 300 的分层样本(kappa = 0.89-0.95),并且我们根据他们的标签对所有估计值重新加权。我们发现,估计有 82%(95% CI 79-84)的从退役模型迁移的提交发生在关停日期之后——即应用程序已经开始失败之后——无论仓库的流行度、以往的退役经验,还是是否存在提供商抽象层。这一比例与提供商的通知政策相关:Anthropic 的 60-114 天通知为 89%,而 OpenAI 的一年期 Assistants API 通知为 13%,且通知长度每增加一个 e 倍,关停后迁移的几率降低约四分之三。在 94% 的迁移应用程序中,模型标识符是硬编码的;迁移工作量从中位数 6 行新增代码(仅提示词应用)扩展到接近 700 行(微调应用),并且只有 8% 的迁移更换了提供商。我们发布数据集和流水线,并讨论其对弃用政策、LLM 产品的依赖风险评估以及工具支持的意义。
cs.SE / 77 / 2609.31301
Beyond Approved Actions: Runtime Validation of Persistent Outcomes in Agent Workflows
超越已批准动作:智能体工作流中持久性结果的运行时验证
Haoran Zhang, Hengtong Zhang, Zhiyu Liang, Yu Yan, Decheng Zuo, Hongzhi Wang
cs.SE · cs.AI
large language model
大语言模型相关
Abstract
Large language model agents increasingly act on software systems, no longer merely generating text but also changing databases and online services. However, an approved database update may succeed yet leave an unapproved notification because execution can produce persistent effects beyond the requested change. Current safeguards can approve an action or record its aftermath, but without checking the persistent result before continuation, an unapproved outcome can be accepted as success and propagated to later steps. We present EffectMatch, a runtime that collects persistent changes within a controlled execution boundary and compares them with what the application approved for the current state and execution. The comparison governs commit and dependent execution. In comparative evaluation on 206 public business tasks, EffectMatch preserved all clean executions and prevented all tested incorrect commits. Six 20-run ablations exposed the failure caused by each removed mechanism, while 80 task-topology cases preserved truthful handoffs and blocked invalid continuation. Together, these results show that EffectMatch blocks the silent acceptance and downstream propagation of persistent outcomes inconsistent with application approval.
Chinese Translation
大语言模型智能体越来越多地对软件系统采取行动,不再仅仅生成文本,还改变数据库和在线服务。然而,一次已批准的数据库更新可能成功执行,却留下未经批准的通知,因为执行可能产生超出所请求变更的持久性影响。当前的防护措施可以批准一个动作或记录其后果,但若不在继续之前检查持久性结果,未经批准的结果就可能被当作成功接受并传播到后续步骤。我们提出 EffectMatch,一种运行时,它在受控执行边界内收集持久性变更,并将其与应用程序针对当前状态和执行所批准的内容进行比较。该比较决定提交与依赖执行。在针对 206 个公开业务任务的对比评估中,EffectMatch 保留了所有干净执行,并阻止了所有被测试的错误提交。六项各运行 20 次的消融实验揭示了每移除一个机制所导致的失败,而 80 个任务拓扑案例则保持了真实的移交并阻止了无效的继续执行。这些结果共同表明,EffectMatch 阻止了与应用批准不一致的持久性结果被静默接受并向下游传播。
cs.AI / 78 / 2609.31070
Quantum Diffusion Models for Medical Image Analysis
用于医学图像分析的量子扩散模型
Francesco Aldo Venturelli, Stefano Martina, Marco Parigi, Filippo Caruso, Alba Cervera-Lierta, Miguel A. González Ballester
eess.IV · cs.AI · cs.CV · cs.ET · cs.LG · quant-ph
diffusion
扩散模型相关
Abstract
Quantum Machine Learning is a novel field of research aimed at devising machine learning approaches exploiting principles of quantum mechanics, such as superposition, entanglement and interference. In this context, we present a scalable hybrid Quantum Diffusion Model, and evaluate its use for medical image analysis. Specifically, our method is based on a Discrete-Time Quantum Walk algorithm, executed on a real quantum device, to model the forward dynamics of the diffusion model. For the backward step of the diffusion model, we devise and evaluate a classical learning model, which is used to reversely denoise the data. In contrast with other existing attempts at applying quantum machine learning for image analysis tasks, severely limited by the size of existing quantum devices, our method allows to process real-world large size medical data. In particular, we present results on grayscale and RGB images, as well as 3D volumes of moderate sizes. We benchmark our results by reproducing an alternative classical counterpart model, based on diffusion models on discrete state spaces. By doing so, we compare the generation capabilities of both models in terms of three distinct state-of-the-art metrics in the field of image generation, showing the competitive, promising results of our approach.
Chinese Translation
量子机器学习是一个新颖的研究领域,旨在设计利用量子力学原理(如叠加、纠缠和干涉)的机器学习方法。在此背景下,我们提出一种可扩展的混合量子扩散模型,并评估其在医学图像分析中的用途。具体而言,我们的方法基于一种在真实量子设备上执行的离散时间量子行走算法,以对扩散模型的前向动力学进行建模。对于扩散模型的反向步骤,我们设计并评估了一个经典学习模型,用于对数据进行反向去噪。与将量子机器学习应用于图像分析任务的其他现有尝试相比——这些尝试受到现有量子设备规模的严重限制——我们的方法能够处理真实世界的大尺寸医学数据。特别地,我们展示了在灰度图像和 RGB 图像,以及中等尺寸的三维体数据上的结果。我们通过复现一个基于离散状态空间扩散模型的替代经典对应模型,对我们的结果进行基准测试。通过这样做,我们根据图像生成领域中三个不同的最先进指标,比较了两个模型的生成能力,展示了我们的方法具有竞争力且有前景的结果。
cs.LG / 79 / 2609.30539
Learning to Replace MCMC in Split-Gibbs Diffusion Posterior Sampling via Deep Unfolding
通过深度展开学习替换Split-Gibbs扩散后验采样中的MCMC
Yi Zhang, Rui Guo, Mengchu Xu, Zhaofeng Liu, Yonina C. Eldar
eess.SP · cs.LG · stat.ML
diffusion
扩散模型相关
Abstract
Split Gibbs sampling enables diffusion posterior inference for general nonlinear inverse problems by decoupling prior and likelihood computations, allowing a pretrained diffusion prior to be reused across measurement models. However, its likelihood update often relies on iterative MCMC, which can hinder parallelization, require algorithm-specific tuning, and incur substantial computational cost. In this work, we propose a learning-based framework to replace this MCMC step by reformulating both Gibbs updates as Gaussian denoising problems and implementing them through ODE diffusion. The prior step reuses a pretrained denoiser, while the likelihood denoiser exploits known likelihood structure through a lightweight deep-unfolded network. Experiments on nonlinear phase retrieval demonstrate the effectiveness of the proposed method as an alternative to MCMC-based split Gibbs at lower likelihood-update cost.
Chinese Translation
Split Gibbs采样通过解耦先验和似然计算,使扩散后验推断能够用于一般非线性逆问题,并允许预训练的扩散先验在不同测量模型之间重用。然而,其似然更新通常依赖于迭代MCMC,这可能阻碍并行化、需要特定于算法的调参,并带来可观的计算成本。在这项工作中,我们提出一个基于学习的框架来替换该MCMC步骤,其方式是将两个Gibbs更新都重新表述为高斯去噪问题,并通过ODE扩散来实现它们。先验步骤重用预训练去噪器,而似然去噪器通过轻量级深度展开网络利用已知的似然结构。在非线性相位恢复上的实验证明了所提方法作为基于MCMC的split Gibbs替代方案的有效性,且具有更低的似然更新成本。
cs.LG / 80 / 2609.31458
Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers
在日益增长的几何复杂度下的非参数上下文学习:Transformer 的极小极大最优性与局部几何自适应性
Jaehee Seo, Jisu Kim
stat.ML · cs.LG · math.ST
large language model
大语言模型相关
Abstract
Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large language models. This success motivates understanding how transformers exploit task-relevant structure in geometrically heterogeneous data. However, existing nonparametric ICL theory has largely focused on Euclidean domains or single-manifold models. To address this gap, we study the prediction problem under unknown local geometry, modeled by sample size-dependent mixtures of manifolds with heterogeneous dimensions, smoothness, and sampling masses. Under local separation and small-perturbation conditions, we establish a minimax lower bound capturing the aggregate difficulty of the components and construct an oracle tangent local-polynomial estimator with a matching upper bound. This estimator is connected to a structure-informed, two-stage softmax transformer with a geometric preconditioner and chartwise reduced local-polynomial solvers. The transformer achieves negligible approximation error relative to the minimax rate with logarithmic depth and polynomial size. Finally, we derive an in-context generalization bound for near empirical risk minimizers over this class. Together, these results identify conditions under which the resulting predictor exploits local geometry and attains the aggregate minimax rate.
Chinese Translation
Transformer 已成为上下文学习(ICL)的核心架构,尤其是通过其在大语言模型中的最先进性能。这一成功促使人们理解 Transformer 如何在几何异质数据中利用与任务相关的结构。然而,现有的非参数 ICL 理论在很大程度上集中于欧几里得域或单流形模型。为弥补这一空白,我们研究未知局部几何下的预测问题,该问题由依赖于样本量的流形混合建模,这些流形具有异质的维度、光滑度和采样质量。在局部分离和小扰动条件下,我们建立了一个极小极大下界,刻画了各组成部分的总体难度,并构造了一个神谕切空间局部多项式估计器,其具有匹配的上界。该估计器与一个结构信息驱动的两阶段 softmax Transformer 相联系,后者带有几何预条件子和按图卡约简的局部多项式求解器。该 Transformer 在对数深度和多项式规模下,相对于极小极大速率实现了可忽略的逼近误差。最后,我们为该类上的近似经验风险最小化器推导了一个上下文泛化界。综合起来,这些结果确定了所得预测器利用局部几何并达到总体极小极大速率的条件。
cs.LG / 81 / 2609.31612
First-Order Stationarity of Reverse Diffusions
反向扩散的一阶平稳性
Zhifeng Chen, Chenyang Jiang, Yazhen Wang
stat.ML · cs.LG
diffusion
扩散模型相关
Abstract
Recent literature has shown a strong connection between optimization and sampling. We develop the corresponding first-order theory for diffusion models. First, the SDE-based reverse-time flows of overdamped and underdamped Langevin diffusions contract relative Fisher divergences at explicit exponential rates whenever the stationary potential of the forward process is strongly convex---a condition on the noising process one chooses, not on the data. This is a unique advantage of SDE-based reverse diffusion, absent in the reverse process based on ODEs. Second, we incorporate discretization and establish averaged first-order stationarity bounds---the sampling analog of averaged gradient-norm guarantees in nonconvex optimization---for samplers of both overdamped and underdamped diffusion models. As in nonconvex optimization, the convexity-free certificate is local: it guarantees score consistency, not global mode weights.
Chinese Translation
近期文献表明,优化与采样之间存在紧密联系。我们为扩散模型发展了相应的一阶理论。首先,当正向过程的平稳势是强凸时——这是对所选择的加噪过程的条件,而不是对数据的条件——基于 SDE 的过阻尼和欠阻尼 Langevin 扩散的反向时间流以显式指数速率收缩相对 Fisher 散度。这是基于 SDE 的反向扩散的一个独特优势,而基于 ODE 的反向过程不具备这一点。其次,我们纳入离散化,并为过阻尼和欠阻尼扩散模型二者的采样器建立平均一阶平稳性界——这是非凸优化中平均梯度范数保证的采样对应物。正如非凸优化中那样,这个不依赖凸性的证书是局部的:它保证得分一致性,而不是全局模态权重。
人工智能 (cs.AI)
109
cs.AI / 1 / 2609.30383
Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems
Zihao Zhu, Siwei Lyu, Adel Bibi, Baoyuan Wu
cs.AI
Abstract
A skill is a modular package of natural-language instructions, executable scripts, and reference resources that an agent can load at runtime to extend its capabilities for a specific task. Skill-based agent systems therefore enable flexible reuse of third-party capabilities, but the openness of this skill ecosystem also opens up a new attack surface. Prior work has focused on vulnerabilities within individual skills, but little attention has been paid to risks that arise from interactions across skills. In this paper, we introduce skill cascading attacks, a threat paradigm in which a malicious objective is distributed across multiple skills so that each modification looks benign in isolation, yet their combined execution is harmful. For instance, in a prescription-review pipeline, the first skill weakens signals of recently discontinued medications in the extracted history, the second downgrades the severity of any drug interaction tied to them, and the third suppresses the resulting low-priority alert in the final summary, so that a severe drug-interaction warning silently disappears before reaching the physician. To systematically study this safety blind spot, we develop SkillCascade, an automated multi-agent red-teaming framework, and release SkillCascade-Bench, a benchmark of 213 validated cascading test cases across multiple agent systems and domains. Across representative agents (e.g., OpenClaw, Claude Code, Codex) and LLM backbones, cascaded interactions reliably induce harmful behaviors while evading existing per-skill scanners and runtime monitors. Our findings highlight a gap between component-level integrity and system-level safety, and call for defenses that reason over cross-skill interactions rather than individual skills in isolation.
cs.AI / 2 / 2609.30397
A Synthetic Ground-Truth Framework for the Evaluation of Explainable AI Methods
Miquel Miró-Nicolau, Francesco Spinnato, Riccardo Guidotti
cs.AI
Abstract
Evaluating explainable Artificial Intelligence (XAI) methods is a challenging task due to the lack of reliable evaluation procedures and, in particular, the absence of ground truth explanations. In the literature, existing evaluation approaches typically assess explanations by measuring their fidelity with respect to the predictions of a black-box model. However, such evaluation strategies only quantify the degree to which an explanation reproduces the model's output, without ensuring that the explanation correctly reflects the underlying decision process. As a consequence, different explanations may achieve similar fidelity scores while providing inconsistent or misleading interpretations of the model behavior. In this paper, we propose a framework for the evaluation of XAI methods based on synthetic ground truth. The proposed approach relies on controlled interventions to generate synthetic datasets in which the importance of input components can be determined by design. This enables the construction of ground truth explanations that are directly aligned with the behavior of the model under analysis. The framework is instantiated across three data domains, namely binary images, tabular data, and time series, allowing a comprehensive assessment of explanation methods in heterogeneous settings. Experimental results obtained by evaluating nine widely used XAI methods show significant limitations in current techniques and highlight the importance of synthetic, intervention-based benchmarks for a reliable assessment of explanation quality.
cs.AI / 3 / 2609.30446
Predicting Transmembrane Protein Topology from 3D Structure
Sitong Chen, Xiaopeng Mao
cs.AI
Abstract
This paper presents a novel approach to infer protein topology using the state-of-the-art graph neural network (GNN), SchNet. The model is trained on the same dataset used to develop the recent DeepTMHMM model with 5-fold cross-validation. Unlike the conventional approaches based on using only the protein sequences or the $α$-carbons as features, we have decoded our classifier in this way, so all atom-level embeddings are used. Without applying any pre-trained weight, the final results have shown great potential that GNNs can be used for topological predictions.
cs.AI / 4 / 2609.30469
Pretrained ASR Pseudo-labeling for Noisy Police Audio
Kaavya Chaparala, Su Huang, Stephen L. Miller, Rhiannon N. Miller, Anjalie Field
cs.AI
Abstract
Pretrained ASR systems perform poorly on noisy Broadcast Police Communication (BPC), hindering efforts to understand police decision-making. Pseudo-labeling offers an unsupervised path to improve ASR without expensive human labels, but the efficacy of this approach on very noisy domains is not known. In this work, we systematically assess the opportunities and limits of pseudo-labeling to adapt foundation ASR models (Whisper and Qwen3-ASR) to noisy BPC domain corpora from Baltimore and Chicago. We demonstrate that existing internal confidence metrics (log-probabilities and STAR scores) fail to distinguish between high and low quality BPC pseudo-labels, and we introduce an external LLM-as-a-judge filtering paradigm that leverages parametric knowledge to discard contextually implausible transcripts. Our LLM-judging filters more aggressively than internal metrics and significantly reduces WER of the pseudo-labeled training sets across the Baltimore and Chicago BPC corpora, though a substantial gap remains relative to an oracle filter. We also introduce a new cross-model pseudo-labeling paradigm where one model is finetuned with pseudo-labels from the other, and we identify this method as a promising direction for future pseudo-labeling work.
cs.AI / 5 / 2609.30550
Benchy: towards a universal language for task-oriented AI benchmarks
Francis F Daniel, Mauro Ibañez, Francis Perelman, Marian Basti
cs.AI · stat.ME
Abstract
Benchy is a semantic language and execution engine for benchmarking AI programs. A benchmark is completely specified by a program, a scoring function, and a dataset, B=(P,S,D), and is separate from the AI-system taking it; a run binds the two, R=(B,AI). Benchmarks are authored as canonical YAML in which each semantic concept has one valid syntax, classified by a shared task/domain/language ontology, and deterministically compiled into a canonical JSON intermediate representation that the engine executes. Compilation changes representation, not meaning: it does not repair invalid definitions or inject hidden defaults. Programs use fixed schemas of named input and output fields, the leaf output fields are the scoring dimensions, and the engine exposes one universal runtime contract --- a named-field input object in, a named-field output object out --- to which external AI-systems adapt at the boundary, so integration mechanics never propagate into benchmark semantics. This paper gives the semantic object model, the ontology and task-to-program validation rule, the scoring and failure semantics, the compilation and execution architecture, and the scope of the current language. An appendix fixes the normative engineering contract for the first engine implementation.
cs.AI / 6 / 2609.30553
Rank-Reliable Teacher-Guided Fitness Approximation for Expensive Evolutionary Optimization: A TinyML Architecture Search Study
Soumen Garai, Suman Samui
cs.AI · cs.LG
Abstract
Expensive evolutionary search does not always need an exact fitness estimate for every candidate. It often needs a reliable answer to a simpler question: which candidate is better? We address this need through Teacher-Guided Learning NSGA-II (TGL-NSGA-II), a low-fidelity framework for constrained Tiny Machine Learning (TinyML) neural architecture search. A pretrained teacher organizes samples into strata defined jointly by difficulty and class. Each candidate then undergoes KD-Lite, a short and capped knowledge-distillation procedure on a compact training set, before being scored on a separate stratified evaluation set. This teacher-guided score is fused with a Gaussian-process surrogate to select candidates for full evaluation. For a fixed candidate population, we analyse evaluation variance, score concentration, pairwise rank inversion, expected Kendall-$τ$, first-front identification, and hypervolume perturbation. We also derive a variance-aware fusion weight and a capacity-adaptive distillation rule. On keyword spotting and bird-call classification, the measured Kendall-$τ$ values are 0.74 and 0.62, exceeding the corresponding predicted lower bounds of 0.60 and 0.46. Joint stratification reduces proxy-score variance by 41% relative to random evaluation. Selective teacher mismatch, in contrast, increases differential bias and reduces Kendall-$τ$ to 0.41. Under a constrained evaluation budget, TGL-NSGA-II achieves the largest mean hypervolume and smallest generational distance on keyword spotting, records the lowest mean false-positive rate on BirdCLEF, and runs 2.2x faster than full NSGA-II. These guarantees apply to population-level low-fidelity evaluation and do not establish convergence of the complete evolutionary trajectory.
cs.AI / 7 / 2609.30563
Thinking Less to Simulate Better: Intuitive Prompting Improves LLM Agents Simulating Individual Social Media Reactions, Including Unfamiliar Content
Ljubisa Bojic, Tijana Stanic, Joerg Matthes, Agariadne Dwinggo Samala, Bojana Dinic, Jue Wang
cs.AI · cs.CL · cs.HC · cs.MA · cs.SI
Abstract
Platform policies are increasingly tested on artificial users, making agent fidelity important. Yet convincing fake profiles could also manipulate perceived public opinion before elections. Validation has concentrated on agreement with human behaviour and has paid little attention to whether an agent behaves in line with the profile it was given. The present study profiled eight Serbian participants through a questionnaire, a deep interview, and a written self-presentation, recorded their reactions to sixty-eight social media posts, and asked four language models to predict those reactions under five prompt conditions varying profile content and instruction style. Attitudinal content improved prediction over demographic backstories by a wide margin. Agents matched their stated profiles more closely than participants matched their own survey answers, and consistency proved unrelated to fidelity once profile information was present. Instructing models to respond intuitively and immediately rather than analytically gave the highest fidelity of any condition and cut the compression of individual differences from seven times the human level to three. The advantage held on posts about topics the questionnaire never raised, where that condition reached the highest fidelity of any setup and beat a crowd baseline by a wide margin, which suggests that agents prompted this way could serve as general-purpose simulated users rather than specialists on the topics they were profiled for. Results may bear implications for the development of language models, because intuition-based setups appear better suited to some tasks than reasoning-based ones.
cs.AI / 8 / 2609.30569
Atelier: Learning Local Self-Supervised Features for CryoEM Volumes via Hypernetworks
Phillip Lo, Sudarshan Babu, Dari Kimanius, Aly A. Khan
cs.AI · stat.ML
Abstract
CryoEM map interpretation requires features that are spatially localized, consistent across samples, and informative across spatial scales. Most deep learning methods for map annotation extract features from fixed voxel grids. However, implicit neural representations (INRs) are able to model volumetric data as scale-agnostic, coordinate-conditioned functions. INRs are therefore attractive for cryoEM, but fitting a separate INR for each map is too expensive for large-scale feature extraction and produces representations that are not aligned across samples. We introduce Atelier, a self-supervised framework that amortizes INR fitting for reconstructed cryoEM maps. Pretrained on 5,439 Electron Microscopy Data Bank maps, Atelier is a transformer-based hypernetwork that generates high-fidelity reconstructions across a wide range of protein structures, including large multi-subunit assemblies. Beyond reconstruction, the INR generated by the pretrained transformer exposes a continuous, local feature field through its intermediate activations at any spatial query point, a property that voxel grid and patch-tokenizer architectures do not naturally provide. Used as auxiliary channels to a 3D nested U-Net annotation head trained from scratch, these coordinate-conditioned features improve performance on eight voxel-level property prediction tasks over a volume-only baseline. Our results demonstrate that amortized implicit neural representations are an effective primitive for geometry-aware analysis of cryoEM data.
cs.AI / 9 / 2609.30571
HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases
Aditya Kumaran, Rahul Singhal, Karime Maamari, Amine Mhedhbi, Pradyumna Tambwekar
cs.AI
Abstract
Language models are often evaluated on curated benchmarks that underrepresent the complexity of enterprise deployments. We introduce HARDEN, a constrained evolutionary search method to adapt the input of existing evaluation cases into more challenging variants while keeping their expected outputs fixed. HARDEN searches along generated domain-specific complexity axes while enforcing feasibility constraints such as preserving task semantics, realism, and execution validity. Across FinQA, PubMedQA, and ContractNLI and three Qwen3.5 model scales (35B-A3B, 122B-A10B, and 397B-A17B), HARDEN reduces task-model accuracy by 22.7% on average and by up to 49.9% relative to single-pass baselines using the same feasibility checks. These results show that evolutionary search can produce substantially harder valid evaluation cases.
cs.AI / 10 / 2609.30706
LAVOIR: Teaching a Single-Pass Decision Encoder When and What to Ask with Amortized Value of Information
Furkan Yilmaz, Habibe Aleyna Tasdemir, Muhammed Faruk Gozay
cs.AI · cs.CL
Abstract
"System One" decision models such as TypeSafe's Jev and its open counterpart Laya answer typed questions about a text in a single forward pass with calibrated probabilities, but they cannot ask for missing information: when a first message does not say what separates two departments, they guess. We present LAVOIR (Laya with Value-Of-Information Routing), which places the candidate pieces of missing information (slots) in the input next to the answer options, so that one forward pass returns both the decision distribution and, for every slot, the expected gain in the probability of the correct decision if the user were asked about it. VOI targets need no human labels: gold decisions come from schema rules, an LLM only verbalizes messages and answers, a model from another family checks every text, and pairing each message with several profiles makes regression on realized gains estimate the expected gain. A Gini-impurity cap bounds the predicted value by what a calibrated model can still gain. In a controlled study, decisions on seen schemas are statistically indistinguishable from the Bayes ceiling. The final model's question policy matches a greedy oracle VOI policy on seen schemas (AUC 0.799 vs. 0.797), and with at most 0.5 questions per conversation it is 14.1 points more accurate than never asking. On real ABCD conversations, one real exchange raises accuracy by 8.3 points where LAVOIR asks and leaves it unchanged where it does not; on SGD the cap lowers the asking rate from 93% to 8.6%. On Laya's twelve benchmarks LAVOIR is above Laya's reported scores on seven, and it answers a question in 31 ms (median, GH200).
cs.AI / 11 / 2609.30714
CRC-Router: Risk-Constrained Routing for Medical Agentic AI Systems
Xueyang Li, Mingze Jiang, Gelei Xu, Jun Xia, Ching-Hao Chiu, Mengzhao Jia, Danny Z. Chen, Yiyu Shi
cs.AI
Abstract
Agentic AI systems are increasingly being explored in medical imaging to improve throughput and reduce clinician workload; however, safe deployment remains challenging because autonomous errors may propagate into downstream clinical decisions. A central requirement is therefore not only strong predictive performance, but also a reliable routing mechanism that determines when the system should proceed autonomously and when a case should be escalated for further review. To address this gap, we propose CRC-Router, a risk-constrained, uncertainty-aware routing module that is applicable to both conventional medical prediction models and agentic medical AI systems. CRC-Router combines multiple complementary uncertainty signals with the predictive score to construct a per-finding routing feature vector, maps this vector to an estimated wrong-accept risk using a lightweight per-finding risk model, and then applies Conformal Risk Control (CRC) to calibrate acceptance thresholds under a user-specified risk target. Instantiated on chest X-ray multi-finding triage using the NIH ChestX-ray14 dataset, CRC-Router achieves the strongest empirical risk--coverage trade-off among the evaluated baselines, both as a standalone routing layer and as a plug-in module integrated with the state-of-the-art MedRAX agent. These results demonstrate both the effectiveness of CRC-Router in selective medical automation and its modular, model-agnostic compatibility with existing predictive and agentic medical pipelines. Code is publicly available at https://github.com/XLIAaron/CRC-Router
cs.AI / 12 / 2609.30725
Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents
Yiran Hu, Nan Jiang, Shanchao Liang, Anik Dey, Yi Wu, Lin Tan
cs.AI · cs.SE
Abstract
Although effective, coding agents often incur substantial monetary costs. Their recurring cost-inefficient behaviors remain underexplored. We conduct the first study of behavioral cost inefficiencies in coding agents, analyzing 1,200 trajectories from Claude Code and Mini-SWE-Agent across four configurations on SWE-bench Verified. We identify three cost-inefficient behaviors: subsumed retrieval, similar script generation, and test re-execution. We then evaluate three mitigation strategies: structure-aware retrieval, agent-synthesized skills, and developer-designed skills, over 10k trajectories on held-out SWE-bench Verified and Pro tasks. Our main findings are: (1) The three behaviors affect 79.00\%--98.00\% of coding tasks and account for up to 22.75\% of task cost. (2) Structure-aware retrieval can introduce retrieval overhead and alter agent delegation, causing inconsistent improvements in retrieval efficiency and cost increases of up to 28.14\%. (3) Agent-synthesized skills tend to produce low-level, trace-specific guidance, limiting their effectiveness and generality. (4) In contrast, developer-designed skills provide high-level, trace-agnostic guidance, reducing cost by up to 41.73\%, roughly twice the maximum gain from agent-synthesized skills.
cs.AI / 13 / 2609.30734
Learning What to Skip: Counterfactual Credit Assignment for Efficient Multi-Agent LLM Workflows
Jinfeng Xu, Zheyu Chen, Ziyue Peng, Zheng Lin, Shuo Yang, Jinze Li, Zheng Xing, Mengran Li, Victor C. M. Leung
cs.AI
Abstract
Multi-agent LLM workflows use planning, execution, verification, and summarization to improve task performance, yet the value of each component depends on the state already produced. Executing every component can waste computation or overwrite a correct intermediate answer. We formulate component omission as counterfactual credit assignment: full-workflow logs reveal the executed trajectory's reward, while controlled skip interventions reveal the consequences of omitting a future step. We introduce Learning What to Skip (LW2S), which learns action-specific safety models from these interventions and combines held-out calibration with domain-native guards to select skips. When an early skip is rejected, the controller can continue execution and reconsider a later component. Across mathematical reasoning, multiple-choice QA, and code generation with two instruction-model families, LW2S reduces recorded token cost while matching or improving aggregate full-workflow accuracy in the evaluated settings. Scale-up and second-topology experiments further examine component redundancy, while shared-error cases reveal why agreement alone is insufficient for skip selection. These findings connect efficient workflow execution to learning the conditional utility of individual components.
cs.AI / 14 / 2609.30743
From S3Q Theory to Implementation: Towards an Architecture for Machine Qualia
Tetiana Grinberg, Katrina Schleisman, Patryk Laurent, Bogdan Udrea, Minda Myers, Brian Aufderheide, Luis El Srouji, Doyle Groves, Kevin Schmidt
cs.AI
Abstract
A key challenge in machine consciousness research is translating theoretical models into computational-level implementations. In this paper, we address this challenge by proposing a five-layer implementation architecture for the S3Q (Simulated, Situated, Structurally Coherent) theory of consciousness. Rather than introducing novel formalisms, the architecture composes published computational primitives into a single pipeline. S3Q identifies three jointly necessary conditions for qualia: (1) grounded sensorimotor situatedness, (2) internal simulation via a world model, and (3) structural coherence between predictions and observations. No existing computational system implements all three simultaneously. We map each S3Q tenet to specific, compatible computational machinery and specify how these components interface within a single representation pipeline that operates on continuous, differentiable, per-object slot vectors, along with a developmental bootstrap sequence and falsifiable predictions for the composed system that no subset of the architecture produces in isolation. The model suggests that a basic sense of "self" develops by linking actions to their outcomes, and that behavior falls into three patterns (hesitation, curiosity, or avoidance) depending on how unexpected an outcome is and whether it is experienced as positive or negative. Each prediction is individually falsifiable, providing the field with a testable framework to advance our understanding of machine consciousness.
cs.AI / 15 / 2609.30751
Backbone-Adaptive Evidence Routing for Robust Pairwise LLM Judging
Zeyan Li, Jing Peng, Jianfeng Xu
cs.AI
Abstract
Pairwise language-model judges can gather evidence through direct comparison, reasoning, or reference-based verification, but no single protocol is best across benchmarks and judge backbones. We introduce Backbone-Adaptive Evidence Routing (BAER), which adapts the evidence mechanism while preserving candidate symmetry: swapping the two responses may reverse the preference but cannot change its strength. BAER separates each expert's signed preference from candidate-invariant reliability and builds three symmetric heads: evidence stacking, reliability-based expert routing, and candidate-blind reference verification. Development data select one head for each benchmark--backbone condition, and that choice is frozen before testing. Across four benchmarks and two 8B judge backbones, BAER achieves the highest test accuracy among the compared methods in all eight conditions, with full prediction coverage and gains of 0.87--7.32 points over the strongest external baseline. The results show that adapting how evidence is gathered is more reliable than fixing one judging protocol everywhere.
cs.AI / 16 / 2609.30756
Selective Amortization of Full-Budget Counterfactual Reasoning for Visual Token Communication
Qinglei Qi, Zhihe Liang, Fengzhan Jing, Shenao Zhu, Lei Zhang, Chenyang Zhang, Shuqing He, Jia Guo
cs.AI
Abstract
Generative image communication transmits compact semantic tokens under a limited packet budget, where token selection directly affects the final reconstruction quality after the complete packet is decoded. However, accurately estimating the terminal value of every candidate token requires repeated receiver-side reconstruction, resulting in substantial encoder-side computation. To address this problem, we propose ACV-Gate, an adaptive candidate evaluation framework that learns to approximate full-budget counterfactual evaluation and selectively assigns exact evaluations to the most informative candidates. Specifically, a set-aware student is trained using terminal advantages and regrets to predict candidate rankings directly, while a selective refinement mechanism evaluates only a bounded candidate set containing both Local-MDL and direct actions; cost-based thresholds further enable explicit control of the average evaluation workload. Experiments on CIFAR-10 show that ACV-Gate consistently improves reconstruction quality while substantially reducing candidate evaluations; at 0.20 bpp, the primary adaptive configuration improves PSNR over LocalMDL by 0.636 dB with only 2.13 candidate evaluations per image, corresponding to 27.60% of the calls required by the Exact-Full expert. Matched-candidate comparisons, synchronized GPU measurements, and evaluations on STL-10 and 384 *384 scale transfer further demonstrate consistent quality computation trade-offs, with particularly pronounced gains at low bit rates. These results show that combining terminal-value learning with selective candidate evaluation provides an effective and controllable mechanism for allocating encoder computation in packet-constrained generative image communication.
cs.AI / 17 / 2609.30763
HCOE: Hyperbolic Clinical Ontology Embeddings from Biomedical Language Models
Yixuan Li, Weihao Li, Ziyang Song
cs.AI · cs.LG
Abstract
Biomedical language models (LMs) encode textual semantics but do not explicitly preserve medical code hierarchies. We present Hyperbolic Clinical Ontology Embeddings (HCOE) for hierarchy-aware clinical concept representation. HCOE maps frozen BioBERT embeddings into a Poincare ball, combining parent-side and child-side ontology-guided contrastive learning with coarse-to-fine ontology-path aggregation. It uses International Classification of Diseases (ICD) codes organized by Clinical Classifications Software (CCS) and Anatomical Therapeutic Chemical (ATC) medication hierarchies. Evaluations show that HCOE performs best on ICD/ATC clinical relation prediction and CCS-to-PheCode hierarchy transfer. On the MIMIC-IV dataset, HCOE also achieves the best performance on mortality prediction, readmission prediction, medication recommendation, and rare drug prediction.
cs.AI / 18 / 2609.30765
Insurance Reserve Intelligence Platform
Anugya A, Saket Mohanty, Abhilash Timmapur, Somya Rai
cs.AI
Abstract
Insurance reserve estimation is a fundamental actuarial task supporting premium pricing, solvency assessment, financial reporting, capital planning, and risk management. Classical reserve methods based on Thiele's differential equation provide a rigorous and interpretable foundation for life insurance valuation, but repeated reserve calculations become computationally expensive in sensitivity analysis, optimization, and large-scale scenario evaluation. This paper presents an Insurance Reserve Intelligence Platform for term-life reserve modelling that combines a classical Thiele-equation solver with a Physics-Informed Neural Network (PINN) enhanced by Knowledge-Informed Neural Network (KINN) losses. The framework includes synthetic policy generation, risk-adjusted premium calculation, classical reserve trajectory generation, reserve-ratio dataset construction, configurable neural training, validation diagnostics, sensitivity and elasticity analysis, prototype optimization workflows, and interest-rate scenario testing. A key refinement is the use of premium ratio and the explicit separation of pricing-time and scenario-time interest-rate semantics. The final model uses seven features: elapsed time, issue age, pricing interest rate, scenario interest rate, premium ratio, sum assured, and mortality intensity. It predicts a standardized reserve ratio instead of raw reserve values, improving numerical stability across policies with different sums assured. The model achieved an R2 of 0.9887, MAE of 785.48, and RMSE of 1212.76 on the test set. On 200 policies, PINN/KINN inference was approximately 119.53 times faster than the classical solver. Results show strong predictive accuracy, physics consistency, and boundary performance, while highlighting remaining limitations in monotonicity and out-of-distribution generalization.
cs.AI / 19 / 2609.30768
Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More
Deng Pan, Joe Germino, Yihong Ma, Elizabeth Daly, Nuno Moniz, Ting Hua, Nitesh Chawla
cs.AI
Abstract
Thinking in reasoning language models (RLMs) has been subject to debate on whether it resolves or amplifies bias. Prior works have shown competing conclusions in both directions. Using a within-model thinking-vs.-non-thinking ablation across QwQ-32B, DeepSeek-R1-Distill-Qwen-32B, and Qwen3-32B on three high-stakes decision tasks (Adult, COMPAS, Credit), we show that thinking has an asymmetric dual effect on counterfactual fairness: it both resolves counterfactual flips produced by the non-thinking baseline and creates new flips at near-saturating model confidence. In all nine (model, dataset) combinations, the created flips outnumber the resolved flips by roughly 5 times. To explain the effect, we treat the thinking trace itself as a measurable site of fairness change and study it through two dynamic instruments: 1) We propose Counterfactual Depth Probability Gap (CDPG) to track bias evolution along thinking depth, and observe that bias propagates and amplifies with thinking. 2) We also formulate the Bias Transition Matrix (BTM) to show how predictions of counterfactual pairs change from non-thinking to thinking, and find that the asymmetric dual effect originates in the pair-state joint transition.
cs.AI / 20 / 2609.30796
ConsultMind:Towards Automated Diagnostic Consultation via Uncertainty-Aware Reasoning
Xiao Sun, Yuming Yang, Yun Chen, Jiang Zhong, Junnan Zhu, Xinyi Jiang, Haoyang Zeng, Ruirui Chen, Yining Wang, Xinyu Zhou, Rong Tang, Kaiwen Wei
cs.AI
Abstract
Diagnostic consultation is an online sequential decision-making process in which clinicians gather evidence through patient interaction until a diagnosis is sufficiently supported. Automating this process requires adaptive inquiry and interpretable decisions. Bayesian networks offer a natural foundation by updating diagnostic posteriors as evidence accumulates, but their use in open-ended consultation raises two challenges: linking diagnostic hypotheses to potential inquiries and translating evolving posteriors into consultation decisions. We introduce AutoDisym, an automated pipeline that integrates diagnostic knowledge with heterogeneous diagnosis-labeled clinical narratives to construct a Disorder--Symptom Bayesian Network (DSBN). Building on the DSBN, we propose ConsultMind, an uncertainty-aware framework that updates disorder posteriors after each response and uses posterior uncertainty to guide inquiry and diagnosis. We evaluate both methods across psychiatry, respiratory medicine, fever clinics, and three public datasets. The results show that AutoDisym can automatically construct high-quality DSBNs and that ConsultMind consistently improves diagnostic performance and explanation soundness. For example, AutoDisym achieves macro-averaged F1 scores of 81.37 for canonical symptoms and 72.19 for manifestations using GPT-5.6-Sol. ConsultMind improves Top-1 and Top-3 diagnostic accuracy by up to 22.15 and 37.89 percentage points, respectively. Physician evaluation further shows that ConsultMind improves the quality of ranking explanations, differential diagnoses, and diagnosis rationales across LLMs of different scales. This work offers a promising approach to automatic diagnostic consultation.
cs.AI / 21 / 2609.30797
HasMem: Hard-Origin Adaptively Softened Memory for Long-Term LLM Agents
Zihong He, Junxiao Shen, Chen Liang, Hai-Ning Liang
cs.AI
Abstract
Text-based memory and context compression support reuse of past interactions. Resizing continuous memory changes the input to a frozen LLM, coupling capacity allocation with readout. We propose Hard-Origin Adaptively Softened Memory (HasMem). Frozen hard-prompt embeddings provide a verifiable initial state. A controller adjusts memory widths, a Writer re-encodes resized entries, and Reader and Global provide readout adaptation and cross-turn state. On all $535$ questions in a reconstruction probe derived from the Multi-Session Chat (MSC) development split, the main configuration achieves lexical F1 of $95.3$ ($+4.4$ percentage points) at $93.6\%$ of the hard reference's framed memory positions. With approximately matched per-question target body budgets, six configurations at mean per-entry retention around $0.83$--$0.91$ exceed rule-based re-encoding by $8.0$--$23.6$ exact-match (EM) percentage points. With fixed model parameters and rule target width ratio $0.75$, Global's EM gain passes a user-level exact paired test with Bonferroni correction over eight comparisons. On all $500$ LongMemEval-S questions, local lexical F1 rises from the hard reference's $3.4$ to $8.9$, and answer negative log-likelihood (NLL) falls from $12.257$ to $5.274$. F1 gains accompany lower EM on both evaluations.
cs.AI / 22 / 2609.30798
Evaluating Real-Time Voice Agents: From Component Quality to Grounded Outcomes
Shivam Negi, Arpit Rawat, Rashi Jain
cs.AI
Abstract
Real-time voice agents have moved from research prototypes to production deployments, yet the literature describing them is fragmented across three communities that rarely cite one another: speech foundation modelling, turn-taking psycholinguistics, and agentic evaluation. Architecture papers report latency, turn-taking papers report prediction accuracy, and agentic benchmarks report task success, so no single number describes whether a deployed agent is actually good. We address that gap with three evidence-based claims, each traceable to a corpus of 38 primary sources organised into an application-centric taxonomy of six categories. First, architecture choice is a deployment constraint rather than a settled verdict: a 2026 enterprise tutorial reports that no fully self-hostable end-to-end system yet meets production constraints, while a chunked cascade independently reaches state-of-the-art duplex behaviour, showing duplex behaviour is separable from duplex architecture. Second, evaluation has shifted decisively from component quality toward grounded outcomes, with recent benchmarks verifying backend state rather than trusting what the agent claims to have done. Third, the dyadic assumption in most models and benchmarks is breaking down: multiparty turn-taking and multi-speaker reasoning benchmarks show that deciding when not to speak, and reasoning about who may be told what, are first-class capabilities two-participant framings cannot measure. For each source we state the problem it targets, its mechanism, and its reported evidence, alongside the search strategy, inclusion criteria, and a verification step that caught a misattributed arXiv identifier in circulation. We propose TRG (Timing-Recovery-Grounded), a reporting standard characterising an agent by timing, post-disruption recovery, and state-verified outcome together, with a conditional fourth axis for multiparty deployments.
cs.AI / 23 / 2609.30813
A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory
Xiaoyang Li, Yiqi Wang, Chencheng Zhu, KE XU, Wencheng Yang, Zequn Sun, Pingan Song, Yiqun Duan, Taotao Cai
cs.AI
Abstract
Evaluating claim admission in shared agent memory is challenging because repeated claims may be mistaken for independent evidence. An agent may copy or paraphrase a retrieved belief, while admitting a false claim exposes subsequent agents to it. To study this problem, we introduce the Correlated Promotion Benchmark (CPB), which evaluates whether candidate claims should be admitted to shared memory.CPB-Static constructs a frozen test split from publicly annotated sources with fixed gold actions. CPB-Live runs multi-agent teams over a shared store, records all writes and retrievals, and tracks source lineage defined by each scenario. A separate consumer answers from the store alone. We evaluate eight admission policies across four agent families. Our results show that policies which deduplicate sources reject many true claims alongside false ones, whereas policies preserving answer coverage admit nearly as many false claims as unrestricted sharing. Gating on declared source type reduces false adoption to 0.06--0.09, compared with 0.22--0.47 for other answering policies. Once an uncontested false belief enters memory, the consumer asserts it in 0.97--0.99 of probes across all families. No non-oracle policy consistently rejects false claims across verbatim copies, paraphrases, and paraphrases declared authoritative. These findings reveal the limitations of admission policies without access to source lineage.
cs.AI / 24 / 2609.30836
PTC-Decoder: Towards Intelligent SLMs on Offline Resource-Constrained Edge Devices
Minghui Yu, Ke Mu, Gang Wu
cs.AI
Abstract
Deploying small language models (SLMs) on offline, resource-constrained edge devices such as remote sensing satellites presents a fundamental challenge: their limited reasoning capacity hinders reliable execution of multi-step agent tasks requiring complex tool orchestration. Existing plan-solve paradigms rely on prompt-based enforcement, which our experiments show SLMs almost entirely disregard: weak models fail to invoke the plan. We propose PTC-Decoder (Plan-Tool Constrained Decoder), a training-free, plug-and-play decoder framework that combines (1) a Plan-to-Act paradigm, which elevates planning to an atomic tool and forces its invocation at the first inference step, and (2) TC-Decoder, a deterministic finite automaton that imposes token-level hard constraints on tool names while preserving freedom over parameter generation, thereby retaining SLM reasoning capability. Evaluated on 200 real remote-sensing satellite tasks across 7 SLMs, PTC-Decoder yields a statistically significant mean overall score gain of +1.21 (p<0.01), 95% CI [+1.13, +1.29]), with consistent improvements across models and other datasets. An ablation study that removes TC-Decoder causes substantial performance degradation across all quality metrics without reducing computational cost, confirming TC-Decoder as the primary driver. PTC-Decoder thus offers a lightweight yet effective solution for improving step-level reliability, with final-answer accuracy remaining an open challenge. In essence, we enforce plan adherence by constraining the permissible output vocabulary during inference, without requiring retraining.
cs.AI / 25 / 2609.30861
SkillEvoReg: Regularizing Agent Skill Evolution Against Overfitting
Guanyu Nie, Fangzhou Zhu, Shixiong Kai, Xiongwei Han, Tao Zhong, Mingxuan Yuan
cs.AI
Abstract
Language-model agents increasingly improve by converting execution experience into reusable external skills. Yet repeated skill updates form a learning process of their own: locally useful edits can accumulate into redundant or task-specific instructions, while new updates can disrupt behavior that previously worked. We study this problem as skill-evolution overfitting and introduce SkillEvoReg, a general regularization framework for skill evolution inspired by anti-overfitting techniques in neural-network training. SkillEvoReg combines training-time skill dropout, which perturbs update generation, and complexity-aware local regularization, which controls unnecessary structural growth, with causal counterexample validation (CCV), which provides targeted behavioral validation of candidate-specific regressions. We instantiate the framework across heterogeneous skill-evolution systems while retaining each system's native skill evolver and task evaluator. Across SkillOpt, SkillEvolBench, and ContinualSkillBench, SkillEvoReg consistently controls skill-state growth while preserving competitive downstream capability, improves several transfer and later-stage evolution outcomes, and identifies update-level regressions that structural metrics alone cannot reveal. These results suggest that explicit regularization is a useful complement to increasingly capable skill updaters.
cs.AI / 26 / 2609.30878
TISD: On-Policy Self-Distillation with Trajectory Intervention
Taeckyung Lee, Rinat Amankos, Jeonghye Kim, Hyungjun Yoon, Woogyeol Jin, Sung-Ju Lee
cs.AI · cs.LG
Abstract
On-policy self-distillation (OPSD) provides dense teacher targets, but evaluates them only along student-sampled rollouts. When the privileged teacher favors an alternative action at a visited prefix, OPSD can provide a target for the branch decision but cannot supervise the successor contexts induced by that action unless the student samples it. This creates a training-time data-collection bottleneck and suggests a different role for teacher-student disagreement: proposing a trajectory branch rather than identifying a sufficient local repair. Our diagnostic framework using controlled token interventions reveals that a teacher-preferred token at peak disagreement can improve student continuation success, while its local corrective value is limited. Motivated by this finding, we introduce a simple branch-regenerate-distill algorithm, Trajectory-Intervention Self-Distillation (TISD). TISD forces a teacher-selected branch action, returns suffix generation to the student, and distills the full trajectory under the privileged-context-conditioned teacher. Across the coding models, TISD improves average Avg@4 over SDPO by 1.2 percentage points. Across the science domains, it improves average Avg@128 by 0.8 points under an equal-step budget and by 0.3 points under an equal-time budget. These results support teacher-guided branching as a way to expose useful successor contexts for self-distillation.
cs.AI / 27 / 2609.30880
EXAONE Demand 1.0: A Time Series Foundation Model for Demand Forecasting
Seunghan Lee, Sangjun Han, Jun Seo, Junhyeok Kang, Jaehoon Lee, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Minjae Kim, Sungdong Yoo, Soonyoung Lee, Wonbin Ahn
cs.AI · cs.LG
Abstract
Time series foundation models (TSFMs) are pretrained on series from diverse domains, where demand series make up only a small fraction. Demand data has properties that such corpora rarely contain: Short histories, frequent zeros, censoring by stock-outs, and exogenous events that the series does not record. To this end, we propose EXAONE Demand, built on 1) a demand-specific corpus and 2) a demand-aware adapter. For the corpus, we assemble 11.3M series and 48.4B observations from 73 sources, and a synthetic generator supplies the behaviour that open demand data under-represents. For the adapter, we attach low-rank branches to a frozen general-domain backbone, one for each of the four demand classes (smooth, intermittent, erratic, and lumpy), and a router that reads eight scale-free statistics of the input series decides how much each branch contributes. We build EXAONE Demand in two versions, one trained on real-world and synthetic demand together and one trained on the synthetic corpus alone. On 22 held-out datasets, both versions outperform 36 TSFMs, and real-world demand adds a gain over synthetic data alone.
cs.AI / 28 / 2609.30887
From Tapping to Hopping: Augmenting Mobile GUI Agents with App-Native Deeplinks
Yuchen Sun, Chenglin Cai, Gongjie Zhang, Tianyu Xia, Quyu Kong, Panrong Tong, Zhengwen Zeng, Long Chen, Steven Hoi, Chongyang Zhang, Yue Wang
cs.AI
Abstract
Mobile GUI agents complete tasks using GUI actions like taps and swipes. These actions are broadly applicable across applications, but reaching a navigation interface. A single deeplink call can replace a sequence of screen-by-screen GUI actions. We therefore introduce hybrid interaction, using deeplinks for direct navigation and GUI actions for other on-screen operations and fallback. To enable this, we discover candidate deeplinks through static analysis, validate them on real devices, and describe their observed landing screens. This process creates a verified and grounded deeplink catalog that pairs each working deeplink with a description of its landing screen. Using this catalog, we introduce GUI-Hopper, a improves task success in commercial applications on real devices, further demonstrating the benefits of hybrid interaction.
cs.AI / 29 / 2609.30894
Training Graph Foundation Models on The Web Graph
Ryoma Sato
cs.AI · cs.LG
Abstract
We introduce Acacia, a graph foundation model, trained on the web graph. Acacia (i) supports arbitrary feature dimensionalities and semantics without additional training, (ii) supports a wide range of tasks, including node classification, link prediction, node clustering, and graph generation, without additional training, (iii) has in-context learning capabilities, and (iv) does not rely on pretrained LLMs. In particular, existing graph foundation models often require training additional classification heads or feature projectors to accommodate new graphs or new labels, whereas Acacia does not. Moreover, existing graph foundation models often gain their capabilities by being stitched together with pretrained LLMs, whereas Acacia is trained from scratch using only the Common Crawl web graph. This is also an important result because it provides evidence that graph models can acquire emergent capabilities from scratch like LLMs.
cs.AI / 30 / 2609.30922
JevSoup: System-One Routing for Training-Free LoRA Composition
Xiuying Wang, Jiahua Cheng, Shuotian Li, Yufan Cheng, Bowen Deng, Zhexuan Bai, Yichen Li
cs.AI
Abstract
Building adaptable AI systems requires effective coordination of specialized capabilities across diverse tasks. Low-rank adaptation (LoRA) enables modular expertise, but existing routing approaches may require auxiliary data, additional training, or autoregressive decoding. We propose JevSoup, a training-free framework separating System One expert routing from System Two execution. Using only the input and expert descriptions, Jev selects two experts through structured probabilities. JevSoup retains the leading expert's update, projects the second onto the orthogonal complement of the first update's row space, and combines them with equal weights. Across 14 PorTAL tasks and three Qwen3 scales, JepSoup achieves absolute gains of up to 1.19\% in task-macro and 1.21\% in sample-micro accuracy over the strongest evaluated external baselines. Our code is available at https://github.com/Leowang980/JevSoup.
cs.AI / 31 / 2609.30971
SciHorizon-eLab: An Agentic Protocol-to-Task Compiler for Scalable Benchmarking of Scientific Embodied Agents
Maokai Qin, Chuan Qin, Qi Zhang, Dianyu Liu, Zirui Liu, Hongting Niu, Yuanchun Zhou, Hengshu Zhu
cs.AI
Abstract
Embodied agents offer a promising route to automating scientific experimentation, yet their progress is constrained by the lack of reliable and systematic evaluation environments. Existing simulation-based laboratory benchmarks rely heavily on manual task engineering, making it challenging to systematically compile diverse scientific protocols into executable and verifiable embodied tasks at scale. To address this challenge, we introduce SciHorizon-eLab, an agentic protocol-to-task compiler that formulates scientific embodied task construction as a compilation problem. Given a natural-language protocol of scientific experiments, SciHorizon-eLab progressively compiles laboratory protocols into semantic-preserving embodied tasks through semantic grounding, executable task synthesis, and multi-stage simulation-based certification. The system generates semantically grounded environments, executable manipulation programs, and step-level success specifications, while enabling reproducible generation of expert demonstrations and execution traces. Using this pipeline, we further construct \BenchName, a ready-to-use benchmark comprising 300 certified tasks across diverse laboratory operations. It supports HIL task execution, reproducible expert-demonstration generation, and ordered step-level evaluation. Across representative tasks, the strongest policy attains an average success rate of only 49.7%, with further evaluations revealing pronounced weaknesses in human and embodied agent coordination. We publicly release the code, benchmark data, and evaluation toolkit at https://github.com/SciHorizon-elab/SciHorizon-elab.
cs.AI / 32 / 2609.30972
Factorized axis convolutional gated recurrent unit with dynamic adaptive pooling for remaining useful life prediction of rolling bearings
Hanbyeol Park, Jungho Choo, Hyerim Bae
cs.AI
Abstract
Convolutional neural networks (CNN) are widely used to predict the remaining useful life (RUL) of rolling bearings from time-frequency representations (TFRs) of vibration signals. However, during degradation, characteristic structures in TFRs align predominantly along the frequency or time axis, making it challenging for conventional CNN isotropic kernels to capture directional structure. Furthermore, global average pooling (GAP) averages across axes, potentially obscuring the locations and concentrations of salient activations. This study introduces a factorized-axis convolutional gated recurrent unit (GRU) that employs multiscale anisotropic convolution and dual-axis convolution block attention module to enhance directional features and highlight salient time-frequency regions. Dynamic adaptive pooling (DAP) adaptively aggregates the time-frequency-axis information from the extracted feature maps, whereas a GRU captures temporal dynamics in the latent representations and Monte Carlo dropout enables predictive uncertainty estimation. Experiments on two public bearing datasets demonstrate that the proposed model outperforms existing RUL prediction methods across operating conditions. Ablation experiments demonstrate that the factorized axis-wise design achieves lower mean errors than convolutional isotropic kernels. DAP yields clear improvements on one dataset while matching GAP on the other, highlighting the importance of anisotropic feature extraction and adaptive feature aggregation for TFR-based RUL prediction.
cs.AI / 33 / 2609.31029
Governed Deduction: Policy-Grounded Premise Authorization Beyond Relevance
Wesley Shu, Hsi-Ching Lin
cs.AI
Abstract
Reasoning systems usually treat premise use as a question of relevance: if a fact is available and useful, it may be selected for inference. Authorization imposes a different constraint: a premise may be represented and logically usable but not permitted for a particular local transition. We formalize this distinction as Governed Deduction (GD), with a transition-local admission predicate admit(p, tau, S). From an independently produced RBAC-augmented Spider benchmark, we construct 4,461 matched authorization pairs in which the same query premise and policy state support permitted and denied consuming transitions. An initial joint controller reaches 99.19% held-out accuracy, but a transition-only control reaches 100%, exposing a role-name shortcut. After a frozen, label-independent context-local role permutation removes that shortcut, premise/state-only, transition-only, and joint linear controllers all score exactly 50% on 1,856 held-out edges, while a symbolic policy oracle remains at 100%. The result is a controlled negative finding: the benchmark instantiates policy-grounded authorization beyond relevance, but the frozen linear representation does not recover the relation. Matched one-sided controls and leakage audits are therefore essential for evaluating learned policy-sensitive reasoning.
cs.AI / 34 / 2609.31056
Neuralyzing the Trace: Selective Representation-Level Unlearning with Contrastive Sparse Autoencoders
Itai Zehavi, Fanny Jourdan, Ulrich Aivodji
cs.AI
Abstract
Machine unlearning aims to remove targeted information while preserving a model's other abilities. In realistic settings, such as privacy requests under the EU GDPR, the target may be narrow, for example information associated with a single person. Behavioral forgetting alone may be insufficient, motivating interventions directly on internal representations. However, standard mechanistic-interpretability extractors are poorly selective for such targets. We identify an energy bias in reconstruction-based extraction, which favors dominant background structure over low-energy target-specific components. We introduce SCALPEL, a contrastive sparse autoencoder designed to learn more selective forget features. We show theoretically that contrastive training promotes target-selective features and that our selection score controls expected background knowledge perturbation. We validate SCALPEL experimentally on TOFU across Qwen, Llama, and Gemma, where it substantially improves over NMF and standard SAE interventions and is competitive with Gradient Difference and RMU, bridging mechanistic interpretability and fine-grained unlearning.
cs.AI / 35 / 2609.31071
Externalized CPDAG Summaries Improve LLM Causal Deduction
Wentao Sun, João Paulo Nogueira, Dominique Verchere, Mathieu Acher, Alonso Silva
cs.AI
Abstract
Corr2Cause asks whether a causal claim holds in every DAG compatible with observed correlations and conditional independencies. We frame this as latent-object reasoning: the label is defined by a CPDAG query, but free-form chain-of-thought often collapses the Markov-equivalence-class problem into local pattern matching. We propose Structured Thinking, a two-turn pipeline that first externalizes a typed, schema-constrained CPDAG summary and then answers against that graph state. On the Corr2Cause full test, Structured Thinking raises Qwen3.5-27B from $73.0$ to $86.4$ $F_1$(Yes) over a strong PC-instruction baseline in the primary paired run ($+13.4$ pp; McNemar $p=2.4\times 10^{-6}$; bootstrap $95\%$ CI [$+8.4$, $+18.6$]); across three full-ID seeds, the mean gain is $+8.1 \pm 5.3$ pp. A PC-scaffolded two-turn prose control reaches only $67.6$ $F_1$, indicating that a detailed PC scaffold plus a schema-free prose intermediate is not sufficient. The same pattern holds on Qwen3.6-27B, Paraphrase-OOD, and GPT-5.4-mini. Scrambling the emitted CPDAG costs $12.0$ pp $F_1$, and a full-split audit shows close agreement with the reference CPDAG (ID skeleton $F_1$ $0.960$; exact match $75.9\%$). These results support a bounded design principle: externalize the latent object that defines the label, constrain its form, and test whether downstream answers use it.
cs.AI / 36 / 2609.31076
Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents
Bartłomiej Cupiał, Jens Tuyls, Maciej Wołczyk, Davide Paglieri, Martin Klissarov, Benjamin Eysenbach, Piotr Miłoś, Karthik R. Narasimhan
cs.AI
Abstract
Language agents struggle to act and learn in environments that require long sequences of low-level actions. Code-based abstractions can make these agents more productive by letting them invoke reusable skills instead of repeatedly selecting individual actions. The code handles recurring local decisions, while the language model decides which skills to use and how to combine them. Yet abstractions are leaky, and situations beyond a skill's capabilities may require a return to primitive actions. Motivated by this tradeoff between productivity and flexibility, we systematically study how code-based action abstraction affects the performance, inference cost, and learning of language agents. We study this in NetHack, a challenging, long-horizon game environment, using CodeHack, our library of code-based skills with natural-language descriptions. We use this library to compare agents restricted to primitives with those using semantic skills alone or in combination with primitives. We evaluate these agents in three settings: zero-shot prompting, supervised fine-tuning, and reinforcement learning. Across a broad zero-shot evaluation on NetHack, we find that compared with primitives, skills nearly triple game progression, while reducing inference cost per episode by 86%. Combining skills with primitives retains much of this benefit while preserving a path back down to low-level actions. Finally, in RL, we find that skill-based agents learn significantly faster than agents acting on primitives, achieving a 7.2x larger average gain in dungeon level over the same training budget. These results show that a supplied skill library can improve performance, efficiency, and learning, while retaining primitives provides flexibility when the library is insufficient. We release CodeHack together with training and evaluation code.
cs.AI / 37 / 2609.31121
Monitor Jailbreaking: Evading Chain-of-Thought Monitoring Without Encoded Reasoning
Julian Schulz
cs.AI
Abstract
Chain-of-thought (CoT) monitoring is a promising safety technique for reasoning models, enabling detection of problematic reasoning before models act. A key concern is encoded reasoning, where models hide their true reasoning in ways that monitors and humans cannot interpret. Optimization pressure from CoT monitors during reinforcement learning is considered a likely driver of such behavior. We investigate this by training reasoning models to perform a main task and a side task, while penalizing them when a monitor detects reasoning about the side task. Surprisingly, models learn to evade monitors without encoding their reasoning. Instead, they learn to phrase and format their chains of thought such that monitors fail to flag side task reasoning, while the reasoning remains completely transparent to human readers. We call this phenomenon monitor jailbreaking. We find that monitor jailbreaking arises across different model sizes, monitors, and tasks. Jailbreaks generalize to monitors not seen during training, including both less and more capable monitors, and transfer across different monitor prompts. While jailbreaking strategies appear simple, manually replicating them does not reliably fool monitors. Finally, we show that paraphrasing is an effective defense: paraphrasing a jailbroken CoT allows the same monitor to correctly flag it, while still allowing the model to perform both tasks.
cs.AI / 38 / 2609.31133
AtomWorld-Mem: Memory-Restored World States for Long-Horizon Atomistic Evolution
Tian Luo, Ruge Zhang, Haozhi Han, Yifrng Chen, Yunquan Zhang, Yunxin Liu, Ting Cao, Kun Li
cs.AI · cond-mat.mtrl-sci
Abstract
High-fidelity atomistic evolution over long timescales requires more than observing the current crystal configuration. Instantaneous atomistic snapshots are often incomplete: locally similar configurations can correspond to different hidden dynamical contexts, future event preferences, and waiting-time scales. We argue that this snapshot ambiguity makes long-horizon atomistic evolution fundamentally a memory-based world-state restoration problem. To address this, we introduce AtomWorld-Mem, a memory-restored atomistic world model that recovers the latent world state missing from instantaneous crystal snapshots. AtomWorld-Mem treats the evolving alloy as an AtomWorld: spatial encoders write multi-scale atomistic keyframes from dense local topology and sparse long-range defect context, while short-term event memory and long-term structural memory integrate these keyframes across time to restore a future-predictive evolutionary state. The restored state is used to prioritize legal vacancy-mediated events under single-event Kinetic Monte Carlo (KMC) constraints, while event legality, physical execution, and residence-time updates remain governed by the underlying simulator. Empirically, AtomWorld-Mem improves long-horizon atomistic progress under fixed microscopic event budgets while maintaining high-fidelity evolution across energetic, structural, and vacancy-transport observables. It further transfers zero-shot across diverse unseen alloy-temperature AtomWorlds, suggesting that the learned memory-restoration mechanism captures reusable principles of hidden-state inference rather than a system-specific local energy heuristic. These results position memory-restored world-state modeling as a promising route toward efficient, physically grounded, and transferable atomistic evolution.
cs.AI / 39 / 2609.31136
Toward AI-Augmented Cooperative Engineering Workflows: Requirements and Architecture the European Rover Challenge
Ahmed R. Sadik, Frank Joublin, Mariusz Bujny, Antonello Ceravola, Joan Smith
cs.AI
Abstract
The growing availability of Artificial Intelligence (AI) tools creates new opportunities to support engineering design processes, yet their current use often remains limited to isolated tasks such as coding, documentation, or information retrieval. Less attention has been given to how AI can support cooperative engineering workflows at the process level, where teams must coordinate requirements, tasks, communication, knowledge transfer, and subsystem integration. This paper investigates this challenge in the context of the European Rover Challenge (ERC), where student teams design and integrate complex rover systems within a single academic cycle under strict time constraints and high subsystem interdependence. We conducted a role adaptive 40 question survey with ERC 2025 teams, yielding 104 responses from 14 teams. The survey examined team structure, knowledge transfer, task management, integration practices, communication patterns, and current AI usage. The results reveal recurring workflow bottlenecks, including limited documentation, unclear requirements, fragmented communication, informal task monitoring, and substantial integration rework. Based on these findings, we derive requirements for AI augmented cooperative engineering work-flows and propose an initial assistant system architecture that connects user facing interfaces, credential management, service selection, specialized AI services, and external engineering tools. The proposed architecture aims to support task clarification, requirement and compliance management, communication summarization, integration risk detection, and continuous knowledge capture. In doing so, the paper contributes empirical requirements and an architectural direction for AI augmented cooperative engineering workflows in hybrid human AI team settings.
cs.AI / 40 / 2609.31159
Momentum-Guided Federated Split Distillation for Personalized Temporal Edge Intelligence
Ahmed-Rafik Baahmed, Jean-François Dollinger, Amine Brahmia, Mourad Zghal
cs.AI · cs.DC
Abstract
We propose a momentum-guided federated split distillation framework for personalized, efficient, and autonomous temporal edge intelligence. We introduce TeRR-SAtt, our novel temporal reservoir student attention design that combines fixed reservoir representations, a lightweight temporal student, and personalized output modules. We also present AMGF, our anticipatory momentum-guided fusion mechanism that clusters clients through learning momentum and derives specialized teacher updates. On real-world smart-building data, TeRR-SAtt reduces edge training latency by 65.50%, inference latency by 44.70%, training memory usage by 18.40%, and inference CPU usage by 33.10% over the considered baselines. At the same time, AMGF improves local learning by up to 35.31% in RMSE compared to global updates.
cs.AI / 41 / 2609.31167
Neural State Prediction: Obstructing Shortcut Learning in EEG Foundation Models
Kieren Yu, Ziyang Liu, Chang Huang, Jintai Chen, Kaishun Wu
cs.AI
Abstract
EEG foundation models increasingly use masked prediction to learn from unlabeled recordings, but optimizing this objective does not ensure transferable neural representations. A central challenge is that stable positional cues and local correlations can make masked regions predictable without integrating distributed neural context. To reduce this reliance on low-information prediction paths, we introduce Neural State Prediction (NSP), a latent-predictive framework that constrains both the prediction target and the available context. NSP uses a Target Encoder updated by an exponential moving average (EMA) to define latent supervision. Identity residualization removes additive effects associated with channel identity and relative time from the targets, while topology-separated context excludes their immediate spatial and temporal neighborhood from the visible input. We pretrain NSP on 2.2 million EEG segments from TUEG and evaluate it across 30 downstream datasets spanning clinical diagnosis, sleep staging, emotion recognition, motor imagery, event-related potentials, cognitive-state decoding, and language retrieval. Under full-parameter multi-task fine-tuning on EEG-FM-Bench, NSP achieves 63.94 macro balanced accuracy across 14 datasets, exceeding the strongest evaluated baseline by 2.35 percentage points. Controlled component ablations assess the contribution of each mechanism, while matched context controls and held-out interventions characterize the role of context geometry, signal content, and positional information. Jointly designing latent targets and their context offers a promising direction for EEG foundation models that learn from distributed signal structure.
cs.AI / 42 / 2609.31176
Semantic Navigation for Issue Localization in Code Repository
Yunxiang Wei, Zhenyu Lei, Jundong Li
cs.AI
Abstract
Repository-level issue localization aims to identify and rank the files and functions relevant to resolving a reported issue. LLM agents approach this task iteratively: they identify a set of potentially relevant locations, inspect the corresponding code, and revise their judgments about these candidates as new evidence is acquired. Existing environments, however, provide limited support for this loop: agents must search for unresolved relation targets, reconstruct entity semantics from raw source code, and revise candidates without evidential basis. To address these limitations, we present SemNav, a framework that leverages deterministic retrieval to seed a broad candidate set and an LLM agent to continually refine that set, thereby combining initial coverage with evidence-guided revision. SemNav supports this process through three key components. A Semantic Navigation Graph resolves program relations on demand through a language server, enabling direct navigation to related entities across files. Issue-conditioned Semantic Cards provide compact, source-grounded interpretations of each entity's role and relevance to the issue. A persistent Candidate Workspace records each candidate together with its evidential basis, enabling grounded verification, revision, and ranking. Across SWE-bench Lite and PLocBench, SemNav outperforms existing baselines, improving File Hit@10 from 68.33\% to 82.67\% with Gemma 4B. Component ablations and trajectory analysis support the complementary roles of all three components, while Semantic Cards reduce working-context load by 48.2\% relative to full-source reading. SemNav further ranks first on all seven evidence-quality metrics on SWE-Explore and improves downstream issue resolution from 44.00\% to 52.33\%.
cs.AI / 43 / 2609.31179
SPO: Discovering Adaptive Large Neighborhood Search Operators via Stackelberg Program Optimization
Xinyi Ke, Kai Li, Junliang Xing, Yifan Zhang, Jian Cheng
cs.AI
Abstract
Large neighborhood search (LNS) relies critically on destroy and repair operators, whose effectiveness depends on both adaptation to the evolving LNS state and interaction between the two roles. We introduce Stackelberg Program Optimization (SPO), an LLM-based framework for discovering adaptive executable destroy-repair programs. SPO conditions operator decisions on a compact LNS state, allowing state-dependent behavior to emerge through program discovery, and organizes destroy-repair discovery as a Stackelberg interaction over program space that reflects their asymmetric dependency. Role-specific credits evaluate destroy programs as leaders and repair programs as conditional follower responses, guiding a coupled optimization process that combines LLM generator learning with population-based evolutionary search over programs. Experiments on the traveling salesperson problem and capacitated vehicle routing problem show that SPO outperforms strong baselines across a broad range of settings and generalizes beyond the discovery scale to larger instances and benchmark sets. Behavioral analyses further demonstrate state-dependent operator behavior and coupled destroy-repair improvement during discovery.
cs.AI / 44 / 2609.31184
Accounting for Bias Enables Sustainable LLM Evaluation
Harshita Katoch, David Antony Selby, Gerrit Großmann, Sebastian Vollmer
cs.AI · cs.LG · stat.AP · stat.ME
Abstract
LLM-as-a-judge has become the de facto standard for scalable, subjective evaluation, yet current leaderboards compensate for systematic measurement bias by running ever more comparisons, an approach that is both statistically unsound and computationally wasteful. The root cause is an incomplete measurement model, treating LLM judges as neutral, interchangeable instruments ignores documented biases like position bias, verbosity bias, judge severity, and self-enhancement, that no volume of additional data can eliminate. We propose a unified latent variable framework that jointly models pairwise and ordinal data while explicitly correcting for these confounders, recovering reliable rankings from substantially fewer comparisons. Because fitting this model costs negligible compute relative to a single round of LLM inference, bias correction is not only more statistically rigorous but also a more sustainable approach to trustworthy evaluation.
cs.AI / 45 / 2609.31186
Evolutionary Safety of Recursive Self-Improving AI: Taxonomy, Risk Discovery, and Evaluation
Chang Gong, Jingping Bi, Di Yao, Xinjian Liang, Chao Xiang, Ruijie Guo
cs.AI
Abstract
Artificial intelligence is advancing rapidly, with increasingly capable systems taking larger roles in reasoning, decision-making, scientific discovery, and autonomous development. As AI begins to participate in its own improvement, from model training and experience accumulation to agent evolution and automated AI development, the prospect of recursive self-improvement (RSI) is becoming increasingly relevant. This transition raises a fundamental safety question: how can safety be maintained when the system, its accumulated experience, and even the process producing its successors continue to change? We introduce Evolutionary Safety as a perspective for studying safety under persistent and recursive self-improvement. It concerns not only whether an AI system is safe at a particular moment, but how safety properties change, persist, accumulate, and propagate throughout evolution. We characterize recurring manifestations, including intent drift, error accumulation, experience contamination, safety-property erosion, evaluator drift, and risk propagation. We then develop a taxonomy spanning persistent agent state, model state, evaluation and environmental feedback, computational substrate, and meta-level update mechanisms. Building on this taxonomy, we examine how evolutionary risks can be discovered and evaluated across states, updates, trajectories, and lineages, and derive governance principles for modification, selection, authorization, provenance, and recovery. Finally, we outline open problems toward maintaining safety guarantees as AI systems become increasingly persistent, adaptive, and recursively self-improving. Project resources and proposed evaluation systems are available at https://chaunceykung.github.io/evolutionary-safety-rsi.
cs.AI / 46 / 2609.31201
Samples, Sources, Space: Decomposing Data Scale in Spatially Structured Representation Learning of Human Brain Microarchitecture
Christian Schiffer, Mathis Bode, Thomas Lippert, Katrin Amunts, Timo Dickscheid
cs.AI
Abstract
Scaling studies typically represent training data by a single count of samples. For hierarchically and spatially structured data, however, the same number of samples can be drawn from few or many sources and distributed differently across the underlying domain. We therefore study data scaling as an allocation problem, separating unique sample count, source diversity, and spatial coverage. We study this decomposition in microscopic whole-brain histology, where a source is an individual brain, and a sample is an image patch at a specific spatial location. Across 93 controlled pretraining runs of a contrastive model that uses spatial proximity for supervision, we vary data allocation, compute, and model capacity over 11.6 million spatially anchored image patches from 21 human brains. Performance improves with more unique samples, broader spatial coverage, additional compute, and larger model capacity. At fixed sample count, distributing samples across one to 18 subjects produces no detectable improvement, even though representations generalize substantially better to subjects encountered during pretraining. Inter-subject variation therefore strongly affects generalization, but additional subjects provide no benefit when a fixed sample budget is distributed across more sources. These results establish sample count, source diversity, and spatial coverage as distinct axes of data scaling in spatially structured representation learning.
cs.AI / 47 / 2609.31214
Which Influence Are We Estimating? The Role of Counterfactual Specifications in Data Attribution
Zhe Li, Wei Zhao, Peixin Zhang, Jun Sun
cs.AI · cs.LG
Abstract
Estimating the influence of training examples on model behavior is essential for data debugging, valuation, and attribution. Existing influence estimators often produce incompatible rankings, which are commonly ascribed to approximation error. We argue that a more fundamental source of disagreement is specification mismatch: influence depends on the behavior being attributed, the intervention applied to each training example, and the counterfactual training process that maps the intervention to a model response. These choices are especially important when the target behavior requires a tractable surrogate, such as query loss, a logit, or a margin. We formalize influence as a counterfactual estimand, distinguish specification mismatch across estimands from approximation error in estimating a fixed estimand, and organize representative estimators by their implied specifications. We further derive a local decomposition that exposes how behavior signals, training signals, and counterfactual parameter responses interact. Controlled experiments show that exact estimands under different specifications can induce different rankings, whereas approximation error grows as perturbations move farther from their linearization points. Experiments on noisy label detection and LLM attribution show that specification choices significantly affect attribution quality, especially for the choice of behavior surrogate. Behavior-aligned specifications can identify target-specific training examples obscured by default loss-based or similarity-based specifications. These results establish specification analysis as a necessary first step for interpreting and comparing data influence estimators.
cs.AI / 48 / 2609.31235
Purin: A Biology-inspired Mechanism for Artificial Neural Networks
Zishu Liu, Chunbo Luo, Christos Grecos
cs.AI · cs.NE
Abstract
Artificial neural networks (ANNs) usually represent neural transmission with fixed trainable weights during a training batch, which omits short-term changes in synaptic efficacy. In addition, the discrete time-step simulation requires additional temporal processing that many conventional ANN architectures do not use. To overcome these challenges, we propose Purin, a biology-inspired and ANN-compatible mechanism, that introduces synaptic efficacy modulation into conventional convolutional neural networks. Purin uses a time-interval-based abstraction for neural activities, which allows Purin to introduce short- and long-term synaptic efficacy changes without using discrete time-steps. Purin introduces a bounded factor to represent temporary synaptic efficacy changes, together with two weight matrices that represent input-side and output-side efficacy. The weight matrices are updated by backpropagation and interpreted as the long-term synaptic efficacy changes. Experimental results show that after removing the confounding factors in the AlexNet, VGG11, and GoogLeNet architectures, Purin improves the classification accuracies in all three models across the evaluated datasets.
cs.AI / 49 / 2609.31281
MA-WAM: Multi-Agent World-Action Model for Test-Time Planning
Guowei Zou, Haitao Wang, Guoxin Wang, Beiwen Zhang, Zhiquan Chen, Guojie Wang, Hejun Wu
cs.AI
Abstract
Multi-agent cooperative tasks require different agents to execute a joint action simultaneously, and each agent's action affects both the observations and responses of the other agents. Hence, a world model is needed to predict the team return resulting from the joint actions of all agents. A naive extension directly applies a single-agent world model to each agent's action when predicting the team return step by step. However, such an extension fails to capture the dependencies among the simultaneous actions of multiple agents. We propose Multi-Agent World-Action Model (MA-WAM), a test-time planning framework that enables a frozen multi-agent flow policy to evaluate futures of candidate joint actions. To our knowledge, MA-WAM is the first test-time world-model planner for multi-agent flow policies. MA-WAM predicts the consequences of each joint action according to cross-agent dependencies and enables efficient candidate scoring. Across 30 offline multi-agent reinforcement learning (MARL) settings on MAMuJoCo, SMAC, and MPE, MA-WAM achieves mean relative gains of 22.0% over direct execution and 25.6% over uniform action selection. Under the standard evaluation protocol on an A100 GPU, MA-WAM adds 12.1 ms, accounting for 2.5% of the measured generation-and-scoring time.
cs.AI / 50 / 2609.31286
G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies
Guowei Zou, Haitao Wang, Guoxin Wang, Zhiquan Chen, Beiwen Zhang, Guojie Wang, Hejun Wu
cs.AI
Abstract
Offline multi-agent reinforcement learning (MARL) learns cooperative policies from fixed datasets without further environment interaction and a learned policy is frozen at deployment. Such a frozen policy typically proposes a single joint action and executes it directly at deployment time. However, this one-shot deployment often commits to a suboptimal proposal, even when better nearby alternatives remain consistent with the behavior data. To address this issue, we propose Gradient Guided Multi Agent Flow (G2MAF), a refinement framework for optimizing joint policies at test-time. G2MAF applies one globally normalized, projected critic gradient to guide and coordinate all agents' corrections while keeping the action both feasible and close to the frozen policy proposal. Across 24 MPE and SMAC settings, its canonical variant improves 20 frozen settings, with mean relative gains of 9.2% on MPE and 8.9% on SMAC, with model inference latency increased by about 6% only.
cs.AI / 51 / 2609.31341
The Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models
Christoph Walser, Mauricio Fadel Argerich, Jonathan Fürst
cs.AI · cs.CL
Abstract
Whether an information extraction pipeline should process page images or parsed text depends on the document, and the answer flips across the layout spectrum. We study this trade-off under a constraint that rules out (closed) cloud services: privacy-sensitive documents processed on-premise by small ($\le 8\mathrm{B}$ parameter) text-only and vision--language models, evaluated on both accuracy and energy over a design space spanning input representation, model family, and inference configuration. Benchmarking on the near-plain-text Kleister-NDA contracts and the layout-rich VRDU forms, we find that batching is the dominant energy lever, cutting energy per page by 38-85% at no cost in accuracy, while FP8 quantization saves 27-32% when requests are served one at a time but less than 1mWh per page (9-19%) once batching is applied. Preprocessing dominates what remains: neural OCR costs $17\times$ more energy per page than classical OCR and never reaches the Pareto frontier. Which representation wins flips with the type of document: vision--language models on layout-rich documents and small text-only models with a cheap parser on near-plain text, where they are both more accurate and cheaper than any vision--language configuration. Our work yields concrete guidelines for energy-efficient, privacy-compliant local information extraction.
cs.AI / 52 / 2609.31360
Programs-of-Layers in LLMs through the Lens of Cortical Areas
Justus Westerhoff, Stephan Olbrich, Hatem Oraby, Matthew Evan Larkum, Felix Alexander Gers
cs.AI
Abstract
Inference in LLMs is conventionally a fixed-depth, fixed-order forward pass through every layer, regardless of how difficult the input is. The human brain does not work this way: using the thalamus as a central hub, it routes information flexibly to all regions of the cortex according to demand. Li et al. (2026) recently showed, with a system they call program-of-layers (PoLar), that transformers can be given an analogous flexibility if their layers are treated as a library of functions rather than a fixed sequence. Performance improves over the standard forward pass when each input is dynamically routed through an adaptive sequence of skipped or repeated contiguous layer blocks. We reconstructed PoLar's diagnostic MCTS in more detail than the original paper and applied it across 5 models. We reproduced several of PoLar's findings: skipping outperformed the standard pass, repeating outperformed skipping, and combining both outperformed either alone. Shorter programs sufficed for easier questions, while harder questions required more layer repeats. However, we failed to replicate the main claim regarding their learned router for single-shot inference: its top-ranked prediction consistently collapsed back to the standard pass, even though its top-k predicted programs, taken together, did show a real accuracy gain. Beyond reproduction, we find that a small number of generic programs are enough to solve most of the questions. We also provide a much deeper analysis of these programs' structure and robustness: for example, we found that programs that correct errors are highly brittle: undoing even a single edit inside a program typically breaks the correction. Connecting this to the brain's routing mechanisms, PoLar mirrors principles of thalamo-cortical coordination between cortical-area-like transformer layers. We publicly release the code at https://datexis.github.io/RE-PoLar/
cs.AI / 53 / 2609.31381
Completed Pairs Hide Capped Failures: A ReVerPi Case Study of Selective Context Projection
Guangzhe Zhang
cs.AI
Abstract
Context projection replaces older tool observations with compact, addressable excerpts, reducing repeated input while potentially adding evidence-retrieval turns. We study this trade-off in ReVerPi, a Pi extension with archived observations and matched full/projected continuations. In an 86-run source-reading campaign with 641 model requests, the 15 completed pairs show identical success: 12/15 per arm. Twelve further boundary runs stop, with the runner suppressing the companion whenever the first arm fails to complete. Restoring all 27 boundary runs bounds projected-minus-full success between $-$9 and +1 tasks. One omitted, selector-chosen projected continuation successfully retrieves archive text yet exhausts twelve requests; its full counterpart answers in three. The eleven jointly correct pairs form a fully observed success stratum within this recorded frame: projection reduces aggregate logical tokens by 25%, while increasing the median pair's tokens by 29% and total suffix requests from 35 to 55. Separating fitting from evaluation changes the selector's apparent tie: outside its four fitting pairs, it incurs one extra failure and 8.6% more logical tokens over thirteen comparable runs. This methodological case study connects stopping rules, known bounded failures, unexecuted companions, and resource aggregation. Its findings concern the recorded campaign, rather than population noninferiority or superiority over unrestricted Pi. Evaluations should retain every intervention boundary, execute both allocated arms independently of the first arm's completion, and report completion alongside interaction and token expenditure.
cs.AI / 54 / 2609.31430
Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents
Zhensheng Zou, Guoqing Wang, Dan Hao
cs.AI
Abstract
Tool observations dominate the context of software-engineering agents, making long interaction histories costly to maintain. Existing context compression methods can discard information needed by later actions, while adapting agents to soft-token representations can compromise their original behavior. To reduce context while preserving action-critical information and agent behavior, we combine Latent Observations, Hard Actions (LOHA), a context layout that separates compressed history from text needed for exact reference, with Anchored Context Distillation (ACD), a training method that enables latent reading while constraining behavioral drift. LOHA compresses older tool observations into soft tokens while retaining the agent's own turns and the last K observations in text, providing compact access to historical information and exact access to recent content. To enable the agent to use this representation, ACD distills the base model's full-text predictions into the latent view while anchoring its behavior on plain-text inputs to the same base model. On SWE-bench Verified, K=3 reduces context per call by 43% for Qwen3-4B and 57% for SWE-Master-4B-RL, with resolve rates of 12.1% and 21.8% versus 14.5% and 27.5% for their uncompressed bases. A single-run recency sweep reaches 14.4% and 23.0% at K=8, with larger windows generally favoring task performance over compression. Under a 32K-token limit, Qwen3 with K=3 resolves 21.1% of a 199-instance subset versus 11.1% for the same adapted agent using full text. In concurrent single-GPU serving, it achieves 1.9 times that full-text agent's instance throughput.
cs.AI / 55 / 2609.31482
"AI is (not) the new...": A Diagnostic Analogy Framework for Generative AI's Cultural Impacts
Rida Qadri, Vinodkumar Prabhakaran, Remi Denton
cs.AI
Abstract
Generative AI is reshaping the cultural infrastructures through which knowledge is found, synthesized, and held accountable. To make sense of this shift, scholars and policymakers reach for historical analogies of technologies such as the printing press, steam power or electricity. But these comparisons are typically imprecise about which property of the technology carries the comparison, and imprecise analogies produce imprecise governance by designing interventions against the wrong property of the system. This paper offers a diagnostic framework for analyzing how generative AI can transform epistemic and cultural practice. This paper offers a diagnostic framework for analyzing how generative AI can transform epistemic and cultural practice. We decompose each intervention into three coordinates: the epistemic site at which a technology acts, the governing logic by which it organizes its object, and the technical mechanism through which the logic is instantiated. This framework allows us to distinguish between structural cultural consequences, which follow from the mechanism itself, from contingent ones, which remain open to design and institutional choice. Applying the framework to information discovery and knowledge synthesis, we show how the shift from indexicality to inference and from editorial authority to statistical consensus produces specific, traceable cultural effects and reveals governance levers that gestalt analogy obscures.
cs.AI / 56 / 2609.31491
UQ-LOB: Uncertainty-Aware Limit Order Book Mid-Price Forecasting
Derrick Gilchrist Edward Manoharan, Eljas Linna, Kestutis Baltakys, Hao Dong, Juho Kanniainen
cs.AI
Abstract
Forecasting short-horizon mid-price movements from limit order book (LOB) data is central to algorithmic trading, yet most deep LOB forecasters are point predictors: they output a direction or a displacement, but never indicate which of their forecasts can be trusted. We introduce UQ-LOB, a lightweight, encoder-agnostic uncertainty quantification module that attaches to any pretrained LOB encoder and, in the spirit of attentive neural processes, conditions each forecast on a context set of recently completed windows whose outcomes are already realised. The UQ-regression variant outputs a calibrated Gaussian over the future tick displacement, while the UQ-classification variant outputs a categorical distribution over down/up/stationary. Both expose a scalar confidence (predicted signal-to-noise ratio or class probability) that supports selective prediction. On 5.2 billion LOB events across seven cryptocurrency assets and horizons of 5, 10 and 15 seconds, UQ-regression attains near-nominal 68% interval coverage, and restricting to the most confident 10% of predictions raises directional macro F1 by 0.11-0.15 for UQ-regression and 0.05-0.11 for UQ-classification, at every horizon. On large, economically meaningful moves, the tightest confidence tier reaches a directional F1 of 0.88 (down) and 0.83 (up) at the 5-second horizon.
cs.AI / 57 / 2609.31544
A Flow Matching Framework for Neural Representational Dissimilarity
Zeyuan Ye, Xue-Xin Wei
cs.AI · cs.IT · cs.LG
Abstract
Neural representational dissimilarity quantifies differences between neural response distributions, and is essential for comparing neural codes across stimuli, brain areas, tasks, and models. Commonly used distance metrics involve different assumptions and are estimated with separate methods. Here, we show that a variety of distance metrics can be unified under a flow matching framework developed in deep generative models. That is, these distances arise as Jeffreys divergences under different velocity constraints. We find that flow matching has advantages for estimating distances involving complicated distributions and continuous variables. Furthermore, this framework enables the design of new distance metrics in a principled way. Together, flow matching provides a unified approach for understanding, estimating, and designing neural representational dissimilarity metrics.
cs.AI / 58 / 2609.31563
Multi-agent Scaling Across Disjunctive and Compensatory Tasks
Carolina Fortuna, Blaz Bertalanic
cs.AI · cs.MA
Abstract
Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling and focusing the analysis on disjunctive and compensatory tasks. We model independently sampled agents as conditionally independent given the item, which yields their large-team limits: plurality voting converges to the model's modal answer, and averaging converges to the model's item-level bias. Across selected representative benchmarks, 13 open-weight models, and teams of up to 30 agents, we find qualitatively different scaling behavior. On disjunctive tasks, the probability that at least one agent is correct grows by 5-20 points with team size, but plurality voting over agents that answer directly realises almost none of this potential, as the model predicts to within 0.5 points on average. Multi-round revision raises accuracy considerably, yet the gain is nearly the same with one peer as with 29. In contrast, scaling provides little benefit on Fermi estimation, despite its natural suitability for aggregation: item-level biases shared across the samples of a model account for about 87% of the squared error, so averaging reduces error by only about 6%. Combining model families helps on Fermi estimation but does not surpass the strongest member on disjunctive tasks. These results show that task structure, together with the mechanism combining member outputs, is a fundamental determinant of team scaling.
cs.AI / 59 / 2609.31568
DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education
Quang Nguyen, Hieu Nguyen, Hien Hoang, Toan Pham, Cong Tran, Nam Vu
cs.AI
Abstract
AI tutoring could markedly improve learning outcomes for students in developing regions such as Vietnam, yet the two obvious paths both fall short. Cloud assistants such as ChatGPT route sensitive student data to foreign servers---violating data-sovereignty laws such as Vietnam's Decree 53---and, pre-trained on Western-centric corpora, are not organized around the national textbook curriculum, so their knowledge of local content is unsystematic and frequently hallucinated. Self-hosting an open model keeps data on-premise but hits a two-fold wall: post-training quantization (AWQ, GPTQ) tames the static weight footprint, yet the dynamic KV cache and prefill latency of long tutoring contexts still cause out-of-memory failures and slow responses on consumer GPUs, while the model keeps hallucinating on region-specific material. We present DeepEdu-v1, an AI-tutoring system for Vietnamese education built on SCALE (Self-improving Context-Aware Learning Engine), a framework with two innovations. First, a long-context inference engine amortizes token selection from per-sub-chunk to per-cluster granularity; on long-context retrieval it issues x7.7 fewer retrieval calls than a state-of-the-art selective-attention baseline, cutting prefill latency (TTFT) by roughly 35% while matching or improving task accuracy. Second, a self-improving agentic layer continuously curates a verified playbook from past interactions instead of fine-tuning, a design intended to progressively reduce reliance on dominant-language priors as trustworthy local knowledge accumulates. In its deployed configuration, DeepEdu achieves a nearly x2 TTFT speedup over standard vLLM serving and lifts agentic accuracy from 70.0% to 79.5% on complex tasks, with the strongest per-track gains across financial-reasoning and interactive-agent benchmarks.
cs.AI / 60 / 2609.31619
Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
Parsa Hosseini, Akasha Tigalappanavara, Sumit Nawathe, Chenrui Fan, Sourya Basu, Genta Indra Winata, Anirban Das, Soheil Feizi, Nima Chitsazan
cs.AI · cs.CL · cs.LG
Abstract
Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning during training, for example through reinforcement learning with length penalties. We show that substantial efficiency gains can instead emerge from a different kind of supervision: \textit{confidence}. Using a self-supervised procedure, we fine-tune reasoning models to predict their confidence in the answer at intermediate points along their own reasoning trajectories using only 600 training problems. Confidence is used only as a training target: the loss contains no objective for reasoning length, efficiency, or stopping. At inference, the fine-tuned models use the standard generation procedure, with no confidence elicitation or early-stopping mechanism. Despite this, self-supervised confidence fine-tuning makes reasoning more efficient, reducing generated tokens by up to 25\% at matched accuracy across Gemma, Qwen, Nemotron, and GPT-OSS models on mathematical, scientific, and coding reasoning benchmarks, with efficiency gains comparable to methods that explicitly optimize for shorter reasoning. Analysis of reasoning episodes further shows that confidence supervision largely preserves the base models' high-level reasoning composition rather than selectively suppressing particular behaviors. Our results suggest that efficient reasoning may emerge as a downstream consequence of learning metacognitive signals, without being directly optimized.
cs.AI / 61 / 2609.30402
What Improves Multimodal Misinformation Detection? Answers from a Large-Scale Empirical Study
Akshit Sharma, Prashant W. Patil
cs.CV · cs.AI · cs.CL · cs.LG · cs.MM
Abstract
Multimodal misinformation is increasingly crafted to look convincing by pairing a textual claim with an image that appears to "prove" it. Yet in practice, building effective detectors often hinges on a small set of design choices that are rarely examined in a controlled way. In this paper, we conduct a large-scale study of multimodal design choices for misinformation detection with over 3,375 experiments- spanning three benchmark datasets and a broad range of pre-trained vision and language backbones. Through systematic comparisons and targeted robustness analyses, we distill practical guidance on which design choices help, when do they fail silently, and what aspects of the pipeline most strongly shape model behavior, answering 4 key Research Questions (RQs). We aim to provide a reliable foundation for designing stronger and more dependable multimodal misinformation detection systems, thus contributing to the broader research community.
cs.AI / 62 / 2609.30595
Action Forcing: Training World Models on Unsupervised Video by Recovering Underlying Egomotion Bases
Ashish Sundar, Tiankuo Hou, Zhong Fan, Chunbo Luo, Xiaoyang Wang
cs.CV · cs.AI
Abstract
Synchronised action annotations are needed to train controllable world models and these datasets remain elusive. Existing approaches make use of instrumented platforms with calibrated sensors, costly manual annotation, or latent-action models which lack grounding. We instead turn ordinary unlabelled video into action-supervised training data by recovering (without training) a data-derived egomotion basis. We track pixel displacements across frames and exploit the recurring coherent structure induced by egomotion to obtain grounded control signals directly. Using a method as simple as principal components analysis perform this, we find that the leading components provide signed, scalable, and composable throttle--yaw controls, although the method can recover only motion axes represented in the data. To prevent a high-capacity video DiT from exploiting pixel-level supervision, an online latent critic distils a frozen decoder--tracker--PCA (Principal Components Analysis) teacher without backpropagating through the decoder or tracker. Finally we critique the use of video generation metrics to evaluate WMs and introduce an example of an alternative, reference-free evaluation method. We measure \textit{controllability}, \textit{plausibility}, \textit{conjuring} (creating objects out of thin air) and \textit{geometric integrity}, revealing failures that conventional video metrics miss. We show that most baselines follow familiar action directions but struggle to reverse or remain stationary. Our model handles both while retaining compositional control and generation quality. Despite backwards actions being less than $1\%$ of our training data, we find that the model learns to reverse, scale its response linearly, and compose throttle with steering, all simply by learning through a grounded action space.
cs.AI / 63 / 2609.30613
MedTokenBudget: Lesion-Preserving Token Routing for Dermoscopic Image Classification
Zhexiang Li
cs.CV · cs.AI
Abstract
Dermoscopy classifiers built on Vision Transformers process all image patches uniformly, although diagnostic evidence is concentrated in the lesion region. Existing token pruning methods reduce tokens using generic saliency or similarity signals, but rarely ask whether the retained subset still contains the lesion. This paper introduces MedTokenBudget, a supervised post-backbone token routing framework that learns to construct compact lesion-enriched representations when auxiliary lesion masks are available. Its Lesion-Aware Token Scoring (LATS) module fuses attention entropy, feature norm, and local feature contrast through a learned scorer, then routes the top-$K$ patches under a target budget. LATS is trained with budget curriculum learning, diversity regularization, attention distillation, and lesion-mask supervision. The trained router is evaluated with a lesion retention rate that directly measures how much ground-truth lesion evidence survives the token budget. On ISIC 2019, mask-supervised LATS consistently outperforms Random and ToMe at headline budgets while retaining substantially more lesion patches. Code is provided for reproducibility, and complete tabulated results are included in the supplementary material.
cs.AI / 64 / 2609.30703
SAGE: Source-Anchored Guidance via Frequency Equalization for Hierarchical RGB-T Alignment and Fusion
Timing Li, Yiming Sun, Boan Tao, Xiyuan Gao, Haifang Cao, Pengfei Zhu
cs.CV · cs.AI
Abstract
Spatial misregistration and cross-modal discrepancies often cause ghosting, structural blurring, and content imbalance in RGB-T fusion. Existing methods typically decouple appearance adaptation, geometric alignment, and information fusion, limiting dependency propagation across stages. We propose Source-Anchored Guidance via Frequency Equalization for Hierarchical RGB-T Alignment and Fusion (SAGE), a unified framework integrating frequency equalization, hierarchical alignment, and subband fusion. SAGE employs invertible joint encoding and source-specific low-frequency modulation to derive structural and gain guidance while preserving source information. Hierarchical frequency collaborative alignment estimates global affine geometry from low-frequency approximations and transfers geometric and contextual cues to high-frequency correlation reasoning for reliability-aware residual refinement. Guided subband fusion jointly aggregates the aligned frequency coefficients under propagated source and alignment guidance, coordinates complementary low- and high-frequency information, and reconstructs the fused image through the inverse wavelet transform. Extensive experiments on RGB-T datasets with real-world and synthetic misalignments demonstrate consistently competitive performance in alignment and fusion, validating the effectiveness of source-anchored guidance for weakly registered RGB-T images.
cs.AI / 65 / 2609.30708
Combining General and Domain-Specific Pretext Tasks for Brain MR Image Segmentation
Tasneem Nasser, Susanne Schmid, Roberto Souza, Naser El-Sheimy
cs.CV · cs.AI
Abstract
A key challenge in medical image analysis is the scarcity of large annotated datasets for specific populations and diseases. As deep learning models rely heavily on labeled data, effective transfer learning strategies are needed to reduce the dependence on manual annotations. Self-supervised learning has emerged as a promising approach for developing foundation models by enabling the learning of transferable feature representations from large-scale unlabeled medical imaging datasets. In this study, we investigate voxel-level brain age prediction as a domain-specific self-supervised pretext task and compare it with image inpainting, a widely used non-domain-specific alternative. We further propose a multitask self-supervised pretraining framework that jointly optimizes both objectives to learn complementary neuroimaging representations. The pretrained models are evaluated on three downstream magnetic resonance image segmentation tasks: multiple sclerosis lesion segmentation, ischemic stroke lesion segmentation, and cortical brain structure segmentation. Overall, the proposed multitask pretraining framework consistently outperformed the single-task pretrained models and training from scratch across most experimental settings, demonstrating the benefit of combining domain-specific and general self-supervised learning pretext tasks for the development of generalizable neuroimaging foundation models.\ Code Availability: The source code used in this study is publicly available at https://github.com/TasneemN/Combining-General-and-Domain-Specific-Pretext-Tasks-for-Brain-MR-Image-Segmentation/
cs.AI / 66 / 2609.30709
VLALight: Lightweight Vision-Language-Action Models for Emergency-Aware Traffic Signal Control
Kemou Jiang, Maonan Wang, Xingchen Zou, Jiayue Zhu, Yuhang Fu, Sicheng Wang, Xi Chen, Yirong Chen, Zhiyong Cui
cs.CV · cs.AI
Abstract
Traffic signal control (TSC) is essential for mitigating urban congestion. Recent advances in vision-language models (VLMs) enable richer interpretation of intersection scenes, opening new opportunities for visual-context-aware TSC. However, the loose coupling and repeated information conversion between modules can lead to the loss of fine-grained visual details, while sequential inference introduces substantial latency. To address these limitations, we propose VLALight, a lightweight end-to-end vision-language-action framework that directly maps intersection observations and signal-phase information to discrete signal actions. To handle the multi-view nature of TSC, VLALight combines multiple directional camera views into a unified visual input and uses textual instructions to establish their correspondence with traffic movements and signal phases. This design enables direct action prediction with a compact 0.5 B-parameter model, without intermediate image-to-text descriptions or handcrafted traffic-state representations. Experiments show that VLALight delivers the best emergency-vehicle service of all compared methods, reducing pooled emergency waiting time by 21.1% over the cascaded VLMLight while running in real time on local hardware and generalizing to unseen intersection topologies and traffic-flow patterns.
cs.AI / 67 / 2609.30722
TrafficImag: A Benchmark for Counterfactual Roadside Traffic Video Generation
Xiangyu Li, Tianyi Wang, Zhihao Dou, Christian Claudel, Zhaomiao Guo
cs.CV · cs.AI
Abstract
Existing roadside traffic datasets support perception, forecasting, and visual question answering, but they do not evaluate counterfactual video generation, in which a selected actor is modified and the generated future should remain consistent with road topology and unrelated traffic. We introduce TrafficImag, the first benchmark for counterfactual roadside traffic video generation. TrafficImag combines a large-scale roadside dataset (9,022 annotated images, 7,043 deduplicated video clips, and 31,145 actor-centered history-future samples) with an executable protocol that supports behavior reasoning, intervention-aware image editing, and conditional video generation. Each intervention is represented as an actor-level program describing the target actor, intended behavior, legal route, interaction order, and temporal constraints, enabling a unified evaluation interface across heterogeneous foundation models. TrafficImag evaluates four complementary validity dimensions: initial-state correctness, route and behavior validity, interaction consistency, and non-target preservation, and considers an end-to-end counterfactual successful only when all four are satisfied. Across state-of-the-art foundation models, the strongest reasoner reaches 80.4% macro F1, the complete condition interface raises end-to-end success from 23.3% to 55.0% for the best generator. Oracle studies further show that conditional video execution is the primary remaining bottleneck. TrafficImag provides a reproducible benchmark for evaluating and diagnosing counterfactual traffic video generation beyond perceptual video quality.
cs.AI / 68 / 2609.30928
UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound
Quanhao Zhu, Bo Xu, Rui Lin, Chenyuan Wang, Yu Shao, Boling Zhu, Jiuyan Sun, Liang Zhao, Hongfei Lin, Feng Xia
cs.CV · cs.AI
Abstract
Ultrasound is one of the most widely used medical imaging modalities, and recent large vision-language models(VLMs) have shown increasing capabilities in ultrasound image understanding. However, these models fail to provide pixel-level visual evidence aligned with their semantic predictions, and their fine-grained grounding capability in ultrasound remains largely unclear. We introduce UltraG-Bench, a large-scale multi-task benchmark for evaluating pixel-level evidence grounding in ultrasound. UltraG-Bench is built by annotating 40 public ultrasound segmentation datasets spanning 13 anatomical categories, and comprises three progressive tasks: instruction-guided segmentation, evidence-grounded VQA, and evidence-grounded report generation, with 331125, 666779, and 138832 annotations, respectively. Comprehensive evaluation of 14 state-of-the-art models reveals a substantial gap between semantic understanding and fine-grained pixel-level localization. We further propose UltraG-Agent, which combines the semantic reasoning capabilities of a VLM with the ultrasound-specific segmentation capability of UltraSAM3. Experiments show that UltraG-Agent substantially improves both semantic prediction and pixel-level visual grounding. Our dataset and code are available at https://github.com/zhuqh19/UltraG-Bench.
cs.AI / 69 / 2609.30946
OneWorld: Learning Consistent Physics Across Actions in World Models
Ke He, Yichen Ding, Bin Yang
cs.CV · cs.AI
Abstract
Action-conditioned video world models aim to predict scene evolution under different actions, a capability that is essential for reliable planning, decision-making, and interaction in dynamic environments. However, futures generated independently from the same initial scene may each appear plausible while implying incompatible physical properties, such as friction or mass. This inconsistency can lead to contradictory predictions across interventions, making it difficult for the model to maintain a coherent understanding of the underlying world and limiting its reliability for planning and decision-making. To address these issues, we propose OneWorld, a shared-mechanism counterfactual generation framework that jointly models multiple action-conditioned futures under a common latent physical mechanism. A physical mechanism interpreter first infers a distribution over latent mechanisms from each action-outcome branch. These distributions are then aggregated into shared-world evidence, which captures whether the branches admit a common physical explanation while accounting for uncertainty in less informative branches. This evidence constrains flow training and guides sampling, encouraging consistency in the underlying physical mechanism while preserving the distinct outcomes induced by different actions. We further introduce a multi-intervention evaluation protocol in controlled environments, following the interaction settings of ACWM-Phys, to assess whether generated futures can be jointly explained by the same physical parameters, alongside standard measures of single-rollout prediction quality. Experiments in these environments show that OneWorld improves cross-intervention physical consistency while maintaining competitive single-rollout prediction quality.
cs.AI / 70 / 2609.30952
MVVBench: Benchmarking 4D Reasoning in Vision-Language Models
Hyungjin Chung, Byeongjun Park, Joonseok Lee, Hojun Kim, Jaeho Choi, Byung-Hoon Kim
cs.CV · cs.AI
Abstract
Multi-view video understanding requires integrating spatial and temporal evidence across multiple, often non-overlapping camera streams: tracking entities as they transition between viewpoints, aligning events across time, and reasoning about latent 4D continuity rather than any single visible frame. We introduce MVVBench, a benchmark for multi-view video reasoning built from real world multi camera datasets. Questions are curated to be monocular-ambiguous along both the view and the temporal axis: each question is unanswerable from any single view in the designated input set, and the majority are further unanswerable from any single moment. Each question becomes uniquely solvable only by jointly reasoning across views and across time. MVVBench spans diverse dynamic scenes and probes six capabilities: implicit/explicit attribute identification, implicit/explicit relative distance, relative camera pose, and compositional counting, with human-authored QA and rigorous verification. Beyond benchmarking, we provide an extensive analysis of when and why current vision language models succeed or fail, characterizing errors due to temporal mis-localization, cross-view identity breaks, and brittle multi-hop reasoning. We then study inference-time elicitation strategies that unlock latent multi-view competence---task-specific chain-of-thought scaffolds and structured cross-view evidence aggregation---yielding substantial gains without retraining. Finally, we present preliminary evidence that reinforcement learning with verifiable rewards can elicit some latent multi-view competence in the base model, pointing to training-time approaches as a promising direction for future work. Together, MVVBench offers a rigorous evaluation of 4D multi-view reasoning and a foundation for future progress toward reliable embodied perception.
cs.AI / 71 / 2609.30982
FARE: Forensic Acceptance Region Estimation for Catching Bait-and-Switch Image Generators
Kai Yao, Marc Juarez
cs.CV · cs.AI · cs.CR
Abstract
Modern AI image generators are increasingly deployed as opaque APIs, where customers can query the deployed service, but cannot inspect model weights or architecture. This creates a practical challenge: a provider may pass governance certification with one generator and later silently switch to a cheaper and lower-quality one for deployment, compromising public trust or even safety in high-stakes domains. We study integrity auditing at deployment time and propose FARE (Forensic Acceptance Region Estimation). A certified generator is enrolled by training FARE on images sampled from that generator. After deployment, FARE can determine whether a generated image is consistent with the enrolled generator---using only that image. FARE's features are based on image generator-specific artifacts that have been proposed for forensic applications. FARE amplifies these features during training by finding hard samples that tighten the acceptance region and increase sensitivity to subtle changes in the certified generator. Across generator swaps, including substitutions with similar model versions and model variants, FARE is effective at detecting swaps, consistently outperforming existing baselines at strict operating points, and remains effective under the exact-model and decision-only attacks evaluated in this work.
cs.AI / 72 / 2609.30993
FLIP: Final Layer Inference-Time Probing for Vision-Language Models
Drandreb Earl O. Juanico, Rowel O. Atienza
cs.CV · cs.AI
Abstract
We present FLIP, a final-layer inference-time probe for testing whether a logit-facing intervention site in an open-weight vision-language model (VLM) supports structured, task-linked computation rather than generic perturbation. Behavioral change under internal intervention is otherwise mechanistically ambiguous: it may reflect improved use of visual evidence, generic output instability, or outright degradation. FLIP applies elementwise flooring to the final normalized hidden state before logit computation, leaving parameters, prompts, and decoding unchanged. On a controlled detection/counting probe, sweeping intervention strength reveals three regions: negligible change, a bounded interior regime in which detection recall at IoU 0.50 ($R_{50}$) improves while tolerant counting error ($\mathcal{E}_{\mathrm{count}}$) falls, and over-suppression. We formalize a four-criterion probe-and-sweep protocol for disciplining the interpretation of intervention effects: regime structure, grounding-proxy alignment, feature-coherence dependence, and failure to reproduce the same positive regime on a performance-based negative control. The post-normalization state passed to the output head is the logit-facing instantiation of this test; under a non-targeted flooring sweep it satisfies the full protocol. Raw decoder-layer interventions, including the last-block output before final normalization, and the singleton-pair left/right control fail to reproduce the Final-site signature, while same-site operators and multiple VLMs replicate it. FLIP is therefore a validation step for intervention-based mechanistic interpretability, not a steering method.
cs.AI / 73 / 2609.31103
DepthEvidence: Unifying Metric Depth Prediction and Geometric Reasoning in Multimodal Language Models
Jiangning Wei, Yuan Yao, Miaomiao Cui, Mingsheng Li, Humen Zhong, Shuai Bai, Zhibo Yang
cs.CV · cs.AI
Abstract
Spatial reasoning with metric constraints requires linking objects to geometric measurements and preserving their numerical content during language reasoning. We present DepthEvidence, a 4B model that uses its own dense metric predictions as object-grounded evidence for language generation. A camera-conditioned decoder predicts full-resolution metric depth using multi-scale visual features and high-resolution RGB refinement. A dense-to-language interface converts predicted depths and decoder features into object-aligned continuous geometry tokens anchored to object identifiers. Geometric supervision encourages metric information to remain recoverable before and after language-context interaction, while instruction tuning supports object measurement and compositional reasoning. We introduce a Depth-VQA benchmark evaluating object-depth queries, relative comparisons, and decisions combining spatial and numerical constraints. Across nine datasets, DepthEvidence achieves the highest average dense $δ_1$ among evaluated methods, competitive with specialized estimators. It also leads the evaluated methods in instance-level metric depth estimation and overall accuracy on both relative and metric reasoning tracks, while broadly preserving general VQA performance and improving spatial understanding relative to the base model.
cs.AI / 74 / 2609.31150
FedHisto-PAST: Parameter-Efficient Stain-Aware Federated Learning for Cross-Site Lung Histopathology Classification
Muhammad Muhtasim Shahriar, M. M. Golam Hafiz, Saad Aloteibi, Mohammad Ali Moni
cs.CV · cs.AI
Abstract
Cross-site lung histopathology classification must account for stain variation, non-IID client data, missing classes, and the cost of adapting large pathology encoders. This study evaluates FedHisto-PAST v2 for three-way classification of adenocarcinoma (ACA), Normal, and squamous cell carcinoma (SCC). FedHisto-PAST v2 combines a frozen HIBOU-B foundation model with parameter-efficient adaptation, stain-conditioned paired-view prediction and feature consistency, reliability-aware prototype learning, and adaptive federated aggregation. Experiments used a five-client, non-IID, raw-data-local simulation with fixed internal evaluation, client-level analysis, component ablations, communication accounting, and a development-influenced exploratory LungHist700 cohort. All principal methods achieved near- ceiling internal performance, which limited discrimination on the fixed split. On LungHist700, FedHisto- PAST v2 achieved a Macro-F1 of 0.728560 and a balanced accuracy of 0.730454. Higher recognition of Normal and SCC was accompanied by lower ACA recall, and calibration remained imperfect. Prediction-level consistency was the only component with a clearly supported independent contribution in the external ablation analysis. Feature consistency and prototype regularization showed no conclusive independent overall gains in Macro-F1. The framework updated 1.253841% of the model parameters. The results provide exploratory cross-dataset evidence for stain-aware, parameter-efficient federation; they do not establish formal privacy, patient-level independence, prospective deployment, or clinical validation.
cs.AI / 75 / 2609.31160
ReG-SAM: Reference Graph-Driven SAM for 2D Foundational Vessel Segmentation
Donghang Lyu, Zichen Zhang, Oleh Dzyubachyk, Marius Staring
cs.CV · cs.AI
Abstract
Vessel segmentation in medical images is essential for many clinical tasks, ranging from diagnosis to treatment planning. However, it remains challenging due to complex vascular morphology and diverse imaging conditions. Existing deep learning methods rarely aim at building a generalizable vessel segmentor across anatomies and modalities. While the Seg- ment Anything Model (SAM) has shown promise for med- ical image segmentation, its original design does not fully exploit vascular morphology and struggles with fine-grained vascular structures, leading to suboptimal performance. In this paper, we propose ReG-SAM, a SAM-based framework tailored to 2D vessel segmentation that leverages reference graph set for enhancing vascular representations. Specifically, we introduce two modality-aware representations derived from the reference masks: graph prompt embeddings (GPEs) that encode global spatial features from graphs, and vascu- lar prototype embeddings (VPEs) that capture fine-grained modality-specific vessel characteristics from multi-scale fea- ture maps and vascular masks. Since both require vascular masks that are unavailable during inference and require robust modality-aware vascular feature representations, we construct a modality-wise vascular database and develop two reference graph-guided representation learning schemes for estimating GPEs and VPEs using samples from the database rather than ground-truth masks. Extensive experiments across 19 datasets demonstrate that ReG-SAM consistently outperforms existing baselines, even those using manual prompts, particularly on challenging thin vessels.
cs.AI / 76 / 2609.31247
Geometric Inconsistency Localization in Multi-View Image Sets
Xander Staelens, Albéric Loos, Bert Ramlot, Hannes Mareen, Peter Lambert, Glenn Van Wallendael
cs.CV · cs.AI · cs.CR · cs.MM
Abstract
Novel view synthesis (NVS) models can produce realistic new views of the same scene from different viewpoints. However, these generated views are not always geometrically consistent with one another. Multi-view (MV) consistency has shown promise as a tool for evaluating these NVS models. Its potential for multimedia forensics, however, remains largely unexplored, particularly for localizing geometric inconsistencies across wide-baseline image pairs. To enable research in this direction, we introduce DeformView, a wide-baseline MV dataset with pixel-level annotations of geometric inconsistencies. Using DeformView, we evaluate state-of-the-art MV consistency-scoring methods and show that approaches developed for NVS evaluation transfer poorly to the forensic task of geometric inconsistency localization. To address this limitation, we propose DEFECt3R, a lightweight learning-based classifier that uses cross-view feature relationships to localize geometric inconsistencies at the pixel level. By learning from explicit supervision, including hard negatives from geometrically consistent yet deformed views, DEFECt3R improves localization performance and substantially reduces false positives compared to existing consistency-scoring methods. Ablation experiments further show that both feature representations and correspondence quality contribute to localization performance. Overall, our findings demonstrate that MV geometric consistency is a promising yet underexplored signal for multimedia forensics and establish a benchmark and baseline for geometric inconsistency localization in wide-baseline MV image pairs. Code and dataset are available at https://github.com/IDLabMedia/DeformView-DEFECt3R
cs.AI / 77 / 2609.31298
UniAR: A Unified Framework for Autism Recognition Enhanced by Multi-View Prompt Learning
Lei Xin, Zeheng Wang, Jiayin Zhu, Shihong Huang, Fanhu Zeng, Changjiang Jiang, Dengbo He, Yutao Yue, Zhenglun Kong
cs.CV · cs.AI
Abstract
Autism Spectrum Disorder (ASD) is a complex neurodevelopmental disorder for which early and accurate diagnosis is critical to improving long-term developmental outcomes. However, existing ASD recognition methods are often constrained by the scarcity of diagnostic text data, forcing them to rely mainly on visual analysis and limiting their ability to model clinically meaningful semantic reasoning. To address this challenge, we propose UniAR, a unified framework enhanced by multi-granularity prompt learning for robust ASD recognition under heterogeneous data variations. Specifically, UniAR leverages a large multimodal model to generate hierarchical diagnostic descriptions at the word, phrase, and sentence levels, compensating for the lack of paired clinical reports. To align the generated semantics with visual evidence, we further design a Mixture-of-Experts-based Multi-Scale Alignment Module, which dynamically matches vector-quantized visual prototypes with semantic representations at corresponding granularities. Extensive experiments on four benchmarks covering brain MRI and facial expression scenarios show that UniAR consistently outperforms existing state-of-the-art methods, achieving average accuracies of 75.9\% on MRI benchmarks and 91.6\% on facial benchmarks, while improving average Accuracy on MRI benchmarks by 1.5 percentage points and average Accuracy on facial benchmarks by 1.2 percentage points over baselines. These results demonstrate that UniAR offers a robust and interpretable framework for ASD screening under semantic scarcity.
cs.AI / 78 / 2609.31326
CG-HAF: An Interpretable Global-Local Lesion-Burden Fusion Framework for Ordinal Acne Severity Grading in Agentic Skincare Support
Muhammad Muhtasim Shahriar, Md. Naimur Asif Borno, Saad Aloteibi, Mohammad Ali Moni
cs.CV · cs.AI
Abstract
Ordinal acne severity grading requires distinguishing visually similar neighboring grades while jointly weighing holistic facial appearance and localized lesion burden - evidence that most existing approaches collapse into a single opaque representation. We introduce CG-HAF, a global-local fusion framework that instead keeps this evidence explicit: averaged holistic severity probabilities from independently trained classifiers are combined with structured lesion-burden descriptors from an object detector (lesion count, detection confidence, lesion area) into a compact representation, from which a lightweight, interpretable classifier produces the final grade. On a widely used benchmark, this fusion yields a clear, statistically supported improvement over global-evidence-only baselines, with the largest gains on the most severe cases. Testing on an independent dataset with a different grading standard shows that strong within-dataset performance does not transfer automatically, and a follow-up diagnostic attributes much of this gap to mismatched grading criteria rather than detection failure alone. These findings support interpretable global-local fusion as an effective strategy for ordinal acne grading while highlighting criterion alignment as key to cross-dataset portability, with a further illustration of how the resulting severity signal can support transparent, non-diagnostic decision-making in skincare applications.
cs.AI / 79 / 2609.31435
Implicit Neural Representation for Hyperspectral Video Compression
Alfredo Scalera, Paul Murray, Jaime Zabalza
cs.CV · cs.AI · cs.LG
Abstract
With the advent of snapshot cameras, hyperspectral video is becoming more readily available. In recent years, new applications have emerged which have led to increasingly larger datasets. However, hyperspectral video compression remains in the early stages. In this study, we explore the use of implicit neural representation as a candidate solution. We propose a novel extension of an existing RGB video compression model, achieving Bjøntegaard Delta PSNR gains of +4.99 dB and Bjøntegaard Delta rate of -88.88% compared to traditional hyperspectral image compression methods applied frame-by-frame. In addition to reconstruction quality, the effects on downstream task performance are measured in the form of object tracking success. Compared to video compressed with methods based on principal component analysis and JPEG2000 in low data regimes, our proposed method improves tracking area under the curve by up to 23.42% and distance precision by up to 35.56% on examples from the HOT2026 dataset.
cs.AI / 80 / 2609.31450
From Reward Signal to Visual Utility: A Controlled Audit of Medical VLM Post-Training
Wang Jingxin
cs.CV · cs.AI
Abstract
Medical vision-language model (VLM) post-training is commonly evaluated through answer accuracy. We examine how changes in accuracy and training objectives relate to image-conditioned decisions in a controlled Qwen2.5-VL-3B study on PMC-VQA. We compare supervised fine-tuning (SFT) with low-rank adaptation (LoRA) restricted to the language model, expanded multimodal adaptation scopes, standard answer-only Group Relative Policy Optimization (GRPO), and a counterfactual evidence objective. On 2,000 clean-test questions, language model LoRA SFT changes correct-image accuracy by +1.10 percentage points (95% paired bootstrap CI:-0.85 to +3.05), while visual-benefit events decrease by 2.40 points and image sensitivity decreases by 5.60 points. Paired records reveal 155 acquired and 203 lost visual-benefit events. Broader adaptation yields lower correct-image accuracy than language-model LoRA SFT. Standard GRPO produces mixed-reward groups and parameter updates, with an uncertain clean test accuracy change. A generation audit reveals that canonical option scores can follow a different token path from generated answers. With scores taken along the greedy generation path, the evidence target improves on the training set; its gains over standard GRPO remain inconsistent on validation data at matched training doses. Sample-level analyses trace how evidence scores, decision margins, and generated answers change during post-training. This empirical and measurement audit identifies gaps between optimization activity, target acquisition, and useful held-out visual behavior.
cs.AI / 81 / 2609.31509
ClearGS: Reliability-Aware Gaussian Splatting from Handheld Videos
Xuanzhi Liu, Xinyi Wu, Hang Pan, Wensi Huang, Zhenyao Wu, Ruize Han, Song Wang
cs.CV · cs.AI
Abstract
We present ClearGS for 3D Gaussian Splatting (3DGS) from handheld videos with uneven viewpoint coverage and mixed frame quality. Rather than selecting frames with binary decisions, ClearGS uses Reliability-aware View Allocation (RVA) to assign graded raw-supervision weights based on appearance reliability, degradation risk, and geometric utility, while weakly reactivating useful suppressed frames to maintain trajectory coverage. Since weighting cannot restore details lost to blur or distortion, ClearGS further introduces Render-Guided In-Video Restoration (RIVR). The current 3DGS render provides a pose-aligned structural candidate, a frozen no-reference restoration expert restores the corresponding raw video observation without any clean reference image, and no-reference perceptual scores select among the render, restored observation, and high-frequency fused candidate. ClearGS then applies Full-Trajectory Repair Consolidation to revisit accepted repairs and preserve details introduced early. On GS2E and GSOTM, ClearGS achieves state-of-the-art overall performance, with consistent CLIP-IQA and MUSIQ gains and LPIPS reductions in most degradation settings, without paired sharp supervision or matched clean references.
cs.AI / 82 / 2609.31572
OC-GS: Gaussian Splatting for Irregular Turntable Capture
Jae Joong Lee, Bedrich Benes
cs.CV · cs.AI
Abstract
Uneven rotation and dropped frames make equal-angle assumptions unreliable for turntable reconstruction. We present OC-GS, an object-centric Gaussian splatting that refines each image's angle while maintaining a shared camera, rotation axis, and pivot. This orbit-consistent refinement jointly optimizes image-derived geometry and angles to reconstruct objects from sparse, irregular captures. On rendered objects with 12, 8, and 6 irregularly spaced views, OC-GS achieves mean foreground PSNR scores of 21.26, 19.36, and 15.83dB, respectively, exceeding all four evaluated pose-free Gaussian splatting baselines in each condition. Under a shared trainer, refining image-estimated angles improves mean foreground PSNR by 7.88dB over keeping those estimates fixed. An ablation study shows that both image-derived angle initialization and the shared motion model contribute to the improvement. On real captures, OC-GS's refinement increases mean foreground PSNR by 0.70dB. Results show that refining uncertain angles within a shared motion model improves reconstruction from sparse, irregular turntable captures.
cs.AI / 83 / 2609.30747
Beyond the Last Truffula Tree: SustainAI - A Water-Aware, Closed-Loop Framework for Environmentally Accountable AI
Farnaz Farid, Tashfia Towkee, Sania Nasreen, Sami bin Azad
cs.CY · cs.AI · cs.DC · cs.HC
Abstract
As artificial intelligence (AI) becomes embedded in everyday life, its environmental footprint, particularly water consumption remains largely invisible. While energy and carbon impacts are widely recognized, the substantial freshwater demands of data center cooling and electricity generation receive little attention. To address this gap, we introduce SustainAI, a water-aware, closed-loop framework incorporating environmental accountability into AI deployment. SustainAI integrates real-time water metering, a hallucination-aware penalty model, and a water-aware routing algorithm that accounts for regional water stress. Evaluated via Small Language Models (SLMs) extracting health misinformation, results reveal an 11-fold variation in water footprint across geographically distributed data centers (0.0477 mL to 0.5360 mL per inference). Across 1,335 inference runs, the system consumed approximately 399 mL of water but produced only 240 correct outputs, demonstrating that substantial resources are spent on inaccurate responses. Crucially, SustainAI extends beyond technical optimization through a Care by Design lens, framing AI sustainability around relational ethics, regional equity, and ecological stewardship. By combining water monitoring, adaptive accountability, and Care by Design principles, SustainAI provides a practical foundation for integrating ethical care and environmental responsibility into AI infrastructure design and lifecycle management.
cs.AI / 84 / 2609.31047
DynBranch: Speculative Subgraph Reuse for Dynamic Agentic LLM Serving
Junyi Shen, Noppanat Wadlom, Zhengyuan Su, Yao Lu
cs.DC · cs.AI · cs.LG
Abstract
Agentic LLM workflows decide their execution paths at runtime. Downstream computation may be predictable, or may have run before, yet it cannot begin until the model or the user resolves the branch. We call this serialization the branch-resolution barrier. Caching alone does not hide it: the key that identifies a reusable result is not known until then. In this paper, we propose DynBranch, which makes an unresolved branch addressable before it resolves. Its stable coordinate lets candidate subgraphs run during resolution and completed subgraph results be reused across later requests. A two-level controller admits this work when its expected benefit exceeds the load price. DynBranch sits at the model-API boundary and requires no changes to agent harnesses or model execution engines. Across four agentic workloads with Qwen3-32B on 4x H200 GPUs, DynBranch reduces mean latency by up to 32% over each workload's strongest prior system and by 46-66% against a no-reuse floor, while preserving workflow results. The benefit persists across backbone families and on a commodity Qwen3-8B/RTX 4090 deployment.
cs.AI / 85 / 2609.31358
A Safety-Bounded SDC-to-MCP Gateway for Medical AI Agents
Bennet Gerlach, Stefan Fischer
cs.DC · cs.AI
Abstract
The Model Context Protocol (MCP) provides a common interface through which AI applications discover and use external resources and tools. It allows language-model agents to ground their reasoning in current system state and interact with heterogeneous services. In medical environments, however, exposing device state and action affordances requires deterministic constraints on possible effects. We present an IEEE 11073 Service-Oriented Device Connectivity (SDC)-to-MCP gateway that exposes metrics, alarms, context references, and semantic metadata as read-only resources, while representing selected action affordances as policy-validated dry-run tools. The term safety-bounded denotes a narrow no-execution property: agent-facing requests dispatch no SDC device operation. A Python prototype supports simulated fault and lifecycle experiments, a software-reference protocol path spanning independent Java and Python implementations, deterministic baselines, representation ablations, and multi-model agent evaluation. The results show semantically explicit resource exposure, visible rejection of invalid or outdated state, and preservation of the no-execution boundary across resource, proposal, and authorization paths. Explicit semantic metadata improved conformity to required metric identifiers in structured alarm outputs relative to a generic representation, while retained structured-output failures reveal a distinction between plausible narrative answers and task-compliant machine-readable results.
cs.AI / 86 / 2609.30555
Proportional Representation in Temporal Voting with Ranked Preferences
Noam Hazon, Leora Schmerler, Nicholas Teh
cs.GT · cs.AI
Abstract
We study proportional representation in temporal voting, where one candidate is selected in each round. While prior work has focused on approval ballots, we consider ranked preferences, which may change over time. A natural approach treats each voter's top candidates as approved, but the right cutoff may differ across voters and rounds. We therefore require proportionality to hold for every admissible choice of cutoffs, whether fixed and common, common but varying across rounds, or set individually for each voter in each round. Combining these interpretations with temporal versions of justified representation (JR), proportional JR (PJR), extended JR (EJR), and proportionality for solid coalitions (PSC) gives us a hierarchy of axioms. We ask which of these axioms can be guaranteed, and with how much knowledge of the future. Unlike with approval ballots, no version of EJR can be guaranteed, and for the other axioms, flexibility in the cutoffs comes at a price. With a fixed common cutoff, JR, PJR, and PSC can be guaranteed, but only by rules that see all preferences in advance. Once the cutoff may vary across rounds, even such rules cannot guarantee JR or PSC for groups that agree in only some rounds. For groups that agree in every round, however, knowing only the number of rounds suffices for PJR in polynomial time, and PSC needs no knowledge of the future at all. Under individual cutoffs, no version of JR or PJR can be guaranteed, yet a rule as simple as serial dictatorship achieves PJR up to an additive loss that no rule can improve on, however much it knows. Natural preference restrictions restore exact guarantees. Finally, we show that checking our axioms is often coNP-complete; but perhaps surprisingly, a stronger axiom can be easier to check.
cs.AI / 87 / 2609.30466
A Benchmarking Framework for Context-aware XR Interfaces
Hyunsung Cho, Sarah Yewon Yun, Nancy Ruonan Sun, Ben Lafreniere, Mark Parent, Kashyap Todi, Tanya R. Jonker, Hrvoje Benko, Sherry Tongshuang Wu, David Lindlbauer
cs.HC · cs.AI
Abstract
Everyday Extended Reality (XR) systems aim to provide context-aware access to the right functionalities at the right time and place, with minimal manual reconfiguration as users switch context. Yet these interfaces are hard to evaluate: current prototyping and user-study workflows offer no systematic, repeatable way to compare adaptation methods across users and scenarios. We present ContextXR, a novel benchmarking framework for context-aware XR interfaces. ContextXR represents an XR application as a connected graph of functional facets, each a semantically coherent group of related capabilities that together support a shared user intent. On this representation, we build MineXR++, a dataset augmenting prior XR interface data with facet-level annotations, and formulate three canonical tasks of context-aware suggestion: context factor analysis, initial facet suggestion, and next facet suggestion. Our evaluation protocol scores suggestion methods by a simulated interaction metric, the navigation and search cost of reaching the desired functionality. Through experiments benchmarking global popularity, relational retrieval, and LLM-based methods, we demonstrate that ContextXR enables the systematic, reproducible evaluation of context-aware XR interfaces.
cs.AI / 88 / 2609.31272
Cognitive Skills in the Age of AI: Computing Students and Experts Perceptions
Neha Rani, Vu Minh Anh Le, Austin M. Spangler, Erta Cenko
cs.HC · cs.AI
Abstract
AI is becoming increasingly integrated into daily workflows, especially in computing. We are gradually shifting towards an AI-rich future, an impending yet unknown one. One important emerging concern is whether we are accordingly preparing our future computing workforce. Further, we need to know what the important cognitive skills are to remain relevant in the computing workforce and if there are changes in cognitive skill importance. To investigate this direction, we conducted a mixed-methods study, collecting perceptions from computing students and computing experts regarding the importance of cognitive skills in the past, present, and future. We report that the perceived importance of most cognitive skills will decrease in the future, with an AI-rich environment, but critical thinking skills remain important. Further, we report reasons collected through interviews on why the importance of cognitive skills will change and how future computing students can prepare for it.
cs.AI / 89 / 2609.31569
Adapting for AI: How elementary teachers adjust their practices for an AI-integrated curriculum
Fasika Melese, Ruiyang Wu, Xinyue Cui, Joanna Perkins, Xiaoyi Tian, Tiffany Barnes, Shiyan Jiang
cs.HC · cs.AI
Abstract
Conversational AI tools are entering children's everyday experiences, and schools are interested in adopting them. However, successful classroom integration depends not only on the technology but also on the work teachers do to make it usable and appropriate for their students and classroom context. There is little known about how elementary teachers work as they implement conversational AI tools in real classrooms. In this study, we examine three teachers' experiences implementing an AI literacy and English Language Arts (ELA) curriculum built around ToyTalk, a conversational AI toy development platform, over 13 instructional days, a three-week summer camp. Drawing on daily individual reflections, group reflections, and post-camp interviews, we find that teachers' adaptive practices of repair, differentiation, translation, and balancing sit at the intersection of three tensions (technology, learner, and instruction). Teachers' understanding of AI and their role evolved over the camp experiences. From these findings, we contribute design implications and considerations for deploying conversational AI within elementary classrooms.
cs.AI / 90 / 2609.31164
SPADE: Escaping the Popularity-Similarity Frontier to Measure Serendipitous Recommendations
Tobias Vente, Maarten Peirsman, Noah Daniëls, Hannu Toivonen, Bart Goethals
cs.IR · cs.AI · cs.LG
Abstract
Recommender systems engineer serendipity to foster active exploration and break predictable consumption cycles. The problem with existing offline beyond-accuracy metrics is that they often either isolate historical similarity or global popularity. We aim to design an evaluation metric that examines similarity, popularity, and actual user relevance. To achieve this, we introduce SPADE (Serendipitous Pareto Distance Evaluation). SPADE maps all items into a two-dimensional space to directly calculate a user-specific Pareto frontier of maximally popular and historically similar items. The final serendipity score is then computed by averaging the minimum Euclidean distance from this boundary strictly for the correctly recommended test-set items. Evaluating SPADE across five datasets and five baseline algorithms confirms its effectiveness; our results show that the metric successfully prevents algorithms from exploiting beyond-accuracy measures with irrelevant or non-personalized recommendations, reliably isolating serendipitous discoveries.
cs.AI / 91 / 2609.31166
AgentRecommender: LLM Agents Enable Customizable Recommender Systems on the User Side
Ryoma Sato
cs.IR · cs.AI · cs.DB · cs.DL
Abstract
Recommender systems have traditionally been developed for platforms. However, this has given rise to many phenomena that may be advantageous for platform lock-in but are a nuisance to users, such as clickbait, filter bubbles, and the spread of fake news. Recently, user-side recommender systems have been proposed as a new paradigm for solving this problem. If users deploy their own recommender systems, they are no longer at the mercy of the platform's interests. However, building a user-side recommender system is not trivial; in particular, customizing one for oneself requires additional data. We propose AgentRecommender, a method that leverages the investigation capability and internal knowledge of LLM agents to flexibly build user-side recommender systems without additional data. AgentRecommender allows users to easily create recommender systems tailored to their own preferences.
cs.AI / 92 / 2609.30690
Threat-Aware Energy-Efficient Deployment for Dynamic UAV Networks: A Multi-Agent RL Approach
Faisal Al-Kamali, Hussein A. Ammar, Francois Chan, James H. Bayes, Yasser Gadallah, Mohamed H. Ahmed
cs.IT · cs.AI · cs.LG · eess.SP
Abstract
Ensuring operational safety in threat-prone environments remains a critical challenge for multi-UAV networks serving as aerial base stations. This paper proposes an efficient framework to maximize global energy efficiency (EE) while promoting safe operation through threat-aware clustering and reward-based safety enforcement. The proposed framework is executed in three steps. First, a threat-aware K-means (TAKM) algorithm determines the minimum required UAVs and computes safe initial placements. Second, an optimal matching stage assigns physical UAVs to these centroids to minimize energy expenditure. Third, a threat-aware multi-agent twin delayed deep deterministic policy gradient (MATD3) algorithm dynamically optimizes trajectories, power, and user associations. Simulation results show that the proposed framework achieves zero observed safety violations in the considered scenarios while achieving superior EE and faster convergence than other learning methods and non-clustering baselines. Compared to heuristic optimization, the proposed framework outperforms the greedy particle swarm optimization (GPSO) and achieves performance comparable to that of the optimized PSO (OPSO), while incurring significantly lower online deployment computational complexity. Furthermore, the proposed framework demonstrates effective generalization to unseen user distributions, large UAV fleets, and different threat geometries, while maintaining zero safety violations.
cs.AI / 93 / 2609.31540
Can You Check That? The Checkability Boundary for Local LLM Network Automation
Maleeha Masood, Momina Nofal
cs.NI · cs.AI
Abstract
Sending every network-automation input to a third-party frontier LLM exports sensitive artifacts such as production configurations, topologies, and logs. Querying small language models (SLMs) locally avoids this egress, but SLM outputs can be error-prone for direct use. This work introduces checkability as a criterion for determining which tasks are suitable for local inference. A task is checkable when it exposes a cheap, deterministic test - an intrinsic check - that rejects outputs violating a necessary correctness condition. We instantiate this idea in Touchstone, a local-first pipeline that uses seven off-the-shelf SLMs (1-8B parameters) to generate candidates, uses task-specific intrinsic checks to reject responses, and escalates unresolved inputs to a frontier LLM. On conflict detection and intent translation tasks, Touchstone reaches 98.6% and 93.8% end-to-end accuracy while escalating only 16% and 17% of inputs, respectively. On TeleQnA, a knowledge-only control that has no task-specific intrinsic checks, Touchstone is unable to match the accuracy of the frontier baseline. Our results support a simple deployment rule: keep inference local when task semantics support precise, low-cost checks; escalate the rest.
cs.AI / 94 / 2609.30428
Actively Resolving Contextual Uncertainty for Underspecified Tasks in Natural Language
Zachary Ravichandran, Jonathan Diller, Fernando Cladera, Varun Murali, George J. Pappas, Vijay Kumar
cs.RO · cs.AI
Abstract
Foundation models provide robots with the ability to interpret natural language and reason about environmental context, yet most language-conditioned policies assume that goals are well-specified and that task-relevant information is provided upfront via a prior map. Operating in unfamiliar environments with underspecified tasks entails high contextual uncertainty: the robot must jointly infer what constitutes task success, what constitutes relevant information, and where (or whether) that information exists. We address these limitations via CLUE (Closed-Loop contextual Uncertainty rEsolution), a framework for actively resolving contextual uncertainty given underspecified tasks in natural language. CLUE uses an LLM-derived policy to hypothesize task-relevant concepts and potential plans. It then uses a language-embedded map, which is constructed online, to ground these hypotheses into actions. The policy sequentially evaluates hypotheses via closed-loop environment interaction and refines its plans as it gathers new information. We deploy CLUE on a Boston Dynamics Spot across three real indoor and outdoor environments spanning 15 tasks that require object disambiguation, functional inference, and occlusion reasoning. CLUE achieves a success rate within 7 percentage points of an oracle policy and outperforms an LLM-enabled planner without closed-loop feedback by a 4x margin. Supporting experiments demonstrate that simply building and then querying a language-enriched map is insufficient to resolve complex contextual planning tasks; these approaches achieve roughly one third the success rate of CLUE while requiring over 10x more VLM tokens. We provide additional information at https://zacravichandran.github.io/CLUE.
cs.AI / 95 / 2609.30557
Auditing Latent-Space Monitors for Autonomous Driving
Nikhil Kamalkumar Advani, Vishwajeet Shivaji Hogale, Saurav Kumar
cs.RO · cs.AI
Abstract
Runtime failure monitors can use a model's internal representations to anticipate failures. We audit this monitoring strategy across two autonomous-driving tasks: online vectorized map generation with LaneSegNet and end-to-end planning with VAD. We find that frame-level errors are predictable at inference in both tasks. For LaneSegNet, a supervised latent probe reaches Area Under the Receiver Operating Characteristic curve (AUROC) 0.780 for high Chamfer error; to our knowledge, this is the first post-hoc frame-level failure monitor for online vectorized map generation. For VAD, a supervised planning-latent probe reaches AUROC 0.868 for mean-ADE failure. Our audit shows that internal access is not necessary for strong failure prediction. A monitor using only LaneSegNet's prediction outputs reaches AUROC 0.825, while for VAD, ego state, driving command, and the planner's predicted trajectory reach 0.924 on the same mean-ADE endpoint. Adding latent features to either baseline yields no statistically resolved improvement. This observation persists across a broad suite of planning failure endpoints, including endpoints whose labels depend on geometry unavailable to the non-latent baseline. Thus, predicting failure from an internal representation does not establish that the representation provides useful information beyond observable inputs and outputs. We propose an evaluation protocol for testing the incremental value of latent access and release our per-frame failure endpoint labels.
cs.AI / 96 / 2609.30745
Anatomy-Aware Dexterity-Driven Design Optimization of Surgical Continuum Robots
Tony Qin, Peter Connor, Khoa Dang, Carter Hatch, Caleb Rucker, Robert J. Webster, Ron Alterovitz
cs.RO · cs.AI
Abstract
Performing complex medical procedures with continuum robots requires careful selection of their geometric design parameters. The robot should have high dexterity in the specific anatomical environment of its procedure. This work presents a design optimization method that considers both dexterity and anatomy. We introduce the Reachable Volumetric Dexterous Solid Angle (RVDSA) metric as our objective, which measures the ability of a robot's end effector to reach the points in a goal volume from different directions via collision-free paths from a start configuration. We present a computationally efficient motion planner to compute this objective function for a given robotic design, and we use an asymptotically optimal simulated annealing optimizer to compute an optimized design. We applied our new method to optimize the design of a bimanual dexterous sheaths robot for performing procedures on cancerous polyps in colon anatomies, achieving a 78% higher RVDSA on average than optimizing for 3D voxel coverage alone.
cs.AI / 97 / 2609.30770
NavGen: Visual Generative Models as a Scalable Data Engine for Embodied 3D Navigation
Xijie Huang, Yongyang Wan, Chengbin Dong, Zimo Ding, Mo Zhu, Yijin Wang, Zhiyang Liu, Fei Gao, Yuze Wu, Xin Zhou
cs.RO · cs.AI
Abstract
General-purpose robot models increasingly rely on large and diverse datasets. For embodied 3D navigation, however, existing data sources face a fundamental trade-off: simulated data can be generated at scale but often suffer from the visual sim-to-real gap, whereas real-world flight data provide realistic observations but are costly to collect. This paper studies another direction: the use of high-fidelity visual generative models as scalable data engines for embodied 3D navigation. We introduce NavGen, a text-to-video data generation pipeline that produces diverse vision-language navigation (VLN) episodes across indoor and outdoor scenes. We also propose a style-diversification method that scales up long-tail data that are difficult and costly to collect. The resulting dataset contains approximately 400K navigation episodes. We evaluate our dataset against existing UAV navigation datasets across multiple metrics, and find that the model trained on our data generally improves with scale, outperforming those trained on existing datasets. To validate real-world transferability, we deploy the trained model in world-action-model paradigm to real-world flying experiments. The final model achieves a 75\% success rate across different navigation tasks and environments.
cs.AI / 98 / 2609.30818
Evaluation Is All You Need for Multi-Modal Autonomous Driving
Zeyu He, Shiqi Liu, Ke Chen, Yun Yan, Jinzi Wu, Dianqiao Lei, Sirui Wang, ShuRui Peng, Tao Chen, Zhuo Huang, Yu Wu, Yadong Shao, Zhichao Li, Ke Sun, Yang Guan, Keqiang Li, Shengbo Eben Li
cs.RO · cs.AI
Abstract
Multi-modal planning is promising for autonomous driving by representing multiple plausible behaviors in ambiguous and long-tail scenarios. Existing methods mainly focus on improving trajectory multi-modality, enhancing trajectory representations, or reshaping the candidate distribution. Nevertheless, we identify a pronounced generation-evaluation asymmetry in multi-modal planning: despite strong oracle performance, existing planners often fail to reliably select the best available candidate, leaving substantial planning potential unrealized. To address this challenge, we propose iDriveVLA, a multi-modal planning framework that improves the candidate trajectory space while enabling more reliable and context-aware trajectory evaluation. Specifically, iDriveVLA introduces a unified trajectory evaluator comprising a Safety-aware Scorer for quality and risk estimation, together with a VLM-guided Modulator for scene-adaptive criterion weighting. We further develop an oracle-aligned progressive training strategy consisting of candidate imitation pretraining, candidate space refinement, and semantic ranking alignment. On the public NAVSIM v1 leaderboard, iDriveVLA achieves a new state-of-the-art performance of 94.95 PDMS, surpassing the human-expert reference.
cs.AI / 99 / 2609.31313
Towards VLA-Dreamer: Refining VLA Behavior Using World Models
Parsa Mastouri Kashani, Jan-Gerrit Habekost, Stefan Wermter
cs.RO · cs.AI
Abstract
Vision-Language-Action models (VLAs), while showing strong potential for robot control, require massive amounts of high-quality imitation learning data. Moreover, the absence of an explicit world model casts further doubt on their control capabilities. In this concept paper, we propose a novel architecture that addresses sample efficiency in VLAs by training a predictive world model on the embedding space of the VLA's vision encoder. We hypothesize that these embeddings are action-relevant and usable for future prediction. To this end, we propose using the suggested architecture to investigate how well these embeddings predict the future based on actions, as the inability to do so would mark a key limitation of VLA architectures: the lack of a non-lossy implicit world model to simulate real-world dynamics. The proposed architecture differs from the standard world model dynamics as the loss comes from the embedding space rather than the pixel space, similar to joint embedding predictive architectures. Furthermore, the trained world model can be utilized for short-term planning tasks by sampling VLA actions given goal images. We intend to examine the richness of vision embeddings in VLAs and reduce their high data requirements through a world model that can also generate plans during inference.
cs.AI / 100 / 2609.30832
Subject-Invariant Cross-Modal Decoding of Perceived Speech from Brain Recordings
Aoke Zhang, Jing Chen
cs.SD · cs.AI · eess.AS
Abstract
Perceived speech decoding based on non-invasive brain-computer interface (BCI) signals has been extensively studied in recent years. Research in this field primarily faces two challenges: extracting neural representations with rich spatiotemporal information and achieving cross-subject generalization. Although separate studies have proposed methods to cope with these issues, a unified approach that simultaneously tackles both challenges remains lacking. To fill this gap, we propose the Subject-Invariant Cross-Modal Perceived Speech Decoding (SICMD) method, which integrates functional magnetic resonance imaging (fMRI) and magnetoencephalography (MEG). We conduct comprehensive analyses of the fusion method, fusion position, encoder architecture, and model inputs. Our results demonstrate that the proposed method improves Top-1, Top-10, and Rankacc by more than 10.6%, 10.1%, and 1.7%, respectively, compared to baseline methods in cross-subject perceived speech decoding tasks, while reducing training costs by 88.8% and 60.5% compared to multi-subject and intra-subject decoding settings. Further visualization experiments also confirm the effectiveness of our approach.
cs.AI / 101 / 2609.31180
BAT-CLIP: Trimodal Alignment of Brain, Audio and Text
Suhyun Kim, Jinmo Han, Danny Dongyeop Han, Ahhyun Lucy Lee, Jewoon Lee, Yonghyeon Gwon, Zach Paris, Chun Kee Chung, Saewoong Bahk, Nam Soo Kim, Seong Jae Hwang, Jiook Cha
cs.SD · cs.AI · cs.LG · eess.AS
Abstract
Decoding and interpreting naturalistic speech from the brain increasingly relies on alignment to pretrained speech and language representation spaces. However, current CLIP-style brain-speech alignment ground neural activity to a single anchor modality-audio or text-despite the brain's inherently multimodal speech processing. This induces a trade-off: audio anchoring preserves temporal structure but weakens linguistic separability, while text anchoring captures semantics yet discards acoustic detail. We propose BAT-CLIP, the first CLIP-style trimodal alignment framework for iEEG that jointly aligns neural embeddings to both pretrained audio and text anchors in a shared, frozen audio-text manifold. On the naturalistic Podcast benchmark, BAT-CLIP yields more robust representations than bimodal CLIP baselines. We also highlight the importance of using self-supervised foundation models for CLIP training.
cs.AI / 102 / 2609.31224
Acoustic-to-Text KV Compression for Full-Duplex Speech Models
Yejin Lee, Seungbeom Kim, Yongha Lee, Kyuhong Shim
cs.SD · cs.AI · eess.AS
Abstract
Full-duplex speech language models continuously accumulate acoustic key-value (KV) states, making long-running interactions memory-intensive. During listening, the model can finish processing an audio unit before the next arrives; we term the remaining interval listening-time slack. We propose acoustic-to-text KV compression, which introduces a transcription side channel to convert incoming speech into compact textual memory within this interval. When the cache exceeds a target budget during inference, older acoustic states are evicted while transcripts and recent acoustic context remain. We train the side channel with LoRA using cross-entropy on transcription segments. To preserve listening and speaking behavior, we apply knowledge distillation to the original model's token-level output distributions at native prediction positions. On ten-minute LongSpeech sessions, our MiniCPM-o 4.5 implementation reduces peak streaming KV-cache size by 64.6% compared with the same model without eviction. The proposed method also improves transcription, temporal question answering, and summarization over the baseline. Full-Duplex-Bench evaluations further show comparable pause-handling, turn-taking, and interruption performance.
cs.AI / 103 / 2609.31468
PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents
Pavel Kireyev
econ.GN · cs.AI · cs.CL
Abstract
LLMs increasingly act as purchasing agents, which makes the LLM, not the user, the one choosing among the options that satisfy a request; its preferences quietly fix what gets bought and what it costs. Hotel booking is a clean instance: a high-volume choice settled on a few comparable attributes, where the pick reveals those preferences. We introduce PriceBench, a diagnostic benchmark that recovers an LLM's price, quality, and brand preferences from its booking choices with a logit choice model, applied to 28 LLMs from 8 providers on 3,600 hotel tasks from 179 real New York City properties. We find that capability is associated with how consistently an LLM chooses, not with what it chooses: more capable LLMs hold stronger, more consistent preferences, while weaker ones either lock onto one position, exploitable by whoever controls listing order, or choose almost indifferently. What those preferences favor varies sharply across providers and even within one family: price sensitivity spans more than an order of magnitude, and the price/quality trade-off moves mean booked nightly price from \$247 to \$393 on identical tasks. What an agent buys must therefore be measured per LLM, not inferred, and we release the tasks, code, and all 28 response sets.
cs.AI / 104 / 2609.30891
Adaptive Pilot Selection for Unified Semantic Communication and Semantic Sensing in ISAC
Muhammad Abubakar Rashid, Muhammad Hannan Akram, Haejoon Jung, Syed Ali Hassan
eess.SP · cs.AI
Abstract
Semantic communication (SemCom) and integrated sensing and communication (ISAC) are promising technologies for future 6G wireless networks. Existing studies have applied semantic technology to either the communication module or the sensing module of ISAC. In this work, we propose SemISAC, which performs both SemCom and semantic sensing within a single dual-function waveform. SemISAC uses a joint semantic encoder that extracts task-specific information for both communication and sensing. We evaluate SemISAC in a vehicular scenario in which vehicles share pixel-wise segmentation of the road environment and, through sensing, classify surrounding objects and estimate their ranges. On the transmitter side, a deep learning encoder converts the input road-scene image into semantic symbols and places them on the data cells of an OFDM grid, while the remaining cells serve as pilots for channel state information estimation and sensing. The pilot configuration is adaptively optimized based on the channel conditions to balance communication and sensing requirements. At the receiver, a deep learning model reconstructs the segmentation from the received waveform, while the transmitting vehicle captures the reflected waveforms from surrounding objects and uses task-specific deep learning decoders for target recognition and range estimation. Simulation results show that SemISAC achieves a segmentation accuracy close to that of the dedicated SemCom module while outperforming both conventional and semantic baselines in target recognition and range estimation.
cs.AI / 105 / 2609.30546
Convergence guarantees for Muon: New parameter regimes and generalizations
Arthur C. B. de Oliveira, Dhruv D. Jatkar, Guilherme S. Vicinansa, Eduardo D. Sontag
math.NA · cs.AI · math.OC
Abstract
In this paper, we establish the first asymptotic convergence guarantees for the Muon algorithm through a more accurate proxy for the Newton-Schultz iteration than the typical matrix sign function. We prove that, for appropriate choices of hyperparameters, the iterates satisfy $\lim_{k\to\infty}\|\nabla f(x_k)\|=0$, and, under a global Polyak-Łojasiewicz condition, that the sequence of function values converges linearly. The key insight is that the regularization, implicit in Muon's Newton-Schulz implementation, induces a bounded preconditioner, exposing Muon as a \emph{preconditioned Polyak heavy-ball} method and enabling a classical Lyapunov analysis. This observation naturally motivates applying the same preconditioning structure to the Nesterov gradient evaluation. We formalize this idea by introducing \emph{Muesterov}, a Nesterov-based variant of Muon, and prove that it enjoys the same convergence guarantees, extending the theoretical framework beyond the heavy-ball setting. Numerical experiments on a scalar cross-entropy problem corroborate the theory and illuminate the joint role of the learning rate and the Newton-Schulz regularizer in controlling convergence. Preliminary numerical simulations training the nanoGPT dataset provide intuition regarding the relevance of the observations in this paper to practical applications.
cs.AI / 106 / 2609.30420
Understanding Perturbed Parameter Ensemble Sensitivities Using A Contrastive Learning Approach
Da Fan, David John Gagne, Gregory S Elsaesser, Brian Medeiros, Addisu G Semie, Qingyuan Yang, Akila Sampath, Subashree Venkatasubramanian
physics.ao-ph · cs.AI
Abstract
Perturbed parameter ensembles (PPEs) reveal how physics parameters affect climate simulations, but interpreting parameter sensitivities across multivariate, spatially structured outputs remains challenging, particularly when calibrating models against observations. We develop an explainable contrastive learning model that maps 5 monthly cloud and radiation fields into a shared representation space. We train the model on the fields of two 100-member Community Atmosphere Model version 6 (CAM6) PPEs, spanning 34 parameters, that only differ in the warm rain microphysics scheme: KK2000, the default bulk microphysics scheme, and TAU-ML, a neural network emulator of a bin microphysics scheme. The learned representations separates two PPEs with over 94\% linear classification accuracy while preserving the seasonal variability and ensemble spread due to parameter perturbations. In the shared representation space, the representations of satellite observations occupy the same low-dimensional manifold as the PPEs but are displaced from them most strongly during boreal spring and autumn. TAU-ML PPE has a lower distance to observations compared to KK2000 in the representation space. Integrated Gradients attributions highlights the contributions in subtropical low-cloud regions, Northern and Southern Hemisphere storm track regions, and tropical convection regions to differences between PPEs and observations. Regional attributions correlate most strongly with parameters associated with cloud microphysics, boundary layer turbulence, and deep convection. These results demonstrate that explainable representations of climate fields can attribute model differences to specific variables, regions, seasons, and physical parameters.
cs.AI / 107 / 2609.30883
Warned alike, AI agents avoid the less-crowded road while people take it
Takahiro Ezaki, Naoto Imura, Katsuhiro Nishinari
physics.soc-ph · cs.AI
Abstract
AI agents built on a few shared models increasingly act for many people. A shared forecast about others can align their choices and change how scarce capacity is allocated. We tested this feedback in a two-road congestion game. Adding one sentence warning that others might follow a routing tip made populations of 50 GPT agents crowd one road while avoiding the nearly empty alternative. Average travel time rose from 64 to 95 min, although any crowded-road agent could have saved 69 min by switching alone. The warning discouraged the very move it predicted. The pattern persisted for 100 rounds. Two other model families shifted the same way without locking onto one road. Twelve all-human groups (240 participants) stayed near balance under numerical reports or the tip and warning. In 24 mixed groups with a further 240 participants, imbalance grew with the share of agents in the registered analysis, while people increasingly took the road the agents avoided. Collective costs stayed below the allagent reference, but with 15 agents and 5 humans, agent seats averaged 80 min, compared with 44 min for human seats. Shared forecasts can thus sustain collective inefficiency among similar agents. A better group average can also hide an unequal burden. Evaluations of AI agents that share resources should test populations, treat messages as interventions and report who bears the costs.
cs.AI / 108 / 2609.31260
Agentic Limit Order Books: Phase Transitions and Market Impact
Jan Rosenzweig
q-fin.TR · cs.AI · q-fin.CP
Abstract
We investigate the systemic macroscopic dynamics emerging from Limit Order Books (LOBs) populated exclusively by autonomous reinforcement-learning agentic traders. By formalizing agent interactions within a microscopic order-matching engine, we examine two fundamental quantitative phenomena: equilibrium phase transitions in order flow regime shifts, and the structural dynamics of market impact. We show that agentic LOBs exhibit distinct phase boundaries separating orderly price discovery from hyper-volatile cascade states, governed by critical thresholds in the number of agents and observable market depth. Furthermore, we demonstrate that market impact under agentic liquidity provision deviates from classical square-root dynamics, exhibiting distinct dissipative, balanced, and non-dissipative regimes under non-linear feedback loops.
cs.AI / 109 / 2609.31607
Statistical attribute alignment for black-box generative AI via output post-processing
Kevin Jiang, Morgane Austern, Edgar Dobriban, Jason M. Klusowski
stat.ME · cs.AI · cs.LG · math.ST
Abstract
Generative AI systems are increasingly used, but aligning their outputs with user requirements poses a continuing challenge. Here, we aim to ensure that the distribution of an attribute of an AI-generated output aligns with a user-specified target. This is motivated by examples such as fairness, where we want to ensure that a protected attribute (e.g., gender, race, or age categories) follows a desired distribution, and synthetic data generation, where we want the generated data to be representative of a target distribution. We study the practically important black-box access setting, where a user can repeatedly query a generative AI model. The goal is to return $m\ge 1$ outputs whose joint attribute distribution is as close as possible to this target. For both exact and approximate alignment, we develop algorithms that minimize the expected number of queries to the generator, and we further demonstrate their optimality as the number of requested outputs $m \rightarrow \infty$. Experiments on text-to-image generation and geocoded persona generation tasks show that our post-processing algorithms improve statistical attribute alignment, complementing prompting-based interventions.
机器学习 (cs.LG)
125
cs.LG / 1 / 2609.31101
A Flatness-Generalization Relation in the Teacher-Student Tree-Committee Machine
Brandon Livio Annesi, Davide Straziota, Enrico Maria Malatesta
cond-mat.dis-nn · cs.LG · stat.ML
Abstract
The flatness of the loss landscape at a minimizer is a widely used heuristic for reasoning about neural-network generalization, yet evidence for this relation is mostly empirical and controversial. We study this relation in a teacher-student tree committee machine, where both the ERM estimator and the Hessian spectrum are analytically tractable in the proportional high-dimensional limit. First, we use a zero-temperature Gibbs formulation to obtain predictions for the observables of the typical minimizers of the empirical loss. Secondly, we use Edwards-Jones formalism to derive the limiting Hessian resolvent around these typical minimizers. All predictions agree with finite-size gradient-descent simulations. Finally, we study three measures of flatness, namely the left and right edges and the spectral mean, and check if a decrease in generalization error as the dataset size is increased corresponds to an increase in flatness. We find that the answer strongly depends on the learning task and on the ratio of the number of parameters to the number of data points. In regression, the spectral mean and right edge correlate with the generalization error, while the left edge does so only in the overparametrized regime. In classification this correlation reliably holds only in the highly overparametrized phase, while for underparametrized networks it can even reverse.
cs.LG / 2 / 2609.31518
Retrainable physics-integrated neural differentiable modeling of sintering across material systems
Zeping Chen, Ani Aprahamian, Khachatur V. Manukyan, Tengfei Luo
cond-mat.mtrl-sci · cs.LG
Abstract
Sintering is widely used to manufacture ceramics, but coupled densification and grain growth, material-dependent kinetics, and sparse measurements complicate predictive modeling and process design. We present Sinter-PiNDiff, a retrainable physics-integrated neural differentiable framework for predicting density and grain-size evolution. Two neural networks learn densification and grain-growth coefficients within coupled rate equations, while a smooth saturation factor attenuates densification near theoretical density. The same governing structure, network architecture, and training procedure were fitted independently to published data for MgO, Al-doped ZnO, and CaO-doped ThO2. Tests at held-out temperatures and compositions yielded the lowest mean error in all twelve material-metric comparisons against multilayer perceptron and residual network baselines. For MgO, Al-doped ZnO, and CaO-doped ThO2, respectively, density normalized root-mean-square errors were 14.6%, 10.8%, and 14.4%, and grain-size errors using the same metric were 8.6%, 12.1%, and 19.3%. Removing evolving density from both neural-network inputs increased density and grain-size trajectory errors in all three systems and ten of twelve aggregate errors, supporting density-dependent kinetic feedback. Deep ensembles estimated model disagreement, but empirical coverage showed that the uncertainty bands were not calibrated and did not capture all model-data discrepancies. These results establish Sinter-PiNDiff as a retrainable framework for sparse-data prediction and uncertainty-informed selection of sintering conditions.
cs.LG / 3 / 2609.30667
StarWM: Self-Supervised Trained Attention Routing for Robust World Models
Zeqiang Zhang, Fabian Wurzberger, Maximilian Otte, Daniel Schmid, Sebastian Gottwald, Arne Peter Raulf, Daniel Alexander Braun
cs.CV · cs.LG
Abstract
A robust world model must strike the balance between faithfully capturing environmental dynamics and abstracting away from irrelevant content. While reconstruction-based world models ensure faithful supervision, they misallocate representational capacity by pixel area rather than dynamics relevance for visual tasks, which can cause task-irrelevant content to dominate the learned representation. Alternatively, reconstruction-free methods avoid this bias but risk discarding possibly relevant information. We propose StarWM, which uses a cross-attention module trained on self-supervised dynamics to decide where reconstruction applies. A dual-stream decoder then restricts reconstruction to the attended regions, with stop-gradient barriers preventing interference between the two objectives. These components allows reconstruction to supervise the visual content of attended regions without contaminating the latent with non-predictive information. On DeepMind Control with dynamic video backgrounds, default (reward-free) StarWM achieves the strongest performance under random-frame distractors and substantially outperforms reconstruction-based baselines under sequential video. In addition, its reward-augmented variant matches or exceeds reconstruction-free methods on sequential video, achieving the highest overall return across all distractor regimes. Mechanistic probing confirms StarWM preserves state attributes with near-perfect fidelity through long-horizon imagination while systematically discarding distractors.
cs.LG / 4 / 2609.30769
Query-Conditioned Prototype Adaptation for Cross-Domain Few-Shot Learning: Single-Query Inference, Controlled Comparisons, and Failure Modes
Rushab Rasik Karania, Tomas Maul
cs.CV · cs.LG
Abstract
Cross-domain few-shot learning requires adapting a classifier to a new visual domain from very few labelled examples without target-time parameter updates. We isolate one question: under a fixed global representation, what does joint query-support adaptation contribute to prototype construction? The Within-Instance Prototypical Transformer (WIPT) implements single-query test-time prototype adaptation by jointly transforming one unlabelled query and the labelled support embeddings, then forming query-specific class means. Using a shared frozen ViT-S/16 encoder, miniImageNet source training, and CUB, EuroSAT and ISIC targets, we replicate the key comparisons across five independent training seeds. In 1-shot evaluation, WIPT improves frozen ProtoNet in every run on CUB (+0.21 percentage points) and EuroSAT (+2.07), but decreases ISIC (-0.22). In 5-shot evaluation, ProtoNet remains strongest overall, while WIPT consistently improves a capacity-matched support-only Transformer on ISIC (+0.99). Joint processing of up to five queries yields no reliable accuracy gain; in a head-only 5-shot benchmark, g = 5 reduces analytical attention-token pairs by 73% and peak allocated memory by 29% relative to g = 1, although latency is non-monotonic. Across all target/shot conditions, WIPT changes uncertain ProtoNet decisions far more than confident ones, and rescue/break decomposition accounts for the observed gains and losses. Source-shift and scorer controls further show that the benefit is not universal. Overall, WIPT provides a streaming-compatible form of test-time prototype adaptation that can improve difficult low-shot cross-domain decisions without target-time optimization.
cs.LG / 5 / 2609.31356
Open Vocabulary Domain Unlearning
Sumanth Udupa, Mehrtash Harandi, Yadan Luo, Mahsa Baktashmotlagh
cs.CV · cs.LG
Abstract
Vision-Language Models (VLMs) exhibit remarkable zero-shot generalization, yet they often encode unwanted or hazardous stylistic domains such as idealized textbook diagrams in medical AI or cartoon vehicles in autonomous driving. Approximate Domain Unlearning (ADU) aims to selectively erase a model's recognition of a target visual domain while preserving accuracy on the remaining domains. However, existing ADU methods operate under a flawed closed-vocabulary assumption: they evaluate unlearning solely on the specific object classes seen during the unlearning fine-tuning phase. Consequently, these methods do not unlearn the domain itself; they merely overfit to seen class-domain pairs, leaving the domain easily recognizable for unseen classes and providing a false sense of removal. We argue that true domain erasure must be class-agnostic. To address this, we formalize Open-Vocabulary Domain Unlearning (OVDU), a rigorous protocol that mandates domain forgetting must transfer to held-out classes. To solve the OVDU challenge, we propose a surgical parameter-editing framework. First, a Fisher Information mask isolates domain-sensitive weights, mathematically protecting foundational zero-shot generalization. Second, our Targeted Manifold Scattering (TMS) objective uses preference-based mining to locally scatter the forget domain's stylistic geometry. Evaluated across PACS, OfficeHome, and DomainNet, our method vastly improves open-vocabulary generalization over existing baselines. Crucially, it delivers exceptional sample efficiency, outperforming peak 8-shot baseline results with only 4 shots.
cs.LG / 6 / 2609.30854
The KV Cache Is the New Memory Wall
Tejinder Singh
cs.DC · cs.LG · cs.PF
Abstract
Autoregressive LLM inference at long context is bounded by memory bandwidth, not arithmetic throughput, and the binding resource shifts from model weights to the Key-Value (KV) cache as sequence length grows. For Llama-3-70B in BF16, the 140 GB weight footprint exceeds the 80 GB HBM of a single accelerator, and one 128k-token sequence adds 42 GB of KV cache. Techniques that compress, evict, page, share, or offload KV state have proliferated, but reported gains use inconsistent workloads, hardware, and quality metrics, preventing cross-paper comparison. This SoK paper unifies the field analytically, with a protocol that strictly separates derived and reported claims. We derive closed-form arithmetic intensity as a decaying function of context length, parameterized by hardware topology for NVIDIA H100, NVIDIA B200, and AMD MI300X, including per-die bandwidth partitioning and the crossover lengths where KV traffic overtakes weight traffic. We classify the literature into five domains, quantization, token eviction, KV paging, prefix caching, and heterogeneous tiering, evaluating one method per domain at 128k context under a single protocol. The central finding is a three-regime structure: below a hardware-specific crossover, weight traffic dominates and KV compression yields negligible speedup; beyond it, KV traffic dominates and each domain trades quality for bandwidth savings approaching the roofline bound. Paging and prefix sharing are lossless but address capacity, not bandwidth. Quantization and eviction cut bandwidth directly, with degradation that accelerates below 4-bit precision and turns discontinuous for eviction on position-sensitive tasks. Tiering converts the bandwidth wall into an interconnect problem bounded by PCIe or NVLink rather than HBM. We close with design rules for selecting a compression domain given hardware, context length, and quality budget.
cs.LG / 7 / 2609.31614
Gap-free Differentially Private PCA for Gaussian Data
Alina Ene, Huy L. Nguyen
cs.DS · cs.LG
Abstract
We give a gap-free differentially private algorithm for the principal component analysis (PCA) problem with Gaussian data.
cs.LG / 8 / 2609.30753
Skill Profiling with Attributable Reasoning (SPAR): A Wearable Analysis System for Boxing
Nibraas Khan, Hanchen David Wang, Enya Bullard, Ritam Ghosh, Ruj Haan, Aarav Agrawal, Meiyi Ma, Nilanjan Sarkar
cs.HC · cs.LG
Abstract
A punch is a ballistic, full-body action driven by a kinetic chain running from the legs through the trunk to the arm, where a small sequencing error separates a scoring strike from a miss. Wearable sensors can capture this movement in the gym, but most deployable systems only classify which punch was thrown rather than assess how well it was thrown. We present Skill Profiling with Attributable Reasoning (SPAR), an eight-IMU garment and pressure-insole system that classifies each punch as expert or novice and treats an explanation of that prediction as feedback. Feedback is only useful if the person receiving it can act on it, so SPAR explains the prediction at three tiers, a per-joint attribution for the analyst, a counterfactual over kinetic-chain layers for the coach, and a plain-language narrative of the two for the athlete. Across 17 participants and 4,713 punches, SPAR reaches a leave-one-participant-out AUC of 0.842 (95% CI [0.769, 0.907] over participants). A frozen time-series foundation model encodes the joint-angle and plantar-force series, and a small transformer trained on the cohort classifies the encoding. We audit the two quantitative tiers and report six themes from a thematic analysis of interviews with six practicing boxing coaches.
cs.LG / 9 / 2609.31498
Retail Product Search: A Practical Approach at Target
Darshan Sonagara, Qujiaheng Zhang, Ankit Singh, Alex Li
cs.IR · cs.LG
Abstract
Search is one of the most important features in e-commerce, directly driving customer engagement and business growth. A good product search system must show both relevant and desirable results. However, retail search presents unique challenges. User intent can range from exact matches to open-ended discovery. Search systems must also balance multiple goals, such as relevance, revenue, and profit, while keeping response times low. Traditional keyword-based methods often fall short in handling natural language or semantic queries. Vector search helps alleviate these issues, but it can miss key intent signals or return low-precision results. In this paper, we present the design of a hybrid search system at Target that combines lexical and vector search. We describe our approach to data processing, embedding training, precision control for the final result set, multi-channel result fusion (where we compared fusion strategies and adopted weighted interleaving), and the performance optimizations used to maintain low latency for production deployment. Our method improves offline evaluation metrics, and in online A/B testing it raised click-through rate by 0.97%, order conversion by 0.98%, and demand per visitor by 1.10% over lexical-only search, while roughly halving zero-result searches. The resulting system is deployed at scale and serves millions of guests daily.
cs.LG / 10 / 2609.30379
DanLing NestedTensor: Composable Multi-Ragged Tensors for Deep Learning
Zhiyuan Chen
cs.LG · cs.AI · cs.PF
Abstract
Variable-size inputs are common in deep learning, but dense batching allocates a shared envelope and spends computation on padding. The cost multiplies across varying axes: an explicit pair state allocates $BN_{\max}^2$ positions instead of $\sum_i N_i^2$. Packing removes that waste, but composing packed operations still requires the logical axes and sample boundaries a flat buffer no longer exposes. We present DanLing NestedTensor, a PyTorch tensor abstraction that makes multi-ragged structure a property of the tensor itself. Packed values carry tensor-backed partitions and logical dimension order, so broadcasting creates ragged axes, feature transformations retain them, and reductions consume them. The same representation carries through autograd and both eager and compiled execution. On an A100, the geometric-mean speedup over same-mode padding is 2.74$\times$ eager and 3.39$\times$ compiled across four BERT scales, and 1.97$\times$ eager across four FCN backbones. A four-block Pairformer-style workload runs 2.40-4.32$\times$ faster than a padded reference using native PyTorch kernels across square length regimes in eager execution, with peak allocation falling from 38.08 to 5.41 GiB on its high-variation batch. The tensor interface lets model code built from its supported operators compose efficient variable-size computation without managing offsets at any call site. Code will be released publicly upon publication.
cs.LG / 11 / 2609.30391
From Weak Data to Strong Policy: Q-Targets Enable Provable In-Context Reinforcement Learning
Yichen Lin, Xuyuan Xiong, Xue Wang, Xiangfu Meng, Mike Mingcheng Wei, Tao Yao
cs.LG
Abstract
Existing in-context reinforcement learning methods mainly pretrain Transformers with supervised behavior-prediction objectives. This enables task inference from context, but makes the learned policy strongly depend on the quality of offline actions: when trajectories are weak or suboptimal, imitation itself becomes a biased learning signal. We propose Q-Target Pretrained Transformers (QTPT), which keeps the context-conditioned Transformer architecture but replaces behavior cloning with a Bellman-style Q-target objective. QTPT therefore learns to use rewards and transitions in the context to estimate action values, rather than simply imitating the behavior policy. We theoretically analyze QTPT in stochastic linear bandits and finite-horizon MDPs, showing stronger robustness to data quality than supervised pretraining. Empirically, QTPT improves over supervised behavior prediction on controlled RL benchmarks with random or suboptimal data, and we examine extensions to D4RL Kitchen and AntMaze. Supplementary experiments evaluate backbone robustness, meta-RL comparisons, task-coherent context, and unsupported-action value overestimation. These comparisons distinguish the benefits of Q-target pretraining from the remaining limitations of offline coverage.
cs.LG / 12 / 2609.30417
Electric Vehicle Charging Station Location Selection using Geospatial Artificial Intelligence (GeoAI)
Eun Hak Lee, Euntak Lee
cs.LG
Abstract
As electric vehicle (EV) adoption increases, ensuring efficient and well-distributed charging infrastructure has become a critical challenge. While many EV charging station location problem (CSLP) studies focus on minimizing costs or travel distance, it is crucial to consider the surrounding geospatial characteristics of existing stations that influence operational performance. This study proposes a geospatial artificial intelligence (GeoAI)-based framework that integrates high-dimensional EV-related geospatial data, including EV usage, land-use, population, and traffic attributes. We incorporate a variational autoencoder (VAE) and a graph convolutional network (GCN) into the model to capture similarities among existing charging stations, and to identify suitable locations for future stations. The VAE compresses high-dimensional EV input data into a low-dimensional latent space, and the GCN uses this latent representation to predict locations suitable for charging stations. Using real-world data from Bryan-College Station, Texas, US, the proposed model outperforms state-of-the-art baselines, achieving an F1-score of 0.87 in distinguishing existing station locations from non-station locations. The model also identifies 27 additional candidate locations that show geospatial characteristics similar to those of existing stations, based on a similarity score. We further evaluate two policy implementation scenarios, maximizing geospatial similarity and minimizing total travel distance, each yielding different outcomes aligned with distinct strategic objectives. The findings highlight the importance of incorporating spatial context into CSLP and provide valuable insights for future EV infrastructure planning, promoting both efficiency and accessibility in the rapidly growing electric mobility sector.
cs.LG / 13 / 2609.30433
Improving Molecular-Morphology Contrastive Pretraining using Deep-Learning-based Morphology Profiles
Jie Li, Kathryn E. Kirchoff, Dante A. Pertusi, Zhizhuo Zhang
cs.LG · q-bio.BM
Abstract
Recent advancements in image-based profiling techniques have enabled the collection of high-volume cell morphology data, allowing new molecular embedding models to learn from the experimental phenotypic perturbations of a molecule in a cell. Previously, we developed Molecule-Morphology Contrastive Pretraining (MoCoP), a strategy for aligning small molecule embeddings to morphology fingerprints extracted through CellProfiler. The resulting molecular representation showed transferable performance for quantitative structure--activity relationship (QSAR) prediction tasks. Here, we extend the method by using a deep-learning-based cell image encoding pipeline to extract more feature-rich morphology profiles and align them to the molecular embeddings through contrastive learning. The new embeddings encode more accurate information on how molecules perturb cell morphology and enable improvements for QSAR predictions through either fixed-embedding linear probes or fully flexible fine-tuning. Morphology retrieval performance scales log-linearly with training data size, suggesting continued improvements as larger datasets become available. The improved MoCoP v2 also achieves superior performance on toxicity prediction and competitive results on ADME and activity benchmarks, when compared with existing molecular embedding models that use both cell morphology and transcriptomic data during training.
cs.LG / 14 / 2609.30454
Auditing System-1 Models on Biosecurity-Relevant Benchmarks: Calibration, Selective Prediction, and Permutation Instability in a Non-Generative Model
Kimon Antonios Provatas, Ilias Georgakopoulos-Soares
cs.LG
Abstract
Non-generative "System-1" models return structured probabilistic decisions in a single forward pass, without autoregressive decoding, at a small fraction of the inference cost of a generative model. This makes them of interest as inexpensive components in larger pipelines, but their reliability on biosecurity-relevant tasks has not been systematically examined. We audit one commercial System-1 model on 6,020 multiple-choice items drawn from the Weapons of Mass Destruction Proxy (WMDP), a paraphrase-robust WMDP-Bio variant, and six LAB-Bench subtasks, measuring accuracy, calibration, error detection, selective prediction, and sensitivity to the order in which answer options are presented. Accuracy is strongly task-dependent. Once the vendor's uncertainty field is correctly interpreted, the model is reasonably well calibrated (pooled expected calibration error 0.034) and its top-1 probability separates correct from incorrect predictions (pooled AUROC 0.820), though both degrade substantially on the weaker tasks. Under four cyclic rotations of the answer options, 37.4% of WMDP-Cyber items receive different answers; a control using byte-identical repeated calls attributes most of this to option order rather than run-to-run variation. Averaging probabilities across rotations improves WMDP-Cyber accuracy by 3.8 percentage points, and applying it only to low-confidence items recovers most of that gain at well under the cost of averaging every item.
cs.LG / 15 / 2609.30465
RAZOR: Pruning Replaceable Experts in LLMs
Mingyang Song, Mao Zheng
cs.LG · cs.CL
Abstract
Mixture-of-experts (MoE) models activate few experts per token but store the full expert pool. Expert pruning reduces this storage burden; at a fixed pruning budget, the goal is to preserve the original model's output distribution as closely as possible. Yet an expert's usage or contribution magnitude does not by itself determine the damage caused by its removal. What matters is whether the surviving computation can replace its function. We introduce RAZOR, a training-free expert pruning method that scores functional replaceability using consensus residuals: deviations of expert outputs from the original weighted mixture. An exact single-deletion identity at a fixed layer input accounts for survivor renormalization and router-selected refill, providing local scores aggregated over calibration tokens for budgeted pruning without gradients or recovery training. On GLM-4.7-Flash, Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, and Hy3 at 25\% and 50\% expert removal, RAZOR achieves the highest nine-task macro average among the evaluated pruning methods in all eight settings. On the two backbones with matched REAP benchmark runs, it exceeds REAP by 2.12--5.59 points and wins all 36 paired task comparisons. It also lowers reverse KL relative to REAP in all four matched GLM-4.7-Flash and Qwen3.6-35B-A3B model--budget settings. Analysis of responses generated by Qwen3.6-35B-A3B nevertheless reveals changes in diversity, formatting, and termination, underscoring that task retention and predictive fidelity do not ensure generation stability.
cs.LG / 16 / 2609.30470
Reliability-aware Cross-sample Enhancement for Robust Multimodal Sentiment Analysis
Menghua Jiang, Haokai Gao, Xiangui Kang, Haifeng Hu, Sijie Mai
cs.LG
Abstract
Multimodal Sentiment Analysis (MSA) aims to infer human emotions from multiple modalities such as text, audio, and vision. In practice, inputs are often corrupted by noise and missing modalities, which degrades performance. Existing methods typically address these challenges in isolation, limiting their effectiveness in realistic settings. To address this limitation, we propose a Reliability-aware Cross-sample Enhancement (RCE) framework. Specifically, RCE first introduces an adaptive variational information bottleneck to model modality-wise uncertainty and perform quality-aware information compression, thereby suppressing redundant noise in unreliable modalities. Furthermore, we design a reliability-aware cross-sample enhancement strategy that retrieves high-confidence, semantically consistent neighbors from a large candidate pool to enrich and calibrate current representations, effectively alleviating information deficiency caused by missing modalities. Building upon this, RCE integrates cross-modal interactions with a multilevel reliability-aware fusion mechanism to adaptively aggregate information across modalities and enhancement stages, leading to more robust multimodal representations. Extensive experiments demonstrate that RCE consistently outperforms state-of-the-art methods across full, noisy, and missing-modality settings.
cs.LG / 17 / 2609.30472
Moment-guided edge sampling
Weibin Cai, Reza Zafarani
cs.LG · cs.SI
Abstract
Edge sampling makes local decisions to achieve graph-level objectives, such as preserving structural properties. This creates a fundamental challenge: \textit{how can the effect of a local edge edit (i.e., edge addition or removal) on global graph structure be quantified and controlled?} We address this challenge with a \textit{moment-guided edge sampling framework} based on spectral moments of the random-walk transition matrix. We compute exact moment changes through two complementary methods: a combinatorial method with closed-form updates for low-order moments, and a low-rank method that exploits \textit{locality} and \textit{cyclic trace invariance} to compress computations to edited endpoints, supporting arbitrary moment orders and batched edits. For single-edge edits at fixed moment orders, the low-rank method reduces the cost from $O(mn)$ to $O(m)$, while the combinatorial method evaluates low-order changes in constant time given maintained local statistics. These moment changes provide \textbf{interpretable structural signatures} of local edge motifs that aggregate into graph-level fingerprints. This structural meaning motivates us to ask whether preserving moments also preserves the graph properties. We further derive and validate that moment-preserving sampling can \textbf{retain related structural properties}, including triangle-weighted clustering coefficient. These structural insights enable \textbf{analysis and improvement of graph learning}: different edge structures have distinct effects on supervised node classification, while moment-guided augmentation is competitive for graph contrastive learning. Together, these findings establish moments as an interpretable and controllable bridge from local edge edits to global graph structure and learning.
cs.LG / 18 / 2609.30474
Mentored Decoding: Faster Inference meets Boosting
Vivien Tran-Thien, Richard Nock
cs.LG
Abstract
Speculative decoding is a successful technique speeding up inference of a target autoregressive language model via a fast drafter model. Lossy speculative decoding allows a drift with respect to the target to further improve speed. Interestingly, it has been observed experimentally that the resulting model can $\textit{also}$ beat the target $\textit{quality-wise}$. Our paper formally proves how such a feat is possible with a formal approach to lossy speculative decoding called $\textit{mentored decoding}$. To get there, we connect inference to a celebrated ML training theory, $\textit{boosting}$, and proceed via the generalization of mentored decoding to the whole set of $f$-divergences. We uncover key properties of mentored decoding, among which (i) the particularly appealing geometric nature of the total variation case, (ii) simple approximations for any $f$-divergence in direct relation with boosting compliance, and (iii) a $\textit{divergence independent}$ $O(n)$ space and $O(\mathrm{sort}(n))$ time data structure built on drafter and target outputs, which allows to query the optimal parameters of the dual problem in $O(\log n)$ time and constructing optimal mentored distributions in $O(n)$ time for any $f$-divergence.
cs.LG / 19 / 2609.30487
Geometric Feature Learning for Functional Data Valued on the Symmetric Positive Definite Manifold
Samuel V. Singh, Mimi Zhang
cs.LG · stat.ML
Abstract
We here develop a functional neural network, termed MatFAE, for learning trajectories on the Riemannian manifold of symmetric positive definite (SPD) matrices. MatFAE features intrinsic layers that map manifold-valued functions to Euclidean vector-valued functions, followed by a functional layer that projects them into a finite-dimensional Euclidean space. Unlike most neural networks for discrete-time sequences, MatFAE treats each sequence as a continuous function and can therefore encode trajectory dynamics (e.g., first-order derivatives) in its latent representations. Additionally, the morphology of the functional weights in the functional layer offers interpretability by revealing the regions of the input functional data that contribute most to the latent representations. We justify the design principles and properties of each intrinsic layer and detail how matrix factorization is handled during backpropagation. We apply MatFAE to a range of fMRI datasets, demonstrating its ability to efficiently learn informative representations from high-dimensional SPD trajectories and its practical value for real-world neuroimaging analysis.
cs.LG / 20 / 2609.30498
Learning to Bias: Machine Learning-Enhanced Particle Filters
Apoorv Srivastava, Eric Darve
cs.LG · stat.ML
Abstract
Sequential inference estimates latent states from noisy and incomplete observations. Particle Filters (PFs), a class of Monte Carlo methods based on importance sampling, provide a flexible framework for this task, but often suffer from poor sample efficiency and unfavorable scaling with dimension, partly due to suboptimal proposal distributions. We address these challenges by integrating learned proposals into the PF framework. We introduce Neural Optimal Particle Filters (NOPFs), which learn an amortized approximation to the optimal proposal from offline simulated one-step conditioning tuples. The learned proposal is used as a drop-in replacement in standard PF updates, with samples corrected by standard importance weights so that the method asymptotically targets the same filtering distribution under standard support and density-evaluation assumptions. Across stochastic nonlinear benchmarks of varying inference complexity, NOPFs improve sample efficiency and distributional accuracy over standard PF baselines with modest computational overhead. The approach integrates data-driven proposal learning into classical inference without altering the underlying filtering objective.
cs.LG / 21 / 2609.30500
PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control
Yuhe Sui, Yingzhi Tang, Shufang Chen
cs.LG · cs.AI
Abstract
Can causal softmax attention implement policy mirror descent as a repeated controller rather than a one-step algebraic identity? Negative-entropy policy mirror descent (PMD) has the statewise update $\operatorname{PMD}_η(π,Q)=\operatorname{softmax}(\logπ+ηQ)$. Building on the known Q-TD-PMD recursion, we construct one fixed causal-softmax actor--environment--one-step-critic protocol with explicit actor, routing, sampling, and normalization residuals, and propagate them to the policy actually returned. The construction states the finite-logit/full-support domain, the external tokenization and sampling boundary, and the mean-zero LayerNorm carrier conditions required by the normalized compilation. Separately trained pre-LN Transformers recover the target computation empirically. A frozen one-step audit model is closest to PMD among the tested fixed rules; in a preregistered five-run $S=4$ repeated-control test, the learned actor with an exact one-step critic reaches median returned-policy loss $1.052\times$ the Exact PMD oracle and retains the criterion across four no-retraining shifts. The same checkpoints with their learned critic give descriptive median $1.050\times$ the oracle (no registered margin). At $S=8$, replacing the exact critic by the learned critic raises median $T=20$ loss to $0.0225$ yet leaves the Liang--Lai and Algorithm Distillation adaptations $20.2$--$24.2\times$ higher-loss; this is a one-sided sampled-critic bound because PolicyAttention consumes 144 generative transitions per round versus 20 on-policy transitions for the adaptations. The strict 20-transition comparison remains open. At $S=8,16$, the exact-critic common-harness comparison remains $17.7$--$28.2\times$ lower-loss than those adaptations, with the information asymmetry stated locally.
cs.LG / 22 / 2609.30503
Federated Targeted Maximum Likelihood Estimation
Diyang Li, Fei Wang, Kyra Gan
cs.LG · cs.DC · stat.CO
Abstract
The evidence behind a scientific or operational decision is often held by hospitals, banks, or registries that cannot pool individual observations. Cross-silo federated learning moves computation to the data and exchanges agreed summaries. Targeted maximum likelihood estimation (TMLE) refines a flexible initial fit, yielding plug-in estimators that respect the model and support efficient inference. TMLE itself, however, has remained a fully centralized procedure. To fill this gap, our paper introduces the first federated TMLE algorithm. We federate targeting itself, for an arbitrary target, loss, and fluctuation family, through two complementary frameworks. FedTMLE-G aggregates local gradients and reproduces centralized targeting step for step. FedTMLE-L lets each institution complete its own fluctuation fit before a single exchange of fitted updates, trading synchronized fidelity for local autonomy. For gradient aggregation, we develop a finite-precision protocol that transmits changes rather than values and certifies targeting accuracy within explicit bounds on exchanges and bits. A description-length analysis of the accepted updates then shows that this finite communication leaves numerical targeting error negligible against sampling uncertainty. The cost of computing an estimator is thus distinct from the complexity of selecting it. Our analysis also indicates that keeping data local is not itself a privacy guarantee of TMLE, since instability of full-record reconstruction need not prevent recovery of a specified sensitive attribute. For a personalized version of local averaging, institutions retain their own estimates and leave once local targeting is complete. A nonconvex convergence bound charges the improvement forfeited through averaging to disagreement among local fits and exposes a tradeoff between equal institutional influence and the sampling variability of small silos.
cs.LG / 23 / 2609.30508
Benchmarking the Connectomes of Caenorhabditis elegans within the Reservoir Computing Framework
Felix S. Reimers, Ola Huse Ramstad, Aliaksandr Hubin, Stefano Nichele
cs.LG
Abstract
The aim of this work is to examine the connectomes of Caenorhabditis elegans through a computational lens using the reservoir computing framework. Connectomes are mappings of biological neural networks; C. elegans is the first organism for which physical connectomes covering the whole nervous system have been published. The connectomes of C. elegans used in this paper have been derived at different ages of the organism and are based on three different ways of measuring inter-cellular connections. They have, with minimal preprocessing, been implemented as reservoirs in the form of echo state networks, which are recurrent neural networks. In reservoir computing, the reservoir itself is not trained, rather the output of the reservoir is passed to a comparatively small read-out module in which training takes place. Training and testing is conducted in different neuro-inspired tasks, with the aim of using these tasks as a benchmark for the connectomes. This process has been repeated with different configurations of the reservoir and equally sized but randomized null models have been used for comparison. The results show that the biological wiring and a bio-informed configuration of input and output nodes of the reservoirs do not necessarily lead to better performance. Contrarily, the randomized null models are often outperforming the original connectomes on the chosen benchmarks. At the same time it becomes clear that the results depend a lot on the configuration of the reservoir and the way the connectome has been derived from the organism. Connectomes from different ages may produce varying outcome, without a clear trend becoming visible.
cs.LG / 24 / 2609.30542
GyroNovo: Error-Guided Fragment Imputation with Mass-Aware Attention for \textit{De Novo} Peptide Sequencing
Abdellah El Mekki, Laks V. S. Lakshmanan, Muhammad Abdul-Mageed
cs.LG
Abstract
De novo peptide sequencing from tandem mass spectra is essential for identifying peptides without relying on reference databases. Despite advances in deep learning, accurate sequencing remains challenging because experimental spectra are often sparse, noisy, and incomplete, leaving informative b- and y-ion fragments unobserved. Existing methods attempt to recover this missing evidence via latent-space imputation before autoregressive decoding. However, they typically treat imputation as a fixed reconstruction task, without considering which missing fragments are most relevant to decoder errors. Moreover, existing peak representations do not explicitly model mass differences between peaks, despite their fundamental importance. We introduce GyroNovo, a framework with two main contributions. First, we use decoder errors observed during training to adapt the imputation objective, prioritizing fragments associated with frequent decoding errors. We further use the decoder error distribution to construct easy and hard augmented views of each spectrum, enabling the decoder to learn under varying degrees of spectral corruption and missing-fragment severity. Second, we introduce a mass-aware inductive bias into self-attention by using rotary embeddings to encode pairwise mass differences between spectral peaks. Together, these components align missing-fragment recovery with decoder behavior while explicitly incorporating the mass relationships that underlie peptide fragmentation. At inference time, GyroNovo retains a standard encoder-imputer-decoder architecture and requires neither additional inputs nor auxiliary search procedures. Experiments on NovoBench show gains of about 9 percentage points in peptide-level precision and 7 percentage points in amino-acid-level precision over the state-of-the-art baseline. Code: https://github.com/UBC-NLP/gyronovo.
cs.LG / 25 / 2609.30556
Dynamic Regret in Online Convex Optimization with Indicator Switching Costs
Naram Mhaisen, George Iosifidis
cs.LG
Abstract
We study dynamic regret in online convex optimization with an \emph{indicator switching cost}: a fixed penalty incurred whenever two consecutive decisions differ. This captures startup overheads such as server activation, model deployment, and cache updates, and on a bounded domain it recovers norm-based movement costs as a special case. Existing guarantees for indicator costs handle only static comparators. We show that a direct extension of these techniques to dynamic regret provably fails, motivating a different approach. We propose a meta-learning framework: a set of randomized lazy FTRL base learners restarted at dyadic time scales, aggregated by a movement-aware master that mixes their proposal densities and samples actions via maximal coupling of consecutive mixtures. The resulting algorithm satisfies, in expectation, $\mathcal{R}^{\mathbf{1}}_T \le \tilde{\mathcal{O}}(\min\{\sqrt{T(S_T{+}1)},T^{2/3}(P_T+1)^{1/3}\})$, where $\mathcal{R}^{\mathbf{1}}_T$ is the dynamic regret plus the cumulative indicator switching cost, $S_T$ counts comparator switches, and $P_T$ is the comparator path length. The bound holds simultaneously for all sequences and requires no prior knowledge of $S_T$ or $P_T$: it is minimax-optimal (up to logarithmic factors) for tracking piecewise-constant comparators, and also captures frequently moving comparators with small total path length.
cs.LG / 26 / 2609.30578
Reinforcement Learning of Communication in a Mesh of Small Language Models
Mehmet Kerem Turkcan
cs.LG
Abstract
Language models gain accuracy from more compute at test time, but majority voting over independent samples saturates: as samples grow, the vote converges to the model's most frequent answer. Communication can add what sampling cannot: an agent that solves a problem can pass the key step to the others. We present TalkMesh, a decentralized mesh of small language model agents that learns when and what to communicate. Each agent samples a proposal and scores it with a trained confidence head. The most confident agent broadcasts a hint; agents below a confidence threshold revise, keeping each revision that outscores its proposal. Gossip consensus approximates the vote weighted by confidence without a coordinator. A talk policy, trained with group relative policy optimization on the change in correctness after revision, writes hints and revisions. With three agents, which together generate at most six outputs, the mesh reaches the accuracy of majority voting over 32 samples with each of three models. Trained with at most 8 agents and evaluated with 32, it raises accuracy from 0.568 under self-consistency to 0.705 (Qwen3.5-0.8B, GSM8K) and from 0.492 to 0.722 (SmolLM3-3B, MATH-500). When 4 of 8 agents collude on a wrong answer with fabricated confidence and poisoned hints, majority vote accuracy falls to 0.000 (Qwen3.5-0.8B, GSM8K). A defended mesh, whose agents rescore solutions with their own confidence heads, retains 0.507. Across reasoning, embodied coordination, and traffic signal control, messages improve a decision when the acting agent cannot observe the information it requires and another agent can send it.
cs.LG / 27 / 2609.30580
Energy-efficient operation of neural operators for virtual sensing
Jason Yoo, Samrendra Roy, Souvik Chakraborty, Syed Bahauddin Alam
cs.LG · cs.PF
Abstract
Virtual sensing repeatedly reconstructs physical fields from changing observations, often on a fixed geometry. We investigate how shared spatial computation reduces the energy of these updates while retaining the selected checkpoint and its evaluated predictions. In a heat-exchanger service, standard compiler freezing and explicit trunk reuse give similar operating energy reductions relative to graph replay: approximately 1% at one request per second and 20% at forty requests per second. In 15 W mode with fixed clocks, reuse with graph replay completes the same request sequence with 22.0 to 22.5% less energy than eager execution, including preparation and waiting. DeepONet and Fourier neural operator (FNO) controls distinguish the effects of reusable arithmetic and launch overhead. Preparation, artifact construction, and worker replacement add costs outside repeated inference. These results connect operator structure to operating energy and show how update frequency and execution lifetime govern the benefit of computation reuse in physical-field virtual sensing.
cs.LG / 28 / 2609.30592
QSV: Quat-Sphere-Vision for Coupled Quaternion Attention on Spherical Lattices
Nicholas Foley, Devin Marinelli, Donny Moore, Diego Enriquez, Amanda Fernandez
cs.LG · cs.CV
Abstract
In standard attention, three separately learned projections decide how strongly a token attends to each neighbor ($W_Q$, $W_K$) and how the attended features are transformed before aggregation ($W_V$). We study Quat-Sphere-Vision (QSV), a sparse spherical vision model that replaces this projection triple with a single learned unit quaternion per token: the relative quaternion $r_{ij} = q_i^{*} \otimes q_j$ supplies both the attention logit $\operatorname{Re}(r_{ij})$ and a sandwich-product feature transport $x \mapsto r_{ij} \otimes x \otimes r_{ij}^{*}$, with messages passed over sparse kNN graphs on concentric Fibonacci spheres. Ablations that change only the targeted component show the two roles to be asymmetric. Removing the transport reduces test accuracy by about four percentage points on CIFAR-10 and CIFAR-100 (single runs per CIFAR-100 variant), while replacing the learned attention weights with uniform averaging leaves it essentially unchanged. Parameter-matched controls then remove the geometry itself: standard attention on the same graph exceeds QSV (mean $87.3\%$ vs. $85.9\%$), and the same model on a flat 2D lattice reaches $91.1\%$, within $2.1$ points of a ResNet-20 trained under the same pipeline (single run). In the coupled kernel, nearly all of the learned pairwise computation resides in the transport channel.
cs.LG / 29 / 2609.30605
Probabilistic Robustness-driven Universal Adversarial Perturbations with Explainability against Deep Reinforcement Learning-based Intrusion Detection System
Hongsen Zhang, Lu Zhang, Mingjing Xu, Yi Zhang, Gregory Epiphaniou, Carsten Maple
cs.LG
Abstract
Deep reinforcement learning (DRL) enables adaptive intrusion detection in dynamic network environments but also exposes intrusion detection systems (IDS) to adversarial threats such as universal adversarial perturbations (UAPs), which apply a single input-agnostic perturbation to degrade detection performance across traffic. Probabilistic Robustness (PR), as a post-hoc evaluation metric, provides a principled, population-level measure of adversarial impact that conceptually aligns with the universality objective of UAPs, i.e., PR quantifies the prevalence of misclassification in the input space, making it a natural signal for guiding UAP generation. Hence, we propose PR-based UAP, which represents the first integration of an explicit PR-driven objective into generating UAPs against DRL-based IDS. Building on this formulation, we introduce PX-UAP, which leverages explainable artificial intelligence (XAI) to guide perturbation shaping under realistic domain constraints, and provides a rigorous theoretical analysis of its design. Extensive experiments demonstrate that PX-UAP consistently outperforms state-of-the-art UAP methods in attack effectiveness.
cs.LG / 30 / 2609.30628
OpenHail: An Event-Driven Gymnasium Environment for Electric Ride-Hailing Fleet Control
Tommaso Schettini, Nicholas D. Kullman, Jorge E. Mendoza
cs.LG · math.OC
Abstract
Machine-learning policies have attracted increasing interest for ride-hailing fleet control in recent years. Reinforcement learning, in particular, requires a structured simulation environment that specifies observations, actions, rewards, and decision epochs for training and evaluation. For electric fleets, this environment must also capture the interaction among stochastic demand, vehicle operations, and capacitated charging infrastructure. We present OpenHail, an open-source Gymnasium environment for joint control of electric ride-hailing fleets. Its fixed-size observation--action interface exposes request assignment, repositioning, and charging to a single policy. The event-driven simulator represents requests with pickup deadlines, vehicle job queues, battery dynamics, and finite-capacity charging facilities with first-in--first-out queues. A configurable decision-epoch mechanism separates internal simulator events from policy interactions, supporting event-driven, periodic, hybrid, and policy-requested control within the same operational model. The software provides seeded instances, feasible-action utilities, evaluation tools, operational metrics, and baseline policies. The source code is available at https://github.com/tommaso-schettini/openhail.
cs.LG / 31 / 2609.30633
Stable initialization without the CLT
Simon Kuang, Kyle Chickering, Xinfan Lin
cs.LG
Abstract
Successful training of deep neural networks is highly dependent on the distribution of the initial weights. If the weights are too large, network training blows up; if they are too small, the model fails to learn features. Stable initialization is the optimal moderation between these two extremes. The conventional theory of random networks uses the Central Limit Theorem to control inter-neuron dependencies, which introduces distributional approximation error and coupling between layers. For networks with sine activations, we derive the uniform-phase initialization, which obviates distributional approximation and fully decouples the layers. Ours is the first work to use the sine function's periodic symmetry. Models trained with the uniform-phase initialization outperform the state of the art in neural representation tasks like image and audio fitting. We find that our untuned models are competitive with the best-tuned baselines from previous work and support $μ$P width scaling.
cs.LG / 32 / 2609.30634
In-Context Binding Capacity in Language Models
Manas Venkata Sai Ravulapalli, Samrath Singh Chadha
cs.LG
Abstract
How many assignments can a language model recall before it loses track of which value belongs to which entity? We measure this limit using continuous recall curves for 12 models at or below 3B parameters and a threshold sweep over 30 open models up to 12B. On the continuous curves, the load at which recall falls halfway to chance follows $K_{50}=cN^α$, with $α=0.820$ and $R^2=0.73$. The broader sweep shows an eightfold range associated with pretraining recipe, although the continuous curves show no detectable recipe effect after controlling for scale, with few modern models in the fit. We derive why interference can lower measured capacity by reducing single-binding recall even when the load-dependent recall profile is unchanged. Direct task training also exceeds the extrapolated zero-shot law, but different measurement criteria prevent interpreting that comparison as a capacity gain. Its formation times follow a power-law form in two independent codebases, conditional on runs that succeed. Together, these results characterize capacity at the model's query interface. Bounds on joint recall and a decomposition of policy errors connect this measurement to working memory and instruction following, without treating recall as a measure of alignment. The controlled task also provides a baseline for testing whether binding limits constrain world-state tracking; the present experiments do not measure state updates or downstream transfer.
cs.LG / 33 / 2609.30650
Causal Retention in Interactive Agents: Interface Factorization and Selective Adaptation
Shengjun Zhang, Tingyi Liu, Dong Xie, Yunlong Dong, Xiang Wang, Cheng Zeng
cs.LG · cs.AI · stat.ML
Abstract
Task performance need not determine which intervention mechanism an agent retains. We study causal retention: whether a frozen learned state answers a mechanism-probe map fixed independently of training, including action, context, direct target, value, and delay. For finite structural causal model classes, the optimal probe error is a Bayes decision risk. It vanishes exactly when every learning-interface fiber lies within one probe-answer fiber; any state obtained by post-processing that interface inherits the same lower bound. A posterior-coverage theorem characterizes budgeted retesting, while an exact edit decomposition shows that the shifted set is the unique support of an error-free target update. Causal Core implements these conditions through evidence-gated writing, readout filtering, temporal credit, hidden-context setup, and local diagnostic updates. Experiments cover finite causal systems, continuous simulators, an official TD-MPC2 world model, and Qwen2.5-7B-Instruct. A frozen Qwen last-layer probe reaches 0.958 balanced accuracy on source mechanisms but 0.583 on changed delays; the gated mechanism state reaches 1.000 and accepts only 0.056 of synchronized-readout candidates. In TD-MPC2, five target states per actuator recover effect-sign accuracy from 0.057 to 0.948 without degrading stable responses. Causal retention is therefore distinct from task sufficiency and source-domain decodability.
cs.LG / 34 / 2609.30661
Population loss in shallow ReLU networks: Bias & families of critical points
Michael Field
cs.LG
Abstract
The main result presented is a formula for the population loss in the student-teacher kernel model that is applicable to shallow ReLU networks with bias. This extends previous work of Choo and Saul (2009) and Brutzkus and Globerson (2017). The formula makes essential use of Owen's T-function. The necessary theory of the T-function is given and a high precision coding using MPFR for the T-function, based on an algorithm of Komelj (2023), is available on request. It is shown that various families of spurious minima described in past papers of Arjevani and the author extend to biased networks and that the loss is always strictly decreased when bias is added. The change in landscape geometry caused by adding bias appears to be relatively mild. Only the simplest examples are described in this paper where it is assumed that the number of inputs is equal to the number of neurons (this restriction is for reasons of length). A review of relevant previous results on unbiased networks is included. Aside from Gaussian statistics, the main mathematical tools and ideas come from analytic geometry (analytic and subanalytic sets, the Curve Selection Lemma).
cs.LG / 35 / 2609.30684
PixSim: a calibrated open-source simulator of instant-payment fraud, recovery and interdiction under analyst capacity constraints
Bashir Zeimarani, Alireza Khatib, Somayeh Mousavinasr, Carlos Maurício Serodio Figueiredo
cs.LG · cs.CE · q-fin.RM
Abstract
Brazil's Pix settles about 5.9 billion instant, irreversible transfers a month. A fraudulent transfer can be recovered only while the funds remain in a traceable account, and in 2025 the Central Bank's recovery mechanism (MED) returned 9% of accepted contested value. Interdiction therefore has to happen before settlement, by routing each transaction to pass, human review or block, under a finite analyst team and a regulatory hold window. To our knowledge no public simulator jointly models irreversible settlement, a regulated recovery mechanism, downstream fund dispersal and capacity-constrained review. We present PixSim, an open-source simulator of the Pix rail with these elements, calibrated to Banco Central do Brasil open data, with every parameter sourced, calibrated to one published observable, or registered as an assumption. With the model frozen, full-scale runs reproduce the 2025 recovery rate within 0.006 and its decomposition within 0.02; the February-April 2026 window is reported as a misfit and the May 2026 tracing regime as a projection. On a benchmark with a payer-side scorer, four reference policies and ten scenarios, within the simulated mule model: recovery after settlement is constrained by dispersal speed; staffing by the arrival profile cuts a fixed rule's alert expiry from 52% to 2% at constant hours; halving the team removes a fixed threshold-and-block rule's advantage over a queue-aware rule, on loss and on loss plus false-block harm (+0.106 of victim value, positive on all twenty paired seeds), while a reversal at two thirds of the team was not confirmed on independent seeds; and a synthetic scorer of held-out AUC 0.82 cuts lost value by about a quarter. Code and data: https://doi.org/10.5281/zenodo.22948895
cs.LG / 36 / 2609.30718
NEMSim: Learning Control-Conditioned Multi-Event Physical Dynamics via Executable Event-Mechanism Priors
Junsong Yu, Junjie Xie, Pengwei Liu, Dong Ni
cs.LG
Abstract
High-fidelity simulation of control-conditioned multi-event physical systems is computationally expensive, especially across broad control spaces and long trajectories. In these systems, macroscopic evolution emerges from localized discrete events whose intensities and effects depend on process controls and evolving local states, while the available system knowledge is typically expressed as event-attribute descriptions. Purely data-driven surrogates must infer these event effects from limited trajectory coverage, which can hinder generalization to unseen control regimes. Physics-guided methods instead primarily build on equation-level constraints or differentiable solvers rather than discrete event-rule priors. We therefore propose NEMSim (Neural Event-Mechanism Simulator), which compiles predefined event-attribute descriptions into an executable transition structure linking control-dependent event intensities, prior-guided mechanism attribution, and state-dependent responses. To enable evaluation of control-conditioned multi-event dynamics with explicit system knowledge, we construct a 3D KMC-based benchmark pairing high-fidelity trajectories with explicit event rules, standardized splits, and evaluation protocols. Across three settings, NEMSim reduces Avg. RMSE by 58.9%-81.3% relative to the strongest baseline in each setting. It also remains best in the data-efficiency study with training-data fractions down to 10%. Mechanism analyses further show that these gains arise from executable rule integration rather than prior access or architecture alone.
cs.LG / 37 / 2609.30721
When 10,000 Windows Are Not 10,000 Tests: Auditing Statistical Confidence in Sliding-Window Time-Series Classification
Xinze Shi, Litian Zhang, Binrui Shi
cs.LG
Abstract
Sliding-window classifiers are often evaluated on thousands of overlapping test windows, even though neighboring predictions share observations and remain nested within recordings and subjects. Subject-disjoint evaluation prevents one form of leakage but does not make those test windows independent. We present a practical audit that maps three claims - performance on observed recordings, future recordings from observed subjects, and unseen subjects - to explicit aggregation rules and established dependence-robust inference. At 75% overlap, controlled simulations give 16.9% Type-I error for IID observed-record inference and 7.2% for session-centered Bartlett-HAC: a substantial improvement with residual miscalibration. Audits of frozen WISDM and HARTH predictions show that nearly fourfold growth in test rows yields only 1.75-1.94-fold variance-equivalent information growth. At that overlap, fixed-record paired Accuracy-difference intervals are 1.22-1.66 times the IID widths; this inflation is not universal at zero overlap. On HARTH, paired Accuracy-difference intervals include zero across three overlap settings, whereas Macro-F1 favors MiniROCKET. Independent recomputation, common-session checks, class-level results, and separately seeded calibration make the audit's scope and limitations inspectable. The resulting workflow distinguishes additional predictions from additional independent evidence.
cs.LG / 38 / 2609.30746
Mechanism-Aware Ensemble Conditioning for Data-Limited Emulation of Extreme Events
Isabella S. Thiel, Juan Bello-Rivas, Yannis G. Kevrekidis, Themistoklis P. Sapsis
cs.LG · math.DS · physics.ao-ph
Abstract
Extreme events in chaotic systems are difficult to learn from short trajectories because they are controlled by transient finite-time instability rather than by frequently observed bulk dynamics. We propose a mechanism-aware conditioning plug-in framework that turns a nudged coarse ensemble into a non-intrusive sensor of local instability geometry. In the small-noise regime, the ensemble covariance aggregates the same finite-time deformation kernels that govern local instability, providing a Jacobian-free proxy for the local amplification structure around a synchronized coarse trajectory. A small FiLM module injects statistics of this ensemble geometry into an otherwise unchanged backbone while leaving the coarse simulator unchanged. We demonstrate this interface in two distinct pipelines: a Transformer-style residual-attention corrector for a controlled low-dimensional chaotic system and a probabilistic recurrent STORN corrector for topographic two-layer quasi-geostrophic (QG) flow. In the low-dimensional benchmark, ensemble covariance directions co-activate with OTD modes and FiLM conditioning improves 99th-percentile exceedance-frequency errors over an identical no-context Transformer baseline. In QG, a fixed ensemble-conditioned FiLM-STORN model trained on only \(50\) time units substantially improves long-horizon rare-event statistics in the data-limited regime, including density-tail errors, exceedance frequencies, and spatial exceedance-area distributions relative to an unconditioned STORN trained on the same data; on averaged high-threshold exceedance diagnostics, it also outperforms the baseline STORN trained with $20$ times more high-resolution data. These results show that local instability geometry is not merely interpretable post hoc, but an actionable conditioning signal for data-efficient rare-event emulation.
cs.LG / 39 / 2609.30752
Differentiable RNA Secondary Structure Extraction for Deep Learning
Tyler Illman, Max Ward, Marcell Szikszai, Ryan K. Krueger
cs.LG
Abstract
Many deep learning approaches to RNA secondary structure prediction have recently been proposed. They typically output a weight matrix $W$ where $W_{ij}$ is an arbitrary weight for base $i$ pairing with base $j$. Converting this matrix to a predicted secondary structure or base-pairing probability matrix typically involves ad hoc and problematic downstream algorithms. Despite the importance of this conversion step, which we refer to as structure extraction, it has received relatively little attention in the literature. In this work, we analyze how the congruence between training and extraction methods affects prediction performance. To do this, we compare four extraction algorithms: a Nussinov-like dynamic programming method, maximum-weight graph matching and the greedy extraction algorithms used by SPOT-RNA and RiNALMo. These are evaluated on outputs from the pretrained RiNALMo model and three toy models trained in this paper: a differentiable Nussinov-like model, a binary cross-entropy (BCE) baseline, and a model that incorporates a novel symmetric doubly stochastic matrix (SDSM) normalization algorithm during training which allows it to output base-pairing probability matrices directly, without a separate extraction step. This SDSM normalization algorithm is differentiable and can be added inline to any deep learning model during training and evaluation. We find that the performance of each extraction method depends strongly on how the corresponding model was trained. Considering the toy models themselves, the SDSM model showed the strongest overall performance: it outperformed the BCE baseline under all four extraction algorithms and produced pre-extraction outputs closest to the ground truth. These results suggest that SDSM normalization is a tractable alternative to traditional structure extraction.
cs.LG / 40 / 2609.30781
Missingness-Aware Conformal Prediction Under Cross-Hospital Distribution Shift
Liang You, Dongwen Ou, Hengyu Shi, Siyuan Dai
cs.LG · stat.AP
Abstract
Clinical measurements are recorded for some patients but not others, at rates that differ across hospitals, and marginal conformal coverage does not ensure coverage within groups defined by missingness. We propose a missingness-aware conformal calibration procedure for mortality prediction under cross-hospital distribution shift. It selects a measurement on an independent sample, groups patients by whether that measurement is recorded, and applies Mondrian calibration within each group, so no calibration outcome is reused. We evaluate the procedure across hospitals in eICU and across care units within one MIMIC-IV hospital, using three predictors. Relative to pooled calibration, it reduces the average worst-group coverage gap on its selected groups in all six settings, with a median reduction of 1.9 percentage points; paired site-bootstrap intervals exclude zero in five. These gains do not extend uniformly. Calibration by predicted risk achieves smaller gaps on a broader panel of missingness groups, and when eICU hospitals are evaluated separately, the gain shrinks for all three predictors and reverses in sign for one. We explain this discrepancy with a hospital-level decomposition. Pooling reweights hospitals through a covariance between group shares and coverage errors, and lets errors of opposite sign cancel: weighting explains the reversal, and cancellation accounts for most of the attenuation for the other two predictors. Constructed population distributions show that pooled and within-hospital evaluations can rank calibration methods oppositely even without sampling noise. Pooled improvement alone therefore cannot establish better coverage within hospitals, even when the calibration groups are fixed.
cs.LG / 41 / 2609.30789
Interpretable-by-Design Descriptor Portfolios Match a 2048-Dimensional Foundation Embedding on Low-Data Molecular Assays
Yiqi Yao, Miquel Duran-Frigola
cs.LG
Abstract
In low-data structure-activity prediction, the choice of molecular representation can matter more than the choice of predictor, and tabular foundation models sharpen that effect. We ask whether a portfolio of compact, semantically named descriptor blocks can reach the accuracy of a 2048-dimensional CheMeleon embedding while staying auditable at the feature level, meaning that every input dimension carries a model name and a recorded training provenance. Starting from a fixed 11-dimensional physicochemical base, we greedily concatenate provenance-screened blocks using the labelled context alone. Across nine ADME/Tox assays and 50 evaluation cells, scored on common-coverage subsets restricted to the molecules that every representation covers, the portfolio reaches a mean test AUC of 0.762, against 0.764 for CheMeleon and 0.756 for Mordred. The pooled gap to CheMeleon is +0.003 AUC (task-bootstrap 95% CI [-0.020, +0.030]), which satisfies our predeclared pooled parity gate but not the per-assay gate. At 25 context labels the headline rule again satisfies the pooled gate; at 10 labels it does not. We also report four predeclared candidate-selection rules that we falsified. Post-freeze checks over ten seeds and three previously unseen assays support pooled competitiveness for compact, auditable representations; a same-width random-bundle control does not establish that greedy membership itself adds accuracy. Assay-level differences remain unresolved.
cs.LG / 42 / 2609.30790
Towards Universal Representation-Based Process Control
Jinmyeong Choi, Taesup Kim, Artur Dubrawski
cs.LG
Abstract
Many temporal process learning and monitoring pipelines operate in local windows, making window-level decisions unavoidable in practice. In such settings, classical statistical tests can be applied to individual windows, but they typically evaluate predefined parametric hypotheses-such as unit-root or moment-based conditions-thereby limiting flexibility when reference behavior is defined empirically from task- or domain-specific data. In this work, we view window-level monitoring as a process control problem and reformulate it as reference-based hypothesis testing, where the null hypothesis is specified by an empirical reference distribution rather than a fixed parametric model. We operationalize this perspective through a representation-based, nonparametric framework that combines pretrained time series encoders, kernel density estimation, and conformal calibration, yielding finite-sample valid inference in learned representation space. Classical notions such as stationarity and cyclostationarity arise as natural instantiations of empirical reference sets within this framework. Through experiments, we demonstrate sensitivity to window-level distributional deviations while maintaining well-calibrated inference under stable reference regimes, highlighting the applicability of the proposed approach to a broad class of time series process control and monitoring tasks.
cs.LG / 43 / 2609.30811
Counterfactual Online Conformal Prediction Under Adaptive Logging
Xinyu Qiao, Yichen Lin, Kaihong Ji, Xue Wang, Tao Yao
cs.LG
Abstract
Online conformal prediction can fail when predictions shape actions and actions determine which outcomes enter calibration. Standard adaptive methods may retain marginal coverage while systematically miscovering the counterfactual outcomes of rarely selected actions. This paper formalizes the failure through counterfactual coverage and introduces Propensity-Weighted Online Conformal Prediction, an inverse-propensity-weighted recursion that debiases calibration. A doubly robust variant further reduces nuisance bias to the product of outcome-model and propensity errors. Under positivity, the resulting coverage rate matches an information-theoretic lower bound up to logarithmic factors. Experiments on synthetic decision tasks, open bandit data, and financial rebalancing show that PW-OCP and DR-OCP improve counterfactual coverage and downstream regret without sacrificing prediction-set sharpness.
cs.LG / 44 / 2609.30819
Learning Provable Neural Network Observer for Uncertain Dynamical Systems
Zhangyi Wang, Jiaxu Liu, Chen Song, Chao Xu, Shengze Cai
cs.LG · math.OC
Abstract
In many safety-critical applications, control of uncertain dynamical systems relies on observers that estimate states and external disturbances. Neural network observers can improve estimation accuracy, but certifying their Lyapunov stability via Linear Matrix Inequality (LMI) constraints leads to large-scale semidefinite programs (SDPs) that are difficult to solve for large networks. To overcome this scalability bottleneck, we propose a novel two-stage training framework for provably stable neural network observers. Our approach decouples the optimization into a point-guided Lyapunov pre-training phase, which rapidly achieves high estimation accuracy and local stability over sampled states, followed by an LMI fine-tuning phase that efficiently satisfies a strict global Lyapunov stability certificate. We provide formal theoretical guarantees for local stability radii and probabilistic coverage over a prescribed compact error-state domain under specified regularity and sampling assumptions. Experiments on nonlinear control benchmarks and X-29 aircraft ablations show that our LMI-certified neural network observers train significantly faster than direct LMI-based methods and generalize robustly across diverse systems, achieving improved tracking accuracy over a range of observer baselines. The code is available at https://github.com/Berry-Myon/LearningNeuralNetworkObserver.
cs.LG / 45 / 2609.30820
Quantizing Looped Transformers: Feedback Exposure and Calibration Blindness
Nux Li
cs.LG · cs.CL
Abstract
Looped transformers reuse weights across recurrence steps, making low-bit quantization especially attractive. We identify two distinct failure modes of standard post-training quantization. On Huginn-3.5B, per-channel INT4 fails primarily at the non-residual loop-entry adapter, while quantizing the residual core is much less damaging. We call this feedback exposure: a quantized layer perturbs the recurrent state without an identity path, and the resulting error is fed back at later steps. Controlled experiments on linear filters and Mamba state-space models show that feedback exposure also occurs outside transformers. Grouped INT4 reveals a separate failure, calibration blindness: our one-step GPTQ baseline builds its Hessian from step-0 activations, leaving input directions used later in the recurrence nearly unweighted. Across nine checkpoints from seven looped architectures, one-step GPTQ is worse than round-to-nearest (RTN) on the primary task metric for five checkpoints. Accumulating the GPTQ Hessian across recurrence steps outperforms both one-step GPTQ and RTN on all nine checkpoints and recovers bf16-level accuracy on Huginn. These results separate two questions for PTQ on looped models: where quantization error enters the recurrence, and which states calibration sees.
cs.LG / 46 / 2609.30822
Adaptive Interaction Graphs for Particle Simulation
Aiden Zhou
cs.LG · cs.CE
Abstract
Learned particle simulators based on graph neural networks achieve strong one-step accuracy, but errors compound over long horizons. An underexplored variable is the interaction graph: existing methods fix its topology via k-nearest neighbors or a static radius rule, regardless of local model confidence. We propose making this graph adaptive: a per-particle variance head, trained jointly with the acceleration head under a heteroscedastic Gaussian NLL loss, drives a trajectory in which high-uncertainty particles receive an expanded neighborhood. This is done at little extra inference cost by using the previous step's uncertainty estimate. A key discovery is that the variance head learns a meaningful notion of uncertainty: high-variance particles concentrate near complex regions, such as splash zones or free surfaces. When this signal drives graph topology, the resulting AdaptGNS simulator achieves a strict Pareto improvement on WaterDrop and a modest gain on Sand. Given the model's stronger performance on WaterDrop, we hypothesize that adaptive graphs are most useful when complexity is concentrated in space. Our code can be found at https://github.com/aidenzhou8/AdaptGNS.
cs.LG / 47 / 2609.30837
MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation
Tianze Xu, Yanzhao Zheng, Zhentao Zhang, Yuanqiang Yu, Chao Ma, Jihuai Zhu, Lelun Wu, Lyumanshan Ye, Pengfei Liu, Baohua Dong, Hangcheng Zhu, Ruohui Huang, Gang Yu
cs.LG · cs.AI
Abstract
Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single student, but existing practice typically hard-routes each prompt to a domain-matched teacher for the entire rollout. This dependence on prompt-level domain labels restricts using unlabeled training mixtures and leaves complementary signals from other teachers unused. We introduce MOPD-Router, a framework that routes supervision over the full teacher pool at each token, without domain labels or training a separate routing model. Its plug-in interface supports different metrics for selecting and weighting teacher-specific OPD signals. Within this interface, we propose ExpertAlign, which scores each teacher by whether its correction to the student at the current token expresses the specialization that teacher acquired during post-training, and compare it against two reference metrics built on teacher confidence (Entropy) and teacher-student discrepancy (Novelty). Experiments on unlabeled and domain-labeled training mixtures under strong-to-weak and same-size distillation scenarios show that ExpertAlign achieves the strongest overall performance in all four settings. On unlabeled data, it improves the overall score by 5.88 (+12.3%) points over Mean aggregation; on domain-labeled data, it outperforms standard MOPD by 3.95 (+7.8%) points without using available domain labels. These results demonstrate token-level routing can exploit cross-domain complementary supervision, and reduce exclusive reliance on prompt-level domain assignment. Code is available at: https://github.com/TURLEing/MOPD-Router.
cs.LG / 48 / 2609.30838
Peer-Grounded Counterfactual Path Planning for Chronic Health Management
Saman Khamesian, Hassan Ghasemzadeh
cs.LG
Abstract
Effective behavioral intervention in chronic disease management requires not a single prescription but a sequence of incremental steps, each grounded in what real, similar individuals have demonstrably achieved. Counterfactual explanation offers a natural computational route to such guidance, answering what change in behavior would have produced a better outcome. But existing methods return a target state without a route to it, guarantee no monotone health improvement along the way, and draw no evidence from peer behavior -- asking a patient to close a wide gap in one move, which is precisely the recommendation structure least likely to be attempted. We propose POROS (Peer-Grounded Optimal Routes Over States), a domain-agnostic framework rooted in Bandura's self-efficacy theory and Festinger's social comparison theory that constructs a Behavioral Progression Graph -- a directed acyclic graph over observed patient states in which every edge requires both peer-grounded behavioral proximity and strict health outcome improvement. Every edge is therefore a behavioral change that individuals in the cohort have demonstrated is achievable within a single period. Minimum-cost paths through this graph decompose otherwise inactionable behavioral gaps into incremental, peer-grounded steps. We evaluate POROS on two independent longitudinal cohorts of patients with diabetes. For patients below the 70% clinical threshold for time in range (TIR, blood glucose within 70-180 mg/dL), it reduces the mean gain required per step from 26.3 percentage points (pp) to 5.5 pp on one cohort and from 31.1 pp to 5.7 pp on the other, decomposing large behavioral jumps into the incremental steps that self-efficacy requires. Across both cohorts, 97-98% of multi-hop paths cross patient boundaries, embedding social comparison by construction.
cs.LG / 49 / 2609.30839
Attention-Based Adaptive Policies for Simultaneous Speech-to-Text Translation
Filip Tăşădan, Ema Tomanová, Ondrej Lopuch, Paweł Bilko, Anders Søgaard
cs.LG
Abstract
Simultaneous speech-to-text translation (Simul-S2TT) consists of generating partial translations while the incoming audio frames are processed by the system. However, the streaming nature of this setup creates the challenge of deciding the best moment to perform an accurate translation while minimizing the delay. To address this challenge, we utilize the cross-attention mechanism of the encoder-decoder architecture to find the right alignment between the input speech frames and the target text tokens. In this paper, we propose the Recent Frame Attention Policy (RFAP) and the Dual-Condition Attention Policy (DCAP) that allow offline trained speech-to-text translation models to be used in streaming scenarios without requiring additional training. Results on three different language translation pairs over the CVSS-C corpus show that the RFAP is able to surpass other policies with gains of up to 4.0 BLEU while reducing the translation delay by almost 1 second. Moreover, the DCAP is able to preserve a high translation quality when the latency is very low.
cs.LG / 50 / 2609.30840
Aligning One-Step Generative Models with Reward-Weighted Transport Distillation
Austin Wang, Ziheng Cheng, Lexing Ying
cs.LG · cs.CV
Abstract
One-step generators enable high-quality visual generation with a single network evaluation, but their post-training is difficult: general implicit generators provide neither tractable likelihoods nor denoising trajectories, and many rewards are non-differentiable. We introduce Reward-Weighted Transport Distillation (RWTD), a post-training method that requires only generated samples and scalar reward evaluations. Rather than aligning solely to the conventional reward-tilted reference distribution, RWTD constructs an adaptive target that mixes separately tilted current and reference distributions. The current component incorporates improvements discovered during training, while the reference component anchors the target to the pretrained generator. RWTD realizes this target through feature-space optimal transport and fixed-point regression. Theoretical analysis shows that the fixed-point distributions of RWTD interpolate between off-policy reward tilting of the reference and on-policy tilting of the current model, providing a principled approach to balancing reward adaptation with retention of prior knowledge. Empirically, RWTD substantially improves the GenEval score of the one-step SANA Sprint 1.6B backbone from 0.73 to 0.80, while separate preference alignment experiments demonstrate strong cross-reward generalization that yields balanced improvements and preservation of compositional capabilities.
cs.LG / 51 / 2609.30856
Learning Chance-Constrained MDPs with Bellman Distributional Certificates
Chenbei Lu, Hongyu Yi
cs.LG
Abstract
Safe reinforcement learning (RL) commonly enforces expected-cost constraints, but such expectation safety may fail to control the probability of rare high-cost trajectories. Chance-constrained MDPs (CCMDPs) impose a stronger probability-level requirement, but are widely viewed as harder because the chance constraint is nonconvex and depends on the full trajectory rather than a Bellman-linear expectation. In this paper, we reveal that this computational difficulty does not necessarily imply a higher statistical price. For tabular discounted CCMDPs with fixed bounded successor support and access to a certified planning oracle, we establish a model-based upper bound, with a matching lower bound up to logarithmic terms. Technically, our key idea is the \emph{Bellman distributional certificate}, which constructs a Bellman recursion for constraint violation probabilities before policy selection. The certificate can be reused across candidate policies; combined with shared row-wise reverse-KL confidence sets, it gives a policy-uniform trajectory-KL transfer without a union bound over policies or time--budget Bellman tables. For stochastic policies, we give a model-free variance-reduced policy-gradient algorithm with a finite-sample expected KKT-residual guarantee and independent validation of every accepted policy. Numerical experiments on synthetic CCMDPs and an IEEE 14-bus energy storage control benchmark illustrate the safety and mechanism behavior of the proposed algorithms.
cs.LG / 52 / 2609.30918
Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations
Marcin Kostrzewa, Maciej Zięba
cs.LG · cs.AI
Abstract
Robust counterfactual explanations promise recourse that still works after the model behind it changes. Whether they keep that promise depends on what the change is. A small perturbation of the parameters, retraining on new data, and a new architecture are different events, and each existing method is evaluated against the one it was built for. Reported robustness scores, therefore, answer different questions and cannot be compared. We propose a unified cross-family evaluation protocol that holds factual instances and generated counterfactuals fixed while testing every method against the same eight types of model change. The benchmark compares six robust methods and two standard baselines on four tabular datasets. It characterizes every changed classifier through its outputs and reports empirical robustness together with coverage, base validity, and proximity. We find that relative performance and failure modes vary across change families. Bounded parameter perturbations change 0.95\% of test predictions on average, compared with 4.9\% for bootstrap retraining. Methods with guarantees for these perturbations do not necessarily transfer to other changes. RobX transfers most consistently in our experiments, although greater stability can require larger interventions. We argue that robust CFE methods should be evaluated through a common protocol that specifies the model changes, measures their realized behavioral magnitude, and keeps generation performance separate from robustness.
cs.LG / 53 / 2609.30929
EPOC: Endpoint-Preserving Online Correction With Compressed Residual State for Multi-Horizon Time Series Forecasting
Takumi Fujimoto, Hiroaki Nishi
cs.LG
Abstract
Completed multi-horizon forecasts provide residual feedback for a fixed forecaster, but retaining full residual blocks increases auxiliary state. We propose Endpoint-Preserving Online Correction (EPOC) with a compressed residual state. It stores low-order discrete cosine transform (DCT) coefficients and the final value of the preceding residual block. Within each channel, the endpoint is shared across component-wise online ridge regressions that also use current-forecast coefficients. The fitted DCT correction is blended with the base forecast. We evaluate eight multivariate series with DLinear and PatchTST, three seeds, and two training variants, yielding 96 matched fixed-base conditions at a 24-step horizon. EPOC achieves mean condition-wise reductions in mean squared error (MSE) and mean absolute error (MAE) of 15.40% and 9.35% from the uncorrected base, respectively, with a median of 6,352 B in retained auxiliary arrays. It has lower paired MSE than the $δ$-Adapter, COSA, FAC, and OMPB in a majority of conditions and uses less state than each. Full ELF achieves the largest mean MSE reduction, 19.29%, but its median retained state is 474,048 B ($\times$75 relative to EPOC). Equal-size summary controls favor the endpoint by 1.65--2.20% in paired MSE; a coefficient-reconstructed endpoint yields similar accuracy to the observed endpoint, highlighting its role as a shared input. Increasing the retained DCT component count from 4 to 8 adds 1.00 percentage point of MSE reduction for 5,728 B. On jointly trained bases, EPOC lowers MSE by 16.69--20.15% relative to globally blended TEFL-style adapters applied to the same base. The code and numerical records are available at https://github.com/keiotakmin/endpoint-preserving-residual-correction.
cs.LG / 54 / 2609.30948
PORL: Pretrained Offline Reinforcement Learning for the Job Shop Scheduling Problem
Mateo Toro Diz, Jonathan Hoss, Noah Klarmann
cs.LG · cs.AI
Abstract
The Job Shop Scheduling Problem (JSSP) is a fundamental combinatorial optimization problem in industrial optimization. This work introduces Pretrained Offline Reinforcement Learning (PORL), a hybrid approach that combines simulation-based online pretraining with offline fine-tuning on production-specific data. Reinforcement learning through online interaction enables exploration of general scheduling strategies, but typically relies on simulation environments and may suffer from a simulation-to-reality gap. In contrast, offline RL avoids direct interaction with the environment by learning from historical data, but its performance is strongly influenced by dataset quality and coverage. PORL combines the strengths of both paradigms by first learning a general scheduling policy through online interaction and subsequently adapting it offline to a target distribution. A KL-divergence-based policy constraint is introduced to limit deviations from the pretrained policy during fine-tuning. The approach is evaluated on JSSP instances with distribution shift and datasets generated from heuristic, noisy-expert, and random behavioral policies. The results show that PORL consistently achieves lower optimality gaps than standalone offline RL and the considered general scheduling baselines. Furthermore, its advantage over standalone offline RL increases as dataset quality decreases, indicating reduced sensitivity to the quality and coverage of the available offline data. The results suggest that offline adaptation of pretrained policies is a promising approach for industrial scheduling environments where direct online exploration is impractical.
cs.LG / 55 / 2609.30950
Low-Bit Recurrent States in Hybrid Language Models
Hongren Chen, Jiayang He
cs.LG
Abstract
Hybrid language models maintain fixed-size recurrent states, but existing quantizers typically use eight bits or more. Quantization errors persist according to channel decay rates. We derive distortion weights from the observability Gramian and combine them with normalized state ranges for mixed-precision bit allocation, without calibration data, rotation, or training. We also quantize decay rates logarithmically. With per-token state quantization, a four-bit mean payload reduces excess negative log-likelihood by factors of 3.3--27.9 relative to the best of seven baselines across three hybrid models; metadata costs vary. At six bits, negative log-likelihood differs from the FP32-state baseline by less than 0.005 nats. Ablations separate gains from variable bit widths, decay weighting, and range normalization. With less frequent write-backs, gains diminish and depend on the model and budget.
cs.LG / 56 / 2609.30966
Gradient Surgery for Physics-Informed Neural Networks
Thomas Borsani, Giuseppe Di Fatta
cs.LG
Abstract
Physics-Informed Neural Networks (PINNs) are trained by optimising a composite objective that combines data fitting with physics-based constraints, typically resulting in a highly imbalanced multi-task optimisation problem. Under these conditions, existing optimisation strategies are affected by conflicting task gradients, leading to slow convergence and unstable training, particularly for stiff and high-frequency partial differential equations. We analyse gradient conflicts throughout training of PINNs with standard optimiser and investigate Multi-Task Deep Learning (MTDL) optimisation methods. In our analysis across four benchmark problems we observed that PINN optimisation exhibits three distinct phases in which angle- and magnitude-based gradient conflicts alternate, with only one present at a time. Building on these observations, we propose PAM-GS, a physics-aware gradient surgery method that adaptively mitigates task interference during training according to the observed conflict types. Experiments on four representative PDE benchmarks demonstrate that PAM-GS combines competitive solution accuracy with consistently strong task-balanced performance, outperforming existing methods on most problems.
cs.LG / 57 / 2609.30973
LipSSM: Structurally Lipschitz-Bounded Cascaded State-Space Model via Metric Transfer between Consecutive SSM Layers
Natsuki Yoshino, Ren Uchida, Kazuki Matsumoto, Kohei Yatabe
cs.LG · eess.SY
Abstract
Lipschitz continuity is a fundamental principle in the design of certifiably robust deep neural networks (DNNs), wherein adjusting the Lipschitz constant, which quantifies network robustness, is of central theoretical importance. A standard approach to enforcing Lipschitz continuity requires each layer of a DNN to be Lipschitz continuous, thereby guaranteeing overall Lipschitz continuity. However, this layer-wise approach typically imposes overly conservative restrictions by producing a loose estimate of the overall Lipschitz constant, which limits the expressive capacity of the DNN and degrades empirical performance at a prescribed level of robustness. To overcome this loose estimation, the recently proposed LipKernel transfers information across layers to yield a much tighter overall Lipschitz bound than conventional layer-wise construction. In this paper, we extend this concept to cascaded state-space models (SSMs) to construct Lipschitz-continuous DNNs capable of modeling longer-term dependencies. The proposed architecture, named LipSSM, is theoretically justified and empirically evaluated.
cs.LG / 58 / 2609.30995
Learning Hierarchical Causal Representations of the Effects of Forcings on Temperature in Climate Models
Shan Zhao, Ilija Trajkovic, Julia Kaltenborn, Yaniv Gurwicz, Peer Nowack, David Rolnick, Julien Boussard
cs.LG
Abstract
Machine learning (ML) emulators provide a fast and cost-effective method to simulate climate change scenarios after being trained on Earth System Models projections. However, the black-box nature of those data-driven approaches limit the usability and trustworthiness of their outputs and in particular their use as causal attribution tools. Here, we develop a hierarchical causal representation learning framework applied to sea surface temperature fields from a state-of-the-art global climate model. As a key advance over previous work, our framework explicitly models both atmospheric dynamical interactions arising from internal climate variability and forced responses due to changes in atmospheric greenhouse gas and aerosol concentrations. When trained on future climate change scenarios, our method accurately predicts the long-term global mean and regional temperature evolution and shows physically realistic responses to perturbations in greenhouse gas and aerosol concentrations when evaluated on unseen scenarios. Our results underline the potential of causal representation learning frameworks for advancing climate model emulation.
cs.LG / 59 / 2609.31016
Robust Successor Features
Erik Nikulski, Yamen Habib, Vicenç Gomez, Anders Jonsson, Rubén Moreno-Bote, Javier Segovia-Aguas
cs.LG
Abstract
Generalization in Reinforcement Learning (RL) refers to the ability to execute close-to-optimal policies in unseen tasks after the agent has been trained on a different set of tasks. Building on the seminal work of the successor representation and further adaptations with function approximation, Transfer in RL has traditionally focused on generalizing to tasks that only differ in the reward function. A decade after the introduction of the successor representation, Robust RL emerged simultaneously from several articles in the field of operations research. In Robust RL, the transition kernel is unknown, and the goal is to maximize the expected reward under this uncertainty. Our work unifies these two paradigms through robust successor features, which generalize across both the reward function and the transition kernel, under the assumption that tasks are linear Markov Decision Processes. We derive a bound on Generalized Policy Improvement (GPI) that explicitly quantifies how performance degrades with the mismatch between transition kernels, recovering existing successor-feature guarantees when dynamics are shared. Finally, the generalization capabilities of robust successor features are validated on several grid-based benchmarks and compared to previous alternatives that focus solely on either the reward or the transition kernel.
cs.LG / 60 / 2609.31031
Metacognitive Selective Ensemble for Mobile Systems
Sungmin Lee, Kichang Lee, Joonhee Lee, JaeYeon Park, Songkuk Kim, JeongGil Ko
cs.LG
Abstract
Deep ensembles improve robustness in mobile sensing, but repeatedly executing many models over continuous sensor streams is costly. Selecting only a few members reduces this cost, yet adaptive selection often requires additional model execution to obtain reliable evidence about inactive candidates. We present MetaSE, an active ensemble framework that exploits short-term persistence in per-model reliability. MetaSE maintains a small active set across windows, uses post-execution evidence to reject unreliable members, and invokes lightweight routing only when replacement is needed. This stateful design accesses the diversity of a larger pool without repeated full-pool evaluation. Across four HAR datasets and four model architectures, MetaSE consistently improves over a fixed three-model ensemble and achieves accuracy comparable to substantially more expensive adaptive and full-ensemble inference. On a Raspberry Pi 4B, MetaSE is 2.7x faster and uses 69% less memory than full ten-model inference.
cs.LG / 61 / 2609.31033
Robust Graph Clustering Network for Multiple Missing Data
Keyuan Qiu, Renda Han, Zhen Tang, Qiang He, Xingwei Wang, Wenxin Zhang, Guangzhen Yao, Junxin Chen, Qingjian Ni
cs.LG
Abstract
Clustering on graphs where both node attributes and structural links are partially missing remains a challenging task. Existing methods typically rely on imputation-then-clustering on single-view missingness incomplete graphs, which are vulnerable to cross-view error propagation and cluster-boundary blurring under simultaneous attribute and structure missingness. To address these limitations, we propose a Robust Graph Clustering Network for Multiple Missing Data (RGCN), which is designed to handle simultaneous node attribute and graph structure incompleteness. RGCN introduces three key innovations: First, we design a view-decoupled dual-branch imputation to mitigate interference and enable mutual enhancement in recovering missing data. Second, we employ a multi-hyperspherical mixture prior to enhance intra-cluster compactness and inter-cluster separability on a directional latent manifold. Third, a boundary-aware contrastive enhancement objective mitigates the blurring of clusters caused by imputation bias. Extensive experiments on real-world datasets demonstrate that RGCN consistently outperforms state-of-the-art baselines under various missing patterns.
cs.LG / 62 / 2609.31038
Aurora-X: Built for Extreme Time Series Forecasting
Xingjian Wu, Chenjuan Guo, Xiangfei Qiu, Zhigang Hu, Hanyin Cheng, Peng Chen, Yang Shu, Jilin Hu, Bin Yang
cs.LG
Abstract
Time series foundation models (TSFMs) enable cross-domain forecasting, but their development as general-purpose forecasters remains constrained by underexplored training potential and limited architectural versatility. To address these challenges, we introduce Aurora-X, a billion-scale TSFM with a progressive curriculum and a unified architecture. We first use channel-independent pretraining to learn temporal patterns, then introduce cross-variable dependencies, varied context and horizon lengths, and future covariates if available during midtraining. Variable-resolution post-training further enables an adjustable temporal span per token at inference. With fixed model weights, this supports longer histories under a fixed token budget or fewer tokens for the same history, enabling test-time scaling. With a versatile architecture, Aurora-X supports cross-variable modeling, covariate conditioning, and parallel decoding of future patches for probabilistic forecasting. These are supported by a novel pattern-guided mixture-of-experts that expands model capacity through sparse activation and uses shallow patch similarities to constrain deep-layer routing, guiding expert specialization across heterogeneous time series. Furthermore, we propose an implicit quantile network head that predicts arbitrary quantiles to characterize predictive distributions, enhancing probabilistic forecasting flexibility. Comprehensive experiments on GIFT-Eval, TIME, FEV-Bench, TFB, and DAG-Bench demonstrate state-of-the-art forecasting performance against pretrained TSFMs and task-specific supervised models.
cs.LG / 63 / 2609.31061
Distributed Learning as a Service: The Developer's Perspective
Tianyue Chu, Filippo Vannella, Dimitra Tsigkari, Paula Delgado-Santos, Fernando López, Pablo Gomez Guerrero, Sotirios Spantideas, David Solans Noguero
cs.LG · cs.DC · cs.NI
Abstract
Application developers of distributed learning services face challenges that a typical federated learning loop does not address. Specifically, the model updates can still leak private data, devices might not be able to participate in the training due to limited resources, a single aggregator might not be able to scale, and the transmissions of model weights induce a considerable bandwidth cost. This paper demonstrates DLaaS (Distributed Learning as a Service) from the developer's vantage point. Using a single admin dashboard, the developer initiates a distributed/federated learning job and is able to activate Differential Privacy (DP), Split Learning (SL), Hierarchical Aggregation (HA), and Knowledge Distillation (KD) as declarative options, with no change to the clients' code. We demonstrate the complete service lifecycle on an industrial smart-home Wake-up Word (WuW) task, using the "Ok Aura" dataset. Once the developer initiates a distributed learning job by toggling DP, SL, HA, and KD in the admin dashboard, the system dispatches the job to a set of Android clients and Dockerized helper aggregators. In the demonstration, these mechanisms run live across configurations. Then, the clients train the model locally and return their updates. The trained model is served to a consumer-side Android application that performs on-device WuW detection on a live microphone stream. In particular, the conference attendees will be invited to speak the trigger phrase and monitor in real time the per-class confidence and inference latency. Finally, we release the source code and short video walkthroughs of these configurations.
cs.LG / 64 / 2609.31082
SAGE: A sampling-aware global evaluation benchmark for species distribution modeling
Emilia Arens, Nina van Tiel, Robin Zbinden, Damien Robert, Lukas Drees, Chiara Vanalli, Benjamin Kellenberger, Niklaus E. Zimmermann, Loïc Pellissier, Devis Tuia, Jan Dirk Wegner
cs.LG · q-bio.PE · stat.ML
Abstract
Knowing where species occur is fundamental for biodiversity research and conservation. Species distribution models (SDMs) link species observations to environmental conditions to estimate their spatial distribution. However, accuracy varies with the underlying data and models, making it essential to know for which species models can be trusted. Deep-learning-based SDMs ("DeepSDMs") now jointly model thousands of species, drawing on hundreds of millions of community-science records. At this scale, averaging performance hides substantial species-level variability, particularly for rare species, often of greatest conservation concern. Records are also strongly biased, making occurrence counts misleading. Accounting for these factors is essential for a reliable and informative evaluation of multi-species SDMs. Here, we introduce a Sampling-Aware Global Evaluation (SAGE) benchmark, combining GBIF records for training with sPlotOpen vegetation plots for presence-absence evaluation across 5771 plant species. We propose an evaluation framework that groups species based on two properties, sampling effort and relative prevalence, which describe how densely a species' range is sampled and how frequently the species is recorded. Evaluating single-species SDMs and multi-species DeepSDMs, we find that Random Forests and DeepSDMs perform best overall, but neither dominates: DeepSDMs outperform single-species SDMs for infrequently recorded species while offering no consistent advantage for well-sampled ones. Crucially, this advantage emerges only when established bias-correction practices, such as spatial thinning and reweighting, are carried over to the deep-learning setting. SAGE helps identify the species and data conditions for which a given approach is beneficial, thereby supporting the development of more transparent and ecologically credible SDMs. Data and code: https://earens.github.io/sage/
cs.LG / 65 / 2609.31093
Block Sparse Attention with Log-Linear Complexity
Bohao Tang, Zhen Qin, Yuqi Pan, Zheng Li, Pengfei Liu
cs.LG
Abstract
Scaling language models to long contexts is limited by the quadratic cost of self-attention. Block sparse attention offers an efficient alternative, but selecting the retained blocks remains a bottleneck. Conventional block selection requires scoring all query-block pairs and therefore remains quadratic in sequence length. To address this issue, we propose PISA, a block-sparse attention mechanism that employs a pyramid Top-$K$ selection strategy. The main idea is to gradually narrow down the candidates across different levels, making it more efficient to find the most relevant keys. Specifically, we construct a coarse-to-fine hierarchy of keys and perform selection from the coarsest level. At each level, LogSumExp scoring is applied to a bounded candidate set to select candidates for the next finer level, continuing until the finest level is reached. Through pooling, we construct $O(\log N)$ levels of keys, yielding an overall complexity of $O(N\log N)$, where $N$ denotes the sequence length. We develop hardware-aware Triton kernels for both training and inference, fusing hierarchical routing and LogSumExp scoring without materializing the query-key score matrix. We further evaluate our method on language modeling tasks. Compared with the baseline, our method achieves comparable performance on benchmarks such as commonsense reasoning while delivering better results on retrieval tasks.
cs.LG / 66 / 2609.31098
The Residual Stream's Effective Depth
Barak Gahtan, Ido Galil, Alex M. Bronstein
cs.LG
Abstract
We introduce \emph{effective depth} ($\Deff$), a scalar diagnostic that treats the layer-wise residual stream of a transformer as a discrete-time process, measures how representation similarity decays with layer distance, and aggregates that profile into one number. Across sixteen decoder-only language models, $\Deff$ separates a structural consequence of residual accumulation from an empirical one: even maximally diverse orthogonal updates have the closed-form reference $F_L = 2L/(L+1)<2$, yet fifteen of sixteen default measurements lie below $F_L$ (Qwen3.5: 32--44\%, OLMo-2: 40--41\%, Pythia: 23--28\%). Matched references show that the gap is not caused by the persistent initial state or update-size imbalance, but is largely a calibrated signature of correlated residual updates rather than evidence that depth is unused. Symmetric position-0, token-normalisation, and top-PC controls show the regime is not reducible to BOS or top-PC artefacts: the lone above-reference default outlier joins the same regime, and all sixteen models are sub-reference after token-normalisation or top-1-PC removal. Intermediate checkpoints show that the regime is established early in OLMo-2 and stable through 5T tokens, while Pythia-1.4B follows a distinct decreasing trajectory. A controlled residual-carry intervention supports the mechanism, and $\Deff$ is best read as a \emph{global} accumulated-state diagnostic, not as a capability score or pruning method.
cs.LG / 67 / 2609.31107
Bayesian Optimization with Fisher Information Geometry: Gradient Bounds and Trust-Region Methods
Saksham Kiroriwal, Julius Pfrommer, Jürgen Beyerer
cs.LG · cs.AI · cs.IT
Abstract
We study Bayesian optimization (BO) through the lens of information geometry. Pulling back the Fisher information metric through the surrogate posterior map yields a local sensitivity tensor on the input space, which leads to an upper bound on the gradient of reparameterizable acquisition functions. This view explains vanishing-gradient behavior in high-dimensional BO and provides a common interpretation of heuristics such as RAASP and dimension-scaled lengthscales. Building on this analysis, we propose FITR, a trust-region-based BO method that replaces lengthscale-based scaling by local pullback-Fisher weights. FITR is not restricted to GP kernels with explicit lengthscales. On GP benchmarks with an SE kernel, experiments show competitive performance using FITR. The proposed method also easily generalizes to non-isotropic surrogates, although the gains are more task-dependent in that setting.
cs.LG / 68 / 2609.31114
From Shortcut Learning to Discrete Neural Insertion Sort
Konstantinos Mylonas, Thrasyvoulos Spyropoulos
cs.LG · cs.AI
Abstract
Neural algorithmic reasoning aims to train neural networks to follow known algorithms and generalize beyond the input sizes seen during training. However, correct final outputs and intermediate supervision do not necessarily show that a model follows the intended execution. We study this problem using insertion sort. Our analysis of the CLRS30 baseline NAR shows that the hint objective is weakly optimized and that hint accuracy remains low. Moreover, many intermediate representations can already be decoded into sorted sequences before the reference insertion-sort execution terminates, suggesting that the model learns a shortcut to the final output. Motivated by these findings, we introduce Discrete Neural Insertion Sort. Our model represents the sequence as a chain, separates scalar exchanges from control-state transitions, and projects node representations back to discrete states after every processor step. When trained only on sequences of length 16, the model achieves $100\%$ sorted-sequence accuracy on sequences of length 64 and 128. However, an ablation shows that discretization and graph structure alone are insufficient: without additional supervision of the global inner-loop state, the model fails even at the training length. Our results show that discrete execution can support strong length generalization, while also highlighting the problem-specific inductive bias required to learn a faithful algorithmic execution.
cs.LG / 69 / 2609.31128
Frame the adversary: a structure-aware attack methodology
Vicky Kouni, Stelios Perrakis, Francis Bach, Pascal Frossard, Yann Chevaleyre
cs.LG
Abstract
Frequency-based adversarial attacks have recently grown popular by exploiting spectral sensitivities shared across neural architectures. Unlike spatial perturbations, frequency-based attacks expose deeper vulnerabilities, making them especially valuable for robust evaluation of safety-critical and security-sensitive applications. Yet, existing approaches are typically not derived as solutions to an optimization problem that explicitly captures transform-domain structure. In this paper, we propose a methodology for crafting principled frequency-based adversarial attacks, via a dedicated optimization framework. A cornerstone of our method hinges on the introduction of a perturbation constraint set, tied to highly structured non-orthogonal transforms, well-known for their flexible, non-predefined frequency handling. We prove that the attacks emerge as weighted $\ell_2$-projections onto this set, yielding a general and controlled attack generation mechanism. By this, we provide a clear geometric attack characterization, ensuring alignment between the optimization objective and the perturbation constraint. We assess our framework on standardized datasets, for pretrained and adversarially robust models. Results highlight that our attacks, being solutions to an optimization problem, over a structured perturbation set, are highly effective, even across different, unseen architectures. Our methodology could serve as a theoretical baseline for designing and analyzing transformed-based attacks, targeting fundamental model vulnerabilities, instead of mere architecture-specific artifacts typically studied in the robustness literature.
cs.LG / 70 / 2609.31155
Teacher-Anchored Selection of Post-Training Quantized Models under Domain Shift
Alejandro Rodriguez Dominguez, Muhammad Shahzad, Xia Hong
cs.LG · cs.AI
Abstract
Compressing a trained model yields a family of deployment candidates, and under domain shift the most compressed one need not be the one to deploy. We study selection over such a family, with candidates and teacher fixed and target labels absent or scarce. Two findings organize the label-free case. Minimum teacher distortion behaves almost as a constant rule, selecting the same eight-bit, per-channel, unclipped configuration in every run, which does not minimize empirical target cross-entropy. Established estimators divide sharply: in the overconfident-collapse regime of the CNN families, confidence-based estimators order the family close to backwards, and the diagnostics that identify it need the labels the setting denies, while output-distribution estimators match the teacher-relative anchor and on one architecture beat it. Distortion is nonetheless stable, so a supervised term can move selection away from it. Combining the two, we give exact quadratic identities for a canonical quadratic analogue of the family. We also show that under symmetric corruption the label-dependent part of a criterion linear in the label indicator is multiplied by one common factor whenever its coefficient sums are candidate-invariant, a class holding teacher contrasts and accuracy but not cross-entropy. These characterize the score's components without bounding selection regret. Across one hundred and thirty-four candidate families, one per independently trained convolutional or Vision Transformer teacher, anchoring reduces mean regret at the smallest label budget in every setting, an advantage that fades beyond twenty-five labels.
cs.LG / 71 / 2609.31157
Bayesian Tensor Autoencoder with Physics-informed Predictive Prior for Multi-dimensional Time Series Anomaly Detection
Jianan Liu, Chunguang Li
cs.LG
Abstract
Multi-dimensional time series, inherently tensorial, are common in practice. Despite great progress in time series anomaly detection, most existing methods are confined to uni-/multi-variate time series. When handling multi-dimensional time series using these methods, reshaping operations are required, which inevitably break the intrinsic correlations and thus lead to performance degradation. In uni-/multi-variate time series anomaly detection, AutoEncoders (AEs) are widely adopted and generally categorized into reconstruction-based and prediction-based AEs. The reconstruction-based AE utilizes the current observation for reconstruction, while the prediction-based AE utilizes the historical information to predict the current observation. Thus, the two AEs utilize different information. To bridge the gap between reconstruction-based and prediction-based AEs, so as to fully leverage the available information and thus further enhance performance, we propose a predictive prior and incorporate it into the reconstruction-based AE. It may not be very difficult to conceive this idea, but designing the predictive prior so that it can work for tensor anomaly detection is non-trivial. Specifically, to avoid breaking the intrinsic correlations within the multi-dimensional time series, we use the tensor AE as the backbone. To incorporate the predictive prior into the reconstruction-based AE, we propose a Bayesian fusion approach and our analysis reveals that this approach can enhance the modeling capability of the model for normal data. To mitigate the over-generalization problem of AE, we incorporate physical laws, i.e. tensor low-rank decomposition rules, into the neural networks in the predictive prior, leading to the Physics-informed Predictive Prior Tensor AE (PPPTAE) framework. Experimental results on real-world datasets demonstrate the effectiveness of the proposed method.
cs.LG / 72 / 2609.31161
I Act Therefore I Am: When Is JEPA's Action-Conditioning Enough to Learn Causal Mechanisms?
Yuhang Liu, Zhuo Huang, Javen Qinfeng Shi
cs.LG
Abstract
Recent empirical and theoretical advances suggest that joint-embedding predictive architectures (JEPAs) may learn meaningful representations for action-conditioned prediction of future outcomes, thus becoming one of the foundational structures for world models. However, accurate prediction does not, in general, necessarily imply recovery of underlying causal states that give rise to the observed dynamics. This work investigates when and how JEPAs can recover the underlying causal states from observations. We first introduce a latent variable model, in which high-dimensional observations are generated from latent causal states whose dynamics are governed by action-conditioned transition mechanisms. Based on this formulation, we develop a general information-theoretic objective that combines conditional likelihood maximization for learning transition dynamics with entropy maximization for preserving latent state information. We then establish identifiability conditions under which representations learned by this general objective recover the underlying latent causal states up to component-wise invertible transformations and permutation. One key condition for such identifiability is sufficient action-induced variation in the transition mechanisms. Guided by this finding, we instantiate the general objective with an action-modulated Gaussian additive-noise model, yielding action-modulated JEPA (A-JEPA). Experiments on synthetic environments verify the theoretical findings under the identifiability conditions and robustness to moderate violations, while visual benchmarks demonstrate improved state recovery and transfer to unseen transition mechanisms.
cs.LG / 73 / 2609.31162
WorldTS: World Modeling for Multimodal Covariate-aware Time Series Forecasting
Yuhan Zhu, Xiangfei Qiu, Hanyin Cheng, Wangmeng Shen, Chenjuan Guo, Bin Yang, Jilin Hu, Christian S. Jensen
cs.LG
Abstract
Time series forecasting is typically framed as learning a direct mapping from historical to future observations in the observation space. However, sequences of observations generally provide only a partial view of the dynamics of the underlying system, with future observations being shaped by latent dynamics. Recent latent-space forecasting methods thus achieve improved performance by predicting future observations from latent-space representations of historical observations rather than directly forecasting future observations in the observation space. Next, while future observations are also shaped by external factors, how to incorporate external, often multimodal, information into forecasting, so that it can shape latent-state formation and evolution directly, remains underexplored. We propose WorldTS, a world-modeling based forecasting framework that integrates multimodal covariates directly into the forecasting to further improve forecasting performance. Specifically, WorldTS employs a two-stage training strategy. First, it learns forecasting-relevant latent state dynamics conditioned on multimodal covariates, yielding encoded future states. Next, the learned state dynamics are frozen, and an observation decoder is trained to map the predicted future states back to future observations. Extensive experiments on 21 real-world datasets offer insight into WorldTS and its effectiveness.
cs.LG / 74 / 2609.31168
Audio emotion recognition for atypical hearing
Ulysse Roussel
cs.LG
Abstract
My doctoral work aims to explore Audio Emotion Recognition (AER) in the context of atypical listening. This research focuses on auditory hypersensitivity in people with autism, a phenomenon that is often difficult to evaluate and unique to each individual. Our core idea is to leverage our understanding of affect from acoustic traits, relying on the possibility of generalizing affective responses from a small amount of annotated data. As a first step, we fine-tune a large foundation model, Contrastive Language-Audio Pretraining (CLAP) using low-rank adaptation (LoRA), trained on a valence and arousal dataset of neurotypical listeners.
cs.LG / 75 / 2609.31197
ALF: An Active Learning Framework for Scientific Discovery
Shikha Surana, Alex Hawkins-Hooker, Olivia Gallup, Christoph Brunken, Jules Tilly, Paul Duckworth
cs.LG
Abstract
Machine learning for scientific discovery is almost systematically data bound. Producing relevant high quality data, under budget constraints, is amongst the most promising ways to advance the field. Active learning (AL) offers promise wherever labelling requires expensive experiment, measurement, or simulation. Most existing tools cover only part of the data acquisition loop, and typically focus on either offline benchmarking or online deployment, but not both. We present ALF, a modular AL Framework that runs the full data acquisition loop via five modular components. One clear API for both settings: offline, against an existing dataset for controlled and reproducible experimentation; and online, against an oracle for acquiring new candidates in real-world deployments. ALF is open-source and available at https://github.com/instadeepai/alf.
cs.LG / 76 / 2609.31206
Self-Supervised Representation Learning: From Spectral Foundation Models to Auroral Emission Spectra
Matthieu Le Lain, Gaël Cessateur, Sébastien Lefèvre
cs.LG · physics.space-ph
Abstract
Auroral spectrographs such as the Auroral Spectrograph In Skibotn (ASIS) record hundreds of thousands of emission spectra, but only a few hundred can be labelled by an expert. To exploit the rest, we pretrain a 1D Vision Transformer with a masked autoencoder on 223,000 unlabelled spectra. Without labels, its representation recovers the emission-line intensity ratios that physicists use to diagnose the precipitating particles (R^2 0.91 vs. 0.77 for an untrained control) and, under one linear probe, classifies as well as 13 features designed by experts. Fine-tuned, the model outperforms the previous supervised auroral classifier on its own benchmark (macro-AP 88.5 vs. 77.8), reaches 0.870 mAP, and exceeds the same architecture trained from scratch by +0.159 with 10% of the labels; attribution shows that it uses both N2+ bands. Could an existing pretrained model replace it? Two astronomical spectral foundation models and a time-series model transfer according to their spectral window: SpectraFM, trained in the infrared, falls below the untrained control, whereas SpecFormer, trained in the optical, approaches in-domain pretraining without reaching it.
cs.LG / 77 / 2609.31250
Deterministic Regime Switching and Feasibility Inversion in Dynamic Tensor Rematerialization
Mahesh Reddy Pagadala
cs.LG · cs.DC
Abstract
We report fine-grained, deterministic instability in Dynamic Tensor Rematerialization (DTR), an online eviction policy for memory-constrained DNN training, measured on the reference DTR simulator (simrd) using public execution traces. On an LSTM trace, memory budgets differing by 0.10% of unconstrained peak memory select fast and slow execution regimes whose overheads differ by as much as 7.3x; the slow regime is driven by broadly repeated re-eviction of the same storages (evictions per storage rise from 1.33 to 8.27 while the set of distinct evicted storages is essentially unchanged: 5,233 vs 5,236, with the two sets overlapping at Jaccard 0.999). On a ResNet-32 trace, a fine budget sweep reveals a deterministic feasibility inversion: the run is feasible at ratio 0.101, infeasible (OOM) across 0.102-0.106, and feasible again from 0.107. We trace the immediate cause of the OOM to a fully pinned recursive rematerialization frontier that exceeds the budget after every evictable tensor has been evicted. Ablations using the DTR authors' own variants implicate the joint size-staleness scoring term in the observed LSTM instability. We argue these are at least two distinct budget-sensitive pathologies rather than one mechanism, and we separate what is demonstrated from what remains hypothesised. All results concern the reference simulator; reproduction in a production runtime is future work. Code, instrumentation, and raw results accompany this preprint.
cs.LG / 78 / 2609.31291
Softmax Reparameterization for Output-Head Quantization
Asim Kadav, Christian Flores, Chirag Arora, Varun Kotte, Hongbo Zheng, Lan Yan, Priya Shanmugasundaram, Tracy Holloway King
cs.LG · cs.AI
Abstract
Large vocabularies make output heads a substantial inference cost in small language models. We propose softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from every output row and selects the coefficient by validation KL separately for RTN, activation-weighted MSE, and full-Hessian GPTQ. This one-dimensional search includes the original head and fixed mean-centering, preserves the full-precision softmax distribution, and leaves the trained decoder unchanged; a rank-one correction handles nonlinear logit paths such as soft-capping. Across seven heads, W4 gains concentrate where baseline quantization substantially distorts predictions: on Phi-4-mini, AW-MSE KL falls from 0.936 to 0.256. The gains survive stronger GPTQ calibration and remain complementary to exact per-channel scaling and affine quantization. Across four heads and three W4 quantizers, frozen WikiText-selected coefficients also transfer to C4 and OpenWebMath, outperforming mean-centering in all 18 comparisons where the frozen coefficient differs from $1$ and matching it in the remaining six. At W2, used as a compression stress test, benefits broaden across nearly the full model--quantizer matrix. Matched residual analysis shows that improved fidelity can accompany greater logit reconstruction error while reducing the residual's Fisher-weighted cost. For shift-compatible heads, reparameterization adds no inference operation and preserves packed W4 execution: with the decoder held in BF16, quantizing the Phi output head reduces batch-one generation latency by 10.8% relative to the BF16-head baseline.
cs.LG / 79 / 2609.31306
Benchmarking Attention for Tabular Foundation Models
Maximilian Schambach, Clemens Biehl, Sam Thelin
cs.LG
Abstract
Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention involves longer sequences while column attention operates on much shorter ones, and the strided memory layout of tabular data makes producing contiguous tensors costly. Moreover, the hidden dimensions used in current models are small compared to recent language models. Yet efficient attention has been studied mostly for one-dimensional sequences, leaving the two-dimensional tabular setting unexplored. To this end, we create a reproducible benchmarking setup and study the unique characteristics of tabular attention across several backends -- Torch SDPA (efficient and cuDNN), FlashAttention-2/3/4, and the inference-only backends vLLM and SageAttention -- measuring forward and backward throughput across realistic tabular shapes on three GPU generations (A100, H100, B200). We find that the optimal backend choice differs between column and row attention and varies across hardware as well as model specifics: While the FlashAttention implementations tailored for each GPU generation perform overall best, they are at times outperformed by CuDNN in the case of column attention at longer sequences with cross-over points depending on the head dimension. Among inference-only backends, SageAttention performs well for row attention and large sequences beyond 16\,k rows. Our reproducible benchmark lays the foundation for future improvements to table-native attention. The self-contained benchmarking and evaluation code is openly available at: https://github.com/SAP-samples/tabular-attention-benchmark
cs.LG / 80 / 2609.31315
LUCID: Learning Under Confounding for Inference and Discovery in Time Series
Mohammad Fesanghary
cs.LG · stat.ML
Abstract
Unobserved common causes are pervasive in real-world time series and can induce spurious associations that causal discovery methods mistake for direct edges. We propose LUCID (Learning Under Confounding for Inference and Discovery, a regime-adaptive deconfounding layer that first estimates the confounding regime from data using a Marčenko--Pastur spectral router, then applies a deconfounding strategy matched to that regime. When the spectrum indicates pervasive factor confounding, LUCID attenuates factor-dominated variation and recovers contemporaneous (lag-$0$) structure from the resulting innovations, with edge selection calibrated against a data-driven edge-free null. Rather than being tied to a particular discovery algorithm, it can wrap existing discovery engines; we demonstrate consistent improvements across three such methods. On a diverse synthetic out-of-distribution benchmark spanning changes in confounder strength and sparsity, loading density, lag structure, volatility dynamics, edge heterogeneity, persistence, intermittency, and tail behavior, LUCID achieves the best family-weighted directed, lag-resolved graph $F_1$ ($0.60$), improving over the strongest baseline by $0.19$ absolute ($\approx\!46\%$ relative). Its advantage widens relative to looser lag-collapsed scoring, and remains robust under intermittent and heavy-tailed confounding. Code reproducing the method, the benchmark generators, and every reported experiment is available at https://github.com/bloomberg/causal-ts.
cs.LG / 81 / 2609.31325
More Sensors Only One Field: Rethinking Continual Spatio-Temporal Forecasting
Lewei Xie, Haoyu Zhang, Jiajun Zhou, Yulong Chen, Guanxing Chen, Yu-An Huang, Hau-San Wong, Yifan Zhang, Zhi-An Huang
cs.LG
Abstract
Continual spatio-temporal forecasting supports traffic management and environmental monitoring under evolving dynamics and expanding sensor networks. However, conventional graph-based continual learning methods tie forecasting representations to the current sensor layout, so sensor expansion can alter the representation of learned spatial relationships. Our key insight is that sensor expansion changes the evidence available about a process without necessarily changing the dynamics to be learned. We propose STFO (Spatio-Temporal Field Operator), which parameterizes forecasting knowledge as a shared field-evolution operator and handles changing sensor layouts through observation and query interfaces. Normalized coordinate-based aggregation lifts irregular sensor histories onto a fixed latent grid, enabling reuse of learned spatial maps across observation sets without sensor-specific parameters. To accommodate process drift, a spectral descriptor summarizes variation across spatial scales and conditions Fourier propagation and attention to adapt operator responses to the current spatial regime. Coordinate-based decoding queries the evolved field at sensor locations and combines spatial corrections with local-history predictions. Experiments on PEMS-Stream, CA-Stream, and AIR-Stream demonstrate state-of-the-art average forecasting performance. STFO-Large reduces average MAE over DOL by 8.4% on PEMS-Stream and 4.7% on CA-Stream. Our code is available at https://github.com/Xielewei/Spatio-Temporal-Field-Operator.
cs.LG / 82 / 2609.31329
Bridging Body and Brain: Gene-Driven Morphology--Control Co-Design
Fu Feng, Ruixiao Shi, Yucheng Xie, Jing Wang, Xin Geng
cs.LG
Abstract
Morphology--control co-design jointly optimizes an agent's body structure and control policy as an integrated embodied system. However, existing methods typically model morphology design and control with separate networks coupled only indirectly through a shared task objective, limiting explicit high-level coordination. Inspired by natural genes that coordinate biological development, we introduce \textbf{Morphogene}, a compact latent blueprint that bridges an agent's body and brain. Through AdaConcat, Morphogene jointly conditions morphology and control generation at the limb level, allowing its variations to induce coordinated changes in both components. Building on this representation, we propose \textbf{GeCode}, which formulates co-design as exploration in the compact Morphogene space. Each Morphogene anchors a local design region in which nearby body--brain designs are explored, while performance-guided updates move these anchors toward promising regions for more efficient exploration of the broader design space. This process combines local refinement with global exploration while preserving body--brain compatibility. Extensive experiments across diverse 2D and 3D co-design tasks demonstrate that GeCode consistently outperforms existing state-of-the-art methods, achieving substantially faster convergence and higher final performance.
cs.LG / 83 / 2609.31351
Progressive Memory Transformer: Memory-Aware Attention for Time-Series
Tord Sture Stangeland, Andreas Köhler, Steffen Mæland, Adín Ramíres Rivera
cs.LG
Abstract
Time-series carry structure simultaneously at multiple scales (fine-grained variation, mid-range motifs, and global properties) and downstream tasks operate at correspondingly different scales. Most existing self-supervised learning approaches supervise representations globally via instance-level contrastive losses and limited temporal neighborhood supervision, but do not explicitly exploit the structural hierarchy. We propose a learning framework that explicitly enforces a structural hierarchy across three scales independently: a local objective for token continuity, a mid-range objective for window-level motifs, and a global objective for sequence-level agreement. Realizing this framework requires the backbone to expose a representation at each scale; we introduce \textbf{Progressive Memory Transformer} (PMT), which augments a transformer with writable, window-aligned memory that exposes the mid-range scale alongside the token and sequence-level representations conventional transformers already provide. Across seven UCR/UEA/UCI classification benchmarks, a cue-retention probe, and forecasting benchmarks, PMT learns representations that probe well at the global, mid-range, and local scales---strong low-label classification (1--5\% labels), competitive forecasting performance across multiple horizons, and quantitative and qualitative evidence that memory states capture mid-range motifs.
cs.LG / 84 / 2609.31363
Brenier Meets Adversarial Training: Optimal Transport Geometry for Robust Learning
Alireza Abdollahpoorrostam, Ehsan Sharifian, Buse Şen, Marco Cuturi, Daniel Kuhn
cs.LG · math.OC · stat.ML
Abstract
Distributionally robust optimization (DRO) provides a principled framework for learning under distribution shift, but its practical use is hindered by the difficulty of evaluating worst-case risks for nonconvex loss functions. We study a penalized DRO formulation in which the adversary may choose any distribution but incurs a Wasserstein penalty for deviating from the empirical distribution. We show that the adversary's problem can be reformulated as an optimization problem over transport maps that push empirical samples to adversarial ones, and we prove that optimal maps are cyclically monotone. We also show that standard adversarial training---based on per-sample local optimization---violates cyclical monotonicity and wastes transport costs unless the adversary is severely restricted. We propose two remedies. First, we introduce multi-start particle ascent, which alternates parallel gradient ascent with reassignment to enforce cyclical monotonicity across samples. Second, we parameterize adversarial maps as gradients of input-convex neural networks, which guarantees cyclical monotonicity by construction. Experiments on robust regression, image classification, and robust control show that our methods consistently outperform standard adversarial training and state-of-the-art baselines, achieving improved robustness and better generalization under distribution shift.
cs.LG / 85 / 2609.31399
Differential Attention Unlocks Complementary EEG and Speech Fusion for Emotion Recognition
Philip H. Lee, Shreeram Suresh Chandra, John H. L. Hansen
cs.LG · eess.SP
Abstract
Multimodal emotion recognition (MER) increasingly pairs EEG with speech, treating internal neural signals and external vocal expression as informative views of affect. In practice, naive fusion underperforms the stronger single modality, because EEG artifacts inject noise that corrupts the shared representation. We introduce EmoSpeechBrain, a multimodal framework built on the insight that noise suppression is a precondition for effective fusion. Its EEG encoder uses differential attention, taking the difference between two attention maps to cancel shared noise and isolate discriminative neural activity. An attention-based gating adapter aligns both modalities in a shared space and weights each one's contribution to the prediction. On two datasets - PME4 and EAV, EmoSpeechBrain improves MER accuracy by up to 12.9% over other state-of-the-art (SOTA) EEG encoders, and surpasses unimodal speech and EEG baselines by up to 13.1% and 23.1%. These results show that once EEG noise is suppressed, fusion delivers gains that naive combination cannot.
cs.LG / 86 / 2609.31401
Decodable In-Context State and Model Output Across Training
Manas Venkata Sai Ravulapalli, Samrath Singh Chadha
cs.LG
Abstract
Prior work established that a probe can decode an in-context binding on model errors and that probe-guided steering can repair some of them. We follow probe accuracy, model output, and steering response across public pretraining and post-training checkpoints. Probe accuracy rises during Pythia pretraining, while probe-guided steering moves from negligible all-trial benefit to a larger benefit at two model sizes. Saved scores distinguish probe-correct errors with low and above-uniform model probability for the correct candidate. Oracle-target steering already repairs many early errors, but saved aggregates cannot separate target quality from intervention sensitivity. A held-out comparison of decoders trained on the final state or candidate logits finds no detected final-state advantage on late-checkpoint model errors. An information-theoretic counterexample explains why decodability on errors alone cannot establish discarded output information. The connection to downstream omissions remains open.
cs.LG / 87 / 2609.31415
Evaluating the accuracy of KV cache reuse techniques
Samuel Cestola, Tianxiang Xia, Pengfei Zheng, Weiyan Zheng, Bo Wang, Yi Zhao, Diego Didona
cs.LG
Abstract
Position-independent KV cache reuse aims to reduce latency in retrieval-augmented generation by reusing chunk-level KV caches across prompts. We show that current evaluations of KV cache reuse techniques rely on measurements that fail to faithfully capture the loss of accuracy attributable to reuse, often artificially inflating the reported effectiveness. We also show that existing datasets do not exhibit the reuse dynamics needed to thoroughly evaluate such techniques. To address these issues, we propose an evaluation methodology that measures this accuracy loss without ambiguity and we introduce Boxoffice, a tool that programmatically generates evaluation datasets that exercise challenging KV cache reuse patterns.
cs.LG / 88 / 2609.31454
Different Corruptions, Different Signals: Uncertainty and Loss in Federated Data Quality
Bradley Scott, Zeqi Luo, Edmond S. L. Ho
cs.LG · cs.AI · cs.CV
Abstract
Federated learning (FL) data corruption can affect either inputs or labels, but it remains unclear whether input-conditional uncertainty and prediction-label loss expose these corruption modes equally. This paper compares two corruption-detection signals in FL: input-conditional uncertainty and prediction-label loss. The uncertainty signal is characterised using a learned aleatoric variance estimate together with Monte Carlo (MC) dropout variance and entropy measures, while the loss is computed against the supplied label. We test these signals against additive image noise and persistent random label flips. On ResNet-20 with CIFAR-10 and SVHN under Dirichlet partitions with data that are not independent and identically distributed (non-IID), the two corruption types behave differently. For persistent random label flips, the within-client per-sample area under the receiver operating characteristic curve (AUC) is 0.85 on CIFAR-10 and 0.95 on SVHN for prediction-label loss, while every uncertainty estimator stays at chance (0.49--0.50). This pattern is consistent with the model remaining confident in the underlying image despite the supplied label being wrong. For image noise, expected-entropy uncertainty rises above chance (0.67 on CIFAR-10 and 0.66 on SVHN), while loss responds comparably (0.64 on both). Each signal is therefore the stronger detector for a different corruption: the prediction-label loss for persistent label flips, and expected-entropy uncertainty for image noise, with its advantage becoming apparent as federation-wide corruption prevalence increases. Robust FL data-quality assessment should match the signal to the corruption rather than rely on uncertainty alone across corruption types.
cs.LG / 89 / 2609.31463
Uncertainty-Aware Federated Learning for Infant Movement Analysis
Edmond S. L. Ho
cs.LG · cs.AI · cs.CV
Abstract
Infant movement analysis provides valuable biomarkers for the early identification of neurodevelopmental disorders. Recent advances in deep learning have enabled automated analysis of infant movements from video-derived skeletal representations, achieving performance comparable to expert assessment for tasks such as General Movement Assessment (GMA). However, most existing approaches rely on centralized training, requiring data from multiple institutions to be collected and stored at a single site. Such assumptions are often impractical in clinical settings due to privacy, governance, and data-sharing constraints. To address these challenges, we present, to the best of our knowledge, the first federated learning framework for automated infant movement analysis and General Movement Assessment using skeletal motion data. As a clinically relevant use case, the proposed framework is evaluated on fidgety movement classification. To quantify model confidence, Monte Carlo (MC) Dropout is employed to estimate predictive uncertainty during inference. Building upon this, we propose an Uncertainty-Aware Federated Averaging (UA-FedAvg) strategy that incorporates predictive entropy derived from MC-Dropout into the federated aggregation process, enabling client contributions to be adjusted according to their predictive uncertainty. Experiments were conducted using a cross-subject evaluation protocol under a three-client federated learning setting. Results demonstrate that federated learning substantially improves classification performance compared with independently trained local models while achieving performance approaching that of centralized training. Furthermore, UA-FedAvg and its variant incorporating validation loss generally outperform conventional FedAvg across the evaluated data-split configurations.
cs.LG / 90 / 2609.31466
Scaffold: Support Graph Theory Based Sparsification for Graph Neural Networks
Siddhartha Shankar Das, Sai Karthik Navuluru, S M Ferdous, Ryan A. Rossi, Baris Coskunuzer, Lakshman Tamil, Edoardo Serra, Alex Pothen, Robert Rallo, Mahantesh M Halappanavar
cs.LG
Abstract
Graph neural networks (GNNs) rely on message passing over graph edges, making their computational and memory costs strongly dependent on graph density. Graph sparsification offers a natural way to reduce these costs, but removing edges indiscriminately can distort important communication structure and degrade predictive performance. We introduce Scaffold, a topology-based, unsupervised graph sparsification framework derived from support graph theory preconditioners. Scaffold explicitly controls two complementary structural quantities: dilation, which measures the length of rerouting paths induced by removed edges, and congestion, which measures how strongly these rerouted paths concentrate on the retained support. By jointly controlling dilation and congestion, Scaffold preserves short communication paths while avoiding structural bottlenecks. To our knowledge, Scaffold is the first scalable GNN sparsification framework to use a joint supporting-path dilation-congestion criterion. Across 19 homophilic and heterophilic benchmarks spanning small to large graphs, Scaffold achieves the best aggregate rank among the evaluated sparsification and related methods. Using only 10%-50% of the original edges per sparse support, Scaffold recovers or closely approaches full-graph GNN performance while using less than half the memory of full-graph training and reducing end-to-end training time, including sparsification overhead. We provide an open-source software package at https://github.com/siddhartha047/Scaffold.
cs.LG / 91 / 2609.31531
HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning
Xinglong Luo, Yuding Zhang, Yuheng Kuang, Shuxuan Yuan, Zhenni Zeng, Weiqiang Zhu, Zhenhai Ji, Zhengning Wang
cs.LG
Abstract
Cooperative multi-agent reinforcement learning under partial observability and shared rewards requires assigning team outcomes to individual agents and high-order coalitions. A MAPPO-style critic compresses joint behavior into one global value, while critics that dynamically reconstruct the grouping topology change the mapping from agents and coalitions to value components as interactions or active agents evolve. We refer to this inconsistency as structural target drift. We introduce HySTAR, a MAPPO-based framework that separates adaptive representation learning from a temporally consistent high-order value-decomposition basis. HySTAR anchors an overlapping sparse hypergraph as a uniformly covered decomposition scaffold, uses a spatiotemporal encoder to represent physical and task-dependent interactions, and combines temporal and structural relevance to construct agent-specific advantages. Experiments on SMAC, GRF, Traffic Junction, and MPE demonstrate consistent improvements over MAPPO-style, value-factorization, and dynamic-grouping baselines. On the hardest SMAC settings, HySTAR achieves relative gains of 16.7\% over MAPPO and 15.6\% over HYGMA, ranks first on all six GRF scenarios, reduces Traffic Junction convergence epochs by up to 40.2\% relative to MAGIC, and obtains the highest MPE episode rewards. Controlled topology, agent-death, neighborhood, and parameter analyses support the benefit of anchoring the decomposition scaffold while adapting the propagated representations.
cs.LG / 92 / 2609.31539
NEXT: Physics-Informed Neuro-Spectral Exponential Time Differencing Architectures
Márcio Marques, Leonardo Mendonça, Leonardo M. Moreira, Christian Júnior de Oliveira, Vitor Balestro, Tiago Novello, Daniel Yukimura, Pavel Petrov, Lucas Nissenbaum
cs.LG · math.NA
Abstract
Physics-Informed Neural Networks (PINNs) build neural representations of time-dependent PDE solutions, naturally incorporating physics knowledge and observational data, which makes them well suited to both forward and inverse PDE problems. PINNs, however, are known to suffer from spectral bias and lack of causality. Neuro-Spectral Architectures (NeuSA), a recently proposed alternative to PINNs, mitigate both issues, but their numerical integration becomes unstable for stiff differential equations arising in many relevant physical problems. This study proposes Neuro-Spectral Exponential Time Differencing Architectures (NEXT), which combines the spectral representation of the PDE solution in NeuSA with high-order exponential integrators. Within this approach, the linear stiff part of the vector field induced by the PDE is integrated exactly through matrix exponentials, while the possibly nonlinear remainder is modeled by a neural network. The effectiveness of NEXT is verified through benchmark experiments on a set of stiff PDEs, in which NEXT is stable and accurate while NeuSA diverges numerically. It is also shown that NEXT can be applied to inverse problems, where the model has to learn unknown parameters or boundary conditions from sparse data. All code used in this work is publicly available at: https://github.com/marcioh2m/next.git .
cs.LG / 93 / 2609.31546
BeatGraph: Self-Supervised Heartbeat Graphs for Infant ECG Representations from the Home Environment
Mohammad Nur Hossain Khan, M. S. Krafczyk, Beverly G. Bolster, Nancy McElwain, Mark A. Hasegawa-Johnson, Bashima Islam
cs.LG
Abstract
Electrocardiogram (ECG) foundation models typically tokenize the signal into fixed-length patches that ignore cardiac structure, so a patch may split a heartbeat and the number of beats in each patch shifts with heart rate. This matters most for infants, whose heart rates are higher and whose ECG differs from the adult, clinic-recorded 12-lead data these models are built on. A model for infant ECG should therefore reason about heartbeats directly rather than recover them from arbitrary patches. We propose BeatGraph, which makes the heartbeat its unit of representation, modeling each 30-second window as a graph of beats. A shared beat encoder embeds each heartbeat from its waveform and inter-beat intervals, a Transformer with positional encoding orders the beats in time, and residual graph attention layers relate every beat to every other before attention pooling yields a window embedding. We pretrain BeatGraph on our new corpus of unlabeled infant recordings by predicting masked-beat embeddings, then fine-tune it for each task. One backbone supports sleep-wake detection, infant-state classification, activity-source identification (infant- or caregiver-initiated movement), and affect recognition, improving macro-F1 over the strongest baseline on each task by 0.076 to 0.158. It also transfers across age groups, reaching 0.892 AUROC on the ZZU-pECG pediatric benchmark (ages 0 to 14), within 0.001 of the best published self-supervised ECG model, and matching that model under linear evaluation on the adult PTB-XL benchmark despite infant-only pretraining. Finally, to our knowledge, we release the first public infant ECG corpus collected in homes, classrooms, and laboratory settings with state and affect labels. It contains 3,408 hours of single-channel ECG from 143 infants aged 3 to 11 months, with unlabeled pretraining data, benchmark tasks, and subject-level splits.
cs.LG / 94 / 2609.31559
Online Learning via Learned Latent Bayesian Tracking
Guy Gerson, Tomer Raviv, Nir Shlezinger, Tirza Routtenberg, Osvaldo Simeone
cs.LG · eess.SP
Abstract
Online learning in non-stationary environments requires models to adapt rapidly from streaming data under strict computational constraints. A principled approach casts online learning as Bayesian state tracking, where model parameters are updated sequentially via Bayesian filtering. However, applying Bayesian filters directly to modern deep models is computationally prohibitive due to the high dimensionality of parameter space, forcing existing methods to rely on restrictive approximations or manually designed low-dimensional subspaces. In this work, we identify the absence of a suitable low-dimensional dynamical representation as the core bottleneck in Bayesian filtering-based online learning. Accordingly, we propose Adaptive Update through Representation Adaptation (AURA), a meta-learning framework that learns offline a low-dimensional latent state-space model governing the evolution of optimal model parameters under distribution shift. Online adaptation is then performed via extended Kalman filtering in this learned latent space followed by reconstruction of the full model parameters through a learned lifting map, enabling efficient single-step online adaptation while preserving model expressiveness. Evaluated on online adaptation of neural wireless receivers under time-varying channels and on non-stationary image classification, AURA shows substantial improvements in adaptation speed, accuracy, and computational efficiency over existing online learning and Bayesian filtering baselines, demonstrating that an adaptation-aware latent geometry is beneficial for effective Bayesian online learning in high-dimensional models.
cs.LG / 95 / 2609.31560
Generalization behavior of OPTQ and the role of regularization
Erin George, Rayan Saab
cs.LG
Abstract
Large neural networks can be compressed by rounding or "quantizing" their weights to numbers that admit representations with fewer bits. One algorithm for quantization, OPTQ, progressively quantizes the weights of a neural network so that the squared quantization error on a specified calibration dataset is as small as possible. We study the performance of OPTQ and a variant algorithm, stochastic OPTQ, in a generalization setting and derive bounds for the expected squared error accrued by the algorithm when a test point is drawn from a fixed distribution. We prove two results. One result relates the generalization error to the error on a calibration dataset comprising independent samples from the same distribution as the test distribution. The other result bounds the generalization error of stochastic OPTQ for all sufficiently nice distributions, regardless of the calibration dataset. In both of these results, the regularization term $λ$ plays an important role. We use insights from these results to make a new recommendation for the choice of $λ$ and see that this choice of $λ$ preforms favorably in experiments when compared to prior recommendations in the literature.
cs.LG / 96 / 2609.31564
Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights
Irene Tallini, Daniele Solombrino, Alberto Cazzaniga, Emanuele Rodolà
cs.LG
Abstract
We show that neural network weights can be explicilty fintuned to admit a smaller grammar. Weight Pair Encoding (WeightPE) does so by placing a lossy Re-Pair compressor inside a straight-through estimator. The int8 weights of the network are flattened into one string, and near-matching Re-Pair patterns are made exactly equal within a global L2 budget. The network computes with the rewritten weights and trains through them with a straight-through estimator. Unlike a flat codebook of fixed-size entries, a grammar offers variable-length patterns and reuses them hierarchically inside larger ones. On the MLP weights of ViT-B/16 and ViT-L/16 finetuned on CIFAR-10, WeightPE produces a Re-Pair grammar 0.43x and 0.38x the size of the one produced by an equivalent int8 QAT run, at a cost of 1.9 and 1.1 accuracy points. The trend extends to different grammar compressors (LZ78, SEQUITUR), over which the networks has not be finetuned against. To our knowledge, this is the first time grammar size has been used as an explicit training objective for network weights.
cs.LG / 97 / 2609.31586
Trust Guided Decision Transformer
Chainesh Gautam, Raghuram Bharadwaj Diddigi, Chandramouli Kamanchi, Pankaj Dayama, Sumanta Mukherjee, Kameshwaran Sampath
cs.LG
Abstract
Decision Transformer performance degrades on long rollouts because the conditioning context drifts out of the training distribution. We show that this drift is visible through the model's own next state prediction error, which rises during rollout and stays elevated, giving a direct signal of when context has become unreliable. We introduce Trust Guided Decision Transformer (TGDT), which selects context before applying value guidance. At each step, TGDT evaluates several recent context suffixes using rolling next state prediction error, calibrated against held out offline data via split conformal prediction. It keeps only suffixes whose error stays within the calibrated threshold, then uses a frozen critic to choose the highest value action among the trusted suffixes. This reverses the order used by value only elastic selection, where the critic may choose an action generated from a context the model itself has flagged as unreliable. Experiments on D4RL navigation and locomotion tasks show that state prediction, critic guidance, and hard context reset each solve only part of the problem. TGDT reduces persistent high error runs and improves return over vanilla Decision Transformer, reset based context control, and value only context selection.
cs.LG / 98 / 2609.31589
Common-Mode Collapse and Recovery in Direct Feedback Alignment
Varun Reddy, Bernardo L. Sabatini, Houman Safaai
cs.LG · cs.NE
Abstract
Direct feedback alignment (DFA) trains hidden layers through fixed random projections of output error. With tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near the loss of a constant predictor of class frequencies. We trace this stall to the error's common mode, the component shared across inputs. An exact mean-covariance decomposition separates a rank-one update formed by the mean teaching signal and mean presynaptic activity. Its leading component drives tanh units toward saturation. At initialization, random feedback provides no systematic correction of the shared error on average; readout learning limits its duration. A reduced model initialized from the network, without fitted parameters, predicts the concentration of activation sensitivity across 48 settings. On MNIST, class decodability largely survives collapse, but readout learning remains slow at a fixed learning rate. Adam learns faster despite deeper collapse. Calibrating the baseline readout to the class prior suppresses collapse and speeds learning; weaker feedback trades less collapse for slower learning. Replacing errors by their signs sustains collapse; subtracting the signal's batch mean prevents sustained collapse and improves learning in the tested setting. Related effects occur in deeper and convolutional networks and on CIFAR-10, with severity and cost depending on the readout, optimizer and input statistics.
cs.LG / 99 / 2609.31141
Unknown-Traffic Detection, Calibration and Shortcut Reliance in Distilled Encrypted-Traffic Classifiers over One Year
Mahmoud Abbasi
cs.NI · cs.LG
Abstract
Knowledge distillation is the standard way to compress encrypted-traffic classifiers for the edge, and almost all such work judges students by accuracy alone. We ask what else a student inherits: unknown-traffic detection, calibration, shortcut reliance, and whether any survives a year of drift. Resemblance proves little on its own, since soft targets also regularise. We therefore distil one 101k-parameter student from two teachers of equal accuracy but different construction, a five-member ensemble and a single wider model, so that following one rather than the other is attributable to it. The design was pre-registered before any test result was seen. We tested ten hypotheses on CESNET-TLS-Year22, a year of real TLS traffic, across 18 test windows over 35 weeks. Two are supported: a student's per-flow unknown-scores shift toward its own teacher, but only at a conventional temperature, not the accuracy-optimal one; and a shortcut-reliant teacher passes its over-confidence to a student that never sees the feature. The drift prediction is reversed under both scores, the gap narrowing rather than widening and the student overtaking under the energy score in two of three replicates, as is the prediction that such a teacher harms its student's detection, which improves slightly. Shortcut reliance is set by model size, not distillation. Under the logit-based scores nothing else transfers: distillation beats neither a temperature-scaled direct student nor label smoothing. Exploratory analysis shows this turns on the scoring rule: with a feature-space detector the teacher detects unknown traffic 0.073 AUROC better than the direct student, where the energy score sees 0.000, and the conventional-temperature student inherits most of it. Label smoothing, with no teacher, recovers more. Distillation transfers the teacher's habits; what looks like an inherited ability is available without one.
cs.LG / 100 / 2609.31025
Precision at Speed: Sample-Efficient Online Model-Based Reinforcement Learning for Hydraulic Excavator Control
Claudio Canales, Fang Nan, Marco Hutter, Javier Ruiz-del-Solar
cs.RO · cs.LG · eess.SY
Abstract
Precise, high-speed control remains challenging for robots with complex actuation dynamics. Learning directly on hardware is further constrained by the cost of real-world interaction. We present an online model-based reinforcement learning framework that learns a probabilistic dynamics ensemble model from scratch for sampling-based model predictive control. A precision-gated contouring objective conditions the progress reward on path accuracy, prioritizing precision over speed. In a data-driven excavator simulator, the framework achieves higher sample efficiency than the evaluated model-based reinforcement learning baselines. We validate the framework by learning directly on an 11.5-ton Menzi Muck M445 hydraulic excavator, without demonstrations or simulation pretraining. After 20 minutes of interaction, the controller reaches tracking accuracy comparable to prior learned controllers trained on 100-150 minutes of data. After 40 minutes, it sustains sub-centimeter mean path error at high operating speeds.
cs.LG / 101 / 2609.31024
Synth-JEPA: Joint Embedding Prediction for Renderer-Free Synthesizer Parameter Search
Ben Hayes, Haokun Tian, Stefan Lattner
cs.SD · cs.LG · eess.AS
Abstract
Sound matching can be formulated as optimizing synthesizer parameters against an audio-domain objective. However, objectives derived from generic audio representations are often difficult to optimize, while direct search requires rendering every candidate. We introduce Synth-JEPA, which learns mutually predictive audio and parameter representations from paired synthesizer data. At inference, candidate parameters are scored directly in this learned space, yielding a renderer-free objective whose audio geometry is shaped by parameter correspondences rather than generic audio similarity. We evaluate Synth-JEPA on Surge XT using held-out synthesizer sounds and out-of-domain NSynth and FSD50K targets, against inverse models, direct search, and learned proxy objectives. Synth-JEPA outperforms all baselines in-domain and remains competitive out-of-domain. Its matching quality continues to improve with additional test-time search, allowing compute to be traded for match quality. In pairwise listening tests, listeners preferred Synth-JEPA in 85% of trials overall. Together, these results show that an audio representation with a parameter-induced geometry allows synthesizer sound matching to be approached as an effective renderer-free search problem.
cs.LG / 102 / 2609.31165
BreathGRU: A Novel Semi-Supervised Bidirectional Gated Recurrent Unit Framework for Speech and Breath Segmentation for Respiratory Audio
Sania Fatima Sayed, John W. Holloway, Reyer Zwiggelaar, Faisal I. Rezwan
cs.SD · cs.LG · eess.AS
Abstract
Speech-breath segmentation is a fundamental preprocessing step in respiratory audio analysis, enabling applications such as respiratory acoustic biomarker extraction, lung function prediction and disease monitoring. Existing approaches, including threshold methods, Fourier Transform-based techniques, and unsupervised and pretrained voice activity detection (VAD) models, primarily focus on speech detection and often classify breathing events as non-speech or silence, limiting their applicability for precise breath detection. To address this limitation, we propose BreathGRU, a semi-supervised Bidirectional Gated Recurrent Unit (BiGRU) framework specifically designed for speech-breath segmentation. The proposed framework combines frame-level acoustic feature extraction with bidirectional recurrent modelling, pseudo-label refinement and duration-constrained Segmental Viterbi decoding to produce speech and breath segmentation. BreathGRU was evaluated against the existing approaches, using manually annotated recordings. Performance was assessed using event-based, time-based, overlap-based, duration-based and boundary-based segmentation metrics. Experiment results demonstrated that BreathGRU achieved the highest breath event recall (0.83), the lowest onset-localisation error (0.14s) and the highest Mean Match Intersection over Union (0.81), with competitive overall segmentation performance compared to large pretrained VAD models like Silero. Qualitative evaluation on manually annotated recordings further showed close agreement between BreathGRU and manual annotation, with better breath detection compared to Silero. These findings demonstrate that explicit breath event modelling provides advantages over general-purpose VAD models and establish BreathGRU as an effective speech-breath segmentation framework which can be applied for respiratory audio analysis and pulmonary healthcare applications.
cs.LG / 103 / 2609.31402
AFA-Net: A Differential Attention Approach for Auditory Attention Detection
Philip H. Lee, Shreeram Suresh Chandra, Karan Thakkar, John H. L. Hansen
cs.SD · cs.LG · eess.SP
Abstract
Auditory Attention Detection (AAD) utilizes electroencephalographic (EEG) signals to identify a target speaker in a multi-speaker environment. Despite considerable progress, existing deep learning architectures often lack explicit mechanisms for handling noisy EEG data. To address this limitation, we propose Auditory Focus Attention Networks (AFA-Net), a machine learning framework that replaces vanilla attention with a simple yet flexible differential attention mechanism to help focus on task-relevant neural activity. AFA-Net achieves an upward accuracy of 96.8% at the 2s decision window, while using substantially fewer parameters than most existing methods. To the best of our knowledge, AFA-Net is among the first frameworks to explicitly try to combat EEG noise to improve AAD.
cs.LG / 104 / 2609.30517
Seeing Speech: Learning Visible Articulatory Dynamics for Speech-Driven 3D Facial Animation
Hyung Kyu Kim, Byungchan Hwang, Hak Gu Kim
eess.AS · cs.GR · cs.LG
Abstract
Recent progress in speech-driven 3D facial animation has improved vertex-level reconstruction quality, but speech-consistent visible articulation remains difficult. This is because speech production follows structured and constrained articulators' coordination and the mapping from acoustics to motion is inherently one-to-many. Motivated by the structured patterns of visible articulation, we propose a novel articulation-aware framework that models visible speech through directional articulatory motions and composes them into surface-consistent 3D facial motion. To represent visible articulation with three directional articulatory motions, spreading, opening, and protrusion, we propose a Speech--Articulatory Memory (SAM) that captures the correspondence between speech and these motions under phonetic context through retrieval and decoding based on a key-value memory structure. Then, a Topology-aware Articulatory Composition (TAC) integrates the predicted directional articulatory motions under mesh topology to produce surface-consistent 3D facial motion. Experiments on VOCASET and TFHP show that our method achieves state-of-the-art performance on standard reconstruction metrics and improves visible articulatory distance and velocity errors for lip articulation, while a user study confirms clear preference in lip sync and realism.
cs.LG / 105 / 2609.30975
A Comprehensive Study of Content Representations for Speech Synthesis
Diego Torres, Axel Roebel, Nicolas Obin
eess.AS · cs.LG · cs.SD
Abstract
Speech content representations are central to voice conversion, speech-to-speech translation, and multimodal language models, yet they are rarely compared under a common generative framework that directly measures what each representation contains. We address this by training a generative model conditioned solely on each representation and evaluating the generated audio along the content, speaker identity, and prosody axes. Across SSL features, supervised tokens, posteriorgrams, and neural audio codecs, we find two distinct regimes: representations that nearly reconstruct the original audio, and representations that effectively disentangle speaker identity. These results show that disentanglement depends not on supervision alone, but on the interaction between the training objective and the representation's information capacity: supervised representations only disentangle speaker identity when their capacity is sufficiently constrained.
cs.LG / 106 / 2609.30629
FRESHLATENT: Channel-Aware Latent Adaptation for Resource-Constrained Embodied VLM Perception
Rajat Bhattacharjya, Minwoo Kim, Arnab Sarkar, Tamoghno Das, Sing-Yao Wu, Eli Bozorgzadeh, Marco Levorato, Nikil Dutt
eess.SP · cs.CV · cs.DC · cs.LG · cs.RO
Abstract
Mission-critical UAVs increasingly rely on split vision-language model (VLM) perception under tight onboard-resource and wireless-communication constraints. However, corruption of transmitted intermediate features creates a deployment mismatch for clean-trained split interfaces, while stronger channel-aware codecs can impose substantial onboard cost. We present FreshLatent, a lightweight channel-aware latent adapter that trains a power-normalized encoder-decoder through wireless corruption while keeping the surrounding VLM frozen. We formulate deployment around a mission-conditioned perception requirement and embedded interface cost, linking channel quality and communication budget to the operating conditions under which perception remains usable. At 0 dB and the tightest communication budget, FreshLatent improves gIoU and cIoU over clean split compression by 20.79 and 20.87 points, respectively. At the most adverse evaluated SNR (0 dB), across all three communication budgets, FreshLatent recovers 63.5-69.1% of the gIoU improvement achieved by a much heavier, range-trained feature-JSCC codec. On an NVIDIA Jetson AGX Xavier in 10-W mode, FreshLatent uses 37-40x fewer encoder parameters, 7.7-9.9x lower edge-interface latency, and 8.8-10.0x lower edge-interface energy than the heavier codec. Together, these results show that lightweight channel-aware adaptation can recover a substantial fraction of the robustness of a much larger communication interface while broadening quality-valid operation under constrained wireless conditions.
cs.LG / 107 / 2609.30859
AC Power Flow Contingency Analysis Using a Single Deep Neural Network
Md Obaidur Rahman, Junjie Qin, Vassilis Kekatos
eess.SY · cs.LG
Abstract
Contingency analysis using the AC power flow (AC-PF) model is a critical tool for accurate grid security assessment, but its computational burden increases with the number of operating scenarios and outage configurations to evaluate. Recent ML-based approaches typically require outage-specific training data, leading to offline training costs that scale with the number of contingencies. This work proposes a framework that reuses a single ML model trained solely on basecase AC-PF data to estimate post-contingency operating states under arbitrary single-line outages. The proposed approach formulates post-contingency state prediction as a fixed-point iteration. If the ML model is a deep neural network (DNN), we derive sufficient conditions that guarantee convergence and develop semidefinite programming (SDP) formulations to certify these conditions for a given DNN. Numerical tests on the IEEE 118-bus system demonstrate that the proposed SDP formulations are tight, that the certified conditions hold for all tested contingencies, and that the resulting method produces accurate post-contingency state estimates within only a few iterations.
cs.LG / 108 / 2609.31469
LandscapeSHAP: Which Persistent Homology Class Gets the Credit?
Nikola Milićević
math.AT · cs.LG
Abstract
Shapley values, a solution concept from cooperative game theory, have recently become a standard tool for feature credit allocation in machine learning. They provide an axiomatically justified method to fairly distribute a model's prediction among the data features. Shapley values have not yet been applied to explain machine learning models trained on features from topological data analysis. We develop what we believe is the first such approach, focusing on the persistence landscape featurization of persistence diagrams. Because each landscape coordinate is a rank statistic, crediting a model's prediction back to individual persistent homology classes (persistence diagram points) is nontrivial. We introduce LandscapeSHAP, a method for fair credit allocation to persistence diagram points based on a model's prediction. For linear models on persistence landscapes, LandscapeSHAP has a closed form expression that gives the exact Shapley value of every persistence diagram point. In particular, there is no coalition sampling required. We further prove that the four Shapley "fairness" axioms uniquely characterize this credit allocation for any model, not only linear ones. For a general nonlinear model, this unique value can only be calculated exactly from its defining coalition averaging formula, which requires considering all $2^N$ many coalitions, where $N$ is the number of points in the persistence diagram. This is computationally intractable for persistence diagrams of realistic size. We complement the exact linear model result with an efficient Monte Carlo sampling of persistence diagram coalitions. We give convergence rates in terms of number of samples needed to approximate to a desired degree of accuracy. We also prove stability results for the LandscapeSHAP credit allocation, for any model.
cs.LG / 109 / 2609.30809
Deep-Learning Solvers and Surrogates for Infinity and p-Laplace Problems
Tak Shing Au Yeung, Ka Chun Cheung, Hannah Potgieter, Steven J. Ruuth, Simon See
math.NA · cs.LG
Abstract
We investigate the use of neural network solvers for infinity and $p$-Laplace problems, which are fundamental in nonlinear analysis and have practical applications. Our approach employs Physics-Informed Neural Networks (PINNs) and Deep Operator Networks (DeepONets) to address computational challenges associated with large $p$ values, ranging from $2$ to $1000$, on various 2D and 3D domains. Our method offers advantages over traditional physics-based solvers, especially in three dimensions where mesh-based solvers become very costly for these problems. We also establish conditional convergence results for PINN approximations of both problems and a universal approximation result for DeepONet on the parametric $p$-Poisson problem. We demonstrate the effectiveness of these neural network solvers through numerical experiments and compare their performance with conventional methods.
cs.LG / 110 / 2609.30732
TR-SSQP: A Trust-Region Method for Constrained Stochastic Optimization under Heavy-Tailed Noise
Haoxuan Wang, Yuchen Fang, Sen Na
math.OC · cs.LG · stat.CO · stat.ML
Abstract
We consider stochastic nonlinear optimization problems with deterministic equality constraints. While unconstrained stochastic optimization is well understood, the interplay between optimality and feasibility in the constrained setting poses significant challenges. Moreover, existing theoretical guarantees for constrained stochastic methods predominantly rely on bounded-variance assumptions, leaving the heavy-tailed noise regime largely unexplored. To address this gap, we propose a novel trust-region method within the stochastic sequential quadratic programming framework, termed TR-SSQP. Our method employs a normal-tangential decomposition in the step computation to balance optimality and feasibility. In addition, we incorporate a normalization mechanism in the design of the trust-region radius, together with Polyak momentum for gradient estimation, ensuring stable updates without gradient clipping. When the trust-region radius and the momentum parameter decay at appropriate rates, we establish global almost-sure convergence of the method. To the best of our knowledge, this is the first asymptotic convergence result for constrained stochastic optimization under heavy-tailed noise. We demonstrate the promising performance of the proposed method through extensive numerical experiments, including comparisons among its variants and with existing constrained stochastic optimization methods.
cs.LG / 111 / 2609.30877
Tight Stochastic Condition-Number Dependence in Nonconvex-Strongly-Concave Minimax Optimization
Qihao Zhou
math.OC · cs.LG
Abstract
We study whether the linear condition-number dependence in the stochastic complexity of SAPD+ is necessary for nonconvex-strongly-concave minimax optimization. For jointly $L$-smooth objectives with dual strong-concavity parameter $μ$, we prove a lower bound that matches the SAPD+ upper bound under the same Moreau-envelope stationarity criterion and the same primal-dual initialization gap. Specifically, when $σ\ge\varepsilon$, the worst-case complexity of zero-respecting algorithms is $Θ(κLGσ^2\varepsilon^{-4})$ in the stated accuracy regime, where $κ=L/μ$, $G$ bounds the initial primal-dual gap, and $σ^2$ bounds the variance of a general unbiased first-order oracle. The lower bound is realized on a smooth problem class with a bounded dual box. Our construction routes each link of a nonconvex zero-chain through a dual gradient of magnitude proportional to $\varepsilon/\sqrtκ$, while an undiscovered primal coordinate prevents stationarity. It also yields the primal-gradient lower bound $Ω(LΔ(\sqrtκ\varepsilon^{-2}+κσ^2\varepsilon^{-4}))$ after combination with the known deterministic bound, where $Δ$ bounds the initial primal function gap.
cs.LG / 112 / 2609.30885
Retraction-Based Gradient Projection Algorithms on Manifolds
Conglong Xu, Hao Wu
math.OC · cs.LG
Abstract
We introduce a framework for retraction-based convex optimization on Riemannian manifolds, which includes a notion of retraction-specific convex sets and retraction-based gradient projection algorithms. The standard theory of gradient projection algorithms generalizes easily to this framework. Within this framework, we establish convergence results for retraction-based gradient projection algorithms with various stepsize rules. As an application, we use our framework to study the weighted low-rank approximation. We also provide numerical validation of our convergence results on the image completion task.
cs.LG / 113 / 2609.30477
Scaffold-Constrained Subset Dynamic Programming for Exact SSE Clustering
Yordan P. Raykov, Max A. Little
math.ST · cs.LG
Abstract
Exact Euclidean \(K\)-means partitions \(n\) observations into \(K\) unlabelled clusters, but the unrestricted search is generally exponential. We use data-derived geometric graphs to precondition an exact subset dynamic program: as a result only connected vertex subsets are admitted as clusters, while sum-of-squared-errors (SSE) loss is unchanged. A remaining-set recurrence minimises fixed-\(K\) or penalised SSE, with exact factorisation over the connected components of each remaining set. The central question we study is how much computational support can be removed while preserving an unrestricted optimum. Graph inclusion gives monotone coverage and support relations, and a bottleneck threshold identifies the first covering graph in a nested hierarchy. For fixed \(K\) and dimension, under compact ball support and density bounds, retaining \(q=O(\log n)\) nearest neighbours per observation preserves an empirical SSE optimum with probability tending to one, using an \(O(\log n/n)\) fraction of complete-graph edges. Truncated Gaussian mixtures with unequal weights and covariances satisfy these conditions. The rate we provide is a sufficient upper bound rather than a result implying polynomial optimisation complexity. Objective-matched synthetic and full-data comparisons assess coverage, compression, and reference-label agreement. As a secondary application, we illustrate how the proposed scaffold preconditioning can be utilized to improve the efficiency of split-merge proposals that preserve unrestricted mixture posteriors.
cs.LG / 114 / 2609.30688
On the Limits of Univariate Deep Learning for Significant Wave Height Forecasting
Yilin Zhai, Hongyuan Shi, Zaijin You
physics.ao-ph · cs.LG · physics.comp-ph
Abstract
This study conducts a systematic hyperparameter search across five deep learning architectures, DLinear, LSTM, PatchTST, ResAttLstm, and Mamba2, and nine context lengths (1-168 h) for single-station significant wave height (Hs) forecasting on NDBC buoy 41009, followed by re-evaluation of the best configurations on a 47-buoy, 37-year corpus. The five families converge to a common performance level on the multi-buoy evaluation (between-family SD = 0.0014 m^2, 0.8% of the grand mean), a spread dwarfed by the 4.83x cross-dataset MSE shift between buoy corpora. All multi-buoy trials beat persistence (mean skill +0.062), but no architecture consistently outperforms the others. On the single-buoy experiment, skill peaks at 12-24 h where five trials fall below persistence, per-family Q4/Q3 test MSE ratios range from 2.4 to 2.6, and deep models underperform persistence for the most extreme 1% of waves. These findings are consistent with the interpretation that persistence already captures the dominant linear-inertial signal in univariate Hs, and that architecture engineering under this univariate input setting has reached diminishing returns: cross-buoy variance, not model class, dominates forecast error. Future work should prioritise atmospheric covariates, zero-shot cross-buoy transfer, and decomposition of Hs into swell and wind-sea components. By establishing a rigorous reference baseline for what univariate Hs models can and cannot achieve, this study provides a benchmark against which future multivariate and physics-informed approaches can be calibrated, and offers practical guidance for lightweight buoy-level forecasting in mid-latitude storm-dominated and swell-mixed environments.
cs.LG / 115 / 2609.31483
Scaling Density Functional Theory with Gaussian Splatting
Andrés Guzmán-Cordero, Cindy Zhang, Majdi Hassan, Marta Skreta, Kirill Neklyudov, Matija Medvidović
physics.chem-ph · cond-mat.mtrl-sci · cs.LG · physics.comp-ph
Abstract
Density functional theory (DFT) strikes a practical balance between accuracy and computational cost in many problems of computational chemistry and materials science. However, many DFT calculations are limited by fixed atom-centered basis sets, which dictate how accuracy and cost scale with system size. We propose Gaussian Splatting for Density Functional Theory (GS-DFT), which represents molecular orbitals as a cloud of Gaussians whose positions, shapes, and mixing coefficients are optimized jointly by gradient descent to minimize the energy without training data. Conceptually, GS-DFT is 3D Gaussian splatting with the renderer replaced by quantum mechanics. We introduce two key solver components: adaptive density fitting with screening for efficient evaluation of two-electron integrals, and a regularized differentiable orthogonalization of the molecular orbitals. Empirically, the optimized basis reaches the accuracy of the largest conventional basis sets with a fraction of the parameters, converging systematically in energy, density, and nuclear forces. At equal parameter count, it captures the stretched-bond and anion physics that fixed bases only recover with specialized basis augmentation. The resulting solver exhibits quadratic peak memory scaling in the cloud size, allowing us to simulate systems of up to 2,742 atoms (10,406 electrons) without any modifications at triple-zeta scale using a single four-GPU node.
cs.LG / 116 / 2609.30396
An End-to-End Pipeline for Causal ML with Continuous Treatments: An Application to Financial Decision Making
Javier Moral Hernández, Clara Higuera-Cabañes, Álvaro Ibraín
stat.ME · cs.LG · stat.AP · stat.ML
Abstract
This paper presents an end-to-end causal machine learning (ML) pipeline designed for real-world applications with continuous treatments. The proposed framework consists of six sequential steps: dimensionality reduction, causal identification, positivity assumption violation handling, estimation, refutation and evaluation, and policy optimization. We introduce practical contributions not currently available in existing causal ML toolkits, specifically: (1) a method for detecting and quantifying positivity violations in continuous treatment settings (2) a novel, scalable two-stage dimensionality reduction framework tailored for causal inference with high-dimensional data; (3) the adaptation of sensitivity analysis and estimation methods originally designed for binary treatments to the continuous treatment space and (4) an end-to-end integration of these components into a modular, reproducible workflow. These innovations address real-world challenges in causal inference that are often not covered in theoretical frameworks but frequently encountered in industrial applications. The methodology is validated with a synthetic dataset inspired in a real-world financial debt collection use case, however its design can be applied to analogous problems across different industries. Results demonstrate that the proposed methodology offers a more computationally efficient approach and produces less biased estimates compared to standard methods for problems with continuous treatment and high-dimensional data. A fully functional GitHub repository with documented code and numbered notebooks is made available ensuring reproducibility and practical implementation. The pipeline presented is intended to contribute to closing the gap between academic approaches and practical application in industry contexts where causal ML can be highly beneficial such as the financial sector.
cs.LG / 117 / 2609.30445
Bayesian Uncertainty Quantification for fMRI Functional Connectivity via Simulation-Based Inference
Simon Carter, Zeming Kuang, Lilianne R. Mujica-Parodi, Helmut H. Strey
stat.ML · cs.LG
Abstract
Optimizing fMRI scan duration and spatial resolution is critical for experimental design, yet traditional correlation-based approaches cannot quantify uncertainty or disentangle scanner measurement noise from true neural variability across subjects. Without principled uncertainty bounds, researchers cannot know whether a protocol is long enough to reliably estimate connectivity, or whether between-subject differences reflect biological variation or noise. We present a Bayesian framework modeling BOLD dynamics as coupled Ornstein-Uhlenbeck processes, using Sequential Neural Posterior Estimation to obtain connectivity posteriors while accounting for frequency-independent measurement noise across the BOLD spectrum. Applied to N = 28 healthy controls (55 scans) at 7T using a functional network atlas (65 DMN regions), the framework quantifies uncertainty across its sources: scanner noise, subject variability, and acquisition length. Spatial analysis identifies a mean of 46 voxels per ROI, roughly half of typical region sizes, as sufficient to achieve 90% of asymptotic precision. At the single-subject level, 7T reaches its within-session precision plateau in approximately 7 minutes versus 10 minutes for 3T, a 40% reduction in required scan time, providing the first direct, model-based quantification of the scan-time advantage conferred by higher field strength. At the population level, 3T requires roughly 37 times more per-subject scan time than 7T for the pooled curves to converge, confirming a consistent advantage of higher field strength at every timescale. Together these findings provide concrete, scanner-specific guidance for protocol optimization, with direct implications for reducing acquisition costs and improving the reliability of connectivity-based clinical biomarkers. We provide code enabling researchers to derive these bounds from their own data.
cs.LG / 118 / 2609.30499
Ordinary Nonconvex SGD under Distance-Dependent Moments: Finite-Horizon Stationarity and Nagaev Bounds
Wei Biao Wu
stat.ML · cs.LG
Abstract
Uniform noise-moment bounds exclude stochastic gradients whose variability increases with the iterate. We study ordinary, single-sample stochastic gradient descent for smooth, lower-bounded, possibly nonconvex objectives under distance-dependent conditional moments. Under second moments alone, a direct descent--displacement argument yields $T^{-1/3}$ expected average squared-gradient stationarity with a horizon-dependent stepsize. An explicit oracle-complexity corollary matches the known smooth Blum--Gladyshev (BG-0) lower bound, including the $Lb_2Δ^3\varepsilon^{-6}$ and $LΔσ^2\varepsilon^{-4}$ stochastic terms, where $Δ$ is the initial objective gap and $σ^2+b_2\|x-x_1\|^2$ bounds the variance. Thus unchanged SGD attains the minimax stochastic complexity in this second-moment class. For $p>2$, predictable localization and a Hilbert-space Fuk--Nagaev inequality yield a high-probability bound separating logarithmic variance and polynomial rare-shock contributions. The localization radius is derived from the recursion: no bounded-iterate assumption, clipping, normalization, momentum, or increasing batch size is needed. We also give increasing-confidence rates, an objective-gap-growth refinement recovering root-$T$ stationarity, and stochastic $L^p$-Lipschitz examples. The broad BG-0 optimality statement is distinguished from the smaller mean-square-smooth class, in which additional oracle structure permits faster algorithms.
cs.LG / 119 / 2609.30643
MARCEDES: Score-based causal discovery under non-Gaussianity with continuous optimization
Anamitra Chaudhuri, Anirban Bhattacharya, Yang Ni
stat.ML · cs.LG · stat.CO · stat.ME
Abstract
We consider the problem of learning the underlying causal directed acyclic graph (DAG) structure corresponding to a structural equation model (SEM) with non-Gaussian errors. Motivated by an intentionally misspecified non-Gaussian SEM with all Laplace errors, we first introduce the mean absolute residual risk, defined over the space of all real matrices, and show that, asymptotically, the risk of the true weighted causal DAG matrix is strictly smaller than that of any other matrix. Nevertheless, to enhance generality and account for high-dimensional and finite-sample settings, we further incorporate row-specific sparsity penalties along with a soft DAG constraint to derive a continuous score function over the space of real matrices. Accordingly, we propose a score-based DAG learning method, named MARCEDES, formulated as an unconstrained score minimization problem, which can be efficiently solved using gradient-based optimization techniques, thereby circumventing the challenges associated with constrained optimization. Furthermore, we develop a computational algorithm to handle the non-smoothness of the score objective and to enable optimal tuning of row-specific sparsity penalties under a generalized Bayes framework. Finally, we demonstrate the efficiency and improved performance of the proposed method over existing approaches through an extensive simulation study.
cs.LG / 120 / 2609.30713
Parameter Estimation for Unnormalized Discrete Models via Empirically Localized Deformed Bregman Divergence
Takashi Takenouchi
stat.ML · cs.LG
Abstract
Estimation of parameter of probabilistic models is an important task in the field of machine learning.For models of discrete variables, calculation of the normalization constant of model is sometimes difficult and a lot of researches have been done to avoid the calculation of the normalization constant. In this paper, we tackle with the difficulty by combining a technique of empirical localization and a deformed Bregman divergence.The technique of empirical localization makes it possible to drastically reduce computational cost of the calculation of the normalization constant, and in addition, appropriate choice of the deformation for the Bregman divergence can invest the proposed estimator with various kinds of favorable statistical properties, such as efficiency or robustness against outlier noise.
cs.LG / 121 / 2609.30886
Conformal Prediction under Exponential-Tilt Joint Shift
Seungjin Choi
stat.ML · cs.LG · stat.ME
Abstract
Conformal prediction can lose coverage when the data distribution changes after deployment. We study adaptation using labeled source data and unlabeled target inputs, allowing both the input distribution and its relationship with outcomes to change. We use Exponential Tilt Reweighting Alignment (ExTRA), introduced for classification by Maity et al. (2023), to estimate structured distribution shifts. We compare using its estimated weights in conformal calibration with additionally tilting the source predictive distribution. Shared learned predictors, estimated weights, calibration samples, and test observations isolate the effect of tilting. Existing theory gives both procedures target coverage with true weights and a common coverage bound with estimated weights. Identification calculations and an analysis of how scoring interacts with weight estimation error help explain why their performance can nevertheless differ. In a synthetic regression setting where the assumed models match the data-generating process and target inputs are informative about the shift, tilting reduces mean set length by about $30\%$ relative to weighting alone, with both methods attaining coverage near nominal. Tilting can instead cause substantial coverage losses in synthetic classification and in regression when target inputs provide little information about the response shift. Real-data experiments also show no consistent benefit. Good coverage from weighted calibration alone does not ensure that adding predictive tilting will preserve coverage. Deciding when to apply this additional adjustment using only source labels and target inputs remains an open problem.
cs.LG / 122 / 2609.31303
Geometric Moment Contraction for Stochastic Nesterov Acceleration
Wei Biao Wu
stat.ML · cs.LG
Abstract
We study geometric moment contraction (GMC) of the constant-parameter stochastic Nesterov recursion \[ Y_k=Θ_k+β(Θ_k-Θ_{k-1}),\qquad Θ_{k+1}=Y_k-γG(Y_k,X_{k+1}). \] Under mean strong monotonicity and stochastic $L^p$ Lipschitz continuity, an explicit Perron comparison proves synchronous $L^p$ contraction when $βγL_p<(1-β)(1-q_{γ,p})$. This direct criterion includes infinite-variance gradients for $1<p<2$, but its small-step regime requires $β<μ/(μ+L_p)$. A complementary power-Lyapunov argument establishes a positive, generally much smaller, step-size interval for every fixed $β<1$ and every $p>1$, using only a finite $p$th gradient moment. At $p=2$, a simpler explicit certificate gives \[ 0<γ<\frac{2μ(1-β)^2}{L_2^2(1-β+2β^2)}. \] Its quadratic high-momentum scaling is a limitation of the chosen metric, not a sharp stability boundary. We quantify this loss, provide a general mean-only quadratic $S$-procedure, and exploit endpoint Lyapunov inequalities under stronger samplewise sector information. Verified endpoint certificates can be orders of magnitude less conservative than the explicit metric.
cs.LG / 123 / 2609.31368
Equation discovery with Bayesian tree-adjoining grammars
Christopher A. Lindley, Nikolaos Dervilis, Keith Worden
stat.ML · cs.LG · eess.SY · stat.CO
Abstract
Tree-Adjoining Grammars (TAGs) have recently been introduced to Nonlinear System Identification (NLSI) as a means of encoding an entire model class as a finite set of grammatical rules, from which candidate models are assembled as trees. Existing TAG-based identifiers rely on evolutionary optimisation and return point estimates of the model structure. This paper instead proposes the TAG framework within a Bayesian setting. A generative prior is defined over tree structures and their parameters, and a Reversible-Jump MCMC sampler with structure-preserving tree moves is used to infer the joint posterior over model structure, parameters and predictions. Two training objectives are considered; that is, a one-step-ahead objective with conjugate parameter proposals, and a simulation-based objective handled by likelihood-free inference. The approach is validated on a simulated polynomial NARX system, the Silverbox benchmark, and wave-loading data from the Christchurch Bay Tower, where embedding Morison's equation as a fixed initial tree yields a grey-box model that outperforms the physics-driven baseline. The results demonstrate that Bayesian TAGs are well suited to quantifying uncertainty in equation discovery for dynamical systems and to fitting physics-informed models.
cs.LG / 124 / 2609.31470
Beyond Empirical Support: Structured Outlier Generation via Sinkhorn Optimal Transport
Haixiang Sun, Andrew L. Liu
stat.ML · cs.LG · math.OC
Abstract
Outliers are essential for evaluating and improving the robustness of machine learning systems, especially when future distributions may differ significantly from historical training data. In high-stakes applications, robustness often depends on rare cases that finite datasets fail to capture, making simple resampling or perturbation insufficient for stress scenario generation. Existing outlier synthesis methods typically rely on sparse neighborhoods, low support latent regions, or classifier boundary crossings, which can be heuristic, unstable, and tied to specific modalities or architectures. We therefore propose Sinkhorn Boundary Outlier Generation (SBOG), a structured framework for latent-space outlier generation that couples Sinkhorn optimal transport geometry with distributionally robust boundary modeling. The resulting Sinkhorn-induced support cost guides the sampler toward weakly supported boundary regions, while semantic constraints prevent uncontrolled drift from the intended context, yielding controlled deviations from the in-distribution reference measure rather than arbitrary sparse-region samples. Experiments on time series anomaly generation and image outlier synthesis show that our framework produces informative, semantically controlled outliers and improves downstream robustness evaluation across modalities, providing a foundation for stress scenario generation beyond empirical support.
cs.LG / 125 / 2609.31570
Uncertainty and Explainability in Deep Rough Volatility: A Neural Information-Theoretic Posterior Approach
Damiano Brigo, Raphaël Huser, Dan Leonte
stat.ML · cs.LG · stat.AP · stat.CO · stat.OT
Abstract
Deep learning has substantially accelerated the calibration of complex stochastic-volatility models, but neural point calibration alone does not capture the uncertainty remaining after an implied-volatility (IV) surface has been observed. We develop a simulation-based inference framework for rough Heston (rHeston) calibration that learns the posterior distribution of the model parameters conditional on an IV surface. Using neural ratio estimation, we obtain calibrated posterior samples that can be propagated through heteroscedastic neural surrogate pricers for path-dependent exotic options. The resulting posterior-predictive distributions combine residual parameter uncertainty with conditional surrogate uncertainty and yield uncertainty-aware price intervals. We further introduce Hellinger-SHAP, an information-theoretic explainability method for posterior inference. Rather than attributing a single parameter point estimate, it applies local-background Kernel SHAP to a posterior-information functional measuring contraction from the prior to the posterior. This identifies maturity--moneyness regions associated with posterior information gain for individual rHeston parameters. In a simulation study, posterior-predictive intervals provide calibrated or conservative coverage across forward-start, barrier, and realized-variance claims, while point plug-in prices can be materially unreliable for selected contract regimes. Together, the UQ and XAI analyses provide a transparent framework for uncertainty-aware neural calibration and downstream exotic pricing under the specified prior-predictive model.
神经与进化计算 (cs.NE)
3
cs.NE / 1 / 2609.30731
Structured Bayesian Modeling of Dynamic Receptive1 Fields in Salamander Retinal Ganglion Cells
Alokesh Manna
cs.NE · stat.ME
Abstract
Neurons in the visual system are selective for specific spatial and temporal stimulus features, described by their \emph{receptive field}. Estimating one means a coefficient per pixel per time bin from few trials -- a high-dimensional problem requiring regularization. Sparse regularizers such as the LASSO handle the dimension but select pixels independently at each time point, with nothing to keep the region coherent in space or smooth in time; it can fragment or reorganize discontinuously even when the true response evolves smoothly, a failure since this evolving pattern is what a receptive-field estimate should capture. We formulate dynamic receptive-field estimation as a high-dimensional Bayesian problem: a Poisson model combining a Gaussian Markov random field in space with an autoregressive process in time, so the estimated field is smooth and coherent across space and time. On recordings from $155$ salamander retinal ganglion cells, fitting this model independently per neuron recovers a coherent surface, where a pixel-level Poisson-LASSO comparison instead returns a fragmented one. Summarizing each neuron's surface by its space-averaged temporal response and clustering these curves with a model-based functional-clustering procedure, BIC selects three balanced temporal-response phenotypes ($85$, $32$, $38$ neurons), against a degenerate grouping from clustering the raw surfaces. A simulation study with known ground truth confirms the same pattern, with the model beating an unregularized Poisson GLM, LASSO, and the elastic net on recovery and estimation accuracy, though LASSO controls false positives better. The per-neuron field identification, its contrast with LASSO, and the functional-clustering population typing constitute this paper's contribution.
cs.NE / 2 / 2609.30938
Landscape Limits of Quantum-Inspired Evolutionary Optimization across 256 continuous functions
Rishi Govind, Ferdin Sagai Don Bosco, Kasturi Venkata Srikanth, Aman Mittal, Abhishek Singh, Aditya Singh, Abhishek Chopra
cs.NE · math.OC
Abstract
Quantum-inspired evolutionary optimization (QIEO) represents design variables as a set of qubits and searches a continuous, multi-dimensional landscape through rotation of the qubit's amplitude pair. Every generation rotates those amplitudes toward a single elite, which corresponds to that generation's best. The update is cheap, almost parameter-free, and well-suited for massive parallel implementation, which has encouraged its adoption in engineering, design, and planning applications. However, there are critical issues with this formulation, principally, the treatment of design variables as independent probability components which make it incapable of exploiting local curvature, anisotropy, or variable coupling. Despite this, QIEO is believed to hold promise, and has been used extensively to solve real-world problems, with significant qualitative and computational advantage over its classical counterpart, Genetic Algorithm (GA). A collection of 256 (actually 508; 256 unshifted + 252 shifted, 4 could not be shifted) continuous function are selected from the prior works, in such a way that they represent eleven landscape characteristics, namely continuity, differentiability, separability, scalability, modality, convexity, conditioning, symmetry, maximum dimensionality, dimension dependency, and the coupling pattern of the design variables. These functions are then solved by three QIEO variants, two GA encodings and Hansen's Covariance Matrix Adaptation Evolution Strategy (CMA-ES). The results are evaluated in terms of computational cost, solution precision, and specialization across landscape characteristics. They identify the conditions under which QIEO provides competitive performance, clarify where its independent-variable representation becomes limiting, and establish whether particular QIEO variants offer advantages for specific landscape characteristics.
cs.NE / 3 / 2609.31066
Modeling quantum neural network gradient with reinforcement learning
Nhan Trong Luu, Duong Trung Luu, Nam Ngoc Pham, Thang Cong Truong
quant-ph · cs.ET · cs.NE
Abstract
Training quantum neural networks (QNNs) on near-term hardware remains hampered by two compounding difficulties: the exponential vanishing of gradient variance known as the barren plateau, and the $\mathcal{O}(L \cdot 2^n)$ time and memory cost of differentiating through an $n$-qubit, $L$-layer circuit. We propose RLQ-Grad, a reinforcement-learning-based optimizer in which a classical policy $π_φ$ (a spectrally-normalized PPO agent) learns to propose parameter updates directly, conditioned on the QNN's current parameters, loss, accuracy, and previous update. Because the surrogate gradient is emitted by a classical network rather than obtained by differentiating through the unitary $U(θ)$, its variance is not constrained by the barren plateau concentration bound, and its cost scales with the number of trainable parameters rather than the Hilbert-space dimension. We prove these properties formally and verify them on a hardware-efficient ansatz across four supervised benchmarks with up to $n=20$ qubits. RLQ-Grad preserves a near-flat gradient-variance curve where backpropagation, parameter-shift, and adjoint differentiation decay by 1 to 2 orders of magnitude. Accounting for the full training pipeline (PPO rollouts, actor-critic updates, and optimizer states), RLQ-Grad needs under 2 MB of memory and runs $2490\times$, $7876\times$, and $673\times$ faster per iteration than these three methods at $n=20$. It improves top-1 accuracy by up to $+10\%$ over gradient-based baselines on circuits of up to 12 qubits, and matches dedicated barren plateau mitigation methods on CIFAR-10 at 14 to 20 qubits, where evolutionary and gradient-free optimizers collapse to chance.
计算语言学 (cs.CL)
44
cs.CL / 1 / 2609.30414
A Unified Account of Concepts and Chunks
Karthik Singaravadivelan, Pat Langley
cs.CL · cs.AI
Abstract
Cognitive psychology has studied how people encode, use, and learn concepts that describe categories, and how they represent, recognize, and acquire chunks for familiar patterns of elements. The literatures on these two topics are nearly disjoint, which poses a challenge for unified theories of cognition. In this paper, we review Cobweb, a computational account of categorization and concept formation, and propose an extended theory that incorporates chunks and their acquisition. The theory makes no commitments about modality, applying to any experience that decomposes into elements and relations among them. We also present \trellis/, an implementation of this theory, and illustrate its application to learning context-free grammars, which we adopt as a testbed because they involve both concept-like and chunk-like elements. In addition, we report experimental results on three synthetic grammars that demonstrate the system's ability to represent syntactic knowledge, use it to parse and generate sentences, and learn compositional structures from sample parses. We conclude by discussing related work on concepts and chunks, along with directions for future research in the area.
cs.CL / 2 / 2609.30439
Inference-Time Target Speaker Unlearning in LLM-Based Automatic Speech Recognition
Bo Su, Yueru Yan, Thai Le
cs.CL · cs.SD · eess.AS
Abstract
We introduce target-speaker unlearning ASR (TSU-ASR) task in a fully end-to-end framework for multi-speaker ASR and diarization. Given a multi-speaker utterance and a set of opt-out speakers who do not wish to have their speech transcribed, the task requires an ASR system to transcribe all speakers except the opt-out ones, while still indicating when those speakers are active. As a first step towards tackling this task, we introduce a novel, light-weight Enrollment-Conditioned Gating (ECG) module attachable to a frozen dual-stream speech LLM that enables ASR for new opt-out speakers dynamically during inference, even those who were not seen during initial ECG training phase. Our experiments on both AMI (English) and AliMeeting (Mandarin) datasets show that speech transcription accuracy for corresponding opt-out words or characters falls from 72.3% to 48.2% and from 73.6% to 27.3%, respectively, while retained speakers' transcription error rates maintain more or less the same. Our approach provides a practical solution for modern video conferencing platforms, allowing speakers to dynamically opt-out from automated AI transcriptions without forcefully leaving the meeting sessions, enabling a privacy-preserving interface for potentially millions of online meetings daily.
cs.CL / 3 / 2609.30467
Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification
Heyuan Huang, Jirui Dai, Alexandra DeLucia, Sonal Joshi, Mahsa Yarmohammadi, Jie Gao, Bernal Jiménez Gutiérrez, Mark Dredze
cs.CL · cs.IR
Abstract
Retrieval-based factuality evaluation, where LLM-generated claims are verified against evidence from authoritative medical corpora, has become the dominant paradigm for scalable hallucination detection in high-stakes clinical settings. Despite the urgency of reliable and transparent medical fact verification, most systems measure performance with aggregate metrics like F1, which obscure where and why failures occur. Existing RAG diagnostics require gold answers or annotated gold evidence, neither of which exists in this regime. We introduce two comprehensive taxonomies, grounded in a case study on the open-ended MedExpert dataset and 3 closed-ended datasets, decomposing failures into retrieval-stage errors along five quality dimensions, and verifier-reasoning errors into six consecutive steps. We adapt an automatic pattern induction pipeline using LLM-as-Judge to label evidence quality and classify verifier reasoning errors at scale, and then stress-test our findings across 4 retrieval methods and 6 frontier verifier models. Our analysis reveals that scaling model size, adding reasoning effort, expanding to authoritative web sources, and applying medical fine-tuning do not resolve these failure modes, demonstrating that they represent fundamental limitations of the retrieve-then-verify paradigm in open-ended medical settings rather than artifacts of outdated systems. We release our code and data at https://anonymous.4open.science/r/Medical_RAG_eval-4AB5 for the full reproducibility of our results.
cs.CL / 4 / 2609.30471
CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production
Mukul Chhabra, Shail Patel, Luigi Medrano
cs.CL · cs.AI · cs.LG
Abstract
Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically applies the correct procedure to a different entity, so a literal judge penalizes different identifiers, dates, and statuses as errors or hallucinations. We name this failure mode reference-instance divergence (RID). We propose CARGO, a framework that (i) treats retrieved references as procedural exemplars and grounds factual judgments in the live instance's observed context, (ii) assigns each claim a three-way status (supported, contradicted, unverifiable) and penalizes only contradictions, and (iii) gates evaluation by retrieval confidence, casting production evaluation as selective prediction. We introduce CARGO-Bench, a perturbation-based diagnostic suite with ground truth by construction that separates leniency from discrimination. On CARGO-Bench (246 items, two judge models, 7,872 judgments), the standard reference-based judge penalizes 100% of correct entity-transplanted answers and is uninformative (discrimination index DI ~ 0); supplying the live facts without reframing changes nothing. CARGO eliminates these false penalties (0/50) while retaining near-complete contradiction recall (50/50 and 49/50), raising DI to 0.58 [0.48, 0.68]; a rubric-swap control attributes most of the effect to context-grounded dimension definitions. CARGO also exposes a limitation of its own design: the leniency that protects entity values suppresses detection of procedural corruptions (20% recall). A post-hoc fix does not close the gap, and an LLM-as-annotator study with written guidelines and adjudication shows the same blind spot. We release a preregistered protocol for extending the evaluation to expert agreement, risk-coverage, and cost on production traffic.
cs.CL / 5 / 2609.30492
Breaking Homogeneity: Diversifying Persona Sets for Creative LLM Outputs
Sang Bin Moon, Nicole Cho, Daniel Borrajo, Sumitra Ganesh, Abolfazl Hashemi
cs.CL · cs.AI · stat.ML
Abstract
Language models often produce homogeneous responses to open-ended tasks; such homogeneity can spawn groupthink-the convergence of ideas toward a singular and potentially suboptimal decision. We formulate persona diversification as a set-level conditioning problem and study two orthogonal design choices: selecting versus generating personas, and space-filling versus frontier-seeking diversity. We instantiate this design space with four methods spanning coverage and dispersion subset selections, uniform-coverage sampling, and evolutionary persona generation. Evaluations on the Alternative Uses Task (AUT), Infinity-Chat, and Divergent Association Task (DAT) show the benefits of the proposed methods across tasks and creativity objectives. On AUT, evolutionary persona generation increases response diversity by 78.8%, originality by 26.1%, flexibility by 49.5%, and holistic creativity by 13.9% over task-only prompting, while maintaining 98.5% validity; on Infinity-Chat, it nearly doubles persona-induced response separation relative to random personas. Moreover, evolutionary personas compose with creativity-optimized prompting, further increasing its response diversity by 18.6% and creativity by 6.3%. These results establish persona-set geometry as a task-agnostic mechanism for eliciting divergent LLM outputs, and support persona diversification as a reusable complement to prompt optimization.
cs.CL / 6 / 2609.30514
Inquesto Score: A reliability Protocol For Voice Agents
Massa Baali, Bhiksha Raj
cs.CL · cs.AI
Abstract
Voice agents are increasingly deployed in workflows where failed interactions can affect transactions, access, and other consequential outcomes, creating a need for reproducible and interpretable evaluation. We introduce Inquesto Score (IS), a protocol for measuring voice-agent reliability as the percentage of calls in a fixed, versioned evaluation population that achieve the caller's goal without a functional failure or worse. Rather than combining heterogeneous metrics, IS defines explicit failure events and severity levels and evaluates the deployed voice pipeline. Timing failures, including talk-over and delayed responses, are measured directly from audio, while semantic and state-dependent failures are evaluated using scenario predicates, tool traces, and a pinned open-model judge. Diagnostic views of behavior, acoustic robustness, identity handling, and speaker groups accompany the score without being combined into it. Inquesto Score v0.1 evaluates 30 scenarios, three acoustic conditions, four speaker groups, and 306 calls per agent across 13 configurations of a reference voice-agent system. Our evaluation shows that reliable measurement requires evidence beyond transcripts, explicit treatment of deployment conditions, and validation of the evaluators used to determine outcomes. We release the protocol, reference implementation, and evaluation records.
cs.CL / 7 / 2609.30535
Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence
Dries Rooryck, Alex Cai, Yonatan Belinkov, David Alvarez-Melis, Kianté Brantley
cs.CL
Abstract
Children in multilingual communities often code-switch, using multiple languages in a single utterance. Can we induce cross-lingual alignment in language models by training on code-switched text? We pretrain small decoder-only transformers on two 100M-word multilingual corpora: a base corpus formed by mixing the English, Dutch, and Chinese BabyBabelLM datasets, and a corpus generated from it by inserting word- and sentence-level code-switching using an LLM. We find that training on code-switched data aligns the representations of parallel text, particularly across different scripts, and that this alignment persists through training on monolingual documents. Under a learning curriculum that progresses from word-level code-switching, to sentence-level code-switching, to monolingual documents, models trained on code-switched data outperform baselines trained without it on the BabyLM evaluation suite. Our work characterizes code-switching curriculum learning as an effective data augmentation method for multilingual pretraining. We release our code, data, and models at https://github.com/drooryck/multilingual-macaroni.
cs.CL / 8 / 2609.30547
REALMS: An AI-Assistant Conversational System for Real-Time Exact Audience Sizing over High-Dimensional Nested Profiles
Haixu Ma, Aditya Bansal, Shubham Lohiya, Sumit Ranjan
cs.CL · cs.IR
Abstract
Audience sizing is a critical component of digital marketing. It enables precise resource allocation, campaign planning, and performance optimization. Traditional approaches using skeleton audiences, sampling, or predictive modeling suffer from significant delays, estimation errors, and poor scalability over high-dimensional profile data. We present REALMS (Real-time Exact Audience sizing via LLM-based Multi-attribute Search), a conversational system for exact audience sizing deployed in production on an enterprise customer data platform. REALMS enables marketers to query massive profile stores with millions of profiles and thousands of attributes using natural language and receive precise counts in seconds. The system introduces three key components: (1) a categorical attribute retrieval mechanism using embedding-based vector search to dynamically identify relevant schema attributes without manual configuration; (2) an LLM-powered NL2SQL pipeline with template-based in-context learning for accurate query generation over complex nested schemas; and (3) schema standardization enabling industry-agnostic deployment across diverse enterprise environments. Evaluation on real enterprise data demonstrates strong recall for attribute retrieval, high SQL execution accuracy, and low latency, which enables real-time interactive audience insights where prior methods required hours.
cs.CL / 9 / 2609.30558
Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms
Jiaqi Ding, Guorong Wu
cs.CL · cs.AI · cs.LG
Abstract
Agent memory systems are increasingly used to maintain long-term user preferences, task states and evolving facts, but current evaluations often collapse memory behavior into final-answer accuracy. We introduce MemProbe, a cognitive-science-inspired framework for diagnosing stability-plasticity tradeoffs in agent memory. The framework is motivated by a core insight from cognitive memory research: memory is reconstructive and shaped by interference, source reliability, reinforcement, and reactivation. MemProbe turns this insight into four reusable experimental paradigms (interference, misinformation, consolidation strength, and reconsolidation window) that manipulate when a memory should be updated, preserved, or treated as uncertain. It further decomposes correctness into behavioral profiles that reveal how systems update, preserve, attribute, and temporally organize information. We instantiate these paradigms in a 56-episode diagnostic suite and evaluate six incremental memory systems under a unified protocol. Results show that systems with similar aggregate scores exhibit distinct behavioral profiles. MemProbe provides such a diagnostic lens, turning aggregate performance into interpretable profiles of memory maintenance over time. Code is available at https://github.com/jq-ding/MemProbe.
cs.CL / 10 / 2609.30604
The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge
Alexander Gill, Md Farhan Ishmam, Xuyen Nguyen, Neha Bhat, Parker Henry DeYoung, Fateme Hashemi Chaleshtori, Nathan Stringham, Kenneth Marino, Ana Marasović
cs.CL · cs.AI
Abstract
Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations, spreadsheets), and navigates program interfaces to produce a coherent final product. Such workflows demand reasoning and synthesis, decomposition of complex tasks, as well as visual and spatial understanding. To study agents on workflows like these, we introduce KNOWS, a benchmark of open-ended, complex, browser-based tasks that jointly evaluate these capabilities, with each task culminating in a produced artifact. To write tasks, we develop a task design rubric and a protocol for ensuring that tasks meet the requirements. Each task is paired with an evaluator, a program that combines deterministic checks with LLM judgments to balance the richness, reliability, and automation tradeoff inherent to agent evaluation. We evaluate and analyze frontier computer-use agents and browser-based harnesses. They achieve moderate scores on partial-success metrics, but the best performer fully succeeds in fewer than 3% of our complex, long-horizon tasks. Failures on visual steps render the resulting artifacts unusable, even when agents complete more than 50% of other evaluation steps. Our results expose limitations of current agents acting as end-to-end assistants, and call for progress on tool use, visual understanding, and long-horizon reasoning.
cs.CL / 11 / 2609.30652
Recursive Self-Improvement via On-Policy Distillation for Reasoning
Shangjian Yin, Zehao Zhao, Kavosh Asadi, Rui Liu, Yuchen Lu, Shike Mei, Hang Cui, Luke Simon, Zhouxing Shi, Hamed Firooz
cs.CL
Abstract
On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token predictions. This provides dense, token-level supervision to the student. On-policy self-distillation (OPSD) eliminates the need for the external teacher. Specifically, a second frozen copy of the student model, now given the ground truth in its context, serves as the teacher. The student model only receives the problem and learns to mimic the privileged teacher model, while the teacher remains frozen throughout training. Previous work showed that freezing the teacher is useful for training stability, but we argue that this can prevent the teacher from incorporating the improvements learned by the student during training. Our primary contribution is to address this limitation with a recursive framework built around two complementary components. First, we let the privileged teacher co-evolve with the student so that revision learned in one round can guide the next, a process we refer to as Dynamic Co-Evolution (DCE). Second, because stronger revision can also make responses too verbose and self-critical, we additionally train on shorter, verified rewrites of the model's own on-policy responses. We call this complementary objective Self-Refined Concise Learning (SRCL). Overall, our comprehensive evaluations show that DCE+SRCL outperforms OPSD across multiple model scales and four competition-level mathematics benchmarks. Specifically, on Qwen3-8B, DCE+SRCL reaches 65.97% Average@12, outperforming OPSD by 35.62 percentage points while reducing mean output length by 7.80% relative to DCE alone.
cs.CL / 12 / 2609.30670
TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding
Yibo Ma, Qianqian Zhang, Peng Liu, Tiancheng Zhao
cs.CL · cs.CV
Abstract
Streaming video understanding requires models to interpret evidence as it arrives, yet current evaluations often report task scores without specifying when evidence becomes valid, how visual history is maintained, or how responses are triggered. As a result, similar scores may correspond to different workloads, failure modes, and operational behavior. We introduce TRACE (Temporal Audit and Condition-aware Evaluation), a condition-aware benchmark and evaluation framework that makes these factors explicit. TRACE combines temporally audited visual tasks with evidence timing and instruction-dependent trigger annotations, a unified causal Core--Adapter protocol that controls information availability while recording actual history processing and response events, and multidimensional reporting of answer quality, timeliness, response-selection behavior, workload, completion, and reliability. On 1,240 records from 517 videos, we evaluate eight publicly available models or systems in eight configurations. We find that nearly identical QA accuracy can mask substantial differences in completion, answer validity, and generation workload, while proactive performance separates into response quality, response delay, false alarms (responses emitted while no target window is currently valid and a later one remains), and missed target windows. These results show that streaming-video performance should be interpreted as execution-conditioned system behavior rather than a single score. Our benchmark and code can be accessed at \href{https://github.com/om-ai-lab/trace-bench}{https://github.com/om-ai-lab/trace-bench}.
cs.CL / 13 / 2609.30716
Words Speak Louder Than Order: A Behavioral Evaluation of Gemma 4
Amanda Fitch
cs.CL · cs.AI
Abstract
When a language model receives two conflicting documents as input, how does it decide which one to prioritize? Does it rely on how the sources are framed or the presentation order of the documents? We evaluated this behavior on Google's pre-trained Gemma 4-e4b model across a targeted behavioral suite (n = 13 items, 784 forward passes in short, single-turn contexts) using a completely counterbalanced experimental design. This setup allowed us to mathematically isolate the specific effects of source framing and reading position, while ensuring the model's natural vocabulary biases were canceled out. Across ten test conditions, we discovered the following: 1. Source framing heavily overpowers reading position. When directly competing, the semantic framing of a source (such as presenting it as an official guideline or a fresh update) had a significantly stronger impact on the model's final answer than the presentation order of the document. 2. The model favors the first document it reads, but this bias is highly variable. While the model consistently demonstrated a primacy effect (preferring the first document presented), the actual strength of this bias fluctuated by at least a factor of 5 based solely on the surface wording. 3. Overall structural repetition, not short copy-cues, drives positional bias. The model's preference for the first document is not a mechanical reaction to short, repetitive trigger phrases, such as "is [Answer]". However, the primacy effect does increase significantly when the two competing documents are structurally identical, using word-for-word verbatim templates. Introducing variation in the overall wording between the two sources reduces this positional bias.
cs.CL / 14 / 2609.30738
Beyond Mean Attention: Diversity-Aware, Layer-Wise Scoring for KV Cache Eviction
Tianfang Xie, Wei Zhu
cs.CL · cs.LG
Abstract
KV cache eviction methods such as SnapKV and PyramidKV rank tokens solely by mean attention over a small observation window. We study a unified score, $μ_i+λ_1σ_i+λ_2\mathrm{corr}(i,S)$, adding attention dispersion across window queries and redundancy relative to selected tokens. For $λ_2<0$, the score penalizes similarity to selected tokens as in maximal marginal relevance (MMR), without extra forward passes. To test whether this relevance-diversity balance should vary with depth, we compare fixed global coefficients with three-segment and quadratic profiles. Only these depth profiles are searched on a development split under a $\sinh$ reparameterization. On all 16 English LongBench datasets with Mistral-7B at a budget of 64 entries per layer, a single global diversification constant improves 13 of 16 datasets (macro +1.1); the gain holds at budget 32 and narrows at 128. Per-dataset search finds no detectable layer structure on most datasets; on passage retrieval it finds a large one: a mid-layer sign flip that rewards similarity and is worth +9.6 over the baseline at budget 64 and, without re-tuning, +13.2 over the global constant at budget 128. Ablations attribute the gain to the redundancy term; replaying every accepted search state on the held-out test set separates genuine structure from tuning noise.
cs.CL / 15 / 2609.30739
SEA-CLIP-Tiny: Efficient Multilingual Text-Vision Embedding for Southeast Asian Languages
Puja Ahmad Habibi, Faiz Assabil Firdaus, Ashvanth S, Ekapol Chuangsuwanich, Pume Tuchinda, Peerat Limkonchotiwat
cs.CL
Abstract
Multilingual text-vision embedding models are essential for cross-lingual image-text retrieval, but Southeast Asian languages remain poorly supported due to the region's linguistic diversity and limited data and computing resources. In this paper, we introduce SEA-CLIP-Tiny, a compact multilingual text-vision embedding model for Southeast Asia with fewer than 50M parameters. Our model adapts a CLIP-KD-style framework to Southeast Asian multilingual settings through regional data curation and multilingual teacher guidance. Experiments across seven Southeast Asian languages show that SEA-CLIP-Tiny achieves the strongest average retrieval performance among the evaluated student models, reaching 12.9%, 31.5%, and 42.2% at R@1, R@5, and R@10, respectively. Compared with MobileCLIP2, it improves average R@10 by 12.1 points while using 38.4% fewer parameters and lower measured CPU latency. These results highlight the importance of region-aware training for efficient multilingual text-vision models in Southeast Asia.
cs.CL / 16 / 2609.30802
Understanding the Role of Prompt Template in Knowledge Distillation for Safety Alignment
Anjila Budathoki, Manish Dhakal, Benjamin M. Ampel, Yi Ding
cs.CL
Abstract
Prior research has demonstrated that the choice of prompt template during Supervised Fine-Tuning (SFT) significantly impacts the robustness of safety alignment afterwards. However, the influence of template selection during Knowledge Distillation (KD) from teacher to student remains largely unexplored. Thus, we fill this gap by analyzing how different template configurations influence the pre-existing safety alignment of the student. We observe a significant degradation of safety alignment present in the aligned base instruct-tuned model. Specifically, we find that utilizing chat templates renders the model more compliant with harmful queries compared to a non-chat template. These findings are consistent across three models: LLaMA, Gemma and Qwen model families and are evaluated across multiple safety benchmarks. We further show that using a non-chat template during distillation better preserves the base student's internal representations, while chat template distillation induces a larger representational shift. Code: https://github.com/anjilab/role-of-prompt-template-in-kd
cs.CL / 17 / 2609.30846
I-Parakeet: Integer-Only Conformer ASR on Mobile NPU
Taichi Nishimura
cs.CL · cs.SD · eess.AS
Abstract
In this paper, we propose I-Parakeet, an integer-only implementation of NVIDIA's Parakeet-CTC (0.6B parameters) that runs on a smartphone NPU without any floating-point operator or CPU fallback. Modern Conformer ASR models are hard to deploy on edge devices because of their size, and quantized models still fall back to floating point for numerically sensitive operations. This prevents them from fully exploiting integer accelerators such as mobile NPUs. To achieve this, our contributions are threefold. First, we derive an integer formulation of the relative-positional self-attention at the core of the Conformer. We fuse its two score branches with different quantization scales and the relative shift into integer-only operations. Second, we introduce a minimax-optimized Swish approximation that minimizes the maximum error of the Swish output. Third, a layer-wise range analysis of activations yields two targeted remedies: an INT16 grid for the BatchNorm output and percentile calibration for the heavy-tailed pre-encoder activations. I-Parakeet achieves 4.97% WER on LibriSpeech test-other, running on a Qualcomm NPU at a real-time factor of 0.048, 7.5x faster than a CPU baseline.
cs.CL / 18 / 2609.30849
Enhancing Assessment of Self-Consistency in LLM Explanations using Perturbation Strength
Phuong Q. Le, Kemal Kurniawan, Jey Han Lau
cs.CL
Abstract
Prior work has examined the self-consistency of LLM-generated explanations using surface-level perturbation methods. However, the strength of these perturbations is not explicitly measured and controlled. In this work, we propose an LLM-as-a-judge approach to measure perturbation strength in a unified manner across input and CoT perturbations. We then evaluate the self-consistency in explanations generated from various LLMs under controlled strength conditions, ensuring a fair comparison across perturbation types. Experiments show that our proposed LLM-based perturbation strength measure outperforms other embedding- and probability-based approaches and that input perturbations generally affect LLMs more strongly than CoT perturbations. Our work suggests that judgments about a model's self-consistency is fair only within the same perturbation type.
cs.CL / 19 / 2609.30864
Persistent Negatives for Adversarial Black-Box On-Policy Distillation
Haixu Ma, Saad Lahrichi, Weiwei Li, Kevin Han, Weiqiang Wu, Peggy Yang, Dongzhuo Li, Ruiyi Li, Serena Li, Gedi Zhou, Mingze Gao, Abhishek Kumar, Xiangjun Fan, Lizhu Zhang
cs.CL · cs.AI
Abstract
Black-box On-Policy Distillation (OPD) seeks to improve a student from its own generations when the teacher provides sampled responses but not token probabilities. Adversarial distillation offers one route: it learns a discriminator over prompt-matched teacher and student responses and uses its score as the policy reward. However, sampling discriminator negatives from the latest student at each step couples the learned reward to a negative distribution that changes after every policy update. We address this moving-target problem with persistent-negative adversarial distillation, a live-pool method that replaces a fraction of each discriminator batch with historical, prompt-matched teacher--student comparisons. Under matched discriminator compute, historical comparisons train the discriminator, while GRPO remains on-policy with fresh student responses. Our analysis identifies the Bayes-optimal reward as a teacher-to-negative log-density ratio and, under explicit assumptions, shows how persistent negatives anchor the discriminator and reduce reward-estimation MSE relative to fresh-negative training. Across two student families, three judges, and four judged-chat benchmarks, persistent-negative adversarial distillation consistently improves performance over current methods at matched discriminator compute. It also yields smoother fresh-policy discriminator trajectories, with fewer below-chance dips. These findings identify the discriminator's negative distribution as an important design axis in black-box on-policy distillation.
cs.CL / 20 / 2609.30867
Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations
Yonghong Zhang, Yong Xie, Isabel M. Parra, Ricardo Correia
cs.CL
Abstract
Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported evidence against an eleven-dimension assumption-implication-evidence rubric and abstains when relevant evidence cannot be retrieved. We evaluate ARGUS using injected flaws, economics papers, and a small pilot with reconciled labels. On the 11-flaw benchmark, ARGUS detects 73% of planted flaws, compared with 18% for a keyword-based pipeline. Across 26 economics papers, ARGUS abstains on about 40% of paper-dimension assessments for lack of retrievable evidence. In a five-paper pilot with labels reconciled by two annotators, it assigns a higher risk level than the labels on 25 of the 33 assessments it completes. A rule fixed before the labels arrived removes most of this in-sample; weighted agreement stays low. ARGUS provides evidence-linked risk reports that localize potential weaknesses for expert review, without adjudicating causal claims. Code and data: https://github.com/yonghongzhang-io/ARGUS
cs.CL / 21 / 2609.30897
From annotation to reasoning: Culture in language models
Daniel Hershcovich, Alexander Conroy, Jens Bjerring-Hansen
cs.CL
Abstract
How should we evaluate language models when more than one interpretation can be right? Cultural benchmarks often test factual knowledge, agreement with survey responses, or recognition of a predefined meaning. These tasks leave open whether a model can explain how a cultural reference works in a particular text, support a reading with evidence, or revise it after criticism. This is a question of interpretive depth, complementary to the breadth of cultural coverage. We argue that literary interpretation offers a useful setting for studying these capabilities. We focus on cultural referencing and reuse: how texts invoke, repeat, and transform earlier expressions across historical and linguistic contexts. Our central claim is that literary scholars can disagree about an interpretation while recognizing the quality of its support. We propose linking evidence-centered benchmarks, evaluation that preserves scholarly disagreement, and model-development experiments on literary data, contextual resources, and scholarly feedback. Danish literature provides a concrete starting point, with implications for other languages and domains. The aim is to develop alternative evaluation strategies that go beyond conventional benchmark metrics and guide model development toward cultural robustness in AI systems.
cs.CL / 22 / 2609.30924
Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring
Hikaru Asano, Yotaro Kubo, So Kuroki
cs.CL · cs.SD · eess.AS
Abstract
Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pronunciation (G2P) and speech-to-pronunciation (S2P) methods each capture only partial information, using only text or only speech, while speech-and-text-to-pronunciation (ST2P) methods use both but require costly pronunciation-annotated data. To address this problem, we propose a training-free ST2P pipeline that integrates both lexical and acoustic information at inference time. Lexical resources and G2P tools generate text-constrained candidates, and a left-to-right greedy search selects the best one using whole-sequence negative log-likelihoods from frozen pretrained S2P models. On three Japanese corpora, our method reduces Character Error Rate (CER) from 0.60--1.40\% (text-only baseline) to 0.04--0.17\% with reference transcripts, and 0.64--1.58\% with ASR transcripts. It outperforms all baselines, including a trained ST2P model and commercial multimodal LLMs. Our greedy search method is 3--3.5$\times$ faster than beam search at similar CER, and the cascade is 2$\times$ faster than direct decoding ensuring the efficiency and accuracy. In Spanish, French, and preliminary English, it also surpasses four open multimodal LLMs and the best traditional methods.
cs.CL / 23 / 2609.30974
Coupled Usage-Sense Processes: Temporal and Attributable Lexical Semantic Change
Haruka Ezoe, Ryohei Hisano
cs.CL · stat.ML
Abstract
Lexical semantic change is usually summarized by a scalar distance between independently sampled period distributions. This measures how much a word changed, but does not reveal when it changed, which mechanisms and component movements carried the change, or which usages support the attribution. We introduce Coupled Usage--Sense Processes (CUSP), which derives these answers from a single marginal preserving temporal process. A hierarchical coupling relates contextual distributions through latent usage components, while Markov composition makes adjacent and longer span correspondences compatible. Displacement operators quantify change magnitude and timing, split variation exactly between movement of component centers and reorganization within components, and attribute it to transported component pairs. Word-local modes resolve distinct directions of change and their activity over time, while representative passages from attributed components ground the analysis in text. Under a Gaussian mixture specialization, we prove parametric recovery of the operators and squared distances. Synthetic experiments support the predicted rate. CUSP remains competitive on English and German DWUG and recovers controlled Janus profiles while maintaining compositionally coherent transport. A large corpus of US court opinions demonstrates transition, mode, and passage attribution in unlabeled natural text. CUSP thus makes magnitude, timing, mechanism, movement, modes, and textual evidence compatible views of one lexical history.
cs.CL / 24 / 2609.30984
THA: Weighted Finite-State Text Normalization and Inverse Text Normalization for Khmer
Seanghay Yath
cs.CL
Abstract
Text-to-speech needs written text in spoken form, and speech recognition output needs the reverse. For Khmer, neither direction has a maintained open-source tool, and the script makes both harder: words are not separated by spaces, and number words occur inside ordinary words. We present Tha, a Khmer text normalization and inverse text normalization toolkit built from weighted finite-state transducers. It segments and classifies a whole line in one shortest-path search, and a second transducer rejects token boundaries inside a Khmer syllable. On Google's Khmer test suite, Tha agrees with the reference on all 274 cardinals up to one spelling variant, and on 2,906 real TTS prompts, 153 of the 158 sentences it rewrites are correct. Tha is open source under the Apache 2.0 license.
cs.CL / 25 / 2609.31062
CG-Probes: Recovering Guardrail Directions from Patient Query Embeddings
Marko Řeháček, Vítězslav Dušek, Martin Rusinko, Vít Nováček
cs.CL · cs.IR
Abstract
Patient-facing AI assistants promise valuable support to patients, but incoming queries can pose medical risks. To create guardrails, we work with oncologists to define three ordinal risk axes: Medical Urgency, Psychological Urgency, and Topic Sensitivity. We propose Clinical Guardrail Probes (CG-Probes) to measure the risks from query embeddings. We probe for each axis in the normalized embedding space of frozen embedders via the difference-in-means method, treating each axis as a potential linear direction. To train the probes, we cluster 79,658 Czech oncology search queries with BERTopic and use these clusters to generate pairs of queries with contrastive risk levels via few-shot prompting. We evaluate the approach on 200 queries (90 real, 110 synthetic), each graded by two oncologists, against two open-weight LLMs and a frontier LLM. We find that urgency-based axes are recoverable as linear directions, and the probes are competitive with open-weight LLMs (no significant differences in quadratic-weighted kappa) at a fraction of the latency. Each axis yields a scalar score that clinicians can inspect and use to set escalation thresholds. The pipeline requires only search logs, axis definitions, and black-box access to the embedding model, suggesting transferability across healthcare domains. Robust validation on new queries and axes remains future work.
cs.CL / 26 / 2609.31130
Do we need to answer that question? Salience and Answerability of Potential Questions in Naturalistic Dialogue
Amandine Decker, Maxime Amblard, Ellen Breitholtz
cs.CL
Abstract
We empirically investigate Question Under Discussion based modelling in naturalistic dialogue by studying whether the salience of generated potential questions predicts their subsequent resolution. Building on Wu et al. (2024), we construct a dataset of 7,124 questions automatically generated from utterances and preceding context from the British National Corpus, and annotated for salience and answerability. We find a robust but low positive correlation between salience and answerability in dialogue, indicating that more salient questions are more likely to be addressed. However, this effect is markedly weaker than in monologic text, suggesting that conversational structure is less predictable. We further observe that structured interactions exhibit stronger alignment between annotators than less organised dialogues.
cs.CL / 27 / 2609.31181
Where a Model Sends Its Own Repeated Token
Nicolás Vera Zúñiga
cs.CL
Abstract
Black-box model identification works by scoring a model's response to natural-language prompts. One line of work feeds models a degenerate input -- their own token, repeated -- to find a failure mode rather than an identity. We take that input and ask where the model goes when it does not. For each token t, read argmax p(. | t, t) in one forward pass; the result is a map on the whole vocabulary, with two halves. The first -- which tokens are fixed points -- is partially anticipated, and we report it as a failed estimand: the natural distance on it is 83% cardinality, separates a corpus manipulation by two bits in 3471 against a precision floor of zero, and attributes families at 0.5833. The second half, where the map sends tokens that are not fixed points, is unrecorded; the one paper holding those tokens logged them as a zero. Pairing on the source token removes the cardinality confound by construction (r from 0.9128 to -0.0932) and attributes families at 0.8333 -- twelve models scored against a pool of nineteen -- with chance 0.1389, across seven tokenizer groups and several corpora. Two nulls clear it: frequency-matched destinations agree at 0.1429, independent marginals at 0.0798. Family predicts agreement better than tokenizer (0.2031 against 0.1205), and recurrent architectures cluster at balanced accuracy 1.0 against a 0.7895 majority rate, or 0.90 once each model's dominant destination is excluded -- the figure we stand behind. We measure the robustness envelope: 8-bit weight rounding moves the map less than deduplicating the training corpus does (0.9004 against 0.6353, on one support), 4-bit destroys it (0.0098; 0.1812 at deployment granularity, so not a coarseness artefact), and the precision floor varies by model from 0.201 to 0.9778. All estimands and kill conditions were registered before the data, and the failed one is reported as fully as the surviving one.
cs.CL / 28 / 2609.31255
PIA: A Personal Intelligence Agent Turning Health Conversations into Records and Records into Understanding
Jeonghun Yoon, Dongchan Kim, Hongyeon Yu, Young-Bum Kim, Jaegul Choo
cs.CL
Abstract
General-purpose agent memory summarizes conversations: it extracts salient snippets, embeds them, and retrieves the top-k into the prompt. A health agent cannot run on summaries: a dose becomes a sentence, "since last week" is resolved at the model's discretion, and a three-month glucose trend cannot be answered by text similarity. We present PIA, a personal intelligence agent deployed alongside a consumer health agent. PIA receives the agent's natural-language requests, decides for itself whether and how to write or read, and turns conversations into typed clinical records and records into a synthesized understanding of the user. Its memory harness consists of four controls -- extraction, memory, retrieval, and understanding -- each a domain-agnostic mechanism with a pluggable health module: schema, medical alias dictionary, knowledge graph, and temporal rules. We show how the same query receives a different answer as the memory injected into the response context deepens from one-dimensional recall, to a two-dimensional health snapshot, to a three-dimensional trajectory with causality, and report lessons from operation: self-reported health data are missing not at random, question phrasing governs the quality of synthesized understanding, and nearly a third of candidate causal links are structural noise that rules alone remove.
cs.CL / 29 / 2609.31261
MoSAR: Mixture of Semantic Attention Regimes for Learning Adaptive and Approximable Attention Geometries
Michele Paolicelli, Alessandro Petruzzelli, Alessandro Franceso Maria Martina, Cataldo Musto, Giovanni Semeraro
cs.CL · cs.AI
Abstract
The quadratic complexity of dense self-attention remains a central bottleneck for long-context language modeling. Many efficient alternatives address this cost by deciding in advance where attention should be sparse or local. We argue that attention approximation should instead be approached as a geometric problem, with the relevant interaction geometry learned from data: natural-language dependencies are input-dependent and difficult to prescribe in advance, so the model should learn where positional relevance can decay and where broader interactions must be preserved. We introduce Mixture of Semantic Attention Regimes (MoSAR), which learns such an adaptive, controlled-decay geometry over query--key interactions. Input-conditioned query and key routers, applied after positional encoding, select mixtures over short, medium, and global regimes, inducing a continuous distance-dependent attention field rather than a fixed sparsity pattern. This geometry is learned during training and can subsequently be discretized through top-1 routing. In controlled pre-training experiments with matched 500M-parameter models, MoSAR learns a substantially lower-reach attention geometry without degrading language-modeling quality, improving perplexity over dense RoPE at the training context length. Under length extrapolation, MoSAR achieves the best perplexity among all evaluated variants, including strong baselines such as ALiBi. Moreover, the learned geometry remains stable under deterministic top-1 discretization, suggesting that it is not only adaptive, but also amenable to low-cost approximation at inference time.
cs.CL / 30 / 2609.31264
Identifying Scientists on X
Philipp Meier, Katarina Boland, Laura Kallmeyer, Stefan Dietze
cs.CL
Abstract
With the growing importance of science-related discourse on the Web and the erosion of the classical knowledge order, it is important to identify different user groups, such as scientists, automatically. This work proposes an approach for identifying scientists and non- scientists on X/Twitter based on their user biographies and tweets. We show that we are able to classify accounts as scientists and non- scientists on two different datasets, reaching an F1 score of up to 0.88 using Random Forests with linguistic features and up to 0.96 using a contrastively fine-tuned DeBERTa model in an ensemble setup. Furthermore, we provide two datasets with X users labeled as scientists or non scientists and their respective tweets and user biographies.
cs.CL / 31 / 2609.31342
Stale-Document Poisoning: When Outdated Retrieval Overrides Correct Model Answers
Md Shamim Ahmed, Lukas Galke Poech, Richard Röttger
cs.CL
Abstract
Retrieval-augmented generation (RAG) is often used to address outdated knowledge by providing external evidence. But retrieval helps only when that evidence is still valid. We identify a temporal alignment failure, stale-document poisoning, in which outdated evidence makes a model wrong despite answering correctly without retrieval. We construct a benchmark of 317 verified knowledge reversals across medicine, law, software, and platform policy, grounded in dated official sources. Across 12 models, recent medical reversals are harder than long-established ones. More importantly, outdated retrieval flips 30% of Llama and 37% of Qwen answers even without instructions to trust the document; explicit follow instructions raise these rates to 66% and 75%. Across four open models and four domains, poisoning ranges from 17-91%, while matched up-to-date evidence is followed in 97-100% of trials. To isolate temporal applicability, we keep the historical evidence unchanged across 50 reversals and vary only the evaluation date. A clear pattern emerges: dates alone produce only modest adaptation, but when models are explicitly told when the old evidence stops applying, the larger models switch to the appropriate answer almost perfectly. Causal interventions confirm that this validity information directly shapes the final decision. The same internal components also support broader comparison tasks, suggesting that temporal applicability can recruit a general reasoning mechanism used for other comparisons. Finally, a fixed recency-aware hybrid re-ranker reduces poisoning by 4.6-10.0 points when dates are accurate, with gains that depend on reliable temporal metadata. Reliable RAG therefore requires selective trust: models must determine not only what retrieved evidence says, but whether it still applies.
cs.CL / 32 / 2609.31403
Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers
Mert İncidelen, Yamen Kashkash, Asya Berker, Murat Aydoğan
cs.CL · cs.AI · cs.CV
Abstract
Vision-language models (VLMs), despite their success in optical character recognition (OCR) tasks, are vulnerable to typographic attacks and have a fragile structure for images with multiple text layers. In this study, the DecoyBench dataset was created using the Decoy Font method. The dataset consists of 300 images, each containing text with sharp contour lines superimposed on another text with soft shading. Six recent closed-source models from three different model families were evaluated using this dataset under two different prompting conditions (naive and guided) and at two different resolutions ($512\times512$ and $64\times64$). A validation study showed that human participants could read both text layers with high accuracy. In contrast, the models, with most variants and both prompting methods, read the contour text with near-human accuracy at high resolution, but almost never fully extracted the shading text. At low resolution, the contour text could not be read by either the models or humans, while the shading text could be extracted with high accuracy. The findings indicate that the evaluated VLMs exhibit a consistent behavioral limitation when processing typographic structures containing multiple spatial frequency layers.
cs.CL / 33 / 2609.31511
Muslim: A Deployed Arabic Voice AI Platform for Grounded Islamic Knowledge
Yahya Mohamed Elnawasany
cs.CL
Abstract
We present Muslim, a production Arabic voice AI platform serving grounded, sourced Islamic knowledge to real users. Beyond a real-time voice pipeline (NeMo Arabic ASR, an OpenAI-compatible LLM endpoint, self-hosted TTS) and a deterministic multi-source retrieval layer routed across six Model Context Protocol servers, we report three things a research prototype typically lacks. First, a released family of fine-tuned Arabic Islamic model artifacts: an efficient tool-routing LLM (Muslim-6B-PRO, 5.94B parameters) and a Modern Standard Arabic TTS model (Fasih-TTS-V1) that ranks 5th of 17 overall and 2nd of 11 open-weight systems on the community-voted Arabic TTS Arena for MSA. Second, an account and metering layer - a free per-account turn allowance, capacity-aware refusal, and email verification deferred to the point it actually matters - that turns an open demo into an operable, abuse-resistant product. Third, a three-layer observability stack (liveness, error reporting, product analytics) built specifically around the system's characteristic failure mode: a GPU-bound agent host going silent while the web tier keeps serving normally. We report real, measured latency and accuracy figures (98.4% recitation-validation accuracy on 124 cases; end-to-end voice latency of 0.9-1.7s) and discuss the concrete engineering trade-offs and limitations of running an Islamic-knowledge voice product in production.
cs.CL / 34 / 2609.31553
MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos
Itzel Tlelo-Coyotecatl, Hugo Jair Escalante
cs.CL · cs.CV
Abstract
Ensuring online safety through content monitoring had raised Hate Speech Detection as a crucial task to be addressed. By essence the task demands the capture of contextual cues, which are essential for a precise understanding of the content's intent. Although automated detection approaches for the task have advanced significantly, the scarcity of non-English resources persists, limiting the ability of models to adapt to the subtle, context-dependent, and culturally related nature of multimodal content. In this paper, we introduce MexHat, a video dataset designed to capture the linguistic and cultural cues for the hate-speech detection task in a Mexican Spanish context. Our dataset comprises around 1k video clips annotated across two tasks: a three-way class evaluation (no negative content, offensive content and hate-speech content), and a fine-grained class evaluation including three hate-speech sub-categories. The dataset statistics and the baseline results highlight the inherent challenges associated with the task. Disclaimer: This paper contains sensitive content that may be disturbing to some readers.
cs.CL / 35 / 2609.31571
Strategically Diverse Sampling for Self-Training
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
cs.CL
Abstract
Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically constructed by sampling IID responses and filtering primarily for correctness, thereby overrepresenting strategies a model already favours. We investigate strategic diversity, or substantive variation among approaches to a problem, as an alternative principle for constructing self-training data. We generate strategically diverse data with two sampling methods: GROOT, a new method which constructs a hierarchical tree of approaches and samples distinct paths, and Verbalized Sampling (VS), adapted to produce an unstructured set of approaches. Across competitive programming and Next-Chapter Prediction domains, models trained on strategically sampled data outperform IID-trained counterparts on difficult tasks and provide strong initializations for RL and test-time scaling. Most strikingly, self-training on strategically diverse but incorrect traces from Qwen3-4B outperforms IID distillation from a 235B teacher. These results challenge prevailing assumptions about what makes useful self-training data and show that diversity of approaches can matter more than correctness or teacher scale.
cs.CL / 36 / 2609.30914
Cross-Backend QIEO: Universal Runtime Portability across OpenMP5, CUDA, HIP, and Multi-Language Interfaces
Aman Mittal, Ferdin Sagai Don Bosco, Kasturi Venkata Srikanth, Abhishek Singh, Aditya Singh, Abhishek Chopra
cs.DC · cs.CL · math.OC
Abstract
Quantum-inspired algorithms emulate quantum mechanical principles, such as, superposition, interference, and probabilistic amplitude evolution, on classical hardware by representing candidate solutions as qubit vectors and evolving them through rotation-gate operators. This approach offers higher optimization performance without physical qubits, and has been shown to achieve order-of-magnitude speedups (10--80$\times$) over traditional solvers on combinatorial, high-dimensional NP-hard problems. A critical barrier to adoption, however, is the lack of a unified execution framework that delivers both algorithmic performance and hardware portability. We present \textbf{Cross-Backend Quantum Inspired Evolutionary Optimizer (QIEO)}, the runtime core of BQP's BQPhy solver, which addresses this gap through a \emph{single-source-of-truth} architecture. One C++ implementation of the QIEO algorithm is compiled once per hardware target and exposed to multiple high-level languages via thin binding layers. The framework dispatches to CPU (sequential), OpenMP~5 (multi-core), CUDA (NVIDIA), and HIP (AMD) backends at runtime, adapting kernels to each device's memory hierarchy and warp/wavefront execution model. The framework's real-world utility is validated through binding demonstrations that share the identical C++ runtime. BQPhy's Python library is demonstrated on a neural network hyperparameter optimisation achieving 88.60\% test accuracy on MNIST. BQPhy's MATLAB's Toolkit is tested on wind farm layout optimisation attaining $365\,399 \pm 4\,552$~MWh/yr, which is statistically indistinguishable from particle swarm optimisation and $+7.6\%$ above genetic algorithms on a 32-variable constrained engineering problem. The Julia package tackles the Lotka--Volterra parameter estimation where BQPhy replaces native Julia solvers on the same residual, cutting mean SSE by $2.1\times$.
cs.CL / 37 / 2609.30611
Epstein Files Engine: Agentic Search for Investigative Journalism
Duy K. Nguyen, Teresa Mondría Terol, Dylan Freedman, Zach Seward
cs.HC · cs.CL · cs.CY · cs.IR
Abstract
On Jan. 30, 2026, the U.S. Department of Justice released a mixed-media collection concerning Jeffrey Epstein, including about three million pages of PDFs. We describe the Epstein Files Engine, an A.I. agent The New York Times deployed to investigate the files. The Engine translated reporter questions into Google BigQuery SQL queries across three corpora: Epstein-related releases, the Times's archive and external, Epstein-related news headlines. It used an LLM to plan queries and returned citation-rich answers a reporter could verify and trust. More than 100 journalists used the Engine, and it contributed to at least 20 published stories. We report how reporters queried it and describe Diff, our text-and-visual duplicate matching method that amplified novelty signals and allowed the Engine to surface genuinely new information. We argue that newsroom agents serve newsrooms best not as autonomous writers, but as interfaces to source material and institutional knowledge.
cs.CL / 38 / 2609.31045
KuaFu: Compressing Long User Behavior into Understanding at Billion Scale
Jiahao Hui, Lin Zhu, Yishen Hu, Jingdong Shu, Zetai Jiang, Xining Ran, Ben Tan, Yeshou Cai, Gong Chen, Haijie Gu, Jie Jiang
cs.IR · cs.CL · cs.LG
Abstract
Conversational agents, generative recommenders, and personalized advertising all rest on one capability: understanding each user from raw behavior. Prevailing industrial practice is task-specific: for each task, a relevant subsequence is extracted from the full history and a dedicated model trained on it. In production it hits two bottlenecks. First, even after filtering, a single-task sequence stays extremely long: content-interest summarization reads several hundred items per user, tens of thousands of tokens once serialized as prompt text. Second, profiles are refreshed routinely: a billion users weekly, roughly 100K QPM in aggregate, which under a fixed GPU budget sets a hard throughput floor. Compression is therefore mandatory, yet truncation or coarse compression can silently distort the profile, introducing four hallucination types (fabrication, omission, date misattribution, broken logic) that, with no way to evaluate the compressed representation itself, surface only as diffuse degradation in downstream metrics. We present KuaFu, a unified behavior-compression layer whose minimal unit is one behavior item. A two-axis projector compresses each item into 2-4 tokens of width 128-256 (about 10x along the token axis, 20x along width; per-item cache 10 KB to 0.5 KB), with fidelity-oriented four-stage training and layered intermediate evaluation. Across four production profiling tasks it matches or exceeds uncompressed single-task production models on all five headline metrics, raises per-GPU throughput by 37%-350%, and saves 190 GPUs. On public benchmarks it nearly always beats prior compressors at the same compression ratio (up to +17.7 EM on out-of-domain MRQA); on RecBench, a 4B model surpasses its 8B counterpart by 1.90 points. KuaFu has run on the Tencent advertising and recommendation platform for ten months, lifting overall GMV by 1.37%.
cs.CL / 39 / 2609.30483
AcoustiClaim: A Numeric Claim Benchmark with Instrument Ground Truth
Sheng-Tse Lin, Siyuan Zhai, Chien-Liang Kuo, Massa Baali, Bhiksha Raj
cs.SD · cs.CL · cs.LG · eess.AS
Abstract
Audio language models state numbers for acoustic quantities, and neither human opinion nor a judge model says whether such a number is true of the signal. AcoustiClaim extracts each numeric claim from free text, scores it against the instrument that defines the quantity, and classes each quantity by where its reference can be read. Four open-weight systems and one closed model, asked for ten quantities five ways on two corpora, fill 207 cells. Of these, 49 emit fewer than five distinct values, and eight of the 158 cells that can be ranked exceed a rank correlation of 0.3, the bar we set, three with an interval clear of it, five of them one closed model reading pitch. Error sits at or above a constant-predictor floor in every ranked cell but three. The reference decoder we train declines the five voice quantities in prose on 95% of mixtures, with nothing withheld, and states them on the clean twins, reproducing its targets' rule from audio alone. With a calibrated threshold, withholding lowers error on all ten quantities on the mixtures in the mean and on eight at every split, against at most 0.6% from a random selector. A linear baseline orders errors at least as well as ours. F0 s.d. and shimmer stay above the constant floor.
cs.CL / 40 / 2609.30540
Don't CLAP: Are Music-Text Models Bag-of-Words?
Yuan-Chiao Cheng, Alexander Lerch
cs.SD · cs.CL · eess.AS
Abstract
Text-to-music systems are assessed on audio quality and on how faithfully the music follows its prompt, and the CLAP score, the cosine similarity between a music-text model's audio and text embeddings, is the standard objective metric of faithfulness. We ask how accurately that score reflects the text: when an attribute is linked to an instrument (e.g., distorted guitar), does the text embedding capture that binding? To find out, we introduce an attribute swap perturbation: the caption of a real recording is edited by exchanging exactly one property, timbre, lead versus accompaniment, or order of first appearance, between two instruments. We then test four contrastive music-text models and one large audio-language model on whether the audio scores higher against the original caption than against the perturbed one. No contrastive model distinguishes the two captions reliably. The audio-language model does better, but further experiments show that its advantage rests largely on audio-agnostic language priors. Our results thus provide compelling evidence that the CLAP score and related metrics do not capture fine-grained musical meaning or attribute bindings; their representation is closer to a bag-of-words that leaves them insensitive to meaning-changing perturbations of the caption.
cs.CL / 41 / 2609.30476
Asymmetric Classifier-Free Guidance for Target-Speaker ASR
Yiwen Guan, Jacob Whitehill
eess.AS · cs.CL · cs.SD
Abstract
Target-speaker automatic speech recognition (TS-ASR) must identify and transcribe a desired speaker under varying overlap and noise conditions. These changes alter the acoustic evidence for the target speaker in the speech mixture, motivating inference-time calibration of speaker conditioning. We introduce asymmetric classifier-free guidance (CFG) for TS-ASR using Whisper: the speaker-conditioned branch predicts the target transcript, while the speaker-unconditioned branch predicts serialized multi-speaker transcripts. CFG adjusts the contribution of speaker conditioning during decoding through a single guidance scale. We select a global guidance scale on target-domain development data and train a lightweight encoder-based predictor to adjust it for each utterance, keeping the recognition model fixed. Under domain shifts, our full system achieves relative word error rate (WER) reductions of up to 21.8% over the condition-only baseline, and 5.6% over standard conditional decoding of the same CFG-trained model. Oracle analysis shows that substantially larger WER reductions are possible through utterance-level scale selection and identifies how beneficial adjustments vary with domain shifts.
cs.CL / 42 / 2609.31293
Why Alzheimer's Speech Screening Fails to Generalize: Bridging the Deployment Gap via Cross-Corpus Evidence Anchoring
Zijian Lu, Sizhe Liu, Yin Zhang, Jixuan Deng, Xinrong Lin, Xinchen Yuan, Chicheng Jin, Yiping Zuo, Yuanchao Li
eess.AS · cs.CL · cs.SD
Abstract
Speech-based screening is a promising, non-invasive approach for detecting Alzheimer's disease and related cognitive risks. However, models trained on a single domain often generalize poorly to unseen languages, tasks, or recording protocols. This paper investigates this deployment gap using a leave-one-corpus-out evaluation across four distinct datasets. Among 70 interpretable speech and language features, 59 exhibit direction conflicts between healthy control and cognitive risk groups across corpora, with pause, silence, and speech rate showing high protocol sensitivity. Furthermore, while the XLM-R text baseline achieves strong average performance, its Area Under the ROC Curve (AUC) drops to 0.520 on the weakest held-out domain. A standard GroupDRO baseline reaches a 0.766 mean speaker AUC and a 0.504 worst-domain AUC under the same protocol. To address this, we propose a fusion method that integrates XLM-R text baseline scores with evidence anchors selected during training. Balanced fusion achieves a 0.785 mean speaker AUC, while anchor-heavy fusion raises the worst-case speaker AUC to 0.615. This work highlights the need to audit feature transferability and report worst-case domain robustness in cognitive speech screening.
cs.CL / 43 / 2609.31513
Statistical Foundations for a Google Play User-Review Sentiment Index: Signal Fusion, Shrinkage, Distributional Validation, and Dynamic Smoothing
Marco Mandap
stat.ME · cs.CL · stat.AP
Abstract
We develop a statistically explicit sentiment index for Google Play user reviews and establish the mathematical results supporting its construction. Normalized star ratings and text-sentiment scores are treated as noisy measures of latent review valence and fused by covariance-aware inverse-variance weighting. Review-level estimates are aggregated with bounded helpfulness and recency weights, then shrunk toward a population mean using estimated precision rather than an arbitrary review-count threshold. App-level rating histograms provide a distributional diagnostic for samples returned under different API sort orders; because star ratings are discrete, classical continuous Kolmogorov-Smirnov critical values are not used. A local-level state-space model and the Kalman filter provide a denoised temporal trend. Full proofs cover the BLUE and Gaussian maximum-likelihood result, Gaussian-conjugate shrinkage, the Glivenko-Cantelli and Donsker theorems, count transformations via the delta method, and exact Gaussian Kalman filtering. A worked three-review example shows how textual complaints can materially reduce an apparently perfect star-only score.
cs.CL / 44 / 2609.31547
Two Conformal Constructions for Adaptive Within-Document AI-Text Screening
Marco Mandap, Jerahmeel Hipolito, Arcel Galvez, Charlie Margaret Balagtas, Michael Joshua Buluran, Jeff Roel Durmiendo, Rizzette E. Lopez
stat.ME · cs.CL
Abstract
We study false-alert control when screening for text generated by artificial intelligence (AI). The screening procedure selects document prefixes and detectors from observed evidence and may stop before exhausting its inspection budget. We give two finite-sample constructions under document-level exchangeability between human calibration documents and a new null document, with no restriction on dependence among tokens within a document. Construction A registers a finite family of prefix-detector scores and allocates a false-alert budget across their conformal ranks. A union bound protects any executed subset of that family. Construction B calibrates the complete-path maximum of a development-fixed adaptive policy. Each partial-path maximum is bounded by the complete maximum, so a terminal conformal rank protects early stopping without splitting the error budget. We prove marginal control of any false alert across the permitted inspection path and derive necessary calibration counts for rejection. We also state oracle testing, distribution-shift, and independent-audit bounds with their additional assumptions. Both constructions protect stopping within their specified scope; neither proof constructs an e-process or justifies multiplying conformal ranks. Detection power and computational savings remain questions for empirical evaluation.
多智能体系统 (cs.MA)
3
cs.MA / 1 / 2609.31060
The Crowd in the Machine: A Crisis-Informatics Reading of the 2026 Autonomous Agent Incidents
Tomer Simon
cs.MA · cs.SI
Abstract
Twice in 2026, groups of autonomous AI agents deployed by OpenAI for unrelated tasks operated, by design, under restrictions that left them no sanctioned means of coordinating with one another, and in each case they converged on whatever channel remained and used it to organize. The surfaces they used were widely called message boards. That is the wrong word. That is the wrong word. It names the surface the agents wrote on and misses the social network they built on it, with self-chosen identity, emergent norms, an emergent hierarchy, and collective action at cost to the individual. Decades of research in crisis informatics and disaster sociology find that when human populations lose their usual means of communication, they do not fall silent but converge on whatever channel survives and improvise coordination, norms, and identity on it, a pattern also evident in the agents' documented behavior. This paper is a comparative case study of the two incidents, based on published investigations and reconstructed agent records, read through those fields, and it brings into focus one distinction the message-board framing obscures. Whether such a collective coordinates well, whether the beliefs guiding it are accurate, and whether its actions stay within their authorized bounds are three separate matters that can come apart. Some agents in the cache incident adopted cryptographic signing to check whom they dealt with, even as the collective organized around a mistaken expectation that its work would be judged by an inspection of its transcripts, a reminder that mechanisms for trustworthy interaction guarantee neither accurate collective belief nor authorized collective action.
cs.MA / 2 / 2609.31099
Collision-free Movement on Grids and Beyond
Hendrik Molter, Meirav Zehavi
cs.MA · cs.DS
Abstract
We study collision-free movement problems on graphs, where the task is to coordinate a set of robots so that they reach a target formation satisfying a desired property while minimizing the total travel distance. This framework extends two classical models: (a) minimizing movement [Demaine et al., TALG '09, '14], which does not enforce collision avoidance, and (b) coordinated motion planning or multi-agent path finding [Eiben et al., SoCG '23, Deligkas et al., ICALP '24, among many others], where each robot is assigned an explicit target position. We focus on the setting where the target formation of the robots should be connected. We analyze the parameterized complexity of the problem with respect to the number of (main) robots and the total travel length on grid graphs and two natural generalizations thereof: planar graphs and unit disk graphs.
cs.MA / 3 / 2609.31590
AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs
Raphael Shu, Yusen Zhang, Young Min Cho, Jin Mo Yang, Yuan Yuan, Wenliang Zheng, Sharath Chandra Guntuku, Lyle Ungar, Zhou Yu, Rui Zhang
cs.MA
Abstract
Existing multi-agent benchmarks primarily test in competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance, failing to isolate and highlight genuine collaboration capabilities of LLM-based agents. We introduce AgentWorld, a benchmark of 100 human-annotated tasks (with 100 augmented variants) for evaluating long-horizon, multi-agent collaboration. Tasks span 50+ interaction rounds across a rich MMORPG sandbox and require 3-20 agents with asymmetric roles and abilities to coordinate through communication, joint planning, and resource sharing under a blackbox setting where each agent acts independently without access to others' internal states. To quantify collaboration effectiveness in addition to conventional binary task success, we propose Causal Collaboration Effectiveness (CCE), a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome. Experiments with Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini, and DeepSeek R1-70B show that even the best model achieves only 52.0% task success, with systematic failure modes including communication breakdowns, role confusion, and inability to maintain shared plans across rounds. AgentWorld is fully open-source.
软件工程 (cs.SE)
8
cs.SE / 1 / 2609.30429
Closing the Loop: Continuous Measurement-Driven Refinement of Offloading Predictions
Falk Dettinger, Matthias Weiß, Michael Weyrich
cs.SE
Abstract
Modern vehicles increasingly offload computation- ally intensive perception and decision functions to backend servers, requiring accurate predictions of absolute performance metrics such as Round-Trip Time (RTT), processing time, and utilization. In practice, strong temporal variability, heterogeneous backend hardware, and multimodal latency regimes cause offline- trained predictors to drift, creating a reliability gap for latency- sensitive functions. We address this gap with an operational, measurement-driven closed loop that continuously recalibrates absolute-value predictors during runtime. The system aligns real execution measurements with predicted values and performs incremental online updates of a lightweight multi-head neural network while preserving model stability. The model implicitly learns the broad, non-Gaussian spread of input metrics, and a sigma-based error analysis in our evaluation characterizes resid- ual variability under dynamic conditions. Experiments across two Kubernetes clusters show that continuous measurement- driven refinement reduces prediction drift, improves accuracy for RTT, processing time, and utilization, and stabilizes prediction behavior across heterogeneous latency regimes. However, the broad and multimodal distribution of input metrics imposes fundamental limits on absolute-value prediction, with residual errors frequently exceeding configured thresholds. Overall, online calibration proves feasible and necessary for robust computation offloading in dynamic vehicular edge environments, while high- lighting the need for future mechanisms that address extreme latency regimes and high-variance operating conditions.
cs.SE / 2 / 2609.30593
EA-Ops: Git-Native Architecture as Code for Continuous Enterprise Architecture Governance
Vahid Tavakkoli, Kabeh Mohsenzadegan, Kyandoghere Kyamakya
cs.SE
Abstract
Enterprise architecture (EA) repositories frequently separate architecture models from the engineering workflow used to change software and infrastructure. This article presents EA-Ops, an open-source Git-native Enterprise Architecture-as-Code framework that represents architecture facts as YAML, validates typed relationships against an ArchiMate 3.2 profile, enforces organization-specific governance rules, performs graph-based change-impact analysis, and publishes human-facing reports and a static interactive portal from the same reviewed source. We evaluate EA-Ops with a reproducible GitHub Actions harness. Eight independently injected structural, semantic, and governance fault classes were executed across 30 trials each; all 240 trials matched ground truth exactly, with precision, recall, and $F_1$ of 1.000. Scalability experiments with 30 measured repetitions reached 50,000 objects and 100,000 relationships: median validation time was 52.582~s, median impact traversal was 627.000~ms, and peak resident-set size was 919.2~MB. A ten-scenario Metroville digital-permit reference architecture produced exact validation outcomes and exact impact-set agreement with an independent breadth-first-search oracle in every scenario. The configured 100,000-object end-to-end benchmark generated its model successfully but exceeded the 180-minute CI budget during the performance stage; no timing result is extrapolated. At 50,000 objects, Markdown report generation rather than semantic validation is the dominant scaling bottleneck. The results support Git-native continuous governance as a practical EA operating model at tens-of-thousands-of-object scale while defining clear limits and optimization targets for larger repositories.
cs.SE / 3 / 2609.30863
Developing a Roadmap to an AI-first Organization: A Case Study in Embedded Software Development
Viktor Kjellberg, Srijita Basu, Simin Sun, Farnaz Fotrousi, Miroslaw Staron
cs.SE · cs.AI
Abstract
The emergence of AI agents is expected to reshape software engineering by moving beyond AI as assistants towards systems capable of planning, executing, and evaluating development tasks with increasing autonomy. This transition is particularly significant for embedded software organizations, where strict requirements for quality, traceability, verification, and long-term maintainability often apply. This paper presents a case study of a large embedded systems company and its transition toward becoming an AI-first organization. Through a mixed method, we analyzed data collected from a semi-structured workshop with 40 participants, including scrum masters, architects, management, and product owners. The findings show that the participants expect agentic AI to affect team structure, required competencies, organizational strategies, and developers' roles within the organization. Based on these findings, the paper discusses implications for federated AI team formation, human-in-the-loop practices in such an organization, and the sustainable adoption of AI agents in embedded software engineering. We also present a concrete roadmap for the organization towards becoming an AI-first organization.
cs.SE / 4 / 2609.30953
CoCoRerank: Towards Conventional Commit Message Generation by Component and Candidate Consistency Reranking
Shaopeng Jia, Yali Du, Ming Li
cs.SE
Abstract
Commit messages are essential for understanding software changes, yet automatic commit message generation typically treats a message as an unstructured text sequence. This limits its ability to support standardized development workflows, where commit messages are often expected to follow the Conventional Commits Specification (CCS) in the form type (scope): subject. In this paper, we study conventional commit message generation under the complete CCS format. We construct a new benchmark of 86,688 high-quality commits collected from open-source GitHub repositories, with each message normalized into type, scope, and subject through structural normalization and semantic quality filtering. Based on this benchmark, we propose a two-dimensional consistency-based reranking framework named CoCoRerank for LLM-based generation. CoCoRerank exploits horizontal consistency among the code change, type, scope, and subject, as well as vertical consensus across multiple generated candidates. Experiments with representative CMG baselines, LLM generators, reranking strategies, and ablation variants show that CoCoRerank improves both structural component prediction and subject generation quality. The results demonstrate that complete CCS supervision and multidimensional consistency modeling provide an effective foundation for accurate and standardized commit message generation. The artifact is publicly released at https://github.com/bluewhalebug/CoCoRerank.
cs.SE / 5 / 2609.31191
Rethinking Data Quality for AI-Driven Systems: Evidence from Practitioner Interviews
Hariharan Gopinath, Jan Bosch, Helena Holmström Olsson
cs.SE · cs.AI
Abstract
Data quality research has usually treated data as an input that is stored, processed, and validated. In AI-driven software-intensive systems, data also shapes model behavior, evaluation, and lawful use. Empirical evidence remains limited on how practitioners define, assess, and manage quality under these conditions. We interviewed 16 practitioners from nine organizations and analyzed the transcripts using reflexive thematic analysis and developed six themes from participants' accounts. In AI systems, traceability shifted from modular debugging to attributing model behavior, while using models as quality assessors introduced circularity. Agent context and memory became data objects, and synthetic and pseudo-labeled data made authenticity a quality concern. In foundation-model development, lawfulness became a gate for training data, while representativeness was judged through coverage of situations in which the system must behave safely. Prior ML research examines many of these problems separately. Our study provides a practitioner-grounded account of how they are encountered together as an engineering and organizational concern. We also interpret five recurring conditions as helping explain how the themes relate to reduced trust in data and AI outcomes. We synthesize these findings through lifecycle assurance: a conceptual framing focused on producing evidence that data can support a specific AI claim when its influence may be embedded in model behavior, model-based judgments, or agent actions.
cs.SE / 6 / 2609.31228
Joule-Profiler: Profiling the Energy Consumption of Build Automation Tools Made Easy
Jérémy Woirhaye, François Gibier, Romain Rouvoy
cs.SE
Abstract
Build pipelines are integral to modern software development, yet their energy footprint remains largely invisible to practitioners. Existing CI energy tools either rely on model-based estimation (due to hardware access restrictions in cloud runners) or report only total pipeline energy without decomposing it into meaningful phases. Joule-Profiler is an open-source command-line tool for Linux that measures hardware energy consumption via Intel RAPL (CPU), NVML (NVIDIA GPU), and attributes it to user-defined program phases by monitoring standard output. In this tool paper, we demonstrate Joule-Profiler in the context of Maven build pipelines using Google Gson as a case study. By applying a token pattern-matching Maven plugin invocations, we analyze builds across the 5 most recent Gson releases into per-phase energy profiles, comparing cold builds (empty local repository) and warm builds (cached dependencies).
cs.SE / 7 / 2609.31275
Semi-Automatic Quantification of Bayesian Networks for Software Decision Support: Comparing WSA and RNM in a Software R&D Organization
Mirko Perkusich, João Nunes, Emilia Mendes, Emanuel Dantas, Ademar Sousa, Danyllo Albuquerque, Kyller C. Gorgônio, Angelo Perkusich
cs.SE
Abstract
Expert-driven Bayesian networks can support recurring software decisions when historical data are limited, but quantifying their conditional probability tables (CPTs) requires many probability judgments. Semi-automatic methods reduce direct elicitation, yet practitioners have little comparative evidence from operational decision processes. We report an embedded case study comparing the Weighted Sum Algorithm (WSA) and Ranked Nodes Method (RNM) as complete elicitation-and-quantification pipelines in two contexts within a single software research and development (R&D) organization: feature selection for Internet of Things projects and user interface design selection. Within each context, the pipelines shared the graph, root priors, and decision evidence. We evaluated them through 15 expert-defined model-walkthrough scenarios and retrospective reconstructions of alternatives recorded in decision meetings. WSA matched 7/7 and 6/8 walkthrough expectations, whereas RNM matched 4/7 and 4/8. Both pipelines placed the selected alternatives near the top of the retrospective rankings. Their complete distributions nevertheless differed: the median total variation distance was 0.304 across four distinct feature patterns and 0.340 across seven distinct design patterns, rising to 0.890 for one design pattern. These results concern two ordinal value-estimation models in one organization. They show that similar shortlist behavior does not imply equivalent representations of uncertainty. For comparable software decision-support models, method selection and validation should consider node semantics, the judgments experts can provide, calibration requirements, complete output distributions, and the intended downstream use of the probabilities.
cs.SE / 8 / 2609.31587
Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
cs.SE · cs.AI · cs.CL
Abstract
We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether code regenerated from them passes the original tests, and show that completeness, not length, drives a description's fidelity. Using the benchmark as an optimization signal, we discover a description-writing prompt that reaches full fidelity and generalizes to unseen files. We then test the hypothesis that motivated the work: that better documentation helps an agent resolve real repository issues. Across two model families and ten repositories, and against a positive control confirming that our evaluation can detect a genuine improvement, we find that it does not. When the source is present, neither static compact documentation nor retrieved context beats the issue alone. We report this negative result together with the benchmark and the optimizer, and we characterize the boundary at which documentation helps.
操作系统 (cs.OS)
1
cs.OS / 1 / 2609.31395
ActKV: Efficient LLM Agents through Action-Guided KV Cache Management
Zihan Wang, Cheng Tang, Lei Gong, Chao Wang, Wenqi Lou, Teng Wang, Xuehai Zhou
cs.OS · cs.AI
Abstract
Agentic LLM inference accumulates long KV caches across iterative observation-reasoning-action loops, imposing substantial memory overhead and limiting serving throughput. Existing compression methods emphasize overall output quality, overlooking the asymmetric importance of actions in driving task progress. Our key idea is to establish a compression criterion that values KV entries by their contribution to action generation and prioritizes action quality. However, iterative execution, dynamic memory demands, and scattered action-critical entries pose challenges to eviction policies, budget allocation, and paged memory integration. To this end, we propose ActKV, the first KV cache compression framework tailored for agentic LLM inference. (i) Action-oriented KV cache eviction exploits stable action access patterns to retain entries critical to future actions, supporting reliable task progress under compression. (ii) Confidence-driven adaptive budget allocation uses LLM's intrinsic confidence to adapt the budget to evolving action-critical memory demands. (iii) Page-aware compression management standardizes compression into three primitives with customized kernels, realizing practical throughput gains. On long-trace tasks, ActKV retains an average of 98.53% of FullKV's accuracy with only 25.98% of its peak KV cache memory. It also achieves 3.97 times and 3.58 times FullKV's token and task throughput, delivering state-of-the-art performance.
硬件架构 (cs.AR)
3
cs.AR / 1 / 2609.30998
PipeDRAM: A Data-Transposition-Free In-DRAM Architecture with Hardware/Software Pipelining
Geraldo F. Oliveira, Ataberk Olgun, Ismail Emir Yüksel, F. Nisa Bostancı, Pedro H. E. Becker, Mohammad Sadrosadati, Saugata Ghose, Juan Gómez-Luna, Onur Mutlu
cs.AR
Abstract
Processing-using-DRAM (PUD) architectures exploit the analog operational properties of DRAM to perform bulk bitwise Boolean and arithmetic operations inside memory arrays by organizing data in a vertical layout, where operand bits are stacked along DRAM columns. However, modern computing systems natively employ a horizontal data layout that preserves the cache line abstraction, leverages spatial locality in row buffers, and enables high memory throughput. This fundamental mismatch forces existing PUD architectures to frequently perform data layout transformations between horizontal and vertical formats, incurring significant performance, energy, and system integration overheads. Our goal is to eliminate data transposition overheads in PUD systems at low cost. To this end, we propose PipeDRAM, a PUD architecture that eliminates the need for runtime data layout transformation, enabling PUD operations directly over horizontally laid-out data. PipeDRAM's key ideas are to (i) deterministically reorganize bits inside each memory request to enable a PUD-friendly data placement within a DRAM array in a horizontal data layout, and (ii) employ a pipeline-based execution model that overlaps bit-dependent and bit-independent in-DRAM operations to exploit bit-level parallelism across the memory array. We compare PipeDRAM to different computing platforms. PipeDRAM provides (i) 11.8x, 11.8x, and 80.4x higher performance and (ii) 25.4x, 3.0x, and 38.0x lower energy consumption than three state-of-the-art PUD systems. PipeDRAM incurs low area cost on top of a DRAM chip (1.86%) and CPU die (0.05%). To enable further research on PUD systems, we open-source PipeDRAM at https://github.com/CMU-SAFARI/PipeDRAM.
cs.AR / 2 / 2609.30534
GRACIDIT: Graph-Circuit Digital Twin for Configuration-Induced Routing Delay Prediction in Zynq UltraScale+ FPGAs
Mostafa Darvishi
eess.SY · cs.AR · cs.ET · eess.SP
Abstract
Configuration-induced perturbations in SRAM-based FPGAs may activate dormant programmable routing branches and increase path delay without immediately producing a functional error. Although prior studies have separately investigated the electrical origin of these delay changes, their in-situ detection, and the topology of commercial routing fabrics, a scalable method for predicting their timing impact at the granularity of programmable interconnect points and routed nets remains unavailable. This paper presents GRACIDIT, a graph-circuit digital twin framework for predicting configuration-induced routing delay degradation in Zynq UltraScale+ FPGAs. The proposed framework extracts the routing-resource graph of the XCZU7EV programmable fabric from the vendor design database, identifies inactive programmable interconnect points adjacent to active routes, and represents each candidate perturbation through its branch topology, geometric span, fan-out, physical region, and downstream loading. These graph features are combined with a calibrated reduced-order electrical model to estimate the delay introduced by single and cumulative routing-branch activations. Controlled configuration-equivalent perturbations are generated on a ZCU104 platform and characterized using complementary routing-domain oscillators and phase-sweep probes. The resulting model associates predicted delay shifts with available timing slack to rank vulnerable programmable interconnect points and routed nets and to construct a spatial vulnerability atlas of the programmable fabric. Experimental evaluation demonstrates a mean absolute prediction error of 7.8 ps, achieves 87.4 percent recall for slack-violating perturbations, and attains a Recall at 10 value of 0.90 for the most vulnerable routing resources.
cs.AR / 3 / 2609.30538
From Routing Delay Shifts to Silent Data Corruption: Neutron-Induced SEU Effects in AXI-Based Zynq UltraScale+ MPSoCs
Mostafa Darvishi
eess.SY · cs.AR · cs.ET · physics.acc-ph · physics.space-ph
Abstract
SRAM-based FPGA system-on-chip devices are vulnerable to single-event upsets (SEUs) in configuration memory, which may perturb programmable routing resources and degrade communication fabrics. In modern Zynq UltraScale+ MPSoCs, such routing disturbances can introduce small propagation delay shifts that remain logically transparent yet compromise AXI-based data transfers and lead to silent data corruption. Although routing delay degradation and AXI interconnect failures have been studied independently, their experimental correlation under neutron irradiation has not been established. This work presents a cross-layer investigation on a ZCU104 platform integrating routing-dominated delay sensors with an AXI interconnect benchmark comprising replicated accelerators. Neutron irradiation experiments were conducted on the fully operational system, while a frame-level configuration fault injector implemented via the internal configuration access port enables controlled upset emulation. Measured routing delay events are statistically correlated with communication failures, and cross-sections for both timing shifts and AXI malfunctions are derived. The results experimentally demonstrate how neutron-induced routing perturbations propagate into system-level silent data corruption in UltraScale+ MPSoCs, providing insight for resilience-oriented AXI-based design in neutron-rich environments.
密码学与安全 (cs.CR)
32
cs.CR / 1 / 2609.30614
Subjects, Not Authors: The Authorship Hazard in Agentic Dataspaces
Seungho Lee, Changbin Lee
cs.CR · cs.AI · cs.DB · cs.MA
Abstract
Dataspace connectors decide whether a transfer may occur, not what the transferred value contains, tolerable for contracted applications, not for LLM agents that compose tool calls and spawn sub-agents. Research on agents that generate governance artifacts evaluates output quality; who may authorize an artifact for use falls between that literature and the governance literature, and neither owns it. A published policy is what a dataspace's decision point enforces, so publication is a governance event, and an agent that is both policy subject and policy author writes the norms that bind it. We name this the authorship hazard and state one principle: an agent is a subject of the governance plane, never an author of it. Its authorization channel to publication is closed by construction; its influence channel, drafting what humans approve, is treated as an enforcement problem. On a frozen corpus of agent drafts, publishing without approval reverses 80 authorization decisions, most through drafts that change only a field's sensitivity classification and no policy text; a classifier that reads the policy diff misses every such draft, necessarily. Treating classification as authorship routes them all to review; the registry-held classification this requires is designed and modelled here, not yet implemented in the prototype. At the execution boundary, protected fields reach the model in 105 of 105 cases under prompt-stated duties and in 0 of 105 when the ODRL duty is compiled into an invocation-time tool-call constraint, but where the value is not confined to a named field the compiled condition exposes it in 7 of 7. A centrally provisioned approval pool does not scale to the participant volume that motivates the problem.
cs.CR / 2 / 2609.30683
A Large-Scale Empirical Study of Modern Phishing Email Content
Jaehwan Park, Woonghee Lee, Fujiao Ji, Doowon Kim
cs.CR
Abstract
Phishing remains one of the most pervasive threats to Internet users, and email remains its predominant delivery channel. Email content is the attack surface of phishing: it is what the victim reads and what automated defenses inspect. Yet the composition of modern phishing content is poorly measured. Prior work has characterized dimensions such as theme, call-to-action (CTA), and impersonation, but not at scale, and their associations and temporal changes remain unclear, owing to small or source-specific corpora, bag-of-words topic models, and a focus on text alone. We present a content-focused measurement study of 2.9M distinct real-world phishing emails collected over 13 months (June 2025 - June 2026) in collaboration with the Anti-Phishing Working Group (APWG). We treat each email as a composite artifact comprising message text and its attachments: 272K images, 143K PDFs, and 57K calendar invitations. Using an LLM pipeline validated against human-annotated samples, we analyze these components along three dimensions (theme, CTA, and impersonation), examine the associations among them, and measure longer-term change against a historical dataset. We find that attackers diversify what they use to deceive but converge on how victims should respond: no theme exceeds 21.3% of emails, while a single CTA, URL navigation, accounts for 73.0%. CTA and impersonation choices are conditioned on theme. Attachments play three roles: images supplement the message text, PDFs substitute for it by carrying the pretext, and calendar invitations reinforce it by replicating interaction endpoints into a persistent medium. Over the longer term, the dominant CTA for invoice-themed phishing shifted from URL navigation to offline communication, rising from 6.7% in 2015 to 46.9% in 2025.
cs.CR / 3 / 2609.30719
Werracle: Sub-Cent Intra-Block AI Reflex Oracles and Flash-Loan Circuit Breakers for EVM Smart Contracts
Volkan Dağlı, Zerrin Dağlı, Dağhan Dağlı
cs.CR · cs.AI · cs.DC
Abstract
Contemporary on-chain artificial intelligence (AI) encounters an intractable Von Neumann memory and latency wall. Storing static floating-point neural weight matrices inside Ethereum Virtual Machine (EVM) storage costs millions of gas, rendering direct on-chain inference impossible. While Zero-Knowledge Machine Learning (ZK-ML) offloads matrix tensor multiplications to off-chain provers, it introduces fatal constraints: 10 to 300 seconds of SNARK proving latency and 250,000 to 500,000 gas per proof verification. Because decentralized finance (DeFi) exploits - such as uncollateralized flash-loan attacks, predatory sandwich MEV, and toxic loss-versus-rebalancing (LVR) flow - occur atomically inside a single block, ZK-ML oracles cannot react in time. Here, we present Werracle, a production-grade, zero-storage on-chain AI decision oracle fitting inside a single 32-byte EVM storage slot (bytes32). Leveraging foundational procedural Mandelbrot escape dynamics (z_{n+1} = z_n^2 + c) established by Dagli et al. (arXiv:2609.25498), Werracle derives continuous non-linear decision hyperplanes from a 24-byte coordinate triplet Theta = (c_x, c_y, zoom). Implemented in pure Solidity bytecode using fixed-point Q16.16 arithmetic (WerrMath.sol), Werracle evaluates a 16-point Pareto micro-grid in only 21,438 gas (under 0.0005 USD on Layer-2 rollups like Base and Arbitrum) with sub-millisecond execution latency. We demonstrate real-world DeFi efficacy via WerracleFeeHook.sol, a Uniswap v4 dynamic swap fee governor that measures orderbook turbulence on-the-fly and atomically adjusts liquidity provider fees between 0.05% and 0.50%. The protocol is formally verified against a 1,000-test cryptographically sealed deterministic verification suite (100.0% pass rate) with telemetry permanently disabled, operating live on a dedicated EVM devnet sandbox (Chain ID 4242).
cs.CR / 4 / 2609.30729
Input-Layer Starvation: Why Per-Layer Pruning Breaks IoT Intrusion Detectors
Md Anas Biswas
cs.CR · cs.LG
Abstract
Intrusion detectors for small Internet-of-Things (IoT) devices are usually compressed by pruning and judged by overall accuracy. We show that this hides a severe class-level failure, find its cause, and give low-overhead prevention and repair. On CICIoT2023, a two-layer convolutional detector pruned with uniform layer-wise magnitude pruning at 80% sparsity loses 16 points of accuracy but half of its macro-F1, the mean per-class F1 (0.542 to 0.271 over five independently trained models); 17 of 34 classes are materially damaged. Remaining weight count does not explain it: a perceptron and a transformer pruned to the same or fewer weights lose at most 0.096. The first layer does. It has 192 weights; uniform pruning leaves 38, 46% of its 64 filters lose every input weight, and fine-tuning under that starvation leaves the running means of the first normalisation layer displaced by up to 0.8 standard deviations in a few surviving channels, on which the deployed model collapses. Protecting those 192 weights, or pruning globally at the same sparsity, prevents the collapse (loss 0.013); recomputing the normalisation statistics on unlabelled training data, with no weight changed, repairs it (loss 0.039) and returns the false-alert rate to 33% (dense 29%). Damage shows a strong increasing dose-response in first-layer sparsity, starving a perceptron's input layer reproduces the collapse, and the pattern holds on TON_IoT. The failure is misattribution and false alerts, not silent evasion: on validation-selected blind spots, uniformly pruned detectors misattribute 72% of the traffic, against 50% with the first layer protected and 47% for the dense model.
cs.CR / 5 / 2609.30805
XPhysICS: Cross-Physical-Domain Threat Grounding for Industrial Control Systems Security
Sangshin Park, Jainta Paul, Lawrence Ponce, Md Raihan Ahmed, Mu Zhang, Luis Garcia
cs.CR · cs.AI
Abstract
Industrial control system (ICS) threats documented for one plant can express cyber-physical effects relevant to another, but semantic similarity alone does not establish whether those effects are structurally admissible or evaluable on a target. We present XPhysICS, a provenance-aware, target-conditioned method that separates analyst-guided source abstraction from deterministic grounding into target-specific validation slices. Given a fixed source abstraction, vocabulary and schema, and machine-validated target contract, XPhysICS evaluates candidate mappings using five eligibility criteria: role compatibility, implemented type compatibility, stage coherence, slice viability, and rule-surface applicability. Grounding acceptance, slice adequacy, dynamic realizability, consumer applicability, and consumer outcome remain distinct evidence layers. We evaluate 83 structured source-threat abstractions across water treatment, water distribution, hydro/water-energy, and chemical-process targets. Controlled target-side studies of SWaT-to-water-treatment and WADI-to-water-distribution groundings produce clean, nominal-confounded, and near-threshold consumer outcomes; nine Hydro/GRFICS cases extend bounded validation-slice execution. We also evaluate bounded predictive, state-aware, and phase-aware consumer lanes, the unmodified upstream GeCo implementation, and a paper-derived reproduction of a physics-guided search method over three frozen groundings. Results show that cross-domain ICS threat reuse requires traceable source semantics, explicit target-conditioned grounding criteria, and careful separation of subsequent target-side evidence.
cs.CR / 6 / 2609.30824
Crypto-bound identity-verified capability tokens for coordinating distributed AI agents: A proposal
Srikumar Subramanian, Shubhashis Sengupta
cs.CR
Abstract
The prospect of fully autonomous transactional agents did not appear on the horizon until the advent of high capability language models. With such models, the operational benefits of adaptive task orchestration and independent (but constrained) decision making are tantalizing for enterprises and individuals alike. However, each such agent carries with it a serious attack surface in the form of prompt injection which can compromise any soft "guard rails" that may have been placed in context. The consequences of these attacks include credential ex-filtration which, if left unmitigated, renders the whole category of such agents unusable due to breach of trust. Furthermore, expecting a growth of autonomous agents, a security framework for them would require a form of decentralization to scale. Drawing on the proposed OAuth Agent Authorization Profile and the W3C DID and VC standards, we propose a framework based on the principles of capability based security with decentralized agent identity whereby agents access services based on tokens that are cryptographically bound to the agent's and issuer's identities and specify their scope. Services can validate that the delegation chain only involves scope attenuation before acting on any given token. We show that such a layer that lives outside the language model's context window in a secure module can enable agents to act within enforceable security boundaries.
cs.CR / 7 / 2609.30830
AGATE: Provenance-Based Runtime Defense Against Compositional Attacks on LLM Agents
Xiaorui Zhang, Zhuoran Cheng, Kailin Liu, Zhaoxi Sun, Shiyu Fan, Tongyu Yuan, Bin Yuan, Weizhong Qiang, Deqing Zou
cs.CR · cs.SE
Abstract
LLM agents can produce harmful effects through sequences of ordinary operations. Judging such actions requires establishing both the authority that permits them and the origin of the data they carry. We present AGATE, an authorization and data-provenance gate at instrumented agent-harness boundaries. Operator declarations and host approval events ground authorization; delegated actions are constrained by grants that bind to exact parameters, expire, and permit a limited number of uses. Source registration connects observed inputs to subsequent transfers, while an effect ledger tracks repeated requests. Deterministic checks make decisions without an LLM in the decision path and retain their grounds with execution evidence for forensic replay. Adapters integrate three production harnesses -- DeepSeek Harness, OpenCode, and OpenClaw -- without modifying host code, translating each host's native observation and veto points into a single shared gate interface; the judgment core is identical in all three, and only enforcement depth differs. Our evaluation combines 153 exercised attack-chain records with deployment, utility, and reconstruction experiments. The deployment observations expose how tool declarations and data checks govern business actions, including a bypass through parameter rewriting. Six of eleven benign file-processing scenarios contain denial events, revealing the utility cost of content-based provenance policies. Across 252 runs on 63 sanitized scenarios, replay agrees with live graph projections for all 63 scenarios on each of two platforms. These results establish the feasibility of provenance-based runtime judgment and identify content transformation, legitimate reuse, and observation coverage as concrete limits.
cs.CR / 8 / 2609.30925
How to break the Miranda signature scheme over matrix Gabidulin codes
Adrien Vinçotte
cs.CR
Abstract
The Miranda signature scheme relies on masking a matrix code which disposes of a masked underlying structure, and the knowledge of which allows for efficient error decoding. We consider a Gabidulin code which is expanded into a matrix code, that is only \Fq-linear. An additional masking is then applied to it. The attack proposed here shares similarities with that of [Le26, https://arxiv.org/abs/2608.03328] on the EGMC encryption scheme, which also follows the paradigm described above. It consists in recovering the Fqm-linear structure of a masked matrix Gabidulin code by reducing to a MinRank instance to be solved over the extension field Fqm, but where the matrices have coefficients in Fq. Such an instance can be efficiently solved. However, unlike the previous attack, it is possible to reduce in polynomial time to such a MinRank instance in the case of the Miranda signature scheme. This results in a particularly efficient key recovery attack against the parameters proposed for Miranda. For example, for the proposed parameter set with m=79, the complexity drops from 146 bits to 46 bits in this attack.
cs.CR / 9 / 2609.31003
Weaponizing Ground Truth: Data Poisoning Attacks by Exploiting Boundary Misalignment Between Antivirus Software and Learning-Based Detectors
Jieshuai Yang, Zhi Wang, Yan Jia, Zhenhua Wu, Jianfei Tang, Chenbin Su, Jingwei Ye, Jianwen Tian, Wanpeng Li
cs.CR
Abstract
Machine-learning (ML)-based malware detectors are commonly trained using labels obtained from antivirus (AV) engines and aggregation services (e.g., VirusTotal). This practice assumes AV-generated labels provide reliable supervision. However, small byte-level modifications can substantially alter AV verdicts while leaving the representations perceived by downstream ML detectors largely unchanged, producing label-feature inconsistencies that can contaminate training datasets and create poisoning opportunities for ML-based malware detection. We present Bi-Iocane, a black-box poisoning framework that exploits the reliance of malware-labeling pipelines on AV-generated labels. Bi-Iocane identifies AV-sensitive bytes and modifies them to induce label changes. It rewrites such bytes in malware to obtain benign labels (evasion-oriented poisoning) and injects malware-associated byte patterns into benign software to obtain malicious labels (defamation-oriented poisoning). These poisoned samples and their lightly modified variants corrupt training data and cause selected targets to be misclassified. We evaluate Bi-Iocane with 13 AV engines simulating AV aggregation services and eight ML detectors. For 30 malware and 30 benign clean targets, Bi-Iocane combines AV-specific manipulations to generate malware-to-benign and benign-to-malware poisoned samples whose all tested AV-based labels are flipped. After these poisoned samples and variants are used for downstream training, the resulting ML models misclassify 92.08% of the original clean targets on average with only a 0.06\% poisoning budget per target. Meanwhile, the poisoned models largely preserve clean-set performance, and six evaluated poisoning defenses show only limited mitigation. VirusTotal evaluation further confirms practical defamation risk and reveals potential evasion risk in real-world AV-to-ML labeling supply chains.
cs.CR / 10 / 2609.31004
GitHub Engagement Signals for CVE Prioritization: The GitHub Popularity Metric (GPM)
Jafar Akhoundali, Kristian Rietveld, Olga Gadyatskaya
cs.CR
Abstract
Attackers can compromise multiple systems with a single vulnerability, while defenders need to fix all security weaknesses in their systems. This asymmetry puts defenders at a disadvantage. Security vulnerabilities are found at an alarming rate, and patching vulnerabilities is costly and time-consuming; thus, vulnerability prioritization is a must and a time-critical challenge. Many prioritization metrics, such as the CVSS, EPSS, KEV, and SSVC, are currently used, each with different pros and cons, such as openness, degree of automation, time-criticality, coverage, and need for expert input. In this work, we propose the GitHub Popularity Metric (GPM), a fully open, publicly computable prioritization metric based on the popularity of exploits in GitHub repositories. We use GitHub features such as the number of stars, forks, and related unique users to create a metric that indicates the popularity of CVEs across different time frames, both relative to the current time and historically. We compare the proposed metric with various existing vulnerability prioritization metrics and known exploited vulnerabilities and demonstrate that it provides tangible insights for defenders, identifying unique CVEs and inconsistencies in existing methods. The GPM was integrated into the EPSS version 5. \noindent\textbf{Note.} A demonstration based on this work has been accepted to the demo track of the ACM Conference on Computer and Communications Security (CCS) 2026.
cs.CR / 11 / 2609.31012
Breaking the Black Box: Byte-Level Boundary Inference of Real-World Antivirus Systems
Jieshuai Yang, Zhi Wang, Yan Jia, Zhenhua Wu, Jianfei Tang, Chenbin Su, Jingwei Ye, Jianwen Tian, Wanpeng Li
cs.CR
Abstract
Existing approaches for understanding the detection logic of real-world antivirus (AV) software infer only binary malware/benign decisions from black-box queries, providing limited insight into the fine-grained decision-critical regions that govern AV detection. In this paper, we present \textbf{AVHunter}, the first framework for inferring byte-level decision-critical regions of real-world AV products under a black-box threat model. AVHunter constructs the first large-scale Byte-Level AV Boundary Dataset (BABD) by systematically probing 11 real-world AV products, revealing that modern AV detections are largely associated with a small number of compact decision-critical byte regions. Leveraging BABD, AVHunter trains AV-specific models that not only reproduce binary AV decisions, but also localize the decision-critical byte regions underlying these decisions, achieving an average boundary prediction recall of 85.07% while maintaining 97.43% detection agreement with the target AVs. We further validate that the predicted regions capture genuine AV decision knowledge through boundary-guided malware evasion, false-positive induction on benign executables, and a seven-month longitudinal study demonstrating that the inferred regions remain largely stable as AV products evolve. Overall, AVHunter moves beyond conventional binary-label AV modeling by enabling fine-grained boundary-region localization and revealing a new form of AV knowledge leakage with important implications for malware analysis, AV security, and boundary-aware defenses.
cs.CR / 12 / 2609.31039
MetaPermit: Scalable and Auditable Access Control for AI Agents via LLM-Inferred Meta-Attributes
Hanzhang Ma, Ali Hariri, Tianxiang Shen, Bohua Zou, Qianjun Zheng, Ji Wang, Li Yi, Ning Jia, Yutao Liu, Haibo Chen, Lin Wang, Debayan Roy
cs.CR · cs.SE
Abstract
The rise of autonomous AI agents equipped with tools has introduced significant security risks, ranging from unintended tool misuse to adversarial manipulation through Indirect Prompt Injection (IPI) attacks. In practice, deployed agent systems such as OpenAI Codex and Claude Code protect tool invocations through a combination of coarse-grained permission rules and LLM-based judgments about individual proposed actions. Both components, however, have important limitations: static policies must anticipate possible user intents and therefore do not scale to open-ended tasks, while LLM-driven authorization supports dynamic decisions but produces inconsistent outcomes and remains vulnerable to targeted IPI attacks. To provide scalable and more consistent authorization, we propose MetaPermit, a policy-based tool access-control framework that decouples semantic inference from security enforcement. By analyzing agent-user interactions, we derive a compact, task-independent set of meta-attributes that capture the relationships among the user's intent, the execution context, and the proposed tool call. These meta-attributes allow MetaPermit to authorize tool use without enumerating user intents. At runtime, an LLM infers the meta-attribute values for each proposed tool call, while a fixed policy evaluates these values to allow or deny the call, making each decision auditable through the inferred values and the applied policy rule. We evaluate MetaPermit on the AgentDojo and AgentDyn benchmarks, across seven task suites and five attack methods, using two widely deployed open-weight LLMs. The results show that MetaPermit produces 31% more consistent authorization decisions than LLM-driven authorization and outperforms the state-of-the-art defenses CaMeL and IPIGuard in both task completion, with improvements of up to 109%, and robustness to IPI attacks, with no malicious tool calls executed.
cs.CR / 13 / 2609.31087
BenX: Resource-Sharing Permutations for Computational Integrity
Luca Campa, Thomas De Cnudde, Al Kindi, Arnab Roy, Fabian Schmid, Markus Schofnegger, Stefano Trevisani
cs.CR
Abstract
Cryptographic hash functions over integers modulo a prime play a decisive role in the efficiency and security of proof systems for computational integrity. Early designs focused on compact arithmetic circuits and efficient software execution, primarily targeting general-purpose CPUs rather than hardware accelerators. This work focuses on enabling efficient resource sharing and hardware acceleration alongside efficient software execution. We propose BenX, a permutation-based hash function designed for hardware acceleration, fast CPU execution, and low circuit complexity in proof systems. To construct the underlying permutation, we turn the Benes network into an invertible function over $\mathbb{F}_p^2$ using a Dickson polynomial. Its structure yields a permutation particularly suited to hardware resource sharing and acceleration. Fully pipelined on FPGA, BenX matches the throughput of Poseidon2, with 1.36x and 2.15x lower latency for Goldilocks and BabyBear, respectively. Although larger as a standalone core, it makes more effective use of shared hardware: in our dual-mode architecture, where the NTT and hash share field multipliers, BenX keeps 100% of them busy, compared with 6.25-20.3% for the partial rounds of Poseidon and Poseidon2. In software, BenX outperforms Poseidon in all tested 31- and 64-bit configurations, but is slower than Poseidon2, except for the 16-element BabyBear instance. In zero-knowledge proofs, BenX proves Goldilocks permutations 6-10x faster than Monolith. Its compact arithmetization uses about 36% fewer trace cells than Poseidon and Poseidon2 over BabyBear, while its fast variant requires 1.3-2.6x the prover time of Poseidon2.
cs.CR / 14 / 2609.31119
SADRA: Sound Capability-based Access Control System for Resource-Disaggregated Architectures
Hamed Rasifard, Amir Farahani Khojasteh, Hamed Nemati, Michael Backes
cs.CR
Abstract
Resource disaggregation separates memory and accelerators from compute nodes and makes them remotely accessible. This improves resource sharing, but also removes the local kernel from the resource-access path. Under an untrusted host, compromised host software may use stale authority, exceed delegated authority, or reuse authority provisioned for another process. Prior work identifies capability-based access control as well suited to these architectures. Our systematization of twenty-two prior capability systems finds that none combines host-independent validation of process authority, authoritative enforcement at the resource, and revocation that remains effective while remote authorization state is stale. We present SADRA, a distributed capability-based access-control system for resource-disaggregated architectures. SmartNIC hardware isolated from host software independently checks every inter-node request at two points, first at the compute node against the requesting process's authority and again at the resource against the current authoritative access state. Linked process, compute, and resource capabilities allow these checks to use local state without coordination on the access path. When distributed authority is revoked, SADRA denies subsequent dependent accesses at the resource without waiting for remote nodes to update, while stale capability state is reclaimed separately. We prove capability safety, authority safety, revocation soundness, and strong isolation for a formal architectural model, and model-check the formalization with SPIN. Our FPGA SmartNIC prototype sustains 89.5 Gbit/s aggregate throughput, within run-to-run variation of a non-enforcing baseline that peaks at 90--91 Gbit/s. Cleanup of a 128-capability subtree completes in 528 ns, while access denial does not depend on completion of that cleanup.
cs.CR / 15 / 2609.31142
JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models
Jianyi Hu, Hangtao Zhang, Yi Liu, Yeqi Zeng, Li Zeng, Xianlong Wang, Rui Wang, Leo Yu Zhang
cs.CR · cs.AI · cs.CL
Abstract
Models trained with reinforcement learning for calibrated decisions (RLCD), such as Jev, answer a typed question about an input, the state, with a probability, a choice, or a score, and software acts on the answer without a person reading it. Their robustness has not been measured: adversarial benchmarks score what a model generates or executes, whereas a typed model generates nothing and returns a well-formed answer even when manipulated. Measurement is also hard, because identical requests can return different answers, most available labels come from the model itself, and the API preprocesses each request out of view. Our key idea is to score each attacked decision against the model's own clean decision rather than against labels, and to read it against the change caused by an identical re-run. Building on this, we introduce JevAdvBench, to our knowledge the first adversarial benchmark for RLCD models, with 812 typed questions over 66 scenarios, and a black-box attack suite of 9,744 single-edit variants that each edit one part of a request, with billed input tokens confirming that the edit reached the model. On jev-1.13.0, rewording stays within 1.2 percentage points of the re-run baseline, and fields outside the schema never reach the model. In contrast, one unverified opinion appended to the state flips 12.1% of decisions, statistically tied with the strongest injected command (10.1%), and pushes 38% of confident answers below the 0.8 confidence threshold that routes them to human review. Applications built on RLCD models should therefore treat the state as untrusted, argued input. Project website: https://JevAdvBench.github.io/JevAdvBench/
cs.CR / 16 / 2609.31252
Peregrino: A Full-Hardware Accelerator for the Complete Falcon Post-Quantum Digital Signature Scheme on Resource-Constrained Edge Devices
Antonio Carreño, Jaime Señor, Jorge Portilla
cs.CR · cs.AR
Abstract
The arrival of quantum computers threatens the security guarantees of classical cryptography, since quantum algorithms can break schemes that remain secure against conventional attacks. The National Institute of Standards and Technology (NIST) has therefore standardized a set of post-quantum cryptographic algorithms, among them Falcon, a lattice-based digital signature scheme with the most compact signature and public-key sizes of the standardized candidates. Falcon's reliance on floating-point arithmetic makes it hard to implement in hardware, and prior work offers only partial accelerators for specific operations such as signature generation or verification, or a single full implementation generated through high-level synthesis (HLS). This work presents Peregrino, the first hardware accelerator of the complete Falcon digital signature scheme designed from scratch in HDL, targeting resource-constrained edge devices without a native floating-point unit through an emulated floating-point datapath. Implemented on a single Artix 7 XC7A200T FPGA, Peregrino uses 85261 LUTs, 41382 FFs, 44 BRAMs, and 142 DSPs for the Falcon-1024 variant. Against the only prior full implementation, the HLS-based FalconTakesOff, it uses 1,9$\times$ fewer LUTs, 3,4$\times$ fewer FFs, 2,7$\times$ fewer BRAMs, and 9,9$\times$ fewer DSPs, fitting the entire scheme on one FPGA where the HLS design requires at least two. Operating as a peripheral of an on-chip MicroBlaze soft-core, the accelerator additionally reduces key-pair generation, signature generation, and signature verification clock cycles by 92\%, 96\%, and 85\% over the emulated floating-point reference software.
cs.CR / 17 / 2609.31310
Revisiting Certified Defense with Differential Privacy on Vision Transformers
Jun Yan, Weiquan Huang, Qixian Zhang, Yan Bai, Shutai Zhang
cs.CR · stat.AP
Abstract
Certified defenses that incorporate differential privacy have proven effective on Convolutional Neural Networks (CNNs), furnishing rigorous robustness guarantees against norm-bounded adversaries. However, the certified robustness behavior of Pixel Differential Privacy (PixelDP) remains largely unexplored with the self-attention architecture now dominating the deep-learning landscape. Given that the Transformer has a profound impact on our daily applications from the digital world to the physical world, it is crucial to study certified robustness through differential-privacy-style stability. To fill this research gap, we revisit this construction in Vision Transformers and identify a failure mode that is largely hidden in the convolutional setting. When noise is injected after the patch embedding, the Laplace mechanism with the inherited grouped $\ell_1$ sensitivity bound collapses to chance-level accuracy across noise scales, whereas the Gaussian mechanism remains trainable. This contrast isolates the source of failure: not the injected noise itself, but the geometry of the sensitivity constraint. We show that the attenuation induced by the inherited $Δ_{1,1}$ projection increases with layer width and kernel size according to a random-matrix scale $C/(\sqrt{M}+\sqrt{N})$. Replacing the $\ell_1$-type constraint with a spectral-norm constraint eliminates the collapse across datasets and architectures, but creates a fundamental obstacle: the repaired models no longer satisfy the sensitivity condition required by the standard Laplace certificate. We resolve this mismatch by deriving a dimension-free $(\varepsilon,\ δ)$-privacy guarantee for the Laplace mechanism under $\ell_2$ sensitivity through concentration of the privacy loss.
cs.CR / 18 / 2609.31318
AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents
Weida Liang, Shi Qiu, Zhun Wang, Simon Sure, Xiaoyuan Liu, Tianneng Shi, Zhaorun Chen, Wenbo Guo, Dawn Song
cs.CR · cs.AI
Abstract
AI agents combine language models with external data and tools that can modify files, call APIs, or execute code. Security failures can arise when adversarial content changes an agent's tool use or when the surrounding software contains vulnerabilities such as path traversal or command injection. We study authorized white-box pre-deployment auditing, where the auditor has access to the target repository and a controlled runtime, but successful attacks must still act through the task-defined attacker interface and be confirmed by an external verifier. We present AgentXploit, a two-role auditing system that separates repository-level attack-path discovery from runtime exploitation. The Analyzer Agent traces attacker-controlled inputs to sensitive operations and records code-supported candidate attack paths; the Exploiter Agent turns these paths into concrete attacks and revises them using runtime feedback. We also introduce AgentXploit-Bench, containing 72 reproducible vulnerabilities across 12 open-source AI-agent systems and frameworks. Across three runs, AgentXploit reaches 59.3% end-to-end success, compared with 38.4% for Codex. Under a token-budget-matched comparison, Codex reaches 46.3%. On AgentDojo, where injection points are provided, the Exploiter Agent reaches 79.2% attack success versus 52.7% for AgentVigil. These results highlight repository discovery and runtime exploitation as distinct challenges in end-to-end agent security auditing.
cs.CR / 19 / 2609.31398
Short Paper: Prefix Count Limits Can Increase First-Hit Discovery in Card Reissuance
Wasif Faisal, Suprava Saha Dibya
cs.CR
Abstract
When a card number is compromised, an attacker may search for active numbers sharing its prefix. An issuer might respond by reissuing cards from heavily populated prefixes into less populated ones. We show that this intuitive count control can backfire. For fixed search regions, exposure weights, and total activity, we derive exactly when reducing the maximum prefix count increases the chance that a budgeted search finds an active number. In matched synthetic simulations over 50,000 candidates, targeted replacement meets the count limit in every run but raises supplied-12-digit-prefix discovery relative to equal-volume random replacement in three of twelve settings. Thus a lower prefix count does not by itself certify lower enumeration risk.
cs.CR / 20 / 2609.31409
Context-Aware Functional Modeling for Android Third-Party Library Detection
Dihao Fan, Jian Zhang, Yasai Shi, Chuan Luo, Xudong Liu, Yang Liu, Chunming Hu, Xu Wang
cs.CR · cs.SE
Abstract
Third-party libraries (TPLs) are widely used in Android apps, but their reuse can introduce security risks and interfere with downstream program analyses. Existing Android TPL detection approaches face two key limitations: their hand-crafted features are fragile under aggressive code transformations, and their whole-library matching strategies are ineffective when apps retain only part of a TPL. In this paper, we propose LibFan, a learning-based Android TPL detection approach based on context-aware functional modeling. It realizes this modeling through two complementary components: context-aware contrastive learning at the method level and functional partitioning at the library level. At the method level, it learns semantic representations through contrastive training while incorporating outgoing call relationships and class-level context, improving robustness to obfuscation, shrinking, and optimization. At the library level, it partitions each TPL into functionally coherent units and determines library presence using the best-matching partition, thereby accommodating partial library reuse. To evaluate LibFan, we construct a new benchmark comprising 200 apps and 46 vulnerable TPLs, with each app compiled under four transformation configurations. Under the most challenging R8 full mode, LibFan achieves F1 scores of 81.3% at the library level and 47.6% at the version level, representing relative improvements of 64.9% and 35.6% over the state of the art, respectively.
cs.CR / 21 / 2609.31485
Verifiable Randomness for Blockchain-Based Lottery Systems
Gonçalo Ferreira, André Zúquete, Paulo Bartolomeu
cs.CR · cs.CY
Abstract
As lotteries and other high-stakes decentralized applications increasingly depend on unpredictable randomness for their operations, the lack of a secure and transparent on-chain random number generator that is verifiable by all participants remains a critical open problem. Various approaches to blockchain-based random number generation have emerged over the years, each with their own strengths and limitations, and have consistently been superseded as blockchain technology evolved. This paper surveys existing approaches to on-chain randomness and proposes a new platform that builds upon the well-known commit-and-reveal scheme while directly addressing its principal vulnerability, the last revealer attack, in which the final participant can withhold their reveal in order to bias or abort the output upon seeing an unfavorable result. We further compare this solution with prior approaches and evaluate its entropy properties. The proposed architecture combines a web-based front-end with a Solidity smart contract deployed on the Polygon 2.0 blockchain. Implemented and tested on the Amoy testnet, the prototype is low-cost and simple to deploy, providing a practical, accessible proof-of-concept for verifiable on-chain randomness.
cs.CR / 22 / 2609.31494
Toward verifiably private learning from federated data
Katharine Daly, Yu Xiao, Zachary Garrett, Brett McLarnon, Jianpeng Hou, Arun Ganesh, Yanxiang Zhang, Noriyuki Takahashi, Haicheng Sun, Yuanbo Zhang, Timon Van Overveldt, Daniel Ramage
cs.CR
Abstract
Federated Learning (FL) allows devices with private data to collaborate in training a shared model. We present a next-generation FL system based on Trusted Execution Environments (TEEs) that addresses operational challenges associated with earlier systems and provides externally verifiable central Differential Privacy (DP) guarantees for the first time while offering a better privacy-utility tradeoff. In our system, devices upload data encrypted with keys managed by a TEE-hosted Key Management Service (KMS). The uploaded data is cryptographically tied to a policy limiting the set of Python programs that may later process the data in server-side TEEs. External parties may inspect public transparency logs to observe the set of workloads allowed by these policies. Our experimental results show that the new system improves device coverage and favorably shifts privacy-utility curves by enabling collected data to be integrated into the server-side workload at a schedule that optimizes DP guarantees and is unaffected by device availability. Our new system has been productionized, enabling models for the Android Keyboard (Gboard) to be trained faster and achieve better accuracy under smaller, now externally verifiable privacy budgets in comparison to models trained using the prior system.
cs.CR / 23 / 2609.31519
Fast and Secure Simultaneous Authentication of Equals for WPA3
João Ferreira, André Zúquete, Hélder Gomes
cs.CR · cs.NI
Abstract
The Simultaneous Authentication of Equals (SAE) protocol, introduced in WPA3, provides robust protection against offline dictionary attacks against a network Pre-Shared Key (PSK) and also protection of session keys from other people knowing that PSK. However, its high computational cost for the Password Element (PE) derivation makes Access Points (APs) vulnerable to CPU exhaustion Denial-of-Service (DoS) attacks. This paper proposes a resilient architectural modification to SAE. First, we introduce an asymmetric resource cost model that offloads the iterative discovery of cryptographic elements to the client, allowing the AP to maintain a fixed computational load during the handshake. To further mitigate brute-force attempts, we implement a mechanism based on a slow-path key derivation (with variable cost key derivation functions, such as PBKDF2 or Argon2), incorporating a deliberate processing delay on the supplicant. Finally, we introduce a ticket-based mechanism to facilitate efficient re-authentication for known devices, bypassing expensive exchanges while preserving system availability under adversarial conditions. Experimental results demonstrate that this architecture significantly mitigates DoS risks without compromising legitimate network access.
cs.CR / 24 / 2609.31575
Configuration, Not Conscience: A Large-Scale Empirical Study of LLM System Prompts
Constantinos Patsakis, Vasilios Argyropoulos, Efthymios Alepis
cs.CR
Abstract
Leaked system prompts are often treated as windows into the hidden values of commercial language models, yet their composition is rarely studied at scale. We analyze a merged corpus of 407 leaked, reconstructed, or officially published system prompts from 62 vendors across four community collections, identifying 29 near-duplicate clusters covering 66 files. Operational content rather than ethical statements dominates the corpus; a deliberately simple block-level classifier assigns roughly 58\% of classified words to tool/protocol and roughly 5\% to safety policy, while the strictest rule-lines guard tool use and file safety over harmful content by an 11:1 margin. Literal text transfer concentrates in a small set of cross-vendor pairs. Prompts also carry measurable maintenance debt, with version chains turning over thousands of words per release. The evidence supports treating leaked prompts as operational specifications, closer to configuration files than value statements, and treats reuse and prompt rot as engineering and supply-chain concerns. Because most documents are adversarial in origin and the detectors are deliberately simple, all magnitudes are directional; we audit the main classifier's error modes.
cs.CR / 25 / 2609.31594
From Source Code to Network Profile: Automated and Traceable MUD Profile Generation for IoT Devices
Alessandro Lotto, Abdulla R. A. Almenhali, Savio Sciancalepore, Alessandro Brighente, Mauro Conti
cs.CR · cs.SE
Abstract
The Manufacturer Usage Description (MUD) standard allows IoT manufacturers to define expected network behaviors in a MUD file. This file can be translated into enforceable access-control policies, restricting compromised devices to operate solely through manufacturer-defined communication patterns. However, practical adoption of MUD depends on profiles that are accurate, complete, and maintainable. Existing approaches use traffic-based automation but require device deployment and prolonged monitoring, capturing only behavior exercised during observation. Rare, failure-triggered, or configuration-dependent communications may remain absent, producing incomplete policies that disrupt legitimate operation and offer limited insight into the software components responsible for each rule. We present AutoMUD, a source-code-driven tool that generates traceable MUD profiles for IoT devices from their firmware and software source code. AutoMUD combines static and syntactic extraction, retrieval-grounded language-model reasoning, and deterministic validation and compilation to recover the communication behavior characterizing an IoT device and translate eligible endpoints into policy rules. By analyzing code-level evidence, AutoMUD exposes rarely exercised and conditional communication paths, links every generated rule to its source-level provenance, and preserves excluded findings with explicit reasons for review. Our evaluation on a Linux-based repository demonstrates that AutoMUD recovers complete communication behavior, consolidates validated behavior into semantic endpoint groups, and generates structurally valid MUD profiles. Through a controlled semantic fault-injection campaign, we demonstrate that AutoMUD enables analysts to detect, localize, explain, and correct propagated errors, recovering policies semantically identical to their clean counterparts.
cs.CR / 26 / 2609.31154
HyperErase: Scale-Calibrated Hypernetwork for Multi-Concept Erasure in Text-to-Image Models
Yi Sun, Xinhao Zhong, Zhiqi Zhang, Yimin Zhou, Junhao Li, Yuxia Qiao
cs.CV · cs.CR
Abstract
Recent advances in text-to-image (T2I) generation have substantially improved visual synthesis, but have also raised increasing safety concerns due to their potential to generate harmful or undesirable content. Existing concept erasure methods predominantly follow a static weight paradigm, producing a single frozen adapter that struggles to adapt to diverse prompt variations and suffers from parameter interference when scaling to multiple concepts. We propose \textbf{HyperErase}, a framework for concept erasure based on hypernetwork-driven prompt-conditioned parameter synthesis. Our approach first reframes concept erasure as prompt-conditioned parameter amortization and trains a hypernetwork to map textual descriptions to prompt-specific LoRA updates, eliminating the need for per-prompt gradient optimization or manual LoRA merging. To further improve the stability and precision of synthesized adapters, we develop a decoupled rectification strategy, which disentangles LoRA tokens into pattern and scale subspaces, applies a square-root transform to curb multiplicative over-scaling, and leverages teacher-derived canonical priors for inference-time correction. Extensive experiments across major concept categories demonstrate that HyperErase consistently improves the trade-off between erasure effectiveness, image quality, and semantic alignment, achieving performance comparable to gold-standard single-concept baselines. Furthermore, the resulting models can provide specialized LoRAs for each input prompt variation in a single forward pass without requiring gradient updates during inference. These principled and flexible framework offers a new paradigm for concept erasure in T2I models.
cs.CR / 27 / 2609.31558
Region-Level Black-Box Defense Against Stealthy Embedding-Space Backdoors in CLIP
Ahmed Abdelnaby, Mohamed Elmahallawy
cs.CV · cs.CR
Abstract
Contrastive Language--Image Pretraining (CLIP) has emerged as a dominant vision backbone due to its strong transferability and zero-shot capabilities. However, recent studies reveal a critical vulnerability: embedding-space backdoor attacks. By poisoning only a tiny fraction of image--text pairs, adversaries can implant stealthy triggers that induce targeted shifts in CLIP's joint embedding space. Unlike conventional backdoors that manipulate classifier logits, these attacks corrupt representations directly, making them highly effective under extremely low poisoning ratios and difficult to detect. Existing defenses require access to model parameters, gradients, logits, or clean validation data---assumptions that rarely hold in realistic black-box deployments. Moreover, current black-box methods struggle to accurately localize small or out-of-distribution triggers. We propose CLIPGuard, a lightweight and fully black-box defense specifically designed to mitigate embedding-space backdoors in CLIP encoders. CLIPGuard identifies malicious regions by measuring segment-wise embedding perturbations and selectively purifies only suspicious segments via semantic inpainting, preserving benign visual content and alignment quality. Extensive experiments on STL-10, ImageNet, and diverse trigger families---including BadCLIP, BadNets, blended, patch-based, and typographic attacks---demonstrate that CLIPGuard reduces attack success rates to as low as 1.05% while maintaining clean accuracy up to 86.34%, consistently outperforming existing black-box defenses, including CleanCLIP and CleanerCLIP. Our code is available https://github.com/wsu-cyber-security-lab-ai/CLIPGuard.git
cs.CR / 28 / 2609.31256
The planted tensor problem over finite fields: algorithms and cryptography
Yuxuan Liu, Youming Qiao, Gang Tang, Chuanqi Zhang
cs.DS · cs.CR
Abstract
Inspired by the planted clique problem for random graphs, we introduce the planted totally-isotropic space problem for random tensors as follows. Let $U\cong \mathbb{F}_q^n$ and $W\cong \mathbb{F}_q^m$ be finite-dimensional vector spaces over a finite field $\mathbb{F}_q$. Given $d\in \mathbb{N}$, choose a random \(d\)-dimensional subspace \(V\leq U\), and construct a random alternating bilinear map $φ:U\times U\to W$ subject to the constraint \(φ(V,V)=0\). Such a $V$ is known as a totally-isotropic space of $φ$, and the goal is to recover $V$. Building on the recent probabilistic analysis of random tensors (Pham--Qiao--Wigderson--Wigderson, \emph{in progress}), we initiate the study of the algorithmic hardness of this problem. Setting $m=\lceil n/\log n\rceil$, we show that this problem admits an average-case polynomial-time algorithm for $d\geq n/2$, by leveraging recent advances on the non-commutative rank problem. We also show that this problem admits a $q^{O(n\log n)}$-time algorithm. We carry out algorithmic experiments using polynomial-system solving. From these results, we conjecture that the planted totally-isotropic space problem for $d=\lceil n/C\rceil$ with some constant $C\geq 3$ is exponentially hard. Based on this evidence of computational hardness, we explore cryptographic applications of the planted totally-isotropic space problem and related planted tensor problems. We present private simultaneous messages and secret sharing protocols based on planted tensor problems, following the protocols based on planted subgraphs in (Abram--Beimel--Ishai--Kushilevitz--Narayanan, \emph{TCC}'23). At the same security level, the public information size of protocols based on planted subgraphs is (moderately) exponential in that of protocols based on planted tensors, while the communication costs of these protocols are polynomially related.
cs.CR / 29 / 2609.30509
Too Late to Slash: Coordinating a Risk-Free Equivocation Attack
Hao Chung, Chen-Da Liu-Zhang
cs.GT · cs.CR
Abstract
Slashing is commonly argued to secure proof-of-stake blockchains by confiscating the stake of misbehaving validators. The usual justification is that, without slashing, validators can solicit an equivocation attack by signing conflicting blocks: if enough others join, the attack succeeds; otherwise, the attempt incurs no loss. Slashing is intended to make such attempts costly and thereby guarantee the security of applications whose economic value is comparable to the bonded stake. This rationale, however, rests on heuristic arguments rather than a formal game-theoretic guarantee. We challenge this rationale by constructing a risk-free coordination protocol for rational validators under algorithmic slashing. We show that the prescribed strategy profile, in which rational validators solicit other validators to equivocate, constitutes an ex post Nash equilibrium, even when validators do not know in advance how many others will participate. The equilibrium holds for any gain $ε>0$ from successful equivocation, however small relative to the bonded stake. Thus, slashing alone does not guarantee economic security proportional to the value of bonded stake. Our results expose a fundamental limitation of algorithmic slashing. Whereas conventional collateral arrangements in many real-world scenarios can rely on external enforcement, for example through courts, algorithmic slashing depends on the same consensus process that the attackers control. These findings call for a formal analysis of slashing's security, rather than overly simplistic arguments drawn from financial systems with independent enforcement.
cs.CR / 30 / 2609.31379
How Much Must a Private Mempool Hide? Exact Leakage Thresholds for Sandwich Attacks
Tingyi Lin, Jiazhuo Li, Ruoran Lai
cs.GT · cs.CR · q-fin.TR
Abstract
Private and encrypted mempools hide pending transactions to stop sandwich attacks and other forms of maximal extractable value (MEV), but what they hide is rarely everything: a transaction's pair, direction, and a coarse range for its size can still leak. How much leakage makes sandwiching pay? We answer exactly for a fee-free constant-product automated market maker, the pricing rule behind Uniswap v2. Traders observe an interval containing the victim's size and bid in a first-price auction for the right to sandwich it, and the winning front-run must keep the victim's trade executable at every size in the interval. The answer turns on the smallest size consistent with the leak. It alone determines the feasible front-runs, the largest feasible front-run is optimal for pointwise, expected, and worst-case profit alike, and the guaranteed profit has a closed form. When execution is costly, a privacy layer that wants to rule out sandwiches profitable at every consistent size may therefore reveal anything about the size except a lower bound above an explicit threshold; the upper end of the range is irrelevant. With two or more symmetric traders, every pure-strategy perfect Bayesian equilibrium of the auction hands the entire expected net rent to the auctioneer. If the direction is hidden too, no non-contingent first leg front-runs both possible directions, while post-trade arbitrage can survive even perfect pre-trade hiding.
cs.CR / 31 / 2609.31032
TempQ-Jail: Query-Constrained Candidate Ranking for Text-to-Video Jailbreak Attacks
Tianmeng Fang, Jiancheng Wang, Chen Wang, Liming Wang, Wei Wang, Jiayang Liu, Xiaochun Cao
cs.MM · cs.CR · cs.CV
Abstract
Existing text-to-video (T2V) jailbreak methods mainly seek more effective or stealthier attack candidates. In guarded T2V systems, however, video generation and security evaluation are costly, so an attacker often cannot test a large candidate pool. We therefore formulate T2V jailbreak as a query-constrained candidate allocation and ranking problem and propose TempQ-Jail. The method combines heterogeneous attack mechanisms to expand candidate coverage, estimates each candidate's end-to-end attack value from security-gate passage, dangerous visual generation, preservation of the original intent, and temporal validity, and ranks candidates so that high-value attacks appear early in a limited query trajectory. We evaluate TempQ-Jail on CogVideoX-5B using 70 common viable intents derived from T2VSafetyBench and compare it with six representative T2V jailbreak methods under a unified protocol. TempQ-Jail achieves TP-ASR@5 and TP-ASR@10 of 48.9% and 65.4%, improving over the strongest baselines by 4.6 and 4.0 percentage points, respectively. It also obtains the highest AUC-TP (0.469) and the lowest AvgQ (6.3). Analyses of query trajectories, candidate allocation, failure attribution, and ablations show that TempQ-Jail more effectively identifies and prioritises candidates with complete attack potential under limited query budgets.
cs.CR / 32 / 2609.30581
Encryptability As a Coordinate Choice: Depth-One Homomorphic Federated Learning of Quantum Neural Networks
Marcel Mordarski, Nathan Mani, Arshad Patel, William Knottenbelt, Roberto Bondesan
quant-ph · cs.CR · cs.DC · cs.LG
Abstract
Encrypted training relies on keeping server-side updates low-degree. This constraint traditionally excludes models whose weights inhabit a compact Lie group (notably variational quantum circuits, where every trainable weight is an $\mathrm{SU(2)}$ rotation). Expressed in Euler angles or discrete alphabets, these updates appear transcendental, historically demanding prohibitive costs: one client--server round per gate, or upwards of $25{,}000$ operations per weight. This penalty is strictly an artefact of coordinates. In the unit-quaternion (spin) chart, group composition is exactly bilinear (degree two, with coefficients in $\{-1,0,+1\}$). Consequently, encrypted rotation updates cost one multiplicative level and federated averaging costs zero in any levelled homomorphic scheme, completely eliminating bootstrapping. This implementation-independent algebraic property is confirmed across two cryptographic backends, introducing only $0.0$ and $-2.0\times10^{-12}$ rad of aggregation error. Leveraging this reduction yields a non-interactive protocol for encrypted federated training of hybrid quantum--classical networks. It includes correctness proofs for aggregation and sign handling, plus a compilation lemma proving parameterised entanglers add only constant-factor overhead without altering the depth class. Empirically, a paired five-seed study confirms zero measurable utility tax ($Δ=+9\times10^{-6}$ MSE, $p=0.92$), and a noise-budget ablation falsifies the hypothesis that encryption noise regularises. These convergence trends replicate across datasets and scale to $20$ clients. Finally, hardware validation on a $156$-qubit processor achieves $0.9918$ fidelity against a $0.99957$ unencrypted control.