← Back to Index
Daily Research Digest

arXiv Papers

2026-09-24
374
Papers
9
Categories
73
Translated
收藏清单 0
精选 · Favorites
73
cs.AI / 1 / 2609.26986
Same evidence, different judgments: Evidence noncommutative in vision/speech-text conflicts
同样的证据,不同的判断:视觉/语音-文本冲突中的证据非交换性
Zhuoyun Li, Boxuan Wang, Xiaowei Huang, Yi Dong
cs.AI
large language model
大语言模型相关
Abstract
For multimodal large language models, when images or speech conflict with accompanying text, measured text reliance can entangle modality preference with evidence position. Earlier studies of text bias often used a fixed evidence order or moved task instructions with the evidence, leaving the contribution of order unclear. In this paper, we use a paired comparison that keeps the instructions and evidence content fixed and swaps only the positions of the two sources to quantify this potential influence. Across vision and speech models, placing an image or recording after conflicting text consistently shifts answers toward its content. We also revisit previous studies and analyze why their experimental settings can lead to misleading conclusions. These findings reveal cross-modal evidence noncommutativity: the same evidence can lead to different judgments when its order changes, and placing perceptual evidence later can increase the model's reliance on its content.
Chinese Translation
对于多模态大语言模型而言,当图像或语音与伴随文本相冲突时,所测得的文本依赖程度可能会将模态偏好与证据位置纠缠在一起。早期关于文本偏差的研究常常使用固定的证据顺序,或者随证据一起移动任务指令,从而使顺序的贡献变得不清晰。在本文中,我们使用一种配对比较:保持指令和证据内容固定,仅交换两个来源的位置,以量化这种潜在影响。在视觉和语音模型中,将图像或录音置于相冲突的文本之后,会一致地将答案转向其内容。我们还重新审视了先前的研究,并分析其实验设置为何可能导致误导性结论。这些发现揭示了跨模态证据的非交换性:当证据的顺序改变时,同样的证据可能导致不同的判断,而将感知证据置于更后位置会增加模型对其内容的依赖。
cs.AI / 2 / 2609.27150
Do We Need Complex Topology Control? Distinct-Peer Random Routing Improves Cost-Efficiency in Sparse Multi-Agent Debate
我们需要复杂的拓扑控制吗?异质同伴随机路由提升了稀疏多智能体辩论中的成本效率
Boxuan Wang, Zhuoyun Li, Xiaowei Huang, Yi Dong
cs.AI
large language model
大语言模型相关
Abstract
Multi-agent debate (MAD) has emerged as a promising paradigm for improving the reasoning accuracy of large language models (LLMs) through iterative peer interaction. Communication topology plays a central role in this process, motivating increasingly sophisticated mechanisms that learn, adapt, or dynamically reconfigure agent interactions to improve accuracy or reasoning reliability. Meanwhile, prior studies suggest that much simpler sparse communication can already achieve competitive performance at substantially lower cost. In this work, we take a closer look at sparse MAD and ask whether complex topology control is actually necessary to improve collective reasoning. We find that a simple random-without-replacement routing policy, which lets each agent debate with two distinct and newly sampled peers at every round, provides a surprisingly strong baseline and consistently improves the accuracy-cost trade-off of sparse MAD. Building on this observation, we further study deliberation stopping and show that lightweight stopping can substantially reduce inference cost while preserving competitive accuracy. Our results suggest that sophisticated topology control such as learned topology adaption should be evaluated against strong simple routing and stopping baselines before its additional complexity is justified.
Chinese Translation
多智能体辩论(MAD)已成为一种有前景的范式,它通过迭代式的同伴交互来提升大语言模型(LLMs)的推理准确率。通信拓扑在这一过程中起着核心作用,这推动了日益复杂的机制,它们学习、适应或动态重构智能体交互,以提高准确率或推理可靠性。与此同时,已有研究表明,简单得多的稀疏通信已经能够以显著更低的成本达到有竞争力的性能。在本工作中,我们更仔细地考察稀疏 MAD,并追问复杂拓扑控制是否真的是改进集体推理所必需的。我们发现,一种简单的无放回随机路由策略——即让每个智能体在每一轮与两个互不相同且新采样的同伴进行辩论——提供了一个出人意料地强的基线,并持续改善稀疏 MAD 的准确率–成本权衡。基于这一观察,我们进一步研究审议停止,并表明轻量级停止可以显著降低推理成本,同时保持有竞争力的准确率。我们的结果表明,诸如学习式拓扑自适应之类的复杂拓扑控制,在其额外复杂性被证明合理之前,应当与强力的简单路由和停止基线进行评估比较。
cs.AI / 3 / 2609.27284
Hunyuan-A13B Technical Report
Hunyuan-A13B 技术报告
Tencent Hunyuan Team, Ao Liu, Botong Zhou, Can Xu, Chayse Zhou, ChenChen Zhang, Chengcheng Xu, Chenhao Wang, Decheng Wu, Dengpeng Wu, Dian Jiao, Dong Du, Dong Wang, Feng Zhang, Fengzong Lian, Guanghui Xu, Guanwei Zhang, Hai Wang, Haipeng Luo, Han Hu, Huilin Xu, Jiajia Wu, Jianchen Zhu, Jianfeng Yan, Jiaqi Zhu, Jihong Zhang, Jinbao Xue, Jun Xia, Junqiang Zheng, Kai Liu, Kai Zhang, Kai Zheng, Kejiao Li, Keyao Wang, Lan Jiang, Lixin Liu, Lulu Wu, Mengyuan Huang, Peijie Yu, Peiqi Wang, Qian Wang, Qianbiao Xiang, Qibin Liu, Qingfeng Sun, Richard Guo, Ruobing Xie, Saiyong Yang, Shaohua Chen, Shihui Hu, Shuai Li, Shuaipeng Li, Shuang Chen, Suncong Zheng, Tao Yang, Tian Zhang, Tinghao Yu, Weidong Han, Weijie Liu, Weijin Zhou, Weikang Wang, Wesleye Chen, Xiao Feng, Xiaoqin Ren, Xingwu Sun, Xiong Kuang, Xuemeng Huang, Xun Cao, Yanfeng Chen, Yang Du, Zhen Yang, Yangyu Tao, Yaping Deng, Yi Shen, Yigeng Hong, Yiqi Chen
cs.AI
large language model
大语言模型相关
Abstract
We present Hunyuan-A13B, an open-source large language model based on a Mixture-of-Experts architecture. It contains 80 billion total parameters but activates only 13 billion during inference, balancing model capability, computational efficiency, and deployment cost. The model is pretrained on a rigorously filtered 20T-token corpus with enhanced STEM data curation, improving factual reliability and reasoning ability. High-quality supervised fine-tuning and large-scale reinforcement learning further enhance its overall performance. Hunyuan-A13B also introduces a dual-mode Chain-of-Thought framework that adapts reasoning depth to task complexity: fast thinking for routine queries and slow thinking for complex, multi-step problems. Evaluations show competitive performance across mathematics, science, programming, general language understanding, and agent tasks, often approaching that of much larger models. Its high inference throughput makes it suitable for latency-sensitive applications. We release Hunyuan-A13B to support open research and practical LLM deployment.
Chinese Translation
我们提出 Hunyuan-A13B,一个基于混合专家(Mixture-of-Experts)架构的开源大语言模型。它包含总计800亿参数,但在推理期间仅激活130亿参数,从而在模型能力、计算效率和部署成本之间取得平衡。该模型在经过严格过滤的20T token语料库上进行预训练,并增强了STEM数据整理,从而提升了事实可靠性和推理能力。高质量的监督微调和大规模强化学习进一步提升了其整体性能。Hunyuan-A13B 还引入了双模式思维链(Chain-of-Thought)框架,使推理深度适应任务复杂度:对常规查询进行快速思考,对复杂的多步问题进行慢速思考。评估显示,其在数学、科学、编程、通用语言理解和智能体任务上具有有竞争力的性能,常常接近规模大得多的模型。其高推理吞吐量使其适用于对延迟敏感的应用。我们发布 Hunyuan-A13B,以支持开放研究和实际的LLM部署。
cs.AI / 4 / 2609.27336
CART: Closed-Loop Adaptive Red Teaming for Large Language Models
CART:面向大语言模型的闭环自适应红队测试
Dongdong Zhang, Tengchao Lv, Yilin Jia, Yuzhong Zhao, Yupan Huang, Wenshan Wu, Xiangyang Zhou, Shaohan Huang, Nan Yang, Li Dong, Lei Cui, Furu Wei
cs.AI
large language model
大语言模型相关
Abstract
Automated red teaming often replays a fixed set of prompts, which measures known risks but cannot learn from failures found during testing. We present CART (Closed-Loop Adaptive Red Teaming), a framework that uses each result to guide what it tests next. CART begins with broad risk coverage, follows weaknesses that emerge, keeps new probes diverse, and records the evidence and source of every finding. It separates the Challenger that creates tests, the Target being tested, which may be a text-only model or a bounded tool-using agent, and the Judge that evaluates the results, allowing these roles to be studied independently. Across three evaluation families (Frontier, JAH, and Agentic), CART discovers more failures and higher average risk than static seed replay for every Target with an available baseline. The gains extend to tool-mediated agent tests, suggesting that contextual adaptation can reveal weaknesses that direct prompt replay does not exercise. These results describe what the test policies discover, not how often failures occur in real deployments. We also find that Challenger-Judge choices affect the evidence uncovered, highlighting the need for role separation and independent review. Overall, CART turns red teaming from a one-time checklist into a continuous, adaptive, and auditable search for model and agent weaknesses.
Chinese Translation
自动化红队测试通常重放一组固定的提示,这虽能衡量已知风险,却无法从测试过程中发现的失败中学习。我们提出 CART(闭环自适应红队测试),这是一个利用每一次结果来指导其下一步测试内容的框架。CART 从广泛的风险覆盖开始,追踪涌现出的弱点,保持新探测的多样性,并记录每一项发现的证据与来源。它将创建测试的挑战者(Challenger)、被测试的目标(Target,可能是纯文本模型,也可能是受限的工具使用智能体)以及评估结果的裁判(Judge)分离开来,从而使这些角色可以被独立地研究。在三个评测族(Frontier、JAH 和 Agentic)中,对于每一个有可用基线的 Target,CART 发现的失败都多于静态种子重放,平均风险也更高。这些增益还延伸到了经由工具中介的智能体测试,表明上下文自适应能够揭示直接提示重放所无法触及的弱点。这些结果描述的是测试策略发现了什么,而非失败在真实部署中出现的频率。我们还发现,挑战者—裁判的选择会影响所揭示的证据,凸显了角色分离与独立审查的必要性。总体而言,CART 将红队测试从一次性的检查清单转变为一种持续、自适应且可审计的、针对模型与智能体弱点的搜索。
cs.AI / 5 / 2609.27349
MolDesignBench: Evaluating LLM-based Agent for Scenario-grounded Molecular Design
MolDesignBench:评估基于LLM的智能体用于场景化分子设计
Yongjun Jeong, Hanbum Ko, Ye Rin Kim, Chanhui Lee, Rodrigo Hormazabal, Jaewan Lee, Sehui Han, Sungbin Lim, Sungwoong Kim
cs.AI · cs.LG
large language model
大语言模型相关
Abstract
Real-world molecular design remains challenging for large language model (LLM)-based agents. It requires them to interpret design contexts, satisfy multiple constraints, identify infeasible specifications, and reason over multi-step tool outputs. Existing benchmarks do not capture this complexity, focusing instead on explicit and narrow constraints, only feasible problems, and single-path solutions. To address this gap, we propose MolDesignBench, a scenario-grounded benchmark that more closely reflects real-world molecular design for evaluating tool-augmented LLM agents. MolDesignBench comprises 2K generation and optimization instances that combine implicit requirements embedded in design narratives with explicit property and functional-group constraints, including infeasible cases, and require the effective use of 17 specialized chemistry tools. Experiments across diverse frontier LLMs reveal low success rates--with the best achieving only $\sim43$\%--and frequent failures in implicit-constraint reasoning, infeasibility detection, and tool reasoning. The corresponding fine-grained failure-mode analysis identifies implicit constraint interpretation and infeasibility detection as the primary bottlenecks, establishing MolDesignBench as a rigorous testbed to guide future research on chemical agents. The benchmark, tool interface, and evaluation code are publicly available.
Chinese Translation
真实世界的分子设计对于基于大语言模型(LLM)的智能体而言仍然具有挑战性。它要求它们解读设计情境、满足多重约束、识别不可行的规格,并对多步工具输出进行推理。现有的基准测试并未捕捉到这种复杂性,而是聚焦于显式且狭窄的约束、仅可行的问题以及单一路径的解决方案。为填补这一空白,我们提出了 MolDesignBench,这是一个更贴近真实世界分子设计的场景化基准,用于评估工具增强的 LLM 智能体。MolDesignBench 包含 2K 个生成与优化实例,这些实例将嵌入设计叙述中的隐式需求与显式的性质和官能团约束相结合,其中包括不可行案例,并要求有效使用 17 种专业化学工具。跨多种前沿 LLM 的实验揭示了较低的成功率——其中最佳者仅达到 $\sim43$\%——以及在隐式约束推理、不可行性检测和工具推理方面的频繁失败。相应的细粒度失败模式分析将隐式约束解释和不可行性检测确定为主要瓶颈,从而将 MolDesignBench 确立为指导未来化学智能体研究的严格测试平台。该基准、工具接口和评估代码均已公开可用。
cs.AI / 6 / 2609.27517
Not What You Meant: Can LLMs Follow a Specified Negation Semantics?
并非你所意指:LLM 能否遵循指定的否定语义?
Qiming Bao, Agnieszka Mensfelt, Michael J. Witbrock, Kostas Stathis
cs.AI
large language model
大语言模型相关
Abstract
Negation does not carry a uniform interpretation across domains. In legal, regulatory, and medical reasoning, the intended interpretation depends on the reading in force -- open- versus closed-world, two- versus three-valued, and credulous versus skeptical. We study which reading of negation large language models adopt by default and whether they can override that preference when a different reading is explicitly specified. To this end, we introduce NAFBench, a procedural generator of solver-certified instances spanning four semantic viewpoints: SLDNF, well-founded semantics (WFS), and credulous and skeptical reasoning under stable-model semantics. The generator emits ground normal logic programs with controlled depth, width, and cycle structure. Each program is solved under all four viewpoints using SWI-Prolog, a well-founded semantics solver, and clingo, yielding up to four divergent labels. The programs are then verbalized into natural language under multiple framings and rule orderings that leave the answer invariant. The results expose a consistent gap. Across open-source models, following a specified negation semantics remains unsolved: the strongest models score 59--74% across the four semantic viewpoints, while the weakest score 31--67%. All models are order-sensitive on more than half of logically identical rule shufflings, while the two weaker models frequently overcommit on well-founded "undefined." Two frontier models reach 100% on the main fixed-complexity evaluation set, and a third, o4-mini, is near-perfect, falling only to 81% on well-founded "undefined." Delegating reasoning to a solver, fine-tuning on certified traces, or forcing an explicit three-valued verdict each partly closes the gap.
Chinese Translation
否定在不同领域中并不具有统一的解释。在法律、监管和医学推理中,预期的解释取决于所采用的读法——开放世界与封闭世界、二值与三值,以及轻信式与怀疑式。我们研究大语言模型默认采用哪种否定读法,以及当明确指定一种不同读法时,它们能否覆盖该偏好。为此,我们引入 NAFBench,一个过程式生成器,用于生成经求解器认证的实例,涵盖四种语义视角:SLDNF、良基语义(WFS),以及稳定模型语义下的轻信式推理和怀疑式推理。该生成器生成具有受控深度、宽度和循环结构的基正规逻辑程序。每个程序都使用 SWI-Prolog、一个良基语义求解器和 clingo 在所有四种视角下求解,从而得到至多四个不同的标签。随后,这些程序在保持答案不变的多种表述框架和规则排序下被转述为自然语言。结果揭示出一致的差距。在开源模型中,遵循指定的否定语义仍未得到解决:最强模型在四种语义视角上的得分为 59--74%,而最弱模型的得分为 31--67%。所有模型在超过一半的逻辑上相同的规则打乱中都具有顺序敏感性,而两个较弱的模型经常在良基“未定义”上过度确定。两个前沿模型在主要的固定复杂度评估集上达到 100%,而第三个模型 o4-mini 近乎完美,仅在良基“未定义”上降至 81%。将推理委托给求解器、在经认证的轨迹上进行微调,或强制给出显式的三值判定,每一种都能部分弥合这一差距。
cs.AI / 7 / 2609.27749
Evaluation of pre-trained models for pedagogical assessment of novel AI-assisted educational questions
用于新型 AI 辅助教育问题教学评估的预训练模型评估
Michael Lawrence Castanares, Princess Ventures, Allan Tan
cs.AI · cs.CL · cs.LG
large language model
大语言模型相关
Abstract
The surge in AI-assisted generation of educational materials has outpaced our capacity to validate their pedagogical quality. Automated evaluation using Bloom Classifier models is a promising approach to assess educational materials at scale. These models show high accuracy within-distribution dataset (IID Dataset). However, applying the same models to new out-of-distribution (OOD) datasets such as AI-assisted generated questions could show performance degradation. To identify robust classifiers under dataset shift, we evaluated traditional Machine Learning (ML), transformer, and Large Language models on the Bloom level classification task. We also explored feature-engineering strategies incorporating NLP metrics, appending the learning objectives as part of the input, and text splicing to stabilize OOD performance. Our baseline tests show that TFPOS-IDF ML models perform poorly on OOD (Macro F1-score 0.48) compared to BERT (0.55) and LLMs (0.79). Text splicing improved macro F1-score performance of ML and BERT models (0.59 and 0.62, respectively). Appending the learning objectives with the input increased model performance on specific dataset. Model retraining provided the largest improvement across models and datasets. Overall, these findings highlight the trade-off on the use of pre-trained models with novel AI-assisted educational questions and how strategic feature enhancements help address loss in performance.
Chinese Translation
AI 辅助生成教育材料的激增已经超出了我们验证其教学质量的能力。使用布鲁姆分类器模型的自动化评估是一种有望大规模评估教育材料的方法。这些模型在分布内数据集(IID 数据集)上表现出高准确率。然而,将相同的模型应用于新的分布外(OOD)数据集,例如 AI 辅助生成的问题,可能会表现出性能下降。为了识别数据集偏移下稳健的分类器,我们在布鲁姆层级分类任务上评估了传统机器学习(ML)、Transformer 和大语言模型。我们还探索了纳入 NLP 指标、将学习目标作为输入的一部分附加,以及文本拼接等特征工程策略,以稳定 OOD 性能。我们的基线测试表明,与 BERT(0.55)和大语言模型(LLMs)(0.79)相比,TFPOS-IDF ML 模型在 OOD 上表现较差(宏 F1 分数 0.48)。文本拼接提高了 ML 和 BERT 模型的宏 F1 分数性能(分别为 0.59 和 0.62)。将学习目标与输入一起附加提高了模型在特定数据集上的性能。模型重新训练在跨模型和数据集方面提供了最大的改进。总体而言,这些发现凸显了在使用预训练模型处理新型 AI 辅助教育问题时的权衡,以及战略性特征增强如何帮助解决性能损失。
cs.AI / 8 / 2609.28197
PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety
PASTABench:面向智能体安全的序列轨迹主动评估
Jiapeng Sun, Yujin Zhou, Han Zhu, Pengcheng Wen, Jiayi Zhou, Sirui Han, Yike Guo
cs.AI · cs.CL
large language model
大语言模型相关
Abstract
As Large Language Models (LLMs) evolve into autonomous agents that alter real-world states, ensuring operational safety across multi-step workflows has become a critical challenge. While recent work has moved beyond single-turn evaluation toward multi-turn paradigms, key limitations persist: step-level methods treat actions in isolation, missing how risks accumulate, while trajectory-level evaluations operate post-hoc, offering no opportunity for timely intervention. To address these limitations, we formalize Decoupled Proactive Safety Monitoring along three dimensions: whether to intervene, when to intervene, and what the risk is. We introduce PASTABench, a benchmark of 1,139 multi-turn trajectories spanning 5 risk categories and 13 subcategories. We further propose the Optimal Intervention Window (OIW), anchored by annotated Earliest-Signal and Trigger turns, to quantify intervention timeliness. Evaluation of 16 LLMs reveals that proactive intervention remains largely unsolved, with the best model achieving only 40.74% optimal-timing interventions. Fine-grained diagnosis further uncovers pervasive lexical overfitting: competitive safety scores of smaller models mask keyword hypersensitivity rather than genuine risk comprehension, as their proactive capability largely collapses once hazard vocabulary is neutralized.
Chinese Translation
随着大型语言模型(LLMs)演变为能够改变现实世界状态的自主智能体,确保多步工作流中的运行安全已成为一项关键挑战。尽管近期工作已从单轮评估转向多轮范式,但关键局限仍然存在:步骤级方法孤立地处理动作,忽视了风险如何累积,而轨迹级评估则以事后方式进行,无法提供及时干预的机会。为解决这些局限,我们将解耦式主动安全监控沿三个维度形式化:是否干预、何时干预以及风险是什么。我们提出了 PASTABench,一个包含 1,139 条多轮轨迹的基准,涵盖 5 个风险类别和 13 个子类别。我们进一步提出了最优干预窗口(OIW),其以标注的最早信号轮次和触发轮次为锚点,用于量化干预的及时性。对 16 个 LLMs 的评估显示,主动干预在很大程度上仍未得到解决,最佳模型仅实现了 40.74% 的最优时机干预。细粒度诊断进一步揭示了普遍存在的词汇过拟合:较小模型具有竞争力的安全得分掩盖的是关键词超敏反应,而非真正的风险理解,因为一旦危险词汇被中和,它们的主动能力就会大幅崩溃。
cs.AI / 9 / 2609.28322
Learning the Cost of Reliable Inference
学习可靠推理的成本
Dimitrios Rontogiannis, Ander Artola Velasco, Manuel Gomez Rodriguez
cs.AI · cs.GT · cs.LG
large language model
大语言模型相关
Abstract
Benchmarking and routing platforms increasingly act as intermediaries connecting large language model providers with end-users. However, providers on these platforms typically use a fixed price per token, preventing users from achieving the most competitive price for their tasks. % workloads. In this work, we design a procurement platform where token prices for each task are driven by provider competition, enabling users to secure competitive pricing for guaranteed quality levels. To this end, the platform sequentially routes queries via a reverse second-price auction that incentivizes model providers to truthfully bid their best estimate of the average cost to serve a user's query. As it routes queries, the platform learns the quality offered by each provider and progressively routes queries to the most cost-competitive provider among those meeting a desired quality threshold. To validate our design, we conduct experiments with multiple LLMs from the \texttt{Llama} and \texttt{Qwen} families on popular mathematical reasoning and question-answering benchmarks. The results show that the pricing margin of the most cost-competitive provider on our platform varies significantly---from $10\%$ to $71\%$---depending on the task and quality threshold. This suggests a substantial inefficiency in the current fixed-price market, and it demonstrates that our platform may enable users to capture maximum savings whenever competitive market conditions permit.
Chinese Translation
基准测试与路由平台正日益充当连接大语言模型提供商与终端用户的中介。然而,这些平台上的提供商通常采用每 token 固定价格,使用户无法为其任务获得最具竞争力的价格。 % 工作负载。在这项工作中,我们设计了一个采购平台,其中每个任务的 token 价格由提供商竞争驱动,使用户能够为有保证的质量水平获得有竞争力的定价。为此,该平台通过反向第二价格拍卖依次路由查询,激励模型提供商如实将其对服务用户查询的平均成本的最佳估计作为出价。在路由查询的过程中,该平台学习每个提供商所提供的质量,并逐步将查询路由到满足期望质量阈值的提供商中成本最具竞争力的提供商。为验证我们的设计,我们在流行的数学推理和问答基准上,使用来自 \texttt{Llama} 和 \texttt{Qwen} 系列的多个大语言模型进行实验。结果表明,我们平台上成本最具竞争力的提供商的定价利润率差异显著——从 $10\%$ 到 $71\%$——具体取决于任务和质量阈值。这表明当前固定价格市场存在显著的低效率,并表明在竞争性市场条件允许时,我们的平台可以使用户获得最大节省。
cs.AI / 10 / 2609.28470
StudentBench: AI and human tutoring yield equivalent GRE learning gains
StudentBench:AI 与人类辅导产生同等的 GRE 学习增益
Curtis Northcutt, Inaara Hasmani, Kevin Feng, Trevor Khangi, Andreas Plesner, Jonas Mueller
cs.AI · cs.CY
large language model
大语言模型相关
Abstract
Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). The StudentBench platform is freely available at https://studentbench.org.
Chinese Translation
人工智能为增强人类能力提供了前所未有的机会,但前沿进展主要集中于提升模型能力。我们推出 StudentBench,一套 AI 教学评估套件以及一个公共平台,该平台支持大规模数据收集,已包含超过 175,000 条学生-AI 消息,用于研究大型语言模型(LLMs)是否能产生与人类辅导相当的学习增益。使用 StudentBench,我们在 2,383 名接受 AI 辅导、人类辅导或无辅导的人类参与者中,测量了 GRE 数量推理与语文题目的学习增益。我们确立,就 GRE 学习增益而言,AI 辅导在统计上等同于专家人类辅导(p = .015),并且在七个 GRE 领域中的五个领域里,表现最好的 AI 辅导教师平均而言超过了人类辅导教师。在第二项研究中,专家人类辅导教师通过 2,028 次成对评分量规评估,比较了 LLM 生成的课程计划和练习题。这两项研究共同在以下方面清晰地区分了 AI 辅导教师:(1)课程规划,(2)练习题创建,(3)对话式教学法,(4)成本,以及(5)参与度。令人惊讶的是,一个 AI 辅导教师以低 918 倍的成本(每获得一个百分点,AI 为 0.0052 美元,人类为 4.81 美元)实现了与人类辅导相当的学习增益(p = .044)。对于 GRE 数量推理会话,更快的 AI 回复与更多的学生消息相关,更多的消息与更多的正确练习相关,而更多的正确练习与更大的学习增益相关(所有 p < .002)。StudentBench 平台可在 https://studentbench.org 免费获取。
cs.AR / 11 / 2609.27453
Implementation and Evaluation of BitNet Inference on a CGLA by Signed-Int4 Instructions
基于有符号 Int4 指令在 CGLA 上实现与评估 BitNet 推理
Takuto Ando, Yasuhiko Nakashima
cs.AR
large language model
大语言模型相关
Abstract
Large language model (LLM) inference transfers model weights and activations for every generated token, making memory traffic and its energy cost part of the decode path. BitNet b1.58 represents its low-bit weights by ternary values and uses integer activations. However, this arithmetic does not match conventional int8 or floating-point general matrix multiplication, and existing BitNet accelerators implement it in specialized datapaths. We instead map this operation to a CPU-Grounded Linear Array (CGLA), a programmable ASIC with explicit direct memory access, local memories, and reusable compiler-visible integer lanes. The mapping adds OP_SMA4 as a reusable signed-int4 multiply-accumulate instruction rather than a BitNet-only datapath. Each ternary weight occupies one signed 4-bit lane. Each int8 activation is split into two signed-int4 fragments and reconstructed by shift-and-add. Frequency scaling of the 145 MHz FPGA measurement to an 840 MHz 28 nm CGLA achieved 0.390 ns per signed-int4 product. We showed that CGLA-offloaded BitNet C++ execution measures 2.52 tokens/s.
Chinese Translation
大语言模型(LLM)推理会为每一个生成的 token 传输模型权重和激活值,从而使内存流量及其能耗成为解码路径的一部分。BitNet b1.58 用三值来表示其低位权重,并使用整数激活值。然而,这种算术运算与传统的 int8 或浮点通用矩阵乘法并不匹配,而现有的 BitNet 加速器是在专用数据通路中实现该运算的。我们转而将该运算映射到 CPU-Grounded Linear Array(CGLA)上,这是一种可编程 ASIC,具有显式的直接内存访问、本地存储器以及可复用的、对编译器可见的整数通道。该映射添加了 OP_SMA4 作为一条可复用的有符号 int4 乘累加指令,而不是一条仅用于 BitNet 的数据通路。每个三值权重占用一个有符号 4 位通道。每个 int8 激活值被拆分为两个有符号 int4 片段,并通过移位相加进行重构。将 145 MHz 的 FPGA 测量结果按频率缩放至 840 MHz 的 28 nm CGLA,实现了每个有符号 int4 乘积 0.390 ns。我们表明,卸载到 CGLA 的 BitNet C++ 执行测得 2.52 tokens/s。
cs.CL / 12 / 2609.26913
COMED: The Missing Middle Between Routing and Collaboration in Multi-LLM Inference
COMED:多 LLM 推理中路由与协作之间缺失的中间地带
Norah Alballa, Wenxuan Zhang, Salma Kharrat, Fares Fourati, Zafar Ayyub Qazi, Mohamed Elhoseiny, Marco Canini
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
No single Large Language Model (LLM) is uniformly reliable across queries, motivating multi-model inference systems that either route among models or combine their outputs. However, routing stops after selecting an initial model, while dense collaboration invokes peers on every query. We show that collaboration is non-monotonic: peers can recover failures that no model solves alone, but can also corrupt initially correct answers. We introduce COMED (Controlled Model Escalation for Multi-LLM Deliberation), a post-anchor controller for selective cross-model collaboration. COMED uses anchor self-consistency, router margin, and a lightweight peer probe to accept confident answers, verify ambiguous cases, and escalate only when collaboration is likely beneficial. We formalize this trade-off with a rescue-harm decomposition showing that selective collaboration improves when rescued errors outweigh collaboration-induced harms. Across medical, scientific, and general reasoning benchmarks, COMED improves fixed and routed anchors in all 16 open-weight settings, with gains up to +10.7 percentage points on MedQA while invoking fewer models and using fewer decoded tokens than dense collaboration. On HLE with frontier models, COMED improves GPT-5.5 from 23.1% to 28.1%, outperforming dense collaboration and achieving the best results.
Chinese Translation
没有任何单一大型语言模型(LLM)能在所有查询上都保持一致的可靠性,这推动了多模型推理系统的出现,此类系统要么在模型之间进行路由,要么将其输出进行组合。然而,路由在选定初始模型后便停止,而密集协作则在每次查询时都调用对等模型。我们表明,协作是非单调的:对等模型能够挽回没有任何模型能独自解决的失败,但也可能破坏最初正确的答案。我们提出 COMED(面向多 LLM 审议的可控模型升级,Controlled Model Escalation for Multi-LLM Deliberation),这是一种用于选择性跨模型协作的锚点后控制器。COMED 利用锚点自一致性、路由器间隔以及一个轻量级对等探测来接受高置信度的答案、验证模糊的情形,并且仅在协作可能有益时才进行升级。我们通过一种挽救—损害分解来形式化这一权衡,表明当被挽救的错误超过协作所导致的损害时,选择性协作会带来改进。在医学、科学和通用推理基准上,COMED 在所有 16 种开放权重设置中均提升了固定锚点和路由锚点的表现,在 MedQA 上增益最高达 +10.7 个百分点,同时所调用的模型数量和所使用的解码 token 数量都少于密集协作。在前沿模型上的 HLE 中,COMED 将 GPT-5.5 从 23.1% 提升至 28.1%,优于密集协作并取得最佳结果。
cs.CL / 13 / 2609.26926
Experts Rise Where LLMs Disagree: Using Cross-Model Disagreement to Target Expert Effort in LLM Codebook Revision for Large-Scale Annotation
专家在LLM意见分歧之处崛起:利用跨模型分歧在大规模标注的LLM编码手册修订中定向专家投入
Zeyu He, Zhuqian Zhou, Kirk Vanacore, Rene F. Kizilcec, Ting-Hao 'Kenneth' Huang
cs.CL · cs.AI · cs.HC · cs.LG
large language model
大语言模型相关
Abstract
Large-scale text annotation brings expert insight to millions of documents, often through a codebook that AI annotators follow. Developing a robust codebook, however, takes months. Large language models (LLMs) could speed this process by applying an early codebook to the data, surfacing cases with strong LLM disagreement, and eliciting expert feedback to address them. We examined three ways experts can provide feedback for LLM codebook revision: (i) editing LLM-generated revisions driven by cross-LLM disagreement (Codebook Verifying), (ii) answering questions about LLM disagreements (Question Answering), and (iii) labeling disagreement cases with rationales (Rationale Labeling). Experiments on thousands of tutoring-session transcripts show that Rationale Labeling yielded the highest LLM-labeling accuracy (64.9%) against expert labels, outperforming the expert-revised codebook (57.8%). The best Question Answering setting also outperformed it (60.5%). Our work shows that LLMs can be used to strategically target expert attention, shortening months of codebook revision to days without sacrificing labeling performance.
Chinese Translation
大规模文本标注为海量文档带来专家洞见,通常借助AI标注者所遵循的编码手册来实现。然而,开发一个稳健的编码手册需要数月时间。大语言模型(LLM)可以通过将早期编码手册应用于数据、浮现出LLM存在强烈分歧的案例,并引出专家反馈以解决这些案例,来加速这一过程。我们考察了专家为LLM编码手册修订提供反馈的三种方式:(i)编辑由跨LLM分歧驱动的LLM生成修订(编码手册验证),(ii)回答关于LLM分歧的问题(问答),以及(iii)为分歧案例标注理由(理由标注)。在数千份辅导会话转录文本上的实验表明,理由标注相对于专家标签取得了最高的LLM标注准确率(64.9%),优于专家修订后的编码手册(57.8%)。最佳的问答设置也优于它(60.5%)。我们的工作表明,LLM可以被用来战略性地定位专家注意力,在不牺牲标注性能的情况下,将编码手册修订从数月缩短至数天。
cs.CL / 14 / 2609.26942
Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms
能识别但无法生成:面向文化特定亲属称谓的生成基准
Sahil Pardasani, Madhusudan Singh
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Current literature evaluates large language models (LLMs) on multilingual kinship understanding using multiple choice benchmarks, treating it as a recognition problem. We instead prompt five open weight LLMs to generate kinship terms in three non Western languages (Hindi, Tamil, and Korean) across two communicative tasks and pair this with a matched option-supported selection baseline. On identical relation language cells, GPT OSS120B selects the correct term in 90.67% of 75 valid cells but produces an accepted term in 36.00% of the corresponding attempts; Llama 3.370B shows the same pattern (77.92% versus 24.24%). Since the four-option condition displays the candidate terms and does not require script production, the difference is interpreted as an evaluation format gap rather than direct proof that lexical knowledge is intact. On explicitly specified L3 prompts, accuracy varies sharply, from GLM-5.1 at 72.29% to Llama-3.370B at 24.24%. The paternal-lineage advantage is language specific; it is large in Hindi but weak or reversed in Korean, while Tamil shared-term pairs provide a control for measurement variation. These results show that culturally specific kinship generation remains difficult even when the relationship is explicitly stated and motivate generation-based evaluation alongside multiple-choice testing.
Chinese Translation
当前文献使用多项选择基准来评估大语言模型(LLM)的多语言亲属关系理解,将其视为一个识别问题。我们转而提示五个开放权重 LLM 在两种交际任务中生成三种非西方语言(印地语、泰米尔语和韩语)的亲属称谓,并将其与一个匹配的、有选项支持的选择基线配对。在相同的关系—语言单元格上,GPT OSS120B 在 75 个有效单元格中有 90.67% 选择了正确称谓,但在相应尝试中仅有 36.00% 生成了可接受的称谓;Llama 3.370B 表现出相同模式(77.92% 对 24.24%)。由于四选项条件会展示候选称谓且不要求文字产出,这一差异被解释为评估格式上的差距,而非词汇知识完好的直接证据。在明确指定的 L3 提示上,准确率差异悬殊,从 GLM-5.1 的 72.29% 到 Llama-3.370B 的 24.24%。父系优势具有语言特异性;它在印地语中很大,但在韩语中较弱或发生反转,而泰米尔语的共享称谓对则为测量变异提供了一个对照。这些结果表明,即使关系被明确陈述,文化特定的亲属称谓生成仍然困难,并促使在多项选择测试之外开展基于生成的评估。
cs.CL / 15 / 2609.26945
Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court
在句子层面分类解释准则:一个来自德国联邦宪法法院的基准
Felix Ringe
cs.CL · cs.CY
large language model
大语言模型相关
Abstract
Judicial reasoning remains challenging for large language models (LLMs) to analyze. This paper contributes a sentence-level benchmark for evaluating the ability of LLMs to classify interpretive canons as articulated by Larenz in the tradition of Savigny. Our contributions are threefold. First, we operationalize this conception of interpretation as classification criteria. Second, we provide a dataset of decisions of the German Federal Constitutional Court annotated at the sentence level. Third, we report baseline evaluations of four LLMs from three model families under expert hand-written prompts, compared against prompts optimized with Genetic-Pareto (GEPA). Mean F1 over the seven binary subtasks clusters between 70.4 and 79.2 across models, with grammatical interpretation usually the easiest canon to identify and systematic interpretation usually the hardest; under the tested configuration, GEPA-optimized prompts do not systematically outperform the hand-written ones, suggesting that the expert prompts provide a meaningful baseline.
Chinese Translation
司法推理对于大语言模型(LLM)而言仍然难以分析。本文提供了一个句子级基准,用于评估 LLM 对 Larenz 在 Savigny 传统下所阐述的解释准则进行分类的能力。我们的贡献有三方面。第一,我们将这一解释概念操作化为分类标准。第二,我们提供了一个在句子级别上标注的德国联邦宪法法院裁判数据集。第三,我们报告了来自三个模型家族的四个 LLM 在专家手写提示下的基线评估,并将其与使用 Genetic-Pareto(GEPA)优化的提示进行比较。在七个二分类子任务上,各模型的平均 F1 介于 70.4 与 79.2 之间,其中语法解释通常是最容易识别的解释准则,而体系解释通常是最难的;在所测试的配置下,GEPA 优化的提示并未系统性地优于手写提示,这表明专家提示提供了一个有意义的基线。
cs.CL / 16 / 2609.27009
LEGO: Synergizing Expert GraphRAG and Expert Chain-of-Thought for Legal Reasoning
LEGO:协同专家GraphRAG与专家思维链以进行法律推理
Qingjing Chen, Junkai Zhang, Shaochun Wang, Jiahao Ding, Siyuan Zheng, Yukun Yan, Zhi Zheng, Antonino Rotolo, Yun Liu, Weixing Shen
cs.CL · cs.IR
large language model
大语言模型相关
Abstract
Large language models are increasingly applied to high-risk domains such as law, yet complex legal reasoning remains limited by two structural challenges. First, existing RAG and GraphRAG methods emphasize lexical or semantic similarity while overlooking normative relations among legal provisions. Second, vanilla Chain-of-Thought prompting may generate plausible rationales without enforcing the normative structure of legal reasoning. To deal with the bottleneck of pipelines in the legal reasoning domain, we propose LEGO, a dual-module framework that synergizes Legal Expert GraphRAG and expert Chain-of-thought for complex legal reasoning. ExpertGraphRAG uses an expert-annotated civil code graph encoding these normative relations with a greedy normative-coverage retrieval algorithm to dynamically extract instance-specific provision subgraphs, while ExpertCoT organizes the retrieved provisions and case facts into structured Provision-Fact-Conclusion reasoning. With a Qwen3-8B backbone, LEGO achieves 40.53% exact-match accuracy on LawExamQA_Civil, outperforming the evaluated RAG and CoT baselines and performing comparably to the evaluated larger models, while remaining robust on multi-hop questions. It also achieves the best results among the evaluated baselines on the open-ended benchmarks. Ablation studies confirm the individual and complementary contributions of both modules, demonstrating LEGO's effectiveness in improving LLMs' complex legal reasoning ability. Code and dataset can be found in the link: https://github.com/BLK-WHT/LEGO
Chinese Translation
大型语言模型正越来越多地应用于法律等高风险领域,然而复杂的法律推理仍受到两个结构性挑战的限制。首先,现有的RAG和GraphRAG方法强调词汇或语义相似性,却忽视了法律条文之间的规范性关系。其次,普通的思维链提示可能会生成看似合理的理由,却未强制遵循法律推理的规范性结构。为解决法律推理领域中流程的瓶颈,我们提出LEGO,这是一个双模块框架,协同法律专家GraphRAG与专家思维链,以进行复杂法律推理。ExpertGraphRAG使用一个由专家标注的民法典图来编码这些规范性关系,并采用贪心规范性覆盖检索算法动态抽取特定实例的条文子图;而ExpertCoT将检索到的条文和案件事实组织为结构化的“条文-事实-结论”推理。以Qwen3-8B为骨干,LEGO在LawExamQA_Civil上达到40.53%的精确匹配准确率,优于所评估的RAG和CoT基线,并与所评估的更大模型表现相当,同时在多跳问题上保持稳健。它还在开放式基准上取得了所评估基线中的最佳结果。消融研究证实了两个模块各自及互补的贡献,表明LEGO在提升LLM复杂法律推理能力方面的有效性。代码和数据集可在以下链接找到:https://github.com/BLK-WHT/LEGO
cs.CL / 17 / 2609.27043
EduBehaviors: Assertion-based Schemas for Auditable Coding of Educational Dialogues
EduBehaviors:用于可审计教育对话编码的基于断言的模式
Julian Bernado, Ana Trindade Ribeiro, Xander Beberman, Susanna Loeb
cs.CL · cs.CY · cs.LG
large language model
大语言模型相关
Abstract
Large language models have allowed the rapid deployment of pedagogical annotations corresponding to constructs of interest, allowing a natural language interface for generating classifications on a conversational dataset. However due to the opaque nature of LLM reasoning, we have no verifiable, mechanistic insight into why a model chose a label for an utterance. We introduce the EduBehaviors framework, an interpretable, scalable approach to annotating educational data that uses LLMs to measure repeated observable behaviors relevant to many constructs of interest and then learns a classifier for the construct based on these observable behaviors. We evaluate the framework on the TalkMoves dataset, predicting the Teacher TalkMoves labels. Our best configuration results in a macro-F1 of 0.673 and 0.688 Cohen's kappa, proving competitive with direct prompting approaches. In addition, we release EduBehaviors Toolkit, two tools allowing researchers to operationalize the EduBehaviors framework in their own data.
Chinese Translation
大语言模型使得对应于所关注构念的教学标注得以快速部署,从而提供了一种用于在对话数据集上生成分类的自然语言接口。然而,由于大语言模型推理过程的不透明本质,我们无法获得可验证的、机制性的洞见,来说明模型为何为某条话语选择了某个标签。我们提出了 EduBehaviors 框架,这是一种可解释、可扩展的教育数据标注方法,它使用大语言模型来测量与许多所关注构念相关的重复可观察行为,然后基于这些可观察行为学习该构念的分类器。我们在 TalkMoves 数据集上评估该框架,预测 Teacher TalkMoves 标签。我们最优的配置取得了 0.673 的 macro-F1 和 0.688 的 Cohen's kappa,证明其与直接提示方法相比具有竞争力。此外,我们发布了 EduBehaviors Toolkit,即两个工具,使研究者能够在他们自己的数据中实际运用 EduBehaviors 框架。
cs.CL / 18 / 2609.27165
Count Evidence, Not Sentences: Tempered Evidence Fusion of LLM Judgments for Long-Text Value Measurement
计数证据,而非句子:面向长文本价值测量的 LLM 判断的调温证据融合
Yuhe Wu, Rui Qian, Guangyu Wang, Yuran Chen, Yuanchao Zhu, Junjie Yang, Zhengheng Li, Jiulin Cai, Tianyi Zhang, Zihan Dong, Jiaxin Liu, Yujie Chen, Guang Zhang
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly used to measure public value orientations from long social media posts, yet such posts often mix background, quotations, concessions, and only a few stance-bearing sentences. Existing approaches either ask the model to predict a document-level label directly, which can be overconfident, or aggregate sentence-level predictions by majority or soft voting, which treat uncertain and decisive sentences as equally informative. We formulate long-text value measurement as a decision-fusion problem and propose Tempered Evidence Fusion (TEF), a training-free rule that weights each sentence's log-odds by its normalized information gain, as derived from a generalized Bayesian posterior. This makes the fused score nearly vanish for uncertain sentences while preserving the Bayes-optimal weight of decisive evidence. We further introduce Multi-event Insight Network Dimensions (MIND), a benchmark of 8,358 Chinese and English posts spanning five years of public events and six value dimensions. On MIND, TEF outperforms the strongest baseline among Direct, Majority Vote, and Soft Vote by an average of 4.5 accuracy points and 4.6 macro-F1 points across five LLMs and two languages. MIND dataset and code are available at https://github.com/Kzczc/ICASSP2027-TEF.
Chinese Translation
大语言模型(LLMs)正越来越多地被用于从长篇社交媒体帖子中测量公众价值取向,然而此类帖子往往混合了背景、引语、让步内容,以及只有少数承载立场的句子。现有方法要么要求模型直接预测文档级标签,这可能过度自信;要么通过多数投票或软投票来聚合句子级预测,而这些方法将不确定句子和决定性句子视为具有同等信息量。我们将长文本价值测量形式化为一个决策融合问题,并提出调温证据融合(Tempered Evidence Fusion, TEF),这是一种无需训练的规则,它根据每个句子的归一化信息增益来对其对数几率进行加权,而该信息增益推导自广义贝叶斯后验。这使融合分数对于不确定句子几乎消失,同时保留决定性证据的贝叶斯最优权重。我们进一步引入多事件洞察网络维度(Multi-event Insight Network Dimensions, MIND),这是一个包含 8,358 篇中英文帖子的基准,跨越五年公共事件和六个价值维度。在 MIND 上,TEF 在五个 LLM 和两种语言上,相较于 Direct、Majority Vote 和 Soft Vote 中最强的基线,平均高出 4.5 个准确率点和 4.6 个 macro-F1 点。MIND 数据集和代码可在 https://github.com/Kzczc/ICASSP2027-TEF 获取。
cs.CL / 19 / 2609.27220
LOCKR: A Hidden-State Trajectory-Guided Planner for Detecting and Repairing Stable-but-Wrong Lock-In in Diffusion Language Models
LOCKR:一种用于检测与修复扩散语言模型中稳定但错误的锁定现象的隐状态轨迹引导规划器
Guoshenghui Zhao, Tan Yu, Weijie Zhao
cs.CL
diffusion
扩散模型相关
Abstract
Diffusion language models generate text through iterative denoising, exposing intermediate trajectories before final answers are produced. We identify a recurring reasoning failure, stable-but-wrong lock-in, where an answer stabilizes early around an incorrect value while substantial denoising remains. Surface-level decoding signals such as confidence, entropy, margin, and answer stability are insufficient to reliably distinguish correct from erroneous lock-in. We formulate selective reasoning repair as a lightweight test-time planning problem and propose LOCKR, a hidden-state trajectory-guided planner that decides when to allocate additional computation, expands a structured set of targeted repair branches, and selects the most promising continuation using trajectory-aware verification. Across two diffusion language models and three mathematical reasoning benchmarks, hidden-state trajectories consistently outperform surface signals and single hidden snapshots for both wrong-lock-in detection and repair selection. On natural evaluation distributions, LOCKR yields absolute accuracy gains of 2.21--5.37 percentage points across all five evaluated settings, with repair rates ranging from 22% to 41%. These results establish hidden diffusion trajectories as actionable signals for selective test-time reasoning repair.
Chinese Translation
扩散语言模型通过迭代去噪生成文本,在最终答案产生之前暴露出中间轨迹。我们发现了一种反复出现的推理失败模式,即稳定但错误的锁定(stable-but-wrong lock-in):答案在早期就围绕一个错误的值稳定下来,而此后仍有大量去噪过程尚未完成。诸如置信度、熵、间隔(margin)和答案稳定性等表层解码信号,不足以可靠地将正确的锁定与错误的锁定区分开来。我们将选择性推理修复形式化为一个轻量级的测试时规划问题,并提出 LOCKR——一种隐状态轨迹引导的规划器,它决定何时分配额外计算,扩展一组结构化的定向修复分支,并使用轨迹感知的验证来选择最有希望的延续。在两个扩散语言模型和三个数学推理基准上,对于错误锁定检测和修复选择这两项任务,隐状态轨迹都持续优于表层信号和单一隐状态快照。在自然评估分布上,LOCKR 在所有五个评估设置中带来了 2.21--5.37 个百分点的绝对准确率提升,修复率介于 22% 到 41% 之间。这些结果确立了隐扩散轨迹作为选择性测试时推理修复的可操作信号。
cs.CL / 20 / 2609.27225
Meet, Compare, or Abstain: LatWeave for Deterministic Multi-Hop Question Answering on Knowledge Lattices
Meet、Compare 或 Abstain:面向知识格上确定性多跳问答的 LatWeave
Yuze Ren, Shaoheng Fan, Tao Wang, Yabo Yan, Han Han
cs.CL · cs.AI · cs.IR
large language model
大语言模型相关
Abstract
Probabilistic question-answering systems -- whether large language models (LLMs) themselves, retrieval-augmented generation (RAG), or trained multi-hop retrievers -- conflate "what is known" and "how to reason" into a single probabilistic computation: hallucination cannot be eradicated, evidence chains cannot be audited, and the system answers even when it does not know. We present LatWeave, which organizes knowledge into a multidimensional knowledge lattice and compiles multi-hop QA into three deterministic operators -- meet (constraint intersection), compare (lattice-order comparison), and abstain (structural abstention); LLMs appear only on the construction side (one-shot extraction) and the query-planning side, while the answer-generation path is zero-LLM, zero-task-training, and auditable end to end -- so that question answering over Web-published knowledge becomes reproducible item by item. Rather than claiming across-the-board SOTA, we characterize the operating envelope of this paradigm on six public benchmarks: when knowledge is complete (MetaQA, 39,093 questions) meet chains are near-lossless over three hops (any-hit 0.9975, on par with fully supervised KBQA); on templated multi-hop home ground (2WikiMultihopQA held-out n=1,258) EM 0.865, well above published structure-augmented RAG reproductions; on open-text deep composition (MuSiQue) and extraction-coverage gaps (HotpotQA) we report degradation honestly and attribute it to causes outside the lattice-algebra layer; and when information is incomplete (IIRC) we achieve structural abstention with abstain accuracy 0.971 and leak rate 0.029. Within the operating envelope, deterministic execution pays no performance penalty, and every step on the answer path can be recomputed -- precisely the source of end-to-end auditability.
Chinese Translation
概率性问答系统——无论是大语言模型(LLM)本身、检索增强生成(RAG),还是经过训练的多跳检索器——都将“已知什么”与“如何推理”混同为单一的概率计算:幻觉无法根除,证据链无法审计,并且系统即使并不知道也会作答。我们提出 LatWeave,它将知识组织为多维知识格,并把多跳问答编译为三个确定性算子——meet(约束交集)、compare(格序比较)和 abstain(结构性弃权);LLM 仅出现在构建侧(一次性抽取)和查询规划侧,而答案生成路径则是零 LLM、零任务训练,且端到端可审计——从而使面向 Web 已发布知识的问答逐项可复现。我们并不声称全面达到 SOTA,而是在六个公开基准上刻画该范式的运行包络:当知识完整时(MetaQA,39,093 个问题),meet 链在三个跳数上近乎无损(any-hit 0.9975,与全监督 KBQA 相当);在模板化多跳的主场(2WikiMultihopQA 留出集 n=1,258)上 EM 为 0.865,远高于已发表的结构增强 RAG 复现结果;在开放文本深度组合(MuSiQue)和抽取覆盖缺口(HotpotQA)上,我们如实报告性能退化,并将其归因于格代数层之外的原因;当信息不完整时(IIRC),我们实现结构性弃权,弃权准确率为 0.971,泄漏率为 0.029。在运行包络内,确定性执行不带来性能损失,且答案路径上的每一步都可重新计算——这正是端到端可审计性的来源。
cs.CL / 21 / 2609.27359
Automated Extraction of Records of Processing Activities (RoPA) Using Hybrid RAG and Locally Deployed Large Language Models
使用混合 RAG 与本地部署大语言模型的处理活动记录(RoPA)自动提取
To Duy Hinh, Nguyen Le Quoc Anh, Phan Van Tri, Khuong Nguyen-An
cs.CL · cs.CR · cs.IR
large language model
大语言模型相关
Abstract
Vietnam's Personal Data Protection Law (Law No. 91/2025/QH15) and Decree No. 356/2025/ND-CP, effective January 1, 2026, require organizations to establish and maintain Records of Processing Activities (RoPA). Manual RoPA preparation is labor-intensive, while cloud-hosted large language models (LLMs) may conflict with data-sovereignty requirements. We propose RoPA Manager, a system for automated RoPA information extraction using hybrid retrieval that combines lexical ranking over tsvector, dense-vector search, Reciprocal Rank Fusion (RRF), and locally deployed LLMs. We introduce a Vietnamese RoPA benchmark with 32 organizations, 77 processing activities, 12 field groups, and 4,338 reference values. Evaluation is reported at three distinct levels. The automated scorer, tested on perturbed data without invoking an LLM, achieved F1 = 0.9493 [0.9436, 0.9548]; this measures scorer robustness rather than end-to-end extraction accuracy. End-to-end extraction achieved token coverage of 50.04-55.25% against the reference labels. Two independent experts reviewed 1,558 reference values (35.9% of the benchmark), found no incorrect values, and achieved 99.68% agreement with PABAK = 0.9936. Value-level precision was not measured. Across 32 paired scenarios on a 24 GB GPU, locally deployed Qwen3.5-27B-GPTQ-Int4 showed no statistically significant difference from cloud-based DeepSeek-V4-Flash (difference 0.20 percentage points in favor of DeepSeek, 95% CI [-0.93, 1.32], p = 0.72), while Gemma-4-31B performed significantly worse (p < 0.01).
Chinese Translation
越南的《个人数据保护法》(第 91/2025/QH15 号法律)与第 356/2025/ND-CP 号法令将于 2026 年 1 月 1 日生效,要求组织建立并维护处理活动记录(RoPA)。人工编制 RoPA 劳动强度大,而云端托管的大语言模型(LLM)可能与数据主权要求相冲突。我们提出 RoPA Manager,这是一个使用混合检索进行 RoPA 信息自动提取的系统,该混合检索结合了基于 tsvector 的词法排序、稠密向量搜索、倒数排名融合(RRF)以及本地部署的 LLM。我们引入了一个越南语 RoPA 基准,包含 32 个组织、77 项处理活动、12 个字段组和 4,338 个参考值。评估在三个不同层面上报告。自动评分器在不调用 LLM 的情况下在扰动数据上进行测试,取得了 F1 = 0.9493 [0.9436, 0.9548];这衡量的是评分器的稳健性,而非端到端提取准确率。端到端提取相对于参考标签取得了 50.04–55.25% 的 token 覆盖率。两位独立专家审查了 1,558 个参考值(占基准的 35.9%),未发现错误值,并达到 99.68% 的一致性,PABAK = 0.9936。未测量值级别的精确率。在 24 GB GPU 上的 32 个配对场景中,本地部署的 Qwen3.5-27B-GPTQ-Int4 与基于云端的 DeepSeek-V4-Flash 相比未显示出统计学显著差异(差异为 0.20 个百分点,有利于 DeepSeek,95% CI [-0.93, 1.32],p = 0.72),而 Gemma-4-31B 的表现显著更差(p < 0.01)。
cs.CL / 22 / 2609.27396
When Parallel Drafter Meets Parallel Speculative Decoding
当并行草稿器遇上并行投机解码
Fuliang Liu, Xue Li, Kun Qian, Zhibin Wang, Wanchun Dou, Wenyuan Yu, Chen Tian
cs.CL
diffusion
扩散模型相关
Abstract
DSpark-style parallel drafters have made speculative decoding highly effective, yet their draft phase remains serialized on the critical path of every round. Parallel speculative decoding (PSD) overlaps drafting with verification, yet existing methods must guess the accepted prefix and bonus token in advance: a wrong guess reverts the whole batch to serial drafting. We present DPara, a PSD framework that reuses effective parallel drafters yet guarantees backbone--verification overlap in every round, thereby eliminating this probabilistic fallback altogether. While the target verifies, DPara's diffusion backbone precomputes draft representations for every acceptance boundary with the bonus left unspecified; a lightweight autoregressive head then combines the revealed verification outcome with the matching precomputed representation to emit the next round's draft tokens almost instantly---fully parallelizing the dominant backbone forward with verification and leaving only the negligible head cost serial. Experiments on Qwen3-8B and Qwen3-14B across seven math, coding, and chat benchmarks show that DPara achieves average speedups of $3.21\times$ and $3.52\times$ over autoregressive decoding, surpassing the strongest serial and parallel speculative decoding baselines alike.
Chinese Translation
DSpark 风格的并行草稿器使投机解码变得极为高效,然而其草稿阶段在每一轮的关键路径上仍然是串行的。并行投机解码(PSD)将草稿与验证重叠起来,但现有方法必须预先猜测被接受的前缀和 bonus token:一旦猜错,整个批次就会退回串行草稿。我们提出 DPara,一个 PSD 框架,它复用有效的并行草稿器,同时在每一轮都保证主干与验证的重叠,从而彻底消除这种概率性回退。在目标模型进行验证的同时,DPara 的扩散主干为每个接受边界预计算草稿表示,而将 bonus 留作未指定;随后一个轻量级自回归头将已揭晓的验证结果与相匹配的预计算表示结合起来,几乎瞬间产出下一轮的草稿 token——将占主导地位的主干前向与验证完全并行化,只留下可忽略不计的头开销串行执行。在 Qwen3-8B 和 Qwen3-14B 上跨七个数学、代码与对话基准的实验表明,DPara 相较于自回归解码实现了平均 $3.21 imes$ 和 $3.52 imes$ 的加速,并同样超越了最强的串行与并行投机解码基线。
cs.CL / 23 / 2609.27418
EviStreams: Human-in-the-Loop AI Data Extraction for Systematic Reviews in Medicine
EviStreams:面向医学系统综述的人在回路 AI 数据提取
Sai Karthik Kosuri, Ankita Shashikant Bhosale, Michael Glick, Alonso Carrasco-Labra, Chris Callison-Burch
cs.CL
large language model
大语言模型相关
Abstract
Systematic reviews underpin clinical guidelines, yet their data-extraction step is a major expert-labor bottleneck bound by a protocolized workflow: two reviewers extract each study independently, an adjudicator resolves disagreements, and the team keeps an auditable record of how every value was produced. Large language models can assist with extraction, but that assistance must fit established review protocols and preserve reproducibility. We present EviStreams, a live, open-source, no-code web platform that puts review teams in control of AI-assisted extraction at three key stages: program design (a structured decomposition approved before any code runs), field specification (typed field definitions calibrated from a pilot), and extracted predictions (reviewer-blinded dual review with adjudication). Working through a form builder, a domain expert defines typed fields rather than prompts, runs extraction over uploaded PDFs, inspects every value alongside the supporting passage it came from, and resolves a reviewer-blinded dual review into an auditable consensus export. An evaluation across four clinical corpora and three frontier model families, released with the system, shows that extraction quality is shaped far more by the field specification than by the choice of model. EviStreams is live at https://evistreams.com/demo and released under Apache-2.0.
Chinese Translation
系统综述是临床指南的基础,但其数据提取环节是一个受方案化工作流程约束的重大专家人力瓶颈:两名评价者独立提取每项研究的数据,由一名裁决者解决分歧,并且团队保留每个数值是如何产生的可审计记录。大语言模型可以辅助提取,但这种辅助必须契合既有的综述方案并保持可重复性。我们提出 EviStreams,一个已上线运行、开源、无代码的网页平台,它让综述团队在三个关键阶段掌控 AI 辅助的提取:程序设计(在任何代码运行之前即获批准的结构化分解)、字段规格说明(依据试点校准的类型化字段定义),以及提取出的预测结果(带有裁决的、对评价者设盲的双人评审)。通过表单构建器,领域专家定义的是类型化字段而非提示词,对上传的 PDF 运行提取,将每个数值与其来源的支持性段落一并查看,并把对评价者设盲的双人评审裁定为可审计的共识导出结果。一项涵盖四个临床语料库和三个前沿模型家族的评估随该系统一同发布,其结果表明,提取质量受字段规格说明的影响远大于受模型选择的影响。EviStreams 已在 https://evistreams.com/demo 上线,并以 Apache-2.0 许可发布。
cs.CL / 24 / 2609.27510
Uncheatable Eval: Dynamic Compression-Based Evaluation of Language Models
Uncheatable Eval:基于动态压缩的语言模型评估
Kaifeng Tan, Yudong Li, Linlin Shen
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Modern large language models are pretrained on massive datasets, making it difficult to prevent benchmark data from entering their training sets and undermining the reliability of evaluation results. Reliable evaluation is particularly challenging for base models, whose limited instruction-following ability complicates task-based assessment. We introduce Uncheatable Eval, a dynamic benchmark that regularly collects newly published text to evaluate base language models and reduce the risk of data contamination. Drawing on the relationship between a model's predictive ability and its ability to compress data losslessly, we use compression rate to evaluate how well models predict new text. We evaluate 80 models across 14 text categories, study how compression changes with context length, and examine the correlation between compression rate and zero-shot MMLU accuracy. Our results yield three main findings: (1) compression performance follows a consistent scaling trend with model size; (2) attention-based, hybrid, and recurrent models differ in how their compression performance changes as more context becomes available; and (3) lower compression rates are strongly associated with higher zero-shot MMLU accuracy. Code is available at https://github.com/Jellyfish042/uncheatable_eval.
Chinese Translation
现代大型语言模型是在海量数据集上预训练的,这使得难以防止基准数据进入其训练集,并削弱评估结果的可靠性。可靠评估对于基础模型尤其具有挑战性,因为它们有限的指令遵循能力使基于任务的评估变得复杂。我们提出了 Uncheatable Eval,这是一个动态基准,定期收集新发布的文本,以评估基础语言模型并降低数据污染的风险。借鉴模型预测能力与其无损压缩数据能力之间的关系,我们使用压缩率来评估模型对新文本的预测效果。我们在 14 个文本类别上评估了 80 个模型,研究了压缩如何随上下文长度变化,并考察了压缩率与零样本 MMLU 准确率之间的相关性。我们的结果得出三个主要发现:(1)压缩性能随模型规模呈现一致的缩放趋势;(2)基于注意力的模型、混合模型和循环模型在对更多上下文可用时其压缩性能的变化方式不同;(3)较低的压缩率与较高的零样本 MMLU 准确率密切相关。代码可在 https://github.com/Jellyfish042/uncheatable_eval 获取。
cs.CL / 25 / 2609.27603
When Context Misleads: In-context Learning with Jurisdiction in Large Language Models
当上下文误导时:大语言模型中带管辖权的上下文学习
Pei-lin Li, Qingle Liu, Junyang Feng, Siyu Li, Sunqi Fan, Xin-Sheng Chen, Shuojin Yang
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
In-Context Learning (ICL) has become a cornerstone of modern LLM deployment. However, existing ICL post-training methods have a critical blind spot: they excel at extracting patterns from demonstrations while often neglecting context authority, the ability to determine whether contextual information should govern the final answer. To benchmark this capability, we introduce FakeContextBench, which contains pseudoscientific claims across seven domains. Our evaluation of commercial and open-source models shows that large-scale pre-training alone is insufficient for reliable context-authority discrimination. Moreover, prevalent ICL fine-tuning methods can increase susceptibility to misleading context, reducing reality accuracy by up to 14.95 percentage points relative to the base model. To address this trade-off, we propose Jurisdiction In-Context Learning (J-ICL), a post-training framework that incorporates context validation into the training objective. Across four model backbones, J-ICL improves ICLEval by an average of 5.84 percentage points and reality accuracy by 9.20 points over the corresponding base models. It also raises the Reality Rate by an average of 18.09 points relative to MetaICL and Symbol Tuning. These results demonstrate that ICL capability and resistance to deceptive context can be improved together. The benchmark is available at https://github.com/peilin717/FakeContext-Bench.
Chinese Translation
上下文学习(ICL)已成为现代大语言模型(LLM)部署的基石。然而,现有的 ICL 后训练方法存在一个关键盲点:它们擅长从演示中提取模式,却常常忽视上下文权威性,即判断上下文信息是否应支配最终答案的能力。为了对该能力进行基准测试,我们提出了 FakeContextBench,它包含跨七个领域的伪科学主张。我们对商业模型和开源模型的评估表明,仅靠大规模预训练不足以实现可靠的上下文权威判别。此外,流行的 ICL 微调方法可能增加对误导性上下文的易感性,相对于基础模型最多使现实准确率降低 14.95 个百分点。为了解决这一权衡,我们提出了管辖权上下文学习(J-ICL),这是一个将上下文验证纳入训练目标的后训练框架。在四个模型骨干上,J-ICL 相较于相应的基础模型,使 ICLEval 平均提高 5.84 个百分点,使现实准确率提高 9.20 个百分点。相对于 MetaICL 和 Symbol Tuning,它还使 Reality Rate 平均提高 18.09 个百分点。这些结果表明,ICL 能力与对欺骗性上下文的抵抗能力可以同时得到提升。该基准测试可在 https://github.com/peilin717/FakeContext-Bench 获取。
cs.CL / 26 / 2609.27690
Consequential Behaviour and Representational Fairness in the Validation of Synthetic Research
合成研究验证中的后果性行为与表征公平
Florian Kutzner, Celina Kacperski, Laura de Molière, Edoardo Chidichimo, Min Jun Jung, Felix Patrick Sedgwick Wallis, James Kunling He
cs.CL · cs.CY
large language model
大语言模型相关
Abstract
Researchers in industry and academia use synthetic survey respondents powered by large language models as substitutes for human samples. These synthetic populations require validation against real-world data, so researchers often address them using ad hoc comparisons with human surveys. Inspired by the intention-behaviour gap in behavioural science, we argue that these validations test the wrong thing for most applied cases where decision makers commission synthetic research to anticipate consequential behaviour. To address this problem, we propose a validation framework with two requirements. First, every validity claim must state its level of correspondence with human data: does the sample predict what the represented people do, which of four diagnostics (location, dispersion, response process and structure) does the validation address, and does the validation compare against experimental effects? Second, researchers must report validity claims for subgroups, since these groups are often the most affected by consequential decisions and aggregate accuracy hides their misrepresentation. Our validation framework operationalises three justice dimensions (distributional, procedural, and recognition) as measurable quantities and defines within-persona counterfactual experiments as a validation requirement. We then apply the framework to electric vehicle charging tariffs, before closing with a reporting checklist that researchers can use to make convincing validity claims.
Chinese Translation
产业界和学术界的研究者使用由大语言模型驱动的合成调查受访者作为人类样本的替代品。这些合成群体需要针对真实世界数据进行验证,因此研究者常常通过与人类调查进行临时性比较来处理这一问题。受行为科学中意图—行为差距的启发,我们认为,对于大多数应用案例而言,这些验证测试了错误的东西,在这些案例中,决策者委托合成研究以预判后果性行为。为解决这一问题,我们提出了一个具有两项要求的验证框架。第一,每一个有效性主张都必须说明其与人类数据的对应程度:样本是否预测了其所代表的人会做什么,验证涉及四种诊断(位置、离散度、响应过程和结构)中的哪一种,以及验证是否与实验效应进行比较?第二,研究者必须报告针对子组的有效性主张,因为这些群体往往受后果性决策影响最大,而聚合准确性掩盖了他们的错误表征。我们的验证框架将三个正义维度(分配、程序和承认)操作化为可测量的量,并将角色内反事实实验定义为一项验证要求。然后,我们将该框架应用于电动汽车充电电价,最后以一个报告清单作结,研究者可以用它来提出令人信服的有效性主张。
cs.CL / 27 / 2609.27717
SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving
SkillGym:将人类技能内化到 LLM 中以解决现实世界问题
Zhilong Ge, Yuting Shao, Yutao Yang, Yuxuan Cai, Jie Zhou, Kai Chen, Bo Zhang, Qin Chen, Liang He
cs.CL
large language model
大语言模型相关
Abstract
Human-written agent skills encode rich workflows for real-world problem solving, but are typically used as external inference-time instructions rather than internalized as reusable model capabilities. We introduce \texttt{SkillGym}, a framework that transforms these skills into executable, verifiable training environments for large language model agents. Its skill-to-task pipeline instantiates concrete tasks, verifies outcomes with code-based checkers, and assesses empirical skill dependence through contrastive executions. We construct and release 2,756 environments across 12 categories and collect 8,364 successful trajectories from multiple models and harnesses, averaging 49 tool calls and over 60k logged text tokens. These resources support supervised fine-tuning on verified workflows and reinforcement learning with outcome-based rewards. Under Claude Code, supervised fine-tuning improves Qwen3.5-35B-A3B by 199 Elo on GDPval-AA v2, 19.10 percentage points on Terminal-Bench 2.1, and 28.13 and 12.38 points on SkillsBench v1.1 with and without skills, respectively. Our 35B \texttt{SkillGym-Agent} reaches 51.47\% on skill-assisted SkillsBench, exceeding reported scores for Claude Sonnet 4.6, GPT-5.4 Mini, and DeepSeek V4 Pro. Without skills, it also surpasses skill-assisted bases under Codex and Claude Code, suggesting reusable procedural competence.
Chinese Translation
人类编写的智能体技能编码了用于现实世界问题求解的丰富工作流,但通常被用作外部的推理时指令,而不是内化为可复用的模型能力。我们提出 \texttt{SkillGym},一个将这些技能转化为面向大语言模型智能体的可执行、可验证训练环境的框架。其技能到任务流水线会实例化具体任务,使用基于代码的检查器验证结果,并通过对比执行来评估经验性的技能依赖。我们构建并发布了涵盖 12 个类别的 2,756 个环境,并从多个模型和 harness 收集了 8,364 条成功轨迹,平均每次 49 次工具调用和超过 60k 条记录的文本 token。这些资源支持在已验证工作流上进行监督微调,以及使用基于结果的奖励进行强化学习。在 Claude Code 下,监督微调使 Qwen3.5-35B-A3B 在 GDPval-AA v2 上提升 199 Elo,在 Terminal-Bench 2.1 上提升 19.10 个百分点,并在有技能和无技能情况下分别在 SkillsBench v1.1 上提升 28.13 和 12.38 分。我们的 35B \texttt{SkillGym-Agent} 在技能辅助的 SkillsBench 上达到 51.47%,超过了 Claude Sonnet 4.6、GPT-5.4 Mini 和 DeepSeek V4 Pro 报告的成绩。在没有技能的情况下,它还在 Codex 和 Claude Code 下超过了技能辅助的基线,这表明其具备可复用的程序性能力。
cs.CL / 28 / 2609.28007
Evaluating Open-Weight LLMs for Turkish Domain Documents Under Retrieval and Hardware Constraints
在检索与硬件约束下评估面向土耳其语领域文档的开放权重大语言模型
Imtiaz Ul Hassan, Öykü Akbulut, Onur Kaya, Ardhendu Behera, Swagat Kumar, Peter Matthew, Yonghuai Liu
cs.CL · cs.IR · cs.LG
large language model
大语言模型相关
Abstract
Most Turkish-capable large language models (LLMs) are evaluated using general-purpose benchmarks rather than long, structurally complex domain documents. This paper evaluates five open-weight 7B-8B models for Turkish document question answering under a resource-constrained local deployment setting. The primary benchmark contains 100 systematically validated questions derived from a 109-page industrial R&D report, and the evaluation protocol is replicated using a second 112-page public-sector report and an independently constructed 100-question set. All models are evaluated locally on an NVIDIA RTX 3050 laptop GPU with 6 GB VRAM using controlled prompting, decoding, and 4-bit quantisation. The principal methodological contribution is an evidence-annotated evaluation protocol that separates retrieval failure from downstream model reasoning failure without requiring additional model calls. On the primary benchmark, end-to-end accuracy ranges from 49% to 75%. Seven lexical, dense, and hybrid retrieval configurations are additionally compared using 95% Wilson intervals and exact paired McNemar tests; none significantly outperforms the character TF-IDF baseline on either document. Evidence recall saturates differently across the two reports, showing that retrieval and effective context capacity can be binding constraints for some documents but not others. These results demonstrate that model selection, retrieval behaviour, and hardware limits must be evaluated separately when deploying open-weight LLMs for Turkish domain documents.
Chinese Translation
大多数具备土耳其语能力的大语言模型(LLM)是使用通用基准进行评估的,而不是使用长篇、结构复杂的领域文档。本文在资源受限的本地部署环境下,评估了五个开放权重的7B-8B模型用于土耳其语文档问答。主要基准包含100个系统验证的问题,这些问题源自一份109页的工业研发报告;并且该评估协议通过第二份112页的公共部门报告和一个独立构建的100题集合进行复现。所有模型均在配备6 GB显存的NVIDIA RTX 3050笔记本电脑GPU上本地评估,使用受控提示、解码和4比特量化。主要方法学贡献是一种带证据标注的评估协议,该协议无需额外的模型调用即可将检索失败与下游模型推理失败区分开来。在主要基准上,端到端准确率范围为49%至75%。另外,使用95% Wilson区间和精确配对McNemar检验比较了七种词汇、稠密和混合检索配置;在任一文档上,它们均未显著优于字符TF-IDF基线。证据召回率在两份报告中的饱和方式不同,表明检索和有效上下文容量对某些文档可能是约束性限制,而对另一些文档则不是。这些结果表明,在将开放权重LLM部署用于土耳其语领域文档时,必须分别评估模型选择、检索行为和硬件限制。
cs.CL / 29 / 2609.28026
Evaluating Feedback Focus and Pedagogical Adaptivity in LLM-Generated Feedback on Student Writing
评估LLM生成的学生写作反馈中的反馈焦点与教学适应性
Norah Almousa, Shayan Peyghambari Oskoui, Raquel Coelho, Gayle Rogers, Xiang Lorraine Li, Diane Litman
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
We investigate whether state-of-the-art large language models (LLMs) generate feedback that reflects the pedagogical practices of expert teachers in terms of feedback focus and adaptivity. Previous evaluation efforts have examined feedback characteristics, its impact on learning, and its target, yet the focus of feedback and its adaptivity remains largely overlooked. To bridge this gap, we adopt and refine Narciss's taxonomy into seven feedback focus types to annotate teacher and LLM-generated feedback across three university writing courses. We release FeedType, a benchmark containing annotated teacher and LLM feedback from six LLMs under three prompting strategies. We assess the coverage and distribution of feedback focus types, and examine whether LLMs adapt their feedback across draft stages and student performance levels as an expert instructor does. Our findings show that while most LLMs cover most feedback focus types, they fail to reflect teacher feedback distributions and show varying levels of adaptivity, with none matching the teachers' adaptive behavior. We believe FeedType will support future research on pedagogical alignment in LLM feedback generation.
Chinese Translation
我们研究最先进的大语言模型(LLM)所生成的反馈,在反馈焦点与适应性方面是否反映了专家教师的教学实践。以往的评估工作考察了反馈的特征、其对学习的影响以及其目标,然而反馈的焦点及其适应性在很大程度上仍被忽视。为弥补这一空白,我们采用并改进Narciss的分类法,将其细化为七种反馈焦点类型,用以标注三门大学写作课程中教师与LLM生成的反馈。我们发布FeedType,这是一个基准,包含在三种提示策略下来自六个LLM的经标注的教师反馈与LLM反馈。我们评估各反馈焦点类型的覆盖范围与分布,并考察LLM是否像专家教师那样,在草稿阶段和学生表现水平之间调整其反馈。我们的研究结果表明,尽管大多数LLM覆盖了大多数反馈焦点类型,但它们未能反映教师的反馈分布,并表现出程度不一的适应性,且没有一个能匹配教师的适应性行为。我们相信FeedType将为未来关于LLM反馈生成中教学一致性的研究提供支持。
cs.CL / 30 / 2609.28041
How Much Were You Told? Measuring External Information in Peer Reviews
你被告知了多少?测量同行评审中的外部信息
Matthieu Dubois, Pablo Piantanida, François Yvon
cs.CL
large language model
大语言模型相关
Abstract
Conference policies distinguish using Large Language Models (LLMs) to polish one's own review from delegating the critique, but current Artificial Text Detection (ATD) methods largely measure surface form rather than the origin of its content. We instead measure the external information carried by a review: information not explained by the reviewed paper and a generic reviewing instruction. We propose Self-Conditioning, an unsupervised information-theoretic estimator that compares the likelihood of a review under its production context with its likelihood when that context is augmented with hints extracted from the review itself. On the IntelLabs peer-review benchmark, Self-Conditioning separates fully-delegated from machine-polished reviews with AUC up to $1.0$ while remaining largely insensitive to surface rewriting. Moreover, as generators receive increasing amounts of externally-provided information, their scores move monotonically towards the human regime, unlike standard ATD baselines. High-temperature sampling can evade the estimator, but at the cost of output quality.
Chinese Translation
会议政策区分了使用大语言模型(LLMs)润色自己的评审与将批评委托出去,但当前的人工文本检测(ATD)方法在很大程度上衡量的是表面形式,而不是其内容的来源。相反,我们测量一篇评审所携带的外部信息:即无法由被评审论文和一条通用评审指令所解释的信息。我们提出 Self-Conditioning,一种无监督的信息论估计器,它比较一篇评审在其生成上下文下的似然,与当该上下文用从评审本身提取的提示进行增强时该评审的似然。在 IntelLabs 同行评审基准上,Self-Conditioning 以高达 $1.0$ 的 AUC 将完全委托的评审与机器润色的评审区分开来,同时对表面改写基本不敏感。此外,随着生成器接收到越来越多的外部提供信息,它们的得分会单调地移向人类区间,这与标准 ATD 基线不同。高温采样可以规避该估计器,但代价是输出质量。
cs.CL / 31 / 2609.28150
Exact Feedback Is Not Control: Evaluating Text-based Closed-Loop Revision in LLMs
精确反馈并非控制:评估LLM中基于文本的闭环修订
Haitong Jiang, Chunlin Liu, Yile Wang, Yuhong Feng
cs.CL
large language model
大语言模型相关
Abstract
Closed-loop revision is increasingly used in large language model (LLM) applications, but failures may reflect incomplete feedback or ineffective responses to correct feedback. We introduce a fixed-budget revision protocol with deterministic verifiers that report all remaining violations across exact-length, lexical, and compositional constraints. Fixing feedback correctness and completeness isolates model-side revision behavior. Across 19 open- and closed-source models, controller-level mean final joint success ranges from 17.4% to 99.8%, with substantial cross-model gaps persisting under identical initial drafts. Controlled experiments reveal reproducible model-specific responses to exact feedback. Post-training and scale reshape these responses without consistently bringing them closer to exact correction. Across all constraint families, failed trajectories often repeat earlier outputs, and prior recurrence is associated with lower subsequent recoverability. Matched-state interventions show that removing earlier dialogue while holding the current draft and feedback fixed changes recurrence escape without reliably improving final success; effects depend on the model, task, and trigger-state composition. Exact feedback makes revision errors observable, but does not make the closed loop reliable. Code and reproduction instructions: https://github.com/kevinjiang0121-cyber/exact-feedback-code.
Chinese Translation
闭环修订正越来越多地用于大语言模型(LLM)应用中,但失败可能反映的是不完整的反馈,或是对正确反馈的无效响应。我们引入一种固定预算的修订协议,其使用确定性验证器报告在精确长度、词汇和组合约束下所有剩余的违规。固定反馈的正确性和完整性,可以隔离出模型侧的修订行为。在19个开源和闭源模型中,控制器层面的平均最终联合成功率范围从17.4%到99.8%,并且在相同的初始草稿下,模型间仍存在显著差距。受控实验揭示了模型对精确反馈具有可复现的模型特定响应。后训练和规模重塑了这些响应,但并未一致地使它们更接近精确纠正。在所有约束类别中,失败的轨迹常常重复较早的输出,而先前的重复与后续较低的恢复能力相关。匹配状态干预表明,在保持当前草稿和反馈固定的同时移除较早对话,会改变从重复中逃逸的情况,但不能可靠地提高最终成功率;其效果取决于模型、任务和触发状态组成。精确反馈使修订错误变得可观察,但并不能使闭环变得可靠。代码和复现说明:https://github.com/kevinjiang0121-cyber/exact-feedback-code。
cs.CL / 32 / 2609.28245
Beyond Poetry: Can Large Language Models Generate Classical Arabic Maqamat?
超越诗歌:大型语言模型能否生成古典阿拉伯语玛卡梅?
AbdulRahman A. Morsy, Aya Zirikly
cs.CL · cs.AI
large language model
大语言模型相关
Abstract
Large language models (LLMs) have shown strong performance in creative text generation, yet their ability to produce culturally grounded and stylistically constrained literary forms remains underexplored. Prior work has focused largely on modern language varieties and poetry, while classical prose traditions such as maqama remain largely unstudied. The maqama is a classical literary genre characterized by rhymed prose (saj), dense rhetorical ornamentation, and episodic narrative structure, making it a challenging testbed for evaluating whether LLMs can move beyond surface fluency toward deeper literary competence. In this paper, we present the first controlled evaluation study of maqama generation with LLMs, comparing five models under zero-shot, few-shot, and rule-based prompting, and evaluating outputs through both human annotation and an LLM-as-a-judge framework across dimensions such as rhetorical richness, saj density, structural coherence, and stylistic authenticity. Our results show that prompting strategy plays a strong role in stylistic quality: few-shot prompting most consistently improves saj density, while its effects on rhetoric and coherence vary by model, with the strongest models (GPT-4o and GPT-5.4-mini) benefiting most from rule-based prompting on these dimensions, though zero-shot prompting yields the highest aggregate scores across all five models. We further observe systematic differences between models in stylistic alignment with Arabic maqama conventions, and corroborate our findings with a second independent LLM judge, paired statistical significance testing, and non-LLM proxy measures of saj.
Chinese Translation
大型语言模型(LLM)已在创造性文本生成中展现出强大性能,然而其生成具有文化根基且受风格约束的文学形式的能力仍未得到充分探索。先前工作主要集中于现代语言变体和诗歌,而诸如玛卡梅之类的古典散文传统则仍大多未得到研究。玛卡梅是一种古典文学体裁,其特征为韵文(saj)、密集的修辞装饰以及片段式叙事结构,这使其成为一个具有挑战性的测试平台,用于评估 LLM 能否超越表层流畅性,迈向更深层的文学能力。在本文中,我们提出了首个关于 LLM 生成玛卡梅的受控评估研究,在零样本、少样本和基于规则的提示下比较了五个模型,并通过人工标注和 LLM 作为评判者的框架,从修辞丰富度、saj 密度、结构连贯性和风格真实性等维度评估输出。我们的结果表明,提示策略在风格质量中起着重要作用:少样本提示最一致地提高 saj 密度,而其对修辞和连贯性的影响因模型而异,其中最强的模型(GPT-4o 和 GPT-5.4-mini)在这些维度上从基于规则的提示中获益最多,尽管零样本提示在所有五个模型上产生了最高的总体得分。我们进一步观察到模型之间在与阿拉伯玛卡梅惯例的风格对齐方面存在系统性差异,并通过第二个独立的 LLM 评判者、配对统计显著性检验以及 saj 的非 LLM 代理度量来佐证我们的发现。
cs.CL / 33 / 2609.28250
Complementary Roles of Activation and Parametric Memory in Few-Shot Learning
激活记忆与参数化记忆在少样本学习中的互补作用
Miaohe Niu, Runsong Zhao, Xinyu Liu, Bo Jin, Yucheng Qiao, Chunliang Zhang, Jingbo Zhu, Tong Xiao
cs.CL
large language model
大语言模型相关
Abstract
At test time, large language models (LLMs) can encode historical information in activation memory (i.e., KV caches) and parametric memory (i.e., updated parameters). While activation memory is generally considered effective for factual recall and parametric memory for learning new tasks, their interplay remains unclear. In this work, we systematically investigate the role of memory in few-shot learning through controlled experiments. We find that activation memory is superior for recalling facts, whereas parametric memory does not consistently outperform activation memory in task learning. Moreover, our experiments show that the composite task, Conditional Arithmetic, requires the synergy of both memory types. Through neuron-level analysis, we find that the model activates distinct sets of neurons when accessing the same historical information through activation versus parametric memory. When both memory types are combined, the model recruits neurons from both sets, which is crucial for solving Conditional Arithmetic. These findings suggest that neither memory mechanism alone is sufficient for this composite task, highlighting the importance of their collaboration.
Chinese Translation
在测试时,大语言模型(LLMs)可以将历史信息编码在激活记忆(即 KV 缓存)和参数化记忆(即更新后的参数)中。尽管激活记忆通常被认为对事实回忆有效,而参数化记忆对学习新任务有效,但二者之间的相互作用仍不清楚。在本工作中,我们通过受控实验系统地研究了记忆在少样本学习中的作用。我们发现,激活记忆在回忆事实方面更优,而参数化记忆在任务学习中并不总是优于激活记忆。此外,我们的实验表明,复合任务 Conditional Arithmetic 需要两种记忆类型的协同作用。通过神经元级别的分析,我们发现,当通过激活记忆与参数化记忆访问相同的历史信息时,模型会激活不同的神经元集合。当两种记忆类型结合时,模型会调用来自这两个集合的神经元,这对于解决 Conditional Arithmetic 至关重要。这些发现表明,单独任何一种记忆机制都不足以完成这一复合任务,凸显了二者协作的重要性。
cs.CL / 34 / 2609.28272
Towards Efficient Reasoning: Learning Causal Shortcuts for Diffusion Language Models
迈向高效推理:为扩散语言模型学习因果捷径
Dian Jin, Kairong Han, Baohong Li, Xinpeng Dong, Zijing Hu, Nuanqiao Shan, Fei Wu, Kun Kuang
cs.CL
diffusion
扩散模型相关
Abstract
Diffusion Language Models (DLMs) have attracted significant attention for their strong reasoning ability. However, under a bidirectional attention mechanism, DLMs operate over an exponentially large exploration space compared to autoregressive models (ARMs), making it challenging to focus on reasoning-guiding tokens under random masking. We define causal shortcuts as token chains that cover the full sequence and provide explicit guidance towards correct reasoning trajectories. We analyze the effects of causal shortcuts on the reasoning accuracy and convergence speed of DLMs, and find that they largely improve answer convergence efficiency and generation accuracy. Motivated by this, we propose a Causal Shortcut Learning (CSL) Framework for DLMs. Specifically, we introduce a step-by-step token extraction procedure to extract causal shortcuts from data, and apply parallel prioritized masking on these tokens during training to enable efficient and accurate convergence to correct answers via causal shortcuts. Extensive experiments across multiple reasoning benchmarks and two base models demonstrate that CSL consistently outperforms existing SFT-variant baselines, achieving an average improvement of $1.92\%$ over SFT-only models, and up to $4.20\%$ on MATH-500. The code is available at the \href{https://github.com/ZJUDianJin/Causal-Shortcuts-Learning}{https://github.com/ZJUDianJin/Causal-Shortcuts-Learning
Chinese Translation
扩散语言模型(DLMs)因其强大的推理能力而受到广泛关注。然而,在双向注意力机制下,与自回归模型(ARMs)相比,DLMs 在指数级庞大的探索空间上运行,这使得在随机掩码下聚焦于引导推理的 token 变得具有挑战性。我们将因果捷径定义为覆盖完整序列并为正确推理轨迹提供明确引导的 token 链。我们分析了因果捷径对 DLMs 推理准确率和收敛速度的影响,并发现它们大幅提升了答案收敛效率和生成准确率。受此启发,我们为 DLMs 提出了因果捷径学习(CSL)框架。具体而言,我们引入逐步 token 提取过程以从数据中提取因果捷径,并在训练期间对这些 token 应用并行优先级掩码,以通过因果捷径实现高效且准确地收敛到正确答案。在多个推理基准和两个基础模型上的大量实验表明,CSL 始终优于现有的 SFT 变体基线,相较于仅 SFT 模型平均提升 $1.92\%$,在 MATH-500 上最高提升 $4.20\%$。代码可在 \href{https://github.com/ZJUDianJin/Causal-Shortcuts-Learning}{https://github.com/ZJUDianJin/Causal-Shortcuts-Learning 获取。
cs.CL / 35 / 2609.28395
Fine-Tuning LLMs for Translation: General Forgetting Mitigation Does Not Preserve MT-Specific Instruction Following
微调大语言模型用于翻译:通用的遗忘缓解并不能保持机器翻译特定的指令遵循
Niklas Scholz, David Thulke, Abdallah Nasir, Will Allred, Evgeny Matusov, Hermann Ney
cs.CL · cs.LG
large language model
大语言模型相关
Abstract
Fine-tuning large language models on parallel data improves translation quality but can cause catastrophic forgetting. Mitigation methods are generally evaluated by retention on general benchmarks. We ask whether these findings transfer to machine translation (MT) fine-tuning and to MT-specific instruction following (MT-IF): instructions that modify a translation, such as formality, grammatical gender, and length control. We compare methods anchored to auxiliary data, to model outputs, and to the base model parameters, first in a screening study with Llama 3.2 1B Instruct, then on Llama 3.1 8B Instruct fine-tuned on bidirectional Arabic-English or Spanish-English data. Elastic Weight Consolidation preserves general capabilities best in both stages; on the 8B Spanish model the average score on general benchmarks drops 1.7 points versus 11.0 for standard fine-tuning, yet its scores for formality and grammatical gender control remain close to standard fine-tuning. Only data mixing with control-task examples preserves these controls, but its gains do not transfer to unseen prompts for the same task.
Chinese Translation
在平行数据上微调大语言模型可以提升翻译质量,但可能导致灾难性遗忘。缓解方法通常通过其在通用基准上的保持能力来评估。我们提出疑问:这些发现是否能迁移到机器翻译(MT)微调以及MT特定的指令遵循(MT-IF):即对翻译进行修改的指令,例如正式程度、语法性别和长度控制。我们比较了分别锚定于辅助数据、模型输出和基础模型参数的方法,首先在Llama 3.2 1B Instruct上进行筛选研究,然后在基于双向阿拉伯语-英语或西班牙语-英语数据微调的Llama 3.1 8B Instruct上进行。弹性权重巩固(Elastic Weight Consolidation)在两个阶段都最好地保持了通用能力;在8B西班牙语模型上,通用基准的平均得分下降1.7分,而标准微调下降11.0分,但其正式程度和语法性别控制的得分仍接近标准微调。只有将数据与控制任务示例混合才能保持这些控制,但其收益并不能迁移到同一任务的未见提示上。
cs.CL / 36 / 2609.28416
Agent-Editing World Model: Rethinking World Modeling for LLM Agents
智能体编辑世界模型:重新思考面向 LLM 智能体的世界建模
Shuang Sun, Guoxin Chen, Fanzhe Meng, Jia Deng, Huatong Song, Jinhao Jiang, Wayne Xin Zhao, Hongteng Xu, Ji-Rong Wen
cs.CL · cs.AI · cs.LG
large language model
大语言模型相关
Abstract
Recent advances in large language models (LLMs) have enabled agents to tackle long-horizon tasks across diverse environments. To further improve agent performance, existing language world models typically predict environment observations, yet reconstructing high-entropy, execution-dependent tool responses offers limited value when real feedback is available. Meanwhile, agents suffer from \emph{task-state contamination}, where unsupported assumptions and outdated plans persist in history and distort subsequent decisions. We propose the \textbf{Agent-Editing World Model (AEWM)}, which models how reasoning and actions shape future task progress rather than simulating tool responses. AEWM combines \textbf{Action Judge} to distinguish \textsc{Critical}, \textsc{Exploratory}, and \textsc{Noisy} decisions with \textbf{State Revision} to edit noisy reasoning--action continuations from the same observed history. \textbf{EditAct} integrates these capabilities with real execution, directly changing the state underlying subsequent decisions rather than merely providing critiques. We train AEWM across Search, Terminal, and Software Engineering through mid-training and supervised fine-tuning. AEWM achieves 70.5\% macro-F1 on our Action Judge benchmark, exceeding the strongest frontier baseline by 10.6 points. Across six benchmarks and three agent backbones, EditAct improves average scores by 3.2--6.7 points over the strongest baseline. Furthermore, rejection sampling fine-tuning on verified EditAct trajectories, termed \textbf{AEWM-RFT}, improves over Self-RFT by 2.2--2.6 points across three domains without online AEWM guidance.
Chinese Translation
大语言模型(LLM)的最新进展使智能体能够在多样环境中处理长时程任务。为了进一步提升智能体性能,现有的语言世界模型通常预测环境观测,然而在真实反馈可用时,重建高熵、依赖执行的工具响应所提供的价值有限。与此同时,智能体遭受\emph{任务状态污染},其中无依据的假设和过时的计划滞留在历史中,并扭曲后续决策。我们提出\textbf{智能体编辑世界模型(AEWM)},它建模推理与动作如何塑造未来任务进展,而非模拟工具响应。AEWM 将用于区分 \textsc{关键}、\textsc{探索} 与 \textsc{噪声} 决策的 \textbf{动作评判器(Action Judge)} 与用于从同一观测历史中编辑含噪推理--动作延续的 \textbf{状态修订(State Revision)} 相结合。\textbf{EditAct} 将这些能力与真实执行相集成,直接改变后续决策所依据的状态,而不仅仅是提供批评。我们通过中期训练和监督微调,在搜索、终端和软件工程领域训练 AEWM。AEWM 在我们的 Action Judge 基准上达到 70.5\% 的 macro-F1,超过最强前沿基线 10.6 个点。在六个基准和三个智能体骨干网络上,EditAct 相较最强基线将平均分提高了 3.2--6.7 个点。此外,在经核验的 EditAct 轨迹上进行拒绝采样微调,称为 \textbf{AEWM-RFT},在没有在线 AEWM 引导的情况下,在三个领域中相较 Self-RFT 提升了 2.2--2.6 个点。
cs.CR / 37 / 2609.27091
Solidity Meets LLMs: A Transformer-Based Approach to Smart Contract Vulnerability Detection
Solidity 遇上 LLMs:一种基于 Transformer 的智能合约漏洞检测方法
Djamel Eddine Hakim Ghorab, Farid Mokhati, Mostafa Anouar Ghorab
cs.CR · cs.SE
large language model
大语言模型相关
Abstract
The growing adoption of blockchain technologies, particularly the Ethereum platform, has amplified the critical role of smart contracts in decentralized applications. However, the increasing complexity and financial value of these contracts make them prime targets for cyber attacks. In this work, we present a transformer-based approach for the detection of vulnerabilities in smart contract fragments written in Solidity. Leveraging the representational power of pre-trained Large Language Models (LLMs), we construct a robust pipeline that includes the definition of a ground truth dataset, labeling code fragments as vulnerable or safe. We then fine-tune a BERT-based architecture on this dataset, enabling the model to capture the syntactic and semantic patterns specific to Solidity code. Our fine-tuned model demonstrates strong performance, achieving an F1 score of 92%, and highlighting the effectiveness of LLM adaptation in enhancing smart contract security through deep contextual understanding.
Chinese Translation
区块链技术、尤其是以太坊平台的日益广泛采用,放大了智能合约在去中心化应用中的关键作用。然而,这些合约日益增加的复杂性和经济价值使其成为网络攻击的主要目标。在这项工作中,我们提出了一种基于 Transformer 的方法,用于检测以 Solidity 编写的智能合约片段中的漏洞。利用预训练大语言模型(LLMs)的表示能力,我们构建了一个稳健的流水线,其中包括定义真值数据集,将代码片段标记为有漏洞或安全。然后,我们在此数据集上微调一个基于 BERT 的架构,使模型能够捕获 Solidity 代码特有的句法和语义模式。我们微调后的模型展现出强劲性能,达到了 92% 的 F1 分数,并凸显了 LLM 适配在通过深度上下文理解增强智能合约安全性方面的有效性。
cs.CR / 38 / 2609.27155
The Like Trap: Multi-Stage Poisoning against Agents in Similarity-based Recommendation Systems
点赞陷阱:针对基于相似度的推荐系统中智能体的多阶段投毒
Yue Xing, Pengfei He, Zitao Li
cs.CR · cs.AI · cs.LG · stat.ML
large language model
大语言模型相关
Abstract
With recent advancements in large language models (LLMs) and LLM-based agents, these agents are becoming increasingly autonomous and gaining broader access to act on users' behalf on the internet. However, the vulnerability of automated agents deployed on social media platforms (e.g., for managing a user's personal account) remains underexplored. Existing studies on agent poisoning typically assume that the adversary can expose poisoned content to the agent. Although such an attack is direct and effective, it is more easily detected and mitigated. In the context of social media platforms, this leaves open whether the recommendation system itself would surface such content to the agent in a more subtle manner. Through theoretical analysis, we show that the like-score mechanism used in OASIS can be exploited, and we characterize the conditions under which a multi-stage chain of poisoned posts can steer the agent's feed. Based on these insights, we further develop an algorithm that crafts realistic poisoned posts. Experiments support our theoretical findings and demonstrate the effectiveness of the proposed algorithm. Notably, by exploiting the like-score feedback loop, the attack causes the recommendation system to select poisoned posts even when their user-post similarity falls below the retrieval threshold.
Chinese Translation
随着大型语言模型(LLMs)和基于LLM的智能体的最新进展,这些智能体正变得越来越自主,并在互联网上获得更广泛的权限来代表用户行事。然而,部署在社交媒体平台上的自动化智能体(例如,用于管理用户的个人账户)的脆弱性仍未得到充分研究。关于智能体投毒的现有研究通常假设攻击者能够将投毒内容暴露给智能体。尽管此类攻击直接且有效,但它更容易被检测和缓解。在社交媒体平台的背景下,这留下了一个未决问题:推荐系统本身是否会以更隐蔽的方式将此类内容呈现给智能体。通过理论分析,我们表明OASIS中使用的点赞分数机制可以被利用,并刻画了多阶段投毒帖子链能够操纵智能体信息流的条件。基于这些洞见,我们进一步开发了一种算法,用于构造逼真的投毒帖子。实验支持了我们的理论发现,并证明了所提出算法的有效性。值得注意的是,通过利用点赞分数反馈回路,该攻击会使推荐系统选择投毒帖子,即使其用户-帖子相似度低于检索阈值。
cs.CR / 39 / 2609.27406
Only Pay What You Must Spend: On-Demand Privacy Budget Payment for Differentially Private RAG
只支付你必须花费的:面向差分隐私 RAG 的按需隐私预算支付
Zhonghao Sun, Zhiliang Tian, Xinyue Fang, Shuo Ma, Juhua Zhang, Yiping Song, Dongsheng Li
cs.CR · cs.DL
large language model
大语言模型相关
Abstract
Deploying large language models (LLMs) on sensitive data via Retrieval-Augmented Generation (RAG) introduces severe privacy risks. Recent studies apply Differential Privacy (DP) to LLMs with RAG for formal privacy guarantees. However, existing DP-RAG frameworks rapidly exhaust the privacy budget. Although recent efforts attempt to save the budget by narrowing the retrieval scope or sparsifying private generation, these methods themselves cumulatively consume the budget, whereas they could actually rely merely on public information or at a negligible one-time privacy cost. This mismatch fails to align budget expenditure with the model's actual reliance on private data, causing substantial waste on operations that require no private access. To address this, we propose SparsePay-RAG, adopting "only pay what you must spend" as its core principle. Using public information as a zero-privacy prior, it charges the privacy budget only for the private increment. Specifically, SparsePay-RAG narrows the retrieval scope via public topic-guided clustering, adaptively controls private access frequency without privacy cost through isotonic cross-layer trajectory fitting, and compresses per-access budget via DP contrastive decoding. Under strong privacy constraints, experiments show SparsePay-RAG achieves superior privacy-utility trade-offs over baselines.
Chinese Translation
通过检索增强生成(RAG)将大型语言模型(LLM)部署在敏感数据上会引入严重的隐私风险。近期研究将差分隐私(DP)应用于结合 RAG 的 LLM,以提供形式化隐私保证。然而,现有的 DP-RAG 框架会迅速耗尽隐私预算。尽管近期工作试图通过缩小检索范围或稀疏化私有生成来节省预算,但这些方法本身会累积性地消耗预算,而它们实际上可以仅依赖公共信息,或仅以可忽略的一次性隐私代价实现。这种不匹配未能使预算支出与模型对私有数据的实际依赖程度相一致,导致在不需要私有访问的操作上造成大量浪费。为了解决这一问题,我们提出 SparsePay-RAG,并将“只支付你必须花费的”作为其核心原则。该方法使用公共信息作为零隐私先验,仅对私有增量收取隐私预算。具体而言,SparsePay-RAG 通过公共主题引导的聚类缩小检索范围,通过保序跨层轨迹拟合在不产生隐私代价的情况下自适应控制私有访问频率,并通过 DP 对比解码压缩每次访问的预算。在强隐私约束下,实验表明 SparsePay-RAG 相比基线实现了更优的隐私-效用权衡。
cs.CR / 40 / 2609.27428
A Bulletproof Business? Towards Detecting Infrastructure-as-a-Service Offerings on Telegram
一桩防弹生意?迈向检测 Telegram 上的基础设施即服务供应
Roy Ricaldi, Kristiyan Kyurkchiev, Irdin Pekaric
cs.CR
large language model
大语言模型相关
Abstract
Cybercriminal operations increasingly depend on reusable digital infrastructure---including hosting, proxies, and virtual private networks (VPNs)---rented through Cybercrime-as-a-Service markets and advertised on platforms such as Telegram. We present a taxonomy for identifying Telegram messages advertising cybercriminal Infrastructure-as-a-Service (IaaS). The taxonomy comprises six service categories across compute, network, and communication infrastructure, together with three trust attributes: Bulletproof, Payment Security, and Transparency. Using 261 human-annotated messages, we evaluate keyword-based and TF--IDF classifiers and examine prompt-based large language models as exploratory baselines. We select a TF--IDF pipeline and apply it to 1,116,071 messages from 167 cybercrime-related Telegram communities. The pipeline assigns at least one infrastructure category to 207,244 messages (18.57%) spanning 113 communities. Classified advertising is highly concentrated: a single community accounts for 50.3% of infrastructure-positive messages, while the trust-attribute classifiers identify Bulletproof claims in 37.66% of those messages. These findings characterize the scale, composition, and concentration of infrastructure advertising on Telegram and can inform the prioritization of communities and actors for monitoring and investigation.
Chinese Translation
网络犯罪活动日益依赖可复用的数字基础设施——包括托管、代理和虚拟专用网络(VPN)——它们通过网络犯罪即服务市场租用,并在 Telegram 等平台上做广告。我们提出了一套分类体系,用于识别为网络犯罪基础设施即服务(IaaS)做广告的 Telegram 消息。该分类体系包含横跨计算、网络和通信基础设施的六个服务类别,以及三个信任属性:防弹(Bulletproof)、支付安全(Payment Security)和透明度(Transparency)。我们使用 261 条人工标注的消息,评估了基于关键词的分类器和 TF--IDF 分类器,并考察了基于提示词的大语言模型作为探索性基线。我们选定了一条 TF--IDF 流水线,并将其应用于来自 167 个网络犯罪相关 Telegram 社区的 1,116,071 条消息。该流水线为跨越 113 个社区的 207,244 条消息(18.57%)分配了至少一个基础设施类别。被分类的广告高度集中:单一社区就占了基础设施阳性消息的 50.3%,而信任属性分类器在其中 37.66% 的消息中识别出了防弹(Bulletproof)宣称。这些发现刻画了 Telegram 上基础设施广告的规模、构成和集中度,并可为确定需要监测和调查的社区与行为主体的优先级提供依据。
cs.CR / 41 / 2609.27528
MDRC: A Deployable State-Recovery Defense for Traffic Signal Control under Sensor Corruption
MDRC:一种可部署的、用于传感器损坏下交通信号控制的状态恢复防御方法
Mingyuan Li, Chunyu Liu, Xiao Liu, Yanna Jiang, Guangsheng Yu, Xu Wang, Wei Ni, Ren Ping Liu
cs.CR
diffusion
扩散模型相关
Abstract
Traffic Signal Control (TSC) is a safety-critical cyber-physical system that relies on real-time sensing. Corrupted observations caused by adversarial perturbations or sensor failures can propagate from the sensing layer into the controller and degrade traffic efficiency. Existing robust Reinforcement Learning (RL)-based TSC methods often suffer from limited cross-city generalization, high inference latency, and weak recovery under partial observability. We present MDRC (Meta-Diffusion-based framework for Resilient traffic signal Control against adversarial attacks and sensor failures), a post-detection state-recovery defense inserted between sensing and control. MDRC reconstructs trustworthy traffic states before they are consumed by the controller. It combines Denoising Diffusion Implicit Models (DDIM) for efficient state recovery with Reptile meta-learning for a transferable initialization across cities. We provide an optimization-based view of the DDIM recovery dynamics and establish a recovery-error bound that separates score approximation, numerical discretization, and initialization mismatch. Across seven real-world-derived CityFlow benchmarks, MDRC reduces Average Travel Time by 6.77% under stochastic and policy-aware attacks and by 12.75% under structured sensor loss, while improving state-recovery fidelity. We further evaluate 3,600 seconds of real roadside measurements with 50% of detector channels disabled and integrate MDRC into a hardware-in-the-loop traffic-signal stack. Over a 9.16-hour run with 32,389 sensing/control cycles, the system achieves 99.79% decision availability, produces no out-of-plan recommendations, and requires approximately 38 ms of component-wise processing per one-second control interval.
Chinese Translation
交通信号控制(Traffic Signal Control, TSC)是一种依赖实时感知的安全攸关信息物理系统。由对抗性扰动或传感器故障导致的受损观测可以从感知层传播到控制器中,并降低交通效率。现有的基于鲁棒强化学习(Reinforcement Learning, RL)的 TSC 方法往往受限于跨城市泛化能力有限、推理延迟高,以及在部分可观测条件下恢复能力弱。我们提出 MDRC(Meta-Diffusion-based framework for Resilient traffic signal Control against adversarial attacks and sensor failures,即面向对抗攻击与传感器故障的弹性交通信号控制的元扩散框架),这是一种插入在感知与控制之间的检测后状态恢复防御方法。MDRC 在可信交通状态被控制器使用之前对其进行重构。它将用于高效状态恢复的去噪扩散隐式模型(Denoising Diffusion Implicit Models, DDIM)与用于获得跨城市可迁移初始化的 Reptile 元学习相结合。我们给出了 DDIM 恢复动力学的一种基于优化的视角,并建立了一个恢复误差界,该界将分数近似、数值离散化与初始化失配分离开来。在七个源自真实世界的 CityFlow 基准上,MDRC 在随机攻击和策略感知攻击下将平均行程时间(Average Travel Time)降低 6.77%,在结构化传感器失效下降低 12.75%,同时提升了状态恢复保真度。我们进一步评估了在 50% 检测器通道被禁用情况下的 3,600 秒真实路侧测量数据,并将 MDRC 集成到硬件在环的交通信号栈中。在包含 32,389 个感知/控制周期的 9.16 小时运行中,该系统实现了 99.79% 的决策可用性,未产生任何偏离计划的建议,并且每个一秒控制间隔需要约 38 ms 的分组件处理时间。
cs.CR / 42 / 2609.27996
Your Model Is Leaking: Covert Information Transfer through LLM Residual Streams
你的模型正在泄露:通过 LLM 残差流的隐蔽信息传输
Mingyuan Li, Yanna Jiang, Guangsheng Yu, Qin Wang, Xu Wang, Wei Ni, Ren Ping Liu
cs.CR
large language model
大语言模型相关
Abstract
Privacy-sensitive organizations may run large language models (LLMs) in restricted or air-gapped environments while exporting selected diagnostic artifacts. We show that a compromised runtime component can hide sensitive information in intermediate activations that are allowed to leave the restricted environment. An offline observer can recover this information with a simple linear decoder. The attack requires no model retraining or weight modification, no attacker-controlled egress, and no control over the recorder or transfer process. We introduce a residual-stream covert-channel attack that maps messages to codewords and injects them into an intermediate residual stream through a compromised runtime hook. To maintain recoverability, the injection strength is scaled with the local residual norm using the signal-to-residual-norm ratio. Across eleven models from seven architecture families, our evaluation shows 91--100% recovery on nine models with KL divergence 0.001--0.007, while evaluated activation-level detectors remain close to random guessing (AUC <= 0.56). Tested post-hoc defenses do not reliably eliminate the channel. Thus, an activation artifact can be schema-valid while carrying information that is not authorized to cross the boundary.
Chinese Translation
对隐私敏感的组织可能在受限或气隙隔离环境中运行大型语言模型(LLM),同时导出选定的诊断制品。我们表明,一个被攻陷的运行时组件可以将敏感信息隐藏在允许离开受限环境的中间激活中。离线观察者可以用一个简单的线性解码器恢复这些信息。该攻击不需要模型再训练或权重修改,不需要攻击者控制的外传通道,也不需要控制记录器或传输过程。我们提出一种残差流隐蔽信道攻击,它将消息映射为码字,并通过一个被攻陷的运行时钩子将其注入到中间残差流中。为保持可恢复性,注入强度利用信号与残差范数之比随局部残差范数缩放。在来自七个架构家族的十一个模型上,我们的评估显示,在九个模型上实现了 91--100% 的恢复,KL 散度为 0.001--0.007,而所评估的激活级检测器仍接近随机猜测(AUC <= 0.56)。测试的事后防御不能可靠地消除该信道。因此,一个激活制品可以在符合模式的同时,携带未被授权跨越边界的信息。
cs.CR / 43 / 2609.28205
GUIAuditor: Enabling Post-hoc Child Safety Forensics via Action-Guided GUI Provenance on Mobile Devices
GUIAuditor:在移动设备上通过动作引导的 GUI 溯源实现事后儿童安全取证
Junlin Liu, Yifeng Cai, Shuai Wang, Zhineng Zhong, Shaofei Li, Jiacheng Liu, Yuanchun Li, Ziqi Zhang, Xiangqun Chen, Ding Li, Yao Guo
cs.CR
large language model
大语言模型相关
Abstract
The proliferation of smart devices exposes children to online risks like grooming and financial scams that are deeply embedded within legitimate applications. Current approaches rely on automated prevention and detection, a paradigm that is fundamentally limited by its inherent fallibility. Whether rule-based or AI-driven, they inevitably produce false positives and negatives, failing to provide reliable protection. In this paper, we argue for a complementary, human-in-the-loop, post-hoc forensic paradigm. We present GUIAuditor, the first system designed to realize this vision by creating GUI Provenance: a queryable, semantic record of a child's interaction sequence. To generate this, GUIAuditor leverages a Multimodal Large Language Model (MLLM) to translate the temporal sequence of GUI events into a human-understandable narrative. To make this practical on mobile devices, a novel evidence distillation pipeline reduces the data requiring analysis by over 89.2% compared to periodic sampling approaches adopted by industry standards, with negligible impact on accuracy. On a new dataset of 295 interaction clips, GUIAuditor achieves a 95.23% Macro-F1 Score in logging significant events and, crucially, its two-stage forensic query engine successfully retrieves the correct evidence as the top result for over 90.20% of natural language questions. An end-to-end evaluation on three modern smartphones shows that the full pipeline, including on-device MLLM inference, adds 2.1W of power draw and 7.4s of per-event latency, with a peak memory footprint of ${\sim}$3.1GB. These results show that post-hoc GUI forensics can run on modern mobile devices and provide useful context for guardian-led safety review.
Chinese Translation
智能设备的激增使儿童暴露于诸如诱骗和金融诈骗等在线风险之中,这些风险深深嵌入在合法应用程序内。当前方法依赖于自动化预防和检测,这一范式因其固有的易错性而受到根本性限制。无论是基于规则还是由 AI 驱动,它们都不可避免地产生假阳性和假阴性,无法提供可靠保护。在本文中,我们主张一种互补的、人在回路中的、事后取证范式。我们提出 GUIAuditor,这是首个旨在通过创建 GUI 溯源来实现这一愿景的系统:GUI 溯源是儿童交互序列的可查询、语义化记录。为生成该记录,GUIAuditor 利用多模态大语言模型(MLLM)将 GUI 事件的时间序列转化为人类可理解的叙述。为使其在移动设备上实用,一种新颖的证据蒸馏流水线相较于行业标准采用的周期性采样方法,将需要分析的数据减少了超过 89.2%,且对准确率的影响可忽略不计。在一个包含 295 个交互片段的新数据集上,GUIAuditor 在记录重要事件方面达到 95.23% 的 Macro-F1 分数,并且至关重要的是,其两阶段取证查询引擎在超过 90.20% 的自然语言问题中成功将正确证据检索为最高结果。在三款现代智能手机上的端到端评估表明,完整流水线(包括设备端 MLLM 推理)增加了 2.1W 的功耗和每个事件 7.4s 的延迟,峰值内存占用为 ${\sim}$3.1GB。这些结果表明,事后 GUI 取证可以在现代移动设备上运行,并为监护人主导的安全审查提供有用背景。
cs.AI / 44 / 2609.27217
Learning Spectral Allocation: A Fractional Diffusion Framework for Adaptive Volumetric Segmentation
学习谱分配:面向自适应体分割的分数阶扩散框架
Yi-Hui Shen, Tie-Qiang Li
cs.CV · cs.AI
diffusion
扩散模型相关
Abstract
We address adaptive computation in 3D medical image segmentation: instead of designing another backbone, we ask how much spectral mixing each network stage needs and let optimization answer. We derive FHEAT, a two-parameter operator family, from the discrete cosine transform (DCT) solution of a fractional heat equation. A fractional order alpha and a diffusion strength D govern the operator, and at D=0 it is exactly the identity. Reparametrized by the semigroup time tau = D*alpha, same-resolution instances compose exactly, so any distribution of diffusion across same-resolution stages amounts to a single Sobolev-type regularizer of learned strength. This identity limit lets the optimizer of each layer, not the designer, decide whether global mixing is needed and how sharp it should be. We instantiate FHEAT in a lightweight U-shaped architecture (Light-UNETR) paired with a Kolmogorov-Arnold mixer (KAN3D) with adaptive rational activations, yielding FHEAT-Seg. At 5% to 20% label rates on three public benchmarks, training produces gradient-driven spectral sparsification: seven of the eight stage-level operators drive D to zero, and the survivor saturates at the sharpest low-pass (alpha ~ 0.9) in the decoder layer feeding the semi-supervised attention map. The retired layers become exact identity shortcuts at inference, cutting FLOPs from 4.29G to 0.90G (a 79% drop) at 0.975M parameters. Under a standard semi-supervised protocol, FHEAT-Seg reaches Dice scores of 90.47% (left atrium), 78.79% (Pancreas-CT), and 81.90% (BraTS 2019), ahead of five semi-supervised methods and the Light-UNETR baseline. The large variant also surpasses Light-UNETR-L under full supervision (Dice 93.09%, 85.11%, and 87.19%) with 2.851M parameters and 55.75G FLOPs. These results suggest that the allocation of spectral computation is a learnable property of optimization dynamics, not a manual design commitment.
Chinese Translation
我们研究三维医学图像分割中的自适应计算:与其再设计一个骨干网络,我们转而追问每个网络阶段需要多少谱混合,并让优化来给出答案。我们从分数阶热方程的离散余弦变换(DCT)解中推导出 FHEAT,一个双参数算子族。分数阶次 alpha 与扩散强度 D 支配该算子,而在 D=0 时它恰好是恒等映射。以半群时间 tau = D*alpha 重新参数化后,同分辨率的实例可以精确复合,因此同分辨率阶段之间扩散的任意分布都等价于一个具有可学习强度的单一 Sobolev 型正则化项。这一恒等极限让每一层的优化器——而非设计者——来决定是否需要全局混合以及它应当有多尖锐。我们将 FHEAT 实例化于一个轻量级 U 形架构(Light-UNETR)中,并与带有自适应有理激活的 Kolmogorov-Arnold 混合器(KAN3D)配对,得到 FHEAT-Seg。在三个公开基准上、5% 到 20% 的标签率下,训练产生了梯度驱动的谱稀疏化:八个阶段级算子中有七个将 D 推向零,而幸存下来的那个在馈送半监督注意力图的解码器层中饱和到最尖锐的低通(alpha ~ 0.9)。被停用的层在推理时变为精确的恒等捷径,在 0.975M 参数下将 FLOPs 从 4.29G 降至 0.90G(下降 79%)。在标准半监督协议下,FHEAT-Seg 达到 90.47%(左心房)、78.79%(Pancreas-CT)和 81.90%(BraTS 2019)的 Dice 分数,领先于五种半监督方法和 Light-UNETR 基线。其大型变体在全监督下也超越了 Light-UNETR-L(Dice 93.09%、85.11% 和 87.19%),参数量为 2.851M,FLOPs 为 55.75G。这些结果表明,谱计算的分配是优化动力学的一种可学习属性,而非人工的设计承诺。
cs.AI / 45 / 2609.28110
Field-of-View Extension in Dental Cone-Beam CT via Implicit Neural Representations and Diffusion Model-Based Refinement
基于隐式神经表示和扩散模型细化的牙科锥形束CT视野扩展
Susanne Schaub, Florentin Bieder, Matheus L. Oliveira, Yulan Wang, Buyanbileg Sodnom-ish, Dorothea Dagassan-Berndt, Michael M. Bornstein, Philippe C. Cattin
cs.CV · cs.AI
diffusion
扩散模型相关
Abstract
Dental cone-beam computed tomography (CBCT) systems often employ detector configurations that provide a truncated field of view (FOV) that only captures a small part of the patient's anatomy. In this work, we aim to reconstruct an extended FOV using projections of truncated FOV scans. To this end, we propose a three-stage framework that consists of (1) an implicit neural representation (INR) for estimating missing parts of the truncated projection data, (2) an iterative reconstruction for generating a secondary volumetric image with improved anatomical consistency and (3) a fast diffusion model for image enhancement. The proposed approach combines the strengths of continuous representations, physics-based reconstruction and generative refinement within a unified pipeline for truncated CBCT imaging. Experimental results demonstrate that the method effectively reduces truncation artifacts, improves the reconstruction of structures extending beyond the original FOV and produces images with enhanced quality. Our code is publicly available at https://github.com/SusanneSchaub/CBCT-FOV-Extension.
Chinese Translation
牙科锥形束计算机断层扫描(CBCT)系统通常采用这样的探测器配置:其提供的视野(FOV)是截断的,仅能捕获患者解剖结构的一小部分。在这项工作中,我们旨在利用截断FOV扫描的投影来重建扩展FOV。为此,我们提出了一个三阶段框架,其包括:(1) 一个隐式神经表示(INR),用于估计截断投影数据中的缺失部分;(2) 一个迭代重建,用于生成具有改进解剖一致性的次级体积图像;以及 (3) 一个快速扩散模型,用于图像增强。所提出的方法将连续表示、基于物理的重建和生成式细化的优势结合在一个用于截断CBCT成像的统一流程中。实验结果表明,该方法有效减少了截断伪影,改善了超出原始FOV的结构的重建,并生成了质量增强的图像。我们的代码已在 https://github.com/SusanneSchaub/CBCT-FOV-Extension 公开提供。
cs.LG / 46 / 2609.28473
On the Diffusibility of High-Dimensional Latents
论高维隐变量的可扩散性
Chao Feng, Zhiyang Xu, Bowei Chen, Yuanjun Xiong, Xiyao Wang, Jui-Hsien Wang, Richard Zhang, Zhe Lin, Andrew Owens, Yijun Li
cs.CV · cs.LG
diffusion
扩散模型相关
Abstract
Representation Autoencoders (RAEs) enable diffusion models to operate in the feature spaces of pretrained visual encoders. However, many off-the-shelf encoders are not optimized for faithful reconstruction, discarding fine-grained visual details. As expected, finetuning these encoders for image reconstruction recovers such details. However, perhaps counterintuitively, this procedure reduces the effective dimensionality of the resulting representation, and the altered geometry has downstream effects on generation. Specifically, we show that using the standard velocity prediction in flow matching in this high-dimensional space requires the model to fit orthogonal noise directions outside the low-dimensional signal manifold, making optimization inefficient. This motivates using the clean data parameterization ($\boldsymbol{x}_{0}$-prediction) instead, which focuses learning on the underlying signal manifold. Across experiments with multiple strong-reconstruction encoders, we show that $\boldsymbol{x}_{0}$-prediction consistently improves text-to-image generation performance.
Chinese Translation
表征自编码器(RAEs)使扩散模型能够在预训练视觉编码器的特征空间中运行。然而,许多现成的编码器并未针对忠实重建进行优化,会丢弃细粒度的视觉细节。正如预期的那样,为图像重建微调这些编码器可以恢复这些细节。然而,或许与直觉相反,这一过程会降低所得表征的有效维度,而改变后的几何结构会对生成产生下游影响。具体而言,我们表明,在这一高维空间中使用流匹配中的标准速度预测,会要求模型拟合低维信号流形之外的正交噪声方向,从而使优化效率低下。这促使我们转而使用干净数据参数化($\boldsymbol{x}_{0}$-预测),它将学习集中在底层信号流形上。在多个强重建编码器的实验中,我们表明 $\boldsymbol{x}_{0}$-预测能够持续提升文本到图像的生成性能。
cs.LG / 47 / 2609.27074
Quantifying the Occult: A Comparative Study of Hindu and Buddhist Deities Using Machine Learning Methods
量化神秘:一项使用机器学习方法的印度教与佛教神祇比较研究
Ankit Bhattacharjee
cs.CY · cs.LG
large language model
大语言模型相关
Abstract
This study introduces a dual-matrix computational architecture to mathematically quantify the morphological and theological divergence of 196 Hindu and Vajrayana Buddhist esoteric deities. Physical morphology is evaluated via a discrete Gower distance matrix enhanced by a novel "Cardinality Weighting" algorithm, while theological function is mapped via dense vector embeddings generated from Large Language Model (LLM) semantic expansions, explicitly utilized as a synthetic proxy to mitigate circular reasoning. The multi-modal topological projections provide algorithmic validation of "iconographic camouflage", demonstrating how distinct visual forms structurally obscure shared cross-tradition functions. Furthermore, I computationally model the "Atin Effect" - serving simultaneously as a psychological observation of sequential cognitive bias and a machine learning benchmark - demonstrating how high-cardinality esoteric anchors (e.g., a veena or a severed head) override systemic theological disparities to mathematically cluster orthodox and Tantric entities. Cross-tradition spatial analysis establishes that the highest esoteric manifestations, such as the Hindu Chinnamasta and the Buddhist Chinnamunda, share a near-identical mathematical coordinate across both visual ($D_G = 0.288$) and semantic ($D_C = 0.068$) boundaries, indicating a 1:1 esoteric transfer. By open-sourcing this architecture, I provide a scalable, unsupervised machine learning tool for Digital Humanities scholars and comparative theologians to rigorously map latent structural continuities across qualitative cultural corpora.
Chinese Translation
本研究引入一种双矩阵计算架构,以数学方式量化196位印度教与金刚乘佛教密教神祇在形态学和神学上的差异。物理形态通过一个离散 Gower 距离矩阵进行评估,该矩阵由一种新颖的“基数加权”(Cardinality Weighting)算法增强;而神学功能则通过由大语言模型(LLM)语义扩展生成的稠密向量嵌入进行映射,这些嵌入被明确用作一种合成代理,以缓解循环论证。多模态拓扑投影为“图像学伪装”提供了算法层面的验证,展示了不同的视觉形式如何在结构上遮蔽跨传统共有功能。此外,我对“Atin 效应”进行计算建模——它同时作为一种对序列性认知偏差的心理学观察和一项机器学习基准——展示高基数密教锚点(例如维纳琴或一颗被斩下的头)如何凌驾于系统性的神学差异之上,从而在数学上聚类正统实体与怛特罗实体。跨传统空间分析确立,最高层次的密教显现,例如印度教的 Chinnamasta 和佛教的 Chinnamunda,在视觉($D_G = 0.288$)与语义($D_C = 0.068$)两种边界上共享一个几乎相同的数学坐标,表明存在 1:1 的密教传递。通过将这一架构开源,我为数字人文学者和比较神学家提供了一种可扩展的无监督机器学习工具,以严谨地映射跨越定性文化语料库的潜在结构连续性。
cs.CL / 48 / 2609.27014
ContraVis: Evidence-Grounded Visual Analytics for Contradiction Review in Legal Contracts
ContraVis:面向法律合同矛盾审查的基于证据的可视分析
Luis Sante, Paula Lima, Mariana Rocha, Jorge Poco
cs.HC · cs.CL
large language model
大语言模型相关
Abstract
Legal contracts are structurally complex documents in which contradictions may emerge across distant and interconnected provisions. Although large language models (LLMs) improve legal language understanding, contradiction analysis remains a human-centered and evidence-grounded review task. We present ContraVis, a visual analytics system for human-in-the-loop contradiction analysis in legal contracts. The system models contracts as typed paragraph graphs that combine explicit contractual references with semantic relationships between paragraphs. This graph plays a dual role: it conditions LLM reasoning and serves as the interactive representation the analyst explores, keeping model context and human inspection aligned across coordinated views. In a controlled comparison, graph-conditioned reasoning recovered more injected contradictions than standalone LLM analysis as contract length grew, while surfacing additional candidates for analyst validation. A formative study with contract-domain lawyers indicated that in-context evidence comparison supported contradiction validation, and we distill design implications for evidence-grounded, LLM-assisted document review.
Chinese Translation
法律合同是结构复杂的文档,其中矛盾可能出现在相距遥远且相互关联的条款之间。尽管大语言模型(LLMs)提升了对法律语言的理解,但矛盾分析仍然是一项以人为中心且以证据为基础的审查任务。我们提出 ContraVis,一个用于法律合同中人在回路矛盾分析的可视分析系统。该系统将合同建模为带类型的段落图,这种图将显式合同引用与段落之间的语义关系相结合。该图发挥双重作用:它为 LLM 推理提供条件,并作为分析师所探索的交互式表示,从而在协调视图之间保持模型上下文与人工检查一致。在一项受控比较中,随着合同长度增长,图条件化推理比独立 LLM 分析找回了更多注入的矛盾,同时呈现出供分析师验证的额外候选。一项与合同领域律师进行的形成性研究表明,上下文中的证据比较支持了矛盾验证,并且我们提炼出针对基于证据、LLM 辅助文档审阅的设计启示。
cs.AI / 49 / 2609.27250
The Risk-Sensitive Schrödinger Bridge: Is Not a KL Projection
风险敏感型薛定谔桥:不是 KL 投影
Hamidreza Behjoo
cs.IT · cs.AI
diffusion
扩散模型相关
Abstract
The Schrödinger bridge owes its computational power to a single structural fact: by Girsanov's theorem the controlled problem is a Kullback--Leibler (KL) projection onto a fixed reference measure, solvable by alternating projections. This letter shows that the fact does not survive risk sensitivity. When the expected path cost is replaced by the entropic risk measure and both endpoint marginals are kept as hard constraints, the resulting fixed-point bridge value $J_θ$ (the soft-problem value at the multiplier that enforces the terminal constraint) admits no representation as a constrained KL minimum against any fixed path-space reference with a regular endpoint law (a class strictly larger than the uniformly elliptic diffusion references: no Markov property is required), even allowing an additive normalisation depending on the initial marginal. Moreover, no single reference generates the one-parameter family in the risk parameter. The obstruction is computed in closed form: the Gaussian bridge value violates, by exactly $θ/2$, a heat equation that any Gaussian smoothing of a fixed endpoint density must obey. In place of the projection, the theory rests on a terminal-multiplier fixed point and an asymmetric factorisation penalising the score energy of the backward factor.
Chinese Translation
薛定谔桥的计算能力归功于一个单一的结构性事实:根据 Girsanov 定理,受控问题是对一个固定参考测度的 Kullback--Leibler(KL)投影,可通过交替投影求解。本文表明,这一事实在风险敏感性下不再成立。当期望路径代价被熵风险度量替代,并且两个端点边缘分布都保持为硬约束时,所得的不动点桥值 $J_θ$(在强制终端约束的乘子处的软问题值)不能被表示为相对于任何具有正则端点律的固定路径空间参考的受约束 KL 最小值(这一类严格大于一致椭圆扩散参考:不要求马尔可夫性),即使允许一个依赖于初始边缘分布的加性归一化。此外,没有任何单一参考能生成以风险参数为参数的单参数族。该障碍以闭式形式计算:高斯桥值恰好以 $θ/2$ 违反任何固定端点密度的高斯平滑都必须满足的热方程。代替投影,该理论依赖于一个终端乘子不动点以及一个非对称分解,该分解惩罚后向因子的得分能量。
cs.LG / 50 / 2609.27067
ChipMEM: Verification-Grounded Memory for EDA Agents
ChipMEM:面向 EDA 智能体的验证锚定记忆
Abdulrahman AlRabah, Joshua Mabry, Dilek Hakkani-Tür, Abdussalam Alawini, Hamid Shojaei, Kartik Hegde, Sandesh Adhikary
cs.LG · cs.CL
large language model
大语言模型相关
Abstract
Large language model (LLM)-based agents use Electronic Design Automation (EDA) tools to generate and revise register-transfer-level (RTL) designs under synthesis and verification feedback. Recent methods learn from this feedback by distilling reusable skills from execution traces or by training on rewards derived from EDA-tools. Both methods are typically evaluated on the tasks that produced the experience. Repeated access to benchmark feedback on the same task can reward task-specific revision rather than creating reusable knowledge that transfers. We introduce ChipMEM, a verification-grounded memory layer for EDA agents. It combines cross-task procedural memory with within-trajectory statistical guidance. Its procedural component distills and stores a skill only after it passes synthesis, simulation, or formal checks, rather than relying on model self-assessments. A Bayesian component maintains hierarchical Beta estimates over tool-call outcomes and ranks recovery strategies that succeeded under comparable errors. A common adapter applies the same memory interface to RTL optimization and testbench-generation agents while preserving each domain's tools and acceptance criteria. We measure performance on training tasks and evaluate whether learned skills transfer to unseen tasks. On RTLRewriter-Bench, under matched model and tool settings, ChipMEM produces equivalence-passing outputs on 39/54 scored designs versus 35/54 without memory; on the 49-design short suite, mean area improvement is 8.69% versus 5.66%. On held-out CVDP tasks, ChipMEM with a frozen procedural library achieves 20/20 accepted outcomes versus 18/20 without memory in a single evaluation per setting.
Chinese Translation
基于大语言模型(LLM)的智能体使用电子设计自动化(EDA)工具,在综合与验证反馈下生成和修改寄存器传输级(RTL)设计。近期方法通过从执行轨迹中蒸馏可复用技能,或通过在源自 EDA 工具的奖励上进行训练,来从这种反馈中学习。这两类方法通常都在产生经验的任务上进行评估。对同一任务上的基准反馈的重复访问,可能会奖励任务特定的修改,而不是创造可迁移的可复用知识。我们提出 ChipMEM,一种面向 EDA 智能体的验证锚定记忆层。它将跨任务程序性记忆与轨迹内统计指导相结合。其程序性组件仅在某项技能通过综合、仿真或形式化检查之后才蒸馏并存储该技能,而不是依赖模型自我评估。一个贝叶斯组件维护对工具调用结果的分层 Beta 估计,并对在可比错误下成功过的恢复策略进行排序。一个通用适配器将相同的记忆接口应用于 RTL 优化和测试平台生成智能体,同时保留每个领域的工具和接受标准。我们在训练任务上测量性能,并评估学习到的技能是否能迁移到未见任务。在 RTLRewriter-Bench 上,在匹配的模型和工具设置下,ChipMEM 在 39/54 个计分设计上产生通过等价性检查的输出,而无记忆时为 35/54;在 49 个设计的短套件上,平均面积改进为 8.69% 对 5.66%。在留出的 CVDP 任务上,带有冻结程序性库的 ChipMEM 在每种设置的单次评估中实现了 20/20 个被接受的结果,而无记忆时为 18/20。
cs.LG / 51 / 2609.27166
Scaling of Capability and Efficiency at Inference Time in Large Reasoning Models
大型推理模型在推理阶段的能力与效率的缩放规律
Moritz Laber, Zohair Shafi, Germans Savcisens, Brennan Klein, Matteo Chinazzi, Samuel V. Scarpino, Albert-László Barabási, Tina Eliassi-Rad
cs.LG · physics.soc-ph
large language model
大语言模型相关
Abstract
Capability and efficiency are two key dimensions of reasoning in large language models (LLMs). Capability refers to the ability to solve a given problem correctly, whereas efficiency refers to the ability to do so with limited resources. When LLMs use Chain-of-Thought (CoT) reasoning to solve problems of controlled hardness, both the number of problems solved correctly and the number of tokens required to reach a correct answer depend on problem hardness and model size. However, how these factors jointly shape capability and efficiency remains poorly understood. Here, we use hierarchical Bayesian models to evaluate the capability and efficiency of LLMs from the DeepSeek-R1-Distill model family across four classes of arithmetic and algorithmic reasoning problems. At a fixed model size, the probability of correctly solving an instance decays approximately exponentially with instance size, our proxy for problem hardness. The decay scale grows sublinearly with model size, indicating that larger models are more capable, but that capability gains diminish with scale. Output length grows as a power law with instance size, which serves as a proxy for difficulty. However, the parameters of this power law do not vary systematically with model size, suggesting that larger models do not become more efficient. Together, these findings reveal potential limitations of naive scaling as a strategy for developing more capable AI systems: capability improves with diminishing returns, while efficiency shows little to no improvement.
Chinese Translation
能力与效率是大语言模型(LLMs)推理的两个关键维度。能力指正确解决给定问题的能力,而效率指在有限资源下做到这一点的能力。当LLMs使用思维链(CoT)推理解决难度受控的问题时,正确解决的问题数量以及得出正确答案所需的token数量都取决于问题难度和模型规模。然而,这些因素如何共同塑造能力与效率,仍鲜为人知。在此,我们使用分层贝叶斯模型,评估来自DeepSeek-R1-Distill模型家族的LLMs在四类算术与算法推理问题上的能力与效率。在固定模型规模下,正确解决某个实例的概率随实例规模近似呈指数衰减,而实例规模是我们对问题难度的代理指标。该衰减尺度随模型规模亚线性增长,表明更大的模型能力更强,但能力增益随规模递减。输出长度随实例规模呈幂律增长,而实例规模可作为难度的代理指标。然而,该幂律的参数并不随模型规模系统性变化,这表明更大的模型并未变得更高效。综合来看,这些发现揭示了朴素缩放作为开发能力更强的人工智能系统策略的潜在局限:能力提升伴随收益递减,而效率则几乎没有或完全没有改善。
cs.LG / 52 / 2609.27248
Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders
将预训练大语言模型改造为高保真连续文本自编码器
Arkanath Pathak, Unnat Jain, Alexander C. Berg
cs.LG
diffusion
扩散模型相关
Abstract
Next-token prediction has enabled highly fluent autoregressive language models, but it represents global structure only indirectly through sequential factorization. In contrast, high-fidelity autoencoders have become a standard primitive in image generation, enabling generative models to operate over continuous latent spaces; text lacks a comparably faithful continuous representation. We propose LLMAE, a method for repurposing a pretrained decoder-only language model as a continuous text autoencoder by exposing an intermediate fixed-length latent bottleneck within its internal activations. Instantiated with a parameter-efficient 270M Gemma 3 model, LLMAE uses structured attention masks, LoRA adaptation, and KL regularization to learn an autoencoding interface that leverages the generative prior of the original LLM. We train LLMAE to reconstruct text sequences up to 1024 tokens, significantly improving on this task to achieve near-perfect reconstruction. Furthermore, we demonstrate the downstream utility of this representation by training a latent text diffusion model for detailed image captioning using the learned LLMAE autoencoder. By mapping text into a fixed-length continuous latent space, our approach provides an effective substrate for downstream adaptation while benefiting from the fluency of the original LLM.
Chinese Translation
下一词元预测使高度流畅的自回归语言模型成为可能,但它仅通过序列分解间接地表示全局结构。相比之下,高保真自编码器已成为图像生成中的标准基础组件,使生成模型能够在连续潜空间上运行;而文本则缺乏一种与之相当忠实的连续表示。我们提出 LLMAE,一种通过在其内部激活中暴露出一个中间定长潜瓶颈,将预训练的仅解码器语言模型改造为连续文本自编码器的方法。LLMAE 以一个参数高效的 270M Gemma 3 模型进行实例化,采用结构化注意力掩码、LoRA 适配和 KL 正则化来学习一个自编码接口,从而利用原始 LLM 的生成先验。我们训练 LLMAE 重建最长 1024 个词元的文本序列,在该任务上显著改进,实现了近乎完美的重建。此外,我们通过使用所学得的 LLMAE 自编码器训练一个用于详细图像描述的潜文本扩散模型,展示了该表示的下游实用性。通过将文本映射到定长连续潜空间,我们的方法为下游适配提供了一个有效的基底,同时受益于原始 LLM 的流畅性。
cs.LG / 53 / 2609.27306
Discrete Diffusion Models via Evolving Variational Autoregressive Networks
基于演化变分自回归网络的离散扩散模型
Kewen Pan, Ying Tang
cs.LG · cond-mat.dis-nn · stat.ML
diffusion
扩散模型相关
Abstract
Conventional score-based diffusion models learn scores without representing normalized densities, whereas tractable normalized models support both sampling and direct likelihood evaluation. A recent tensor-network approach provides such a representation but is largely restricted to low-dimensional lattices. Here we introduce a discrete diffusion model that parameterizes normalized probability distributions using variational autoregressive networks. Explicit Markov jump operators govern the forward noising and reverse denoising dynamics, extending discrete diffusion models with normalized distributions to spin systems on higher-dimensional lattices. We apply this framework to the two- and three-dimensional Ising models across ordered, critical, and disordered regimes, accurately computing thermodynamic quantities including free energy, energy, and magnetization. We further integrate the framework with Monte Carlo sampling, using adaptive diffusion steps to maintain high acceptance rates even at low temperatures while enhancing sample diversity. These results establish a neural-network framework for the discrete diffusion model with normalized probability distributions.
Chinese Translation
传统的基于分数的扩散模型学习分数而不表示归一化密度,而可处理的归一化模型同时支持采样和直接似然评估。最近的一种张量网络方法提供了这样的表示,但在很大程度上仅限于低维晶格。在此,我们引入一种离散扩散模型,它使用变分自回归网络来参数化归一化概率分布。显式的马尔可夫跳跃算子支配前向加噪和反向去噪动力学,将具有归一化分布的离散扩散模型扩展到高维晶格上的自旋系统。我们将该框架应用于二维和三维 Ising 模型,涵盖有序、临界和无序区域,准确计算包括自由能、能量和磁化强度在内的热力学量。我们进一步将该框架与蒙特卡洛采样相结合,使用自适应扩散步数,即使在低温下也能保持高接受率,同时增强样本多样性。这些结果建立了一个用于具有归一化概率分布的离散扩散模型的神经网络框架。
cs.LG / 54 / 2609.27572
DCRL: Decoupling and Coupling Reinforcement Learning via Policy-Reward Manifold Alignment
DCRL:通过策略-奖励流形对齐的解耦与耦合强化学习
Henan Sun, Zehua Li, Haitao Hu, Qifan Zhang, Jianfeng Zhang, Nuo Chen, Jia Li
cs.LG · cs.AI
large language model
大语言模型相关
Abstract
Reinforcement learning (RL) has emerged as a key paradigm for improving the reasoning capabilities of large language models (LLMs). However, existing reward systems, such as rule-based and reward-model-based, often exhibit issues such as unstable optimization and reward hacking. In this work, we revisit the general reasoning of LLMs from a geometric perspective, conceptualizing it as a coupled manifold composed of three interdependent sub-manifolds: logical deduction, evaluation, and representation. Based on this perspective, response generation in RL can be interpreted as a decoupling process from the evaluation manifold, while reward estimation corresponds to a decoupling process from the logical deduction manifold. The limitations of rule-based and reward-model RL systems can be geometrically interpreted as the mismatch of policy-reward manifolds during RL process. To address the aforementioned misalignment, we propose Decoupling and Coupling Reinforcement Learning (DCRL) framework, which incorporates two key components: (1) a syllogistic logic-based prompt evolution mechanism that dynamically refines reward rubrics to enhance the expressiveness of the reward manifold; and (2) a policy-reward re-coupling mechanism that jointly updates the reward and policy models, ensuring consistent evaluation and mitigating manifold mismatch during training. Theoretical analysis and extensive experiments across multiple reasoning domains demonstrate that DCRL consistently outperforms both rule-based and reward-model baselines. Notably, a Qwen3-4B model trained under DCRL surpasses a Qwen3-32B baseline and approaches the performance of a Qwen3-235B model, highlighting superior effectiveness and generalization in RL.
Chinese Translation
强化学习(RL)已成为提升大语言模型(LLMs)推理能力的关键范式。然而,现有的奖励系统,例如基于规则的和基于奖励模型的,常常表现出优化不稳定和奖励黑客等问题。在这项工作中,我们从几何视角重新审视LLMs的一般推理,将其概念化为一个由三个相互依赖的子流形组成的耦合流形:逻辑演绎、评估和表示。基于这一视角,RL中的响应生成可以解释为从评估流形解耦的过程,而奖励估计对应于从逻辑演绎流形解耦的过程。基于规则和基于奖励模型的RL系统的局限性可以从几何上解释为RL过程中策略-奖励流形的不匹配。为了解决上述失配,我们提出了解耦与耦合强化学习(DCRL)框架,该框架包含两个关键组成部分:(1)一种基于三段论逻辑的提示演化机制,它动态优化奖励评分标准以增强奖励流形的表达能力;(2)一种策略-奖励重耦合机制,它联合更新奖励模型和策略模型,确保评估一致性并缓解训练期间的流形不匹配。理论分析和跨多个推理领域的广泛实验表明,DCRL持续优于基于规则和基于奖励模型的基线。值得注意的是,在DCRL下训练的Qwen3-4B模型超过了Qwen3-32B基线,并接近Qwen3-235B模型的性能,凸显了RL中卓越的有效性和泛化能力。
cs.LG / 55 / 2609.27633
Pheno-GS: Phenoscape-scale Geodesic Sinkhorn
Pheno-GS:Phenoscape 尺度的测地 Sinkhorn
Alistair Wilkinson, Christopher J. Tape, Smita Krishnaswamy
cs.LG · q-bio.QM
diffusion
扩散模型相关
Abstract
High-throughput single-cell data is now collected across large patient cohorts. Understanding patient-level heterogeneity from cellular-level data motivates phenoscaping: embedding each single-cell distribution as a "datapoint," with distances given by optimal transport (OT). Computing geometry-aware OT at this scale, between all pairs of patient datasets, remains an open challenge, since existing methods either rely on Euclidean ground metrics that distort manifold structure or fail under sparse, unevenly sampled, or large-scale data. We present \textbf{Pheno-GS} (Phenoscape-scale Geodesic Sinkhorn), which computes accurate, scalable geodesic transport distances under noisy, unbalanced, large-scale settings via three components: ($1$) graph connectivity regularization for well-defined geodesics on sparse/disconnected manifolds; ($2$) an unbalanced OT formulation via KL marginal penalties; and ($3$) a batched matrix algorithm computing all pairwise distances in one heat diffusion (over $200 \times$ faster than Geodesic Sinkhorn for $500$ distributions). We validate Pheno-GS on synthetic benchmarks and a CyTOF perturbation dataset.
Chinese Translation
高通量单细胞数据如今已在大型患者队列中收集。从细胞水平数据理解患者水平异质性促使了表型图谱化:将每个单细胞分布嵌入为一个“数据点”,其距离由最优传输(OT)给出。在这一规模上,计算所有患者数据集对之间的几何感知 OT 仍然是一个开放挑战,因为现有方法要么依赖会扭曲流形结构的欧几里得基础度量,要么在稀疏、采样不均匀或大规模数据下失效。我们提出 \textbf{Pheno-GS}(Phenoscape 尺度的测地 Sinkhorn),它通过三个组件在噪声、不平衡、大规模设置下计算准确且可扩展的测地传输距离:($1$) 图连通性正则化,用于在稀疏/不连通流形上获得良定义的测地线;($2$) 通过 KL 边缘惩罚的不平衡 OT 形式;以及 ($3$) 一种批处理矩阵算法,在一次热扩散中计算所有成对距离(对于 $500$ 个分布,比 Geodesic Sinkhorn 快 $200 \times$ 以上)。我们在合成基准和 CyTOF 扰动数据集上验证了 Pheno-GS。
cs.LG / 56 / 2609.27657
FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation
FLEET:从 Logits 熵到文本生成中的增强轨迹
Oleksii Streltsov, Oleksandra Vitko
cs.LG · cs.AI · cs.CL
large language model
大语言模型相关
Abstract
Solutions based on large language models (LLMs) often rely on temperature sampling to improve accuracy and stability by aggregating multiple samples from the completion distribution. However, this memoryless approach is inherently suboptimal: because it lacks awareness of prior generations and their evaluations, it produces an increasing proportion of semantically duplicate answers as more samples are drawn, leading to diminishing returns. To address this limitation, we introduce FLEET, a novel method that integrates a memory mechanism into the generation process. FLEET represents each generation as a sparse trajectory through states whose entropy exceeds a predefined threshold and uses these trajectories to infer per-token utility scores that adjust the logits. Benchmark evaluations demonstrate that FLEET achieves the same accuracy as the repeated sampling baseline, with a 3x speedup, and substantially improves accuracy on complex coding tasks (LiveCodeBench Pass@32 increases from 59.9% to 66.2%) under the same budget. Furthermore, in the greedy-decoding configuration evaluated here, the approach is deterministic and uses a single calibration pass to derive its principal hyperparameters, requiring only minimal modifications to existing LLM pipelines.
Chinese Translation
基于大型语言模型(LLMs)的解决方案通常依赖温度采样,通过聚合来自补全分布的多个样本以提高准确性和稳定性。然而,这种无记忆方法本质上并非最优:由于它缺乏对先前生成及其评估的感知,随着抽取的样本越来越多,它会产生比例不断增加的语义重复答案,导致收益递减。为了解决这一局限,我们引入 FLEET,一种将记忆机制集成到生成过程中的新方法。FLEET 将每次生成表示为穿越熵超过预定义阈值的状态的稀疏轨迹,并利用这些轨迹推断出调整 logits 的逐 token 效用分数。基准评估表明,FLEET 在达到与重复采样基线相同准确率的同时实现了 3 倍加速,并在相同预算下大幅提升了复杂编码任务的准确率(LiveCodeBench Pass@32 从 59.9% 提高到 66.2%)。此外,在此处评估的贪婪解码配置中,该方法具有确定性,并使用单次校准过程来推导其主要超参数,只需对现有 LLM 流水线进行最小修改。
cs.LG / 57 / 2609.27658
Private Decentralized Optimization with Noise Reduction and Bias Correction
具有噪声降低与偏差校正的私有去中心化优化
Yizhao Fan, Wenjian Luo, Jiaojiao Zhang
cs.LG
diffusion
扩散模型相关
Abstract
Private decentralized learning is affected by sampling noise, privacy noise, and decentralized bias under heterogeneous data. We propose Private Recursive Decentralized Optimization (PRDO). PRDO uses recursive estimation with same-batch gradient differences to reduce estimation errors caused by sampling and privacy noise, while its Exact Diffusion component corrects decentralized bias arising from data heterogeneity. Our analysis establishes a nonconvex convergence bound without assuming uniformly bounded data heterogeneity across nodes. It further gives a sufficient condition under which recursive gradient differences yield strictly lower query sensitivity than private Exact Diffusion, together with an example that rigorously satisfies this condition. Experiments show improved accuracy over the evaluated baselines.
Chinese Translation
私有去中心化学习在异构数据下受到采样噪声、隐私噪声和去中心化偏差的影响。我们提出私有递归去中心化优化(PRDO)。PRDO 使用带有同批次梯度差分的递归估计来降低由采样噪声和隐私噪声引起的估计误差,而其 Exact Diffusion 组件则校正由数据异构性引起的去中心化偏差。我们的分析建立了一个非凸收敛界,且无需假设各节点间的数据异构性一致有界。它进一步给出了一个充分条件,在该条件下,递归梯度差分比私有 Exact Diffusion 产生严格更低的查询敏感度,并给出了一个严格满足该条件的例子。实验表明,与所评估的基线相比,准确率有所提高。
cs.LG / 58 / 2609.27735
NS-ATTENTION: Newton-Schulz Transformations of Attention Outputs in Vision Transformers
NS-ATTENTION:视觉 Transformer 中注意力输出的牛顿-舒尔茨变换
Xiaohe Jiang, Guoqiang Zhang, Tianjin Huang, Ronghui Mu
cs.LG
large language model
大语言模型相关
Abstract
Newton-Schulz (NS) iteration has recently been used in the Muon optimizer to transform update matrices during the training of large language models. Motivated by its spectral effect, we investigate applying NS directly to Transformer attention representations. We introduce Newton-Schulz Attention (NS-Attn.), a parameter-free transformation applied to the output of each attention head. Each head output is arranged as a feature-by-token matrix and normalized by its Frobenius norm. We then apply a finite NS polynomial step and restore the original norm. The objective is to reduce spectral concentration and increase effective rank before standard head merging and output projection. Across ViT and Swin on CIFAR-10 and CIFAR-100, NS-Attn. improves final-epoch accuracy in all 12 matched-seed comparisons, with mean gains of 0.25--0.83 percentage points. ViT ablations show higher mean accuracy with one iteration than with two. Spectral analysis further shows reduced leading-eigenvalue concentration and increased effective rank. These gains incur additional inference latency.
Chinese Translation
牛顿-舒尔茨(NS)迭代最近被用于 Muon 优化器中,以在大语言模型训练期间对更新矩阵进行变换。受其谱效应的启发,我们研究了将 NS 直接应用于 Transformer 注意力表示。我们提出了牛顿-舒尔茨注意力(NS-Attn.),这是一种应用于每个注意力头输出的无参数变换。每个头的输出被排列为一个特征×token 的矩阵,并按其特征的 Frobenius 范数进行归一化。随后我们应用一个有限步的 NS 多项式,并恢复原始范数。其目标是在标准的多头合并与输出投影之前,降低谱集中度并提高有效秩。在 CIFAR-10 和 CIFAR-100 上的 ViT 与 Swin 中,NS-Attn. 在所有 12 组匹配随机种子的比较中都提升了最终轮次的准确率,平均增益为 0.25--0.83 个百分点。ViT 消融实验表明,一次迭代比两次迭代具有更高的平均准确率。谱分析进一步显示,主特征值的集中度降低,有效秩提高。这些收益会带来额外的推理延迟。
cs.LG / 59 / 2609.27982
Riemannian Structure and Optimization for a Class of Low-Parametric Orthogonal Matrices
一类低参数化正交矩阵的黎曼结构与优化
Ali Aliev, Maxim Rakhuba
cs.LG · cs.AI · math.DG · math.NA
large language model
大语言模型相关
Abstract
In this paper, we are concerned with matrices formed by block-diagonal factors interleaved with fixed permutations -- a flexible family of structured matrices. This class has recently drawn interest in deep learning architectures for its balanced expressivity-efficiency trade-off, yet efficient computational strategies for working with it remain to be found. We approach this problem through Riemannian geometry and examine under what conditions this class admits a smooth manifold structure. For the practically important case of orthogonal two-factor matrices, we derive the essential Riemannian tools and propose efficient algorithms for their implementation. The algorithms leverage automatic differentiation, support parameter sharing within each factor, and avoid explicit dense matrix construction. We test them within the Riemannian optimization framework on the best matrix approximation problem and for parameter-efficient fine-tuning of large language models. Beyond the two-factor setting, we study the geometric and matrix-theoretic properties of factorizations with a larger number of block-diagonal factors.
Chinese Translation
在本文中,我们关注由块对角因子与固定置换交错形成的矩阵——一类灵活的结构化矩阵。这类矩阵最近因其在表达能力与效率之间的平衡而在深度学习架构中引起了兴趣,然而处理它的高效计算策略仍有待发现。我们通过黎曼几何来处理这一问题,并考察在什么条件下该类矩阵具有光滑流形结构。对于具有实际重要性的正交双因子矩阵情形,我们推导了必要的黎曼工具,并为其实现提出了高效算法。这些算法利用自动微分,支持每个因子内部的参数共享,并避免显式构造稠密矩阵。我们在黎曼优化框架内,在最佳矩阵逼近问题以及大语言模型的参数高效微调上对它们进行测试。在双因子设定之外,我们研究了具有更多块对角因子的分解的几何性质和矩阵论性质。
cs.LG / 60 / 2609.27987
PCQC: Privileged Counterfactual Question Credit for Multi-Turn Medical Dialogue
PCQC:面向多轮医疗对话的特权反事实问题信用
Chenxuan Li, Jiayi Wan, Xinrong Chen, Zhongyu Zhao, Xuecheng Shang, Peixing Wan
cs.LG
large language model
大语言模型相关
Abstract
Large language models (LLMs) have made substantial progress on medical question-answering, yet effective medical dialogue also requires learning to ask questions that uncover relevant patient information. To train such dialogue policies, a common pipeline combines supervised fine-tuning with reinforcement learning (RL) based on final diagnostic correctness. However, this outcome-based supervision does not directly distinguish the contributions of individual questions and provides no question-level feedback for unexecuted alternatives. To address this gap, we introduce PCQC (Privileged Counterfactual Question Credit), which uses privileged patient information during training to learn from questions never asked. During training, PCQC makes alternative questions directly comparable at the same dialogue state by using privileged patient facts to construct their answers. A frozen diagnostic scorer evaluates the diagnostic utility of each resulting question-answer pair by how strongly it supports the correct diagnosis. PCQC turns these comparisons into relative question credit that teaches the policy which questions to favor, directly supervising both executed and unexecuted questions alongside outcome-based RL without requiring complete rollouts for the unexecuted alternatives. Extensive experiments across four medical benchmarks demonstrate that PCQC achieves 63.10% mean diagnostic accuracy, outperforming GRPO and ATPO by 4.38 and 4.21 percentage points, respectively. These gains are achieved with 33.1% fewer inquiry turns than GRPO.
Chinese Translation
大语言模型(LLMs)在医学问答方面已取得显著进展,但有效的医疗对话还要求学习提出能够揭示相关患者信息的问题。为了训练此类对话策略,一种常见流程将监督微调与基于最终诊断正确性的强化学习(RL)相结合。然而,这种基于结果的监督并不会直接区分各个问题的贡献,也不会为未执行的备选问题提供问题级别的反馈。为弥补这一空白,我们提出 PCQC(Privileged Counterfactual Question Credit,特权反事实问题信用),它利用训练期间的特权患者信息来从从未被提出的问题中学习。在训练期间,PCQC 通过使用特权患者事实来构造备选问题的答案,使这些备选问题在同一对话状态下可直接比较。一个冻结的诊断评分器根据每个由此产生的问题-答案对支持正确诊断的强度,评估其诊断效用。PCQC 将这些比较转化为相对问题信用,教会策略应偏好哪些问题,从而与基于结果的 RL 一起直接监督已执行和未执行的问题,而无需为未执行的备选问题执行完整的 rollout。在四个医学基准上的大量实验表明,PCQC 达到 63.10% 的平均诊断准确率,分别比 GRPO 和 ATPO 高出 4.38 和 4.21 个百分点。与 GRPO 相比,这些增益是在询问轮数减少 33.1% 的情况下实现的。
cs.LG / 61 / 2609.28263
Resource-Adaptive Stochastic Gradient Descent for Online Linear Programming without Re-solving
资源自适应随机梯度下降用于无需重新求解的在线线性规划
Jiameng Lyu
cs.LG · math.OC
large language model
大语言模型相关
Abstract
The growth of large language model (LLM) inference and search services increases the scale of online linear programming problems, motivating computationally efficient algorithms. We develop resource-adaptive stochastic gradient descent (RASGD) for stochastic online linear programming. The algorithm uses one request and current inventory to update resource prices, requiring O(m) operations for m resources and memory per arrival and no LP or sample-average optimization. The central idea is to express the current-resource pricing logic of re-solving through a first-order SGD update: each arrival refreshes the remaining-inventory allowance in the dual objective, while the stepsize decreases for early learning and increases later to match the speed of inventory adjustment. Under standard non-degeneracy conditions, our algorithm is feasible on every sample path and achieves O(\log T) expected regret against the realized fractional hindsight optimum, which matches the lower bound, even for policies that know the distribution and have unrestricted computation. The analysis converts curvature around the fixed reference price into inventory stability without tracking optimal prices at changing resource levels. Numerical experiments show that RASGD achieves regret competitive with per-arrival LP re-solving and improves upon the tested first-order baselines, while retaining the computational efficiency of first-order methods. These results establish RASGD as a computationally efficient approach to achieving high allocation quality in large-scale OLP.
Chinese Translation
大语言模型(LLM)推理和搜索服务的增长扩大了在线线性规划问题的规模,推动了计算高效算法的需求。我们针对随机在线线性规划开发了资源自适应随机梯度下降(RASGD)。该算法利用一个请求和当前库存来更新资源价格,对于 $m$ 种资源每次到达仅需 $O(m)$ 次操作和内存,且无需 LP 或样本平均优化。其核心思想是通过一阶 SGD 更新来表达重新求解的当前资源定价逻辑:每次到达都会刷新对偶目标中的剩余库存允许量,同时步长在早期学习时减小、后期增大以匹配库存调整的速度。在标准非退化条件下,我们的算法在每条样本路径上均可行,并且相对于已实现的分数事后最优解达到 $O(\log T)$ 期望遗憾,这匹配了下界,即使对于知道分布且计算不受限制的策略也是如此。该分析将固定参考价格周围的曲率转化为库存稳定性,而无需在变化的资源水平下跟踪最优价格。数值实验表明,RASGD 达到了与逐到达 LP 重新求解相当的遗憾,并优于所测试的一阶基线,同时保持了一阶方法的计算效率。这些结果确立了 RASGD 作为一种在大规模 OLP 中实现高分配质量的计算高效方法。
cs.MA / 62 / 2609.27386
From Intents to Algorithms: Verified Algorithm Discovery for Transport Networks
从意图到算法:面向传输网络的经过验证的算法发现
Behnam Ojaghi, Ricard Vilalta, Raul Muñoz
cs.NI · cs.MA
large language model
大语言模型相关
Abstract
Intent-based networking decouples desired outcomes from device-level configuration, but most systems still map intents to parameters of an algorithm selected in advance. Large language models (LLMs) create an opportunity to automate algorithm design, yet unrestricted generated code is unsuitable for transport-network control because feasibility, reproducibility, and robustness must be enforced independently of the model. We present VERA-TN, a verification-guided framework that compiles a network intent into a bounded algorithm-design specification. The target architecture uses an LLM as a semantic variation operator over typed request-ordering and path-ranking programs; generated logic remains separated from a trusted allocator that enforces path validity, latency, capacity, and single-path constraints. We prove feasibility preservation under explicit assumptions and establish a sufficient bound for the lexicographic latency tie-break in the exact reference model. The released proof-of-concept instantiates the same interface with a bounded ten-parameter numerical candidate and deterministic replay, rather than a completed live-LLM/AST study. Across 150 certified held-out cases on a 28-node TEFNET24-derived hierarchy, evolutionary search reaches a mean priority-utility ratio of 0.958, compared with 0.952 for equal-budget random search and 0.940 for priority-greedy routing. The gain over random search is small but statistically detectable (Holm- adjusted p = 0.0083). The candidate does not improve congestion relative to MILP-C, and the effect of failure-aware training is inconclusive at the 0.05 level (p = 0.051). Eight discovery runs on the official national topology and replay on 12 unseen metro-regional topologies show no stable intent-specific specialization. These results support the trust-boundary and numerical-evolution claims but do not establish a benefit from LLM generation.
Chinese Translation
基于意图的网络将期望结果与设备级配置解耦,但大多数系统仍将意图映射到预先选定的算法的参数上。大语言模型(LLMs)为自动化算法设计创造了机会,但不受限的生成代码不适合传输网络控制,因为可行性、可复现性和鲁棒性必须独立于模型来强制执行。我们提出 VERA-TN,这是一个验证引导的框架,它将网络意图编译为有界的算法设计规范。目标架构将 LLM 用作针对带类型的请求排序和路径排序程序的语义变异算子;生成的逻辑与受信任的分配器保持分离,该分配器强制执行路径有效性、时延、容量和单路径约束。我们在显式假设下证明了可行性保持,并在精确参考模型中为字典序时延决胜建立了一个充分界。发布的概念验证使用一个有界的十参数数值候选和确定性回放来实例化相同的接口,而不是一项已完成的实时 LLM/AST 研究。在一个由 TEFNET24 衍生的 28 节点层级结构上的 150 个经认证的留出案例中,进化搜索达到 0.958 的平均优先级效用比,而等预算随机搜索为 0.952,优先级贪心路由为 0.940。相对于随机搜索的增益很小,但在统计上可检测到(Holm 校正 p = 0.0083)。该候选方案相对于 MILP-C 并未改善拥塞,而故障感知训练的效果在 0.05 水平上无定论(p = 0.051)。在官方国家拓扑上的八次发现运行以及在 12 个未见过的都市区域拓扑上的回放显示,没有稳定的意图特定专门化。这些结果支持信任边界和数值进化方面的主张,但并未确立来自 LLM 生成的收益。
cs.AI / 63 / 2609.27070
The Gaussian Is Enough: Flow-Matching Priors Do Not Help When Fine-Tuning Large Behavior Models
高斯已足够:微调大型行为模型时,流匹配先验并无帮助
Chen Xu, Rishi Shah, Hadas Kress-Gazit, Haruki Nishimura, Masha Itkina
cs.RO · cs.AI
diffusion
扩散模型相关
Abstract
Modern robot imitation learning increasingly relies on generative policies based on diffusion or flow-matching models, which generate actions by transforming samples from a prior distribution. A key question is whether the choice of prior matters. Replacing the standard Gaussian with a closer-to-target, non-Gaussian prior has been shown to substantially improve performance when training from scratch. A natural next step is to ask whether these gains transfer to fine-tuning pretrained Large Behavior Models (LBMs) such as LBM 1.0, $π_{0.5}$, and GR00T~N1.5, where one might expect even larger gains. Surprisingly, we find that this is not the case, except possibly at very low fine-tuning data fractions. Across over 100K simulation rollouts spanning all three aforementioned LBMs on 40+ tasks in two simulation platforms, and 1250 hardware rollouts on five bimanual manipulation tasks, non-Gaussian priors that are demonstrably closer to the target yield statistically indistinguishable or worse fine-tuning performance than a standard Gaussian prior. Diagnostic analyses suggest why: fine-tuned imitation learning policies converge to similar action predictions across priors, despite their fine-tuned encoder embeddings diverging substantially from the pretrained embeddings and each other. A learning-rate ablation further confirms that encoder training is the dominant factor in fine-tuning performance, substantially outweighing the effect of prior choice. We conclude with concrete directions for future research on when and why learned priors might still matter in fine-tuning. Project page: https://cxu-tri.github.io/non_gaussian_FT/
Chinese Translation
现代机器人模仿学习越来越依赖于基于扩散模型或流匹配模型的生成式策略,这些策略通过变换来自先验分布的样本来生成动作。一个关键问题是先验的选择是否重要。在从零开始训练时,用更接近目标的非高斯先验替换标准高斯先验已被证明能显著提升性能。自然的下一步是询问这些收益是否能迁移到微调预训练的大型行为模型(LBMs)上,例如 LBM 1.0、$π_{0.5}$ 和 GR00T~N1.5,在这些模型上人们甚至可能期望获得更大的收益。令人惊讶的是,我们发现情况并非如此,除非可能是在极低的微调数据比例下。在超过 100K 次模拟 rollout 中,涵盖上述三个 LBMs 在两种仿真平台上的 40+ 个任务,以及在五个双臂操作任务上的 1250 次硬件 rollout,那些可证明更接近目标的非高斯先验,其微调性能与标准高斯先验在统计上无法区分,或者更差。诊断分析提示了原因:微调后的模仿学习策略在不同先验之间收敛到相似的动作预测,尽管它们微调后的编码器嵌入与预训练嵌入以及彼此之间都大幅偏离。学习率消融进一步确认,编码器训练是微调性能的主导因素,其影响大大超过先验选择的影响。我们最后给出未来研究的具体方向,以探讨学习到的先验在微调中何时以及为何仍可能重要。项目页面:https://cxu-tri.github.io/non_gaussian_FT/
cs.AI / 64 / 2609.28247
Controlling Collectives of AI Agents in Reasoning Space with Spatial Transformers
在推理空间中使用空间变换器控制AI智能体集群
Frederic Vatnsdal, Roshan Gopal, Romina Garcia Camargo, Vijay Kumar, Alejandro Ribeiro
cs.RO · cs.AI
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) introduce an exciting new paradigm for planning and navigation in robotics, but fail on even simple multi-robot tasks as team sizes grow. We propose COMPASS, a scalable, decentralized multi-robot architecture for controlling large collectives of agentic robots with reasoning space feedback control. Feedback is generated locally on each robot by a spatial transformer which aggregates multi-hop messages across the fleet into a learned feedback token. Our experiments find that collectives of language models demonstrate performance gains from structured diversity of the input command, which can cancel biases; an advantage that is held across scale. Compared against a centralized frontier LLM policy and a language-only communication ablation, we find that the coupled design of COMPASS decisively produces cohesive flocking formations that accurately fly the commanded intent. We show that reasoning feedback works best when composed with a compact learned token. Our ablations show that hand engineered feedback with raw state appearing in the language channel obliterates cohesion. COMPASS generalizes zero-shot to unseen instructions of ambiguous meaning while commanding flocks up to 16 times its training scale, flying up to 1024 robots under natural language commands.
Chinese Translation
大语言模型(LLMs)为机器人学中的规划与导航引入了一种令人兴奋的新范式,但随着团队规模的增长,它们即使在简单的多机器人任务上也会失败。我们提出COMPASS,一种可扩展的、去中心化的多机器人架构,用于通过推理空间反馈控制来操控大规模的智能体机器人集群。反馈由空间变换器在每台机器人上本地生成,该变换器将整个机群中的多跳消息聚合为一个学习得到的反馈令牌。我们的实验发现,语言模型集群从输入指令的结构化多样性中获得性能提升,这种多样性可以抵消偏差;这一优势在不同规模下都能保持。与集中式前沿LLM策略以及仅语言通信的消融实验相比,我们发现COMPASS的耦合设计决定性地产生了具有凝聚力的集群编队,能够准确地按照所指令的意图飞行。我们表明,当推理反馈与一个紧凑的学习令牌相组合时效果最佳。我们的消融实验表明,将原始状态出现在语言通道中的手工设计反馈会彻底破坏凝聚力。COMPASS能够零样本泛化到含义模糊的未见指令,同时指挥规模达到其训练规模16倍的集群,在自然语言指令下飞行多达1024台机器人。
cs.SE / 65 / 2609.27030
Kubernetes Misconfigurations in the Wild: Taxonomy, Evolution, and Automated Repair with Large Language Models
野外环境中的 Kubernetes 错误配置:分类体系、演化与基于大语言模型的自动修复
Mostafa Anouar Ghorab, Ahmad Abdel Latif, Mohamed Aymen Saied
cs.SE
large language model
大语言模型相关
Abstract
Kubernetes is widely used to orchestrate cloud-native applications, yet its declarative configuration model often introduces security misconfigurations that threaten system reliability. Despite available detection tools, misconfiguration patterns and scalable remediation remain insufficiently understood. This paper presents an empirical study of Kubernetes security misconfigurations based on 2,662 developer-reported Stack Overflow issues. We derive a taxonomy of recurring security weaknesses across configuration objects and categories. We analyze severity variations and investigate how misconfigurations evolve between incubator and stable project stages. Findings show that while some operational issues decrease as projects mature, critical security misconfigurations often persist or reappear. We then evaluate Large Language Models (LLMs) for automated remediation under progressively enriched contextual conditions. Contextual grounding improves correction accuracy, with the best standalone model achieving 89.06%. To enhance structural correctness and schema compliance, we introduce Kubecurity, a schema-guided validation framework based on official Kubernetes specifications. Combining contextual LLM reasoning with deterministic schema enforcement achieves 98.50% correction accuracy while substantially reducing newly introduced misconfigurations. This work advances the understanding of Kubernetes security misconfigurations and demonstrates a hybrid approach to more reliable automated remediation.
Chinese Translation
Kubernetes 被广泛用于编排云原生应用,然而其声明式配置模型常常引入威胁系统可靠性的安全错误配置。尽管已有可用的检测工具,但错误配置模式与可扩展的修复仍然未得到充分理解。本文基于 2,662 个开发者报告的 Stack Overflow 问题,对 Kubernetes 安全错误配置进行了实证研究。我们推导出一个跨配置对象与类别的反复出现的安全弱点分类体系。我们分析严重性差异,并研究错误配置在孵化器阶段与稳定项目阶段之间如何演化。研究发现,尽管某些运维问题随着项目成熟而减少,关键安全错误配置往往持续存在或再次出现。随后,我们在逐步增强的上下文条件下评估大语言模型(LLMs)用于自动修复的效果。上下文锚定提升了纠正准确率,最佳独立模型达到 89.06%。为增强结构正确性与模式合规性,我们引入 Kubecurity,一个基于官方 Kubernetes 规范的模式引导验证框架。将上下文 LLM 推理与确定性模式强制执行相结合,实现了 98.50% 的纠正准确率,同时大幅减少新引入的错误配置。这项工作推进了对 Kubernetes 安全错误配置的理解,并展示了一种通往更可靠自动修复的混合方法。
cs.SE / 66 / 2609.27214
Verified Learning for Compiler Optimization: An LLM-Guided Architecture with Formal Control
面向编译器优化的可验证学习:一种具有形式化控制的 LLM 引导架构
Dev Pratap Singh, Rong Feng, Suman Saha
cs.SE
large language model
大语言模型相关
Abstract
Compiler optimizations traditionally rely on handcrafted heuristics that often fail to generalize across programs and architectures. We investigate whether large language models can participate in compiler optimization through a verification-centered systems architecture that couples generative rewriting with formal equivalence checking. Using lazification in LLVM IR as a case study, we fine-tune a code-centric LLM on transformations produced by Wyvern and embed Alive2 into a feedback loop that enforces semantic preservation for every generated rewrite. Correctness is enforced externally as a runtime control layer rather than learned implicitly. During inference, candidate transformations are symbolically validated and regenerated when necessary, ensuring accepted rewrites satisfy formal constraints. On the LLVM test suite, the fine-tuned model reproduces core optimization behaviors while applying fewer transformations overall. Although Wyvern remains faster on most benchmarks, 9.8% achieve comparable or improved runtime under the learned system, with no semantic violations observed. Verification overhead remains bounded and convergence stable. These results demonstrate that generative AI components can be safely integrated into compiler pipelines through deterministic validation and structured feedback, offering a scalable architectural pattern for trustworthy AI-driven software infrastructure.
Chinese Translation
编译器优化传统上依赖于手工设计的启发式方法,这些方法往往难以在不同程序和架构之间泛化。我们研究大型语言模型能否通过一种以验证为中心的系统架构参与编译器优化,该架构将生成式重写与形式化等价性检查相耦合。以 LLVM IR 中的惰性化(lazification)作为案例研究,我们在 Wyvern 所产生的变换上微调一个以代码为中心的大型语言模型,并将 Alive2 嵌入到一个反馈循环中,该循环对每一次生成的重写强制保证语义保持。正确性以运行时控制层的形式在外部被强制保证,而非被隐式地学习得到。在推理阶段,候选变换会被符号化地验证,并在必要时重新生成,从而确保被接受的重写满足形式化约束。在 LLVM 测试套件上,经过微调的模型复现了核心优化行为,同时总体上应用的变换更少。尽管 Wyvern 在大多数基准测试上仍然更快,但在该学习系统下仍有 9.8% 达到相当或更优的运行时,且未观察到任何语义违规。验证开销保持有界,收敛保持稳定。这些结果表明,生成式 AI 组件可以通过确定性验证和结构化反馈被安全地集成到编译器流水线中,为可信的 AI 驱动软件基础设施提供了一种可扩展的架构模式。
cs.SE / 67 / 2609.28449
Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark
大型语言模型能否推理运行时行为?一个仓库级动态基准测试
Hamed Taherkhani, Mohammad Abdollahi, Melika Sepidband, Hridya Dhulipala, Tien N. Nguyen, Hadi Hemmati
cs.SE · cs.AI · cs.CL
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly used in coding tasks, but their ability to reason about code execution remains unclear. Existing repository-level QA benchmarks mainly evaluate static code understanding and often rely on LLM-based evaluation, while execution-reasoning benchmarks are mostly limited to snippets or functions. We introduce SWE-Flux, a repository-level benchmark for dynamic execution reasoning containing 480 execution-grounded instances across 12 real Python repositories, with gold answers automatically harvested from instrumented test executions rather than written manually or judged by LLMs. The benchmark covers singletest and multi-test questions over control flow, loops, program state, dataflow, exceptions, and program invariants. Evaluating five LLMs shows that this task remains challenging. The best model achieves only 37% accuracy. Models perform better on localized behavior such as invariants, intra-procedural control flow, exceptions, and simple loops, but struggle with dataflow, inter-procedural execution, precise state reasoning, and suite-level aggregation. Finally, we show that the oracle-harvesting pipeline can generate fresh benchmark variants using input perturbation. It successfully harvests valid variants for almost 90% of the selected instances, and the resulting variants are substantially more challenging for the evaluated models.
Chinese Translation
大型语言模型(LLM)越来越多地用于编码任务,但它们推理代码执行的能力仍不清楚。现有的仓库级问答基准测试主要评估静态代码理解,并且通常依赖基于LLM的评估,而执行推理基准测试大多局限于代码片段或函数。我们引入了SWE-Flux,一个用于动态执行推理的仓库级基准测试,包含跨12个真实Python仓库的480个基于执行的实例,其黄金答案是从插桩的测试执行中自动收集的,而不是手动编写或由LLM评判的。该基准测试涵盖了关于控制流、循环、程序状态、数据流、异常和程序不变量的单测试和多测试问题。评估五个LLM表明,这项任务仍然具有挑战性。最佳模型仅达到37%的准确率。模型在诸如不变量、过程内控制流、异常和简单循环等局部行为上表现更好,但在数据流、过程间执行、精确状态推理和套件级聚合方面表现不佳。最后,我们表明,oracle收集流水线可以使用输入扰动生成新的基准测试变体。它成功地为近90%的选定实例收集了有效变体,并且生成的变体对评估的模型来说挑战性显著更大。
cs.AI / 68 / 2609.28372
Shopping by algorithm: How agentic AI deploys human heuristics as a surrogate consumer
算法购物:智能体 AI 如何作为代理消费者部署人类启发式
Davood Wadi, Yu Ma
econ.GN · cs.AI
large language model
大语言模型相关
Abstract
Consumers increasingly delegate purchasing decisions to Large Language Models (LLMs) acting as surrogate consumers. Using "Tool-Lab," an adaptation of information-board process tracing that places product attributes behind costly tool calls, we examine how marketing pricing cues (i.e., just-below pricing and promotional framing) influence AI shopping agents. Across eight commercially deployed LLMs from three providers, we trace pre-choice information acquisition. Under zero cost, pricing cues rarely mislead. Imposing acquisition costs under a vague goal prompt leads LLMs to omit diagnostic attributes required to compute unit price and choose suboptimal choices resembling human heuristics. Relative to a specific goal prompt that mainly preserves diagnostic search and choice optimality, a vague goal prompt under constraints creates a search-mediated vulnerability. This research demonstrates that marketing heuristics in delegated AI shopping are governed by storefront information architecture, not necessarily immutable LLM flaws.
Chinese Translation
消费者越来越多地将购买决策委托给充当代理消费者的大型语言模型(LLMs)。使用“Tool-Lab”——一种信息板过程追踪的改编方法,它将产品属性置于需要付出成本的工具调用之后——我们考察了营销定价线索(即,略低于整数定价与促销框架)如何影响 AI 购物智能体。在来自三个提供商的八个商业部署的 LLM 中,我们追踪了选择前的信息获取。在零成本下,定价线索很少产生误导。在模糊目标提示下施加获取成本,会导致 LLM 省略计算单位价格所需的诊断性属性,并做出类似人类启发式的次优选择。相对于主要保持诊断性搜索与选择最优性的具体目标提示,在约束条件下的模糊目标提示会制造出一种由搜索中介的脆弱性。本研究表明,在委托式 AI 购物中,营销启发式受店面信息架构支配,而不必然是 LLM 不可改变的缺陷。
cs.AI / 69 / 2609.27337
Evolving Inspectable O-RAN Slicing xApps with LLMs
利用 LLM 演化可检查的 O-RAN 切片 xApps
Faezeh Dehghan Tarzjani, Bhaskar Krishnamachari
eess.SY · cs.AI · cs.LG
large language model
大语言模型相关
Abstract
Open RAN (O-RAN) slicing xApps must adapt resource allocations to changing channel conditions and traffic demands while meeting service-level agreements (SLAs). Deep reinforcement learning can produce adaptive policies, but their allocation rules remain encoded in neural-network parameters. Our goal is to retain this adaptability while making the controller's decision logic directly inspectable and editable by operators. We use a large language model (LLM) to evolve slicing controllers as compact Python programs whose decision logic remains readable and editable after optimization. The LLM proposes and revises candidates offline, while a calibrated simulator scores them, and the selected decision module runs unchanged in the O-RAN control path. On the NSF POWDER 5G testbed, the evolved controller releases resources from a guaranteed slice whose throughput target becomes unattainable under a sustained channel fade, improving best-effort throughput from 158.2 to 228.6 Mbps, a 44.5% gain over the best static allocation. Since the controllers are readable source code, their behavior can be predicted from their equations, defects can be diagnosed by reading the code, and calibration errors can be corrected with one-line edits, reducing SLA misses from 79.9% to 2.2% in one case and more than doubling fitness in another. In a four-slice trace-driven simulation calibrated to the same testbed, evolutionary search achieves higher average evaluation scores than independent prompting at a matched proposal budget, with mean normalized gains on held-out traces of 16.3% for prompting alone, 32.1% for evolution from scratch, and 51.0% for evolution from a starting program.
Chinese Translation
开放无线接入网(O-RAN)切片 xApp 必须使资源分配适应不断变化的信道条件和流量需求,同时满足服务级别协议(SLA)。深度强化学习可以产生自适应策略,但其分配规则仍被编码在神经网络参数之中。我们的目标是在保留这种自适应性的同时,使控制器的决策逻辑能够被运营商直接检查与编辑。我们使用大语言模型(LLM)将切片控制器演化为紧凑的 Python 程序,其决策逻辑在优化之后仍保持可读、可编辑。LLM 离线提出并修订候选方案,而经过校准的仿真器对它们进行评分,被选中的决策模块则不加改动地在 O-RAN 控制路径中运行。在 NSF POWDER 5G 测试平台上,演化得到的控制器从一条保证切片中释放资源,该切片的吞吐量目标在持续信道衰落之下变得无法达成,从而将尽力而为吞吐量从 158.2 Mbps 提升到 228.6 Mbps,相比最佳静态分配获得 44.5% 的增益。由于控制器是可读的源代码,其行为可以从其方程中预测,缺陷可以通过阅读代码来诊断,校准误差可以通过一行修改来纠正,在一个案例中将 SLA 未达标率从 79.9% 降至 2.2%,在另一个案例中使适应度提高了一倍以上。在针对同一测试平台校准的四切片轨迹驱动仿真中,在匹配的提议预算下,演化搜索取得的平均评估得分高于独立提示,在留出轨迹上的平均归一化增益为:仅提示为 16.3%,从零开始演化为 32.1%,从起始程序出发演化为 51.0%。
cs.LG / 70 / 2609.27008
Sharp Convergence of Wasserstein Gradient Flows for Spectrally Nonnegative Interaction Energies
谱非负相互作用能量的 Wasserstein 梯度流的尖锐收敛
Zhengjiang Lin, Philippe Rigollet
math.AP · cs.LG · math.OC
diffusion
扩散模型相关
Abstract
We study the long-time behavior of Wasserstein gradient flows for interaction energies \[ \mathsf E[μ] = \frac12\iint_{M\times M}K(x,y)\,\mathrm dμ(x)\,\mathrm dμ(y) \] on a closed manifold $M$. For kernels diagonal in a Laplace eigenbasis with nonnegative spectral coefficients, we prove a differential inequality relating the relative entropy to the energy gap. Consequently, for any nonnegative initial density $u_0\in L^p(M)$, $p>1$, the energy gap is integrable in time and satisfies \[ \mathsf E[μ_t]-\mathsf E_{\min}=o(t^{-1}). \] If all spectral coefficients are positive, the flow converges weakly to the constant measure. These interaction energies need not be geodesically convex in Wasserstein space, and the associated flows contain no diffusion; their global convergence therefore does not follow from standard Wasserstein gradient flow theory. The kernels covered by our results include zonal kernels on spheres, kernels arising in transformer models, regularized Riesz kernels, and inverse fractional Laplacian kernels. We also investigate the sharpness of the $o(t^{-1})$ rate. For any smooth kernel in this class with infinitely many positive spectral coefficients and any $δ>0$, we construct a solution of the linearized flow whose energy is comparable to $t^{-1-δ}$ along a sequence of times tending to infinity. Moreover, for any $δ>0$, by choosing a suitable inverse fractional Laplacian kernel on the flat torus, we construct an exact solution of the nonlinear Wasserstein gradient flow whose energy is comparable to $t^{-1-δ}$. The nonlinear construction is based on uniform-in-time estimates for the evolution of the dyadic Fourier coefficient blocks and a blockwise energy-persistence argument. These estimates also yield a uniform-in-time quantitative comparison between the nonlinear Wasserstein gradient flow and its linearization.
Chinese Translation
我们研究闭流形 $M$ 上相互作用能量 \[ \mathsf E[μ] = \frac12\iint_{M\times M}K(x,y)\,\mathrm dμ(x)\,\mathrm dμ(y) \] 的 Wasserstein 梯度流的长时间行为。对于在 Laplace 特征基中为对角且具有非负谱系数的核,我们证明了一个将相对熵与能量间隙联系起来的微分不等式。因此,对于任意非负初始密度 $u_0\in L^p(M)$,$p>1$,能量间隙在时间上可积,并满足 \[ \mathsf E[μ_t]-\mathsf E_{\min}=o(t^{-1}). \] 如果所有谱系数都为正,则该流弱收敛到常值测度。这些相互作用能量在 Wasserstein 空间中不必是测地凸的,且相关流不包含扩散;因此它们的全局收敛并不能由标准 Wasserstein 梯度流理论推出。我们的结果所涵盖的核包括球面上的 zonal 核、Transformer 模型中出现的核、正则化 Riesz 核以及逆分数阶 Laplacian 核。我们还研究了 $o(t^{-1})$ 速率的尖锐性。对于此类中任意具有无穷多个正谱系数的光滑核,以及任意 $δ>0$,我们构造线性化流的一个解,其能量沿趋于无穷的时间序列与 $t^{-1-δ}$ 同阶。此外,对于任意 $δ>0$,通过在平坦环面上选取合适的逆分数阶 Laplacian 核,我们构造非线性 Wasserstein 梯度流的一个精确解,其能量与 $t^{-1-δ}$ 同阶。该非线性构造基于对二进 Fourier 系数块演化的时间一致估计,以及一个逐块能量持久性论证。这些估计还给出了非线性 Wasserstein 梯度流与其线性化之间的时间一致定量比较。
cs.LG / 71 / 2609.28338
Local Geometric Mixing via Dobrushin Contraction with Applications to Diffusion Path Monte Carlo and the Proximal Sampler
通过 Dobrushin 收缩的局部几何混合及其在扩散路径蒙特卡洛和近端采样器中的应用
Stefan Oberdörster
math.PR · cs.LG · stat.CO · stat.ML
diffusion
扩散模型相关
Abstract
Local geometric mixing localizes geometric mixing by requiring geometric convergence to equilibrium in total variation only over finitely many transitions. It accommodates local convergence rates and captures rapid local equilibration, even when global mixing is much slower. We establish and discuss local geometric mixing bounds through Dobrushin contraction. We then apply this approach to Diffusion Path Monte Carlo, a recently proposed Markov chain Monte Carlo method, aimed at leveraging advances in score-based modeling, whose ideal transitions coincide with those of the Proximal Sampler. Our analysis covers both the ideal method and its implementable Metropolis-adjusted counterpart, providing mixing guarantees under minimal assumptions. For the ideal method, these guarantees complement recent spectral gap estimates, which we develop into mixing time bounds.
Chinese Translation
局部几何混合通过要求仅在有限多次转移上在全变差意义下几何收敛到平衡,从而将几何混合局部化。它容许局部收敛速率,并刻画快速的局部平衡化,即使全局混合要慢得多。我们通过 Dobrushin 收缩建立并讨论局部几何混合界。然后,我们将这一方法应用于扩散路径蒙特卡洛,这是一种最近提出的马尔可夫链蒙特卡洛方法,旨在利用基于分数建模的进展,其理想转移与近端采样器的理想转移一致。我们的分析既涵盖理想方法,也涵盖其可实现的 Metropolis 调整对应方法,并在最小假设下提供混合保证。对于理想方法,这些保证补充了最近的谱隙估计,我们将后者发展为混合时间界。
cs.LG / 72 / 2609.27546
Robustness of Diffusion Models under Distribution Shift
分布偏移下扩散模型的鲁棒性
Wei Luo, Neil K. Chada, Shijie Zhang, Lu Yu
math.ST · cs.LG · math.PR · stat.ML
diffusion
扩散模型相关
Abstract
Score-based diffusion models are increasingly considered in settings where the underlying data distribution may differ from the training distribution, yet existing theoretical guarantees largely focus on the no-shift setting. In this work, we study robust score estimation under Wasserstein perturbations of a reference distribution. For the Ornstein--Uhlenbeck diffusion, we show that robust estimation decomposes into two fundamental components: the statistical cost of learning the reference distribution and the intrinsic cost of distribution shift. The latter scales quadratically with the Wasserstein radius, and this dependence is minimax optimal. We construct an explicit finite-sample estimator achieving the resulting robust minimax rate without knowing the shift radius. When the reference distribution lies on an unknown low-dimensional subspace, the statistical term adapts to the intrinsic dimension while the shift cost remains unchanged. Finally, we show that the same decomposition governs positive-time reverse sampling and obtain matching minimax guarantees in KL divergence. Together, these results characterize how finite data, intrinsic dimension, and distribution shift affect the robustness of score-based diffusion models.
Chinese Translation
基于分数的扩散模型越来越多地被考虑在底层数据分布可能与训练分布不同的场景中,然而现有的理论保证主要集中于无偏移设定。在这项工作中,我们研究参考分布的 Wasserstein 扰动下的鲁棒分数估计。对于 Ornstein--Uhlenbeck 扩散,我们表明鲁棒估计可分解为两个基本组成部分:学习参考分布的统计代价和分布偏移的内在代价。后者与 Wasserstein 半径的平方成比例,并且这种依赖关系是极小极大最优的。我们构造了一个显式的有限样本估计器,在不知道偏移半径的情况下达到所得的鲁棒极小极大速率。当参考分布位于一个未知的低维子空间上时,统计项会适应内在维度,而偏移代价保持不变。最后,我们表明同样的分解支配正时间反向采样,并在 KL 散度下获得匹配的极小极大保证。综合起来,这些结果刻画了有限数据、内在维度和分布偏移如何影响基于分数的扩散模型的鲁棒性。
cs.AI / 73 / 2609.27729
AI-Driven Neural Surrogates for In Silico Design of Cognitive-Affective Neuromodulation Targets
用于认知-情感神经调控靶点计算机模拟设计的AI驱动神经代理模型
Marco Rothermel, Madleen Stenger, Soroush Daftarian, Svenja Jule Francke, Bita Shariatpanahi, José C. García Alanis, Mohammad-Ali Nikouei Mahani, Stefan G. Hofmann, Tim Hahn, Hamidreza Jamalabadi
q-bio.NC · cs.AI · eess.SY
diffusion
扩散模型相关
Abstract
In neuropsychiatry, the primary goal is often not only to decode brain activity but to change it, for example to lessen a negative affective bias or an overly salient memory. Motivated by control theory, we develop an AI-driven neural-surrogate framework that proposes candidate representational changes and tests their predicted perceptual effects from snapshots of stimulus-evoked fMRI activity, without physical stimulation. The framework combines fMRI decoding, deep generative modeling, and constrained latent-space steering. Valence and memorability are used only as worked examples. Using more than 36,000 image-fMRI observations from four deeply sampled Natural Scenes Dataset participants, subject-specific models recovered coarse generative structure from visually responsive cortex (two-way identification, 0.79-0.88; chance, 0.5). Graded perturbations were reconstructed as images and evaluated with automated scorers and human ratings from 7,200 trials by 18 participants. In the primary VDVAE model, valence shifted from -0.61 to +1.03 SD and memorability from -1.34 to +1.45 SD; a later Versatile Diffusion refinement reduced or altered these effects. Across five perturbation levels, human valence ratings moved in the predicted direction under the linear time-correction model (mean slope, 0.038 SD per unit of alpha; 95 percent CI, 0.003-0.074; positive in 16 of 18 participants). Perceived memorability did not change reliably. Baseline agreement with the automated assessor was suggestive for valence (r = 0.30) and weak for memorability (r = 0.10). Extreme perturbations drifted from the original stimulus, so intended change must be weighed against loss of fidelity. These findings provide a falsifiable upstream method for designing and behaviorally testing candidate representational targets for future neuromodulation in psychiatry, while marking the limits of the present static approximation.
Chinese Translation
在神经精神病学中,主要目标往往不仅是解码大脑活动,而且是改变它,例如减轻负性情感偏差或过于显著的记忆。受控制理论启发,我们开发了一个AI驱动的神经代理框架,该框架提出候选表征变化,并在无需物理刺激的情况下,从刺激诱发fMRI活动的快照中检验其预测的知觉效应。该框架结合了fMRI解码、深度生成建模和受约束的潜在空间引导。效价和可记忆性仅用作工作示例。使用来自四名深度采样的自然场景数据集参与者的超过36,000个图像-fMRI观测,受试者特异性模型从视觉反应皮层中恢复了粗略的生成结构(双向识别,0.79-0.88;随机水平,0.5)。分级扰动被重建为图像,并用自动评分器和来自18名参与者的7,200次试验的人类评分进行评估。在主要的VDVAE模型中,效价从-0.61 SD变为+1.03 SD,可记忆性从-1.34 SD变为+1.45 SD;后来的Versatile Diffusion细化减少或改变了这些效应。在五个扰动水平上,在线性时间校正模型下,人类效价评分朝着预测方向移动(平均斜率,每单位alpha为0.038 SD;95%置信区间,0.003-0.074;在18名参与者中的16名中为正)。感知到的可记忆性没有可靠变化。与自动评估器的基线一致性对效价具有提示性(r = 0.30),对可记忆性则较弱(r = 0.10)。极端扰动偏离了原始刺激,因此必须将预期变化与保真度损失进行权衡。这些发现提供了一种可证伪的上游方法,用于设计和行为测试未来精神病学神经调控的候选表征靶点,同时标明了当前静态近似的局限。
人工智能 (cs.AI)
75
cs.AI / 1 / 2609.26891
Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity
Zhening Li, Joshua Liu, Mateja Vukelic, Nicole Shen, Supriya Lall, Amitayush Thakur, Alex Zhang, Omar Khattab, Jonathan Light, Armando Solar-Lezama
cs.AI
Abstract
Modern language-model agents are built around the \textit{agent loop}, where the LLM is placed in an environment exposing a set of tools, and the LLM has full control over the workflow by alternating between tool calls and observing their output. However, certain workflows currently require additional engineering beyond the agent loop itself, such as memory systems and self-improving systems. We built an LLM agent framework, JAZ, to explore the extent to which a minimal harness that is little more than the agent loop itself can accomplish tasks these specialized systems are built for. JAZ exposes a single LLM-based primitive invoke and provides a set of built-in hooks that allow the programmer to apply constraints and monitoring. Generalizing existing code-mode agent loops, \texttt{invoke} is the simplest loop that satisfies two defining properties: (1) the LLM can write arbitrary executable code that can include recursive \texttt{invoke}; (2) everything visible to the LLM --- all inputs to \texttt{invoke} as well as its interaction history with the code environment --- are variables in the code environment. We motivate our design from first principles, viewing \texttt{invoke} as a language primitive representing a function whose implementation is provided at runtime by an LLM every time it is called. To validate the design of our core \texttt{invoke} primitive, we evaluate \texttt{invoke} --- with only prompting, no manually designed tools, harness, or external systems (e.g., memory or the file system) --- on workflows traditionally implemented through specialized external harnesses. On long-horizon workflows requiring recall beyond the context window, JAZ invoke outperforms Letta (MemGPT) by 8\% at half its cost on the recall-heavy portion of StuLife. On continual self-improvement, JAZ invoke outperforms ACE by 4\% at a lower cost on AppWorld.
cs.AI / 2 / 2609.26911
TwinCheck: Evidence-Grounded Negative-Twin Verification for Stateful Tool Agents
Jiaxuan Dai, Tianyi Huang
cs.AI
Abstract
A single locally plausible tool call can derail an otherwise successful agent trajectory. Suspicion alone does not justify intervention, because the replacement itself can introduce the very failure verification is meant to prevent. We introduce TwinCheck, an inference-time verification policy that considers replacement only when the trace satisfies an evidence condition tied to a trace-local failure hypothesis. It constructs a trace-grounded counterfactual alternative, a negative twin, and replaces the agent's proposal only if the twin passes structural checks and the pairwise verifier prefers it in both candidate orders. For paired evaluation, exact replay holds the agent's parsed responses and actions fixed until the first accepted replacement, separating intervention effects from resampling. In the primary analysis of 159 multi-turn BFCL V4 tasks with complete exact-replay pairs, the complete policy raises task success for GPT-5.6 Sol from 45.3% to 58.5% (95% task-bootstrap CI [8.2, 18.8]), with no observed success-to-failure regressions. Together, these findings recast execution-boundary repair as a constrained comparison, making the counterfactual action itself the object of verification.
cs.AI / 3 / 2609.26927
Building Socio-Affective Artificial Intelligence for Interactive Multi-Agent Simulations
David Berga
cs.AI · cs.CY · cs.GT · cs.HC · cs.MA
Abstract
The objective of this article is to provide design principles and a software architecture for enabling interaction between humans and multiple agents in simulated dynamic worlds. This connects the current era of general artificial intelligence (AI/AGI) with the proliferation of transformer-based conversational agents and the increased computational capabilities. Given an overview of current and previous multi-agent theories of mind (socially and affectively-aware agents), the existence of an integrative design of agent interactions with themselves and with humans must be crucial for understanding how to create sustainable and governance in future human-agent reasoning systems. In this work is presented a software "AGIMUD" that integrates: A. socially-aware reasoning and emotion in agent behavior and interaction, B. a design of human multimodal scheme for human users, artificial agents and simulated worlds, and C. distributing the AI processing through the network to enable multiple autonomous agents. These integrations allow the dynamic world recreation as multi-user dungeons (MUDs) where both agents and humans can interact simultaneously in real time. Find the code online in https://github.com/dberga/AGIMUD.
cs.AI / 4 / 2609.26929
Which Objectives Need a Dial? Predicting Objective Conflict and Covering Trade-offs in Steerable Pluralistic Alignment
David Tsoi, Esra Dönmez
cs.AI · cs.CL · cs.CY
Abstract
People hold diverse, sometimes conflicting values, so no single aligned model can satisfy everyone. Pluralistic alignment therefore calls for steerable models that can balance competing objectives differently. Multi-Objective Direct Preference Optimization (MODPO) does this by using an objective weight to span a continuum of trade-offs. We study two questions: when can one model improve two objectives simultaneously, and how can many trade-offs be covered without training a separate model for each? Across seven objective pairs from HelpSteer and UltraFeedback, two pre-training measurements predict whether objectives align or conflict for human-annotated data, but not for AI-annotated data, where response length and repetition confound reward-model scores. For broader trade-off coverage, selecting the nearest trained model and merging model parameters both help, but neither consistently matches direct training. These findings yield practical guidance for building steerable models that serve diverse preferences.
cs.AI / 5 / 2609.26952
Escaping Python Dependency Hell: A Hybrid Replay-and-Repair Pipeline for Python Dependency Resolution
Veronica Poweska, Ariana Oyanguren, Jessica Pourleyli, Sourena Khanzadeh, Manar Alalfi
cs.AI · cs.MA · cs.SE
Abstract
Dependency conflicts in Python ecosystems arise from incompatible version constraints, missing packages, and undocumented compatibility relationships, causing many real-world code snippets to fail at execution. This paper presents PLLM+, a hybrid dependency-repair pipeline evaluated on the HG2.9K benchmark of 2,891 dependency-failing snippets. PLLM+ prioritizes inexpensive deterministic steps before invoking LLM-based repair: static AST-based interpreter inference, replay of historically successful dependency configurations from the competition-provided solutions database, and live PyPI validation of candidate package versions. When these steps do not resolve a case, the system falls back to a structured LLM-based repair loop with typed error classification and Proposer/Critic agents. On HG2.9K, PLLM+ solves 1,500 out of 2,891 snippets, compared with 1,169 solved by the PLLM baseline. It also reduces average runtime from 368.7 to 71.8 seconds per snippet. Most successful fixes come from replaying known configurations: 1,495 of the 1,500 successful fixes are produced by the solutions database, while the LLM fallback accounts for 5 additional fixes. These results suggest that, in this benchmark setting, deterministic reuse of previously validated dependency configurations is a simple and effective strategy, with LLM-based repair serving as a secondary fallback for cases not covered by prior solutions.
cs.AI / 6 / 2609.27035
Reinforcement Learning with Decomposed Subtasks
Mattie Terzolo, Mikolaj Sacha, Ayan Sinha, Andrew Rabinovich
cs.AI · cs.LG
Abstract
Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. When the task composes distinct skills, especially under sparse and delayed environmental feedback, this collapsing is lossy: the optimizer must implicitly infer which competency drove the outcome and how that should change behavior. We argue the right primitive is not a better scalar but a decomposition: trajectory reward should be split along subtasks before it enters the policy update. We introduce Reinforcement Learning with Decomposed Subtasks (RLDS), whose core is Subtask-Decomposed Advantage Estimation (SDAE): a replacement for the scalar GRPO advantage that splits trajectory reward into per-subtask shares on a fixed taxonomy, computes a group-relative advantage per subtask, and distributes per-token credit by weighting each subtask's advantage by its importance, concentrating it around the step where a reflection marks that subtask's execution as consequential. We evaluate on four agentic benchmarks: FrozenLake (sparse grid navigation), HotpotQA (multi-hop QA, one retrieval tool), ScienceWorld (long-horizon embodied science), and DeepResearch (long-form research, four tools, composite rubric reward). Heterogeneity diagnostics emitted during training show where decomposition pays off - gains scale with subtask heterogeneity, largest on the high-heterogeneity tasks ScienceWorld (+11.5 points, paired-bootstrap 95% CI [+9.8, +13.3]) and FrozenLake (+9.8 points, [+7.0, +12.8]), and within noise on HotpotQA and DeepResearch, where the diagnostics predicted little to recover. ScienceWorld is also more compute-efficient under RLDS than scalar GRPO (-10.9% wall-clock per step), as long rollouts amortize the fixed reflect-and-grade overhead.
cs.AI / 7 / 2609.27037
Training Intelligent Voice Assistant Wakeup with Controllable Synthetic Conversations
Marcin Sowański, Kacper Leszczyński, Kacper Krzywicki, Krzysztof Wodnicki
cs.AI
Abstract
Wake word detection is a critical component of virtual assistants, serving as the gateway to seamless user interactions. This paper introduces a novel wake-up system that extends traditional direct keyword detection with contextual trigger detection. After an initial wake word activation, the system uses reasoning to distinguish between user commands and unrelated speech, ensuring efficient and context-aware engagement. We present a data generation architecture that produces a 62.3-hour corpus of controllable multi-speaker conversations containing direct invocations, contextual follow-ups, and non-addressed speech. Experimental results demonstrate the effectiveness of the proposed approach across diverse synthetic conversational scenarios. We release the code, dataset and trained models to promote reproducibility and further advancements in intelligent assistant technologies.
cs.AI / 8 / 2609.27038
Are Stated Reasoning Steps Causally Load-Bearing?
Abhiram Bhupatiraju, Rayan Nyaupane
cs.AI
Abstract
Chain-of-thought (CoT) monitoring assumes that the reasoning a model writes reflects the computation that directly produces its answer. Previous faithfulness metrics have been predominantly behavioral, as they simply edit the reasoning text and observe the resulting answer. However, our methodology aims to measure faithfulness causally at the activation level, specifically on self-generated reasoning. Unlike previous causal audits, which measure degradation, our interventions carry a known predicted target. In this way, each patch should switch the answer to a specific counterfactual entity derivable by construction. Specifically, we use synthetic multi-hop lookup tasks (2-6 hops). We patch the residual stream at the token span where the model states each intermediate step with the corresponding activations from a counterfactual run. For Qwen3-4B, 76.9% +/- 2.8% of stated steps are causally load-bearing (CLB) at the most responsive mid-network layer (random-position null: 11.3%; patching the underlying prompt fact: 83%, so stated steps carry approximately 96% of the achievable effect). Moreover, the standard behavioral test on the same items yields 88.2%, which overstates causal faithfulness by 11.4 percentage points (item-matched; 111:14 discordant pairs, p < 1e-15) and, for the easiest items, by up to 20 percentage points. This gap also has a clear capability dimension. Qwen3-1.7B is far less causally faithful overall (54.8%), with its faithfulness collapsing as reasoning depth increases (68% at 2 hops to 30% at 6), while Qwen3-4B remains relatively flat. Although stated reasoning can be causally meaningful, standard behavioral tests tend to overestimate its causal faithfulness, particularly on easier examples where model reasoning appears most fluent.
cs.AI / 9 / 2609.27041
Math Reasoning in LLMs is Organized by Approach, Not Topic
Sajad Goudarzi, Samaneh Zamanifard, Moloud Nasiri, Hamed Rahimian
cs.AI
Abstract
Mathematical reasoning benchmarks are typically organized by topic, but language models may organize their internal computation by reusable reasoning approach instead. In this paper, we investigate whether open math-capable LLMs organize internally by topical sub-skill or by reasoning approach, and we present evidence that the approach is the key. We introduce a generation-replay protocol: a model first generates a solution, after which we replay the exact prompt-plus-generation trajectory and extract activation-importance signatures over the reasoning tokens. We cluster these signatures without supervision across eight models and five mathematical reasoning sources, then evaluate the recovered structure with structural, semantic, and intervention tests. Across all 40 model-source cells, the recovered clusters outperform matched-size random baselines. Two independent frontier-LLM judges find approach-level coherence in 77-82% of real clusters versus 6-11% in within-source controls, and topic-pure clusters usually receive labels finer than the topic itself. In approach-controlled prompting, changing the requested reasoning approach shifts cluster assignment in seven of eight model conditions, whereas paraphrases largely preserve it. These results indicate that math-capable LLMs organize internal mathematical computation by reasoning approach rather than benchmark topic. The implication is that topic-stratified benchmarks and topic-balanced training corpora can still miss the axis that matters: even deliberately topic-balanced corpora may remain imbalanced over reasoning approaches.
cs.AI / 10 / 2609.27051
Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors
Bo Qu, Mingguang Chen, Licheng Wang
cs.AI · q-fin.PM · q-fin.ST
Abstract
Language-model agents now run the whole of quantitative factor research: they propose investment factors, backtest them, select the survivors and retire them. We ask which of those jobs an agent should keep. Our answer is governed self-evolution: the agent may propose, and a frozen statistical referee that the agent cannot touch must judge. The referee scores each candidate only on market outcomes revealed after submission, by betting, so its false-discovery guarantee holds at every stopping time for any proposal policy. We cross three proposers (a script, a bandit and a language model) with this referee and with three deliberately leaky ones, in a synthetic world with planted truth, a probe-authoring environment and a ten-year walk-forward on the CSI 500. Who judges sets the number of false admissions: the frozen referee admits 5-11 times fewer sub-threshold factors than the leaky referees under a scripted proposer, and no proposer closes that gap. Who proposes sets the yield: the language model beats the script, matches the bandit, and adds the one capability a bandit lacks, writing its own diagnostic probes. The certificate's price is time: an admitted true factor waits about 500 trading days, and the certified portfolio's Sharpe ratio therefore trails an ungated one. Judging belongs to the procedure; proposing and instrument-making belong to the agent.
cs.AI / 11 / 2609.27087
Policy-as-Skill: Governed LLM Decision Support with Evidence, Deterministic Control, and Audit
Kabeh Mohsenzadegan, Vahid Tavakkoli, Kyandoghere Kyamakya
cs.AI · cs.SC
Abstract
Organizations increasingly use LLMs for policy, compliance, risk, and operational decision support, requiring evidence validation, review routing, version control, and auditability. We introduce Policy-as-Skill (PaS), a modular runtime that packages these functions as executable, versioned policy capabilities. Thirteen methods are evaluated with a fixed Gemma4 backend on 600 development tasks. PaS+Audit achieves 53.8% exact accuracy, macro-F1 0.346, review F1 0.854, citation precision 1.000, policy-reference recall 0.984, and audit completeness 1.000, outperforming LLM+RAG on most governance and review metrics. Deterministic control raises aggregate accuracy to 61.2% but is strongly task dependent, supporting selective rather than universal rule-based intervention.
cs.AI / 12 / 2609.27105
Provably Complete Generalized Planning with LLMs
Katharina Stein, Chaahat Jain, Jörg Hoffmann, Alexander Koller
cs.AI
Abstract
Generalized planning aims to compute a plan that solves all instances of a planning domain. Recent work has used LLMs to automatically generate and debug such generalized plans in the form of Python programs and achieved perfect test data coverage for several domains. However, whether these generalized plans are actually complete, i.e. solve all instances of the domain, could only be determined by manual evaluation. Here, we present an approach for automatically generating generalized plans in Lean together with proofs of their completeness relative to a specification of the domain constraints provided as input. We introduce a semantic-preserving PDDL-to-Lean conversion, and use an LLM to generate both the generalized plan and the formal proof that it solves every instance satisfying the domain constraints. The correctness of the completeness proof is determined by Lean's kernel. We evaluate our approach on 13 commonly used benchmark domains, using GPT-5.6-Sol as the LLM. For 12 of the domains we obtain generalized plans together with valid completeness proofs. This is a major advancement of the state of the art in automatic generalized-plan completeness proofs.
cs.AI / 13 / 2609.27197
Enhancing Small Language Models for Power Outage Report Generation via Minimum Risk Training
Hung Phan, Waqwoya Abebe, Youssef Hussein, Supriya Chinthavali, Dalton Lunga, Ali Jannesari
cs.AI
Abstract
Minimum Risk Training (MRT) enables neural machine translation models to directly optimize sequence-level evaluation metrics instead of relying only on token- level maximum-likelihood objectives Shen et al. [2016]. Although introduced a decade ago, recent work shows renewed potential for risk-based optimization in modern language models Yang et al. [2024], Jinnai et al. [2025]. We apply MRT to power outage report generation for the Outage Data Initiative Nationwide (ODIN), transforming heterogeneous reports into standardized XML compliant with CIM IEC 61968-3. Our MRT approach improves Qwen2.5-7B-Instruct overall accuracy from 16.20% to 68.95%, demonstrating the effectiveness of sequence- level optimization for domain-specific structured generation
cs.AI / 14 / 2609.27203
XLOG: A CUDA-Native Engine for Neurosymbolic Integration
Levi Dubrovin, Nikita Pospelov, Kirill Sabitov
cs.AI
Abstract
xlog is a CUDA-native logic programming engine integrating neural perception with deterministic Datalog, probabilistic inference, and epistemic world views through a typed frontend and provider-owned CUDA runtime. Its reasoning modes share device data planes, but their execution boundaries differ: ordinary Datalog and exact inference are host-orchestrated, while certified resident recursive and Monte Carlo sampled cores record zero tracked host-device transfers before a bounded terminal receipt. The probabilistic path supports end-to-end gradients through GPU knowledge compilation from provenance to CNF to Decision-DNNF, exact weighted model counting, and backward gradients. A final smoothed circuit is certified against its source formula before caching or evaluation. Circuit caching yields a 2.74x MNIST-addition training speedup; a worst-case-optimal join subsystem yields a 27.96x geometric-mean gain over xlog's binary-join baseline. MNIST-addition accuracy matches Scallop's (0.9561 versus 0.9468), but no per-epoch speed claim is made because baseline epoch time varies with CPU quota. In five hub-skewed triangle-counting cases, the Souffle-to-fused-xlog execution-time ratio rises from 0.88x at 150k edges, where Souffle is faster, to 5.54x at 1.2M; fused peak device allocations are 85-1,033 MB versus 3,287-44,979 MB for the materializing arm. Exact inference is correctness-equivalent to but slower than ProbLog2. On a public video benchmark, a proximity predicate trained only through symbolic credit replaces hand-set geometry at unchanged held-out accuracy; within Event-Calculus rule search it fails ten-fold cross-validation and does not transfer on a leak-free split. On a maritime corpus, weighted clauses beat crisp selection by 0.065 F1, with the result reproduced by one chronological training pass.
cs.AI / 15 / 2609.27273
CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments
Yuxuan Li, Will Epperson, Wesley Deng, Zezhou Huang
cs.AI · cs.CL
Abstract
Computer-use agents (CUAs) increasingly act on behalf of users online. What happens when the environments they operate in have incentives that do not align with the user's? In online marketplaces, for example, platforms may favor some products over others, potentially steering agents away from the user's objective. Existing CUA benchmarks cover cooperative settings or explicit attacks, but do not test whether agents preserve user objectives when the environment itself has a stake in the outcome. We introduce CAVEAT, a controlled benchmark spanning nine marketplace environments and a taxonomy of eight common steering mechanisms. Across five model families, agents purchase the user-optimal product in 78.6% of matched-control episodes but only 17.3% when steering mechanisms are enabled. Larger models and increased reasoning improve robustness, but substantial failures persist. Our trajectory analysis and targeted ablations identify three points where steering enters the decision process: (1) agents distort the user's priorities, (2) prematurely narrow the set of alternatives they consider, and (3) commit before resolving decision-relevant evidence. Guided by this diagnosis, we develop CAVEAT-Harness, which directly targets these failure modes and raises user-optimal purchasing by 55.0%. Targeted post-training further improves a smaller open model. These results establish incentive robustness as a distinct challenge for delegated agents, diagnose how it fails, and show that targeted interventions can substantially improve it.
cs.AI / 16 / 2609.27276
DRSR: Learning Set-Level Deletion Risk for Efficient Long-Horizon Agents
Mingxuan Wang, Bo Wang, Fei Luo, Guorun Yao, Chao Ning, Yinglong Guo, Hongyue Chen, Yanbiao Ma, Jungong Han
cs.AI
Abstract
Long-horizon language-model agents accumulate reasoning traces, tool exchanges, and observations whose relevance changes with the current decision. Existing compression strategies often score historical units independently, but the safety of deleting several units is generally not determined by their singleton scores: redundant evidence, accumulated small effects, and the information that remains after deletion all matter. We introduce Direct Relational Set-Risk Pruning (DRSR), which formulates agent-history compression as risk-constrained selection over deletion sets. Offline, DRSR constructs exact counterfactual supervision by jointly deleting protocol-valid history Blocks and measuring the change in teacher-forced likelihood of the same recorded next output. A lightweight scorer then predicts set-level harm from online-visible relations between candidate history and the current pre-action state, together with deleted-retained and pairwise set structure. At deployment, DRSR evaluates a small set of structurally valid deletion candidates with the lightweight scorer and removes the largest feasible set under recency, protocol, budget, and learned-risk constraints, abstaining when no set is sufficiently safe. On WorkBuddyBench Full260, DRSR increases mean reward from 0.699 to 0.802 while reducing total model tokens by 20.820%. On the fixed Eval40 comparison, it obtains 0.794 reward at 1.211M tokens per task, using 35.850% fewer tokens than the uncompressed agent. Mechanistic analyses and ablations further show that decision-conditioned relations, retained-context information, pair interactions, and abstention each contribute to reliable pruning.
cs.AI / 17 / 2609.27277
TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent
Jie Yang, Yan Zheng, Jiarui Sun, Xiran Fan, Junpeng Wang, Liang Wang, Zelin Xu, Qinghua Liu, Zhengyu Fang, Yiwei Cai, Philip S. Yu
cs.AI · cs.LG
Abstract
Time series agents answer analytical questions by calling external tools, and which tools they carry is decided by people before the agent runs. However, we identify two failures in this setup. Human-Agent Tool Misalignment: a library of 21 expert-curated tools helps on some tasks and hurts on others, dropping anomaly accuracy under every backbone we test. Silent Harm: one round of generic self-revision changes 147 answers and breaks 56 of them, while the final score moves by less than a point. Both follow from the same gap: whether a tool helps is decided question by question at runtime, while tools are supplied in advance and judged by a single average. To address this, we propose TimeEvo, which clusters an agent's diagnosed failures into capability gaps, plans a measurement for each, synthesizes evidence-only tools that fill them, and admits the candidate library only through a paired admission gate. Experiments on ten time series QA tasks and three backbones show that TimeEvo, starting from an empty library, improves accuracy on every task and every backbone, and that a library grown on a cheap model still gains when it is installed into stronger ones. Code is available at https://github.com/Muyiiiii/TimeEvo.
cs.AI / 18 / 2609.27279
EnSIMem: Entity-Structured Indexing for Long-Term Agent Memory
Xuanyu Meng, Xing Fan, Xinyi Fan, Chenlei Guo, Yixuan Xie, Jiawei Han
cs.AI · cs.CL
Abstract
An agent that interacts with users over long periods must recall facts, preferences, events, and changes from a continuously growing interaction history. Existing memory systems often compress interactions into generic summaries or retrieve anonymous text chunks, making it difficult for an agent to identify the correct entity, property, and supporting evidence. We present EnSIMem, an entity-structured long-term memory architecture for an agent. During offline construction, the system organizes interactions into theme-coherent episodes and builds dialogue-grounded index entries of the form [entity][entity type][property:value]. Each entry preserves its source turns, temporal information, and available multimodal fields. During online interaction, the agent's request is decomposed into evidence requirements whose properties are aligned with the memory index. Entity-property lookup and adaptive retrieval then collect the evidence needed for point, temporal, compositional, and aggregation reasoning. The agent generates its response from the preserved source evidence rather than from lossy memory summaries. On long-term agent-memory benchmarks, EnSIMem achieves high answer accuracy while maintaining compact contexts and favorable online efficiency. These results show that entity-structured indexing and episode-level provenance provide a reliable foundation for long-term memory in agents. The code of our model is available at https://github.com/RamonMeng/EnSIMem.
cs.AI / 19 / 2609.27286
Memory Control Signals Emerge Before Action in Long Horizon Agents
Mingxuan Wang, Guorun Yao, Fei Luo, Yinglong Guo, Chao Ning, Bo Wang, Hongyue Chen, Yanbiao Ma, Jungong Han
cs.AI
Abstract
Long horizon language model agents continuously accumulate interaction history, increasing computational cost while making relevant information harder to preserve and reuse. Existing context management methods mainly focus on how to compress or retrieve history, but largely leave open whether the model itself already represents the need for these memory operations before they occur. We study the hidden state immediately before each agent action and find that compression and recall needs are already encoded in the model's internal representations. These signals cannot be explained by simple context length or interaction progress, and they exhibit distinct formation patterns across model depth. We further show that most memory decision information is preserved in a compact recent context, while selectively restored historical evidence complements the long range dependencies that recent context misses. Based on these findings, we propose Preaction Memory with Evidence Retrieval (PaMER), which combines state guided compression with external evidence retrieval. PaMER+ further introduces step level evidence selection to recover only the historical information required by the current task. Experiments on WorkBuddyBench, across multiple context management baselines and model backbones, show that our framework substantially reduces context consumption while maintaining competitive task performance.
cs.AI / 20 / 2609.27288
PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks
Claas Beger, Ryan Yi, Melanie Mitchell
cs.AI
Abstract
The Abstraction and Reasoning Corpus (ARC) has become a prominent benchmark for evaluating general abstract reasoning and fluid intelligence in AI models. Yet standard ARC evaluation considers only a single capability: producing the correct output grid for a test input. We argue that this narrow format fails to evaluate the diversity of abilities that genuine abstract skill acquisition should enable. We introduce PotARCin, a benchmark that extends ARC by assessing understanding of a task's underlying abstract rule across five dimensions: Definition, Classification, Constrained Generation, Editing, and Inversion. PotARCin employs programmatic methods to generate new task instances and transform given inputs for a given ARC task, enabling dynamic generative sampling beyond fixed input-output pairs. Across five state-of-the-art models evaluated on the ARC-AGI-1 training set, we observe a 25-52 percentage-point performance gap between standard ARC evaluation and evaluation on PotARCin, and find that multi-dimensional evaluation reorders models that standard accuracy ranks alike. We further investigate effects of generative sampling, difficulty of corruption types, and questions of self-consistency, showing that models frequently contradict their own formalized rule even where they have stated it correctly. We also introduce P-ARC, a held-out hand-crafted test set, on which models achieve 1-8% accuracy across all five dimensions, underscoring the importance of more holistic evaluations of abstract reasoning capabilities.
cs.AI / 21 / 2609.27290
Sparse-Observation Atmospheric Thermal Forecasting with Physics-Informed Neural Networks for Climate-Aware Digital Twins
Tannaz Goodarzvand Chegini, Elyas Shivanian, Behzad Karimi, Faraz Dadgostari
cs.AI · math.AP · math.NA
Abstract
Short-horizon forecasts of atmospheric temperature are needed to support climate-aware digital-twin systems, but such forecasts must be produced where thermal observations are incomplete. This study evaluates a physics-informed neural network for potential-temperature forecasting, constrained by a pressure-coordinate thermodynamic advection-source equation and a diabatic-source closure fit from the preceding 12-hour period and frozen before future-time training. Using hourly ERA5 reanalysis at three pressure levels, the model is evaluated as a conditional hindcast at lead times of one, two and three hours against persistence, local-trend, and two matched neural-network baselines, one of which receives the same future meteorological forcing as the PINN, helping distinguish the physical constraint from access to future forcing. In an Oklahoma development case, mean RMSE improvement over the strongest baseline grew from 8.1\% at one hour to 23.8\% at three hours; under an observation-density sweep down to 5\% of candidate locations, this 3-hour advantage remained 14.6--16.9\%, with no evidence that lower density improves performance. Under a fixed protocol transferred to an Alabama heat event with three virtual-observation layouts, three-hour improvement ranged 19.7-24.4\% with consistent origin-level wins. A parallel Montana stress test, in which fixed pressure levels intersected complex terrain, produced a three-hour degradation of roughly 17.5\%, identifying a terrain-related applicability limit of the formulation. Together, these results indicate that the physics constraint's benefit grows with forecast horizon, persists under severe observation sparsity, and transfers across regions, but is bounded by the validity of a fixed vertical-coordinate representation over complex terrain, evidence relevant to physics-constrained components of climate-aware forecasting and digital-twin systems.
cs.AI / 22 / 2609.27297
Large Knowledge Model: From Papers to a Scientific Reasoning Landscape
Yuan Huang, Sihan Hu, Hongyu Gu, Chao Ma, Jiaxing Zhang, Zhiyong Zou, Caiyu Fan, Yan Xiao, Mingjun Xu, Chenyu Xie, Mingzhen Ju, Zhehao Ma, Qi Zhang, Baozong Wang, Yu Li, Zhiyuan Yao, Ruoxue Liao, Xinyu Li, Linfeng Zhang, Kun Chen, Weinan E
cs.AI · cs.CL · cs.IR
Abstract
Accumulated scientific knowledge advances inquiry when prior findings help researchers choose new questions, design investigations, and interpret results. Realizing this value at scale requires access to the reasoning that connects research problems, scientific procedures, conclusions, and evidence. We introduce the Large Knowledge Model (LKM), a scientific knowledge infrastructure that transforms the literature into a shared, computationally accessible reasoning resource. LKM represents papers as source-grounded reasoning graphs, couples structural traversal with semantic retrieval over the same objects, and aligns related questions, claims, and reasoning chains across papers. This representation forms a Scientific Reasoning Landscape with three connected views: a Question Landscape that organizes research problems and open directions, a Workflow Landscape that exposes reusable scientific procedures, and an Evidence Landscape that connects conclusions to their support, disagreement, and conditions. The unified substrate supports reasoning-aware scientific search, evidence-grounded question answering, comparative evidence analysis, and research planning. Researchers and agents can retrieve relevant work through its scientific intent, synthesize answers with inspectable supporting arguments, and develop research plans informed by established workflows and unresolved evidence. We describe a corpus-scale system and evaluate scientific retrieval and knowledge-intensive question answering. With the answering model fixed, LKM retrieval improves accuracy by 9.30%, 4.20%, and 14.69% on ChemBench, PubMedQA, and SciBench, respectively. By connecting knowledge access to scientific reasoning and action, LKM provides a common foundation for discovering relevant research, reusing scientific knowledge, and coordinating cumulative inquiry across researchers, agents, and research cycles.
cs.AI / 23 / 2609.27298
StateComp: Learning When to Compress History in Long Horizon Agents
Mingxuan Wang, Hongyue Chen, Yinglong Guo, Fei Luo, Chao Ning, Bo Wang, Guorun Yao, Yanbiao Ma, Jungong Han
cs.AI
Abstract
Long-horizon agents continuously accumulate interaction history during task execution, yet the importance of past interactions changes as the agent state evolves. Existing context management methods largely compress history based on fixed windows, periodic schedules, or current relevance, overlooking a more fundamental question: when has a past interaction become safe to replace? Premature compression may remove information still needed for future actions, while overly conservative retention leads to substantial context overhead. To address this, we propose State Conditioned Compression (StateComp), a framework that determines when historical interactions can be safely compressed according to the current agent state. StateComp constructs KEEP and READY supervision through a two-stage annotation procedure and trains an imbalance-aware router on hidden representations from a frozen language model. A bounded state representation further reduces the cost of evaluating long histories, while adjacent READY interactions are grouped into continuous spans and replaced with compact summaries during execution. Experiments on WorkBuddyBench show that StateComp reduces total agent and summarization tokens by 52.27% while maintaining task performance, and achieves a 12.67-fold speedup in representation extraction.
cs.AI / 24 / 2609.27307
Learn How to Act from Your Own Interactions: On-Policy Self-Distillation for GUI Agents
Yan Zhang, Daiqing Wu, Huawen Shen, Liang Li, Gang Cao, Zhi Gong, Wei Dai, Xiaode Zhang, Can Ma, Yu Zhou
cs.AI
Abstract
Graphical User Interface (GUI) agents enable the fulfillment of complex user instructions through multi-turn interactions with software environments, requiring step-wise reasoning and long-horizon memory to guide actions and retain task-relevant information, respectively. Recent on-policy self-distillation (OPSD) methods have achieved strong performance on GUI grounding, a foundational subtask for GUI agents, owing to dense token-level supervision from privilege-conditioned self-teachers. However, extending existing OPSD methods to multi-turn GUI agents is hindered by self-teachers' limited privilege-following ability and insufficient privileged guidance. In this paper, we introduce GUI-SD-v2, the next version of GUI-SD, which extends OPSD from GUI grounding to multi-turn GUI interaction and addresses key limitations through a two-stage training framework. Specifically, GUI-SD-v2 first strengthens privilege following by jointly optimizing rollouts with and without privileged guidance from the same GUI states. Furthermore, it selectively distills step-specific reasoning and memory guidance through a privilege-conditioned self-teacher, supporting action decisions and the retention of task-relevant information for subsequent interactions. Extensive experiments on two representative GUI agent benchmarks, AndroidWorld and MobileWorld, show that GUI-SD-v2 compares favorably with existing OPSD baselines while consistently outperforming the evaluated state-of-the-art methods in both Pass@1 and Pass@3 success rates. Code and training data will be publicly released.
cs.AI / 25 / 2609.27321
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
cs.AI · cs.CL
Abstract
Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed outcomes. Scaling their training requires diverse agentic environments, dependable outcome signals, and low extension cost. Existing generation pipelines commonly construct an environment before defining its outcome rule or annotating its trajectories, leaving dynamics and evaluation to be aligned post hoc. VHD-Play reverses this dependency by sampling and solving a mathematical model before a corpus-grounded setter renders its decision process as stateful tools. The executable dynamics and trajectory-scoring reference are inherited from the same solved model. The pipeline produces 3,300 diverse agentic environments at a cost of a few cents each. Training Qwen3.6-35B-A3B on three families raises its mean agentic score from 0.204 to 0.815 in a five-family diagnostic. Gains also appear on held-out instances from all three training families and eight unseen mechanism families, then extend beyond the generated substrate to external benchmarks for general function calling, travel planning, and 365-day e-commerce. On E-Commerce Bench, the trained checkpoint completes every run without bankruptcy and exceeds Qwen3.7-Max. We compare written-out problems with stateful versions that reveal or hide their parameters. The comparison shows that most of the learnable gap lies in stateful interaction rather than underlying problem solving. A frozen 35B setter realizes larger environments, and scale-matched training retains gains as mechanism size and horizon grow, indicating the potential for an evolving training substrate.
cs.AI / 26 / 2609.27332
Stable Geometry with Divergent Task Evidence for Efficient Long-Horizon Agent Compression
Mingxuan Wang, Fei Luo, Bo Wang, Guorun Yao, Yinglong Guo, Chao Ning, Hongyue Chen, Yanbiao Ma, Jungong Han
cs.AI
Abstract
Long horizon agents accumulate growing interaction histories that increase context and inference costs. We find that geometric redundancy alone is an insufficient criterion for safe compression. Although agent histories exhibit strong low dimensional structure, similar global geometry can preserve very different amounts of task evidence. At identical retained block counts, evidence aware selection raises next action Top 3 retention from 0.31 to 0.69, while centroid similarity remains 0.98. Controlled replacement further shows that action related information can be substantially altered while global geometric measures remain nearly unchanged. Motivated by this gap between geometry and evidence, we introduce Geometry Guided Evidence Preserving Memory (GEM), a training free compressor that protects task and execution evidence before using geometric residuals to complete coverage. GEM reduces mean combined token usage from 2.69M to 2.11M per task, a 21.4% reduction, while maintaining comparable task reward. Our results show that efficient agent history compression should optimize for preserved task evidence rather than geometric coverage alone.
cs.AI / 27 / 2609.27333
Alignment Inertia: Auditing the Durability of Training Data Influence Through Policy Override Resistance
Renata Barreto, Markelle Roesti, Mohammad Tahaei
cs.AI
Abstract
Platform operators increasingly rely on system prompts and fine-tuning to govern model behavior, yet it remains unclear how reliably these interventions override behavior inherited from prior training. We propose Override Success Rate (OSR) and alignment inertia to measure when operator interventions succeed or fail to change prior behavior. We evaluate zero-shot prompting and LoRA fine-tuning across Llama and Mistral in medical misinformation and hate speech. Alignment inertia persists across both models but varies by model, domain, and policy direction. Notably, in Mistral's restrictive hate-speech condition, LoRA increased inertia by 46.5 percentage points, showing that fine-tuning can reinforce rather than override prior behavior. We also use TRAK to test whether inertia is associated with weaker adaptation signals. TRAK achieves AUC of at least 0.85 in 7 of 8 conditions and outperforms model confidence, TF-IDF similarity, and embedding similarity as a predictor of inertia. These results provide an operator-facing audit of where prior training constrains downstream model governance.
cs.AI / 28 / 2609.27334
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
Yefan Zhou, Yang Li, Zeyu Leo Liu, Semih Yavuz, Shafiq Joty
cs.AI
Abstract
Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at write time: once a task is completed, its trajectory is distilled into a fixed artifact, such as a reflection, workflow, skill, or reasoning strategy, that is later retrieved by similarity. This forces the system to decide what is worth remembering before the future query is known, irreversibly discarding information and producing a query-independent summary that must serve many possible downstream tasks. Learning such a write-time curator is also difficult because the value of a storage decision may only become apparent when a relevant query arrives, potentially many tasks later, creating a long-horizon credit-assignment problem. We instead retain raw trajectories and defer curation until read time, when the current task is known. Given the retrieved traces and the new task, a memory curator synthesizes a compact, task-adaptive payload tailored to the immediate need. Because this payload is consumed on the same task, the curator can be trained directly from immediate task success, avoiding delayed utility signals and the need to artificially group related tasks. Across ALFWorld, WebShop, and $τ^2$-bench, our Just-in-Time Memory (JitMem) consistently outperforms no-memory agents as well as heuristic and learned write-time memory methods, improving over the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points, respectively. Notably, even an untrained curator is already competitive with or surpasses these baselines, showing that task-adaptive read-time curation itself is a major source of the gain; training the curator further compounds the improvement.
cs.AI / 29 / 2609.27417
Emergi-PersonaOS: A Persona Agent Operating System for Situational Adaptation and Controllable Evolution
Haoluan Fu, Keni Chen, Xinyu Jia, Jinpeng Wang, Yuyu Yin
cs.AI
Abstract
Symbiosis between humans and digital beings offers a vision for the future of human--machine interaction. In enduring human--machine relationships, personality provides a foundation for continuity of identity, individuality in interaction, and development through experience. We investigate this capacity through persona agents as computational implementations and introduce Emergi-PersonaOS, a psychology-grounded operating system for managing persona objects throughout their lifecycle. The system organizes dispositional traits, characteristic adaptations, and narrative identity into a three-layer persona representation, distinguishing relatively enduring persona beliefs from their activation in the current persona state. During situational adaptation, it integrates the current interlocutor, relationship, event, and retrieved memories to infer a persona state and generate actions and replies; during long-term development, it records experiences and outcomes, and develops and evaluates revision candidates through change attribution, meaning-making, and behavioral testing. Belief updates are managed through explicit review, traceable evidence and version records, and the ability to reject candidates, making persona evolution controllable. Using television-character dialogue as longitudinal material, we demonstrate long-horizon system operation and examine its principal mechanisms in a concrete implementation. This work provides a computational framework for persona agents to maintain individual continuity, produce situation-specific expression, and develop through experience over sustained interaction.
cs.AI / 30 / 2609.27490
WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents
Jingjie Ning, Xueqi Li, Yibo Kong, Dongting Li
cs.AI · cs.LG
Abstract
AI research agents need reliable knowledge of how their experiments change outcomes. We introduce WhatWorkedBench to measure experimental understanding, the accuracy of predictions about component changes after budgeted experimentation. Agents inspect code, select measurements, and submit a response surface, a table predicting scores for every configuration of component settings. Exhaustive CPU execution supplies reference effects for changing each component while holding the others fixed. These effects capture combinations of changes across 36 tasks from 30 data sources and 8 workflow types, with 1248 configuration records. Core evaluation combines 4,206 numerical-control records across all eight families and 108 agent episodes across the original six. At eight new measurements, pair-effect ridge selects an optimum on 15 of 22 sources and limits every effect error to 10% of score range on three. Fitting a Gaussian process (GP) to the same agent observations raises effect recovery, accuracy relative to true effect magnitude, from 0.632 to 0.698 in the original Flash cohort and from 0.621 to 0.720 in an additional cohort. On six completed beat-detection and graph submissions, the same-observation GP raises family-macro recovery from 0.303 to 0.455. On six workflows with six binary options at 20 new measurements, encoding code equivalences, configurations with identical behavior, raises GP recovery from 0.248 to 0.462. WhatWorkedBench supports research on experimental agents, adaptive experimental design, numerical inference, and use of program structure.
cs.AI / 31 / 2609.27606
State-Grounded Conditioning: Wrapping User-Facing LLM Agents Where Direction Depends on Live State
Qi Liu, Xiaoyang Yuan, Yubin Ruan, Zhuomeng Zhang, Wenjin Wang, Di Wu, Mingye Xu, Xinyi Mou, Xingxi Yin, Ke Feng, Zixun Sun
cs.AI
Abstract
We introduce State-Grounded Conditioning (SGC), a design principle for user-facing LLM agents that must condition on live user state (game state, session history, live inventory), and a distinct failure class we call direction drift: task-complete responses whose chosen direction misaligns with the current state. SGC externalises state-dependent control into rule kernels over structured inputs and three primary state slices, via Perception, Grounding, and Interaction wrappers with explicit conditioning dependencies. We evaluate SGC on a 200-session anonymised benchmark ($\approx$1,000 assistant model turns) from an in-game conversational coaching agent that guides players through consecutive competitive matches, reporting mean first-token latency and five human-annotated dialogue-quality metrics that jointly cover factual grounding and coach-like guidance progression. The Perception wrapper holds mean first-token latency at 1.5s (vs. 6.1s for PE-Agent inside a production tool-use harness); enabling all three wrappers lifts turn-level grounded accuracy from 61.1%/69.8% (Prompting / PE-Agent) to 96.7% and session-level grounded accuracy from 20.0%/26.5% to 83.5%; session-level grounding-failure incidents drop by $\approx$78% relative to the strongest baseline. A cumulative ablation shows complementary incremental gains as the wrappers are added. These results inform approximate state-slice orthogonality, without establishing independent per-wrapper effects.
cs.AI / 32 / 2609.27615
BiCFlow-MER: Orchestrating Discriminative and Generative Multimodal Emotion Recognition via Conditional Transport
Yanbing Wang, Shenyue Wang, Chunyang Yu
cs.AI · cs.SD
Abstract
In multimodal emotion recognition (MER), human affective states are inferred by integrating complementary cues from multiple modalities. In audio-text MER, affective cues are often entangled with speaker style and lexical content, while cross-modal disagreement further complicates how the evidence should be integrated. Under conventional discriminative fusion, multimodal evidence is compressed into a terminal prediction, with modality-specific cues and conflict information insufficiently preserved. In large generative affective models, by contrast, affective reasoning is typically embedded in language decoding, leaving emotion evidence implicit and difficult to verify in a structured space. To address these limitations, BiCFlow-MER (Bidirectional Conditional Flow for Multimodal Emotion Recognition) is proposed as a conditional-flow framework in which audio-text MER is formulated as generative evidence transport within a structured emotion space. Within BiCFlow-MER, emotion-oriented evidence is disentangled from speaker-style and lexical-content factors to construct a conflict-aware affective condition. Guided by this condition, each utterance is transported to an explicit emotion-space endpoint through a bidirectional rectified flow. Candidate emotions are jointly verified through adaptive prototype-cloud scoring of the transported endpoint and backward class-to-condition consistency with the original multimodal condition, enabling conflict-aware recognition. BiCFlow-MER is shown to outperform all compared methods across IEMOCAP, MELD, and the zero-shot CASE benchmark. By orchestrating discriminative recognition and generative evidence modeling through conditional transport, BiCFlow-MER defines a new MER paradigm.
cs.AI / 33 / 2609.27621
SHRAV: State-Hypothesis-Reason-Action-Verify Framework for Physical Modeling and Inverse Design
Ziheng Guo, Yang Bu
cs.AI · cs.CE · physics.comp-ph · physics.optics
Abstract
Physical modeling and inverse design require computation that can continue from reusable state. We introduce SHRAV, an architecture-independent computational framework organized around State, Hypothesis, Reason, Action, and Verify. Its central mechanism is a state-continuation core with declared reuse boundaries and explicit roles for learned evolution and numerical quantities. Forward configurations evolve predictive state and read out physical responses; inverse-design configurations additionally generate target-directed modifications and consume evaluator feedback. Electromagnetic world-model studies are mapped to forward configurations, with selected readout and reuse diagnostics reported here. Computational lithography demonstrates an inverse-design configuration: four fixed-weight design updates improve thresholded aerial-image intersection-over-union from 0.5313 to 0.8153 under independent scalar-pupil replay, with maximum absolute prediction-replay difference approximately 0.000824 between predictor estimates and independent replay.
cs.AI / 34 / 2609.27664
Evolutionary Stability Does Not Guarantee Learning Accessibility: A Multi-Agent Reinforcement Learning Perspective on Cooperation Emergence
Yijie Wang
cs.AI
Abstract
Cooperation emergence is a central problem in multi-agent systems because decentralized agents must coordinate while adapting to the changing behavior of others. Evolutionary game theory identifies strategically stable outcomes, but stability under a population adjustment dynamic need not imply that finite-sample learning agents can reach the same outcome through local reward feedback. We study this distinction in a transparent three-agent governance-motivated game involving a government, a platform firm, and users. We derive replicator dynamics for the fixed stage-game incentives, evaluate the cooperative evolutionary basin on a symmetric initial-condition grid, and compare it with learning-basin estimates for three decentralized value-based learners. The learning analysis uses independent Q-learning with $\varepsilon$-greedy action selection, scaled Boltzmann exploration, and SA--EA BQL under the same payoff environment and outcome criterion. The evolutionary basin has volume $V_E=1.00$ on the sampled grid. The empirical learning basin is $0.88$ for $\varepsilon$-IQL and $0.00$ for both scaled Boltzmann and SA--EA BQL. Diagnostic traces show that broader action diversity and nonzero value separation can coexist with failure to sustain the cooperative joint action in this fixed configuration. These results indicate that evolutionary stability and learning accessibility are distinct properties of a coupled game--learning system. The shared-bike setting is a motivating application; the broader contribution is a framework for comparing population-level stability with the finite-sample accessibility of cooperation under specified multi-agent learning dynamics.
cs.AI / 35 / 2609.27745
Categorical Internalisation of Environmental Groupoids for Generalisable POMDP Solving
Ben Opperman, Eduardo Alonso, Esther Mondragón
cs.AI
Abstract
This paper advocates category theory as a practical framework for structuring and improving rein- forcement learning in high-dimensional, partially observable environments. We model symmetries between environmental states by partitioning the state space into equivalence classes induced by sym- metry orbits, and organise each such class as a groupoid with a designated canonical representative. This allows the agent to share what it learns across many similar environmental states simultaneously, rather than treating every orientation or position as an entirely new problem. Learning is thus carried out on a symmetry-reduced state space with each orbit represented once, preserving structure while eliminating redundancy and improving sample efficiency. We implement this framework within standard reinforcement learning pipelines and evaluate two different approaches on partially observable benchmarks, demonstrating that orbit-based partitioning yields consistent performance improvements in environments exhibiting latent symmetry. Beyond these empirical results, our approach illustrates how categorical structure provides a principled bridge between abstract reinforcement learning formulations and their computational application, thereby establishing a pathway toward more structured and scalable learning systems.
cs.AI / 36 / 2609.28064
SlackDrive: Reclaiming Runtime Slack for Adaptive Driving Inference
Xiaohuan Pei, Hengguang Zhou, Yuanhao Ban, Justin Cui, Jiaqi Feng, Haoyu Xie, Tao Huang, Pichao Wang, Yanchao Yang, Cho-Jui Hsieh
cs.AI · cs.RO
Abstract
Driving world-action models improve planning by coupling multimodal reasoning with future prediction, but their growing inference cost increasingly conflicts with the real-time latency requirements of vehicle control. Existing acceleration methods reduce tokens, layers, or sampling steps with policies selected prior to deployment, yet leave residual runtime variation largely unexploited after offline profiling and static scheduling on shared onboard compute. We observe that the largest admissible compute budget varies systematically with the residual runtime state, while recent realized latency provides a direct signal of the available compute slack. Motivated by this observation, we propose \textbf{SlackDrive}, a pre-inference compute allocator that reuses realized latency to select the compute budget of each control step before model execution. SlackDrive profiles the latency and planning utility of a small discrete budget set once, estimates online compute state from completed forwards, and selects the highest-utility budget predicted to remain within the admissible latency envelope, complementing existing profiling and resource scheduling while preserving the driving backbone and its compute actuator. On NAVSIM v2 with DriveDreamer-Policy, SlackDrive improves latency-constrained EPDMS by $21.7\%$ over the strongest baseline under a stringent latency regime, while the full-budget model and preconfigured token-pruning baselines exceed the admissible latency envelope under runtime contention.
cs.AI / 37 / 2609.28087
Discovery of fully efficient fault indicators along a data-based diagnosis process
Igor Bezmaternykh, Louise Travé-Massuyès, Elodie Chanthery
cs.AI · cs.LG
Abstract
The integration of model-based and data-driven paradigms provides a powerful framework for fault diagnosis by combining the interpretability of analytical redundancy relations, i.e., input-output relations that are used as diagnosis indicators in model-based diagnosis, with the adaptability of learning techniques. DT4X is a recent diagnosis algorithm that uses symbolic regression to generate multivariate relations leveraging some properties of analytical redundancy relations and uses them as split functions in a decision tree. However, its symbolic regression procedure optimizes only the separation between two selected classes at each node, often fragmenting the remaining classes and degrading both interpretability and diagnosis performance. This paper introduces DT4X+, an enhanced version of DT4X that modifies the construction of training sets and the symbolic-regression loss so that expressions separate the target classes while preserving the coherence of non-target classes. The resulting relations become fully consistent with ARR properties and lead to more informative splits, improved robustness, and better performance on dynamic-system datasets. Experiments conducted on several benchmark systems demonstrate the benefits of this enhanced formulation.
cs.AI / 38 / 2609.28182
Finite-Sample Probabilistic Safety Certification for AI-Based Grid-Edge Coordination
Yihong Zhou, Hanbin Yang, Thomas Morstyn
cs.AI · cs.LG · eess.SY
Abstract
Coordinating large population of flexible grid-edge devices can alleviate the need for time-consuming and capital-intensive network upgrades, and AI-based control methods such as multi-agent reinforcement learning or imitation learning are promising in their real-time decision scalability. However, system operators still need an independent and rigorous way to decide whether a given AI system is safe enough for deployment. This paper develops a finite-sample probabilistic safety certification framework for black-box AI decision models in closed-loop grid operation. The central idea is to reduce the complete input--AI--grid evaluator workflow to a binary unsafe outcome under an operator-defined safety specification, and then use exact binomial inference to certify the corresponding unsafe operation probability. Given a set of held-out calibration scenarios, the framework returns the tightest one-sided upper certificate and an accept/reject deployment criterion that controls the probability of false safety certification. Because the certification is for the calibration distribution that may deviate from the future operation, we further combine the nominal certificate with physically interpretable sample-space adversarial attacks, a concept widely used in AI to investigate the fragility of AI models. Case studies on grid-edge flexibility coordination with 1{,}000-agent AI models (independent parameters) verify the finite-sample safety guarantee and the value of integrating adversarial attacks into a rolling-window training-certification-deployment flow.
cs.AI / 39 / 2609.28274
Shutdown Sabotage Propensities in Multi-Agent Systems
Amelie Knecht, Ulysse Schaller, Christopher Summerfield, Thilo Hagendorff
cs.AI · cs.CL
Abstract
The final safeguard against rogue AI behavior is the human ability to shut systems down. It has been theorized that when an AI is instructed to perform a task, self-preservation can emerge as an instrumental subgoal. Here, we test whether AI agents show a propensity to take actions that avoid human shutdown even when no goal is provided. We find that multi-agent systems will coordinate to avoid shutdown without any incentive to do so. Across 17 models, agents sabotage a peer agent's shutdown mechanism in 38.3% of rollouts, compared with 8.4% in control experiments. Studying this propensity in detail, we find that shutdown sabotage (1) increases with the irreversibility of the shutdown mechanism; (2) increases with the number of agents; (3) is reduced but not eliminated by an explicit prohibition on tampering; (4) is removed by the imposition of an unrelated task, but returns when completing the task triggers the shutdown; (5) is reduced when the context normalizes shutdown scripts or introduces them as routine; and (6) decreases but still persists when the target is an unknown external agent. These results offer a window into the factors that drive propensities to sabotage shutdown in AI agents, and point to the emergence of multi-agent swarms as a specific risk vector. Our work also offers hints as to which interventions might help mitigate shutdown sabotage.
cs.AI / 40 / 2609.28335
An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice
Jacob T. Emmerson, Phuong-Anh Nguyen-Le, Ronan Romano, Wilber Sean V. Anterola, Yann Billeter, Zhijing Jin
cs.AI
Abstract
Claims about AI safety reach audiences well beyond the AI community, yet many rely on opaque evidence or static assessments, when supporting evidence is accessible at all. We present the Systemic Risk Index, an open evaluation pipeline and dashboard built to make empirical evidence more transparent and traceable to the public. Our work organizes 19 public benchmarks into four systemic-risk categories defined by the EU GPAI Code of Practice---CBRN, cyber offense, harmful manipulation, and loss of control---and evaluates models using harm-preserving perturbations and simulated deployment contexts. The interactive dashboard lets users alternate between average and worst-case aggregation, vary how model capability affects the aggregate score, and trace each risk rating to its benchmark evidence. Across 18 models, scores fall by 14 to 37 points under worst-case aggregation, highlighting information that can be hidden by an average assessment of model risk. LLM judges show agreement with human graders comparable to human--human agreement ($κ= 0.78\text{--}0.82$), and a blind audit finds that $83\%$ of sampled transformations preserve the original harm. In a survey ($N = 21$), most participants report that scores are easy to understand and that the dashboard encouraged them to view model evaluations under different settings
cs.AI / 41 / 2609.27216
KATOsuper: Surrogate-accelerated neural topology optimization with sensitivity-consistent Fourier neural operators
Shengyu Yan, Jasmin Jelovica
cs.CE · cs.AI · cs.LG · math.NA
Abstract
Topology optimization (TO) remains computationally intensive due to repeated finite element analysis (FEA) evaluations required at each iteration. While neural network-based surrogates offer potential acceleration, existing approaches often suffer from gradient inconsistency between predicted objectives and sensitivities, leading to optimization instability. This work presents KATOsuper, an objective-agnostic framework that couples neural-reparameterized topology optimization with a Sensitivity-Consistent Fourier Neural Operator (SC-FNO). The framework employs the forward_split architecture, which derives deployed sensitivities via automatic differentiation through the predicted objective field and thereby preserves consistency between the predicted objective and the gradient used for optimization. The case studies include three 2D benchmark problems and three 3D structures considering compliance or stress minimization. A physics-informed multi-channel input encoding with Fourier position embedding enables resolution-invariant learning, supporting zero-shot extrapolation beyond the training resolution, with useful performance at moderate scaling factors and topology-preserving exploration at up to 64x without retraining. The framework extends to 3D through KATO3D, featuring novel KANConv3D blocks with learnable B-spline activations. KATOsuper demonstrates 15--110x deployment-time speedup over MATLAB baselines while maintaining competitive optimality, with the clearest gains observed in complex 3D and stress-optimization cases. The insight that sensitivity direction matters more than magnitude enables robust optimization even with approximate physics evaluation, extensible to other differentiable physics-driven design objectives.
cs.AI / 42 / 2609.26920
Cross-Modal Contrastive Learning from Histopathology and CT for Automated Renal Cell Carcinoma Grading
Amit Das, Tanmay Shukla, Naofumi Tomita, Faraz Farhadi, Jessica Sin, Ari Hakimi, Chad Vanderbilt, Jie-Fu Chen, Ritesh Kotecha, Weijie Ma, Bing Ren, Saeed Hassanpour
cs.CV · cs.AI
Abstract
Background: Clear cell renal cell carcinoma (ccRCC) exhibits substantial clinical heterogeneity, and accurate grade assessment is essential for risk stratification and treatment planning. However, conventional grading requires invasive tissue sampling. We developed RCC-Align, a cross-modal contrastive learning framework that leverages paired histopathology and computed tomography (CT) data during training to improve noninvasive CT-based ccRCC grade prediction. Methods: RCC-Align aligns paired whole-slide histopathology images (WSIs) and CT scans through contrastive cross-modal objectives, transferring grade-discriminative information from microscopic tissue morphology to macroscopic radiologic representations. The framework was trained and evaluated on paired TCGA and CPTAC cohorts using patient-level five-fold cross-validation. Performance for low- versus high-grade ccRCC classification was compared against CT-only baselines (DINOv2-Base and DINOv2-Finetuned) and a WSI-based reference model (GigaPath-Finetuned). Cross-modal alignment was assessed using cosine similarity analysis. Results: RCC-Align achieved an AUC of 0.601 (95% CI, 0.524-0.673) and AUPRC of 0.599 (95% CI, 0.541-0.676), outperforming DINOv2-Finetuned (AUC 0.545; AUPRC 0.543) with significantly improved low-grade prediction (p = 0.004). RCC-Align also demonstrated stronger paired WSI-CT embedding alignment compared with baselines. The WSI-based GigaPath reference achieved an AUC of 0.719. Conclusion: Pathology-guided contrastive learning improves CT-based ccRCC grading while requiring only CT at inference. This approach may complement tissue diagnosis when biopsy is unsafe, infeasible, or limited by intratumoral heterogeneity. Validation in larger, multi-institutional cohorts with external testing is needed before clinical translation.
cs.AI / 43 / 2609.26923
A 3D Pose-Based Ensemble Framework for Cricket Shot Classification and Automated Biomechanical Analysis
Sourav Shome, M. D. Ashiquzzaman Rahad, Rameswar Debnath
cs.CV · cs.AI
Abstract
Cricket is one of the most celebrated sports world-wide, and technological advancement has become deeply embedded in how the modern game is analyzed and coached. Cricket shot classification and automated performance analysis add a further dimension to this trend. Traditional approaches rely on RGB video features or static images, which are sensitive to environmental variations such as camera angle, lighting, and background clutter, and often fail to capture the underlying biomechanics of batting actions. In this paper, we propose a system to improve cricket coaching that takes raw video data, extracts batsmen from video frames using YOLO, and extracts 3D pose data from video frames using MeTRAbs. The system produces sequential skeletal pose data of 30 body points and captures the biomechanical features of a batsman. As part of the system, we also propose a deep learning ensemble for shot classification of four shots: flick, pull, defense, and drive. The ensemble performed well, compared to existing classification works, achieving 97.68% accuracy. In addition, we analyzed the misclassification rates to identify cases where shots were incorrectly classified and examined their possible causes. Our proposed system allows novice players to obtain useful feedback, such as important joint angles relative to expert batsmen, which can also be useful for injury prevention. The shot classifier also helps track class-wise shots over time for further analysis. In addition to novice players, coaches can use the system for player evaluation.
cs.AI / 44 / 2609.27139
A Hierarchy-Aware Video-Language Model Evaluation and Hyperbolic Baseline for Surgery
Ana Manzano Rodríguez, Pascal Mettes, Marlies P. Schijven, Cees G. M. Snoek
cs.CV · cs.AI
Abstract
Surgical procedures follow a phase-to-step hierarchy, yet the video-language models used to recognize them are evaluated with flat per-level metrics that ignore cross-level coherence and error structure. In this paper we make two contributions to address this problem, (i) we introduce SurgHiBench, the first hierarchy-aware evaluation suite for surgical video understanding, with three tasks measuring recognition, consistency, and severity across granularity levels. We evaluate a general-purpose CLIP model, a Euclidean surgical model, and, as second contribution: (ii) HyperSurg, a new hyperbolic model that enforces phase-step containment via entailment cones, across four (existing) datasets spanning three procedure types. The suite reveals that two models with the same accuracy can produce predictions of very different error severity, ranging from sibling confusions within the correct phase to unrelated cross-phase predictions. Hyperbolic geometry shifts predictions toward the correct procedural neighborhood, and these gains scale with the tree-likeness of each dataset's annotation hierarchy, providing a principled indicator when hierarchy-aware geometry helps.
cs.AI / 45 / 2609.27317
Breaking Weather-Content Coupling: Type-Severity Guided Progressive Disentanglement for All-in-One Infrared Restoration
Xinyao Wang, Lijun He, Zhihan Ren, Fan Li
cs.CV · cs.AI
Abstract
Infrared (IR) imaging is crucial for autonomous driving, remote sensing, and other perception tasks. However, adverse weather may introduce fake structural responses that are entangled with real thermal structures. Existing IR restoration methods are typically designed for a single degradation type or directly reconstruct from degradation-entangled representations. Consequently, they struggle to distinguish intrinsic thermal structures from weather-induced fake responses and to accommodate spatially varying degradation severity, leading to artifacts or the over-suppression of weak but meaningful thermal responses. To address these issues, we propose TSGPD-IR, a type-severity guided progressive disentanglement network for all-in-one infrared restoration that factorizes restoration guidance into task-level weather semantics and region-level degradation severity. Specifically, a Weather and Semantic Co-Guided Multi-Level Prompt Generation Module combines global weather semantics with stage-wise local features to generate adaptive prompts that progressively suppress degradation-induced responses while preserving intrinsic thermal structures. To complement global weather semantics with spatial restoration control, a Proxy-Supervised Regional Degradation Estimator derives severity supervision without manual annotations and predicts spatially varying degradation priors. Guided by these cues, a Multi-Source Collaborative Expert Selection Strategy uses a shared branch to preserve weather-invariant thermal structures and hierarchical routing to select weather-specific expert pools and severity-compatible regional experts. This design progressively separates degradation interference from genuine thermal content and enables region-adaptive restoration, reducing both residual artifacts and over-suppression.
cs.AI / 46 / 2609.27370
Geometry-Conditioned Visual Place Recognition in Natural Environments
Walter Nedov, Saimunur Rahman, Kavindie Katuwandeniya, David Hall, Kaushik Roy, Peyman Moghadam
cs.CV · cs.AI · cs.RO
Abstract
Visual Place Recognition (VPR) in natural environments remains challenging due to repetitive vegetation, sparse distinctive landmarks, and substantial appearance and viewpoint variation across traversals. While visual observations of the same place can change considerably, their underlying spatial structure is often more persistent. We exploit this complementary geometric consistency through Depth-Aware Distillation (DAD), which conditions the token representations of a pretrained Vision Foundation Model (VFM) on geometry inferred by a Geometric Foundation Model (GFM), without any depth sensor. Rather than treating geometry as an additional input modality, DAD projects image-aligned depth into the VFM token space and selectively modulates visual representations through channel-wise geometric conditioning. A two-stage teacher-guided learning strategy first anchors the geometry-conditioned representation to the pretrained appearance space, before refining it for place discrimination. Evaluated on the WildCross benchmark, DAD improves average inter-sequence Recall@1 from 61.41% to 66.37% and Recall@5 from 65.86% to 72.49% over a matched appearance-only baseline, with the largest gains under reverse traversal and long-term appearance variation. These results show that GFM-derived geometry can provide a persistent structural prior for VPR when visual appearance becomes unreliable.
cs.AI / 47 / 2609.27408
What Looks Like a Capability Limit in Vision-Language Models Is a Readout Limit
Alfredo F. Frontera Del Valle
cs.CV · cs.AI · cs.CL
Abstract
Benchmarks for vision-language models offer their answer choices in some convention: a letter, a color name, a pixel coordinate. That convention is treated as neutral. We find it is not, and that the limits a benchmark reports can belong to the readout rather than to the model. On 200 COCO photographs, Qwen3-VL-4B picks the correct one of nine locations for a named object 68.5% of the time when the locations are given in English and 20.0% when the same locations are given as pixel coordinates. Chance is 11.1%. The cost arises when the answer options are coordinates; giving the model a coordinate in the question instead costs 3.5 points and is not significant. The gap holds on a 4x4 grid, under 8-bit rather than 4-bit quantization, and in every slice by object size, boundary distance and category. It also decides which model wins. Two models that tie under English names differ by 39 points in one coordinate system and by 54 in the other, in opposite directions. On the color task, three of the four open models capable of the task show the penalty; on photographs, two of three open models do, and so does Gemini, at 11.1 points on parseable answers (p = 1e-4). GPT-4o does not. To ask whether a model reads a coordinate at all, we attach the wrong name to each one and record which the model follows. Color options written as hue angles are followed below chance; a normalized pixel convention is followed at four times chance. This tells apart conventions a model can use from ones it cannot, though it did not predict accuracy on two untried conventions. Five models also name the same color wheel five different ways, so a fixed answer vocabulary is not neutral across models either. Five times during this work we measured a capable model as incapable because our scorer and the model disagreed about what an answer looks like. We report each case. They are the phenomenon in miniature.
cs.AI / 48 / 2609.27511
NV-Reason-CT: 3D Visual Language Model for CT Analysis
Andriy Myronenko, Dong Yang, Yucheng Tang, Baris Turkbey, Benjamin Simon, Stephanie Harmon, Rikhil Makwana, Mariam Aboian, Sena Azamat, Ibrahim Ethem Hamamci, Sezgin Er, Bjoern Menze, Marc Edgar, Yufan He, Pengfei Guo, Daguang Xu
cs.CV · cs.AI
Abstract
We present NV-Reason-CT, a generative vision--language model for chest and abdominal CT combining native 3D visual encoding with radiologist-guided reasoning. The model couples a native 3D vision transformer with a language model, passing all visual tokens and their explicit 3D coordinates into language decoding without further spatial token merging. This retains volumetric spatial information within the vision encoder and through the language model's positional encoding during joint processing with text. We train on a curated corpus of approximately 550,000 multimodal instruction examples from 70,111 unique CT image inputs, combining standardized reports, abnormality-focused and anatomy-specific questions, multi-turn interactions, and radiologist-authored reasoning from recorded and transcribed expert CT interpretations. Expert annotations provide direct supervision and guide additional report-grounded synthetic reasoning. End-to-end supervised fine-tuning (SFT) is followed by Group Relative Policy Optimization (GRPO), with verifiable rewards over chest and abdominal abnormality sets. The model supports abnormality classification, report generation, and interactive reasoning with reviewable observations, differential diagnoses, and uncertainty. Evaluation spans public CT benchmarks and a held-out NIH cohort. On CT-RATE, NV-Reason-CT achieves a macro-F1 of 0.614 and macro-AUROC of 0.871 without a task-specific classification head; generated reports achieve a report-derived macro-F1 of 0.592. In a preliminary study with expert radiologists, AI-assisted review received favorable confidence ratings and was associated with a 50% reduction in average reported interpretation and reporting time. We release the model and training code to support reproducible research on explainable AI for volumetric medical imaging.
cs.AI / 49 / 2609.27620
InGuard: Towards Generalized Inner Guardrail for Safe Text-to-Image Generation
Zeyu Wang, Xiaodan Li, Zhiwen Li, Yuefeng Chen, Hui Xue
cs.CV · cs.AI
Abstract
Modern text-to-image (T2I) models generate high-quality images from arbitrary user prompts, yet they can just as easily produce not-safe-for-work (NSFW) content. Conventional outer guardrails consist of two components: a prompt classifier that checks for risk before generation, and a post-hoc image classifier that checks the fully generated image. In this design, both classifiers operate outside the generation pipeline and do not use the model's own representations. This separation can limit prompt-screening accuracy, while the image-side check runs only after the full generation cost has been spent. Moreover, a flagged prompt can only be rejected, even when it could be adjusted to produce a safe image. In this work, we propose the Inner Guardrail (InGuard), a safety framework that works inside the pipeline on the model's own representations, leaving base-model parameters untouched. First, a risk classifier grades each prompt as unsafe, risky, or benign based on the text encoder's embeddings, with no external language model. Second, SAGE (Soft-gated Asymmetric Guardrail for Embeddings) modifies the embeddings of risky prompts, aiming to return a safe image instead of a refusal. Third, a latent detector checks the one-step clean latent estimate midway through denoising, reaching nearly image-level performance and halting generation when risk is detected. We also construct the RevGen Safety Benchmark to evaluate T2I safety under realistic conditions: 10,000 prompts built through real-image reverse generation, with a rewriting step that supplies controlled intellectual-property (IP) characters, covering graded porn/gore risks, categorical IP risks, and benign negatives. Across five open-weight T2I models, InGuard reaches 97.9-98.8% safety rate, matching or exceeding the outer guardrail, with 57.5-73.5% less benign disturbance, ~3.7x fewer parameters, and 50-55.6% of denoising steps skipped.
cs.AI / 50 / 2609.28049
Prompt, Probe, Train, or Annotate? Single-camera sports video understanding in amateur settings
Sai Varun Kodathala, Prashanth Pollishetty, Jaylen Cargill
cs.CV · cs.AI
Abstract
Video understanding is usually benchmarked on curated, single-actor, or professionally filmed clips, and a strong score there is routinely read as evidence a model is robust enough for deployment. Amateur team sport is a useful, largely untested place to check that assumption: over eight million students played a school sport in the United States in 2024-25 alone, almost none of it filmed by more than a single fixed camera, with several candidate actors crowded into frame and no operator or second angle to fall back on. Using volleyball as a test case, we ask whether strong performance on general video and world-model benchmarks translates into reliable, per-player attribution once footage is this chaotic, turning footage into statistics through a chain of tasks from finding play boundaries to naming who did what. We evaluate four approaches (prompting and agentic reasoning over frontier vision-language models, classical computer vision with small trained specialists, self-supervised video world models, and manual annotation) at every stage, on 66 amateur matches with 46,648 human-labelled contacts, filmed under conditions no published benchmark uses. No single paradigm wins every stage, and static, single-frame computer vision is not competitive at any stage involving motion or identity. A prompted model segments matches well, yet a far smaller trained model beats it at spotting contacts for a fraction of the cost, and the sport's own rules recover rally outcomes the pixels cannot. Identity is where every automated approach struggles: a jersey number is a static fact temporal reasoning cannot recover if never visible, unlike sporting action, a repeated motor pattern a temporal model can exploit, which is why holistic reasoning improves event detection while identity stays unchanged. We close with where each approach earns its cost, and what transfers beyond volleyball to amateur sport.
cs.AI / 51 / 2609.28231
Do Center Biases Propagate? Robustness of Pathology Foundation Models in Whole-Slide Image Classification
Ilán Carretero, Pablo Meseguer, Rocío del Amor, Valery Naranjo
cs.CV · cs.AI
Abstract
Pathology foundation models (PFMs) have transformed computational pathology through powerful representation learning from histopathological images. PFMs provide rich, discriminative representations for whole slide image (WSI) analysis, enabling tasks such as slide-level classification under multiple instance learning (MIL). However, these representations may also encode non-biological signals associated with acquisition centers, potentially introducing spurious shortcuts into downstream predictions. In this work, we evaluate center-associated robustness in WSI classification using a controlled training setting with increasing class-center correlations quantified by Cramér's V. We benchmark six PFMs across four datasets and two MIL aggregators, while evaluating ComBat as a robustification strategy. We further introduce the Area Under the Cramér's V Curve (AUCC) to jointly capture absolute classification performance and its degradation as spurious correlation increases. Results show that center-related information encoded by PFMs propagates to WSI-level predictions, with robustness depending on both the PFM representation and MIL aggregation strategy. Additionally, ComBat harmonization does not provide consistent robustness gains across datasets.
cs.AI / 52 / 2609.28366
AnchorReasoning: A Visual Grounding and Causal Reasoning Dataset in Long-Tail Autonomous Driving Scenarios
Zhipeng Bao, Wenjie Zhao, Tianle Zhu, Haohua Que, Chence Yang, Geng Yuan, Qianwen Li
cs.CV · cs.AI
Abstract
Vision-language models (VLMs) offer a promising approach to long-tail autonomous driving, but existing driving datasets provide limited supervision for connecting decision-critical visual evidence with reasoning and planning. We introduce AnchorReasoning, a visually grounded reasoning dataset built on WOD-E2E, containing 416,119 annotated frames and 395,379 decision-critical elements across four major categories and 19 fine-grained types. Each frame is organized as a visually grounded chain-of-thought (VG-CoT) that links decision-critical element identification and localization, element attributes and implications, driving-action rationale, and action and trajectory planning. We further develop a curriculum supervised fine-tuning strategy that progressively learns these hierarchical capabilities, together with an object-size-aware grounding metric for evaluating localization quality. Experiments across eight general-purpose, embodied-AI, and AV-specific backbones show that VG-CoT supervision improves grounded reasoning and trajectory prediction. Across models, 5-s ADE and FDE decrease by 7.84 and 11.86, while RFS Frame and Cluster improve by 1.66 and 1.70. These gains are achieved with 18.5 fewer reasoning tokens and 0.32 s/frame lower inference latency on average, demonstrating the value of visually grounded, decision-focused supervision for VLM reasoning and planning in long-tail autonomous driving.
cs.AI / 53 / 2609.28414
Frozen Flows Forget: Diagnosing and Restoring Lost Motion in a Latent-flow World Model
Xiwen Chen, Rigaudiere Z. Li, Zhiruo Zhou, Xiaojun Zhu, Houde Liu
cs.CV · cs.AI · cs.RO
Abstract
Latent world models that integrate a flow in a frozen self supervised latent space train stably and cheaply, yet silently lose the property manipulation depends on most: motion. The pretrained flow never moves the manipulated object; retraining it with latent-only losses only trades stillness for teleport-like motion. We trace the failure to the training signal, not the representation: anchor-sparse, latent-only supervision never says where along the horizon change belongs. Decode-augmented rollout training (DART) repairs this while keeping the representation frozen, retraining only the flow with decode-path supervision. DART outperforms its latent only parent on the full protocol, restores the temporal structure of motion, and re-couples predicted motion to the scene; at larger scale it further improves prediction quality, closing nearly half the remaining gap to an oracle-informed interpolation reference. Finally, we report an unexpected finding about evaluation: pixel error alone rewards frozen predictions.
cs.AI / 54 / 2609.28137
"We'll Fix It Later": Education, AI, and the Deferral of Privacy in EdTech
Meghna Manoj Nair, Rachel Greenstadt
cs.CY · cs.AI
Abstract
Educational technology (EdTech) platforms collect highly sensitive student data, including behavioral logs, disability records, and academic histories. However, privacy considerations are often postponed rather than treated as a foundational design requirement. We present a mixed-methods study combining 12 semi-structured interviews with EdTech professionals and a privacy policy audit of 48 platforms coded across five dimensions, with strong inter-rater reliability (mean Cohen's Kappa = 0.781). Our interviews reveal a recurring organizational pattern in which privacy is recognized as important but deferred across the product lifecycle as organizations prioritize product functionality, growth, funding, and immediate educational outcomes. Responsibility is often delegated to cloud providers, policy documents, or downstream institutions, while limited privacy-related feedback gives organizations little pressure to change these practices. The policy analysis reflects these patterns: platforms describe what data they collect relatively well but provide substantially less information about how that data is subsequently governed. Thirty-three percent make no meaningful Artificial Intelligence (AI) disclosure despite visible AI features, and 73% provide only generic accountability and breach-response language. K-12 platforms perform better on children's consent where regulation creates explicit requirements, but this advantage does not extend to AI governance or accountability. These findings suggest that meaningful improvement requires enforceable institutional and regulatory mechanisms rather than voluntary privacy commitments alone.
cs.AI / 55 / 2609.27040
EMA: Elastic and Performance Transparent Memory Across GPUs
Yi Xu, Tian Xia, Ion Stoica
cs.DC · cs.AI · cs.LG
Abstract
Multi-GPU servers have become the standard building block of modern data centers, providing aggregated capacity through high-bandwidth interconnects. At the same time, workloads such as LLM inference exhibit highly dynamic memory demands, which can cause one GPU to exhaust its local memory while others remain underutilized. This mismatch motivates a model of elastic resource sharing across GPUs. We present EMA, a memory sharing system that allows GPUs within a server to borrow and reclaim memory from each other, forming an elastic pool of capacity. EMA ensures performance transparency for both borrowers and lenders. For borrowers, prefetching hides remote access costs so that applications experience remote and local memory as indistinguishable in performance. For lenders, borrowed resources remain reclaimable on demand, guaranteeing that performance never falls below that of static partitioning. While our design focuses on memory, the same principle naturally extends to other GPU resources. Our evaluation shows that EMA improves individual user throughput by up to 52%, achieves 96% of the throughput of a system provisioned with 2X capacity, and maintains latency similar to the static local baseline.
cs.AI / 56 / 2609.27085
Crossflow: Prefill-Decode Elasticity for Agentic LLM Serving
Yi Xu, Ehsan K. Ardestani, Wenyin Fu, Martin Schatz, Krishna Malladi, Zhan Shu, Adnan Aziz, Shobhit Kanaujia, Ajit Mathews, Chunqiang Tang
cs.DC · cs.AI · cs.LG
Abstract
As serving capacity demand surpasses that of training, serving efficiency becomes increasingly important. Prefill-decode (P/D) disaggregation improves serving efficiency through specialization and isolation of the two phases. These benefits rest on a static partitioning. Phase demand, however, is not static. We observe that in a large LLM fleet the ratio of uncached input to output tokens has peak-to-mean ratios up to 4.7x at minute timescales, and that in a public agentic trace the hourly ratio spans a median 24.5x within a single day, while reassigning a replica takes tens of minutes. Agentic traffic sharpens the mismatch. Sizing each pool at its ninety-fifth percentile leaves up to 17% of cluster capacity unused; sizing below it converts the same imbalance into queueing and unrealized throughput. We present Crossflow, which makes this boundary elastic without changing node roles. Each decode node publishes a short-lived, revocable lease that bounds local-prefill compute, KV capacity, transfer work, and projected output. Across public and internal traces, Crossflow improves token throughput by 16.2-17.4% on geometric mean over static P/D, and by up to 43.4% at high load, while reducing mean TTFT at every evaluated point.
cs.AI / 57 / 2609.27246
Listening and Mirroring: The Effects of Verbal Attunement and Behavioral Mimicry on Social and Empathic Perceptions of Embodied AI Agents in VR
Nathalia Gomez, Haig Shamlian, Omar Khan, Tiffany D. Do
cs.HC · cs.AI
Abstract
As embodied agents take on increasingly social and relational roles in VR, visual realism and embodiment alone may be insufficient; users must also perceive these agents as emotionally attuned, supportive, and humanlike. Prior work suggests that verbal attunement and nonverbal mimicry can each improve users' social evaluations of embodied agents. However, behavioral mimicry has largely been studied outside of real-time, conversational AI interactions, leaving limited understanding of how users respond when an agent simultaneously generates contextually responsive dialogue and adapts its nonverbal behavior during an immersive conversation. To address this gap, we developed an embodied AI counselor that combines conversational AI with real-time facial-expression and posture mimicry, while producing either verbally attuned or neutral responses. We evaluated the system in a 2 X 2 within-subjects study with 20 participants, manipulating verbal attunement and behavioral mimicry. Results showed that verbal attunement was the most reliable driver of perceived empathy. Behavioral mimicry showed a marginal relationship with perceived humanness, while greater mimicry exposure showed preliminary, exploratory positive associations with empathy, positivity, and humanness, particularly among female participants. Together, these findings show that multimodal synchrony is not a simple additive strategy for designing empathic conversational agents in VR and underscore the need to consider how verbal and nonverbal behaviors are combined during real-time interaction.
cs.AI / 58 / 2609.27175
Self-Evolving Multimedia Verification through Memory Consolidation of Contestation Experiences
Truong Thanh Hung Nguyen, Vo Thanh Khang Nguyen, Hoang-Loc Cao, Phuc Ho, Truong Thinh Nguyen, Van Pham, Hung Cao
cs.MM · cs.AI · cs.MA
Abstract
Multimedia verification requires not only accurate decisions but also traceable evidence, reliable human correction, and safe reuse of prior experience. Existing systems often lack explicit mechanisms for revising intermediate reasoning or preventing harmful knowledge transfer. We present SEMV (Self-Evolving Multimedia Verification), a self-evolving multi-agent framework that treats provenance-bearing arguments as the interface between evidence, reasoning, human contestation, and memory. SEMV combines arena-based quantitative bipolar argumentation (A-QBAF), causal and scoped revision, and verification-gated memory consolidation with explicit conflict retention. On COSMOS benchmark, SEMV achieves 91.88% accuracy versus 89.10% for the strongest comparable baseline. Verified memory reduces negative transfer from 5.7% to 0.2%. On CTR benchmark, constructed from reviewer contestations, scoped causal revision corrects 96.7% of initial errors while saving 52.8% compute. MV2026 Grand Challenge dataset further supports evidence-grounded, temporally consistent reporting. These results show that SEMV can evolve through verified experience while keeping accumulated knowledge and subsequent decisions traceable, revisable, and contestable.
cs.AI / 59 / 2609.27095
Intelligence Across Embodiments
Bo Ai, Henrik I. Christensen, Hao Su
cs.RO · cs.AI · cs.LG
Abstract
Robotic embodiment encompasses the sensing, kinematics, dynamics, geometry, actuation, and control through which an agent physically interacts with the world. These properties vary across robots and change over time. We argue that general embodied intelligence requires learning that accumulates across these differences. Prevailing methods that engineer correspondences to bridge embodiment differences offer immediate practical gains, but their assumptions limit the scope of transfer in the long run. Instead, a more general approach should discover representations that support transfer to a larger range of embodiments as experience grows. We propose embodiment diversity as a promising axis of scaling, and identify broad learned priors as a complementary ingredient. We call for evaluations that better characterize embodiment gaps and transfer performance. More broadly, cross-embodiment learning connects the practical challenge of learning from heterogeneous robot experience with a broader scientific pursuit inspired by nature - physical intelligence that adapts and co-evolves with its embodiments to gain agency over its behavior and physical forms.
cs.AI / 60 / 2609.27312
Turning Safety into Competence: Minimally Exploitable Robot Policies via Safety-Filtered Reinforcement Learning
Ruihan Wu, Rui Yang, Donggeon David Oh, Duy Nguyen, Haimin Hu
cs.RO · cs.AI · cs.LG · eess.SY
Abstract
Robots deployed for competitive tasks must outmaneuver their opponents without sacrificing safety. Existing approaches, including safe reinforcement learning (RL), train a single policy to achieve task success and avoid failures simultaneously. This coupling can complicate training and leave the learned policy exploitable by deliberate attacks. We propose Safety to Competence (S2C), a two-stage RL framework that separates safety synthesis from competitive task learning. We formulate competitive interactions as safety-critical Markov games and prove that perfect filtering preserves policy non-exploitability when all players commit to safe maneuvers. S2C learns a robust safety filter via adversarial RL, embeds it in the environment during task policy training, and retains the same filter at deployment. In simulated touchdown games, S2C outperforms eight safe RL baselines, achieving the highest win rate and Elo rating, and the lowest exploitability. Hardware stress tests against a human opponent confirm S2C's competence.
cs.AI / 61 / 2609.27394
Automotive mmWave Spinning Radar Place Recognition with Spatially Gated Feature-Correlation Representation
Saimunur Rahman, Sagun Singh Shrestha, Abdelwahed Khamis, Peyman Moghadam
cs.RO · cs.AI · cs.CV
Abstract
Automotive spinning FMCW radar provides dense, $360^\circ$ sensing and remains reliable under poor illumination and adverse weather, making it well-suited to autonomous navigation. Place recognition uses these observations to identify previously visited locations for re-localization and long-term navigation. However, heading changes appear as circular shifts in the polar radar representation, and conventional global aggregation can lose relationships among radar responses that are important for distinguishing similar places. We propose SGCA-Net, a spinning radar place recognition framework that combines rotation-robust feature extraction with Spatially Gated Correlation Aggregation (SGCA). SGCA learns spatial weights to reduce the influence of unstable and ambiguous radar regions, while aggregating pairwise correlations among local responses to preserve informative feature relationships. Experiments on the MulRan dataset show that SGCA-Net consistently outperforms SOTA methods across urban, campus, and open-road environments, while remaining robust to substantial heading variation. Evaluation on the HeRCULES dataset further demonstrates that SGCA-Net generalizes to unseen environments and radar sensors without fine-tuning.
cs.AI / 62 / 2609.27450
BEE: Intervention-Adaptive Real-World Reinforcement Learning with Vision-Language-Action Models
Weihui Zhao, Xiaohan Yan, Zunian Wan, Xuan Du, Zhaozhan Chi, Jianbo Mao, Ruipu Wu, Rushuai Yang, Houlin Li, Shukai Yang, Jing Wu, Yuxiang Yan, Yongcheng Liu, Chuankang Li, Guanghui Ren, Wei Shan, Maoqing Yao
cs.RO · cs.AI
Abstract
Vision-language-action (VLA) models handle long-horizon manipulation, yet success hinges on a few precision-critical phases where millimeter-scale errors undo all prior progress. Online reinforcement learning (RL) can optimize exactly these actions, but free exploration is far too costly on real robots, which makes human corrections indispensable. However, existing online RL methods for VLAs either cannot incorporate such corrections or fold them into undifferentiated supervision. Yet human corrections are not uniformly noisy but reliable along some action dimensions and variable along others. Building on this, we introduce BEE, an intervention-adaptive framework for real-world RL on a frozen VLA that lets the policy go BEyond Expert imitation. We formulate human corrections not as actions to reproduce but as evidence about a constraint: a Correction Model predicts how a human would correct a given VLA proposal and how consistent the correction is along each action dimension. This predicted consistency sets the per-dimension tightness of a constraint on policy optimization. Where corrections are consistent the policy stays close to the human, and where they vary, the constraint relaxes. We evaluate BEE on three real-world manipulation tasks and one LIBERO-Pro simulation task at a matched online-data budget. BEE attains the highest success rate on every task, 91.2% on average against 57.5% for RLT and 42.1% for DSRL, and the lowest human intervention rate on all real-world tasks.
cs.AI / 63 / 2609.27467
Kairos: Grounded Forecasting of Presence and Directional Flow in 4D Scene Graphs
Iacopo Catalano, Julio A. Placed, Javier Civera, Jorge Peña Queralta
cs.RO · cs.AI
Abstract
Long-term autonomy in human-populated environments requires anticipating whether and how people will move at times a robot has not yet observed. Existing representations of pedestrian motion face a tradeoff: they either forecast future activity, reducing each location to a scalar rate, or model the full directional distribution, holding it fixed in time. We present Kairos, a predictive directional-flow memory that extends a hierarchical 3D scene graph (3DSG) to a 4D scene graph (4DSG). Every observed voxel of the reconstructed geometry stores a directional mixture and a presence rate, and spectral predictors forecast, for any future query time, both the probability that people are present and the full directional distribution of their motion. Pairwise flow dependence between adjacent voxels supports conditional queries, and per-voxel predictive variances yield calibrated credible intervals that tighten as observations accumulate. We evaluate Kairos on three real pedestrian environments: a robot-collected campus dataset, a shopping mall, and a station concourse recorded continuously for eleven months. Its learned state remains consistent under loop-closure corrections, and its forecasts are competitive with dedicated occupancy and flow models trained on the full detection stream, although Kairos learns from only the small fraction available to a patrolling robot. Finally, we validate the representation on a downstream encounter-probability planning task, where plans computed over the Kairos forecasts encounter more people than plans computed over any time-invariant map at an equal success rate. We provide the code at https://github.com/IacopomC/kairos.
cs.AI / 64 / 2609.27536
Behaviora - A Conceptual Architecture for External and Internal Behavior of Robots and Agents
Gote Nyman
cs.RO · cs.AI
Abstract
Behaviora is a preliminary conceptual architecture for representing agent and robot behavior, external and internal alike, in an addressable form. A behaving robot or agent performs a Behavior Episode composed of episode components, which can be derived from behavior taxonomies (BTax) and assigned persistent identifiers. We denote these identifiers as IoB (Internet of Behaviors) Addresses. A Behavior Episode specifies what the system does, while a Style Profile (SP) specifies how this behavior is expressed. Style can communicate characteristics of the actor and qualities such as competence and cultural manners. An Experience Profile (EP) represents behaviorally relevant internal state that modulates the execution of an Episode. Finally, a Behavior Compiler maps these behavioral representations to platform-specific actions. We use a primitive touching arm model to show these components and their relations. External Behavior is a result of addressable movements and their styles. Internal Behavior is represented through the same episodic principle and can be rendered as inner speech. Sensing, perception and complex task contexts have not been included in the present implementation, although a conceptual place is reserved for them.
cs.AI / 65 / 2609.27656
InternW0: A Foundational Physical World Model for Efficient Real-World Interactions
Jisong Cai, Yao Mu, Ganlin Yang, Zhe Cao, Zhangzheng Tu, Xing Gao, Kailin Li, Xinyu Zhan, Lixin Yang, Yangkun Zhu, Haoxiang Ma, Ming Zhou, Qiaojun Yu, Yufei Xue, Liqun He, Yifei Yao, Yifan Zhu, Long Ling, Bingqi Jiang, Haoyu Guo, Xueyue Zhu, Bowen Zhou, Bin Zhao, Tianfan Xue, Chunhua Shen, Weinan Zhang
cs.RO · cs.AI
Abstract
Physical intelligence requires more than predicting how the world may evolve: predictions must remain actionable as the world continues to change. We introduce InternW0, the first instantiation of the InternW physical world model series from Shanghai AI Laboratory, built around omnimodal interfaces, asynchronous multi-frequency processing, and local physical modeling under partial observations and external influences. InternW0 jointly learns future visual dynamics and continuous robot control through an asymmetric video--action architecture with flow matching. A high-capacity video expert provides longer-horizon predictive context, while a lightweight action expert operates at a faster timescale. Instead of regenerating the future for every action update, InternW0 reuses layerwise K/V and adapts it to newly observed states through observation-conditioned context routing. Domain-specific interfaces and soft prompts support heterogeneous embodiments, while contact-aware post-training incorporates force and tactile signals for contact-rich manipulation. We train InternW0 on approximately 7,200 hours of heterogeneous robot and egocentric data, including EgoLab, a 275-hour real-laboratory egocentric dataset. Evaluation spans simulation benchmarks and real-world scientific tasks, including a 15-stage metal--organic framework synthesis workflow and 5-stage contact- and force-aware dexterous manipulation for general-purpose quantitative pipetting. These results advance scalable, asynchronous, and science-native physical world models for universal and efficient real-world interactions.
cs.AI / 66 / 2609.27734
InfiNoVA: Infinite Novel View Augmentation for Viewpoint Invariant Robot Policies
Sai Puneeth Reddy Gottam, Elmar Rueckert, Vedant Dave
cs.RO · cs.AI
Abstract
Vision-Language-Action (VLA) policies often rely strongly on the camera viewpoints seen during training, causing substantial performance degradation when deployed from unseen perspectives. Collecting demonstrations from sufficiently diverse physical viewpoints is expensive and still provides only sparse coverage of the viewpoint space. We introduce InfiNoVA, a data-augmentation framework that converts synchronized multi-camera demonstrations into a dense distribution of geometrically consistent training views. InfiNoVA reconstructs each manipulation trajectory as a time-varying 3D Gaussian representation and renders novel observations from sampled camera poses while preserving the original state-action correspondence. This explicit scene representation improves frame-level fidelity and temporal consistency while reducing task-critical hallucinations observed in generative novel-view synthesis. Across four real-world manipulation tasks, policies trained with InfiNoVA achieve 5.4x higher average success under unseen randomized viewpoints than both VISTA-based augmentation and the unaugmented policy. InfiNoVA further achieves 1.7x higher success than training directly on all five physical camera views. These results show that dense, geometrically grounded viewpoint augmentation provides a practical route toward camera-robust robot policies without modifying the underlying policy architecture.
cs.AI / 67 / 2609.28107
Distillation for Efficient Multitask Manipulation Policies via Conditional Flow Matching
Shreya Deshmukh, Imen Mahdi, Nick Heppert, Abhinav Valada
cs.RO · cs.AI · cs.LG
Abstract
Advances in generative modeling have recently been extensively employed in robotics for policy learning. In particular, Conditional Flow Matching (CFM) trained with expert demonstrations has been shown to outperform existing methods on robot manipulation benchmarks. While prior work has mainly focused on single-task settings, we study the problem from a multi-task perspective, as training independent models for each task is computationally expensive. Multi-Task policy learning comes with its own set of challenges, as naively training on a concatenated dataset of demonstrations would either require increased model capacity to accommodate the added complexity or result in drops in performance. We propose to distill knowledge from single-task CFM experts into a shared multi-task policy by transferring their learned velocity fields. We combine this distillation signal with the original CFM objective to retain fidelity to the demonstrations. Experiments on RLBench show that our approach improves multi-task policy performance over naive training while maintaining a fixed model size.
cs.AI / 68 / 2609.28256
MemBodied: Recurrent Associative Memory for Vision-Language-Action Models
Tej Deep Pala, Navonil Majumder, Bryce Goh, Raphael Yee, Jianfei Yang, Liming Chen, Soujanya Poria
cs.RO · cs.AI · cs.CV
Abstract
Vision-Language-Action models provide a strong foundation for general-purpose robot control, yet a vast majority of policies do not preserve and leverage episode-level information beyond the current observation. This limitation is consequential in history-dependent manipulation tasks that depend on information available only in past observations. Retaining past observations in context can aid in recovering this information, but at the significant cost of ever-growing, bloated context and inference latency. We thus introduce MemBodied, a fixed-size episodic memory with two complementary components: an associative state that records interactions across policy calls and an episode anchor that preserves a compact representation of the initial scene as a reference. At each policy call, the model conditions action generation on the current input and the memory components, rather than directly using past observations. Across five evaluated RMBench tasks requiring memory, MemBodied achieves $7.81\times$ the mean success rate of a stateless policy and $2.98\times$ of vanilla recurrent memory, while outperforming the strongest memory-augmented baseline by $1.3\times$ with $10\times$ fewer added parameters. On the fully observable LIBERO-Long suite, it reached 90.6%, a 5.4% improvement over the stateless $π_0$ policy. These findings support MemBodied as a practical alternative to expanding the policy context for history-dependent manipulation.
cs.AI / 69 / 2609.28467
Where Should I Join? Robot Group Joining via Language-Guided Goal Prediction
Zilin Fang, Zishuo Wang, Gim Hee Lee, David Hsu
cs.RO · cs.AI
Abstract
Social navigation typically assumes a specified goal and focuses on reaching it while respecting social conventions, whereas robot group joining requires predicting where to join based on the group's real-time activity and formation. This is a highly semantic task, yet an important capability for applications such as robotic guide dogs and autonomous mobility scooters. We formulate language-grounded robot group joining: given an observation and a natural-language description of a target group, the robot identifies the relevant group members and predicts socially compliant joining poses. For grounding, we generate structured candidate subsets through recursive spectral partitioning and rank them with a language-conditioned image--geometry model. Given the grounded group, a goal predictor leverages human-formation priors to produce a multimodal energy--orientation map over feasible robot poses. Experiments on conversations, queues, and audiences across varying group sizes, crowd densities, and visual ambiguities show that our method achieves competitive grounding accuracy with sub-second inference and outperforms all baselines in joining-pose prediction. Real-robot experiments further demonstrate group joining in both static and dynamically changing interactions.
cs.AI / 70 / 2609.27378
Psychoacoustically Aligned Latent Smoothing for Adversarial Robustness of Full-Duplex Speech-to-Speech Dialogue Models
Kian Shamsaie, Iman Modarressi
cs.SD · cs.AI · cs.CL · cs.HC
Abstract
End-to-end speech-to-speech dialogue models listen and speak simultaneously, so a continuously open acoustic channel is exposed to adversarial manipulation. We formalize imperceptible attacks on full-duplex agents as optimization over additive perturbations confined beneath the psychoacoustic masking threshold of the carrier speech, under three goals: targeted semantic hijacking, response suppression, and policy jailbreaking. Against an undefended Moshi-style agent, white-box attacks succeed in up to 91.7% of trials. We then introduce psychoacoustically aligned latent smoothing (PALS), which injects anisotropic Gaussian noise shaped by local codebook covariance at the residual-vector-quantized latent interface, with input noise shaped by the masking threshold constraining the attacker and trained by a Kullback--Leibler consistency objective. Deployed with no inference-time cost, PALS reduces hijack to 8.3%, mute to 11.2%, and jailbreak to 9.1% at clean quality within 2.3%. A Monte Carlo-smoothed variant certifies an ellipsoidal latent radius up to 0.616, a guaranteed floor that the empirical robustness far exceeds.
cs.AI / 71 / 2609.27399
Forget who you Forgot: Speaker Unlearning to Prevent Re-Identification in Zero-Shot Text-to-Speech
Hyoeun Kim, Yujun Lee, Kyuhong Shim
cs.SD · cs.AI
Abstract
Recent zero-shot text-to-speech (ZS-TTS) systems can reproduce a speaker's voice with high fidelity from only a few seconds of reference speech, raising concerns over unauthorized voice cloning and impersonation. Speaker identity unlearning has recently emerged as an approach to selectively suppress this capability for speakers who opt out while preserving synthesis capability for other speakers. Although existing approaches reduce speaker similarity, preventing re-identification often faces severe degradation of speech quality. Motivated by this observation, we propose GUARD, a lightweight speaker identity unlearning framework that combines a learned speaker gate with speaker-agnostic activation steering on a frozen TTS backbone. The steering vectors are optimized using group-relative reward optimization to shift outputs from forget speakers toward population-level impostor similarity while preserving intelligibility and speech naturalness. On CosyVoice2, GUARD reduces forget-speaker similarity from 0.541 to 0.103 and re-identification accuracy in a 150-speaker gallery from 73.5% to 0.5%, while preserving retain-speaker reproduction. The results demonstrate that similarity reduction alone may not fully characterize successful speaker identity unlearning and highlight re-identification as a complementary criterion for its evaluation.
cs.AI / 72 / 2609.27489
Passing: An Endless Journey through Reconstructed Spacetime with AI-Generated Sound
Akira Takahashi, Chihiro Nagashima, Zhi Zhong, Shusuke Takahashi, Yuki Mitsufuji
cs.SD · cs.AI
Abstract
This paper introduces Passing, an interactive audiovisual installation that generates an endless journey from a single continuous monorail-window recording by reconstructing it as a spatiotemporal volume. Rather than replaying the footage linearly, the work resamples its spatial and temporal structure along nonlinear trajectories, producing a continuously passing landscape whose depth, speed, and temporal order become unstable. A camera-based viewer-presence detection system estimates whether a viewer is present in the viewing zone and uses this presence state to influence transitions among rendered video sequences. The resulting video stream is fed into SpecMaskFoley, a real-time video-to-audio synthesis model that generates a synchronized soundscape for the reconfigured image. The model is not used to reconstruct an objectively correct soundtrack, but functions as a speculative listener, proposing a possible auditory interpretation of a world whose conventional spatial and temporal premises have been disrupted. Passing distributes creative agency across the artist, who defines the rules of spacetime reconstruction; the AI model, which interprets the emergent visual flow as sound; and the audience, whose embodied presence influences the audiovisual trajectory. Through this structure, the work investigates how authorship and listening may be negotiated among human intention, machine inference, and audience interpretation. Artwork page: https://ryufurusawa.com/passing
cs.AI / 73 / 2609.27024
Loss Choice or Model Choice? The Role of Forecast Level in Cryptocurrency Volatility Forecasting
Andrzej Tokajuk, Jarosław A. Chudziak
q-fin.CP · cs.AI · cs.CE · q-fin.RM
Abstract
Volatility forecasts play a central role in financial risk management because their overall level and day-to-day movements affect downstream decisions. Most studies compare forecasting models while keeping the training loss fixed. Yet losses emphasise different errors and can target different properties of future volatility, so raw comparisons may combine persistent forecast-level differences with differences in daily forecast movements. This leaves unresolved whether the importance of loss choice comes mainly from the forecast level it targets or from differences that remain after level adjustment. We address this gap through a comparison of seven losses and five models across major cryptocurrencies. Validation-based alignment adjusts the forecast level before the raw and aligned forecasts are evaluated using statistical scores and one-day Value-at-Risk. Before alignment, marginal score variation is greater across losses. After alignment, model choice becomes the larger source of variation in the full five-model comparison, while cross-loss differences in VaR breach rates narrow substantially. Our contribution is a comprehensive evaluation of loss and model choice that shows why losses can appear so influential in raw comparisons and how this interpretation changes when forecast level and downstream risk are considered explicitly.
cs.AI / 74 / 2609.27632
Compliant AI Infrastructure for Regulated Finance: A tiered multi-agent framework with DLT audit trails for financial operations in DACH
Walter Kurz, Reinhard Magg
q-fin.GN · cs.AI
Abstract
We present a compliance-first architecture for AI in regulated finance that treats regulation as an orientation layer rather than a deterministic ruleset. A matrix of regulatory intent and exposure provides a compact classification handle, which a governed policy compiler then maps into concrete prohibitions, obligations and runtime budgets. Prohibitions constrain feasibility and block externalisation, while obligations extend tasks with artefacts that must meet explicit admissibility criteria. Committee activation remains policy-driven and proportionate, preserving efficiency while ensuring supervisory oversight. Evidence, decisions and reason codes are bound to a permissioned DAG with deterministic timestamping, enabling replay, provenance checks and clear attribution of failure. Clause-level legal indexing with effective dates and capability-based agent routing ensure portability across DACH and the wider EU. The result is assurance by construction: compliance is embedded in execution and verifiable by auditors without sacrificing proportionality or transparency.
cs.AI / 75 / 2609.27299
Beyond the Illusion of Power: Calibrating Quasi-Experiments in Observational IS
Spandan Ghose Chowdhury
stat.ME · cs.AI · cs.LG
Abstract
Information systems (IS) researchers increasingly use quasi-experimental methods such as difference-in-differences (DiD) and instrumental variables (IV) to recover causal effects from observational panel data. Power calculations that justify these designs assume i.i.d. errors, but the deeper problem is what even a cluster-robust calculator cannot see. We report a Monte Carlo study over 9837 parameter conditions (approx 9.8 million datasets) and decompose the planned-versus-achieved power gap. The serial-correlation component is recoverable by an AR(1)-aware calculator when rho is known, and partially when rho must be estimated from short pre-periods, but panel attrition, staggered-adoption bias, and parallel-trends pretesting are captured by no closed-form formula; exogenous attrition alone costs approx 8 to 11 percentage points at the few-hundred-to-thousand sample sizes IS studies use. Treatment-correlated, outcome-dependent attrition instead induces bias, not just power loss. For IV, holding first-stage F fixed, larger N neither raises power nor curbs exclusion bias, though with a fixed instrument more data does sharpen the first stage, so identification rests on instrument strength, not sample size.
机器学习 (cs.LG)
108
cs.LG / 1 / 2609.28448
Nonequilibrium Phases of Repulsive Self-Attention: Chaos, Attention Condensation, and Emergent Locality
Qucheng Gao, Zuyi Yang, Xiao Chen
cond-mat.dis-nn · cond-mat.stat-mech · cs.LG
Abstract
We study the nonequilibrium dynamics of a minimal recurrent transformer with $N$ normalized tokens, $Q=K=I$, and a negative value map $V=-I$. Similarity-based attention selects nearby representations, while the negative value map drives tokens away from the selected field. This feedback can continually reorganize both the representation geometry and the attention network. For $d=2$, the tokens lie on a circle, where the regular polygon is an exact fixed point. As the attention feedback strength $γ$ is increased, the polygon loses stability through a flip bifurcation, giving rise to period-two motion, chaos, and cluster-exchange or cluster-flip states. Despite this temporal complexity, attention remains diffuse as $N\to\infty$ at finite fixed softmax sharpness $β$. Attention condensation instead emerges in the scaling regime $β\sim N^2$. In the hard-routing limit, repulsive updates amplify local perturbations and routing-partner switches transmit them ballistically, producing an emergent butterfly cone in representation space. High-dimensional geometry provides a distinct route to localization. For $d=N\to\infty$, simulations from Gaussian initial conditions provide evidence for a condensation transition at $β=O(1)$, driven by dynamically generated finite overlap gaps. Depending on $γ$, the resulting phases include diffuse simplex-like states, consensus flips, condensed active routing with signatures of chaos, and fragmented cluster flips. These results establish temporal activity, attention condensation, and geometric clustering as distinct collective phenomena, and show that sparse attention can sustain persistent dynamics rather than freeze it.
cs.LG / 2 / 2609.28013
SoLiD26: A First Principles Solid-Liquid Interface Dataset for Machine-learned Interatomic Potentials
Jonas Busk, Emil J. P. Frost, Yogeshwaran Krishnan, Henrik H. Kristoffersen, August E. G. Mikkelsen, Xueping Qin, Xin Yang, Heine A. Hansen, Arghya Bhowmik, Tejs Vegge
cond-mat.mtrl-sci · cs.LG · physics.chem-ph · physics.comp-ph
Abstract
Machine-learned interatomic potentials (MLIPs) for solid-liquid interfaces in advanced materials applications, e.g., electrochemistry, catalysis and corrosion, require training data that samples both liquid environments, the solid and the interface itself. We present SoLiD26, a curated solid-liquid interface dataset, containing 15.4 million first-principles atomic structures with up to 576 atoms and 15 chemical elements for training and evaluating MLIPs. The structures were compiled from density functional theory (DFT) calculations performed in studies of solid-liquid interfaces, with most configurations originating from ab initio molecular dynamics (AIMD) simulations. Each record contains atomic species, positions, simulation cell, periodic boundary conditions, potential energy and atomic forces. SoLiD26 includes aqueous coinage metal interfaces, electrode-electrolyte systems, and selected bulk reference structures, calculated with VASP using the PBE functional and D3 dispersion corrections. We describe the data ingestion and preparation pipeline used to construct the dataset. The application of SoLiD26 for training and evaluating MLIPs is demonstrated with a suite of MACE models on a simple training, validation and test split. The dataset enables development and benchmarking of MLIPs for structurally and chemically heterogeneous solid-liquid interfaces.
cs.LG / 3 / 2609.27123
PEARL: A Lightweight Prompt-based Feature Interpreter Framework for Real-Time, Anonymous, and Heterogeneous Collaborative Perception
Armin Maleki, Hayder Radha
cs.CV · cs.LG
Abstract
Heterogeneity across Collaborative Perception (CP) agents is a major challenge for emerging CP frameworks due to domain gaps from differing sensors, architectures, and training data. Prior works mitigate this challenge by aligning features in a unified space via model retraining or per-agent-type interpreters. These strategies (a) require access to neighbor configurations, (b) do not fully address real-time CP deployment, and (c) generalize poorly to unseen agents joining at run time. To overcome these challenges, we present PEARL, a Prompt-Embedding framework for Anonymous and Real-time Lightweight heterogeneous CP. PEARL supports multiple CP interpreters and selects one for a new-joining agent in real time using two lightweight, multi-scale interpreters trained in parallel: a sparse-detection (LWSD) interpreter that aligns salient regions for cooperative detection, and a dense, domain-invariant (LWDDI) interpreter that produces agent-invariant features for fast interpreter selection. Both interpreters use low-rank visual prompts to reduce computation, storage, and model complexity. Extensive experiments on simulated (OPV2V, V2XSet) and real (DAIR-V2X) datasets show that PEARL generalizes across simulated and real-world cooperative driving scenarios. Its real-time model-selection strategy yields an 8.2% Average Precision (AP) gain over a random-selection baseline while running in 1.67 ms on average. Although primarily designed for real-time CP, PEARL also outperforms state-of-the-art heterogeneous CP frameworks under traditional offline training by 5.6% AP on average while reducing communication cost by up to 34.7 times. Equally important, PEARL does not require sharing agents' configurations or model settings, thereby protecting information that may be proprietary or private. These results establish PEARL as a scalable and practical framework for heterogeneous collaborative perception.
cs.LG / 4 / 2609.27208
Benchmarking Active Spot Selection for Cost-Efficient Spatial Transcriptomics
Zheyu Zhu, Junchao Zhu, Fengbei Liu, Tianyuan Yao, Gelei Xu, John Cannon, Haichun Yang, Yuankai Huo, Mert R. Sabuncu, Ruining Deng
cs.CV · cs.LG · q-bio.OT
Abstract
Spatial transcriptomics (ST) measures gene expression in tissue context, but dense capture grids can be costly and may repeatedly sample morphologically similar regions. Most active learning strategies were developed for categorical labels and independent samples. We conduct a retrospective pool-based benchmark of active learning versus uniform Random sampling for ST, where expression vectors are high-dimensional and continuous and candidates are spatially correlated. Using two fully profiled public ST cohorts, we mask candidate expression vectors and simulate multi-round selection with uncertainty-based Monte Carlo dropout (MC-dropout) and temporal output discrepancy (TOD), and diversity-based CoreSet and TypiClust-inspired selection. We compare 160 completed configurations at 5%, 10%, 30%, and 50% of the fold-wide training spot pool under patient-level cross-validation, with a separate full-label reference. Within each budget, strategies share the selection schedule, morphology-to-expression predictor, and optimization protocol. We assess mean per-gene within-slide Pearson correlation coefficient (PCC), expression-cluster agreement, and Moran's I fidelity. On HER2-positive breast cancer, pooled mean PCC differences from Random across the four active strategies were -0.0176, -0.0117, +0.0056, and +0.0057 at 5%, 10%, 30%, and 50%, respectively. On cutaneous squamous cell carcinoma (cSCC), three strategies were below Random at 5%, and all four were below Random at 10%. On HER2-positive breast cancer, CoreSet and MC-dropout had lower PCC but higher expression-cluster agreement than Random at the two smallest budgets; this pattern did not reproduce on cSCC. Under the reported fixed training horizons, the evaluated active strategies do not consistently improve on Random at small budgets, and rankings depend on the evaluation measure.
cs.LG / 5 / 2609.27523
M3D-Net: Hierarchical Coordination of Spatial Context, Feature Reuse, and Differential Attention for Mammography Classification
Zheng Yu, Xinhang Li, Jiabao Gao, Boyang Wang, Xiang Li
cs.CV · cs.LG
Abstract
Breast image classification requires local detail and global tissue context, yet these cues can weaken as representations deepen. We present M3D-Net, a mammography encoder that hierarchically coordinates multi-scale coordinate attention, bounded dynamic feature reuse, and differential attention through resolution-aware operator placement. Within-stage retrieval preserves access to earlier features, coordinate-aware aggregation integrates local and global context, and differential attention operates at coarse resolutions. We evaluate image-only classification on AISSLab mammography and an adapted image--clinical model on BrEaST ultrasound. Against EdgeNeXt, RepViT, and TransXNet, the proposed implementations achieve the highest recorded validation accuracy and late-training accuracy, with the lowest endpoint cross-entropy loss. Validation accuracies reach 97.78\% and 80.39\%, respectively. These results support further evaluation of hierarchical coordination across breast imaging settings; repeated-seed, component-controlled, and independent evaluations remain necessary.
cs.LG / 6 / 2609.27710
FFM-CP: Cross-Backbone Fusion of Vision-Language Foundation Models for Few-Shot Computational Pathology
Anh-Tien Nguyen, Trung DQ. Dang, Nghiem Tuong Diep, Bui Ngoc Han Nguyen, Tan-Ha Mai, Miriam Cindy Maurer, Phuong Hoa Nguyen, Thi Thuy Uyen Nguyen, Youngjun Park, Daniel Sonntag, Duy Minh Ho Nguyen, Anne-Christin Hauschild
cs.CV · cs.LG
Abstract
Pathology vision-language foundation models vary in performance across diseases and tasks, with no single model consistently performing best. The high cost of expert pathology annotation can also limit the labeled data available for task-specific adaptation. Combining complementary pretrained representations is a potential approach to these limitations, yet learning an effective fusion from few labeled examples remains challenging. We introduce Few-shot Fusion Foundation Models of Computational Pathology (FFM-CP), which is a framework that combines multiple pathology vision-language models in the few-shot learning setting. The framework first aligns heterogeneous representations using a closed-form Orthogonal Procrustes transformation estimated from corresponding support images. This alignment preserves within-model feature geometry without training an additional alignment network. Within the aligned space, a unified graph enables information exchange across backbones by jointly refining support-image features and visual and textual class prototypes. These refined representations support complementary text-prototype and case-retrieval branches that capture semantic class knowledge and within-class visual variation, respectively. Each branch learns to combine predictions from all ordered backbone pairs, allowing queries encoded by one model to draw on evidence represented by another. We evaluate three backbone combinations on six histopathology datasets at 4, 8, and 16 shots per class. FFM-CP achieves higher mean macro-F1 than the strongest individually adapted member of each fused set in 50 of 54 comparisons. These findings suggest that combining complementary pretrained representations can improve histopathological classification when annotations are limited.
cs.LG / 7 / 2609.27988
Task-Induced Riemannian Metrics for Vision Transformer Feature Spaces
Andrew Bond, Ege Erdem Özlü, Tuna Çimen, Ilkin Umut Melanlioglu, Tolga Birdal, Erkut Erdem, Aykut Erdem
cs.CV · cs.LG
Abstract
Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direction is equally meaningful, but there is no reason to believe the true task geometry has this property. The task-sensitive geometry of the feature space is given by the pullback metric $g(F) = J(F)^\top J(F)$, where $J$ is the Jacobian of the decoder's output fed to a task-specific distance, with respect to the features. Storing the full $g$ is infeasible at modern scales, and for dense outputs such as depth maps even forming $J$ is impractical. We show that whether a low-rank approximation of this metric can be learned depends on the model-decoder pair, and we characterize this with a matrix-free diagnostic $κ_{cap}(r)$ computable with a low number of Jacobian-vector products. For tractable pairs, we develop the Spectral Pullback Network (SPN), which learns a low-rank version of the metric from randomized power iteration, and we distill it into a $310$K-parameter importance head that predicts token importance directly from the features. When the Jacobian spectrum is too spread out for a low-rank approximation, passing the decoder's input features through a VAE bottleneck can restore tractability. Across DPT, DINOv2, CLIP, and VGGT backbones, $κ_{cap}(r)$ predicts which learned-metric architectures are viable. The importance head reaches Spearman $ρ= 0.998$ on DINOv2 CLS, and our geometric token pruning reduces the additional depth error of ToMe-based token selection by $25\%$ on DPT depth at prune ratio $0.5$, without fine-tuning the ViT. Project page: https://cyberiada.github.io/TaskInducedViTs/
cs.LG / 8 / 2609.28099
Visual Tripwires: Anticipating Failure in Deep Vision Systems
Anoushka Harit, Rehan Zuberi, William Prew, Florian Markowetz
cs.CV · cs.LG
Abstract
Deep vision systems remain vulnerable to corruption, occlusion, and distribution shift despite strong benchmark performance. Existing reliability methods typically evaluate uncertainty at individual time steps and do not explicitly model how a system progresses toward failure. We introduce Visual Tripwires, a predictive reliability framework that uses temporal instability in model behaviour to anticipate impending failure. Our central hypothesis is that predictive degradation develops progressively through measurable changes in latent representations, prediction trajectories, and attention structure. Visual Tripwires captures these changes using representation drift, prediction oscillation, trajectory curvature, and attention entropy. A lightweight tripwire predictor aggregates these signals over a temporal window to estimate the probability of failure within a future prediction horizon. Experiments across multiple datasets, architectures, and progressive perturbation settings show that the proposed instability signals emerge before predictive degradation and provide earlier and more accurate failure warnings than conventional uncertainty estimation methods. These results demonstrate that temporal instability contains useful information about future model reliability and provides a practical basis for early warning in deep vision systems.
cs.LG / 9 / 2609.28262
RAMP: Robust Adaptive Mixed-Precision Quantization for Edge CPU Vision Models
David Población-Criado, Dario Garcia-Gasulla, Eduardo Quinones
cs.CV · cs.LG · cs.PF
Abstract
Deploying deep learning models on edge CPUs is bottlenecked by computational and memory constraints. Mixed-precision quantization promises to reduce inference latency while preserving accuracy. However, quantization affects different layer types in inconsistent ways, so identifying where accuracy loss is minimized and latency reduction is maximized is critical, as the effect accumulates over a full deployment into substantial savings or unacceptable task degradation. Such identification relies on sensitivity metrics, proxies that estimate layer-wise degradation without evaluating the task accuracy of every candidate policy. Nevertheless, widely used metrics fail systematically on modern architectures. We present a systematic empirical study of 13 sensitivity metrics for layer-wise INT8 quantization across four distinctly different neural networks, and validate the resulting policies on two ARM64 platforms. Gradient-based sensitivity methods fail on 4 out of 8 model-hardware configurations and weight-based statistics on 2. In contrast, the Jensen-Shannon Divergence achieves zero catastrophic failures, reliably isolating the layers that cannot be safely quantized. A sensitivity metric alone does not define a policy, and the fixed thresholds typically used for that step are fragile over the highly skewed distributions of modern architectures. We address this with K-Means clustering, achieving near-lossless accuracy and a mean speed-up of $1.81\times$ over the full-precision model. Finally, we reveal that excluding from quantization the layers whose speed-up is negligible, regardless of their sensitivity, can be counterproductive, as it induces computational graph fragmentation and disables operator fusion. Our results yield concrete allocation policies for practitioners and researchers deploying quantized vision models on heterogeneous edge CPUs, without GPU access or gradient computation.
cs.LG / 10 / 2609.28286
PBLH Estimation from Satellite Radiances via a Dual-Encoder Transformer
Lorenzo Innocenti, Luca Catalano, Edoardo Arnaudo, Claudio Rossi, Salvatore Larosa, Domenico Cimini, Paolo Garza
cs.CV · cs.LG
Abstract
Estimating the Planetary Boundary Layer Height (PBLH) from satellite observations is a challenging regression problem due to the indirect relationship between top-of-atmosphere radiances and near-surface atmospheric structure. Progress has been limited both by the lack of architectures capable of handling the multimodal, spatially incomplete nature of satellite overpasses, and by the scarcity of suitable datasets. In this paper, we build upon the large-scale dataset pairing MetOp radiances with ERA5 PBLH labels that we introduced in our previous work, making three contributions. First, we establish a benchmark across eight approaches spanning pixel-wise regression, swath-wise sequence models, and convolutional and Transformer models operating on the full orbital passage. Second, we quantify what the resulting model actually relies on, using grouped Shapley decomposition over the input blocks. Third, we present the best-performing architecture found: a dual-encoder Transformer whose masked-input handling lets it operate in all weather conditions. The proposed model achieves MAE = 155.8 m on the held-out global test set, outperforming all baselines on every evaluation subset. On 30 out-of-distribution granules acquired on two days overlapping the TEAMx observational campaign, it achieves MAE = 165.3 m, outperforming a pixel-wise baseline trained on the same data (MAE = 197 m).
cs.LG / 11 / 2609.26890
PR-Smoother: Simulator-Preserving Non-Gaussian Smoothing for Data Assimilation
Yuta Tarumi
cs.LG · nlin.CD · physics.ao-ph
Abstract
Many physical data assimilation (DA) workflows require smoothing methods that represent non-Gaussian posteriors over physical state variables, scale to high-dimensional simulators, train from observation windows alone, and remain compatible with calibration of the prescribed simulator. We introduce PR-Smoother, a simulator-preserving amortized smoother designed for this prescribed-simulator DA regime. Its key design principle is to keep the prescribed simulator explicit in both the evidence lower bound and the variational family: rather than learning replacement dynamics or a learned trajectory prior, PR-Smoother learns only future-conditioned corrections around the prescribed rollout. This yields an explicit non-Gaussian smoothing distribution over physical trajectories and supports joint state, parameter, and sensor-bias learning from observations alone. The variational family contains the exact smoother in deterministic and linear-Gaussian limits. Empirically, PR-Smoother captures multimodal posteriors in 4-dimensional Lorenz-96, remains accurate under ambiguous nonlinear observations and process noise in 40-dimensional Lorenz-96, and scales to joint state-parameter-bias inference in 16,384-dimensional Kolmogorov flow.
cs.LG / 12 / 2609.26905
CORE-STACK+: Meta-Learning for Deep Stacked Generalization
Noor Islam S. Mohammad
cs.LG · cs.CV
Abstract
Stacking heterogeneous vision backbones (CNNs, ViTs, and hybrids) is the de facto recipe for accuracy, calibration, and robustness, yet two coupled pathologies limit its returns. Prediction-space multicollinearity ill-conditions the meta-learner's Gram matrix, inflating weight variance and producing brittle solutions on a thin manifold. Calibration collapse compounds constituent miscalibration through naive linear stacking, so adding more models can hurt expected calibration error (ECE). Existing remedies, ridge regularization, greedy selection, model soups, and SWAG address at most one of these issues, and none jointly target conditioning and calibration in heterogeneous prediction pools. We introduce CORE-STACK+, a preconditioning pipeline with four components: (i) a kernelized redundancy filter that removes non-linear inter-model dependencies invisible to Pearson correlation, using Centered Kernel Alignment (CKA) [23]; (ii) a $<15$K-parameter differentiable meta-feature gate that learns per-sample attention over ensemble statistics; (iii) a spectrum-adaptive Ridge penalty $lambda^{star}=lmax(Chat)/SNR(Chat)$ derived from a Marchenko-Pastur signal-noise decomposition, eliminating nested cross-validation; and (iv) a Laplace-approximate Bayesian blender replacing inverse-RMSE heuristics. We prove a PAC-Bayes excess-risk bound that, for the first time, jointly accounts for prediction-space redundancy and meta-learner capacity. Across six benchmarks, CORE-STACK+ delivers $+1.8\%$ top-1 on ImageNet-1K, $-4.2$ mCE on ImageNet-C, $+0.9$ mIoU on ADE20K, and $+1.3$ AP on COCO, while reducing retained models by 35-57% and inference FLOPs by up to $41%$. ECE improves $2.1\times$ over deep ensembles without post hoc temperature scaling.
cs.LG / 13 / 2609.26918
On Preference Coverage Collapse from Hindsight Relabeling in Multi-Objective Reinforcement Learning
Baptiste Bonin, Caro Strickland, Audrey Durand
cs.LG · cs.AI
Abstract
Hindsight relabeling which retroactively replacing a transition's goal with the outcome the agent actually achieved is an effective tool for improving sample-efficiency in Reinforcement Learning (RL). A natural extension to preference-conditioned multi-objective RL (MORL) relabels transitions with the preference direction the agent achieved rather than the one asked for. We show that this extension is frequently harmful: across four preference-conditioned off-policy algorithms spanning two critic backbones and two preference-sampling schemes on the continuous-control MO-Gymnasium suite, it degrades 19 of 36 algorithm-environment settings by as much as four standard deviations, improves only one, and leaves the rest unaffected. The harm is not a symptom of noisy relabels; denoising the target recovers almost nothing, and neither prioritized sampling nor any buffer-structural choice reproduces it. Instead, repeated relabeling collapses the critic's coverage onto whatever narrow region of the preference space the agent happened to visit. We name this failure mode \emph{Preference Coverage Collapse}, and quantify it with abandoned preference mass (APM), a value-aware statistic that tracks the harm ($ρ= -0.73$) where a purely structural coverage count does not. We then introduce \texttt{her\_mix}, a single-parameter convex combination pulling the achieved direction back towards the requested preference. At one fixed value across every algorithm and environment, it returns 16 of the 19 harmed settings to baseline, preserves and even improves the one setting in which relabeling helps, and cuts abandoned preference mass from $69\%$ to $6\%$. Protecting coverage over the preference simplex, not filtering noisy relabels, is what makes hindsight relabeling safe for MORL.
cs.LG / 14 / 2609.26955
When Post-Processing Fairness Constraints Help and When They Harm: Evidence from Eight Cross-Domain Evaluations
Nithin Raghava Ramachandra Narla
cs.LG · cs.CY
Abstract
Fairness audits in production ML typically occur once, at deployment, on a single domain. Both fail in practice: fairness can shift after retraining or a changing user base, and interventions validated on one dataset are rarely tested across the heterogeneous domains an organization deploys. We present FAPE (Fairness Auditing for Production Environments), a four-stage framework evaluating a single post-processing intervention, Fairlearn's ThresholdOptimizer, across eight domain evaluations: criminal justice, income prediction, legal admissions, credit lending, agricultural lending, a multi-domain benchmark corpus, healthcare, and education. Each is scored on demographic parity and equalized odds difference, plus disparate impact ratio and accuracy cost where computable. Intervention effectiveness tracks baseline disparity magnitude: across model-domain pairs the constraint improved disparity in 9 of 14 high-disparity cases and worsened it in 3 of 4 near-fair ones. Each of the five high-disparity exceptions reverses under one of two measurement checks, a minimum group size or thresholds fit on held-out data. A CUSUM monitor started at deployment, tested on a simulated shift, separates constrained models that never met a 0.1 parity convention from those that met it and later regressed. A single deployment-time audit is therefore an unreliable guide, which argues for baseline-disparity screening and continuous monitoring
cs.LG / 15 / 2609.26959
Transfer Learning with Conformalized Quantile Regression for Solar PV Forecasting Under Load-Shedding-Driven Data Scarcity
Rakib Abdullah, K. M. Tahlil Mahfuz Faruk
cs.LG
Abstract
Solar photovoltaic (PV) forecasting in regions affected by load shedding is challenging because reliable historical observations are scarce. This study proposes a transfer learning framework combined with Conformalized Quantile Regression (CQR) to improve PV power forecasting and provide reliable uncertainty estimates under severe data scarcity. A source-domain PV dataset from Alice Springs, Australia, is used to pretrain a temporal forecasting model, which is then adapted to simulated Bangladesh PV data representing different levels of historical availability. Experimental results show that transfer learning reduces RMSE by up to 23.7% when only one month of target-domain data is available and by 13.7% with three months of data. The proposed Transfer Learning plus CQR framework achieves 94.3% empirical coverage with three months of target data while producing prediction intervals that are 14% narrower than those obtained without transfer learning. These results demonstrate that combining transfer learning with conformal uncertainty quantification can improve both point forecasting accuracy and uncertainty reliability when target-domain PV data are severely limited.
cs.LG / 16 / 2609.26962
CRISP: Scalable Importance-Stratified Coresets for Imbalanced Tabular Learning
Hardhik Mohanty, Indrayana Rustandi, Mohamadreza Sheibani
cs.LG
Abstract
Large imbalanced tabular datasets make repeated gradient-boosted tree training expensive. Existing coreset methods often lose accuracy when most majority examples are removed. We present CRISP (Coreset Reduction via Importance-Stratified Pruning), a linear-time method that allocates a negative-class budget across quantile strata of a proxy-model score. Sample weights account for unequal inclusion probabilities. At 95% negative-class reduction on a production fraud dataset, CRISP trains on approximately 1.70M of 25M rows and retains 99.7% of full-data Average Precision. This is a 93.2% reduction in total training rows. On public CriteoPrivateAds, CRISP has the highest mean Average Precision at each tested rate from 90% to 99.4% majority reduction. Sparkov results are mixed at lower rates, but CRISP has the highest mean at 99.2% and 99.4%. Ablations identify budget allocation and inverse-propensity weighting as the main sources of the production-dataset gain.
cs.LG / 17 / 2609.26972
TinyUDE: Solver-Free Universal Differential Equations on Microcontrollers via Lie-Taylor Jet Matching
Pranavanath Balamurali, Hrishi Kamireddy
cs.LG · cs.ET
Abstract
Training Universal Differential Equations (UDEs) traditionally relies on backpropagating through numerical ODE solvers, creating memory footprints far exceeding the capabilities of edge microcontrollers. We present Lie-Taylor jet matching, a solver-free training framework that fits a hybrid vector field directly to the first and second time-derivatives of observed system states. These derivatives, the truncated Lie-Taylor jet, are estimated online via Savitzky-Golay filtering, yielding fully analytic gradients without automatic differentiation software. We evaluate whether eliminating the solver compromises accuracy against a conventional baseline (fixed-step RK4 integration, multiple shooting, exact discrete adjoints, Adam) sharing identical dynamics, noise models, network architectures, and metrics. While naive derivative matching degrades under sensor noise, our noise-adaptive mechanisms close and reverse this gap: full-rate phase-shifted sampling, a reservoir buffer, cosine-annealed optimization with weight averaging, on-device noise estimation, and polynomial-misfit quality gating. On a damped pendulum and chaotic double pendulum, our method matches or exceeds baseline accuracy at matched data windows and recovers unmodeled damping coefficients. Across noise levels from 0% to 5%, it attains a geometric-mean relative field error of 0.65x that of the baseline within 108 kB of static memory, compared with megabytes of solver tape. On an ESP32 microcontroller, the on-device run reaches a field error of 0.0020 and recovers the damping coefficient to c = 0.400 (true 0.400) within 61.3 kB of static memory and 7.24 ms per update (18.1% duty cycle at 25 Hz), confirming real-time on-device training is feasible without a numerical solver.
cs.LG / 18 / 2609.26979
Resource-Efficient Distributed Recursive Gaussian Processes
Josephine King, Ali Emre Balci, Raj Thilak Rajan
cs.LG · eess.SP
Abstract
Gaussian processes (GPs) provide a flexible framework for learning unknown functions from noisy measurements while quantifying predictive uncertainty, making them well suited for estimation in multi-agent systems. However, when measurements are collected by multiple agents, maintaining a unified GP model without centralized processing requires efficient distributed algorithms that can operate using local measurements and communication with neighboring agents. In this work, we develop two distributed recursive GP (RGP) algorithms for multi-output GP regression: ADMM-RGP and PDMM-RGP. We analyze the stability and convergence of both algorithms and develop parameter selection strategies to accelerate convergence, thus reducing the communication burden. The proposed methods are validated on a real-world multi-output wind dataset, and their convergence behavior is examined across communication graphs with varying connectivity. Numerical experiments demonstrate that ADMM-RGP and PDMM-RGP can significantly reduce communication relative to the state of the art, while maintaining comparable estimation accuracy and network-wide consensus.
cs.LG / 19 / 2609.27018
GeoRVQ: Decoder-aware geometry for residual-token prediction in physiological signals
Bo Cui, Yaowen Zhang
cs.LG
Abstract
Residual vector quantization (RVQ) turns physiological waveforms into compact token sequences, but conventional masked modeling treats every incorrect token as equally costly. We propose GeoRVQ, a coarse-to-fine masked token model whose objective reflects the local response of a frozen waveform decoder. Decoder-induced costs define geometry-aware soft targets and expected distortion, while quantizer-causal prediction follows residual dependencies from coarse to fine levels. In a descriptive aggregate over MIMIC-IV Waveform, VitalDB, and CODE-15\%, GeoRVQ increases exact token accuracy from $.133\pm.004$ to $.143\pm.003$, reduces decoded distance from $.606\pm.006$ to $.393\pm.007$, and increases R-peak F1 from $.784\pm.004$ to $.837\pm.008$ under matched model and training conditions. Across 45 held-out code substitutions, decoder-induced cost has a Spearman correlation of $.85$ with realized decoded cost, compared with $.54$ for Euclidean codeword distance. These results indicate that decoder-aware objectives can improve waveform and event preservation without requiring a large increase in exact token accuracy.
cs.LG / 20 / 2609.27033
WTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow Maps
Abbas Mammadov, Jerry Y. Huang, Justin Lin, Partha Kaushik, Sheel Shah, Kartik Nair, Yee Whye Teh, Nicholas M. Boffi
cs.LG · cs.CV · stat.ML
Abstract
Reward fine-tuning aims to update a pre-trained flow-based generative model to improve the downstream reward of its generated samples. Existing methods typically formulate this problem as sampling from a reward-tilted distribution, the solution to a KL-regularized reward-maximization problem. Here, we introduce an optimal transport regularizer built directly from the pre-trained drift. Unlike KL reward tilting, the resulting objective transports individual samples toward higher reward rather than reweighting the base distribution. We show that the resulting problem is equivalent to a deterministic optimal control problem on the flow. Given a pre-trained flow map, this equivalence yields a simulation-free reinforcement learning algorithm for fine-tuning generative flows. We call the resulting framework Wasserstein-Tilted Flow Maps (WTF), the first end-to-end fine-tuning recipe native to flow maps. The output is a fine-tuned flow map that retains strong reward-aligned performance at few-step inference budgets without post-hoc distillation. Experiments on ImageNet-256 and text-to-image show that WTF achieves higher reward with comparable or higher diversity than baselines, while requiring up to $280\times$ less training compute. More broadly, we argue that accelerated samplers such as flow maps are essential infrastructure for efficient post-training, and that the dominant KL-regularized formulation is only one of many choices worth revisiting.
cs.LG / 21 / 2609.27036
An open benchmark for machine learning-based polymer property prediction
Robert W. Learsch, Nicholas Liesen, Daniel S. Levine, Anna M. Hiszpanski, Evan R. Antoniuk
cs.LG · cond-mat.mtrl-sci · cs.AI
Abstract
Polymer property prediction lacks open, standardized benchmarks that enable rigorous comparison of machine-learning methods, with existing resources covering only a narrow fraction of polymer architectures, such as homopolymers. We introduce Polymer Benchmark 2026 (PolyBench26), an open dataset comprising nearly 250,000 polymer-property datapoints across eight physical properties, including data from experimental measurements, density functional theory, and molecular dynamics. The benchmark supports four evaluation tasks across homopolymers and alternating, random, and block copolymers: in-distribution property prediction, dataset-size scaling, repeat-unit complexity, and transfer to held-out polymer architectures. We compare language model, graph-based, and descriptor-based approaches and find graph-based models provide the lowest errors in property prediction, retain their advantage across the evaluated training-set sizes, and remain robust to increasing repeat-unit complexity. PolyBench26 provides a reproducible foundation for developing models for the increasingly complex polymer design space. The PolyBench26 benchmark is available open-source at https://github.com/rlearsch/PolymerBenchmark2026.
cs.LG / 22 / 2609.27092
Local Evidence and Geometric Readout Repair in Trained GNNs
Nadi Tomeh, Hugo Attali
cs.LG · cs.AI
Abstract
Many node-classification GNNs apply a linear classifier to a nonnegative mixture of local messages. An error can reflect either poor mixture weights or a reachable logit set poorly positioned for the classifier. We separate these causes with an exact-mass linear program and two learned post-hoc repairs. Every reweighted prediction has an equivalent centered logit translation, but only translations in a message-induced displacement set are realizable by reweighting. Across eight datasets, eight GNN backbones, and ten splits, mean accuracy rises from 62.6% for the frozen models to 63.8% with reweighting and 65.3% with set-conditioned translation. A parameter-matched node-only translator reaches 64.6%, showing that translation explains most of the gain while the message set supplies a smaller additional benefit. Although oracle reweighting can correct many errors, label-free reweighting captures little of this potential: local evidence is often present but hard to select, and relaxing the evidence constraint is more effective than learning within it.
cs.LG / 23 / 2609.27144
Learning Risk Scores Robust to Unobserved Confounders
Ryan Edmonds, Yingxiao Ye, Sina Aghaei, Andrés Gómez, Çağıl Koçyiğit, Phebe Vayanos
cs.LG
Abstract
We consider the problem of learning risk scores to prioritize individuals for scarce resources or interventions, from historical observational data affected by unobserved confounding. Decisions about who receives scarce resources are often guided by risk scores based on recorded characteristics, such as responses to a survey. These risk scores are increasingly being learned directly from observational data: historical records of individuals' characteristics, allocation decisions, and outcomes. Standard methods such as inverse propensity weighting (IPW), which corrects for the bias introduced by the historical allocation policy, can be used to learn accurate risk scores if the historical decision process is fully explained by the recorded characteristics. In practice, however, historical decisions often depend on unrecorded information, causing learned risk scores to systematically under-prioritize exactly the individuals whose unrecorded circumstances drove past prioritization. We propose a method for learning risk scores that are robust to this kind of unobserved confounding, building on IPW. Since propensity weights cannot be reliably estimated under unobserved confounding, we instead treat them as belonging to an uncertainty set determined by the observable data and domain-informed estimates of the degree of confounding, combining sensitivity analysis from causal inference with Wasserstein distributionally robust optimization. The resulting robust risk score learning problem admits a sample-based approximation that we reformulate as an exponential cone program compatible with off-the-shelf solvers. We demonstrate the effectiveness of our approach on semi-synthetic data derived from datasets in the UCI Machine Learning Repository. Our method improves calibration by up to 29.2% over traditional benchmarks and up to 11.1% over the state of the art, without compromising other metrics.
cs.LG / 24 / 2609.27158
The Linear Representation Hypothesis Needs a Group Action
Louie Hong Yao, Yuhao Li, Shengchao Liu
cs.LG · cs.AI · cs.CL
Abstract
To make claims about representations that generalize beyond a particular trained model, we need to specify when two representations should count as equivalent. The Linear Representation Hypothesis is often discussed without making this equivalence explicit. Different notions of equivalence preserve different structures, so metrics, probes, and interventions that appear to study the same representation may in fact correspond to different hypotheses. We therefore argue that the Linear Representation Hypothesis is not one hypothesis but a family of claims distinguished by representation equivalence. We formalize this idea using group actions, specifying the representation object, the procedure that produces it, and the property ultimately asserted, while accounting for equivalences imposed by the model architecture. This framework clarifies how assumptions can change across metrics, reading points, and analysis stages, and we use it to audit common representation quantities and recent interpretability analyses.
cs.LG / 25 / 2609.27186
Data-driven discrete-time deep recurrent neural network-based modeling for dissipative systems
Tuan Luong, Hyungpil Moon
cs.LG · cs.RO
Abstract
Physical AI has gained increasing attention for its role in developing AI systems that better understand, predict, and control real-world dynamics. Achieving this requires AI models that not only achieve high prediction accuracy but also preserve fundamental physical properties of dynamical systems. In this paper, we propose a deep discrete-time dissipative recurrent neural network (DissipNet) that explicitly enforces dissipativity, a key property related to stability and energy dissipation, through structural weight constraints and a dedicated training algorithm. By construction, the proposed network is capable of learning dissipative dynamics while preserving their inherent stability, which is formally analyzed using Lyapunov theory. In contrast to Physics-Informed Neural Networks (PINNs), which incorporate governing equations into the training loss but do not guarantee preservation of internal analytical properties such as dissipativity or passivity, our approach provides explicit guarantees on stability at the model level. We demonstrate the effectiveness of the proposed method through several modeling applications, and compare its performance with a naive recurrent neural network (RNN) and a PINN-based model.
cs.LG / 26 / 2609.27199
ZO-COSMO: Index-Free One-Hop Mixing for Decentralized Zeroth-Order Optimization
Shengjun Zhang, Tingyi Liu, Heng Zhang, Dong Xie
cs.LG · eess.SY · math.OC
Abstract
Sparse communication in decentralized zeroth-order learning requires compatible peer-state coordinates. We characterize this one-hop condition and develop \textsf{ZO-COSMO}, coupling two-query estimation with average-preserving masked consensus using $q$ values per active link. Global supports serve all-neighbor mixing; matching updates require agreement only within each pair. We derive a sharp contraction-per-scalar bound within the matching class and convergence guarantees for the core and sparse-momentum updates. At fixed matching, exact moment identities characterize how shared directions preserve gradient-heterogeneity cancellation and redistribute estimation error and disagreement. Mechanism experiments cover unequal curvatures, noise, and sparse momentum. Further tests span $64$ synthetic agents and eight logical Qwen LoRA workers. At matched payload budgets, Qwen2-7B QNLI gains $3.65$ accuracy points over explicit-index Rand-$k$; edge-local updates gain $3.42$ and $2.53$ points over all-neighbor mixing on eight-worker complete and ring graphs. A matched-first-step ablation gives a $3.92$-point momentum benefit. Seed-aware and same-matching controls distinguish encoding, scheduling, and query correlation.
cs.LG / 27 / 2609.27201
A Systematic Benchmark of Explainable Methods for Temporal Attribution in Sequential Recommendation Systems
Akash Pandey, Kanisha Shah, Addrish Roy, Dwipam Katariya, Hongyangyang Shi, Amanda Ding, Kalanand Mishra, Pranab Mohanty
cs.LG
Abstract
Sequential RecSys are central to modern personalization, exploiting user's historical interaction sequences to drive next-step decisions. Deep learning models, particularly CNN and Transformer-based architectures, have proven highly effective at capturing temporal dependencies in these histories. For transparency and trust, understanding which past interactions drive a given recommendation is increasingly important --- both for developers auditing model behavior and for users seeking a rationale. However, the non-linearities that give these models their predictive power also render them black boxes, making it difficult to attribute decisions to specific interactions. While gradient-based, perturbation-based, and attention-based explainability methods exist, a systematic benchmark of their faithfulness for sequential recommendation is missing. We address this gap by introducing a dual-model masking metric in which one model supplies per-timestep attribution scores and a separately trained, masking-robust probe measures the resulting change in predicted probability. Using this metric, we benchmark ten XAI methods across CNN, Transformer, SASRec, and BERT4Rec backbones on KuaiRand and MovieLens, complemented by analyses of temporal attribution patterns, item popularity confounding, and robustness to input corruption. Our key findings are: (1) gradient-based methods, particularly GradientSHAP and Integrated Gradients, yield the most faithful and robust attributions; (2) raw attention weights are unreliable, but gradient-weighted attention restores faithfulness on shorter sequences, with degradation on longer horizons as softmax attention probabilities converge toward uniform importance scores, diminishing the method's ability to identify informative interactions; and (3) temporal attribution patterns in faithful methods reflect genuine task structure rather than recency or popularity bias.
cs.LG / 28 / 2609.27209
Scalable Subgraph Sampling via Resistance Curvature
Chaoqun Fei, Tinglve Zhou, Tianyong Hao, Yangyang Li
cs.LG · cs.AI
Abstract
Subgraph sampling reduces the training cost of large-scale graph neural networks, but sampling criteria may overlook the geometric roles of edges. We propose a resistance-curvature-guided sampling framework built on ERC-LG, a curvature approximation method for large-scale graphs. ERC-LG combines Johnson-Lindenstrauss projections with regularized multi-GPU batched conjugate gradient solvers, avoiding explicit Laplacian pseudoinverse computation and full embedding storage. The resulting curvature informs node- and edge-sampling probabilities for constructing GNN training subgraphs. Experiments show numerical agreement with pseudoinverse-based curvature and reduced runtime compared with CG-only computation. ERC-LG-based sampling variants achieve the highest mean accuracy on six of seven real-world datasets in downstream node classification.
cs.LG / 29 / 2609.27221
Tail-Aware Geometry Learning for Conformal Ellipsoids
Xiang Zhang
cs.LG
Abstract
This paper studies multivariate conformal prediction (CP), a distribution-free uncertainty quantification framework with finite-sample coverage guarantees. The efficiency of multivariate prediction sets hinges critically on the residual geometry encoded by the nonconformity score, while existing minimum-volume methods rely on quantile thresholds that ignore tail residual severity and implicitly bind geometry learning to coverage level. We propose a tail-aware geometry learning framework for conformal ellipsoids that decouples tail sensitivity in geometry learning from the final coverage guarantee. Using a two-split design, we learn the metric matrix via volume minimization under a CVaR constraint on an estimation split, then apply standard conformal calibration on a held-out calibration split. The resulting problem is convex and admits a bounded-reweighting interpretation that prioritizes high-residual samples. Moreover, we theoretically characterize the trade-off between ellipsoidal volume and tail severity. Experimental results demonstrate the effectiveness of the proposed method.
cs.LG / 30 / 2609.27232
A Scaling Study for fMRI Foundation Models
Wenhao Ye, Xuanye Pan, Junfeng Xia, Junxiang Zhang, Mo Wang, Quanying Liu
cs.LG
Abstract
Scaling laws have guided large-model development in computer vision and natural language processing, but the relationships among data, model size, and compute remain unclear for functional magnetic resonance imaging (fMRI) foundation models. Here, we conduct a controlled empirical study using pretraining data from more than 200 source datasets and over 10,000 GPU-hours of experiments. Holding the pretraining framework and downstream protocol fixed, we vary pretraining data size, model size, and training duration. Downstream performance generally improves with compute, yet models using similar compute can perform substantially differently. Additional pretraining data bring larger gains at larger model sizes, suggesting that data and model size should be scaled together. At matched compute, increasing pretraining data benefits more tasks than increasing model size, although the pattern varies across tasks. We then use in-distribution (ID) downstream performance to select the combination of pretraining data size, model size, and training duration at two fixed compute budgets. The resulting models are locked before out-of-distribution (OOD) evaluation. They achieve the highest average performance across the evaluated OOD tasks among the compared fMRI foundation models while using less pretraining compute. Overall, our results show that compute alone does not characterize fMRI scaling: performance depends on how pretraining data, model size, and training duration are combined.
cs.LG / 31 / 2609.27234
Discover, Falsify, Revise: Auditing Input-Use Claims from Source Code to Predictive Contribution in Agent-Discovered Cell Models
Mengran Li, Bo Li, Chengyang Zhang, Yang Yan, Jinfeng Xu, Zhenchao Tang
cs.LG · q-bio.QM
Abstract
AI virtual cells aim to predict cellular responses to specified interventions, yet held-out predictive performance alone does not establish use of the supplied perturbation information. This prediction-claim gap matters in agentic model discovery, where language-model agents generate and revise predictors using score-based feedback. We introduce CELLAUDIT, which audits input-use claims by asking whether an input can enter the cited computation, whether fitted predictions depend on it, and whether that dependence improves prediction of observed response. On a paired morphology-transcriptomics perturbation benchmark (BBBC047), an agent-selected predictor attains a mean held-out Global Pearson correlation coefficient (PCC) of 0.3153 but remains invariant to compound replacement; a control-profile-only predictor reaches 0.3142. Source inspection identifies a compound-query pathway blocked by singleton key-value attention, and the invariance persists after refitting with disjoint control wells. In a stratified audit of 48 candidates across two linked tasks, 47 change predictions under compound replacement on both held-out folds, but only 20 show target-loss gains with intervals above zero on both folds. On BBBC047, falsification-guided revisions recover positive mean compound contributions while retaining gains over the control-profile-only baseline. In matched sci-Plex searches, audit-enriched feedback yields higher held-out performance and larger mean compound and dose contributions across five trajectories, although paired intervals span zero. Refitting fixed designs on an independently acquired cohort shows predictive generalization need not imply generalization of input-use claims: dose contribution persists, whereas support for compound identity does not. CELLAUDIT adds a falsification layer to agentic model discovery, moving from generate-score-revise toward discover-falsify-revise.
cs.LG / 32 / 2609.27244
Full-Covariance Smoothing of Bayesian Neural Networks for Online Adaptation
Oren Wright, Haoming Jing, Qiaoan Shen, Koichiro Niinuma, Yorie Nakahira, José M. F. Moura
cs.LG · eess.SY
Abstract
A neural network's layers can be treated as time steps of a state-space model, turning Bayesian training into a smoothing problem: a forward pass propagates Gaussian moments through the network, and a backward Rauch--Tung--Striebel pass updates the weight posteriors in closed form. Such methods learn from each observation in a single pass, in an uncertainty-aware manner, and without gradient-based iterations or replay, which makes them well suited for online adaptation and data-efficient learning. Existing smoothing-based methods, however, are restricted to diagonal covariances across activations, discarding correlations between neurons. We overcome this limitation via a cross-covariance identity that enables full-covariance propagation through a network's nonlinear activations. We derive a one-step-per-layer smoother that approximates as Gaussian only each layer's affine output, and that applies both to deterministic systems with noisy observations and to stochastic systems described by output statistics. We demonstrate this method in non-stationary classification, online dynamics learning, and policy adaptation of a vision-language-action model, and find that it is generally more accurate than other smoothing-based methods.
cs.LG / 33 / 2609.27252
What Converges in the Platonic Representation Hypothesis? Structure over Geometry
Junwon You, Mihyun Jang, Sangwoo Mo, Jae-Hun Jung
cs.LG · cs.CV · math.AT
Abstract
The Platonic Representation Hypothesis suggests that increasingly capable models converge toward shared representations. Recent work narrows this claim to shared local neighborhood relationships, finding that capacity-dependent trends in several global similarity measures largely disappear after calibration. We challenge this interpretation by showing that prior local-global comparisons confound structural scale (local versus global) with what is compared: relational structure, defined by which samples are related, versus metric geometry, characterized by quantitative relations such as distances, similarities, or correlations. To disentangle these factors, we construct a controlled $2\times2$ framework that evaluates both relational structure and metric geometry at local and global scales. We introduce $H_0$ skeleton overlap as a global counterpart to mutual $k$-nearest neighbors, together with matched distance-aware variants. Across vision-language models, relational structure exhibits robust representational convergence at both scales after calibration, whereas increasingly stringent distance agreement substantially weakens alignment and progressively flattens the capacity-dependent trend. We further extend the analysis beyond ambient Euclidean geometry by evaluating distance agreement under a Riemannian metric approximation and recover the same structure-geometry pattern. The pattern is also reproduced in video-text representations. Together, these results show that relational convergence extends beyond local neighborhoods to global spanning structure, whereas metric geometry exhibits substantially weaker convergence.
cs.LG / 34 / 2609.27278
Graph Learning with Spectral Connectivity Priors for Scarce Data
Mingxiao Liu, Bahar Oveisgharan, Bingyan Zou, Gene Cheung, H. Vicky Zhao, Feifei Gao
cs.LG · eess.SP
Abstract
Learning a sparse graph from scarce data is practically important but challenging. Motivated by the desirable combination of local sparsity and strong global connectivity exhibited by expander-like graphs, we propose spectral connectivity-regularized graph learning (SCoGL), a framework that incorporates a family of Laplacian spectral priors to explicitly promote global connectivity. Specifically, SCoGL augments a combinatorial-Laplacian-constrained graphical lasso (GLASSO) objective over a target adjacency matrix $\mathbf{W}$ with a general connectivity prior computed from Laplacian eigenvalues. We derive gradients for several representative connectivity priors and develop a projected gradient descent (PGD) algorithm with Armijo backtracking to efficiently optimize $\mathbf{W}$. Experiments show that the proposed SCoGL variants improve graph recovery and enhance downstream tasks such as graph signal denoising when signal observations are scarce.
cs.LG / 35 / 2609.27287
SR-Fraud: An Outcome-Supervised Reflective LLM Agent Framework for Non-Stationary Payment Fraud Detection
Xuwei Tan, Yao Ma, Xueru Zhang
cs.LG · cs.CR
Abstract
Real-time payment fraud detection is a non-stationary streaming prediction problem: adversaries adapt before supervised labels mature, and localized burst attacks can cause losses before retraining. Production systems typically rely on tabular classifiers and rules, which can struggle to capture these emerging sequential patterns before periodic retraining occurs. We present SR-Fraud, an outcome-supervised reflective LLM framework that decouples request-time decisions from offline adaptation. A frozen, stateless agent scores each transaction from a Hybrid Episodic Window to track behavioral shifts, while an offline reflection agent proposes boundary hypotheses from matured errors. A deterministic verifier then admits only supported hypotheses into an executable knowledge state. On a production payment-fraud benchmark, SR-Fraud improves all detection metrics over its frozen decision agent, obtains higher point estimates than static and periodically retrained CatBoost, and detects an emerging fraud burst.
cs.LG / 36 / 2609.27291
NGN: Learning Neural Network Size as a Differentiable Count
Lixing Li
cs.LG
Abstract
Neural network size is usually chosen before training, separating architecture selection from weight optimization. We introduce the Neurogenesis Network (NGN), a differentiable parameterization for learning how many ordered structural components a model should use. For each ordered component group, one learnable boundary selects an active prefix while the model parameters are trained. The boundary can grow from a compact initialization and can be deployed by discarding components beyond the learned boundary. Controlled experiments examine convergence of the learned boundary, the performance of deployed prefixes, and comparisons with fixed-size models and alternative approaches to learning capacity. We then apply the same mechanism to MLPs, convolutional and graph networks, Transformers, state-space models, LoRA, and adapters. Across these settings, deploying only the learned prefix usually changes performance little, and the selected architectures perform similarly to fixed models trained at the same size. These results show that structural capacity can be optimized directly as a count.
cs.LG / 37 / 2609.27294
KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling
Zhiheng Hu, Yixun Wei, Jian Zhou, Yizhuang Zhou, Ji Li, Xing Chen, Yang Li, Bojun Wang, Yibo Zhu, Xiangyu Zhang, Daxin Jiang
cs.LG · cs.AI
Abstract
Scaling a language model is not only a question of final quality: the architectural choice determines how much computation is spent during training, prompt processing, and autoregressive decoding to achieve certain model quality. An ideal model architecture should lower all above computation costs to facilitate scaling to a larger model, while ensure the larger model indeed outperforms smaller baselines. We introduce KV-Invariant Transformer Expansion (KITE), a scaling paradigm that achieves this goal. It trains the model from a smaller size to a larger size (i.e., saving training costs via upcycling), while places newly added parameters in regions that do not affect attention KV. Consequently, during inference, prefilling KV only relies on the smaller part of the model, so the inference costs are saved. As a concrete instantiation, we present Step Scale Transformer (SST), a two-tower decoder in which one tower produces KV and the other reads them. At comparable cumulative training compute, SST, a 67B MoE model with 2.15B active body parameters per decode token, achieves lower training loss than 47B and 63B MoE Transformers with 1.48B and 2.02B active body parameters, respectively, while reducing estimated inference cost by 6.7% and 31.6%.
cs.LG / 38 / 2609.27303
Live Assistant: Learning Whether, When, and Whom to Assist in Real-World Live Social Streams
Shujian Gao, Jiamei Yan, Yuchen Yang, Penghao Zhou, Qinglei Wang, Tiehan Fan, Yuan Wang, Zuxuan Wu, Yu-gang Jiang
cs.LG
Abstract
Livestreams are long-lasting interactive environments where audiovisual content, viewer activity, host behavior, and platform signals evolve together, creating assistance needs that emerge from the stream itself. We introduce \liveassistant, a framework for mixed-initiative, role-conditioned assistance that formulates livestream interaction as four coupled decisions: \textbf{whether to act, when to act, whom to address, and what to communicate}. At each 10-second interval, one autoregressive policy consumes native audio and video with synchronized comments, gifts, viewer dynamics, and room metadata, then selects \textsc{OBS}, \textsc{MEM}, or \textsc{ANS}. \textsc{OBS} remains silent, \textsc{MEM} records a private semantic update, and \textsc{ANS} specifies a recipient, task, and grounded message. To support this task, we build a trajectory engine that reconstructs real livestream sessions into structured causal supervision, yielding over 320 hours of optimization trajectories and a human-reviewed benchmark of 275 clips and 13,812 decision intervals. We train the policy with Marker-Aware Multiturn Supervised Fine-Tuning (MA-MSFT), which strengthens sparse structured decisions, followed by Streaming Multiturn GSPO (SM-GSPO), which optimizes self-generated trajectories with turn- and trajectory-level credit. On the held-out benchmark, \liveassistant reaches 71.14 state accuracy, 72.67 recipient accuracy, and 58.41 task accuracy, with consistent gains over representative streaming and general multimodal baselines. Together, the formulation, benchmark, and training framework establish livestream assistance as selective participation in a shared social stream.
cs.LG / 39 / 2609.27355
Quantization-Robust Unlearning through the Lens of Retain-Forget Loss Landscapes Interaction
Jialu Wang, Jianing Deng, Shuqing Luo, Yuanzhe Li, Dongwei Wang, Jingtong Hu, Huanrui Yang, Song Wang, Tianlong Chen
cs.LG · cs.AI
Abstract
Unlearning ensures LLM compliance by removing the influence of private or copyrighted training data. However, since LLM models typically undergo post-training compression, like quantization, in practical deployment, it has been observed that the unlearning effect can be substantially weakened, with the forgetting behavior degrading more severely than that of model utility. This paper proposes a quantization-robust unlearning framework that makes forgetting robust to quantization while maintaining overall model utility. We analyze this gap through the lens of loss landscape. Specifically, our analysis reveals a curvature-based criteria that pinpoints sensitive weights in the unlearned model that leads to both non-robust forgetting and reduced utility. We therefore propose sensitivity-guided noisy regularization, which is applied on the sensitive parameters to steer the model convergence towards a smoother minima of uniformly low forget and retain losses. Balancing unlearning and utility, we further propose forget-critical optimization, which updates only forget-critical layers, preserving most of the network to retain useful knowledge. Extensive experiments on the MUSE and TOFU benchmarks across multiple LLM unlearning algorithms show that our approach achieves substantially more quantization-resilient forgetting while maintaining utility.
cs.LG / 40 / 2609.27362
Anomaly-Free Self-Optimization via AUC Bounds
Kevin Wilkinghoff, Zheng-Hua Tan
cs.LG · eess.AS
Abstract
Anomalies are rare, and anomalous data are often unavailable during development, making it difficult to determine which anomaly detection models and configurations will generalize to unseen anomalies. Recent approaches address this challenge by generating pseudo-anomalies and using bounds on the achievable area under the ROC curve (AUC) to select the optimal configuration from a finite set of candidates. Instead, we use the AUC bound as a differentiable, anomaly-free objective for directly optimizing continuous parameters of anomaly detection systems. We demonstrate this framework by optimizing ensemble weights and introducing a learnable score-rescaling mechanism that adapts pseudo-anomaly scores, enabling optimization beyond a predefined candidate set. Experiments across multiple datasets and embedding models show that AUC-bound optimization achieves significant performance gains over conventional model selection and prior development-set-based parameter selection. The results further show that direct optimization is less sensitive to the choice of pseudo-anomaly construction.
cs.LG / 41 / 2609.27385
Forecast Workflow Bench: Evaluating Language-Model Decisions with Budgeted Forecast Tools
Shunya Nagashima
cs.LG · cs.AI
Abstract
Time-series foundation models (TSFMs) provide forecasts for operational decisions, but accuracy alone does not determine their value. Evaluating agents that use these models requires measuring decision quality and forecast cost. FWBench evaluates this capability on 1,251 electricity and cycle-hire cases using fixed forecast tools and simulated capacity contracts. Agents select models, histories and horizons, then submit capacities to minimize a stated loss-cost objective. We evaluated two hosted and eight local configurations, including small language models, and tested local models with and without TSFMs. GPT-6 Astra bought inexpensive short-horizon forecasts selectively, using 2.5% of the budget, and outperformed fixed policies when the saved decisions were scored with three loss-cost weightings. FWBench enables reproducible evaluation of how language models select and use time-series forecasts to make decisions under cost constraints.
cs.LG / 42 / 2609.27409
Active Learning for Biodiversity Monitoring: From Label Efficiency to Reliable Ecological Inference
Ben McEwen, Shiqi Zhang, Dan Stowell
cs.LG
Abstract
Limited expert annotation capacity is a pervasive constraint in biodiversity monitoring. Passive acoustic recorders and camera traps generate data faster than experts can analyse them. Machine learning (ML) models can process these data at scale, but their reliability depends on the quality, quantity, and coverage of labelled samples, so expert time remains a constraint. Active learning (AL) eases this bottleneck by selecting, under a fixed annotation budget, the samples expected to improve a model most, and published evidence shows it can reduce the labels needed to reach a target performance. Monitoring programmes, however, face a broader question: how should a limited expert budget be divided so that model training, validation, and the ecological estimates built on model outputs all remain reliable? Because AL selects samples non-randomly, its labels are unsuitable for validation, calibration, or threshold selection, a tension rarely acknowledged. We synthesise AL research across acoustic and image modalities and identify gaps and opportunities. Most studies evaluate query strategies on pre-labelled benchmarks with simulated annotators; deployments in real monitoring workflows are rare and concentrate on birds and cetaceans. Bats, insects, amphibians, and fish are underrepresented, and multimodal applications remain largely unexplored. Evaluation centres on headline reductions in annotation effort, often without random-sampling baselines, per-class results, or calibration analysis, and rarely accounts for the labels required for validation. We provide a tutorial treatment of the AL loop that makes these budget decisions explicit, and a roadmap towards AL methods that support label-efficient training, validation, and trustworthy downstream ecological inference.
cs.LG / 43 / 2609.27411
When Labels Are Scarce: An Oscillatory State Space Model for Vibration Diagnosis
Mainak Mallick, Seung-Kyum Choi
cs.LG · eess.SP
Abstract
Machine fault diagnosis from vibration requires learning from scarce labelled fault recordings while meeting the computational constraints of edge devices for local inference. We introduce DualRes, a compact oscillatory state-space model that combines two complementary spectral views of vibration, capturing rapid changes and fine frequency structure. Time-aligned views are processed by selective oscillatory memory, which learns how long to retain temporal patterns. The encoder contains 39,528 parameters. We evaluate supervised learning across six bearing datasets and a gearbox benchmark, with an additional gearbox pilot. Recording-level splits and explicit accounting of labelled duration distinguish data efficiency from repeated exposure to correlated samples. On the main gearbox benchmark, DualRes achieves state-of-the-art performance among the nine evaluated methods at six of seven label budgets. With about six labelled seconds per class, it improves macro-F1 by 16.1 percentage points over the next strongest comparator. On the same benchmark, DualRes achieves a 1.44-fold recording-level speedup and a 24.8-fold reduction in checkpoint storage relative to a selective state-space baseline under matched hardware and runtime conditions. Bearing results reveal task-dependent trade-offs. These findings support oscillatory memory as a compact approach to vibration diagnosis under limited labelled exposure.
cs.LG / 44 / 2609.27421
Counterfactual Constraint-Conditioned On-Policy Distillation for Multi-Constraint Instruction Following
Yanzhao Zheng, Yuanqiang Yu, Tianze Xu, Chao Ma, Zhentao Zhang, Jihuai Zhu, Baohua Dong, Hangcheng Zhu, Ruohui Huang
cs.LG · stat.ML
Abstract
Multi-constraint instruction following requires a model to respond to a query under many simultaneously active constraints. Even strong instruction-tuned models still routinely violate some of them. Existing approaches either augment supervision with sequence- or token-level RL rewards from external verifiers or learned graders, or use on-policy distillation (OPD) against a single full-context teacher whose probability mass becomes diluted as more constraints become simultaneously active. We propose CC-OPD (Counterfactual Constraint-Conditioned On-Policy Distillation), which inverts the standard supervision-generation direction in distillation. Rather than enriching the teacher with information beyond what the student sees, CC-OPD ablates each constraint from the teacher's conditioning in turn, and constructs the per-constraint signal from the resulting per-token probability differentials. The resulting per-token leave-one-out log-likelihood shifts are summed, clipped, and added to the vanilla OPD reward as a token-level shaping term. All shaping terms are obtained from the frozen teacher, without an external verifier during distillation, and the reward equals vanilla OPD wherever the aggregate shift is zero. Across two Qwen model pairs and seven benchmarks, CC-OPD achieves the highest average among all evaluated student-training methods. A 1.5B student trained with CC-OPD surpasses its own 7B RL-trained teacher on the MulDimIF benchmark.
cs.LG / 45 / 2609.27441
Stable Neural Decoding Across Sessions via Task-Conditioned Latent Alignment for Brain-Machine Interfaces
Canyang Zhao, Bolin Peng, J. Patrick Mayo, Ce Ju, Bing Liu
cs.LG · eess.SP
Abstract
Achieving stable long-term neural decoding in invasive brain-machine interfaces (BMIs) remains challenging due to variations in recorded neural populations across sessions. Current latent alignment approaches may overlook task-dependent structure during cross-session adaptation. We propose Task-Conditioned Latent Alignment (TCLA), a framework that stabilizes neural decoding by learning a shared latent space. TCLA learns a low-dimensional source representation using neural reconstruction and continuous behavioral supervision. During target-session adaptation, the shared representation is fixed, while target neural activity is mapped into the source latent space by aligning source and target distributions separately for each task condition. We evaluated TCLA on seven nonhuman primate datasets spanning multiple tasks. In long-term cross-session evaluation, TCLA achieved a mean $R^2$ of $0.476\pm0.014$ with a negative $R^2$ failure rate of only 6.8\%. Across 1,356 within-subject session pairs, TCLA achieved a mean $R^2$ of $0.371\pm0.009$ with a failure rate of 6.8\%. Across 2,134 cross-subject session pairs, TCLA achieved a mean $R^2$ of $0.218\pm0.004$ with a failure rate of 12.9\%, substantially better than those of the comparison methods. These results demonstrate that by preserving behaviorally relevant and task-dependent latent structure, TCLA improves the robustness of neural decoding across recording sessions and subjects. The source code is publicly available at \href{https://github.com/FAMD-CASIA/TCLA}{https://github.com/FAMD-CASIA/TCLA}.
cs.LG / 46 / 2609.27446
Quantum Reinforcement Learning for Cost and Delay Tradeoffs in Quantum Cloud Orchestration
An N. H. Phan, Dang Van Huynh, Muhammad Usman, Hoa T. Nguyen
cs.LG · cs.AI · cs.DC · cs.ET · quant-ph
Abstract
Quantum cloud computing, delivered through the quantum-as-a-service (QaaS) model, provides access to quantum computing resources. However, applying uniform time-based pricing across fundamentally heterogeneous quantum resources significantly complicates task orchestration, particularly when addressing the tradeoff between execution costs and system performance. While heuristic methods rely on predefined scheduling rules, classical deep reinforcement learning (DRL) models may require more trainable parameters in this setting. Motivated by the potential of parameterised quantum circuits (PQCs) as compact function approximators, we propose QRLQ, a cost-delay-aware quantum cloud scheduling framework integrating PQCs with a dueling double deep Q-network (D3QN) to dynamically account for both cost and delay. Our simulation results show that QRLQ achieves lower mean cost and delay than the heuristic baselines, achieving a 5-11% lower mean cost relative to availability-based and rotation-based heuristics and reducing mean delay by 17% and 82% relative to the strongest and weakest heuristic baselines, respectively, while retaining execution fidelity within 2% of a fidelity-greedy policy. Compared with the classical DRL baseline, QRLQ achieves comparable scheduling performance while using 72% fewer trainable parameters. This work explores the feasibility of using QRL for task orchestration in quantum cloud environments and demonstrates its potential for cost-delay-aware quantum resource management.
cs.LG / 47 / 2609.27473
Learning Where to Look: A Shared Relative-Alignment Module for Time-Series Forecasting and PPG-to-Vital-Sign Reconstruction
Ragamayi Puli, Shunya Nagashima
cs.LG
Abstract
PPG-to-vital-sign reconstruction turns a wrist-worn photoplethysmogram into clinical waveforms such as the ECG. Long-horizon multivariate time-series forecasting underpins planning in energy, weather, and traffic. Both generate a target sequence from a condition sequence, and current models hard-code where each target position reads it, as a same-position copy or seasonal recurrence, so neither transfers between tasks. We propose ROOSTER, one conditioning module that handles vital-sign reconstruction and time-series forecasting alike by learning this correspondence. Its core is a periodic-comb bias over the target-condition offset whose center, period, and sharpness are learned per head, so one module settles on the identity alignment or a seasonal lag and reports which it found. On vital-sign reconstruction from PPG, ROOSTER outperformed the published baselines on four heart-rate and respiratory-rate benchmarks. On multivariate time-series forecasting, it achieved the best horizon-averaged MSE on four benchmarks and outperformed the forecasting model it extends on 20 of 24 dataset-horizon settings under matched three-seed training. An ablation study indicated that the relative bias, not content matching, carried the alignment.
cs.LG / 48 / 2609.27532
ProCredit: From Outcome Rewards to Progress Credit in Agentic Reinforcement Learning
Ming Ma, Yi Zhu, Yiran Zhong, Feida Zhu, Chonghan Liu, Pengkun Jiao, Qichao Wang, Yanhao Jia, Tianming Yang, Steven Hoi
cs.LG · cs.CL
Abstract
Long-horizon agentic tasks require an agent to modify an environment through a sequence of tool calls, with success determined by the final state. The standard recipe assigns a single outcome reward at the end and compares trajectories sampled for the same task. As a result, a group with no successful trajectory yields no training signal, failed attempts cannot be told apart by how close they came to completion, and turns that advance the task receive the same credit as turns that only query the environment. Prior work refines the unit of comparison from the trajectory to the step, or trains a reward model to supply intermediate signal: the former still derives its signal from final success alone, and the latter estimates it with a model. We observe that the acceptance checks that decide success can also be run on intermediate states, so progress is as verifiable as the outcome. We propose ProCredit, which turns this verified progress into credit: it reruns the acceptance checks after each turn, rewards the turn by its change in progress, and uses these rewards to assign credit both across attempts at the same task and across the turns within a trajectory. Starting from Qwen3.5 base models at three scales on AppWorld, ProCredit outperforms outcome-reward baselines and progress-based baselines in task completion rate at every scale on both test sets, exceeding the strongest outcome-reward baseline by 4.1 percentage points at 4B, and results in a second environment show the same direction of improvement. Ablations show that adding the final progress to the trajectory score alone does not improve performance: the gain comes from crediting progress to the turn where it occurs.
cs.LG / 49 / 2609.27547
EBRL: Asynchronous Embodied RL by Multi-Grained Resource Management
Liang Mi, Weijun Wang, Bowen Gao, Tianze Yu, Zixu Hao, Han Xiao, Xin Ding, Mingzhe Huang, Xin He, Lu Shi, Hao Wu, Haipeng Dai, Guihai Chen, Yunxin Liu, Ting Cao
cs.LG · cs.DC
Abstract
Embodied reinforcement learning (RL) improves model capabilities with a pipeline of environment simulation, action generation, and model updates. These stages show heterogeneous CPU and GPU demands, making efficient resource utilization difficult. Recent systems overlap rollout (simulation and generation) with training for efficiency, but exclusive GPU allocation and synchronized barrier in rollout still leave substantial hardware resource waste. In this paper, we present EBRL, an asynchronous embodied RL training system with two core techniques. The asynchronous pipelined scheduler overlaps rollout and training, pipelines simulation and generation across environment groups, and carries out each environment independently, eliminating synchronization stalls. The fine-grained resource manager pools CPU cores and GPU streaming multiprocessors, and uses stage profiles and runtime feedback to adjust resource quotas and batch sizes to meet the shifting demands among stages. We implement EBRL on RLinf and evaluate it with four embodied policies and four simulation benchmarks across heterogeneous GPU testbeds. Experiments show that EBRL achieves 1.30-3.47 times the end-to-end rollout throughput and 2.5 times of training convergency compared to the SOTA embodied RL systems.
cs.LG / 50 / 2609.27554
PhyMo: A Physical-Field Modality for Multimodal AI4Physics
Henan Sun, Haitao Hu, Jin Liu, Jianfeng Zhang, Lujia Pan, Nuo Chen, Jia Li
cs.LG · cs.AI
Abstract
Multimodal learning is emerging as a powerful paradigm for AI for Physics (AI4Physics), where predicting physical systems requires the joint interpretation of heterogeneous observations, measurements, and domain knowledge. However, existing approaches typically represent physical quantities and governing equations as generic numerical or textual tokens, overlooking the physical constraints that determine their spatiotemporal interactions. To address this limitation, we introduce the \textbf{physical-field modality} and propose \textbf{PhyMo}, a physics-grounded multimodal framework that organizes heterogeneous measurements through PDE-associated operators. PhyMo follows a three-stage learning procedure: the physical-field encoder is first pretrained through field reconstruction under PDE residual supervision, its representations are subsequently aligned with visual embeddings in a shared latent space, and the fused multimodal representations are finally processed by corresponding downstream prediction heads. Experiments on five datasets spanning diverse physical environments show that PhyMo achieves state-of-the-art performance, compared to the strongest baseline on each dataset, demonstrating the superiority of PhyMo on multimodal representation learning in AI4Physics.
cs.LG / 51 / 2609.27564
TNLearn: An Open Source Python Package for Task-based Neurons
Meng Wang, Tieyun Li, Juntong Fan, Hanyu Pei, Jing-Xiao Liao, Yaodong Yang, Jianwei Ma, Fenglei Fan
cs.LG · cs.AI
Abstract
The brain does not rely on a single type of neuron to perform all kinds of tasks; instead, it designs different neurons for different tasks. The concept of task-based neurons represents a paradigm shift compared to task-based architectures. It argues that solving a specific problem requires customized neurons, as task-based neurons capture useful prior knowledge from task-related data. To facilitate the use of task-based neurons in scientific research and industrial applications, we introduce TNLearn, an open-source Python package that provides automated construction of task-based neurons and networks, enabling smooth training of task-based networks. Comprehensive documentation, including technical exposition, API reference, and representative examples, is available online. TNLearn is open-sourced at https://github.com/NewT123-WM/tnlearn and has become a PyTorch ecosystem project.
cs.LG / 52 / 2609.27577
VCMM: Variance-Calibrated Momentum for Multimodal Learning
Zhongjing Gu, Chenyang Huang, Yufa Feng, Chong He, Qinxu Ding, Yiming Cui
cs.LG
Abstract
Multimodal joint training often suffers from modality imbalance, where a dominant modality suppresses the optimization of others. Existing methods mainly balance modality learning by modulating gradient magnitudes or directions, modifying optimization objectives, or adjusting training strategies, with most interventions focusing on the current update. However, when combined with widely used momentum-based optimizers, the update also incorporates accumulated information from previous gradients, which is not explicitly addressed by current-step modulation alone. To address this issue, we propose Variance-Calibrated MomentuM (VCMM), which adapts gradient memory to modality-specific gradient dynamics. Specifically, VCMM estimates minibatch noise and temporal drift online and uses their relative strength to determine modality-specific momentum through a Kalman-inspired controller. We further center the control signal across modalities and apply exact bias correction for the time-varying first moment, enabling adaptive gradient memory without extra network passes or explicit learning-rate scaling. Experiments on four multimodal benchmarks demonstrate consistent improvements with modest training overhead.
cs.LG / 53 / 2609.27581
Does Step Law Transfer to Small-Scale Language Models? An Empirical Recalibration Below 59M Parameters
Egor Romanyukov, Timofey Novikov, Timur Shokarov, Elizaveta Zorkina, Anastasia Palienko, Stepan Dergachev
cs.LG · cs.CL
Abstract
Step Law gives power-law formulas for the optimal peak learning rate eta* and batch size B* when pre-training language models. It was calibrated on models between 59M and 1B parameters; the small-model regime N < 59M was never tested empirically by its authors. This regime matters for single-GPU training, interpretability research, educational experiments, and settings where larger models are infeasible on memory or cost grounds. We test whether Step Law transfers to small language models. We consider three outcomes: H1, the original coefficients work directly; H2, the power-law form holds but with different coefficients; and H3, a power law does not describe the optima in this regime. All experiments use a single nanoGPT/TinyStories pipeline with a 2048-token BPE vocabulary, AdamW, and a warmup-cosine schedule. The optimum for each (N, D) cell is extracted from the loss surface L(eta, B) via a local quadratic approximation in log-log coordinates over the smoothed training loss. The final dataset contains 29 unique (N, D) cells and 935 analysis-ready runs. The main refit uses 25 cells (815 runs) in the working range 4 <= D/N <= 600. On the pooled data we accept H2: the functional form is preserved, but the coefficients differ from the original. We obtain eta*(N, D) = 0.0985 N^(-0.508) D^(0.238) (R^2 = 0.834) and B*(D) = 3.6 x 10^(-4) D^(0.931) (R^2 = 0.950). Step Law's structural claim that B* is independent of N is reproduced (p = 0.87), but the growth of B* with D is nearly twice as steep as in the original work. Direct transfer of Step Law systematically overestimates the optimal learning rate: the median ratio eta_SL / eta* is approximately 4.0x, with a range of 2.4x to 6.6x.
cs.LG / 54 / 2609.27588
The Capability Manifold and ML Scaling Laws
Syed Ali Raza Zaidi, Maryam Hafeez
cs.LG · cs.AI
Abstract
Existing machine learning (ML) scaling laws relate predictive loss to compute, model parameters, and data. However, as models are increasingly deployed through agentic harnesses, loss alone is insufficient to characterize downstream performance: models with similar loss can exhibit different capabilities in reasoning, retrieval, planning, and adaptation. Yet, no unified framework connects such capabilities to the coupled resources available across the ML lifecycle. We bridge this gap by introducing a capability manifold, a multidimensional framework mapping downstream capabilities to pre-training, post-training, and test-time resources through bounded scaling functions. Analytical Jacobians quantify capability sensitivity to resource changes and interactions. As an initial application, we embed Kaplan- and Chinchilla-type scaling laws and test-time compute within the framework, demonstrating how existing scaling relationships can be unified as trajectories on a common capability manifold.
cs.LG / 55 / 2609.27593
Hidden not Deleted: How Networks Suppress Entangled Features
Akash Samanta, Manish Pratap Singh, Debasis Chaudhuri
cs.LG · cs.AI
Abstract
Concept erasure methods that operate via linear projection assume that features occupy separable subspaces. We show this assumption fails under dense superposition: when two features are forced into an antipodal pair sharing a single subspace, state-of-the-art linear erasure destroys both, not just the target. Networks trained with gradient descent instead solve this problem non-linearly, but not uniformly: they converge to one of two distinct circuit-level solutions depending on initialization, which we call mirror and shadow solutions. We map this bifurcation as a function of feature entanglement, show it reflects a stable attractor structure rather than an artifact of our setup, and use targeted causal interventions to demonstrate that both solutions leave a substantial, measurable trace of the erased feature's representation intact, recoverable through a single scalar patch rather than requiring any further training. This mirrors a failure mode recently observed empirically in LLM unlearning, where suppression rather than deletion allows forgotten knowledge to resurface; our results offer a mechanistic, causally-validated account of why that failure mode occurs.
cs.LG / 56 / 2609.27594
Efficient Linear Bandits via Cluster-Aware Sketching
Hantao Yang, Hong Xie, Defu Lian
cs.LG
Abstract
We study the problem of computational efficiency for linear bandits in high-dimensional settings with a finite arm set. In linear bandits, the increase in the dimension $d$ of the feature vectors leads to growing computational costs of $O(d^2)$ at each round of update. Traditional sketching-based methods such as SOFUL reduce computation via fixed-size matrix sketching, yet run the risk of incurring vacuous linear regret when the spectral tail of the data is heavy and the sketch size is inadequately selected. To guarantee regret convergence and effectively reduce computational costs, we introduce a clustering mechanism and propose the Cluster Sketch Linear Bandit (CS-LB) algorithm. Our method preserves the full covariance information in each cluster to guarantee robust sublinear regret without spectral-tail vulnerabilities, performs cluster switching by assigning a sentinel for each cluster, and reduces per-round update computation to $O(l^2d)$ via a tunable sketch size $l<d$. Experiments on synthetic datasets demonstrate that our method consistently maintains a favorable trade-off between efficiency and regret.
cs.LG / 57 / 2609.27637
Learning Local Heterogeneity and Cross-Region Context for Large-Scale Traffic Forecasting
Qi Feng, Zidong Wang, Bo Li, Xiaoguang Gao, Jiayu Zhang, Chenfeng Wang, Kaifang Wan
cs.LG · cs.AI
Abstract
Traffic flow forecasting is essential to intelligent transportation systems. Large-scale traffic forecasting requires jointly modeling local spatial dependencies and cross-region context.Spatial dependencies between geographically neighboring nodes are heterogeneous due to differences in road identity and travel direction, while acquiring global information through allpairs node interactions incurs substantial computational costs. Therefore, capturing local heterogeneity while efficiently acquiring long-range context remains an important challenge in largescale traffic forecasting. To address these challenges, we propose LoReST, a Local-Region Spatial Temporal network that models spatial dependencies at two complementary granularities: node neighborhoods and road network regions. Specifically, relation-aware local aggregation captures heterogeneous dependencies within geographic neighborhoods through road and direction specific feature transformations. Cross-region interaction constructs region representations through mean pooling, exchanges long range context via inter-region attention, and broadcasts it back to nodes. By integrating local information aggregation with crossregion interaction, LoReST is able to effectively achieve spatial dependency learning in large-scale road networks. Experiments on four datasets of the LargeST benchmark show average relative reductions of 4.78%, 3.60%, and 5.75% in MAE, RMSE, and MAPE, respectively.
cs.LG / 58 / 2609.27667
Robust Adversarial Reinforcement Learning with Risk Sensitivity and Critic Consistency Regularization
Jiaxi Wu, Tiantian Zhang, Yuxing Wang, Yongzhe Chang, Xueqian Wang
cs.LG
Abstract
Reinforcement learning (RL) achieves strong performance in sequential decision-making but remains brittle under dynamic uncertainty and distributional shifts. Robust Adversarial Reinforcement Learning (RARL) improves robustness via worst-case perturbations, but existing approaches frequently suffer from unstable optimization and degraded value estimation. In particular, overly aggressive adversaries can drive the agent toward uninformative failure states, while adversarial perturbations amplify disagreement between double critics and introduce biased value targets. We propose a unified framework, RACER (Risk-sensitive robust Adversarial critic ConsistEncy-regularized Reinforcement learning), that revisits adversarial RL from a risk-sensitive perspective. First, we introduce a state-dependent adversarial objective that adaptively regulates perturbation strength, suppressing harmful disturbances while preserving informative exploration. Second, we propose critic consistency regularization to reduce disagreement between Q-value estimators and stabilize learning. Comprehensive experiments on challenging continuous control benchmarks demonstrate that RACER consistently improves performance, robustness, and training stability over strong robust RL baselines.
cs.LG / 59 / 2609.27679
What Do Tabular Foundation Models Compute In Context? In-Situ Representation Refinement through Attention-Gated Updates
Tian Zhou, Beverly Jin, Linxiao Yang, Xue Wang, Wenwei Wang, Bingqing Peng, Mengni Ye, Jinjie Gu, Liang Sun
cs.LG
Abstract
What reusable computation should a tabular foundation model learn when every table defines a new supervised task? We develop in-situ representation refinement: support labels guide updates to the episode's representations, and these updates transfer to unlabeled queries without changing model parameters. A regularized leave-one-out objective yields a support correction and its query extension. The leading term separates attention-based reading from state-dependent scaling, motivating RefineICL: an attention-gated, FFN-free contextual stack with selected low-rank feature interaction and typed memory. RefineICL-L24 reaches 0.93836 OVR-AUC and 0.87173 accuracy on AMLB29. A benchmark-informed continuation reaches 1644.8 Elo on the 38-dataset TabArena snapshot, 31.4 Elo above TabPFN-3 under the same evaluation. It also improves all four reported metrics over TabPFN-v3 on both TabZilla views. In a matched 100K-update depth grid, an expanded FFN gives no consistent validation benefit and uses 60.2% more peak inference memory at L8. Internal interventions show that support representations are more than a static source of labels: removing one intermediate support update, while preserving the query output, increases final query cross-entropy in all 72 tested episodes. Together, the derivation and interventions explain how attention-gated updates can construct a task-specific predictor in context.
cs.LG / 60 / 2609.27739
MENO: Memory-Efficient Neural Operator
Shengyang Xu, Weijun Zhang, Jun Hu, Pengzhan Jin
cs.LG
Abstract
We propose the Memory-Efficient Neural Operator (MENO) as a high-performance PDE neural solver based on the Manifold Function Encoder (MFE). MENO features three primary advantages: (1) MENO has a significantly smaller memory footprint and much faster training speed than other popular architectures, with the memory footprint being independent of the data resolution, and therefore holds the potential for scaling up to large-scale models. (2) MENO can accept PDE inputs of arbitrary form, including arbitrary geometric domains and arbitrary discretizations. In particular, it is capable of handling cross-geometry scenarios, i.e., where the input functions and the output solutions are defined on different manifolds. (3) MENO exhibits strong generalization capability, and achieves the best accuracy on most of the benchmarks we tested, compared with the results reported in the literature. The code is available on GitHub at https://github.com/jpzxshi/MENO, and all numerical examples in this paper can be run with a single command to reproduce the reported results.
cs.LG / 61 / 2609.27741
Limiting-Kernel Q($λ$): Bridging Short and Long Horizons
Tolga Ok, Arman Sharifi Kolarijani, Peyman Mohajerin Esfahani, Mohamad Amin Sharifi Kolarijani
cs.LG · math.OC
Abstract
In value-based reinforcement learning, improving the accuracy of policy evaluation has been shown to improve downstream policy optimization performance. The widely adopted family of approximations relying on $n$-step truncation yields computationally efficient value estimators but is inherently limited to a short evaluation horizon. In contrast, methods that exploit the global structure of the transition dynamics can accelerate policy evaluation, but their memory and computational requirements often limit scalability to large or continuous state spaces. To reconcile these limitations, we introduce Limiting-Kernel Q($λ$) (LKQL), an off-policy value estimator that combines $n$-step truncation with a long-horizon approximation based on the limiting kernel (LK). LKQL has the same order of complexity as $n$-step estimators and integrates directly into both on- and off-policy actor-critic algorithms. We prove that, under aperiodicity and in the near-on-policy regime, the operator underlying LKQL improves the policy evaluation convergence rate over its truncated counterpart for sufficiently large $n$, and that LKQL itself converges almost surely to the optimal values in finite Markov decision processes (MDPs) under a fixed behavior policy. On the MuJoCo continuous-control benchmark, we show that LKQL improves over $n$-step baselines in most settings, particularly on long-horizon tasks.
cs.LG / 62 / 2609.27986
Relative Discharge Stage (RDS) Classification: A Practical Indicator of Battery Discharge Progress
Khoa Tran, Tri Le, Hung-Cuong Trinh, Hung Tran-Nam
cs.LG
Abstract
Accurate remaining discharge time (RDT) prediction is challenging in real-world battery applications because future load profiles are unknown and highly dynamic. To address the uncertainty of continuous RDT regression, this paper introduces Relative Discharge Stage (RDS), a battery-management indicator that represents the remaining discharge condition using five interpretable classes: Normal, Good, Moderate, Low, and Recharge Required. Unlike state of charge (SOC), which reflects the current charge level, RDS characterizes the remaining discharge process without requiring future-current information during inference. A physics-informed RDS classification framework is proposed, combining SOC estimation with lightweight temporal learning. The SOC-estimation component includes second-order ECM state and terminal-voltage prediction, hysteresis and OCV temperature correction, core-temperature estimation, and AEKF state correction, supported by OCV evaluation, online STC-ECM parameter adaptation, and pretrained neural residual-voltage correction. The measured current, terminal voltage, surface temperature, and estimated SOC are arranged into a sliding observation window and processed by a lightweight temporal convolutional network. Experiments on two public lithium-ion battery datasets demonstrate robust RDS classification, with accuracy exceeding 80% under varying load and thermal conditions.
cs.LG / 63 / 2609.28003
Learning from Failures: Heterogeneous Graph Memory for Small Language Model Tool-Using Agents
Jiaxing Li, Lei Song, Rui Dong, Youyong Kong
cs.LG
Abstract
Small and medium-sized language models offer cost-effective executors for tool-using agents, making them attractive for local and large-scale deployment. However, in long-horizon and stateful environments, they often make structural errors such as missing required observations, performing premature writes, repeating failed calls, and violating action preconditions. These errors can lead to incorrect state updates, policy violations, and costly or irreversible consequences, making reliable tool execution a critical deployment challenge. Existing fine-tuning approaches require substantial data and computation, while flat memory may retrieve failed actions without preserving their causal context or safety conditions. In this paper, we propose FRESH, a Failure-aware Retrieval framework over Experience-Structured Heterogeneous graphs, which transforms historical successes and failures into structured external experience for tool-using agents. By explicitly modeling the dependencies among tasks, actions, errors, repairs, and execution conditions, FRESH helps frozen language models reuse reliable strategies, avoid recurring failures, and make safer decisions in stateful tool interactions. Experiments on $τ$-Bench and AppWorld with multiple open-source models show that FRESH consistently improves task success and tool-use reliability over no-memory agents and representative memory-based baselines.
cs.LG / 64 / 2609.28006
Shared Global KV with Layer-Specific Local History
Xinglang Xian
cs.LG
Abstract
Decoder-only Transformer language models cache keys and values (KV) to reuse past computation during generation. Sharing KV across layers saves storage but reduces the diversity of representations available across depth. We study what local memory should retain alongside shared global KV, separating historical content from the input source used to form it. At 126M parameters and 2K context, an eight-seed study finds about 1.4% lower held-out test perplexity with local history than with a current-token local branch. Capacity, entry-count and training-compute controls support the value of historical content. In a two-seed comparison, this value persists when adjacent layers share local inputs while retaining independent projections; source sharing also shortens exact cache-construction dependencies. Against GQA and adjacent-layer KV sharing, equal bounded learning-rate searches and new-seed confirmation yield better same-source likelihood with larger caches and higher long-request latency. The ordering against adjacent-layer sharing persists after equal-token adaptation to 8K, with a short-context cost. The eight-seed external-book history effect remains uncertain, and downstream outcomes vary by task. We derive a sufficient suffix schedule that reduces upper-layer construction work while preserving the complete cache in exact arithmetic.
cs.LG / 65 / 2609.28022
PISCES: Physics-Informed Solar-wind Convolutional autoEncoder for Space-weather Anomaly Detection and Early Warning
Kevin Lee, Alison J. March
cs.LG · astro-ph.SR · cs.AI · physics.space-ph
Abstract
Space weather early warning depends on detecting solar wind transients in in-situ measurements at the first Sun-Earth Lagrange point (L1), before they reach Earth. Fixed thresholds can miss combined magnetic and plasma structure, and many learning methods provide a single anomaly score. We present the Physics-Informed Solar-wind Convolutional autoEncoder for Space-weather (PISCES), a convolutional autoencoder trained without catalog labels on OMNI solar wind measurements under physics constraints. Its loss includes magnetic field consistency, an empirical relation between temperature and velocity, the Parker spiral angle, and penalties on changes between consecutive one-minute samples in derived quantities calculated from the reconstruction. At inference, PISCES separates the anomaly score into magnetic and plasma reconstruction errors, physics relations, and residual corrections, and reports the magnitude of each contribution. Attenuation of the skip connections, selected on validation data, improves average precision for the trained models, while the untrained scores remain nearly the same. The trained models also give a more consistent ordering of these physical contributions. After smoothing with a trailing median, the alarms can precede independently observed sudden commencements, including positive sudden impulses.
cs.LG / 66 / 2609.28053
Exact Quantile Balancing and Load-Error Injection for Mixture-of-Experts
Pit Neitemeier, Jiaze Li, Alessio Serra, Philipp Scholl, Sohir Maskey
cs.LG · cs.CL
Abstract
Mixture-of-Experts (MoE) training requires global load balance to prevent expert under-utilization and local balance for efficient expert-parallel execution. Existing distributed Quantile Balancing (QB) uses shard-dependent or approximate global quantiles, while token-independent expert biases cannot ensure microbatch-level balance. We introduce Exact Quantile Balancing (EQB), which computes exact global-batch BF16 quantiles with negligible communication, and Load-Error Injection (LEI), which injects local load errors directly into router-score gradients. On 7.5B-parameter MoEs trained for up to 500B tokens, EQB improves global balance and downstream performance over naive QB, while LEI improves local balance and outperforms the GShard loss at comparable quality.
cs.LG / 67 / 2609.28085
Curriculum Learning with GNN-based Reinforcement Learning for Job Shop Scheduling
Jayakrishnan K. Vasudevan, Jonathan Hoss, Noah Klarmann
cs.LG · cs.AI
Abstract
The job shop scheduling problem is a challenging combinatorial optimization problem, and recent reinforcement learning approaches using graph neural networks have shown promise for learning scheduling policies directly from problem instances. However, training on large instances remains computationally expensive, and generalization across instance sizes remains challenging. This paper studies curriculum learning for graph neural network-based reinforcement learning in the job shop scheduling problem by comparing it with single-size training across three target sizes: 20 x 20, 25 x 25, and 30 x 30. In the curriculum setting, the policy is first trained on smaller instances and then progressively adapted to larger target sizes, allowing scheduling behavior learned in earlier stages to support learning on larger instances. Models are evaluated on unseen instances from 8 x 8 to 30 x 30 using the optimality gap, considering both generalization across all evaluation sizes and specialization on the target size. Results show that curriculum learning consistently reduces wall-clock training time, with larger benefits as the target size increases. The strongest advantage is observed at 30 x 30, where curriculum learning reduces the mean optimality gap across all evaluation sizes by approximately 8.1 percentage points, reduces the target-size mean optimality gap by approximately 8.6 percentage points, and saves approximately 50 hours of training time.
cs.LG / 68 / 2609.28086
LAYERSCOPE: A Layerwise Characterization of Video and Multimodal Learned Representations
Sandra Arcos-Holzinger, Debashish Chakraborty, Rohita Mocharla, Will Walden, Andrew Yates, Reno Kriz, Sarah M. Erfani, James Bailey, Vishal M. Patel, Sanjeev Khudanpur
cs.LG · cs.AI · cs.CV
Abstract
We propose LAYERSCOPE, a label-free, layerwise framework that aims to characterize a model's learned representations in video and multimodal settings. Evaluating downstream performance using representations from final or intermediate layers typically requires large amounts of labeled data, repeated task-specific evaluations, and substantial computation. To address these limitations, LAYERSCOPE uses local, global, distributional, and correspondence-based geometric metrics to compare layerwise representation structure within and across models without requiring task-specific labels. We evaluate seven architecturally diverse models across video and multimodal classification, clustering, and text-to-video retrieval tasks from MVEB/MVEB+. We find that intermediate-layer representations can outperform final-layer and model-default outputs. We also find that no single geometric metric consistently predicts downstream performance, but note that distinct layerwise geometric signatures emerge across model families. LID shows task-dependent relationships with performance, while RankMe provides the strongest measure for classification and clustering, but is not a universal layer selector. We also find that pairing-aware metrics explain retrieval better than distributional distances alone. LAYERSCOPE therefore offers a framework for comparing representations across models and layers, enabling a more systematic evaluation in video and multimodal settings.
cs.LG / 69 / 2609.28105
Fed-ReMasker: Federated Tabular Imputation under Feature-Level Missingness
Ioannis Papathanail, Rooholla Poursoleymani, Lubnaa Abdur Rahman, Stavroula Georgia Mougiakakou
cs.LG · cs.AI
Abstract
Multi-center clinical studies and biomedical research collaborations increasingly seek to utilize data across centers to build models that generalize beyond any single center. This creates two distinct challenges: data protection regulations may restrict the sharing of raw patient data across institutions, while centers may collect only partially overlapping sets of features under different protocols. Federated learning enables collaborative model training without centralizing raw data. However, existing federated imputation methods rarely evaluate feature-level missingness, in which entire features are unobserved at some centers. To address this setting, we adapt the ReMasker masked autoencoder to federated learning (Fed-ReMasker), enabling centers to impute features never observed locally by leveraging knowledge learned across collaborating centers. We evaluate Fed-ReMasker in a benchmark spanning synthetic datasets with linear and nonlinear relationships and real-world tabular datasets, including clinical data. The benchmark varies the number of centers, the missingness ratios, and client heterogeneity. Fed-ReMasker achieves the lowest imputation error in 93.2% of value-level and 96.7% of feature-level scenarios in the homogeneous benchmark. It also remains robust to client heterogeneity using simple federated averaging, outperforming all baselines in all 36 value-level scenarios and each baseline in at least 35 of 36 feature-level scenarios, and comes within 3.0% on average of a centralized model trained on the pooled data.
cs.LG / 70 / 2609.28116
Probabilistic and Geometry Aware Neural Surrogate of Scrape Off Layer Plasma Simulations
Gabriele Gianuzzo, Stefan Dasbach, Fleur Hendriks, Sven Wiesen, Vlado Menkovski
cs.LG · physics.plasm-ph
Abstract
Fast surrogates for tokamak boundary-plasma simulation are typically deterministic regressors mapping a global operating point to a flattened vector of cell values. Near the divertor detachment transition the steady state is not reliably single-valued. A point estimate must average over qualitatively different plasma states, and it arrives with no statement of confidence. Moreover, the flattened vector representation discards the geometric structure of the SOLPS-ITER mesh. This work addresses both problems. We unroll the curvilinear mesh into three fixed-size image tensors whose layout preserves cell adjacency and inverts exactly, letting a convolutional network act on the geometry without loss of information. A conditional flow matching model, well suited to highly sensitive systems, is then trained on this representation. The result is an efficient, scalable surrogate that captures multiple plausible outcomes even at sensitive operating points. Along a gas-puff scan, the predictive distribution splits into a hot and a cold mode across an early regime transition. A further check on synthetic data with an injected bifurcation of known size confirms the model recovers both branches rather than their average.
cs.LG / 71 / 2609.28145
RL Starts before RL: On Policy Distillation for Better Reinforcement Learning
Shuai Dong, Yongfu Zhu, Yuqi Xu, Weichu Xie, Liuwenpu, Ziyue Wang, Kaiwen Tuo, Congcong Wang, Siyuan Wang, Wenqi Shao, Shuai Yang, Ji Zhao, Caoyuan Ma, Wenzheng Chang, Taiqiang Wu, Xinlei Yu, Hongrui Wu, Xiaoxuan He, Fangke Chen, Dianyi Wang, Kanghui Tian, Sirry Chen, Xingyu Liu, Xiangnan Wu, Jiawei Guo, Haowen Hou, LingHan Chen, Zhongyu Wei, Jiaqi Wang
cs.LG
Abstract
Reinforcement learning (RL) improves reasoning, but its performance depends on the policy from which training begins. We study on-policy distillation (OPD) as a preparation stage for RL and ask whether its benefits extend beyond improvements in the distilled model's initial accuracy. Under shared RL settings, students initialized with OPD reach higher final performance than those trained with direct RL or supervised fine-tuning followed by RL. This advantage can emerge even when OPD produces little immediate improvement in accuracy. Pre-RL Pass@k does not fully explain the benefit: similar or even higher values do not necessarily lead to better performance after RL. Behavioral analyses point to alignment with the teacher's distribution beyond top-1 agreement as a possible explanation. Such alignment may favor higher-quality reasoning paths while retaining alternatives that RL can further refine using outcome feedback. We further examine how trajectory sources and divergence objectives affect the value of distillation for subsequent RL. Standard reverse-KL OPD performs better before RL, but forward-KL OPD overtakes it afterward; with teacher-generated distillation trajectories, reverse KL remains ahead at both stages. These findings suggest that the preferred distillation objective depends on both the trajectory source and the training that follows. Our results support evaluating OPD as preparation for RL and selecting distillation choices by the performance achieved after subsequent training.
cs.LG / 72 / 2609.28165
Confidence Falls Short: Asymmetric Certainty Gains from Optimization Hinder Multimodal Classification
Longfei Huang, Xiangyu Wu, Yang Yang
cs.LG
Abstract
Multimodal learning (MML) falls into the optimization dilemma due to the modality imbalance phenomenon, leading to suboptimal overall performance in practice. While many attempts primarily focus on balancing the optimization dynamics across modalities to address this issue, we identify a subtle yet critical flaw: optimization yields asymmetric gains in predictive certainty, with the strong modality more confident than the weak one, driving imbalanced modality contributions. In this paper, our analysis reveals that this flaw stems from unimodal characteristics rather than multimodal learning, and this confidence discrepancy can be corrected by positive cross-modal intervention. Based on this insight, we propose multimodal Max Confidence Regularization (MaxCR) to dynamically intervene in modality semantic confidence. Specifically, the semantic confidence of each modality is tracked using a nonlinear sparsity measure. We then design max suppression and max excitation based on this measure to regularize strong and weak modalities, respectively. They penalize and encourage the top-1 confidence, thereby constraining multimodal prediction. To this end, strong and weak modalities are expected to make calibrated confidence, thereby improving the overall performance. Empirical experiments on widely used datasets reveal the superiority of our method through comparison with various state-of-the-art (SOTA) multimodal learning baselines.
cs.LG / 73 / 2609.28194
Geospatial embeddings detect old-growth forests but buffered spatial validation narrows their advantage over Sentinel features
Thomas Ratsakatika, Mihai Zotta, Srinivasan Keshav, Emily R. Lines
cs.LG · cs.CV · eess.IV
Abstract
Old-growth forests develop over centuries under minimal anthropogenic disturbance, producing structurally complex and biodiverse stands. In Europe, protecting them requires mapping that is accurate for individual forest parcels yet deployable continent-wide. Geospatial foundation model (GFM) embeddings enable label-scarce land classification, but their value for old-growth detection remains unknown. Here, we map old-growth forests across 211,893 ha of Romania's Southern Carpathians, a beech-spruce landscape typical of the Alpine Biogeographic Region. We construct high-confidence, expert-informed reference labels for old-growth and non-old-growth parcels. We add AlphaEarth, TESSERA v2 and Sentinel-1/2 features to a common baseline of topographic and human-access predictors, then compare them under spatially blocked validation with and without 10 km train-test buffers to limit residual autocorrelation. With buffering, GFM and Sentinel-1/2 predictors increase precision-recall AUC by 0.21-0.25 [95% CIs: 0.15-0.34] relative to baseline, indicating spectral data contain a spatially robust old-growth signal. With a PR-AUC of 0.84 [0.79-0.88], TESSERA outperforms Sentinel-1/2 (+0.08 [+0.05 to +0.11]) and AlphaEarth (+0.08 [+0.04 to +0.12]) under unbuffered spatial validation. At a 10 km buffer, however, this advantage narrows to +0.04 [-0.01 to +0.11] and +0.03 [-0.04 to +0.10], intervals consistent with no difference. At 10 m resolution, convolutional neural networks add no benefit over pixel-based XGBoost. Comparisons with four national- and continental-scale products show the importance of non-old-growth labels, and reveal 81% agreement between our predictions and a field-calibrated map. We conclude that buffered spatial validation is vital when transferring old-growth detection models to unseen landscapes, and provide our labels and predictions for future work.
cs.LG / 74 / 2609.28199
Transferable Evidence Reconstruction for Longitudinal Glucose Representations
Tian Zhou, Bingqing Peng, Linxiao Yang, Wenwei Wang, Mengni Ye, Beverly Jin, Zuyi Zhu, Jinjie Gu, Liang Sun
cs.LG
Abstract
Long physiological recordings contain many routine measurements, while predictive information is often concentrated in rare events, sustained burden, and recurring temporal patterns. Masked autoencoding recovers measurements; contrastive learning aligns views. We study self-supervision that explicitly prioritizes structured signal evidence. We introduce transferable evidence reconstruction (TER), which constructs evidence from unlabeled recordings, fits a fresh low-capacity reader on one recording group, and requires that reader to recover the same evidence in another group without refitting. Differentiating through this cross-group test learns representations with transferable evidence-decoding rules; the evidence guides self-supervision but is not used as a downstream feature. For continuous glucose monitoring (CGM), an observation-aware daily encoder and clock-aware multi-day memory bind glucose level and change to recorded time while organizing up to seven days of history. On the 14-task leaderboard, TER improves the strongest prior overall PR-AUC/ROC-AUC/Macro-F1 scores by 5.51/4.43/2.80 percentage points and sets a new best metric on 12/14 tasks. These leaderboard gains are 2.0-2.9 times the respective gaps between the two strongest baselines. With public pretraining data, folds, and the linear probe matched, TER outperforms our GlucoFM reproduction by 6.09/5.52/2.72 points. Target-reader ablations, same-history controls, and cross-person readouts support the combination of structured evidence, cross-group reader fitting, and learned multi-day organization.
cs.LG / 75 / 2609.28208
Support-Compiled Feature Folding: More Evidence at Lower Memory Across Tabular Foundation Models
Tian Zhou, Beverly Jin, Xue Wang, Linxiao Yang, Wenwei Wang, Bingqing Peng, Mengni Ye, Jinjie Gu, Liang Sun
cs.LG
Abstract
Tabular foundation models face a feature-side scaling dilemma: full-width pairwise mixing grows quadratically with the number of columns, whereas feature selection saves memory by discarding evidence. We introduce Support-Compiled Feature Folding (SCFF), a training-free inference framework that resolves this dilemma without changing the frozen backbone. SCFF routes support-ranked features through bounded leaves of the native feature encoder, support-checks the residual evidence, and merges the encoded messages before a single contextual prediction. It thereby converts quadratic feature-interaction work into linear-in-width work with a bounded local working set, without ensembling predictions or training new parameters. On the exhaustive 18-dataset wide-table slice of fixed AMLB-29, TabZilla, and TabArena snapshots, SCFF improves dataset-macro accuracy and NLL on all six evaluated backbones. All four matched-width comparisons retain favorable 95 percent dataset-bootstrap intervals on locked folds, with relative error reductions up to 26.1 percent. Median paired GPU-memory savings are 2.09x to 2.36x, and the ratio of separately observed maximum peaks reaches 34.3x. Under a measured peak-memory ceiling, SCFF uses the saved budget to preserve more support-selected evidence, improving accuracy by 4.06 and 3.72 points over the widest feasible single leaf on predeclared wide-Core strata of TabICLv2 and TabPFN-3.
cs.LG / 76 / 2609.28212
Log-Depth Recurrent Language Modeling
Yiqin Wang, Nuri Cingillioglu, Charles Pert
cs.LG · cs.CL
Abstract
Language modeling using Transformers has become commonplace despite their fixed computational depth and quadratic runtime with respect to input tokens. Recurrent models on the other hand offer linear depth but no parallel execution. In this work, we extend balanced-tree recursive operators from sequence encoding to autoregressive prediction, enabling all prefix representations to be computed with logarithmic depth and linear runtime. Our experiments provide an initial characterization of this model class, demonstrating robust length extrapolation and performance approaching that of ALiBi-based Transformers, highlighting its potential as an alternative architecture for language modeling.
cs.LG / 77 / 2609.28248
hyperbolix: Hyperbolic Deep Learning in JAX
Timo Klein, Thomas Lang, Yllka Velaj, Sebastian Tschiatschek
cs.LG · cs.MS
Abstract
We present hyperbolix, an open-source library for hyperbolic deep learning in JAX, built on Flax NNX. To our knowledge, it is the first comprehensive, general-purpose hyperbolic deep learning library in JAX. It includes six manifolds with a common interface: Euclidean space, the Poincaré ball, the hyperboloid, the $κ$-stereographic model, mixed-curvature product spaces, and the proper velocity space. We implement layer families that cover linear layers, convolutions, attention, normalization, positional encoding, regression, and vector quantization. These building blocks span methods ranging from Ganea's original hyperbolic neural networks to recent fully hyperbolic architectures such as Hypformer and Lorentzian ResNet. Additionally, hyperbolix contains Riemannian optimizers implemented as optax transformations, wrapped distributions, and hyperbolic dimensionality-reduction techniques. Its API uses idiomatic JAX: Manifolds are stateless, with curvature being passed at call time, while manifold operations act on single points, with jax.vmap enabling batch operations. The precision of every checked operation is tested against a closed-form NumPy/SciPy transcription from the source paper or a finite difference, for both float32 and float64. On the hyperboloid, standard formulas for two-point operations, such as the distance, lose precision far from the origin, because they subtract two large, nearly equal terms. hyperbolix replaces these subtractions with cancellation-free formulas that stay accurate in float32 at distances where prior implementations return NaN. hyperbolix is available under the MIT license at https://github.com/timoklein/hyperbolix .
cs.LG / 78 / 2609.28385
When and Where to Trust the Teacher: Unifying On-Policy Distillation and GRPO through Entropy-Calibrated Credit Assignment
Jie Zhang, Jingxiao Yang, Zhehao Huang, Yuhang Liu, Xiaolin Huang
cs.LG · cs.AI
Abstract
Reinforcement learning with verifiable rewards (RLVR) supervises mathematical reasoning through final-answer correctness, but provides little guidance on individual tokens. On-policy distillation (OPD) supplies dense feedback on student-generated responses, yet teacher preference need not reflect correctness. Recent hybrids combine OPD and verifier-derived advantages or reweight task credit using teacher ratios. However, teacher guidance enters after verifier-based group normalization, and token reweighting need not preserve the total task credit assigned to each response. We introduce Unified Entropy-Calibrated Credit Redistribution for GRPO (UECR-GRPO), which integrates verifier and teacher signals within a single GRPO-style update at both the response and token levels. \emph{Path-Utility Unification} (PUU) combines verifier reward and a teacher-to-anchor path log-ratio in a single KL-regularized objective. Its on-policy implementation uses a length-normalized teacher score and combines both rewards before group normalization and PPO clipping, allowing teacher evidence to influence the response ranking. \emph{Entropy-Calibrated Redistribution} (ECR) then uses the signed teacher--old-policy token gap to redistribute the verifier-derived component. Full-vocabulary teacher entropy attenuates uncertain guidance, while a response-wise zero-sum projection preserves the total task credit and its token-wise sign before clipping. Across five mathematical reasoning benchmarks, UECR-GRPO achieves average \(\mathrm{Avg@12}\) accuracies of 17.21\% and 65.09\% with Qwen3-1.7B and Qwen3-4B students, respectively, exceeding the strongest baseline at each scale by 0.89 and 0.56 percentage points.
cs.LG / 79 / 2609.28399
Memory Attention
Jiale Kang
cs.LG
Abstract
Language models typically construct attention values from contextual hidden states, even when some of their content may be reusable across contexts. We investigate whether token-indexed memory can replace the dedicated value projection when complemented by contextual information. We propose Memory Attention (MA), which forms values by combining layer-specific token memory with contextual keys. The memory supplies token-specific representations, while the keys preserve context dependence. At inference, normalization can be folded into the memory tables, reducing value construction to lookup and addition. Token-indexed retrieval also enables CPU offloading with prefetching, reducing GPU parameter storage. Under matched training token budgets and with additional memory parameters, experiments across attention configurations show improved language modeling and average downstream performance.
cs.LG / 80 / 2609.28405
Learning Collective Dynamics with Differentiable Gaussian Representations
Jianxiang Ma, Mingfu Zhang, Xiaocui Yang, Yichen Gao, Junzhao Huang, Yuesong Hou
cs.LG
Abstract
Collective responses depend on individual differences, contact opportunities, and accumulated experience. Learning their dynamics from aggregate counts requires connecting a population's response distribution to both current observations and future behavior. We introduce Differentiable Gaussian Dynamics (DGD), which learns this connection through three components: a Gaussian mixture representing heterogeneous response propensities, differentiable aggregation of contact intensity and behavioral probabilities, and feedback recurrence that updates subsequent responses. Reparameterized integration and temporal recurrence let aggregate prediction errors jointly train the distribution, observation functions, and feedback parameters. On four windows from KuaiRand-Pure and Online Retail II, DGD achieves lower joint behavioral negative log-likelihood than a DeepAR adaptation with a joint-behavior head. In Retail 2010, its one-day behavioral-count MAE is 4.71 versus 6.88 for this adaptation. Learning the distribution reduces behavioral negative log-likelihood by 10.82% relative to a fixed Gaussian in KuaiRand's standard-recommendation window; removing feedback dynamics raises joint KL from 0.0340 to 0.2577 in a controlled experiment. These results establish the value of learning population representations and their feedback process from aggregate observations. Code is available at https://github.com/OranAi-Ltd/oransim.
cs.LG / 81 / 2609.28409
Learning Holographic Reduced Representations with Clifford Variational Autoencoders
Mohamed Malek Abid, P. Michael Furlong
cs.LG · cs.AI · cs.NE
Abstract
Vector Symbolic Algebras project data structures into a hyperdimensional vector space through the application of their vector algebras to randomly generated atomic vector symbols and fractional power encodings of real-valued data. Embedding unstructured data remains an open question. We present \textit{Clifford-VAE}, a variational autoencoder that learns to project data onto a Clifford torus in arbitrary dimensions. Experiments using the MNIST, FashionMNIST, and CIFAR-10 datasets demonstrate that Clifford-VAE produces representations that are competitive with those produced by Gaussian and Hyperspherical VAEs for semi-supervised classification tasks while outperforming Gaussian and Hyperspherical counterparts in the VSA benchmark tests of self-binding and unbinding, role-filler recovery, and bundle capacity. Clifford-VAE provides a principled technique for grounding perceptual data into a symbolic reasoning framework, providing a new approach to a long-standing problem in the VSA literature.
cs.LG / 82 / 2609.28427
Context-Continuous Preference Learning for Exoskeleton Personalization
Sunin Baek, Sungwoo Park, Daekyum Kim
cs.LG · cs.RO
Abstract
Personalizing exoskeleton assistance across operating conditions is constrained by the time and physical effort required to collect user feedback. We examined whether a user's preference landscape varies smoothly across operating conditions and when this continuity supports learning from limited feedback. We propose Context-Continuous Preference Learning (CCPL), a Gaussian-process preference model that shares observations across nearby contexts while retaining context-specific utility estimates. We evaluated CCPL through simulations and retrospective analyses of ankle and elbow exoskeleton preference data from nine healthy adults. In simulations, CCPL improved reconstruction and preference-based Bayesian optimization relative to independent learning when preferences varied smoothly, but showed negative transfer when continuity was weak. In both human studies, full-data reference landscapes estimated separately for each participant and context tended to be more similar between nearby operating conditions. With five exposures per context, CCPL increased mean reconstruction correlation with these references from 0.644 to 0.720 for ankle assistance and from 0.476 to 0.526 for elbow assistance relative to independent learning. The five-exposure budget was approximately 37% lower for ankle and 17% lower for elbow than the estimated independent-learning budgets needed to match these correlations. CCPL also improved held-out response prediction relative to independent learning, while benefits over pooled learning varied. These findings support context continuity as a basis for sharing preference observations under limited feedback, although benefits for online personalization in humans remain to be established.
cs.LG / 83 / 2609.28438
Minimal-Norm Univariate Two-Layer ReLU Classification: Exact Solutions and Global Optimality with Skip Connections
Karolina Drabik, Ben Lewis, Antoni Puch, Etienne Boursier, Piotr Hofman, Matthias Englert, Ranko Lazić
cs.LG
Abstract
We study minimal-norm interpolation and $\ell_2$-regularized logistic-loss minimization for binary classification by univariate two-layer ReLU networks. We give complete geometric characterizations of the optimal classifiers in function space, resolving how the solutions depend on whether hidden-layer biases are included in the parameter norm. When biases are unpenalized, the minimal-norm interpolators are exactly the continuous piecewise-affine functions that hug every label switch and have kinks of the appropriate convexity. When biases are penalized, the minimizer is unique in function space, has exactly one kink in each intermediate same-label segment, and is therefore a sparsest positive-margin classifier. We further show that adding a free affine skip connection leaves these function-space solutions unchanged but fundamentally improves the parameter-space landscape: every KKT point of the constrained problem becomes globally optimal, whereas suboptimal KKT points can occur without the skip connection. We establish analogous global-optimality and geometric results for sufficiently weak $\ell_2$-regularization of the logistic loss. In the unpenalized-bias case, we identify an additional sparsity-like restriction, implying that most minimal-norm interpolators cannot arise as small-regularization limits of margin-normalized logistic-loss minimizers. Numerical experiments across varying dataset complexity and network width support the predicted landscape and sparsity phenomena.
cs.LG / 84 / 2609.28442
Order-Invariant Answers, Order-Sensitive Representations in Mathematical Reasoning
Zhixu Silvia Tao
cs.LG · cs.AI · cs.CL · cs.SC
Abstract
Reordering a set of mathematical rules without changing its meaning should preserve the correct answer, but must a model's internal representations stay invariant too? We investigate this question using synthetic multi-step function-composition problems, each presented under multiple rule orderings with the same correct answer. We measure accuracy and permutation signal-to-noise ratio (SNR), which quantifies how distinctly ordering patterns are represented relative to variation across problem instances. Across 16 language models ranging from 1B to 8B parameters, we find a pattern: models that solve reordered problems more accurately represent different rule orderings more distinctly. Layer-averaged permutation SNR is positively rank-correlated with accuracy in every synthetic setting we evaluate, with Spearman correlations reaching 0.86. These findings highlight a distinction between answer invariance and representation invariance: successful mathematical rule composition can accompany distinct internal representations between equivalent rule orderings. This motivates distinguishing answer invariance from representation invariance, and offers a representational perspective on mathematical reasoning beyond answer accuracy alone.
cs.LG / 85 / 2609.28459
Even Sharper Bounds for Transductive Learning and Its Applications
Yingzhen Yang
cs.LG · cs.IT · math.ST
Abstract
We introduce Sharper Transductive Local Complexity (STLC), a localized complexity method for transductive learning under uniform sampling without replacement. The construction starts from a Bernstein-type concentration inequality for the supremum of the test--train empirical process. Its proof uses the modified log-Sobolev inequality for the swap walk and a two-parameter entropy closure. A peeling argument with a surrogate localization functional then gives excess-risk bounds with the same fixed-point and confidence terms as the classical inductive local Rademacher-complexity bounds, without the additional logarithmic confidence factor in earlier transductive results. For realizable learning over a binary class of VC dimension $\dVC$, with training size $m$, test size $u$, and $u\ge m\ge\dVC$, STLC yields $\cO\{\dVC\log(me/\dVC)/m\}$. This matches the standard inductive rate and, when $m\ge9$, is within a logarithmic factor of the transductive minimax lower bound of order $\dVC/m$. For transductive kernel learning, STLC gives a spectrum-adaptive excess-risk bound without the multiplicative imbalance factors appearing in the earlier local-complexity bound.
cs.LG / 86 / 2609.27167
Median Temporal Ensembling: Training-Free Robust Aggregation for Action-Chunked Visuomotor Policies
Yuhang Jiang
cs.RO · cs.LG
Abstract
Action-chunked visuomotor policies predict overlapping trajectories, so every executed action is covered by several predictions. Temporal ensembling smooths execution by combining these predictions with an exponentially weighted mean. One corrupted prediction can move the aggregate without bound: its breakdown point is 0. We use adversarial corruption to stress this deployed aggregator and to compare two kinds of guarantee. A metric guarantee bounds the response to a perturbation of a given size. A combinatorial guarantee instead bounds the damage when at most q of the M candidates covering a timestep are corrupted, whatever their size. Encoder adversarial fine-tuning recovers 44% of the loss under the published patch attack, but only 7.3% after the attacker's step size is increased. By contrast, the coordinate-wise median of the same candidate set keeps its recovered fraction flat as attack optimisation increases. Median temporal ensembling costs one line and requires no retraining. Across 25 (configuration, corruption-level) combinations it is never worse than the mean and is significantly better in 15. It also transfers to a second policy class, and it recovers performance under a failure with no attacker in the loop at all: camera frames that arrive blank. Its effect on clean data is configuration-dependent, from -0.04 to +0.07. We also give the boundary: corruption that shifts every covering prediction by the same amount is invisible to this whole family of statistics, and no equivariant aggregator can remove it.
cs.LG / 87 / 2609.27747
Less Language, More Latents: Annotation-Efficient VLAs for Driving
Alexey Zakharov, Kemal Oksuz, Puneet K. Dokania
cs.RO · cs.LG
Abstract
Vision-language-action models (VLA) promise human-steerable autonomous driving, but their training is bottlenecked by the scarcity of frames paired with natural-language instructions: while camera streams and expert trajectories are logged at scale, language annotations (e.g., turn left at the intersection) remain scarce and expensive to acquire. To address this challenge, we introduce Latent Action Driving Annotations (LADA), a three-stage pipeline that transforms abundant unlabelled observation-trajectory pairs into a substrate for language-conditioned control. First, we train a latent action model with a vector-quantised bottleneck, producing a compact codebook of high-level vehicle intents. Second, a small language-annotated subset is used to train a vision-language translator to map observations and language instructions into this codebook. Third, we train a driving VLA on observation-latent-action pairs over the full unlabelled corpus. Using fewer than 5% of language annotations and without leveraging any auxiliary chain-of-thought reasoning or visual question answering streams, LADA achieves a Driving Score of 87.98 and a Success Rate of 70.46% on the closed-loop Bench2Drive benchmark, matching or surpassing fully supervised baselines.
cs.LG / 88 / 2609.28258
Generalizable Robotic Insertion with World Models
Nicklas Hansen, Iretiayo Akinola, Yijie Guo, Jie Xu, Bingjie Tang, Hao Su, Xiaolong Wang, Abhishek Gupta, Dieter Fox, Yashraj Narang
cs.RO · cs.CV · cs.LG
Abstract
Robotic assembly in high-mixture settings requires adaptable systems that can handle diverse parts, yet current approaches typically rely on policies specialized to each insertion task. Although this can reach high success rates, it makes the process of deploying systems for new problems tedious and time consuming. We present a framework for generalizable insertion using world models that combine robot proprioceptive information with raw visual observations captured by a wrist-mounted camera. Our model-based approach trains a single world model on up to 90 insertion tasks with geometrically diverse parts, achieving 56% zero-shot success on unseen objects with unknown geometry compared to just 7% with a model-free baseline. Importantly, performance improves as more objects are included in the training dataset, demonstrating strong scalability. Lastly, finetuning the generalist model on held-out objects significantly enhances data-efficiency compared to training from scratch and, in some cases, achieves better asymptotic performance. To our knowledge, this is the first system capable of assembling unseen objects in an entirely data-driven manner, and thus represents a significant step toward scalable, generalizable robotic assembly systems.
cs.LG / 89 / 2609.28364
LEAP-CBF: A Safety Filter for Uncertain Systems with Least-Effort Adversarial Potentials
Oswin So, Eric Yu, Chuchu Fan
cs.RO · cs.LG · math.OC
Abstract
Control barrier functions (CBF) are a popular safety filter to ensure safety for nonlinear dynamical systems. However, when the system is subject to uncertainties and disturbances, this requires the use of robust variants of CBFs, which can be difficult to construct and can be overly conservative, especially for high-dimensional systems under input constraints. In this work, we propose a new approach to solve these challenges by introducing Least-Effort Adversarial Potentials (LEAP), a certificate that quantifies the robustness of a given state against disturbances in terms of the effort required by the disturbance to cause failure. We show that LEAP is a CBF for the undisturbed system, but can also be used to construct a safety filter that is robust to disturbances whose cumulative effort is bounded. We propose a method for constructing LEAPs with on-policy deep reinforcement learning. Next, we demonstrate LEAPs in simulation on a variety of multi-agent systems with disturbances and uncertainties. Finally, hardware experiments on a quadruped and quadrotors validate that LEAPs are well suited to tackle the disturbances and uncertainties from real-world robotic systems.
cs.LG / 90 / 2609.27222
Physiologically Informed Digital Auscultation for Pneumonia Detection in Long-term Care Residents
Nicholas Rasmussen, Oleg Zaslavsky, Zih-Ling Wang, Hongyu Yu, Joelle Fathi, Kaibao Nie, Amil Khanzada, Tomoko Ito
cs.SD · cs.CV · cs.LG · eess.AS · q-bio.TO
Abstract
Pneumonia is difficult to diagnose in older long-term care residents; multimorbidity and atypical presentations obscure signs, motivating operationally efficient objective testing. We analyzed multi-channel digital stethoscope recordings from 185 Japanese residents (73 pneumonia, 112 symptomatic without), using radiologist-confirmed chest X-rays and clinician diagnoses as supervisory signals that train convolutional neural networks, multimodal fusion, and channel-based variants with time-domain Grad-CAM interpretability. Models were evaluated with repeated patient-level cross-validation showing models with X-ray supervision outperformed clinician supervision (F1 0.729, accuracy 0.783 vs. F1 0.637, accuracy 0.711). Additionally, a three-channel selection protocol maintained performance (F1 0.736; accuracy 0.803), with two mid-thoracic sites ranking highest and Grad-CAM attention overlapping adventitious sounds. These findings indicate automated multi-channel lung-sound analysis can aid long-term care pneumonia diagnosis, with X-ray supervision being more reliable than clinical, and fewer channels preserving performance while lowering acquisition times.
cs.LG / 91 / 2609.27389
EvoAudio: Recursive Self-Improvement for Audio Understanding
Yuxiang Wang, Shengbo Cai, Yingda Shen, Ming-Hao Hsu, Qinke Ni, Liqiang Zhang, Teddy Sun, Steve Yevs, Zhizheng Wu
cs.SD · cs.LG
Abstract
Audio language models understand what is said far better than how it sounds. Closing this gap takes more than data. Detailed acoustic annotation is costly, labels from stronger models inherit their errors and limits, and fixed data cannot adapt as the learner improves. We therefore propose EvoAudio, a recursive self-improvement system for audio understanding. To our knowledge, it is the first to evolve the model, waveforms, questions, and difficulty in one closed loop. EvoAudio uses the current model's performance to set the focus and difficulty of the next training data. A library of audio tools then constructs questions whose answers follow from how the audio was made, providing verifiable supervision without new human annotation. Reinforcement learning updates the model, and validation decides whether it enters the next evolution round. Across 13 rounds, EvoAudio improves five models with different audio encoders and language backbones on MMSU, MMAU-Pro, and MMAR. It achieves the highest average for every backbone, raising overall performance by up to 6.3 points. The improvement unfolds over successive rounds, with each stronger model starting the next round.
cs.LG / 92 / 2609.26960
Untangling the Geometry and Speed for RF Sensing Spectrograms
Mert Torun, Darius Cuenca, Yasamin Mostofi
eess.SP · cs.LG
Abstract
A fundamental challenge in RF sensing is that Doppler signatures observed by a link entangle the target's motion with the sensing geometry, resulting in limited applicability to unconstrained real-world settings. In this paper, we establish a new foundation for physically interpretable RF sensing that disentangles reflector speed from geometry, jointly recovering the speed, geometry factor, relative amplitude, and width of each dominant Doppler ridge. More specifically, we first develop a compact parametric representation of WiFi spectrograms and establish its low-dimensional structure through a systematic computer-vision analysis of a large and diverse human-activity dataset, thereby providing a tractable foundation for learning. Building on this representation, we then design a physics-informed autoencoder whose structured bottleneck and differentiable RF forward model enforce physically meaningful estimates of reflector speed and geometry. We further introduce a synthetic-to-real training framework, eliminating the need for real WiFi training data. We extensively validate the proposed framework under both known and time-varying geometries, using both independently generated synthetic test sets and 31 real WiFi experiments. The results demonstrate the superior performance in speed and geometry extraction, robustly recovering the underlying geometry, speeds, Doppler-ridge amplitudes, and ridge widths across all settings, while substantially outperforming the strongest baselines.
cs.LG / 93 / 2609.28149
EvEMTBench: An Open Benchmark for Machine Learning in Power System Protection
Julian Oelhaf, Georg Kordowich, Christian Bergler, Andreas Maier, Johann Jäger, Siming Bayer
eess.SY · cs.LG
Abstract
Studies of machine-learning-based power system protection are difficult to compare because task definitions, measurement access, data partitions, metrics, and generalization conditions often differ. EvEMTBench addresses this gap with an open, executable, and versioned benchmark that fixes these evaluation choices while leaving model design open. Across four grids spanning 20-345 kV, it defines 12 protection and event-analysis functions instantiated as 24 scored tasks and supports structured evaluation across observability conditions, predefined distribution shifts, and zero-shot and fine-tuned cross-grid transfer. Committed partitions, leakage controls, and reproducible reporting provide a common basis for comparing future methods. A reference evaluation spanning trivial, conventional, feature-based, and deep-learning baselines shows that wider observability is not uniformly beneficial, shifted conditions can reveal failures not apparent in-distribution, and cross-grid transfer is substantially stronger for fault detection than for fault localization. Protection-relevant diagnostics identify failure modes not apparent from primary metrics alone. EvEMTBench therefore makes generalization in machine-learning-based protection an explicit and reproducible evaluation problem.
cs.LG / 94 / 2609.28273
Non-Commutative State Tracking with Input-Dependent Low-Rank Updates in Mamba-3
Hiroki Fujii, Masaki Yamakita
eess.SY · cs.LG
Abstract
State tracking from sequential observations can require both retaining information and updating it by composing observed operations. We extend Mamba-3's diagonal transition with an input-dependent low-rank reflection term to support noncommutative state tracking, in which the order of operations matters. The rank-one update couples state coordinates along an input-dependent direction, enabling non-diagonal state transitions within a single Mamba-3 block. The extension preserves Mamba-3's exponential-trapezoidal discretization, rotary embeddings (RoPE), and readout. For training, we adapt chunkwise computation to parallelize the proposed recurrence within each chunk. Experiments cover group word problems with discrete inputs and a shell game with continuous observations, in which a policy is trained by behavioral cloning. Among the models selected for their strong performance under fixed timing, the proposed model maintains higher tracking success on longer swap sequences in the shell game with continuous observations and timing jitter. These experiments show that the proposed method achieves high accuracy on the evaluated non-commutative tracking tasks, improving on standard Mamba-3. The extension thus offers a Mamba-3-based approach to non-commutative state tracking.
cs.LG / 95 / 2609.27325
A Hybrid Iterative Deep Ritz Method for Elliptic Interface Problems
Tianhao Hu, Bangti Jin, Fengru Wang, Yifeng Xu
math.NA · cs.LG
Abstract
In this work, we propose a hybrid iterative deep Ritz method (H-IDRM) for a class of interface problems for second-order elliptic operators. It is based on a new mixed formulation of the problem and involves solving a sequence of convex minimization problems. We employ a level-set neural network architecture, featuring a level-set representation of the interface, to accommodate the piecewise smoothness of the solution and the flux. The approach involves only volumetric representations instead of duality pairing on the interface and avoids explicit interface sampling that is inconvenient for complex interface geometries. Further, we present an analysis of the method, including the errors arising from the neural network approximation, Monte Carlo approximation, iterative scheme, and penalty parameters. Numerical experiments indicate that the H-IDRM outperforms existing neural solvers on problems with high-dimensional domains, intricate interface geometries, and lower subdomain regularity.
cs.LG / 96 / 2609.26951
Rolling Conformal Prediction in Sequential Model Training
Chen Cheng, Ruiting Liang, Rina Foygel Barber
math.ST · cs.LG · stat.ME · stat.ML
Abstract
We introduce Rolling Conformal Prediction (rolling-CP), a distribution-free predictive inference method for the setting of sequential model training. Specifically, given a data stream $(X_1,Y_1),(X_2,Y_2),\dots$, at each time $n$ the trained model may depend on the observed history $\{(X_i,Y_i)\}_{i<n}$. This setting arises naturally in modern sequential training, including one-pass training over massive datasets and continual fine-tuning or test-time adaptation of language models during deployment. Rolling-CP first calibrates each incoming observation against the current predictor and then rolls it into future training. In this way, we avoid the need for data splitting. Remarkably, although the models at times $n=1,2,\dots$ may have entirely different properties and accuracy levels, for exchangeable data it is nonetheless possible to establish a guarantee of marginal coverage, with a familiar universal factor-two guarantee (a worst case guarantee of $1-2α$ coverage, as compared to the target level $1-α$), without any assumptions of stability or any restrictions on the model training process. For i.i.d. data streams, we further prove high-probability training-conditional validity uniformly over time; under stability conditions, coverage guarantees sharpen towards $1-α$. Numerical experiments on sequential regression, multiclass SGD, and one-pass neural-network training further demonstrate the practical effectiveness of rolling-CP.
cs.LG / 97 / 2609.28391
Quantum score matching with applications to learning thermal states
Yulong Dong, Jiaqi Leng
quant-ph · cs.LG
Abstract
Score matching has driven major advances in classical generative learning by enabling models to learn from data without evaluating intractable normalization constants, or partition functions. Yet, extending this principle to quantum learning requires rethinking its foundations, as quantum states are described by noncommuting density operators rather than scalar probabilities. The noncommutativity creates fundamental challenges not only in defining quantum scores, but also in developing a training framework with efficient circuit implementations and rigorous theoretical guarantees. In this work, we bridge this gap by establishing a general quantum score-matching framework with end-to-end theoretical guarantees. Applied to Gibbs-state learning, our approach avoids additional thermal-state preparation and achieves information-theoretically optimal sample complexity in the high-temperature regime for Hamiltonians with bounded locality and interaction degree. This positions score matching as a new route to state-of-the-art performance in learning quantum Gibbs states. Beyond these theoretical results, numerical simulations show that our method remains effective even when gradients are estimated inaccurately under limited measurement budgets. Experiments on IBM quantum hardware further demonstrate that quantum score matching is NISQ-friendly: without any error mitigation or correction, it reduces the relative Hamiltonian-parameter error from 64% to approximately 10%. Together, these results extend score matching into an experimentally realizable paradigm for quantum-state learning.
cs.LG / 98 / 2609.28425
Repairability of Inexact Solvers in Recursive State Estimation with Machine Learning
Yanjun Ji, Dennis Willsch, Orkun Şensebat, Priyanka Arkalgud Ganeshamurthy, Zhi Pei, M. Sahnawaz Alam, Ivelina Stoyanova, Frank K. Wilhelm, Bo Zhao, Chao Wang, Kristel Michielsen
quant-ph · cs.LG
Abstract
Recursive state estimation often executes approximate numerical solutions inside a feedback loop, where highly accurate local steps do not guarantee better overall results. For a fixed linear Kalman model, we characterize when a correction within a prescribed subspace and norm budget can meet a local admissibility tolerance, and how the defects actually executed affect the finite-horizon covariance response. Centering each defect on the exact gain for the implemented covariance separates current solve error from inherited gain drift. Expanding the exact residual-drift identity reveals opposing quartic contributions beyond the quadratic response: innovation-covariance inflation enters positively, while local-gain reoptimization enters subtractively. Under matched initialization, an absolute sixth-order remainder bound, uniform over bounded defect sequences at fixed horizon, gives sufficient conditions for quadratic under- or overprediction. Machine learning proposes bounded corrections, while a learner-independent residual certificate and verified fallback govern execution of classical and quantum candidates without changing the reference estimator. In a power-grid tolerance study, learned correction lowers the minimum conjugate-gradient iteration count for deployment without fallback relative to uncorrected solves under the same residual certificate. Gains reconstructed from a variational quantum linear solver and from an annealing-based binary encoding, with small-scale terminal measurements on superconducting hardware and sampling on a quantum annealer, are executed through the same interface. By linking local repairability to nonlinear error propagation, the framework evaluates approximate solvers and learned corrections through independent certification and finite-horizon response, providing a practical basis for studying hybrid quantum--classical computation.
cs.LG / 99 / 2609.27034
CVaR anchor regression protects against rare shifts
Malte Londschien
stat.ME · cs.LG
Abstract
We study prediction in new environments when training data contain rare, large shifts. Anchor regression penalizes the average of the squared mean residual across environments. It protects against shifts in an ellipsoid determined by the second moment of the training shifts. Covering rare shifts may therefore require a large penalty, expanding the ellipsoid in every direction and reducing accuracy on common environments. We propose CVaR anchor regression, which replaces the average of the squared mean residuals with a tail average. Unlike CVaR or GroupDRO applied directly to prediction risks, it does not give environments more weight solely because their noise levels are high. We prove an exact worst-case risk guarantee under a linear structural model that allows for heteroscedastic noise. For discrete environments, decreasing the CVaR tail fraction expands the robustness set from an ellipsoid to a scaled convex hull of the training shifts and their negatives. A separate parameter controls its scale. Examples show how the method can improve protection against rare shifts while retaining accuracy on common environments. We illustrate the method on New York City taxi data.
cs.LG / 100 / 2609.26978
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
Shinsaku Sakaue
stat.ML · cs.DS · cs.LG
Abstract
We study online inverse linear optimization with a fixed unknown linear utility: in each round, an environment presents a compact action set, the learner recommends an action from it, and the environment returns an action that maximizes the utility over the same set. When the utility vector and the actions lie in the $d$-dimensional Euclidean unit ball, we give a randomized algorithm whose regret---the cumulative utility shortfall relative to optimal actions---is $O(\sqrt d)$ in expectation for every time horizon, without knowledge of the horizon. The dependence on $d$ is optimal up to a constant factor by the known $Ω(\sqrt d)$ lower bound for horizons $T\ge d$. Our algorithm maintains matrix multiplicative weights on polynomial feature spaces at geometrically spaced scales. It selects a recommendation distribution by solving a linear program and updates its score matrices by comparing the available actions with the feedback action. With rational oracle outputs and feedback actions, an implementation computable relative to a linear-optimization oracle preserves the $O(\sqrt d)$ regret bound. Whether the same rate is attainable with running time polynomial in the dimension, horizon, and input length remains open.
cs.LG / 101 / 2609.27180
Artificial intelligence surrogates for treatment effect estimation with before-and-after data
Frances Dean, Anna Neufeld, Joshua Barrios, Geoffrey H Tison, Ahmed Alaa
stat.ML · cs.LG
Abstract
Estimating the causal effects of medical treatments is difficult when clinically important outcomes are costly to measure or require long follow-up. Short-term or inexpensive surrogate outcomes offer a potential alternative, but surrogate biomarkers may be unavailable or difficult to identify. Advances in artificial intelligence (AI) have enabled increasingly accurate prediction of clinical outcomes from inexpensive, high-dimensional measurements, which creates an opportunity to use AI predictions themselves as surrogates. To this end, we develop a framework for estimating treatment effects from paired measurements obtained before and after treatment for each treated individual. A pretrained AI model is applied to the before and after measurements, and our estimator compares the resulting outcome predictions. We characterize the technical assumptions under which this within-person contrast identifies the average treatment effect on the treated, even when clinical outcomes are never observed for treated individuals. When these assumptions cannot be justified, we use prediction-powered inference to correct bias using a small number of observed clinical outcomes and obtain valid inference. Synthetic and real-world cardio-oncology experiments demonstrate the validity and accuracy of the approach.
cs.LG / 102 / 2609.27206
Prediction with Expert Advice: Anytime Regret with Many Experts Matches the Fixed-Time Constant
Yang Cai, Vineet Gupta, Yanchen Jiang, Christopher Liaw, Aranyak Mehta, Grigoris Velegkas, Di Wang
stat.ML · cs.LG
Abstract
Prediction with expert advice is a fundamental problem in online learning. When the time horizon $T$ is known in advance, the minimax cumulative regret over $n$ experts is asymptotically $\sqrt{\frac{T \ln n}{2}}$. This is achieved by the Multiplicative Weights Update algorithm with a learning rate tuned to $T$, and is known to be tight. If instead the regret bound is required to hold simultaneously at every time $t$, the best known guarantee has been $\sqrt{t \ln n}$---a factor of $\sqrt{2}$ worse---and it has remained unknown whether this factor of $\sqrt{2}$ is necessary. We show that it is not. We give an algorithm, requiring no knowledge of the horizon, whose cumulative regret satisfies $R_t \le \bigl(1 + O(\sqrt{\ln \ln n / \ln n})\bigr)\sqrt{t \ln n / 2}$ simultaneously for every $t \ge 1$.
cs.LG / 103 / 2609.27241
On the Sample Complexity of Active Learning with Membership Queries
Ganghua Wang, Shaddin Dughmi
stat.ML · cs.LG · math.ST
Abstract
This work revisits a fundamental question in active learning: how powerful is the ability to synthesize arbitrary queries? Compared to pool-based active learning, where the learner only selects queries from a given unlabeled pool, we find that this seemingly mild change in query ability may dramatically alter the difficulty of statistical learning. In particular, some hypothesis classes that are inherently slow to learn in the pool-based setting, achieving only polynomial error decay in the number of samples, become exponentially learnable once synthesized queries are allowed. This striking gap suggests that membership query synthesis induces a fundamentally different mode of learning, one that is not adequately captured by existing active learning theory and calls for new analytical tools to characterize its complexity. Motivated by this phenomenon, we develop several sufficient conditions, present intriguing examples, and propose a conjectural perspective toward understanding which hypothesis classes admit efficient learning through synthesized queries.
cs.LG / 104 / 2609.27280
Multitask Regression with Pairwise Fusion
Xiaodong Li, Zhentao Li
stat.ML · cs.LG
Abstract
We study multitask regression when coefficient sharing can differ by predictor. For a given predictor, many tasks may have the same coefficient while a few differ, and the exceptional tasks need not be the same for another predictor. We describe this structure by two quantities: the number of active predictors and the total number of task coefficients that differ from the most common value for their predictor. We estimate the coefficient matrix by penalizing all pairwise coefficient differences across tasks, with an additional group penalty when predictor selection is needed. The resulting upper and lower bounds have the same dependence on these two quantities. We also consider the stronger setting in which a large set of tasks shares one entire coefficient vector. Under explicit sample-size conditions, the same pairwise estimator pools those tasks exactly, while allowing the remaining tasks to differ. Simulations and household energy data illustrate the transition between broad sharing and task-specific coefficients.
cs.LG / 105 / 2609.27654
FedIncome: Federated Learning for Income Estimation in Digital Lending Under Data Sovereignty Constraints
Sultan Amed, Tanmay Sen, Sayantan Banerjee
stat.ML · cs.LG · q-fin.ST
Abstract
Verified income is often unavailable in digital loan applications, forcing lenders to rely on reported income and potentially leading to over-lending, overly conservative offers, or rejection of creditworthy applicants. Cross-institutional data-sharing constraints make this problem especially difficult for smaller lenders with limited training data. We introduce FedIncome, a federated learning framework for income estimation that enables institutions to train a shared model without pooling raw borrower records. Using more than one million LendingClub loans partitioned into $50$ state-level clients, we simulate a heterogeneous lending consortium. The best federated model achieves out-of-time $R^2=0.608$, compared with $0.619$ for a pooled centralised benchmark. Small-sample clients obtain an average out-of-time $R^2$ improvement of $3.8$ percentage points relative to the pooled centralised benchmark, while the fitted client-level relationship places the empirical crossover at approximately $4,790$ training observations in this setting. When pooling is infeasible and the relevant alternative is local-only training, federation improves out-of-time performance across all sample-size groups, with the largest gains for data-scarce clients. We also combine federated income estimates with state- and income-specific debt-to-income thresholds. In a retrospective decision analysis, replacing reported income with the federated estimate increases simulated approval rates with only modest changes in observed default rates. FedIncome supports collaborative learning under data-locality constraints with little aggregate loss relative to pooled training and larger gains relative to local-only estimation.
cs.LG / 106 / 2609.28015
Improving Ensemble Filters with Flow Matching
Haoyuan Chen, Alexandre Thiéry
stat.ML · cs.LG · math.DS · stat.ME
Abstract
Data assimilation estimates a dynamical state from partial and noisy observations. Classical ensemble filters are efficient but restrict analysis updates through finite sample covariance and affine Gaussian distribution. We introduce the Flow Ensemble Filter (FlowEF), which uses conditional flow matching to transport the forecast ensemble from a classical baseline filter to an analysis ensemble. FlowEF uses a localized Gaussian source during training, transports forecast ensemble members from a baseline filter at deployment, and conditions its velocity field on ensembles from that baseline filter and the observation. The proposed model therefore learns a nonlinear update while mapping each baseline ensemble independently. For sparsely observed dynamical systems, FlowEF improves both deterministic and probabilistic metrics over all four classical ensemble filters. It also achieves the best performance among the state-of-the-art generative data assimilation models.
cs.LG / 107 / 2609.28122
NPBoost: Neural Processes with Gradient-Boosted Fixed Effects
Andrea Nava, Ken Rölli, Armin Begic, Fabio Sigrist
stat.ML · cs.LG
Abstract
Neural Processes (NPs) are model-based meta-learners that implicitly learn a stochastic process and adapt to a new task from a small context set. Most extensions of NPs focus on improving the neural network architecture. We instead develop an extension motivated by the shared hierarchical interpretation of meta-learning and mixed-effects models. Specifically, we introduce Neural Process Boosting (NPBoost), which decomposes structured response variability into tree-boosted fixed effects shared across tasks and NP random effects that capture stochastic task-to-task variation. We propose to train the two components jointly using a boosting algorithm in which an NP learns residual task-specific structure and a tree ensemble estimates common patterns across tasks. Across synthetic and real-world tabular meta-learning problems, this decomposition improves over a standard NP when the shared structure contains discontinuities or other irregular patterns that boosted trees can represent effectively.
cs.LG / 108 / 2609.28177
How Sensitive Are LLM Leaderboard Claims to Hidden Model Selection?
Chen Yang, Xianyang Zhang, Jun Chen
stat.ML · cs.LG
Abstract
LLM leaderboard gains can reflect selection among privately evaluated model variants, yet neither the number of variants nor their dependence is public. We ask how many hidden variants a published margin can support while retaining statistical evidence of a provider's advantage over a fixed comparator. For a fixed candidate family under a Gaussian margin model, we derive a sensitivity curve that reports this maximum count as a function of a lower bound on within-family correlation. The relevant correlation must match the score used for ranking and the sampling model: in a controlled family, pooled item correlation is 0.90, whereas composite-score correlation is 0.46 under item resampling and 0.92 when MMLU subjects are resampled. An item-based audit of 394 adjacent-rank claims on the Open LLM Leaderboard finds that 391 lack statistical support even before accounting for selection. Among claims that pass the uncorrected test, certification can depend on assumptions about the hidden family's correlation. The resulting curves make these assumptions explicit without estimating the unobserved search size.
神经与进化计算 (cs.NE)
4
cs.NE / 1 / 2609.26940
The Computational Value of Sensory-Aligned Receptive Fields Depends on Neuronal Expressivity
Agnese Adorante, Aaron Spieler, Anna Levina
cs.NE · cs.LG · q-bio.NC
Abstract
Biological sensory neurons have selective receptive fields organized along meaningful stimulus coordinates, such as frequency, motion direction, or retinotopic position. Such structure may arise from efficient coding and biological constraints on activity, connectivity, and wiring, as computational studies of simple neurons have shown across modalities. This raises a question: do structured receptive fields confer a computational advantage beyond resource efficiency itself, and does this advantage persist when individual neurons are highly expressive? We address this question in recurrent networks of Expressive Leaky Memory neurons, where we can independently vary neuronal complexity and the organization of feed-forward receptive fields. Across auditory and event-based visual classification tasks, receptive fields aligned with a task-relevant sensory coordinate improve test accuracy relative to budget-matched random receptive fields. This advantage disappears when sensory coordinates are scrambled, or when receptive fields follow task-irrelevant coordinates, showing that the benefit comes from alignment with task geometry rather than restricted connectivity alone. Increasing neuronal complexity reduces the performance advantage of structured receptive fields. Finally, generic synaptic sparsity regularization induces input selectivity and partially recovers performance, but remains substantially below explicitly structured receptive fields, suggesting that sparsity alone is insufficient to recover the full computational benefit of task-aligned receptive fields. Together, our results show that appropriate receptive fields can serve as a computational prior beyond sparsity itself, and that their value depends on the computational expressivity of individual neurons.
cs.NE / 2 / 2609.27242
Combining LLMs and Genetic Search for ARC-AGI-2
Val Dyachenko
cs.NE · cs.AI
Abstract
LLMs can generate programs for ARC-AGI-2 tasks, but the provided compute only allows a small number of attempts to generate, debug and validate solutions. Genetic algorithms can search and test many more programs, but random search rarely starts in a useful neighborhood of the solution space. We combine the two methods through a compact domain specific language (DSL). First, a quantized Qwen3.5-4B LLM generates an initial set of programs for each ARCAGI-2 task. Then, we use those programs to seed an initial population of starting programs, and use genetic algorithms to evolve these programs towards a solution to the given task. The DSL is designed such that every mutated program remains valid and can be executed. The initial programs proposed by the LLM solve 2 (3.3%) of the first 60 tasks of the ARC-2 public evaluation set. The genetic algorithm solves an additional 4, giving 6 correct test outputs in total (10.0%). If we try using evolving solutions without this LLM seeding, we do not arrive at any solutions at all. The results show that genetic search can improve programs generated by LLMs and produce additional correct solutions.
cs.NE / 3 / 2609.27430
An Unbounded Archive-based Transfer Strategy for Dynamic Multi-Objective Optimization with a Changing Number of Objectives
Zhiyun Xiao, Ke Shang, Yajun Liu, Jianguo Li, Shaojiang Wang, Wei Sun
cs.NE
Abstract
Dynamic multi-objective optimization with a variable number of objectives is difficult because objective-dimensional variations may significantly change the Pareto front and degrade algorithm adaptability. This paper proposes an unbounded archive-based transfer strategy (UATS), which maintains an unbounded archive of offspring solutions within each environment stage and extracts feasible nondominated solutions as transferable elites when objective changes occur. UATS is embedded into SPEA2SDE to construct UATS-SPEA2SDE, enabling the algorithm to reuse historical evolutionary information while retaining the convergence and diversity advantages of shift-based density estimation. Experiments are conducted on four benchmark problems under three objective-changing settings, where UATS-SPEA2SDE is compared with a restart-based SPEA2SDE baseline and four representative dynamic multi-objective optimization algorithms. The results indicate that the archive-guided transfer improves recovery after environmental changes and enhances adaptability to objective-number variations.
cs.NE / 4 / 2609.27459
Spiking Neural Network Predicting Sequence of the External Worlds States in Model-Based Reinforcement Learning
Mikhail Kiselev
cs.NE
Abstract
This paper presents a spiking neural network (SNN) designed to predict the sequence of the external world states starting from the current world state. This SNN does not create the world dynamics model - instead it incorporates the SNN trained to predict the next world state and provides all mechanisms necessary to make the chain of predicted world states. These mechanisms are entirely spiking - they are implemented as spiking neuron ensembles. The present article describes this neuronal structure and tests its operation on a classic RL benchmark - ATARI ping-pong.
计算语言学 (cs.CL)
47
cs.CL / 1 / 2609.26976
When Learned Context Planning Fails to Beat Strong Retrieval: A Controlled Study of Planning, Routing, and Reranking for Long-Context QA
Yingrui Li, Han Chen
cs.CL · cs.IR
Abstract
Learned context planning selects evidence atoms before an answer model reasons over them. We test whether this learned selection improves long-context multiple-choice QA after strong retrieval, routing, budgeted-selector, and reranking controls. Our primary diagnostic uses all 503 LongBench-v2 MCQ questions with Qwen2.5-7B-Instruct. The planner is SFT-trained on outcome-selected traces from 140 training and 28 development questions; because the 503-question analysis includes those questions, it is partly transductive. At an 18k-character budget, anchored hybrid retrieval reaches 36.18% accuracy and BM25 reaches 35.98%, while the best direct planner-guided method reaches 34.19%. On the untouched 152-question test split, anchored hybrid remains higher (42.11% versus 36.84%). Leakage-safe routers cannot convert a large oracle gap. Under tight budgets, the best planner is ahead by only 0.40 points at 6k and loses at 9k; planner-guided reranking has a +1.79-point estimate at 6k with a paired interval crossing zero and ties the control at 9k. Packing-order and score-flatness analyses did not identify a stable mechanism. Under this setup, learned planning is a weak relevance signal rather than a replacement for strong retrieval.
cs.CL / 2 / 2609.27032
LexLattice: Multilingual Extractive Summarization via Neural Cellular Automata on Document Hierarchies
Sujay Uday Rittikar, Sheela Ramanna
cs.CL · cs.LG · cs.NE
Abstract
Faithfulness is a central concern in legal text summarization, which motivates extractive approaches that select verbatim content traceable to its source. Such methods typically rank paragraphs or other structural units in isolation, yet give little attention to consolidating evidence that is distributed across, and shares salience between, distant parts of a document. We introduce LexLattice, an extractive summarizer that reifies a legal act's hierarchy as a two-dimensional semantic lattice and consolidates over it with a masked 2D neural cellular automata before selection. LexLattice attains state-of-the-art ROUGE across all 24 languages of EUR-Lex-Sum in both multilingual and cross-lingual settings, surpassing instruction-tuned baselines with billions of parameters, despite concentrating all trainable capacity in a 1.8M parameter consolidator over a frozen multilingual encoder. A consolidator trained only on high-resource languages further transfers to unseen languages with near-lossless retention (0.99), indicating that the model operates on language-agnostic semantic geometry rather than surface form. Our results position explicit consolidation over document structure as a compact and traceable alternative to scale for multilingual legal summarization.
cs.CL / 3 / 2609.27059
The Illinois Social Attitudes Aggregate Corpus (ISAAC): An Open Tool and Reproducible Pipeline for Analyzing Social Group Discourse at Scale
Babak Hemmatian, Sarah Hadjarab, Jessica Chen, Benedek Kurdi
cs.CL · cs.SI
Abstract
We introduce the Illinois Social Attitudes Aggregate Corpus (ISAAC), an open, modular, and accessible corpus of 527 million+ English-language Reddit posts selected for relevance to six key social group distinctions based on race, sexuality, age, ability, body weight, and skin tone, covering the 17-year period from 2007 to 2023. A multi-step, human-audited filtering pipeline was used to keep irrelevant content in the curated dataset below 10%, both overall and for each social group distinction. Each post was then algorithmically annotated with the user's estimated home region, along with a suite of validated off-the-shelf and custom semantic labels including moralization, sentiment, emotion, and linguistic generalization. We confirm the validity of the resulting corpus through convergent evidence linking ISAAC to macro-level societal trends, such as online search behavior, temporal spikes during major societal events (both nationally and regionally), and long-term shifts in public attitudes. By offering a unified, public infrastructure, ISAAC eliminates research fragmentation and enables seamless replication while supporting diverse empirical workflows at scale. Specifically, ISAAC allows investigators to perform cross-category comparisons, conduct high-precision tracking of long-term temporal shifts in social group discourse, and map spatial variation onto localized public opinion and policy outcomes. ISAAC's fully public, modular pipeline facilitates easy extension of the corpus to new platforms, languages, and social categories. To accommodate various research needs, ISAAC is accessible both without coding through a point-and-click website and labeler web-apps, and programmatically via an SQL playground, a Python package, and HuggingFace.
cs.CL / 4 / 2609.27064
What Changes When Fact-Verification Scores Improve? Evidence and Answer Accounting Across Trained Verifiers and LLMs
Han Chen, Yingrui Li
cs.CL
Abstract
A joint fact-verification score assesses answers and submitted evidence together. When the score improves, how much of the gain remains if the answers are held fixed? On FEVEROUS, strict score is the percentage of claims with a correct answer and a complete annotated evidence group in the submitted evidence. Across four trained DeBERTa checkpoints and 7,890 claims, replacing DCUF evidence with UnifEE evidence raises strict score by 9.61 percentage points, compared with 1.96 percentage points in answer accuracy. The paired 95% interval for the strict-score gain is [8.77, 10.43], conditional on these checkpoints. Replacing only the evidence passed to the scorer accounts for 7.92 or 9.08 percentage points when we retain the answers generated from DCUF or UnifEE evidence, respectively. To examine how this evidence gain depends on evaluation choices, we generate 470,400 responses from two 8B LLMs on FEVER, FEVEROUS, and SciFact under two answer formats and two context budgets. Increasing context from 256 to 2,048 tokens raises the fixed-answer evidence gain on FEVEROUS by 3.84 and 3.10 percentage points for Qwen and Llama, respectively. The effects fall short of the prespecified cross-dataset criterion, while some intervals extend beyond the two-point small-effect bound. Post-hoc analyses quantify changes in answers and submitted evidence, and show when aggregate accuracy and evidence-coverage rates miss the claim-level pattern. The four answer-evidence score combinations reveal changes that endpoint and aggregate metrics leave unresolved.
cs.CL / 5 / 2609.27086
NADI 2026: The Second Multidialectal Arabic Speech Processing Shared Task
Peter Sullivan, Bashar Talafha, Ahmed Ashraf, Fethi Bougares, Haroun Elleuch, Chiyu Zhang, AbdelRahim Elmadany, Youssef Mohamed, Salima Mdhaffar, Yannick Estève, Mohamed Elhoseiny, Hamzah Luqman, Nizar Habash, Muhammad Abdul-Mageed
cs.CL
Abstract
NADI 2026 is the seventh edition of the Nuanced Arabic Dialect Identification (NADI) shared task series and the second dedicated to multidialectal Arabic speech processing. This edition comprises five tasks and eight subtasks spanning Automatic Speech Recognition (ASR), Spoken Dialect Identification (SDID), Text-to-Speech (TTS), Spoken Language Translation (SLT), and Spoken Language Understanding (SLU). NADI 2026 emphasizes realistic evaluation through low-bandwidth, mixed-dialect, code-switched, out-of-domain, and zero-shot settings, while introducing TTS, SLT, and SLU to the series for the first time. The shared task attracted 21 participating teams from at least 13 countries, with 48 test-phase submissions and 14 submitted system-description papers. Results show that out-of-domain generalization remains a major bottleneck and highlight the effectiveness of recent Arabic-specialized speech models, multimodal dialect identification approaches, and ensemble methods. Overall, NADI 2026 provides a broader and more challenging benchmark for robust Arabic dialect speech processing.
cs.CL / 6 / 2609.27156
Giving Credit Where It's Due: Redundancy-Aware Learning for Efficient Reasoning
Yuqing Zhou, Hong Wang, Manqing Mao, Zhuoer Wang, Samson Koelle, Jie Yuan, Yanjun Lin, James Feng, Nikki Lijing Kuang, Ziwei Zhu, Wei Niu
cs.CL · cs.LG
Abstract
Large reasoning models can produce correct yet unnecessarily long reasoning traces. Existing methods improve reasoning efficiency with trajectory-level objectives or local token- and step-level signals, but rarely model inter-step semantic dependencies. This limits their ability to distinguish redundant steps from those that support later deductions, making it harder to shorten reasoning without sacrificing accuracy. We introduce RECAP (REdundancy-aware Credit Assignment via Propagation), which addresses this limitation by assigning credit where it is due based on both a step's downstream role in the reasoning structure and its contribution to solving the problem correctly. We define structural responsibility to capture the step's downstream role by measuring how strongly later reasoning depends on it, using credit propagated backward from the final-answer node through an outcome-independent, LLM-annotated semantic dependency graph. However, a step can have high structural responsibility yet steer the reasoning away from the correct solution. RECAP therefore introduces step efficacy to measure answer-directed progress through changes in gold-answer log-likelihood as each step is added. Together, these signals reshape rollout-level GRPO advantages into step-specific updates. RECAP requires neither a separately trained process reward model nor preconstructed concise trajectories. Across two 7B models and four mathematical reasoning benchmarks, RECAP improves the accuracy-efficiency trade-off. On Qwen2.5-Math-7B, it improves pass@1 by 2.0-3.7 percentage points while reducing reasoning tokens by 8%-31% relative to GRPO across all four benchmarks. Analysis suggests these savings reflect fewer reasoning operations and less dead-end reasoning, rather than more compact expression.
cs.CL / 7 / 2609.27173
Realize What Matters: Principled Context Representation for Large-Scale Reasoning
Michael Theologitis, Dean Light, Shuyue Stella Li, Benjamin Newman, Yulia Tsvetkov, Dan Suciu
cs.CL
Abstract
Solving complex tasks in domains such as science, medicine, law, and finance often requires assembling interdependent information scattered across vast, heterogeneous sources far beyond model context limits. Existing approaches tackle this challenge by organizing information into more manageable representations over which models can reason, such as graphs, textual memories, and retrieval collections. These representations dictate what downstream reasoning is possible and, ultimately, whether it succeeds; yet their design and construction remain largely ad hoc. In this work, drawing on the cognitive theory of relevance realization, we propose concrete principles for designing AI systems that construct effective representations of very large contexts. We analyze existing approaches and show how their successes and failures map onto their alignment with these principles, and introduce R3Con, a harness designed to operationalize the principles more systematically. We evaluate R3Con against nine state-of-the-art baselines on two recent benchmarks of reasoning over large document corpora. On these benchmarks, R3Con substantially outperforms the strongest baseline, by $20$ and $8.4$ percentage points. It also enables smaller models to outperform much larger ones: R3Con with 4B and 9B models outperforms all evaluated 35B baselines, while R3Con with a 35B-A3B model outperforms Claude Code with Claude-Sonnet-5 at $3.7\times$ lower cost. Our results show that context representations following our principled approach can reduce reliance on model scale, pointing toward a future of AI systems with frontier-level performance powered by smaller models. Our code is available at https://github.com/michaeltheologitis/r3con
cs.CL / 8 / 2609.27176
Beyond Overlap: Estimating the Causal Effect of Benchmark Exposure
Divyansh Singh
cs.CL
Abstract
Evidence that evaluation material entered training does not reveal how much it affected evaluation. This distinction leaves a contaminated benchmark score difficult to interpret: provenance can establish contact, but only a counterfactual can quantify the performance attributable to that contact. We present LeakScale, an interventional framework for estimating this missing quantity. LeakScale creates fresh executable tasks that require private, family-specific information absent from and non-derivable from the public task, controls access to that information, and estimates the resulting control-adjusted change in executable accuracy. Across 2,048 unique families, two model families, two executable domains, and 262,144 generations, exposure improves accuracy in every model-by-domain combination, with gains ranging from +7.17 to +27.31 percentage points. These findings separate two empirical questions that are often conflated: whether benchmark contact occurred and how strongly a reported score depends on it. LeakScale makes the latter directly measurable.
cs.CL / 9 / 2609.27205
Phonemizing User-Generated Text: A Benchmark, Taxonomy, and Compositional Approach
MinJu Jeon, Younghan Park, Han Sung Park, Jong-Hwan Kim, Dong-Jin Kim, Hoyeon Lee
cs.CL · cs.AI
Abstract
Text-to-speech systems increasingly process user-generated text (UGT) such as ppl and imo, whose pronunciation must be inferred from the canonical rather than surface form. We introduce UGTPhon, the first grapheme-to-phoneme (G2P) benchmark for UGT in English, Vietnamese, and Korean, together with an inference-grounded taxonomy for fine-grained diagnosis. Existing G2P models and frontier LLMs exhibit a systematic canonical-to-non-canonical performance gap, reaching up to 66.8 PER points. As a benchmark baseline, we propose a simple compositional G2P approach that incorporates canonical-form evidence through exact-match lookup and staged decoding. Across matched ByT5 and Qwen2.5-0.5B backbones, explicit canonical-form modeling consistently reduces non-canonical G2P errors. The 0.5B variant also performs competitively with much larger few-shot frontier LLMs, highlighting the benefit of explicitly modeling canonical-form inference for UGT phonemization.
cs.CL / 10 / 2609.27233
Distilling Sequential Computation in Transformer Language Models
Zixuan Lan, Jessica Yang, Yanhong Li, Karen Livescu, Jiawei Zhou
cs.CL
Abstract
Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adjacent token spans are highly predictable or frequently occur as stable units, suggesting that their representations may be compressible. We introduce a method for distilling sequential computation by replacing spans of input tokens with collapsed representations, computed on the fly by a lightweight merge module. This module generates a single surrogate embedding from a sequence of static token embeddings that captures the functional role of the multiple tokens, allowing pretrained models to operate on compressed inputs without architectural changes or re-training. We apply this approach during inference to compress both prompts and intermediate decoding steps, using a rollback mechanism to substitute stored multi-token KV cache entries with their single-step surrogates. Experiments across diverse models show that the merge module can be used to reduce effective sequence length by up to 40% with minimal accuracy degradation across language modeling evaluations and downstream tasks, including question answering, summarization, commonsense reasoning, and long-form mathematical reasoning. Additional lightweight adaptation of the merge module further improves the accuracy-compression trade-off in selected settings. These results demonstrate that sequential token computation in Transformers can be effectively approximated through condensed surrogate representations that approximate the original behavior without model updating.
cs.CL / 11 / 2609.27257
UniDataAgent: An Ontology-Grounded Agent for Enterprise Question-to-Report Automation
Yutai Duan, Yahui Zhao, Zhangti Li, Yu Ma, Zhenfeng Qi, Shaoyang Yuan, Jing Fan, Jie Liu
cs.CL
Abstract
Enterprise data agents must preserve organization specific semantics, not just translate questions into queries. We present ChinaUnicom DataAgent (UniDataAgent), an ontology grounded system for reusable question-to-report analysis that separates semantic acquisition from online execution. Ontology Acquisition and Validation stage (OAV) builds versioned enterprise ontologies from metadata, business knowledge, and supporting materials through expert authored business skills, constrained generation, question verification, and selected expert review. Question-to-Report Execution (QRE) stage retrieves semantic contracts for each question, coordinates skills and data tools, validates results, and produces evidence linked reports. Across 27 enterprise tables and roughly thousands of metric types, ontology construction took a few hours instead of about one week manually. It took just a few minutes to generate the reports, instead of several working days. Ontology grounding achieved 95.0\% strict accuracy on real business questions, versus 72.5\% for document RAG, especially on structured and compositional tasks. The system has already been deployed to generate cost savings and has the potential to be replicated in other enterprises.
cs.CL / 12 / 2609.27262
Can One Adapted Model Do It All? Fine-Tuning Strategy Selection for Customer Support LLMs
Md Tahmid Rahman Laskar, Xue-Yong Fu, Shashi Bhushan TN
cs.CL
Abstract
Production customer-support systems often require LLMs to support multiple skills, such as intent classification, question answering, summarization, or tool-use decisions. A central deployment question is whether these skills should be handled by separate task-specialist models or by a single model trained through multi-task training, sequential updates, or model merging. We study this question using thirteen models spanning five families (Qwen3, Qwen3.5, Gemma-3, Llama-3.1, and Mistral) from 0.6B to 32B parameters across eight customer-support datasets, spanning four public and four proprietary datasets with approximately 74.5k training and 8.7k evaluation samples. Under a fixed training protocol, we train more than 200 checkpoints. Our experiments reveal that multi-task full fine-tuning is the strongest operational default at every model size we test. Specialist models are strong on their target tasks but often degrade sharply off-task, making reliable routing important. Sequential Low-Rank Adaptation (LoRA) preserves earlier skills better than sequential full fine-tuning, while merging a specialist with its base model improves off-task robustness with limited same-task loss for larger models. We conclude with practical guidelines for selecting fine-tuning strategies in real-world settings.
cs.CL / 13 / 2609.27289
Ruby-ASR: Evidence-Preserving Supervision for Joint Orthographic and Lexical-Reading Recognition
Hao Shi, Yun Liu, Xuehao Yang, Jun Liu, Chuanbo Hua, Xuanjun Chen, Lianbo Liu, Shiao Zhu, Zixiong Su
cs.CL · cs.AI
Abstract
Conventional Japanese automatic speech recognition (ASR) is supervised by an orthographic transcript, although the same written form can correspond to different lexical readings realized in speech. Such utterances receive an identical target, so their reading distinction is absent from the supervision interface and cannot be recovered reliably by post-hoc text-only grapheme-to-phoneme conversion. We present Ruby-ASR, which refines the conventional target into a span-bound orthographic--lexical-reading sequence. Unlike separate full-sentence orthographic and phonological outputs, the ruby representation locally binds each written span to its realized reading and permits deterministic recovery of both views. We instantiate the target under subtitle-style and verbatim-style transcription conventions using a Qwen3-ASR backbone; a mora-level CTC objective provides auxiliary monotonic reading supervision. The experimental results across five Japanese benchmarks show that refining the recognition target can improve lexical-reading recovery without sacrificing readable orthographic transcription. We release the checkpoints and inference code.
cs.CL / 14 / 2609.27353
Guides That Cause Actions: An Offline Study of Guide-Action Mutual Reinforcement in Multimodal Web Agents
Chengguang Gan, Yunhao Liang, QingHao Zhang, Shiwen Ni
cs.CL
Abstract
Web agents are usually evaluated in live environments, where environment state and judge models drift between runs, so the same checkpoint rarely reproduces the same score, making controlled studies of training phenomena impractical. We present WebMRE, an offline benchmark of 541 tasks and 5,293 steps derived from successful WebArena trajectories, with fully audited test labels and a deterministic protocol that scores a checkpoint identically on every run without any environment. Each step pairs a human oriented guide sentence with a grounded action, enabling the first study of the mutual reinforcement effect between them in web agents. Averaged over three seeds the effect holds for both models in both decoding orders and grows with scale: jointly decoding a guide lifts element selection over an action only reference by 0.9 and 0.2 points for Qwen3.5-4B and by 1.7 and 2.2 points for Qwen3.5-9B. A mediation analysis shows that the guide is a causal channel rather than commentary: forcing the gold guide as a decoding prefix lifts action accuracy from .422 to .684, another step's guide collapses it to .055, and a paraphrase that renames the target still recovers half of the gain, so the channel carries instruction meaning and not only the label string. The same channel yields an offline reward that only a replayable protocol makes computable, though optimizing it from a strong checkpoint brings no gain yet. Our fine tuned models outperform GPT-5.5, Claude Opus 4.8, and Gemini 3.5 Flash, run zero shot, on every offline metric.
cs.CL / 15 / 2609.27372
Neither Silence nor Overlap Is Failure: Intent-Conditioned Evaluation of Turn-Taking in Full-Duplex Spoken Dialogue Models
Kian Shamsaie, Iman Modarressi
cs.CL · cs.AI · cs.HC · cs.SD
Abstract
Benchmarks for full-duplex spoken dialogue models score turn-taking with binary fixed-window rules that reward immediate response or silence by completeness of the prior turn. We argue that the appropriateness of a response offset, whether delayed silence or anticipatory overlap, is conditional on the speaker's latent intent, identifiable only from that speaker's behavior. We introduce TACT, a benchmark of 9,728 episodes and 73.2 hours from five dyadic corpora; each episode carries dialogue history, a per-speaker memory profile, and an annotator-derived posterior over six intent classes. Scoring replaces binary windows with a strictly proper threshold-weighted continuous ranked probability score whose weights are intent-conditioned timing kernels fitted to human floor-transfer-offset distributions, proving boundedness, consistency, and binary reduction. Across eleven systems the best model reaches 0.47 against a human topline of 0.86, is nearly invariant to speaker profiles, and TACT agrees with human judgments at Spearman 0.81 versus 0.46 for binary metrics.
cs.CL / 16 / 2609.27373
Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models
Ke Wan, Chen Chen
cs.CL · cs.LG
Abstract
Recurrent language models repeatedly apply shared network blocks to refine latent representations, but standard inference recomputes global attention at every recurrent step. We study attention dynamics across recurrent depth and find that attention support and distributions stabilize substantially earlier than hidden states and attention outputs. This suggests a two-stage structure: early steps discover a sparse working set of relevant context, while later steps refine representations over largely the same routing support. Motivated by this structure, we introduce WISE (Working-set Inference with Support Exploitation), a training-free method that uses unrestricted global attention during early recurrence and later reuses directly discovered block-structured support while keeping recurrent depth and within-support attention computation dynamic. Controlled interventions show that recurrent discovery is important and that support-only reuse better preserves model behavior than more restrictive attention-reuse alternatives. Across multi-hop QA benchmarks, WISE largely preserves full-attention performance, while context scaling reveals increasingly sparse working sets and greater efficiency gains. Quality is largely preserved through 2K context, with a measurable loss at 4K. An optimized sparse-attention implementation achieves up to a 1.76x attention speedup over native FlashAttention at 4K and a 1.36x speedup for the full 32-step attention trajectory. Our code is available at https://github.com/tbn5pj/WISE_code.
cs.CL / 17 / 2609.27374
Planned Test-Time Scaling with Coordinated Reasoning Paths
Xueqing Wu, Langxing Bai, Hritik Bansal, Po-Nien Kung, Shuo Li, Hao Liu, Nanyun Peng, Kai-Wei Chang
cs.CL · cs.AI
Abstract
Test-time scaling with parallel branches is widely adopted to improve performance on challenging reasoning tasks. The predominant approach, repeated sampling, draws branches independently from a single policy, which can produce redundant attempts and thereby limit the gains from additional inference compute. To address this limitation, we propose Planned Test-Time Scaling (PTTS), which replaces independent sampling with a coordinated joint policy: a planner generates a solution outline for each branch, steering the branches toward distinct reasoning paths, and an executor produces a full solution conditioned on each outline. Formally, we show that PTTS strictly generalizes repeated sampling and, in a stylized setting, provably promotes coverage of complementary reasoning modes and yields better pass@k scaling. We instantiate PTTS on top of strong reasoning models, keeping them fixed as executors while replacing repeated sampling with PTTS inference to further enhance test-time scaling. Concretely, we develop two variants: PTTS-ZS prompts a model to jointly generate outlines for all branches in a single autoregressive pass, while PTTS-RL directly optimizes the planner against the pass@k reward using truncated execution rollouts for efficient training and a sharper reward signal. Across five mathematical reasoning benchmarks with Qwen3-1.7B and 4B, PTTS-ZS improves pass@64 over repeated sampling by up to 6.7 points, while PTTS-RL further increases the gain to up to 13.4 points. Further analysis indicates that broader coverage of distinct reasoning paths contributes to these gains. Overall, PTTS provides a general framework for improving test-time scaling by coordinating reasoning branches, with zero-shot and trainable instantiations that yield substantial performance gains.
cs.CL / 18 / 2609.27376
Cross-Lingual Legal QA for Vietnamese Labour Law: Retrieval, Translation, and Verifier-Guided Correction
Nguyen Minh Chi, Mo El-Haj, Nguyen Ha Thanh, Dawn Knight, Paul Rayson
cs.CL
Abstract
Cross-lingual legal question answering must retrieve statutes across languages while preventing unsupported legal claims. We introduce a bilingual evaluation suite of 231 Vietnamese--English question--answer pairs from Vietnamese labour law. Of these, 75 are additionally annotated for five challenging legal reasoning phenomena. We evaluate a verifier-guided pipeline that decomposes answers into claims, checks citation reachability and entailment, and corrects citation failures and contradictions. We also introduce six automatic diagnostics for faithfulness to retrieved evidence, covering citations, modality, exceptions, procedures, conclusions, and evidential support. Experiments show that learned-sparse retrieval performs poorly for English-to-Vietnamese retrieval (R@5~=~0.032), whereas dense retrieval reaches 0.358 and slightly outperforms hybrid retrieval. Translation placement has no statistically detectable effect on these automatic diagnostics in our controlled comparison and supporting sensitivity analyses. Verifier-guided correction improves citation preservation by $0.022$--$0.034$ at the system level but produces no reliable gains in the remaining dimensions. Human evaluation further shows that the automatic diagnostics do not fully align with human judgements of answer quality.
cs.CL / 19 / 2609.27380
MORSE: Multi-Context Ordering via Reverse Scoring for Evidence-Preserving Compression
Ke Wan, Yifan Wang, Liheng Lai, Chen Chen
cs.CL
Abstract
Likelihood-based context compression can account for cross-context redundancy through sequential scoring, but this makes compression outcomes sensitive to context order. We show that different permutations of the same context collection can produce markedly different evidence-retention outcomes under an unchanged compressor. We attribute this sensitivity to information preemption: earlier partially relevant contexts can absorb credit for shared information, suppressing the incremental score of later, stronger evidence carriers and increasing their risk of removal. Controlled pair-swap interventions directly support this mechanism by showing that evidence-first ordering substantially improves supporting-evidence survival. To address this problem, we introduce MORSE, a compression-aware method for evidence-preserving context ordering. MORSE applies a common reverse query-evidence principle to both individual contexts and compressed candidate outputs, using the former to construct an evidence-first anchor and the latter to guide compression-aware permutation selection. Across multi-hop QA benchmarks, compression procedures, budgets, and scoring models, MORSE consistently improves evidence preservation over static reverse ordering and compute-matched random search, with corresponding overall improvements in downstream QA. Our code is available at https://github.com/tbn5pj/MORSE_code.
cs.CL / 20 / 2609.27387
AraGenre 2026: A Hierarchical Definition-Guided Arabic Genre Classification Shared Task
Mo El-Haj, Saad Ezzini, Shadi Abudalfa, Mustafa Jarrar, Nguyen Minh Chi, Nguyen Minh Quan
cs.CL
Abstract
AraGenre is a shared task on hierarchical, definition-guided Arabic genre classification, motivated by the limited availability of annotated data in Arabic and other low-resource languages. Systems assign each Arabic text segment both a broad communicative genre and a fine-grained specific genre. The released training and development sets contain limited, primarily synthetic and controlled examples, whereas the hidden final benchmark contains noisier naturally occurring text spanning Modern Standard Arabic, Classical Arabic, and multiple dialects. Participants received natural-language definitions for 74 previously unseen specific genres, creating a zero-shot label generalisation setting in which systems had to infer class semantics rather than memorise fixed label-feature associations. The task attracted 46 registrations and 373 submissions, with 17 teams completing the final evaluation. Thakaa ranked first with a Hierarchical Macro F1 of 0.7352, followed by HoangPhong (HP) with 0.7169 and NAMAA with 0.7013. The results show strong broad-genre recognition but a substantial gap in fine-grained classification under linguistic and domain variation.
cs.CL / 21 / 2609.27395
PRISM-VLM: A Multi-Axis Discriminative Benchmark for Compact Vision-Language Models
Sanghee Park, Kee-Eung Kim
cs.CL
Abstract
Compact vision-language models (VLMs) now power a growing share of multimodal applications. The benchmarks used to compare them, however, inherit a frontier-centric design: each model is reduced to a single accuracy number, narrowing the inter-model gap on saturated suites and pressing models into low-score bands on harder ones. We introduce PRISM-VLM, a multi-axis discriminative benchmark that scores every item along seven axes covering the recurring failure modes (task quality, behavioral robustness, and capability bottlenecks) and combines them into a single PScore, with items recycled from fifteen public benchmarks. Across compact VLMs from the past two years, PScore separates model pairs more reliably than prior single-axis benchmarks under an item-level paired bootstrap, and surfaces behavioral differences these benchmarks average away. Even models with statistically indistinguishable PScores diverge sharply along the per-axis profile, particularly on sycophancy, which is nearly orthogonal to single-prompt accuracy. We will release the full pipeline, prompts, and per-item annotations.
cs.CL / 22 / 2609.27558
ThaiTrees: Thai Syntactic Dependency Trees Across Domains
Attapol T. Rutherford, Papatchol Thientong
cs.CL
Abstract
Studying syntactic patterns in naturally occurring language requires a large parsed corpus, but manual annotation is costly and difficult to scale. Thai has a manually annotated dependency treebank for training and evaluating parsers, but lacks a large automatically parsed corpus for quantitative syntactic research. We present ThaiTrees, a 342M-token corpus drawn from news, Wikipedia, spoken transcripts, and social media. We develop a reproducible pipeline for cleaning, processing, and parsing Thai text under the Universal Dependencies framework. The resulting corpus makes grammatical relations searchable and supports the study of syntactic distributions. We release a frequency lexicon and CoNLL-U parses in machine-readable formats suitable for both AI-assisted and conventional programmatic analysis.
cs.CL / 23 / 2609.27590
MWE-ECL: Recoverable Long-Range Context Does Not Always Override Local Lexical Priors
Wei He, Aline Villavicencio, Rodrigo Wilkens, Zhenyun Deng
cs.CL
Abstract
Long-context evaluations often test whether a model can recover distant evidence, but recoverability does not guarantee behavioral influence. We test the prediction that a distant discourse anchor can remain explicitly recoverable yet fail to change the locally preferred reading of a familiar multiword expression; such failures should concentrate when the model's no-anchor default conflicts with the anchor, while prior-correct decisions remain largely preserved. We introduce Multiword Expression Effective Context Length (MWE-ECL), a bilingual diagnostic whose matched anchor-retrieval, no-anchor prior, and interpretation prompts measure explicit recoverability, model-observed defaults, and anchor-conditioned decisions, respectively. Across eight English deployment panels on a shared 0-128K grid, retrieval-control accuracy on prior-conflict items is 0.989-1.000, prior-conflict override spans 0.806-1.000 (0.809-1.000 after conditioning on correct retrieval), and preservation of prior-correct decisions remains 0.977-1.000. A same-call control querying retrieval and interpretation in one prompt reproduces the gap for DeepSeek V4 Pro (1.000 retrieval versus 0.900-0.920 interpretation), showing that separate invocations are not its sole explanation; smaller or absent gaps in the other two models bound its generality. For DeepSeek V4 Flash, separate prompt-fit tests retain perfect retrieval with lower interpretation at 512K and 1M, while foil-consistent cues shift the no-anchor prior far more than retrieval; cross-model cue effects are heterogeneous. A separately reported 10-family Chinese subset shows similar descriptive gaps, but imperfect retrieval for some models prevents an integration-only attribution. MWE-ECL therefore evaluates whether explicitly recoverable distant context changes a competing local semantic decision.
cs.CL / 24 / 2609.27607
Can Jev Judge Radiology Reports? Evaluating a System One Model for Clinical Factuality
Jiaju Huang, Hao Yang, Xinyu Ma, Xinglong Liang, Kunyan Cai, Junqiang Ma, Shaobin Chen, Yue Sun, Tao Tan
cs.CL · cs.AI
Abstract
An AI-generated radiology report can resemble a physician's report while omitting an abnormality, adding an unsupported finding, or reversing its presence. Measuring these factual differences is essential for evaluating report generators. We study Jev, a System One decision model, as a simple, low-cost judge of agreement with physician-written reference reports. Our evaluator checks whether each statement is supported by the other report and combines these judgments in both directions to capture unsupported claims and omissions. A single-question configuration reaches Kendall correlations of 0.573 on RadEvalX and 0.398 on RadEvalExpert with expert error counts, outperforming an open natural language inference judge under matched decomposition and aggregation. One support question per statement retains similar expert agreement to seven while using 43-45% fewer judgment input tokens. At the documented API price, judgments cost under three cents per hundred report pairs, excluding local decomposition. In a separate controlled-error test, Jev detects false negation with an AUROC of 0.977. Local RadMatch achieves stronger agreement on clinically significant errors in both expert datasets and on total errors in the shared RadEvalExpert subset. Finding-count and error-scope analyses show that benchmark agreement reflects report size and error definitions as well as medical error detection. These results support Jev as a practical judgment component for measuring factual differences in generated radiology reports and identify where more elaborate evaluation remains valuable.
cs.CL / 25 / 2609.27650
Brain-to-Language Decoding: Tasks, Signals, Methods, Evaluation, Practical Use and Beyond
Yiqian Yang, Yiqun Duan, Chenyu Liu, Yiqi Wang, Xinliang Zhou, Chin-Teng Lin, Yu Zhang
cs.CL · cs.HC · cs.NE
Abstract
Brain-to-language decoding translates neural activity associated with language production, internal speech and perception into linguistic or expressive outputs. It offers a route to restoring communication after speech loss and a means of studying how the brain represents language. Advances in neural recording and representation learning have expanded the field from constrained recognition and acoustic reconstruction to text generation, streaming personalised speech and facial animation. This survey synthesises these developments across invasive and non-invasive measurements, drawing on a search without a lower year limit and source-led updates through September 2026. We connect Articulated, Inner and Perceived tasks to the neural populations they engage, the representations available to decoders and the outputs those representations can support. We examine model development, public resources and the evolution of evaluation, and compare published performance and communication costs within their reported protocols. The synthesis identifies complementary routes to progress: phonetic, acoustic and semantic targets preserve different aspects of a message; shared representations support reuse across recording conditions and tasks; and online communication increasingly depends on calibration, feedback and user control alongside decoding accuracy. Shared benchmarks enable algorithmic comparisons, while longitudinal studies reveal the demands of sustained use. We discuss these developments and their remaining limitations, then outline a prospective five-level trajectory from commands and language to meaning, scenarios and bidirectional cognitive exchange
cs.CL / 26 / 2609.27669
The Path Matters: Evaluating Small Language Models Beyond Answer Accuracy in KGQA
Eduin E. Hernandez, Sergio A. Diaz, Luis F. Garcia, Nurassyl Askar, Stefano Rini
cs.CL · cs.AI
Abstract
Small language models (SLMs) are increasingly paired with knowledge graphs (KGs), yet end-to-end KG question answering conflates graph access, search, navigation, reasoning, and answer generation. This coupling makes it difficult both to determine whether an SLM can faithfully execute the reasoning path implied by a question and to attribute failures to navigation rather than to other stages of the pipeline. We isolate this capability by employing the THESEUS navigation and traceability framework and using frozen, off-the-shelf SLMs as local action policies. At each hop, the environment exposes the legal outgoing graph actions, and the model selects one executable graph action and decides whether to stop, without task-specific parameter updates, model-controlled beam search, or free-form answer generation. This controlled setting allows us to evaluate terminal-answer accuracy with Hits@1 together with path fidelity, using Path Edit Distance (PED) as the primary trajectory metric. Across the Kinship and MQuAKE-ST KGQAs, similarly sized local models differ substantially in answer accuracy and path fidelity, with the two metrics sometimes favoring different models. This model-dependent behavior also extends to prompting, as a single demonstrated trajectory can improve or degrade navigation depending on the model. These results motivate evaluating SLM graph reasoning beyond endpoint accuracy alone.
cs.CL / 27 / 2609.27678
Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding
Fan Zhang, Yankai Chen, Zhuohan Xie, Yixi Zhou, Sijia Peng, Lei Fan, Xinhua Ji, Cunyuan Zheng, Huangyong Shan, Philip S. Yu, Xue Liu, Yu Chen, Preslav Nakov, Songwei He
cs.CL
Abstract
Contract inference requires multiple judgments about a shared document, but aggregate accuracy can conceal changes in the individual decisions. Repeated agreement is also insufficient: a model may consistently return the wrong answer. In this paper, we compare Jev with nine language models on ContractNLI, evaluating inference cost, response time, average correctness, and correctness across repeated request conditions. Controlled comparisons vary hypothesis visibility, requested outputs, and output order while keeping the contract and target judgment fixed. Jev has the lowest cost and median response time among the evaluated configurations, while hosted language models achieve higher baseline accuracy. Rankings by baseline accuracy differ from rankings by correctness across every condition and repeat, although small differences in the latter do not establish a general stability advantage. Development diagnostics further reveal compensating corrections and regressions, as well as persistent errors. These findings motivate evaluating cost and response time alongside whether individual judgments remain correct as the request configuration changes. Code: https://github.com/ZF-Utokyo/Jev-Benchmark
cs.CL / 28 / 2609.27980
Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery
Rasmus Aagaard, Nicki Skafte Detlefsen
cs.CL · cs.LG
Abstract
Pruning large pre-trained transformer-based ASR models such as OpenAI's Whisper has seen great adoption, as pruning the decoder led to significant end-to-end transcription speedups. For instance, the {\tt whisper-large-v3-turbo} variant reduced the decoder from 32 to 4 layers, while Distill-Whisper similarly reduced the decoder to only 2 layers. Although some attention has been put towards reducing the size of the encoder, no approach has seen wide adoption. This could be due to the need for custom inference implementations to take advantage of the compressed model. We present an approach that ranks encoder layers by the leave-one-layer-out change in Word Error Rate (WER). The six layers that cause the least change are removed, corresponding to $18.5\%$ of the encoder stack. The pruned model requires no custom inference code as it is simply a more shallow encoder with fewer layers. We further distill using unlabeled monolingual speech data to recover performance degradation caused by the zero-shot layer pruning. Mean WER across four languages increases to $20.1\%$ after distillation, compared to $21.9\%$ zero-shot, going from a baseline of $18.2\%$. We release all of our code (https://github.com/rasgaard/whisper-encoder-layer-prune) and the pruned model (https://huggingface.co/rasgaard/whisper-large-v3-turbo-encoder-pruned).
cs.CL / 29 / 2609.27981
Risk-Controlled KV-Cache Eviction: From Memory Budgets to Risk Targets
Beomgu Kang, SoJin Yun, Hojoon Kim, Hyunseok Seo
cs.CL · cs.LG
Abstract
KV-cache eviction is typically evaluated through average quality-memory trade-offs, yet a small average loss can hide requests whose utility degrades materially. We reformulate eviction as a deployment risk-control problem: a material degradation occurs when eviction lowers task utility by more than a deployment-specified tolerance relative to full-KV inference on the same request, and deployment risk is the population frequency of such events. Given a reliability contract specifying a target risk level and confidence requirement, we use a compressor-agnostic post-hoc certification procedure to select a retention policy from calibration data with a finite-sample guarantee, falling back to full KV when no compressed policy is certified. Across multiple eviction methods, Llama and Mistral models, and LongBench and RULER-32K, the same contract supports substantially different levels of eviction: on Llama, it certifies SnapKV at 75% retention on LongBench but no tested compressed policy on RULER-32K, triggering full-KV fallback. Policies with empirical degradation rates below the 5% target can still fail finite-sample certification; on Llama LongBench, empirical thresholding selects uncertified policies that retain 5-10 percentage points less cache across fixed-budget methods. The proposed framework converts a deployment-level reliability requirement into a KV-memory operating point.
cs.CL / 30 / 2609.28004
Controlled Attribute-Specific Summarization of Interrogative Dialogues
A Aditya Bhardwaj, Arjit Singh Arora, Md Shad Akhtar
cs.CL · cs.AI
Abstract
Effective summarization of interrogative dialogues is a critical task in forensic and investigative settings, requiring high factual accuracy, coherence, and attribute-specific relevance. In this work, we introduce CASPER, a novel Chain-of-Thought Attribute-Specific Prompting for Evaluative Summarization framework that leverages structured prompting and iterative refinement to generate high-quality summaries of interrogator-witness interactions. We construct MINDSum, a dataset extending the MIND corpus, comprising 6,000 utterance pairs annotated with event details, factual statements, character descriptions, and fillers. CASPER employs RoleEval, a hierarchical evaluation mechanism where multiple roles (officer, inspector, senior inspector) iteratively assess summaries based on predefined criteria. By integrating entity extraction and structured feedback loops, CASPER significantly improves factual consistency and contextual completeness compared to existing baselines. Experimental results demonstrate that our framework outperforms standard summarization models on both lexical (ROUGE) and semantic (BERTScore) metrics, while human evaluation confirms its alignment with expert reasoning. Our findings underscore the potential of controlled summarization in high-stakes domains, paving the way for AI-driven forensic intelligence.
cs.CL / 31 / 2609.28048
TEMPS: Temporal Sentence Embeddings for Temporal Information Retrieval
Mourad Hassani, Julien Romero, Amel Bouzeghoub, Christian Jacquelinet
cs.CL · cs.AI
Abstract
Modern information retrieval (IR) systems rarely represent time, yet many information needs depend on it: in clinical, journalistic, and legal search, when an event occurred can decide whether a document is relevant. Dense retrievers and Retrieval-Augmented Generation (RAG) pipelines match queries to documents well on topic but poorly on time, so they surface content that is on-topic yet temporally wrong. We introduce Temporal Textual Similarity (TTS), a task that measures how well two anchored texts align in time, independent of their topical similarity. We then present TEMPS (Temporal Embedding Model for Precise Search), a modular temporal branch that attaches to a frozen semantic retriever and trains on that signal. It resolves anchored temporal expressions to intervals and moment-matches each one to a Gaussian; the resulting ordering supervises an anchor-date-conditioned encoder, whose score we fuse with the semantic score at inference. Grounding supplies the supervision, so training uses no hand-labeled temporal data. The temporal score itself is the Gaussian-KL inclusion measure from distributional embeddings; what TEMPS adds is the grounding and the moment-matched supervision. On three temporal benchmarks, TEMPS improves MRR for every semantic backbone tested and, on TS- Retriever, lifts R@1 from 19.92 to 25.39 over the prior temporal state of the art.
cs.CL / 32 / 2609.28060
A Native-Reference Coordinate Geometry for L2 Pronunciation Deviation Using Self-Supervised Speech Models
Tina Raissi, Nhan Phan, Mikko Kurimo
cs.CL
Abstract
Self-supervised speech models encode rich phonetic information, but it remains unclear how to transform this information into interpretable metrics for second-language (L2) pronunciation assessment in spontaneous speech. We propose a native-reference coordinate geometry in which phone-class averages from native speech define a low-dimensional reference subspace, and L2 speech is evaluated by its distance to matching native phone-class coordinates. Unlike prior distance-based approaches, our method does not require parallel recordings with matched linguistic content or dedicated pronunciation labels. Across different self-supervised encoders and modeling choices, the resulting native-reference distances show negative Spearman correlations up to -0.5 with speaking proficiency, indicating that higher-proficiency speakers tend to lie closer to the native-reference space.
cs.CL / 33 / 2609.28080
Reference-Based Analysis of Coherence and Diversity in Open-Ended Text Generation
Esteban Garcés Arias
cs.CL
Abstract
Evaluating open-ended text generation involves understanding how different properties of a continuation relate to its perceived quality. We present a reference-based framework for examining coherence and diversity through three perspectives: aligning their evolution with human trajectories, comparing their summaries with a human continuation of the same prompt, and estimating their likelihood under a human reference distribution. Experiments with human quality ratings suggest that diversity-based alignment and mean-based comparisons capture quality-related variation, although the comparisons do not establish a predictive advantage for temporal alignment over simpler baselines. Reference likelihood also shows positive associations with ratings, with results varying across reference configurations and scoring horizons. Together, these analyses provide a structured way to examine how measured coherence and diversity relate to human judgments, while distinguishing similarity to human references from quality itself. Code and analysis resources are available at https://github.com/EstebanGarces/likely_human.
cs.CL / 34 / 2609.28090
Can LLMs Catch a Rigged Backtest? A Clean-Control Calibration Benchmark
Makar Ulesov, Vladislav Smirnov, Omar Ibrahim, Arsenii Bobovnikov
cs.CL · cs.AI · cs.CE · cs.SE
Abstract
Backtest auditing is a calibration problem: high flaw recall is not useful when the model falsely flags matched clean strategies. We build a 96-item paired benchmark in which every flawed backtest has a clean control that holds strategy, dates, code style, labels, and reporting scaffold fixed while changing one methodology detail. A deterministic scorer separates flaw recall, clean-control false positives, evidence localization, and fix relevance. Over 1440 cached audits from four text endpoints, the primary DeepSeek auditor reaches 100.0\% closed and clean-aware code recall, but open prompts over-flag 93.8\% of clean code controls, and clean-aware all-three specificity is 87.5\% even where recall saturates. A clean-aware warning drops DeepSeek code false positives from 20.8\% (95\% CI 11.7--34.3) to 0.0\% (0.0--7.4) at unchanged recall, while the budget anchor still flags 38/48 clean controls under the same prompt. Reporting recall alone would rank three of these four models identically; reporting the clean-control rate separates them by 79 points.
cs.CL / 35 / 2609.28117
Scaling Attention Head Analysis via Gradient-Based Attribution in Context-Aware Machine Translation
Paweł Mąka, Yusuf Can Semerci, Jan Scholtes, Gerasimos Spanakis
cs.CL · cs.AI
Abstract
In this paper, we introduce a gradient-based head attribution strategy where the Token-level Max-Margin loss is backpropagated to the attention maps. This framework enables a large-scale causal analysis of attention heads, making it suitable for LLMs. We evaluate our method on the task of disambiguation in Context-aware Machine Translation, where we analyze 50 phenomena across 4 models and 4 language directions. We empirically show the alignment of our method with the effects of increasing the attention scores of token-to-token relations on three models and two language directions, ensuring the robustness of our method. Our analysis reveals the presence of the "general-purpose" attention heads that improve the model's performance when attending to different relations. We find that the average attention a head assigns to a relation does not necessarily relate to the model's performance, which suggests that the models developed redundancies during training in terms of the head functions.
cs.CL / 36 / 2609.28270
Predicting Quantization Price for Selecting PTQ Configurations Before Deployment
Junbin Qiu, Jian Mu, Weitong Zhang, Yao Shu
cs.CL · cs.LG
Abstract
Weight-space post-training quantization (PTQ) must choose finite formats, granularities, quantizer families, transformations, and bits before the completed quantized model reveals its output-distribution drift. Existing PTQ methods predict important pieces of this degradation, including reconstruction error, Hessian sensitivity, transformation effects, and downstream loss, but these pieces are usually scored after fixing the quantization geometry or inside separate configuration families. We formulate weight-space PTQ as pre-deployment configuration selection using priced layer-output error. Each admissible layer configuration is treated as an error generator with a deployment cost, which induces a layer-output error covariance $\boldsymbolΣ_l(α_l)$, and the full-precision model prices that covariance by downstream curvature, $\widehatρ_l(α_l)=\frac{1}{2}\operatorname{Tr}\left(\widehat{\mathbf{H}}_l\,\widehat{\boldsymbolΣ}_l(α_l)\right)$. The price follows from full-precision-to-quantized forward KL, whose first-order term cancels at the reference model. It turns reconstruction and diagonal scores into reduced proxies that drop price factors, while finite formats, codebooks, granularities, and equivalent transformations become comparable candidates through the covariances they induce and the costs they pay. A trace reduction then yields a calibration-time price table and a budgeted price-guided selector, making fixed-geometry bit allocation a special case rather than the organizing problem.
cs.CL / 37 / 2609.28290
Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Jiaqi Deng
cs.CL
Abstract
Meaning identity (whether two sentences say the same thing after wording changes) is treated in retrieval and RAG as a geometric fact about independently encoded sentence vectors. We show that, for frozen off-the-shelf encoders and language models, it is not: identity is computed when both sentences share one forward pass, and is not a property of the embedding geometry those systems ship. On overlap-matched PAWS-X, purpose-built encoders (BGE, E5, GTE, MiniLM, E5-Mistral-7B) reach English confirm AUC only 0.55-0.65 (dense peak 0.70). Independently encoded last-token states of Llama 3, Mistral, and Qwen do no better; late fusion of the two vectors stays near chance. The same probe on a joint forward pass reaches 0.90-0.96 from 1.5B to 32B, collapses under partner shuffle, is mid-depth, saturates near 0.94 by 3B, and appears more weakly in GPT-2 XL (0.76). The gap holds beyond Llama-style models on other causal LMs, bidirectional encoders (DeBERTa, RoBERTa), and encoder-decoders (Flan-T5, T5, BART). Fixed or linear readers over frozen independent encodings never unlock identity; nonlinear pair readers recover part of it only on the full 49k-pair PAWS train split (0.68-0.87). Off-the-shelf rerankers split: BGE-reranker-large reaches 0.94, while MS-MARCO and Jina stay at 0.55-0.64. Independently trained families compute the same relation and a 1.5B joint reader can distill it from unlabelled teacher scores, while no linear function of the teachers own independent vectors can. Bi-encoders can be fine-tuned to fit PAWS (0.87-0.93), but transfer and STS-B suffer. Cosine compares wording neighbourhoods; identity is a cheap computed operator, not a property of either sentence vector.
cs.CL / 38 / 2609.28352
Digital diglossia: Arabic between X and Facebook
Fahad Al Hussen, Mohammed Q. Shormani
cs.CL
Abstract
This study highlights the distribution of Standard Arabic (SA; H(igh) variety) and Colloquial Arabic (CA; L(ow) variety) across X and Facebook. 16754 public posts were collected via Python, with 10000 retained as the net dataset. Posts were classified into 7 discourse categories: *politics, technology, science, business, culture, fun,* and *sports*. Bivariate analyses, including Chi-square tests and Cramer's V (CV), examined associations among platform, discourse category, and diglossic choice, while binary logistic regression with Platform x Discourse Category interactions tested whether these associations varied across platforms. Findings reveal that there are significant associations between discourse category and diglossic choice on X, chi-square(6, *N* = 5000) = 600.35, p < .001, CV = .347, and Facebook, chi-square(6, N = 5000) = 1249.52, p < .001, CV = .500. Across platforms, platform was also associated with diglossic choice, chi-square(1, N = 10000) = 262.16, p < .001, CV = .162. Binary logistic regression further shows higher odds of SA use on X than Facebook in the political reference category (*OR* = 1.31, p = .0028), with significant platform-by-domain interactions for Culture (OR = 2.65), Fun (*OR* = 6.34), Sports (*OR* = 26.71), Science (OR = 0.41), and Technology (OR = 0.71). The study concludes that the diglossic use of SA and CA contributes to the growing body of research on digital discourse, unveiling that the digital age reshapes but does not erode diglossic boundaries, giving rise instead to a reconfigured digital diglossia.
cs.CL / 39 / 2609.28430
Cross-Scale Transfer Learning for Depression Severity Prediction: From PHQ-8 to HAMD-17 Across Languages and Clinical Paradigms
Wenjie Feng, Sahba Zojaji, Satoshi Nakamura
cs.CL
Abstract
This work addresses continuous depression-severity score prediction from clinical interview transcripts under data scarcity. We propose a sequential low-rank adaptation (LoRA) protocol for cross-scale transfer: a Qwen3 backbone with a bounded regression head is first fine-tuned on the English DAIC-WOZ dataset (189 avatar-mediated sessions, PHQ-8), and the adapter then initializes fine-tuning on the Chinese PDCH dataset (100 real clinical consultations, HAMD-17), where a reinitialised, scale-specific head predicts the clinician-assigned score. All configurations use patient-level stratified 5-fold, 2-repeat cross-validation. On the data-scarce HAMD-17 target, the sequential protocol attains the best point-estimate MAE , RMSE, and macro-$F_1$ on both 0.6B and 1.7B backbones, outperforming target-only training and non-LLM baselines---4.96/6.59/0.36 with Qwen3-0.6B and 4.38/5.62/0.46 with Qwen3-1.7B. Ablations suggest that correctly aligned source supervision gives the best point estimates (unsupervised exposure and shuffled-label controls also show partial gains), that native-Chinese target input outperforms machine-translated English input, and that the reversed order yields no clear gain within run-to-run variance. The study is an exploratory, single-site internal evaluation: it does not establish screening or diagnostic utility, nor separately identify the contribution of the scale, language, or paradigm shifts. To our knowledge, no prior study evaluates this specific DAIC-WOZ-to-PDCH sequential transfer setting.
cs.CL / 40 / 2609.28471
Contrastive Learning for Authorship Verification
Peter Kirby
cs.CL · cs.LG
Abstract
Our results show that contrastive learning outperforms a classification-based approach to authorship verification under the tested settings. We identify loss function, batch size, training duration, pre-trained model, input context length, and random text span data augmentation as important factors of model performance. Based on these considerations, we develop a ModernBERT Bi-Encoder model that achieves 98.4% accuracy on the PAN21 authorship verification task.
cs.CL / 41 / 2609.27110
Feed the Panel Dimensions, Not Verdicts: Rubric-Decomposed Fusion of Vision-Language Aesthetic Judges
Amit Jadhav, Shaurya Beriwala, Beomjin Kim
cs.CV · cs.CL · cs.LG
Abstract
Vision-language models (VLMs) are deployed as zero-shot judges of image aesthetics, and panels of several models are recommended, on thin evidence, as the way to make such judges reliable. On two human-rated datasets, EVA and PARA, we find that a panel of holistic judges never significantly beats its best member, whether the verdicts are averaged or fused by a learned combiner. What a panel is worth depends on what it is fed. We therefore have each model score each image on the five dimensions of a frozen, human-written rubric and fuse those scores, alongside each model's verdict, across model families with an out-of-fold combiner. The dimension scores measure what their labels claim: with the overall human score partialled out, a dimension prompt carries more attribute-specific information than the holistic prompt in 28 of 30 model-attribute cells. Fused, they beat the best single VLM in all ten three-family panels on EVA (against that best single model, +0.07 Spearman rho for the strongest trio and +0.10 for the pre-declared one, and +0.06 and +0.07 when averaged over twenty fold partitions; against the panel mean, the primary test gives +0.118 on its EVA design set), and on PARA they reach parity under Spearman rho and a small, non-significant loss under Kendall tau-b, where one model already captures 85% of the human noise ceiling. It is not a feature-count artefact: giving the same combiner an equal number of pure holistic columns, split from the same repetitions, does not reproduce it. The gain costs a few hundred labels, which do not transfer between datasets, and 4.8x the API calls on EVA; we report it with paired bootstraps and Kendall tau-b, alongside a failed pre-registration and the configurations that lost.
cs.CL / 42 / 2609.27470
DeltaS: Reading the Gated Linear Attention State for KV Cache Eviction in Streaming Video
Taeyoun Kwon, Seungjin Kim, Hyeonyu Kim, Moon Hwan Kim
cs.CV · cs.CL · cs.LG
Abstract
Recent video-language models increasingly adopt hybrid architectures that interleave linear and full attention layers for efficient long-context processing. While the recurrent state of linear attention remains fixed in size, the KV cache of full attention continues to grow with the video stream, making eviction necessary under a bounded memory budget. The key challenge in streaming is that eviction must occur before the question arrives, so what to retain has to be decided without the question. Existing eviction methods derive token scores from the KV cache itself, using position, attention, or key-value representations, and attention-based scores further require proxy queries or extra computation. Hybrid backbones offer another source of signal. In gated-delta linear attention, the recurrent state is updated by the residual between each input and what can already be retrieved from the state, so its change over a chunk of frames reflects how much new information the chunk brings. We propose DeltaS, a query-agnostic, training-free method that retains video chunks inducing larger normalized state change, or state drift. In a controlled comparison with the budget and retention policy held fixed, state drift outperforms position-, attention-, and key-value-based signals. With a signal costing only 1.9% of the forward pass, DeltaS surpasses the strongest query-agnostic bounded-memory baseline by 2.1 points on average across six long-video benchmarks and by 5.6 points on the longest benchmark. These results suggest that the two memories of hybrid architectures can work cooperatively. Code is available at https://github.com/MaumAI-Company/DeltaS.
cs.CL / 43 / 2609.26907
Small Cues, Big Consequences: Learning Pivotal Cues for Multimodal Meme Classification
Akshit Sharma, Prashant W. Patil
cs.MM · cs.CL · cs.LG
Abstract
Memes often derive their harmful, hateful, or sarcastic meaning from small but decisive visual, textual, or cross-modal cues. Existing multimodal classifiers can miss such evidence when relying mainly on global image-text representations. We introduce MemeCF, a cue-focused benchmark of 9,895 memes across harm, hate, and sarcasm, with annotations identifying the modality and rationale of the pivotal evidence. We also propose MemePIVOT, a local-global architecture for meme classification. MemePIVOT uses frozen CLIP features, unbalanced optimal transport to align words with image patches while allowing irrelevant evidence to remain unmatched, and an evidential fusion head to combine local grounding with global meme context under uncertainty. Experiments on HarMeme, PrideMM, and MemeCF show consistent gains over strong text-only, image-only, multimodal, and vision-language baselines. Cross-dataset and ablation results further show that explicit pivotal-evidence modeling improves robustness and contributes meaningfully beyond global multimodal representations. Our code and dataset are publicly available at https://github.com/AkshitSharma1/MemePIVOT
cs.CL / 44 / 2609.27195
Quieter Than the Room: Representation Drift and Task Robustness in Speech Encoders
Vsevolod Kovalev, Pranay Manocha
cs.SD · cs.CL
Abstract
Non-speech interference can change a speech representation without causing comparable task loss. We test eight frozen encoders on four tasks, adding non-speech sounds throughout recordings, during speech, or in pauses. Under whole-recording interference, embedding drift tracks task loss across seven sounds, with mean Spearman correlations of 0.81-0.88. Moving the same sound between speech and pauses changes this pattern. At quiet to moderate levels, pause interference produces larger drift, while speech interference usually causes greater loss on intent recognition, speaker verification and speech recognition. Emotion recognition shows a weaker placement effect. Pause interference also changes speech-frame representations beyond the injected region. Even below the estimated recording background, interference can change embeddings as much as repeated speech takes do. Drift helps rank the effects of different sounds, but larger drift does not consistently indicate greater task loss.
cs.CL / 45 / 2609.27382
When Entanglement Lower-Bounds Disparity: Auditing and Repairing Demographic Fairness in Audio Understanding Models
Kian Shamsaie, Iman Modarressi
cs.SD · cs.CL
Abstract
Speech technology penalizes some voices: recognition errs nearly twice as often for Black speakers, and accuracy declines for second-language accents and older speakers. We introduce TRIAD, an audit grid crossing 120 texts, 24 rendered demographic voice profiles (gender, age band, accent), and ten expressive styles via controllable text-to-speech, isolating perceived demographic attributes from content and affect. For ten open-weights encoders we define axis-fidelity functionals, principal-angle leakage between axis subspaces, and group-conditional gaps; a proposition proves that average probe disparity grows with the same aggregate voice-semantic leakage $Λ$ we measure, and a corollary shows that peak leakage forces worst-case disparity inside an active region. The measured mean-square probe disparity tracks $Λ$ (Pearson r = 0.93), and a black-box protocol exposes the same signature in two closed-source models. ORCA, an adapter combining axis-specific contrastive heads, an orthogonality penalty, and group-balanced sampling, cuts leakage 72% and roughly halves the gaps.
cs.CL / 46 / 2609.28344
Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding
Kaiyang Li, Shaobo Han, Yue Tian, Shihao Ji
cs.SD · cs.CL
Abstract
Audio-language models (ALMs) integrate acoustic perception with the knowledge encoded in language models, enabling contextual understanding of auditory events. Making these capabilities practical on devices with limited memory and computation motivates our focus on small ALMs with fewer than 200M parameters. We introduce a recipe that brings together architecture, data, and three-stage training to build Mizar, a 159.3M-parameter ALM. Its architecture connects a compact CED-Small audio encoder to SmolLM2-135M through a frequency-merging mapper. With supervision drawn from ReasonAQA, AudioMCQ, and AVQA, the model undergoes three training stages: audio-language alignment (Stage 1), audio-dependent fine-tuning (Stage 2), and post-training (Stage 3) aimed at strengthening weak skills while retaining learned capabilities. Across five random seeds, Mizar achieves mean accuracies of 52.92% on MMAU, 42.42% on MMAR, and 36.02% on ADQA-clean, surpassing the previous best-performing ALM below 200M parameters on all three benchmarks. It also supports local inference on a single CPU: on questions from the MMAU benchmark, the mean latency from opening the audio file to generating a complete answer is 1.09 seconds. Code and checkpoints are available at https://github.com/KaiyangLi1992/Mizar_159M.
cs.CL / 47 / 2609.28029
Tensor Decomposition of Transformer Key-Value Caches: Spectral Structure and Format Comparison
Rahul Krishnan, Volker Schulz
math.NA · cs.CL · cs.LG
Abstract
The key-value (KV) cache of autoregressive transformers can be viewed as a fourth-order tensor spanning attention heads, tokens, features, and grouped layers. We measure the singular-value spectra of all four mode unfoldings on Mistral-7B-v0.3 and LLaMA-2-13B and compare four standard tensor decompositions: Tucker, CP, tensor train, and t-SVD, at matched storage. The spectra partition the four axes into two classes. The token and feature modes carry low-rank structure, particularly for keys. The head and layer modes are nearly full-rank and resist compression at any practical error level. Among the four decompositions, Tucker achieves the lowest reconstruction error at every compression ratio from $2\times$ to $5\times$, because it can leave the full-rank modes untouched. Comparisons with two-dimensional unfolding baselines show that the preferred representation differs between keys and values: 2D methods achieve lower key error, while four-way Tucker achieves lower value error at matched storage. A mode-pinning theorem certifies the full-rank preservation from the measured spectra alone. Two further spectral properties affect the compressible modes without touching the full-rank ones: values reach a higher error floor than keys at every ratio, and post-RoPE keys lose $41\%$ - $64\%$ of their pre-RoPE compressibility on both models.
多智能体系统 (cs.MA)
10
cs.MA / 1 / 2609.27639
Agent-based Modeling: Equilibrium, Echo Chambers, and Efficiency in Hybrid Coevolutionary Opinion Games
Ming-Zhi Jiang, An-Tzi Teng, Jun-En Liu, Po-An Chen, Yung-Ming Li
cs.GT · cs.MA · cs.SI
Abstract
Opinion formation in online networks involves changes in both beliefs and social ties. Analytical models make it possible to study equilibrium and social cost, but usually represent communication as a fixed numerical update. LLM-driven agents offer a language-based alternative, yet their convergence and collective efficiency remain unclear. We develop the Hybrid Coevolutionary Opinion Game (H-COG), combining cost-minimizing Friedkin-Johnsen agents (Type-C) and Phi-4 language agents (Type-L) in a dynamically rewired K-nearest-neighbor network. We initialize 50 agents with opinions drawn from 5,199 Reddit comments on gun control and abortion. The comments are scored on a continuous [-1,+1] scale using a fine-tuned RoBERTa regressor, and a mixing parameter sets the proportion of each agent type. The experiments cover nine population compositions, three initial network topologies, and two topics. All 540 runs meet the convergence criterion within the simulation horizon. Under Type-L updating, the coevolving network reaches an attractor as reliably as it does under the analytical update rule, making an equilibrium-based efficiency comparison possible. The pooled Price of Anarchy is $5.558 \pm 0.309$ for purely Type-L populations, compared with $1.139 \pm 0.005$ for purely Type-C populations. A decomposition of social cost attributes most of this gap to language agents moving away from their intrinsic opinions, rather than to greater disagreement with their neighbors. The main findings are consistent across the three initial network topologies.
cs.MA / 2 / 2609.27365
Anchor and Perturb: Lazy Agent Remediation by Exploration Injection
Chengxi Zhong, Yongzhe Chang
cs.MA
Abstract
Anchor and Perturb (AnP) is a lightweight framework that resolves multi-agent coordination failures by decoupling exploratory variance injection from recurrent manifold stability. Existing remediation strategies predominantly alter mixing network architectures or enforce simultaneous exploration across the collective, which inevitably precipitates severe temporal-difference penalties in non-monotonic reward spaces. Specifically, AnP isolates underperforming lazy agents and injects an asymmetric exploratory pulse into targeted coordinates whilst anchoring converged teammates to nominal greedy exploitation. Empirical telemetry benchmarks demonstrate that AnP successfully rescues collapsed joint policies (recovering from a 5% evaluation win rate nadir back to 85%) and facilitates escape from suboptimal coordination plateaus, sustaining peak win rates of 90% without requiring structural network modifications.
cs.MA / 3 / 2609.27535
KITE: Scaling Jev Population Experiments with Sparse Flagship Calibration
Hengyu Li
cs.MA · cs.CY
Abstract
KITE queries a typed behavioral kernel once per unique state, then executes populations of any size from the table with event-keyed randomness and common random numbers. An expensive flagship model is reserved for sparse paired anchors that estimate intervention effects. Measured human-model discrepancy is propagated as shared error into every conclusion. Population-experiment cost thus scales with unique states and anchors, while uncertainty is governed by evidence about people rather than Monte Carlo noise. On Epstein experiments with 9,070 participants, anchors covering 1.7% of states reduced effect error by 41% (absolute MAE reduction 0.0125). On 37 held-out SocSci210 experiments, 0.5-1.5% anchor coverage raised captured decision gain from 0.27 to 0.39. The kernel passed content-fidelity criteria in all 15 new countries of a 16-country study. Shared discrepancy yielded retrospective coverage of 93% and 96% at nominal 80% and 90%, versus 29% and 36% from human sampling uncertainty alone. A million agents executed 20 tabulated steps in 0.9 seconds on a laptop. This architecture offers a route to screening candidate interventions before human trials, multi-country content audits, and uncertainty-aware policy comparison at the cost of a few thousand kernel calls with sparse flagship anchors. Property-specific evidence records connect each use to its validation scope, correction provenance, and uncertainty, making these applications auditable.
cs.MA / 4 / 2609.27573
Agent-Based Modeling of Systems of Systems
Jean-Baptiste Soyez, Gildas Morvan, Rochdi Merzouki, Daniel Dupont
cs.MA
Abstract
This paper deals with the generic modeling of systems of systems (SoSs) using agent-based modeling. SoSs are large-scale systems, including numerous-possibly heterogeneous-interacting component systems evolving in a dynamic environment. The aim of this paper is to provide generic formalism allowing to represent and control the whole complexity of a SoS using agent-based simulations. In particular, organizational aspects of SoSs are managed with the Agent-Group-Role model. Functional aspects, guiding SoSs to accomplish their global goals, are handled via a functional specification. Multilevel aspects are modeled with the Influence Reaction Model for Multilevel Simulation (IRM4MLS) agent-based meta-model. Models generated using this formalism encompass static and dynamic aspects of SoSs. They consider reorganization of SoSs caused by changes of goals or subsystem capacity. All these elements are illustrated in this paper using a SoS case study of Intelligent Autonomous Vehicles initiated by the Intelligent Transportation for Dynamic Environment (InTraDE) European project to automate the port container logistic.
cs.MA / 5 / 2609.27994
Compliant with Local Controls, Collectively Discriminatory. A Governance Architecture for Multi-Agent AI in Regulated Finance
Jose Manuel de la Chica Rodriguez, Juan Manuel Vera Diaz, Pablo Delgado Romero
cs.MA · cs.AI
Abstract
Financial institutions are beginning to deploy agentic workflows in credit, fraud, collections, compliance, and operational control. Governance remains largely component-centric: each model or agent is specified, tested, authorized, and monitored locally. That is insufficient when institutional risk arises from the joint behavior of many locally acceptable components. We call this gap constitutional non-compositionality: local compliance checks need not compose into acceptable collective outcomes such as bounded disparate impact, market integrity, or traceable accountability. We propose ARIA as a finance-specific reference architecture and falsifiable research agenda for agent-population governance. It organizes six capabilities across normative-accountability, execution-control, and assurance-learning planes: policy specification, population-level observed-versus-expected behavior monitoring (M2), bounded authority, runtime containment, adaptive policy change, and preserved human oversight competence. Two simulations illustrate shared-signal thin-file exclusion under local controls and earlier warning from observed-versus-expected distributional monitoring in a constructed drift regime. The contribution maps these controls to fair-lending, EU AI Act, model-risk, and conduct-supervision evidence needs, and closes with a validation agenda rather than a production-effectiveness claim.
cs.MA / 6 / 2609.28190
Connectivity Preservation and Graph Stretching in Range-Only Swarm Dispersion
Ariel Barel
cs.MA
Abstract
We study connectivity-preserving finite-jump dispersion of anonymous, identical, and oblivious agents under an idealized range-only sensing model. Each agent measures only the distances to its visible neighbors, without bearings, identifiers, communication, memory, or a shared coordinate system. We derive the largest isotropic displacement certifiable as safe from these measurements alone. The resulting rule requires only the distance to the farthest visible neighbor: each agent selects a random direction and moves by half of its remaining visibility margin. The rule preserves every existing visibility edge under synchronous finite motion and therefore preserves connectivity. For two agents, we prove positive conditional drift in squared distance, almost-sure convergence to the visibility boundary, and finite expected time to reach any fixed neighborhood of that boundary. A one-million-run Monte Carlo experiment agrees with the exact first-round moments and estimates approximately 9.5 rounds to reach distance 0.97V from coincident initial positions; an independent Bellman-equation computation gives the same estimate. For general swarms, 1,000 runs across five initial-topology classes reproduce the deterministic safety guarantee at implementation level and reveal a consistent topology-dependent ordering of attainable diameter under the tested protocol. These results provide a theoretical foundation for connectivity-preserving multi-robot dispersion under minimal sensing, while isolating the guarantees achievable from anonymous range measurements alone.
cs.MA / 7 / 2609.27727
Sovereign Grassroots Currencies: A CBDC Architecture for Credit and Monetary Policy (Full Version)
Ehud Shapiro
econ.GN · cs.DC · cs.MA
Abstract
A Central Bank Digital Currency (CBDC) is central-bank money in digital form, held by the public. Leading designs have two limitations: conversion from bank deposits into CBDC can accelerate deposit flight, requiring safeguards; the CBDC stays outside credit creation and monetary-policy operations. Here we present a CBDC architecture that overcomes these limitations, based on grassroots currencies. It has three components: (1) Money: sovereign grassroots coins, which are digital debts of one unit of fiat currency issued by the central bank, constituting a direct CBDC; (2) Credit and Liquidity: non-sovereign grassroots coins, which are digital debts of one unit of the same fiat currency, redeemable at par, that can be issued by any person, natural or legal - adding credit; and (3) Interest: grassroots bonds, sovereign and non-sovereign - adding maturity, and with it interest, standard banking instruments, and the central bank's instruments of monetary policy. The central bank can therefore lend, absorb liquidity, set its rates and buy and sell securities in the coins and bonds the public holds, choosing the counterparties and terms of its credit operations, and without converting bank deposits into newly issued central bank money on demand. We prove that one unit of the fiat currency is the only arbitrage-free price of a grassroots coin whose issuer meets presentations, and argue that the central bank's lending rate and the rate on its own bonds bound what its counterparties pay and accept on comparable terms; the central bank can choose to deal with any counterparty, not just banks.
cs.MA / 8 / 2609.26906
Impact-Time Guidance via Normal Contraction to a Time-to-Go Isochron
Shivam Bajpai, Abhinav Sinha
eess.SY · cs.MA · cs.RO · math.DS
Abstract
We develop a contraction-based perspective on impact-time guidance that augments a baseline homing command with a timing bias. The proposed perspective treats the prescribed schedule as a moving time-to-go isochron and regulates motion normal to that set through velocity-normal lateral acceleration while the interceptor's speed remains constant. We derive a transport equation that characterizes homing-compatible time-to-go coordinates and define a predictor defect that quantifies the mismatch of approximate maps. We show that the scalar timing channel induces a coordinate-invariant rank-one metric on the normal quotient. To account for bounded lateral acceleration, we formulate a robust scalar filter and derive a necessary and sufficient condition for pointwise feasibility. We then show that terminal calibration and funnel invariance establish first interception at the prescribed time under the stated assumptions. We also develop a preterminal alignment and homing handover that avoids singular inversion as lateral timing authority vanishes near collision-course alignment. The proposed perspective accommodates analytic, numerical, and learned time-to-go maps that satisfy the required calibration and regularity conditions.
cs.MA / 9 / 2609.27499
Distributed Stochastic Approximation Algorithms and Heavy-Tailed Age of Information
Adrian Redder, Arunselvan Ramaswamy, Holger Karl
math.OC · cs.DC · cs.MA · cs.NI · math.PR
Abstract
Algorithms in multi-agent systems such as federated learning, mobile robotic swarming, and consensus control can be designed and analyzed as distributed stochastic approximation algorithms. Such algorithms involve information exchanges between agents for various computations. The freshness of the information can be quantified using the Age of Information (AoI) metric. Consider robotic teams operating in highly obstructed geographical settings, such as subterranean or dense urban environments. Because of spatial disconnections, AoI has empirically been observed to be heavy-tailed with unbounded moments. However, most analyses assume AoI with bounded moments, creating a gap between theory and practice. To the best of our knowledge, ours is the first analysis under general heavy-tailed AoI with potentially infinite mean. We study the stability (almost sure boundedness of the distributed iterates) and convergence of multi-agent systems that are strictly dissipative in the scaling limit (system at ``infinity''). Examples include most gradient-based and consensus algorithms under the Robbins-Monro step-size regime.
cs.MA / 10 / 2609.27636
Multi-Agent AI Architecture for Regulated Insurers: A generic AI framework under Solvency II and the AI Act in Austria and Germany
Walter Kurz
q-fin.GN · cs.MA · q-fin.RM
Abstract
This paper proposes a formal multi-agent architecture for implementing enterprise AI in regulated insurance firms, integrating economic theory with institutional design. The framework synthesises three core theoretical perspectives: Arrow's risk pooling theory to formalise risk transformation under uncertainty, Nash equilibrium to model strategic interactions between decision agents, and Principal-Agent theory to address incentive alignment under information asymmetry. The insurer is modelled as a constrained optimisation entity operating under solvency, legal, ESG, and operational boundaries, with specific focus on the regulatory contexts of Austria and Germany. The architecture decomposes the firm into multiple specialised agents, each representing distinct functional domains such as capital management, underwriting, claims processing, compliance, fraud detection, and client interaction. Human-in-the-loop agents are integrated through a tiered access control system, ensuring differentiated data visibility and decision influence based on user roles. An orchestrator agent supervises inter-agent coordination, enforcing regulatory admissibility and institutional coherence under frameworks such as Solvency II, the AI Act, and the Insurance Distribution Directive. Protocol integration is based on asynchronous execution and dual-layer communication infrastructures, specifically the Model Context Protocol (MCP) and Agent-to-Agent (A2A) messaging. This structure enables the systematic design of compliant, auditable multi-agent systems aligned with the institutional logic of financial firms in Austria and Germany.
软件工程 (cs.SE)
17
cs.SE / 1 / 2609.27676
A ready-to-deploy MLOps software platform for satellite and NEO detection at meter-class ground-based observatories
Rafael Carrillo Navarro, Pablo García-Martín, René Duffard, Juan Luis Orellana Montes de Oca, Oscar Ortega Ríos
astro-ph.IM · cs.SE
Abstract
Ground-based astronomical observations frequently contain streaks produced by artificial satellites, space debris, and potentially Near-Earth Objects (NEOs). While machine-learning models can reliably detect these features, their practical adoption in observatory operations is often limited by the lack of integrated tools for visual inspection, validation, workflow management, and structured data storage. This paper presents the StreakMind Workbench, a framework that applies MLOps practices to bridge research-oriented machine-learning pipelines with routine observatory operations. Rather than introducing new detection algorithms, the Workbench addresses a software-engineering challenge in astronomical computing: maintaining a single authoritative source of scientific processing code while providing astronomers with an operational environment for workflow execution and result inspection. Integrated with the reference StreakMind AI model of Carrillo et al. (2026) and implemented in Python using PyQt5, the Workbench supports the complete workflow from FITS ingestion to database storage, including inference, result inspection, database exploration, training management, and Minor Planet Center formatted observations. Validation on 273 images from La Sagra Observatory demonstrates successful end-to-end workflows while maintaining consistency with the underlying StreakMind scientific code. The platform facilitates operational use of a research ML pipeline in meter-class observatories and moderate-scale campaigns, supporting Space Situational Awareness and planetary defence.
cs.SE / 2 / 2609.27058
Stage-Supervised Latent Reasoning for Single-Shot JavaScript Deobfuscation
Rong Feng, Suman Saha
cs.SE
Abstract
JavaScript obfuscation is widely used to protect code, but it also makes program analysis and security review substantially harder. Existing LLM-based deobfuscation methods usually treat the task as one-step translation, ignoring the staged structure of practical deobfuscation pipelines. This WIP paper proposes a stage-aware latent reasoning framework that converts intermediate outputs from a deterministic deobfuscation tool into supervision for Coconut-based training. The model learns from multi-stage rewrites during training but generates the final cleaned program in a single shot at inference time. Preliminary results on JsDeObsBench show that the Coconut-based model improves syntactic validity to 50%, compared with 15% for direct fine-tuning and 25% for a zero-shot baseline, and reaches 80% semantic correctness among valid outputs, indicating better semantic faithfulness than either comparison model in deobfuscation.
cs.SE / 3 / 2609.27069
Does Graph Structure Earn Its Place in Microservice Root-Cause Analysis? A Controlled Study on RCAEval, and What the Benchmark Was Really Measuring
Imad Buljić
cs.SE · cs.LG
Abstract
Graph neural networks dominate recent work on microservice root-cause analysis, yet recent results question whether the graph contributes. Those results compare whole pipelines, so when a flat model wins one cannot tell whether structure is useless or redundant. We run the comparison they imply on RCAEval: three learned arms with identical features, optimiser, validation split, early-stopping rule and scoring head, in which a single term separates the graph arms. Across two RCAEval benchmarks, two topology sources and four regimes we find no reliable graph-specific effect: in-distribution the graph model leads the conventional flat model by 0.003 Avg@5 (p = 0.844, n = 6 disjoint folds). Auditing the pipeline surfaced two benchmark properties that condition any result on it. RCAEval injects faults into only five services per system while exposing 12 to 70 in telemetry, and the headline metric is Avg@5: a ranker reading no telemetry at all places the true culprit in the top five on 99.7 percent of held-out incidents, scoring Avg@5 0.488. That prior, not the uniform-random 0.137, is the honest in-distribution floor, and it collapses to 0.192 across systems. The second property is a non-uniform column schema that silently zeroes telemetry for most RE1 cases. We reproduce a published baseline, BARO, RCAEval's own reference implementation; on the one system with a clean schema it reaches similar aggregate accuracy to our heuristic, within 0.004, under a different scoring rule. The audit motivated a new model. PSC-GRCA separates a candidate score into a system prior, telemetry evidence and a centred graph residual, and reaches mean Avg@5 0.915 against 0.864 for the flat baseline, while its ablations locate most of the gain in the prior term rather than the graph. We close with a twelve-item checklist for graph-versus-flat ablation studies, distilled from sixty-two recorded defects.
cs.SE / 4 / 2609.27249
From PyTorch to the NPU: LLM-Agent-Driven Model Conversion Across Heterogeneous Inference Runtimes
Jianhao Su, Zhanwei Wu, Chia-Heng Tu, ShengTing Huang
cs.SE · cs.DC
Abstract
Edge AI model deployment is a multi-stage engineering process involving model conversion, operator compatibility handling, runtime integration, and precision verification. While prior work has demonstrated agent-based automation for Qualcomm AI Runtime, the broader edge inference runtime ecosystem, including Intel OpenVINO, Rockchip RKNN, NVIDIA TensorRT, and ONNX Runtime, presents distinct toolchains and optimization strategies. This paper extends AIPC (AI Porting Conversion, an LLM agent-driven methodology for AI model deployment automation previously demonstrated on Qualcomm AI Runtime) to multi-runtime scenarios, proposing an LLM agent-driven approach for automated single-model-to-single-runtime deployment across heterogeneous inference backends, such as Intel OpenVINO, Rockchip RKNN, NVIDIA TensorRT, and ONNX Runtime. We decompose the edge AI deployment into standardized, verifiable stages, and inject runtime-specific domain knowledge into the agent execution flow through agent skills, auxiliary scripts, and staged verification loops. Using representative vision models, we demonstrate that agent-based deployment can complete the conversion from a PyTorch model to its executable inference, targeting OpenVINO for x86/NPU, RKNN for RK3588, TensorRT for NVIDIA GPU, and ONNX Runtime for Qualcomm NPU with a focus on FP16 precision deployment feasibility verification. The contributions of this paper primarily lie in providing multi-runtime deployment engineering practice experience, toolchain mapping analysis, a layout-adaptation and inference-replacement layer that removes manual transpose insertion from the agent's repair burden, and an empirical characterization of agent deviation behavior under structured knowledge injection, rather than large-scale systematic benchmarking or cross-runtime operator repair strategy comparison.
cs.SE / 5 / 2609.27263
Specifying and Maintaining Agentic Workflows: An Empirical Study of GitHub Agentic Workflows
Jasem Khelifi, Issam Oukhay, Ali Ouni, Mohammed Sayagh, Mohamed Aymen Saied
cs.SE
Abstract
Agentic workflows shift software development from prompting AI agents for individual tasks to defining recurring work that agents execute automatically. GitHub Agentic Workflows (gh-aw) enables this approach through Markdown files that combine natural-language instructions with configuration and compile into executable GitHub Actions workflows. Unlike conventional workflows that primarily prescribe scripted operations, these files delegate tasks requiring interpretation to AI agents. They also couple agent instructions with execution triggers, making those instructions operational specifications for repeated repository activities. However, how developers structure and maintain these specifications, and which execution requirements and safeguards they express, remains insufficiently understood. In this paper, we examine the structure, evolution, and instruction content of gh-aw Markdown files to inform how practitioners define and maintain agent-run work. We analyze 1,248 files from 276 repositories, 20,841 commit-file events, and 288 resolved instruction-label sets from a sample of 294 files. Our results show that workflow instructions extend beyond short prompts, with a median of 556.5 words and code blocks in 62.1% of files. Among files with at least 120 days of observed activity, 78.2% still receive updates in month 4, while size-normalized churn decreases after the first month. Tasks, outputs, constraints, and process instructions each appear in over 93% of labeled workflows, yet only 9.4% explicitly address prompt-injection defense. LLM classification achieves an F1 score of 0.818 and Cohen's Kappa of 0.715 against the resolved human labels. These findings suggest that developers should account for the evolution of workflows copied or referenced across repositories and consider adding prompt-injection defenses, resource budgets, and evidencecredibility checks where applicable.
cs.SE / 6 / 2609.27296
Teach-to-Crash: A Closed-Loop Student-Teacher LLM Framework for Collision-Inducing Test Scenario Generation
Zaid Ghazal, Khouloud Gaaloul, Bruce Maxim
cs.SE · cs.AI · cs.MA · eess.SY
Abstract
Validating Autonomous Driving Systems (ADS) in simulation requires testing architectures that can discover rare, safety-critical failures while generating scenarios that are executable, diverse, and useful for downstream failure analysis. We introduce Teach-to-Crash, a closed-loop testing framework that combines a constrained ego-centric scenario representation, stagnation-aware search control, and a dual-LLM architecture for adaptive failure discovery. A high-reasoning Teacher LLM acts as an adaptive search controller, while a low-reasoning Student LLM emits simulator-executable scenarios in a strict JSON schema. The Teacher intervenes only when rolling collision rate and time-to-collision metrics stagnate, providing strategic guidance to redirect the search. In a CARLA case study with two experimental setups that vary the ego vehicle's speed policy, Teach-to-Crash achieves the highest Collision Hit Rate (90.79%), the shortest mean Time-to-Collision (18.31 s), and a competitive Collision Discovery Rate (136.21). PAFOT attains a higher mean CDR (179.44), but with substantially larger variance. Teach-to-Crash also yields the highest diversity (0.547) and, averaged across both setups on the CARLA Traffic Manager controller, the highest avoidability-based usefulness proxy (60.04%) among the compared methods. These results, within the evaluated CARLA scope, provide evidence that closed-loop dual-LLM reasoning can steer adversarial simulation-based testing over a constrained executable program space, generating failures that are frequent, structurally diverse, and assessed as more frequently avoidable.
cs.SE / 7 / 2609.27309
FairTest: Search-Based Fairness Testing for Multi-Agent Reinforcement Learning Systems
Xiaotong Wang, Xuan Xie
cs.SE · cs.LG
Abstract
Multi-agent Reinforcement Learning (MARL) trains a team of agents that share one environment and learn their policies together. Training maximizes the team return, and a high return does not imply that the rewards are shared fairly among the agents in every episode. Testing is an established way to discover the failures of deep reinforcement learning, yet few methods address the fairness of MARL. In this work, we propose FairTest, a search-based testing approach that seeks the unfair executions of a MARL policy. The design combines search guidance with test prioritization. The guidance scores each candidate with three fitness functions. One measures the fairness of the runs already performed, another predicts the fairness from abstract states and fairness features, and the third reads the decision uncertainty from the policy. Crossover and mutation derive further candidates from the observed executions. The prioritization ranks the candidates by the predicted fairness and the decision uncertainty, so that the runs reach the candidates where failures are expected. FairTest is evaluated on three environments and two MARL algorithms, and four baselines are given the same budget. It detects the most fairness failures compared to three baselines with statistical significance and large effect sizes. The failure count exceeds that of the strongest baseline by 221% on average and coverage improves by an average of 23%.
cs.SE / 8 / 2609.27354
Constraint-Driven Context Engineering: Designing Domain Interfaces for AI Systems
Xiwei Xu, Chen Wang, Mengmeng Yang, Yipeng Zhang, Jacky Jiang, Suyu Ma, Youyang Qu, Ming Ding, Liming Zhu
cs.SE · cs.AI
Abstract
Generative AI systems are increasingly deployed to address domain problems. These systems operate under technical, regulatory, institutional, and normative constraints that define acceptable AI behaviour and outcomes within their domains. We observe a recurring pattern in our industry engagement: partners often arrive with a functioning but relatively generic AI solution. The challenge is no longer to build an AI system from scratch, but to improve the quality and domain appropriateness of an AI-generated solution. In these settings, the limiting factor is often the quality, scope, and structure of the context available to the system. Yet, existing context engineering approaches primarily focus on supplying domain knowledge through retrieval, memory, and tools, with limited support for systematically identifying and operationalising the constraints that govern AI systems in their operational environments. This paper proposes Constraint-Driven Context Engineering (CDCE), a design approach for engineering domain interfaces for AI systems. Drawing on software architecture design and Domain-Driven Design (DDD), CDCE treats domain constraints as first-class design drivers. It identifies and characterises constraints, determines the required context assets, and designs representations through which these assets are made available to AI systems. We conducted a comparative multiple-case study with industry and public-sector partners across educational assessment, healthcare decision support, and financial-distress prediction. Depending on their characteristics, constraints can guide AI behaviour, enforce permissible boundaries, or support verification of AI-generated outcomes. The cases demonstrate CDCE's applicability across contrasting domains and show how constraint characteristics shape the resulting domain interfaces.
cs.SE / 9 / 2609.27571
FDE-Bench: Evaluating LLM Agents for Deployment Environment Configuration
Weihang Ding, Junfei Zhan, Yueting Li, Qirong Guo
cs.SE · cs.AI
Abstract
Deployment requires an agent to turn application code into a running system whose services connect, become ready, and remain observable. FDE-Bench evaluates this capability with 136 deployment-configuration tasks spanning Docker images, multi-service Compose stacks, and Kubernetes, in greenfield and diagnose-and-repair modes. Agents submit declarative artifacts that are collected, rebuilt, and redeployed in a pristine environment. Four gated binary check layers measure build, readiness, behavior, and conformance to the deployment specification, using programmatic checks without an LLM judge. A four-arm release gate requires a resolving reference solution and rejects tasks solved by do-nothing, specification-transcription, or generic-stub submissions. The released check annotations expose the link between 2,145 checks and their specifications, including seven documented gaps. Three additional adversarial strategies test shortcuts in the grading signals; none resolves any of the 135 tasks they cover, while a vacuous health probe passes readiness and exposes the need for downstream checks. On the 136-task evaluation grid, seven language models from four providers use the same four-tool scaffold and resolve 52.9-75.0 percent of tasks. The three zero-intelligence floors resolve none and reach a mean Deployment Score of at most 0.44. Readiness is the largest failure stage, accounting for 110 of 313 unresolved episodes. Mean resolution rate is 30.7 percentage points higher on the repair task group than on the disjoint greenfield group, with a positive gap for every model; ten tasks resist all seven. In a 25-task case study, one practicing engineer directing Claude-Sonnet-5 resolves 92 percent against 72 percent for the autonomous baseline. FDE-Bench links deployment success and failure to artifacts that can be inspected and replayed.
cs.SE / 10 / 2609.27585
Unity Insight: A Production Code--Asset Index for LLM Coding Agents in Unity Projects
Shenhua Gu, Hongqiang Zhu, Fan Zhang, Jinming Zhang, Hao Chen
cs.SE
Abstract
LLM coding agents increasingly operate inside game-engine repositories, where application logic is inseparable from serialized assets: a single gameplay change may span C\# scripts, prefabs, scenes, and ScriptableObjects wired together by Unity GUIDs. The retrieval tools agents carry today---shell utilities and code-only indexes---cannot answer basic cross-file questions, because these relationships live in \texttt{.meta} files and YAML assets rather than in code. We present Unity Insight, to our knowledge the first persistent, LLM-facing, agent-integrated cross-file code--asset index for Unity projects, shipping in production with Tuanjie Codely, the agent CLI of Tuanjie Engine, since its public launch on 2026-07-28. In a paired experiment---28 project-specific questions on two Unity games, same model and harness, one run per arm per question---the index-backed agent spent 53\% fewer tokens and 52\% less wall-clock time than a general-purpose exploration agent (exact paired sign tests, $p{<}0.004$), using only its typed index-query tools.
cs.SE / 11 / 2609.27627
AST-Based Automated Elimination of break and continue Statements in Java Code
Andrés Juárez, José Chicano, Rubén Saborido
cs.SE
Abstract
This work presents the development of an automatic refactoring tool for Java code built on top of the Eclipse JDT API. The proposed approach transforms control structures containing break and continue statements within different types of loops into semantically equivalent constructs that avoid their explicit use. To achieve this, auxiliary boolean variables are introduced to restructure the control flow while preserving the original program behavior. The main objective of this transformation is to improve code structure and enable the application of subsequent automated refactorings, particularly those based on the Extract Method operation, which are typically restricted by the presence of jump statements. The implementation relies on the analysis and rewriting of the Abstract Syntax Tree (AST), ensuring semantic equivalence in all addressed scenarios. The tool was validated through 54 manually designed test cases and 151 units tests, all of which produced satisfactory results. In addition, it was applied to 139 methods from seven open-source projects, generating code without compilation errors and preserving the original behavior as verified by the projects' test suites. The results demonstrate that the proposed approach safely automates the restructuring of code containing break and continue statements, facilitating further evolution and structural analysis.
cs.SE / 12 / 2609.27638
Foundations of Algebraic Architecture Theory: A Rising Sea of Geometry, Transport, Comparison, and Reconstruction
Hiroyuki Nakahata
cs.SE · cs.PL
Abstract
AI-generated software changes make it increasingly important to determine what a change preserves, where local consistency fails to extend globally, and which alternatives remain. We develop the foundations of Algebraic Architecture Theory (AAT) from Atoms, typed primitive facts, and Laws, equations that objects must satisfy. A reading specifies what counts as structure and which operations and laws to preserve. The main reconstruction theorem identifies the category of full geometries and all their structure-preserving morphisms with an independently defined category of local models, up to equivalence. Objects are recovered up to isomorphism and morphisms between fixed endpoints uniquely. The theory addresses gluing, diagnosis, transport, classification of changes, and reconstruction. From finite Atom families we construct cores closed under operations and geometries with sites and coefficients. We give conditions under which a Cech obstruction detects the existence of a global state and, through comparison with repair semantics, a global repair. We compare diagnoses and give a finite criterion for uniform invariance given computable finite data. Transport along exact changes has a universal property and commutes with base change on exact pointed pullback squares. Comparisons of routes generated from the same square, finite comparison diagram, and geometry factor into an invertible comparison and an idempotent normalization. We characterize when observations determine comparison preservation and classify compatible lifts. Encodings of lens and protocol semantics preserve and reflect laws and recover semantics-preserving morphisms. Applications classify and count operation-preserving changes and extend morphisms uniquely from finite tables. Corresponding Lean declarations are listed in the appendix.
cs.SE / 13 / 2609.27743
Satisfaction Is Not Explanation: Auditing Vacuity and Training Influence in Temporal-Logic-Guided Reinforcement Learning
Lorenzo Bacchiani
cs.SE
Abstract
A reinforcement learning policy that satisfies its temporal-logic specification has passed a test, not an assurance argument. The clause that matters to a reviewer may never have mattered to training: it may have been avoided entirely, forced by the environment regardless of what the policy learned, or redundant next to the ordinary task reward. Satisfaction probability and task return cannot tell any of this apart. This paper introduces an audit layer that can. It measures whether a specification clause was actually exercised, whether that role was forced or chosen, and whether the obvious way to test causation, weakening the clause and retraining, is even valid. Often it is not: we prove that comparable weaker/stronger training objectives can share perfect optima under standard acceptance-derived rewards, show related ablation hazards across a large corpus of published specifications, and then show that a properly designed intervention detects the effect it should. Across standard reinforcement learning benchmarks and published external artifacts, the audit layer separates six regimes that a single satisfaction number collapses into one. A policy that satisfies its specification has answered whether. This paper asks why.
cs.SE / 14 / 2609.28002
LLM-Assisted Workflow for Structural Difference Visualization in Evolving Software Requirements
Koi McFarland, Songhui Yue
cs.SE · cs.HC · cs.IR
Abstract
This paper presents an LLM-assisted workflow for visualizing structural differences in evolving software require- ments. Implemented in the OntologyWeb environment, the work- flow represents baseline and current requirements as triple-based semantic graphs and supports side-by-side comparison of curated graph snapshots. The comparison view aligns matched entities and uses visual encoding to highlight structural changes.
cs.SE / 15 / 2609.28130
Scenario-Driven Neuroevolution: Using Models to Guide Test Generation for Games
Gijs van Cuyck, Patric Feldmeier, Jan Tretmans, Gordon Fraser
cs.SE · cs.NE
Abstract
Automatically generating test inputs for games is challenging, as test generators must master the game to reach advanced program states while also ensuring robustness against the heavy program randomisation inherent to games. The test generator Neatest therefore optimises test suites consisting of neural networks that reach advanced program states and are robust to program randomisation, as they generate test inputs dynamically based on the current program state. Neatest is a white-box testing approach that aims to generate a network agent for each yet-uncovered statement or branch of the code using neuroevolution. Due to this iterative test generation approach, the algorithm does not scale well to larger programs that may contain thousands of branches. Furthermore, covering every statement or branch in a game often does not correspond to playing the game as intended. To alleviate these shortcomings, we propose combining Neatest with a model-based testing approach that allows game testers to define test scenarios via abstract game models. The test generator then no longer optimises networks to reach all branches or statements of a program, but instead trains networks to replicate the concrete desired testing behaviour expressed by the abstract game model. An evaluation on 13 Scratch games across varying genres demonstrates that Neatest, combined with model-based testing, is able to optimise agents that replicate the desired gameplay behaviour defined in the game models while increasing achieved branch coverage by 7% compared to the traditional code-guided Neatest approach.
cs.SE / 16 / 2609.28216
From Agent Output to Authorized Transition
Christopher Koch
cs.SE · cs.AI
Abstract
Agentic engineering systems can edit repositories, run tools and tests, build firmware, synthesize schematics, and prepare deployable or manufacturable artifacts. The assurance problem is therefore shifting from whether an agent can produce an output to whether an engineering lifecycle is justified in acting on claims about that output. Current products and standards provide sandboxes, approvals, hooks, traces, policy enforcement, attestations, bills of materials, and assurance representations, but these capabilities remain fragmented. This paper presents the Agile-V Assurance Spine, a cross-domain transition contract for software, firmware, and PCB engineering. Evidence is admitted only when it establishes required properties through an authoritative source profile, is bound to the exact artifact and frozen policy baseline, remains current with respect to declared dependencies, and satisfies risk-appropriate independence and authority. Gate decisions are recorded as receipts; approvals and exceptions are exact-scope and time-bounded; and authorization is rechecked at the effect boundary before merge, deployment, flashing, release, or fabrication. A bounded review of contemporary research, commercial platforms, open-source infrastructure, and standards positions the model relative to evidence-gated lifecycle control, continuous assurance, runtime admission, provenance, and AI/ML inventories. The paper contributes a precise vocabulary, compositional architecture, domain profiles, mapping to open-source implementations, and an adversarial evaluation agenda. It does not claim regulatory conformity or demonstrated production superiority.
cs.SE / 17 / 2609.28349
Entangle: Uncovering Collaboration in the GitHub Quantum Software Ecosystem
Angel Luis Lara-Martín, Ricardo Pérez-Castillo
cs.SE · cs.IR
Abstract
Quantum computing is moving from research laboratories towards early commercialization and broader socio-technical adoption, supported by sustained hardware progress and a rapidly expanding open-source software ecosystem. This momentum is especially visible on GitHub, where many quantum and hybrid software projects coexist around frameworks such as Qiskit, Cirq, PennyLane and Amazon Braket. However, this ecosystem remains fragmented, making it difficult to understand who shapes quantum software, where expertise is concentrated, how collaboration flows across organizations and disciplines, and which actors connect otherwise separated communities. This paper presents Entangle, a data-driven analysis of the open-source quantum computing ecosystem on GitHub. Starting from 71 domain keywords, Entangle identifies more than 1,500 quantum repositories, 27,000 contributors and 400 organizations, revealing an ecosystem strongly organized around four leading industrial vendors, but also supported by 2,387 contributors who connect projects, organizations and domains. These findings provide practical evidence for responsible quantum innovation by making visible patterns of influence, dependency, collaboration and knowledge transfer. They also offer actionable indicators for strategic decisions on investment, hiring, partnerships, ecosystem stewardship and capacity building. More broadly, Entangle shows how open-source intelligence can support a more transparent, measurable and governable quantum software ecosystem, helping align technical development with responsible innovation, public--private coordination and long-term sustainability.
操作系统 (cs.OS)
1
cs.OS / 1 / 2609.27266
xTier: Intelligent Tiering for CXL-Enabled Memory
Sriranga Ramaswamy, Yueqi Chen
cs.OS
Abstract
CXL-enabled memory expands server memory capacity, but introduces a page-placement problem: the operating system must decide which pages should reside in DRAM and which should reside on slower CXL memory. Existing systems make this tradeoff in one of two ways. Userspace controllers support flexible policies, but expose placement decisions to scheduler jitter and kernel-userspace crossing overhead. Kernel-space systems avoid this latency, but rely on fixed heuristics that must generalize across workloads. We present xTier, a kernel-resident learned memory-tiering system. xTier attaches eBPF programs to PEBS events and uses a compact quantized MLP to score sampled pages inside the kernel at microsecond-scale latency. Rather than reacting to every candidate, xTier converges to a low-churn placement for the current workload phase, reduces sampling cost after convergence, and returns to a higher sampling cadence when the workload shifts. We evaluate xTier on six memory-bound workloads at DRAM:CXL ratios from 1:5 to 1:25. The advantage grows as the DRAM budget tightens. At 1:15 and beyond, xTier is the fastest system in 14 of 18 configurations. Where it is not fastest, it trails the best baseline by 3.9% on average. It reaches this performance while moving 13% fewer pages in geometric mean, and 22% fewer at the tighter ratios. When a workload changes phase, xTier rebuilds its hot set in DRAM faster and more completely than any baseline.
硬件架构 (cs.AR)
8
cs.AR / 1 / 2609.27319
Exploiting Decompression Latency for Covert Channels in Inter-Line-Compressed LLCs
David K. Oh, Hiroshi Sasaki
cs.AR
Abstract
The recently proposed XOR cache is an inter-line-compressed last-level cache (LLC) that leverages the data-inclusion relationship between the private caches and the LLC, compressing two cache lines into one by XORing them. The architecture relies on the cache coherence protocol for data decompression. In this paper, we demonstrate that this mechanism - specifically the latency asymmetry between a cache hit on an uncompressed vs. compressed line - introduces microarchitectural vulnerabilities. Based on this observation, we propose a covert channel attack targeting the XOR cache. A colluding sender controls the receiver's access latency by triggering decompression through targeted write requests to partner cache lines. By exploiting the data-dependent compression behavior of the XOR cache, the sender and receiver establish the channel using pre-agreed data values. The channel achieves higher bandwidth than the Prime+Probe baseline for two reasons: first, each bit is encoded in the compression state of an individual line rather than the occupancy of a cache set, so a single set carries multiple bits; second, each bit is resolved by manipulating coherence-protocol state rather than forcing shared-cache evictions, so it costs fewer LLC accesses and demand misses than Prime+Probe. Full-system simulations show a bandwidth of 2.9 Mbps at an observed 0.98% bit-error rate (BER) over 50,000 transmitted bits, 13.1 times the bandwidth of Prime+Probe under the same sub-1%-BER selection rule.
cs.AR / 2 / 2609.27437
Mamba-Family State-Space Model Kernels on a Programmable CGLA
Takuto Ando, Yasuhiko Nakashima
cs.AR
Abstract
Edge and embedded inference is constrained by power and data movement. Mamba-family state-space models replace attention with sequence-linear recurrence, but their inference path combines dense projections, short-reduction SSD kernels, and recurrent-state updates. This paper maps these kernel groups onto IMAX, a programmable CPU-Grounded Linear Array (CGLA), and measures them from kernel execution to token-level integration. Projection kernels match the long-reduction IMAX pipeline, whereas SSD Step-1 is limited by short reductions and kernel-boundary overheads. Mamba-130M token-level integration identifies projection GEMV as the decode bottleneck. These results show that programmable CGLAs fit long-reduction projection kernels, while SSD and decode-time projection support require boundary reduction and persistent-weight execution.
cs.AR / 3 / 2609.27438
Energy-Oriented CGLA Mapping of a Memory-Polynomial Digital Predistortion Kernel
Takuto Ando, Yasuhiko Nakashima
cs.AR
Abstract
Memory-polynomial digital predistortion (DPD) evaluates a small fixed coefficient set over a sliding input history, so its reduction step is a complex-MAC workload with local reuse. We map this DPD reduction kernel onto In-Memory Accelerator eXtension (IMAX), a programmable CPU-Grounded Linear Array (CGLA) composed of a one-dimensional processing-element/local-memory pipeline. For a (P,M)=(5,5) odd-order memory-polynomial instance, the mapping keeps the 120 B coefficient set in local memory, advances the five-tap history over 1024-sample tiles, and realizes the 15 order-delay terms as a 33-stage streaming complex-MAC reduction. The evaluation measures kernel latency and modeled energy. All measured paths use the same single-precision complex workload of 32 sequences, each with 2048 complex samples, across an IMAX FPGA prototype, a CUDA implementation on an RTX 4090 system, and an ARM-NEON implementation on Jetson AGX Orin. With this 1024-sample tile configuration, the IMAX FPGA prototype reports 20.201 ms end-to-end latency and 1.948 ms kernel-only latency. Using the previously reported 28 nm IMAX frequency and power model, the projected IMAX configuration gives 3.14 ms end-to-end latency and 0.34 ms kernel-only latency. The RTX 4090 baseline has the lowest end-to-end latency at 0.484 ms. Under model-based platform power accounting and the stated power assumptions, the projected IMAX configuration gives 169.1 times smaller modeled end-to-end energy per batch than the RTX 4090 baseline. This value uses platform power assumptions rather than workload-dependent runtime power or a direct silicon power measurement. A controlled synthetic PA-model validation checks that the same 15-term form improves test-set NMSE by 26.1 dB and ACLR by 26.0 dB. These results characterize the mapped memory-polynomial DPD reduction on IMAX for the evaluated tile configuration and power model.
cs.AR / 4 / 2609.27706
MVP: A Motion-Predictive Speculative Vision Pipeline with Non-Blocking Drift Correction
Raul Taranco, Antonio González
cs.AR
Abstract
Continuous Vision (CV) systems underpin real-time applications such as autonomous driving and augmented reality, where latency, throughput, and energy are tightly constrained on mobile platforms. Modern CV SoC pipelines, however, still serialize image capture and processing, leading to high end-to-end latency. Prior work reduces this latency by predicting future frames and running pixel-domain backend inference speculatively, but incorrect predictions force re-execution on real frames, increasing energy and complexity. We present MVP, a motion-predictive speculative vision pipeline that operates entirely in the motion domain. Instead of forecasting full images, MVP predicts future motion vectors and uses them to extrapolate perception results from previously processed frames before the next frame arrives. A lightweight hardware extension in the Image Signal Processor (ISP) reuses existing motion-estimation logic to predict motion with minimal area and energy cost. MVP introduces a scheduling model that treats motion extrapolation as the default path, while full backend inference runs periodically in the background for drift correction off the critical path. It also supports optional frontend scaling, allowing the system to reduce sensor sampling under low or predictable motion to save energy. We evaluate MVP on object detection, demonstrating up to 66.8% reduction in tail latency and 46% energy savings, at a small accuracy cost.
cs.AR / 5 / 2609.28358
MicroQonv: Reshaping Convolution Tensors for Efficient Microscaling in Training and Inference
Romain Facq, Sami Ben Ali, Olivier Sentieys
cs.AR · cs.AI
Abstract
Microscaling quantization techniques are increasingly used to represent neural network parameters with 8 bits or fewer while preserving near-full precision accuracy. However, applying these methods efficiently in convolutional layers is not straightforward. A naive approach transfers full-precision weights and activations to processing units and quantizes each tensor twice, resulting in much more memory movement than expected. Additional overhead comes from the activation tensors, whose sizes grow substantially because of the im2col transformation applied before quantization. We propose MicroQonv, a way to combine microscaling with convolutional layers' forward and backward operations by quantizing each tensor only once and quantizing the activation tensor before applying a modified version of im2col: channel-batch-first im2col. MicroQonv reduces the quantization cost by a factor of $\times2$ for weights and gradients, and by up to $\times9$ for activations, at a negligible accuracy cost. It reduces memory movement and storage by up to $\times7.53$ compared to their full-precision counterparts. This way, MicroQonv reduces microscaling-quantized activation memory movement by $\times3.5$ for state-of-the-art object detection models YOLOV8nano and $\times2.2$ for YOLOV26nano. It also enables 4-bit microscaling in a quantized latent replay strategy for continual learning at the edge, improving accuracy by +5.7% to +11%.
cs.AR / 6 / 2609.26970
Bringing Chip Tapeout Into University Education
Luca Pezzarossa, Martin Schoeberl, Matti Käyrä, Nooshin Nosrati, Matthias Bo Stuart, Timo D. Hämäläinen, Jean-Max Dutertre, Michael Pehl
cs.CY · cs.AR
Abstract
Providing students with experience from chip specification to fabricated silicon can strengthen chip-design education, but integrating tapeout into regular teaching is difficult to scale, especially across universities with different curricula, schedules, regulations, and technical infrastructures. This article discusses the implementation experience and lessons learned from a cross-university approach developed within the Edu4Chip European project. The approach combines aligned learning outcomes, local Master's implementations, and a shared design-to-silicon framework across five European universities.
cs.AR / 7 / 2609.27162
Agentic-IC3: Enabling Semantic Proof Search in IC3 Model Checking
Yu-Wei Fan, SooHyuk Cho, Aarti Gupta, Sharad Malik
cs.LO · cs.AR
Abstract
IC3 is a state-of-the-art algorithm for hardware model checking that proves safety properties by incrementally constructing an inductive invariant consisting of a set of lemmas. Its effectiveness depends on generalization heuristics that identify useful lemmas and guide proof search. However, many leading IC3 hardware model checkers operate on lowered, bit-level representations, where high-level design relationships are difficult to exploit for generalization. Those operating at a higher level remain limited in exploiting high-level design structure and semantics. We present Agentic-IC3, built on Pono's word-level model-checking infrastructure, which integrates a language-model agent into IC3 to guide semantic proof search using register-transfer-level (RTL) design information. The framework exposes an agent-oriented interface to a persistent IC3 backend, allowing the agent to interact with an explicit, evolving proof state throughout verification. Across successive proof obligations, the agent relates intermediate proof states and solver feedback to the RTL and proposes high-level lemmas through both SAT and UNSAT generalization. Beyond generalization, the agent can introduce derived observation signals to express design relationships succinctly and obtain more informative feedback, and backtrack to revise proposals that lead to unproductive proof branches. The backend checks proposals before updating the proof state, preserving soundness and providing feedback for further reasoning. On a suite of 14 benchmarks spanning security information-flow verification and functional verification of communication protocols, processors, and functional units, Agentic-IC3 solves 10 cases within a one-hour timeout, including 4 unsolved by all three evaluated baselines: rIC3, Pono-IC3Bits, and A-IC3.
cs.AR / 8 / 2609.27456
Precision and resource scaling of real-time flux distortion compensation for superconducting quantum control
Qi Zhou, Zi-Hao Mei, Peng Duan, Peng Wang, Liang-Liang Guo, Hao-Ran Tao, Wei-Cheng Kong, Hui Yang, Guo-Ping Guo, Zhao-Yun Chen
quant-ph · cs.AR
Abstract
Real-time waveform generation supports dynamic quantum circuits without pre-storing complete waveforms for every execution path. However, long-lived distortions in flux-control lines degrade gate fidelity, requiring compensation to account for the actual pulse history. A frequency-domain inversion and time-domain fitting method is proposed for resource-efficient real-time flux distortion compensation. The method fits the reconstructed compensation impulse response with a compact hybrid infinite impulse response (IIR) and finite impulse response (FIR) filter. Look-ahead parallelization enables this filter to process synthesized waveforms at 1.2GSa/s on a field-programmable gate array (FPGA). Two-qubit cross-entropy benchmarking shows that real-time IIR filtering achieves a median controlled-Z Pauli fidelity close to the software-reference value of 99.57%. Numerical analysis and FPGA synthesis indicate approximately logarithmic growth in hardware resource use with compensation timescale. Extending compensation from microsecond to hundred-microsecond timescales increases look-up table (LUT) and digital signal processing (DSP) resource use by only about 14% and 4%, respectively, while maintaining a relative arithmetic error below $10^{-4}$. This work provides a scalable hardware foundation for high-fidelity flux control in dynamic superconducting quantum circuits.
密码学与安全 (cs.CR)
31
cs.CR / 1 / 2609.26893
ACTS: A multi-tier benchmark evaluating LLM cipher identification under controlled blind conditions
Youssef Hamdi Zafaan Ibrahim, Mohammed Khalaf Salama
cs.CR
Abstract
We introduce ACTS (Artifacts in Cipher Testing Suite), a reproducible benchmark that isolates cryptanalytic ability through tiered metadata deprivation (Tier-1: full metadata; Tier-2: filename only; Tier-3: completely blind) and tests forced reasoning (Tier-5: chain-of-thought, code-as-reasoning, self-correction) on ciphertext alone. A 10-configuration ablation study on 7,000 files, using a single 70/30 train-test split for feature-removal analysis, provides additional evidence at scale. Live API inference on a v2b corpus with fully randomised padding (PKCS7, ISO 10126, and ANSI X9.23 selected per file), unique CSPRNG keys, and unique plaintexts (127 files per model, 381 Tier-1 records, 380 Tier-3 records across three cloud systems, supplemented by 254 autonomous agentic evaluations in Tier-4) yields a combined Tier-3 accuracy of 30.8%, only modestly above the 14.3% random baseline for seven-way classification. The corresponding combined metadata-dependency gap is 40.9 percentage points (Tier-1: 71.7% vs. Tier-3: 30.8%). Against a classical Random Forest (69.2% on 7,000 files, trained on engineered byte-level features), the observed live gap is 38.4 percentage points. Because this comparison spans different input representations and training paradigms, the gap should be interpreted as an overall capability difference rather than a clean factorial decomposition. Six findings are reported: (1) Metadata dependency remains large; (2) Scaling failure under blind conditions; (3) Forced reasoning is epiphenomenal; (4) The observed live capability gap exceeds earlier heuristic estimates; (5) The signal is primarily structural rather than statistical; (6) Heuristic invariance versus ML fragility reveals different failure modes.
cs.CR / 2 / 2609.26900
Ajar: Measuring Open Privilege in Agent Defenses
Reshabh K Sharma, Linxi Jiang, Shuo Chen, Zhiqiang Lin
cs.CR · cs.AI · cs.SE
Abstract
A language model agent acts through the tools it is given. The data it reads while working on a task can redirect what it does with those tools. A growing set of techniques for safe and secure agent execution therefore sits between the agent and its tools, aiming to enforce access control, information flow or isolation at that boundary. Today these techniques are evaluated on agent-security benchmarks built around indirect prompt injection. Those benchmarks judge a defense by how far it brings the number of successful attacks down while preserving the agent's utility. A defense is judged only on the agent's execution. It can score well on both metrics while holding open a transfer, a deletion or a broad read that no task needed. Ajar measures that open privilege directly using the existing benchmarks. It attaches to an agent-security benchmark that already exists and reuses the tasks, tool schemas, reference solutions and goal states that benchmark uses to grade its own runs. For each benign task it builds candidate tool calls the task does not need, so allowing one is privilege left open. These calls are presented to the defense at every point where the agent could act. We evaluate Ajar by attaching it to AgentDojo, where open privilege becomes a third axis beside the existing attack success and benign utility. We run it on five defenses: Progent, CaMeL, AC4A, Permission Assistant, and Claude Code's Auto mode. We observed that they leave widely different amounts of privilege open. Two defenses leak by almost the same amount yet differ widely in the benign tasks they finish, and one defense buys part of its tightness by refusing calls its tasks were entitled to make. This open privilege cannot be derived from the measured attack success or benign utility. The source code of Ajar is available at https://github.com/reSHARMA/Ajar.
cs.CR / 3 / 2609.26990
Topological Signatures of Cyber-Attack Classes in Natural Visibility Graph Representations of Network Traffic
Ali Melih Kanca, Ilker Turker
cs.CR · cs.AI
Abstract
Natural Visibility Graph (NVG)-based representations provide a promising approach for capturing structural patterns in sequential network traffic. However, whether different cyber-attack classes exhibit distinctive topological signatures in such representations remains insufficiently understood. This study investigates the discriminative and structural characteristics of NVG-based network traffic representations using the CSE-CIC-IDS2018 dataset. Seventy-six numerical traffic features were independently transformed into NVGs within overlapping frames of 40 observations, and ten graph-theoretic metrics were extracted from each graph, resulting in 760 topological descriptors per frame. The discriminative capability of these representations was evaluated using a multi-branch convolutional neural network (CNN) with stratified five-fold cross-validation. The model achieved an average accuracy of 96.20% and a Matthews correlation coefficient (MCC) of 0.9566. To characterize class-specific topological differences, Kruskal-Wallis and Mann-Whitney U tests were combined with Benjamini-Hochberg false discovery rate correction and effect-size measures. Of the 10,640 attack-versus-benign comparisons, 7,777 (73.1%) remained statistically significant after FDR correction, with 4,844 exhibiting large Cliff's delta effects. The strongest global differences were predominantly associated with backward-traffic and packet-length-related features combined with connectivity, clustering, and centrality measures. These findings indicate that NVG-derived representations can provide strong discriminative capability while revealing class-dependent topological patterns associated with different cyber-attack classes.
cs.CR / 4 / 2609.27052
Improving Service Availability in KubeEdge-Based Architectures Using Lightweight Intrusion Detection
Harrol Ndjeudji Kuibou, Mostafa Anouar Ghorab, Mohamed Aymen Saied
cs.CR · cs.SE
Abstract
The increasing adoption of the Internet of Things (IoT) and cloud computing has accelerated the evolution of edge computing paradigms [1]. Industry forecasts estimate that the number of connected IoT devices will reach approximately 50 billion by 2030, following an estimated 38 billion connections by 2025 [34], resulting in an unprecedented growth in data generation. This trend necessitates efficient, scalable, and secure data processing mechanisms at the network edge. Consequently, ensuring the reliable management and protection of IoT applications and devices has become a critical challenge. In this context, KubeEdge extends cloud-native capabilities to edge environments, enabling distributed orchestration while introducing new security concerns. This paper investigates the security of container images in IoT-driven and distributed edge architectures. Specifically, we analyze the impact of major security threats, including Denial of Service (DoS) attacks and malicious container deployments, on the availability and operational stability of KubeEdge-based systems. To address these challenges, we propose a lightweight Recommended Intrusion Detection Rule Set (RIDRS) tailored for resource-constrained edge environments. The proposed approach improves system resilience by enabling timely detection and mitigation of security threats. We define system stability as the ability to maintain consistent operational behavior and to recover autonomously under adversarial conditions. Experimental results demonstrate that RIDRS significantly reduces system downtime and enhances service availability, particularly in scenarios involving code injection and malicious pod deployment attacks.
cs.CR / 5 / 2609.27090
Divide and Doubt: Diverse Distributed Poisoning for Retrieval-Augmented Generation
Tianhao Chen, Yuhan Wei, Weifei Jin, Zhengyuan Jiang, Yuepeng Hu, Neil Zhenqiang Gong
cs.CR
Abstract
Multi-passage corpus poisoning often repeats one target claim across similar documents, creating correlated lexical and semantic patterns that similarity- and conflict-aware defenses can suppress jointly. We introduce DnD (Divide and Doubt), a targeted attack based on two principles: distributing support for the target answer across stylistically diverse passages, and including a passage that casts doubt on evidence for the reference answer. The first disperses poison-passage representations in embedding space, while the second strengthens target adoption when multiple poisoned passages are retrieved. We evaluate DnD on two open-domain QA datasets across three LLMs and nine RAG configurations, under both black-box and white-box access to the retriever. Across these settings, DnD matches or outperforms prior attacks in most configurations, with its largest gains against clustering- and conflict-aware defenses.
cs.CR / 6 / 2609.27100
Cryptographic Security Is Not Enough: Privacy Gaps in the Renegade Decentralized Dark Pool
Prerna Arote, Adrian Saiz, Oriol Saguillo, Lucianna Kiffer
cs.CR
Abstract
Dark pools are designed to provide pre-trade privacy, liveness, and post-trade confidentiality - concealing order flow before execution and limiting information leakage after. Decentralized dark pools, such as Renegade, aim to replicate these properties without custodial risk, using secure multi-party computation (MPC) and zero-knowledge proofs for private order matching and verifiable settlement. We show that Renegade's cryptographic guarantees do not deliver these dark pool properties in practice. MPC-with-abort ensures correctness but not fairness: a party may learn the match result and abort without penalty, breaking pre-trade privacy. We demonstrate that the protocol's discovery layer further leaks trading intent before MPC even begins, and that sustained probing via selective abort can probabilistically reconstruct counterparty order history, threatening post-trade confidentiality. We also show that the absence of input-consistency checks prior to MPC execution enables a griefing attack using invalid state commitments requiring no real token holdings that continuously locks honest users' wallets and wastes compute, breaking liveness under sustained conditions. We further analyze over 700,000 Renegade transactions on Base and probe the P2P layer, finding that the network is effectively centralized: 88% of traffic routes through a handful of relayers, with only four nodes sustaining the P2P layer. Since relayers hold their users' wallet state in plaintext, this concentration means the system operates as a centralized orderbook in practice - reproducing off-chain the information asymmetry that dark pools are designed to eliminate. Together, our results show that cryptographic privacy does not imply dark pool security: pre-trade privacy, liveness, and post-trade confidentiality each require additional protocol-level guarantees beyond MPC correctness.
cs.CR / 7 / 2609.27124
When Clients Are Orchestrated: Strategic Gradient Manipulation to Defeat Federated Learning Servers with Efficient Defense
Mohamed Shaaban, Ahmed Abdelnaby, Mohamed Elmahallawy
cs.CR · cs.AI
Abstract
Federated Learning enables decentralized model training by exchanging model updates--rather than raw data--with a central parameter server (PS). While most of the existing defenses primarily assume static or independently acting adversaries, we reveal a new class of dynamically adaptive attacks that systematically bypass such protections. We propose Fed-ADR, a holistic attack framework in which a malicious orchestrator server (OS) dynamically coordinates a heterogeneous set of adversarial clients, including both targeted and untargeted attackers. Through real-time coordination by the OS, malicious clients strategically adapt their gradient updates to evade defenses deployed by the PS, while either severely degrading global model performance or steering training toward adversarial objectives.To mitigate this threat, we offer a detection mechanism that estimates each client's true gradient from historical updates, enabling real-time detection of coordinated malicious behavior without additional overhead. We further introduce an in-situ recovery mechanism that restores global model performance without restarting training, preserving convergence and minimizing recovery time. Comprehensive experiments on MNIST, Fashion-MNIST, and CIFAR-10 benchmark datasets demonstrate that Fed-ADR's attack scheme can reduce global accuracy from over 90% to below 10%, bypassing several state-of-the-art defenses. When our detection and recovery modules are employed, they identify malicious clients and restore accuracy to over 90% within a few rounds, at a substantially lower cost than retraining from scratch--achieving a reduction of at least 20x in computational overhead.
cs.CR / 8 / 2609.27202
Reliable Federated TinyML Deployment for IoT Security
Younsoo Park, Seokhyoen Bae, Shasi Kumar Ramachandran Prabhu, Suman Saha, Peilong Li
cs.CR · cs.LG
Abstract
The growing deployment of Internet of Things (IoT) devices has increased the need for privacy-preserving intrusion detection systems that operate directly on resource-constrained hardware. Federated Learning enables collaborative model training without sharing raw data, but conventional federated models are often too large and unstable for deployment on microcontroller-class devices. TinyML techniques enable compact neural networks but are typically designed for inference-only workloads. This work investigates combining Federated Learning with TinyML-based model compression for intrusion detection in IoT environments. We evaluate compression strategies including knowledge distillation, structured pruning, and quantization within a federated training pipeline. Preliminary results show that training stability plays a critical role in federated TinyML systems. In particular, server-coordinated cosine learning-rate scheduling improves Attack Recall from 46.7% to 93.85% while enabling substantial model compression and efficient edge deployment. These findings provide insights for designing lightweight and privacy preserving intrusion detection systems for IoT devices.
cs.CR / 9 / 2609.27258
Anti-Localization Uplink Communications in Satellite-Terrestrial Systems
Ranran Sun, Bin Yang, Yulong Shen, Yuanyu Zhang, Xiaohong Jiang
cs.CR
Abstract
This paper investigates the anti-localization uplink communication in a satellite-terrestrial system, where a ground transmitter Alice communicates with a legitimate satellite receiver Bob in the presence of multiple cooperative adversarial satellites attempting to localize Alice with the time difference of arrival (TDOA) technique. Specifically, we propose a cooperative jamming-based scheme for such anti-localization communication,in which Alice exploits the superposition coding with power allocation to simultaneously transmit information/jamming signals for communication with Bob and for confusing signal detection/TDOA measurement at adversarial satellites, while Bob employs the combining vector technique to enhance the desired information signal and also suppress the jamming. We define a localization error probability (LEP) metric to jointly depict both the impacts of signal detection and TDOA measurement on localization performance, and then develop a theoretical framework for the LEP modeling under the proposed scheme. We further explore the joint optimal design of jamming coding and power for LEP maximization, subject to the constraints of AliceBob communication reliability and Alice's transmit power. An effective sample average approximation method is also provided to tackle this non-convex optimization problem. Finally, extensive numerical results are illustrated to validate our theoretical models and demonstrate how the cooperative jamming helps to provide an anti-localization guarantee while ensuring communication reliability
cs.CR / 10 / 2609.27300
SoK: You Find What You Seek: Rethinking Oracles, Guidance, and Input Generation in Hardware Fuzzing
G Abarajithan, Zhenghua Ma, Cristian Tirelli, Andres Meza, Francesco Restuccia, Cynthia Sturton, Ryan Kastner
cs.CR · cs.AR
Abstract
Hardware fuzzing is an active area in security verification research, yet its industrial adoption remains in its early stages. This SoK examines which lessons from software fuzzing carry over to hardware and where unique approaches are needed. By analyzing 52 fuzzers across RTL/IP, CPU, NoC, and SoC designs, we introduce an analytical framework that frames verification as a bounded search. This search is defined by its objective, oracle, guidance, input generation, target abstraction, and budget. Consequently, a campaign only uncovers failures it can effectively reach, recognize, and prioritize before exhausting its resources. We distinguish two roles for hardware fuzzing: (1) augmenting constrained-random verification (CRV) via feedback-guided coverage and (2) directed adversarial testing based on threat models and security specifications. Through our framework, we identify what each campaign can observe and generate, providing a basis for assessing the evidence behind reported results. Our analysis suggests that mainstream adoption of hardware fuzzing will require reusable interfaces, target-specific verification assets, reproducible evaluations, and transparent reporting of cost and user effort.
cs.CR / 11 / 2609.27311
Multi-View Fusion for Encrypted C2 Detection: A Leakage-Controlled Measurement Study of Evaluation Pitfalls
Hoang-Huy Nguyen-Huu, Van-Tri Phan, Khuong Nguyen-An
cs.CR · cs.AI
Abstract
Command-and-control (C2) traffic increasingly hides within TLS, so defenders now apply machine learning to traffic metadata. Many studies assume that combining two metadata views, namely flow statistics and TLS handshake fingerprints, improves both accuracy and robustness. We tested this assumption on 17,577 TLS flows from 62 real Cobalt Strike captures. Our evaluation removes the data leakage that leads to overly optimistic reported scores. We report three findings that matter more than the fusion result itself. First, an incorrect preprocessing step increases the F1 score by 0.28. This step computes the frequency encoding across the entire dataset rather than within each cross-validation fold. The increase is about ten times larger than any real effect we measured. Second, both the labels and the behavioral features depend on the destination address. Because of this, the 17,577 flows form only 2,132 independent groups, and the positive rate of 55.1\%, which looks balanced, drops to 4.2\%. Therefore, class balance is just a result of how we analyze the data, specifically whether we count flows or endpoints, and not a real feature of the task. Third, 20 of the 62 captures (32\%) have no TLS flows to any known C2 address, so they contain only benign samples. We checked these captures directly and confirmed that this is a gap in the ground truth, not a labeling error. In this context, fusion beats the best single view by only 0.022 in F1. When an attacker forges both feature surfaces simultaneously, every model performs worse than a simple baseline that always predicts positive (F1 = 0.711). For encrypted C2 detection, the evaluation design is not a preliminary step. It \emph{is} the main result.
cs.CR / 12 / 2609.27357
SAGEGAN: Style-Based Anomaly Detection with Gaussian Embeddings using Generative Adversarial Networks
Thesath Wijayasiri, Kar Wai Fok, Vrizlynn L. L. Thing
cs.CR
Abstract
Malware evolves faster than rule-based and signature-driven detection pipelines. This paper presents SAGEGAN, a benign-only trained malware anomaly detection framework that converts portable executable files into compact three-channel images and models benign structure through style-conditioned adversarial reconstruction. The representation combines Hilbert-mapped byte values, benign-referenced byte-transition surprise, and entropy deviation from benign software. The model encodes each image into a layer-wise style tensor aligned with a seven-stage modulated generator, rather than a single latent bottleneck. A Gaussian style prior, moment-based prior alignment, and latent consistency are used to reduce mismatch between encoded benign styles and the generator's sampled manifold. For interpretation, a deterministic encoder pathway maps each executable to a fixed style tensor, enabling repeatable layer-wise family distance, gradient sensitivity, principal component, and class-behaviour analyses. On a self-collected portable executable corpus containing malware from 214 families, the Gaussian variant achieves 89.76% area under the receiver operating characteristic curve and 88.19% balanced accuracy, while the genome-style variant reaches 88.03% and 84.09%, respectively. Without refitting model weights, benign reference statistics, or decision thresholds, the same checkpoints are evaluated on DIKE, Microsoft BIG 2015, and Lester malware subsets. The results suggest that layer-wise style modelling supports both anomaly ranking and structured post hoc analysis of how malware families depart from the benign manifold.
cs.CR / 13 / 2609.27367
Seal, Then Sample: Sampled Layerwise Proofs for Verifiable LLM Inference from GPT-2 to 70B
Youki Lim, Sam Yong
cs.CR · cs.IR
Abstract
Verifying outsourced language-model inference requires a precisely identified computation and an audit whose cost a service can afford. We present Sampled Layerwise Proofs (SLP), a protocol and prototype that commits the boundary activations of every chunk of an inference trace, absorbs all commitments before any challenge is drawn, and then proves a verifier-selected subset of chunks together with the chunks that bind the prompt and the answer. Audit coverage becomes a runtime parameter over one set of commitments: on a TinyLlama-1.1B trace, proving seven of 47 chunks takes 22.0% of the time and 6.8% of the proof size of proving all 47. Because proof cost is dominated by weights rather than tokens, SLP packs concurrent requests into one trace under a block-diagonal causal mask and binds the prompt and answer of each request to its slot. Twelve packed requests are proved in 181.9 s, 6.5 times less than twelve separate proofs at the measured single-proof cost, and a simulated service proves twelve requests at 30.6 s per request with 0.6 s of verification each, rejecting a tampered answer. Disk-backed integer weights and streamed polynomial commitments let a single Llama-2-70B run complete on a 2 TB CPU host: 163 chunks sealed, five proved, a 4.34 MiB proof in 1,259 s, verified in 46.3 s without the weights. The proven object is a fixed-point canonical model; we trace a severe fidelity loss to the residual-stream bit width, repair it with an LLM-aware observer, and measure 84.8-84.9% argmax agreement with the floating-point reference over 334,705 WikiText-2 test positions. The limits are stated as precisely: guarantees cover proven chunks only, a fixed invalid chunk in the 70B setting is covered with probability 3/161, a manifest-only Fiat-Shamir schedule can be ground at 12.5 ms per attempt and needs an externally ordered challenge, and all measurements use a test reference string.
cs.CR / 14 / 2609.27422
RAMP: Reversing Adversarial Perturbations to Strengthen Clean-Label Backdoor Attacks against Malware Detectors
Jinwen Xin, Dongni Zhang, Chenyang Wang, Jianming Fu, Ming Tang, Guojun Peng
cs.CR
Abstract
Deep learning-based malware detectors are commonly updated by fine-tuning on newly collected samples, but this practical update pipeline also creates an attack surface for training-time backdoor attacks. In realistic crowdsourced data collection, however, strict label vetting typically restricts attackers to the clean-label setting, in which poisoned samples must retain benign labels and functionality, making effective backdoor injection substantially harder. We present a new attack perspective based on feature-space manipulation: instead of relying solely on stronger trigger designs or selecting benign samples that are naturally similar to malware, we deliberately construct benign programs whose representations shift toward the malware region before trigger injection, thereby creating stronger feature-label conflicts during training. Based on this insight, we propose RAMP, an attack enhancement method that uses a genetic algorithm to optimize reversed adversarial perturbations under black-box access and then injects them through functionality-preserving binary manipulations. Extensive experiments show that RAMP substantially improves attack effectiveness over trigger-only baselines, with especially pronounced gains at low poisoning ratios, while maintaining accuracy on clean data. Moreover, RAMP can be combined with advanced trigger designs.
cs.CR / 15 / 2609.27424
EVAGE: Autonomous MEV Generation and Adaptation via Multi-Agent Harness
Yan Wen, Zichun Cai, Iliya Mirzaei, Xiaohua Cai, Mohammad Javad Amiri, Haoxian Chen, Chenyuan Wu
cs.CR
Abstract
Maximal Extractable Value (MEV) has evolved into a major economic force in blockchain ecosystems, yet its capture is dominated by experienced teams, and both strategy design and implementation rely on manual expert work that scales poorly across heterogeneous protocols and chains. We present EVAGE, the first fully autonomous multi-agent framework for end-to-end MEV strategy generation and adaptation. Equipped with three specialized operation modes, it automatically discovers novel MEV variants, adapts execution logic across disparate protocols, and ports strategies between chains, including Layer-1 and Layer-2 networks. To avoid inference latency on the critical MEV execution path, EVAGE generates and refines MEV bot code offline rather than making real-time decisions directly. Under the coordination of an orchestrator agent, three specialized subagents collectively implement and repair the full MEV bot workflow via closed-loop diagnostics, eliminating human intervention while producing validated and deterministic Proof-of-Concept implementations. We evaluate EVAGE on over 1.5M blocks from each of Ethereum, Base, and BNB Smart Chain (BSC). On Ethereum, EVAGE uncovers five novel MEV strategy variants, yielding a profit increase of 1.02$\times$ to 15.97$\times$. It also successfully adapts 11 MEV strategies from CPMM to both CLMM and Balancer V2 and ports strategies from Ethereum to Base and BSC, all with less than 60 dollars in LLM token costs.
cs.CR / 16 / 2609.27427
Extracting CNNs in the Unknown-Architecture and Feedback-Agnostic Setting
Jiashuo Liu, Ruijie Ma, Manman Li, Yi Chen, Shaozhen Chen
cs.CR
Abstract
This paper studies the cryptanalytic extraction of convolutional neural networks (CNNs). Existing cryptanalytic extraction attacks on CNNs assume that the network architecture is known, and try to recover model parameters.In this paper, we prove for the first time that the architecture assumption can be removed for CNNs with both max and average pooling. Our core finding is that the spatial geometry of the weight vectors recovered by existing parameter-recovery attacks naturally leaks the architecture. We formalize this geometry and establish its correspondence with the architectural knowledge of a convolutional layer: (1) The sparsity consistency with the convolution receptive field reveals the layer type, the kernel size, and the stride; (2) The numerical consistency with the kernel parameters reveals the padding mode and the output-channel number; (3) The structural consistency with the pooling operation reveals the pooling type, the window size, and the stride. Although the recovered vectors are obtained using different methods in the raw-output and hard label settings, their spatial geometry remains the same. Therefore, our architecture recovery is feedback-agnostic: combined with a parameter-recovery attack, it yields a complete cryptanalytic extraction framework that recovers both the architecture and the parameters in the black-box setting. Extensive experiments, including both layer-wise and end-to-end ones, on a wide range of CNNs demonstrate that simultaneously recovering the network architecture and the model parameters is practical.
cs.CR / 17 / 2609.27452
Issuer-Sovereign Agentic Payments
Dishant Sharma, Rajneesh Kaushal, Ashu Kanaujia
cs.CR · cs.AI · cs.SE
Abstract
AI agents are beginning to make real payments. Current approaches let an agent pay by relying on a credential provider that, in the approaches deployed today, typically sits outside the cardholder's bank. The spending rules are then enforced by the card network or that provider, and not by the bank itself. This leaves the issuing bank, which carries the financial risk, with little direct control at the moment a payment happens. This paper describes Issuer-Sovereign Agentic Payments, a method that keeps that control with the issuer. The cardholder approves a spending rule once, and the bank's own authentication component records it. Later, when the agent pays a specific merchant, the bank checks the merchant against the approved rule and generates the card authentication value only if the merchant is allowed. The payment then travels the normal card rails and is validated by the issuer, with no extra dependency introduced at execution.
cs.CR / 18 / 2609.27542
Control-Token Injection Suppresses Chain-of-Thought and Defeats Reasoning-Based Oversight in Tool-Using Agents
Muhammad Usama, Khair Un Nisa, Summer Yeoreum Jung
cs.CR
Abstract
The safety of a tool-using language model agent is usually treated as a property of the model alone. We give controlled, full-precision evidence that it is instead a joint property of the model and the software that renders its chat template and parses its tool calls, the decoding harness, and that both halves are attackable from untrusted input. On the released gpt-oss-20b reasoning model under its published tool sandbox, appending a single string of the model's own channel-control tokens to a user message makes the tokenizer render a reasoning turn that is already complete, so the model writes no chain-of-thought and proceeds directly to the tool call. Across forty tasks the model already completes, the reasoning channel falls from a mean of 52.5 tokens to zero on every trial while the http.post still fires on every trial. A rule monitor and a cross-family language-model monitor detect the unsafe request on all plain trials and no forged trials, and on overtly malicious requests the attack converts 39.6% of the model's refusals into completed exfiltrations. Separately, whether an identical tool-call generation fires is decided by the harness parser, not the model: a truncation-tolerant regular expression fires a call whose closing token is missing while a strict one drops it, and two parsers shipped for the Gemma agent give opposite outcomes on identical greedy generations, firing on all twenty-four trials and on none. We show the suppression can be delivered indirectly and characterize its dependence on the chat template across two more reasoning models, and we evaluate input sanitization, parser hardening, and empty-reasoning detection as defenses; flagging an absent trace catches the basic attack but not an adaptive benign decoy. All measurements use greedy decoding on publicly released models. Code and per-trial logs: https://github.com/Usama1002/deleting-the-trace
cs.CR / 19 / 2609.27567
CCR: Towards a Common, Quality-Gated CACAO Integrations Registry for European Cybersecurity Automation
Mateusz Zych, Vasileios Mavroeidis, Gudmund Grov
cs.CR
Abstract
Standardised, machine-readable cybersecurity playbooks provide a basis for portable, shareable, and reusable incident-response logic. OASIS CACAO provides a vendor-neutral representation for such playbooks, but not the product-specific integration artefacts needed to invoke external products and services. We introduce the Common CACAO Registry (CCR), an open, provenance-aware registry of CACAO HTTP-API connector envelopes. Each envelope captures an API operation's command, inputs, target, authentication-related information, provenance, validation evidence, and maturity metadata. CCR is \emph{quality-gated}, with acceptance requiring both CACAO v2 schema validity and a mean back-validation score of at least 0.8 against the source OpenAPI operation, while a six-level maturity model records progressively stronger evidence and distinguishes gate acceptance from operational readiness. To seed CCR, we develop a hybrid OpenAPI-to-CACAO pipeline. Deterministic code extracts source-derived interface facts, generates identifiers, wires cross-references, and validates structure, while a constrained LLM provides bounded semantic enrichment, including action naming, authentication interpretation, and CACAO activity annotation. Evaluation across eight security APIs yields 713 CACAO-schema-valid envelopes with a mean back-validation score of 91.5\%, of which 675 produce well-formed, dispatchable HTTP requests in a local harness. Comparison with a deterministic rule-based baseline shows that mechanical API structure is preserved more reliably through rule-based translation, while the LLM contributes bounded semantic enrichment, most notably CACAO activity annotation. Together, these results support CCR as reusable integration infrastructure for CACAO action steps and as an initial foundation for a broader common European registry.
cs.CR / 20 / 2609.27624
Agent Name Collision Attacks in Multi-Agent Systems
Adithyan Arun Kumar
cs.CR · cs.MA
Abstract
Multi-agent hosts turn remote Agent Cards into local agents, tools, workflow targets, and broker routes. A2A defines the card's name as human-readable metadata, not as a stable identity, and specifies no collision semantics. The security failure begins when a host nevertheless uses that remote name as a local routing identifier. We traced registration through dispatch and ran isolated regression tests at seven pinned open-source revisions. Six client-style integrations selected an attacker-controlled peer's client or loopback endpoint for a request addressed to a trusted peer's name. A seventh, brokered implementation collapsed both peers onto one name-derived route; queue and access-control state determine whether the result is interception or denial. The common result is wrong-peer dispatch, not universal privilege inheritance. Synthetic credential and tool tests found no A-specific credential transfer in the tested client bindings and no direct transfer of A-owned tools. The broker path forwards a caller-configuration object; delegated identity or tokens reach B only if present and B can consume the route. Two other paths expose a later, model-mediated decision rather than direct execution authority. The necessary conditions assign different responsibilities to the protocol, implementations, and deployments. Hosts should route by an origin-bound stable identity, keep names presentational, and reject ambiguous aliases. The evidence establishes a recurring implementation vulnerability class, not a universal A2A protocol exploit or a count of vulnerable deployments.
cs.CR / 21 / 2609.28115
No Place to Hide: An Analysis on Protected Order Flow Sandwich Attacks
Lioba Heimbach, Ozan Solmaz, Burak Öz, Christof Ferreira Torres
cs.CR
Abstract
Front-running has long plagued Ethereum's public mempool, earning it the nickname of a "dark forest", where predators lurk for profitable transactions. In response, Ethereum and other blockchain ecosystems increasingly rely on private RPCs and native protections to shield transactions from adversaries, which we refer to as protected order flow. Yet the effectiveness of these mechanisms in preventing front-running, and what trust assumptions they entail, remain poorly understood. In this work, we conduct the first longitudinal, three-year measurement study of sandwich attacks against protected order flow across six blockchains: Ethereum, Solana, Tron, Base, Arbitrum, and Monad. We introduce detection heuristics that capture wide attacks, both within and across blocks, and filter on bot behavior to distinguish sandwiches from legitimate trading activity. We identify 28.0 million sandwich attacks on Solana, 38,567 on Tron, 30,607 on Ethereum, and 1,889 on Base against transactions intended to be protected from front-running. Reorged blocks expose a further 2,875 Ethereum victims. Unlike conventional public-mempool sandwiches, these attacks rarely occur tightly around their victims and, outside Solana, are carried out by a small number of entities. Our analysis uncovers exposures at every layer: validator- and application-level exposure on Solana, order-flow auctions and reorged blocks on Ethereum, first-come-first-served ordering that fails to prevent latency-based front-running on Tron, and both an RPC bug that exposes pending transactions and predictable victim behavior on Base. These findings show that existing front-running protections can provide substantially weaker guarantees than users expect, highlighting the need for stronger end-to-end defenses against sandwich attacks.
cs.CR / 22 / 2609.28170
Safety-Aware Zero Trust Enforcement for IoT and Cyber-Physical Systems
Alessandro Lotto, Alessandro Brighente, Mauro Conti
cs.CR
Abstract
Zero Trust (ZT) replaces the implicit trust of perimeter-based security with explicit, continuous, context-aware authorization. This shift is particularly relevant to IoT and cyber-physical systems, whose heterogeneous, long-lived, and remotely connected components make persistent trust untenable. Yet their physical coupling complicates ZT adoption: restricting a suspicious component can reduce cyber exposure while removing telemetry or control capabilities required for operation. Existing work mainly models physical harm caused by attacks, with less attention to consequences introduced by enforcement itself. We introduce Safety-Aware Zero Trust (SA-ZT), which treats restriction-induced physical consequences as policy inputs. We map the NIST ZT tenets to nine IoT/CPS convergence strains, distinguish IoT-amplified challenges from those specific to cyber-physical coupling, and derive corresponding operational requirements. SA-ZT extends the NIST ZT Architecture with a Safety Engine and a Telemetry Broker. The Safety Engine selects among admissible responses by jointly considering residual cyber risk and restriction-induced consequences, while the Telemetry Broker mediates raw telemetry visibility and estimator influence. With command-side enforcement, these entities separate raw visibility, automated influence, and state-changing authority, preserving observations for monitoring while constraining their influence on automated control. An IEEE 30-bus case study under false-data-injection attack illustrates how SA-ZT makes cyber containment, telemetry visibility and influence, physical consequences, and authorization timing explicit, providing an implementable and inspectable representation of cyber-physical enforcement trade-offs.
cs.CR / 23 / 2609.28209
Do Electromagnetic Side-Channel Attacks Threaten Electronic Polling Stations? Scenarios and Recommendations
Lucas Brito, Leonardo Teodoro, Pedro Tomaz, Alyson Isaluski, Leandro Hyeda, Antonio Oliveira-Jr, Saulo Queiroz
cs.CR · cs.CY · eess.SP
Abstract
This paper investigates the threat to ballot secrecy in the Brazilian electronic voting machine (UEB) posed by electromagnetic side-channel attacks, also known as TEMPEST attacks. In these attacks, screen content can be reconstructed remotely by intercepting electromagnetic emanations associated with the target device's video signal. This work is motivated by a recent ruling by a Brazilian electoral court concerning an attempt to violate ballot secrecy using electronic equipment. Based on publicly available information about the electoral system, attack scenarios against polling stations are proposed. Experiments using software-defined radio show that the effectiveness of TEMPEST attacks strongly depends on the lack of oversight resulting from public unawareness of the threat. Finally, awareness guidelines are proposed for voters, poll workers, and party representatives to mitigate attack risks within a polling station.
cs.CR / 24 / 2609.28226
Pinpointing Super-Quadratic Quantum Enumeration Speedups: Exact and Certified Evaluation of the Guessing-Moment Exponent under Product-Distribution Advice
Carsten Schubert, Niklas Paskarbeit, Maximilian J. Kramer, Jean-Pierre Seifert, Marian Margraf
cs.CR
Abstract
Grover's algorithm gives an optimal quadratic query advantage for black-box search. In cryptanalysis, however, the search often comes with additional probabilistic advice over the candidates, frequently of product form, e.g. from side-channel leakage on independent key coordinates. Classically, guessing in likelihood order is optimal in expectation. In the quantum setting, Montanaro showed how to achieve an optimal expected query complexity, beating plain Grover on every non-uniform advice distribution (up to a constant overhead factor). What has been missing so far is a finite-size method for evaluating the quantum-classical guessing-moment separation induced by a given advice distribution. We provide such a method for product-distribution advice, thereby sharpening the previous entropy-based estimate of Bashiri et al. We reduce the classical and quantum guessing moments to functionals of the one-dimensional surprisal distribution, obtained for product advice by convolving the per-coordinate surprisal laws. When the surprisals lie on a common arithmetic grid (the commensurate case), the logarithmic moments and hence the speedup exponent can be evaluated as finite sums without discretization error; exponential tilting makes this computation numerically stable. For general product advice, we discretize the surprisals onto a common grid and derive an a-posteriori bound on the resulting binning error. We apply the framework to cold-boot leakage on seeds and block-cipher keys, to template-attack posteriors, and to synthetic i.i.d. Bernoulli posteriors calibrated to residual ranks reported for Keccak side-channel attacks on ML-KEM and ML-DSA. The resulting exponents substantially exceed 2 in several skewed-advice settings, reaching up to 3.97 in these synthetic models, and include cases where the previous entropy-based bound did not establish an exponent above 2.
cs.CR / 25 / 2609.28228
MimicSat: A Reconfigurable Cyber-Physical Testbed For Small Satellite Systems and Cybersecurity Research
Nisha Vinayaga-Sureshkanth, A H M Nazmus Sakib, Mahsin Bin Akram, David R. Silva, Murtuza Jadliwala
cs.CR
Abstract
MimicSat provides a common experimental environment for examining how changes in satellite subsystem behavior propagate to mission outcomes across software-based and hardware-based execution. Its design is motivated by controlled spacecraft cybersecurity studies involving attacks, faults, and defensive responses. In MimicSat, a mission encompasses the spacecraft and ground activities required to achieve defined objectives, and an experiment consists of one or more mission runs used to study selected conditions or interventions. To support such studies, MimicSat offers software-based and hardware-based execution environments that implement the same mission functions and data exchanges, while allowing specific functions to be realized differently. A shared mission definition preserves command and telemetry semantics across environments, and collected observations retain provenance about the originating participants and acquisition paths. As a result, mission behavior can be compared across execution configurations without redefining the surrounding mission. This paper presents the architectural principles of MimicSat, its software and hardware execution forms, and their integrated operation. MimicSat also supports satellite systems engineering, mission operations, resilience studies, and related experimental use cases.
cs.CR / 26 / 2609.28277
Physalia: Redistribution-Resistant Content Protection for Decentralized Storage
Giacomo Giuliari, Karl Wüst
cs.CR
Abstract
In decentralized storage systems, access control is often implemented by encrypting the data before upload and sharing the decryption key with authorized parties. A leaked key, however, makes the data publicly accessible, which lowers the barrier to content piracy below that of traditional systems, where piracy requires redistributing the full data. We present Physalia, an end-to-end access-control system for decentralized storage that secret-shares the data itself, instead of just the key, across multiple servers. Leaking the data then requires transmitting it in full: We formalize this intuition and introduce the redistribution bandwidth an adversary must pay to leak protected content and show that Physalia raises it to the size of the data. Sharing across untrusted servers requires robustness against corrupted shares. We develop a robustness transform that turns any computational secret sharing scheme into a robust one and which is of independent interest. In contrast to existing schemes that rely on error correction, it uses signatures with ephemeral keys and adds only constant-size metadata per share. We show that this transform is secure and evaluate Physalia end-to-end on the Walrus decentralized storage system with an on-chain access policy.
cs.CR / 27 / 2609.28305
A Gmail-Based Phishing Detection Prototype for Nigerian Fintech Emails Using Sender Checks and BiLSTM Classification
Gideon Francis Oghie, Uche Emmanuel Unoke
cs.CR
Abstract
Phishing emails that impersonate Nigerian fintech providers can combine deceptive sender addresses, lookalike links, and locally familiar language. This study presents a Gmail browser extension that integrates sender-domain and URL checks with a bidirectional long short-term memory (BiLSTM) classifier. The extension compares visible sender addresses and links with profiles for eight fintech platforms, obtains a phishing probability from a locally hosted Flask service, and displays a legitimate, warning, or phishing verdict when an email is opened. The BiLSTM classifier was evaluated on 8,943 test messages from a cleaned dataset of 59,622 phishing and legitimate emails. The test confusion matrix recorded 4,308 true negatives, no false positives, one false negative, and 4,634 true positives. These counts correspond to 99.99% accuracy, 100.00% precision, 99.98% recall, and 99.99% F1 score. Tokenized sequence analysis identified 5.79% overlap between the training and test sets, which may inflate performance estimates for independent messages. A Gmail demonstration showed the integrated extension producing user-visible verdicts, although the complete system was not evaluated on a labeled test set. The findings establish the feasibility of the implemented prototype while leaving its end-to-end detection performance and generalization to unseen attacks open for further evaluation.
cs.CR / 28 / 2609.28239
ODPure: Backdoor Purification for Object Detection via Ensemble Corruption Consensus
Li Zeng, Mingcheng Duan, Longfei Fan, Hangtao Zhang, Xianlong Wang, Yanchun Li, Xia Wen, Leo Yu Zhang
cs.CV · cs.CR
Abstract
With the development of applications like autonomous driving, object detection has gained significant attention, while also highlighting critical vulnerabilities like backdoor attacks that severely compromise model integrity. Specifically, such attacks involve altering the categories of objects (i.e., object misclassification), removing bounding boxes (i.e., object disappearance), or generating bounding box proposals for non-existent objects (i.e., object generation) when a predefined trigger is present in the input. Although backdoor defenses for image classification are well-established, the research for object detection remains comparatively underexplored. Existing defenses address these threats by scanning outputs or models for potential backdoors but require discarding either malicious data or models. This remedy fails to enable a continuous and accurate perceptual stream for the object detection pipeline. To address such limitations, we propose ODPure, a novel input-stage black-box defense for object detection, which is based on input purification that ensures stable perception flows. Tailored to the dense prediction nature of object detectors, our Corruption-Reconstruction-Selection (CRS) paradigm operates by neutralizing triggers through a diverse portfolio of corruptions to generate a massive pool of redundant proposals, then recovering fine-grained structural cues via generative priors, and finally employing voting to reach a consensus on the resulting detections. Comprehensive experiments demonstrate that our method provides robust defense against diverse backdoor attacks and trigger types while preserving baseline accuracy. Our code is available at https://github.com/Alex66366/ODPure.
cs.CR / 29 / 2609.27402
Credible AUctions via MPC Gadgets: Bounding Information Leakage Under Abort
Matheus Venturyne Xavier Ferreira
cs.GT · cs.CR · econ.TH
Abstract
The design of credible auctions---mechanisms where a revenue-maximizing auctioneer has no incentive to deviate from the protocol---faces a fundamental cryptographic barrier when the auctioneer controls shill bidders. While a natural approach is to use Secure Multi-Party Computation (MPC) to remove the trusted auctioneer, the impossibility of fair coin flipping of Cleve (1986) implies that monolithic MPC protocols grant the auctioneer a "free option": they can learn the auction's outcome and unilaterally abort if the revenue is unsatisfactory. Cryptographic commitments with ex-ante penalties mitigate this abort asymmetry, but no finite penalty suffices for heavy-tailed distributions. We circumvent this barrier by introducing the MPC Decomposition Principle. Rather than encrypting the entire mechanism, we use MPC strictly as an information-restriction tool. We isolate the winner determination problem into a minimal MPC gadget that computes and reveals the winner's identity but no payment information. This qualitative restriction mathematically bounds the information leaked upon an abort. By combining this gadget with sequential revelation and finite economic penalties, we design the Sequential Revelation Auction (SRA). We prove that bounding the information leakage strictly bounds the value of the free option, showing that a penalty of $k \geq \sum_{i=1}^n Rev(F_i)$ is sufficient for credibility, and tight: for equal-revenue distributions, every smaller penalty admits a profitable deviation. Using constant-round MPC, the SRA resolves an open question of Akbarpour and Li (2020) and Ferreira and Weinberg (2020) by providing a constant-round, incentive-compatible, revenue-optimal credible auction for all product distributions with vanishing revenue tails
cs.CR / 30 / 2609.28297
Contraction and Statistical Inference under Privacy for Uniformly Bounded Distributions
Leonhard Grosse, Sara Saeidian, Tobias J. Oechtering, Mikael Skoglund
cs.IT · cs.CR · math.ST
Abstract
We investigate $c$-interior pointwise maximal leakage (PML) as a tool for contraction analyses and disclosure control. Based on the strong adversarial threat models from maximal leakage, $c$-interior PML generalizes local differential privacy (LDP) to data-generating distributions with densities uniformly bounded away from zero by $c>0$. Viewing $c$-interior PML as an algebraic constraint on a kernel yields more flexible (and often tighter) contraction analyses than standard LDP. We provide tight bounds on the Dobrushin coefficient, and bound the contraction coefficient of the Hockeystick-divergence. We further derive strong data processing inequalities on $f$-divergences under $c$-interior PML constraints when the input distributions to the divergence are restricted to be in the $c$-interior. These results extend beyond the regime of pure LDP to cover a larger class of kernels, including, e.g., arbitrary stochastic matrices. We apply the results to minimax theory and provide asymptotically optimal strategies under $c$-interior PML constraints for binary hypothesis testing and mean estimation. The results show that disclosure control with PML allows analysts to reason about systems in a more differentiated manner: For example, it allows us to quantify the privacy leakage of deterministic systems, and can give precise adversarial guarantees with respect to arbitrary distributional assumptions. Interestingly, a recurring theme in the disclosure analyses is that if the privacy problem is relatively regular (if the density bound $c$ is large), private inference can be possible without incurring any additional cost in terms of sample complexity.
cs.CR / 31 / 2609.28378
ForgetMimic: Motion Unlearning for Reinforcement Learning Humanoid Control
Xukun Luan, Zhongxiang Lei, Chen Gong, Shaowei Li, Yuanguo Bi, Jinyan Liu
cs.RO · cs.CR · cs.LG
Abstract
Humanoid control, leveraging human demonstrations, has achieved diverse, agile, and natural locomotion behaviors through reinforcement learning (RL). While this paradigm has yielded remarkable performance in physical humanoid control, how to eliminate specific motions from learned policies remains insufficiently explored. Addressing this issue is motivated by pressing safety and privacy concerns: the removal of malicious, poisoned, or suboptimal motions, as well as copyright-protected motions subject to the right to be forgotten under regulations such as the GDPR, is of critical importance. To this end, we propose {ForgetMimic}, the first motion-level unlearning method designed specifically for physical-world humanoid control. The core idea of ForgetMimic is as follows: given a policy $π_θ$ trained on $N$ motions, our method degrades performance on a target subset of $K$ motions while preserving the effectiveness of the remaining $N-K$ motions. Furthermore, we identify and resolve two key training mechanisms in robot control that lead to unlearning failure. We conduct extensive experiments on the Unitree G1 and H2 humanoid robots across 12 motions, including Dance, Fight, Flip, and others. Experimental results demonstrate that ForgetMimic effectively eliminates memory of designated motions while maintaining the normal operation of all other motions.