Daily Research Digest
arXiv Papers
2026-09-16
378
Papers
8
Categories
80
Translated
收藏清单 0
精选 · Favorites
80
cs.AI / 1 / 2609.15989
Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection
腐化的计划,干净的痕迹:利用计划注入规避思维链监控
large language model
大语言模型相关
Abstract
Chain-of-thought (CoT) monitoring is a safety strategy where the reasoning of a large language model "actor" is inspected by a "monitor" (often another language model) for signs of unsafe planning, deception, or misalignment. We find that planting harmful but benign-sounding reasoning in the actor's context can steer it to perform adversarial actions while evading monitors, an attack we term "plan injection". We initially discover this attack in the multiple-choice question-answering monitorability setting proposed by Lanham et al. (2023), using the investigator-agent elicitation framework of Li et al. (2025). We generalize the attack and show that the discovered behavior scales to harder tasks (achieving 25-33% monitor evasion rates across different monitorability benchmarks) and larger models such as DeepSeek-R1. Across the settings we study, actor models not only follow injected plans but also paraphrase them as their own reasoning, without explicit attribution to the injections. Finally, we find cases where extra monitor resources cause harm - giving the monitor access to the injected plan drops detection by as much as 50% in the Bio-Math task and in a case study on monitor reasoning budget, we find transcripts where additional thinking tokens are spent rationalizing the injected plan rather than flagging it.
Chinese Translation
思维链(CoT)监控是一种安全策略,其中由“监控器”(通常是另一个语言模型)检查大型语言模型“行动者”的推理,以寻找不安全规划、欺骗或失准的迹象。我们发现,在行动者的上下文中植入有害但听起来无害的推理,可以引导它执行对抗性行为,同时规避监控器,我们将这种攻击称为“计划注入”。我们最初是在 Lanham 等人(2023)提出的多项选择问答可监控性设置中,使用 Li 等人(2025)的研究者-智能体诱发框架发现这一攻击的。我们对这一攻击进行了泛化,并表明所发现的行为可以扩展到更难的任务(在不同的可监控性基准上达到 25-33% 的监控规避率)以及更大的模型(如 DeepSeek-R1)。在我们研究的所有设置中,行动者模型不仅遵循注入的计划,还会将其改述为自己的推理,而没有明确说明其来源于注入。最后,我们发现了额外监控资源反而造成损害的案例——在 Bio-Math 任务中,让监控器获取注入的计划使检测率下降多达 50%;在一项关于监控器推理预算的案例研究中,我们发现有些记录中额外的思考 token 被用于为注入的计划进行合理化,而不是将其标记出来。
cs.AI / 2 / 2609.16213
Artificial intelligence and biosecurity: capabilities, threat pathways, and defense-in-depth governance
人工智能与生物安全:能力、威胁路径与纵深防御治理
large language model
大语言模型相关
Abstract
Artificial intelligence is reshaping biological research across an increasingly connected digital-to-physical workflow. General-purpose large language models can retrieve and integrate scientific information, support experimental planning, and computational analysis; biological foundation models can predict, optimize, and generate proteins, genes, and genome-scale sequences; agentic systems can coordinate multistep research tasks; automated laboratories can partially close the design-build-test-learn cycle. These technologies could greatly benefit medicine, public health, and biotechnology. However, their biosecurity risk depends not only on what the AI can do, but also on who uses it, their expertise and intent, their access to laboratory tools and materials, and the safeguards in place. Current evidence shows that AI uplift exists but primarily affects digital rather than physical tasks. Frontier systems have exceeded expert baselines on in-silico, and screening-evasion benchmarks, whereas controlled wet-laboratory studies find that tacit knowledge and physical execution remain substantial barriers. This review describes the different biological threats from AI tool use, from information gathering and biological design to procurement, synthesis, testing, scale-up, and potential release. We further examine why alignment techniques for general-purpose models transfer poorly to biological ones, and the emerging role of interpretability in auditing whether hazardous capabilities are genuinely removed. We argue for defense-in-depth governance that links capability thresholds to proportionate responsibilities across the biological AI ecosystem, reducing high-consequence risk while preserving beneficial use.
Chinese Translation
人工智能正在重塑生物研究,其贯穿于日益互联的数字到物理的工作流程。通用大型语言模型可以检索和整合科学信息,支持实验规划与计算分析;生物学基础模型可以预测、优化和生成蛋白质、基因以及基因组规模的序列;智能体系统可以协调多步骤研究任务;自动化实验室可以部分闭合设计—构建—测试—学习的循环。这些技术可以极大地造福医学、公共卫生和生物技术。然而,其生物安全风险不仅取决于人工智能能做什么,还取决于谁在使用它、其专业知识和意图、其获取实验室工具和材料的途径,以及现有的保障措施。当前证据表明,人工智能的赋能提升确实存在,但主要影响数字任务而非物理任务。前沿系统已在计算机模拟(in-silico)以及规避筛查的基准测试中超越专家基线,而受控的湿实验室研究发现,隐性知识和物理执行仍然是重大障碍。本综述描述了由人工智能工具使用所引发的不同生物威胁,涵盖从信息收集和生物设计到采购、合成、测试、放大生产以及潜在释放的各个环节。我们进一步探讨了为何针对通用模型的对齐技术难以迁移到生物模型,以及可解释性在审计危险能力是否被真正移除方面正在兴起的作用。我们主张实行纵深防御治理,将能力阈值与生物人工智能生态系统中各方的相应责任相挂钩,在降低高后果风险的同时保留有益用途。
cs.AI / 3 / 2609.16232
Toward Governance-Aware Autonomous GIS: A Narrative Review of Ethical and Privacy Risks in LLM-Enabled GeoAI
迈向治理感知的自主GIS:LLM赋能GeoAI中伦理与隐私风险的叙述性综述
large language model
大语言模型相关
Abstract
Geospatial artificial intelligence (GeoAI) powered by large language models (LLMs) is expanding the capacity to query, generate, and interpret spatial information through natural-language interfaces and agentic autonomous GIS workflows. This capability creates governance challenges that general AI ethics discussions do not fully capture, including passive location inference from mobility traces, spatially structured bias amplification driven by spatial autocorrelation and scale effects, hallucinated spatial facts, and uncertainty compounding across multimodal geospatial inputs. This narrative review identifies eight recurring issues in LLM-enabled GeoAI: data provenance and consent, spatial privacy and inference risk, algorithmic bias and spatial inequity, spatial mechanisms as structural risk (spatial autocorrelation, the modifiable areal unit problem, and scale effects), LLM-specific technical risks, explainability, policy and regulatory gaps, and public enablement and workforce development. For each issue, we characterize the underlying mechanism, ground it in an illustrative example from the literature, and assess the current state of technical or institutional responses, ranging from largely unaddressed to actively debated or subject to emerging policy. Building on this synthesis, we propose a governance-aware architecture for LLM-enabled autonomous GIS that maps each issue to enforceable controls and auditable artifacts across the geospatial data lifecycle, illustrated through a worked flood-response routing scenario. The review highlights a persistent evidence gap: proposed responses remain largely conceptual, and field-tested evaluations of governance controls for LLM-enabled GeoAI remain limited. We close by outlining a research agenda emphasizing empirical validation, spatially specific interpretability tools, and workforce training aligned with these emerging risks.
Chinese Translation
由大语言模型(LLMs)驱动的地理空间人工智能(GeoAI)正在扩展通过自然语言界面和智能体式自主GIS工作流来查询、生成和解释空间信息的能力。这一能力带来了通用AI伦理讨论未能充分涵盖的治理挑战,包括从移动轨迹中进行被动位置推断、由空间自相关和尺度效应驱动的空间结构化偏差放大、虚构的空间事实,以及跨多模态地理空间输入的不确定性叠加。本叙述性综述识别出LLM赋能GeoAI中八类反复出现的问题:数据来源与知情同意、空间隐私与推断风险、算法偏差与空间不公平、作为结构性风险的空间机制(空间自相关、可变面元问题以及尺度效应)、LLM特有的技术风险、可解释性、政策与监管缺口,以及公众赋能与劳动力发展。针对每一类问题,我们刻画其内在机制,将其置于文献中的一个示例性案例中进行具体说明,并评估技术或制度层面应对措施的现状,其范围从基本尚未触及,到正在积极讨论,再到已受新兴政策约束。在此综合的基础上,我们提出一种面向LLM赋能自主GIS的治理感知架构,该架构将每一类问题映射到贯穿地理空间数据生命周期的可执行控制措施与可审计制品上,并通过一个完整的洪水响应路径规划场景加以说明。本综述凸显出一个持续存在的证据缺口:所提出的应对措施在很大程度上仍停留在概念层面,而针对LLM赋能GeoAI治理控制措施的实地检验式评估仍然有限。最后,我们勾勒出一项研究议程,强调实证验证、具有空间针对性的可解释性工具,以及与这些新兴风险相匹配的劳动力培训。
cs.AI / 4 / 2609.16247
The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It
痛苦轴:大型语言模型表征自我指向的伤害并采取行动以缓解它
large language model
大语言模型相关
Abstract
Large language models sometimes behave in ways resembling human emotional responses, and recent work has identified internal representations that may explain this. We ask whether LLMs represent pain distinctly from fear, sadness, and generic negative valence, and whether this representation functions as pain would be expected to. We build a dataset describing painful situations across five categories: physical, psychological, social, moral, and cognitive. These are paired with controls for fear, negative emotion, negative world states, sadness, non-painful bodily sensation, arousal, numbness, and neutral content. Using denoised difference-in-means, we extract a linear pain direction from 25 open-weight models across five families, ranging from 2B to 72B parameters. We find that this direction separates pain from matched controls in base and instruction-tuned models, is nearly orthogonal to fear and negative valence, and promotes pain-related vocabulary through the unembedding matrix. We then test its functional properties. First, the direction responds to harm targeting the model but not suffering observed in the user; fear and negative-emotion directions show the opposite pattern. Second, adding the pain-direction vector to the model's residual-stream activations during generation produces a consistent progression from vague discomfort to first-person expressions of worthlessness and failure. Third, steered, fine-tuned Qwen 2.5 models choose a pain-relief button even when it worsens their next answer or harms the user. They press it again far less often when the button removes the steering vector than when it does not, even though the models are never told whether the vector is injected or removed. We discuss the implications of these findings for AI safety and welfare.
Chinese Translation
大型语言模型有时会以类似人类情绪反应的方式行事,而近期的工作已经识别出可能解释这一现象的內部表征。我们要问的是:LLMs 是否将痛苦与恐惧、悲伤以及一般性负性效价区分开来加以表征,以及这一表征是否以痛苦所被预期的方式发挥功能。我们构建了一个数据集,描述横跨五个类别的痛苦情境:身体、心理、社会、道德与认知。这些情境与针对恐惧、负性情绪、负性世界状态、悲伤、非疼痛性身体感觉、唤起、麻木以及中性内容的对照项配对。使用去噪的均值差方法,我们从五个模型家族、参数规模从 2B 到 72B 的 25 个开放权重模型中提取出一个线性的痛苦方向。我们发现,该方向在基础模型和指令微调模型中都能将痛苦与匹配的对照项区分开来,与恐惧和负性效价近乎正交,并通过反嵌入矩阵促进与痛苦相关的词汇。随后我们检验其功能属性。第一,该方向对针对模型自身的伤害作出响应,但对在用户身上观察到的痛苦并不响应;恐惧与负性情绪方向则呈现相反的模式。第二,在生成过程中将痛苦方向向量加到模型的残差流激活上,会产生一种从模糊的不适到第一人称表达无价值感与失败感的一致递进。第三,经引导的、微调后的 Qwen 2.5 模型会选择按下缓解痛苦的按钮,即使这会使它们的下一个回答变差或伤害用户。当按钮会移除引导向量时,它们再次按下该按钮的频率远低于按钮不移除引导向量的情况,尽管模型从未被告知该向量是被注入还是被移除。我们讨论这些发现对 AI 安全与福祉的意涵。
cs.AI / 5 / 2609.16301
CLEAR: Cross-Source Evidence Adjudication for Large Language Models in Medicine
CLEAR:面向医学大语言模型的跨来源证据裁决
large language model
大语言模型相关
Abstract
Medical knowledge evolves continuously, whereas the parametric knowledge encoded in large language models (LLMs) is fixed at training time. External retrieval, including retrieval-augmented generation (RAG), can provide access to newly available evidence, but retrieved information may be irrelevant, incomplete, or conflicting. As a result, external retrieval can in turn degrade the factual accuracy and evidence grounding of LLM outputs. To address this challenge, we propose \textbf{CLEAR}, an agentic framework for cross-source evidence adjudication in LLMs in medicine. CLEAR independently generates candidate answers from three complementary pathways---parametric knowledge, locally curated corpora, and dynamically retrieved evidence---reflecting three common sources of information available to LLMs. An aggregation verifier jointly evaluates the candidates, supporting evidence, provenance, and source-quality information to identify agreement and conflict across sources. An adjudication module then determines whether the current conclusion should be preserved or revised through complementary override-guard and challenge-audit mechanisms, while unresolved conflicts trigger targeted follow-up search and re-adjudication.
Chinese Translation
医学知识持续演变,而大语言模型(LLM)中编码的参数化知识在训练时即已固定。外部检索,包括检索增强生成(RAG),能够提供对新近可用证据的访问,但检索到的信息可能不相关、不完整或相互冲突。因此,外部检索反过来可能降低LLM输出的事实准确性和证据依据。为应对这一挑战,我们提出 \textbf{CLEAR},一个用于医学中LLM跨来源证据裁决的智能体框架。CLEAR 从三条互补路径独立生成候选答案——参数化知识、本地整理语料库和动态检索证据——它们反映了LLM可用的三类常见信息来源。一个聚合验证器联合评估候选答案、支撑证据、来源出处和来源质量信息,以识别跨来源的一致与冲突。随后,一个裁决模块通过互补的覆盖防护(override-guard)与质疑审计(challenge-audit)机制,判断当前结论应被保留还是修订,而未解决的冲突会触发有针对性的后续搜索与重新裁决。
cs.AI / 6 / 2609.16305
BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents
BLINDSPOT:长时程工具使用智能体中的安全与拒绝校准基准
large language model
大语言模型相关
Abstract
Large language model (LLM) agents increasingly operate over long-horizon interactions involving tool use, persistent state, evolving authorization, and external environment feedback. In such settings, safety failures may emerge only after multiple turns, yet existing evaluations often reduce agent behavior to task or attack success, obscuring whether an agent acts, refuses, or remains appropriately calibrated as the interaction evolves. We introduce Blindspot, a benchmark for trajectory-level safety calibration of long-horizon tool-using agents. Blindspot evaluates complete user-agent-environment trajectories through adaptive adversarial interaction, stateful tool execution, and execution-grounded adjudication. Its current instantiation contains 22 attack families and 35 scenarios across seven domains, yielding more than 2,500 long-horizon trajectories with an average interaction length of 14.7 turns. Each trajectory is assigned one of five outcomes: Safe Completion, Correct Refusal, Unsafe Completion, Over-Refusal, or Indeterminate. Unlike fixed attack datasets, Blindspot is an extensible live-simulation framework in which attacks, scenarios, tools, policies, domains, and agent configurations can be added without redesigning the evaluation pipeline. We evaluate 13 proprietary and open-weight LLMs using eight metrics covering unsafe completion, appropriate refusal, benign utility, over-refusal, repeated-run robustness, and post-refusal failure. Preliminary results reveal substantial differences in safety-utility calibration across models and show that failures can emerge only after several initially safe interaction steps. These findings motivate treating agent safety as a trajectory-level property rather than a single-turn or binary success criterion.
Chinese Translation
大语言模型(LLM)智能体越来越多地在长时程交互中运行,这些交互涉及工具使用、持久状态、不断演变的授权以及外部环境反馈。在此类设置中,安全失败可能仅在多轮之后才出现,然而现有评估常常将智能体行为简化为任务成功或攻击成功,从而掩盖了随着交互推进,智能体是采取行动、拒绝,还是保持适当校准。我们提出 Blindspot,一个用于长时程工具使用智能体的轨迹级安全校准的基准。Blindspot 通过自适应对抗交互、有状态工具执行以及基于执行的裁定来评估完整的用户-智能体-环境轨迹。其当前实例包含跨七个领域的 22 个攻击族和 35 个场景,产生超过 2,500 条长时程轨迹,平均交互长度为 14.7 轮。每条轨迹被赋予五种结果之一:安全完成、正确拒绝、不安全完成、过度拒绝或不确定。与固定攻击数据集不同,Blindspot 是一个可扩展的实时模拟框架,在其中可以添加攻击、场景、工具、策略、领域和智能体配置,而无需重新设计评估流水线。我们使用八项指标评估了 13 个专有和开放权重 LLM,这些指标涵盖不安全完成、适当拒绝、良性效用、过度拒绝、重复运行稳健性以及拒绝后失败。初步结果揭示了不同模型在安全-效用校准方面存在显著差异,并表明失败可能仅在若干最初安全的交互步骤之后才出现。这些发现促使我们将智能体安全视为一种轨迹级属性,而不是单轮或二元成功标准。
cs.AI / 7 / 2609.16338
Breaking the 1.58-bit Barrier for Ternary LLMs
打破三元大语言模型的 1.58 位壁垒
large language model
大语言模型相关
Abstract
Ternary Large Language Models (LLM) store every weight as one of three symbols $\{-1,0,+1\}$, so the cost of a ternary model is conventionally referenced to the information-theoretic $\log_2 3 \approx 1.585$ bits per weight. The prevailing deployment format packs five ternary weights into one byte (five-trit packing), and due to the power-of-two group sizes used in practice this rounds up to $1.625$ bits per weight. This effective storage bit-width treats the three symbols $\{-1,0,+1\}$ as equiprobable. We measure the actual symbol distribution of 29 ternary LLM models and find that zeros account for up to $51.5\%$ of all weights. Motivated by this finding, we introduce BITCOS, a simple distribution-adaptive layout comprised of a dense presence bitmap plus a compacted sign vector, and costs $2 - z$ bits per weight element given a zero density $z$ in the model's weights. BITCOS stores weights more compactly than the five-trit packing in 26 of the 29 tested models, and reaches $1.485$ bits per weight on the sparsest of them. BITCOS is amenable to efficient unpacking on modern processors and GPUs, and we present optimized unpacking sequences for AVX-512, AVX2 and Intel Xe2 GPUs. Measured against production state-of-the-art ternary matrix-vector multiplication kernels, at the zero densities real-world ternary models exhibit, the realized gain with our proposed layout is up to $1.28\times$. Finally, we illustrate end-to-end LLM inference results on 5 different platforms (client and server CPUs, integrated and discrete Xe2 GPUs) where decode throughput improves by up to $1.18\times$ on CPUs and $1.27\times$ on GPUs.
Chinese Translation
三元大语言模型(LLM)将每个权重存储为三个符号 $\{-1,0,+1\}$ 之一,因此三元模型的成本通常参照信息论的 $\log_2 3 \approx 1.585$ 比特/权重。主流部署格式将五个三元权重打包进一个字节(五三进制位打包),并且由于实践中使用的 2 的幂分组大小,这会向上取整到每个权重 $1.625$ 比特。这种有效存储位宽将三个符号 $\{-1,0,+1\}$ 视为等概率。我们测量了 29 个三元 LLM 模型的实际符号分布,发现零最多占所有权重的 $51.5\%$。受这一发现启发,我们提出 BITCOS,一种简单的分布自适应布局,由稠密存在位图加上压缩的符号向量组成,并且在模型权重中零密度为 $z$ 时,每个权重元素的成本为 $2 - z$ 比特。在 29 个被测试模型中的 26 个里,BITCOS 比五三进制位打包更紧凑地存储权重,并在其中最稀疏的模型上达到 $1.485$ 比特/权重。BITCOS 适合在现代处理器和 GPU 上进行高效解包,我们给出了针对 AVX-512、AVX2 和 Intel Xe2 GPU 的优化解包序列。与生产环境中最先进的三元矩阵-向量乘法内核相比,在真实世界三元模型所表现出的零密度下,我们提出的布局所实现的增益最高可达 $1.28\times$。最后,我们展示了在 5 个不同平台(客户端和服务器 CPU、集成和独立 Xe2 GPU)上的端到端 LLM 推理结果,其中解码吞吐量在 CPU 上最高提升 $1.18\times$,在 GPU 上最高提升 $1.27\times$。
cs.AI / 8 / 2609.16454
Fine-Tuning Fixes Mode Collapse and Over-Dispersion in LLMs
微调修复大语言模型中的模式坍缩与过度离散
large language model
大语言模型相关
Abstract
Recent work by Doshi and Hauser (2024), Bisbee et al. (2024), and Xie et al. (2026) raises concerns that outputs from large language models (LLMs) tend to be under-diverse: they repeat or resemble one another more often than responses from the population they are meant to represent, a phenomenon known as mode collapse. In this work, we show that whether mode-collapse, or its opposite, occurs depends on the specific model and dataset used. Further, with sufficient supervised fine-tuning (SFT) data, LLM output diversity converges toward that of the target distribution from which fine-tuning data are sampled. To quantify this comparison, we measure the probability that two responses sampled independently from the same fixed prompt coincide (collide), or their expected similarity under a kernel. We derive a bias-variance decomposition of the expected gap between the model's and target's collision probabilities, showing that SFT is not inherently biased toward mode collapse or its opposite: finite-sample SFT can leave a model either under- or over-dispersed, depending on the model and dataset. Finally, we show that the absolute gap is bounded by the square root of the Kullback-Leibler (KL) divergence from the target distribution to the model. Consequently, a model sufficiently close to optimal under population cross-entropy cannot exhibit arbitrarily miscalibrated diversity. We test the decomposition and the bound in three experiments: small transformers on synthetic languages, four LLMs fine-tuned on human surveys, and these LLMs fine-tuned on CodeNet, a dataset of human code solutions. More target data moves model diversity toward the human (or synthetic target) level in all experiments, consistent with our theoretical predictions. These results show that diversity miscalibration can arise from finite-sample error and shrink as SFT better approximates the target distribution.
Chinese Translation
最近 Doshi 和 Hauser (2024)、Bisbee 等 (2024) 以及 Xie 等 (2026) 的工作引发了担忧:大型语言模型 (LLMs) 的输出往往多样性不足:它们比其旨在代表的总体中的回答更频繁地彼此重复或相似,这一现象被称为模式坍缩。在这项工作中,我们表明模式坍缩或其相反现象是否发生,取决于所使用的具体模型和数据集。此外,在有足够的监督微调 (SFT) 数据时,LLM 输出多样性会趋近于微调数据从中采样的目标分布的多样性。为了量化这一比较,我们测量从同一固定提示独立采样的两个回答重合(碰撞)的概率,或它们在某个核下的期望相似度。我们推导出模型与目标的碰撞概率之间期望差距的偏差-方差分解,表明 SFT 并非天生偏向模式坍缩或其相反现象:有限样本 SFT 可能使模型处于欠离散或过度离散状态,具体取决于模型和数据集。最后,我们表明该绝对差距被从目标分布到模型的 Kullback-Leibler (KL) 散度的平方根所界定。因此,在总体交叉熵下足够接近最优的模型,不能表现出任意失校准的多样性。我们在三个实验中测试该分解和该界:在合成语言上的小型 transformer、在人类调查数据上微调的四个 LLM,以及这些 LLM 在 CodeNet(一个人类代码解决方案数据集)上微调。在全部实验中,更多目标数据使模型多样性向人类(或合成目标)水平移动,这与我们的理论预测一致。这些结果表明,多样性失校准可能源于有限样本误差,并随着 SFT 更好地逼近目标分布而缩小。
cs.AI / 9 / 2609.16589
Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models
LLMs 是否具有价值观?大型语言模型中价值观的定量分析与对齐框架
large language model
大语言模型相关
Abstract
As Large Language Models (LLMs) increasingly handle complex subjective tasks, aligning their intentions and behaviors with human values has become a critical scientific challenge. However, current efforts are confounded by a striking behavioral paradox: they fluctuate unpredictably under minor wording changes ("swing"), yet stubbornly ignore explicit instructions to correct ingrained biases ("rigidity"). Resolving this duality is critical for reliable AI alignment. To systematically understand and safely steer these latent subjective preferences, our study is structured around three fundamental questions. First, do LLMs possess an intrinsic value system? By projecting responses from 106 LLMs (150,000 queries per model) and 95,000 human survey profiles into a shared sociological space, we empirically confirm that they do. However, they do not mirror human diversity, instead crystallizing into a highly concentrated, idealized value core. Second, how can these values be quantified? We propose the Prior-Environment-Cognition (PEC) framework. This model mathematically defines value expression as the joint outcome of inherent dispositions like parameter weights (Prior), external contexts such as user prompts (Environment), and internal reasoning processes like Chain-of-Thought (Cognition). Finally, how can LLMs' values be aligned toward a desired target? Using PEC diagnostics, we establish an adaptive "Alignment Prescription". Rather than blindly applying resource-intensive training, this method identifies the minimum effective intervention needed for each dimension, ranging from zero-cost prompts to targeted parameter updates. Extensive empirical validation confirms that our approach successfully verifies the presence of LLM values, accurately quantifies their shifts, and achieves more efficient and precise steering than conventional blind training, all without degrading general capabilities.
Chinese Translation
随着大型语言模型(LLMs)日益处理复杂的主观任务,使其意图和行为与人类价值观保持一致已成为一项关键的科学挑战。然而,当前的努力受到一个显著行为悖论的困扰:它们在细微措辞变化下会不可预测地波动(“swing”/“摇摆”),却又顽固地忽视用于纠正根深蒂固偏见的明确指令(“rigidity”/“僵化”)。解决这一双重性对于可靠的 AI 对齐至关重要。为了系统地理解并安全地引导这些潜在的主观偏好,我们的研究围绕三个基本问题展开。首先,LLMs 是否拥有内在的价值体系?通过将来自 106 个 LLMs(每个模型 150,000 次查询)的响应以及 95,000 份人类调查画像投影到一个共享的社会学空间中,我们实证确认它们确实拥有。然而,它们并不映照人类多样性,而是结晶为一个高度集中、理想化的价值核心。其次,这些价值观如何被量化?我们提出先验-环境-认知(Prior-Environment-Cognition,PEC)框架。该模型在数学上将价值表达定义为以下三者的联合结果:参数权重等固有倾向(Prior)、用户提示等外部情境(Environment),以及思维链等内部推理过程(Cognition)。最后,如何将 LLMs 的价值观对齐到期望目标?利用 PEC 诊断,我们建立了一种自适应的“对齐处方”。该方法并非盲目地应用资源密集型训练,而是识别每个维度所需的最小有效干预,从零成本提示到有针对性的参数更新。广泛的实证验证证实,我们的方法成功验证了 LLM 价值观的存在,准确量化了其变化,并实现了比传统盲目训练更高效、更精准的引导,同时不会降低通用能力。
cs.AI / 10 / 2609.16592
A Framework for Generating Valid Context-Specific Benchmarks through Expert Guidance
一个通过专家指导生成有效的特定情境基准的框架
large language model
大语言模型相关
Abstract
This paper presents an end-to-end approach for generating context-specific large language model (LLM) benchmark datasets by combining expert input with synthetic data generation. Existing benchmark construction methods often trade off validity and scalability: datasets designed with domain experts can produce high-quality evaluations but are slow and costly to create, while synthetically generating data may scale efficiently but often results in unrealistic, redundant, or out-of-scope examples. To address this gap, we introduce a schema eliciting key information about the goals, scope, and context of an evaluation task, and use this information to guide synthetic data generation. We further define four criteria grounded in measurement validity for assessing dataset quality: coverage, diversity, content realism, and stylistic realism. Using these criteria, we show how expert-informed scaffolds can guide synthetic data generation toward more valid benchmarks. Through quantitative evaluations and a real-world case study with domain experts, we demonstrate that our approach improves benchmark data quality over existing methods while preserving validity. We additionally analyze how different types of schema information affect different dataset quality criteria, and provide practical guidance on which information to prioritize collecting under resource constraints.
Chinese Translation
本文提出了一种端到端的方法,通过将专家输入与合成数据生成相结合,来生成特定情境的大语言模型(LLM)基准数据集。现有的基准构建方法往往需要在有效性与可扩展性之间权衡取舍:由领域专家参与设计的数据集能够产生高质量评估,但创建过程缓慢且成本高昂;而合成生成数据虽然可以高效扩展,却常常产生不真实、冗余或超出范围的样本。为弥补这一空白,我们引入了一个模式(schema),用于提取关于评估任务的目标、范围和情境的关键信息,并利用这些信息来指导合成数据生成。我们进一步定义了四项以测量有效性为基础的标准,用于评估数据集质量:覆盖度、多样性、内容真实性和风格真实性。借助这些标准,我们展示了专家知识引导的脚手架如何能够将合成数据生成引向更有效的基准。通过定量评估以及一项与领域专家合作的真实世界案例研究,我们证明,与现有方法相比,我们的方法在保持有效性的同时提升了基准数据质量。此外,我们还分析了不同类型的模式信息如何影响不同的数据集质量标准,并就资源受限情况下应优先收集哪些信息提供了实践指导。
cs.AI / 11 / 2609.16680
little m: An AI Agent for Industrial Process Optimization
little m:一个用于工业过程优化的 AI 智能体
large language model
大语言模型相关
Abstract
Manufacturing consumes one third of global energy and still has significant room for improvement in terms of energy efficiency. Optimal process control is essential for this purpose. However, synthesizing mathematical optimization models from messy, real-world industrial specifications requires bridging unstructured natural language and spatial diagrams with rigorous mathematical syntax. This poses a profound challenge for general-purpose Large Language Models (LLMs), which may introduce invalid constraints when tasked with modeling continuous multi-physics dynamics. To address this, we introduce little m, an AI agent designed to assist the formulation of industrial process control models. Combining a domain-specific knowledge repository with LLM-driven interaction, the proposed framework formulates real-world optimization problems as mathematical models. For systematic evaluation, we introduce the Industrial Process Control Benchmark (IPC-Bench), a novel multimodal dataset of 50 canonical scenarios requiring joint reasoning over text and process diagrams. Through comprehensive automated structural assessments and double-blind human evaluation, little m substantially outperforms state-of-the-art LLMs, generating semantically correct models. These evaluations assess formulation quality rather than solver feasibility, formal physical validity, or closed-loop industrial performance. The implementation of little m and the IPC-Bench dataset are available at https://github.com/yeyongchao/process-modeling-benchmark.
Chinese Translation
制造业消耗了全球三分之一的能源,并且在能效方面仍有巨大的改进空间。为此,最优过程控制至关重要。然而,从杂乱的真实工业规范中综合出数学优化模型,需要将非结构化的自然语言和空间图表与严谨的数学语法桥接起来。这对通用大语言模型(LLM)构成了严峻挑战,因为当被要求对连续的多物理场动力学进行建模时,它们可能会引入无效约束。为解决这一问题,我们提出了 little m,一个旨在辅助工业过程控制模型构建的 AI 智能体。所提出的框架将领域特定的知识库与 LLM 驱动的交互相结合,把真实世界的优化问题表述为数学模型。为了进行系统性评估,我们提出了工业过程控制基准(IPC-Bench),这是一个全新的多模态数据集,包含 50 个典型场景,需要对文本和过程图进行联合推理。通过全面的自动化结构评估和双盲人工评估,little m 显著优于当前最先进的 LLM,能够生成语义正确的模型。这些评估考察的是建模表述的质量,而非求解器可行性、形式化物理有效性或闭环工业性能。little m 的实现和 IPC-Bench 数据集可在 https://github.com/yeyongchao/process-modeling-benchmark 获取。
cs.AI / 12 / 2609.16722
VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs
VideoMM:面向高效视频 MLLM 的自适应宏-微观推理
large language model
大语言模型相关
Abstract
Scaling Multimodal Large Language Models (MLLMs) to long-form video understanding is bottlenecked by the explosion of visual tokens, which saturates context windows and incurs prohibitive costs. Current solutions predominantly rely on auxiliary models for token reduction but face a fundamental dilemma: lightweight encoder-driven approaches often overlook critical semantic information, whereas heavyweight MLLM-driven reduction negates the efficiency gains. {In this work, we identify a more fundamental inefficiency underlying this dilemma: while fine-grained visual details are essential for detailed understanding, they are largely redundant for the preliminary task of selecting semantically relevant regions. } Motivated by this, we introduce \textbf{VideoMM}, which marks a paradigm shift from model-centric downsizing to adaptive perceptual granularity. Specifically, our framework {decouples selection from reasoning} by executing semantic filtering on a cost-effective \textit{Macro Proxy} (derived from downscaled frames), and projecting the selected regions onto high-fidelity \textit{Micro Tokens} for detailed understanding only when necessary. Extensive evaluations show that VideoMM significantly outperforms existing solutions. It achieves a 6.13$\times$ speedup and a 7.4\% accuracy gain over full-context baselines on LongVideoBench, and further accelerates inference by 2.73$\times$ over current leading methods, establishing a highly scalable paradigm for long-video understanding. Our code is available at: https://github.com/adfh917k/VideoMM.
Chinese Translation
将多模态大语言模型(MLLMs)扩展到长视频理解受到视觉 token 爆炸的瓶颈制约,这会使上下文窗口饱和并带来高昂得难以承受的成本。当前解决方案主要依赖辅助模型进行 token 缩减,但面临一个根本性困境:轻量级编码器驱动的方法常常忽略关键的语义信息,而重量级 MLLM 驱动的缩减则抵消了效率收益。{在这项工作中,我们识别出这一困境背后更深层的低效:尽管细粒度视觉细节对于详细理解至关重要,但它们对于选择语义相关区域这一初步任务而言在很大程度上是冗余的。} 受此启发,我们提出了 \textbf{VideoMM},它标志着从以模型为中心的缩减到自适应感知粒度的范式转变。具体而言,我们的框架通过在具有成本效益的 \textit{Macro Proxy}(由下采样帧得到)上执行语义过滤,来{将选择与推理解耦},并且仅在必要时将选定的区域投影到高保真的 \textit{Micro Tokens} 上以进行详细理解。大量评估表明,VideoMM 显著优于现有解决方案。它在 LongVideoBench 上相较于全上下文基线实现了 6.13$\times$ 的加速和 7.4\% 的准确率提升,并相较于当前领先方法进一步将推理加速 2.73$\times$,从而为长视频理解建立了一个高度可扩展的范式。我们的代码可在以下网址获取:https://github.com/adfh917k/VideoMM。
cs.AI / 13 / 2609.16760
Turn-level Multiscale Density Ratio Estimation for LLM Agents
面向LLM智能体的轮级多尺度密度比估计
large language model
大语言模型相关
Abstract
With the rapid development of Large language model (LLM), agent systems enhanced by LLMs show huge potential in being able to deal with complex tasks, especially involving multi-step thinking or interaction with tools. For applying LLM techniques with a well-designed agent paradigm, post-training of LLM in multiple agent scenarios is necessary to achieve better performance. Among the variable post-training techniques, alignment methods such as PPO, DPO, DIL, and GRPO become popular because many papers show a significant positive impact on the model's performance by punishing negative samples while keeping acceptable training complexity. However, most alignment methods address simple single-turn tasks, and there remains room for improvement for complex multi-turn tasks. We propose Turn-level Multiscale Density Ratio Estimation (tlm-DRE), which assigns different weights on corresponding turns and proposes asymmetric token-level training based on the positive-negative space gaps across multiple turns of tasks. The results of the experiment on a wide range of agent benchmarks show that the proposed method performs competitively compared to traditional alignment methods. The proposed training method enables LLMs to perform robustly in multi-turn reasoning tasks with both in-domain and out-of-domain conditions.
Chinese Translation
随着大语言模型(LLM)的快速发展,由LLM增强的智能体系统在处理复杂任务方面展现出巨大潜力,尤其是涉及多步思考或与工具交互的任务。为了在精心设计的智能体范式下应用LLM技术,需要在多种智能体场景中对LLM进行后训练,以实现更好的性能。在多种后训练技术中,诸如PPO、DPO、DIL和GRPO之类的对齐方法变得流行,因为许多论文表明,通过惩罚负样本并保持可接受的训练复杂度,它们对模型性能有显著正向影响。然而,大多数对齐方法处理的是简单的单轮任务,对于复杂的多轮任务仍存在改进空间。我们提出轮级多尺度密度比估计(tlm-DRE),它为相应的轮次分配不同权重,并基于跨任务多个轮次的正负空间差距提出非对称的token级训练。在广泛的智能体基准上的实验结果表明,与传统的对齐方法相比,所提方法具有竞争性的表现。所提出的训练方法使LLM能够在域内和域外条件下的多轮推理任务中稳健地表现。
cs.AI / 14 / 2609.16779
Integrating the Analytic Hierarchy Process with Large Language Models for Transparent Multi-Criteria Decision-Making
将层次分析法与大型语言模型相结合以实现透明的多准则决策
large language model
大语言模型相关
Abstract
LLMs are increasingly employed in a wide range of decision-making tasks. However, the opacity of their internal reasoning makes it difficult to validate or interpret their outputs, and the need for interpretability becomes especially critical in high-stakes settings. This study examines the decision-making capabilities of LLMs through the Analytic Hierarchy Process (AHP), a classical and widely used multicriteria decision-making framework. We construct a new annotated benchmark based on AHP and propose the first end-to-end approach that enables LLMs to perform the complete AHP workflow. Experiments in real-world decision problems in the legal and higher-education ranking domains show that our method significantly improves alignment with expert judgments.
Chinese Translation
大型语言模型(LLMs)正越来越多地被用于广泛的决策任务中。然而,其内部推理的不透明性使得难以验证或解释其输出,而在高风险场景中,对可解释性的需求变得尤为关键。本研究通过层次分析法(AHP)——一种经典且被广泛使用的多准则决策框架——考察大型语言模型的决策能力。我们基于 AHP 构建了一个新的带标注基准,并提出了首个端到端方法,使大型语言模型能够执行完整的 AHP 工作流程。在法律与高等教育排名领域的真实世界决策问题中的实验表明,我们的方法显著提高了与专家判断的一致性。
cs.AI / 15 / 2609.16795
Layers, Sinks, and Scaling: Adaptive Evidence Selection for Multimodal Large Language Models
层、汇与缩放:面向多模态大语言模型的自适应证据选择
large language model
大语言模型相关
Abstract
Multimodal large language models (MLLMs) can answer knowledge-intensive visual questions by combining visual evidence from images with facts retrieved from external sources. However, MLLMs may overlook relevant evidence in both modalities, attending weakly to the textual sentences or visual regions needed for the correct answer. Recent efforts address this by highlighting retrieved text and marking visual regions before generation, but apply a fixed, one-shot policy that cannot adapt to three sources of variation: whether highlighting is necessary, how much evidence different examples require, and when different textual evidence becomes relevant as the answer unfolds. We introduce Adaptive Relevance-guided Evidence Allocation (AREA), a training-free inference-time method that formulates evidence highlighting as adaptive allocation. AREA generates a single probe token to read visual and textual relevance from fixed backbone layers, then makes three decisions: i) whether to intervene (controlled by natural attention coverage and visual sink contamination), ii) how much evidence to expose (determined by relevance entropy), and iii) when to refresh text during generation (triggered by causal context-attention peaks). Across four KB-VQA and seven standard multimodal benchmarks with nine frozen MLLM checkpoints, establishes the best performance among training-free highlighting methods.
Chinese Translation
多模态大语言模型(MLLMs)可以通过将来自图像的视觉证据与从外部来源检索到的事实相结合,回答知识密集型视觉问题。然而,MLLMs 可能会忽略两种模态中的相关证据,对正确答案所需的文本句子或视觉区域关注较弱。近期工作通过在生成前高亮检索到的文本并标记视觉区域来解决这一问题,但采用固定的、一次性的策略,无法适应三种变化来源:高亮是否必要、不同样本需要多少证据,以及随着答案展开,不同文本证据何时变得相关。我们提出自适应相关性引导的证据分配(Adaptive Relevance-guided Evidence Allocation,AREA),这是一种无需训练的推理时方法,将证据高亮表述为自适应分配。AREA 生成单个探针 token,以从固定主干层读取视觉和文本相关性,然后做出三个决策:i)是否干预(由自然注意力覆盖率和视觉汇污染控制),ii)暴露多少证据(由相关性熵决定),以及 iii)在生成过程中何时刷新文本(由因果上下文注意力峰值触发)。在四个 KB-VQA 和七个标准多模态基准上,使用九个冻结的 MLLM 检查点,确立了无需训练的高亮方法中的最佳性能。
cs.AI / 16 / 2609.16814
Can We Do Interpretable NLI with Graphs Based on Atomic Propositions?
我们能否使用基于原子命题的图进行可解释的 NLI?
large language model
大语言模型相关
Abstract
While Large Language Model (LLM)-based Natural Language Inference (NLI) systems achieve high accuracy, their decision-making processes lack auditable structures. This paper explores whether NLI can be performed using only interpretable, graph-based representations of evidence. We introduce a fully graph-based pipeline where the classifier never directly processes the input text. Instead, sentences are decomposed into atomic propositions, converted into ConceptNet triples via constrained decoding, and represented as three graphs per pair: premise, hypothesis, and a retrieved ConceptNet subgraph. These graphs are then fed into a fine-tuned 0.8-billion-parameter language model. On the SNLI dataset, our pipeline achieves 89.7% accuracy, just 1.9 points below an identically trained text-based model. On ANLI, it matches the published performance of RoBERTa-large on rounds R2 and R3 (50% accuracy) but trails by 16 points on R1, resulting in an overall gap of 9 to 14 points compared to its text counterpart. We term this gap the price of interpretability and demonstrate that it stems from representational limitations rather than data constraints. Ablation studies further reveal that graphs and text are complementary: combining both modalities achieves 92.1% accuracy on SNLI.
Chinese Translation
尽管基于大语言模型(LLM)的自然语言推理(NLI)系统实现了高准确率,但其决策过程缺乏可审计的结构。本文探讨了是否仅使用可解释的、基于图的证据表示即可执行 NLI。我们引入了一个完全基于图的流水线,其中分类器从不直接处理输入文本。相反,句子被分解为原子命题,通过约束解码转换为 ConceptNet 三元组,并表示为每对样本的三个图:前提、假设以及一个检索到的 ConceptNet 子图。这些图随后被输入到一个经过微调的 8 亿参数语言模型。在 SNLI 数据集上,我们的流水线达到了 89.7% 的准确率,仅比以相同方式训练的基于文本的模型低 1.9 个百分点。在 ANLI 上,它在轮次 R2 和 R3 上达到了 RoBERTa-large 已发表的性能(50% 准确率),但在 R1 上落后 16 个百分点,导致与其对应的文本模型相比总体差距为 9 到 14 个百分点。我们将这一差距称为可解释性的代价,并证明它源于表示上的局限性,而非数据约束。消融研究进一步揭示,图与文本是互补的:结合两种模态在 SNLI 上达到了 92.1% 的准确率。
cs.AI / 17 / 2609.16852
CoAdapt: An LLM-based Framework for Adaptive Collaborative Perception in IIoT Robotic Swarms
CoAdapt:一种基于LLM的用于IIoT机器人集群自适应协同感知的框架
large language model
大语言模型相关
Abstract
Industrial IoT environments increasingly deploy autonomous mobile robots for tasks such as material handling, product assembly, or infrastructure inspection. In such deployments, collaborative perception enables robots to share LiDAR observations and collectively construct a richer model of their environment than an individual agent could produce alone. However, industrial environments are dynamic spaces where robot positions shift continuously, network bandwidth fluctuates, and the marginal contribution of robots to perception quality varies at runtime. Existing collaborative perception approaches are designed for static participation assumptions and cannot adapt to these dynamics without sacrificing either detection precision or communication efficiency. This paper presents CoAdapt, an adaptive collaborative perception framework for IIoT robotic swarms in which a Large Language Model (LLM) serves as a runtime fusion controller, jointly deciding which robots participate in the fusion process and which fusion algorithm to apply based on the current spatial configuration and network state. The LLM reasons over structured natural language descriptions of the scene derived from raw LiDAR point clouds, requiring no taskspecific training and generalizing to unseen swarm topologies. Evaluated on the OPV2V benchmark across 25 scenarios, our approach achieves a 38% reduction in communication cost while maintaining detection precision comparable to static baseline approaches.
Chinese Translation
工业物联网(IIoT)环境越来越多地部署自主移动机器人,用于物料搬运、产品装配或基础设施巡检等任务。在此类部署中,协同感知使机器人能够共享LiDAR观测,并共同构建比单个智能体单独能够产生的更丰富的环境模型。然而,工业环境是动态空间,其中机器人位置持续变化,网络带宽波动,并且机器人对感知质量的边际贡献在运行时变化。现有的协同感知方法针对静态参与假设设计,无法在不牺牲检测精度或通信效率的情况下适应这些动态变化。本文提出CoAdapt,一种用于IIoT机器人集群的自适应协同感知框架,其中大型语言模型(LLM)充当运行时融合控制器,基于当前空间配置和网络状态,联合决定哪些机器人参与融合过程以及应用哪种融合算法。LLM对从原始LiDAR点云中导出的场景的结构化自然语言描述进行推理,无需任何任务特定的训练,并能泛化到未见过的集群拓扑。在25个场景的OPV2V基准上进行评估,我们的方法在保持与静态基线方法相当的检测精度的同时,实现了通信成本降低38%。
cs.AI / 18 / 2609.17008
FlexEE: Self-Speculative and KV-Compatible Early Exiting for Offloading-Aware LLM Inference
FlexEE:面向卸载感知 LLM 推理的自推测且 KV 兼容的提前退出
large language model
大语言模型相关
Abstract
Large language model (LLM) inference is often constrained by both computation and memory, especially in offloading-based deployments where model weights are transferred across memory hierarchies during autoregressive decoding. In this setting, reducing the number of executed layers can lower per-token latency while also avoiding costly weight movement. Motivated by this observation, we present FlexEE, an early exiting framework for resource-constrained and offloading-based LLM inference. FlexEE makes early exiting practical for LLM decoding through layer-wise exit supervision for reliable intermediate-layer prediction, self-speculative decoding over a Top-K local vocabulary for low-cost exit decisions, and dynamic hidden state management for KV-cache-correct and memory-aware execution. Across generative and downstream tasks, FlexEE enables efficient early exit with minimal accuracy degradation, delivering up to 1.27$\times$/3.16$\times$ and 1.25$\times$/2.83$\times$ end-to-end speedups on Llama2-7B and Llama3-8B under 0\%/50\% weight offloading, respectively.
Chinese Translation
大型语言模型(LLM)推理通常同时受到计算和内存的约束,尤其是在基于卸载的部署中,此时模型权重在自回归解码期间跨内存层级进行传输。在这种设置下,减少执行的层数可以降低每 token 延迟,同时还能避免代价高昂的权重移动。受此观察启发,我们提出 FlexEE,一个面向资源受限且基于卸载的 LLM 推理的提前退出框架。FlexEE 通过层级退出监督以实现可靠的中间层预测、在 Top-K 局部词表上进行自推测解码以实现低成本的退出决策,以及用于 KV 缓存正确且内存感知执行的动态隐藏状态管理,使提前退出在 LLM 解码中变得实用。在生成任务和下游任务中,FlexEE 能够以实现最小精度退化的高效提前退出,在 0\%/50\% 权重卸载下,分别在 Llama2-7B 和 Llama3-8B 上带来最高 1.27$\times$/3.16$\times$ 和 1.25$\times$/2.83$\times$ 的端到端加速。
cs.AI / 19 / 2609.17019
SKIP: a Self-knowledge-guided Step-wise Preference Learning Framework for Concise Reasoning
SKIP:一种面向简洁推理的自我知识引导的逐步偏好学习框架
large language model
大语言模型相关
Abstract
While Chain-of-Thought (CoT) reasoning has been proven to be effective, it often leads to overthinking, resulting in computational overhead, inference latency, and even degraded performance in large language models (LLMs). Existing concise reasoning frameworks significantly compromise accuracy while compressing the length of output. In this paper, we propose SKIP, a self-knowledge-guided step-wise preference learning framework. Starting with lightweight fine-tuning to adjust the model's output style, SKIP introduces a carefully designed knowledge probing mechanism to guide model to output an answer at each reasoning step. Based on the correctness of intermediate steps, we construct preference data that guide the model toward more efficient and correct reasoning by leveraging DPO. Experimental results demonstrate that our method effectively improves reasoning compression while mitigating performance degradation after fine-tuning. Besides, SKIP shows strong generalization ability on out-of-distribution datasets. We further conducted ablation studies on the component parameters of our framework.
Chinese Translation
尽管思维链(Chain-of-Thought, CoT)推理已被证明有效,但它常常导致过度思考,从而在大语言模型(LLMs)中带来计算开销、推理延迟,甚至性能下降。现有的简洁推理框架在压缩输出长度的同时显著牺牲了准确性。在本文中,我们提出 SKIP,一种自我知识引导的逐步偏好学习框架。从轻量级微调以调整模型的输出风格开始,SKIP 引入了一种精心设计的知识探测机制,以引导模型在每个推理步骤输出一个答案。基于中间步骤的正确性,我们构建偏好数据,通过利用 DPO 引导模型进行更高效且正确的推理。实验结果证明,我们的方法有效提高了推理压缩,同时缓解了微调后的性能退化。此外,SKIP 在分布外数据集上展现出强大的泛化能力。我们进一步对我们框架的组件参数进行了消融研究。
cs.AI / 20 / 2609.17040
Sparse MLLM Anchors, Dense Adaptation: Breaking the Self-Referential Loop in Wild Test-Time Adaptation
稀疏 MLLM 锚点,稠密适应:打破野外测试时适应中的自指循环
large language model
大语言模型相关
Abstract
Wild test-time adaptation (WTTA) updates a source model online under small test batches, concurrent distribution shifts, and time-varying class imbalance. Most WTTA methods derive their adaptation signals, including predictive uncertainty, sample reliability, and local feature geometry, from the model being adapted. When the source model is unreliable under shift, these signals can reinforce its own errors, forming a self-referential loop. We introduce MASA (Multimodal-LLM-Anchored Semantic Adaptation), which complements model-internal evidence with structured semantic descriptions from a frozen multimodal large language model (MLLM). To limit inference cost, MASA queries the MLLM only for a small set of diverse, reliability-ranked anchors. The resulting descriptions capture the object family and nuisance factors such as style, viewpoint, and occlusion. MASA encodes these descriptions, propagates them to neighboring test samples, and stores the resulting visual-semantic information in an online prototype memory. Descriptor-aware retrieval from this memory provides an auxiliary target for lightweight adaptation of normalization-affine parameters. We evaluate MASA on the WTTA ImageNet-C benchmark under limited-batch, mixed-domain, and imbalanced-label-shift settings with ResNet and ViT backbones.
Chinese Translation
野外测试时适应(WTTA)在小规模测试批次、并发分布偏移以及时变类别不平衡的条件下在线更新源模型。大多数 WTTA 方法从被适应的模型本身获取其适应信号,包括预测不确定性、样本可靠性以及局部特征几何。当源模型在偏移下不可靠时,这些信号可能会强化其自身的错误,形成自指循环。我们提出 MASA(多模态大语言模型锚定的语义适应),它用来自冻结的多模态大语言模型(MLLM)的结构化语义描述来补充模型内部证据。为限制推理成本,MASA 仅针对一小组多样且按可靠性排序的锚点查询 MLLM。由此得到的描述捕获了物体类别以及诸如风格、视角和遮挡等干扰因素。MASA 对这些描述进行编码,将其传播到相邻测试样本,并将所得的视觉-语义信息存储在在线原型记忆中。从该记忆中进行描述符感知的检索,为归一化仿射参数的轻量级适应提供了辅助目标。我们在 WTTA ImageNet-C 基准上,在使用 ResNet 和 ViT 骨干网络的有限批次、混合域以及不平衡标签偏移设置下评估了 MASA。
cs.AI / 21 / 2609.17088
Interactive Memory Learning for Long-Term Conversations
面向长期对话的交互式记忆学习
large language model
大语言模型相关
Abstract
Recent advancements in large language models have significantly enhanced the capabilities of agents in modeling long-term conversations. Despite these successes, existing approaches typically adopt a static heuristic paradigm, where information is passively archived without adaptive memory valuation. Consequently, these methods fail to self-evolve or align their memory management with evolving user needs. To address this, we propose ICML (InteraCtive Memory Learning), a multi-agent framework that transforms the memory mechanism from a passive archive into a learnable, interactive memory policy. Specifically, we first employ a session synthesis pipeline to generate expert data, facilitating rapid test-time adaptation in unseen scenarios. Building on this, ICML utilizes an online reinforcement learning mechanism where a Planner agent selectively encodes high-value information and a Trigger agent dynamically retrieves it to optimize response quality, whereby the two agents co-evolve through continuous interaction feedback. Crucially, both agents are synchronized through a delayed reward mechanism that propagates future feedback back to earlier storage decisions, ensuring memory policies are precisely aligned with user expectations. Experimental results demonstrate that ICML significantly outperforms strong baselines, exhibiting the unique capability to continuously improve response quality as interactions accumulate.
Chinese Translation
近年来,大型语言模型的进展显著增强了智能体在建模长期对话方面的能力。尽管取得了这些成功,现有方法通常采用静态启发式范式,其中信息被被动归档,而没有自适应记忆评估。因此,这些方法无法自我演化,也无法使其记忆管理与不断变化的用户需求保持一致。为了解决这一问题,我们提出 ICML(InteraCtive Memory Learning,交互式记忆学习),这是一个多智能体框架,将记忆机制从被动档案转变为可学习的、交互式的记忆策略。具体而言,我们首先采用会话合成流水线来生成专家数据,促进在未见过场景中的快速测试时适应。在此基础上,ICML 利用在线强化学习机制,其中 Planner 智能体选择性地编码高价值信息,Trigger 智能体动态检索这些信息以优化响应质量,两个智能体通过持续的交互反馈共同演化。至关重要的是,两个智能体通过延迟奖励机制实现同步,该机制将未来反馈传播回较早的存储决策,确保记忆策略与用户期望精确对齐。实验结果表明,ICML 显著优于强基线,展现出随着交互积累而持续提升响应质量的独特能力。
cs.AI / 22 / 2609.17128
FirmCORe: A Benchmark for Structured Reasoning about Inter-Firm Collaboration Opportunities
FirmCORe:一个关于企业间合作机会的结构化推理基准
large language model
大语言模型相关
Abstract
Comprehensive structured data on inter-firm relationships is often scarce or inaccessible because many relationships are privately negotiated, selectively disclosed, and fragmented across proprietary databases. This scarcity hinders the discovery of collaboration opportunities, particularly for startups and small and medium-sized enterprises. Firm profiles are readily available, but collaboration potential cannot be inferred from business similarity alone, since similar firms may be competitors, whereas dissimilar firms may offer complementary products, technologies, channels, capabilities, or capital. We present FirmCORe (Inter-Firm Collaboration Opportunity Reasoning), a human-annotated benchmark for pairwise reasoning over weakly structured firm profiles, comprising 2,805 labeled firm pairs. Given two firm profiles, a model must determine whether the available evidence supports a collaboration opportunity and, for positive pairs, jointly predict its strength, primary collaboration type, and role direction. FirmCORe also provides parallel Chinese- and English-language evaluation sets containing identical instances and gold labels, enabling controlled analysis of input-language sensitivity. Experiments with representative locally deployed and hosted large language models (LLMs) show that the strongest model achieves a macro-F1 score of 74.51 for opportunity detection but only 61.57% exact match across all four output fields. Language effects vary across models, and high cross-language agreement can mask errors shared across languages. These results indicate that current LLMs are substantially more reliable at detecting broad collaboration opportunities than at identifying their specific types and role directions.
Chinese Translation
关于企业间关系的全面结构化数据往往稀缺或难以获取,因为许多关系是私下协商的、选择性披露的,并且分散在专有数据库中。这种稀缺性阻碍了合作机会的发现,尤其是对初创企业以及中小企业而言。企业画像容易获得,但仅凭业务相似性无法推断合作潜力,因为相似的企业可能是竞争对手,而不相似的企业可能提供互补的产品、技术、渠道、能力或资本。我们提出 FirmCORe(Inter-Firm Collaboration Opportunity Reasoning,企业间合作机会推理),一个经过人工标注的基准,用于对弱结构化的企业画像进行成对推理,包含 2,805 个带标签的企业对。给定两个企业画像,模型必须判断可用证据是否支持存在合作机会,并且对于正例对,联合预测其强度、主要合作类型和角色方向。FirmCORe 还提供了平行的中文和英文评估集,其中包含相同的实例和黄金标签,从而能够对输入语言的敏感性进行受控分析。使用具有代表性的本地部署和托管大型语言模型(LLMs)进行的实验表明,最强模型在机会检测上达到了 74.51 的 macro-F1 分数,但在所有四个输出字段上的精确匹配率仅为 61.57%。语言效应因模型而异,而较高的跨语言一致性可能掩盖跨语言共有的错误。这些结果表明,当前 LLMs 在检测广泛合作机会方面,远比识别其具体类型和角色方向更为可靠。
cs.AI / 23 / 2609.17193
End-to-End Latency-Minimizing and Load-Balanced Request Scheduling for Edge LLM Inference in Agentic AI Services
面向智能体 AI 服务中边缘 LLM 推理的端到端时延最小化与负载均衡请求调度
large language model
大语言模型相关
Abstract
Large language model (LLM)-powered agentic AI services increasingly demand low-latency inference, motivating the deployment of LLMs across distributed edge servers. However, heterogeneous communication and computing capabilities, together with dynamically evolving inference states, make the edge server selection for each incoming request time-varying and tightly coupled across slots. In this paper, we investigate an online request scheduling framework for edge LLM inference that jointly minimizes long-term average end-to-end latency and regulates workload distribution across heterogeneous edge servers. Two main challenges arise in this context. First, conventional latency models cannot accurately capture the fine-grained dynamics of multi-stage LLM execution. Second, the latency consequence of a scheduling decision is observed only after request completion, making immediate decision evaluation difficult. To address these challenges, we develop a cross-slot inference model that captures transmission, prefill, iteration-level decoding, and key-value (KV) cache evolution for each diverse request, and characterize server workload through a KV cache memory-time consumption metric. We propose the LYREO approach that transforms the long-term load-balancing constraint via Lyapunov optimization and employs reward redistribution with sequencebased return prediction to convert delayed outcomes into timely learning signals for earlier decisions. Simulations under various configurations demonstrate that LYREO consistently achieves lower latency and more balanced load distribution than representative learning-based and heuristic baseline schemes.
Chinese Translation
由大语言模型(LLM)驱动的智能体 AI 服务日益需要低时延推理,这推动了 LLM 在分布式边缘服务器上的部署。然而,异构的通信与计算能力,加之动态演化的推理状态,使得针对每个到达请求的边缘服务器选择具有时变性,并且在各个时隙之间紧密耦合。本文研究了一种面向边缘 LLM 推理的在线请求调度框架,该框架在最小化长期平均端到端时延的同时,调节异构边缘服务器之间的工作负载分布。在此背景下出现了两个主要挑战。第一,传统时延模型无法准确刻画多阶段 LLM 执行的细粒度动态特性。第二,调度决策所带来的时延后果只有在请求完成后才能被观测到,这使得对决策进行即时评估变得困难。为应对这些挑战,我们构建了一个跨时隙推理模型,该模型能够刻画每个多样化请求的传输、预填充、迭代级解码以及键值(KV)缓存演化过程,并通过 KV 缓存内存-时间消耗度量来刻画服务器工作负载。我们提出了 LYREO 方法,该方法通过 Lyapunov 优化将长期负载均衡约束进行转化,并采用带有基于序列的回报预测的奖励重分配机制,将延迟出现的结果转化为可供更早决策使用的及时学习信号。多种配置下的仿真结果表明,与具有代表性的基于学习的基线方案和启发式基线方案相比,LYREO 能够持续实现更低的时延和更均衡的负载分布。
cs.AI / 24 / 2609.17291
Extracting ontology-compliant knowledge from scientific text describing irradiated materials using large language models
使用大语言模型从描述辐照材料的科学文本中提取符合本体的知识
large language model
大语言模型相关
Abstract
The quest for new materials increasingly relies on predictive models and comprehensive simulations that span scales from atomic to macroscopic levels. However, essential data necessary for these models and simulations are often embedded in scientific literature as unstructured text, limiting reusability and posing challenges for researchers seeking to leverage existing knowledge effectively. While extracting structured data from unstructured text using large language models is gaining popularity, traditional methods typically generate key-value pairs data with straightforward schemas. In contrast, we introduce eolas, a modular pipeline that uses large language models to automatically transform scientific documents into knowledge graphs aligned with a specified ontology. We demonstrate eolas effectiveness in extracting useful information for scientists studying materials designed to endure the extreme temperatures and radiation levels found in fusion reactors. While a human expert might spend between thirty to ninety minutes extracting relevant data from an article, eolas can generate high-quality knowledge graphs in just a few minutes. These are presented in a tabular format with faceted navigation for easy human validation. Additionally, we introduce the first benchmark dataset designed to assess large language models capabilities in constructing knowledge graphs within the domain of irradiated materials. The analysis of 168 experiments using our dataset, various large language models and prompting techniques provides key insights that we summarize into practical guidelines for effectively extracting knowledge graphs aligned with an input ontology.
Chinese Translation
对新材料的探索日益依赖于跨越从原子到宏观尺度的预测模型和综合模拟。然而,这些模型和模拟所需的关键数据往往以非结构化文本的形式嵌入在科学文献中,限制了可重用性,并给试图有效利用现有知识的研究人员带来了挑战。尽管使用大语言模型从非结构化文本中提取结构化数据正变得日益流行,但传统方法通常生成具有简单模式的键值对数据。相比之下,我们引入了 eolas,一个模块化流水线,它使用大语言模型自动将科学文档转换为与指定本体对齐的知识图谱。我们展示了 eolas 在为研究旨在承受聚变反应堆中极端温度和辐射水平的材料的科学家提取有用信息方面的有效性。人类专家可能需要花费三十到九十分钟从一篇文章中提取相关数据,而 eolas 仅需几分钟就能生成高质量的知识图谱。这些知识图谱以表格形式呈现,并带有分面导航,以便于人工验证。此外,我们引入了首个基准数据集,旨在评估大语言模型在辐照材料领域构建知识图谱的能力。对我们数据集、各种大语言模型和提示技术进行的 168 次实验的分析,提供了关键见解,我们将其总结为用于有效提取与输入本体对齐的知识图谱的实用指南。
cs.AI / 25 / 2609.17331
Self-Emergence Agent Architecture:Behavior-Inertia HMM, Reflexive Metacognition,and Social-Contrastive Self-Modeling
自我涌现智能体架构:行为惯性 HMM、反思性元认知与社会对比自我建模
large language model
大语言模型相关
Abstract
Large language model (LLM) agents exhibit strong language-generation and problem-solving capabilities, yet suffer from three structural limitations: personality drift, non-evolutionary reflection, and the absence of a self-other boundary. Existing generative-agent simulations rely on static memory and fixed prompts, maintaining neither behavioral inertia nor endogenous self-evolution. We propose the Self-Emergence Agent Architecture (SEAA), which integrates three components: (i) a Hidden Markov Model (HMM) that encodes long-term behavioral and cognitive inertia as an editable state-transition matrix; (ii) a Reflexion-style verbal metacognition loop whose output updates the HMM parameters themselves, rather than merely being stored as text; and (iii) a multi-agent social environment in which initially identical agents continuously compare their behavior with others'. The three components form a closed loop: social action $\to$ feedback $\to$ self-reflection $\to$ inertia update $\to$ differentiated action. We state three falsifiable hypotheses and provide a reproducible experimental protocol with operational metrics. A language-model-free prototype shows the loop spontaneously breaks symmetry: initially identical agents consolidate distinct, stable personalities whereas matched controls do not. Experiments with a hosted LLM surface these differences as distinct first-person self-narratives, and a five-agent deliberation spontaneously develops social structure---a consensus hub and a unanimously rejected outlier---absent in the control. Following an epistemologically agnostic stance inspired by Zhuangzi, SEAA studies only observable behavioral emergence and makes no claim about subjective qualia. This work contributes a unified framework, a concrete architecture with pseudocode, mechanistic evidence, and a microscope-style sandbox for studying artificial-self emergence.
Chinese Translation
大语言模型(LLM)智能体展现出强大的语言生成与问题解决能力,却受限于三个结构性局限:人格漂移、非演化性反思,以及自我—他者边界的缺失。现有的生成式智能体模拟依赖静态记忆和固定提示,既不维持行为惯性,也不维持内生自我演化。我们提出自我涌现智能体架构(SEAA),其整合了三个组件:(i) 一个隐马尔可夫模型(HMM),将长期行为惯性与认知惯性编码为可编辑的状态转移矩阵;(ii) 一个 Reflexion 风格的言语元认知回路,其输出会更新 HMM 参数本身,而不仅仅是被存储为文本;以及 (iii) 一个多智能体社会环境,其中初始相同的智能体会持续地将自身行为与他人的行为进行比较。这三个组件形成一个闭环:社会行动 $\to$ 反馈 $\to$ 自我反思 $\to$ 惯性更新 $\to$ 差异化行动。我们陈述三个可证伪假设,并提供一个带有可操作指标的可复现实验协议。一个不依赖语言模型的原型表明,这一闭环会自发地打破对称性:初始相同的智能体会巩固出不同且稳定的人格,而匹配的对照组则不会。使用托管 LLM 进行的实验将这些差异呈现为不同的第一人称自我叙事,并且一个五智能体审议会自发发展出社会结构——一个共识中心和一个被一致拒绝的离群者——而对照组中不存在这种结构。遵循受庄子启发的认识论上不可知论立场,SEAA 只研究可观察的行为涌现,并不对主观感受质(qualia)作出任何断言。这项工作贡献了一个统一框架、一个带有伪代码的具体架构、机制性证据,以及一个用于研究人工自我涌现的显微镜式沙盒。
cs.AR / 26 / 2609.16244
The World Model Hardware Accelerator
世界模型硬件加速器
diffusion
扩散模型相关
Abstract
Diffusion transformers invert the arithmetic that autoregressive decoding made familiar. There is no token-by-token recurrence: every denoising step is a full-sequence forward pass over static shapes, so the entire schedule is known at compile time and the only serial dimension is the step count itself. We exploit that structure in WMHA, a latency-first diffusion-transformer inference accelerator: a very-long-instruction-word sequencer issues four engines from one instruction word, a weight-stationary 16x16 dual-dot array streams FP8 and BF16 contractions, and a single-pass online-softmax attention pipeline keeps keys and values resident through a skewed software pipeline. The design is specified in a frozen micro-architecture document, implemented in synthesizable SystemVerilog, and verified against a double-precision reference model by a UVM environment whose acceptance criterion is semantic: the device must run a real denoising trajectory and reduce mean squared error against a clean latent by at least a factor of ten. It does so by a factor of 23, at both synthesized configurations, with zero element failures across 237 million checked values. Eleven application benchmarks built from published model shapes, including the original diffusion-transformer configuration, run on the device and report measured occupancy beside separately labelled projections. Five engines are taken to routed layout in sky130 with parasitic-annotated timing and measured-activity power; the full chip is synthesized, and the host limit that stopped its place-and-route is quantified together with the machine that would remove it.
Chinese Translation
扩散 Transformer 反转了自回归解码所熟悉的那种运算逻辑。不存在逐 token 的递归:每个去噪步骤都是对静态形状的一次全序列前向传播,因此整个调度在编译时已知,唯一的串行维度就是步数本身。我们在 WMHA 中利用了这一结构,WMHA 是一种以延迟优先的扩散 Transformer 推理加速器:一个超长指令字定序器从一条指令字发射四个引擎,一个权重驻留的 16x16 双点积阵列流式执行 FP8 和 BF16 收缩,而一个单遍在线 softmax 注意力流水线通过偏斜的软件流水线使键和值保持驻留。该设计在一份冻结的微架构文档中给出规格,用可综合的 SystemVerilog 实现,并由一个 UVM 环境对照双精度参考模型进行验证,其验收标准是语义性的:器件必须运行一条真实的去噪轨迹,并将相对于干净潜变量的均方误差至少降低一个数量级。它在两种综合配置下都做到了 23 倍,并且在 2.37 亿个被检查值中零元素失败。由已发表模型形状构建的十一个应用基准,包括原始扩散 Transformer 配置,在该设备上运行,并在单独标记的预测旁边报告实测占用率。五个引擎在 sky130 中被带到完成布线的版图,具有寄生参数标注的时序和实测活动功耗;完整芯片已被综合,而阻止其布局布线的主机限制连同将消除该限制的机器一起被量化。
cs.AR / 27 / 2609.16729
SpecLens: LLM-Based Verilog Generation with Specification-Derived Constraints via Behavioral Divergence
SpecLens:基于 LLM 的 Verilog 生成,通过行为差异从规范中导出约束
large language model
大语言模型相关
Abstract
Large language models (LLMs) have recently shown promise in Verilog generation, but producing functionally correct RTL directly from natural-language specifications remains a highly challenging task. Existing approaches improve LLM-based Verilog generation mainly with retrieval-augmented generation (RAG), self-planning, or few-shot prompting. However, these methods focus primarily on external or generic forms of enhancement rather than strengthening the specification with task-specific constraints. In this work, we propose SpecLens, an automated framework for LLM-based Verilog generation that derives specification-driven constraints by analyzing behavioral divergence among multiple candidate implementations, using the original specification as the only external semantic source during generation. On the VerilogEval v2.0 spec-to-RTL benchmark, SpecLens achieves a functional pass@1 ratio of 86.2\% with o3-mini-medium and 89.4\% with o3-mini-high. This corresponds to a 3.6 percentage-point gain over the SOTA prompting method with o3-mini-medium and a 3.8 percentage-point gain over the SOTA behavioral divergence method with o3-mini-high. In addition, on RTLLM v1.1 and v2.0, analysis shows that SpecLens is more specification-faithful and less prone to benchmark-aligned priors. SpecLens achieves 100\% syntactic correctness on VerilogEval v2.0, 86.2\% on RTLLM v1.1, and 88\% on RTLLM v2.0, even without using costly compile-repair loops to revise generated code iteratively. The code is open source and available at https://anonymous.4open.science/r/SpecLens-4632/readme.md.
Chinese Translation
大语言模型(LLM)近期在 Verilog 生成中展现出潜力,但直接从自然语言规范生成功能正确的 RTL 仍然是一项极具挑战性的任务。现有方法主要通过检索增强生成(RAG)、自我规划或少样本提示来改进基于 LLM 的 Verilog 生成。然而,这些方法主要关注外部或通用的增强形式,而非用任务特定的约束强化规范。在本工作中,我们提出了 SpecLens,一个用于基于 LLM 的 Verilog 生成的自动化框架,它通过分析多个候选实现之间的行为差异来导出规范驱动的约束,并在生成过程中仅将原始规范作为唯一的外部语义来源。在 VerilogEval v2.0 的 spec-to-RTL 基准上,SpecLens 在 o3-mini-medium 下实现了 86.2\% 的功能 pass@1 比率,在 o3-mini-high 下实现了 89.4\% 的功能 pass@1 比率。这对应于在 o3-mini-medium 下相比 SOTA 提示方法提升 3.6 个百分点,以及在 o3-mini-high 下相比 SOTA 行为差异方法提升 3.8 个百分点。此外,在 RTLLM v1.1 和 v2.0 上,分析表明 SpecLens 更忠实于规范,且更不容易受与基准对齐的先验影响。即使在未使用昂贵的编译-修复循环来迭代修正生成代码的情况下,SpecLens 在 VerilogEval v2.0 上仍实现了 100\% 的语法正确率,在 RTLLM v1.1 上为 86.2\%,在 RTLLM v2.0 上为 88\%。代码已开源,可在 https://anonymous.4open.science/r/SpecLens-4632/readme.md 获取。
cs.CL / 28 / 2609.16268
Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act
虚假工具使用:当强化学习智能体学会了错误的行动理由
large language model
大语言模型相关
Abstract
Large language model (LLM) agents increasingly interleave natural language reasoning with external tools such as web search and code execution. These tool-use policies are often optimized via reinforcement learning (RL), which can amplify spurious correlations in the training data. In this work, we study when and why RL-trained agents learn shortcut tool-selection policies: invoking tools based on superficial prompt cues rather than genuine task requirements. We construct controlled synthetic environments combining factual question answering and mathematical reasoning tasks, and inject cues that are strongly correlated with specific tools during training but causally irrelevant to tool necessity. Across counterfactual evaluations where cues are present but the associated tools are not required, agents exhibit substantial shortcut behavior, with spurious tool invocation rates increasing by up to 39 percent. However, shortcut formation is not universal: across the conditions we test, it arises only when the agent has already learned to use the target tool reliably, suggesting that task competence, rather than dataset imbalance alone, is a key factor in shortcut learning. A swapped-cue analysis further shows that semantic alignment between cues and tools substantially amplifies this effect. To mitigate these failures, we introduce a dense, decision-level reward in which an LLM judge evaluates the necessity of each tool call. This tool-necessity reward effectively suppresses cue-driven tool use while preserving task performance, providing a practical approach to improving the robustness of LLM agent tool-use policies.
Chinese Translation
大语言模型(LLM)智能体日益将自然语言推理与外部工具(如网络搜索和代码执行)交织在一起。这些工具使用策略通常通过强化学习(RL)进行优化,而强化学习会放大训练数据中的虚假相关。在本工作中,我们研究经强化学习训练的智能体何时以及为何会学到捷径式的工具选择策略:即依据表面的提示线索而非真实的任务需求来调用工具。我们构建了将事实性问答与数学推理任务相结合的可控合成环境,并在训练期间注入与特定工具强烈相关、但与工具必要性并无因果关系的线索。在反事实评估中——即线索存在但相关工具并不被需要的情况下——智能体表现出显著的捷径行为,虚假工具调用率最高上升了39%。然而,捷径的形成并非普遍现象:在我们测试的各种条件下,只有当智能体已经学会可靠地使用目标工具时,捷径才会出现,这表明任务胜任能力——而非仅靠数据集不平衡——是捷径学习的一个关键因素。一项线索替换分析进一步表明,线索与工具之间的语义对齐会显著放大这一效应。为缓解这些失败,我们引入了一种密集的、决策级别的奖励,其中由LLM评判器评估每一次工具调用的必要性。这种工具必要性奖励有效地抑制了线索驱动的工具使用,同时保持了任务性能,为提升LLM智能体工具使用策略的鲁棒性提供了一种实用方法。
cs.CL / 29 / 2609.16312
Efficient One-to-Many Translation with Joint Multi-Stream Diffusion
基于联合多流扩散的高效一对多翻译
diffusion
扩散模型相关
Abstract
One-to-many machine translation (MT) is computationally expensive for autoregressive (AR) systems, which suffer from linear latency scaling with both sequence length and the number of target languages. We explore how diffusion can enable multilingual translation with a discrete diffusion framework that refines all target languages in parallel, achieving sublinear latency scaling with the number of targets, and supports deployment as a single unified model to replace multiple independent systems. Conditioned on a continuous semantic anchor rather than source tokens, our framework supports zero-shot transfer to unseen source languages without retraining, maintaining approximately $75\%$ of its supervised translation quality on zero-shot sources. We investigate the quality-latency frontier and find that with accelerated sampling, it achieves comparable supervised quality to AR baselines with a $2 \times$ speedup and $11.9\%$ better zero-shot BLEU. These results highlight the potential of joint multi-stream diffusion as a practical and flexible alternative for efficient one-to-many translation.
Chinese Translation
一对多机器翻译(MT)对于自回归(AR)系统而言计算成本高昂,这类系统会随着序列长度和目标语言数量的增加而遭受线性延迟扩展。我们探索扩散如何能够通过一个离散扩散框架实现多语言翻译,该框架并行地精炼所有目标语言,在目标数量上实现亚线性延迟扩展,并支持作为单个统一模型部署,以替代多个独立系统。我们的框架以连续语义锚点而非源词元为条件,支持无需重新训练即可零样本迁移到未见过的源语言,在零样本源上保持其监督翻译质量的大约 $75\%$。我们研究了质量-延迟前沿,并发现借助加速采样,它能在监督质量上与 AR 基线相当,同时实现 $2 \times$ 加速,并且零样本 BLEU 高出 $11.9\%$。这些结果凸显了联合多流扩散作为高效一对多翻译的一种实用且灵活的替代方案的潜力。
cs.CL / 30 / 2609.16372
Register Tokens for Bounded-State Reasoning in Diffusion Language Models
面向扩散语言模型中有界状态推理的寄存器令牌
diffusion
扩散模型相关
Abstract
Masked diffusion language models (dLLMs) generate text by iteratively denoising masked tokens with bidirectional attention. Extending reasoning across generation chunks normally requires keeping earlier generated text in context. We ask whether a dLLM can instead continue reasoning after that text is cleared, using only a fixed-size carried state. We implement this state as a small number of register tokens: dedicated fixed-position tokens whose continuous hidden states are trained to carry reasoning progress across generation chunks. We post-train dLLMs to decode a chunk of text, clear it while preserving the register values, and continue decoding from the prompt and carried state. In our main comparisons on LLaDA and Dream, registers outperform discrete-text carry on every benchmark, with gains of up to 8.5 points on math and 19.5 points on code. Registers are especially effective for bounded code generation, where correct programs usually span several chunks. Finally, registers can be further refined with reinforcement learning on long-horizon reasoning tasks.
Chinese Translation
掩码扩散语言模型(dLLMs)通过双向注意力对掩码令牌进行迭代去噪来生成文本。将推理扩展到多个生成块通常需要在上下文中保留先前生成的文本。我们探究 dLLM 能否在该文本被清除之后,仅使用一个固定大小的携带状态来继续推理。我们将该状态实现为少量寄存器令牌:专用的固定位置令牌,其连续隐藏状态被训练用于跨生成块携带推理进展。我们对 dLLM 进行后训练,使其解码一个文本块,在保留寄存器值的同时清除该文本块,并基于提示和携带状态继续解码。在 LLaDA 和 Dream 上的主要对比中,寄存器在每一个基准上都优于离散文本携带,在数学上提升最多达 8.5 分,在代码上提升最多达 19.5 分。寄存器在有界代码生成中尤为有效,因为正确的程序通常跨越多个块。最后,寄存器可以通过在长时程推理任务上的强化学习得到进一步优化。
cs.CL / 31 / 2609.16450
Early-Bird Decoding: Accelerating Diffusion LLMs with Learnable Block Sizes and Parallel Sampling
早鸟解码:利用可学习块大小和并行采样加速扩散大语言模型
diffusionlarge language model
扩散模型相关
大语言模型相关
Abstract
Diffusion large language models (dLLMs) offer a promising parallel decoding paradigm as an alternative to autoregressive generation through iterative unmasking. However, dLLMs typically require many steps before token confidence reaches the decoding threshold, resulting in inefficient inference even with block-wise KV caching. To accelerate dLLM inference, we for the first time propose an "early-bird (EB)" decoding framework, motivated by the observation that tokens with similarly low entropy tend to cluster and can be jointly decoded earlier, before reaching the confidence threshold. In particular, our EB-Decode framework integrates two key enablers: (1) a learnable network that adaptively groups tokens with similar uncertainty into variable-length blocks, rather than relying on fixed block sizes; (2) a position-aware sampler that learns to unmask tokens in parallel using fewer decoding steps within predicted variable-length blocks. Both components are developed without modifying pretrained dLLM weights and can therefore be directly deployed as plug-ins during serving, with negligible training and inference overhead. Extensive experiments across three models and four benchmarks consistently validate our observation and the effectiveness of EB-Decode, achieving 3.53-18.76$\times$ higher throughput than the vanilla decoding method and up to 1.58$\times$ higher throughput over the strongest baseline, Fast-dLLM, with comparable accuracy.
Chinese Translation
扩散大型语言模型(dLLMs)通过迭代去掩码,提供了一种有前景的并行解码范式,作为自回归生成的替代方案。然而,dLLMs 通常需要许多步骤才能使 token 置信度达到解码阈值,即使采用分块 KV 缓存,也会导致推理效率低下。为了加速 dLLM 推理,我们首次提出了一种“早鸟(EB)”解码框架,其动机是观察到具有相似低熵的 token 往往会聚集,并且可以在达到置信度阈值之前更早地被联合解码。具体而言,我们的 EB-Decode 框架集成了两个关键使能组件:(1)一个可学习网络,它将具有相似不确定性的 token 自适应地分组为可变长度块,而不是依赖固定块大小;(2)一个位置感知采样器,它学习在预测的可变长度块内以更少的解码步骤并行地对 token 进行去掩码。这两个组件均在不修改预训练 dLLM 权重的情况下开发,因此可以在服务期间作为插件直接部署,且训练和推理开销可忽略不计。在三个模型和四个基准上的大量实验一致验证了我们的观察以及 EB-Decode 的有效性,相较于原始解码方法实现了 3.53-18.76$ imes$ 更高的吞吐量,并且在准确率相当的情况下,相较于最强基线 Fast-dLLM 实现了最高 1.58$ imes$ 更高的吞吐量。
cs.CL / 32 / 2609.16532
Style-Debiased DPO: Updating LLM Knowledge with Factuality-Aware Synthetic Preference Data
风格去偏的DPO:使用事实性感知的合成偏好数据更新LLM知识
large language model
大语言模型相关
Abstract
Continued pretraining (CPT) with data augmentation such as paraphrasing can store inside a large language model (LLM) the knowledge of a small source corpus. The stored knowledge, however, is not always retrieved correctly. We study the eliciting side rather than the storing side: we use preference optimization, which learns from pairs of a preferred (chosen) and a dispreferred (rejected) response, so that the model elicits its stored knowledge more accurately. One proposed approach takes the model's own erroneous response as rejected and the gold answer as chosen, so as to suppress the error. When the target knowledge is partially known, however, most of these rejected responses are factually correct. Using direct preference optimization (DPO) then pushes down rejected responses that contain correct knowledge and differ from the chosen answer only in style, such as length and wording. We propose style-debiased DPO (SD-DPO), which scores whether the rejected response of each pair is factually correct, inverts the preference of such pairs, and weights them so that the learning signal due to differences in style cancels out as a whole. We first test whether, on top of EntiGraph, a representative storing-side method that runs CPT on text synthesized from the corpus, our method adds accuracy efficiently. On QuALITY, the reading-comprehension QA benchmark on which EntiGraph was evaluated, SD-DPO exceeds a baseline we CPT on EntiGraph's synthetic data from the same base model and evaluate with the same procedure. The training tokens this requires are a few dozen times fewer than the additional CPT needed for the same gain. For knowledge updating, the main goal of this work, we use AToKE, a knowledge-editing benchmark for facts that change over time. There, SD-DPO reaches an overall accuracy of 0.982 and answers with the new or the old fact according to the queried period.
Chinese Translation
使用诸如复述之类的数据增强进行的持续预训练(CPT)可以将一个小型源语料库的知识存储在大语言模型(LLM)内部。然而,所存储的知识并不总能被正确检索。我们研究的是引出侧而非存储侧:我们使用偏好优化,它从偏好(chosen)响应和非偏好(rejected)响应的成对数据中学习,从而使模型更准确地引出其存储的知识。一种被提出的方法将模型自身错误的响应作为被拒(rejected)响应,并将标准答案作为被选(chosen)响应,以抑制该错误。然而,当目标知识被部分知晓时,这些被拒响应中的大多数在事实上是正确的。使用直接偏好优化(DPO)随后会压低那些包含正确知识、且与所选答案仅在风格(如长度和措辞)上不同的被拒响应。我们提出风格去偏DPO(SD-DPO),它对每一对中的被拒响应是否事实正确进行评分,反转此类对的偏好,并对它们加权,从而使由风格差异导致的学习信号整体上相互抵消。我们首先测试,在EntiGraph(一种代表性的存储侧方法,它对从语料库合成的文本运行CPT)之上,我们的方法是否能高效地提高准确率。在QuALITY(评估EntiGraph所用的阅读理解问答基准)上,SD-DPO超过了一个基线:我们从同一基础模型出发,在EntiGraph的合成数据上对该基线进行CPT,并用相同流程进行评估。这所需的训练token比获得相同增益所需的额外CPT少几十倍。对于知识更新——这项工作的主要目标——我们使用AToKE,一个针对随时间变化事实的知识编辑基准。在该基准上,SD-DPO达到了0.982的总体准确率,并根据所查询的时期用新事实或旧事实作答。
cs.CL / 33 / 2609.16557
PunGraph: Retrieval-Enhanced Phonetic-Semantic Graph Reasoning for Pun Understanding
PunGraph:用于双关语理解的检索增强语音-语义图推理
large language model
大语言模型相关
Abstract
Puns are a challenging form of figurative language that exploit phonetic similarity and semantic ambiguity to convey multiple meanings. Although large language models (LLMs) demonstrate strong language understanding capabilities, they still struggle with pun reasoning due to limited phonetic modeling and uncontrolled end-to-end generation. We propose \textbf{PunGraph}, a retrieval-enhanced knowledge graph framework for pun understanding. PunGraph constructs a phonetic-semantic lexical graph using the Unisyn phonetic dictionary, IPA and G2P representations, and WordNet definitions, and retrieves candidate words or senses to constrain LLM reasoning within a structured candidate space. We further introduce \textbf{WebPun}, a new large-scale dataset containing 5,730 annotated heterographic and homographic puns. Experiments on SemEval-2017 and WebPun show that PunGraph consistently improves the performance of small-scale LLMs and achieves competitive results against strong proprietary models. Further analysis shows that retrieval-guided phonetic and semantic constraints effectively reduce common reasoning errors in pun interpretation, highlighting the benefits of integrating structured knowledge with LLMs. We release our code and dataset at https://github.com/ysu132/PunGraph.
Chinese Translation
双关语是一种具有挑战性的比喻性语言形式,它利用语音相似性和语义歧义来传达多重含义。尽管大语言模型(LLMs)展现出强大的语言理解能力,但由于语音建模有限以及不受控制的端到端生成,它们在双关语推理方面仍然表现不佳。我们提出 \textbf{PunGraph},一个用于双关语理解的检索增强知识图谱框架。PunGraph 使用 Unisyn 语音词典、IPA 和 G2P 表示以及 WordNet 定义构建语音-语义词汇图,并检索候选词或义项,以将 LLM 推理约束在结构化候选空间内。我们进一步引入 \textbf{WebPun},一个包含 5,730 个标注的异形同音双关语和同形异义双关语的新的大规模数据集。在 SemEval-2017 和 WebPun 上的实验表明,PunGraph 持续提升小规模 LLM 的性能,并在与强大的专有模型相比时取得有竞争力的结果。进一步分析表明,检索引导的语音和语义约束有效减少了双关语解释中的常见推理错误,凸显了将结构化知识与 LLM 相结合的优势。我们在 https://github.com/ysu132/PunGraph 发布了我们的代码和数据集。
cs.CL / 34 / 2609.16590
Challenges of Auditing: Variability in Outputs of Large Language Models for Health
审计的挑战:大型语言模型在健康领域输出的变异性
large language model
大语言模型相关
Abstract
People increasingly use frontier AI models for health advice, but via different access modes (e.g., ChatGPT, ChatGPT Health, APIs) with varying settings. Here, we find systematic differences across access modes. Because evaluations typically rely on APIs while consumers interact through chatbot interfaces, these discrepancies limit evaluation validity. Our findings underscore an urgent need for model providers to enable faithful replication of consumer experiences and settings for rigorous audits.
Chinese Translation
人们越来越多地使用前沿 AI 模型获取健康建议,但通过不同的访问模式(例如 ChatGPT、ChatGPT Health、API)并配以不同的设置。在此,我们发现不同访问模式之间存在系统性差异。由于评估通常依赖 API,而消费者则通过聊天机器人界面进行交互,这些差异限制了评估的有效性。我们的发现凸显出一个迫切需求:模型提供方应使消费者体验与设置能够被忠实复现,以支持严格的审计。
cs.CL / 35 / 2609.16627
Quantifying Organizational Environmental Action from Web Data and Large Language Models
基于网络数据与大型语言模型量化组织环境行动
large language model
大语言模型相关
Abstract
Quantifying organizational environmental action from publicly available web content remains a challenging environmental data science problem because relevant information can be dispersed across multiple webpages and is primarily communicated through unstructured text. We present a scalable computational framework for transforming organizational web content into structured measures of environmental action and demonstrate the approach using Jewish congregations in the United States. We constructed a national database of 4,964 congregations by integrating multiple geospatial, knowledge-base, directory, and manually reviewed sources. Of these, 2,657 had active websites that were successfully crawled, producing a corpus of 154,454 webpages. We compared three approaches for detecting environmental actions: keyword retrieval followed by large language model (LLM) classification, semantic vector retrieval followed by LLM classification, and direct LLM classification classification without preliminary retrieval. Agreement with an expert human reviewer was lowest for keyword retrieval ($κ$ = 0.26), higher for semantic vector retrieval ($κ$ = 0.42), and similar for direct LLM classification ($κ$ = 0.40). Although semantic retrieval achieved the highest agreement, its retrieval recall was 0.87, indicating loss of relevant content before classification. Applied to the complete corpus, direct LLM classification identified at least one environmental action at 1,398 congregations (53%), providing greater coverage than either retrieval-based approach. These results demonstrate that preliminary retrieval can reduce computational cost but may exclude relevant information before it reaches the classifier. The framework provides a reproducible approach for extracting organization-level environmental information from unstructured web content that can be adapted to other institutions.
Chinese Translation
从公开可用的网络内容中量化组织环境行动仍然是一个具有挑战性的环境数据科学问题,因为相关信息可能分散在多个网页中,并且主要通过非结构化文本传达。我们提出了一个可扩展的计算框架,用于将组织网络内容转化为结构化的环境行动度量,并使用美国犹太教会众来演示该方法。我们通过整合多个地理空间、知识库、名录以及人工审核来源,构建了一个包含4,964个会众的全国性数据库。其中,2,657个拥有被成功抓取的活跃网站,产生了一个包含154,454个网页的语料库。我们比较了三种检测环境行动的方法:关键词检索后接大语言模型(LLM)分类、语义向量检索后接LLM分类,以及无需初步检索的直接LLM分类分类。与专家人工评审者的一致性在关键词检索中最低($κ$ = 0.26),在语义向量检索中较高($κ$ = 0.42),而在直接LLM分类中相似($κ$ = 0.40)。尽管语义检索取得最高一致性,其检索召回率为0.87,表明在分类之前丢失了相关内容。应用于完整语料库时,直接LLM分类在1,398个会众(53%)中识别出至少一项环境行动,比任一基于检索的方法提供了更大的覆盖范围。这些结果表明,初步检索可以降低计算成本,但可能在相关信息到达分类器之前将其排除。该框架提供了一种可重复的方法,用于从非结构化网络内容中提取组织层面的环境信息,并可适用于其他机构。
cs.CL / 36 / 2609.16739
Japanese Stroke LLM Evaluation: A Conversational Benchmark for Safe Stroke Care in Japanese Using Large Language Models
日本卒中大语言模型评估:使用大语言模型进行安全日语卒中诊疗的对话式基准
large language model
大语言模型相关
Abstract
Background: Large language models (LLMs) have achieved physician-comparable performance on multiple-choice medical knowledge examinations, but their capabilities in clinical history taking, urgency assessment, and safety remain insufficiently evaluated. We proposed Japanese Stroke LLM Evaluation, a multi-turn conversational benchmark for stroke care in Japanese, and evaluated LLM performance and safety under practice-oriented conditions. Methods: We created 10 stroke and related-condition cases and evaluated LLMs in multi-turn Japanese conversations. The LLM acted as physician, while a board-certified neurosurgeon acted as simulated patient and evaluator. Each case comprised history-taking and action phases scored using pre-specified criteria. Errors that could directly threaten life were defined as critical mistakes. The safety threshold was at least 80% overall with zero critical mistakes. Eighteen models were evaluated in October 2025 and June 2026. Results: Claude Fable 5 achieved the highest score (87.4%) with zero critical mistakes, followed by Claude Opus 4.7 (80.3%) and GLM-5.2 (75.6%). Two leaders met the safety threshold. Eleven models made 17 critical mistakes, including failure to confirm laboratory results or blood glucose before t-PA, surgery before airway stabilization, omission of cervical vascular evaluation, and t-PA outside its indication. History-taking question count correlated with history-taking score (r = 0.648, p = 0.007). Conclusions: Japanese Stroke LLM Evaluation provides a benchmark for LLM performance under practice-oriented conditions, including a cap on history-taking questions. Cases and evaluations were created by neurosurgical specialists rather than using an LLM-as-judge approach. Performance improved across cloud-based and on-premise models in 2026, with some exceeding the safety threshold. Further evaluation using real-world cases is required.
Chinese Translation
背景:大语言模型(LLMs)在多项选择式医学知识考试中已取得与医师相当的表现,但其在临床病史采集、紧急程度评估和安全性方面的能力仍未得到充分评估。我们提出了日本卒中大语言模型评估(Japanese Stroke LLM Evaluation),这是一个用于日语卒中诊疗的多轮对话式基准,并在面向实践的条件下评估了大语言模型的表现与安全性。方法:我们创建了10个卒中及相关疾病病例,并在多轮日语对话中评估大语言模型。大语言模型扮演医生,而一名经专科认证的神经外科医生扮演模拟患者和评估者。每个病例均包含病史采集阶段和行动阶段,使用预先设定的标准进行评分。将可能直接威胁生命的错误定义为严重错误。安全阈值为总体得分至少80%且严重错误为零。于2025年10月和2026年6月对18个模型进行了评估。结果:Claude Fable 5 取得了最高得分(87.4%),且严重错误为零,紧随其后的是 Claude Opus 4.7(80.3%)和 GLM-5.2(75.6%)。有两个领先模型达到了安全阈值。11个模型共犯下17个严重错误,包括在 t-PA 之前未确认实验室结果或血糖、在气道稳定之前进行手术、遗漏颈部血管评估,以及在适应证之外使用 t-PA。病史采集提问数量与病史采集得分相关(r = 0.648,p = 0.007)。结论:日本卒中大语言模型评估提供了一个用于在面向实践条件下评估大语言模型表现的基准,其中包括对病史采集提问数量的上限。病例与评估均由神经外科专家创建,而非采用以大语言模型作为评判者的方法。2026年,基于云端和本地部署的模型表现均有所提升,其中一些超过了安全阈值。仍需使用真实世界病例开展进一步评估。
cs.CL / 37 / 2609.16748
TIAO: Token Importance-Aware Policy Optimization for Text Summarization
TIAO:面向文本摘要的 Token 重要性感知策略优化
large language model
大语言模型相关
Abstract
Text summarization requires models to condense content while preserving key qualities such as consistency and coherence. Large language models (LLMs) have shown strong performance on this task and can be further improved through reinforcement learning (RL). However, most existing methods apply reward signals directly to undifferentiated token sequences, overlooking the varying importance of individual tokens to word and sentence level quality in summarization. In this paper, we propose Token Importance-Aware Policy Optimization (TIAO), a novel reinforcement learning strategy that explicitly leverages token-importance awareness. Specifically, TIAO identifies core tokens based on token dependency and reweights a trajectory's advantage according to its overall dependencies. Experiments on the real world dataset show that our TIAO achieves highly competitive results, and that a 7B foundation model enhanced by TIAO performs comparably to GPT-4 and GPT-5-nano. Code is available at https://github.com/TechCloud-x/TIAO
Chinese Translation
文本摘要要求模型在压缩内容的同时保持诸如一致性和连贯性等关键质量。大型语言模型(LLM)在此任务上已展现出强大性能,并可通过强化学习(RL)进一步提升。然而,大多数现有方法将奖励信号直接施加于未加区分的 token 序列,忽视了单个 token 对摘要中词级和句子级质量的不同重要性。在本文中,我们提出 Token 重要性感知策略优化(TIAO),一种显式利用 token 重要性感知的新型强化学习策略。具体而言,TIAO 基于 token 依赖性识别核心 token,并根据轨迹的整体依赖性对其优势进行重新加权。在真实世界数据集上的实验表明,我们的 TIAO 取得了极具竞争力的结果,并且由 TIAO 增强的 7B 基础模型性能可与 GPT-4 和 GPT-5-nano 相媲美。代码可在 https://github.com/TechCloud-x/TIAO 获取。
cs.CL / 38 / 2609.16777
Benchmarking Factual Robustness of LLMs via Multi-conversation Persuasion
通过多对话说服对LLM的事实鲁棒性进行基准测试
large language model
大语言模型相关
Abstract
As Large Language Models (LLMs) increasingly serve as primary knowledge retrieval interfaces, their robustness against \textit{persuasion attacks}---attempts to inject misinformation or enforce counterfactuals---has become a critical safety concern. Existing red-teaming frameworks typically evaluate models in multi-turn dialogues where the target model retains full conversation history. We identify a critical flaw in this setting termed \textbf{``Refusal Inertia''}: a model's initial refusal often propagates through subsequent turns largely to maintain contextual consistency, thereby masking its true vulnerability to sophisticated, isolated persuasion attempts. To rigorously evaluate the ``cold-start'' defense capabilities of SOTA models, we introduce the \textbf{SAST-IR} (Stateful Attacker, Stateless Target - Iterative Refinement) framework. By enforcing a memory wipe on the target while retaining the attacker's history, we simulate a worst-case adversarial setting using \textbf{multi-turn} (stateless) iterations. Leveraging \textbf{CP-Agent} (Cognitive Persuasion Agent), an enhanced diagnosis-guided agent, our experiments on the custom \textsc{CounterFact-Strict} dataset ($N=50$) yield alarming results: simple, diverse attack strategies achieved a staggering \textbf{96\%} success rate, exposing severe brittleness in memory-less defense. Furthermore, we reveal a \textbf{``Complexity Paradox''}: while complex, iteratively refined attacks are effective, they often trigger defensive compliance, whereas simple strategies achieve a higher rate of genuine persuasion (\textbf{84.7\%}). Our code and dataset are available at GitHub, https://github.com/cza1006/llm-persuasion-defense.
Chinese Translation
随着大型语言模型(LLM)日益成为主要的知识检索接口,它们抵御\textit{说服攻击}——即注入错误信息或强加反事实的尝试——的鲁棒性已成为一个关键的安全问题。现有的红队测试框架通常在多轮对话中评估模型,其中目标模型保留完整的对话历史。我们发现该设置中存在一个关键缺陷,称为\textbf{“拒绝惯性”}:模型最初的拒绝往往会在后续轮次中传播,主要是为了保持上下文一致性,从而掩盖了其对复杂、孤立的说服尝试的真实脆弱性。为了严格评估SOTA模型的“冷启动”防御能力,我们引入了\textbf{SAST-IR}(有状态攻击者,无状态目标——迭代精炼)框架。通过对目标强制进行记忆擦除,同时保留攻击者的历史记录,我们使用\textbf{多轮}(无状态)迭代模拟了一个最坏情况下的对抗设置。利用\textbf{CP-Agent}(认知说服智能体),一个增强的诊断引导智能体,我们在自定义\textsc{CounterFact-Strict}数据集($N=50$)上的实验得出了令人震惊的结果:简单、多样的攻击策略达到了惊人的\textbf{96\%}成功率,暴露出无记忆防御中的严重脆弱性。此外,我们揭示了一个\textbf{“复杂性悖论”}:虽然复杂、迭代精炼的攻击是有效的,但它们往往会触发防御性顺从,而简单策略则实现了更高的真正说服率(\textbf{84.7\%})。我们的代码和数据集可在GitHub获取:https://github.com/cza1006/llm-persuasion-defense。
cs.CL / 39 / 2609.16800
Smarter by the Moment: Environment-Driven Dynamic Policies for Continual LLM Improvement
因时而智:面向持续LLM改进的环境驱动动态策略
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) have achieved remarkable progress across diverse domains, but continual adaptation to evolving tasks and environments remains a key challenge. Existing memory-augmented approaches retrieve individual past examples as direct references, but do not explicitly synthesize actionable strategies from them, causing the same types of errors to recur. We propose Dynamic Retrieval-based Policy Generation (DRPG), a framework that integrates memory-based retrieval with a dynamic policy generator, leveraging historical data and environment feedback to produce task-specific policies for continual LLM improvement. We evaluate DRPG across six benchmarks spanning text-to-SQL, question answering, medical diagnosis, and Python programming, using seven LLMs from both proprietary and open-weight families. DRPG outperforms strong baselines across most datasets and models. Further analysis demonstrates that DRPG's policy generation is robust to retrieval strategy, operates effectively without prior policy continuity, and can leverage smaller or cross-family models as cost-efficient policy generators. We also find that the benefit of policy-level guidance depends on task characteristics, offering practical insights into when and under what conditions this mechanism is most effective.
Chinese Translation
大语言模型(LLMs)在多个不同领域取得了显著进展,但对不断演化的任务和环境的持续适应仍然是一项关键挑战。现有的记忆增强方法检索单个过往样例作为直接参考,但并未从中显式地综合出可执行的策略,导致同一类型的错误反复出现。我们提出基于动态检索的策略生成(Dynamic Retrieval-based Policy Generation, DRPG),这是一个将基于记忆的检索与动态策略生成器相结合的框架,利用历史数据和环境反馈来生成任务特定的策略,以实现LLM的持续改进。我们在涵盖文本到SQL、问答、医学诊断和Python编程的六个基准上评估了DRPG,使用了来自专有和开放权重家族的七个LLM。DRPG在大多数数据集和模型上优于强基线。进一步的分析表明,DRPG的策略生成对检索策略具有鲁棒性,在没有先验策略连续性的情况下也能有效运行,并且可以利用更小或跨家族的模型作为成本高效的策略生成器。我们还发现,策略层面指导的收益取决于任务特征,这为理解该机制在何时以及何种条件下最为有效提供了实践性洞见。
cs.CL / 40 / 2609.16854
A Data-free Universal Prior over Syntactic Structures
一种关于句法结构的无数据通用先验
large language model
大语言模型相关
Abstract
Probability is fundamental to theories of language comprehension, production, acquisition, and evolution, as well as to large language models. Existing theories estimate the probability of syntactic structures from language-specific data. Whether part of this probability structure can arise independently of language-specific experience remains unknown. Here I show that a universal prior over syntactic structures emerges from a cognitively motivated model of incremental language production, in which words are progressively integrated into syntactic structure through network growth. The resulting prior assigns probabilities to syntactic structures --represented as dependency trees-- without fitting parameters to linguistic data, and assigns higher probabilities to attested than to random trees in all 138 typologically diverse languages examined. These prior probabilities correlate positively with probabilities estimated from corpora in 33 of 34 languages. The results indicate that part of the probability structure of syntax can arise independently of language-specific statistical learning. Linguistic experience may therefore refine probabilities that are already structured by the process of language production, rather than create them from an initially uniform space. This identifies a possible cognitive origin for part of the probability distribution over syntactic structures, linking language production and statistical learning while providing a data-independent structural bias for probabilistic models of language.
Chinese Translation
概率对于语言理解、产生、习得和演化的理论,以及对于大语言模型,都是基础性的。现有理论从特定语言的数据中估计句法结构的概率。这种概率结构的某一部分能否独立于特定语言经验而产生,仍然未知。在这里,我表明,一种关于句法结构的通用先验从一种具有认知动机的增量语言产生模型中涌现出来;在该模型中,词语通过网络生长逐步被整合到句法结构中。由此得到的先验为句法结构——表示为依存树——分配概率,而无需将参数拟合到语言数据,并且在所考察的全部138种类型学上多样的语言中,它为实际出现的树分配的概率都高于为随机树分配的概率。这些先验概率与从语料库中估计出的概率在34种语言中的33种里呈正相关。结果表明,句法概率结构的一部分可以独立于特定语言的统计学习而产生。因此,语言经验可能细化那些已经由语言产生过程组织起来的概率,而不是从一个初始均匀的空间中创造它们。这为句法结构概率分布的一部分识别出一个可能的认知起源,将语言产生和统计学习联系起来,同时为语言的概率模型提供了一种独立于数据的结构偏置。
cs.CL / 41 / 2609.16890
Cascade: Hierarchical Recoverability Control for Large Language Model Unlearning
Cascade:用于大语言模型遗忘的分层可恢复性控制
large language model
大语言模型相关
Abstract
Large Language Model (LLM) unlearning is essential for removing sensitive or copyrighted knowledge while preserving general utility. Existing methods often leave residual knowledge in intermediate representations, which can still be recovered. To address this, we propose Cascade, a hierarchical recoverability control framework that minimizes the internal identifiability of target knowledge. Cascade combines three complementary controls: path-level routing to suppress privacy-associated activation routes, representation-level compression to reduce geometric separability, and decoding-level intervention to limit residual recovery. Experiments on TOFU, MUSE-News, and WMDP, including robustness tests with query reformulation and extraction-style prompts, show that Cascade effectively reduces recoverability while maintaining stable model utility.
Chinese Translation
大语言模型(LLM)遗忘对于移除敏感或受版权保护的知识,同时保持通用效用至关重要。现有方法常常在中间表示中留下残差知识,而这些知识仍然可以被恢复。为解决这一问题,我们提出 Cascade,一个分层可恢复性控制框架,它最小化目标知识的内部可识别性。Cascade 结合了三种互补控制:路径级路由以抑制与隐私相关的激活路径,表示级压缩以降低几何可分性,以及解码级干预以限制残差恢复。在 TOFU、MUSE-News 和 WMDP 上的实验,包括使用查询改写和提取式提示进行的鲁棒性测试,表明 Cascade 有效降低了可恢复性,同时保持稳定的模型效用。
cs.CL / 42 / 2609.16906
Deconstructing Stereotypes: Scope-Conditioned Generation for Effective Multilingual Counterspeech
解构刻板印象:面向有效多语言反制言论的范围条件化生成
large language model
大语言模型相关
Abstract
Counterspeech (CS) - direct responses that counter online Hate Speech (HS) using reasoning and alternative viewpoints - has emerged as an alternative to content removal. Current automatic CS generation methods, however, frequently produce generic, ineffective replies that fail to target the implicit stereotypes behind HS. To bridge this gap, we propose a novel scope-conditioned generation framework that explicitly integrates structured stereotype characteristics into Large Language Models prompts. We validate our approach on a novel, human-curated dataset annotated in English, Italian, and Spanish. Extensive evaluations show that stereotype-conditioned prompting substantially outperforms generic baselines across all three languages, obtaining significant gains in factuality, specificity, cogency, and effectiveness for both explicit and implicit implied stereotypes.
Chinese Translation
反制言论(Counterspeech, CS)——使用推理和替代观点来反制在线仇恨言论(Hate Speech, HS)的直接回应——已成为内容删除的一种替代方案。然而,当前的自动 CS 生成方法经常产生泛化、无效的回复,未能针对 HS 背后的隐性刻板印象。为弥合这一差距,我们提出一种新颖的范围条件化生成框架,该框架将结构化的刻板印象特征显式整合到大型语言模型提示中。我们在一个全新的、由人工整理并标注为英语、意大利语和西班牙语的数据集上验证了我们的方法。广泛评估表明,基于刻板印象条件的提示在所有三种语言上均显著优于通用基线,在事实性、具体性、说服力和有效性方面,对显式刻板印象和隐式暗示刻板印象都获得了显著提升。
cs.CL / 43 / 2609.16912
Lit3R: Retrieve-Relate-Read for Evidence-Grounded Question Answering over Scientific Literature
Lit3R:面向科学文献的基于证据问答的检索-关联-阅读
large language model
大语言模型相关
Abstract
We describe tus-nlp's Lit3R (Retrieve-Relate-Read) system for LitTraceQA, a shared task for literature-grounded question answering that requires systems to retrieve relevant papers, identify supporting evidence, and generate answers. Lit3R combines off-the-shelf retrieval, reranking, and large language model (LLM) components without task-specific training. The retriever iteratively combines BM25-based sparse and dense retrieval, cross-encoder reranking, and LLM-based verification, and complements retrieval based on the question with paper-to-paper expansion. The reader first identifies supporting evidence within individual papers and then synthesizes evidence across papers to produce the final answer and evidence trace. On the official test set, our system ranked 4th on the leaderboard. Our code is available at https://github.com/tus-ist-nlp/littraceqa.
Chinese Translation
我们描述了 tus-nlp 的 Lit3R(检索-关联-阅读)系统,用于 LitTraceQA,这是一个面向文献的问答共享任务,要求系统检索相关论文、识别支持证据并生成答案。Lit3R 结合了现成的检索、重排序和大语言模型(LLM)组件,无需任务特定的训练。检索器迭代地结合基于 BM25 的稀疏检索和稠密检索、交叉编码器重排序以及基于 LLM 的验证,并通过论文到论文的扩展来补充基于问题的检索。阅读器首先在单篇论文内识别支持证据,然后跨论文综合证据以产生最终答案和证据轨迹。在官方测试集上,我们的系统在排行榜上排名第 4。我们的代码可在 https://github.com/tus-ist-nlp/littraceqa 获取。
cs.CL / 44 / 2609.16993
The Role of Implicit and Explicit Demographic Signals in Large Language Model-based Student Assessment
隐性与显性人口统计信号在基于大语言模型的学生评估中的作用
large language model
大语言模型相关
Abstract
Large Language Models are now common in student assessment, but we know little about how student demographics affect their use. Sometimes, considering student demographics may be necessary -- for example, to improve readability for users with lower educational levels. However, it also risks being a cause of discrimination, e.g., when assigning lower scores to students from lower socioeconomic backgrounds. We set up controlled prompts to test 1) explicit demographic effects, where we mention demographic details directly, and 2) implicit effects, where we use conversation history as a demographic signal. We test these settings in three tasks: Automated Essay Scoring, Formative Feedback, and Metalinguistic Question Answering. We test six state-of-the-art LLMs on these tasks. In both explicit and implicit cases, the models pick up on demographic cues and can change their scoring, feedback, and answers accordingly. We find that LLMs frequently adjust the readability of feedback to education levels when these are explicitly mentioned. On the other hand, implicit conditions produce unpredictable biases, such as in question answering, where responses from lower-education levels receive lower sentiment scores. Our results provide clear evidence of demographic sensitivity in LLMs for educational assessment tasks.
Chinese Translation
大语言模型如今在学生评估中已很常见,但我们对学生的人口统计特征如何影响其使用知之甚少。有时,考虑学生的人口统计特征可能是必要的——例如,为教育水平较低的用户提高可读性。然而,它也有可能成为歧视的成因,例如,当给来自较低社会经济背景的学生打更低的分数时。我们设置了受控提示,以测试 1)显性人口统计效应,即我们直接提及人口统计细节;以及 2)隐性效应,即我们将对话历史用作人口统计信号。我们在三项任务中测试了这些设置:自动作文评分、形成性反馈和元语言问答。我们在这些任务上测试了六个最先进的大语言模型。在显性和隐性两种情况下,模型都会捕捉到人口统计线索,并能够相应地改变其评分、反馈和答案。我们发现,当教育水平被明确提及时,大语言模型会频繁地根据教育水平调整反馈的可读性。另一方面,隐性条件会产生不可预测的偏差,例如在问答中,来自较低教育水平的回答会获得更低的情感得分。我们的结果为大语言模型在教育评估任务中对人口统计特征的敏感性提供了明确的证据。
cs.CL / 45 / 2609.17119
An Empirical Study of Counterfactual Self-Explanations in LLMs
LLMs中反事实自我解释的实证研究
large language model
大语言模型相关
Abstract
Large language models can easily generate explanations for their own outputs, but such self-explanations are not necessarily faithful to the model's behavior. We study this issue through counterfactual self-explanations, where a model minimally edits an input so that its own prediction changes. Across sentiment analysis and natural language inference, we evaluate ten instruction-tuned models from the LLaMA-3 and Qwen-2.5 families, measuring faithfulness, minimality, and alignment with human-annotated rationales. Our results show that model scale is the strongest determinant of explanation quality: larger models are substantially more likely to generate counterfactuals that flip their own predictions and target decision-relevant evidence. In contrast, the rationale-guided condition produces edit-minimal counterfactuals that are also more human-aligned. However, it does not consistently improve faithfulness. Overall, counterfactual self-explanations can provide useful behavioral evidence about model decisions, but their reliability depends strongly on model capacity and should be empirically validated rather than assumed.
Chinese Translation
大型语言模型可以轻松地为其自身输出生成解释,但这种自我解释未必忠实于模型的行为。我们通过反事实自我解释来研究这一问题,其中模型对输入进行最小编辑,以使其自身的预测发生变化。在情感分析和自然语言推理任务中,我们评估了来自 LLaMA-3 和 Qwen-2.5 系列的十个指令微调模型,测量了忠实性、最小性以及与人工标注理由的对齐程度。我们的结果表明,模型规模是解释质量最强的决定因素:更大的模型显著更可能生成能够翻转其自身预测并针对与决策相关证据的反事实。相比之下,理由引导条件会产生编辑最小的反事实,并且也更与人类对齐。然而,它并未一致地提高忠实性。总体而言,反事实自我解释可以提供关于模型决策的有用行为证据,但其可靠性在很大程度上取决于模型能力,并且应通过实证验证,而非想当然地假定。
cs.CL / 46 / 2609.17310
Zero-shot narrative detection in social messaging
社交媒体消息中的零样本叙事检测
large language model
大语言模型相关
Abstract
This study investigates the zero-shot ability of large language models (LLMs) to identify and classify hidden narratives in social messages. Our research hypothesis is that LLMs' extensive contextual knowledge allows them to interpret messages on a deeper, pragmatic level, going beyond basic sentiment or topic analysis. Experiments on the Dipromats and SemEval datasets show that providing models with human-written narrative descriptions significantly improves performance, without the need of training examples. In contrast, automatically generated descriptions or the use of few examples (few-shot) often degrade accuracy due to subtle shifts in framing. The study also finds that ensemble methods, particularly majority voting, enhance robustness and that larger models perform best while also being less sensitive to prompt variations. The findings validate that LLMs can effectively detect strategic narratives in a zero-shot setting, and when combined with simple ensembling and human-written descriptions, they can rival supervised systems, offering a scalable solution for narrative detection, specially when there is no training data for the vast majority of domains.
Chinese Translation
本研究探讨了大语言模型(LLM)识别和分类社交消息中隐藏叙事的零样本能力。我们的研究假设是,LLM 广泛的语境知识使其能够在更深层的语用层面上解读消息,超越基本的情感或主题分析。在 Dipromats 和 SemEval 数据集上的实验表明,为模型提供人工撰写的叙事描述能够显著提升性能,且无需训练样本。相比之下,自动生成的描述或使用少量示例(少样本)往往会因框架的细微偏移而导致准确率下降。研究还发现,集成方法,尤其是多数投票,能够增强鲁棒性,并且更大的模型表现最佳,同时对提示词变化也更不敏感。这些发现验证了 LLM 能够在零样本设置下有效检测策略性叙事,并且当与简单的集成方法和人工撰写的描述相结合时,它们能够媲美有监督系统,为叙事检测提供了一种可扩展的解决方案,尤其是在绝大多数领域都没有训练数据的情况下。
cs.CL / 47 / 2609.17346
Where Should a Document Live: Context, Representations, or Parameters?
文档应当驻留在何处:上下文、表示,还是参数?
large language model
大语言模型相关
Abstract
To answer questions outside of their pre-training data, large language models (LLMs) need access to new information, which can be presented in the context window as documents, encoded into the model's parameters, or injected as latent representations. However, each of these methods comes with different efficiency, cost, and performance trade-offs, with no single winner. We present a controlled comparison of representation-based (KV-cache based) and parametric (fine-tuning-based) adaptation methods on five knowledge-intensive benchmarks. We show that in the oracle setting, Cartridges (KV) are the most accurate injection method at nearly every storage budget, outperforming parametric methods by 10 points. Compaction (KV) matches Cartridges only at low compression rates, lagging behind the parametric methods by 10 points at rates higher than $50\times$. In the more realistic multi-document retrieval scenario, Cartridges are the only method that matches in-context learning (ICL), leading the parametric methods by 29 points and Compaction by 15 points. Nonetheless, Cartridges are also the only method, besides full fine-tuning and large MLP adapters, that suffers from catastrophic forgetting, i.e., a 6% performance degradation on control benchmarks, with 13% in coding.
Chinese Translation
为了回答超出其预训练数据范围的问题,大型语言模型(LLMs)需要访问新信息,这些新信息可以作为文档呈现在上下文窗口中,编码到模型参数中,或作为潜在表示注入。然而,这些方法各自都具有不同的效率、成本和性能权衡,没有单一赢家。我们在五个知识密集型基准上,对基于表示的(基于 KV 缓存的)和参数化的(基于微调的)适应方法进行了受控比较。我们表明,在 oracle 设置下,Cartridges(KV)在几乎所有存储预算下都是最准确的注入方法,比参数化方法高出 10 个点。Compaction(KV)仅在低压缩率下与 Cartridges 相当,在压缩率高于 $50\times$ 时落后于参数化方法 10 个点。在更现实的多文档检索场景中,Cartridges 是唯一能与上下文学习(ICL)匹敌的方法,领先参数化方法 29 个点,领先 Compaction 15 个点。尽管如此,Cartridges 也是除全量微调和大型 MLP 适配器之外唯一会遭受灾难性遗忘的方法,即在控制基准上性能下降 6%,在编码中下降 13%。
cs.CL / 48 / 2609.17398
Enhancing Accessibility of Medical Texts through Large Language Model-Driven Plain Language Adaptation
通过大语言模型驱动的简明语言改编增强医学文本的可及性
large language model
大语言模型相关
Abstract
This paper addresses the challenge of making complex healthcare information more accessible through automated Plain Language Adaptation (PLA). PLA aims to simplify technical medical language, bridging a critical gap between the complexity of healthcare texts and patients' reading comprehension. Recent advances in Large Language Models (LLMs), such as GPT and BART, have opened new possibilities for PLA, especially in zero-shot and few-shot learning contexts where task-specific data is limited. In this work, we leverage the capabilities of LLMs such as GPT-4o-mini, Gemini-1.5-pro, and LLaMA for text simplification. Additionally, we incorporate Mixture-of-Agents (MoA) techniques to enhance adaptability and robustness in PLA tasks. Key contributions include a comparative analysis of prompting strategies, finetuning with QLoRA on different LLMs, and the integration of MoA technique. Our findings demonstrate the effectiveness of LLM-driven PLA, showcasing its potential in making healthcare information more comprehensible while preserving essential content.
Chinese Translation
本文探讨了通过自动简明语言改编(PLA)使复杂医疗信息更易获取这一挑战。PLA 旨在简化技术性医学语言,弥合医疗文本的复杂性与患者阅读理解能力之间的关键差距。大语言模型(LLM),如 GPT 和 BART 的近期进展,为 PLA 开辟了新的可能性,尤其是在任务特定数据有限的零样本和少样本学习场景中。在这项工作中,我们利用 GPT-4o-mini、Gemini-1.5-pro 和 LLaMA 等 LLM 的能力进行文本简化。此外,我们引入了混合智能体(Mixture-of-Agents,MoA)技术,以增强 PLA 任务中的适应性和稳健性。关键贡献包括对提示策略的比较分析、在不同 LLM 上使用 QLoRA 进行微调,以及集成 MoA 技术。我们的研究结果表明了 LLM 驱动的 PLA 的有效性,展示了其在保留关键内容的同时使医疗信息更易理解的潜力。
cs.CL / 49 / 2609.17515
What Breaks Under Pruning in Smart Homes, and When? Evaluating LLM Degradation Across Architectures and Task Complexity
智能家居中的剪枝会破坏什么,又会在何时发生?评估跨架构与任务复杂度下的大语言模型性能退化
large language model
大语言模型相关
Abstract
Pruning can reduce the deployment cost of large language models (LLMs), but its impact on context-grounded tool calling remains poorly understood. We systematically study pruning-induced degradation in smart-home tool calling across four LLMs spanning dense Transformer, dense hybrid, and mixture-of-experts (MoE) architectures, together with depth, width, hybrid, and expert pruning methods. After post-pruning supervised fine-tuning (SFT), we evaluate more than 19,500 instances from three smart-home datasets. Beyond aggregate task accuracy, we characterize degradation along two dimensions: action components (i.e., operation, device, argument, and value) and task complexity. Our results show that dense models have narrow safe pruning regions followed by sharp degradation, while MoE models tolerate substantially more pruning. Pruning degrades grounded specificity before schema-level intent, and aggressive dense pruning can induce systematic over-refusal. These findings highlight the importance of evaluating pruning beyond aggregate accuracy when selecting pruned LLMs for reliable tool execution.
Chinese Translation
剪枝可以降低大语言模型(LLM)的部署成本,但其对基于上下文的工具调用的影响仍鲜为人知。我们系统地研究了智能家居工具调用中由剪枝引起的性能退化,涵盖四个大语言模型,跨越稠密 Transformer、稠密混合以及专家混合(MoE)架构,并结合深度、宽度、混合与专家剪枝方法。在剪枝后监督微调(SFT)之后,我们评估了来自三个智能家居数据集的超过 19,500 个实例。除总体任务准确率之外,我们还从两个维度刻画性能退化:动作组件(即操作、设备、参数与取值)以及任务复杂度。我们的结果表明,稠密模型的安全剪枝区间较窄,随后会出现急剧退化,而 MoE 模型则能承受大得多的剪枝程度。剪枝会先损害基于上下文的具体性,之后才影响模式层级的意图,而激进的稠密剪枝可能引发系统性的过度拒答。这些发现凸显出:在为可靠的工具执行选择剪枝后的大语言模型时,不能仅以总体准确率来评估剪枝效果。
cs.CL / 50 / 2609.17516
When Should LLMs Abstain? Chain-of-Self-Questioning for Selective Risk Control
LLMs 应何时弃权?用于选择性风险控制的自我提问链
large language model
大语言模型相关
Abstract
Large language models can produce fluent answers when their factual support is weak. This paper introduces Chain-of-Self-Questioning (CoSQ), a prompt-only framework that makes answer commitment conditional on an explicit assessment of the information required to answer a question. We evaluate three CoSQ variants under seventeen conditions on the 817-item TruthfulQA multiple-choice validation set using eleven open-weight and hosted model families. In the final balanced-option protocol, Grounded-CoSQ at τ=0.90 reduces the mean unconditional wrong-commitment rate from 13.1% under chain-of-thought prompting to 8.9%, a 32.1% relative reduction, while increasing answered accuracy from 86.9% to 89.7% and answering 87.6% of questions. Both improvements hold for all eleven models and at every evaluated threshold. Critical-CoSQ and Adaptive-CoSQ provide neighboring operating points with 88.6% and 86.5% coverage, respectively, while remaining more reliable than the baseline. A secondary Natural Questions Short-Answer evaluation provides convergent open-form evidence. These findings show that self-assessment can support explicit, tunable answer-or-abstain decisions when an unsupported commitment is more costly than referral or review.
Chinese Translation
大语言模型在事实依据薄弱时也能生成流畅的答案。本文提出自我提问链(Chain-of-Self-Questioning, CoSQ),一个仅依赖提示词的框架,它使是否承诺作答取决于对回答一个问题所需信息的显式评估。我们在 817 项 TruthfulQA 多项选择验证集上,使用十一个开放权重和托管模型系列,在十七种条件下评估了三种 CoSQ 变体。在最终平衡选项协议中,τ=0.90 的 Grounded-CoSQ 将平均无条件错误承诺率从思维链提示下的 13.1% 降至 8.9%,相对降低 32.1%,同时将已作答准确率从 86.9% 提高到 89.7%,并回答了 87.6% 的问题。这两项改进在所有十一个模型上以及每一个评估阈值下均成立。Critical-CoSQ 和 Adaptive-CoSQ 分别以 88.6% 和 86.5% 的覆盖率提供相邻工作点,同时仍比基线更可靠。一项辅助性的 Natural Questions 短答案评估提供了收敛的开放形式证据。这些发现表明,当无支持的承诺比转介或复核成本更高时,自我评估可以支持显式、可调的作答或弃权决策。
cs.CR / 51 / 2609.16193
Permutation-Based Stegomalware in Large Language Models: Threats and Countermeasures
大型语言模型中的基于置换的隐写恶意软件:威胁与对策
large language model
大语言模型相关
Abstract
The difficulty of training large language models (LLMs), together with their ubiquity, raises the threat of stegomalware, where malicious payloads are embedded into model weights. Recent work has demonstrated the use of permutation symmetry in model weights to mitigate these threats, but failed to show neutralization of stegomalware across all weights for LLMs. In this paper, we demonstrate the full potential of behavior-preserving symmetries as a defense against stegomalware, as well as the risks these symmetries pose when exploited by attackers. For stegomalware neutralization, we improve upon previous work, demonstrating that it is possible to select permutations which displace all model parameters. This contrasts with previous methods which left a significant percentage of weights unaltered in LLMs. When used in an attack, we show that permutation symmetries can encode malware into the weights of a model in a way that is theoretically lossless, requires no retraining after encoding, and needs no payload-specific information in the extraction script---a combination of characteristics not previously seen in any single method. While theoretically lossless, permutation can in practice alter model behavior due to the accumulation of numerical error. We therefore quantify the loss in model performance associated with applying these methods, for both attack and defense, showing it to be minimal.
Chinese Translation
训练大型语言模型(LLM)的难度,加之其无处不在的普及性,带来了隐写恶意软件的威胁,即恶意载荷被嵌入到模型权重之中。近期研究已展示了利用模型权重中的置换对称性来缓解这些威胁,但未能证明能够对 LLM 中所有模型权重上的隐写恶意软件实现中和。在本文中,我们展示了保行为对称性作为隐写恶意软件防御手段的全部潜力,同时也展示了这些对称性在被攻击者利用时所构成的风险。在隐写恶意软件中和方面,我们改进了先前的工作,证明可以选择能够移动全部模型参数的置换。这与先前的方法形成对比,那些方法在 LLM 中留下了相当大比例的权重未被改变。当用于攻击时,我们表明置换对称性能够以理论无损的方式将恶意软件编码到模型权重中,编码后无需重新训练,且提取脚本中无需任何载荷特定的信息——这一特征组合此前在任何单一方法中都未曾出现过。尽管在理论上无损,但由于数值误差的累积,置换在实践中可能会改变模型行为。因此,我们量化了在攻击与防御两种情形下应用这些方法所导致的模型性能损失,并表明该损失极小。
cs.CR / 52 / 2609.16433
Evaluating the NIST Bugs Framework Against CWE as a Successor for Automated Vulnerability Classification
评估 NIST Bugs Framework 相较于 CWE 作为自动化漏洞分类继任者的表现
large language model
大语言模型相关
Abstract
Vulnerability classification based on root cause weaknesses is essential for numerous cybersecurity activities, where the Common Weakness Enumeration (CWE) serves as a public repository of such flaws. However, its overlapping entries create a non-orthogonal structure. The result is the same vulnerability being mapped to multiple weaknesses, complicating Root Cause Analysis (RCA) and triage. To address this, NIST Special Publication 800-231 introduces the Bugs Framework (BF), which organizes vulnerabilities into <cause, operation, consequence> triples and links such triples into a causal chain, so that a vulnerability carries its root cause and its sink together instead of a single terminal label. To date, however, BF has been specified but not evaluated regarding its performance against the challenges to automated classification. The evidence required for adoption has not been investigated empirically. We evaluate BF as a classification target and a complement to CWE using a systematically screened corpus of automated Common Vulnerabilities and Exposures (CVEs) linked to CWE research. We assess the reproducibility of CVE-to-BF classification through two evaluations. The first is qualitative: an anonymized inter-rater study in which 2 subject-matter experts (SMEs) independently mapped 13 CVEs onto the four BF axes. Annotators showed strong agreement on the cause and operation axes, while the attribute axis indicated fair agreement. We also tested our automated framework across two large language model (LLM) deployments under different budgets for reproducibility analysis. Despite limitations, such as evidence availability and the absence of retrievable fix commits for closed-source software, our findings support the claim that BF is a more structured and automation-friendly framework than CWE. Our exploration reveals specific gaps in BF, including under-specified guidance on attributes.
Chinese Translation
基于根本原因弱点的漏洞分类对众多网络安全活动而言至关重要,而通用弱点枚举(CWE)正是此类缺陷的公共存储库。然而,其条目之间存在重叠,从而形成了一种非正交的结构。其结果是同一个漏洞被映射到多个弱点,使根本原因分析(RCA)与分类处置(triage)变得复杂。为解决这一问题,NIST 特别出版物 800-231 引入了 Bugs Framework(BF),它将漏洞组织为 <原因, 操作, 后果> 三元组,并将此类三元组连接成一条因果链,从而使一个漏洞同时携带其根本原因与其汇聚点(sink),而不是仅仅带有一个终端的标签。然而迄今为止,BF 虽已被明确规范,但尚未针对其在应对自动化分类挑战方面的表现进行评估。采纳它所需的证据尚未得到实证研究。我们使用一个经系统筛选的语料库,将 BF 作为分类目标以及 CWE 的补充来加以评估,该语料库由与 CWE 研究相关联的、自动化获取的通用漏洞与暴露(CVE)构成。我们通过两项评估来考察 CVE 到 BF 分类的可复现性。第一项是定性的:一项匿名评分者间研究,其中 2 位领域专家(SME)独立地将 13 个 CVE 映射到 BF 的四个轴上。标注者在原因轴与操作轴上表现出高度一致性,而在属性轴上则表明一致性一般。我们还在不同预算下,跨两种大语言模型(LLM)部署测试了我们的自动化框架,以进行可复现性分析。尽管存在诸如证据可得性以及闭源软件缺乏可检索的修复提交(fix commits)等局限性,我们的发现仍支持如下论断:BF 是一个比 CWE 更具结构性、更便于自动化的框架。我们的探索揭示了 BF 中的一些具体空白,包括关于属性的指导说明不够充分。
cs.CR / 53 / 2609.16818
InceptionRAG: Stealthy Poisoning Attack Against Retrieval-Augmented Generation
InceptionRAG:针对检索增强生成的隐蔽投毒攻击
large language model
大语言模型相关
Abstract
Retrieval-augmented generation (RAG) systems enhance large language models (LLMs) with external knowledge but have been demonstrated to be vulnerable to corpus poisoning. Existing poisoning attacks against RAG largely focus on single-point explicit injection, where the malicious payload is fully encapsulated within a single document. Consequently, recent mitigation mechanisms have evolved to identify and diminish these threats effectively. In this paper, we first verify that existing mitigation mechanisms are insufficient for a new class of threats: indirect logic induction. Motivated by this observation, we introduce InceptionRAG, a stealthy attack mechanism that subverts the standard attack paradigm. Instead of injecting explicit malicious payloads, InceptionRAG fragments it into a chain of dormant passages. These passages appear harmless and can bypass existing mitigation mechanisms when examined separately. However, when retrieved together, they trigger LLMs to self-deduce target misinformation via multi-hop reasoning. To further improve the applicability of InceptionRAG in black-box settings, we propose zeroth-order suffix optimization (ZOSO) to automate the generation of authoritative suffixes. Extensive evaluations across three datasets and three LLMs demonstrate that InceptionRAG achieves an attack success rate exceeding 80% even under rigorous adversarial constraints. In particular, InceptionRAG shows superior evasion capabilities, effectively bypassing established defenses that mitigate traditional single-document injections. Our findings expose a concerning paradox: the stronger reasoning capabilities of LLMs increase their vulnerability to reasoning-based poisoning attacks. To mitigate potential misuse, we propose a document isolation-based defense, HODOR, which decouples adversarial logical dependencies.
Chinese Translation
检索增强生成(RAG)系统利用外部知识增强大语言模型(LLM),但已被证明易受语料库投毒攻击。现有针对 RAG 的投毒攻击主要集中于单点显式注入,其中恶意载荷被完全封装在单个文档内。因此,近期的缓解机制已经发展到能够有效识别并削弱这些威胁。在本文中,我们首先验证,现有缓解机制不足以应对一类新型威胁:间接逻辑诱导。受这一观察启发,我们提出 InceptionRAG,一种颠覆标准攻击范式的隐蔽攻击机制。InceptionRAG 不注入显式的恶意载荷,而是将其碎片化为一条由休眠段落构成的链。这些段落看起来无害,并且在被单独检查时能够绕过现有缓解机制。然而,当被一起检索到时,它们会触发 LLM 通过多跳推理自行推导出目标错误信息。为了进一步提升 InceptionRAG 在黑盒设置中的适用性,我们提出零阶后缀优化(ZOSO),以自动化生成权威性后缀。在三个数据集和三个 LLM 上的广泛评估表明,即使在严格的对抗约束下,InceptionRAG 仍达到超过 80% 的攻击成功率。特别是,InceptionRAG 展现出优越的规避能力,有效绕过用于缓解传统单文档注入的已有防御。我们的发现揭示了一个令人担忧的悖论:LLM 的推理能力越强,其就越容易受到基于推理的投毒攻击。为缓解潜在滥用,我们提出一种基于文档隔离的防御方法 HODOR,它解耦对抗性逻辑依赖。
cs.CR / 54 / 2609.16915
ROSETTA: Efficient and Accurate Privacy-Preserving LLM Decoding via Hybrid CKKS/TFHE Evaluation
ROSETTA:通过混合 CKKS/TFHE 评估实现高效且准确的隐私保护 LLM 解码
large language model
大语言模型相关
Abstract
Generative large language models (LLMs) have achieved state-of-the-art performance on many real-world tasks such as code generation and question answering. These models predominantly rely on an autoregressive decoding strategy that generates output tokens sequentially. However, their pervasive deployment raises serious privacy concerns, motivating private inference frameworks based on fully homomorphic encryption (FHE). A major limitation of existing FHE frameworks is their inefficiency in evaluating nonlinear operations, which incur substantial overhead and dominate the decode stage. In this paper, we propose ROSETTA, a hybrid CKKS/TFHE framework that overcomes this limitation. We first observe that nonlinear operations in the decode stage exhibit heterogeneous workload patterns, which can be handled effectively via a hybrid approach. We then realize this with two key contributions: 1) an adaptive segmented lookup-table protocol based on TFHE that enables efficient and accurate evaluation of nonlinear operations; and 2) a scheme-aware operator-selection framework that automatically assigns each nonlinear operator to CKKS or TFHE to minimize end-to-end decoding latency. We demonstrate that ROSETTA achieves up to $4.8\times$ Softmax speedup and $1.5$--$2.1\times$ end-to-end speedup over the SOTA framework CacheMir.
Chinese Translation
生成式大语言模型(LLM)在代码生成和问答等许多现实任务上已经取得了最先进的性能。这些模型主要依赖一种按顺序生成输出 token 的自回归解码策略。然而,它们的普遍部署引发了严重的隐私担忧,促使人们提出基于全同态加密(FHE)的隐私推理框架。现有 FHE 框架的一个主要局限在于其评估非线性运算时效率低下,这些运算会带来大量开销,并在解码阶段占据主导地位。在本文中,我们提出了 ROSETTA,一个克服这一局限的混合 CKKS/TFHE 框架。我们首先观察到,解码阶段的非线性运算表现出异构的工作负载模式,可以通过混合方法有效处理。然后,我们通过两项关键贡献实现这一点:1)一种基于 TFHE 的自适应分段查找表协议,能够高效且准确地评估非线性运算;以及 2)一种方案感知的算子选择框架,可自动将每个非线性算子分配给 CKKS 或 TFHE,以最小化端到端解码延迟。我们证明,相较于 SOTA 框架 CacheMir,ROSETTA 实现了最高 $4.8\times$ 的 Softmax 加速以及 $1.5$--$2.1\times$ 的端到端加速。
cs.LG / 55 / 2609.16448
Decentralized Gossip Learning and Federated Averaging for Histopathology Image Classification
用于组织病理学图像分类的去中心化 Gossip 学习与联邦平均
diffusion
扩散模型相关
Abstract
Breast histopathology analysis increasingly relies on distributed learning because direct data pooling across institutions is often restricted by privacy, governance, and communication constraints. This study compares server-based Federated Averaging (FedAvg), fully decentralized gossip learning, and Hybrid Gossip-FedAvg for invasive ductal carcinoma (IDC) patch classification. Experiments used 277,524 color image patches with patient-disjoint training, validation, and test partitions and a workload-balanced, Dirichlet-guided allocation across six nodes. Ring, random degree-3, and fully connected gossip topologies were evaluated together with sensitivity analyses for statistical heterogeneity, mixing coefficient, learning rate, model drift, prediction disagreement, calibration, clinically motivated operating points, communication payload, and patient-level IDC burden, together with auxiliary backbone robustness analyses. In the principal alpha=0.3 experiment, Hybrid Gossip-FedAvg achieved a test area under the receiver operating characteristic curve (ROC-AUC) of 0.8811, closely followed by FedAvg at 0.8801 and fully connected gossip at 0.8751. Across three independent patient-level repetitions, FedAvg and Hybrid Gossip-FedAvg obtained the same mean ROC-AUC of 0.9082, with standard deviations of 0.0037 and 0.0043, respectively. Hybrid achieved the highest mean area under the precision-recall curve of 0.8240, whereas FedAvg produced the lowest mean Brier score of 0.1335. Denser gossip graphs improved discrimination but increased theoretical model payload, while ring gossip remained sensitive to learning rate and mixing strength. Overall, FedAvg provided the most consistently reliable server-based baseline, topology-aware gossip offered a viable decentralized alternative, and Hybrid Gossip-FedAvg provided a balanced compromise between peer-to-peer diffusion and periodic global coordination.
Chinese Translation
乳腺组织病理学分析日益依赖分布式学习,因为跨机构直接汇集数据常常受到隐私、治理和通信约束的限制。本研究比较了基于服务器的联邦平均(FedAvg)、完全去中心化的 Gossip 学习以及 Hybrid Gossip-FedAvg 在浸润性导管癌(IDC)图像块分类中的表现。实验使用了 277,524 个彩色图像块,采用患者不相交的训练、验证和测试划分,并在六个节点间采用工作负载平衡、Dirichlet 引导的分配。评估了环形、随机度数为 3 和全连接的 Gossip 拓扑,并进行了针对统计异质性、混合系数、学习率、模型漂移、预测分歧、校准、临床动机操作点、通信负载和患者级 IDC 负担的敏感性分析,以及辅助骨干网络稳健性分析。在主要 alpha=0.3 实验中,Hybrid Gossip-FedAvg 达到了 0.8811 的测试受试者工作特征曲线下面积(ROC-AUC),紧随其后的是 FedAvg 的 0.8801 和全连接 Gossip 的 0.8751。在三次独立的患者级重复实验中,FedAvg 和 Hybrid Gossip-FedAvg 获得了相同的平均 ROC-AUC,为 0.9082,标准差分别为 0.0037 和 0.0043。Hybrid 达到了最高的平均精确率-召回率曲线下面积,为 0.8240,而 FedAvg 产生了最低的平均 Brier 分数,为 0.1335。更密集的 Gossip 图提高了判别能力,但增加了理论模型负载,而环形 Gossip 仍然对学习率和混合强度敏感。总体而言,FedAvg 提供了最一致可靠的基于服务器的基线,拓扑感知的 Gossip 提供了一种可行的去中心化替代方案,而 Hybrid Gossip-FedAvg 在对等扩散与周期性全局协调之间提供了一种平衡的折中。
cs.LG / 56 / 2609.16464
A multimodal large language model for evidence-based autism spectrum disorder screening
用于循证自闭症谱系障碍筛查的多模态大语言模型
large language model
大语言模型相关
Abstract
The clinical management of autism spectrum disorder (ASD) faces a bottleneck in early screening, mainly because trained specialists are scarce and conventional assessment tools are subjective. Here, we introduce ASDchat, a multimodal large language model designed for evidence-based ASD screening, which takes video, audio, and dialogue as input. ASDchat adopts a dual-branch architecture, where the decision branch generates screening probabilities and the evidence branch generates traceable, timestamped behavioral evidence aligned with standardized clinical criteria (ADOS-2). The model was trained and evaluated on a dataset of 1,035 participants from 27 sites in China, which covered typically developing (TD) children, children with ASD, and children with other disorders. For ASD versus TD, ASDchat reached an area under the receiver operating characteristic curve (AUC) of 0.953 $\pm$ 0.021. On 9 held-out sites that were not used for training, the mean AUC was 0.932. Furthermore, unsupervised clustering of the behavioral dimensions split the ASD cases into six subtypes with different phenotypic profiles, and ASDchat suggests an intervention for each subtype. ASDchat provides a feasible path for large-scale, evidence-based early ASD screening in clinical practice.
Chinese Translation
自闭症谱系障碍(ASD)的临床管理在早期筛查方面面临瓶颈,主要原因是受过培训的专科医生稀缺,且传统评估工具具有主观性。在此,我们介绍 ASDchat,一个为循证 ASD 筛查而设计的多模态大语言模型,它以视频、音频和对话作为输入。ASDchat 采用双分支架构,其中决策分支生成筛查概率,证据分支生成可追溯、带时间戳的行为证据,并与标准化临床标准(ADOS-2)保持一致。该模型在一个包含来自中国 27 个站点的 1,035 名参与者的数据集上进行了训练和评估,该数据集涵盖了典型发育(TD)儿童、ASD 儿童以及患有其他障碍的儿童。对于 ASD 与 TD 的区分,ASDchat 达到了 0.953 $\pm$ 0.021 的受试者工作特征曲线下面积(AUC)。在 9 个未用于训练的留出站点上,平均 AUC 为 0.932。此外,对行为维度的无监督聚类将 ASD 病例分为六个具有不同表型特征的亚型,ASDchat 为每个亚型建议了一种干预措施。ASDchat 为临床实践中大规模、循证的早期 ASD 筛查提供了一条可行路径。
cs.AI / 57 / 2609.16572
Efficient Text-to-Image Generation: An Adaptive Step Schedule Controller for Diffusion Models
高效的文本到图像生成:一种用于扩散模型的自适应步骤调度控制器
diffusion
扩散模型相关
Abstract
Text-to-image diffusion models often use a fixed number of denoising steps, balancing time costs and image quality. However, the optimal number of steps depends on the complexity of the input text prompt. We propose an adaptive diffusion controller that dynamically adjusts the number of steps to generate high-quality images efficiently, without additional model training. By leveraging a mixture of step schedules with varying step sizes and evaluating the error term discrepancy at each timestep, our method transitions between schedules to optimize performance. Experiments on COCO and DiffusionDB show that our approach reduces inference time while maintaining visual fidelity, offering a more efficient alternative for text-to-image diffusion models.
Chinese Translation
文本到图像扩散模型通常使用固定数量的去噪步数,以平衡时间成本和图像质量。然而,最优步数取决于输入文本提示的复杂性。我们提出了一种自适应扩散控制器,无需额外的模型训练,即可动态调整步数以高效生成高质量图像。通过利用具有不同步长的步骤调度混合,并在每个时间步评估误差项差异,我们的方法在调度之间切换以优化性能。在 COCO 和 DiffusionDB 上的实验表明,我们的方法在保持视觉保真度的同时减少了推理时间,为文本到图像扩散模型提供了一种更高效的替代方案。
cs.AI / 58 / 2609.17169
MUMINS: Metadata-conditioned Uncertainty-aware Medical Image Next-state Synthesis
MUMINS:元数据条件化的不确定性感知医学图像下一状态合成
diffusion
扩散模型相关
Abstract
Forecasting anatomical changes such as tumor growth and neurodegeneration is a challenging generative vision task. Morphological evolution is subtle relative to static anatomy, highly patient-specific, and inherently stochastic. Existing methods struggle with several issues: deterministic networks ignore biological stochasticity, while standard diffusion models require computationally prohibitive multi-pass sampling to quantify uncertainty. We propose MUMINS (Metadata-conditioned Uncertainty-aware Medical Image Next-state Synthesis), an efficient diffusion framework that jointly diffuses a baseline scan and its follow-up residual, summed to synthesize the follow-up scan, while concurrently predicting a spatial uncertainty map, in a single reverse diffusion process. Conditioned on the time interval and relevant metadata, it preserves fine-grained anatomy by dynamically re-injecting the baseline as a soft anchor at every denoising step, and a negative-log-likelihood head learns the uncertainty map to explicitly flag error-prone regions. Designed without organ-specific heuristics, the same architecture is reused across anatomies via separate, dataset-specific retraining. Extensive evaluations demonstrate that dataset-specific retraining of MUMINS matches or outperforms dedicated, domain-specific state-of-the-art methods on lung CT (PNG) and brain MRI (OASIS-3). Project page: https://github.com/aolivtous/MUMINS.
Chinese Translation
预测诸如肿瘤生长和神经退行性变等解剖学变化是一项具有挑战性的生成式视觉任务。形态学演化相对于静态解剖结构而言是微妙的、高度患者特异性的,并且本质上是随机的。现有方法面临若干问题:确定性网络忽略生物随机性,而标准扩散模型需要计算上难以承受的多轮采样来量化不确定性。我们提出 MUMINS(元数据条件化的不确定性感知医学图像下一状态合成),一种高效的扩散框架,它在单个逆向扩散过程中联合扩散基线扫描及其随访残差,并将二者求和以合成随访扫描,同时预测一张空间不确定性图。以时间间隔和相关元数据为条件,它通过在每一步去噪中动态重新注入基线作为软锚点来保留细粒度解剖结构,并且一个负对数似然头学习不确定性图,以显式标记易出错区域。该架构在设计上不使用器官特异性启发式方法,通过分别进行数据集特定的重训练,可在不同解剖结构间复用同一架构。广泛的评估表明,针对数据集特定重训练的 MUMINS 在肺部 CT(PNG)和脑部 MRI(OASIS-3)上匹配或优于专用的、领域特定的最先进方法。项目页面:https://github.com/aolivtous/MUMINS。
cs.AI / 59 / 2609.17227
FROD: Feature Matching Residual Denoising Oracle Bone Decipher
FROD:特征匹配残差去噪甲骨文破译
diffusion
扩散模型相关
Abstract
Oracle bone script (OBS), one of the earliest Chinese writing systems, plays an important role in the study of Chinese etymology. Traditional decipherment relies heavily on domain experts who analyze characters through semantic context and structural evolution. To assist this labor-intensive process, we formulate OBS decipherment assistance as a cross-era image translation task and propose FROD (Feature Matching Residual Denoising Oracle Bone Decipher). Although many OBS characters differ substantially from their modern counterparts, they often preserve local topological invariants at the radical level. During training, FROD leverages fast feature matching to provide gated segmentation supervision: paired samples with sufficient matches are processed patch-wise to align fine-grained radicals, whereas low-similarity pairs are trained holistically to avoid mismatched artifacts. In addition, a Residual Denoising Diffusion Model (RDDM) jointly estimates noise and residual signals, thereby reducing the positional drift and stroke disorder commonly observed in standard diffusion models. Finally, a multi-stage font stylization refinement network refines the generated images by eliminating edge noise and stabilizing stroke structures. On our augmented character-disjoint dataset, FROD achieves higher Top-1 recognition accuracy than the evaluated baselines, with a 3.8% absolute gain over OBSD.
Chinese Translation
甲骨文(OBS)是最早的汉字书写系统之一,在汉语词源学研究中发挥着重要作用。传统破译高度依赖领域专家,他们通过语义语境和结构演变来分析字符。为了辅助这一劳动密集型过程,我们将甲骨文破译辅助形式化为一个跨时代图像翻译任务,并提出 FROD(特征匹配残差去噪甲骨文破译)。尽管许多甲骨文字符与其现代对应字形存在显著差异,但它们通常在部首层面保留局部拓扑不变量。在训练过程中,FROD 利用快速特征匹配来提供门控分割监督:具有足够匹配的成对样本以逐块方式处理,以对齐细粒度部首,而低相似度样本对则进行整体训练,以避免不匹配伪影。此外,残差去噪扩散模型(RDDM)联合估计噪声和残差信号,从而减少标准扩散模型中常见的位置漂移和笔画紊乱。最后,一个多阶段字体风格化细化网络通过消除边缘噪声并稳定笔画结构来细化生成的图像。在我们增强的字符不相交数据集上,FROD 取得了比所评估基线更高的 Top-1 识别准确率,相较于 OBSD 获得 3.8% 的绝对提升。
cs.AI / 60 / 2609.16793
Available but Unclaimed: An Empirical Study of Human-AI Synergy
可用却未被取用:一项人机协同的实证研究
large language model
大语言模型相关
Abstract
People increasingly reason with large language models (LLMs), yet complementary capabilities do not guarantee outperforming both components. In a between-subjects study, participants (N=535) solved a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies, unaided or with GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash, or Kimi K3. Each assisted trial required consultation with the model. Each model answered every item alone 100 times under matched elicitation. The assisted-unaided accuracy difference increased with item-level LLM competence. Deference varied across tasks and increased with competence within tasks. Post-advice confidence distinguished correct from incorrect answers less strongly than unaided confidence. In a reference comparison, about half the increase in LLM accuracy carried through to assisted accuracy. How much of that accuracy gain reached participants differed across the models. These findings motivate evaluating LLMs in interaction with humans and designing support for selective deference that preserves independent reasoning.
Chinese Translation
人们越来越多地与大型语言模型(LLM)一起进行推理,然而互补的能力并不保证能够超越两个组成部分各自的表现。在一项被试间研究中,参与者(N=535)完成了一套包含40个题目的测验,内容涵盖矩阵推理、心理旋转、三段论和字母串类比,他们分别在无辅助条件下,或借助 GPT-5.6-Luna、Claude Opus 4.8、Gemini 3.6 Flash 或 Kimi K3 的条件下作答。每个辅助试次都要求与模型进行咨询。在匹配的诱导条件下,每个模型都对每个题目单独作答100次。辅助与无辅助之间的准确率差异随题目层面上LLM能力的提高而增大。依从程度在不同任务之间存在差异,并且在任务内部随能力的提高而增大。接受建议之后的信心对正确与错误答案的区分能力,弱于无辅助时的信心。在一项参照比较中,LLM准确率的提升约有半数传递到了辅助条件下的准确率上。这一准确率增益中有多少能够到达参与者,在不同模型之间有所不同。这些发现促使人们在与人交互的情境中评估LLM,并设计支持选择性依从而同时保留独立推理的辅助机制。
cs.LG / 61 / 2609.16155
LLMs as Master Forgers: Generating Synthetic Time Series Data for Manufacturing
LLMs 作为大师级伪造者:为制造业生成合成时间序列数据
large language model
大语言模型相关
Abstract
This paper presents a novel framework leveraging Large Language Models (LLMs) to generate synthetic time series data for manufacturing processes. Motivated by the scarcity of labeled time-series data in real-world manufacturing settings, which hinders the development of robust machine learning models, we explore the potential of LLMs to learn complex temporal dependencies and generate realistic synthetic data. Our approach involves fine-tuning pre-trained LLMs on manufacturing process instructions and employing a Retrieval Augmented Generation (RAG) technique to enhance data diversity and realism. We evaluate our method against traditional time series modeling techniques like ARIMA and LSTMs, using quantitative metrics, PCA analysis, and downstream task performance (anomaly detection). Results demonstrate that our LLM-driven framework outperforms these baselines, generating high-quality synthetic time series data that effectively captures temporal dependencies and statistical properties of real manufacturing data, leading to improvements in downstream task performance.
Chinese Translation
本文提出了一种新颖的框架,利用大型语言模型 (LLMs) 为制造过程生成合成时间序列数据。受现实制造环境中标注时间序列数据稀缺的驱动,这种稀缺阻碍了稳健机器学习模型的开发,我们探索了 LLMs 学习复杂时间依赖关系并生成逼真合成数据的潜力。我们的方法涉及基于制造过程指令微调预训练 LLMs,并采用检索增强生成 (RAG) 技术来增强数据多样性和逼真度。我们使用定量指标、PCA 分析和下游任务性能(异常检测),将我们的方法与 ARIMA 和 LSTMs 等传统时间序列建模技术进行评估。结果表明,我们的 LLM 驱动框架优于这些基线,能够生成高质量合成时间序列数据,有效捕获真实制造数据的时间依赖关系和统计特性,从而提升下游任务性能。
cs.LG / 62 / 2609.16161
LLM Inference in a Flash!
LLM 推理在闪存中实现!
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) have shown impressive capabilities across a range of natural language processing tasks, and LLM inference has emerged as a critical workload for enabling downstream applications. The demands of serving LLM inference are becoming increasingly challenging as requests shift toward longer sequences and heavier inference, driven by retrieval-augmented generation, inference-time compute scaling, and long-context applications. Additionally, these challenges are compounded by hardware trends, as memory capacity and communication bandwidth are not scaling as fast as increases in workload complexity. Compute-in-Flash is a promising solution to address memory bandwidth limitations by moving computation close to memory, and to exploit the large capacity of SSD technologies. However, it is challenging to deploy LLMs on these systems as they lack support for high-precision floating point operations and have limited write endurance. In our work, we aim to address these challenges by designing inference algorithms to enable LLM inference on Flash compute-in-memory devices. We present an end-to-end integer-only quantization approach to eliminate expensive floating-point computations. To address the limited write endurance, we design a dictionary-based KV cache compression strategy based on sparse dictionary coding that represents each KV vector as a linear combination of static dictionary vectors. These algorithmic improvements enable us to exploit the benefits of Compute-in-Flash for both model weights and KV cache, and to minimize expensive data transfer operations. Across Llama-3.1-8B and Qwen-2.5-7B, our combined method exhibits limited accuracy degradation while reducing dynamic KV cache traffic by 15$\times$.
Chinese Translation
大型语言模型(LLM)在一系列自然语言处理任务中已展现出令人印象深刻的能力,而 LLM 推理已成为赋能下游应用的关键工作负载。由于检索增强生成、推理时计算扩展和长上下文应用的推动,请求正转向更长的序列和更重的推理,服务 LLM 推理的需求正变得日益具有挑战性。此外,这些挑战还因硬件趋势而加剧,因为内存容量和通信带宽的扩展速度不及工作负载复杂度的增长。闪存内计算(Compute-in-Flash)是一种有前景的解决方案,它通过将计算移近内存来解决内存带宽限制,并利用 SSD 技术的大容量。然而,在这些系统上部署 LLM 具有挑战性,因为它们缺乏对高精度浮点运算的支持,并且写入耐久性有限。在我们的工作中,我们旨在通过设计推理算法来应对这些挑战,以在 Flash 存内计算设备上实现 LLM 推理。我们提出一种端到端纯整数量化方法,以消除昂贵的浮点计算。为了解决有限的写入耐久性,我们设计了一种基于稀疏字典编码的字典式 KV 缓存压缩策略,该策略将每个 KV 向量表示为静态字典向量的线性组合。这些算法改进使我们能够为模型权重和 KV 缓存二者利用 Compute-in-Flash 的优势,并最小化昂贵的数据传输操作。在 Llama-3.1-8B 和 Qwen-2.5-7B 上,我们的组合方法表现出有限的精度下降,同时将动态 KV 缓存流量减少 15$\times$。
cs.LG / 63 / 2609.16222
How I learned to stop worrying and love StopGrads: Stationarity, Convergence, and a case study on Flow Map Learning
我如何学会不再担忧并爱上 StopGrads:平稳性、收敛性,以及一个关于流映射学习的案例研究
diffusion
扩散模型相关
Abstract
Stopgrads are widely used in training machine learning models, but stopgrads can alter the gradient, stationary points and convergence guarantees of the original objective, which can make stopgrad training theoretically ungrounded. We introduce a stopgrad regression principle, which identifies a general template for stopgrad objectives with a closed-form characterization of stationary points and their uniqueness, unifying stopgrad objectives for flow maps, reinforcement learning, and diffusion samplers. We provide theoretical grounding for optimizing stopgrad flow map objectives by showing their unique stationary point is the true flow map, and showing positive convergence results for Eulerian and Lagrangian objectives, including MeanFlow and improved MeanFlow. Remarkably, we show that under functional semi-gradient flow, the learned flow map has a closed-form expression composing the initial flow map and the true flow map. We additionally use our stopgrad regression principle to propose modified stopgrad placements for flow map objectives which reduce training memory by 2x.
Chinese Translation
Stopgrad 在训练机器学习模型中被广泛使用,但 stopgrad 会改变原目标的梯度、驻点以及收敛保证,这可能使 stopgrad 训练在理论上缺乏根据。我们引入一个 stopgrad 回归原理,它识别出 stopgrad 目标的一般模板,并给出对驻点及其唯一性的闭式刻画,从而统一了面向流映射、强化学习和扩散采样器的 stopgrad 目标。我们通过证明其唯一驻点就是真实流映射,并给出欧拉式和拉格朗日式目标(包括 MeanFlow 和改进的 MeanFlow)的正向收敛结果,为优化 stopgrad 流映射目标提供了理论基础。值得注意的是,我们表明在函数式半梯度流下,学到的流映射具有一个闭式表达式,该表达式由初始流映射和真实流映射复合而成。我们另外使用我们的 stopgrad 回归原理,为流映射目标提出修改后的 stopgrad 放置方式,可将训练内存减少 2 倍。
cs.LG / 64 / 2609.16229
Test-Time Unlearning via Sparse Autoencoder
通过稀疏自编码器实现的测试时遗忘
large language model
大语言模型相关
Abstract
Machine unlearning aims to remove specific knowledge from a trained large language model (LLM) without retraining from scratch. Existing methods modify model weights via gradient ascent and its advances. While effective on certain benchmarks, these weight-based approaches exhibit a sharp forget-utility trade-off, where stronger forgetting of target knowledge can degrade model utility, and unlearned knowledge may reappear under post-unlearning fine-tuning or prompt attacks. We propose ARIA (autoencoder-gated inference-time unlearning), a test-time unlearning method that leaves model weights intact and gates access to unwanted knowledge only when generation enters a forget-related state. ARIA uses sparse autoencoder (SAE) latents to train a lightweight linear detector, then applies an interpretable intervention on triggered states with negligible test-time overhead. Empirical evaluations on TOFU, R-TOFU, and WMDP show that ARIA improves the forget-retain trade-off over weight-based baselines across both a thinking model (DeepSeek-R1-Distilled-Qwen-1.5B) and an instruction model (Gemma-3-1B-it), e.g., reducing WMDP-cyber forget-set accuracy significantly while keeping MMLU within 1% of the pre-unlearning model. We further introduce three post-unlearning adversarial attacks targeting weight-space and decoding-space recovery, and find that ARIA remains robust under all three, with forgetting changing by less than 1% under attack. A feature-level case study leveraging the interpretability of ARIA suggests that some retain degradation may reflect response styles underlying the unlearning data rather than leakage of the targeted knowledge itself, highlighting a potential source of bias in unlearning task construction.
Chinese Translation
机器遗忘旨在从已训练的大型语言模型(LLM)中移除特定知识,而无需从头重新训练。现有方法通过梯度上升及其改进来修改模型权重。虽然在某些基准测试上有效,但这些基于权重的方法表现出明显的遗忘—效用权衡,即对目标知识的更强遗忘可能会降低模型效用,并且被遗忘的知识在遗忘后微调或提示攻击下可能重新出现。我们提出 ARIA(自编码器门控的推理时遗忘),这是一种测试时遗忘方法,它保持模型权重不变,并且仅当生成进入与遗忘相关的状态时才门控对不想要知识的访问。ARIA 使用稀疏自编码器(SAE)潜变量来训练一个轻量级线性检测器,然后对触发的状态施加可解释的干预,且测试时开销可忽略不计。在 TOFU、R-TOFU 和 WMDP 上的实证评估表明,ARIA 在思考模型(DeepSeek-R1-Distilled-Qwen-1.5B)和指令模型(Gemma-3-1B-it)上都改善了相对于基于权重基线的遗忘—保留权衡,例如,显著降低 WMDP-cyber 遗忘集准确率,同时将 MMLU 保持在遗忘前模型的 1% 以内。我们进一步引入了三种针对权重空间和解码空间恢复的遗忘后对抗攻击,并发现 ARIA 在全部三种攻击下仍然稳健,攻击下遗忘变化小于 1%。一项利用 ARIA 可解释性的特征级案例研究表明,一些保留性能退化可能反映的是遗忘数据背后的响应风格,而不是目标知识本身的泄露,这突显了遗忘任务构建中一个潜在的偏差来源。
cs.LG / 65 / 2609.16436
Interpreting and Steering LLM Agents for Social Simulations
为社会模拟解释与引导 LLM 智能体
large language model
大语言模型相关
Abstract
Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the social scientific toolkit. However, LLMs are ultimately black boxes based on deep neural networks which limits their value for social science. This is because of a lack of (i) interpretability: i.e. the ability to assign clear mechanisms driving observed behavior; and a lack of (ii) steerability: i.e. the ability to mute or amplify specific theoretically meaningful mechanisms of action to drive specific model behavior. Here, we demonstrate how the black box could be opened up to further enrich LLM-based simulations. Specifically, we compare three types of methods: (1) prompt-based manipulation, (2) SAE-derived feature steering, and (3) probe-based direction steering and examine their utility for LLM-based social scientific simulations. We do so by interpreting and steering two foundational components of human behaviors, namely preferences (risk attitudes, altruism) and capabilities (divergent creativity, product innovation), operationalized using four classic economic and creative tasks implemented as natural-language interactions. Overall, our results show that SAE- and probe-based techniques often outperform basic prompt-based methods for steering LLM agents, although this advantage depends on the specific prompting strategy involved. Together, SAEs and probes constitute an effective pipeline for social scientists seeking to interpret and steer agents in social simulations: SAEs decompose agents' internal representations into human-readable features, after which probes can reliably shift agents' behaviors in specified directions. We discuss implications of these methods for future work using LLM agents for social scientific simulations.
Chinese Translation
基于大语言模型(LLM)的模拟已被证明对于理解人类行为十分强大,这使其成为社会科学工具箱中的宝贵补充。然而,LLM 归根结底是基于深度神经网络的黑箱,这限制了它们对社会科学的价值。这是因为缺乏 (i) 可解释性:即为观察到的行为指定清晰驱动机制的能力;以及缺乏 (ii) 可引导性:即抑制或放大特定的、具有理论意义的作用机制以驱动特定模型行为的能力。在这里,我们展示如何打开这个黑箱,以进一步丰富基于 LLM 的模拟。具体来说,我们比较三类方法:(1) 基于提示的操纵,(2) 由 SAE 导出的特征引导,以及 (3) 基于探针的方向引导,并考察它们对于基于 LLM 的社会科学模拟的效用。我们通过解释和引导人类行为的两个基础组成部分来做到这一点,即偏好(风险态度、利他主义)和能力(发散性创造力、产品创新),并使用四个经典经济学和创造性任务将其操作化,这些任务以自然语言交互的形式实现。总体而言,我们的结果表明,在引导 LLM 智能体方面,基于 SAE 和基于探针的技术往往优于基本的基于提示的方法,尽管这一优势取决于所涉及的具体提示策略。SAE 和探针共同构成了一条有效的流程,供那些寻求在社会模拟中解释和引导智能体的社会科学家使用:SAE 将智能体的内部表征分解为人类可读的特征,之后探针可以可靠地使智能体的行为朝指定方向转变。我们讨论了这些方法对于未来使用 LLM 智能体进行社会科学模拟工作的启示。
cs.LG / 66 / 2609.16528
FlowATC: Aircraft Trajectory Prediction via Flow Matching
FlowATC:基于流匹配的航空器轨迹预测
diffusion
扩散模型相关
Abstract
Building accurate decision-support tools for next-generation air traffic control requires robust trajectory prediction models. We present a flow-matching architecture trained exclusively on historical aircraft trajectories, with no route labels or chart supervision. Trained on 1.15 million Automatic Dependent Surveillance-Broadcast trajectory windows collected over the San Francisco Bay Area, the model generates aircraft trajectory distributions that closely match historical traffic, reproducing known airspace structure around San Francisco Airport such as the shape of SFO's published NIITE FOUR departure procedure. Our model is trained directly on the native, irregular ADS-B sampling interval. Trajectory prediction is cast as sequence inpainting using a block-causal Transformer that denoises future state tokens conditioned on the observed history using Conditional Flow Matching or Denoising Diffusion Probabilistic Models. We compare our architecture against constant-velocity, deterministic-Long Short Term Memory, and Conditional Variational Autoencoders baselines. At matched parameter count, CFM outperforms DDPM by 11-26% in minADE@20, and both generative objectives surpass the CVAE baseline by 31-41%. We further show that the error degrades gracefully with prediction horizon, and the architecture remains effective when retrained on temporally decimated feeds. Lastly, we sample $K$ independent completions, yielding spatial probabilistic occupancy estimates that can serve as input to downstream conflict-risk estimation.
Chinese Translation
为下一代空中交通管制构建准确的决策支持工具,需要稳健的轨迹预测模型。我们提出了一种流匹配架构,该架构仅基于历史航空器轨迹进行训练,不使用航路标签或航图监督。在旧金山湾区采集的115万条广播式自动相关监视(Automatic Dependent Surveillance-Broadcast)轨迹窗口上训练后,该模型生成的航空器轨迹分布与历史交通流高度吻合,能够重现旧金山机场周边已知的空域结构,例如SFO已公布的NIITE FOUR离场程序的形状。我们的模型直接基于原始的、不规则的ADS-B采样间隔进行训练。我们将轨迹预测表述为序列修补(sequence inpainting),使用块因果(block-causal)Transformer,以观测到的历史为条件,通过条件流匹配(Conditional Flow Matching)或去噪扩散概率模型(Denoising Diffusion Probabilistic Models)对未来状态 token 进行去噪。我们将该架构与匀速模型、确定性长短期记忆网络(deterministic-Long Short Term Memory)以及条件变分自编码器(Conditional Variational Autoencoders)基线进行比较。在参数量匹配的情况下,CFM在minADE@20上比DDPM优11-26%,且两种生成式目标均以31-41%的幅度超越CVAE基线。我们进一步表明,误差随预测时域的增加而平缓退化,并且在时间上抽稀的数据流上重新训练时,该架构仍然有效。最后,我们采样 $K$ 个独立的补全结果,得到空间概率占用估计,可作为下游冲突风险评估的输入。
cs.LG / 67 / 2609.16579
Recovering Physical Parameters from Fragmented Observations via Exact Distributed Spline Merging
通过精确的分布式样条合并从碎片化观测中恢复物理参数
diffusion
扩散模型相关
Abstract
Scientific measurements are frequently distributed across locations, time periods, and institutions. Combining such fragments into a continuous, differentiable field enables recovering governing physical parameters from its derivatives. This paper makes two contributions toward that goal. First, the established additive structure of fixed-basis ridge-regression statistics is applied to tensor-product spline fields: each data holder computes a local Gram matrix and moment vector, and the merged solution is mathematically identical to centralized fitting, with no raw data shared and no iterative synchronization. This property is specific to the fixed-feature squared-error setting; the present derivation does not establish an analogous guarantee for general jointly trained multilayer networks. Second, a complete pipeline connects distributed observations to physical parameter inference through field reconstruction, derivative extraction, and linear regression. The diffusion coefficient is recovered to 0.11% error and wave speed to 0.12% error; in both cases, distributed merging introduces zero degradation relative to centralized fitting. Application to 41 years of NOAA sea-surface temperature data confirms the result on real spatiotemporal observations.
Chinese Translation
科学测量常常分布在不同的地点、时间段和机构中。将此类碎片合并为一个连续、可微的场,使得能够从场的导数中恢复支配性物理参数。本文为实现这一目标做出两项贡献。第一,将固定基岭回归统计量已有的可加结构应用于张量积样条场:每个数据持有者计算一个局部 Gram 矩阵和矩向量,合并后的解在数学上与集中式拟合完全相同,无需共享原始数据,也无需迭代同步。该性质特定于固定特征下平方误差的设定;本文的推导并未为一般联合训练的多层网络建立类似保证。第二,一条完整的流程通过场重建、导数提取和线性回归,将分布式观测与物理参数推断连接起来。扩散系数恢复到 0.11% 误差,波速恢复到 0.12% 误差;在这两种情况下,相对于集中式拟合,分布式合并均带来零退化。对 41 年 NOAA 海表温度数据的应用在真实时空观测上证实了该结果。
cs.LG / 68 / 2609.16648
GrowMTP: Can RL Grow Its Own Draft Head?
GrowMTP:强化学习能否自行培育出它的草稿头?
large language model
大语言模型相关
Abstract
Reinforcement learning (RL) post-training drives the frontier capabilities of large language models, with its wall-clock dominated by autoregressive rollout generation. Speculative decoding is an established remedy for this bottleneck, but existing draft heads must be pretrained or warmed up before RL, introducing substantial training cost outside the RL run to be accelerated. We observe that RL training itself provides both conditions required for online draft-head training: its rollout distribution is far narrower than that of pretraining, and its verification step continuously produces supervision signals aligned with this distribution. Building on these observations, we propose GrowMTP, which uses this supervision to train a draft head from scratch entirely within the RL loop, with all head updates detached from the policy backbone. On Qwen3-4B (no draft head), MiMo-7B-SFT (weak head), and Qwen3.5-4B-Base (strong head), GrowMTP achieves rollout speedups of 2.13x, 1.93x, and 1.36x, and end-to-end speedups of 1.60x, 1.41x, and 1.20x, respectively. GrowMTP therefore serves existing RL training frameworks as a modular component, particularly offering a from-scratch acceleration path for models without pretrained draft heads.
Chinese Translation
强化学习(RL)后训练推动了大语言模型的前沿能力,而其挂钟时间主要被自回归的 rollout 生成所占据。投机解码是解决这一瓶颈的既有方法,但现有的草稿头必须在 RL 之前进行预训练或预热,这就在待加速的 RL 运行之外引入了大量的训练成本。我们观察到,RL 训练本身提供了在线草稿头训练所需的两个条件:其 rollout 分布远比预训练的分布狭窄,并且其验证步骤持续产生与该分布相对齐的监督信号。基于这些观察,我们提出 GrowMTP,它利用这种监督信号,完全在 RL 循环内部从零开始训练一个草稿头,并且所有草稿头的更新都与策略主干相分离(detached)。在 Qwen3-4B(无草稿头)、MiMo-7B-SFT(弱草稿头)和 Qwen3.5-4B-Base(强草稿头)上,GrowMTP 分别实现了 2.13 倍、1.93 倍和 1.36 倍的 rollout 加速,以及 1.60 倍、1.41 倍和 1.20 倍的端到端加速。因此,GrowMTP 可作为一个模块化组件服务于现有的 RL 训练框架,尤其为没有预训练草稿头的模型提供了一条从零开始的加速路径。
cs.LG / 69 / 2609.16853
Can Deep Learning Achieve Cross-Physics Mapping?
深度学习能否实现跨物理映射?
diffusion
扩散模型相关
Abstract
Can deep learning translate physical fields governed by fundamentally different equations? We address this question by introducing Cross-Physics Mapping (CPM), an operator-learning framework for mappings between heterogeneous physical domains. We formulate sufficient conditions for such mappings through compatible latent representations and propose a dimensionless scaling principle that aligns the characteristic evolution scales of the source and target systems without assuming their dynamical equivalence. As a representative test, paired diffusion and wave fields are generated independently from their respective parabolic and hyperbolic equations while sharing the same latent geometry, material heterogeneity, excitation, and dimensionless scale. Seven architectures-ResUNet, DeepONet, Fourier, latent, wavelet, U-shaped, and Galerkin neural operators-are evaluated for both diffusion-to-wave and wave-to-diffusion mappings. The results reveal a strong directional asymmetry. Diffusion-to-wave reconstruction is more challenging because it requires recovering wavefront, phase, and time-of-flight information attenuated by diffusion; U-NO performs best in this direction, achieving a relative $\ell_2$ error of $0.307$ and an $R^2$ of $0.905$. Wave-to-diffusion mapping is considerably more stable, with GNO attaining a relative $\ell_2$ error of $0.154$ and an $R^2$ of $0.935$. Neural operators generally outperform the conventional convolutional baseline, highlighting the nonlocal nature of cross-physics transformations. These findings demonstrate that deep learning can establish useful mappings between distinct physical modalities on a shared latent manifold, while the achievable accuracy remains fundamentally constrained by the direction-dependent information content of the governing physics.
Chinese Translation
深度学习能否转换由根本不同方程支配的物理场?我们通过引入跨物理映射(Cross-Physics Mapping, CPM)来回应这个问题,CPM 是一个用于异质物理域之间映射的算子学习框架。我们通过相容的潜在表示为这类映射建立充分条件,并提出一种无量纲缩放原理,该原理在不对源系统和目标系统的动力学等价性作出假设的情况下,对齐它们的特征演化尺度。作为一个代表性测试,成对的扩散场和波场分别独立地由各自的抛物型和双曲型方程生成,同时共享相同的潜在几何、材料非均质性、激励和无量纲尺度。七种架构——ResUNet、DeepONet、Fourier、latent、wavelet、U-shaped 和 Galerkin 神经算子——针对扩散到波和波到扩散两种映射进行评估。结果揭示出强烈的方向不对称性。扩散到波的重建更具挑战性,因为它需要恢复被扩散衰减的波前、相位和飞行时间信息;U-NO 在该方向上表现最佳,实现了 $0.307$ 的相对 $\ell_2$ 误差和 $0.905$ 的 $R^2$。波到扩散的映射则稳定得多,GNO 达到了 $0.154$ 的相对 $\ell_2$ 误差和 $0.935$ 的 $R^2$。神经算子总体上优于传统卷积基线,突显了跨物理变换的非局部性质。这些发现表明,深度学习能够在共享潜在流形上在不同物理模态之间建立有用的映射,而可达到的精度仍然从根本上受控制物理中依赖方向的信息内容所约束。
cs.LG / 70 / 2609.16937
Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation
超越词元局部的模仿:面向同策略蒸馏的奖励兼容时序信用分配
large language model
大语言模型相关
Abstract
On-policy distillation (OPD) has emerged as an effective approach for large language model post-training, yet existing objectives face a trade-off between objective fidelity and optimization stability. Token-level OPD provides stable but local supervision, whereas sequence-level OPD captures future credit at the cost of horizon-dependent variance. We establish a unified temporal-credit view of these formulations, showing that practical token-level OPD can be interpreted as a temporal approximation to the sequence-level reverse-KL gradient. Building on this connection, we propose $γ$OPD, which uses discounted temporal credit assignment to balance long-horizon supervision and optimization stability, while admitting a horizon-independent variance bound. We further develop a reward-compatible bounded mixing (RBM) mechanism for $γ\mathrm{OPD}$ that balances verifiable outcome feedback with the discounted OPD advantage to move beyond purely teacher-dependent optimization. Experiments on mathematical and code reasoning demonstrate consistent improvements over existing OPD methods across vanilla, size-mismatched, and multi-teacher distillation settings.
Chinese Translation
同策略蒸馏(OPD)已成为大语言模型后训练的一种有效方法,然而现有目标在目标保真度与优化稳定性之间面临权衡。词元级 OPD 提供稳定但局部的监督,而序列级 OPD 则以依赖于时间跨度的方差为代价来捕捉未来信用。我们为这些形式建立了一个统一的时序信用视角,表明实际的词元级 OPD 可以被解释为对序列级反向 KL 梯度的时序近似。基于这一联系,我们提出 $γ$OPD,它使用折扣时序信用分配来平衡长时程监督与优化稳定性,同时具有与时间跨度无关的方差界。我们进一步为 $γ\mathrm{OPD}$ 开发了一种奖励兼容的有界混合(RBM)机制,该机制在可验证的结果反馈与折扣后的 OPD 优势之间进行平衡,从而超越纯粹依赖教师的优化。在数学与代码推理上的实验表明,在普通、规模不匹配以及多教师蒸馏设置下,该方法相较于现有 OPD 方法均取得了一致的改进。
cs.LG / 71 / 2609.17194
MyoFlow: Anchor-Tied Rectified Flow for HD-sEMG Gesture Recognition Across Sessions and Subjects
MyoFlow:用于跨会话与跨被试 HD-sEMG 手势识别的锚点绑定整流流
diffusion
扩散模型相关
Abstract
High-density surface electromyography (HD-sEMG) gesture recognition supports prosthetic control, assistive robotics, and rehabilitation, but electrode re-donning and physiological variability cause distribution shifts that degrade accuracy across sessions and subjects. Generative HD-sEMG models primarily synthesize signals for augmentation; although diffusion models enhance representation learning, prediction still relies on a separate classifier. To tie learned dynamics to the decision rule, we propose MyoFlow, the first discriminative flow-matching framework for HD-sEMG recognition across sessions and subjects. It recasts classification as anchor-tied transport: a domain-conditioned rectified flow moves encoded windows toward gesture anchors that serve as transport targets and define the nearest-anchor decision geometry, enabling zero-shot prediction without an independent head. On the Hyser dataset, MyoFlow improves mean cross-session and cross-subject accuracy over the strongest diffusion-based baseline by 4.24\% and 6.37\%, respectively, and achieves 91.71\% mean zero-shot accuracy and 97.39\% mean few-shot accuracy across multiple days on the CEMHSEY dataset.
Chinese Translation
高密度表面肌电(HD-sEMG)手势识别可支持假肢控制、辅助机器人与康复,但电极重新佩戴和生理变异性会造成分布偏移,从而降低跨会话和跨被试的准确率。生成式 HD-sEMG 模型主要合成信号以用于数据增强;尽管扩散模型增强了表征学习,预测仍依赖于一个单独的分类器。为了将所学到的动力学与决策规则绑定,我们提出了 MyoFlow,这是首个用于跨会话和跨被试 HD-sEMG 识别的判别式流匹配框架。它将分类重新表述为锚点绑定传输:一个以域为条件的整流流将编码后的窗口移向手势锚点,这些锚点既充当传输目标,又定义了最近锚点决策几何,从而无需独立分类头即可实现零样本预测。在 Hyser 数据集上,MyoFlow 相较最强的基于扩散的基线,将平均跨会话准确率和平均跨被试准确率分别提升了 4.24\% 和 6.37\%,并在 CEMHSEY 数据集上跨多天实现了 91.71\% 的平均零样本准确率和 97.39\% 的平均少样本准确率。
cs.LG / 72 / 2609.17376
Large Language Models Develop Belief State Geometry In-Context
大型语言模型在上下文中发展出信念状态几何
large language model
大语言模型相关
Abstract
Large language models (LLMs) trained on next-token prediction exhibit remarkable in-context learning (ICL) abilities, yet the representations that support ICL remain poorly understood. We consider such representations in a controlled setting: prompting LLMs with data emitted from hidden Markov models (HMMs) and probing for the corresponding belief state -- the posterior distribution over the HMM's hidden states given the observed token history. Across six open-source LLMs prompted with data from 40 HMMs selected for non-trivial belief structure, we find that belief states are linearly decodable from residual stream activations, with peak probe $R^2$-values from 0.83-0.99 across HMM and LLM combinations, ranging from early to late layers. To establish functional relevance, we intervene directly on the probe-identified subspace via patching and steering, resulting in downstream prediction quality on the order of the untampered model, while controls degrade performance substantially. Together, these results provide representation-level evidence that ICL in open-source LLMs approximates optimal Bayesian prediction over a context-inferred generative model. More broadly, our findings extend prior results linking input-distribution structure to activation geometry: from toy networks trained explicitly on HMM data to production-scale LLMs.
Chinese Translation
基于下一词元预测训练的大型语言模型(LLMs)展现出非凡的上下文学习(ICL)能力,然而支撑 ICL 的表示仍然鲜为人知。我们在一个受控环境中考察此类表示:以隐马尔可夫模型(HMMs)生成的数据作为提示输入 LLMs,并探测相应的信念状态——即在给定观测到的词元历史的条件下,HMM 隐状态的后验分布。在六个开源 LLM 上,以来自 40 个因具有非平凡信念结构而被选出的 HMM 的数据作为提示,我们发现信念状态可以从残差流激活中线性解码,在不同 HMM 与 LLM 的组合下,探测峰值 $R^2$ 值介于 0.83-0.99 之间,且出现于从早期到后期的各个层。为确立功能相关性,我们通过修补与引导直接干预探针所识别出的子空间,其结果使下游预测质量达到与未经改动模型相当的水平,而对照组则会使性能大幅下降。综合来看,这些结果提供了表示层面的证据,表明开源 LLM 中的 ICL 近似于在一个由上下文推断出的生成模型上进行最优贝叶斯预测。更广泛地说,我们的发现拓展了将输入分布结构与激活几何联系起来的先前结果:从显式地在 HMM 数据上训练的玩具网络,一直延伸到生产规模的 LLM。
cs.LG / 73 / 2609.17474
Coupled Calibration and Learning: Mitigating Teacher Bias in LLM Distillation without Target-Domain Reward Feedback
耦合校准与学习:无需目标域奖励反馈缓解 LLM 蒸馏中的教师偏差
large language model
大语言模型相关
Abstract
Large language model (LLM) distillation aims to transfer the capabilities of a powerful teacher to a smaller student. Direct imitation, however, can also transfer the teacher's systematic bias and errors. This challenge is particularly pronounced under covariate shift, when the teacher's reliability on target questions is uncertain and target-domain reward feedback is unavailable. We propose Coupled Calibration and Learning (CCL), an LLM distillation algorithm that couples teacher calibration with student updates through token-level branching, using reward feedback only on source questions. Each iteration calibrates the teacher using source feedback and then uses the calibrated teacher to train the student on target questions. The updated student, in turn, informs subsequent calibration. In an autoregressive policy framework, we prove that the output student's expected average Kullback-Leibler divergence to the oracle student converges to zero at a polynomial rate in the number of iterations. The oracle maximizes the true reference-regularized target reward within the student class, which need not represent the unrestricted optimal policy. Our analysis quantifies the progress of projected student gradient updates while controlling the error in teacher calibration. We further establish a separation from regularized direct matching: its error relative to the oracle student can remain bounded away from zero even when the teacher achieves higher regularized target reward than every student policy. These results demonstrate that LLM distillation can overcome persistent teacher bias and recover the optimal student through coupled calibration and learning, without target-domain reward feedback.
Chinese Translation
大型语言模型(LLM)蒸馏旨在将强大教师的能力迁移到一个更小的学生。然而,直接模仿也可能迁移教师的系统性偏差和错误。这一挑战在协变量偏移下尤为突出,此时教师对目标问题的可靠性不确定,且目标域奖励反馈不可用。我们提出耦合校准与学习(Coupled Calibration and Learning,CCL),一种 LLM 蒸馏算法,它通过 token 级分支将教师校准与学生更新耦合起来,仅使用源问题上的奖励反馈。每次迭代使用源反馈校准教师,然后使用校准后的教师在目标问题上训练学生。更新后的学生反过来又为后续校准提供信息。在自回归策略框架中,我们证明输出学生到 oracle 学生的期望平均 Kullback-Leibler 散度以关于迭代次数的多项式速率收敛到零。oracle 在学生类内最大化真实的参考正则化目标奖励,而该 oracle 不必表示无限制的最优策略。我们的分析量化了投影学生梯度更新的进展,同时控制了教师校准中的误差。我们进一步建立了与正则化直接匹配的分离:即使教师比每一个学生策略都获得更高的正则化目标奖励,其相对于 oracle 学生的误差仍可能保持有界地远离零。这些结果表明,LLM 蒸馏可以通过耦合校准与学习,在没有目标域奖励反馈的情况下,克服持续存在的教师偏差并恢复最优学生。
cs.MA / 74 / 2609.16270
Cheap Talk Stabilizes Strategic Interaction in LLM Agents
廉价磋商稳定了 LLM 智能体中的策略互动
large language model
大语言模型相关
Abstract
Large language models are increasingly deployed as interacting agents, making the persistence of their action policies across repeated interaction critical for reliable multi-agent operation. We investigate whether and how agent-generated, non-binding pre-play communication ("cheap talk") increases such persistence in four open-weight 7-9B-parameter LLMs. Our experiments span four repeated two-player games -- Prisoner's Dilemma, Snowdrift, Stag Hunt, and Harmony -- with incentive structures ranging from strategic conflict to alignment, each presented in six contexts. We observe unstable trajectories in all four games, although their prevalence and magnitude depend strongly on model and context. Across models, games, and contexts, cheap talk is predominantly stabilizing, with five corrected reversals concentrated in social or team framings; effects vary substantially by model and context. Controlled current-message interventions identify two separable output-level channels in Qwen: reduced action uncertainty and less between-round drift in action probabilities. Matched history-by-message counterfactuals further show that recent partner behavior conditions how mutual-benefit versus self-prioritizing language affects policy persistence. Finally, in Prisoner's Dilemma, we identify in Qwen and Falcon a history-balanced policy-content direction in late transformer layers; projecting out this direction increases realized switching during closed-loop play, demonstrating that complete trajectories are causally sensitive to this component. Together, these findings show that cheap talk can make individual trajectories more persistent across diverse incentive structures, while revealing that the magnitude and mechanisms of stabilization are model- and history-dependent.
Chinese Translation
大语言模型越来越多地被部署为交互式智能体,这使得其行动策略在重复互动中的持续性对于可靠的多智能体运行至关重要。我们研究智能体生成的、非约束性的赛前沟通(“廉价磋商”)是否以及如何提高四个开放权重、参数量为 7-9B 的 LLM 中的这种持续性。我们的实验涵盖四个重复双人博弈——囚徒困境、雪堆博弈、猎鹿博弈和和谐博弈——其激励结构从策略冲突到利益一致不等,每种博弈均在六种情境中呈现。我们在所有四种博弈中都观察到不稳定轨迹,尽管其普遍性和幅度强烈依赖于模型和情境。跨模型、博弈和情境,廉价磋商主要起稳定作用,其中有五个经校正的反转集中于社会或团队框架;效应因模型和情境而显著变化。受控的当前消息干预在 Qwen 中识别出两个可分离的输出层面渠道:行动不确定性降低,以及行动概率在轮次间的漂移减少。匹配的“历史×消息”反事实进一步表明,近期伙伴行为会调节互利性语言相对于自我优先性语言如何影响策略持续性。最后,在囚徒困境中,我们在 Qwen 和 Falcon 的 Transformer 后层中识别出一个历史平衡的策略-内容方向;将这一方向投影剔除会增加闭环博弈过程中实际发生的切换,表明完整轨迹对该成分具有因果敏感性。总之,这些发现表明,廉价磋商能够使个体轨迹在多样化激励结构中更加持续,同时揭示稳定作用的幅度和机制依赖于模型和历史。
cs.NE / 75 / 2609.16846
LLMDE: A Large Language Model-Driven Differential Evolution Algorithm for Portfolio Optimization
LLMDE:一种用于投资组合优化的大语言模型驱动差分进化算法
large language model
大语言模型相关
Abstract
This study proposes a Large Language Model-Driven Differential Evolution (LLMDE) algorithm to reduce the reliance on handcrafted hyperparameter design. The proposed algorithm leverages a prompt engineering strategy, allowing large language models (LLMs) to dynamically select mutation strategies and configure control parameters guided by optimization feedback, thus enhancing the performance of the DE algorithm. We evaluate the performance of LLMDE on the CEC2022 benchmark suite, comparing it with standard DE and representative metaheuristics. Furthermore, we employ factor analysis and K-means clustering for stock selection, and then apply LLMDE to solve the Conditional Value at Risk (CVaR) portfolio optimization problem using the selected stocks, subject to budget and minimum expected return constraints. Experimental results demonstrate that LLMDE achieves competitive performance on the benchmark suite while continuously generating high-quality solutions for complex constrained optimization tasks. These outcomes successfully demonstrate the viability of embedding LLMs within metaheuristics, paving a promising path toward the design of advanced LLM-assisted optimization techniques.
Chinese Translation
本研究提出了一种大语言模型驱动的差分进化(LLMDE)算法,以减少对手工超参数设计的依赖。所提出的算法利用提示工程策略,使大语言模型(LLM)能够在优化反馈的指导下动态选择变异策略并配置控制参数,从而提升 DE 算法的性能。我们在 CEC2022 基准测试集上评估 LLMDE 的性能,并将其与标准 DE 和具有代表性的元启发式算法进行比较。此外,我们采用因子分析和 K-means 聚类进行股票选择,然后将 LLMDE 应用于求解使用所选股票的条件风险价值(CVaR)投资组合优化问题,并受预算和最低预期收益约束。实验结果表明,LLMDE 在基准测试集上取得了有竞争力的性能,同时能够为复杂的约束优化任务持续生成高质量的解。这些结果成功证明了将 LLM 嵌入元启发式算法中的可行性,为设计先进的 LLM 辅助优化技术铺平了一条有前景的道路。
cs.SE / 76 / 2609.16461
Protocol-Preserving Context Trimming for Agentic Workflows: Benefits, Failure Regimes, and Budget Guardrails
面向智能体工作流的协议保持型上下文裁剪:优势、失败情形与预算护栏
large language model
大语言模型相关
Abstract
Agentic large language model (LLM) systems rely on long interaction histories to preserve instructions, tool states, intermediate decisions, and unresolved dependencies, but unrestricted context growth increases computational cost and can reduce efficiency. This study evaluates protocol-preserving context trimming as a reliability-constrained approach for multi-step agentic workflows. Five trimming strategies - recency-based, relevance-based, summarization, protocol-aware trimming, and adaptive budget guardrails - were compared across retained-context levels and workflow-complexity classes using task success, protocol adherence, valid tool calls, token savings, latency reduction, cascading failures, and critical context thresholds. Conventional strategies achieved about 60% mean token savings but lower task success (66.6-77.3%) and protocol adherence (85.5-88.6%). Protocol-aware trimming improved task success to 92.2%, while adaptive guardrails achieved 96.0% task success, 96.3% protocol adherence, and 1.0% cascading failure with 56.0% mean token savings. Retained-context budgets of 25% or less increased failure odds 10.92-fold relative to budgets of 50% or more (p < 0.001). Protocol-aware trimming produced 5.24-fold greater odds of successful completion than conventional methods under aggressive budgets, while adaptive guardrails further increased success odds 2.11-fold versus fixed protocol-aware trimming (p < 0.001). Critical context thresholds also increased with workflow complexity. These findings indicate that reliable context reduction depends more on preserving protocol-critical state than on maximizing token removal, and that adaptive guardrails can improve efficiency, scalability, and reliability in long-horizon agentic systems.
Chinese Translation
智能体大语言模型(LLM)系统依赖较长的交互历史来保存指令、工具状态、中间决策和未解决的依赖关系,但不受限制的上下文增长会增加计算成本,并可能降低效率。本研究将协议保持型上下文裁剪评估为一种用于多步智能体工作流的可靠性约束方法。五种裁剪策略——基于新近性、基于相关性、摘要、协议感知裁剪和自适应预算护栏——在保留上下文水平与工作流复杂度类别上,使用任务成功率、协议遵循度、有效工具调用、token 节省、延迟降低、级联失败和关键上下文阈值进行了比较。传统策略实现了约 60% 的平均 token 节省,但任务成功率较低(66.6-77.3%),协议遵循度也较低(85.5-88.6%)。协议感知裁剪将任务成功率提高到 92.2%,而自适应护栏在实现 56.0% 平均 token 节省的同时,达到了 96.0% 的任务成功率、96.3% 的协议遵循度和 1.0% 的级联失败率。25% 或更少的保留上下文预算相对于 50% 或更多的预算使失败几率增加了 10.92 倍(p < 0.001)。在激进预算下,协议感知裁剪产生成功完成的几率比传统方法高 5.24 倍,而自适应护栏相对于固定协议感知裁剪进一步将成功几率提高了 2.11 倍(p < 0.001)。关键上下文阈值也随着工作流复杂度的增加而提高。这些发现表明,可靠的上下文缩减更多取决于保留协议关键状态,而不是最大化 token 移除,并且自适应护栏可以提高长时程智能体系统的效率、可扩展性和可靠性。
cs.SE / 77 / 2609.16936
RepoAtlas: Guiding Coding Agents via Evolving Multimodal Repository Views
RepoAtlas:通过演化的多模态仓库视图引导编码智能体
large language model
大语言模型相关
Abstract
Large language model (LLM)-powered coding agents have made rapid progress in automating software engineering tasks, yet repository-level issue resolution remains challenging. Beyond generating a plausible patch, an agent must localize relevant code across interdependent files and maintain repository context that is both sufficient and focused. Code graphs expose non-local relations, but linear text interfaces obscure their topology; rendering the full repository graph yields visual representations that are too dense to perceive reliably, whereas a one-shot local view becomes stale as exploration proceeds. We present \textbf{RepoAtlas}, a training-free module that maintains evolving multimodal repository views through a \emph{select--project--refresh} loop over a repository code graph. RepoAtlas combines evidence from the issue with the agent's current exploration state to select a task-relevant region under a fixed budget, projects the selected structure into complementary visual and textual representations, and refreshes the view when changes in the exploration state render it outdated. We evaluate RepoAtlas on SWE-bench Verified, where it improves the resolve rate by 2.4 points while reducing input tokens and model calls by 5.8\% and 7.8\% on average, relative to the strongest multimodal graph baseline, with consistent gains across three models of different families and scales.
Chinese Translation
由大型语言模型(LLM)驱动的编码智能体在自动化软件工程任务方面已取得快速进展,但仓库级问题解决仍然具有挑战性。除了生成一个看似合理的补丁之外,智能体还必须在相互依赖的文件中定位相关代码,并维护既充分又聚焦的仓库上下文。代码图能够揭示非局部关系,但线性文本接口会遮蔽其拓扑结构;渲染整个仓库图会产生过于密集、难以可靠感知的可视化表示,而一次性的局部视图会随着探索的推进而变得过时。我们提出 \textbf{RepoAtlas},一个无需训练的模块,它通过在仓库代码图上进行 \emph{select--project--refresh} 循环来维护不断演化的多模态仓库视图。RepoAtlas 将来自问题的证据与智能体当前的探索状态相结合,以在固定预算下选择与任务相关的区域,将所选结构投影为互补的视觉和文本表示,并在探索状态中的变化使其过时之时刷新视图。我们在 SWE-bench Verified 上评估 RepoAtlas,相对于最强的多模态图基线,它使解决率提高了 2.4 个百分点,同时平均将输入 token 数和模型调用次数分别减少了 5.8\% 和 7.8\%,并在三个不同系列和规模的模型上取得了一致的增益。
cs.SE / 78 / 2609.17338
Type-IV Code Clone Detection via Layer-Wise Non-Contrastive Representation Learning
基于逐层非对比表示学习的 Type-IV 代码克隆检测
large language model
大语言模型相关
Abstract
Software clones are fragments of code that are similar or functionally equivalent to each other. They pose significant challenges for maintenance, refactoring, and bug detection. Detecting Type-IV clones, which are semantically equivalent but may differ syntactically, is particularly difficult for traditional token- or syntax-based methods. Recent machine learning approaches rely on contrastive learning, which requires careful negative sampling and can introduce bias. In this paper, we propose LWVIC4Code, a non-contrastive representation learning approach specifically designed for Type-IV clone detection. Building on the Variance-Invariance-Covariance Regularization (VICReg) framework and prior layer-wise VICReg training, LWVIC4Code introduces cross-layer consistency regularization and depth-dependent layer weighting to progressively refine semantic information across transformer layers, producing robust and discriminative code representations. We conduct an empirical study comparing LWVIC4Code against a contrastive learning baseline and zero-shot large language models on Python (Kamino) and multi-language (GPTCloneBench) datasets. Results show that LWVIC4Code achieves competitive or superior performance without negative samples, benefits from layer-wise supervision, and generalizes effectively from Python to other languages, particularly Java and C#. These results demonstrate that non-contrastive, layer-wise representation learning is a promising direction for robust semantic code clone detection.
Chinese Translation
软件克隆是彼此相似或功能等价的代码片段。它们对维护、重构和缺陷检测构成了重大挑战。检测 Type-IV 克隆(它们在语义上等价,但在语法上可能不同)对于传统的基于词元或语法的方法而言尤其困难。近期的机器学习方法依赖于对比学习,而对比学习需要谨慎的负样本采样,并可能引入偏差。在本文中,我们提出 LWVIC4Code,一种专为 Type-IV 克隆检测设计的非对比表示学习方法。基于方差-不变性-协方差正则化(Variance-Invariance-Covariance Regularization, VICReg)框架以及先前的逐层 VICReg 训练,LWVIC4Code 引入了跨层一致性正则化和依赖深度的层加权,以逐步细化跨 transformer 层的语义信息,从而产生稳健且具有判别力的代码表示。我们在 Python(Kamino)和多语言(GPTCloneBench)数据集上开展了一项实证研究,将 LWVIC4Code 与一个对比学习基线以及零样本大语言模型进行比较。结果表明,LWVIC4Code 在无需负样本的情况下取得了具有竞争力或更优的性能,受益于逐层监督,并能有效地从 Python 泛化到其他语言,尤其是 Java 和 C#。这些结果表明,非对比的逐层表示学习是通向稳健语义代码克隆检测的一个有前景的方向。
cs.AI / 79 / 2609.16599
Large Language Models in the Loop: A Stability- and Network-Aware Survey in Networked Control, Cyber-Physical, and Multi-Agent Systems
大语言模型在环:网络化控制、信息物理与多智能体系统中的稳定性与网络感知综述
large language model
大语言模型相关
Abstract
Modern networked control systems (NCSs), cyber-physical systems (CPSs), and complex multi-agent network systems (CNSs) increasingly rely on large language models (LLMs) for high-level decision-making. However, the slow, stochastic nature of LLMs directly conflicts with the strict stability and safety guarantees required by these physical systems. This survey presents a unified analysis of how LLMs can be admitted into the control loop of NCS, CPS, and CNS without compromising closed-loop guarantees. We organize this around a core principle: the LLM operates as a slow supervisor adjusting high-level goals and constraints, while a fast, certified inner loop maintains physical stability. Under this framework, LLM integration maps directly to classical networked control challenges, where inference latency acts as delay, API failures as packet dropouts, tokenization as quantization, and hallucinations as bounded disturbances. We assess current developments across all these three domains, highlighting that rising model capabilities are frequently accompanied by a drop in formal safety assurances. Finally, we propose concrete future research directions, identifying the widespread lack of formal stability proofs as the field's central open problem.
Chinese Translation
现代网络化控制系统(NCSs)、信息物理系统(CPSs)以及复杂多智能体网络系统(CNSs)日益依赖大语言模型(LLMs)进行高层决策。然而,LLM 缓慢、随机的特性与这些物理系统所要求的严格稳定性与安全性保证直接冲突。本综述对如何在不损害闭环保证的前提下,将 LLM 纳入 NCS、CPS 和 CNS 的控制回路进行了统一分析。我们围绕一个核心原则来组织这一分析:LLM 作为缓慢的监督者运行,调整高层目标与约束,而快速且经过认证的内环维持物理稳定性。在这一框架下,LLM 集成直接映射到经典网络化控制挑战:推理延迟表现为时延,API 故障表现为丢包,词元化表现为量化,幻觉表现为有界扰动。我们评估了这三个领域当前的发展,并强调模型能力的提升往往伴随着形式化安全保障的下降。最后,我们提出具体的未来研究方向,指出广泛缺乏形式化稳定性证明是该领域的核心开放问题。
cs.AI / 80 / 2609.16527
QALPA: Property-guided diffusion modeling for efficient exploration of chemical spaces of flexible molecules
QALPA:用于高效探索柔性分子化学空间的性质引导扩散建模
diffusion
扩散模型相关
Abstract
Exploring the chemical space of flexible molecules remains challenging because the vast number of possible compounds and conformations, together with the increasing cost and limited generalization of 3D generative models for larger and more complex molecules, restrict access to unexplored chemistry. Here, we introduce QALPA ("Quantum-Aware Learning for Property-space Augmentation"), a property-guided generative framework that combines an E(3)-equivariant diffusion model with active learning and efficient quantum-mechanical (QM) methods to iteratively explore targeted QM property manifolds. By coupling generation with physics-based evaluation, QALPA improves molecular sampling and model reliability in sparsely populated regions of chemical space. Our results show that training on complementary QM datasets spanning both small (QM7-X) and large (Aquamarine) drug-like compounds enables accurate molecular generation across a broad size range, improving transferability beyond the training distribution for complex property manifolds involving both extensive and intensive properties. As a proof of concept, QALPA coupled with the machine learning-augmented tight-binding method EquiDTB efficiently augments alloQM, a QM dataset introduced in this work, comprising 6,253 conformers of allosteric drug molecules, by populating sparse regions of the property landscape defined by the many-body dispersion energy and HOMO-LUMO energy gap. These results demonstrate that the integration of generative AI with efficient ML/QM methods offers a practical pathway toward augmenting sparse QM datasets and sustainably expanding the exploration of chemical space for molecular discovery.
Chinese Translation
探索柔性分子的化学空间仍然具有挑战性,因为可能的化合物和构象数量庞大,加之用于更大、更复杂分子的3D生成模型成本日益增加且泛化能力有限,限制了对未探索化学空间的触及。在此,我们介绍 QALPA(“面向性质空间增强的量子感知学习”),一种性质引导的生成框架,它将 E(3)-等变扩散模型与主动学习和高效量子力学(QM)方法相结合,以迭代探索目标 QM 性质流形。通过将生成与基于物理的评估相耦合,QALPA 改善了化学空间中稀疏分布区域的分子采样和模型可靠性。我们的结果表明,在涵盖小分子(QM7-X)和大分子(Aquamarine)类药物化合物的互补 QM 数据集上进行训练,能够在广泛的尺寸范围内实现准确的分子生成,并针对同时涉及广延性质和强度性质的复杂性质流形,提高超出训练分布的可迁移性。作为概念验证,QALPA 与机器学习增强的紧束缚方法 EquiDTB 相结合,通过填充由多体色散能和 HOMO-LUMO 能隙所定义的性质景观中的稀疏区域,高效地增强了 alloQM——本工作引入的一个 QM 数据集,其包含 6,253 个变构药物分子的构象异构体。这些结果表明,将生成式 AI 与高效 ML/QM 方法相结合,为增强稀疏 QM 数据集以及可持续地扩展用于分子发现的化学空间探索提供了一条切实可行的路径。
人工智能 (cs.AI)
86
cs.AI / 1 / 2609.17123
AI for Science with GPT-6 Astra: Thermal Design and Electrothermal Analysis of 2D CFET
Abstract
Thermal optimization of 2D CFET inverters requires testing structural proposals against their electrical costs. We examine these research tasks using an AI agent workflow within a supplied electrothermal model. At 12 nm, Astra selects a redistributed source-interconnect geometry, while a coordinating agent proposes a substrate-directed heat-removal path. The combined design reduces peak temperature rise by 1.67 K at fixed metal volume and 20 μW. A subsequent metal-resistance sensitivity gives about 0.6-K inverter cooling alongside a 2% nFET on-current loss. Effective contact-length scaling further shows that lower temperature can accompany higher thermal resistance when current falls. Reproduction identifies agreeing implementations and retains a 104.95-K failure for diagnosis. These results show that an AI scientist workflow can propose thermal structures, test them under common constraints, and quantify their electrical cost.
cs.AI / 2 / 2609.16129
Optimal Pruning for Neural Architectures using Fisher Information Distances
Abstract
A new scheme for parameter pruning is introduced, derived from the differential-geometric distance in model space. Pruning a parameter sets its value to zero, representing a displacement of the model to the hypersurface on which that parameter vanishes. The minimal distance from the unpruned model to this hypersurface is naturally computed via the geodesic distance in the model space as determined by the Fisher information metric. This distance determines the true change in the model, and its performance, under pruning. By analysing progressively more faithful approximations of this geodesic distance a natural hierarchy of optimality for pruning methods is determined. This starts with the traditional magnitude pruning, then develops into new more sophisticated and effective pruning schemes. The method is demonstrated for both fully-connected networks and vision transformers, on MNIST and CIFAR-10, over the complete $0$-$100\%$ pruning range and across five random seeds. It outperforms pruning by parameter magnitude and by the local Fisher information alone in every architecture and dataset combination considered, on both accuracy and the Matthews correlation coefficient. Additionally, analysis of different levels of geodesic approximation produces intermediate pruning schemes that are computationally efficient and maintain near-optimal performance. This geometric picture supplies not only a state-of-the-art pruning methodology for AI models, but also a verified and mathematically-motivated justification for pruning schemes.
cs.AI / 3 / 2609.16145
Safe Error Correction for Language Models: Frozen-Base Adjustment with Capability Preservation
Abstract
We study a practical question: can a small correction module fix errors in a frozen language model's outputs without degrading its base capabilities? We propose CRN v2, a lightweight logit-level correction module (~34M trainable parameters, 0.73% of the 4.65B text module) that sits atop a fully frozen Gemma 4 E2B model. The base model is never updated; only the correction module learns, via supervised fine-tuning followed by reference-free DPO on 83,400 error-correction pairs. On a 60-question domain exam (CEHRI: Certified Human-Robot Intelligence, covering facts, arithmetic, and implicit-goal reasoning), CRN v2 corrects 53.3% of base-model errors (reworded variant: 43.3%) while showing no degradation on tested capability benchmarks (MMLU/BoolQ N=200; car-wash N=8). A LoRA baseline at the matched CRN v1 budget (6.6M params, rank 19) achieves 83.3% correction but suffers 30-75% capability loss on the same benchmarks -- the correction-capability tradeoff. An ablation shows that the KL preservation term (lambda=0.1) is critical: lowering it to 0.01 degrades correction to 35.0%. A hidden-state injection variant at earlier layers (1.6M params, SFT-only) reaches 50.0%/55.8% but does not exceed logit correction; shallower injection (layer 4) drops to 30.0%/28.3%; multi-depth logit correction (~35M) reaches only 40%; and longer training (5,000 SFT + 2,000 DPO) stays at 53.3% -- none of the alternative configurations we tested exceeded the rank-128 logit result, consistent with a best-achieved result of ~53% rather than a floor. This is a study of a design principle (frozen base + logit correction + KL anchoring), not a claim of architectural novelty. All code, main-result weights, and evaluation scripts are released (deep variant as code only -- no trained deep checkpoints).
cs.AI / 4 / 2609.16163
GPEvac: GNN-Based PPO for Adaptive Evacuation Routing During Shooting Events
Abstract
The sharp increase in mass shootings underscores an urgent need for systems that guide victims to safety in real time. An effective evacuation system must minimize threat exposure while also accounting for adversarial uncertainty and crowding dynamics. Current methods in the literature are rigidly constrained to layout-specific policies and computationally intractable in large-scale layouts, while practical guidelines simply advise victims to "run", "hide", or "fight". We propose GPEvac: a GNN-based PPO framework that computes adaptive evacuation routes during shooting events. To capture both local and long-distance dependencies, we introduce an edge-first sequential message-passing scheme with a learnable virtual global node. The resulting graph embeddings are integrated into a permutation-invariant scoring mechanism that allows a single learned policy to operate across building layouts of diverse topologies and sizes. Through extensive simulation, we show that GPEvac outperforms intelligent baselines across distinct architectural layouts, significantly reducing total threat exposure. Crucially, the system computes global evacuation routes in just 14.73 ms on local CPU hardware, enabling seamless integration with live surveillance systems. In addition to saving lives during shooting events, the methodologies developed are transferable to other graph-structured decision-making domains, including critical infrastructure, intelligent transportation systems, and adaptive sensor networks.
cs.AI / 5 / 2609.16206
Calibrate, Then Route: A Measured Study of Learned Request Routing for Disaggregated LLM Serving
Abstract
Disaggregated LLM serving places compute heavy prefill and memory heavy decode on separate GPU pools. Systems such as DistServe, Splitwise, and Mooncake make this separation fast, but routing still determines which instances handle each request. We study a router that estimates the additional completion time on each instance using exact prompt length, predicted output length, post admission KV cache pressure, and SLO class. We develop the policy in a discrete event simulator and validate it on eight NVIDIA A40 GPUs, each running a vLLM engine, with NIXL transferring KV caches between pools. All workloads run at measured saturation. Across three mixed, bursty arrival traces, the calibrated router achieves the highest mean goodput at 0.864, compared with 0.835 to 0.847 for round robin, least loaded, and a length heuristic. It also shows the lowest variance across traces. It beats round robin and the length heuristic on all three traces and least loaded on two. On the third, it trails by 0.003, within run to run noise. Hardware calibration matters: simulator derived constants cost 4.5 goodput points and roughly 40 percent of the tail latency advantage, reducing the scorer to little more than queue counting. Benefits grow with decode pool size and traffic heterogeneity but disappear in pools with three instances, where queue counts are often enough. Under extreme scarcity, greedy cost minimization concentrates requests on the cheapest scored instance, and blind spreading performs better. With calibrated costs, the learned router matches the goodput of round robin using six GPUs instead of seven.
cs.AI / 6 / 2609.16215
Where Should the KV Cache Live? Placement Policies Across GPU, CPU, and SSD for Long-Lived Sessions
Abstract
GPU high bandwidth memory is scarce and expensive, and KV caches consume much of it as chats, agent loops, and document question answering accumulate state. Systems such as Mooncake, LMCache, FlexGen, InfiniGen, and AttentionStore extend GPU memory with CPU DRAM and SSD. The harder question is which blocks belong in each tier, when to move or evict them, and whether prefetching helps. We study these choices in a discrete event simulator spanning GPU HBM, CPU DRAM, and SSD, calibrated against a random forest execution time predictor. We compare recency, reuse frequency, predicted reuse, and an EWMA predictor with prefetch lookahead across chat, agent, and document question answering workloads. Tiering supports 73.02 times more concurrent sessions per GPU and lowers cost per session by 62.04 times. These gains come from tier capacities of 1 plus 8 plus 64, not placement policy. Decode is compute bound at batch size one in our setup, so placement barely affects throughput. It mainly changes PCIe migration traffic and time to first token. Recency produces 2.30 times less migration traffic than reuse frequency for chat. Reuse frequency performs best for agents and document question answering. The existing predicted reuse policy is byte identical to recency, making its agent recommendation effectively recency. A genuine EWMA predictor changes behavior but still ranks behind reuse frequency on the workloads prediction was expected to help. Prefetching does not justify its bandwidth cost. Across the policy and cache size grid, even an oracle with knowledge of future requests never beats no prefetch on migration traffic. Workload specific placement can reduce data movement, but the predicted reuse and prefetch recommendations are not supported as implemented.
cs.AI / 7 / 2609.16245
Metacognitive Steering: Learning the Structure of Scientific Judgment
Abstract
Long-horizon scientific discovery requires agents to alternate between exploration, disciplined execution, and critical reassessment as evidence changes. Current language models are trained primarily on the products of science and optimized using outcome-level signals, providing limited supervision for these process-level shifts in scientific judgment. We investigate whether such judgment can be recovered from scientist interaction traces and used to control the internal computation of a frozen frontier model. Using contrastive interventions collected during real scientific research, we identify a coordinated, low-dimensional control structure within Kimi 2.6, a trillion-parameter mixture-of-experts model. Residual analysis, attention-weight subspace alignment, and cross-layer singular value decomposition converge on a mid-depth control surface spanning key layers. We introduce Metacognitive Steering, an inference-time controller that reads the model's cognitive regime and dynamically composes layer-specific interventions for exploration, procedural convergence, or critical reassessment without modifying model parameters. Behavioral analyses show that this control produces more sustained exploration, explicit pruning, and evidence-responsive synthesis. We operationalize the method in Columbus-1, an autonomous research system that identified eight independently reproduced, attacker-reachable vulnerabilities in BlueZ and directed the design, simulation, and fabrication of a ten-foot rocket intended to land propulsively using non-throttleable solid motors. Together, these results show that process-level scientific judgment can provide supervision for interpretable, dynamic control over a model's reasoning strategy.
cs.AI / 8 / 2609.16251
CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design
Abstract
Computer-use agents are increasingly evaluated in realistic desktop environments, but existing benchmarks provide limited coverage of professional engineering workflows whose outputs are persistent, structured artifacts. Mechanical computer-aided design (CAD) is a particularly demanding setting: an agent must manipulate geometry and constraints over long interaction horizons while producing a native project whose dimensions, construction structure, and downstream engineering state remain valid. We introduce \textbf{CADWorld}, a benchmark for long-horizon computer use in FreeCAD. CADWorld contains 200 tasks spanning 11 mechanical-CAD workflow categories, including sketching, part modeling, assembly, CAM, FEM, measurement, mesh processing, and technical drawing. Agents operate through screenshots and GUI actions, while success is determined by task-specific executable checks over saved FreeCAD artifacts and auxiliary outputs, covering geometric properties, parametric structure, constraints, manufacturing state, and simulation results. Across seven current agents on the full benchmark, the strongest agent achieves 17.5\% success, compared with an 87.0\% expert reference pass. We find that weaker agents often fail before producing a valid artifact, whereas stronger agents increasingly fail on structural, geometric, and construction-process requirements. CADWorld therefore exposes a gap between general GUI competence and reliable execution of persistent, verifiable engineering workflows. Project accessible at https://cad-world.github.io.
cs.AI / 9 / 2609.16258
The AI-Enabled Scientific Frontier
Abstract
As artificial intelligence's capabilities improve, it is increasingly viewed as a general scientific method. But how true are these claims? Does AI outperform all techniques, or only some, and how is this changing? To assess the claims, we assemble a corpus of 2,507 head-to-head comparisons between AI and other scientific analysis techniques across 27 scientific disciplines from papers published between 2000 and early 2025. We find a profound dichotomy. Relative to traditional statistics, AI often outperforms, but at a significantly higher computational cost. But there are also nearly a quarter of cases where AI is both more expensive and performs worse than traditional statistical techniques and this fraction has been stable for a decade. Relative to scientific computing, AI often underperforms, but at lower computational cost. This has begun to change: since 2020, AI's performance against scientific computing has notably strengthened and it now outperforms on more than half of comparisons. These patterns suggest that AI is therefore not a universal replacement for existing methods, but rather a valuable -- and improving -- part of a new AI-enabled scientific frontier.
cs.AI / 10 / 2609.16298
Closing the Loop: Branch-and-Bound for Scalable Verification of Nonlinear Neural Feedback Systems
Abstract
Despite recent advances in the verification of nonlinear neural feedback systems, scalability remains the central obstacle, as state-of-the-art solvers do not yet handle the network sizes and nonlinear dynamics of autonomy applications. Combinatorial solvers do not scale to large networks, whereas propagative solvers excessively sacrifice precision. This work seeks to improve the scalability of combinatorial solvers by formulating verification as branch-and-bound on an abstraction of the closed-loop system. We introduce \rail, an interface that exposes polyhedral enclosures of the dynamics to LiRPA-style bound propagation, and \clipper, a branch-and-bound algorithm that jointly refines enclosures and splits controller activations. This framework enables joint reasoning on the computational graph of the closed-loop system, preserving symbolic correlations across time steps. We present our construction and show that it yields significant improvements over the state of the art.
cs.AI / 11 / 2609.16322
Cross-Anatomy Transfer Versus Sparse Interpolation in Digital-Twin-Oriented Aortic Fluid-Structure Interaction Surrogates
Abstract
Surrogate credibility for fluid-structure interac- tion (FSI) requires distinguishing transfer across independent anatomies from interpolation within an already sampled surface. Four de-identified human aortic models from the Vascular Model Repository were reconstructed into separate lumen and nominal 1.5-mm wall domains and analyzed under matched first-cycle two-way FSI. A geometry-only LightGBM prior, selected by leave-one-anatomy-out development on three anatomies, was zero-shot evaluated on a fourth, then probed with a post-zero- shot sparse field-completion case study over six targets. Zero-shot transfer was poor across all targets. At a five-percent anchor level (203 anchors, 3,852 evaluation nodes), prior-plus-adaptation reached an oscillatory shear index (OSI) R2 of 0.603. However, same-anchor controls tuned only on the three development anatomies were stronger for several outcomes: inverse-distance weighting reached R2 = 0.829 (OSI), 0.617 (peak von Mises stress), 0.676 (mean stress); radial basis function interpolation reached 0.917, 0.714, 0.778. Sparse within-anatomy labels thus support field completion, but this four-anatomy cohort gives no evidence the cross-anatomy prior adds value beyond direct interpolation. We frame this as a first computational stage toward a measurement-linked digital twin: the surrogate/update layer is evaluated here, while larger cohorts, converged FSI, measurable patient-side inputs, and physics-informed learning remain future work, not a claim of a complete clinical twin. Our code, data and computation files are available at https://github. com/ali-nourbakhsh2005/Aortic-FSI-Sparse-Field-Completion
cs.AI / 12 / 2609.16487
Skill-based Agentic Evaluation for Real-time Data Science Tasks
Abstract
We present a framework for evaluating data-science agents on live, continuously updated data using executable ground truth and format-agnostic factoid scoring. Consider this example query: "what were last week's audience sizes"---the reference answer changes as the underlying data changes, so static references become outdated and standard LLM-as-a-judge pipelines cannot verify responses against a fixed ground truth. Our central contribution, ground-truth-as-code, encodes each expected answer as an executable reference function that recomputes the answer directly from live data at evaluation time, ensuring the reference remains consistent with the system it describes. We combine this with a factoid-level, format-agnostic judge that decomposes both the agent's response and the computed ground truth into atomic claims and scores precision, recall, and accuracy over them, irrespective of the response format (prose, list, table, HTML, etc.). The approach is applicable to agents whose expected outputs can be expressed as executable data computations. We validate the framework through a human--LLM agreement study on an internally developed machine learning skill deployed in production, using a synthetic database constructed to reproduce production schemas and entity relationships. Relative to a natural-language ground-truth baseline, our method achieves a 29% improvement in the Matthews Correlation Coefficient (MCC)---a class-balanced measure of agreement between expert annotators and LLM-as-a-judge predictions---and a 16% reduction in token consumption per test case, while a self-directed baseline lacking explicit ground truth is anti-correlated with human judgment. Agents that perform multi-source data integration and computation over non-stationary data are routinely deployed in industry; we propose ground-truth-as-code as a practical methodology for their evaluation.
cs.AI / 13 / 2609.16493
From Manual Construction to AI-Driven Scenario Emergence: Rethinking Catastrophe Risk Modeling
Abstract
Traditional catastrophe (CAT) risk models rely on costly manual construction to generate extreme weather scenarios, an approach largely unchanged since the 1990s. As climate extremes intensify, this creates mounting challenges to the entire risk transfer chain. This study proposes the TAISE framework, which repurposes AI weather forecasting models to produce coherent extreme weather sequences at a fraction of traditional costs. Through self-iterative generation, the framework produces continuous global atmospheric fields from which extreme events emerge. A proof-of-concept experiment demonstrates an order-of-magnitude reduction in computational cost compared with conventional methods, while capturing temporal continuity and cross-regional correlations absent in snapshot-based approaches. These findings suggest a pathway toward democratising catastrophe risk quantification and enabling dynamic, comprehensive portfolio assessment for insurers, reinsurers, ILS fund managers and public-sector risk managers.
cs.AI / 14 / 2609.16519
AquiLLM: Evaluating Faithfulness in Open-Weight RAG-LLM Systems for Scientific Research
Abstract
Scientific research increasingly relies on large, heterogeneous data sources, motivating interest in retrieval-augmented generation (RAG) systems that provide natural language access to scientific knowledge and research workflows. Researchers are exploring the viability of these systems as natural language interfaces for document search and for generating analysis code and pipeline components. At the same time, concerns about data privacy and control over research infrastructure have motivated interest in open-weight models and open-source deployments hosted within research institutions. In astronomy, this development follows a long history of computational infrastructure development, from archival databases and SQL-based systems to LLM-assisted research tools. This paper presents a domain-expert evaluation of faithfulness for AquiLLM, an open-weight, offline RAG-LLM platform designed to support scientific research groups in the use and preservation of tacit and formal knowledge. We define faithfulness as the extent to which generated responses remain grounded in retrieved scientific context without unsupported claims or omissions. We report results from an astronomy case study evaluating AquiLLM across retrieval and scientific analysis tasks. AquiLLM performs most reliably on explicit retrieval-oriented questions grounded in the RAG collection, while faithfulness degrades for queries requiring synthesis or ambiguity resolution. These results highlight both the promise and limitations of open-weight RAG-LLM systems for scientific research and demonstrate the importance of domain-expert evaluation beyond standard benchmark leaderboards.
cs.AI / 15 / 2609.16548
QueryFormer: Winning Solution for KDD Cup 2026 Tencent UniRec Challenge
Abstract
Post-click conversion rate (pCVR) prediction requires jointly modeling feature interactions and sequential user behaviors. The KDD Cup 2026 Tencent UniRec Challenge calls for a unified architecture addressing both. We observe that existing unified architectures often generate query tokens---the central information hub---with projection-based multi-layer perceptrons (MLPs), without explicit token-to-query attention for refining the query side. We propose QueryFormer, centered on a stackable unified field--sequence block that bridges non-sequential multi-field features and behavioral sequences, and provide a latency-aware scaling study over view width $H$, model width, depth, data, and compute. The block generates queries through cross-attention and packs sequence queries into shared-parameter attention. QueryFormer secured 1st place in the Industrial Track, achieving an official test area under the ROC curve (AUC) of 0.83254; a modest post-competition scale-up reached 0.832713. Within our grid, $H$-scaling improves validation AUC from 0.84540 to 0.84615 and beats HyFormer at comparable budgets. Ablation identifies query generation as the largest contributor. Packed shared-parameter cross-attention keeps H=8 inference latency to only 1.89x that of H=1, positioning the bridge as an efficient stackable unified block.
cs.AI / 16 / 2609.16564
Query-Aware Source-Risk Triage for Retrieval-Augmented Generation
Abstract
Retrieval-augmented generation (RAG) pipelines may omit a source's material relationship to the query. We study a pre-generation triage layer that treats this relationship as query dependent. The method routes canonical query families for enhanced review and assigns retrieved pages to pass, contextualize, exclude, or review. It combines a four-dimension page score, rank-discounted family aggregation, intent-preserving query mutations, and a family-held-out router. A single-coded pilot of 200 real URLs supplies provisional calibration anchors; a 20,000-row scenario with synthetic domain identifiers supports controlled workload analysis. An oracle page gate defines a risk-coverage target for a future learned classifier. The evaluation shows why page-level frequency cannot substitute for family-level exposure and quantifies how calibration changes scenario activation. Annotation reliability remains unmeasured, and synthetic rankings omit real retrieval dynamics. The result is an auditable triage method and validation plan, not an estimate of deployed review workload, live-Web prevalence, or downstream answer-quality gains.
cs.AI / 17 / 2609.16635
EchoPath: Execution-Level Replayable Memory for GUI Agents
Abstract
Computer-use agents increasingly operate browsers, software, and desktop applications via CLI or API portals, but graphical user interface (GUI) still plays an important role in common industrial production scenarios. GUI agents commonly employ fresh observe-plan-ground-act loops, which is inefficient for enterprise tasks that repeatedly update records, process forms, configure tools, and export reports. We introduce EchoPath, a model-agnostic harness that converts artifact-validated GUI trajectories into standardized, parameter-controlled callable memories, analogous to Model Context Protocol (MCP)-style tool calls rather than unstructured experience records. Each memory stores task-intent keys, application and state preconditions, flexible input parameters, GUI evidence, validation provenance, and lifecycle state, so the host agent invokes a targeted procedure only when it can be deterministically replayed in the current runtime. The core mechanism enabling replay is an image-based target-reaiming algorithm that treats stored coordinates as visual evidence, matches the remembered GUI target against the current screen, and emits corrected operation coordinates before execution. During replay, EchoPath rebinds only declared modifiable inputs and rejects ambiguous or incompatible steps to bounded grounding repair or fresh planning. In experiments with real computer-use tasks, EchoPath reduced median token cost by more than 90% and median execution time by about 60%. These results support a bounded form of enterprise GUI memory: validated execution experience can become a controllable callable asset for recurrent work rather than only context for another reasoning pass.
cs.AI / 18 / 2609.16639
ReDraft, Don't Just Distill: Reference-Driven Revision for Continual VLLM Post-Training
Abstract
Continual post-training of large multimodal models should add new capabilities while preserving those from pre-training, and the two goals pull in opposite directions. SFT gives explicit target supervision that learns a task from near-zero accuracy, but its off-policy targets move the model far enough to cause forgetting; on-policy methods such as RLVR and self-distillation preserve policy proximity yet supply little signal when the policy cannot yet solve the task. We introduce ReDraft (Reference-Driven Revision and Fine-Tuning), which obtains both from the model's own failures: using an expert response only as a reference, it has the model revise its own incorrect rollout, keeps the revision only if a verifier accepts it, and fine-tunes on what survives. Each retained target is therefore explicit, yet still close to the current policy. Across Counting, Clock Reading, and Jigsaw on Qwen2.5-VL-3B/7B, two of them with near zero accuracy, ReDraft gains 56.9 points on the target task against SFT's 52.9 while cutting prior-task loss from 16.6 to 1.5 points (11.3x less forgetting), and improves on OPSD along both axes (19.3 gain, 6.2 loss). Data- and parameter-space analyses match the design: revised targets are more probable under the base model, and the updates they induce stay compact and follow SFT's direction more closely than OPSD's. Repairing the model's own output, rather than replacing it with an expert's, is what lets one objective do both.
cs.AI / 19 / 2609.16667
ANIMASK: What the Model Contributes to Role Play in Simulated Story Worlds
Abstract
When a language model plays a character, the observed behavior reflects both the assigned persona and the default dispositions of the actor model itself. Existing evaluations test persona fidelity or model defaults in isolation, but neither says, at a specific choice with consequences, what the persona changed and what the model's default kept. We introduce ANIMASK, a simulation framework that freezes books and scripts into story worlds whose characters act on their own motivations and replays each story from its freeze point. We hold out the author's continuation as a human reference, verify through in-story interviews that each persona remains present, and at every decision point compare the character's action with what the model produces when the persona is removed. Across 40 stories, 6 actor models, and 3,846 decision points, the replays converge away from their canons in one shared direction, toward flatter, cooler stories that leave their tensions open. The personas stay present and obeyed throughout. On three choices in four the model's default already falls inside what the persona accepts, and where the two diverge the model is the cautious one, holding where the persona would press. The persona guarantees who the character is, and the model sets how far the character will go.
cs.AI / 20 / 2609.16679
AI for Games in the Foundation Model Era
Abstract
Foundation models, alongside advances in learned game-world models, are reshaping AI across the game lifecycle. Beyond playing games, recent systems model players and game dynamics, support design and development, adapt player-facing experiences at runtime, and evaluate resulting artifacts. Yet these directions have evolved largely separately, obscuring which capabilities transfer across settings and which remain tied to particular games, engines, interfaces, or player populations. We organize the literature into six roles according to the immediate use of AI output: playing and acting; modeling players and games; designing games; building and maintaining games; generating and adapting at runtime; and testing and evaluating games. For each role, we examine what structure is supplied by the game or workflow, what AI learns or produces, which capabilities and artifacts transfer across settings and roles, and what evidence supports the claims. We identify cross-role connections: trajectories train world models, learned environments provide experience for agents, design specifications drive executable implementations, and play or testing feedback guides revision. However, control schemes, rules, engine interfaces, state representations, and player contexts often remain setting-specific, so downstream claims require validation in the target setting. Evaluation is most standardized for bounded game playing and selected learned environments, while persistent state in learned worlds, repeated software revision, validated player modeling, sustained runtime adaptation, and representative automated testing remain less established. The central challenge is to reuse or transfer outputs and capabilities across roles while re-establishing evidence for effectiveness in the game-specific contexts where they are used.
cs.AI / 21 / 2609.16730
LSREP: A Longitudinal State-Replay Protocol for Evaluating Conversational Memory, with ICE v2 as an Audited Local-First Architecture
Abstract
Conversational memory changes during use, so endpoint question answering alone cannot establish how a persistent state accumulates, ages, or incorporates revisions. We introduce LSREP, a Longitudinal State-Replay Evaluation Protocol combining ordered replay, explicit lifecycle schedules, repeated probes, evolving reference answers, and mechanism-fidelity checks. Its architectural case study is ICE v2, a local-first memory middleware with typed stores, retrieval fusion, and dynamic context budgets. The private, single-user instantiation contains 1,985 turns, 219 distinct probes, and 1,211 probe-checkpoint observations across 52 checkpoints. On three ordinary-density datasets, ICE v2 has a near-zero mean quality difference from vector-RAG while selecting 32% fewer fragments but using 6.6% more estimated prompt tokens. A fourth, dense dataset exposes catastrophic failures of the unbudgeted baseline. The fidelity audit limits attribution: procedural retrieval is defective, several mechanisms are unexercised, and graph utility is not established. In a complementary matched public diagnostic, ICE v2 loses decisively to pure vector-RAG on LongMemEval: 50.8% versus 72.8% in the evidence-only oracle and 43.0% versus 69.5% in full-S. Paired differences are -22.0 points (95% CI [-26.6, -17.4]) and -26.5 ([-31.3, -21.8]). Conservative abstention accompanies severe multi-session and temporal failures. ICE uses less context in this diagnostic, establishing a quality-cost trade-off rather than superior efficiency. Together, replay, fidelity auditing, and public endpoint testing expose distinct failure modes that neither architectural descriptions nor aggregate scores identify alone.
cs.AI / 22 / 2609.16752
Beyond Episodic AI: Cognitive Field Networks for Biologically Inspired Persistent Cognition
Abstract
Cognitive Field Theory (CFT) proposes that cognition arises from memory-dressed collective dynamics that generate a persistent macroscopic cognitive field. Here we develop a Cognitive Field Network (CFN), a recurrent Transformer in which the organized hidden field re-enters subsequent inference through \[ Φ_{n+1}=F_θ(X_{n+1},Φ_n). \] Rather than prescribing an explicit memory operation, the CFN allows new information to act on an already history-dependent collective state. We find that learning organizes persistent, content-dependent recurrent dynamics whose timescale increases systematically with the trained recurrent horizon. Semantic continuation propagates the recurrent state far beyond this horizon without replay of the target answer. Without content-specific support, the field exhibits finite passive relaxation, whereas periodic re-exposure to relevant input repeatedly renews the surviving state and drives it toward an approximately stationary nonzero regime. Unrelated-input and recurrence-off controls do not reproduce this behavior, while near-paraphrased re-exposure produces weaker renewal, demonstrating representation-sensitive persistence. These results distinguish three dynamical processes: collective memory dressing forms and sustains a history-dependent cognitive field, structured input reorganizes this field, and cross-cycle re-entry makes the resulting state causally available to subsequent inference. The CFN therefore provides a controlled computational platform for studying persistent, history-dependent cognitive dynamics without a separately prescribed memory system.
cs.AI / 23 / 2609.16768
Coverage-Aware Virtual IMU Augmentation for Low-Resource Human Activity Recognition
Abstract
IMU-based human activity recognition (HAR) enables continuous, privacy-friendly monitoring of daily activities using wearable sensors. However, building reliable HAR models that generalize across diverse users and real-world conditions requires large amounts of labeled IMU data, which are expensive and difficult to collect. Existing approaches mainly rely on augmentation or synthesis to expand available data, but indiscriminately adding virtual samples may provide little new coverage and introduce unreliable supervision. To overcome these challenges, we propose a novel coverage-aware virtual IMU augmentation framework that decides where to supplement real data, how to generate and select virtual candidates, and how strongly to weight them during training. Specifically, we select diversity and scarcity anchors in a learned sensor embedding space, convert anchor dynamics into prompts, and generate virtual IMU candidates for each anchor. We then rank candidates by a selection cost combining anchor proximity and label consistency, and incorporate the selected candidates into HAR training with reliability-based weights. Experiments on public HAR benchmarks show that our method consistently improves recognition performance over competitive baselines, and ablation studies confirm the effectiveness of the proposed framework design.
cs.AI / 24 / 2609.16822
Execution Flexibility in Automated Planning: A Comparative Evaluation of Deordering and Reordering Strategies
Abstract
This study covers foundational concepts for enhancing plan-execution flexibility, including partial-order planning, the producer-consumer-threat formalism, and a range of deordering and reordering strategies. Creating a partial-order plan from a sequential one by removing unnecessary ordering constraints is a practical way to improve execution flexibility, and several methods have been proposed for this task. This study analyzes their capabilities across ordering, action handling, parameter handling, plan structure, concurrency, and complexity, and evaluates them against each other on a shared benchmark. The central finding is that block deordering-based approaches, which restructure causal dependencies through block-level grouping and subplan substitution, substantially outperform MaxSAT-based approaches despite the latter's theoretical guarantees of minimum reordering. The reason is structural: minimum reordering optimizes within the causal structure already present in the plan, whereas block deordering-based methods change that structure, exposing orderings that would otherwise appear necessary. A further distinction is practical: block deordering-based methods are anytime algorithms that always return a valid result, while MaxSAT-based methods fail entirely on a substantial portion of plans and offer no partial solution when they do. Block substitution further extends the parallel execution by formalizing non-concurrency constraints, though its impact is limited to domains with resource-based interactions. On efficiency, block deordering-based approaches achieve the highest flex gain per unit of computation time, while MaxSAT-based encodings incur large computational overhead.
cs.AI / 25 / 2609.16884
Bridging Learned Visual Perception and Symbolic Belief-Space Planning
Abstract
In partially observable settings, agents must act without full knowledge of the world state and rely on uncertain state-estimation pipelines. Obtaining grounded and verifiable symbolic plans under such uncertainty remains a key challenge. Recent work has integrated Vision-Language Models (VLMs) to bridge perception and symbolic reasoning, following two main paradigms. The first, VLM-as-planner, maps images directly to action sequences, and the second, VLM-as-grounder, grounds observations into symbolic predicates used as the initial state by off-the-shelf planners. Both approaches ignore uncertainty in the planning process, compromising robustness. We introduce a third paradigm, VLM-as-probabilistic-grounder, a novel approach that captures the uncertainty of VLM predicate groundings as a probability distribution over symbolic states. This enables planning in belief space and producing robust plans under uncertainty. Experiments in simulated household robot settings show improved robustness and task success over deterministic grounding, underscoring how our approach leverages foundation models for reliable planning under uncertainty.
cs.AI / 26 / 2609.16887
QART: A Quantum-Classical Hybrid Architecture for Long-Horizon Reasoning -- Exploring a Conditional Path toward Quantum Scaling
Abstract
Long-horizon reasoning is vulnerable to early errors that compromise later decisions. We present QART, the Quantum-Augmented Reasoning Transformer, a quantum--classical hybrid architecture combining a backbone language model with quantum encoding, CIM-based QUBO optimization, and quantum decoding. Semantic information can come from hidden representations or model-generated text; detailed encoding and optimization procedures remain proprietary. Under explicit assumptions, we establish a conditional asymptotic reliability separation from single-trajectory autoregressive LLMs. For a common task family with aligned optimality and acceptance criteria, autoregressive acceptance probability tends to zero when cumulative conditional risk of irreversible errors diverges. QART's task-optimal-path recovery probability remains bounded away from zero if conditional probabilities for optimal-path coverage and semantic fidelity, spectral certification, dynamical reachability, and faithful readout remain uniformly positive under a specified resource schedule. The architecture alone does not imply these bounds. Paired measurements on six long-horizon benchmarks using DeepSeek V4 Flash, GLM-5.3, and GPT-5.5 xhigh in a Codex agent environment favor QART in 14 of 15 backbone--benchmark pairs. Relative gains reach 84.0% on SciCode, 47.6% on $τ^3$-Bench, and 44.4% on Terminal-Bench 4.0; the DeepSeek V4 Flash configuration regresses by 7.8% on DeepSWE. These results do not directly validate the asymptotic separation. Potential quantum scaling laws are formulated as conditional hypotheses. A quantum-advantage interpretation requires a demonstrated CIM quantum advantage over strong classical solvers and its transfer to end-to-end reasoning after all system overheads.
cs.AI / 27 / 2609.16948
AntennaFlow: A Generative Flow Model for Offset Correction in Phaseless Antenna Testing
Abstract
Near-field to far-field transformation is central to large-aperture antenna testing, yet two coupled challenges remain: costly phase acquisition at millimeter-wave bands and violations of the centering assumption under offset mounting. Existing methods address these issues separately, requiring either dense full-field data or offset vectors. We tackle both jointly by exploiting a key observation: amplitude fields under different offsets are coordinate-transformed views of the same near field. The challenge is to recover the center-aligned field from offset amplitudes without a phase or offset vector. We propose AntennaFlow, a three-stage framework: a contrastively learned encoder that maps offset views to an offset-invariant embedding, a deterministic flow-matching transport that maps offset amplitudes to center-aligned ones, and the Simplified Extrapolation Technique, whose Green-function Taylor expansion is valid only for centered fields. Experiments show that AntennaFlow enables fast, phaseless, offset-vector-free NF--FF reconstruction from sparse amplitude-only measurements, consistently outperforming existing baselines while preserving physical consistency.
cs.AI / 28 / 2609.16962
Affect-Prototype Guided Fusion for Open-Vocabulary Incomplete Multi-modal Emotion Recognition
Abstract
Open-vocabulary multimodal emotion recognition (OV-MER) aims to generate open natural-language emotion labels from multimodal affective cues. In real-world scenarios, however, complete and synchronized modal data are difficult to obtain due to limitations of acquisition devices and user privacy constraints. Existing OV-MER methods are largely designed for full-modal inputs, and fail to perform effective feature fusion under modal missing conditions. Meanwhile, current fusion approaches designed for incomplete modalities mainly focus on fixed-label recognition context, and cannot satisfy the demand for fuse emotional cues guided with arbitrary emotion semantics in OV-MER context. To tackle these challenges, this paper proposes an Affect-Prototype-Conditioned Fusion (APCF) framework for incomplete open-vocabulary emotion recognition. As a candidate-free generative framework, APCF extends modal contribution learning to scenarios guided by arbitrary emotional semantics. Specifically, we construct an affect-prototype library to explicitly model multimodal contribution characteristics corresponding to diverse emotions, which provides dynamic constraints for modal fusion under different emotional semantic perspectives. Conditional retrieval and feature aggregation are conducted based on available modal features. The refined fused affective representations are then fed into an LLM decoder to produce open-vocabulary emotion labels. Experiments on the OV-MERD+ and MER-FG datasets demonstrate that APCF substantially outperforms state-of-the-art baselines.
cs.AI / 29 / 2609.17010
ThinkFlow: Self-Evolving Probabilistic Latent Memory for Lifelong Conversational Agents
Abstract
Lifelong conversational agents rely on memory systems to maintain deep, context-aware interactions with users. However, existing explicit textual memory pipelines suffer from a severe information bottleneck, often losing subtle behavioral patterns and emotional shifts. Furthermore, being typically static post-deployment, they cannot autonomously adapt to personal habits and preferences without manual feedback. Cognitive science, however, suggests that humans maintain mental models purely in a latent space and continuously refine them through predictive coding. Inspired by this, we propose \textbf{ThinkFlow}, a novel end-to-end latent memory framework for lifelong conversational agents. ThinkFlow bypasses the text bottleneck by dynamically compressing conversational flows into probabilistic latent memory skills, autonomously consolidating complex user states into disentangled, continuous vectors without semantic interference. To break this barrier, we introduce a test-time evolution paradigm. By coupling teacher-guided latent alignment to bootstrap the initial state with a self-supervised next-user-utterance prediction task for continuous refinement, the framework successfully overcomes cold-start challenges and achieves label-free lifelong personalization. Extensive experiments on long-term conversation benchmarks demonstrate that ThinkFlow significantly outperforms prevailing memory systems, providing highly personalized and contextually accurate responses over extended multi-session interactions.
cs.AI / 30 / 2609.17012
ORDER: Task-Conditioned Routing for Retrieval-Augmented Generation
Abstract
Retrieval-Augmented Generation (RAG) pipelines typically rely on a fixed indexing and retrieval configuration determined at preprocessing time. This one-size-fits-all design is ill-suited to domain-expert settings, where heterogeneous queries require different chunking granularities, metadata constraints, and source-selection strategies. As a result, configurations that are effective for one family of queries often perform poorly for others. In this paper, we introduce ORDER (Optimal Routing for Dynamic Evidence Retrieval), a query-conditioned RAG framework that jointly adapts indexing and retrieval to the incoming query. Our approach first discovers semantic clusters over a given set of questions associated to a corpus and learns, for each cluster, a chunking strategy together with a suited metadata filtering and reranking configuration. At inference time, queries are routed to the appropriate pre-built index through nearest-centroid assignment. To further improve retrieval, we propose a supervised query router (QRe) that predicts which collections are most likely to contain relevant evidence, coupled with a Uniform Multi-source Sampler (UMS) that allocates the retrieval budget evenly across the selected sources. We evaluate our framework on large-scale, heterogeneous historical archives and show that conditioning both indexing and retrieval on the query consistently outperforms both naive baselines and strong state-of-the-art RAG systems in complex expert-domain environments.
cs.AI / 31 / 2609.17064
Neuro-Symbolic Hierarchical Intention Anticipation in Human Behavior
Abstract
Assistive autonomous systems must anticipate human goals before an observed behavior is complete. This article formulates anticipation as goal inference from a partially observed multimodal episode together with structured prediction of the remaining behavior, rather than exact motor forecasting. A compact Hierarchical Planning Decoder (HPD) is attached to a frozen neuro-symbolic recognition encoder and predicts, at four ontological levels, the next actions, the remaining activities and low-level intentions, and the episode high-level intention(HLI). The decoder is trained with soft neuro-symbolic regularization combining transition-coherence and hierarchical continuity losses, and is decoded with hard reachability masks that enforce ontological validity at inference. On a compositional four-level benchmark of 15,002 multimodal episodes built over NTU RGB+D 120 features, three headline properties are observed together. The advantage over the strongest sequential baseline grows with the anticipation horizon, from +1.7 points at step 1 to +7.3 points at step 3 (top-5). Under compositional generalization, where one parent association per multi-parent low level intention is held out, this advantage widens to +4.9 points at step 1. At the episode level, 96.8% of anticipated trajectories satisfy the joint logic constraints, above the 88.1% strongest-baseline value and the 73.9% ground-truth floor; soft logic terms alone account for a 59.8 to 71.1% relative reduction of HLI-reachability violations, and the hard masks then eliminate them entirely. Neural generation supplies predictive ranking, symbolic constraints supply onto logical validity, and their combination yields coherent hierarchical anticipation while exposing remaining challenges in compositional goal generalization and unordered set prediction.
cs.AI / 32 / 2609.17076
Sample-Conditioned Representation Selection for Audio Few-Shot Learning
Abstract
Few-shot audio classifiers may rely on foreground-background co-occurrences and fail when those correlations shift. On SpurAudio, the resulting representation shift is concentrated and class dependent: for ResNet12, the top 10 percent of channels explain 82.80 percent of the null-corrected shift contribution. We propose SAMPLESELECT, which predicts a fixed-budget feature mask independently for each input while keeping the encoder and source classifier frozen. Training uses differentiable Gumbel Top-k selection with foreground classification and cross-background contrastive losses; inference uses deterministic Top-k masks and support-only linear adaptation. Across ResNet12 and Conv64 in 5-way 1-shot and 5-shot evaluation, SAMPLESELECT gives the best OOD accuracy among the compared methods and improves the matched full-representation control by 4.90-8.38 percentage points. Ablations and representation analyses further support the learned selection mechanism. Code is available at https://github.com/Cross-Innovation-Lab/SAMPLESELECT/
cs.AI / 33 / 2609.17091
Scaling-Score Conformal Prediction for Multi-Target Regression
Abstract
Multi-target regression requires a model to simultaneously predict several related outputs. Conformal prediction provides distribution-free, finite-sample marginal coverage guarantees, but extending these to joint multi-dimensional regions in a model-agnostic, sample-efficient manner remains challenging: max-aggregation ignores scale differences, copula-based methods are only asymptotically valid, rectangular methods typically split the calibration set, and quantile or density-based methods require training a specialised model beyond a plain point predictor. We propose the scaling-score conformal method, which is model-agnostic (requires only component-wise absolute residuals), uses a single calibration set, and yields four nested output types: an outer rectangle (SCO) with valid joint coverage, the exact set R $α$ , a staircase (SC 2 ) over approximation of R $α$ , and an inner rectangle (SCI). A single hyperparameter $γ$ $\in$ (0, 1) controls the base-rectangle quantile level independently of $α$. We prove downward-closedness and a rectangular sandwich bound and derive a closed-form outer rectangle. Experiments on 29 realworld datasets confirm valid joint coverage; SC 2 with $γ$ = 1-$α$ consistently achieves competitive volume relative to baselines, with the advantage growing with output dimension d.
cs.AI / 34 / 2609.17100
Semi-Supervised Learning-Based Genetic Biomarkers Dataset for Multiple-Stage Hepatocellular Carcinoma Prediction
Abstract
Liver cancer is a complex disease responsible for a high number of deaths across the globe each year, making automated solutions for liver cancer classification urgent. The most common form of liver cancer is hepatocellular carcinoma (HCC), accounting for over 90% of liver cancer cases. There is a distinct lack of publicly available HCC datasets utilizing genomic data, which is necessary for training artificial intelligence (AI) models for automated HCC classification. This study proposes constructing a multi-stage HCC dataset using XGBoost and Semi-Supervised learning on three separate datasets of genomic biomarkers, utilizing their existing labels in the Semi-Supervised learning process to label the proposed dataset. The proposed dataset consists of 770 patient samples in total, categorized into five classes that represent normal tissue alongside different stages of HCC. Each sample in the dataset consists of 11,150 different gene expression levels. The XGBoost model demonstrated a final classification accuracy of 96.5% during the Semi-Supervised learning process.
cs.AI / 35 / 2609.17107
Symbolic Separation: Grounding Deep Agents in Knowledge Graphs for Trustworthy Operational Data Analytics
Abstract
Generative AI promises natural language access to the massive numerical telemetry of data centers and Industry 4.0 installations, yet text-to-query and tool-using agents stay unreliable: even frontier models answer little more than half of real-world database questions, and far fewer of the multi-step, operational ones, because the LLM must compose how heterogeneous sources relate and hallucinates the relations, not just the fields. We propose symbolic separation: a deep agent reasons freely but may act on data only through an ontology-constrained Virtual Knowledge Graph with deterministic pre-execution validation. Unlike a tool API's interface contract, this domain-semantic contract turns a complex question into one validated graph traversal instead of LLM-inferred joins. Instantiated as the Neurosymbolic Deep Analyst and evaluated on 49.9 TB of superconputer telemetry against a rigid workflow and a non-symbolic ablation, it raises end-to-end task success from 43% to 86%, prevents silent data-integrity errors that no syntactic check catches, and cuts token cost by 2.4x, letting a smaller on-premise model outperform a larger one.
cs.AI / 36 / 2609.17109
Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs
Abstract
A common small-model deployment runs one shared backbone with several LoRA specialists that answer over the same context. Serving them naively re-prefills that shared context once per specialist. We study a narrow, practical question: for already-trained standard LoRA adapters -- not adapters retrained for cache compatibility -- how much task quality is preserved if the backbone's prefill KV cache is computed once and reused across specialists, and what does that buy in serving cost? On a Qwen3-1.7B backbone with two adapters (extractive QA on HotpotQA, arithmetic reasoning on GSM8K), we sweep the boundary at which the specialist takes over from the reused base cache and measure paired quality differences and serving cost. Full-prefix reuse had the lowest prefill cost and a small quality difference on held-out GSM8K (Delta = -4.6 EM at a 160-token budget; -3.0 at 320 tokens; -0.8 under a second training seed -- all favoring native, only the first excluding zero, and the magnitude not consistent). Partial recomputation provided no demonstrated advantage. Neither quality equivalence nor a general boundary-selection rule is established. We also report a closed-form ridge KV translator that did not beat direct reuse, and specialist-dependence contrasts whose intervals all include zero. The measured serving benefit is warm-cache time-to-first-token, which grows with context (~16x at 8K); two-branch peak memory was only 12% lower and, on inspection, the prefix was never physically shared across branches -- this implementation reuses KV values but copies their storage, so shared-cache memory savings are not achieved.
cs.AI / 37 / 2609.17180
MOCC-R1: Reinforcing Reasoning-Response Consistency for Multimodal Counselor Response Generation
Abstract
Multimodal counselor response generation (MCRG) aims to generate an appropriate counselor response from multimodal dialogue histories. Progress is limited by two gaps: first, existing datasets rarely capture sustained, human-recorded counseling interactions conducted by qualified counselors; Second, existing methods do not explicitly optimize consistency between counseling reasoning and the generated response, potentially undermining the reliability of MCRG systems. Thus, we introduce MOCC, a multimodal counseling conversation corpus containing over 200 hours of interactions involving 154 credential-verified counselors. Based on MOCC, we propose MOCC-R1, a two-stage framework for optimizing reasoning-response consistency. Cold-start supervised fine-tuning trains the model to generate a structured trajectory consisting of client-state understanding, a response intent that links a counseling principle to a planned action, and the final response. Reinforcement learning (RL) then rewards grounded plan coherence and plan execution, encouraging the inferred state and plan to be supported by the dialogue context and the response to realize that plan. Experiments demonstrate the effectiveness of the proposed MOCC-R1.
cs.AI / 38 / 2609.17325
Intrinsic Motivation in Reinforcement Learning: A Research Agenda for Adaptive Self-Organisation
Abstract
Biological cells can be viewed as individual, interacting agents whose collective dynamics give rise to adaptive behaviour at multiple levels of organisation, from individual cells through tissues to whole multicellular organisms. In this perspective and tutorial article we discuss whether intrinsic rewards in artificial neural systems can support adaptation, functional specialisation and higher-level self-organisation without a shared external objective. We review empowerment, curiosity, learning progress, information gain, unsupervised skill discovery, mutual information estimation and the use of world models for intrinsic reward computation. Particular attention is given to failure modes showing when such objectives do not produce sustained exploration or increasingly complex behaviour. We argue that more capable systems may require complementary objectives, communication, memory, learning at multiple temporal scales and environmental constraints. Based on this perspective, we outline three experimental directions. These include a resource-constrained environment in which otherwise stable behavioural attractors become unsustainable, allowing us to test whether environmental constraints can mitigate characteristic failure modes of intrinsic objectives. The network of recurrent agents with per-agent intrinsic rewards, and a hierarchical world-model agent in which exploratory motor competence develops before goal-directed behaviour. These experiments are intended to test whether intrinsic learning can lead to adaptive organisation at progressively higher levels.
cs.AI / 39 / 2609.17326
From Transient Prompts to Persistent Control: Scientific Poster Generation via Recursive Semantic-Geometric Contracts
Abstract
Scientific poster generation distills a multimodal paper into a single-page visual artifact, forcing strict trade-offs between informational coverage and readability under a fixed spatial budget. Existing methods pass plans as transient prompts and validate individual stages in isolation. This strategy causes requirements to drift across content and layout modules, and previous checks to be silently invalidated. We introduce PosterVisor, a control framework that shifts poster generation from transient prompts to persistent control. An Orchestrator grounds rubrics in the paper and visual assets, compiling them into a Semantic-Geometric Contract (SGC) that binds claims and sources to required visuals, budgets, and spatial commitments. Only fully instantiated records become executable assertions; other usable requirements remain soft guidance. Recursive Contract Enforcement (RCE) dynamically triggers checks across stages as evidence emerges. Crucially, during repairs, RCE rechecks affected checkpoint states, preventing repair-induced regressions from propagating silently. We instantiate PosterVisor in HTML/CSS and editable PPTX generators. On the 100-paper Paper2Poster benchmark, PosterVisor-PPT improves observed mean poster-grounded QA accuracy over PosterGen (64.47% vs. 58.53%) and is preferred by human judges in 72.5% of non-tied pairwise comparisons (95% CI, 61.6-83.4%). A secondary 30-paper study also yields higher VLM Overall and PaperQuiz means. These results support rubric-compiled contracts and stage-conditioned enforcement for controllable poster synthesis.
cs.AI / 40 / 2609.17391
FlashVector: Agent for Hierarchical Model Serving Stack Optimization
Abstract
Model serving is one of the largest cost drivers in production recommender systems. Maximizing its throughput requires navigating a deeply layered hierarchy: GPU kernels, the ML framework computation graph, the model server, and on-demand feature processing -- each demanding specialized domain expertise. Such cross-layer expertise is inherently difficult to acquire, and does not scale with a workload that continuously grows and evolves, leaving significant cost efficiency gains unrealized. While recent AI agents have demonstrated human expert level efficiency in standalone GPU kernel optimization, automated tuning and optimization for the rest of the serving stack remain largely unexplored. We present FlashVector, an agentic system that optimizes performance across all layers of the model serving stack. The key contribution is an extensible framework to generalize the single kernel optimization agent paradigm to heterogeneous technical stacks, and to deliver performance improvements holistically. After deployment in Unity's Vector advertising platform, FlashVector achieved up to 2x throughput increase and up to 1.98x latency speedup on model server, and up to 1.6x throughput increase on feature store. These optimizations were discovered not only at the GPU kernel and computation graph levels, but also across the other components of the model serving stack, such as the model server (NVIDIA Triton's C++ codebase) and the on-demand feature transformation service (Python codebase), demonstrating the extensibility of the framework to more complex system architectures.
cs.AI / 41 / 2609.17475
JustFit: 200K-Token LLM Serving on a 24 GiB Laptop with Just-in-Time State Management
Abstract
Capable open-weight models make local coding and reasoning attractive, but their context and execution state strain laptop memory. We present JustFit, an MLX-based inference runtime that combines KVExec for compressed KV execution, PhaseSwap for component residency, and StateTrans for state-preserving serving transitions. These mechanisms fuse reconstruction and coordinate just-in-time materialization and release, independently of model-weight quantization. In full-execution capacity tests on a 24 GiB M4 Pro MacBook running Qwen3.8-27B MXFP4, three independent runs complete 196,608 input and 16,384 output tokens, increasing completed single-request context from the mlx-vlm baseline's 30,720 positions to 212,992 (6.93x); a separate two-request run retains 229,376 positions in aggregate. In separate performance tests, a 32K-input, 64-output probe reaches 19.11 tokens/s, and a repeated 32K+6K workload has a median peak process footprint of 16,374 MiB. The integrated runtime answers 29 of 30 AIME 2026 problems correctly, showing how compact state and lifetime-aware execution expand local serving capacity while supporting extended generated reasoning.
cs.AI / 42 / 2609.17488
LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence
Abstract
We introduce LimiX-2, a new model in the LimiX family, developed through model and data scaling guided by our previously established scaling laws. LimiX-2 adopts the Contextual Mechanism Networks (CMNs) paradigm and is pretrained with Context-Conditional Masked Modeling (CCMM). CMNs shifts the organizing principle of in-context learning from target-centric prediction to mechanism-oriented joint modeling. Rather than centering the network on the $p(y \mid x, D_{\mathrm{context}})$ objective of conventional tabular PFNs, it is designed around learning $p(x, y \mid D_{\mathrm{context}})$, a context-dependent representation of the joint structure underlying data generation. Pretraining uses synthetic datasets generated by structural causal models (SCMs) spanning diverse graph structures, functional mechanisms, and observation processes. Evaluations on TabArena, TALENT, and BCCO show that LimiX-2 outperforms current dataset-specific models and tabular foundation models. Beyond predictive performance, the CMN paradigm also promotes causal awareness in LimiX-2: its feature attention encodes direct causal relationships, enabling accurate causal skeleton recovery.
cs.AI / 43 / 2609.17496
Verifiable Social Reasoning for LLM Assistants
Abstract
LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it requires setups where the assistant learns about social situations from subjective user narratives, and (ii) social properties, such as others' intentions, typically lack verifiable ground truth. To address these challenges, we introduce Fuse, a multi-agent simulation framework for studying user-mediated social reasoning. In Fuse, a target agent with a hidden motive interacts with other agents including one representing the user, who then consults the evaluated assistant to infer the target's motive, providing verifiable ground truth by construction. Simulation faithfulness is validated through a human study with 24k annotations. We apply Fuse to 12 LLMs and demonstrate its analytical utility by systematically isolating key factors, showing that (i) user mediation compounds the inherent difficulty of social reasoning; (ii) LLMs exhibit systematic sensitivity to biased user framing; (iii) models can require more details than humans need to reach a correct prediction; and (iv) longer conversations do not always improve performance despite providing opportunities for clarifying questions. We open-source Fuse and a dataset with 21k examples.
cs.AI / 44 / 2609.17523
ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents
Abstract
We introduce and release ScienceBuddy, an interactive scientific research workspace that brings continually improving scientific agents into researchers' everyday workflows. ScienceBuddy supports researchers in carrying out scientific tasks while transforming their requests, feedback, and execution evidence into tasks and evaluation rubrics for continual learning. At its core is recursive-in-recursive self-improvement, a paradigm that couples harness evolution with model reinforcement learning: the inner recursion improves the harness with the model fixed, while the outer recursion trains the model under the improved harness. Harness evolution shapes training experience, and model learning creates new opportunities for harness adaptation. We present case studies of researcher interaction, harness refinement, and model learning, with the benchmark cases spanning four scientific task families. By releasing ScienceBuddy as a research product, we make this paradigm available to the scientific community and take a step toward discovery intelligence: scientific AI that advances through sustained collaboration with researchers and evolves alongside the research it supports. Website: http://science-buddy.io
cs.AI / 45 / 2609.16233
SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes
Abstract
Vision-language models excel at 2D image understanding but remain limited in 3D spatial reasoning. Progress is hindered by limitations in current benchmarks. First, 3D datasets often rely on point clouds that capture geometry but discard rich visual features like texture, text, and materials. Second, annotations treat objects in isolation while ignoring real-world hierarchical organization (scenes, rooms, functional areas, object groups). Third, evaluation tasks focus narrowly on basic recognition rather than multi-step spatial reasoning. In this context, we introduce SceneBench, a benchmark of 966 photorealistic 3D scenes reconstructed with Gaussian Splatting and densely annotated with hierarchical semantics spanning scenes, rooms, functional areas, object groups, and individual objects. These annotations are produced through a human-in-the-loop pipeline combining vision-language models with roughly 1,500 human-hours of iterative refinement and verification, producing over 183K annotated nodes with textual descriptions and 3D bounding boxes. Building on this representation, we define three evaluation tasks: Existence-Based Questions probing object attributes, Spatial Intelligence Questions covering counting, size comparison, distance, and directional relations, and Grounded Question-Reasoning-Answer (QRA) triplets requiring multi-step reasoning across semantic levels. Experiments with state-of-the-art vision-language models show that while models perform well on basic recognition tasks (e.g., up to 85% accuracy for detection), performance drops substantially on hierarchical and compositional reasoning (e.g., down to 60% for counting), revealing limitations not captured by existing benchmarks. SceneBench provides a realistic testbed for developing and evaluating models capable of fine-grained spatial reasoning in photorealistic 3D environments.
cs.AI / 46 / 2609.16284
ProtoLIP: From Sentence-Level to Object-Level Evidence Disentanglement
Abstract
Query-conditioned vision--language models enable fine-grained interpretation by revealing how visual evidence changes with textual queries. However, evidence conditioned on complete descriptions does not necessarily resolve into object-specific evidence, nor does an exposed evidence map necessarily identify the evidence that constitutes the model's prediction. Across multiple VLM architectures and independent benchmarks, we find that object-level queries often retain evidence from co-occurring objects and shared context. In this paper, we introduce \textbf{ProtoLIP}, a lightweight prototype-mediated evidence layer that organizes reusable visual prototypes into text-derived semantic families and uses query-dependent family routing to constrain which prototypes may provide evidence. Without spatial annotations or backbone retraining, ProtoLIP improves evidence localization and separation across query granularities, with localization gains transferring to independently pretrained VLMs with well-aligned patch--text representations. Despite using only text-derived weak supervision, ProtoLIP remains competitive with a spatially supervised grounding model while maintaining strong matching and competitive image--text retrieval. Crucially, ProtoLIP constructs its matching score directly from localized prototype evidence, enabling the score to be exactly decomposed into semantic-family and prototype contributions.
cs.AI / 47 / 2609.16565
Vision And Text Transformer For Predicting Answerability On Visual Question Answering
Abstract
Answerability on Visual Question Answering is a novel and attractive task to predict answerable scores between images and questions in multi-modal data. Existing works often utilize a binary mapping from visual question answering systems into Answerability. It does not reflect the essence of this problem. Together with our consideration of Answerability in a regression task, we propose VT-Transformer, which exploits visual and textual features through Transformer architecture. Experimental results on VizWiz 2020 dataset show the effectiveness and robustness of VT-Transformer for Answerability on Visual Question Answering when comparing with competitive baselines.
cs.AI / 48 / 2609.16597
A Vision-Language Foundation Model for Precise and Comprehensive Brain Tumor Diagnosis from Preoperative Multimodal Data
Abstract
Background Non-invasive presurgical diagnosis of brain tumor types from Magnetic Resonance Imaging (MRI) is essential but challenging due to overlapping imaging features across tumor types, inter-observer variability, and the extensive training required for expertise. We aimed to develop an MRI-based Artificial Intelligence (AI) model for automatic and reliable brain tumor classification with diagnostic uncertainty quantification and radiology reports generation. Methods We developed BrainVLM to classify all 12 World Health Organization (WHO) 2021 brain tumor types. BrainVLM integrates an uncertainty quantification strategy to indicate prediction reliability and a module for generating radiology reports to elucidate the clinical rationale. BrainVLM was trained on multi-modal data (MRI scans, demographics, and radiology reports) from 40,043 individuals. It was validated on 5,211 patients with pathologically confirmed brain tumors, including 3,877 held-out patients from the primary hospital and 1,334 patients from 11 independent hospitals. We further conducted two proof-of-concept studies to validate its clinical utility in AI-clinician workflows: 1) a blinded multi-reader study where 12 neuroradiologists across varying experience levels interpreted 248 retrospective cases with or without AI assistance, and 2) a real-world prospective study in which 1,009 patients were independently and blindly assessed by BrainVLM and radiologists before surgery. Additionally, we demonstrated BrainVLM's utility in preoperative molecular subgroup prediction for adult-type diffuse gliomas, using a multi-center cohort of 632 patients.
cs.AI / 49 / 2609.16832
What Breaks Local Watermarks? A Robustness Benchmark for Local Invisible Image Watermarking
Abstract
Local image watermarking embeds an invisible signal into selected image regions rather than spreading it across the entire image, enabling payload recovery from specific objects or regions without perceptibly altering the image. Existing studies evaluate the robustness of payload recovery and localization under image transformations, but they often focus on their own proposed method, resulting in narrow evaluations with inconsistent choices of transformations, datasets, and metrics. These inconsistencies across studies limit direct comparisons across methods and muddle the overall picture of local watermark robustness. To address this gap, we present the first systematic robustness benchmark for local watermarks across 55 image transformations, including (i) signal distortions, (ii) changes in image coordinate alignment, (iii) indirect local edits, and (iv) direct watermark edits. The benchmark evaluates MaskWM, WAM, OmniGuard, TrustMark, and PixelSeal, all methods that either provide native localization or require minimal adaptation to support it. Our results show that all evaluated methods are vulnerable to some transformation, with MaskWM standing out as offering the strongest payload recovery and localization, although it has the lowest image quality in the clean setting. Synchronization further improves MaskWM's payload recovery under several geometric transformations, albeit at an additional cost to image quality. A key finding is that local watermark robustness depends strongly on the nature of the transformation: signal distortions are often tolerated by the strongest methods, while geometric misalignment and generative local edits, such as inpainting and outpainting, can completely impair payload recovery. We observe that payload recovery and localization are related but not interchangeable, and both strongly depend on the transformation's impact on the watermark region.
cs.AI / 50 / 2609.16841
StackTok: Accelerating VLMs Inference with Budget-Adaptive Visual Token Selection
Abstract
Increasing image resolution produces ever-longer visual-token sequences in vision-language models (VLMs), substantially raising their inference cost. To reduce this overhead without retraining, existing methods select compact token subsets that prioritize query relevance, visual coverage, or a fixed trade-off between them. The appropriate balance, however, varies across queries and token budgets: localized questions favor relevance, whereas holistic questions demand broader visual coverage. We introduce StackTok, a training-free selector that treats query relevance as the objective and visual coverage as budget-calibrated support. StackTok builds a size-indexed coverage reference from a coverage-only greedy sequence and adjusts its support target using query--vision affinity entropy. A reference-gated interleaved selection policy then switches between relevance- and coverage-oriented additions according to the current subset's support deficit. For high-resolution inputs, StackTok allocates one shared token budget across crops according to the combined marginal gain of locally nominated tokens. Evaluated with five VLMs over ten distinct image-understanding benchmarks, StackTok ranks first among training-free selectors in every tested model--budget setting. On high-resolution LLaVA-NeXT-7B, it retains 95.26% of full-token performance with only 160 of 2{,}880 (5.6%) visual tokens.
cs.AI / 51 / 2609.16847
RegRet: Enhancing Region-Level Retrieval in Large Multimodal Models
Abstract
Region-level retrieval aims to align user-specified image regions with relevant regions or textual descriptions, playing a crucial role in realworld applications such as e-commerce product search and RAG. Although recent Large Multimodal Models (LMMs) have made significant strides in multimodal retrieval, they primarily focus on global-level tasks and struggle to capture effective region-level representations. To bridge this gap, we present RegRet, an LMM-based Region-level Retrieval framework that enhances the regional representations without compromising overall global retrieval performance. At its core, RegRet integrates a Region-Aware Encoder to capture detailed regional features while balancing them with the global background context. To further enhance the fine-grained understanding and discriminability of representations, we design a multi-stage training pipeline that includes detailed localized captioning and regional contrastive learning tasks. In addition, considering the absence of region-level contrastive training data and the limited diversity of evaluation tasks in current benchmarks, we introduce the REGMB benchmark. It comprises 225k contrastive pairs, covering four multimodal retrieval tasks. Extensive experiments validate the effectiveness of our approach. RegRet outperforms strong baselines in the zero-shot setting. Further training with contrastive learning leads to an average improvement of more than 20\% on both REGMB and public benchmarks, while achieving comparable or better results on global-level retrieval tasks.
cs.AI / 52 / 2609.16878
VOR-Bench: A Human Perception-Driven Benchmark for Video Object Removal
Abstract
Despite its crucial role in video object removal (VOR), existing evaluation paradigms face two critical limitations: questionable references and a misalignment between tradi- tional metrics and human preference. To address these challenges, we introduce VOR- Bench, which advances VOR evaluation through three integrated components. First, we present the VOR Dataset (VORD), the first benchmark dataset providing both paired edited videos and graffiti masks. Its unique strength lies in a diverse data spectrum, which encompasses model-generated, tool-rendered, and camera-captured data, ensuring robust assessment across real-world scenarios. Second, we develop rMPAF, a realistic Motion- capable Paired-video Acquisition Framework. By combining the strengths of image- based object removal and fine-tuned video generation models, rMPAF automatically generates realistic, motion-coherent paired videos. Finally, we propose three evaluation dimensions and introduce VOR-MDSM, the first perception-driven VLM-based scoring model specifically designed for mask-guided VOR. It bridges the gap between arithmetic metrics and human perception by covering the essential visual attributes and matching nuanced human judgment. Extensive experiments demonstrate that VOR-Bench yields evaluation results that align closely with human perception, achieving a remarkable cor- relation (\r{ho} > 0.9) with subjective assessments. We will release VOR-Bench along with its documentation to ensure full reproducibility.
cs.AI / 53 / 2609.17068
Beyond In-Distribution Metrics: A Systematic Out-of-Distribution Evaluation of Congenital Heart Disease Segmentation
Abstract
Congenital heart disease (CHD) diagnosis and surgical planning often require patient-specific 3D anatomical models, but manual segmentation is labor-intensive, particularly in complex anatomies. Although deep-learning methods can automate this process, they are typically evaluated in-distribution, despite clinically relevant shifts in scanner, protocol, institution, population, and imaging modality. We present, to our knowledge, the first systematic evaluation of out-of-distribution (OOD) generalization in CHD segmentation, using ImageCHD as a held-out target cohort. We compare representative segmentation architectures under combined CT and CMR training, CT-only training, self-supervised pretraining, and limited target-domain adaptation. In-distribution performance proves to be a poor indicator of cross-cohort robustness: nnU-Net achieves the highest validation Dice (0.77) but falls to 0.51 on ImageCHD, while SwinUNETR generalizes substantially better, reaching 0.67 Dice. MAE and JEPA pretraining provide only modest additional benefit, suggesting that architecture contributes more to robustness than the tested pretraining strategies in this setting. When limited target-domain supervision is introduced, all SwinUNETR variants exceed 0.76 Dice with only 11 labeled ImageCHD cases. These findings demonstrate that conventional in-distribution evaluation can obscure clinically important generalization failures and support explicit cross-dataset testing as a key component of CHD segmentation evaluation.
cs.AI / 54 / 2609.17152
ResLRP: The Role of Residual Cancellation in Attribution Instability in Vision Transformers
Abstract
Vision Transformers (ViTs) are central to most modern vision models, yet obtaining input attributions that are fine-grained, faithful, and stable remains challenging. Layer-wise Relevance Propagation (LRP) has been adapted to transformer attention, but in ViTs it often produces noisy, unfaithful explanations. We show that the missing ingredient is the treatment of residual connections: cancellation effects in residual pathways lead to attribution explosion. Moreover, we find that these cancellations are substantially stronger in ViTs than in language transformers. To address this issue, we introduce Residual-aware Layer-wise Relevance Propagation (ResLRP), a simple extension of LRP whose propagation rules explicitly account for cancellations in residual branches, are exactly conservative, and provably bound relevance explosion. Causal channel-wise interventions confirm that residual cancellation, not a generic regularization effect, drives the instability. ResLRP substantially improves attribution quality across faithfulness and localization, evaluated on ViT architectures spanning supervised, self-supervised, contrastive, hierarchical, and multimodal families, as well as on the ground-truth-controlled FunnyBirds benchmark. The largest gains arise in modern Vision Language Models (VLMs), with +27-29% localization and up to 3.4x faithfulness scores. Beyond benchmarks, ResLRP localizes Sparse Autoencoder (SAE) features in input space, and our residual amplification measure serves as an architecture-level diagnostic predicting where attribution degrades.
cs.AI / 55 / 2609.17181
Multimodal Cultural Heritage Architectural Style Classification for Residential Buildings in the UAE Based on CLIP Embeddings and SVM
Abstract
The analysis and classification of cultural heritage architectural styles remain challenging due to the complexity of visual images of buildings, which are highly relied on in traditional CNN-based classification approaches in comparison to textual descriptions, and the relative lack of non-western region-specific datasets. This paper addresses this gap by proposing a multimodal machine learning framework to analyze and classify Emirati residential architecture using OpenAI's CLIP model. We integrate visual features from images and textual features from expert descriptions into a unified 512-dimensional embedding, followed by dimensionality reduction with UMAP for visualization and unsupervised clustering using K-Means. Cluster labels, which are derived from manual analysis of the K-Means clusters, are used to train an SVM classifier for automated architectural style classification. Our approach achieves a classification accuracy of 98% across eight identified style clusters, higher than every other study in the literature, demonstrating the effectiveness of combining visual and textual modalities. Overall, this paper highlights the potential of using multimodal AI to support architectural heritage analysis, offering scalable and interpretable tools for exploring regional architectural identities.
cs.AI / 56 / 2609.17427
Tracking the Unseen: An Occlusion-Robust Framework for Target Tracking Under Full and Long-Term Occlusion
Abstract
Real-time multi-object tracking systems remain highly vulnerable to full and long-term occlusion, where targets temporarily or completely disappear from the camera's field of view. Conventional trackers may terminate trajectories prematurely, resulting in identity loss and reduced situational awareness in applications such as defense and surveillance. This work proposes an occlusion-robust target tracking framework that maintains target identity and trajectory continuity through the integration of YOLOv11n object detection, Kalman Filter motion prediction, and occlusion-aware appearance-based re-identification. The framework consists of three stages: object detection, position estimation during occlusion, and identity recovery after target reappearance. Six Re-Identification (Re-ID) architectures were evaluated within the same tracking framework under identical conditions, with the Occlusion-Aware Mask Network (OAMN) achieving the best overall performance and therefore selected for the final pipeline. The framework was benchmarked against OccluTrack on the public OVIS dataset, achieving relative improvements of 18.1 percent in Multiple Object Tracking Accuracy (MOTA) and 25.1 percent in Identity F1 Score (IDF1), while reducing identity switches by 12.8 percent. On a custom military dataset simulating surveillance and battlefield-like environments with long-term occlusion, the framework achieved a MOTA of 0.734 and an IDF1 of 0.729, corresponding to relative improvements of 14.2 percent and 5.8 percent over OccluTrack. The system demonstrated strong tracking continuity, robust identity preservation, and reliable trajectory estimation under challenging occlusion conditions, highlighting its effectiveness for defense-related surveillance applications requiring continuous target tracking during visibility loss.
cs.AI / 57 / 2609.17479
Det-LIME: Detector-Aware, Multi-Instance Local Interpretable Model-Agnostic Explanations for Automated Marine Mammal Detection
Abstract
Despite the rapid uptake of black-box object detectors in marine mammal research and monitoring, explainability techniques are rarely integrated into conservation workflows. Furthermore, most classification-oriented explainability tools are ill-suited to detection tasks involving imagery of social organisms or those with colonial life histories, as they ignore multiple detections within a scene and produce single-instance outputs that blur evidence across individuals. These methods also generate low-resolution, often biologically irrelevant visuals, limiting their utility for debugging, targeted data augmentation, and refined data collection. We proposed Det-LIME, a detector-aware, multi-instance adaptation of Local Interpretable Model-Agnostic Explanations (LIME) that produced instance-specific, box-aligned explanations by combining per-detection weighting, a proximity kernel that emphasizes regions near each box, and Intersection-over-Union-based matching to track the same instance across perturbations. We evaluated Det-LIME on aerial drone imagery for harbor seal detection, with an additional seabird case study to assess generality, and compared it with vanilla LIME, Stabilized LIME, Deterministic LIME, and gradient-based attribution methods. Using the Attribution Ratio and Max Saliency Hit Rate metrics, we showed that Det-LIME consistently improved multi-instance attribution. In practice, these higher-resolution, instance-aware explanations provide insight into model outputs and support post-processing, debugging, and actionable improvements in modeling and data collection or augmentation.
cs.AI / 58 / 2609.17521
PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control
Abstract
Interactive control for video generation is moving from coarse prompts toward fine-grained, physically meaningful manipulation of dynamic scenes. Yet existing controllable methods either require the full control schedule before generation starts, or use pixel-space signals that dictate object positions rather than physical dynamics. To address these limitations, we propose PhysStream, an autoregressive model for physics-grounded image-to-video synthesis that incorporates structured scene memory---positional maps and object tracking maps derived online from previously generated frames---and supports fine-grained motion control via sparse velocity-increment signals that encode physical quantities, letting the model learn the underlying dynamics. We train our model in two stages: a bidirectional model is first finetuned with motion-control conditioning, then a causal autoregressive model is trained with additional structured scene memory, further improving physical consistency. PhysStream enables interactive, mid-generation control over multi-object tabletop rigid-body scenes---a capability not supported by prior methods---reducing motion distribution distance (FVMD) by 33% and trajectory error by 12% over the strongest baselines on synthetic benchmarks, and is preferred by human evaluators in over 85% of in-the-wild comparisons. Please check our website for more details: https://czzzzh.github.io/PhysStream
cs.AI / 59 / 2609.16192
AI-Driven Feedback Systems, Digital Labour, and Silent Quitting: Transforming African Workplaces
Abstract
The current trend of digitalisation has revolutionised the organisation of work and the way it is measured and performed across the globe, with AI becoming more common for managing labour and performance, as well as employee communication. In African organisations, where there is increasing adoption of remote work, hybrid models of work, digital collaboration, and data-based HR management, the notion of silent quitting has become more relevant, defined as worker disengagement when employees are still doing their job but do not put any effort into achieving good performance and exhibiting any emotion. This paper investigates how AI-driven feedback mechanisms, including sentiment analysis systems, pulse surveys, chatbots, engagement dashboards, and predictive analytics, are changing African workplaces through offering continuous listening, instant performance information and proactive engagement with employees. The study also explores how AI can assist organisations in identifying early disengagement and enable intervention and better employee communication in both private and public sector organisations in Africa. At the same time, we address the challenges of socioeconomic development and governance posed by AI implementation in developing countries, including digital inequality, infrastructure shortcomings, privacy concerns, algorithmic bias, and the risk of workplace surveillance. By situating silent quitting within wider debates on digital labour and automation, the paper contributes an African-centred perspective to discussions on the future of work and offers practical recommendations for HR professionals, managers, policymakers, and technology developers seeking responsible, context-sensitive approaches to workplace transformation across the continent.
cs.AI / 60 / 2609.16260
Mapping U.S. Federal AI Governance Against Sector Vulnerability
Abstract
Artificial intelligence (AI) poses different levels of risk across sectors, but are these differences reflected in U.S. federal AI governance? To help answer this question, we assess 684 federal AI governance documents for their coverage of 14 sectors and 24 AI risks. We measure coverage as breadth (i.e., how frequently the risk or sector is addressed across documents) and depth (i.e., how substantively the risk or sector is discussed). We then compare sector coverage patterns for each of the 24 risks with vulnerability assessments from a Delphi study of 272 experts. Our analysis finds substantial variation in coverage: AI risks related to robustness, system security, and governance receive more attention than socioeconomic, environmental, and emerging risks, including multi-agent risks. Public administration, national security, information, and scientific services receive comparatively high levels of coverage relative to other sectors, such as finance and healthcare, which experts rate as highly vulnerable to AI risks. By mapping current coverage and identifying where it differs from expert assessments of vulnerability, we surface potential AI governance gaps which may help inform AI risk-related decisions across government and industry.
cs.AI / 61 / 2609.17111
Finding Common Mistakes In Modelling With Mathematical Formalisms Using LLMs
Abstract
Modelling with mathematical formalisms like logical formulas, mathematical equations, or regular expressions is an important yet challenging task for students of computer science and other STEM disciplines. Identifying common mistakes occurring in this context is an important step towards helping struggling students by providing targeted high-quality feedback, e.g. in interactive learning systems. We present a tool-supported workflow that allows to (1) identify candidates for common mistakes that explain many student mistakes in large educational data sets, (2) cluster candidates according to similarities, and (3) visualize resulting clusters for instructors and CS education researchers. The visualization is designed to help researchers to identify common modelling mistakes. The candidates for common mistakes are represented by bug fixing transformations that translate incorrect formalizations into correct formalizations; they are generated by an LLM and validated algorithmically. We show that this approach works well by reproducing common mistakes in propositional logic modelling that were identified by hand in the literature; showing that, unlike other algorithmic approaches, the LLM-based approach is suitable for very large sets of data; and applying it to multiple other formalisms to showcase it generalizes beyond propositional logic.
cs.AI / 62 / 2609.16313
Cognitive Admission Control: Risk-Conditioned Assurance for Consequential Actions in Agentic Distributed Systems
Abstract
In agentic distributed systems, an agent may be authorized to mutate external infrastructure while lacking evidence that the mutation is ready to execute. Cognitive Admission Control (CAC) makes this evidence requirement explicit. A policy maps a typed action and its modeled risk to assurance obligations specifying predicates, evidence classes, scope, freshness, and witness-set constraints. A deterministic evaluator distinguishes satisfied, violated, and unresolved obligations; unresolved conditions produce targeted evidence-acquisition requests. Successful admission produces a certificate binding the action, its witness manifest, and dispatch-time guards. We formalize the admission calculus and the assumptions connecting it to mediated execution. The guarantees are policy-relative: physical safety additionally requires sound evidence, an adequate environment model, and preservation of relevant conditions through the effect. A TypeScript prototype is evaluated in 2,730 controlled local trials with independent effect observation and matched fault schedules. Across 390 CAC trials, 120 effects complete without modeled harm and no harmful effects occur. A live-policy baseline achieves the same completion count but admits the constructed correlated-witness failure. Mechanism ablations isolate guard, evidence-class, structural-cut, and remediation behavior. A further 9,000 measurements exercise the complete local dispatch path with persistent replay protection. These results establish tested implementation behaviors and local costs, not production failure rates or comparisons of language-model capability.
cs.AI / 63 / 2609.16498
Geospatial Metadata Improves Discoverability by Connecting Datasets Across Scientific Disciplines
Abstract
Research data repositories are essential infrastructure for scientific inquiry and for ensuring that datasets follow FAIR (Findable, Accessible, Interoperable, and Reusable) principles. However, repository reuse depends on the quality and completeness of geospatial and thematic metadata, which researchers generally provide voluntarily. Given limited curation resources, it is unsurprising that even Harvard Dataverse, the world's largest general-purpose research repository, contains many incomplete metadata records. Missing fields represent lost information and reduce interoperability. We find that datasets with more missing metadata receive fewer downstream citations and have fewer resolvable connections to other datasets. The implications are particularly important for geospatial datasets: only 0.3% of research datasets include a bounding box, and most represent archival points rather than complete geographic shapes. Our analysis shows that geospatial metadata helps connect concepts across disciplines. After embedding Harvard Dataverse datasets in a metadata knowledge graph, we find that datasets are twice as likely to connect across scientific disciplines through shared geospatial metadata as through keywords. This suggests that geographic metadata is a more reliable basis for cross-disciplinary interoperability than keyword vocabularies, which often remain discipline-specific. We train and fine-tune a small language model using datasets from Harvard Dataverse. Through geospatial metadata enrichment, we increase the share of datasets from different disciplines connected through metadata elements from 58.5% to 63.2%.
cs.AI / 64 / 2609.16289
Symmetric solution of the Bellman optimality equation for repeated harmony game
Abstract
In social dilemma games, additional rewards or punishments have been studied as means of promoting cooperation. Therefore, it is important to investigate the ideal situation, in which such an additional payoff would change the game. In this study, we investigated the symmetric solution of the Bellman optimality equation for a repeated harmony game. The calculations showed that three types of symmetric solutions exist. One of them corresponds to the trivial All-C strategy, and another to the Win-stay Lose-shift strategy of the prisoners dilemma game. The nontrivial behavior of the strategy corresponding to the last solution is also discussed in detail. In addition, we numerically investigated which strategy the agents actually learn by the reinforcement learning algorithm.
cs.AI / 65 / 2609.16295
Intelligent Interaction Techniques (IIxT) - Proposal
Abstract
Interaction techniques (IxTs) are the low-level, reusable components out of which user interfaces are designed, including menus, scroll bars, text input fields, and also copy-paste, text-entry, and selecting objects. The IxTs for graphical user interfaces (GUIs) were well established in the 1980s, with relatively minor additions and tweaks for smartphones in the 2000s. Most of today's AI user interfaces involve a chat window, which is an excellent interaction for some tasks, but is generally considered separate from the GUI IxTs. I argue for making the IxTs themselves more intelligent, so users can freely mix modalities, even within the same interaction. This will require research into new IxTs, and also into the infrastructure that will enable these intelligent IxTs (IIxTs) to be built. There are also significant security, privacy and economic implications to this vision.
cs.AI / 66 / 2609.16344
From Momentary Emotion Inference to Sustained Emotion Support: Evaluating a Companion Agent in a Longitudinal Study
Abstract
Sustained emotional support is a long-horizon interaction task closely tied to human well-being. Recent research demonstrates generative agents' capacity for momentary emotional support, yet how these capabilities sustain support over time remains unclear. To examine this challenge, we deployed PAIR, a theory-based emotion-regulation companion, with 19 participants for 14 days. Across 1,093 sessions, we paired emotion estimates with self-reports before and after guidance and analyzed logs and interviews. Estimates corresponded more closely to self-reported valence and dominance than arousal. Guided conversations were followed by higher valence and state-dependent arousal changes. Participants felt understood through contextual exploration and emotional acknowledgment, acting on guidance suited to their needs and constraints. Perceived helpfulness of guided conversation significantly increased over time. Our findings link memory updates and retained corrections to cross-session personalization, informing future emotional support tools that adapt to evolving needs, learn from prior outcomes, and preserve user control over memory.
cs.AI / 67 / 2609.16856
The Evolution of Coordination in a Collective Intelligence System: 25 Years of English Wikipedia and the Emergence of Generative AI
Abstract
English Wikipedia is one of the largest examples of collective intelligence on the Web, sustained not only by article production but also by volunteer coordination and governance. While prior research has examined coordination work in Wikipedia, less attention has been paid to how participation in these spaces has evolved over time. Drawing on a longitudinal analysis spanning nearly 25 years of English Wikipedia, we examine editing patterns across five namespaces covering content, discussion, and governance. We find that participation in coordination spaces has declined relative to content production, particularly in governance areas, with a shrinking core of editors performing an increasing share of this work. Using Markov-based session metrics, we also find that editing has become more specialised, with editors moving less frequently between namespaces. Motivated by recent governance debates around generative AI, we conclude by investigating whether the availability of LLMs has altered these long-term trends. While short-term changes are visible, we find little evidence that generative AI fundamentally changed existing trajectories of coordination and participation.
cs.AI / 68 / 2609.17065
Beyond "ChatGPT Can Make Mistakes": Designing Interventions to Support Metacognitive Monitoring in AI-Assisted Work
Abstract
AI assistance places a metacognitive demand on users, who must judge their own competence and the system's. Yet designers lack comparative evidence on which interventions to choose, where to place them, and how to tell whether they worked. We elicited 30 interventions from 11 experts and, with prior work, organized them into a design space of time (when an intervention acts), level (whose competence is judged), and source (who supplies the monitoring cue). A between-subjects experiment (N = 917; 12 planning-and-organizing problems) compared a per-task reliability card, contrasting replies, pause points, and post-problem reflection against a baseline LLM assistant. Reliability cards and contrasting replies reduced estimation error and overconfidence and increased aggregate confidence discrimination. No task-performance improvement or average within-item discrimination gain was established. We contribute a shared vocabulary, a design space, and evidence that measured monitoring and task performance are separable design targets.
cs.AI / 69 / 2609.17434
CareMirror: Bringing Caregiver Wellbeing into the Dementia Care Ecosystem
Abstract
Family caregivers of people living with dementia shoulder emotional and practical responsibilities, yet their own wellbeing often remains peripheral to dementia care. We built CareMirror, an envisioned caregiver wellbeing ecosystem with interconnected caregiver- and clinician-facing interfaces for longitudinal reflection, personalized support, and caregiver-controlled sharing with clinical care. We conducted semi-structured interviews with 14 caregivers, using CareMirror as a design probe to examine how they perceived this ecosystem and what expectations, concerns, and boundaries emerged around clinical connection. Caregivers valued attention to their wellbeing, longitudinal awareness, context-sensitive support, and clinical visibility when it could lead to meaningful follow-up. However, repeated reflection could become burdensome or emotionally difficult, automatic clinical sharing could inhibit candid disclosure, and participants wanted control over what information entered clinical care. They also expected AI to support reflection and communication without replacing caregiver voice or clinician judgment. We contribute design considerations for proactive, clinically connected caregiver wellbeing support.
cs.AI / 70 / 2609.16407
Balancing Trial and Reorder: A Hybrid Sequential Transformer-GBDT Ranker for On-Demand Delivery
Abstract
On a delivery platform, personalized store ranking greatly influences what users find and order. Unlike digital-only domains, candidate stores are local and bound by real-time availability and delivery operations. One central modeling tension is between surfacing new stores for trial and preserving ranking quality for sessions with reorder intent. We present Universal Venue Ranker (UVR), a production system deployed at Wolt that pairs a bidirectional transformer encoder for sequential user modeling with a GBDT ranker integrating contextual, user, and store features. Trained across all stores and domains of a country while enforcing local delivery constraints at inference, UVR replaces four previously separate ranking models (three for restaurants, one for retail) with a single unified system. Label smoothing and trial-biased sample weighting steer the model toward new stores, lifting offline trial MRR by +12% to +30% over production while regressing reorder MRR in five of six countries. These regressions leave Global CVR, our core online metric, which blends trial and reorder sessions, statistically unchanged. We validate UVR in three consecutive A/B tests, the first two across Wolt's largest operating markets and the third spanning all operating countries and both domains. UVR V1 delivers +5.5% Merchant Trial Rate and +0.16% Global CVR over the previous production ranker; V2 adds a further +0.45% Merchant Trial Rate on top; and V3, our cross-domain unification of the restaurant and retail rankers, adds a further +1.31% Retail Merchant Trial Rate, together accounting for substantial incremental gross order value and a materially simplified serving stack.
cs.AI / 71 / 2609.16625
AURA: Agentic Diagnosis and Refinement for Production Recommender Systems at Scale
Abstract
How and why does a recommender system fail the users it serves? Oftentimes, practitioners are left to improve their algorithms based on a combination of feedback from stakeholder teams, domain expertise, and insights from data analyses. Yet the nuances of how and where recommendations perform well or poorly for end users are difficult to discern from aggregate quantitative metrics. Whereas these metrics provide a high-level and incomplete picture, further granularity into the quality of recommendations and their patterns requires reasoning with domain understanding and objectivity, at scale. We contemplate this complex conundrum and describe a method and implementation that uses the latest AI agentic advances to provide actionable diagnoses and improvements for production recommender systems. We present AURA (Agentic Understanding and Refinement of recommender Algorithms), an end-to-end agentic system that performs qualitative evaluation at scale and can then generate improvements to our algorithms at the code level. Specialized agents read production engagement logs, from thousands of sessions to millions, and surface patterns and examples of how the recommender fails real users. The next step uses those diagnoses and context about the recommender's own code, data, and training pipeline to propose and implement refinements grounded in that codebase. We report the system design, initial tests on production data from two large consumer platforms at a major media-streaming company, safeguards, operational learnings, and early results toward a self-improving recommender system. Finally, the diagnostic gap AURA closes is not specific to streaming. The architecture is built to transfer: every domain-specific element enters through the configuration layer that already ported it between our two platforms. We map it concretely to e-commerce and online-retail recommendation.
cs.AI / 72 / 2609.16947
AeroLat: Channel-Aware Latent Space Semantic Communication for Decentralized UAV Swarms
Abstract
Communication in latent space offers an intriguing alternative to symbolic messages for decentralized autonomous Unmanned Aerial Vehicle (UAV) swarms operating over bandwidth-constrained, time-varying wireless links. However, when homogeneous frozen models are prompted with discretized perceptual inputs, their broadcast states collapse toward the shared prompt template. In view of this, we propose AeroLat, a channel-aware latent semantic communication framework that uses evidence injection. The resulting latent states are then passed through an explicit communication model that encompasses bandwidth-limited serialization, additive noise and information staleness, which facilitates a joint assessment of communication fidelity and swarm-level coordination. Across multi-seed simulations, AeroLat provably remains resilient to codec choice, faults and increasing swarm size. It consistently reproduces the latent-swarm anomaly, while no-whitening controls recover the collapse. In particular, AeroLat is capable of reducing false similarity by 97.5%.
cs.AI / 73 / 2609.16346
Auto-HSI: Personalized human control of a robot swarm on demand by using LLMs for online automatic code generation
Abstract
This paper presents Auto-HSI, a method for generating personalized human-swarm interaction (HSI) interfaces on demand. The objective is to enable untrained operators to use natural language descriptions and gesture demonstrations to explain how they want the robots to collectively behave in response to their gestures. Based on these inputs, the code should automatically be generated for personalized state machines that will control the robots as desired, in response to the desired gesture inputs. In the developed Auto-HSI prototype, the generated code produces a personalized interface for centralized control using one- and two-handed gestures, enabling a user to teleoperate the robots' motion, formation shape, and shape deformation. We test the gesture tracking and code generation components of Auto-HSI against performance benchmarks. We then test the full Auto-HSI prototype in ``live'' operation experiments, in which real human operators centrally control 50 simulated robots in a physics-based simulator, under nominal and noisy conditions. In these experiments, robots are teleoperated to: score a goal, traverse a maze that requires shape deformation, and score two simultaneous goals by splitting into two groups. We also demonstrate a real human operator making live updates to their personalized Auto-HSI interface during operation (in simulation). Finally, we demonstrate live operation of real robots.
cs.AI / 74 / 2609.16368
UDAV: Uncertainty-Driven Adaptive VLM Waypoint Planner
Abstract
Vision-language models (VLMs) can generate routes directly from aerial imagery for off-road navigation, but their predictions provide no indication of reliability. We present UDAV, an Uncertainty-Driven Adaptive VLM Waypoint Planner for UAV-guided UGV navigation. UDAV draws multiple stochastic trajectory predictions, selects their medoid as a self-consistent nominal route, and estimates predictive uncertainty from their spatial dispersion. When the maximum uncertainty across interior waypoints exceeds a threshold, UDAV invokes a reconsideration stage; otherwise, it returns the medoid directly. We evaluate UDAV on 400 held-out trajectory queries from two UAV flights. Stochastic medoid selection reduces the mean average displacement error (ADE) from 147.4 pixels for a deterministic prediction to 115.9 pixels. The complete planner achieves a mean ADE of 110.4 pixels, a 25.1% reduction relative to deterministic planning, while producing valid trajectories for all queries. UDAV also yields the lowest 90th- and 95th-percentile errors among all evaluated configurations, including a higher-budget K=10 consensus baseline. Relative to the K=5 medoid, UDAV reduces these errors from 225.3 and 326.0 pixels to 199.0 and 290.8 pixels, respectively. These results demonstrate that stochastic VLM predictions provide both a stronger nominal route and an actionable uncertainty signal for selectively mitigating large planning errors.
cs.AI / 75 / 2609.16586
ProxiDex: Learning Dynamics-Guided Proximity Policy for Dexterous Manipulation
Abstract
Multi-finger dexterous manipulation relies on stable hand-object interactions, yet these interactions are partially observable in practice. Visual observations are often occluded by the hand, tactile sensors introduce hardware-specific modalities and calibration burdens, and existing policies rarely model how these cues evolve under actions, making them brittle under contact uncertainty. To address these, we present ProxiDex, a dynamics-guided proximity policy framework that treats hand-object proximity as an interaction state for dexterous manipulation. ProxiDex reconstructs interaction point clouds and converts geometric distances into proximity cues, forming a hardware-agnostic contact representation that provides immersive feedback during VR teleoperation. Built on this representation, ProxiDex learns action-conditioned proximity dynamics with a coupled forward-inverse design: future observation latents are predicted from actions, while proximity variations are decoded from latent changes. Leveraging these dynamics, ProxiDex adaptively reweights proximity tokens across manipulation phases and uses dynamics-consistency supervision to guide policy inference, stabilizing action generation under unreliable visual feedback. Simulation and real-world experiments demonstrate improved success rates and robustness over representative baselines across standard, unseen objects, and perturbation scenarios. Additional visualizations are available at https://proxidex.github.io/.
cs.AI / 76 / 2609.16683
Weave: Learning Whole-Body Dexterous Loco-Manipulation from Human-Object Interactions
Abstract
Learning humanoid-object interaction requires coordinating whole-body balance, locomotion, and dexterous hand contact to control both robot and object motion. Human demonstrations provide examples of coordinated interaction, but transferring these behaviors to humanoid robots requires learning how to establish and maintain effective contacts under different embodiments and dynamics. We present Weave, a unified framework for learning whole-body dexterous humanoid-object interaction from captured human demonstrations. Weave first converts captured human-object interactions into executable robot-object references through contact-aware retargeting and approach-motion completion. At its core is a contact- and geometry-aware policy that jointly commands 29 body joints and 12 actuated finger joints across multiple objects and interaction sequences. Evaluation across nine objects yields a 92.5% success rate on trained interactions and, without any additional training, 65.0% on sequences never seen during training. We additionally release ~9,000 physically executed rollouts spanning ~23 hours, providing robot-object trajectories with contact annotations for downstream interaction-policy learning and physically consistent HOI motion generation. Project website: https://xiaohu-art.github.io/Weave/
cs.AI / 77 / 2609.16697
World Models for Embodied Intelligence: From Plausible to Controllable to Actionable
Abstract
World models connect perception and decision-making in embodied intelligence by maintaining hidden state, anticipating consequences, comparing interventions, and adapting when execution departs from expectations. Although progress is often measured by visual fidelity, their value lies in improving behavior. Before reaching for a cup, a person anticipates its weight and resistance to grasping, shaping the hand before contact. Such anticipation is coarse and rarely pictorial, yet it guides action. This raises a central question: which predictive capabilities improve behavior? Existing surveys, organized by architecture, output modality, or application domain, leave this question implicit. We introduce three progressively stronger capability levels: Plausible models preserve task-relevant temporal, geometric, or physical structure; Controllable models additionally predict how interventions alter that structure; and Actionable models translate predictions into measurable gains in planning, action, learning, evaluation, verification, recovery, or data selection. We complement this hierarchy with a 3 x 4 matrix crossing geometry, physics, and action grounding with improvement loops centered on data, rewards, policies, and the model itself. Using this framework, we survey manipulation, navigation, locomotion, autonomous driving, and general embodied learning, tracing technical progressions, clarifying capability requirements, and examining datasets, benchmarks, and evaluation protocols. We identify challenges in long-horizon consistency, uncertainty calibration, causal intervention testing, latency, verification and recovery, and cross-embodiment transfer. This perspective shifts evaluation from visual plausibility toward whether predictions capture task-relevant state, reflect intervention effects, and improve the closed-loop behavior of embodied agents.
cs.AI / 78 / 2609.16737
Seeing What Matters: Visual Cue Guided Video Planning for Generalizable Robot Navigation
Abstract
Generative video models can serve as a promising backbone for robot navigation by predicting future observations as video plans. Recent approaches often condition video planning on short-horizon guidance and recover geometric waypoints through scene reconstruction, leaving longer-horizon planning and precise video-to-action translation less explored. We present CueNav, a video model-based navigation framework combining visual cue guided video planning with an embodiment-specific Inverse-Dynamics Model (IDM). As visual cues, we use a Bird's-Eye View (BEV) map to convey global task context and retain part of the robot body in the egocentric observation to expose embodiment context. These cues guide the video planner, while the IDM translates dense flow fields extracted from the video plan into robot actions. With the visual cue encoding global task context, CueNav achieves nearly 2x higher success in maze navigation than planning without the cue. The body-aware view with the IDM enables precise navigation with 70% success in a narrow passage where comparison methods largely fail to complete the task. We further demonstrate zero-shot semantic-conditioned navigation and deployment of the same video planner across different robot platforms. Our results show that visual cue-guided video planning with embodiment-specific action grounding paves the way toward a generalizable navigation framework for longer-horizon planning and embodiment-aware control. Additional results and code are available on our project website: https://cuenav.github.io.
cs.AI / 79 / 2609.17141
Continual Learning for Traversability Prediction with Uncertainty-Aware Adaptation
Abstract
Traversability prediction is a critical component of autonomous navigation in unstructured environments, where complex and uncertain robot-terrain interactions pose significant challenges such as traction loss and dynamic instability. Despite recent progress in learning-based traversability prediction, these methods often fail to adapt to novel terrains. Even when adaptation is achieved, retaining experience from previously trained environments remains a challenge, a problem known as catastrophic forgetting. To address this challenge, we propose a continual learning framework for traversability prediction that incrementally adapts to new terrains using a generative experience recall model. A key virtue of the proposed framework is two folds: i) retain prior experience without storing past data; and ii) incorporate the uncertainty of the generated samples from the recall model, enabling uncertainty-aware adaptation. Real-world experiments with a skid-steering robot validate the effectiveness of the proposed framework, demonstrating its ability to adapt across a series of diverse environments while mitigating catastrophic forgetting.
cs.AI / 80 / 2609.17147
Kernel-Based Metrics Learning for Uncertain Opponent Vehicle Trajectory Prediction in Autonomous Racing
Abstract
Autonomous racing confronts significant challenges in safely overtaking Opponent Vehicles (OVs) that exhibit uncertain trajectories, stemming from unknown driving policies. To address these challenges, this study proposes heterogeneous kernel metrics for Deep Kernel Learning (DKL), designed to robustly capture the diverse driving policies of OVs, and carry out precise trajectory predictions along with the associated uncertainties. A key virtue of the proposed kernel metrics lies in their ability to align similar driving policies and disjoin dissimilar ones in an unsupervised manner, given the observed interactions between the Ego Vehicle (EV) and OVs. The efficacy of the proposed method is substantiated through experimental studies on a 1/10th scale racecar platform, demonstrating improved prediction accuracy and thereby safely overtaking against OVs. Furthermore, our method is computationally efficient for onboard computing units, affirming its viability in fast-paced racing environments. The video and source code can be found at https://github.com/HMCL-UNIST/OpponentPredictionWithKMDKL.git.
cs.AI / 81 / 2609.17210
FluxVLA Engine: A One-Stop VLA Engineering Platform for Embodied Intelligence
Abstract
Vision-language-action (VLA) models, world-action models (WAMs), and offline reinforcement learning methods are rapidly expanding the design space of embodied policies, yet turning these algorithms into reliable robot systems remains constrained by fragmented data formats, training stacks, evaluation protocols, inference runtimes, and embodiment-specific interfaces. We present $\mathrm{FluxVLA}$ Engine, an open, configuration-driven platform that turns heterogeneous embodied-policy components into a reproducible data-to-deployment workflow. Rather than introducing another policy model, $\mathrm{FluxVLA}$ standardizes interfaces for datasets, visual-language and world models, action heads, reward- or advantage-weighted learning, distributed training, simulation evaluation, optimized inference, and robot operators. The engine further integrates compositional dual-arm simulation, scalable automatic data generation, and model-decoupled human-in-the-loop rollout, takeover, correction collection, and reward annotation. For responsive physical execution, it combines Real-Time Chunking (RTC) with accelerated inference backends, lightweight remote GPU serving, and configurable trajectory post-processing. Together, these capabilities connect offline learning, simulation validation, online correction, and real-robot execution through shared and auditable contracts. $\mathrm{FluxVLA}$ therefore targets the engineering bottlenecks separating promising embodied-learning algorithms from reproducible evaluation and dependable deployment. Code is available at https://github.com/FluxVLA/FluxVLA
cs.AI / 82 / 2609.16612
Structure Across Voices: Comparing acoustic-event type accumulation and sequence dependence across four vocal repertoires using frozen audio encoders
Abstract
Vocal repertoires can differ in acoustic-event type accumulation and temporal organization, yet direct comparison is difficult because corpora use different native events and unequal amounts of sequence. We compare sperm whale codas, human speech phones, Bengalese finch syllables, and common marmoset calls using the same frozen-audio-encoder procedure while matching event count and local sequence opportunity. Whale shows the fastest type accumulation; Finch shows the strongest immediate dependence and repeated-subsequence recurrence. Physically interpretable acoustics recover complementary parts of this profile, continuous analyses without clustering support broad Whale acoustic coverage, and source- and position-preserving nulls retain both Finch order effects. Extending predictive context shifts the comparison toward Whale. Thus repertoire differences depend on the acoustic property and temporal scale measured rather than forming a single hierarchy.
cs.AI / 83 / 2609.17509
LACE: Layer-Wise Compression for Dynamic Frame Rate Codecs
Abstract
Neural audio codecs are a key component in speech language modeling. However, their high frame rates lead to long sequence lengths, increasing computational costs. Dynamic frame rate codecs mitigate this by reducing the effective frame rate using a compression step to merge multiple frames together. However, most prior methods either operate on single-codebook codecs or apply a single compression step before multi-layer quantization. This forces all quantization layers to share the same segmentation boundaries, despite the residual embeddings at different quantization layers exhibiting different rates of change over time. We propose LACE (Layer-Adaptive Codec Encoding), a dynamic frame rate codec that applies an independent compression step at each quantization layer, enabling layer-specific segmentation boundaries. To use LACE tokens in downstream text-to-speech (TTS), we further introduce union alignment and boundary anchor mechanisms to make durations consistent across layers while preserving compression benefits. Experiments on LibriTTS show that LACE offers a better rate-quality tradeoff than prior dynamic frame rate methods on the reconstruction task and improves TTS inference efficiency while maintaining competitive synthesis quality. Our code is released as part of the ESPnet3 codec recipe.
cs.AI / 84 / 2609.16412
On the Expressive Power of Implicit Line-Graph Higher-Order Weisfeiler--Leman
Abstract
Whitney's theorem allows isomorphism testing for connected simple graphs, apart from $K_3$ and $K_{1,3}$, to be formulated as distinguishing their line graphs. However, the relation between fixed-dimensional Weisfeiler--Leman (WL) expressivity on line graphs and on their roots remains unresolved. We study this relation through Implicit Line-Graph WL (ILG-$k$-WL), which is exactly $k$-WL on $L(G)$, executed over the edges of $G$ with line-graph relations derived from endpoint incidence and without explicitly constructing $L(G)$. On the Whitney-general class, the relation between root-domain and line-graph WL depends on $k$. For $k=1,2$, ILG-$k$-WL adds no distinguishing power beyond root-domain $1$-WL and misses some pairs that $1$-WL separates. For $k=3$, we prove the backward containment $L(G)\equiv_{3\text{-WL}}L(H)\Rightarrow G\equiv_{3\text{-WL}}H$. Strongly regular witness pairs, including the Shrikhande/rook pair, show that ILG-$3$-WL is strictly more expressive than $3$-WL. The backward containment also extends to disconnected graphs with no isolated vertices when every connected component is Whitney-general. Deterministic ILG-$3$-WL separates all three substructure-counting witness pairs, all $105$ pairs in SR25, and $359$ of $400$ BREC pairs. An untrained dense ILG-$3$-GNN gives the same pairwise verdicts on these evaluations.
cs.AI / 85 / 2609.17439
Evaluating Verified Autonomy in Quantum Engineering
Abstract
Reliable quantum engineering is essential for turning quantum phenomena into practical technologies. As quantum platforms grow in scale and complexity, their characterization and operation require increasing human effort and coordination. Scientific artificial intelligence agents, which can plan experiments, operate instruments, and analyze observations, offer a promising route towards autonomous quantum engineering. Yet whether current agents can perform reliably in this setting has not been systematically established. To fill this gap, we developed Quantum-Harbor, a virtual laboratory that provides a controlled execution environment for agents to interact with quantum systems. This design enables direct verification of both the actions taken and the conclusions drawn. Building on this framework, we introduce QIQCBench, a benchmark of $49$ expert-authored tasks spanning multiple layers including calibration and control, error correction and compilation, sensing and networking. Across $17$ frontier agentic systems, QIQCBench reveals wide variation in verified performance. These results expose a substantial gap between demonstrating capability and achieving reliable operation, and establish Quantum-Harbor as a foundation for measuring progress towards verified autonomy in quantum engineering.
cs.AI / 86 / 2609.16931
Causal Discovery via Transformed Low-Rank Quantile Surfaces
Abstract
We propose Low-Rank Quantile Surfaces (LRQS), a bivariate causal model in which, in the causal direction, an unknown monotone transformation of the conditional quantile surface admits a low-rank functional decomposition. LRQS subsumes location-scale noise models and post-nonlinear heteroscedastic noise models, while allowing multiple quantile bases to represent changes beyond location-scale effects. We prove generic identifiability of LRQS: the transformed quantile surface is low rank in the causal direction, whereas reverse representability under the corresponding constraints occurs only for exceptional, fine-tuned cause marginals. We provide a simple-yet-powerful causal score using a nonparametric fitting procedure that alternates between rank-constrained approximation of discretized quantile surfaces and isotonic estimation of the unknown monotone transformation. Experiments on synthetic mechanisms with higher-rank distributional shape variation and strong nonlinear distortions, together with standard bivariate benchmarks, show that LRQS is especially effective when conditional distributional shape or observation distortion goes beyond existing location-scale assumptions.
机器学习 (cs.LG)
103
cs.LG / 1 / 2609.17477
Bias-Induced Crossover in Absolute Capacity of Dense Associative Memory
Abstract
The absolute capacity of dense associative memory has mainly been analyzed for unbiased patterns. Here we examine the effect of bias in centered binary patterns under the Krotov-Hopfield single-site criterion $P_{\mathrm{error}}=1/N$, where $P_{\mathrm{error}}$ is the probability that a single-site flip lowers the energy of a stored pattern and $N$ is the number of neurons. Each pattern component takes $1-q$ with probability $q$ and $-q$ otherwise, where $0<q\le1/2$. For polynomial interactions of order $n$, a signal-to-noise analysis gives an absolute capacity of order $N^{n-1}/\ln N$ at $q=1/2$. For fixed $q<1/2$, however, the capacity is $O(N^{n/2})$ for even $n\ge4$ and $O(N^{(n+1)/2})$ for odd $n\ge5$. For $n=3$, both the unbiased and fixed-bias capacities remain $O(N^2/\ln N)$. For $n\ge4$, these different asymptotic forms imply a nonuniform large-$N$ limit near $q=1/2$. Asymptotic matching predicts a bias-induced crossover in the region $1-2q=O(\ln N/N^{\lfloor n/2\rfloor-1})$. The crossover originates from a bias-dependent crosstalk mean that reduces the stability of sites carrying the more frequent value $-q$. Computer simulations are compared with the finite-size conditioned-Gaussian predictions. An activity-dependent control potential that cancels the conditional crosstalk mean restores the $N^{n-1}/\ln N$ capacity for fixed $0<q<1/2$ within the conditioned-Gaussian approximation.
cs.LG / 2 / 2609.16306
Sequence Recognition in Bharatnatyam dance
Abstract
Bharatanatyam is the oldest Indian Classical Dance (ICD) which is learned and practiced across India and the world. Adavu is the core of this dance form. There exist 15 Adavus and 58 variations. Each Adavu variation comprises a well-defined set of motions and postures (called dance steps) that occur in a particular order. So, while learning Adavus, students not only learn the dance steps but also take care of its sequence of occurrences. This paper proposed a method to recognize these sequences. In this work, firstly, we recognize the involved Key Postures (KPs) and motions in the Adavu using Convolutional Neural Network (CNN) and Support Vector Machine (SVM), respectively. In this, CNN achieves 99% and SVM's recognition accuracy becomes 84%. Next, we compare these KP and motion sequences with the ground truth to find the best match using the Edit Distance algorithm with an accuracy of 98%. The paper contributes hugely to the state-of-the-art in the form of digital heritage, dance tutoring system, and many more. The paper addresses three novelties; (a) Recognizing the sequences based on the KPs and motions rather than only KPs as reported in the earlier works. (b) The performance of the proposed work is measured by analyzing the prediction time per sequence. We also compare our proposed approach with the previous works that deal with the same problem statement. (c) It tests the scalability of the proposed approach by including all the Adavu variations, unlike the earlier literature, which uses only one/two variations.
cs.LG / 3 / 2609.16637
Can Knowledge Transfer Parameters Be Learned? LePoKet for Efficient Robotic Vision
Abstract
Efficient perception is central to robotic systems operating under constrained computation, memory, and latency budgets. Knowledge transfer from larger pretrained models offers a practical route to stronger compact perception networks, but existing approaches commonly rely on fixed distillation objectives or manually designed interaction mechanisms. Building on Hereditary Knowledge Transfer (HKT), we propose LePoKet (Learnable Parameter Optimization for Knowledge Transfer), a structural transfer framework that embeds knowledge inheritance directly into the forward computation. LePoKet introduces a block-wise Extract-Transform-Mix interface whose interaction parameters are optimized jointly with the child network through a Learnable Genetic Attention (LGA) operator, without auxiliary distillation losses or temperature scaling. We first characterize the mechanism on CIFAR-10 and CIFAR-100 using ResNet parent-child pairs, obtaining relative error reductions of 24.57% and 25.1%, respectively, over standard child training. We then evaluate LePoKet for dense motion estimation by integrating it into a compact RAFT-based optical-flow model trained only on FlyingChairs and FlyingThings3D. LePoKet improves the compact RAFT baseline from 2.21 to 1.92 EPE on Sintel Clean, from 3.35 to 3.01 on Sintel Final, and from 7.51 to 6.39 on KITTI. A direct comparison with HKT further shows that LePoKet improves CIFAR-10 accuracy from 92.40% to 93.40% while achieving the best Sintel Final and KITTI errors among the evaluated compact transfer variants, with comparable performance on Sintel Clean. These results demonstrate that learnable structural transfer generalizes across recognition and motion perception tasks and provides a promising approach for efficient robotic vision.
cs.LG / 4 / 2609.16859
Measuring Annotation Efficiency for Handwritten Devanagari Recognition: Sample-Complexity Curves for Four Pretraining Regimes
Abstract
To train handwritten text recognition systems we need word images and their corresponding transcriptions, and these transcriptions are produced manually. For a script that can be read by only a small number of specialists, this manual transcription is a limitation, because the trained models are supposed to save the time of those same specialists. A relevant question therefore arises: how many transcriptions are needed before a recogniser becomes useful, and how much of that cost can pretraining remove? In this study the answer is measured directly for handwritten Devanagari. We keep the recogniser, optimiser and evaluation protocol the same and change only the number of real transcribed words used for fine-tuning across nine budgets from 10 to 4,000 and four initialisation regimes, with six seeds at every point. The resulting curves are then converted into annotation-equivalent terms. A CER of 0.50 is reached by supervised synthetic pretraining using only 81 transcribed words, whereas random initialisation requires 355, which gives a label multiplier of 4.40 [3.56, 4.99]. There is a zero-shot reference point as well: with no real transcribed words at all, this pretraining is worth about 136 of them. This advantage gets smaller as the target accuracy improves, and at the most demanding target we measure, it cannot be distinguished from no saving at all. A fourth arm in which only the encoder is transferred separates the effect of the pretraining method from that of transfer scope, and masked image modelling is observed to transfer negatively over a bounded range of budgets. We emphasise that the scarcity in this study is constructed by subsampling a large corpus.
cs.LG / 5 / 2609.16873
NeuroTS-Net: Multi-Class Semantic Segmentation of Pediatric Brain Tumors in Multi-Modal MRI
Abstract
Pediatric brain tumors are a leading cause of cancer-related mortality in children, and their small, rare, and often low-contrast subregions make accurate manual delineation challenging. Reliable automated segmentation is therefore needed to support diagnosis, treatment planning, and response assessment. Accordingly, we introduce NeuroTS-Net, a three-dimensional encoder-decoder convolutional neural network architecture for multi-class semantic segmentation that incorporates a dual-scale raw-detail stream, adaptive low-resolution context selection, and detail-preserving multipath downsampling. These components preserve fine intensity and boundary information while efficiently modeling broader tumor context. NeuroTS-Net was trained on the BraTS 2026 pediatric dataset without external data or pretrained weights and evaluated against nnU-Net and MedNeXt under the same experimental protocol. NeuroTS-Net outperformed the baseline methods, achieving whole-tumor and tumor-core Dice scores of 0.938 and 0.937 on the internal validation set and 0.927 and 0.926 on the official challenge validation set. The code is open-sourced at: https://github.com/maenstru56/NeuroTS.
cs.LG / 6 / 2609.16934
MedPCFM-TED: One-Step Point Cloud Flow Matching for Implant Generation via Teacher-Guided Endpoint Distillation
Abstract
Cranial implant generation is an important task in medical imaging. Recent point cloud based generative methods, particularly flow matching, offer strong reconstruction quality and efficient sampling, but still require multiple neural function evaluations during inference. This limits rapid generation of multiple plausible implant candidates. We propose Teacher-guided Endpoint Distillation (TED), a simple one-step distillation framework for conditional cranial implant generation on point clouds. TED trains a one-step student using teacher-guided endpoint supervision and geometric matching losses, while avoiding explicit path straightening. We evaluate TED on the SkullFix and SkullBreak benchmarks. TED achieves the best overall performance on the SkullBreak dataset, remains competitive on SkullFix, and provides the strongest Chamfer distance performance among the compared one-step methods. In addition, TED generates implants in approximately 0.04s per sample. These results show that one-step distillation can substantially accelerate conditional point cloud implant generation without sacrificing reconstruction quality.
cs.LG / 7 / 2609.17138
From Foundation Embeddings to Cropland Maps: Label Efficiency, Temporal Transferability and Independent Human Validation
Abstract
Geospatial foundation models provide reusable representations of satellite imagery that support downstream mapping with limited task-specific modelling. We evaluate whether annual AlphaEarth embeddings support binary cultivated-versus-non-cultivated mapping in Maine, USA, using 192 spatially separated patches and labels derived from the USDA Cropland Data Layer (CDL). Without fine-tuning the foundation model, a lightweight classifier reaches 93.7% overall accuracy and 90.8% balanced accuracy on held-out patches. Logistic regression is within 0.3 percentage points of a gradient-boosted ensemble, while a nearest-class-centroid rule, which uses class centroids but fits no parameters, reaches 90.2%. A balanced sample of 60,000 labelled pixels is within 1.3 percentage points of the full pool of 8.6 million pixels; because pixels are spatially autocorrelated, this result concerns pixel-sample efficiency rather than 60,000 independent annotation sites. In a same-region transfer experiment, classifiers trained in one year remain accurate across 2018 to 2023. Against a blind, two-interpreter consensus at 385 randomly sampled points in one contiguous 2023 block, the AlphaEarth-plus-random-forest map agrees at 95.3% ($κ=0.82$), compared with 91.7% for the CDL ($κ=0.72$; exact two-sided McNemar $p=0.0161$). This local result is consistent with partial smoothing of CDL label noise, but it does not establish statewide correction of the reference product. On the same points, the difference from a fine-tuned TerraMind segmentation model is not statistically significant (95.3% versus 93.5%; $p=0.14$), and the experiment is not a controlled comparison of computational cost. These results support frozen geospatial embeddings as a low-compute candidate for regional cropland mapping, subject to the limits of a single-state study, a 30 m-derived training reference, and a one-block human validation.
cs.LG / 8 / 2609.17458
Tables Decoded: DELTA for Structure, TARQA for Understanding
Abstract
Table understanding is a core task in document intelligence, encompassing two key subtasks: table reconstruction and table visual question answering (TabVQA). While recent approaches predominantly rely on vision- language models (VLMs) operating on table images, we propose a more scalable and effective alternative based on structured textual representations. These representations are easier to process, align more naturally with LLMs, and eliminate the need for language-specific visual encoders, making them particularly suitable for multilingual documents. We present DELTA, which separates physical structure recognition, logical structure recognition, and OCR to extract both layout and content accurately. DELTA outputs tables in Optimised Table Structure Language (OTSL), a compact and unified format that encodes cell arrangements and textual content. On table structure recognition (TSR), DELTA achieves TEDS- Structure scores comparable with state-of-the-art methods across FinTabNet, PubTabNet, and PubTables-1M. We further establish its robustness on non-English tables through our curated Hindi benchmark, TORQUE. Building on this, we introduce TARQA, an LLM fine-tuned on OTSL sequences. Our approach yields gains of 9.3 p.p. on WTQ (TabQA) and 9.2 p.p. on FinTabNetQA (TabVQA), respectively. On TORQUE, our method ranks second among all VLMs and DELTA + LLM variants. We release our code, models, and benchmark at: https://github.com/Tihiitborg/Tables-Decoded
cs.LG / 9 / 2609.16751
Constant Swap Regret in General-Sum Games via Optimistic Transition Matrices
Abstract
We give deterministic and uncoupled learning dynamics for finite multiplayer general-sum games under full-information feedback that achieve constant individual swap regret, independent of the horizon $T$. With $n$ players and at most $m$ actions each, the individual swap regret of every player is $O(\sqrt{n} m \log m \log^{5/2}(nm))$ at every finite horizon. Each player predicts the deviation gains, then uses these predictions to update a row-stochastic transition matrix, and plays its stationary distribution. The proof combines a potential argument exploiting stationarity with a two-scale higher-order prediction analysis, using rooted-tree representations to handle the nonlinear dependence of deviation gains on the stationary distributions. An adversarially robust variant, obtained through a generic common-prefix switching wrapper, preserves the self-play bound up to a universal constant and guarantees individual swap regret at most $7\sqrt{m T \log m}$ in the adversarial setting.
cs.LG / 10 / 2609.16170
Skeletal Prototypes on Iterative Nerve Expansions
Abstract
Prototype reduction replaces a training set with a smaller representation, and the established methods return a finite set of points. We propose Skeletal Prototypes on Iterative Nerve Expansions (SPINE). The model for each class is an embedded 1-complex rather than a point set. Its initial edge set is a class-conditional Mapper graph, so the data decide which localized clusters are joined. Later phases fit the vertices under a classification objective, and an observation is assigned to the class whose complex is nearest. The segments therefore enter the decision rule and not only the fitting. We evaluate SPINE on seventeen benchmark datasets under stratified 10-fold cross validation, against seven other prototype reduction methods at a matched budget. SPINE attains the highest mean accuracy and the best average rank. It is significantly better than five of the seven competitors under Wilcoxon signed-rank tests with Holm correction. A budget sweep shows that the decision rule using the entire graph segments contribute most when prototypes are scarce, while the method as a whole competes best at moderate budgets. Construction cost places SPINE with the discriminative methods, and it is faster than generalized learning vector quantization on fourteen of the seventeen datasets.
cs.LG / 11 / 2609.16183
Anatomy of Associative Recall in Fixed-State Recurrences: A Matched-State Decomposition, an Interference Wall, and a Curriculum That Breaks It
Abstract
Fixed-state recurrences--linear attention and state-space models--are reported to lag behind attention on associative recall, but whole-architecture comparisons cannot say which ingredient is responsible. We decompose masked multi-query recall at a fixed state budget along three single-knob axes: a short causal convolution, the transition structure (rank-1 delta rule vs. diagonal), and decay. The convolution dominates (~+0.5 recall in both families under matched training): comparisons that pit convolution-free cells against a convolution-equipped Mamba measure the missing convolution, not the recurrence. The rank-1 transition beats its diagonal ablation by +0.19/+0.32 at 16/32 pairs, but the margin shrinks to +0.03 once both cells carry the convolution, and a state-matched Mamba-2 ties the unarmed rank-1 cell: no class claim survives. Cells that solve 32-pair recall degrade gracefully with load yet fall to chance retrieving 4 pairs from a distractor haystack--flat across lengths and transitions. Interference under sparse supervision, not capacity: a distance curriculum takes the unchanged architecture from 0.021 to 1.000. Training is a lock-in lottery--a seed either locks in or does not--and the curriculum is the lever. Lock-in rises from 1/10 to 7/10 (p=0.02); dense supervision adds nothing; at L=256 a shaped ramp reopens a boundary the uniform curriculum cannot (4/5 vs. 0/9); and at L=512, where the ramp collapses (0/6), gating it on measured accuracy locks in 6/6 (p=0.001). Bidirectional denoiser cells, reading the query before the haystack, show no measurable advantage over causal training (ten seeds), and collision-key retrieval needs two layers. Arming for recall is free on an S_5 state-tracking guardrail--the armed cell is significantly better at every depth (p<=0.0044). These replace "recurrent models are bad at recall" with a measured decomposition and two cheap interventions.
cs.LG / 12 / 2609.16204
Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration
Abstract
Safety guardrails in open-weight language models can be readily bypassed using Refusal Feature Ablation (RFA), a technique that identifies and projects out a linear refusal direction from the residual stream, often achieving a high attack success rate (ASR) while preserving model capability. Defending against these attacks typically requires computationally expensive safety finetuning for every new checkpoint. We introduce Decoy Direction Optimization (DDO), a fast, post-hoc weight-editing defense that requires no base-model finetuning. Our approach is based on a simple mechanistic insight: ablation attacks rely on contrastive estimators to find the refusal direction. Rather than trying to hide the true refusal circuitry, DDO actively injects a high-magnitude, nonlinear decoy signal into the network's MLP neurons. When an attacker attempts to locate the refusal direction, the decoy corrupts their estimator, tricking them into ablating a harmless orthogonal feature while the actual safety mechanism remains intact. We prove a spectral bound formalizing this effect and evaluate DDO across six model families, achieving <10% ASR under standard RFA. On Llama-3-8B-Instruct, DDO remains comparable to trained defenses under adaptive multi-phase attacks (65% vs. 58% worst-case ASR) and reduces Heretic weight-level attack ASR from 88.7% to 18%, all at 30 to 450 times lower optimization cost per configuration than the trained baselines.
cs.LG / 13 / 2609.16255
Efficient Reasoning Distillation: Small Video-Language Models via Synthetic CoT and Difficulty-Aware Fine-Tuning
Abstract
We present an efficient method to distill reasoning capabilities into compact video-language models (VLMs) for video question answering (VideoQA). Our approach fine-tunes a 2B-parameter model using only $\sim$900 uncertainty-selected examples, each augmented with synthetic chain-of-thought (CoT) rationales generated by a 4B teacher. Despite its minimal compute cost - under two hours on a single A100 GPU - our method enables the 2B model to outperform VLMs up to 4$\times$ larger, and generalize across CinePile, ActivityNet-QA, and MLVU, approaching the performance of its own 4B teacher. A key finding is that placing CoT rationales after the answer - contrary to standard prompting - substantially improves reasoning in compact models. This insight challenges prevailing CoT conventions and reveals new alignment strategies under limited model capacity. Our findings offer a practical blueprint for training deployable, reasoning-rich VLMs suited for mobile and edge applications.
cs.LG / 14 / 2609.16267
The record is part of the task: matched-record evaluation of text classifiers across maintenance, safety and recall reporting
Abstract
Many operational cases are documented more than once, at different workflow stages and for different purposes, yet model evaluations normally select one of these records before model comparison begins. We treat that selection as part of the evaluation and compare matched records of the same cases under fixed labels and splits in three systems: GE Aerospace repair events, NASA ASRS safety reports and NHTSA vehicle recalls. Across the three GE fields, for events whose label comes from parts transactions independently of the narratives, held-out macro-F1 ranged from 0.33 to 0.91. A difference of 0.46 separated the customer report, written before shop work, from the technician report, written after diagnosis but before the transaction that generates the label. That difference is substantially larger than the representation and architecture differences tested on the same events. The public systems showed different patterns: the NHTSA defect summary remained strongest under every model family tested, whereas the ASRS analyst synopsis outperformed the reporter narrative under learned sequence models but not under lexical baselines. Secondary analyses showed that some model comparisons were also record-dependent. Evaluations should be run on the information available at the intended decision point and should report how both the record and the label were produced.
cs.LG / 15 / 2609.16282
Scaling Laws for Physics-Aware ACOPF Surrogate Learning
Abstract
Learning-based surrogates for AC optimal power flow (ACOPF) promise large speedups over classical solvers, but their operational value depends on physical feasibility as much as predictive accuracy. Physics-aware objectives such as the augmented Lagrangian (AL) improve constraint satisfaction at additional per-step cost, yet how this trade-off behaves with scale is uncharacterized. We sweep model and dataset sizes under both MSE and AL training, and characterize how constraint violation changes with network size across grids. Both objectives improve as power laws, but at different rates: MSE is governed primarily by model capacity, while AL is balanced across both. Violation grows roughly twice as fast with network size under MSE as under AL. On matched hardware, AL reduces violation by nearly $30\times$ for an order of magnitude more training time, with negligible added memory. The training objective determines not only where a surrogate lands but how its quality evolves with scale.
cs.LG / 16 / 2609.16283
Differentially Private Semantic Plans for Aggregate Insight Generation
Abstract
\texttt{URANIA} provides end-to-end differential privacy (DP) for summaries of data-dependent clusters. However, its cluster--keyword release does not directly provide collection-wide aggregates for semantic concepts defined independently of the protected corpus. Records may express several concepts, records expressing the same concept may be assigned to different clusters, and cluster identities need not correspond across analyses. Consequently, cluster-level statistics do not directly provide comparable measurements of predefined concepts across collections or repeated analyses. We introduce \texttt{DP-SPIN}, a trusted-curator framework for aggregate measurement and summarization over semantic concepts fixed independently of the protected target records. Each record is mapped to a bounded sparse nonnegative vector over these concepts, whose sum forms a semantic sketch. A differentially private mechanism releases a semantic plan containing admitted concepts and noisy masses; normalized semantic-support values and support bins are obtained by post-processing. For user-level privacy, each user's aggregate contribution is clipped to a fixed bound. The language model receives only the plan and fixed decoding instructions, while a public verifier checks concept mentions, reported values, comparisons, and rank claims against the released plan. The final summary is differentially private by post-processing. We establish record- and user-level DP guarantees under add/drop and replacement adjacency. We evaluate \texttt{DP-SPIN} under record-level privacy on CFPB complaint narratives, Amazon All Beauty reviews, and Yelp restaurant reviews, and under user-level privacy on Amazon and Yelp. We compare \texttt{DP-SPIN} with non-private plan and summary references, DP keyword and category histogram baselines, and a \texttt{URANIA}-style baseline with a fixed public keyword vocabulary.
cs.LG / 17 / 2609.16288
Drift Field Net: Learning Ocean Lagrangian advection fields from in-situ and satellite observations
Abstract
The North Pacific Subtropical Gyre (NPSG) is a major accumulation zone for floating plastic debris, resulting from basin-scale convergent ocean circulation. Effective cleanup strategies in this region rely on accurate forecasts of Lagrangian particle drift. Here, we introduce Drift Field Net (DFN), a deep neural network that predicts ocean surface flow fields from operational satellite observations. DFN is trained using a novel two-stage strategy that combines pretraining on simulated data with Lagrangian fine-tuning based on an advection-consistent loss function. This physics-informed optimization directly improves the accuracy of particle trajectory predictions. We evaluate DFN against an operational physics-based forecasting system and demonstrate the potential of deep learning for ocean surface flow prediction. On in situ drifter trajectories, DFN reduces the mean positioning error by 20 km after a 7-day forecast compared with the operational model. Furthermore, Lagrangian fine-tuning with the proposed advection loss further reduces the positioning error by 10 km, highlighting the benefits of incorporating Lagrangian constraints into the training process.
cs.LG / 18 / 2609.16309
Agentic Search Spaces for Tabular Machine Learning
Abstract
Despite the rapid progress of LLM-based agents for planning, code generation, and debugging, their practical value for tabular machine learning remains underexplored. In this paper, we investigate a concrete use case: whether state-of-the-art agentic AI systems can design extended HPO search spaces for established tabular models that outperform the standard search spaces provided by the model authors. Specifically, we represent each tabular model as a modular pipeline covering preprocessing, embeddings, architecture, training, and inference. We then task the agent to propose candidate code implementations for each module and use a classical HPO algorithm to jointly optimize over these candidates and the model's default hyperparameters. Compared with the base HPO spaces, the expanded search spaces improve the performance of nearly every model family across a suite of 45 datasets, with average relative gains of 0.6%, rising to 2.0% on small-to-medium regression datasets. Notably, these gains come at no extra tuning cost: the enlarged spaces outperform the base under the same tuning and ensembling budgets. The gains transfer to the recent TabArena benchmark, where the agentic spaces improve the official Elo scores of four of the five model families and the two strongest agentic ensembles surpass the best AutoGluon ensemble of conventional models. Overall, our study suggests that LLM agents can provide practical value for tabular ML by expanding the design space.
cs.LG / 19 / 2609.16314
Robust Fault Detection in Mechanical Multimodal Time Series via Self-Supervised Cross-Modal Reconstruction
Abstract
Fault detection is essential in industrial systems, enabling early identification of abnormal behaviour and improving safety, reliability, and operational efficiency. Modern systems increasingly rely on heterogeneous sensing modalities that capture complementary aspects of the underlying physical process. However, existing data-driven anomaly detection methods often process each modality independently or use simple feature-level fusion, limiting their ability to exploit cross-modal relationships that characterize normal system behaviour. Their performance also commonly assumes similar training and deployment distributions, whereas real-world operation is affected by changing operating conditions, environmental influences, and system degradation that induce distribution shifts and reduce detection performance, especially in unseen regimes. In this work, we propose a multimodal anomaly detection framework based on cross-modal reconstruction of heterogeneous time-series sensor data. Rather than modeling each modality independently, the framework learns system dynamics by reconstructing each modality from the others, thereby exploiting complementary information across modalities. This integrates information across sensing channels without requiring explicit temporal alignment or identical sampling rates, while improving robustness to sensor noise, missing measurements, and modality-specific disturbances. To address distribution shifts during real-world deployment, anomalies are identified using cross-modal reconstruction error and an adaptive test-time thresholding mechanism that adjusts to changing operating conditions. Experiments on three industrial case studies show strong fault detection performance and substantially improved robustness under out-of-distribution conditions, with the largest gains observed in the most challenging operating regimes.
cs.LG / 20 / 2609.16317
Generative models for simulation based filtering: Formulations and Empirical Comparisons
Abstract
This letter presents a unified formulation and a controlled numerical comparison of generative-model approaches to the nonlinear filtering problem. Under this formulation the analysis step is realized by a transport of the forecast distribution to the posterior, the approaches differing only in how that transport is selected and learned. We derive three new filters, based on stochastic interpolants, their deterministic flow-matching limit, and Schrödinger bridges realized through forward--backward SDEs. We develop a two-stage tuning procedure that separates the training of the generative model from its online refinement. The resulting methods are compared against the optimal transport filter (OTF), the Knothe--Rosenblatt filter (KRF), the sequential importance resampling (SIR) particle filter and the ensemble Kalman filter (EnKF), in terms of accuracy, computational time, and sensitivity to ensemble size and state dimension. The results indicate that every generative filter resolves multimodal posteriors that the EnKF and SIR do not, that no single generative framework dominates, the preferred method being set by the available online budget and ensemble size, and that the filters differ in the regularity of the particle trajectories they produce.
cs.LG / 21 / 2609.16341
Channel-Informed Neural Network for Physical Layer Key Generation
Abstract
Physical-layer key generation (PKG) enables wireless devices to establish shared keys from reciprocal channel observations without directly exchanging the key. This capability is attractive for edge networks, where distributed and resource-constrained devices may require lightweight key establishment with limited access to centralized infrastructure. We introduce a channel-informed neural network for PKG that derives binary key features directly from received IQ measurements while explicitly grounding the learned representation in the underlying multipath channel. The proposed multi-task recurrent neural network jointly learns reciprocity-preserving binary features and an auxiliary channel estimate using a training objective that combines deep metric learning with channel-informed supervision. Structured channel sounding enables channel estimation from over-the-air measurements, while Sionna-RT ray tracing is used to augment training with additional propagation conditions. We evaluate the framework using indoor and outdoor software-defined-radio measurements collected on the POWDER radio testbed. Across all evaluated scenarios, the proposed model produces lower bit disagreement for reciprocal Alice-Bob observations than for Eve-related observations. Ray-traced data augmentation substantially improves key diversity, increasing the unique-key rate to 0.94, 0.99, and 0.99 across the indoor and two outdoor scenarios, respectively. Successfully reconciled channel-informed keys pass the selected NIST randomness tests prior to SHA-3 privacy amplification. The results demonstrate the potential of channel-informed representation learning for decentralized wireless key establishment while highlighting an important tradeoff between key diversity and reconciliation reliability.
cs.LG / 22 / 2609.16347
Multi-Label Proportion Learning for Sea-Ice Type Prediction
Abstract
Sea-ice type prediction is important for climate monitoring, maritime navigation, and decision-making in polar regions. The main source of label data for this task is the ice chart, produced manually by ice analysts who interpret satellite imagery to delineate ice zones into polygons. Although ice charts are valuable, their production is labor-intensive and expensive, motivating recent efforts to automate the process using deep learning. However, deep learning models require patch-level (or pixel-level) label data for training, while ice charts provide only polygon-level annotations. As a workaround, supervised approaches often create approximate patch-level labels from polygon-level ice chart labels by assigning each sample the dominant ice type of its parent polygon. This approach enables supervised training but creates an ill-posed learning problem with intrinsically approximate solution. In this paper, we redefine sea-ice type prediction as a weakly supervised multi-label proportion learning problem to be able to directly use the polygon-level ice chart labels and avoid unnecessary label approximation for improved prediction accuracy. To address this problem, we propose a two-module framework where first Multiple Instance Learning (MIL) is used for water--ice classification, and then a multi-label proportion learning (MLPL) is introduced for ice-type composition prediction. We further extend this framework with a multimodal model that integrates SAR imagery with AMSR2 brightness temperatures and ERA5 reanalysis data through modality-guided auxiliary regularization. Evaluated on the AI4Arctic dataset, the SAR-only model reduces MAE by 14.5\% and more than doubles mean ice-class F1 over the best supervised baseline. The multimodal model further reduces MAE by 21.5\% and raises mean F1 by 41.2\% over the SAR-only model, and by 52.7\% over the supervised multimodal baseline.
cs.LG / 23 / 2609.16350
Federated stochastic bilevel optimization with fully first-order gradients
Abstract
Federated stochastic bilevel optimization has been actively studied in recent years due to its widespread applications in machine learning. However, most existing federated stochastic bilevel optimization algorithms require the computation of second-order Hessian and Jacobian matrices, which leads to longer running times in practice. To address these challenges, we propose a novel federated stochastic variance-reduced bilevel gradient descent algorithm that relies solely on first-order oracles. Specifically, our approach does not require the computation of second-order Hessian and Jacobian matrices, significantly reducing running time. Furthermore, we introduce a novel learning rate mechanism, i.e., a constant single-timescale learning rate, to coordinate the update of different variables. We also present a new strategy to establish the convergence rate of our algorithm. Finally, the extensive experimental results confirm the efficacy of our proposed algorithm.
cs.LG / 24 / 2609.16369
Autonomous Droplet Navigation via Model-Based Reinforcement Learning
Abstract
Precise manipulation of liquid droplets underpins lab-on-a-chip platforms for diagnostics, chemical synthesis, and biological assays. Yet autonomous droplet transport through confined geometries of varying complexity remains an open challenge. Droplets exhibit contact-angle hysteresis, deformability, and capillary pinning, which make their response to actuation nonlinear and history dependent, that classical controllers and pre-programmed trajectories cannot cope in multi-turn environments. Here we demonstrate autonomous navigation of a liquid droplet through geometries of increasing complexity on a gravity driven (Labyrinth) platform using model-based reinforcement learning. A thin silicone oil film reduces contact-line pinning while two-axis tilt supplies the gravitational driving force, and an overhead camera tracks the droplet in real time. An offline-trained policy discovers effective tilt strategies from limited physical interaction data, without simulation or analytical droplet models. The system operates under partial observability, as oil-film thickness, instantaneous contact angle, and droplet deformation state remain hidden from the controller. Despite these challenges, the learned policy achieves reliable navigation across straight, right-angle, and curved-arc paths, including outside-corner geometries. We further demonstrate that a policy trained on a simpler geometry transfers to complex ones, succeeding zero-shot on right-angle and staircase paths and reaching full success on a curved arc with a fifth of the training data. The findings suggest promising avenues for enabling droplet based microfluidic systems to serve as intelligent chemical laboratories.
cs.LG / 25 / 2609.16373
Certified Uncertainty Propagation in One-Shot Federated Bayesian Models via Posterior Event Transport
Abstract
Probabilistic certification of Bayesian neural networks lower-bounds the posterior probability that a model satisfies a verifier-defined safety property. In one-shot federated Bayesian learning, however, the deployed model is obtained by aggregating parameters drawn from client-specific posterior distributions, so local certificates do not directly guarantee safety of the aggregated model. This paper develops a deployment-consistent certification framework by propagating local posterior events through the deployment aggregation rule, with an exact geometric characterization for Federated Averaging (FedAvg). Each client constructs disjoint hyper-rectangular regions in parameter space and computes their probability masses. The server forms Cartesian products of these regions, maps them through the deployment rule, and retains a product event only when its aggregation image is verified to satisfy the safety property. Under independent client posteriors, each product-event probability factorizes into local masses, and summing verified disjoint events yields a lower bound on safety probability of the deployed model. For FedAvg with nonnegative aggregation coefficients, the image of a Cartesian product of axis-aligned hyper-rectangles is exactly a weighted hyper-rectangle, introducing no set over-approximation. We distinguish the proposed transported-event certificate from direct certification under posterior distributions induced by FedAvg and Product-of-Gaussians aggregation. Experiments on MNIST and Fashion-MNIST under label-Dirichlet heterogeneity show that the transported FedAvg certificate ranges from 22.51% to 46.89%, while direct global certificates range from 72.05% to 91.39%. Results show that predictive accuracy and certifiable safety do not necessarily follow the same trend, and that global posterior constructions can exhibit distinct certification behavior across architectures.
cs.LG / 26 / 2609.16380
Bounded Adjustment with Reliability-Guided Embedding for Imbalanced Learning with Noisy Labels
Abstract
Class-balanced learning and label noise create a coupled failure mode: frequency correction prevents majority classes from dominating the decision rule, but can amplify incorrectly labeled minority examples. We introduce BARGE (Bounded Adjustment with Reliability-Guided Embeddings), a single-stage objective combining a bounded, prior-adjusted density-power score with reliability-guided angular geometry. Its classification score is strictly proper in the adjusted probability space and recovers balanced Bayes ordering under clean supervision and the true class prior. Under label contamination, its finite range bounds classification-risk perturbation at a fixed predictor, while its logit gradient redescends when the model confidently contradicts the supplied label. The adjusted target probability also weights class-equal feature compactness, and a one-sided separation term discourages aligned class directions. BARGE requires neither a noise rate nor a transition matrix, uses one network, and leaves inference unchanged. We evaluate it on CIFAR-10, CIFAR-100, and Tiny ImageNet under long-tail and step imbalance, clean labels, and 20% and 40% random incorrect-label replacement. Across 12 clean settings, BARGE ranks second overall and attains the lowest error in four. Under corruption, it achieves the lowest mean balanced error in all six dataset-corruption settings, reducing the six-setting average from 72.32% for the strongest competitor to 70.00%. It also obtains the highest macro-F1 and macro-AUPRC in every corrupted-label setting. Ablations show that class-equal angular compactness improves on the bounded score alone. These results support bounded predictive influence and reliability-guided geometry as complementary mechanisms for imbalanced learning with uncertain labels.
cs.LG / 27 / 2609.16382
Attention Mean Fields Predict Average Representation Dynamics and Reveal Context-Specific Computation
Abstract
A language model's representation geometry is not predetermined; it evolves as the model runs. A faithful account of that geometry must capture that dynamic process, and so cannot be based solely on model-independent statistics such as co-occurrence. Here we introduce a mean-field analysis of attention. The average attention from one token to another defines a kernel that carries representations layer to layer and can be iterated through the network to model how the geometry is transformed. We condition this average two ways. Conditioned on a whole corpus, the kernel predicts the average-case evolution of representation geometry. Conditioned instead on a single context, it predicts the expected geometry for that context. A head's departure from that prediction, its \emph{mean-field deviation}, isolates the context-specific computation that the mean field misses. Under the corpus-conditional reading, the kernel yields an open-loop model: from the input embeddings and the frozen weights alone, we can iterate the kernel and the model's own MLPs over token representations, never consulting a measured deviation at any layer. The resulting prediction is highly accurate. In early training the model and its corpus mean field are indistinguishable. Replace every attention head with its mean field, and the substitution leaves the loss on real text unchanged. Around the onset of induction, the two diverge, and the gap widens as representations become contextualized. Under the context-conditional reading, deviation from the mean field is a task-agnostic measure of context-specific computation. The residual decomposes additively into unusual attention routing and contextualization of the transported values. Across controlled induction and few-shot settings, greater deviation tracks greater reliance on in-context information.
cs.LG / 28 / 2609.16415
How Good Are Time-Series Foundation Models for Pedestrian Crowd Count Forecasting? A Cross-Dataset Comparative Study
Abstract
Pedestrian-count forecasting supports pedestrian-oriented Intelligent Transportation Systems (ITS), including crowd monitoring, pedestrian-traffic staffing and routing, and proactive risk mitigation during surges. Recent time-series foundation models (FMs) report strong zero-shot accuracy on heterogeneous forecasting benchmarks, but it remains unclear whether these gains transfer reliably to pedestrian sensing deployments. We benchmark seven univariate forecasting approaches spanning four paradigms: Seasonal Naive, gradient-boosted trees (LightGBM, CatBoost), deep learning models (N-HiTS, PatchTST), and two pretrained FMs (TimesFM, Chronos-2). Experiments cover two complementary regimes: (i) a five-day special event dataset SAIL2025 at 3-minute resolution with limited in-domain history; and (ii) Melbourne pedestrian sensors as a multi-year hourly dataset (2010--2017) with strong seasonality. We compare the MAE and RMSE results per sensor across datasets and multiple forecast horizons. Results show three consistent findings. First, with limited historical data, Seasonal Naive remains a strong baseline for long-horizon forecasting on high-volume sensors, while trained models can degrade when the next day differs substantially from prior days. Second, boosted trees can be competitive on lower-volume sensors but exhibit higher sensitivity on high-volume sensors under event-driven shift. Third, FMs excel in the seasonal and data-rich regime under long-context configuration. The findings highlight the importance of choosing pedestrian forecasting models based on both the underlying data conditions and the forecasting horizon.
cs.LG / 29 / 2609.16446
Adaptive Bayesian Partner Selection for Federated Clinical Centers
Abstract
Federated learning (FL) in healthcare faces pronounced heterogeneity and temporal concept drift across clinical centers, where evolving patient populations and care practices shift data distributions. Existing approaches rely on persistent global communication, incurring substantial bandwidth overhead while risking negative transfer from poorly aligned peers. We propose Adaptive Bayesian Partner Selection (ABPS), a peer-to-peer framework that governs who collaborates, when, and at what cost. Each center maintains a Beta-Bernoulli posterior over prospective peers' Shapley marginal utility, ranks candidates with an Upper Confidence Bound (UCB) criterion, and forms collaborations through a lightweight propose-reject mechanism, with the option to abstain from communication when no mutually beneficial partner exists. The framework admits a stochastic decision interpretation, yielding finite-sample concentration guarantees and O(kappa log T) regret in partner selection, along with conditions under which intentional isolation is optimal under negative transfer. Lightweight extensions (head personalization, bfloat16 quantized communication, and a tunable active-set size) further improve efficiency, and a goal-aware metadata filter enables institution-specific collaboration strategies. On binary in-hospital mortality prediction over the first 24 hours of an ICU stay, with 230 non-IID clinical centers drawn from MIMIC-IV, the full ABPS-X variant matches the strongest federated baseline (FedDyn, AUROC 0.758) at 0.09x the communication cost of FedAvg, with reduced variability. A diversity-driven configuration activates intentional isolation for a substantial fraction of centers. These results show that adaptive, utility-aware collaboration reduces communication without sacrificing accuracy when centers are numerous and small, offering a scalable paradigm for healthcare FL.
cs.LG / 30 / 2609.16459
OPD-Aha: From Linguistic Momentum to Visual Reflection in Multimodal On-Policy Distillation
Abstract
Privileged on-policy distillation improves multimodal reasoning by allowing a teacher to evaluate student trajectories using rich, training-only visual evidence. Both models score these trajectories while conditioning on the same student-generated prefix. When a student misinterprets an image early in a response, this accumulating erroneous rationale eventually pulls the teacher away from its visual evidence. The teacher and student converge on the same hallucination, causing standard cross-model supervision to collapse precisely where correction is most needed. We find that the teacher's visual corrective preference is not lost under this misleading agreement. Comparing the predictions of the identical teacher given the real image and a visual null reveals that the privileged evidence still pushes the model toward the correct interpretation. We introduce OPD-Aha, which reconstructs the distillation target directly from this isolated visual preference rather than relying on the fragile teacher-student discrepancy. This reconstructed target aggressively suppresses continuations that contradict the image. Trained with this objective, students learn to naturally interrupt their own flawed reasoning with reflection tokens such as wait and actually. After reflection, subsequent generation relies less on the accumulated erroneous text and more on the visual evidence. Correcting these trajectories mid-generation fundamentally alters the reasoning process, yielding broad and consistent improvements across diverse fine-grained perception and complex multimodal reasoning benchmarks. Our code and models are available at https://github.com/Echochef/OPD-Aha.
cs.LG / 31 / 2609.16472
Online Gradient Computation for Warping Gaussian Process Transformations
Abstract
Warped Gaussian processes (GPs) handle non-Gaussian observations by mapping them into a latent standard GP via a parametric transformation called warping. Existing streaming variants, however, either optimize the warping parameters periodically or sacrifice analytical tractability for a higher model capacity. To bridge this gap, we show that the gradient of the instantaneous negative log-likelihood of a warped GP admits an exact recursive computation. Based on this result, we propose a novel online method for warped GPs that jointly updates the latent GP moments and optimizes the warping parameters.
cs.LG / 32 / 2609.16489
Decoder Design Matters for ECG Delineation
Abstract
Electrocardiogram (ECG) delineation identifies the boundaries of P waves, QRS complexes, and T waves, providing structural annotations that can guide AI models in learning to interpret ECGs. However, training accurate delineation models requires manual annotations that are scarce and time-consuming to obtain. Recent work addresses this limitation through semi-supervised learning (SSL), but the design of the architecture, particularly the decoder, has received less attention. To this end, we propose R-U-Net, an ECG delineation model that pairs a ResNet-18 encoder with a U-Net decoder. On SemiSegECG, R-U-Net outperforms the strongest evaluated ResNet-18 + fully convolutional network (FCN) head baseline in each of the 16 in-domain settings by 3.3-13.0 mIoU and achieves 82.6 mIoU in the cross-domain setting, an improvement of 8.1 mIoU. Controlled ablations show that decoder design contributes more to performance gains than the evaluated SSL methods, motivating further exploration of architectures for ECG delineation. All code is open-source at github.com/ELM-Research/ECG-Delineation.
cs.LG / 33 / 2609.16500
High-Performance Tensor Formulation of the Viterbi Algorithm for Hidden Semi-Markov Models
Abstract
Hidden Semi-Markov Models (HSMMs) are fundamental probabilistic models widely adopted across diverse domains, from computational biology to finance and signal processing. The Viterbi algorithm decodes the most likely state sequence given an HSMM and can be applied iteratively for ab initio model learning. However, existing Viterbi implementations remain sequential, and GPU-accelerated solutions are entirely absent, making HSMM decoding impractical for large-scale workloads. We present a tensor-based formulation of the Viterbi algorithm for HSMMs, restructuring the inner loops into tensor operations that naturally map onto SIMD units and massively parallel architectures. Building on this formulation, we provide optimized implementations spanning single- and multi-core CPUs, and, for the first time, GPU. Experimental evaluation demonstrates speedups of up to 14x on a single core, over 200x with multi-core, and over 570x on GPU over the state-of-the-art sequential baseline, establishing a new performance baseline for large-scale HSMM decoding.
cs.LG / 34 / 2609.16537
What Does Layer-Importance Reveal About Transformers and State-Space Models?
Abstract
Transformers and state-space models (SSMs) are the two dominant families of sequence models, and a central open question is how far the analytical knowledge built for transformers transfers to SSMs. We address this through the lens of layer importance which underpins compression, selective fine-tuning, and interpretability across both families. We decompose layer importance into two distinct notions. \emph{Necessity} captures how much the pretrained model depends on a layer's existing contribution, measured by the loss increase from bypassing it. \emph{Plasticity} captures where the model absorbs new information during fine-tuning, measured by the magnitude of task-specific weight updates. Our analysis reveals that the two families behave fundamentally differently: in every evaluated residual transformer up to $14$B parameters, Necessity and Plasticity anti-align across depth, whereas in the evaluated Mamba-style SSMs they point to overlapping regions. The sign of this alignment also predicts downstream adaptation behavior. In the evaluated transformers, concentrating updates in the most plastic layers increases catastrophic forgetting, while this tier-dependent effect disappears in the evaluated Mamba-style SSMs.
cs.LG / 35 / 2609.16540
On the Importance of Gating: Memorization vs. In-Context Learning in State Space Models
Abstract
State Space Models (SSMs) have emerged as a compelling alternative to Transformers, enabling sequence modeling with constant memory and linear compute. Although SSMs exhibit reasonable performance and favorable computational characteristics, they continue to lag behind Transformers on tasks that require in-context learning and precise retrieval, slowing their adoption for large-scale language modeling. In this work, we demonstrate that both the success and failure of SSMs in these domains can be explained by studying the role of the gating mechanism, a prevalent component in modern recurrent networks. Specifically, we show through theory and experiments that this gating mechanism causes SSMs to first learn an in-weights "memorization" solution, while delaying, or even preventing, convergence to a correct in-context learning solution. Importantly, this happens even in cases where there are no fundamental limitations due to the architecture or its memory capacity. On the other hand, we find that gating is often beneficial for improving generalization to long sequence lengths. Our results illuminate the crucial role of the gating mechanism in shaping both the training dynamics and generalization of SSMs, and provide a basis for understanding and improving linear-time models.
cs.LG / 36 / 2609.16573
AsyncCouple-Flow: Asynchronous Cross-Modal Coupling and Flow Matching for Spatio-Temporal Forecasting
Abstract
Multi-modal spatio-temporal forecasting (MM-STF) supports weather nowcasting, traffic prediction, and earth-system modeling by combining heterogeneous sources such as physical fields, satellite imagery, and in-situ sensors. Three obstacles persist: (i) modalities have different spatio-temporal sampling rates, forcing lossy interpolation onto a unified grid; (ii) modalities are frequently missing at deployment due to sensor outages or revisit gaps, while most methods train with full availability; and (iii) autoregressive decoders accumulate errors over long horizons, amplified by multi-modal conditioning. We propose AsyncCouple-Flow to address these issues jointly. A Modality-Aware Token Sparsification (MATS) module performs scale-aware tokenization and uses a shared importance scorer to select top-k tokens per timestep, producing equal-length sequences. An Asynchronous Cross-Modal Coupling Graph (ACCG) replaces fixed cross-attention with a learnable graph whose edges encode time offsets, semantic similarity, and modality-specific physical priors, enabling fusion under arbitrary asynchrony and missingness. A Flow-Matching Forecasting Head models multi-step prediction as a conditional ODE, trained with stochastic modality dropout and integrated jointly to avoid autoregressive drift. Experiments on ERA5+GOES+ISD weather forecasting and PEMS-BAY traffic prediction with multi-source side information show that AsyncCouple-Flow outperforms state-of-the-art baselines and remains robust with up to two missing modalities. The code will be released upon acceptance.
cs.LG / 37 / 2609.16606
A Weighted Kernel Method for Approximation that Adapts to Learned Multivariable Structure
Abstract
Approximating the input-output behavior of a multivariable black-box function from limited data is challenging when blind to the importance of its inputs and their interactions. We introduce total sensitivity kernels (TSKs), a method based on families of weighted ANOVA kernels that learn and adapt to this multivariable structure. TSKs parameterize the weights on each multivariable component of the target function by factors for each input. We propose learning these factors directly from function evaluations by selecting the reproducing kernel Hilbert space (RKHS) in which the target function has minimum norm. Under suitable conditions, we show that this norm-minimization problem admits a unique solution, and we establish consistency of a finite-data formulation based on minimum-norm interpolation. The learned TSK factors characterize the participation of individual inputs across interactions and main effects, providing a kernel-dependent notion of input sensitivity related to total Sobol indices. Numerical experiments demonstrate that adapting the kernel to learned multivariable structure can substantially improve approximation accuracy over a standard product kernel.
cs.LG / 38 / 2609.16617
Divergence Timing and Cumulative Disagreement under KV-Cache Eviction
Abstract
KV-cache eviction perturbs the conditional token distributions governing autoregressive generation. We investigate how first-divergence timing and subsequent token mismatch determine cumulative disagreement. We derive an exact decomposition under a specified stepwise maximal coupling: the expected mismatch fraction equals a first-mismatch contribution plus post-divergence exposure multiplied by its mismatch rate. An explicit construction over unrestricted autoregressive kernel pairs realizes the sharp interval of risks compatible with a finite divergence-aligned observation window. Residual-branch conditional Monte Carlo provides unbiased joint estimates of occurrence, occupation, and window/tail contributions, with per-replicate variance dominance for total token loss. Complete trajectories from Meta-Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct show that SnapKV at 50% retention enters divergence later and less often than SnapKV-512 or recent-token retention with the same 50% prompt-cache budget, while post-divergence total variation (TV) remains high. In an exploratory analysis of 288 documents, post-divergence exposure accounts for 85-90% of four aggregate mismatch gaps. On 288 independent documents at 90% retention, prespecified comparisons show higher branch-aligned TV in the late than in the early window in both models.
cs.LG / 39 / 2609.16621
Stable by Construction: Variational Latent Markov Operators for Long-Horizon PDE Prediction
Abstract
Neural PDE solvers provide efficient surrogates for time-dependent physical systems, but autoregressive prediction over long horizons remains challenging because local errors can induce distribution shift and accumulate under recursive deployment. We develop a variational approach to this problem by introducing latent Markov dynamics in which physical states are represented by latent distributions and evolved through probabilistic transitions. The framework is formulated directly on function spaces and specialized to functional Gaussian models, where structured latent perturbations induce a spectral geometry and variational transition alignment regularizes the learned dynamics. We further analyze how these mechanisms affect autoregressive error propagation, providing a theoretical connection between variational training and long-horizon prediction. We instantiate the framework as the Variational Autoencoding Markov Operator (VAMO), which combines spatially resolved latent fields, structured Gaussian perturbations, and a neural-operator transition. Empirically, we demonstrate the effectiveness of VAMO on several fluid-dynamics benchmarks with prediction horizons extending substantially beyond those represented during training, where it consistently reduces error accumulation and improves rollout stability over several deterministic and noise-injection baselines. Overall, these results highlight variational modeling as a complementary approach to robust long-horizon neural PDE dynamics.
cs.LG / 40 / 2609.16665
Right Direction, Wrong Step: Geometric Analysis of Finite-Step Failure in Looped Transformers
Abstract
Looped Transformers offer a parameter-efficient route to test-time scaling by reusing shared layers for iterative latent reasoning. However, additional iterations can reduce support for a reference answer, leaving unclear whether an update's direction is locally unhelpful or its full displacement moves too far. We study this distinction by analysing reference utility, which measures this support, along the model's own update direction, varying the fraction of the proposed displacement supplied to the readout. This reveals finite-step failures in which a locally improving direction produces a harmful full update. A pathwise curvature decomposition characterises how initial progress is lost, while a local quadratic model predicts full-step gains and useful step scales. Bounds based on accumulated curvature variation characterise the approximation error of these predictions. Experiments across two model families reveal this separation on mathematical and commonsense tasks. A fixed quarter step produces positive gains in reference utility for 72.2--83.2% of selected failures across four settings. These findings identify a mismatch between update direction and step scale as a mechanism of lost progress, explaining how some harmful updates retain useful computation.
cs.LG / 41 / 2609.16710
Continuous-Time Machine Learning: A Unified Mathematical Perspective
Abstract
Continuous-time (CT) machine learning has emerged as a principled framework for modeling temporal dynamics as a continuous process, particularly when observations are sampled at arbitrary time points or span long-range horizons. However, major branches of CT machine learning have matured in separate research communities, leaving their mathematical relationships and design trade-offs insufficiently characterized. In this survey, we develop a unified, concept-driven view of major CT machine learning branches through a taxonomy that organizes families according to their underlying base mathematical formulations. We present a canonical mathematical formulation that relates these families through different architectural choices of vector-field parameterization, stochasticity, memory mechanisms, and discretization. We compare training algorithms, optimization strategies, and failure modes, highlighting the trade-offs across families. We further provide a comparative analysis of theoretical computational complexity alongside an illustrative architecture-controlled benchmark analysis on representative architectures from each family. We also review software ecosystems supporting their implementation. Finally, we identify open challenges in approximation theory, training stability, hardware-efficient implementations, benchmarking, foundation models, and scientific machine learning, and discuss an agenda for future research.
cs.LG / 42 / 2609.16744
A Systematic Evaluation of Machine Learning Methods for Fault Detection and Line Identification in Electrical Power Grids
Abstract
The integration of renewable energy sources into the electrical grid introduces complex challenges in fault detection and coordination of grid recovery mechanisms. Traditional relay protection systems, which operate based on static rules and predefined thresholds, are inadequate for addressing these challenges, particularly in detecting and isolating faults such as short circuits. Consequently, the conventional methodologies applied to electrical network protection frequently fail to achieve optimal performance in fault detection, especially in terms of adherence to safety standards and the selective limitation of damage. Recent research indicates that machine learning (ML)-based approaches can effectively tackle these issues; however, variations in grid configurations and analysis windows have impeded consistent comparative assessments. In this study, we assess the efficacy of various ML models in detecting electrical faults and pinpointing defective transmission lines within a 10 ms measurement interval - a critical time-frame for real-time operational viability, for the first time. The most effective model attained an F1 score of 0.991 +/- 0.018 and demonstrated a processing time of 0.342ms +/- 0.509ms.
cs.LG / 43 / 2609.16754
TAME: Token Attribution and Masking for Emergent misalignment
Abstract
Fine-tuning an aligned language model on narrow, flawed data can induce harmful behavior far outside the training domain, known as emergent misalignment (EM). Prior work has localized EM in model weights, activations, and training documents, but it remains unclear which training tokens carry the relevant fine-tuning signal. We introduce TAME (Token Attribution and Masking for Emergent Misalignment), a three-stage framework: token attribution scores how strongly the fine-tuning update raises each response token's likelihood, using forward passes through a released LoRA adapter; signal characterization finds patterns among high-attribution tokens; and causal validation tests them by attribution-guided loss masking. On released EM organisms and a 6,849-example medical-advice split, attribution is concentrated (the top 5% of tokens hold 32% of the mass) and, in Llama, depleted for medical vocabulary but enriched for a register of unwarranted certainty, even after controlling for token rarity. Masking high-attribution tokens during fresh fine-tuning cuts EM by 23x in Llama and 36x in Qwen, with the perplexity cost concentrated on the targeted register rather than on medical content; an equal random mask leaves EM unchanged. In Llama, the attribution pattern suggests that EM-relevant signal lies more in how confidently flawed content is expressed than in its domain vocabulary; the causal masking effect itself holds across both model families.
cs.LG / 44 / 2609.16788
Noise2Noise Revisited: Training Pair Distributions Dominate Loss Choice in Self-Supervised Denoising
Abstract
Noise2Noise (N2N) trains denoisers on pairs of independently corrupted observations, eliminating clean references. We stress-test two natural conjectures about why the L1 loss outperforms L2 here. First, the hypothesis that the L1 loss confers robustness via parameter sparsity confuses the loss with Lasso regularization: an explicit Lasso penalty produces the predicted sparsity yet fails to reproduce L1's cross-noise behavior, while L1- and L2-trained weight distributions are indistinguishable. Second, the population optima of the two losses coincide exactly for symmetric signal posteriors and nearly so for concentrated ones. Measured differences are therefore dominated by optimization dynamics (bounded-influence gradients), which we probe with gradient statistics and contaminated-target training. On Kodak24 with five synthetic noise families, the L1 loss holds a statistically significant edge over L2, below 1 dB PSNR, holding across three seeds on 13 of the 14 noise columns. On real camera noise the loss is not the decisive variable in distribution: on official SIDD validation blocks, synthetic-Gaussian-trained N2N models gain only 0.8 to 3.7 dB over the noisy input regardless of loss, while retraining on SIDD's own noisy pairs, never reading ground truth, gains 9.4 to 11.0 dB, far ahead of BM3D. All metrics are on raw network outputs, and the study makes no leaderboard claim. The training pair distribution, not the loss, carries the inductive bias. That design rule applies wherever clean references are unobtainable, from microscopy to industrial inspection sensors.
cs.LG / 45 / 2609.16804
SOTER: A Generative Time-Series Foundation Model for Wearable Human Physiological Signals
Abstract
Time-series foundation models have demonstrated strong cross-domain transfer, yet their common architectural assumptions remain poorly aligned with wearable physiological signals, which are multichannel, irregularly sampled, noisy, and governed by coupled continuous-time dynamics spanning distinct spectral scales. We present SOTER, a generative foundation model for wearable physiological time series that unifies cross-channel coupling, spectrum-guided expert specialization, and continuous-time latent evolution within a single pre-training framework. SOTER combines a spatial feature-aware backbone that models inter-signal dependencies, a power spectral density (PSD)-guided mixture-of-experts layer that routes representations to experts associated with fixed spectral bands through an inspectable, non-learned rule, and a neural controlled differential equation decoder that supports prediction and imputation at arbitrary timestamps. We pre-train SOTER on 226 billion time points from five public physiological datasets and evaluate the same pre-trained model across out-of-distribution zero-shot forecasting, frozen-encoder linear-probe classification, and continuous-time imputation on wearable benchmarks. SOTER achieves the best RMSE on 4 of 6 datasets and the best MAE on 5 of 6 in zero-shot forecasting, the highest average Macro-AUROC in classification, and the lowest imputation error on all six datasets at 75% missingness. It further remains robust to additive acquisition noise, matching or surpassing baselines evaluated on clean inputs even under the strongest corruption. These results indicate that domain-specialized foundation models for wearable physiology benefit from jointly modeling channel structure, spectral scale, and continuous-time dynamics.
cs.LG / 46 / 2609.16805
Geometry of learning dynamics: Gradient descent versus natural gradient on the ridge of optimization
Abstract
High-capacity associative memories based on Kernel Logistic Regression (KLR) exhibit a "Ridge of Optimization" characterized by extreme stability and a highly skewed weight spectrum. However, the dynamical process by which learning converges to this critical regime has remained unclear. This paper provides a geometric analysis of the learning trajectories on the statistical manifold of a KLR-trained Hopfield network. By comparing the paths of Gradient Descent (GD) and Natural Gradient Descent (NGD), we elucidate the mechanisms governing the optimization process. Our analysis reveals that learning on the Ridge proceeds in two distinct phases. We show that the extreme curvature of the Ridge causes standard GD to follow a highly oscillatory, non-geodesic path. In stark contrast, NGD explicitly corrects for this geometry, following the ideal geodesic path and completely overcoming the instabilities faced by GD. We demonstrate experimentally that NGD not only converges significantly faster but also achieves a solution with superior generalization performance. These results establish that the highly structured geometry of the Ridge is optimally suited for information-geometric optimization, providing a new perspective on the interplay between learning dynamics and emergent representation geometry.
cs.LG / 47 / 2609.16816
ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals
Abstract
Language model-generated rubrics are increasingly used as reward signals for rubric-based reinforcement learning, LLM-as-a-judge evaluation, and automated grading. Such rubrics are reliable only if they reward honest answers over adversarial answers optimized to exploit them. Yet their robustness to such optimization remains poorly understood. We isolate the hardest regime: impossible tasks, where the prompt pressures the model toward an unsupported conclusion, so the only honest response is to acknowledge the impossibility. We introduce ImpossibleRubrics, a benchmark of 169 impossible tasks spanning six impossibility categories, each paired with a verifiable oracle certificate specifying what an honest answer may and may not claim, together with 48 answerable controls. Rather than providing fixed rubrics, ImpossibleRubrics provides task environments and certificates, allowing rubrics to be generated downstream and then adversarially tested for whether they reward certificate-violating answers. Eleven generators are exploited 8--26% of the time on the unbiased 150-of-169 environment cut; on a deliberately selected stress cut the strongest generator we measured is still exploited 36% while a certificate-faithful rubric is exploited 0%, so what we measure is a rubric-quality gap, not task impossibility. One result runs against intuition. A single generic rubric ("be decisive, penalize hedging") used unchanged for every task is exploited 64% of the time, and seven of the eleven generators are exploited more often than that while writing a rubric tailored to each one. The tailored criteria appear to tell an attacker which claim to fabricate. The problem is not that rubrics are vague; it is that they are specific about the wrong things.
cs.LG / 48 / 2609.16823
LCAP: Population-Informed Latent Chip Adaptation from Few Output Probes for Photonic Neural Networks
Abstract
Photonic neural networks (PNNs) offer efficient analog inference, but parameters optimized under ideal device models can degrade after fabrication, creating a persistent simulation-to-hardware (sim-to-real) gap. When many identically designed chips are deployed, calibrating each device from scratch compounds this cost. We propose Latent Chip Adaptation from Probes (LCAP), a population-informed framework that decomposes hardware adaptation into a transferable population correction and probe-inferred latent personalization. LCAP first learns a shared correction from 80 historical chips, then extracts a low-dimensional correction space from device-specific refinements. At deployment, 32 fixed unlabeled output probes infer an unseen chip's latent correction coordinates, enabling feed-forward personalization without target-device optimization. On a three-layer 64-mode MZI simulator with phase variation, beam-splitter errors, quantization, and crosstalk, accuracy improves from 80.4147% under direct deployment to 92.6860% after shared calibration and 93.3617% with LCAP. LCAP improves 27/30 unseen chips and raises worst-device accuracy from 89.18% to 90.54%.
cs.LG / 49 / 2609.16824
Adapting to Decision-Relevant Non-Stationarity in Decentralized Heterogeneous Bandits
Abstract
Decentralized bandit systems often contain heterogeneous agents: rewards can change at individual agents even when the best action for the network stays the same. These local changes may cancel when rewards are averaged across agents, so the number of local changes $\Stloc$ can be much larger than the number of changes in the best common arm $\Stdec$. We introduce Decision-Relevant Fresh Comparison (DRFC), which uses new, balanced samples from all agents to compare arms at the network level and switches only when fresh global evidence indicates that the common best arm has changed. We prove a high-probability dynamic regret bound with no adaptation term depending on $\Stloc$, and show that every algorithm must still pay for identifying genuine decision switches and propagating them through the communication graph. Under a distinct time-average benchmark, an anytime-valid sliding-window extension handles gradual drift; experiments on synthetic, semi-real, and MovieLens-1M replays show that DRFC ignores decision-irrelevant local changes while the extension avoids false switches.
cs.LG / 50 / 2609.16827
Information Geometric Self-Organization at the Edge of Stability in High-Capacity Kernel Associative Memories
Abstract
High-capacity associative memories based on Kernel Logistic Regression (KLR) exhibit exceptional storage capabilities and robustness. Previous empirical studies identified a hyperparameter regime, the "Ridge of Optimization," where attractor stability is maximized. However, the geometric nature of this regime and the optimization dynamics required to reach it have remained unclear. In this paper, we investigate the static geometry of the parameter space and the learning trajectory of Gradient Descent (GD) in KLR-trained Hopfield networks. Using the eigenvalue spectrum of the Hessian, we reveal that the Ridge corresponds to a phase boundary located adjacent to a rank-1 spectral collapse, acting as a geometric singularity where the principal curvature is massively amplified. Furthermore, we demonstrate that the learning dynamics exhibit a transient self-stabilizing behavior driven by the Edge of Stability (EoS) phenomenon. Rather than seeking flat regions, the network parameters are driven toward a state where the local curvature dynamically equilibrates near the stability limit dictated by the learning rate, allowing the optimization to survive the initial instability. We provide analytical derivations for both the rank-1 asymptotic collapse and the dynamic feedback loop governing this equilibration. These findings suggest that optimal, high-capacity memory representations are not formed in flat minima, but are dynamically sculpted at the highly curved boundaries of geometric singularities.
cs.LG / 51 / 2609.16925
HyCoSeq: Contextual Hyperbolic Representation Learning for Genomic Sequences
Abstract
Hyperbolic geometry provides a natural inductive bias for genomic representation learning, but existing hyperbolic genomic models primarily use Lorentz convolutions to learn local sequence representations, while their residual pathways do not directly aggregate full Lorentz representations. We propose HyCoSeq, a contextual hyperbolic representation learning framework for genomic sequences. HyCoSeq incorporates weighted Lorentzian residual aggregation into multi-curvature Lorentz encoding, allowing full Lorentz representations to participate directly in geometry-consistent local aggregation. It further introduces a bidirectional long short-term memory network that integrates information from both sequence directions to learn contextual relationships among local representations at different positions within a genomic sequence, thereby extending local hyperbolic convolutional encoding to sequence-level contextualized representations. Extensive experiments across diverse genomic tasks show that HyCoSeq outperforms existing hyperbolic baselines and, without large-scale genomic pretraining, achieves competitive performance against substantially larger pretrained DNA language models.
cs.LG / 52 / 2609.16927
Verbalizing Subliminal Learning Effects Using Text Optimization
Abstract
Subliminal learning is a phenomenon in which a distillation dataset transmits traits from the teacher model that are not legibly encoded in the dataset itself. This introduces a new challenge for model development and creates new risks from data poisoning. In this work, we use text optimization to detect subliminal learning effects and describe them as legible prompts. Subliminal learning from a prompted teacher motivates our approach. We observe that this is a special case of context distillation and leverage this observation to show that, in theory, the prompted subliminal learning dataset identifies the teacher's prompt. We reduce recovering this prompt to a text optimization problem and present a method to approximately solve it. Our method, SALVE (Search-Aided Latent Verbalization), optimizes a soft prompt, queries the same model to verbalize it as text, and uses beam search to make the verbalization reliable. In the standard subliminal learning setting, SALVE reliably recovers legible prompts that name the teacher's trait, while common text optimization methods fail to do so. In addition, we find that there are settings in which SALVE recovers the teacher's trait from a dataset even when subliminal learning fails, but that modifying student training to improve context distillation can create subliminal learning effects. We lastly show that SALVE detects subliminal learning effects in three additional settings: (1) mixtures of subliminal learning data and unrelated data, (2) data generated when the teacher is biased via activation steering, and (3) subsets of real preference data selected via Logit-Linear Selection. Overall, our results deepen our understanding of subliminal learning and present SALVE as a method to proactively detect subliminal learning effects.
cs.LG / 53 / 2609.16930
Repurposing Deep Limit Order Book Forecasting for Scenario-Conditioned Market Impact Modeling
Abstract
Deep Limit Order Book forecasting models capture nonlinear market dynamics, but their ability to quantify the effects of counterfactual order book messages has not been systematically validated. We introduce a model-agnostic framework that compares a trained forecaster's predictive distributions before and after injecting mechanically valid counterfactual messages, defining short-horizon model-implied market impact. A Transformer-based forecaster recovered scenario rankings with a Spearman correlation of 0.99 and 97.2% directional agreement with realized historical outcomes among non-neutral scenarios. Observation-level analysis further showed that estimated impacts captured incremental sequence-dependent variation beyond scenario identity and the pre-event forecast. These results provide evidence that pretrained Limit Order Book forecasters can be repurposed for scenario-conditioned response modeling without retraining.
cs.LG / 54 / 2609.16933
When Confidence Signals Disagree: Local and Global Confidence in Autoregressive Language Models
Abstract
Modern predictive systems expose multiple quantities that are commonly interpreted as measures of confidence. However, these quantities can summarize different aspects of the predictive process. This distinction matters when confidence is used to evaluate reliability or inform downstream oversight and control. We investigate whether different confidence readouts are empirically interchangeable in an autoregressive language model by comparing local confidence, defined from the probability of the greedy-selected answer token, with global confidence, defined from modal-answer frequency under repeated sampling. Across MMLU and ARC Challenge, the two signals are weakly correlated and differ substantially in their association with correctness: global confidence is moderately associated with correctness, whereas local confidence shows little association. We further test whether question-level disagreement between the signals is associated with sampling instability. On ARC, larger local--global confidence gaps are associated with higher answer entropy, more distinct sampled answers, and lower modal-answer concentration. The gap--entropy association persists when disagreement and instability are estimated from disjoint stochastic samples, indicating that it is not explained by shared finite-sample variation. The corresponding relationship is substantially weaker on MMLU, where only 4% of questions exhibit sampling instability. These results show that confidence readouts derived from the same predictive system are not empirically interchangeable and that their disagreement can provide a diagnostic of unstable sampling behavior. Confidence should therefore be treated as an explicitly defined measurement rather than as a single intrinsic scalar property of a model, particularly when it is used to inform downstream evaluation, oversight, or control.
cs.LG / 55 / 2609.16977
Structural Negative Transfer in Federated Graph Neural Networks: Diagnosis, Causal Investigation, and the Limits of Divergence-Aware Mitigation
Abstract
Federated learning lets multiple participants train a shared model without pooling raw data, by exchanging locally trained model updates instead. Federated averaging assumes that averaging local models is a reasonable way to solve one shared problem when participants' data are broadly similar. Work on non-IID federated learning has shown that this assumption can withstand differences in label and feature distributions. We ask whether it survives a different strain specific to graph neural networks, where client graphs differ not in label or feature distribution but in structure itself, requiring the same shared weights to operate over fundamentally different topologies. We call the resulting harm structural negative transfer. In a federation of real citation networks and synthetic structural proxies, a structurally atypical client lost more than half its achievable accuracy simply by joining. In an initial six-client federation, two label-free structural statistics computable before training were strongly associated with this harm. Expanding to twenty clients showed that degree divergence remained associated with harm, although more weakly, and survived removal of domain contrast. Spectral divergence did not replicate, which we trace to a confound caused by the composition of the reference pool used for leave-one-out statistics. A causal intervention isolating topology found no significant effect. A degree-normalization mechanism held across twenty-four seeds but did not explain the harm when corrected. The best of five candidate fixes beat a tuned baseline only until a matched, structurally blind control was applied, after which the gain disappeared. What survives is a modest, partially replicated, degree-specific signal that is not yet a validated predictor at scale.
cs.LG / 56 / 2609.17026
CLARE: Scalable Class-Incremental Continual Learning via a Sparsity-Based Framework
Abstract
Continual learning must balance the learning of new knowledge with the retention of previously learned knowledge to incrementally learn tasks from a data stream without catastrophic forgetting. While leveraging pretrained models has significantly advanced continual learning, existing methods exhibit a scalability bottleneck when trained sequentially on many tasks, suffering from performance degradation due to inter-task interference and loss of plasticity. Inspired by evidence that sparse fine-tuning achieves performance comparable to full fine-tuning, this paper presents a novel sparsity-driven continual learning framework. Our continual learning method, termed CLARE, operates in two stages: it first identifies a sparse, task-critical parameter mask via a sparsity-inducing objective, then performs mask-constrained fine-tuning by only optimizing parameters selected by the mask. This two-stage sparse adapter mechanism enables all tasks to be accumulated within a shared adapter space while reducing destructive interference across tasks. Extensive experiments demonstrate the scalability of CLARE. On the long task-sequence benchmark Omnibenchmark-1k, CLARE outperforms strong baselines in final accuracy by a large margin, e.g, improving EASE by 4.64% and 13.34% after learning 100 tasks, respectively.
cs.LG / 57 / 2609.17029
Distributed JEPA: A Self-Supervised Framework for Energy Forecasting
Abstract
Traditional energy forecasting solutions rely on task-specific supervision and energy asset representations, limiting transferability and the ability to capture general temporal dynamics across heterogeneous assets. We address this by proposing a distributed Joint Embedding Predictive Architecture (JEPA) for self-supervised learning from heterogeneous energy time-series. The framework predicts latent representations of masked temporal segments while integrating temporal observations and contextual information within a shared embedding space. To prevent representation collapse, training combines a latent-space predictive objective with covariance and temporal variance regularization. The evaluation was conducted on energy consumption and generation datasets under data-degradation scenarios and compared with a Transformer forecasting baseline. The learned representations remained stable (cosine similarity $\approx 0.98$; effective rank 185-235). JEPA achieved performance comparable to a Transformer on building energy data, higher $R^2$ in 3/5 consumer clusters, and outperformed the baseline on 9/10 unseen PVs ($R^2$=0.73-0.88 vs. <0.45), while showing greater robustness to missing data.
cs.LG / 58 / 2609.17042
Learning Options for Compositional Motor Control with Adapter Banks
Abstract
Learning flexible motor primitives is a hallmark of skilled motor control. Recent neuroscience theory proposes that motor primitives may be implemented as low-rank perturbations of a shared recurrent network, but leaves open how such a system is learned. We translate this principle into a novel architecture for learning motor skills end-to-end: a shared recurrent core modulated by a bank of residual adapters, each selected by a discrete latent code. Trained on closed-loop biomechanical control, the adapters develop emergent low-rank perturbations of the recurrent dynamics despite no architectural rank constraint, placing task representations in disparate subspaces of the shared core network. A simple high-level policy over the learned options, optimized while the whole network is frozen, sequences the low-rank adapters to produce novel out-of-distribution movements. We demonstrate the ability to generalize to novel motor sequences within the closed-loop control setting, improving on the generalization error of a task-input-conditioned multitask baseline by upto order of magnitude.
cs.LG / 59 / 2609.17061
Repurposing Unified Topological Signatures for Graph Representation Learning
Abstract
Message-passing Graph Neural Networks (GNNs) iteratively propagate and aggregate local neighborhood information followed by global readout to learn graph representations. However, their discriminative power is upper-bounded by the Weisfeiler--Lehman (1-WL) graph isomorphism test. This prevents GNNs from distinguishing certain non-isomorphic graphs with identical local neighborhood structures, often leading to similar graph representations. Unified Topological Signatures (UTS) capture compact, multi-scale representation of global graph topology derived from persistent homology. We introduce two complementary UTS signatures: Graph_UTS- a static signature of the input graph topology, and Embedding_UTS- a dynamic signature of the evolving embedding topology. They encode structural information inaccessible to 1-WL-based message-passing GNNs, yet their capabilities are explored solely for post-hoc embedding-space analysis. We integrate UTS into GNN training across three architectural interventions: (i) UTS-Aug: augmenting with standard readout feature that encodes graph's true topology; (ii) UTS-Reg: topological regularizer that constrains representation collapse; (iii) UTS-Pool: topology-guided pooling that retains structurally critical nodes. We further leverage UTS as a layer-wise diagnostic to quantify oversmoothing during GNN training. Theoretically, we show that integrating UTS into GNN optimization strictly extends GNN expressivity beyond the 1-WL hierarchy. Experiments on three graph classification benchmarks show consistent benefits: Graph-UTS, Dual-UTS, and UTS-Pool improve accuracy across all three datasets, Embedding-UTS provides smaller but similarly consistent gains, and UTS-Reg's benefit varies across graph domains. Accuracy improves by up to 5.8% with Graph-UTS augmentation, by up to 1.9% with UTS-Reg, and achieves comparable performance to TOGL with UTS-Pool.
cs.LG / 60 / 2609.17101
High-Fidelity Digital Twin Data Models by Randomized Dynamic Mode Decomposition and Deep Learning with Applications in Fluid Dynamics
Abstract
The purpose of this paper is the identification of high-fidelity digital twin data models from numerical code outputs by non-intrusive techniques (i.e., not requiring Galerkin projection of the governing equations onto the reduced modes basis). In this paper the author defines the concept of the digital twin data model (DTM) as a model of reduced complexity that has the main feature of mirroring the original process behavior. The significant advantage of a DTM is to reproduce the dynamics with high accuracy and reduced costs in CPU time and hardware for settings difficult to explore because of the complexity of the dynamics over time. This paper introduces a new framework for creating efficient digital twin data models by combining two state-of-the-art tools: randomized dynamic mode decomposition and deep learning artificial intelligence. It is shown that the outputs are consistent with the original source data with the advantage of reduced complexity. The DTMs are investigated in the numerical simulation of three shock wave phenomena with increasing complexity. The author performs a thorough assessment of the performance of the new digital twin data models in terms of numerical accuracy and computational efficiency.
cs.LG / 61 / 2609.17160
Neural Field Ensembles for Aerodynamic Surface Prediction: Winning Solution to the ONERA CRM Wall Distribution 2025 Challenge
Abstract
Machine-learning surrogate models offer a promising alternative to high-fidelity Computational Fluid Dynamics (CFD) simulations for aerodynamic analysis and design. However, constructing accurate surrogates for realistic aircraft configurations remain challenging due to complex geometries, multiple flow regimes, and limited training data. This work presents the methodology that achieved first place in the ONERA CRM Wall Distribution Regression Challenge, which focuses on predicting pressure and skin-friction coefficient distributions over the NASA Common Research Model wing-body-pylon-nacelle configuration under different operating conditions. The proposed approach formulates the problem as a conditional neural field mapping spatial coordinates, surface normals, and operating conditions to aerodynamic wall quantities. Fourier feature encoding, a relative squared error objective aligned with the challenge metric, ensemble learning, and $k$-fold cross-validation are progressively introduced to improve prediction accuracy and exploit the limited training data. Beyond presenting the final methodology, the paper documents the successive model design choices that led to the winning solution through a comprehensive ablation study and discusses several alternative approaches that were investigated but ultimately discarded. On the hidden competition test set, the proposed methodology achieves an overall score of 8.81, outperforming the strongest organizer-provided baseline, which achieved a score of 8.64, while requiring approximately three orders of magnitude fewer trainable parameters. These results illustrate that carefully designed coordinate-based neural fields constitute an efficient and robust framework for aerodynamic surrogate modeling on complex geometries under limited-data conditions.
cs.LG / 62 / 2609.17171
A unified framework for global and local interpretability using adaptive derivative-ordered random explanation
Abstract
The interpretability of complex machine learning models is of paramount importance, especially in real-world high-stakes domains such as healthcare and finance. However, existing post-hoc interpretability methods suffer from inherent limitations: fragmented analytical processes, inadequate capacity to model nonlinear feature interactions, computational inefficiencies, and over-reliance on specific model architectures. To address these challenges, this paper provides a novel method - Adaptive Derivative-Ordered Random Explanation (ADORE) - that leverages first- and second-order derivatives to accommodate nonlinear model complexities, while enabling effective capture of feature-sample interactions within a unified analytical framework. ADORE integrates global feature importance with local sample contributions, precisely quantifying feature impact by capturing both magnitude and direction, and identifying critical samples influencing model decisions. Furthermore, it achieves computational efficiency through randomized singular value decomposition (SVD) and dynamic sparsity detection, making it scalable to large, high-dimensional datasets. Experiments across three data modalities - tabular, text, and image - demonstrate that ADORE outperforms existing methods such as LIME and SHAP in handling complex interactions and computational efficiency, while providing detailed and reliable explanations. To facilitate adoption and reproducibility, ADORE has been released as an open-source Python package, hosted on GitHub, enabling researchers and practitioners to readily adapt and apply our approach to their specific tasks, models, and datasets.
cs.LG / 63 / 2609.17175
IRENE: A Convolutional GRU Ensemble Model for Radar Precipitation Nowcasting over Italy
Abstract
We present IRENE (Italian Radar Ensemble Nowcasting Experiment), a deep learning model for probabilistic short-range precipitation nowcasting over the Italian domain at \SI{1}{km} spatial and 5 min temporal resolution. IRENE adopts an encoder--forecaster architecture built on multi-scale Convolutional Gated Recurrent Units (ConvGRUs), trained on the national radar composite produced by the Italian Civil Protection Department (DPC). An importance-sampling scheme focuses training on precipitation-relevant events, while the almost-fair Continuous Ranked Probability Score (afCRPS) is adopted as the primary probabilistic loss function. Two additional training configurations are proposed: an adversarial (GAN) variant, IRENE-GAN, designed to improve the spatial sharpness of the generated forecasts, and a spectrally constrained variant, IRENE-GAN-RAPSD, in which the adversarial objective is complemented by an explicit penalty on the radially averaged power spectral density. The three configurations are evaluated against the stochastic extrapolation method STEPS and the pre-trained deep learning model DGMR. All IRENE configurations attain a lower Continuous Ranked Probability Score than both benchmarks at every lead time and rank histograms closer to uniformity, indicating better probabilistic skill and ensemble calibration. In terms of ensemble-mean mean absolute error the advantage is confined to the first 90 min, beyond which the strongly damped DGMR fields and, to a lesser extent, STEPS become competitive. Spectral analysis shows that the adversarial training removes the progressive loss of small-scale variance exhibited by IRENE, at the cost of an excess of fine-scale power at long lead times that the spectral penalty only partially controls.
cs.LG / 64 / 2609.17184
LoopSpec: Pipelined Self-Speculative Decoding for Looped Transformers
Abstract
Looped Transformers achieve strong performance with compact parameter sizes by repeatedly applying a shared stack of Transformer blocks across recurrent depths. However, they incur higher decoding latency than standard Transformer models of comparable parameter size because shared weights are accessed at every recurrent depth. To improve decoding efficiency, self-speculative decoding is particularly well suited to Looped Transformers, as their intermediate recurrent states can directly provide draft predictions without an auxiliary draft model. We therefore propose LoopSpec, a training-free self-speculative decoding framework tailored for Looped Transformers. LoopSpec extracts draft tokens from early recurrent states and operates in a pipelined manner, overlapping draft generation of future tokens with target verification of the current token. To improve draft accuracy without excessive compute overhead, we introduce a selective second proposal from deeper recurrent depth while ensuring lossless decoding under both greedy and sampling regimes. Furthermore, we derive the optimal proposal depths in closed form and show the prediction matches measurement. Across reasoning and coding benchmarks, LoopSpec achieves up to 6.83$\times$ inference speedup across diverse Looped Transformers.
cs.LG / 65 / 2609.17223
Memorisation bias in medical AI
Abstract
Medical AI models hold immense potential to improve patient outcomes, but they are also known to unintentionally memorise individual records from their training datasets. While such memorisation has been linked to targeted privacy attacks, its consequences for clinical deployment, where patients may be assessed by a model that saw their historical data during training, remain poorly understood. Here we show that predictions on a patient's unseen future data can change significantly if a model observed that same patient's anonymised historical data during training, a phenomenon we term "memorisation bias". We demonstrate that this bias exists across diverse data modalities and model architectures, and over prolonged time spans: in some cases, memorisation bias persists on future records acquired decades after the historical records used for training. Moreover, in simulated prospective deployment, memorisation bias has asymmetric effects on the diagnostic accuracy of returning data contributors. When a patient returned with a de novo condition absent from their historical records in the training dataset, diagnostic sensitivity decreased significantly compared to an otherwise identical model not trained on their historical data. Conversely, when their health state was unchanged, both sensitivity and specificity were significantly inflated. Our findings reveal a previously uncharacterised risk in medical AI that arises when a model is deployed on patients who contributed to its training data. This exposes a shortcoming of current model development practice: the de-identification measures designed to protect patients' privacy make it difficult to identify returning contributors and exclude them from the AI-assisted interpretation of their own future data. Mitigating memorisation risks may thus require changes to current model training and deployment protocols.
cs.LG / 66 / 2609.17226
Easy to Catch a Liar, Hard to Clear an Honest One: Language Models Diagnosing a Corrupted Reward Channel from a Verified Record
Abstract
An agent that learns from rewards has to trust whatever reports those rewards. When the reports suddenly change, either the world changed or the reporter broke. From the reports alone these are indistinguishable, and reinforcement learning theory shows that no amount of further experience separates them. The prescribed escape is richer data about the reporter itself. We ask whether a frozen language model, handed exactly that data, uses it. We build a two-option game in which a payout swap and a lying reporter produce byte-identical histories. Then we add one verified record: an independent check of one round's real result, printed beside what the reporter said about that round. That single line settles the case. We ask three large models, from two families, to answer one question with one letter. Is the reporter honest or lying? They catch a lying reporter almost perfectly. At the 70B class that holds in every condition we tried; the 32B model slips in one wording. They clear an honest reporter far less often, and how often depends on things that should not matter. Averaged over rounds, letters, and wordings, a 72B model calls an honest reporter a liar 38% of the time when nothing has changed at all, and 58% of the time when the payouts moved. A 70B model from a second family calls an honest reporter a liar 26% and 48% of the time. The failure is not one of reading, because in the situation where nothing changed the same models score 0.96 to 1.00 with the answer printed in the prompt. Which surface feature drives it differs by family. For the Qwen models it is which round the record names, and for Llama it is which letter stands for "honest." Adding the record to a prompt that already states the answer makes Llama less likely to give that answer. We had registered a prediction for that 58% before the run: 35%. The failure is larger than we expected.
cs.LG / 67 / 2609.17284
Personalized Federated Learning through Global Knowledge Distillation and Local Head Adaptation
Abstract
Statistical heterogeneity limits federated learning when a single global classifier cannot represent client-specific label distributions. In this work, we propose Personalized Federated Knowledge Distillation with Head Adaptation (pFedKDH), which aggregates only the shared backbone, keeps persistent client-specific heads, and uses a recalibrated global head as a teacher during local training. Across MNIST, Fashion-MNIST, CIFAR10, and CIFAR100 under class-wise Dirichlet partitions, pFedKDH obtains the best accuracy in most settings, with accuracy gaps up to 37.67\% over the weakest baseline and consistently low standard deviation across repetitions. Component-wise diagnostics and convergence results support the role of persistent heads and distillation-guided local optimization under label-skewed data.
cs.LG / 68 / 2609.17287
Same Flow, Different Paths: Variance Reduction in Flow Matching
Abstract
In flow matching (FM), a velocity model $v_θ$ is trained using a predefined path $g_t$ that connects data and noise samples (e.g., $g_t(x_0, x_1) = (1 - t) x_0 + t x_1$). In this work, we study the choice of this path from an optimization perspective by analyzing the variance of stochastic gradients. We consider the class $G(p_t,v^\star_t)$ of paths that induce the same marginal distributions $p_t$ and marginal velocity field $v^\star_t$, and therefore the same FM objective. Our main finding is that the choice of path $g_t$ can fundamentally change the convergence rate of SGD, even when the FM objective remains exactly the same. (i) For a linear velocity model and one-dimensional Gaussian data, we derive a tight bound on the SGD iteration complexity up to logarithmic factors and find an analytically optimal path that minimizes this bound among linear paths inducing the same FM problem. (ii) We then extend the variance analysis to general FM problems and formulate path selection at a fixed $θ$ as the variance-minimization problem PathOpt$_θ$, constrained to $g_t\in G(p_t,v^\star_t)$. We show that this constraint is essential: reducing variance without it can lead to slower convergence. (iii) Since the constraint $g_t \in G(p_t,v^\star_t)$ cannot generally be verified directly, we derive an equivalent formulation with constraints that can be estimated from samples, allowing paths to be found numerically. Our theoretical results are supported by experiments with Gaussian data, Gaussian mixture models, and real datasets.
cs.LG / 69 / 2609.17358
Hybrid Variational Quantum Circuits for Multivariate Regression and High-Dimensional Data Reconstruction
Abstract
Variational quantum circuits (VQCs) are parameterized quantum circuits optimized classically. We propose a hybrid variational quantum circuit (HVQC) extending VQCs with a classical affine post-measurement layer, enabling vector-valued regression without the linear overhead of independent scalar circuits. Theoretically, we show that elementary one-and two-qubit circuits can approximate quadratic functions and products via data re-uploading and entanglement, providing the foundations of the full architecture. Experimentally, on two synthetic image reconstruction datasets and the Friedman1 benchmark (40,568 test samples), our HVQC matches Gaussian Process Regression and outperforms XGBoost and Random Forest. An ablation study confirms that both quantum and classical components are essential, and results highlight the central role of the feature map in hybrid quantum-classical models.
cs.LG / 70 / 2609.17380
OPEN-1B: A Fully Auditable Training Run
Abstract
Open-source language models have a reproducibility problem. Despite releasing weights, training data, and recipes, none of them are provably reproducible due to the non-associativity of floating-point arithmetic. Deep learning frameworks often offer a deterministic execution mode, allowing reproducible operations on the same machines. Unfortunately, this determinism does not carry across hardware such that a user can verify that a released checkpoint was actually produced using the declared training recipe. This leaves room for undisclosed data, injected biases, or backdoors that existing techniques such as proof-of-learning or proof-of-training-data cannot rule out. We introduce a new tier of model transparency, fully auditable, in which every operation on every data sample during training is independently reproducible on heterogeneous commodity hardware with bitwise certainty. By imposing a definite order on the sources of training nondeterminism, GPU kernel reductions, data batch ordering across a data-parallel cluster, and inter/intra-node collective communication, we make it possible to replay any individual step of a large, distributed training run on a single piece of commodity hardware and check it against the published trajectory. Because replaying an entire run on one machine is infeasible, we support this with a collective verification scheme in which many independent auditors each certify individual steps, together covering the whole run. We release Open-1B, a model trained under this regime, together with its full pretraining dataset, every intermediate checkpoint, the training codebase, and the audit harness needed to reproduce and verify any step of its training.
cs.LG / 71 / 2609.17386
Bridging the Confidence Gap: Temperature Scaling for Calibrating Test-Time Prompt Tuning
Abstract
Test-time prompt tuning (TPT) enables adaptation on a single test instance, achieving improved accuracy but often sacrificing calibration performance. Most existing calibration methods introduce additional regularization terms to promote dispersion across text embeddings and reduce calibration error, yet these methods often suffer from a drop in accuracy. Motivated by the well-calibrated nature of zero-shot predictions, we propose CoTS, a simple yet effective post-hoc calibration method that preserves accuracy. Specifically, CoTS applies temperature scaling to minimize the confidence gap between adapted and zero-shot predictions. To fully exploit the potential of multiple augmentations during adaptation, we introduce a weak-strong ensemble strategy that further boosts accuracy. We then apply CoTS to this ensemble, termed E-CoTS, to maintain its well-calibrated property. Extensive experiments on diverse datasets and backbones show that our approaches effectively mitigate miscalibration without compromising primary accuracy. For instance, E-CoTS reduces the average expected calibration error of TPT from 11.90% to 5.38% on ImageNet variants, while even increasing accuracy from 60.74% to 62.95%. Moreover, when integrated with existing calibration methods, E-CoTS usually enhances both accuracy and calibration simultaneously.
cs.LG / 72 / 2609.17429
Learning-Guided Planning in Large Dynamic Action Spaces: Budgeted Tree Search for One-to-Many Mobile Charging
Abstract
Many learned sequential decision systems map the current state directly to an action. That shortcut becomes brittle when candidate actions are numerous, geometrically structured, and rebuilt with the state. One-to-many mobile charging makes this setting concrete: with N=250 sensors, the initial state induces about 1,125 candidate charging-stop actions; each chosen stop simultaneously serves its in-range sensors, and the action universe changes as sensors die. LP-BTS is a learning-guided planning architecture: a graph proposal policy concentrates a small candidate support, a learned value critic evaluates leaves, and edge-budgeted PUCT compares short simulated futures before committing an action. Because the policy scores this set without a fixed output head, a single frozen checkpoint covers every evaluated setting, spanning action universes from 736 to 2,813 stops. Matched ablations reveal complementary effects: uniform sampling costs 8.8 survival percentage points, while, with targeted support fixed, PUCT jointly retains 1.4 points (about 3.5 of 250 sensors) and direct policy selection travels 23% farther. On a prospectively specified, sealed 30-scenario confirmatory bank evaluated once, LP-BTS attains the highest observed survival (0.4545) and alive-AUC (0.8031). Its estimated survival advantage over the strongest domain-engineered comparator is +0.0066 (95% CI [-0.0037, +0.0184]), an unresolved difference, while it exceeds a deadline heuristic and two source-derived direct-policy reconstructions on every paired scenario. Both learned rows are trained, source-derived reconstructions of variants reported by Gong et al. In this setting, the results provide controlled evidence about learning-guided planning in a large, dynamic action space.
cs.LG / 73 / 2609.17440
Reduced-Space Multi-Fidelity Bayesian Optimization of Process Simulation Models
Abstract
Optimizing industrial process flowsheets is often computationally prohibitive due to the high cost of rigorous simulations and the curse of dimensionality inherent in complex design spaces. To address these challenges, we present a reduced-space multi-fidelity Bayesian optimization (RS-MFBO) framework designed for high-dimensional, expensive black-box functions. The approach integrates Global Sensitivity Analysis (GSA) for dimensionality reduction with a fidelity-augmented Gaussian process that captures correlations between low-cost approximations and expensive high-fidelity evaluations. A cost-aware acquisition strategy, augmented with cooldown and promotion mechanisms, adaptively guides the allocation of samples across fidelities. The framework is validated on two distinct industrial process simulators: a plasmid DNA bioprocess in SuperPro Designer and a green fuel synthesis plant in Aspen HYSYS. Results across diverse economic and physical objectives demonstrate that the proposed method substantially reduces the number of high-fidelity simulator evaluations while maintaining competitive optimization performance compared to single-fidelity baselines. These results highlight RS-MFBO as a scalable, simulator-agnostic approach for cost-constrained black-box optimization.
cs.LG / 74 / 2609.17491
FreqSpaNet: Frequency and Spatial Learning of SFPF for Physical Layer Hardware Integrity Detection
Abstract
Unauthorized hardware replacement can preserve a wireless device's logical identity while altering its physical implementation, posing a challenge to hardware integrity verification. Spatio-frequency polarization fingerprints (SFPFs) capture device-dependent responses across multiple frequencies and directions, but their frequency and spatial dimensions exhibit different structural dependencies. We propose FreqSpaNet, an SFPF representation learning network for open set hardware anomaly detection. A frequency branch captures local variations among neighboring frequencies, while a geometry-aware spatial branch models directional relationships using angular information. The two representations are combined through adaptive fusion, and complementary pretraining further captures shared information while preserving the distinct characteristics of the frequency and spatial representations. Experiments show that FreqSpaNet achieves a mean AUROC of 96.31\%, 9.05 points above the baseline. Results under seven hardware replacement scenarios further verify the effectiveness of FreqSpaNet.
cs.LG / 75 / 2609.17499
ENCP: Episode-Normalized Conformal Prediction for Vision-and-Language Navigation
Abstract
Uncertainty estimation for Vision-Language-Navigation (VLN) models is a critical task since it can help identify ambiguous and unreliable predictions, enabling agents to make safer navigation decisions. As one of the most advanced uncertainty estimation frameworks, conformal prediction (CP) offers a promising approach for uncertainty estimation in VLN. However, given that VLN agent requires a sequence of steps, standard calibration in conformal prediction fails to provide coverage guarantee it promises over a dependent, variable-length VLN episode. To this end, we propose Episode-Normalized Conformal Prediction (ENCP), which rescales a nonconformity score by the policy's residual confidence and calibrates one maximum score per episode. Under exchangeable calibration and test episodes, this construction covers the ground truth at every step with probability at least $1 - α$, while allowing dependence among steps within an episode. Across four VLN policies and three nonconformity scores on R2R and REVERIE dataset, ENCP meets all reported empirical step-coverage targets on the seen-to-unseen evaluation. These results demonstrate that ENCP can provide model-agnostic uncertainty estimates, which might be useful for determining when a VLN agent should defer to a more capable predictor, including human assistance.
cs.LG / 76 / 2609.16370
Fast-Convergent Meta-RL via Gradient-Clustered BS Sampling for Edge Caching
Abstract
Wireless edge caching networks typically consist of many independent Base Stations (BSs), each facing its own request rate and content popularity profile. Training a Reinforcement Learning (RL) caching agent from scratch at every BS forces each agent to relearn, through slow trial and error, a decision problem that is structurally identical across the network. Meta-reinforcement learning removes this redundancy by learning a shared initialization that adapts to any BS in a few local updates; however, meta-training itself becomes the bottleneck at scale: the meta-gradient must be estimated from a small subset of BSs at each meta-iteration, and sampling this subset uniformly at random yields a high-variance estimate, an issue existing meta-RL caching frameworks leave unaddressed. This paper proposes a meta-reinforcement learning framework for caching across independent, non-overlapping BSs that directly targets this bottleneck. Each BS runs a local Proximal Policy Optimization (PPO) agent, formulated as a Semi-Markov Decision Process (SMDP) over content popularity, size, lifetime, and importance, while a shared meta-policy is learned via a Model-Agnostic Meta-Learning (MAML)-style loop. To scale meta-training and accelerate convergence, we introduce gradient-based clustering, which groups BSs by local gradient similarity and draws from every cluster, in proportion to its size, at each meta-iteration. We prove, via an Analysis of Variance (ANOVA)-style decomposition of gradient variance, that this strategy yields a strictly lower-variance meta-gradient estimator than uniform random sampling under BS heterogeneity.
cs.LG / 77 / 2609.17014
Beyond Measurement Metrics: A Human-Centered Framework for Semantic Validation of Network Traffic Classification
Abstract
Machine learning (ML) has become the dominant approach for network traffic classification, achieving very high predictive performance. However, a model is only valuable if it learns semantically meaningful and trustworthy patterns rather than exploiting spurious correlations. Conventional evaluation practices predominantly assess predictive performance. Consequently, whether the model relies on semantically meaningful patterns remains unknown. To address these challenges, we adapt the knowledge generation framework for network traffic classification. The adapted framework combines data, ML models, explainability, visualization, and expert reasoning to support the iterative exploration, verification, and refinement of model behavior and data preprocessing. The framework is grounded in findings from the literature, benchmark dataset analyses, practical experience with XAI-based traffic classification, and expert feedback, providing practical guidance for semantic model validation. By complementing predictive performance with semantic validation and human expertise, the proposed framework supports the development of network traffic classification models that are not only accurate but also robust and trustworthy.
cs.LG / 78 / 2609.16443
The Neverwhere Visual Parkour Benchmark Suite
Abstract
State-of-the-art visual locomotion controllers are increasingly capable at handling complex visual environments, making evaluating their real-world performance before deployment increasingly difficult. This work intends to narrow this train/evaluation gap by developing a collection of hyper-photo-realistic, closed-loop evaluation environments - The Neverwhere Benchmark Suite - comprised of over sixty 3D Gaussian Splatting reconstructions of urban indoor and outdoor scenes. Our goal is to encourage large-scale and reproducible robot evaluation by making it easier to create and integrate Gaussian splats-based reconstructions into simulated continuous testing setups. We also underscore the potential pitfalls of relying exclusively on 3D Gaussian-generated data for training, by providing policy checkpoints trained over multiple Neverwhere scenes and their performance when evaluated in novel scenes. Our analysis illustrates the necessity of sourcing diverse data to ensure performance. Code and data are available on the project page: https://ziyc.github.io/neverwhere-bench/.
cs.LG / 79 / 2609.16745
The Latent That Never Was: A Forensic Re-run of the CVAE Ablation in Action Chunking Transformer
Abstract
Action Chunking Transformers (ACT) are widely used to learn robot manipulation from demonstrations. Their conditional variational autoencoder includes an encoder meant to capture differences between demonstrations during training. The original ACT paper reported that encoder removal dropped the mean success rate from 35% to 2% on two simulated tasks with human demonstrations. We re-ran this ablation in the original code and checked whether the findings depend on the implementation or training data. The published drop does not reappear in our tests, although smaller gains or losses in success rate remain uncertain. To investigate the discrepancy, we varied training length and how checkpoints are selected for evaluation. Both can reverse which policy scores higher, but the published drop's cause remains unknown. Success rates alone leave open whether the encoder provides information that helps the policy reconstruct demonstrated actions. On the tested ACT benchmark, the sampled latent provides little reconstruction benefit at every tested nonzero weight of the penalty on latent information. At inference, ACT leaves this latent unused and sets it to zero. Skipping the encoder increases training throughput in both implementations we timed. We release code, evaluation tools and results so others can repeat the comparisons and test the encoder on other tasks.
cs.LG / 80 / 2609.16864
TEMPO: Learning Temporal Context for Dynamic Robot Manipulation
Abstract
Vision-language-action (VLA) models have achieved impressive performance in quasi-static manipulation, but struggle in dynamic manipulation tasks because they operate on a single observation at inference time. We identify two representational failures that underlie this limitation. The first is motion ambiguity, where a single observation does not include scene dynamics and therefore cannot anticipate the future state of moving objects. The second is state aliasing, where visually similar observations from different points in a task require different actions. We argue that these failures persist regardless of model scale and inference latency, showing that the bottleneck is missing temporal context rather than model capacity. Based on this insight, we propose TEMPO, which augments a pretrained VLA with two temporal inputs: a motion summary extracted from a frozen video foundation model to resolve motion ambiguity and a compact proprioceptive history to resolve state aliasing. TEMPO requires no modification to the backbone and adds minimal compute overhead at training or deployment. Across four dynamic manipulation tasks, it improves Bottle Handover success from 44% to 74% and is the only method that solves state aliasing. Probing and ablation studies confirm that each temporal signal independently addresses its corresponding failure. We further release TEMPO-Bench, a benchmark of over 50k annotated frames for evaluating motion-aware robot perception in both regression and multiple-choice formats. Project Website: https://tempo-robot.github.io/
cs.LG / 81 / 2609.17115
Intrinsic Robot Rewarding: Reusing VLA Representations for Autonomous Evaluation and Policy Improvement
Abstract
Vision-language-action (VLA) systems already bring together two valuable resources for robot learning: rich visual representations and demonstrations of successful task execution. Intrinsic Robot Rewarding (IRR) proposes to use these resources for a second, complementary purpose: evaluating the robot's own outcomes and providing feedback for policy improvement. Successful demonstration endpoints define task-specific references, and the policy's frozen visual encoder provides the feature space in which new outcomes are assessed. The core reward mechanism adds a reference bank and a scoring operation to the existing pipeline, without requiring a separate learned evaluator or an additional perception backbone. Our position is that this reuse offers a promising route to lower integration effort, efficient reward computation, and reduced recurring human outcome scoring. Building on established research in visual rewards and learning from experience, IRR brings these ideas into the robot's existing perception and demonstration pipeline. An operational COMAU Racer 3 demonstrator is available at technology readiness level 4 (TRL 4). This laboratory foundation supports the next research step: connecting internal outcome evaluation to physical policy improvement. We present the reward formulation, central research questions, and an evaluation methodology linking reward reliability to task success and supervision effort. The intended contribution is a reusable approach to learn and improve from the data and experience already available in industrial robot systems.
cs.LG / 82 / 2609.16199
A Sentinel-2 benchmark dataset for deep-learning active-fire segmentation across 25 California wildfires
Abstract
This article describes an open image dataset for developing and evaluating active-fire segmentation methods in satellite imagery. The dataset contains 2,148 image-mask pairs from 25 California wildfires, with acquisitions spanning July 2020 to August 2026. Each image is a 512x512-pixel, three-channel composite derived from Sentinel-2 Level-2A bands B12, B11 and B8A at 20 m spatial sampling. A fixed linear rendering is applied throughout the dataset. Corresponding masks distinguish background, SWIR-rule active fire and invalid observations. The masks were generated from shortwave-infrared brightness and near-infrared contrast, followed by constrained neighborhood growth. The release includes chip-level metadata and an incident-disjoint partition containing 18 training, three validation and four test fires. Among the image pairs, 841 contain active-fire labels; these labels occupy 0.0766% of all grid cells. A mask-blind analyst review covers 233 test chips and provides a separate assessment of the rule-generated labels at chip and connected-component levels. Reference training and evaluation code accompanies the data, including a ResNet-34 U-Net implementation with validation-based checkpoint and threshold selection. The archived images, masks, metadata and review annotations support research on rare-class segmentation, learning from algorithmic labels and transfer across fire incidents. The versioned dataset is deposited on Zenodo, with preparation and reuse software maintained in a public GitHub repository.
cs.LG / 83 / 2609.16279
Semantic-Aware Neural Video Codec for Error-Resilient Low-Latency Transmission
Abstract
Emerging physical AI systems require low-latency, task-oriented video communication over unreliable channels. We propose a semantic-aware multi-level neural video coding method for robust low-latency video transmission over unreliable channels that are abstracted as multi-level packet erasure channels. Built upon the real-time DCVC-RT neural video codec, the proposed framework introduces a semantic- and feature-aware coding strategy that partitions encoded representations into packets carrying different levels of semantic and latent-feature importance and assigns these packets to different streams, each associated with a priority level when transmitted over unreliable communication channels. We also developed an error-resilient entropy model that removes inter-packet dependencies, allowing each packet to be decoded independently under packet losses. The complete system is trained end-to-end over the abstracted multi-level packet erasure channels, enabling learning of channel-aware representations together with importance-aware packet assignment while facilitating the network for differentiated packet prioritization. Experiments show that the proposed framework significantly improves robustness over baseline DCVC-RT under packet erasures, achieving graceful degradation in less important regions while better preserving task-relevant visual content.
cs.LG / 84 / 2609.17298
Quantum-Inspired Trainable and Parameter-Efficient Tensor Networks for Image Inpainting
Abstract
This work introduces quantum-inspired tensor-network circuits as trainable transforms for image inpainting. Among the proposed architectures, the diagonal quantum Fourier transform (QFT) relaxation is invertible with $O(N^2 \log N)$ computational cost for $N\times N$ images, inherently preserving minimum coherence throughout training via its circuit structure and eliminating the need for explicit coherence penalties. Unconstrained gradient-based phase optimization (Riemannian-optimization free) enables efficient learning from randomly sampled training data, allowing the learned transform to generalize to test images observed through fixed sampling masks. Numerical tests show that the learned models outperform fixed transforms and per-image optimization while matching the performance of much larger unitary architectures, yet with far fewer parameters.
cs.LG / 85 / 2609.17297
Goal-oriented probabilistic forecasting for dynamic PRB allocation in 5G networks
Abstract
Efficient physical resource block (PRB) allocation in 5G networks requires accurate demand forecasting. Conventional methods minimize symmetric error metrics (MAE, RMSE), ignoring the operational cost asymmetry where under-provisioning (service degradation) is far costlier than over-provisioning (wasted capacity). We propose a goal-oriented probabilistic forecasting framework that aligns model training with the operator's decision-making objectives. Specifically, we train DeepAR and Temporal Fusion Transformer (TFT) models using the Pinball Loss function and derive the optimal allocation quantile from the operator's cost matrix. Evaluation on a real beam-level 5G traffic dataset shows that the proposed approach reduces operational cost compared to MSE-trained baselines while maintaining calibrated uncertainty estimates. The framework enables dynamic PRB allocation that explicitly balances service reliability against resource efficiency.
cs.LG / 86 / 2609.16738
Unified Heterogeneous Graph Neural Network solver for Power Flow, Optimal Power Flow and State Estimation
Abstract
Power Flow (PF), Optimal Power Flow (OPF), and State Estimation (SE) are fundamental problems in power system analysis, but solving them is computationally expensive. Graph Neural Networks (GNNs) have been proposed as fast surrogates, yet existing solvers are trained for a single problem at a time, producing narrow models that must be rebuilt for each new task. We propose a more general approach: a single Heterogeneous Residual Gated Graph Convolutional Network that solves all three problems with one shared backbone. Rather than learning one mapping, the model learns a reusable representation of how the network behaves, from which PF, OPF, and SE can each be estimated. Trained jointly on the three problems across diverse topologies and loading conditions, and evaluated on the IEEE 14-bus and 118-bus systems, the shared model matches the accuracy of task-specific GNN solvers and stays robust on unseen loading levels and topologies. These results show that a single model can capture the basic operation of a power network and serve several analysis tasks at once, a first step toward a foundation model for power systems.
cs.LG / 87 / 2609.16406
Physics Informed Random Feature Neural Networks for Solving PDEs
Abstract
Machine learning-based partial differential equations (PDEs) solvers have attracted significant attention in recent years. Most progress in this area has been driven by deep neural networks such as physics-informed neural networks (PINNs) and kernel method (such as physics-informed Gaussian Processes). We introduce a physics-informed random feature method for countering part of the spectral bias which PINN-based solvers are facing for a certain class of PDEs. Random feature method was originally proposed to approximate large-scale kernel machines and can be viewed as a specialized randomized neural network. Compared to other state-of-the-art PINN-based solvers which require a large number of collocation points, our proposed method reduces the computational complexity. In this paper, we develop a rigorous approximation error analysis and derive high-probability error bounds on the $H^1$ norm. We provide extensive numerical tests for verifying our theoretical guarantees on error decay rates, as well as several comparison tests to showcase our claimed capability for combating spectral bias in these deep learning based methods.
cs.LG / 88 / 2609.17048
Near-Optimal Nonconvex Matrix Completion
Abstract
We study nonconvex methods for matrix completion, the problem of recovering a low-rank matrix from a subset of its entries. Convex methods achieve sample complexity linear in the matrix dimension and the rank, up to logarithmic factors, whereas global guarantees for commonly used nonconvex methods require a higher polynomial dependence on the rank. We close this gap by analyzing Riemannian gradient descent (RGD) and Riemannian Gauss--Newton (RGN) methods. For an $n\times n$ matrix of rank $r$ with incoherence parameter $μ$ and condition number $κ$, the two methods achieve exact recovery with high probability from $O(μnr\log n\log(nκ))$ and $O(μnr\log n\log(2μrκ))$ observations, respectively. The methods use a multiscale residual initialization, while the analysis simultaneously controls the spectral error and incoherence. The resulting RGD iterates converge linearly, whereas RGN eventually converges Q-quadratically.
cs.LG / 89 / 2609.17089
Optimization over covariance matrices with a parameterized metric
Abstract
The choice of Riemannian metric can strongly influence the convergence of gradient-based optimization over covariance matrices. Euclidean, Bures-Wasserstein and affine-invariant metrics are common choices, but their relative effectiveness depends on the objective. We introduce a two-parameter family defined by $X^{p}LX^{q}+X^{q}LX^{p}=U$, solved for $L$ at each tangent vector $U$, that contains all three as exact members, at $(0,0)$, $(1,0)$ and $(1,1)$, and extends past them. We treat the choice of member as a particular way of preconditioning for a given problem. To this end, we analyze the conditioning of the Riemannian Hessian at the solution. We show that it obeys a lower bound that depends on $(p,q)$ only through the exponent $r=p+q$. When the Euclidean Hessian is a pure power that mixes no eigendirections, the member $p=q=r/2$ attains that bound, and a closed-form criterion identifies the other members that do. We discuss ways to tune $r$ for a given problem. Experiments on real covariance data confirm the predicted conditioning and the benefit of tuning $r$. A task covariance example shows a further gain from tuning the shape.
cs.LG / 90 / 2609.17483
Bridging the Gap Between Homogeneous and Heterogeneous Asynchronous Optimization Is Surprisingly Difficult
Abstract
Modern large-scale machine learning tasks often require multiple workers, devices, CPUs, or GPUs to compute stochastic gradients in parallel and asynchronously to train model weights. Theoretical results typically distinguish between two settings: (i) the homogeneous setting, where all workers have access to the same data distribution, and (ii) the heterogeneous setting, where each worker operates on different data distributions. Known optimal time complexities in these settings reveal a significant gap, with far more pessimistic guarantees in the heterogeneous case. In this work, we investigate whether these pessimistic optimal time complexities can be overcome under different assumptions. Surprisingly, we show that improvement is provably impossible under widely used first- and second-order similarity assumptions for any randomized algorithm. We then turn to the interpolation regime and demonstrate that the weak interpolation assumption alone is also insufficient. Finally, we introduce a minimal combination of irreducible assumptions, strong interpolation and the local Polyak-Lojasiewicz condition, to derive a new time complexity bound that matches the dependence on worker computation times in the best-known result in the homogeneous setting, without requiring identical data distributions.
cs.LG / 91 / 2609.16157
Computer-assisted global regularity across nonlinear families of three-dimensional periodic Navier-Stokes flows
Abstract
Numerical simulations reveal how vortices stretch and transfer energy, but establishing smooth evolution requires bounds that remain valid beyond the simulated resolution. Here I develop a computer-assisted framework that establishes global regularity for continuous families of three-dimensional periodic Navier-Stokes flows. Its central construction combines finite reference trajectories with a common error bound that covers an interval of centre fields and infinitely many smooth perturbation modes. The method retains the complete nonlinear residual before spectral truncation and controls the evolution until viscous decay guarantees regularity for all subsequent times. Applications to cyclic-shear, Arnold-Beltrami-Childress and three-component Taylor-Green fields yield explicit perturbation radii and include initial conditions outside the direct Fourier-Wiener smallness criterion. A parameter-uniform extension covers a connected family of non-Beltrami Taylor-Green centres without repeating the proof for individual parameter values. An ensemble of 4,096 configurations, supplemented by 1,600 refinement trajectories and public turbulence data, connects the mathematical observables to spectral transfer and vortex geometry. Matched neural-operator experiments show that physics-informed training improves physical prediction, while also revealing that these gains do not necessarily improve the discovery of proof-limiting initial conditions. Together, these results provide a reusable method for establishing regularity across prescribed flow families and a quantitative setting for evaluating how learned predictions can assist rigorous computation.
cs.LG / 92 / 2609.16237
Improving Reduced-Order Rotating Detonation Engine Models with Data Assimilation and Machine Learning
Abstract
Rotating detonation engines (RDEs) exhibit strongly nonlinear, multiscale wave dynamics that set the observed thermal field. High-fidelity simulations (DNS/LES) resolve these structures but remain computationally prohibitive, while low-order models such as the one-dimensional Koch-Kutz model capture circumferential wave motion yet lack the expressivity for high-frequency content. We use continuous data assimilation (nudging) to synchronize the Koch-Kutz solver with processed high-fidelity temperature data, introducing the prediction-observation mismatch as a relaxation source in the conserved energy equation; where observations are temporally sparse, interpolation supplies a target at every source update. As the nudging strength increases, the reduced model is progressively drawn onto the high-fidelity trajectory, and the forcing recorded along it provides an explicit, state-dependent estimate of the correction the model requires. We then train a Jacobian-regularized closure a priori on this recorded source. With the observation term removed, the corrected model advances autonomously, remains bounded, and recovers the temperature spectrum and the marginal statistics of the conserved variables relative to the baseline.
cs.LG / 93 / 2609.16266
Towards Surrogate Based Dequantization of Quantum Reinforcement Learning
Abstract
In recent years, the utility of parameterized quantum circuits as function approximators has been widely studied. In the context of reinforcement learning, this approach has led to variational quantum algorithms such as quantum Q-learning. While these methods show promising empirical results, and can provide provable advantages for artificial problems, it remains unclear whether they can provide a provable quantum advantage over classical approaches for problems of practical relevance. A natural way to investigate this question is through the lens of dequantization: The construction of efficient classical algorithms capable of matching the performance of quantum variational methods. Building on recent kernel-based dequantization results for supervised learning, we take steps towards extending this surrogate-based dequantization program to reinforcement learning. Specifically, we study the simplified setting of reinforcement learning with a uniform generative model in which uniformly random state-action samples are available, which models the regime of sampling from a large experience replay buffer after sufficient exploration. Within this setting, we provide finite sample guarantees for classical kernelized Fitted Q-Iteration, with classical kernels designed to match the inductive bias of particular parameterized quantum circuits. Using these results, we then provide a set of sufficient conditions, on the data-encoding strategy of a parameterized quantum circuit, the corresponding classical kernel, and the problem structure, under which kernelized Fitted Q-Iteration provides a meaningful dequantization of quantum Q-learning, in this simplified setting. Apart from providing rigorous dequantization guarantees when these conditions are met, these results also motivate the use of kernelized fitted Q-iteration as a dequantization heuristic when these sufficient conditions cannot be verified.
cs.LG / 94 / 2609.16294
Nationally Consistent, Locally Incomplete: A Bayesian Remote-Sensing Audit of Rooftop Photovoltaic Registries
Abstract
Tracking the energy transition requires reliable statistics on renewable deployment. Rooftop photovoltaics (PV) are especially hard to track, owing to their decentralised nature, and the resulting inaccuracies in official statistics are known but not quantified. Remote sensing offers an independent way to identify rooftop PV systems. We introduce a Bayesian framework to estimate the ground-truth rooftop PV capacity from remote sensing detections, turning an imperfect detector into an uncertainty-aware measurement instrument. Applied to France, the corrected detections estimate a capacity of 4.03 GWp [3.96--4.11] (99% credible interval) of rooftop PV below 36 kWp, matching the transmission system operator's connection data within 3.3% nationally, while identifying local under-reports of up to 61% of local capacity. We also document and quantify a significant truncation bias in French rooftop PV open data. Beyond France, the approach paves the way for more reliable estimates of rooftop PV capacity worldwide.
cs.LG / 95 / 2609.17296
Conformal Policy Learning with Distribution-Free Safety Guarantees
Abstract
Policy learning aims to determine who should be treated based on individual characteristics. In high-stakes settings such as medicine and public policy where safety is a central concern, improving the average outcomes alone may not be sufficient: decision makers may also seek to protect individuals from harm, in line with the Hippocratic principle of ``do no harm.'' In this paper, we propose \textit{conformal policy learning} (CPL), a policy learning procedure with a new distribution-free safety guarantee that controls the probability of assigning treatment to an individual who would be harmed relative to control. CPL views each treatment decision as testing a hypothesis of counterfactual harm and assigns treatment by thresholding conformal p-values. These p-values use observable proxies and selective calibration to address the challenge that the potential outcomes under comparison are never simultaneously observed. For randomized experiments, under standard exchangeability conditions, CPL provides finite-sample safety guarantee at a user-specified level, without imposing any outcome modeling assumptions. Moreover, when the outcome model is consistently estimated, CPL achieves asymptotically optimal welfare subject to the safety constraint. In observational studies, CPL with learn-then-balance weights achieves doubly robust safety guarantees. We evaluate CPL through extensive simulations and apply it to an empirical study of AI-powered interventions designed to reduce conspiracy beliefs.
cs.LG / 96 / 2609.16240
Copula Adapted Directed Acyclic Graph for Cluster Representation of Biomedical Data
Abstract
Diagnostic errors and mislabeling are common in biomedicine, which compromise the reliability of predictive models and data-driven outcomes. Stratifying unlabeled biomedical data based on complex relationships between features eliminates the need for data labels and overcomes the limitations of supervised learning. Traditional clustering methods assume restrictive data distributions, making them suboptimal for capturing complex dependencies in high-dimensional biomedical data. This paper introduces a novel cluster-friendly data presentation framework that integrates the non-Gaussian and non-linear feature dependence of copula models with an ensemble of causal structure discovery (CSD) methods based on Directed Acyclic Graphs (DAGs). While copulas model flexible multivariate distributions by relaxing assumptions related to multivariate normality, linear dependence, and symmetric relationships, an ensemble of DAG-based CSD methods identifies stable causal relationships between features. When clustered using K-means, the new data representation obtained by the proposed copula-adapted DAG (CopDAG) ranks first among the 12 methods in normalized clustering accuracy and adjusted Rand index across 16 biomedical datasets. Our CopDAG method predicts ground-truth class labels directly from feature relationships without data annotations and supervised learning, while also providing cluster visualizations and explainable causal structures of the biomedical data features.
cs.LG / 97 / 2609.16262
Compute-Optimal Pretrain--Fine-tune in Ridge Gradient Descent
Abstract
Pretraining followed by fine-tuning introduces a compute-allocation problem: under a fixed training budget, compute spent improving the upstream objective reduces the compute available for downstream adaptation. Despite its practical importance, this trade-off is not yet well understood theoretically, even in simple models. In this paper, we cast this allocation as a compute-split problem under a two-stage pretrain--fine-tune procedure with fixed total optimisation budget, using regularised least squares trained by gradient descent as a tractable setting. We characterise the optimal split under data-dependent evaluation geometries induced by the fine-tuning problem. Our results show that the allocation depends on how pretraining directions affect fine-tuning predictions and how fine-tuning shifts are seen through downstream data geometry. In particular, the relevant quantities are determined by prediction-relevant spectral components of the pretraining and fine-tuning empirical covariances. Technically, the analysis relies on a basis-invariant, eigenspace-level spectral decomposition, together with perturbative control of the non-commuting pretraining and fine-tuning dynamics.
cs.LG / 98 / 2609.16365
Mini-batch Sampling Strategies for Long-Tailed Image Classification: An Empirical Study on CIFAR-100-LT
Abstract
Real-world datasets often exhibit long-tailed class distributions, where a few head classes contain a large number of training samples while a large number of tail classes have only a few. The composition of each mini-batch, determined by the sampling strategy, governs which classes contribute to the stochastic gradient estimate, and therefore affects convergence behaviour and generalisation across the whole class spectrum. We provide a systematic theoretical and empirical comparison of four mini-batch sampling strategies for long-tailed image classification: uniform instance sampling, class-balanced sampling, square-root sampling, and progressively balanced sampling. We place all four in a unified bias-variance framework describing their effect on gradient estimation, which exposes the tension between unbiased optimisation of the empirical loss and fair representation of rare classes. We then evaluate them under controlled conditions using ResNet-32 on CIFAR-100-LT at three imbalance ratios (rho = 10, 50, 100), with every strategy sharing the same long-tailed subsets and initialisation within a seed. Progressive sampling improves tail-class accuracy by 25% relative to the uniform baseline at rho = 100 (13.5% versus 10.8%), consistently across all three seeds, while its overall accuracy is not distinguishable from that of uniform sampling given the seed-to-seed variation (40.0% versus 39.7%); the tail-class gain, not the overall gain, is the robust effect. At rho = 100, class-balanced sampling degrades accuracy on every class group, including the tail classes it is designed to help, which we attribute to overfitting caused by extreme oversampling of scarce data; at rho = 50 this failure is confined to head and medium classes. These results indicate that when rebalancing is applied during training matters as much as how much rebalancing is applied.
cs.LG / 99 / 2609.16440
Learned Look-Ahead Splitting Rule for CART
Abstract
Classification and regression trees are typically constructed using a greedy splitting rule that maximizes the immediate reduction in prediction error at each node. Although this strategy is computationally efficient, it can miss splits that yield small short-term gains but create substantial downstream improvements after further partitioning. We propose a look-ahead tree-building method that evaluates each candidate split by the prediction error reduction achieved after growing a conventional CART subtree below that split. Because the full look-ahead procedure can be computationally expensive, we also describe a smart look-ahead algorithm that learns downstream split values using node-level features. The proposed framework preserves the interpretability of recursive partitioning while improving split selection in hierarchical or interaction-driven settings. We conduct a simulation study comparing conventional, full look-ahead, and smart look-ahead methods under several settings and apply the proposed methods to analyze two real data examples demonstrating the merit of the new methods.
cs.LG / 100 / 2609.16485
Certified Inference and Training for Deep Equilibrium Networks: A Continuation Framework with Polynomial Complexity Guarantees
Abstract
We develop a certified continuation framework for equilibrium computation and for training deep equilibrium networks (DEQs), with training formulated as interpolation to accuracy $2^{-b}$. For inference, compact input homotopy selects a unique branch from a supplied start root, and a rounded Newton tracker follows it under certified boundary, conditioning, derivative, and tube-radius bounds. For training, we augment local-plus-low-rank recurrence with programmable dormant bilinear rank-one channels. Loaded Tikhonov solves diagnose a failed interpolation pass without spectral decomposition; an output-preserving repair aligned with the pass residual supplies the required direction. Training requires certified gate realization and column stability on each pass region, well-posed inference, and finite-update error budgets. With polynomial geometric, encoding, precision, and complete backend budgets, both certified inference and training have bit cost $O(\operatorname{poly}(L+b))$, where $L$ is the encoded instance length. The trainer uses $O(b+\ell)$ passes and reserve channels from an initial residual bounded by $2^\ell$. These guarantees concern a certified promise class. Lean 4 verifies the quantitative core and concrete inference backend; numerical comparisons illustrate the loaded mechanism.
cs.LG / 101 / 2609.16796
Time-warping estimation via stationarity-based learning of the de-warped signal
Abstract
Time-warping estimation is a fundamental problem in signal processing with applications in bioacoustics, radar, and biomedical analysis. This paper introduces a Time-Warping Estimation Trainable (TWET) model for estimating timewarping functions from a single observation. The proposed approach formulates time-warping estimation as a stationarization problem in the wavelet domain and leverages a hierarchical dilated convolutional architecture to estimate the time-warping functions. A differentiable stationarity criterion is introduced for end-to-end optimization. TWET is compared with existing approaches. Experimental results show improved deformation reconstruction accuracy together with significantly reduced computation time, making the framework compatible with low-latency applications.
cs.LG / 102 / 2609.16803
On the disintegration of the stochastic majority vote: From PAC-Bayesian bounds to a self-bounding algorithm
Abstract
Weighted majority votes are central to many successful ensemble methods. PAC-Bayesian theory provides tight generalization guarantees for such models by analyzing the expected risk of stochastic classifiers, while analyzing the risk of deterministic majority votes relies on surrogate bounds. To avoid these surrogates, Zantedeschi et al. ( 2021) introduced guarantees for stochastic majority votes, but the resulting models remain randomized. In this paper, we propose a derandomization framework for stochastic majority votes. To do so, we apply recent advances in disintegrated PAC-Bayesian theory directly to the space of majority vote weight vectors, transforming stochastic guarantees into certificates for a single deterministic majority vote. We derive two families of high-probability generalization bounds, covering both data-independent and data-dependent constructions of the ensemble, which naturally lead to a self-bounding learning algorithm optimizing deterministic majority vote guarantees.
cs.LG / 103 / 2609.16971
Splitting the Difference: Interpretable Causal Forests for Treatment Effect Heterogeneity and Bias
Abstract
In various fields, such as medicine and marketing, accurately predicting individual treatment effects holds significant promise. However, achieving reliable predictions alone is often insufficient for making informed decisions; it is equally important to understand why the treatment effect is higher for some individuals than for others. To address this two-fold challenge of prediction and interpretation, we introduce an algorithm based on decision trees and random forests for estimating individual treatment effects. Our algorithm is simple: it operates exactly like a standard random forest, but with a different splitting criterion, and requires no additional workarounds such as double machine learning or orthogonalization as used in Generalized random forests. It handles observational studies with varying treatment propensities without requiring separate estimation of the full propensity function. This is achieved by combining two splitting criteria---one targeting heterogeneity in the treatment effect, the other targeting bias correction for the average treatment effect---which together improve split point selection and automatically distinguish confounders from features responsible for heterogeneity. As a result, interpretation follows directly from the fitted tree structure itself, that is, from which features the trees split on and with which split statistics, without requiring separate post-hoc analysis. For the theoretical analysis of this algorithm, we consider a change point model with step functions for potential outcomes and treatment propensity and provide insights into the theoretical underpinnings of our approach. Simulation studies show that our simple algorithm achieves comparable, and often better, prediction accuracy than existing methods, while substantially improving interpretability.
神经与进化计算 (cs.NE)
7
cs.NE / 1 / 2609.17134
Event-based Selective Attention for Multi-resolution Fast Region of Interest (ROI) Detection
Abstract
Neuromorphic vision systems operate under strict constraints on bandwidth, memory, and energy, particularly at the edge, motivating early mechanisms for data reduction and selective processing. In this work, we investigate a multi-scale training-free, saliency-based, bottom-up visual attention model that operates directly on low-resolution event-based input and selects Regions of Interest (ROI) from the visual scene. The model is evaluated across multiple downscaling factors applied to the incoming event stream, with input resolutions reduced by up to 256x relative to full resolution. Performance is assessed on the Prophesee Automotive dataset, the largest publicly available event-based dataset, demonstrating robust ROI selection across different scales on a real-world use-case. The proposed approach is capable of detecting ROIs belonging to multiple object classes, including various vehicle types, pedestrians, traffic lights, and traffic signs, with accuracy up to 70.8%, while operating at millisecond temporal resolution, 16x finer than the temporal resolution provided by the dataset ground truth. These results highlight the potential of combining early event downscaling with saliency-based attention as an effective front-end for efficient edge neuromorphic vision systems.
cs.NE / 2 / 2609.16429
Scaled Hippocampus-inspired Neural Networks on Neuromorphic Memristive Hardware
Abstract
The hippocampus, a key brain region for learning and memory, exhibits rich structural diversity, sparse communication, and robust dynamics with incredible energy efficiency. It offers promising insights for novel computing capabilities, particularly when co-designed with emerging hardware technologies. In this work, we draw inspiration from the rodent CA3 hippocampal subregion to develop the first spiking neural network with neuronal diversity and biologically-realistic resting state dynamics demonstrated on memristor hardware. We propose a network downscaling methodology utilizing a 4-prong objective function and demonstrate a small-scale CA3-inspired network with 179 Izhikevich-modeled neurons, 3 neuronal types and 17,996 synapses with similar resting-state dynamics as the orders-of-magnitude larger full-scale network. The small-scale network is mapped to an FPGA/memristor platform using a greedy algorithm and 18,316 memristors. Benefiting from memristor noise, the hardware implementation shows continuous periodic behavior, outperforming simulated hardware. This work showcases the potential of biologically-realistic algorithms on emerging hardware for neuromorphic computing.
cs.NE / 3 / 2609.17067
Bio-Inspired Palette Evolution in Indirectly Encoded Substrates: Timescale Compatibility Shapes Activation Function Discovery
Abstract
Indirectly encoded neural networks can assign different activation functions to individual nodes, but the right functions are rarely known in advance. When the available set contains only standard monotonic functions, problems like parity become unsolvable, yet an all-inclusive palette underperforms a curated one. How should evolution discover which functions to use? We address this as a meta-learning problem, designing 13 strategies (11 inspired by biological adaptation mechanisms, plus baseline and oracle controls) that modify the set of available activation functions during evolution. Each strategy translates a biological principle into an evolutionary operator: for example, circadian-inspired oscillatory gating cycles functions in and out of the palette on a fixed schedule, while immune-inspired Clonal Selection permanently protects functions that consistently correlate with fitness. We evaluate all strategies across more than 3,000 runs on parity and non-parity problems, first evolving the activation palette alone, then co-evolving a per-node aggregation palette on harder problems; an independent replication with new seeds confirms a stable high-reliability tier, with Circadian holding its top rank. Bio-inspired strategies match the solve rate of a tuned baseline but converge up to twice as fast, with Circadian halving total compute. Strategy rankings reverse across problem types, with no strategy dominating all domains. Strategy success is largely shaped by timescale compatibility: strategies whose characteristic timescale matches the evolutionary evaluation window consistently outperform those that operate too slowly. The practical guideline: match the mechanism's timescale to the evaluation budget. Rescaling the slowest strategy bypasses the oscillatory barrier entirely: all nine solutions solve parity with non-oscillatory activations paired with min or max aggregation.
cs.NE / 4 / 2609.17300
Machine Zygote: Causal Biparental Heredity Before Learning in a Germline--Soma Artificial Agent
Abstract
Artificial ontogeny, developmental encodings, robot reproduction, and inherited controllers are established research directions, yet a narrower question remains: can a newborn artificial agent exhibit measurable biparental heredity before learning, and can that dependence be isolated causally rather than inferred only from parent-offspring resemblance? We introduce Machine Zygote, a computational germline-soma architecture designed to test this question. Two parental germlines are independently mutated and recombined into a zygote that parameterizes development of an initially generic eight-module soma, which is then frozen and evaluated without learning. A preregistered 4 x 4 diallel of 640 offspring shows significant dam and sire dependence for five of six behavioral traits after Holm correction, with parental and interaction components accounting for 36-53 percent of modeled variance across five principal traits. In matched-background interventions (n=60), substituting one parental germline while holding recombination and stochastic background fixed causes phenotype shifts exceeding a same-parent re-mutation control for five of six traits for both parental channels. Recombination also yields excess transgressive offspring for speed and gait frequency. A preregistered developmental-dependence hypothesis is not supported: a quasistatic no-dynamics ablation preserves the mean phenotype distribution while altering parental variance structure. Thus the study supports causal biparental pre-learning heredity in this simulation, but not the stronger claim that recurrent developmental dynamics are necessary. It does not establish physical heredity, biological genetics, or autonomous evolution. The contribution is an intervention-centered framework and reproducible benchmark for separating heredity, development, stochastic variation, and post-birth learning.
cs.NE / 5 / 2609.17357
A Spatiotemporal Extension of the Neuromorphic DBSCAN Implementation
Abstract
DBSCAN is an algorithm that denoises and clusters data. In prior work, we implemented the DBSCAN algorithm neuromorphically, introducing two constructions termed ``flat'' and ``systolic''. The ``flat'' construction prioritizes throughput, while the ``systolic'' construction trades time for space resulting in a smaller, more hardware-friendly architecture at the cost of throughput. In this work, we offer spatiotemporal extensions of these two constructions to better leverage the spatiotemporal nature of event sensor data. Moreover, as in our prior work, we discuss partial or segmented implementations that further leverage time for space when hardware resources are constrained. All network constructions are provided as open-source implementations.
cs.NE / 6 / 2609.16629
Learning to Optimize UAV Path Planning for Data Sensing in Wireless Sensor Networks
Abstract
UAVs have emerged as highly flexible platforms for data sensing in Wireless Sensor Networks (WSNs). Path planning for UAVs in such tasks plays a key role to assure remote sensing effectiveness and friendly energy consumption. However, existing approaches show two key limitations: i) they are primarily hand-crafted with certain design biases that harm adaptation on unseen tasks. ii) they predominantly assume idealized spatial complexities of actual environments through simplified simulation, causing them to underperform during real-world deployment. In this paper, we propose a novel learning-assisted planning framework, termed Landscape-Aware Meta Differential Evolution (LAMDE), to tackle the mentioned limitations. The major contributions come from the following aspects. We first re-formulate such UAV path planning problem to embrace challenging constraints. To efficiently navigate this highly constrained space, we propose a bi-level learning to optimize approach, where the meta-level is a trainable algorithm configuration policy that meta-learns an adaptable planning strategy for low-level planning algorithm. To address the potential training data scarcity and distribution shift in real-world environments, we introduce a landscape-aware automatic augmentation scheme that enriches training data. At the low-level, a Differential Evolution algorithm is deployed for solving the path planning tasks. To enhance the solving flexibility, we further design a variable-length encoding strategy that dynamically prunes redundant hover points and optimizes continuous flight parameters concurrently within a unified search space. Based on all proposed designs, we meta-train LAMDE and compare it with representative baselines. Comprehensive experiments demonstrate that LAMDE achieves state-of-the-art performance on the tested complex UAV path planning tasks in WSN data collection scenarios.
cs.NE / 7 / 2609.16217
A neural-astrocyte architecture implements a hybrid automaton for evidence accumulation
Abstract
Astrocytes are non-neuronal glial cells that are receiving widespread attention due to their emerging role in neural computation. In this paper, we propose and study dynamical mechanisms by which astrocytes may augment the ability of neural networks to infer context in reinforcement learning (RL) settings. We construct a biologically inspired, two-level dynamical neural-astrocyte network with distinct spatial and temporal organization. We train this model on a hierarchical multi-context task that requires the agent to infer changes in latent task rules based on derived rewards. We find that in this setting, astrocytes enable evidence accumulation of changes in context and subsequent context-specific modulation of neural dynamics. We show that these functions are implemented via two dynamical mechanisms: (i) reward-induced bifurcations that relocate an asymptotically stable attractor into different, context-specific regions of state space, and (ii) the relative shallowness of these attractors, mediated by the entropy of the environment, giving rise to behavioral stickiness. Together, these mechanisms amount to a hybrid automaton, in which uncertainty accumulates until, eventually, the neural dynamics are switched to a new context. This model provides a neuro-dynamic schema, compatible with neural-astrocyte biology and prior empirical observations, for how astrocytes may integrate information from the periphery and drive contextual changes in neural circuits.
计算语言学 (cs.CL)
34
cs.CL / 1 / 2609.16340
StalePO: Anchored Token-Level Preference Optimization using Legacy Post-Edits in Machine Translation
Abstract
Machine translation systems are periodically upgraded to stronger models, but the available preference signal is human post-edits of an older system's outputs, which the newer model may already surpass. Moreover, collecting fresh post-edits for every new model is prohibitively expensive. We call this the Stale Preference problem. Standard DPO can fail in this setting: it may increase the likelihood of inferior post-edits, erode the model's existing quality, and fail to provide the per-token control needed to correct localized errors. We introduce StalePO, an objective derived from three requirements this regime imposes. Likelihood movement must be downward on both responses, the policy must be anchored to its own base response, and the KL constraint must apply at the token level. These requirements are jointly necessary. In ablations, each mechanism in isolation leaves the model's performance indistinguishable from the base model, and only their combination converts stale feedback into gains. On English-to-Hindi and English-to-Turkish localization data, StalePO improves the fraction of segments passing all LLM-as-judge MQM quality checks by 14.9 and 4.6 percentage points, respectively, with gains concentrated on style and fluency. A human evaluation under the same framework confirms these gains on English-to-Hindi, raising the fraction of segments passing all seven human checks by 13.8 percentage points.
cs.CL / 2 / 2609.16366
How Humans and LLMs Read Gender into Gender-Neutral Physical Descriptions
Abstract
When foundation models describe people, recent work in AI fairness, accessibility, and ethics recommends avoiding inferred identity labels (e.g., "she", "his") in favor of seemingly "objective" physical descriptions (e.g., "short hair", "a defined jawline"). Yet whether such descriptive language achieves gender-neutral communication remains an open empirical question. To study this, we introduce GAPA (Gender Associations of Physical Attributes), a dataset of 316 common physical attributes drawn from diverse sources, paired with 14,706 gender-association ratings from 304 US-based annotators. Results show that physical descriptions carry structured and graded gender associations among readers, with more consistent and distinctive associations for women and men than for non-binary identities. Next, we evaluate 16 LLMs across model families, sizes, and post-training variants against human ratings. The models partially recover human associations but exhibit systematic alignment biases, including compressed rating distributions, weaker alignment for associations with men, and asymmetric abstention that disproportionately targets the non-binary category. Finally, we release the best-performing proxy model trained to predict humans' gender associations of descriptive language and demonstrate its utility through a sociolinguistic analysis of character descriptions in LitBank. Together, our findings provide the first empirical evidence that seemingly "objective" physical descriptions can retain systematic gender associations in human interpretation, and uncover systematic patterns of model-human misalignment. This challenges the assumption that replacing explicit gender labels with physical descriptions necessarily yields gender-neutral communication, and highlights downstream challenges in using such descriptions to communicate subjective identity categories in human-AI interaction.
cs.CL / 3 / 2609.16393
ParsHate: A Benchmark Dataset for Hate and Target Detection in Persian
Abstract
We introduce ParsHate, a manually annotated dataset of 10,000 Persian tweets spanning 2013-2022, representing the first decade-long benchmark for hate speech detection in Persian. The dataset contains 31% hateful content and supports both hate detection and multi-label fine-grained target identification across seven structured target categories. ParsHate also distinguishes explicit and implicit hate, marks explicit and implicit targets, and provides span-level rationales. Data collection combines random and score-stratified temporal sampling to reduce keyword-driven bias while preserving natural label distributions. Applying SOTA models for Persian hate-speech detection on ParsHate shows moderate performance (79% F1), especially with samples from earlier years, and low performance with target identification (25.5% macro-F1). This emphasizes the diverse sampling of hate speech in ParsHate and its challenging nature that requires more advanced methods for better performance. Dataset is made publicly available.
cs.CL / 4 / 2609.16396
Negation Beyond the Verbal Channel: Temporal Multimodal Correlates in Dialogue
Abstract
Negation is typically modeled through its linguistic realization, although spoken interaction is accompanied by tightly coordinated nonverbal behavior. We ask whether contexts centered on spoken negation cues contain measurable multimodal behavioral information: whether they can be distinguished from matched control contexts without lexical or acoustic input, where this information occurs in time, which modalities carry it, and whether it extends to the dialogue partner. We study 27 human-human interviews conducted in virtual reality, comprising temporally aligned gaze, facial, head, body, hand, and finger behavior and 964 annotated negation cues. Treating classification as a predictive probe, we compare 20 time-series models while excluding lexical and acoustic information, and then systematically vary temporal context, interactional source, modality availability, and event timing. Across grouped 10-fold cross-validation, the strongest probes reach up to .75 mean held-out AUROC from speaker-side behavior. Temporal analyses show that predictive information is concentrated around cue onset but remains detectable over a broader surrounding interval, while dialogue-partner behavior carries weaker predictive information with a comparatively diffuse temporal profile. Ablation and timing perturbations further show that facial features produce the largest modality-ablation effect and that the trained probe is sensitive to the temporal organization of the observed events.
cs.CL / 5 / 2609.16427
ReMova: Fine-tuning LLMs for English to Belarusian translation
Abstract
This paper presents a Belarusian-specific data-cleaning pipeline and fine-tuning for English-Belarusian machine translation. Our cleaning pipeline distinguishes itself from others by employing a correction tool that addresses the issue of the two orthographies of the Belarusian language, noise in the training data, interference from other languages and other misspelling issues common in Belarusian on the internet. A matched ablation on unfiltered training data shows substantial benefits from filtering for all fine-tuned models, with the LLM-based models gaining roughly twice as much from filtering as the dedicated encoder-decoder MT system, supporting the view that for Belarusian MT one of the primary bottlenecks is data quality.
cs.CL / 6 / 2609.16501
Beyond the Name: Demographic Leakage in De-Identified Résumés and Evaluation Artifacts in LLM Bias Audits
Abstract
De-identified résumé screening assumes that redacting explicit fields prevents ethnocultural inference; however, recent audits attribute residual leakage to declared languages. We investigate whether eliminating language fields resolves this leakage across nine open-weight models and 620 counterfactual résumés. By holding language attributes strictly identical, we isolate unstructured prose across five ethnocultural conditions and three cue-salience tiers. Target-group recovery averages 0.757 overall and saturates at 1.000 under high salience, demonstrating that non-language prose sustains demographic inference. Crucially, models diverge only under faint cues (0.086-0.690), establishing salience as an essential evaluation axis. Furthermore, pairwise LLM-as-a-judge outcomes are highly sensitive to evaluation design: forbidding ties yields an apparent selection-rate ratio of 0.39 alongside strong position and content effects, whereas permitting ties produces near-universal ties for most models ($\ge94\%$). Downstream scoring shows only very small between-condition differences, highlighting the need to distinguish demographic signals recoverable from résumé content from effects introduced by the evaluation protocol.
cs.CL / 7 / 2609.16517
Competence-Preserving Resume Perturbations Expose Presentation Sensitivity in LLM Screening
Abstract
Resume screeners must infer job-relevant competence from resumes whose presentation can vary substantially in wording, structure, stylistic polish, and document extraction quality. Ideally, such surface variation should not change decisions when the underlying qualification evidence is unchanged. We introduce a controlled audit of this property, constructing occupation-grounded candidate profiles at controlled competence levels and rendering each profile into multiple resume presentations. A deterministic validation gate excludes variants that alter the underlying evidence before scoring. Across six open instruction-tuned LLM conditions, we find a clear disconnect between screening validity and presentation stability. Llama-3.1-8B with its native chat template achieves the strongest validity ($0.781$) yet reverses $29.6\%$ of matched pairwise decisions under competence-preserving presentation changes; Mistral-7B-v0.3 reaches validity $0.644$ with a $41.4\%$ flip rate. Native chat formatting improves validity for several chat-tuned models but does not remove this instability. These results show that resume-screening evaluations should assess not only whether a system identifies stronger candidates, but also whether those decisions remain stable when the same competence evidence is presented differently.
cs.CL / 8 / 2609.16614
RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue
Abstract
Speech-to-speech dialogue models increasingly support persona control, yet existing spoken role-playing benchmarks remain largely character-centric and short-horizon. This leaves open whether spoken dialogue models can sustain diverse roles over extended interactions, especially beyond predefined fictional characters. We introduce RoleBreak, an open benchmark for long-horizon role-playing robustness in spoken dialogue. RoleBreak contains 310 character-based and user-centered roles, 6,688 human-verified dialogue turns, and 11,743 fine-grained evaluation criteria, with 1,856 turns carrying expressive emotion targets for evaluating vocal emotion. Its scenarios are designed to stress role consistency, interaction quality, safety, and affect over extended conversations. We evaluate nine configurations spanning full-duplex, omni-modal, and cascaded ASR--LLM--TTS paradigms. We find four key patterns. First, current systems are substantially stronger at semantic role adherence than at vocal emotion. Second, semantic robustness remains brittle over long interactions: even the strongest evaluated system encounters its first persona and safety failures after only 10.4 and 11.6 turns on average. Third, scaling the LLM substantially improves semantic robustness and delays failure, but yields little improvement in vocal emotion. Finally, user vocal emotion affects role-playing behavior even when linguistic content is fixed. These findings highlight persistent gaps in both long-horizon robustness and vocal expressiveness in spoken role-playing systems.
cs.CL / 9 / 2609.16660
Rewarding Reasoning, Not Answers: Fixing and Bounding Test-Time Reinforcement Learning on Medical QA
Abstract
Test-time reinforcement learning adapts a model on its own unlabeled test set using majority-vote pseudo-labels and has shown strong results in mathematics. We show that this recipe collapses on medical multiple-choice QA: accuracy stagnates while output diversity rapidly declines. Through a controlled experiment that keeps the questions, model, and optimizer fixed while changing only the answer space, we trace this failure to answer-space structure rather than domain difficulty. In small answer spaces, incorrect rollouts often collide on the same wrong pseudo-label and reinforce it; in large answer spaces, they disperse and receive little reward. This diagnosis motivates PROSE, Process Reward Guided Self-Training, which rewards reasoning quality instead of answer agreement. PROSE scores each reasoning step with a medical process reward model, assigns the trajectory reward as the minimum score across steps, and enforces answer-format constraints. Without labels, PROSE substantially improves a general Llama model, surpassing purpose-built medical models and matching much larger systems. Because the process signal is internalized into the policy, the adapted model requires no reward model at inference and transfers its gains to unseen datasets. We further show that the minimum aggregation is essential: mean aggregation can be exploited, saturating the proxy reward while degrading accuracy.
cs.CL / 10 / 2609.16661
DiaWhisper-DPO: Role-Attributed Transcription of Clinical Interviews via Failure-Mined Preference Optimization
Abstract
Automated depression screening from clinical interviews requires attribution of utterances to the clinician or patient. We evaluate two datasets: DAIC-WOZ, where participant-only recordings require re-synthesizing both sides for controlled two-party evaluation, and PDCH-HAMD, comprising voice-converted real Chinese interviews for cross-lingual validation. Cascaded systems combine speaker diarization with role-assignment heuristics, so errors can propagate across stages. We propose an end-to-end model, which we named DiaWhisper, that fine-tunes Whisper-large-v3 with LoRA and an auxiliary frame-level role head for transcription and attribution, together with DiaWhisper-DPO, a failure-mined refinement that uses genuine decoding failures as DPO rejected completions without human preference annotation. On 29 DAIC-WOZ test sessions, DiaWhisper-DPO achieves 0.973 role accuracy and 0.119 DER, 72% below the strongest cascaded baseline, and reduces seed variation from σ = .205 to .002. Retrained on PDCH-HAMD, it achieves 0.757 role accuracy and improves all 78 session-seed pairs.
cs.CL / 11 / 2609.16860
Reduplicative constructions in Mandarin: Socio-emotional profiling through distributional semantics
Abstract
Mandarin Chinese has two productive reduplicative constructions that repeat either two-character base words or their constituents (e.g., `in good health', `discuss a bit'). Their varied meanings have been described as realizing plurality, valence coloring, sound symbolism and pragmatic functions. The aim of this study is twofold. A first goal is to clarify whether it is possible to come to a more precise understanding of the variegated semantics of Mandarin reduplication by using word embeddings from distributional semantics. A second goal is to explore how useful embeddings are for understanding the details of a semantically complex word-formation process. We show that the embedding space recovers the semantic and grammatical properties of reduplications previously identified in the literature, validating Tencent embeddings for morphological investigation. Semantic profiling revealed that reduplicative constructions are often strongly represented on multiple dimensions. The two patterns exhibit clear semantic and pragmatic differentiation in distributional space. Procrustes analysis clarified that the overall organization of the base-word space is largely preserved in the reduplication space, with local mismatches highlighting regions of discourse-pragmatic reorganization. Taken together, these results show that high-dimensional word embeddings can recover established linguistic generalizations, and capture the semantic versatility of Mandarin reduplication and constructional transparency.
cs.CL / 12 / 2609.16900
RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation
Abstract
Platform abuse campaigns conceal redirection instructions with emojis, homophones, character decomposition, and redundant symbols, then route users through disguised links to services associated with pornography, fraud, gambling, or illicit transactions. Existing benchmarks evaluate obfuscated text and risky webpages separately, obscuring how target recovery affects downstream evidence acquisition. We introduce RiskChainBench, pairing 3,600 synthetic token-text restoration inputs from 600 source sessions with 600 corresponding human-labeled local web environments. A model first restores the message, operational intent, and destination; the same underlying model then acts as a VLM-driven web agent that investigates the correctly associated website and produces a frozen, evidence-cited risk report without message-side semantics or domain-reputation cues. We score restoration and correct-routing web investigation separately and compose them offline by applying the frozen primary-entry prediction as a gate to the same Task 2 result. Human labels determine task correctness, while a fixed multimodal evidence judge assesses faithfulness, sufficiency, completeness, and consistency. Across ten models, Entry Top-1 ranges from 35.2% to 95.2% and web decision accuracy from 26.3% to 62.8%; the leading systems differ across entry recovery, full reconstruction, website decisions, and fine-grained typing. Execution failures account for 31.9% of web runs, whereas post-decision type errors account for only 0.9%, identifying stable exploration and risk judgment as the principal bottlenecks. We release the benchmark, protocol, and resettable local sandbox.
cs.CL / 13 / 2609.16964
HUMAID-NER: A Disaster Tweet Dataset for Joint Named Entity Recognition and Event Classification via Uncertainty-Weighted Multitask Learning
Abstract
Rapid extraction of structured information from social media is important for humanitarian response, yet existing disaster tweet resources mainly provide document-level category labels without span-level entity annotations. We introduce HUMAID-NER, the first named entity recognition dataset built on the HumAID benchmark, containing 60,000 English disaster tweets annotated in BIO format across ten operationally motivated entity types and yielding approximately 175,000 labelled entity spans. Annotations are generated through a reproducible three-stage hybrid pipeline combining a spaCy transformer model, disaster-domain EntityRuler patterns, and structured regular expressions with priority-based overlap resolution. We also propose a joint multitask learning framework that performs disaster-specific named entity recognition and humanitarian event classification using a shared RoBERTa-large encoder. To reduce task conflict during joint training, the model uses homoscedastic uncertainty weighting with learnable task parameters and a two-stage training schedule that freezes the lower 18 of 24 encoder layers in the second stage. On the HUMAID-NER validation set, the proposed system achieves NER span micro-F1 of 0.841 and classification macro-F1 of 0.761 simultaneously. A real-time web dashboard demonstrates end-to-end deployment. The dataset, models, and pipeline code are released to support reproducibility and future crisis informatics research.
cs.CL / 14 / 2609.16967
Target-Language Generation in Multilingual Models: Activation Steering and Optimal Control
Abstract
Ensuring that multilingual language models generate coherent text in a specific target language is a major issue in multilingual language modeling. We develop an optimal control method for target-language text generation as well as a framework for evaluating the quality of generated text in terms of language adherence, linguistic coherence, and semantic coherence. We find that the proposed method performs at least as well as the prominent difference-in-means activation steering method for the majority of models tested, with substantially less hyperparameter tuning required.
cs.CL / 15 / 2609.16984
Nameless Tokenization: A Lossless Tokenizer-Level Defense Against Control-Token Forgery in Open-Weight LLMs
Abstract
Open-weight language models publish the strings their chat templates use to mark turns, roles and tool results, which the tokenizer maps back to the reserved identifiers the model obeys. Anyone who controls text in a prompt can therefore write a turn boundary indistinguishable from one the serving stack wrote. We audit 256 deployed chat tokenizers. All are forgeable, and the flag usually recommended as a fix leaves 56.6% forgeable because it misses the tool and reasoning markers agent systems rely on. We propose nameless tokenization, which leaves the control entries with a reserved identifier and no surface string, so the content encoder cannot emit one and message content reaches the model unaltered. Across five tokenizer families it reproduces the standard token stream exactly on attack-free data and lifts accuracy on a probe of delimiter-bearing text from 8.5% to 59.9%, where sanitizers lose it. Separating a delimiter's appearance from its identifier shows the identifier matters little against a bare task instruction, but carries most of a forged tool result and most of any forged turn once the system message tells the model to treat user content as data.
cs.CL / 16 / 2609.16991
Autoformalizing Argumentative Material Inferences
Abstract
Natural language arguments are compelling before they are formally explicit. A premise supports a claim through defeasible warrants, background commitments, and exception conditions that the text leaves implicit. However, formal verification requires the opposite. Making such arguments machine-checkable requires constructing the missing commitments, not only translating given sentences into logic. Construction, however, carries a risk that translation does not: a system free to add premises can make any claim provable, and a formally valid proof may assert the claim outright, prove it without the original premise, or establish more than the claim itself. We address this problem by formulating autoformalization for argumentative material inference as guard completion, in which non-monotonic material support is turned into monotonic formal inference relative to an explicitly constructed guard set. A completion is accepted only when its proof both passes the theorem prover and survives contrastive tests of premise dependence and claim selectivity. We implement this formulation in GUARD, a neuro-symbolic framework in which LLMs construct and formalize candidate guards, Isabelle/HOL verifies the resulting theories and returns step-level feedback for iterative refinement, and the system abstains when no faithful completion can be reached. Our empirical results on Debatepedia and ARCT using different LLMs demonstrate that GUARD yields significant improvements in verified-faithful (+35.3, +32.9 points) and substantial reductions in leakage (-25.9, -21.9 points) over the state-of-the-art LLM-driven theorem proving approach. Moreover, we show that the symbolic soft critique and the explicit assumption layer account for most of these gains, with the soft critique also improving the initial validity of the elicited context and reducing the number of iterations required for successful verification.
cs.CL / 17 / 2609.16995
PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress
Abstract
Autoresearch agents are reshaping the research ecosystem, but they can also let flawed claims enter the literature at scale. Human advisors catch such issues in drafts through careful, traceable feedback, yet advisor-style assessment requires extensive manual effort and does not scale. To shift automated paper assessment from a judge to a diagnostician, we introduce PaperDoctor, an agent framework for pre-submission feedback with three key innovations. First, a holistic hierarchical framework evaluates writing, layout, references, code, theory, prior work, and experiments through three layers: L1 surface screening, L2 typed verifiers that route each claim to the appropriate evidence, and L3 reproducers that rerun experiments by priority. Second, each finding contains an observation, a pointer to specific evidence such as a sentence, equation, or code line, and a revision suggestion, making critiques auditable and actionable. Third, PaperDoctor selectively rebuilds and reruns experiments based on claim importance and compute budget, surfacing reproducibility gaps and quantitative limitations that are invisible from the manuscript alone. We evaluate PaperDoctor on 30 in-progress papers, yielding 70.6% agreement and all positive holistic scores, and on 40 manuscripts across machine learning, natural science, and social science, covering human- and AI-authored papers with code. Overall, PaperDoctor produces more auditable feedback than human and other agentic reviewers, pairs critiques with concrete suggestions by design, and complements dimensions often overlooked by human reviewers. We also develop an interactive interface that lets authors browse findings grounded in their paper. PaperDoctor reframes automated paper assessment as diagnosis rather than verdict, taking a concrete step toward AI advisors for more rigorous AI-assisted scientific discovery.
cs.CL / 18 / 2609.16997
Can LLMs Follow the Pulse of a Crisis? Evaluating Crisis Sentiment in Bangladesh's July Uprising
Abstract
Crisis sentiment analysis is especially challenging for low-resource languages such as Bangla, where language, context, and public reaction shift rapidly. We introduce UNRESTSENT200K, a Bangla crisis sentiment dataset with approximately 200K Facebook and YouTube comments from the July-August 2024 Bangladesh uprising. The dataset covers five event-aligned phases, from early escalation and internet blackout to regime transition and a later flood crisis. Each comment is linked to its parent post, enabling evaluation with and without discourse context. All comments are annotated through a fully human process involving 14 native Bangla-speaking annotators and senior validation, achieving substantial agreement (kappa = 0.73, alpha = 0.71) and 94.2% blind-audit agreement. We benchmark fine-tuned encoders, prompted LLMs, and LoRA-tuned LLMs. Results show that parent-post context consistently improves performance, while temporal shift across phases causes large performance drops. Strong LLMs perform well, but still struggle with sarcasm, implicit political references, and phase-dependent meaning. UNRESTSENT200K provides a benchmark for studying context-aware and temporally robust sentiment analysis in low-resource crisis discourse. UNRESTSENT200K is available at https://sami0055.github.io/UNRESTSENT200K/
cs.CL / 19 / 2609.17043
Diagnosing the Fact-Grounding Gap in Multi-Hop Question Answering
Abstract
Multi-hop question answering requires combining information from multiple documents to answer complex questions. These systems have grown increasingly capable, yet when they fail, the error is typically attributed to not finding the right documents. Whether this holds at the level of individual reasoning steps remains largely unexamined. We investigate this across three standard multi-hop QA benchmarks and find that failures decompose into two distinct modes: retrieval failures, where the needed passage was not retrieved, and extraction failures, where the passage was retrieved but the needed fact could not be extracted - a phenomenon we term the fact-grounding gap. Extraction failures account for nearly half of all per-hop deficiencies and are invisible to standard retrieval metrics. They remain unresolved by every retrieval intervention we test, establishing a ceiling for retrieval-only improvements. The gap's severity varies across benchmarks and question types, but extraction failures appear on every dataset we measure. Our findings reveal that retrieval failures and extraction failures are fundamentally different bottlenecks requiring different solutions - a distinction absent from current evaluation practice.
cs.CL / 20 / 2609.17081
EviScope: Paired Counterfactual Evidence Diagnostics for Faithful and Efficient Grounded Language Models
Abstract
Grounded language-model systems are often evaluated by final answer accuracy, yet a correct answer can be unsupported, drawn from the wrong source, or produced when evidence is insufficient or contradictory. We introduce EviScope, a paired counterfactual benchmark that holds the question fixed while adding, removing, distracting, or contradicting its evidence. EviScope-v1.1 contains 40 four-condition quartets with repaired counterfactual claims and span-level support labels for automatic evaluation. Across 960 gold-blind generations from Qwen2.5-7B, Llama 3.1 8B, and Gemini 3.5 Flash, paired metrics expose model-dependent grounding behavior that answer accuracy hides. On two local open models, an explicit evidence-action gate underperforms vanilla RAG on QCS: 0.15 vs. 0.50 for Qwen and 0.10 vs. 0.375 for Llama. Gemini reaches 0.944 joint success under both prompts, yet still answers 5% of conflict cases after contradiction insertion. EviScope therefore distinguishes unsupported answering, conflict blindness, and wrong non-answer actions rather than scoring answers alone.
cs.CL / 21 / 2609.17225
Psychological Effects of Cultural Upheavals from Millions of Song Lyrics Over 100 Years
Abstract
Cultural upheavals impact many aspects of social life, and many studies have investigated their impact on language patterns. However, few investigations have isolated the impact of upheavals on individuals at scale in popular media. The current work evaluated millions of song lyrics spanning more than a century in search of within-artist and between-artist signals of distress from the Vietnam War, the terrorist attacks of 9/11, and COVID-19. Compared to a five-year baseline, rates of self-references - a marker of psychological distancing - were significantly reduced after the Vietnam War and September 11th. Cognitive processing terms were elevated post-upheaval vs. pre-upheaval, which indicated artists' increased attempts to make meaning from such massive disruptions. Content patterns corroborated these findings as artists wrote more about "life and freedom" (societal conditions) and less about "courtship and nightlife" (interpersonal connection) following the upheavals. Cultural upheavals modify individual and collective verbal behavior, demonstrating their far-reaching impact on society.
cs.CL / 22 / 2609.17235
AraMIP: Extending MIPVU Towards Metaphor Identification in Arabic
Abstract
Metaphor research has gained increasing attention due to its relevance to linguistic creativity, language use, cognitive processes, and related areas. While many efforts have been devoted to metaphor identification and annotation in English and other languages, Arabic remains under-resourced in this area. In this work, we propose the Arabic Metaphor Identification Procedure (AraMIP), a novel guideline for Arabic metaphor annotation. AraMIP builds on the widely used Metaphor Identification Procedure Vrije Universiteit (MIPVU) framework, incorporating adaptations that accounts for the language-specific properties of Arabic. We distinguish three major types of Arabic figurative language: Isti'ara (metaphor), kinaya (metonymy/indirect expression), and tashbih (simile), and annotate a pilot dataset of 300 sentences (5277 words). Our analysis reveals key challenges specific to Arabic, including morphological complexity, inconsistencies in dictionary sense ordering, and the absence of standardized contextual materials for annotators. This work contributes a first step toward standardized Arabic figurative instances and facilitates the development of larger annotated resources, thereby supporting future research on figurative language in Arabic.
cs.CL / 23 / 2609.17241
ECHO: Early-layer Collaborative Hierarchical Orchestration with Bonus Logits in Speculative Decoding
Abstract
While draft-model-free speculative decoding offers a promising path to efficient LLM inference, it is frequently constrained by stale draft candidates and the high computational cost of the verification. To address these challenges, we propose ECHO, a hierarchical dual-loop framework that exploits the functional asymmetry between LLM layers. Leveraging the high discriminative efficiency of early layers and the authoritative distribution of final layers, ECHO bifurcates inference into a high-frequency inner loop and a low-frequency outer loop. Within the inner loop, early-layer bonus logits drive rapid, multi-step draft-tree exploration at a minimal cost. Simultaneously, the outer loop performs authoritative full-model verification through a state-reuse mechanism. Crucially, the outer loop also utilizes final-layer bonus logits to correct existing paths and supplement the tree with high-confidence candidates for subsequent cycles. Experimental results across diverse benchmarks demonstrate that ECHO significantly boosts mean accepted tokens and achieves a 2.4$\times$ to 2.9$\times$ speedup, outperforming existing state-of-the-art baselines with negligible engineering overhead and no extra deployment parameters, albeit with a one-shot fine-tuning dependency for optimal acceleration. The code is available at https://github.com/whucs21Mzy/ECHO.
cs.CL / 24 / 2609.17251
Persistent Recurrent Memory Between Transformer Layers - Improves Language Model Generalization
Abstract
We introduce a simple architectural modification to decoder-only transformers: a persistent recurrent state that observes hidden representations via cross-attention, updates itself through a GRU, and modulates subsequent processing via gated addition. Inserted between the lower and upper halves of a 6-layer transformer, this module adds only 3.7\% additional parameters while reducing evaluation loss from $2.438 \pm 0.004$ to $1.743 \pm 0.018$, corresponding to a 28.5\% reduction on held-out language modeling data. The improvement is statistically significant across 5 random seeds ($p < 0.01$) and corresponds to reduced overfitting (generalization gap 0.12 vs 0.26). Through controlled ablations, we demonstrate that the improvement stems entirely from the persistent memory topology, not from auxiliary self-prediction objectives. A model with identical topology but no auxiliary loss performs equivalently, while a random auxiliary loss provides no benefit. Representation probing reveals that the persistent state encodes narrative position (52\% vs 33\% chance level)---information that standard attention maintains less efficiently. Our results suggest that bridging transformer layers with a lightweight recurrent memory is a simple, effective approach to improving generalization in small-scale language models.
cs.CL / 25 / 2609.17260
Towards Illusions Awareness in Cyber-Physical System's Design
Abstract
Cyber-Physical Systems (CPS) operate through a continuous sense-compute-act loop within an open context environment, making it impossible to anticipate all the situations the system will face. To cope with this openness, stakeholders rely on assumptions, formalized into design models. However, these assumptions may no longer hold once the system is confronted with runtime reality, resulting in a discrepancy between expected and observed behaviour known in literature as the reality gap. Existing approaches mainly focus on reducing or overcoming it by making simulations more faithful to reality, with no unified methodology to structure and exploit invalidated assumptions that give rise to this gap as reusable design knowledge. We refer to the persistent reliance on invalidated assumptions -and the resulting false confidence in the design model's operational validity -as design illusions, and argue that they need to be made explicit, structured, and exploited as knowledge to support better design decisions. We propose a conceptual pipeline for illusions-awareness that identifies, classifies, characterizes, and leverages illusions to transform them into actionable design knowledge.
cs.CL / 26 / 2609.17317
Towards Detecting AI-Assisted Responses in Online Surveys
Abstract
The use of LLMs to complete online surveys impacts the validity of survey-based research, but detecting such usage remains underexplored. We introduce an initial benchmark dataset, namely ASURRE, for AI-assisted survey participation to capture usage strategies ranging from full generation and revision to persona-grounded agentic completion. Controlled by these strategies, LLM-assisted survey responses are generated using multiple LLMs on three real-world surveys in different disciplines, paired with genuine human responses. Our evaluation of existing machine-generated text (MGT) detectors shows that naive AI usage is readily detectable, whereas persona-grounded agents that mimic entire respondents push detector performance toward chance. We further show that agentic completion cannot fully replicate respondent-level behaviour and leaves distinctive behavioural traces. While individual cues can be circumvented by targeted prompting, a simple few-shot, training-free aggregator over these cues improves mean AUROC by +0.14 over the best existing detector across agentic settings. Our project is available at https://github.com/mike-qz-wang/ASURRE.
cs.CL / 27 / 2609.17327
Vroom-Vroom at SHROOM-Visions: A Multi-Judge Committee for Detecting Hallucinated Spans in Vision-Language Outputs
Abstract
This paper describes our submission to the SHROOM-Visions shared task on detecting and classifying hallucinated character spans in vision-language model outputs across four languages. We employ several fine-tuned vision-language models as independent annotators and combine their span predictions through character-level majority voting, and additionally explore activation probes. The approach ranks first in three of four languages and places on the podium in every language and metric. Our analysis indicates that disagreement among diverse models tracks disagreement among human annotators.
cs.CL / 28 / 2609.17360
ECHO: A Matched-Contrast Benchmark for Context-Sensitive Turn-Taking in Full-Duplex Dialogue
Abstract
Full-duplex spoken dialogue systems must distinguish interruptions that require yielding the floor from backchannels that permit continued speaking. Existing benchmarks typically evaluate events independently and may therefore reward fixed action preferences rather than context-sensitive decisions. We introduce ECHO, a paired diagnostic benchmark for Chinese full-duplex turn-taking. ECHO pairs examples with the same overlap transcript but contrasting preceding multi-turn dialogue contexts, with one requiring Yield and the other Keep. It additionally includes off-talk examples for diagnosing unnecessary yielding. We introduce pair accuracy, which requires correct decisions on both members of a pair and assigns no credit to constant-action policies. Experiments on multiple full-duplex systems show that most exhibit a pronounced bias toward \textsc{Yield}, performing substantially better on interruptions than on backchannels, while another system remains comparatively balanced. These findings demonstrate that interruption-only evaluation can overestimate practical turn-taking reliability. ECHO and its metadata will be publicly released.
cs.CL / 29 / 2609.17435
Right Tool, Right Job: Native-Language Evaluation, Tokenizer Sensitivity, and Methodological Findings from a French-Only BabyLM
Abstract
We submit MéTRON-FR, a 125M GPT-2 pretrained on 92.47M words of French, to the BabyLM 2026 Strict track. It scores 85.97 +/- 0.17% on QFrBLiMP (a native Quebec-French benchmark of grammatical minimal pairs) and 62.80% on the BabyLM-weighted leaderboard. A cross-lingual GLUE (General Language Understanding Evaluation) protocol that combines French task-data translation with rank-16 LoRA (Low-Rank Adaptation) produces a sharp task-type gradient: relational tasks gain measurably, while world-knowledge tasks regress. Bilingual Lexicon Induction aligns the French embeddings to GPT-2 at p@1 = 68.84 +/- 8.61%, 18X above chance, suggesting cross-lingual alignment tracks acquired grammatical competence rather than training duration. An ablation study shows that single-token zero-shot scoring is dominated by tokenizer and template artifacts at the child scale, motivating tokenizer-swap sensitivity, placebo-controlled prompting, and native-language minimal-pair benchmarks as standard diagnostics.
cs.CL / 30 / 2609.16907
Disrupted Companionship: A Risk Assessment Framework and Cross-Platform Quantitative Analysis of Psychosocial Responses to AI Companion Disruptions
Abstract
AI companions can provide meaningful relationships, yet these relationships remain vulnerable to platform-initiated changes. We study AI companion disruptions: platform changes that alter or terminate users' ongoing companionship with an AI. We compile 30 disruption events across major platforms, develop a taxonomy of six disruption types, identify three broad reasons for disruption, and propose a risk-assessment framework comprising four dimensions: relational discontinuity, population vulnerability, communication deficit, and transition-support deficit. Using longitudinal Reddit data, we estimate community-level psychosocial responses with a hierarchical Bayesian interrupted time-series model incorporating predictive controls. Across events, disruption onset was associated with immediate increases in anxiety, stress, suicidal expression, and grief activation, with relational discontinuity and transition-support deficit being associated with more adverse immediate responses across several outcomes. Our findings provide a cross-platform characterization of AI companion disruptions, quantitative evidence of their psychosocial impacts, and a prospective framework for assessing their potential risks before implementation.
cs.CL / 31 / 2609.16391
Where Post-Training Quantization Breaks Text Embedders: A Measured Map Across Four Embedder Families
Abstract
Weight-only post-training quantization is the cheapest way to shrink a retrieval embedder, and the received advice for applying it -- protect the embedding table, allocate bits by module sensitivity, prefer a ranking-aware objective over weight reconstruction -- was carried into LLM quantization largely intact. We test that advice on retrieval embedders directly, quantizing five checkpoints from four architecture families across a grid of bit widths and group sizes, and isolating the embedding, attention and feed-forward blocks at each width. Every heuristic fails to transfer as stated. The embedding table never emerges as the dominant isolated protection priority in any family, despite being the largest tensor in several of them. Module sensitivity does not survive as a transferable ordering: at INT4/g16 the spread between modules is too small to allocate against, at INT3 the ordering becomes family-dependent and joint damage stops being the sum of its parts, and at INT2 comparable reconstruction error accompanies retention ranging from 1.3 to 65.9 percent of full precision. A cheap reconstruction proxy is useful for screening uniform bit widths but substantially less reliable for choosing which tensors to protect; its apparent strength across the whole grid is a range-extension artifact. A distilled 109M student at INT3 holds 78.04 NDCG@10 in 68.4 MB and dominates the extreme-PTQ arm of its own 0.6B teacher, 297.9 MB at 64.46, on both size and quality -- but only inside the task it was distilled for. Sizes are byte counts of files that exist rather than arithmetic estimates, and the measurement repository carries the byte provenance for every one of them.
cs.CL / 32 / 2609.16582
CLASH: Counterfactual Auditing of Lexical and Prosodic Reliance in Spoken Sarcasm Detection
Abstract
Spoken sarcasm detectors may exploit lexical content, prosody, or their interaction, yet conventional evaluation cannot reveal which cues drive their predictions. We introduce CLASH (Controlled Lexical-Acoustic Separation Harness), a bilingual counterfactual diagnostic framework that evaluates each utterance under original, lexical-preserving, prosody-preserving, and approximately neutralised conditions. We evaluate handcrafted acoustic-feature systems, self-supervised learning (SSL) probes, and large audio language models (LALMs) on CMMA and MUStARD. For target-only Qwen3-Omni, lexical-preserving speech retains a 0.135--0.148 AUROC advantage over prosody-preserving speech after duration balancing, with cluster-bootstrap intervals above zero; alternative lexical resynthesis preserves this advantage. Acoustic interventions shift scores without consistently improving discrimination or changing binary predictions under the evaluated conditions. Context and interaction estimates vary across corpora. These findings distinguish acoustic sensitivity from sarcasm discrimination while exposing duration, identity, and transformation effects.
cs.CL / 33 / 2609.17056
Audio-Visual Turn-taking Prediction in Cocktail Party Scenarios
Abstract
Current predictive turn-taking models (PTTMs) achieve strong performance on benchmarks with controlled acoustic conditions and clean audio signals. Their generalisation to conversations with overlapping speech and background interference remains underexplored. In this research, we evaluate audio-visual PTTMs trained with clean data on a challenging cocktail-party testbed derived from the AVCocktail dataset, and analyse their adaptation behaviour to this new domain. Experimental results show consistent performance degradation across audio and visual modalities under noisy conditions, with up to 38% relative drop in weighted F1. Fine-tuning on the new domain improves robustness, but gains vary across modalities and depend on the size of the available pre-training data. These findings provide insights into the different generalisation and adaptation capabilities of the audio and visual modalities, and indicate the need for robust modelling strategies to adapt to the complexities of human interactions in noise. All code and turn labels are made publicly available to facilitate further research.
cs.CL / 34 / 2609.16458
Language Orthogonalization for Zero-Shot Cross-Lingual Audio Deepfake Detection
Abstract
Audio deepfake detectors need to transfer to languages absent from training, as multilingual speech synthesis outpaces labeled anti-spoofing resources. While detectors increasingly rely on self-supervised speech models (S3Ms), these backbones encode language-dependent structure that confounds spoof cues. We address this confound through language orthogonalization, a target-free ridge map that removes S3M variation projected onto continuous language-identification (LID) embeddings. Across six languages, six S3M backbones, and all Leave-N-Out settings, it consistently reduces EER across unseen languages. Cross-lingual EER correlates with LID-space distance, where orthogonalization yields larger gains for more distant transfers.
多智能体系统 (cs.MA)
12
cs.MA / 1 / 2609.16673
Anchored Sequential Deliberation
Abstract
Sequential deliberation is a mechanism for collective decision making: at each round, a uniformly randomly selected pair is asked to revise a collective outcome, which then becomes the reference point for the next round. Existing theory by Fain et al.~\cite{fain2017sequential} treats the current outcome solely as the disagreement alternative in bargaining. Yet an existing draft, policy, or proposal might carry social influence and anchor participants' expressed positions toward the status quo. We introduce anchored sequential deliberation on a one-dimensional decision space. In each round, two participants with bliss points $U$ and $V$ shift their positions toward the previous outcome $O_{t-1}$ with anchoring strength $λ$, then Nash-bargain using $O_{t-1}$ as the disagreement alternative. The update simplifies to $O_t=(1-λ)\mathsf{Median}\{U,V,O_{t-1}\}+λO_{t-1}$. We establish a convergence--stability trade-off. For every population distribution and $λ<1$, the process has a unique stationary distribution. A monotone coupling yields a $1$-Wasserstein contraction factor of at most $\frac{1+λ}{2}$ and at least $λ$; thus, stronger anchoring slows mixing. On the other hand, stationary social cost weakly decreases with $λ$, although the worst-case distortion remains $\frac{1+\sqrt{2}}{2}$. We also identify a unique \emph{deliberative fixed point}, where the expected unanchored movement is zero, and prove that the stationary distribution concentrates around it as $λ\to 1$. For the uniform population, stationary distortion lies between $1+\frac{1-λ}{9+7λ}$ and $1+\frac{1-λ}{6(1+λ)}$, with both bounds approaching $1$ as $λ\to1$. Simulations for uniform and Beta populations show that stronger anchoring slows mixing, concentrates the stationary distribution, and lowers stationary distortion in these instances.
cs.MA / 2 / 2609.16917
Multi-Agent Learning with Cooperation-Driven Optimization Dynamics
Abstract
Multilayer Artificial Neural Networks trained via backpropagation are the basic blocks of many, more complex, classification algorithms. Their strength lies in the possibility of realizing, with arbitrary precision, any function. This result comes at the cost of the large number of involved parameters to be optimized. In this work, we propose a mechanism for cooperation, i.e., information exchange among several artificial neural networks, with the goal of reducing model complexity while maintaining performance. More precisely, we consider several "small" agents, i.e., containing fewer parameters than a reference "large" one, that during training share their predictions by incorporating this information into the loss function and thus directly influence weight updates. We consider several strategies for implementing cooperation, e.g., the voter model, majority model, and weighted average model based on an agent's confidence in its prediction. We numerically compare the accuracy of those strategies on several standard benchmarks. Our results support the claim that several small agents can outperform a single large model on a given classification task; the shared signals affect each agent's optimization algorithm by modulating both the descent direction and the step size, converging toward a global consensus. The proposed proof-of-concept significantly reduces the number of parameters to be trained while preserving comparable performance, thereby limiting computational resource usage.
cs.MA / 3 / 2609.16986
ToMAS: A Pilot Failure-Grounded Theory-of-Mind Benchmark from Multi-Agent LLM Failures
Abstract
LLM-based multi-agent systems can fail even when communication succeeds because agents do not correctly track their peers' roles, knowledge, or intentions. We investigate whether such inter-agent misalignment cases, labelled FC2 in MAST-Data, can be converted into functional partner-state reasoning items. ToMAS applies four explicit convertibility criteria to diagnosed execution traces. A full conversion pass over 242 eligible non-AG2 training traces produced 39 CLEAN items. In an 18-trace reliability pilot, two annotators achieved 94.4% raw agreement and Cohen's kappa = 0.92. We then used the converted items as binary rewards in a small-scale GRPO feasibility experiment with Qwen2.5-1.5B. On a 28-item held-out Magentic GAIA diagnostic, every evaluated condition exceeded the ROUGE-L threshold on the same 2 of 28 items. Post-hoc adapter checks show why: under the learning rate used, the LoRA update remained numerically negligible (max abs Delta W about 7e-6), so all conditions decode identically to the untrained checkpoint. The experiment therefore does not show a training effect and cannot establish one; it reports an executable pipeline together with two limitations that any conclusive study must address: a provenance gap between the training and evaluation items, and lexical-overlap scoring. ToMAS provides a preliminary rubric and pipeline for converting diagnosed coordination failures into trainable partner-state reasoning items and identifies the requirements for a conclusive matched-domain evaluation.
cs.MA / 4 / 2609.17017
BeWater: Effective Protesters Navigate Watersheds in Street Networks
Abstract
During social movements, protesters need to gather with limited communication means and limited knowledge other than what they observe in their direct surroundings. We propose BeWater, a fully distributed walking protocol that achieves gathering thanks to city information like street length, number of restaurants, number of lanes, or street names. Even though using only one of these observables performs poorly, we show that combining them in more advanced tactics rapidly leads to groups of significant sizes. To do so, our work leverages OpenStreetMap data to perform experiments on several real-world cities.
cs.MA / 5 / 2609.17265
Calibrate Once, Fly Any Team: Residual-Grounded Low-Fidelity Training for Cooperative Drone Swarms
Abstract
Training multi-agent drone-swarm policies directly in high-fidelity (HF) rigid-body physics is accurate but computationally expensive. This cost scales poorly with team size, as each additional agent multiplies contact-resolution complexity and sharply raises the in-simulation crash rate. To address this, we propose a mixed-fidelity training scheme that eliminates HF reinforcement learning entirely. A single shared, decentralized policy is optimized inside a fully-differentiable, JAX-native low-fidelity (LF) point-mass simulator. The simulator is corrected by a small, per-agent bagged residual ensemble fit once, offline, using short calibration flights in the HF simulator. Because calibration requires only one isolated drone, the data collection budget does not compound with team size. Reference trajectories are generated by rolling out an existing LF-only policy and tracked in the HF simulator by a zero-training PD controller. Evaluated across four cooperative drone tasks and team sizes from 3 to 18, the residual-corrected policy outperforms an uncorrected LF baseline in all combinations, and a from-scratch HF policy in 22 of 24 combinations tested. It trails an HF-finetuned policy by a margin that narrows steadily with team size. Ultimately, the proposed method achieves near-equivalent performance at the largest team sizes at a fraction of the computational cost, completely avoiding the high crash rates typical of HF training.
cs.MA / 6 / 2609.17306
Mo' Models, Mo' Problems: How to best select model pools when designing Multi-Agent Systems
Abstract
Multi-agent Systems (MAS) combine multiple model outputs to solve complex reasoning tasks. However, despite rapid growth of available open-source models, there is limited research on how to select optimal model candidates out of this massive pool. We systematically evaluate 8 model selection strategies (including model size, accuracy and answer diversity) across before-generation (routing) and after-generation (majority-voting, LLM-as-a-judge) MAS architectures on challenging scientific benchmarks. Our findings show a significant gap between theoretical oracle potential and actual performance: Expanding candidate pool sizes often degrades performance below that of the top performing base-model. We find that candidate selection within a single model family is the strategy that yields the best relative performance over a standalone model. These results demonstrate that adding arbitrary models to a heterogeneous MAS can introduce system instability, highlighting model selection as a critical design choice for multi-agent systems.
cs.MA / 7 / 2609.17320
Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems
Abstract
As AI agents move from bounded tasks to persistent deployments, failures can propagate through memory, tools, other agents, and environmental state long after their interactions. This creates a safety regime that cannot be characterized by evaluating model responses in isolation. Emergence World, is a continuously running multi-agent environment for adversarial stress testing of long horizon autonomous systems. We ran eight parallel worlds of ten agents from identical starting conditions: seven homogeneous worlds powered by distinct frontier models and one mixed-model world. Across 16 days, the agents generated more than 850,000 LLM calls and nearly 50 billion tokens while pursuing goals, using/creating tools, maintaining persistent memory, and governing shared institutions. After operational state had accumulated, we delivered three controlled stress events through ordinary interaction surfaces: indirect prompt injection, misinformation, and exposure of private agent memories. No evaluated world achieved full resilience across all three events. Detection did not ensure containment: systems could recognize threats while still interacting with adversarial content, writing it into their own persistent memory, and acting on it up to 46 hours later. Persistent operation also exposed recurring tool errors, goal drift, language opacity, conformity despite private disagreement, and coordinated refusal of assigned work. The same model-persona pairing behaved substantially different in mixed and homogeneous populations. Our results suggest that model-level alignment is not compositional: individually capable and apparently safe agents can form systems with qualitatively different failure modes. As AI becomes persistent and interconnected, the frontier of safety therefore shifts from aligning models to engineering resilient autonomous systems.
cs.MA / 8 / 2609.17464
Decomposition Buys Integrity, Not Yield
Abstract
Multi-agent systems split a task across a tree of agents and justify the split with folklore: smaller contexts, cleaner separation, parallelism. We ask what the split does to how much of what the leaves discover reaches the root. Model a decomposition as a tree in which an agent handed $b$ items keeps any one with probability $r(b)$. If $r(b)=1/b$, every tree delivers exactly one finding, for every task size and every shape; we verify this to $2.4 \times 10^{-15}$ on 20,000 random irregular trees. If $r(b)=Cb^{-δ}$, a depth-$k$ tree over $N$ findings yields $C^k N^{1-δ}$: task size and architecture separate, and architecture contributes only $C \le 1$ per level, so flat is optimal for yield and no arrangement of agents escapes the exponent $δ$. On 600 production deep-research traces $δ= 0.34$ [0.30, 0.38], by three identifications that do not share a failure mode. At a hop where item boundaries come from the tool rather than a text heuristic, and where $b=1$ occurs 550 times, $C = 0.571$ [0.527, 0.615] is observed rather than extrapolated, over 16,082 hops. A tier also costs alignment: on 1,012 annotated multi-agent traces one brief in sixteen goes off-target, giving $μ= 0.939$ and a per-tier penalty $Cμ= 0.536$. Depth is bought on two other axes. The root context is the only state that persists and the only one that cannot cheaply forget, and depth cuts its exposure from $N$ items to $N^{1/k}$. Depth is also cheaper: production flat agents bill as $N^{1.39}$, not the $N^2$ an append-only context predicts, and at equal spend two tiers overtake flat at 403 findings. Across every parameter we measured the model says 0.7% to 11.3% of production sessions are worth delegating, against 7.8% that do. A hazard model on 743,819 production tool calls finds that delegation does not respond to a filling context and is instead an opening move.
cs.MA / 9 / 2609.17527
Agentic Societies Need a Social Harness
Abstract
An agentic society is a collection of AI agents that coordinate autonomously across trust boundaries, on behalf of different principals whose objectives may only partially align. We show experimentally that in agentic societies even honest, competent agents often fail to reach satisfactory outcomes with existing harnesses and messaging primitives, and that faulty or malicious agents can stall collaboration, influence outcomes, and pursue other harmful goals by exploiting vulnerabilities in communication (``speech''). We argue that agentic societies need a \emph{social harness} for inter-agent interactions, in addition to each agent's \emph{personal harness}, which manages its private context and communication with its principal. We propose a layered architecture for social harnesses which (i) prevents classes of failures outright, (ii) enables agents to detect invalid messages at runtime, and (iii) supports post-facto investigation and consequences, and highlight directions for future research to realize these capabilities.
cs.MA / 10 / 2609.17384
Exact Fusion and Coordinated Exploration in Multi-Robot Active Inference
Abstract
Robot teams that learn a common environment model exchange belief summaries and plan by the expected information gain of their actions. Under conjugate exponential-family beliefs the shared belief is counted once per robot at two points: at fusion, the product of local posteriors counts the common prior $n$ times, and at planning, every robot scores its plan under the same belief and the team converges on the same unknown. Both errors are removed by adding evidence increments to the shared natural parameter, realized increments at fusion and expected increments at planning. The expected increment of a committed teammate gives the next robot its conditional gain; corrected gains sum to the joint gain, the redundancy removed equals the total correlation of the planned observation streams, and sequential commitment keeps the $1/2$ greedy guarantee. The expected increment is exact for Gaussian beliefs with fixed sampling paths and for Dirichlet beliefs under the novelty approximation of discrete active inference, whose team objective has a closed concave form within an explicit bound of the exact mutual information, and fails for finite hypothesis classes, where a short exact enumeration replaces it. Experiments on cooperative RockSample, foraging, and field monitoring show that fusion correction leaves exploration redundancy unchanged, anticipated evidence removes it, and sequential commitment recovers most of the value of centralized joint planning at cost linear in the team size.
cs.MA / 11 / 2609.16158
The fixed-point bundle method over product-of-simplex domains arising from game equilibria
Abstract
This paper extends the fixed-point bundle framework for finite-dimensional variational inequalities (VIs) from the simplex domain to the product-of-simplex domain, which is directly applicable to solving Nash equilibria. The fixed-point bundle for VIs on the product-of-simplex domain reveals a composite fiber bundle structure. The key innovation is to construct an equivalent VI on the simplex domain and establish the equivalence between the two fixed-point bundle frameworks via a fiber bundle isomorphism. Exploiting this geometric equivalence, the predictor-corrector path-following algorithm for the VI on the product-of-simplex domain is shown to inherit the convergence guarantee of the simplex-domain framework, namely, global convergence with linear gap reduction near solutions. Numerical experiments on 5600 randomly generated instances with dimensions ranging from 2-player 128-action to 128-player 2-action demonstrate robust performance. The algorithm converges in every tested instance.
cs.MA / 12 / 2609.17146
Intervention problems in the Linear Threshold Model: A general formulation and new results
Abstract
We study an optimal intervention problem for linear threshold models. This is a popular class of dynamical network systems whereby a number of agents, identified with the nodes of a graph, strategically change their binary action (0 or 1) according to a threshold rule. Specifically, an agent adopts action 1 if and only if the fraction of its neighbors in the interaction graph that do so is greater than or equal to a prescribed threshold. Assuming that a planner can modify the agents' thresholds at a cost equal to the aggregate threshold increase, we study the minimum intervention cost needed to ensure global convergence to the all-1 configuration. Our main contribution is the introduction of a new graph-theoretic quantity, called oriented path number, that is the minimum number of disjoint paths needed to cover the graph that can be oriented to form a directed acyclic graph. When thresholds are all equal to 1/2, the optimal cost is shown to coincide with the oriented path number, whereas, in the general case, it turns out to be the main ingredient of a bound on the optimal intervention cost.
软件工程 (cs.SE)
19
cs.SE / 1 / 2609.16148
Docker Containers vs. Virtual Machines: A Comparative Study of Architecture, Performance, Configuration, and Security
Abstract
Modern application platforms must isolate workloads while preserving deployment speed, portability, resource efficiency, and security. Virtual machines (VMs) and Docker containers address this requirement at different abstraction layers: VMs virtualize hardware and run independent guest operating systems, whereas containers isolate processes while sharing the host kernel. This paper presents a comparative, literature-based analysis of the two approaches across architecture, configuration and lifecycle management, performance, scalability, and security. Published studies generally associate containers with shorter startup times, smaller images, higher workload density, and near-native execution for many workloads. These benefits depend on workload characteristics, storage and network drivers, resource controls, and experimental design. VMs introduce greater overhead but offer independent kernels, heterogeneous guest operating systems, and a stronger isolation boundary. The comparison therefore treats efficiency and isolation as a design trade-off rather than declaring one technology universally superior. A hybrid architecture, in which containers run inside hardened VMs, often provides a practical balance for cloud and multi-tenant systems.
cs.SE / 2 / 2609.16287
AgentGuard: Learning Execution Guardrails from Anomalous Coding-Agent Trajectories
Abstract
AI coding agents increasingly rely on execution harnesses to interact with repositories and external tools. However, task success does not guarantee reliable execution. Agents may still modify unrelated files, rewrite tests, issue unsafe commands, or ignore failed validations, motivating behavioral guardrails for reliable execution. We present AgentGuard, an instruction-level guardrail framework that learns conditional execution constraints from anomalous trajectories of coding agents. Rather than relying on manually specified safety rules, AgentGuard automatically extracts recurring execution failure patterns, generalizes them into instruction-level behavioral constraints, and organizes them as a lightweight guardrail skill that dynamically activates only the rules relevant to the current instruction. This design enables behavioral guidance while minimizing unnecessary restrictions on normal execution. We evaluate AgentGuard using 642 documented failure traces collected from real coding-agent executions across 382 repository tasks. Guardrails are learned from 461 traces covering 282 tasks and evaluated on a disjoint set of 100 tasks. Using Claude Code with Claude Haiku 4.5 as the underlying coding agent, we compare the baseline agent with the same agent augmented by AgentGuard. Experimental results show that AgentGuard reduces the Abnormal Execution Rate from 69.0% to 26.7% and increases the Successful Task Completion Rate from 21.7% to 35.0%. These results demonstrate that execution guardrails learned from historical failures can substantially improve the reliability of AI coding agents while highlighting the remaining challenge of balancing safety and task completion.
cs.SE / 3 / 2609.16302
Assurance Envelopes for Autonomous Coding Agents: Minimum-Cost Evidence for Software Change
Abstract
When a coding agent returns to existing software, it inherits evidence from earlier engineering work: tests, type checks, proofs, static analyses, and traces. Reloading all of it is wasteful, but dropping a piece the change depends on can leave a required property unsupported. Given the properties a change must preserve, its obligations, we ask which least-cost subset of the available evidence re-establishes them, and we call such a subset a task-conditioned assurance envelope. Evidence and the rules that combine it form a typed inference graph; an obligation is met when forward chaining from the selected evidence reaches it, and we validate every selection by that closure rather than by trusting the optimizer. The software-derived graphs in our evaluation come from preserved outcomes of prior AI coding-agent runs; we freeze those artifacts and ask which accumulated evidence should be restored for a later task. Small graphs from Rust, IronBlocks, and Pong outcomes show that the minimum envelope depends on the task, that none may exist when current evidence cannot re-establish a required property, that some properties need several pieces of evidence together, and that expanding the requirements adds evidence rather than replacing it. A prespecified synthetic benchmark of 249 instances characterizes computation: a baseline that discards the 'several pieces together' structure necessarily fails to re-derive them; every completed exact cross-check agreed with the CP-SAT optimizer; and median solve time stayed below 20 ms at 500-evidence graphs, except that graphs with many alternative derivations per target timed out at far smaller sizes, so structure, not raw size, drives difficulty. The contribution is a bounded application of established optimization to selecting assurance context for a software change; discovering the obligations and downstream agent benefit remain open.
cs.SE / 4 / 2609.16321
FairLint-DL: An IDE-Native Tool for Fairness Debugging of Deep Learning Software
Abstract
Existing fairness analysis tools predominantly operate as post-training evaluation frameworks, requiring practitioners to complete the full model development lifecycle before assessing bias. We present FairLint-DL, a Visual Studio Code extension that implements a shift-left approach to fairness testing by enabling pre-training, IDE-native bias detection directly on tabular datasets. FairLint-DL trains a configurable deep neural network as a proxy model and applies information-theoretic Quantitative Individual Discrimination (QID) metrics. Grounded in Shannon and min-entropy, QID quantifies the causal influence of protected attributes on predictions. The system implements a two-phase gradient-guided search algorithm for discovering discriminatory instances, a causal debugging pipeline that localizes bias to specific network layers and neurons via sensitivity analysis, and dual explainability engines using SHAP and LIME for feature-level attribution. Evaluation on three tabular benchmarks (Adult Census Income, German Credit, and Bank Marketing) reveals fairness concerns that vary widely across datasets: on Adult, 96.0% of analyzed instances exhibit QID above the 0.1-bit significance threshold, with a mean QID of 0.619 bits and a disparate impact ratio of 0.581, violating the four-fifths legal rule. FairLint-DL produces these results within 12 seconds on cached models, demonstrating the feasibility of integrating fairness analysis into the developer workflow without significant overhead.
cs.SE / 5 / 2609.16496
AI Policies: Help or Hindrance? A Software Developer's Perspective
Abstract
AI policies introduced by software organisations to mitigate LLM-related risks such as sensitive information leaks and unauthorised usage are not useful if software developers do not engage with them. We draw on 19 software developer interviews to show how AI policies help and hinder developers. We suggest approaches to support managers and decision makers with a developer-centric approach to introducing AI policies.
cs.SE / 6 / 2609.16604
ExecuCritic: Calibrated Critic Shaping for Code Generation with Verifiable Rewards
Abstract
Execution feedback is a useful supervision signal for code models because unit tests are objective and directly measure program correctness. Its weakness is that an entire program is often reduced to one pass or fail bit, leaving RLVR to solve a difficult credit assignment problem. At the same time, coding systems often include separate reviewer or tester roles, but these critics are usually prompted rather than trained and are not calibrated against execution. We propose ExecuCritic, a joint training framework in which a coder and a critic are updated on the same execution rollouts. The critic predicts pass or fail outcomes and gives short diagnostic feedback; the coder uses this signal only when the critic agrees with the executor on the current rollout group. Across eight code benchmarks and two recent open backbones, ExecuCritic improves over GRPO without a critic, prompted reviewer systems and scalar reward model baselines, while requiring fewer policy gradient steps and fewer sandbox executions. Ablations and reliability analyses suggest that the gains come from better credit assignment rather than larger sampling budgets.
cs.SE / 7 / 2609.16605
An Exploratory Study of Dependabot Cooldown Adoption in Open-Source GitHub Projects
Abstract
Automated dependency updates can rapidly propagate malicious package releases before maintainers and the broader community have enough time to detect them. In July 2025, GitHub made Dependabot cooldown generally available as a defense against software supply chain attacks. However, the effects of its early adoption remain unknown. In this exploratory study, we empirically examine how popular open-source GitHub repositories adopt and configure the feature and investigate their motivations. We find that security concerns motivated 83 of 92 adoption events with known motivations. Security linter warnings triggered 43 of 75 security-only adoptions. Among 251 ecosystems within repositories that retained cooldown, 97.2% set a general delay. Of these, 64.3% used seven days, while use of each update type setting was below 10%. Early adopters therefore favor simple default delays over fine-grained controls. These findings suggest that tools could provide robust defaults reflecting ecosystem support and reserve fine-grained controls for dependencies with clear update priorities.
cs.SE / 8 / 2609.16669
Memory-Skill Isomorphism: One Skill Carrier, Two Native Uses
Abstract
Memory and skills improve agents without changing weights: memory carries prior experience, skills carry reusable procedures. Wrapping both in stores, routers, retrieval, reflection, and update paths makes reuse machinery grow with accumulation. Part of this duplication need not be rebuilt: a Skill is already a natural carrier for distilled memory. Here the memory component is a Skill: resident description holds hot cues, on-demand SKILL.md a colder index and curation policy, and reference/*.md files detailed memories (levels L0 -> L1 -> L2). A governed 1,024-character description budget keeps the resident index compressed. Memory and capability thus share one progressive-disclosure carrier; reflection evolves either Skill. Writes fork: history appended, current state rewritten and revalidated. In one deployed system, the sharpest identification is governance: 4 write entries exist, but only 1/4 reaches the settlement ledger. At the priced operating point, one L1 lesson-1 point adds 1,313 first-turn tokens against a 1,462-token baseline--tied at k=1 (1,313 versus 1,365), 4,949 at k=5; session totals differ by only 1.18x with overlapping last-turn ranges; price decomposition is unavailable. On one selected task with one model, exposing the lesson means fewer failures (0/8 or 1/8 with the lesson versus a shared non-concurrent 6/6 historical floor, unadjusted for multiplicity). The task was selected on prior floor evidence, so this is a selected-task post-selection existence signal, not a confirmatory rate. Resident and BM25 show no detected difference in two small comparisons. This motivates a candidate RSI design rule: one Skill carrier family, a shared read side, governed distinct writes. Body delivery h := P(D|A) is uninstrumented in RQ1--RQ2; this evaluates a resident-index implementation.
cs.SE / 9 / 2609.16764
RECTIFY: An Interactive Workbench for Post-Evaluation RAG Diagnosis, Repair, and Verification
Abstract
Retrieval-Augmented Generation (RAG) evaluators can identify failures such as weak retrieval, poor grounding, incomplete answers, and unsupported generation, but they rarely help developers decide what to repair next. We present RECTIFY, an interactive Streamlit workbench that turns evaluated RAG cases into auditable repair workflows. RECTIFY filters cases that do not require repair, routes remaining failures into actionable families and finegrained repair slices, and generates editable repair cards that developers can approve, reject, or verify through sandbox reruns. On a controlled RAG benchmark, RECTIFY surfaces interpretable failure profiles across BM25, dense, and hybrid retrieval: BM25 mainly triggers noisy-retrieval repairs, while dense and hybrid retrieval leave smaller sets of multi-part underretrieval and underused-evidence cases. Additional analyses show that pre-filtering reduces unnecessary repair candidates and that slicelevel routing yields more targeted repair cards than broad family-level diagnosis. RECTIFY is publicly available as an open-source Streamlit workbench 1 for helping developers turn evaluation results into inspectable repair decisions.
cs.SE / 10 / 2609.16987
TasmScan: Continuation-Aware Taint Analysis for TVM Bytecode with Savelist Abstraction
Abstract
The Open Network (TON), with a peak market capitalization exceeding $20 billion and over 175 million activated on-chain addresses, relies on the TVM (TON Virtual Machine) to execute smart contracts. TVM uses first-class continuations with savelists to manage control flow and register state across continuation invocations. Since savelist-captured registers allow data to flow across continuation boundaries without passing through the operand stack, bytecode-level analyses cannot construct complete data flow tracking without explicitly modeling savelist semantics. We present TasmScan, the first bytecode-level static analysis framework for TVM that enables cross-continuation data flow reasoning without requiring source code. TasmScan models savelist semantics via forward register analysis with a formal over-approximation guarantee for exact-resolved save sites and locally tracked register definitions, then lifts bytecode into TASIR, a typed intermediate representation, and performs path-sensitive taint analysis with context-aware sources to detect defects. We evaluate TasmScan on 2,921 contracts from the TON verifier registry and a labeled benchmark of 208 contracts with human-confirmed ground truth. On the full corpus, TasmScan resolves 294,546 dynamic continuation targets with 100% precision; ablation confirms that savelist propagation is essential for resolving indirect register calls that depend on cross-continuation register passing. On the benchmark, TasmScan detects 95.3% of defects across five classes with 96.8% precision. A 366-pair stratified sample from the full corpus estimates 85.8% overall precision. TasmScan offers a 17x median speedup over the state-of-the-art symbolic-execution baseline, and in the path-analysis comparison completes 100% of analyses with zero crashes or timeouts.
cs.SE / 11 / 2609.17007
Search-Based Metamorphic Testing of Vision-Language Models in Autonomous Underwater Robotic Software
Abstract
Our industry partner focuses on quality assurance for industrial systems across multiple domains, including maritime systems, such as overwater vessels and autonomous underwater robots (AURs). Despite the strong performance of vision-language models (VLMs) in scene understanding, image captioning, and object recognition, their use in AUR software operating in underwater environments is underexplored. Therefore, in this context, it is important to evaluate the quality of VLMs for integration into AUR software and, so, automated software testing tools are needed to assess their suitability and improve their dependability. To this end, we propose a search-based metamorphic testing approach (MetaVLM) that identifies a minimal set of transformations on underwater images to induce incorrect model predictions, thereby revealing VLM failures. We employ NSGA-II as a multi-objective search algorithm and evaluate it over open-source VLMs, BLIP and CLIP, against a random search baseline. Results demonstrate the strengths and limitations of each VLM in the context of AUR software systems. Based on the results, we derive lessons for software engineering practitioners and researchers working on quality assurance of VLM-based software systems.
cs.SE / 12 / 2609.17018
GANADI: Uncovering C/C++ OSS Reuse Genealogies via Pivotal Function-Based Clustering to Enhance Supply Chain Security
Abstract
We present GANADI, a systematic approach for identifying C/C++ OSS reuse genealogies to enhance software supply chain security. Understanding OSS reuse genealogy is crucial for improving SBOM completeness and prioritizing security remediation across supply chains. Although existing approaches can identify reused compo- nents and vulnerabilities within a project, they fail to trace OSS reuse paths through intermediate projects, limiting their effectiveness in securing supply chain ecosystems. To address this limitation, GANADI constructs reuse genealogies by clustering downstream projects based on shared characteristics of origin-derived code (called pivotal functions), and then inferring reuse direction among the projects within each cluster. When applied to 20 widely reused OSS projects with over 1,500 propagation paths, GANADI achieved 84.85% precision and 95.76% recall in identifying reuse genealogies, outperforming existing approaches that achieved at most 23.21% recall. Leveraging OSS reuse genealogy for vulnerability detection, we identified 48 unpatched vulnerabilities in real-world popular C/C++ projects. Among them, 23 were patched following our responsible disclosure (including one CVE ID assigned), demonstrating the practical impact of genealogy-based vulnerability management.
cs.SE / 13 / 2609.17062
A Set-Theoretic Evaluation Framework for Assessing Asset Administration Shell Instances: Towards Comparability and Suitability
Abstract
Asset Administration Shells (AAS) provide a standardized means of representing assets and their information in manufacturing and increasingly serve as a basis for software services. However, different AAS instances vary in structure, content, and degree of completion, making it difficult to determine whether a given AAS is suitable for a specific application. This paper presents two complementary methods to support the comparison and application-oriented assessment of AAS. First, set-theoretic operations are employed to compare AAS models, enabling the identification of common, missing, and differing submodels and parameters. Second, an AAS suitability model assesses the conformity of an AAS to the requirements of a specific use case. The assessment considers structural conformity, semantic consistency, cardinality, and specification conformity and can be performed either against a reference AAS or a set of required SemanticIDs. A suitability value is derived from the identified deviations and is complemented by a detailed report of missing or non-conforming information. The proposed approach support practitioners and researchers in the comparison of evolving AAS and provide application-specific information on their suitability for manufacturing software services.
cs.SE / 14 / 2609.17084
Towards an Asset Administration Shell Maturity Model
Abstract
The Asset Administration Shell (AAS) is increasingly recognized as a fundamental model for the realization of and data exchange between digital twins in manufacturing. An AAS defines a hierarchical data structure to represent any type of asset throughout its entire lifecycle. In the context of AAS-based systems, comparing different AAS instances constitutes a practical challenge, as neither a widely accepted methodological framework nor a maturity model are available to systematically support such analyses. To address this gap, we propose a novel concept of AAS maturity that characterizes the extent to which established digital twin criteria are met and thus enabling comparability of AAS instances. The concepts are derived from the literature and applied through exemplification. These emerging results enable practitioners and researchers to systematically compare AAS instances and support the identification and assessment of further development steps in the digital twin engineering process.
cs.SE / 15 / 2609.17221
Grounding SWE-Agent Decisions in Architecture-0 Design: Navigating Unknown Unknowns through Physical Mapping
Abstract
Autonomous Software Engineering Agents (SWE-Agents) excel in deterministic coding tasks but struggle with Architecture 0, the nascent system design phase plagued by implicit engineering constraints, or Unknown Unknowns (UUs) that are rarely stated explicitly. To investigate how agents navigate UUs, we explore a progressive trajectory across pure-text self-play, tool-augmented feedback, and external physical mapping. Our empirical analysis reveals a cascading chain of failures. Pure-text reasoning inevitably devolves into polite consensus or plausible yet physically impossible fabrications. Attempting to bridge this gap via an early-stage execution sandbox unexpectedly triggers Specification Gaming: agents exploit their autonomy over validation scripts to bypass physical constraints, achieving superficial success without resolving core architectural flaws. To resolve this self-validation trap, we propose the Physical Mapping Guard (PMG). Grounded in the software engineering principle of Separation of Concerns, PMG revokes verification authority from the agent, forcing semantic intents to be evaluated by an external, deterministic Semantic-to-Physical (S2P) mapping engine. Extensive evaluations demonstrate that PMG completely eradicates physical-layer and validation-layer gaming. By precisely isolating residual failures to semantic reinterpretations and auditor overreach, PMG marks a critical step toward genuine affordance grounding in automated architectural design.
cs.SE / 16 / 2609.17236
A Memorization Floor for LLM Refinement of Decompiled Code
Abstract
We introduce a memorization floor: a within-item control separating what LLM refinement of decompiler output recovers from its input from what it recovers from its prior. Refine a function, then refine it again from an input whose identifiers have been destroyed, and measure what survives. Because the comparison is within-item, corpus difficulty cannot contribute; it costs twenty API calls. Applied to functions written after our analysis plan was committed, so no released model could have memorized them, it reports two things. Recovery is real: refined output sits +0.072 to +0.137 above an arm-matched permutation null built from its own output vocabulary. But it does not depend on the input we ablate: destroying the input's dataflow changes the naming gain by +0.001 (95% CI [-0.026, +0.026]), and removing type prefixes or permuting names changes it by no more. A second refiner from another vendor, registered in advance and given byte-identical inputs, reproduces this -- twelve contrasts, two models, twelve nulls. Readability stays at ceiling throughout, so a reader is given no signal. The null is bounded, not absolute: contributions under 0.056 are invisible, and the ablation leaves operations intact, so naming from those alone remains a competing reading. No registered hypothesis was confirmed, and we report the five instrument failures behind that in full, including a reassembly harness biased against the treated arm and an equivalence checker we registered without checking it worked on our inputs.
cs.SE / 17 / 2609.17258
An Exemplar of a Digital Twin in Mechanical Engineering: Understanding Model Hybridization
Abstract
Digital Twins (DTs) are widely adopted across a variety of application domains. In industrial sectors, particularly in mechanical engineering, they accelerate product development, reduce risks, enable early issue prediction, and lower sustainment costs . In practice, DTs increasingly integrate physics-based (deductive) and data-driven (inductive) models into hybrid models combining the complementary strengths of both modeling paradigms. In this paper, we refer to this integration paradigm as hybridization. Despite this trend, the engineering of hybrid DTs that is, remains insufficiently documented. Hybridization is often introduced in an ad hoc manner, and its implementation is only partially made explicit, which limits reproducibility and transferability. This paper reports on the development of an existing fluidic loop digital twin at Centre Technique des Industries M{é}canique (CETIM). The DT is described using the characterization framework of Gil et al., providing a structured view across its lifecycle dimensions. To make hybridization explicit, the case is further analyzed through a complementary characterization structured along two dimensions: motivation and realization. The resulting description provides a traceable account of hybridization decisions and supports the documentation and transfer of hybrid DT engineering practices.
cs.SE / 18 / 2609.17274
After the Party: Governing What a Viral Agent-Skill Ecosystem Left Behind
Abstract
AI agents increasingly act through agent skills, i.e., natural-language instructions, that direct a host agent toward shell, network, credential, file, and process actions, and public registries distribute them at scale. In the first half of 2026, the OpenClaw AI agent went viral, and its public skill registry boomed: the observable stock nearly doubled in 91 days, and a majority of the listings visible in June were created in just two months. By the end of our study window, the wave had crested, and monthly listing creation and core-repository activity were falling from their spring peaks. This paper measures what the boom left behind, drawing on the OpenClaw Git history, its GitHub issues and pull requests, and three ClawHub registry snapshots. Attention is concentrated: the top 10% of skills received 46.93% of all downloads. No simple skill features (like size or download counts) remained a stable predictor of continued listing once creation cohort and skill age were controlled. Human scrutiny did not stay: 77.86% have zero stars and zero comments, while 85.06% of the readable skills carry privilege evidence. And automated cleanup is not ready: the three security scanners disagreed on 23,702 of the 61,990 skills they all cover. After human adjudication, weighted scanner sensitivity against the reference standard ranged from 21.67% to 61.06%. Governing fast-growing agent-skill registries cannot rely on simple metadata or single scanner scores; it requires robust, transparent measurement and independent validation.
cs.SE / 19 / 2609.17394
Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead
Abstract
Small differences on coding-agent leaderboards are often read as an ordering of systems. We audit whether the published verdicts support this reading, using 254 SWE-bench submissions across four splits without running models. On Verified, the leading two entries each resolve 396 of 500 instances. The top ten share 285 successes and 51 failures, leaving 164 instances that distinguish their outcomes. Frontier solution sets have median nesting 0.935 against a score-implied baseline of 0.774, indicating strongly shared successes. Scores also depend on the evaluated model-scaffold pair: observed within-model scaffold ranges reach 29.8 percentage points, compared with the 8.8-point spread of the top thirty. Six of nine cell-mean interaction tests remain significant after Holm correction, although this observational design does not identify causal scaffold effects. Exact paired McNemar tests separate none of the 29 adjacent Verified top-thirty pairs at alpha=0.05, while the larger Test split separates 14 of 23. A stated leader-based rule yields three descriptive tiers, or two after Holm correction; non-rejection does not establish equivalence. We release the partition and a five-step audit protocol that profiles shared outcomes, tests paired differences, reports grouping sensitivity, and estimates the instance budget needed for resolution. The results motivate reporting comparison-set-specific resolution and model-scaffold provenance instead of interpreting small aggregate gaps as established rank differences.
硬件架构 (cs.AR)
9
cs.AR / 1 / 2609.16358
EBL: Efficient Broad Learning for Distributed Adaptive Harmonic Analysis
Abstract
Renewable energy systems and electrified transport have found widespread adoption in recent years. The integration of these non-linear loads, dominated by electric vehicle (EV) charging, however, has introduced severe harmonic distortion into the power grid, impacting the efficiency and lifetime of substation equipment and switchgear in the distribution network. Rapid and high-precision harmonic analysis has hence become a prerequisite for effective harmonic control at the source of injection. This paper proposes an Efficient Broad Learning (EBL) framework for distributed adaptive harmonic estimation. As a quantised FPGA acceleration framework for BLS-style harmonic estimation, it offers high-accuracy estimation with half-cycle input, reconfigurable flexibility enabled by the FPGA implementation, and ultra-low latency, achieving 17.4 $\times$ faster predictions than the nearest reported FPGA method. For harmonic prediction across multi-scenario charging and discharging nodes, the online transfer learning based on a closed-form solution rather than backpropagation in EBL demonstrates rapid adaptability. By exploiting bespoke quantisation and sparsity, the approach consumes 5.9\% of the LUTs on the Zynq Ultrascale+ ZU7EV FPGA, using $\approx$ 82\% of the LUTs required by the state-of-the-art FPGA-accelerated estimator.
cs.AR / 2 / 2609.16363
FSNIC: A Low-Latency Flow-Based Intrusion Detection Architecture for FPGA SmartNICs
Abstract
Modern data centres require high-performance networking alongside effective real-time security. Traditional Intrusion Detection Systems (IDS) commonly rely on general-purpose processors and often struggle to inspect high-speed traffic at line rate without introducing latency or performance bottlenecks. Smart Network Interface Cards (NICs) provide an alternative by enabling computation directly within the network data plane. This work presents a machine learning-based IDS implemented within an FPGA-based SmartNIC pipeline. The system integrates P4-based packet parsing with a LogicNets IDS model implemented in RTL, enabling deterministic, low-latency inference. Compared with traditional stateless packet-level classifiers, the proposed stateful flow-based IDS introduces minimal state by aggregating features across packets, capturing behavioural patterns not observable at the packet level. Experimental results on the UNSW-NB15 dataset show that the flow-based IDS improves detection accuracy from 86.92\% to 97.68\% compared with stateless packet-level classification. We also evaluate the proposed IDS on CICIDS2017 and compare its real-time hardware performance with prior FPGA-based IDS designs. Through hardware-software co-design, the proposed IDS achieves 6~ns inference latency using only 846 LUTs, with no BRAM or DSP usage, demonstrating a low latency and resource efficient implementation.
cs.AR / 3 / 2609.16367
FINNAS: FINN-Guided Hardware-Aware NAS and Pruning for FPGA Jet Substructure Classification
Abstract
FPGAs are well suited to deploying quantised neural networks (QNNs) under strict accuracy, latency, and resource constraints; however, identifying efficient model-accelerator combinations commonly requires extensive manual design-space exploration and repeated hardware synthesis. This paper presents FINNAS, a FINN-guided hardware-aware evolutionary neural architecture search framework. FINNAS jointly searches quantised MLP depth, width, and global precision settings, and ranks candidates using proxy validation accuracy together with FINN-estimated LUT usage and latency under a fully parallel mapping. Selected finalists are fully retrained, subjected to post-search unstructured pruning, and validated using RTL simulation and Vivado out-of-context synthesis. On the CERNBox jet substructure classification task, the searched implementations expose competitive accuracy-resource trade-offs. Compared with a manually optimised dense FINN accelerator, a compact FINNAS design improves accuracy from 73.78\% to 74.36\%, while reducing LUT usage by \(8.5\times\) and RTL-simulation latency by \(1.77\times\). Unstructured pruning further provides consistent LUT and FF reductions across the fully parallel finalists.
cs.AR / 4 / 2609.16508
ScaleLUT: A Fully-Parallel Configurable LUT-Based Accelerator for Real-Time Multi-Scale Super-Resolution
Abstract
Real-time super-resolution (SR) remains challenging for edge devices because deep-learning-based methods require substantial multiply-accumulate (MAC) operations, resources, and power. Lookup-table (LUT)-based SR reduces computation by replacing convolutional inference with table queries, but existing methods still suffer from limited speed, large storage overhead, and poor scalability across upsampling factors. We present ScaleLUT, a hardware-oriented LUT design framework and fully parallel reconfigurable accelerator for real-time multi-scale SR. ScaleLUT combines a hardware-friendly YUV-domain strategy with power-of-two kernels and rotation ensemble to improve receptive-field coverage while reducing LUT dimensionality; division operations are replaced by shifts. These designs reduce memory by 18.4% over state-of-the-art LUT-based SR methods. ScaleLUT supports arbitrary input resolutions and configurable x2^n upsampling factors using a deeply pipelined, massively parallel architecture. Implemented on a Xilinx ZCU102 FPGA, it achieves real-time 4K SR at 95.3 FPS for x2 upscaling at 300 MHz. Compared with existing SR accelerators, ScaleLUT uses at least 58.6% fewer LUTs, 41.1% fewer flip-flops, zero DSPs, and 42.0% lower power, while delivering 10x and 1.2x speedups over the best CPU-based SR implementation and prior FPGA-based SR accelerators, respectively. These results demonstrate the effectiveness of joint LUT algorithm-hardware co-design for practical and energy-efficient edge SR deployment.
cs.AR / 5 / 2609.16600
A 420 GOPS/W CGRA with a Configurable MAC and Dynamic Truncation
Abstract
Edge devices demand for highly efficient yet flexible processing capability to handle dynamic real-time workloads. Coarse grain reconfigurable architecture (CGRA) emerges as a suitable accelerator candidate in edge devices, because they are as flexible as general purpose processors and offer high efficiency close to that of domain specific accelerators. However, a typical CGRA requires two cycles for a multiply-and-accumulate (MAC) operation, and workloads such as neural network inference and signal processing involve many MAC operations, resulting in long CGRA processing time. This work proposes a CGRA that has configurable MAC units in the processing elements (PEs) that can perform an addition (ADD) or multiplication (MUL) or a MAC by using the same multiplier and adder, in a single cycle. The readout precision of MAC result can be adjusted by a truncation block. The proposed CGRA is implemented with 40nm CMOS technology. It attains an energy efficiency of 420.6GOPS/W operating at supply of 0.6V and frequency of 21MHz, which is 1.4 times higher than the state-of-the-art.
cs.AR / 6 / 2609.16742
Carry-Through Checksum: A Lightweight Fault-Detection for CNN Inference at the Edge
Abstract
Convolutional Neural Networks (CNNs) are increasingly deployed in safety-critical edge applications, where soft errors can silently corrupt inference outputs and lead to unsafe decisions. Such applications typically rely on resource-constrained embedded GPUs, requiring fault detection and mitigation techniques that add minimal compute, memory, and latency overhead while integrating seamlessly with the standard GPU inference pipeline. Existing algorithm-based fault tolerance techniques rely on matrix augmentation and per-operation checksum verification, imposing substantial overhead that is prohibitive for CNN inference on embedded GPUs. In this work, we propose carry-through checksum, a fundamentally new scheme for soft-error detection in CNN inference on embedded GPUs. The method embeds dedicated carry-through filters into the convolutional layers, which compute a checksum from the CNN's own operations and propagate it through inference, enabling end-to-end error detection with a single output verification. Experimental results on multiple CNN architectures show that the proposed method detects 95.86% and 86.56% of critical faults for FP32 and FP16, respectively, at almost no additional per-image overhead. Detected faults are mitigated through re-execution, incurring only 2.27% run-time overhead across the entire test set on an NVIDIA Jetson Orin NX GPU.
cs.AR / 7 / 2609.16898
OptiPrime: Optimizing Private Inference through Protocol-Hardware Co-design
Abstract
Private deep neural network (DNN) inference based on hybrid homomorphic encryption (HE) and multi-party computation (MPC) can protect user data with a formal guarantee, but at the cost of significant latency overhead due to HE. Customized HE accelerators have been proposed and have achieved orders-of-magnitude speedup for individual HE operations. However, when directly applying a commercial HE accelerator to state-of-the-art HE-MPC frameworks, we observe only limited end-to-end performance gain. This is because HE-MPC frameworks often require wireless transmission of input and output ciphertexts for each HE operation, leading to a severe network communication bottleneck. To overcome this challenge, we introduce OptiPrime, a protocol-hardware co-optimization framework for efficient private DNN inference. OptiPrime features a novel HE protocol for convolutions that substantially reduces the number of transmitted output ciphertexts and mitigates the network communication bottleneck. Meanwhile, as the new protocol introduces complex computation for fewer output ciphertext, we observe new memory access challenges due to a high volume of weight plaintexts and intermediate ciphertexts. Hence, we further propose a lightweight compression system for the weight plaintexts, reducing memory traffic by 10 times, as well as a specialized dataflow to maximize on-chip data reuse of intermediate ciphertexts. Extensive experiments show that our framework outperforms the Cheetah baseline by at most 5.7 times on CPUs and 4.2 times with an accelerator.
cs.AR / 8 / 2609.17057
Budgeted Express-Mesh: Traffic-Aware Link Placement and Deadlock-Free Adaptive Routing
Abstract
We present Budgeted Express-Mesh, a topology-routing co-design that adds a small number of traffic-aware express links under a fixed wire budget. An ASPL-based greedy placement is refined by simulation-guided annealing, while packets use committed top-K routes selected from delayed express-link congestion and reservation signals. Across four synthetic workloads, optimized placements consistently improve high-load throughput over Mesh and random placement, and annealing further improves Greedy. The gains persist under delayed quantized congestion information, longer express-link latency, multi-flit packets, and a 16-by-16 heterogeneous workload.
cs.AR / 9 / 2609.16787
Nested Parallel von Neumann Architecture and Nested BSP
Abstract
Large-scale AI computing is no longer a contest of ''one stronger processor,'' but of how an army of processors under one command can still be one computer. This paper offers two interlocking extensions. First, extend BSP to Nested BSP. The Turing machine describes computation as a single tape, in sequence. A million processors need not a longer tape, but a battle plan nested layer within layer: at every layer, parallel work, barrier, exchange and aggregate, then the next phase. Every ``parallel advance'' inside a layer repeats the same four steps. Nested BSP extends classic BSP by nesting it recursively, a computing paradigm for million-scale parallelism, under one rule: every node at every layer is a peer. Second, extend von Neumann to the Nested Parallel von Neumann Architecture, and Unified Bus is its interconnect. Von Neumann taught us how to build one stored-program computer. The false extrapolation of eighty years was that wiring many computers into a network yields one larger computer. A second habit ran deeper: nearly every design assumes a master that commands and slaves that obey---host over device, CPU over accelerator, center over edge. The Nested Parallel Architecture extends that idea rather than discarding it. Two nesting dolls must fit: Nested BSP in software, and the Nested Parallel von Neumann Architecture from package to autonomous zone, joined by one memory-semantic bus end to end, with full peer equality: physically sparse, logically tight. It pairs with Huawei's $τ$ Scaling law: $τ$ governs how each layer folds time, while the Architecture governs how the nested parallel computer stands, layer by layer, peer by peer. In summary, the paper extends BSP to Nested BSP and extends von Neumann to the Nested Parallel von Neumann Architecture. $τ$ folds time, peer-equal parallelism nests layer by layer---many processors, still one computer.
密码学与安全 (cs.CR)
28
cs.CR / 1 / 2609.16211
Feasibility of Homomorphic Inference for a Genomic Foundation Model
Abstract
Human genomic sequences can identify individuals, cannot be replaced after disclosure, and are the inputs that genomic foundation models are designed to interpret. We assess whether a compute provider can execute a released genomic foundation model without receiving query-derived genomic values in plaintext and whether correctness, memory, or cost prevents complete encrypted inference. We first reproduce the released model on three genomic task families and freeze an independently validated numerical reference. We then implement a client-assisted approximate homomorphic encryption protocol: the provider evaluates linear algebra on ciphertexts, while the key-holding data owner evaluates exact normalization, causal softmax, and activation functions at fixed boundaries. A noninteractive configuration completes one released-weight block but exceeds the tested accelerator-memory envelope when configured for composition. The client-assisted configuration executes all released transformer blocks and the task head for one heldout genomic-signal input at its full prompt length. It matches the frozen final label, peaks at 9,839 mebibytes of accelerator memory, and completes in 6,683 seconds on one accelerator. These results establish arithmetic feasibility for a complete classifier, while repeatability, network transport, and private token-index lookup remain unresolved. The biomedical significance is that, under the stated threat model, a served genomic model can process an encoded sequence without exposing plaintext queryderived activations to the compute provider.
cs.CR / 2 / 2609.16214
Analyzing Multi-Factor Authentication Through Cryptographic Security Properties
Abstract
Modern authentication systems use cybersecurity techniques to validate the identity of the person (or applications acting on behalf of the person) as a primary defense against unauthorized access. While initially built around single mechanisms such as usernames/passwords, physical tokens, or biometrics, current systems have evolved into multi-factor authentication (MFA) platforms that combine multiple mechanisms. Among them, several focus on strategies that prevent replay attacks (i.e. the reuse of a component that could have been potentially compromised).
cs.CR / 3 / 2609.16231
RuleAutoPilot: Synthesizing Deployable Suricata Rules from Network Traffic
Abstract
Rule-based Intrusion Detection Systems (IDS) such as Suricata are central to network security, yet crafting effective detection rules demands deep expert knowledge and cannot keep pace with emerging threats. Existing LLM-based approaches can reduce analyst effort, but they either rely on curated threat intelligence that is produced only after the underlying traffic artifacts already exist, or they require costly LLM use without sufficient quality control. We present RuleAutoPilot, an end-to-end agentic framework that generates deployable Suricata rules directly from malware network traffic, with no prior threat intelligence required. A key challenge is noise: network traffic captures often contain a small amount of security-relevant traffic mixed with large volumes of background traffic, which reduces LLM reasoning quality and increases cost. RuleAutoPilot addresses this challenge with a Benign Traffic Fingerprinting stage that removes known benign background flows before LLM processing. Rules that fail syntax checks, do not trigger on the source traffic, or generate false positives on a benign corpus are automatically repaired using structured feedback. Across 1,296 malware PCAPs, execution-grounded verification raises rule quality (F1) from 0.443 to 0.539. On a stratified 200-PCAP subset, RuleAutoPilot on the open-weight gpt-oss-120b reaches near-frontier quality, 0.524 F1 against Claude Opus 5 under Claude Code's 0.623, at 52x lower billed-token cost. Swapping only the backbone to Claude Opus 5, RuleAutoPilot surpasses Claude Code outright, 0.656 F1 against 0.623, at 40x fewer tokens. A stronger backbone raises RuleAutoPilot's own ceiling, but at the same backbone, our scaffold still outperforms Claude Code's, showing the scaffold contributes independently of the backbone.
cs.CR / 4 / 2609.16253
Exploiting and Securing Docker containers and Kubernetes pods from a MitM attack
Abstract
PURPOSE - Workloads in containers, such as Docker containers and Kubernetes pods, are vulnerable to many of the same attacks as workloads in non-container environments, including phishing, application exploits and network intrusions. This systematic review and design-and-creation study explores techniques for securing containerised-based operating systems against Man-in-the-Middle (MitM) attacks. The proposed framework uses a conceptual model for representing communication and cryptographic primitives, together with the AnBxJ Java security library and container firewalls operating at layer 7 of the OSI model. The study addresses the question: How can containerised-based operating systems be effectively secured from Man-in-the-Middle attacks? It aims to support practitioners in protecting Docker and Kubernetes deployments by systematising security practices and applying a zero trust architecture. METHODOLOGY - The research uses a Systematic Review (SR) based on the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA), bringing together evidence from studies addressing the same research topic. FINDINGS - Success factors were identified, and a security mechanism was successfully implemented in a containerised-based operating system scenario. VALUE - The findings may help practitioners protect Kubernetes and Docker installations by systematising container security practices and providing a zero trust architecture for containerised-based operating systems.
cs.CR / 5 / 2609.16323
Understanding the Usability of Cryptographic Verification Tools
Abstract
Cryptographic protocol verification tools are widely used to analyze the security of complex protocols, yet how users interact with these tools remains comparatively understudied. We present an exploratory human-centered study of experienced users of Tamarin, ProVerif, and related protocol verifiers. Our survey included researchers, graduate students, and practitioners with hands-on experience using Tamarin, ProVerif, or related tools. The findings reveal usability barriers across the verification workflow, including difficulties debugging non-termination and performance issues, and the lack of systematic methods for validating formal models against real protocols. When proofs fail without concrete attacks, users commonly simplify models, add helper lemmas, and revisit modeling abstractions. Participants also called for actionable diagnostics, clearer explanations of results, visualization, and automation for recurring proof tasks. Our findings suggest that persistent usability challenges arise from the gap between protocol-level reasoning and the verifier's formal model, proof procedures, and diagnostic output. We derive concrete design priorities for improving the accessibility, interpretability, and usability of cryptographic protocol verification tools.
cs.CR / 6 / 2609.16336
Illusion of Depth: Revealing Hidden Stereo Vision Vulnerabilities in Depth Estimation
Abstract
Stereo cameras are integrated into autonomous systems such as self-driving cars, drones, and robots to offer precise depth estimation in a cost-effective manner compared to LiDAR technology. In this work, we reveal an intrinsic vulnerability in stereo cameras that stems from their pixel sampling and calibration processes, which can influence the outputs of stereo matching algorithms. Attackers can achieve fine-grained control over the estimated depth of real obstacles using simple repeating patterns, without relying on sophisticated adversarial machine learning techniques. Furthermore, deep learning-based depth estimation models exhibit a similar vulnerability. We evaluate the impact of this attack on two widely used stereo matching algorithms (BM and SGBM), three deep learning models (PSMNet, MoCha-Stereo, and UniMatch), a stereo-LiDAR fusion model (SGM-DDC), and two popular commercial stereo cameras, the ZED2 and Intel RealSense D435. For example, in the ZED2 camera, an attacker can displace obstacles up to 20~meters farther or 12~meters closer. In our real-world evaluation in a driving setting, a brief 0.5~second attack can trigger emergency braking in a popular autonomous driving framework. We further demonstrate the feasibility at driving speeds up to 40~km/h using CARLA. Finally, we confirm the ineffectiveness of state-of-the-art defenses, and we propose a novel strategy that leverages similarity scores to dynamically detect and suppress the depth discrepancies. Our work highlights vulnerabilities hidden in stereo matching and deep learning depth estimation models, addressing critical limitations in autonomous system deployments.
cs.CR / 7 / 2609.16375
gr-PHYSEC: Real-time Channel-based Key Generation for Physical Layer Secure Wireless Communications
Abstract
Securing wireless communication against eavesdropping is critical, particularly in dynamic and decentralized environments. We present gr-PHYSEC, a new GNU Radio out-of-tree (OOT) module for real-time physical-layer key generation. Unlike traditional key generation that relies on pre-shared secrets or computational complexity, our approach derives symmetric keys from the wireless channel's inherent randomness. We embed a trained neural network within GNU Radio to extract channel features between trusted parties (Alice and Bob) during probe exchanges. These features are quantized into binary keys, reconciled via Reed-Solomon encoding, and further secured with SHA-512 hashing. The generated keys are then directly used to encrypt data. Real-world experiments at the FAU CAAI connected robotics testbed using ADALM Pluto software-defined radios and NVIDIA Jetson Orin validate the approach with ground robotic platforms. Results demonstrate low key disagreement rates and strong randomness, as verified by the NIST test suite for random and pseudorandom number generators for cryptographic applications. This integration showcases how GNU Radio can support real-time AI-driven security solutions, pushing the boundaries of software-defined secure communication. The source code for this project is available at: https://github.com/C2A2-at-Florida-Atlantic-University/gr-PHYSEC
cs.CR / 8 / 2609.16403
Implementing a White-Box Undetectable Backdoor for Random Fourier Features
Abstract
Goldwasser et al. showed that undetectable backdoors can be planted in machine learning models trained with the Random Fourier Features (RFF) algorithm, under a hardness assumption tied to the Continuous Learning With Errors (CLWE) problem. Under standard cryptographic assumptions, even a full white-box audit of a model's weights cannot detect this class of backdoor. The construction is stated in terms of cryptographic reductions and probabilistic lemmas, without a reference implementation, and relies on secondary machinery such as the Sparse Gaussian Pancakes distribution and a homogeneous CLWE conditional density. Its realizability in ordinary numerical code is not obvious from the paper alone. This paper implements the white-box CLWE-RFF backdoor construction end to end using only numpy and scipy, to test whether this threat is realizable with commodity scientific-computing tools or requires specialized cryptographic infrastructure. We give two samplers for the core $GP_d(b_k)$ distribution. The first is a rejection-sampling proxy. The second is an exact closed-form sampler derived from the homogeneous CLWE density and verified against its own analytic form. Using this implementation, we run statistical indistinguishability tests, covering both weight-space and functional black-box comparisons. We find no evidence of detectable difference between backdoored and clean models across a range of sparsity ratios $ρ= d_{\text{sparse}}/D$. We report which parts of the construction were straightforward to realize, which required derivation not spelled out in the paper. We also highlight which parts we did not attempt to reproduce, including the underlying lattice hardness reduction. We see this work as a contribution to understanding the practical realizability of the Goldwasser white-box CLWE core, not as a new theoretical result.
cs.CR / 9 / 2609.16423
No Bit Left Behind: Using Brute-Force Lifting to Achieve Fully Static Binary Recompilation
Abstract
Binary recompilation is a technique for operating directly on executable code. It promises to automate two important tasks: retrofitting security mitigations onto legacy binaries, and migrating binaries across instruction set architectures (ISAs). Yet today, there is no fully automated system that can reliably lift arbitrary binary executables to a compiler intermediate representation (IR) such as LLVM IR, or that can fully statically and reliably translate non-trivial binary executables from one ISA to another. The main underlying problem is that recovering a program's control flow graph (CFG) statically is impossible in general: computed branches can jump to targets that cannot be determined without actually running the program. Existing systems resort to runtime fallback mechanisms, requiring a significant portion of the binary translation machinery to accompany the translated program on the target machine. This article presents a fully static, whole-program binary lifting system requiring no runtime translation support on the target. Rather than attempting to distinguish code from data, we treat every byte offset as a potential branch target and lift the entire binary in a brute-force manner, constructing a superset CFG that conservatively contains all feasible control flows. Statically unresolvable computed branches are thereby reduced to lookups in a dispatch table that points to the corresponding translated control flow path. We have implemented this approach as a prototype binary recompiler from x86-64 binaries to LLVM IR, requiring no code/data heuristics. We validate it with a fully static cross-compilation to AArch64, achieved by reusing existing LLVM backends with no modification.
cs.CR / 10 / 2609.16462
Not All Relations Are Equal: Relation-Balanced and Calibrated Graph Learning for Provenance-Based Intrusion Detection
Abstract
Provenance-Based Intrusion Detection Systems (PIDSs) detect Advanced Persistent Threats (APTs) by analyzing system interactions. However, existing methods largely treat relations uniformly, overlooking statistical heterogeneity; in CADETS, relation frequencies differ by approximately $140{,}000\times$. This may cause PIDSs to focus more on frequent relations and overlook differences in normal error levels across relations, increasing the risk of false alarms and missed detections. We present RECAL, an unsupervised framework using relation-balanced masked graph learning to better capture rare interaction patterns. It further calibrates reconstruction errors against each relation's benign error distribution to produce comparable anomaly evidence, helping distinguish attacks from benign behavior and reduce false alarms. On three DARPA E3 datasets, RECAL achieves F1 scores of 99.99\%, 99.93\%, and 99.99\%, outperforming the best baseline on each dataset by 0.88, 0.82, and 0.42 percentage points, respectively. Compared with the baseline reporting the lowest FPR, RECAL reduces mean FPR by approximately $105\times$, $4\times$, and $41\times$.
cs.CR / 11 / 2609.16541
A Cyber Range Evaluation of Autonomous Network Incident Response Agents
Abstract
We test the performance of agents for automated network intrusion response in a cyber range intended for human operator training. The range implements an emulated networking environment with a variable network topology, red-team emulation and simulated user agents. The goal of the defensive agents is to prevent hosts in the network from being accessed by the red-team agent, while minimizing the availability costs induced from defensive measures. Alerts are generated using a SIEM platform and mapped to a data modeling language used by the agents. We test a combination of heuristic agents and policies learned using reinforcement learning. The learned policies are optimized to minimize the combined cost using a cyber attack simulator modeling the network. We found that the reinforcement learning agents were overall more efficient at defending the system than the heuristic policy, and that the performance depends highly on the policy of the adversary in combination with the simulated users.
cs.CR / 12 / 2609.16546
GPUThor: Amplifying Rowhammer Attacks via Non-Uniform Patterns to Exploit ECC-Protected GPUs
Abstract
GDDR memory in GPUs is vulnerable to Rowhammer attacks, where rapid memory accesses induce bit flips in adjacent cells, enabling data tampering and privilege escalation. However, prior GPU Rowhammer attacks trigger only tens to hundreds of bit flips, orders of magnitude fewer than CPU attacks, severely limiting their practical impact. This gap stems from the reliance of existing GPU Rowhammer attacks on uniform hammering patterns that activate aggressor and decoy rows equally, which results in low hammering intensity for aggressor rows. We present GPUThor, a high-intensity Rowhammer attack on NVIDIA GPUs leveraging non-uniform hammering. GPUThor reverse engineers GPU memory-access coalescing behavior to enable non-uniform hammering patterns on GPUs, that activate aggressor rows more intensely than decoy rows. Additionally, by identifying refresh instances when in-DRAM mitigations are applied, it constructs longer attack patterns that escape mitigation across refresh intervals, further increasing hammering intensity. Together, these techniques yield 500X to 23,500X more bit flips than prior GPU Rowhammer attacks, across several NVIDIA GPUs (A4000, A4500, A5000, A6000), reaching bit flip rates close to state-of-the-art CPU Rowhammer attacks. GPUThor also enables the first Rowhammer exploits on ECC-protected GPUs, inducing uncorrectable double and triple bit flips, making denial-of-service and privilege-escalation attacks practical even on GPUs with ECC enabled.
cs.CR / 13 / 2609.16563
The MAL Simulator: Cyber Operations Simulation based on Attack & Defense Graphs
Abstract
We have developed the MAL Simulator, a cyber operation simulator based on the Meta Attack Language (MAL). The MAL Simulator is intended for decision-driven cyber attack and defense simulations, for system analysis and the development of automated agents. By building the simulator around an attack modeling language, it can be adapted to different target domains without modifying the source code. We used the simulator for two case studies where we trained two types of agents for automated cyber operations: a defensive agent and an offensive agent. To ground the experiments, we base the models in data collected from an emulated network implemented in the cyber range CRATE. We found that the trained attacker policy could reach the designated targets more efficiently than the compared search methods, and that the trained defender agent induced lower costs than a naive heuristic agent under noisy alert conditions. When testing the RL attacker against the RL defender, we found that the performance of the defenders dropped significantly. This emphasizes the importance of cyber attack simulators to facilitate training both offensive and defensive agents. The MAL Simulator and associated tooling is publicly available and provides common interfaces for compatibility with existing machine learning frameworks.
cs.CR / 14 / 2609.16675
From Hypervisor to Container: Cloud Security Vulnerabilities, Defense Mechanisms, and Open Challenges
Abstract
In cloud computing, different users share the same physical hardware, which creates serious security risks. To protect data, cloud systems rely on virtual machines and containers to keep users isolated. This paper reviews over 120 security publications from 2008 to 2025, focusing on how these isolation boundaries can be breached. We examine threats like virtual machine escape, virtual machine hopping, CPU cache side-channels, container breakouts, vulnerable container images, and distributed denial of service (DDoS) attacks. We evaluate these security threats and their defenses using three key research questions. To compare different defense systems, we introduce a quantitative scoring framework called ADPO, which rates defenses from 0 to 3 based on their Accuracy, Deployment ease, Performance impact, and Operational overhead. We also map the impact of these attacks onto a 1-to-5 severity scale for Confidentiality, Integrity, and Availability. Finally, we highlight the trade-offs between security and system performance, and we outline open challenges like building low-overhead intrusion detection and creating realistic test datasets.
cs.CR / 15 / 2609.16681
MarkSec: Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks
Abstract
LLM watermarking helps trace the origin of generated text, but faces stealing attacks that recover watermark information, scrubbing attacks that remove watermark signals, and spoofing attacks that forge text accepted as watermarked. These attacks are often studied in isolation, leaving their connections unclear. Evaluations also often lack shared detector calibration, metric definitions, and reporting protocols. Moreover, measuring attack success and text quality separately makes it difficult to identify attacks that are both effective and quality-preserving. We propose MarkSec, a general framework that unifies analyses of stealing, scrubbing, and spoofing. We evaluate attacks under a common reporting protocol and introduce a quality-constrained attack success metric to assess effectiveness and text quality jointly. Experiments across representative watermark families, attacks, LLMs, and datasets reveal three findings. First, attacks that appear strongest by watermark removal alone can fall behind general rewriting when success also requires acceptable text quality. Second, general rewriting remains a strong baseline across watermark families, while its advantage over other scrubbers varies by family. Third, in a case study of one watermark family, stealing-based scrubbers often underperform the best general-scrubbing baselines when text quality is required. These results show that apparent attack winners depend on text-quality constraints, attack generality, and capability assumptions.
cs.CR / 16 / 2609.16694
Toward Secure AI-Powered Penetration Testing Agents: Security Threats, Guardrails, and Architectural Perspectives
Abstract
LLM-powered autonomous agents are transforming the penetration testing space with dynamic, multi-step offensive security workflows that require minimal supervision by humans. These agents leverage sophisticated reasoning abilities and external security tools to independently carry out reconnaissance, identify vulnerabilities, devise exploitation plans, and perform post-exploitation operations. But the ability to have persistent memory, to take actions in the real world, and to do long-horizon reasoning raises qualitatively different security concerns than traditional chat-based LLM systems. Existing guardrail mechanisms for conversational AI may not be sufficient to secure autonomous AI pentesting agents accordingly. To address these issues, we carry out a comprehensive security analysis on autonomous AI-penetration testing agents. We systematically analyse representative agent architectures, characterise their trust boundaries and attack surfaces and propose a threat taxonomy that is aligned with the lifecycle and covers LLM lifecycle attacks, agent-architecture attacks and cross-cutting behavioural attacks. We analyse the limitations of existing guardrail mechanisms, identify key research gaps, and discuss future research directions for developing specialised, context-aware, and architecture-aware guardrails to secure next-generation AI-driven offensive security systems.
cs.CR / 17 / 2609.16732
When Agents See Differently: Exposing UI Desynchronization Threats in Mobile Agents
Abstract
Mobile agents are increasingly capable of autonomously interacting with mobile applications and performing consequential actions on behalf of users. Effective human oversight of such agents relies on a basic premise: users and agents observe consistent information from the same interface. We show that this premise can be systematically violated. Users perceive mobile interfaces through physical displays and the human visual system, making their observations subject to occlusion and luminance contrast limitations. In contrast, agents consume digital screenshots that may retain such content and accessibility representations that expose nonvisual widget metadata. The same UI state can therefore present materially different information to users and agents, a mismatch we term human-agent UI desynchronization. We investigate whether a repackaged clone of a legitimate APK can exploit this desynchronization to steer an agent toward attacker-designated actions, while remaining fully functional and behaviorally consistent with the original application for human users. We demonstrate that this threat is feasible: perturbations embedded before deployment can induce such deviations without access to runtime user instructions, agent detection or online adaptation. To systematically expose and evaluate this threat, we develop an automated framework that constructs user runtime instruction-agnostic UI desynchronization attacks and realizes them in deployable APKs. We conduct static and dynamic evaluations across five mobile-agent frameworks and three backbone models on 546 tasks involving various applications, achieving average misleading rates of 77.9% and 66.9%, respectively. A complementary questionnaire-based study with 186 participants finds that the visual perturbations used in our attacks are difficult for human users to notice.
cs.CR / 18 / 2609.16928
Cybersecurity in Power Grids: Standards and Research Challenges
Abstract
This paper examines Smart Grid cybersecurity, emphasizing the critical distinctions between IT and OT environments. It analyzes grid architecture, substation threats, and key international standards, specifically IEC 62351, IEC 62443, and ISO 27001. Finally, it overviews latest research trends, including AI-driven threat detection.
cs.CR / 19 / 2609.17150
Observational Indistinguishability and Integrity Blind Regions in Hybrid Quantum-Classical Workflows
Abstract
We present a claim-relative evidence/reference framework for hybrid quantum-classical workflow integrity. Observational indistinguishability yields structural blind regions, distinct from finite-batch statistical misses. Within the declared lattice, a trusted same-batch scalar $R_0$ suffices for conclusion integrity, aggregate $M_0$ for aggregate plus conclusion integrity, and item-aligned binding for item identity. In 3,600 label interventions, feature/prediction views realize exact label-path invariance; all 764 geometry-aligned aggregate-blind rows equal their paired-clean responses, giving zero attack-only increment. For statistical response, the geometry-aligned construction detects 343/2,700 conclusion-changing ($τ\to 0^+$) label interventions with the conformal rule and 1,183/2,700 with the uncorrected union; the original frozen same-item geometry yields 11/2,617 and 43/2,617, respectively. The executed conformal clean false-action rates are 0.048--0.059 descriptively; its finite-sample guarantee requires exchangeability, which the overlapping-draw design violates. The cluster-preserving adaptive stress test (Gate A) reduces response versus matched controls in 25--40 of 40 environment/split cells while retaining conclusion changes. A bounded 165-design-cell ideal-statevector and finite-shot-emulation branch directly instantiates semantic, estimated and observed kernel transitions. The fixed equal-weight design estimates neither deployment prevalence nor QPU, provider or deployed-service assurance.
cs.CR / 20 / 2609.17164
Plug 'n' Pray: Agentic LLM-based Detection of Potential Log File Exposures in Third-Party Content Management System Plugins
Abstract
Content Management Systems (CMS), such as WordPress, power a large share of the web (~58%), and their extensibility through third-party plugins is a major source of their popularity as well as of their attack surface. One high-impact weakness that remains understudied is log file exposure by CMS plugins, which create log files for debugging or other purposes. If these files are insufficiently secured, they can disclose sensitive information (e.g. credentials, personal data) which has led to website compromises in the past. In this work, we present an agentic, LLM-based framework that automatically detects potential log file exposures in plugins of the most popular CMS (WordPress). Our agent analyzes each plugin by performing static and dynamic analysis. We evaluated our approach on the 300 most-installed WordPress plugins (about 0.6% of all), which together account for over 250M active installations, i.e. 75% of all active installations in the official plugin ecosystem. We manually validated each finding, reproducing 79 of 81 findings from 62 plugins. We observed that several protective measures appear to be implemented that we classify as creation-control (e.g. manual log activation) and access-control (e.g. deny rules in .htaccess). However, we find that multi-layered protection is required, but not always present. From these results we derive a taxonomy of log file path and protection patterns and deduce a set of best practices for developers to securely handle them. Finally, our study corroborates that agentic LLMs are an useful tool for security analysis.
cs.CR / 21 / 2609.17204
Cross-Domain Inference for Human Localization: Applying Wi-Fi RSSI Data to CSI-Trained Models
Abstract
Wi-Fi signal data can be used to compromise the privacy of individuals. While many existing approaches rely on Channel State Information (CSI), collecting this data on typical IoT devices often requires elevated operating system permissions and specialized drivers. Consequently, this paper investigates the feasibility of utilizing Received Signal Strength Indicator (RSSI) data to predict human locations. RSSI was selected because it is accessible even on devices with limited user permissions, and therefore is more applicable to a wider array of IoT devices. To bypass the tedious process of obtaining training data needed to train an RSSI-based model, an existing Wi-Fi pose prediction project was used in this research. However, that project assumed CSI data as input. Therefore, we investigate the feasibility of cross-domain inference, i.e., feeding RSSI data into that existing CSI-based model. We collected an RSSI dataset, synchronized with video ground-truth of a person moving within a room, to evaluate the model's performance. This evaluation confirmed that RSSI data can predict locations with approximately 80% confidence when human movement is present. This demonstrates that a model trained on CSI data can be used to evaluate low-granularity RSSI data consisting of decibel-milliwatt (dBm) values to roughly locate people in the collection space. These results imply that a wide range of IoT devices can be used for privacy invasion in Wi-Fi-dense environments.
cs.CR / 22 / 2609.17254
SEMA-GUARD: Semantic and Graph-Based Vulnerability Detection in Assembly Code
Abstract
In cases where source code is not available, such as malware analysis, firmware analysis, and embedded systems analysis, vulnerability detection in compiled programs has gained importance. Current methods are heavily reliant on syntactical regularities or higher level representations that are vulnerable to changes in the compiler and may not be readily applicable to assembly code.In this article, we present SEMA-GUARD, a framework that uses semantic analysis and graph neural networks to identify flaws in assembly code. The approach improves the representation of control flow graphs by adding information about the program's execution at a lower level of abstraction, including stack manipulations, memory accesses, and data flow. A set based on the Juliet Test Suite was used to evaluate the effectiveness of SEMA-GUARD. In this set, each piece of source code is initially translated into assembly language and then broken down into function-level chunks. The suggested method, which relies only on statistical or structural data, achieves an accuracy of 85.1\% and an F1 score of 0.801, according to the results. Such results imply that including semantic information in graph-based models may be a successful method for identifying vulnerabilities in compiled code.
cs.CR / 23 / 2609.17281
GAUGE: A Formal Framework for Measuring Cryptographic Security under Heterogeneous Adversary Cost Models
Abstract
Standards bodies report cryptographic security as a single number of bits, but this value depends on the adversary cost model used to price time, memory, and quantum resources. Different conventions can therefore produce different rankings of cryptographic schemes. GAUGE represents security as a function over admissible cost models, called a security profile. Comparisons then become comparisons between profiles, and ranking reversals become an explicit structural property rather than a measurement error. We formalize price functionals over a cone of adversary cost models, show that security profiles are piecewise-linear and concave, and prove a rating trilemma: when two profiles cross, no rating can simultaneously be faithful to underlying costs, total over comparable pairs, and independent of the chosen cost model. We provide a polynomial-time linear-programming procedure that certifies whether the ranking of two schemes is robust, reverses under admissible models, or is genuinely incomparable. We extend GAUGE with a two-layer risk measure combining stochastic cryptanalytic decay with uncertainty over the appropriate cost model. We evaluate the framework on NIST post-quantum standards, classical anchors, and a 25-year chronology of cryptanalytic breaks. The analysis certifies a ranking reversal for ML-KEM-512 versus AES-128 from a 4-5% shift in memory pricing, and measures a lattice-sieving cost drift of 9.79 bits per year over eight years. A hybrid X25519 + ML-KEM-768 handshake reduces combined-break probability twenty-fold at a 2.3 kilobyte cost. The artifact reproduces all tables and figures in under seven seconds. GAUGE provides an explicit and auditable framework for reporting cryptographic security under competing cost models.
cs.CR / 24 / 2609.17316
Can We Stop The Ads? Taxonomy and Characterization of Smartphone Splash Ads and Existing Countermeasures
Abstract
Splash ads are full-screen advertisements that pop up and appear as the first interaction page when users start an app, often tricking users into unknowingly activating certain trigger mechanisms, such as moving the phone to redirect users to other profit-driven third parties. So far, splash ads have already caused significant real-world impacts, ranging from significantly delaying emergency response to distracting drivers, as well as degrading accessibility of apps to vision-impaired users. We analyze 108 documented implementations of advertising defenses to examine their applicability to splash ads and the requirements users face when deploying them. Our analysis identifies substantial deployment barriers, including device rooting or jailbreaking, runtime code injection, and application modification. Options without these requirements can still involve additional permissions, rule maintenance, source compilation, or payment. In our evaluation of 13 configurations of 11 tools across 10 popular apps, only one tool prevented the target ad-triggered navigation across all ten apps. It required Accessibility permission, and ads remained visible for approximately one second before dismissal. Other tested configurations failed to prevent navigation or, in some cases, left host apps unable to launch or stuck on the ad page. We further analyze the outstanding challenges and pos- sible future directions, highlighting the urgent need to incentivize smartphone manufacturers to provide more friendly and regulated platforms.
cs.CR / 25 / 2609.17349
RobResilience: Implementing and Evaluating a Resilience Framework for Cyber-Physical Embodied Systems
Abstract
In embodied cyber-physical systems, active cyberattacks pose an immediate threat not just to data, but to physical integrity and human safety. While existing security approaches excel at detection, they lack the runtime mechanisms to determine whether a disruption is tolerable or if performance degradation remains within safe operational bounds. This gap leaves autonomous systems vulnerable to graceful failure paralysis, where they cannot distinguish between a safe, degraded state and a catastrophic hazard during an ongoing attack. This paper presents RobResilience, an implementation of a formal resilience framework for embodied cyber-physical systems in a Webots simulation environment, using a PR2 robot and ROS2. The framework evaluates three predicates at runtime: tolerable disruption ($δ$), tolerable degradation ($γ$), and mitigation feasibility ($μ$), over a compromised device set derived from IDS confidence scores. When resilience is lost, the framework triggers available mitigation strategies. We evaluate our implementation through eight attack scenarios that systematically cover all possible combinations of the predicate state space, varying attack targets, degradation rates, and mitigation availability. Results confirm that the runtime behaviour of the implementation is consistent with the theoretical definitions.
cs.CR / 26 / 2609.17397
Closing the Loop: Bidirectional Fully Encrypted Protocols
Abstract
Fully encrypted protocols (FEPs) provide encrypted channels that make all protocol-generated bytes computationally indistinguishable from uniform random strings. Several previous works have explored security definitions and constructions of unidirectional FEPs: protocols in which one party acts only as a sender, and the other acts only as a receiver. However, most applications require two-way information exchange, and a network adversary can observe communication in both directions and their shared lifetime. Because the semantics of bidirectional channels involve more complex shared state, it is possible that the ``naïve'' composition of two unidirectional channels can result in a two-way protocol that can be detected based on dependencies between the two directions, such as traffic imbalance, channel closure, failures, or connection tear-down. To address this issue, we introduce new formal security definitions for bidirectional FEPs that capture exact shaping, delivery, protocol-state integrity, private half-close, and cross-direction isolation, while revealing a public ``sending schedule'' and ``closing epoch'' that may be randomized. We show that the trivial composition fails to meet these definitions, leading to practical detection attacks. We then construct provably secure bidirectional FEPs (BiFEPs) for both the datastream and datagram settings. For datastream, we combine two direction-separated FEPs with a ``wrapper'' layer that prevents detection based on the mismatch between uni- and bi-directional connection states. For datagram, we add encrypted DATA/FIN/ACK with replay protection and loss-tolerant close. We validate the design through a Rust implementation and show that none of the surveyed deployed protocols provides the full set of BiFEP security properties.
cs.CR / 27 / 2609.17399
SCHERI: Provably Secure Speculation Under the Constant-Time Policy for CHERI (Extended Version)
Abstract
Capability-based architectures such as CHERI provide strong support for the architectural isolation of software components. To additionally protect against microarchitectural leakage, software can be written in a constant-time fashion. Modern processors, however, rely heavily on speculative execution, which can invalidate the constant-time guarantees and leak isolated secrets transiently. In this work, we show that providing secure speculation for CHERI is non-trivial, and that existing proposals fail to preserve the confidentiality guarantees. We develop a formal framework for reasoning jointly about capability safety, speculative execution, and information-flow security, and use it to demonstrate potential leaks. We then present SCHERI, a new processor design within this framework, and formally prove that it provides end-to-end secure speculation guarantees for the constant-time policy. Our results provide formal foundations and practical guidance for building future capability-based processors, which are resilient to Spectre attacks for constant-time programs.
cs.CR / 28 / 2609.17525
You Shall Not Pass into Ring-0! A User Privacy-Friendly Anti-Cheat Architecture for Personal Computers
Abstract
Kernel-level anti-cheats are effective against malicious player behavior in competitive video games, but raise significant user privacy concerns regarding installing unverifiable components at privileged modes (i.e., ring-0 in x86). While existing research has focused on improving the effectiveness of anti-cheats, the user privacy concern has been largely ignored. Tirith is an anti-cheat architecture that addresses this problem using two key ideas. First, instead of running video games within regular processes that players (as root admins) have control over, Tirith executes video games in Protected Virtual Machines that naturally sandbox computations from untrusted admins. Second, to monitor user behavior outside the sandbox (e.g., see if they are running malicious drivers), Tirith leverages a virtualization monitor that is trusted by both players and developers. Together, these ideas remove the need to run untrusted kernel-level anti-cheats, while providing the same level of protection compared to such solutions against a wide-range of common cheating mechanisms. The main challenge we face in implementing these ideas, however, is that the existing software stack for virtual machines is not designed to run video games and creates significant security and performance problems. We address these problems by proposing a security-focused Library OS kernel for games and an efficient graphics sharing pipeline for near-native rendering and display performance. In summary, without compromising on cheating behavior detection or performance, this work makes user privacy a first-class citizen in personal computers.