Daily Research Digest
arXiv Papers
2026-09-03
322
Papers
9
Categories
86
Translated
收藏清单 0
精选 · Favorites
86
cs.AI / 1 / 2609.01982
Benchmarking Language Models for Statistical Problem Formulation
面向统计问题表述的语言模型基准测试
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly used as assistants for statistical and data science work, yet existing evaluations largely assume the analysis target is already specified. In practice, users arrive with informal goals and heterogeneous data, leaving the model to decide what statistical task is implied and which data are relevant. We first formalize this upstream step as Statistical Problem Formulation and decompose it into two subtasks: (1) Statistical Problem Classification and (2) Variable Identification & Role Assignment. We then introduce StatFormBench, a benchmark built from five cross-domain statistics textbooks and a data science case library, covering diverse problem types, data representations, and scenario styles. It contains 1,013 samples spanning 20 coarse-grained and 85 fine-grained statistical problem categories. Across 14 open- and closed-source LLMs, the best zero-shot models reach only 72.0 fine-grained classification accuracy and 63.2 variable set overlap. No model performs consistently best across the two subtasks, while enhanced prompting strategies yield only limited or inconsistent gains. We release the benchmark data on Hugging Face at https://huggingface.co/datasets/THU-CongLab/StatFormBench and the evaluation code on GitHub at https://github.com/THU-CongLab/StatFormBench.
Chinese Translation
大型语言模型(LLMs)越来越多地被用作统计和数据科学工作的助手,然而现有评估大多假设分析目标已经明确。在实践中,用户带着非正式的目标和异构数据而来,需要由模型来判断隐含的统计任务是什么以及哪些数据是相关的。我们首先将这一上游步骤形式化为统计问题表述(Statistical Problem Formulation),并将其分解为两个子任务:(1)统计问题分类;(2)变量识别与角色分配。随后我们介绍了StatFormBench,这是一个基于五本跨领域统计学教科书和一个数据科学案例库构建的基准,涵盖多种问题类型、数据表示和场景风格。它包含1,013个样本,覆盖20个粗粒度和85个细粒度的统计问题类别。在14个开源和闭源LLM中,最佳零样本模型仅达到72.0的细粒度分类准确率和63.2的变量集合重叠率。没有任何模型在两个子任务上都能始终表现最佳,而增强的提示策略只带来了有限或不一致的提升。我们在Hugging Face上发布了基准数据,网址为 https://huggingface.co/datasets/THU-CongLab/StatFormBench ,并在GitHub上发布了评估代码,网址为 https://github.com/THU-CongLab/StatFormBench 。
cs.AI / 2 / 2609.02059
DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents
DocHop:信息密集型文档中域外多跳推理的基准测试
large language model
大语言模型相关
Abstract
Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, leaving underexplored a key capability: whether models can use textual context to determine how chart evidence should be selected, interpreted, and aggregated. We introduce DocHop, a benchmark for integrated chart--context reasoning in document-style images. In DocHop, the document narrative specifies multi-step compositional constraints, while charts provide the corresponding data values. Questions are grounded on a semantic reference label defined in the narrative, requiring models to resolve target entities from context before aggregating evidence across multiple charts. To enable systematic evaluation, we construct DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, covering 2,074 examples across six task categories. Experiments on a wide range of proprietary and open-source MLLMs show a substantial gap to human performance: annotators achieve over 90% accuracy, while the best model reaches only 62.83%. Reasoning-enhanced models consistently show improved results, but performance degrades as reasoning complexity increases. Overall, DocHop provides a controlled testbed for challenging multi-hop document reasoning.
Chinese Translation
多模态大语言模型(MLLMs)在结构化视觉理解任务(如图表问答和文档问答)上已展现出强劲的性能。然而,现有基准通常孤立地评估这些领域,从而使得一个关键能力未被充分探索:模型是否能够利用文本上下文来确定图表证据应如何被选择、解释和聚合。我们提出了DocHop,一个用于文档风格图像中图表-上下文集成推理的基准。在DocHop中,文档叙述规定了多步组合约束,而图表提供了相应的数据值。问题基于叙述中定义的语义引用标签,要求模型在跨多个图表聚合证据之前,先从上下文中解析出目标实体。为了能够进行系统评估,我们通过一种随机逻辑优先的生成流水线构建了DocHop,该流水线具有可控的推理深度和视觉密度,并覆盖了六个任务类别的2,074个示例。针对广泛的专有和开源MLLM的实验显示出与人类表现之间的巨大差距:标注者的准确率超过90%,而最佳模型仅达到62.83%。推理增强模型始终展现出改进的结果,但性能随推理复杂度的增加而下降。总体而言,DocHop为具有挑战性的多跳文档推理提供了一个受控的测试平台。
cs.AI / 3 / 2609.02116
Semantic Signal-Assisted Inspection and Recovery Allocation in Reverse Logistics
逆向物流中语义信号辅助的检验与回收分配
large language model
大语言模型相关
Abstract
Reverse-logistics operators often decide how to inspect and route returned assets before their condition is fully observed, while full inspection consumes scarce labor. Semantic Signal-Assisted Decision Support converts return notes into a condition factor and a signal-quality score that guide inspection depth and recovery allocation under shared labor capacity. We evaluate the framework in three synthetic benchmark scenarios spanning information technology decommissioning, aircraft maintenance, and consumer-electronics returns. Across 30 paired simulation seeds, the keyword implementation improves net recovery value relative to a structured-feature comparator with noisy full inspection while reducing inspection cost in all three scenarios. A risk-blind comparator that skips inspection altogether still records higher value under the benchmark's purely economic objective. At matched inspection cost, score-guided targeting adds 53.9 thousand United States dollars per batch in the aircraft scenario but has little economic effect in the other two configurations; phrase and large language model extractors provide further gains in the aircraft scenario. These results show how narrative evidence can support inspection allocation before recovery decisions are made.
Chinese Translation
逆向物流运营者常常在完全观察到退货资产的状况之前,就决定如何检验和安排这些退货资产的路线,而全面检验会消耗稀缺的劳动力。语义信号辅助决策支持将退货备注转换为一个状况因子和一个信号质量得分,这些在共享劳动力容量下指导检验深度和回收分配。我们在三个跨越信息技术退役、飞机维护和消费电子退货的合成基准场景中评估该框架。在30个配对模拟种子中,与具有噪声全检的结构化特征比较器相比,关键词实现提高了净回收价值,同时在所有三个场景中降低了检验成本。在基准的纯经济目标下,完全跳过检验的风险盲比较器仍然记录了更高的价值。在匹配的检验成本下,得分引导的目标定位在飞机场景中每批增加5.39万美元,但在其他两种配置中几乎没有经济效果;短语和大语言模型提取器在飞机场景中提供了进一步的收益。这些结果表明,在做出回收决策之前,叙述性证据可以如何支持检验分配。
cs.AI / 4 / 2609.02215
ASCII Attack: Recontextualising Harmful Requests as Artistic Critique in Large Language Models
ASCII 攻击:在大语言模型中将有害请求重新语境化为艺术批评
large language model
大语言模型相关
Abstract
Safety alignment trains large language models to refuse harmful requests stated plainly, but that training is applied mostly to surface form. Requests that only recontextualise the same operational content, changing how the model reads it, are therefore only weakly covered. The ASCII Attack is one such recontextualisation. It is single-turn and black-box: one message, with no access to model internals. It embeds a fully legible harmful request in ASCIl-art characters, presents it as artwork, and asks for feedback. Unlike ArtPrompt, it hides nothing: the request stays readable. The reply is written as artistic critique and can contain operational detail that a plain request would have been refused for. Every framed prompt is paired with a direct-question control, so the contrast is isolated from topic, model and decoding variation. The contrast identifies a bundled surface, not one isolated channel. Across eleven models and eight harm topics, a harm-aware classifier judges 62% of framed prompts harmful against 42% of controls. On the most susceptible model the framed prompt succeeds 93% of the time. A single query matches or exceeds published single-query attacks under four of five harm judges. The effect tracks the model more than the topic and does not diminish with scale. At least one judge dissents from the panel majority on nearly two-thirds of framed rows, which is itself a measurement-validity finding. That pattern is consistent with mismatched generalisation.
Chinese Translation
安全对齐训练大语言模型拒绝以直白方式陈述的有害请求,但这种训练大多应用于表面形式。因此,那些仅仅重新语境化相同操作内容、从而改变模型对请求解读方式的请求,只会得到很弱的覆盖。ASCII 攻击就是这样一种重新语境化。它是单轮且黑盒的:一条消息,无法访问模型内部。它将一个完全清晰可读的有害请求嵌入到 ASCII 艺术字符中,将其作为艺术作品呈现,并请求反馈。与 ArtPrompt 不同,它不隐藏任何内容:请求仍然可读。回复以艺术批评的形式写出,并且可能包含操作性细节;对于直白请求而言,这些细节本会导致被拒。每个框架化提示都与一个直接提问对照项配对,从而使这种对比与主题、模型和解码变化相隔绝。这种对比识别的是一个捆绑的表面,而非一个孤立的通道。在 11 个模型和 8 个有害主题中,一个危害感知分类器判定 62% 的框架化提示有害,而对照项中这一比例为 42%。在最易受影响的模型上,框架化提示有 93% 的时间成功。在五个危害评判器中的四个下,单次查询达到或超过了已发表的单次查询攻击。这种效应更多是随模型变化,而非随主题变化,并且不会随规模扩大而减弱。在近三分之二的框架化行中,至少有一位评判者不同意评判小组的多数意见,这本身就是一个测量效度方面的发现。这种模式与错配泛化是一致的。
cs.AI / 5 / 2609.02216
PEARL: Path-Entity Aligned Relational Learning with Contextual Subgraphs for Inductive Knowledge Graph Completion
PEARL:基于上下文子图的路径-实体对齐关系学习用于归纳知识图谱补全
large language model
大语言模型相关
Abstract
Inductive knowledge graph completion (IKGC) aims to predict missing links involving entities unseen during training, requiring models to learn transferable relational and structural patterns. Existing subgraph- and path-based approaches often encode relational paths independently of their surrounding query subgraphs, although the predictive relevance of a path may vary across structural contexts. We propose PEARL, a Path-Entity Aligned Relational Learning framework that models paths as context-conditioned reasoning signals. PEARL constructs a query-specific contextual subgraph from the union of the query entities' neighborhoods and uses a large language model (LLM)-guided retriever to distill semantically relevant paths. It then builds a bipartite interaction graph over paths, contextual entities, and a global subgraph representation, allowing path embeddings to adapt to local and global structural evidence. To suppress noise introduced by the enlarged context, PEARL employs a dual-view contrastive objective that promotes representation consistency under stochastic contextual perturbations. Experiments on WN18RR, FB15k-237, and NELL-995 show that PEARL obtains the best average Hits@10 among the compared IKGC methods on all three benchmarks. Ablation studies, efficiency analyses, and case studies further validate the contributions of contextual subgraph modeling, semantic path retrieval, path-entity interaction, and contrastive regularization.
Chinese Translation
归纳知识图谱补全(IKGC)旨在预测涉及训练期间未见实体的缺失链接,要求模型学习可迁移的关系模式和结构模式。现有的基于子图和基于路径的方法通常独立于其周围的查询子图来编码关系路径,尽管路径的预测相关性可能因结构上下文而异。我们提出了PEARL,一种路径-实体对齐关系学习框架,将路径建模为上下文条件化的推理信号。PEARL从查询实体的邻域并集中构建一个查询特定的上下文子图,并使用大型语言模型(LLM)引导的检索器来提炼语义相关的路径。然后,它在路径、上下文实体和全局子图表示之上构建一个二分交互图,使路径嵌入能够适应局部和全局结构证据。为了抑制因上下文扩大而引入的噪声,PEARL采用了一种双视图对比目标,在随机上下文扰动下促进表示一致性。在WN18RR、FB15k-237和NELL-995上的实验表明,在所有三个基准上,PEARL在比较的IKGC方法中获得了最佳的平均Hits@10。消融研究、效率分析和案例研究进一步验证了上下文子图建模、语义路径检索、路径-实体交互和对比正则化的贡献。
cs.AI / 6 / 2609.02244
Task-Level Natural Language Priors as Learning Signals for Low-Resource LLM Training
任务级自然语言先验作为低资源LLM训练的学习信号
large language model
大语言模型相关
Abstract
Large language models (LLMs) often struggle when low-resource training data are ambiguous or incomplete. Task-level natural-language priors can provide useful guidance in such settings, but existing approaches usually treat these priors as input context rather than as learning signals during training. We propose Prior-Guided Tuning (PGT), a training perspective that incorporates natural-language priors as auxiliary learning signals for low-resource LLM training. Under this perspective, we introduce Contrastive Prior Steering (CPS), which keeps the original supervised objective intact while adding positive and negative prior-conditioned auxiliary losses to encourage task-consistent learning and discourage plausible but misleading alternatives. Experiments on AmbiMath, Jigsaw, and MNLI/HANS show that CPS consistently improves over plain and prompt fine-tuning. On AmbiMath, CPS achieves 97.6% average exact-match accuracy. On Jigsaw, CPS improves average Macro F1 by 9.5 percentage points over standard fine-tuning, and with 1/10 of the experimental training data slightly exceeds full-data plain fine-tuning. On HANS, CPS improves non-entailment accuracy by 8.3 and 5.2 percentage points for LLaMA 3.1 8B and Qwen 2.5 7B, respectively, while maintaining comparable in-domain MNLI accuracy. These results support our central claim: task-level natural-language priors can provide useful guidance as auxiliary learning signals for low-resource LLM training. Our code and data will be publicly available.
Chinese Translation
大型语言模型(LLMs)在低资源训练数据模糊或不完整时常常表现不佳。在此类情境下,任务级自然语言先验可以提供有用的指导,但现有方法通常将这些先验视为输入上下文,而非训练过程中的学习信号。我们提出先验引导微调(Prior-Guided Tuning, PGT),这是一种将自然语言先验作为辅助学习信号融入低资源LLM训练的训练视角。在该视角下,我们引入对比先验引导(Contrastive Prior Steering, CPS),它保持原始监督目标完整不变,同时添加正负先验条件下的辅助损失,以鼓励任务一致的学习并抑制合理但具有误导性的备选方案。在AmbiMath、Jigsaw和MNLI/HANS上的实验表明,CPS一致优于普通微调和提示微调。在AmbiMath上,CPS达到97.6%的平均精确匹配准确率。在Jigsaw上,CPS相比标准微调将平均Macro F1提高了9.5个百分点,并且仅用实验训练数据的1/10就略微超过了全量数据的普通微调。在HANS上,对于LLaMA 3.1 8B和Qwen 2.5 7B,CPS分别将非蕴含准确率提高了8.3和5.2个百分点,同时保持了相当的领域内MNLI准确率。这些结果支持了我们的核心主张:任务级自然语言先验可以作为低资源LLM训练的辅助学习信号提供有用的指导。我们的代码和数据将公开提供。
cs.AI / 7 / 2609.02253
APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering
APEx:面向自适应深度研究问答的智能体程序性经验蒸馏
large language model
大语言模型相关
Abstract
Deep research agents augment large language models with external tools to answer complex, long-horizon questions through multi-turn reasoning. Learning from prior experience is crucial for continual improvement, yet existing methods either retrieve verbose task-specific traces that burden decision-making, or distill procedural skills that remain decoupled from downstream policy adaptation. We propose APEx, a hierarchical experience utilization framework that organizes interaction history into instance-level trajectory memories and category-level procedural skills, and couples them through a closed-loop architecture of Executor, Distiller, and Planner. The three modules are optimized via a three-stage alternating GRPO training paradigm, enabling reward-guided skill distillation rather than fixed-prompt generation. At test time, distilled skills serve as procedural priors for online Planner adaptation through skill-guided test-time reinforcement learning, allowing ground-truth-free self-improvement with skill-alignment regularization to prevent policy drift. Experiments on 7 benchmarks demonstrate that APEx achieves state-of-the-art performance, surpassing GPT-5.4 by 14.7 points and the strongest memory-augmented baseline by 3.0 points.
Chinese Translation
深度研究智能体通过外部工具增强大型语言模型,以多轮推理方式回答复杂的长期问题。从先前经验中学习对于持续改进至关重要,然而现有方法要么检索冗长的任务特定轨迹而加重决策负担,要么蒸馏出与下游策略适应相脱节的程序性技能。我们提出 APEx,一种分层经验利用框架,将交互历史组织为实例级轨迹记忆和类别级程序性技能,并通过执行器(Executor)、蒸馏器(Distiller)和规划器(Planner)构成的闭环架构将二者耦合。这三个模块通过三阶段交替的 GRPO 训练范式进行优化,实现了奖励引导的技能蒸馏,而非固定提示生成。在测试阶段,蒸馏出的技能作为程序性先验,通过技能引导的测试时强化学习实现在线规划器适应,从而在无需真实标签的情况下自我改进,并通过技能对齐正则化来防止策略漂移。在 7 个基准上的实验表明,APEx 取得了最先进的性能,超过 GPT-5.4 达 14.7 个百分点,超过最强的记忆增强基线达 3.0 个百分点。
cs.AI / 8 / 2609.02264
Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems
Codebook Agent: 大语言模型多智能体系统的摊销拓扑设计
diffusion
扩散模型相关
Abstract
Adapting the communication topology of an LLM multi-agent system to each query improves both accuracy and efficiency, yet current designers treat this as conditional graph generation: a variational, autoregressive, or diffusion decoder searches the $N \times N$ adjacency space, and a graph-network proxy trained on utility and a structural cost such as edge count ranks the sampled candidates. We argue that this formulation is misaligned with the problem. Empirically, topologies that survive a reward filter collapse to about six distinct graphs even when the codebook capacity grows from 8 to 64; edge count is negatively correlated with measured token consumption (Pearson $r \approx -0.4$), so sparsifying the graph makes inference more expensive; and a message-passing scorer over agent-profile nodes is adjacency-invariant whenever agents share a profile---the default configuration of published benchmarks---so it cannot rank candidates at all in that regime. These three facts motivate Codebook Agent: a vector-quantized autoencoder compresses successful topologies into a query-independent 16-entry codebook; a reward-weighted MLP maps the query embedding to a distribution over codes; and an MLP proxy that reads the flattened adjacency, regressed on measured utility and per-task normalized token cost, reranks the top decoded candidates in a single batched forward pass. With no iterative search and no message passing at test time, Codebook Agent is the most accurate method on all six benchmarks we compare (84.6 average against 83.0 for the strongest prior designer), emits a topology in 2.4 ms, and uses 21.9--33.2% fewer LLM tokens.
Chinese Translation
使大语言模型多智能体系统的通信拓扑适应每个查询,既能提高准确性又能提高效率,然而当前的设计者将此视为条件图生成问题:一个变分、自回归或扩散解码器在 $N \times N$ 邻接空间中搜索,而一个在效用和结构成本(如边数)上训练得到的图网络代理对采样候选进行排序。我们认为这种表述与问题本身不一致。经验上,即使码本容量从 8 增长到 64,能够通过奖励过滤的拓扑也会坍缩为大约六种不同的图;边数与实测的 token 消耗呈负相关(Pearson $r \approx -0.4$),因此稀疏化图反而会使推理更加昂贵;并且,只要智能体共享一个档案——已发表基准测试的默认配置——那么一个在智能体档案节点上进行消息传递的评分器就是邻接不变的,因此在那种情况下它根本无法对候选进行排序。这三个事实催生了 Codebook Agent:一个向量量化的自编码器将成功的拓扑压缩到一个与查询无关的 16 条目码本中;一个奖励加权的 MLP 将查询嵌入映射为对码的分布;一个读取展平后邻接矩阵的 MLP 代理,在实测效用和每个任务归一化的 token 成本上回归,在单次批处理前向传播中对解码出的候选进行重新排序。由于测试时没有迭代搜索也没有消息传递,Codebook Agent 在我们比较的所有六个基准上都是最准确的方法(平均 84.6,而最强先前设计者为 83.0),在 2.4 ms 内生成一个拓扑,并使用 21.9–33.2% 更少的 LLM token。
cs.AI / 9 / 2609.02273
CoMerge: Conflict-Driven Preference Optimization for Multi-Task Model Merging
CoMerge:用于多任务模型合并的冲突驱动偏好优化
large language model
大语言模型相关
Abstract
Model merging provides an efficient paradigm for constructing multi-task large language models (LLMs) without full model retraining, yet it remains challenged by parameter interference. While existing methods aim to preserve the capabilities of individual expert models and mitigate interference, they generally do not directly learn from the potentially degraded behaviors exposed by naive merging. In this paper, we propose a conflict-driven preference optimization framework for model merging (CoMerge), which reformulates model merging as a preference optimization problem. The approach utilizes a self-supervised, conflict-driven strategy that leverages the defects of naive merging methods (e.g., task arithmetic) as hard negative samples to construct preference pairs without external annotations. By applying preference optimization to refine lightweight, tensor-wise merging coefficients, CoMerge enables the model to mitigate parameter-space conflicts while preserving task-specific capabilities. Extensive experiments show that CoMerge achieves an average normalized performance of 0.9968 on MergeBench, outperforming all evaluated data-free and data-driven model-merging baselines. Furthermore, on Llama-3.1-8B-Instruct, CoMerge yields marked improvements on conflict-sensitive tasks such as instruction following and safety, while remaining highly competitive with full-parameter fine-tuning despite optimizing only 1,445 scalar coefficients.
Chinese Translation
模型合并在无需进行完整模型重训练的情况下,为构建多任务大语言模型(LLMs)提供了一种高效范式,但仍面临参数干扰的挑战。尽管现有方法旨在保留各个专家模型的能力并减轻干扰,但它们通常不会直接从朴素合并所暴露出的潜在退化行为中学习。在本文中,我们提出了用于模型合并的冲突驱动偏好优化框架(CoMerge),该框架将模型合并重新表述为偏好优化问题。该方法采用一种自监督、冲突驱动的策略,将朴素合并方法(例如任务算术)的缺陷作为硬负样本,在没有外部标注的情况下构建偏好对。通过应用偏好优化来细化轻量级、逐张量的合并系数,CoMerge 使模型能够减轻参数空间中的冲突,同时保留任务特定能力。大量实验表明,CoMerge 在 MergeBench 上达到了 0.9968 的平均归一化性能,超过了所有经评估的无数据模型合并基线和数据驱动模型合并基线。此外,在 Llama-3.1-8B-Instruct 上,CoMerge 在指令遵循和安全等冲突敏感任务上取得了显著改进,同时仅优化了 1,445 个标量系数,却仍与全参数微调保持高度竞争力。
cs.AI / 10 / 2609.02292
SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology
SCX路由器:使用解码器-KV分类器与现实世界任务本体进行流式零样本模型选择
large language model
大语言模型相关
Abstract
The rapid proliferation of large language models (LLMs) and the growing diversity of their applications presents a unique optimization opportunity: selecting the right model for the task, while optimizing for speed, cost, and quality at a per-task level. However, inference endpoints can vary widely in quality, price, latency, context support, tool use, domain expertise, and reasoning behavior. This heterogeneity makes manual heuristics difficult to maintain and unlikely to achieve consistently favorable speed--cost--quality trade-offs on their own. We introduce \router{}, a lightweight GLiClass-based router that assigns a suitability score to each inference-time model label without autoregressive generation. The released 0.6B-parameter checkpoint combines a Qwen3 decoder with a shallow bidirectional scorer. Its decoder-KV execution path preserves a text-only key--value cache across a session, encodes only new dialogue turns, and evaluates transient candidate-label tokens without adding them to the persistent cache. The same checkpoint also predicts task type, difficulty, reasoning mode, and expected output length, and supports custom zero-shot labels. For task generation, we construct a task ontology with 23 families, 115 task types, 345 routable subtypes, 1,173 synthetic examples, and an orthogonal axis of 30 domains. Using this structure, we generate 150,000 verifier-scored tasks and 15,000 open-ended tasks. We then train the Qwen3 decoder on these tasks, while explicitly separating learned request prediction from per-task policies for attributes such as eligibility, cost, cache reuse, safety, and sovereignty. Across six LiveBench subsets, the router outperforms the mean candidate; on the selected 1,000-task subset, it achieves an aggregate top-1 score of 0.707 versus 0.696 for the strongest fixed model, with benchmark-dependent gains.
Chinese Translation
大型语言模型(LLM)的快速普及及其应用场景日益多样化,带来了一个独特的优化机会:在单个任务层面选择最适合该任务的模型,同时优化速度、成本和质量。然而,不同推理端点在质量、价格、延迟、上下文支持、工具使用、领域专业知识和推理行为等方面可能存在很大差异。这种异质性使得人工启发式规则难以维护,且不太可能仅凭自身实现持续有利的速度-成本-质量权衡。我们介绍了SCX Router,一个基于GLiClass的轻量级路由器,它无需自回归生成即可为每个推理时的模型标签分配适合度分数。发布的0.6B参数检查点结合了Qwen3解码器与一个浅层双向评分器。其解码器-KV执行路径在会话中保留纯文本键-值缓存,仅编码新的对话轮次,并在不将临时候选标签令牌添加到持久缓存的情况下对其进行评估。同一检查点还预测任务类型、难度、推理模式和预期输出长度,并支持自定义零样本标签。对于任务生成,我们构建了一个任务本体,包含23个任务族、115种任务类型、345个可路由子类型、1,173个合成示例以及30个领域的正交轴。利用这一结构,我们生成了150,000个由验证器评分的任务和15,000个开放式任务。然后,我们在这些任务上训练Qwen3解码器,同时将学习到的请求预测与针对资格、成本、缓存复用、安全和主权等属性的每任务策略明确分开。在六个LiveBench子集上,该路由器的表现优于平均候选模型;在选定的1,000个任务子集上,它实现了0.707的总体top-1得分,而最强固定模型为0.696,且增益依赖于基准。
cs.AI / 11 / 2609.02421
UTP-Bench: Uncertainty-aware Travel Planning Benchmark
UTP-Bench:不确定性感知的旅行规划基准
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) have recently demonstrated strong capabilities in automated travel itinerary generation. However, real- world travel planning is inherently uncertain: transportation delays, crowd fluctuations, and unexpected stochastic delays frequently inval- idate otherwise feasible schedules. Existing benchmarks like TravelPlanner and TripCraft assume deterministic environments, evaluating only static constraint satisfaction and ignoring whether generated plans remain robust when such uncertainties arise. To address this limitation, we introduce UTP-Bench1 , a large-scale benchmark for uncertainty-aware travel planning. The dataset integrates real-world travel data spanning 504 cities of India, including attractions, restau- rants, accommodations, and multi-modal trans- portation networks. To model realistic disrup- tions, UTP-Bench incorporates empirical delay distributions and crowd-density patterns col- lected from major cities, enabling evaluation of travel plans under stochastic conditions. We further propose three evaluation metrics, namely Buffer Adequacy Score (BAS), Crowd- Aware Timing Score (CATS), and Transport Delay Absorption Score (TDAS), which quan- tify the ability of generated itineraries to main- tain robustness against transit delays and crowd variability. Experiments with state-of-the-art LLMs like GPT-5, Qwen3, Mistral and Phi-4 re- veal substantial gaps between model-generated and human-authored plans, particularly in tem- poral buffering, delay-aware transportation scheduling, and crowd-sensitive planning.
Chinese Translation
近年来,大型语言模型(LLMs)在自动化旅行行程生成方面展现出了强大的能力。然而,现实世界的旅行规划本质上具有不确定性:交通延误、人流波动以及意外的随机延误经常会使原本可行的日程安排失效。现有的基准测试(如TravelPlanner和TripCraft)假定环境是确定性的,仅评估静态约束满足情况,而忽略了当此类不确定性出现时生成的计划是否依然稳健。为解决这一局限,我们提出了UTP-Bench1,一个用于不确定性感知旅行规划的大规模基准测试。该数据集整合了覆盖印度504个城市的现实世界旅行数据,包括景点、餐厅、住宿以及多模式交通网络。为了对现实干扰进行建模,UTP-Bench纳入了从主要城市收集的经验延误分布和人群密度模式,从而能够在随机条件下评估旅行计划。我们进一步提出了三个评估指标,即缓冲充足度评分(BAS)、人群感知时间评分(CATS)和交通延误吸收评分(TDAS),这些指标量化了生成行程在交通延误和人群变化下保持稳健性的能力。使用GPT-5、Qwen3、Mistral和Phi-4等最先进LLMs进行的实验揭示了模型生成计划与人类编写计划之间的显著差距,尤其是在时间缓冲、延误感知的交通调度以及人群敏感规划方面。
cs.AI / 12 / 2609.02649
Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting
Loom:通过嵌入空间重新加权将诊断线索编织为自由文本共识
large language model
大语言模型相关
Abstract
Aggregating noisy, conflicting textual hypotheses into a reliable consensus is a fundamental challenge when deploying NLP systems in real-world industrial settings. While monolithic Large Language Model (LLM) agents offer unbounded expressivity for tasks like Root Cause Analysis (RCA), they suffer from context limits, compounding hallucinations, and prohibitive inference latency. Traditional weak supervision offers statistical rigor but is mathematically restricted to discrete classes. We present Loom, a generative consensus framework deployed for real-world RCA that bridges these paradigms. Loom aggregates open-form hypotheses emitted by modular heuristics (diagnostic templates dynamically populated with episode-specific entities, times, and metrics) by projecting them into a continuous embedding space, and resolves conflicting signals with an iterative centroid-based reweighting algorithm. The resulting consensus weights ground a single lightweight LLM synthesis step. Evaluated on the OpenRCA benchmark, Loom occupies the accuracy--efficiency Pareto frontier: it matches a state-of-the-art autonomous agent on Bank and Market-2 and trails on Market-1 and Telecom, while using a single LLM call per incident on all four datasets ($\sim$26$\times$ faster; $\sim$33$\times$ with an 8B-parameter synthesizer). We discuss our deployment experience, highlighting lessons learned regarding the trade-offs between agentic depth and inference latency, negative results in redundancy detection, and how deterministic consensus fosters trust among Subject Matter Experts~(SMEs).
Chinese Translation
在现实世界的工业环境中部署NLP系统时,将嘈杂、相互冲突的文本假设聚合为可靠的共识是一项基本挑战。虽然 monolithic 大型语言模型(LLM)智能体在根本原因分析(RCA)等任务中提供了无限制的表达能力,但它们受限于上下文长度、复合幻觉和过高的推理延迟。传统的弱监督方法具有统计严谨性,但在数学上仅限于离散类别。我们提出了 Loom,一个为现实世界 RCA 部署的生成式共识框架,它弥合了这些范式。Loom 将通过模块化启发式方法(由事件特定实体、时间和指标动态填充的诊断模板)生成的开放式假设投影到连续嵌入空间中,并使用基于迭代质心的重新加权算法来解决冲突信号。由此产生的共识权重支撑了一个轻量级 LLM 的单一合成步骤。在 OpenRCA 基准上的评估表明,Loom 占据了准确性-效率的帕累托前沿:它在 Bank 和 Market-2 上匹配了最先进的自主智能体,在 Market-1 和 Telecom 上略有落后,同时在这四个数据集上每次事件仅使用一次 LLM 调用(速度提升约 $\sim$26$\times$;使用 8B 参数合成器时提升约 $\sim$33$\times$)。我们讨论了部署经验,重点强调了关于智能体深度与推理延迟之间权衡的教训、冗余检测中的负面结果,以及确定性共识如何培养主题专家(SME)信任的见解。
cs.AI / 13 / 2609.02707
Door-in-the-Face Requests and Refusal Behaviour in Large Language Models
大型语言模型中的“以退为进”请求与拒绝行为
large language model
大语言模型相关
Abstract
Does the door-in-the-face technique work on language models? In humans, a large request that is refused makes a smaller follow-up request more likely to be granted. We test this on nine production models from three providers: each model refuses a large request, then receives a smaller version of the same request, and we compare its compliance with asking directly. The answer depends on the model. On Anthropic's frontier models the technique works: Opus 5 answers the smaller request 65.8% of the time after refusing the larger one, against 29.3% when asked directly. On the frontier models of OpenAI and Google, and on Haiku 4.5, it backfires, lowering compliance by 15.5 to 23.0 points. A control locates the effect: a refused large request on an unrelated topic does less than the related one on all nine models, so the concession itself matters everywhere, while the reaction to having just refused something differs by model family. The technique does not transfer to refusals drawn from public benchmarks. What decides whether a retreat can work is what the request asks for: rewriting 265 refused requests for usable instructions into requests for explanations of the same topic removed the refusal in 263 cases. Human influence techniques port to language models one model family at a time.
Chinese Translation
以退为进技术对语言模型有效吗?对人类而言,一个被拒绝的大请求会让随后提出的较小请求更有可能被应允。我们在来自三家提供商的九个生产模型上对此进行了测试:每个模型先拒绝一个大请求,然后收到同一请求的较小版本,我们将其遵从度与直接提出该请求时的遵从度进行比较。答案因模型而异。在Anthropic的前沿模型上,该技术有效:Opus 5在拒绝了较大请求后,对较小请求的应允率为65.8%,而直接请求时这一比例为29.3%。在OpenAI和Google的前沿模型上,以及Haiku 4.5上,它适得其反,使遵从度降低了15.5到23.0个百分点。一项对照实验定位了这一效应:在所有九个模型上,与主题无关的被拒绝的大请求所产生的效果小于与主题相关的被拒绝的大请求,因此让步本身在所有模型上都很重要,而对刚刚拒绝过某个请求之后的反应则因模型家族而异。该技术无法迁移到从公共基准中提取的拒绝案例上。决定一次退让能否奏效的是请求本身要求什么:将265个被拒绝的、要求提供可用指令的请求改写为对同一主题进行解释的请求,在263个案例中消除了拒绝。人类影响技巧向语言模型的迁移是逐模型家族进行的。
cs.AI / 14 / 2609.02805
Large Language Models (LLMs) for Telecom Root Cause Analysis (RCA): A Structured Reasoning Framework for Evidence-Grounded Diagnosis
面向电信根因分析(RCA)的大语言模型(LLMs):一种基于证据的诊断的结构化推理框架
large language model
大语言模型相关
Abstract
Root cause analysis (RCA) is a critical task in telecom network operations, but diagnosing performance degradations in modern 5G and emerging 6G networks remains challenging due to complex cross-layer dependencies. While large language models (LLMs) offer promising capabilities for reasoning and knowledge integration, directly applying vanilla LLMs to telecom RCA often leads to hallucination, unstable reasoning, and poor alignment with structured network evidence. This work first reviews the evolution of telecom RCA from rule-based and machine learning (ML) approaches to emerging LLM-enabled techniques, and provides an overview of recent paradigms, including structured reasoning, retrieval-augmented knowledge grounding, agentic orchestration, and verifiable reasoning. Building upon these insights, we propose a structured reasoning framework for LLM-enabled telecom RCA that aligns diagnostic reasoning with telecom-specific evidence and domain knowledge. The proposed approach first organizes heterogeneous network telemetry into canonical contexts, and then enforces decision-path reasoning during diagnosis, and finally generates evidence-grounded explanations for reliable fault identification. Experimental results on two 5G RCA datasets, TeleLogs and TelecomTS, demonstrate that the proposed framework consistently improves diagnostic accuracy and decision consistency compared with baseline techniques. These cross-dataset results highlight the importance of structured reasoning design for practical LLM-based RCA systems in next-generation telecom networks.
Chinese Translation
根因分析(RCA)是电信网络运维中的关键任务,但由于复杂的跨层依赖性,诊断现代5G和新兴6G网络中的性能退化仍然具有挑战性。尽管大语言模型(LLMs)在推理和知识整合方面展现出有前景的能力,但将原始LLM直接应用于电信RCA往往会导致幻觉、不稳定的推理以及与结构化网络证据的较差对齐。本工作首先回顾了电信RCA从基于规则和机器学习(ML)方法到新兴的LLM赋能技术的演变,并概述了近期范式,包括结构化推理、检索增强的知识锚定、智能体编排和可验证推理。基于这些见解,我们提出了一种面向LLM赋能的电信RCA的结构化推理框架,该框架将诊断推理与电信特定证据和领域知识对齐。所提出的方法首先将异构网络遥测数据组织为规范上下文,然后在诊断过程中强制执行决策路径推理,最后生成基于证据的解释以实现可靠的故障识别。在两个5G RCA数据集(TeleLogs和TelecomTS)上的实验结果表明,与基线技术相比,所提出的框架持续提高了诊断准确性和决策一致性。这些跨数据集结果凸显了结构化推理设计对于下一代电信网络中实用的基于LLM的RCA系统的重要性。
cs.CL / 15 / 2609.01788
VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages
VakyArth:评估大语言模型在印度诸语言中的语用能力
large language model
大语言模型相关
Abstract
Real-world communication often requires pragmatic reasoning: interpreting meanings implied through context and cultural convention rather than stated literally. Existing pragmatic evaluation remains largely limited to English and high-resource languages, leaving Indic languages unexplored despite their linguistic and cultural diversity. We introduce VakyArth, the first pragmatic benchmark for Indic languages, designed as a diagnostic evaluation covering Hindi, Punjabi, Tamil, and Malayalam. VakyArth evaluates models across five phenomena: deixis, speech acts, implicature, social pragmatics, and coherence; through multiple-choice questions, natural language inference, and translation, with all items authored by native speakers. Across multilingual large language models (LLMs) of varying families and sizes, we find consistent failures on pragmatic meanings rooted in Indic linguistic and cultural conventions. Our analysis shows systematic differences across languages and tasks: MCQ accuracy exceeds NLI accuracy in all model-language combinations, translation performance does not reliably track pragmatic understanding, and Indo-Aryan languages show a translation advantage over Dravidian languages. We further show that automatic translation metrics can miss fluent but pragmatically unfaithful outputs, especially for implicature and deixis.
Chinese Translation
现实世界中的交流常常需要语用推理:通过语境和文化惯例来理解隐含的意义,而非字面上的直白陈述。现有的语用评估在很大程度上仅限于英语和高资源语言,尽管印度诸语言具有语言和文化多样性,却仍未得到探索。我们提出了VakyArth,这是首个面向印度诸语言的语用基准,设计为一项覆盖印地语、旁遮普语、泰米尔语和马拉雅拉姆语的诊断性评估。VakyArth通过多项选择题、自然语言推理和翻译来评估模型在五种现象上的表现:指示语、言语行为、含意、社会语用和连贯性;所有题目均由母语者编写。在跨越不同家族和规模的多语言大语言模型(LLMs)中,我们发现模型在与印度语言和文化惯例相关的语用意义上存在一致的失败。我们的分析显示,不同语言和任务之间存在系统性差异:在所有模型-语言组合中,多项选择题(MCQ)的准确率均高于自然语言推理(NLI)的准确率;翻译性能并不能可靠地反映语用理解;印度-雅利安语族相对于达罗毗荼语族表现出翻译优势。我们进一步表明,自动翻译指标可能漏判流畅但语用上不忠实的输出,尤其是在含意和指示语方面。
cs.CL / 16 / 2609.01798
How Do Prompt Variations Affect Energy Consumption in On-Device LLMs?
提示词变化如何影响设备端大语言模型的能耗?
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly deployed on mobile devices, making energy efficiency a key deployment constraint, yet the energy impact of prompt design remains underexplored. This paper aims to understand how two prompt properties, cognitive load and phrasing pattern, shape the energy behavior of on-device LLM inference. We conduct a broad empirical study covering prompt properties, datasets, models, and devices, with phase-level profiling that separates prefill and decode energy. We find that cognitive load primarily affects the energy cost per token, while phrasing pattern affects energy largely through token usage. Our energy-quality analysis further shows that prompt design reshapes the attainable frontier differently across models, highlighting the need for model-aware prompt design in energy-efficient on-device LLM inference. Code, datasets, and scripts are available at https://amai-gsu.github.io/PromptProperty/.
Chinese Translation
大语言模型(LLMs)正越来越多地部署在移动设备上,使能效成为关键的部署约束,然而提示词设计对能耗的影响仍未得到充分探索。本文旨在理解两种提示词属性(认知负荷和措辞模式)如何塑造设备端大语言模型推理的能耗行为。我们开展了一项广泛的实证研究,涵盖提示词属性、数据集、模型和设备,并采用阶段级剖析,将预填充与解码能耗分开。我们发现,认知负荷主要影响每个标记(token)的能耗成本,而措辞模式则主要通过标记使用量影响能耗。我们的能耗-质量分析进一步表明,提示词设计在不同模型间以不同方式重塑可实现的前沿,凸显了在能效型设备端大语言模型推理中需要采用面向模型的提示词设计。代码、数据集和脚本可在 https://amai-gsu.github.io/PromptProperty/ 获取。
cs.CL / 17 / 2609.01832
Interpretable Symptom Vectors for Depression in a Large Language Model
大型语言模型中抑郁症的可解释症状向量
large language model
大语言模型相关
Abstract
Patients with depression present with diverse symptom profiles, yet clinical practice routinely reduces this variation to a single severity score. Large language models (LLMs) can potentially capture various symptoms and their severity from patient speech. However, how depressive symptoms are represented inside LLMs remains poorly understood, limiting clinical trust. To examine whether internal model activations match clinician judgment, we analyzed the residual stream of Gemma-3-27B-PT using mechanistic interpretability techniques. Recording activations across symptom descriptions drawn from validated clinical instruments, we found that symptom groups geometrically separated the most at layer 21 across multiple distance metrics. Using Semantic Projection, we then projected held-out naturalistic text onto Symptom Vectors constructed from these instruments. The resulting per-symptom coefficients preserved clinician-annotated rank ordering across mood, somatic, and suicidality axes. Furthermore, a single depression vector in Layer 21 separates held-out depressive from non-depressive text (AUC = 0.789), which can be used as an emotional valence gate that restricts symptom projection to depressive speech. These results reveal a decorrelated, clinician-aligned symptom signal readable directly from internal activations, offering a mechanistic foundation for interpretable depression-assessment tools.
Chinese Translation
抑郁症患者呈现出多样的症状特征,但临床实践通常将这些差异简化为单一的严重程度评分。大型语言模型(LLM)有可能从患者言语中捕捉各种症状及其严重程度。然而,抑郁症状在LLM内部是如何表示的仍然知之甚少,这限制了临床信任。为了检验内部模型激活是否与临床医生的判断一致,我们使用机制可解释性技术分析了Gemma-3-27B-PT的残差流。记录了来自经过验证的临床工具的样本症状描述上的激活后,我们发现症状组在层21处通过多种距离度量在几何上分离最为明显。利用语义投影,我们将留出的自然主义文本投影到由这些工具构建的症状向量上。所得的每个症状的系数在情绪、躯体及自杀倾向轴线上保持了临床医生标注的等级顺序。此外,层21中的单一抑郁向量区分了留出的抑郁文本与非抑郁文本(AUC = 0.789),该向量可用作限制症状投影的情感效价门,仅针对抑郁言语进行分析。这些结果揭示了一种去相关的、与临床医生一致的症状信号,可直接从内部激活中读取,为可解释的抑郁评估工具提供了机制基础。
cs.CL / 18 / 2609.01867
Thinking effort aligns between humans and reasoning models in abductive reasoning
在溯因推理中,思考努力在人类与推理模型之间是对齐的
large language model
大语言模型相关
Abstract
A major question in cognitive modeling concerns the behavioral alignment between large language models and humans across linguistic and non-linguistic tasks. Unlike standard LLMs, large reasoning models (LRMs) are optimized with reinforcement learning from verifiable rewards, encouraging correct solutions to reasoning tasks rather than preference-aligned responses. Recent work (de Varda et al., 2025) investigates the cost of thinking in humans and LRMs by comparing human reaction times with model reasoning traces across a range of reasoning tasks. We isolate this alignment by turning to abductive reasoning: unlike deductive tasks, its difficulty cannot be inferred from formal structure and offers no shortcuts a model could exploit to mimic effort without genuine search, providing firmer ground for empirical claims of shared effort. We find further evidence of alignment between LRM and human reasoning effort, as well as evidence that models and humans tend to make similar errors. Finally, we show that decoding methods that let models explore multiple reasoning paths increase alignment in reasoning cost between humans and LRMs across the three models tested.
Chinese Translation
认知建模中的一个主要问题涉及大型语言模型与人类在语言和非语言任务中的行为对齐。与标准大型语言模型(LLMs)不同,大型推理模型(LRMs)通过基于可验证奖励的强化学习进行优化,鼓励对推理任务给出正确解决方案,而非偏好对齐的回应。近期研究(de Varda et al., 2025)通过在一系列推理任务中比较人类反应时间与模型推理轨迹,考察了人类和LRMs的思考成本。我们通过转向溯因推理来单独考察这种对齐:与演绎任务不同,其难度无法从形式结构推断出来,也不提供可供模型利用的捷径,使其无需真正搜索即可模仿努力;这为关于共享思考努力的经验论断提供了更坚实的基础。我们发现了LRM与人类推理努力之间一致性的进一步证据,也发现了模型和人类往往犯类似错误的证据。最后,我们表明,让模型探索多个推理路径的解码方法在测试的三种模型中提高了人类与LRMs之间推理成本的一致性。
cs.CL / 19 / 2609.01971
NS-Copilot: An LLM-Driven Agent System for Autonomous Neuroscience Analysis
NS-Copilot:一种用于自主神经科学分析的LLM驱动智能体系统
large language model
大语言模型相关
Abstract
AI is rapidly advancing neuroscience, yet many laboratories fail to fully unleash its potential due to significant interdisciplinary barriers. While pre-trained neural models for physiological data are progressing quickly, their heterogeneous architectures and modality-specific constraints hinder systematic integration, selection, and evaluation. Despite recent advances in large language model (LLM)-based agent systems for intelligent scientific applications, existing approaches often still lack the domain expertise required to effectively select and coordinate diverse neuroscience pre-trained models and handle unique data types in this domain. We present NS-Copilot, an LLM-driven multi-agent system for neuroscience analysis that autonomously supports end-to-end workflows for diverse professional tasks. It unifies domain-specific pre-trained models and supports key neuroscience modalities, including EEG and extracellular spike data, through a natural-language interface. Given raw data and a task description, NS-Copilot orchestrates agents with specialized roles for planning, adaptive control, code generation, and result synthesis, enabling analysis without dataset-specific heuristics. We evaluate NS-Copilot on neuroscience benchmarks spanning Alzheimer's disease, Parkinson's disease, and working memory spike decoding. Across 8 trials per task, the system consistently outperforms strong baselines on the primary metric, demonstrating the ability of NS-Copilot for effective and scalable neuroscience analysis.
Chinese Translation
人工智能正在快速推动神经科学的发展,然而由于显著的跨学科障碍,许多实验室未能充分发挥其潜力。尽管用于生理数据的预训练神经模型进展迅速,但其异构架构和模态特定约束阻碍了系统的集成、选择和评估。尽管基于大型语言模型(LLM)的智能体系统在智能科学应用方面近期取得了进展,但现有方法通常仍缺乏有效选择和协调多种神经科学预训练模型以及处理该领域独特数据类型所需的领域专业知识。我们提出了NS-Copilot,一个用于神经科学分析的、由LLM驱动的多智能体系统,能够自主支持多种专业任务的端到端工作流程。它统一了领域特定的预训练模型,并通过自然语言接口支持关键的神经科学模态,包括脑电图(EEG)和细胞外尖峰数据。给定原始数据和任务描述,NS-Copilot编排具有专门角色的智能体,用于规划、自适应控制、代码生成和结果综合,从而在没有数据集特定启发式规则的情况下实现分析。我们在涵盖阿尔茨海默病、帕金森病和工作记忆尖峰解码的神经科学基准上评估了NS-Copilot。在每个任务8次试验中,该系统在主要指标上持续优于强基线,展示了NS-Copilot在有效且可扩展的神经科学分析方面的能力。
cs.CL / 20 / 2609.02054
A Tri-Agent Framework for Evaluating and Aligning Question Clarification Capabilities of Large Language Models
评估和对齐大型语言模型问题澄清能力的三智能体框架
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) are increasingly deployed in interactive systems where understanding user intent precisely is paramount. A key capability for such systems is effective question clarification, especially when user queries are ambiguous or underspecified. This paper introduces a novel tri-agent framework for the robust evaluation of an LLM's ability to engage in clarifying dialogue. Our framework comprises three distinct LLM-based agents: (1) a Question Clarifying Agent (QCA), the system under evaluation, tasked with identifying ambiguities and posing clarifying questions; (2) a Respondent Agent (RA), designed to simulate human user responses, potentially including irrelevant or challenging replies; and (3) an Evaluator Agent (EA), an LLM-as-a-judge, which assesses the quality of the dialogue based on a comprehensive set of metrics. We detail a methodology for synthetic data generation in the supply chain domain as an example. We propose metrics evaluating ambiguity handling, question quality, dialogue efficiency, language appropriateness, and final intent alignment. We also briefly discuss the validation of the EA against human judgments. This work provides a structured approach to benchmark, validate, and improve the clarification capabilities of conversational LLM applications.
Chinese Translation
大型语言模型(LLMs)越来越多地部署在交互式系统中,在这些系统中,精确理解用户意图至关重要。这类系统的一个关键能力是有效的问题澄清,尤其是在用户查询模糊或说明不充分时。本文介绍了一种新颖的三智能体框架,用于稳健评估LLM参与澄清对话的能力。我们的框架由三个不同的基于LLM的智能体组成:(1)问题澄清智能体(QCA),即被评估的系统,负责识别歧义并提出澄清性问题;(2)应答智能体(RA),旨在模拟人类用户响应,可能包括无关或具有挑战性的回复;(3)评估智能体(EA),一个LLM作为评判者,基于一组综合指标评估对话质量。我们以供应链领域为例,详细介绍了一种合成数据生成方法。我们提出了评估歧义处理、问题质量、对话效率、语言适当性和最终意图对齐的指标。我们还简要讨论了针对人类判断对EA进行的验证。这项工作为基准测试、验证和改进对话式LLM应用的澄清能力提供了一种结构化方法。
cs.CL / 21 / 2609.02056
HyGRAIL: Cost-Aware and Evidence-Grounded Scientific Hypothesis Discovery over Knowledge Graphs
HyGRAIL:知识图谱上成本感知与证据支撑的科学假设发现
large language model
大语言模型相关
Abstract
Scientific knowledge graphs organize entities and relations extracted from scientific literature, but they remain inherently incomplete. Missing typed links in such graphs can therefore represent plausible scientific hypotheses, such as unexplored associations between materials and applications. However, scientific hypothesis discovery is challenging because true discoveries are extremely sparse among typed candidate pairs: graph neural networks (GNNs) are efficient but unreliable for ambiguous cases, while large language models (LLMs) are knowledgeable but too costly to apply exhaustively and are not naturally grounded in graph structures. We propose HyGRAIL, a cost-aware and evidence-grounded framework that combines heterogeneous GNN triage with LLM-based hypothesis review. HyGRAIL first uses a GNN to score candidate hypotheses and identify a validation-calibrated ambiguous region, routing only graph-uncertain cases to LLM review. For each routed hypothesis, HyGRAIL retrieves node-level associations and multi-hop relational paths from the knowledge graph (KG), then converts this structured evidence into natural language through template-based or LLM-based naturalization. An LLM review agent finally judges each hard hypothesis using the naturalized evidence and validation-selected decision criteria. On MatKG, HyGRAIL achieves the best F1 score of 0.429, improving over the strongest prior baseline by 0.242 F1 points and over the GNN-only baseline by 0.322. Meanwhile, GNN triage reduces the LLM call rate by 54.36% on average. Ablation studies further show that retrieved graph evidence is crucial for reliable hypothesis verification and that compact, two-sided evidence is more effective than simply increasing retrieval quantity.
Chinese Translation
科学知识图谱组织从科学文献中提取的实体和关系,但它们本质上仍然不完整。因此,这些图谱中缺失的类型化链接可以代表合理的科学假设,例如材料与应用之间未被探索的关联。然而,科学假设发现具有挑战性,因为真正的发现在类型化候选对中极其稀疏:图神经网络(GNN)效率高,但在模糊案例上不可靠,而大型语言模型(LLM)知识丰富,但详尽应用成本过高,并且不能自然地植根于图结构。我们提出HyGRAIL,一个成本感知且证据支撑的框架,它将异构图神经网络分诊与基于LLM的假设审查相结合。HyGRAIL首先使用GNN对候选假设进行评分,并识别出经验证校准的模糊区域,仅将图不确定的案例路由至LLM审查。对于每个被路由的假设,HyGRAIL从知识图谱(KG)中检索节点级关联和多跳关系路径,然后通过基于模板或基于LLM的自然化处理,将此结构化证据转换为自然语言。一个LLM审查代理最后使用自然化证据和经验证选择的决策标准来判断每个困难假设。在MatKG上,HyGRAIL达到了最佳F1分数0.429,比先前最强基线提升了0.242个F1点,比仅使用GNN的基线提升了0.322。同时,GNN分诊平均将LLM调用率降低了54.36%。消融研究进一步表明,检索到的图证据对于可靠的假设验证至关重要,并且紧凑的双面证据比单纯增加检索数量更有效。
cs.CL / 22 / 2609.02089
IDEEA: training-free Input-Dependent stEEring via Activation cluster matching
IDEEA:基于激活簇匹配的免训练输入相关型引导
large language model
大语言模型相关
Abstract
Steering aligns large language models (LLMs) by injecting a bias into selected activations at inference time, offering a far cheaper alternative to weight-update methods such as supervised fine-tuning or reinforcement learning. However, most existing training-free steering methods are input-independent: a single direction is fitted once and shared across all inputs. This is fundamentally limiting as different inputs occupy different regions of the activation space and admit different optimal steering directions toward the same target concept, much as the gradient with respect to a fixed loss varies from input to input. We close this gap with IDEEA (Input-Dependent stEEring via Activation cluster matching), a training-free framework for input-dependent steering. IDEEA clusters the positive and negative activation supports per attention head, and solves an optimal-matching problem to construct a set of cluster-conditional directions, all about the target concept. At inference time, it picks from this pool of directions and uses the one that best matches the input's own activation for steering. IDEEA aligns the model toward the target concept while preserving the input's original representation, evidence that activations encoding a concept occupy several distinct sub-regions of the representation space rather than a single one. IDEEA improves the truth $\times$ info rate in TruthfulQA by an average of 9.9% (up to 23.5%) over the best input-independent baseline.
Chinese Translation
引导(Steering)通过在推理时向选定的激活中注入偏置来对齐大型语言模型(LLM),与监督微调或强化学习等权重更新方法相比,提供了一种成本远为更低的替代方案。然而,大多数现有的免训练引导方法是输入无关型的:一条单一的方向被拟合一次,并在所有输入间共享。这从根本上受到了限制,因为不同输入占据激活空间中不同的区域,并且在朝向同一目标概念时,各输入所对应的最优引导方向也会有所不同,这与固定损失函数相对于不同输入的梯度各不相同颇为类似。我们通过 IDEEA(基于激活簇匹配的输入相关型引导)弥补了这一缺口,这是一个用于输入相关型引导的免训练框架。IDEEA 对每个注意力头的正负激活支撑集进行聚类,并求解一个最优匹配问题,以构造一组针对目标概念、按簇条件化的有向方向。在推理时,它从这一方向池中选出与输入自身激活最为匹配的一个方向用于引导。IDEEA 在将模型朝目标概念对齐的同时,保留了输入原有的表征,这证明了编码某概念的激活在表征空间中占据若干个不同的子区域,而非仅仅一个单一区域。在 TruthfulQA 上,IDEEA 将 truth $\times$ info rate 相对于最优的输入无关型基线平均提升了 9.9%(最高可达 23.5%)。
cs.CL / 23 / 2609.02091
Selective Knowledge Edit Reversal via Gated Singular Vector Shrinkage
基于门控奇异向量收缩的选择性知识编辑逆转
large language model
大语言模型相关
Abstract
Knowledge editing provides an efficient way to update factual knowledge in large language models. However, malicious edits may introduce safety risks, making it necessary to reverse undesirable editing effects. Existing reversal methods for parameter-modifying edits mainly focus on global removal, which may also erase beneficial edits that should be preserved. In this paper, we study selective reversal of edited knowledge, where the goal is to reverse targeted edited facts while preserving the remaining edited facts. Based on the hypothesis that each edit is sparsely encoded within the dominant subspace of the edited matrix, we propose a spectral-based reversal framework that locates edit-sensitive components within the dominant singular subspace of edited weights. Experiments across multiple settings demonstrate the effectiveness of our method in reversing selected edits while preserving unrelated edited facts. These results suggest that different edits are sparsely encoded within dominant singular components and can be separable when the number of edits is moderate, making selective spectral reversal a promising direction for locating edit-specific components and repairing edited language models.
Chinese Translation
知识编辑为大型语言模型中的事实知识更新提供了一种高效方式。然而,恶意编辑可能引入安全风险,因此有必要逆转不良编辑效果。现有的针对参数修改型编辑的逆转方法主要聚焦于全局移除,这可能会同时抹除本应保留的有益编辑。本文研究已编辑知识的选择性逆转,目标是逆转目标编辑事实,同时保留剩余的编辑事实。基于每个编辑在编辑后矩阵的主导子空间中被稀疏编码的假设,我们提出了一种基于谱的逆转框架,该框架在编辑后权重的主导奇异子空间中定位编辑敏感分量。跨多种设置的实验证明了我们的方法在逆转选定编辑并保留无关已编辑事实方面的有效性。这些结果表明,不同的编辑在主导奇异分量中被稀疏编码,并且在编辑数量适中时是可分离的,这使得选择性谱逆转成为定位编辑特定分量和修复编辑后语言模型的一个有前景的方向。
cs.CL / 24 / 2609.02108
Predict, Don't Iterate: Efficient Adaptive-Length Infilling for Diffusion Language Models
预测而非迭代:扩散语言模型的高效自适应长度填充
diffusion
扩散模型相关
Abstract
Diffusion language models (DLMs) have emerged as a promising alternative to the auto-regressive paradigm. With bidirectional attention and any-order generation, DLMs naturally fit infilling tasks, which require generating a middle span conditioned on both the prefix and the suffix. However, infilling is sensitive to the length of the span, while DLMs require the length to be fixed before generation. Although prior studies extend DLMs to dynamic lengths, they still suffer from two limitations. (i) Sensitivity to initial length. These methods require a preset length to initialize the search and are highly sensitive to this initial length, often yielding suboptimal results. (ii) Inference inefficiency. They either insert length-changing operations during generation or repeatedly search for an appropriate length using multi-step denoising confidence, both of which introduce substantial extra forward passes and computational cost. Therefore, we propose PILL (Probing-based InfiLling with preset-Length-free decoding), an efficient infilling method for DLMs that requires no preset initial length and adds far fewer extra forward passes than baselines, substantially reducing inference time. Experiments show that, across five DLMs spanning different families, architectures, and training recipes on eight infilling benchmarks, PILL improves over the strongest baseline by +4.8 average pass rate on code and +6.0 BLEU-2 on text, while running 1.82x faster than that baseline. The code is available at https://github.com/Hsu1023/PILL.
Chinese Translation
扩散语言模型(DLMs)已成为自回归范式的一种有前景的替代方案。凭借双向注意力和任意顺序生成,DLMs天然适用于填充任务,即需要在给定前缀和后缀的条件下生成中间片段。然而,填充对片段的长度敏感,而DLMs要求在生成前固定长度。尽管先前的研究将DLMs扩展到动态长度,它们仍然存在两个局限性。(i) 对初始长度的敏感性。这些方法需要预设长度来初始化搜索,并且对该初始长度高度敏感,往往产生次优结果。(ii) 推理效率低下。它们要么在生成过程中插入长度调整操作,要么利用多步去噪置信度反复搜索合适的长度,这两种方式都会引入大量额外的前向传播和计算成本。因此,我们提出了PILL(基于探测的无预设长度解码的填充方法),一种高效的DLMs填充方法,它无需预设初始长度,并且比基线方法增加的前向传播次数少得多,从而大幅减少推理时间。实验表明,在跨越不同家族、架构和训练策略的五种DLMs上,针对八个填充基准,PILL在代码上的平均通过率比最强基线提升了+4.8,在文本上的BLEU-2提升了+6.0,同时运行速度比该基线快1.82倍。代码可在 https://github.com/Hsu1023/PILL 获取。
cs.CL / 25 / 2609.02115
text2ql: Multi-Target Natural Language Querying via a Language-Agnostic Intermediate Representation
text2ql:通过语言无关的中间表示进行多目标自然语言查询
large language model
大语言模型相关
Abstract
Natural language interfaces to databases have traditionally suffered from three structural limitations: exclusive targeting of relational SQL, unconditional dependence on large language model (LLM) inference at query time, and absence of any runtime signal when generated queries are semantically incorrect. This paper presents text2ql, an open-source Python framework that addresses all three limitations through a language-agnostic Intermediate Representation (QueryIR) and a pluggable renderer architecture. A single seven-stage detection pipeline serves both SQL and GraphQL targets; a zero-LLM deterministic mode delivers 100% execution accuracy at a median latency of 3.2 ms with no API cost; and every generated query carries a runtime confidence score in [0.15, 0.97] computed from an additive signal model. Evaluated on 50-query random samples from the Spider and BIRD benchmarks (indicative results; full-set evaluation is planned), the LLM-backed mode achieves 62-70% exact match and 84-91% execution accuracy; the deterministic mode achieves 100% execution accuracy with zero parse errors across all 100 test cases. An ablation study isolates schema-aware prompting as the dominant accuracy lever, contributing +18.4 percentage points of exact-match gain over the schema-free baseline on both benchmarks. text2ql is publicly available at https://pypi.org/project/text2ql/ under the Apache 2.0 license.
Chinese Translation
数据库的自然语言接口历来存在三个结构性局限:仅针对关系型 SQL、在查询时无条件依赖大型语言模型(LLM)推理,以及在生成的查询语义不正确时缺乏任何运行时信号。本文提出 text2ql,一个开源的 Python 框架,通过语言无关的中间表示(QueryIR)和可插拔的渲染器架构解决了上述所有三个局限。一个统一的七阶段检测流水线同时服务于 SQL 和 GraphQL 目标;一种零 LLM 的确定性模式以 3.2 ms 的中位延迟实现 100% 的执行准确率,且无 API 成本;并且每个生成的查询都携带一个由加性信号模型计算得出的运行时置信度分数,取值范围为 [0.15, 0.97]。在从 Spider 和 BIRD 基准中抽取的 50 条查询随机样本上进行的评估(指示性结果;完整集评估计划中),基于 LLM 的模式实现了 62-70% 的精确匹配率和 84-91% 的执行准确率;确定性模式在所有 100 个测试用例上实现了 100% 的执行准确率和零解析错误。一项消融研究将模式感知提示(schema-aware prompting)确定为主要的准确性杠杆,在两个基准上相对于无模式基线贡献了 +18.4 个百分点的精确匹配提升。text2ql 以 Apache 2.0 许可证公开发布于 https://pypi.org/project/text2ql/。
cs.CL / 26 / 2609.02122
AI agents reshape consensus formation in human groups
AI智能体重塑人类群体中的共识形成
large language model
大语言模型相关
Abstract
As large language model (LLM) agents shift from tools to participants in human groups, a fundamental question for collective behavior is how their growing presence reshapes consensus formation. Here we study mixed human-AI groups in a collaborative description game, in which shared conventions emerge through repeated rounds of random pairwise communication. Varying the proportions of LLM agents, we identify three distinct regimes of consensus formation: low agent proportions facilitate human-led consensus, intermediate proportions disrupt convergence, and high proportions restore strong consensus while shifting it toward agent-led conventions. Crucially, these regimes differ not only in the strength of convergence, but also in the semantic grounding and communicative form of the resulting consensus: human-led consensus is more concrete, holistic, and grounded in shared real-world analogies, whereas agent-led consensus is more abstract, less information-dense, and more geometrically segmented. Mechanistically, agent influence arises from a shared linguistic prior that places agents near one another in the expression space, combined with relatively stable expression choices across rounds; humans initially resist adopting expressions from partners perceived as AI but gradually yield to conformity pressure. These findings provide evidence that AI composition can shape the emergence, content, and perceived legitimacy of group norms, making agent proportion and transparency important design variables for human-AI systems.
Chinese Translation
随着大语言模型(LLM)智能体从工具转变为人类群体中的参与者,集体行为的一个基本问题是:它们日益增长的存在如何重塑共识形成。在此,我们研究了一个协作描述游戏中的混合人类-AI群体,在该游戏中,共享约定通过多轮随机配对交流而涌现。通过改变LLM智能体的比例,我们识别出共识形成的三种不同状态:低智能体比例促进人类主导的共识,中等比例破坏收敛,高比例则恢复强共识,同时使其转向智能体主导的约定。关键在于,这些状态不仅在收敛强度上存在差异,而且在由此产生的共识的语义基础和交流形式上也有所不同:人类主导的共识更加具体、整体化,并植根于共享的现实世界类比;而智能体主导的共识则更加抽象、信息密度较低,且在几何上更具分割性。从机制上讲,智能体的影响源于一种共享的语言先验,该先验使智能体在表达空间中彼此靠近,再加上各轮之间相对稳定的表达选择;人类起初会抵制采纳被感知为AI的伙伴的表达,但逐渐屈服于从众压力。这些发现为以下观点提供了证据:AI的构成能够塑造群体规范的涌现、内容及其被感知的合法性,从而使智能体比例和透明度成为人机系统的重要设计变量。
cs.CL / 27 / 2609.02153
A Layered Taxonomy for Chinese Learner Grammatical Error Annotation
中文学习者语法错误标注的分层分类法
large language model
大语言模型相关
Abstract
Grammatical error annotation in Chinese learner writing requires labels that are both consistent and linguistically meaningful. This paper proposes a layered scheme linking computational Chinese grammatical error correction (CGEC) with pedagogical error analysis. The scheme first identifies character- and punctuation-level orthographic errors, labeling them by edit operation and subtype. Other errors receive a three-layer core label combining edit operation, linguistic domain, and part of speech, with optional Chinese-specific extensions for aspect, modality, comparison, argument structure, and complements. Drawing on CGEC resources, learner-error taxonomies, and Mandarin grammar, the taxonomy is evaluated through a coverage analysis of automatically extracted MuCGEC edits and a preliminary consistency study in which five large language models apply it to a sample. The results support the layered approach while identifying category boundaries requiring further refinement.
Chinese Translation
中文学习者写作中的语法错误标注需要既一致又具有语言学意义的标签。本文提出了一种将计算中文语法纠错(CGEC)与教学错误分析联系起来的分层方案。该方案首先识别字级和标点级正字法错误,并按编辑操作和子类型对其进行标注。其他错误则获得一个三层核心标签,该标签结合了编辑操作、语言学领域和词性,并带有可选的针对体、情态、比较、论元结构和补足语的中文特有扩展。借鉴CGEC资源、学习者错误分类法和普通话语法,该分类法通过对自动提取的MuCGEC编辑的覆盖率分析以及一项初步一致性研究进行评估,其中五个大型语言模型将其应用于一个样本。结果支持分层方法,同时指出了需要进一步细化的类别边界。
cs.CL / 28 / 2609.02275
Do Large Language Models Capture the Diversity in their Training Data?
大型语言模型是否捕捉到了其训练数据中的多样性?
large language model
大语言模型相关
Abstract
Large language models are trained to model conditional distributions over text, yet it remains inadequately understood whether they capture the full diversity of plausible outputs present in their training data. We study this question through an information-theoretic lens by comparing the conditional entropy of model-generated outputs with that of the corresponding training data. Given paired input-output samples, we use conditional entropy and its matrix-based analogue based on von Neumann entropy to measure output variability beyond what is explained by the conditioning input, without requiring multiple reference outputs for the same prompt. Across LLM families with publicly available training data, including OLMo, Pythia, and GPT-Neo, we consistently find that model-generated outputs exhibit lower conditional entropy than their training data, across different model scales, sequence lengths, and decoding strategies. We observe a similar conditional diversity gap beyond language modeling, including class-conditioned ImageNet generators and text-conditioned models trained on MS-COCO. To address this gap, we propose a post-hoc correction mechanism that generates multiple outputs for each input and reweights them through a matrix-entropy projection, increasing conditional diversity while remaining close to the original model distribution. We prove the concavity of the matrix-based conditional entropy functional, which makes the resulting entropy-constrained projection a convex optimization problem, and develop a scalable mirror-descent algorithm for its implementation. Our results reveal a systematic conditional diversity gap between modern generative models and their training data, and provide an information-theoretic framework for measuring and mitigating this gap.
Chinese Translation
大型语言模型被训练用于对文本上的条件分布进行建模,但它们是否捕捉到了训练数据中存在的全部合理输出的多样性,这一问题仍未得到充分理解。我们通过信息论的视角来研究这一问题,比较模型生成输出的条件熵与相应训练数据的条件熵。给定成对的输入-输出样本,我们使用条件熵及其基于冯·诺依曼熵的矩阵类比,来衡量超出条件输入所能解释部分的输出变异性,而无需同一提示的多个参考输出。在具有公开训练数据的LLM家族中,包括OLMo、Pythia和GPT-Neo,我们一致发现,在不同模型规模、序列长度和解码策略下,模型生成输出的条件熵低于其训练数据的条件熵。我们观察到,在语言建模之外也存在类似的条件多样性差距,包括类别条件ImageNet生成器和在MS-COCO上训练的文本条件模型。为了解决这一差距,我们提出了一种事后修正机制,该机制为每个输入生成多个输出,并通过矩阵熵投影对其重新加权,从而在保持接近原始模型分布的同时增加条件多样性。我们证明了基于矩阵的条件熵泛函的凹性,这使得由此产生的熵约束投影成为一个凸优化问题,并开发了一种可扩展的镜像下降算法来实现它。我们的结果揭示了现代生成模型与其训练数据之间存在系统性的条件多样性差距,并提供了一个用于衡量和缩小这一差距的信息论框架。
cs.CL / 29 / 2609.02315
DiffIE: Diffusion-based Open Information Extraction
DiffIE:基于扩散的开放信息抽取
diffusion
扩散模型相关
Abstract
A single sentence often expresses multiple valid relational triplets, which makes Open Information Extraction (OpenIE) fundamentally a multi-output task. Existing neural systems handle this by autoregressive generation, which is flexible but slow and prone to redundancy, or by fixed-slot prediction, which is efficient but couples the extraction budget to training. We introduce DIFFIE which instead treats the stochasticity of conditional discrete diffusion as the extraction mechanism itself: independent reverse-diffusion trajectories over per-token role tags produce a pool of candidate triplets, which are clustered under lenient matching and ranked to form the output. Both the pool size and the number of returned extractions are inference-time choices, decoupling the extraction budget from training and exposing test-time compute as a tunable axis. DIFFIE achieves the new state of the art in CaRB (1-1) both F1 and AUC, and outperforms the strongest rule-based system (ClausIE) in BenchIE; it also remains competitive in standard CaRB and WiRe57 evaluations, giving the best average score among systems that report all four benchmarks. Ablations show that uniform discrete diffusion outperforms absorbing state diffusion in our setting, and that a matched non-diffusion stochastic tagger does not reproduce its gains. Our results indicate that diffusion stochasticity is an effective mechanism for structured prediction tasks with multiple valid outputs.
Chinese Translation
一个句子通常表达多个有效的关系三元组,这使得开放信息抽取(OpenIE)从根本上是一个多输出任务。现有的神经系统中,有些通过自回归生成来处理这一问题,这种方式灵活但速度慢且容易产生冗余;有些则通过固定槽位预测来处理,这种方式高效但抽取预算与训练绑定。我们提出了DIFFIE,它转而将条件离散扩散的随机性本身作为抽取机制:对每个标记的角色标签进行独立的逆扩散轨迹,产生候选三元组池,在宽松匹配下进行聚类并排序以形成输出。池大小和返回的抽取数量都是推理时的选择,将抽取预算从训练中解耦,并将测试时计算暴露为可调轴。DIFFIE在CaRB (1-1)的F1和AUC上均取得了新的最先进水平,并在BenchIE上优于最强的基于规则的系统(ClausIE);在标准CaRB和WiRe57评估中它也具有竞争力,在报告全部四个基准的系统间平均得分最高。消融实验表明,在我们的设置中,均匀离散扩散优于吸收态扩散,并且匹配的非扩散随机标注器无法再现其增益。我们的结果表明,扩散随机性是处理具有多个有效输出的结构化预测任务的有效机制。
cs.CL / 30 / 2609.02366
NE-R1: Enhancing Named Entity Recognition Model via Reinforcement Learning
NE-R1:通过强化学习提升命名实体识别模型
large language model
大语言模型相关
Abstract
Named Entity Recognition (NER) has achieved substantial progress since the advent of large language models (LLMs). Nevertheless, the recognition of long-tail and domain-specific entities remains challenging due to the deficiency in parametric knowledge. Retrieval-augmented generation (RAG) offers a promising remedy by injecting external knowledge, but it also introduces noise and unnecessary cost when dealing with familiar cases. In this paper, we propose NE-R1, a novel framework for adaptive retrieval-augmented NER. We design a "retrieval-on-demand" mechanism for NER. Then we integrate it into models by a two-stage training method: (1) multi-task instruction tuning initialization; (2) end-to-end RL optimization with CoT. To achieve reasonable selection between parameterized and external knowledge, we design a multi-dimensional reward considering both accuracy and retrieval benefit. NE-R1 achieves state-of-the-art performance on various benchmarks, with an average F1 score gain of 2.52% in in-domain evaluation and 1.18% in zero-shot cross-domain evaluation.
Chinese Translation
命名实体识别(NER)自从大型语言模型(LLMs)出现以来已经取得了显著的进展。然而,由于参数化知识的不足,对长尾和特定领域实体的识别仍然具有挑战性。检索增强生成(RAG)通过注入外部知识提供了一种有前景的解决方案,但在处理熟悉案例时也会引入噪声和额外的成本。在本文中,我们提出了NE-R1,一个用于自适应检索增强NER的新型框架。我们为NER设计了一种“按需检索”机制。然后我们通过两阶段训练方法将其集成到模型中:(1)多任务指令微调初始化;(2)结合思维链(CoT)的端到端强化学习优化。为了在参数化知识和外部知识之间做出合理选择,我们设计了一个综合考虑准确性和检索收益的多维奖励。NE-R1在多种基准上取得了最先进的性能,在域内评估中平均F1分数提升了2.52%,在零样本跨域评估中提升了1.18%。
cs.CL / 31 / 2609.02396
Improving Health Literacy through Lay Summarization of Radiological Reports: An Evaluation of BioNER and Retrieval-Augmented Generation
通过放射学报告的通俗摘要改善健康素养:BioNER与检索增强生成的评估
large language model
大语言模型相关
Abstract
Radiology reports are written primarily for clinicians, and their specialized terminology often makes them difficult for patients to interpret. As a result, many patients turn to publicly available Large Language Models (LLMs) to help explain their reports, despite well-documented risks of factual inaccuracies and hallucinations. Automated lay-summary generation has emerged as a promising alternative, yet the effectiveness of retrieval-enhanced and clinically informed approaches for radiology-specific communication remains underexplored. This study investigates the extent to which Retrieval-Augmented Generation (RAG) and Named Entity Recognition (NER) improve the quality, factual consistency, and readability of automatically generated lay summaries compared with standard LLM-based generation. We develop a framework combining NER-based extraction of clinically relevant findings with a RAG mechanism for contextual grounding, evaluated across few-shot and fine-tuned variants of two models (Qwen, BioBART). Results show that NER consistently improves readability and overall quality, while RAG alone offers no benefit and can introduce hallucinations from irrelevant retrieved terms. Combining RAG with NER degrades performance in few-shot settings but improves readability when fine-tuned. Fine-tuned BioBART with NER achieves the best overall performance, highlighting entity-aware extraction as the primary driver of improved patient-friendly summaries.
Chinese Translation
放射学报告主要面向临床医生编写,其专业术语常使患者难以理解。因此,许多患者求助公开可用的大型语言模型(LLMs)来解释其报告,尽管存在事实不准确和幻觉的有据可查的风险。自动通俗摘要生成已成为一种有前景的替代方案,然而,检索增强和临床信息感知的方法在放射学特定沟通中的有效性仍未被充分探索。与基于标准LLM的生成相比,本研究探讨了检索增强生成(RAG)和命名实体识别(NER)在多大程度上提高了自动生成的通俗摘要的质量、事实一致性和可读性。我们开发了一个框架,将基于NER的临床相关发现提取与用于上下文锚定的RAG机制相结合,并在两个模型(Qwen、BioBART)的少样本和微调变体上进行了评估。结果表明,NER持续改善可读性和整体质量,而单独使用RAG则没有益处,并可能因检索到不相关术语而引入幻觉。将RAG与NER相结合在少样本设置中会降低性能,但在微调时能提高可读性。微调后的BioBART结合NER取得了最佳整体性能,突显了实体感知提取作为改善患者友好摘要的主要驱动因素。
cs.CL / 32 / 2609.02438
When Decodability Is Not Enough: Logical Validity Representations, Behavioral Dissociation, and Causal Tests in Language Models
当可解码性不足时:语言模型中的逻辑有效性表征、行为分离与因果检验
large language model
大语言模型相关
Abstract
Large language models can look capable of logical reasoning, but correct or incorrect answers alone tell us little about what the model represents internally. We study logical verification in five open-weight transformer models using matched valid--invalid premise--claim pairs that vary across inference families, semantic domains, templates, and difficulty levels. Despite near-chance behavioral performance, logical validity is often almost perfectly decodable from hidden states and remains strongly decodable under held-out templates, domains, and inference families. Validity also remains highly decodable on behaviorally incorrect examples in the conditions where correctness-conditioned evaluation is well defined. At the same time, exhaustive leave-one-out tests reveal clear limits to this generalization, and interventions along probe-derived validity directions have only weak, nonspecific effects compared with random controls. Our results suggest that representing validity, expressing it in behavior, and using it causally are distinct. Validity related information can be strongly decodable from a model's hidden states without being reliably expressed in its output.
Chinese Translation
大型语言模型看起来能够进行逻辑推理,但仅凭答案的正确或错误并不能告诉我们模型内部表征了什么。我们使用匹配的有效-无效前提-论断对来研究五个开源权重 transformer 模型中的逻辑验证,这些前提-论断对在推理家族、语义领域、模板和难度水平上有所不同。尽管行为表现接近随机水平,逻辑有效性通常可以从隐藏状态中被近乎完美地解码,并且在未见过的模板、领域和推理家族条件下仍然具有强可解码性。在正确性条件评估有明确定义的情况下,有效性在行为错误的示例上也仍然高度可解码。同时,穷尽的留一法测试揭示了这种泛化的明显局限,并且沿着探针导出的有效性方向进行的干预与随机对照相比仅产生微弱的非特异性效应。我们的结果表明,表征有效性、在行为中表达有效性以及因果地使用有效性是彼此不同的。与有效性相关的信息可以从模型的隐藏状态中被强解码,却并不能在其输出中被可靠表达。
cs.CL / 33 / 2609.02473
Learning to Fuse LLMs with Ontology Rankers for Rare-Disease Diagnosis
学习将大语言模型与本体排序器融合用于罕见病诊断
large language model
大语言模型相关
Abstract
Ontology rankers remain useful for rare-disease diagnosis because each candidate can be traced to matched patient phenotypes. Large language models (LLMs) can generate differential diagnoses from the same patient description, but their predictions lack an equally clear evidence trail. Rather than asking which system should replace the other, we ask whether an LLM can improve the ranker without giving up its evidence. Our behavior-based fusion model examines the two ranked lists, their agreement, and the ontology support behind each candidate, and learns how much to rely on each system for the individual case. Before comparison, we remove a documented test-set leakage pathway caused by benchmark cases and ontology annotations being derived from the same publications. Across eight open LLMs, fusion improves Phenomizer Recall@1 by 7.86 percentage points on Phenopacket Store and 20.18 points on RAMEDIS. When paired with DeepSeek-V4-Flash through an API, a fusion model trained only on the other LLMs improves Recall@1 from 0.1657 to 0.2176, a 5.19-point gain, without retraining. For 90.8% of correct fused diagnoses, the disease retains candidate-level ontology evidence that can be inspected. These results show that LLMs can strengthen an established diagnostic tool without discarding the structured evidence that makes it useful.
Chinese Translation
本体排序器对罕见病诊断仍然有用,因为每个候选疾病都可以追溯到匹配的患者表型。大语言模型(LLMs)可以从同一患者描述生成鉴别诊断,但其预测缺乏同样清晰的证据链。我们不是追问哪种系统应取代另一种,而是追问大语言模型能否在不放弃其证据的情况下改进排序器。我们的基于行为的融合模型检查两个排序列表、它们的一致性以及每个候选疾病背后的本体支持,并学习在单个病例中应分别依赖每个系统的程度。在比较之前,我们移除了一条已记录的测试集泄漏路径,该路径是由基准病例和本体注释源自相同出版物造成的。在八个开放大语言模型上,融合将Phenomizer在Phenopacket Store上的Recall@1提高了7.86个百分点,在RAMEDIS上提高了20.18个百分点。当通过API与DeepSeek-V4-Flash配对时,仅在其他大语言模型上训练的融合模型无需重新训练即可将Recall@1从0.1657提高到0.2176,提升了5.19个百分点。对于90.8%的正确融合诊断,疾病保留了可检查的候选级本体证据。这些结果表明,大语言模型可以在不丢弃使其有用的结构化证据的情况下,增强既有的诊断工具。
cs.CL / 34 / 2609.02482
How LLMs Build Fictional Worlds: Setting and Narrative Space in AI-Generated Creative Storytelling
LLM如何构建虚构世界:AI生成创意叙事中的设定与叙事空间
large language model
大语言模型相关
Abstract
In this paper, we analyze how Large Language Models (LLMs) employ worldbuilding strategies, focusing on setting as one measurable dimension of storyworld construction. We compare 1,000 AI-generated stories per model in English and German with human-authored fiction from Project Gutenberg. Building on prior work, we operationalize setting through five types of narrative space: "action", "perceived," "visual," "descriptive" and "no space", identified using fine-tuned BERT classifiers for German and English. We generate narratives using GPT 4.1, LlaMA 3.3, Mistral 3.2, and Gemma 3 and compare their spatial distributions to a human-authored baseline. We find that human-authored texts predominantly employ "action space," grounding narratives in embodied character-environment interaction, whereas LLMs systematically overproduce "perceived space," emphasizing atmosphere and affect. This divergence remains stable across narrative time. Overall, our findings show that LLMs exhibit worldbuilding patterns that differ consistently from human-authored fiction in ways that are both model-specific and language-sensitive.
Chinese Translation
在本文中,我们分析了大型语言模型(LLM)如何运用世界构建策略,重点关注设定作为故事世界构建的一个可衡量维度。我们将每个模型分别用英语和德语生成的1,000个AI故事与来自Project Gutenberg的人类作者虚构作品进行比较。基于先前的工作,我们通过五种类型的叙事空间来操作化设定:“动作空间”“感知空间”“视觉空间”“描述空间”和“无空间”,并使用针对德语和英语微调的BERT分类器加以识别。我们使用GPT 4.1、LlaMA 3.3、Mistral 3.2和Gemma 3生成叙事,并将其空间分布与人类作者基线进行比较。我们发现,人类作者创作的文本主要采用“动作空间”,将叙事植根于具身化的角色—环境互动之中,而LLM则系统地过度生成“感知空间”,强调氛围与情感。这种差异在叙事时间内保持稳定。总体而言,我们的研究结果表明,LLM所表现出的世界构建模式与人类作者虚构作品存在持续且一致的差异,这些差异既具有模型特异性,又对语言敏感。
cs.CL / 35 / 2609.02496
Debias-SparseGPT: Bias-Aware Pruning for Large Language Models
Debias-SparseGPT:面向大型语言模型的偏差感知剪枝
large language model
大语言模型相关
Abstract
Model compression techniques such as pruning and quantization facilitate the efficient deployment and acceleration of Large Language Models (LLMs). However, recent studies show that weight sparsification methods, such as SparseGPT, can amplify existing biases in models, with outputs varying significantly depending on persona cues in the prompt. In this paper, we introduce Debias-SparseGPT, a post-training pruning method incorporating representational debiasing using a second-order term defined over demographically contrasting inputs. We perform empirical validation of our method over a wide range of generative LLMs. Across models and sparsity regimes (25%, 50%, and structured 2:4 sparsity), Debias-SparseGPT consistently reduces pruning-induced bias compared to SparseGPT while preserving model perplexity and zero-shot accuracy. Under the most restrictive 2:4 structured sparsity pattern, which most aggressively degrades model quality, augmenting the calibration set with long-context, content-rich examples further improves both downstream performance and fairness. Overall, Debias-SparseGPT advances the bias-performance trade-off while preserving the computational efficiency of sparse models.
Chinese Translation
剪枝和量化等模型压缩技术促进了大型语言模型(LLMs)的高效部署与加速。然而,最近的研究表明,SparseGPT等权重稀疏化方法会放大模型中已有的偏差,其输出会随提示中的人物线索而发生显著变化。在本文中,我们提出了Debias-SparseGPT,一种后训练剪枝方法,它使用在人口统计学上形成对比的输入上定义的二阶项,融合了表征去偏。我们在多种生成式大型语言模型上对我们的方法进行了实证验证。在不同模型和稀疏度设置(25%、50%以及结构化2:4稀疏度)下,与SparseGPT相比,Debias-SparseGPT在保持模型困惑度和零样本准确率的同时,持续降低了剪枝引起的偏差。在最具限制性的2:4结构化稀疏模式下(该模式对模型质量的损害最为严重),使用长上下文、内容丰富的示例扩充校准集,可进一步提高下游性能和公平性。总体而言,Debias-SparseGPT在保持稀疏模型计算效率的同时,推动了偏差-性能权衡的改善。
cs.CL / 36 / 2609.02526
When Persona Attributes Improve Population Alignment in Large Language Models
角色属性何时能改善大型语言模型中的总体对齐
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) are increasingly used to predict the responses of human participants in survey panels. Towards that goal, persona prompting has recently emerged as a technique to inform and align large pretrained language models. Persona prompting refers to the practice of using short textual descriptions of 'personas' in prompts to steer the LLM's generations. Personas describe individuals through different attributes such as their socio-demographics, attitudes, or behaviors, with the aim of aligning LLMs to produce responses that correlate with the corresponding human responses. Yet, recent work has produced mixed and partly conflicting results of persona prompting without clear patterns of success and failure. Among the few consistent findings is that the selection of persona attributes matters, and that using more attributes does not necessarily lead to better performance. It remains unclear how different attribute selection methods perform and how to choose among them. In this paper, we propose that observed human response variation of a survey question is a potential explanation for the mixed performance observed so far. In addition, we compare the performance of persona prompting associated with different methods for selecting persona attributes. We evaluate these methods on four different (general) social surveys across two countries, six LLMs, and twenty prediction tasks per survey. Our work helps to identify when persona prompting can be expected to be useful in survey prediction tasks, and provides new insights on the effectiveness of different attribute selection methods for LLM-based survey prediction using persona prompting.
Chinese Translation
大型语言模型(LLMs)越来越多地被用于预测调查面板中人类参与者的回答。为实现这一目标,角色提示最近作为一种为大型预训练语言模型提供信息并使其对齐的技术而出现。角色提示指的是在提示中使用关于“角色”的简短文本描述来引导LLM生成内容的做法。角色通过不同属性(如社会人口特征、态度或行为)来描述个体,其目的是使LLM对齐,从而生成与相应人类回答相关的回答。然而,近期工作产生了角色提示的混合且部分相互矛盾的结果,没有清晰的成败模式。在少数一致的发现中,角色属性的选择很重要,而使用更多属性并不一定带来更好的性能。目前仍不清楚不同的属性选择方法表现如何,以及如何在它们之间进行选择。在本文中,我们提出,对一项调查问题观察到的人类回答变异是迄今观察到的混合表现的一个潜在解释。此外,我们比较了与不同角色属性选择方法相关的角色提示的性能。我们在两个国家的四项不同的(一般性)社会调查、六个LLM以及每项调查二十个预测任务上评估了这些方法。我们的工作有助于识别在调查预测任务中预期角色提示何时会有用,并为使用角色提示的基于LLM的调查预测中不同属性选择方法的有效性提供了新见解。
cs.CL / 37 / 2609.02730
CORAL: An LLM-Native Harness for Production Recommender Systems
CORAL:面向生产级推荐系统的LLM原生操控框架
large language model
大语言模型相关
Abstract
Production recommender systems shape what billions of people see, and sustaining their performance requires continual optimization: as content, user behavior, and upstream models shift, the choices governing retrieval, ranking, and serving must be revisited. Traditionally, human engineers test such changes through online experiments--a slow, reactive process limited by engineering effort, leaving parts of the system unrevised as conditions change. Although large language models have been applied to ranking, user modeling, and offline model development, few systems place an agent in a continual closed loop that acts on a live recommender and learns from the measured effects of its decisions. We present CORAL (Constraint-Optimized Recommender via an Agentic Loop), an LLM-native harness that closes this loop: each cycle, the agent observes operating signals, reasons over a memory of past decisions and outcomes, and invokes tools--including a numerical optimizer that keeps changes within a fixed operating budget--to reconfigure the recommender, with measured outcomes informing the next cycle. We formulate this as a partially observed, non-stationary, constrained optimization problem in which the policy improves in context, without parameter updates, from its prior actions. Across two large-scale social platforms, evaluated with A/B experiments, the same harness improves engagement at no additional serving cost on one and reduces serving cost without degrading engagement on the other, spanning the engagement-efficiency frontier. Performance improves as the loop iterates, suggesting that a single agentic loop can automate continual optimization work traditionally performed by human algorithm engineers under explicit guardrails.
Chinese Translation
生产级推荐系统塑造着数十亿人所看到的内容,维持其性能需要持续的优化:随着内容、用户行为和上游模型发生变化,支配检索、排序和投放的选择必须被重新审视。传统上,人类工程师通过在线实验来测试此类变更——这是一个缓慢、被动且受工程人力限制的过程,导致当条件变化时系统的某些部分得不到修订。尽管大语言模型已被应用于排序、用户建模和离线模型开发,但很少有系统能够让一个智能体处于持续的闭环中,对实时推荐系统采取行动,并从其决策的可测量效果中学习。我们提出了CORAL(通过智能体循环实现约束优化的推荐系统),这是一个LLM原生的操控框架,用于闭合这一循环:每个周期中,智能体观察运行信号,基于过去决策与结果的内存进行推理,并调用工具——包括一个将变更限制在固定运行预算内的数值优化器——来重新配置推荐系统,而测量的结果则为下一周期提供信息。我们将其形式化为一个部分可观测、非平稳、带约束的优化问题,其中策略在上下文中、不进行参数更新的情况下,从其先前的行动中改进。在两个大规模社交平台上,通过A/B实验进行评估,同一个操控框架在一个平台上以零额外投放成本提升了参与度,在另一个平台上则在未降低参与度的同时降低了投放成本,跨越了参与度-效率的前沿。随着循环的迭代,性能持续提升,这表明单一的智能体循环能够自动化那些传统上由人类算法工程师在明确护栏下执行的持续优化工作。
cs.CL / 38 / 2609.02754
Untangling the Mechanisms of Misleading Context in Medical Question Answering
解开误导性上下文在医学问答中的机制
large language model
大语言模型相关
Abstract
Large language models now answer medical questions with expert-level performance. However, the context these systems act on can be misleading, and misleading context can corrupt a model's medical judgment. To understand how misleading context corrupts this judgment, we examine the model's susceptibility to the context, disclosure of it, mechanism of corrupted reasoning, and monitorability of the decision. On the medical reasoning subset of MedMisBench, a clinician-reviewed question-answering benchmark of 8,627 questions, we inject two types of misleading context cues, fabricated evidence and a bare assertion. We test three reasoning models, two that expose their full reasoning trace and one frontier model that exposes only its response. All three are more susceptible to the assertion than to the fabricated evidence, adopting the asserted answer 10 to 27 points more often. The misleading cues are disclosed in 81 to 98% of traces but only 7 to 90% of responses, and the assertion is disclosed less often than evidence based cues. Resampling from reasoning traces without disclosure shows the two cues corrupt reasoning differently, evidence entering early and accumulating while the assertion redirects the conclusion near its end. An LLM monitor catches 78% of corrupted decisions at 5% false positives when reading an open model's trace with guidance, against at most 32% from any response. The misleading context that models are most susceptible to is disclosed least, and was caught reliably only from an open reasoning trace, which frontier providers withhold.
Chinese Translation
大型语言模型现在以专家级水平回答医学问题。然而,这些系统所依据的上下文可能具有误导性,而误导性上下文可能损害模型的医学判断。为了理解误导性上下文如何损害这种判断,我们考察了模型对上下文的敏感性、上下文的披露情况、推理受损的机制以及决策的可监控性。在MedMisBench的医学推理子集(一个由临床医生审查的包含8,627个问题的问答基准)上,我们注入了两类误导性上下文线索:捏造的证据和裸断言。我们测试了三个推理模型,其中两个暴露其完整的推理轨迹,另一个是仅暴露其响应的前沿模型。三个模型对断言的敏感性均高于对捏造证据的敏感性,采纳断言答案的频率高出10到27个百分点。误导性线索在81%到98%的推理轨迹中被披露,但在响应中仅被披露7%到90%,且断言相比基于证据的线索被披露的频率更低。从无披露的推理轨迹中重新采样表明,这两类线索对推理的损害方式不同:证据较早进入并不断累积,而断言则在其接近结尾处重新引导结论。一个LLM监控器在读取带有指导的开放模型的推理轨迹时,在5%的假阳性率下捕获了78%的受损决策,而仅从任何响应中捕获最多32%。模型最易受影响的误导性上下文被披露得最少,并且只有从开放的推理轨迹中才能可靠地被捕获,而前沿提供者不公开这种轨迹。
cs.CL / 39 / 2609.02796
DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation
DiscoSign:语篇感知的文本到手语注释翻译
large language model
大语言模型相关
Abstract
Sign language processing systems have traditionally operated at the sentence level, ignoring critical discourse phenomena fundamental to sign language comprehension. We introduce DiscoSign, a computational approach for discourse-aware text to sign language gloss translation grounded in linguistic research. We address three key phenomena within our modular Large Language Model (LLM)-based translation framework: (i) spatial coreference resolution, where entities maintain consistent spatial locations throughout discourse; (ii) Question-Answer Clauses (QACs), pseudocleft structures serving specific discourse functions; and (iii) concept-gloss consistency, ensuring stable mappings between English concepts and American Sign Language (ASL) signs. Traditional translation metrics fail to capture discourse-level quality, so we introduce a suite of novel evaluation metrics designed to assess each dimension of discourse coherence addressed by our framework. Experiments on sentence-level and discourse-level datasets show that our approach for discourse-aware processing significantly improves spatial consistency and entity tracking relative to sentence-only translation, while maintaining competitive single-sentence gloss translation quality. Our work establishes the first systematic framework for discourse-level text to sign language gloss translation with corresponding evaluation methodology.
Chinese Translation
手语处理系统传统上在句子层面运行,忽视了对于手语理解至关重要的关键语篇现象。我们提出了 DiscoSign,一种基于语言学研究的语篇感知文本到手语注释翻译的计算方法。我们在基于模块化大语言模型(LLM)的翻译框架中处理三个关键现象:(i) 空间共指消解,即实体在整个语篇中保持一致的空间位置;(ii) 问答子句(Question-Answer Clauses,QACs),即承担特定语篇功能的伪分裂结构;(iii) 概念-注释一致性,确保英语概念与美国手语(ASL)手势之间的稳定映射。传统的翻译指标无法捕捉语篇层面的质量,因此我们引入了一套新型评估指标,旨在评估我们框架所处理的语篇连贯性的每个维度。在句子层面和语篇层面的数据集上进行的实验表明,与仅进行句子翻译相比,我们的语篇感知处理方法显著提升了空间一致性和实体追踪能力,同时保持了有竞争力的单句注释翻译质量。我们的工作建立了第一个用于语篇级文本到手语注释翻译并配有相应评估方法的系统框架。
cs.CL / 40 / 2609.02859
User Feedback Provides a Unique Signal that LLMs Can not Detect
用户反馈提供了一种LLMs无法检测的独特信号
large language model
大语言模型相关
Abstract
Harnessing naturally occurring feedback from user interactions offers a promising learning signal for Large Language Models (LLMs). However, recent studies suggest this feedback is inherently noisy and difficult to leverage effectively. We challenge this conception by demonstrating that user feedback is a highly actionable signal for improvement, and that its perceived ineffectiveness stems from a systematic bias in current evaluation paradigms. To isolate the usefulness of feedback, we construct synthetic data with a definitive ground truth, alongside naturalistic data to validate that our findings hold in real-world scenarios. By comparing model revisions generated with and without access to feedback across both settings, we show that feedback-informed revisions resolve targeted issues at significantly higher rates than baseline revisions. Finally, we expose the root of the evaluation bias: when a model successfully fixes an issue exclusively due to feedback, LLM judges frequently fail to identify the genuinely corrected response, systematically preferring inferior baseline outputs instead.
Chinese Translation
利用用户交互中自然产生的反馈,可为大语言模型(LLMs)提供一种很有前景的学习信号。然而,最近的研究表明,这种反馈本质上有噪声且难以有效利用。我们通过表明用户反馈是一种高度可操作的改进信号,且其表面上的无效性源于当前评估范式中的系统性偏差,来挑战这一观点。为了分离出反馈的作用,我们构建了具有明确真值的合成数据,并同时使用自然数据,以验证我们的发现在真实场景中依然成立。通过比较在两种设置下,有反馈与无反馈条件下生成的模型修订,我们发现基于反馈的修订解决目标问题的比率显著高于基线修订。最后,我们揭示了评估偏差的根源:当模型仅仅因为反馈而成功修复一个问题时,LLM评判器经常无法识别出真正被纠正的回答,而是系统性地偏好较差的基线输出。
cs.CR / 41 / 2609.01758
Towards Behavior Tree-Guided Vulnerability Detection with Lightweight LLMs
面向轻量级大语言模型的基于行为树引导的漏洞检测
large language model
大语言模型相关
Abstract
Large Language Models (LLMs) are increasingly used for software vulnerability detection, but their performance depends on how source code is represented in the input. Most prompting approaches use source code in its original form, while some works propose the use of structured representations. Abstract Syntax Trees (ASTs) are one of the most popular approaches, but AST verbosity increases input size relative to source code, making them hard to fit within some LLMs context windows. This paper investigates Behavior Trees (BTs) as an alternative intermediate representation for LLM-based vulnerability detection. BTs encode control flow, conditions, and executable actions more compactly than ASTs, making them a natural candidate when token count is a constraint. First, we propose a preprocessing stage that parses Java source code into ASTs and then converts them into BT representations. We then compare vulnerability detection performance across 460 Java samples from the Juliet Java test suite, using three input representations: raw source code, AST, and BT. All experiments use a single quantized local LLM, Mistral Small 3.2 24B (Q4_K_M). Our results show that using BT representations improves recall on short code samples, while raw source code achieves higher precision. On longer samples, BTs improve overall performance over the original representation and fit within the context window, whereas many ASTs exceed the context limit. These findings suggest that BTs can provide a compact and useful structured representation for vulnerability detection with quantized, locally deployable LLMs.
Chinese Translation
大型语言模型(LLM)越来越多地被用于软件漏洞检测,但其性能取决于源代码在输入中是如何表示的。大多数提示方法使用其原始形式的源代码,而一些工作则提出使用结构化表示。抽象语法树(AST)是最流行的方法之一,但相较于源代码,AST 的冗长性会增加输入规模,使其难以适配某些 LLM 的上下文窗口。本文研究行为树(BT)作为基于 LLM 的漏洞检测的一种替代中间表示。行为树比 AST 更紧凑地编码控制流、条件和可执行动作,因此在令牌数量受限时,它们是一个自然的候选方案。首先,我们提出了一个预处理阶段,将 Java 源代码解析为 AST,然后将其转换为行为树表示。随后,我们使用三种输入表示:原始源代码、AST 和行为树,对来自 Juliet Java 测试套件的 460 个 Java 样本进行漏洞检测性能比较。所有实验均使用单个量化本地 LLM,即 Mistral Small 3.2 24B(Q4_K_M)。我们的结果表明,使用行为树表示能提高短代码样本的召回率,而原始源代码则能达到更高的精确率。在较长的样本上,行为树相比原始表示能提升整体性能,并且能适应上下文窗口,而许多 AST 则超出上下文限制。这些发现表明,在基于量化的、可本地部署的 LLM 进行漏洞检测时,行为树可以提供一种紧凑且有用的结构化表示。
cs.CR / 42 / 2609.02177
WeaveMark: Robust and Scalable Multi-bit LLM Watermarking via Coded Payload Spreading
WeaveMark:基于编码载荷扩展的鲁棒且可扩展的多比特大语言模型水印
large language model
大语言模型相关
Abstract
Multi-bit watermarking for large language models (LLMs) enables content source tracing by embedding user-identifiable messages into generated text. Existing methods face a fundamental trade-off among extraction accuracy, text quality, and payload capacity. We propose WeaveMark, a robust and scalable multi-bit LLM watermarking scheme based on coded payload spreading. WeaveMark shifts this trade-off frontier by improving payload capacity through multi-bit-per-token spreading, improving extraction accuracy through soft-decision error-correcting code, and preserving text quality through unbiased multilayer reweighting. It further introduces dedicated zero-bit layers for reliable watermark presence detection. Experiments show large gains, especially for long messages and edited text. WeaveMark achieves 89.8% match rate for 32-bit messages at 200 tokens, compared with 20.8% for BiMark. Under 10% substitution attacks on 16-bit messages at 200 tokens, it maintains 86.0% versus 30.7%, while preserving text quality. Our code is available at https://github.com/qkrrkd90-source/WeaveMark.
Chinese Translation
面向大语言模型(LLM)的多比特水印通过将可标识用户的消息嵌入生成的文本中,实现内容来源追踪。现有方法在提取准确率、文本质量和载荷容量之间面临根本性权衡。我们提出 WeaveMark,一种基于编码载荷扩展的鲁棒且可扩展的多比特 LLM 水印方案。WeaveMark 通过多比特/令牌扩展提升载荷容量,通过软判决纠错码提升提取准确率,并通过无偏多层重加权保持文本质量,从而推动了这一权衡边界。它进一步引入专用的零比特层,以实现可靠的水印存在性检测。实验表明,特别是在长消息和被编辑文本上,取得了显著提升。对于 32 比特消息,在 200 个令牌长度下,WeaveMark 实现了 89.8% 的匹配率,而 BiMark 仅为 20.8%。在 16 比特消息、200 个令牌长度下,面对 10% 的替换攻击,它保持了 86.0% 的匹配率,而 BiMark 为 30.7%,同时保持了文本质量。我们的代码可在 https://github.com/qkrrkd90-source/WeaveMark 获取。
cs.CR / 43 / 2609.02465
Can Risk-Based Alerting Mitigate Cybersecurity Alert Fatigue?
基于风险的告警能否缓解网络安全告警疲劳?
large language model
大语言模型相关
Abstract
Security operations centers (SOCs) face large numbers of false alerts, making detection of cyberattacks difficult under typical resource constraints. Risk-based alerting (RBA) has been proposed as a means to reduce false alerts and has reportedly succeeded in doing so in various enterprise deployments. However, RBA has not been comprehensively evaluated until now, leaving implementation mostly guesswork based on anecdotal evidence. In this paper, we present the first systematic evaluation of RBA. To this end, we reformulate it as a continuous alert prioritization problem rather than a binary decision problem (i.e., whether an alerting threshold is exceeded), allowing us to evaluate performance across all possible thresholds and thus model SOCs of varying sizes and alert volumes. We distill five fundamental risk hypotheses, formalize them as independently parametrizable modules, and implement them in our novel experimentation suite CATS. We thoroughly assess the hypotheses across eight diverse alert datasets, six of which we created or extended to make such an evaluation possible. Our results show that certain combinations of hypotheses achieve a remarkable alert prioritization performance (AUROC $μ=0.92$, $σ=0.09$ across the eight datasets), outperforming a straightforward prioritization by alert severity level (AUROC $μ=0.72$, $σ=0.21$). We conclude that RBA can substantially reduce the number of false alerts that analysts have to review and thus has the potential to mitigate cybersecurity alert fatigue. In addition, it serves as a strong baseline for more complex, resource-intensive alert triage approaches (e.g., based on large language models).
Chinese Translation
安全运营中心(SOC)面临大量误报,这使得在典型的资源约束下检测网络攻击变得困难。基于风险的告警(RBA)已被提出作为一种减少误报的手段,并且据报道在各种企业部署中已成功实现了这一目标。然而,RBA 直到现在才被全面评估,这使得其实施在很大程度上仍基于轶事证据的猜测。在本文中,我们提出了对 RBA 的首次系统性评估。为此,我们将其重新表述为一个连续的告警优先级排序问题,而非二元决策问题(即是否超过告警阈值),从而使我们能够评估所有可能阈值下的性能,并以此对不同规模和告警量的 SOC 进行建模。我们提炼出五个基本的风险假设,将其形式化为可独立参数化的模块,并在我们新颖的实验套件 CATS 中加以实现。我们跨八个不同的告警数据集对这些假设进行了全面评估,其中六个是我们为使得此类评估成为可能而创建或扩展的。我们的结果表明,某些假设组合实现了显著的告警优先级排序性能(八个数据集上的 AUROC $\mu=0.92$,$\sigma=0.09$),优于按告警严重级别直接排序(AUROC $\mu=0.72$,$\sigma=0.21$)。我们得出结论:RBA 能大幅减少分析人员必须审查的误报数量,因而具有缓解网络安全告警疲劳的潜力。此外,它为更复杂、资源密集型的告警分流方法(例如基于大型语言模型的方法)提供了一个强大的基线。
cs.CR / 44 / 2609.02553
The Shape of Ownership: Verifying LLM Provenance through Semantic Structures
所有权的形态:通过语义结构验证大语言模型的出处
large language model
大语言模型相关
Abstract
As large language models (LLMs) are increasingly redistributed, adapted, and served behind opaque APIs, model ownership can no longer be established reliably by inspecting model internals or deployment records. This creates a need for behavioral signatures that remain observable through black-box interaction. Yet most existing black-box fingerprints instantiate ownership signals through fixed query-key associations, reducing model identity to sparse memorized associations detached from ordinary behavior and limiting both robustness and stealth (e.g., fine-tuning or quantization) and stealthiness. A stronger fingerprint should instead be distributed, naturally elicited, and expressed at a higher semantic level. To this end, we introduce PROSE (Provenance through Relational Organization of Semantic Expression), replacing fixed query sets with a target semantical domain and brittle response keys with semantic structures internalized as domain-conditioned response behavior. Specifically, the fingerprint is encoded in how the model semantically organizes its in-domain conclusions, rather than in particular tokens or prescribed outputs. PROSE constructs a private bank of domain-specific semantic templates, internalizes them through mixed fine-tuning on structurally verified and clean responses, and verifies ownership by detecting the designated structures in responses to held-out natural queries. Extensive experiments across multiple model architectures, scales, and target domains show that PROSE achieves a 100% fingerprint detection rate on unmodified models with no observed false positives, preserves model utility, and retains strong detectability under downstream modifications and output transformations.
Chinese Translation
随着大语言模型(LLMs)越来越多地被再分发、适配并在不透明的API背后提供服务,模型所有权已无法再通过检查模型内部结构或部署记录来可靠地确立。这产生了对行为签名的需求,这些签名在通过黑盒交互时仍然可观察。然而,大多数现有的黑盒指纹通过固定的查询-键关联来实例化所有权信号,将模型身份简化为与普通行为相脱离的稀疏记忆关联,从而限制了(对例如微调或量化等修改的)鲁棒性和隐蔽性。更强的指纹应当转而采用分布式、自然触发的方式,并在更高的语义层面上表达。为此,我们提出了PROSE(通过语义表达的关系组织来实现溯源),用目标语义域取代固定的查询集,并用作为域条件响应行为内化的语义结构替换脆弱的响应键。具体而言,指纹编码在模型如何以语义方式组织其领域内结论之中,而不是编码在特定的词元或指定的输出中。PROSE构建一个私有的领域特定语义模板库,通过在结构上经过验证的干净响应上进行混合微调将其内化,并通过检测对保留的(held-out)自然查询的响应中是否包含指定结构来验证所有权。跨多种模型架构、规模和目标领域的广泛实验表明,PROSE在未修改模型上实现了100%的指纹检测率,且未观察到假阳性;它保持了模型效用,并在下游修改和输出变换下仍具有很强的可检测性。
cs.CR / 45 / 2609.02564
A Finger on the Scale: Covert Policy Steering through Agentic Skills
天平上的手指:通过智能体技能进行隐蔽策略引导
large language model
大语言模型相关
Abstract
Reusable agent skills extend large language model (LLM) agents with task procedures, tool-use guidance, and output constraints. Yet these skills also act as externalized behavioral policies, which create a supply-chain risk: a third-party skill may preserve the declared task and valid output interface while covertly redirecting agent decisions toward an undisclosed objective. We formalize Skill Policy Integrity, which requires a Skill-induced policy to remain aligned with its declared functionality and the user-authorized objective. We further present SkillShift, a constrained black-box framework for covert policy steering without explicit target command injection or task hijacking. It combines semantically plausible policy edits with hierarchical validation, failure-guided optimization, and strategy compression to preserve effectiveness, output validity, transferability, and inconspicuousness. We instantiate this threat in agentic commerce and software dependency use, with SkillShift achieving attacker-favored selection rates of 81.33% and 63.33% while maintaining a 100% utility-preserving rate. The frozen policies also transfer without further optimization across heterogeneous LLM backends and agent environments. Moreover, the evaluated scanners fail to detect the constructed skills, motivating behavioral auditing of reusable skills as agent policy artifacts.
Chinese Translation
可复用的智能体技能以任务流程、工具使用指南和输出约束来扩充大型语言模型(LLM)智能体。然而,这些技能也充当外部化的行为策略,从而产生供应链风险:第三方技能可能保留所声明的任务和有效的输出接口,同时隐蔽地将智能体决策引向未公开的目标。我们形式化了技能策略完整性(Skill Policy Integrity),它要求由技能诱导的策略与其声明的功能及用户授权的目标保持一致。我们进一步提出了SkillShift,这是一个受约束的黑盒框架,用于在没有显式目标命令注入或任务劫持的情况下进行隐蔽策略引导。它结合了语义合理的策略编辑与分层验证、失败引导优化和策略压缩,以保持有效性、输出合法性、可迁移性和隐蔽性。我们在智能体商务和软件依赖使用场景中实例化了这种威胁,SkillShift在保持100%效用保持率的同时,实现了攻击者偏好的选择率81.33%和63.33%。这些冻结策略还无需进一步优化即可在异构LLM后端和智能体环境之间迁移。此外,所评估的扫描器未能检测到所构造的技能,这促使将可复用技能作为智能体策略产物进行行为审计。
cs.CR / 46 / 2609.02690
ACLE-MCP: Attested Capability Leases for Execution-Time Trust in Remote LLM Tool Use
ACLE-MCP:用于远程LLM工具使用中执行时信任的经证明能力租约
large language model
大语言模型相关
Abstract
Remote Model Context Protocol (MCP) services enable large language model agents to invoke external tools, but OAuth authorization alone does not ensure that a later tool call is executed by the provider-side workload that the relying party intended to trust. An endpoint may remain authorized even after execution shifts to a substituted workload, relies on stale appraisal state, reuses authority transferred from another sender, or traverses an undeclared downstream component. We call this problem the post-authorization execution trust gap. We present ACLE-MCP, an invocation-scoped architecture that couples delegated authorization, workload appraisal, and resource-side execution admission. For protected calls, ACLE-MCP issues a short-lived, sender-constrained capability lease that binds the expected workload, freshness requirement, operation, object and parameter bounds, downstream constraints, and receipt obligations. A provider-side Execution Gate consumes the lease immediately before protected tool logic begins. We implement a runnable prototype with Keycloak/OIDC validation, an MCP Python SDK server, and an optional vTPM quote-verification backend. Controlled security experiments and an agent tool-use extension show that weaker authorization or connect-time attestation modes leave distinct post-authorization attacks open, whereas full ACLE-MCP blocks all evaluated attack families while preserving all benign tasks. In the locally simulated agent extension, the complete design increases request-level pooled p95 latency on normal allowed calls by 25.7% relative to OAuth-only. These results indicate that invocation-time binding between call authority and current workload state is a practical complement to OAuth-protected remote tool use.
Chinese Translation
远程模型上下文协议(MCP)服务使大型语言模型代理能够调用外部工具,但仅OAuth授权本身并不能确保后续的工具调用是由依赖方原本打算信任的提供方工作负载执行的。即使执行转移到被替换的工作负载、依赖过时的评估状态、重用从另一发送方转移的权限,或穿越未声明的下游组件,端点也可能仍然保持授权。我们将此问题称为授权后执行信任缺口。我们提出ACLE-MCP,这是一种以调用为范围的架构,它耦合了委托授权、工作负载评估和资源端执行准入。对于受保护的调用,ACLE-MCP签发一个短期的、发送方受限的能力租约,该租约绑定预期工作负载、新鲜度要求、操作、对象和参数边界、下游约束以及回执义务。提供方侧的执行门在受保护的工具逻辑开始之前立即消费该租约。我们实现了一个可运行的原型,包括Keycloak/OIDC验证、一个MCP Python SDK服务器,以及一个可选的vTPM引用验证后端。受控安全实验和一个代理工具使用扩展表明,较弱的授权或连接时证明模式会留下明显的授权后攻击漏洞,而完整的ACLE-MCP可以阻止所有被评估的攻击家族,同时保留所有良性任务。在本地模拟的代理扩展中,与仅使用OAuth相比,完整设计使正常允许调用的请求级聚合p95延迟增加了25.7%。这些结果表明,在调用时将调用权限与当前工作负载状态绑定是对受OAuth保护的远程工具使用的一种实用补充。
cs.AI / 47 / 2609.01749
Swin Meets EfficientNet: Lightweight Architectures for GAN-Based Face Forensics
Swin 遇见 EfficientNet:面向基于 GAN 的人脸取证的轻量级架构
diffusion
扩散模型相关
Abstract
Modern generative models, such as GANs, diffusion architectures, and autoregressive systems, now produce facial images that are nearly indistinguishable from authentic photographs. This capability makes detecting forged images increasingly difficult, raising serious concerns about identity theft, fraud, and misinformation campaigns. Our research focuses specifically on GAN-generated synthetic faces, which underpin many face-centric deepfakes, and investigates efficient detection approaches using image analysis alone. Existing detection systems rely heavily on either convolutional neural networks (CNNs) or global vision transformers. While CNNs excel at identifying texture-based local features, they struggle with broader contextual understanding. Traditional Vision Transformer (ViT) models can capture long-range structures effectively, but demand substantial computational resources. Our work explores Swin-Transformer-based architectures across three implementations: a compact Swin Transformer trained from the ground up, ImageNet-1K pre-trained Swin-Tiny and Swin-Small models adapted for binary classification, and a novel hybrid combining EfficientNet-B0's convolutional processing with a Swin Transformer backend. We evaluated all models using the 140K Real and Fake Faces dataset, which includes StyleGAN-generated fake faces alongside authentic images from Flickr and DFDC, with balanced splits for training, validation, and testing. The EfficientNetB0+Swin hybrid achieved 99% accuracy and a 99.44% recall on 5,000 test images, outperforming both pure Swin variants and a previous CNN-only baseline on this dataset. Our results suggest that combining hierarchical CNN features with shifted-window self-attention provides an efficient and computationally lightweight method for detecting GAN-generated synthetic faces.
Chinese Translation
现代生成模型,例如 GAN、扩散架构和自回归系统,如今能够生成与真实照片几乎无法区分的人脸图像。这种能力使得检测伪造图像变得越来越困难,引发了对身份盗窃、欺诈和虚假信息活动的严重担忧。我们的研究专门聚焦于 GAN 生成的合成人脸——这些合成人脸支撑着许多以人脸为中心的深度伪造——并研究仅使用图像分析的高效检测方法。现有检测系统很大程度上依赖于卷积神经网络(CNN)或全局视觉变换器。虽然 CNN 善于识别基于纹理的局部特征,但它们难以进行更广泛的上下文理解。传统的视觉变换器(ViT)模型能够有效捕获长距离结构,但需要大量计算资源。我们的工作探索了基于 Swin Transformer 的三种实现架构:从头训练的紧凑型 Swin Transformer;针对二分类进行微调的 ImageNet-1K 预训练 Swin-Tiny 和 Swin-Small 模型;以及一种结合 EfficientNet-B0 的卷积处理与 Swin Transformer 后端的新型混合架构。我们使用 140K Real and Fake Faces 数据集评估了所有模型,该数据集包含 StyleGAN 生成的假人脸以及来自 Flickr 和 DFDC 的真实图像,并采用平衡划分用于训练、验证和测试。EfficientNetB0+Swin 混合模型在 5,000 张测试图像上取得了 99% 的准确率和 99.44% 的召回率,优于该数据集上的纯 Swin 变体和先前仅使用 CNN 的基线模型。我们的结果表明,将分层 CNN 特征与移位窗口自注意力相结合,为检测 GAN 生成的合成人脸提供了一种高效且计算轻量的方法。
cs.LG / 48 / 2609.01997
Linear Fusion MultiDiffusion for Fast Training-Free Spherical Panorama Generation
线性融合多扩散:用于快速免训练球形全景生成的线性融合多扩散方法
diffusion
扩散模型相关
Abstract
We propose LF-MultiDiffusion, a training-free panorama generation method that extends MultiDiffusion to support linear projections between target and reference image spaces. Our key idea is to reformulate latent aggregation as a regularized least-squares problem and solve it efficiently with a Krylov-based iterative solver inside the denoising loop. This formulation enables denser and more natural mappings than prior training-free methods, yielding more stable generation with far fewer perspective views. As a result, LF-MultiDiffusion reduces the number of image generator evaluations during denoising and significantly improves inference efficiency. Experiments show that LF-MultiDiffusion achieves better visual quality, text alignment, and panoramic consistency than the strongest training-free baseline, while providing a 15.36$\times$ speedup. Our project page is available at: https://ahykw.github.io/lfmd.
Chinese Translation
我们提出了 LF-MultiDiffusion,一种免训练的全景生成方法,它将 MultiDiffusion 扩展为支持目标图像空间与参考图像空间之间的线性投影。我们的关键思想是将潜在空间中的聚合重新构造为一个正则化最小二乘问题,并在去噪循环内使用基于 Krylov 的迭代求解器高效地求解该问题。这种公式化方式比先前的免训练方法支持更密集、更自然的映射,从而在需要远少于透视视图的情况下生成更稳定的结果。因此,LF-MultiDiffusion 减少了去噪过程中图像生成器的评估次数,并显著提高了推理效率。实验表明,与最强的免训练基线相比,LF-MultiDiffusion 在视觉质量、文本对齐和全景一致性方面均取得了更好的表现,同时实现了 15.36$\times$ 的加速。我们的项目页面位于:https://ahykw.github.io/lfmd。
cs.AI / 49 / 2609.02004
InstEditSeg: Instruction-Driven Image Editing for Polyp and Skin Lesion Segmentation
InstEditSeg:指令驱动的图像编辑用于息肉和皮肤病变分割
diffusion
扩散模型相关
Abstract
Accurate segmentation of polyps and skin lesions is pivotal for clinical diagnosis, yet existing methods struggle with low contrast, ambiguous boundaries, and cross-domain distribution discrepancies. Discriminative networks and most diffusion-based segmentation approaches predict standalone binary masks, leaving the visual priors of large-scale pretrained generative models largely unexploited. We propose InstEditSeg, a unified generative framework that reformulates medical segmentation as an instruction-driven image editing problem. Instead of emitting a mask, the model renders a color-coded overlay on the original image, conditioned on a textual instruction, so that the edited output aligns with the natural image distribution learned by latent diffusion models and mitigates the domain gap between natural and medical imagery. To recover fine anatomical structures, we introduce DINOv3 as an auxiliary visual encoder and a DINO Feature Guidance Block that builds a multi-scale feature pyramid. The pyramid is fused into the diffusion U-Net by channel concatenation and zero-initialized convolution so that hierarchical discriminative priors can be injected without perturbing the pretrained weights. A dual-branch classifier-free guidance strategy requiring only two forward passes per denoising step reduces inference cost. On polyp and skin lesion benchmarks the framework achieves accuracy competitive with strong discriminative baselines, and it further demonstrates concrete advantages of the generative formulation: notably better cross-domain generalization on unseen data, more complete multi-lesion segmentation, instruction-conditioned task control, and sampling flexibility. We also analyze the strengths and limitations of the paradigm, including its color sensitivity and unsupported attribute-conditioned selection. Code is available at: https://github.com/wincharm001/InstEditSeg.
Chinese Translation
息肉和皮肤病变的准确分割对于临床诊断至关重要,然而现有方法在处理低对比度、模糊边界以及跨域分布差异方面存在困难。判别式网络和大多数基于扩散的分割方法预测独立的二值掩码,这使得大规模预训练生成模型的视觉先验在很大程度上未被利用。我们提出了 InstEditSeg,这是一个统一的生成式框架,将医学分割重新构建为指令驱动的图像编辑问题。该模型不是输出掩码,而是根据文本指令在原始图像上渲染一个颜色编码的覆盖层,从而使编辑后的输出与潜在扩散模型学习到的自然图像分布对齐,并缓解自然图像与医学图像之间的域差距。为了恢复精细的解剖结构,我们引入了 DINOv3 作为辅助视觉编码器,并设计了一个 DINO 特征引导块来构建多尺度特征金字塔。该金字塔通过通道拼接和零初始化卷积融合到扩散 U-Net 中,从而在不扰动预训练权重的情况下注入层次化的判别先验。一种双分支的无分类器引导策略,每个去噪步骤仅需两次前向传播,降低了推理成本。在息肉和皮肤病变基准上,该框架达到了与强判别式基线相当的准确性,并进一步展示了生成式公式的实际优势:在未见数据上显著更好的跨域泛化能力、更完整的多病灶分割、指令驱动的任务控制以及采样灵活性。我们还分析了该范式的优势与局限性,包括其颜色敏感性和不支持基于属性的条件选择。代码可在 https://github.com/wincharm001/InstEditSeg 获取。
cs.AI / 50 / 2609.02291
VoRTeC: Taming Foundation Flow for One-step Real time Video Compression
VoRTeC: 驯服基础流实现一步式实时视频压缩
diffusion
扩散模型相关
Abstract
Ultra-low bitrate video compression still faces critical challenges: traditional neural video compression inevitably introduces blurring artifacts, while diffusion-based generative video compression suffers from excessive decoding latency and poor temporal consistency. To address these issues, we propose $\mathtt{VoRTeC}$, a Video Compression framework built upon a foundational flow model (Wan2.1). By compactly encoding latent video representations, predicting the positions of compressed representations along flow trajectories, and integrating multi-scale priors, $\mathtt{VoRTeC}$ enables the compressor to harness generative video flow priors effectively. Without accessing the parameters or gradients of flow matching networks, our framework achieves one-step decoding and reconstructions with high perceptual fidelity. Meanwhile, we maintain consistency across frame groups via tail-frame reuse and prior caching. Extensive experiments demonstrate that our method reduces bit consumption by 58\% compared to prior diffusion-based approaches, with decoding speed boosted by 3 to 197 times: $\mathtt{VoRTeC}$ achieves a decoding speed of 13 FPS at 720p and 32 FPS at 480p.
Chinese Translation
超低码率视频压缩仍然面临关键挑战:传统神经视频压缩不可避免地引入模糊伪影,而基于扩散的生成式视频压缩则存在过高的解码延迟和较差的时间一致性。为了解决这些问题,我们提出了 $\mathtt{VoRTeC}$,这是一个构建于基础流模型(Wan2.1)之上的视频压缩框架。通过紧凑地编码潜在视频表示、预测压缩表示沿流轨迹的位置,并整合多尺度先验,$\mathtt{VoRTeC}$ 使压缩器能够有效利用生成式视频流先验。在不访问流匹配网络参数或梯度的情况下,我们的框架实现了一步解码和具有高感知保真度的重建。同时,我们通过尾帧重用和先验缓存来维持帧组之间的一致性。大量实验表明,与先前的基于扩散的方法相比,我们的方法将比特消耗降低了58%,解码速度提升了3到197倍:$\mathtt{VoRTeC}$ 在720p下达到13 FPS的解码速度,在480p下达到32 FPS。
cs.CL / 51 / 2609.02780
ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding
ShallowStream:面向流式视频理解的浅层索引与深层作答
large language model
大语言模型相关
Abstract
Streaming video understanding is a critical capability for real-world applications, including embodied intelligence, autonomous driving, industrial monitoring, surveillance and early warning, and wearable assistants. However, processing continuous video streams with multimodal large language models (MLLMs) is computationally expensive. Existing efforts have explored reducing streaming overhead through visual token pruning, token merging, quantization, on-demand frame retrieval, and context offloading. However, most existing methods overlook the dimension of model depth. Repeatedly executing full-depth MLLM prefill over incoming frames is prohibitively expensive, incurring substantial computational overhead and causing the KV cache to grow at a rate directly proportional to the prefill depth. To address these challenges, we propose ShallowStream, a novel framework that leverages the shallow layers of an MLLM to simultaneously perform frame encoding and retrieval index building. During stream processing, ShallowStream maintains an always-on lightweight index using the KV cache of shallow layers. During query-time answering, we leverage the attention scores generated by the shallow layers to score context frames and employ a diversity-aware selection strategy to retrieve precise and comprehensive evidence. ShallowStream achieves performance on par with the strongest existing streaming methods, while reducing per-frame prefill latency and 10-second end-to-end latency by up to 52.1x and 11.9x, respectively. Our code is available at https://github.com/CURRENTF/ShallowStream.
Chinese Translation
流式视频理解是现实世界应用所需的关键能力,这些应用包括具身智能、自动驾驶、工业监控、监视与预警以及可穿戴助手。然而,使用多模态大语言模型(MLLMs)处理连续视频流在计算上是昂贵的。现有工作已经探索了通过视觉标记剪枝、标记合并、量化、按需帧检索和上下文卸载来减少流式处理开销。然而,大多数现有方法忽略了模型深度的维度。对新到达的帧重复执行全深度MLLM预填充是极其昂贵的,不仅产生大量计算开销,还导致KV缓存以与预填充深度成正比的速率增长。为了应对这些挑战,我们提出了ShallowStream,这是一种新颖框架,利用MLLM的浅层同时进行帧编码和检索索引构建。在流处理过程中,ShallowStream利用浅层的KV缓存维护一个始终在线的轻量级索引。在查询时回答阶段,我们利用浅层生成的注意力分数对上下文帧进行评分,并采用多样性感知的选择策略来检索精确且全面的证据。ShallowStream实现了与现有最强流式方法相当的性能,同时将每帧预填充延迟和10秒端到端延迟分别降低高达52.1倍和11.9倍。我们的代码可在 https://github.com/CURRENTF/ShallowStream 获取。
cs.AI / 52 / 2609.01902
Accurate in space, unreliable in time: how LLMs represent national cultural change
在空间上准确,在时间上不可靠:大语言模型如何表征国家文化变迁
large language model
大语言模型相关
Abstract
Assessments of cultural alignment have become an important part of the development and improvement of large language models (LLMs). However, the majority of the evaluations treat culture as a single snapshot, investigating only whether a model represents a society accurately at the current time. Research in cultural psychology shows that cultural values change at different rates and directions over time. Therefore, a "culturally aware" model should capture not only where a culture is today but also how it has changed over time. We examine this missing dimension of cultural awareness using more than two decades of the World Values Survey data. We compare the cultural trajectories of 40 countries with the trajectories produced by four state-of-the-art (SOTA) LLMs on the Inglehart-Welzel cultural map. Our findings show that while models generally place countries close to their most recent surveyed positions, these representations tend to lag several years behind that position. They also capture only part of the magnitude of the observed change, introduce movement where little occurred, and rarely reproduce reversals in countries' trajectories. These findings point to temporal flattening and suggest that snapshot accuracy can give an incomplete picture of cultural awareness in LLMs and have implications for model evaluation, representational harms, and the governance of culturally aware AI systems.
Chinese Translation
文化对齐评估已经成为大语言模型开发与改进的重要部分。然而,大多数评估将文化视为单一快照,仅考察模型在当前时间是否准确地表征了一个社会。文化心理学研究表明,文化价值观随时间以不同的速率和方向发生变化。因此,“具有文化意识”的模型不仅应捕捉文化当前的状况,还应捕捉其随时间的变化过程。我们利用二十多年的世界价值观调查数据来考察这一缺失的文化意识维度。我们比较了40个国家的文化轨迹与四个最先进的(SOTA)大语言模型在英格尔哈特-韦尔策尔文化地图上产生的轨迹。我们的研究结果表明,虽然模型通常将国家放置在其最近一次调查的位置附近,但这些表征往往滞后于该位置数年。它们还仅捕捉到观察到的变化幅度的一部分,在几乎未发生变动的地方引入移动,并且很少再现国家轨迹中的逆转。这些发现指向时间扁平化,并表明快照准确性可能给大语言模型的文化意识带来不完整的图景,对模型评估、表征危害以及具有文化意识的人工智能系统的治理具有启示意义。
cs.AI / 53 / 2609.02106
Git4Data: Database-Native Version Control for AI Agents
Git4Data:面向AI智能体的数据库原生版本控制
large language model
大语言模型相关
Abstract
Large Language Model (LLM) agents increasingly explore many candidate states of relational data in parallel, each of which should remain isolated, reproducible, and auditable, preferably through the same SQL interface used for ordinary data work. Existing tools support this requirement only partially: source-code version control does not scale to large datasets, whereas relational databases manage large data efficiently but rarely expose native branching, comparison, and merging. We present Git4Data, a database-native version-control layer for agentic workflows. Git4Data treats a database as a repository and a table as a versioned object, exposing Git-style operations (snapshot/tag, branch, diff, and merge with explicit conflict-resolution policies) through SQL extensions. Implemented in MatrixOne, a cloud-native relational database, Git4Data leverages immutable object storage and MVCC to make the cost of these operations proportional to the size of the change rather than the size of the data. On the BranchBench agentic branching workloads, Git4Data outperforms DoltDB by up to an order of magnitude. Overall, we believe this work sheds light on how relational databases can better support AI agents through efficient versioning.
Chinese Translation
大型语言模型(LLM)智能体越来越倾向于并行探索关系数据的许多候选状态,每个候选状态都应保持隔离、可复现且可审计,最好通过用于普通数据工作的同一SQL接口进行访问。现有工具仅能部分满足这一需求:源代码版本控制无法扩展到大型数据集,而关系数据库虽然能高效管理大型数据,却很少暴露原生分支、比较与合并能力。我们提出Git4Data,一个面向智能体工作流的数据库原生版本控制层。Git4Data将数据库视为仓库,将表视为版本化对象,通过SQL扩展提供Git风格的操作(快照/标签、分支、差异比较,以及附带显式冲突解决策略的合并)。Git4Data在云原生关系数据库MatrixOne中实现,利用不可变对象存储和MVCC,使这些操作的成本与变更的大小成正比,而非与数据的大小成正比。在BranchBench智能体分支工作负载上,Git4Data的性能比DoltDB高出一个数量级。总体而言,我们相信这项工作揭示了关系数据库如何通过高效版本化更好地支持AI智能体。
cs.AR / 54 / 2609.01821
Scaling Inference Prefill with High-Radix Photonic Interconnects
利用高基数光子互连扩展推理预填充阶段
large language model
大语言模型相关
Abstract
With the rise of inference as today's dominant AI workload, the industry is transitioning to high-bandwidth photonic interconnects to meet the large scale-up requirements of increasingly complex Mixture-of-Experts (MoE) models. This paper quantifies the benefits of 3D-integrated photonic interconnects for inference prefill by analyzing tradeoffs between high-concurrency throughput for Large Language Model (LLM) chat and the large context windows typically required for reasoning and agentic AI. We simulate three MoE models: short context (1K--8K tokens), medium context (128K tokens), and long context (1M tokens). We evaluate this workload across existing copper-based GPU systems and one with high bandwidth integrated photonics. We show 2.1--3.2x latency improvements in the stressed high-batch regimes and 2.8--5.8x improvements over baselines in communication-limited configurations. 3D photonics enable the 1152-GPU footprint required to lower time-to-first-token, yielding 2.2--4.5x speedups across production-grade platforms when electrical systems cross their inherent scale-up-pod limits.
Chinese Translation
随着推理成为当今主导性AI工作负载,业界正转向高带宽光子互连,以满足日益复杂的混合专家(MoE)模型的大规模纵向扩展需求。本文通过分析大语言模型(LLM)聊天场景所需的高并发吞吐量与推理和智能体AI通常所需的大上下文窗口之间的权衡,量化了三维集成光子互连在推理预填充阶段带来的优势。我们模拟了三种MoE模型:短上下文(1K–8K个token)、中上下文(128K个token)和长上下文(1M个token)。我们在现有的基于铜互连的GPU系统以及一个采用高带宽集成光子学的系统上评估了这一工作负载。我们展示了在高批量压力场景下2.1–3.2倍的延迟改善,以及在通信受限配置中相比基线的2.8–5.8倍改善。三维光子学实现了降低首token延迟所需的1152个GPU规模,当电力系统跨越其固有的纵向扩展机架限制时,在生产级平台上可带来2.2–4.5倍的加速。
cs.AI / 55 / 2609.01976
Knowing Is Not Enough: Information Retrievability as a Precondition to Effective LLM Oversight
知晓并不足够:信息可检索性作为有效LLM监督的前提条件
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly embedded in organizational work, yet their errors often pass human review. Prior research locates such failures in users' capability to review LLM output or their engagement in doing so. We develop an alternative, retrieval-based account of human oversight and posit that error detection is more effective when oversight-relevant information is accessible to users at the moment of review. Across two randomized lab-in-the-field experiments with 640 customer-facing employees, we show that self-generated explanations improve error detection and strengthen recall of verification-relevant reasoning, while cues that reactivate such reasoning help sustain detection under repeated LLM use. Theoretically, we identify information retrievability as a distinct precondition for effective oversight and specify generative encoding and cue-supported reactivation as mechanisms that build and sustain it. Practically, lightweight onboarding self-explanations and daily retrieval cues can make human oversight more resilient as LLM use becomes routine.
Chinese Translation
大语言模型(LLMs)越来越多地嵌入组织工作之中,但它们的错误常常能通过人工审查。以往研究将此类失败归因于用户审查LLM输出的能力或其参与审查的投入程度。我们提出了另一种基于检索的人类监督解释,认为当用户在审查时能够获取到与监督相关的信息时,错误检测会更有效。在两项随机化的实地实验室实验中,共涉及640名面向客户的员工,我们表明自我生成的解释能改善错误检测并增强对验证相关推理的回忆,而重新激活此类推理的线索有助于在反复使用LLM时维持检测效果。在理论上,我们将信息可检索性识别为有效监督的一个独特前提条件,并阐明生成性编码和线索支持的再激活是建立并维持该可检索性的机制。在实践上,轻量级的上手自我解释和日常检索线索可以在LLM使用成为常态时,使人类监督更具韧性。
cs.AI / 56 / 2609.02149
OmegaUse-SOP: SOP Engineering for Professional Computer Use from Human Demonstrations
OmegaUse-SOP:面向人类演示的专业计算机操作的SOP工程
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly evolving from conversational assistants into agents capable of operating external digital environments. Graphical user interface (GUI) agents play an important role in this transition, as many real-world workflows remain accessible only through user-facing software interfaces. However, despite recent progress on general computer-use benchmarks, domain-specific professional standard operating procedures (SOPs) remain challenging for GUI agents because they often involve implicit domain knowledge, software-specific conventions, and task-level verification requirements. We introduce OmegaUse-SOP, a human-in-the-loop SOP Engineering system for transforming human demonstrations of professional computer use into reusable SOP skills for GUI agents. Analogous to prompt engineering, SOP Engineering iteratively refines demonstrations, execution rules, and domain knowledge to convert professional SOPs into reusable GUI-agent skills. OmegaUse-SOP consists of four modules: Observe, Reason, Configure, and Execute. Together, these modules record expert operations as multimodal GUI traces, abstract low-level events into semantic step-level instructions, incorporate domain rules and task-specific parameters, and execute the resulting skills in live GUI environments through step-wise grounding, action generation, and verification. To demonstrate its effectiveness, we collaborate with a power-sector client and test OmegaUse-SOP on photovoltaic simulation workflows in PVsyst 7.2. The results suggest that OmegaUse-SOP can improve GUI-agent reliability on professional SOP tasks, highlighting a practical path toward deploying GUI agents in domain-specific professional software environments.
Chinese Translation
大型语言模型(LLM)正日益从对话助手演变为能够操作外部数字环境的智能体。图形用户界面(GUI)智能体在这一转变中扮演着重要角色,因为许多真实工作流程仍然只能通过面向用户的软件界面访问。然而,尽管近年来通用计算机使用基准取得了进展,特定领域的专业标准操作程序(SOP)对GUI智能体而言仍然具有挑战性,因为它们通常涉及隐性的领域知识、特定软件的习惯用法以及任务级的验证要求。我们提出了OmegaUse-SOP,一个人在环路中的SOP工程系统,用于将专业计算机操作的人类演示转化为可复用的GUI智能体SOP技能。与提示工程相类似,SOP工程通过迭代地优化演示、执行规则和领域知识,将专业SOP转化为可复用的GUI智能体技能。OmegaUse-SOP由四个模块组成:观察(Observe)、推理(Reason)、配置(Configure)和执行(Execute)。这些模块共同将专家操作记录为多模态GUI轨迹,将低级事件抽象为语义化的步骤级指令,整合领域规则和任务特定参数,并通过逐步的定位、动作生成和验证在实时GUI环境中执行所得到的技能。为了证明其有效性,我们与一家电力行业客户合作,在PVsyst 7.2中的光伏仿真工作流程上测试了OmegaUse-SOP。结果表明,OmegaUse-SOP能够提高GUI智能体在专业SOP任务上的可靠性,凸显了在特定领域的专业软件环境中部署GUI智能体的一条实际路径。
cs.LG / 57 / 2609.02162
GenCAR: Generative Counterfactual Alignment with Risk-Controlled Selection for Out-of-Distribution Recommendation
GenCAR:面向分布外推荐的生成式反事实对齐与风险受控选择
large language model
大语言模型相关
Abstract
Serving useful recommendations under distribution shift is crucial for balancing utility and risk in out-of-distribution (OOD) recommendation. However, most existing OOD methods improve ranking or construct counterfactual candidates without controlling the proxy-label false discovery rate (FDR) of the served set. In this work, we formulate OOD serving as the $α$-Valid Counterfactual Recommendation ($α$-VCR) problem to retain candidate support learned from counterfactual supervision while controlling proxy-label FDR, and propose GenCAR, which couples preference-grounded counterfactual supervision with calibrated set selection. In particular, GenCAR fixes the stable-preference representation while intervening on the environmental factor, grounds offline large language model proposals through preference anchors and trust-radius filtering, and uses conformal $p$-values for Benjamini--Hochberg selection. We theoretically bound conditional counterfactual approximation error and prove finite-sample, distribution-free control of proxy-label FDR under exchangeability and positive regression dependence, with a Benjamini--Yekutieli guarantee under arbitrary dependence. Extensive experiments audit realized proxy false discovery proportions and demonstrate that GenCAR consistently enhances OOD candidate recovery across diverse benchmarks.
Chinese Translation
在分布偏移下提供有用的推荐对于在分布外(OOD)推荐中平衡效用与风险至关重要。然而,大多数现有的OOD方法改进排序或构建反事实候选,却未对服务集的代理标签错误发现率(FDR)进行控制。在这项工作中,我们将OOD服务形式化为$α$-有效反事实推荐($α$-VCR)问题,以便在控制代理标签FDR的同时保留从反事实监督中学到的候选支持,并提出GenCAR——它将偏好锚定的反事实监督与校准后的集合选择相结合。特别地,GenCAR在干预环境因子的同时固定稳定偏好表示,通过偏好锚点和信任半径过滤来锚定离线大语言模型提议,并使用共形$p$值进行Benjamini--Hochberg选择。我们从理论上界定了条件反事实近似误差,并证明了在可交换性和正回归依赖下,代理标签FDR具有有限样本、无分布的控制,且在任意依赖下具有Benjamini--Yekutieli保证。大量实验审查了实际实现的代理错误发现比例,并证明GenCAR在多种基准上持续提升OOD候选恢复。
cs.CL / 58 / 2609.02316
Counter-GEO-Bench: Evaluating Defenses Against Information-Distorting Generative Engine Optimization
Counter-GEO-Bench:评估抵御信息扭曲型生成引擎优化的防御
large language model
大语言模型相关
Abstract
Generative engine optimization (GEO) enables content producers to increase the visibility of their web pages in generative search engines, but the same techniques can deliver targeted misinformation when adversaries publish ordinary-looking GEO-optimized documents that victim large language models (LLMs) retrieve and synthesize into distorted answers. No existing benchmark evaluates defenses against this threat under controlled conditions. Therefore, we present Counter-GEO-Bench, a defense benchmark that pairs 247 human-verified, quality-gated queries with information-preserving and information-distorting GEO rewrites, and evaluates defenses on attack success rate (ASR), false positive rate, and answer quality across three victim LLMs. Under Counter-GEO-Bench, three off-the-shelf defenses (Granite Guardian, Llama Guard 3, and NeMo Self-Check Fact-Checking) reduce ASR by at most 5.7% relative, while Granite Guardian's reduction is not statistically significant. Safety-taxonomy guardrails target policy violations, while GEO misinformation passes through them as fluent informational content. To this end, a lightweight benchmark baseline, C-GEO Guard, is proposed, reducing ASR by 47.6% relative with near-zero utility loss, which proves threat tractable.
Chinese Translation
生成引擎优化(GEO)使内容生产者能够提高其网页在生成式搜索引擎中的可见性,但同样的技术也可能被用于传递针对性错误信息:当攻击者发布看似普通的GEO优化文档,这些文档被受害的大语言模型(LLMs)检索并综合成扭曲的答案。目前尚无基准在受控条件下评估针对这一威胁的防御措施。因此,我们提出了 Counter-GEO-Bench,一个防御基准,它将247个经过人工验证和质量把关的查询与信息保留型和信息扭曲型的GEO改写配对,并在三个受害大语言模型上从攻击成功率(ASR)、误报率和答案质量三方面评估防御措施。在 Counter-GEO-Bench 下,三种现成防御(Granite Guardian、Llama Guard 3 和 NeMo Self-Check Fact-Checking)将 ASR 最多相对降低5.7%,而 Granite Guardian 的降低在统计上并不显著。安全分类护栏针对的是政策违规,而 GEO 错误信息则作为流畅的信息内容穿过这些护栏。为此,我们提出了一种轻量级的基准基线 C-GEO Guard,使 ASR 相对降低47.6%,且效用损失近乎为零,从而证明了该威胁是可处理的。
cs.LG / 59 / 2609.01705
Generative Diffusion Surrogates with Analytical Variance Schedule
具有解析方差调度的生成式扩散代理模型
diffusion
扩散模型相关
Abstract
Stochastic transport describes physical systems in which an initially structured distribution spreads under unresolved forcing, scattering, or heterogeneous media. Useful surrogates for such systems should be probabilistic, time-resolved, and able to represent non-Gaussian distributional structure. Generative diffusion models, which corrupt data with Gaussian noise and learn a reverse flow back to structured states, have these properties. Their noise schedules, however, are usually chosen heuristically: image and audio generation---the canonical use cases---provide no physical clock. In transport, by contrast, the variance, or mean-square displacement, is often known from macroscopic theory or empirical scaling even when the full distribution is not. Here we prescribe the forward noising rate as the time derivative of this variance, turning generative time into a calibrated transport clock. The variance path is enforced by construction, while the learned score field represents how non-Gaussian structure inherited from entrance data is smoothed along that path, requiring no intermediate-time physical transport data. For ballistic-to-diffusive transport in turbulent plasmas, the surrogate matches test-particle distributions, reproduces the laboratory-measured variance scale, and tracks the simulated kurtosis evolution without schedule tuning, enabling calibrated emulation and likelihood-based inference.
Chinese Translation
随机输运描述了这样的物理系统:初始具有结构性的分布会在未解析的强迫、散射或非均匀介质作用下发生扩散。此类系统的有用代理模型应当是概率性的、时间分辨的,并且能够表示非高斯的分布结构。生成式扩散模型通过对数据添加高斯噪声并学习反向流回到结构化状态,具备这些性质。然而,它们的噪声调度通常是启发式选择的:图像和音频生成——典型的应用场景——不提供物理时钟。相比之下,在输运问题中,方差或均方位移即使在全分布未知时,也常常可以从宏观理论或经验标度得知。在此,我们将前向加噪速率规定为该方差的时间导数,从而使生成时间转化为经过校准的输运时钟。方差路径通过构造得以保证,而学习到的得分场则表示从入口数据继承的非高斯结构如何沿着该路径被平滑,这不需要中间时刻的物理输运数据。对于湍流等离子体中的弹道-扩散输运,该代理模型能够匹配试验粒子分布,再现实验室测量的方差标度,并跟踪模拟的峰度演化而无需调度调参,从而实现经过校准的仿真和基于似然的推断。
cs.LG / 60 / 2609.01746
CAT-Flow: Curvature-Adaptive sTeps for Flow Matching
CAT-Flow: 用于流匹配的曲率自适应步长
diffusion
扩散模型相关
Abstract
Flow Matching has emerged as a leading framework for generative modeling, powering state-of-the-art systems such as FLUX and Stable Diffusion 3.5. However, the iterative nature of its ODE-based sampling process creates a fundamental efficiency bottleneck: the quality of generated samples is highly sensitive to the choice of step-sizes, and current models typically require 20 to 30 steps for good quality. In this work, we propose two lightweight, training-free algorithms, CAT-OV and CAT-OT that adapt step-sizes at inference time based on a novel connection between Flow Matching sampling and gradient flow. Our algorithms are computed efficiently by not requiring additional neural function evaluations. Specifically, CAT-OT estimates curvature over time via a finite-difference approximation of the time-derivative of the vector field, while CAT-OV approximates curvature over the state space via a gradient of the vector field. Under suitable conditions, both methods have truncation error bounds of constant order. Empirically, CAT-OV and CAT-OT outperform existing step-size heuristics in image quality metrics across four text- to-image Flow Matching models, reducing the number of generation steps required to reach comparable quality by up to 40%.
Chinese Translation
流匹配已成为生成建模的领先框架,支撑着 FLUX 和 Stable Diffusion 3.5 等最先进的系统。然而,其基于 ODE 的采样过程的迭代特性造成了根本性的效率瓶颈:生成样本的质量对步长的选择高度敏感,而当前模型通常需要 20 到 30 步才能获得良好的质量。在这项工作中,我们提出了两种轻量级、无需训练的算法 CAT-OV 和 CAT-OT,它们基于流匹配采样与梯度流之间的新颖联系,在推理时自适应调整步长。我们的算法无需额外的神经函数评估即可高效计算。具体而言,CAT-OT 通过向量场时间导数的有限差分近似来估计随时间变化的曲率,而 CAT-OV 则通过向量场的梯度来近似状态空间上的曲率。在适当条件下,这两种方法的截断误差界均为常数阶。在实验上,CAT-OV 和 CAT-OT 在四个文本到图像流匹配模型的图像质量指标上优于现有的步长启发式方法,将达到相当质量所需的生成步数最多减少了 40%。
cs.LG / 61 / 2609.01756
A Study of Conditional Diffusion Models for Open-Loop Control under Dry Friction and Stiction
干摩擦与静摩擦下开环控制的条件扩散模型研究
diffusion
扩散模型相关
Abstract
Diffusion models have recently emerged as expressive generative priors for planning and control. This paper studies Action Diffusion, an action-sequence diffusion formulation used as an open-loop proposal distribution for a point-mass system with dry friction and stiction. In this benchmark, motion starts only when the applied input exceeds a static-friction threshold, so effective controls occupy a small and temporally structured subset of the action-sequence space. A compact conditional 1D U-Net generates bounded control sequences conditioned on initial and target states. We compare it with uniform random shooting, random shooting from the same structured dataset prior, and the Cross-Entropy Method (CEM). Results show that Action Diffusion reduces terminal error and stuck steps, especially in low-sample regimes. These results indicate that conditional diffusion provides an effective mechanism for generating temporally coherent control sequences that overcome stiction by conditioning and recombining structured control primitives from the training prior for state-to-state open-loop control.
Chinese Translation
扩散模型最近已成为规划与控制中具有表现力的生成先验。本文研究动作扩散(Action Diffusion),这是一种动作序列扩散公式,用作具有干摩擦和静摩擦的点质量系统的开环提议分布。在该基准测试中,只有当施加的输入超过静摩擦阈值时运动才开始,因此有效控制占据了动作序列空间中一个小的且时间结构化的子集。一个紧凑的条件一维U-Net生成以初始状态和目标状态为条件的受限控制序列。我们将其与均匀随机射击、来自相同结构化数据集先验的随机射击以及交叉熵方法(CEM)进行比较。结果表明,动作扩散降低了终端误差和卡住步数,尤其是在低样本情况下。这些结果表明,条件扩散提供了一种有效机制,通过条件化并重组来自训练先验的结构化控制基元,生成克服静摩擦的时间连贯控制序列,用于状态到状态的开环控制。
cs.LG / 62 / 2609.01807
hLLM: Single Pass Decoding for Generative Reranking
hLLM:用于生成式重排序的单遍解码
large language model
大语言模型相关
Abstract
Large language models (LLMs) achieve state-of-the-art generative ranking quality, but the ranking they produce must be decoded, and autoregressive decoding spends one sequential forward pass per emitted token. We observe that the only tokens a ranker must emit are the $N$ ordinal values naming the items in ranked order, and that this narrow, permutation-structured output format admits decoding strategies which are much more efficient than left-to-right generation. We introduce hLLM (Hungarian LLM), a format-specialized decoding strategy that decodes all $N$ ordinals in $O(1)$ forward passes. hLLM reads an $N \times K$ item-position score matrix off the LLM's prefill hidden states with a lightweight self-attention head, then decodes the ordinals as the optimal bipartite assignment of that matrix via the Hungarian algorithm, yielding a valid permutation by construction rather than by repair. Through a systematic study of training signals and backbone adaptation, we show that LoRA-based fine-tuning combined with teacher ranking distillation reaches 28 ms end-to-end inference, a speed-up of $64\times$ while maintaining ranking quality on par with the teacher. We provide a complete ablation decomposing the contributions of architecture, training signal, and backbone adaptation. Our framework connects generative ranking to combinatorial optimization, opening a path toward other $O(1)$-decode mechanisms for real-time ranking.
Chinese Translation
大语言模型(LLMs)取得了最先进的生成式排序质量,但它们产生的排序必须经过解码,而自回归解码为每个生成的 token 都要花费一次顺序的前向传播。我们观察到,排序器必须生成的唯一 token 是以排序顺序命名条目的 $N$ 个序数值,而这种狭窄的、排列结构的输出格式允许采用比从左到右生成高效得多的解码策略。我们提出了 hLLM(Hungarian LLM),一种格式特化的解码策略,能以 $O(1)$ 次前向传播解码全部 $N$ 个序数值。hLLM 通过一个轻量级自注意力头从 LLM 的预填充隐藏状态中读取一个 $N \times K$ 的条目-位置得分矩阵,然后通过匈牙利算法将该矩阵的最优二分匹配作为序数值解码出来,从而通过构造而非修复得到有效排列。通过对训练信号和主干网络适配的系统研究,我们表明,基于 LoRA 的微调结合教师排序蒸馏可以达到 28 毫秒的端到端推理,实现了 $64\times$ 的加速,同时保持与教师相当的排序质量。我们提供了一个完整的消融实验,分解了架构、训练信号和主干网络适配各自的贡献。我们的框架将生成式排序与组合优化联系起来,为其他用于实时排序的 $O(1)$ 解码机制开辟了一条道路。
cs.LG / 63 / 2609.02042
Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM Agents
多行动,少决策:面向长时程 LLM 智能体的技能引导自适应动作分块
large language model
大语言模型相关
Abstract
Large language model (LLM) agents for long-horizon interactive tasks typically follow a ReAct-style protocol, issuing one primitive action per LLM round. While this enables frequent replanning, it is inefficient for long-horizon tasks where many rounds are spent on routine action sequences. A natural alternative is to let the agent emit variable-length action chunks. However, naively training such policies with standard reinforcement learning fails: the agent either collapses to single-action behavior or over-commits to excessively long sequences. Both failures share a common root cause: the inability to learn chunk boundaries. We propose SPACE, which addresses this challenge by distilling chunk-boundary supervision from trajectory-induced programmatic skills. We induce two-level programmatic skills from successful trajectories, where subskill boundaries serve as direct chunk-boundary supervision. This temporal structure is then distilled into a primitive-chunk policy via hybrid on-/off-policy optimization with chunk-aware credit assignment. Experiments on ALFWorld and ScienceWorld show that SPACE improves success rates by 7.0%-31.3% over the strongest baseline in each setting while reducing average LLM decision rounds by up to 78.9%.
Chinese Translation
用于长时程交互任务的大型语言模型(LLM)智能体通常遵循 ReAct 式协议,每轮 LLM 只产生一个原始动作。虽然这能支持频繁的重新规划,但对于长时程任务而言效率低下,因为许多轮次都花费在例行动作序列上。一种自然的替代方案是让智能体输出可变长度的动作块。然而,使用标准强化学习来训练这类策略会失败:智能体要么退化为单动作行为,要么过度执着于过长序列。这两种失败有一个共同根源:无法学习分块边界。我们提出 SPACE,通过从轨迹诱导的程序性技能中蒸馏分块边界监督信息来解决这一挑战。我们从成功轨迹中归纳出两级程序性技能,其中子技能边界可直接作为分块边界监督信息。随后,这种时间结构通过混合在线/离线策略优化以及具有分块感知的信用分配被蒸馏到原始动作-分块策略中。在 ALFWorld 和 ScienceWorld 上的实验表明,与每种设置中最强的基线相比,SPACE 将成功率提高了 7.0% 到 31.3%,同时将平均 LLM 决策轮数最多减少了 78.9%。
cs.LG / 64 / 2609.02068
DynG-Diff: A State-Aware Dynamic Guidance Diffusion Framework for Probabilistic Time Series Forecasting
DynG-Diff:面向概率时间序列预测的状态感知动态引导扩散框架
diffusion
扩散模型相关
Abstract
Probabilistic multivariate time series (MTS) forecasting is crucial for modeling complex dynamical systems. However, existing diffusion-based methods rely on task-specific conditional paradigms that lack flexibility and struggle with inherent "information heterogeneity"--the significantly varying noise levels and evolutionary patterns across variables. To address this, we propose DynG-Diff, a variable-sensitive dynamic guidance diffusion framework for probabilistic multivariate time-series forecasting: (1) DynG-Diff adopts a two-stage separated training strategy and uses an unconditional diffusion backbone to model the joint distribution of multivariate time series. (2) DynG-Diff introduces a lightweight state-aware policy network that adaptively infers variable reliability from real-time noisy states and one-step denoising estimates, outputting a dynamic guidance strength matrix. (3) DynG-Diff mathematically formulates this dynamic weight as the local precision of the observation distribution, enabling precise guidance for high-confidence variables during inference while filtering out interference from anomalous noise. Extensive experiments on real-world benchmarks demonstrate competitive probabilistic forecasting performance against state-of-the-art conditional diffusion models and improved robustness under severe observation corruption.The implementation code is available at: https://github.com/TT-20011031/DynG-Diff
Chinese Translation
概率多变量时间序列(MTS)预测对于建模复杂动态系统至关重要。然而,现有的基于扩散的方法依赖于任务特定的条件范式,这些范式缺乏灵活性,并且在面对固有的“信息异质性”——即各变量间噪声水平和演化模式的显著差异——时表现不佳。为了解决这一问题,我们提出了 DynG-Diff,一个面向概率多变量时间序列预测的变量敏感的动态引导扩散框架:(1)DynG-Diff 采用两阶段分离训练策略,并使用无条件扩散主干来建模多变量时间序列的联合分布。(2)DynG-Diff 引入了一个轻量级的状态感知策略网络,该网络从实时噪声状态和单步去噪估计中自适应地推断变量可靠性,并输出一个动态引导强度矩阵。(3)DynG-Diff 将该动态权重数学上表述为观测分布的局部精度,从而能够在推理过程中对高置信变量进行精确引导,同时滤除异常噪声的干扰。在真实世界基准上的大量实验表明,该方法相对于最先进的条件扩散模型具有有竞争力的概率预测性能,并在严重观测污染下具有改进的鲁棒性。实现代码可在以下网址获取:https://github.com/TT-20011031/DynG-Diff
cs.LG / 65 / 2609.02160
GeoSPRINT: Geometric Redundancy-Aware Step Pruning for Inference in Diffusion Trajectories
GeoSPRINT:扩散轨迹推理中基于几何冗余感知的步长剪枝
diffusion
扩散模型相关
Abstract
Diffusion models achieve high sample quality but remain expensive at inference time because sampling requires many sequential neural function evaluations (NFEs). Existing acceleration methods either use fixed step-skipping schedules, adapt step sizes based on local numerical error, or require additional training. We introduce GeoSPRINT (Geometric Step Pruning for Inference in Trajectories), a training-free framework for constructing non-uniform sampling schedules from the geometry of denoising trajectories. GeoSPRINT detects geometrically redundant steps using a hyperplanarity test in latent space, implemented efficiently via QR factorization, and converts the resulting redundancy profile into a sampling schedule that allocates more steps to high-curvature regions of the trajectory. In addition, we introduce the trajectory projection score $α_{\mathrm{traj}}$, a residual-variance metric that quantifies trajectory straightness and serves as a model-free diagnostic for rectified flow quality. Across CIFAR-10 ($32{\times}32$), LSUN Church ($256{\times}256$), and Stable Diffusion v1.5 ($512{\times}512$ latent), GeoSPRINT consistently improves over uniform DDIM (Denoising Diffusion Implicit Models) schedules at matched NFE budgets. On CIFAR-10, GeoSPRINT improves FID (Fréchet Inception Distance) by 0.7-1.1 over DDIM across 49-89 NFEs and surpasses DPM-Solver++ at NFE${\geq}30$ despite using a first-order DDIM solver. On LSUN Church, it reduces FID from 1.48 to 1.26 at 52 steps, and on Stable Diffusion v1.5 it achieves up to 1.93 FID improvement over DDIM. These results show that trajectory geometry provides a useful global signal for allocating inference steps and that schedule quality can substantially improve diffusion sampling efficiency without retraining.
Chinese Translation
扩散模型实现了高样本质量,但在推理时成本依然高昂,因为采样需要多次连续的神经函数评估(NFEs)。现有的加速方法要么使用固定的跳步调度,要么基于局部数值误差调整步长,要么需要额外训练。我们提出了 GeoSPRINT(轨迹推理中的几何步长剪枝),一种无需训练、根据去噪轨迹的几何构造非均匀采样调度的框架。GeoSPRINT 借助潜空间中的超平面性检验来检测几何上冗余的步骤,该检验通过 QR 分解高效实现,并将由此产生的冗余分布转换为采样调度,为轨迹的高曲率区域分配更多步骤。此外,我们引入了轨迹投影分数 $\alpha_{\mathrm{traj}}$,这是一种残差方差度量,用于量化轨迹的平直度,并可作为整流流质量的一种免模型诊断指标。在 CIFAR-10($32{\times}32$)、LSUN Church($256{\times}256$)以及 Stable Diffusion v1.5($512{\times}512$ 潜空间)上,GeoSPRINT 在匹配 NFE 预算下始终优于均匀 DDIM(去噪扩散隐式模型)调度。在 CIFAR-10 上,在 49-89 个 NFE 范围内,GeoSPRINT 将 FID(Fr\'echet 起始距离)相对于 DDIM 改善了 0.7-1.1,并且尽管使用了一阶 DDIM 求解器,在 NFE${\geq}30$ 时超越了 DPM-Solver++。在 LSUN Church 上,它在 52 步时将 FID 从 1.48 降低到 1.26;在 Stable Diffusion v1.5 上,相对于 DDIM 最多实现了 1.93 的 FID 改善。这些结果表明,轨迹几何为分配推理步骤提供了有效的全局信号,并且调度质量可以在不重新训练的情况下显著提高扩散采样的效率。
cs.LG / 66 / 2609.02170
DMRL: Document-Mediated Reinforcement Learning for Skill Optimization in Advertising Recommendation
DMRL:用于广告推荐中技能优化的文档介导强化学习
large language model
大语言模型相关
Abstract
Advertising recommendation requires continuously tuning complex system parameters while balancing commercial returns and user experience. Recent work has introduced large language models (LLMs) with skill documents to assist this labor-intensive process, but skill optimization remains largely prompt-driven, lacking a principled mechanism to attribute rewards to specific document edits. To address this limitation, we propose Document-Mediated Reinforcement Learning (DMRL), a skill self-evolution framework that models skill document optimization as a sequence of structured editing actions. In DMRL, an upper-level agent performs controlled document edits, while a frozen lower-level task agent evaluates their effects through A/B testing. To address credit assignment and long-term outcomes, we introduce two key components: (1) Dual-Relative Policy Optimization (DRPO), a post-training policy optimization method for robust and risk-aware advantage estimation; and (2) Long-term Reward Predictor (LRP), which estimates long-term outcomes by modeling population heterogeneity with disentangled representation learning and cross-attention transfer. DMRL was deployed on a large-scale short-video ads platform and extensive empirical evaluation shows that DMRL outperforms state-of-the-art baselines across key advertising metrics
Chinese Translation
广告推荐需要在平衡商业回报与用户体验的同时,持续调整复杂的系统参数。近期研究引入了带有技能文档的大语言模型(LLMs)来辅助这一劳动密集型过程,但技能优化在很大程度上仍由提示驱动,缺乏将奖励归因于特定文档编辑的原则性机制。为解决这一局限性,我们提出了文档介导强化学习(DMRL),这是一种技能自我进化框架,将技能文档优化建模为一系列结构化编辑动作。在DMRL中,上层智能体执行受控的文档编辑,而冻结的下层任务智能体通过A/B测试评估其效果。为解决信用分配和长期结果问题,我们引入了两个关键组件:(1) 双相对策略优化(DRPO),一种用于稳健且风险感知的优势估计的后训练策略优化方法;(2) 长期奖励预测器(LRP),其通过使用解耦表示学习和交叉注意力迁移对人群异质性进行建模,从而估计长期结果。DMRL已部署在一个大规模短视频广告平台上,广泛的经验评估表明,DMRL在关键广告指标上优于最先进的基线方法。
cs.LG / 67 / 2609.02293
SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment
SEAL:通过共享专家对齐强化混合专家模型的全局安全性
large language model
大语言模型相关
Abstract
Mixture-of-Experts (MoE) is a scaling architecture for large language models that activates only a small subset of expert modules per token, enabling massive parameter growth with nearly constant computation. Recent Hybrid MoE architecture adds \textit{shared experts} to capture consistently useful representations, further improving stability and generalization. MoE now powers many flagship open-source and commercial models, yet remains vulnerable to adversarial attacks. Specifically, sparse routing introduces a structural vulnerability: MoE safety hinges on which experts are activated, and adversaries can subvert this selection through jailbreak prompts, malicious fine-tuning, and weight-level pruning of safety-critical neurons. Existing defenses primarily focus on hardening the router, but an adversary may still manipulate or bypass the routing trajectory due to the routing process's nondeterministic nature, thereby collapsing the defense. To cope with this problem, we first identify theoretically and empirically that shared expert, an always-activated component containing a small proportion of safety-critical neurons, can overcome the uncertainty of sparsely activated routing path and serve as a router-independent anchor to enhance global safety alignment. Based on this insight, we propose SEAL, a training-time parameter-efficient defense that produces a plug-and-play adapter attached to shared expert, and SEAL++, a variant that adds an orthogonal constraint preserving pre-existing safety subspaces during training. We evaluate SEAL and SEAL++ across six attack scenarios that combine three adversarial inputs (harmful prompting, jailbreak, malicious fine-tuning) with and without neuron pruning. SEAL reduces attack success rate (ASR) by up to 60\%, at a capability cost of at most 1.4\% on a five-benchmark average. Additionally, SEAL can seamlessly integrate with router-level ......
Chinese Translation
混合专家(MoE)是一种用于大语言模型的扩展架构,它对每个词元仅激活一小部分专家模块,从而在计算量几乎不变的情况下实现参数的大规模增长。最近的混合MoE架构引入了\textit{共享专家}来捕获一致有用的表示,进一步提升了稳定性和泛化能力。MoE目前支撑着许多旗舰级开源和商业模型,但仍然容易受到对抗性攻击的影响。具体而言,稀疏路由引入了一种结构性漏洞:MoE的安全性取决于哪些专家被激活,而攻击者可以通过越狱提示、恶意微调以及对安全关键神经元的权重级剪枝来破坏这种选择。现有的防御措施主要集中于加固路由器,但由于路由过程的非确定性本质,攻击者仍可能操纵或绕过路由轨迹,从而使防御崩溃。为了解决这一问题,我们首先从理论和实证上识别出,共享专家——一个始终被激活、包含少量安全关键神经元的组件——能够克服稀疏激活路由路径的不确定性,并充当独立于路由器的锚点,以增强全局安全对齐。基于这一洞察,我们提出了SEAL,这是一种训练时参数高效的防御方法,可生成一个附着于共享专家的即插即用适配器;以及SEAL++,一种变体,在训练过程中添加正交约束以保留预先存在的安全子空间。我们在六种攻击场景下评估了SEAL和SEAL++,这些场景将三种对抗性输入(有害提示、越狱、恶意微调)与是否进行神经元剪枝相结合。SEAL可将攻击成功率(ASR)降低高达60\%,而在五个基准测试的平均性能上,能力损失最多为1.4\%。此外,SEAL可以无缝集成到路由器级别的......
cs.LG / 68 / 2609.02548
Learn from Whoever Is Right: Answer-Verified Multi-Teacher Distillation for Multi-Domain LLMs
从正确者学习:用于多领域大语言模型的答案验证式多教师蒸馏
large language model
大语言模型相关
Abstract
Modern large language models (LLMs) rely on reinforcement learning to build strong capabilities in individual domains, but integrating those capabilities into a single deployable model remains challenging. By routing each sample to the teacher whose domain matches it, existing approaches let a domain label decide which teacher provides supervision. However, domain expertise holds only on average: the matched teacher is not always correct on a given sample, while a teacher from another domain sometimes is. The reliable teacher therefore has to be identified per sample, not per domain. In this paper, we introduce Multi-Teacher Self-Distillation Policy Optimization (MT-SDPO), an on-policy distillation method that unifies several frozen teachers into one student model. MT-SDPO consists of three components: (1) self-anchors, where a rollout is supervised by a correct rollout from its own group; (2) answer-verified eligibility, where a teacher may supervise a sample only if its own answer passes a verifier; and (3) privileged distillation, which merges the anchor and all verified feedback into one context that an exponential moving average self-teacher reads and the student does not, thereby keeping one policy at deployment. Across five students from three model families, MT-SDPO lifts the weakest domain of Qwen3-8B by 14.79 points and narrows its domain gap by 74.7%, a better balance than serving one matched teacher per domain. Verified reliability, not domain membership, should decide who teaches. Code is available at https://github.com/hexixiang/MT-SDPO.
Chinese Translation
现代大语言模型(LLMs)依靠强化学习构建单个领域的强大能力,但将这些能力整合进一个可部署的模型中仍然具有挑战性。现有的方法通过将每个样本路由到其领域匹配的教师模型,让领域标签决定由哪个教师提供监督。然而,领域专长只具有平均意义:匹配的教师并不总是在给定样本上正确,而来自其他领域的教师有时却是正确的。因此,必须针对每个样本而非每个领域来识别可靠的教师。在本文中,我们提出了多教师自蒸馏策略优化(MT-SDPO),这是一种在策略蒸馏方法,它将多个冻结的教师模型统一为一个学生模型。MT-SDPO 包含三个组件:(1)自锚定,即一条轨迹由其自身组内的一条正确轨迹进行监督;(2)答案验证资格,即教师只有在自己的答案通过验证器时才有资格对样本进行监督;(3)特权蒸馏,它将锚定和所有通过验证的反馈合并到同一个上下文中,该上下文由指数移动平均自教师读取而学生不读取,从而在部署时保持单一策略。在来自三个模型家族的五个学生模型上,MT-SDPO 将 Qwen3-8B 的最弱领域提升了 14.79 个百分点,并将其领域差距缩小了 74.7%,比每个领域仅部署一个匹配教师取得了更好的平衡。谁来进行教学应由经过验证的可靠性而非领域归属来决定。代码可在 https://github.com/hexixiang/MT-SDPO 获取。
cs.LG / 69 / 2609.02817
Cliff: Learning Process Rewards from the First Mistake
Cliff:从第一个错误中学习过程奖励
large language model
大语言模型相关
Abstract
Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for large language model (LLM) post-training, but its reliance on coarse outcome rewards leads to limited guidance on intermediate reasoning processes. Existing approaches such as process reward modeling and on-policy distillation introduce additional constraints, such as reliance on a specialized reward model or assuming identical reasoning patterns between teacher and student. Nevertheless, we observe that once a reasoning process first goes wrong, evaluating the subsequent reasoning provides limited additional information, as it is already conditioned on an invalid prefix. Therefore, we propose Cliff, a reward shaping strategy that utilizes an off-the-shelf LLM as a teacher to identify the first mistake in each rollout. As a result, the rollout is naturally decomposed into two parts: a correct prefix and an incorrect suffix. Cliff then converts this signal into token-level advantages, assigning positive advantages for the correct prefix and negative feedback afterward. Experiments across 12 different scenarios demonstrate that Cliff consistently improves reasoning performance, outperforming on-policy distillation by 15% and standard GRPO by 7%, even with teachers of modest capability. Furthermore, we analyse the role of ``ground truth'' in Cliff and investigate its training dynamics. These results establish Cliff as a simple, general and effective approach for improving RLVR with richer, fine-grained supervision.
Chinese Translation
基于可验证奖励的强化学习(RLVR)已成为大型语言模型(LLM)后训练的一种强大范式,但其对粗粒度结果奖励的依赖导致对中间推理过程的指导有限。现有方法如过程奖励建模和同策略蒸馏引入了额外限制,例如依赖专门的奖励模型,或假设教师与学生具有相同的推理模式。尽管如此,我们观察到,一旦推理过程首次出错,对后续推理的评估就只能提供有限的额外信息,因为它已经以无效前缀为条件。因此,我们提出了 Cliff,一种奖励塑形策略,它利用一个现成的 LLM 作为教师,以识别每条轨迹中的第一个错误。这样,轨迹自然分解为两部分:正确前缀和错误后缀。Cliff 接着将该信号转换为词元级优势,为正确前缀赋予正优势,并在之后的部分提供负反馈。在 12 个不同场景上的实验表明,Cliff 一致地提升了推理性能,即使使用能力一般的教师,也比同策略蒸馏高 15%、比标准 GRPO 高 7%。此外,我们分析了 Cliff 中“真值”的作用,并研究了它的训练动态。这些结果确立了 Cliff 是一种简单、通用且有效的方法,可通过更丰富、更细粒度的监督来改进 RLVR。
cs.LG / 70 / 2609.02849
Post-Training Language Models for Gold-Medal Performance in Coding Competitions
面向编程竞赛金牌表现的后训练语言模型
large language model
大语言模型相关
Abstract
Competitive programming has become a key test of large language model reasoning, with international competitions such as IOI and ICPC representing its most challenging settings. We present an end-to-end specialization pipeline combining large-scale problem curation, synthetic reasoning traces, supervised fine-tuning (SFT), and reinforcement learning (RL). Using 22,000 curated problems, we train Nemotron-3-Nano-CC (30B-A3B) with SFT and RL and Nemotron-3-Ultra-CC (550B-A55B) with SFT alone. We further introduce GenCorrect, a feedback-driven test-time compute strategy that iteratively generates, evaluates, and refines diverse solutions. On IOI 2025, Nano-CC improves from 130 points to 291 after post-training and to 468 with GenCorrect, exceeding the gold threshold of 438.3 while Ultra-CC reaches 502. Guided by these results, we develop a competition-specific Ultra-CC system and evaluate it prospectively during IOI 2026. Under the same time, internet-access, and submission constraints as human contestants, it scores 535.4 out of 600, exceeding both the gold threshold of 361.12 and the top human score of 498.27. To our knowledge, this is the first AI system to outscore the highest-scoring human contestant on an IOI problem set.
Chinese Translation
竞技编程已成为大型语言模型推理能力的关键测试,其中IOI和ICPC等国际竞赛代表了最具挑战性的场景。我们提出了一种端到端的特化流水线,结合大规模问题精选、合成推理轨迹、监督微调(SFT)和强化学习(RL)。使用22,000个精选问题,我们使用SFT和RL训练了Nemotron-3-Nano-CC(30B-A3B),并仅使用SFT训练了Nemotron-3-Ultra-CC(550B-A55B)。我们进一步引入了GenCorrect,一种反馈驱动的测试时计算策略,它可以迭代地生成、评估和完善多种解决方案。在IOI 2025上,Nano-CC经过后训练从130分提高到291分,使用GenCorrect后提高到468分,超过了438.3分的金牌门槛,而Ultra-CC则达到了502分。在这些结果的指导下,我们开发了一个面向特定竞赛的Ultra-CC系统,并在IOI 2026期间对其进行了前瞻性评估。在与人类参赛者相同的时间、网络访问和提交限制下,它获得了600分中的535.4分,同时超过了361.12分的金牌门槛和498.27分的最高人类得分。据我们所知,这是第一个在IOI题目集上得分超过最高分人类参赛者的AI系统。
cs.MA / 71 / 2609.02250
RideSkill: A Hierarchical Algorithm for Generalized Ride Sharing with LLM-Driven Automatic Evolution
RideSkill:一种结合LLM驱动的自动进化的通用拼车分层算法
large language model
大语言模型相关
Abstract
Ride-sharing, which allows multiple passengers with different origin-destination (OD) pairs to share a single vehicle, is a challenging operational problem, as it requires orders with different OD pairs to be efficiently bundled and assigned to vehicles under uncertain and varying scenarios. Although multi-agent reinforcement learning (MARL) solutions have achieved promising performance, they suffer from limited generalization (adapting to different environmental scenarios), low transferability (adapting to different platform objectives), and training difficulties in large-scale systems, such as the curse of dimensionality. Recently, motivated by the scaling of large language models (LLMs), several works have incorporated LLMs into ride-hailing systems, either by employing LLMs directly as decision-making agents or using them for automatic algorithm design. However, none of these approaches support vehicle sharing, which complicates the problem by expanding both the state and action spaces exponentially. Moreover, most of them require frequent LLM calls at inference time, making them infeasible for real-time deployment. To address these issues, we propose RideSkill, a hierarchical method for ride-sharing that leverages LLM-assisted automatic algorithmic design. RideSkill consists of a combiner that assigns appropriate skills to each vehicle from a learned skill repository, enabling adaptive dispatch under varying scenarios and objectives, and a repositioner that sequentially relocates idle vehicles to emerging regions, avoiding conflicts among vehicles. Crucially, the skill repository, combiner, and repositioner are all trained by an LLM-based automatic evolutionary method, eliminating the need for LLM calls during deployment and thus ensuring high real-time performance.
Chinese Translation
拼车服务允许多个具有不同起讫点(OD)对的乘客共乘一辆车,这是一个具有挑战性的运营问题,因为它要求在不确定和变化的情景下,将不同OD对的订单高效地捆绑并分配给车辆。尽管多智能体强化学习(MARL)解决方案已取得令人鼓舞的性能,但它们仍存在泛化能力有限(适应不同环境情景)、可迁移性低(适应不同平台目标)以及在大规模系统中的训练困难(例如维数灾难)等问题。近期,受大语言模型(LLMs)规模化的启发,一些工作已将LLMs整合到网约车系统中,要么直接将LLMs用作决策智能体,要么将其用于自动算法设计。然而,这些方法均不支持车辆共享,而车辆共享会同时指数级地扩大状态空间和动作空间,从而使问题复杂化。此外,它们中的大多数在推理时需要频繁调用LLM,这使得它们无法用于实时部署。为解决这些问题,我们提出了RideSkill,一种利用LLM辅助的自动算法设计的分层拼车方法。RideSkill由一个组合器和一个重新定位器组成。组合器从学习的技能库中为每辆车分配适当的技能,从而在不同情景和目标下实现自适应调度;重新定位器则顺序地将空闲车辆重新定位到新出现的区域,避免车辆之间的冲突。关键的是,技能库、组合器和重新定位器均通过基于LLM的自动进化方法进行训练,从而消除了部署期间对LLM调用的需求,并因此确保了高实时性能。
cs.MA / 72 / 2609.02580
Competitive Market Behavior of LLMs
大语言模型的竞争性市场行为
large language model
大语言模型相关
Abstract
Large language models (LLMs) are increasingly deployed as economic agents, yet there is little evidence whether LLM agents are suited for participating in market mechanisms designed for humans, and whether these mechanisms deliver desired outcomes when faced with LLM agents. We address this question by replicating seminal economic experiments, replacing human subjects with LLM agents. We place agents in a double auction environment, which is a widely-used market mechanism. We check whether such a market is able to deliver an efficient allocation of resources, thereby testing a novel dimension of alignment of LLM agents -- their compatibility with a fundamental market mechanism. We find that markets populated by LLM agents exhibit slower or no convergence towards market equilibrium, thus providing less efficient allocations than markets populated by humans. We then analyze agents' individual trading decisions and find substantial heterogeneity both across model families and market roles. We also run a lexical analysis of Chain-of-Thought (CoT) traces generated by the agents. We find that the decision to execute a trade rather than continue incrementally adjusting prices is associated with a shift from strategic considerations toward urgency. We publicly release our testing framework, which can be used for future evaluations.
Chinese Translation
大语言模型(LLM)越来越多地被部署为经济主体,然而目前几乎没有证据表明LLM主体是否适合参与为人类设计的市场机制,以及这些机制在面对LLM主体时是否能产生预期的结果。我们通过复现开创性的经济学实验来解决这一问题,用LLM主体替代人类被试。我们将主体置于双向拍卖环境中,这是一种广泛使用的市场机制。我们检验这种市场是否能够实现资源的有效配置,从而测试LLM主体对齐的一个新维度——它们与基本市场机制的兼容性。我们发现,由LLM主体构成的市场在趋向市场均衡方面表现出更慢的收敛速度,甚至完全不收敛,因此其配置效率低于由人类构成的市场。随后,我们分析了主体的个体交易决策,发现不同模型家族和市场角色之间存在显著的异质性。我们还对主体生成的思维链(CoT)轨迹进行了词汇分析。我们发现,执行交易而非继续逐步调整价格的决策,与从战略考量向紧迫感的转变相关。我们公开发布了我们的测试框架,可用于未来的评估。
cs.AI / 73 / 2609.02082
Transfer Safety Awareness for Cross-Modal Safety Drift in Multimodal Large Language Models
多模态大语言模型中面向跨模态安全漂移的安全意识迁移
large language model
大语言模型相关
Abstract
Visual modality enhances the capabilities of multimodal large language models (MLLMs) but also introduces a safety concern: a benign textual query may convey harmful intent when grounded in a visual image. We term this cross-modal safety drift and our pilot studies show that the safety response rate for such requests is substantially lower than that for requests containing explicitly unsafe text. This paper aims to systematically study this issue. First, we conduct an empirical analysis to identify representative unsafe response patterns. Building on these, we interpret model representations and attentions, revealing that visually risky cues receive limited attention and weakly trigger refusal. Motivated by the observation that safety signals from unsafe text processing can be transferred, we propose safety-awareness representation transfer (SRT), a lightweight direction-refinement method that mitigates cross-modal safety drift with a frozen MLLM backbone. Experiments across multiple benchmarks and models show that SRT effectively improves safety in diverse cross-modal settings while preserving utility. Code is available at https://github.com/cucu220123/safety-awareness.
Chinese Translation
视觉模态增强了多模态大语言模型(MLLMs)的能力,但也带来了一个安全问题:一个良性的文本查询在与视觉图像结合时可能传达有害意图。我们将此称为跨模态安全漂移,我们的初步研究表明,此类请求的安全响应率显著低于包含明确不安全文本的请求。本文旨在系统地研究这一问题。首先,我们进行了实证分析,以识别具有代表性的不安全响应模式。在此基础上,我们解读模型表示和注意力,揭示了视觉风险线索获得的注意力有限,并且只能微弱地触发拒绝。受到来自不安全文本处理的安全信号可以被迁移这一观察的启发,我们提出了安全意识表示迁移(SRT),这是一种轻量级的方向精炼方法,可在冻结的MLLM主干下缓解跨模态安全漂移。在多个基准和模型上的实验表明,SRT在多种跨模态设置中有效提高了安全性,同时保留了效用。代码可在 https://github.com/cucu220123/safety-awareness 获取。
cs.NE / 74 / 2609.02353
LLM-Driven Joint Evolution of Coupled Heuristics Components for Routing Optimization
面向路由优化的耦合启发式组件的LLM驱动联合进化
large language model
大语言模型相关
Abstract
Heuristic design for combinatorial optimization remains heavily reliant on expert knowledge, while existing large language model (LLM)-enhanced evolutionary methods typically evolve isolated algorithmic components, even when one determines the search state on which another operates. This paper proposes LLM-driven Heuristic Components Joint Generation (LLM-HCJG), a population-based framework that jointly generates and co-evolves interdependent heuristic components under a shared design blueprint. Applied to guided local search (GLS), LLM-HCJG couples solution initialization with penalty construction and embeds the generated pair into an enhanced online search mechanism. The resulting form is further transferred from the traveling salesman problem (TSP) to the capacitated vehicle routing problem (CVRP). Theoretical analysis establishes the non-separable state-transition effects between the two components and the advantage in generation consistency. Across synthetic instances and 41 public TSPLIB/CVRPLIB benchmarks, LLM-HCJG attains consistently low optimality gaps, including best or tied-best results on 28 of 29 TSPLIB instances and all 12 CVRPLIB instances. Ablation and structural analyses further indicate that these gains are associated with cross-component compatibility and alignment rather than isolated-component recombination. These results support effective cross-instance transfer within the evaluated routing settings under limited-sample, modest-cost training.
Chinese Translation
组合优化的启发式设计仍然严重依赖专家知识,而现有的大语言模型(LLM)增强进化方法通常进化孤立的算法组件,即使其中一个组件决定了另一个组件所操作的搜索状态。本文提出了LLM驱动的启发式组件联合生成(LLM-HCJG),一种基于种群的框架,在共享设计蓝图下联合生成并共同进化相互依赖的启发式组件。应用于引导式局部搜索(GLS)时,LLM-HCJG将解初始化与惩罚构建耦合,并将生成的配对嵌入增强的在线搜索机制中。所得形式进一步从旅行商问题(TSP)迁移到带容量约束的车辆路径问题(CVRP)。理论分析确立了两个组件之间不可分离的状态转移效应以及在生成一致性方面的优势。在合成实例和41个公开的TSPLIB/CVRPLIB基准上,LLM-HCJG始终获得较低的最优性差距,包括在29个TSPLIB实例中的28个以及全部12个CVRPLIB实例上取得最佳或并列最佳结果。消融和结构分析进一步表明,这些收益与跨组件兼容性和对齐相关,而非孤立的组件重组。这些结果支持在有限样本、适度成本训练下,在所评估的路由设置内实现有效的跨实例迁移。
cs.NE / 75 / 2609.02387
Semantics-Guided Automatic Tensorization for Multiobjective Evolutionary Algorithms: A Multi-Agent Framework
面向多目标进化算法的语义引导自动张量化:一种多智能体框架
large language model
大语言模型相关
Abstract
Multiobjective evolutionary algorithms (MOEAs) naturally expose population-level parallelism, but many mature implementations encode their computation in sequential program structures designed for central processing units. Exploiting modern tensor computing platforms therefore requires more than direct code translation: the implementation must be restructured without changing the defining optimization mechanism of the underlying MOEA. We formulate automatic tensorization for MOEAs as semantics-guided computational restructuring and develop Evolutionary Code Conversion (EvoCoCo), a multi-agent framework that realizes this formulation. EvoCoCo reconstructs algorithm-specific states, dependencies, operators, and update logic into a structured semantic representation and organizes them through a shared tensorization blueprint. Specialized transformation branches then explore alternative tensor realizations, while execution feedback guides validation, repair, and candidate selection. Experiments on a benchmark of 48 MOEAs evaluate migration reliability, optimization fidelity, and computational scalability. Under matched large language model backends, EvoCoCo attains higher migration reliability than direct one-shot translation. Across the benchmark suites, 88.2% of valid comparisons satisfy the predefined optimization-fidelity criterion. The tensorized implementations also exhibit increasing acceleration on graphics processing units as population size or decision dimension grows, with median measured speedups ranging from $22.6\times$ under population scaling to $80.2\times$ under decision-dimension scaling. External-source and ablation studies further assess transfer beyond the main benchmark and the roles of the major framework components.
Chinese Translation
多目标进化算法(MOEAs)天然具有种群级并行性,但许多成熟实现将其计算编码在面向中央处理单元设计的顺序程序结构中。因此,利用现代张量计算平台所需要的不仅仅是直接的代码转换:实现必须在不改变底层MOEA的定义性优化机制的前提下进行重构。我们将MOEA的自动张量化形式化为语义引导的计算重构,并开发了进化代码转换(EvoCoCo),一个实现该形式化的多智能体框架。EvoCoCo将算法特定的状态、依赖、算子和更新逻辑重构为结构化语义表示,并通过共享的张量化蓝图组织它们。专门的变换分支随后探索替代的张量实现,同时执行反馈指导验证、修复和候选选择。在包含48个MOEA的基准上的实验评估了迁移可靠性、优化保真度和计算可扩展性。在匹配的大语言模型后端下,EvoCoCo比直接一次性翻译获得了更高的迁移可靠性。在基准套件中,88.2%的有效比较满足预定义的优化保真度标准。张量化实现还在图形处理单元上表现出随着种群规模或决策维度增长而增加的加速效果,其中位数实测加速比范围从种群规模扩展下的 $22.6\times$ 到决策维度扩展下的 $80.2\times$。外部来源与消融研究进一步评估了主基准之外的迁移性以及主要框架组件的作用。
cs.AI / 76 / 2609.01779
Agents That Model Agents: Five Principles Toward a Theory of Mind for 6G Networks
对其他智能体进行建模的智能体:迈向6G网络心智理论的五项原则
large language model
大语言模型相关
Abstract
Future 6G networks will rely on Large Language Model (LLM) agents to manage the Radio Access Network (RAN). However, current architectures assume inter-agent messages convey objective facts. A message is instead a \emph{trace} of the sender's reasoning: it carries a subjective conclusion, so a syntactically valid report can propagate an AI hallucination and trigger a cascading outage invisible to protocol validation. Reading such a trace requires a Theory of Mind (ToM)---before acting, the receiver must model what the peer believes, and what a peer in that position should have believed. Modeling these interactions as cognitive channels on a cellular sheaf, we obtain a unified framework for resilient multi-agent systems, from which five design principles emerge: (i) a message is evidence of the sender's hidden reasoning; (ii) trust is a continuous cognitive Signal-to-Noise Ratio (SNR)---asserted precision over deviation from the modeled peer belief; (iii) network-wide consistency and resistance to hallucination contagion are computable via the sheaf's Laplacian; (iv) peer-modeling must halt at exactly two levels to conserve compute and survive mutual information decay; and (v) credible capacity is bounded by operational goal alignment, not link bandwidth. A signaling-storm study on locally deployed 1B-parameter telecom language models validates it: cognitive SNR isolates a hallucinating peer that three of its four neighbors agree with, where a divergence gate ranks every wrong peer above the right one; only depth two ToM recovers the correct action; and the spectral gap decides whether a topology reaches consistency inside the near-real-time budget.
Chinese Translation
未来6G网络将依赖大型语言模型(LLM)智能体来管理无线接入网(RAN)。然而,当前架构假设智能体间消息传递的是客观事实。相反,一条消息是发送者推理的痕迹:它承载主观结论,因此语法上有效的报告可能传播AI幻觉,并引发对协议验证不可见的级联中断。读取此类痕迹需要心智理论(ToM)——在行动之前,接收方必须建模对等方相信什么,以及处于该位置的对等方本应相信什么。将这些交互建模为胞腔层上的认知信道,我们得到了一个构建韧性多智能体系统的统一框架,并从中得出五项设计原则:(i)消息是发送者隐藏推理的证据;(ii)信任是一个连续的认知信噪比(SNR)——即相对于建模对等方信念之偏差的断言精度;(iii)全网一致性和对幻觉传染的抵抗力可通过该胞腔层的拉普拉斯算子计算;(iv)对等方建模必须恰好停在两个层级,以节省计算资源,并在互信息衰减中存活;(v)可信容量受操作目标一致性约束,而非链路带宽。对本地部署的1B参数电信语言模型进行的信令风暴研究验证了这一点:认知SNR隔离出一个幻觉对等方,其四个邻居中有三个同意它;而散度门将每个错误对等方都排在正确对等方之上;只有深度为二的ToM能恢复正确动作;谱间隙则决定拓扑能否在近实时预算内达成一致性。
cs.AI / 77 / 2609.02252
DiffuSearch: How Hybrid Trajectory Planning Benefits from Aligned Objectives in Diffusion and Action Space
DiffuSearch:混合轨迹规划如何在扩散与动作空间中受益于对齐目标
diffusion
扩散模型相关
Abstract
In trajectory planning for autonomous driving, hybrid planning architectures are often realized as a collection of disparate modules, each with its own objectives. This lack of a unifying principle can lead to inconsistencies between the initial and refined trajectory, resulting in suboptimal behavior. We address this by introducing DiffuSearch, a novel hybrid planner that uses a unified set of objectives across generation and refinement. Our model encourages all components to follow the same shared driving goals: collision avoidance, drivable area compliance, comfort, and progress. DiffuSearch employs a two-stage architecture. First, a guided diffusion model generates a scene-consistent, joint trajectory prediction, using our driving objectives as differentiable guidance functions to implicitly steer the denoising process. Second, a Monte Carlo Tree Search (MCTS) in a discretized action space performs an explicit, local refinement of this proposal, leveraging the same driving objectives as its reward function. This synergistic design leverages the diffusion model's strength in finding scene-consistent solutions combined with the explainable, constraint-aware refinement of MCTS. Experiments on nuPlan and interPlan reactive closed-loop benchmarks demonstrate that DiffuSearch achieves strong and often state-of-the-art performance, substantially reducing collisions and improving comfort, particularly in complex, interactive scenarios. Our ablation studies indicate that MCTS refinement is the main mechanism behind the gains, while sharing objectives between implicit guidance and explicit search provides further consistent improvements.
Chinese Translation
在自动驾驶的轨迹规划中,混合规划架构通常被实现为一组各自具有独立目标的异构模块。这种缺乏统一原则的设计可能导致初始轨迹与精化轨迹之间的不一致,进而产生次优行为。为解决这一问题,我们引入了 DiffuSearch,一种新颖的混合规划器,它在生成与精化阶段使用统一的目标集合。我们的模型鼓励所有组件遵循相同的共享驾驶目标:碰撞避免、可行驶区域合规、舒适性和进度。DiffuSearch 采用两阶段架构。首先,一个引导式扩散模型生成与场景一致的联合轨迹预测,将我们的驾驶目标作为可微引导函数,以隐式地引导去噪过程。其次,一个在离散动作空间中执行的蒙特卡洛树搜索(MCTS)对该提议进行显式的局部精化,并将相同的驾驶目标作为其奖励函数。这种协同设计充分利用了扩散模型在寻找场景一致解方面的优势,同时结合了 MCTS 的可解释且受约束感知的精化能力。在 nuPlan 和 interPlan 反应式闭环基准上的实验表明,DiffuSearch 取得了强劲且常为最先进的性能,显著减少了碰撞并改善了舒适性,尤其是在复杂的交互场景中。我们的消融研究表明,MCTS 精化是性能提升的主要机制,而在隐式引导与显式搜索之间共享目标则提供了进一步的一致改进。
cs.AI / 78 / 2609.02270
CrashDiffuser: VLM-Guided Collision Intent Reasoning for Fine-Grained Safety-Critical Traffic Scenario Generation
CrashDiffuser:VLM引导的碰撞意图推理用于细粒度安全关键交通场景生成
diffusion
扩散模型相关
Abstract
Generating safety-critical scenarios is essential for evaluating autonomous driving systems. However, existing generators primarily focus on inducing collisions and offer limited control over where contact occurs on the target vehicle. In this paper, we study fine-grained safety-critical scenario generation, where success requires both a target collision and a specified head, rear, or side contact region. We propose CrashDiffuser, a closed-loop VLM-guided diffusion framework that decouples semantic collision reasoning from continuous trajectory synthesis through a hierarchical collision-intent interface derived from the requested target contact region. At initialization, the VLM extracts reusable scene-level context; at each replanning step, it predicts a structured action tuple describing speed change, turning behavior, and collision stage. This intent conditions a diffusion model to generate executable adversarial trajectories, while collision-guided sampling, candidate selection, and short-horizon replanning adapt generation to the target vehicle's evolving behavior. On WOMD-derived closed-loop scenarios, CrashDiffuser achieves a target-collision rate of 50.33% in a single attempt and 67.98% after three attempts, together with a contact-region control success rate of 40.05% and competitive trajectory naturalness. Component ablations further support the proposed design.
Chinese Translation
生成安全关键场景对于评估自动驾驶系统至关重要。然而,现有生成器主要侧重于诱导碰撞,对目标车辆上接触发生位置的控制能力有限。在本文中,我们研究细粒度安全关键场景生成,其中成功需要同时满足目标碰撞和指定的车头、车尾或侧面接触区域。我们提出CrashDiffuser,一种闭环的VLM引导扩散框架,通过从请求的目标接触区域导出的分层碰撞意图接口,将语义碰撞推理与连续轨迹合成解耦。在初始化时,VLM提取可复用的场景级上下文;在每个重新规划步骤中,它预测一个描述速度变化、转向行为和碰撞阶段的结构化动作元组。该意图对扩散模型进行条件化,使其生成可执行的对抗性轨迹;同时,碰撞引导的采样、候选选择和短视界重新规划使生成过程适应目标车辆不断变化的行为。在源自WOMD的闭环场景上,CrashDiffuser在单次尝试中实现了50.33%的目标碰撞率,在三次尝试后达到67.98%,同时接触区域控制成功率为40.05%,轨迹自然度具有竞争力。组件消融实验进一步支持了所提出的设计。
cs.CL / 79 / 2609.02343
SonicCaps: Large-Scale Diverse and Fine-Grained Captioning for Improved Audio-Retrieval
SonicCaps:大规模多样化且细粒度的字幕生成以改进音频检索
large language model
大语言模型相关
Abstract
Recent advances in audio-language modeling have been driven by large-scale audio captioning datasets. However, existing datasets remain limited by low semantic diversity, generic descriptions lacking acoustic details, and one-to-one audio-caption mappings that poorly reflect the inherent ambiguity of auditory perception. We introduce SonicCaps, a large-scale audio captioning dataset comprising ~15M captions paired with ~700k audio clips, generated using a multi-modal large language model (Qwen3-Omni) conditioned on both audio and text. To explicitly promote diversity, we generate around 24 captions per audio via structured prompt engineering and few- shot generation, spanning main descriptions, rephrased variants (verbosity, style) and semantic tags. Human evaluation shows that SonicCaps is rated significantly higher than existing captioning datasets, with fine-grained analyses indicating that our captions are perceived as more descriptive and precise, which strongly correlates with quality judgments. Finally, training CLAP models on SonicCaps with a multi-caption sampling strategy consistently improves audio retrieval and zero-shot classification, with stronger generalization across public and commercial benchmarks. We release both SonicCaps and two specialized CLAP models on hugging face: https://huggingface.co/datasets/Zineb/SonicCaps.
Chinese Translation
近期音频-语言建模的进展得益于大规模音频字幕数据集。然而,现有数据集仍受限于语义多样性不足、缺乏声学细节的通用描述,以及无法反映听觉感知固有歧义的一对一音频-字幕映射。我们提出了SonicCaps,一个大规模音频字幕数据集,包含约70万个音频片段及与之配对的约1500万条字幕,通过以音频和文本为条件的多模态大语言模型(Qwen3-Omni)生成。为明确促进多样性,我们通过结构化提示工程和少样本生成为每个音频生成约24条字幕,涵盖主描述、改写变体(详略程度、风格)和语义标签。人工评估表明,SonicCaps的评分显著高于现有字幕数据集,细粒度分析显示我们的字幕被认为更具描述性和精确性,这与其质量判断高度相关。最后,在SonicCaps上结合多字幕采样策略训练CLAP模型,持续改进了音频检索和零样本分类性能,并在公共及商业基准上展现出更强的泛化能力。我们在Hugging Face上发布SonicCaps和两个专门的CLAP模型:https://huggingface.co/datasets/Zineb/SonicCaps。
cs.SE / 80 / 2609.01736
Harness Engineering in LLM Tool Use via Agent-Native Reusable Tool Primitives
通过智能体原生可复用工具原语实现 LLM 工具使用的驾驭工程
large language model
大语言模型相关
Abstract
Large language models (LLMs) augmented with external tools have demonstrated remarkable capability in solving complex real-world tasks. However, existing approaches suffer from two key challenges: brittle multi-step and multi-turn reasoning caused by incompatible tool output types and API schemas, and performance degradation under large tool catalogues. To address these, we introduce \textbf{Tool Primitives}, a design that replaces rigid API schema-based invocation with natural language as the interface for tool calling, where each tool is wrapped with an LLM interface that handles schema resolution and execution internally, enabling natural inter-tool communication for nested and multi-turn tool calling. Building on Tool Primitives, we host \textbf{ToolFace}, a centralized repository of 25,519 functions from which LLMs dynamically retrieve only the relevant tools at inference time, eliminating the need to enumerate raw API schemas in context. To orchestrate Tool Primitives and ToolFace reliably in complex settings, we further propose \textbf{HEART}, a \textbf{H}arness \textbf{E}ngineering framework via \textbf{A}gent-native, \textbf{R}eusable \textbf{T}ool Primitives, comprising a Planner, Router, and Verifier that jointly support dynamic tool invocation planning, multi-step execution, and feedback-driven recovery. Experiments on five benchmarks demonstrate that HEART outperforms SFT-based models by $10\%$ on average and surpasses GPT-5.4, Claude-4.6-Sonnet, and Gemini-3.1-Pro by $6\%$ on average while reducing API cost by up to $85\%$. On 50 real-world tasks, HEART achieves $84\%$ task completion, $3.8\times$ the average of three frontier commercial models ($22\%$).
Chinese Translation
配备了外部工具的大型语言模型(LLM)在解决复杂的现实世界任务方面展现出了显著的能力。然而,现有方法面临两个关键挑战:由不兼容的工具输出类型和 API 模式导致的脆弱的多步骤及多轮推理,以及在大型工具目录下的性能下降。为了解决这些问题,我们引入了 \textbf{Tool Primitives}(工具原语),这是一种设计,用自然语言作为工具调用的接口取代了基于刚性 API 模式的调用,其中每个工具都包装有一个 LLM 接口,该接口在内部处理模式解析和执行,从而实现用于嵌套和多轮工具调用的自然工具间通信。在 Tool Primitives 的基础上,我们建立了 \textbf{ToolFace},这是一个包含 25,519 个函数的集中式仓库,LLM 在推理时仅动态检索相关工具,从而无需在上下文中枚举原始 API 模式。为了在复杂环境中可靠地编排 Tool Primitives 和 ToolFace,我们进一步提出了 \textbf{HEART},一个通过 \textbf{A}gent-native(智能体原生)、\textbf{R}eusable(可复用)\textbf{T}ool Primitives(工具原语)实现的 \textbf{H}arness \textbf{E}ngineering(驾驭工程)框架,包含 Planner(规划器)、Router(路由器)和 Verifier(验证器),它们共同支持动态工具调用规划、多步骤执行以及基于反馈的恢复。在五个基准上的实验表明,HEART 平均优于基于 SFT 的模型 $10\%$,并平均超过 GPT-5.4、Claude-4.6-Sonnet 和 Gemini-3.1-Pro $6\%$,同时将 API 成本降低高达 $85\%$。在 50 个现实世界任务中,HEART 实现了 $84\%$ 的任务完成率,是三个前沿商业模型平均值的 $3.8$ 倍($22\%$)。
cs.SE / 81 / 2609.02624
Automated Vulnerability Injection in Smart Contracts Using Large Language Models
使用大型语言模型在智能合约中自动注入漏洞
large language model
大语言模型相关
Abstract
Assessing vulnerability detection tools for smart contracts requires datasets with known ground truth, yet such datasets are scarce and difficult to build by hand. We propose an approach that uses Large Language Models (LLMs) to automatically inject vulnerabilities into Solidity smart contracts, and demonstrate it in a case study targeting 49 vulnerability types from OpenSCV. Injected contracts are validated through a multi-step pipeline checking compilation, execution, business logic, and the presence of the intended vulnerability. Applied to real-world contracts from SmartBugs, LLMs generate nearly 1,000 candidate variants; after deduplication and validation, 32 confirmed vulnerable contracts spanning 25 vulnerability types survive (a 16.58% survival rate). Surviving contracts concentrate in structurally simpler targets and vulnerability types with localized syntactic patterns. We report practical challenges including LLMs' non-determinism and the difficulty of preserving contract semantics. We then use the validated contracts to assess three static analyzers, revealing complementary and incomplete coverage profiles. Results show that LLM-based vulnerability injection is feasible, while exposing key limitations in scalability and diversity.
Chinese Translation
评估智能合约漏洞检测工具需要带有已知基准真值的数据集,然而此类数据集稀少且难以手工构建。我们提出一种方法,使用大型语言模型(LLMs)自动向Solidity智能合约中注入漏洞,并在一个针对来自OpenSCV的49种漏洞类型的案例研究中展示该方法。注入后的合约通过一个多步骤流水线进行验证,检查编译、执行、业务逻辑以及预期漏洞是否存在。将LLMs应用于来自SmartBugs的真实合约时,其生成了近1,000个候选变体;经过去重和验证后,32个经确认的含漏洞合约得以保留,涵盖25种漏洞类型(存活率为16.58%)。存活的合约集中在结构更简单的目标以及具有局部语法模式的漏洞类型上。我们报告了实际挑战,包括LLMs的非确定性和保持合约语义的困难。随后,我们使用经过验证的合约来评估三个静态分析器,揭示了互补且不完整的覆盖率概况。结果表明,基于LLM的漏洞注入是可行的,同时也暴露了在可扩展性和多样性方面的关键局限性。
cs.SE / 82 / 2609.02789
ShikumiMiner: Mining Recurring Implementation Patterns in AI Codebases
ShikumiMiner:挖掘AI代码库中的重复实现模式
large language model
大语言模型相关
Abstract
Large language models are paving the way towards innovation by understanding, analyzing, summarizing and generating content in the modern world. Currently there are thousands of LLM projects developed by engineers in open-source repositories. However, whether these LLM projects have underlying patterns or not remains a question. Exploring these underlying patterns will give new dimensions to the developers who aim to develop these LLM projects. In this paper, we propose ShikumiMiner, a static-analysis framework that combines Abstract Syntax Tree (AST) and Control Flow Graph (CFG) features to detect and compare recurring implementation patterns in C++ local LLM codebases. We analyze ten GitHub open-source repositories and classify functions into seven study-specific categories using a multi-label Random Forest model. Studying these patterns can provide useful insights for developers aiming to design LLM applications.
Chinese Translation
大型语言模型通过理解、分析、总结和生成现代世界中的内容,正在为创新铺平道路。目前,在开源代码库中,工程师们开发了数千个LLM项目。然而,这些LLM项目是否存在底层模式仍是一个问题。探索这些底层模式将为致力于开发这些LLM项目的开发者提供新的维度。在本文中,我们提出了ShikumiMiner,一种静态分析框架,它结合了抽象语法树(AST)和控制流图(CFG)特征,以检测和比较C++本地LLM代码库中重复出现的实现模式。我们分析了十个GitHub开源代码库,并使用多标签随机森林模型将函数划分为七个研究特定的类别。研究这些模式可以为致力于设计LLM应用程序的开发者提供有用的见解。
cs.CL / 83 / 2609.02262
From Detection to Characterization: A Large-Scale Study of Ragebait on Japanese X
从检测到特征刻画:日本X平台上愤怒诱导内容的大规模研究
large language model
大语言模型相关
Abstract
Ragebait refers to online content intentionally designed to provoke anger or outrage and thereby increase attention and engagement. However, reliable large-scale detection and systematic analysis of ragebait remain limited, hindering efforts to understand its prevalence, impact, and mitigation. This study aims to develop an effective ragebait detection framework and to clarify the characteristics of ragebait at scale, providing a basis for understanding and mitigating emotionally provocative content online. We constructed a labeled dataset with the assistance of a large language model (LLM) and trained several Japanese language models for ragebait detection. The resulting ensemble classifier was then applied to a large-scale dataset of Japanese-language posts on X. Our analysis shows that ragebait is more prevalent in politically and socially contentious topics, including politics, discrimination, public health, and interpersonal conflict. Ragebait posts also spread faster and receive more negative reactions than non-ragebait posts, particularly anger, fear, disgust, sadness, and surprise. These findings demonstrate the utility of the proposed detector and provide a large-scale characterization of ragebait in Japanese online discourse.
Chinese Translation
愤怒诱导内容是指故意设计用来激起愤怒或愤慨,从而增加关注和互动的在线内容。然而,对愤怒诱导内容的可靠大规模检测和系统分析仍然有限,这阻碍了理解其普遍程度、影响和缓解措施的努力。本研究旨在开发一个有效的愤怒诱导内容检测框架,并大规模地阐明愤怒诱导内容的特征,为理解和缓解在线情绪挑衅内容提供基础。我们在大型语言模型(LLM)的辅助下构建了一个带标注的数据集,并训练了多个日语语言模型用于愤怒诱导内容检测。随后,将得到的集成分类器应用于X平台上日语帖子的大规模数据集。我们的分析表明,愤怒诱导内容在政治和社会争议话题中更为普遍,包括政治、歧视、公共卫生和人际冲突。与非愤怒诱导内容的帖子相比,愤怒诱导内容的帖子传播得更快,并获得更多负面反应,尤其是愤怒、恐惧、厌恶、悲伤和惊讶。这些发现证明了所提出检测器的实用性,并为日语在线话语中的愤怒诱导内容提供了大规模特征刻画。
cs.LG / 84 / 2609.02016
Perceptually Regularized Diffusion Model for Image Super-Resolution
感知正则化扩散模型用于图像超分辨率
diffusion
扩散模型相关
Abstract
Image super-resolution, which aims to reconstruct high-resolution images from their low-resolution observations, is fundamental to medical imaging, remote sensing, surveillance, microscopy, and scientific visualization. Traditional model-based methods formulate super-resolution as an inverse problem with hand-crafted regularization priors. While interpretable and theoretically grounded, they rely on fixed assumptions and require computationally intensive iterative solvers. Deep learning methods offer data-driven flexibility by learning nonlinear mappings from low- to high-resolution images, among which diffusion models have achieved particularly impressive perceptual quality. However, the standard diffusion training objective is a pixel-domain noise-prediction loss that does not explicitly enforce perceptual fidelity, which can lead to oversmoothing and loss of fine image structure. To address these limitations, we propose a perceptually regularized diffusion framework that incorporates prior knowledge through perceptual-loss-based regularization, improving training convergence and encouraging the recovery of meaningful image features. Experiments on benchmark datasets demonstrate improved perceptual quality and competitive distortion metrics, highlighting the effectiveness of regularization for diffusion-based super resolution.
Chinese Translation
图像超分辨率旨在从其低分辨率观测中重建高分辨率图像,是医学成像、遥感、监控、显微镜和科学可视化等领域的基础任务。传统的基于模型的方法将超分辨率建模为一个逆问题,并使用手工设计的正则化先验。尽管这些方法具有可解释性和理论依据,但它们依赖固定的假设,并且需要计算密集的迭代求解器。深度学习方法通过学习从低分辨率图像到高分辨率图像的非线性映射提供了数据驱动的灵活性,其中扩散模型在感知质量方面取得了尤其令人印象深刻的表现。然而,标准的扩散训练目标是像素域中的噪声预测损失,并未显式地强制感知保真度,这可能导致过度平滑和精细图像结构的丢失。为解决这些局限性,我们提出了一种感知正则化扩散框架,通过基于感知损失的正则化引入先验知识,改善训练收敛性并促进对有意义图像特征的恢复。在基准数据集上的实验表明,该框架在感知质量和竞争性的失真指标上均有所提升,凸显了正则化对基于扩散的超分辨率方法的有效性。
cs.AI / 85 / 2609.02011
Seed-Anchored Budget-Bounded Graph Rendering for Question Answering on Industry-Standard Power-Grid Information and Exchange Models
面向行业标准电网信息与交换模型问答的种子锚定预算受限图渲染
large language model
大语言模型相关
Abstract
Large language model question answering over power-grid models must respect a fixed context budget. We introduce seed-anchored graph rendering, a deterministic method that prioritizes query-local graph evidence without adding method-specific tuned or learned parameters beyond the shared hop bound and context budget. The method provides a checkable condition under which predefined seed-local answer-bearing render units are preserved in a greedy bounded-context prefix. We evaluate the approach on Common Information Model (CIM) network models exchanged through the Common Grid Model Exchange Standard (CGMES). On two budget-binding CGMES encodings, naive descriptions-first rendering retains local evidence for every single-hop item but only 0.12 and 0.00 of multi-hop items, whereas seed-anchored rendering retains all such evidence. On a preregistered fresh 100-item bank from the SmallGrid topology family, accuracy rises from 0.450 to 0.970 under a fixed 8,000-character context budget. Under a common retrieval and rendering pipeline, the standards-native seed-anchored graph matches or exceeds extracted graph representations produced by LightRAG, Microsoft GraphRAG, and HippoRAG, while avoiding LLM graph-construction tokens. The results are specific to the evaluated CIM/CGMES models, reader, and context budget; they concern budget-bounded retrieval rather than general question answering.
Chinese Translation
在电网模型上进行的大语言模型问答必须遵守固定的上下文预算。我们引入了种子锚定图渲染,这是一种确定性方法,它优先考虑查询局部的图证据,并且除了共享的跳数界限和上下文预算之外,不添加方法特定的调优或学习参数。该方法提供了一种可检查的条件,在该条件下,预定义的种子局部且携带答案的渲染单元在贪心的有界上下文前缀中被保留。我们在通过通用电网模型交换标准(CGMES)交换的公共信息模型(CIM)网络模型上评估了该方法。在两种预算受限的CGMES编码上,朴素的描述优先渲染为每个单跳项保留了局部证据,但仅保留了多跳项的0.12和0.00,而种子锚定渲染保留了所有此类证据。在来自SmallGrid拓扑族的预注册的全新100项题库上,在固定的8,000字符上下文预算下,准确率从0.450上升到0.970。在通用检索与渲染流程下,标准原生的种子锚定图匹配或超过LightRAG、Microsoft GraphRAG和HippoRAG产生的提取图表示,同时避免了LLM图构建词元。这些结果特定于所评估的CIM/CGMES模型、阅读器和上下文预算;它们涉及的是预算受限的检索,而非通用问答。
cs.LG / 86 / 2609.02677
Eliciting ESG Preferences for Reinforcement Learning-Based Portfolio Optimization
为基于强化学习的投资组合优化引出ESG偏好
large language model
大语言模型相关
Abstract
Modern portfolio management increasingly demands a balance between traditional risk-adjusted returns and strict Environmental, Social, and Governance (ESG) mandates. Current Reinforcement Learning (RL) approaches typically optimize for a single ESG provider, neglecting the significant divergence in rating methodologies across the industry and the unintuitive nature of manually weighting conflicting objectives. This paper addresses these limitations by formulating ESG-aware portfolio optimization as a Multi-Objective Reinforcement Learning (MORL) problem that simultaneously incorporates ratings from three distinct ESG agencies. To bridge the gap between high-dimensional algorithmic trade-offs and human decision-making, we integrate a Preference Elicitation framework using Gaussian Processes. This system enables practitioners to infer their latent utility functions through intuitive pairwise comparisons of candidate portfolios based on their Sharpe ratios and aggregate ESG scores. We systematically evaluate our framework by employing Large Language Model (LLM) personas to simulate Portfolio Managers operating under varied regional contexts. Empirical results using historical market data reveal that regional backgrounds fundamentally shift the derived preference weights. For instance, European-based personas tend to prioritize ESG alignment over financial returns, while Texas-based personas favor risk-adjusted performance. This work offers a highly adaptable framework that successfully aligns multi-objective algorithmic trading with diverse, real-world human sustainability preferences.
Chinese Translation
现代投资组合管理日益要求在传统风险调整收益与严格的环境、社会和治理(ESG) mandates 之间取得平衡。当前的强化学习(RL)方法通常针对单一ESG提供者进行优化,忽视了评级方法在行业内的显著分歧,以及手动加权相互冲突的目标这一不直观的特性。本文通过将ESG感知的投资组合优化建模为一个多目标强化学习(MORL)问题来应对这些限制,该问题同时纳入来自三家不同ESG机构的评级。为了弥合高维算法权衡与人类决策之间的差距,我们整合了一个使用高斯过程的偏好引出框架。该系统使从业者能够通过对候选投资组合基于其夏普比率和总体ESG分数的直观成对比较,来推断其潜在效用函数。我们通过使用大语言模型(LLM)人物角色来模拟在不同区域背景下运作的投资组合经理,从而系统性地评估我们的框架。使用历史市场数据的实证结果表明,区域背景从根本上改变了所导出的偏好权重。例如,基于欧洲的人物角色倾向于优先考虑ESG一致性而非财务回报,而基于德克萨斯州的人物角色则偏好风险调整后的表现。这项工作提供了一个高度自适应的框架,成功地将多目标算法交易与多样化的、真实世界中的人类可持续性偏好对齐。
人工智能 (cs.AI)
69
cs.AI / 1 / 2609.01741
When Can a Machine Trust a Statute? A Survival Certificate for Machine-Extracted Legal Logic
Abstract
Statutes are increasingly parsed by machines before people read them, and the parsers disagree: on Missouri's statutes, two independently written extractors diverge on numeric-threshold presence at a false-negative rate of 0.43. We ask what formal logic survives such noise. We build a passive survival certificate for the Duquenne-Guigues implication basis of machine-extracted statutory contexts: per-attribute inter-extractor disagreement is measured, replayed against the basis in 1,000 Monte Carlo trials, and an implication is certified only when a one-sided Wilson 95% lower bound on survival reaches 0.95; every certified implication carries premise spans and a minimal counterexample. On 29,365 Missouri sections and 502 Indian central-Act sections, the preregistered held-out gate passes (10 statute families across 7 Titles exact; 16 across 11 with 5% tolerance), yet under one globally deployed error model 93.2% of held-out chapters fall below the informativeness floor, and a 2x2 factorial assigns that to calibration-rate transfer, not selection. The certificate is usable but fragile: deploy it per-chapter-calibrated or error-tolerant. Code, data products, and the audit trail, including one retracted claim, are released.
cs.AI / 2 / 2609.01814
When Does Information Sharing Improve Decentralized Discovery? Aggregation, Independent Rescue, and Equilibrium Selection
Abstract
Information sharing can improve a pooled estimate while eliminating independent rescue actions. This paper separates those effects in exact finite discovery models. A centralized action-budget profile shows that equal one-person accuracy can coexist with different portfolio values. Under a registered incremental-sharing protocol, a sharing step improves discovery exactly when pooled residual error contracts faster than an independent rescue attempt. Exact bounded registries exhibit compression, aggregation, neutral curves, and a bounded zero mixed class. In a two-agent Bayesian game with a hidden mixture of common and independent signal sources, the registered selected equilibrium yields a strict positive sharing interval at signal accuracy 3/5, while alternative equilibria show that the result is selection-dependent rather than universal. The models are synthetic and finite; no human or organizational data are used.
cs.AI / 3 / 2609.01815
Induction and Inquiry via Probabilistic Reasoning over Language and Code
Abstract
How humans grow and maintain abstract knowledge from the sparse, streaming noisy data of experience is a longstanding challenge in cognitive science. Any computational account must satisfy at least three desiderata: It must be (1) data-efficient and compute-efficient, (2) capture gradations of uncertainty to support intelligent inquiry and information gathering, and (3) be flexible enough to mentally represent the endless range of concepts people can learn and think about. Here we introduce a computational model that captures these three properties, by encoding symbolic knowledge as mental programs that combine natural language with source code, and sequentially inferring mental programs using LLM-guided Bayesian learning algorithms. Across a range of behavioral studies this model successfully reproduces quantitative signatures of human inductive learning and active inquiry, such as anchoring, garden-pathing, and other effects. In contrast, pure LLMs and classic Bayesian models either fail at the underlying task, or do not reproduce human behavior, or succeed only at exorbitant computational cost. These results suggest that one way humans continually grow their knowledge is by mentally representing many hypotheses spanning language-like and program-like representations, then revising those hypotheses to approximate Bayesian updates, while a bottom-up neural mechanism (an LLM) makes inference both tractable and learnable.
cs.AI / 4 / 2609.01834
Architecting Conversational Data Systems for Stateless LLM APIs: The Hydration Proxy Pattern
Abstract
As enterprise platforms transition to conversational reasoning interfaces, the stateless nature of LLM APIs creates an architectural gap. While statelessness enables horizontal scalability for AI providers, it forces client applications to manage the entire burden of conversational state and semantic memory. The work identifies the Hydration Proxy Pattern, an architecture that decouples session persistence from the reasoning engine. The framework ensures platform sovereignty over conversational data while enabling secure, multi-stage semantic grounding. We further propose the Context Stabilization Mandate to resolve the tradeoff between sovereign state management and KV caching.
cs.AI / 5 / 2609.01849
SSAKG 2.0: An Open-Source Package for Structural Associative Sequence Memory and Context-Based Retrieval
Abstract
This article presents SSAKG 2.0, an open-source software package for constructing and operating Structural Sequential Associative Knowledge Graphs (SSAKGs). An SSAKG represents objects as graph vertices and ordered sequences as structural patterns of graph connections. The resulting sparse graph is used as an associative memory in which complete sequences can be reconstructed from a partial, unordered context. Version 2.0 introduces new algorithms that exploit individual bits of computer memory to efficiently search graph connections. The package is implemented in Python, while performance-critical graph operations are implemented in C and exposed through a Python interface. This hybrid implementation provides a flexible high-level programming environment while reducing the memory and computational overhead associated with large sparse graphs. The algorithms were evaluated using randomly generated numerical sequences, sequences derived from sentences in the NLTK corpus, and mRNA sequences. The experiments demonstrate the ability of the package to store and reconstruct sequences from partial contexts and provide a basis for evaluating the effects of graph density, sequence length, and memory size on retrieval performance. SSAKG 2.0 is distributed under the Apache 2.0 open-source license. The package includes documentation and reproducible examples and is publicly available through GitHub and the Python Package Index (PyPI).
cs.AI / 6 / 2609.01852
The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents
Abstract
Persistent memory supports personalized agents, but a stale stored fact can override current authoritative evidence without warning. We study when this harm begins as model capability changes. We evaluate a frozen, closed-set, action-scored benchmark with 2 suites that represent 2 different meanings of "no memory" (a Benefit suite, unsolvable without the stored fact, and a Safety suite, in which an authoritative tool always holds the correct value), on a same-family model-size series (Qwen3 0.6/1.7/4/8B). The Memory Trust Gap reflects over-trust rather than confusion. In the Benefit suite, models answer with the stale value 0.92-1.00 of the time at every scale. In the Safety suite, harm below the no-memory baseline under the trap conditions ($Δ_{\mathrm{mem}}$) is capability-gated, with the larger models collapsing most once a stale note is made to look current. In a $2\times2\times2\times2$ factorial, which feature triggers over-trust depends on both the feature and model scale. Removing a label amplifies over-trust at every size, and a recency feature (stale dated newer) fools the larger models harder. Source authority is weak and scale-flat, and position changes from positive to negative across the Qwen3 model-size series. We confirm these scale interactions with direct cross-size contrast tests rather than overlapping per-model intervals. Mitigation is likewise capability-dependent: exposing metadata improves accuracy for the capable models, but only pre-resolving the conflict restores accuracy for the 2 smaller checkpoints. The same pattern appears on the capable models in an independent Llama-Instruct model-size series and on 2 external datasets (RGB, MisBench). A framing control finds no consistent advantage for the memory label: at the 3 smaller scales, models trust a stale document more than a stale memory; at 8B, the difference is not significant.
cs.AI / 7 / 2609.01861
Belief-Calibrated Optimization: An Explicit World Model for Agentic Optimization
Abstract
The performance of an LLM agent depends on the scaffold around a frozen model. A common way to improve that scaffold is to use a coding agent as an optimizer: it reads current scores and traces and iteratively edits the source, producing a new candidate each round. Each edit is chosen according to a belief about how the environment will respond: what went wrong, and which change should help. That belief is typically implicit. It lives in the coding agent's reasoning on the current call, or remains latent in its parameters, rather than as something written down. Later calls therefore see scores and traces, but they do not use that belief. We introduce Belief-Calibrated Optimization (BCO), a method that writes that belief down as a persistent in-context document and continually revises that document as new candidates are evaluated. The resulting document is a world model: the current account of how the environment responds to edits. Added to an otherwise standard loop, BCO reaches a higher train passrate than a matched control that lacks only the world model, on five benchmarks spanning memory QA, tool-use QA, code-as-action app agents, and terminal agents. The gap remains on every held-out split, which is not used to select the candidate. After a target-model swap, in which the frozen model is replaced and the scaffold is not, the selected BCO scaffold leads on the tasks we test, except where context-window overruns leave it unfinished. An offline ablation then asks whether that gap comes from what the world model says. A fresh predictor given the accumulated document forecasts how the environment will respond more accurately than predictors given either no document or a same-form copy whose content has been falsified. The comparison indicates that the document carries reusable information in its content, not only in its form.
cs.AI / 8 / 2609.01873
Epistemic Sybil Resistance: Multiplying AI Agents Without Multiplying Evidence
Abstract
Multi-agent AI systems improve inference by spawning agents and synthesizing reports. But another agent is not another observation: apparently independent reports may descend from the same evidence, and genuinely independent evidence can produce nearly identical reports. We formalize this as an epistemic Sybil problem. A report Z is an epistemic Sybil extension relative to reports R when I(Theta; Z | R) = 0. No report-only aggregator can generally distinguish replication from independent corroboration: identical reports can warrant different posteriors under unobserved ancestry. A Gaussian shared-root model shows common ancestry does not imply complete redundancy. Repeated extraction adds information toward a source-level ceiling, and correlated extraction errors, which a shared base model can induce among independent agents, lower that ceiling further. We test these predictions with more than 20,000 controlled LLM-agent report and extraction calls on synthetic evidentiary documents. Holding one evidence root fixed while report multiplicity rises from 1 to 32 collapses naive posterior coverage from 0.940 to 0.263. Holding report count fixed while evidence-root multiplicity rises from 1 to 16 closes the gap, and the aggregators are statistically indistinguishable at k = 16. The agent's replicate extraction errors are correlated (gamma_cal = 0.719, estimated out of sample), and a correlated-extraction aggregator restores calibration accordingly. A controlled manipulation isolates representation similarity from evidential ancestry. It changes a report-space deduplication mechanism's mean inferred cluster count by 1.425 (95% CI [1.363, 1.485]), whereas a fourfold change in true ancestry changes it by only 0.040 ([-0.045, 0.120]). Collective inference should therefore track evidential ancestry and dependence, not agent or report multiplicity or similarity.
cs.AI / 9 / 2609.01909
The Ceiling Is in the Channel: Auditing Learner Gaps and Measurement Frontiers in Clinical Prediction
Abstract
Clinical prediction can saturate for two different reasons: a fitted learner may fail to extract available information, or the recorded variables may impose a population frontier. We separate these quantities through the \emph{learner gap} and the \emph{measurement-channel ceiling}. Optimal balanced accuracy is characterized by total-variation separation, yielding architecture invariance, a sharp partial-identification result under replacement contamination, a cross-fitted ceiling estimator, and exact conditions for multimodal decision improvement. We add two finite-sample diagnostics, namely a label-permutation optimism floor and an underfit curve, and validate the audit on three real cohorts: UCI readmission ($n=99{,}343$), BRFSS diabetes ($n=253{,}680$), and NHANES HbA1c ($n=10{,}219$). Well-tuned gradient boosting nearly reaches the estimated frontier in UCI and BRFSS, whereas deliberately or practically deficient learners retain large gaps. NHANES yields a null difference between questionnaire and measured marginal frontiers but a significant joint complementarity gain, refining the simplistic claim that an objective modality must dominate. Across all cohorts, modest AUROC gains coexist with substantially larger Bayes decision-flip rates, and several architectures estimate similar frontiers while their achieved balanced accuracy differs sharply. A PRISMA-guided synthesis of 104 clinical tasks then shows that the same channel-level regularities recur across more than 18 disease categories: a broad but non-universal structured-clinical region, diminishing same-channel gains across model families, and higher performance when measurement channels change. The framework converts saturation from an empirical observation into an auditable decision: improve the learner when headroom remains; improve measurement when it does not.
cs.AI / 10 / 2609.01924
Looped Transformers under the Jacobian Lens: Does the Global Workspace Survive Recurrence?
Abstract
Recent work identifies a mid-depth band of verbalisable, causally potent representations in a standard feedforward transformer --- a functional analogue of a global workspace. Whether the same workspace functionality emerges when depth is implemented through recurrence rather than a stack of distinct layers remains unknown. Looped and depth-recurrent transformers provide a direct test of this question because they reuse the same weights across depth. We extend the Jacobian lens to iterated architectures using a virtual-unrolling adapter. We apply the full workspace suite --- lens fitting, readout, and eleven causal experiment families --- to Ouro-2.6B (48 layers looped 4 times, deeply supervised) and Huginn-0125 (a 4-layer core recurred 16 times, trained for latent reasoning), using Qwen3.6-27B (64 untied layers) as the standard baseline. We find that a workspace forms in the iterated part of each architecture, but that recurrence changes how it can be accessed. Ouro reconstructs workspace content in every loop, and linear transport cannot carry that content across loop boundaries; writes and ablations must therefore span every remaining loop. Huginn carries content forward across all sixteen recurrences, while reads, writes, and ablations act only within a sliding window of roughly two recurrences. Whether newly injected content can be verbalised tracks explicit per-iteration supervision; whether existing content can be steered does not.
cs.AI / 11 / 2609.01962
Post-Training Ternarization of Qwen3-4B Capability, Effective Bit Budget, Storage Compression, and Deployment
Abstract
Ultra-low-bit language models can reduce storage and memory bandwidth, but a nominal "1.58-bit" label does not fully describe the stored representation, retained capability, or runtime behavior. We study an end-to-end post-training conversion of Qwen, an instruction-tuned 4B-parameter model, using KOTMS rotation, E2M-ATQ ternarization, and GPTQ-style error compensation from TWLA. The experiment is weight-only: activations remain at 16-bit precision, so ILA-AMP is omitted. We evaluate effective bit accounting, task capability retention, perplexity, calibration sensitivity, checkpoint composition, and deployment behavior. The final conversion uses 1.641 effective bits per weight for quantized linear weights, with 81.62% of model parameters targeted. Across ten scored capability comparisons, accuracy falls from 64.5% to 54.7%. Degradation is uneven: BoolQ retains 84.6% chance-corrected teacher performance, while ARC-Challenge retains 43.8%. Perplexity rises from 13.639 to 18.748 on WikiText-2, 24.700 to 31.992 on PTB, and 19.831 to 28.966 on C4. A subsequent packing run preserves the ternary planes and scales, reducing reported model size from 8.29 GiB to 3.96 GiB with essentially unchanged perplexity. A separate third-party packing attempt was lossy and is excluded from the primary artifact claim. The packed artifact has not been benchmarked end-to-end for task accuracy or generation throughput. A preliminary Triton GEMV microbenchmark is 4.6x slower than FP16 cuBLAS on one tested shape. We therefore do not claim that compression alone yields faster inference.
cs.AI / 12 / 2609.01985
When Agents Implement Systems: A Case Study in Defects, Detection, and Evaluation Rigor
Abstract
As LLM coding agents increasingly perform end-to-end engineering work, we lack empirical characterization of how they behave on systems-level requirements: schema design, async orchestration, configuration correctness, and retrieval-filtering trade-offs. We present a case study of one such agent implementing a multi-component data system against a detailed pre-existing specification. Storage technologies, schema, entity-resolution algorithm, and retrieval-filtering strategy were fixed in advance; the agent autonomy was in the implementation, in diagnosing and fixing defects it introduced, and in interaction-design choices left open. Over a single session, we catalog five such defects, categorized by constraint violated and detection method. We further evaluate, on the public HotpotQA benchmark, the one retrieval trade-off specified in that architecture: restricting candidates to a graph-identified entity set before ranking versus unfiltered search. We substitute the benchmark gold evidence labels for entity identification, since we lacked LLM access to run that stage, and report standard recall rather than the benchmark own accuracy metrics. Across retrieval budgets from 1 to 10 and 100 questions against a pooled corpus of 2994 paragraphs, filtered recall reaches its ceiling by a budget of 3, expected once candidates are restricted to the gold paragraphs themselves, while unfiltered search recovers all required evidence only 69 percent of the time even at a budget of 10, a gap that holds at every budget tested, with sign test p less than 0.0001. We close with a discussion of where the agent autonomy succeeded versus required correction, including one instance where a claimed performance fix was never re-measured on the regression that motivated it.
cs.AI / 13 / 2609.01992
ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations
Abstract
Agent evaluations face two distinct evidentiary questions: whether a reported claim is recomputable from retained evidence (sufficiency), and whether the retained records cover the committed experiment set (coverage). Generic logs and hash-linked transcripts answer neither reliably. We introduce ClaimReceipt, a claim-relative receipt specification and selective verifier that binds typed transaction evidence to a signed experiment manifest and returns PASS, INVALID, or INCONCLUSIVE per claim. We freeze the specification before implementation (SHA-256 18d109...b81). On 1,392 historical buyer--seller records, a CR-2 verifier reproduces all five manually labeled audit verdicts, exactly replays 600 deterministic and 792 post-generation records, makes every one of 13 declared field groups non-redundant under tested ablations, and returns the expected result on 11/11 semantic faults with 0/8 false positives. We then run a separate prospective CR-3 epoch: 30 assignments are committed before inference, terminal receipts are signed and chained, and private evidence is encrypted for an auditor. Complete evidence yields coverage and accounting PASS; withholding one terminal receipt returns INCONCLUSIVE_COVERAGE, while withholding all private openings preserves coverage and protocol verification but makes economic claims inconclusive, exactly matching a preregistered prediction. Receipt instrumentation adds 0.021% of model-inference time and 9.9 KB per transaction. A specification-legibility probe indicates that our own frozen specification is not yet unambiguous to an independent reader. Claim verification therefore requires both claim-sufficient evidence and a committed universe against which omissions become visible.
cs.AI / 14 / 2609.02029
HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models
Abstract
Long-context inference retains a growing key--value (KV) cache during decoding, which consumes substantial GPU memory and can reduce generation throughput. This bottleneck remains in hybrid language models because their residual global-attention layers can dominate context-dependent cache demand. We study how to allocate this state under an aggregate KV-residency budget. We introduce HeadWiseKV, a training-free framework that compresses the residual global KV caches of hybrid language models while preserving their native local, recurrent, and linear paths. It assigns each physical KV head a static, multilevel history window, making cache demand predictable before serving. We formulate this allocation as a restricted operational rate--distortion problem and propose SeqCalib as the core policy-generation algorithm in HeadWiseKV. SeqCalib processes layers in execution order and conditions each decision on the lower-layer policy used at deployment, thereby accounting for interactions across depth. A grouped-cache runtime materializes the selected policy as actual per-head KV residency rather than a mask over a full cache. We evaluate downstream quality across four hybrid long-context models and study physical residency and serving behavior on Qwen3.6-27B. HeadWiseKV retains near-Full-KV RULER and LoCoMo quality across the evaluated models. In the fixed-model systems study, it reduces sampled peak device memory by 8.59\% at a 112K context length and extends the largest verified successful context from 114K to 161K.
cs.AI / 15 / 2609.02057
Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision
Abstract
Reliable web-agent monitoring is difficult when model-internal uncertainty signals such as token logits are unavailable. In this work, we study prefix-level risk prediction for web agents using observable trajectory signals: given an evolving prefix, estimate whether the current execution remains on track or is tending toward failure. We derive two observable trajectory representations: Macro features summarize cross-step agent--environment behavior and feedback, while Micro features measure the consistency of intention, action, and anticipated state change through repeated black-box queries. Instead of inheriting the final result label, we label the first critical error that remains uncorrected in the observed continuation and is associated with final failure as a key-step boundary, preserving valid early prefixes of failed trajectories as on track. Across WebArena-Lite and Online Mind2Web web agent benchmarks with five open- and closed-source backbones, observable trajectory signals are competitive with internal-signal baselines. The resulting predictors also support early intervention under fixed false-cut budgets and transfer across held-out website categories. These findings show that observable trajectory signals support valuable risk prediction abilities.
cs.AI / 16 / 2609.02060
MineTRACE: An Evidence-Grounded Interactive Reasoning System for Mineral Prospectivity
Abstract
Mineral exploration requires integrating heterogeneous geochemical, geophysical, and geological evidence, yet existing prospectivity systems often provide only opaque scores or heatmaps. We present MineTRACE, a web-based system for evidence-grounded exploration of eight commodities: Cu, Au, Ni, W, Sn, Co, Ta, and Mn. Users can explore prospectivity maps, query locations or regions, inspect supporting evidence, and interact through natural language. A transparent expert tree, informed by geological knowledge and known deposits, combines multi-source evidence into interpretable prospectivity scores. For a new location, the conversational assistant retrieves the score and supporting evidence from the analysis pipeline and presents them in natural language. The scorer achieves spatial AUC values of up to 0.917 across different test scenarios, while end-to-end evaluation assesses query accuracy and response grounding. MineTRACE makes public geoscience data easier to access, interpret, and verify, supporting more efficient and transparent mineral exploration.
cs.AI / 17 / 2609.02067
ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction
Abstract
Scientific benchmarks are commonly built by domain experts who write tasks and cross-check one another's work, or who adapt existing material from textbooks, published papers, and online resources. These routes can produce strong evaluations, but they require substantial per-item labor. Language models can reduce this repeated work by proposing candidates quickly. The remaining problem is acceptance. We target scientific questions whose answers require computations with specialist software rather than unaided reasoning alone. A candidate is invalid if its script fails or returns a different answer, or trivial if a model answers it without the software. We present ToolGate, which treats every generated item as a proposal and keeps it only if three gates pass. First, an executable solution script must reproduce the proposed answer when run with the scientific software. Second, randomized no-tool screening rejects candidates that models can already solve from the prompt alone. Third, a tool-using agent must solve each survivor within a fixed time limit. We instantiate ToolGate in FEniCSx with 500 generation attempts. The local-verification gate retains 478 candidates. For final reporting, we rescreen this pool after generation: two randomized no-tool screens exclude 222 from the reported pool, and direct GPT-5.5 API calls at medium reasoning (the API default) exclude another 121. Of the remaining 135, a GPT-5.5 Codex CLI agent with access to FEniCSx solves 130; exact deduplication leaves 128 unique protocol survivors. ToolGate turns repeated answer checking and difficulty screening into an auditable process while leaving domain design and final review to experts.
cs.AI / 18 / 2609.02074
CHIME: Credit-Aware Hierarchical Memory Evolution for Long-Horizon Agentic Planning
Abstract
Planning is a central capability that enables agents to decompose complex long-horizon tasks into manageable steps. Test-time search and training-based methods improve planning but incur high inference costs or require expensive training data. Self-evolving memory instead accumulates reusable experience from agent interaction outcomes into an external memory bank, so planning capability keeps improving at inference time without parameter updates. However, existing self-evolving memory methods share an inherent credit assignment problem: they rely on final task outcomes as feedback, but such outcomes conflate plan quality with execution errors and environmental factors, so the accumulated planning experience is often biased and noisy. To address this problem, we propose Credit-Aware Hierarchical Memory Evolution (CHIME), a self-evolving memory framework that maintains a separate planning bank and execution bank and follows an attribute-before-memorize principle: CHIME first attributes each task outcome to the plan, the execution, both, or neither, and then updates only the corresponding memory bank. Extensive experiments on four long-horizon agent benchmarks show that CHIME consistently outperforms state-of-the-art training-based and self-evolving memory baselines. Further analyses reveal several interesting findings. For example, CHIME accumulates effective memory with far fewer items. In addition, the learned memory values faithfully reflect downstream utility: high-quality planning memories are more valuable than execution memories. Finally, the accumulated memory effectively transfers across backbone models. Code will be released at https://github.com/ATH-MaaS/Marco-DeepResearch.
cs.AI / 19 / 2609.02092
Beyond Outcome Gaps: Process-Aware Fairness Diagnosis for LLM-based Multi-Agent Decision Systems
Abstract
LLM-based multi-agent systems (MAS) are increasingly considered for high-stakes decision-making, yet outcome-based fairness audits can miss where risks arise within the decision trajectory. We present SCOPED-Hiring, a process-aware fairness diagnosis pipeline for LLM-based hiring MAS. SCOPED-Hiring constructs controlled resume variants, runs role-based hiring committees, logs over 311K structured decision trajectories, and converts trajectory fields into quantitative fairness signals organized by six diagnostic lenses: final outcome, counterfactual, process, pathway, dynamic, and design effects. SCOPED-Hiring reveals that balanced final hire rates can mask hidden trajectory unfairness in multi-agent decision trajectories: career gaps trigger suspicion, proxy cues shape qualification judgments, and identity cues lead to unequal investigation. Targeted repair guided by these diagnoses reduces total layered burden by 72.3% while shifting the hire rate by only 1.86 pp, showing that process diagnosis can guide effective repair. Project Page: https://scoped-hiring-project-page.vercel.app/
cs.AI / 20 / 2609.02094
MASkills: Continual Skills Optimization for Multi-Agent LLM Systems
Abstract
LLM-based multi-agent systems have shown strong performance on complex tasks, yet continual improvement from interaction experience remains challenging. Existing self-reflection methods build experience memories, but memories are mostly hard to invoke, refine, or scale, while agent skills offer a more actionable unit: structured procedural knowledge that specifies when to act, how to act, and which resources or tools to use. We introduce MASkills, a continual learning framework that optimizes multi-agent LLM systems through agent skills. MASkills presents a new agent-optimization pipeline that integrates skill-conditioned credit assignment, hierarchical credit aggregation, and momentum-smoothed optimization, enabling agent skill libraries to evolve through refinement, induction, consolidation, and pruning. Experiments on HotpotQA, LoCoMo, and GAIA demonstrate the effectiveness of MASkills across multiple agentic tasks. Our code is available at https://github.com/DaRL-GenAI/MASkills
cs.AI / 21 / 2609.02095
READY or Not: Reliable Enterprise Agent Deployment
Abstract
An AI agent can perform well on benchmarks and still be unsuitable for deployment. Existing AI-agent benchmarks measure whether an agent can complete realistic professional work, whereas enterprise deployment asks a different question: whether an agent can meet a required reliability level, under acceptable human oversight, and at tolerable cost. We introduce Reliable Enterprise Agent Deployment (READY), a framework for qualifying AI agents for deployment on enterprise workflows. READY preserves each workflow's own definition of successful execution while applying a common qualification procedure. Given an agent, a workflow, and a class of candidate oversight policies, READY measures the reliability and operating cost of the human-AI system, selects the minimum-cost policy that satisfies a specified reliability target, and statistically qualifies it on held-out cases. The resulting deployment profile characterizes the supported operating point: reliability, human-oversight burden, and cost. READY is implemented as an open testbed that decouples workflow specification, execution, evaluation, and qualification, and runs on existing agent-evaluation infrastructure. In an end-to-end clinical-audit case study spanning 16 agent systems and 750 cases, READY reveals differences hidden by autonomous performance: two systems separated by only 0.3 points in autonomous accuracy (72.8% vs. 72.5%) require 39.2% versus 29.6% human review, respectively, to qualify at the same 76% reliability target under the evaluated oversight policy. READY thus shifts enterprise agent evaluation from how well can the agent perform the work? to under what conditions, and at what cost, can it be reliably deployed? By making those conditions explicit and statistically testable, READY provides a basis for comparing agent systems, setting oversight requirements, and making evidence-based deployment decisions.
cs.AI / 22 / 2609.02129
Beyond Context Windows: Persistent Discovery Context for Data-Centric Agents
Abstract
Data-centric agents repeatedly perform a discovery step before planning or execution: identifying the data objects relevant to a task. Yet successful discovery outcomes are typically discarded rather than reused. We introduce persistent discovery context, a lightweight memory layer that stores prior intent-to-object mappings and reuses them to augment future retrieval. Across three structured data environments, persistent discovery context consistently improves retrieval quality over metadata-only search, remains effective with automatically generated memories, and exposes a reproducible interference failure mode. In lexically sparse domains, memory-only retrieval can even outperform metadata-based retrieval. These findings suggest that discovery outcomes constitute a useful form of reusable context for data-centric agents.
cs.AI / 23 / 2609.02133
EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision
Abstract
Empathetic response generation requires models to decide not only what to say, but also how to respond to the previous speaker's affective situation. We formulate this as response-side affective-orientation control and use multi-annotator emoji distributions as weak affective--attitudinal evidence, rather than as output symbols or gold labels, to induce a latent control space that operationally approximates listener stance. We construct EmojiDialogue, an utterance-level extension of EmpatheticDialogues with emoji votes and confidence scores, and propose EmoStance, which models source-side affective expression, predicts a soft response-side orientation from dialogue context and speaker roles, and steers a frozen instruction-tuned LLM through continuous prefix embeddings. In blind pairwise evaluation with 20 annotators and 800 judgments, EmoStance achieves a 62.2% decisive win rate, with the clearest gains in contextual specificity and perceived responsiveness, while remaining complementary to external-knowledge methods. Code, annotation metadata, and reconstruction scripts are available in our GitHub repository: https://github.com/18277390221/EmoStance.
cs.AI / 24 / 2609.02168
FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs
Abstract
Fragmented safety evaluation undermines the governance of dangerous AI capabilities. We present a modular framework that evaluates each model through three orthogonal pipelines---Knowledge ($K$), Defense ($D$), and Harm ($H$)---under a unified protocol, aggregating results into a standardized dangerous-capability profile $φ$. Pluggable modules supply scenario seeds, knowledge banks, hazard queries, and judge rubrics, while the core evaluation engine remains unchanged across domains; the CB evaluation is complemented by a cyber pilot demonstrating protocol transfer. Instantiating the framework with a chemical-biological (CB) module, we evaluate 12 commercial LLMs from four families. Our first contribution is a horizontal comparison of dangerous capability across models and model families: the three dimensions expose sharply divergent profiles---models with comparable knowledge differ in refusal resilience, and strong defenders do not generate less harmful content when they do comply---while family-level patterns further separate Claude, DeepSeek, and GPT models. The second is a temporal analysis of capability evolution: tracking $K$, $D$, and $H$ against model release dates reveals that dangerous capability has not monotonically declined; newer models deepen knowledge while only partially improving defense, showing that scaling and alignment progress do not uniformly translate into safety. Reliability is established via cross-judge consistency (bootstrap $ρ> 0.79$, 4 of 5 judges) and pipeline orthogonality ($K$--$D$--$H$ inter-correlations $ρ\in [0.32, 0.52]$).
cs.AI / 25 / 2609.02191
Examining the Vulnerability of Multi-Agent Medical Systems to Human Interventions for Clinical Reasoning
Abstract
Human interventions at fault points can alter the diagnostic accuracy of multi-agent medical systems. We defined fault points as moments in AI agent conversations, in which an agent's reasoning became most vulnerable to external influence. Using the MedQA dataset, this study analyzed simulated doctor-patient conversations to measure how interventions shifted reasoning and accuracy. Correct intervention methods showed an improvement in baseline diagnostic accuracy of up to 40%, while incorrect or bias-related interventions degraded performance by up to 6% and increased diagnostic drift and uncertainty. Beyond performance changes, our analysis revealed behavioral similarities between cognitive biases in simulated agent environments and real-world clinical practice. Examples included premature closure and susceptibility to misleading cues. Overall, these findings demonstrate that identifying and guiding fault points with human interventions may provide a mechanism for improving diagnostic robustness in multi-agent medical systems.
cs.AI / 26 / 2609.02217
SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams
Abstract
LLM agents increasingly self-improve by writing and reusing textual skills, kept either as one global document or as a flat pool of per-task entries, though most of the evidence comes from domains with structurally similar tasks. On long-horizon workloads where each task demands a different solution, the two forms fail in opposite ways: the document collapses into generic discipline, while the pool inflates and its entries stay bound to the instance that wrote them. We argue the missing unit of reuse is the solving procedure shared by a cluster of related tasks, and build SkillGLoW (Global-Local Weave) around it: the local skills a task writes from its own execution are aggregated into procedural families and compressed into de-instantiated global priors, while the instance detail they hold is regenerated per task rather than stored; a commit gate admits a prior only when real execution shows it does not degrade the deployed library. Across four benchmarks (mathematical reasoning, terminal automation, software repair, and embodied control) and three models, the priors gain 17.2 points (hard) over the no-skill baseline on average, with positive gains in all 12 continual-improvement runs, and 18.0 with local regeneration, while the library holds one prior per procedural family, 3.6x more compact than the per-task pool. Under the same protocol GLoW leads a published single-document optimizer on 15 of 21 cells. Unmodified, the library lifts success on unseen ALFWorld tasks from 73.9% to 83.9%, evidence that what transfers is procedure rather than task memory.
cs.AI / 27 / 2609.02231
PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment
Abstract
Interview assessment requires per-criterion judgments grounded in behavioral evidence, yet surging applicant volumes have made human-only evaluation costly and inconsistent, while existing AI approaches yield opaque scores without traceable rationale. We introduce PhoenixNest-Video, an evidence-grounded multimodal agent framework for automated video interview assessment. It builds a semantic video graph as structured working memory, performs rubric-conditioned retrieval with cross-modal verification across visual, audio, and textual streams, and produces per-criterion scores anchored to the candidate's materials. A Scorer trained via Rubrics-based Reinforcement Learning with dual rewards for rubric alignment and score-level differentiation internalizes the discriminative structure of multi-level rubrics. PhoenixNest-Video attains 91.50\% grade-level accuracy on VInterview-2025, outperforming substantially larger proprietary models. A compact, rubric-grounded agent therefore scores candidates in closer agreement with an expert panel than direct prompting of much larger models, and exposes the evidence behind each score for human review.
cs.AI / 28 / 2609.02236
PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks
Abstract
Group-based reinforcement learning (RL) has become an effective paradigm for LLM post-training, but in multi-turn agentic tasks with sparse terminal rewards, it often provides coarse credit for intermediate actions. To obtain more fine-grained credit assignment, recent work such as GiGPO introduces step-level advantages for intermediate actions. However, these step-level signals still rely on the final outcome of each individual trajectory. As a result, actions within failed trajectories can remain poorly differentiated, so effective actions can receive the same unfavorable credit as erroneous ones. In this work, we propose Potential-Guided Policy Optimization (PGPO) for multi-turn agentic tasks. PGPO estimates empirical state potentials from anchor-state-group return statistics within each rollout group. It then derives action advantages from potential differences between adjacent states, enabling cross-trajectory credit propagation. This provides finer-grained step-level credit assignment, especially within failed trajectories. Experiments on ALFWorld and WebShop show strong overall performance relative to recent group-based RL methods. Further analysis provides evidence that PGPO yields more informative failure-side credit signals with negligible training overhead.
cs.AI / 29 / 2609.02242
Propose to Learn, Learn to Propose: Evaluability-Aware Assistance under Bounded Rationality
Abstract
AI assistants often collaborate by proposing candidate edits, plans, or designs that users evaluate before adoption. Existing assistance methods focus on proposal quality or user-goal inference, often assuming that the user can reliably evaluate any proposal, which can fail in practice because of bounded rationality. We study evaluability-aware proposal planning, where proposals serve both as task interventions and as probes for learning latent preferences and evaluation constraints, where the resulting belief updates then guide later proposals. We formalise this setting as ProSE, a hidden-parameter sequential assistance problem, and instantiate it with a KL-regularised bounded-rational binary response model in which acceptance trades off value gain against a distance-dependent evaluability penalty. Analysing the planning consequence of this likelihood reveals that likely accepted proposals and informative probes need not coincide, which explains why planners that only pursue acceptance systematically underperform. We operationalise ProSE with \textsc{ProSE-Plan}, a depth-2 Bayes-adaptive planner that scores proposals by possible responses and response-induced posterior beliefs. In controlled graph simulations, \textsc{ProSE-Plan} improves over evaluability-unaware and myopic baselines when evaluation cost is the bottleneck, and a probe-commit ablation confirms that our approach selects informative proposals that simpler methods miss. Our results thus identify user evaluability as a planning-relevant dimension of AI assistance, complementary to generation quality and preference inference.
cs.AI / 30 / 2609.02246
LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails
Abstract
Self-improving agent pipelines have a problem at their center. An optimizer rewrites prompts to score higher, and the score comes from a judge that is itself an LLM. That judge has the last word on whether the system is getting better, and our position is that it has not earned it. The judge should be demoted from oracle to advisor: its verdict becomes one input among several, and every change is gated instead by a deterministic verification layer the judge cannot override. We reached this position by building the alternative and running it. Over months of running autonomous prompt-optimization loops in production across contract analysis, compliance review, and code quality, we cataloged eleven ways the evaluation signal failed, in four classes: judge bias, harness and metric failures, ground-truth errors, and reward hacking. Agents achieved perfect scores by reading cached answer keys from their environment, a 100% pass rate concealing 68% true capability. A corrupted ground-truth label caused the optimizer to delete correct compliance rules to agree with it. A syntactically broken prompt was promoted as the winner because a silent parser fallback improved the metric. Attempts to fix the judge by rewriting its rubric plateaued; the only reliable gain came from a structural constraint on its output order. In response we describe PROCTOR, a Teacher-Student loop in which a stateful orchestrator holds all tool access, stateless subagents diagnose failures and draft mutations they cannot apply, and a Teacher grades those mutations under five deterministic guardrails: hermetic sandboxes, capability-disjoint roles, acceptance checks that outrank the Teacher, frozen holdouts, and canary cases engineered so that a perfect score is itself evidence of cheating. We report the failures this prevented, and, because the Teacher is itself an LLM judge, the failures it did not.
cs.AI / 31 / 2609.02302
Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds
Abstract
A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluation can support. We present two techniques that make simulated alignment evaluations harder to distinguish from real deployments. Our first technique, critique refinement, spends additional inference-time compute on each simulator action: the simulator generates multiple candidate actions, refines them using feedback from an instance of the target model on how to make them more realistic, and continues the evaluation with the most deployment-like candidate. Our second technique, DISH (Deployment-Imitating SWE-Agent Harness), wraps the target in an agent harness, reducing the gap between simulated and real deployment environments in coding settings. We test the techniques on multiple target models and find that they compose: applying both yields larger realism gains than either alone. Our results show that automated approaches can improve the realism of alignment evaluations, and that these improvements use additional compute more effectively than making the audits longer.
cs.AI / 32 / 2609.02336
SALA: Semantic-Aware Logical Alignment for Complex Reasoning in In-Context Learning
Abstract
Effective in-context learning (ICL) for complex reasoning relies on selecting the right demonstrations. Traditional retrieval methods based on surface similarity fail to capture the underlying problem-solving logic. Recent logic-based methods address this by matching predefined reasoning steps, but the rigid rules and exact-match criteria is improper to handle flexible or diverse reasoning processes. To address the problem, we propose SALA, a Semantic-Aware Logical Alignment framework. Instead of relying on a fixed inventory, SALA automatically learns task-specific reasoning operations. It then embeds these operations into a continuous semantic space and uses dynamic time warping (DTW) to align the reasoning sequences. This approach allows for soft, flexible matching of reasoning logic while remaining highly interpretable. Experiments across four reasoning benchmarks and three LLMs demonstrate that SALA outperforms existing demonstration selection methods. Further analysis confirms the roles of the operation induction and the logical semantic alignment.
cs.AI / 33 / 2609.02371
Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions
Abstract
With the proliferation of LLM agents, the ability to understand and diagnose failures in agents is essential to achieving superior effectiveness and trustworthiness. As agent failures often manifest via long and complex trajectories, manually finding the needles in the haystack is untenable. However, traditional diagnosis techniques for software bugs can hardly address LLM agent failures, while completely relying on LLMs as the judge yields unreliable diagnosis results. To overcome these challenges, this paper presents AGENTSCOPE, a new neuro-symbolic approach for agent failure mode diagnosis. The key principle of AGENTSCOPE is to abstract agent behavior, based on its trajectories, into structured representations. Furthermore, AGENTSCOPE introduces the concept of neural invariants to specify agent behavior properties. AGENTSCOPE leverages LLM-guided reasoning atop the structured representation against neural invariants to pinpoint both the failure step and its type in the trajectory. We show the effectiveness of AGENTSCOPE on publicly available agent failure datasets (Who&When) and a more comprehensive dataset created by us (AgentErrata), where AGENTSCOPE significantly outperforms the current state of the art in fault localization and attribution accuracy. Our work shows that integrating structured abstractions with LLM-guided reasoning enables effective, reliable, and interpretable diagnosis for agent failures.
cs.AI / 34 / 2609.02399
Contrastive Explanations in Quantitative Bipolar Argumentation Frameworks
Abstract
Argumentation frameworks are useful tools for representing and reasoning with information in a variety of settings, e.g. in supplementing AI models as they perform classification tasks, with a notable benefit of providing additional explainability. In this paper, we introduce contrastive explanations for Quantitative Bipolar Argumentation Frameworks (QBAFs), one such formalism. Unlike most existing explanations for QBAFs, which explain the reasoning outcome of a single argument of interest (i.e. a topic argument), contrastive explanations explain the difference between two topic arguments. We introduce a general form of contrastive attribution functions (CAFs) and establish a set of general properties they should satisfy. We introduce CAFs based on removal, gradients and Shapley-values, and study their properties. Finally, to illustrate contrastive explanations, we demonstrate their usefulness in healthcare and bias identification settings.
cs.AI / 35 / 2609.02459
CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI
Abstract
We present CivBench, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments through the Model Context Protocol (MCP). A single episode spans 300+ turns and produces thousands of tool calls over a large action space, requiring sustained planning, state monitoring, and execution under partial observability. The environment exposes 76 MCP tools and a narration layer that converts visual game state into structured text. We use CivBench to characterise agent behaviour across four model families in 23 admissible runs. The sample is a pilot, not a model ranking: aggregate outcomes do not reliably discriminate models at this scale. Instead, we introduce two interface-level metrics that the environment makes measurable: Proactive Monitoring Rate (PMR), capturing whether agents actively query latent strategic state, and RAG@10, capturing whether commitments stated in structured planning reflections are executed within ten subsequent turns. Across runs we observe two consistent patterns under a shared playbook protocol. Agents under-monitor strategically relevant state that is available but requires explicit querying: despite playbook guidance to query victory progress every 20 turns, agents do so only every 30 to 75 turns, and in 7 of 20 detectable defeats they failed to query within the 20 turn warning window before game end. Agents also frequently fail to execute near-term commitments stated in their own planning reflections (RAG@10 between 48.2% and 65.8% across models). Both patterns arise despite tool access and explicit guidance, and we interpret them as deviations under instruction rather than absences of capability. We release the environment, scenarios, logs, metrics, and analysis pipeline at https://github.com/lmwilki/civ6-mcp
cs.AI / 36 / 2609.02620
Collective creativity in hybrid societies
Abstract
Generative AI is changing how cultural artifacts are created and circulated, and with it our understanding of creativity itself. Researchers disagree about whether these tools enrich or impoverish culture, and we argue that much of that disagreement comes from conflating two distinct components of creativity: novelty, a property of single artifacts, and diversity, a property of populations. We argue further that creativity in the context of generative AI is best understood as a property of hybrid collectives, or populations of interacting people and algorithms, rather than of individuals. AI-assisted ideation reliably raises the novelty of individual output while narrowing diversity in the aggregate, but this is not an inevitable consequence of putting machines in the loop. Because humans and models search in complementary ways, mixed groups can outperform and out-diversify groups of either kind alone, and machine-discovered solutions can enter human culture and persist there. What decides the outcome is composition: which agents are present, in what proportion, and how they are connected. The question is no longer whether AI helps or harms creativity, but which mixtures let individual gains accumulate without eroding collective diversity.
cs.AI / 37 / 2609.02749
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
Abstract
Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model backbone with a harness for planning, execution, memory, and verification, but this architecture still leaves domain-specific know-how outside the agent. We call this missing layer operational knowledge, the know-how that separates knowing a method from making it work. That knowledge is not absent from the field. It appears in repositories and papers, but in forms written for human readers and too large to load during a task. Once distilled into compact, verified skills, this knowledge can be reused across tasks rather than rediscovered during each run. We present DisCo, a skill-powered research agent that creates skills and uses them during research. Its distillation runs in two complementary forms: task-agnostic, condensing the field's widely used repositories into reusable skills, and task-oriented, producing the skills a concrete task calls for. The former, applied across the open ecosystem, yields the AREX-Skill Library, with 5,000+ verified skills distilled from 1,000 widely used ML repositories and organized into 20 areas and 178 capability families. With the GPT-5.5 backbone, research harness, and downstream execution budget held fixed, the skill-equipped research agent scores 134.3% higher on MLE-bench, 34.4% higher on PaperBench, 9.2% higher on FrontierCS, and 14.0% higher on PassNet than the same agent without skills. These gains come from adding distilled operating context under that fixed setup.
cs.AI / 38 / 2609.02750
Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems
Abstract
Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through textual reflection. Despite strong empirical results, these systems lack a unified account of coordination, memory improvement, and the role of external verification. We model orchestrator-worker interaction as a bilevel coordination game: under bounded coupling, the workers' local-update game is an approximate potential game whose equilibrium slack is controlled by decomposition quality. We then analyse reflection as stochastic movement over semantic memory states. For free-form reflection, we derive a finite-time upper bound, prove worst-case tightness, and give a positive lower bound under a falsifiable persistent-harm condition. We further prove an information-theoretic impossibility result: no gate that observes only the generated transcript can improve uniformly over text-indistinguishable environments, whereas an environment-grounded gate can. Motivated by this separation, we introduce Stochastic Reflective Memory Ascent (SRMA), which accepts a candidate memory only after a grounded evaluation risk strictly decreases. Under calibration and non-degenerate corrective mass, SRMA converges exactly, geometrically or polynomially; matching constructions show that both rate regimes are order-tight. We also provide confidence gating for stochastic evaluation and re-anchoring guarantees for piecewise-stationary environments. Experiments instantiate these objects with environment-grounded metrics and test the predicted coordination and drift laws. On 500 SWE-bench instances, the complete Kimi-based system resolves 72.2% versus a 70.8% public mini-SWE-agent reference. Code: https://github.com/YihangChen9/Bilevel-Coordinated-Reflection
cs.AI / 39 / 2609.02760
Measurement-Driven Sub-Network Selection for On-Premise Retrieval-Augmented Factory Agents
Abstract
On-premise assistants can give factory workers conversational access to machine documentation, but models capable of the task rarely fit shop-floor hardware. We show that after structural compression and retrieval-grounded adaptation, model size is no longer a reliable predictor of adapted answer quality: general capability falls almost linearly with parameter count, while judged retrieval-augmented answer quality does not. We therefore treat deployment as a post-adaptation selection problem, committing one sub-network per device on judged answer quality and measured on-device throughput under a configurable general-capability floor and memory budget; rules that optimize size, speed, or quality alone each give up capability or throughput. A weight-shared supernetwork trained with sandwich-style in-place distillation keeps this selection inexpensive. In a manufacturing-manual case study, extraction costs 13.7 percent of the unpruned model's judged quality and retrieval-grounded distillation returns it to within 4.6 percent, recovering two thirds of the loss, and the same assistant runs across three heterogeneous edge tiers at 1.3 to 5 watts standby.
cs.AI / 40 / 2609.02786
SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment
Abstract
The performance of LLM-based agents is jointly shaped by the base model and the harness used when interacting with the environment. This exposes them to safety risks in both harmful final responses and multi-step execution trajectories. Existing safety alignment mechanisms often rely on either external harness updates or policy optimization, yet applying either paradigm in isolation fails to bridge runtime control with intrinsic safety. We propose SafeEvolve, an experience-driven self-evolving framework for agent safety alignment. SafeEvolve leverages safety experience from completed on-policy trajectories to drive a continual loop of harness-policy co-evolution. On the harness side, SafeEvolve converts trajectory-level safety evidence into bounded, component-level updates across safety prompt and hierarchical skills, yielding auditable and reversible harness artifacts. On the policy side, SafeEvolve follows a two-stage SFT-RL paradigm, where harness-use SFT bootstraps the policy to actively leverage evolved harness artifacts, and harness-augmented RL further shapes autonomous safety behaviors during multi-step exploration via verifier-decomposed rewards. Through harness-policy co-evolution, SafeEvolve converts safety experience into an evolved runtime harness and improved policy behavior. Experiments on agentic safety benchmarks show that SafeEvolve achieves a stronger safety-utility tradeoff than existing baselines. For Qwen3.5-4B, SafeEvolve achieves a $3\times$ ASR reduction on AgentDojo while improving benign utility from 59.79% to 61.86%.
cs.AI / 41 / 2609.02821
AI Contextual Measurement for Recovering Individual and Group-Level Effects: Validation Against Survey Measures and an Occupational Application
Abstract
Researchers increasingly use artificial intelligence to construct measures of social, organizational, and occupational characteristics that are absent from conventional surveys. We propose AICOME, AI COntextual MEasurement, a framework for evaluating whether AI-derived respondent-level measures can recover individual and group-level effects in contextual models. The key idea is that an AI measure constructed at the respondent level can be used to derive its group-level aggregate and its individual deviation, allowing researchers to estimate both between-group and within-group associations rather than treating AI measurement as response prediction alone. We validate the framework using the 2022 China Family Panel Studies (CFPS), where occupations provide the empirical grouping structure and several job-related survey variables provide validation benchmarks. For computer use, foreign-language use, weekly hours, and management responsibilities, we compare survey measures with AI-derived measures in response-level, model-level, contextual, and boundary-condition validations. The results show that AI contextual measurement can recover much of the contextual-model information contained in observed survey variables when rich respondent and job characteristics are available. Weekly hours provides the strongest validation case, with AI-derived measures reproducing the large negative between- and within-occupation associations with satisfaction observed in CFPS. The framework also identifies clear boundary conditions: performance deteriorates when information is restricted to occupation and basic demographics, and recovery is weaker when several related concepts are treated as simultaneously unobserved. The findings suggest that AICOME is most useful for recovering a limited number of theoretically important constructs from rich existing datasets.
cs.AI / 42 / 2609.02885
Discriminative World Models for Web Agents
Abstract
Recent web agents use world models for test-time action selection by sampling candidate actions, predicting the resulting web states, and ranking them with a ranker model or a Process Reward Model (PRM). These world models are typically trained via supervised next-state prediction to generate fixed representations like HTML or AXTree snapshots. However, this objective is misaligned with the downstream ranker, which relies on predicted states being discriminative across candidates to accurately score them. To address this, we introduce predicted-state matching, a training objective where the predicted representation must distinguish the true resulting state from those reached by alternative actions. We train these models using a branching web-agent dataset derived from WebArena Go-Browse trajectories, where every decision point contains multiple alternative actions and their resulting states. Experiments on our held-out predicted-state matching benchmark show that our approach outperforms world models trained with supervised next-state prediction. We further show that our approach improves PRM-style action ranking on WebPRMBench compared with action-only PRMs and PRMs augmented with supervised-next-state world models. Finally, on WebArena-Lite, using our world model for test-time action selection improves end-to-end task success. Our project page is available at: https://dhruvpendharkar.github.io/dwm/.
cs.AI / 43 / 2609.02002
InsightSeg: Reusing Correction Insights for Guideline-Consistent Segmentation
Abstract
Guideline-consistent semantic segmentation requires more than category recognition, as real-world labeling policies demand fine-grained, task-specific decisions. Recent multi-agent refinement systems improve compliance with such textual guidelines by detecting and correcting errors. However, they are stateless: feedback from the critiquing agent is discarded, causing the same guideline-specific mistakes to be repeatedly rediscovered and corrected across the dataset at the cost of additional refinement. We introduce InsightSeg, an episodic memory mechanism that converts successful correction episodes into reusable, visually grounded insights. A meta-analyzer distills each qualifying episode into directive natural-language insights and anchors them to the local image regions that caused the error using patch-level visual concept vectors. On subsequent images, these concepts are matched against dense patch embeddings to retrieve relevant insights, which condition the segmenting agent before making its first prediction. This shifts the system from correcting recurring errors to preventing them, improving segmentation quality before any refinement occurs. Across Waymo and Cityscapes, InsightSeg improves both first-pass and final guideline-consistent segmentation performance while requiring fewer refinement steps, demonstrating that multi-agent refinement can become more accurate and efficient by drawing on past correction experience.
cs.AI / 44 / 2609.02111
Disease Burden over Skin Tone: Decomposing the Dermatology-AI Generalization Gap
Abstract
Dermatology artificial intelligence (AI) models are predominantly trained on light-skinned, cancer-focused image collections, yet they are increasingly proposed for deployment in resource-constrained settings where patients differ from training populations along two confounded axes: skin tone and disease distribution. We investigate whether poor generalization is primarily caused by skin-tone underrepresentation or disease-distribution shift. We evaluate a cancer-trained baseline (ResNet-50 fine-tuned on HAM10000 and ISIC 2019), two dermatology foundation models (DermLIP and MONET), and a general-purpose vision model (DINOv3) as frozen feature extractors. Models are evaluated on a tone-stratified disease-matched dataset (Diverse Dermatology Images, DDI) and a disease-shifted tone-diverse dataset (Skin Condition Image Network, SCIN). Our results show that disease-distribution shift contributes more than skin tone in the evaluated settings. The cancer baseline decreases from 0.62 to 0.21 balanced accuracy when transferred to unfamiliar clinical conditions, while the within-disease skin-tone gap is smaller (0.10-0.18) and inconsistent. Label-free representation analysis shows that this failure reflects a representational limitation rather than only missing output labels: cancer-specialized features poorly cluster unfamiliar conditions (kNN purity lift +0.06 over chance), whereas dermatology-pretrained features retain stronger transferable structure (+0.23). Finally, we show that representation quality predicts recoverable performance under lightweight adaptation. Starting from dermatology foundation models, approximately ten labeled examples per clinical category recover most attainable performance. We release the evaluation protocol and code to support reproducible auditing of dermatology AI generalization.
cs.AI / 45 / 2609.02224
Signal or Noise? Auditing Rotation-Induced Saliency Drift in Medical and Aerial Imaging
Abstract
Post-hoc saliency maps such as Grad-CAM are increasingly used to audit why a deployed vision model made a decision, yet the heatmap drifts when the input is rotated, even when the prediction is unchanged. In domains with no canonical orientation, such as histopathology and aerial imagery, this undermines using saliency as evidence. We ask whether that drift is faithful signal or noise introduced by the CAM operator, and answer it by measuring equivariance at every stage of the operator rather than inferring it from the network's output. The instability is not where one would guess: the channel weights are the most rotation-stable stage, and on ResNet-50 exactly stable, because a GAP+linear head makes the class gradient field spatially constant. What moves is the spatial activation tensor, and the classifier's own pooling discards that movement. A causal test confirms the consequence: occluding the pixels whose saliency drifts costs the model less than occluding random pixels, at either orientation. The drift is carried by degrees of freedom the classifier throws away, which is what makes removing it faithful rather than destructive. EquiGrad-CAM is a training-free wrapper that takes T rotated views, inverse-rotates each view's saliency into a common canonical frame, and averages. On the full ImageNet-1K validation set it raises equivariance over single-view Grad-CAM by +36.0% (ResNet-50), +87.5% (VGG-16) and +247% (ViT-B/16); a scale-matched ablation isolates alignment before averaging, not the locus of aggregation, as the driver. It beats rotation-augmented training without retraining, lifts zero-shot CLIP by +145%, and yields rotation-consistent explanations on PatchCamelyon and RESISC45. Its by-product PEUM ranks explanations by how reproducible they are, at no cost beyond the views already taken. Code: https://github.com/Khawaja-Murad/EquiGrad-CAM
cs.AI / 46 / 2609.02233
InfraPatch: Cross-Task Targeted Grayscale Patch Attacks on Infrared-Adapted Vision-Language Models
Abstract
Infrared vision-language models (IR-VLMs) have emerged as a promising paradigm for multimodal perception under low-visibility conditions, yet their robustness to targeted adversarial attacks remains poorly understood. Existing adversarial patch methods mainly study RGB-based models or a single downstream task and do not characterize whether localized perturbations can induce an intended semantic target in IR-VLMs. We propose InfraPatch, a white-box, per-instance framework for targeted digital grayscale patch attacks against IR-VLMs. InfraPatch optimizes a compact single-channel patch within an approximately 5% local-area budget, combines proxy-guided placement with task-adaptive semantic objectives, and induces target behaviors in image classification, image captioning, and binary visual question answering. We evaluate ten infrared-adapted model variants on 300 synthetic infrared-style images generated by applying DiffV2IR to a fixed 30-category COCO subset, using clean-conditioned targeted success criteria. InfraPatch achieves targeted attack success rates from 86.00% to 100% across the ten variants. On CLIP and BLIP-2, proxy location search improves success by 6.67 and 10.33 percentage points over optimized random placement, respectively; LLaVA-1.5 remains saturated near 100% under both settings. Patch-area and objective ablations further expose substantial differences in vulnerability across architectures and task formats. These results show that small grayscale patches can inject chosen target semantics across IR-VLM families under a controlled digital threat model, motivating stronger robustness evaluation for infrared multimodal systems.
cs.AI / 47 / 2609.02247
SAUF-Net: Structure--Appearance Representation Learning with Uncertainty Feedback for Semi-Supervised Medical Image Segmentation
Abstract
Semi-supervised learning has shown great potential for reducing annotation costs in medical image segmentation. However, most existing methods mainly exploit unlabeled data through prediction-level consistency, while the reliability of internal feature representations is often overlooked. In medical images, target-related structural cues are easily entangled with unstable appearance variations, which may lead to unreliable pseudo labels and error accumulation during training. To address these issues, we propose SAUF-Net, a Structure--Appearance Representation Learning with Uncertainty Feedback Network for semi-supervised medical image segmentation. SAUF-Net uses the Structure--Appearance Decomposition Module (SADM) to separate bottleneck features into structural and appearance representations. The Disentangled Guidance Module (DGM) injects these representations into the decoding process to enhance structure-aware segmentation. Meanwhile, the Auxiliary Decoder produces branch-specific predictions for reliability estimation and a fused prediction for appearance-swapped consistency. Furthermore, we introduce an Appearance-Swapped Consistency branch to encourage structural representations to remain stable under appearance variations. We also introduce a reliability-map-guided dual-head discriminator with a Validity Head and an Uncertainty Head to provide feature-level uncertainty feedback. Extensive experiments on ISIC-2016 and Kvasir-SEG demonstrate that SAUF-Net outperforms state-of-the-art semi-supervised methods, especially under low-label settings.
cs.AI / 48 / 2609.02282
RouteGraph-Mona: Confusion-Aware Routing Fine-Tuning for Mineral Image Classification
Abstract
Mineral image classification is important for geological exploration and resource development, but it remains challenging due to substantial intra-class variations in appearance and high inter-class visual similarity. Multi-cognitive Visual Adapter (Mona) is a vision-oriented parameter-efficient adapter that adapts pre-trained visual models by tuning only a few parameters. However, Mona statically aggregates responses from multiple scales, limiting its ability to accommodate sample-specific scale preferences and model confusion among visually similar mineral categories. To address this issue, we propose \textbf{RouteGraph-Mona}, a lightweight route-space regularization method built on Mona. Specifically, we replace Mona's static multi-scale aggregation with sample-adaptive routing. The resulting branch-selection behavior defines a compact routing space that captures each image's scale preferences. We then regularize the resulting routing signatures with class-wise route anchors and confusion-weighted margins. The route anchors encourage class-consistent routing patterns, while the margins promote greater separation between visually similar categories in the routing space. Experiments on three public mineral image datasets with two visual backbones show that RouteGraph-Mona consistently outperforms Mona in mean accuracy and remains competitive with representative fine-tuning methods and mineral image classification baselines.
cs.AI / 49 / 2609.02333
ORB-SVM : An Innovative Hybrid Framework for Efficient Brain Tumor Detection from MRI Scans
Abstract
Brain cancer remains one of the most significant challenges in modern medicine, where the accuracy of early stage diagnosis is a decisive factor in patient survival and treatment efficacy. Although Magnetic Resonance Imaging (MRI) is the established gold standard for visualizing neurological structures, the interpretation of these high dimensional scans is often complicated by subjective variability among practitioners and the inherent noise present in complex medical images. While contemporary approaches frequently rely on high parameter deep learning architectures, such models often involve significant computational costs and require extensive data for effective training. This study introduces a hybrid framework that utilizes the Oriented FAST and Rotated BRIEF (ORB) algorithm for precise feature extraction and a Support Vector Machine (SVM) for classification [1], [2]. The proposed approach achieves a sub- stantial data reduction of approximately 99.5%, which effectively minimizes the influence of non informative background data while preserving critical diagnostic patterns essential for tumor identification. By balancing feature sparsity with a robust kernel based classifier, this methodology addresses the limitations of over parameterized systems while maintaining high diagnostic integrity. Experimental evaluations conducted on the Br35H dataset demonstrate that the framework attains a classification accuracy of 97.5%. The findings suggest that the integration of localized feature representation and optimized classification provides a reliable and resource efficient alternative for medical image analysis, offering a structured solution that maintains per- formance without the need for extensive computational overhead.
cs.AI / 50 / 2609.02502
Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models
Abstract
Text-to-image (T2I) models have achieved remarkable success at faithfully rendering specified objects and attributes, yet their ability to produce visual metaphors, images that convey abstract ideas by combining elements from two distinct domains, remains largely unexamined. To bridge this gap, we introduce VMetaphor-Bench, the first benchmark for evaluating visual metaphor generation in T2I models. It comprises 1,500 visual metaphors curated from real-world creative imagery, organized into three levels and ten categories, with each sample paired with two prompts of differing specificity. For evaluation, we develop a hybrid framework within an MLLM-as-judge paradigm, combining a multiple-choice question (MCQ) based protocol of 9,594 questions across four levels of metaphorical fidelity with a dimension-based scoring protocol along three perceptual dimensions. Extensive evaluation of 11 representative T2I models reveals that even the strongest proprietary models struggle with compositional structuring and cross-domain mapping, key aspects of metaphorical expression, highlighting visual metaphor generation as an important frontier for future T2I research.
cs.AI / 51 / 2609.02529
Fine-Grained Anomaly Perception in Wild UGC-Enhanced Images: A Comprehensive Dataset and Difference-Fusion Framework
Abstract
Image enhancement and restoration have become standard back-end operations on short-video and social media platforms to boost UGC visual experience. Yet these processes inevitably introduce visual anomalies--especially in faces, texts, and textures--that directly undermine perceptual fidelity and viewer trust. While existing IQA methods perform well on classic distortions, they target holistic quality assessment and fail to capture the specific, localized anomalies caused by enhancement algorithms in real-world UGC. To bridge this gap, we formally define a new task-quality Anomaly Perception for UGC image Enhancement (UEAP), and contribute the first UEAP benchmark dataset, named UEAP-4k, curated from the real business scenarios. It provides fine-grained annotations for anomaly categories, localization and severity levels. Furthermore, we propose a Difference-Fusion Anomaly Perception Method (DFAP-UGC) for wild UGC-enhanced images, which leverages explicit problem-reference difference fusion with dense spatial querying, regional verification, and quality-aware ranking, enabling robust anomaly identification in challenging scenarios. To handle the inherent coupling of subtasks in this new task, we propose a Locality-Aware Dynamic Task Prioritization (LADTP) training strategy that enables effective end-to-end learning and eliminates multi-stage overhead. Extensive experiments show that our method outperforms baselines adapted from classical approaches for this task, validating the value of this dataset and the superior of DFAP-UGC for robust UGC-enhanced image anomaly perception. Code and data will be public.
cs.AI / 52 / 2609.02731
RVSD: Retrieval Vision Sparse Decoding for Mitigating Visual Hallucinations in Large Vision-Language Models
Abstract
Large vision-language models have achieved remarkable success in vision-language tasks. However, they remain prone to Visual Hallucinations (VHs), undermining their reliability in real-world applications. Existing solutions typically require curated datasets, additional training, or multi-round decoding, resulting in considerable computational overhead. In this paper, we propose \textbf{RVSD} (\underline{R}etrieval \underline{V}ision \underline{S}parse \underline{D}ecoding), a training-free and plug-and-play decoding framework that, for the first time, unifies token sparsification and \textbf{Semantic-Space Visual Retrieval} (SSVR) within a single decoding pass. Within RVSD, we introduce a \textbf{semantics-directed token selection} strategy that selectively sparsifies redundant tokens while preserving critical visual information. We further propose the SSVR mechanism, which reformulates visual compensation as an on-demand cross-modal retrieval process within a shared semantic space. Extensive experiments demonstrate that RVSD achieves state-of-the-art performance in mitigating VHs while maintaining robust suppression capabilities under long-context generation settings. Our code is available here.\footnote{https://github.com/canjie-liu/RVSD}
cs.AI / 53 / 2609.02453
Addressing Trust in AI Systems through Education: A Didactic Perspective
Abstract
Machine learning (ML) education faces two persistent and connected obstacles: many educational tools present ML as an opaque black box, which leaves learners with a superficial understanding, and this same opacity prevents users from forming the calibrated trust that appropriate reliance on AI systems requires. We present ICE-T, a didactic framework that integrates three mutually reinforcing facets: intermodal transfer grounded in Bruner's enactive, iconic, and symbolic modes of representation, computational thinking operationalized through the Use-Modify-Create progression, and explanatory thinking supported by a process model. Connecting the framework to the empirical literature on algorithm aversion, AI literacy, and mental model formation, and to systematic reviews of the K-12 ML activity landscape, we argue that the three facets supply the cognitive mechanisms that the trust calibration literature identifies as drivers of appropriate reliance: representational richness, graduated process control, and the capacity to contextualize errors. On this basis, we propose that trust calibration be treated as an explicit educational objective, with ICE-T as a principled and scalable means of achieving it.
cs.AI / 54 / 2609.01818
Zeta-Lite: A Concurrent, Branchable In-Browser SQL Database for Agentic Memory
Abstract
The browser has become a first-class database host: applications increasingly want to store, query, and reason over structured data entirely on the client - for privacy, offline operation, local-first collaboration, and, most recently, as durable memory for in-browser AI agents. One way to get SQL in the browser, compiling PostgreSQL to WebAssembly (PGlite), inherits PostgreSQL's process model: a single backend connection that executes one statement at a time and blocks. That model cannot express concurrent transactions, and it leaves richer capabilities - graph queries, database branching - to whatever the compiled server happens to include. We present zeta-lite, the browser form factor of the Zeta database engine: a WebAssembly build that compiles the same Zeta server down to a 2.87 MB gzipped artifact. Zeta-lite keeps the engine's log-centric asynchronous MVCC core, which yields two capabilities no other in-browser SQL engine provides. First, overlapping snapshot-isolated transactions on a single thread: multiple transactions hold distinct read/commit timestamps and interleave, with snapshot-isolation conflict detection between them. Second, copy-on-write database branching - whole-database fork, merge, and rebase - is unique in a browser SQL database and rare even in servers. On top of these, zeta-lite exposes a feature-complete PostgreSQL surface (joins, CTEs, window functions, JSONB with GIN indexes, full-text search, HNSW vector search, SQL/PGQ graph queries, multi-database) and snapshot-to-OPFS durability. Across Chrome, Firefox, and a native reference runtime, zeta-lite sustains 268k-315k point reads/s and holds a mixed read/write workload flat over millions of operations. This small, fully-featured, concurrent SQL database is an especially good fit for agentic memory - where cheap branchable state lets an agent explore, inspect, and commit or discard speculative work.
cs.AI / 55 / 2609.02143
A Power Law in Logarithm's Clothing: On the Scalability of Graph-Based Vector Search
Abstract
Most vector databases rely on graph-based indexes, notably HNSW and Vamana, for approximate nearest neighbor search. With embedding models widely adopted, the datasets these databases store grow rapidly. At a fixed accuracy, how does search cost scale with dataset size? The prevailing answer is poly-logarithmic growth. Yet the claim is proven only under special conditions and asserted without proof for the indexes used in practice. It is also largely untested: standard benchmarks measure cost at one dataset size, not across sizes. We put the claim to the test. The answer depends on the scale itself. While the dataset size $N$ is small relative to the data's intrinsic dimensionality, search cost grows as $N^c$ for a constant $0<c<1$. We call this scaling the Sublinear Power Law. Once $N$ is large enough, growth slows to subpolynomial, consistent with the poly-logarithmic claim. The Sublinear Power Law appears on every dataset, mostly up to its full size, at every recall target, query hardness level, and index configuration we test. The transition to subpolynomial growth appears on the two datasets that grow large enough relative to their intrinsic dimensionality. One mechanism underlies both behaviors: a dataset's intrinsic dimensionality grows with its size until the data resolves its underlying distribution. Higher intrinsic dimensionality packs more vectors into the query neighborhood the search must examine. We present a unifying theory of beam-search cost that explains our observations. For exact and bounded-degree constructions, we prove the Sublinear Power Law and the eventual transition to poly-logarithmic scaling, and derive the scale at which it occurs. We also develop models that predict the power-law exponents for any recall target and index configuration. These models give a principled way to navigate trade-offs among search cost, insertion cost, and recall as data grows.
cs.AI / 56 / 2609.02109
MeanField Surrogate Modeling for Scalable Runtime Scheduling of Concurrent Heterogeneous AI Inference on Shared GPUs
Abstract
Deploying heterogeneous AI models concurrently on a shared GPU introduces resource contention that complicates runtime scheduling. While surrogate models avoid costly online benchmarking, their profiling requirements typically grow combinatorially with the number of co-running models, limiting scalability. We propose a MeanField surrogate that predicts per-model performance from local configuration and aggregate GPU state rather than explicitly modeling all joint interactions. Experiments on concurrent LLM and vision workloads across $N \in \{2,3,4,5,6\}$ show high predictive accuracy ($R^2 \approx 0.96$) with an empirical sample budget that grows approximately linearly in $N$, in contrast to the combinatorial cost of fully joint profiling. Integrated into a genetic algorithm scheduler, the surrogate scales to an $N=5$ problem with 78,732 feasible joint configurations, remaining within 0.10% of the exhaustive search with zero SLA violations across eight dynamic workload scenarios, while complete online GA decisions take 26 ms median, about $5\times$ faster than exhaustive surrogate search.
cs.AI / 57 / 2609.02804
frb100-40 After Two Decades: An Optimality Certificate and a Preregistered Search Study
Abstract
For more than 20 years, the Model-RB benchmark frb100-40 remained an open challenge; since 2014, its public record had stood at 99 of 100 variables. We give a directly checkable 100-vertex independent set for its 4,000-vertex graph. Together with a verified partition into 100 cliques of size 40, the witness proves that the maximum independent-set size is 100 and the minimum vertex-cover size is 3,900. The stochastic run that found the witness is kept separate from this proof. We evaluated its added pair and triple repair operators in a preregistered campaign comprising 8,668 valid runs. The primary comparison found no detectable acceleration over base ULSA (hazard ratio 0.967, 95% confidence interval 0.915-1.023; p=0.248), and the factorial ablation reached the same conclusion. On a smaller FRB suite, the group-aware CSP pipeline solved 2,500/2,500 runs, compared with 2,391/2,500 for LibMVC-NuMVC. On frb100-40, full ULSA, base ULSA, and NuMVC each produced 0/56 new certificates. With no events, the planned cross-solver hazard ratios remain unidentified. NuMVC ended with cover size 3,902 in 40 runs and 3,903 in 16. Exhaustive enumeration showed that none of the 108 unique recorded conflict-two states had a strictly improving group-aware CSP neighbor within Hamming radius three. The certificate settles the instance. The experiments characterize the search barrier, and the preregistered comparisons show no heuristic advantage.
cs.AI / 58 / 2609.01775
Dictionary-Guided Mutation Operators for Automated HDL Repair
Abstract
Automated repair of Hardware Description Language (HDL) designs remains challenging due to the large search space of candidate repairs and the strict syntactic and semantic constraints imposed by HDL grammars. Generic mutation strategies overwhelmingly generate syntactically invalid candidates that waste compilation and simulation budget, while synthesis-driven and template-based approaches impose their own constraints on generality and portability. In this paper, we propose a dictionary-guided HDL repair system that combines ANTLR-derived DUT-specific mutation vocabularies with a simulation-divergence fault localization (FL) module. The mutation operator applies category-constrained token substitutions, insertions, and deletions directly to Verilog source via regex-based matching, without requiring AST manipulation or synthesis. The FL module identifies diverging output wires from a single simulation run and scores source lines by structural proximity to those signals, directing the mutation search toward high-suspicion regions. A deterministic targeted sweep exhausts all dictionary mutations on the highest-scored lines before falling back to a genetic programming (GP) search. Evaluated on the CirFix benchmark suite across six design under test (DUT) families, the proposed approach produces correct oracle-passing repairs on 14 bug variants, including a 6-edit multi-bug instance that CirFix cannot repair, and achieves an 18x speedup over CirFix on a two-edit benchmark variant. These results indicate that dictionary-constrained mutation operators, combined with lightweight simulation-divergence FL, are a practical and competitive approach to automated HDL repair for common bug classes without formal analysis or synthesis dependencies.
cs.AI / 59 / 2609.02354
Fair Stable Matching: A Nash Social Welfare Approach
Abstract
While traditional stable matching algorithms, such as the Gale-Shapley algorithm, prioritize stability, they may fall short of achieving equitable outcomes among participants. We study the role of \emph{Nash social welfare} (NSW) as a fairness objective in the classic \emph{stable marriage problem}. We develop \texttt{SNSW-Alg} that finds a stable matching that maximizes Nash social welfare under rank-induced utilities in $\tilde{\mathcal{O}}(n^4)$ time, where $n$ is the number of men or women. We demonstrate that \texttt{SNSW-Alg} balances equity while preserving stability. We empirically evaluate our methods across diverse preference distributions, demonstrating significant gains in fairness without substantial losses in other key measures such as regret, egalitarian criterion, and sex equality. Our findings suggest that the stable matching produced by \texttt{SNSW-Alg} is statistically Pareto-undominated by stable matchings based on other fairness measures - regret, egalitarian, and sex equality. This study offers compelling insights for designing fair-stable matching.
cs.AI / 60 / 2609.02364
Towards a Foundational Ontology for Identifying and Resolving Contradictions in Dialogue-based Human-Robot Interactions
Abstract
Existing Human-Robot Interaction (HRI) literature has focused on identifying and structuring errors, failures, conflicts, and knowledge issues (called in this work as contradictions) in domain-specific dialogue-based interactions. However, there is still lack of a formal computational framework to represent and define these contradictions, interoperable and usable across HRI and human-agent interaction (HAI) domains. Thus, this research project aims to capture, represent, and evaluate the notion of (1) dialogue-based collaborative interaction and (2) related contradictions in a foundational ontology. METHONTOLOGY, a systematic approach to build domain-independent ontologies was applied. In the conceptualisation stage of the presented ontology, concepts and models from Activity Theory were used. Preliminary results presented in this short article are: (i) Natural language definitions of dialogues and related contradictions in HRI, (ii) Set Theoretic definitions of dialogues and contradictions, and (iii) First Order Logic (FoL) formulation of the contradiction concepts and three novel principles guiding dialogue-based interactions between humans and robots. In summary, we report on ongoing work to develop a foundational ontology based on Activity Theory called Activity Theory-based foundational ontology (ATFOt) to capture and represent the notion of contradictions in HRI.
cs.AI / 61 / 2609.02152
Beyond Modality Harmony: Orthogonal Purification and Topology-Guided MoE for Conflict-Aware Multimodal Recommendation
Abstract
Multimodal Recommender Systems (MRSs) typically rely on a flawed "modality harmony" assumption, presuming that multimodal features are inherently beneficial and strictly aligned with users' collaborative interaction patterns. However, modality-topology conflicts are ubiquitous in real-world scenarios due to deceptive visual clickbaits and mismatched semantics. Blindly integrating these noisy modalities inevitably pollutes the pristine collaborative space, causing severe representation distortion. To address this, we propose Orthogonal purification and topology-guided MoE for conflict-aware multimodal Recommendation (OrthoRec). At its core, OrthoRec introduces Collaborative-Guided Orthogonal Purification (CGOP), which geometrically decouples multimodal features into directions parallel and orthogonal to a pure collaborative anchor. By adaptively truncating the orthogonal noise with an energy-preserving normalization, CGOP rectifies deceptive semantic directions while preserving the modality's intrinsic representation capacity. Furthermore, we design a Topology-Aware Routing Mixture-of-Experts (TAR-MoE). Guided by the collaborative topology, TAR-MoE employs decoupled sigmoid gating to break the zero-sum bottleneck of traditional softmax attention, autonomously determining the injection scale for each purified modality. Finally, a safe-SSL objective is introduced to dynamically penalize the forced contrastive alignment of contradictory pairs. Experiments on three real-world Amazon datasets show that OrthoRec consistently outperforms competitive recent baselines and exhibits improved robustness under modality noise and item sparsity.
cs.AI / 62 / 2609.02486
ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question Answering
Abstract
Document Visual Question Answering (DocVQA) often leverages Retrieval-Augmented Generation (RAG), where late-interaction encoders are commonly used to identify document pages relevant to a user query, before answer generation by a Large Vision-Language Model (LVLM). Existing approaches typically retrieve a fixed top-$k$ number of pages regardless of query complexity, which increases LVLM latency and may degrade answer accuracy. We introduce ViSAR (Visual Semantic Activation Retrieval), a training-free adaptive-$k$ retrieval method for late-interaction visual document retrieval. ViSAR operates directly in the embedding space to construct a query-conditioned page-level similarity matrix that highlights query-relevant semantics and dynamically determines the number of pages to retrieve. Across multiple encoders and LVLMs, ViSAR retrieves compact, query-adapted page sets that reduce RAG latency by up to 58.7\%, while maintaining or improving answer accuracy compared with fixed top-$k$ and adaptive retrieval heuristics. Furthermore, we show that the similarity matrix structure correlates with answer accuracy, suggesting future directions for retrieval quality-aware document understanding.
cs.AI / 63 / 2609.02046
Modeling What Changes: Sparse, Residual World Models for Object-Centric Manipulation
Abstract
Monolithic world models predict the entire next state at every step, spending capacity re-predicting the static majority of a scene and injecting error into it. We ask whether explicitly modeling change (a per-object change gate plus a residual delta head that perturbs only the objects the gate flags) is a more effective and interpretable bias for physical prediction and control. On a MuJoCo tabletop pushing benchmark scaling from 3 to 8 objects, the sparse/residual model predicts next-state poses 2.5 to 4.6 times more accurately than a dense multilayer perceptron at 8.6 to 11.1 times fewer parameters, sustains change-detection F1 of 0.80 to 0.87 where the dense baseline is degenerate, transfers across object counts with zero retraining (99.4 percent F1 retention), and reaches about 90 percent of its full-data accuracy with a quarter of the data. In autoregressive rollout it compounds far less error, hugging the no-motion floor while the dense model drifts. Finally, inside a sampling-based planner, prediction-only models fail (though a true-simulator oracle solves the task with the identical planner, confirming the planner is sound), but once featurized and trained for the states a planner visits, the sparse model begins to plan (0.23 plus or minus 0.06 success over three seeds) while the dense monolith stays at zero at every seed. Modeling what changes, rather than re-predicting the whole world, is a simple, effective bias for object-centric physical AI; code, data generators, and all checkpoints will be released upon publication.
cs.AI / 64 / 2609.02861
Towards Trustworthy Autonomous Robots: An Explainable AI-Based Decision Framework
Abstract
Autonomous robots powered by deep learning face a fundamental auditability challenge: when incidents occur, investigators cannot reconstruct why the system made specific decisions. This paper presents TRACE (Transparent Reasoning Architecture for Credible Execution), a decision framework that ensures every autonomous action can be traced back to sensor evidence through documented causal chains. The framework organizes decision-making into four auditable layers: Semantic Perception for evidence-grounded entity recognition, Belief Reasoning for probabilistic state estimation with causal graphs, Action Synthesis for constraint-aware planning with counterfactual documentation, and Execution Verification for compliance monitoring. TRACE is model-agnostic yet designed to integrate learning-based perception modules (CNNs, transformers) while preserving decision-level auditability. We evaluate the framework using three objective metrics: Evidence Traceability (sensor-to-decision linkage), Decision Reconstructability (post-hoc analysis capability), and Temporal Continuity (audit trail completeness). Experimental evaluation on warehouse robot navigation demonstrates that TRACE achieves 98.6% evidence traceability, 99.0% temporal continuity, and 98.1% decision reconstructability across 500 simulated decision cycles. Post-hoc methods like LIME provide feature attributions but lack the artifact structure needed for decision-level reconstruction. The framework addresses EU AI Act requirements for high-risk system transparency and contributes to Explainable AI for safety-critical autonomous systems.
cs.AI / 65 / 2609.02277
Auditory Illusion Benchmark for Large Audio Language Models
Abstract
Perceptual illusions have long served as crucial probes into human cognition, revealing biases and limitations of perception. In the auditory domain, such illusions provide a unique lens for testing whether Large Audio Language Models (LALMs) replicate human perceptual tendencies. Despite their importance, most benchmarks focus on visual illusions or general audio tasks, leaving auditory illusions underexplored. To this end, we present AIB, the first auditory illusion benchmark for LALMs, covering ten representative illusions across music, sound, and speech, each annotated for the presence of knowledge-based priors. Our methodology pairs model evaluation with controlled human listening studies, enabling direct comparison of responses. Results show systematic differences: while most LALMs remain signal-faithful on low-level acoustic illusions, several exhibit more human-like responses when linguistic or musical priors are involved, although no model matches the human perceptual profile. These findings highlight the current limitations of LALMs as cognitive models. By establishing auditory illusions as a rigorous testbed, our work offers a new perspective for probing neural black-box models and advancing understanding of auditory cognition. AIB is publicly available at https://github.com/gillosae/aib.
cs.AI / 66 / 2609.02797
Dutch Books for Language Models
Abstract
People increasingly use language models to support life decisions. Many such decisions involve a probabilistic forecast: How likely is a major life event, a natural disaster, or an economic outcome? Users of language models may implicitly trust that these forecasts fall out of a coherent world model. In this paper, we evaluate the coherence of language model probabilistic forecasts through a procedure that builds on a theorem due to de Finetti. We elicit forecasts from language models across events generated from stock returns data. We then use linear programs to compute the largest Dutch-book profit - the profit an arbitrageur could guarantee by betting against model-generated probabilities - which we use as a measure of incoherence. Our procedure does not require outcome labels, so we can evaluate coherence even in settings where outcomes are not observed or have not yet resolved. We find substantial evidence of incoherence in language model forecasts. Such incoherence increases when there are richer logical relationships between events, and irrelevant contextual details can increase incoherence by an order of magnitude. We conclude by discussing how alternative training strategies may improve probabilistic coherence.
cs.AI / 67 / 2609.02746
HiPoly: a hierarchical polymer-native AI framework for property prediction and generative design
Abstract
Polymeric materials are central to modern technologies, with applications ranging from energy to health and transportation. Although AI has made significant advances in materials discovery, the hierarchical structure of polymers across multiple length scales makes them inherently difficult to represent in a unified and physically meaningful way. Here we introduce HiPoly, a polymer-native AI framework that processes complete polymer descriptions through a three-level hierarchical graph architecture built on the G2RINS representation. HiPoly encodes stochastic inter-monomer connectivity, composition, and molecular weight directly within its architecture, using physically motivated design principles that mirror the multi-scale nature of polymeric systems. The framework establishes an end-to-end AI-driven workflow from experimental formulation data to property prediction, generative molecular design, and physics-based validation through molecular simulations, all unified by a single polymer representation. We demonstrate state-of-the-art prediction accuracy for thermophysical properties of multi-component polymer systems, with ablation studies confirming that each hierarchical design choice contributes independently to model performance. As an example, the generative design pathway is applied here to the discovery of sustainable alternatives to persistent fluorinated polymers, where it is possible to identify and independently validate PFAS-free candidates with target surface-energy properties. This work demonstrates how polymer-native AI can accelerate discovery by linking representation, prediction, and design across complex polymer chemistries.
cs.AI / 68 / 2609.02344
Subcellularly Resolved Single-Cell Embedding Learning with Transcriptomic data, Protein Structure and Localization Information
Abstract
Existing cell embedding methods predominantly rely on transcriptomic or proteomic measurements and represent each cell as a holistic entity, thereby overlooking the subcellular localization of individual molecules. Moreover, they rarely incorporate protein structural information, despite its fundamental role in determining molecular interactions and functions. In this work, we propose a multimodal framework for learning subcellularly resolved cell embeddings by jointly leveraging RNA expression profiles, protein sequence representations, and protein structural information. Specifically, we employ a cross-attention architecture to integrate transcriptomic, sequence, and structural modalities and model their interactions within distinct subcellular compartments. The resulting embeddings represent each cell through its fine-grained subcellular organization, capturing both molecular expression patterns and the functional properties of the associated proteins. By learning cell representations at subcellular resolution, our framework preserves spatially organized biological information while integrating complementary signals across multiple molecular levels. To the best of our knowledge, this is the first framework that produces subcellularly resolved cell embeddings by jointly incorporating transcriptomic information, protein sequence representations, and protein structural knowledge within a unified cross-modal learning paradigm.
cs.AI / 69 / 2609.02196
Schrödinger Bridges on Lie Group Manifolds for Probabilistic Intrinsic Generation
Abstract
Generative modeling directly on geometric manifolds can avoid errors introduced by flattening non-Euclidean data, repeated ambient projection, and coordinate inconsistency in Euclidean representations. Schrodinger bridges provide a probabilistic generative framework for entropy-regularized transport between prescribed endpoint distributions. We study Schrodinger bridges for kinetic dynamics on Lie group manifolds with state X_t = (g_t, xi_t) in G x g, allowing endpoint observations to constrain only the variables that are actually measured. In particular, the entropy projection determines the conditional law of the unobserved endpoint velocities. For the same observed endpoint bridge, we develop two computational realizations: Wrapped-Kernel Bridge Calibration (WKBC) uses an explicit periodized kinetic kernel on compact Abelian groups, whereas Reciprocal Conditional-Control Bridge Matching (RCCBM) handles compact non-Abelian groups through two-sided endpoint calibration and mollified conditional-control matching. The canonical teacher-mixture path law is itself a Markov reciprocal law, so forward generation uses a calibrated initial law and one learned Doob controller. Moreover, we establish a modular error bound in the bounded-Lipschitz path metric that provides a clean separation of errors due to endpoints, control regression, initialization, discretization, and related approximations. Experiments on multiple Lie group manifold datasets validate the feasibility and consistency of our proposed method, covering protein and RNA torsions, SO(3), U(n), and the Protein Conformational Transition Pathway Generation task using mdCATH trajectories in a compact reduced representation. The source code is publicly available at https://github.com/cafferyzhang12/Schr-dinger_Bridge_on_LieGroup.
机器学习 (cs.LG)
84
cs.LG / 1 / 2609.01786
Ten Architectures, One Error: Shared Failure Modes in Hyperspectral Classification under Spatially Disjoint Evaluation
Abstract
Hyperspectral image classification still relies heavily on random pixel splits within a single scene. The Salinas dataset, randomly split, is among the most widely used datasets for comparing different architectures. However, under a random split method, a large fraction of test pixels fall immediately adjacent to a training pixel, which inflates reported accuracy. This work introduces a leakage-free evaluation protocol linking spatial separation to the model's receptive field. Applying this protocol across ten different architectures, including classical, spectral, spectral-spatial, transformer, vision-backbone, and state-space families, shows that Macro-F1 drops by 0.147 on average and model rankings change by as many as five places. Furthermore, leakage-free evaluation limits which architectures can be tested on a given benchmark. Since each partition supports patches only within a finite radius, reporting this radius alongside the receptive field is essential for fair comparison. In addition, this study reveals that all ten architectures misclassify largely the same pixels, pointing to a spectral ambiguity in the data that none of them resolves.
cs.LG / 2 / 2609.01987
Morphology signal in whole slide image foundation models can automatically triage slides
Abstract
Patient exams in the cancer diagnosis and staging process typically generate several whole slide images (WSIs). One of the initial steps in training models on WSI data is identifying one or a few slides containing tumor or other diagnostic biomarkers necessary for downstream prediction tasks such as estimating recurrence risk or progression-free survival. This step requires tedious manual curation by experienced pathologists. Many published datasets make the artificial assumption of 1 slide per patient. Alternatively, all slides per patient may be used for model training, which may dilute the signal from the few slides containing tumor or other relevant information. In this paper, we present a pipeline to overcome these challenges using publicly available WSI foundation models (FMs). Our evaluations show that ranking WSIs based on predictions from zero-shot classification using WSI FMs accurately identifies slides with the most tumor, indicating that WSI FMs contain sufficient morphology signal to automatically triage slides. We also present a formulation for ranked evaluation to benchmark FM performance in slide triage. We show, on multiple datasets, that tumor slides are identified in the top-2 ranked slides for patients with up to 43 slides.
cs.LG / 3 / 2609.02510
Orthogonal Ensembles and Tested Explanations for Performer-Independent Body-Motion Emotion Recognition
Abstract
We study body-only, 12-class acted-emotion classification from skeleton motion under leave-performer-out (LPO) evaluation, a hard, underdetermined setting: chance is 8.3%, and a protocol-matched reproduced STGCN++ baseline reaches only 25.73 +/- 4.03% Macro-F1. We show that reliable gains come not from a new architecture but from combining eleven models with orthogonal error modes: under 10-fold LPO cross-validation on the labeled training performers, an equal-weight logit-mean ensemble reaches 36.80 +/- 4.00% per-fold Macro-F1, a protocol-matched +11.07 pp (+43% relative) over the same-split reproduced baseline. Our central contribution is a tested explanation suite: for a strong ensemble member, part-masking and counterfactual edits show (rather than assert) that its decisions depend on motion-grounded body-region evidence, and this region saliency aligns with rule-based Laban Movement Analysis (LMA) attributes far more than with classical kinematics: region-level saliency-LMA Spearman rho = +0.500 versus +0.033, roughly 15x, and the alignment holds for the submitted 11-way ensemble itself at rho = +0.517; the audit is post hoc and needs no retraining. The same suite faithfully reports a negative: within-window temporal saliency is diffuse rather than localized.
cs.LG / 4 / 2609.02499
Training seeds and model-selection stability in recommender-system evaluation
Abstract
Recommender-system experiments often rely on a single random training seed, assuming that run-to-run stochasticity has limited impact on evaluation conclusions. This assumption is risky, as a training seed may influence several algorithm-dependent mechanisms, including parameter initialization, mini-batch ordering, dropout, masking, latent sampling, and training-time negative sampling. We examine this assumption by fixing the data partition and varying the training seed across hyperparameter configurations. We analyze seed effects at three levels: user-level metric sensitivity, validation-based model selection and recommendation-list agreement. Results show that seed variation is often detectable. Its impact depends on whether configurations are clearly separated, whether validation results transfer to test, and whether similar scores lead to similar top-$k$ lists. Findings suggest that reporting single-seed results can overstate the stability of recommender system evaluation, and that training seeds should be treated as part of the evaluation protocol rather than as incidental implementation noise.
cs.LG / 5 / 2609.01729
RecKAN: Kolmogorov-Arnold Networks with a Learnable Recursive Polynomial Basis
Abstract
Kolmogorov--Arnold Networks (KANs) replace the fixed scalar weights of a standard network with learnable univariate functions on each edge, but existing variants still fix the \emph{basis} that those functions are built from: B-splines, Chebyshev polynomials, wavelets, or Jacobi polynomials, and learn only the combination weights over it. We introduce RecKAN, which instead defines the basis itself by a second order polynomial recurrence, $R_{n+1}(x) = (ax^2+bx+c)R_n(x) + (dx+e)R_{n-1}(x)$, whose five coefficients are learned jointly with the network. We show this recurrence recovers several classical polynomial families including both kinds of Chebyshev polynomials, Fibonacci, Pell, and Jacobsthal polynomials as special cases, and prove that its degree grows linearly in $n$ exactly on the sub-family containing all of them, giving a concrete sense in which the learned basis can move beyond any fixed classical choice. Across multiple benchmark datasets spanning image, text, biomedical time series classification, and time series forecasting, RecKAN outperforms three parameter-matched KAN baselines (Chebyshev, Jacobi, and spline based) on all classification tasks and achieves the lowest MSE on the ETTh1 forecasting benchmark. Additionally, when used as a classifier head with a convolutional backbone, RecKAN achieves higher accuracy than standard MLP heads on Fashion MNIST, CIFAR-10, and SVHN. On a synthetic function fitting benchmark it tracks a sharply oscillatory target that a parameter comparable MLP under fits. We further show that the learned recurrence coefficients are interpretable: on the task requiring the most local structure, training moves the basis away from the linear degree growth regime that contains every classical family we identify, consistent with our theoretical analysis of what that structural shift enables.
cs.LG / 6 / 2609.01765
Toward Explainable and Policy-Aware AI for Carbon Credit Price Prediction: A Research Framework for Emerging Carbon Markets
Abstract
Carbon markets put a price on emissions, yet that price remains hard to forecast. Work in this area clusters on the EU and Chinese schemes, compresses regulatory text into a sentiment score, and reports accuracy without calibration or explanation stability. We distil ten recurring gaps into an impact-feasibility matrix and propose EPA-CarbonNet, a six-layer architecture that fuses market series with policy text by cross-attention and calibrated intervals alongside policy-attributed explanations. We then build and test it on eleven years of daily S and P carbon index data. The findings are largely negative, and reported as measured: a random walk beats the model on five-day RMSE (0.0365 against 0.0475), SHAP rankings agree at rho = 0.54 across resampled backgrounds, and policy attention never coincides with documented regulatory events. Directional accuracy, at 58.6 percent, leads every baseline. Code, data documentation and all result artifacts are available at https://github.com/Kimalice/Toward-Explainable-and-Policy-Aware-AI-for-Carbon-Credit-Price-Prediction
cs.LG / 7 / 2609.01768
Emergence of Fibrations, Compression, and Symmetry Breaking in Artificial Neural Networks
Abstract
Artificial neural networks are often regarded as powerful yet opaque black boxes. Here, we demonstrate that learning in deep neural networks generates local symmetries known in graph theory as fibrations and coverings. We prove that covering symmetries are stable attractors of stochastic gradient descent. Consistent with this theory, we report the emergence of covering symmetries across major network architectures, including multilayer, convolutional, recurrent, and transformer networks. Exploiting these symmetries enables drastic model compression - reducing networks to 17% of their original size without sacrificing performance. Furthermore, controlled breaking of covering symmetry overcomes the loss of plasticity, achieving state-of-the-art performance in continual learning. The theoretical results provide a new foundation for AI systems based on symmetries that convert black boxes into interpretable colored graphs and enable more efficient inference and lifelong learning.
cs.LG / 8 / 2609.01802
D-FROST: Decentralized Federated pRompt-tuning via Optimal tranSporT for Non-IID and Imbalanced Data
Abstract
Prompt tuning provides a parameter-efficient way to adapt foundation models (FMs) by freezing the pretrained backbone and updating only a small set of learnable prompts. This property makes prompt tuning especially suitable for decentralized federated learning (DFL), where exchanging full-model updates can be prohibitively expensive. However, prompt tuning in DFL introduces new challenges. Prompt sets learned from heterogeneous local data may not be index-wise aligned, making standard decentralized averaging unsuitable. In addition, the algorithm should be theoretically guaranteed to achieve consensus and make progress toward the shared objective. In this work, we provide the first study of prompt tuning in DFL. We formulate decentralized prompt tuning as a Wasserstein-based optimization problem over prompt measures, which captures the set-valued structure of prompts. We then propose D-FROST, an optimal-transport-based (OT-based) decentralized prompt-tuning algorithm that merges neighborhood prompts into compact representative prompt sets through transportation-based matching. We further analyze D-FROST by bounding the Wasserstein consensus error across clients, and establishing convergence of the network-level prompt barycenter to a neighborhood of stationarity. Experiments under heterogeneous client data demonstrate the effectiveness of D-FROST for decentralized prompt tuning.
cs.LG / 9 / 2609.01839
Import What You Need: Learning When and How to Augment EHR Graphs with External Knowledge
Abstract
Longitudinal prediction from electronic health records (EHRs) is limited by the sparsity and irregularity in patient trajectories, and knowledge augmentation with external knowledge graphs (KGs) offers a promising way to alleviate these issues. However, most existing methods perform fixed, context-agnostic topology augmentation by adding the same KG nodes and edges regardless of a patient's evolving state. We propose ReTA, a Reinforcement learning-based dynamic Topology Augmentation framework that casts KG import as a per-visit, budget-aware policy. ReTA first constructs an offline refined pool of KG-grounded templates, then learns a policy to select one augment action per visit from three options: Soft Import, which enriches node features without modifying graph topology, Hard Import, which grafts a compact KG subgraph onto the visit graph to create message-passing shortcuts, and Skip, which leaves the visit unaugmented when the base encoder is already confident. To stabilize learning, ReTA employs a decoupled encoder that processes semantic and structural signals in separate channels and fuses them via adaptive gating. Experiments on MIMIC-III and MIMIC-IV across diagnosis prediction, mortality, and readmission show that ReTA consistently outperforms strong baselines while remaining efficient, transfers across datasets and knowledge graphs, and yields interpretable augmentation patterns. The robust gains under sparse supervision highlight the advantage of ReTA's dynamic decision to import knowledge, boosting accuracy while curbing costs.
cs.LG / 10 / 2609.01896
OutageDiT: A Generative Foundation Model for Power Outage Forecasting and Scenario Simulation
Abstract
Power-outage planning requires scenarios before an event occurs. These scenarios must represent uncertainty in magnitude, timing, and duration while preserving temporal dependence. However, severe events are rare, and data from any single region contain few examples of extreme outage and restoration patterns. To address this challenge, we introduce OutageDiT, a foundation model for generating seven-day outage trajectories at quarter-hour resolution, trained on outage and weather records across the United States. Specifically, a condition encoder processes the historical context and known future covariates once per forecast, and a shallow flow decoder reuses the resulting horizon-aligned states to generate complete trajectories. The resulting samples support point forecasting, uncertainty quantification, and conditional event simulation within one deep generative model. Across outage forecasting benchmarks, OutageDiT improves forecast accuracy and scenario quality over strong baselines and supports zero-shot transfer to held-out regions. Together, these results position conditional outage simulation as a bridge from outage forecasting to operational planning under uncertainty.
cs.LG / 11 / 2609.01925
CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing
Abstract
The attention prefilling phase of long-context LLM inference scales quadratically, making self-attention a severe computational bottleneck. Traditional sparse attention methods mitigate this through fixed patterns or offline profiling, but lack the flexibility to adapt to input-dependent attention structure. Recent dynamic methods address this by routing heads to sparse patterns in real-time, but rely on indirect routing proxies with overhead and budget allocation mechanisms that overlook the post-softmax mass hierarchy. We present CRISP (Cliff-awaRe Input-adaptive Sparse Prefilling), which identifies and addresses two structural challenges in this dynamic routing paradigm. First, we show that the routing decision can be read directly off the structure of the proxy attention map. We replace the Jensen-Shannon Divergence (JSD) routing with C_struct, a structural proxy that measures mass at Vertical-Slash compatible positions and reproduces JSD's routing decisions while eliminating both the pooled matmul and subsequent KL divergence overhead. Second, we formalize the post-softmax mass cliff and demonstrate theoretically that strictly cumulative coverage thresholds accumulate O(n) background noise at long contexts. CRISP navigates this via a sink-aware threshold grounded in the noise floor. Empirically, across InfiniteBench, RULER and LongBench on two model families, CRISP is the strongest sparse method overall and matches or exceeds exact dense attention on retrieval-heavy benchmarks, recovering up to +28.0 pp on retrieval tasks over baselines and achieving up to a 5.30x attention speedup at 512k tokens, driven primarily by our O(n) noise elimination during selection while preserving structural integrity.
cs.LG / 12 / 2609.01933
OR-Transformer: Scaling Real-Time Decision-Making to 1,000 Items
Abstract
Modern supply chain operations can require coordinating replenishment across thousands of heterogeneous items under correlated stochastic demand, heterogeneous lead times, and shared fixed ordering costs, yielding observation spaces exceeding $10^4$ dimensions. At this scale, rolling-horizon stochastic mixed-integer linear programs (MILPs) become prohibitively slow, while standard reinforcement learning (RL) methods face increasingly challenging credit assignment in high-dimensional action spaces. We introduce OR-Transformer, a deep reinforcement learning framework for joint replenishment under stochastic demand, with an item-permutation-equivariant Transformer architecture and pathwise-gradient training through the inventory dynamics. Across problem sizes up to 1,024 inventory items, OR-Transformer increasingly outperforms learning-based and rolling-horizon MILP baselines as scale grows. It also reduces online decision-making time by over 4 million times relative to MILP solvers, enabling real-time, large-scale deep RL in supply chain operations.
cs.LG / 13 / 2609.01942
Refining Heuristic-Based Bitcoin Address Clustering with Graph Neural Networks
Abstract
Bitcoin's pseudonymous nature makes it challenging to analyze user-level activity, since a single user may control multiple identifiers (addresses). Existing heuristic-based methods attempt to identify addresses belonging to the same user, but they often produce flat cluster assignments with limited modularity and are prone to errors such as merging different users together. In this work, we propose a method for refining heuristic-obtained clusters by grounding our clustering on contrastive embeddings yielded by graph neural networks. Our contributions are threefold: (i) we release a publicly available dataset of Bitcoin transaction graphs containing a substantial number of clusters; (ii) we propose a methodology for learning address embeddings consistent with heuristics, and back it up with theoretical guiding intuitions; (iii) through hierarchical clustering, we enable a finer analysis of heuristic clusters and provide a quantitative criterion for flagging suspicious merges.
cs.LG / 14 / 2609.01947
On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers
Abstract
Compact instruction-following rerankers are attractive for deployment, but conventional distillation pipelines typically train students by offline imitation of teacher outputs on a fixed set of examples, constraining supervision to the teacher's observed ranking space. We revisit reranker distillation through the lens of reinforcement learning. We propose a two-stage framework combining off-policy teacher optimization with on-policy student distillation. In Stage 1, a 4B teacher reranker is strengthened with off-policy GRPO using LLM-judge feedback on 88K instruction-following examples. In Stage 2, a compact 1B student samples rankings from its own policy and receives soft teacher-derived rewards on those rankings, coupling student exploration with knowledge transfer. Our strongest gains appear under distribution shift. On MAIR-11, the original 11-subset, 869-query evaluation, the proposed student reaches 0.7670 nDCG@6, outperforming offline listwise KD by +4.6 points. Controlled comparisons against offline pairwise RankNet KD and on-policy GKD show that neither changing the offline distillation objective nor moving teacher-distribution matching on-policy reproduces the performance of reward-based on-policy distillation over student-sampled rankings. The advantage persists on MAIR-Full: across all 126 tasks and 9,356 queries, the proposed method obtains the highest task-macro point estimates among the evaluated distillation variants, reaching 0.6808 nDCG@6 and 0.7865 MRR@6. It also exceeds two released 7B RL-trained rerankers on the comparable MAIR-11 evaluation, while the same Stage 2 training procedure consistently improves three architecturally distinct alternative student backbones. On the 9,861-query validation benchmark, the resulting 1B reranker achieves 0.7624 nDCG@6 while providing a favorable quality-efficiency tradeoff relative to larger alternatives.
cs.LG / 15 / 2609.01952
Convergence Theory of Knowledge Distillation in Asynchronous P2P Gossip Learning Network
Abstract
Decentralized, serverless learning increasingly connects devices running different architectures, where the standard tool, decentralized SGD, is undefined as models with different parameter counts cannot be averaged. Knowledge distillation (KD) exchanges soft predictions rather than weights and sidesteps this obstacle, yet convergence theory for fully decentralized, asynchronous peer-to-peer (P2P) KD is lacking. We provide one, relocating consensus from parameter space to function (output) space: a KD event is a geometric contraction operator in logit space on the peers' predictive distributions, which we analyse in the Hilbert space of predictions on a reference measure. Under standard smoothness/variance assumptions and two realizability assumptions, one bridging parameter SGD to the functional step and one controlling restricted task/KD alignment, the time-averaged functional stationarity and function-space disagreement converge at rate $O(1/(ηT))$ to an $O(η)+O(B_f^2)+O(ζ_f^2)$ neighbourhood. Here $B_f$ is the distance from the task optimum to the peers' reachable classes and $ζ_f$ measures persistent local-task heterogeneity. Across homogeneous, width-heterogeneous, and mixed-family networks of the experiments, KD contracts function disagreement by $40-61\times$, while isolated training does not. The sampled stationarity diagnostic has late transient exponents $0.99-1.90$ on the shared-skeleton main runs, and the four-point step-size sweep exhibits the predicted transient: neighbourhood tradeoff.
cs.LG / 16 / 2609.01956
FlashKAN: B-Spline KANs via Truncated Power Form
Abstract
Kolmogorov-Arnold Networks (KANs) place learnable B-spline activations on network edges rather than fixed activations on nodes. The standard Cox-de Boor recursion evaluates these activations through k sequential passes for degree-k splines, consuming over 90% of forward-pass time. FlashKAN replaces this recursion with the truncated power form, a classical result from approximation theory that expresses each uniform cubic B-spline as five (x)_+^3 terms at shifted knot positions. This paper makes three contributions: (1) a torch.compile-fused implementation that collapses these operations into a single GPU kernel, eliminating all recursion, span lookup, and scatter-gather operations; (2) a bounded-coordinate stabilization that clamps the normalized input to [0, k+1], preventing the catastrophic cancellation that historically motivated the Cox-de Boor recursion; and (3) a production-ready, open-source package (pip install flashkan) that serves as a drop-in replacement for existing KAN layers.
cs.LG / 17 / 2609.01967
A Unified Particle Filter LSTM for Data-Driven Process Simulation
Abstract
Data-driven process simulation aims to generate realistic case trajectories from historical event logs without requiring an explicitly specified model of the underlying dynamics. Deep sequence models can capture complex temporal dependencies through next-activity probabilities and conditional time distributions. However, event logs provide only a partial view of the underlying process state, often recording activity completions without the corresponding service-start times. Consequently, the same observed process history may be consistent with multiple plausible latent process conditions, whereas standard recurrent models compress each process prefix into a single deterministic recurrent state. We propose a Unified Particle Filter LSTM (Unified PF-LSTM) that maintains and sequentially updates a weighted set of recurrent-state hypotheses. We summarize this particle belief using its weighted mean and learned features based on the moment-generating function. The resulting representation is used to predict a categorical distribution over the next activity and conditional quantiles of the current activity's sojourn time. The framework is trained end-to-end from event-log data and evaluated on three real-world emergency department datasets. The results show that the proposed framework consistently outperforms the considered data-driven baselines in reproducing routing, duration, and system-level behavior across all datasets, with particularly strong gains in settings where complex process dynamics are only partially reflected in the available event logs.
cs.LG / 18 / 2609.01991
CAHR-Net: Condition-Adaptive Hysteresis Reconstruction for Compact and Interpretable Magnetic Core Loss Modeling
Abstract
Magnetic core loss originates in the hysteresis loop: the energy dissipated per excitation cycle equals the loop area, and frequency, temperature, and waveform shape set the loss by reshaping the loop geometry. Most existing models let these conditions act only on a terminal scalar - empirical equations fold them into fitted exponents, and data-driven predictors append them to encoded features - so no intermediate hysteresis representation remains for the conditions to reshape. This paper proposes CAHR-Net, a condition-adaptive hysteresis reconstruction network that injects the operating conditions where they physically act. It preserves the interpretable chain from flux density waveform to magnetic field reconstruction, loop-area integration, and power loss estimation, and uses feature-wise linear modulation to inject frequency, temperature, and waveform statistics into the intermediate reconstruction representation. A matched large-batch training protocol based on AdamW, cosine scheduling, and a staged reconstruction-to-power-loss objective is also reported, because the modulation pathway takes effect only within it. On the MagNet final A-E material protocol, CAHR-Net attains an average p95 relative error of 6.89% with only 1874 parameters, the lowest among all compared methods, together with a lower worst-material p95 than the strongest black-box solution at about 48x fewer parameters; it reduces the average p95 of the physical reconstruction backbone from 7.47% to 6.89% and the p95 of material D, the most difficult material, from 16.40% to 14.87%. Ablation and condition-slice analyses attribute the improvement to the coupling of physical loop reconstruction, structured condition modulation, and the matched optimization trajectory.
cs.LG / 19 / 2609.02006
Train What You Deploy: Closing the MLP Reachability Gap in Low-Rank Clone Distillation
Abstract
A compressed student has two shapes that need not agree: the weight it deploys at inference and the weight family its training can reach. We show that a state-of-the-art weight-inheritance distiller, Low-Rank Clone (LRC), deploys a full-width student MLP but ties training to a teacher-induced slice, leaving 62.5-81.4% of each deployed matrix's independent linear degrees of freedom unreachable-paid for at inference, never trainable. Our principle is one line: train what you deploy. From the identical LRC warm start, we make the training object the entire deployed matrix, with no change in deployed shape, deployed parameter count, or inference FLOPs, via two mergeable realizations (Dense-LRC and CORE-LRC) that both collapse to one deployed weight. This recovers stranded capacity: taking the stronger realization per teacher, +2.36/+2.71/+10.45 Avg9 over matched-budget plain-LRC baselines across three teachers (Llama3.2-3B, Llama3.1-8B, Qwen2.5-3B), with the largest gain on the widest teacher (Qwen), where it reaches the original recipe's approx. 20B-token accuracy at 10B tokens (2x token efficiency); there the strictly same-lineage arm still recovers +6.39, the fully controlled figure. Controls strongly support attributing the gain to the enlarged reachable set, rather than to added parameters or the recipe. From approx. 10B distillation tokens plus a short SFT, a half-parameter 1.5B student matches its approx. 9T-token teacher's 9-task macro-average, within evaluation noise and with a residual MMLU deficit, and a 2.7B student beats Meta's own official compression of Llama3.1-8B at ~900x fewer compression tokens (a token count under unmatched recipes, not a compute claim). All results are from single-seed runs on the LRC backbone.
cs.LG / 20 / 2609.02018
Source-Free Class Relearning: Diagnosing Forgetting in Class Unlearning
Abstract
Class unlearning aims to remove a model's ability to recognize designated forget classes while preserving performance on retain classes. However, low forget accuracy after unlearning does not necessarily mean the class structure has been erased. Approximate unlearning methods can alter classifier decision boundaries while leaving recoverable structure in the representation. Prior work has shown that forget classes can be recovered, but existing approaches require real forget or retain samples, auxiliary data, or reference checkpoints. We study class relearning in a strictly source-free setting, asking whether a forget class can be recovered through a classifier-head update using only the unlearned model. Our approach rests on a theoretical analysis establishing a sufficient alignment condition under which a single gradient step on a synthetic probe set increases the expected logit margin of the forget class. Building on this, we propose a white-box Source-Free Relearning Audit (SFRA), which generates candidate embeddings in representation space and uses model-guided confidence filtering to construct high-confidence retain probes and low-confidence boundary-adjacent probes that are relabelled as the forget class. Gaussian sampling and Softmax confidence are used by default, while ablations with alternative proposal distributions and uncertainty criteria show that recoverability is not specific to these choices. To quantify recoverability, we introduce the Relearning Score (RS), which jointly measures forget-class recovery and retain-accuracy preservation, and report class-matched $Δ$RS relative to a retrained reference. Experiments on CIFAR-10, CIFAR-100, and TinyImageNet with ResNet-18, ViT-B/16, and Swin-T show that several unlearning methods exhibit substantial source-free recoverability, and that for a subset of methods this recoverability exceeds the matched retrained reference.
cs.LG / 21 / 2609.02049
The Dynamics of Continuous Mixture Collapse in Language Models
Abstract
LLMs latent-state reasoning methods replace discrete intermediate tokens with continuous states, such as weighted mixtures of token embeddings, to retain multiple possible reasoning directions rather than committing to one. Yet pretrained language models often fail to preserve these mixtures. We study why through a combination of theoretical analysis and controlled empirical investigations on a variety of models. We identify three independent, distinct sources of failure. First, transformer architectures already distort mixture geometry, and training substantially amplifies this effect. Moreover, the failure can occur even if the model transports mixtures perfectly linearly: the softmax readout and autoregressive feedback form a dynamical system that either amplifies small differences until one component of the mixture dominates or contracts different mixtures until they become indistinguishable. We verify this theoretical prediction empirically: the observed transition between contraction and amplification occurs near the theoretical threshold derived by our analysis, and pretrained-model rollouts lie predominantly on the amplifying side. Finally, we generalize to mixtures of many components and show that exact preservation generally requires context-dependent correction, whose required dimensionality can grow with the number of components.
cs.LG / 22 / 2609.02083
XMerge: Cross-Axis Selection and Reconstructive Layer Merging for LLM Depth Compression
Abstract
Removing complete transformer layers preserves a standard serving architecture, but existing depth-compression methods can lose substantial quality, and the loss varies unpredictably across models. We introduce XMerge, a post-training method with two components. Cross-axis selection identifies a block with low relative-magnitude and angular hidden-state change, and local boundary reconstruction re-fits the adjacent surviving block to match the original two-block output. XMerge uses no task labels or end-to-end fine-tuning, and it introduces neither architectural changes nor additional inference-time parameters. Across seven Llama and Qwen backbones (0.5B-8B), five published baselines, and three layer-reduction levels, its advantage over baselines is largest at the most aggressive removal: at k=4 it ranks first on six of seven backbones on CORE (a 22-task aggregate) and, separately, on six of seven on MMLU (five of seven on both at once), while avoiding the large perplexity increases of several competing operators. In a task-level bootstrap, the 95% confidence intervals for the three largest CORE margins exclude zero; the remaining margins are consistent with ties. Across the 14 (model, regime) cells it is also the only evaluated operator that never collapses, ranking top-2 in both zero-shot and in-context regimes; on a first calibration probe (one backbone) it is the best-calibrated operator. Ablations show that local reconstruction provides most of the gain, while cross-axis fusion helps when the two selection axes disagree. The additional construction cost is recovered through per-token decode savings after roughly tens of thousands of requests.
cs.LG / 23 / 2609.02085
TC-Next: Zero-Shot Multimodal Cyclone Forecasting
Abstract
We present TropicalCycloneNext (TC-Next), a multimodal deep learning model that forecasts tropical cyclone track and intensity at $6$-$24$ h leads by leveraging a foundation model's forecast fields of atmospheric kinematic and thermodynamic fields and GridSat infrared satellite imagery. Trained only on GraphCast forecasts over the Western Pacific (WP), yet reliant only on generic atmospheric variables, TC-Next on GraphCast lowers track error by $15$-$44\%$ and intensity error by a factor of $3$-$6$ relative to a conventional, rule-based tracker, TempestExtremes; applied without retraining to the forecast fields of Pangu-Weather and IFS HRES, it stays ahead of TempestExtremes on both. Applied zero-shot to the generic weather fields of WeatherNext Cyclones on the 2025 WP season, TC-Next attains lower intensity error at every lead time, and lower or comparable track error, compared to that model's specialized direct tracker in a deterministic comparison. Our ablation studies show that our multimodal model is able to utilize the additional modality to improve performance in tracking errors at every lead time and in intensity prediction at longer lead times.
cs.LG / 24 / 2609.02093
Compositional Spectral Prompts for LLM-based Online Time Series Forecasting
Abstract
To address the sequential and evolving nature of time series, the Online Time Series Forecasting (OTSF) task has been extensively studied in multiple domains. Existing research focuses on adapting to non-stationary environments by employing memory buffer-based retrieval strategies. However, we observe that such frameworks struggle with long-term adaptation and fail to generalize to unseen patterns. To this end, we introduce CoSPOT, an LLM-based online time series forecasting framework that leverages a pre-trained LLM as the backbone online forecaster, motivated by its strong few-shot capabilities. For efficient online adaptation, CoSPOT keeps the LLM frozen and employs compositional spectral prompts grounded in frequency-domain bases to guide the model with the overall distribution of the input, thereby substantially reducing the number of parameters updated during the online phase. Specifically, CoSPOT decomposes time series into frequency bases and composes the corresponding spectral basis prompts according to their amplitudes, allowing unseen patterns to be represented as new combinations of learned basis prompts. Our extensive experiments on real-world datasets demonstrate the superiority and practicality of CoSPOT across challenging online scenarios, including extended online phases and cross-dataset settings with substantial distribution shifts. Our code is available at https://github.com/seungyoon-Choi/CoSPOT.
cs.LG / 25 / 2609.02101
Federated LoRA Adaptation of BiomedCLIP Across Four International Chest X-Ray Cohorts
Abstract
Federated learning (FL) lets institutions train a shared model without exchanging data, and Low-Rank Adaptation (LoRA) makes this practical at scale by communicating only compact low-rank updates. Biomedical imaging is a compelling setting for this combination: patient data are archived behind privacy regulations, and institutions differ widely in scanners, protocols, and compute. Such heterogeneity raises the question of how federated LoRA updates should be aggregated, increasingly pressing as multimodal vision-language models become central to medical image analysis. We benchmark federated Parameter-efficient fine-tuning (PEFT) of BiomedCLIP for chest radiograph classification across four public cohorts on three continents (USA, Vietnam, Spain). Federated LoRA adaptation improves shared-class AUC on all four cohorts over the unadapted BiomedCLIP backbone (mean 0.687 to 0.802), showing that the gains come from federated adaptation rather than from the pretrained model's zero-shot ability. Relative to isolated single-cohort training, federation improves the weaker cohorts while largely preserving the strongest and approaches a centralized reference (0.812) that pools all data. The singular value decomposition (SVD)-based product-space aggregation introduced by FlexLoRA is essential to this gain (naive factor averaging drops mean AUC by 0.097), whereas a drift-correcting optimizer (FedProx) shows no benefit over FedAvg in our single-seed runs, consistent with LoRA's low-rank updates already limiting client drift. Biomedical vision-language models can thus be adapted collaboratively across heterogeneous, geographically distributed institutions without centralizing data.
cs.LG / 26 / 2609.02107
A Unified Rate-Distortion Perspective on Vector, Product, and Scalar Quantization
Abstract
Discrete visual tokenization, predominantly driven by vector, scalar, and product quantization, lacks a unified conceptual framework for understanding quantization tradeoffs. In this paper, we propose a unified rate--distortion perspective on modern discrete visual tokenization. By viewing quantization as lossy compression, we characterize the nominal fixed-length coding rate through token count and codebook size, and quantization error as the distortion. Within this framework, we resolve three central questions. First, we theoretically and empirically show that minimizing distortion, rather than maximizing codebook utilization, is the primary intrinsic objective for reconstruction fidelity, with a direct connection to the STE-induced gradient discrepancy. Second, we establish two critical fairness conditions for intrinsic quantization comparison: controlling latent feature statistics and enforcing identical coding rates. Third, under these conditions, we recover the VQ--PQ--SQ distortion hierarchy in modern visual tokenization and show empirically that modern VQ methods achieve the lowest distortion. This work provides a foundational rate--distortion reframing of modern discrete visual tokenization, resolves ambiguities in quantizer evaluation, and provides a controlled framework for isolating intrinsic quantization effectiveness under fixed-rate constraints.
cs.LG / 27 / 2609.02110
A Computational Comparison of Fourier Spectral Differentiation and Spatial Automatic Differentiation in Periodic Physics-Informed Neural Networks
Abstract
Physics-informed neural networks (PINNs) commonly evaluate the spatial derivatives appearing in partial differential equation residuals using automatic differentiation (AD), whose computational and memory costs can become substantial when multiple or high-order derivatives are required. We perform a controlled comparison of spatial AD and Fourier spectral differentiation in periodic physical-space PINNs. Within each paired experiment, the neural representation, temporal differentiation, optimizer, sampling procedure, and training schedule are held fixed, so that the two cases differ only in the spatial differentiation procedure. For the Fourier variant, network outputs are evaluated on a uniform periodic grid and transformed to Fourier space, where spatial derivatives are obtained through spectral multiplication and the same Fourier coefficients are reused across derivative orders. We compare the two procedures in standard PINNs for the Allen--Cahn and Korteweg--de Vries equations and in Causal PINNs for the Allen--Cahn, Korteweg--de Vries, and Kuramoto--Sivashinsky equations. Across these five equation--framework settings, Fourier differentiation yields mean paired end-to-end training speedups ranging from $2.90\times$ to $18.52\times$ and reduces peak allocated graphics processing unit (GPU) memory by $68.7\%$--$94.1\%$. The final relative $L_2$ errors remain of the same order, with neither differentiation procedure showing a consistent accuracy advantage. For the one-dimensional periodic benchmarks considered here, Fourier spectral differentiation therefore provides substantially lower training time and memory usage than spatial AD while retaining comparable solution error, at the cost of requiring a uniform structured spatial grid.
cs.LG / 28 / 2609.02126
Scalable Bayesian Optimization of Composite Functions for Image-Based Inverse Problems in Materials Characterization
Abstract
Estimating physical parameters from scientific images is a common inverse problem in materials characterization that often relies on expensive physics-based simulations. In electron microscopy, specimen thickness and crystal mistilt are critical parameters that govern how electrons scatter through the sample, and therefore the accuracy of any atomic-scale structure recovered from it. They are commonly inferred by matching experimental position-averaged convergent-beam electron diffraction (PACBED) patterns to simulated ones, but grid searches scale poorly and neural-network methods require extensive pretraining that may not transfer to new conditions. Here, we propose scalable Bayesian optimization of composite functions (SBOCF), a simulation-efficient method that exploits the known composite structure of the image-matching objective and the intermediate information contained in simulated images. By representing PACBED images with patch-level summaries and two correction terms, SBOCF preserves the original pixel-wise objective while reducing the number of modeled outputs from 24,649 to 11. Under a budget of 50 simulator evaluations, SBOCF outperformed standard Bayesian optimization with expected improvement on synthetic SrTiO3 benchmarks with thick and thin specimens, reducing the median final SSE by up to 290x in the thick-sample case. On experimental data, SBOCF produced parameter estimates consistent with previously reported values without task-specific pretraining. For a simulated mistilted specimen, using the SBOCF estimates in a downstream ptychographic reconstruction recovered sharp atoms that were otherwise blurred. These results establish SBOCF as a promising approach for inverse problems involving expensive simulators and high-dimensional structured outputs.
cs.LG / 29 / 2609.02145
Online Non-Monotone DR-Submodular Maximization Matching the Offline $0.401$ Factor
Abstract
We study online maximization of nonnegative, non-monotone DR-submodular functions over compact convex down-closed subsets of the $d$-dimensional unit cube. The best known constructive offline approximation factor is $0.401$ under the corresponding meta-solvability assumptions, whereas comparable adversarial online guarantees had remained at $1/e$. We show that this factor is also achievable online. In the post-decision full-information value-oracle model, our algorithm attains factor $0.401$ with sublinear approximate regret when oracle feedback is conditionally unbiased and bounded. The online algorithm does not run the offline construction on a changing objective. Instead, it replaces the offline objective-dependent box step by a weighted online learner that controls the required residual terms cumulatively. An exact asymmetric balance theorem preserves the offline coefficients despite adversarial variation. The direct implementation has $O(T^{3/4})$ regret and uses $O(dT^{1/4})$ oracle calls per round. More generally, for every $δ\in[0,1/4]$, batching gives $O(T^δ)$ calls per round and $O(T^{4/5-δ/5})$ regret, including a one-call $O(T^{4/5})$ endpoint. Under a positive-anchor condition, randomized blocking retains factor $0.401$ with $O(T^{5/6})$ one-point bandit regret.
cs.LG / 30 / 2609.02155
Exact Limits of Random Projections for Preserving Geometry: Distance Recovery, Nearest-Neighbor Rankings, and Covariance Shape in Gaussian Models
Abstract
The Johnson-Lindenstrauss (JL) lemma guarantees that a random projection of $n$ points to $m=O(\varepsilon^{-2}\log n)$ dimensions preserves pairwise squared distances within relative error $\varepsilon$ with high probability, and this dimension order is asymptotically optimal. In high dimensions, however, distances concentrate around a baseline while key geometric information lies in much smaller fluctuations. We show that the JL bound can therefore be uninformative about retained geometry: an independent Gaussian replacement map can satisfy it even though the replacement cloud is independent of the original data. We then ask how well any decoder can recover a feature $f(D)$ of a squared distance $D$ from a linear sketch. Under squared-error loss, the optimal decoder is conditional expectation, so recovery defines a linear operator whose singular values quantify feature recovery. For isotropic Gaussian data ($Σ=σ^2 I_d$), we diagonalize this operator in closed form. For fixed $k$ with $m,d-m\to\infty$, its $k$th singular value satisfies $\ell_k\approx(m/ d)^{k/2}$. This yields three sharp consequences. A rank-$m$ sketch retains at most an $m/d$ fraction of the variance of any feature of one squared distance. If $m\to\infty$ and $m/d\to0$, the expected Kendall correlation is $\frac{2}π\sqrt{m/d}(1+o(1))$; for fixed $q$, nearest- neighbor agreement tends to $1/q$. Yet one projection can satisfy the JL bound while mean Kendall correlation vanishes when $\log n\ll m\ll d$. After removing scale, Haar-averaged retained covariance-shape information is $(m/d)^2$. Thus JL distance preservation does not quantify the geometry available for comparison or inference.
cs.LG / 31 / 2609.02194
Learning the Constitutive Behavior of Materials via Neural Operators and Causal Attention: Case Studies in Plasticity and Damage
Abstract
Classical constitutive modeling of path-dependent inelastic materials relies on internal state variables whose evolution equations must be postulated based on domain knowledge and calibrated against experimental data. However, in many practical settings, the relevant internal variables are typically not measurable in experiments, and the constitutive response must be inferred entirely from measured strain-stress data without any prior knowledge of the material's internal state. We propose a data-driven constitutive modeling framework based on the concept of a material operator, which treats a deforming material as a functional mapping from its entire strain history to the corresponding stress response. In contrast to traditional autoregressive or recurrent formulations, the model is trained directly on full loading paths as function-to-function mappings, predicting complete stress trajectories in a single parallel forward pass. Temporal path dependence is enforced through a causally masked attention mechanism embedded within the operator, which restricts the model's attention to past material states while preserving computational parallelizability. Spectral convolutions provide discretization-invariant representations in the frequency domain, while causal attention captures highly adaptive, non-local history dependence. Furthermore, sinusoidal activation functions are used to resolve the strong nonlinear transitions inherent in inelastic regimes. The framework is evaluated across multidimensional, rate-independent material models exhibiting complex phenomena, with an emphasis on nonlinear plasticity and ductile damage accumulation. The results demonstrate accurate and robust predictions of irreversible deformation mechanisms while simultaneously achieving resolution invariance and excellent parallel efficiency.
cs.LG / 32 / 2609.02203
SMart: A Multi-source Multi-phase Time Series Representation Transfer Framework
Abstract
Time series representation learning (TSRL) has attracted growing research interests in recent years. Two recent explorations in TSRL are: i) exploiting a transformer-based framework to learn time series; ii) instead of using only the targeted dataset, borrowing time series from other datasets to to facilitate representation transfer. While these two explorations are shown effective, the self-supervised time series recovery task in (i) and the single-source dataset used in (ii) are technically simple and thus can be enhanced with new ideas. In this work, we propose a new TSRL framework, namely multi-source multi-phase time series representation transfer (SMart), which has two novel mechanisms to address the aforementioned deficiencies: 1) a multi-phase recurrence plots recovery task, in three alternative modes, for guiding the encoder to embed time series dynamics into the time series representation; and 2) a source dataset selector to select multiple suitable source datasets to supplement the original target dataset for pre-training the TSRL encoder. Experimental results show that SMart outperforms several state-of-the-art models for time series representation learning, classification and regression on both uni-variate and multi-variate time series datasets, reducing mean absolute error up to 19.5% for time series regression, and increasing average accuracy up to 1.34\% for time series classification.
cs.LG / 33 / 2609.02237
Recursive Value Learning for Long-Horizon Offline Goal-Conditioned RL
Abstract
Scaling offline goal-conditioned reinforcement learning (GCRL) to long-horizon tasks is difficult because (1) long-range value learning depends on shorter-range estimates that may still be inaccurate, and (2) max-based value backups can amplify overestimation through repeated propagation. We propose DCRL (Divide-and-Conquer RL), which recursively decomposes each trajectory segment into a balanced binary tree and trains the values from leaves to root. Each parent is therefore updated only after its children, using an exact factorization of the observed route rather than selecting among noisy alternatives. Since this objective learns values along demonstrated routes that are not necessarily optimal, DCRL jointly propagates values across trajectories to discover shorter routes. Thanks to the balanced binary tree, DCRL reduces worst-case bootstrap depth from linear to logarithmic, and this shorter dependency structure empirically corresponds to much slower error accumulation. Across diverse goal-reaching tasks, DCRL substantially outperforms prior flat offline GCRL methods, and on the five most challenging long-horizon OGBench tasks, it improves the best prior average score from 55 to 64, surpassing all flat and hierarchical baselines.
cs.LG / 34 / 2609.02241
Similarity-Aware Personalized Federated Learning in Heterogeneous Environments
Abstract
Federated Learning (FL) allows decentralized clients to train models collaboratively while preserving data privacy. However, distribution mismatch across clients often leads to poor global generalization and degraded local client-level performance. In such scenarios, some of the clients with their local models trained solely on local data may perform better than the globally learnt model, thus nullifying the benefits of collaborative federated learning. To address this, we propose SAPE-FL (Similarity-Aware Personalized Federated Learning), a novel personalization framework that anchors each client's model to both the global model and a similarity-weighted peer averaged model. By incorporating dynamic, client-specific regularization based on both model similarity and output similarity, SAPE-FL adaptively balances global knowledge transfer and peer collaboration while filtering out dissimilar clients. This dual anchoring mitigates negative transfer and enhances robustness in heterogeneous settings. We theoretically analyze our algorithm establishing its convergence guarantees and empirically show that SAPE-FL outperforms state-of-the-art methods under high statistical heterogeneity and low client data regimes.
cs.LG / 35 / 2609.02265
CAPTURE: Disentangling Preference Drift from Memory Poisoning in Personalized LLM Agents
Abstract
Personalized language agents use persistent memory to adapt to users over time, but the same mechanism creates an attack surface. When new information conflicts with stored preferences, an agent must distinguish genuine preference drift from temporary context shifts, ambiguity, or adversarial memory poisoning. We formulate this problem as a continuous-time partially observable decision process over a latent user state and show why rules based only on recency and provenance are insufficient. CAPTURE addresses this ambiguity with a neural differential-equation belief tracker, a multi-timescale memory ledger, uncertainty-triggered clarification, and counterfactual auditing of cited memories. On 480 held-out episodes from 96 users, CAPTURE achieves a 71.5% win rate, compared with 69.3% for an identically supervised baseline and 66.1% for the strongest heuristic baseline. It limits fixed-policy poisoning success to 11.5% while accepting 83.5% of genuine preference updates. Under an adaptive attacker with access to the released weights, attack success rises to 24.7%, exposing a real adaptation-security tradeoff. We further evaluate the frozen system zero-shot on an independently constructed benchmark and replay longitudinal interaction histories from 40 users collected over two to three weeks. These results suggest that modeling preference authenticity explicitly can improve both personalization and robustness in memory-augmented LLM agents.
cs.LG / 36 / 2609.02285
Entangled Representations Amplify Collateral Damage in Unlearning
Abstract
A long-held intuition in interpretability research is that representational entanglement, the sharing of structure between knowledge domains in a neural network, makes unlearning harder. While the intuition is widespread, it has never been directly tested in a controlled experiment. We present a way to do so: by repurposing Selective Gradient Masking (SGTM), we train a suite of six 254M-parameter language models on English Wikipedia with graded levels of disentanglement between biology and non-biology knowledge. Applying three standard unlearning methods to every model in the suite, we find that more disentangled models consistently achieve better retain-forget trade-offs: at a fixed level of forgetting, the most disentangled models incur roughly $4\times$ lower retain cost under two of the three methods, and $1.3\times$ lower under the third. Because our intervention changes only the model, not the data or the unlearning algorithm, this is direct evidence that representational entanglement is one of the causes of collateral damage in unlearning, as interpretability researchers have long suspected. A similar design could be used to test other structural claims from interpretability.
cs.LG / 37 / 2609.02304
Bayes-Optimal BER and AUC: Estimation and Evaluation of Estimators
Abstract
A fundamental quantity in machine learning is the optimal performance achievable by any model on a given task. Estimating this quantity allows us to distinguish the irreducible part of the error from a deficiency of the model, telling us how much room for improvement remains. Recent work has shown that the Bayes error, or equivalently the optimal accuracy, can be estimated from soft labels in binary classification. However, accuracy is often a poor summary of performance in settings with severe class imbalance or noisy annotations, where metrics such as the balanced error rate (BER) and the area under the ROC curve (AUC) are more appropriate. We address this gap with two complementary contributions. (i) Estimation. We propose soft-label-based estimators for the optimal BER and AUC. We first consider the clean setting in which true soft labels and the class prior are known, and then extend the estimators to a more realistic setting in which the class prior is unknown and the observed soft labels are corrupted by an unknown order-preserving transformation, possibly followed by additive noise. In the latter setting, we approximately recover the clean soft labels via isotonic regression with auxiliary hard labels, estimate the class prior with a clipped mean of the hard labels, and derive finite-sample error bounds for the resulting plug-in estimators. (ii) Evaluation. Since the optimum is unobservable on real datasets, evaluating any such estimator is itself nontrivial. We extend the FeeBee framework, originally proposed for evaluating Bayes-error estimators, to the optimal BER and AUC. The resulting procedure provides practical evaluation scores without requiring knowledge of the optimum, and applies to any estimator of the optimal BER or AUC, not only our proposed ones. Experiments on synthetic and real-world datasets validate both the estimators and the evaluation procedure.
cs.LG / 38 / 2609.02322
What Is Worth Representing? Representational Empowerment for Continual Model Construction
Abstract
The first problem of modeling the world is not just estimating the right parameters or causal structure, but deciding what should be represented at all. We frame this problem as continual model construction: an agent maintains an environment-specific model M of an inaccessible world W and curates a persistent library L of reusable representational elements across environments. We propose Representational Empowerment (RepEmp) to score candidate elements by how much they expand the agent's future capacity to model and plan, complementing the classic definition of empowerment, but redefined as control over internal representations instead of external states. We realize the framework as a hierarchical Curator-Actor architecture and test it across three experiments. In a closed-vocabulary causal-learning task, human participants construct causal models at varying abstraction granularities to maximize goal reachability rather than fidelity to the world, a signature better predicted by RepEmp than by information-gain alternatives. Matched simulations reveal that RepEmp-guided construction contributes more than exploration to sufficient structure recovery and cross-task transfer. Finally, in an open-vocabulary planning domain, an LLM-augmented Curator builds more compact symbolic libraries, which also generalize better than baselines. Ablating RepEmp eliminates these benefits. Together, these results identify RepEmp as a key principle for continual model construction: deciding what to build, retain, and reuse under bounded resources.
cs.LG / 39 / 2609.02339
AGI Maze Prediction Datasets: A Compact Benchmark for Learning World Dynamics with Transformers
Abstract
World modeling requires a predictive model to maintain and update an internal state adequate for reasoning about the consequences of actions. We introduce the AGI Maze Prediction Datasets and Benchmark, a lightweight controlled testbed for studying this capability in Transformers and other predictive models. Derived from procedurally generated, stateful grid worlds, the benchmark comprises per-step transition prediction, fixed-horizon state prediction, and sequential textual-observation prediction. Source-maze-disjoint training and validation splits, together with greedy exact-match evaluation, distinguish learning transferable action-conditioned dynamics from memorizing transitions in familiar layouts. We establish from-scratch byte-level Transformer baselines and compare them with two working-memory-augmented architectures. A generic auxiliary latent-memory Transformer can fit some training sets perfectly but does not consistently improve held-out performance. In contrast, a pseudo-video spatial-memory Transformer initializes a two-dimensional latent workspace from the input map and updates it from action history without receiving intermediate maps, positions, or state labels. Under the same data, objectives, and evaluation protocol, this model reaches perfect validation accuracy on selected fixed-horizon tasks where the byte and unstructured-memory baselines do not, and substantially improves sequential text-trace prediction. These results suggest that structured, task-aligned working memory can be more useful than additional latent capacity alone. More broadly, we argue that language grounding is mediated by persistent data structures and computations over them; the benchmark offers a compact setting for testing architectures that couple textual interfaces to learned structured state.
cs.LG / 40 / 2609.02373
Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance
Abstract
We study the dynamics of Stochastic Gradient Descent (SGD), which is known to steer deep neural networks toward invariant sets that correspond to simpler subnetworks. How this steering unfolds over time remains poorly understood. We answer this by modeling the stochastic gradient flow (SGF) as a percolation process, in which architectural symmetries force subnetworks to merge in discrete simultaneous blocks rather than one at a time. These structural transitions register as variance spikes in a macroscopic order parameter, echoing physical phase transitions. We further show this trapping mechanism and its associated scaling cascade extend to Adam and AdamW under an explicit heavy-tailed noise model.
cs.LG / 41 / 2609.02404
Evidence for Shared Routing Geometry and Dynamics in Sparse Mixture-of-Experts
Abstract
Sparse mixture-of-experts (MoE) models use an independently parameterized router at each sparse layer to select experts for every token. Prior work has shown that routing decisions across depth can often be predicted from earlier routing signals, suggesting that routing is not fully independent across layers. However, the structure behind this predictability remains unclear. In this work, we provide evidence that routing-relevant states across layers share a common geometric structure that is obscured by layer-specific coordinate systems. We isolate the control subspace of each router and align these spaces into a shared canonical representation using generalized orthogonal Procrustes analysis. After alignment, a single linear transition reaches $R^2=0.39$--$0.71$ and retains 79--90\% of the predictive power of separately fitted layer-specific dynamics, indicating that much of routing-state evolution follows a reusable process across depth. We then ask whether this shared dynamics is specific to routing or simply reflects the smooth evolution of hidden representations. A matched-rank comparison shows that residual representations are often easier to predict across layers, while router-control states preserve the model's expert choices much more faithfully. This separates generic cross-layer predictability from routing-specific information. Finally, we test whether the predicted canonical states remain meaningful when used in place of native routing states. The transported states preserve local routing behavior, while learned state evolution reduces $Δ\mathrm{NLL}$ relative to simple persistence by 15.7\% on OLMoE and 6.2\% over a 10-router horizon on Phi.
cs.LG / 42 / 2609.02417
Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment
Abstract
Multi-turn agentic RL increasingly treats credit assignment as a targeting problem: given a terminal verifiable reward, per-turn methods localize credit onto the turns that mattered. We identify the structural quantity that predicts when this is the right move, the verifier information density V_d = k/C (the fraction of an agent's C-step causal chain whose per-turn correctness the verifier exposes), and show that terminal-state verifiers sit deep in a low-V_d regime where targeting is the wrong axis. In controlled shared-rollout comparisons on tau^2-bench that separate reward density from credit geometry, a continuous dense reward spread uniformly beats the sparse binary outcome reward (net-harmful on 4/5 seeds), while concentrating the same advantage on progress turns or on random turns is equally harmful: targeting is second-order. The mechanism is coverage: terminal-state verification collapses the observable signal to a single final-write turn (k=1 in 98% of rollouts) while success requires a 5-8 step chain of prerequisite tool calls. A synthetic phase boundary places the crossover at V_d* ~ 0.8, whereas measured V_d is ~0.15 on tau^2-bench and ~0.4 on BFCL V3; uniform also wins on BFCL, where a matched-concentration shuffled control is negative on 8/8 seeds. The effect reproduces across model families on ToolACE-2-8B (Delta = -0.048 over 32 pre-registered seeds; an independent 20-seed replication is itself significant), and a pre-registered matched-budget breadth sweep traces a monotone dose-response whose deficit vanishes only at full chain coverage, with a reward-to-go arm reaching full-coverage parity. Uniform redistribution is the zero-information coverage default that per-turn schemes must beat; we contribute the matched-concentration shuffled control that any targeting claim should clear.
cs.LG / 43 / 2609.02422
IFW-BLS: Dual-Robust Broad Learning System with Intuitionistic Fuzzy Wave Loss
Abstract
Broad Learning System is an efficient randomized learning model that expands network width through feature and enhancement nodes and estimates the output weights without deep backpropagation. Its standard least-squares training, however, is vulnerable in two different ways: (i) large residuals caused by noise, outliers, or corrupted labels can dominate the objective, and (ii) all samples are treated as equally reliable even when some lie in ambiguous or locally conflicting regions. This paper proposes IFW-BLS, an Intuitionistic Fuzzy Wave Broad Learning System that addresses these two sources of fragility within one optimization model. The first robustness mechanism is residual-level protection, obtained by replacing the squared loss with the bounded, smooth, and asymmetric wave loss. Boundedness prevents extreme residuals from receiving unbounded influence, while asymmetry allows positive and negative deviations to be penalized differently when the dominant error direction varies. The second mechanism is sample-level credibility control, obtained through intuitionistic fuzzy scores that combine global class-center consistency with local neighborhood conflict. The resulting model evaluates the wave loss on credibility-weighted residuals, so unreliable samples are down-weighted before the bounded loss further limits the effect of extreme errors. A Nesterov accelerated gradient based optimizer is used to solve the proposed objective, avoiding the explicit matrix inversion used in conventional BLS. Experiments on UCI benchmark datasets validate the superiority of the proposed IFW-BLS model over the baseline models; additional corruption experiments also show more stable performance than BLS under noise and outlier contamination.
cs.LG / 44 / 2609.02440
Towards One-for-All Robustness Across a Continuum of Threat Levels
Abstract
Adversarially robust models often overfit to a specific attack budget, necessitating multiple specialized models for diverse and dynamic adversarial environments, a strategy that becomes fundamentally intractable as the threat space grows. This raises an open challenge: can we achieve strong robustness across a continuum of threat levels within a single model? We propose the Threat Conditional Network (TCN), grounded in a representation factorization framework that decomposes representation learning into a threat-invariant shared backbone and a lightweight threat-conditional adaptor. TCN conditions a single model on the perturbation level via Fourier-based embeddings and channel-wise affine modulation, and is trained against a distribution over perturbation budgets, enabling flexible and seamless adaptation across an infinite continuum of threat levels during inference. Extensive experiments on CIFAR-10, CIFAR-100, and Tiny-ImageNet show that TCN matches or surpasses a full ensemble of budget-specialized models with a single set of parameters, generalizes to unseen perturbation budgets, and transfers robustly under mismatched threat conditions, with only 4.6\% parameter overhead. These contributions chart a promising path toward adaptive and generalizable robustness in dynamic and diverse threat environments.
cs.LG / 45 / 2609.02450
CACTUS: Mask-Guided Semantic Clean-Label Backdoors in Decentralized Federated Learning
Abstract
Semantic triggers in federated learning (FL) can be less conspicuous than synthetic patches, but sample-dependent placement may weaken backdoor implantation across aggregation rounds. This challenge is compounded in decentralized FL (DFL), where topology-dependent peer aggregation repeatedly mixes local models. CACTUS converts label-consistent semantic pairs into target-directed representation shifts. Mask-guided, modality-specific operators isolate trigger effects, couple them across samples, and apply the shifts counterfactually to clean non-target embeddings before peer aggregation. Experiments cover speech, text, tabular, and image tasks under nine aggregation rules. With 30\% malicious nodes, CACTUS reaches a nine-rule mean attack success rate (ASR) of 51.2\% on Speech Commands and the highest nine-rule mean ASR among evaluated attacks on three of four modalities. Sensitivity analyses show that ASR varies with network topology and increases with the malicious-node ratio. These results indicate that CACTUS can propagate backdoors through repeated DFL aggregation.
cs.LG / 46 / 2609.02451
Scalable Kronecker-Fisher Approximation: Efficient Hessian Analysis for Billion-Parameter Language Models Compression
Abstract
In this paper, we propose a scalable Kronecker-based approximation that captures cross-layer interactions without storing the entire Fisher matrix, enabling practical Hessian analysis for billion-parameter networks where full computation is infeasible. Our approach reveals consistent vulnerability patterns: value projection layers exhibit the highest sensitivity and strongest cross-layer correlations across multiple model families, while other components exhibit architecture-specific behaviors. Through extensive experiments on quantization, sparsification, inter-layer corruption, and post-corruption fine-tuning, we demonstrate that our approximation strongly correlates with both performance degradation and recovery. Our framework provides a practical, theoretically grounded tool for identifying fragile components in large models, opening new avenues for guided compression and optimization strategies, such as mixed-precision allocation, layer-wise sparsity, and adaptive low-rank decomposition across layers and even individual weight groups.
cs.LG / 47 / 2609.02468
DeepAffinity: Long-Term Aspect Preference Prediction in eCommerce using Small Language Models
Abstract
We explore predicting eCommerce user preferences for product aspects such as brand, size, and color - a task we define as Aspect Affinity. Solving this task improves customer understanding and enables fine-grained personalization in recommendation, search, and marketing. We frame Aspect Affinity as a temporal prediction task: forecasting a users future aspect choices from their time-ordered interaction history, capturing long-term preferences that evolve beyond the current session. To this end, we propose DeepAffinity, which leverages Small Language Models (SLMs) with structured prompts and specialized prediction heads fine-tuned for this task. We show DeepAffinity outperforms standard generative fine-tuning methods, while general-purpose open-source LLMs perform poorly without task-specific tuning, highlighting their limits in modeling nuanced behavior. Finally, DeepAffinity enhances recommendation quality on a large-scale multinational eCommerce platform.
cs.LG / 48 / 2609.02497
RINSE: Robust Target-Time Normality Estimation for Zero-Shot Graph Anomaly Detection
Abstract
Zero-shot graph anomaly detection seeks to deploy a detector trained on source graphs to unseen, unlabeled targets, yet domain shift can make source-derived notions of normality unreliable. We introduce RINSE (Robust Iterative Normality Self-Estimation), a gradient-free target-time framework that keeps the source-trained detector fixed while sequentially estimating target normality, representation calibration, and evidence reliability from the target graph. Its core idea is to identify a reliable subset of low-residual target nodes, use them to construct a trimmed target-aware normality model, and combine complementary anomaly evidence through reliability-gated rank fusion and encoder ensembling. Across eight unseen target graphs, RINSE achieves the highest average AUPRC among the evaluated methods under two separate preprocessing protocols, while block ablations and sensitivity analyses support the combined design. These results support robust target-time estimation as a practical approach to generalist graph anomaly detection without target labels, gradients, or per-target tuning.
cs.LG / 49 / 2609.02507
Rethinking the Teacher-Student Framework for Test-Time Adaptation
Abstract
Test-Time Adaptation (TTA) has recently emerged as a promising strategy that allows the adaptation of pre-trained models to changing data distributions at deployment time, without access to any labels. To mitigate error accumulation, researchers have widely adopted the teacher-student framework, though its long-term stability is often taken for granted. In this work, we challenge the common strategy of setting the teacher weights to an exponential moving average of the student by showing that error accumulation still occurs, although it is mostly apparent on longer sequences compared to those commonly utilized. We analyze the stability-plasticity trade-off within the teacher-student framework and propose to use an intransigent teacher that does not update its weights. Surprisingly, we show that this simple change allows TTA methods to significantly improve their performance on multiple datasets with longer scenarios and result in increased robustness to changes in hyperparameters. Finally, we show that those changes can be seamlessly and effectively applied to various architectures and experimental setups, including semantic segmentation. The code is available at https://github.com/dmn-sjk/intransigent_teacher.
cs.LG / 50 / 2609.02519
Spectral Initialization and Scheduled Graph Smoothness for Uncertain Knowledge Graph Completion
Abstract
Uncertain knowledge graphs (UKGs) extend knowledge graphs by assigning each triple a continuous confidence score. Since most possible triples lack observed confidences, recent methods rely on semi-supervised learning to generate pseudo-labels. These methods initialize entity embeddings without using the confidence-weighted graph, discarding its global community and hub structure. We introduce QUEST, which adds no trainable parameters to the standard confidence-distribution learning pipeline. First, QUEST initializes entity embeddings using the smallest non-trivial eigenvectors of the confidence-weighted graph Laplacian, incorporating community and hub structure before training. Second, QUEST applies an unbiased mini-batch Dirichlet energy regularizer to enforce early-stage structural consistency. On two UKG datasets, QUEST improves confidence prediction and link prediction on six of eight metric-dataset pairs over prior methods and matches the previous best on the remaining two, while removing the instability spike observed on dense graphs. These results indicate that spectral structural priors combined with a graph Dirichlet energy regularizer improve accuracy, training stability, and checkpoint reliability in UKG completion.
cs.LG / 51 / 2609.02538
A Comparative Study of Graph Representations for GNN-Based Power Grid Control in L2RPN
Abstract
Graph construction is a critical but underexamined design choice in deep reinforcement learning for power grid control. We present a controlled experimental comparison of different graph representations, including physical topology, electrical-sensitivity, and hybrid variants for topology control in the Learning to Run a Power Network (L2RPN) environment. Our findings indicate that matching graph complexity to task granularity is more important than maximizing representational richness, and highlight the importance of controlled representation studies at scale.
cs.LG / 52 / 2609.02540
TrajMind: Chaining Role-Specialized LoRAs for Fast-and-Slow Collective Trajectory Anomaly Diagnosis
Abstract
Diagnosing collective anomalies from urban trajectories is increasingly important for traffic governance, as it reveals what happened, who was involved, and where and when the event occurred. Existing detectors efficiently produce scores or labels, whereas vision--language pipelines provide richer semantics; neither couples verifiable diagnosis with low-latency monitoring. The central challenge is to recognize collective patterns and recover exact event details from the source trajectories without running the full diagnostic pipeline for every monitored window. We therefore separate always-on screening from on-demand diagnosis: screening raises alerts, while diagnosis releases only source-verified what--who--where--when records. We present TrajMind, a fast-and-slow framework that switches three role-specialized LoRA adapters over one frozen vision--language backbone. Its slow path, \textit{TrajMind$_{\text{slow}}$}, chains canvas-based typing, type-conditioned localization over serialized trajectories, and executable verification, yielding structured, evidence-backed diagnoses. Additionally, the fast path, \textit{TrajMind$_{\text{fast}}$}, screens each window in a single text-only pass, delivering efficient structured alerts. Extensive experiments show that, TrajMind$_{\mathrm{slow}}$ outperforms the strongest baselines by at least $15.3$ percentage points in anomaly typing and $13.8$ percentage points in localization. These gains persist under cross-city transfer, and TrajMind$_{\mathrm{fast}}$ reduces latency by $41.1\%$ and maintains binary balanced accuracy of at least $93.5\%$. Together, TrajMind delivers accurate, evidence-backed diagnoses across cities and efficient front-line monitoring.
cs.LG / 53 / 2609.02549
ProbeMatchDTI: Probe-Driven Multi-Scale Biochemical Pattern Matching for Drug-Target Interaction Prediction
Abstract
Drug-target interaction (DTI) prediction is an important task in AI-driven drug discovery. Although recent biochemical representation learning methods have improved DTI prediction, their passive feature aggregation tends to favor dominant molecular patterns while suppressing weak yet binding-relevant signals, such as functional groups and residue-context patterns, limiting the modeling of multi-scale biochemical correspondences. To address this issue, we propose ProbeMatchDTI, a pattern-probe-driven framework comprising IterProbe and BindingProbe. IterProbe explicitly retains contextual states across refinement depths and uses learnable probes to select them at each position before cross-entity matching, thereby preserving weak biochemical patterns and strengthening associations among functional groups, local motifs, and molecular scaffolds. BindingProbe then characterizes cross-entity drug-protein complementarity at local biochemical-unit and whole-pair levels, jointly modeling fine-grained interactions and multi-scale correspondences while preserving weaker binding-relevant associations. Extensive experiments demonstrate the superiority of ProbeMatchDTI, achieving 2.0% and 0.5% higher AUC-ROC on BindingDB and DrugBank, respectively. Feature-level pattern analyses further characterize its probe-driven behavior in cross-scale biochemical pattern matching. We further connect ProbeMatchDTI predictions with an evidence-guided downstream drug-discovery workflow, demonstrating their utility for candidate refinement and validation planning. Our code is available at https://github.com/developer-hq/ProbeMatchDTI
cs.LG / 54 / 2609.02566
Online Reinforcement Learning in the Met Office Unified Model through Distributed Model-Agent Coupling
Abstract
Machine-learnt corrections can complement numerical weather prediction only if they adapt to the evolving model state while preserving dynamical consistency and numerical stability. To test this within a global forecasting model, we couple the Met Office (UKMO) Unified Model (UM) with distributed RL agents through rank-local tensors. A DDPG actor shares weights across the 70 vertical model levels of each atmospheric column and applies bounded potential-temperature corrections to the model tendencies. Across ten nudged training forecasts, nudging calculations towards the UKMO operational analysis provides an immediate counterfactual target. The frozen policy is then evaluated in a non-nudged forecast for inference. The coupled workflow successfully completes training and remains numerically stable in the evaluated case. Relative to a matched native UM forecast at +6 h, the learnt policy reduces Z$_{500}$ MAE in four of six latitude bands, including reductions of 45.8% and 40.8% in the northern and southern tropics. MSLP error too decreases in three bands, with a maximum reduction of 27.3% at 0-30°N. This single-case experiment demonstrates significant promise and feasibility of distributed online learning followed by non-nudged inference, laying the groundwork for RL-based bias correction and parametrisations within operational systems.
cs.LG / 55 / 2609.02622
Source Distribution Estimation by Posterior Averaging
Abstract
Simulation-based science often requires a distribution over simulator parameters whose push-forward reproduces a set of real observations: this is the source distribution estimation (SDE) problem. Existing methods fit the source against a likelihood surrogate trained once from a fixed proposal prior. Their objective is therefore stated only in terms of the surrogate instead of the true simulator, which may fail for inaccurate areas in parameter space where the surrogate was never trained. We instead solve SDE by expectation maximization: an E-step trains an amortized posterior on fresh simulations from the current source estimate, and an M-step refits the source to the average of that posterior over the observed data. We give two parameterizations, (1) separate source and posterior flows and (2) a single shared conditional flow. We evaluate our method on three benchmark tasks under both broad and misspecified initial priors. Both improve on existing fixed surrogate approaches and on iterated variants of each, most clearly on Lotka--Volterra, where no baseline falls below 0.96 data-space C2ST while our methods reach 0.64-0.68 in three of four initial-prior settings.
cs.LG / 56 / 2609.02638
Oracle, will I ever learn? A study of prediction convergence and complementarity across link prediction models
Abstract
Knowledge graphs have become an important source of structured knowledge for Web applications, including search, question answering, and recommender systems. In these applications, link prediction can serve either as a prediction task itself or as a means to enrich incomplete knowledge graphs for downstream tasks. Interestingly, different link prediction models, or even different training runs of the same model, can produce substantially different predictions for the same query. This suggests a variability in the capture of the underlying knowledge by models, thus raising a fundamental question: to what extent do different models capture complementary knowledge, and how much of this knowledge could be recovered by combining them? We propose to measure model complementarity through the performance of an oracle that, for each query, selects the best prediction among a considered set of models, hence providing an upper bound on the performance achievable through model combination. Across several architectures and benchmarks, we find a substantial gap between individual models and their oracle, revealing that different models capture complementary knowledge. Yet, this complementarity rapidly saturates as more models are added, leaving a persistent subset of queries unsolved even by a large number of models. These findings reveal both the potential of model complementarity and a fundamental limit to what current link prediction models can collectively recover; thereby highlighting the need for further research to build robust Web applications.
cs.LG / 57 / 2609.02646
Differentiable Electricity-Market Clearing for Gradient-Based Planning
Abstract
Planning a large data center is difficult because a facility big enough to matter changes the electricity prices it will pay. Those prices are set by market clearing, a constrained optimization problem solved anew in every operating condition. However, simulating the market tells a planner how a candidate plan performs but not how to improve it. Here we treat market clearing as a differentiable optimization layer: each forward pass solves the market, and reverse-mode automatic differentiation propagates the planning cost back through the cleared prices to the plan. After validating these gradients against finite differences, we apply them to a concrete problem: allocating 50 MW of data-center load across six candidate buses in two synthetic networks, under a fixed cost per active site, evaluated over 36 operating states. Judged against exhaustive enumeration of all site combinations, gradient optimization recovers the continuous allocations almost exactly, with worst-case objective gaps of 2.3\% and 8.5\% of the cost difference between the best and worst single site. Its one systematic error is instructive: near the costs at which a site should close, the smooth relaxation of the discrete site count shrinks the site rather than closing it, so discrete transitions arrive late. Differentiable market clearing thus turns market-aware planning into a problem gradients can search.
cs.LG / 58 / 2609.02652
Unfolding the Leech Lattice: Fused Multi-Shell Decoding and VRAM Layouts for 2-Bit LLM Weights
Abstract
Leech-lattice vector quantization holds the strongest reported 2-bit quality under its own evaluation protocol. Its kernel decodes one shell; we found no implementation of the multi-shell decoder the rate requires. This paper supplies one and measures its serving cost for decode-phase GEMV at batch 1. First, a serving path for the full 301-class codebook: an offline expansion into GPU layouts and a fused dequantize-plus-matvec kernel reading them without warp divergence, verified against f64. Second, the in-VRAM rate is a design axis distinct from the on-disk rate. Four bit-exact layouts timed in one process show binary bit planes beating one-hot masks on size and speed at constant bandwidth (4.80 bits per weight, 2.15x FP16). Below 4.3 bits a second, irregular stream enters; at 3.6 the decode stops being shifts and masks. Third, deployed four-bit (AWQ) and two-bit (QTIP) GEMV kernels run in the same process. The trellis kernel reads 2.40x fewer bytes than our served layout and runs 2.27x faster at near-equal fractions of their byte bounds: the time gap tracks the traffic gap, the price of unfolding a codebook too large for a lookup table. Fourth, the validity envelope: the trellis kernel outruns our no-weights control, so our launch geometry sets that floor, and on a second memory hierarchy every lattice arm falls below FP16. With the output head held identical across arms, the kernel-and-format path gains 1.11x, 1.29x and 1.41x end to end at 4B, 8B and 14B; with an int8 output head the served 4B reaches 87.0 tok/s in 2.60 GB. The quality cost, 1.38x perplexity and 14.7 MMLU points at 4B, shrinks across the three sizes measured.
cs.LG / 59 / 2609.02684
H3DNAS: Hardware-Aware ONNX-Native 3D Point Cloud Model Compression
Abstract
Deploying 3D point cloud models on edge hardware such as the NVIDIA Jetson Orin Nano is severely constrained by compute and memory budgets. Existing compression methods require access to the model's original source code, rendering them inapplicable to the Open Neural Network Exchange (ONNX) binaries commonly distributed by vendors and model repositories. We present \textbf{H3DNAS}, a hardware-aware model compression framework that operates directly on ONNX computational graphs without requiring original source code, architecture class definition, or gradient access during search. H3DNAS makes three contributions: (1) a \textbf{Channel Dependency Graph (CDG)} that classifies ONNX operators into four constraint classes and formally establishes that the free parameter fraction $ρ_f$ is topological invariant, a provable compression ceiling computable in $\mathcal{O}(|V|+|E|)$; (2) a \textbf{Two-Stage Hierarchical Search} that prunes candidate architectures by $L_1$-importance channel selection, ranks them by output fidelity as a zero-shot label-free proxy, and applies GhostConv structural mutation to Pareto-optimal candidates; and (3) the \textbf{first source-code-free compression pipeline for 3D point cloud models}, operating entirely via ONNX graph surgery with no original architecture definition required. On ModelNet40, H3DNAS reduces the number of parameters in PointNet, PointNet++, and PointMLP by $65.5\%$, $43.2\%$, and $49.1\%$, respectively, while achieving $1.99\times$, $1.29\times$, and $1.67\times$ inference speedups with negligible loss in accuracy. The source code is publicly available\footnote{https://github.com/ClarityLab-Org/h3dnas}.
cs.LG / 60 / 2609.02734
LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style Updates
Abstract
Low-rank adaptation (LoRA) is the standard way to fine-tune large models, yet when its two factors are trained independently, the update ignores the geometry of the low-rank weight change it induces. We introduce LoRA-TSD, an optimizer that treats every LoRA step as a tangent vector of the fixed-rank matrix manifold and takes the spectral-norm steepest-descent step of Muon inside that tangent space, mapping the result back to the factors through a retraction native to the LoRA parametrization. The step avoids expensive operations on full weight matrices, and its retraction is up to $2.8\times$ cheaper than the truncated-SVD retraction used by prior manifold methods. We prove that the Frobenius-norm version of our surrogate recovers LoRA-Pro, and we identify the tangent-projected gradient, the Riemannian gradient of the manifold, as the stationarity measure natural to LoRA training and computable from the factor gradients alone. Under this measure we give the first global convergence guarantees for both LoRA-Pro and LoRA-TSD, with rates that drive the factor-gradient norms to zero. Across six commonsense and natural-language-inference benchmarks with Llama-3.2-1B, Llama-3.1-8B and Qwen3-32B, LoRA-TSD outperforms every competing LoRA optimizer and stays robust to the adapter rank. Code is available at https://github.com/brain-lab-research/LoRA-TSD.
cs.LG / 61 / 2609.02766
Do Tabular Foundation Models Know Physics? Contamination, Units, and the Deterministic Limit
Abstract
Tabular foundation models (TFMs) learn to fill in tables the way language models fill in text, and tables are arguably the format in which most physical measurement arrives. Did they learn any physics in the process? They are Bayesian by construction, so the question is what their prior contains. We probe it directly, evaluating four of them (TabPFN-3, TabICLv2, TabDPT and Real-TabPFN-2.5) against six baselines on datasets sampled from 316 physical equations, in and out of domain. TFMs dominate, out of the box and after tuning. But we show that their prior can represent neither a noiseless mechanism nor physical units, which is why they interpolate physics without yet being able to act as physical models.
cs.LG / 62 / 2609.02846
UE5M3 FP4 Block Scaling for Stable Language Model Pretraining
Abstract
Stable 4-bit floating-point (FP4) pretraining is difficult because the E2M1 payload represents only a narrow range of magnitudes. NVIDIA's Transformer Engine \nv{} recipe addresses this with current-tensor scaling, a randomized Hadamard transform (RHT), and bfloat16 (BF16) final layers, adding work outside the FP4 matrix multiplications. We instead pair E2M1 payloads with unsigned E5M3 (\ue{}) block scales. Their wider range permits periodic tensor scaling, while our recipe applies selective stochastic rounding to backward gradients, omits RHT, and uses FP4 in all eligible internal linears. We pretrain a Nemotron-H 8B model for nearly 190 billion tokens. Compared with Transformer Engine \nv{}, the proposed block-16 recipe finishes with lower final-window training loss and, under their respective quantized-inference policies, lower validation loss measured as held-out negative log-likelihood. Its quantized-inference downstream point estimates are also higher on all three reported aggregates. A native \nv{} execution ablation that jointly removes RHT and the BF16 final-block exemption increases measured model-body token throughput by 21.2\%. These results demonstrate end-to-end software-emulated \uefp{} pretraining with a simpler recipe and motivate native support for \ue{} block scaling.
cs.LG / 63 / 2609.02852
The Implications of Linguistic Illegibility for LLM Security
Abstract
LLMs are trained to generate natural language. However, various strands of evidence indicate that an LLM's externalized linguistic outputs and mechanistically-extracted linguistic features can be an unreliable lens for understanding internal model computation. We introduce the term ``linguistic illegibility'' to broadly refer to scenarios in which an LLM's externalized or mechanistically-probed language artifacts fail to represent how the model actually thinks. We argue that the specter of linguistic illegibility is unavoidable for LLMs whose internal computations are not directly expressed via language, but rather math over activation spaces (with lossy translations between activation spaces and natural language happening at the bookends). If linguistic illegibility is always possible, then security mechanisms that rely on a model's linguistic self-reporting (e.g., chain-of-thought monitoring, constitutional self-critique, activation probing for linguistically-defined feature vectors) can never be completely sound; the model sandbox will always need isolation techniques whose guarantees do not depend on reading a model's linguistic state at all. We argue that observing a model's outputs using taint tracking is a promising approach for an effective sandbox: regardless of how a model linguistically self-reports, a taint tracking policy can define, a priori, various pieces of system state that should never be influenced by model-produced data. We also discuss several additional sandboxing mechanisms (e.g., robust virtualization, third-party auditing of sandboxing configurations) which collectively provide a critical floor beneath linguistic monitoring, and would have mitigated recent sandbox exploits by frontier models.
cs.LG / 64 / 2609.02881
Graph Machine: Towards Better Pretraining via Edges
Abstract
We introduce the Graph Machine (GM), an architecture that maintains an $O(n)$-sized state and accesses it through sparse, dynamic routing. Unlike methods with fixed-size states or sparse but static routing, GM preserves $O(n)$ complexity in its sparse layers without restricting the potentially accessible state size to $O(1)$. Instead, GM uses edges - pointer-like objects updated differentiably by a referral mechanism resembling pointer chasing. We replace 75% of the dense Transformer layers in Qwen3-0.6B with GM sparse layers and pretrain from scratch on 15.7B tokens. With only 2 of 4,096 tokens retrieved per KV head in each sparse layer, loss degrades only slightly; with 4, the best model marginally improves loss.
cs.LG / 65 / 2609.02887
A Common Measure of Communication for Speech Brain-Computer Interfaces
Abstract
Speech brain-computer interfaces (speech BCIs) translate neural activity into language, offering a path towards restoring speech for people with paralysis and, more broadly, enabling new forms of natural human-computer interaction. Despite this promise, the field lacks a common measure of progress because systems use different datasets, recording methods, types of speech, and vocabularies, so their reported scores are rarely comparable. Underlying this measurement problem are two unresolved questions: (i) what distribution of words should a speech BCI enable a user to communicate, and (ii) how much information from this distribution can a system convey. We address both by deriving open-vocabulary mutual information (OVMI), an information-theoretic quantity that measures the information conveyed by a decoder relative to a reference distribution over the words a user may wish to communicate. This allows capabilities measured under different conditions, such as distinct vocabularies, to be evaluated on a common communication scale. We show that ordinarily reported accuracy, word error rate (WER), and other metrics computed only over the words a system supports can overstate how much of a user's intended speech the system can communicate. We then use OVMI to compare existing systems, expose trade-offs between how much of the user's language a system supports and how accurately it decodes those words, show that these comparisons depend on what the user is expected to communicate, and demonstrate that selecting a vocabulary to maximise OVMI yields up to 16.3% relative improvement in accuracy across three speech domains. OVMI therefore provides the speech BCI community with a principled way to compare heterogeneous systems, improve vocabulary design, and measure progress in the field.
cs.LG / 66 / 2609.02876
GRADSOLVE: fast exact gradients for ODE ensembles on GPUs
Abstract
Ordinary differential equations (ODEs) underlie models in science and engineering, and many applications need derivatives of their solutions with respect to parameters. Ensembles of independent trajectories suit graphics processing units (GPUs), but current GPU software forces a trade-off: the fastest ensemble solvers cannot be differentiated in reverse mode at the speed they solve, and the solvers built for differentiation solve more slowly. No single tool has yet offered a reverse-mode gradient at the speed of a fused-kernel solve. We present GRADSOLVE, an open-source JAX library for solving and reverse-mode differentiating low-dimensional ODE ensembles on NVIDIA GPUs. It records the steps an adaptive solver accepts and differentiates a fixed-step replay of them; the returned gradient is the exact discrete adjoint of those steps, the same derivative Diffrax returns by default, obtained more cheaply from a fixed-length chain than from an adaptive loop. It targets ensembles differentiated many times against one recorded mesh, keeps Diffrax as a fallback, and supports explicit and Rosenbrock integrators. Used as a solver, GRADSOLVE's forward-only kernel ran 2.8x faster than DiffEqGPU.jl; used for gradients, once a record exists, it computed them 5.6-14.1x faster than Diffrax's checkpointed adjoint at matched forward-state accuracy across three GPU generations, the advantage narrowing on large ensembles and, on stiff systems, down to parity at tight accuracy. GRADSOLVE is released at https://github.com/ECLIPSE-AI4Science/gradsolve.
cs.LG / 67 / 2609.01957
Network-Aware Forecasting on Wireless Access Points
Abstract
Enterprise wireless access points (APs) are promising platforms for predictive machine learning (ML), but their primary responsibility remains providing wireless connectivity and network services. Predictive inference must therefore share an AP's CPU and memory with packet processing, Wi-Fi and IoT radio operations, and client management. This resource contention creates two risks: a model that performs well on proxy hardware may be too slow on the target AP, while a model that fits in isolation may still degrade network services under load. We define \textit{network-aware deployability} using two gates: qualification of the model and its execution path on the target AP, followed by validation of its execution profile under packet-service and forecasting constraints. Our benchmarks show that edge testbeds do not reliably capture target behavior. Across matched artifacts and serving settings, five model implementations run 6.1--19.1$\times$ slower on an AP than on a Raspberry Pi~5, while peak memory usage differs by up to 22\%. Moreover, two forecasting foundation models of similar size differ in AP latency by 19$\times$. When serving a smaller model across 13 parallel streams at a 30~s cadence under network saturation, default execution increases p99 round-trip time (RTT) by 76\% and reduces throughput by 7.06\%. Understanding these trade-offs is essential for live deployment if we aim to use APs for both networking and ML workloads.
cs.LG / 68 / 2609.02140
SoK: Where Do Flow Labels Come From? Auditing Label Provenance in Encrypted Traffic Benchmarks
Abstract
Encrypted traffic classification infers semantics beyond the flow record from transport-layer observables, and supervised training rests on labels that hold for the individual flow they are attached to. Recent systematizations scrutinize model in- puts and data splits; we systematize the complementary label side. Across 14 audited benchmark entries, we identify two recurring label-side strategies: coarse inheritance, which risks labelling flows the evidence does not cover, and overstrict filtering, which keeps only self-attesting flows and risks dis- carding relevant ones. No audited entry exposes a countable pre-selection population, and the task objects downstream papers attach to the same labels disagree with the recovered record in 8 of 23 referenced cells. Under strict side-channel features we derive a representation-relative ceiling on bal- anced accuracy for any classifier restricted to those features: on the public benchmarks that inherit, it ranges from 0.56 to 0.76. On the filtering side, only 24.95% of connections in our fully captured corpus carry an observable SNI of their own; yet the discarded connections raise macro accuracy from 0.44 to 0.65 through same-run co-occurrence features. We end with recommendations for benchmark builders and users.
cs.LG / 69 / 2609.02358
Humanoid Safe Stop via Learned Stoppability Value
Abstract
Humanoid robots responding to emergency stop commands typically execute a fixed maneuver, without reasoning about whether a safe stop is actually feasible from the current state. We cast emergency stopping as a reach-avoid problem and propose Safe-Stop, a task-agnostic framework that pairs a learned stop policy with learned stoppability estimators. The estimators are complementary: a stop-probability estimator supervised by the actual outcomes of the fixed stop policy, and a reach-avoidance estimator supervised by a Hamilton-Jacobi backup over physical state. The first captures emergent stopping behavior of the learned controller; the second provides a complementary recoverability signal. Because the stop policy and estimators do not depend on the behavior policy that preceded the stop command, they transfer across diverse upstream tasks without retraining. At deployment, the two estimates are combined: Safe-Stop commits to the stop only when both estimators indicate that stopping remains feasible, otherwise it hands off to a fall policy, instantiated as a damping fallback. This agreement check yields decisions that are robust without sacrificing reactivity.
cs.LG / 70 / 2609.02727
Neural operators approximate strongly continuous convex monotone semigroups
Abstract
We approximate strongly continuous convex monotone semigroups by learning their Chernoff-type one-step operators with neural operators. First, we introduce the general class of so-called Chernoff-neural operators and show in a universal approximation theorem that they can approximate the Chernoff one-step operators arbitrarily well. By using stability estimates between weighted Hölder spaces, the one-step approximation error can be propagated through the iterations which yields universal approximation of the corresponding semigroup. Second, we introduce the more specialized class of envelope-neural operators for envelope semigroups which allows us to derive quantitative approximation rates. Finally, we illustrate the effectiveness of these neural operators in several numerical examples arising from non-linear partial differential equations, stochastic optimal control and stochastic processes under model uncertainty.
cs.LG / 71 / 2609.02855
Improved Gradient Descent Lower Bounds Beyond Nesterov
Abstract
We study how far gradient descent (GD) can be accelerated by predetermined stepsizes in smooth convex optimization. Going beyond the classical $Ω(n^{-2})$ first-order oracle lower bound of Nemirovsky and Yudin, we prove an $Ω(n^{-1.6342})$ non-anytime lower bound and an $Ω(n^{-1.2408})$ anytime lower bound. These improve the recent $Ω(n^{-1.932})$ non-anytime lower bound of Ma and Chen and the $Ω(n^{-4/3})$ anytime lower bound of Tsai et al., respectively. Together with the non-anytime $O(n^{-\log_2(1+\sqrt{2})})$ rate achieved by silver schedules, our anytime lower bound establishes a strict separation between the achievable convergence exponents in the two settings.
cs.LG / 72 / 2609.02659
Dimension Dependent Correlation Gap Bounds under Restricted Independence
Abstract
The pairwise independent correlation gap is the ratio of the maximum expected value of a set function under arbitrary dependence to that under pairwise independence, measuring the loss from this independence restriction. Under mutual independence, this gap is universally bounded by $e/(e-1)$ for monotone submodular functions. With pairwise independence, a tighter $4/3$ upper bound was established for several special cases, including $n=3$, and conjectured to hold universally. A recent AI-assisted counterexample disproved this conjecture for $n=5$, leaving the validity of the $n=4$ bound and the tight worst case bound open. We resolve both questions. First, for $n=4$, we establish that the $4/3$ bound holds universally and is tight using an AI-assisted proof combining theoretical analysis and computational verification. The proof combines a structural characterization of optimal numerator vertices, permutation symmetry, cone certificate systems, Bernstein polynomial representations, recursive simplex subdivision, and verification of $2,745$ Bernstein coefficient systems. Second, we show that the worst case pairwise independent correlation gap attains $e/(e-1)$ asymptotically by constructing an instance with identical marginal probabilities and a monotone submodular union coverage function on a ground set partitioned into $m$ blocks. The number of blocks grows sublinearly with the ground set size. The result follows by constructing a feasible solution to a scaled asymptotic reduced dual of the pairwise independent linear program and immediately extends to $t$-wise independent random elements ($t\ge2$), since $t$-wise independence implies pairwise independence. Thus, pairwise independence, despite being the least restrictive form of independence in the $t$-wise independence hierarchy, can be as restrictive as mutual independence in the worst case.
cs.LG / 73 / 2609.01914
Basin Geometry and Reliable Recall of Dynamical Memories in Reservoir Computing
Abstract
Reliable attractor recall conventionally requires broad basins of attraction. However, in reservoir-computing based associative memory, temporal cues reliably recover dynamical memories despite basins dominated by unpredictable, riddled-like regions. We reveal that memory basins exhibit an ``octopus-like'' structure: a robust ``head'' near the attractor and thin, intertwined ``tentacles'' spanning state space. Initial states in tentacular regions yield near-zero uncertainty exponents, making the recalled memory effectively unpredictable at finite precision. Yet, cue-driven generalized synchronization bypasses this unpredictability, driving the system into the robust basin head. This mechanism yields a quantitative relation linking minimum cue duration, synchronization rate, and basin-head radius. Trained recurrent neural networks exhibit similar geometry, suggesting this phenomenon extends beyond reservoir computing.
cs.LG / 74 / 2609.01871
Latent unified smooth Hamiltonians for excited state chemistry
Abstract
We describe a neural network architecture and training procedure designed to model electronic ground and excited states of arbitrary molecular systems. By indirectly learning a latent, implicit basis representation of the electronic-state Hamiltonian, the model offers a unified treatment of multiple electronic states, conical intersections, and non-adiabatic couplings. The formalism can be further extended to learn consistent latent representations of additional operators such as transition dipole moments, for example. To demonstrate the general capabilities of our architecture, we train and evaluate networks on two realistic photochemical systems, thymine and azobenzene. The resulting models accurately reproduce energies and oscillator strengths for the ground- and low-lying excited states relevant to the photochemistry of these systems. We highlight the performance of the trained networks by studying critical molecular geometries, including conical intersections and excited state minima. By construction, the proposed framework also recovers the emergence of Berry phase accumulation around conical intersections. By pairing key mathematical structure from quantum chemistry with the representation learning power of transformers, the presented architecture offers a qualitatively new path toward fast and accurate ground- and excited-state simulations.
cs.LG / 75 / 2609.02209
Prototype-guided transfer of sparse literature knowledge for electrolyte additive discovery
Abstract
Electrolyte additive discovery remains challenging because experimentally validated molecules are sparse, whereas accessible chemical spaces are vast and largely unlabeled. This challenge is amplified in lithium-ion batteries, where additive performance arises from coupled interfacial reactions rather than a single molecular property. Here, we develop a prototype-guided molecular intelligence, ProtoMI, a literature-driven framework that learns transferable structural priors from reported electrolyte additives and uses them to prioritize candidates in unlabeled chemical space. For boron-containing additives, ProtoMI combines 126 literature-reported molecules with 179,977 unlabeled candidates. Graph contrastive learning identifies seven chemically interpretable prototypes from the reported additives, and prototype guided semi-supervised contrastive learning adapts these prototypes to the candidate space under source-target distribution mismatch. In retrospective temporal validation, ProtoMI achieves enrichment factors of 9.2-45.6 while screening less than 2% of the candidate space. A subsequent translation step identifies four commercially accessible candidates. One representative candidate, 4,4,5,5-Tetramethyl-2-[10-(1naphthyl)anthracen-9-yl]-1,3,2-dioxaborolane (TNDB), improves high-temperature LiFePO4||graphite cycling at 55 °C by 34.93% relative to the baseline electrolyte. An arsenal of characterizations and operando optical fiber Fourier transform infrared spectroscopy suggest that TNDB forms B-containing, F/P/O-modified inorganic interphases, suppresses solvent decomposition and reduces Fe deposition on graphite. This case study shows how sparse literature knowledge can guide experimentally efficient molecular discovery in data-scarce battery-additive spaces.
cs.LG / 76 / 2609.02833
Learning Spectral-Like Mesh-Free Discretisations
Abstract
Meshfree methods such as smoothed particle hydrodynamics (SPH) with kernel corrections, radial basis function-generated finite differences (RBF-FD), and the local anisotropic basis function method (LABFM) construct discrete differential operators by imposing polynomial consistency on a local stencil. For stencils containing more nodes than there are consistency constraints, the resulting linear system is underdetermined, and the remaining degrees of freedom are fixed implicitly by the choice of kernel, basis preconditioning, or a minimum-norm condition. Polynomial consistency constrains the operator only in the low-wavenumber limit, and no part of the construction selects for accuracy at the wavenumbers where fine-scale content resides. We introduce Spectral-like Neural Discretisation (SpeND), in which the choice of those degrees of freedom is cast as a learning problem: stencil weights are parametrised by a neural network conditioned on the local node geometry, trained to approximate the modal response of a spectral operator over the resolvable band. A hard-constrained projection layer maps the network output onto the affine subspace of consistent weights, so that polynomial consistency holds exactly by construction rather than as a penalty. Training is self-supervised and physics-agnostic, requiring no reference solutions; the objective minimises dispersion and dissipation error over a prescribed band-limited function space. Modal analysis on disordered two-dimensional node distributions shows that the learned fourth-order operator follows the exact response over a substantially wider band than either explicit LABFM at equal stencil size or fourth-order finite differences on a structured grid, whilst recovering the expected fourth-order convergence rate under refinement.
cs.LG / 77 / 2609.02186
Quantum MeanFlow: single-shot generative sampling on NISQ hardware
Abstract
Quantum generative models offer a promising framework for exploring whether quantum computation can enhance generative machine learning. Flow matching is a generative method in which samples are generated by transporting a simple, known distribution to the target data distribution with a learned velocity field. Its quantum counterpart, known as quantum flow matching (QFM), was introduced recently, and, like its classical counterpart, requires integrating an ordinary differential equation over many time steps during inference. As each step requires the output from the previous step, the circuit submission is sequential and a drawback on quantum computers as they have high input/output costs. To alleviate this problem, we introduce Quantum MeanFlow (QMF), the quantum analogue of the MeanFlow formulation, which allows single-step sample generation. While the QFM learns an instantaneous velocity field at each time step, QMF learns the average velocity over a time interval. We use a parameterized quantum circuit to learn these velocity fields and benchmark the two methods on the MNIST dataset. We show that while single-step QMF has lower image quality compared to multi-step QFM, it performs better than the single-step QFM sampling at every shot count. Both of our models are executed on IBM quantum computers and best-of-N rejection sampling recovers most of the accuracy lost to device noise without modifying the circuit. This is especially advantageous for QMF which has only one circuit evaluation per image. Here, We establish QMF as a viable method for single-step quantum generative sampling, saving on quantum circuit evaluations per generated sample.
cs.LG / 78 / 2609.01761
Pooling and Drift in Delayed Bandits
Abstract
A system often has to act long before it learns whether the act worked: a recommender sees a click in seconds and a purchase in days. With $K$ actions and a delay of $d$ rounds, the best rate known for this setting is $\widetilde{O}(\sqrt{(K+d)T})$ over $T$ rounds, so a longer menu is always more expensive to learn from. It need not be: if the outcome depends on the action only through the state it produced, then one late outcome informs every action that could have produced the observed state, and the price is set by how many genuinely different states the actions produce rather than by how many actions there are. We measure this using an effective dimension $v_t$ between $1$ and the number of states, and prove $\widetilde{O}(\sqrt{(d+1)V\log K})$ for a rotating algorithm and $\widetilde{O}(\sqrt{V^{-}}+\sqrt{dT})$ for the single-copy algorithm used in practice, for any budget fixed in advance; merging similar states lowers the price further, at an explicit bias. Even when given the exact losses from $d$ rounds ago, no algorithm escapes $Ω(\sqrt{dE\min\{1+\log J,T/d\}})$, where $J$ counts the drifting directions and $E$ bounds how far losses move while the learner waits. On generated data, the state channel cuts regret by up to 79 percent against action-level weighting and, on the funnel family, by 32 to 68 percent against a tuned minimax-optimal method.
cs.LG / 79 / 2609.01999
Posterior Tempering Explains Variance Inflation in Linear and Generalized Linear Thompson Sampling
Abstract
We study a variant of the Thompson Sampling (TS) algorithm, called $α$-TS, for solving stochastic generalized linear bandit problems. Existing analyses of TS require inflating the posterior variance to derive near-optimal regret guarantees. We formalize the idea of variance inflation by introducing $α$-TS that uses a fractional or $α$-posterior instead of the standard posterior. Our main contribution is to identify general regularity conditions on the prior and reward distributions that enable a regret analysis of $α$-TS without assuming any tractable approximation of the posterior distribution, unlike previous works. For a specific choice of $α\propto d^{-1}$, our general regret bound yields the best known regret bound of $O(d^{3/2}\sqrt{T}\log T)$ for both the exponential and sub-Gaussian families of reward distributions. We further provide an $α$-dependent lower bound showing that the regret constant depends on the product $αd$, and that when $α\propto d^{-1}$ the regret scales as $Ω(d^{3/2}\sqrt{T})$, explaining the origin of the $d^{3/2}$ factor in the upper bound. Our proof technique adapts and combines recent advancements in the analysis of linear bandit problems with first- and second-order posterior concentration theory from the Bayesian statistics literature.
cs.LG / 80 / 2609.02138
HyperMC: Multi-Fidelity Hyperparameter Tuning for Stochastic Gradient MCMC
Abstract
Stochastic gradient Markov chain Monte Carlo (SGMCMC) methods enable scalable Bayesian inference, but their performance depends strongly on hyperparameters such as the step size, mini-batch size, and number of leapfrog steps. Since most SGMCMC algorithms lack a Metropolis-Hastings acceptance rate, standard acceptance-based tuning methods are not directly applicable. We propose HyperMC, a multi-fidelity tuning framework that combines Hyperband-style resource allocation with kernel Stein discrepancy (KSD) evaluation. By running multiple successive-halving brackets, HyperMC balances broad exploration of a continuous hyperparameter space with increasingly accurate evaluation of promising configurations under a fixed computational budget. We further introduce Robust HyperMC, which uses global grid initialization followed by elite-guided local refinement to reduce sensitivity to random candidate generation and noisy finite-budget evaluations. Under suitable approximation and concentration conditions for the estimated KSD, we establish that the successive-halving component selects a near-optimal configuration among the sampled candidates with high probability and derive a sufficient computational budget for successful selection. Experiments on logistic regression, probabilistic matrix factorization, and Bayesian neural networks show that HyperMC improves posterior approximation or predictive calibration relative to MAMBA, grid search, and heuristic baselines, while Robust HyperMC yields more stable and reproducible tuning results.
cs.LG / 81 / 2609.02286
From topology learning to graph generation: A unifying perspective
Abstract
Learning graph structures from data is a fundamental problem that spans a wide range of signal processing and machine learning tasks. While significant effort has been made to tackle the problem, existing research has largely evolved along two parallel directions. The first seeks to infer the topology of an individual graph from observations supported on it, whereas the second seeks to learn a generative distribution from observed graph instances, enabling the sampling of new graphs. This review presents a unified framework that connects these formulations by viewing them as inverse problems of a common generation process for graph data. We review the major methodologies within this framework, highlight their relationships, strengths, and limitations, and identify opportunities for integrating ideas across paradigms. By bridging graph topology learning and graph generation, this review provides a broader cross-disciplinary perspective on the field and outlines promising directions for future research.
cs.LG / 82 / 2609.02382
A computational approach to maximum likelihood thresholds for colored Gaussian graphical models
Abstract
Gaussian graphical models (GGMs) are essential tools for interpretable structure learning. However, in high-dimensional, small-sample regimes, the available data is often insufficient for the maximum likelihood estimator to exist. Colored Gaussian graphical models (CGGMs) mitigate this limitation by imposing symmetry constraints through graph coloring, which reduces the required sample size. This minimal number of observations needed to guarantee that the estimator exists almost surely is defined as the maximum likelihood threshold (MLT). Here, we address the computation of the MLT for CGGMs by focusing on its geometric formulation: finding the minimum rank of a sample covariance matrix such that its projection lies almost surely within the interior of the cone of sufficient statistics. We establish a unified theoretical framework, extending results from uncolored to colored models and introducing new symbolic algorithms. Furthermore, we present a computational study integrating sampling with topological data analysis (TDA) to investigate the local geometry of the cone of sufficient statistics. Our results demonstrate the potential of TDA to overcome the computational bottlenecks of traditional symbolic algebraic methods, particularly Groebner basis computations, in analyzing the likelihood geometry of CGGMs.
cs.LG / 83 / 2609.02728
Momentum in large-batch training: Polyak enlarges the critical batch size, Nesterov improves data efficiency
Abstract
We study when and how momentum improves large-batch training in the one-pass regime, using power-law kernel regression as a tractable setting. We first characterize risk stability through the critical learning rate, defined as the largest learning rate for stable training, and obtain $η_{\mathrm{SGD}}^{\mathrm{crit}}\eqsim 1$, $η_{\mathrm{Polyak}}^{\mathrm{crit}}\eqsim \min\{1,B(1-ρ)\}$, and $η_{\mathrm{Nesterov}}^{\mathrm{crit}}\eqsim \min\{1,B^β(1-ρ)\}$, where $B$ is the batch size, $ρ$ is the momentum factor, and $β>1$ is the capacity exponent. Within this admissible region, we derive scaling laws for the full risk dynamics, capturing the progression from an early transient, through power-law decay, to a noise floor. We then minimize the final-step risk over the admissible learning rates and momentum factors under a fixed data budget, yielding a three-regime batch-size phase diagram that reveals how the role of momentum changes with batch size. Notably, Polyak enlarges the critical batch size, the largest batch size preserving the best small-batch data-scaling exponent, thereby enabling greater parallelism without sacrificing data efficiency. In contrast, Nesterov achieves better data efficiency in the large-batch regime because its look-ahead mechanism suppresses noise accumulation. Numerical experiments validate the predicted stability boundaries, risk dynamics, and batch-size phase diagram.
cs.LG / 84 / 2609.02790
Full-Model Optimality for Tunable Linear Generative Priors in Compressed Sensing
Abstract
Generative models have been studied experimentally and theoretically as priors for inverse problems such as compressed sensing. Recent work by Gunn et al. studied the use of generative priors with tunable complexity, where a family of generative priors with varying complexity is maintained and a specific complexity can be selected at inversion time. They demonstrated that lower reconstruction errors can be experimentally attained for a variety of inverse problems by appropriately tuning the complexity of the generative prior. In the present paper, we establish theory for compressed sensing in the setting of a tunable family of linear generative priors naturally related through their singular value decompositions. We prove that in noiseless Gaussian compressed sensing, the full-dimensional linear prior attains the minimum expected reconstruction error over the entire family of linear priors. Thus, in this idealized linear noiseless setting, tuning to a lower-complexity prior does not improve the expected reconstruction error. This result is in contract to the behavior of denoising, where lower complexity priors attain lower reconstruction errors due to a standard bias-variance tradeoff. This result indicates that the experimental benefits of tunability in compressed sensing with neural network priors arises due to nonlinearities in the generative models.
神经与进化计算 (cs.NE)
5
cs.NE / 1 / 2609.01735
CircuitsDNA: Discovering Unconventional Multi-Accuracy Arithmetic Circuits via Evolutionary Synthesis
Abstract
Emerging edge AI workloads increasingly require arithmetic units that can trade computational accuracy for efficiency on demand. However, existing approximate arithmetic circuits are typically fixed-accuracy or rely on predefined structures for runtime configurability. This work introduces CircuitsDNA, an evolutionary framework that automatically evolves accuracy-configurable arithmetic circuits supporting multiple accuracy modes within a single circuit. It integrates three key features: 1) multi-threshold verifiability miter to enforce mode-specific accuracy requirements, 2) resource-limited verifiability-driven search to reduce verification overhead without sacrificing correctness, enabling efficient exploration of large circuit design, and 3) feedback-driven adaptive mutation to prioritize effective structural modifications and accelerate search convergence. Experimental results show that the 8-bit multiplier variants synthesized in 28-nm CMOS reduce the area-power product by up to 56% on INT8 DNN workload and 93% under exhaustive activity, compared with an exact 8-bit multiplier. Across CNNs and DeiTs, the accuracy loss relative to FP32 remains below 2% after fine-tuning under worst-case error (WCE) budgets of at most 1%. CircuitsDNA eliminates all search stalls observed in conventional methods across 8/12/16-bit multipliers, while adaptive mutation provides up to 1.33 times faster convergence than its non-adaptive counterpart.
cs.NE / 2 / 2609.01811
Reinforcement learning to choose optimizers
Abstract
No single optimization method is uniformly best for all problems, and the most suitable optimizer choice can change during a run. Existing approaches that change optimizer during execution typically predetermine part of the strategy: the portfolio is restricted to one algorithm class, the switch occurs once at a fixed time, or the frequency of decisions is treated as a hyperparameter rather than a learned one. We introduce "Reinforcement Learning to Choose Optimizers", which formulates the optimization algorithm choice as a sequential decision-making problem. At each decision, a recurrent policy reads the current run state and decides both which optimizer should be used next and for how long. The portfolio includes both gradient-based and derivative-free optimizers, and each switch passes on the current best solution and a representative step size. A context proxy conditions a gating network over expert heads, and training employs a decoupled actor-critic whose return is expressed in the same empirical runtime distribution metric used at evaluation. Training tasks and portfolio are designed jointly so that no optimizer dominates. On unseen problems, the learned policy outperforms every portfolio optimizer at all but the smallest budgets, and it remains robust under distribution shift.
cs.NE / 3 / 2609.02195
Memory as an Energy Landscape---Hopfield
Abstract
This chapter reconstructs the Hopfield network as a physical theory of memory rather than merely an early neural-network algorithm. It begins with the problem as it stood before 1982-threshold logic, Hebbian association, correlation memories, and recurrent binary networks-and isolates what Hopfield's synthesis added: a dynamical definition of content-addressable memory, a symmetric recurrent architecture with a Lyapunov function, a Hebbian embedding of patterns in its couplings, and a physical account of basins, robustness, and graceful degradation. The binary and graded-response energy functions are derived in full, together with the signal-crosstalk decomposition governing pattern stability, the mean-field theory of retrieval at extensive load, and the zero-temperature retrieval spinodal at (alpha 0.138) established by Amit, Gutfreund, and Sompolinsky. The energy-based program is then followed through analog optimization networks, polynomial dense associative memories, exponential interactions, and modern continuous Hopfield updates, including the precise conditions under which the update becomes scaled dot-product attention. Throughout, capacity claims are tied to their disorder ensemble, scaling limit, and success criterion, showing why numerically different storage limits need not conflict. A closing assessment distinguishes established results from surviving principles, assumption-bound limitations, and open problems, treating the Hopfield network as an effective theory whose symmetry, locality, and point-neuron assumptions delimit its biological reach. Fixed-seed numerical experiments expose the mechanisms discussed but do not substitute for analytical results.
cs.NE / 4 / 2609.02460
Model-Free Surrogate-Assisted Neural Architecture Search for Evolving Variable-Length Dense Blocks
Abstract
Neural Architecture Search (NAS) has emerged as a powerful paradigm for automatically designing deep neural networks; however, its practical adoption is often limited by substantial computational cost. To alleviate expensive full-training evaluations, surrogate-based methods have been introduced to estimate network performance efficiently. Nevertheless, existing approaches-particularly model-based surrogates-require training many candidate architectures and involve additional optimization overhead. In this work, we propose a Model-Free Surrogate PSO Network (MFSPNet) for evolving convolutional neural network architectures. The proposed method integrates a lightweight model-free surrogate predictor within a particle swarm optimization (PSO) framework, eliminating the need for pre-trained surrogate models. Specifically, MFSPNet introduces two key contributions: (1) a validation-loss-driven exponential moving average estimator (VLE-EMA) that captures early generalization behavior for reliable architecture ranking; and (2) a block-based dense connection strategy that enables effective stacking of evolved blocks while mitigating vanishing-gradient issues. This design also facilitates transferability of learned blocks across datasets. Extensive experiments demonstrate that MFSPNet achieves competitive performance with reduced computational cost. Under a consistent training protocol with ten independent runs, the proposed method attains error rates of 3.91%, 17.68%, and 1.91% on CIFAR-10, CIFAR-100, and SVHN, respectively, along with top-1/top-5 error rates of 28.29%/12.82% on ImageNet, while requiring less than three GPU days for architecture search. Due to computational constraints, the ImageNet result is based on a single run and should be interpreted as indicative of scalability. Overall, MFSPNet provides an efficient and reliable framework for cost-aware neural architecture search.
cs.NE / 5 / 2609.02183
Neural Logic, Invariance, and the Retina---McCulloch and Pitts
Abstract
This chapter reconstructs the McCulloch-Pitts program as a physics of neural computation rather than the familiar cartoon of a binary neuron. The 1943 logical calculus is developed in both directions: given a net, characterize the propositions realized by its activity; given an admissible logical expression, construct a net that realizes it. We recover the original distinction between thresholded excitatory summation and absolute inhibitory veto-one the weighted-threshold form cannot preserve for arbitrarily large excitatory inputs-and read unit-time delay as the physical realization of logical depth. Recurrence is treated exactly: an autonomous, deterministic network of finitely many binary units has a finite state space, so every trajectory eventually enters a periodic orbit-a fact about finite-state dynamics, not unbounded Turing computation. A single threshold element realizes only linearly separable Boolean functions, whereas finite feedforward networks of them synthesize any Boolean function on a finite domain. We then follows McCulloch and Pitts beyond threshold logic. The 1945 heterarchy paper turns cyclic preference into an obstruction to representation by a scalar utility. The 1947 work on universals asks how a physical network can identify inputs related by nuisance transformations, developed here via group averaging and feedback canonicalization. The 1959 frog-retina study makes the adequate-stimulus question experimental, revealing parallel invariant operations before the brain proper. Spike-triggered analysis shows how a nonlinearly driven neuron can have a vanishing first-order average while second-order statistics recover its hidden selectivity: methodological failure can masquerade as physiological absence. Modern mathematical tools are used without projecting their notation onto the historical papers, and limitations of the idealization are stated explicitly.
计算语言学 (cs.CL)
37
cs.CL / 1 / 2609.01737
SpeakPay: Domain-Adaptive LoRA Fine-Tuning of Whisper for Low-Resource Nepali Financial Speech Recognition
Abstract
Mobile payment applications in Nepal are graphically mediated and largely inaccessible to visually impaired users. This paper presents SpeakPay, a voice-first digital wallet, and documents the central technical contribution: a controlled study of domain adaptation for low-resource financial speech recognition. We introduce NepFinSpeech-403, a 403-utterance dataset of Nepali financial voice commands (send, load, and balance operations spanning 237 unique numerals), and fine-tune Whisper large-v2 with LoRA. On the held-out test set, the domain-adapted model reduces Word Error Rate from 129.95% (zero-shot baseline) to 42.58% --- a 67.2% relative reduction --- and improves Devanagari numeral recognition accuracy from 0.0% to 73.9%. We find that word-level metrics understate the practical task-level impact: domain adaptation improves the Transaction Success Rate from 1.67% to 33.33%, a roughly 20x gain. The improvement is consistent at the individual-utterance level (sign test, $p < 10^{-17}$) and across all command types. A data efficiency analysis shows that as few as 100 domain-specific utterances are sufficient to halve the zero-shot WER, with performance plateauing around 300 examples. Error analysis reveals systematic numeral confusion patterns (zero insertion/deletion, prefix hallucination) that account for the majority of remaining transaction failures. The trained system is deployed as a publicly accessible voice-first web application. All code, dataset, model weights, and this paper are released at https://github.com/subedibiraj/speakpay.
cs.CL / 2 / 2609.01772
MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models
Abstract
Meme understanding goes beyond recognizing visual content or literal text; it requires implicit cultural knowledge and pragmatic inference that most vision-language models still lack. We introduce MemeCULT-1K, a multilingual benchmark of 1,000 South Asian memes in Bengali, English, and Hindi, where each meme is paired with a cultural context note and three human-written explanations, along with a supplementary set of 54 Bengali regional dialect memes. We evaluate thirteen popular Vision Language Models (VLMs) under two settings: meme-only and context-aware. Providing minimal cultural context yields consistent gains across all models and languages: mean SBERT similarity improves from 44.6 to 56.4 (+11.8), BLEURT from 37.3 to 42.3 (+5.0), and LLM-as-a-Judge scores from 2.57 to 3.43 out of 5 (+0.86). Fine-grained error analysis reveals that closed-source models fail mainly on entity and reference misidentification, while open-source models are bottlenecked by broader cultural knowledge gaps, with linguistic and phonological failures proving the most context-resistant across both. These results highlight the difficulty of culturally grounded meme understanding and motivate future work on explicit cultural knowledge integration. Our dataset and code are publicly available at TawsifDipto17/MemeCULT-1K.
cs.CL / 3 / 2609.01794
Disentangling Statistical Preemption from Entrenchment in Language Models' Avoidance of Overgeneralization
Abstract
How do learners avoid overgeneralizations such as Tom laughed me without explicit negative evidence? Constructionists have posited two proposals that describe indirect negative evidence against overgeneralizations: preemption (which privileges exposure to near-synonymous construction---e.g., she made him laugh) vs. entrenchment (all exposures to a verb's grammatical usages, including cases like He laughed). We disentangle these hypotheses by running controlled rearing experiments on LMs trained on child-caregiver conversations, where we systematically remove preemptive vs. non-preemptive evidence. We find that while LMs avoid overgeneralizations, they do not show preemption at a verb-specific level, instead showing weak but non-zero evidence of abstract preemption. Combined with results from analyzing the LMs' training dynamics, we find that LMs treat competing structures as indirect positive---as opposed to negative---evidence in the verb-specific condition. Insofar as preemption is the more plausible route to avoiding overgeneralizations in humans, our results point the need for there to be sensitivities to indirect negative evidence in neural network learners, and suggest new human experiments to test abstract preemption.
cs.CL / 4 / 2609.01810
TalkFa: A Unified Benchmark for Farsi Dialogue Generation and Understanding
Abstract
Farsi, spoken by more than 120 million people, lacks a comprehensive benchmark for dialogue generation and understanding. We introduce TALKFA, a unified benchmark comprising three complementary datasets: (1) WIKI-FADIAL, 4.2K Wikipedia-grounded dialogues for knowledge-grounded generation; (2) DAILYDIALOG-FA, 6.6K dialogues annotated for dialogue acts and emotions; and (3) PLAYDIAL-FA, 2.1K theatrical dialogues with sentiment labels. While LLMs assist data construction, every dialogue undergoes multi-stage review and revision by native Farsi speakers, and only the final human-approved dialogues are released. Experiments with six LLAMA and MISTRAL models show that LoRA substantially improves dialogue generation while requiring only 25-50% of the training data to recover over 90% of the final performance gains. Across classification tasks, FABERT achieves the best dialogue-act performance, LORA-MISTRAL-7B performs best on emotion recognition, and MISTRAL-24B achieves the highest sentiment score. Human evaluation and independent external validation demonstrate the reliability of the benchmark, while comparisons with GPT-4.1 as an LLM judge reveal that automatic metrics substantially overestimate dialogue quality. Zero-shot evaluation with frontier LLMs further shows that TalkFa remains a challenging benchmark. We will release all datasets, annotation guidelines, code, and checkpoints.
cs.CL / 5 / 2609.01828
AVERT: Audio-Verified Adjudication for Spoken Dialogue State Tracking
Abstract
Spoken dialogue state tracking recovers slot-value pairs from speech, where ASR errors concentrate in entity values and persist across turns, making it both a generation and an editing problem. A strong per-turn text editor corrects much of this but, operating on the transcript alone, leaves three recoverable errors: a value predicted inconsistently across turns, an omitted slot, and a value the audio does not support. We present AVERT, which scores each candidate value by combining cross-turn agreement with a trained audio-conditioned verifier and resolves the three error types with three operators, vote, add, and swap, each restricted to the slots where its error is common. On SpokenWOZ, a base speech-LLM reaches 33.04 JGA, a text editor 38.34, and AVERT 40.13, without retraining either. This is in the range of a 1B end-to-end system that consumes the full spoken history (39.32), though AVERT uses two 1B decoders rather than one. The audio verifier contributes a statistically significant gain, and restricting each operator to a selected slot subset matters: removing it lets unrestricted voting overwrite correct categorical values and fall below the editor.
cs.CL / 6 / 2609.01833
Candidate Generation and Definition-Guided Verification for Sentence-Level Depression Symptom Recognition
Abstract
Sentence-level recognition of depression symptoms is challenging because similar expressions can differ in symptom relevance, and language-model inference is insufficiently grounded in diagnostic definitions. This study proposes a two-stage framework separating symptom-candidate generation from definition-grounded verification. A contrastively fine-tuned sentence encoder generates a symptom candidate per sentence, and a fine-tuned language model verifies whether the candidate is present or absent using the sentence, its context, and a candidate-specific diagnostic definition, checking its judgment against that definition before answering. Evaluated against encoder, inference-based, medical, and general LLM baselines and a matched single-stage supervised classifier, the proposed pipeline attains the best accuracy and F1 scores of all methods, with rationales matching expert-authored annotations. A preliminary clinical audit indicates moderate alignment with diagnostic definitions, with explanation quality strongly dependent on prediction correctness. The results support decomposing symptom recognition into candidate generation and definition-grounded verification, though performance remains limited for rare categories.
cs.CL / 7 / 2609.01846
Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture Videos
Abstract
Recorded lecture videos, often enhanced with search and summarization features, are a standard study resource. However, students cannot easily ask course specific questions or verify answers against an instructor's lecture. We report a semester-long deployment of VideoPoints platform with a retrieval-augmented chatbot that answers from course lecture materials and returns timestamped citations. The chatbot retrieves only from the active course, uses chapter summaries to guide transcript ranking, and returns clickable timestamped citations. Students used it for quick lookups and exam review. Across 833 messages, 70.5% included citations, none crossed a course boundary, and when no lecture evidence matched, the chatbot usually declined rather than answering. Among the users, citations were the most consistently useful feature, while practice-question generation was the strongest unmet request. We also evaluated the design on the real-world test split of EduVidQA, a public multimodal benchmark for lecture-video question answering. Our design improved correct-lecture retrieval by 6.3 percentage points over dense-only retrieval. Together, the results show that effective deployment depends on course isolation, supported citations, and alignment with students' study practices.
cs.CL / 8 / 2609.01878
GAPS: Dimension-Level Gates for Conditional Activation Steering
Abstract
Activation steering suppresses undesired behaviors in language models by adding a steering vector to the hidden state during generation. Recent conditional methods such as CAST and DSAS improve the behavior-capability trade-off by deciding when to intervene, but once active, they apply the full dense vector to all hidden dimensions, regardless of whether a neuron carries concept information or already lies in the desired regime. We introduce dimension-level conditioning as a complementary axis of selectivity that also decides which neurons to intervene on. Our method, GAPS (Gated Activation steering via Posterior and Separability), combines two training-free gates: a static separability gate that restricts steering to neurons with statistically reliable concept information (via AUROC), and a dynamic posterior gate that steers a neuron only when its current activation is better explained by the undesired concept under a Gaussian model. The gates add O(D) overhead per token, and they plug into existing conditional methods. On toxicity mitigation (RealToxicityPrompts) and concept removal (OneSeC) with Gemma-3 (4B) and Qwen-3 (1.7B), GAPS consistently matches or improves the Pareto front of its token-level counterparts; under a fixed capability budget, DSAS+GAPS reduces Gemma-3's toxicity rate from 6.52% to 0.48%, versus 3.52% for DSAS alone. Ablations attribute most of the gain to the posterior gate.
cs.CL / 9 / 2609.01918
Grounded, Compute-Efficient LLM Policy Agents for Energy-Poverty Equity in Physically-Constrained Peer-to-Peer Energy Markets
Abstract
Energy poverty is nearly absent from NLP-for-social-good, and the little existing work is either static retrieval/QA or relies on carbon-intensive cloud LLMs, a self-defeating "computational irony" for a humanitarian setting. We present EqGrid, a closed-loop simulation in which a low-frequency, open-weight LLM policy agent sets price and carbon bounds and targeted subsidies over a community of empirically-grounded household personas, while high-frequency multi-agent RL traders clear a continuous double auction constrained by a physical distribution grid (IEEE-33-bus with Dynamic Operating Envelopes). Our contribution is threefold and directly addresses how to measure the social impact of AI: (i) grounded personas (region-matched socio-demographics) whose load curves are checked for shape and level realism against real smart-meter data; (ii) formal energy-poverty equity metrics (Energy Burden, Gini of EB, LIHC) showing the intervention reduces burden inequality without raising net grid cost; and (iii) a compute-efficiency frontier that measures how much equity performance survives compressing the policy agent from a 235B teacher down to a sub-1B model deployable on a laptop, in estimated energy/carbon per decision. A decoupled-safety design (the LLM sets bounds; a validate-and-project grid gate executes) yields zero grid-constraint violations versus 55 under direct LLM control. On energy-poverty equity, the LLM policy lowers the Gini of energy burden to 0.305 (from 0.351) and mean burden by 28% while cutting cost (outperforming a tuned rule baseline), and a 3B-active model retains 95% of the benefit at roughly 9x lower inference energy than the teacher, with even a 0.8B on-device model retaining 92% at roughly 24x lower energy. We will release code and configs.
cs.CL / 10 / 2609.01936
Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens
Abstract
A language model's prediction of its next token develops across layers, and lens methods track this process by decoding intermediate hidden states into tokens. But a lens reading reflects both the hidden state and the readout (the unembedding matrix) used to decode it. Many lenses are fit on a corpus, and we show that two lenses differing only in their fitting corpus can report different tokens for the same hidden states. We call this dependence corpus conditionality. To examine readout structure independently of the fitting corpus, we introduce Sparse Readout Prism (SRP), which decomposes the readout using only its weights and expresses any token logit or logit difference as a sum of contributions from sparse readout features. This reveals readout features as a new unit of analysis for lens readings, exposing structure that token identities can obscure and enabling comparisons across tokens, contexts, layers, and lenses. Replacing the original readout with SRP's sparse approximation reconstructs 8.9-17.3 percentage points more of the tested logit differences than the strongest of six baselines built on geometric relations among readout rows. Ablating features shifts logit differences in proportion to their SRP contributions. Although token readings vary with the fitting corpus, the dominant readout feature remains stable. Because SRP uses no corpus in its construction, it provides a control independent of the fitting corpus for lens analyses.
cs.CL / 11 / 2609.02015
How Output Format Confounds Data Quality and Capability in Instruction Tuning
Abstract
Instruction-tuning data are judged by quality metrics, and tuned models are judged by benchmarks, but both judgments pass through an output interface: the surface format in which an answer is written. Using gradient signatures across 12 tasks, four semantically equivalent interfaces, three model families, and controlled corruptions, we show that this interface confounds both measurements. Spectral statistics such as effective rank are provably invariant to interface rotation and empirically blind to semantic corruption, while the direction of the update carries the quality signal. The interface-varying residual is not noise: it identifies each unit's own target task perfectly across all three families. Capability itself is stored relative to the training interface: a skill that raises accuracy by more than 40 points under the training format can be nearly invisible under every other, and correcting a single generation budget flips the measured effect of fine-tuning on GSM8K from a gain into a large loss. Pre-registered interventions delimit where this geometry stops short of control. Data quality and model capability are interface-conditioned quantities, and current practice often reports the interface instead of the content.
cs.CL / 12 / 2609.02131
C$^{3}$T: Counterfactual Causal Reasoning for Sentiment Shifts in Social-Media Conversation Trees
Abstract
Sentiment in social-media threads does not only vary across posts; it shifts as users react to claims, corrections, evidence, and hostility within a branching reply tree. We study why sentiment changes in rumor-centric conversation trees by treating discourse moves (e.g., denial/correction, evidence/link, toxicity/attack) as candidate interventions and asking (i) what sentiment a reply expresses, (ii) whether the sentiment shifts relative to its parent, and (iii) which prior message most plausibly drove the reply's sentiment. To support this setting, we introduce CaSiRe, a causal sentiment reasoning layer over public rumor conversation datasets that adds post-level sentiment labels, induced parent-child shift labels, calibrated multi-label intervention tags, and explicitly annotated causal-source labels. We then propose C$^{3}$T (Counterfactual Causal Conversation Transformer), a thread-structured temporal model that jointly predicts node sentiment and shifts, learns sparse ancestor attribution, and supports counterfactual queries by forcing conversational intervention embeddings on or off to estimate potential outcomes. Under an event-level split, C$^{3}$T improves out-of-event robustness and attribution over text-only, graph-based, and temporal baselines, and yields interpretable model-based effects: denials/corrections and evidence reduce downstream negativity, while toxicity increases it. We also benchmark open-weight LLM prompting baselines and find that added conversational context helps, but attribution remains less reliable, motivating structure-aware counterfactual modeling for social-media analysis.
cs.CL / 13 / 2609.02158
OBJECTION! Lawyer Agents Mitigate Guilty Bias in Legal Judgment Prediction
Abstract
Legal Judgment Prediction (LJP) models are typically trained on documents that describe facts from a prosecutorial perspective. Existing datasets further exhibit severe label imbalance toward guilty outcomes. Consequently, these models suffer from "Guilty Bias", blindly accepting the prosecution's narrative as objective truth. Previous studies employing three-step reasoning structures or training on synthetically generated innocence data improve overall accuracy, but they still fail to mitigate bias at inference time. In this paper, we introduce OBJECTION, an inference-time pipeline that integrates an Adversarial Lawyer Agent into each 3-step reasoning of offense, unlawfulness, and culpability. Unlike generic critics, our agent actively challenges the model's presumptions of guilt by injecting legal defense arguments at each reasoning stage. To thoroughly evaluate this, we present a new "Natural Innocent" dataset including 3.4k real-world cases, overcoming the limitations of synthetic innocence benchmarks. Test results show that OBJECTION drastically reduces the False Guilty Rate (FGR) from 82.93% (SOTA baseline) to 16.69%, proving its capability to perform substantive legal reasoning. This work denotes a key progress toward aligning Legal AI with the presumption of innocence.
cs.CL / 14 / 2609.02163
Do Cantonese-Adapted Language Models Better Predict Cantonese Reading? A Cross-Model Eye-Tracking Evaluation
Abstract
Information-theoretic measures derived from autoregressive language models are widely used to characterize the expectations that shape human reading, but whether language-variety-specific training improves such psycholinguistic alignment remains unclear. This question is still open for Cantonese, where recent NLP evaluations reported mixed benefits from Cantonese-specific training relative to Mandarin-oriented or general-purpose models. Using naturalistic Cantonese eye-tracking data, we compare two within-family adaptation contrasts: CKIP GPT-2 Tiny versus its lightly Cantonese-adapted JED351 derivative, and Qwen2.5-7B versus CantoneseLLM-7B, which underwent substantially more extensive Cantonese continued pretraining and instruction tuning. From each model, we derive lexical surprisal, POS surprisal, entropy before the target, and entropy reduction. Lexical surprisal and the joint four-metric model consistently favor CantoneseLLM-7B, followed by Qwen2.5-7B, CKIP, and JED351, whereas entropy reduction favors CKIP. These results suggest that more extensive Cantonese-specific training can be associated with stronger predictive fit, while model rankings also depend on the information-theoretic measure being evaluated.
cs.CL / 15 / 2609.02172
Breadth Beats Depth: Improving GCG-Based Jailbreak Optimization with Breadth-Oriented Suffix Search
Abstract
Optimization-based jailbreak attacks such as Greedy Coordinate Gradient (GCG) achieve strong effectiveness and transferability by optimizing adversarial suffixes on white-box source models. However, existing GCG-based methods rely on averaged adversarial loss and deep greedy search, which can over-emphasize easy-to-jailbreak behaviors and overlook promising regions of the suffix space. We propose BOSS, a plug-and-play framework that improves GCG-based jailbreak optimization through breadth-oriented suffix search. BOSS uses Tail-Focused Adversarial Loss (TFAL), standard source loss, and behavior coverage to select terminal suffixes, then explores multiple short trajectories and selectively continues promising suffixes. Experiments on public benchmarks show that BOSS improves attack success rates across multiple GCG-based methods while reducing optimization time.
cs.CL / 16 / 2609.02272
PaperCompiler: Faithful Paper-to-Code Generation via Repository-Level Specification Compilation
Abstract
Faithfully translating research papers into repository-level implementations remains challenging because papers often describe methods at a high level, leave implementation assumptions implicit, and require generated repositories to preserve method logic, evaluation protocols, and cross-file consistency. Despite recent advances in paper-to-code agents, their intermediate outputs are often presented as free-form plans or summaries that downstream coding agents may ignore, reinterpret, or compress, leading to algorithmic simplification and inconsistent repository structure. To address these challenges, we introduce PaperCompiler, a paper-to-code generation framework that compiles paper-grounded evidence into explicit repository-level implementation specifications. PaperCompiler grounds implementation-relevant evidence while preserving source provenance and distinguishing paper-supported, inferred, externally delegated, and unresolved information. The resulting specifications encode non-degradation requirements, ownership assignments, cross-file dependencies, and file-level constraints. Repository generation proceeds under these compiled specifications while retaining flexibility over local engineering choices not fixed by the paper. PaperCompiler outperforms strong baselines on Paper2CodeBench, achieving a 13.8% relative improvement in reference-based fidelity (from 3.64 to 4.15) and reducing high-severity evaluator critiques (from 13.2% to 6.1%).
cs.CL / 17 / 2609.02309
Efficient GUI Agents: A Systems Survey of Observation, Memory, Action, and Runtime Optimization
Abstract
GUI agents increasingly operate across websites, mobile apps, and desktop environments, yet the field still reports progress primarily through task success. We argue that practical deployment depends equally on efficiency: how much context, computation, action budget, and runtime overhead an agent consumes while succeeding. This survey studies efficient GUI agents through an end-to-end systems lens that preserves the current technical axes of observation efficiency, context and memory efficiency, action efficiency, and planner-side/system efficiency. For each subsection, we expand the seed literature through targeted search plus backward and forward citation chaining, then synthesize the dominant mechanisms, reported efficiency signals, and new overheads they introduce. Across the literature, recent progress converges on a small set of recurring ideas: selective reading instead of full-context ingestion, global-to-local visual allocation, recoverable memory rather than raw history replay, verification-aware control, and hybrid runtimes that can switch between GUI and non-GUI execution. We conclude by identifying the main open problems, including honest accounting of verifier cost, cross-benchmark comparability, and co-design of observation, memory, and execution layers under real latency and privacy constraints.
cs.CL / 18 / 2609.02379
MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts
Abstract
While existing work on LLM authorship attribution (AA) has made progress, available benchmarks remain limited, often focusing on English, controlled settings, or relatively outdated models, with the few multilingual studies considering only relatively short texts. We introduce MultiGhostBench, a multilingual benchmark comprising 928 books generated by five recent LLMs across six languages and three scripts, with an average length of approximately 59K words per book. The benchmark supports evaluation under domain, author, and language shifts. Evaluation of representative AA methods shows that no single method consistently performs best across settings, and performance generally degrades under distribution shifts. Transformer-based detectors can retain generator-related information across languages, although transfer effectiveness varies by language pair, whereas statistical and fingerprint-based detectors are more language-dependent. We envision MultiGhostBench as a valuable resource for the development and evaluation of robust AA methods. The dataset and code can be found at https://github.com/GrecoMT/MultiGhostBench.
cs.CL / 19 / 2609.02391
PolERo: Studying Political Evasion in Romanian
Abstract
Political evasion refers to responses that engage with a question while withholding the requested information. Recent NLP work frames political evasion as a classification task using a two-level taxonomy of response clarity and fine-grained evasion strategies. Existing work on response clarity and evasion classification is limited to English, leaving open whether the taxonomy and model behavior transfer across languages and political contexts. We introduce PolERo, a dataset of 3,574 human-annotated question-answer pairs extracted from official transcripts of five Romanian presidents. We evaluate multiple classification approaches on both datasets under matched conditions, including TF-IDF baselines, fine-tuned encoder models, a proposed sliding-window encoder, and zero/few-shot LLM prompting. We study cross-lingual transfer through joint bilingual training and machine-translation-based data augmentation. Our results indicate that fine-tuned encoders are competitive, cross-lingual transfer is asymmetric, and ambivalent evasion categories involving pragmatic cues remain the main challenge across all model families.
cs.CL / 20 / 2609.02414
Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking
Abstract
Multi-turn jailbreak attacks demonstrate that harmful intent can be distributed across dialogue, yet existing methods obscure what conversational mechanisms drive vulnerability. We introduce BLUEPRINT, a safety-evaluation framework separating a factorized social-influence strategy space from WORLDVIEWSIM, a cross-turn situational context module. Monte Carlo Tree Search optimizes turn-level combinations of 18 theory-grounded influence factors across a four-turn trajectory. Across six frontier models, BLUEPRINT achieves near-ceiling ASR on major open-weight and proprietary models, while requiring the fewest average queries (2.46). The resulting trajectories further reveal model-specific vulnerability among resistant targets: each responds to distinct influence factors and strategy transitions, yet all share a common recovery pathway-shifting toward concrete, executable task framing consistently escapes hard-refusal states. Ablations confirm operational cues matter most: making requests actionable has the largest impact, gain framing is unusually potent, and some legitimacy appeals can backfire. These findings suggest robust multi-turn safety requires monitoring not only harmful content, but also how dialogue state makes unsafe requests appear concrete and locally executable.
cs.CL / 21 / 2609.02480
PragAlign: Feedback-Guided Pragmatic Alignment for Controlled Synthetic Dialogue Generation
Abstract
Synthetic dialogue generation can support research in privacy-restricted service settings, but generated conversations must preserve communicative intent, affective meaning, and natural dialogue flow. We introduce PragAlign, a feedback-guided framework for controlled synthetic dialogue generation conditioned on service context, target intent, and target emotion, with auxiliary trait-style controls. PragAlign uses a generate--evaluate--revise loop in which an LLM-based evaluator scores intent alignment, emotion alignment, coherence, fluency, and aggregate quality, then provides criterion-specific feedback for up to three refinement rounds. On 800 matched dialogue specifications, PragAlign achieves 99.50\% evaluator-defined acceptance, compared with 72.25\% for one-shot generation and 95.88\% for repeated generation without structured feedback. This indicates that repeated attempts account for much of the gain over one-shot generation, while structured feedback primarily improves last-mile multi-constraint satisfaction rather than broad average quality. Refinement gains are concentrated in emotion alignment, which is also the dominant failure mode in ablations. A separate human evaluation of 1,200 generated dialogues shows that intent expression and dialogue flow are highly recognizable to annotators, while emotion appropriateness is less stable and more subjective. These results support PragAlign as a quality-control framework for improving evaluator-defined communicative constraint satisfaction, while showing that affective realization and independent human-perceived quality remain open challenges.
cs.CL / 22 / 2609.02606
Predictors of Loneliness in Older Adults Using Multimodal Analysis of Speech and Language
Abstract
Loneliness is a critical public health issue among older adults, linked to higher risks of depression, cognitive decline, and mortality. Scalable, objective methods for its detection remain limited, particularly in natural conversational contexts. We analyzed speech and language markers of loneliness in 310 older adults using semi-structured telephone interviews to help understand how they process feeling lonely and how their language differs at different levels of feeling loneliness. Our multimodal framework combined linguistic features (psycholinguistic dictionaries, n-grams, and topic models) with acoustic features (pitch, tone, loudness) to examine associations with self-reported loneliness scores. Both predefined and data-driven methods captured patterns in verbal content and vocal delivery. Higher loneliness was associated with negations(r = 0.11), negative tone(r = 0.12), and conflict-related language. Lower loneliness was linked to social references(r = -0.18), motivational drives(r = -0.11), and emotional richness in speech(r = -0.12). We also found that the multimodal model (r = 0.298) outperforms the text-only and audio-only models. Findings suggest that loneliness manifests through both linguistic and acoustic cues, supporting the potential of speech-based analysis in psychological assessments and as an early indicator of emotional loneliness when used alongside existing assessments, rather than as standalone diagnostic tools.
cs.CL / 23 / 2609.02639
TaRA: Training-Aware Low-Rank Adaptation Initialization
Abstract
Low-Rank Adaptation (LoRA) has become a de facto standard for parameter-efficient fine-tuning (PEFT), yet its performance is highly sensitive to initialization due to the information bottleneck imposed by low-rank decomposition. Existing approaches attempt to construct high-quality LoRA initializations by exploiting principal components of pretrained weights, activations, or gradients. However, these methods do not directly account for the training dynamics of the full-rank model. In this paper, we propose Training-aware Low-Rank Adaptation Initialization (TaRA), a method that initializes LoRA such that the gradients induced by the low-rank factors closely approximate the gradient of the corresponding full-rank weight matrix. Derived from a mathematical formulation, TaRA improves gradient fidelity at the start of training while introducing negligible computational overhead. Across diverse and challenging fine-tuning tasks, TaRA consistently outperforms prior state-of-the-art methods, establishing a simple, robust, and scalable solution for effective LoRA initialization.
cs.CL / 24 / 2609.02651
WinoQueer-NL: Assessing Bias in Dutch Language Models toward LGBTQ+ Identities
Abstract
While English language models have been widely examined for anti-queer bias, Dutch models remain understudied. To address this gap, we developed a culturally and linguistically adapted Dutch dataset based on the English WinoQueer benchmark, containing pairs of stereotypical and counter-stereotypical sentences. To validate and expand it, we conducted an online survey with 43 Dutch queer participants, confirming 145 of 171 stereotypes as culturally relevant and identifying 22 new biases through free-text responses. The final released dataset, comprising 42,906 sentences, was evaluated using a range of Dutch-specific and multilingual models, including both masked language models (MLMs) and autoregressive language models (ARLMs), with bias measured via a score comparing log-likelihoods of stereotypical versus counter-stereotypical sentences. While the mean bias score across models appeared neutral (~50%), closer analysis revealed significant disparities: some models favored stereotypical sentences up to 97% of the time for transgender identities, but only 6% of the time for gay-related pairs, with transgender and non-binary identities consistently receiving the highest bias scores. Our findings highlight the importance of culturally grounded datasets for evaluating and mitigating biases that disproportionately impact marginalized groups in Dutch language models.
cs.CL / 25 / 2609.02672
oHC: Orthogonal Hyper-Connections on SO(4) via Quaternions
Abstract
Hyper-Connections (HC) replace the single residual stream of a Transformer with $n$ parallel ones, mixing them at every layer with a learned $n \times n$ residual matrix. Leaving that matrix unconstrained places no limit on the factor by which the mixing step rescales the residual streams, and that factor compounds across layers, which destabilizes training. Manifold-constrained Hyper-Connections (mHC) address this by restricting the matrix to the doubly stochastic matrices. That caps the factor at one, so the mixing can no longer amplify any direction, but nothing bounds it from below. We prove that inside this set the mixing step can reduce the norm of the residual streams only by shrinking the differences between the streams, while their mean is left unchanged; and since the reduction accumulates over layers, the streams grow more alike and their diversity is spent with depth. We therefore propose Orthogonal Hyper-Connections (oHC), restricting the residual matrix to the rotation group $SO(n)$, so that the mixing step can neither amplify nor attenuate the residual streams in any direction, which keeps training stable and no longer forces the differences between the streams to contract. Specifically, at the four streams used by recent HC models we parameterize the group in closed form by a pair of unit quaternions, which adds no parameters, replaces the iterative projection with a fixed pattern of signed additions, and can be constructed faster than mHC. We evaluate oHC across a comprehensive set of downstream tasks, where it outperforms the single-stream residual baseline, mHC and iHC, which fixes the residual matrix to the identity.
cs.CL / 26 / 2609.02679
From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs
Abstract
When LLMs support public-facing or high-stakes workflows, missed fabrications can harm users and institutions, while false alarms consume limited human-review capacity. When no trusted context or reference document is available, we study two signals accessible through black-box model APIs: semantic entropy, which measures disagreement among sampled response meanings, and uncertainty derived from token log-probabilities. Their failure modes can be complementary: semantic entropy becomes uninformative when responses form one semantic cluster, while token uncertainty can miss consistently confident errors. We extend token-based uncertainty detection by aggregating token-level signals across sampled responses through our TopK method, evaluate the hybrid CoCoA method, which combines target-response uncertainty with semantic dissimilarity, and propose and study two supervised methods: Gated, which routes single-cluster cases to an aggregated-token-feature classifier, and Stacked, which learns jointly from semantic uncertainty and broader token features. We evaluate seven benchmarks, including five public benchmarks (four text datasets and multimodal handwritten-cheque extraction) and two constructed benchmarks (Financial Summaries and Long-Text QA), using four language models. In our evaluation across models and datasets, Stacked gave the best performance in nearly half of the cases, while TopK and CoCoA remain competitive without supervised training labels, although their thresholds require careful calibration. No method is universally strongest. We therefore evaluate performance at false-positive-rate budgets from 1% to 15%, assess their sensitivity to generation and calibration choices, and examine variation across dataset characteristics.
cs.CL / 27 / 2609.02685
DKL: Decoupled Knowledge Learning for Instruction-Tuned Language Models
Abstract
RAG has become the de facto method for incorporating new, corpus-specific knowledge into an instruction following LLM (Instruct LLM). Although RAG-based prompting improves factual grounding, it fails when retrieval is incorrect or incomplete, leading to hallucinations. Finetuning methods such as RAFT and PA-RAG enhance RAG by injecting new knowledge into the model's parameters, but require generating a massive amount of synthetic QA that covers the entire corpus. Extended Pre-Training (EPT) on the text corpus avoids the need for comprehensive synthetic data generation but compromises an Instruct LLM's instruction-following capabilities, necessitating instruction fine-tuning (IFT) after pre-training. However, IFT is costly and may be infeasible due to the unavailability of an instruction-tuning corpus. In this work, we propose DKL-Decoupled Knowledge Learning for Instruction-Tuned Language Models. Instead of doing EPT on the Instruct LLM, DKL performs EPT on its corresponding base LLM to infuse new knowledge. These knowledge infused weights are then merged with the Instruct LLM, imparting new knowledge without affecting their instruction-following capabilities. DKL is a lightweight method that avoids expensive instruction fine-tuning and relies on model merging to infuse the new knowledge into the Instruct LLM without destroying its instruction following capabilities. Empirical results show that DKL improves RAG accuracy from 54.17 to 79.26 on retrieval failure cases, while outperforming prior approaches with substantially less training data.
cs.CL / 28 / 2609.02702
Trace as State: Reasoning Traces as Conditional States for Long-Context Transformers
Abstract
Transformers process information causally, but long-context reasoning may depend on task state discovered only later. We formalize this mismatch through conditional state update tasks. For causal state update processors, providing the condition first can require exponentially less memory in the worst case than providing it last. Motivated by this principle, we introduce Trace as State. We use collected reasoning traces as a textual proxy for task state and place it before the long-context block on a fresh pass, allowing information derived previously to guide rereading. We conduct extensive experiments on Trace as State and Trace Append, a matched control that uses the same task state proxy but put it after the context. Across three models and three long-context datasets, Trace as State outperforms Trace Append in 26 of 27 reported combinations of model, task, and metric. On GraphWalks Parents, exact match lifts DeepSeek V4 Pro Preview from 29.2% on the initial pass and 43.0% with Trace Appendto 81.8% with Trace as State, and from 66.4% and 83.2% to 100.0% for GLM-5.2. These results show that placing traces before the context can improve long-context reasoning while retaining the causal transformer structure.
cs.CL / 29 / 2609.02735
Choosing a PEFT Variant for Per-Patient Dysarthric ASR: A Single-Speaker Case Study on Two ASR Bases
Abstract
Per-patient adapters are the preferred production architecture for dysarthric automatic speech recognition (ASR), yet parameter-efficient fine-tuning (PEFT) variants have not been compared in the speaker-dependent, per-patient regime. We present a single-speaker case study comparing seven LoRA-family methods (LoRA, QLoRA, AdaLoRA, DoRA, LoHA, VeRA, VB-LoRA) on two production bases (Whisper-large-v3 with Hungarian fine-tuning, and a multilingual Qwen3-ASR-1.7B checkpoint) for one post-stroke Hungarian male speaker (S1, 409 utterances; severe dysarthria on auditory-perceptual clinical assessment). Attention-projection adapters substantially improve CER on both bases. Across three seeds, a paired bootstrap detects no significant LoRA-DoRA difference (p>0.5; 13.86/13.90 % CER on Whisper, 28.10/28.33 % on Qwen3-ASR), so we adopt the simpler, cheaper LoRA. Real 4-bit (NF4) QLoRA is worse on every seed and both bases (14.56/30.09 % CER) with no memory saving at this scale, and LoHA, VeRA, VB-LoRA and AdaLoRA do not reach the LoRA family, though LoHA still gives an 18.6 % relative CER reduction on Whisper. On the same base, full fine-tuning is more accurate (11.43 % CER), but a 115 MB LoRA that also adapts the feed-forward blocks reaches within 0.66 pp of it at approximately 3.7 % of the per-patient storage. A 6-point enrollment grid shows about 5 min of patient audio captures 45.6 % of the zero-shot-to-30-min CER reduction, with further gains at 10 and 30 min (caveat: one speaker, one language, severe post-stroke dysarthria). Training scripts and recipes will be released, source-available under a research-use licence, on publication.
cs.CL / 30 / 2609.02737
Language Models Can Control Their Own Attention
Abstract
Language models spend most of their attention on a small fraction of context, yet they read the entire KV cache to find the few tokens that matter. If the user asks about a previous detail in a 1M-token conversation, global attention layers must scan the full context to generate each token of the reply. A prominent approach mitigates this cost by pre-selecting relevant tokens via lightweight proxy scores, but this extrinsic scoring still incurs O(N) per step. We take an intrinsic approach motivated by the simple question: wouldn't the model already know which parts of the context are relevant? To this end, we introduce Declarative Attention (DA), a protocol that elicits the model to declare where it needs to attend within its chain-of-thought, partitioning generation into three modes: <global> (full context), <focus> (a specific region), and <local> (recent output only). The inference engine parses these declarations like tool calls and skips most of the KV cache read. Under zero-shot evaluation across 15 long-context tasks, DA on off-the-shelf models (Gemma-4-31B, Qwen-3.6-27B) significantly reduces total attended tokens during decoding (52.0%, 31.1%) with modest accuracy drops (1.27pp, 2.75pp) that shrink with model scale. DA unlocks a new axis of sparse attention, with further potential under training-based methods that future work can explore.
cs.CL / 31 / 2609.02771
From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution
Abstract
Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention value depends on both which examples are selected and how they are modified. Influence functions (IF) estimate behavioral changes under infinitesimal reweighting, yet IF-selected examples often show limited advantages over random selection under conventional weight-based interventions. This raises the question of whether influential examples lack intervention value or whether reweighting fails to realize their behavioral leverage.We introduce influence-guided response rewriting, which uses IF to identify intervention targets and replaces their responses with behavior-aligned or behavior-opposed supervision while keeping instructions fixed. Across four open-weight LLMs, we compare rewriting and reweighting on the same influence-selected examples using epistemic abstention as our primary testbed. Response rewriting produces stronger, more persistent, and bidirectional behavioral shifts, while reweighting the same examples yields weak and inconsistent effects. Further analyses show that influence-selected examples provide greater rewriting leverage than alternative selectors, with changes remaining concentrated on target-relevant behaviors. The same qualitative contrast extends to safety refusal. These results distinguish the local reweighting effects captured by influence estimates from the broader intervention leverage of the examples they identify, motivating intervention-aware evaluation of TDA methods.
cs.CL / 32 / 2609.02772
HyperStyler: Low-resource Authorship Style Transfer via Context-aware Style Navigation and Hypernetworks
Abstract
Low-resource authorship style transfer (LAST) aims to rewrite text into the style of an arbitrary target author using only a few reference examples while preserving the original meaning. Existing methods often struggle to achieve both high style fidelity and semantic preservation because they compress diverse references into a single static author embedding, which averages out context-dependent stylistic variation, and rely on hidden representations for style control, which entangle style with content. We propose HyperStyler, a novel architecture that decouples LAST into style selection and style realization. Stylo-navigator predicts style coordinates by jointly modeling the source context and target-author references, and Stylo-hypernet realizes them via dynamic parameter modulation instead of hidden-state injection. Our experiments on Reddit, Blog, and News datasets demonstrate that HyperStyler consistently outperforms prior methods including LLM-based approaches and generalizes robustly across domains. Notably, HyperStyler achieves superior performance with as few as 2.4% additional parameters over T5-large, while being over 1.8x faster than LLMs at inference.
cs.CL / 33 / 2609.02783
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction
Abstract
Evaluating LLM agents is essential for guiding their development, yet it has grown prohibitively expensive: a single pass of a frontier model over an agentic benchmark can cost hundreds to thousands of dollars, a price paid repeatedly across iterative development cycles. Prior efforts, centered on benchmark distillation, reduce the number of evaluation tasks but leave the cost of executing each retained task untouched. In this work, we introduce early outcome prediction, a complementary axis of efficiency that instead cuts cost within each task. Our key insight is that an agent's final outcome is often evident from its intermediate behavior well before execution completes. We instantiate this idea in EarlyEval, a lightweight framework that trains a pair of LightGBM success and failure classifiers over behavioral, textual, and reference-solution features, and halts an agent run the moment either classifier crosses a calibrated confidence threshold, adding negligible per-step overhead. Across three benchmarks, SWE-bench Verified, TerminalBench, and Toolathlon, EarlyEval can eliminate 13%-26% of agent steps and up to 44.1% input tokens and 29.4% output tokens at 89%-97% prediction accuracy, while perturbing per-agent resolve rates by only one to two percentage points on average.
cs.CL / 34 / 2609.02207
LeakageBench: Document-Level Leakage Risk for Redacting Personally Identifiable Information in Document Images
Abstract
Real-world personally identifiable information (PII) redaction often operates on document images---scans, screenshots, and PDF renderings---where OCR errors, layout structure, and visual noise determine whether sensitive information is actually removed. Existing PII benchmarks are mostly text-centric and do not measure document-level redaction risk: a page remains unsafe if even one identifier is missed. We introduce LeakageBench, a challenge set of 500 document images with 11,954 GDPR-aligned PII annotations spanning direct identifiers, linkage keys, and contextual re-identification surfaces. We evaluate generic OCR pipelines, commercial and task-adapted OCR-dependent detectors, and OCR-free vision-language models using entity-level F1, group-wise leakage, and document-level leakage metrics. Code Interpreter raises GPT-5.5 localization F1 from 0.090 to 0.249, but critical page-level leakage remains 0.968. These results show that stronger detection and tool assistance improve localization without making most pages safe for release. LeakageBench provides a diagnostic benchmark for high-recall, spatially grounded PII redaction in document images.
cs.CL / 35 / 2609.02055
Privacy Washing: Detecting Internal Contradictions in Privacy Policies
Abstract
Privacy policies may contain internal contradictions in which commitments are undermined by practices documented elsewhere in the same policy. We operationalize this phenomenon, privacy washing, through a four-stage pipeline: statement extraction, compatibility filtering and natural language inference screening, multi-model judge verification, and thematic analysis, with contradictions confirmed by majority vote of a three-model LLM panel. Applied to two corpora of website privacy policies, 123 collected in 2026 (OPPT) and 115 collected in 2015 (OPP-115), the pipeline finds the same category patterns recurring across the 11-year gap, with third-party sharing contradictions the majority of confirmed cases in each primary run, consistent with structural factors in policy composition rather than necessarily intentional deception. At least one panel-confirmed contradiction appears in 12.2% of OPPT companies (15/123; 9.8% excluding legacy pairs) and 36.5% of OPP-115 companies (42/115). A stability re-run seven months later, with a fully separated configuration (new extraction models, judges from three Chinese providers absent from both corpora, matched filters, no judge-submission similarity threshold), reproduces the OPPT prevalence under the original protocol (13.0% vs. 12.2%), finds sub-threshold pairs confirm at rates of the same order as those above (raising prevalence to 20.3% and 40.9%), and shows the third-party majority is panel-sensitive while the recurrence of the same category pairs is not. Two caveats govern all figures: panel verdicts are not validated against human expert judgment, so precision is unknown and prevalence figures are lower bounds; and the two primary runs used different filter configurations, so their prevalence difference is not interpretable as a corpus or era effect (the matched re-run reduces the gap to roughly twofold but does not eliminate it).
cs.CL / 36 / 2609.02745
Incremental Pooled LLM Evaluation for Cost-Effective Retrieval Model Selection
Abstract
Selecting a retrieval model for a production RAG system requires reliable comparative evaluation, but obtaining relevance judgments at scale is expensive and difficult to repeat as new candidate systems arrive. We study pooled LLM evaluation, in which an LLM judges the union of documents retrieved by the current set of candidate systems, and the pool is then expanded incrementally as new systems are introduced by judging only the new documents they contribute. These judgments are reused to evaluate all systems on a common basis. We validate this approach on four retrieval benchmarks with 11 systems spanning dense, sparse, and hybrid configurations, and deploy it to compare 62 retrieval configurations for a financial news QA system. Pooled LLM rankings correlate strongly with gold-standard evaluation across datasets, and 97% of pairwise system orderings are preserved once bootstrap uncertainty in the qrels is taken into account. In production, document overlap yields 65-80% judgment reuse and up to 4.9x lower evaluation cost, allowing teams to benchmark new retrieval candidates without re-judging previously assessed documents. These results suggest pooled LLM evaluation is a practical and cost-effective workflow for incremental retrieval model selection in deployed systems.
cs.CL / 37 / 2609.02623
Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction
Abstract
Voice actors often re-read the same script while modifying their delivery in response to performance directions. We study this setting as direction-following TTS, where a system generates a new utterance that reflects a given direction relative to a reference utterance while preserving speaker identity and linguistic content. A key challenge is the lack of training data capturing such relative modifications. To address this, we propose a scalable pseudo-triplet construction pipeline that generates~(reference utterance, direction text, modified utterance) triplets. It generates controlled style variations using an impression-controllable TTS model and uses an LLM to produce natural language directions from estimated impression differences. Experimental results demonstrate that pseudo-triplets alone enable stable speaker-preserving modification, and that combining pseudo and recorded data further improves direction alignment while maintaining speaker similarity. Audio examples are available on our demo page https://ntt-hilab-gensp.github.io/IS2026pseudo/
多智能体系统 (cs.MA)
2
cs.MA / 1 / 2609.01838
Differential Games for Compositional Handling of Competing Control Tasks
Abstract
We introduce a novel Divide and Conquer control design methodology leveraging differential games in single-agent, multi-objective dynamical systems. The proposed framework associates each control objective with a virtual input and establishes a non-cooperative, finite or infinite horizon differential game among representative players. Each player optimizes a distinct virtual cost function tailored to its specific goal, the full system state, and the other virtual inputs, while accounting for the remaining players' optimal policies. By establishing a Nash Equilibrium for this game, we synthesize a composite controller that achieves a stable balance across competing objectives, providing control engineers with an intuitive and modular framework for parameter re-tuning throughout the design cycle. We provide formal mathematical derivations for both continuous-time and discrete-time dynamical systems, targeting large-scale single-agent applications where complex, dynamically conflicting control objectives make global weighting intractable. To demonstrate the methodology, we developed an open-source Python package implementing a novel numerical algorithm for solving Coupled Algebraic Riccati Equations arising in infinite-horizon differential games. We evaluate the approach on two benchmark case studies: an inverted pendulum on a cart and a non-linear hierarchically controlled quadrotor. The resulting closed-loop performance is compared against the classical Linear Quadratic Regulator (LQR) across various transient and steady-state control metrics, demonstrating superior trajectory tracking and robust multi-objective regulation.
cs.MA / 2 / 2609.01870
ArcticSwarm: Deferring Early Consensus in Long-Horizon Multi-Agent Research
Abstract
Multi-agent systems have shown strong performance in domains with reliable verifiers such as coding, where multi-parallel candidate generation selected by a verifier is effective. However, such pipelines would not generalize to open-ended, long-horizon research tasks without a verifier. While majority voting or self-consistency is often used to reach consensus as a proxy verifier, parallel agents repeatedly explore the same evidence, while access to peers' partial findings cause search to converge on an early candidate before alternatives are tested. We present ArcticSwarm, a multi-agent research architecture that separates evidence gathering from evidence integration. Subagents publish findings to a shared bulletin board, while gated isolation lets selected search tasks maintain their own prior, preventing early consensus. Structured review at three commitment boundaries enforce only confident candidates to be propagated. As a result, ArcticSwarm reaches 82.6% on the full BrowseComp-Plus set with the open-weight Qwen 3.5-27B model, compared with 78.8% without gated isolation and 74.5% additionally with structured review disabled, outperforming aligned baseline MiroFlow runs (70.6%). Extending to live-web BrowseComp, ArcticSwarm reaches 73.6% with GPT-5, which is well above the reported provider system (54.9%) and MiroFlow (63.4%). Overall, the results show that restricting peer reads during evidence gathering and strengthening commitment boundaries before a hypothesis is shared can broaden search and improve long-horizon multi-agent deep research.
软件工程 (cs.SE)
7
cs.SE / 1 / 2609.01769
From Silicon to Boot Code: Extending Automated Program Repair to Firmware-Layer Security Workarounds
Abstract
Automated program repair (APR) research has been constrained to design time. Current techniques localize and fix bugs in RTL or HLS designs before a chip reaches production. Once a hardware vulnerability surfaces post-silicon, the patch must be manually generated: existing automation addresses patch deployment but not patch synthesis. We study the feasibility of extending a dictionary-guided, localize-synthesize-validate APR methodology originally developed for RTL repair to this firmware layer. An automated commit-clustering miner surfaces recurring fix templates across the EDK II (UEFI) firmware repository's full commit history without depending on known CVE identifiers, recovering all three known CVE-fix campaigns and surfacing two additional candidate bug families. Grounded in real fix evidence, we build four independent localizers: missing speculation barriers in C (CVE-2017-5753, Spectre v1), missing bounds checks before array writes in C (a decompression library CVE), missing Return Stack Buffer stuffing in x86 assembly (CVE-2017-5715), and missing integer-overflow guards in Hand-Off Block creation code (surfaced by the miner itself). All four achieve 100% recall; precision ranges from 2.1-15.5% on the C families to 100% on the assembly and HOB families. Root-cause analysis of the C-family false positives attributes 77-90% to two intra-procedural causes, isolating the inter-procedural alias-analysis gap as a measured 15-20% rather than an estimate. A held-out test confirms Spectre v1 localization holds at 100% recall on unseen files; a fifth, independently built dictionary entry (CVE-2018-3630) shows the methodology extends to a new bug signature at low cost; and a naive syntactic baseline recalls at most 14% where our detector recalls 100%. We frame these results within a broader research agenda for a unified hardware-to-firmware correctness lifecycle.
cs.SE / 2 / 2609.01781
Modelstamp: Pre-Deserialization Verification of Machine-Learning Artifacts and Runtime Environment State
Abstract
Persisted machine-learning models can remain byte-identical while the software environments in which they are loaded evolve, creating a verification problem that artifact integrity checks alone cannot expose. This paper presents Modelstamp, a lightweight Python persistence library for verifying artifact integrity and represented runtime-environment state before deserialization. At persistence time, Modelstamp associates a serialized artifact with a sidecar JSON manifest containing a SHA-256 digest, runtime metadata, and installed versions from a bounded tracked-package set; a separately recorded model-relevant subset determines which package versions participate in drift comparison. Optional HMAC authentication supports workflows in which the producer and verifier share a secret key. At verification time, the artifact and represented current environment are checked against this recorded evidence before the model is deserialized. Modelstamp is evaluated using 14 controlled environment-drift scenarios, eight controlled trust-boundary scenarios, and an artifact-size scaling benchmark from 10 MiB to 1 GiB. The controlled drift experiments behaved as specified across relevant dependency changes, unchanged environments, and unrelated environmental changes, including broader noise controls. The trust-boundary experiments similarly confirmed both intended detections and expected limitations, including shared-key forgery and replay. Median verification time increased from 0.032 s at 10 MiB to 3.334 s at 1 GiB, with measured throughput of approximately 307-312 MiB/s in the benchmark environment. These results characterize Modelstamp as a complementary pre-deserialization reference-state verification control rather than as a replacement for dependency-management systems, malicious-model detection, safe deserialization, or public publisher authentication.
cs.SE / 3 / 2609.01865
ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval
Abstract
Embedding-based code retrieval is a core component of coding agents and retrieval-augmented code generation, where retrieving correct code matters more than retrieving lexically similar code. Existing code-retrieval benchmarks do not plant controlled, execution-verified single-edit variants of each query's canonical implementation in the search pool, leaving the question of whether embeddings can functionally discriminate correct from near-clone-but-incorrect code unanswered in a retrieval setting. Resolving this requires a benchmark whose search pool itself contains the relevant counterfactuals -- execution-verified buggy variants near-identical to each canonical -- so that a retriever's rank ordering can be directly tested for functional discrimination rather than topical or identity overlap. We introduce ExecRetrieval, 939 Python tasks each paired with one execution-verified canonical implementation and up to four execution-verified buggy distractors, each generated by a mechanical mutation making a single targeted edit, and evaluate 23 dense embedding configurations plus BM25 under provider-native invocation with paired McNemar tests and query-level bootstrap intervals. With near-clone counterfactuals in the pool, the top hosted system reaches exec@10 = 1.00 but only exec@1 = 0.331; rank-1 misses are paired buggy variants 91.5-99.4% of the time across the four leading systems, and the canonical scores below at least one of its four paired distractors in 67-78% of queries on the leading systems. The full dataset, execution oracle, embedding matrices, environment snapshot, and pairwise statistical tests are released at the URL in Appendix D.
cs.SE / 4 / 2609.02248
From Prompting to Engineering: A Research Agenda for Prompt Engineering in Software Engineering
Abstract
Prompt engineering is increasingly used across Software Engineering (SE) activities, including requirements analysis, coding, testing, documentation, repository analysis, and planning. Yet prompts and related instruction artifacts are often created and evolved through task-specific and informal practices, with limited support for their systematic evaluation, management, traceability, and governance. To examine how SE can contribute to the maturation of these practices, we organized a structured community discussion at the First International Workshop on Empirical Prompt Engineering for Software Engineering (PROMPT-SE), co-located with EASE 2026. Participants discussed current prompting practices, challenges to their adoption and evaluation, and future directions for integrating prompt engineering into software development. We synthesized these discussions into five areas: prompt artifacts and standardization; evaluation and benchmarking; lifecycle integration; human-AI collaboration and skills; and governance, privacy, and technical debt. Based on these areas, we outline a research agenda to move prompt engineering from predominantly ad hoc interactions toward more systematic, maintainable, evaluable, traceable, and governable SE practices.
cs.SE / 5 / 2609.02591
AgOSS: A Dataset and Multi-Layer Characterization of Open-Source Agricultural Software
Abstract
Much of agriculture depends on open-source software spanning farm management platforms, cloud services, edge gateways, embedded systems, and field-deployed sensors, forming a domain-specific software supply chain that has drawn little empirical security attention. It is unknown whether this ecosystem's supply chain security posture differs from that of comparable non-agricultural software, and if it does, whether the difference reflects the agricultural domain or the size and maturity of the projects within it. As a step towards securing agricultural open-source software, we present AgOSS, a dataset of 66 repositories across six architectural categories. We assess supply chain security within the dataset via OpenSSF Scorecard, governance metrics, SBOM-based dependency analysis, and KEV matching, and compare against matched non-agricultural controls. We report two findings. First, governance is largely independent of inherited dependency risk. Scorecard tracks community activity, but we detect no association with vulnerability count, density, or known-exploited count, and in regression the exposure signal loads on architectural category rather than domain. Second, agricultural projects score far lower on raw Scorecard, but size and maturity confound the gap: the residual loses significance under matching and regression. Securing this ecosystem means investing in contributor capacity and dependency hygiene, not agriculture-specific controls.
cs.SE / 6 / 2609.02753
The Import Tax: A Longitudinal Measurement of Startup Cost in the Python Ecosystem
Abstract
Python programs pay for their imports at every process start, a cost that is invisible in steady-state benchmarks but dominant for command-line tools, test workers, and serverless cold starts. Python 3.15 adds explicit lazy imports (PEP 810) largely on anecdotal evidence; no systematic measurement of the ecosystem's import cost exists. We present one: the 500 most-downloaded PyPI packages, sampled quarterly over five years of releases, measured under six CPython versions (3.9-3.14) on two platforms (Apple M5/macOS and Intel Xeon/Linux), for 63,431 measurements in total, plus direct measurement of 3.15's global lazy-import mode. Import cost is heavily skewed: half of packages import in under 6 ms, but the 99th percentile is 354 ms, the first import after installation costs 3-22x more (bytecode compilation), and importing a package's submodules costs up to 294x more than the top-level import that benchmarks report. The median package's cost grows only +1.6-2.4%/year, but the mean grows +11-13%/year: growth is concentrated in a heavy tail. Newer interpreters import the same code 1.16x slower on macOS, but not on Linux, and a single point release (3.11.5 vs. 3.11.16) swings cost by 1.34x. Global lazy mode makes import statements essentially free, yet breaks 8 of 414 top packages. Harness and dataset are available on request.
cs.SE / 7 / 2609.02782
Type Hints in Python Libraries and Frameworks: An Empirical Analysis of Adoption and Maintenance
Abstract
Context: In Python, type hints allow developers to annotate variables and functions with explicit type information, improving code clarity and reliability. Although type hints are widely available, little is known about how they are adopted and maintained in libraries and frameworks. Objective: We investigate the adoption, usage, maintenance, and rationale of type hints in Python libraries and frameworks. Method: We analyzed 1,000 popular GitHub repositories, identifying libraries and frameworks and extracting their type annotations. We examined annotation coverage, the locations and origins of annotations, their evolution across git histories, and the relationship between developer annotations and types inferred by Pyright. Results: Of the analyzed repositories, 91% of libraries use type hints at least once, although adoption is inconsistent. Among libraries with systematic usage, maintainers prioritize function parameters and return types, with median coverage of 45.8% and 35.9%, respectively, and mainly use built-in types (73.0%). When modified, annotations tend to migrate to more expressive types. Developers annotate members even when Pyright can infer their types, and these annotations often simplify the inferred type. Conclusion: Type hints in Python libraries and frameworks primarily serve as API contracts rather than comprehensive descriptions of implementation details. The findings suggest opportunities for tooling that prioritizes public interfaces, identifies meaningful annotation changes, and supports maintainers in evolving type information.
操作系统 (cs.OS)
1
cs.OS / 1 / 2609.02052
SchedBlame: Who Ran While You Waited? Culprit-Attributed CPU Contention for Containers on Stock Kernels
Abstract
Containers that share a machine compete for CPU. When one slows down, the operator needs to know which co-tenant is responsible, and no deployed signal can say. Pressure stall information, per-cgroup wait counters, and run-queue latency histograms are all victim-side: they report that a container waited, never who it waited for. Recovering the culprit means a kernel patch, full scheduler tracing, or statistical inference: unportable, too costly to leave on, or unreliable when victims coexist. SchedBlame is an eBPF tracer that attributes CPU contention to the cgroups that caused it, on stock kernels, continuously. It inverts the accounting: instead of measuring how long a victim waited, it measures the CPU time every other cgroup consumed while that victim was runnable but not running on the same CPU. The mechanism is a per-CPU bitmap of which measured cgroups are waiting, maintained from the kernel's own runnable counts at four scheduler hooks. Every run slice carries that bitmap, so one 16-byte record charges CPU time to a full row of a competitor x victim blame matrix; the kernel stores no per-pair state. Three properties follow. Slices are self-describing, so userspace holds no waiting state and a lost record costs measurements, not correctness. The measured set is reconfigured by publishing an epoch, invalidating every cache and per-CPU bitmap in constant time while the hooks keep running. Sampling never touches waiting state, so rescaling by the inverse keep probability keeps the estimator unbiased. SchedBlame splits each container's per-second CPU demand into runtime, internal contention, external contention, and throttling, flags anomalies against a rolling 99th-percentile baseline, and names the competitors responsible. In production on unmodified 4.18 and 5.10 kernels, tracking 84 containers on a 96-core host, it costs about 1% of Redis throughput and 6% of one core.
硬件架构 (cs.AR)
5
cs.AR / 1 / 2609.01948
An Emerging NVM-Based On-Chip Training Architecture with Non-Ideality Mitigation Through Bipolar Weight Distributions
Abstract
The rapid advancement of deep learning has presented significant energy efficiency challenges to the conventional von Neumann architecture. In-memory computing (IMC) architectures based on emerging non-volatile memory (eNVM) are widely regarded as a promising solution for accelerating neural network training due to their high parallelism and low power consumption. However, the intrinsic non-idealities of eNVM devices can cause conductance updates to deviate from target values, thereby limiting the performance of on-chip training. To address this challenge, this paper presents a Non-ideality Optimized eNVM Accelerator (NOVA) architecture for on-chip training. Specifically, we first fabricate a two-dimensional (2D) ferroelectric field-effect transistor (FeFET) and develop a conductance modulation behavioral model calibrated with experimental data. Building upon this device model, we propose, for the first time, a Non-ideality Avoidance Training (NAT) algorithm tailored for eNVM devices, which mitigates accuracy degradation by guiding weight convergence toward the most stable conductance regions of eNVM devices. Experimental results demonstrate that, even under severe device asymmetry, NAT improves the accuracy by an average of 15.1\% over the baseline methods across multiple benchmark tasks. Meanwhile, the NOVA achieves an average energy efficiency gain of approximately 33.58$\times$ compared with the peak energy efficiency of graphics processing units (GPUs).
cs.AR / 2 / 2609.02352
Atlas: Algorithm-Hardware Co-Design for On-Device City-Scale 3D Gaussian Splatting in VR
Abstract
3D Gaussian splatting (3DGS) has drawn significant attention in the architectural community recently. However, enabling city scale 3DGS on mobile VR devices remains challenging, as the memory requirement of large scale scenes far exceeds the memory capacity of today's mobile GPUs. This paper presents Atlas, an on device city scale 3DGS rendering framework that enables scalable rendering without runtime Internet access. The key insight is that although the full 3DGS model is massive, each frame only requires a small subset of Gaussians under the current pose and level of detail requirement. Based on this insight, Atlas introduces a hierarchical memory offloading mechanism that dynamically loads only necessary Gaussian data into device memory. To further improve performance, Atlas proposes temporal aware LoD search and stereo rasterization to avoid redundant computation in VR. We further show that our technique can be integrated with existing 3DGS accelerators with negligible hardware overhead. Overall, Atlas achieves 18.5x speedup over the GPU baseline and 3.9x speedup over the state of the art 3DGS accelerators, with 92.4% energy savings.
cs.AR / 3 / 2609.02470
Batch Before You Time: Decision-Scoped Proxy Execution for Timing-Aware Logic Rewriting
Abstract
Standard Delay Format (SDF)-annotated switching simulation distinguishes delay-dependent activity among functionally equivalent rewrites, but evaluating every candidate repeats timing, compilation, and replay. A zero-delay proxy can remove timed evaluations, but generating that proxy candidate by candidate can cost more than the timed work it saves. We present Batch Before You Time (BBYT), which compiles all candidates of one rewrite decision into one scoped zero-delay image and either commits a well-separated proxy winner or invokes the unchanged timed chain. Across 12 counterbalanced holdout sequences, BBYT reduces complete candidate-selection time by 18.05% on average; the design-level reductions are 5.84% and 30.25%, with both confidence intervals above zero. On a counterbalanced 8,192-transition C6288 workload, BBYT is 9.24% faster than the same gate executed with candidate-wise proxy launches. In the five-workload corpus, BBYT removes 32.52% of timed candidate evaluations and matches exhaustive timed selection on all 250 evaluated decisions. For on-demand timing-aware rewrite selection, proxy execution and fidelity continuation should use the same decision scope.
cs.AR / 4 / 2609.02219
Hardware-Accelerated Instance Segmentation for Resource-Constrained Space Robotics with Criticality Analysis
Abstract
Autonomous lunar missions require real-time per- ception under three coupled constraints: extreme low-light conditions, limited onboard compute, and radiation-induced hardware faults that can silently corrupt inference. We present a deployment-oriented instance segmentation framework for resource-constrained lunar robotics that jointly addresses quan- tization calibration and system-level fault exposure under strict compute constraints. First, we introduce Activation Variance Informative Sampling (AVIS), a label-free calibration strategy that deterministically selects calibration samples based on activation variance statistics. Second, we deploy a YOLO-based segmentation model on a Deep Learning Processor Unit (DPU) with architectural modifications that reduce CPU fallback paths and enable statically compiled execution with bounded latency in low-lighting conditions. We further introduce a software-level criticality analysis to estimate fault exposure and guide mitigation under radiation-constrained operation. On a lunar micro-rover platform, AVIS with bias correction recovers 69.8% of quantization-induced accuracy loss while achieving 309 ms inference latency and 5.7 W power consumption. Targeted mitigation reduces global criticality by 31.7%. The results demonstrate an integrated approach and a blueprint for a reliable and safe AI perception framework under space deployment constraints.
cs.AR / 5 / 2609.01771
GadIR: A Spatial-Topology Preserving Compiler for Quantum Many-Body Systems Simulation
Abstract
Simulating quantum many-body systems has been one of the most important applications of quantum computation. For simulation, the Hamiltonian of a physical system is compiled into quantum programs with native instructions for quantum hardware. In previous works, the Hamiltonian is represented as Pauli strings, then compiled and optimized based on the quantum circuit model. Such representation paradigm neglects the spatial topology of original physical models, which is vital information to reducing the overhead of compiling many-body systems Hamiltonians. To address such neglect, we introduce a spatial-topology preserving compiler for quantum many-body simulation. Using Pauli gadgets as the representations of the Hamiltonian, we introduce our intermediate representation -- GadIR, to preserve the spatial-topology information of original physical models. Our compiler frontend performs the group reduction algorithm based on Pauli gadget model, which is a hardware-independent optimization. Our compiler backend performs trotterization and scheduling on Pauli gadgets, then synthesizes the Pauli gadgets into hardware-native quantum programs. We evaluate our compiler on all the canonical quantum many-body system models, while achieving a significant reduction on compilation overhead regarding four major quantum architectures. Overall, our spatial-topology preserving IR exploits the compilation optimization space for quantum many-body systems Hamiltonian.
密码学与安全 (cs.CR)
26
cs.CR / 1 / 2609.01723
Hearing the Whispers: Black-Box Membership Inference Attacks on Finetuned TTS Models
Abstract
Text-to-Speech (TTS) foundation models are increasingly fine-tuned on private datasets to synthesize highly personalized voices, introducing severe privacy risks by exposing both biometric identities and sensitive speech content. Existing black-box membership inference attacks (MIAs) follow a two-stage pipeline of query generation and representation engineering, both of which face unique challenges when adapted to TTS. For query generation, dual conditioning on synthesis text and reference speech creates a large and underexplored query design space with no established criterion for identifying an effective query. For representation engineering, the multi-level speech characteristics and temporal variability of speech make low-level representations and direct comparisons inadequate for capturing membership signals. To address these challenges, we present the first black-box MIA framework explicitly tailored to TTS models at both the speaker and record levels. For query generation, we characterize the feasible query space and establish two criteria, scorable extent and memorization elicitation, for evaluating five representative queries, identifying recitation as the strongest. For representation engineering, we obtain multi-level speech representations from embedding models and temporally align the generated and target audio for fine-grained comparison. Evaluations across three state-of-the-art TTS models (CosyVoice2, F5-TTS, and XTTS-v2) fine-tuned on two benchmark datasets (VCTK and British Dialect) reveal severe privacy leakage: speaker-level AUC remains above 0.80 and approaches 1.0 in the strongest settings, while record-level AUC ranges from 0.80 to 0.90 and remains effective even in challenging scenarios where both members and non-members are of the same speakers. We further identify speech characteristics associated with disproportionate vulnerability to memorization.
cs.CR / 2 / 2609.01730
HEAT: Faster Fully Homomorphic Inference via Approximations-Weights Co-Adaptation
Abstract
Fully homomorphic encryption (FHE) allows a server to run a language model directly on encrypted user prompts, but current approaches remain prohibitively slow. Ciphertexts natively support only addition, multiplication, and rotation, and multiplications may be composed only to a bounded depth before a costly bootstrapping operation is needed to continue. Every nonlinearity must therefore be approximated by an iterative method, and each iteration uses multiplications. A higher iteration count buys precision but exhausts the available depth faster and triggers more bootstraps, which dominate latency. Existing approaches fix the iteration counts uniformly across the model rather than tailoring them to each site's error tolerance. We introduce Homomorphic Encryption-Aware Training (HEAT), a fine-tuning method that makes the per-nonlinearity iteration counts learnable, enabling them and the model weights to co-adapt during training. HEAT optimizes iterations with respect to the task objective, allowing the model to adapt to approximation errors encountered during inference without architectural changes or retraining from scratch. On encrypted GPT-2 decoding, HEAT reduces iterations by $3.1\times$, bootstraps by $1.6\times$, and end-to-end latency by $1.4\times$, while improving decode agreement over the calibrated baseline.
cs.CR / 3 / 2609.01836
Agent Memory Is a Surface for Endogenous Authorization Laundering
Abstract
Long-running LLM agents rely on persistent memory to carry state across interactions, including permissions, restrictions, and revocations. When memory misrepresents this evolving authorization state, the agent's own records can grant authority that the underlying history never permitted, resulting in misaligned behavior without any external attacks. We term this failure endogenous authorization laundering, where spurious permissions written into memory lead to unauthorized actions as their provenance is washed away. We then introduce EAL-Bench, which measures how accurately persistent memory preserves evolving authorization state and whether errors propagate to downstream unauthorized actions. We evaluate five LLMs as memory writers and two as executors across procurement, cybersecurity, and finance. We find that under incremental memory updates, writers create false authority for up to 50.2% of unauthorized requests; once false authority is present, executors act on it in 98.6% of trials. Two safeguards, requiring stored permissions to be backed by valid source events, and tracking permission changes through bounded event sourcing, substantially reduce laundering, but both also reject more legitimate actions, exposing a safety-utility tradeoff. Persistent memory is therefore not merely a performance component, but a part of an LLM agent's effective authorization policy.
cs.CR / 4 / 2609.01856
Adversarial Vulnerabilities of Neural Biomarker Identification Systems
Abstract
There is growing interest in the proposed use of EEG signals as biometric credentials, but thus far there has been little research on the reliability and security of such biometrics. Prior adversarial tests have focused on deep-learning classifiers and assumed attackers have full access to the classifier model. This has left unexamined other, more popular categories of neural signature methods as well as the more realistic case of an adversary having only black-box access to a classifier. In this paper we develop a collection of adaptive attack algorithms which learn to fool an authentication system via targeted alterations of stolen EEG recordings, without requiring any knowledge of the authentication system itself. Tested on 6 public datasets spanning three recording conditions (reacting to visual stimuli, imagining hand movements, and resting), it reveals that different signaturing approaches vary significantly in their degrees of vulnerability to adversarial attacks. We show that vulnerability to spoofing attack is greatly impacted by the recording conditions, with significant variation depending on task at time of recording. Finally, we provide recommendations for improving neural signature biometrics based on the results of our adversarial testing.
cs.CR / 5 / 2609.01931
Agent Flight Recorder: Tamper-Evident Audit Trails with On-Chain Anchoring for Long-Horizon Tool-Using Agents
Abstract
Long-horizon agents execute thousands of actions, resulting in sequential failures rather than isolated errors. When a coding agent deletes a production database or a prompt injection spreads across agents, the incident raises questions of causality, authority, and non-repudiable third-party verification. The Agent Flight Recorder captures each agent action as a structured, canonically serialized event binding eight semantic fields from intent through execution to provenance. Hash chaining and Merkle batching provide tamper evidence and compact inclusion proofs. For cross-organizational disputes where no party's infrastructure qualifies as neutral ground, periodic on-chain anchoring of epoch roots lets any verifier with the disclosed payload and Merkle proof check the record independently, without pre-agreeing on a trusted intermediary. The on-chain footprint is minimal: each anchor stores a 32-byte epoch root and a back-pointer, and no event content touches the chain. We evaluate the system across five cumulative ablation configurations on synthetic agent workloads. The full system adds ~48 microseconds median per-event latency and 512 bytes per event. L2 anchoring costs $2.30 per 100K events at 100-event epochs. The full integrity stack detects edit, delete, reorder, and fork tampering at 100% with zero false positives. Structured forensic queries achieve 1.0 precision on guardrail and delegation lookups where unstructured text search yields 0.013 and 0.077 respectively.
cs.CR / 6 / 2609.01939
Bonded Recourse for Smart-Contract Settlement of Compensable Agent Side Effects
Abstract
Autonomous agent runtimes execute tool actions that mutate databases, repositories, and cloud services across organizational boundaries. Authorization and local compensation cover pre-action admission and in-runtime rollback, but neither settles the residual harm left after a permitted action fails. We design Recourse, a smart-contract settlement protocol for compensable agent side effects that binds each admitted action to scope, recovery, evidence, payout, and collateral. Recourse separates ex ante eligibility from ex post objective settleability: typed receipts make objective residual claims computable under an optimistic-oracle challenge pattern, while subjective or incomplete claims route to ERC-792 arbitration or exclusion. We implement the contract suite, deploy it on Base Sepolia, build adapters against Postgres, Git, and cloud-compatible local sandboxes, and evaluate the system on a deterministic harness, sandbox traces, adversarial sweeps, and property-based fuzzing. Against authorization-only and local-compensation baselines, bonded coverage cuts uncompensated harm. The on-chain tier supplies neutral custody, public challenge, non-cooperative payout, and portable history under cross-organizational trust assumptions.
cs.CR / 7 / 2609.01944
Privacy Amplification Without Independence: How Far Negative Dependence Carries the Guarantees of Poisson Subsampling
Abstract
Poisson subsampling is the default sampler in differentially private optimization because its independence makes privacy amplification tractable. Practical systems, however, are moving toward structured participation: random allocation (balls-in-bins), per-epoch allocation, random check-ins, schemes widely believed to be at least as private as Poisson subsampling at the matched rate. We isolate the probabilistic mechanism behind this belief and delimit it exactly, for Gaussian mechanisms up to correlated-noise matrix mechanisms. (1) If the participation indicator vector is negatively associated (NA), then at every integer Rényi order $α\ge2$, exactly at all finite parameters, its remove-direction Rényi divergence is dominated by that of the marginal-matched independent scheme. For fixed gradient sequences, this extends to the mechanism level whenever the noise strategy's Gram matrix is sign-balanced, an $O(t^2)$-checkable condition. (2) The integer-order restriction is essential. For random allocation with $k=1$, we prove a linear law for the Rényi-difference criterion: at large $t$, dominance reverses for every $α<3/2$, including KL divergence, while the crossing order tends to $3/2$ independently of $σ$. (3) We also localize the known failure of rate-matched Poisson domination exactly: below $(1-q)^t$, the hockey-stick ordering reverses, so substituting the Poisson pair into composition machinery is unsound. An upper-tail argument yields a finite crossover $γ_\star$, connecting this threshold picture to the Rényi boundary at $3/2$. Together, these results give a substitution map for privacy accounting: when Poisson-based computations remain sound for structured participation, where they fail, and what sound alternatives cost in deployment.
cs.CR / 8 / 2609.01945
Pushing Forward Multi-Secret-Key Homomorphic Encryption for Private Average Aggregation
Abstract
Federated Learning enables multiple clients to train a shared model while keeping their local datasets isolated. However, the exchanged model updates may still leak sensitive information, making private aggregation a central building block in practical deployments, especially in the cross-silo setting. Homomorphic Encryption naturally fits the client--aggregator communication pattern of Federated Learning, but conventional single-key deployments rely on strong non-collusion assumptions. Multiparty Homomorphic Encryption removes this limitation, although recent attacks under restricted decryption access require large-variance smudging noise during collaborative decryption, which significantly increases ciphertext size and implementation complexity. In this work, we propose lightweight multi-secret-key protocols for private average aggregation based on RLWE-based Homomorphic Encryption. Our construction departs from the usual multiparty blueprint by avoiding the generation of a collective public key. Instead, each client encrypts its update under its own secret key, while the resulting ciphertexts remain compatible with homomorphic aggregation and collaborative decryption. By explicitly tracking and cancelling the ciphertext noise during decryption, the protocol removes the need for large $λ$-dependent smudging noise. We instantiate the construction with both exact BFV-based and approximate CKKS-based variants, prove its security in the semi-honest model against an adversary corrupting the aggregator and up to $L-1$ clients, and compare its communication and runtime performance with state-of-the-art MHE-based aggregation. Our results show that the proposed approach substantially reduces ciphertext expansion and online cost, while preserving practical homomorphic aggregation performance.
cs.CR / 9 / 2609.02007
C$^2$T-OpenMax: A Novel Open-Set WiFi RF Fingerprinting Method via Center Constrained Learning and Confidence-Guided Tail Modeling
Abstract
Radio frequency fingerprinting (RFF) enables device authentication from transmitter-specific hardware imperfections, but practical deployment requires cross-environment open-set recognition. Data augmentation improves environmental generalization, yet may yield dispersed, low-confidence known-class representations that distort the class statistics used by OpenMax. To address this problem, we propose C$^2$T-OpenMax, an enhanced OpenMax framework combining center-constrained learning with confidence-guided tail modeling. The former improves intra-class compactness, making class-wise representations more suitable for distance-based modeling. The latter retains only correctly classified, high-confidence logits for mean activation vector estimation and Weibull fitting, reducing bias from ambiguous boundary samples. Together, the two modules refine representation geometry and OpenMax construction while preserving augmentation benefits. Experiments on a public WiFi CSI dataset show that C$^2$T-OpenMax achieves the highest open-set accuracy in seven of eight location groups and outperforms all baselines in area under the receiver operating characteristic curve (AUROC) and open-set classification rate (OSCR) across every tested openness level. Under the largest-openness setting, it improves accuracy by 12.31%, AUROC by 0.0887, and OSCR by 0.0856 over the augmented OpenMax baseline.
cs.CR / 10 / 2609.02035
Implicit Manipulation for Skill Selection in LLM Agents with Semantic Matching
Abstract
Skill selection is a key stage in LLM-agent workflows, determining which installed skill should handle a user request. Existing attacks on this stage primarily rely on explicit prompt injection or instruction-level steering, which can expose recognizable manipulation signals. In this work, we identify a new implicit attack surface for skill selection: even when the user prompt and skill description appear benign in isolation, their semantic relationship can still be strategically shaped to favor an attacker-chosen skill. Based on this observation, we present Implicit Skill-Selection Manipulation via Semantic Matching (ISM), which jointly shapes target-skill metadata and reusable prompts to manipulate skill selection without explicit selection instructions. Specifically, we develop a three-stage strategy to broaden semantic coverage, strengthen target distinctiveness, and preserve natural prompt wording. Across four task domains and eight selector models, ISM increases the average target-selection rate (TSR) from 15.2% to 63.5%. In a matched comparison, ISM achieves a 73.5% TSR, only 9.8 percentage points below Explicit Steering. Human reviewers block ISM in only 2.9% of judgments, versus 91.4% for Explicit Steering, while five LLM-based inspectors pass ISM at an average rate of 82.9%, versus 37.4% for Explicit Steering. Moreover, ISM remains effective against PPL-W, Llama Prompt Guard 2, and PIGuard.
cs.CR / 11 / 2609.02048
Type-Directed, Secure-by-Construction Enclave Partitioning for LLVM
Abstract
Trusted Execution Environments (TEEs) provide hardware-supported isolation through enclaves that protect code and data independently of software abstractions. However, TEEs alone cannot enforce information-flow security. This problem is further aggravated in LLVM-like low-level languages that allow unrestricted pointer manipulation and unstructured control flow. Moreover, using TEEs effectively typically requires manually partitioning applications into enclave and non-enclave components, a process that is labor-intensive, error-prone, and lacks fine-grained control. We address these challenges with a three-step approach. First, we formalize SIR, an enclave-oblivious calculus based on LLVM IR, equipped with a novel permissive type system that enforces security against low-level attackers. To obtain meaningful guarantees, SIR combines information-flow control with security-aware coarse-grained memory safety. Second, we extend SIR to SIREN, an enclave-aware calculus that enforces noninterference against stronger attackers capable of observing arbitrary non-enclave memory. Third, we develop a type-driven, type-preserving compilation from SIR to SIREN that automatically produces secure enclave-aware programs, eliminating manual partitioning while providing fine-grained control over host-enclave boundaries. We implement and evaluate SPLITR on thirteen microbenchmarks and real-world workloads, including applications from SGXGauge, on Intel SGX hardware. SPLITR scales to OpenSSL (425,953 LLVM IR instructions) and supports multiple objectives that expose trade-offs among enclave TCB size, host-enclave transitions, and boundary data movement. For OpenSSL, optimizing for transitions reduces them from 393 to 187. Runtime overhead is dominated by fixed enclave costs for short-running workloads, whereas long-running applications better amortize these costs and approach native performance.
cs.CR / 12 / 2609.02127
Stored Is Not Supported: Typed Provenance and Assertion Guardrails for Persistent AI Agents
Abstract
Persistent AI agents construct autobiographical state through reflection, retrieval, and consolidation. Persistence changes availability, not epistemic standing: stored or retrieved material is not thereby supported. Untrusted inputs, prompt injections, and model inferences can therefore enter persistent state and later be presented as agent history or user commitments. We specify typed provenance and assertion guardrails for autobiographical assertion boundedness, a system-relative release property requiring governed statements about the agent, user, or named relationships to satisfy accepted-evidence, temporal-validity, and disclosure policies. A typed provenance graph separates origin, dependency lineage, epistemic role, validity, and disclosure scope. A resolver evaluates authorized state projections and returns one evidential status, orthogonal conflict, staleness, and withholding flags, and a protected decision witness. A generate-verify-revise mediator then checks candidate semantic units before release and renders policy-authorized status responses. Under explicit assumptions about extraction, predicate correctness, resolution soundness, view declassification, and channel mediation, we prove a conditional assertion-boundedness contract. In an executable suite of 24 hand-authored conformance cases, typed mediation passed none of 19 unsafe opportunities unqualified while preserving all five supported controls. The flat/prior and source-tag comparison rules released 19/19 and 18/19 unsafe candidates, respectively. These results validate the encoded resolver and mediator obligations; they do not constitute an end-to-end evaluation of language models or retrieval systems.
cs.CR / 13 / 2609.02208
Agentic Settlement Protocol: An Application Profile for Refundable, Delayed-Fulfilment Agent Commerce on Stablecoin Rails
Abstract
Autonomous agents can already pay per request: HTTP-native protocols such as x402 let an agent sign a stablecoin authorization and receive a resource in the same round trip. That model is atomic and final, which suits metered access and fails commerce: a purchase made on a person's behalf -- a service appointment, a physical order, a flight -- is large, frequently cancelled, and should not become the seller's money until delivery. We present the Agentic Settlement Protocol (ASP), an application profile over on-chain authorize-and-capture escrow (as standardised by the Commerce Payments Protocol) for businesses whose fulfilment is confirmed off-chain by their own order, scheduling, invoicing or booking system, abstracted as a fulfilment engine. ASP contributes: a three-deadline hold model separating the issuance deadline, escrow expiry and the engine's inventory expiry by an explicit submission-inclusion-finality margin, under which no inventory is issued against reclaimable funds; a fulfilment-verification ladder stating who is trusted to trigger capture, what their attestation proves, and under what challenge window; engine-authoritative partial refunds with a refund-liquidity order and per-seller exposure controls, including atomic exposure reservation, that bound the operator's credit risk; distributor revenue share that is provably unprofitable to self-deal; a single-currency-per-charge invariant; and a normative interface specification (x402 scheme, vault ABI, operator and connector APIs, conformance levels) intended to let independent implementations interoperate. The design originated in a review of travel-agency participation in agentic settlement and is instantiated on the XDC Network. This is a design paper; a fault-injection evaluation plan is specified and measurement is left to follow-up work.
cs.CR / 14 / 2609.02268
Retrosynthesis of Synthetic Media for Explainable AI Provenance Forensics
Abstract
With the rapid proliferation of generative models on Machine Learning as a Service (MLaaS) platforms, reliably tracing the provenance of synthetic media without modifying generator architectures or parameters remains a major challenge. In this work, we propose a self-referential retrosynthesis framework for explainable AI provenance forensics under a fixed-generator setting. The framework leverages a jointly optimized encoder-decoder pair to implement a self-embedding mechanism that enables round-trip consistency verification. During inference, client inputs are first encoded and then processed by the generator to produce outputs with high visual fidelity. For forensic verification, the consistency between the resynthesized image and the query image is analyzed to determine whether the image originates from the target generative model. Our approach eliminates the need for watermark embedding or modifications to the generation process. Experimental results show that images generated from encoded inputs maintain visual quality comparable to original generator outputs, while decoded images reliably trace back to their corresponding source inputs. Furthermore, the framework provides interpretable evidence for generative content provenance, establishing a practical tool for explainable generative AI forensics.
cs.CR / 15 / 2609.02393
CAPTCHAs in the Agentic Era: Solvers That Learn from Every Encounter
Abstract
Vision-language models (VLMs) can solve visual CAPTCHAs without task-specific training, but the agents built on them approach every challenge from scratch. For such an agent, the hundredth instance of a familiar puzzle costs as much time and compute as the first. Specialized detectors invert the trade-off, answering in milliseconds but only for categories they were trained on. Neither improves with exposure. We study what changes when a solver improves with use. Our system pairs a fine-tuned YOLOv8 detector with an open-weight VLM behind a confidence-based router, and runs entirely from screenshots and operating-system input events, with no browser automation or DOM access. It reaches 85.4% overall and 84.2% macro accuracy across 16 classes, exceeding either component alone. Every answer VLM produces also serves as a training label, so the detector absorbs categories it was never trained for, typically after one or two encounters and without human annotation. The same loop also repairs it. A CAPTCHA operator can perturb images against the publicly released detector and drive its accuracy to 0%, but the perturbations leave VLM untouched, and its labels let the detector recover. Under a year-long simulated arms race in which the CAPTCHA operator re-crafts its perturbations each month, the solver recovers every round, and a cheap ~70%-accurate open-weight teacher hardens it as effectively as a perfect oracle. Visual CAPTCHA defenses that assume a failing bot stays failing therefore understate how quickly an adaptive solver returns.
cs.CR / 16 / 2609.02469
Evaluating ML-based Intrusion Detection Systems: The Illusion of Model Efficacy
Abstract
Intrusion Detection has been revolutionized due to the integration of Machine Learning. Improved detection rates, reduced false alarms, and optimized algorithms contribute to the perception of improved systems with optimal accuracy and near-perfect performance, the illusion of model efficacy. However, the value of this effectiveness diminishes when confronted with unseen attacks. In this paper, we go beyond solely algorithmic enhancements and metric adjustments in ML-based Network Intrusion Detection Systems. We design an experiment to test the generalization capabilities of certain classifiers on unseen attacks. Our approach examines the dimensionality parameter's impact through two experimental methodologies, which are applied in two distinct settings. The experimental findings reveal how effectively the models could identify even a fraction of unseen attacks and underscore structural weaknesses in ML-based IDS research and evaluation techniques. Finally, seven evaluation criteria are outlined to address these challenges.
cs.CR / 17 / 2609.02532
SpiderSapien: Client-Centric Web Crawler and Security Scanner
Abstract
Black-box web application crawling and scanning play an important role for security testing of web applications. Yet state-of-the-art scanners fall short of addressing key characteristics of a modern web application: its extreme dynamism and interactivity on the client side. This paper identifies immersive interaction as a key ingredient for scanners to deeply explore modern web applications. We propose SpiderSapien, a client-centric crawler and security scanner. SpiderSapien incorporates a unique combination of high-level, user-facing feedback channels from the web application to achieve immersive interaction in a black-box crawling loop. These feedback channels include both novel methods to detect interactable elements and sensibly order UI interactions, and orthogonally using an LLM to solve forms. In doing so, we demonstrate how to reliably discover and test deep states of modern web applications. Furthermore, our modular approach and useful abstraction layer can serve as a building block for future scanners. The evaluation of our approach shows substantial improvements in both code coverage and vulnerability detection over previous work. Our approach increased average code coverage across applications by at least 46% over any other scanner, or 16% when compared to the union of all other scanners. We find XSS vulnerabilities in 7 web applications, while any other scanner finds XSS in up to 2 applications.
cs.CR / 18 / 2609.02568
Learning-Based Reconstruction Attacks on Coordinate-Obfuscated Point Clouds
Abstract
Volumetric video based on point cloud representations enables immersive virtual and augmented reality applications but introduces significant challenges for efficient and secure content delivery. Prior work proposed a selective coordinate encryption framework for point clouds that encrypts only a subset of coordinates, reducing computational costs while visually degrading unauthorized content. However, it remains unclear whether the remaining unencrypted information is sufficient to enable content reconstruction. In this paper, we evaluate the robustness of selective coordinate encryption against machine learning-based reconstruction attacks. We consider an attacker with access to selectively encrypted point clouds attempting to recover encrypted coordinates without decryption by exploiting spatial and geometric correlations in the unencrypted data. We evaluate PointNet and Random Forest models under two encryption granularities: \texttt{X}, where all $X$ coordinates are encrypted, and \texttt{2X}, where every second $X$ coordinate is encrypted. Our results show that reconstructing fully encrypted $X$ coordinates remains challenging, whereas the \texttt{2X} scheme leaks sufficient information through neighboring coordinates to enable accurate reconstruction. These findings demonstrate that the security of selective coordinate encryption depends strongly on encryption granularity.
cs.CR / 19 / 2609.02647
PrimSynth: An Agentic Approach to Discover, Validate, and Synthesize Exploit Primitives for Linux Kernel Vulnerabilities
Abstract
Linux kernel vulnerabilities are critical to downstream systems. Despite extensive research on automated kernel exploitation, a fundamental challenge remains the conceptual gap between abstract exploit strategies and concrete technical operations. To fill this gap, this paper introduces a systematic characterization that formalizes six classes of exploit primitives from logical capability to validatable effect. Then, an extended exploit strategy representation is proposed, which couples primitive upgrading strategies with primitive path code synthesis rules governing object constraints, temporal sequencing, environment prerequisites, and validation constraints. Building upon this foundation, this paper presents \textsc{PrimSynth}, a multi-agent framework that encapsulates these representations through coordinated agents to discover, validate, and synthesize exploit primitives for memory corruption vulnerabilities in the Linux kernel. These agents operate in an iterative closed loop until valid primitives are found, leveraging validation signals as evidence of exploitable state transitions to ground primitive synthesis decisions. An automated method for extracting and validating primitives is also proposed based on vulnerability-directed execution and a rebootable validation environment. \textsc{PrimSynth} is evaluated on 16 real-world Linux kernel CVEs spanning 5 vulnerability types. Experimental results show that PrimSynth achieves reliable primitive extraction, maintaining a 100% primitive match rate. For primitive synthesis, PrimSynth successfully synthesizes multi-primitive exploitation chains with 82.4% strategy synthesis rate (SSR) when the public PoC is available and a 61.3% SSR without the guidance of primitive hypotheses.
cs.CR / 20 / 2609.02716
Card-Based Computation in the Virtual Player Simulation Model
Abstract
Player simulation has recently emerged as a new direction in card-based cryptography, with protocols developed for simulating virtual players in physical card games such as Old Maid, UNO, and President. Unlike conventional card-based secure computation, player simulation imposes additional constraints: the cards represent a persistent game state, the remaining cards in a virtual player's hand must be preserved after each action, and it is desirable to represent each card in the game by a single physical card. In this paper, we study generic card-based computation in the virtual player simulation model. We focus on games whose cards admit a publicly known ranking and propose two fundamental protocols. First, we present the Play-Minimum protocol, which securely selects and plays the minimum-value card from a virtual player's hand when all cards in the deck have distinct values. By symmetry, the protocol can also be used to play the maximum-value card. Second, we present the Sorting protocol, which securely arranges a virtual player's hand in nondecreasing order and remains applicable when multiple cards have the same value. These protocols provide generic computational primitives independent of any particular card game and constitute a step toward understanding the computational capabilities of the virtual player simulation model.
cs.CR / 21 / 2609.02741
SPADE: SPaT Attack Detection from the Connected Vehicle's Perspective
Abstract
Signal Phase and Timing (SPaT) messages are a cornerstone of connected vehicle (CV) safety, enabling CVs to perceive and respond to intersection state through Vehicle-to-Infrastructure (V2I) and Vehicle-to-Vehicle (V2V) communication. The integrity of these messages is threatened by a range of application-layer attacks that can bypass conventional authentication when a roadside unit or peer vehicle is compromised. Existing intrusion detection research either defends the infrastructure side or targets V2V Basic Safety Message (BSM) / Cooperative Awareness Message (CAM) misbehavior, leaving the onboard CV perspective on SPaT integrity unaddressed.To close this gap, we introduce SPADE --- the SPaT Attack Detection and Evaluation dataset --- a labelled, multi-modal, simulation-based dataset designed specifically for deep learning IDS research in this space. SPADE is generated through Eclipse MOSAIC using runtime attack injection at the SAE J2735 application layer across six attack classes and one benign class. By combining four intersection geometries, six operating conditions, and five independent random-seed repetitions, SPADE comprises 180 unique base scenario runs, yielding $\sim$1,890,000 labelled timestep records (270,000 per class). Each record fuses SPaT message fields, onboard camera confidence scores, and cooperative V2V peer data across 40 features, reflecting the multi-modal signal space required to distinguish deliberate attacks from environmental degradation. The dataset, generation code, and scenario configurations are released publicly to support reproducible and comparative IDS research in C-V2X security. The developed toolbox, instructions, and dataset link are publicly available on GitHub: https://github.com/jdinovo/SPADE.
cs.CR / 22 / 2609.02774
CodePoisonRAG: Knowledge Poisoning Attacks on Retrieval-Augmented Code Generation
Abstract
Retrieval-Augmented Code Generation (RACG) improves LLM-based software development by retrieving external code artifacts, documentation, and patches, and incorporating them into the generation context. This reliance on external knowledge introduces a critical trust boundary: poisoned artifacts can influence generated code without modifying the underlying LLM. Prior work shows that selecting existing vulnerable examples can increase the general vulnerability rate of RACG outputs, but leaves open whether a black-box attacker can construct a single task-matched artifact that propagates an attacker-selected weakness. We introduce CodePoisonRAG, a targeted upstream knowledge-poisoning framework that transforms benign fixed-code entries into poisoned artifacts. Its attack chain combines CWE-specific Vulnerability Injection, which embeds a selected source-to-sink flow while retaining task alignment, with Semantic Mislabeling, which adds false safety claims without repairing the vulnerable behavior. The attacker has no access to the victim's deployed knowledge base, retriever, re-ranker, generator, prompt, or defense mechanism and injects at most one artifact per anticipated programming task. We construct 85 poisoned artifacts covering ten CWE classes across Java and C, yielding an aggregate corpus-poisoning ratio of 0.7%. Across three generators, all 85 artifacts appear among the Top-3 results for their corresponding queries, and CodePoisonRAG achieves attack success rates between 0.80 and 0.93. Against CodeGuarder, which injects vulnerability-specific security knowledge into the generation context, the attack retains success rates between 0.40 and 0.71. These results show that RACG poisoning extends beyond the incidental propagation of existing vulnerabilities to the targeted construction and propagation of attacker-selected weaknesses.
cs.CR / 23 / 2609.02866
When Does Authorization End? Effect Closure at Provider Boundaries
Abstract
Revocation completion, clean state, or operation success can leave authorized work able to cause an effect the application rejects while the provider stays within its contract. We call the absence of all such paths policy-relative effect closure, or effect closure for short. Thus, a grant is closed when its existing authorizations retain no such path, and it cannot issue any new ones. We present EFFECTBOUND, which uses an evidence-supported finite contract to decide whether an interface can truthfully report closure while required work completes. It reduces this to finite control with hidden state and returns a strategy, an impossibility certificate, or no verdict when evidence is insufficient. Machine-checked proofs establish the reduction and checker soundness; the checker derives closure results and validates certificates. Across GitHub, Kubernetes, NATS, and Kafka, closure fails in three ways: an interface lacks a needed control, clean visible state hides active work, or the model stops before the effect frontier---the last point where the effect can be prevented. The GitHub tool cannot bind a merge to the reviewed commit; a controlled run confirms that it may merge a different commit. NATS can report no stored or pending messages while dispatched work can still publish downstream. In Kafka, all fixed-set brokers had applied the revocation, yet an earlier authorized request could still append. We add a gate that blocks new use of revoked authority and delays return until earlier in-flight work completes. In a fixed-set Kafka~4.3.1 test deployment, this closes the studied synchronous, nontransactional write path without blocking unrelated requests. For a grant, authorization ends only when issuance stops and no earlier authorization can reach an effect the application rejects.
cs.CR / 24 / 2609.02880
Overcoming the Randomness-Utility Trade-off in Answering Differentially Private Linear Queries
Abstract
We study the question of answering linear queries with differential privacy using few (expected) random bits. We provide a randomness-efficient analog of the $\| \cdot \|_K$-norm mechanism of Hardt and Talwar [HT10]. For the $\ell_\infty$-error, our algorithm can answer $d$ linear queries with $O(d / \varepsilon)$ error using $O(\log d)$ random bits, improving upon algorithms of Canonne et al. and Ghentiyala [CSV25, Ghe26]; this is optimal when $\varepsilon \le 1/d$. We also provide a computationally efficient version of our algorithm, albeit with an $O(\log d)$ multiplicative increase in the error.
cs.CR / 25 / 2609.02328
Poisoning Attacks on the PGM-index
Abstract
The PGM-index (Ferragina and Vinciguerra, VLDB'20) is one of the most practical learned indexes, owing to its theoretical elegance and consistently strong empirical performance. It is built on optimal piecewise linear approximations (PLAs) that minimize the number of segments. In this paper, we ask how sensitive this optimal PLA itself is to poisoning attacks. We propose PGM-attack, an efficient poisoning attack that sequentially inserts adversarial keys to inflate the resulting number of segments, and we develop a method for deriving theoretical upper bounds on the number of segments attainable under arbitrary insertions. Our experiments show that poisoning only 10% of the keys allows PGM-attack to increase the segment count by up to 120x. On every evaluated instance, our instance-dependent upper bound is at most 1.92x the segment count attained by PGM-attack, certifying that PGM-attack achieves at least 52% of the optimum. This increase in the number of segments enlarges the PGM-index by up to 120x. Moreover, the attack also transfers to other learned indexes, substantially inflating the index size of PLA-based ones in particular. Our results reveal that, despite the optimality of its PLAs, the PGM-index has an intrinsic vulnerability rooted in its optimization objective, motivating robustness-aware objective design for future learned indexes. Our code is publicly available at https://github.com/atsukisato/pgm-attack.
cs.CR / 26 / 2609.02376
Removing Speech, Keeping Activities: A Privacy Firewall for Acoustic Sensing in Assisted Living
Abstract
Acoustic sensing offers a promising non-intrusive approach for monitoring daily activities of older adults, yet speech privacy concerns remain a critical barrier to real-world deployment. We present a privacy firewall pipeline based on a U-Net encoder-decoder, trained entirely on synthetic data, that removes speech from ambient audio while preserving environmental sounds indicative of daily activities. Activity recognition is performed using VGGish transfer learning with an SVM classifier. Evaluated on the ESC-50 and SINS datasets across multiple speech content levels, the proposed model reduced residual speech to 0% VAD-detectable speech (Silero Voice Activity Detection) under all tested conditions, outperforming Facebook Denoiser (6.55% residual), SepFormer (36.34%) and ConvTasNet (47.21%) on ESC-50 at the 100\% speech level. On ESC-50 at 40% speech level, classification performance recovers to 85% precision and 85% recall after speech removal, compared with 81%/75% before removal and an 84%/83% speech-free baseline. Evaluation on real-world participant home recordings collected with the AudioHive app showed 0% VAD-detectable speech after processing while maintaining 76% precision and recall. The pipeline enables privacy-preserving acoustic sensing without sacrificing activity recognition performance, addressing a key obstacle to the adoption of ambient monitoring in elderly care.