Skip to content
TILens What matters today in tech
Theme
Topics - AI
Calendar · AUG 2026
M08 Aug 2026
2
9
819 13 14 15 16
17 18 19 20 21 22 23
24 25 26 27 28 29 30
31

Daily edition · AI

The daily ledger

TILens turns technical updates into a focused daily brief: official releases, trusted reporting, and practitioner analysis, deduplicated and organized by topic.

12 Aug 2026 edition
AI

Scaling AI agents with trustworthy data

Business and technology leaders need no convincing that the time of agentic AI is here. Organizations are rapidly adopting agents, and few executives doubt the technology’s potential to transform work. But many…

Source: MIT Technology Review AI MIT Technology Review Insights
AI

Nonlinear multi-study sparse factor analysis

arXiv:2601.18128v2 Announce Type: replace-cross Abstract: High-dimensional data often exhibit variation that can be captured by lower-dimensional factors. For high-dimensional data from multiple studies, one goal is to…

Source: arXiv cs.LG Gemma E. Moran, Anandi Krishnan
AI

Coordinating the Unknown Lipschitz Constant in Multiplayer Bandits

arXiv:2608.10526v1 Announce Type: new Abstract: Motivated by decentralized applications, we study cooperative multi-agent bandits in continuous (Lipschitz) action spaces when the Lipschitz constant is unknown. We…

Source: arXiv cs.LG Ricardo Parada, Chenzhang Zhao, William Chang
AI

MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games

arXiv:2602.24188v2 Announce Type: replace-cross Abstract: We present a scalable and verifiable methodology for evaluating language models in multi-turn interactions, using a suite of collaborative games that require…

Source: arXiv cs.LG Jacob Eisenstein, Fantine Huot, Adam Fisch, Jonathan Berant, Mirella Lapata
AI

Coachable agents for interactive gameplay

arXiv:2607.00642v2 Announce Type: replace-cross Abstract: Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing to…

Source: arXiv cs.LG Roberto Capobianco (Sony AI, Zurich, Switzerland), Harm van Seijen (Sony AI, North America, various locations), Nolan D. Bard (Sony AI, North America, various locations), Neil Burch (Sony AI, North A…
AI

SearchArt: Training Long-Horizon Search Agent with Scalable Synthetic and Verified Task

arXiv:2607.24850v2 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) have enabled search agents to autonomously tackle complex tasks across extended search and reasoning horizons.…

Source: arXiv cs.LG Lang Mei, Xiaohan Yu, Chong Chen, Liyan Liu, Xiangnan Chen, Jinchao Ma, Chao Feng, Li Huang, Siyu Mo, Sichen Kang, Yunkun Xu, Zhihan Yang, Zhujun Xue, Jingren Zhang, Qing He, Yingdi Huang, Hao Jiang,…
AI

Learning Disease-Sensitive Latent Interaction Graphs From Noisy Cardiac Flow Measurements

arXiv:2602.23035v2 Announce Type: replace Abstract: Cardiac blood flow patterns contain rich information about disease severity and clinical interventions, yet current imaging and computational methods fail to capture…

Source: arXiv cs.LG Viraj Patel, Marko Grujic, Philipp Aigner, Theodor Abart, Marcus Granegger, Deblina Bhattacharjee, Katharine Fraser
AI

Efficient Hypergradient Descent for Inverse Reinforcement Learning

arXiv:2608.11052v1 Announce Type: new Abstract: Inverse reinforcement learning (IRL) aims to recover a reward function under which the resulting policy reproduces the behavior observed in expert demonstrations. A…

Source: arXiv cs.LG Nikita Sevriukov, Anna Barabanova, Uliana Gagarina, Karina Ivanova, Sofiia Kasaeva, Ilya Levin, Marina Sheshukova
AI

Sheaf-Based Federated Representation Learning

arXiv:2608.10016v1 Announce Type: new Abstract: Heterogeneous federated systems require agents to learn and exchange informative representations despite differences in data distributions, sensing modalities, model…

Source: arXiv cs.LG Gabriele D'Acunto, Enrico Grimaldi, Valeria Avino, Mario Edoardo Pandolfo, Leonardo Di Nino, Sergio Barbarossa, Paolo Di Lorenzo
AI

Diffract: Spectral View of LLM Domain Adaptation

arXiv:2608.10850v1 Announce Type: new Abstract: We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction, code, and natural…

Source: arXiv cs.LG Nikita Borodin, Maria Krylova, Artem Zabolotnyi, Dmitry Aspisov, Egor Shikov, Nikita Tyuplyaev, Oleg Travkin, Roman Alferov, Dmitry Vinichenko
AI

Iterative Erasure Count Is Not an Affine-Invariant Concept Dimension

arXiv:2608.10566v1 Announce Type: cross Abstract: How many directions does a neural representation use to encode a concept? A common answer repeatedly erases probe directions and reports the stopping count or cumulative…

Source: arXiv cs.LG Tingan Jin, Shuhang Dong, Haosong Li, Chung-Hsien Chou
AI

TimePre: Bridging Accuracy, Efficiency, and Stability in Probabilistic Time-Series Forecasting

arXiv:2511.18539v3 Announce Type: replace Abstract: We propose TimePre, a simple framework that unifies the efficiency of Multilayer Perceptron (MLP)-based models with the distributional flexibility of Multiple Choice…

Source: arXiv cs.LG Lingyu Jiang, Lingyu Xu, Peiran Li, Dengzhe Hou, Qianwen Ge, Dingyi Zhuang, Shuo Xing, Wenjing Chen, Xiangbo Gao, Ting-Hsuan Chen, Xueying Zhan, Xin Zhang, Ziming Zhang, Zhengzhong Tu, Michael Zielew…
AI

Retrieval-Corrected Conformal Prediction for Time Series

arXiv:2608.10553v1 Announce Type: new Abstract: Conformal prediction (CP) provides distribution-free prediction intervals for fixed forecasters, but its standard calibration procedure is often inefficient for time…

Source: arXiv cs.LG Sangjin Jin, Kangmin Kim, Junhyeong Lee, Yongjae Lee
AI

Bayesian Symbolic Regression with Entropic Reinforcement Learning

arXiv:2608.09617v2 Announce Type: replace Abstract: Symbolic regression is the problem of finding an algebraic expression describing a stochastic dependence of a target variable on a set of inputs. Unlike forms of…

Source: arXiv cs.LG Oussama Boussif, Mohammed Mahfoud, Younesse Kaddar, Moksh Jain, Sida Li, Damiano Fornasiere, Xiaoyin Chen, Yoshua Bengio, Esmeralda S. Whitammer
AI

Time-Series Foundation Model Embeddings for Remaining Useful Life Estimation

arXiv:2606.11990v3 Announce Type: replace Abstract: Remaining Useful Life (RUL) prediction is essential for industrial predictive maintenance, yet many learning-based approaches rely on extensive feature engineering or…

Source: arXiv cs.LG Amir El-Ghoussani, Michele De Vita, Ronald Naumann, Vasileios Belagiannis
AI

FiGuRO: Intrinsic Dimension Estimation for Multi-Modal Data

arXiv:2608.10857v1 Announce Type: new Abstract: Determining the complexity, or Intrinsic Dimension (ID), of data is fundamental to efficient and interpretable representation learning. This is particularly challenging in…

Source: arXiv cs.LG Viktoria Schuster, Sana Tonekaboni, Caroline Uhler
AI

Scheduling Mixed RL Rollouts Beyond Prefix Locality

arXiv:2608.11152v1 Announce Type: cross Abstract: Modern reinforcement learning (RL) post-training pipelines for large language models (LLMs) increasingly combine rollout workloads across multiple domains and feedback…

Source: arXiv cs.LG Zetao Hong, Song Yuan, Yuanhao Ding, Yibo Zhu, Daxin Jiang, Zhibin Wang, Chen Tian
AI

Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions

arXiv:2607.19378v3 Announce Type: replace Abstract: Subquadratic alternatives to attention require compromises when applied to multi-dimensional data: standard convolutions lack global receptive fields and input…

Source: arXiv cs.LG David R. Wessels, Farhad Ramezanghorbani, Alireza Moradzadeh, David W. Romero, Olivia Viessmann, Maksim Zhdanov, John St. John, Ken Janik, David M Knigge, Yucheng Tang, Erik J Bekkers, Saee Gopal Pal…
AI

Quantum Incremental Learning with Mixed State Prototypes

arXiv:2608.10464v1 Announce Type: cross Abstract: Incremental learning models are required to learn new classes sequentially without catastrophic forgetting, while operating under parameter and memory constraints. In…

Source: arXiv cs.LG Yu Wu, Qianli Zhou, Xinyang Deng, Wen Jiang, Kang Hao Cheong, Witold Pedrycz
AI

Knowledge-Guided 3D CT Generation: A Conditioning-Centric Taxonomy

arXiv:2608.09992v1 Announce Type: cross Abstract: Controllable generation guided by external knowledge is a key requirement in modern generative deep learning applications, enabling the synthesis of samples with…

Source: arXiv cs.LG Francesca Pia Panaccione, Eugenio Lomurno, Matteo Matteucci
AI

Scaling Self-Play with Self-Guidance

arXiv:2604.20209v2 Announce Type: replace Abstract: LLM self-play algorithms are notable in that, in principle, nothing bounds their learning: a Conjecturer model creates problems for a Solver, and both improve…

Source: arXiv cs.LG Luke Bailey, Kaiyue Wen, Kefan Dong, Tatsunori Hashimoto, Tengyu Ma
AI

High-Dimensional Calibration from Swap Regret

arXiv:2505.21460v2 Announce Type: replace Abstract: We study online calibration of multi-dimensional forecasts over an arbitrary convex set $P \subset \mathbb{R}^d$ relative to an arbitrary norm $|\cdot|$. We connect…

Source: arXiv cs.LG Maxwell Fishelson, Noah Golowich, Mehryar Mohri, Jon Schneider
AI

TideRL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling

arXiv:2608.10402v1 Announce Type: new Abstract: Reinforcement learning (RL) for large language models is moving toward multi-turn agentic workloads, where rollout tasks repeatedly pause for external environments, resume…

Source: arXiv cs.LG Yanyu Ren, Xizheng Wang, Xiao Liu, Bowen Lv, Hanchen Zhang, Shudan Zhang, Hanyu Lai, Shuai Wang, Li Chen, Dan Li, Jie Tang
AI

More Accurate, Less Human: Gestalt Grouping in Vision Models

arXiv:2608.10195v1 Announce Type: cross Abstract: Human vision organizes what it sees into wholes: same-colored points group into series, similar marks cohere into categories, and shapes complete into recognizable…

Source: arXiv cs.LG Sudhanva Manjunath Athreya, Sai Phani Kumar Malladi
AI

HyperShape: Hyperelasticity Across Diverse Shapes

arXiv:2608.09938v1 Announce Type: cross Abstract: Hyperelastic deformations are highly sensitive to domain geometry and boundary conditions, making generalization across both a critical capability for neural operators…

Source: arXiv cs.LG Leo Widmer, Sidaty El Hadramy, St\'ephane Cotin, Philippe Claude Cattin
AI

Measuring Semantic Abstractness of SAE Features via Nonlocality

arXiv:2608.10537v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the corresponding…

Source: arXiv cs.LG Chuqiao Lin, Shivaji Sondhi, Xiao-Liang Qi
AI

AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents

arXiv:2607.26998v3 Announce Type: replace-cross Abstract: Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by tools.…

Source: arXiv cs.LG Ruoyu Wang, Heng Zhao, Renjie Wu, Mengnan Zhao, Zhixuan Chu, Wanyu Lin, Tianhang Zheng
AI

MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale

arXiv:2608.10333v1 Announce Type: new Abstract: LLM agents execute heterogeneous sequences of model calls within a single task: some invocations require careful reasoning, while others are structured steps such as…

Source: arXiv cs.LG Yuhang Yao, Zeyu Wang, Wanyi Chen, Tongyun Yang, Yuhang Han, Jie Xiao, Chengke Bao, Tianyi Zhao, Lynn Ai, Eric Yang, Tianyu Shi
AI

On The Statistical Limits of Self-Improving Agents

arXiv:2510.04399v3 Announce Type: replace-cross Abstract: We develop a learning-theoretic framework for analyzing self-improving agents by decomposing self-modification into five axes. Within this framework, we prove a…

Source: arXiv cs.LG Charles L. Wang, Keir Dorchen, Peter Jin
AI

Multi-Granular Rationale-Guided Molecular LLM for Property Prediction

arXiv:2608.10480v1 Announce Type: cross Abstract: Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discovery. Molecular LLMs represent a…

Source: arXiv cs.LG Junwoo Park, Minyoung Shin, Cheol Soon Lee, Sujee Lee
AI

Regression and Classification with Single-Qubit Quantum Neural Networks

arXiv:2412.09486v2 Announce Type: replace-cross Abstract: The literature reflects a mutually beneficial relationship between machine learning and quantum computing, where progress in one field frequently drives…

Source: arXiv cs.LG Leandro C. Souza, Bruno C. Guingo, Gilson Giraldi, Renato Portugal
AI

Emergent Neural Network Mechanisms for Generalization to Objects in Novel Orientations

arXiv:2109.13445v3 Announce Type: replace-cross Abstract: The capability of Deep Neural Networks (DNNs) to recognize objects in orientations outside the distribution of the training data is not well understood. We…

Source: arXiv cs.LG Avi Cooper, Xavier Boix, Daniel Harari, Spandan Madan, Hanspeter Pfister, Tomotake Sasaki, Pawan Sinha
AI

MRIComp4Flow: Compression of 3D Brain MRI for Training Multi-Modal Generative Models

arXiv:2608.10291v1 Announce Type: cross Abstract: Large-scale multi-modal MRI datasets impose substantial storage and I/O costs, limiting the training of 3D generative models on commodity infrastructure. While lossy…

Source: arXiv cs.LG Lisa K. Fischer, Mykhailo Riabets, Daniel Rueckert, Benedikt Wiestler, Anke Meyer-Baese, Sandeep Nagar
AI

TACTICL: Task-Aware Compression of Tabular ICL Models

arXiv:2608.10837v1 Announce Type: new Abstract: The strong performance of foundation models for tabular tasks comes at substantial inference costs. Distilling models into task-specific architectures reduces model size…

Source: arXiv cs.LG Mykhailo Koshil, Matthias Feurer, Katharina Eggensperger
AI

Measuring and Reducing WebGPU Dispatch Overhead for LLM Inference

arXiv:2608.08730v2 Announce Type: replace Abstract: Large Language Models are deployed to multiple types of environments, from internet browsers to edge devices, and WebGPU serves as a modern cross-platform standard.…

Source: arXiv cs.LG J\k{e}drzej Maczan
AI

Generalized Linear Markov Decision Process

arXiv:2506.00818v2 Announce Type: replace-cross Abstract: Offline reinforcement learning for longitudinal studies often faces two linked challenges: rewards may be binary or bounded, and reward observations may be…

Source: arXiv cs.LG Sinian Zhang, Kaicheng Zhang, Ziping Xu, Zongqi Xia, Jue Hou, Tianxi Cai, Doudou Zhou
AI

ELMER: Evolutionary Language Model that Explores and Refines

arXiv:2608.10196v1 Announce Type: new Abstract: Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space. Syntactic edit size is an unreliable proxy: a…

Source: arXiv cs.LG Matthew Siper, Ahmed Khalifa, Julian Togelius
AI

Why Do Safety Guardrails Degrade Across Languages?

arXiv:2605.17173v2 Announce Type: replace-cross Abstract: Large language models exhibit safety degradation in non-English languages. Standard evaluation relies on Jailbreak Success Rate (JSR), which confounds several…

Source: arXiv cs.LG Max Zhang, Ameen Patel, Sang T. Truong, Sanmi Koyejo
AI

Quantifying the noise sensitivity of the Wasserstein metric for images

arXiv:2510.01015v3 Announce Type: cross Abstract: Wasserstein metrics are increasingly adopted as similarity scores for images. We consider the sensitivity of Wasserstein metrics with respect to pixel-wise additive…

Source: arXiv cs.LG Erik Lager, Gilles Mordant, Amit Moscovich
AI

ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation

arXiv:2608.10905v1 Announce Type: new Abstract: On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Existing methods use…

Source: arXiv cs.LG Ximo Zhu, Ruiqi Liu, Rong Wang, Ping Wu, Xiang Zheng, Wenzhuo Xu, Xubin Yao, Zhiyuan Yan, Bo Li, Jun Gao, Xiaolei Lv
AI

SeFaR: Semantic Feature-aware Robustness Testing of Deep Neural Networks

arXiv:2608.10289v1 Announce Type: cross Abstract: Deep neural networks are increasingly deployed in safety-critical domains as perception modules, where failures are often caused due to rare and under-represented…

Source: arXiv cs.LG Nusrat Jahan Mozumder, Divya Gopinath, Corina Pasareanu, Matthew Dwyer
AI

UniScale: Synergistic Entire Space Data and Model Scaling for Search Ranking

arXiv:2603.24226v4 Announce Type: replace-cross Abstract: Recent advances in Large Language Models (LLMs) have inspired a surge of scaling research in industrial search, advertising, and recommendation systems. However,…

Source: arXiv cs.LG Liren Yu, Caiyuan Li, Feiyi Dong, Tao Zhang, Zhixuan Zhang, Dan Ou, Haihong Tang, Bo Zheng
AI

Detecting Soft Skills in ML Engineering Roles CVs

arXiv:2608.10046v1 Announce Type: new Abstract: Soft skills shape collaboration among ML engineers, data scientists, and software engineers building ML-enabled systems, yet what we know about them comes almost entirely…

Source: arXiv cs.LG Aidin Azamnouri, Nouran Ayad, Justus Bogner, Stefan Wagner
AI

Inverse Design of Inorganic Compounds with Generative AI

arXiv:2604.11827v2 Announce Type: replace-cross Abstract: Machine learning is revolutionizing chemistry. Beyond the value of predictive models accelerating virtual screening, generative AI aims at enabling inverse…

Source: arXiv cs.LG Hannes Kneiding, Luc\'ia Mor\'an-Gonz\'alez, Nishamol Kuriakose, Ainara Nova, David Balcells
AI

Seq2Synth: Benchmarking Temporal Fidelity in Synthetic Sequential Tabular Data

arXiv:2607.15606v2 Announce Type: replace Abstract: Synthetic sequential tabular data are increasingly used for privacy-preserving data sharing and data-driven research, but evaluating their fidelity remains difficult…

Source: arXiv cs.LG Kiwan Kwon, Kangmin Kim, Hojin Lee, Yeseong Jung, Hyeongwoo Kong, Vamsi K. Potluru, Saerom Park, Yongjae Lee
AI

InSight-doc: Agentic Visual Perception for Long-Document Understanding

arXiv:2608.10628v1 Announce Type: cross Abstract: Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone to context rot. In this work, we propose…

Source: arXiv cs.LG Kaican Li, Weiyan Xie, Lewei Yao, Jiannan Wu, Lanqing Hong, Yongxiang Huang, Nevin L. Zhang
AI

Physics-Informed Condition Monitoring of SiC Power Modules

arXiv:2608.08363v2 Announce Type: replace-cross Abstract: Silicon carbide (SiC) power modules are increasingly deployed in automotive traction inverters, where condition monitoring is essential to prevent in-service…

Source: arXiv cs.LG Mattia Scarpa, Evgeny Kusmenko, Francesco Toso, Mattia Bruschetta, Ruggero Carli, Simon Achatz
AI

Representation and Invariance in Reinforcement Learning

arXiv:2112.07752v5 Announce Type: replace-cross Abstract: Researchers have formalized reinforcement learning (RL) in different ways. If an agent in one RL framework is to run within another RL framework's environments,…

Source: arXiv cs.LG Samuel Alexander, Arthur Paul Pedersen
AI

Do AI weather models miss extremes?

arXiv:2608.09972v1 Announce Type: cross Abstract: First-generation AI weather models are often reported to underperform at extremes, mostly in reanalysis-based evaluations of deterministic regression systems. We verify…

Source: arXiv cs.LG Marvin Vincent Gabler, Roberto Molinaro, Niall Siegenheim, Henry Martin, Mark Frey, Niels Poulsen, Philipp Seitz, Olivier Lam
AI

Stochastic Emulation of a Fully Coupled Preindustrial E3SMv3 Simulation

arXiv:2608.10277v1 Announce Type: cross Abstract: We present a stochastic coupled emulator of E3SM version 3, built on the SamudrACE framework, which couples an atmosphere emulator (ACE2) with a full-depth ocean…

Source: arXiv cs.LG Elynn Wu, James P. C. Duncan, Troy Arcomano, Jeremy McGibbon, Oliver Watt-Meyer, Christopher S. Bretherton, Naser Mahfouz, Claudia Tebaldi, Luke Van Roekel, Andrew Roberts, Wuyin Lin, Finn Rebassoo,…
AI

URS: A Unified Neural Routing Solver for Cross-Problem Zero-Shot Generalization

arXiv:2509.23413v3 Announce Type: replace Abstract: Multi-task neural routing solvers have emerged as a promising paradigm for their ability to solve multiple vehicle routing problems (VRPs) using a single model.…

Source: arXiv cs.LG Changliang Zhou, Canhong Yu, Shunyu Yao, Xi Lin, Zhenkun Wang, Yu Zhou, Qingfu Zhang
AI

Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders

arXiv:2608.11197v1 Announce Type: new Abstract: Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. Their analysis…

Source: arXiv cs.LG Nikolai Bolik, Lennart St\"opler, Artur Andrzejak
AI

How Robust Are LLMs to Vietnamese Dialects?

arXiv:2608.10414v1 Announce Type: cross Abstract: Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday communication frequently involves regional dialects that preserve…

Source: arXiv cs.LG Minh Tran, Trinh Chau, Thanh-Nhan Le, Nam Tran, Luan Thanh Nguyen, Cuong Dang, Duc Hoang
AI

A Systematic Sample Size Analysis of ML-Based Path Loss Prediction for LPWAN

arXiv:2608.11083v1 Announce Type: cross Abstract: Low Power Wide Area Networks like LoRa are increasingly deployed for smart city applications, requiring accurate path loss prediction for effective network planning.…

Source: arXiv cs.LG Robert Bitterling, Christian Nettersheim, J\"orn Hees, Michael Rademacher
AI

Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety

arXiv:2512.06227v3 Announce Type: replace-cross Abstract: Real-world indicators play an important role in many Natural Language Processing (NLP) applications, such as life events for mental health analysis and risky…

Source: arXiv cs.LG Junyu Mao, Anthony Hills, Talia Tseriotou, Maria Liakata, Aya Shamir, Dan Sayda, Dana Atzil-Slonim, Natalie Djohari, Pamela Ugwudike, Mahesan Niranjan, Stuart E. Middleton
AI

Proteo-R1: Reasoning Foundation Models for De Novo Protein Design

arXiv:2605.02937v2 Announce Type: replace Abstract: Deep learning in de novo protein design has achieved atomic-level fidelity. However, existing models remain largely non-deliberative: they directly synthesize…

Source: arXiv cs.LG Fang Wu, Weihao Xuan, Heli Qi, Hanqun Cao, Heng-Jui Chang, Zeqi Zhou, Haokai Zhao, Ma Jian, Carl Ma, Yu-Chi Cheng, Kuan Pang, Xiangru Tang, Zehong Wang, Guanlue Li, Hanchen Wang, Kejun Ying, Pan Lu,…
AI

On Effectiveness and Efficiency of Agentic Tool-calling and RL Training

arXiv:2606.00135v2 Announce Type: replace Abstract: Tool-calling is a central component of modern large language model (LLM) agents, equipping them with skills beyond their parametric knowledge. This paper studies…

Source: arXiv cs.LG Tong Liu, Cheng Qian, Matej Cief, Yuan He, Daniele Dan, Nikolaos Aletras, Gabriella Kazai
AI

Lipschitz Dueling Bandits over Continuous Action Spaces

arXiv:2604.00523v2 Announce Type: replace Abstract: We study for the first time, stochastic dueling bandits over continuous action spaces with Lipschitz structure, where feedback is purely comparative. While dueling…

Source: arXiv cs.LG Mudit Sharma, Shweta Jain, Vaneet Aggarwal, Ganesh Ghalme
AI

BreastMammo and DenseMammo: Benchmarks for Mammography Domain Generalization

arXiv:2608.10271v1 Announce Type: cross Abstract: Breast density classification is a critical component of breast cancer risk assessment, yet AI models often struggle to generalize across clinical sites due to…

Source: arXiv cs.LG Hongyi Pan, Gorkem Durak, Halil Ertugrul Aktas, Andrea Mia Bejar, Mustafa Ege Seker, Nebile Alibeyoglu, Rumeysa Guclu, Rana Gunoz Comert Bozkurt, Sibel Ozkan Gurdal, Neslihan Cabioglu, Beyza Ozcinar,…
AI

Delays in Spiking Neural Networks: A State Space Model Approach

arXiv:2512.01906v3 Announce Type: replace Abstract: Spiking neural networks (SNNs) are biologically inspired, event-driven models suited for temporal data processing and energy-efficient neuromorphic computing. In SNNs,…

Source: arXiv cs.LG Sanja Karilanova, Subhrakanti Dey, Ay\c{c}a \"Oz\c{c}elikkale
AI

Adaptive Supervised Anchoring for On-Policy Self-Distillation

arXiv:2608.07935v2 Announce Type: replace Abstract: On-policy self-distillation (OPSD) adapts a language model by distilling guidance from a frozen teacher on trajectories sampled from the student. Its effectiveness,…

Source: arXiv cs.LG Meilin Yang (Renmin University of China, Beijing, China), Zixuan Ding (Renmin University of China, Beijing, China), Jianhao Nie (Renmin University of China, Beijing, China), Weite Zhang (Renmin Unive…
AI

Risk-Averse Wasserstein Distributionally Robust Online Learning

arXiv:2602.20403v2 Announce Type: replace Abstract: We study distributionally robust online learning, where a risk-averse learner updates decisions sequentially to guard against worst-case distributions drawn from a…

Source: arXiv cs.LG Guixian Chen, Salar Fattahi, Soroosh Shafiee
AI

HoloQ-VLA: Uniform W4A4 Quantization of Vision-Language-Action Models

arXiv:2605.28803v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models unify perception, reasoning, and control in a single policy, but their multi-billion-parameter backbones and diffusion-based…

Source: arXiv cs.LG Xinyu Wang, Mingze Li, Sicheng Lyu, Dongxiu Liu, Kaicheng Yang, Ziyu Zhao, Yufei Cui, Xiao-Wen Chang, Peng Lu
AI

ConTact: Contact-First Antibody CDR Design via Explicit Interface Reasoning

arXiv:2605.21600v3 Announce Type: replace Abstract: Computational antibody CDR design methods condition on antigen structure to generate binding loops. Yet, the existing architectures conflate two fundamentally distinct…

Source: arXiv cs.LG Mansoor Ahmed, Spencer VonBank, Nadeem Taj, Sujin Lee, Naila Jan, Murray Patterson
AI

Toward Human Rights Benchmarking for LLMs: A Pilot Methodology

arXiv:2608.10268v1 Announce Type: new Abstract: Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation benchmark exists to assess whether…

Source: arXiv cs.LG Savannah Thais, Wm. Matthew Kennedy, Abhigyan Acherjee, Matilda Wysocki, Malcolm Langford, Caitlin Kraft Buchman
AI

Market Design for AI: Beyond the Copyright Binary

arXiv:2606.12260v3 Announce Type: replace-cross Abstract: How can we design a market of human-generated content for use in training AI models that both enables technological progress and preserves individual incentives…

Source: arXiv cs.LG Yan Dai, Maryam Farboodi, Negin Golrezaei, Sepehr Shahshahani
AI

ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions

arXiv:2608.10621v1 Announce Type: new Abstract: Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formulate safety…

Source: arXiv cs.LG Xinzhe Huang, Biwu Yao, Kedong Xiu, Mengnan Zhao, Di Wang, Puning Zhao, Tianhang Zheng
AI

Information Bottleneck under Perfect Privacy

arXiv:2608.11003v1 Announce Type: cross Abstract: In this work, we study the information bottleneck under perfect privacy, with particular emphasis on the active-rate regime, where the representation-rate constraint is…

Source: arXiv cs.LG Junle Zhong, Mohamad Assaad, Sreejith Sreekumar
AI

MARCO: Click-Intent Decomposition for Calibrated Ads Conversion Prediction

arXiv:2608.10562v1 Announce Type: new Abstract: Not all clicks are equal. Industrial ads ranking decouples conversion probability into click-through rate (CTR) and post-click conversion rate (CVR), yet treats every…

Source: arXiv cs.LG Shiwen Shen, Xiru Huang, Liang Luo, Jianbo Sun, He Lyu, Zihang Fu, Ivonne Xu, Zhizhuo Li, Zhengyu Zhang, Pei-Ju Sung, Yunmiao Wang, Zixuan Wang, Zhengli Zhao, Qiang Jin, Mike Jermann, Mingda Li, Yang…
AI

$\beta$-VAEs as Effective Theories: Tolerance-Dependent Dimension

arXiv:2608.10599v1 Announce Type: new Abstract: In a $\beta$-VAE, increasing the regularization strength acts as a spectral cutoff by collapsing low-utility latent coordinates. In the linear Gaussian VAE, the collapse…

Source: arXiv cs.LG Johannes Hirn
AI

Observational Policy Ranking for SMB Financial Guidance from Multi-Action Accounting Logs

arXiv:2608.10050v1 Announce Type: new Abstract: Small and medium-sized businesses need timely financial guidance, yet historical accounting logs record self-selected and often co-occurring business changes rather than…

Source: arXiv cs.LG Shrutendra Harsola, Vignesh Subrahmaniam, Vikas Raturi, Kamalika Das, Xiang Gao, Kratika Gupta, Ruocheng Guo, Padmaja Jonnalagedda, Ananya Pramod, Sricharan Kumar
AI

Infra-Bayesian Reinforcement Learning Agents Outperform Classical RL For Worst-Case Robustness

arXiv:2605.23146v3 Announce Type: replace Abstract: Classical reinforcement learning assumes the agent interacts with a fixed environment whose behavior does not depend on the agent's policy. This assumption breaks down…

Source: arXiv cs.LG Manish Aryal, Faiyaz Azam, Agnivo Banerjee, Syed Mahir Ahamed, Sai Sidhanth Manoharan Jayanthi, Allegra Laro, Cl\'ement Legentilhomme, Andrew Lin, Florian Lorkowski, Marina P\'erez del Valle, Radman…
AI

Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control

arXiv:2608.10703v1 Announce Type: new Abstract: Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making.…

Source: arXiv cs.LG Haoze Liu, Run Liu, Haiying Xu, Jiahui Han, Siyuan Fang, Siyu Yan, Huiqi Deng, Guanchu Wang, Na Zou
AI

Diffusion-Based Impedance Learning for Contact-Rich Manipulation Tasks

arXiv:2509.19696v4 Announce Type: replace-cross Abstract: Learning-based methods excel at robot motion generation but remain limited in contact-rich physical interaction. Impedance control provides stable and safe…

Source: arXiv cs.LG Noah Geiger, Tamim Asfour, Neville Hogan, Johannes Lachner
AI

Population-Level Generative Modeling for Ranking Data

arXiv:2608.08422v2 Announce Type: replace-cross Abstract: Ranking data arise in scientific and machine learning applications, including recommendation systems, information retrieval, voting, marketing, and AI preference…

Source: arXiv cs.LG Zhaoyang Shi
AI

Status Association Does Not Reliably Predict Decision Leakage

arXiv:2608.10089v1 Announce Type: cross Abstract: Bias evaluations often move too quickly from evidence that a model encodes a social association to claims that the same association will alter consequential decisions.…

Source: arXiv cs.LG Abdullah X
AI

Procedural Fairness Failures in RLHF from Preference Averaging

arXiv:2608.10126v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous preferences into a single reward model, assuming preference homogeneity. When preferences are…

Source: arXiv cs.LG M P V S Gopinadh, Karthik Kamuju, Kummari Avinash, John Joshua, Srinivasa Raju Rudraraju
AI

DACRI: Decision-Aware Causal Intervention Ranking for Critical Supply Chains

arXiv:2608.11154v1 Announce Type: new Abstract: Detecting or attributing a supply-chain disruption is not the same as selecting the intervention that maximizes recoverable net value. We present CriticalSCM-Bench v1, a…

Source: arXiv cs.LG Shiqi Huang, Jiani He, Dingyan Shang, Yihua Xu, Jize Li, Yan Lyu, Lashimi Muraleedharan Nair
AI

CHORUS: Complementary Experts for High-Coverage Testbench Stimulus Generation

arXiv:2608.10090v1 Announce Type: cross Abstract: Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than textual imitation alone. Hardware…

Source: arXiv cs.LG Hejia Zhang, Sheng Lu, Zhongming Yu, Chia-Tung Ho, Brucek Khailany, Jishen Zhao
AI

How to Verify Consistency of Probabilistic Claims

arXiv:2608.11181v1 Announce Type: cross Abstract: When a probabilistic predictor answers many conditional-probability queries, are its answers self-consistent, and can this be verified in polynomial time? This problem…

Source: arXiv cs.LG Orr Paradise, Oliver Richardson, Yoshua Bengio, Shafi Goldwasser
AI

Uncertainty-Aware Ensemble Deep Randomized Neural Networks for Classification

arXiv:2608.10007v1 Announce Type: new Abstract: The current state-of-the-art (SOTA) deep randomized neural networks, such as deep Random Vector Functional Link (dRVFL) and ensemble deep RVFL (edRVFL), treat all training…

Source: arXiv cs.LG M. Sajid, A. Quadir, A. Rahaman, P. N. Suganthan, M. Tanveer
AI

OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories

arXiv:2608.08557v2 Announce Type: replace-cross Abstract: Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing recipe…

Source: arXiv cs.LG Changhao Xiang, Shilin Zhang, Zheng Ma, Kanzhi Cheng, Ruize Ma, Yi Feng, Jianbing Zhang, Zhi Wang, Zhen Wu, Xinyu Dai, Lewei Lu
AI

Can Bayesian Optimization Efficiently Find a Strong Single Expert in Neural Thickets?

arXiv:2608.10867v1 Announce Type: new Abstract: Gradient-free post-training has emerged as a compelling alternative to gradient-based optimization for large language models (LLMs), but existing approaches remain costly.…

Source: arXiv cs.LG Nigel Bastian Cendra, Abdelhamid Ezzerg, Fernando Julio Cendra, Jeremias Knoblauch, Jakob Zeitler
AI

CADET: Context-Conditioned Ads CTR Prediction With a Decoder-Only Transformer

arXiv:2602.11410v2 Announce Type: replace Abstract: Click-through rate (CTR) prediction is fundamental to online advertising systems. While Deep Learning Recommendation Models (DLRMs) with explicit feature interactions…

Source: arXiv cs.LG David Pardoe, Neil Daftary, Miro Furtado, Aditya Aiyer, Yu Wang, Liuqing Li, Tao Song, Lars Hertel, Young Jin Yun, Senthil Radhakrishnan, Zhiwei Wang, Tommy Li, Khai Tran, Ananth Nagarajan, Ali Naqvi…
AI

Stay or Stray - A Dynamical Systems Viewpoint of Popularity Bias

arXiv:2608.10474v1 Announce Type: cross Abstract: Popularity bias in recommendation systems arises when a majority user class generates disproportionate interaction data, causing the system to increasingly favour it…

Source: arXiv cs.LG Sarvesh Shashidhar, Lankireddy Prabhat, Arpit Agarwal, D. Manjunath, Karan Bhukar, Tanmay Khandelwal
AI

MoE Proxy Models for Low-Cost Failure Reproduction and Diagnosis in LLM RL Post-Training

arXiv:2608.10823v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training of large language models (LLMs) is computationally intensive and involves complex system pipelines with substantial debugging…

Source: arXiv cs.LG Yikai Wang, Chuansai Zhou, Yuhang Zhou, Weiqiang Wu, Cong Wu, Yue Deng, Ben Feng, Mingming Zhu, Beirong Zhou, Zhibin Wang, Sheng Zhong, Chen Tian, Wangze Zhang
AI

AgForce Enables Antigen-conditioned Generative Antibody Design

arXiv:2605.21610v2 Announce Type: replace Abstract: Antibody design methods condition on antigen structure to generate complementarity-determining regions (CDR), yet a systematic evaluation of baseline methods reveals…

Source: arXiv cs.LG Mansoor Ahmed, Murray Patterson
AI

Efficient Uncoupled Learning Dynamics with $\tilde{O}\!\left(T^{-1/4}\right)$ Last-Iterate Convergence in Bilinear Saddle-Point Problems over Convex Sets under Bandit Feedback

arXiv:2602.21436v2 Announce Type: replace-cross Abstract: In this paper, we study last-iterate convergence of learning algorithms in bilinear saddle-point problems, a preferable notion of convergence that captures the…

Source: arXiv cs.LG Arnab Maiti, Claire Jie Zhang, Kevin Jamieson, Jamie Heather Morgenstern, Ioannis Panageas, Lillian J. Ratliff
AI

Robust Reputation-Driven Crowdsourced Federated Learning

arXiv:2608.08574v2 Announce Type: replace Abstract: Crowdsourced Federated Learning (CrowdFL) extends traditional federated learning by enabling open and heterogeneous participation through a crowdsourcing paradigm. In…

Source: arXiv cs.LG Mouhamed Amine Bouchiha, Gregory Blanc
AI

EweAcT: Ewe behaviour aligned to accelerometer data for activity monitoring in extensive grazing systems

arXiv:2608.09943v1 Announce Type: cross Abstract: Monitoring livestock behaviour under extensive conditions would provide valuable insights to assess animal adaption to environmental perturbations in agroecological…

Source: arXiv cs.LG Lucile Riaboff (GenPhySE, INRAE), Ny Aina Andriamampandry (GenPhySE, GenPhySE), Jean-Fran\c{c}ois Bompa (GenPhySE, GenPhySE), Mathias Aletru (GenPhySE, GenPhySE), Christian Durand (UEF), S\'ebastien…
AI

Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives

arXiv:2608.11093v1 Announce Type: new Abstract: Cross-view feature matching aims to establish reliable correspondences across images with large viewpoint variations. Over the past decade, the field has evolved from…

Source: arXiv cs.LG Songlin Du, Xiaoyong Lu, Zeyu Wu, Xiaobo Lu, Guobao Xiao, Bin Fan, Jiayi Ma, Takeshi Ikenaga
AI

Recovering Wasted Compute in Autoresearch Agents

arXiv:2608.10424v1 Announce Type: cross Abstract: A slew of recent works develop agents for solving research problems end-to-end, a paradigm increasingly referred to as autoresearch. Such agents have inspired large…

Source: arXiv cs.LG Au Kwok Chun, Abhigyan Acherjee, Amrutha Rao, Zaiqian Chen, Kazem Meidani, C. Bayan Bruss, Micah Goldblum
AI

On the Condition Number Dependency in Bilevel Optimization

arXiv:2511.22331v4 Announce Type: replace-cross Abstract: Bilevel optimization minimizes an objective function, defined by an upper-level problem whose feasible region is the solution of a lower-level problem. We study…

Source: arXiv cs.LG Lesi Chen, Kaiyi Ji, Jingzhao Zhang
AI

The Parser Already Knows: Lightweight Bias Correction in Constrained Decoding

arXiv:2608.10137v1 Announce Type: cross Abstract: Grammar Constrained Decoding (GCD) forces Language Models (LMs) to produce syntactically valid outputs by masking out non-conforming tokens at each step. However, rigid…

Source: arXiv cs.LG I\c{s}{\i}l \"Ozg\"u, Yaoxuan Wu, Guy Van den Broeck, Miryung Kim
AI

GLAM: Efficient Continual Learning at Scale via Grouped LoRA Adapter Merging

arXiv:2509.13211v4 Announce Type: replace Abstract: The ability to learn continuously over time remains a major challenge for modern machine learning systems, even in the era of Foundation Models. While the rich…

Source: arXiv cs.LG Irene Testa, Luigi Quarantiello, Eric Nuertey Coleman, Samrat Mukherjee, Julio Hurtado, Vincenzo Lomonaco
AI

FACT: Failure-Aware Causal Training for World-Action Models

arXiv:2608.10232v1 Announce Type: cross Abstract: Recent world-action models (WAMs) show that co-training policies with future prediction can provide physical priors for action generation. Building on the…

Source: arXiv cs.LG Quanquan Peng, Yutong Liang, Rui Yan, Nicklas Hansen, Xiaolong Wang
AI

V-FiLLM: Verified Financial LLM Reasoning Benchmark

arXiv:2608.11047v1 Announce Type: cross Abstract: While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data remains comparatively less…

Source: arXiv cs.LG Alicia Larsen, Victoire Laurent, Aulia Kharis Rakhamsari, Lara Turgut, Nino Antulov-Fantulin
AI

BREAD: Baseline-Referenced Explanations for Anomaly Diagnosis

arXiv:2608.10587v1 Announce Type: new Abstract: Artificial Intelligence (AI)-based prospective anomaly detection methods are increasingly deployed in high-dimensional and nonlinear settings. Among these approaches,…

Source: arXiv cs.LG Jiaqi Qiu, Rob Goedhart, Jannis Kurtz, Inez M. Zwetsloot
AI

BooST: Bridging Semantics and Motions for Efficient Skill Transfer

arXiv:2608.10600v1 Announce Type: cross Abstract: Skill abstraction---the process of learning reusable and temporally extended behaviors---has emerged as a key paradigm for improving sample efficiency and generalization…

Source: arXiv cs.LG Jusuk Lee, Daesol Cho, Jonghun Shin, Seungyeon Yoo, Jonghae Park, Taekbeom Lee, H. Jin Kim
AI

Instance-Adaptive Online Multicalibration

arXiv:2605.09273v3 Announce Type: replace Abstract: We study online multicalibration beyond the worst-case. We give a single, efficient algorithm which dynamically interpolates between benign and worst-case sequences by…

Source: arXiv cs.LG Zhiming Huang, Jamie Morgenstern, Aaron Roth, Claire Jie Zhang
AI

LVCG: Learning ECG Representations in the Latent Vectorcardiogram Space

arXiv:2605.31249v2 Announce Type: replace Abstract: Electrocardiography (ECG) is a cornerstone of cardiac assessment, making the learning of informative ECG representations fundamental to tasks ranging from disease…

Source: arXiv cs.LG Bosong Huang, Panzhen Zhao, Zengxiang Li, Patricia Lee, Wei Jin, Alan Wee-Chung Liew, Ming Jin, Shirui Pan
AI

KKL Observer Synthesis for Nonlinear Systems via Physics-Informed Learning

arXiv:2501.11655v3 Announce Type: replace-cross Abstract: This paper proposes a novel learning approach for designing Kazantzis-Kravaris or nonlinear Luenberger (KKL) observers for autonomous nonlinear systems. The…

Source: arXiv cs.LG M. Umar B. Niazi, John Cao, Matthieu Barreau, Karl Henrik Johansson
AI

Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks

arXiv:2608.10357v1 Announce Type: new Abstract: Long-horizon tool-using agents must reason over user goals, domain policies, tool calls, simulator state, and delayed verifiable rewards. Reinforcement learning (RL) is a…

Source: arXiv cs.LG Zelei Cheng, Amritansh Mishra, Sambit Sahu, William Campbell
AI

Auto-exploration for online reinforcement learning

arXiv:2512.06244v4 Announce Type: replace Abstract: The exploration-exploitation dilemma in reinforcement learning (RL) is a fundamental challenge to efficient RL algorithms. Existing algorithms for finite state and…

Source: arXiv cs.LG Caleb Ju, Guanghui Lan
AI

Smooth Flow Matching for Synthesizing Functional Data

arXiv:2508.13831v4 Announce Type: replace-cross Abstract: Functional data, i.e., random functions observed over a continuous domain, are increasingly available in areas such as biomedical research, health informatics,…

Source: arXiv cs.LG Jianbin Tan, Anru R. Zhang
AI

Do Judges Behave Like Algorithms?

arXiv:2608.10400v1 Announce Type: new Abstract: What if judges already behave like algorithms? As artificial intelligence and algorithms are deployed in many settings, including the judicial system, many have debated…

Source: arXiv cs.LG Riya Manchanda, Eric Chen, Chloe Zhu, Cynthia Rudin, Brandon Garrett, Songman Kang
AI

BPG: Balancing Plasticity and Generalization for Domain Incremental Learning

arXiv:2608.10804v1 Announce Type: cross Abstract: Deep neural networks excel in various tasks but struggle to generalize across evolving data distributions, leading to significant performance degradation under domain…

Source: arXiv cs.LG Qiang Wang, Songlin Dong, Shaokun Wang, Jizhou Han, Xiang Song, Chenhao Ding, Yuhang He, Yihong Gong
AI

Threshold Structure of Optimal Policies in Restart POMDPs

arXiv:2608.10936v1 Announce Type: cross Abstract: We study a Restart POMDP (Partially Observable Markov Decision Process) on a general Borel state space, where the controller either lets the hidden state evolve…

Source: arXiv cs.LG Konstantin Avrachenkov, Alexey Piunovskiy, Yi Zhang
AI

Temporal Straightening for Latent Planning

arXiv:2603.12231v3 Announce Type: replace Abstract: Learning good representations is essential for latent planning with world models. While pretrained visual encoders produce strong semantic visual features, they are…

Source: arXiv cs.LG Ying Wang, Oumayma Bounou, Gaoyue Zhou, Randall Balestriero, Tim G. J. Rudner, Yann LeCun, Mengye Ren
AI

Optimized Sequential Testing for Binary Ensemble Classifiers

arXiv:2606.15237v1 Announce Type: cross Abstract: Ensemble classifiers are predictive models that combine the results of simpler base models, often by majority vote. A classic example is random forests, which combine…

Source: arXiv cs.LG Joseph Kalman, Amit Moscovich
AI

A lower bound for stepsize-based acceleration of gradient descent

arXiv:2608.10418v1 Announce Type: cross Abstract: Recent work has shown that, for smooth convex optimization, plain gradient descent can be accelerated from its textbook convergence rate of $O(T^{-1})$ (where $T$…

Source: arXiv cs.LG Jianhao Ma, Yuxin Chen
AI

Benchmarking Time Series Generation Methods for Privacy-Preserving Forecasting

arXiv:2608.10891v1 Announce Type: new Abstract: Time series forecasting in privacy-sensitive domains often requires training models on released data rather than original observations. Synthetic time series generation…

Source: arXiv cs.LG Luis Amorim, Vitor Cerqueira, Moises Santos, Paulo J. Azevedo, Carlos Soares
AI

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning

arXiv:2608.10204v1 Announce Type: new Abstract: Safe reinforcement learning maximizes reward subject to safety constraints. For Constrained Markov Decision Processes, the linear-programming view over occupancy measures…

Source: arXiv cs.LG Chenhua Fan, Jiahui Zhu, Yuhang Zhang, Honghao Wei
AI

Spherical Flows for Sampling Categorical Data

arXiv:2605.05629v4 Announce Type: replace-cross Abstract: We study the problem of learning generative models for discrete sequences in a continuous embedding space. Whereas prior approaches typically operate in…

Source: arXiv cs.LG Jannis Chemseddine, Gregor Kornhardt, Gabriele Steidl
AI

Path Integral Value Matching for Linear Quadratic Stochastic Optimal Control

arXiv:2608.10777v1 Announce Type: new Abstract: Linear Quadratic Stochastic Optimal Control (LQ-SOC) establishes a fundamental framework for steering noisy dynamical systems and has recently gained renewed interest in…

Source: arXiv cs.LG Bangyan Liao, Chenglei Yu, Yuchen Yang, Chuanrui Wang, Zhisheng Song, Peidong Liu, Tailin Wu
AI

Optimistic Rates for Multiclass PAC Learning

arXiv:2608.10869v1 Announce Type: new Abstract: Worst-case multiclass bounds do not become smaller when the best classifier is already nearly correct: what is missing is an optimistic rate, a guarantee whose fluctuation…

Source: arXiv cs.LG Xiaoyu Li, Andi Han, Jiaojiao Jiang, Junbin Gao
AI

PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing

arXiv:2608.10109v1 Announce Type: new Abstract: Social media has become a major venue for multilingual communication, where users frequently mix multiple languages within a single utterance. Although code-mixed corpora…

Source: arXiv cs.CL Ghazal Kalhor, Zahra Jafari, Amirarsalan Shahbazi, Behnam Bahrak
AI

StreamFlow: Dynamic Memory Flows for Streaming Video Understanding

arXiv:2608.10949v1 Announce Type: cross Abstract: Streaming video understanding requires multimodal large language models (MLLMs) to preserve relevant evidence from continuously evolving streams under strict causality…

Source: arXiv cs.CL Muxin Fu, Yifan Zhang, Wentao Zhang, Fangming Guo, Qian Chen, Guibin Zhang, Shuicheng Yan, Bo An
AI

Most biomedical publications show signs of LLM-assisted writing

arXiv:2608.10715v1 Announce Type: new Abstract: Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing can be valuable by removing…

Source: arXiv cs.CL Lena Holzwarth, Rita Gonz\'alez-M\'arquez, Dmitry Kobak
AI

From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop

arXiv:2608.11171v1 Announce Type: new Abstract: The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six…

Source: arXiv cs.CL Rahul Gupta, Abhinav Mohanty, Anaelia Ovalle, Anil Ramakrishna, Anubrata Das, Apurv Verma, Jwala Dhamala, Ninareh Mehrabi, Tharindu Kumarage, Yada Pruksachatkun, Yang Trista Cao, Kai-Wei Chang, Aram…
AI

EVIL-Detect for NLPCC 2026 Shared Task 6: LLM-Generated Text Detection

arXiv:2608.10698v1 Announce Type: new Abstract: The rapid development of large language models (LLMs) has increased the need for reliable detection of LLM-generated text, especially in realistic Chinese scenarios…

Source: arXiv cs.CL Hongrui Bao, Hangyu Rong, Zhuoshang Wang, Yubing Ren, Yanan Cao
AI

Bayesian-Agent: Posterior-Guided Skill Evolution Across LLM Agent Harnesses

arXiv:2606.08348v2 Announce Type: replace Abstract: LLM agents increasingly rely on prompts, tools, memory, SOPs, skills, and harness feedback, yet current self-evolution pipelines often update these assets through…

Source: arXiv cs.CL Xiaojun Wu, Cehao Yang, Honghao Liu, Xueyuan Lin, Wenjie Zhang, Zhichao Shi, Xuhui Jiang, Chengjin Xu, Jia Li, Jian Guo
AI

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes

arXiv:2606.11470v2 Announce Type: replace Abstract: Reasoning has become central to how Large Language Models (LLMs) are evaluated and interpreted, spanning Chain-of-Thought (CoT), mathematical problem-solving,…

Source: arXiv cs.CL Avinash Anand, Mahisha Ramesh, Avni Mittal, Ashutosh Kumar, Rishitej Reddy Vyalla, Erik Cambria, Zhengkui Wang, Timothy Liu, Aik Beng Ng, Simon See, Rajiv Ratn Shah
AI

Edge Phoneme Recognition for Children's Speech through Age-Aware Training

arXiv:2608.10206v1 Announce Type: cross Abstract: Detecting phonemes from children's speech has historically been difficult due to the scarcity of training data, and unique characteristics of children's speech. During a…

Source: arXiv cs.CL Matthew Arboleda, Ryan Arboleda, Sophie Haak, Sam Hjelmeset, Andrew Franck, Bingrui Yang, Jose Bustamante Ortiz, Yuanrong Shen, Joel Walsh
AI

Narrative Keyframing for Generative Creative Writing

arXiv:2608.10337v1 Announce Type: cross Abstract: We introduce narrative keyframing, an interaction technique for AI-assisted creative writing that lets writers specify different types of narrative constraints at…

Source: arXiv cs.CL Chao Zhang, Abe Davis
AI

Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems

arXiv:2608.10218v1 Announce Type: cross Abstract: AI agents are becoming more autonomous and increasingly interconnected, exposing them to new emergent risks arising from agent-to-agent interaction. One such risk is the…

Source: arXiv cs.CL Vassilis Papadopoulos, McNair Shah, Sam Zimmerman, Jack Lindsey
AI

Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora?

arXiv:2608.10690v1 Announce Type: new Abstract: Pretraining corpus composition shapes LLM capabilities, but it often remains hidden even when model weights are released. Prior work has inferred corpus mixtures or traced…

Source: arXiv cs.CL Qingjie Zhang, Xingzhang Ren, Zixuan Chen, Jinfeng Li, YueFeng Chen, Yitong Yang, Hui Xue, Dayiheng Liu, Han Qiu
AI

TEAMMix: Taxonomy Enrichment Augmentation and Minority-augmented Mixing Strategy for LLM-enhanced Weak-Supervised Hierarchical Text Classification

arXiv:2608.11044v1 Announce Type: new Abstract: Hierarchical Text Classification (HTC), as a critical text mining task, faces challenges such as complex label hierarchies and class imbalance. Existing methods based on…

Source: arXiv cs.CL Jian Zhang, Zhuohao Yang, Songlin Lei, Bangli Liu, Ziwei Wang, Xufeng Weng, Gehan Amaratunga, Yu Lin, Hongwei Wang
AI

OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents

arXiv:2608.08775v2 Announce Type: replace Abstract: Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in…

Source: arXiv cs.CL Andrea Caciolai, Pere-Llu\'is Huguet Cabot, Chierh Cheng, Albert Ventayol-Boada, Gabriel Mejia Gonzalez, Christophe Ropers, Lucas Bandarkar, Sebastian Ruder, Darlene Sakakihara, Elliot Yun, Pierre An…
AI

InternAgentHarness: A Scalable Synthetic Environment for Enhancing LLM Agentic Abilities

arXiv:2508.08636v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly expected to act as generalist agents capable of solving complex real-world problems. Training such agents, however,…

Source: arXiv cs.CL Xiaozhe Li, Yongkang Chen, Shujian Deng, Peiji Li, Yichuan Ma, Huaxi Huang, Qiye Cai, Tianyi Lyu, Le Ma, Linyang Li, Qipeng Guo, Dahua Lin, Kai Chen
AI

Language corpora for the Dutch medical domain

arXiv:2604.25374v2 Announce Type: replace Abstract: Background: Dutch medical corpora are scarce, limiting NLP development. Methods: We translated English datasets, identified medical text in generic corpora, and…

Source: arXiv cs.CL B. van Es
AI

Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse

arXiv:2608.10810v1 Announce Type: new Abstract: Emotion understanding in discourse requires reasoning beyond surface sentiment because speakers often convey affect through indirect, implicit, polite, ironic, or…

Source: arXiv cs.CL Zhenyan Zheng, Yunyao Zhang, Junxi Sheng, Junqing Yu, Zikai Song
AI

ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models

arXiv:2608.09548v2 Announce Type: replace Abstract: Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question…

Source: arXiv cs.CL Yilin Jiang, Xiaorong Zhu, Fei Tan, Zicheng Zhang, Kaiyi Huang, Yang Yu, Zexuan Fei, Yiming Luo, Keqian Li, Hao Hao, Guangtao Zhai, Aimin Zhou
AI

Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences

arXiv:2608.10908v1 Announce Type: cross Abstract: As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the judgment crisis. While…

Source: arXiv cs.CL Martina Ianaro, Guilherme Fernandes, Maurizio Gabbrielli, Joao Magalhaes
AI

From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models

arXiv:2608.10444v1 Announce Type: new Abstract: Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains. This progress primarily…

Source: arXiv cs.CL Si'an Xie (Beijing University of Posts and Telecommunications), Jiaxun Liu (Peking University), Biao Yang (Kuaishou Technology), Wei Yuan (Kuaishou Technology), Fan Yang (Kuaishou Technology), Tingti…
AI

Mitigating Context Interference for Reliable and Efficient Search Agents

arXiv:2608.10743v1 Announce Type: new Abstract: Recent research empowers Large Language Models (LLMs) as multi-turn search agents to iteratively retrieve and generate outputs until complex tasks are solved. However, the…

Source: arXiv cs.CL Boyang Xue, Bin Wu, Shuofei Qiao, Sheng Wang, Rui Wang, Yiming Du, Hongru Wang, Jeff Z. Pan, Emine Yilmaz, Kam-Fai Wong, Aldo Lipani
AI

Data Attribution of Emergent Misalignment with Persona Features

arXiv:2608.11025v1 Announce Type: new Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic…

Source: arXiv cs.CL Clemens Vetter, David Kacz\'er, Lucie Flek, Florian Mai
AI

Simplex Relaxation for Discrete Diffusion

arXiv:2608.10615v1 Announce Type: new Abstract: Discrete diffusion models for categorical generation are defined by a corruption kernel, which determines the intermediate state space and the associated reverse…

Source: arXiv cs.CL Jinya Sakurai, Patrick Pynadath, Satoshi Hayakawa, Jaehong Yoon, Xulei Yang, Nancy F. Chen, Xun Xu
AI

LLM Agents Factory: Retrieval of Domain-Specific LLM Agents

arXiv:2608.09934v1 Announce Type: new Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practical deployment is often limited by…

Source: arXiv cs.CL Vitalii Belov, Artyom Sosedka, Andrey Sakhovskiy, Elizaveta Kovtun, Artyom Boyarskikh, Semen Budennyy
AI

No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

arXiv:2503.05061v3 Announce Type: replace Abstract: Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as business and finance.…

Source: arXiv cs.CL Michael Krumdick, Charles Lovering, Varshini Reddy, Seth Ebner, Chris Tanner
AI

Faster Superword Tokenization

arXiv:2604.05192v2 Announce Type: replace Abstract: Byte Pair Encoding (BPE) is a widely used tokenization algorithm, whose tokens cannot extend across pre-tokenization boundaries, functionally limiting it to…

Source: arXiv cs.CL Craig W. Schmidt, Chris Tanner, Yuval Pinter
AI

Multiplayer Nash Preference Optimization

arXiv:2509.23102v4 Announce Type: replace-cross Abstract: Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences. However,…

Source: arXiv cs.CL Fang Wu, Xu Huang, Weihao Xuan, Zhiwei Zhang, Yijia Xiao, Guancheng Wan, Xiaomin Li, Bing Hu, Peng Xia, Jure Leskovec, Yejin Choi
AI

Cost-Efficient Estimation of General Abilities Across Benchmarks

arXiv:2604.01418v2 Announce Type: replace Abstract: Thousands of diverse benchmarks have been developed to measure the quality of large language models (LLMs). Yet prior work has demonstrated that LLM performance is…

Source: arXiv cs.CL Michael Krumdick, Adam Wiemerslage, Seth Ebner, Charles Lovering, Chris Tanner
AI

FaithformBench: Benchmarking Faithfulness of Mathematical Chain-of-Thought Autoformalisation

arXiv:2608.10916v1 Announce Type: new Abstract: Autoformalisation (AF) systems map natural language reasoning steps into formal statements in a proof assistant such as Lean. We consider how to assess the faithfulness of…

Source: arXiv cs.CL Rob Cornish, Iacopo Ghinassi, Po-Hung Yeh, Shuqi Liu, Qiyuan Xu, Haoxuan Yin, Dominik Wagner, Wenda Li, Yee Whye Teh, Luke Ong
AI

LiFT: How to Enable In-Context Learning for Longitudinal Modelling

arXiv:2604.16382v2 Announce Type: replace Abstract: Longitudinal NLP tasks such as mental health monitoring and stance evolution require modeling temporally ordered text to track persistence and detect change. Such…

Source: arXiv cs.CL Iqra Ali, Talia Tseriotou, Mahmud Elahi Akhter, Yuxiang Zhou, Maria Liakata
AI

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

arXiv:2608.10692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple…

Source: arXiv cs.CL Junjie Ye, Zhuohui Sheng, Shaofan Liu, Yulun Zhu, Wenjie Fu, Dingwei Zhu, Ming Zhang, Yujiong Shen, Weichao Wang, Xin Zhao, Shihan Dou, Tao Gui, Qi Zhang, Xuanjing Huang, Pluto Zhou
AI

Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design

arXiv:2608.10299v1 Announce Type: new Abstract: Agentic systems are increasingly expected to improve after deployment, yet single-entity self-evolution is often bounded by a static learning context, such as fixed tasks…

Source: arXiv cs.CL Qing Zong, Jiayu Liu, Junhao Shen, Zecong Tang, Linsi Wu, Yuxuan Liu, Rui Wang, Zhaowei Wang, Weiqi Wang, Cheng Qian, Xiusi Chen, Yangqiu Song
AI

OpenPM: Auditable Point-in-Time Evaluation for LLM Portfolio-Management Agents

arXiv:2608.09988v1 Announce Type: cross Abstract: Large language models are increasingly used to read markets, assess risk, and allocate capital. However, reported results for LLM trading agents can be inflated by…

Source: arXiv cs.CL Xinying Cai, Minghao Guo, Jiahe Liu, Jiaojiao Han, Bangwei Guo, Yitao Long, Yuxuan Chen, Bohan Wu, Dimitris N. Metaxas, Raymond Li
AI

FlexSQL: Flexible Exploration and Execution Make Better Text-to-SQL Agents

arXiv:2605.02815v2 Announce Type: replace Abstract: Text-to-SQL over large analytical databases requires navigating complex schemas, resolving ambiguous queries, and grounding decisions in actual data. Most current…

Source: arXiv cs.CL Quang Hieu Pham, Yang He, Ping Nie, Canwen Xu, Davood Rafiei, Yuepeng Wang, Xi Ye, Jocelyn Qiaochu Chen
AI

VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?

arXiv:2608.10408v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world visualization…

Source: arXiv cs.CL Mizanur Rahman, Arshia Azimlu, Shadikur Rahman, Md Tahmid Rahman Laskar, Amran Bhuiyan, Shafiq Joty, Enamul Hoque Prince
AI

DuplexWorld: Can voice agents help you get through the day?

arXiv:2608.10716v1 Announce Type: cross Abstract: Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the…

Source: arXiv cs.CL Aryan Vijay Bhosale, Harshit Rajgarhia, Akhil Pothanapalli, Asif Shaik, Abhishek Mukherji, Dinesh Manocha
AI

Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR

arXiv:2608.10670v1 Announce Type: new Abstract: At corpus sizes typical of low-resource dialects, single-run comparisons can yield gains that do not replicate. We show this for Garhwali, an under-resourced Indo-Aryan…

Source: arXiv cs.CL Karamvir Singh Batra, Prathamjyot Singh, Ashima Sood, Jasmeet Singh, Sahil Sharma
AI

The Illusion of Cross-Lingual Safety in Low-Resource Languages

arXiv:2608.11146v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. However, this…

Source: arXiv cs.CL Abigail Oppong, P Sam Sahil, Tadesse Destaw Belay, Maryam Ibrahim Mukhtar, Esmael Ahmed Abdu, Tassallah Abdullahi, Jessica Oparebea, Saminu Mohammad Aliyu, Idris Abdulmumin, Abubakar Juma Chilala, Ni…
AI

Multimodal QUD: Inquisitive Questions from Scientific Figures

arXiv:2604.23733v2 Announce Type: replace Abstract: Discourse comprehension in complex documents often involves continuously posing and resolving Questions Under Discussion (QUDs). While QUD frameworks have so far…

Source: arXiv cs.CL Yating Wu, William Rudman, Venkata S Govindarajan, Alexandros G. Dimakis, Junyi Jessy Li
AI

Evo-Bench: Can Language Models Improve Agent Harness?

arXiv:2608.09096v2 Announce Type: replace Abstract: Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is…

Source: arXiv cs.CL Lisheng Huang, Chen Yang, Hao Zhou, Huatong Song, Zongchao Chen, Ran Le, Yang Song, Wayne Xin Zhao, Tao Zhang
AI

Self-Knowledge Retrieval Augmented Generation Framework for Patent Matching

arXiv:2608.11030v1 Announce Type: cross Abstract: Patent retrieval and matching based on large language models (LLMs) play a vital role in intellectual property protection. However, due to the complex structure of…

Source: arXiv cs.CL Jian Zhang, Songlin Lei, Zhuohao Yang, Bangli Liu, Ziwei Wang, Xufeng Weng, Gehan Amaratunga, Yu Lin, Hongwei Wang
AI

Living-Harness Is an Interactive-Agent Evolver

arXiv:2607.26598v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents may recover from a failure within an episode or after a retry, yet the same execution failure can recur in later tasks because…

Source: arXiv cs.CL Yuetian Du, Yucheng Wang, He Xu, Jiexu Xu, Shanwen Tan, Bing Zhao, Boyu Yang, Zhijie Xu, Ming Kong, Hu Wei, Jie Liu, Qiang Zhu
AI

Multilingual Embedding Probes Fail to Generalize Across Learner Corpora

arXiv:2604.07095v2 Announce Type: replace Abstract: Do multilingual embedding models encode a language-general representation of proficiency? We investigate this by training linear and non-linear probes on hidden-state…

Source: arXiv cs.CL Laurits Lyngbaek, Ross Deans Kristensen-McLachlan
AI

HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models

arXiv:2506.03922v4 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains. However, current benchmarks for evaluating MLLMs…

Source: arXiv cs.CL Zhaolu Kang, Junhao Gong, Jiaxu Yan, Wanke Xia, Yian Wang, Ziwen Wang, Huaxuan Ding, Zhuo Cheng, Wenhao Cao, Zhiyuan Feng, Siqi He, Shannan Yan, Junzhe Chen, Xiaomin He, Chaoya Jiang, Wei Ye, Kaidong…
AI

Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics

arXiv:2608.10678v1 Announce Type: new Abstract: Chinese web pollution has surfaced in LLMs, motivating audits of upstream Chinese corpora. However, auditing such corpora faces three challenges: (1) their web-scale size…

Source: arXiv cs.CL Qingjie Zhang, Ziqi Tang, Jie Zhang, Gelei Deng, Jinfeng Li, YueFeng Chen, Yitong Yang, Hui Xue, Tianwei Zhang, Han Qiu
AI

Beyond Screenshots: Evaluating VLMs' Understanding of UI Animations

arXiv:2604.26148v3 Announce Type: replace-cross Abstract: AI agents operating on user interfaces must understand how interfaces communicate state and feedback to act reliably. As a core communicative modality,…

Source: arXiv cs.CL Chen Liang, Xirui Jiang, Naihao Deng, Eytan Adar, Anhong Guo
AI

RadFusion: Towards Threshold-Controllable Radiology Report Generation

arXiv:2608.10505v1 Announce Type: cross Abstract: Automated radiology report generation is advancing rapidly in response to the shortage of radiologists, yet unlike a perception model, existing generation models offer…

Source: arXiv cs.CL Ying Jin, Noel C. F. Codella, John Corring, Mu Wei, Dinei Florencio, Eric Horvitz
AI

The Hidden Puppet Master: Predicting Human Belief Change in Manipulative LLM Dialogues

arXiv:2603.20907v5 Announce Type: replace Abstract: As users increasingly turn to LLMs for practical and personal advice, they become vulnerable to subtle steering toward hidden incentives misaligned with their own…

Source: arXiv cs.CL Jocelyn Shen, Amina Luvsanchultem, Jessica Kim, Kynnedy Smith, Valdemar Danry, Kantwon Rogers, Hae Won Park, Maarten Sap, Cynthia Breazeal
AI

Do LLMs Benefit From Their Own Words?

arXiv:2602.24287v2 Announce Type: replace Abstract: In multi-turn conversations, large language models typically condition on the full conversation history: both past user prompts and assistant responses. We revisit…

Source: arXiv cs.CL Jenny Y. Huang, Leshem Choshen, Wei Sun, Omar Khattab, Ram\'on Fernandez Astudillo, Mehul Damani, Tamara Broderick, Jacob Andreas
AI

Simulating Organized Group Behavior: New Framework, Benchmark, and Analysis

arXiv:2604.09874v2 Announce Type: replace Abstract: Simulating how organized groups (e.g., corporations) make decisions (e.g., responding to a competitor's move) is essential for understanding real-world dynamics and…

Source: arXiv cs.CL Xinkai Zou, Yiming Huang, Zhuohang Wu, Jian Sha, Nan Huang, Longfei Yun, Jingbo Shang, Letian Peng
AI

TRIBE: Predicting Team Performance via Communication Behavior Ensembles

arXiv:2608.06926v1 Announce Type: cross Abstract: Designing autonomous agents that effectively assist human teams hinges on understanding team dynamics, often without task specific knowledge. We present TRIBE, a domain…

Source: arXiv cs.CL Ali Jalal-Kamali, Nikolos Gurney, David V. Pynadath, Fred Morstatter
AI

Calibrating Post-Training Feature Shifts for LLM Data Contamination Detection

arXiv:2608.10462v1 Announce Type: new Abstract: Large language models (LLMs) are trained on massive and largely undisclosed corpora that may contain copyrighted or privacy-sensitive content. Data contamination detection…

Source: arXiv cs.CL Zhen Yang (The University of New South Wales), Mengqi Wang (The University of New South Wales), Gengda Zhao (The University of New South Wales), Mo Zhou (The University of New South Wales), Jianwei W…
AI

Evaluating Rational Contracting in Natural Language

arXiv:2608.10475v1 Announce Type: cross Abstract: The emergence of language-based AI agents promises to transform the scope of machine economic activity. Instead of just proposing bids or following hard-coded protocols,…

Source: arXiv cs.CL Bhavyesh Sajja, Max Kleiman-Weiner, Roger Zimmermann, Tan Zhi-Xuan
AI

ReLTEx: Reliable LLM-based Taxonomy Expansion

arXiv:2608.10970v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have demonstrated strong capabilities in generating semantically relevant concepts and relations, making them promising…

Source: arXiv cs.CL Zeinab Ghamlouch, Mehwish Alam
AI

No Single Best Model for Diversity: Learning a Router for Sample Diversity

arXiv:2604.02319v3 Announce Type: replace Abstract: When posed with prompts that permit a large number of valid answers, comprehensively generating them is the first step towards satisfying a wide range of users. In…

Source: arXiv cs.CL Yuhan Liu, Fangyuan Xu, Vishakh Padmakumar, Daphne Ippolito, Eunsol Choi
AI

Multiclass Sentiment Analysis for Identifying Political Viewpoints

arXiv:2608.11049v1 Announce Type: new Abstract: The rapid growth of social media has created vast amounts of political discourse, which provides valuable opportunities to analyze public opinions and identify different…

Source: arXiv cs.CL Girma Yohannis Bade, Olga Kolesnikova, Jose Luis Oropeza, Grigori Sidorov
AI

Authorship Verification of Transcribed German-Language Videos

arXiv:2607.29168v2 Announce Type: replace Abstract: Authorship Verification (AV) represents an important subfield of digital text forensics and addresses the fundamental question of whether two texts were written by the…

Source: arXiv cs.CL Oren Halvani, Sophie Titze
AI

When Reranking Hurts: Uncertainty-Based Gating for Few-Shot Reranking

arXiv:2606.31087v3 Announce Type: replace Abstract: Few-shot selection typically assumes that reranking retrieved examples always improves performance. We challenge this view by identifying that the expensive reranking…

Source: arXiv cs.CL Orian Dabod, Amir DN Cohen, Gabriel Stanovsky
AI

Assessing Reliability of BERT-Based Models on Question Answering Tasks

arXiv:2608.10806v1 Announce Type: new Abstract: Reliability estimation of large language models is in many cases as crucial as their accuracy, as reliable models are more trustworthy, robust, and suitable for practical…

Source: arXiv cs.CL Pooja Yadav, Priyanka Harjule, Basant Agarwal, Marko Robnik \v{S}ikonja
AI

SynBoost: A Synergistic Framework for Fast Sampling of Diffusion Models

arXiv:2506.13058v2 Announce Type: replace-cross Abstract: Diffusion probabilistic models (DPMs) have demonstrated remarkable success in visual generation. However, their iterative sampling mechanism results in slow…

Source: arXiv cs.AI Hu Yu, Hao Luo, Xueyang Fu, Jie Huang, Fan Wang, Feng Zhao
AI

Memory-Augmented Reinforcement Learning Agent for CAD Generation

arXiv:2605.19748v2 Announce Type: replace Abstract: Automatic generation of computer-aided design (CAD) models is a core technology for enabling intelligence in advanced manufacturing. Existing generation methods based…

Source: arXiv cs.AI Yin Xiaolong, Liu Yu, Shen Jiahang, Lu Xingyu, Ni Jingzhe, Fan Fengxiao, Sang Fan
AI

Blast Radius

arXiv:2608.07440v2 Announce Type: replace Abstract: Agentic coding faces growing problems of affordability and wasted tokens. We introduce Blast Radius, a predictive memory management layer that estimates an incoming…

Source: arXiv cs.AI MY Pitsane, Hope Mogale
AI

ComBodied Agents: a New Paradigm of Human-Centric Agentic AI

arXiv:2608.10915v1 Announce Type: new Abstract: After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the…

Source: arXiv cs.AI Qianggang Ding, Xingyao Wang, Rui Feng, Zhibin Wang, Feixiang Wang, Kelong Mao, Hao Sun, Zhiyao Luo, Jiankai Tang, Lei Li, Jiadong Guo, Minheng Ni, Weicong Lin, Chenxi Yang, Hongxiang Gao, Zhenghua C…
AI

RLMOpt: Adaptive Prompt Optimization via Recursive Language Models

arXiv:2608.10471v1 Announce Type: new Abstract: Prompt optimizers automate the search for prompts that improve language-model performance, but existing methods rely on a predefined optimization procedure: the algorithm…

Source: arXiv cs.AI Subhash Bangalore Satheesha, Nirvik Pande, Deepthi Duddempudi, Bharath Dandala
AI

Situation Graph Prediction for User Perspective Modeling

arXiv:2602.13319v2 Announce Type: replace Abstract: Perspective-aware AI requires modeling evolving internal states---goals, emotions, contexts---not merely preferences. Progress is limited by a data bottleneck: digital…

Source: arXiv cs.AI Jisung Shin, Daniel Platnick, Marjan Alirezaie, Hossein Rahnama
AI

REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems

arXiv:2608.10669v1 Announce Type: new Abstract: Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks. Adversarial inputs can exploit interactions between the…

Source: arXiv cs.AI Zixing Chen, Xingyuan Liu, Jie Zhu, Huaixia Dou, Shuo Jiang, Junhui Li, Lifan Guo, Feng Chen, Chi Zhang
AI

Compositional Benchmark Synthesis for Hierarchical Human Action Recognition

arXiv:2608.10765v1 Announce Type: new Abstract: Recognizing human behavior across levels of abstraction, from atomic actions to long-horizon intentions, requires data annotated along a semantic hierarchy. Large corpora…

Source: arXiv cs.AI Farnaz Soleimani (LISSI), Abdelghani Chibani (LISSI), Yacine Amirat (LISSI), Ghazaleh Khodabandelou (LISSI)
AI

RTSKG: Building a Rail Transit Station Knowledge Graph Dataset

arXiv:2608.11080v1 Announce Type: new Abstract: Rail transit systems play a vital role in urban mobility and economic development. As key components of such systems, rail transit stations function as critical transport…

Source: arXiv cs.AI Shutong Zhu, Tianxing Wu, Runfeng Liu, Yuang Gu, Xuan He, Yuan Zhu
AI

Rethinking Text-Based Image Retrieval in Specific Domain

arXiv:2608.10524v1 Announce Type: cross Abstract: Driven by the rapid advancement of vision-language representation learning, Text-based Image Retrieval (TBIR) has made notable progress. However, existing benchmarks are…

Source: arXiv cs.AI Jingyang Tan, Sheng Yang, Yuanpeng Chen, Jian Wang, Nianjin Ye, Chen Xing, Lanpeng Jia
AI

Self-Correcting Long-Horizon Search Agents via Tree-Structured Memory

arXiv:2608.10676v1 Announce Type: new Abstract: Large language model (LLM)-based search agents answer questions through multi-step interactions with external environments. However, providing complete execution…

Source: arXiv cs.AI Aijun Yang, Qianxue Guo, Ziyi Huang, Yuxuan Chen, Shiyou Qian, Jian Cao
AI

Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning

arXiv:2608.10403v1 Announce Type: new Abstract: Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due to insufficient…

Source: arXiv cs.AI Xincong Hu (Nanjing University), Lei Ou (Nanjing University), Maosen Li (Yinwang Intelligent Technology Co., Ltd), Jingtao Zhang (Yinwang Intelligent Technology Co., Ltd), Liguo Hou (Yinwang Intellig…
AI

FITTER: Vocabulary-Agnostic Cross-Domain Inference on Temporal Knowledge Graphs

arXiv:2608.10668v1 Announce Type: new Abstract: Temporal knowledge graphs are central to many uses of the Semantic Web, but existing completion methods assume the entities, relation names, and timestamps to be reasoned…

Source: arXiv cs.AI Jiaxin Pan, Mojtaba Nayyeri, Osama Mohammed, Daniel Hernandez, Rongchuan Zhang, Cheng Cheng, Steffen Staab
AI

Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training

arXiv:2608.10522v1 Announce Type: cross Abstract: While vision-language models dominate medical representation learning, unstructured text lacks the dense, quantitative diagnostic phenotypes inherent in structured…

Source: arXiv cs.AI Yingsheng Liu, Haiming Li, Jingmin Zhu, Jiajun Sun, Victoria Mar, Monika Janda, H. Peter Soyer, Zongyuan Ge, Zhen Yu
AI

DiffImaginE: Imagine to Verify Entity Types with Diffusio

arXiv:2608.03025v2 Announce Type: replace Abstract: Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence.…

Source: arXiv cs.AI Feng Zhang, Feiyu Han, Rongxin Yang, Yang Liu, Yancheng Chen, Rui Wang, Yingguang Yang, Tian Xueyun, Chongyang Zhang, Hao Zheng, Xu Kefu, Congjing Ran, Fuhai Chen, Bin Chong
AI

TongGuOCR: A Layout-Aware and Token-Augmented OCR MLLM for Chinese Historical Documents

arXiv:2608.07917v2 Announce Type: replace Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval,…

Source: arXiv cs.AI Zhongheng Zhou, Yi Sun, Huiguo He, Yuyi Zhang, Peirong Zhang, Yulin Fang, Dezhi Peng, Minghui Liao, Lianwen Jin
AI

JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures

arXiv:2602.17162v3 Announce Type: replace Abstract: Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature". While effective at…

Source: arXiv cs.AI Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Nicole Bussola, Simon Lee, Shane O'Connell, Dung Hoang, Maris…
AI

Hierarchical Compositionality for An Assistive AI Agent

arXiv:2608.10330v1 Announce Type: new Abstract: AI agents are increasingly being developed to assist humans in various applications, and Large Language Models and other deep network architectures are considered to be…

Source: arXiv cs.AI Tianyi Fu, Mohan Sridharan
AI

Improving TensorSketch Using Complex Random Variables

arXiv:2608.10523v1 Announce Type: cross Abstract: \texttt{TensorSketch} by~\cite{pham2013fast,kar2012random} provides efficient sketching algorithms for high-dimensional polynomial kernels $\vec{x}^{\otimes p} \in…

Source: arXiv cs.AI Amit Sharma, Mohammad Azhar Khan, Rameshwar Pratap, Keegan Kang
AI

Reinforcement Learning-Based Laser Cutting Machine Parameter Optimization

arXiv:2608.10549v1 Announce Type: new Abstract: Achieving high accuracy in laser-based cutting of optical films requires careful tuning of parameters such as focal length and laser power beam, adjusted according to the…

Source: arXiv cs.AI Khanh Quan Pham, Majid Kundroo, Geunwoo Ban, Seongho Bae, Taehong Kim
AI

MEGA: Self-Evolving Agent Optimization Infrastructure via Wisdom Graph

arXiv:2608.10504v1 Announce Type: new Abstract: As coding agents increasingly handle implementation, the central challenge shifts from building individual agents to building an infrastructure that systematically…

Source: arXiv cs.AI Jung Hwan Lee, Kyu Ho Lee, Gwang Hoon Yoo
AI

Inferential Capability Does Not Determine Legal Scope

arXiv:2608.10601v1 Announce Type: cross Abstract: Two instruments of EU digital law place inference at their centre and mean different things by it. Article 3(1) of the AI Act uses the capability to infer…

Source: arXiv cs.AI Nicola Fabiano
AI

Field-Localized Forgery Detection for Digital Identity Documents

arXiv:2605.09089v2 Announce Type: replace-cross Abstract: Digital onboarding and eKYC systems used by banks, fintech platforms, telecom providers, and other third-party services commonly verify users by comparing an…

Source: arXiv cs.AI Abhishek Kumar, Riya Tapwal, Carsten Maple, Mark Hooper
AI

Towards Unified Dynamic Face Landmark Detection

arXiv:2608.10346v1 Announce Type: cross Abstract: Although advancements in face landmark detection (FLD) methods continue to push performance boundaries, they overlook two major functional limitations: (1) different…

Source: arXiv cs.AI Sebastian Regalado, Varshanth R. Rao, Ruowei Jiang, Parham Aarabi, Igor Gilitschenski
AI

Agentic Instruction Data Selection: Let DataMaster Interpret Your Intent

arXiv:2608.10579v1 Announce Type: new Abstract: Although existing instruction data selection methods have introduced various metrics, the inherent complexity of real-world datasets makes it impractical for any single…

Source: arXiv cs.AI Fanqi Zhou, Qiaosheng Chen, Zixian Huang, Gong Cheng
AI

Closed-Loop LLM Co-Pilots for Digital Agriculture

arXiv:2608.09949v1 Announce Type: new Abstract: This study evaluates the application of Large Language Models (LLMs) in complex biological systems, evolving from data analysis to autonomous, AI-guided experimentation.…

Source: arXiv cs.AI Serge Kernbach
AI

Conversational Orchestration for Organic 6G

arXiv:2608.10714v1 Announce Type: cross Abstract: The Organic 6G vision of a network of networks spanning an edge-cloud continuum complemented by non-terrestrial resources requires, to realize its promise, service…

Source: arXiv cs.AI Masoud Shokrnezhad, Tarik Taleb
AI

DIMOS: Disentangling Instance-level Moving Object Segmentation

arXiv:2606.12826v2 Announce Type: replace-cross Abstract: Moving instance segmentation (MIS) attracts increasing attention due to its broad applications in traffic surveillance, autonomous driving, and animal tracking.…

Source: arXiv cs.AI Hongxiang Huang, Hongwei Ren, Xiaopeng Lin, Yulong Huang, Zeke Xie, Bojun Cheng
AI

Generating Attacks for LLMs with GFlowNets

arXiv:2608.10171v1 Announce Type: new Abstract: The rapid advancement of Large Language Models (LLMs) has facilitated their ubiquitous integration into various domains, leading to widespread adoption. However, this…

Source: arXiv cs.AI Berkay Ozcam, Irem Onen, Mehmet Fatih Amasyali, Emin Islam Tatli
AI

IO Factory: Simulating AI-Enabled Influence Campaigns at Scale

arXiv:2608.10920v1 Announce Type: new Abstract: We introduce IO Factory, an AI-driven framework for simulating information and influence campaigns as fully integrated, traceable processes. The threat of digital…

Source: arXiv cs.AI Lukasz Olejnik, Wenchao Dong, Jonas R. Kunst, Signe Riemer-S{\o}rensen, Tobias Herb, Meeyoung Cha, Daniel Thilo Schroeder
AI

Rationale-Guided Learning for Multimodal Emotion Recognition

arXiv:2608.10448v1 Announce Type: new Abstract: Multimodal emotion recognition in conversation (MERC) requires understanding complex interactions between verbal and non-verbal cues. However, most existing approaches…

Source: arXiv cs.AI Sujung Oh, Jung Uk Kim, Sangmin Lee
AI

A Single Atom in Front of a Mirror is a Universal Reservoir Computer

arXiv:2608.10382v1 Announce Type: cross Abstract: Universal approximation in reservoir computing is typically associated with a class of reservoirs. We show that universality can be associated with a single reservoir,…

Source: arXiv cs.AI Peter J. Ehlers, Phi Hung Nguyen, Kanu Sinha, Noelle Daigle, Travis W. Sawyer, Hendra I. Nurdin, Daniel Soh
AI

Entropy-Centric Explainable AI for Remote Sensing Image Segmentation

arXiv:2608.11064v1 Announce Type: cross Abstract: Artificial intelligence (AI) has become a powerful approach to solving complex problems in critical domains. Many concerns arise regarding the decision-making process of…

Source: arXiv cs.AI Ali Saleh, Abdul Karim Gizzini, Mohamad Ghassany, Ali J. Ghandour
AI

Persistent Recursive Worlds Enable Autonomous Software Evolution

arXiv:2608.10450v1 Announce Type: cross Abstract: Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems preserve continuity through…

Source: arXiv cs.AI Beichen Huang, Zhenyu Liang, Bowen Zheng, Ran Cheng
AI

sLTN: Structural Logic Tensor Networks

arXiv:2608.11136v1 Announce Type: new Abstract: Logic Tensor Networks (LTN) provide a neurosymbolic framework in which first-order logic is interpreted through tensor operations, enabling logical constraints to be…

Source: arXiv cs.AI Davide Rinaldi, Luciano Serafini
AI

From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection

arXiv:2605.27944v2 Announce Type: replace Abstract: With rapid advances in audio-visual generative models, reliable forgery detection becomes increasingly critical. Existing methods for audio-visual deepfake detection…

Source: arXiv cs.AI Ke Liu, Jiwei Wei, Wenyu Zhang, Shuchang Zhou, Ruikun Chai, Yutao Dai, Chaoning Zhang, Yang Yang
AI

X2C: A Large-Scale Benchmark for Nuanced Humanoid Facial Expression Imitation

arXiv:2505.11146v4 Announce Type: replace-cross Abstract: Fine-grained facial expression transfer from humans to humanoid agents presents a unique pattern recognition challenge due to the significant domain gap between…

Source: arXiv cs.AI Peizhen Li, Longbing Cao, Xiao-Ming Wu, Runze Yang, Xiaohan Yu
AI

Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data

arXiv:2608.11022v1 Announce Type: cross Abstract: Model Cards and Data Cards have demonstrated the value of structured, human-readable documentation for machine learning artifacts, capturing their context, parameters,…

Source: arXiv cs.AI Nicola Giuseppe Marchioro, Gabriele Padovani, Amal Gueroudji, Rafael Ferreira da Silva, Wesley Brewer, Valentine Anantharaj, Sandro Fiore, Renan Souza
AI

Toward a Theory of Value in AI Alignment

arXiv:2608.10327v1 Announce Type: new Abstract: Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a rise in harms spanning from…

Source: arXiv cs.AI Andrew Smart, Shazeda Ahmed, Jackie Kay, Jimmy Tobin, Kris Shrishak, Abeba Birhane
AI

Hybrid Token Compression for Vision-Language Models

arXiv:2512.08240v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) rely on hundreds of visual tokens, leading to high computational and memory costs. Existing compression methods face a trade-off:…

Source: arXiv cs.AI Jusheng Zhang, Xiaoyang Guo, Tongyu Mo, Qinhan Lv, Wenhao Chai, Jian Wang, Keze Wang, Liang Lin
AI

Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning

arXiv:2608.11204v1 Announce Type: cross Abstract: Learning reliable surgical manipulation policies is bottlenecked by the scarcity of action-labeled demonstrations: teleoperated surgical robot (e.g., dVRK) trajectories…

Source: arXiv cs.AI Wenrui Bao, Tianyun Jiang, Zhiben Chen, Ser-Nam Lim, Peter D. Peng, Yuzhang Shang
AI

Expert-Guided g-computation with Large Language Models for Estimating Causal Effects on Timings: Applications to Hospital Quality Improvement

arXiv:2608.10339v1 Announce Type: cross Abstract: Hospital quality improvement (QI) programs routinely face multiple candidate interventions to optimize hospital flow, but existing methods struggle to estimate and rank…

Source: arXiv cs.AI Patrick Vossler, Jialin Ouyang, F. Richard Guo, Anran Huang, Ali Shojaie, Lucas Zier, Fan Xia, Jean Feng
AI

Automating and Scaling Behavioral Scientific Research on AI Agents

arXiv:2608.10030v1 Announce Type: new Abstract: As AI agents are increasingly deployed in complex environments, understanding their behaviors becomes critical. Yet behavioral scientific research on AI agents remains…

Source: arXiv cs.AI Soo Yong Lee, Jongha Lee, Jaewan Chun, Hyunjin Hwang, Fanchen Bu, Ziv Ben-Zion, Taekwan Kim, Denny Borsboom, Jaemin Yoo, Kijung Shin
AI

Optimal Stopping of Self-Refining Foundation Models

arXiv:2608.10729v1 Announce Type: cross Abstract: Foundation models can improve their outputs through a self-refinement process driven by external feedback. In this process, the model is embedded in an iterative loop…

Source: arXiv cs.AI Kim Hammar, Tansu Alpcan, Emil C. Lupu
AI

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning

arXiv:2608.10513v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language…

Source: arXiv cs.AI Caoyuan Ma, Wenpu Liu, Weichu Xie, Tian Gu, Shilei Zhao, Lingxi Min, Shuai Dong, Yuqi Xu, Ji Zhao, Ziyue Wang, Wenzheng Chang, Taiqiang Wu, Yongfu Zhu, Wenqi Shao, Yinqiang Zheng
AI

GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning

arXiv:2608.10494v1 Announce Type: new Abstract: Earth observation (EO) agents construct scientifically valid tool workflows and ground their conclusions in current geospatial evidence. This is challenging because EO…

Source: arXiv cs.AI Xin Xiao, Jiang Zhong, Junnan Zhu, Yingchao Feng, Peijin Wang, Yidan Zhang, Kaiwen Wei
AI

HoosierHelp: Benchmarking LLM Agents for Social Service Navigation

arXiv:2608.09946v1 Announce Type: cross Abstract: Social service navigation requires connecting help-seeking individuals to resources that satisfy their needs and specific constraints. Although LLM agents offer a…

Source: arXiv cs.AI Yiyang Li, Weixiang Sun, Tianyi Ma, Kaiwen Shi, Zheyuan Zhang, Yanfang Ye
AI

Frozen Brain-MRI Foundation Models Are Site Fingerprints

arXiv:2608.10295v1 Announce Type: cross Abstract: Frozen foundation-model (FM) embeddings are increasingly used as off-the-shelf brain-MRI representations, on the assumption that they capture anatomy. We audit what they…

Source: arXiv cs.AI Saman Rahbar
AI

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution

arXiv:2608.08311v2 Announce Type: replace-cross Abstract: We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that…

Source: arXiv cs.AI Anton Razzhigaev, Andrei Gritsaev, Andrei Kaznacheev, Nikita Dragunov, Roman Yampolskiy, Andrei Kuznetsov
AI

ImpactHO: Importance-Aware KV Cache Transfer for Multi-User Edge LLM Handover

arXiv:2608.10545v1 Announce Type: cross Abstract: Edge LLMs must preserve inference continuity when a user hands over between edge nodes, requiring key-value (KV) cache transfer to the target node. However, simultaneous…

Source: arXiv cs.AI Minwoo Kim, Soochang Song, Namyoon Lee, Bang Chul Jung, Yongjune Kim
AI

AttriMem: Attribution-Guided Process Feedback for Agent Memory Construction

arXiv:2607.21106v3 Announce Type: replace Abstract: Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what information to extract,…

Source: arXiv cs.AI Qinfeng Li, Yuntai Bao, Xinyan Yu, Hongze Chen, Yanming Liu, Huifeng Zhu, Yier Jin, Jintao Chen, Wenqi Zhang, Xuhong Zhang
AI

LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning

arXiv:2601.10129v2 Announce Type: replace-cross Abstract: Current multimodal latent reasoning often relies on external supervision (e.g., auxiliary images), ignoring intrinsic visual attention dynamics. In this work, we…

Source: arXiv cs.AI Linquan Wu, Tianxiang Jiang, Yifei Dong, Haoyu Yang, Fengji Zhang, Shichaang Meng, Ai Xuan, Linqi Song, Jacky Keung
AI

Graphical Models of False Information and Fact Checking Ecosystems

arXiv:2208.11582v2 Announce Type: replace-cross Abstract: The wide spread of false information online, including misinformation and disinformation, has become a major problem for our highly digitised and globalised…

Source: arXiv cs.AI Haiyue Yuan, Enes Altuncu, Shujun Li, Can Baskent, Jason R. C. Nurse
AI

ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling

arXiv:2608.10928v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) improve performance by allocating additional inference-time compute to generate extended chain-of-thought reasoning. However, recent studies…

Source: arXiv cs.AI Vaibhav Singh, Soumya Suvra Ghosal, Sarvesh Gharat, Soumyabrata Pal, Ramasuri Narayanam, Dinesh Manocha
AI

Pricing Access to Dynamic Information Services

arXiv:2510.09859v5 Announce Type: replace-cross Abstract: A provider sells a \emph{dynamic information service}---a real-time, capacity-constrained process that resolves a customer's uncertainty---to customers who…

Source: arXiv cs.AI Weijie Zhong
AI

Interpreting Language Model Hidden States at Scale

arXiv:2608.10260v1 Announce Type: new Abstract: Lens methods interpret large language models (LLMs) by mapping intermediate activations to the output vocabulary, revealing how next-token predictions develop through the…

Source: arXiv cs.AI Jordan Pettyjohn, Mansi Sakarvadia, Nathaniel Hudson, Daniel McKenzie, Kyle Chard, Ian Foster
AI

Self-evolving Agentic Customer Support System at LinkedIn

arXiv:2608.10224v1 Announce Type: new Abstract: Enterprise support agents operate in rapidly changing environments where policies, product capabilities, and knowledge bases evolve continuously, making static assistants…

Source: arXiv cs.AI Chih Hui Wang, Mengdie Tu, Qianyun Zhang, Wei Wu, Lili Zhou, Mingqi Shen, Changshuai Wei
AI

EvoMem: Memory-Augmented Evolution for Code Optimization

arXiv:2608.10795v1 Announce Type: new Abstract: Successful mutation strategies in evolutionary code search may contain reusable knowledge that is useful beyond a single run, and in some cases may transfer across related…

Source: arXiv cs.AI Viktor Volkov, Valentin Khrulkov, Andrey V. Galichin, Danil Sivtsov, Nikita Glazkov, Olga Volkova, Konstantin Pchelin, Iaroslav Bespalov, Dmitry V. Dylov, Petr Anokhin, Ivan Oseledets
AI

Causality Sum Rules in Conventional Scattering Matrices

arXiv:2608.10427v1 Announce Type: cross Abstract: Scattering matrices are the standard experimental and computational description of photonic and electromagnetic devices. Passivity is explicit in the conventional…

Source: arXiv cs.AI Ning Han, Rui Zhao, Shuxing Yang, Mingzhu Li, Hongsheng Chen, Yihao Yang
AI

FedCGR: Federated Cross-Domain Generative Recommendation

arXiv:2608.10929v1 Announce Type: new Abstract: Cross-domain recommendation (CDR) transfers preference knowledge across related domains, but federated deployment makes cross-domain alignment difficult because the…

Source: arXiv cs.AI Zhuodong Liu, Hugen Lv, Xiangyu Li, Bohan Guo, Peiyu Hu
AI

Thought-Level Beam Search for Reasoning

arXiv:2608.08020v2 Announce Type: replace Abstract: Test-time compute scaling is a primary driver of performance in large reasoning models (LRMs), but extreme inefficiency bounds current approaches, shifting the…

Source: arXiv cs.AI Lijie Yang, Hongyin Luo, Jiawei Zhao, Tri Dao, Ravi Netravali
AI

Small Foundation Models of Human Cognition and Behaviour

arXiv:2608.05224v3 Announce Type: replace Abstract: Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models…

Source: arXiv cs.AI Nick Oh, Fernand Gobet
AI

Patients With Personality: Realistic Patient Simulation through Controlled Diversity and Selective Disclosure

arXiv:2606.17441v2 Announce Type: replace-cross Abstract: Simulating realistic patient interactions is a key requirement to testing clinical applications of LLMs at scale without time-consuming and expensive user…

Source: arXiv cs.AI Moritz Schlager, Friederike Jungmann, Samuel Schmidgall, Philipp Raffler, Franziska Hartl, Eva Wende, Paula Ro{\ss}m\"uller, Conrad Ketzer, Avinatan Hassidim, Dale R. Webster, Yossi Matias, Yun Liu,…
AI

CuSearch: Curriculum Rollout Sampling via Search Depth for Agentic RAG

arXiv:2605.11611v3 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for training agentic retrieval-augmented generation (RAG) systems from…

Source: arXiv cs.AI Jianghan Shen, Siqi Luo, Xinyu Cheng, Jing Xiong, Yue Li, Jiyao Liu, Jiashi Lin, Yirong Chen, Junjun He
AI

MESA:Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory

arXiv:2608.10108v1 Announce Type: new Abstract: Long-horizon agents accumulate trajectories spanning hundreds of interleaved reasoning, action, and observation steps, where answering a query may depend on evidence…

Source: arXiv cs.AI Beidi Zhao, Yaoqi Chen, Yuru Feng, Menghao Li, Qianxi Zhang, Baotong Lu, Jianan Lu, Zhirui Wang, Xinjiang Wang, Shusen Xu, Zengzhong Li, Xiaoxiao Li, Qi Chen
AI

The Epistemic Politics of AI Anthropomorphism

arXiv:2608.00961v2 Announce Type: replace-cross Abstract: AI anthropomorphism is typically treated as a problem of user misperception requiring institutional correction. Users who engage in sustained or relational…

Source: arXiv cs.AI Donna M Bye, Levin Kuhlmann
AI

SDDBMs: Soft Denoising Diffusion Bridge Models

arXiv:2608.08594v2 Announce Type: replace Abstract: Diffusion bridge models leverage Doob's \(h\)-transform to construct stochastic transports between arbitrary endpoint distributions, and have shown strong potential in…

Source: arXiv cs.AI Shiyi Qi, Kun He, Mingmou Liu
AI

Entropy-based Code Adversarial Translation for Real-world Repository Migration

arXiv:2608.09273v2 Announce Type: replace Abstract: LLMs have demonstrated strong capabilities in code generation and automated program repair, but migrating an entire repository rarely produces a runnable application…

Source: arXiv cs.AI Yushun Tang, Yisen Cao, Zhicheng Chen, Lin Peng, Junkang Mao, Fengyi Song, Yantao Jia
AI

MaskFlow: Precise, Consistent and Seamless Regional Image Editing

arXiv:2608.06929v2 Announce Type: replace-cross Abstract: Regional image editing has attracted considerable attention for its spatial controllability. Although instruction-based and mask-reference-based editing methods…

Source: arXiv cs.AI Rui Xu, Yang Yong, Shunzi Yang, Ruihao Gong, Chengtao Lv
AI

Exploring Semantic Stability Across Reviews in the Linux Kernel

arXiv:2608.10101v1 Announce Type: cross Abstract: Code review is credited with substantially changing a patch's code between its first submission and the version that eventually lands. However, prior work typically…

Source: arXiv cs.AI Lucas Ciziks, Paulo Meirelles, Marco Aur\'elio Gerosa
AI

Closing a 17-Year Gap: Algorithmic Detection and Empirical Prevalence of Rank Reversal in Multi-Criteria Decision Analysis

arXiv:2508.00129v2 Announce Type: replace Abstract: Rank Reversal, where the relative order of alternatives changes in ways that violate axioms of rational decision-making, is a well-documented threat to the reliability…

Source: arXiv cs.AI Juan Bautista Cabral, Gonzalo Giarda, Diego Nicol\'as Gimenez Irusta, Paula Pacheco, Alvaro Roy Schachner, Agust\'in Borda
AI

Eleven Years of BRACIS: A Meta-Scientific Study of the Brazilian Conference on Intelligent Systems

arXiv:2608.09964v1 Announce Type: cross Abstract: The Brazilian Conference on Intelligent Systems (BRACIS) is the main national venue for Artificial Intelligence research in Brazil, hosted by the Brazilian Computer…

Source: arXiv cs.AI Thales Sales Almeida, Giovana Kerche Bon\'as, Thiago Laitz, Jo\~ao Guilherme Alves Santos, Hugo Abonizio, Roseval Malaquias Junior, Marcos Piau, Celio Larcher, Ramon Pires, Rodrigo Nogueira
AI

What We Know about Responsible AI Practices in Industry: A Half Decade of Empirical Research

arXiv:2608.10431v1 Announce Type: cross Abstract: Responsible AI (RAI) has become a central concern for technology companies, regulators, and the public. How industry practitioners interpret, implement, and sustain RAI…

Source: arXiv cs.AI Wesley Hanwen Deng, Agathe Balayn, Andrew Selbst, Jason I. Hong, Motahhare Eslami, Kenneth Holstein, Hanna Wallach, Jennifer Wortman Vaughan, Solon Barocas
AI

CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification

arXiv:2604.01687v3 Announce Type: replace Abstract: Anthropic proposes the concept of skills for LLM agents to tackle multi-step professional tasks that simple tool invocations cannot address. A tool is a single,…

Source: arXiv cs.AI Hanrong Zhang (Steve), Shicheng Fan (Steve), Henry Peng Zou (Steve), Yankai Chen (Steve), Zhenting Wang (Steve), Jiayu Zhou (Steve), Chengze Li (Steve), Wei-Chieh Huang (Steve), Yifei Yao (Steve), Ke…
AI

A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa

arXiv:2608.11053v1 Announce Type: cross Abstract: The application of computer vision in agriculture has shown significant potential for improving crop monitoring and precision farming. However, many existing approaches…

Source: arXiv cs.AI Ismail Ismail Tijjani, Sunusi Muhammad Ibrahim, Amina Ibrahim Khaleel, Lanre Olusegun Akinola, Fatima Isa Jibrin, Muhammad Bashir Aliyu, Abdullahi Abdussalam Dalhat, Abdullahi Suiudeen
AI

TRACE: Trustworthy Retrieval-Augmented Conversational Engine

arXiv:2608.10176v1 Announce Type: new Abstract: Public service chatbots are expected to deliver recommendations from an underlying public service directory, while also making sure that the recommendations respect…

Source: arXiv cs.AI Touseef Hasan, Laila Cure, Souvika Sarkar
AI

Comprendia: AI-Augmented Code Comprehension

arXiv:2608.10290v1 Announce Type: cross Abstract: Comprendia is an Eclipse plugin that integrates structural dependency visualization with LLM-powered code explanation on a shared interactive graph for Java program…

Source: arXiv cs.AI Costain Nachuma, Minhaz F. Zibran
AI

MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models

arXiv:2608.10635v1 Announce Type: cross Abstract: Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain challenging. Existing…

Source: arXiv cs.AI Yuan Wang, Hualiang Wang, Yixin Chen, Songtao Jiang, Shujian Gao, Jiaming Lin, Siming Fu, Jian Wu, Zuozhu Liu
AI

CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation

arXiv:2608.09790v2 Announce Type: replace Abstract: Online credit card discussions provide a natural setting for studying how consumers communicate about financial products. Simulating these discussions requires more…

Source: arXiv cs.AI Yaoning Yu, Kai-Min Chang, Ye Yu, Yi-Chia Wang, Haojing Luo, Haohan Wang
AI

MAP-Graph: Provenance-Aware Shared Memory for Multi-Agent Workflows

arXiv:2608.10509v1 Announce Type: new Abstract: Shared memory helps language-model agents reuse information across long workflows, yet relevant evidence may not be admissible for a particular agent or action. Because…

Source: arXiv cs.AI Yiqi Wang, Zihao Yan, Jiaqi Zhang, Zhangkai Wu, Mingkai Zheng, Zequn Sun, Yanming Zhu, Taotao Cai
AI

Pretrained Optimization Model for Zero-Shot Black Box Optimization

arXiv:2405.03728v3 Announce Type: replace-cross Abstract: Zero-shot optimization involves optimizing a target task that was not seen during training, aiming to provide the optimal solution without or with minimal…

Source: arXiv cs.AI Xiaobin Li, Kai Wu, Yujian Betterest Li, Xiaoyu Zhang, Handing Wang, Jing Liu
AI

CARE: Confidence-Aware Reasoning for Reliable Medical VQA

arXiv:2608.10964v1 Announce Type: cross Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual question answering,…

Source: arXiv cs.AI Yuetian Du, Yucheng Wang, Zhenyuan Chen, Luyuan Chen, Rongyu Zhang, Jinjian Zhang, Wei Zhou, Zhijie Xu, Ming Kong, Zhan Zhou, Jie Liu, Qiang Zhu
AI

Token-Based Detection of Spurious Correlations in Vision Transformers

arXiv:2509.04009v2 Announce Type: replace-cross Abstract: Due to their powerful feature association capabilities, neural network-based computer vision models have the ability to detect and exploit unintended patterns…

Source: arXiv cs.AI Solha Kang, Esla Timothy Anzaku, Wesley De Neve, Arnout Van Messem, Joris Vankerschaver, Francois Rameau, Utku Ozbulak
AI

FUSE: Frame-Unified Stress Estimation from Facial Video

arXiv:2608.10442v1 Announce Type: cross Abstract: Automatic stress detection from facial video offers a practical path to non-intrusive affect monitoring, yet existing video-based approaches commonly decompose full…

Source: arXiv cs.AI Stefanos Gkikas, Thomas Kassiotis, Yang Guo, Guangliang Li, Giorgos Giannakakis
AI

Modelling Geographic Atrophy Progression using Implicit Neural Representations

arXiv:2608.10807v1 Announce Type: cross Abstract: Age-related Macular Degeneration (AMD) is the major cause of blindness in the Western world. Its late dry phase is characterised by irreversible atrophic areas, namely…

Source: arXiv cs.AI Simone Sarrocco, Paul Friedrich, Florentin Bieder, Christina Bornberg, Philippe Valmaggia, Peter Maloca, Philippe Cattin
AI

Contextual Information Policy Optimization for Search Agents

arXiv:2608.06128v3 Announce Type: replace Abstract: Search agents extend large language models beyond static parametric memory by enabling them to acquire and use external evidence during multi-step reasoning. For…

Source: arXiv cs.AI Xingyu Guo, Wei Chen, Linlin Yang, Baochang Zhang
AI

MIRA: Medical Image Reflection for Agentic Diagnosis

arXiv:2608.10827v1 Announce Type: cross Abstract: Medical visual agents can use tools to inspect images and retrieve external knowledge, but indiscriminate tool use may introduce noisy or misleading evidence. Reliable…

Source: arXiv cs.AI Shengzhi Wang, Jun Yang, Kai Wu, Xiaozhong Ji, Yiwen Ye, Ziyang Chen, Mingliang Xiong, Wen Fang, Mingqing Liu, Mengyuan Xu, Miaoxuan Shan, Caiyan Liu, Bin He, Qingwen Liu
AI

Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware Reformulation

arXiv:2608.10805v1 Announce Type: cross Abstract: Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in replacement for standard convolutions, expanding a network's receptive field exponentially…

Source: arXiv cs.AI Amit Aflalo, Shahaf E. Finder, Roy Amoyal, Eran Treister, Oren Freifeld
AI

GitSkills: A Dataset of Agent Skills on GitHub

arXiv:2608.10906v1 Announce Type: cross Abstract: An agent skill is a folder containing a SKILL.md file with instructions for a language-model agent, optionally accompanied by scripts and reference files. The agent…

Source: arXiv cs.AI Giuseppe Destefanis, Daniel Graziotin, Matteo Vaccargiu, Marco Ortu
AI

IndexTTS 2.5 Technical Report

arXiv:2601.03888v5 Announce Type: replace-cross Abstract: In prior work, we introduced IndexTTS 2, a zero-shot neural text-to-speech foundation model comprising two core components: a transformer-based Text-to-Semantic…

Source: arXiv cs.AI Yunpei Li, Xun Zhou, Jinchao Wang, Lu Wang, Yong Wu, Siyi Zhou, Yiquan Zhou, Yining Wang, Yaogen Yang, Zhetao Hu, Shiyao Duan, Jiacheng Xu, Jingchen Shu, Bin Xia
AI

Towards Efficient Reasoning in LLM-Based Recommender Systems via Model Merging

arXiv:2608.10447v1 Announce Type: cross Abstract: Large language model-based recommender systems are increasingly adopting slow-thinking models that generate step-by-step reasoning before making predictions, often…

Source: arXiv cs.AI Linh Dieu Le, Tong Chen, Shazia Sadiq, Hongzhi Yin, Ming Jin, Junliang Yu