Skip to content
TILens What matters today in tech
Theme

Daily edition · AI

The daily ledger

TILens turns technical updates into a focused daily brief: official releases, trusted reporting, and practitioner analysis, deduplicated and organized by topic.

18 Aug 2026 edition
AI Labs AI

ChatGPT Ads expands across Europe

ChatGPT Ads is expanding to 31 European markets. Learn how advertisers can reach people as they explore, compare options, and make decisions.

Source: OpenAI News
Python AI Django

Mojo🔥 is now open source

Mojo🔥 is now open source Mojo🔥 is now open source The Mojo programming language has been promising an open source release since May 2023. Last week they shipped their 1.0 and today they have followed through on that…

Source: Simon Willison's Weblog
AI Tools GitHub AI

langchain-ai/langchain: langchain-openai==1.5.2

Changes since langchain-openai==1.5.1 release(openai): 1.5.2 (#39719) fix(openai): preserve reasoning item boundaries (#39278) release(openai): 1.5.2a1 (#39709) feat(openai): extract gateway metadata from response…

Source: LangChain Releases github-actions[bot]
AI Tools GitHub AI

langchain-ai/langchain: langchain-openai==1.5.2a1

Initial release release(openai): 1.5.2a1 (#39709) feat(openai): extract gateway metadata from response headers when available (#39706) chore(openai): update snapshots (#39657) fix(openai): support o-series models in…

Source: LangChain Releases github-actions[bot]
AI

We still don’t know how people are really using AI

AI companies like Anthropic and OpenAI regularly publish reports on how people are using products like Claude and ChatGPT, but they only release the data they want us to see, AI researchers say. “There is no independent…

Source: MIT Technology Review AI Eileen Guo
AI Research AI

Advancing Open and Reproducible Relational Learning: RelArena-$\alpha$, TabPFN-Rel and RPI

arXiv:2608.16319v1 Announce Type: new Abstract: This first release of Prior Labs in relational learning shows our continued commitment to open science. We open-source three pieces of software that we expect to…

Source: arXiv cs.LG Adrian Hayler, Klemens Fl\"oge, Alan Arazi, Rishabh Ranjan, Jure Leskovec, Felix Birkel, Brendan Roof, Anurag Garg, Kristina Collins, Lydia Sidhoum, Jonas K\"ubler, Siyuan Guo, Oscar Key, Jan Hendrik…
AI Research AI

Koopman early warning signals for bifurcation and rate-induced tipping

arXiv:2608.14716v1 Announce Type: cross Abstract: Abrupt transitions in complex systems are often preceded by early warning signals. However, most indicators rely on the notion of critical slowing down and do not…

Source: arXiv cs.LG Juan Nathaniel, Carla Roesch, Derek DeSantis, Parvathi Kooloth, Hang Fan, Valerio Lucarini, Anastasia Romanou, Pierre Gentine
AI Research AI

Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search

arXiv:2608.15669v1 Announce Type: new Abstract: Scientific discovery often involves optimising expensive-to-evaluate objectives over vast, structured, and open-ended hypothesis spaces, such as molecules, protein…

Source: arXiv cs.LG Zhongwei Yu, Yan Song, Xue Yan, Anjie Liu, Xingyu Lu, Yihang Chen, Huichi Zhou, Siyuan Guo, Luoyang Sun, Sihan Chen, Xiangning Yu, Jun Wang
AI Research AI

Doubly robust nearest neighbors in factor models

arXiv:2211.14297v4 Announce Type: replace-cross Abstract: We introduce and analyze an improved variant of nearest neighbors (NN) for estimation with missing data in latent factor models. We consider a matrix completion…

Source: arXiv cs.LG Raaz Dwivedi, Sabina Tomkins, Predrag Klasnja, Susan Murphy, Devavrat Shah
AI Research AI

Conditional Evaluation of Language Models with Cheap Auxiliary Signals

arXiv:2608.16210v1 Announce Type: new Abstract: Aggregate accuracy hides where models succeed and fail. Estimating conditional performance profiles from gold labels alone is expensive, while cheap auxiliary signals such…

Source: arXiv cs.LG Zhi Zhang, Lingfeng Lyu, Yue Kang, Doudou Zhou
AI Research AI

Transfer Learning of Keystroke Dynamics for Cross-Device User Authentication

arXiv:2608.16334v1 Announce Type: new Abstract: Keystroke dynamics (typing patterns) can be used as a behavioural biometric modality for user authentication, with applications such as fraud prevention. While the…

Source: arXiv cs.LG Nuwan Kaluarachchi, Sevvandi Kandanaarachchi, Kristen Moore, Arathi Arakala, Conrad Sanderson
AI Research AI

LLM-based Framework for Generating and Verifying Parallel DEVS Statecharts

arXiv:2608.14956v1 Announce Type: new Abstract: The development of models demands sound modeling and simulation knowledge as well as domain knowledge. Every model should accurately represent a system's dynamics and be…

Source: arXiv cs.LG Vamsi Krishna Vasa, Hessam S. Sarjoughian, Edward J. Yellig
AI Research AI

CrevasseSeg: A Label-Efficient UAV Crevasse Segmentation Framework

arXiv:2608.15790v1 Announce Type: new Abstract: Crevasse mapping from uncrewed aerial vehicle (UAV) imagery matters for glaciological research and for field safety in glaciated terrain. Yet, pixel-level annotation of…

Source: arXiv cs.LG Steven Wallace, William D Harcourt, Richard Hann, Aiden Durrant, Somayajulu Sripada, Georgios Leontidis
AI Research AI

NPU Offloading of a Frozen Visual Encoder for Robot Policy Training

arXiv:2608.15002v1 Announce Type: cross Abstract: When a robot policy is trained for a new task or dataset, its visual encoder can be frozen and only its action generation module trained, reducing training cost.…

Source: arXiv cs.LG Hyojun Yun, Seungjae Won, Hyungpil Moon
AI Research AI

One Score, Two Decisions: Selective Prediction on the Rare-Disease Tail

arXiv:2608.14683v1 Announce Type: new Abstract: Given a patient's clinical findings, a diagnostic system ranks possible diseases and must decide when to endorse its first prediction or defer it for review. This decision…

Source: arXiv cs.LG Zhaoyang Jiang, Zhizhong Fu, Yunsoo Kim, Zicheng Li, Xuanqi Peng, Fei Teng, Jiacong Mi, Honghan Wu
AI Research AI

Dynamic Pricing and Advertising with Demand Learning

arXiv:2304.14385v4 Announce Type: replace-cross Abstract: We consider a novel pricing and advertising framework in which a seller not only sets the product price but also designs flexible advertising schemes to…

Source: arXiv cs.LG Shipra Agrawal, Yiding Feng, Wei Tang
AI Research AI

Look Before You Lift: Visual and Quantitative Diagnostics for Topological Deep Learning

arXiv:2608.15388v1 Announce Type: new Abstract: Topological deep learning (TDL) methods rely on lifting raw data into higher-order discrete domains such as simplicial complexes, cell complexes, and hypergraphs. In…

Source: arXiv cs.LG Mathilde Papillon, Guillermo Bern\'ardez, \'Alvaro Ball\'on Barreiro, Marco Montagna, R\'emi Devaux, Antoine Jardin, Nina Miolane
AI Research AI

FDA-Opt: Federated Fine-Tuning via Dynamic Update Schedules

arXiv:2505.04535v4 Announce Type: replace Abstract: Federated Learning (FL) enables the utilization of vast, previously inaccessible data sources. At the same time, pre-trained Language Models (LMs) have taken the world…

Source: arXiv cs.LG Michael Theologitis, Vasilis Samoladas, Antonios Deligiannakis
AI Research AI

Time-Correlated Video Bridge Matching

arXiv:2510.12453v3 Announce Type: replace Abstract: Diffusion models excel in noise-to-data generation tasks, providing a mapping from a Gaussian distribution to a more complex data distribution. However, they struggle…

Source: arXiv cs.LG Viacheslav Vasilev, Arseny Ivanov, Nikita Gushchin, Maria Kovaleva, Alexander Korotin
AI Research AI

Knowing When to Defer: Selective Prediction for Responsible Knowledge Tracing

arXiv:2509.21514v4 Announce Type: replace Abstract: Research on Knowledge Tracing (KT) models traditionally focuses on improving predictive accuracy. However, responsible real-world deployment requires models to know…

Source: arXiv cs.LG Joshua Mitton, Prarthana Bhattacharyya, Ralph Abboud, Simon Woodhead
AI Research AI

Spectral Rank Certification for Foundation Model Adapters

arXiv:2608.15351v1 Announce Type: new Abstract: Nominal LoRA rank is a design parameter; calibrated spectral evidence is a separate inferential quantity. This article develops a finite-sample framework for inferring…

Source: arXiv cs.LG Mohammed Ahnouch, Lotfi Elaachak
AI Research AI

Solvable Sokoban Without a Solver via Diffusion

arXiv:2608.15958v1 Announce Type: cross Abstract: Deciding whether a Sokoban puzzle is solvable is PSPACE-complete (Culberson, 1997): solutions can be exponentially long and there is no short certificate to check.…

Source: arXiv cs.LG Sina Baghal
AI Research AI

Reference-free logged energy-oracle recovery for neural approximations of symmetric coercive variational problems: conforming Riesz reconstruction and archive-level selection

arXiv:2608.16473v1 Announce Type: new Abstract: Neural PDE training yields a finite checkpoint archive, yet its logged energy errors are inaccessible without the exact solution, while loss-based selection does not…

Source: arXiv cs.LG Karim Bounja, Lahcen Laayouni, Boujemaa Achchab, Abdeljalil Sakat
AI Research AI

UniTAC: Universal Task-Aware Compression via Weighted Distortion Measures

arXiv:2608.16696v1 Announce Type: new Abstract: Physical AI systems such as autonomous vehicles and robots rely on timely exchange of high-dimensional sensory signals under tight bandwidth, latency, and energy budgets.…

Source: arXiv cs.LG Homa Esfahanizadeh, Matin Mortaheb, Jinfeng Du, Harish Viswanathan
AI Research AI

FinFraudBench: A Heterogeneous Graph Benchmark for Financial Fraud Detection

arXiv:2608.15177v1 Announce Type: new Abstract: The increasing complexity of digital financial systems has reshaped financial fraud detection from isolated transaction classification into relational risk reasoning over…

Source: arXiv cs.LG Yixuan Chen, Hongyu Zhan, Jie Sheng, Weiyu Han, Shuai Chen, Tianyi Zhang, Xiao Tan, Jun Xia
AI Research AI

Pointer Networks with Q-Learning for Combinatorial Optimization

arXiv:2311.02629v5 Announce Type: replace Abstract: We introduce the Pointer Q-Network (PQN), a hybrid neural architecture that integrates model-free Q-value policy approximation with Pointer Networks (Ptr-Nets) to…

Source: arXiv cs.LG Alessandro Barro
AI Research AI

The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks

arXiv:2608.16630v1 Announce Type: cross Abstract: Repository-scale coding requires an agent to keep tests, imports, configuration, and migration rules consistent within a bounded context window. We model this as…

Source: arXiv cs.LG Bardia Mohammadi, Lars Klein, Aman Chadha, Akhil Arora, Laurent Bindschaedler
AI Research AI

Scale-Consistent Posterior Dynamics for Diffusion Inverse Problems

arXiv:2608.15144v1 Announce Type: cross Abstract: Posterior sampling with a pretrained diffusion prior is governed by a conditional score whose intermediate likelihood component is generally intractable. We begin from…

Source: arXiv cs.LG Zhaoqiang Liu, Tongyao Pang, Ruibing Wang, Yang Zheng
AI Research AI

IP Protection in the Era of Visual Generative AI: A Survey

arXiv:2608.14730v1 Announce Type: cross Abstract: The rapid evolution of visual generative AI has introduced a wide range of intellectual property risks, spanning the unauthorized learning, reproduction, extraction,…

Source: arXiv cs.LG Zhuan Shi, Shunchang Liu, Alireza Dehghanpour Farashah, Qian Yang, Han Yu, Cao Yang, Chaochao Chen, Yuping Yan, Yaochu Jin, Golnoosh Farnadi, Lingjuan Lyu
AI Research AI

Multi-Granularity Sentiment Integration for LLM-Based Multimodal Sentiment Analysis

arXiv:2608.16201v1 Announce Type: new Abstract: Multimodal sentiment analysis (MSA) aims to predict sentiment polarity and intensity from heterogeneous inputs such as text, audio, and vision. While large language models…

Source: arXiv cs.LG Shanshan Lin, Yuesheng Wu, Chao Chen, Yizhe Yang, Zhihao Chen, Zexian Yang, Xiangwen Liao
AI Research AI

A Theoretical Framework for Statistical Evaluability of Generative Models

arXiv:2604.05324v3 Announce Type: replace Abstract: Statistical evaluation aims to estimate the generalization performance of a model using held-out i.i.d. test data sampled from the ground-truth distribution. In…

Source: arXiv cs.LG Shashaank Aiyer, Yishay Mansour, Shay Moran, Han Shao
AI Research AI

LACE-SVD: Loss-Aware SVD with Cumulative Error Correction for LLM Compression

arXiv:2607.03057v2 Announce Type: replace Abstract: The rapid growth in the parameter scale of large language models (LLMs) has created a strong demand for efficient compression techniques. As a hardware-agnostic and…

Source: arXiv cs.LG Zhuowen Liu, Longkun Hao, Shiyu Feng, Xiaowen Chang, Ruiqun Li, Changqun Li
AI Research AI

Spectral Saliency for Machine Unlearning

arXiv:2608.15548v1 Announce Type: new Abstract: Machine unlearning (MU) aims to remove the influence of specific training data while preserving model utility. As the name suggests, MU can be viewed as the inverse of…

Source: arXiv cs.LG Cedar Site Bai, Amber Yijia Zheng, Raymond A. Yeh, Brian Bullins
AI Research AI

Tail-Aware Top-$k$ On-Policy Distillation

arXiv:2608.14728v1 Announce Type: new Abstract: On-policy distillation (OPD) has emerged as an effective paradigm for transferring knowledge between language models, where a student is trained to align its next-token…

Source: arXiv cs.LG Huipeng Huang, Hongxin Wei
AI Research AI

The Null Token Knows: Reducing Message-Free Hallucination in ASR and NMT

arXiv:2608.15940v1 Announce Type: cross Abstract: Modern encoder-decoder systems can produce fluent text even when their input contains no recoverable message. We study this failure in ASR and NMT through the models'…

Source: arXiv cs.LG Kirill Borodin, Vasiliy Kudryavtsev, Ivan Viakhirev
AI Research AI

Comprehensive language-image pre-training for 3D medical image understanding

arXiv:2510.15042v3 Announce Type: replace-cross Abstract: In the 3D medical image domain, vision-language pre-training is used to create vision-language encoders (VLEs) that can support radiologists by retrieving…

Source: arXiv cs.LG Tassilo Wald, Ibrahim Ethem Hamamci, Yuan Gao, Sam Bond-Taylor, Harshita Sharma, Maximilian Ilse, Cynthia Lo, Olesya Melnichenko, Anton Schwaighofer, Noel C. F. Codella, Maria Teodora Wetscherek, Kla…
AI Research AI

Optimal Lower Bounds for Networked Information Aggregation

arXiv:2608.15472v1 Announce Type: new Abstract: The problem of networked information aggregation, studied in Kearns et al. (2026), involves a group of learners situated on the vertices of a directed acyclic graph $G$,…

Source: arXiv cs.LG Ambar Pal
AI Research AI

Near-Equilibrium Propagation training in nonlinear wave systems

arXiv:2510.16084v3 Announce Type: replace Abstract: Backpropagation learning algorithm, the workhorse of modern artificial intelligence, is notoriously difficult to implement in physical neural networks. Equilibrium…

Source: arXiv cs.LG Karol Sajnok, Micha{\l} Matuszewski
AI Research AI

LLMs for Zero-Shot Threat Detection via Structured Risk Indicators

arXiv:2608.16508v1 Announce Type: cross Abstract: We propose a two-stage large language model (LLM) framework for zero-shot detection of insider threats and advanced persistent threats (APTs) from heterogeneous security…

Source: arXiv cs.LG Abdullah Alghamdi, Siamak Layeghy, Marius Portmann
AI Research AI

Learning to Unlearn: Machine Unlearning via Learning the Unlearning Behaviors

arXiv:2608.16700v1 Announce Type: new Abstract: Various machine unlearning techniques have been developed in response to privacy legislation requirements, enabling individuals to exercise their legal right to have their…

Source: arXiv cs.LG Hang Zhang, Kaifeng Zhang, Yixiao Ma, Weijie Xu, Ye Zhu, Kai Ming Ting
AI Research AI

Fiber Fingerprints of Hidden Learning-State Dynamics

arXiv:2608.15976v1 Announce Type: new Abstract: A learning system can occupy execution states that are indistinguishable under every declared present-behavior readout yet respond differently to future training. We…

Source: arXiv cs.LG Qinyou Wang
AI Research AI

Degeneracy Counting Quantum Algorithm using Decoherence

arXiv:2608.14941v1 Announce Type: new Abstract: Counting the global optima of a classical optimization problem is a #P-hard task. We develop the canonical thermal pure quantum (CTPQ) state-based degeneracy counting…

Source: arXiv cs.LG Malay Marut Das, Mark A. Novotny, Yaroslav Koshka
AI Research AI

Towards a theory of inference-time alignment with unknown rewards

arXiv:2608.15402v1 Announce Type: new Abstract: Generative model alignment has received broad interest, and significant progress has been made in supervised fine-tuning and inference-time computation. Yet, alignment has…

Source: arXiv cs.LG Steve Hanneke, Hongao Wang, Mingyue Xu
AI Research AI

Correlation Clustering with Random Partial Information

arXiv:2608.16315v1 Announce Type: cross Abstract: Correlation clustering is a fundamental unsupervised learning problem. On complete graphs, both the min-disagreement and min-max objectives admit constant-factor…

Source: arXiv cs.LG Rajath Rao K. N., Jens Schl\"oter, Sami Davies, Amira Ouchene, Yasamin Nazari
AI Research AI

Differentiable Thermodynamic Phase-Equilibria for Machine Learning

arXiv:2603.11249v4 Announce Type: replace Abstract: Accurate prediction of phase equilibria remains a central challenge in chemical engineering. Physics-consistent machine learning methods that incorporate thermodynamic…

Source: arXiv cs.LG Karim K. Ben Hicham, Moreno Ascani, Jan G. Rittig, Alexander Mitsos
AI Research AI

Efficient Neural-Network-Based High-Resolution Radiative Transfer for CO___ Retrieval, and Application to Interferometric Sensing

arXiv:2608.14645v1 Announce Type: new Abstract: Studying climate change requires reducing uncertainties in CO2 and CH4 emission estimates to better distinguish anthropogenic from natural sources, which motivates…

Source: arXiv cs.LG Jordan Lontsi Tedongmo (CB), Yann Ferrec (CB, IFUMI), Laurence Croiz\'e (CB, IFUMI), Pablo Mus\'e (CB, IFUMI), Gabriele Facciolo (CB), Andr\'es Almansa (MAP5 - UMR 8145, IFUMI)
AI Research AI

Learning Varying Physical Therapist-Patient Interactions for Robot-mediated Upper Limb Task-Specific Training

arXiv:2608.15995v1 Announce Type: cross Abstract: Upper extremity motor function recovery is positively linked to Task-Specific Training (TST) and sufficient therapy dosage. Rehabilitation robots can increase TST dosage…

Source: arXiv cs.LG Jia Quan Loh (Human Robotics Laboratory, Department of Mechanical Engineering, The University of Melbourne), Vincent Crocher (Human Robotics Laboratory, Department of Mechanical Engineering, The Univ…
AI Research AI

Q-Regularized Generative Auto-Bidding: From Suboptimal Trajectories to Optimal Policies

arXiv:2601.02754v3 Announce Type: replace Abstract: With the rapid development of e-commerce, auto-bidding has become a key asset in optimizing advertising performance under diverse advertiser environments. The current…

Source: arXiv cs.LG Mingming Zhang, Na Li, Zhuang Feiqing, Hongyang Zheng, Jiangbing Zhou, Wang Wuyin, Sheng-jie Sun, XiaoWei Chen, Junxiong Zhu, Lixin Zou, Chenliang Li
AI Research AI

Equilibrium Forcing: Adaptive Video Generation Without Noise Conditioning

arXiv:2608.14706v1 Announce Type: cross Abstract: Standard autoregressive video generation algorithms based on Diffusion and Flow Matching rely on rigid training objectives and static sampling schedules, limiting…

Source: arXiv cs.LG Hansen Jin Lillemark, Alex Rojas, Zachary Novack, Runqian Wang, Yilun Du, Yian Ma, Taylor Berg-Kirkpatrick, Rose Yu
AI Research AI

Fast Test-Time Refinement for Robust Learned Image Compression

arXiv:2608.15113v1 Announce Type: cross Abstract: Learned image compression (LIC) has demonstrated remarkable rate-distortion (RD) performance in benign settings. However, the high representational capacity endowed by…

Source: arXiv cs.LG Jiaming Liang, Chi-Man Pun, Weisi Lin
AI Research AI

Graph Machine Learning: An Opportunity for Power Systems

arXiv:2608.16494v1 Announce Type: new Abstract: Modern power systems face growing operational complexity driven by the integration of renewable energy sources, decentralization, and the need for real-time…

Source: arXiv cs.LG Martin Sadric, Sebastian P\"utz, Christian Nauck, Veit Hagenmeyer, Frank Hellmann, Dirk Witthaut, Benjamin Sch\"afer
AI Research AI

6G Native AI and Channel Foundation Models

arXiv:2608.14591v1 Announce Type: cross Abstract: The integration of artificial intelligence (AI) and wireless communications is widely regarded as a core objective of sixth-generation (6G) systems. However, both the…

Source: arXiv cs.LG Shugong Xu, Jun Jiang, Yuan Gao
AI Research AI

Mint-Agent: Introducing Finance-Native Agentic Foundation Models

arXiv:2608.16386v1 Announce Type: cross Abstract: Financial agents must do more than recall domain knowledge: they must be both reliable, executing precise operations over grounded evidence, and executive, sustaining…

Source: arXiv cs.LG Agent Team, B. Zhang, Yaze Geng, Lei Tang, Yaoyang Yi, Zonghan Wu, Yifan Hu, Kun Wang, Qingsong Wen, Yilei Shao
AI Research AI

SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA

arXiv:2509.25459v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) show promise in generating long-form scientific explanations that synthesize evidence and connect multiple factors. However, in…

Source: arXiv cs.LG Haozhou Xu, Dongxia Wu, Matteo Chinazzi, Ruijia Niu, Rose Yu, Yi-An Ma
AI Research AI

Hardware-in-the-Loop Phase-Aware CNN for Real-Time 5G Channel Estimation

arXiv:2608.14709v1 Announce Type: cross Abstract: This demo presents real-time AI-based uplink channel-estimation inference using data collected from a hardware-in-the-loop 5G platform. The data-collection setup…

Source: arXiv cs.LG Javad Zolfaghari-Bengar, Rakibul Rony, Elisa Gomez-de-Lope, Alejandro Villena-Rodriguez, Abhinav Mahadevan, Nicolas Kourtellis
AI Research AI

BrainLinear: A Linear Model for Brain Network Analysis in Sparse Tangent Subspaces

arXiv:2608.15266v1 Announce Type: cross Abstract: Functional connectome analysis examines brain-region interactions to understand and identify disorders such as autism spectrum disorder and Alzheimer's disease. Existing…

Source: arXiv cs.LG Sijing Wu, Dongyuan Li, Miaoting Huang, Weiwei Ye, Ying Zhang, Feng Xia, Renhe Jiang
AI Research AI

Speculative Rollback Correction for Quality-Diverse Web Agent Imitation

arXiv:2606.12485v2 Announce Type: replace Abstract: Training interactive web agents through imitation learning from expert trajectories has emerged as a highly effective approach. However, determining the optimal timing…

Source: arXiv cs.LG Longkun Hao, Hongyu Lin, Hao Li, Zhuowen Liu, Zhichao Yang, Haojie Hao, Dongshuo Huang, Haitao Yang, Hongyu Ge, Ming jie Xie, Yanjun Wu, Zi Hao Yin, Yan Bai, Yihang Lou
AI Research AI

Global Convergence of DGM and PINN Algorithms for Solving Nonlinear PDEs

arXiv:2607.24726v2 Announce Type: replace Abstract: The Deep Galerkin Method (DGM) and Physics Informed Neural Networks (PINNs) have become widely-used methods for solving partial differential equations (PDEs) in the…

Source: arXiv cs.LG Justin Sirignano, Konstantinos Spiliopoulos, Samuel Cohen
AI Research AI

SCPP: A Unified Python Library for Soft Clustering

arXiv:2607.19620v2 Announce Type: replace Abstract: In this paper, we present SCPP (Soft Clustering Python Package), an open-source Python framework for soft clustering. SCPP establishes a canonical,…

Source: arXiv cs.LG Kiyan Rezaee, Morteza Ziabakhsh, Artin Bahrampour, Seyed Mohammad Ghoreishi, Asal Khaje, Ali Sajedifar, Manny Chalak, Ava Zerafatangiz, Sadegh Eskandari
AI Research AI

WANDR: A Benchmark for Wide and Deep Research

arXiv:2608.14747v1 Announce Type: new Abstract: WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, challenging data-collection tasks for research agents. Each task requires a system to discover a large set…

Source: arXiv cs.LG Vitaliy Polshkov, Marcin Pitera, Jeremy Yang, Kirill Priemko, Maksim Gaiduk, Aleksandr Nikolenko, Denis Bykov, Clare Southern, Denis Yarats, Jerry Ma
AI Research AI

Domain-Specific Text Embedding Models for Entity Resolution

arXiv:2608.16161v1 Announce Type: cross Abstract: General-purpose text embedding models are designed to capture semantic similarity but are not optimised for distinguishing entity records that represent the same…

Source: arXiv cs.LG Khajesh Sapram, Srivardhani Raju, Kishore Konda
AI Research AI

MiNO: Cotangent-bundle propagator learning for PDEs

arXiv:2608.15187v1 Announce Type: new Abstract: Scientific machine learning for partial differential equations commonly targets solution fields, as in physics-informed neural networks, or solution maps, as in neural…

Source: arXiv cs.LG Gnankan Landry Regis N'guessan, Bum Jun Kim
AI Research AI

A Unified Geometric Framework for Developmental Analysis of Spatial Transcriptomic Data

arXiv:2608.15306v1 Announce Type: cross Abstract: High-throughput single-cell and spatial transcriptomic technologies provide high-resolution snapshots of heterogeneous cellular states, but their destructive nature…

Source: arXiv cs.LG Mary Chriselda Antony Oliver, Kaitlyn Hohmeier, Tuyen Tran, Alejandra Castillo, Caroline Moosm\"uller, Shiying Li
AI Research AI

Workspace Topology as an Attack Vector in Agentic Coding Assistants

arXiv:2608.14876v1 Announce Type: cross Abstract: Agentic coding assistants are finding widespread use, not just in new code development but in quickly ingesting and leveraging third-party code. This opens up a risk of…

Source: arXiv cs.LG Alexandre G. R. Day, Pradeep Yadlapalli, Sriram Venkatapathy, Thomas Paniagua, Nick Raines, Sahil Wadhwa, Himanshu Kumar, Andy Luo, Sudeep Panyam, Rikhiya Ghosh, Pranab Mohanty, Giri Iyengar
AI Research AI

Information Geometry of Message Passing

arXiv:2608.15922v1 Announce Type: new Abstract: We show that the natural-gradient stationary condition of variational inference has an edge-local form on a Forney-style factor graph. We start from the Bethe free energy…

Source: arXiv cs.LG Mykola Lukashchuk, Kyrylo Yemets, Alex Ledbetter, \.{I}smail \c{S}en\"oz
AI Research AI

Enhancing Differentially Private Linear Regression via Public Second-Moment

arXiv:2508.18037v2 Announce Type: replace Abstract: Leveraging information from public data has become increasingly crucial in enhancing the utility of differentially private (DP) methods. Traditional DP approaches…

Source: arXiv cs.LG Zilong Cao (The School of Mathematics, Northwest University), Hai Zhang (The School of Mathematics, Northwest University)
AI Research AI

Spinning Conformal Correlators from Neural Networks

arXiv:2608.15001v1 Announce Type: cross Abstract: We construct spinning conformal fields from neural networks and the embedding formalism, computing their two-, three- and four-point functions in examples, building on…

Source: arXiv cs.LG Manas Dogra, James Halverson, Joydeep Naskar
AI Research AI

Explaining Reinforcement Learning Decisions in Self-adaptive Systems

arXiv:2608.14620v1 Announce Type: new Abstract: Reinforcement Learning (RL) has been extensively used in autonomous and self-* systems, but RL policies, especially deep RL ones relying on neural networks, lack…

Source: arXiv cs.LG Jasmina Gajcin, Juan C. Rosero, Ivana Dusparic
AI Research AI

Quantifying Depth Sufficiency in Residual Neural Networks: A First-Order Criterion

arXiv:2608.14664v1 Announce Type: new Abstract: How can we determine whether a trained neural network is already deep enough? We study this under a fixed function-preserving residual-growth protocol specifying insertion…

Source: arXiv cs.LG Zeyu Liu, Jinhao Zhang, Yunquan Zhang, Guangming Tan, Xiang Gao, Fangming Liu, Daning Cheng
AI Research AI

Model Inversion Attacks: A Survey of Approaches and Countermeasures

arXiv:2411.10023v3 Announce Type: replace Abstract: Deep neural networks have enabled numerous studies and applications on both Euclidean data, such as images and text, and non-Euclidean data, such as graphs. Because…

Source: arXiv cs.LG Zhanke Zhou, Jianing Zhu, Fengfei Yu, Xuan Li, Xiong Peng, Tongliang Liu, Bo Han
AI Research AI

Spectral Gaps of Hit-and-Run and Coordinate Hit-and-Run

arXiv:2608.16878v1 Announce Type: cross Abstract: For any convex body $\mathcal{K}\subset\mathbb{R}^{n}$ containing a unit ball, the spectral gap of Hit-and-Run is $\Omega(1/(n^2 C_{\mathsf{PI}}))$, where…

Source: arXiv cs.LG Yunbum Kook, Santosh S. Vempala
AI Research AI

ROC-n-reroll: How verifier imperfection affects test-time scaling

arXiv:2507.12399v3 Announce Type: replace Abstract: Test-time scaling aims to improve language model performance by leveraging additional compute during inference. Many works have empirically studied techniques such as…

Source: arXiv cs.LG Florian E. Dorner, Yatong Chen, Andr\'e F. Cruz, Fanny Yang
AI Research AI

Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling

arXiv:2608.14652v1 Announce Type: new Abstract: The development of 0.1$^{\circ}$ global weather forecasting models based on machine learning (ML) is constrained by the limited availability of high-resolution data, as…

Source: arXiv cs.LG Yang Zhao, Peisong Niu, Tian Zhou, Ziqing Ma, Guanlong Ma, Rong Jin, Huiling Yuan, Liang Sun
AI Research AI

Multi-Bin Batching for Increasing LLM Inference Throughput

arXiv:2412.04504v2 Announce Type: replace-cross Abstract: As large language models (LLMs) grow in popularity for their diverse capabilities, improving the efficiency of their inference systems has become increasingly…

Source: arXiv cs.LG Ozgur Guldogan, Jackson Kunde, Kangwook Lee, Ramtin Pedarsani
AI Research AI

RigidBench: Evaluating Rigid-Body Physics in Video Generation Models

arXiv:2608.15555v1 Announce Type: cross Abstract: Video models are increasingly used to predict what happens next in a scene, yet the metrics commonly used to compare their outputs say little about whether the predicted…

Source: arXiv cs.LG Swarnim Jain, Shangzhe Wu
AI Research AI

Mitigating Rubric Interference in LLM Judges via On-Policy Self-Distillation

arXiv:2608.14684v1 Announce Type: new Abstract: LLM judges increasingly evaluate responses against fine-grained rubric checklists. When a sample requires multiple rubrics, current methods typically assess each in a…

Source: arXiv cs.LG Dingyao Yu, Tong Zhang, Yutao Mou, Yunxiao Zhang, Wei Ye, Shikun Zhang
AI Research AI

P2E-VQ: ECG-linked representation augmentation for PPG via discrete patch retrieval

arXiv:2608.14656v1 Announce Type: new Abstract: Photoplethysmography (PPG) is widely used in consumer wearables because of its low cost and ease of acquisition. However, unlike electrocardiography (ECG), PPG measures…

Source: arXiv cs.LG Zhongli Wu, Zhuangzhi Gao, He Zhao, Feixiang Zhou, Fu Wang, Jinru Ding, Yuankai Wang, Hongyi Qin, Gregory Y. H. Lip, Bil Kirmani, Yalin Zheng
AI Research AI

Scaling Laws for Dynamic Mini-Batch SGD in Sketched Linear Regression

arXiv:2605.24316v3 Announce Type: replace Abstract: Mini-batching is central to large-scale optimization, yet its role in statistical scaling laws remains limited. We study one-pass and multi-pass batch SGD for sketched…

Source: arXiv cs.LG Ziyan Chen, Zhongzhu Zhou, Ding-Xuan Zhou
AI Research AI

Multi-Feature Riemannian Hypergraph for Online Test-Time Adaptation of Motor Imagery Brain-Computer Interface

arXiv:2608.16134v1 Announce Type: new Abstract: In clinical motor imagery brain-computer interface (MI-BCI) decoding, cross-day transferability and online operation remain two critical challenges. Hypergraphs can…

Source: arXiv cs.LG Siqi Li (Peking University, Chinese Institute for Brain Research, Beijing), Zhi Li (NeuCyber Neurotech), Tong Liu (NeuCyber Neurotech), Shuai Zhang (NeuCyber Neurotech), Yanfei Jia (Beijing Medical U…
AI Research AI

AMPLIFAI: A Multiphase CT Dataset for Benchmarking Clinical Reasoning in LI-RADS Assessment of Liver Lesions

arXiv:2608.14778v1 Announce Type: cross Abstract: Hepatocellular carcinoma (HCC) is the third leading cause of cancer-related mortality worldwide, with early detection improving survival from <20\% to >70\%. The…

Source: arXiv cs.LG Pranav Kulkarni, Nikhil Shah, Amritansh Suryavanshi, Jana Delfino, James Tonascia, Jade Wong-You-Cheong, Barton Lane, Joseph Chirico, Jeffrey D. Hirsch, Ang Li, Heng Huang, Florence X. Doo
AI Research AI

Joint MDPs and Reinforcement Learning in Coupled-Dynamics Environments

arXiv:2603.06946v2 Announce Type: replace Abstract: Many distributional quantities in reinforcement learning are intrinsically joint across actions, including distributions of gaps and probabilities of superiority.…

Source: arXiv cs.LG Ege C. Kaya, Mahsa Ghasemi, Abolfazl Hashemi
AI Research AI

Unifying Graph Neural Networks Through a Common Layer Equation

arXiv:2608.16097v1 Announce Type: new Abstract: Graph neural networks are commonly described through family-specific equations whose notation obscures shared computations and structural differences. We introduce a…

Source: arXiv cs.LG Sai Karthik Navuluru, Siddhartha Shankar Das, Bo Ni, Hongjie Chen, Yu Wang, Baris Coskunuzer, Nesreen K. Ahmed, Franck Dernoncourt, Mahantesh Halappanavar, Tyler Derr, Ryan A. Rossi, Lakshman Tamil
AI Research AI

In Defense of OCTA: The Reconstruction-Utility Gap in OCT-to-OCTA Synthesis

arXiv:2608.15626v1 Announce Type: new Abstract: Optical coherence tomography angiography (OCTA) images retinal blood flow, giving capillary-perfusion and foveal-avascular-zone biomarkers that grade diabetic-retinopathy…

Source: arXiv cs.LG Michael Chertok, Alon Tiosano, Orly Gal-Or, Lior Kramarski, Einav Baharav Shlezinger, Irit Bahar, Lior Wolf
AI Research AI

OceanDepths: A Global Dataset of Paired Subsurface and Surface Ocean Observations

arXiv:2608.16373v1 Announce Type: new Abstract: Despite comprising over 70\% of its surface, the world's oceans are critically underobserved compared to the land surface or the atmosphere.Understanding the global ocean…

Source: arXiv cs.LG Simon Donike, Ruben Cartuyvels, Antonino Ian Ferola, Elisa Carli, Diego Fernandez Prieto, Marie-Helene Rio
AI Research AI

Rethinking Reverse KL as Adaptive Entropy Distillation

arXiv:2608.14685v1 Announce Type: new Abstract: Knowledge distillation (KD) is widely used to transfer the capabilities of large language models (LLMs) to smaller students, but existing objectives often struggle to…

Source: arXiv cs.LG Shizhen Li, Zhiyu Shen, Yuyin Lu, Yunhe Pang, Jielin Song, Yanghui Rao, Fu Lee Wang
AI Research AI

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens

arXiv:2506.03096v2 Announce Type: replace-cross Abstract: Contrastive language-image pre-training aligns features of text-image pairs in a common latent space via distinct encoders for each modality. While this approach…

Source: arXiv cs.LG Christian Schlarmann, Francesco Croce, Nicolas Flammarion, Matthias Hein
AI Research AI

Prompting is not enough: supervised baselines and leakage control for measuring shared decision-making with LLMs in pediatric encounters

arXiv:2608.14792v1 Announce Type: cross Abstract: Objectives: To determine whether zero-shot prompting of a large language model (LLM) is sufficient to detect shared decision-making (SDM) behaviors in real clinical…

Source: arXiv cs.LG Bernardo Modenesi, Jody Lin, Kimberly Kaphingst, Angela Zhu, Maya Wheeler, Peilu Zhang, Angela Fagerlin
AI Research AI

A Privacy Study of Sparse Collaborative Inference

arXiv:2608.16236v1 Announce Type: new Abstract: Collaborative inference (CI) splits a model between an edge device and a server, whereby the client computes an intermediate activation, transmits it, and the server…

Source: arXiv cs.LG Maximilian Andreas Hoefler, Karsten Mueller, Wojciech Samek
AI Research AI

SAUL: Sharpness-Aware Augmented-Lagrangian Unlearning

arXiv:2608.16249v1 Announce Type: new Abstract: Machine unlearning in Large Language Models (LLMs) faces a critical trade-off between erasing target knowledge and preserving general utility. We propose SAUL…

Source: arXiv cs.LG Jaewan Choi, Junyoung Yang, Sangdon Park
AI Research AI

Invariant Pretraining for Robust Code Representations

arXiv:2608.15412v1 Announce Type: new Abstract: Encoder-based code representation models remain widely deployed for discriminative tasks such as clone detection and code classification, where their small size and low…

Source: arXiv cs.LG Yifeng He, Yundi Xu, Christopher Castro Gaw Gonzalo, Zili Wang, Hao Chen
AI Research AI

Multilook Coherent Imaging: Theoretical Guarantees and Algorithms

arXiv:2505.23594v2 Announce Type: replace-cross Abstract: Multilook coherent imaging is a widely used technique in applications such as digital holography, ultrasound imaging, and synthetic aperture radar. A central…

Source: arXiv cs.LG Xi Chen, Soham Jana, Christopher A. Metzler, Arian Maleki, Shirin Jalali
AI Research AI

Le Critique: Privileged Value Functions for LLM Reinforcement Learning

arXiv:2608.16739v1 Announce Type: new Abstract: Reinforcement learning algorithms for Large Language Models (LLMs) are largely distinguished by their variance reduction strategy. Group-relative methods like GRPO reduce…

Source: arXiv cs.LG Siddarth Venkatraman, Matthieu Dinot, Laurence Aitchison
AI Research AI

ShadowNet for Data-Centric Quantum System Learning

arXiv:2308.11290v2 Announce Type: replace-cross Abstract: Understanding the dynamics of large quantum systems is hindered by the curse of dimensionality. Statistical learning offers new possibilities in this regime…

Source: arXiv cs.LG Yuxuan Du, Yibo Yang, Tongliang Liu, Zhouchen Lin, Bernard Ghanem, Dacheng Tao
AI Research AI

Learning Optimal Dynamic Matching via Graph Neural Networks

arXiv:2607.28925v2 Announce Type: replace Abstract: Dynamic matching markets require decisions about whom to match and when: matching now yields value but removes participants who may create better future opportunities.…

Source: arXiv cs.LG Genta Okada, Shunya Noda, Junpei Komiyama, Akira Matsushita
AI Research AI

In-Context Learning to Assess Built Environment Impacts on Perceived Neighborhood Walkability Among Mobility-impaired Older Adults

arXiv:2608.14663v1 Announce Type: new Abstract: As global populations age, enhancing neighborhood walkability through inclusive urban design is important for mitigating built environment (BE) barriers that discourage…

Source: arXiv cs.LG Houhao Liang, Kresimir Friganovic, Joanne Kua, Noor Hafizah Ismail, Su Su, Bryan Yijia Tan, Navrag B. Singh, Panos Mavros
AI Research AI

Memory-Bounded Continuation of Greedy Sampling for Continual Anomaly Detection

arXiv:2608.15277v1 Announce Type: cross Abstract: Greedy sampling produces a compact yet representative summary of normal data, which is essential for reliable anomaly detection that relies on measuring distance from…

Source: arXiv cs.LG Yoon Gyo Jung, Jaewoo Park, Kuan-Chuan Peng, Seongdeok Bang, Octavia Camps
AI Research AI

CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?

arXiv:2608.16829v1 Announce Type: new Abstract: Video world models approximate the stochastic distribution of physical outcomes through generative sampling, but existing benchmarks score individual generations or…

Source: arXiv cs.LG Jonathan Sadeghi, Jenny Seidenschwarz, Jesse Allardice, Sirish Srinivasan, Benjamin Graham, Jeffrey Hawke
AI Research AI

KOALA: Koopman Operator Learning for WiFi-Based Anticipatory Hum

arXiv:2608.15815v1 Announce Type: new Abstract: WiFi Channel State Information (CSI) has emerged as a privacy-preserving alternative to cameras for human pose estimation. However, existing approaches treat pose…

Source: arXiv cs.LG Quang-Anh N. D., Duc Pham Minh, Thao Phuong Pham, Minh Anh Nguyen, Huan X. Nguyen, Tuan Dang
AI Research AI

ReliaGate: Reliability Routing for Low-Stakes Wearable Stress Prediction

arXiv:2608.15951v1 Announce Type: new Abstract: We study when a wearable stress system should surface a prediction rather than change it. In low-stakes reflection and summary settings, aggregate accuracy is insufficient…

Source: arXiv cs.LG Jaden Moon, Yu Wu, Arvind Pillai, Andrew Campbell
AI Research AI

Efficient Safety Alignment of Language Models via Latent Personality Traits

arXiv:2607.07918v2 Announce Type: replace Abstract: Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alternatives. Latent Adversarial…

Source: arXiv cs.LG Mohamed Amine Merzouk, Nolan Smyth, Damiano Fornasiere, Linh Le, David Williams-King, Adam Oberman
AI Research AI

Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment

arXiv:2608.15693v1 Announce Type: cross Abstract: Running large AI models on resource-constrained edge devices requires model compression to reduce model size and computation. What compresses well, however, need not…

Source: arXiv cs.LG Subhransu Das, Jiaming Cheng, Arnav Kumar, Sadia Afrose, Mingzhe Han, Michael Silagy, Shreya Palande, Brijesh Soni, Rajiv Ramnath
AI Research AI

The Distributional View of Knowledge Distillation

arXiv:2608.15215v1 Announce Type: cross Abstract: Token-level knowledge distillation (KD) matches two conditional distributions per position, yet the standard objectives compare them pointwise: a Kullback-Leibler…

Source: arXiv cs.LG Gordei Verbii, Juho Lee
AI Research AI

Lipschitz Bandits with Arbitrary Feedback Delays

arXiv:2608.15036v1 Announce Type: new Abstract: The Lipschitz bandit problem extends the traditional multi-armed bandit framework to continuous action spaces by assuming that the reward functions satisfy a Lipschitz…

Source: arXiv cs.LG Yuhao Liu, Yu Chen, Longbo Huang
AI Research AI

LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review

arXiv:2608.14626v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safety guarantees remain significantly weaker in low-resource and…

Source: arXiv cs.LG Valdini Douglace Lemofouet, Blessing Ngozi Uzor, Paula Chikaodinaka Anyanwu, Danielle Blanche Kapsa, Sukairaj Hafiz Imam, P Sam Sahil, Abigail Oppong, Tassallah Abdullahi, Clemencia Siro, Idris Abdul…
AI Research AI

ClawGym II: Exploring Black-Box RL on Agent Harness

arXiv:2608.16798v1 Announce Type: cross Abstract: Agent harnesses have substantially improved performance on long-horizon tasks by coordinating agent interactions with the environment. However, reinforcement learning…

Source: arXiv cs.LG Huatong Song, Fei Bai, Ming Yang, Renyuan Li, Jia Deng, Jujie He, Zhange Zhang, Daixuan Cheng, Yan Xing, Qi Yun, Xuxing Chen, Danyang Li, Feng Chang, Chuan Hao, Ran Tao, Jian Yang, Bryan Dai, Wayne X…
AI Research AI

PL-Guard: Probabilistic Logic Reasoning for LLM Guardrails

arXiv:2608.15673v1 Announce Type: new Abstract: Large language model guardrails can be viewed as policy-consistency problems: a system must determine which policy-relevant facts hold in a prompt-response pair and what…

Source: arXiv cs.LG Satchit Chatterji, Shihan Wang, Giovanni Sileno, Erman Acar
AI Research AI

Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

arXiv:2608.15869v1 Announce Type: cross Abstract: Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating…

Source: arXiv cs.LG Xiaoyu Zhu, Xinke Deng, Suresh Taddewadikar, Arnab Kumar Mondal, Zhongyu Jiang, Ian Fasel, Joerg Liebelt
AI Research AI

Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning

arXiv:2608.14851v1 Announce Type: cross Abstract: Learning and skill mastery require extensive and deliberate practice. In many learning settings, producing high-quality pedagogical materials can require a high level of…

Source: arXiv cs.LG Allen Nie, Anirudhan Badrinath, Nicholas Tomlin, Timothy Dai, Carissa Yip, Rose E Wang, Emma Brunskill, Chris Piech
AI Research AI

Improving the matrix multiplication exponent with modern optimization and AlphaEvolve

arXiv:2608.16884v1 Announce Type: cross Abstract: The current best bounds on the matrix multiplication exponent $\omega$ are obtained through a refinement of the laser method called combination loss analysis (Duan et…

Source: arXiv cs.LG Emilien Dupont, Marvin Eisenberger, Borislav Kozlovskii, Abbas Mehrabian, Francisco J. R. Ruiz, Abigail See, Renfei Zhou, Josh Alman, Virginia Vassilevska Williams, Matej Balog
AI Research AI

Turning spectra into images improves plant trait retrieval with 2D-CNNs

arXiv:2608.16661v1 Announce Type: cross Abstract: Hyperspectral reflectance spectroscopy enables non-destructive estimation of plant functional traits, yet current deep learning approaches process spectra as…

Source: arXiv cs.LG Javier Lopatin, Teja Kattenborn, Eya Cherif, Sebasti\'an Moreno
AI Research AI

Second-Moment Memory in Coordinatewise Adam

arXiv:2608.15824v1 Announce Type: new Abstract: Adam retains a moving average of past squared gradients in its denominator, but the optimization cost of this memory is not well understood. We show that second-moment…

Source: arXiv cs.LG Jeonseong Kim
AI Research AI

Decentralized Federated Learning by Partial Message Exchange

arXiv:2603.01730v4 Announce Type: replace Abstract: Decentralized federated learning (DFL) has emerged as a transformative server-free paradigm that enables collaborative learning over large-scale heterogeneous…

Source: arXiv cs.LG Shan Sha, Shenglong Zhou, Xin Wang, Lingchen Kong, Geoffrey Ye Li
AI Research AI

AutoSR: Automatic Symbolic Regression by Searching Research States

arXiv:2608.16876v1 Announce Type: cross Abstract: We introduce Automatic Symbolic Regression (AutoSR), a fully automated system that instantiates Research-Space Symbolic Regression by searching persistent scientific…

Source: arXiv cs.LG Kejia Zhang, Youran Sun, Xinyu Ren, Chugang Yi, Haizhao Yang
AI Research AI

Quantum Large Language Models via Tensor Network Disentanglers

arXiv:2410.17397v2 Announce Type: replace-cross Abstract: We introduce a framework for seamlessly integrating quantum computing into pretrained large language models (LLMs). The key idea is to construct a hybrid…

Source: arXiv cs.LG Borja Aizpurua, Fernando Loren, Saeed S. Jahromi, Sukhbinder Singh, Roman Orus
AI Research AI

Developing an Offshore Machine Learning Surface Layer Scheme

arXiv:2608.14935v1 Announce Type: cross Abstract: Turbulent fluxes between the surface and the atmosphere are typically parameterized using empirically fit relationships. Here we test machine learning techniques for…

Source: arXiv cs.LG Susan Dettling, Sue Ellen Haupt, Thomas Brummet, Patrick Hawbecker, Branko Kosovi\'c, David John Gagne
AI Research AI

Crystal-structure design by agentic AI in a language of motifs

arXiv:2608.15900v1 Announce Type: cross Abstract: Data-driven materials discovery interpolates more reliably than it extrapolates and seldom reaches new structure types. We present MatEvolve, an agentic-AI framework…

Source: arXiv cs.LG Dinh-Khiet Le, Minh-Quyet Ha, Hong-Phuc Vu-Dinh, Takashi Miyake, Hiori Kino, Hieu-Chi Dam
AI Research AI

Forward Pass Domain Adaptation (Without Cross-Layer Backpropagation)

arXiv:2608.14563v1 Announce Type: new Abstract: Forward-Pass-Only MLP training (FPO) adapts large language models without a backward pass through the model body, achieving 2.7--3.2x the throughput of standard…

Source: arXiv cs.LG Rivaan Patil, Simon Dennis, Hao Guo, Kevin Shabahang
AI Research AI

RT-Lynx: Putting GEMM Sparsity in the Right Place for Diffusion Models

arXiv:2605.26632v3 Announce Type: replace Abstract: Diffusion Transformers (DiT) achieve strong performance in image generation but incur substantial inference costs. While prior work has reduced this cost via…

Source: arXiv cs.LG Xing Cong, Hanlin Tang, Kan Liu, Tao Lan, Lin Qu, Chenhao Xie
AI Research AI

Non-Crossing Deep Quantile Regression for Distributional Survival Prediction

arXiv:2608.16864v1 Announce Type: cross Abstract: In survival analysis the way covariates act on the risk of an event often differs between early and late failure times, yet hazard- and mean-based summaries collapse…

Source: arXiv cs.LG Shuai Huang, Zhe Qu, Zhaowei Hua, Guohao Shen, Rui Tang, Hongtu Zhu
AI Research AI

Generative Learning of Separatrices

arXiv:2608.14743v1 Announce Type: new Abstract: The identification and reconstruction of the boundaries separating basins of attraction in multistable, multidimensional dynamical systems presents a fundamental challenge…

Source: arXiv cs.LG Ellis R. Crabtree, Dimitris G. Giovanis, Anastasia Georgiou, George Datseris, Ioannis G. Kevrekidis
AI Research AI

Iso-Riemannian Optimization on Learned Data Manifolds

arXiv:2510.21033v3 Announce Type: replace-cross Abstract: We develop a theory of iso-Riemannian optimization for problems constrained to learned data manifolds, a setting in which classical Riemannian optimization - and…

Source: arXiv cs.LG Willem Diepeveen, Melanie Weber
AI Research AI

A 2-Block Architecture for Real-Time EEG Gait Decoding: A Pilot Study

arXiv:2608.02083v2 Announce Type: replace Abstract: Closed-loop lower-limb exoskeleton control via Electroencephalography (EEG) remains limited by motion artifacts, low signal-to-noise ratio, and binary gait…

Source: arXiv cs.LG Shantanu Sarkar, Saurabh Prasad, Jose L. Contreras-Vidal
AI Research AI

Helios 2.0: A Robust, Ultra-Low Power Gesture Recognition System Optimised for Event-Sensor based Wearables

arXiv:2503.07825v3 Announce Type: replace-cross Abstract: We present an advance in wearable technology: a mobile-optimized, real-time, ultra-low-power event camera system that enables natural hand gesture control for…

Source: arXiv cs.LG Prarthana Bhattacharyya, Joshua Mitton, Ryan Page, Owen Morgan, Oliver Powell, Benjamin Menzies, Gabriel Homewood, Kemi Jacobs, Paolo Baesso, Taru Muhonen, Richard Vigars, Louis Berridge
AI Research AI

Robust Privacy: Inference-Stage Privacy through Certified Robustness

arXiv:2601.17360v3 Announce Type: replace Abstract: An adversary observing a model's released prediction can infer sensitive attributes of the queried input, or even reconstruct representatives of the model's training…

Source: arXiv cs.LG Jiankai Jin, Xiangzheng Zhang, Zhao Liu, Wenzhuo Xu, Dongdong Yang, Deyue Zhang, Quanchen Zou
AI Research AI

ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems

arXiv:2608.15424v1 Announce Type: cross Abstract: The rapid adoption of large language models has enabled the development of clinical multi-agent systems (MAS) capable of integrating multimodal patient data and…

Source: arXiv cs.LG Rakesh Sharma, Sydney Pugh, Cameron Beeche, Pankhuri Singhal, Rachel Wu, Margaret Eby, Jeffrey Duda, James Gee, Kyra O'Brien, Hersh Sagreiya, Marina Serper, Victoria Gershuni, Angela Bradbury, Anurag…
AI Research AI

Sequential Batch Learning in Finite-Action Linear Contextual Bandits

arXiv:2004.06321v2 Announce Type: replace Abstract: We study the sequential batch learning problem in linear contextual bandits with finite action sets, where the decision maker is constrained to split incoming…

Source: arXiv cs.LG Yanjun Han, Zhengqing Zhou, Zihao Hu, Jose Blanchet, Peter W. Glynn, Yinyu Ye, Zhengyuan Zhou
AI Research AI

Deep learning-based computed tomography (CT) derived body composition classifier for colorectal cancer patients

arXiv:2608.15712v1 Announce Type: cross Abstract: Background: Accurate body composition analysis using Computed Tomography (CT) scans is essential for assessing skeletal muscle area (SMA) and skeletal muscle density…

Source: arXiv cs.LG Eve Harling (James Watt School of Engineering, College of Science & Engineering, University of Glasgow, Glasgow, UK), Chattarin Pumtako (Academic Unit of Surgery, School of Medicine, College of Medic…
AI Research AI

Leveraging Machine Unlearning for Cost-Efficient Preference Alignment

arXiv:2504.06659v2 Announce Type: replace Abstract: Despite advances in Preference Alignment (PA) for Large Language Models (LLMs), mainstream methods like reinforcement learning with human feedback face notable…

Source: arXiv cs.LG Xiaohua Feng, Yuyuan Li, Huwei Ji, Jiaming Zhang, Li Zhang, Tianyu Du, Chaochao Chen
AI Research AI

Resource-Efficient QUBO Formulation for Anchored Currency Arbitrage

arXiv:2608.15889v1 Announce Type: cross Abstract: Currency arbitrage (CA) involves trading currencies in cycles to exploit discrepancies in market valuations. Quadratic unconstrained binary optimization (QUBO) involves…

Source: arXiv cs.LG Eric A. F. Reinhardt, Adam J. Hauser
AI Research AI

Phase-Aware CNN for Real-Time 5G/6G Channel Estimation with Hardware-in-the-loop Validation

arXiv:2608.14676v1 Announce Type: cross Abstract: In 5G/6G wireless systems, accurate and timely channel estimation is critical to ensure reliable communication under complex, fast-changing radio conditions. This work…

Source: arXiv cs.LG Javad Zolfaghari-Bengar, Rakibul Rony, Elisa Gomez-de-Lope, Alejandro Villena-Rodriguez, Abhinav Mahadevan, Nicolas Kourtellis
AI Research AI

Q-based Variational Inverse Reinforcement Learning

arXiv:2608.16888v1 Announce Type: new Abstract: The development of safe and beneficial AI requires that systems can learn and act in accordance with human preferences. However, explicitly specifying these preferences by…

Source: arXiv cs.LG Ondrej Bajgar, Peter Tisnikar, Alessandro Abate, Konstantinos Gatsis, Maike Osborne
AI Research AI

Optimal Watermark Localization in Mixed-Source Large Language Model Texts

arXiv:2608.14906v1 Announce Type: cross Abstract: Watermarking provides a principled way to authenticate text generated by large language models (LLMs). In practice, however, the final text may be mixed-source, with…

Source: arXiv cs.LG Jose H. Blanchet, T. Tony Cai, Xiang Li, Hao Liu, Qi Long, Weijie J. Su
AI Research AI

CoM$^3$eT: A foundation model for medical image analysis through federated, multidimensional context integration

arXiv:2608.16268v1 Announce Type: cross Abstract: Medical foundation models improve generalization when training AI models with limited labeled data, but remain confined to a single specialty, such as pathology or…

Source: arXiv cs.LG J. Raphael Sch\"afer, Kai Geissler, Till Nicke, Chiara Tappermann, Karoline Heber, Eike Petersen, Habib Mergan, Lars Ole Schwen, Nick Weiss, Annika Gerken, Jan Hendrik Moltz, Tom Bisson, Isil Dogan O…
AI Research AI

SchurQuant: Groupwise Discrete Optimization for Layer-Wise LLM Quantization

arXiv:2608.15567v1 Announce Type: new Abstract: Weight-only post-training quantization (PTQ) enables the deployment of large language models under tight memory budgets, but accuracy often collapses at 2-3 bits. Existing…

Source: arXiv cs.LG Gunjun Lee, Sehwan Son, Younjoo Lee, Byungjun Kim, Jung Ho Ahn
AI Research AI

Randomly initialized autoencoders: fixed points and edge-of-chaos

arXiv:2608.14638v1 Announce Type: new Abstract: In this paper we study autoencoders, a special class of deep neural nets (DNNs) whose performance can be characterized via their fixed points. This perspective naturally…

Source: arXiv cs.LG Leonid Berlyand, Roman Sarapin, Yitzchak Shmalo, Victor Slavin, Sasha Sodin
AI Research AI

T-LLM Compiler: Trusted LLM-based Code Optimization and Verification Framework

arXiv:2608.14953v1 Announce Type: cross Abstract: Recent advances in Large Language Models (LLMs) have opened opportunities to apply high-level code transformations to the field of code optimization, and it has since…

Source: arXiv cs.LG Zahra Fazel, Sunanda Gamage, Shayan Shirahmad Gale Bagi, Amir H. Ashouri, Tomasz S. Czajkowski, Bryan Chan, Reza Azimi, Yaoqing Gao
AI Research AI

Learning reshapes power-law anisotropy in internal representations

arXiv:2608.15239v1 Announce Type: new Abstract: Power-law anisotropy in internal representations has been observed across a wide range of biological and artificial neural systems, from state-of-the-art language models…

Source: arXiv cs.LG Asahi Nakamuta, Jun-nosuke Teramae
AI Research AI

MAPLE: MoE Adaptive Plug-and-play Layer-wise Expert allocation

arXiv:2608.15299v1 Announce Type: new Abstract: Sparsely-activated Mixture-of-Experts (MoE) Transformers universally fix the same number of routed experts across all layers, a convention that ignores the well-documented…

Source: arXiv cs.LG Lie Li, Wen Li, Junxiao Shen, Gusheng Hu
AI Research AI

Does 1/2-Tsallis-INF Also Work Well for Best-Arm Identification?

arXiv:2608.15365v1 Announce Type: new Abstract: Regret minimization (RM) and best-arm identification (BAI) are two fundamental objectives in multi-armed bandits. Among regret-minimizing algorithms, $1/2$-Tsallis-INF is…

Source: arXiv cs.LG Jingxin Zhan, Yuze Han, Zhihua Zhang
AI Research AI

MDP Planning as Policy Inference

arXiv:2602.17375v3 Announce Type: replace Abstract: We formulate episodic Markov decision process (MDP) planning as Bayesian inference over policies. The primary contribution is conceptual: the policy itself is treated…

Source: arXiv cs.LG David Tolpin
AI Research AI

Projection-based multifidelity linear regression for data-scarce applications

arXiv:2508.08517v2 Announce Type: replace-cross Abstract: Surrogate modeling for systems with high-dimensional quantities of interest remains challenging, particularly when training data are costly to acquire. This work…

Source: arXiv cs.LG Vignesh Sella, Julie Pham, Karen Willcox, Anirban Chaudhuri
AI Research AI

PLeDO: Pain Level Detection for Osteoarthritis from EMR Data

arXiv:2608.15719v1 Announce Type: cross Abstract: Osteoarthritis (OA) is a progressive chronic joint disease resulting in a breakdown of articular cartilage and bone when damaged joint tissues are not able to normally…

Source: arXiv cs.LG Yuhao Chen, Jiahao Cai, Nafiz Sadman, Farhana Zulkernine, John Queenan, David Barber
AI Research AI

LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks

arXiv:2608.14927v1 Announce Type: cross Abstract: Multi-agent large language model (LLM) systems can improve reasoning by spending more computation, but deployment requires deciding when extra collaboration is worth its…

Source: arXiv cs.LG Chih-Hsuan Yang, Jingyan Jiang, Cheng-Hau Yang, Vikram Vasudevan, Huihuo Zheng, Venkatram Vishwanath, Rajeev Thakur
AI Research AI

SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation

arXiv:2608.04419v2 Announce Type: replace Abstract: On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can assign insufficient probability…

Source: arXiv cs.LG Zikun Qu, Min Zhang, Mingze Kong, Zhiwei Shang, Zhengyu Chen, Yikun Ban, Shuang Qiu, Zhongxiang Dai
AI Research AI

Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values

arXiv:2607.14345v4 Announce Type: replace Abstract: People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage: the information they…

Source: arXiv cs.LG Jan Betley, Johannes Treutlein, Jan Dubi\'nski, Harry Mayne, Karol Ga{\l}\k{a}zka, Niels Warncke, Anna Sztyber-Betley, Owain Evans
AI Research AI

The canonical facets of multi-separator polytopes

arXiv:2608.16861v1 Announce Type: cross Abstract: We initiate a polyhedral study of the graph multi-separator problem proposed by Irmai et al. (2024) as an alternative to the lifted multicut problem for application to…

Source: arXiv cs.LG Bjoern Andres, Silvia Di Gregorio, Jannik Irmai, Lucas Fabian Naumann, Shengxian Zhao
AI Research AI

Experimentally Extending Quantum Kernel Learning to Quantum Data by NMR

arXiv:2412.09557v3 Announce Type: replace-cross Abstract: Quantum kernel learning (QKL) promises efficient machine learning by encoding feature maps onto exponentially large Hilbert spaces inherent in quantum systems.…

Source: arXiv cs.LG Vivek Sabarad, Vishal Varma, T. S. Mahesh
AI Research AI

Online Convex Optimization with Dueling Feedback

arXiv:2608.15050v1 Announce Type: new Abstract: We study online convex optimization with dueling (pairwise comparison) feedback, where the learner observes only a binary preference between two queried points. While…

Source: arXiv cs.LG Yiyang Lu, Hareshkumar Jadav, Mohammad Pedramfar, Ranveer Singh, Vaneet Aggarwal
AI Research AI

From Monte Carlo to neural networks approximations of boundary value problems

arXiv:2209.01432v4 Announce Type: replace-cross Abstract: In this paper we study probabilistic and neural network approximations for solutions to Poisson equation subject to Holder data in general bounded domains of…

Source: arXiv cs.LG Lucian Beznea, Iulian Cimpean, Oana Lupascu-Stamate, Ionel Popescu, Arghir Zarnescu
AI Research AI

A Novel Fourier Feature Network for Solving Partial Differential Equations

arXiv:2608.14733v1 Announce Type: new Abstract: Building on the foundation of single-hidden-layer neural networks, Fourier Feature Networks (FENs) are proposed, which incorporate Fourier features using $\cos$, $\sin$,…

Source: arXiv cs.LG Qihong Yang, Zhijie Su, Yangtao Deng, Qiaolin He
AI Research AI

Local Gains and Fixed-Assignment Set Losses in Shared Set Decoders

arXiv:2608.14717v1 Announce Type: cross Abstract: A query-relation deletion can improve the edited slot while reducing the utility of the prediction set that contains it. We study this tension in two related ResNet-50…

Source: arXiv cs.LG Ze Zhang, Yang Zhang
AI Research AI

Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer

arXiv:2608.15690v1 Announce Type: cross Abstract: Text-to-audio-video (T2AV) generation models produce a video and its soundtrack from a textual description, but offer no control over whose voice speaks in the output.…

Source: arXiv cs.LG Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko, Alexey Letunovskiy, Ivan Kirillov, Kirill Chernyshev, Denis Dimitrov
AI Research AI

GATTA: Graph Active Learning with Test-Time Augmentation

arXiv:2608.15084v1 Announce Type: new Abstract: Test-time augmentation (TTA) has proven effective for improving model robustness and uncertainty estimation in computer vision, yet its application to graph-structured…

Source: arXiv cs.LG Zsombor B\'anfi, Andr\'as G\'ezsi, Andr\'as Formanek
AI Research AI

Belayer: Efficient Fault Tolerance for LLM Agentic RL Training

arXiv:2608.14635v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly trained with reinforcement learning in long-horizon, sandboxed environments. Unlike conventional RL, agentic RL…

Source: arXiv cs.LG Jiecheng Zhou, Qinghao Hu, Peng Sun, Xingcheng Zhang, Weiming Zhang
AI Research AI

Generative Model Unlearning: A Survey through Target Events, Unlearning Operators, and Evaluation Protocols

arXiv:2507.19894v2 Announce Type: replace Abstract: With the rapid advancement of generative models, privacy, copyright, safety, and reliability risks have attracted growing attention. To mitigate these risks, machine…

Source: arXiv cs.LG Xiaohua Feng, Jiaming Zhang, Fengyuan Yu, Chengye Wang, Li Zhang, Kaixiang Li, Yuyuan Li, Lingjuan Lyu, Chaochao Chen, Jianwei Yin
AI Research AI

Generalized Linear Bandits with Memory

arXiv:2608.15848v1 Announce Type: cross Abstract: We study generalized linear bandits with memory, an endogenous non-stationary setting in which rewards depend on past actions through a finite memory matrix. Building on…

Source: arXiv cs.LG Heesang Ann, Hyunjun Choi, Taehyun Hwang, Younghoon Shin, Haeju Cheong, Min-hwan Oh
AI Research AI

Operator-Theoretic Generalization Bounds for Multitask Deep Learning

arXiv:2608.15982v1 Announce Type: new Abstract: We develop operator-theoretic generalization bounds for deep multi-output function classes by representing network layers as Koopman composition operators on vector-valued…

Source: arXiv cs.LG Mahdi Mohammadigohari, Thomas Borsani, Giuseppe Di Fatta
AI Research AI

Layers Matter: Why Continual Learning Regularization Should Be Layer-Adaptive

arXiv:2608.15901v1 Announce Type: new Abstract: Continual learning regularizers like EWC fight forgetting by penalizing changes from previous-task parameters with per-parameter importance, typically diagonal Fisher…

Source: arXiv cs.LG Brian B. Moser, Ahmed Anwar, Tobias Christian Nauen, Shishir Muralidhara, Federico Raue, Ren\'e Schuster, Stanislav Frolov, Andreas Dengel
AI Research AI

Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking

arXiv:2608.14808v1 Announce Type: cross Abstract: When a user question is underspecified, a capable model should recognize that its context is insufficient, identify the missing information, ask for it, and respond only…

Source: arXiv cs.LG Yepeng Huang, Jiawen Zhang, Michelle Dai, Xiaorui Su, Shanghua Gao, Zi Wang, Marinka Zitnik
AI Research AI

Geometry of Forgetting: Representation Flux in Continual Learning

arXiv:2608.15854v1 Announce Type: new Abstract: Catastrophic forgetting remains a fundamental obstacle to continual learning, where neural networks lose previously acquired knowledge while learning new tasks. Existing…

Source: arXiv cs.LG Maksim A. Kazanskii
AI Research AI

Efficient Coreset Selection via K-Nearest Neighbor Graphs

arXiv:2608.16270v1 Announce Type: new Abstract: Coreset selection reduces the cost of model training by replacing a large training set with a small representative subset. Existing gradient-approximation coreset methods…

Source: arXiv cs.LG Yingfan Liu, Leiyu Zhang, Jiadong Xie, Mingzhe Wang, Jeffrey Xu Yu, Jiangtao Cui
AI Research AI

Machine Learning Approaches to Decoding Topological Quantum Codes

arXiv:2608.15760v1 Announce Type: cross Abstract: Decoding is an essential component of quantum error correction (QEC), translating stabilizer measurement outcomes into corrective actions that suppress logical errors…

Source: arXiv cs.LG Changwon Lee, Tak Hur, Jeongwoo Jae, Daniel K. Park
AI Research AI

Feature-Aware (Hyper)graph Generation via Next-Scale Prediction

arXiv:2506.01467v4 Announce Type: replace Abstract: Graph generative models perform well on small-scale structured data but struggle to scale to large, complex structures. Hierarchical approaches improve scalability but…

Source: arXiv cs.LG Dorian Gailhard, Enzo Tartaglione, Lirida Naviner, Jhony H. Giraldo
AI Research AI

PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data

arXiv:2608.16419v1 Announce Type: new Abstract: Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning traces. Here we show that cellular…

Source: arXiv cs.LG Zhenchao Tang, Xiaogang Xu, Tianxu Lv, Jiahui Guan, Jiale Zhou, Haohuai He, Zhi Song, Hanbo Huang, Jiehui Huang, Jiafei Wu, Zhe Liu
AI Research AI

Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation

arXiv:2608.15949v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled their use as conversational recommender systems (CRS), demonstrating strong recommendation accuracy and…

Source: arXiv cs.LG Cedar Site Bai, Duanshun Li, Zhenyu Liao, Sheikh Sarwar, Huiyuan Chen, Yuan Chen, Changhe Yuan, Haiyang Zhang, Qilin Qi
AI Research AI

DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on $60 GPUs

arXiv:2608.14614v1 Announce Type: new Abstract: As AI datacenters retire functional GPUs, vast quantities of still capable accelerators enter secondary markets. This paper investigates whether these retired GPUs can…

Source: arXiv cs.LG Zeyu Cao, Xuan Guo, Cheng Zhang, Cheuk Hang Lau, Ilia Shumailov, Yiren Zhao
AI Research AI

Convolution Smoothed Quantile Regression for XGBoost

arXiv:2608.15290v1 Announce Type: cross Abstract: The increasing availability of large and complex datasets across many scientific disciplines has led to widespread adoption of machine learning (ML) for prediction.…

Source: arXiv cs.LG Mandy Yao (University of Toronto), Meredith Franklin (University of Toronto)
AI Research AI

Convex Optimization with Nested Evolving Feasible Sets

arXiv:2605.07386v2 Announce Type: replace Abstract: \emph{Convex Optimization with Nested Evolving Feasible Sets (CONES)} is considered where the objective function \(f\) remains fixed but the feasible region evolves…

Source: arXiv cs.LG Karthick Krishna M., Haricharan Balasundaram, Rahul Vaze
AI Research AI

On Stopping Rules and Spatial Adaptation for CART

arXiv:2608.15649v1 Announce Type: cross Abstract: The popular CART algorithm for regression trees combines a greedy splitting rule with a stopping rule, but while the splitting rule has been well studied, the…

Source: arXiv cs.LG Zineng Xu, Yuchao Cai, Yan Shuo Tan
AI Research AI

The Optimal Sample Complexity of Multiclass and List Learning

arXiv:2604.24749v3 Announce Type: replace Abstract: While the optimal sample complexity of binary classification in terms of the VC dimension is well-established, determining the optimal sample complexity of multiclass…

Source: arXiv cs.LG Chirag Pabbaraju
AI Research AI

On the Principles Behind Neural Network Optimizers

arXiv:2608.16760v1 Announce Type: new Abstract: Reliable optimization is central to neural network (NN) training, yet Adam, the default optimizer for modern LLMs, rests on a fragile foundation. This thesis develops a…

Source: arXiv cs.LG Yushun Zhang
AI Research AI

LiD-GLM: Lipschitz-constrained Deep Generalized Linear Models

arXiv:2608.16340v1 Announce Type: cross Abstract: The combination of traditional statistical models and neural network (NN) components into semi-structured hybrid models is an intriguing approach to construct models…

Source: arXiv cs.LG Tom Splittgerber, Niklas Koenen, Marvin N. Wright, Werner Brannath
AI Research AI

Stitch the Fragments: One-Shot Hierarchical Federated Clustering

arXiv:2601.06404v2 Announce Type: replace Abstract: Federated Clustering (FC) faces a critical bottleneck in real-world scenarios, i.e., global clusters are rarely intact, often fragmenting into incomplete,…

Source: arXiv cs.LG Shenghong Cai, Zihua Yang, Yang Lu, Mengke Li, Yuzhu Ji, Yiqun Zhang, Yiu-Ming Cheung
AI Research AI

A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing

arXiv:2607.16183v2 Announce Type: replace Abstract: To help address the escalating energy and latency demands of machine-learning workloads, we introduce a blueprint for an energy-efficient and fast thermodynamic…

Source: arXiv cs.LG Owen Lockwood, J\'er\'emy B\'ejanin, Joost Bus, Christopher Chamberland, Patrick Huembeli, Frank Sch\"afer, Guillaume Verdon
AI Research AI

VirnyFlow: Optimizing ML Pipelines for Accuracy, Fairness, and Stability at Scale

arXiv:2506.01584v2 Announce Type: replace Abstract: Developing machine learning (ML) systems for real-world deployment requires navigating context-dependent trade-offs among accuracy, fairness, stability, and other…

Source: arXiv cs.LG Denys Herasymuk, Anastasiia Mozghova, Nazar Protsiv, Vladyslav Sydorak, Julia Stoyanovich
AI Research AI

Macroeconomic Forecasting with Large Language Models

arXiv:2407.00890v5 Announce Type: replace-cross Abstract: This paper presents a comparative analysis evaluating the accuracy of Large Language Models (LLMs) against traditional macro time series forecasting approaches.…

Source: arXiv cs.LG Andrea Carriero, Davide Pettenuzzo, Shubhranshu Shekhar
AI Research AI

Sufficient Dimesion Reduction via Generalized Stein's Lemma

arXiv:2608.15121v1 Announce Type: cross Abstract: Sufficient dimension reduction (SDR) seeks the minimal subspace of the predictors that captures the full conditional distribution of the response, which is known as the…

Source: arXiv cs.LG Ye Tian
AI Research AI

Zero-Shot Instruction Following in RL via Structured LTL Representations

arXiv:2602.14344v2 Announce Type: replace Abstract: We study instruction following in multi-task reinforcement learning, where an agent must zero-shot execute novel tasks not seen during training. In this setting,…

Source: arXiv cs.LG Mathias Jackermeier, Mattia Giuri, Jacques Cloete, Alessandro Abate
AI Research AI

Learning Multi-Timescale Interventions under Safety and Resource Constraints

arXiv:2508.03875v4 Announce Type: replace Abstract: Many sequential decision problems offer qualitatively different ways of influencing the environment: some interventions act immediately, whereas others induce…

Source: arXiv cs.LG David Mguni, Wanrong Yang, Jing Dong, Jing Peng, Ziquan Liu, Muhammad Salman Haleem, Baoxiang Wang, Dominik Wojtczak
AI Research AI

Evolving Executable Pipeline Programs for AutoML with Language Models

arXiv:2608.16416v1 Announce Type: new Abstract: Automated machine learning (AutoML) systems search for pipelines within a space of preprocessing operators, learners, and hyper-parameters specified in advance: they can…

Source: arXiv cs.LG Sofoklis Kitharidis, Cor J. Veenman, Jan N. van Rijn, Thomas B\"ack, Niki van Stein
AI Research AI

FluxBin: Flexible LUT-based Ultra-low-bit LLM Inference by Algorithm-Kernel Synergy

arXiv:2608.15602v1 Announce Type: new Abstract: While binary quantization theoretically promises extreme compression and acceleration for Large Language Models (LLMs), existing research often overlooks the necessity of…

Source: arXiv cs.LG Qingyao Yang, Runming Yang, He Xiao, Wendong Xu, Junyu Chen, Haobo Liu, Chenchen Ding, Ruihan Hu, Yik-Chung Wu, Ngai Wong
AI Research AI

Hide&Seek: Learning to Explain in an End-to-End Differentiable Network

arXiv:2608.16689v1 Announce Type: cross Abstract: Instance-wise feature selection is a valuable tool for interpreting labeled data and the predictions of black-box models. In contrast to global feature selection…

Source: arXiv cs.LG Tal Ellinson, Hadi Mohasel Afshar, Sally Cripps
AI Research AI

GEO-Flag: Detecting and Measuring GEO-Optimized Web Content

arXiv:2608.16824v1 Announce Type: new Abstract: Generative Engine Optimization (GEO) modifies web content to increase its likelihood of being selected and cited by generative search engines. This can give strategically…

Source: arXiv cs.LG Junjie Chu, Ye Leng, Mingjie Li, Yun Shen, Xinyue Shen, Yang Zhang
AI Research AI

Learning to Price with Persuasion

arXiv:2608.16699v1 Announce Type: cross Abstract: Motivated by modern marketplaces, where the platform or the seller routinely gathers detailed user profiles, we study a novel learning theoretic model that…

Source: arXiv cs.LG Maria-Florina Balcan, Tejas Pagare, Karan Singh
AI Research AI

Annealed Softmax Greedy in Many-Armed Bayesian Bandits

arXiv:2605.31034v3 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards and group-based policy optimization methods update a stochastic policy by sampling multiple completions per prompt and…

Source: arXiv cs.LG William Overman, Mohsen Bayati
AI Research AI

p-Spin Glass Network Efficient Single-Batch Continual Learning

arXiv:2608.14774v1 Announce Type: new Abstract: Modern sequence models heavily rely on massive memory footprints and large-batch stochastic optimization, barriers that restrict sample efficiency and continual learning.…

Source: arXiv cs.LG Vladimer Khasia
AI Research AI

In-Context Source and Channel Coding

arXiv:2601.10267v2 Announce Type: replace Abstract: Separate Source-Channel Coding (SSCC) remains attractive for text transmission due to its modularity and compatibility with mature entropy coders and powerful channel…

Source: arXiv cs.LG Ziqiong Wang, Tianqi Ren, Rongpeng Li, Zhifeng Zhao, Honggang Zhang
AI Research AI

Periodic Topological Deep Learning for Polymer Design and Discovery

arXiv:2605.26833v2 Announce Type: replace Abstract: Polymers underpin applications across energy, healthcare, and materials science, yet their vast chemical space makes systematic discovery challenging. Most machine…

Source: arXiv cs.LG Yasharth Yadav, Tze Kwang Gerald Er, Atsushi Goto, Kelin Xia
AI Research AI

FedADB: Class Anchor-Driven Dual-Branch Federated Learning for Mitigating Forgetting

arXiv:2608.15310v1 Announce Type: cross Abstract: Multimodal data collected by heterogeneous devices are used for collaborative training, where federated learning (FL) serves as a key paradigm for effective distributed…

Source: arXiv cs.LG Zhenyan Liu, Hua Zhang, Haoran Gao, Qi Li, Hongliang Zhu, Huiyu Zhou, Zongliang Shen, Yanxin Xu, Jiahui Wang
AI Research AI

One-shot Robust Federated Learning of Independent Component Analysis

arXiv:2505.20532v2 Announce Type: replace Abstract: This paper studies robust one-shot aggregation for distributed and federated Independent Component Analysis (ICA). In this setting, each client computes a local ICA…

Source: arXiv cs.LG Dian Jin, Xin Bing, Yuqian Zhang
AI Research AI

Do Language Models Consistently Encode the Current Year?

arXiv:2608.15507v1 Announce Type: cross Abstract: A consistent concept of the current time is important for temporal reasoning, yet how language models represent the current time is not well understood. We contribute…

Source: arXiv cs.LG Suze van Adrichem, Aditi Bhaskar, Diyi Yang, Christopher Potts, Jing Huang
AI Research AI

PathFinder: Joint Decompositions of Linked Multimodal Datasets

arXiv:2608.14951v1 Announce Type: new Abstract: Low-rank matrix decompositions can uncover patterns and structure in data and have a number of different applications across many disciplines. Extensions to "joint"…

Source: arXiv cs.LG Ying-Qiu Zheng, Alex Fung, Stephen M Smith, Rogier B Mars, Saad Jbabdi
AI Research AI

Proteus: Incremental Memory Activation for Long-Context Sequence Modeling

arXiv:2608.16844v1 Announce Type: new Abstract: The quadratic cost of attention-based sequence models for long contexts has motivated a growing line of research on memory-based models that can compress context into a…

Source: arXiv cs.LG Reza Bayat, Ali Behrouz, Vahab Mirrokni, Aaron Courville
AI Research AI

Language models suffer from a curse of ambiguity

arXiv:2608.15448v1 Announce Type: cross Abstract: Large language models increasingly rely on sampling as a driver of their own improvement, making the fidelity of their learned distributions more critical than ever.…

Source: arXiv cs.LG Nicolas Zucchet, Hyun Dong Lee, Scott Linderman
AI Research AI

SubZero+: Efficient Zeroth-Order LLM Fine-Tuning via Large Learning Rates

arXiv:2608.15665v1 Announce Type: new Abstract: Zeroth-order (ZO) optimization enables backpropagation-free fine-tuning of large language models, but existing ZO methods suffer from high-variance gradient estimators,…

Source: arXiv cs.LG Ziming Yu, Shuyao Xiao, Xingyu Zhao, Sike Wang, Pan Zhou, Peiyu Zang, Xiangda Yan, Yongjie Yang, Jia Li
AI Research AI

Generalised Transportability via Causal Abstractions

arXiv:2608.15645v1 Announce Type: new Abstract: Transporting a causal conclusion from a source study population to a target one is a fundamental problem in causal inference. The theory of transportability provides a…

Source: arXiv cs.LG Yorgos Felekis, Paris Giampouras, Fabio Massimo Zennaro, Theodoros Damoulas
AI Research AI

A Generative Deep Learning Workflow for Inverse Molecular Design of Fuels

arXiv:2504.12075v4 Announce Type: replace Abstract: In the present work, a generative deep learning framework combining a Co-optimized Variational Autoencoder (Co-VAE) with quantitative structure-property relationship…

Source: arXiv cs.LG Kiran K. Yalamanchi, Pinaki Pal, Balaji Mohan, Abdullah S. AlRamadan, Jihad A. Badra, Yuanjiang Pei
AI Research AI

ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow

arXiv:2607.27924v3 Announce Type: replace Abstract: In the physical world we inhabit, space and time are fundamentally continuous. However, existing machine learning paradigms for world modeling are largely confined to…

Source: arXiv cs.LG Dongxiu Liu, Haoyi Niu, Peng Cheng, Yuan Gao, Xirui Kang, Sangli Teng, Koushil Sreenath, Xianyuan Zhan
AI Research AI

The Limits of Binding in Dual Encoders

arXiv:2608.15971v1 Announce Type: new Abstract: Dual-encoder models such as CLIP score an image-caption pair by a single inner product of two independently computed unit vectors, and fail at binding, often scoring near…

Source: arXiv cs.LG Kin Ian Lo
AI Research AI

NRCD: An Open Database of Collegiate Running with Unified Performance Standardization

arXiv:2608.14776v1 Announce Type: new Abstract: Collegiate running in the United States generates thousands of race results annually in cross country and track and field, yet no large-scale dataset has been publicly…

Source: arXiv cs.LG Jonathan A. Karr Jr., Ryan M. Fryer, Ben Darden, Nicholas Pell, Kayla Ambrose, Evan Hall, Ramzi K. Bualuan, Nitesh V. Chawla
AI Research AI

RouteTS: Frequency-Time Routing for Time Series Forecasting

arXiv:2608.14682v1 Announce Type: new Abstract: Real-world time series inherently intertwine global periodic structures with localized non-stationary variations. Existing approaches process these heterogeneous dynamics…

Source: arXiv cs.LG Gaofeng Lin, Lei Duan
AI Research AI

Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic

arXiv:2608.16273v1 Announce Type: new Abstract: Foresight-England (Foresight-E) is the first national-scale generative foundation model of electronic health records (EHRs), developed as a research pilot strictly for…

Source: arXiv cs.LG Simon Ellershaw, Christopher Tomlinson, Zeljko Kraljevic, Spiros Denaxas, Harry Hemingway, Cathie Sudlow, Angela M. Wood, Anoop D. Shah, Richard Dobson
AI Research AI

Adaptive Optimization via Momentum on Variance-Normalized Gradients

arXiv:2602.10204v2 Announce Type: replace Abstract: We introduce MVN-Grad (Momentum on Variance-Normalized Gradients), an Adam-style optimizer that improves stability and performance by combining two complementary…

Source: arXiv cs.LG Francisco Patitucci, Aryan Mokhtari
AI Research AI

INSPIRE: A Benchmark for Instruction-Aware Speech Retrieval

arXiv:2608.16203v1 Announce Type: cross Abstract: Existing speech retrieval systems rely on fixed similarity matching and cannot adapt to diverse user intents. We introduce INSPIRE, the first benchmark for…

Source: arXiv cs.CL Chen-An Li, Hung-yi Lee
AI Research AI

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models

arXiv:2606.05177v2 Announce Type: replace Abstract: Existing multimodal safety benchmarks focus solely on visual inputs and cannot assess Omni Large Language Models (LLMs) that process vision, audio, and text. We…

Source: arXiv cs.CL Manh Luong, Tamas Abraham, Junae Kim, Amar Kaur, Rollin Omari, Gholamreza Haffari, Trang Vu, Lizhen Qu, Dinh Phung
AI Research AI

Personalized Auto-Research: Towards a True AI Co-Scientist

arXiv:2608.14881v1 Announce Type: cross Abstract: AI co-scientists that generate hypotheses, retrieve related work, design experiments, execute code, and draft full papers are beginning to change how research is carried…

Source: arXiv cs.CL Bo Ni, Franck Dernoncourt, Hongjie Chen, Yu Wang, Nesreen K. Ahmed, Zhengzhong Tu, Tyler Derr, Ryan A. Rossi
AI Research AI

DA-RAC: Distance-Aware Calibration of LLM Judges for Trustworthy AI Auditing

arXiv:2608.14950v1 Announce Type: new Abstract: Generative AI systems are increasingly producing real-world artifacts, however their efficacy and validity are often evaluated via context-free LLM-scoring. These judges…

Source: arXiv cs.CL Cheng Wu, Vishal Anand, Jaya Krishna Mandivarapu, Xiya Liu, Rui Zhuang
AI Research AI

RamseyGadgets: A Graph Construction Dataset for LLMs

arXiv:2608.14999v1 Announce Type: new Abstract: Constructing special graphs is an important task within graph theory and computer science. Many popular graph constructions are the result of a comprehensive exploration…

Source: arXiv cs.CL Zohair Raza Hassan, Deepak Pandita
AI Research AI

Discovering Conceptual Metaphors Across Topics and Media Types

arXiv:2608.06652v2 Announce Type: replace Abstract: Conceptual metaphors guide our thinking and actions by allowing us to reason about more abstract experiences (e.g., paying taxes) in terms of more concrete or embodied…

Source: arXiv cs.CL Alexandria Leto, Rohan Das, Juan V\'asquez, Abram Handler, Maria Leonor Pacheco
AI Research AI

HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam

arXiv:2602.13964v4 Announce Type: replace Abstract: Humanity's Last Exam (HLE) has become a widely used benchmark for evaluating frontier large language models on challenging, multi-domain questions. However,…

Source: arXiv cs.CL Weiqi Zhai, Zhihai Wang, Jinghang Wang, Boyu Yang, Xiaogang Li, Xander Xu, Bohan Wang, Peng Wang, Xingzhe Wu, Anfeng Li, Qiyuan Feng, Yuhao Zhou, Taolin Han, Wenjie Luo, Yiyuan Li, Xiang Zheng, Yaxua…
AI Research AI

Hypergraph as Language

arXiv:2605.21858v2 Announce Type: replace Abstract: Large language models (LLMs) have recently shown strong potential in modeling relational structures. However, existing approaches remain fundamentally graph-centric:…

Source: arXiv cs.CL Mengqi Lei, Guohuan Xie, Shihui Ying, Shaoyi Du, Jun-Hai Yong, Chuan Shi, Ling Tian, Siqi Li, Yue Gao
AI Research AI

Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents

arXiv:2608.15071v1 Announce Type: cross Abstract: Learning from experience is critical for developing capable, self-improving large language model (LLM) agents. Existing methods typically extract knowledge from…

Source: arXiv cs.CL Tianxin Wei, Zhan Shi, Minhua Lin, Bing He, Zewen Liu, Yisi Sang, Yuanchen Bei, Xuying Ning, Jiaru Zou, Ting-Wei Li, Xiao Lin, Yanjun Zhao, Chi Wang, Benoit Dumoulin, Dakuo Wang, Jingrui He, Hanqing…
AI Research AI

VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding

arXiv:2608.14718v1 Announce Type: cross Abstract: Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading models have already…

Source: arXiv cs.CL Fan Zhang, Guangming Yao, Jinyang Wu, Hao Wu, Zheng Lian, Xinyu Geng, Jingdong Chen, Yi Yuan, Pheng-Ann Heng
AI Research AI

CulTrace: Tracing Internal Cultural Reasoning in Large Language Models

arXiv:2508.08879v3 Announce Type: replace Abstract: The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of models' hidden representations of…

Source: arXiv cs.CL Haeun Yu, Arnav Arora Seogyeong Jeong, Nadav Borenstein, Siddhesh Pawar, Jisu Shin, Jiho Jin, Junho Myung, Alice Oh, Isabelle Augenstein
AI Research AI

Efficient Code Embeddings from Code Generation Models

arXiv:2508.21290v2 Announce Type: replace Abstract: jina-code-embeddings is a novel code embedding model suite designed to retrieve code from natural language queries, perform technical question-answering, and identify…

Source: arXiv cs.CL Daria Kryvosheieva, Saba Sturua, Michael G\"unther, Han Xiao
AI Research AI

MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations

arXiv:2608.15844v1 Announce Type: new Abstract: Long-horizon, multi-agent language model (LM) simulations are widely proposed for studying social behavior, yet instruments to measure whether persona-conditioned agents…

Source: arXiv cs.CL Sky Ng (Eliza), Brihi Joshi (Eliza), Ishan Gupta (Eliza), Shirley Huang (Eliza), Zonglin Di (Eliza), Yun Shen (Eliza), Qianfeng Wen (Eliza), Yifan Simon Liu (Eliza), Ruoqi Gao (Eliza), Yilan (Eliza)…
AI Research AI

Schema-Agnostic Graph Reasoning Agent for Hybrid Knowledge Graphs

arXiv:2608.15834v1 Announce Type: cross Abstract: Tool-calling LLM agents navigate unfamiliar codebases with a handful of generic primitives for listing, reading and searching files (ls, cat, grep). A knowledge graph…

Source: arXiv cs.CL Marius Dragic, Ruben Ifrah, Alexandre Rio
AI Research AI

Do Value Vectors in Deep Layers Need Context from the Residual Stream?

arXiv:2606.02780v4 Announce Type: replace Abstract: The success of the transformer architecture as the backbone of modern LLMs is in large part due to its use of attention layers. An attention layer follows the standard…

Source: arXiv cs.CL Muyu He, Yuchen Liu, Qingya Huang, Li Zhang
AI Research AI

Neurosymbolic Embodied Agents

arXiv:2608.16794v1 Announce Type: cross Abstract: Language and vision-language models generate plausible embodied plans but do not guarantee executability, as their outputs can violate environment dynamics or act on…

Source: arXiv cs.CL Mohammad Albinhassan, Yuming Feng, Alessandra Russo, Pranava Madhyastha
AI Research AI

LLMs Get Smarter from Targeted Synthetic Multilingual Data

arXiv:2608.15964v1 Announce Type: new Abstract: Language-specific competency (LSC) is the phenomenon of a language model performing better or worse depending on the language of the prompt. In other words, a language…

Source: arXiv cs.CL Ishika Agarwal, Arkajyoti Charaborty, Tanner Sorensen, Neha Gupta, Andreas Stolcke
AI Research AI

A Pilot Study of Autocompleting Tokenizers

arXiv:2608.15080v1 Announce Type: new Abstract: Modern input methods routinely rely on autocomplete to omit information that can be recovered from local context. Inspired by these autocomplete-assisted writing systems,…

Source: arXiv cs.CL Samuel Wexler, Mark Hopkins
AI Research AI

Douyin Multimodal Embedding Model Technical Report

arXiv:2608.02148v2 Announce Type: replace-cross Abstract: Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vectors, it powers industrial search and…

Source: arXiv cs.CL Haonan Chen, Chu Li, Zhicheng Wang, Yuanwei Liu, Yuanjiang Wang, Shaohua Jiang, Zhicheng Dou
AI Research AI

PLSQLBench: Benchmarking LLM Systems for Executable Procedural Database Programming

arXiv:2608.15931v1 Announce Type: new Abstract: We present PLSQLBench, to our knowledge the first benchmark for evaluating whether LLMs can write executable PL/SQL programs, with correctness measured through…

Source: arXiv cs.CL Marianne Menglin Liu, Leonid Boytsov, Daniel W. Peterson, Pramuditha Perera, Rongguang Wang, Sai Ashish Somayajula, Syed Hamza Rafique, Rohit Saini, Shubham Pathak, Sujeeth Bharadwaj, Tao Sheng, Grah…
AI Research AI

A Declarative-Procedural Perspective on Expert Routing in Bilingual Mixture-of-Experts Language Models

arXiv:2608.15102v1 Announce Type: new Abstract: We investigate whether Mixture-of-Experts (MoE) language models develop linguistically structured expert routing during bilingual language acquisition. Inspired by the…

Source: arXiv cs.CL Amrit Gopinath (Sri Sivasubramaniya Nadar College of Engineering, Chennai, India), Raghul (Sri Sivasubramaniya Nadar College of Engineering, Chennai, India), Durairaj Thenmozhi (Shiv Nadar Universit…
AI Research AI

BabelSteering: Multilingual Safety Alignment via English Steering Vectors

arXiv:2608.16577v1 Announce Type: new Abstract: Large language models (LLMs) are deployed globally in high-stakes settings, yet most safety research and alignment efforts remain concentrated on English. Thus, users…

Source: arXiv cs.CL Emma V. Stein, Dominik Meier, Terry Ruas, Jan Philip Wahle, Bela Gipp
AI Research AI

GALA: Generation-Aware Cross-Modal Alignment for Text-to-Time-Series Synthesis

arXiv:2608.13741v2 Announce Type: replace Abstract: Synthesizing time series from natural language is emerging as the most expressive form of controllable time series generation. However, existing text-conditioned…

Source: arXiv cs.CL Haochen Zhang, Gengwei Zhang, Laura Yao, Nicholas Konz, Tianlong Chen
AI Research AI

STAGE: Controlled Objective Admission for Multi-Preference LLM Alignment

arXiv:2608.16553v1 Announce Type: new Abstract: Multi-preference alignment is often framed as scalarization: combine reward dimensions, then optimize. This leaves a temporal decision underspecified: when should each…

Source: arXiv cs.CL Yongqi Tong, Zhenyu Zhang, Ruirui Wang, Kewei Fu, Shaoqing Lin, Sijie Dong, Jiang-Ming Yang, Xin Zhang, Jianshe Li
AI Research AI

FollowUpBot: An LLM-Based Conversational Robot for Automatic Postoperative Follow-up

arXiv:2507.15502v1 Announce Type: cross Abstract: Postoperative follow-up plays a crucial role in monitoring recovery and identifying complications. However, traditional approaches, typically involving bedside…

Source: arXiv cs.CL Chen Chen, Jianing Yin, Jiannong Cao, Zhiyuan Wen, Mingjin Zhang, Weixun Gao, Xiang Wang, Haihua Shu
AI Research AI

BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences

arXiv:2607.27366v2 Announce Type: replace Abstract: While data synthesis for large language models (LLMs) is prevalent, it primarily targets domains with verifiable answers, overlooking open-ended humanities and social…

Source: arXiv cs.CL Ru Peng, Haokai Xu, Xijun Gu, Tianyu Zhao, Zhiting Fan, Yawen Zeng, Yihong Zhuang, Jinyang Zhang, Kexin Yang, Jian Wu, Hao Chen, Junyang Lin, Dayiheng Liu, Junbo Zhao
AI Research AI

Semantic Space of Parts of Speech

arXiv:2608.15443v1 Announce Type: new Abstract: Parts of speech categorization is understood in the European linguistic tradition as crisp categorization, which is also reflected in corpus linguistics, where each…

Source: arXiv cs.CL Ji\v{r}\'i Mili\v{c}ka, Ivan Kraus, Arnold Stanovsk\'y, Anna Vyslou\v{z}ilov\'a, Barbora \v{S}t\v{e}p\'ankov\'a, Lenka F\'arov\'a, Vojt\v{e}ch Cink, \v{S}\'arka Dohnalov\'a
AI Research AI

SEER: Long-Context Reasoning via Selective Visual-Text Compression

arXiv:2608.15962v1 Announce Type: new Abstract: Long-context reasoning remains computationally expensive for large language models due to the quadratic complexity of attention over text tokens. Visual-text compression…

Source: arXiv cs.CL Jiawei Xu, Zhilin Zhai, Jinrui Fang, Ruohan Xu, Mingfei Lu, Yi Zhang, Guanchu Wang, Tianlong Chen, Ying Ding
AI Research AI

Beyond Single Object: Learning 3D Relations with Large Language Models

arXiv:2608.15710v1 Announce Type: cross Abstract: We address a fundamental gap in 3D-LLMs: existing models focus on single-object/scene description, struggling with detailed, inter-object comparison. We propose a…

Source: arXiv cs.CL Kohsuke Ide, Ryousuke Yamada, Yue Qiu, Xianzheng Ma, Yoshihiro Fukuhara, Hirokatsu Kataoka, Yutaka Satoh
AI Research AI

Inference-Time Mitigation of Adversarial Political Bias in Large Language Models

arXiv:2608.14629v1 Announce Type: new Abstract: As Large Language Models (LLMs) become the mainstay for information retrieval and summarization tasks, ensuring that they are always non-partisan and invulnerable to…

Source: arXiv cs.CL Tejaswi V. Panchagnula, Bruce Coburn, Bryce J. Dietrich, Robert X. Browning, Edward J. Delp, Fengqing Zhu
AI Research AI

Large language model-assisted discovery of cohorts from scientific literature

arXiv:2608.15909v1 Announce Type: cross Abstract: Background: Planning multi-study analyses requires identifying cohorts with the relevant participants, phenotypes, and data modalities. This process commonly relies on…

Source: arXiv cs.CL Moritz Sturm, Lisa M. Berg, Inken Berg, Harishny Sarma, Jasmin Hartmann, Denissa Girschik, Gemma Roig, Christine M. Freitag, Andreas G. Chiocchetti
AI Research AI

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA

arXiv:2411.15041v2 Announce Type: replace-cross Abstract: Advanced Multimodal Large Language Models (MLLMs) struggle with recent Knowledge-based Visual Question Answering (VQA) tasks, such as INFOSEEK and…

Source: arXiv cs.CL Tao Zhang, Ziqi Zhang, Zongyang Ma, Yuxin Chen, Zhongang Qi, Chunfeng Yuan, Bing Li, Junfu Pu, Yuxuan Zhao, Zehua Xie, Jin Ma, Ying Shan, Weiming Hu
AI Research AI

MoRFI: Monotonic Sparse Autoencoder Feature Identification

arXiv:2604.26866v2 Announce Type: replace Abstract: Large language models (LLMs) acquire most of their factual knowledge during the pre-training stage, through next token prediction. Subsequent stages of post-training…

Source: arXiv cs.CL Dimitris Dimakopoulos, Shay B. Cohen, Ioannis Konstas
AI Research AI

BengaliMCQ: Automatic Generation and Answer Prediction of Academic Multiple-Choice Questions in a Low-Resource Language

arXiv:2608.15547v1 Announce Type: new Abstract: Traditional retrieval-augmented generation (RAG) frameworks process documents without attending to their hierarchical structure, leading to poor performance, especially in…

Source: arXiv cs.CL Abu Tarabin Surzo, A. K. M. Nihalul Kabir, Sm Azmain Faysal, Ariana Haque Ami, Lawrence Amlan Gomes, Farig Sadeque
AI Research AI

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

arXiv:2608.12253v2 Announce Type: replace Abstract: Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach…

Source: arXiv cs.CL Simon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou, Dilara Soylu, Sergey Levine, Christopher D. Manning, Weiyan Shi
AI Research AI

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers

arXiv:2607.23811v2 Announce Type: replace-cross Abstract: Siri Expressive Voices synthesize rich, configurable speech in real time and entirely on device, powered by AFM 3 Core Advanced, Apple's most powerful on-device…

Source: arXiv cs.CL Dongseong Hwang, Prasanth Yadla, Kaan Elgin, Shifas Padinjaru Veettil, Sivanand Achanta, Dipjyoti Paul, Ramya Rasipuram, Tyler Johnson, Emad Soroush, Chung-Cheng Chiu, Zhifeng Chen
AI Research AI

MobileMem: Learning from a Year of Mobile Experiences

arXiv:2608.13606v2 Announce Type: replace-cross Abstract: The next generation of AI agents is increasingly moving beyond systems that answer isolated questions toward persistent personal assistants that can understand,…

Source: arXiv cs.CL Xinle Deng, Yida Xue, Xiangyuan Ru, Yijun Chen, Buqiang Xu, Mingjun Mao, Xinjie Liu, Haoming Xu, Shuofei Qiao, Mengru Wang, Chen Jiang, Yuchen Eleanor Jiang, Lizhong Wang, Jason Wang, Li Zeng, Haofen…
AI Research AI

Vision Language Models Cannot Plan, but Can They Formalize?

arXiv:2509.21576v2 Announce Type: replace Abstract: The advancement of vision language models (VLMs) has empowered embodied agents to accomplish simple multimodal planning tasks, but not long-horizon ones requiring long…

Source: arXiv cs.CL Muyu He, Yuxi Zheng, Yuchen Liu, Zijian An, Bill Cai, Jiani Huang, Lifeng Zhou, Feng Liu, Ziyang Li, Li Zhang
AI Research AI

Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning

arXiv:2608.16554v1 Announce Type: new Abstract: Answer-only reinforcement learning (RL) trains reasoning models to solve fully specified problems, but many realistic queries omit a premise needed for a unique answer. In…

Source: arXiv cs.CL Yongqi Tong, Zhenyu Zhang, Zimi Liu, Kewei Fu, Mingli Song, Haofei Zhang, Junshao Zhang, Hong Zhu, Jiang-Ming Yang, Xin Zhang, Jianshe Li
AI Research AI

DSPrompt: Dynamic Soft Prompt Defense Against M-RAG Corruption

arXiv:2608.16536v1 Announce Type: cross Abstract: Multimodal Retrieval Augmented Generation (M-RAG) is increasingly vulnerable to adversarial attacks where malicious data are crafted to produce embeddings that align…

Source: arXiv cs.CL Chang Liu, Yuni Lai, Mingyue Cui, Cong Tian, Yunyan Zhang, Xian Wu, Kai Zhou, Bin Xiao
AI Research AI

Language Models that Think, Chat Better

arXiv:2509.20357v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) trains language models to use long chain-of-thought reasoning (CoT) in domains like mathematics and code with…

Source: arXiv cs.CL Adithya Bhaskar, Xi Ye, Danqi Chen
AI Research AI

PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows

arXiv:2607.06008v3 Announce Type: replace-cross Abstract: While Large Language Model (LLM) agents excel at monolingual long-horizon planning and tool use, enterprise workflows inherently require processing multilingual…

Source: arXiv cs.CL Hongliang Li, Yijin Liu, Zhiwei Zhang, Zihe Liu, Xinyue Lou, Jinan Xu, Fandong Meng, Kaiyu Huang
AI Research AI

TaoLive Digital Avatar Agent Technical Report: Training Agents to Evolve with Their Harness

arXiv:2608.15763v1 Announce Type: new Abstract: AI-powered digital-avatar streamers in live e-commerce must answer product questions, engage viewers, and execute changing business strategies in real time. This requires…

Source: arXiv cs.CL TaoLive AIGC LLM Team, Yuhan Sun, Wenhao Lin, Yongdong Luo, Yibo Hu, Meiguang Jin, Junfeng Ma, Weihang Pan, Jiaxin Zhao, Zulong Chen
AI Research AI

Reasoning-Based Personalized Generation for Users with Sparse Data

arXiv:2602.21219v2 Announce Type: replace Abstract: Large Language Model (LLM) personalization holds great promise for tailoring responses by leveraging personal context and history. However, real-world users usually…

Source: arXiv cs.CL Bo Ni, Branislav Kveton, Samyadeep Basu, Subhojyoti Mukherjee, Leyao Wang, Franck Dernoncourt, Sungchul Kim, Seunghyun Yoon, Zichao Wang, Ruiyi Zhang, Puneet Mathur, Jihyung Kil, Jiuxiang Gu, Nedim L…
AI Research AI

When Do Concepts Become Functionally Sufficient During Language-Model Training?

arXiv:2608.15323v1 Announce Type: new Abstract: Understanding a model and its learning mechanisms in depth requires identifying when its internal structures become useful, rather than simply looking at the final state.…

Source: arXiv cs.CL Raphael Bernas, Paul G. Chevalier, Fanny Jourdan, C\'eline Hudelot
AI Research AI

From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents

arXiv:2608.16002v1 Announce Type: new Abstract: Reliable uncertainty quantification (UQ) is essential for deploying large language model (LLM) agents in complex interactive environments. Existing UQ methods largely rely…

Source: arXiv cs.CL Zhengzhao Ma. Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun
AI Research AI

D2-ScaleAgent: Dual-Dimensional Scaling for Long Document Understanding

arXiv:2608.16417v1 Announce Type: new Abstract: Multi-modal retrieval-augmented generation (RAG) is a key technique for visually rich long document understanding. Existing multi-modal RAG methods are progressively…

Source: arXiv cs.CL Hao Zhang, Longrong Yang, Lunhao Duan, Ziyang Wang, Qing-Guo Chen, Shanshan Zhao
AI Research AI

Honeyquest for LLMs: Rethinking Cyber Deception for AI Attackers

arXiv:2606.21037v2 Announce Type: replace-cross Abstract: The empirical foundation of cyber deception relies on human-centered hypotheses, but the rapid emergence of autonomous, AI-enabled attackers challenges whether…

Source: arXiv cs.CL Kerri Prinos, Lilianne Brush, Cameron Denton
AI Research AI

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

arXiv:2505.14107v5 Announce Type: replace Abstract: The emergence of groundbreaking large language models capable of performing complex reasoning tasks holds significant promise for addressing various scientific…

Source: arXiv cs.CL Yakun Zhu, Zhongzhen Huang, Linjie Mu, Yutong Huang, Wei Nie, Jiaji Liu, Shaoting Zhang, Pengfei Liu, Xiaofan Zhang
AI Research AI

Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning

arXiv:2608.16333v1 Announce Type: new Abstract: On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories. This approach has achieved strong empirical…

Source: arXiv cs.CL Changhui Sun, Lanbo Liu, Hang Lei, Tong Ling, Jiahang Xie, Zhiyong Zheng, Yujia Wang, Hao Liu, Feng Xiao, Lu Liu, Yanlong Du, Zifeng Cheng, Ziwei Jiang, Qing Gu
AI Research AI

Mitigating Bias in Locally Constrained Decoding via Tractable Proposals

arXiv:2606.01926v2 Announce Type: replace Abstract: Generations from large language models often fail to conform to desired constraints such as JSON schema. Existing locally constrained decoding (LCD) approaches enforce…

Source: arXiv cs.CL Meihua Dang, Linxin Song, Honghua Zhang, Jieyu Zhao, Guy Van den Broeck, Stefano Ermon
AI Research AI

HarmProfile: Characterizing Harmful Distributions in Frontier LLMs

arXiv:2608.14577v1 Announce Type: new Abstract: Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis. Consequently,…

Source: arXiv cs.CL Zhouyuan Ma, Yutao Wu, Hanxun Huang, Xiang Zheng, Xiao Liu, Yixin Cao, Zuxuan Wu, Xingjun Ma, Yu-Gang Jiang
AI Research AI

Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models

arXiv:2608.16647v1 Announce Type: new Abstract: On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains…

Source: arXiv cs.CL Zhaoyi Li, Deyang Kong, Yuan Wei, Evan Yang, Ranran Shen, Mahardika Krisna Ihsani, Ming Yang, Wei Zhang, Chuan Hao, Jian Yang, Ran Tao, Bryan Dai, Shikun Zhang, Wei Ye, Ying Wei, Defu Lian
AI Research AI

Analyzing Speech Condition Effects in Dysarthric ASR: A Layer-wise Probing Study

arXiv:2608.01865v2 Announce Type: replace Abstract: Automatic speech recognition (ASR) performance degrades sharply on dysarthric speech, yet how disordered articulation reshapes a model's internal representations is…

Source: arXiv cs.CL Darwin Jelestin Muthu, Navya Gupta, Wei Lin Tay, Zhengchen Zhang, Daniel Wang Zhengkui, Rong Tong
AI Research AI

jina-vlm: Small Multilingual Vision Language Model

arXiv:2512.04032v4 Announce Type: replace Abstract: We present jina-vlm, a token-efficient 2.4B parameter vision-language model that achieves state-of-the-art multilingual VQA performance among open 2B-scale VLMs. The…

Source: arXiv cs.CL Andreas Koukounas, Georgios Mastrapas, Florian H\"onicke, Sedigheh Eslami, Guillaume Roncari, Han Xiao
AI Research AI

SocialCoach: Personalized Social Skill Learning with Agentic Tutoring and Practice

arXiv:2606.04155v2 Announce Type: replace-cross Abstract: Social skills such as negotiation and leadership are crucial for personal and professional success in today's interconnected world. However, scalable and…

Source: arXiv cs.CL Tianfu Wang, Max Xiong, Jianxun Lian, Hongyuan Zhu, Zhengyu Hu, Yuxuan Lei, Linxiao Gong, Dapeng Hu, Xiaofang Li, Peiting Tsai, Nicholas Jing Yuan, Qi Zhang
AI Research AI

Iterative Self-Learning for Expressive Text-to-Speech Synthesis

arXiv:2608.15910v1 Announce Type: cross Abstract: Expressive text-to-speech (TTS) systems that use explicit conditioning labels provide direct and interpretable control over expressive attributes, in contrast to…

Source: arXiv cs.CL Nicholas Sanders, Gustav Eje Henter, Simon King, Korin Richmond
AI Research AI

LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents

arXiv:2606.06087v2 Announce Type: replace Abstract: Agent systems increasingly use textual skills to encode reusable task procedures, but injecting these skills into the prompt at every step incurs substantial context…

Source: arXiv cs.CL Aofan Yu, Chenyu Zhou, Tianyi Xu, Zihan Guo, Rong Shan, Zhihui Fu, Jun Wang, Weiwen Liu, Yong Yu, Weinan Zhang, Jianghao Lin
AI Research AI

Model Hypnosis: Strong control of AI via additive subliminal effects

arXiv:2608.16834v1 Announce Type: new Abstract: We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually weak and seemingly irrelevant cues in the prompt can be…

Source: arXiv cs.CL Enric Boix-Adsera, Benedict Tessler
AI Research AI

Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization

arXiv:2608.16539v1 Announce Type: cross Abstract: Large Audio Language Models (LALMs) have made rapid progress on standardized benchmarks, yet their deployment in practical media workflows, curation, archival indexing,…

Source: arXiv cs.CL Tony Alex, Wish Suharitdamrong, Sara Atito, Armin Mustafa, Muhammad Awais, Philip J. B. Jackson, Jiankang Deng, Ismail Elezi
AI Research AI

FTA-Mem: Fact-Time-Affect Anchored Memory for Low-Density Long-Term Dialogue

arXiv:2608.16303v1 Announce Type: new Abstract: Long-term emotional-support agents require memory mechanisms for personalized understanding across sessions. However, emotional-support dialogue is often low-density:…

Source: arXiv cs.CL Chang Liu, Shuyi Zhang, Changsheng Ma, Yongfeng Tao, Minqiang Yang, Bin Hu
AI Research AI

Using the Mimi codec for metalinguistic representations

arXiv:2608.15799v1 Announce Type: new Abstract: In this paper, we focus on the dictionary of 2048 tokens used in Mimi semantic token codebook, the neural codec of the Moshi language model. We show that the ABX…

Source: arXiv cs.CL Artem Saloev, Erin Pacquetet, Nicolas Ballier
AI Research AI

CAPO: Constraint-Aware Prompt Optimization for LLM Agents

arXiv:2608.16068v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as agents that rely on system prompts to use tools and complete tasks. Such deployments impose distinct operational…

Source: arXiv cs.CL Victor Ye Dong, Reid Pryzant, Yi Liu, Jian Jiao
AI Research AI

Hallucination Span Detection with Input-Side Evidence Alignment

arXiv:2608.15804v1 Announce Type: new Abstract: Hallucinations remain a major obstacle to the reliable use of large language models (LLMs) in conditional text generation. Existing methods primarily assess the factuality…

Source: arXiv cs.CL Miyu Yamada, Yuki Arase
AI Research AI

HalluTracer: Hallucination Detection via Depth-Averaging Truth Signals

arXiv:2608.16353v1 Announce Type: new Abstract: Even well-aligned large language models confidently generate factually incorrect text, making hallucination a persistent reliability risk in high-stakes deployments. These…

Source: arXiv cs.CL Zhihao Guo, Zonghan Wu, Huan Huo, DaYong Ye, Junwei Zhang, Weiran Yao, Zhiwei Liu, Qingsong Wen, Yilei Shao
AI Research AI

On the Role of Directionality in Structural Generalization

arXiv:2607.02307v2 Announce Type: replace Abstract: Several SLOG test categories explicitly involve directional distinctions (modifier position shifts, argument extraction positions), yet AM-Parser, the previous SOTA,…

Source: arXiv cs.CL Zichao Wei
AI Research AI

Propaganda Forensics: Recovering the Generation Pipeline of an AI-Driven Influence Campaign

arXiv:2608.15746v1 Announce Type: cross Abstract: We present a forensic analysis of the generation pipeline behind a recent AI-driven influence campaign. We introduce PROPAGIA, a corpus of 2,646 propagandist French…

Source: arXiv cs.CL Benjamin Icard, Elouan Vuichard, Louis Lefebvre, Lila Sainero, Thomas Girault, Alice Breton, Tanguy Launay, Gauvain Bourgne, Morgane Casanova, Guillaume Gadek, Victor Kl\"otzer, Michel Le Nouy, Guill…
AI Research AI

The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese

arXiv:2607.10745v2 Announce Type: replace Abstract: This paper presents the first ChineseBabyLM Challenge, organized as part of NLPCC 2026. The challenge asked participants to train language models from scratch using no…

Source: arXiv cs.CL Siyuan Song, Zhiheng Qian, Yunhao Zhang, Linyang He, Xiaozhe Ji, Yingxin Lin, Hongao Zhu, Chongtian Shao, Chuhan Lang, Luan Li, Rui Wang, Renfen Hu, Shaonan Wang, Hai Hu
AI Research AI

Wiktionary as a Crowdsourced Lexicon for English Dialects

arXiv:2608.15641v1 Announce Type: new Abstract: This paper evaluates Wiktionary as an ethically crowdsourced lexicon for English dialects. We took a two-phase approach, providing an in-depth descriptive analysis of the…

Source: arXiv cs.CL Sidney Wong
AI Research AI

Writing Style Similarity Reflects Academic Genealogy

arXiv:2608.14843v1 Announce Type: new Abstract: As authorship attribution systems are increasingly deployed to detect ghostwritten and AI-generated papers, their errors can support accusations against legitimate…

Source: arXiv cs.CL Cameron Manzo
AI Research AI

NARRATE: A Multimodal Real-World Australian Driving Dataset for Human-Centred Explanations in Automated Driving

arXiv:2608.14767v1 Announce Type: cross Abstract: Automated vehicles must explain their decisions in ways that passengers can understand, monitor, and trust. Existing language-annotated driving datasets are mostly…

Source: arXiv cs.CL Ashkan Yousefi Zadeh, Zishuo Zhu, Xiaomeng Li, Andry Rakotonirainy, Sebastien Glaser, Ronald Schroeter, Patricia Delhomme, Zahra Mehraban
AI Research AI

I-CALM: Incentivizing Confidence-Aware Abstention for LLM Selective Answering

arXiv:2604.03904v2 Announce Type: replace Abstract: Large language models (LLMs) often produce confident but incorrect answers, in part because standard evaluation incentives reward guessing over expressing uncertainty.…

Source: arXiv cs.CL Haotian Zong, Binze Li, Yufei Long, Sinyin Chang, Jialong Wu, Gillian K. Hadfield
AI Research AI

jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers

arXiv:2605.08384v4 Announce Type: replace Abstract: In this work, we introduce GELATO (Geometry-preserving Embeddings via Locked Aligned TOwers), a novel approach to multimodal embedding models. We build on the…

Source: arXiv cs.CL Florian H\"onicke, Michael G\"unther, Andreas Koukounas, Mohammad Kalim Akram, Saba Sturua, Han Xiao
AI Research AI

QA-Merging: Query-Adaptive Reasoning via Layer Selective Model Merging

arXiv:2601.03506v2 Announce Type: replace Abstract: Recent large reasoning models (LRMs) have achieved strong performance on complex reasoning tasks by generating a long chain-of-thought (Long-CoT). However, such…

Source: arXiv cs.CL Zhaofeng Zhong, Wei Yuan, Tong Chen, Liang Qu, Xiangyu Zhao, Quoc Viet Hung Nguyen, Hongzhi Yin
AI Research AI

Agentic Test-Time Scaling for WebAgents

arXiv:2602.12276v2 Announce Type: replace-cross Abstract: Test-time scaling has become a standard way to improve performance and boost reliability of neural network models. However, its behavior on agentic, multi-step…

Source: arXiv cs.CL Nicholas Lee, Lutfi Eren Erdogan, Chris Joseph John, Surya Krishnapillai, Michael W. Mahoney, Kurt Keutzer, Amir Gholami
AI Research AI

ContextClaim: A Context-Driven Paradigm for Verifiable Claim Detection

arXiv:2603.30025v2 Announce Type: replace Abstract: Automated fact-checking pipelines typically begin with a filtering stage that decides which claims are worth verifying, given that the later evidence retrieval and…

Source: arXiv cs.CL Yufeng Li, Rrubaa Panchendrarajan, Arkaitz Zubiaga
AI Research AI

Structural Generalization on SLOG without Hand-Written Rules

arXiv:2604.26157v4 Announce Type: replace Abstract: Structural generalization in semantic parsing requires systems to apply learned compositional rules to novel structural combinations. Existing approaches either rely…

Source: arXiv cs.CL Zichao Wei
AI Research AI

Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans

arXiv:2608.16514v1 Announce Type: cross Abstract: Human visual search is serial: the fovea must land on a candidate to confirm it, and those landings form a scanpath. Whether multimodal large language models (MLLMs),…

Source: arXiv cs.CL Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Ulas Bagci, Alessandro Bruno
AI Research AI

Logical Embeddings for Argument Analysis

arXiv:2608.15325v1 Announce Type: new Abstract: We propose a new framework for machine-learning-oriented argument analysis tasks. Our proposal involves replacing traditional contextualized word embeddings used in most…

Source: arXiv cs.CL Leander Heldring, Santiago Torres
AI Research AI

Subliminal Steering: Stronger Encoding of Hidden Signals

arXiv:2604.25783v2 Announce Type: replace Abstract: Subliminal learning describes a student language model inheriting a behavioral bias by fine-tuning on seemingly innocuous data generated by a biased teacher model.…

Source: arXiv cs.CL George Morgulis, John Hewitt
AI Research AI

DanceOPD: On-Policy Generative Field Distillation

arXiv:2606.27377v3 Announce Type: replace-cross Abstract: Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), local editing, and global editing. However,…

Source: arXiv cs.CL Wei Zhou, Xiongwei Zhu, Zelin Xu, Bo Dong, Lixue Gong, Yongyuan Liang, Meng Chu, Leigang Qu, Lingdong Kong, Wei Liu, Tat-Seng Chua
AI Research AI

Sequential LLM Release Facilitates Manipulation in Regulated Markets

arXiv:2601.11496v3 Announce Type: replace-cross Abstract: AI agents increasingly mediate bargaining, negotiation and persuasion for people and firms. Such markets extend software-mediated commerce, but add a governance…

Source: arXiv cs.CL Eilam Shapira, Moshe Tennenholtz, Roi Reichart
AI Research AI

TRACE-BN: Transferring Bangla-English Tutoring Behavior to a Sub-1B Offline Language Model

arXiv:2608.15223v1 Announce Type: new Abstract: Bangla-English tutoring requires more than producing a correct translation: learners also need explanations of grammar differences, awareness of their likely errors, and…

Source: arXiv cs.CL Khan Raiyan Ibne Reza, Sanjana Aktar Maria, Mohammad Tushar Abdullah, Asfee Bhuiyan Leen, Sumaiya Tabassum Nimi
AI Research AI

$R^3$-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets

arXiv:2608.16033v1 Announce Type: new Abstract: In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use…

Source: arXiv cs.CL Peisong Wang, Zhiwei Ma, Bowen Liu, Feixue Liu, Aochuan Chen, Chenyi Zi, Hongchuan Zeng, Yuhan Li, Jia Li
AI Research AI

Scaling Manual-Grounded Appliance Manipulation with Data Synthesis and Unified Planning

arXiv:2608.15863v1 Announce Type: cross Abstract: Operating household appliances requires long-horizon planning that is state-dependent and robust to disturbances, yet existing large models fall short, as no…

Source: arXiv cs.CL Yuxing Long, Lei Kang, Ziyan Yu, Yuzheng Gao, Bin Cheng, Jiyao Zhang, Xiaoqi Li, Haolin Yang, Dongjiang Li, Hui Shen, Hao Dong
AI Research AI

Language Models Compare Quantities Using Number-specific and Unit-specific Heuristics

arXiv:2606.03982v2 Announce Type: replace Abstract: Quantities with measurement units, such as 110 cm and 1.2 m, require language models (LMs) to combine a numeral with a symbolic unit scale. Here, we study how LMs…

Source: arXiv cs.CL Mutsumi Sasaki, Go kamoda, Ryosuke Takahashi, Kosuke Sato, Kentaro Inui, Keisuke Sakaguchi, Benjamin Heinzerling
AI Research AI

QUMem: Personalized Memory for Query-Conditioned User-State Inference in LLM Agents

arXiv:2608.16168v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly use external memory systems to support personalization by drawing on long and evolving interaction histories, in which user…

Source: arXiv cs.CL Heng Wang, Yifei Li, Lingling Zhang, Pengyu Li, Xinyu Che, Xinyu Zhang, Zesheng Yang
AI Research AI

What to Forget in Unlearning? Forget Set Curation for Language Models

arXiv:2608.14855v1 Announce Type: new Abstract: Machine unlearning aims to remove targeted data or behaviors from a trained model without retraining from scratch. Yet most evaluations assume that the examples to forget…

Source: arXiv cs.CL Animesh Jha, Arpandeep Khatua, Youssef Allouah, Sanmi Koyejo
AI Research AI

Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns

arXiv:2607.11621v2 Announce Type: replace-cross Abstract: Aphasia following stroke commonly produces systematic naming errors with characteristic profiles, but whether general-purpose language models not designed for…

Source: arXiv cs.CL Yong Yang, Xiang Guan, Sophie Arheix-Parras, Saeed Ahmadi, Roger Newman-Norlund, Leonardo Bonilha, Christopher Rorden, Julius Fridriksson, Rutvik H. Desai, Srihari Nelakuditi
AI Research AI

Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents

arXiv:2608.15008v1 Announce Type: new Abstract: Memory is becoming core infrastructure for long-horizon LLM agents, yet existing evaluations offer limited guidance on which memory substrate, namely the underlying medium…

Source: arXiv cs.CL Wei-Chieh Huang, Weizhi Zhang, Yuchen Wu, Yankai Chen, Eric Hanchen Jiang, Wooseong Yang, Yiwei Yang, Henry Peng Zou, Hanrong Zhang, Ying Nian Wu, Haolun Wu, Kai-Wei Chang, Philip S. Yu, Xue Liu, Ayl…
AI Research AI

IndicQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages

arXiv:2608.16344v1 Announce Type: new Abstract: Indic quality estimation (QE) and automatic post-editing (APE) data is spread across separate releases, so no single resource supports training and evaluation across tasks…

Source: arXiv cs.CL Diptesh Kanojia, Archchana Sindhujan, Sourabh Deoghare, Daria Sokova, Shenbin Qian, Girish Koushik, Tharindu Ranasinghe, Constantin Or\u{a}san, Chrysoula Zerva, Ricardo Rei, Fr\'ed\'eric Blain, Andr\…
AI Research AI

Misconception Diagnosis From Student-Tutor Dialogue: Generate, Retrieve, Rerank

arXiv:2602.02414v2 Announce Type: replace Abstract: Timely and accurate identification of student misconceptions is key to improving learning outcomes and pre-empting the compounding of student errors. However, this…

Source: arXiv cs.CL Joshua Mitton, Prarthana Bhattacharyya, Digory Smith, Thomas Christie, Ralph Abboud, Simon Woodhead
AI Research AI

NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representations in Vision-Language Models

arXiv:2608.15425v1 Announce Type: cross Abstract: Vision-language models (VLMs) achieve strong performance on high-level multimodal tasks, yet numerosity perception, a cognitive ability that emerges in human infants…

Source: arXiv cs.AI Yiming Fu, Fangjun Li, Xiujin Liu, Ruidong Ma, Hang Yu, Zhichen Lu, Kanwei He, Alessandro Di Nuovo, Angelo Cangelosi, Zhegong Shangguan
AI Research AI

Position: AI Lock-In Is in Progress, and We Must Be Prepared

arXiv:2608.14565v1 Announce Type: new Abstract: AI safety research has mainly focused on two areas: technical alignment (ensuring AI systems produce human-aligned outputs) and the regulation of generative AI's societal…

Source: arXiv cs.AI Jaeho Kim, Seokhyun Lee, Jieun Lee, Changhee Lee
AI Research AI

From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems

arXiv:2608.15127v1 Announce Type: cross Abstract: Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state.…

Source: arXiv cs.AI Chaokun Chang, Yukun Zhou, Kaihua Fu, Dakai An, Tianyu Feng, Hanfeng Lu, Sheng Yao, Pu Guo, Yinghao Yu, Yizhou Shan, Bo Li, Binhang Yuan, Wei Wang
AI Research AI

Audio-Visual Segmentation via Depth-Guided Collaborative Modeling

arXiv:2608.16285v1 Announce Type: cross Abstract: Audio-Visual Segmentation (AVS) is a fundamental task in multimodal perception that performs pixel-level segmentation of sounding objects in videos by leveraging both…

Source: arXiv cs.AI Zhaojin Fu, Yuyang Hong, Qi Yang, Zili Wang, Kun Ding, Shiming Xiang, Bin Fan
AI Research AI

OTel: Building Domain-Specialized Telecom LLM Foundations for Intelligent Networks

arXiv:2608.15436v1 Announce Type: new Abstract: Frontier AI models have advanced rapidly, but they still struggle with telecom-specific tasks. We present Open Telco (OTel), an open telecom AI resource with derived…

Source: arXiv cs.AI Farbod Tavakkoli, Roderic Paulk, Jorden Terrazas, Kenneth Church, Mark Austin, Louis Powell, Gregory Diamos, Lina Bariah, Syed Ali Raza Zaidi, Maryam Hafeez, Ali Maatouk, Imtiaz Karim
AI Research AI

When Do LLMs Apply the Wrong Law? Diagnosing LLM Failures in Temporal Legal Reasoning

arXiv:2608.14610v1 Announce Type: new Abstract: Legal reasoning tasks such as legal judgment prediction (LJP) require identifying the temporally correct version of the law governing a case -- a capability we term…

Source: arXiv cs.AI Yiqian Huang, Shuyuan Zheng, Qianying Liu, Shaowen Peng, Yuntao Kong, Kotaro Funakoshi, Chuan Xiao, Manabu Okumura, Yang Cao
AI Research AI

Baseline-Relative Counterfactual Refinement for Bit-Aware Visual Token Communication

arXiv:2608.16192v1 Announce Type: new Abstract: Generative visual-token communication reduces transmission load by sending only selected discrete tokens and reconstructing missing content at the receiver. However,…

Source: arXiv cs.AI Jia Guo, Xiaohan Zhao, Changwang Liu, Shuqing He, Chenyang Zhang, Bingchuan Zhao, Jinqi Zhu
AI Research AI

AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research

arXiv:2608.11216v2 Announce Type: replace Abstract: World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across…

Source: arXiv cs.AI Marjan Moodi, Xuankang Zhu, Fernando De Mesentier Silva, Harold Chaput, Mohammad Reza Taesiri
AI Research AI

CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs

arXiv:2608.14791v1 Announce Type: new Abstract: Abductive reasoning, often characterized as inference to the best explanation, is central to explanation under uncertainty, from everyday sense-making and investigation to…

Source: arXiv cs.AI Moein Salimi, Danial Parnian, Shaygan Adim, Amirmohammad Ebrahiminasab, Nima Alighardashi, Parsa Gholami, Sahand Akramipour, Mahdi Jafari Siavoshani, Mohammad Hossein Rohban
AI Research AI

UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations

arXiv:2608.15930v1 Announce Type: new Abstract: Foundation GUI agents can automate complex digital tasks, but deployment is hindered by scarce and biased training data, ambiguous prompts, and unreliable execution.…

Source: arXiv cs.AI Zihan Ding, Longxu Dou, Qi Gao, Xiangwu Guo, Shengchao Hu, Zilong Huang, Zihang Jiang, Lei Ke, Mengcheng Lan, Weixian Lei, Hanxuan Li, Honglin Li, Xiyun Li, Zaitang Li, Leowei Liang, Xin Luo, Haozhe…
AI Research AI

HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL

arXiv:2608.16837v1 Announce Type: cross Abstract: Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) foundation models are not…

Source: arXiv cs.AI Langzhe Gu, Chengkai Hou, Meng Li, Xinhua Wang, Jiaming Liu, Xinyuan Lv, Bowei Zhang, Shuanghao Bai, Guangrun Li, Jingyang He, Gaole Dai, Ziluo Ding, Zhiyuan Xu, Kuan Cheng, Jian Tang, Zhengping Che,…
AI Research AI

Constitutive Priors for Machine Intelligence: A Legitimacy Theory of the Artificial Physical World

arXiv:2608.15147v1 Announce Type: new Abstract: Machine intelligence has conquered the symbolic world but stalled at the physical one. The stall is structural: physical AI faces a cold-start deadlock -- no intelligence…

Source: arXiv cs.AI Jiang Jiang (Persagy Science and Technology Co., Beijing, China), Yifu Sun (Persagy Science and Technology Co., Beijing, China), Qi Shen (Persagy Science and Technology Co., Beijing, China)
AI Research AI

OGX: An Open-Source, Vendor-Neutral Generative AI Application Server

arXiv:2608.14580v1 Announce Type: new Abstract: OGX (Open GenAI Stack) is an open-source AI application server and Python library that implements the APIs of major frontier labs (OpenAI, Anthropic, Google) with…

Source: arXiv cs.AI Francisco Javier Arceo, S\'ebastien Han, Matthew Farrellee, Charlie Doern, Yuan Tang, Derek Higgins, Varsha Prasad Narsing, Gordon Sim, Sumanth Kamenani, Ben Browning, Raghotham Murthy
AI Research AI

Large Language Models and their Awareness of Mechanics and Spatial Geometry

arXiv:2608.14615v1 Announce Type: new Abstract: Large Language Models (LLMs) perform well on established code-generation and mathematical-reasoning benchmarks, but their capabilities in mechanics and spatial geometry,…

Source: arXiv cs.AI Johannes Gerstmayr, Sebastian Weyrer, Tobias M\"oltner, Peter Manzl, Michael Pieber
AI Research AI

Augmenting Text to Increase Translation Difficulty

arXiv:2608.15932v1 Announce Type: new Abstract: As state-of-the-art machine translation models saturate standard benchmarks, the field needs more challenging evaluations to distinguish between models of varying quality.…

Source: arXiv cs.AI William Kalikman, \v{S}imon Sukup, Michal Te\v{s}nar, Vil\'em Zouhar
AI Research AI

Software Engineering for AI-driven Building Operation

arXiv:2608.16237v1 Announce Type: cross Abstract: Building operations are energy-inefficient. Artificial Intelligence (AI)-driven control systems promise benefits through optimization and predictive control, but…

Source: arXiv cs.AI Philipp Zech, Sascha Hammes, Johannes Weninger, J\"urgen Pannosch, Gernot Steidl
AI Research AI

AstronOS: A Unified Execution Model and Runtime for Long-Horizon Agentic Systems

arXiv:2608.16381v1 Announce Type: new Abstract: Agentic systems often organize execution and state around a single conversation, model invocation, or agent instance, even when real work spans many calls and stages. We…

Source: arXiv cs.AI Zhenhang Nie (iFLYTEK Co., Ltd., Hefei, China), Gui Zheng (iFLYTEK Co., Ltd., Hefei, China), Xudong Sun (iFLYTEK Co., Ltd., Hefei, China), Tailong Zhu (iFLYTEK Co., Ltd., Hefei, China), Bin Zhang (iF…
AI Research AI

Dear Algo: A Precision-First Agentic Intent Layer for Unified Search and Recommendation

arXiv:2608.15877v1 Announce Type: new Abstract: Search and recommendation serve a shared discovery objective but encode intent differently. We study this boundary through Dear Algo on Threads, a deployed product where…

Source: arXiv cs.AI Rui Wang, Jiazhou Wang, Zheng Wei, Chenglin Lu, Fangcheng Sun, Ivy Sun, Jin Sun, Hui Geng, Lillian Zhang, Chao Yang, Lei Chen, Shahin Sefati, Reem Helou, Joe Zhou, Babak Shakibi, Yiyi Pan, Bi Xue, Ho…
AI Research AI

Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems

arXiv:2608.14667v1 Announce Type: new Abstract: Large language model-based agents are increasingly deployed as collaborators in scientific discovery yet most current work focuses on the autonomous capabilities of "AI…

Source: arXiv cs.AI Patrick Emami, Sameera Horawalavithana, Truc Nguyen, Gihan Panapitiya, Bruno Jacob, Siddhisanket Raskar, Saumya Sinha, Jared D. Willard, Andrew Glaws, Nithin Somasekharan, Ling Yue, Brian Lu, Shaowu…
AI Research AI

MELD: A Protocol for Merging Knowledge Across Distributed Agentic Memories

arXiv:2608.16357v1 Announce Type: cross Abstract: Autonomous agents share a transport and can call each other's tools, but they cannot share what they know: no protocol lets two agents' memories reconcile a fact phrased…

Source: arXiv cs.AI Lauri Lov\'en, Jaakko Sauvola, Jukka Riekki, Sasu Tarkoma
AI Research AI

Beyond Direct Access: Resource Hijacking in LLM Agents

arXiv:2608.15108v1 Announce Type: cross Abstract: Large language model agents are increasingly connected to high-value resources such as computing infrastructure, credentials, usage budgets, identities, private…

Source: arXiv cs.AI Puyu Zeng, Qibing Ren
AI Research AI

Multi-Agent Closed-Loop Reasoning for Organic Structure Elucidation from Multimodal Spectra

arXiv:2608.14720v1 Announce Type: cross Abstract: Following the molecular discovery and synthesis revolutions, scalable automated structure elucidation from routine spectroscopic data remains an outstanding challenge.…

Source: arXiv cs.AI Bingsen Xue, Zhuojun Jiang, Jianhao Zhang, Mingcheng Gu, Yizhe Yuan, Yongtai Zhuo, Yifan Zhang, Li Wang, Ya Su, Yue Yuan, Jiang Liu, Xueqian Kong, Cheng Jin
AI Research AI

Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents

arXiv:2608.16578v1 Announce Type: new Abstract: AI agents increasingly operate as part of interacting systems rather than in isolation. As agents exchange information and jointly make decisions, their interactions can…

Source: arXiv cs.AI Batu El, Jinhee Paeng, Fatih Dinc, Shiye Su, Mete Erdogan, Aneesh Pappu, Haotian Ye, Wanjia Zhao, Surya Ganguli, James Zou
AI Research AI

RISE: Roadside Infrastructure Sequence Understanding across 3D Tracking and Structured Vision-Language Reasoning

arXiv:2608.16480v1 Announce Type: cross Abstract: We present RISE (Roadside Infrastructure Sequence Understanding and Evaluation), a framework spanning metric 3D tracking and structured vision-language reasoning in…

Source: arXiv cs.AI Yanbo Jiang, Haotian Zheng, Jiahao Wang, Hanxiao Ren, Yitao Xu, Yining Xing, Zehong Ke, Hao Cheng, Yiqian Tu, Jinhao Li, Zhiyuan Xuan, Fang Zhang, Jianqiang Wang
AI Research AI

GRIP: Grounded Reasoning via Information-Restricted Premises

arXiv:2608.16776v1 Announce Type: new Abstract: High-capacity encoders in retrieval-augmented generation (RAG) can let the query dominate the latent state, leaving retrieved evidence functionally irrelevant. We call…

Source: arXiv cs.AI Lirui Teng
AI Research AI

Euclid-Omni : A Unified Neuro-Symbolic Framework for Plane Geometry

arXiv:2608.14585v1 Announce Type: new Abstract: Euclidean geometry is a compelling testbed for AI reasoning, as it demands the combination of intuitive diagram understanding, axiomatic deduction, and algebraic…

Source: arXiv cs.AI Zhaoyu Li, Hangrui Bi, Youyuan Zhang, Wenjie Ma, Zenan Li, Zhaolei Zhang, Xujie Si, Kaiyu Yang
AI Research AI

HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction

arXiv:2608.16222v1 Announce Type: cross Abstract: Humanoid intelligence requires learning over an extremely diverse space of whole-body motions and physically grounded interactions. However, existing embodied datasets…

Source: arXiv cs.AI Jiahao Ji, Ji Ma, Runhan Zhang, Runyi Yu, Wenjia Wang, Weiheng Chi, Qianqian Peng, Weichao Yan, Yongfei Gu, Ye Tian, Ting Wu, Longwei Li, Chun Yuan, Ruoli Dai, Lei Han
AI Research AI

TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation

arXiv:2608.16765v1 Announce Type: cross Abstract: Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g.,…

Source: arXiv cs.AI Haoran Wang, Chaofan Ma, Ran Yi, Lizhuang Ma
AI Research AI

A cross-modal generative model for incomplete and degraded prostate MRI with multicentre clinical validation

arXiv:2608.16233v1 Announce Type: cross Abstract: Missing or degraded sequences can limit prostate multiparametric MRI. We developed MSCNet, a sequence-conditioned cross-modal generative framework for reconstructing…

Source: arXiv cs.AI Siyuan Ma, Liang He, Mengying Zhu, Yi Chai, Mengyao Lyu, Haowei Wang, Qizhen Lan, HaoBo Sun, Qixin Zhang, Jingli Chen, Xiaobing Wei, Jiaming Liu, Guiqin Liu, Qianwen Zhang, Yang Liu, Dacheng Tao, Gua…
AI Research AI

MoE Router-Guided Clustering for Heterogeneous Federated Instruction Tuning

arXiv:2608.15311v1 Announce Type: new Abstract: Federated instruction fine-tuning enables Large Language Models (LLMs) to adapt to decentralized, privacy-sensitive data without requiring data sharing. Recent…

Source: arXiv cs.AI Ankita Sharma, Bahar Farahani, Sanaz Rahimi Moosavi, Amir Rrahmani, Farshad Firouzi, Krishnendu Chakrabarty
AI Research AI

Translating finite-domain integer constraint models to CP/SMT/ILP/PB/SAT solvers with CPMpy

arXiv:2608.15143v1 Announce Type: new Abstract: Constraint solving is a declarative approach for solving combinatorial satisfaction and optimization problems. The user specifies their problem through constraints and…

Source: arXiv cs.AI Tias Guns, Ignace Bleukx, Hendrik Bierlee, Jo Devriendt, Emilio Gamba, Orestis Lomis, Wout Piessens, Thomas Sergeys, Dimos Tsouros, Wout Vanroose, H\'el\`ene Verhaeghe
AI Research AI

Understanding Cognition-Induced Risks in Agentic AI Systems

arXiv:2608.15304v1 Announce Type: new Abstract: Frontier agentic systems powered by large language models (LLMs) exhibit human-like patterns of cognition. As these systems become deeply integrated across different…

Source: arXiv cs.AI Guanchu Wang, Qinuo Li, Mengnan Du, Xia Hu, Bowen Zhou
AI Research AI

Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability

arXiv:2604.06628v2 Announce Type: replace Abstract: A prevailing narrative in LLM post-training holds that supervised finetuning (SFT) memorizes while reinforcement learning (RL) generalizes. We revisit this claim for…

Source: arXiv cs.AI Qihan Ren, Peng Wang, Ruikun Cai, Shuai Shao, Dadi Guo, Yuejin Xie, Yafu Li, Quanshi Zhang, Xia Hu, Jing Shao, Dongrui Liu
AI Research AI

Does the Proof Prove It That Way? Faithful Formalization of Elements Proofs

arXiv:2608.15432v1 Announce Type: new Abstract: In formal verification, both the autoformalization of statements and automated proof search have been studied extensively. While automated proof search can produce a…

Source: arXiv cs.AI Tadd Mao, Tianjun Zhong, Dhruva Arekar, Yuming Feng, One An, Jiani Huang, Xujie Si, Ziyang Li
AI Research AI

X$^2$Localizer: Cross-grained Alignment for Progressive Cross-view Video Geo-localization

arXiv:2608.16658v1 Announce Type: cross Abstract: Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their corresponding geo-tagged aerial images. However, CVG approaches rely on…

Source: arXiv cs.AI Zichao Zeng, Weijia Fan, Yufan Chen, June Moh Goo, Junwei Zheng, Ruiping Liu, Kunyu Peng, Jiaming Zhang, Rainer Stiefelhagen, Jan Boehm
AI Research AI

Quipu: A Governed Bitemporal Knowledge Graph Store

arXiv:2608.16813v1 Announce Type: new Abstract: Agents now write knowledge graphs, but knowledge-graph stores still carry defaults set when humans curated them: accept writes now and clean later, keep one time axis or…

Source: arXiv cs.AI Steve Brown
AI Research AI

Budget-Aware Tool Use Enables Effective Agent Scaling

arXiv:2511.17006v2 Announce Type: replace Abstract: Scaling test-time computation has been extended from language model reasoning to tool-augmented agents, where scaling involves not only thinking in tokens but also…

Source: arXiv cs.AI Tengxiao Liu, Zifeng Wang, Jin Miao, I-Hung Hsu, Jun Yan, Jiefeng Chen, Rujun Han, Fangyuan Xu, Yanfei Chen, Ke Jiang, Samira Daruki, Yi Liang, William Yang Wang, Tomas Pfister, Chen-Yu Lee
AI Research AI

Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs

arXiv:2608.16391v1 Announce Type: cross Abstract: As large language models become increasingly widespread, third-party providers that deploy open-weight models have become an important part of the ecosystem. Auditing…

Source: arXiv cs.AI Xiangfan Wu, Zonghao Ying, Huiyu Wu, Xing Zheng, Huangsheng Cheng, Xiaorong Shi, Jing Guo
AI Research AI

Visualizing Uncertainty-to-Action Composition for Human Oversight

arXiv:2608.16428v1 Announce Type: cross Abstract: Artificial intelligence systems often disclose uncertainty, yet they rarely make clear what response that uncertainty should trigger. Most uncertainty visualizations…

Source: arXiv cs.AI Chisom Anyabolu, Akshat Dubey, Georges Hattab
AI Research AI

Runtime Observability for Heterogeneous Attention Memory

arXiv:2608.05863v2 Announce Type: replace Abstract: Modern models no longer keep a plain KV cache: latent caches, learned sparse selectors and recurrent states each carry the model's memory in a different form, and each…

Source: arXiv cs.AI Fanzhe Wei, Li Liu, Ziyang Wang, Chenyu Wang
AI Research AI

Incoherent by Design? On the Moral Self-Consistency of LLMs

arXiv:2608.15354v1 Announce Type: new Abstract: LLMs are increasingly used in morally sensitive contexts, yet it is unclear whether they apply ethical principles consistently across situations. A model that can state a…

Source: arXiv cs.AI Pegah Nokhiz, Aravinda Kanchana Ruwanpathirana, Helen Nissenbaum
AI Research AI

Kozuchi Agent: A Language-Agnostic Open-Weight Agent for Software Repair

arXiv:2608.15579v1 Announce Type: cross Abstract: Industrial software-engineering teams increasingly need LLM agents that turn bug reports into correct patches, yet benchmark-scale operation adds long horizons, tool-use…

Source: arXiv cs.AI Mehdi Bahrami, Kosaku Kimura, Satoshi Munakata, Satoshi Nakashima, Yu Ishikawa, Kosuke Maeda, Nao Soma, Kenichi Kobayashi, Keisuke Miyazaki, Keizo Kato, Shigeki Fukuta, Tatsuo Kumano, Nobutaka Imamur…
AI Research AI

LAPF: LLM-Agent-Based Path Finder Using the UAVScenes Dataset

arXiv:2608.15175v1 Announce Type: cross Abstract: Uncrewed aerial vehicles (UAVs) are increasingly deployed for autonomous navigation in complex outdoor environments, where dynamic conditions and mission requirements…

Source: arXiv cs.AI Yousef Emami, Mohammadhossein Homaei, Hao Zhou, Miguel Guti\'errez Gait\'an, Atefeh Hajijamali Arani, Rui Zhang
AI Research AI

Information-Theoretic Causal Modelling of Semiconductor Process Dynamics

arXiv:2608.14678v1 Announce Type: cross Abstract: With the progress of the semiconductor industry toward increasingly complex compute devices and tighter process tolerances, advanced process control has become crucial.…

Source: arXiv cs.AI Daniel S{\o}rensen, Giorgio Melchiorre, Sudip Bandyopadhyay, Sandip Halder, Roel Wuyts, Bappaditya Dey
AI Research AI

Small Models Scout Bottleneck Order for Large-Model Data Control

arXiv:2608.14936v1 Announce Type: new Abstract: Small proxy models are commonly used to identify data mixtures for larger-scale training. We ask whether their training trajectories reveal another transferable structure:…

Source: arXiv cs.AI Seungmin Choi, Jiwon Sung, Muhammad Umer, Abhiram Rao Gorle, Guijin Son, Youngjae Yu, John M. Cioffi
AI Research AI

Intent-Driven Situation Tracking for User-Centric Multi-Turn Agents

arXiv:2608.15755v1 Announce Type: new Abstract: User-centric multi-turn agents must act on an evolving task situation shaped by changing user intents, accumulated tool-grounded facts, missing information, and execution…

Source: arXiv cs.AI Meiling Tao, Yiling Tao, Peng Wang
AI Research AI

Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

arXiv:2608.14566v1 Announce Type: new Abstract: Recent work on evaluating the moral competence of large language models (LLMs) has focused primarily on what we call the moral value problem, i.e., whether model outputs…

Source: arXiv cs.AI Aidan Kierans, Ritam Dutt, Kaley Rittichier, Shiri Dori-Hacohen, Avijit Ghosh
AI Research AI

TAHB: A Comprehensive Benchmark for Text-Attributed Hypergraph Learning

arXiv:2608.15055v1 Announce Type: new Abstract: Hypergraphs effectively model higher-order groupwise relationships beyond pairwise interactions, while pretrained language models (PLMs) and large language models (LLMs)…

Source: arXiv cs.AI David Yoon Suk Kang, JungHyun Kim, Juhyun Jeon, Sang-Wook Kim
AI Research AI

Robo-Dopamine 2.0: History-Conditioned and OOD-Aware Process Reward Modeling for Robotic Manipulation

arXiv:2608.15680v1 Announce Type: cross Abstract: Vision-language-action (VLA) models improve robotic manipulation but remain vulnerable to compounding errors, scene changes, and off-trajectory states. Reinforcement…

Source: arXiv cs.AI Yijie Xu, Haopeng Jin, Run Zhou, Shengbang Liu, Sixiang Chen, Hongyang Cheng, Sicheng Hu, Peterson Co, Jinwen Luo, Huajie Tan, Shanghang Zhang
AI Research AI

NebulaVLA: A Dual-Frequency Vision-Language-Action Model With Guide Action for Robotic Manipulation

arXiv:2608.16503v1 Announce Type: cross Abstract: Real-world deployment of Vision-Language-Action (VLA) models is often bottlenecked by efficiency-performance trade-offs, cross-embodiment generalization, and execution…

Source: arXiv cs.AI Cong Zhao, Shuai Tian, Xu Zhang, Baocheng Ni, Xinguo Song, Xueying Sun, Shu Jiang, Shouchang Yang, Bo Tang, Jin Deng, Ge Zhu, YongCheng Wang, Jin Xu, Ri Yang
AI Research AI

Physiological World Models for Human State Transitions

arXiv:2608.15309v1 Announce Type: new Abstract: Continuous multimodal sensing now allows human physiology to be observed throughout daily life rather than only during occasional clinical visits. However, most health…

Source: arXiv cs.AI Chongyang Zhang, Rendong Wang, Hao Zheng, Hanwen Zhang, Yang Liu, Xiaolong Wei, Bin Chong
AI Research AI

ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models

arXiv:2608.15145v1 Announce Type: new Abstract: Large Language Models (LLMs) have been increasingly adopted in Text-to-SQL systems, yet SQL errors remain a major obstacle in real-world Text-to-SQL inference pipelines.…

Source: arXiv cs.AI Xinmei Huang, Jie Song, Peng Li, Fuxin Jiang, Jing Zhang, Tieying Zhang, Jianjun Chen, Chenming Liu, Tao Yang, Maoyin Liu, Wenda Li, Hong Chen, Cuiping Li
AI Research AI

GSBF: Gaussian Splatting for Environment-Aware Beamforming

arXiv:2608.05896v2 Announce Type: replace Abstract: Beamforming plays a key role in multiple-input-multiple-output (MIMO) communication systems. However, conventional beamforming design normally requires accurate…

Source: arXiv cs.AI Yijie Bian, Wei Guo, Zixin Wang, Shenghui Song, Jun Zhang, Khaled B. Letaief
AI Research AI

JarvisBench: Always-on Intelligence Between Humans and Agents

arXiv:2608.14870v1 Announce Type: new Abstract: Long-horizon agents can execute continuously, but human attention remains intermittent and scarce. This creates a bidirectional coordination problem: users may need…

Source: arXiv cs.AI Chen Chen, Zhehuai Chen
AI Research AI

Process-Constituted Intelligence: A Shared Criterion for Humans and Machines

arXiv:2608.16213v1 Announce Type: new Abstract: Intelligence is constituted by \textit{process} (iterative activity through which output emerges), not in the output itself. Generative AI (GenAI) is trained on…

Source: arXiv cs.AI Michael J. Richardson, Ayeh Alhasan, Cassandra Crone, M. Paula Diaz Monfort, Patrick Nalepka, Mark Dras, Rachel W. Kallen, David M. Kaplan
AI Research AI

Bounded Agents: Delegation Security for Multi-Agent AI Systems

arXiv:2608.15888v1 Announce Type: new Abstract: LLM-based agents can act on behalf of a user to access cloud services, call tools, or invoke agents. At session start, the agent's permissions are set but remain static,…

Source: arXiv cs.AI Xabier Muruaga
AI Research AI

Time to Reason: Scalable Neurosymbolic Learning for LTLf via Fuzzy Semantics

arXiv:2608.16443v1 Announce Type: new Abstract: Neurosymbolic (NeSy) Artificial Intelligence aims to integrate Deep Learning (DL) architectures with symbolic reasoning. While initial NeSy approaches have targeted mainly…

Source: arXiv cs.AI Riccardo Andreoni, Andrei Buliga, Alessandro Daniele, Paolo Felli, Chiara Ghidini, Marco Montali, Massimiliano Ronzani
AI Research AI

Dynamic Multi-Byte Prediction With Hierarchical Language Models

arXiv:2608.15454v1 Announce Type: new Abstract: Byte-level hierarchical language models (LMs) have recently emerged as a robust alternative to their popular counterparts that use subword tokenization. However,…

Source: arXiv cs.AI Abraham Toluwase Owodunni, Chibuzor Okocha, Christan Grant, Tomasz Limisiewicz, Sachin Kumar
AI Research AI

Pre-training Visual Dexterity in Simulation

arXiv:2608.15917v1 Announce Type: cross Abstract: Large-scale pre-training has made robot policy fine-tuning increasingly data-efficient, but this progress has largely been driven by datasets and embodiments built…

Source: arXiv cs.AI Sarthak Kamat, Adam Rashid, Satvik Sharma, Aseem Doriwala, Chelsea Finn, Phillip Isola, C. Karen Liu
AI Research AI

ALKEMIE Agent: an autonomous platform for computational materials design

arXiv:2608.15776v1 Announce Type: cross Abstract: Despite the powerful multi-scale modeling methods and high-throughput infrastructures established in the materials community, real material computation workflows remain…

Source: arXiv cs.AI Hongfu Huang, Yuzhe Li, Ao Xu, Bo Liu, Changrui Wang, Kan Tang, Ning Yang, Shengxian Liu, Hanyu Liu, Pengpeng Zhang, Linggang Zhu, Fengkai Liu, Yichen Lu, Tong Zhao, Naihua Miao, Jian Zhou, Zhimei Sun
AI Research AI

Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce

arXiv:2608.14825v1 Announce Type: cross Abstract: Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety literature…

Source: arXiv cs.AI Zeyuan Li (Massachusetts Institute of Technology), Lukas Petersson (Andon Labs), Alessandro Acquisti (Massachusetts Institute of Technology), Michiel A. Bakker (Massachusetts Institute of Technology)
AI Research AI

PDDLCoder: Agentic PDDL Generation for LLM-Assisted Symbolic Planning

arXiv:2608.16637v1 Announce Type: new Abstract: LLMs remain unreliable for long-horizon planning, often generating logically inconsistent or non-applicable plans. Recent hybrid methods instead translate natural language…

Source: arXiv cs.AI Veit Laule, Jiangtao Shuai, Manfred Hauswirth, Sonja Schimmler
AI Research AI

An Evaluation Framework for National AI Regulation

arXiv:2608.15417v1 Announce Type: cross Abstract: Governments use laws, institutions, funding programs and nonbinding guidance to shape how AI is developed and used. Comparing these national approaches is difficult. A…

Source: arXiv cs.AI Kaushik Sanjay Prabhakar, Tarun Adarsh R S, Amal Dhivyan Gregory, Sreeparvathy Sajeev, Utkarsh Tomar, Avyay M Casheekar
AI Research AI

Towards Risk-free AI Agent Deployment

arXiv:2608.16411v1 Announce Type: cross Abstract: LLM-based agents are rapidly moving from research prototypes into the core business processes of organizations, but these agents pose deployment risks to security,…

Source: arXiv cs.AI Yintong Huo, Rangeet Pan, Abhik Roychoudhury
AI Research AI

Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models

arXiv:2606.03988v3 Announce Type: replace Abstract: Vision language models (VLMs) excel at many tasks but still struggle with spatial reasoning when critical information is not directly observable. Many such problems…

Source: arXiv cs.AI Mahtab Bigverdi, Linjie Li, Weikai Huang, Yiming Liu, Jaemin Cho, Tuhin Kundu, Chris Dongjoo Kim, Zelun Luo, Jieyu Zhang, Linda Shapiro, Ranjay Krishna
AI Research AI

Picking the Right Image to Classify: Reliable-Input Selection in Teledermatology

arXiv:2608.16198v1 Announce Type: cross Abstract: Dermatology models face distribution shifts in teledermatology settings, where submitted images differ from the training data in lighting, angle, distance, focus, and…

Source: arXiv cs.AI Fabian Gr\"oger, Marco Weishaupt, Philippe Gottfrois, Simone Lionetti, Linda Wermelinger, Nipun Ranasekara, Ludovic Amruthalingam, Alexander A. Navarini, Marc Pouly
AI Research AI

Automating and Scaling Behavioral Scientific Research on AI Agents

arXiv:2608.10030v2 Announce Type: replace Abstract: As AI agents are increasingly deployed in complex environments, understanding their behaviors becomes critical. Yet behavioral scientific research on AI agents remains…

Source: arXiv cs.AI Soo Yong Lee, Jongha Lee, Jaewan Chun, Hyunjin Hwang, Fanchen Bu, Ziv Ben-Zion, Taekwan Kim, Denny Borsboom, Jaemin Yoo, Kijung Shin
AI Research AI

Academic League of Artificial Intelligence - An Integrative Perspective of Teaching, Research, and Extension

arXiv:2608.13447v2 Announce Type: replace Abstract: Academic leagues have become important mechanisms for promoting extracurricular education and strengthening the integration between universities and society. This…

Source: arXiv cs.AI Alison R. Panisson, Maria Eduarda W. M. Vianna, Italo Firmino da Silva, Heitor Henrique da Silva, Rafaela Fernandes Savaris, Bernardo Pandolfi Costa, Martin Augusto Gagliotti Vigil, Jim Lau, Agenor H…
AI Research AI

CG-GLORE: A Conjugate Gradient-Based Global-Local Regularization Network for Sparse-View CT Reconstruction

arXiv:2608.15246v1 Announce Type: cross Abstract: Sparse-view computed tomography (CT) reduces radiation dose by acquiring fewer projection views, but the resulting inverse problem is highly ill-posed and often produces…

Source: arXiv cs.AI Tran Xuan Hieu Le, Doanh C. Bui, Vu Trung Duong Le, Hoai Luan Pham, Khang Nguyen, Mai K. Nguyen, Tu Bao Ho, Yasuhiko Nakashima
AI Research AI

A Responsible Artificial Intelligence Framework for Groundwater Modeling

arXiv:2608.15657v1 Announce Type: new Abstract: The rapid development and widespread application of artificial intelligence (AI) have sparked intense discussions on how to deploy responsible AI systems in a manner…

Source: arXiv cs.AI Chong Chen, Yulu Zhang, Qingxi Guo, Yihan Liu
AI Research AI

VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

arXiv:2608.15265v1 Announce Type: new Abstract: Constructing an interactive 3D open world from a user query is important. However, existing methods are primarily evaluated on idealized, simple queries, making it…

Source: arXiv cs.AI Yansong Ning, Jingwen Ye, Zhongkai Wu, Yang Sun, Yiqin Zhu, Xingyi Li, Weidong Zhang, Hao Liu
AI Research AI

TRCA: Transition-wise Rubric Credit Assignment for Long-horizon LLM Agents

arXiv:2608.16156v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions…

Source: arXiv cs.AI Huan Zhang, Mingju Chen, Dongxu Zhou, Can Lv, Heng Chang, Sen Cui, Faguo Wu, Shiji Zhou
AI Research AI

MLLM-Guided Semantic Correction for Text-to-Video Generation

arXiv:2608.16513v1 Announce Type: cross Abstract: Recent advances in diffusion models and Transformer architectures have led to significant progress in text-to-video generation. However, these models often suffer from…

Source: arXiv cs.AI Junhao Chen, Zheqi Lv, Keting Yin, Shengyu Zhang, Zhou Zhao, Feiyang Chen, Xinyu Duan, Baoxing Huai, Fei Wu
AI Research AI

Characterising cardiac tissue properties with graph neural networks

arXiv:2608.15843v1 Announce Type: cross Abstract: Characterising electrophysiological properties of cardiac tissue efficiently and accurately from spatially sparse intracardiac measurements is clinically important for…

Source: arXiv cs.AI Ching-En Chiu, Yoo Ri Kim, Magdi Saba, Danilo Mandic, Marta Varela
AI Research AI

A Regulatory Placebo? The Systemic Failure of Mandatory GenAI Labeling

arXiv:2608.16470v1 Announce Type: cross Abstract: We examine the worldwide trend of mandatory labeling of generative artificial intelligence(GenAI) as a reactive, symbolic form of legislation triggered by technological…

Source: arXiv cs.AI Jingyi Chen, Chaofan Bu, Shibo Yan, Xuesong Li
AI Research AI

BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics

arXiv:2608.16211v1 Announce Type: new Abstract: Long-horizon agents are beginning to automate complete workflows that produce code, reports, and research artifacts. Medical imaging workflows are multi-stage and…

Source: arXiv cs.AI Junqi Liu, Yufan He, Yexiao He, Pengfei Guo, Dong Yang, Andriy Myronenko, Can Zhao, Hanrong Ye, Tianhao Qi, Yuyin Zhou, Daguang Xu, Yucheng Tang
AI Research AI

Physics-informed VAE-EVT for Tail Aware Radio Map Prediction

arXiv:2608.15314v1 Announce Type: new Abstract: Ultra-reliable low-latency communication (URLLC) requires precise identification of spatial regions where the signal-to-noise ratio (SNR) falls below an outage threshold.…

Source: arXiv cs.AI Amanda Sheron Gamage, Niloofar Mehrnia, James Gross
AI Research AI

Position: AI Governance Needs ISO-like Interoperability Protocols, Not Just Laws

arXiv:2608.14568v1 Announce Type: new Abstract: As Artificial Intelligence (AI) systems become deeply integrated into critical global infrastructure, the urgency for robust governance frameworks has intensified.…

Source: arXiv cs.AI Azmine Toushik Wasi, Mst Rafia Islam, Mahfuz Ahmed Anik, Taki Hasan Rafi, Md Manjurul Ahsan, Dong-Kyu Chae
AI Research AI

Decoupled Temporal Encoding for Generative Recommendation

arXiv:2608.16274v1 Announce Type: cross Abstract: Positional encoding is a fundamental component of Transformer-based generative recommendation models, where user histories are modeled as autoregressive item sequences.…

Source: arXiv cs.AI Pengfei Jia, Jingjian Wang, Jingmao Li, Ge Zhang, Feng Shi
AI Research AI

Hierarchical Adaptive Feature Refinement Network for VHR Remote Sensing Image Segmentation

arXiv:2608.15647v1 Announce Type: cross Abstract: Semantic segmentation of very-high-resolution (VHR) remote sensing imagery increasingly benefits from strong pretrained hierarchical encoders, yet exploiting their…

Source: arXiv cs.AI Shuaishuai Cao, Meng Tang, Shuwei Peng, Xuan Liu, Min Huang, Jie Chen, Jiacheng Niu, Yong Chen, Edore Akpokodje, Hui Lin
AI Research AI

Drive, Pack, Fly: The Travelling Thief Problem with Drone

arXiv:2608.16435v1 Announce Type: new Abstract: In collection operations, accumulating payload progressively slows the vehicle, imposing a cumulative penalty on routing efficiency. An onboard drone can offset this…

Source: arXiv cs.AI Kabir Murjani, Abhay Sobhanan
AI Research AI

VCE-Skill: Enhancing Skill Self-Evolution with Version-Change Experience

arXiv:2608.16544v1 Announce Type: cross Abstract: Agents increasingly rely on reusable skills to encode task knowledge, tool-use procedures, and validation rules. Existing skill self-evolution methods primarily revise…

Source: arXiv cs.AI Jianming Chen, Xuanbin Ye, Yawen Wang, Junjie Wang, Qing Wang, Fanjiang XU
AI Research AI

Learning Agent Execution for KV-Cache Management in Agentic Serving

arXiv:2608.14624v1 Announce Type: new Abstract: Multi-agent LLM systems have emerged as an important deployment paradigm for AI services, where each user request is decomposed into a sequence of specialized agents.…

Source: arXiv cs.AI Rui Zhang, Chaeeun Kim, Shaoting Feng, Kuntai Du, Yuhan Liu, Yi Zhong, Cheng-Wei Ching, Junchen Jiang, Liting Hu
AI Research AI

A survey of AI-generated voices and their detection

arXiv:2608.15411v1 Announce Type: new Abstract: The ability of artificial intelligence (AI) models to generate highly realistic human voices has advanced rapidly. These technologies power accessibility tools, virtual…

Source: arXiv cs.AI Chengzhe Sun, Tianle Yang, Siwei Lyu
AI Research AI

WeSCE: A Benchmark for Measuring Security Drift in LLM-Driven Code Editing

arXiv:2608.15092v1 Announce Type: cross Abstract: In this work, we introduce WeSCE, a benchmark for quantifying security drift in code editing under weak-security constraints, where tasks specify only functional…

Source: arXiv cs.AI Zhiyu Zhang, Tingyue Wen, Senke Sun, Dengxiang Liang, Enhao Huang
AI Research AI

Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement

arXiv:2608.16628v1 Announce Type: new Abstract: Modern Multimodal Retrieval-Augmented Generation (M-RAG) systems are fundamentally limited by the binary connectivity paradigm of traditional simple graphs, which fails to…

Source: arXiv cs.AI Shenao Chen, Yidan Xu, Xiangmin Han, Rundong Xue, Duanpo Wu, Yuhan Gao, Chenggang Yan, Yue Gao
AI Research AI

Cross-Domain Industrial Fault Detection by Causal Mechanism Monitoring

arXiv:2608.14666v1 Announce Type: new Abstract: Unsupervised fault detection in industrial systems is dominated by reconstruction based methods that monitor individual sensor marginal distributions. This misses coupling…

Source: arXiv cs.AI Dhiraj Neupane, Mohamed Reda Bouadjenek, Richard Dazeley, Sunil Aryal
AI Research AI

SysEvolve: An AI-native, safe, autonomous adversarial attack-defense co-evolutionary system

arXiv:2608.15012v1 Announce Type: cross Abstract: The rapid advancement of large language models (LLMs) has created a growing asymmetry in cybersecurity, where attack accelerates toward autonomous execution while…

Source: arXiv cs.AI Yuhan Meng, Shaofei Li, Jionghao Huang, Jiandong Jin, Puyi Wang, Hanlin Jiang, Anis Yusof, Peng Jiang, Zhenkai Liang, Yao Guo, Ding Li
AI Research AI

Synchronized Logit Steering: Real-world Steganography

arXiv:2608.14697v1 Announce Type: new Abstract: Steganography in large language models offers a way to embed hidden messages within natural-sounding text. Existing token and logit-level methods typically require the…

Source: arXiv cs.AI Andrew Rufail, Aadi Dash, Onir Narahari, Ethan Mui, Mahi Gajare, Prakhar Tiwari, Shrija Makapothula, Nick Cui
AI Research AI

Reasoning-supported Robustness Validation of Automotive E/E Components

arXiv:2608.16421v1 Announce Type: new Abstract: This paper presents an ontology-supported approach to tackle the complexity of the Robustness Validation (RV) process of automotive electrical/electronic (E/E) components.…

Source: arXiv cs.AI Jan Novacek, Alexander Viehl, Oliver Bringmann, Wolfgang Rosenstiel
AI Research AI

Orbital AI Computing: Carbon Tradeoffs Across Satellite Scale

arXiv:2608.14557v1 Announce Type: cross Abstract: Low Earth Orbit (LEO) computing is emerging for low-latency, globally distributed AI services, enabled by advances in satellite constellations and reusable launch…

Source: arXiv cs.AI Nisha Sarwar, Lei Jiang, Fan Chen
AI Research AI

Graph Neural Assisted Actor-Critic for Latency-Efficient Edge Vision System

arXiv:2608.16142v1 Announce Type: cross Abstract: UAV on-board vision systems are widely used for different activities, including monitoring in no-fly zones. In this case, the vision-equipped UAV streams a video to a…

Source: arXiv cs.AI Alam Noor, Luis Almeida, Kai Li, Jiyan Wu, Miguel Guti\'errez Gait\'an, Eduardo Tovar
AI Research AI

TDD-Agent: Test-Driven Reasoning for Code Generation

arXiv:2608.16742v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable progress in code generation, yet ensuring correctness in complex, repository-level tasks remains challenging.…

Source: arXiv cs.AI Hongyue Yu, Kefan Li, Jiakun Li, Hongzheng Chai, Yuan Yuan, Rui He, Junyi Wei
AI Research AI

DeepInsight II: One Trace from Benchmark to Robot

arXiv:2608.16556v1 Announce Type: new Abstract: Across a Physical AI stack, evaluation maturity is inversely aligned with deployment risk: foundation models enjoy mature, standardized harnesses, while the embodied…

Source: arXiv cs.AI Siyi Li, Yuchen Kang, Wuliang Wang, Zhengjie Zhang, Jiangpin Liu, Jianhao Yao, Jie Chen
AI Research AI

LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures

arXiv:2608.15242v1 Announce Type: new Abstract: When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers…

Source: arXiv cs.AI Yunfei Zhang, Boyu Feng, Changhua Pei, Zexin Wang, Zhihuang Peng, Xinlong Liu, Hengyue Jiang, Difeng Ma, Jiayi Zhang, Yongzhou Yao, Yanan Zhao, Fei Sun, Yintong Huo, Zhaoyang Liu, Jingjing Li, Gaogan…
AI Research AI

ReasonCast: Agentic Demand Forecasting with Selective Semantic Reasoning

arXiv:2608.15291v1 Announce Type: new Abstract: Demand forecasting increasingly requires combining two complementary sources of information: historical sales reveal recurring numerical dynamics, while future promotions,…

Source: arXiv cs.AI Ziyue Yang, Chaolin Xu, Yijing Wang, Tiankai Gu, Hui Yang, Yanhong Lin, Kaiyuan Liu, Fei Xiao
AI Research AI

ARENA: Automated Red-Teaming for Large Audio Language Models

arXiv:2608.15578v1 Announce Type: cross Abstract: Large audio-language models (LALMs) make it possible to interact with language models through speech, music, and environmental sound, but they also introduce a safety…

Source: arXiv cs.AI Jiaming He, Zhicong Huang, Tian Jin, Zhen Sun, Cheng Hong, Yi Yu, Wenbo Jiang, Xudong Jiang
AI Research AI

Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents

arXiv:2608.08601v2 Announce Type: replace Abstract: To anticipate socio-technical risks from AI agents, organizations need taxonomies to classify them. However, existing AI risk taxonomies focus on broad risks and do…

Source: arXiv cs.AI Gabriele La Malfa, Lakmal Meegahapola, Edyta Bogucka, Jie M. Zhang, Michael Luck, Elizabeth Black, Daniele Quercia
AI Research AI

Evaluating Agentic Code Repair Capabilities in Distributed Systems

arXiv:2608.14863v1 Announce Type: cross Abstract: LLM-based coding agents have advanced rapidly on single-process SWE tasks, with frontier models now clustering in the high-70s on SWE-bench Verified. Distributed-system…

Source: arXiv cs.AI Yibo Yan, Huijuan Wang, Junzhou He, Yizhuo Liang, Shaoyu Wang, Huanchen Sun, Seo Jin Park
AI Research AI

A Machine-Learned Comorbidity Index

arXiv:2606.17450v2 Announce Type: replace Abstract: Traditional comorbidity scores (e.g., Charlson and Elixhauser) are widely used for risk adjustment and patient stratification, but they have two key limitations: (i)…

Source: arXiv cs.AI Suleman Baloch, Kishlay Jha, Alberto M. Segre, Philip M. Polgreen, Bijaya Adhikari
AI Research AI

DriveCache: Action-Aware Caching for Driving World Model Inference

arXiv:2608.16354v1 Announce Type: new Abstract: Driving video generation models support autonomous-driving development by predicting controllable future scenes for simulation, planning evaluation, and offline data…

Source: arXiv cs.AI Jianchun Yang, Jian Liang, Xianda Guo, Pinhan Fu, Yanlun Peng, Conglang Zhang, Wenke Huang, Mang Ye
AI Research AI

Who Leads Now? Token-Level Modality Arbitration for Chart-to-Code Generation

arXiv:2608.15510v1 Announce Type: new Abstract: Chart-to-code generation requires a model to read the fine-grained visual details of a chart and write executable code that reproduces it. Existing chart-to-code methods…

Source: arXiv cs.AI Qinghao Fu, Yarong Wang, Shunlei Ning, Yilin Wang, Shunwen Bai, Xinda Wang, Jiaotuan Wang, Yinan Nie, Wei Zhou
AI Research AI

A concentration result for multilayer feedforward neural networks

arXiv:2608.15335v1 Announce Type: new Abstract: We consider for an arbitrary fixed $\rho$ and for each positive integer $n$ a multilayer feedforward artificial neural network with $\rho$ layers, $n$ neurons in the first…

Source: arXiv cs.AI Vera Koponen
AI Research AI

FabriMAE I Trust Myself? Self-Evaluating VLA Action Generation with Markov Attention Entropy

arXiv:2608.16697v1 Announce Type: new Abstract: Vision-Language-Action models (VLAs) integrate visual perception, language instruction, and action generation into end-to-end policies across heterogeneous architectures.…

Source: arXiv cs.AI Aniri, Chen Yilin, Jinhe Bi, Junfei Guo, Donglai Ran, Xu Bian, Zengjie Jin, Yujun Wang, Yijun Tian, Volker Tresp, Fei Shen, Tat-Seng Chua, Yunpu Ma
AI Research AI

DiffImaginE: Imagine to Verify Entity Types with Diffusion

arXiv:2608.03025v4 Announce Type: replace Abstract: Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence.…

Source: arXiv cs.AI Feng Zhang, Feiyu Han, Rongxin Yang, Yang Liu, Yancheng Chen, Rui Wang, Yingguang Yang, Tian Xueyun, Chongyang Zhang, Hao Zheng, Xu Kefu, Congjing Ran, Fuhai Chen, Bin Chong
AI Research AI

DeCo-MIL: Debiased Counterfactual Reasoning for Long-Tailed Whole Slide Image Analysis

arXiv:2608.14719v1 Announce Type: cross Abstract: Multiple instance learning (MIL) is widely used for weakly supervised whole slide image (WSI) analysis. However, under long-tailed distributions, MIL-based WSI analysis…

Source: arXiv cs.AI Xiaoxiao Li, Xitong Ling, Jiawen Li, Weiming Chen, Zhenyang Cai, Xidong Wang, Tian Guan, Benyou Wang, Yonghong He
AI Research AI

ParaTempo: Efficient Parallel Reasoning via Temporal Confidence

arXiv:2608.16425v1 Announce Type: new Abstract: Parallel reasoning improves the accuracy and robustness of large reasoning models by exploring multiple solution paths, but its computational cost grows with reasoning…

Source: arXiv cs.AI Xuteng Zhang, Wenhao Zeng, Xiaodong Gu, Chao Hu, Haotian Lin, Yuling Shi, Min Wang, Beijun Shen
AI Research AI

Geometric Self-Supervised Pre-training for Neural Combinatorial Optimization

arXiv:2608.00270v2 Announce Type: replace Abstract: Neural Combinatorial Optimization (NCO) techniques have emerged as a highly efficient alternative to traditional exact algorithms for solving routing problems such as…

Source: arXiv cs.AI David Aguado, Daniel Fuertes, Carlos R. del-Blanco, Fernando Jaureguizar
AI Research AI

A Human-Centred Approach to Benchmarking LLMs for Parenting Advice

arXiv:2608.14622v1 Announce Type: new Abstract: People are increasingly using large language models (LLMs) to seek advice, including for parenting. Parenting is a critical and socially sensitive domain. Thus, evaluating…

Source: arXiv cs.AI Yunke Zhao, Isobel Voysey, Alastair van Heerden, Rob Hughes, Jun Zhao
AI Research AI

A Policy Algebra for Trust-Preserving Agentic AI Execution

arXiv:2608.16402v1 Announce Type: new Abstract: Large language model-based agentic frameworks primarily optimize capability: whether an agent can reason, retrieve information, call tools, delegate work, and complete a…

Source: arXiv cs.AI Bhaskar Tripathi, Anurag Kumar, Ramendra Kumar, Bhavesh Gadhe
AI Research AI

JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills

arXiv:2608.16465v1 Announce Type: new Abstract: Automated red-teaming has produced a growing collection of attack strategies, yet they typically remain scattered across prompts and workflows, making them difficult to…

Source: arXiv cs.AI Xiaoyu Wen, Jiajia Li, Zhida He, Peng Yu, Chenxu Wang, Han Qi, Ziyuan Zhou, Cheng Jin, Ying Wen, Xingcheng Xu, Shuyue Hu, Tianhang Zheng, Chaochao Lu, Qiaosheng Zhang
AI Research AI

The Fragility of Strategic Thinking in Large Language Models

arXiv:2510.10813v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly applied to domains that require reasoning about other agents' behavior, such as negotiation, policy design, and market…

Source: arXiv cs.AI Enric Junque de Fortuny, Veronica Roberta Cappelli
AI Research AI

Position: Medical AI Neglects Real Treatment Outcomes

arXiv:2608.14598v1 Announce Type: new Abstract: Medical AI has rapidly improved its ability to perform diagnostic and prognostic tasks that lead to treatment decisions. But understanding of treatment itself is still…

Source: arXiv cs.AI Shiva Kaul, Anjum Khurshid
AI Research AI

FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification

arXiv:2604.04074v4 Announce Type: replace Abstract: Large language model (LLM)-based reviewing systems typically assess manuscripts in isolation, leaving literature- and code-dependent claims difficult to verify. We…

Source: arXiv cs.AI Ling Yue, Chaoqian Ouyang, Hang Xu, Ruijun Huang, Yuchen Liu, Libin Zheng, Wei Liu, Shaowu Pan, Shimin Di, Min-Ling Zhang
AI Research AI

SCOPE: Score-Isolated Agentic Optimization for Video World Models

arXiv:2608.15043v1 Announce Type: new Abstract: Video world models are increasingly used as simulators for planning and embodied decision making, yet improving them at inference time introduces a subtle evaluation…

Source: arXiv cs.AI Yuhua Jiang, Jiaming Wang, Qingbin Liu, Feifei Gao
AI Research AI

Unified Pedestrian Path Prediction Using Inverse Reinforcement Learning

arXiv:2608.15929v1 Announce Type: new Abstract: Pedestrian path prediction is crucial for enhancing the safety of autonomous vehicles and advanced driver-assistance systems. Previous studies explored different…

Source: arXiv cs.AI \v{S}imon Sukup, Ariyan Bighashdel, Pavol Jancura
AI Research AI

Advanced modelling and data analytics in aviation

arXiv:2608.14746v1 Announce Type: new Abstract: The aviation industry characterized by its stringent safety standards has seen a growing need for innovative approaches to enhance safety measures. Despite the vast…

Source: arXiv cs.AI Aziida Nanyonga