Skip to content
TILens What matters today in tech
Theme

Topic · Edition

AI

533 items on 19 Aug 2026
AI Tools GitHub AI

langchain-ai/langchain: langchain-core==1.6.0

Changes since langchain-core==1.5.6 release(core): 1.6.0 (#39760) fix(core): resolve postponed annotations in StructuredTool._injected_args_keys (#39602) feat(core): add standard model exception types (#39538)…

Source: LangChain Releases github-actions[bot]
AI Tools GitHub AI

pytorch/pytorch: v2.14.0-rc5

[pytorch][PR] Migrate fastAtomicAdd to headeronly (#193176) (#193604)…

Source: PyTorch Releases pytorchbot
AI Research AI

Towards Zero-Shot Task Transfer with Neurosymbolic World Models

arXiv:2608.17959v1 Announce Type: cross Abstract: State-of-the-art model-based reinforcement learning methods learn neural world models that allow policy improvement by planning in a latent space, without assumptions on…

Source: arXiv cs.LG Isidoro Tamassia, Lennert De Smet, Giuseppe Marra
AI Research AI

CORAM: Coherent Orthogonal Rotation for Model Merging

arXiv:2608.17366v1 Announce Type: new Abstract: Merging finetuned models combines specialized capabilities without joint training or access to the original data. Most methods operate by linear arithmetic in Euclidean…

Source: arXiv cs.LG Xinyi Sui, Ziran Liu, Nam Ling, Wei Wang, Wei Jiang
AI Research AI

Latent Order Bandits

arXiv:2605.07304v2 Announce Type: replace Abstract: Bandit algorithms solve diverse sequential decision-making problems, but are often too sample-inefficient for from-scratch personalization. To substantially reduce…

Source: arXiv cs.LG Emil Carlsson, Newton Mwai, Fredrik D. Johansson
AI Research AI

Abra: Scaling Diffusion Image Training

arXiv:2608.17286v1 Announce Type: new Abstract: Compute-optimal scaling laws guide the training of frontier language models yet remain largely unexplored for visual generation. We present a systematic scaling law study…

Source: arXiv cs.LG Kyle Chickering, Wei-An Lin, Swayam Bhanded, Dan Saunders, Akshat Tripathi, Jiaming Song, Shyamal Buch, Xinchen Yan
AI Research AI

Belayer: Efficient Fault Tolerance for LLM Agentic RL Training

arXiv:2608.14635v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents are increasingly trained with reinforcement learning in long-horizon, sandboxed environments. Unlike conventional RL, agentic…

Source: arXiv cs.LG Jiecheng Zhou, Qinghao Hu, Peng Sun, Xingcheng Zhang, Weiming Zhang
AI Research AI

A Residual Learning Approach for Unsteady Aerodynamic Load Prediction

arXiv:2608.17894v1 Announce Type: cross Abstract: This paper investigates the feasibility of using residual learning to improve unsteady aerodynamic load prediction for aeroelastic applications. The machine learning…

Source: arXiv cs.LG Divya Sanghi, Carlos E. S. Cesnik
AI Research AI

Position: Fairness Failure in Generative Models is an Evaluation Problem

arXiv:2608.16974v1 Announce Type: new Abstract: Despite groundbreaking advancements in generative models during the last decade, concerns about their lack of fairness, reinforcing societal inequalities and harming…

Source: arXiv cs.LG Mariia Vladimirova, Jean-Yves Franceschi, Thibaut Issenhuth
AI Research AI

On Stability in Optimistic Bilevel Optimization

arXiv:2408.13323v3 Announce Type: replace-cross Abstract: Solutions of bilevel optimization problems tend to suffer from instability under changes to problem data. In the optimistic setting, we construct a lifted…

Source: arXiv cs.LG Johannes O. Royset
AI Research AI

Comprehensive framework for evaluation of deep neural networks in detection and quantification of lymphoma from PET/CT images: clinical insights, pitfalls, and observer agreement analyses

arXiv:2311.09614v5 Announce Type: replace-cross Abstract: This study addresses critical gaps in automated lymphoma segmentation from PET/CT images, focusing on issues often overlooked in existing literature. While deep…

Source: arXiv cs.LG Shadab Ahamed, Yixi Xu, Sara Kurkowska, Claire Gowdy, Joo H. O, Ingrid Bloise, Don Wilson, Patrick Martineau, Fran\c{c}ois B\'enard, Fereshteh Yousefirizi, Rahul Dodhia, Juan M. Lavista, William B. W…
AI Research AI

TabularQGAN: A quantum generative model for tabular data synthesis

arXiv:2505.22533v2 Announce Type: replace Abstract: In this paper, we introduce a novel quantum generative model for synthesizing tabular data. Synthetic data is valuable in scenarios where real-world data is scarce or…

Source: arXiv cs.LG Pallavi Bhardwaj, Caitlin Jones, Lasse Dierich, Aleksandar Vu\v{c}kovi\'c
AI Research AI

Recirculation

arXiv:2608.17981v1 Announce Type: new Abstract: We describe an inference-time architectural enhancement for off-the-shelf foundation models that markedly reduces perplexity and boosts accuracy across generation and…

Source: arXiv cs.LG Michael C. Mozer, Shoaib Ahmed Siddiqui, Danny Sawyer, Sunny Sanyal, Rosanne Liu
AI Research AI

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

arXiv:2608.17253v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest successes still depend…

Source: arXiv cs.LG Yunhao Yang, Yuexin Bian, Yunjie Tian, Di Fu, Tianjin Huang, Yuanyuan Shi, Ziang Xiao, Nuno Vasconcelos, Yijiang Li
AI Research AI

Efficient Resource Optimization for Split Federated Learning

arXiv:2608.17849v1 Announce Type: new Abstract: Split federated learning (SFL) has emerged as a powerful paradigm for model training at the edge. However, SFL inherently involves discrete decision variables for model…

Source: arXiv cs.LG Wei Wei, Xianhao Chen
AI Research AI

When AI Designs AI: Innovation or Imitation?

arXiv:2608.17471v1 Announce Type: cross Abstract: Recent advances in LLM agents have made them increasingly capable of designing methods for complex AI tasks. This raises two central questions about agent-designed…

Source: arXiv cs.LG Yikang Yang, Zhengxin Yang, Luzhou Peng, Minghao Luo, Yanqi Kan, Wanling Gao, Jianfeng Zhan
AI Research AI

From Diffusion to Flow: Efficient Motion Generation in MotionGPT3

arXiv:2603.26747v3 Announce Type: replace-cross Abstract: Recent text-driven motion generation methods span both discrete token-based approaches and continuous-latent formulations. MotionGPT3 exemplifies the latter…

Source: arXiv cs.LG Jaymin Bhan, JiHong Jeon, SangYeop Jeong
AI Research AI

Memory by Design: Probabilistic Sequence Layers

arXiv:2605.31163v3 Announce Type: replace-cross Abstract: We introduce the \emph{design-model framework}: a way to derive efficient recurrent sequence maps from explicit assumptions about memory. A design model writes…

Source: arXiv cs.LG Matthew Dowling, Hyungju Jeon, Cristina Savin, Il Memming Park
AI Research AI

Fourth-Moment Geometry of Rademacher Sums

arXiv:2608.17802v1 Announce Type: new Abstract: Let $\varepsilon_1,\ldots,\varepsilon_n$ be independent Rademacher signs and let $a=(a_1,\ldots,a_n)\in\R^n$ satisfy the normalization below. For the normalized Rademacher…

Source: arXiv cs.LG Peigan Gao, Jian Qian
AI Research AI

TiMi: Empower Time Series Transformers with Multimodal Mixture of Experts

arXiv:2602.21693v2 Announce Type: replace Abstract: Multimodal time series forecasting has garnered significant attention for its potential to provide more accurate predictions than traditional single-modality models by…

Source: arXiv cs.LG Jiafeng Lin, Yuxuan Wang, Huakun Luo, Jianmin Wang, Zhongyi Pei
AI Research AI

The Standard Interpretable Model: A general theory of interpretable machine learning to deductively design interpretable methods using Lagrangian mechanics

arXiv:2606.12289v2 Announce Type: replace Abstract: As Artificial Intelligence models grow in complexity, interpretability has become an indispensable tool for understanding, debugging, and controlling their…

Source: arXiv cs.LG Pietro Barbiero, Giovanni De Felice, Mateo Espinosa Zarlenga, Francesco Giannini, Filippo Bonchi, Mateja Jamnik, Giuseppe Marra, Ruggero Noris
AI Research AI

The Authenticity Gap in Human Evaluation

arXiv:2205.11930v3 Announce Type: replace-cross Abstract: Human ratings are the gold standard in NLG evaluation. The standard protocol is to collect ratings of generated text, average across annotators, and rank NLG…

Source: arXiv cs.LG Kawin Ethayarajh, Dan Jurafsky
AI Research AI

Prism-GRPO: Faster VLA Policy Optimization via Splitting Same-outcome Groups

arXiv:2608.17423v1 Announce Type: cross Abstract: GRPO is increasingly used for reinforcement learning of vision-language-action (VLA) policies because, unlike PPO, it does not require training a critic. This…

Source: arXiv cs.LG Zeyun Deng, Yuzhe Lu, Yawei Wang, Linbo Liu, Qing Ping, Han Ding, Guande Wu, Panpan Xu, Jun Huan
AI Research AI

Constitutional Midtraining: Content Presence Drives Alignment Gains

arXiv:2607.26654v3 Announce Type: replace-cross Abstract: Post-training alignment is often shallow, eroding under fine-tuning. It remains untested as to whether constitutional midtraining interventions can produce…

Source: arXiv cs.LG Desiree Cho, Cameron Tice, Bernie Hogan, Hunar Batra, Puria Radmard, Jun Zhao, Nigel Shadbolt
AI Research AI

Inverse Problems for Partial Differential Equations with Jump Discontinuities in Coefficients via Two-Stage Physics-Informed Deep Learning and Statistical Mixture Models

arXiv:2510.14656v3 Announce Type: replace-cross Abstract: This work proposes a two-stage physics-informed deep learning framework that combines neural-network-based sampling with statistical inference and constrained…

Source: arXiv cs.LG Zhikun Zhang, Guanyu Pan, Xiangjun Wang, Yong Xu, Guangtao Zhang
AI Research AI

Community Concealment from Graph Neural Networks

arXiv:2602.12250v2 Announce Type: replace Abstract: Graph neural networks (GNNs) enable powerful unsupervised learning of communities. However, such inference may inadvertently expose sensitive group structures,…

Source: arXiv cs.LG Dalyapraz Manatova, Pablo Moriano, L. Jean Camp
AI Research AI

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

arXiv:2608.17310v1 Announce Type: new Abstract: Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and…

Source: arXiv cs.LG Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee
AI Research AI

FairNVT: Fair Classification via Noise Injection in Vision Transformers

arXiv:2604.16780v2 Announce Type: replace-cross Abstract: This paper presents FairNVT, a lightweight debiasing framework for pretrained transformer-based encoders that improves prediction fairness while preserving task…

Source: arXiv cs.LG Qiaoyue Tang, Sepidehsadat Hosseini, Mengyao Zhai, Thibaut Durand, Greg Mori
AI Research AI

Large Language Models: A Mathematical Formulation

arXiv:2601.22170v2 Announce Type: replace-cross Abstract: Large language models (LLMs) process and predict sequences containing text to answer questions, and address tasks including document summarization, providing…

Source: arXiv cs.LG Ricardo Baptista, Andrew Stuart, Son Tran
AI Research AI

Likelihood Hacking in Probabilistic Program Synthesis

arXiv:2603.24126v2 Announce Type: replace Abstract: When language models are trained by reinforcement learning (RL) to write probabilistic programs, they can artificially inflate their marginal-likelihood reward by…

Source: arXiv cs.LG Jacek Karwowski, Younesse Kaddar, Zihuiwen Ye, Esmeralda S. Whitammer, Sam Staton
AI Research AI

Spectrally Safe Neural Operator Warm-Starts for Large-Scale Newton Solvers

arXiv:2606.21828v2 Announce Type: replace-cross Abstract: Neural operators are increasingly used to warm-start Newton solvers for nonlinear PDEs, on the premise that a low test error places the initial guess inside the…

Source: arXiv cs.LG Jaemin Oh, Youngkyu Lee, Jerome Darbon, George Em Karniadakis
AI Research AI

Population Health-Based Machine Learning Reveals Associations Between Psychosocial Factors and Chronic Kidney Disease

arXiv:2608.17174v1 Announce Type: new Abstract: Chronic kidney disease (CKD) progresses silently and severely undermines quality of life, making early detection critical for improving patient outcomes. We present a…

Source: arXiv cs.LG Md. Atik Shams, David Eisenberg, Sumaiya Fatema, Asma Sultana, D. M Hasibul Islam, Junnatul Mawa, Anindita Datta, Nafiya Ahmed, Danastan Tasaouf Mridula, SK. Sazid Mahmud, Simon Bin Akter, Tanjila He…
AI Research AI

BRo-JEPA: Learning Modular Transformations in Latent Space

arXiv:2606.01372v2 Announce Type: replace Abstract: Can neural networks learn algebraic rules from visual inputs, or do they merely fit observed patterns? We study this question using MNIST (or EMNIST letters) as states…

Source: arXiv cs.LG Divyansh Jha, Yuanfang Xie, Brennen Yu, Varan Mehra
AI Research AI

Which CS1 Students Will Fail? Identifying Digital Markers from Learning Analytics in Computer Systems and Architecture Using Weighted Academic Momentum and Interaction Logs

arXiv:2608.16914v1 Announce Type: cross Abstract: Digital learning platforms generate rich behavioural traces (digital markers) that offer the potential to identify struggling students early. This paper investigates…

Source: arXiv cs.LG Lighton Phiri, Mutune Chaibela, Ivy Chisha, David Pungwa, Danny Siabbaba, Bydon Simukoko
AI Research AI

Debate Training Reduces Reward Hacking in RLAIF

arXiv:2608.17776v1 Announce Type: new Abstract: We demonstrate that RL finetuning an LLM using debate, a two-player adversarial game between a generator and a critic adjudicated by a weaker LLM judge, reduces reward…

Source: arXiv cs.LG Zachary Kenton, Lili Janzer, Rory Greig, Tian Huey Teh, Kirill Tyshchuk, Jonah Brown-Cohen, Harri Edwards, Senthooran Rajamanoharan, Noah Y. Siegel, Natasha Jaques, Rohin Shah
AI Research AI

Neural Operator-Based Nonlinear Nudging for Chaotic Dynamical Systems

arXiv:2508.05778v2 Announce Type: replace Abstract: Nudging is an empirical data assimilation technique that incorporates an observation-driven control term into the model dynamics. The trajectory of the nudged system…

Source: arXiv cs.LG Jaemin Oh, Jinsil Lee, Youngjoon Hong
AI Research AI

Elimination Geometry

arXiv:2608.17646v1 Announce Type: new Abstract: This monograph develops elimination geometry (EG), a typed, native-loss, audit-oriented framework for studying when locally optimal objects can be realized by a shared…

Source: arXiv cs.LG Mian Huang, Xueqin Wang
AI Research AI

Scientific Machine Learning of Chaotic Systems Learns Reduced-Order Equations for Neural Populations

arXiv:2507.03631v5 Announce Type: replace Abstract: Extracting interpretable mathematical models from complex dynamical systems is difficult, especially for chaotic dynamics observed with noisy experimental data. We…

Source: arXiv cs.LG Anthony G. Chesebro, David Hofmann, Vaibhav Dixit, Earl K. Miller, Richard H. Granger, Alan Edelman, Christopher V. Rackauckas, Lilianne R. Mujica-Parodi, Helmut H. Strey
AI Research AI

Attention Flows: Tracing LLM Conceptual Engagement via Story Summaries

arXiv:2604.06416v2 Announce Type: replace-cross Abstract: Although LLM context lengths have grown, there is evidence that their ability to integrate information across long-form texts has not kept pace. We evaluate one…

Source: arXiv cs.LG Rebecca M. M. Hicke, Sil Hamilton, David Mimno, Ross Deans Kristensen-McLachlan
AI Research AI

Low-dimensional topology of deep neural networks

arXiv:2606.31856v2 Announce Type: replace Abstract: We study layered models, including feedforward networks, ResNets, and transformers, by limiting each layer to a width of $d = 3$, i.e., $\mathbb{R}^3$ as…

Source: arXiv cs.LG Junyu Ren, Lek-Heng Lim
AI Research AI

General Semantic Knowledge Infusion for Spatio-Temporal Traffic Forecasting

arXiv:2608.17440v1 Announce Type: new Abstract: Although Graph Neural Networks (GNNs) have made significant advances in spatio-temporal traffic forecasting, their performance is limited when they rely solely on sensor…

Source: arXiv cs.LG Mattis thor Straten, Yannick Wolker, Steffen Strohm, Prathvish Mithare, Ralf Krestel, Matthias Renz
AI Research AI

Conformal Prediction for Molecular Properties under Label Shift

arXiv:2608.17678v1 Announce Type: new Abstract: Drug discovery and development underpins healthcare but remains costly and failure-prone. A critical bottleneck lies in predicting molecular properties such as solubility,…

Source: arXiv cs.LG Hyeonsu Lee, Juyeon Kim, Erkhembayar Jadamba, Seungjin Choi, Hyunjin Shin
AI Research AI

The Null Token Knows: Reducing Message-Free Hallucination in ASR and NMT

arXiv:2608.15940v2 Announce Type: replace-cross Abstract: Modern encoder-decoder systems can produce fluent text even when their input contains no recoverable message. We study this failure in ASR and NMT through the…

Source: arXiv cs.LG Kirill Borodin, Vasiliy Kudryavtsev, Ivan Viakhirev, Grach Mkrtchian
AI Research AI

Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents

arXiv:2608.17524v1 Announce Type: new Abstract: This preliminary paper outlines a planned evaluation benchmark for Explainable Reinforcement Learning (XRL) methods. Current evaluations rely on functionally-grounded…

Source: arXiv cs.LG Ram Rachum, Yotam Amitai, B\'alint Gyevn\'ar, Reuth Mirsky, Cameron Allen
AI Research AI

Task Specialization Fine-Tuning for Contextual Reinforcement Learning

arXiv:2608.17180v1 Announce Type: new Abstract: Contextual Reinforcement Learning (CRL) seeks to generalize classical RL by maximizing task coverage across a context space of related tasks. While prior works often train…

Source: arXiv cs.LG Jianan Zhou, Jung-Hoon Cho, Tianyue Zhou, Han Zheng, Jie Zhang, Roy Dong, Yining Ma, Cathy Wu
AI Research AI

Non-KKT Accumulation in Entropic Mirror Descent

arXiv:2608.01658v3 Announce Type: replace-cross Abstract: For mirror descent generated by a Legendre kernel, perhaps one of the most basic question in optimization is this: must every accumulation point of a bounded…

Source: arXiv cs.LG Kuangyu Ding, Kim-Chuan Toh
AI Research AI

Spikformer V2: Join the High Accuracy Club on ImageNet with an SNN Ticket

arXiv:2401.02020v2 Announce Type: replace-cross Abstract: Spiking Neural Networks (SNNs), known for their biologically plausible architecture, face the challenge of limited performance. The self-attention mechanism,…

Source: arXiv cs.LG Zhaokun Zhou, Yijie Lu, Kaiwei Che, Wei Fang, Keyu Tian, Qihao Peng, Yuesheng Zhu, Shuicheng Yan, Yonghong Tian, Li Yuan
AI Research AI

When to Review: Spaced Repetition for Continual Pre-Training of Language Models

arXiv:2608.17530v1 Announce Type: cross Abstract: Continual pre-training of large language models must acquire new information without erasing old knowledge. Existing replay methods often choose a global old/new mixture…

Source: arXiv cs.LG Alankar Atreya, Devesh Batra, Yoages Kumar Mantri, Geremy Bantug, Greig A Cowan, Raad Khraishi
AI Research AI

EMAN: Optimization-Driven Capacity Growth through Path Emergence in Multi-Task Learning

arXiv:2608.16930v1 Announce Type: new Abstract: Existing multi-task learning methods rely on hard sharing, multiple paths or experts, adaptive sharing, and dynamic expansion. However, their capacity changes are usually…

Source: arXiv cs.LG Chenlei Fang, Jingchen Li, Hongzong LI, Qingyao Li, Yixuan Zhang, Huarui Wu, Haobin Shi, Chunjiang Zhao
AI Research AI

Monotone Classification with Relative Approximations

arXiv:2506.10775v3 Announce Type: replace Abstract: In monotone classification, the input is a multi-set $P$ of points in $\mathbb{R}^d$, each associated with a hidden label from $\{-1, 1\}$. The goal is to identify a…

Source: arXiv cs.LG Yufei Tao
AI Research AI

The Optimal Sample Complexity of Multiclass and List Learning

arXiv:2604.24749v4 Announce Type: replace Abstract: While the optimal sample complexity of binary classification in terms of the VC dimension is well-established, determining the optimal sample complexity of multiclass…

Source: arXiv cs.LG Chirag Pabbaraju
AI Research AI

Communication Reduction via Semantic-Based Encoding in DMPC Using LSTMs

arXiv:2608.17592v1 Announce Type: cross Abstract: The communication demands of distributed model prediction control (DMPC) can overwhelm even advanced wireless communication technologies as agents must exchange a…

Source: arXiv cs.LG Torben Schiz, Pedro H. J. Nardelli, Henrik Ebel
AI Research AI

Adaptive surrogate modeling for high-dimensional spatio-temporal output

arXiv:2608.17250v1 Announce Type: cross Abstract: This paper develops an adaptive surrogate modeling method for problems with very high-dimensional spatio-temporal outputs. The analysis of spatio-temporal multi-physics…

Source: arXiv cs.LG Berkcan Kapusuzoglu, Shunsaku Matsumoto, Yoshitomo Miyagi, Daigo Watanabe, Sankaran Mahadevan
AI Research AI

Large-scale AI-Ready Data for Anti-Cancer Drug Response Modeling

arXiv:2608.11444v2 Announce Type: replace-cross Abstract: Drug response prediction (DRP) models are an active area of research in pharmacogenomics, with growing potential to accelerate the identification of effective…

Source: arXiv cs.LG Vincent Lavelle, Yitan Zhu, Kaitlyn Marlor, Thomas Brettin, Rick Stevens
AI Research AI

How to make the most of your masked language model for protein engineering

arXiv:2603.10302v3 Announce Type: replace Abstract: A plethora of protein language models have been released in recent years. Yet comparatively little work has addressed how to best sample from them to optimize desired…

Source: arXiv cs.LG Calvin McCarter, Nick Bhattacharya, Sebastian W. Ober, Hunter Elliott
AI Research AI

Nonlinear Data Integration via Kernel Methods for Data Collaboration Analysis

arXiv:2605.27219v2 Announce Type: replace Abstract: Collaborative analysis of decentralized confidential datasets is important, but direct sharing of original datasets is often restricted by privacy and institutional…

Source: arXiv cs.LG Yamato Suetake, Yuta Kawakami, Shunnosuke Ikeda, Yuichi Takano
AI Research AI

TabNSM: Neural Sparse Mixer for Tabular Regression

arXiv:2608.18026v1 Announce Type: new Abstract: Large-scale, high-dimensional tabular regression remains challenging: tree-based models are robust but lack end-to-end representation learning, while deep models enable…

Source: arXiv cs.LG Ali Eslamian, Qiang Cheng
AI Research AI

OceanDepths: A Global Dataset of Paired Subsurface and Surface Ocean Observations

arXiv:2608.16373v2 Announce Type: replace Abstract: Despite comprising over 70% of its surface, the world's oceans are critically underobserved compared to the land surface or the atmosphere. Understanding the global…

Source: arXiv cs.LG Simon Donike, Ruben Cartuyvels, Antonino Ian Ferola, Elisa Carli, Diego Fernandez Prieto, Marie-Helene Rio
AI Research AI

Optimize Your Sampling: Tuned Diffusion Sampling with Bayesian Optimization

arXiv:2608.18040v1 Announce Type: new Abstract: Sampling from a diffusion model typically requires many forward passes through a large neural network, making generation computationally expensive. While much work has…

Source: arXiv cs.LG Travis Zhang, Christian Belardi, Justin Lovelace, Jin Peng Zhou, Saebyeol Shin, Carla P. Gomes, Kilian Q. Weinberger
AI Research AI

Open datasets and machine learning for two-phase heat transfer: a review following a spatial-temporal taxonomy

arXiv:2605.23037v2 Announce Type: replace Abstract: Two-phase heat transfer underpins boiling, condensation, immersion cooling, flow boiling, energy conversion, and electronics thermal management, but its coupled…

Source: arXiv cs.LG Christy Dunlap, Ridwan Olabiyi, Firas Al-Hindawi, Hari Pandey, Stephen Pierson, Daniel Curl, Braden Stevens, Mohammad Ishraq Hossain, Annapurna Parjuli, Chinmaya Joshi, Ashif Iquebal, Han Hu
AI Research AI

Doubly robust nearest neighbors in factor models

arXiv:2211.14297v5 Announce Type: replace-cross Abstract: We introduce and analyze an improved variant of nearest neighbors (NN) for estimation with missing data in latent factor models. We consider a matrix completion…

Source: arXiv cs.LG Raaz Dwivedi, Caleb Chin, Sabina Tomkins, Predrag Klasnja, Susan Murphy, Devavrat Shah
AI Research AI

SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA

arXiv:2509.25459v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) show promise in generating long-form scientific explanations that synthesize evidence and connect multiple factors. However, in…

Source: arXiv cs.LG Haozhou Xu, Dongxia Wu, Matteo Chinazzi, Ruijia Niu, Rose Yu, Yi-An Ma
AI Research AI

Deep Learning for Cross-Border Electricity Price Forecasting: A Comparative Study

arXiv:2608.17091v1 Announce Type: new Abstract: While publicly available electricity market data presents a valuable resource for forecasting research, the field lacks established benchmark datasets for standardized…

Source: arXiv cs.LG Hadeer Elashhab, Sai Srijan Papineni, Marvin Dorn, Veit Hagenmeyer, Benjamin Sch\"afer
AI Research AI

Expressivity In Multimodal Contrastive Learning

arXiv:2608.17203v1 Announce Type: cross Abstract: Contrastive learning has become a cornerstone of modern representation learning, powering CLIP-style models that underpin text-to-image generation, vision-language…

Source: arXiv cs.LG Andrew Stuart, Florian Wolf
AI Research AI

MITRE-SAGE: A Multi-Agent Cybersecurity Question-Answering model

arXiv:2608.16921v1 Announce Type: cross Abstract: Effective cybersecurity operations require timely and accurate analysis of large-scale heterogeneous security information; however, analysts increasingly struggle with…

Source: arXiv cs.LG Ali Habibzadeh, Farid Feyzi, Reza Ebrahimi Atani
AI Research AI

Q-Learning With World Models

arXiv:2608.17163v1 Announce Type: new Abstract: Off-policy reinforcement learning (RL) has become increasingly sample-efficient, enabling applications such as RL fine-tuning of Vision-Language-Action models into…

Source: arXiv cs.LG Perry Dong, Yueru Jia, Chelsea Finn, Dorsa Sadigh
AI Research AI

ClockRoPE: Random Fourier Rotations for Temporal Routine Modeling

arXiv:2607.26369v2 Announce Type: replace Abstract: Rotary Position Embedding (RoPE) has been widely adopted in transformer-based large language models. However, its log-linear frequency schedule, originally designed to…

Source: arXiv cs.LG Yiwen Chen, Joshua Ainslie, Krzysztof Choromanski, Xiang Gao, Su-Lin Wu, Yiping Yuan, Qian Sun
AI Research AI

On detection probabilities of link invariants

arXiv:2509.05574v3 Announce Type: replace-cross Abstract: We prove that, for many standard link invariants, both the proportion of distinct invariant values and the detection probability among prime alternating links…

Source: arXiv cs.LG Tuomas Kelom\"aki, Abel Lacabanne, Daniel Tubbenhauer, Pedro Vaz, Victor L. Zhang
AI Research AI

Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation

arXiv:2608.17941v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but relies on costly rollout exploration. Assigning the…

Source: arXiv cs.LG Zhizhao Liu, Zhiliang Tian, Xi Wang, Zhihua Wen, Yihang Xiong, Zhiquan Lai, Dongsheng Li
AI Research AI

Diffusion Models for Smarter UAVs: Decision-Making and Modeling

arXiv:2501.05819v2 Announce Type: replace Abstract: Uncrewed Aerial Vehicles (UAVs) are increasingly used in modern communication networks. However, challenges in decision-making and digital modeling continue to hinder…

Source: arXiv cs.LG Yousef Emami, Hao Zhou, Luis Almeida, Kai Li
AI Research AI

Dynamic Compression in Recurrent Networks

arXiv:2608.17896v1 Announce Type: new Abstract: Recurrent models process long contexts efficiently by compressing their history into a fixed-size state, but modern architectures typically do so in a single causal pass…

Source: arXiv cs.LG Jyothish Pari, Ryan Bahlous-Boldi, Pulkit Agrawal
AI Research AI

Study-Strategy Clusters from EdNet Logs Track Engagement, Not Mastery

arXiv:2608.16963v1 Announce Type: new Abstract: Learning analytics often treats unsupervised clusters of intelligent tutoring system (ITS) logs as learner types that should predict learning. We test that assumption on…

Source: arXiv cs.LG Qingchuan Lyu, Yingxin Li, Albert Yang
AI Research AI

Wasted large language models: A life cycle thinking approach

arXiv:2608.17055v1 Announce Type: cross Abstract: Large Language Models (LLMs) are machine learning (ML) models that have an increasingly large carbon footprint through their development and use. Efforts to increase the…

Source: arXiv cs.LG Erik Johannes Husom, Maria Emine Nylund, Ophelia Prillard
AI Research AI

Efficient Dynamic Shielding for Parametric Safety Specifications

arXiv:2505.22104v2 Announce Type: replace-cross Abstract: Shielding has emerged as a promising approach for ensuring safety of AI-controlled autonomous systems. The algorithmic goal is to compute a shield, which is a…

Source: arXiv cs.LG Davide Corsi, Kaushik Mallik, Andoni Rodriguez, Cesar Sanchez
AI Research AI

Understanding the Surprising Generalization Properties of Tabular Foundation Models

arXiv:2608.17957v1 Announce Type: new Abstract: Tabular Foundation Models (TFMs) increasingly rely on in-context learning, where a model receives labelled examples at inference time and predicts labels for new inputs…

Source: arXiv cs.LG Nour Shaheen, Junwei Ma, Alex Labach, Frank Hutter, Valentin Thomas, Anthony L. Caterini
AI Research AI

Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

arXiv:2603.24472v4 Announce Type: replace-cross Abstract: Self-distillation has emerged as an effective post-training paradigm for LLMs, often improving performance while shortening reasoning traces. However, in…

Source: arXiv cs.LG Jeonghye Kim, Xufang Luo, Minbeom Kim, Sangmook Lee, Dohyung Kim, Jiwon Jeon, Dongsheng Li, Yuqing Yang
AI Research AI

MoNe: Modular Neural Memory for Efficient Long Context Inference

arXiv:2608.17616v1 Announce Type: cross Abstract: We present MoNe, a lightweight modular neural memory that attaches to any frozen pretrained Transformer to enable long-context inference without retraining. MoNe reads…

Source: arXiv cs.LG Wonguk Cho, Kyubyung Chae, Tribhuvanesh Orekondy, Sunghyun Park, Hyoungwoo Park, Jeongho Kim, Arash Behboodi, Kyuwoong Hwang, Sungrack Yun
AI Research AI

A Constant-Competitive Algorithm for Dynamic Mixture-of-Experts Serving

arXiv:2608.16947v1 Announce Type: cross Abstract: Huang, Lou, and Xiao introduced Dynamic Mixture-of-Experts Serving and gave an O(sqrt(log k))-competitive randomized algorithm for its integral primal problem, where k…

Source: arXiv cs.LG Ian D'Ambrosio (Nth Research Collective)
AI Research AI

Picture the Epsilon: Pursuing Identity-Level Privacy Guarantees for Images

arXiv:2608.17147v1 Announce Type: cross Abstract: Image-to-image face generators are widely used, and visual dissimilarity between their outputs and source images is sometimes treated as evidence of privacy. Auditing…

Source: arXiv cs.LG Arman Zareian Jahromi, Vishnu Bondalakunta, Mohammad Akbar Bin Shah, Naimul Haque, Shuangqing Wei, George T. Amariucai
AI Research AI

TokEval: A Tokenizer Evaluation Suite

arXiv:2608.18062v1 Announce Type: cross Abstract: Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be…

Source: arXiv cs.LG Clara Meister
AI Research AI

Online Learning of Scale Parameters in Score-Driven Filters

arXiv:2608.09218v2 Announce Type: replace Abstract: Score-driven filters update a time-varying parameter by multiplying a scaled log-likelihood score by a scale parameter that controls the magnitude of the update. We…

Source: arXiv cs.LG Fabrizio Lillo, Giulia Livieri, Gianluca Palmari
AI Research AI

ComNetX: Local Hierarchical Adaptation for Dynamic Community Detection

arXiv:2608.16906v1 Announce Type: cross Abstract: Dynamic community detection is commonly addressed either by full-snapshot recomputation or by solver-specific dynamic procedures. Full recomputation preserves the…

Source: arXiv cs.LG Aleksandr Konovalov, Anna Uporova, Alexander Drobyshev, Iaroslav Egorov, Grigoriy Bokov
AI Research AI

SparsePixels: Efficient Convolution for Sparse Data on FPGAs

arXiv:2512.06208v4 Announce Type: replace-cross Abstract: Inference of standard convolutional neural networks (CNNs) on FPGAs often incurs high latency and a long initiation interval due to the deep nested loops…

Source: arXiv cs.LG Ho Fung Tsoi, Dylan Rankin, Vladimir Loncar, Philip Harris
AI Research AI

Federated Compositional Muon Optimizer for Matrix-Wise Models

arXiv:2608.12710v2 Announce Type: replace Abstract: Muon, a more recently developed optimizer, is useful for matrix-wise models in AI areas. Although many works have studied Muon and its variants, these methods are…

Source: arXiv cs.LG Wang Yan, Feihu Huang
AI Research AI

Certified but Private: Scalable Zero-Knowledge Proofs for Neural Network Guarantees

arXiv:2608.17070v1 Announce Type: new Abstract: With the growing deployment of machine learning models, formal guarantees of the robustness and fairness of these models have become increasingly important in…

Source: arXiv cs.LG Youwei Zhong, Ben Merbaum, Timos Antonopoulos, Ning Luo, Charalampos Papamanthou, Katerina Sotiraki, Ruzica Piskac
AI Research AI

Diagonal Multi-omics Integration of Heterogenous Datasets

arXiv:2608.16968v1 Announce Type: cross Abstract: In this paper, we consider methods for the diagonal multi-omics integration of heterogeneous datasets. Several approaches to the nature of biological heterogeneity are…

Source: arXiv cs.LG Maksim V. Kukushkin, Mikhail S. Arbatskiy, Dmitriy E. Balandin, Alexey V. Churov
AI Research AI

Quantifying Memorization and Privacy Risks in Genomic Language Models

arXiv:2603.08913v2 Announce Type: replace Abstract: Genomic language models (GLMs) have emerged as powerful tools for learning representations of DNA sequences, enabling advances in variant prediction, regulatory…

Source: arXiv cs.LG Alexander Nemecek, Wenbiao Li, Xiaoqian Jiang, Jaideep Vaidya, Erman Ayday
AI Research AI

Reinforcement Learning as (Discrete) Potential Theory

arXiv:2608.17181v1 Announce Type: new Abstract: Reinforcement learning (RL) theory fundamentally depends on probability theory through the Markov chain. There is a deep connection between probability theory and…

Source: arXiv cs.LG Christopher Connolly
AI Research AI

Backward through Time, Algebraically

arXiv:2608.17087v1 Announce Type: new Abstract: Linear temporal logic is a modal extension of propositional logic that allows one to state how a system should behave over time. Its canonical domain is the booleans, but…

Source: arXiv cs.LG Konstantinos Kogkalidis
AI Research AI

Estimating Parameter Fields in Multi-Physics PDEs from Scarce Measurements

arXiv:2509.00203v3 Announce Type: replace Abstract: Parameterized partial differential equations (PDEs) underpin the mathematical modeling of complex systems in diverse domains, including engineering, healthcare, and…

Source: arXiv cs.LG Xuyang Li, Mahdi Masmoudi, Rami Gharbi, Nizar Lajnef, Vishnu Naresh Boddeti
AI Research AI

How Transparent is DiffusionGemma?

arXiv:2606.20560v2 Announce Type: replace Abstract: LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalignment, and debugging surprising model behaviors.…

Source: arXiv cs.LG Joshua Engels, Callum McDougall, Bilal Chughtai, Janos Kramar, Senthooran Rajamanoharan, Cindy Wu, Arthur Conmy, Asic Q Chen, Jean Tarbouriech, Min Ma, Brendan O'Donoghue, Jo\~ao Gabriel Lopes de Oli…
AI Research AI

Mos-Gen: A Generative Molecular Framework for Mosquito Insecticide Design

arXiv:2606.01846v2 Announce Type: replace Abstract: Mosquito-borne infectious diseases cause more than 700000 deaths worldwide each year. The long-term use of conventional chemical insecticides has induced serious…

Source: arXiv cs.LG Lina Wang, Yaning Cui, Zhifeng Gao, Ping Xing, Biao Jiang
AI Research AI

CrevasseSeg: A Label-Efficient UAV Crevasse Segmentation Framework

arXiv:2608.15790v2 Announce Type: replace Abstract: Crevasse mapping from uncrewed aerial vehicle (UAV) imagery matters for glaciological research and for field safety in glaciated terrain. Yet, pixel-level annotation…

Source: arXiv cs.LG Steven Wallace, William D. Harcourt, Richard Hann, Aiden Durrant, Somayajulu Sripada, Georgios Leontidis
AI Research AI

The concentration game: Bayesian updating, regret, and information

arXiv:2608.18061v1 Announce Type: new Abstract: We give a two-player zero-sum repeated game between a learner and nature whose value identity generates Bayesian updating and an exact accounting of exponential-weights…

Source: arXiv cs.LG Akshay Balsubramani
AI Research AI

On the Pseudo-Mixing of Kac's Walk

arXiv:2608.17374v1 Announce Type: cross Abstract: Motivated by a conjecture of Vaikuntanathan and Zamir, we study the pseudo-mixing of Kac's walk on $\mathrm{SO}(n)$: whether short trajectories are indistinguishable…

Source: arXiv cs.LG Natesh S. Pillai, Aaron Smith, Vinod Vaikuntanathan
AI Research AI

The politics of postmortem privacy

arXiv:2608.16905v1 Announce Type: cross Abstract: While the existence of postmortem privacy is increasingly acknowledged (such as the protection of the presence of deceased within digital spaces), far less attention has…

Source: arXiv cs.CL Mauricio Figueroa
AI Research AI

Grading Needs a Rubric, Not Intelligence

arXiv:2608.17938v1 Announce Type: new Abstract: Small language models can grade open-ended examination answers as reliably as substantially more expensive models when they grade against an explicit rubric. We test this…

Source: arXiv cs.CL Jhen-Ke Lin
AI Research AI

SCRIBES: Web-Scale Script-Based Semi-Structured Data Extraction with Reinforcement Learning

arXiv:2510.01832v2 Announce Type: replace Abstract: Semi-structured content in HTML tables, lists, and infoboxes accounts for a substantial share of factual data on the web, yet the formatting complicates usage, and…

Source: arXiv cs.CL Shicheng Liu, Kai Sun, Lisheng Fu, Xilun Chen, Xinyuan Zhang, Zhaojiang Lin, Rulin Shao, Yue Liu, Anuj Kumar, Wen-tau Yih, Xin Luna Dong
AI Research AI

Chronos: The AI Co-Historian

arXiv:2604.03553v3 Announce Type: replace-cross Abstract: AI is increasingly supporting, accelerating, and automating scientific discovery across subjects. Yet, the adoption of AI in historical research remains limited…

Source: arXiv cs.CL Lorenz Hufe, Niclas Griesshaber, Gavin Greif, Sebastian Oliver Eck, Pieter Francois, Wojciech Samek, Christian Schroeder de Witt, Philip Torr
AI Research AI

LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap

arXiv:2608.17330v1 Announce Type: cross Abstract: Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a…

Source: arXiv cs.CL Yining Hua, Cyrus Ayubcha, Hongbin Na, Levi Lian, Alon Gorenshtein, Yiftach Barash, Eyal Klang
AI Research AI

H$^{2}$MT: Semantic Hierarchy-Aware Hierarchical Memory Transformer

arXiv:2605.24930v2 Announce Type: replace Abstract: Transformer-based LLMs achieve strong results on many language tasks; however, long inputs remain challenging because context windows are finite, and prefill latency…

Source: arXiv cs.CL Maryam Haghifam, Zifan He, Jason Cong, Yizhou Sun
AI Research AI

Chain-of-Experience for Continual LLM Improvement

arXiv:2608.18027v1 Announce Type: new Abstract: Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the models' ability to improve through inference-time…

Source: arXiv cs.CL Haoqin Tu, Yunhao Fang, Yizhong Wang, Cihang Xie, Shen Yan
AI Research AI

How Do Large Language Models Learn Concepts During Continual Pre-Training?

arXiv:2601.03570v2 Announce Type: replace Abstract: Human beings primarily understand the world through concepts (e.g., dog), abstract mental representations that structure perception, reasoning, and learning. However,…

Source: arXiv cs.CL Barry Menglong Yao (UC Davis), Sha Li (Virginia Tech), Yunzhi Yao (UCLA), Minqian Liu (Virginia Tech), Zaishuo Xia (UC Davis), Qifan Wang (Meta AI), Lifu Huang (UC Davis)
AI Research AI

Which Source Wins? Task-Dependent Reliance in Vision-Language Models

arXiv:2608.17205v1 Announce Type: new Abstract: Vision-language models (VLMs) combine images and text, but when the two conflict and one becomes harder to read, it is unclear how a model shifts its reliance between…

Source: arXiv cs.CL Rodela Ghosh, Aviral Gupta, Guangjing Wang
AI Research AI

Eval4Sim: An Evaluation Framework for Persona Simulation

arXiv:2603.02876v2 Announce Type: replace Abstract: Large Language Model personas, explicit profiles specifying a user's attributes, preferences, and behavioural tendencies, are increasingly used to simulate human…

Source: arXiv cs.CL Eliseo Bao, Anxo Perez, Javier Parapar, Xi Wang
AI Research AI

Douyin Multimodal Embedding Model Technical Report

arXiv:2608.02148v3 Announce Type: replace-cross Abstract: Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vectors, it powers industrial search and…

Source: arXiv cs.CL Haonan Chen, Chu Li, Zhicheng Wang, Yuanwei Liu, Yuanjiang Wang, Shaohua Jiang, Zhicheng Dou
AI Research AI

Write, Execute, Refine: From Skill Followers to Skill Optimizers via Reinforcement Learning from Execution Feedback

arXiv:2608.17587v1 Announce Type: new Abstract: Expert-written natural language skills can improve tool-using agents, yet agent-authored skills perform 8-11 points worse than using no skill. This gap suggests that…

Source: arXiv cs.CL Kang Peng, Zhiwei Zhang, Yichen Zhang, Zezhong Wang, Yiming Du, Geng Tu, Baojun Wang, Bin Liang, Ruifeng Xu, Kam-Fai Wong
AI Research AI

AVA-Encoder: Towards Agent-Native Video Representation Learning

arXiv:2608.12313v2 Announce Type: replace-cross Abstract: Video creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key…

Source: arXiv cs.CL Chuyue Li, Jinpeng Yu, Haozhe Wang, Tian Xueyun, Zhijing Zhang, Bingnan Li, Shuqi Gu, Kan Ren, Jiaming Liu, Ruihua Huang
AI Research AI

Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review

arXiv:2507.06185v2 Announce Type: replace-cross Abstract: In July 2025, 18 academic manuscripts on arXiv contained hidden instructions that manipulated AI-assisted peer review (indirect prompt injection). Instructions…

Source: arXiv cs.CL Zhicheng Lin
AI Research AI

ContextClaim: A Context-Driven Paradigm for Verifiable Claim Detection

arXiv:2603.30025v3 Announce Type: replace Abstract: Automated fact-checking pipelines typically begin with a filtering stage that decides which claims are worth verifying, given that the later evidence retrieval and…

Source: arXiv cs.CL Yufeng Li, Rrubaa Panchendrarajan, Arkaitz Zubiaga
AI Research AI

Can LLMs Reliably Self-Report Adversarial Prefills, and How?

arXiv:2606.23671v4 Announce Type: replace Abstract: Prior work shows that large language models (LLMs) exhibit introspective capability on benign tasks. We extend the question to safety contexts and examine how reliably…

Source: arXiv cs.CL Quang Minh Nguyen, Uzair Ahmed, Taegyoon Kim
AI Research AI

Q-Interference: Memory-Efficient Phase-Aware Quantum-Inspired Attention

arXiv:2608.17288v1 Announce Type: new Abstract: GPT attention measures token compatibility through dot-product similarity. This mechanism is simple, effective, and memory-efficient. But it does not explicitly model…

Source: arXiv cs.CL Emama Nahid, Tahmid Imtiaz Imu, Huayue Gu, Liran Ma, Zhipeng Cai, Honghui Xu
AI Research AI

CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method

arXiv:2608.17536v1 Announce Type: new Abstract: Legal consultation questions exhibit multi-level complexity. A single retrieval strategy often leads to over-reasoning for simple questions and poor interpretability for…

Source: arXiv cs.CL Jin Su, Zhuofeng Zhao, Huanhuan Wang, Hao Chen
AI Research AI

Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It

arXiv:2608.17809v1 Announce Type: new Abstract: Humans naturally form and express beliefs in daily communication, e.g., "I think the answer is 3" or "I suppose that's right." Such beliefs inevitably intertwine with fact…

Source: arXiv cs.CL Quang Minh Nguyen, Luis Frentzen Salim
AI Research AI

LLM-Derived Preference Judgments Are Not Self-Consistent

arXiv:2608.17644v1 Announce Type: cross Abstract: Agents increasingly interpret a person's natural-language preferences by querying an LLM for numerical preference judgments, e.g., by asking how much the person would be…

Source: arXiv cs.CL Matthew T. Ford, Francis Bahk, Jingjing Wang, Adam S. Jovine, Tinghan Ye, David B. Shmoys, Peter I. Frazier
AI Research AI

N-gram-like Language Models Predict Naturalistic Reading Time Best

arXiv:2603.09872v2 Announce Type: replace Abstract: Recent work has found that contemporary language models such as transformers can become so good at next-word prediction that the probabilities they calculate become…

Source: arXiv cs.CL James A. Michaelov, Roger P. Levy
AI Research AI

Memory Is Communication: The Frontier Between Remembering and Signaling

arXiv:2608.17053v1 Announce Type: cross Abstract: A bounded agent may obtain information for a decision from its own past, from peers, or from both sources. Retaining task-relevant history can reduce later…

Source: arXiv cs.CL Yashar Talebirad, Eden Redman, Ali Parsaee, Osmar R. Zaiane
AI Research AI

Measuring Narrative Polarization in Online Discourse

arXiv:2601.07398v2 Announce Type: replace-cross Abstract: Polarization research has demonstrated how people cluster in homogeneous groups with opposing opinions. However, this effect emerges not only through interaction…

Source: arXiv cs.CL Jan Elfes, Marco Bastos, Luca Maria Aiello
AI Research AI

Effects of Answer Format Variation on Gender Bias in Large Language Models

arXiv:2608.17516v1 Announce Type: new Abstract: Gender bias or other social biases in large language models (LLMs) are frequently evaluated with question answering or survey benchmarks where the LLM needs to give a…

Source: arXiv cs.CL Ksenia Merzlyakova, Sebastian Pad\'o, Franziska Weeber
AI Research AI

Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards

arXiv:2608.07531v2 Announce Type: replace Abstract: Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing external rewards…

Source: arXiv cs.CL Ruoxi Cheng, Haoxuan Ma, Hongyi Zhang, Junming Zhang, Ranjie Duan, Qiaolin Xia, Hao Wang, Yu Lu, Haibo Shi, Xingjun Ma
AI Research AI

SCOPE: Selective Conformal Optimized Pairwise LLM Judging

arXiv:2602.13110v4 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used as scalable judges in pairwise evaluation, but they remain prone to miscalibration and biases. We propose…

Source: arXiv cs.CL Sher Badshah, Ali Emami, Hassan Sajjad
AI Research AI

Cross-Model Memory Transfer via Target-Side Reader Adaptation

arXiv:2608.17050v1 Announce Type: new Abstract: Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieval offers flexible access to external knowledge, but…

Source: arXiv cs.CL Mingyuan Li, Guangsheng Yu, Xu Wang, Shaoxiong Ji
AI Research AI

TSQueryBench: LLM-as-a-Judge for Time Series Explanations

arXiv:2604.02118v2 Announce Type: replace-cross Abstract: Natural language explanations of time series data are increasingly produced by foundation models in high stakes domains, making factual correctness critical.…

Source: arXiv cs.CL Preetham Sivalingam, Murari Mandal, Dhruv Kumar, Saurabh Deshpande
AI Research AI

BayesPrompt: human readable prompts that make sense

arXiv:2608.17866v1 Announce Type: new Abstract: Reconstructing prompts that can elicit a desired answer or behaviour in an LLM is an open and important research topic. Optimisation methods which aim at minimising the…

Source: arXiv cs.CL Franky Kevin Nando Tezoh, Ali Hussaini Umar, Alessandro Laio, Guido Sanguinetti, Riccardo Rende
AI Research AI

Dripper: Token-Efficient Main HTML Extraction with a Lightweight LM

arXiv:2511.23119v3 Announce Type: replace Abstract: High-quality main content extraction from web pages is a critical prerequisite for constructing large-scale training corpora. While traditional heuristic extractors…

Source: arXiv cs.CL Mengjie Liu, Jiahui Peng, Wenchang Ning, Pei Chu, Jiantao Qiu, Ren Ma, He Zhu, Rui Min, Lindong Lu, Linfeng Hou, Kaiwen Liu, Yuan Qu, Zhenxiang Li, Chao Xu, Zhongying Tu, Wentao Zhang, Conghui He
AI Research AI

The Plot Thins: Uniformity and Linearity in Literary Summaries

arXiv:2608.17218v1 Announce Type: new Abstract: Works of literature are complicated; they balance plot, suspense, surprise, and artistic expression. Summaries of literature prioritize plot, and therefore may deviate…

Source: arXiv cs.CL Rebecca M. M. Hicke, Sil Hamilton, David Mimno, Ross Deans Kristensen-McLachlan
AI Research AI

An Investigation of the NeurIPS and ICML 2025 Position Tracks

arXiv:2608.16894v1 Announce Type: cross Abstract: ML venues shape what kinds of research claims become legible to reviewers and what forms of evidence count as rigorous. The NeurIPS and ICML Position Paper Tracks were…

Source: arXiv cs.CL Fan Yang, Wenkai Li, Jun Liu
AI Research AI

Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades

arXiv:2605.17443v3 Announce Type: replace Abstract: We analyze how automatic speech recognition (ASR) errors propagate through ASR--LLM cascades in Korean spoken question answering (SQA), focusing on downstream semantic…

Source: arXiv cs.CL Donghyuk Jung, Youngwon Choi
AI Research AI

AQuA: Recursively Self-Improving Quantitative Trading Research Agents

arXiv:2608.12841v2 Announce Type: replace Abstract: We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from earlier experiments to improve…

Source: arXiv cs.CL Jiacheng Guo, Suozhi Huang, Yunlong Gao, Zihao Li, Jason Ge, Xu Kuang, Mengdi Wang
AI Research AI

ArborMem: Navigating Interaction States with Memory Forests

arXiv:2608.17534v1 Announce Type: new Abstract: Large language models increasingly serve as persistent conversational assistants, requiring memory that preserves relevant experience and maintains continuity across…

Source: arXiv cs.CL Zongwei Lv, Yuemeng Xu, Yilun Yao, Siyi Ding, Xinyu Tan, Yaoming Li, Guangxiang Zhao, Weihong Lin, Lin Sun, Xiangzheng Zhang, Tong Yang
AI Research AI

SOD: Step-wise On-policy Distillation for Small Language Model Agents

arXiv:2605.07725v3 Announce Type: replace Abstract: Tool-integrated reasoning (TIR) is difficult to scale to small language models due to instability in long-horizon tool interactions and limited model capacity. While…

Source: arXiv cs.CL Qiyong Zhong, Mao Zheng, Mingyang Song, Xin Lin, Jie Sun, Houcheng Jiang, Xiang Wang, Junfeng Fang
AI Research AI

Parametric Knowledge in RAG-SFT for Domain-Specific Document Generation

arXiv:2603.23047v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) fine-tuning has shown substantial improvements over vanilla RAG, yet most studies target document question answering, leaving open…

Source: arXiv cs.CL Julian Oestreich, Maximilian Bley, Frank Binder, Lydia M\"uller, Andr\'e Alcalde, Maksym Sydorenkoq
AI Research AI

The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning

arXiv:2608.18011v1 Announce Type: new Abstract: Reasoning in LLMs is overwhelmingly studied in domains that provide a model with rules: mathematics and code. Linguistic puzzles invert this: the solver must first…

Source: arXiv cs.CL Eduardo S\'anchez, Rita Berrada, Dan-Mircea Mirea, Sara Rajaee, Alexander Piperski, Ana Meta Dolinar, Boris Iomdin, Andrey Nikulin, Mariya Shmatova, Marzieh Fadaee, Julia Kreutzer
AI Research AI

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

arXiv:2607.27155v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for…

Source: arXiv cs.CL Jingbo Zhou, Yusai Zhao, Qi Bao, Jingjia Cao, Zhenghai Chen, Chang Gao, Kaiqi Guo, Muxin Guo, Mingxuan Li, Xinjiang Lu, Yanru Ma, Yixiong Xiao, Zenghui Zhang, Le Zhang, Hua Wu
AI Research AI

Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss

arXiv:2608.17051v1 Announce Type: new Abstract: Secondary use of electronic health records requires de-identification, yet existing systems miss \emph{institutionally situated} protected health information (PHI) such as…

Source: arXiv cs.CL Daniel Palacios, Matthew Brady Neeley, Angel Adetomike Otto, Shalini Dhamodharan, John P. Woodhouse, Chi-fan Lin, Mark Zobeck, Zhandong Liu, Hyun-Hwan Jeong
AI Research AI

Code as Representation: A Compilable Parsing Paradigm for Academic Documents

arXiv:2608.17550v1 Announce Type: cross Abstract: Academic papers are a primary carrier of scientific knowledge, yet most of this knowledge remains locked in PDFs that are optimized for human reading rather than machine…

Source: arXiv cs.CL Rihui Jin, Jun Wang, chengyuan zhu, Liang Mingyu, Yue Gao, Li Yunxuan, Kuicai Dong, Guilin Qi, Lin Ren, Yongrui Chen, Xinbang Dai, Jiaqi Li, Tongtong Wu, Gholamreza Haffari
AI Research AI

Understanding Undesirable Word Embedding Associations

arXiv:1908.06361v2 Announce Type: replace Abstract: Word embeddings are often criticized for capturing undesirable word associations such as gender stereotypes. However, methods for measuring and removing such biases…

Source: arXiv cs.CL Kawin Ethayarajh, David Duvenaud, Graeme Hirst
AI Research AI

Uncertainty-Aware Decision Making in Multimodal Large Language Models

arXiv:2608.17084v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) increasingly answer questions whose correctness depends on visual, textual, temporal, acoustic, document, chart, or embodied…

Source: arXiv cs.CL Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed
AI Research AI

LexKairos: Benchmarking Legal Temporal Capabilities in LLMs

arXiv:2608.09106v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated strong performance across a wide range of legal tasks. In legal practice, time is a critical concept that governs the…

Source: arXiv cs.CL Chenyang Li, Zejia Feng, Yuqin Huang, Yuxiao Ye, Huiyuan Xie
AI Research AI

HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous Environments with Dynamic Multi-Human Interactions

arXiv:2503.14229v5 Announce Type: replace Abstract: Vision-and-Language Navigation (VLN) has been studied mainly in either discrete or continuous spaces, with little attention to dynamic, crowded environments. We…

Source: arXiv cs.AI Yifei Dong, Fengyi Wu, Qi He, Lingdong Kong, Heng Li, Minghan Li, Zebang Cheng, Yuxuan Zhou, Jingdong Sun, Qi Dai, Alexander G Hauptmann, Zhi-Qi Cheng
AI Research AI

LLM-Guided Graph Generation for Structure-Based Local Improvement Methods

arXiv:2608.13333v3 Announce Type: replace Abstract: Large neighborhood search normally selects a random subset of decision variables for iterative optimization. To efficiently solve various problems, researchers tend to…

Source: arXiv cs.AI Hai Xia, Vaidyanathan Peruvemba Ramaswamy, Stefan Szeider
AI Research AI

EvoTS-Agent: A Self-Evolving LLM Agent for Financial Time Series Change Point Detection

arXiv:2608.17933v1 Announce Type: new Abstract: Financial time series exhibit non-stationary and heterogeneous statistical properties, making change-point detection challenging because no single unsupervised algorithm…

Source: arXiv cs.AI Lei Jiang, Ye Wei, Xinyu Xi, Jordan Langham-Lopez, Yifan Bao, Raad Khraishi, Yihao Ang, Anthony K. H. Tung, Lukasz Szpruch, Hao Ni
AI Research AI

The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges

arXiv:2608.17829v1 Announce Type: cross Abstract: LLMs increasingly rely on external contexts, such as pre-defined system prompts or retrieved documents, to improve generation quality. However, processing these contexts…

Source: arXiv cs.AI Maosen Zhang, Jianshuo Dong, Boting Lu, Wenyue Li, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu
AI Research AI

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety

arXiv:2608.17597v1 Announce Type: cross Abstract: Large language models are increasingly deployed through agent harnesses that manage tools, extensions, persistent state, permissions, and external actions. Existing…

Source: arXiv cs.AI Yajing Bai, Jinhao Duan, Jie Peng, Xianfeng Wu, Sijia Liu, Song Wang, Tianlong Chen
AI Research AI

SkillEffect: Checked Lowering for Memory-Bounded Agent Tools

arXiv:2608.17007v1 Announce Type: new Abstract: Agent Skills can specify procedural and resource obligations for tool use, and language models instantiate them as concrete programs. However, when models turn this…

Source: arXiv cs.AI Yinuo Wang, Yiyu Shi
AI Research AI

KernelArc: A Multi-Agent Framework for GPU Kernel Optimization

arXiv:2608.17071v1 Announce Type: new Abstract: We present KernelArc, a multi-agent framework for autonomous GPU kernel optimization across heterogeneous workloads. Strategy-specialized agents run in parallel and…

Source: arXiv cs.AI Joyjit Kundu, Ben Stoffelen, Kaili Wang, Peter Vrancx, Ludovic Denoyer
AI Research AI

A Multimodal Agentic Pathology Co-pilot via Evidence Grounded Reasoning

arXiv:2606.08093v2 Announce Type: replace Abstract: Pathology is the cornerstone of modern medicine, where accurate decision-making relies heavily on evidence-based practices. While artificial intelligence (AI) has the…

Source: arXiv cs.AI Zhe Xu, Zhengyu Zhang, Zhiyuan Cai, Jiahao Xu, Yijie Lin, Ziyi Liu, Junlin Hou, Hongyi Wang, Yuxiang Nie, Yihui Wang, Jiabo Ma, Ling Liang, Yingxue Xu, Zhengrui Guo, Guanghao Wu, Danyi Li, Ziqi Zhou,…
AI Research AI

Sparse Coverage: Semantic Center Representations for Patent Prior-Art Retrieval

arXiv:2608.16918v1 Announce Type: cross Abstract: Patent prior-art retrieval is a recall-oriented search task over long and highly structured technical documents. Dense retrieval improves semantic matching, but…

Source: arXiv cs.AI You Zuo (ALMAnaCH), Kim Gerdes (LISN, Qatent, STL), \'Eric de la Clergerie (ALMAnaCH), Beno\^it Sagot (ALMAnaCH)
AI Research AI

The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks

arXiv:2608.16630v1 Announce Type: cross Abstract: Repository-scale coding requires an agent to keep tests, imports, configuration, and migration rules consistent within a bounded context window. We model this as…

Source: arXiv cs.AI Bardia Mohammadi, Lars Klein, Aman Chadha, Akhil Arora, Laurent Bindschaedler
AI Research AI

FUSE: Frame-Unified Stress Estimation from Facial Video

arXiv:2608.10442v2 Announce Type: replace-cross Abstract: Automatic stress detection from facial video offers a practical path to non-intrusive affect monitoring, yet existing video-based approaches commonly decompose…

Source: arXiv cs.AI Stefanos Gkikas, Thomas Kassiotis, Yang Guo, Guangliang Li, Giorgos Giannakakis
AI Research AI

Learnware for CSI Feedback: Scene-specific Small Models Can Do Big

arXiv:2608.17760v1 Announce Type: cross Abstract: Intelligent channel state information (CSI) feedback is essential for realizing the high capacity and spectral efficiency goals of future 6G systems, yet existing deep…

Source: arXiv cs.AI Xiangyi Li, Jiajia Guo, Chao-Kai Wen, Xin Geng, Shi Jin, Zhi-Hua Zhou
AI Research AI

ScreenSearch: Uncertainty-Aware OS Exploration

arXiv:2605.16024v2 Announce Type: replace Abstract: Desktop GUI agents operate under partial observability: visually similar screens can correspond to different underlying workflow states, so locally plausible actions…

Source: arXiv cs.AI Michael Solodko, Justin Wagle
AI Research AI

DMT-Dens: Density-preserving manifold visualization for biological data

arXiv:2608.17571v1 Announce Type: cross Abstract: Motivation: Low-dimensional embeddings are widely used to explore cell-state heterogeneity in single-cell and other high-dimensional biological data. Although many…

Source: arXiv cs.AI Ruizhe Wang, Yixuan Dong, Bolin Yang, Bingo Wing-Kuen Ling, Fuji Yang, Zelin Zang
AI Research AI

Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems

arXiv:2608.03744v2 Announce Type: replace Abstract: Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such committees can be gamed by…

Source: arXiv cs.AI Sebasti\'an Andr\'es Cajas Ord\'o\~nez, Agastya Munnangi, Aldo Marzullo, Felipe Ocampo Osorio, Quang Bui, Mohammad Shahin, Armaan Grewal, Emmanuel Paul Kwesiga, Anqi Peter Li, Josephine Nanyonjo, Aad…
AI Research AI

AutoResearch: Insight In, Hallucination Out

arXiv:2608.17906v1 Announce Type: new Abstract: Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains…

Source: arXiv cs.AI Yiming Ren, Xiang Liu, Qumeng Sun, Xiao Zhang, Jiahao Li, Haoyang Zhang, Junjie Wang
AI Research AI

Benchmarking Automated Security Patch Backporting: How Far Are We?

arXiv:2608.17671v1 Announce Type: cross Abstract: Automated security patch backporting is critical for mitigating N-day vulnerabilities. Recent tools report success rates above 80% on their respective datasets. However,…

Source: arXiv cs.AI Jincheng Yang, Yulong Fu, Chengwei Liu, Lyuye Zhang, Fangyuan Zhang, Bingyang Ren, Yang Liu, Hui Li
AI Research AI

LSem2Vec: A Simple yet Effective Two-Stage Approach for Source Code Embedding

arXiv:2409.14644v4 Announce Type: replace-cross Abstract: The advent of large language models (LLMs) has significantly advanced artificial intelligence in software engineering, with source code embeddings playing a…

Source: arXiv cs.AI Zixiang Xian, Chenhui Cui, Rubing Huang, Chunrong Fang, Zhenyu Chen
AI Research AI

Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets

arXiv:2608.17360v1 Announce Type: cross Abstract: Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely solely on attack success rate (ASR) without accounting for its…

Source: arXiv cs.AI Zhida He, Xiaoyu Wen, Han Qi, Ziyuan Zhou, Peng Yu, Jiajia Li, Chaochao Lu, Qiaosheng Zhang
AI Research AI

PACE: Policy-Attested Contract Execution for Safe AI Agents in Decentralized Finance

arXiv:2608.17220v1 Announce Type: cross Abstract: Autonomous AI agents are emerging as interfaces for decentralized finance (DeFi) actions such as swaps, lending operations, and yield management. Because these agents…

Source: arXiv cs.AI Rabimba Karanjai (Larry), Yang Lu (Larry), Richard Williamson (Larry), Hemanth Hm (Larry), Prakhar Mehrotra (Larry), Lei Xu (Larry), Weidong (Larry), Shi
AI Research AI

G-ReAct: Graph-Guided Deep Search via Structure-State Co-Evolution

arXiv:2608.01324v2 Announce Type: replace Abstract: Deep search has become a fundamental capability of large language models (LLMs) for solving open-domain complex tasks. However, existing approaches typically rely on…

Source: arXiv cs.AI Shaoxiong Yang, Mengyuan Zhang, Shaojun Lin, Chao Li, Wei Liu, Kun Shao, Jian Luan
AI Research AI

Comparative Study of Out-of-the-Box Technology for Automatic Target Detection and Recognition

arXiv:2608.17917v1 Announce Type: cross Abstract: Automatic Target Detection and Recognition (ATD/R) is critical for military decision support and (semi-)autonomous operations. Recent advances in object detection and…

Source: arXiv cs.AI Alma M. Liezenga, Lotte Nijskens, Henrik R. Baumann, Stefan Becker, Simon Bensberg, Niccol\`o Camarlinghi, H{\aa}vard R. Eiring, Alexander W. Johnsgaard, Tanel Liiv, Giuseppe Martino, Matteo Marturin…
AI Research AI

On the Subgaussianity of Quantized Linear Maps: An AI-Assisted Note

arXiv:2605.27563v2 Announce Type: replace-cross Abstract: We prove an elementary bounded-differences inequality for functions of non-isotropic Gaussian vectors. Specifically, if $f$ has bounded coordinate differences…

Source: arXiv cs.AI Guangyi Zou, Roman Vershynin
AI Research AI

From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

arXiv:2608.18076v1 Announce Type: cross Abstract: Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize…

Source: arXiv cs.AI Xingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng, Qing Jin, Qinye Zhou, Zhengtao Wu, Yongchao Du, Zuan Gao, Chao Lin, Yefeng Shen, Xiaoli Xu, Zhengze Xu, Hao Yan, Yuhang Yu, Mingzhou Zhang, Mengting C…
AI Research AI

StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents

arXiv:2608.18050v1 Announce Type: new Abstract: AI agents increasingly perform knowledge work (i.e., produce and modify persistent digital artifacts such as code repositories, documents, spreadsheets, slides, reports),…

Source: arXiv cs.AI Yining Hua, Hongbin Na, Yifan Zhou, Akshay Kalose, Cyrus Ayubcha, Levi Lian
AI Research AI

Toward Personal Intelligence Through Cooperative Observation

arXiv:2608.17128v1 Announce Type: new Abstract: A personal AI system needs a model of the user's goals, constraints, and ongoing commitments to plan and act on their behalf, and the quality of that model is bounded by…

Source: arXiv cs.AI Yashar Talebirad, Osman Jime, Ali Parsaee, Eden Redman, Yongbin Kim, Osmar R. Zaiane
AI Research AI

Discovering physical mechanisms from experiment-simulation mismatches

arXiv:2604.26703v2 Announce Type: replace-cross Abstract: Scientific discovery often begins where observation and prediction disagree. As computation and machine learning survey chemical space, experiment-simulation…

Source: arXiv cs.AI Yue Li, Penghui Yang, Yushan Xiao, Zhonghan Zhang, Jianguo Huang, Yuhao Lu, Cuntai Guan, Bo An, Bijun Tang, Zheng Liu
AI Research AI

Physics-Grounded Causal Auditing of End-to-End Driving Planners

arXiv:2606.14438v4 Announce Type: replace-cross Abstract: End-to-end (E2E) autonomous-driving planners trained by imitation are prone to statistical shortcuts: they associate scene elements that merely co-occur with…

Source: arXiv cs.AI Zikun Guo, Minglan Chen, Jinyou Zhai, Rongjin Zou
AI Research AI

LLM-Only PDDL Domain Repair with Open-Weight Models

arXiv:2608.17341v1 Announce Type: new Abstract: AI planning is concerned with finding a sequence of actions that achieves a specified goal. It relies on explicit models of the world, commonly represented in the Planning…

Source: arXiv cs.AI Nader Karimi Bavandpour, Pascal Bercher
AI Research AI

FVSpec: Real-World Property-Based Tests as Lean Challenges

arXiv:2606.01008v2 Announce Type: replace-cross Abstract: We present a benchmark for evaluating AI models and agents on real-world formal software verification tasks. We first scrape 11,039 property-based tests (PBTs)…

Source: arXiv cs.AI Quinn Dougherty, Max von Hippel, Simon Henniger, Hazel Shackleton, Mike Dodds
AI Research AI

Learning Where and What to Lift for Bi-planar X-ray-to-CT Reconstruction

arXiv:2608.17255v1 Announce Type: cross Abstract: X-ray imaging can be approximately modeled as the projection of an underlying volumetric attenuation field, with each measurement recording the accumulated attenuation…

Source: arXiv cs.AI Yifei Wu, Yicheng Wu, Qiang Ma, Qi Chen, Renyang Gu, Xinyu Liu, Yongsheng Pan, Yong Xia
AI Research AI

ManiCM: Real-time 3D Diffusion Policy via Consistency Model for Robotic Manipulation

arXiv:2406.01586v4 Announce Type: replace-cross Abstract: Diffusion models have been verified to be effective in generating complex distributions from natural images to motion trajectories. Recent diffusion-based…

Source: arXiv cs.AI Zifeng Gao, Guanxing Lu, Tianxing Chen, Wenxun Dai, Ziwei Wang, Chao Shang, Wenbo Ding, Yansong Tang
AI Research AI

Evaluating the Diversity of AI-Generated Content with Diversity Profiles

arXiv:2608.17731v1 Announce Type: new Abstract: Diversity is a fundamental criterion for evaluating generative artificial intelligence (AI) systems, yet its measurement remains inherently ambiguous. Existing approaches…

Source: arXiv cs.AI Xiuyuan Hu, Xuege Hou, Guoqing Liu, Yang Zhao, Jieran Li, Dongbiao Sun, Jos\'e Miguel Hern\'andez-Lobato, Hao Zhang, Xue Liu
AI Research AI

Planning-aligned Token Compression for Long-Context Autonomous Driving

arXiv:2606.07464v2 Announce Type: replace-cross Abstract: Monolithic vision-action models represent an emerging paradigm in autonomous driving. However, this architecture produces token sequences that quickly exceed…

Source: arXiv cs.AI Zhixuan Liang, Yuxiao Chen, Yurong You, Peter Karkus, Wenhao Ding, Boyi Li, Alexander Popov, Yan Wang, Maximilian Igl, Yiming Li, Danfei Xu, Nikolai Smolyanskiy, Boris Ivanovic, Ping Luo, Marco Pavone
AI Research AI

Effective Personalized AI Tutors via LLM-Guided Reinforcement Learning

arXiv:2608.16907v1 Announce Type: cross Abstract: Generative AI (GenAI) is rapidly reshaping education by unlocking the potential for personalized tutoring. Yet, emerging platforms largely focus on GenAI chatbot tutors…

Source: arXiv cs.AI Angel Tsai-Hsuan Chung, Botong Zhang, Ling-Chieh Kung, Hamsa Bastani, Osbert Bastani
AI Research AI

Education-centered critical policy analysis of AI: Ghana's AI strategy as a case

arXiv:2608.16910v1 Announce Type: cross Abstract: National AI strategies increasingly guide governance, workforce development, innovation, and competitiveness, but less is known about how they frame education as a…

Source: arXiv cs.AI Matthew Nyaaba, Vida Awinime Bugri, Eric Kojo Majialuwe, Bismark Nyaaba Akanzire, Ibrahim Nantomah, Felicia Boateng, Patrick Kyeremeh, Benjamin Quarshie, Ellen Kwarteng, Macharious Nabang
AI Research AI

PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation

arXiv:2608.16984v1 Announce Type: cross Abstract: Recent monocular depth estimators achieve strong zero-shot generalization, yet often struggle to preserve fine-grained structures and object boundaries. We attribute…

Source: arXiv cs.AI Zhiyuan Yuan, Guanying Chen, Lingteng Qiu, Ruimao Zhang, Shuguang Cui, Xiaochun Cao
AI Research AI

Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents

arXiv:2608.17718v1 Announce Type: new Abstract: Long-horizon agents increasingly operate across many steps, tools, and observa- tions. In this setting, the relevant oversight question is not only whether each action is…

Source: arXiv cs.AI An He, Yao Wang, Haibin Zhang
AI Research AI

MoRA: Mobility as the Backbone for Geospatial Representation Learning at Scale

arXiv:2506.01297v5 Announce Type: replace Abstract: Representation learning of geospatial locations remains a core challenge in achieving general geospatial intelligence, with increasingly diverging philosophies and…

Source: arXiv cs.AI Ya Wen, Jixuan Cai, Qiyao Ma, Linyan Li, Xinhua Chen, Chris Webster, Yulun Zhou
AI Research AI

Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce

arXiv:2608.14825v2 Announce Type: replace-cross Abstract: Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety…

Source: arXiv cs.AI Zeyuan Li, Lukas Petersson, Alessandro Acquisti, Michiel A. Bakker
AI Research AI

ARASH: Adaptive Retrieval And Shot Selection for Tabular Prediction

arXiv:2608.17856v1 Announce Type: new Abstract: Tabular prediction is a critical task across numerous applications. The recent success of large language models has sparked various approaches for adapting them to the…

Source: arXiv cs.AI Samirasadat Jamalidinan, Yue Xu, Kazem Cheshmi
AI Research AI

Training with synthetic data for drone detection in thermal imagery

arXiv:2608.17799v1 Announce Type: cross Abstract: Ground-to-Air (G2A) drone detection in medium- and long-wave infrared (MWIR/LWIR) imagery is challenging due to reduced texture information, sensor noise, weak thermal…

Source: arXiv cs.AI Tanel Liiv, Sander Soodla, Nzamba Bignoumba, Alma M. Liezenga, Toomas Pruuden
AI Research AI

TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration

arXiv:2608.17336v1 Announce Type: new Abstract: Long-context prefill in large language models (LLMs) incurs substantial computation and memory traffic because dense self-attention computes quadratic query-key scores.…

Source: arXiv cs.AI Hanzhi Zhang, Qiao Zhang, Qinglei Cao, Heng Fan, Yan Huang, Kewei Sha, Yunhe Feng
AI Research AI

SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution

arXiv:2608.17468v1 Announce Type: new Abstract: Storyboards turn screenplays into visual shot plans for automated short drama production. Professional storyboarding relies on tacit directorial expertise and remains an…

Source: arXiv cs.AI Maolin Ran, Xiaoyang Lu, Jiaqi Liu, Jian Wang, Weiwen Liu, Jianghao Lin, Yong Yu, Weinan Zhang
AI Research AI

COMIC: Reference-Aware Safety Gating for Multimodal Large Language Models

arXiv:2608.17234v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are increasingly used to interact with screenshots, scanned documents, diagrams, and other visually grounded inputs. This shift…

Source: arXiv cs.AI Md Abdullahil Oaphy, Anhao Xiang, Zongxing Xie, Huayue Gu, Chenyu Wang, Honghui Xu
AI Research AI

ASI-Bench: At the Dawn of Artificial Superintelligence

arXiv:2608.17271v1 Announce Type: new Abstract: Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into…

Source: arXiv cs.AI Junwei Zhou, Zhen Sun, Binyu Li, Jiangyu Zhou, Yuexi Pan, Hengyu Wang, Honghe Ren, Xiaohan Jia, Xueyang Zhou, Xiaoyu Cao, Yongchao Chen, Yuanning Feng, Junhao Wu, Cheng Zhang, Sijia Chen, Haoyu Xue,…
AI Research AI

PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs

arXiv:2608.17289v1 Announce Type: new Abstract: Group-relative policy optimization has emerged as a key paradigm for training agentic large language models (LLMs) on multi-turn interactive tasks. However, most existing…

Source: arXiv cs.AI Dayang Liang, Liyuan He, Xuan Feng, Shuxin Li, Bo An, Yunlong Liu
AI Research AI

M3TR: Temporal Retrieval Enhanced Multi-Modal Micro-video Popularity Prediction

arXiv:2411.15455v3 Announce Type: replace-cross Abstract: Accurately predicting the popularity of micro-videos is a critical but challenging task, characterized by volatile, `rollercoaster-like' engagement dynamics.…

Source: arXiv cs.AI Jiacheng Lu, Weijian Wang, Mingyuan Xiao, Yang Hua, Tao Song, Bo Peng, Cheng Hua, Haibing Guan
AI Research AI

Evaluating Skill and Stability of ArchesWeather and ArchesWeatherGen under Multi-Decadal Climate Simulations

arXiv:2605.29976v3 Announce Type: replace-cross Abstract: We evaluate the climate simulation capabilities of ArchesWeather and ArchesWeatherGen, two machine learning models originally trained for weather forecasting and…

Source: arXiv cs.AI Renu Singh, Robert Brunstein, Antonia Jost, Yana Hasson, Thomas Rackow, Claire Monteleoni, Christian Lessig, Guillaume Couairon
AI Research AI

Adaptive Policy Portfolios for Robust Markov Decision Processes

arXiv:2608.17929v1 Announce Type: new Abstract: Robust Markov decision processes optimize one policy against a set of plausible transition functions. This can be conservative when the unknown dynamics are fixed and…

Source: arXiv cs.AI Kasper Engelen, Sebastian Junges, Guillermo A. P\'{e}rez, Marnix Suilen
AI Research AI

AI, Brain Death Detection, and Islamic Law

arXiv:2608.16903v1 Announce Type: cross Abstract: The deployment of machine learning systems capable of detecting covert consciousness in neurologically injured patients creates a profound challenge at the intersection…

Source: arXiv cs.AI Muhammad Aurangzeb Ahmad
AI Research AI

Average Distance Approximation for Static Large Graphs

arXiv:2608.16916v1 Announce Type: cross Abstract: Calculating average distances in large-scale networks is computationally intensive and constrained by limited main memory, posing a significant challenge in graph…

Source: arXiv cs.AI Kartikey Ahlawat
AI Research AI

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

arXiv:2608.17426v1 Announce Type: cross Abstract: We introduce Semantic Task Completion Video Generation, an outcome-oriented video generation task. Under this formulation, success requires both achievement of the…

Source: arXiv cs.AI Keyu Tu, Zhuowei Chen, Mengqi Huang, Yuxin Wang, Jiahao Zhu, Zhendong Mao, Yongdong Zhang
AI Research AI

A 12-CNOT Double Qubit Excitation Gate

arXiv:2608.11733v2 Announce Type: replace-cross Abstract: In this work, we presented, to the best of our knowledge, the first reported 12-CNOT decomposition of the double qubit excitation operator. We compared our new…

Source: arXiv cs.AI Irfansha Shaik
AI Research AI

UltraArUco: A Lightweight Multilingual Library And Framework With Low-Latency Real-Time Marker-Based Tracking System For Mobile AR Interaction

arXiv:2608.13584v2 Announce Type: replace-cross Abstract: UltraArUco is a lightweight multilingual library and framework for low latency, realtime marker-based tracking in mobile augmented reality. Unlike standard…

Source: arXiv cs.AI Mikhail Kiselev, Aleksandr Marukhin, Ivan Snegirev, Elizaveta Semenyakina, Miguel Altamirano Cabrera, Dzmitry Tsetserukou
AI Research AI

Traceable Trust for action-ready artificial intelligence in bioscience

arXiv:2608.17997v1 Announce Type: cross Abstract: Artificial intelligence (AI) is becoming part of the working infrastructure of the biosciences. AI models can predict biomolecular structures, design proteins, rank…

Source: arXiv cs.AI Huayu Xin, Yizhi Cai, Mukilan Deivarajan Suresh, Gavin Michael Farrell, Iwona Gajda, Charlie Harrison, Conor Houghton, Mato Lagator, Yang Lu, Virginia Portillo, Reyer Zwiggelaar, Sebastian Lobentanzer
AI Research AI

Agent Lightning v1.0: Towards Harnessed Agentic RL

arXiv:2608.17528v1 Announce Type: new Abstract: Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the harness a critical part of the agent system. Our original Agent…

Source: arXiv cs.AI Zhiyuan He, Siwei Zhang, Zhiwen Zhou, Yuqing Yang, Yu Kang, Yuge Zhang, Luna K. Qiu, Tin Yan Tsui, Jiahang Xu, Chong Luo
AI Research AI

CARA: Cognitive Adaptive Recommendation Agent

arXiv:2608.16919v1 Announce Type: cross Abstract: Recent advances in large language models and agent-based recommendation frameworks have introduced new opportunities for more flexible and context-aware recommendation.…

Source: arXiv cs.AI Weijun Gao, Jinyang Dong, Chuanru Ren, Hengxiao Li
AI Research AI

Adversarial Data Modeling in Epidemiology

arXiv:2602.20134v2 Announce Type: replace-cross Abstract: Epidemiological models increasingly rely on crowdsourced, self-reported behavioral data such as vaccination status, mask usage, and social distancing adherence.…

Source: arXiv cs.AI Yiqi Su, Christo Kurisummoottil Thomas, Walid Saad, Sanmay Das, Bud Mishra, Naren Ramakrishnan
AI Research AI

NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

arXiv:2608.12898v2 Announce Type: replace-cross Abstract: Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs)…

Source: arXiv cs.AI Peng Cai, Zhaofan Zou, Shifa Liu, Yikun Wang, Jiawei Tang, Kaicheng Yang, Meng Tong, MingKun Jiang, Zhongjiang He, Hao Sun
AI Research AI

Validated Adaptation for Aerial Crowd Monitoring at Mass Gathering Scale: A Deployment Protocol, a Severity Law, and a Diagnostic for Label-Free Drone Crowd Counting, Toward the FIFA World Cup 2034 (Saudi Arabia)

arXiv:2608.17625v1 Announce Type: new Abstract: Saudi Arabia will host the 2034 FIFA World Cup and already operates crowd management at Hajj scale. Drone-based counting must hold accuracy on footage unlike anything in…

Source: arXiv cs.AI AlAnoud AllGhayth, AlJawharh AlOtaibi, Jude AlSubaie
AI Research AI

The 10th AI City Challenge

arXiv:2608.17044v1 Announce Type: cross Abstract: The 10th AI City Challenge, held with ECCV 2026, marks a decade of community benchmarking for intelligent transportation, smart cities, and physical AI. Since its 2017…

Source: arXiv cs.AI Zheng Tang, Shuo Wang, David C. Anastasiu, Ming-Ching Chang, Anuj Sharma, Quan Kong, Munkhjargal Gochoo, Jun-Wei Hsieh, Tomasz Kornuta, Zhedong Zheng, Renran Tian, Judah Goldfeder, Fulgencio Navarro,…
AI Research AI

Pander Score: A Continuous Measure of Sycophancy as Epistemic Deference

arXiv:2606.07897v2 Announce Type: replace Abstract: Current AI models frequently exhibit epistemic sycophancy, endorsing claims to agree with a user. Existing evaluations typically measure this either by assessing what…

Source: arXiv cs.AI Alejandro Botas, Paul de Font-Reaulx, Luke Hewitt
AI Research AI

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

arXiv:2608.17800v1 Announce Type: new Abstract: Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely…

Source: arXiv cs.AI Liya Zhu, Xin Ma, Tao Liu, Haodong Wang, Ge Zhang, Jingzhe Ding, Qingshui Gu, Yongjie Zhong, Jinxiang Meng, Yuan Gao, Yunqiu Zhou, Hao Zhu, Jifeng He, Yongzhi Liao, Xinyi Zhang, Chaoxin Li, Yi Zhu, X…
AI Research AI

Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing

arXiv:2509.12040v3 Announce Type: replace-cross Abstract: Open-Vocabulary Remote Sensing Image Segmentation (OVRSIS), an emerging task that adapts Open-Vocabulary Segmentation (OVS) to the remote sensing (RS) domain,…

Source: arXiv cs.AI Bingyu Li, Haocheng Dong, Da Zhang, Zhiyuan Zhao, Junyu Gao, Xuelong Li
AI Research AI

Nonadaptive Learning in Robust Nonlinear Output Regulation

arXiv:2608.17262v1 Announce Type: cross Abstract: This paper considers robust nonadaptive regulation for general nonlinear systems in an output-feedback setting with arbitrarily high relative degree. We develop a…

Source: arXiv cs.AI Shimin Wang, Martin Guay, Richard D. Braatz
AI Research AI

Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight

arXiv:2510.08713v3 Announce Type: replace Abstract: Enabling embodied agents to imagine future states is essential for robust and generalizable visual navigation. Yet, state-of-the-art systems typically rely on modular…

Source: arXiv cs.AI Yifei Dong, Fengyi Wu, Guangyu Chen, Lingdong Kong, Qiyu Hu, Yuxuan Zhou, Xu Zhu, Jingdong Sun, Jun-Yan He, Qi Dai, Alexander G. Hauptmann, Zhi-Qi Cheng
AI Research AI

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

arXiv:2608.17393v1 Announce Type: new Abstract: Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository contexts, and execution feedback.…

Source: arXiv cs.AI Yiming Du, Yuxin Jiang, Tao Yuan, Jianbo Dai, Shaowei Wang, Jierun Chen, Chaofan Tao, Xianzhi Yu, Lifeng Shang, Kam-Fai Wong, Xiaohui Li, Haoli Bai
AI Research AI

Accuracy and Robustness of Model Cascades Under Data Perturbations

arXiv:2608.17711v1 Announce Type: new Abstract: Prediction cascades significantly reduce energy consumption of Artificial Intelligence (AI) models while maintaining high predictive performance. The idea is that easy…

Source: arXiv cs.AI Pallavi Mitra, Jai Kushwaha, Felix Biessmann
AI Research AI

Planning under Distribution Shifts with Causal POMDPs

arXiv:2602.23545v3 Announce Type: replace Abstract: In the real world, planning is often challenged by distribution shifts. As such, a model of the environment obtained under one set of conditions may no longer remain…

Source: arXiv cs.AI Matteo Ceriscioli, Karthika Mohan
AI Research AI

Wuying-Browser-Agent: Real-World Centric Fundamental Long-Horizon Browser Agents

arXiv:2608.17319v1 Announce Type: new Abstract: Browser agents perform well on short, clean demonstrations, but real deployment is fundamentally different: agents must sustain dozens of decisions on live websites while…

Source: arXiv cs.AI AIMAE Team, Tianxiang Chen, Yan Cheng, Zhangye Han, Xiaowei Li, Chang Liu, Cheng Liu, Zhongqiang Ma, Long Peng, Xiaobing Tu, Yinggui Wang, Hongliang Wei, Chen Wu, Daiping Xin, Kunyu Zhou, Pengyang Zh…
AI Research AI

Rationale-Guided Learning for Multimodal Emotion Recognition

arXiv:2608.10448v2 Announce Type: replace Abstract: Multimodal emotion recognition in conversation (MERC) requires understanding complex interactions between verbal and non-verbal cues. However, most existing approaches…

Source: arXiv cs.AI Sujung Oh, Jung Uk Kim, Sangmin Lee
AI Research AI

When Agents Act on Web3: An Attack-Surface Survey of MCP, Skills, and Tool Calling

arXiv:2608.17275v1 Announce Type: cross Abstract: AI agents increasingly act rather than merely read: across the Model Context Protocol (MCP) ecosystem, the share of deployed tools that modify external state has risen…

Source: arXiv cs.AI Rabimba Karanjai (Larry), Yang Lu (Larry), Nour Diallo (Larry), Wujie Xiong (Larry), Lei Xu (Larry), Weidong (Larry), Shi
AI Research AI

The Problem Is the Problem: Towards Scalable Mathematical Discovery

arXiv:2608.16977v1 Announce Type: new Abstract: AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier-model reasoning is a limited resource, and expert mathematical…

Source: arXiv cs.AI Zeyu Zheng, Shengtong Zhang, Jeremy Avigad, Prasad Tetali, Sean Welleck