Skip to content
TILens What matters today in tech
Theme

Topic · Edition

AI

540 items on 13 Aug 2026
AI GitHub

langchain-ai/langchain: langchain-anthropic==1.5.6

Changes since langchain-anthropic==1.5.5 release(anthropic): 1.5.6 (#39622) fix(anthropic): normalize tool_search_tool_result blocks (#39621) fix(anthropic): correct model profile data for Fable 5, Sonnet 5, Opus 4.1…

Source: LangChain Releases github-actions[bot]
AI

How kids feel about AI, in their own words

When we set out to talk to kids about artificial intelligence, we thought we knew what we’d hear. We expected some to tell us they were using it to cheat a little, the way Millennials and Gen Xers opened up CliffsNotes…

Source: MIT Technology Review AI Jen Swetzoff, Keeley McNamara
AI

Agentic Bayesian Optimization through Surrogate-Augmented Autoresearch

arXiv:2608.00316v2 Announce Type: replace Abstract: Bayesian optimization (BO) has become the standard tool for sample-efficient optimization and owes its efficiency to uncertainty-aware search driven by generic…

Source: arXiv cs.LG Paul Brunzema, Louis Tiao, Nhat Le, Kevin De Angeli, Yao Xuan, Djordje Gligorijevic
AI

Dion3: Full-Stack Orthogonal Updates

arXiv:2608.11612v1 Announce Type: new Abstract: The Muon optimizer incurs a significant overhead cost due to its cubic-time Newton-Schulz orthogonalization step. When weights are sharded, communication overhead…

Source: arXiv cs.LG Noah Amsel, Jack Zhang, Kwangjun Ahn, Ali Naeimi, Austin Feng, Berlin Chen, Tri Dao, John Langford
AI

Kernel Methods for Learning Operators with Multiple Inputs and Outputs

arXiv:2608.11831v1 Announce Type: new Abstract: Learning mappings between infinite-dimensional objects is a central challenge in scientific machine learning. We introduce a general kernel-based encoder-decoder framework…

Source: arXiv cs.LG Adrien Weihs, Chunyang Liao, Jingmin Sun, Hayden Schaeffer
AI

Dueling Deep Q-Learning for Intrusion Detection

arXiv:2608.11291v1 Announce Type: cross Abstract: Intrusion detection systems (IDS) and automated systems for detecting and reporting cyber threats, are commonly handled via supervised machine learning methods. Though…

Source: arXiv cs.LG Logan Luna (Georgia Institute of Technology), Matthew P. Berkowitz (Embry-Riddle Aeronautical University), Laxima Niure Kandel (Embry-Riddle Aeronautical University), Sirio Jansen-S'anchez (Embry-Rid…
AI

Towards Truly Unsupervised Evaluation of Feature Selection

arXiv:2608.12057v1 Announce Type: new Abstract: Feature selection is one of the most important and fundamental tasks in data mining, tackled by a family of methods with an established set of evaluation techniques to…

Source: arXiv cs.LG Hafiz Saud Arshad, Muhammad Rajabinasab, Arthur Zimek
AI

NAE: Normalizing AutoEncoder

arXiv:2608.12084v1 Announce Type: new Abstract: We consider the setting of Normalizing flows with approximate inverses, an established paradigm spanning both full-dimensional ($d=D$) and bottleneck ($d

Source: arXiv cs.LG Muhammad Abdur Rafae, Niels Landwehr
AI

Better, Faster, Stronger: Programmatic Skill Learning Best Reduces Agent Cost

arXiv:2608.11338v1 Announce Type: cross Abstract: Recently, the practice of augmenting LLM agent capability with skills has gained prevalence. We explore the cost effective adaptation of agents to novel domains by means…

Source: arXiv cs.LG Zixi Huang, Xiheng Wang, Andrew Wang, William Jurayj, Bernal Jim\'enez Guti\'errez, Daniel Khashabi, Nicholas Andrews
AI

Policy-as-logic for robust reasoning over rules

arXiv:2608.11905v1 Announce Type: cross Abstract: In many practical applications of generative AI systems, from tax rules to airline baggage allowance, responses to natural language queries must respect written policies…

Source: arXiv cs.LG Rahul Nair, Bastian Lipka, Elizabeth Daly
AI

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

arXiv:2608.12036v1 Announce Type: cross Abstract: AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood.…

Source: arXiv cs.LG Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhan…
AI

OrderMoE: An expert similarity driven distributed edge MoE inference

arXiv:2607.17154v2 Announce Type: replace-cross Abstract: Although mixture-of-experts, MoE, models have been increasingly adopted to scale large language models with moderate computation cost, it remains challenging to…

Source: arXiv cs.LG Xin Yuan, Ning Li, Quan Chen, Wenchao Xu, Song Guo
AI

Diffusion-Based Data-Driven Assortment Optimization

arXiv:2608.11419v1 Announce Type: new Abstract: Assortment optimization is a fundamental problem in revenue management, typically addressed using parametric choice models such as the multinomial logit (MNL) and its…

Source: arXiv cs.LG Junyi Liao, Xiaohui Jiang, Zhengwei Tong, Ethan X. Fang, Vahid Tarokh
AI

Do Judges Behave Like Algorithms?

arXiv:2608.10400v2 Announce Type: replace Abstract: What if judges already behave like algorithms? As artificial intelligence and algorithms are deployed in many settings, including the judicial system, many have…

Source: arXiv cs.LG Riya Manchanda, Eric Chen, Chloe Zhu, Cynthia Rudin, Brandon Garrett, Songman Kang
AI

Modeling Spectral Energy Shifts in Spatio-Temporal Graph Anomaly Detection

arXiv:2606.00304v2 Announce Type: replace Abstract: Graph anomaly detection methods aim to distinguish anomalous nodes. While prior methods characterize anomalies through increased variation in the spectral energy…

Source: arXiv cs.LG Yilin Liu, Hongchao Zhang, Taylor T. Johnson, Ahmad F. Taha, Meiyi Ma
AI

Chain-of-Thought Shows the Path to a Tree: Realizing Branching Complexity

arXiv:2608.11716v1 Announce Type: new Abstract: Chain of Thought (CoT) lifts the expressive ceiling of bounded-depth Transformers, with characterizations tying the number of CoT steps to circuit complexity classes. What…

Source: arXiv cs.LG Debanjan Dutta, Anish Chakrabarty, Swagatam Das
AI

Redistribution-based Cost Inference Improves Sparse Safe Offline RL

arXiv:2608.12306v1 Announce Type: new Abstract: Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level stop-feedback: a binary signal at…

Source: arXiv cs.LG Ebenezer Gelo (University of the Witwatersrand), Geraud Nangue Tasse (University of the Witwatersrand), Steven James (University of the Witwatersrand), Benjamin Rosman (University of the Witwatersran…
AI

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation

arXiv:2608.11698v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolation methods such as ExOPD…

Source: arXiv cs.LG Yang Sun, Lichao Ma, Houyuan Qin, Yuxin Liu, Hanyang Lu, Yao Zhu, Pinlong Cai, Guohang Yan
AI

Improving Performance of Spike-based Deep Q-Learning using Ternary Neurons

arXiv:2506.03392v3 Announce Type: replace Abstract: We propose a new ternary spiking neuron model to improve the representation capacity of binary spiking neurons in deep Q-learning. Although a ternary neuron model has…

Source: arXiv cs.LG Aref Ghoreishee, Abhishek Mishra, John Walsh, Anup Das, Nagarajan Kandasamy
AI

Earth observation embeddings are effective sub-grid descriptors for probabilistic weather downscaling

arXiv:2608.12271v1 Announce Type: new Abstract: Global weather reanalyses and forecasts resolve the evolving atmospheric state on coarse grids, but site-specific applications require predictions at arbitrary locations…

Source: arXiv cs.LG Pedro Sousa (Department of Computer Science, University of Cambridge), Will Tebbutt (Department of Engineering, University of Cambridge), Sadiq Jaffer (Department of Computer Science, University of C…
AI

Clustered Randomized Smoothing for Stochastic Prediction Functions

arXiv:2608.12037v1 Announce Type: new Abstract: Modern stochastic predictors can model rich, multi-modal outcome distributions. However, this expressive power comes with challenges in ensuring robust predictions $-$ a…

Source: arXiv cs.LG Eduardo Figueiredo, Frederik Mathiesen, Julian Schumann, Jens Kober, Arkady Zgonnikov, Luca Laurenti
AI

Probably Approximately Correct Maximum A Posteriori Inference

arXiv:2601.16083v2 Announce Type: replace Abstract: Computing the conditional mode of a distribution, better known as the maximum a posteriori (MAP) assignment, is a fundamental task in probabilistic inference. However,…

Source: arXiv cs.LG Matthew Shorvon, Frederik Mallmann-Trenn, David S. Watson
AI

Patch-based Memory Gate Model in Time Series Foundation Model

arXiv:2509.18751v4 Announce Type: replace Abstract: Recently reconstruction-based deep models have been widely used for time series anomaly detection, but as their capacity and generalization capability increase, these…

Source: arXiv cs.LG Samuel Yoon, Jongwon Kim, Juyoung Ha, Young Myoung Ko
AI

Optimized Deferral for Imbalanced Settings

arXiv:2604.27723v2 Announce Type: replace Abstract: Learning algorithms can be significantly improved by routing complex or uncertain inputs to specialized experts, balancing accuracy with computational cost. This…

Source: arXiv cs.LG Corinna Cortes, Anqi Mao, Mehryar Mohri, Yutao Zhong
AI

Temperature-Driven Sequential Modeling for the Prediction of Annual Power Conversion Efficiency Profiles of Organic Photovoltaic Materials: Douala Case Study

arXiv:2608.11261v1 Announce Type: cross Abstract: Organic photovoltaic (OPV) materials are promising candidates for distributed solar energy in tropical regions, yet existing virtual screening tools report static power…

Source: arXiv cs.LG Steve Cabrel Teguia Kouam, Rockefeller Rockefeller, Raoult Dabou Teukam, Jean-Pierre Tchapet Njafa, Patrick Sorrel Mvoto Kongo, Jean-Pierre Nguenang, Serge Guy Nana Engo
AI

ADEPT: A Unified Framework for Deep Learning Test Adequacy

arXiv:2608.12144v1 Announce Type: cross Abstract: Over the past decade, many test adequacy metrics have been proposed for deep learning that characterize test dataset adequacy from different perspectives, e.g., neuron…

Source: arXiv cs.LG Yidi Kao, Shawn Burnham, Tommi Rose Fahy, Ali Ghanbari
AI

Weightless Fine-Tuning: Personalizing LLMs via Logit-Space Transport

arXiv:2608.11342v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) is a standard approach for adapting LLMs to a target distribution, but in settings such as personalization, where each author requires…

Source: arXiv cs.LG Bohan Zhang, Anqi Ni, Yixin Wang, Paramveer S. Dhillon
AI

On Data-Driven Koopman Representations of Nonlinear Delay Differential Equations

arXiv:2604.03086v2 Announce Type: replace-cross Abstract: This work establishes a rigorous bridge between infinite-dimensional delay dynamics and finite-dimensional Koopman learning, with explicit and interpretable…

Source: arXiv cs.LG Santosh Mohan Rajkumar, Dibyasri Barman, Kumar Vikram Singh, Debdipta Goswami
AI

Pretraining large language models with MXFP4 on Native FP4 Hardware

arXiv:2605.09825v4 Announce Type: replace Abstract: Why does full-pipeline FP4 training of large language models often diverge, even when forward activations and activation gradients remain stable? We address this…

Source: arXiv cs.LG Musa Cim, Sarthak Arora, Poovaiah Palangappa, Miro Hodak, Ravi Dwivedula, Meena Arunachalam, Mahmut Taylan Kandemir
AI

Task- and dataset-specific information in protein language models

arXiv:2608.12090v1 Announce Type: new Abstract: Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of…

Source: arXiv cs.LG Roman Joeres, Ilya Senatorov, Olga V. Kalinina
AI

Confidence Calibration of Deep Learning Systems

arXiv:2608.12100v1 Announce Type: new Abstract: In high-stakes applications, reliable confidence estimates are as important as the predictions themselves. Confidence calibration ensures that predicted probabilities…

Source: arXiv cs.LG Coby Penso
AI

Hardware-Aware Deployment of Joint SAR Compression and Despeckling on FPGA

arXiv:2608.11271v1 Announce Type: cross Abstract: Next-generation Synthetic Aperture Radar (SAR) missions will generate data far faster than they can downlink, making onboard data reduction essential for near-real-time…

Source: arXiv cs.LG C\'edric L\'eonard, Francescopaolo Sica, Martin Schulz
AI

FQTree: Fine-grained Quantization and Hardware Generation of Boosted Decision Trees

arXiv:2608.12140v1 Announce Type: cross Abstract: Boosted decision trees (BDTs) are widely used in latency-critical applications, but efficient hardware deployment remains challenging. Existing designs often rely on…

Source: arXiv cs.LG Zhiqiang Que, Chang Sun, Haiyang Wang, Dinesh Pamunuwa, Roshan Weerasekera, Qijia Tang, Bakhtiar Zadeh, Wayne Luk, Maria Spiropulu
AI

Long-Horizon Forecasting of Complete Financial Statements with Forma

arXiv:2608.11327v1 Announce Type: new Abstract: Specialist training beats generalist scale when forecasting financial statements. To our knowledge, no prior work jointly forecasts complete financial statements beyond…

Source: arXiv cs.LG Travis L. Johnson, Jiannan Jiang, Soumyabrata Chaudhuri, Yihao Chen, Lauren Falvey, Donal O'Cofaigh
AI

Draw This First

arXiv:2608.12064v1 Announce Type: cross Abstract: We invert the typical formulation of sketch generation: instead of drawing strokes in order, we predict a 2D field that defines the order in which strokes are drawn. We…

Source: arXiv cs.LG Dazhi Zhong, Rowan Bradbury, Grant Davis
AI

Learning Multi-Timescale Interventions under Safety and Resource Constraints

arXiv:2508.03875v3 Announce Type: replace Abstract: Many sequential decision problems offer qualitatively different ways of influencing the environment: some interventions act immediately, whereas others induce…

Source: arXiv cs.LG David Mguni, Wanrong Yang, Jing Dong, Jing Peng, Ziquan Liu, Muhammad Salman Haleem, Baoxiang Wang, Dominik Wojtczak
AI

From Monolithic to Modular: Segment-level Automatic Prompt Optimization

arXiv:2608.11219v1 Announce Type: cross Abstract: Automatic Prompt Optimization (APO) often rewrites prompts monolithically, which can improve one behavior while degrading others. We present SAPO, a segment-level APO…

Source: arXiv cs.LG Nikita Kulin, Viktor Zhuravlev, Artur Khairullin, Sergey Muravyov, Ilya Makarov, Daniil Sukhorukov, Ekaterina Averkova
AI

Robustness of AI-Art Detectors under Generator Shift

arXiv:2608.11643v1 Announce Type: cross Abstract: Text-to-image generative models have advanced rapidly, with modern Diffusion Transformer architectures producing images that are increasingly difficult to distinguish…

Source: arXiv cs.LG Shivank Singh Thakur, Meien Li, Mark Stamp
AI

Towards the Harness of Embodied Agents

arXiv:2608.11246v1 Announce Type: cross Abstract: The success of coding agents has established the harness as a paradigm: what an agent achieves depends not on the model alone, but on the infrastructure around it. We…

Source: arXiv cs.LG Qi Wang, Tianyi Wang, Chengyang Li, Shikun Ban, Yurun Chen, Yizhong Ge, Jason Qin, Chengtai Li, Wentao Zhu
AI

ScreenShot: A Foundation Model for Few-Shot Combination Drug Screening

arXiv:2608.12219v1 Announce Type: new Abstract: Treating patients with combinations of drugs reduces the risk of resistance to any individual drug. Finding effective combinations is difficult because the large search…

Source: arXiv cs.LG Antoine de Mathelin, Christopher Tosh, Wesley Tansey
AI

Accelerating Time Series Foundation Models with Speculative Decoding

arXiv:2511.18191v2 Announce Type: replace Abstract: Time series forecasting drives operational decisions under tight latency budgets, and autoregressive time series foundation models (TSFMs) increasingly deliver the…

Source: arXiv cs.LG Pranav Subbaraman, Fang Sun, Jinxi Yu, Yue Yao, Huacong Tang, Xiao Luo, Yizhou Sun
AI

BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents

arXiv:2511.20597v2 Announce Type: replace Abstract: The integration of artificial intelligence (AI) agents into web browsers introduces security challenges that go beyond traditional web application threat models. Prior…

Source: arXiv cs.LG Kaiyuan Zhang, Mark Tenenholtz, Kyle Polley, Jerry Ma, Denis Yarats, Ninghui Li
AI

Local Cluster Cardinality Estimation for Adaptive Mean Shift

arXiv:2508.12450v2 Announce Type: replace Abstract: This article presents an adaptive mean shift algorithm in which every parameter used at a point is derived from that point's own distance distribution. The distance…

Source: arXiv cs.LG \'Etienne Pepin
AI

Attractor Image-Based Deep Learning of Arterial Pulse Waves for Age Classification

arXiv:2608.12117v1 Announce Type: new Abstract: Arterial pulse waveform morphology evolves with age, reflecting structural and functional changes in the cardiovascular system. Thus, vascular age is a valuable surrogate…

Source: arXiv cs.LG Sara Vardanega, Patrick Segers, Philip Aston, Ernst Rietzschel, Jordi Alastruey, Manasi Nandi
AI

Click2Poly: A VLM for vector mapping buildings and walls

arXiv:2608.11424v1 Announce Type: new Abstract: Accurate vector mapping of buildings and walls is critical for geospatial applications but remains a labor-intensive process. While recent deep learning methods have…

Source: arXiv cs.LG Nicolas Girard, Jawher Ben Abdallah, Arno Gobbin, Liuyun Duan, Sacha Lepretre
AI

LoRAQuant: Mixed-Precision Quantization of LoRA to Ultra-Low Bits

arXiv:2510.26690v3 Announce Type: replace Abstract: Low-Rank Adaptation (LoRA) has become a popular technique for parameter-efficient fine-tuning of large language models (LLMs). In many real-world scenarios, multiple…

Source: arXiv cs.LG Amir Reza Mirzaei, Yuqiao Wen, Yanshuai Cao, Lili Mou
AI

Soft-Attention Improves Skin Cancer Classification Performance

arXiv:2105.03358v4 Announce Type: replace-cross Abstract: In clinical applications, neural networks must focus on and highlight the most important parts of an input image. Soft-Attention mechanism enables a neural…

Source: arXiv cs.LG Soumyya Kanti Datta, Seyed Mohammad Abuzar Hashemi, Sargur N. Srihari, Mingchen Gao
AI

Reducing Symmetry Increase in Equivariant Neural Networks

arXiv:2608.12010v1 Announce Type: new Abstract: Equivariant Neural Networks (ENNs) have empowered numerous applications in scientific fields. Despite their remarkable capacity for representing geometric structures, ENNs…

Source: arXiv cs.LG Ning Lin, Jiacheng Cen, Anyi Li, Wenbing Huang, Hao Sun
AI

Ranking vs. Assignment: The Metric Mismatch in Multi-View Object Association

arXiv:2606.02022v2 Announce Type: replace-cross Abstract: Multi-view object association is an important computer vision problem that underlies many multi-camera perception tasks. While this task is naturally formulated…

Source: arXiv cs.LG Matvei Shelukhan, Timur Mamedov, Aleksandr Chukhrov, Karina Kvanchiani
AI

SoftWater: Class-Aware Rate Allocation for Softmax Quantization

arXiv:2608.12026v1 Announce Type: new Abstract: Post-training quantization pipelines routinely leave the softmax output layer in high precision. Yet in small LLMs with modern vocabularies, the head holds 15--30\% of all…

Source: arXiv cs.LG Joao V. Cavalcanti, Ashia C. Wilson
AI

A Factor Graph Approach to Scalable Multi-Output Gaussian Process Regression

arXiv:2608.11917v1 Announce Type: new Abstract: Multi-output Gaussian process regression scales cubically in the number of observations times outputs, and dense kernel-matrix methods need bespoke handling whenever…

Source: arXiv cs.LG Wouter W. L. Nuijten, Esther G. van Pelt, Albert Podusenko, \.Ismail \c{S}en\"oz, Wouter M. Kouw
AI

MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

arXiv:2608.11616v1 Announce Type: cross Abstract: Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only…

Source: arXiv cs.LG Hojun Choi, Jaeyo Shin, Suin Lee, Hyunjung Shim
AI

LLM Router: Rethinking Routing with Prefill Activations

arXiv:2603.20895v3 Announce Type: replace-cross Abstract: Existing routers rely on semantic query features or handcrafted features, which often fail to capture model-specific failures or intrinsic task difficulty. We…

Source: arXiv cs.LG Tanay Varshney, Annie Surla, Michelle Xu, Gomathy Venkata Krishnan, Maximilian Jeblick, David Austin, Neal Vaidya, Davide Onofrio
AI

RelShap: Relationally Consistent Shapley Explanations

arXiv:2608.11508v1 Announce Type: new Abstract: Machine learning pipelines commonly flatten relational data into single-table representations, discarding structural constraints. Widely used Shapley value-based feature…

Source: arXiv cs.LG Seungeun Lee, Joao Fonseca, Julia Stoyanovich
AI

Continual Learning in Transition

arXiv:2608.06216v2 Announce Type: replace Abstract: Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechanisms, e.g., training…

Source: arXiv cs.LG Zhiyan Hou, Dan Zhang, Tao Feng, Liyuan Wang, Wei Li, Xiangzhao Hao, Hongyan An, Junfeng Fang, Haokai Ma, Zhaohui Xu, Xinyu Tang, Haiyun Guo, Jinqiao Wang, Tat-Seng Chua
AI

Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL

arXiv:2608.11669v1 Announce Type: new Abstract: Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models on tasks with no deterministic…

Source: arXiv cs.LG Minglai Yang, Xinyu Guo, Utkarsh Tyagi, Mian Zhang, Razvan Dumitru, Sunjie Hou, Yunzhong He, Daniel Yue Zhang, Ying Liu
AI

Why AI Detection Fails for Academic Integrity

arXiv:2608.11256v1 Announce Type: new Abstract: Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct. In a…

Source: arXiv cs.LG Jonathan A. Karr Jr, Grigorii Khvatskii, Ting Hua, Nitesh V. Chawla
AI

Spectral graph clustering with inhomogeneous latent geometry

arXiv:2608.11321v1 Announce Type: cross Abstract: We study spectral clustering in the presence of a confounding latent geometry. The leading eigenvectors may then be dominated by the latent geometry rather than by the…

Source: arXiv cs.LG Konstantin Avrachenkov, Lucas S. Sibemberg, Alexander Van Werde
AI

MMLA: How Memory Lets the Past Shape the Future

arXiv:2606.28876v3 Announce Type: replace-cross Abstract: Proposal. Long context can replay history, but it does not decide which completed observations deserve authority. MMLA formalizes a bounded resident memory…

Source: arXiv cs.LG Junyi Zou, Avrova Donz
AI

Basin: Efficient and Extensible Numerical Optimization in Rust

arXiv:2608.11279v1 Announce Type: new Abstract: Basin is a numerical optimization library for the Rust programming language. Numerical optimization is the task of finding the inputs that minimize a function, and it is a…

Source: arXiv cs.LG Johan Larsson
AI

Dynamics Models for Offline Hyperparameter Selection in Real-World RL

arXiv:2608.11349v1 Announce Type: new Abstract: A key obstacle to deploying reinforcement learning in real-world systems is hyperparameter selection, particularly when simulators are unavailable and online…

Source: arXiv cs.LG Jordan Coblin, Han Wang, Martha White, Adam White
AI

Clinical Feasibility of Low-Magnification Fluorescence Imaging for Breast Cancer Margin Detection Using Texture Analysis and Deep Learning

arXiv:2608.11317v1 Announce Type: cross Abstract: High-resolution images of unprocessed surgical breast tissue can be obtained using microscopy with ultraviolet surface excitation (MUSE). This technique is considered a…

Source: arXiv cs.LG Pouya Afshin, Tianling Niu, Tongtong Lu, David Helminiak, Julie Jorns, Mollie Patton, Tina Yen, Donghye Ye, Bing Yu
AI

GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs

arXiv:2608.11674v1 Announce Type: new Abstract: On-policy rollout methods such as GRPO are central to post-training of large language models, yet they frequently suffer from training instabilities, cross-task capability…

Source: arXiv cs.LG Kai Yang, Jingwei Xu, Wanyu Wang, Kai-Yuan Guo, Zhenbo Yu, Yi Wang, Yu Qiao
AI

Dual-Primal Graph VAEs for Noisy Label Aggregation

arXiv:2608.11473v1 Announce Type: new Abstract: Inferring the ground-truth from noisy crowdsourced labels is an important theoretical and practical problem. Neural network-based methods offer an alternative to classical…

Source: arXiv cs.LG Patrick Stinson, Nikolaus Kriegeskorte
AI

Adaptive Online Learning with LSTM Networks for Energy Price Prediction

arXiv:2510.16898v2 Announce Type: replace Abstract: Accurate prediction of electricity prices is crucial for stakeholders in the energy market, particularly for grid operators, energy producers, and consumers. This…

Source: arXiv cs.LG Salih Salihoglu, Ibrahim Ahmed, Afshin Asadi
AI

An Efficient Near-Optimal Algorithm for Adversarial $m$-Set Bandits

arXiv:2608.12231v1 Announce Type: new Abstract: We study adversarial combinatorial bandits with $m$-set actions, where at each round the learner selects $m$ out of $d$ items and observes only the aggregate loss of the…

Source: arXiv cs.LG Francesco Bacchiocchi, Tommaso Cesari, Roberto Colomboni
AI

Mind the Gap: Structure-Aware Consistency in Preference Learning

arXiv:2604.27733v2 Announce Type: replace Abstract: Aligning Large Language Models (LLMs) with human intent, whether through explicit reward modeling or direct methods such as DPO, fundamentally relies on minimizing a…

Source: arXiv cs.LG Mehryar Mohri, Yutao Zhong
AI

ED-CSP: Crystal Structure Prediction from Electron Diffraction

arXiv:2608.06448v2 Announce Type: replace Abstract: Recovering a periodic 3D crystal structure from sparse, unindexed electron diffraction (ED) observations is a challenging generative inverse problem. Existing ED-based…

Source: arXiv cs.LG Germain Poloudenny, Ya\"el Fr\'egier, Arnaud Demorti\`ere
AI

Reliable Inference in Edge-Cloud Model Cascades via Conformal Alignment

arXiv:2510.17543v3 Announce Type: replace Abstract: Edge intelligence enables low-latency inference via compact on-device models, but assuring reliability remains challenging. We study edge-cloud cascades that must…

Source: arXiv cs.LG Jiayi Huang, Sangwoo Park, Nicola Paoletti, Osvaldo Simeone
AI

Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing

arXiv:2608.11660v1 Announce Type: cross Abstract: Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes…

Source: arXiv cs.LG Tianci Liu, Zihan Dong, Tianchun Li, Yi-Chung Chen, Qiming Cao, Xingchen Wang, Shiyang Wang, Zichen Miao, Linjun Zhang, Haoyu Wang, Jing Gao
AI

A Modular Agentic Framework for Synthetically Constrained Multi-Objective Hit-to-Lead Optimization

arXiv:2608.11483v1 Announce Type: cross Abstract: Hit-to-lead optimization requires iterative design of hit analogs across competing potency, selectivity, physicochemical, pharmacokinetic, safety, and synthetic…

Source: arXiv cs.LG Kelvin P. Idanwekhai, Enes Kelestemur, Benjamin Strickland, Matthew Hart, Steini Davidsson, Angelos Angelopoulos, Ron Alterovitz, Marcello DeLuca, Alexander Tropsha
AI

Let it Cook: Learning to Wait in Sequential Decision Making

arXiv:2608.11511v1 Announce Type: new Abstract: In sequential decision making, an agent typically observes its environment and acts at every timestep. However, such active participation may not always be necessary;…

Source: arXiv cs.LG Christopher Watson, Arjun Krishna, Dinesh Jayaraman, Rajeev Alur
AI

Federated Learning for Distributed CNC Tool Wear Prediction

arXiv:2608.11281v1 Announce Type: new Abstract: Tool wear prediction is an important task in CNC machining, where accurate monitoring of tool condition supports product quality and process reliability. Machine learning…

Source: arXiv cs.LG Afsana Khan, Morris Stallmann, Marcin Pietrasik, Charis Kouzinopoulos, Anna Wilbik
AI

FLARE++: Low-rank attention with dynamic attention routing

arXiv:2608.11519v1 Announce Type: new Abstract: Full self-attention is a strong token mixer for PDE surrogates on irregular domains, but its quadratic cost limits its use on high-resolution problems. Efficient…

Source: arXiv cs.LG Vedant Puri, Yongjie Jessica Zhang, Levent Burak Kara
AI

MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning

arXiv:2608.11749v1 Announce Type: new Abstract: Multi-objective optimization (MOO) has demonstrated significant success in multi-task learning by mitigating task conflicts through gradient manipulation. However, most…

Source: arXiv cs.LG Shiji Zhou, Kunlin Lyu, Lei Zhang, Ruodong Wang, Yifan Sun
AI

Disentangling the Expressivity of RoPE

arXiv:2608.11909v1 Announce Type: new Abstract: Two accounts recur in explanations of the success of rotary position embeddings (RoPE). Expressivity studies associate periodic position information with modular…

Source: arXiv cs.LG Selim Jerad, Anej Svete, Jiaoda Li, Ryan Cotterell
AI

CAM-Guided Saliency Cutout and Image-Based Malware Classification

arXiv:2608.11634v1 Announce Type: cross Abstract: Dropout regularization is commonly used to reduce overfitting by removing parts of a neural network during training. For Convolutional Neural Networks (CNN), cutouts…

Source: arXiv cs.LG Yasaman Ebrahimi, Martin Jurecek, Mark Stamp
AI

Grounding Large Language Models as Generalizable Policies in Network Control

arXiv:2512.11839v2 Announce Type: replace Abstract: Designing generalizable control policies that operate reliably under changing conditions is essential for robust network services in modern digital infrastructure. Yet…

Source: arXiv cs.LG Duo Wu, Linjia Kang, Zhimin Wang, Fangxin Wang, Wei Zhang, Chongbo Sun, Xuefeng Tao, Wei Yang, Le Zhang, Wenwu Zhu, Peng Cui, Zhi Wang
AI

Tight Nonasymptotic Local Convergence of Sinkhorn-Knopp

arXiv:2608.11760v1 Announce Type: cross Abstract: We revisit the Sinkhorn-Knopp (SK) algorithm for the matrix scaling problem. Despite extensive literature on the global convergence of SK and its variants, its local…

Source: arXiv cs.LG Wenzhi Gao, Zhaonan Qu, Yinyu Ye, Madeleine Odell
AI

ODE-Based Transformer Decoders for Iterative Sign Language Translation

arXiv:2608.11352v1 Announce Type: cross Abstract: Sign language translation has achieved strong results with Transformer architectures, yet recent improvements largely rely on scaling model capacity at the cost of…

Source: arXiv cs.LG Tu\u{g}\c{c}e K{\i}z{\i}ltepe, Hacer Yalim Keles
AI

Planar Symmetric Pattern Generation

arXiv:2606.02073v2 Announce Type: replace Abstract: Generating objects with specific symmetries is essential in various real-world scenarios. However, adapting existing 2D continuous representations to enforce planar…

Source: arXiv cs.LG Ning Lin, Luxi Chen, Huaguan Chen, Jiacheng Cen, Chongxuan Li, Wenbing Huang, Hao Sun
AI

FACTR 2: Learning External Force Sensing for Commodity Robot Arms Improves Policy Learning

arXiv:2606.12406v2 Announce Type: replace-cross Abstract: Contact-rich manipulation requires force sensitivity, but many robot arms lack dedicated force sensors due to their high cost. We present Neural External Torque…

Source: arXiv cs.LG Steven Oh, Jason Jingzhou Liu, Tony Tao, Philip Han, Kenneth Shaw, Satoshi Funabashi, Ruslan Salakhutdinov, Deepak Pathak
AI

Reducing Per-Sample Interference in Stochastic Optimization

arXiv:2607.16261v2 Announce Type: replace Abstract: Modern optimizers combine gradients from the current mini-batch with historical optimization state, such as momentum or adaptive moments. While effective, this…

Source: arXiv cs.LG Apostolos Avranas
AI

Representation Finetuning for Continual Learning

arXiv:2603.11201v3 Announce Type: replace Abstract: The world is inherently dynamic, and continual learning aims to enable models to adapt to ever-evolving data streams. While pre-trained models have shown powerful…

Source: arXiv cs.LG Haihua Luo, Xuming Ran, Tommi K\"arkk\"ainen, Huiyan Xue, Zhonghua Chen, Qi Xu, Fengyu Cong
AI

TESLA: Taylor Expansion of Sinusoidal Learnable Activations

arXiv:2608.11970v1 Announce Type: new Abstract: The parity problem--deciding whether the number of ones in a binary vector is odd or even--remains challenging for standard neural networks due to linear inseparability…

Source: arXiv cs.LG Daehwa Ko, Jaehyeon Kim, Seunghyun Ham, Jay Hoon Jung
AI

Deep Activity Model: A Generative Approach for Human Mobility Pattern Synthesis

arXiv:2405.17468v3 Announce Type: replace Abstract: Human mobility plays a crucial role in transportation, urban planning, and public health, but current approaches face important limitations. Existing deep learning…

Source: arXiv cs.LG Xishun Liao, Qinhua Jiang, Brian Yueshuai He, Yifan Liu, Chenchen Kuai, Jiaqi Ma
AI

Variable Selection in the Context of AI Fairness

arXiv:2608.11251v1 Announce Type: cross Abstract: Fairness in AI systems has become more important with recent regulatory demands, such as the EU AI Act. Traditional approaches often do not take into account…

Source: arXiv cs.LG Ivan Luciano Danesi, Chiara Frigerio, Fabio Maccaferri, Giorgio Alessandro Motta, Pietro Zecca
AI

AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses

arXiv:2608.12307v1 Announce Type: new Abstract: Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters, through teacher forcing, on-policy…

Source: arXiv cs.LG Cheng Qian, Wenting Zhao, Liangwei Yang, Heng Wang, Jielin Qiu, Heng Ji, Silvio Savarese, Huan Wang, Shelby Heinecke
AI

Probing and steering biology across Boltz-1s trunk-diffusion boundary

arXiv:2608.11475v1 Announce Type: cross Abstract: AlphaFold3-class structure predictors pair a representational trunk, which processes sequence and context, with a diffusion module, which generates atomic coordinates.…

Source: arXiv cs.LG Piotr Jedryszek, Tongmeng Xie, Adam Winnifrith, Alexander Hasson, Weronika \'Slesak, George Wicks, Toby Winnifrith, Oliver M. Crook
AI

Look What the Probes Dragged In! Real-World Chest X-ray Shortcuts in MedCLIP

arXiv:2608.12086v1 Announce Type: cross Abstract: Vision-language models, such as contrastive language-image pre-training (CLIP)-based approaches, have reached state-of-the-art (SOTA) results in medical artificial…

Source: arXiv cs.LG Nikolette Pedersen, Regitze Sydendal, Veronika Cheplygina, Th\'eo Sourget
AI

Prompt-Driven Exploration

arXiv:2607.08837v2 Announce Type: replace Abstract: Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers. Standard methods inject stochasticity in the…

Source: arXiv cs.LG Sunshine Jiang, John Marangola, David Zhang, Raghuram Kowdeed, Ruiyang Luo, Nitish Dashora, Richard Li, Pulkit Agrawal, Zhang-Wei Hong
AI

Large-scale AI-Ready Data for Anti-Cancer Drug Response Modeling

arXiv:2608.11444v1 Announce Type: cross Abstract: Drug response prediction (DRP) models are an active area of research in pharmacogenomics, with growing potential to accelerate the identification of effective anticancer…

Source: arXiv cs.LG Vincent Lavelle, Yitan Zhu, Kaitlyn Marlor, Thomas Brettin, Rick Stevens
AI

Distillation of Foundation Models for Time-dependent PDEs

arXiv:2608.11937v1 Announce Type: new Abstract: Foundation models for time-dependent partial differential equations (PDEs) are trained on large and diverse collections of physical systems and can generalize effectively…

Source: arXiv cs.LG Daniel Musekamp, Boshra Ariguib, Andrei Manolache, Mathias Niepert
AI

Analytic Bridge Diffusions for Controlled Path Generation

arXiv:2605.02961v2 Announce Type: replace Abstract: Most modern bridge-diffusion methods achieve finite-time transport by specifying an interpolation, Schrodinger-bridge, or stochastic-control objective and then…

Source: arXiv cs.LG Michael Chertkov
AI

DREvo: Distilling Recalibrated Historical Experience for Harness Self-Evolution

arXiv:2607.26722v2 Announce Type: replace-cross Abstract: Harness plays a critical role in large language model agent performance, and building a high-performing harness requires substantial expert effort. Therefore,…

Source: arXiv cs.LG Hanghui Guo, Weijie Shi, Zhangze Chen, Shengxiang Xu, Yishu Wang, Yimei Zhang, Wangze Ni, Jia Zhu, Shimin Di
AI

FarSky: Task-Aware Latent-Space Coupling for Generative Intra-Hour Solar Forecasting

arXiv:2608.11254v1 Announce Type: new Abstract: Accurate solar irradiance forecasting is essential for the reliable integration of photovoltaic power into modern electricity grids. All-sky imagers (ASI) provide…

Source: arXiv cs.LG Yann Fabel, Bijan Nouri, Milon Miah, Niklas Blum, Luis F. Zarzalejo, Julia Kowalski, Robert Pitz-Paal
AI

Defending against Model Extraction for GNNs with Model Reprogramming

arXiv:2608.11495v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) serve as the backbone for high-stakes applications in Machine-Learning-as-a-Service (MLaaS). Still, their black-box deployment exposes them to…

Source: arXiv cs.LG Yan Wen, Zhenyi Wang, Heng Huang
AI

Gaussian Meta-Space Augmentation for Stacking Ensembles in Multimodal IPMN Risk Stratification

arXiv:2608.11472v1 Announce Type: cross Abstract: Pancreatic cancer is among the most lethal malignancies; risk stratification of intraductal papillary mucinous neoplasms (IPMNs) offers a crucial opportunity for early…

Source: arXiv cs.LG Max A. Nelson, Eminenur Sen Tasci, Zhixiang Wang, Zongwei Zhou, Halil Ertugrul Aktas, Andrea M. Bejar, Elif Keles, Ziliang Hong, S{\i}tk{\i} Safa Taflan, Muhammed Enes Tasci, Frank H. Miller, Michael…
AI

MaSRead: Content-Addressed Reading of Replicated Latent Stores

arXiv:2608.11218v1 Announce Type: cross Abstract: Independent agents that reason in latent space can share computed state as key-value cache fragments rather than text. Merged by a conflict-free replicated data type,…

Source: arXiv cs.LG Carlos Baquero, Lu\'is Brito, Jo\~ao Resende
AI

Small-Scale Experiments: Are We There Yet?

arXiv:2608.11859v1 Announce Type: new Abstract: Scaling laws promised cost-effective experiments; six years later, they have yet to fully deliver. Instead, researchers have found them unreliable at small scales…

Source: arXiv cs.LG Nicholas Lourie, Kyunghyun Cho, Karen Ullrich, Sanae Lotfi
AI

Unifying Physical Backpropagation

arXiv:2608.11585v1 Announce Type: cross Abstract: Physical computing systems exploit device dynamics for computation, but their gradient-based optimization is challenging: backpropagation through a digital twin suffers…

Source: arXiv cs.LG Cyrill B\"osch, Yigithan Gediz, Hakan T\"ureci
AI

Program Semantic Inequivalence Game with Large Language Models

arXiv:2505.03818v3 Announce Type: replace Abstract: Large Language Models (LLMs) can achieve strong performance on everyday coding tasks, but they can fail on complex tasks that require non-trivial reasoning about…

Source: arXiv cs.LG Antonio Valerio Miceli-Barone, Vaishak Belle, Ali Payani
AI

Adaptation of Generalist Robot Policies with Minimal Data

arXiv:2608.11363v1 Announce Type: cross Abstract: A central goal in robot learning is to move beyond task-specific human data collection toward robots that improve through autonomous interaction. Yet fully autonomous…

Source: arXiv cs.LG Shreyas Kowshik, Sreyas Venkataraman, Leo Wang, Niharika Pant, Max Simchowitz, Aviral Kumar
AI

The Advective Fisher-Rao Geometry of Deterministic Measure Transport

arXiv:2608.12111v1 Announce Type: cross Abstract: A novel advective Fisher-Rao metric is introduced for optimization tasks on paths of probability measures governed by the continuity equation. This metric is shown to…

Source: arXiv cs.LG Benjamin Gess, Johannes M\"uller
AI

Air Quality Station Simulation via LSTM and Attention-Based Modelling

arXiv:2608.11839v1 Announce Type: new Abstract: Poor air quality in urban areas is driven by a complex chain of processes and presents a significant public health concern. To better understand and control the mechanisms…

Source: arXiv cs.LG Alexander Kostadinov, Petar O. Hristov, Dessislava Petrova-Antonova
AI

Generative Learning for Quantum Measurement Design

arXiv:2608.11396v1 Announce Type: cross Abstract: Extracting quantum information from a quantum state is a fundamental task of quantum computation, often requiring the estimation of many non-commuting observables under…

Source: arXiv cs.LG Jun Dai, Olivier Nahman-L\'{e}vesque, Guillaume Rabusseau, Hong-Ye Hu, Cunlu Zhou
AI

Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling

arXiv:2608.11829v1 Announce Type: new Abstract: On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. It is commonly believed to enable the student model to distill…

Source: arXiv cs.LG Xinmu Ge, Zizhuo Zhang, Yu Huang, Jianing Zhu, Lin Yuan, Wanli Gu, Weichang Wu, Weiran Huang, Xiaolu Zhang, Bo Han, Jun Zhou, Jiangchao Yao
AI

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

arXiv:2608.12253v1 Announce Type: cross Abstract: Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach…

Source: arXiv cs.LG Simon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou, Dilara Soylu, Sergey Levine, Christopher D. Manning, Weiyan Shi
AI

TradingMoE: Routing the Right Experts in Evolving Markets

arXiv:2608.11785v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong potential for financial analysis and trading, but direct trading remains challenging because the predictive capabilities…

Source: arXiv cs.LG Chang Zhou, Xingtong Yu, Minbin Huang, Zhennan Wu, Yuan Fang, Hong Cheng, Xinming Zhang
AI

AutoGrable: What Is a Good Graph for a Table?

arXiv:2608.11431v1 Announce Type: new Abstract: Graph learning presupposes a graph, and tables and relational databases do not come with one. Applying a GNN to them requires deciding which entities become nodes, which…

Source: arXiv cs.LG Tamara Cucumides, Floris Geerts
AI

FunnelCausalNet: Funnel-aware Joint Conversion-Revenue Uplift for Multi-tier Coupon Allocation

arXiv:2608.11675v1 Announce Type: new Abstract: Coupon campaigns seek to lift both conversion and revenue, but gross merchandise value (GMV) follows a deterministic funnel from conversion to conditional order value and…

Source: arXiv cs.LG Yu Zhang (AMap Alibaba Group, Beijing, China), Zhihan Wang (AMap Alibaba Group, Beijing, China), Guanlin Chen (AMap Alibaba Group, Beijing, China), Min Jiang (AMap Alibaba Group, Beijing, China), Shu…
AI

Forecasting Side Effects of Activation Steering

arXiv:2608.11227v1 Announce Type: cross Abstract: Activation steering modifies a language model by adding a learned direction to its hidden activations, enabling targeted behavioral changes without retraining. While…

Source: arXiv cs.LG Chong Yong Ong, Alson Wei Jie Sim, Peixin Zhang, Jun Sun
AI

SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives

arXiv:2509.13450v3 Announce Type: replace-cross Abstract: We introduce SteeringSafety, a benchmark for evaluating representation steering methods across nine safety perspectives spanning 18 datasets. While prior work…

Source: arXiv cs.LG Vincent Siu, Nicholas Crispino, David Park, Nathan W. Henry, Zhun Wang, Yang Liu, Dawn Song, Chenguang Wang
AI

Benchmarking LLM Judges for Mobile Agent Evaluation

arXiv:2608.11434v1 Announce Type: cross Abstract: Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains…

Source: arXiv cs.CL Ziqiang Wan, Li Gu, Zhixiang Chi, Zhi Liu, Seyed Mehdi Ayyoubzadeh, Yuanhao Yu, Yang Wang
AI

Asymptotic Risk Calibration for Selective Question Answering

arXiv:2608.12008v1 Announce Type: new Abstract: Large language models (LLMs) may generate fluent but incorrect answers, making uncertainty quantification important for reliable question answering. However, heuristic…

Source: arXiv cs.CL Shufan Lin, Sijin Dong
AI

Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?

arXiv:2608.08283v2 Announce Type: replace Abstract: Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of…

Source: arXiv cs.CL Osvaldo Quinjica, Eric Bennett, Xinchen Yang, Andrew Schonebaum, Marine Carpuat
AI

MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques

arXiv:2608.11755v1 Announce Type: cross Abstract: Long-form song generation models continue to improve in duration, structural integrity, and acoustic complexity, making reliable aesthetic rewards increasingly important…

Source: arXiv cs.CL Jiabao Zhuang, Changhao Jiang, Hanchen Wang, Jiahao Chen, Zhixiong Yang, Zhenghao Xiang, Yifei Cao, Jiajun Sun, Hui Li, Ming Zhang, Tao Ji, Tao Gui, Qi Zhang, Xuanjing Huang
AI

Explicit Boundary Markers for Subword Vocabularies

arXiv:2608.08847v2 Announce Type: replace Abstract: Subword tokenizers represent many common words twice in space-using writing systems, once with a leading space and once without. The two entries have separate…

Source: arXiv cs.CL Sander Land, Clara Meister
AI

Text Corpora as Concept Fields: Black-Box Hallucination and Novelty Measurement

arXiv:2605.05103v3 Announce Type: replace Abstract: We introduce the \textbf{Concept Field} of a text corpus: a local drift field with pointwise uncertainty, estimated in sentence-embedding space from the deltas between…

Source: arXiv cs.CL Nicholas S. Kersting, Vittorio Castelli, Chieh Ting Yeh, Xinzhu Wang, Saad Taame, Khaoula Allak
AI

Investigating Learner-Aware Design of LLM-Generated Educational Feedback

arXiv:2602.11650v2 Announce Type: replace Abstract: Although large language models (LLMs) show promise for generating educational feedback, it remains unclear how feedback should be designed (e.g., tone and information…

Source: arXiv cs.CL Momoka Furuhashi, Kouta Nakayama, Noboru Kawai, Takashi Kodama, Saku Sugawara, Kyosuke Takami
AI

Multimodal QUD: Inquisitive Questions from Scientific Figures

arXiv:2604.23733v3 Announce Type: replace Abstract: Discourse comprehension in complex documents often involves continuously posing and resolving Questions Under Discussion (QUDs). While QUD frameworks have so far…

Source: arXiv cs.CL Yating Wu, William Rudman, Venkata S Govindarajan, Alexandros G. Dimakis, Junyi Jessy Li
AI

Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models

arXiv:2608.11742v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) have emerged as a competitive alternative to autoregressive language models, offering the potential for substantially faster…

Source: arXiv cs.CL Yushi Ye, Xu Chen, Haoyun Jiang, Jinsong Lan, Haihong Tang, Bo Han, Ivor Tsang, Yanfeng Wang, Bo Zheng, Jiangchao Yao
AI

Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness

arXiv:2608.09900v2 Announce Type: replace Abstract: Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow,…

Source: arXiv cs.CL Tadanobu Chuyo Kamijo, Ori Rottenstreich, Javier Conde, Gonzalo Mart\'inez, Pedro Reviriego
AI

Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs

arXiv:2608.11573v1 Announce Type: new Abstract: Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language models (LLMs). In this work, we…

Source: arXiv cs.CL Vu Duc Anh, Nhat M. Hoang, Do Xuan Long, Cong-Duy Nguyen, Ponhvoan Srey, Luu Anh Tuan
AI

TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation

arXiv:2608.11236v1 Announce Type: new Abstract: Roleplay evaluation should do more than assign a single score: it should reveal which role requirements were tested, which failed, and which dialogue evidence supports the…

Source: arXiv cs.CL Jiahui Zhang, Ziwei Zhang, Yipeng Wang, Yibo Liu, Haozhou Pang, Yikai Hu, Hongyan Ren, Lan Zhou, Qi Gan, Kai Sheng
AI

Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill

arXiv:2608.11924v1 Announce Type: new Abstract: Turning a research idea into a complete paper requires more than text generation: the system must retrieve literature, design and execute experiments, revise claims…

Source: arXiv cs.CL Zhuoyang Qian, Biao Wu, Yiran Wang, Chris D Yan, Desan Dai, Liangwei Zheng, Jin Jiang, Junsheng Zhang, Wenhao Wang
AI

SAG: SQL-Retrieval Augmented Generation with Query-Time Dynamic Hyperedges

arXiv:2608.12129v1 Announce Type: new Abstract: While retrieval-augmented generation (RAG) has proven effective at giving LLMs access to external knowledge, mainstream dense-retrieval implementations remain inherently…

Source: arXiv cs.CL Yuchao Wu, Junqin Li, XingCheng Liang, Yongjie Chen, Yinghao Liang, Linyuan Mo, Guanxian Li
AI

Easper: An Accessible ASR Pipeline for Language Documentation

arXiv:2608.11629v1 Announce Type: new Abstract: Audio transcription is a critical bottleneck in language documentation. While multilingual Automatic Speech Recognition (ASR) models like Whisper offer solutions, field…

Source: arXiv cs.CL Aso Mahmudi, Ting Dang, Ekaterina Vylomova, Nick Thieberger
AI

Gloss-Free Representation Learning for Cross-Dataset Sign Spotting

arXiv:2608.11332v1 Announce Type: new Abstract: Sign-language research for resource-constrained languages is often limited by the cost of dense linguistic labels such as glosses, temporal boundaries, and sign order.…

Source: arXiv cs.CL O\u{g}uz Akif T\"ufekcio\u{g}lu, Ezgi Ekin, Mustafa Kaan \c{C}evik, Hacer Yalim Keles
AI

Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models

arXiv:2605.27101v2 Announce Type: replace-cross Abstract: A key capability for video understanding is reliably linking subjects to events across time, yet whether Video Large Language Models (VideoLLMs) actually achieve…

Source: arXiv cs.CL Oscar Chew, Serhii Honcharenko, Qian-Hui Chen, Patricia Lu, Dishant Zaveri, Khoa D. Doan, Kuan-Hao Huang
AI

TEMPER: Testing Emotional Perturbation in Quantitative Reasoning

arXiv:2604.07801v2 Announce Type: replace Abstract: Large language models are trained and evaluated on quantitative reasoning tasks written in clean, emotionally neutral language. However, real-world queries are often…

Source: arXiv cs.CL Atahan Dokme, Benjamin Reichman, Larry Heck
AI

AVA-Encoder: Towards Agent-Native Video Representation Learning

arXiv:2608.12313v1 Announce Type: cross Abstract: Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the…

Source: arXiv cs.CL Chuyue Li, Jinpeng Yu, Haozhe Wang, Tian Xueyun, Zhijing Zhang, Bingnan Li, Shuqi Gu, Kan Ren, Jiaming Liu, Ruihua Hua
AI

RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation

arXiv:2608.12099v1 Announce Type: cross Abstract: We present RT-SEMamba, a fully causal speech enhancement (SE) model built upon causal time-frequency Mamba blocks. Unlike Transformer-based architectures that rely on a…

Source: arXiv cs.CL Rong Chao, Sung-Feng Huang, Moreno La Quatra, Sabato Marco Siniscalchi, Wen-Huang Cheng, Szu-Wei Fu, Yu Tsao
AI

VICBench: A Multi-Language Benchmark for Code Vulnerability Detection

arXiv:2608.12246v1 Announce Type: cross Abstract: Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that first introduce…

Source: arXiv cs.CL Jin Lu, Xuening Han, Yang Zhong, Lin Tan, Kevin Luo, Andrew Gacek, Neha Rungta
AI

Is Convergence Inevitable? Tracing Output Homogeneity Back to Base Models

arXiv:2608.11426v1 Announce Type: new Abstract: The lack of diversity in LM content is widely attributed to the alignment process, but how and where exactly in the pipeline this collapse begins is unknown. We argue that…

Source: arXiv cs.CL Alexandrine Fortier, Hazel Chen, Peter West
AI

ReXrank: A Public Leaderboard for AI-Powered Radiology Report Generation

arXiv:2411.15122v2 Announce Type: replace-cross Abstract: AI-driven models have demonstrated significant potential in automating radiology report generation for chest X-rays. However, there is no standardized benchmark…

Source: arXiv cs.CL Xiaoman Zhang, Hong-Yu Zhou, Xiaoli Yang, Oishi Banerjee, Juli\'an N. Acosta, Mohammed Baharoon, Josh Miller, Ouwen Huang, Pranav Rajpurkar
AI

Harnessing agent memory to build lifelong AI partners for materials scientists

arXiv:2608.11224v1 Announce Type: cross Abstract: Materials research advances through accumulated experience - scripts that work, protocols that are trusted, warnings attached to failed calculations or experiments, and…

Source: arXiv cs.CL Siyu Liu, Bo Hu, Beilin Ye, He Cao, David J. Srolovitz, Tongqi Wen
AI

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance

arXiv:2608.11694v1 Announce Type: new Abstract: A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways the same problem could be asked,…

Source: arXiv cs.CL Shailja Thakur, Sungeun An, Chad DeLuca, Hima Patel
AI

QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving

arXiv:2608.12121v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) repeatedly prefills identical text chunks across queries, incurring redundant computations. Position-Independent Caching (PIC)…

Source: arXiv cs.CL Yilin Liu, Rui Meng, Wangze Ni, Jianxin Yan, Heng Cao, Libin Zheng, Peng Cheng, Jinfei Liu
AI

LLM-Powered Automatic Translation and Urgency in Crisis Scenarios

arXiv:2602.13452v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly proposed for crisis preparedness and response, particularly for multilingual communication. However, their suitability…

Source: arXiv cs.CL Belu Ticona, Antonis Anastasopoulos
AI

Self-Harness: Harnesses That Improve Themselves

arXiv:2606.09498v2 Announce Type: replace Abstract: The performance of LLM-based agents is jointly shaped by their base models and the harnesses that mediate their interaction with the environment. Because different…

Source: arXiv cs.CL Hangfan Zhang, Shao Zhang, Kangcong Li, Chen Zhang, Yang Chen, Yiqun Zhang, Lei Bai, Shuyue Hu
AI

Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning

arXiv:2608.11408v1 Announce Type: new Abstract: Prior white-box studies show that large language models can retain latent traces of target knowledge after unlearning, even when the knowledge is no longer expressed in…

Source: arXiv cs.CL Zirui Song, Huaxing Liu, Xiang Wang, Shuai Li, Xinye Li, Lang Gao, Jinghui Zhang, Zheng Lu, Fengxian Ji, Xiaojun Chang, Xiuying Chen
AI

TELLME: Test-Enhanced Learning for Language Model Enrichment

arXiv:2608.11788v1 Announce Type: new Abstract: Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models. However, CPT has consistently been accompanied by…

Source: arXiv cs.CL Minjun Kim, Inho Won, Hyeonseok Lim, MinKyu Kim, Junghun Yuk, Wooyoung Go, Jongyoul Park, Jungyeul Park, KyungTae Lim
AI

Hybrid Gated Attention

arXiv:2608.11805v1 Announce Type: new Abstract: Gated attention is an effective approach to mitigate attention sinks and enhance the representational capacity of attention. To further extend its effectiveness-efficiency…

Source: arXiv cs.CL Zekun Zhou, Ruobing Xie, Lanrui Wang, Weixuan Sun
AI

Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs

arXiv:2608.11624v1 Announce Type: new Abstract: Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions. As…

Source: arXiv cs.CL Nimet Beyza Bozdag, Emre Can Acikgoz, Gokhan Tur, Dilek Hakkani-T\"ur
AI

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction

arXiv:2608.11772v1 Announce Type: new Abstract: Self-correction is particularly useful when a failure constrains the next repair. Coding agents benefit from this property because compilers, tests, and execution traces…

Source: arXiv cs.CL Pan Wang, Yihao Hu, Hang Wang, Zirui Lv, Xin Zhang, Jianshe Li, Jiang-Ming Yang, Wei Wu, Yongqi Tong
AI

Harness-G: A Graph-Structured Harness for Search Agents

arXiv:2607.27652v3 Announce Type: replace Abstract: Reinforcement learning (RL) search agents commonly model retrieval as free-form natural-language query generation and optimize multi-turn interactions using…

Source: arXiv cs.CL Yanning Hou, Haoyuan Chen, Sihang Zhou, Xiaoshu Chen, Xirui Liu, Duanyang Yuan, Lingyuan Meng, Siwei Wang, Quan Liu, Jian Huang
AI

Marco-Voice Technical Report

arXiv:2508.02038v5 Announce Type: replace Abstract: This paper presents a multifunctional speech synthesis system that integrates voice cloning and emotion control speech synthesis within a unified framework. The goal…

Source: arXiv cs.CL Fengping Tian, Chenyang Lyu, Xuanfan Ni, Haoqin Sun, Qingjuan Li, Zhiqiang Qian, Haijun Li, Longyue Wang, Zhao Xu, Weihua Luo, Kaifu Zhang
AI

Causal Agent based on Large Language Model

arXiv:2408.06849v3 Announce Type: replace-cross Abstract: The large language model (LLM) has achieved significant success across various domains. However, the inherent complexity of causal problems and causal theory…

Source: arXiv cs.CL Kairong Han, Kun Kuang, Ziyu Zhao, Junjian Ye, Fei Wu
AI

Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse

arXiv:2608.10810v2 Announce Type: replace Abstract: Emotion understanding in discourse requires reasoning beyond surface sentiment because speakers often convey affect through indirect, implicit, polite, ironic, or…

Source: arXiv cs.CL Zhenyan Zheng, Yunyao Zhang, Junxi Sheng, Junqing Yu, Zikai Song
AI

Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed

arXiv:2608.11981v1 Announce Type: new Abstract: Small Language Models (SLMs) have emerged as a more efficient alternative to traditional Large Language Models (LLMs), offering promising potential in resource-constrained…

Source: arXiv cs.CL Haokun Lin, Kaijie Zhu, Haobo Xu, Yichen Wu, Zhichao Lu, Qingfu Zhang, Zhenan Sun
AI

Stigma and Support in Online Sexual Violence Narratives on Reddit

arXiv:2608.11433v1 Announce Type: new Abstract: Online communities increasingly provide spaces where survivors of sexual violence can share their experiences and seek support. Although prior research has examined stigma…

Source: arXiv cs.CL Shirlene Rose Bandela, Karan Bindal, Vaibhav Garg, Rezvaneh Rezapour
AI

RevCRN: Reversible Analog Computation using Chemical Reaction Networks

arXiv:2608.11362v1 Announce Type: cross Abstract: The computability of real numbers and functions using Turing Machines has been a central area of theoretical computer science since the mid-20th century. In the late…

Source: arXiv cs.CL Saptarshi Biswas, James I. Lathrop, Rana D. Parshad
AI

Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation

arXiv:2608.12125v1 Announce Type: cross Abstract: As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of…

Source: arXiv cs.CL Akash Kundu, Emanuel Tewolde, Ratip Emin Berker, Samuel F. Brown, Vincent Conitzer
AI

CAR: Query-Guided Confidence-Aware Reranking for Retrieval-Augmented Generation

arXiv:2605.04495v2 Announce Type: replace Abstract: Retrieval-augmented generation (RAG) relies on evidence ranking to determine what information is exposed to the generator, yet existing retrieval and reranking methods…

Source: arXiv cs.CL Zhipeng Song, Yizhi Zhou, Xiangyu Kong, Jiulong Jiao, Xuezhou Ye, Chunqi Gao, Xueqing Shi, Yu Wang, Yuhang Zhou, Heng Qi
AI

Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

arXiv:2608.12149v1 Announce Type: new Abstract: We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two architecture-aligned morphologies: MAs consistently spike…

Source: arXiv cs.CL Zunhai Su, Bohan Sun, Xialie Zhuang, Shuibai Zhang, He Xiao, Jing Xiong, Hengyuan Zhang, Zhongzhu Zhou, Tiantian Zhang, Ngai Wong, Chuan-Wei Kuo
AI

Structuring the Space of Perspectives

arXiv:2608.12113v1 Announce Type: new Abstract: The same event can be reported from different perspectives depending on the experiences, background, and beliefs of the writer or speaker. A variety of NLP areas engage…

Source: arXiv cs.CL Agnese Daffara, Sebastian Pad\'o, Tanise Ceron
AI

Claim-Level Reliability Assessment for Efficient Test-Time Reasoning

arXiv:2608.11994v1 Announce Type: cross Abstract: We propose claim-level falsification as a principle for test-time scaling and instantiate it through Claim-Level Reliability Assessment (CLR), a training-free framework…

Source: arXiv cs.CL Sen Xu, Wei Wang, Shixi Liu, Jixin Min, Yingwei Dai, Zhibin Yin, Yirong Chen, Junlin Zhang
AI

Self-Evolving Embodied Agents via Skill-Harness Evolution

arXiv:2608.11350v1 Announce Type: new Abstract: Embodied agents are increasingly built as systems around foundation models, where performance depends not only on model weights but also on the skills, context, action…

Source: arXiv cs.CL Peidong Wang, Zhiming Ma, Ying Chang, Xufang Luo, Xiaocui Yang, Shi Feng, Yuqing Yang, Dongsheng Li
AI

Explainability in Practice: A Survey of Explainable NLP Across Various Domains

arXiv:2502.00837v3 Announce Type: replace Abstract: Natural Language Processing (NLP) is now embedded in critical sectors including healthcare, finance, and customer relationship management, where models such as GPT-4o,…

Source: arXiv cs.CL Hadi Mohammadi, Robert A. Bagheri, Anastasia Giachanou, Daniel L. Oberski
AI

On Weak Bisimilarities in CCSK

arXiv:2608.11531v1 Announce Type: new Abstract: In the context of CCSK, a reversible extension of CCS, we study different notions of bisimilarity (strong/weak, forward-only/reversible) and highlight their differences…

Source: arXiv cs.CL Baptiste Vall\'ee, Ivan Lanese
AI

DORA Explorer: Improving the Exploration Ability of LLMs Without Training

arXiv:2604.17244v2 Announce Type: replace Abstract: Large language model (LLM) agents for sequential decision-making struggle to produce diverse outputs. This leads to insufficient exploration, suboptimal solutions, and…

Source: arXiv cs.CL Priya Gurjar, Md Farhan Ishmam, Kenneth Marino
AI

LookBack: Where and How to Score LVLM Responses via Visual Reference Usage

arXiv:2608.11847v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) integrate visual perception with language generation, enabling responses that span image understanding and complex reasoning.…

Source: arXiv cs.CL Beomsik Cho, Jinhyeong Kim, Dongseok Lee, Jaehyung Kim
AI

BLADE: Better Language Answers through Dialogue and Explanations

arXiv:2604.03236v2 Announce Type: replace-cross Abstract: Large language model (LLM)-based educational assistants often provide direct answers offering little incentive for students to explore or engage with course…

Source: arXiv cs.CL Chathuri Jayaweera, Phoebe Huang, Bonnie J. Dorr
AI

Foresight Without Seeing: Latent Futures for World Action Models

arXiv:2608.11605v1 Announce Type: new Abstract: World Action Models (WAMs) couple future visual prediction with robot action generation, enabling policies to model how the physical world evolves during interaction.…

Source: arXiv cs.AI Jiakai Huang, Zhongbo Wu, Zheng Zhang, Zihan Wang, Shan You, Tao Huang
AI

Designing Agentic AI-Based Screening for Portfolio Investment

arXiv:2603.23300v2 Announce Type: replace-cross Abstract: We introduce a new agentic artificial intelligence (AI) platform for portfolio management. Our architecture consists of three layers. First, two large language…

Source: arXiv cs.AI Mehmet Caner, Agostino Capponi, Nathan Sun, Jonathan Y. Tan
AI

Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL

arXiv:2608.12063v1 Announce Type: cross Abstract: Integrating locomotion and manipulation is essential for robot autonomy, but scaling standard Reinforcement Learning (RL) to complex tasks is severely bottlenecked by…

Source: arXiv cs.AI Martin Schuck, Maks Sorokin, Simone Manni, Duy Ta, Angela P. Schoellig, Marco Hutter, Simon Le Cleac'H, Jan Br\"udigam
AI

OpenAg: Democratizing Agricultural Intelligence

arXiv:2506.04571v3 Announce Type: replace Abstract: Agriculture is undergoing a major transformation driven by artificial intelligence (AI), machine learning, and knowledge representation technologies. However, current…

Source: arXiv cs.AI Srikanth Thudumu, Jason Fisher
AI

A 12-CNOT Double Qubit Excitation Gate

arXiv:2608.11733v1 Announce Type: cross Abstract: Effective implementation of high-level quantum gates is essential for practical quantum computing. To the best of our knowledge, we present the first reported 12-CNOT…

Source: arXiv cs.AI Irfansha Shaik
AI

Making AI-Generated Feedback Matter: From Provision to Student Enactment

arXiv:2608.11625v1 Announce Type: new Abstract: Feedback processes strongly influence student learning, yet their educational value depends on addressing two distinct challenges: providing high-quality, timely, and…

Source: arXiv cs.AI Omar Alsaiari, Nilufar Baghaei, Jason M. Lodge, Dragan Ga\v{s}evi'c, Naomi Winstone, Hassan Khosravi
AI

Machine Learning-Based Cyber Defense for Cloud Infrastructure: An Adaptive Deep Q-Network Architecture for Intelligent Intrusion Detection and Automated Threat Mitigation

arXiv:2608.12190v1 Announce Type: cross Abstract: With the increasing complexity of cyber assaults in cloud environments, adaptable security solutions are needed that can support real-time detection and autonomous…

Source: arXiv cs.AI Md Yassir Mottalib, Md Yousuf, Eklachur Rahman Bhuiyan, S M Ahsan Habib, Sonjoy Kumar Dey, Md. Salahuddin Gazi, Molay Kumar Roy, Asaduzzaman Anik
AI

A Neighborhood Attention Transformer Network for Enhanced 3D Segmentation of the Left Anterior Descending Artery

arXiv:2608.12274v1 Announce Type: cross Abstract: Background: Accurate segmentation of the Left Anterior Descending (LAD) artery in 3D free-breathing, non-contrast CT is critical for cardiac dose sparing in thoracic…

Source: arXiv cs.AI Rafi Ibn Sultan, Chengyin Li, Yiannos Demetriou, Ahmed I. Ghanem, Joshua P. Kim, Justine Cunningham, Hassan Bagher-Ebadian, Dongxiao Zhu, Kundan S. Thind
AI

SportD: How do VLMs physically strategize?

arXiv:2607.14616v3 Announce Type: replace Abstract: Vision-language models (VLMs) can describe a scene, but can they act well within one? We study whether VLMs can make sound strategic decisions, using soccer as an…

Source: arXiv cs.AI Jasin Cekinmez, Addison J. Wu, Haotian Xia, Kyumin Andrew Shim, Anay Putty, Jinglin Xiao, Zhuohan Liu, Leo Liu, Weining Shen
AI

AI Guardrail Survival under Single-Cycle Agentic Self-Summarization

arXiv:2608.11392v1 Announce Type: cross Abstract: Long-running agents periodically compact their context, replacing the transcript with a model-generated summary.Recent work shows that dropping a standing safety…

Source: arXiv cs.AI Ted Kwartler, Alan Aqrawi, Arian Abbasi
AI

Methodologies for Improving the Quality of AI Tutoring in K-12 Education

arXiv:2608.11259v1 Announce Type: cross Abstract: Many AI tutors leverage large language models (LLMs) today. Given that LLMs are opaque black boxes, robust evaluation and live experimentation to measure the impact of…

Source: arXiv cs.AI Tushar Udeshi, Anna Khazenzon, Kabir Khan, Nick Breen, RJ Corwin, Chris DiGiano, Kodi Weatherholtz, Marek Zaluski
AI

Evaluating LLM Generated Detection Rules in Cybersecurity

arXiv:2509.16749v1 Announce Type: cross Abstract: LLMs are increasingly pervasive in the security environment, with limited measures of their effectiveness, which limits trust and usefulness to security practitioners.…

Source: arXiv cs.AI Anna Bertiger, Bobby Filar, Aryan Luthra, Stefano Meschiari, Aiden Mitchell, Sam Scholten, Vivek Sharath
AI

On Benchmarking Human-Like Intelligence in Machines

arXiv:2502.20502v2 Announce Type: replace Abstract: Recent advances in Artificial Intelligence (AI) have yielded powerful computational models that, by learning from vast amounts of human-generated data, are…

Source: arXiv cs.AI Lance Ying, Katherine M. Collins, Lionel Wong, Ilia Sucholutsky, Ryan Liu, Adrian Weller, Tianmin Shu, Thomas L. Griffiths, Joshua B. Tenenbaum
AI

HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting

arXiv:2608.11692v1 Announce Type: new Abstract: Autonomous logistics sorting systems (ALSS) are an important industrial application of embodied AI, which requires joint planning over spatially disjoint camera views. We…

Source: arXiv cs.AI Xikai Sun, Cangtian Zhou, Kebin Liu, Ke Ma, Xu Wang, Zaishu Chen, Haotian Wang, Li Liu, Yunhao Liu
AI

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

arXiv:2607.11175v2 Announce Type: replace Abstract: The growing ability of large language models and vision-language models to jointly interpret and reason over images and text is reshaping medical imaging AI, moving it…

Source: arXiv cs.AI Chunzheng Zhu, Lei Tian, Bohan Tan, Ziqi Zhou, Yuxuan Sun, Yijun Wang, Chengchao Lv, Yilin Wen, Yijun He, Jinghao Lin, Yihang Chen, Chee Wei Tan, Qianshan Wei, Lei Zhao, Bin Pu, Kenli Li, Yuan Xue, J…
AI

A Conceptual Framework for Refining Influence Knowledge from Simulation Evidence in Cyber-Physical Systems

arXiv:2608.11221v1 Announce Type: new Abstract: Cyber-physical systems (CPS) are typically developed by multiple stakeholders who produce artefacts tailored to their specific domains of expertise. The behaviour of these…

Source: arXiv cs.AI Barbara da Silva Oliveira (UniCA, Laboratoire I3S - COMRED, KAIROS), Julien Deantoni (UniCA, Laboratoire I3S - COMRED, KAIROS), Nicolas Ferry (Laboratoire I3S - COMRED, KAIROS, UniCA)
AI

Gaze Target Estimation Anywhere with Concepts

arXiv:2608.11367v1 Announce Type: cross Abstract: Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, multi-stage pipelines that…

Source: arXiv cs.AI Xu Cao, Houze Yang, Vipin Gunda, Zhongyi Zhou, Tianyu Xu, Adarsh Kowdle, Inki Kim, James M. Rehg
AI

CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations

arXiv:2608.12002v1 Announce Type: new Abstract: Agents are increasingly considered for automating network operations and maintenance, where engineers must diagnose network faults, optimize configurations to enhance…

Source: arXiv cs.AI Xingyu Yan, Tingting Dai, Antonio De Domenico, Mohamed Sana, Nicola Piovesan, Changchang Li, Bowen Liu, Kun Jiang, Mengjie Zhang, Dingcheng Shan, Jing-Cheng Pang, Chenwei Wu, Sijie Wu, Lianying Chao,…
AI

On the Definition of Intelligence

arXiv:2507.22423v3 Announce Type: replace Abstract: To engineer AGI, we should first capture the essence of intelligence in a species-agnostic form that can be evaluated, while being sufficiently general to encompass…

Source: arXiv cs.AI Kei-Sing Ng
AI

LLMs in Process Diagram Engineering: From Optimal PFDs to Validated P&IDs

arXiv:2608.11220v1 Announce Type: new Abstract: Nowadays, the creation of a process flow diagram (PFD) and its subsequent transformation into a piping and instrumentation diagram (P&ID) is predominantly performed…

Source: arXiv cs.AI Timur Zakarin, Sergei Voitov, Sergei Shumilin, Evgeny Burnaev
AI

Instruction Alignment for Binary Code Representation Learning

arXiv:2608.11766v1 Announce Type: cross Abstract: Binary code representation learning is a fundamental problem in software security and reverse engineering. Existing methods mainly learn function-level embeddings that…

Source: arXiv cs.AI Huaijin Wang, Shuai Wang
AI

Fingerprinting Text-to-Image Diffusion Models via Collapsed Generation

arXiv:2608.11732v1 Announce Type: cross Abstract: Proprietary text-to-image diffusion models are increasingly distributed as hosted services and downloadable checkpoints, making their intellectual property (IP)…

Source: arXiv cs.AI Yuanmin Huang, Chen Chen, Geng Hong, Xiaoyu You, Hui Xue, Zhenxing Qian, Mi Zhang, Min Yang
AI

Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams

arXiv:2608.12262v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. For example, OpenAI Prism is a free workspace for…

Source: arXiv cs.AI Weihao Bo, Shan Zhang, Yanpeng Sun, Jie Liu, Yongke Yao, Jinhao Du, Wei He, Kai Zou, Zechao Li, Jingdong Wang
AI

Making Gaussian Kolmogorov-Arnold Networks Reliable and Accurate

arXiv:2604.21174v3 Announce Type: replace-cross Abstract: Kolmogorov-Arnold Networks (KANs) replace fixed activations with learnable univariate edge functions whose behavior depends strongly on the chosen basis.…

Source: arXiv cs.AI Amir Noorizadegan, Sifan Wang, Leevan Ling
AI

Herding End-to-End Autonomous Driving via Neuro-Symbolic Safety Guards

arXiv:2608.11451v1 Announce Type: cross Abstract: Modern end-to-end driving agents can achieve high average performance yet still violate basic traffic rules that a human driver would never miss. The reason is…

Source: arXiv cs.AI Sim\'on Pati\~no Idarraga, Erick Silva, Rehana Yasmin, Ali Shoker
AI

ArtiFact: A Large-Scale Multi-Modal Cultural Heritage Dataset

arXiv:2606.09648v2 Announce Type: replace-cross Abstract: Multi-modal data management has emerged as a central research topic in the database community, spanning data integration, semantic query processing, and data…

Source: arXiv cs.AI Luciano Duarte, Olga Ovcharenko, Sebastian Schelter
AI

Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence

arXiv:2608.12290v1 Announce Type: cross Abstract: Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and reliability presents…

Source: arXiv cs.AI Aman Tyagi, Hemanth Boinpally, Jonathan Chen, Douglas Gebert, Steven Hickson
AI

Cavity-Enhanced Collective Quantum Processing with Polarization-Encoded Qubits

arXiv:2605.10473v2 Announce Type: replace-cross Abstract: We introduce a cavity-enhanced optical architecture for collective quantum processing in which logical qubits are encoded in the polarization subspace of…

Source: arXiv cs.AI Kamil Wereszczy\'nski, J\'ozef Cyran, Adam Brzezowski, Dawid Za{\l}u\.zny, Robert Potoniec, Kasper Wi\'sniowski, Agnieszka Michalczuk
AI

Self-evolving network verifiers

arXiv:2608.11340v1 Announce Type: cross Abstract: Symbolic network verifiers can reason about correctness across vast spaces of routing inputs and failures, but only for the protocols and features an expert has encoded…

Source: arXiv cs.AI Ioannis Protogeros, Tibor Schneider, Laurent Vanbever
AI

CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications

arXiv:2608.11588v1 Announce Type: new Abstract: Mobile GUI agents remain brittle when deployed to applications absent from source training. We study novel-app generalization under a limited target interaction budget and…

Source: arXiv cs.AI Linqiang Guo (Peter), Li Gu (Peter), Zihuan Jiang (Peter), Zhixiang Chi (Peter), Siobhan Reid (Peter), Ziqiang Wang (Peter), Yuanhao Yu (Peter), Wei Liu (Peter), Yang Wang (Peter), Tse-Hsun (Peter),…
AI

InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

arXiv:2608.11234v1 Announce Type: new Abstract: Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity. Recent advances in AI agents create a timely…

Source: arXiv cs.AI Yuan Gao (Wanxiang), Zeren Yang (Wanxiang), Junnan Li (Wanxiang), Shawn (Wanxiang), Zhong, Ahmed Dajani, Mai Zheng, Andrea Arpaci-Dusseau, Remzi Arpaci-Dusseau
AI

Governing Agentic AI in FinTech

arXiv:2608.11344v1 Announce Type: cross Abstract: Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little oversight. Yet…

Source: arXiv cs.AI Henry Han
AI

G0.5: One Autoregressive Stream for Robot Reasoning and Action

arXiv:2608.11739v1 Announce Type: cross Abstract: The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action expert. This makes the VLM a…

Source: arXiv cs.AI Yicheng Liu, Zibin Dong, Baijun Ye, Tianyuan Yuan, Tao Jiang, Anqi Yang, Shicheng Cao, Haonan Liu, Yue Sun, Zihan Guo, Xiao Liu, Dong Ke, Changxun Pan, Chenru Wu, Tailai Cheng, Xiaoshu Ren, Xinlei Zh…
AI

Inferential Capability Does Not Determine Legal Scope

arXiv:2608.10601v2 Announce Type: replace-cross Abstract: Two instruments of EU digital law place inference at their centre and mean different things by it. Article 3(1) of the AI Act uses the capability to infer…

Source: arXiv cs.AI Nicola Fabiano
AI

VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

arXiv:2608.12282v1 Announce Type: new Abstract: Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation. We…

Source: arXiv cs.AI Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder, Siyu Huo, Raavi Gupta, Abhinav Jain, Praveen Venkateswaran, Abdulhamid Adebayo, Danish Contractor
AI

ComBodied Agents: a New Paradigm of Human-Centric Agentic AI

arXiv:2608.10915v2 Announce Type: replace Abstract: After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether…

Source: arXiv cs.AI Qianggang Ding, Xingyao Wang, Rui Feng, Zhibin Wang, Feixiang Yao, Kelong Mao, Hao Sun, Zhiyao Luo, Jiankai Tang, Lei Li, Jiadong Guo, Minheng Ni, Weicong Lin, Chenxi Yang, Hongxiang Gao, Zhenghua Ch…
AI

CORE-3D: Context-aware Open-vocabulary Retrieval by Embeddings in 3D

arXiv:2509.24528v4 Announce Type: replace-cross Abstract: Object retrieval from a scene has become a new trend of research due to its numerous applications. Recent approaches achieve zero-shot, open-vocabulary 3D…

Source: arXiv cs.AI Mohamad Amin Mirzaei, Pantea Amoie, Ali Ekhterachian, Matin Mirzababaei, Babak Khalaj
AI

Proportional Committee Elections with Positive and Negative Votes

arXiv:2503.01985v2 Announce Type: replace-cross Abstract: In the classic committee election setting each voter approves a subset of candidates and the goal is to select $k$ winners based on these preferences. A central…

Source: arXiv cs.AI Sonja Kraiczy, Georgios Papasotiropoulos, Grzegorz Pierczy\'nski, Piotr Skowron
AI

memorywire: A Vendor-Neutral Wire Format for Agent Memory Operations

arXiv:2606.01138v4 Announce Type: replace-cross Abstract: Agent-memory frameworks -- mem0, Letta/MemGPT, Cognee, Zep/Graphiti, MemoryOS, MemTensor -- each ship their own SDK, storage layout, and operational vocabulary.…

Source: arXiv cs.AI Thamilvendhan Munirathinam
AI

Agent Safety Should Be a Runtime Contract

arXiv:2608.11274v1 Announce Type: cross Abstract: The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitutional AI. We argue this is structurally…

Source: arXiv cs.AI Albus W. Ng, Yi Han, Jusheng Zhang, Wenhao Wang
AI

LiDAR-based 3D Change Detection at City Scale

arXiv:2510.21112v3 Announce Type: replace-cross Abstract: High-definition 3D city maps enable city planning and change detection, which is essential for municipal compliance, map maintenance, and asset monitoring,…

Source: arXiv cs.AI Hezam Albaqami, Haitian Wang, Xinyu Wang, Muhammad Ibrahim, Zainy M. Malakan, Abdullah M. Algamdi, Mohammed H. Alghamdi, Ajmal Mian
AI

REVERE: Reflective Evolving Research Engineer

arXiv:2603.20667v2 Announce Type: replace-cross Abstract: Existing prompt-optimization techniques rely on local signals, causing poor generalization across tasks. In addition, they also rely on weak update mechanisms,…

Source: arXiv cs.AI Balaji Dinesh Gangireddi, Aniketh Garikaparthi, Manasi Patwardhan, Arman Cohan
AI

Keep the Future, Drop the Rollout: RIFT for World Action Models

arXiv:2608.11521v1 Announce Type: cross Abstract: World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation…

Source: arXiv cs.AI Chushan Zhang, Jinguang Tong, Xuesong Li, Yikai Wang, Hongdong Li
AI

Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations

arXiv:2608.12299v1 Announce Type: cross Abstract: Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial intelligence. Its purpose is intuitive: it converts…

Source: arXiv cs.AI AmirHossein Eshghi, Hamid Saadatfar, Seyyed Ali Hoseini, AmirMohsen Eshghi, Siavash Arjomand Bigdel
AI

Towards Human Motion World Models via Executable Behaviour Representations

arXiv:2604.18064v2 Announce Type: replace Abstract: Human motion world models should capture motion's intentionality by being executable: adaptable to different actions and capable of assessing motion quality. To…

Source: arXiv cs.AI Rimvydas Rubavicius, Manisha Dubey, N. Siddharth, Subramanian Ramamoorthy
AI

Credo: Declarative Control of LLM Pipelines via Beliefs and Policies

arXiv:2604.14401v2 Announce Type: replace Abstract: Agentic AI systems are becoming commonplace in domains that require long-lived, stateful decision-making in continuously evolving conditions. As such, correctness…

Source: arXiv cs.AI Duo Lu, Andrew Crotty, U\u{g}ur \c{C}etintemel
AI

XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication

arXiv:2608.11676v1 Announce Type: new Abstract: Heterogeneous multi-agent LLM systems, where agents are powered by different model families, can outperform homogeneous configurations by reducing redundant reasoning…

Source: arXiv cs.AI Wooseong Yang, Wei-Chieh Huang, Weizhi Zhang, Yu Wang, Philip S. Yu, Junhyun Lee
AI

Empowering Children to Create AI-Enabled Augmented Reality Experiences

arXiv:2508.08467v2 Announce Type: replace-cross Abstract: Despite their potential to enhance children's learning experiences, AI-enabled AR technologies are predominantly used in ways that position children as consumers…

Source: arXiv cs.AI Lei Zhang, Shuyao Zhou, Amna Liaqat, Tinney Mak, Brian Berengard, Emily Qian, Andr\'es Monroy-Hern\'andez
AI

Persistent Recursive Worlds Enable Autonomous Software Evolution

arXiv:2608.10450v2 Announce Type: replace-cross Abstract: Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems preserve continuity…

Source: arXiv cs.AI Beichen Huang, Zhenyu Liang, Bowen Zheng, Ran Cheng
AI

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence

arXiv:2608.11341v1 Announce Type: new Abstract: Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a distant ambition into a mission architecture of…

Source: arXiv cs.AI Brian Wang, Bin Feng, Xiaoman Pan, Chenyang An, Felix Liu, Tangqi Fang, Gongbo Sun, Lingfeng Shen, Ning Wang, Handuo Zhang, Feng Chen, Fuchao Yang, Xiang Wang, Jiacheng Lin, Siting Li, Zixuan Liu, Ch…
AI

VQ-bench: A Composable Vector Quantization Framework

arXiv:2608.11240v1 Announce Type: new Abstract: Vector quantization is an old problem but has recently become central to AI infrastructure. It is therefore experiencing a surge of renewed engineering and research…

Source: arXiv cs.AI Ashwin Padaki, Amir Ingber, Edo Liberty
AI

HCGRec: Hint-Conditioned Generative Recommendation with Semantic IDs

arXiv:2608.11980v1 Announce Type: cross Abstract: Semantic-ID generative recommenders represent each item as a short sequence of discrete semantic tokens and predict the next item by autoregressively generating this…

Source: arXiv cs.AI Kangning Zhang, Haotian Fang, Xukun Luo, Hao Yin, Yang Gao, Peng Yan, Weiwen Liu, Weinan Zhang, Yong Yu
AI

Learning from Online User Feedback for Shopping Agents

arXiv:2608.11604v1 Announce Type: new Abstract: Large language model-based shopping agents are increasingly deployed in real-world e-commerce platforms, generating massive amounts of user interaction logs that provide…

Source: arXiv cs.AI Haobo Zhang, Kelong Mao, Sulong Xu, Simiu Gu, Zhicheng Dou
AI

AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research

arXiv:2608.11216v1 Announce Type: new Abstract: World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across…

Source: arXiv cs.AI Marjan Moodi, Xuankang Zhu, Fernando De Mesentier Silva, Harold Chaput, Mohammad Reza Taesiri
AI

Small Data Explainer -- The impact of small data methods in everyday life

arXiv:2507.11773v2 Announce Type: replace-cross Abstract: The emergence of breakthrough artificial intelligence (AI) techniques has led to a renewed focus on how small data settings, i.e., settings with limited…

Source: arXiv cs.AI Maren Hackenberg, Sophia G. Connor, Fabian Kabus, June Brawner, Ella Markham, Mahi Hardalupas, Areeq Chowdhury, Rolf Backofen, Anna K\"ottgen, Angelika Rohde, Nadine Binder, Harald Binder, the Collab…
AI

Geometry-aware Incremental Neural Operator for Long-Horizon PDE prediction

arXiv:2608.11237v1 Announce Type: new Abstract: Neural operators have shown strong potential for learning solution operators of partial differential equations (PDEs). However, long-horizon autoregressive prediction…

Source: arXiv cs.AI Jiaquan Zhang, Shuxu Chen, Haifan Meng, Yi Lu, Zhihan Lyu, Fan Mo, Wei Dong, Yang Yang, Chaoning Zhang
AI

Tools as Continuous Flow for Evolving Agentic Reasoning

arXiv:2605.07339v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in orchestrating tools for reasoning tasks. However, existing methods rely on a step-wise…

Source: arXiv cs.AI Tairan Huang, Siyu Shang, Qiang Chen, Xiu Su, Yi Chen
AI

Teaching agentic AI to learn expert reasoning for rare disease diagnosis

arXiv:2606.16149v3 Announce Type: replace Abstract: Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer; off-the-shelf large language models (LLMs) rank the correct disease first…

Source: arXiv cs.AI Minh-Ha Nguyen, Erica Gray, Bryce A. Schuler, Kevin W. Byram, Chih-Ting Yang, Fan Ma, Hua Xu, Wu-Chen Su, Chao Yan, Wei-Qi Wei, Adam Wright, Lisa Bastarache, Josh Peterson, Lingyao Li, Siyuan Ma, Und…
AI

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

arXiv:2608.11727v1 Announce Type: new Abstract: When a coding agent obeys a rule, it may simply have been going to do that anyway. Existing instruction-following benchmarks cannot tell the difference: they concentrate…

Source: arXiv cs.AI Zining Huang, Haoran Que, Hong Zeng, Ge Zhang, Zuo Wang, Jin Chen, Haodong Wang, Zhongfei Hou, Changxin Pu, Shen Yan, Wenhao Huang
AI

No One to Blame: A Framework of Constitutive AI Unaccountability

arXiv:2608.12104v1 Announce Type: cross Abstract: The increasing deployment of autonomous, agentic AI systems challenges traditional accountability mechanisms. Existing research predominantly frames AI accountability…

Source: arXiv cs.AI Long Hoang Nguyen, Eva Sp\"athe, Sebastian Lins, Ali Sunyaev
AI

LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs

arXiv:2608.11231v1 Announce Type: new Abstract: LLM serving is increasingly accelerated by position-independent caching (PIC). Existing PIC methods, however, are built for full-attention models, where a token-indexed KV…

Source: arXiv cs.AI Yirui Liu, Ruoling Qi, Longwen Wang, Xuaner Wu, Jian Chen, Yuxin Jin, Jiawei Shao, Xuelong Li
AI

Backdoor Decontamination Dynamics in LLM Agents

arXiv:2608.11295v1 Announce Type: cross Abstract: Open-weight LLM agents are vulnerable to backdoors installed during fine-tuning, which may be undetectable if the trigger conditions are never met during testing.…

Source: arXiv cs.AI Gabriel Huang, Abhay Puri, L\'eo Boisvert, Alexandre Drouin, Perouz Taslakian, Spandana Gella, Christopher Pal
AI

Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability

arXiv:2608.11238v1 Announce Type: new Abstract: Retrieval-augmented generation improves the factuality of large language models by grounding responses in retrieved evidence, yet existing evaluation frameworks struggle…

Source: arXiv cs.AI Jeonghwan Choi, Taewon Yun, Minjeong Ban, Gyeonghun Sun, Jae-Gil Lee, Hwanjun Song
AI

How Organizations Use AI: Evidence from ChatGPT

arXiv:2608.12236v1 Announce Type: cross Abstract: We study how organizations use frontier generative AI by linking ChatGPT Enterprise account records to usage, worker roles, task classifications, and public-company…

Source: arXiv cs.AI Aaron Chatterji, David Holtz, Neel Rakholia, Prasanna Tambe, Gawesha Weeratunga