Skip to content
TILens What matters today in tech
Theme

Daily edition · AI

The daily ledger

TILens turns technical updates into a focused daily brief: official releases, trusted reporting, and practitioner analysis, deduplicated and organized by topic.

15 Aug 2026 edition
Python AI Django

CORS Chat

Tool: CORS Chat I built this today (with GPT-5.6-Sol xhigh) to help test Qwen 3.8 27B running in LM Studio on both my M5 MacBook Pro and an NVIDIA DGX Spark. It provides a web UI for exercising an…

Source: Simon Willison's Weblog
Python AI Django

Northern Gannet

Northern Gannet, in Pillar Point Harbor, CA, US This is Morris. Morris is a local celebrity: the only known Northern Gannet (Morus bassanus) in the entire Pacific Ocean. They showed up in the Farallon Islands off the…

Source: Simon Willison's Weblog
AI Research AI

Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence

arXiv:2608.12743v1 Announce Type: new Abstract: Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents,…

Source: arXiv cs.AI Haokai Zhang, Yuhang Ding, Yunshu Zhou, Xinze Du, Shengtao Zhang, Zhiyue Zhao, Yuling Xi, Hao Chen
AI Research AI

Dimensional Balance Improves Large Scale Spatiotemporal Prediction Performance

arXiv:2605.18793v3 Announce Type: replace-cross Abstract: Accurate spatiotemporal pattern analysis is critical in fields such as urban traffic, meteorology, and public health monitoring. However, existing methods face…

Source: arXiv cs.AI Jing Chen, Shixiang Pan, Yujie Fan, Haocheng Ye, Haitao Xu, Wenqiang Xu
AI Research AI

Labels Are Not Endpoints: Treatment Leakage and Construct Validity in MCP Agent Security Evaluation

arXiv:2608.12880v1 Announce Type: cross Abstract: Security evaluations of tool-using agents often equate stored labels with behavioral facts. We audit a preserved campaign by tracing 10,200 execution rows to 180…

Source: arXiv cs.AI Rana Muhammad Ahmed (Department of Computer Science, Bahria University, Islamabad, Pakistan), Sabahat Abbas (Department of Computer Science, Bahria University, Islamabad, Pakistan)
AI Research AI

A Compositional Theory of Curvature in Probabilistic Circuits

arXiv:2608.12869v1 Announce Type: cross Abstract: Probabilistic Circuits (PCs) are generative models that support exact inference and, unlike deep neural networks, admit an exact and tractable measure of loss-surface…

Source: arXiv cs.AI Hrithik Suresh, Sahil Sidheekh, Shelar Parth Vijay, Yasir Z, Sriraam Natarajan, Narayanan Chatapuram Krishnan
AI Research AI

AHD Agent: Agentic Reinforcement Learning for Automatic Heuristic Design

arXiv:2605.08756v2 Announce Type: replace Abstract: Automatic heuristic design (AHD) has emerged as a promising paradigm for solving NP-hard combinatorial optimization problems (COPs). Recent works show that large…

Source: arXiv cs.AI Haoze Lv, Ning Lu, Ziang Zhou, Yew-Soon Ong, Shengcai Liu
AI Research AI

Trie Automata for Constrained Decoding over Large Finite Sets

arXiv:2608.12574v1 Announce Type: new Abstract: Large language models increasingly need to generate structured outputs that conform to predefined schemas, with one common constraint being selection from a finite set of…

Source: arXiv cs.AI Xingzi Xu, Karim Bouyarmane
AI Research AI

Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs

arXiv:2608.12675v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is widely used to improve the performance of Large Language Models (LLMs) in answering user queries. Existing privacy research on RAG…

Source: arXiv cs.AI Saleh Almohaimeed, Saad Almohaimeed, Mousa Jari, Fahad Alotaibi, Khalid A. Alobaid
AI Research AI

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

arXiv:2608.13558v1 Announce Type: new Abstract: Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to…

Source: arXiv cs.AI Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu
AI Research AI

Into the ORBIT for Time Series: Training Regimes for Foundation Models

arXiv:2608.13262v1 Announce Type: cross Abstract: Time series foundation models (TSFMs) have advanced primarily through architectural innovation, while training regimes for large-scale heterogeneous corpora remain…

Source: arXiv cs.AI Hongjie Xia, Yiding Liu, Yifan Hu, Peiyuan Liu, Zewei Dong
AI Research AI

LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation

arXiv:2608.13136v1 Announce Type: cross Abstract: With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention. Existing approaches enable LLMs to retrieve…

Source: arXiv cs.AI Chenrun Wang, Mingxuan Zhu, Tiancheng Huang, Wenjie Li, Yujie Zhang, Zichen Zhu, Zhiying Zou, Kai Yu, Lu Chen
AI Research AI

MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents

arXiv:2608.12428v1 Announce Type: new Abstract: Memory is a core component of AI agents, enabling them to accumulate experience, maintain personalization, and adapt over long-term interactions. However, existing memory…

Source: arXiv cs.AI Kaichao Liang, Yuqi Cui, Hao Kong, Xinyuan Huang, Guohaotian Hou, Qingcan Kang, Liang Chen, Yiyang Yin, Ke Ye, Jiaquan Guo, Da Chen, Lingan Zeng, Yixing Peng, Rong Yao, Shixiong Kai, Mingxuan Yuan
AI Research AI

GEM: A Generative Embedding Model Bridging Reasoning and Retrieval

arXiv:2608.13200v1 Announce Type: cross Abstract: Modern LLMs excel at reasoning and instruction following, enabling users to express complex and diverse information needs. However, conventional retrievers largely rely…

Source: arXiv cs.AI Zhili Shen, Craig Macdonald
AI Research AI

Concept Drift Detection and Adaptive Retraining of Malware Classification Models

arXiv:2608.13465v1 Announce Type: cross Abstract: Concept drift refers to changes over time in the statistical properties of data, as compared to the data that was used to train a learning model. Machine learning models…

Source: arXiv cs.AI Christofer Washington Berruz Chungata, Martin Jurecek, Katerina Potika, William B. Andreopoulos, Mark Stamp
AI Research AI

HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

arXiv:2608.13555v1 Announce Type: cross Abstract: Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors…

Source: arXiv cs.AI Dairu Liu, Zekun Qi, Jiayu Zeng, Ruixi Yu, Yu Guan, Yintianrun Zhang, Xuchuan Chen, Sikai Liang, Zekai Li, Chenghuai Lin, Xinqiang Yu, Wenyao Zhang, He Wang, Li Yi
AI Research AI

TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems

arXiv:2608.13221v1 Announce Type: new Abstract: The evaluation of LLM reasoning is moving from final-answer accuracy to process-level assessment, yet existing methods still fail to capture how models plan reasoning…

Source: arXiv cs.AI Shunwen Bai, Ziping Ma, Chaoyang Zhang, Yarong Wang, Jiale Liu, Zhen Qin, Qingpei Guo
AI Research AI

Constitutional On-Policy Safe Distillation

arXiv:2606.03089v3 Announce Type: replace-cross Abstract: On-policy self-distillation (OPSD) has emerged as an efficient post-training paradigm by using a teacher conditioned on privileged information to provide dense…

Source: arXiv cs.AI Ming Wen, Yuxuan Liu, Kun Yang, Yunhao Feng, Zhuoer Xu, Yuhao Sun, Shiwen Cui, Xiang Zheng, Yi Liu, Xingjun Ma, Yu-Gang Jiang
AI Research AI

DiffImaginE: Imagine to Verify Entity Types with Diffusion

arXiv:2608.03025v3 Announce Type: replace Abstract: Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence.…

Source: arXiv cs.AI Feng Zhang, Feiyu Han, Rongxin Yang, Yang Liu, Yancheng Chen, Rui Wang, Yingguang Yang, Tian Xueyun, Chongyang Zhang, Hao Zheng, Xu Kefu, Congjing Ran, Fuhai Chen, Bin Chong
AI Research AI

DAPD: Dual-Anchored Policy Distillation

arXiv:2608.01735v2 Announce Type: replace Abstract: On-policy (self) distillation (OPSD) is increasingly adopted for language-model post-training. It strengthens the teacher with privileged information but can induce a…

Source: arXiv cs.AI Jianyu Wu, Yizhou Wang, Encheng Su, Chen Tang, Shixiang Tang
AI Research AI

The AI Accountability Ecosystem in the Era of Language Models

arXiv:2608.12320v1 Announce Type: cross Abstract: This article reviews and updates the framework for accountability in AI based on account- ability ecosystems. We update the framework in light of the latest developments…

Source: arXiv cs.AI Chris Percy, Artur d'Avila Garcez
AI Research AI

General Probabilities of Causation with Causal Knowledge

arXiv:2608.12657v1 Announce Type: new Abstract: Probabilities of causation (PoCs) characterize individual causal responses that cannot be directly observed and therefore generally require partial identification. Tian…

Source: arXiv cs.AI Xin Shu, Zhen Lei, Ang Li
AI Research AI

Sign Language Video Synthesis via Loss-Guided Multi-Expert GANs

arXiv:2608.13368v1 Announce Type: cross Abstract: This preliminary technical report presents a framework for sign language video synthesis using a loss-guided multi-expert Generative Adversarial Network (GAN) to enhance…

Source: arXiv cs.AI Dingzhan Nong, Zhihao Ren, Ziqi Li, Tim Lo
AI Research AI

DiffGRM: Diffusion-based Generative Recommendation Model

arXiv:2510.21805v2 Announce Type: replace-cross Abstract: Generative recommendation (GR) is an emerging paradigm that represents each item via a tokenizer as an n-digit semantic ID (SID) and predicts the next item by…

Source: arXiv cs.AI Zhao Liu, Yichen Zhu, Yiqing Yang, Xiao Lv, Guoping Tang, Rui Huang, Qiang Luo, Ruiming Tang, Kun Gai, Guorui Zhou
AI Research AI

Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples

arXiv:2608.13341v1 Announce Type: cross Abstract: Infrared (IR) spectroscopy is widely used for chemical sensing, but extracting reliable chemical information from spectra remains challenging. Conventional…

Source: arXiv cs.AI Yusen Tan, Yixuan Chen, Zheng Fang, Pan Liu, Yifan Li, Qinyu Guo, Zhedong Lin, Yuqiang Li, Xiangxiang Zeng, Tong Wang, Jun Xia
AI Research AI

FSGR: Mitigating Token Frequency Bias for Fair SID-Based Generative Recommendation

arXiv:2608.12845v1 Announce Type: cross Abstract: Semantic ID (SID)-based generative recommendation has recently achieved remarkable success. However, existing methods suffer from a previously overlooked fairness issue,…

Source: arXiv cs.AI Yuchen Zheng, Sihan Xu, Jingwen Yang, Xiangrui Cai, Haiwei Zhang, Xiaojie Yuan
AI Research AI

CEON: Circular Economy Ontology Network

arXiv:2606.02253v2 Announce Type: replace Abstract: Increasing the circularity of resource use in our society has been recognized as a path to sustainability, i.e., transitioning into a more circular economy. There are…

Source: arXiv cs.AI Huanyu Li, Els de Vleeschauwer, Robin Keskis\"arkk\"a, Mikael Lindecrantz, Mina Abd Nikooie Pour, Ying Li, Ben De Meester, Patrick Lambrix, Eva Blomqvist
AI Research AI

VALG: An Agentic System for ML Theory Research

arXiv:2608.13060v1 Announce Type: new Abstract: Machine learning theory studies learning procedures through mathematical setups in which the data model, training protocol, oracle access, loss, metric, and randomness…

Source: arXiv cs.AI Dechen Zhang, Xuan Tang, Xinxiang Yin, Xingwu Chen, Jian Qian, Difan Zou
AI Research AI

Uniform Herding: Exemplar Replay with Representation Refresh

arXiv:2608.13061v1 Announce Type: new Abstract: As the feature representation changes, replay must preserve the earlier classes. However, only a bounded active exemplar set can be replayed. We propose Uniform Herding,…

Source: arXiv cs.AI Krishna Subedi
AI Research AI

SDS-LoRA: Overcoming Anisotropic Gradient Scaling in Low-Rank Adaptation

arXiv:2606.16454v2 Announce Type: replace-cross Abstract: Low-Rank Adaptation (LoRA) enables efficient adaptation of large pretrained models to downstream tasks by parameterizing weight updates with low-rank matrices.…

Source: arXiv cs.AI Junghun Oh, Sungyong Baik, Kyoung Mu Lee
AI Research AI

What Makes a Peer? Valuation-Anchored Similarity in Private Markets

arXiv:2608.12594v1 Announce Type: cross Abstract: As more investors contemplate private markets and contend with limited transparency, sparse disclosures, and infrequent transactions, identifying economically meaningful…

Source: arXiv cs.AI Sebastian Frank, Jingrao Lyu, Max Jarmey, Preetha Saha, Mingshu Li, Sweet Kaur, Sola Akinola, Dhagash Mehta
AI Research AI

Towards Context-Aware Clinical Motion Understanding in Daily Living at Home: Freezing of Gait Detection with Egocentric Vision

arXiv:2608.13283v1 Announce Type: new Abstract: Understanding motion in daily living requires context beyond kinematics, because similar inertial patterns during activities of daily living (ADLs) can reflect intentional…

Source: arXiv cs.AI Vayalet Stefanova, Diwas Lamsal, Margot Genbrugge, Maxim Yudayev, Christian Schlenstedt, Moran Gilat, Bart Vanrumste, Benjamin Filtjens
AI Research AI

Training AI Scientists to Replicate Research

arXiv:2608.13331v1 Announce Type: cross Abstract: The replicability of papers is a cornerstone of scientific knowledge, ensuring the reliability of existing results and providing a base for further experiments. The act…

Source: arXiv cs.AI Damon Falck, Samer Sabri, Anja Surina, Thom Foster, Anya Sims, Sam Devlin, Dylan Rogers, Tantum Collins, Kaloyan Aleksiev, Louis Kirsch, Edward Hughes
AI Research AI

AI and Consumer Rights in India Working Paper

arXiv:2608.12863v1 Announce Type: new Abstract: As AI systems proliferate in consumer facing applications, questions about liability for AI related harms remain unresolved. This working paper examines whether India's…

Source: arXiv cs.AI Omir Kumar, Sriya Sridhar, Vibhav Mithal, Balaraman Ravindran
AI Research AI

Rules or Character? Scaling Laws for AI Safety Design

arXiv:2608.13345v1 Announce Type: new Abstract: Artificial Intelligence (AI) safety systems combine character shaping (e.g., Reinforcement Learning from Human Feedback [RLHF], Constitutional AI), which modifies…

Source: arXiv cs.AI Satoshi Takahashi, Nobuji Kouno, Masaaki Komatsu, Ryuji Hamamoto
AI Research AI

Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories

arXiv:2608.12847v1 Announce Type: new Abstract: Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or…

Source: arXiv cs.AI Yifei Li, Heng Wang, Lingling Zhang, Muye Huang, Xinyu Zhang, Jiashuai Liu, Hang Yan, Rongman Xu
AI Research AI

Vero: Can AI Agents Build Formally Verified Software Repositories?

arXiv:2608.13522v1 Announce Type: cross Abstract: AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated code. Verified code generation, in which an agent…

Source: arXiv cs.AI Zhe Ye, Hantao Lou, Yuechun Sun, Peiyang Song, Zhengxu Yan, Timothe Kasriel, Qingyang Zhang, Kaiyu Yang, Soonho Kong, Jingxuan He, Dawn Song
AI Research AI

TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability

arXiv:2608.09538v2 Announce Type: replace-cross Abstract: We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation. TCS-Bench…

Source: arXiv cs.AI Vincent Cohen-Addad, Dimitris Paparas, Ernest van Wijland, Max Springer, Julien Canitrot-Paradis, Honghao Lin, David Woodruff, Adarsh Kumarappan, Rajesh Jayaram, Rudrajit Das, Lalit Jain, Ola Svensso…
AI Research AI

Learning to Adapt Cross-Domain Preferences via Meta-LoRA for LLM Personalization

arXiv:2608.12389v1 Announce Type: new Abstract: Cross-domain zero- or few-shot personalization aims to generate user-preferred responses in unseen conversational domains from only a handful of target-domain…

Source: arXiv cs.AI Xuefei Wang, Jun Han, Zixuan Wang, Qingkai Zeng, Xiao Wang, Ruijie Wang, Jianxin Li
AI Research AI

Keep the Future, Drop the Rollout: RIFT for World Action Models

arXiv:2608.11521v2 Announce Type: replace-cross Abstract: World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action…

Source: arXiv cs.AI Chushan Zhang, Jinguang Tong, Xuesong Li, Yikai Wang, Hongdong Li
AI Research AI

Annealed Softmax Greedy in Many-Armed Bayesian Bandits

arXiv:2605.31034v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) and group-based policy optimization methods such as GRPO update a stochastic policy by sampling multiple…

Source: arXiv cs.AI William Overman, Mohsen Bayati
AI Research AI

MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination

arXiv:2608.13476v1 Announce Type: new Abstract: We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for…

Source: arXiv cs.AI Saisha Shetty, Satvik Tripathi, Austin Lin, Colin Zhao, Theodore Kim, Don Enwerem, Jacinta Arnold, Shahriar Faghani, Tessa S Cook
AI Research AI

Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval

arXiv:2608.12843v1 Announce Type: cross Abstract: Text-based person anomaly retrieval aims to retrieve pedestrians exhibiting anomalous behaviors from a large image gallery using natural language descriptions. Compared…

Source: arXiv cs.AI Huu-An Vu, Cam Tu Tran Thi, Thanh Toan Le Ngo, Hoang Vo, Do Trung Hieu, Hieu Dinh Trung Pham, Khang Minh Le, Huy Minh Nhat Nguyen
AI Research AI

Measuring Curriculum-Labor Market Alignment at the Scale of a Program Portfolio

arXiv:2608.12356v1 Announce Type: cross Abstract: A college offering several overlapping computing degrees implicitly assumes that its programs are differentiated in line with how the labor market segments computing…

Source: arXiv cs.AI Sherzod Turaev, Saja Aldabet, Mary John, Namya Musthafa, Mamoun Awad, Nazar Zaki, Khaled Shuaib
AI Research AI

Thought-Aware KV Cache Compaction for Reasoning via Adaptive Attention Matching

arXiv:2608.12331v1 Announce Type: cross Abstract: Reasoning language models generate lengthy chain-of-thought (CoT) sequences whose key-value (KV) cache grows linearly and becomes a memory bottleneck during decoding.…

Source: arXiv cs.AI Yang Liu, Bin Chong, Chongyang Zhang, Hao Zheng, Jiayu Liang, Xu Kefu
AI Research AI

Time-Series Forecasting in Safety-Critical Environments: An Open-Source Package for EU-AI-Act-Compliant Development / Zeitreihenprognose in sicherheitskritischen Umgebungen: Ein Open-Source-Paket f\"ur die KI-VO-konforme Entwicklung

arXiv:2604.23859v3 Announce Type: replace Abstract: With spotforecast2-safe we present an integrated Compliance-by-Design approach to Python-based point forecasting of time series in safety-critical environments. A…

Source: arXiv cs.AI Thomas Bartz-Beielstein, Eva Bartz
AI Research AI

DMDIntel: Interpreting Large Language Models via Dynamic Mode Decomposition

arXiv:2608.13048v1 Announce Type: new Abstract: In this work, we introduce DMDIntel which uses dynamic mode decomposition (DMD) to make the predictions made by LLMs in a classification task interpretable. It develops an…

Source: arXiv cs.AI Amogh Joshi, Animesh Mukherjee, Sergey Utyuzhnikov
AI Research AI

Vision-Language Models are Fragile Multilingual Associators

arXiv:2608.12333v1 Announce Type: cross Abstract: Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the…

Source: arXiv cs.AI Ritabrata Chakraborty, Rajatsubhra Chakraborty, Shivakumara Palaiahnakote, Angelo Cangelosi, Umapada Pal
AI Research AI

Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware Reformulation

arXiv:2608.10805v2 Announce Type: replace-cross Abstract: Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in replacement for standard convolutions, expanding a network's receptive field…

Source: arXiv cs.AI Amit Aflalo, Shahaf E. Finder, Roy Amoyal, Eran Treister, Oren Freifeld
AI Research AI

Recursive Synthesis for Long-Horizon Terminal Tasks

arXiv:2608.05466v3 Announce Type: replace Abstract: High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because each task must…

Source: arXiv cs.AI Zhongzhi Li, Yucheng Shi, Zongxia Li, Ruhan Wang, Anhao Li, Zixun Huang, Junyao Yang, Lei Ke, Ninghao Liu, Haitao Mi, Leowei Liang
AI Research AI

MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

arXiv:2608.13463v1 Announce Type: cross Abstract: Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and difficulty levels. We propose…

Source: arXiv cs.AI Daniel Perkins, John Squires, Janou Milligan, Chandra Raskoti, Linda Ungerboeck
AI Research AI

Beyond Handcrafted Security: Towards Self-Evolving Defense for LLM Agents

arXiv:2608.12977v1 Announce Type: cross Abstract: The expanding operational capabilities of large language model (LLM) agents introduce sophisticated security threats. Runtime defenses have emerged as an effective…

Source: arXiv cs.AI Jiajun Ruan, Peiyang Li, Yukun Chen, Fengting Li, Chao Feng
AI Research AI

SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries

arXiv:2608.12654v1 Announce Type: new Abstract: Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment. The steering decision is the pre-commit choice at…

Source: arXiv cs.AI Oguz Serdar, Cuneyt Mertayak
AI Research AI

Yes, Q-learning Helps Offline In-Context RL

arXiv:2502.17666v5 Announce Type: replace-cross Abstract: Existing offline in-context reinforcement learning (ICRL) methods have predominantly relied on supervised training objectives, which are known to have…

Source: arXiv cs.AI Denis Tarasov, Alexander Nikulin, Ilya Zisman, Albina Klepach, Andrei Polubarov, Nikita Lyubaykin, Alexander Derevyagin, Igor Kiselev, Vladislav Kurenkov
AI Research AI

Unmasking Conversational Bias in AI Multiagent Systems

arXiv:2501.14844v3 Announce Type: replace-cross Abstract: Detecting biases in the outputs produced by generative models is essential to reduce the potential risks associated with their application in critical settings.…

Source: arXiv cs.AI Erica Coppolillo, Giuseppe Manco, Luca Maria Aiello
AI Research AI

DiG-bench: Discovery in Games

arXiv:2608.12593v1 Announce Type: new Abstract: Discovery---formulating novel generalizations---is a central part of the scientific process. Despite its importance, there is a gap in the current AI benchmark landscape,…

Source: arXiv cs.AI Ruairidh M. Battleday, Kai Sandbrink, Jimi Cullen-Drohan, Zihan Yan, Timothy Muller, Clare Maguire, Ales Kubicek, Fraser Greenlee-Scott, Sukrit Sumant, Tri Dao, J\"urgen Schmidhuber, Michal Valko, Jo…
AI Research AI

Interpretable Causal Discovery via Causal-Effect Constraints

arXiv:2608.12640v1 Announce Type: cross Abstract: Causal discovery aims to uncover the underlying causal relationships given data generated from a system. The goal, however, is not merely to predict causal edges given…

Source: arXiv cs.AI Cixuan Zhang, Guy Van den Broeck, Benjie Wang
AI Research AI

LOB-ID: Evaluating Synthetic Market Data by Inception Distances

arXiv:2608.13082v1 Announce Type: cross Abstract: Generative models of limit orderbook (LOB) data have advanced rapidly, but their evaluation often focuses on stylised facts and selected market statistics. These…

Source: arXiv cs.AI Andreea Bacalum, Zhuohan Wang, Ollie Olby, Martin Garaj, Namid Stillman
AI Research AI

Novels generated by language models show compressed formal variation

arXiv:2608.12630v1 Announce Type: cross Abstract: While large language models can generate entire novels, there is little information about the level of formal variation in their output over many generations. Rather…

Source: arXiv cs.AI Mehdy Sedaghat Payam, Justin Quinn
AI Research AI

SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control

arXiv:2511.07820v4 Announce Type: replace-cross Abstract: Despite the rise of billion-parameter foundation models trained across thousands of graphical processing units (GPUs), similar scaling gains have not been shown…

Source: arXiv cs.AI Zhengyi Luo, Ye Yuan, Tingwu Wang, Chenran Li, Fernando Casta\~neda, Sirui Chen, Zi-Ang Cao, Jiefeng Li, David Minor, Qingwei Ben, Jinhyung Park, David Sami, Zi Wang, Xingye Da, Runyu Ding, Cyrus Hog…
AI Research AI

Demand Transfer Estimation at Scale via Restricted Logit Modeling

arXiv:2608.12680v1 Announce Type: cross Abstract: Item demand forecasting is an integral component of store assortment optimization. Existing literature focuses on learning a suitable customer choice model and using…

Source: arXiv cs.AI Lakshya Garg, Deep Narayan Mishra, Swapnil Yadav, Haoan Wang, Sujal Alugubelli, Karthik Kumaran, Anupriya Sharma
AI Research AI

CAPRI: Contract-Aware Proof Repair for Isabelle

arXiv:2608.13459v1 Announce Type: cross Abstract: We address the use of large language models (LLMs) to help discover Isabelle proofs. An Isabelle build establishes that the submitted theory is accepted, but not that an…

Source: arXiv cs.AI Jim Woodcock, Gabriel Leite, Augusto Sampaio, Ran Wei
AI Research AI

Algebraic Decomposition Theory for Transformer Length Generalization

arXiv:2608.13433v1 Announce Type: cross Abstract: Transformer-based language models are known to sometimes generalize to sequences longer than seen during training, but we lack a precise characterization of which tasks…

Source: arXiv cs.AI Andy Yang, Blerta Veseli, Corentin Barloy, Micha\"el Cadilhac, Andreas Krebs, Charles Paperman, Howard Straubing, Michael Hahn
AI Research AI

A Q-learning-based QoS-aware multipath routing protocol in IoMT-based wireless body area network

arXiv:2604.15489v2 Announce Type: replace-cross Abstract: The Internet of Medical Things (IoMT) enables intelligent healthcare services but faces challenges such as dynamic topology, energy constraints, and diverse QoS…

Source: arXiv cs.AI Mehdi Hosseinzadeh, Roohallah Alizadehsani, Amin Beheshti, Hamid Alinejad-Roknyd, Lu Chen, Mohammad Sadegh Yousefpoor, Efat Yousefpoor, Muneera Altayeb, Thantrira Porntaveetus, Sadia Din
AI Research AI

ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval

arXiv:2608.12720v1 Announce Type: cross Abstract: While Large Language Model (LLM) agents increasingly rely on long-term memory for persistent interactions, the retrieval mechanisms governing this memory are rarely…

Source: arXiv cs.AI Haolong Chen, Liang Zhang, Zhuo Li, Lei Xue, Guanrxu Zhu
AI Research AI

PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research

arXiv:2512.19799v2 Announce Type: replace Abstract: Advances in LLM reasoning and tool use have enabled agentic science, yet frontier theoretical and computational physics remains challenging because research requires…

Source: arXiv cs.AI Tingjia Miao, Wenkai Jin, Jinxin Tan, Muhua Zhang, Xianghe Pang, Zexi Liu, Yuwen Du, Tian Jin, Tu Guo, Zhengliang Zhang, Jingkun Liu, Yuelin Hu, Jiejun Zhang, Yunjie Huang, Yuhan Wang, Wenbo Li, Yinu…
AI Research AI

Do LLMs Know Their Vulnerable Scenarios?

arXiv:2607.23496v2 Announce Type: replace Abstract: Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards.…

Source: arXiv cs.AI Ziheng Peng, Huiqi Deng, Haoran Jing, Xuankun Rong, Jiahui Han, Xiting Wang, Na Zou, Xia Hu
AI Research AI

Learning Latency-Aware Orchestration for Multi-Agent Systems

arXiv:2601.10560v2 Announce Type: replace-cross Abstract: Multi-agent systems (MAS) coordinate multiple LLM-powered agents through structured workflows, gaining reasoning power but incurring high inference latency from…

Source: arXiv cs.AI Xi Shi, Mengxin Zheng, Qian Lou
AI Research AI

AI Guardrail Survival under Single-Cycle Agentic Self-Summarization

arXiv:2608.11392v2 Announce Type: replace-cross Abstract: Long-running agents periodically compact their context, replacing the transcript with a model-generated summary. Recent work shows that dropping a standing…

Source: arXiv cs.AI Ted Kwartler, Alan Aqrawi, Arian Abbasi
AI Research AI

Scaling Time Series Classification via XAI-Driven Data Reduction

arXiv:2607.15774v3 Announce Type: replace-cross Abstract: Explainable AI (XAI) for time series has seen significant algorithmic growth, but its utility in providing measurable performance gains for downstream tasks…

Source: arXiv cs.AI Davide Italo Serramazza, Thach Le Nguyen, Georgiana Ifrim
AI Research AI

Coordinated incentives in AI-generated misinformation governance

arXiv:2608.07070v2 Announce Type: replace-cross Abstract: With the rapid diffusion of AI-generated content, AI-driven misinformation is becoming increasingly pervasive and difficult to govern, undermining information…

Source: arXiv cs.AI Qin Li, Gui Zhang, Minyu Feng, Matjaz Perc, Attila Szolnoki
AI Research AI

Designing AI Pipelines for Decision-Ready ITSM Intelligence

arXiv:2608.12670v1 Announce Type: new Abstract: IT service management (ITSM) systems accumulate large volumes of heterogeneous ticket data that are difficult for sales and executive stakeholders to convert into…

Source: arXiv cs.AI Archan Dutta, Yash Dharmadhikari, Marat Valiullin, Rahul Guha, Alexander Liss
AI Research AI

Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents

arXiv:2607.14573v4 Announce Type: replace Abstract: Payment integration is a demanding repository-level software task: agents must select a suitable product, implement coordinated client-server flows, verify payment…

Source: arXiv cs.AI Shiyu Ying, Xuejie Cao, Yingfan Ma, Yuanhao Dong, Wenyu Chen, Bowen Song, Lin Zhu
AI Research AI

Position: Reasoning is a Learnable Rule-Based Process

arXiv:2608.12325v1 Announce Type: new Abstract: Autonomous reasoning is among the most scientifically and economically motivating topics in AI today. Historically the purview of symbolic AI, recent advances have mainly…

Source: arXiv cs.AI Rachel Lawrence, Jacqueline Maasch
AI Research AI

EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory

arXiv:2608.13113v1 Announce Type: cross Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video…

Source: arXiv cs.AI Weitao Chen, Hu Jiaxin, Xie Tianyidan, Yang Li, Yuyi Qian, Banghao Xu, Ziheng Tang, Shenyi Wang, Mingyue Yu, Duo Li, Jiacheng Shi, Gao Wang, Zhan Xu, Zhicheng Qiu, Xuanfu Li, Jian Yang, Lanjun Wang,…
AI Research AI

Assessment Design in the GenAI Era: The X1-X2-X3 Assessment Pattern for Testing Students' AI Literacy, Learning Outcomes, and Reflection

arXiv:2608.12351v1 Announce Type: cross Abstract: Generative artificial intelligence (GenAI) has challenged the validity of unsupervised online assessment, especially in technical subjects where plausible answers can be…

Source: arXiv cs.AI Riasat Islam (School of Electronic Engineering and Computer Science, Queen Mary University of London, London, United Kingdom), Thomas Roelleke (School of Electronic Engineering and Computer Science,…
AI Research AI

Humans are Missing from AI Coding Agent Research

arXiv:2608.12355v1 Announce Type: cross Abstract: Recent progress in AI coding agent research has led to rapid improvements in agents' ability to autonomously perform complex software engineering tasks, from editing…

Source: arXiv cs.AI Zora Z. Wang, John Yang, Kilian Lieret, Alexa Tartaglini, Valerie Chen, Yuxiang Wei, Zijian Wang, Lingming Zhang, Karthik Narasimhan, Ludwig Schmidt, Graham Neubig, Daniel Fried, Diyi Yang
AI Research AI

QuoteBench: How Matched Scores Can Hide Command-Path Failures

arXiv:2608.13547v1 Announce Type: new Abstract: LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish…

Source: arXiv cs.AI Shangao Li, Yao Zhang, Volker Tresp, Yuanyuan Yang
AI Research AI

CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport

arXiv:2608.13226v1 Announce Type: cross Abstract: While 3D Vision-Language Models (3D VLMs) have demonstrated remarkable spatial reasoning capabilities, they suffer from massive visual token counts that create severe…

Source: arXiv cs.AI Peng Ling, Yingda Yin, Lingting Zhu, Weikai Chen, Shengju Qian, Zeyu Hu, Xin Wang, Wenming Yang
AI Research AI

Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence

arXiv:2608.12645v1 Announce Type: new Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but…

Source: arXiv cs.AI Justin Zhao, Himaghna Bhattacharjee, Hannah Korevaar, Bhaktipriya Radharapu, Khalid El-Arini
AI Research AI

How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

arXiv:2608.13267v1 Announce Type: cross Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with…

Source: arXiv cs.AI Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo, Wei Zhao
AI Research AI

RadarGen: Automotive Radar Point Cloud Generation from Cameras

arXiv:2512.17897v2 Announce Type: replace-cross Abstract: We present RadarGen, a diffusion model for synthesizing realistic automotive radar point clouds from multi-view camera imagery. RadarGen adapts efficient…

Source: arXiv cs.AI Tomer Borreda, Fangqiang Ding, Sanja Fidler, Shengyu Huang, Or Litany
AI Research AI

AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement

arXiv:2608.12329v1 Announce Type: cross Abstract: Progress on AI for psychosis-risk assessment is limited by a data-access bottleneck. Real clinical interviews are difficult to share because of privacy, governance, and…

Source: arXiv cs.AI Guilherme C. Oliveira, Stephanie Fong, Zimu Wang, Clarice Lee, Xiangyu Zhao, Duy Khoa Pham, Duong Nhu, Yiwen Jiang, Jiahe Liu, Zhongxing Xu, Dwarikanath Mahapatra, Dominic Dwyer, Zongyuan Ge
AI Research AI

Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

arXiv:2608.12345v1 Announce Type: new Abstract: Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure remains unmeasured. We introduce…

Source: arXiv cs.AI Yash Tripathi, Silu Sharma, Sai Sidhanth Manoharan Jayanthi, Shivank Garg, Lin Li
AI Research AI

PatientAct: Theory-Grounded Mental Health Client Simulation

arXiv:2608.12750v1 Announce Type: cross Abstract: LLM-based simulated clients are increasingly used to train novice counselors, evaluate LLM therapists, and generate synthetic data. However, current simulators produce…

Source: arXiv cs.AI Sahand Sabour, TszYam NG, Yaqian Chen, Guanqun Bi, Jialu Zhao, Minlie Huang
AI Research AI

Zoom In, Reason Out: Efficient Far-field Anomaly Detection in Expressway Surveillance Videos via Focused VLM Reasoning Guided by Bayesian Inference

arXiv:2604.23724v4 Announce Type: replace-cross Abstract: Expressway video anomaly detection is important for traffic safety, but remains challenging across diverse scenes, particularly for far-field vehicles with…

Source: arXiv cs.AI Xiaowei Mao, Bowen Sui, Weijie Zhang, Yawen Yang, Shengnan Guo, Shilong Zhao, Jiaqi Lin, Tingrui Wu, Youfang Lin, Huaiyu Wan
AI Research AI

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

arXiv:2608.13545v1 Announce Type: cross Abstract: Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to…

Source: arXiv cs.AI Fanfei Li, Jana Zeller, Manuel Prada-Corral, Thadd\"aus Wiedemer, Prasanna Mayilvahanan, Ryan Cotterell, Wieland Brendel
AI Research AI

Auditable Agents

arXiv:2604.05485v2 Announce Type: replace Abstract: LLM agents call tools, query databases, delegate tasks, and trigger external side effects. Once an agent system can act in the world, the question is no longer only…

Source: arXiv cs.AI Yi Nian, Aojie Yuan, Haiyue Zhang, Jiate Li, Li Li, Xiyang Hu, Hua Wei, Xiongye Xiao, Chaowei Xiao, Yue Zhao
AI Research AI

Exploring Sparsity for Parameter Efficient Fine Tuning Using Wavelets for Vision

arXiv:2505.12532v3 Announce Type: replace-cross Abstract: Efficiently adapting large pretrained models is critical under tight compute and memory budgets. While Parameter-Efficient Fine-Tuning (PEFT) methods like LoRA…

Source: arXiv cs.AI Ahmet Bilican, M. Ak{\i}n Y{\i}lmaz, A. Murat Tekalp, R. G\"okberk Cinbi\c{s}
AI Research AI

Research Assistant: AstraZeneca's Agentic System for R&D

arXiv:2608.12395v1 Announce Type: new Abstract: We describe Research Assistant, an internal LLM-based system developed at AstraZeneca to help scientists and clinicians explore biomedical questions across a broad range…

Source: arXiv cs.AI Piotr Grabowski, Mohamed Alameen, Jorge Bretones, Sabina Cardell, Miguel Carmona, Gavin Edwards, Ben Grainger, Sameh Hassan, Erik Jansson, Artur Kuziakhmetov, Albert Maristany, Hebatallah Mohamed, An…
AI Research AI

Governing Agentic AI in FinTech

arXiv:2608.11344v2 Announce Type: replace-cross Abstract: Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little…

Source: arXiv cs.AI Henry Han
AI Research AI

Generative Universal Multimodal Retrieval with Dual-role Identifiers

arXiv:2608.12987v1 Announce Type: cross Abstract: Generative information retrieval (GIR) has emerged as a compelling alternative to the conventional index-retrieve-then-rank retrieval pipeline by training a generator to…

Source: arXiv cs.AI Kaipeng Li, Haitao Yu, Xuanchen Zhou
AI Research AI

From Observation to Intervention: Memory in Brains and Large Language Models

arXiv:2608.12377v1 Announce Type: cross Abstract: Brains and large language models (LLMs) are fundamentally different memory systems, but they can be compared through shared functional questions: where memory-related…

Source: arXiv cs.AI Morteza Salehjahromi, Shayan A. Zadegan, Amgad Muneer, Jia Wu
AI Research AI

AQuA: Recursively Self-Improving Quantitative Trading Research Agents

arXiv:2608.12841v1 Announce Type: cross Abstract: We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from earlier experiments to improve…

Source: arXiv cs.AI Jiacheng Guo, Suozhi Huang, Yunlong Gao, Zihao Li, Jian Ge, Xu Kuang, Mengdi Wang
AI Research AI

It's How You Ask: Gender-Associated Linguistic Bias in LLMs

arXiv:2608.13328v1 Announce Type: cross Abstract: Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more…

Source: arXiv cs.AI Katherine Van Koevering, Anjalie Field
AI Research AI

Operationalizing Cyber Threat Intelligence with GraphRAG

arXiv:2608.13050v1 Announce Type: cross Abstract: When a security researcher publishes a report on a cyberattack, detection engineers are supposed to turn it into working detection rules. In practice, most automated…

Source: arXiv cs.AI Atul Kabra, Prakhar Paliwal, Manjesh K. Hanawal
AI Research AI

Moral Hazard in Multi-Agent Language Models

arXiv:2607.23982v4 Announce Type: replace-cross Abstract: Cooperation can fail when socially valuable effort is costly, hard to observe, and benefits mainly someone else. Building on Holmstr\"om's model of moral hazard…

Source: arXiv cs.AI Dane Malenfant
AI Research AI

A Distributional Robustness Margin For Pathology Foundation Models

arXiv:2607.25497v2 Announce Type: replace-cross Abstract: Pathology foundation models encode non-biological variation introduced by tissue preparation, staining and scanning, enabling shortcut learning that undermines…

Source: arXiv cs.AI Cl\'ement Grisi, Jeroen van der Laak, Geert Litjens
AI Research AI

Cueless EEG imagined speech for subject identification: dataset and benchmarks

arXiv:2501.09700v2 Announce Type: replace-cross Abstract: Electroencephalogram (EEG) signals have emerged as a promising modality for biometric identification. While previous studies have explored the use of imagined…

Source: arXiv cs.AI Ali Derakhshesh, Zahra Dehghanian, Reza Ebrahimpour, Hamid R. Rabiee
AI Research AI

Synthetic Persona Pretraining: Alignment from Token Zero

arXiv:2608.13482v1 Announce Type: cross Abstract: As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical. Today, alignment, and…

Source: arXiv cs.AI Julian Minder, Viktor Moskvoretskii, Raghav Singhal, Difan Jiao, Andy Arditi, Shaobo Cui, Yiderigun Borjigin, Kartik Bali, Stefan Krsteski, Harsh Raj, Huu Nguyen, Jannik Brinkmann, Ashton Anderson, R…
AI Research AI

iARCS: Iterative Agentic RL for Controllable 3D Scene Generation

arXiv:2608.06161v2 Announce Type: replace Abstract: Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism…

Source: arXiv cs.AI Saugat Adhikari, Ashok Prasad Neupane, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel
AI Research AI

SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback

arXiv:2608.13120v1 Announce Type: new Abstract: Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the…

Source: arXiv cs.AI Qianxi Yan, Chunrong Chen, Jiuzhou Zhao, Min Zhang, Yongzhou Xu, Xiaochuan Xu
AI Research AI

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces

arXiv:2608.12585v1 Announce Type: new Abstract: Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong training signals during…

Source: arXiv cs.AI Congchao Wang, Diwakar Singh, Qiaozi Gao, Spyros Matsoukas, Yang Liu, Mahdi Namazifar
AI Research AI

NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

arXiv:2608.12898v1 Announce Type: cross Abstract: Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have…

Source: arXiv cs.AI Peng Cai, Zhaofan Zou, Shifa Liu, Yikun Wang, Jiawei Tang, Kaicheng Yang, Meng Tong, Zhongjiang He, Hao Sun
AI Research AI

Physics-informed distribution of relaxation times estimation and latent-space condition monitoring of solid oxide fuel and electrolysis cells from electrochemical impedance spectroscopy

arXiv:2608.13305v1 Announce Type: cross Abstract: Estimating the distribution of relaxation times (DRT) fromelectrochemical impedance spectroscopy (EIS) is an ill-posed inverse problem that is highly sensitive to…

Source: arXiv cs.AI \v{Z}an Gorenc, \v{Z}iga Gradi\v{s}ar, Felix M\"utter, Vanja Suboti\'c, Pavle Bo\v{s}koski
AI Research AI

Cat-DPO: Category-Adaptive Safety Alignment

arXiv:2604.17299v3 Announce Type: replace-cross Abstract: Aligning large language models with human preferences must balance two competing goals: responding helpfully to legitimate requests and reliably refusing harmful…

Source: arXiv cs.AI Tiankai Yang, Yi Nian, Xinyuan Li, Ruiyao Xu, Henry Peng Zou, Kaize Ding, Xiyang Hu, Yan Liu, Yue Zhao
AI Research AI

Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models

arXiv:2608.12391v1 Announce Type: cross Abstract: Graph reasoning provides a promising testbed for evaluating the reasoning ability of large language models (LLMs), as graph instances can be programmatically generated,…

Source: arXiv cs.AI Fali Wang, Ali Al-Lawati, Iliyas Bektas, Jinxuan Fang, Alek Melenski, Tianxiang Zhao, Yao Ma, Suhang Wang
AI Research AI

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems

arXiv:2606.00367v2 Announce Type: replace-cross Abstract: Reinforcement learning with scalar rewards is widely used for aligning machine-learning systems with user preferences. But, pairwise preferences are often more…

Source: arXiv cs.AI Jonathan Cola\c{c}o Carr, Prakash Panangaden, Doina Precup, Benjamin Van Roy
AI Research AI

PIPES: Securing Agent Perception with Provenance and Priors

arXiv:2608.12789v1 Announce Type: cross Abstract: Tool-using agents consume external data from sources with different levels of trust, yet tool responses rarely identify who produced each component or what it should…

Source: arXiv cs.AI Sanjay Kariyappa, Severin Klingler, G. Edward Suh
AI Research AI

InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers

arXiv:2608.12915v1 Announce Type: cross Abstract: The rapid growth of LLM inference is shifting sustainability concerns from one-time training to continuous serving, where infrastructure decisions shape energy use,…

Source: arXiv cs.AI Nicoletta Tsiopani, Moysis Symeonides, George Pallis, Marios D. Dikaiakos
AI Research AI

Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents

arXiv:2608.12851v1 Announce Type: new Abstract: Self-improving LLM agents convert successful trajectories into persistent cross-task state. An unsafe success can thereby become reusable policy after its triggering input…

Source: arXiv cs.AI Xutao Mao, Liangjie Zhao, Xiang Zheng, Cong Wang
AI Research AI

MatchMiner-AI: Open-source, Privacy-preserving Cancer Clinical Trial Matching using Artificial Intelligence

arXiv:2412.17228v4 Announce Type: replace Abstract: Background: Clinical trials are essential to advancing cancer treatments, but fewer than 10% of adults with cancer enroll in therapeutic trials. Open-source AI trial…

Source: arXiv cs.AI Jennifer Altreuter, Pavel Trukhanov, Morgan A. Paul, Michael J. Hassett, Irbaz B. Riaz, Muhammad Umar Afzal, Arshad A. Mohammed, Ayub Umair, Huan He, Chueh Husan Hsu, Sarah Sammons, James Lindsay, Em…
AI Research AI

The Hidden Evolution of Disguised Visual Context inside the VLM

arXiv:2606.20077v2 Announce Type: replace-cross Abstract: Visual tokens enter Large Language Models (LLMs) as raw, foreign signals. How they are transformed into meaningful representations and interact with the language…

Source: arXiv cs.AI Wish Suharitdamrong, Tony Alex, Xiatian Zhu, Muhammad Awais, Sara Atito
AI Research AI

Deliberate Practice: Learning Robot Skills under a Budget

arXiv:2608.13415v1 Announce Type: cross Abstract: We consider the problem of autonomously learning robot skills under a limited practice budget for sequential tasks. We propose an active skill learning algorithm,…

Source: arXiv cs.AI Shivam Vats, Sudarshan Harithas, Mete Tuluhan Akbulut, Arvind Raghunathan, George Konidaris
AI Research AI

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation

arXiv:2608.11698v2 Announce Type: replace-cross Abstract: On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolation methods such as…

Source: arXiv cs.AI Yang Sun, Lichao Ma, Houyuan Qin, Yuxin Liu, Hanyang Lu, Yao Zhu, Pinlong Cai, Guohang Yan
AI Research AI

FluctlightDB: A Memory Model of Data for AI Agents

arXiv:2608.12365v1 Announce Type: cross Abstract: For fifty years, data systems have answered two questions. The relational model asked which records match a predicate; the vector model asked which vectors lie nearest a…

Source: arXiv cs.AI Ganesh S
AI Research AI

Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments

arXiv:2608.12368v1 Announce Type: new Abstract: Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human…

Source: arXiv cs.AI Octavian M. Machidon, Alina L. Machidon, Vojko Strahovnik, Mateja Centa Strahovnik, Jonas Miklav\v{c}i\v{c}, Marko Robnik \v{S}ikonja
AI Research AI

SE(3)-MeanFlow: Few-Step Protein Backbone Generation on Lie Groups

arXiv:2607.27431v4 Announce Type: replace-cross Abstract: Generative modeling of protein backbones promises the de novo design of proteins with prescribed structural and functional properties. Existing diffusion and…

Source: arXiv cs.AI Yikun Bai, Binghang Lu, Yikai Liu, Elaheh Akbari, Soheil Kolouri, Linxuan Wang, Ping He, Shuchan Wang, Ruqi Zhang, Guang Lin
AI Research AI

MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

arXiv:2608.11616v2 Announce Type: replace Abstract: Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only…

Source: arXiv cs.AI Hojun Choi, Jaeyo Shin, Suin Lee, Hyunjung Shim
AI Research AI

Learning to Recover Task Experts from a Multi-Task Merged Model

arXiv:2606.26902v2 Announce Type: replace Abstract: Multi-task model merging aims to consolidate several task-specific experts into a unified model, yet static merging consistently suffers from parameter interference.…

Source: arXiv cs.AI Jinwook Jung, Taegyu Kim, Kumju Jo, Sungyong Baik
AI Research AI

Tracing Provenance and Detecting Tampering with Complementary LLM Watermarks

arXiv:2608.12713v1 Announce Type: cross Abstract: Watermarking LLM-generated text is an important task for tracing its provenance. Existing LLM watermarks preserve provenance under editing, but this same robustness…

Source: arXiv cs.AI Xiaoyan Feng, Yanjun Zhang, He Zhang, Leo Yu Zhang, Shirui Pan
AI Research AI

StorySpark: Module-wise Evolutionary Search for Story Premise Generation

arXiv:2608.12336v1 Announce Type: cross Abstract: A story premise is the creative spark from which a full narrative can grow. Yet LLM-based story generation has mostly emphasized later-stage planning, controllability,…

Source: arXiv cs.AI Yang Yang, Zining Zhong, Qian Cao, Jindong Li, Boyun Xu, Kaishen Yuan, Menglin Yang, Yutao Yue
AI Research AI

Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing

arXiv:2608.13156v1 Announce Type: new Abstract: Pre-norm is the standard normalization placement in modern Transformers because it facilitates joint optimization of full-depth models. We ask whether this preference…

Source: arXiv cs.AI Sheng Ren, Yadong Wang, Naiqiang Tan, Jiangang Kong, Jun Fang, Rui Liu, Jun Wang, Kai Chen, Lipeng Liang, Xiang Chen
AI Research AI

Safe Exploration via Policy Priors

arXiv:2601.19612v4 Announce Type: replace-cross Abstract: Safe exploration is a key requirement for reinforcement learning (RL) agents to learn and adapt online, beyond controlled (e.g. simulated) environments. In this…

Source: arXiv cs.AI Manuel Wendl, Yarden As, Manish Prajapat, Anton Pollak, Stelian Coros, Andreas Krause
AI Research AI

LLM-Guided Graph Generation for Structure-Based Local Improvement Methods

arXiv:2608.13333v1 Announce Type: new Abstract: Large neighborhood search normally selects a random subset of decision variables for iterative optimization. For efficiently solving different problems, researchers tend…

Source: arXiv cs.AI Hai Xia, Vaidyanathan Peruvemba Ramaswamy, Stefan Szeider
AI Research AI

From Caveman to Expert Analyst: Energy Consumption of Variable LLM Tasks

arXiv:2608.12350v1 Announce Type: cross Abstract: The energy demand growth and environmental impacts of artificial intelligence (AI) have generated substantial interest in supplying sufficient low-cost electricity for…

Source: arXiv cs.AI Diego Manya, Ethan I. Thorpe, Ji Zhang, Myranda Shirk, Jiamian He, Angel Hsu, Michael P. Vandenbergh
AI Research AI

Do Transformers Need Three Projections? Systematic Study of QKV Variants

arXiv:2606.04032v3 Announce Type: replace-cross Abstract: Transformers have become the standard solution for various AI tasks, with the query, key, and value (QKV) attention formulation playing a central role. However,…

Source: arXiv cs.AI Ali Kayyam, Anusha Madan Gopal, M Anthony Lewis
AI Research AI

Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting

arXiv:2608.12590v1 Announce Type: new Abstract: Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation…

Source: arXiv cs.AI Haifan Gong, Shiyu Chen, Bodong Wang, Yuqi Wang, Shijie Wang, Guoliang You, Xinyu Xiong, Haowei Wang, Mingzhi Mao, Dexing Kong, Qinghua Liu, Wei Lou, Fei Chen, Guanbin Li
AI Research AI

On the Expressive Power of Transformers

arXiv:2608.12671v1 Announce Type: new Abstract: Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today. Because of their ubiquity and computational capability,…

Source: arXiv cs.AI Phokion Kolaitis, Rik Sengupta
AI Research AI

Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies

arXiv:2608.12679v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in discovery domains such as math and science. The usual approach is to present the problem to the model and use its…

Source: arXiv cs.AI Conor F. Hayes, Elliot Meyerson, Kajetan Schweighofer, Roberto Dailey, Babak Hodjat, Risto Miikkulainen, Xin Qiu
AI Research AI

AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)

arXiv:2608.13492v1 Announce Type: new Abstract: This report presents an improved version of AlayaWorld. While the backbone architecture, chunk-wise autoregressive generation scheme, and training data remain unchanged…

Source: arXiv cs.AI AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Z…
AI Research AI

SchemaLink: An Intelligent Web Editor for LinkML Schema Curation

arXiv:2608.12529v1 Announce Type: cross Abstract: Motivation: LinkML is a suitable language for the representation of the structural and content constraints of different kinds of biomedical data. Even if it is a quite…

Source: arXiv cs.AI Emanuele Cavalleri, Paolo Perlasca, J. Harry Caufield, Justin Reese, Christopher J. Mungall, Marco Mesiti
AI Research AI

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

arXiv:2608.13560v1 Announce Type: cross Abstract: Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process centered on a…

Source: arXiv cs.AI Yaxin Luo, Haobin Jiang, Jialv Zou, Xu Huang, Wenhao Yan, Haodong Li, Zhengrong Yue, Jing Li, Xiaofu Chen, Xiaohan Zhao, Jiacheng Liu, Jiacheng Cui, Zhiqiang Shen, Xiaotong Li
AI Research AI

FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving

arXiv:2608.12932v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far too high for real-time control.…

Source: arXiv cs.AI Zekai Li, Yihao Liang, Hongfei Zhang, Jian Chen, Yesheng Liang, Zhijian Liu
AI Research AI

Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning

arXiv:2608.12372v1 Announce Type: new Abstract: AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers. This position paper argues that in many settings, particularly…

Source: arXiv cs.AI Vijay Keswani, Breanna K. Nguyen, Cyrus Cousins, Vincent Conitzer, Walter Sinnott-Armstrong, Jana Schaich Borg
AI Research AI

A Hierarchical Energy-Based Model for Multimodal Cognition

arXiv:2608.12398v1 Announce Type: cross Abstract: We propose IM-LEPP (Integrated Multimodal Latent Energy-based Predictive Processing), a hierarchical, energy-based model of multimodal cognition that extends a…

Source: arXiv cs.AI Subir Varma
AI Research AI

The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis

arXiv:2608.12677v1 Announce Type: new Abstract: Detecting infection-related behavioral changes in mosquitoes from video data is challenging because mosquitoes are small, move rapidly and irregularly, and are affected by…

Source: arXiv cs.AI Danial Sharifrazi, Saadat Behzadi, Julakha Jahan Jui, Mojtaba Mohammadi, Nouman Javed, Roohallah Alizadehsani, Prasad N. Paradkar, Asim Bhatti
AI Research AI

Novel Knowledge-Guided Generative Methods for Synthetic Transcriptomic Data

arXiv:2608.13256v1 Announce Type: cross Abstract: As biomedical research increasingly relies on data-intensive tools, the quality and utility of datasets are critical. Challenges such as imbalances, biases, and ethical…

Source: arXiv cs.AI Francesca Pia Panaccione, Sofia Mongardi, Marco Masseroli, Pietro Pinoli
AI Research AI

@skills: Attention is all you have

arXiv:2608.12610v1 Announce Type: new Abstract: There are 56,804 public agent skills today, and teams write many more privately. The dominant delivery model is installation: once installed, a skill's description remains…

Source: arXiv cs.AI Li Yin (Atlas), Zhi Li (Atlas), Zhan Shi (Atlas), Haoran Zhang (Atlas), Haebin Seong (Atlas), Zhangyang (Atlas), Wang
AI Research AI

Query Timing Produces Opposite Positional Biases Between LLMs and Humans

arXiv:2608.12387v1 Announce Type: cross Abstract: Positional biases such as recency and primacy effects have been documented in large language models (LLMs), yet the underlying mechanism by which these models make their…

Source: arXiv cs.AI Jasin Cekinmez, Addison J. Wu, Thomas L. Griffiths
AI Research AI

Memorization Diagnostics for Code LLMs Should be Scale-Aware

arXiv:2608.12771v1 Announce Type: cross Abstract: The extent to which large language models for code rely on memorization over genuine understanding remains highly debated. While current literature frequently reports…

Source: arXiv cs.AI Prateek Kumar Rajput, Abdoul Aziz Bonkoungou, Alberick Euraste Djir\'e, Xunzhu Tang, Yewei Song, Iyiola Emmanuel Olatunji, El Hacen Diallo, Jacques Klein, Tegawend\'e F. Bissyand\'e
AI Research AI

Academic League of Artificial Intelligence - An Integrative Perspective of Teaching, Research, and Extension

arXiv:2608.13447v1 Announce Type: new Abstract: Academic leagues have become important mechanisms for promoting extracurricular education and strengthening the integration between universities and society. This paper…

Source: arXiv cs.AI Alison R. Panisson, Maria Eduarda W. M. Vianna, Italo Firmino da Silva, Heitor Henrique da Silva, Rafaela Fernandes Savaris, Bernardo Pandolfi Costa, Martin Augusto Gagliotti Vigil, Jim Lau, Agenor H…
AI Research AI

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

arXiv:2608.13417v1 Announce Type: new Abstract: Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current…

Source: arXiv cs.AI Yiwei Li, Wanli Yang, Hexiang Tan, Xiangzhou Huang, Zhengyu Chen, Ziran Li, Borun Chen, Shanglin Lei, Huaisheng Zhu, Hao Tian, Fei Sun, Xunliang Cai, Jingang Wang
AI Research AI

vToken: Token-Level Virtualization for Reclaimable KV Caches

arXiv:2608.13263v1 Announce Type: new Abstract: Large language model serving faces a critical memory bottleneck: the KV cache grows with sequence length and batch size. PagedAttention uses fixed-size memory blocks to…

Source: arXiv cs.AI Yuanhang Gao, Xiangrui Yang, Yuanfeng Chen, Hongjia Chen, Qianru Lv, Wenfei Wu, Dongsheng Li
AI Research AI

AutoQuREO: A Framework for Automated Quantum Resource Estimation and Optimization

arXiv:2608.12936v1 Announce Type: cross Abstract: As quantum computing progresses from proof-of-principle demonstrations toward practical utility, a significant impediment is the need to augment algorithmic feasibility…

Source: arXiv cs.AI Harshkumar Oza, Aritra Sarkar, Syed Naqi Abbas, Rahul Bhowmick, Aryan Prakash, Prateek P Kulkarni, Krishna Kumar Sabapathy
AI Research AI

Error-Aware Reverse Auction Mechanism for Large Language Model Routing

arXiv:2608.12719v1 Announce Type: cross Abstract: Routing each query to a cost-effective large language model (LLM) is critical for balancing quality and cost, yet most routers rely on a centralized task center to…

Source: arXiv cs.AI Haolong Chen, Zhengyuan Xin, Liang Zhang, Lei Xue, Guangxu Zhu