Skip to content
TILens What matters today in tech v0.0.9
Theme

Topic - Edition

AI Research

100 items on 29 Aug 2026
AI Research AI Python

4 Claude Skills Every Data Scientist Needs in 2026

Four skills worth adding to your workflow today if you don't want to be left behind The post 4 Claude Skills Every Data Scientist Needs in 2026 appeared first on Towards Data Science.

Source: Towards Data Science Haden Pelletier
AI Research AI

Robust Code RL via Faulty-Code-Driven Test case Synthesis and Dense Reward Shaping

arXiv:2608.24135v2 Announce Type: replace Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) is pivotal for enhancing LLM code generation, yet its efficacy is often hindered by insufficient test case…

Source: arXiv cs.AI Yiwen Zhang, Xiaodong Yan, Zhenyu Huang, Deng Zhao, Liang Jiang, Qing Cui, Zujie Wen, Zhiqiang Zhang, Jun Zhou
AI Research AI

Can You Say This for Me? Speaking Up by Proxy in Co-Located Discussion

arXiv:2608.26185v1 Announce Type: new Abstract: Equal participation in co-located discussion is important for effective collaboration, yet people often hold back when they anticipate negative interpersonal or…

Source: arXiv cs.AI Yue Shen, Rehema Abulikemu, Ryan P. McMahan, Yan Chen
AI Research AI

UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on

arXiv:2608.05745v2 Announce Type: replace-cross Abstract: Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics. Dominant approaches…

Source: arXiv cs.AI Yushe Cao, Shikun Feng, Fei Shen, Haikuo Peng, Jianqiang Xia, Yiheng Zhu, Dianxi Shi, Chun Yu
AI Research AI

MOMO: A framework for seamless physical, verbal, and graphical robot skill learning and adaptation

arXiv:2604.20468v3 Announce Type: replace-cross Abstract: Industrial robot applications require increasingly flexible systems that non-expert users can easily adapt for varying tasks and environments. However, different…

Source: arXiv cs.AI Markus Knauer, Edoardo Fiorini, Maximilian M\"uhlbauer, Stefan Schneyer, Promwat Angsuratanawech, Florian Samuel Lay, Timo Bachmann, Samuel Bustamante, Korbinian Nottensteiner, Freek Stulp, Alin Albu…
AI Research AI

PACEShop: Evaluating Personalized, Actionable, Compositional, and Evidence-grounded Shopping Assistants

arXiv:2608.26180v1 Announce Type: cross Abstract: Shopping assistants are shifting from ranked product lists toward structured decision support, where systems must synthesize shopper context, product evidence, and…

Source: arXiv cs.AI Weimin Lyu, Chen Luo, Guangrui Li, Yaochen Xie, Dhineshkumar Ramasubbu, Arief Koesdwiady, Wanqiu Long, Hansu Gu, Yutong Chen, Zheshen Wang, Dakuo Wang, Yi Liu
AI Research AI

CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models

arXiv:2601.04778v2 Announce Type: replace-cross Abstract: Video-language models (VLMs) achieve strong multimodal understanding but remain prone to hallucinations, especially when reasoning about actions and temporal…

Source: arXiv cs.AI Tobia Poppi, Burak Uzkent, Amanmeet Garg, Lucas Porto, Garin Kessler, Yezhou Yang, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara, Florian Schiffers
AI Research AI

Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics

arXiv:2508.11017v3 Announce Type: replace-cross Abstract: Large language models (LLMs) struggle with cross-lingual knowledge transfer: they sometimes hallucinate when asked in one language about facts expressed in a…

Source: arXiv cs.AI Carter Blum, Katja Filippova, Ann Yuan, Asma Ghandeharioun, Julian Zimmert, Fred Zhang, Jessica Hoffmann, Tal Linzen, Martin Wattenberg, Lucas Dixon, Mor Geva
AI Research AI

LLM-Specific Utility for Retrieval-Augmented Generation

arXiv:2510.11358v4 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) is typically optimized for topical relevance, yet its success ultimately depends on whether retrieved passages are useful…

Source: arXiv cs.AI Hengran Zhang, Keping Bi, Jiafeng Guo, Jiaming Zhang, Shuaiqiang Wang, Dawei Yin, Xueqi Cheng
AI Research AI

ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices

arXiv:2608.26204v1 Announce Type: cross Abstract: Computer Use Agents (CUAs) are increasingly deployed to navigate mobile and desktop applications on behalf of users, yet no benchmark comprehensively evaluates whether…

Source: arXiv cs.AI Joy Chen, Alejandro Castillejo Munoz, Pierluca D'Oro, Yuxuan Sun, Chloe Evans, Joseph Tighe
AI Research AI

Evaluating AI Generated Summaries for Cancer Patients

arXiv:2608.26154v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly being integrated into digital health platforms to generate summaries of complex medical data. Although these models can…

Source: arXiv cs.AI Muhammad Aurangzeb Ahmad, Kim Shyu, Leon Oliver, Fergus Sleight, Paul Landau
AI Research AI

CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases

arXiv:2608.27391v1 Announce Type: new Abstract: LLMs are increasingly able to answer complex questions about enterprise-scale document collections. But evaluation is hard: companies don't want to share internal…

Source: arXiv cs.AI Sil Hamilton, Albert Yu Sun, Oscar J. Romero, Carl-Leander Henneking, David Mimno, Bishan Yang, Igor Labutov
AI Research AI

Decoupling Planning and Control for Instructable Agents

arXiv:2608.26788v1 Announce Type: new Abstract: Recent work shows that pre-trained, instruction-tuned vision-language models (VLMs) perform well at mapping from instructions and observations to high-level plans, but…

Source: arXiv cs.AI Zineng Tang, Kelsey R. Allen, Sjoerd van Steenkiste, Ishita Dasgupta, Alane Suhr

100 items available