Two BM25 optimizations and benchmark configuration changes inspired by PlanetScale's TIN benchmarks make ParadeDB's text search faster without changing its document identifiers.
I am currently working on an open-source web app store project. It utilizes a hybrid architecture: a fast server-side rendering (SSR) engine that is subsequently taken over by a lightweight JavaScript program on the…
arXiv:2609.39225v1 Announce Type: cross Abstract: Argument structure prediction (ASP) constructs complete argument structures from discourse by identifying argumentative units and their relations. While recent work has…
arXiv:2609.39358v1 Announce Type: cross Abstract: A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys, argues that this bound limits…
arXiv:2609.38374v1 Announce Type: cross Abstract: Representing scientific papers as points in a space lets us search for similar papers and inquire about how fields relate to one another and drive innovation. Beyond…
arXiv:2609.39209v1 Announce Type: cross Abstract: Pulling a fixed set of fields out of documents that arrive in many formats and under drifting schemas is usually done with hand-written byte patterns, which break…
arXiv:2608.13786v2 Announce Type: replace Abstract: Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant studies, yet the quality of retrieved evidence and…
arXiv:2609.39319v1 Announce Type: new Abstract: Generative retrieval has emerged as a general retrieval paradigm, representing items with discrete Semantic IDs (SIDs) and retrieving them through autoregressive…
arXiv:2609.39696v1 Announce Type: new Abstract: Synthetic dialogues generated by LLM pipelines now serve as complete conversational-recommendation benchmarks: an LLM listener talks to an LLM recommender, and the track…
arXiv:2606.15911v2 Announce Type: replace-cross Abstract: This paper focuses on automatically generating informative ad descriptions in sponsored search. Unlike ad titles which are usually optimized to attract user…
arXiv:2609.37574v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly used to enrich user queries in information retrieval (IR) so that a standard retriever such as BM25 can bridge vocabulary…
arXiv:2609.38946v1 Announce Type: cross Abstract: Generative AI search and AI overviews are transforming access to information and news, renewing concerns that readers will encounter a narrower range of topics and have…
arXiv:2609.38822v1 Announce Type: new Abstract: Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills,…
arXiv:2605.23916v2 Announce Type: replace Abstract: AI agents often pick tools from registries, where each tool's provider writes its description. We ask whether sales language in those descriptions changes which tool…
arXiv:2603.02561v2 Announce Type: replace Abstract: Attention mechanism remains the defining operator in Transformers since it provides expressive global credit assignment, yet its quadratic cost in sequence length N…
arXiv:2609.38455v1 Announce Type: new Abstract: While recent multimodal recommender systems have demonstrated the effectiveness of incorporating visual and textual information to improve downstream performance, most…
arXiv:2609.38473v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is now the standard way to ground Large Language Models (LLMs) in external knowledge, yet the design space of retrieval pipelines is…
arXiv:2608.05543v5 Announce Type: replace Abstract: We present omni-macos, a search engine that embeds text, code, documents, images, audio and video into one representation space and runs its encoder, index and store…
arXiv:2609.39327v1 Announce Type: new Abstract: Generative retrieval reformulates recommendation as the generation of discrete item tokens. However, scaling this paradigm to real-world recommender systems reveals two…