Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality

May 05, 2025

Xueguang Ma, Luyu Gao, Shengyao Zhuang, Jiaqi Samantha Zhan, Jamie Callan, Jimmy Lin

Figure 1 for Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality

Figure 2 for Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality

Figure 3 for Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality

Figure 4 for Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality

Share this with someone who'll enjoy it:

Abstract:Recent advancements in large language models (LLMs) have driven interest in billion-scale retrieval models with strong generalization across retrieval tasks and languages. Additionally, progress in large vision-language models has created new opportunities for multimodal retrieval. In response, we have updated the Tevatron toolkit, introducing a unified pipeline that enables researchers to explore retriever models at different scales, across multiple languages, and with various modalities. This demo paper highlights the toolkit's key features, bridging academia and industry by supporting efficient training, inference, and evaluation of neural retrievers. We showcase a unified dense retriever achieving strong multilingual and multimodal effectiveness, and conduct a cross-modality zero-shot study to demonstrate its research potential. Alongside, we release OmniEmbed, to the best of our knowledge, the first embedding model that unifies text, image document, video, and audio retrieval, serving as a baseline for future research.

* Accepted in SIGIR 2025 (Demo)

View paper on

Share this with someone who'll enjoy it:

Title:Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality

Paper and Code