CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Tevatron /browsecomp-plus BrowseComp-Plus BrowseComp-Plus is a new benchmark for Deep-Research system, isolating the effect of the retriever and the LLM agent to enable fair, transparent comparisons of Deep-Research agents. The benchmark sources challenging, reasoning-intensive queries from OpenAI's BrowseComp. However, instead of searching the live web, BrowseComp-Plus evaluates against a fixed, curated corpus of ~100K web documents from the web. The corpus includes both human-verified evidence documents… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/browsecomp-plus.textquestion-answeringn<1K37 likes43k downloads9mo agoHugging Face02Tevatron /browsecomp-plus-corpus BrowseComp-Plus Project Page | Paper | Code BrowseComp-Plus is a new benchmark for Deep-Research system, isolating the effect of the retriever and the LLM agent to enable fair, transparent comparisons of Deep-Research agents. The benchmark sources challenging, reasoning-intensive queries from OpenAI's BrowseComp. However, instead of searching the live web, BrowseComp-Plus evaluates against a fixed, curated corpus of ~100K web documents from the web. The corpus includes both… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/browsecomp-plus-corpus.textquestion-answering100K<n<1M18 likes31k downloads1y agoHugging Face03Tevatron /wiki-ss-corpusimage1M<n<10M6 likes9k downloads2y agoHugging Face04Tevatron /msmarco-passagetext100K<n<1M10 likes1.9k downloads7mo agoHugging Face05Tevatron /wiki-ss-corpus-newimage1M<n<10M1 likes963 downloads2y agoHugging Face06Tevatron /beir-corpus BeIR Corpus text10M<n<100M0 likes880 downloads5mo agoHugging Face07Tevatron /wikipedia-nqtext10K<n<100K7 likes593 downloads5y agoHugging Face08Tevatron /msmarco-passage-corpustext1M<n<10M13 likes573 downloads1y agoHugging Face09Tevatron /wikipedia-triviatext10K<n<100K3 likes556 downloads5y agoHugging Face10Tevatron /audiocaps-corpusaudio10K<n<100K0 likes445 downloads1y agoHugging Face11Tevatron /colpali-corpus Overview This dataset is a transformed version of the ColPali (vidore/colpali_train_set) dataset, formatted to be compatible with Tevatron. image100K<n<1M0 likes349 downloads2y agoHugging Face12Tevatron /beirtext10K<n<100K1 likes347 downloads4y agoHugging Face13Tevatron /wiki-ss-nqtext10K<n<100K4 likes327 downloads2y agoHugging Face14nthakur /cornstack-php-v1-tevatron-1Mtext100K<n<1M0 likes310 downloads1y agoHugging Face15Tevatron /bge-ir-prompttext1M<n<10M0 likes271 downloads2y agoHugging Face16Tevatron /bge-irtext1M<n<10M1 likes238 downloads2y agoHugging Face17Tevatron /scifact SciFact text1K<n<10K2 likes237 downloads4mo agoHugging Face18Tevatron /wikipedia-curatedtext1K<n<10K1 likes205 downloads5y agoHugging Face19Tevatron /AgentIR-dataThis dataset contains the DR-Synth generated WebShaper data for training AgentIR-4B. Paper: AgentIR: Reasoning-Aware Retrieval for Deep Research Agents Code: https://github.com/texttron/AgentIR Project Page: https://texttron.github.io/AgentIR/ Model: AgentIR-4B Dataset Details Each instance contains: query_id: {webshaper_query_id}_turn{i}, where webshaper_query_id is the original id in WebShaper, and i is the turn number during agent rollout when constructing the data. query:… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/AgentIR-data.texttext-retrieval1K<n<10K9 likes182 downloads29d agoHugging Face20nthakur /cornstack-6-langs-v1-tevatron-6Mtext1M<n<10M0 likes178 downloads1y agoHugging Face21Tevatron /wikipedia-wqtext1K<n<10K0 likes155 downloads5y agoHugging Face22Tevatron /browsecomp-plus-md-toc-gpt5.4-nano BrowseComp-Plus Structured 100k Corpus This dataset is a drop-in, official-format variant of the BrowseComp-Plus 100k corpus used in our RISE Agent experiments. It keeps the same row count, document ids, URLs, and column names as the original BrowseComp-Plus corpus, but replaces each document's text field with a structured version that adds a generated table of contents and section headings. Files data.parquet: the corpus in the same three-column schema as the… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/browsecomp-plus-md-toc-gpt5.4-nano.textquestion-answering100K<n<1M0 likes133 downloads3mo agoHugging Face23Tevatron /pixmo-docs-corpusimage100K<n<1M0 likes123 downloads2y agoHugging Face24Tevatron /reasonir-data-hnThis is the dataset forked from https://huggingface.co/datasets/reasonir/reasonir-data with hard negatives expanded by BM25. Please follow the original License. text10K<n<100K3 likes88 downloads1y agoHugging Face25Tevatron /docmatix-ir Docmatix-IR Docmatix is originally a large dataset designed for fine-tuning large vision-language models on Visual Question Answering tasks. It contains a substantial collection of PDF images (2.4M) and a vast set of questions (9.5M) related to these images. However, many of the questions in the Docmatix dataset are not suitable for open-domain question answering. To address this, we have converted Docmatix into Docmatix-IR, a training set suitable for training document visual… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/docmatix-ir.textquestion-answering1M<n<10M15 likes74 downloads2y agoHugging Face26nthakur /cornstack-python-v1-tevatron-1Mtext100K<n<1M0 likes70 downloads1y agoHugging Face27nthakur /cornstack-go-v1-tevatron-1Mtext100K<n<1M0 likes63 downloads1y agoHugging Face28Tevatron /msrvtt-corpustext10K<n<100K0 likes56 downloads1y agoHugging Face29Tevatron /brightThis is the dataset forked from https://huggingface.co/datasets/xlangai/BRIGHT. Please follow the original License. image1K<n<10K0 likes47 downloads1y agoHugging Face30Tevatron /webshaper-fineweb-1m-corpustext1M<n<10M0 likes46 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.