CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01illuin-conteb /narrative-qa ConTEB - NarrativeQA This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It stems from the widely used NarrativeQA dataset. Dataset Summary NarrativeQA (literature), consists of long documents, associated to existing sets of question-answer pairs. To build the corpus, we start from the pre-existing collection documents, extract the text, and chunk them (using LangChain's… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/narrative-qa.text10K<n<100K5 likes1.2k downloads1y agoHugging Face02illuin-conteb /mldr-conteb-train ConTEB - MLDR (training) This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It stems from the widely used MLDR dataset. Dataset Summary MLDR consists of long documents, associated to existing sets of question-answer pairs. To build the corpus, we start from the pre-existing collection documents, extract the text, and chunk them (using LangChain's RecursiveCharacterSplitter with a… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/mldr-conteb-train.text100K<n<1M0 likes875 downloads1y agoHugging Face03illuin-conteb /squad-conteb-train ConTEB - SQuAD (training) This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It stems from the widely used SQuAD dataset. Dataset Summary SQuAD is an extractive QA dataset with questions associated to passages and annotated answer spans, that allow us to chunk individual passages into shorter sequences while preserving the original annotation. To build the corpus, we start from the… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/squad-conteb-train.text10K<n<100K0 likes650 downloads1y agoHugging Face04illuin-conteb /covid-qa ConTEB - Covid-QA This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It focuses on the theme of Healthcare, particularly stemming from articles about the COVID-19 pandemic. Dataset Summary This dataset was designed to elicit contextual information. It is built upon the COVID-QA dataset. To build the corpus, we start from the pre-existing collection documents, extract the text, and chunk… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/covid-qa.text1K<n<10K1 likes603 downloads1y agoHugging Face05illuin-conteb /geography ConTEB - Geography This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It focuses on the theme of Geography, particularly stemming from Wikipedia pages of cities around the world. Dataset Summary This dataset was designed to elicit contextual information. To build the corpus, we collect Wikipedia pages of large cities, extract the text, and chunk them (using LangChain's… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/geography.text10K<n<100K2 likes504 downloads1y agoHugging Face06illuin-conteb /football ConTEB - Football This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It focuses on the theme of Sports, particularly stemming from Footballer Wikipedia pages. Dataset Summary This dataset was designed to elicit contextual information. To build the corpus, we collect Wikipedia pages of famous footballers, extract the text, and chunk them (using LangChain's RecursiveCharacterSplitter with… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/football.text1K<n<10K1 likes480 downloads1y agoHugging Face07illuin-conteb /mldr-conteb-eval ConTEB - MLDR (evaluation) This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It stems from the widely used MLDR dataset. Dataset Summary MLDR consists of long documents, associated to existing sets of question-answer pairs. To build the corpus, we start from the pre-existing collection documents, extract the text, and chunk them (using LangChain's RecursiveCharacterSplitter with a… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/mldr-conteb-eval.text10K<n<100K2 likes457 downloads1y agoHugging Face08illuin-conteb /squad-conteb-eval ConTEB - SQuAD (evaluation) This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It stems from the widely used SQuAD dataset. Dataset Summary SQuAD is an extractive QA dataset with questions associated to passages and annotated answer spans, that allow us to chunk individual passages into shorter sequences while preserving the original annotation. To build the corpus, we start from the… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/squad-conteb-eval.text100K<n<1M1 likes440 downloads1y agoHugging Face09illuin-conteb /insurance ConTEB - Insurance This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It focuses on the theme of Insurance, particularly stemming from a document of the EIOPA entity. Dataset Summary Insurance is composed of a long document with insurance-related statistics for each country of the European Union. To build the corpus, we extract the text of the document, and chunk it (using LangChain's… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/insurance.textn<1K1 likes417 downloads1y agoHugging Face10illuin-conteb /esg-reports ConTEB - ESG Reports This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It focuses on the theme of Industrial ESG Reports, particularly stemming from the fast-food industry. Dataset Summary This dataset was designed to elicit contextual information. It is built upon a subset of the ViDoRe Benchmark. To build the corpus, we start from the pre-existing collection of ESG Reports, extract… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/esg-reports.text1K<n<10K1 likes169 downloads1y agoHugging Face11illuin-conteb /squad-chunked-par-1000text1K<n<10K0 likes53 downloads2y agoHugging Face12illuin-conteb /squad-chunked-par-175text10K<n<100K0 likes48 downloads2y agoHugging Face13illuin-conteb /chunked-mldr-big-100text1M<n<10M0 likes47 downloads2y agoHugging Face14illuin-conteb /squad-chunked-par-150text10K<n<100K0 likes44 downloads2y agoHugging Face15illuin-conteb /maritime-qatextn<1K0 likes43 downloads2y agoHugging Face16illuin-conteb /squad-chunked-par-300text1K<n<10K0 likes43 downloads2y agoHugging Face17illuin-conteb /squad-chunked-par-250text1K<n<10K0 likes43 downloads2y agoHugging Face18illuin-conteb /us_constitutiontextn<1K0 likes42 downloads2y agoHugging Face19illuin-conteb /squad-chunked-par-200text10K<n<100K0 likes40 downloads2y agoHugging Face20illuin-conteb /football_oldtext1K<n<10K0 likes39 downloads2y agoHugging Face21illuin-conteb /squad-chunked-500text100K<n<1M0 likes38 downloads2y agoHugging Face22illuin-conteb /squad-chunked-par-500text10K<n<100K0 likes37 downloads2y agoHugging Face23illuin-conteb /squad-chunked-par-8000text1K<n<10K0 likes37 downloads2y agoHugging Face24illuin-conteb /squad-chunked-100text100K<n<1M0 likes32 downloads2y agoHugging Face25illuin-conteb /squad-chunked-par-125text10K<n<100K0 likes32 downloads2y agoHugging Face26illuin-conteb /football-maxtextn<1K0 likes30 downloads2y agoHugging Face27illuin-conteb /squad-shorttext10K<n<100K0 likes27 downloads2y agoHugging Face28illuin-conteb /mldr-evaltext10K<n<100K0 likes27 downloads2y agoHugging Face29illuin-conteb /squad-chunked-1000text10K<n<100K0 likes25 downloads2y agoHugging Face30illuin-conteb /squad-chunked-8000text10K<n<100K0 likes25 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.