datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
narrative-qa
ConTEB - NarrativeQA
This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It stems from the widely used NarrativeQA dataset.
Dataset Summary
NarrativeQA (literature), consists of long documents, associated to existing sets of question-answer pairs. To build the corpus, we start from the pre-existing collection documents, extract the text, and chunk them (using LangChain's… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/narrative-qa.mldr-conteb-train
ConTEB - MLDR (training)
This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It stems from the widely used MLDR dataset.
Dataset Summary
MLDR consists of long documents, associated to existing sets of question-answer pairs. To build the corpus, we start from the pre-existing collection documents, extract the text, and chunk them (using LangChain's RecursiveCharacterSplitter with a… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/mldr-conteb-train.squad-conteb-train
ConTEB - SQuAD (training)
This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It stems from the widely used SQuAD dataset.
Dataset Summary
SQuAD is an extractive QA dataset with questions associated to passages and annotated answer spans, that allow us to chunk individual passages into shorter sequences while preserving the original annotation. To build the corpus, we start from the… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/squad-conteb-train.covid-qa
ConTEB - Covid-QA
This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It focuses on the theme of Healthcare, particularly stemming from articles about the COVID-19 pandemic.
Dataset Summary
This dataset was designed to elicit contextual information. It is built upon the COVID-QA dataset. To build the corpus, we start from the pre-existing collection documents, extract the text, and chunk… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/covid-qa.geography
ConTEB - Geography
This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It focuses on the theme of Geography, particularly stemming from Wikipedia pages of cities around the world.
Dataset Summary
This dataset was designed to elicit contextual information. To build the corpus, we collect Wikipedia pages of large cities, extract the text, and chunk them (using LangChain's… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/geography.football
ConTEB - Football
This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It focuses on the theme of Sports, particularly stemming from Footballer Wikipedia pages.
Dataset Summary
This dataset was designed to elicit contextual information. To build the corpus, we collect Wikipedia pages of famous footballers, extract the text, and chunk them (using LangChain's RecursiveCharacterSplitter with… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/football.mldr-conteb-eval
ConTEB - MLDR (evaluation)
This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It stems from the widely used MLDR dataset.
Dataset Summary
MLDR consists of long documents, associated to existing sets of question-answer pairs. To build the corpus, we start from the pre-existing collection documents, extract the text, and chunk them (using LangChain's RecursiveCharacterSplitter with a… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/mldr-conteb-eval.squad-conteb-eval
ConTEB - SQuAD (evaluation)
This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It stems from the widely used SQuAD dataset.
Dataset Summary
SQuAD is an extractive QA dataset with questions associated to passages and annotated answer spans, that allow us to chunk individual passages into shorter sequences while preserving the original annotation. To build the corpus, we start from the… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/squad-conteb-eval.insurance
ConTEB - Insurance
This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It focuses on the theme of Insurance, particularly stemming from a document of the EIOPA entity.
Dataset Summary
Insurance is composed of a long document with insurance-related statistics for each country of the European Union. To build the corpus, we extract the text of the document, and chunk it (using LangChain's… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/insurance.esg-reports
ConTEB - ESG Reports
This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It focuses on the theme of Industrial ESG Reports, particularly stemming from the fast-food industry.
Dataset Summary
This dataset was designed to elicit contextual information. It is built upon a subset of the ViDoRe Benchmark. To build the corpus, we start from the pre-existing collection of ESG Reports, extract… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/esg-reports.squad-chunked-par-1000squad-chunked-par-175chunked-mldr-big-100squad-chunked-par-150maritime-qasquad-chunked-par-300squad-chunked-par-250us_constitutionsquad-chunked-par-200football_oldsquad-chunked-500squad-chunked-par-500squad-chunked-par-8000squad-chunked-100squad-chunked-par-125football-maxsquad-shortmldr-evalsquad-chunked-1000squad-chunked-8000
