conteb
Datasets
All datasets matching “conteb”narrative-qa
ConTEB - NarrativeQA
This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It stems from the widely used NarrativeQA dataset.
Dataset Summary
NarrativeQA (literature), consists of long documents, associated to existing sets of question-answer pairs. To build the corpus, we start from the pre-existing collection documents, extract the text, and chunk them (using LangChain's… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/narrative-qa.mldr-conteb-train
ConTEB - MLDR (training)
This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It stems from the widely used MLDR dataset.
Dataset Summary
MLDR consists of long documents, associated to existing sets of question-answer pairs. To build the corpus, we start from the pre-existing collection documents, extract the text, and chunk them (using LangChain's RecursiveCharacterSplitter with a… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/mldr-conteb-train.squad-conteb-train
ConTEB - SQuAD (training)
This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It stems from the widely used SQuAD dataset.
Dataset Summary
SQuAD is an extractive QA dataset with questions associated to passages and annotated answer spans, that allow us to chunk individual passages into shorter sequences while preserving the original annotation. To build the corpus, we start from the… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/squad-conteb-train.covid-qa
ConTEB - Covid-QA
This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It focuses on the theme of Healthcare, particularly stemming from articles about the COVID-19 pandemic.
Dataset Summary
This dataset was designed to elicit contextual information. It is built upon the COVID-QA dataset. To build the corpus, we start from the pre-existing collection documents, extract the text, and chunk… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/covid-qa.geography
ConTEB - Geography
This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It focuses on the theme of Geography, particularly stemming from Wikipedia pages of cities around the world.
Dataset Summary
This dataset was designed to elicit contextual information. To build the corpus, we collect Wikipedia pages of large cities, extract the text, and chunk them (using LangChain's… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/geography.football
ConTEB - Football
This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It focuses on the theme of Sports, particularly stemming from Footballer Wikipedia pages.
Dataset Summary
This dataset was designed to elicit contextual information. To build the corpus, we collect Wikipedia pages of famous footballers, extract the text, and chunk them (using LangChain's RecursiveCharacterSplitter with… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/football.
