CoolFace
8 results

conteb

illuin-conteb /narrative-qa ConTEB - NarrativeQA This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It stems from the widely used NarrativeQA dataset. Dataset Summary NarrativeQA (literature), consists of long documents, associated to existing sets of question-answer pairs. To build the corpus, we start from the pre-existing collection documents, extract the text, and chunk them (using LangChain's… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/narrative-qa.text10K<n<100K5 likes1.2k downloads1y agoHugging Faceilluin-conteb /mldr-conteb-train ConTEB - MLDR (training) This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It stems from the widely used MLDR dataset. Dataset Summary MLDR consists of long documents, associated to existing sets of question-answer pairs. To build the corpus, we start from the pre-existing collection documents, extract the text, and chunk them (using LangChain's RecursiveCharacterSplitter with a… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/mldr-conteb-train.text100K<n<1M0 likes875 downloads1y agoHugging Faceilluin-conteb /squad-conteb-train ConTEB - SQuAD (training) This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It stems from the widely used SQuAD dataset. Dataset Summary SQuAD is an extractive QA dataset with questions associated to passages and annotated answer spans, that allow us to chunk individual passages into shorter sequences while preserving the original annotation. To build the corpus, we start from the… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/squad-conteb-train.text10K<n<100K0 likes650 downloads1y agoHugging Faceilluin-conteb /covid-qa ConTEB - Covid-QA This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It focuses on the theme of Healthcare, particularly stemming from articles about the COVID-19 pandemic. Dataset Summary This dataset was designed to elicit contextual information. It is built upon the COVID-QA dataset. To build the corpus, we start from the pre-existing collection documents, extract the text, and chunk… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/covid-qa.text1K<n<10K1 likes603 downloads1y agoHugging Faceilluin-conteb /geography ConTEB - Geography This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It focuses on the theme of Geography, particularly stemming from Wikipedia pages of cities around the world. Dataset Summary This dataset was designed to elicit contextual information. To build the corpus, we collect Wikipedia pages of large cities, extract the text, and chunk them (using LangChain's… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/geography.text10K<n<100K2 likes504 downloads1y agoHugging Faceilluin-conteb /football ConTEB - Football This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It focuses on the theme of Sports, particularly stemming from Footballer Wikipedia pages. Dataset Summary This dataset was designed to elicit contextual information. To build the corpus, we collect Wikipedia pages of famous footballers, extract the text, and chunk them (using LangChain's RecursiveCharacterSplitter with… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/football.text1K<n<10K1 likes480 downloads1y agoHugging Face