datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mfaqWe present the first multilingual FAQ dataset publicly available. We collected around 6M FAQ pairs from the web, in 21 different languages.mqaMQA is a multilingual corpus of questions and answers parsed from the Common Crawl. Questions are divided between Frequently Asked Questions (FAQ) pages and Community Question Answering (CQA) pages.20Q20QbLLeQA
Dataset Card for bLLeQA
Dataset Summary
The bLLeQA dataset is a parallel bilingual version of the French bLLeQA dataset, which is extended to Dutch by adding the Dutch version of included legislations, and translating the questions.
The bLLeQa dataset consists of 25,982 statutory articles in French and Dutch from Belgian law and 1,461 legal questions labeled with relevant articles from the corpus.
Supported Tasks and Leaderboards
document-retrieval: The… See the full description on the dataset page: https://huggingface.co/datasets/clips/bLLeQA.
